Portraits
Human ears are among the most structurally complex features on the face, made up of intricate cartilage folds—the helix, antihelix, tragus, and concha—that create a distinct pattern of ridges and shadows. When a real camera captures an ear, these folds cast small, consistent shadows because they follow the physical logic of light hitting a three-dimensional, ridged surface, and every ear is asymmetrical in a believable, organic way. Diffusion-based image generators, however, learn to produce faces from millions of training images where ears are frequently partially obscured by hair, turned away from camera, or blurred by depth of field. This means the model has much less reliable data about ear geometry compared to eyes, noses, and mouths, which dominate portrait training data. As a result, generated ears often come out smoothed, waxy, or structurally simplified, missing the natural cartilage ridges and shadow transitions a real ear would have. This inconsistency is a common giveaway in AI portraits, since the rest of the face can look convincing while the ear reveals a loss of anatomical detail the model never fully learned to render correctly.
Hands
Human hands have a very specific and mechanically constrained structure: each finger is separated by a distinct web of skin at the base, and the knuckle joints create predictable shadow gaps even when fingers are pressed closely together during a grip. A real photograph captures this because light physically cannot pass through solid flesh, so even tightly clenched fingers show subtle creases, shadow lines, and consistent joint spacing dictated by bone structure. Diffusion-based generators, however, don't build images from an understanding of skeletal anatomy -- they statistically assemble pixels based on patterns seen in training data, and hands are notoriously difficult because they appear in countless orientations, lighting conditions, and grips. When fingers overlap or curl around an object like a plant stem, the model often loses track of where one digit ends and another begins, resulting in fingers that blend into a single mass, extra digits appearing at the edge of a hand, or knuckles that don't line up with the number of visible fingertips. This is why hands remain one of the most reliable places to spot AI-generated imagery, even as skin texture and lighting have become nearly photorealistic.
Urban Landscapes
Text in the real world is built from precise, repeatable letterforms that follow strict rules of spacing, kerning, and spelling because a human designer or sign manufacturer created them with intent. A camera photographing that sign simply records those exact shapes with optical fidelity, so letters stay crisp and legible even under motion blur or low light, since the underlying geometry never changes. Diffusion-based generators, however, don't understand language as a system of symbols; they learn to paint textures that statistically resemble letters based on patterns in training data, without any true concept of spelling or grammar. This means they can convincingly mimic the glow, font style, and neon color of signage while still producing nonsense strings, doubled letters, or fused characters. The problem gets worse with reflective or wet surfaces, where the model must also hallucinate a mirrored version of the same invalid text, compounding the inconsistency. Because sign text is often small and stylized, it's one of the most reliable places to spot synthetic imagery, since even highly realistic lighting, reflections, and architecture can be undermined by a single illegible or garbled word painted onto a storefront.
Food & Texture
When a chef drizzles sauce onto a plate, gravity, viscosity, and the randomness of hand movement create asymmetrical pools with irregular edges, varying thickness, and unpredictable branching where the liquid thins out. Real sauce also interacts with the plate's texture and the food's shape, pooling unevenly and catching light differently depending on depth. Diffusion models generate images by learning statistical patterns from millions of photos, and they often default to symmetry or repeating decorative motifs because such patterns are common and 'safe' looking in training data, especially for plated food photography where circular symmetry is aesthetically pleasing. This causes the model to produce a sauce pattern that loops back on itself in a near-mirrored fashion, lacking the organic asymmetry of a real fluid dispersal. A camera capturing an actual dish would show subtle inconsistencies: thicker pools where sauce settled, thinner streaks where it was dragged by a spoon, and light reflecting unevenly off the sauce's varying thickness. The absence of this natural randomness, replaced by a too-perfect repeating loop, is a subtle but telling sign of synthetic generation.
Animals
A catchlight is the tiny bright reflection visible in an eye, created when light sources and the surrounding environment bounce off the eye's curved, wet surface. In real photography, these reflections are physically constrained by the eye's small radius of curvature, meaning any reflected scene is compressed into a tiny, blurry, and often unrecognizable smear of light and shadow. Actually rendering a coherent, sharply defined scene, like a person standing in a field, within a curved reflective surface only a few millimeters wide requires precise ray-tracing of the environment relative to the eye's exact position and curvature. Diffusion models don't simulate optics this way; they generate images by predicting plausible pixel patterns based on training data, so they often insert recognizable shapes into reflective surfaces because such shapes appeared in similar contexts during training, without regard for whether the geometry or scale makes physical sense. This results in reflections that look conceptually clever but violate the basic physics of how curved surfaces compress and distort reflected imagery, a subtle mismatch that becomes obvious once you examine the eye closely.
Crowds
Human hands are among the hardest things for image generators to render correctly because they are complex, articulated structures with many small joints, overlapping surfaces, and constantly changing poses. A real camera captures a hand exactly as it is in that instant: five distinct fingers with consistent width, correct joint spacing, and shadows that fall naturally based on the actual 3D geometry of skin and bone. Photographic optics simply record light bouncing off real anatomy, so proportions and finger separation are always physically accurate, even in low light or partial motion blur. Diffusion-based AI models, however, generate images by statistically predicting pixel patterns based on training data rather than simulating real anatomy. Hands involve fine, high-frequency detail packed into a small area, and small errors in the generative process compound easily, leading to fingers that merge, bend unnaturally, or vanish entirely. This is especially common when a hand is partially obscuring an object or holding something, since the model must infer occluded structure without a true understanding of underlying bone and joint mechanics, often resulting in blurred or fused digits instead of clean anatomical separation.