Portraits
Human teeth are individual structures separated by thin, dark gum lines and subtle vertical grooves that catch light differently depending on curvature, moisture, and angle. A real camera lens, especially at close range with a wide aperture, resolves these fine gaps as crisp micro-shadows because each tooth is a distinct reflective surface with its own edge and highlight. Diffusion models generate images by statistically averaging patterns learned from millions of photos, and mouths are notoriously difficult because they combine specular highlights, wet surfaces, and repetitive small shapes packed close together. Instead of rendering each tooth as a separate geometric object, the model often produces a smooth, continuous white mass with only faint suggestions of separation, since it lacks true 3D understanding of dental structure. This results in teeth that look glossy and blended rather than crisply defined, especially under bright, direct lighting where real dentition would show sharp contrast between enamel highlights and the darker interdental shadows. The effect is often more noticeable in wide, open-mouthed smiles, where more tooth surface is exposed and averaging errors compound across a larger visible area.
Hands
Human hands have a strict, predictable skeletal structure: four fingers plus a thumb, each with consistent joint spacing and a logical path back to the palm. When a camera captures a hand gripping an object, every knuckle bend and tendon line follows the physics of bone and tissue compressing under pressure, and each finger's silhouette connects continuously to the hand's mass. Diffusion-based generators don't understand anatomy as a skeletal system; they generate images by predicting plausible pixel patterns based on training data, which means fingers can be duplicated, fused, or left oddly floating without a clear joint-to-palm connection. This happens because hands are complex, high-detail, frequently-occluded shapes with huge pose variability, making them one of the hardest structures for these models to learn consistently. The result is often a finger that bends at an implausible angle, appears twice, or seems to hover disconnected from the rest of the hand's musculature. A real photograph would never produce this kind of structural ambiguity, because actual bones and tendons impose hard physical constraints on how fingers can bend and where they must attach.
Urban Landscapes
When a camera photographs a line of parked cars, each vehicle has a distinct, continuous body panel edge defined by real-world manufacturing tolerances: a sharp gap of shadow between bumpers, a consistent paint reflection that follows the curvature of one car's metal at a time, and headlight or window glass that catches light in a physically predictable way based on the sun's actual position. Reflections on a car's hood or window are essentially a mirror of the environment, so they shift smoothly and logically as the eye moves along the vehicle's surface. Diffusion-based generators build images by gradually resolving noise into plausible shapes based on patterns learned from millions of photos, rather than simulating actual object boundaries or light physics. Because repeated objects like a row of parked cars share similar colors, angles, and proportions, the model can lose track of where one object ends and the next begins, causing panels, mirrors, or fenders to blend together or duplicate awkwardly. This is especially common with mid-distance clusters of similar objects, where the model prioritizes overall scene plausibility over precise object-level geometry.
Food & Texture
When a camera photographs a pile of cut fruit, every piece is an independent physical object shaped by a unique knife stroke, ripeness, and angle of repose, so no two slices will ever share an identical silhouette, curvature, or fiber pattern. Diffusion models, however, don't simulate individual objects; they generate images by repeatedly denoising a canvas based on learned statistical patterns of what 'mango slice' or 'kiwi wedge' tends to look like. Because the model is drawing from a compressed sense of the 'average' appearance of a food item rather than physically arranging distinct instances, it often reuses very similar shapes, curves, or seed patterns across a scene, especially when many similar objects are clustered together, like in a fruit salad. This creates a subtle but detectable rhythm of repetition that a real photograph would never produce, since real fruit is cut by hand and falls into a bowl or plate randomly. Lighting, juice glisten, and seed placement may also echo suspiciously between pieces because the generator applies the same latent texture pattern more than once instead of rendering each object from first principles.
Animals
Many crepuscular hunters, including foxes, cats, and some other mammals active at dawn and dusk, have evolved vertically elongated slit pupils rather than round ones. This shape lets the eye rapidly and precisely modulate incoming light across a huge range, from bright daylight to near darkness, by closing to a thin vertical line or opening into a fuller aperture, while also enhancing depth perception for judging pounce distances. A camera photographing such an animal simply records whatever pupil geometry biology has already built, since a lens has no influence over the eye's internal muscles. Diffusion models, however, learn eye appearance from vast pools of training images dominated by human portraits and generic circular-pupil animal photos, so they default to statistically common round or slightly ambiguous pupil shapes even when generating species that should show slits. Because the model is pattern-matching pixel statistics rather than simulating actual iris musculature or species-specific anatomy, it can blend or smooth the pupil into an anatomically incorrect form, producing a subtle but telling inconsistency once you look closely at the eye's true shape.
Crowds
Wheeled objects like strollers, bicycles, and carts are notoriously hard for diffusion models because they require rigid, radially symmetric geometry with consistent spoke counts, axle alignment, and perspective-correct ellipses for each wheel. A real camera captures a wheel as a perfect circle or ellipse depending on viewing angle, with spokes converging precisely at a single hub point and consistent metal or plastic reflections following the curve of the rim. Diffusion models generate images through statistical pattern-matching across millions of training photos, but they lack any actual understanding of mechanical structure or physics, so when two overlapping wheels, folding hinges, and thin metal tubing all compete in a small, partially occluded area of the frame, the model often blends them into ambiguous, semi-melted shapes with mismatched spoke counts or wheels that don't sit flush against the ground. This is especially common in busy, cluttered scenes where the object is small relative to the frame, since the model allocates less coherent structure to peripheral details compared to the main subjects it was more heavily trained to render accurately, like faces or torsos.