Portraits
When a camera photographs fabric like a wool suit jacket, the weave pattern, seam lines, and any small details like a lapel pin or button follow consistent physical logic: light falls off predictably along the fabric's grain, stitching lines run parallel and continuous, and hardware like pins or clasps have clean, resolved edges because the lens focuses them the same as everything else at that depth of field. Diffusion models, however, don't understand these objects as physical structures—they generate images by statistically guessing pixel patterns based on training data, which often results in small, semantically meaningless shapes appearing where a camera would show clean fabric or a well-defined object. This is especially common in areas like lapels, collars, and cuffs, where fine repetitive detail must be reconstructed without a real underlying model of stitching or fabric behavior. The result is a soft blotch or rectangular smear that reads as almost like an object but resolves into nothing coherent under close inspection, revealing the absence of any true material or geometry behind it.
Hands
When two real hands overlap in a photograph, each finger maintains a continuous, unbroken silhouette defined by consistent shadow lines, nail beds, and knuckle creases that a camera lens faithfully records because it is simply measuring light bouncing off real, separate surfaces. Depth cues like occlusion edges, parallax, and shifts in focus naturally separate overlapping fingers even when they are inches apart. Diffusion models, however, build images from statistical patterns learned across millions of photos rather than by modeling actual three-dimensional hand geometry. When two hands cross in the frame, the generator has no true skeletal or spatial understanding of where one finger ends and another begins, so it often blends contours, merges knuckle lines, or invents an extra joint where two overlapping digits should remain visually distinct. This is especially common in complex, multi-hand compositions where the model must reconcile several limbs at once, since hands already have notoriously variable, high-frequency detail that pushes generative networks toward smoothing or merging ambiguous regions instead of resolving them into anatomically correct, separate structures.
Urban Landscapes
Bicycles are mechanically precise objects: their frames form clean triangular geometries, wheels are perfect circles seen in consistent perspective, and spokes radiate with even spacing because they are manufactured to tight tolerances. A real camera captures this faithfully because optics simply project straight tubes and circular rims according to well-understood perspective rules, and any blur or distortion follows predictable lens physics tied to focal length and depth of field. Diffusion-based image generators, however, don't understand mechanical structure the way a machinist or a camera does — they build images from statistical patterns learned across millions of photos, guessing at how bars, wheels, and small hardware should connect rather than truly modeling rigid geometry. This makes bicycles, fences, and other repetitive metal structures a common weak point, since overlapping thin lines in the background confuse the model into fusing separate objects, bending frame tubes at impossible angles, or producing wheels that don't sit correctly on the ground plane. The result often looks plausible from a distance but falls apart upon close inspection of how the parts actually connect.
Food & Texture
When ceramic glaze chips through years of use, the damage is a product of countless random impacts, temperature changes, and scratches, so every chip differs wildly in size, depth, and edge shape, with no two ever looking alike. A macro photograph of a worn plate would show this chaotic irregularity clearly, often with subtle color variation at the fracture line where the glaze thins and the clay body shows through unevenly. Diffusion models learn 'distressed ceramic' as a repeating visual motif rather than a physical process, so they tend to generate chips that share a similar scale, curvature, and rhythm around a rim, almost like a stamped pattern rather than organic decay. Because the model is essentially pattern-matching textures it has seen in training data rather than simulating material fracture, it often smooths over the random asymmetry that real wear produces. This kind of repetition is a common giveaway in AI-generated food photography, since props like plates, table grain, and cloth folds are exactly the kind of secondary background detail where generative models take visual shortcuts.
Animals
Goats, like many grazing animals, have horizontal rectangular pupils that stay remarkably consistent in shape and orientation between both eyes, since this design evolved specifically to give them a wide panoramic view of the horizon for spotting predators. A real camera capturing a goat's face straight-on will show both pupils as mirror images of each other, with the same slit width, curvature, and angle, because the underlying eye anatomy is symmetrical and governed by the same muscles and lighting conditions. Diffusion models generate each eye somewhat independently, piecing together patterns learned from thousands of training photos rather than simulating an actual eyeball structure. Without a true anatomical template enforcing left-right symmetry, these models often produce pupils that differ subtly in shape, tilt, or size between the two eyes. This is a common tell in AI-generated animal portraits: the irises and catchlights may look convincing individually, but comparing the two eyes side by side reveals inconsistencies that a real photograph, bound by physical anatomy and consistent lighting geometry, would never produce.
Crowds
Street signage in real photographs is produced by precise manufacturing processes: letters are cut from reflective vinyl or stamped metal using standardized fonts, so every character has consistent stroke width, spacing, and legibility even at a distance or under motion blur. A camera simply records whatever letterforms already exist in the physical world, so text stays coherent no matter the angle or lighting. Diffusion-based image generators, however, don't understand written language as a symbolic system-they learn text as visual texture and pattern statistics from millions of training images. This means they can convincingly reproduce the general shape, color, and placement of a sign, but the actual glyphs often morph into repeated or malformed letters, since the model is essentially guessing plausible-looking shapes rather than spelling real words. This weakness is especially visible on small, distant signage like street markers, license plates, or storefront text, where the model has less contextual information to anchor the correct spelling, resulting in dreamlike, almost-readable but ultimately nonsensical text strings.