Portraits
When a camera photographs a bookshelf out of focus, the bokeh still preserves the underlying structure of each object: a book spine remains a rectangular block of fairly uniform color with a hard edge where it meets its neighbor, and any lettering, even blurred into a smear, follows the physical geometry of the original glyphs because it was actually printed there and optically defocused, not invented. Real out-of-focus text degrades in a predictable way tied to the lens's circle of confusion, spreading letterforms symmetrically while keeping their color and rough stroke pattern intact. Diffusion models, in contrast, don't photograph real books; they generate plausible-looking texture from statistical patterns learned across millions of images. When asked to render distant, blurry text, they often produce a repetitive, patternless smear of color bands that never resolves into coherent letterforms, because the model has no underlying 'true' text to blur, only a learned impression of what shelves generally look like. This results in spines that shift color abruptly, repeat suspiciously, or contain scribble-like marks that mimic text without ever forming real words, a telltale sign of synthetic generation rather than authentic photographic defocus.
Hands
When a hand grips a rigid object like a phone, each finger bends independently at its own knuckle, creating distinct creases, slightly different lengths, and visible gaps or shadows between adjacent fingers where light can't reach. A camera lens captures this because real skin, bone, and tendon interact with light in physically consistent ways — every crease catches a highlight and every gap casts a small shadow, following the actual joint structure underneath. Diffusion models, however, don't build a hand from a skeleton outward; they generate pixels by statistically guessing what 'looks like a hand near an object' based on training images, without any true underlying anatomical model. This means fingers often blur into one another, knuckles disappear, or an extra fold of skin appears where a joint should be, because the model is pattern-matching texture and shading rather than simulating bones and tendons. The result is a soft, waxy, webbed appearance where fingers should be crisply separated, especially in tight grips where overlapping fingers create complex occlusion that generative models still struggle to resolve convincingly.
Urban Landscapes
Bicycles are mechanically simple but geometrically strict: the frame is a closed triangle of tubes welded at precise angles so the front fork, head tube, and handlebars align on a single steering axis, and the chain, pedals, and rear wheel share a fixed drivetrain line. A real camera simply records whatever light bounces off this rigid structure, so every tube stays straight, every joint meets at a physically weld-able angle, and perspective distortion follows predictable lens optics. Diffusion models, however, don't understand rigid mechanical linkages; they generate images by refining noise based on statistical patterns learned from millions of photos, without any internal skeleton or physics engine enforcing that a fork must terminate at a wheel axle or that a handlebar stem must be a single continuous tube. This means bicycles, ladders, chairs, and other lattice-like objects are notorious failure points — the model can produce a tube that curves through space in a way no metal could bend, or a wheel that appears to hover slightly off its axle, or a frame that would collapse if it existed. These small logical breaks in load-bearing structure are one of the most reliable tells of synthetic imagery.
Food & Texture
When dough ferments, wild yeast and bacteria release carbon dioxide unevenly, creating air pockets that vary wildly in size, shape, and spacing depending on hydration, gluten development, and shaping technique. A real photograph of cut sourdough will show this chaotic randomness: some holes huge and irregular, others tiny and clustered near the crust, with no repeating rhythm. Diffusion models generate texture by learning statistical patterns from thousands of training images, and they often default to a learned 'average' texture that looks locally convincing but repeats at a larger scale because the model lacks true physical simulation of gas expansion in dough. This produces crumb structures that look suspiciously symmetrical or rhythmically spaced across multiple slices, almost like a printed pattern rather than the product of biological fermentation. A camera simply records whatever irregular structure existed in the real loaf, so genuine bread photography never shows this kind of cross-slice uniformity unless the loaves were baked from an identical, tightly controlled recipe and even then natural variance in bubble formation would still occur due to microscopic differences in gluten strands and moisture during proofing.
Animals
Whiskers are stiff, translucent hairs rooted in specialized follicles, and a real camera captures them as continuous, thin lines with a consistent width that only tapers gently toward the tip, often with a tiny catchlight where they cross a light source. Because they're rigid, they hold a clean, unbroken curve from root to tip, and depth-of-field causes them to blur uniformly when they leave the focal plane, but they never simply dissolve mid-strand. Diffusion models generate images by iteratively refining noise based on statistical patterns learned from millions of photos, without any understanding of hair as a physical structure with a root and a fixed thickness. This means thin, high-frequency details like whiskers are notoriously hard to render consistently — the model may start a whisker convincingly but lose track of it as the denoising process moves through different regions of the image, causing it to blur, thin out unnaturally, or disappear entirely before reaching a logical endpoint. This inconsistency, especially near a focal point where sharpness should be highest, is a common signature of synthetic imagery struggling with fine linear textures.
Crowds
AI-generated images in the "Crowds" category often contain subtle inconsistencies in lighting, texture, or fine anatomical detail that reveal their synthetic origin, but this round hasn't had a human or vision-model pass yet to pin down exactly which detail gives it away. (needs human QA)