← Back to OneIsFake

How to spot AI

Last updated 26 August 2026

Tells are probabilistic, era-dependent and constantly being patched out. What follows is what we can actually support — with sources where they exist, and an honest label where a tell is really just folklore.

Read this first

Nothing on this page authenticates anything. A tell shifts the odds; it does not prove origin. The only things that come close to proof are provenance metadata and finding the original — both covered below. Use the rest to train your eye, not to accuse people.

Strong

Documented in research or vendor material, and reproducible.

Moderate

Real, but noisy — it shifts sample by sample and model by model.

Folklore

Widely repeated, weakly evidenced. Useful as a prompt to look closer, never as proof.

Images

Modern image models fail at consistency rather than at rendering. Stop hunting for ugly pixels and start checking whether the scene obeys its own rules.

Lighting and shadow logic

Moderate

Shadows fall in directions that disagree with each other, or a reflection is missing the thing that should be in it. Diffusion models compose plausible regions, not a single physical light setup.

Background nonsense

Moderate

The subject is flawless and the background quietly falls apart: a doorframe that changes width, a crowd member with a fused arm, a railing that stops existing behind a shoulder.

Repeated micro-texture

Moderate

Skin pores, fabric weave, gravel or foliage tile or smear at 100% zoom. Detail is generated, not captured, so it lacks the irregularity of a sensor.

Too-clean composition

Moderate

Perfect symmetry, centred subject, no lens dirt, no motion blur, no awkward crop. Real photographs carry evidence of a person holding a camera.

Anatomy of the small stuff

Moderate

Not hands any more — teeth counts, earring pairs, watch faces, glasses arms, jewellery that merges into skin.

Frequency-domain fingerprints

Strong

Generators leave statistical traces in the frequency spectrum that classifiers can learn. This is real and strong, but it needs tooling — it is invisible to the naked eye.

Frequency-based detection of generated images (arXiv)

Garbled embedded text

Moderate

Signage, labels and book spines used to dissolve into pseudo-letters. Current models advertise accurate text rendering, so this now only catches older or open-weight checkpoints.

OpenAI on 4o image generation and text renderingBlack Forest Labs on FLUX.2

Video

Video has more surface to get wrong, because it has to be consistent over time as well as within a frame. Scrub frame by frame; that is where synthetic clips break.

Temporal flicker

Strong

Textures, freckles, logos and hairlines shimmer or re-draw between frames. Whole detection families are built on frame-to-frame inconsistency.

Temporal artifacts in generated video detection (arXiv)

Object permanence failures

Strong

A cup, a hand, a passer-by leaves frame and comes back changed — or never comes back. Objects occluded mid-shot are the most reliable place to look.

Physics that almost works

Strong

Cloth, hair, liquid and crowds move with the right vibe and the wrong mass. Collisions resolve too softly; things pass through each other slightly.

Lip-sync and speech drift

Folklore

Mouth shapes lag or over-articulate, and consonants that need lip closure (b, p, m) are the usual giveaway. Commonly cited, less formally quantified.

Camera behaviour

Moderate

Perfectly smooth motion with no operator error, or a pan whose parallax does not match the depth of the scene.

Text

Text is the hardest medium to call, and the one where confident guessing does the most damage. Detectors are unreliable and biased; treat every judgement as provisional.

Lexical fingerprints

Strong

Measured overuse of words like “delve”, “intricate”, “underscore” and “boasts” in text written since 2023. A real, quantified shift in published English.

Science Advances: measurable LLM word-frequency shiftCOLING 2025 study

Rhetorical scaffolding

Moderate

Tidy tricolons, “it's not just X, it's Y”, a summary paragraph nobody asked for, and balanced hedging that refuses to land on a position.

Uniform rhythm

Moderate

Sentence lengths cluster; paragraphs are the same size. This is what perplexity-based detectors measure, and light paraphrasing defeats it.

Fabricated specifics

Moderate

Citations, quotes, statistics and URLs that look right and do not resolve. A strong signal of an unchecked model — but humans invent sources too.

Em dashes

Folklore

The internet's favourite tell and one of its weakest: the human base rate was always high, and vendors have tuned the behaviour down.

Tells by vendor and model generation

From the GPT-4o era to today’s frontier models. Each entry describes the default look of a family, which prompting can override entirely — and every one of these descriptions is era-dependent by definition.

OpenAI — DALL·E 3 era

2023 – 2024

Illustrative, high-saturation, slightly airbrushed. Prompt-faithful but literal: everything asked for is present, centred and lit like a stock photo. Embedded text collapses.

CaveatSuperseded by native GPT-image generation in 2025, which fixed most of the text and anatomy artifacts.

OpenAI — GPT-image / 4o era

2025 onwards

Photographically convincing, legible in-image text, coherent multi-object scenes. Residual habit: a faintly warm, evenly exposed studio quality and unnaturally tidy scene logic.

CaveatVendor-claimed improvements; the remaining 'tidy' impression is a heuristic, not evidence.

OpenAI — Sora / Sora 2

2024 – 2025

Sora 1 clips drift: morphing limbs, background continuity loss after a few seconds. Sora 2 adds audio and much better physical plausibility, pushing failures into long-shot continuity and crowd behaviour.

CaveatClip length and prompt complexity change the failure rate more than the model version does.

Google — Imagen 3 / 4

2024 onwards

Neutral, documentary-leaning colour and very clean edges. Fewer stylistic fingerprints than Midjourney, which makes it harder, not easier.

CaveatOutputs may carry SynthID watermarking, which tooling can detect even when the eye cannot.

Google — Gemini 2.5 Flash Image (“Nano Banana”)

2025 onwards

Built for editing, so the tell is usually the edit: a subject preserved perfectly while the surrounding lighting, grain or perspective does not quite match the plate.

CaveatEdited real photos are a genuinely mixed case — part real capture, part generation.

Google — Veo 2 / Veo 3

2024 onwards

Strong camera language and native audio. Watch for physics that resolves too gently and for ambience that is generically 'correct' rather than specific to the place.

CaveatIndependent frame-level comparisons across video models are still thin.

Anthropic — Claude 3 to 4.5

2024 onwards

Text only. Careful, structured, heavily hedged prose; explicit caveats and balanced 'on the one hand' framing. Rarely commits to a bold claim without qualification.

CaveatStyle is prompt-steerable, so this describes defaults only. Anthropic ships no public image or video generator.

DeepSeek — V3 / R1 and later

2024 onwards

Long, exhaustively structured answers with heavy enumeration. Early reasoning models sometimes leaked deliberation phrasing (“wait, let me reconsider”) into final output.

CaveatOpen weights mean output style depends heavily on who deployed it and how.

Moonshot — Kimi K1.5 / K2 and later

2025 onwards

Fluent, long-context, list-friendly English. Occasional register mismatches and calques from Chinese-language training data in longer pieces.

CaveatThe clearest folklore entry here — widely observed by practitioners, not formally studied.

Black Forest Labs — FLUX.1 / FLUX.2

2024 onwards

Excellent prompt adherence and skin rendering. Practitioners report a recognisable skin and micro-contrast signature and slightly plasticky highlights, especially in the fast [schnell] variant.

CaveatThe 'FLUX look' is community consensus rather than published research.

Midjourney — v5 to v7

2023 onwards

The strongest house style of any generator: cinematic rim light, shallow depth of field, dramatic haze, and a narrow beauty standard for faces. Aesthetic homogeneity is a documented property.

Caveatv7 explicitly improved hand and body coherence, so anatomy-first checking is no longer effective on it.

Stability AI — SD 1.5 / SDXL / SD3.5

2022 onwards

The classic artifacts survive here: fused fingers, warped ears and teeth, duplicated background limbs, and mushy detail away from the subject. Old checkpoints are still in daily use.

CaveatSpecific anatomy quirks per checkpoint are practitioner lore, not measured findings.

For release dates and the full lineup, see model generations & vendors.

What no longer works

Half the advice circulating online is from 2023. These are the tells that have been patched out, watered down or were never as good as advertised.

  • Count the fingersHand and limb coherence was the target of a wave of fixes through 2024 and 2025, and Midjourney v7 shipped explicit hand improvements. Still useful on old open-weight checkpoints, useless on frontier models. [Midjourney v7 release notes]
  • Look for gibberish text in the imageAccurate in-image text is now a headline feature of both GPT-image and FLUX.2. [OpenAI: 4o image generation][Black Forest Labs: FLUX.2]
  • Em dashes mean a machine wrote itHuman writers have always used them heavily, and vendors have tuned the behaviour down. Base rates make this close to useless as evidence.
  • The word “delve”Real and measurable — but the effect has a half-life. Once a tell becomes famous, prompts and models adapt and its frequency drops.
  • Run it through an AI detectorText detectors are unreliable and are documented to misclassify non-native English writing as machine-generated. A detector score is not evidence. [Stanford HAI: detectors are biased against non-native writers]

How to verify properly

C2PA Content Credentials

An open, cryptographically signed manifest travelling with a file, recording which tool made or edited it. Adopted by OpenAI, Google, Adobe and camera makers.

Limit Metadata is stripped by screenshots, re-encoding and most social platforms, and absence of credentials proves nothing.

C2PA specification

SynthID

Google's imperceptible watermark applied across Gemini, Imagen and Veo output in text, image, audio and video, detectable by its own verification tooling.

Limit Only covers Google-generated media, and research has demonstrated removal and evasion attacks.

Google DeepMind: SynthID

OpenAI provenance metadata

OpenAI attaches C2PA Content Credentials to images produced in ChatGPT and the API.

Limit OpenAI itself describes these as helpful indicators rather than a guarantee, and they survive only while the metadata is intact.

OpenAI: C2PA in images

Reverse image and context search

Still the highest-yield check available to a person with no tooling: find the earliest appearance, the original crop, and whether the event happened at all.

Limit Fails on genuinely novel generations that were never published anywhere else.

Google Images

Now go practise

Reading about tells is not the same as recognising them at speed. Play a few rounds and find out which of these you can actually apply under a timer.