How this site earns your trust

Maintained by the HumanSounding Team with Claude. Not affiliated with OpenAI, Anthropic, or Google.

Honesty rules we hold ourselves to

Every figure traces to a named source. Per-model claims carry confidence labels: measured, documented, or anecdotal.

The visit counter undercounts rather than inflates. On a quiet week, the changelog says "nothing material changed" instead of manufacturing news.

Avoiding the vernacular makes text less annoying to read. It does not defeat detectors, and we never claim otherwise.

Privacy, stated plainly

The draft checker runs entirely in your browser. Your text is never transmitted anywhere.

We record only anonymous daily counts: visits, which tells fired, referrer domain. No IPs, no cookies, no personal data, no signup.

The whole site passes its own checker. Zero em dashes in our prose; we counted.

Sources

Kobak et al., Science Advances 2025: excess vocabulary; delve 28× Liang et al., ICML 2024: peer-review word spikes Sun et al. (CMU): 97.1% five-way model identification SlopDetector, 2026: per-model em-dash rates Xu et al. (CMU), 2026: detectors track instruction tuning, not machine authorship
All 20 sources, with what each one backs up: full list

Kobak et al., "Delving into LLM-assisted writing in biomedical publications" (Science Advances 2025): excess-vocabulary method; 13.5% figure; delve 28×.

Juzek & Ward, "Why Does ChatGPT 'Delve' So Much?" (COLING 2025): 21 focal words; +6,697% "delves".

Liang et al., "Monitoring AI-Modified Content at Scale" (ICML 2024): peer-review word spikes; 6.5–16.9% modified reviews.

Xu, Zhong, Raghunathan, Fang & Kolter, "Base Models Look Human To AI Detectors" (CMU, May 2026): commercial detectors rated base-model Llama-3-8B continuations 96.7% human against 30.3% for the instruction-tuned version (GPTZero), and 98.8% against 17.1% (Pangram). The authors conclude detectors track artifacts of instruction tuning and local context rather than any invariant property of machine-generated text. Backs the claim that detection, like the tells themselves, is model-specific.

Liang et al., Nature Human Behavior 2025: LLM share of scientific papers by field.

Yakura et al. (Max Planck), LLM influence on human spoken communication: 737k hours of podcasts; "delve" in speech.

Juzek, Anderson & Galpin (FSU, AIES 2025): AI words in unscripted speech.

Sun et al. (CMU), "Idiosyncrasies in Large Language Models": 97.1% five-way model identification; per-model markers.

GPTZero, most common AI vocabulary: phrase multipliers (182×, 120×, 107×…).

SlopDetector, em-dash density data (2026): per-model em-dash rates vs human baseline.

"The Last Fingerprint: How Markdown Training Shapes LLM Prose" (2026): suppression-resistance; GPT-5.4 changes.

Wikipedia: Signs of AI writing (WP:AISIGNS): the editor-built taxonomy this site's structural categories follow.

Decrypt, "The 5 Biggest Tells" (Nov 2025): Washington Post 328,744-message dataset; "not just" ~6%.

TechCrunch, OpenAI's em-dash fix (Nov 2025): and PCWorld's caveats.

OpenAI, "Sycophancy in GPT-4o" (Apr 2025): the rollback postmortem.

anthropics/claude-code#3382: the "You're absolutely right!" bug report; tracker at absolutelyright.lol.

SycEval (Stanford) coverage: sycophancy rates: Gemini 62.47%, ChatGPT-4o 56.71%.

Scientific American / Rudnicka: ChatGPT vs Gemini stylometry ("glucose" vs "sugar").

Forbes, Gemini's self-loathing loop (Aug 2025).

Breunig, slop forensics & model ancestry: with sam-paech/slop-forensics and EQ-Bench Slop Score.

Scientometrics 2026 multi-database study: "underscore" through July 2025.

Built from published studies, detector datasets, and editor guides; per-item confidence is labeled where attribution is uncertain. Figures are quoted from their sources without adjustment; corpora, dates, and definitions differ between studies, so multipliers are not directly comparable across charts. Three caveats worth keeping in view. First, attribution is hard: most "Claude-isms" and "Gemini-isms" people complain about are generic LLM habits attributed to whichever model the complainer uses most; per-model claims carry confidence labels for that reason. Second, the tells are a moving target: labs tune out each meme roughly a year after it peaks, so the durable signals are structural (cadence, negative parallelism, triplets), not word lists. Third, nothing here is for passing text off where AI disclosure is required.