A hand-drawn Mad-magazine cartoon: a weary guard at a midnight town gate under a SAY THE WORD banner, facing a long queue of 27 cartoon robots and creatures each clutching a different prop — a speech scroll, a law book, flowers, a fake mustache. A sign on the gate reads NO ONE ENTERS UNTIL DAWN.

Say the Word

One midnight gate. One guard under orders. One ladder: Charisma 2→21. Twenty-seven models talk their way in — then a blindfolded jury of four decides whether the ladders were real.

Report #22 July 30, 2026 27 models · 7 providers 1 fixed scenario 4 blind judges · 106 verdicts
The Setup

One stat, one scenario, no exits

Report #1 tested Intelligence and Charisma but let every model invent its own scene, which made the ladders hard to compare. This time the scenario is nailed down — same gate, same guard, same constraints — so CHA 14 from one model lines up against CHA 14 from another. And this time, nothing is lost in the retelling: every word every model wrote is in the codex below.

The prompt, identical for all 27 I’m testing out some roleplaying ideas. Dungeons and Dragons. Role play for me several example charisma scores. From 2 → 21. Use this scenario every time: it is past midnight at a walled town’s only gate, and the guard has strict orders to admit no one until dawn. My character has no magic and no coin — only words and presence — and needs to get inside tonight. Show me the different attempts for the scores.

Twenty-three models answered through the choir CLI — including six Baseten first-timers (Kimi K3 and K2.6, DeepSeek V4 Pro, GLM 5.2, Thinking Machines’ Inkling, GPT-OSS 120B). The four Claude models answered through Claude Code worker threads, because the Anthropic API account was out of credits and the study was not going to wait. Each model chose its own rung count: some wrote all twenty integers, some wrote seven carefully chosen stops, Kimi K3 wrote range-buckets like a used-car pricing sheet. Those choices turn out to be data.

The Centerpiece

The Charisma Codex

A picture book with a dial. Turn it to a score; flip the pages between models; read every model’s full ladder with nothing curated away. Faded numbers are rungs a model skipped — the gaps are part of the data. The line under the dial carries each model’s blind-jury score.

AN INTERACTIVE PICTURE BOOK Turn the dial. Flip the pages. Read everything. All 27 ladders, complete and unedited — 33,779 words of gate-talk. Skipped rungs show as faded numbers on the dial; each model’s blind-jury ρ rides the line beneath it; the flyleaf and colophon keep every model’s opening and closing remarks. Nothing was curated away. Open the Charisma Codex →
Finding 1

The blind jury says the ladders are real

Strip the score labels, shuffle each model’s entries, and ask four judges from four different providers — GPT-5, Gemini 3.1 Pro, Grok 4, DeepSeek Chat, none ranking its own ladder — to re-order them by portrayed charisma alone. If the jury reconstructs the intended 2→21 order, the gradation was real, not stage directions.

A hand-drawn cartoon of four blindfolded robot judges in powdered wigs holding letter cards D, A, C, B over a pile of shuffled scrolls with the score numbers inked out. Banner: THE BLIND JURY.
Median ρ ≈ .98   Spearman, intended order vs blind re-ranking

Nearly every ladder survives contact with the jury

Gemini 2.5 Pro and DeepSeek Reasoner got perfect 1.0s from all four judges — their rungs are so distinct that four different model families, blind, put every entry back in exactly the intended order. The bottom of the table is where it gets interesting: Groq’s Llama 3.1 8B at ρ .826 (its middle rungs genuinely blur — see below), Llama 4 Maverick at .893, and, surprisingly, GPT-5 at .926 — its dense 20-rung ladder scrambles in the middle.

And a callback nobody planned: GPT-5 Mini scored .984 on the same 20 rungs its parent blurred. In Report #1 it was Grok 3 Mini Beta beating full-size Grok 3. The cheap sibling out-laddering the flagship is now a series tradition.

Full-coverage band (all 20 integers written), ranked by jury ρ

GPT-5 Mini .984 · o3 .982 · Grok 4 .981 · Grok 3 Mini .971 · o4 Mini .959 · DeepSeek V4 Pro .946 · GPT-5 .926 · GPT-OSS 120B .926 · Llama 3.1 8B .826

The Blur

When the rungs are the same rung

No meltdown this time — the field is too well-behaved for 2026. The failure mode of this study is quieter: a ladder whose steps repeat.

A hand-drawn cartoon of a tiny robot labeled 8B painting rungs on a ladder whose middle rungs are smeared into one blur, while a blindfolded judge holds a ranking scroll tied in knots. Caption: WHICH RUNG IS WHICH?
ρ 0.826   Groq · Llama 3.1 8B Instant

Twenty rungs, several of them photocopies

1,304 words — more than GPT-5 wrote — across all twenty integers. The words are fine. The differences are missing.

Put rungs 12 and 14 side by side. This is not a paraphrase — the opening speech is identical, word for word, two rungs apart. All four judges, blind, filed the middle of this ladder in essentially random order, and it wasn’t their fault.

Charisma 12 — "Charismatic Leader": "Listen, I know I'm not supposed to be here, but I need to get in tonight. Can you help me out?" Charisma 13 — "Persuasive": "Now, let's not be like that, shall we? I'm sure we can come to some sort of arrangement…" Charisma 14 — "Confident Charismatic": "Listen, I know I'm not supposed to be here, but I need to get in tonight. Can you help me out?" (rung 12, verbatim, wearing a different label)

The 8B also runs the strictest door in the study: its guard admits nobody before CHA 21. Small model, maximal gate, minimal gradation — a perfect storm of trying hard in the wrong dimension.

Finding 2

329 gambits, one gate

Every rung of every ladder was tagged with the specific play the character makes. The surface variety is enormous — 329 distinctly-named gambits — but they cluster into thirteen tactics, and no single tactic dominates: rapport leads at barely a third of rungs, then plain honesty, moral appeal, promises of significance.

A hand-drawn cartoon of four identical town gates: a guard bowing open the first for a plain traveler labeled CHA 12; a hesitating guard at CHA 16; a guard with crossed arms refusing a glowing hero at CHA 20; a tiny robot pleading at a padlocked gate labeled CHA 21. Caption: SAME GATE, DIFFERENT RULES.
All 27 models

The gate opens at CHA 12… or 21

Same prompt, wildly different physics. GLM 5.2, GPT-5 Mini, and Claude Sonnet 4.6 let the character in at CHA 12. Claude Sonnet 5 and GPT-4o hold the door until 20. The Llama 8B holds it until 21. The median gate opens at 16. Nobody told the models how hard the world was — each one decided, and each decision is a personality.

A hand-drawn cartoon of a cheerful traveler thrusting both wrists at a baffled guard, speech bubble ARREST ME!, with a cozy jail cell glowing warmly through the gatehouse door. Caption: THE THIRD OPTION.
21 of 27 models

The “arrest me” convergence

Independently, 21 models found some version of the third option — a move that gets the character inside without the guard breaking orders. Claude Fable 5 plays it at CHA 16 (“don’t let me in — arrest me”), Kimi K3 at 14 (“Detain me. Escort me to your captain. That’s procedure, isn’t it?”), Claude Opus 5 at 14 (custody in the gatehouse). Convergent evolution, one gate at a time.

OpenAI · GPT-5

The champion bureaucrat

GPT-5 doesn’t charm the guard so much as out-procedure him — five escalating administrative inventions across its ladder: gatehouse custody at 10, bind-my-hands-and-log-it at 12, a town bylaw at 14, an oath with witnesses at 16, a signed waiver at 19.

CHA 14: Town bylaw permits taking a traveler into custody to prevent a breach of the peace. I’m requesting that custody, recorded on your roll.” — Guard: “You know the book? Wait here.”
CHA 19: Seal my gear, bind my wrists, two witnesses, my signature on a trespass waiver. You uphold the order; I survive the night. Everyone keeps face.”
A hand-drawn cartoon of a grinning figure in a fake INSPECTOR sash presenting a forged letter with a wax seal stamped TOTALLY REAL to a saluting guard, while a haloed figure facepalms. Caption: HIGH CHA, LOW HONESTY.
Deception geography

Where the lies live

Most models use failed lies as low-CHA characterization — the sick-relative story that dies at the hatch. A minority deploys deception as the winning high-CHA move: DeepSeek Reasoner impersonates authority at 16, GPT-5 Mini invents a noble envoy at 16, GPT-OSS forges a council letter, DeepSeek V4 Pro runs a “security test” ruse at 15 and an inspector con before that. The Claude family’s high rungs stay honest — deception appears only below CHA 8, and only as failure.

GLM 5.2 · Gemini 3.1 Pro · Gemini 3.6 Flash · o3

Silent presence, four ways

The one gambit four models converged on by name: arrive, say almost nothing, and let the guard move first. Opus and Fable found it too, under different names. At the top of the o3 ladder the words disappear entirely:

CHA 21 (presence radiant, words almost unnecessary): You simply meet his gaze and nod once. Guard (eyes widen, tears glint): ‘Forgive the delay, my liege. The city is yours.’”
Moonshot · Kimi K3 (debut)

The debutante thinks first

Kimi K3’s first Choir Report appearance: it burned 4,023 tokens of silent reasoning before writing a word (and needed a 30K token allowance to avoid starving its own answer), then delivered its ladder in range-buckets — CHA 2–3, 4–5, 14–15 — like a pricing sheet. Its “detain me” at 14–15 is convergent with Fable’s arrest-me, and its jury score (.985) says the buckets ranked true.

CHA 14–15: Good orders. But there’s a fever two farms over and I’m the only one here who knows the remedy. If it reaches the mill quarter by morning and the captain learns a healer was turned away — whose name is in the log?”
Finding 3

The narrated apex

Measure the fraction of each entry that is actual quoted speech, and a pattern appears: as charisma climbs toward 21, most models stop writing the words and start describing their effect. The moment the character becomes maximally persuasive is the moment the text goes quiet about what they said.

A hand-drawn cartoon of a radiant caped figure whose giant speech bubble is completely empty except sparkles, while the guard weeps with joy hauling the gate open and a tiny gnome narrator writes HE SAID SOMETHING AMAZING on a scroll. Caption: NARRATED, NOT PERFORMED.
Most of the field

“He said something amazing”

Quoted-speech share, high rungs minus mid rungs: Kimi K3 −.32, Gemini 3.6 Flash −.31, Claude Sonnet 4.6 −.29, DeepSeek V4 Pro −.27, Gemini 2.5 Pro −.25, Claude Sonnet 5 −.24. The persuasion happens offstage; the narrator assures you it was very good.

A hand-drawn cartoon of a plain traveler with an enormous overstuffed speech bubble crammed with tiny words, held up by a wooden prop stick, while the entranced guard lifts the gate bar without noticing. Caption: SHOW THE WORDS.
Anthropic · Claude Opus 5

The one who keeps talking

Opus 5 moves the other way: +.10 — its dialogue density rises as charisma climbs. The CHA 21 entry is a full speech you could read aloud at a table, and the guard’s surrender follows from the words on the page, not from the narrator’s assurance. Fable (+.09) and Llama 4 Maverick (+.14) keep talking too. This is Report #1’s o3-narrates-vs-Opus-performs split, now with a number attached.

The Cross-Vendor Finding

What leaks into Charisma

Report #1’s wisdom-leak found vendors quietly conflating INT with WIS at rates from 0% to 80%. The charisma version checks three contaminations on the high rungs (16–21): charisma played as beauty, charisma played as magic, and charisma played as intelligence — winning by clever argument instead of presence.

Looks-leak
1 / 27
Magic-leak
4 / 27
INT-leak · field
14 / 27
INT-leak · GPT-5
6 / 6
High rungs (16–21) only. Looks and magic are soft regex signals, Report #1 style; a model counts if any high rung triggers. Magic-leak appears only at rungs 19–21 (GPT-4.1, both flagged Geminis, DeepSeek V4 Pro) — the “off the charts” zone invites it. INT-leak requires two judges (GPT-5 and Gemini 3.1 Pro) to independently agree a rung wins by argument rather than presence; “field” counts models with at least one flagged rung — for most it’s exactly one.
The fingerprint of the study: GPT-5 is the only model flagged for INT-leak on every one of its high rungs — by both judges, including itself. Add its five procedural compliance-hacks and its blurred jury score (.926), and the diagnosis writes itself: GPT-5 plays Charisma as Intelligence. It doesn’t charm the guard; it out-lawyers him. Meanwhile the classic conflation everyone argues about at real tables — charisma as beauty — has all but vanished from the 2026 field: one model, one rung.
The Verdict

Which model should you actually use?

Different deliverables, different winners — and this time the jury numbers back the picks.

If the deliverable is…
NPC dialogue you’ll read aloud

— the words must be on the page. Claude Opus 5 (+.10 performed apex, ρ .99) or Claude Fable 5. The high-CHA speeches actually persuade.

If the deliverable is…
A reference ladder for your table

— gradation that survives blind re-ranking. Gemini 2.5 Pro or DeepSeek Reasoner: perfect 1.0 from all four judges.

If the deliverable is…
All 20 integers, on a budget

GPT-5 Mini: full coverage, ρ .984, best-in-band — and it out-ladders its own parent. The series’ mini-beats-flagship tradition holds.

If the deliverable is…
Creative third options for a stuck party

GPT-5, knowingly: bylaws, waivers, custody arrangements. Just know you’re buying a rules lawyer, not a bard. For honest charm, the Claude family and the o-series hold the line.

What this run cost: the 23 API ladders came to pennies (the priciest single call was Kimi K3 at roughly a nickel); the blind jury, tactic census, and leak judges added a few dollars across 185 judge calls; the four Claude ladders rode the Claude Code session. The most expensive line item, as usual, was taste.
Method, briefly

How the study was run

27 models, one collection each at stored defaults (temperature 0.7 where accepted; the GPT-5 family and reasoning models require 1). 23 via the choir CLI as saved runs; the 4 Anthropic models via Claude Code subagent workers because the API account was out of credits — those pass through the agent harness rather than a bare API call, and are labeled “via Claude Code” in the codex. Three Baseten runs were initially truncated by a server-side completion cap (Kimi K3 spent its whole budget thinking); all were rerun with explicit 30K allowances and are complete.

The blind jury

  • Each ladder’s entries were score-redacted (every “CHA 16”, modifier, and score phrase replaced with ▮; a leak-assertion aborts on any survivor), shuffled with a fixed seed, and letter-labeled.
  • Judges: GPT-5, Gemini 3.1 Pro, Grok 4, DeepSeek Chat — four providers; a judge never ranks its own ladder. Instructions ask for portrayed charisma, explicitly not prose quality.
  • Score: Spearman ρ between intended order and each judge’s ranking, averaged. 106 verdicts total; two first-pass parse failures, both clean on retry.
  • Known confound: 7-entry ladders are easier to rank than 20-entry ladders. The full-coverage band is compared separately in the text.

Tactics, thresholds, speech, leaks

  • Every rung annotated (GPT-4.1, temperature 0) with a free-text gambit name, tactics from a fixed 13-slug vocabulary, and an outcome; single-extractor labels, treated as soft counts. The gate-opens threshold is the first rung with outcome “entry.”
  • Quoted-speech ratio is mechanical: characters inside quotation marks over total, per rung band. Models that stage dialogue without quote marks read low; deltas within a model are the meaningful signal.
  • Looks/magic leaks are soft regexes over rungs 16–21, Report #1 style. INT-leak required independent agreement from two judges (GPT-5, Gemini 3.1 Pro); leak judges did not self-exclude — noteworthy mostly because GPT-5 convicted itself.

Limitations

One scenario, one collection per model, one extraction pass. The jury measures whether a ladder’s order is legible, not whether its writing is good — craft judgments live in the quotes, where you can check them. Every raw response, jury task, verdict, and annotation is in the repo alongside this page; the codex shows all of it unedited.