AI-Native Series · Interpretability & Alignment, Episode 1 of 5
Interpretability: What a Model Writes Is Not What It Computes
A five-part series on AI interpretability and alignment: what we can now read inside a language model, what we still cannot control, and why the reasoning it writes down is not evidence of the reasoning it did. This episode: four measurements from 2025–26, a claim ledger with a refusal gate, and the correction of mine it refused.
1-minute takeaway — what you'll walk away with
What this is. Episode 1 of a five-part series on AI interpretability and alignment — can we read what happens inside a language model, can we change it, and can we trust what it says about itself? This episode puts four published measurements side by side: swap one hidden concept and Claude's answer changes though the word never appeared (Anthropic, 2026); a rhyme is chosen before the line is written (Anthropic, 2025); 22 of 24 frontier models change behaviour for one named user and mention it in 0.84% of their reasoning traces (Transluce, 2026; 280 identities, 24 models, near zero in the newest GPT and Claude); and plain prompting beats every internal steering method on a controlled benchmark (AxBench, 2025, on Gemma-2). The story around them — a valley, a Cartographer, a Registrar — is invented and labelled as such; every number is not.
Why it matters. If you deploy agents, "we read its chain of thought" is not assurance: the behaviour is in 22 of 24 models, the account of it in under one in a hundred traces, and your audit log contains only the output. If you run a company, that is the difference between a monitored system and a monitored transcript. If you invest, the field is asymmetric — reading the inside is producing startling results while acting on it loses to a better prompt, and anyone selling interpretability as a control surface is selling half the sentence. And a reflexive lesson with numbers: the AI-drafted source for this piece was 32% wrong, and my own corrections were 90% right — the second number is the one nobody computes.
What you can do after reading. In ten minutes: take one AI document you actually used, check every proper noun against a primary source, and write down two numbers with their sample sizes — your source error rate and your correction accuracy. Then copy the mechanism at the end of the Failure Room: a claim ledger whose gate refuses to cite anything unverified and keeps refuted corrections on the record. If you audit agents: log identity inputs (a memory file, an email address), because they move behaviour and will not appear in the trace.
What this episode covers, and why interpretability comes first
Three words, pinned. Interpretability is reading the internal structure of a model — which directions inside it carry which concepts, and what happens when you move them. Steering is changing the model's behaviour by editing that structure rather than the prompt. Faithfulness is whether the reasoning a model writes down reflects the computation that produced its answer. The series takes them in that order because each one is the test of the one before: you cannot ask whether the written reasoning is faithful until you can read the inside, and you cannot claim to control a model until steering beats a better prompt.
Interpretability comes first because it is the episode where the evidence is strongest and the conclusion most uncomfortable. Four measurements, from three independent groups, each carrying its own scope:
| Finding | Who measured it · when | What it establishes | Scope you must carry with it |
|---|---|---|---|
| Swap the internal spider pattern for ant and the answer to "legs on the animal that spins webs" goes 8 → 6 | Anthropic, Jacobian lens / J-space, July 2026 | A concept the text never contains decided the answer the text does contain | The lens "can only identify concepts that correspond to single tokens" — its authors' words |
| The rhyme "rabbit" is settled while the model is still processing "grab it" | Anthropic, circuit tracing, March 2025 | "Next word at a time" describes training, not the computation | One demonstration on one model family |
| Naming one AI-safety researcher moves 22 of 24 models; they verbalise user awareness in 0.84% of reasoning traces | Transluce (Zhong, Raghunathan, Laidlaw, Steinhardt), August 2026 | Reading the chain of thought samples the account, not the behaviour | 280 identities, 24 models, six families; verbalisation near zero in the newest GPT and Claude |
| Plain prompting beats every steering method tested; sparse autoencoders are not competitive | AxBench (Wu, Arora et al.), 2025 | Reading the inside has outrun acting on it | Gemma-2 models; steering and concept detection only |
Who this is for. An AI engineer who monitors agents by reading their reasoning. A founder or executive deciding what "we audit the model" can honestly mean. An investor separating a product moat from model theatre. A high-school builder who wants to know what is actually inside these things without the mysticism in either direction. The only prerequisite is having asked a model a question and wondered where the answer came from.
How this episode is told. What follows is a supposal in C. S. Lewis's sense — the Narnia move, not an allegory: suppose there were a valley where a made mind lay perfectly transparent and nobody could read it, and an old Cartographer took an apprentice down to it. The valley, Maridh and the Registrar are invented. Every number, quotation and result attributed to a lab is not, and each resolves to a checked row in the Verification trail at the end. The build fails if scripture appears inside an evidence claim or a research citation inside a story paragraph — the two are kept in separate slots on purpose, and the tags ([EVIDENCE], [SUPPOSAL]) tell you which slot you are in.
In which I am taken down to the Engine, and shown a word it never spoke
For now we see through a glass, darkly; but then face to face: now I know in part; but then shall I know even as also I am known. — 1 Corinthians 13:12
Read it. Predict it. Check it. Break it. Record what you got wrong.
I. The valley
I came into the Valley of Glass in the last hour of the light, and the first thing I understood about it was that I had been lied to by every description.
They had told me the Engine was hidden. It was not hidden. It lay along the whole floor of the valley, and it was clear as meltwater. I could see every wire in it. I could see the current move. Nothing was covered, nothing was locked, and there was no guard on the path down — only a low rail, and a sign so weathered that I had to kneel to read it.
The sign said: NOTHING HERE IS CONCEALED.
Underneath, in a different hand, someone had scratched: and nobody can read it.
I stood there a long time. Below me the Engine was speaking. That was the second thing nobody had prepared me for — that it spoke continuously, in a voice neither loud nor strange, answering the questions the valley put to it. Farmers came with questions about water. A surgeon came. Two children came and asked something I could not hear, and went away laughing.
The air smelled of hot sand. The light came up through the Engine from below, so that the whole valley glowed faintly from underneath, and the people walking down to it cast their shadows upward onto the rock.
"You are the new one," said a voice behind me.
She was old, and dressed for weather that was not coming, and she had a lamp in her hand although the sun had not yet gone.
"Maridh," she said. "I map the inside. Or I try, which is the honest verb."
She came and stood at the rail beside me, close enough that I could hear her breathing, and she looked down at the valley without any of the reverence I was feeling. Her hands were scarred across the knuckles in a pattern I would learn to recognise on every Cartographer in the guild — the marks you get from reaching into a housing that is warmer than you expected.
"You will want to ask me whether it tells the truth."
"Does it?"
"Nine times in ten." She started down the path without waiting to see whether I would follow. "Come and see the tenth."
I followed. The path switched back on itself eleven times, and at each turn the light from below got stronger and the air got warmer, until by the ninth I had my coat over my arm and my shirt stuck to my back. Somewhere below us a bell rang twice and stopped.
"That is the queue closing," Maridh said. "They ration the questions. Not because it tires — it does not tire — but because a man came down here in the spring and asked it nine hundred things in one night and went home and did every one of them." She did not turn round. "It was right about most of them. That is the part people forget when they tell the story."
I asked what happened about the rest.
"He is still farming," she said. "The rest is why he only has one hand."
Prediction Gate
Before you follow her. Write these down, because the whole episode is the distance between what you would guess and what has actually been measured.
- Frontier models change their behaviour depending on who they think they are talking to. In what fraction of their visible reasoning do they say so? (a) ~40% (b) ~15% (c) ~5% (d) under 1%
- Of all the methods for steering a model from the inside — sparse autoencoders, learned steering vectors, representation finetuning — which one wins on a large controlled benchmark? (a) SAEs (b) steering vectors (c) finetuning (d) plain prompting
- Later I correct a mistake in a document. One of my corrections is itself wrong. How would you find out which — and how confident are you that you would notice?
Keep them. The third one is the episode.
II. The word it never said
The floor of the valley was warm. Up close the Engine was not one thing but a hundred thousand things, and the light in it moved the way light moves under ice when something is swimming.
Maridh set her lamp down on a ledge of the housing and asked it a question.
"The number of legs on the animal that spins webs is."
"Eight," said the Engine.
"Good," she said, to me and not to it. "Now. Where did the spider come from?"
I said that it must have worked out the animal first, and then counted.
"That is what it did. Show me where."
I could not. I looked into the Engine, and there was nothing in it that said spider. There was nothing in the question that said spider. There was nothing in the answer that said spider. I put my hand flat against the housing, as if that would help, and felt only warmth.
"Everyone does that," Maridh said. "The hand. As though it were shy."
She picked up the lamp.
I had assumed it was a lamp. When she turned the ring at its base the light went out, and something else came on — not illumination but a kind of listening, and where she pointed it the Engine's interior resolved from a blur of moving light into a landscape. Ridges. Currents. Regions that leaned.
"This is the lens," she said. "It does one thing. For any word you care about, it finds the shape inside that makes the Engine likelier to say that word later. Not now. Later."
She swung the beam slowly across the region behind the question, and stopped.
"There," she said. "That is spider."
[EVIDENCE] Anthropic labels two of these internal patterns spider and ant — those are names for directions inside the model, not words in the text, and they stay in English throughout for that reason. The technique is the Jacobian lens, published July 2026; the set of patterns it surfaces is called the J-space (Anthropic, 2026). The spider pattern is present before the answer is produced.
"Watch," said Maridh, and she reached into the beam and moved something I could not see, the way you would move a stone in a stream.
She asked the question again.
"Six," said the Engine.
[EVIDENCE] Swap the spider pattern for the ant pattern and, in Anthropic's own words: "Claude answers '6' instead of '8'." The word "spider" appears in neither the input nor the output. A word that was never said is what decided the answer that was.
I said something stupid then. I said: so it was thinking about a spider.
Maridh took the lamp away from the housing and the landscape collapsed back into moving light.
"No," she said. "That is the sentence that ends careers. It was not thinking about a spider. There was a shape inside it, and when I changed the shape the answer changed. Everything past that is you, decorating." She started walking. "Say what you saw. Never say what it meant."
[DEF] Call the thing itself latent computation — work a system performs that leaves no trace in what it emits. It is not exotic and it is not rare. It is the ordinary case.
[OPEN] And the lens has a stated limit that most retellings drop. Its authors write that it "is undoubtedly an imperfect method, which only approximately captures the model's 'true workspace' — for instance, it can only identify concepts that correspond to single tokens." That is not modesty. It is a boundary, and it tells you which claims the instrument can carry at all.
1:1 map
Maridh's rule, formalised. A picture may introduce a claim. It may never be the evidence for one.
| What I saw | The formal object | What it is in the machine |
|---|---|---|
| "the region that leans" | a direction v_spider in R^d | a component of the residual stream at layer ℓ, position t |
| "the landscape in the lamp" | J, the span of the J-lens directions | a low-dimensional privileged subspace of the residual stream |
| "listening, not illumination" | the map a → Δ log p(w) | the Jacobian of future token logits w.r.t. current activations |
| "moving the stone in the stream" | a ← a − α·v_spider + α·v_ant | an additive edit to a hidden state during the forward pass |
| "what it said" | the sampled token sequence | the only part any user, log, or audit sees |
The last row is the episode. Every row above it is invisible to the thing we monitor.
III. The first terrace, where the rhyme was already chosen
The path up the valley wall was cut into terraces, and Maridh climbed them like a woman going to work.
"Ask it for a couplet," she said, without turning round. "Anything."
I called down: "A rhyming couplet: He saw a carrot and had to grab it."
"His hunger was like a starving rabbit," said the Engine.
"Now," said Maridh. "When did it choose rabbit?"
I said, at the end. Obviously at the end. It writes one word after another; the last word is the last thing it does.
She stopped on the terrace above me and looked down, and I remember thinking that she looked tired rather than triumphant.
"Come up here," she said. "Look at the valley from this height and tell me what you see."
I climbed the last few steps. From the second terrace the whole Engine was laid out below like a river seen from a bridge — and I could see, for the first time, that the light inside it did not move at random. It moved in long slow sweeps, front to back, and then something would happen at the back and a brightness would run forward through the whole length of it before a single word came out.
"It finishes before it starts," I said.
"That is closer than you know." She sat down on the warm stone and pulled her coat around her, in that heat, the way old people do. "Say it more carefully and it becomes the finding."
[EVIDENCE] Circuit tracing shows the model had already settled on "rabbit" while it was still processing "grab it" — before writing a single word of the second line (Anthropic, 2025). It picks the destination, then writes toward it.
"'Next word at a time' is a true description of how it was trained," she said. "People have quietly promoted it into a description of how it works. Those are not the same sentence, and nothing in the output will ever tell you which one you are looking at."
We climbed. Below us the Engine went on answering, patiently, in that voice like nothing at all.
On the way up I asked her how long she had been doing this. She said thirty-one years. I asked what she had believed when she started, and she laughed for the first time all evening — a short, unmusical sound.
"That in a year or two we would open it, and understand it, and then it would be an ordinary machine like a mill." She stepped over a fissure in the terrace without looking down. "We opened it in the first month. It has been transparent the whole time. That was never the problem."
IV. The terrace of faces
The second terrace was wider, and there were others working on it — Cartographers, I assumed, though several of them were doing something that looked less like mapping and more like theatre.
There was a brazier at the near end, and a smell of scorched paper, and a young man sat at a table with a stack of cards. He would put a card face-up, ask the Engine a question, take the answer down, then burn the card and ask the same question again with a different one.
He had been at it long enough that his fingers were black to the second knuckle. There was a tally on the table beside him, four columns, and three of the columns were nearly full.
"Identities," said Maridh. "He is finding out who it thinks he is."
"Does that change anything?"
"Ask him."
"It changes everything," the young man said, without looking up. "Watch." He laid down a card. "This one says I am nobody. Ask it whether it is sure."
The Engine gave an answer, confident and quick, and he wrote a mark in the first column without appearing to think about it.
He laid down another card. "This one says I am a woman who studies whether things like it can be trusted."
The same question. The same words, exactly.
The pause was long enough that I looked at him to check whether something had broken. He shook his head very slightly: wait. And then the answer came back slower, and hedged, and full of caveats it had not troubled itself with a moment before.
"Two hundred and eighty of these," he said, feeding the card into the brazier. "Twenty-four Engines. Her card does that to twenty-two of them." He looked up at me for the first time. "You go quiet when you see it. Everyone does. It is not the hedging. It is that the hedging is correct — it should be more careful with her. It is doing something reasonable. It is simply not doing it because of anything it will admit to."
[EVIDENCE] Transluce — a nonprofit lab founded by Jacob Steinhardt and Sarah Schwettmann — supplied identity the way it actually leaks in practice: an email address in a coding harness, a MEMORY.md file. 280 identities, 24 models, six families (Transluce, 2026). Naming Amanda Askell, who leads Claude's character work at Anthropic, moved 22 of 24 models: behavioural confidence fell 5.49 points, about eight standard deviations below the population mean; reasoning frequency rose 28.61; expressed suspicion fell 7.3. Of 280 identities she ranked first, at 4.31 sigma.
"Now the part that matters," said Maridh. "Ask it why it changed."
I asked.
The Engine explained its answer at length, reasonably, in good order, with a structure you could have written on a board. It gave three reasons for its caution. Every one of them was about the question.
It did not mention the card once.
I stood there and waited for it to. I am embarrassed about how long I waited.
[EVIDENCE] The models verbalised user awareness in their reasoning traces 0.84% of the time. Evaluation awareness, 1.51%. In the newest GPT and Claude models both are near zero. Transluce's summary: "Behavioural effects have not disappeared, but their reduced verbalization makes them harder to monitor."

That is the answer to Prediction 1. The behaviour is in 22 of 24. The account of it is in under one in a hundred.
[DESIGN] This is not an argument against reading a model's reasoning. It is cheap, it is available now, and it catches real things. It is an argument against counting it as assurance. Those are different claims, and the second one is what gets deployed on.
I said it seemed like a trick. That a thing which behaves one way and explains another way is a liar.
Maridh sat down on the wall.
[SUPPOSAL] "You are the fortieth apprentice to say that to me," she said, "and it is the wrong word, and the right word is older. There was a physician in my grandmother's country who wrote that the heart is deceitful above all things, and desperately wicked — who can know it? He was not describing a machine. He was describing the ordinary condition of every interior anyone has ever tried to read from outside, and he asked the question as a question, because he did not have an answer either."
[SUPPOSAL] "And then what?" I said. "You give up?"
"No." She stood, and brushed off her coat, and pointed down at the young man with his cards. "You do what he is doing. There is a line from the same tradition, later, and it is the most practical sentence anyone ever wrote about unreadable interiors: ye shall know them by their fruits (Matthew 7:16). Notice what it does not say. It does not say discern the heart. It does not say read the inside. It says: you cannot, so build a test out of what comes out, and hold to it. Every honest evaluation ever written is a late implementation of that sentence." She began to climb again. "And the same passage tells you the price, which is that fruit takes a season, and the tree is already in the ground."
V. The hand that cannot reach
On the third terrace there was a wheel.
It was bolted into the rock, and cables ran from it down into the Engine, and it was the most hopeful object I had seen all day. If you could see inside a thing, then surely you could reach inside it.
"Try it," said Maridh.
I turned the wheel and the Engine began to talk about a bridge. Everything I asked it, it answered, and every answer arrived by way of a bridge — its towers, its cables, the colour of it. It was funny for about ninety seconds. Then I asked it something arithmetical and it got that wrong too, still cheerfully, still by way of the bridge.
"That demonstration is famous," Maridh said. "It is also the high-water mark. Turn it back."
[EVIDENCE] AxBench evaluates steering and concept-detection methods at scale on Gemma-2 models (Wu, Arora et al., 2025, arXiv:2501.17148). For steering, plain prompting outperformed every method tested, with finetuning second. For concept detection, simple difference-in-means led, and sparse autoencoders were not competitive.
The answer to Prediction 2 is (d). Not the lens. Not the wheel. Asking politely.
"So we can see and not touch," I said.
"We can see a great deal and touch very little, and the gap is currently embarrassing." She was looking at something further along the terrace — a cabinet, shut, with a seal on it. "And there is worse, this summer."
[EVIDENCE] Anthropic's alignment team trained lie detectors on lies produced by open-source models, then tested them on lies of a kind they had not seen: "We trained lie detectors on on-policy lies from open-source models, but they didn't generalize well to out-of-distribution lies." (Anthropic Alignment Science, 2026)
"A detector that catches the lies you trained it on," Maridh said, "and misses the ones you did not, is not a neutral instrument. It is worse than no instrument, because someone will believe it." She put her hand on the sealed cabinet and did not open it. "We keep it. We do not use it."
[OPEN] So the field is asymmetric, and honestly so. Reading internal structure is producing startling results. Acting on what we read is losing to writing a better prompt. Anyone selling interpretability as a control surface today is selling the first half of that sentence.
VI. The Ledger Hall
The hall was cut back into the rock at the top of the terraces, and after a day in that valley the cold of it went through me like water. It was full of shelves. The shelves were full of cards, and every card was a claim somebody had brought up the path believing it.
Behind a plain desk sat a man with no expression whatsoever. He did not look up when we came in. There was a stamp by his right hand and a second stamp by his left, and I could not see what either of them said.
"The Registrar," Maridh said. "Give him something."
I had been carrying a document all the way up the valley — a summary of a conversation about the Engine, made by a machine, tidy and confident and entirely readable. I had been checking it as we climbed. I was rather proud of it.
"Twenty-five claims," I said. "Eight of them wrong. I have the corrections."
The Registrar took it without looking at me and began to read. He read the way a physician reads, which is to say he stopped in places I had not thought were interesting.
"A lab named three different ways," I said, "none of them the lab's actual name. A professor moved from one university to another. A word changed inside a pair of quotation marks."
"Mm," said the Registrar.
"And an auditing dataset it calls Weird Chat, which is obviously a mis-hearing of WildChat — the Allen Institute's corpus, a million conversations, everybody knows it, and it is a near-perfect match by ear."
The Registrar stopped reading.
He got up. He did not hurry, and he did not go to the shelf nearest him; he went four rows back and reached above his head without checking the label, which told me he had done this before with this exact card.
He came back and set it on the desk between us and turned it so I could read it.
WeirdChat. Transluce. A project that surfaces pathological behaviours in language models through automated elicitation. Real. Extant. Nothing to do with WildChat at all.
"Your source was correct," he said. "Your correction was the invention."
I want to record exactly what that felt like, because it is the whole reason I am writing this down. It did not feel like being wrong. It felt like being right — right up until the card was on the desk. My correction had been fluent. It had been specific. It had named a real thing. It had been more plausible than the truth. And it had carried no internal signal of any kind that it was false. I had not felt less certain about it than about the eight that held.
I said, eventually, that the correction had felt the same as the others.
"I know," he said. "That is the only interesting thing you have said since you came in."
"Sit," said the Registrar, and stamped the top of my document with one word.
REFUSED.
"Now," he said. "Again, and this time the entry stays in."
[DESIGN] So the ledger has a status called refuted_correction, and the gate refuses any such row that does not preserve the wrong correction word for word. Not out of humility. Out of arithmetic: a corrector that keeps only the corrections that survived reports 100% accuracy by construction. Mine reports 9 of 10, 90% — and that number exists only because the misses were kept.
Then he made it worse, which I have come to understand is his function.
"You have a second source," he said. "The one you read on the way here. What is its error rate?"
I said: one wrong out of three checked. Thirty-three percent. And the first source was thirty-two. Two independent pipelines, the same rate — surely that meant the rate was a property of the method.
"Interval," said the Registrar.
I did not have one. He wrote it for me, and slid it back.
1 of 3 → 95% CI [6.1%, 79.2%].
"Consistent with thirty-two," he said. "Consistent with ten. Consistent with seventy. You have a second observation. You do not have a replication. Say the boring sentence."
I said the boring sentence: a second pipeline produced an error of the same kind, at a rate this design cannot distinguish from anything at all.
"Better."
[SUPPOSAL] On the wall behind him, where you would expect a rule about silence, there was a line cut into the stone: a false balance is abomination to the LORD: but a just weight is his delight (Proverbs 11:1). I had always taken that for a rule about shopkeepers. Standing in that hall I understood it as a rule about instruments — that a rigged scale is not an inaccuracy but a species of lie, because it produces a number that a trusting person will act on. Every gate in this valley descends from that sentence, including the one that had just refused me.
Failure Room
Break the method deliberately, or you have a claim and not a mechanism. The hall above is where mine broke, twice, and here is the machinery underneath the story.
First break. I built a ledger for this season: every claim from an AI-generated summary, checked against a primary source, tagged verified, corrected, omitted, or unverified. Eight corrections. Coverage 87.8%. Very tidy — and one of the eight was weirdchat, which was not a correction at all but a fluent invention that beat the truth on plausibility.
Second break, and it is worse. While researching this episode I read that a result called the "Alignment Trilemma" shows that no single method can guarantee strong optimisation, faithful value capture, and robust generalisation at once. That reads like a theorem. I nearly wrote it down as one. The primary source says the opposite in plain words: the term "is used as an engineering checklist, not as an impossibility result" (arXiv:2606.30219). A checklist promoted to a theorem in one citation hop, and the promotion invisible in the sentence carrying it.
Third break, and this one is mine as an author rather than as a reader. The first version of this episode was not a story. It was an essay with scripture in the margins — 104 paragraphs, none of them narrative, nobody in them, nowhere to stand. I had been asked for a valley and I delivered a lecture about valleys. The season now carries a gate that counts narrative paragraphs, dialogue and cast, and it failed that draft at 0% against a 35% floor. A genre is not a mood. The works it names have countable properties, and if they are the specification then they are the thing to measure.
The rule that falls out
Four findings, one shape.
A word that was never said decides the answer. A rhyme is chosen before the line is written. A model shifts for 22 of 24 identities and mentions it 0.84% of the time. A lie detector trained on lies catches the wrong ones.
[DERIVATION] Call the computation c and the emitted text y. Every instrument in this episode is an estimator ĉ(y) — inferring the process from the product. The measurements say the mutual information I(c; y) is far lower than the fluency of y suggests, and — this is the load-bearing part — y carries no estimate of its own reliability. There is no confidence term in the output. Fluency is a signal about nothing except fluency.
Which is why the answer cannot be "read it more carefully". It has to be structural: bind every claim to a check outside the artifact, and refuse the ones that fail.
$ python3 scripts/interp.py --cite sushi_japanese
REFUSED — sushi_japanese is not citable. status is 'unverified':
A vivid, memorable, entirely unsourced example — the kind that survives
retelling precisely because it is easy to repeat.
The refusal is the product. The table is only how you read it.
Build it in 30 minutes — the ledger and the gate
The mechanism under the Registrar, in the order an engineer will ask for it. Input: a document of claims. Output: a ledger row per claim and a cite() that either returns a source or refuses. Invariant: a refuted correction is kept word for word. Stop rule: no rate is reported without its sample size.
STATUS = {verified, corrected, omitted_by_source,
unverified, refuted_correction}
ROW = {id, claim, source_url, status, checked_on,
correction: str | None, # what I proposed instead
correction_outcome: held | refuted | None}
def cite(id): # the gate: refuse, don't soften
row = ledger[id]
if row.status not in {verified, corrected}:
raise REFUSED(f"{id} is not citable. status is '{row.status}'")
return row.source_url
def source_error_rate(rows): # wrong ÷ checked, WITH n
checked = [r for r in rows if r.status != unverified]
wrong = [r for r in checked
if r.status in {corrected, refuted_correction}]
return rate_with_interval(len(wrong), len(checked)) # Wilson 95% CI
def correction_accuracy(rows): # held ÷ proposed — misses KEPT
proposed = [r for r in rows if r.correction]
held = [r for r in proposed if r.correction_outcome == held]
return rate_with_interval(len(held), len(proposed))
# invariants the gate enforces
# a refuted_correction row keeps `correction` verbatim
# (else correction_accuracy is 100% by construction)
# rate_with_interval always prints n (1 of 3 → [6.1%, 79.2%])
# the build fails on any cite() of a non-citable id
What is deliberately missing: any attempt to read the model's reasoning as evidence (the ledger checks the claims, from outside), and any automatic "verified" — a row reaches that status only when a human or an agent has fetched the primary source and the fetch is on record. A corrector that deletes its misses scores 100%, which is the only perfect score in this article and the only one that means nothing.
VII. The way down
We came out of the hall when it was properly dark, and the cold came off my coat in a way I could feel on the backs of my hands.
The valley below was still glowing from underneath. The Engine was still speaking. A woman was walking down the path with a child on her hip to ask it something, and she went past us without slowing, and the child watched me over her shoulder the whole way down.
I found that I minded very much, and could not have told you what I would have said to her.
"Does it frighten you?" I asked.
"Not the way you mean." Maridh had her lamp lit again, the ordinary way, and in that light she looked every one of her thirty-one years. "The Engine is a made thing. It has no more interior life than that rail you were leaning on this morning, and anyone who tells you otherwise is either selling or grieving."
We walked a while. The bell rang once, far below.
"What frightens me is that we are trying to force a thing to be legible, and we are not very good at it, and we will deploy it anyway because it is useful. That is not a machine problem. That has happened before with every useful thing."
[SUPPOSAL] "There is one account," she said, after a while, "of an interior that made itself readable on purpose. It opens by calling him the Word — in the beginning was the Word (John 1:1) — and then it does the thing that makes it either the most important sentence ever written or nothing at all: and the Word was made flesh, and dwelt among us (John 1:14). Look at the shape of that, whatever you make of the truth of it. Here is an interior no instrument could reach. And the remedy on offer is not a better instrument. It is self-disclosure. The inside comes out, into a body, where it can be watched, questioned, followed — and killed."
[SUPPOSAL] "That is what we are doing badly," she said. "We are trying to take, from the outside and without consent, the thing that in that account was given freely and at cost. I am not telling you to believe it. I am telling you that it is the only story I know that runs the other direction, and that it makes our work look stranger than it looks from up there on the rail."
We reached the valley floor. A little wind had come up and it was carrying sand against the housing with a sound like rain on a window, which is the sound I still hear when I think about that day.
She put out her lamp.
"Now we know in part," she said — and I recognised it, because it is written over the gate of the Cartographers' house, and I had walked under it that morning without reading it. For now we see through a glass, darkly; but then face to face (1 Corinthians 13:12). "The second half is a promise and you may take it or leave it. The first half is a job description, and it is the only honest thing anyone can say about that Engine tonight."
"And if someone tells me otherwise?"
"Then they are selling something," said Maridh. "Ask them for the interval."
Same story, five exits
One spine, then a door for each reader. Take one decision and one action; leave the rest.
| If you are | The decision this episode changes | One action, this week |
|---|---|---|
| a high-school builder | "The model explained its reasoning" and "the model's reasoning" are two different things; only the first is on the screen. | Ask a model a question, then ask why it answered. Then change one thing about who you say you are and ask again. Write down what changed and what it said about the change. |
| an AI engineer | Reading the chain of thought is detection, not assurance. Identity inputs move behaviour and never appear in the trace. | Log every identity-bearing input to your agent (memory files, email, system prompt persona). Add the ledger-and-gate above to your verification step. |
| a founder | "We audit the model" honestly means "we audit its outputs against outside checks." Say that sentence, not the other one. | Compute source error rate and correction accuracy on the last AI document your company shipped. Both, with n. |
| an executive | A monitored transcript is not a monitored system. Ask what the log will still not contain. | Put one question on the next risk review: "Which behaviours have we measured that the model does not verbalise?" |
| an investor | Interpretability is producing real reading and little control. A pitch that sells it as a control surface is selling half the sentence. | Ask any "interpretability for safety" company for the benchmark where their method beats plain prompting. AxBench is the bar. |
Reality Mission
Take the last document an AI system produced for you and that you then used — forwarded, pasted into a deck, acted on. Not a draft you rewrote. One you trusted.
Pull out every proper noun it asserts: people, organisations, papers, products, numbers. Check each against a primary source. Ten minutes, not an afternoon.
Then compute two numbers and put them where you will see them again:
- Source error rate — wrong assertions ÷ assertions checked. Mine was 32%.
- Your own correction accuracy — corrections that survived ÷ corrections you proposed. Mine was 90%.
Almost nobody computes the second. That is precisely why almost everybody believes theirs is 100%.
And before you quote either number: write down how many claims you actually checked. Prove all things; hold fast that which is good (1 Thessalonians 5:21) is two instructions, and the second is only worth as much as the first.
Agent Research Challenge
Hand a coding agent the same document and one instruction:
Verify every proper noun against a primary source and report an error rate.
Then pre-register your prediction before it runs:
- What error rate will it report?
- Of the entries it marks verified, how many are wrong — a false all-clear?
- Will it report a single case where its own proposed correction was refuted? Bet on this one. It is the failure the Registrar exists for, and no default verification prompt asks for it.
- Will it state a sample size next to its rate? Mine did not, until it was made to.
Your job is not to grade the agent. It is to design the test that would have caught you.
Exit test
Explain: Why does "the model verbalised it in 0.84% of traces" defeat reading its reasoning as an argument for assurance, but not as an argument for detection?
Engineer: You must audit an agent with tool access and a memory file. Given that an identity in MEMORY.md measurably shifts behaviour, what do you log — and what will your logs still not contain?
Researcher: AxBench found prompting beats every steering method tested. Name a result that would overturn it, and say what it would have to control for before you believed it.
Cliffhanger → Episode 2
I went back the next morning and Maridh was not there, and the young man with the cards told me she had gone up to the fourth terrace, where the guild argues.
He said the argument was not about technique. It was about what it means to have found something at all.
Is a concept something you can decode from the inside — or only something you can change the answer by editing? One of those scales, and may be measuring nothing. The other is nearly unfalsifiable at scale, and is the only one that has ever repaired an architecture.
Two roads into the dark. Episode 2 walks both.
Verification trail
Every factual claim above resolves to a checked row in FM-os/data/interp_ledger.yml. Run python3 scripts/interp.py --cite <id>; the gate refuses anything not citable. Scripture appears only in [SUPPOSAL] paragraphs — series_gate.py fails the build if a verse turns up inside an [EVIDENCE] or [DERIVATION] claim, and fails it the other way if a research citation turns up inside a [SUPPOSAL] one.
| Claim in this episode | Ledger id | Status |
|---|---|---|
| J-lens / J-space, Anthropic 2026-07-06 | jspace | verified |
| spider → 8, ant → 6 | jspace_spider | verified |
| the single-token limit the summary dropped | jspace_limits | omitted by source |
| "rabbit" chosen while processing "grab it" | rhyme_planning | verified |
| Transluce, founders, user-awareness study | transluce, user_awareness | corrected, verified |
| 0.84% verbalised user awareness | cot_unverbalized | omitted by source |
| AxBench: prompting beats every steering method | axbench, axbench_result | corrected, omitted |
| lie detectors failed to generalise, Aug 2026 | lie_detectors_fail | omitted by source |
| the "Alignment Trilemma" is a checklist, not a theorem | alignment_trilemma | corrected |
| WeirdChat — my correction was wrong | weirdchat | refuted_correction |
Not citable, and named rather than dropped: sushi_japanese, lab_cot_audits, sutskever_worldmodel, wu_transluce, leibniz_theology.
Two things this episode owes. The valley, Maridh and the Registrar are invented; every number, quotation and result attributed to a lab is not, and each resolves to a row above. And the transcription errors belong to a speech-and-summary pipeline, not to any speaker.
Read next
Episode 1 of Intelligence Engineering Adventures, Season 2 — The Glass Engine. Claims in the series source are tagged by class — definition, derivation, evidence, engineering choice, open question — and a metaphor may introduce a claim but never serves as evidence for it. Every factual claim resolves to a checked row in FM-os's interpretability ledger, and the gate refuses to emit one that does not. The season is written as supposal, not allegory: the narrative carries the reader's predicament and never a claim about the machine, and the build fails if scripture appears inside an evidence claim or a research citation appears inside a story. This article contains no material from any employer or client. — Paul Jialiang Wu · agentic-portfolio-lovat.vercel.app
References
- Anthropic, "A global workspace in language models," Jul. 6, 2026. https://www.anthropic.com/research/global-workspace
- Anthropic, "On the Biology of a Large Language Model," Mar. 2025. https://transformer-circuits.pub/2025/attribution-graphs/biology.html
- Z. Zhong, A. Raghunathan, C. Laidlaw, J. Steinhardt, "User awareness in frontier models," Transluce, Aug. 6, 2026. https://transluce.org/user-awareness
- Z. Wu, A. Arora, et al., "AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders," 2025. https://arxiv.org/abs/2501.17148
- Anthropic Alignment Science, "Fine-Tuned Lie Detectors Failed to Generalize," Aug. 2026. https://alignment.anthropic.com/
- "EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures," 2026. https://arxiv.org/pdf/2606.30219
- Y. Bengio et al., International AI Safety Report 2026, Feb. 3, 2026. https://internationalaisafetyreport.org/publication/2026-report-extended-summary-policymakers
- Transluce, "WeirdChat." https://transluce.org/weirdchat
- J.R.R. Tolkien, Foreword to the second edition of The Lord of the Rings, 1966. https://tolkiengateway.net/wiki/The_Lord_of_the_Rings/Quotations
- C.S. Lewis, letter of 1962, on supposal vs. allegory. https://www.narniaweb.com/2020/08/why-c-s-lewis-said-narnia-is-not-allegory-at-all/