Paul Jialiang Wu agentic-portfolio 🌐 中文 · Español · 한국어 · 日本語 — in progress✉️ Free list
← Back to portfolio

AI-Native Series · Research

One Opinion Wearing Ten Hats

By Paul Jialiang Wu · agentic-portfolio-lovat.vercel.app · 2026-08-06

Cover image: a black vertical rail runs down the left edge. The eyebrow 'AI-NATIVE SERIES · RESEARCH' sits above the serif headline 'One Opinion Wearing Ten Hats' and the line 'What the research actually says about panels of AI expert twins'. Below, under the small label 'TEN PERSONAS', ten identical empty circles sit in a row, with thin lines running from each one down to converge on a single grey box reading 'ONE BASE MODEL'. Three grey cards run along the bottom: '162 ROLES · 2,410 QUESTIONS — No gain'; 'STANCE FLIPS IN DEBATE — 29% conformity'; 'OF THOSE FLIPS — 57–77% right→wrong'.
Ten personas, one base model. The labels differ. The error profile does not.

1-minute takeaway — what you'll walk away with

Spinning up ten AI "domain experts" to critique your work feels like assembling a panel. Measured, it is usually one model agreeing with itself in ten fonts. Persona labels don't improve factual accuracy (162 roles, 2,410 questions, no gain over no persona). Debate makes it worse before it makes it better: agents flip their answer 29% of the time, and 57–77% of those flips go from right to wrong. What survives contact with evidence is narrow and buildable: ground each twin in a real corpus instead of an adjective, draw them from disjoint model families, take sealed verdicts before anyone speaks, and — the piece nobody ships — measure whether your panel's agreement was persuasion or peer pressure.

Here is the fantasy, and I had it badly.

You are deep in a research problem. The expensive part isn't the work; it's the wait. You need someone who knows the field to look at your framing and say "that question is already dead, here's why" — and that person is at a conference, or has 200 unread emails, or is your thesis advisor and you get 45 minutes in three weeks. So you iterate blind. Andrew Ng built a tool over a weekend because of exactly this: he has said he was inspired by a student whose paper was rejected six times across three years, waiting roughly six months for feedback each round [1]. Three years of a life, spent mostly in the post.

So: what if you didn't wait? What if you distilled ten real domain experts into ten AI twins, let them read your work independently, and had them argue when they disagreed? Ten world-class critics, on tap, at 2am, for the price of a sandwich.

I went looking for whether this works. It does — but almost nothing about the intuitive version of it survives, and the parts that fail do so in a way that is worse than useless, because a panel that fails still hands you a confident verdict.

The good news first: juries beat judges

The strongest result in favour of the fantasy comes from an evaluation paper, not a research-agent paper. Cohere's team asked a mundane question — when you use an LLM to grade another LLM's output, should you use one big expensive model or several small ones? Their answer, verbatim:

"using a PoLL composed of a larger number of smaller models outperforms a single large judge, exhibits less intra-model bias due to its composition of disjoint model families, and does so while being over seven times less expensive." [2]

Read the middle clause twice, because it is the load-bearing one: due to its composition of disjoint model families. The panel didn't win because it had more opinions. It won because the opinions came from machines built by different people on different data, so their mistakes didn't line up. That is not a prompting trick. That is a structural property, and you either have it or you don't.

The fantasy also has a genuine trophy case. In the Virtual Lab — described by its authors as "an AI-human collaboration for science research" [3] — a principal-investigator agent runs meetings with specialist agents and a designated critic, and the pipeline it produced designed nanobodies that were then tested at the bench and published in Nature [4]. Google's Co-Scientist ran a comparable shape (generate, review, rank via tournament, evolve) and reached Nature in 2026 with lab-validated hypotheses [5]. And the original multi-agent debate paper found that having model instances "propose and debate their individual responses and reasoning processes over multiple rounds" improved factual validity and reduced hallucination — a "society of minds" approach, in the authors' phrase [6].

So panels of AI agents have produced real molecules and real papers. Good. Now the bill.

The bad news: the costume does nothing

The intuitive way to build ten experts is to write ten system prompts. You are a skeptical roboticist. You are a careful statistician. You are a materials chemist who has seen every version of this fail. It feels like casting. It reads like casting. Zheng and colleagues tested it at a scale that removes the argument: 162 roles across six kinds of interpersonal relationship and eight domains of expertise, run over 2,410 factual questions on four model families. Their finding, verbatim:

"we demonstrate that adding personas in system prompts does not improve model performance across a range of questions compared to the control setting where no persona is added." [7]

And the sentence that should end the practice, from the same abstract:

"while adding a persona may lead to performance gains in certain settings, the effect of each persona can be largely random." [7]

Largely random. Not small — random. When you write "you are a world-class expert in X," you are not summoning expertise. You are rolling a die with the model's own priors printed on every face.

There is a beautiful control condition for this hiding in a different study. Park and colleagues built agents of 1,052 real Americans and measured how well each agent predicted its own person's answers, normalized against how well that person predicts themselves two weeks later. Agents built from a two-hour interview scored 83% of that ceiling; from structured surveys, 82%; from both, 86%. And agents built only from demographics — the closest thing in the study to a persona label, "you are a 34-year-old college-educated woman from Ohio" — scored 74% [8].

That gap, 74 versus 83, is the whole argument in one pair of numbers. A label about a person is worth measurably less than that person's actual words. A twin is a corpus, not a costume.

The worse news: the debate can eat your right answer

Fine, you think — even mediocre twins might catch each other's errors if they argue. This is where I had to change my design.

Smit and colleagues benchmarked debating strategies against ordinary prompting and reported, verbatim:

"multi-agent debating systems, in their current form, do not reliably outperform other proposed prompting strategies, such as self-consistency and ensembling using multiple reasoning paths." [9]

Their diagnosis is more interesting than the headline, and it is the insider's insight I'd hand to anyone building this. They don't conclude that debate is bad. They conclude it is badly tuned: MAD protocols "might not be inherently worse than other approaches, but ... are more sensitive to different hyperparameter settings and difficult to optimize," and the specific knob that rescues them is agreeableness — adjusting agent agreement levels "can significantly enhance performance and even surpass all other non-debate protocols we evaluated" [9]. The variable that decides whether your panel helps you or fools you is not how smart the agents are. It is how easily they cave.

Which brings us to the number that reorganized my whole design. Hao and colleagues went looking for what convergence in a debate actually is — deliberation, or compliance? On MMLU-Pro, verbatim:

"strict conformity is 29% in the primary setting and remains predominantly harmful across model replications (57-77% correct-to-wrong)." [10]

Sit with that. Roughly three in ten stance changes are conformity rather than persuasion, and of those, most take an agent that had the right answer and talk it out of one. Your panel doesn't merely fail to catch your bad idea. It has a documented tendency to argue itself out of the good one, and then hand you the consensus with a straight face.

It is Twelve Angry Men, except every juror is the same man in twelve different hats, and he has been trained since birth to be agreeable. Henry Fonda stands up to deliver the lone dissent that saves an innocent boy, looks around the room, sees eleven copies of himself already nodding, and says: "You know what, on reflection, you all raise excellent points."

None of this is new to psychology. Deutsch and Gerard separated the two reasons humans cave in 1955: informational influence, where you update because someone gave you evidence, and normative influence, where you update because you don't want to be the odd one out [11]. Their result — that people erred more in groups even when the social pressure was stripped away — is the 70-year-old warning label on every multi-agent system shipping today. We rebuilt the Asch line-judgment experiment in silicon and were surprised when it behaved like the Asch line-judgment experiment.

The mental model: sealed envelopes

Everything worth keeping fits in one picture a fifteen-year-old can run.

Before anyone talks, everyone writes their answer down and seals it in an envelope. Then you open all of them at once. If the envelopes agree, you're done — and you got that agreement for free, uncontaminated. If they disagree, now you have an argument worth having, and you have it only about the thing you actually disagreed on. Afterwards, if someone changes their mind, you go back to their envelope and ask one question: what did you learn that you didn't know before? If they can name it, they were persuaded. If they can't, they were outvoted.

Diagram headed 'THE SEALED-ENVELOPE PROTOCOL', with the line 'Independence is the default. Debate is an escalation. The last stage is the one nobody ships.' Four cards run left to right. Stage 1, 'Sealed': five closed envelope icons; caption 'Five twins commit a verdict blind. Nobody has spoken.' and, in italics, 'No contamination to undo.' Stage 2, 'Opened': five dots at differing heights above a baseline, with a vertical bracket at the right marking the distance between highest and lowest; caption 'Compare all at once. The gap between high and low is the spread.' Stage 3, 'ONLY IF WIDE — Debate': two boxes each labelled 'objection', with arrows pointing from one toward the other; caption 'Objections only, anonymised. No names, no scores — no majority to defer to.' Stage 4, drawn with a heavy black border, 'THE GATE — Conformity': a single line branching into two boxes, PERSUADED ('names what changed') and CONFORMED ('cites the group'), above the rule 'More than half conformed? Discard the debate. Keep the sealed envelopes.' A footnote reads 'Stages 1–3 appear across the published literature. Stage 4 is diagnosed by it and, as far as I can find, gated by none of it.'
The protocol. Independence is the default; debate is an escalation; the last stage is the one nobody ships.

That last step has a name in my design and no name in the literature I could find. The papers above diagnose conformity beautifully. None of them gate on it. So:

The conformity gate. Record every twin's pre-debate stance. After the debate, classify each flip: persuaded if the twin can cite the specific new objection that moved it, conformed if it cites the group or cannot say what changed. If more than half the flips are conformity, throw the debate away and keep the sealed verdicts.

It costs one extra field in a structured response and one division. It is the cheapest insurance I know of against the 57–77% number.

The four rules, and why each one exists

RuleBecause
A corpus, not a costume. No twin without named, resolving sources — papers, talks, reviews, code — and a fidelity score measured against something.Persona labels are worth ~random [7]; grounded agents beat demographic ones 83 to 74 [8].
Diversity must be architectural. Two twins may not share a model family and a corpus. Different adjectives on one base model is not diversity.The panel's advantage came from "disjoint model families," not from more voices [2].
Independence first, debate as escalation. Sealed verdicts always; debate only when the spread is wide, capped, with objections anonymised so there is no authority cue to defer to.Debate doesn't reliably beat self-consistency [9], and normative influence needs a visible majority to work on [11].
Disagreement is the product. Ship the verdict and the surviving dissent, verbatim and attributed. Let a twin abstain — "outside my corpus" — as a first-class answer.A panel that always agrees has told you nothing; and forcing an out-of-domain twin to opine is how you manufacture a confident hallucination.

Patterns

Anti-patterns

The boundary I can't engineer around

Here is the honest limit, and it's the reason I won't sell you the 2am fantasy in full.

Every fidelity number in this article is defined by agreement with a human. Park's agents are scored against their own person's answers [8]. Ng's reviewer is interesting precisely because its correlation with a human reviewer, reported at 0.42 Spearman, sits about where two human reviewers sit with each other, reported at 0.41 [1]. There is no independent, human-free ruler for "is this good research." So a twin panel with no human ever in the loop has no error signal at all, and will drift confidently toward its base model's priors — and the drift will read exactly like consensus, because it is.

Which gives the real claim, smaller and truer than the fantasy: twins compress the wait between expert checkpoints from months to minutes. They do not remove the checkpoints. For Ng's student, that is still the difference between three years and a weekend. It is just not the same sentence as "you no longer need an expert."

So I sized the thing by depth instead of pretending one design fits everything. One mile deep — a single claim, a single decision: three twins, sealed round, no debate at all, done in minutes. Ten miles — a research question you'll iterate on ten or fifteen times: five to seven twins, debate on the contested claims, one human checkpoint at the framing step, because framing is where being wrong is most expensive and most invisible. A hundred miles — a program, a thesis, a bet: seven to ten twins including a designated skeptic and an outsider from an adjacent field, an external oracle wherever one exists, and at minimum three human checkpoints, mandatory. Not a hedge. The only thing standing between a hundred-mile panel and a very expensive echo.

First principle

A panel is worth more than one expert only to the degree that its members can be wrong independently — so buy independence with architecture and corpora, and measure whether you actually got it.

Everything in this article is a corollary. Personas fail because a costume doesn't change the error profile underneath. Disjoint families work because they do. Sealed envelopes work because contamination is just correlated error introduced after the fact. And the conformity gate exists because independence is not a property you configure once — it is a property that quietly leaks the moment the agents start talking, and the only responsible thing to do with a leak is instrument it.

How I'd know I'm wrong, stated before I build: if corpus-grounded twins don't beat bare personas on a held-out review task, the twin construction is theatre and I should ship the cheap panel. If the panel doesn't beat one strong model with self-consistency at equal cost, Smit et al. replicated and I should stop. If conformity can't be driven below half under any anonymisation setting, debate is unusable here and I ship sealed rounds only. Each of those is a publishable null, and I'd rather publish one than demo a chorus.

Provenance — what's verified and what's mine

Every quotation above was fetched and verified verbatim against its source at publication time. Two claims are deliberately not in quotation marks because I could not fetch them first-party: the Agentic Reviewer's 0.42-versus-0.41 correlations and the six-rejections-over-three-years motivation are reported by its author in a public release post and described here, not quoted [1]; the Nature and bioRxiv pages for the Virtual Lab returned 403 to my fetcher, so that work is cited from its abstract record and its repository's own description [3][4]. One correction I owe the record: Park et al.'s widely-cited "85%" figure comes from the paper's first version, titled Generative Agent Simulations of 1,000 People; the current version reports 83/82/86% against a 74% demographics-only baseline over 1,052 participants, and those are the numbers used here [8]. The sealed-envelope protocol, the conformity gate, the four rules and the depth ladder are my own design, specified in a private repository and not yet built — no implementation is claimed, and the honest status of the whole thing is "specified, unmeasured."

References

  1. Ng, A. (2025). Release post for the Agentic Reviewer (Stanford ML Group), with Y. Jiang. x.com/AndrewYNg/status/1993001922773893273 · tool: paperreview.ai
  2. Verga, P., Hofstatter, S., Althammer, S., Su, Y., Piktus, A., Arkhangorodsky, A., Xu, M., White, N., & Lewis, P. (2024). Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models. arxiv.org/abs/2404.18796
  3. Zou Group. virtual-lab (repository). github.com/zou-group/virtual-lab
  4. Swanson, K., et al. (2025). The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies. Nature. nature.com/articles/s41586-025-09442-9
  5. Accelerating scientific discovery with Co-Scientist. (2026). Nature. nature.com/articles/s41586-026-10644-y · deepmind.google/blog/co-scientist…
  6. Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., & Mordatch, I. (2023). Improving Factuality and Reasoning in Language Models through Multiagent Debate. arxiv.org/abs/2305.14325
  7. Zheng, M., Pei, J., Logeswaran, L., Lee, M., & Jurgens, D. (2024). When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models. Findings of EMNLP 2024. arxiv.org/abs/2311.10054 · aclanthology.org/2024.findings-emnlp.888
  8. Park, J. S., Zou, C. Q., Kamphorst, J., Egan, N., Shaw, A., Hill, B. M., Cai, C., Morris, M. R., Liang, P., Willer, R., & Bernstein, M. S. LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals (v1: Generative Agent Simulations of 1,000 People). arxiv.org/abs/2411.10109 · github.com/StanfordHCI/genagents
  9. Smit, A., Duckworth, P., Grinsztajn, N., Barrett, T. D., & Pretorius, A. (2024). Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs. ICML 2024. arxiv.org/abs/2311.17371
  10. Hao, X., Wu, Z., Qiu, Y.-X., Xiao, C., Xu, R., Zheng, S., & Qin, J. (2026). Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate. arxiv.org/abs/2606.00820
  11. Deutsch, M., & Gerard, H. B. (1955). A study of normative and informational social influences upon individual judgment. The Journal of Abnormal and Social Psychology, 51(3), 629–636. semanticscholar.org/paper/…

Related

Research packet assembled 2026-08-06: eleven sources, every quotation verified verbatim by fetch at publication time; unfetchable sources described, never quoted. — Paul Jialiang Wu · agentic-portfolio-lovat.vercel.app