Paul Jialiang Wu agentic-portfolio ✉️ Free list
← Back to portfolio

AI-Native Series · ANDON · Agentic PR Review and CI/CD · Episode 3 of 5

PR Review Rules as Data: the Rule That Cannot Drift

Agentic PR review and CI/CD, episode 3. In 1956 Western Electric published the rules for reading a control chart so that a line worker and an engineer would read the same chart the same way. Most repositories keep their review rules in prose, where every reviewer reads them differently and an AI reviewer paraphrases them. This episode is the D in ANDON. It moves one repository's PR review rules into three data files, renders the human page from them, and adds a check, run in CI on every PR, that fails the moment the page and the data disagree. On 2026-09-09, three days after the check shipped, I made it go red on purpose.

By Paul Jialiang Wu · agentic-portfolio-lovat.vercel.app · 2026-09-09 · Episode 3 of 5

Cover: white ground with a black left rail. Eyebrow ANDON · AGENTIC PR REVIEW & CI/CD · EPISODE 3 OF 5 · DATA, NOT PROSE above the two-line serif title PR Review Rules as Data: the Rule That Cannot Drift, two grey lines: Western Electric wrote the rules down in 1956 so everyone read the chart the same way. A review rule that lives in prose stops being true; one that lives in a file can be checked. Three grey cards: WHAT IT IS, the D in ANDON, built, three files, one renderer, one check; WHO IT IS FOR, anyone with a contributing guide, engineers, founders, the curious student; WHAT YOU LEAVE WITH, a rules-as-data contract, source, render, check, plus the 5 Whys. A black band: THE ONE FINDING · ONE REPOSITORY · 2026-09-09 — 29 rules in one file, docs table rendered from it, one hand-edited id turned the check red. Prose can be wrong for a year and nobody knows. A file can be wrong for one commit.
The one finding, with its scope: one repository, one day, one hand edit.

1-minute takeaway — what you'll walk away with

What this is. Episode 3 of ANDON, a five-part series on agentic PR review and CI/CD. This one is the rule data, not prose, built into a real repository as PR review rules as data: the merge rules, the size thresholds and the response policy each live in one JSON file; the reviewer, the gates and the human-facing docs page all read from those files; and a CI check re-renders the docs table from the data and fails if the page has drifted.

Why it matters. For an engineer: a rule in prose is enforced by whoever remembers it, which in an agent-era repository is increasingly nobody, and the model will paraphrase it into something adjacent. For a founder: the 2026 policy census found that most top repositories now permit AI contributions and about half require disclosure; in my reading of those policies they are prose, and a policy only a human can interpret does not scale to a flood. For a student: this is the same move Western Electric made in 1956 and Shewhart made in 1931, applied to a contributing guide.

What you can do after reading. Put every review rule, threshold and exemption in a data file with an id, a level, the gate that enforces it and a written reason. Render the human page from the file and never type it. Add a check that re-renders and compares, then break it by hand once so you have seen it red.

ANDON — the series spine
  1. A · Assess — The PR flood is measured, not felt episode 1
  2. N · Never silent — Jidoka for repositories episode 2
  3. D · Data, not prose — The rule that cannot drift this episode
  4. O · Only in the envelope — Blast radius is a path list episode 4
  5. N · Negative test — A gate is installed only once it has gone red episode 5

What this episode covers, and why PR review rules as data come third

Episode 1 measured the flood and put the machine gate first. Episode 2 gave every gate a third state so that "could not run" is posted rather than silent. Both episodes assumed something this one has to earn: that the rules the gates enforce are the same rules the humans think they are enforcing. In most repositories that assumption is false in a quiet way. The rules live in a contributing guide, a pull-request template, a wiki page and a few reviewers' habits; the CI configuration enforces some subset of them with numbers that appear nowhere in the guide; and the AI reviewer, if there is one, gets a prompt that paraphrases the guide from memory. Three readers, three rulebooks.

The agent era makes this expensive rather than merely untidy. The 2026 census of 281 AI-contribution policies in the top two thousand GitHub repositories found that 83.3 percent permit AI use, 48.8 percent require disclosure, and 14.9 percent forbid it [1]. My own observation, outside the census: the policies I have read are prose that a human has to interpret per PR. A study of 23,247 agent-authored PRs found that 1.7 percent had a high inconsistency between what the PR said and what the code did, and the series' first episode already recorded that mis-described PRs still merge at 28 percent [2]. When the description can be wrong and the reviewer is often a bot, the rules that decide the merge cannot be a paragraph.

So this episode's rule is: a rule is data first, and prose is rendered from it. The worked example is the same repository as episodes 1 and 2, where 29 review rules, four size thresholds and a response policy now live in three files, and where, three days after the drift check shipped, I edited one rule id by hand to prove the check could go red. It did, on a local run, with the exit code recorded.

The oldest version of the rule: write it down so two people read it the same way

The 300-year record is a record of moving judgment from a person's head into an object two people can agree on. Honoré Blanc's demonstration at Vincennes on 8 July 1785, as reconstructed by Roser, was that "he disassembled 50 locks, picked pieces for 25 locks at random, and … assembled them again" [3]: interchangeability proven by an experiment anyone could repeat, not by an assertion. Bomford's 1821 Ordnance instructions called for "two plugs for the bore diameter, with one to pass through the barrel and the other not to enter" [4]: a rule that is a pair of physical objects, and a pass that has no opinion in it. Joseph Whitworth's standards came before limits did; the Science Museum's note on his wire gauge records that "The concept of tolerance or specification by limits of acceptability had not at that time been established." [5] A standard says what the right value is. A limit says what you do when the value is wrong. Gates need limits.

Shewhart supplied the definition the rest of the century ran on. In 1931 he wrote that "A phenomenon will be said to be controlled when through the use of past experience, we can predict, at least within limits, how the phenomenon may be expected to vary" [6]. Control is prediction within written limits. And then Western Electric's 1956 handbook wrote down the decision rules for reading his chart. The stated reason survives in the secondary record as one sentence: "Their purpose was to ensure that line workers and engineers interpret control charts in a uniform way." [7] That sentence is this episode. The chart already existed; what 1956 added was the rulebook, so that the person at the machine and the person in the office stopped disagreeing about what the chart said.

The same move keeps recurring. Juran's handbook was first published in 1951 [8]; ISO 9001 was released in 1987 as the descendant of a military specification [9]; and in 2009 the surgical safety checklist, a rule written as items on a card, was measured across 7,688 patients in eight hospitals: deaths fell from 1.5 to 0.8 percent and complications from 11 to 7 percent [10]. None of those are software. All of them are the same claim: a rule that lives in an artefact outperforms the same rule living in expertise, because the artefact is read the same way every time.

The kitchen picture from episode 1 still holds. A single cook keeps the recipe in their head. A restaurant writes it on a card. A chain writes it as a specification, in grams and degrees, because a thousand kitchens must read it the same way and none of them may improvise. The card is not a lesser form of the recipe. It is the only form that can be checked.

KitchenRepository
The spec card: grams, degrees, the one allowed substitution, signed offThe three data files: rules with id, level, gate and check; thresholds; the response policy with its exemptions and their reasons
The printed menu, generated from the card, never hand-typedThe docs page's rules table, rendered from the JSON between two markers
The nightly count that the menu still matches the cardThe drift check in CI: re-render, compare, exit 1 on any difference
Someone changes the menu with a marker penThe hand edit of one rule id that turned the drift check red on 2026-09-09 (a local run; the same command CI runs)
Table titled THE SERIES SPINE · FOUR HORIZONS · WHAT EACH ONE SETTLED ABOUT GATES with columns HORIZON, LANDMARKS (VERIFIED), THE PRINCIPLE A 2026 GATE INHERITS. Rows: 300 years (1785 to 2009): Blanc's 50 locks, the 1821 Ordnance bore plugs one to pass the other not to enter, the 1896 Toyoda loom that stops itself, the 1924 Shewhart control-chart memo, the 1950 andon cord, 1956 Western Electric rules, the 2009 WHO checklist; principle: go or no-go, never looks fine, the machine halts itself, anyone may pull the cord, rules written so two people read the chart the same way. 30 years (1999 to 2023): Beck and Fowler daily integration, Bacchelli and Bird 2013, bors, McIntosh 2014, Micco 2017 84 percent flakes, Google 2018 24 lines under 4 hours, SLSA and sigstore 2021, merge queue 2023; principle: small, fast, always-green main, test the merge result, separate flake from fault, a green build is not provenance. 30 months (2024-03 to 2026-08): Devin, Copilot agent, Octoverse 518.7M PRs, GitHub 25M to 90M per month, Stripe 1,000 agent PRs per week, enterprise 2.09x with human review 89 to 68 percent, agent PRs merge 68 to 80 percent, mis-described PRs merge at 28 percent, kernel humans sign; principle: the bottleneck moved to judging, provenance is a trailer nobody checks, no one formalised a review SLA, the merge queue is next. 30 days (2026-08-09 to 09-09), boxed in black: 09-01 Copilot can approve, 09-02 Agent Merge, 08-12 CodeRabbit 143M, 09-07 policy census 83.3 permit 48.8 disclose 14.9 forbid; principle: the AI reaches the merge button, whether it may press it is a repository setting, so the setting is the design. Footer: every landmark is in the series dossier with a verification status.
The series spine, unchanged from episode 1. This episode lives on the first row, at 1956, and on the last, at the policy census.

The 30-year version: Google wrote its reviewer standard down, and it is one sentence

The software lineage of the same idea is shorter and has one canonical example. Google's public engineering practices state the reviewer's standard in a single sentence: "reviewers should favor approving a CL once it is in a state where it definitely improves the overall code health of the system" [11]. That sentence is prose, and it works at Google because it sits inside a review culture that requires review for almost every change and a tooling culture built around small changes [12]. It is the recipe card of a restaurant with a very good kitchen. Copy the sentence into a repository with two maintainers and an AI reviewer and it becomes a control chart with no rules again: the two maintainers will read "definitely improves" differently, and the model will read it a third way on every run.

What the smaller repository needs is not a better sentence. It needs the sentence decomposed into things that can be checked: which conditions are machine-decidable and block, which are judgments the AI may raise with a citation, and which only a human may answer. That decomposition is a data structure, and once it is a data structure the docs page, the reviewer's prompt and the gates can all be generated from it instead of paraphrasing it.

The build: three files, one renderer, one check

1. The rules file

The repository's merge rules live in one JSON file, 29 rules as of the day of writing. Each rule carries six fields: an id, a level, a title, the source document the rule came from, the gate that enforces it, and the check the gate makes. The level is one of three. Ten rules are block: tenant isolation and write permissions, no secrets in the diff, a model change without a migration, a bilingual doc changed on one side only, the self-test contract of episode 5. These are enforced by scripts and a red one cannot merge. Thirteen rules are ai: the questions the AI reviewer may raise, each with a citation, including the three engineering principles the team added in one PR on 2026-09-06, object-oriented design, loop engineering and graph engineering. Six rules are human: the questions no script and no model may answer, the business scenario, the quality owner, the approval node, the closure criteria, the metrics, and whether to merge now, which only the domain co-founder can judge.

The same file carries the auto-merge envelope from the precursor article: the path prefixes and suffixes inside which a fully green, fully measured, finding-free PR may merge without a human, and the deny list that wins over any allow. The envelope is data for the same reason the rules are. When someone asks "can the AI merge a change to the workflow directory", the answer is a line in a file, not a memory of a conversation.

2. The thresholds and the policy

The size gate's numbers live in a second file. Warn at 400 effective lines, block at 3,000 or 60 files, block an agent-authored PR at 1,500, and count effective lines after excluding lockfiles, minified assets, snapshots, images, fonts and generated case data. The file's own documentation field cites the basis: the 2006 finding that review effectiveness collapses past a few hundred lines, and Google's small-change norm. A threshold that cites its reason can be argued with in a PR. A constant in a script cannot.

The response clock from episode 2 reads a third file: 48 hours to a first response, seven days to a decision, the maintainer list, the comment markers that do not count as a response, and the exempt PRs. Each exemption carries a written reason: the legacy PR from before the gates existed, and the deliberately open negative-test PR. An exemption without a reason is refused by the loader. That sentence was false when this article's reviewer read the first draft: the requirement lived in the file's own documentation field, as prose, and the loader accepted an empty string. The episode's thesis failing inside its own example is the kind of finding the reviewer-not-maker rule exists for, and the check was added and merged the same day, with a self-test that proves it refuses.

3. The renderer and the check

The human page that explains the review rules to the domain team is generated. One command renders the rules file as a Markdown table; the table is pasted between two HTML comment markers on the docs page; everything outside the markers is prose a human wrote, and everything inside is data a script wrote. A second command re-renders the table and compares it with what is between the markers. Any difference prints a one-line instruction to regenerate and exits with status 1. That command runs in CI on every PR and is the third structural criterion of the loop audit from episode 2.

On 2026-09-09 I proved it the only way a gate can be proved: I opened the docs page, found the row for the loop-engineering rule inside the markers, changed its id from R12 to R12x, and ran the check. It printed the drift message and exited 1. I restored the file and ran it again. It exited 0. That pair of exit codes is the evidence for the claim in the title; without it, "the rule cannot drift" would itself be prose.

Diagram titled DATA, NOT PROSE · THREE FILES ARE THE RULES · EVERYTHING ELSE IS RENDERED FROM THEM. Top row, SOURCE OF TRUTH · CHANGED ONLY BY PR, three black cards: ai_review_rules.json, 29 rules, G1 to G10 block, R1 to R13 AI, H1 to H6 human, each with id, level, title, source, gate, check, plus auto_merge envelope with allow and deny prefixes, deny wins; gate_thresholds.json, size warn 400, block 3,000 lines, 60 files, agent-authored block at 1,500, exclude lockfiles, snapshots, images, case data, provenance markers; sla_policy.json, first response 48 h hard, decision 7 d soft, maintainers, comments that do not count, exempt PRs each with a written reason, label names, dashboard title. Arrows down to the middle row, READ BY CODE · NEVER RETYPED, three grey cards: the reviewer plus the docs page, ai_review.py loads the rules into the prompt, render-rules prints the Markdown table pasted between rules begin and end markers; the size and provenance gates, check_pr_ready.py reads the thresholds, check_provenance.py reads the markers, no number lives in a script; the response clock, pr_sla.py classifies against the policy, a threshold change is a diff a human reads, the loader refuses an exemption with no reason. A dashed box: THE CHECK THAT MAKES PROSE HONEST · RUNS IN CI ON EVERY PR, ai_review.py check-doc re-renders the table from the JSON and compares it to the docs page, different means exit 1; proven on 2026-09-09, one rule id hand-edited inside the markers R12 to R12x gives red, restored gives green, the loop audit lists it as S3 and F3. Footer quote: Western Electric 1956, rules written down so that line workers and engineers interpret control charts in a uniform way.
Three files are the rules; the reviewer, the gates and the docs page are all rendered from them; one check keeps the page honest.

What changed when the rules became data

Three things happened in the repository that would not have happened with a prose guide. First, the rules acquired a count. Twenty-nine is a number you can watch: the loop audit's compounding criterion reads it from git history and reports that it went from 0 to 29 between the 1 September baseline and the day of writing, alongside the tests and the lessons. A guide has no count, so it has no trend. Second, the rules acquired a source column. Every one of the 29 names the document it came from; the loop-engineering rule, for instance, cites the lessons file and the response proposal. When the domain co-founder asks why the reviewer flagged a missing self-test, the answer is a row, not a recollection. Third, the rules became something the AI reads verbatim. The reviewer's prompt is built from the same file the docs page is rendered from, so the model is handed the same 29-row table on every run, 13 of which are its to answer, and the humans can see exactly what it was handed. The precursor article's rule that every finding must carry a citation now has a partner: every question the model is asked carries an id.

What did not change is worth saying. The prose around the table, the part that explains the three layers to a new team member with the kitchen analogy, is still prose, and it is still allowed to be. Data replaces prose where prose was pretending to be a rule. It does not replace explanation.

Five whys: why a review rule must be data before it is a rule

#Why?BecauseEvidence
1Why not keep the rules in the contributing guide?Because a guide is read by each reviewer differently and by the model differently on each run, so the repository has as many rulebooks as readers.Western Electric's stated purpose for writing the rules down [7]; Google's one-sentence standard works only inside Google's training and tooling [11]
2Why does the agent era make that worse?Because the human reader is increasingly absent and the PR description can be wrong: most AI policies permit AI use and, in my reading, are prose; and mis-described PRs still merge.[1, 2]; episode 1's coverage numbers
3Why render the docs from the data rather than write both?Because two hand-written copies of a rule are two rules that will diverge, and the divergence has no signal; a rendered copy has exactly one source.The drift check exists only because the page is rendered; a hand-written page cannot be compared to anything
4Why must the check run in CI and not in a release script?Because episode 2's rule applies: a check that runs somewhere the merge decision does not look has not been taken. Drift discovered at release time has already been merged.The check is a CI step and a loop-audit criterion; the hand-edit proof produced exit 1 in the same run shape CI uses
5Why is a written reason required on every exemption and threshold?Because Shewhart's "controlled" means predictable within written limits; a limit without a reason cannot be revised, only overridden, and an override is how prose grows back inside the data.[6]; the thresholds file cites its 2006 and Google sources; the policy's two exemptions each carry a reason

The root, in one sentence: a rule nobody can diff is a rule nobody can prove is still in force, and in a repository where the reviewer may be a model, a rule that cannot be proved in force is not a rule.

Build it in 30 minutes: the rules-as-data contract

Thirty minutes is the contract's budget for a reader who copies it; it is not a time I measured.

ElementSpecification
Rules fileOne JSON file. Each rule: id, level ∈ {block, ai, human}, title, source (the document or decision it came from), gate (the script or step that enforces it, or "human"), check (what is actually tested). No rule without a gate; a rule with gate "none" is a wish, and wishes go in the prose.
Thresholds fileEvery number a gate compares against, with a _doc field naming the basis. Exclusion lists for generated files. No numeric literal in a gate script.
Policy fileResponse times, the maintainer list, comment markers that do not count, exemptions as a map from id to written reason. The loader refuses an empty reason; a self-test proves it does.
RendererOne command that prints the human table from the rules file. The docs page contains two markers; the table lives between them; nothing else does.
Drift checkOne command that re-renders and compares the text between the markers. Different ⇒ print the regenerate instruction, exit 1. Runs in CI on every PR.
ConsumersThe AI reviewer builds its prompt from the rules file. The gates load the thresholds. The response bot loads the policy. None of them contains a copy.
ProofBefore the check is called live, hand-edit one id between the markers, run the check, record exit 1, restore, record exit 0. Keep both exit codes in the changelog.
Deliberately missingNo rule engine, no schema language, no YAML anchors, no templating beyond one table. Each would be a second place for a rule to live.
# The whole contract fits in the shape of the file and one comparison (ids real; text abridged).
{
  "rules": [
    {"id": "G1",  "level": "block", "title": "tenant isolation + write permission",
     "source": "AGENTS S1/S2", "gate": "check_tenant_perm.py", "check": "every query filtered by company; every write behind a role"},
    {"id": "R12", "level": "ai",    "title": "loop engineering",
     "source": "docs/05 §8.24 · proposal §3.4", "gate": "ai_review.py", "check": "does the change close a loop or open one?"},
    {"id": "H1",  "level": "human", "title": "business scenario",
     "source": "docs/01", "gate": "human", "check": "domain co-founder only"}
  ],
  "auto_merge": {"allow_prefixes": ["docs/", "plan/"], "deny_prefixes": [".github/", "scripts/"]}
}

# ai_review.py --check-doc
rendered = render_rules(load("scripts/ai_review_rules.json"))
current  = between(DOC.read_text(), "<!-- rules:begin -->", "<!-- rules:end -->")
sys.exit(0 if rendered.strip() == current.strip() else 1)   # different ⇒ red

Same story, five exits

ReaderThe one decisionThe one action
StudentA rule you cannot check is advice.Take one rule from any guide you follow and write it as id, level, gate, check.
EngineerData first; render the prose.Move your review rules into one file, generate the docs table from it, and add the compare-and-exit check to CI.
FounderYour AI policy is prose, and prose does not scale to a flood.Ask which of your contribution rules a script could check today, and put those in a file this week.
Executive1956 was the year the rules left the engineers' heads.Ask to see the file your review rules live in. If the answer is a document, you have the chart and not the rules.
InvestorA team whose rules have a count has a trend you can read.Ask how many review rules a team has and when the number last changed.

Patterns, anti-patterns, and the first principle

Patterns that held. One file per kind of rule. Six fields per rule, with the gate and the source required. Thresholds with a documented basis. Exemptions that cannot be written without a reason. A rendered page between markers. A check that compares and exits. A hand-edit proof before the check is trusted.

Anti-patterns this replaced. A contributing guide that listed rules no script enforced. Size limits that existed only as constants inside a gate script. A reviewer prompt that paraphrased the guide. A docs table that would have been typed by hand, with nothing to compare it against.

The mechanism, named. There is exactly one copy of every rule, and every other appearance of it is a function of that copy; the drift check is the function applied twice and compared, so a difference can only mean someone bypassed the function.

The first principle, in one sentence. A rule is the thing you can diff; everything else is commentary on it.

Reality mission

Episode 2 promised that by episode 3 the loop audit would have been re-run on a green main branch with its report committed. Done: the audit merged with its first report, its own two silent defaults were fixed the same day, the main branch's run went green, and the refreshed report landed in a docs-only PR that the envelope classified correctly and a human merged. Twenty-one criteria pass, none fail, two stay unmeasured: the AI has still not measured a PR, and Tier-1 has still not merged one. The three co-founder PRs still carry the first-response label, and the board still says so.

Season scoreboardEpisode 1 (2026-09-09, 14:00 UTC)Now (2026-09-09, 17:00 UTC)
Oldest open PR without a first response57 hours (three PRs)still those three; labelled, on the board
Gates with a self-test that proves they can go red0 of 69 of 9, the nine the audit's S5 criterion enumerates, enforced by a runner in CI
Agent-assisted commits since 1 September carrying a trace trailernot measured11 of 21; the provenance gate now requires it on new commits
Audit criteria unmeasurednot yet an audit2 of 23, both waiting on things outside the code

By episode 4, the auto-merge envelope in the rules file will have been tested against a real docs-only PR, with the branch-protection state reported honestly: on the free plan the envelope classifies but cannot yet merge, and the article will say which.

Read next

Episode 3 of ANDON · Agentic PR Review and CI/CD. Every number is scoped where it appears; the kitchen is a model, not a measurement. The three data files, the renderer, the check and the hand-edit proof exist in a private repository; their shapes are reproduced in the contract so nothing depends on access to it. The four-horizon dossier this episode cites from is logged in the portfolio's plans directory.

References

  1. Hora, Robbes, Zacchiroli, "We Permit the Use of AI, but […]", arXiv:2609.07542, 2026-09-07. arxiv.org
  2. Gong et al., Message-Code Inconsistency in Agent-Authored PRs, arXiv:2601.04886, 2026-01-08. arxiv.org
  3. Roser, 230 Years of Interchangeable Parts, AllAboutLean, 2015-07-08. allaboutlean.com
  4. Raber, Malone, Gordon, Cooper, Conservative Innovators and Military Small Arms: Springfield Armory 1794–1968, NPS 1989/2006, p. 138. npshistory.com
  5. Science Museum Group, Whitworth standard wire gauge, object co59414. collection.sciencemuseumgroup.org.uk
  6. Shewhart, Economic Control of Quality of Manufactured Product, Van Nostrand, 1931 (archive.org scan). archive.org
  7. Western Electric rules, Wikipedia (citing the Western Electric Statistical Quality Control Handbook, 1956). en.wikipedia.org (secondary; the handbook was not fetched)
  8. Juran Institute, History of Quality Management System. juran.com
  9. Quality Digest, 25 Years of ISO 9001, 2012-04-12. qualitydigest.com
  10. Harvard Gazette on Haynes et al., NEJM 360:491-9, 2009-01. news.harvard.edu
  11. Google Engineering Practices, The Standard of Code Review. google.github.io
  12. Software Engineering at Google, chapter 9, Code Review, 2020. abseil.io