AI-Native Series · The Visionboard
The Visionboard Scoreboard: What Moved, What Did Not, and Three Futures That Carry Their Own Falsifier
1-minute takeaway — what you'll walk away with
What this is. The last episode grades the series with the instrument the series is about. Three readings of the same Visionboard, from production, across the days the five episodes shipped, on every gauge the design exposes: signatures, metrics, tripwires, odds, cards, ledgers, the solver's top three. Then three futures for the board, each stated with the observation that would prove it wrong, and the one schema change shipped with this episode.
Why it matters. A series about honest instruments that ended with a flattering summary would refute itself. The reading is that nothing on the board moved in the days the series was written, while everything around it did: five articles, a correction, a new sensor path, a translation runner. If you ship loops for a living, this is what an honest close looks like, and the futures are the three bets that separate a real loop from a described one.
What you can do. Copy the scoreboard shape: fixed gauges, dated readings from the production system only, and the first reading kept on the page even when it was wrong. Then write your own three futures with falsifiers, and give each a date.
The Visionboard — the series spine
- The Visionboard: Life GPS rebuilt with the human above the loop episode 1
- Life GPS, 2019 → 2026: what the first planner got right and what killed it episode 2
- Life GPS for one person: the daily director's act, and where the agent may not go episode 3
- Life GPS for a small team: roles not persons, load as a constraint episode 4
- The Visionboard scoreboard: what moved, what did not, and three futures with falsifiers this episode
What the Visionboard scoreboard is, who it is for, and the rule it has to obey
The season spec, written on 6 September before any episode shipped, set one test for the whole series: every episode ends in a change to the live board or its data, the series' scoreboard is the production reading re-measured at each ship date, and if the numbers do not move by episode 5, the series says so. This is episode 5. The numbers did not move. Here is the saying so.
The episode is for anyone who ships a loop and then has to report on it: a founder with a roadmap, an engineer with a control system, a student who built the 12-line week from episode 2 and wants to know what "did it work" looks like when the honest answer is "not yet." It has no analogy, declared on purpose. A scoreboard is direct exposition, and dressing it up would be the flattering the design forbids.
One rule governs every number below, and it was earned on day one of the series. Episode 1's first published reading said the board's card feed had gone dark on day 0. That number came from a local build that cannot read the private ledgers. The correction is on episode 1's page, and the rule it produced is in the season spec: every live number is read from production, never from a local build. The wrong reading is kept in the table below, marked, because an instrument's log includes its miscalibrations.
The scoreboard: three readings of one board
[MEASURED]Read the rows in two groups. The first group is the board's authority gauges: signed weights, signed metrics, signed milestones. All zero at every reading. Sixteen nodes, three tracks, twelve milestones, and not one signature in 34 days, including the two days in which the human above the loop was writing five articles about signing. Episode 1's reality mission was one signature before the weekly review. It has not happened. Episode 3's was tonight's journal filed as a card. The relationship track still shows zero cards in fourteen days.
The second group is the board's sensor gauges, and they tell a more useful story. The wrong first reading saw no ledgers at all. The corrected reading saw five of seven. Today's reading sees six of eight, because episode 3 added the relationship track's repository to the list the board reads, and that repository turned out to have a ledger with three cards in it, older than the fourteen-day window but real. The card count went from 211 to 213. Financial cards in the last fourteen days went from twelve to nine, which is not decay, it is the window rolling forward past the late-August burst. The solver's top three re-ordered as a result: the research node moved to second, because it is the only track where a move can still change the odds from null to a number.
[MEASURED]What moved is entirely around the board. Five episodes shipped in two days, each through the same gates. One correction merged the day after the error, with the error kept on the page. One sensor path added, one ledger newly readable. A translation runner built by a concurrent session, with its secret set, and zero of the twenty translated pages delivered yet, because the runner's first nightly run has not happened. And the schema change below.
What the readings settle, and what they do not
[EVIDENCE]They settle the design's central claim in the direction nobody wanted. Episode 1 argued that the human above the loop is a position you have to keep occupying; episode 3 built a protocol because the principle had already failed once; the scoreboard shows the principle failing a second time, during the very days the protocol was being written. Mitchell, Ghosh and Passi's five words, oversight degrades the overseer, now have a third data point on this board, and it is the author.
They also settle something about the instrument. Every gauge stayed honest under pressure to flatter: the odds did not creep above the floor because articles were being published, the tripwires did not clear because attention was being paid, and the solver's re-ordering had a reason that can be read off the data. An instrument that reports "no change" while its owner is working hardest is doing exactly what the design asked.
[OPEN]What they do not settle is why. Two candidate explanations survive the data. One: the tray is below the fold and nothing in the day forces a visit, so the interface is the bottleneck. Two: signing is a decision about which offer to price first, which cohort is the first cohort, which research bet to back, and those are the decisions the owner has been avoiding by building instead, so the interface is irrelevant. The season spec named both on day one. The futures below are built so that thirty days will tell them apart.
Three futures, each with the observation that would prove it wrong
| Future | What ships | The bet | Falsifier (dated) |
|---|---|---|---|
| 1. Signature-first | The signing tray becomes the first thing the board asks for on open, one proposal, yes or no with a reason, before the map renders; the pre-commit note from episode 3 is timestamped in the same step. | The interface is the bottleneck. | By 2026-10-07: if signatures are still 0 with the tray first, the interface was never the problem, and future 1 is retired. |
| 2. A sensor that cannot be silent | A per-track sensor-health tripwire: a track whose ledgers were unreachable, or whose nodes filed nothing in seven days, fires a named ticket; the journal card from episode 3 is the relationship track's designed feed. | Silence is a design defect, not a human one. | By 2026-10-07: if the relationship track has a sensor-health ticket every week and still zero cards, the feed was never the problem; the owner is. |
| 3. The double loop | The board measures whether a signed metric predicts a moved metric: for every signature, the delta on that track four weeks later, computed in code, with the setpoint re-calibrated when signed metrics keep not moving. This is the self-improvement gap named in this site's own notes: strong observation, strong evaluation, no actuator. | Signing is where the loop closes, so signatures should predict movement. | By 2026-12-01: if signed metrics move no more than unsigned ones over three months, signing is ceremony, and the human-only field needs a different design. |
Notice that future 3 cannot even be tested until future 1 or 2 produces a signature. That ordering is the honest shape of the next quarter.
The schema change shipped with this episode
[DESIGN]Episode 4 owed one thing before this episode: owner and signed_by on every node, with the self-signature refusal in the edit vocabulary and a test proving it. It ships here. Every node in the goal graph now carries an executor of record; a signature records who signed; and on a board with two or more people, the person who owns a node may not sign its weight, and a signature that does not say who is refused. On a one-person board the rule is dormant and the file says so, because a rule that silently blocked the only signer on a board with zero signatures would be a joke the instrument is not allowed to make.
// packages/core/src/goal-edits.ts — the refusal, verbatim in spirit
if (people.length >= 2) {
if (!by || !people.some((p) => p.id === by)) return refuse("signWeight needs `by` — on a team the signature must say who");
if (n.owner && n.owner === by) return refuse(`${by} owns ${id} and may not sign its weight — the signer is never its executor`);
}
n.weight = w; n.weight_signed = true; if (by) n.signed_by = by;
Six tests pin it: the owner may sign on a one-person board and is recorded; on a team, the executor of record is refused with the episode-4 sentence; another person signs and signed_by says who; a signature with no signer is refused; an unknown signer is refused; unsigning clears the record. The change is small on purpose. The team version of the board is a design until a second person exists, and the schema is the smallest thing that makes that design true rather than argued.
Build it in 30 minutes: the scoreboard contract
GAUGES (fixed for the life of the loop; add one only with a dated note)
authority : signed weights · signed metrics · signed milestones
sensor : cards total · newest card · cards per track in 14 d · ledgers reachable
instrument: odds per track (clamped) · tripwires fired · solver method · top-k
READINGS
one row per ship date; source = the PRODUCTION system, owner-authenticated
a wrong reading is kept, marked, with the cause (here: a local build without the token)
WHAT MOVED (two columns, never merged)
on the board : the gauges above, with deltas
around the board: articles, fixes, features, runners — real work, but not the loop
FUTURES
each: what ships · the bet it encodes · the observation that proves it wrong · a date
order them by dependency (a future that needs a signature cannot be tested before one exists)
INVARIANT
"no change" is a valid and reported outcome; a scoreboard that cannot say it is a brochure
Same story, five exits
| Reader | The decision this episode hands you | The one action |
|---|---|---|
| Student | "Did it work" is a table with dates, not a feeling at the end | Add a scoreboard to the 12-line week: three readings, one per check-in |
| Engineer | Fixed gauges, production source, wrong readings kept and marked | Write the gauge list before the next feature; refuse to add a gauge mid-quarter without a note |
| Founder | Work around the loop is not the loop; your board will tell you which you did | Sign one thing tomorrow. If you cannot, write down the decision you are avoiding instead |
| Executive | A team that publishes "no change" is more trustworthy than one that never has to | Ask for the scoreboard, not the roadmap, at the next review |
| Investor | Three futures with dated falsifiers is a plan; three futures without is a pitch | Ask which future is retired if the number does not move by the date |
Patterns, anti-patterns, and the mechanism
Patterns. Gauges fixed before the work. Readings from production only. The wrong reading kept and marked. "Around the board" in its own column. Futures with falsifiers and dates, ordered by dependency.
Anti-patterns. Counting articles as movement. Reading the instrument from the machine that cannot see the sensor. Retiring a gauge because it is embarrassing. A future without the observation that kills it.
The mechanism. A scoreboard is a pre-registered experiment on the loop's owner: fixing the gauges and the source before the work removes the degrees of freedom that let a summary flatter.
The first principle, in one sentence. A loop is real exactly when its owner can publish "it did not close" and the instrument that says so is the one they built.
Exit test
- Which gauges moved between the corrected day-32 reading and day 34, and why is each one not evidence the loop closed? (cards 211 → 213 and ledgers 5/7 → 6/8: sensor reach, not authority; the top-3 re-order: the window rolled)
- Why is future 3 untestable today?
- What does the series' first wrong reading have to do with the rule under which this scoreboard was read?
Reality Mission
Mine, and it is the same one it has been since episode 1, now with a date the board will hold me to: one signature on the financial leading metric, with its falsifier, before the review on Monday 14 September. Future 1 ships the week after. The next reading of this scoreboard is on 2026-10-07, and it will be posted whatever it says.
Yours: take the loop you own, write its gauges, and read them from the production system tonight. If the honest row is "no change," publish that row. It is the row this whole series was written to be able to publish.
Read next
Episode 5 of The Visionboard. Claims are tagged by class — measurement, evidence, design, open question. No analogy, declared. Every number above is logged with its timestamp and source in the repository's research dossier under docs/plans/; the season spec's 10X test is quoted verbatim from the version committed on 2026-09-06.
References
- This repository,
docs/plans/2026-09-06-visionboard-season-spec.md(§0, the 10X test; §3a, the measurement rule added after the correction) anddocs/plans/2026-09-06-visionboard-research-dossier.md(§C, the three readings with sources and timestamps; §C2, the 2019 re-run). - Production endpoints, owner-authenticated, read 2026-09-05 and 2026-09-07:
/api/goals(16 nodes, 0 signed; 3 leading metrics, 0 signed; 10 milestones, 0 signed) and/api/palace(213 cards; 6 of 8 ledgers reachable; newest card 2026-08-30); solver run withlib/gps.mjson that data. - Margaret Mitchell, Avijit Ghosh, Samir Passi, "AI Agents Push Humans Out of the Loop," arXiv:2608.23642, 2026 — "Oversight degrades the overseer."
- Paul Wu, The Visionboard, episode 1 (2026-09-06, corrected 2026-09-07); episode 2; episode 3; episode 4.
- This repository,
packages/core/src/goal-edits.ts(owner, signed_by, the separation-of-duties refusal) andscripts/test-goal-edits.mjs(the six tests), shipped with this episode;content/goals.json(peopleandowneron every node).