AI-Native Series · Strategy OS
Your Strategy Deck Is Already Dead. We Made Strategy Pass CI Instead.
1-minute takeaway — what you'll walk away with
Most strategy dies as a deck: numbers with no receipts, forecasts dressed as facts, no mechanism to notice when the world changes. We built a strategy operating system that gates strategy like code — make check for documents, evidence-or-silence, humans holding every decision — then stress-tested it on a real enterprise engagement. 6 strengths, 7 weaknesses, receipts included.
A strategy operating system where every number needs a receipt, forecasts can't dress up as facts, and "ready" is a gate, not a vibe — stress-tested on a real enterprise engagement, scored honestly: 6 strengths, 7 weaknesses, 1 spec defect we found in our own design.
McKinsey Quarterly just put a number on the anxiety: enterprise LLM spend tripled in twelve months while token prices collapsed, 93% of surveyed organizations blew their AI budgets, and AI is heading toward a quarter of enterprise IT spend. Their punchline — "tokens are not value; tokens are the bill" — is really a punchline about strategy: most organizations cannot say which bets are working, what they truly cost, or what evidence their roadmap rests on.
And the instrument we use to manage those bets is still... a deck. A deck is a snapshot of thinking, frozen at the moment of maximum confidence, with numbers whose origins nobody can trace and no mechanism to notice when the world invalidates them. Ask a deck "why is this readiness score 3.2, and what would change your mind?" and you get silence.
So we built the alternative, and then did something most builders skip: we tried to break it on a real engagement and published the scoreboard.
Phase 0: twenty deliverables before one line of engine code
The project is a strategy operating system — a private meta-repo where a strategy is not a document but a connected, testable object: purpose → evidence → diagnosis → options → decision → architecture → execution → metrics → adaptation, with every link inspectable.
Before writing any engine code, we produced twenty pre-implementation artifacts — a portfolio inventory, a landscape report across nine tool categories, a framework map with 25+ framework cards, an ontology, three architecture options, seven decision records, a milestone plan — and here's the move that made it real:
We gave documents a CI pipeline. A fail-closed gate script checks every artifact: missing file, hollow content, absent required section — each a named failure.make check: 26/26. And the one thing the gate never does is infer permission — human approval lives in aStatus:line only a human edits, relayed verbatim.
Four design rules carry the system:
- Evidence or silence. No provenance ⇒ not evidence. An unsourced claim can exist only as a labeled assumption — "the model said so" is representable, and appropriately weak.
- Statements wear their class. Fact, claim, assumption, inference, forecast — a forecast can never render in the visual register of a fact. This is enforced, not requested.
- Humans exclusively author decisions. AI researches, models, red-teams, drafts. The decision node structurally cannot be written by an agent.
- Famous ≠ valid. Our framework research found evidence quality roughly inverted relative to fame — several of the most popular strategy tools have the weakest empirical support. So frameworks live as data with evidence ratings, and low-evidence frameworks are routed to ideation roles, never verdict roles.
The stress test: a real engagement, run by hand
Specs are cheap. So before building the engine, we ran the full methodology manually on a real, confidential enterprise engagement (details stay private — that's the point of the confidentiality rules working). Evidence register with provenance, falsifiable diagnosis, three genuinely different options, tradeoffs with stakeholder responses, a 90-day plan, the works.
What held (6 strengths, receipts in the repo): the classification discipline caught that one source was a summary-of-a-summary and forced an honest downgrade; the diagnosis schema rejected the first draft because it was a goal list wearing a strategy costume; and when temptation said "just write the chosen strategy," the ontology said no — the contract honestly reads Chosen Strategy: NONE, because no human had decided yet.
What broke (7 weaknesses, same repo): hand-classifying every statement was the slowest, most skippable step — unautomated, provenance becomes theater; the markdown evidence register can't answer "if this assumption dies, what else is in question?" at scale; and the sharpest one — the methodology recommended an architecture that mirrors itself, and had no adversarial pass to challenge its own priors. We logged that as a standing risk and added a mandatory red-team step. A system that grades everything else must be gradeable.
One genuine spec defect surfaced (a missing entity type our ontology needs), two new risks entered the register, and two core hypotheses moved from "hypothesis" to "partially supported" — with the evidence that moved them.
The mental model, runnable by a 15-year-old
A strategy is a claim factory. Every claim needs a receipt. When a receipt expires, the strategy must notice — by itself. If your strategy can't tell you what would change its mind, it's not a strategy; it's a mood with formatting.
The agentic-economics angle makes this urgent rather than aesthetic: McKinsey's own data shows agentic cost behaves as a distribution — the same task can vary in cost by a factor of ~30, and roughly 60% of an agentic task's cost is checking and repairing, not the first answer. Any AI roadmap built on point-estimate ROI is structurally wrong. Distributions, tail costs, and forecasts that get scored for calibration aren't rigor decoration — they're the minimum honest unit.
What's next
The repo stays private for now, by design. The implementation gate — human-owned, currently and honestly marked NOT APPROVED — decides when specification becomes engine. When it opens, the first vertical slice runs one strategic question end to end: evidence in, auditable strategy report out, every number answering "why?" with a chain.
Strategy that can't pass a gate is a deck. We'd rather ship the gate first.
AI-Native Series · Paul Jialiang Wu · love12xfuture · The engagement referenced is confidential; every claim about the system itself is verifiable in the project's gate logs and ledgers.