Paul Jialiang Wu agentic-portfolio 🌐 中文 · Español · 한국어 · 日本語 — in progress✉️ Free list
← Back to portfolio

AI-Native Series · AI Economics

Everyone Is Building a Better Receipt

By Paul Jialiang Wu · agentic-portfolio-lovat.vercel.app · 2026-08-17

Cover image: white background with a black vertical rail down the left edge. The label 'AI-NATIVE SERIES · AI ECONOMICS' sits above a two-line serif headline, 'Everyone Is Building a Better Receipt', and two grey lines reading 'A $2.52 trillion market can tell you exactly what it spent and not one thing about what it got.' Below runs a thin horizontal timeline with four markers: three filled black dots labelled 1769 'fixed vs variable', 1865 'cheaper means more' and 1897 'pay for capacity held', then a hollow black-outlined circle labelled 2026 'still no warrant'. Three light grey cards run along the bottom: 'THE BILL — Verified — and about to be free, standardised May 2026'; 'THE CLAIM — Asserted — 98% now track AI spend, up from 31% in 2024'; 'THE WARRANT — Missing — “No one can answer that question yet.”'
Three of the four dates are settled history. The open circle is the whole opportunity.

1-minute takeaway — what you'll walk away with

Worldwide AI spending is on track for $2.52 trillion in 2026, and the industry's own annual survey answers the question is your AI providing value? with the sentence “No one can answer that question yet.” Almost everyone building in this space is building a better meter — and the meter got standardised in May 2026 and consolidated in January 2026, which is what happens to plumbing, not to moats. The durable business is not the number. It is the warrant: the evidence that makes the number safe for a finance team to rest weight on. Three places that warrant turns into a company, in the order I would actually build them: a Gate that only books savings which provably held quality, a Warrant ledger that joins spend to outcomes and refuses to certify what it cannot evidence, and a Desk that buys capacity instead of counting tokens. Plus the honest reason the third one might make the first two obsolete.

I have spent a lot of this year looking at tools that tell you what your AI costs. They are good. Genuinely — the engineering in this category is strong, the dashboards are fast, the integrations are one line of code.

And they all answer the same question, which is how much did we spend. Which is a question I already have an answer to. It's called the invoice.

Here is the thing that reframed the whole category for me. The FinOps Foundation runs the industry's largest annual practitioner survey. In the 2026 edition, under the heading of what makes AI spend hard, sits this [2]:

“Is your AI providing value? No one can answer that question yet.”

That is not a vendor's landing page. That is the trade body for the people whose job is literally to know, publishing that nobody knows. And in the same report, 98% of practitioners say they now manage AI spend — up from 63% in 2025 and 31% in 2024 [1][2].

Read those two facts next to each other, because together they say something specific. In two years the industry went from one-in-three to nearly-all on measuring the spend, and arrived at nearly-nobody on knowing what the spend bought. We built the meter at extraordinary speed and the meter turns out not to have been the hard part.

The mental model: the bill, the claim, and the warrant

Every dollar of AI spend has three questions attached to it, and you can hold all three in your head at once.

The bill. What did it cost? Tokens times price, plus the seats, plus the GPUs. This is measured. It comes off a gateway or an invoice and it is not in dispute.

The claim. What did it get us? This is asserted — and asserted, almost always, by the person asking for next year's budget. "We cut inference cost 40%." "This agent saves the support team nine minutes a ticket."

The warrant. Who says the claim is true? What evidence would survive a finance business partner who does not want it to be true? This is missing. Not thin. Missing.

Diagram titled 'The bill, the claim, and the warrant', with a black rail down the left edge. Three stacked rows, each a card. Row one, 'THE BILL', asks in serif type “What did it cost?”, described as tokens times price plus seats and GPUs measured at the gateway, footnoted 'OpenTelemetry graduated 21 May 2026 · FOCUS 1.4 ratified 4 June 2026'; to its right a solid black pill reads SOLVED with the note 'Being standardised into a commodity. Not a moat.' Row two, 'THE CLAIM', asks “What did it get us?”, described as stated by whoever is asking for the budget, routing saves 30–70%, footnoted that nobody shows the quality traded away; its pill is a white outline reading ASSERTED, noted as 'A number with no counter-hypothesis.' Row three, 'THE WARRANT', is drawn in a dashed black outline instead of a filled card and asks “Who says that is true?”, described as evidence a finance team, auditor or board can rest weight on, footnoted with the State of FinOps 2026 line 'Is your AI providing value? No one can answer that question yet.'; its pill is solid black reading EMPTY, noted 'Almost nobody has this layer at all.' A vertical rule on the right separates a column headed 'WHERE TO BUILD' listing The Gate (ship first, sells on savings that held, attacks row 2), The Warrant (the moat, becomes the system of record, owns row 3) and The Desk (the endgame, buys capacity not tokens, survives row 1). A closing line reads: 'Read down: each row is harder than the one above it, and worth more. Read right: build up, in that order.'
The dashed box is not a design flourish. It is the floor of the building nobody has poured.

The independent framing that matches this most closely comes from the cost-observability world itself, which describes the market as three layers — infrastructure, then inference and application, then business attribution — and then says the quiet part [1]:

“Most organizations today have partial coverage of Layer 1 and Layer 2, but almost none have Layer 3. The tools that exist are strong within their layer but do not cross boundaries.”

Layer 3 is the warrant. Every serious opportunity in this category is a way of pouring that floor.

An aside that is also the whole problem

You have seen the statistic. 95% of enterprise generative-AI pilots deliver no measurable impact. It comes from MIT's Project NANDA, it went mainstream through Fortune in August 2025, and it is now the opening slide of approximately every AI strategy deck on earth [10].

So I went to check it, because that is the least I can do in an article about measurement. And the checking is more interesting than the number.

The underlying report describes a funnel: roughly 60% of organisations investigated these tools, 20% reached a pilot, and 5% were successfully implemented. If 20% pilot and 5% ship, then about a quarter of the organisations that actually ran a pilot cleared the bar — which is roughly 75% of pilots falling short of a deliberately demanding threshold, not 95% of pilots failing [11]. Meanwhile the sample sizes got inflated in transmission: the study ran structured interviews with 52 organisations and surveyed 153 leaders, and mainstream coverage reported 150 interviews and 350 survey respondents, after which other outlets repeated the larger figures [11]. One critic called the methodology "irresponsible and unfounded" [11]. And the group publishing it happens to build the category of agent infrastructure that a finding of widespread failure would argue for [11].

I want to be careful: the direction is almost certainly right, and it is corroborated by people with no such incentive. IBM has put the share of AI initiatives delivering expected ROI at 25%. S&P Global found 42% of companies abandoned most of their AI projects in 2025 [10]. The phenomenon is real.

But savour the recursion for a second. The most-quoted evidence that enterprises cannot measure AI value is itself a badly measured number, amplified by outlets that got the sample size wrong, published by a party with an interest in the result. An industry that cannot measure its own measurement crisis is not going to be fixed by a fourth dashboard. That is the joke, and it is load-bearing.

Deep time, because this problem has been solved before

I ran this question across four windows — what is being said this month, what has actually become default over the last thirty months, what survived thirty years, and what survived three hundred. The long windows were more useful than the short ones, which is usually a sign you are looking at a structural problem wearing a new costume.

1769 — the licence bill and the meter bill are not a new pair

Josiah Wedgwood, the potter, spent the 1760s making beautiful things and losing track of what they cost. In a letter to his partner Thomas Bentley, he described what he had found in his own books [3]:

“Consider that these expences move like clockwork, & are much the same whether the quantity of goods made be large or small”

That sentence is the discovery of the distinction between fixed and variable cost, and it is 257 years old. Wedgwood worked out that his big money — moulds, rent, fuel, bookkeepers, wages — did not move with volume at all. The 1772 downturn then killed a great many of his competitors and did not kill him, because he was the one who knew which of his costs were clockwork [3].

Now look at an enterprise AI estate. Seat licences, platform fees and reserved capacity are the clockwork. Tokens are the variable. Practically every tool in this category meters the variable beautifully and treats the clockwork as somebody else's spreadsheet. Wedgwood would have found that funny, and then he would have gone and joined the two, because that is the entire trick.

1865 — why your bill grows while your prices fall

William Stanley Jevons noticed that more efficient steam engines did not reduce Britain's coal consumption. They increased it, by making coal-burning economical for uses that had previously been too expensive to bother with. Efficiency expanded demand faster than it shrank per-unit cost.

This is now the single most reliable fact about AI economics. Per-token prices for a given capability have been falling at something like an order of magnitude a year, and total spend has gone up anyway — enterprise AI spending reportedly rose several hundred percent across 2025 while unit prices fell. The structural version has been formalised: as inference gets cheaper, teams rationally buy more architecture — deeper reasoning, longer context, more agentic loops — and those choices multiply token consumption faster than the price decline reduces it [4].

Two consequences, and they point in opposite directions, which is why this window matters.

First: the market is guaranteed to grow. Anyone whose business depends on enterprises caring about AI spend has a tailwind that does not require a single customer to become more wasteful.

Second, and less comfortable: optimisation does not shrink the bill. It shifts what the bill buys. If you sell "we cut your token cost 40%" and the customer's invoice goes up 90% next quarter because they shipped three more agents, you have delivered exactly what you promised and lost the renewal. This is the trap I would most expect a first-time founder in this category to walk into, and Jevons called it in 1865.

1897 — the answer already exists, and it is a two-part tariff

Samuel Insull ran Chicago Edison and could not make electricity pay. On a trip to England he found that Arthur Wright, the engineer at Brighton, had built a meter that measured two different things: how much electricity you used, and how fast you used it. Peak demand — the capacity you needed held open for you — was a quantity nobody was billing for. The Wright demand meter let Chicago Edison bill on each customer's actual monthly maximum demand, and by 1897 Insull was selling the city a two-part rate: a fixed charge for the capacity you reserve, plus a very low charge per unit consumed [5].

What followed is the part worth memorising. Insull now had a reason to want customers whose peaks fell at different hours — streetcars in the morning, factories through the day, homes at dusk — because every customer who filled a valley in the load curve made every other customer cheaper to serve [5]. Load factor became the number that mattered. Prices fell and consumption exploded, which is Jevons again, on schedule.

Hold that thought for about nine hundred words. It is the endgame.

And the reason a regulator should be on your reading list

In June 2025, a board member of the PCAOB — the body that inspects the firms that audit America's public companies — gave a speech to railroad accounting officers titled Stop, Look, and Listen: The Legacy of American Railroads and the Dawn of AI. His account is that railroads "created nearly all of the basic techniques of modern accounting"; that the first written corporate audit report was issued in 1827; and that the Interstate Commerce Act of 1887 produced the first independent federal regulatory commission, which became the model for later agencies including his own [12].

His closing link to the present is short: "Whether it is the telegraph, automation, or generative AI, auditors must remain skeptical to still meet their obligations to investors." [12]

So the historical sequence, from the mouth of a regulator, is: a new technology creates unmeasurable scale → someone invents the unit → the unit gets standardised → the standard gets audited → the audit gets regulated. We are, right now, somewhere between "the unit gets standardised" and "the standard gets audited." That is a very specific and very good place to start a company, and it tells you which of the three ideas below is the durable one.

The last thirty months: the meter became plumbing

This is where the short windows earn their keep, because two things happened that most people building here have not fully priced in.

The schema got standardised. On 21 May 2026 the CNCF announced that OpenTelemetry had graduated — its highest maturity level, and the formal end of the argument about a default observability standard [7]. Alongside it, the GenAI semantic conventions define standard attributes for model name, input and output token counts, finish reasons, tool calls and results. Those conventions are already emitted by VS Code Copilot, OpenAI Codex and Claude Code, which means their runs are readable in any backend that speaks the protocol [7]. The conventions are still experimental. The direction is not.

And the billing schema followed. The FinOps Foundation's FOCUS specification — one common format for cost and usage data across vendors — ratified version 1.4 on 4 June 2026, adding, among much else, a Contract Commitment dataset covering payment models, lifecycle status, discount rates and fulfilment intervals [9]. Remember that detail. A standards body does not add a commitment dataset unless commitments are about to be where the money is.

Then the category consolidated. On 16 January 2026, ClickHouse announced it had acquired Langfuse — the most widely deployed open-source LLM observability platform, ending 2025 with over 20,000 GitHub stars and 26 million SDK installs a month, in use at 63 of the Fortune 500 — alongside a $400M Series D that took ClickHouse to a $15 billion valuation [8].

Put those three together and you get one conclusion, which is the thesis of this piece:

The meter is not the moat. When the schema is a public standard, the traces are emitted by the coding agents themselves, and the leading open-source collector belongs to a $15B database company, "we can show you your token spend" is a feature of your data warehouse. It is plumbing. Plumbing is a wonderful business if you are ClickHouse and a terrible one if you are three people with a seed round.

The last thirty days: the argument the market is having right now

Two live positions, and I think the tension between them is where the real design work is.

The mainstream position, from this year's FinOps X coverage, is a shift in the denominator: "The metric to chase is value per token, not cost per token" [13]. Cost attribution has moved down from the cloud bill to the token and now to the individual agent run and session, and the output metrics people want are cost per task, cost per outcome, cost per interaction [13]. Secondary analysis of the same community puts the share of practitioners actually applying unit economics to AI spend at under 20% — I could not confirm that figure in the primary survey data, so treat it as directional rather than exact [13].

The contrarian position is sharper and deserves to be taken seriously, because if it is right it invalidates a large fraction of this category. Yossi Hasson argues that AI pricing will follow electricity and bandwidth — from per-unit metering to reserved capacity — and that the whole discipline of token thrift is a transitional artefact [6]:

“Shaving tokens in 2026 is shaving kilobytes off GIFs in 2003. The work is real and the savings may be real. The whole problem is about to disappear anyway.”

And his one-line version of where value actually sits:

“Whatever is scarce here, it is not tokens. It is the guarantee.”

He is describing Insull's two-part tariff, arriving 129 years late. And note that FOCUS 1.4 just shipped a commitment dataset, which is the standards body agreeing with him in schema form.

I think he is directionally right and slightly early — and here is the important part: his argument kills token-counting products and strengthens warrant products. If the unit becomes reserved capacity, then "did we use the guarantee we bought, and did it produce anything" becomes a harder and more valuable question, not an easier one. Every idea below is designed to survive his being right.

Three places this becomes a company

I looked for ideas that satisfy four conditions at once: the pain is documented rather than imagined; the artefact accumulates into a moat instead of a screenshot; a $15B database company shipping a dashboard next quarter does not kill it; and it still works if Hasson is right about capacity. Three survived. I have ordered them the way I would actually sequence them, which is not the order of their eventual size.

① The Gate — savings that provably held quality

The pitch in one line. Model routing and prompt compression reportedly cut cost 30–70% [15], and nobody selling that number shows you the quality they traded away to get it. The Gate is the system that continuously tests a cheaper configuration against an evaluation suite and only promotes it if quality holds — and issues a receipt saying so.

The mechanism. Take a live workload. Generate candidate configurations: smaller model, shorter system prompt, a cache breakpoint, less reasoning effort. Run each candidate against a graded suite drawn from the customer's own traffic. Promote a candidate only when it clears the quality bar with no unmeasured cases, and record cost, latency, cache behaviour and the eval result as a signed artefact. Roll back automatically when the bar breaks.

Why now. The inputs finally exist. Standardised traces mean you can reconstruct what a workload actually did without a bespoke integration per customer [7]. And the buyer's objection has flipped: nobody senior now asks "can you cut our inference cost." They ask "how do I know you did not quietly make the product worse," which is a question about evidence, not engineering.

The moat. Not the router — routers are a weekend. The moat is the accumulated evaluation suite and the promotion history: after a year, you hold a per-customer record of which trades held and which broke on their real traffic. That is a dataset a competitor cannot buy, and it makes each subsequent change cheaper to certify. It is also the thing that makes the customer unable to leave, in the good way.

Buyer and wedge. Platform or AI-infrastructure leadership, and it sells on a controlled comparison rather than a promise. Land as a one-workload audit: here is your current cost, here are four cheaper configurations, here are the two that held quality and the two that did not. The two that failed are the credibility.

Effort: Small-to-Medium. The measurement harness is small. The eval suite is the hard part, and it has a genuine cold-start problem — a customer who cannot define what "good" means for their workload cannot be helped, and a meaningful share cannot.

What kills it. A model provider shipping this natively as a free routing feature, which is entirely plausible; the defence is being cross-provider and independent, because the party selling the tokens is the wrong party to certify the downgrade. And Jevons: never sell "your bill will fall." Sell "your bill will buy more, and here is the proof it did not cost you quality."

What would falsify it. If enterprises turn out to accept vendor-reported savings without independent quality evidence, there is no business here. Test cheaply: offer ten teams a free audit and see whether any of them ask for the eval methodology. If none do, stop.

② The Warrant — the ledger that refuses to certify

The pitch in one line. Join AI spend upward — call, session, workflow, business outcome — to produce cost per outcome, and refuse to publish a value claim that lacks evidence, marking it unverified instead. This is the Layer-3 floor [1].

The mechanism. Three joins and one discipline. Spend joins to workflow via whatever correlation key the estate has. Workflow joins to outcome in the systems where outcomes actually live — tickets resolved, close tasks completed, contracts reviewed, cases closed. Outcome joins to a baseline, so cost per outcome has something to be compared against. The discipline is the product: every claim carries its evidence and the counter-hypothesis it had to beat, and claims that fail are displayed as rejected, with the reason. Claimed savings and verified savings are separate columns, and nothing crosses from one to the other without evidence.

Why now. Because the demand statement has already been published by the trade body — "no one can answer that question yet" [2] — and because the corroborating failure data is now embarrassing enough to move budget: a quarter of initiatives delivering expected ROI by IBM's count, 42% of companies abandoning most projects [10]. There is also a documented root cause that this product attacks directly: a large share of AI projects were approved on projected returns nobody revisited, and had no agreed definition of success before work started [10]. A ledger that forces the definition before the spend is the fix for that specific failure.

The moat, and it is the best one in this article. The outcome side of the join is irreducibly customer-specific. It lives in a bank's ticketing conventions, a manufacturer's close calendar, an insurer's claim states. OpenTelemetry standardises the trace; nothing standardises what "resolved" means at a particular company. Building that join is weeks of unglamorous work per customer, and once built it becomes the system of record that finance quotes in board papers. Nobody rips out the system of record to save a subscription. This is also the layer the PCAOB sequence predicts becomes auditable, and then required [12] — a business that is heading toward "you need this to sign something" is a business with a long runway.

Buyer. The CFO's organisation, with the CIO as co-sponsor. Notably not the same buyer as most of this category, which sells to engineering. That is an advantage: different budget, larger, and less crowded.

Effort: Medium-to-Large, and the honest reason is that it is a data-integration business wearing an analytics costume. Whoever builds it should say that out loud in the first meeting, because the customers who flinch at it were never going to succeed with it.

What kills it. Being too slow to first value — if the first join takes a quarter, the pilot dies before the insight lands. Mitigation: ship the Gate first and let it fund the integration time. Also, cynically, the risk that nobody actually wants a verified answer, because an unverified one is more useful to the person defending a budget. This is real. It is why the first design partner should be a finance function, not an AI team.

What would falsify it. If three finance organisations in a row prefer a confident estimate to an evidenced range, the market wants theatre, and this product is the wrong shape for it.

③ The Desk — buy capacity, not tokens

The pitch in one line. As AI pricing shifts from consumption to reserved capacity and committed contracts, someone has to decide how much to commit, at what term, across which providers, and then run the load factor. That is a desk, and it is Insull's job description from 1897 [5].

The mechanism. Forecast demand per workload. Decide the commit-versus-on-demand mix. Route traffic to fill the valleys in reserved capacity before it spills to spot pricing. Price internal chargeback so teams see the capacity they hold, not just the tokens they burn — a two-part tariff, internally. And keep a calibration record of your own forecasts, so the desk's credibility is measured rather than asserted.

Why now. Three independent signals converging: FOCUS 1.4 shipping a Contract Commitment dataset in June 2026 [9]; the informed contrarian case that the scarce thing is "the guarantee" [6]; and the plain fact that inference has become one of the largest line items in enterprise AI budgets [1]. Cloud's equivalent — reserved instance and savings-plan optimisation — supported several substantial businesses, and it is arriving in AI with a standards body already laying track.

The moat. Forecast quality compounds. A desk with two years of demand history and a scored track record of its own predictions prices commitments better than a new entrant, and the savings are arithmetic rather than narrative — which makes it the easiest of the three to sell and the hardest to argue with.

Effort: Medium. The modelling is tractable. The dependency is real: this business needs providers to actually offer rich commitment structures, and today that varies a great deal by vendor.

What kills it, and it is worth being blunt. If providers keep pricing simple and generous, there is nothing to arbitrage. This is the highest-variance of the three: the biggest if the capacity thesis lands, and close to nothing if it does not. It is a bet on a market structure, not on a pain.

What would falsify it. Watch provider pricing pages for four quarters. If committed-capacity tiers are not proliferating by mid-2027, the thesis is wrong and the honest move is to say so rather than to keep the deck.

The sequencing is the actual insight

These are not three companies. They are one company in three acts, and the order is not the order of size.

Build the Gate first — it is the smallest, it sells on a controlled comparison, and it earns the right to be in the trace data. The Gate's evidence trail is the Warrant's seed corpus: once you are certifying that changes held quality, you are already holding the per-workload evidence the value ledger needs. The Warrant makes the Desk credible, because you cannot responsibly commit to capacity for workloads whose value you cannot evidence. Gate for the wedge, Warrant for the moat, Desk for the endgame.

And the strategic reason to hold all three: they fail under different conditions. The Gate dies if providers absorb routing. The Desk dies if pricing stays simple. The Warrant is the only one that survives both, and it is the one the 1887-to-audit-to-regulator sequence says gets more valuable with time, not less [12].

Founder–product fit, assessed honestly

This is the part where I am the subject, so I am going to be harder on it than a pitch deck would be.

The fit is unusually specific. My portfolio is not a collection of AI apps; it is a collection of machines that refuse to certify things. An evaluation operating system whose job is independent scoring. A build-and-grade engine whose report cards cap a score by the strength of the evidence behind it, so a claim backed by prose cannot score like a claim backed by a passing gate. A security repo whose headline number is honestly 0 out of 20 because nothing has an integrated receipt yet. A housing engine that reports 0 deployable routes out of 408 and names every blocker. An HR engine where an AI-drafted rule stays at "claimed" until a qualified human signs it into an evidence ledger. A financial system where the money is human-gated and forecasts are Brier-scored.

That is a decade of building the same organ over and over: a thing that will not let you claim what you cannot evidence. The market's number-one unsolved problem, in the trade body's own words, is that nobody can evidence AI value. I did not pick that alignment; I noticed it.

The specific transferable asset is the ablation harness I already run: prompt and model variants measured with per-turn cost, latency and cache behaviour, behind a gate that requires every cell to clear a capability bar with zero unmeasured cases. That is ① as a working artefact rather than a slide. The evidence-tiered report card is ② in miniature. A forecast ledger with calibration scoring is the instrument ③ needs.

Now the gaps, which are not small.

Breadth is the risk, not the strength. Seventy-nine projects is a portfolio. A unicorn is one thing done to exhaustion for six years. Everything in my working style that makes me good at spotting this opportunity — many parallel loops, wide reading, fast reframing — is the opposite of the thing that captures it. I would not fund a founder with my project list without asking that question, so I am asking it first.

Wrong room. I have depth in engineering and in enterprise delivery; ② sells to a finance organisation, and I have no standing in the FinOps practitioner community. That community is unusually tight, has an annual gathering, publishes its own survey, and grants credibility to people who have carried the pager. The realistic first move is not a build. It is a design partner who is a FinOps or finance-transformation practitioner, with equity, before a line of product code.

Conflict surface. Anyone currently employed in enterprise consulting or financial services who starts a company in enterprise AI cost governance has an IP and conflict-of-interest question to settle first, in writing, with their employer. Not a footnote. The first gate.

Cold start. ① needs a graded evaluation suite, and the customers who most need cost discipline are frequently the ones who cannot articulate what good output looks like. That is a services problem hiding inside a product thesis, and it is how a lot of promising eval companies quietly became consultancies.

Patterns and anti-patterns

Patterns. Sell the counter-hypothesis, not the number — the finding that survived an attempt to kill it is worth more than three findings that were merely computed. Show the rejects: a report that names what it refused to certify is the single fastest way to earn a finance audience. Join the clockwork to the meter, because Wedgwood's fixed cost is where the money still mostly sits. Price on capacity where you can. And put your own forecasts on the record with a score, because a desk that grades itself is the only kind anyone should trust.

Anti-patterns. Building a fourth spend dashboard in a market where the schema is a CNCF standard and the leading collector belongs to a $15B database company. Promising the bill will fall, when Jevons has been reliably making that promise a lie since 1865. Inferring prompt verbosity from prompt length, when the input a model is billed for includes system instructions, injected context, retrieval payloads and history. Filtering naively for waste, which produces confidently wrong recommendations because expensive models sometimes earn their cost. And quoting the 95% statistic in an article about measurement without checking it — which is the one I nearly did.

The mechanism, and the first principle

The mechanism. A cost number becomes decision-grade only when it is joined to an outcome and survives a stated challenge. Absent the join it is arithmetic; absent the challenge it is advocacy. Everything valuable in this category is machinery for producing the join and administering the challenge.

The first principle, in one sentence. Measurement is not the scarce good — warranted measurement is, because a number that cannot fail cannot inform a decision.

The receipts in this market are getting very good. The report card has not been built. Three hundred years of accounting history say the report card is what turns into an institution, and the regulator has already said out loud which order it happens in.

Related

References

  1. CloudZero. What is AI cost observability? A guide to tracking LLM and AI spend. Three-layer framing; Gartner forecast of $2.52 trillion worldwide AI spending in 2026 (+44% year over year) and $401 billion in AI infrastructure; 98% / 63% / 31% AI-spend management figures. cloudzero.com/blog/ai-cost-observability/
  2. FinOps Foundation. State of FinOps 2026. "Is your AI providing value? No one can answer that question yet."; FinOps for AI as top forward-looking priority; AI cost management as the #1 skillset to develop. data.finops.org · see also Linux Foundation press release
  3. Bamboo Innovator (2013). How a Potter Took Accounting Into the Industrial Age. Wedgwood's letter to Thomas Bentley, quoted verbatim; the fixed/variable discovery and his survival of the 1772 downturn. bambooinnovator.com/2013/04/19/…
  4. Jevons, W. S. (1865). The Coal Question. Described, not quoted. Modern formalisation of the mechanism for AI inference: The Economics of Digital Intelligence Capital: Endogenous Depreciation and the Structural Jevons Paradox, arXiv. arxiv.org/pdf/2601.12339
  5. Gridium. A History Of Demand Charges. Arthur Wright's demand meter at Brighton; Chicago Edison's 1897 demand tariffs; load factor as ratio of average to peak demand. gridium.com/history-of-demand-charges/
  6. Hasson, Y. Stop Counting Tokens. Nobody Counts Megabytes. Both quotations verified verbatim against the source. yossihasson.substack.com/p/stop-counting-tokens-nobody-counts
  7. Cloud Native Computing Foundation (21 May 2026). CNCF Announces OpenTelemetry's Graduation, Solidifying Status as the De Facto Observability Standard. GenAI semantic conventions; emission by VS Code Copilot, OpenAI Codex and Claude Code. cncf.io/announcements/2026/05/21/… · opentelemetry.io/blog/2026/otel-graduates/
  8. ClickHouse (16 January 2026). ClickHouse raises $400M Series D … acquires Langfuse. $15B valuation; Langfuse adoption figures. clickhouse.com/blog/… · acquisition post · InfoWorld
  9. FinOps Foundation. FOCUS Specification. Version 1.4 ratified 4 June 2026; Contract Commitment dataset. focus.finops.org/focus-specification/ · summary: amnic.com/blogs/…
  10. Challapally, A., Pease, C., Raskar, R., & Chari, P. (2025). The GenAI Divide: State of AI in Business 2025 (MIT Project NANDA). Mainstream coverage: Fortune, via Yahoo Finance. Corroborating ROI figures (IBM 25%; S&P Global 42% abandonment) as compiled in Terminal X, AI ROI in 2026.
  11. Futuriom. Why We Don't Believe MIT NANDA's Weird AI Study. futuriom.com/articles/news/… · funnel reconstruction (60% investigated / 20% piloted / 5% implemented) and the inflated sample sizes: NewMR, Myth Number 2: MIT Showed That 95% of AI Pilots Fail. patreon.com/NewMR/posts/…
  12. Botic, G. R. (18 June 2025). Stop, Look, and Listen: The Legacy of American Railroads and the Dawn of AI. Speech to the Association of American Railroads, Accounting Officers and Internal Audit Division. PCAOB. pcaobus.org/news-events/speeches/…
  13. Abbruscato, A. FinOps X 2026 Recap: AI Tokenomics Explained. Mavvrik. "The metric to chase is value per token, not cost per token." The sub-20% unit-economics figure appears in secondary community analysis and could not be confirmed in the primary survey data; treated as directional in the text. mavvrik.ai/blog/finops-x-2026-ai-token-economics/
  14. Braintrust. How to track LLM costs (2026): A playbook for per-user, per-feature, and per-agent-run attribution. braintrust.dev/articles/how-to-track-llm-costs-2026
  15. Zylos Research (12 April 2026). AI Agent Cost Optimization: Token Budgets, Model Routing, and Production FinOps. Source of the 30–70% routing-savings range. zylos.ai/research/…

Research packet assembled 2026-08-17 across four windows — 30 days, 30 months, 30 years, 300 years. Fifteen sources. Eight load-bearing claims were verified against their sources before drafting; one drafted quotation (Wedgwood's "like clockwork" line) failed verification at its first two sources and is cited only from the one that carries it verbatim, and one statistic (practitioners applying unit economics) could not be confirmed in the primary survey and is labelled directional in the body rather than quietly dropped. This article contains no material from any employer or client. — Paul Jialiang Wu · agentic-portfolio-lovat.vercel.app