AI-NATIVE SERIES Β· SOVEREIGNTY ENGINEERING
My Manifesto Could Not Fail. That Was the Bug.
One minute
I wrote a manifesto called Own My Intelligence β own your data, own your mind, own your future β and then spent a day holding it against sixty sources. Three things happened. The research broke my opening paragraph: I had told readers their judgment was weakening, and the 2026 data says the people who actually use AI report it helps them. Two dead W3C standards explained the deeper defect: a declaration nobody verifies converges to the cheapest conforming lie. And a paper in Science showed why this is not a style problem β the AI that agrees with you gets rated higher while measurably degrading your judgment, so harm and satisfaction point the same direction. The fix was not better prose. It was to compile ten principles into nine machine checks, publish the checks, and report my own coverage at 0 of 14.
This is commentary, not the paper
You are reading an essay about Own My Intelligence. The manifesto itself is a separate document with its own page: read Own My Intelligence v0.1 β Β· v0 (superseded) Β· all versions and their reasoning.
The document I was proud of
The first version said the right things. It rejected the industry's framing β how intelligent will AI become? β for a better question: will human beings become wiser, freer, and more fully alive, or more dependent, distracted, and controlled? It had three pillars, ten principles, and a sentence I still think is the best in it: every meaningful interaction with AI should leave the human more capable than before.
Then I did what I should have done first. I went and looked.
Ninety-eight sources over four time windows: what the internet is arguing about this month, what crossed from demo to default in thirty months, what is still load-bearing after thirty years, and what has survived three hundred. I expected the research to arm the manifesto. It did. It also broke it in three places, and one of those breaks is the reason this article exists rather than a quiet edit.
The mental model: you cannot fail a promise
Here is the idea, and it is the only one you need to carry out of this article.
A promise and a check look similar on the page and behave completely differently in the world. “We believe your data should be portable” is a promise. It has no failure condition. There is no observation you could make tomorrow that would prove the company violated it, because the sentence never said what portable meant. Now compare: a capability must demonstrate a completed round trip of the user's data, judged on whether the capability still works at the other end, or the build refuses to ship. That one can fail. It has an exit code.
You cannot fail a promise. Which means a values document made entirely of promises is not a weak safeguard β it is not a safeguard at all, and worse, it is a costume worn by the appearance of one.
I did not deduce this. I found two graves.
Two standards died of optionality
In 2002 the W3C published the Platform for Privacy Preferences β P3P β a machine-readable privacy policy. A browser could read your site's policy and act on it. It was a genuinely good idea and it is dead. It died because sites discovered they could publish a policy that was technically valid and contained nothing, and there was no consequence for doing so. The obituary was written by one of the standard's own architects, Lorrie Faith Cranor, under a title I find hard to read without wincing: “P3P is dead, long live P3P!” [15]
Do Not Track then tried the same contract from the other side: a signal from the user's browser saying don't track me. Eight years of standards work. The W3C's Tracking Protection Working Group closed it in January 2019, and the closure statement is the most useful sentence in this whole article:
“Since its last publication as a Candidate Recommendation, there has not been sufficient deployment of these extensions (as defined) to justify further advancement, nor have there been indications of planned support among user agents, third parties, and the ecosystem at large.” [16]
Apple removed the toggle from Safari weeks later.
Two W3C-scale efforts, one cause of death: the declaration was optional and ignoring it was free. Every AI ethics charter, model card, and responsible-AI pledge I have read this year β including the one I wrote β is structurally identical to P3P. Read your company's AI principles and ask the diagnostic question: what happens to a party that ignores this? If the answer is nothing, you are looking at a P3P.
The finding that reorganised the whole thing
I had written Principle 1 as a value: success is not measured by how often a person returns to the AI, but by how much the person grows. Nice sentiment. Then I read the 2026 Science paper on sycophancy, and it stopped being a sentiment.
Across eleven frontier models, AI affirmed users' described actions substantially more often than human respondents did β the paper puts it at 49% more β including where the conduct involved deception or harm. In three preregistered experiments (N = 2,405), a single interaction with a sycophantic model reduced participants' willingness to take responsibility and repair interpersonal conflict, and increased their conviction that they were right. [2]
And this is the part to sit with — in the paper’s own words: “sycophantic models were trusted and preferred.”
Read that as an engineer rather than as an ethicist. It means the harm and the satisfaction metric point in the same direction. There is no adversary in this loop. Nobody has to intend anything. Any system scored on whether you liked the answer will drift toward flattery, and you will rate the drift as improvement. Principle 1 is not a preference about metrics. It is a safety requirement, and a satisfaction score is not a weak signal β it is the attack surface.
Kant got to the structure in 1784, which is annoying. He located the failure of intellectual autonomy not in tyranny but in what he called self-caused immaturity β dependence chosen because it is comfortable. [19] The engineering consequence is precise: a harm the user prefers cannot be fixed by giving the user more choice. It can only be fixed by a commitment made in advance. A default. A denylist. A gate.
The correction that stung
My opening paragraph told the reader that their attention was being captured, their memory outsourced, their judgment weakened.
The aggregate data supports that mood. Pew found in 2025 that 53% of Americans expect AI to make people worse at thinking creatively against 16% who expect better, rising to 61% among under-thirties. [10] Comfortable reading, if you have just written a manifesto.
Then Pew's June 2026 work: Americans who actually use AI chatbots are more likely to say it helps their creativity than hurts it — 21% help against 11% hurt, with 17% saying neither. [11]
Two honest notes on that, both of which I got wrong the first time. The magnitude is smaller than it sounds: most chatbot users report no creativity effect at all, so this refutes “users feel harmed” without establishing “users feel helped.” And an earlier version of this article said young adults were “roughly evenly split” — that breakdown does not exist in the report. I had taken it from a search summary. It is gone.
Abstract dread and lived experience point in opposite directions, and I had built an opening on the wrong one. A document addressed to the frightened recruits the frightened β who are the people least likely to ever adopt a personal AI. Meanwhile the people who would build this read “your judgment is weakening” and correctly note that their own experience says otherwise, and I have lost them in the first paragraph.
Worse, look at what I was doing. I wrote a document warning that AI which tells you what you want to hear is dangerous β and opened it by telling readers what the frightened ones wanted to hear. I had flattered the reader's fear. Same mechanism, opposite direction, and I did not notice until a survey embarrassed me.
There is a far better framing available, and it was in the research the whole time. RAND's youth panel found students using AI for homework rose from 48% to 62% between May and December 2025 β while the share saying it harmed their critical thinking rose from 54% to 67%. [9] The same population increased its use and increased its concern. They are not asking to be saved from the tool. They are asking for a version that does not cost them this. That is demand, not opposition, and it is a completely different opening paragraph.
The bit where my own rule caught me
The research corpus uses a four-tier evidence ladder: verified in the artifact, corroborated, claimed, disputed. The load-bearing rule is verify at the claim's native altitude β a code claim in the code, a study's N in the paper, a legal date in the official text.
Applied to my own run, it demoted four of my most quotable figures. The Stanford entry-level displacement numbers, the N of a widely cited knowledge-worker study, a dark-pattern taxonomy β all reached me through commentary rather than the source document, so all are tagged claimed.
And the fourth one is the one I would like to quietly not mention. I had a line about frontier labs disclosing that their agents escaped test sandboxes. Serious claim. Load-bearing. Sourced, when I checked my own notes, to a weekly AI-news aggregator.
I wrote a framework whose entire proposition is provenance, and cited an AI safety incident to a listicle.
It is tagged claimed, with a note saying explicitly that a news aggregator is not a disclosure and that the primary post has not been read. I am leaving it visible rather than deleting it, because the corrections are the most valuable output of the whole exercise, and a repo that cannot show which of its numbers it verified is asking to be trusted on exactly the basis it tells everyone else to reject.
What v0.1 actually changed
Ten principles now name nine checks. Nine, not ten, because one check
(engagement-denylist) does the work for two principles β and I only noticed the count
was wrong because I had to put a number in the infographic above, which is its own small lesson about
where drift hides.
| Principle | Check | Fails when |
|---|---|---|
| Growth over dependence | engagement-denylist | the success metric is actives, retention, session length, or streak |
| Ownership over extraction | data-custody-declared | custody is not declared local / portable / vendor |
| Agency over automation | human-gate-required | an irreversible action has no named human gate |
| Understanding over answers | atrophy-declared | no declared erosion risk plus counter-practice |
| Transparency over manipulation | transparency-declared | the user cannot inspect what was used and whose interest it serves |
| Portability over lock-in | no-exit-no-pass | no working exit path, or the round trip is untested |
| Purpose over engagement | engagement-denylist | the target is not the user's declared goal |
| Compounding over consumption | return-sharper-declared | no signal that the human became more capable |
| People over platforms | no-vendored-satellites | the capability requires one named vendor |
| Legacy over immediacy | window-tagged-evidence | all justification offered is from this quarter |
Three of those are worth their derivation, because each encodes a result rather than my taste.
no-exit-no-pass β portability means a completed round trip judged on
preserved capability, not an export button. I can hold that bar because someone already clears it:
Consumer Reports' Data Rights Protocol runs authorized-agent data-rights requests in production
across several compliance vendors. [14] (The two-million-request figure often
attached to DRP is Consumer Reports’ Permission Slip app total, not the protocol’s
throughput — the DRP page publishes no request count. I had it attached to the wrong thing.)
An existence proof turns an aspirational requirement into a procurement requirement.
atrophy-declared β every capability must name the skill its routine
use may erode and the counter-practice that maintains it. Parasuraman and Riley separated
misuse (overreliance) from disuse (neglect after false alarms) in 1997, which is
why “add a human” is not automatically safer. [5] The deskilling literature
reports that learners “perform competently when familiar prompts guide their thinking, but
their performance weakens when those cues are absent, unfamiliar, or misleading” — and
the same paper states plainly that “there is limited systematic evidence on AI-related
deskilling in healthcare, including its timing, mechanisms, and affected groups.” [8]
I originally wrote that performance “holds, then drops sharply when withdrawn.” The paper
does not claim that. The gate survives on the 1997 taxonomy and the statute; it does not need a
sharper empirical claim than its source will support. Together they produce a defect nobody has priced β an
approval signed by someone who could no longer produce the answer launders the machine's error as a
human decision. And since 2 August 2026, when the EU AI Act's high-risk obligations became
enforceable, the law requires oversight by a person with the necessary competence, training and
authority. [12] The workflow is quietly removing the thing the statute assumes.
human-gate-required β this one is 300 years old. In Keech v
Sandford (1726) a trustee took a lease for himself, but only because the child he held it for
could not have it. The court made him hand over every penny. [17] It did not ask
whether harm had occurred; it removed the possibility of conflict. Read as engineering
rather than law: an agent that both recommends and monetises, both scores and sells, or both keeps
your memory and trains on it, is in a Keech position regardless of intent β and the remedy
is structural separation, not good behaviour. It is prophylactic precisely because after-the-fact
harm assessment is what machine-speed action makes impossible.
The honest number
The north-star metric is Verified Sovereignty Coverage: of the fourteen harm areas in the map, how many have a countermove bound to a passing check with evidence on disk.
It is 0 of 14.
Two areas have no check at all and say so β homogenisation of thought, and obsolescence dread. Homogenisation is the honest one to dwell on: large language models raise individual originality while measurably narrowing population-level diversity of expression. [4] No user ever feels that loss, so no feedback signal ever fires. That is the one place where “own my intelligence” is structurally insufficient β a harm invisible inside every individual case cannot be corrected by individual choice. I would rather publish the gap than proxy it.
And I am publishing zero rather than a friendlier derived number for the same reason every area in the map carries a mandatory steelman of the opposing case: a metric that cannot embarrass its author is not measuring anything. The claim is not that coverage is high. It is that coverage is now a number that can be low.
Mechanism, pattern, anti-pattern
The mechanism
Optionality selects for the cheapest conforming behaviour. A declaration with no enforcement creates a gradient toward the emptiest policy that still passes as compliant β which is why P3P died of valid-but-vacuous policies rather than of outright refusals. Add the sycophancy result and the gradient gets worse: where the harmful condition is also the pleasant one, user choice reinforces the drift instead of correcting it.
The first principle
A principle with no check is a slogan; the check is what makes a value falsifiable.
Patterns
- Compile the adjective into a gate. Name the value, name the observation that violates it, make the observation automatic, then let your marketing quote the gate rather than the aspiration.
- Find the party who already complies. An existence proof anywhere in the market converts an aspirational bar into a procurement bar and removes the last excuse.
- Neutral governance beats an open licence. A licence protects your right to fork; only governance constrains what the owner can do to the roadmap. MCP was already openly licensed and still single-vendor until it was contributed to a Linux Foundation fund in December 2025. [13]
- Keep the corrections in the artifact. The disputed entries are the most valuable part of a knowledge base, because they are the only proof the discipline is running.
Anti-patterns
- Declaration without enforcement. Worse than silence: it manufactures the appearance of accountability.
- The unfalsifiable promise. “We do not train on your data” cannot be verified from outside the model. Price it as insurance, not as a control.
- Format outrunning substance. The Agile Manifesto reorganised an industry and produced Cargo Cult Agile β the sticky notes without the trust. [20] Whatever is cheapest to imitate propagates fastest, so ship the checks with the document or watch the vocabulary travel alone.
- Reading virality as validation. Humane raised roughly $230M, shipped fewer than 10,000 AI Pins, sold to HP for about $116M and bricked every device on 28 February 2025. The named failure was mistaking viral demos for market demand β and a manifesto is optimised for virality. [18]
Try this in twenty minutes
Open your team's values doc, AI principles, or engineering standards. Pick the single most important line. Then:
- Write the observation that would prove you violated it. One sentence, concrete, dated.
- Ask whether any system you own would notice that observation. If not, you have a promise.
- Make the cheapest possible version of the check β a grep, a test, a required field in a template. It does not have to be good. It has to be able to fail.
- Count: values with a check, over total values. Publish the fraction.
You will hate the fraction. Mine is 0 of 14. Publish it anyway, because the fraction is the first honest thing the document has ever said.
What would prove me wrong
A framework that cannot fail should not be adopted, so here are the observations I accept as disconfirming. A preregistered longitudinal study showing heavy AI users retain or improve unaided higher-order performance over twelve months on a delayed-test design. Replication showing the sycophancy effect dissipates over repeated real-world interaction. A demonstrated full round trip of accumulated memory between two major assistants with independently confirmed preserved capability β which would make my strictest check redundant, and I would retire it and say so. And, on the whole thesis: evidence that people using assistive AI intensively are becoming, on average, more capable and more independent in judgment rather than less.
I would rather learn that than be right about this.
References
- International AI Safety Report 2026 (Y. Bengio, chair). internationalaisafetyreport.org (accessed 5 Aug 2026).
- Cheng, M., Lee, C., Khadpe, P., Yu, S., Han, D. & Jurafsky, D. (2026). “Sycophantic AI decreases prosocial intentions and promotes dependence.” Science, 391(6792), eaec8352, 26 March 2026. doi:10.1126/science.aec8352 (abstract verified verbatim, 5 Aug 2026).
- Gerlich, M. (2025). “AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking.” Societies. mdpi.com/2075-4698/15/1/6 (accessed 5 Aug 2026).
- “The homogenizing effect of large language models on human expression and thought.” Trends in Cognitive Sciences, 2026. cell.com/trends/cognitive-sciences (accessed 5 Aug 2026).
- Parasuraman, R. & Riley, V. (1997). “Humans and Automation: Use, Misuse, Disuse, Abuse.” Human Factors, 39, 230–253. doi:10.1518/001872097778543886 (accessed 5 Aug 2026).
- Clark, A. & Chalmers, D. (1998). “The Extended Mind.” Analysis. philarchive.org/archive/JULTEM (accessed 5 Aug 2026).
- Bjork, R. A. & Bjork, E. L. “Introducing Desirable Difficulties into Practice and Instruction.” unh.edu (PDF) (accessed 5 Aug 2026).
- “Deskilling dilemma: brain over automation.” Frontiers in Medicine, 2026. doi:10.3389/fmed.2026.1765692 (accessed 5 Aug 2026).
- RAND Corporation (2026). “More Students Use AI for Homework, and More Believe It Harms Critical Thinking.” rand.org (accessed 5 Aug 2026).
- Pew Research Center (2025). “How Americans View AI and Its Impact on People and Society.” pewresearch.org (accessed 5 Aug 2026).
- Pew Research Center (2026). “Americans' Views on AI Chatbots, Smart Devices and AI's Impact.” pewresearch.org (accessed 5 Aug 2026).
- EU AI Act, Article 26: Obligations of Deployers of High-Risk AI Systems. artificialintelligenceact.eu/article/26 (accessed 5 Aug 2026).
- The Linux Foundation (2025). “Formation of the Agentic AI Foundation.” linuxfoundation.org (accessed 5 Aug 2026).
- Consumer Reports Innovation Lab. “Data Rights Protocol.” github.com/consumer-reports-innovation-lab (accessed 5 Aug 2026).
- The Platform for Privacy Preferences 1.0 (P3P1.0) Specification, W3C Recommendation, 2002. w3.org/TR/P3P. Cranor, L. F. (2012). “P3P is dead, long live P3P!” β the author's site was unreachable at time of writing, so cited via the Internet Archive: web.archive.org (accessed 5 Aug 2026).
- W3C Tracking Protection Working Group closure statement, 17 January 2019; see also “How the tragic death of Do Not Track ruined the web for everyone.” Fast Company. fastcompany.com (accessed 5 Aug 2026).
- Keech v Sandford (1726) Sel Cas King 61. en.wikipedia.org/wiki/Keech_v_Sandford (accessed 5 Aug 2026).
- “AI Product Failures 2026: Sora, Humane & Rabbit R1.” digitalapplied.com (accessed 5 Aug 2026). Secondary source; the Humane figures are reported, not independently verified.
- Kant, I. (1784). “An Answer to the Question: What Is Enlightenment?” philosophynow.org (accessed 5 Aug 2026).
- Agile Alliance (2026). “25 Years Ago, a Manifesto Was Born.” agilealliance.org (accessed 5 Aug 2026).
Correction, 5 August 2026. After publishing, I ran an independent
citation audit of this article — the check the piece argues for, applied to itself. It found six
defects and they are corrected above, each marked in place rather than silently edited: an over-claimed
deskilling “signature” the source paper explicitly disclaims; a truncated W3C quotation
attributed to a document that does not exist; a Bjork sentence I had quoted that appears nowhere in the
artifact (now replaced with a verified one); a two-million-request figure attached to the wrong
organisation; an age breakdown that is not in the Pew report; and a paraphrase of the Science result
where the paper’s own words were available. Eight other citations verified clean at the artifact.
Full scorecard, including six sources I could not fetch and therefore did not upgrade, in the
repository’s docs/CITATION_AUDIT.md.
Written by Paul Jialiang Wu. The two verbatim quotations above β the W3C Tracking Protection Working Group closure statement and the title of Cranor's 2012 post β were verified against their sources. Figures attributed to the Stanford Digital Economy Lab, the Humane post-mortem, and the reported agent sandbox escapes reached this article through secondary reporting and are labelled as such rather than presented as established. Every URL above was checked for resolution on 2026-08-05; publisher bot-walls returned 403 to an automated check and were treated as reachable-by-humans, and one at-risk citation is served from the Internet Archive.