AI-Native Series Β· Learning Engineering
I fed the Bitter Lesson to my agents. It ate my study method first.
1-minute takeaway β what you'll walk away with
Sutton's Bitter Lesson says general methods riding compute beat hand-crafted human knowledge β and your study habits are on that same graph: highlighting and rereading are the hand-crafted priors (fast start, hard ceiling); drawing, teaching, testing and spacing are the general methods that keep climbing. Agents just made the good loop nearly free. One mental model, ten verified references, real ink, and the one step machines never take: clicking Publish.

In 2019 I read Rich Sutton's The Bitter Lesson and did what educated people do with important essays: I nodded, I highlighted, I told a colleague it was "so true." Quick test, seven years later β could I recite three sentences of it? I could not. My highlighter could not either, and it was there the whole time.
The essay, in one breath
Sutton's opening line, exactly as written[1]:
Chess fell to search, not grandmaster intuition. Go fell to self-play, not proverbs. Speech fell to statistics, not phonemes; vision to learned features, not hand-drawn edges β four overtakes Sutton walks through himself[1]. Each time, researchers who had lovingly encoded human insight watched a compute-hungry general method blow past them the moment hardware caught up. The animation replays the overtake, per field β grey is the hand-crafted approach; orange arrives late and rude:
heuristics β search
patterns β self-play
phonemes β learning
features β CNNs
seventy years, four fields, one plot twist β repeated until bitter
The joke I took seven years to get
Here is the part I'd like back-dated to 2019. The essay warns, verbatim[1]:
Now look at how I was studying it. Highlighting β my hand-picked guesses about which sentences matter. Summarizing β my hand-crafted compression. Nodding β my favorite zero-compute inference. My entire study method was hand-crafted human knowledge about my own mind. I was disobeying the essay as a technique for absorbing the essay. The Bitter Lesson has a clause for people like me; I'd highlighted it.
The one mental model of this article: your study methods are on the two-curves graph too. Intuitive methods β highlight, reread, summarize β rise fast and plateau at the ceiling of your own introspection. Generative methods β draw it, teach it, test it, space it β start slower and keep climbing, because they convert time (and now compute) into memory the way search and learning convert compute into capability. Same graph. You are on it.
What a century of evidence says (this part is not vibes)
Cognitive science has run this exact tournament since before computers, and the general methods keep winning there too. Rereading loses to retrieval: a week after study, recallers retained 61% versus 40% for re-readers (Roediger & Karpicke, 2006[2]) β anticipated by Gates' classroom recitation experiments in 1917[3]. Drawing beats staring: learner-generated drawing beat controls in 26 of 28 studies, median effect d = 0.40[4]. Teaching beats consuming: in Fiorella & Mayer's review of generative strategies, explaining to someone else ranks among the strongest interventions measured[5]. Spacing beats cramming, verbatim since 1885[6]:
And here is the insider's twist, and it stings: psychology has known all of this roughly forever, and education mostly ignored it. Frank Dempster's 1988 American Psychologist paper is literally titled "The Spacing Effect: A Case Study in the Failure to Apply the Results of Psychological Research"[7] β the field's best-replicated finding, filed under unused. We have had the correct study algorithm since 1885. We kept rereading anyway, because rereading feels like learning. The fluency illusion is the phlogiston of studying, and your highlighter is a human prior with excellent marketing.
What changed in the last 30 months is the price of running the correct algorithm. A Harvard RCT: students with a pedagogy-tuned AI tutor learned more than twice as much in less time than the same course's active-learning classroom[8]. A World Bank RCT in Nigeria: six weeks of guardrailed GPT-4 tutoring produced gains equivalent to 1.5β2 years of business-as-usual schooling[9]. Both teams credit the pedagogy, not the model. Compute didn't replace the generative loop. It made it cheap.
So I ran the experiment on the essay itself
One sentence to my agent router: "learn The Bitter Lesson deeply, with an animation for each key idea, and sketch its two curves by hand." The router β a deterministic program whose verdicts are pytest cases, inside my allin-anything super-repo β named three organs, and each one's own test suite ran green before it was trusted: a teach-back learning loop (15 tests + smoke), an executable animation-craft linter (100/100), and penecho, a penβdigital canvas (200 tests at a pinned commit). Then the pipeline did what the science orders: it drew, it animated, it structured everything for teach-back. This figure is real ink β my strokes on penecho's canvas, exported by its own renderer:
the same figure as motion β the grey curve saturates; the orange one is still accelerating when the frame ends
Two more of the five animations my agents authored, because mechanisms deserve motion. Sutton's two methods that scale[1] β search as a tree that widens with budget, learning as a ball that finds the bottom without being told where the bottom is:
search widens with budget (blue) Β· learning descends the loss surface (orange)
And the anti-pattern, animated β why priors become a ceiling. One learner grows inside the box of what we managed to encode and hits the lid; one grows against compute and doesn't:
the box IS the plateau β for AI systems and for study methods alike
The chain executed to the end and then did the one thing I insist machines never do: it stopped. Publishing this page was a human click β mine. My agents drew the graph of their own superiority, animated it in five scenes, and then waited politely for permission, because the human gate is welded into the chain definition and a test fails if anyone removes it. The machines are winning, but they are winning with manners. I have never felt so redundant and so essential in the same afternoon.
Pattern: bet on methods that get better when the budget doubles β in AI systems (Sutton's search and learning) and in your own studying (generation and spacing).
Anti-pattern: encoding your introspection as the pipeline β visual features in 2005, grammar rules in 1985, your highlighter last Tuesday. Early wins, hard ceiling.
Mechanism: generative acts (drawing, teaching, self-testing) force retrieval and reconstruction, which is what writes durable memory β exactly as search and learning convert raw compute into capability without human hand-holding. Agents make the generative loop nearly free to run; they do not replace it.
Run it yourself this week (the SMART part)
Specific: pick ONE dense text that has bounced off you before. Actionable: demand generation in your first sentence β "learn X deeply, with an animation per key idea, and a hand sketch of its central figure" β and route to tools that make you produce, not consume; if a tool claims to help you learn, ask what its equivalent of a test suite is. Measurable: seven days from now, recite before you reread β Gates' move, 1917[3]; if you can't reconstruct the central figure from memory, the method failed, not you. Timeboxed: the whole loop fits in one afternoon; mine did, and I kept the receipts[10]. Relevant: you are reading an essay about compute beating priors, in the decade that made compute nearly free. The graph is not waiting for you.
References (each link verified at publication)
- Sutton, R. (2019). The Bitter Lesson. All quotes verbatim from the essay text (fetched and matched character-for-character at publication).
- Roediger, H. L. & Karpicke, J. D. (2006). Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention. Psychological Science 17(3) β 61% vs 40% retention at one week.
- Gates, A. I. (1917). Recitation as a Factor in Memorizing. Archives of Psychology, β40.
- Schwamborn et al. / Fiorella & Mayer (2016). The generative drawing principle in multimedia learning β drawing beat controls in 26 of 28 studies, median d = 0.40.
- Fiorella, L. & Mayer, R. E. (2015). Learning as a Generative Activity. Cambridge University Press β the eight generative strategies, teaching among the strongest.
- Ebbinghaus, H. (1885/1913). Memory: A Contribution to Experimental Psychology, ch. 8 β quote verbatim from the Ruger & Bussenius translation.
- Dempster, F. N. (1988). The Spacing Effect: A Case Study in the Failure to Apply the Results of Psychological Research. American Psychologist 43(8).
- Harvard Gazette (2024), on Kestin et al.'s RCT. Professor tailored AI tutor to physics course. Engagement doubled. β 194 students, more than 2Γ learning in less time.
- World Bank (2025). From chalkboards to chatbots in Nigeria β six-week guardrailed GPT-4 RCT; gains equivalent to 1.5β2 school years.
- The run's receipts β test counts, exit codes, pinned commits, the AutoRunner journal β in github.com/wjlgatech/allin-anything (walkthrough: article-to-animated-understanding); live demo at allin-anything-demo.vercel.app.
Paul Jialiang Wu Β· agentic-portfolio-lovat.vercel.app Β· the raw six-session artifact version of this page ships inside the allin-anything repo; this essay is what it taught me