Learning engineering · four research windows · one live case study
Hard ideas bounce. A 300-year recipe, run by agents, makes them stick.
1-minute takeaway — what you'll walk away with
Four research windows — 30 days, 30 months, 30 years, 300 years — agree: AI study tools optimize consumption, but durable learning lives on the generation side (spacing 1885, retrieval 1917, teach-back, drawing). The graveyard (learning styles, speed reading) shows popularity never predicted survival. Point AI at generation, demand receipts, keep the last click human.
Last week I fed Rich Sutton's Bitter Lesson — the essay arguing that general methods riding compute beat hand-crafted human knowledge — to a stack of AI agents. The stack learned it by doing exactly what the essay predicts: not by my hand-crafting a summary, but by running general, verifiable machinery. The recursion is funny. What the machinery actually did is not new at all. It is a study method with receipts going back three hundred years — and almost every AI learning tool shipping today skips it.
The consumption trap (what the last 30 days show)
Look at any current tool roundup and the pattern is identical: upload the paper, receive a summary, a podcast, a video overview, flashcards. NotebookLM's 2023→2026 evolution — Deep Research, audio overviews in 80 languages, cinematic video, auto-generated quizzes — is genuinely impressive engineering, and students now treat it "less like a chatbot and more like a personalized research environment." But notice the direction of every arrow: content flows toward you, pre-chewed. That is consumption. And the uncomfortable finding from a century of learning science is that consumption is precisely where durable learning doesn't happen. Re-reading — the manual version of receiving a summary — is the strategy that loses in study after study.
What survived 300 years (and what died loudly)
1 · Spaced repetition (measured 1885, age 141). Ebbinghaus, from his own syllable experiments: "with any considerable number of repetitions a suitable distribution of them over a space of time is decidedly more advantageous than the massing of them at a single time" (Memory, 1885). Replicated in 839 assessments across 317 experiments (Cepeda et al., 2006) and now standard in health-professions education. Survived the print→software substrate change (SuperMemo 1985, Anki 2006).
2 · Retrieval practice (formalized 1917, age 109). Gates' recitation studies had children look up from the page and recite; a century later Roediger & Karpicke (2006) measured the same effect in adults: a week out, recallers retained 61% vs 40% for re-readers. The oldest voice and the modern one agree across 89 years.
3 · Learning by teaching (aphorism 1801, evidence base modern). Joubert's "to teach is to learn twice" became the measurable protégé effect: in Fiorella & Mayer's review of generative strategies, explaining to someone else ranks among the strongest interventions tested.
4 · Drawing it yourself (exemplar 1810s, age ~210). Faraday — bookbinder's apprentice, no formal mathematics — bound his lecture notes "with detailed illustrations, thoroughly indexed", and historians treat the notebooks as the active organization of creative science, not mere record-keeping. The modern measurement: learner-generated drawing beat controls in 26 of 28 studies, median d = 0.40.
What the last 30 years measured, and the last 30 months changed
The last three decades turned the survivors into numbers — effect sizes, meta-analyses, Fiorella & Mayer's eight generative strategies (summarize, map, draw, imagine, self-test, self-explain, teach, enact) — and one genre shift: 3Blue1Brown (2015, now 7.85M subscribers) proved that animating the mechanism of an idea could carry graduate mathematics to millions, then open-sourced the manim engine so the craft itself became executable and shareable.
Then the last 30 months moved the delivery cost toward zero. In a Harvard RCT, students using a pedagogy-tuned AI tutor learned more than twice as much in less time than the same course's active-learning classroom (Kestin et al., 194 students) — and the authors credit the tuning to pedagogical best practices, not the model. In Edo State, Nigeria, a World Bank RCT of six weeks of guardrailed GPT-4 tutoring produced gains equivalent to 1.5–2 years of business-as-usual schooling, placing it among the most cost-effective education interventions ever rigorously tested — again with prompts designed to promote reasoning, not answers. The pattern in both: AI amplifies whichever side of the consumption/generation line you point it at. Point it at generation.
One intent, four organs, every claim gated
So here is the method I now run, as one sentence to a router. My allin-anything super-repo indexes a family of agent repos, each with a machine-verified status. The intent was:
A deterministic router (its verdicts are pytest cases, not vibes) named three organs, and each one's own test suite ran green before it was trusted: master-anything — the teach-back session structure (15 tests + smoke) · animate-anything — executable animation craft, a 10-check style linter scoring 100/100 · penecho — a pen→digital canvas, 200 tests at a pinned commit. The chain then produced five animated sessions, and this figure — drawn as real strokes, exported by the canvas's own renderer:
the two curves of the Bitter Lesson — human priors plateau, compute keeps going — as real ink
Map what the pipeline did onto the survivors: it drew (Faraday's move, d = 0.40), it structured the content as teach-back sessions (Joubert's move, the protégé effect), it animated the mechanism rather than the conclusion (Sanderson's move), and the artifact ends with the source's own quiz-able claims (Gates' move). The one thing no organ did: publish. That was my click — the human gate every irreversible step deserves.
The four outcomes, honestly mapped
You made contact with the idea through your hand — strokes on a canvas, not scroll gestures. Evidence line: generative drawing, 26/28 studies positive; enactment in Fiorella & Mayer's eight.
Each session forces the mechanism into your own words and pictures before moving on. Evidence line: retrieval 61%-vs-40%; teach-back / protégé effect among the strongest generative strategies.
The router itself is a connection engine: the same intent touched a math-animation repo, a pen canvas, and a learning loop — ideas cross-referenced across your whole tool family. Evidence line: Mayer's multimedia principle — words + corresponding pictures beat words alone.
Every organ carried an exit code. You believe the artifact because you watched its gates pass — conviction is understanding plus receipts. Evidence line: the anti-portfolio — fields that skip verification (learning styles, speed reading) die on contact with it.
Run it yourself (with or without my repos)
- Pick one dense source — a paper, a briefing, a chapter that has bounced off you before.
- Say one intent that demands generation: "learn X deeply, with an animation per key idea, and a hand sketch of its central figure." The verbs matter more than the tool.
- Route to organs, not to a summarizer — a session-structurer (teach-back), a motion tool (mechanism, not decoration), and something your hand touches (stylus, pen, whiteboard).
- Demand receipts — if an AI tool claims it helps you learn, ask what its equivalent of an exit code is. No gate, no trust: that rule is 300 years of graveyard talking.
- Keep the last step human — publishing, deciding, acting. The Nigeria RCT worked because of guardrails, not despite them.
- Space the return trip — revisit the artifact in a week and recite before you re-read. Ebbinghaus has been right since 1885.
The recursive punchline
The Bitter Lesson says general methods that leverage computation beat hand-crafted human knowledge. The study method that survived 300 years is exactly that kind of general method — draw, teach, test, space, regardless of subject — and agents have now made its cost of execution collapse, the way GPUs collapsed the cost of search and learning. Hand-crafted summaries are the "human priors" of studying: early wins, hard ceiling. The generative loop is the compute curve. You know how that graph ends — you watched me draw it.
🗞 last30days: NotebookLM 2026 feature set and student usage patterns (Glasp, DigitalOcean) · 📅 last30months: Harvard AI-tutor RCT, World Bank Nigeria RCT · 🌳 last30years: Roediger & Karpicke 2006, Fiorella & Mayer, 3Blue1Brown/manim · 🏺 last300years: Ebbinghaus 1885, Gates 1917, Faraday's notebooks (Royal Institution), plus the graveyard: Thorndike & Woodworth 1901, Pashler et al. 2008. Honest limits: the AI-tutor evidence is under three years old and both RCTs credit their pedagogical guardrails — treat "AI makes learning better" without a design description the way you'd treat a learning-styles pitch. Case-study receipts (test counts, exit codes, the ink export) are in the allin-anything repo's walkthroughs and CI history.
Paul Jialiang Wu · agentic-portfolio-lovat.vercel.app · case study: The Bitter Lesson, learned all-in · github.com/wjlgatech/allin-anything