Paul Jialiang Wu agentic-portfolio 中文 Español 한국어 日本語✉️ Free list

Learning engineering · four research windows · one live case study

Hard ideas bounce. A 300-year recipe, run by agents, makes them stick.

By Paul Jialiang Wu · agentic-portfolio-lovat.vercel.app · 2026-08-01 · case study: The Bitter Lesson, learned all-in

1-minute takeaway — what you'll walk away with

Four research windows — 30 days, 30 months, 30 years, 300 years — agree: AI study tools optimize consumption, but durable learning lives on the generation side (spacing 1885, retrieval 1917, teach-back, drawing). The graveyard (learning styles, speed reading) shows popularity never predicted survival. Point AI at generation, demand receipts, keep the last click human.

Last week I fed Rich Sutton's Bitter Lesson — the essay arguing that general methods riding compute beat hand-crafted human knowledge — to a stack of AI agents. The stack learned it by doing exactly what the essay predicts: not by my hand-crafting a summary, but by running general, verifiable machinery. The recursion is funny. What the machinery actually did is not new at all. It is a study method with receipts going back three hundred years — and almost every AI learning tool shipping today skips it.

The consumption trap (what the last 30 days show)

Look at any current tool roundup and the pattern is identical: upload the paper, receive a summary, a podcast, a video overview, flashcards. NotebookLM's 2023→2026 evolution — Deep Research, audio overviews in 80 languages, cinematic video, auto-generated quizzes — is genuinely impressive engineering, and students now treat it "less like a chatbot and more like a personalized research environment." But notice the direction of every arrow: content flows toward you, pre-chewed. That is consumption. And the uncomfortable finding from a century of learning science is that consumption is precisely where durable learning doesn't happen. Re-reading — the manual version of receiving a summary — is the strategy that loses in study after study.

What survived 300 years (and what died loudly)

🏺 last300years · window 1726–2026 · 4 survivors ranked · anti-portfolio attached

1 · Spaced repetition (measured 1885, age 141). Ebbinghaus, from his own syllable experiments: "with any considerable number of repetitions a suitable distribution of them over a space of time is decidedly more advantageous than the massing of them at a single time" (Memory, 1885). Replicated in 839 assessments across 317 experiments (Cepeda et al., 2006) and now standard in health-professions education. Survived the print→software substrate change (SuperMemo 1985, Anki 2006).

2 · Retrieval practice (formalized 1917, age 109). Gates' recitation studies had children look up from the page and recite; a century later Roediger & Karpicke (2006) measured the same effect in adults: a week out, recallers retained 61% vs 40% for re-readers. The oldest voice and the modern one agree across 89 years.

3 · Learning by teaching (aphorism 1801, evidence base modern). Joubert's "to teach is to learn twice" became the measurable protégé effect: in Fiorella & Mayer's review of generative strategies, explaining to someone else ranks among the strongest interventions tested.

4 · Drawing it yourself (exemplar 1810s, age ~210). Faraday — bookbinder's apprentice, no formal mathematics — bound his lecture notes "with detailed illustrations, thoroughly indexed", and historians treat the notebooks as the active organization of creative science, not mere record-keeping. The modern measurement: learner-generated drawing beat controls in 26 of 28 studies, median d = 0.40.

⚰️ The era-matched graveyard — methods that dominated louder than any survivor and died under testing: formal discipline ("Latin trains the mind"), the 19th-century consensus, killed by Thorndike & Woodworth's 1901 transfer experiments; learning styles, a billion-dollar industry when Pashler, McDaniel, Rohrer & Bjork (2008) found no adequate evidence for matching instruction to style; speed reading and sleep learning, both refuted by basic psychophysics. The lesson for 2026: popularity has never predicted survival in this field. Ask any new AI study tool the graveyard question.

What the last 30 years measured, and the last 30 months changed

The last three decades turned the survivors into numbers — effect sizes, meta-analyses, Fiorella & Mayer's eight generative strategies (summarize, map, draw, imagine, self-test, self-explain, teach, enact) — and one genre shift: 3Blue1Brown (2015, now 7.85M subscribers) proved that animating the mechanism of an idea could carry graduate mathematics to millions, then open-sourced the manim engine so the craft itself became executable and shareable.

Then the last 30 months moved the delivery cost toward zero. In a Harvard RCT, students using a pedagogy-tuned AI tutor learned more than twice as much in less time than the same course's active-learning classroom (Kestin et al., 194 students) — and the authors credit the tuning to pedagogical best practices, not the model. In Edo State, Nigeria, a World Bank RCT of six weeks of guardrailed GPT-4 tutoring produced gains equivalent to 1.5–2 years of business-as-usual schooling, placing it among the most cost-effective education interventions ever rigorously tested — again with prompts designed to promote reasoning, not answers. The pattern in both: AI amplifies whichever side of the consumption/generation line you point it at. Point it at generation.

One intent, four organs, every claim gated

So here is the method I now run, as one sentence to a router. My allin-anything super-repo indexes a family of agent repos, each with a machine-verified status. The intent was:

"learn The Bitter Lesson deeply, with an animation for each key idea, and sketch its two curves by hand."

A deterministic router (its verdicts are pytest cases, not vibes) named three organs, and each one's own test suite ran green before it was trusted: master-anything — the teach-back session structure (15 tests + smoke) · animate-anything — executable animation craft, a 10-check style linter scoring 100/100 · penecho — a pen→digital canvas, 200 tests at a pinned commit. The chain then produced five animated sessions, and this figure — drawn as real strokes, exported by the canvas's own renderer:

Hand-drawn on penecho: the human-knowledge curve plateaus while the compute curve crosses it and keeps rising

the two curves of the Bitter Lesson — human priors plateau, compute keeps going — as real ink

Map what the pipeline did onto the survivors: it drew (Faraday's move, d = 0.40), it structured the content as teach-back sessions (Joubert's move, the protégé effect), it animated the mechanism rather than the conclusion (Sanderson's move), and the artifact ends with the source's own quiz-able claims (Gates' move). The one thing no organ did: publish. That was my click — the human gate every irreversible step deserves.

The four outcomes, honestly mapped

Experience
You made contact with the idea through your hand — strokes on a canvas, not scroll gestures. Evidence line: generative drawing, 26/28 studies positive; enactment in Fiorella & Mayer's eight.
Understanding
Each session forces the mechanism into your own words and pictures before moving on. Evidence line: retrieval 61%-vs-40%; teach-back / protégé effect among the strongest generative strategies.
Connections
The router itself is a connection engine: the same intent touched a math-animation repo, a pen canvas, and a learning loop — ideas cross-referenced across your whole tool family. Evidence line: Mayer's multimedia principle — words + corresponding pictures beat words alone.
Conviction
Every organ carried an exit code. You believe the artifact because you watched its gates pass — conviction is understanding plus receipts. Evidence line: the anti-portfolio — fields that skip verification (learning styles, speed reading) die on contact with it.

Run it yourself (with or without my repos)

  1. Pick one dense source — a paper, a briefing, a chapter that has bounced off you before.
  2. Say one intent that demands generation: "learn X deeply, with an animation per key idea, and a hand sketch of its central figure." The verbs matter more than the tool.
  3. Route to organs, not to a summarizer — a session-structurer (teach-back), a motion tool (mechanism, not decoration), and something your hand touches (stylus, pen, whiteboard).
  4. Demand receipts — if an AI tool claims it helps you learn, ask what its equivalent of an exit code is. No gate, no trust: that rule is 300 years of graveyard talking.
  5. Keep the last step human — publishing, deciding, acting. The Nigeria RCT worked because of guardrails, not despite them.
  6. Space the return trip — revisit the artifact in a week and recite before you re-read. Ebbinghaus has been right since 1885.

The recursive punchline

The Bitter Lesson says general methods that leverage computation beat hand-crafted human knowledge. The study method that survived 300 years is exactly that kind of general method — draw, teach, test, space, regardless of subject — and agents have now made its cost of execution collapse, the way GPUs collapsed the cost of search and learning. Hand-crafted summaries are the "human priors" of studying: early wins, hard ceiling. The generative loop is the compute curve. You know how that graph ends — you watched me draw it.

🔬 provenance · four windows, one contract

🗞 last30days: NotebookLM 2026 feature set and student usage patterns (Glasp, DigitalOcean) · 📅 last30months: Harvard AI-tutor RCT, World Bank Nigeria RCT · 🌳 last30years: Roediger & Karpicke 2006, Fiorella & Mayer, 3Blue1Brown/manim · 🏺 last300years: Ebbinghaus 1885, Gates 1917, Faraday's notebooks (Royal Institution), plus the graveyard: Thorndike & Woodworth 1901, Pashler et al. 2008. Honest limits: the AI-tutor evidence is under three years old and both RCTs credit their pedagogical guardrails — treat "AI makes learning better" without a design description the way you'd treat a learning-styles pitch. Case-study receipts (test counts, exit codes, the ink export) are in the allin-anything repo's walkthroughs and CI history.

Paul Jialiang Wu · agentic-portfolio-lovat.vercel.app · case study: The Bitter Lesson, learned all-in · github.com/wjlgatech/allin-anything