Paul Jialiang Wu

AI Lead · Founder · Investor · Recreator

I build open AI operating systems that enable people and companies to grow 12X — so my ceiling becomes your floor.

🤖 Ask my agent

Compounding everything

12X Future

Aim

12X the future, 12% the loop now

Begin from the End — Result-Oriented System Engineering (ROSE): every venture starts as a one-page end-state — the 12X result, its metric, its date — written before anything is built. Then the arithmetic takes over: (1+0.12)^22 ≈ 12.1, so a 12X future is ~22 compounding loops, each moving a named metric by ≥12% now, measured — or it wasn't a loop, it was motion.

the TRUE test ↓

T Transferable & TransformativeEnd-state-first plus the compounding identity (1.12^22 ≈ 12.1) transfers to revenue, research, and relationships alike — any 12X dream decomposes into a designed finish line and ~22 measured loops.

R Reusable & RefinableTwo reusable artifacts: the one-page end-state (result, metric, date, owner) and the loop contract (name the metric, move it ≥12%, show the measurement) — both refinable every re-plan.

U Understandable & U-loopOne breath: begin from the end at 12X, then demand 12% from every loop — because anything less never compounds there.

E Experienceable & ExperimentableWrite the end-state page for your current project today, then run one loop this week and read its measured delta. No +12% reading = motion, not progress.

For youYou never start without the one-page finish line — and you hold both numbers at once: 12X on the wall, +12% measured in every loop.

For an agentA skill that rewrites any goal as its one-page end-state (metric, date, owner), plus a workflow gate: a loop only counts when its named metric shows the +12% delta — measured, never claimed.

Build

Amazing Art × Great Science × Deep Engineering

Big dreams are earned three ways. Amazing art: elegant beauty that strikes awe — simple form, profound truth. Great science: a falsifiable hypothesis. Deep engineering: a tested, shipped system. Every bet carries all three, or it's a wish — then vision sets the heading, telemetry compounds it, and each week one open gap gets closed.

the TRUE test ↓

T Transferable & TransformativeArt-plus-hypothesis-plus-system transfers across domains: the same triple that ships a trading engine ships a biomarker pipeline — and awe travels farther than any spec sheet.

R Reusable & RefinableThree reusable artifacts: the falsifiable claim, the test suite, and the weekly gap-review — every failed hypothesis still compounds the map, every closed gap raises the floor.

U Understandable & U-loopOne sentence: earn the bet three ways — beauty so elegant it strikes awe, truth so simple it runs profound, a system so tested it ships — then let the destination pull while the sensors correct.

E Experienceable & ExperimentableWrite your biggest bet's falsifiable claim and its failing test today; close one open gap this week and check that next week starts from higher ground.

For youYou hold every big dream to three bars — elegant beauty that strikes awe, a claim that could fail, a system that ships — then let the future set the heading and close one gap a week.

For an agentA plugin that refuses to open a build until the bet carries both artifacts (the hypothesis that could fail, the test that would show it) — feeding the telemetry-fed re-solver workflow that recomputes the best next moves from actuals.

Love

Love God · Love Family · Enable Friends

The base case and the multiplier: faith and family are floors, not variables — and friends don't just get encouragement, they get enablement: shipped tools, real introductions, honest coaching.

the TRUE test ↓

T Transferable & TransformativeThe ordering — God, family, friends — transfers to every decision at every scale; it is the one constraint set that never gets re-optimized away.

R Reusable & RefinableReturned to daily, not outgrown: the base case is re-affirmed each morning and each high-stakes decision re-checks against it.

U Understandable & U-loopRecursion needs a base case: this is the value that stops the optimization from eating the optimizer.

E Experienceable & ExperimentableExperienceable in the calendar and the ledger: time with God and family is scheduled first, and each month one friend receives enablement you can point at — a shipped tool, a real introduction, honest coaching.

For youYou keep faith and family as floors, not variables — and you enable friends with artifacts, not just applause.

For an agentA guardrail hook: before any high-stakes autonomous action, surface who this serves — God, family, friends — and stop when the answer is nobody.

↑ Back to top

Built in the open (last 12 months)

Projects

anyagent

private

Build, grade, improve & accountably ship any agent app from one sentence — a tested OOP engine with a closed RefineLoop and swappable seams. (Private.)

Python12026-08-17

FM-os

public

🛠️ FM-os: the living, SLM-first map of foundation-model operations — pre-training, post-training, fine-tuning & RL. Curated repos, courses, papers & jobs, auto-refreshed weekly.

< > Codeno hosted app yet
Python32026-08-17

GodView-anything

private

The god's-eye view of what top-rated repos actually teach — verified, weekly-synced, machine-audited. Three provenance-pinned pillars: transferable knowledge, agentic tooling, community.

🔒 source and app are private
Python12026-08-17

rsi-os

public

🤖 The most comprehensive, community-driven resourcefor Automated AI Research — papers, tools, people, lbs, roadmaps. Featuring Recursive Lab, Sakana AI, DepMind & beyond.

Python32026-08-17

graph-engineering-anything

private

🕸️ the Graph Engineering Operating System — the model finds text; the graph finds reality

🔒 source and app are private
Python2026-08-17

eval-anything

public

🧪 the Evaluation Operating System — OEC (Observe → Evaluate → Control) for foundation models, agents, and business objectives

< > Codeno hosted app yet
JavaScript12026-08-17

design-anything

public

Any design intent in (text, picture) → execution-ready 3D blueprint out — construction- and 3D-print-verified. Game design · simulation · architecture (3DCP+AI) · interior · landscape.

< > Codeno hosted app yet
Python2026-08-17

eval-anything-archive

private

ARCHIVE (private, do not publish) — pre-rewrite history of eval-anything. Superseded by the public repo; retained only because it contains an unreleased sibling product's walkthrough in old commits an

🔒 source and app are private
JavaScript2026-08-17

agentic-portfolio

private

Agentic personal portfolio — Next.js + CopilotKit site whose on-page AI agent answers about my work, powered by a free-LLM survival chain (NVIDIA NIM→Groq→Gemini). Built in the open.

HTML12026-08-14

agentic-marriage

private

PRIVATE. A verified route from a marriage pain point to a named human and one committed step. Knowledge (emotional/spiritual/legal), agentic tooling, indexed communities — under fail-closed safety gat

🔒 source and app are private
Python2026-08-13

kingdom-come

public

Open-source predictive formation intelligence for seminary & ministry training teams.

Python12026-08-13

loop-engineering-anything

public

A loop orchestrator that turns any target into a self-improving, agent-native CLI: generate → judge → refactor → re-judge to convergence.

< > Codeno hosted app yet
Python32026-08-13

MorganStanley-os

private

PRIVATE - Accenture engagement operating system + FDE-os case study. Strategy and method only; no client material.

🔒 source and app are private
Python2026-08-12

agentic-HR

private

PRIVATE — deployability decision engine for HR AI. Collaboration with Lisa Tolle (Mentors US). Not legal advice.

🔒 source and app are private
Python2026-08-11

agentic-house

private

PRIVATE. Landed-cost + permit-path decision engine for factory-built housing (China sourcing, US deployment).

🔒 source and app are private
Python2026-08-10

FDE-os

public

Forward-Deployed-Engineer OS — an operating system for shipping agentic solutions into the field.

HTML12026-08-10

longevity-loop

public

🧬♻️ An AI-native, build-in-public compounding loop for aging science: falsifiable question → open data → public verifier → honest write-up → share → compound. No wet lab required.

< > Codeno hosted app yet
Python12026-08-10

loop

private

High-performance loop runtime (Rust). (Private / experimental.)

Rust2026-08-07

18 shown · swipe or use ◀ ▶

↑ Back to top

Long-form on LinkedIn

Writing

AI Economics · Strategy2026-08-17

Everyone Is Building a Better Receipt

Worldwide AI spending is on track for $2.52 trillion in 2026, and the FinOps Foundation's own 2026 survey answers the question "is your AI providing value?" with the sentence "No one can answer that question yet." In two years the field went from 31% to 98% on measuring AI spend and arrived at almost nobody on knowing what the spend bought. So almost everyone in this category is building a better meter — and the meter stopped being defensible twice this year: OpenTelemetry graduated CNCF on 21 May 2026 with GenAI conventions already emitted by Copilot, Codex and Claude Code, and ClickHouse acquired Langfuse on 16 January 2026 alongside a $400M round at a $15B valuation. The mental model is three questions per dollar: the bill (measured), the claim (asserted by whoever wants the budget), and the warrant (missing). Deep time says the warrant is what becomes an institution — Wedgwood found fixed versus variable cost in a 1769 letter ("these expences move like clockwork"), Jevons explained in 1865 why your bill rises as prices fall, and Insull was selling a two-part capacity tariff by 1897, which is exactly where AI pricing is heading. A PCAOB board member has already narrated the sequence: railroads invented modern accounting, then the audit, then the regulator. Three companies fall out of that — a Gate that only books savings which provably held quality, a Warrant ledger that refuses to certify what it cannot evidence, and a Desk that buys capacity instead of counting tokens — plus an honest founder-fit assessment in which the biggest named risk is my own project count. Also: the 95% statistic everyone quotes about unmeasurable AI value turns out to be badly measured itself.

Share:
AI for Good · Faith2026-08-13

The Cheapest Thing in This Mission Is the Truck

I costed a Cybertruck as a mobile platform to build companies, teach kids and serve a city, and expected the hard question to be financial. It was not. A mobile ministry platform is a 287-year-old design — Wesley started field preaching in 1739 and had it organised into travelling circuits by 1746 — which means it demonstrably works on foot, and the vehicle has to prove it multiplies the mission rather than enables it. The running order was settled in 1865 too: soup, soap, salvation, in that order, and reversing it gets you an audience with a soup pot as a prop. The real risk sits where nobody was looking. A 2002 randomized trial of 1,138 adolescents found that those whose mentoring relationships ended very quickly reported "decrements in several indicators of functioning" — worse off than the children who applied for a mentor and were left on the waiting list. Starting is not the neutral act; stopping is the intervention. Plus the metric trap a large federal character education study exposed (activity counts rose, nothing else moved), the head of the Office of Justice Programs and OJJDP's acting administrator warning in print that Scared Straight made kids up to 28% more likely to offend than youths who did not participate, and why every property that made it great television made it harmful practice. Fifteen sources — and an honesty ledger: three drafted quotations failed verbatim verification, then four independent review passes failed the piece and found thirty-seven more defects between them, including an author I had invented outright in my own reference list, and one court citation that was wrong three times running. All are named in Provenance rather than quietly fixed.

Share:
LinkedInAug 3, 2026

270 impressions View analytics

270 impressions View analytics

Share:
Curriculum Engineering2026-08-03

I Compiled 58 Snowflake Courses Into an Agent. It Failed Me First.

Snowflake's own org ships a strong agentic toolchain — an MCP server, cocoplus (verbatim: "an Agentic Operating System for Snowflake Coco"), Cortex Code skills — and every one of them operates on the WAREHOUSE. None operates on the CURRICULUM, so no machine can answer "what should this person learn, and did they?" I compiled learn.snowflake.com into a provenance-pinned knowledge base (58 courses, 430.25h, 12 distinct exams from 19 pages, 94 records each carrying the URL and sha256 of the page it was parsed from) and built the missing layer on top: a pathfinder that orders a route using Snowflake's own published track order where it exists, and a readiness gate where claimed completions score zero — claim all 58 courses with no badge and coverage is 0.0, NO-GO, exit 2. The useful output is the four things it refuses to know: the role-to-course mapping does not exist upstream (0 of 12 role pages publish one), Specialty exam pricing is unpublished so the parser refuses to borrow the $175/$375 sitting next to it, one track publishes no sequence, and nobody's real competence is measurable from a catalogue at all. Then it failed my own work three times: the integrity gate caught an orphan competency on its first run (8/9), a duplicate-counting bug scored a well-prepared candidate 21% until alternatives collapsed it to 55%, and one line of HTML-escaping had been silently rendering eleven shipped knowledge viewers blank. Anchored on Miller's 1990 assessment pyramid and on four honest words Snowflake prints on its own exam page: "assumed but not tested." First principle: a system that cannot fail cannot certify.

Share:
Physical AI2026-07-17

Machines Can Finally Watch Video. The Hard Part Is Teaching Them to Understand Motion.

A Vision-Language Model that captions a frame is a tourist with a camera; one that understands motion is a witness who can testify — and the gap between them is where most of the value in video AI lives. A mental model a 15-year-old can run: a VLM watching video is a brand-new student driver, and three things turn the student into a driver, none of them a bigger model. (1) Time-glue — spatiotemporal reasoning that binds the same object across frames so 30 snapshots become one scene with a start and end, not per-frame captions. (2) A strict examiner — an agentic-eval harness that grades localization + temporal + narrative consistency against reality and refuses to pass a claim it can't back with evidence (no evidence ⇒ no claim; a benchmark you can't reproduce is a vibe, not a score). (3) A practice-test factory that can't cheat — the 'AI training AI' curation flywheel where the model proposes labels on raw footage, only clips where independent passes agree survive, and a human gates promotion, governed by maker ≠ checker so the system can't launder its own hallucinations into ground truth. The punchline: everyone rents the same multimodal brain, so the edge is the loop — fine-tune → evidence-gated eval → human-gated curation → fine-tune — not the model. The model is the cheapest part; the loop is the moat. The same lesson recurs across driving video, protein folding, and aging biology. Tooling open at github.com/wjlgatech.

Share:
LinkedInJun 24, 2026

344 impressions View analytics

Activate to view larger image,

Share:

↑ Back to top

Proof, not claims — a résumé audited against real artifacts, then closed into a verified one

Resume Verification

Verify it yourself

Skeptical? Paste a résumé — mine, or any — and watch each claim audited live against real public GitHub. Honest by design: unprovable claims come back unverified, never rubber-stamped.

Your run is shown here in your browser; it doesn’t change the published proof.

58/100
Corroboration index

3/6 claims corroborated by real artifacts, 1 partial, 2 need an external source.

3 Corroborated🟡 1 Partially corroborated 2 Unverified — needs external source

top gap: Teaches a Silicon Valley AI-architect cohort.A cohort syllabus crediting Paul as instructor, or a participant testimonial.

Sample receipts (built from public repos). Ask the agent to verify a real résumé. · source: “(Sample receipts — paste your own to the agent: “verify this résumé: …”.) Paul Jialiang Wu…”

↑ Back to top

What people who've worked with me say — imported from LinkedIn, or written here (I approve each one; you can post yours to LinkedIn in a click)

Recommendations

I worked with Paul at Genentech where he was the Principal Data Scientist on my team. His technical knowledge was always astounding and his understanding and direction for the possibilities of Machine Learning using our available data gave our team direction and the confidence to explore new projects. He was a pleasure to work with and I recommend him for any team in need of a technical data/ML lead.

Senior Machine Learning Operations Engineer @ Genentech - People Insights Roche

Rajesh worked with Paul Jialiang on the same team

I have been working with Paul on a GraphML project in the past 6 months. Paul is a passionate leader with a solid technical background. He always try his best to inspire and influence others, and to encourage others to share their ideas. Besides, he can build and organize the workflow properly. There is no doubt that Paul will be a great asset to your team.

Data Scientist - People Insights at Roche

Paul Jialiang was senior to Zhao but didn’t manage Zhao directly

Paul has a great energy and is a good story teller. He leaded a focus group with high-tier people and we worked together to put our business to the next level, 10X-style.

Fondateur de CAIRN-CREA | Je fais du contenu organique & pub qui convertit pour marques DTC

Louis worked with Paul Jialiang on the same team

I am very glad to have spoken with Paul — he's a champion for self-improvement and mental health, and is also a great person to talk to!

Building Multimodal AI in Healthcare

William was Paul Jialiang’s client

I had the pleasure of attending Paul's presentation on GraphML in a conference, he is an amazing Data Scientist and a brilliant instructor of latest technologies in Data Science and Machine Learning. He explains very complex concepts, turns challenging topics look easy by breaking them into clear and concise steps. He tries to teach tough concepts by correlating them with real world scenarios which makes him a great story teller. His principle of improving 10x in 3 months is astonishing. He can guide people with any years of experience to be successful in their Data Science journey. I hope to work with him very closely and continue learning from him in the future. Paul would be a great asset for any organization!

On

Mohith was Paul Jialiang’s client

I had the pleasure of working with Paul at Galvanize, he is a brilliant and gifted Data Scientist who truly understands how to get the best out of people. Paul always made sure to make the students feel at ease, even when he was teaching challenging topics. He always made sure that his students got most of the class and he was always fun to co-teach with. The biggest strength in Paul, we all in the team admired was his ability to deal with conflicting priorities in high-pressure situations. He never lost his cool and was always a people's person. Paul would be a true asset for any company who need a game changing Data Scientist.

On

Swathi worked with Paul Jialiang on the same team

Paul is one of the best instructors in data science. He is super knowledgable and breaks down complex concepts into clear and understandable topics. He is a dedicated teacher and I learned so much from him in a short amount of time and I hope to continue learning from him in the future.

Senior Data Scientist @ Ent Credit Union | AI/ML

Paul Jialiang was senior to Bahar but didn’t manage Bahar directly

Paul was a substitute instructor for our Data Science Immersive program and he was absolutely amazing! He was easy to work with, the students loved his teaching style, and he showed a great depth of Data Science knowledge and teaching ability. I hope to collaborate with him again in the future and I would highly recommend him for any position on your team.

Strategy & Operations | Tech Partner Ecosystem Development

Kristen worked with Paul Jialiang but on different teams

It was my pleasure to be in a class that his teacher is Paul Jialiang, he has all the skills that I feel it is really important that the teacher should have, it was easy for him to make the concept clear and what I really like about him, is he make us feel good about ourself spotting the light on our hard work and what we know and learned,, I felt happy after your class thanks again Paul

On

Paul Jialiang was senior to Marwah but didn’t manage Marwah directly

9 shown · swipe or use ◀ ▶

↑ Back to top

Score any job against past experience · current skillset · future mission/values/vision — held to a golden-set accuracy

Role Fit

Score a role against me

Paste a job-posting URL (Ashby/Greenhouse/Lever, fetched live) or the JD text. The agent scores fit across past experience, current skillset, and future mission/values/vision — and tells you, honestly, where it doesn’t fit.

Why trust this

8/8 within one band · 100%

This scorer is itself held to the standard it holds JDs to: it was run over a 8-example golden set with human-assigned fit labels and agreed within one band 100% of the time (6 exact). Verify, don’t vibe.

No role scored yet — paste a posting above to see the fit breakdown.

↑ Back to top

Knowledge graphs + agentic tooling (skills, plugins, workflows, hooks, bundles) I've built — full for me, a summary card for visitors

Agentic Library

AnyAgent — build · grade · improve · ship any agent apprepo🔒 private

AnyAgent is a tested, object-oriented engine designed to build, grade, improve, and ship agent applications from natural language.

Owner-only knowledge graph + tooling. Unlock owner mode to view the full package.

Retrieved, not guessed: 3 competency clusters, 9 concept/tool nodes grounded in real sources (9 Wikipedia/GitHub definitions), and 6 skills from the ESCO taxonomy — off-domain hits filtered out. Every node carries a real definition; edges are real relationships where an open KG had them.

13 concepts · 12 links9 skills · 2 provenexplore the full graph →

Retrieved, not guessed: 6 competency clusters, 23 concept/tool nodes grounded in real sources (21 Wikipedia/GitHub definitions), and 8 skills from the ESCO taxonomy — off-domain hits filtered out. Every node carries a real definition; edges are real relationships where an open KG had them.

30 concepts · 29 links12 skills · 4 provenexplore the full graph →

Retrieved, not guessed: 5 competency clusters, 17 concept/tool nodes grounded in real sources (15 Wikipedia/GitHub definitions), and 9 skills from the ESCO taxonomy — off-domain hits filtered out. Every node carries a real definition; edges are real relationships where an open KG had them.

23 concepts · 22 links12 skills · 3 provenexplore the full graph →

Transformers scale COMPUTE (Mixture-of-Experts routes tokens to experts) but have no native primitive for looking knowledge UP. DeepSeek's Engram adds that missing primitive: it modernizes classic N-gram embeddings into a deterministically-addressed table with O(1) lookup — a second, complementary axis of sparsity (static memory) alongside MoE's conditional compute. The paper finds a U-shaped scaling law: for a fixed budget there's an interior optimum splitting capacity between compute and memory — going all-in on either side loses. Under iso-parameter and iso-FLOPs constraints, Engram-27B consistently beats MoE baselines on knowledge, reasoning, code, and math. Mechanistically, offloading static recall to Engram frees the early layers from pattern reconstruction, preserving effective depth for reasoning — and the huge embedding tables offload to host memory with minimal inference overhead. Why it matters for you: 'conditional memory as a sparsity axis' is a transferable design move — separate what a system should COMPUTE from what it should LOOK UP, and size each.

10 concepts · 11 links2 skills · 0 provenexplore the full graph →

↑ Back to top

Four growth vectors — deepen · widen · lengthen · heighten — plus who to reach. Drafted for approval.

Next Projects

Runs weekly · last scouted 6/29/2026

Four growth vectors

Deepen

more fundamental, seminal — to the roots

first-principles / foundational research

Widen

new applications, features, markets

Ansoff · Innovation Ambition Matrix

Lengthen

evolve it to robustness, scale, commodity

McKinsey Three Horizons · Wardley evolution

Heighten

generalize, abstract, compress the mechanism

abstraction laddering · compression (MDL)

deepensos, loop-engineering-anything, company-os

Enhance Self-Improving Agentic Operating Systems

Further develop and refine the existing self-improving agentic operating systems to increase their efficiency and capabilities

first step: Integrate idle detection and SMARC output-quality verification from sos into loop-engineering-anything

The collaborator candidates were selected based on their relevance to the widenInterests and verified strengths of the builder.

Drafted for your approval — nothing is sent automatically. · model groq:llama-3.3-70b-versatile

↑ Back to top

A fifth vector — emergent projects from bisociation across the fleet. Scored in code; drafted for approval.

Bridges

Runs weekly · last searched 8/17/2026

Foundation Model Powered Agents

promising · 82

anyagent × FM-os

Autonomous Agents with Foundation Model Capabilities

pivot: Agent Training

first step: Integrate FM-os with anyagent for enhanced training

Graph-Based Agent Engineering

promising · 72

anyagent × graph-engineering-anything

Agents with Enhanced Graph-Based Reasoning

pivot: Graph Representation

first step: Integrate graph-engineering-anything with anyagent for graph-based agent engineering

Evaluation-Driven Agent Development

promising · 71

anyagent × eval-anything

Agents with Enhanced Evaluation Capabilities

pivot: Agent Evaluation

first step: Integrate eval-anything with anyagent for evaluation-driven development

Design-Driven Agent Development

plausible · 61

anyagent × design-anything

Agents with Enhanced Design Capabilities

pivot: Agent Design

first step: Integrate design-anything with anyagent for design-driven development

Loop-Based Agent Engineering

plausible · 51

anyagent × loop-engineering-anything

Agents with Enhanced Loop-Based Reasoning

pivot: Loop Representation

first step: Integrate loop-engineering-anything with anyagent for loop-based agent engineering

Forward-Deployed Agent Development

plausible · 50

anyagent × FDE-os

Agents with Enhanced Deployment Capabilities

pivot: Agent Deployment

first step: Integrate FDE-os with anyagent for forward-deployed development

Longevity-Driven Agent Development

speculative · 39

anyagent × longevity-loop

Agents with Enhanced Longevity Capabilities

pivot: Agent Longevity

first step: Integrate longevity-loop with anyagent for longevity-driven development

AI-Native Agent Development

speculative · 30

anyagent × ai-native-os

Agents with Enhanced AI-Native Capabilities

pivot: AI-Native Representation

first step: Integrate ai-native-os with anyagent for AI-native development

Emergent projects from bisociation across your fleet (Swanson A–B–C + conceptual blending), scored in code. Most combinations are noise — the ranking is the value. Drafted for you; nothing auto-built.

↑ Back to top