Paul Jialiang Wu, PhD
Forward Deployed AI Engineer · Agentic Systems & Multi-Agent Orchestration · Enterprise AI in Production
wjlgatech@gmail.com · github.com/wjlgatech · linkedin.com/in/paul-jialiang-wu-phd · live portfolio ↗
SUMMARY
High-agency, founder-mindset engineer who ships bespoke agentic AI inside customer environments. As a Forward Deployed Engineer on an OpenAI enterprise engagement and technical lead on Accenture’s Physical AI team, I take generative AI from rapid prototype to production-grade reality — architecting the connective tissue between frontier models (Gemini-class VLMs, GPT-class models, NVIDIA Cosmos) and live customer infrastructure, then hardening it with evaluation and observability pipelines for accuracy, safety, latency, and cost-per-request. Creator of open-source multi-agent systems (loop-engineering-anything, super-u) that generate MCP-/agent-native tooling, grade it against reality with an independent referee, and refactor to convergence under hard safety gates. PhD-trained, with 8+ years delivering production ML/AI for Fortune 500 clients; fluent in multi-agent orchestration (ReAct, self-reflection, hierarchical delegation), RAG over structured and unstructured data, and the LLM-native metrics that decide whether an agent is enterprise-ready.
RECENT OPEN SOURCE — LAST 30 DAYS (JUNE 2026\)
12 active repositories, 280+ commits this month — all public at github.com/wjlgatech. Spanning agentic loops, model-quality evaluation, and token/latency optimization.
loop-engineering-anything*Self-Improving Agentic Systems · Creator · 88 commits this month*
- Open-source loop orchestrator (Python, 286 passing tests, CI on 3.11–3.13) that turns any API or codebase into a self-improving, agent-native CLI by driving a generate → judge → refactor → re-judge loop to convergence — walk away, return to a measurably better tool plus an audit report of how it improved.
- Quality is set by an independent referee, not model self-report (a “maker ≠ checker” evaluation pipeline): 5-dimension A–F grading with a hard safety gate that makes unsafe output a terminal, unshippable state regardless of other scores.
- Multi-signal convergence policy — plateau detection \+ regression rollback \+ iteration/token budget — prevents recursive agent degradation and enforces cost limits; every run, grade, and learning is persisted to SQLite for cross-run tracing.
- Fleet layer coordinates many loops under one goal with dependency-ordered execution and automatic feedback routing, escalating only high-judgment forks (safety block, stuck worker, gated result) to a human.
super-u*Multi-Agent Human-Upgrade Platform · Creator · 59 commits this month*
- Multi-agent application (FastAPI, spec-/guardrail-driven CI) across three agentic layers — Creator, Skillfy, Digital Twin — behind a single LLM-provider seam (Anthropic or any OpenAI-compatible host, incl. NVIDIA NIM) for portable model swaps.
- Skillfy is a RAG-style ingestion-to-tool pipeline: ingest text / website / PDF (incl. arXiv) / YouTube transcript / folder → triage gate (quality floor \+ overlap dedup against a knowledge graph) → edge-gated extraction → export as an installable agent tool (SKILL.md / MCP).
- Digital Twin is an observability layer for agent alignment/drift: dual heuristic \+ LLM-judge scoring against a vision kernel, plus per-call efficiency telemetry that meters every LLM call — the heuristic floor any model or prompt change must beat.
More shipped this month — same generate → judge → ship discipline:
- FDE-os — Forward-Deployed-Engineer OS: runnable Claude MCP server, RAG-grounded guide, and an eval-loop \+ criteria-scorer behind a live agentic webapp. *(25 commits)*
- loop — Desktop AI overlay with multi-LLM failover tiers that detect a dead model before the user does, plus one-click local (Ollama) inference. *(17 commits)*
- dreammaketrue — Any source → interactive knowledge graph; engineered 10× perceived AI latency via streaming, caching, and model pre-warm. *(17 commits)*
- cli-judge — LLM-as-judge scoring harness for CLI agents — a reproducible Definition-of-Done scorecard for ranking model/agent output quality. *(15 commits)*
- agentic-portfolio — Conversation-editable portfolio (owner-gated, voice input) with a self-healing, class-independent web harvester. Live app ↗ *(10 commits)*
- Google5Days-Agentic-Engineering — Developer-enablement curriculum with a deployable guide agent (swappable Gemini/Claude backends) and a tested hands-on track. *(7 commits)*
- External OSS: merged a self-healing 429-fallback fix into NousResearch/hermes-agent — honoring api\_key\_env across the fallback chain so agents recover from rate limits instead of failing.
EXPERIENCE
Applied AI Scientist / Forward Deployed Engineer — Accenture, Physical AI TeamJan 2024–Present
*Forward-Deployed Delivery · OpenAI-DSI Engagement*
- Primary technical interface for an executive-level enterprise AI deployment — recognized by a Senior Manager for upholding a ‘no surprises’ standard and earning go-to status as first point of contact, built on deep model understanding and contextual judgment at decision points.
- Led technical discovery and white-glove deployment: translated ambiguous customer requirements into shipped agentic workflows, and co-built with customer engineering teams to instill rigorous development practices and drive end-user adoption.
*Agentic Systems at Scale · WorkflowX / Agenticom*
- Designed and shipped a multi-agent business-process automation engine for enterprise clients: event-sourced state management for concurrent-agent consistency, idempotent retry across distributed agents, and full audit trail — live autonomous workflows with zero-downtime replanning.
- Solved the core agentic failure mode — agents over partially observable state must replan without losing context — with an event-sourced consistency model that doubles as a granular tracing / observability surface.
*Evaluation & Foundation Models · Multimodal / Physical Domains*
- Built a multi-model evaluation framework comparing 6+ vision-language models (incl. Gemini-class VLMs, NVIDIA VSS, Cosmos-Predict2) on accuracy, latency under real-world load, failure modes on edge cases, and integration cost — the basis for production model selection.
- Shipped a confidence-calibrated selective-verification pipeline — a closed-loop observability system that routes only high-uncertainty outputs to human review, cutting annotation cost \~90% while preserving coverage where model confidence is low.
- Technical lead on 4 simultaneous production systems: synthetic data generation, zero-shot safety detection, adaptive evaluation under distribution shift, and hybrid vision-language reasoning over live sensor streams.
Principal Data Scientist — GenentechJun 2021–Dec 2022
- Built and shipped production ML for biomedical research at scale, including 3 open-source PyPI packages for multimodal learning (tf\_multimodal, fastai\_multimodal, fast\_tfrs) adopted by the broader research community for large-scale signal fusion.
- Drove an 85% reduction in unproductive process time via AI workflow optimization; 5-star-rated delivery across all engagements.
Principal Data Scientist — Galvanize Inc.Sep 2019–Jun 2021
- Technical lead across 13+ end-to-end ML systems for Fortune 500 clients — time-series forecasting, NLP, anomaly detection, graph optimization, and large-scale recommendation — owning scoping, stakeholder alignment, technical discovery, and production delivery; 5-star-rated.
SELECTED PROJECT
DataCenterAR — Spatial AI deployed to Fortune 500 field techniciansAccenture Physical AI Team
- Agentic spatial-AI system (Meta Quest XR, real-time monocular depth, affine registration, temporal consensus filtering) that turns physical environments into an agent-navigable coordinate space; architecture evolved from rule-based to fully agentic — clients reported 10x improvement, ‘a coach, not a menu.’ Deployed to Walmart, Costco, and hyperscale data-center clients.
PUBLICATIONS & RESEARCH
SCWM: Self-Calibrating World ModelsNeurIPS 2026 (under review)
- Wu, P.J. et al. — online Bayesian calibration for sim-to-real transfer; 40% reduction in transfer error vs. domain-randomization baseline.
Physical AI: The Next Frontier in AI and RoboticsPreprints.org · Apr 2026
- Wu, P.J. et al. — DOI: 10.20944/preprints202604.0549.v1.
Failure Benchmarking of NVIDIA's VSS Tool: Insights from Vision-Language EvaluationWorking paper · first author · in preparation (no venue selected)
- Wu, P.J., Shah, A., Hosseini, P., Shahab, S. — Center for Advanced AI, Accenture. A five-mode VLM failure taxonomy made executable over deterministic synthetic stimuli, with the Temporal Grounding Score defined as data. Baseline VSS fails 5/5 modes; two frontier models pass 5/5 and the smallest reproduces 3/5 — so the failures track capability, not only late-fusion architecture.
Adapting the BARE Framework for Synthetic Data Generation in Vision-Language ModelsWorking paper · first author · in preparation (no venue selected)
- Wu, P.J., Shah, A., Hosseini, P., Shahab, S. — Center for Advanced AI, Accenture. Base-model diversity \+ instruction-tuned correctness as a twin-gated loop. Mode collapse reproduced (0.23 vs 0.62); base-model hallucination did not. Our own first negative was retracted as underpowered (0/18 ⇒ 95% CI [0, 16.7%]) and re-resolved at n=66 and n=72 as a genuine one.
Eval gates as training-signal factories (RewardForge)Working paper · unpublished, not submitted
- Every honest eval gate both certifies quality and labels preference data. 162 pairs from certified gates, zero human labels → DPO from scratch → LoRA → held-out hallucination 0.398→0.287; five seeds against a pre-committed criterion, treatment \+0.098±0.028 vs shuffled-label control −0.046±0.055.
Evidence-gated benchmarking of agentic tool use (MCP-Arena)Working paper · unpublished, not submitted
- 12 tasks × 6 failure-mode categories over real MCP servers, asserting the tool-call trace rather than the prose, so a fluent answer with no tool calls scores zero by construction; ships a hallucinating mock that must score 0 in CI.
Predicting a frontier lab's post-training stack, scored rather than assertedWorking paper · 12 predictions registered, Brier UNSCORED
- A ten-stage reconstruction of a frontier lab's training stack read from its publications, with 12 falsifiable claims registered 2026-08-13 and resolving by 2026-12-31 against the eventual technical report. 0 of 12 resolved, so the Brier score is reported UNSCORED rather than estimated.
The zero that wasn't evidence — power discipline for gates that report negativesMethods note · unpublished, not submitted
- An audit of our own false negative and the rule it produced: wherever a gate reports a negative it must state the n and the smallest effect that n could have resolved — with a runnable tool (Wilson intervals, rule of three, minimum detectable difference), because a rule without one is a slogan.
TECHNICAL SKILLS
Agentic Systems & Multi-Agent: Multi-agent orchestration (ReAct, self-reflection, hierarchical delegation), MCP servers, event-sourced state management, agentic replanning, idempotent retry, full audit trail; LangGraph / CrewAI / ADK-class patterns
Generative AI & Foundation Models: Gemini-class VLMs, GPT-class models, Qwen-VL, NVIDIA Cosmos / VSS; prompt & context engineering, multi-model evaluation, confidence calibration, distribution-shift detection
Data, Retrieval & RAG: RAG architectures, vector databases, embeddings, structured \+ unstructured data pipelines, document / PDF / web / transcript ingestion, dedup against knowledge graphs
Evaluation & Observability: Independent-referee evaluation pipelines, multi-dimension grading, LLM-native metrics (tokens/sec, cost-per-request, latency), granular tracing, regression rollback, safety gating
Cloud & Infrastructure: Google Cloud Platform (hands-on, ML projects) & Gemini model evaluation; Azure ML (production); NVIDIA AI stack (NIM, Cosmos, VSS); Python, PyTorch, FastAPI, Docker, Kubernetes, CI/CD, SQLite / Postgres, React, WebSocket
Forward-Deployed Delivery: Technical discovery, executive stakeholder alignment, white-glove deployment, prototype → production, customer engineering enablement, field-pattern → reusable module / product feedback loop
EDUCATION
Yale University — National BioMed Fellow, Computational Immunology
Georgia Institute of Technology — PhD, Bioinformatics
University of South Carolina — MS, Mathematics & Computer Science
Sun Yat-Sen University — BS, Applied Mathematics