The Path to AGI: from pattern repositories to auditable, governable agency.
Co-Editor-in-Chief, ACM Books ·
Founder and CEO, QuadriumAI ·
Director, Stanford AGI Lab
Adjunct Professor, Computer Science, Stanford University (2019–2026) ·
Advisor, Clinical Mind AI Lab, Stanford Medicine
Director of Research, Google (2006–2012) · Fellow of ACM and IEEE
System-2 Reasoning: From Semantic Anchoring to Causal Intelligence, ACM Books, July 2026.
A lecture tour across the trilogy's themes: anchoring, debate, causal audit, regret, transactions, memory, TRACE and the Quadrivium, physical AGI, and wisdom. 2026.
Feedback, memory, and causal reasoning for embodied AGI. Symposium on Humanoid Robotics & Sovereign AI, with Oussama Khatib and Hiroshi Ishiguro. February 2026.
Multi-LLM Agent Collaborative Intelligence. First edition 2024; acquired and published by ACM Books, December 2025.
One prior cannot see its blind spots; committees of priors can. Debate, critique, and transactional planning.
ACM Books 2025Top seller
Semantics is not in the box; it is in the binding. Anchoring, causal audit, and regret as first-class objectives.
ACM Books 2026
Memory you can trust, effects you can audit, and a wisdom layer that governs what is optimized.
Forthcoming 2027
What an LLM is (a pattern repository) and where semantics comes from: anchoring strength S = ρd − dr − log k, threshold-governed regime shifts, and the boundary between model and system. ICML 2026arXiv 2025
CRIT scoring, SocraSynth structured opposition with contentiousness as behavior control, and EVINCE's information-theoretic controller: committees beat their best member, measurably. The DIKE–ERIS checks-and-balances framework extends structured opposition into ethical governance. ICML 2025NeurIPS SafeAI 2024IEEE 2023100+ citations
Rung collapse, sycophancy, and Wise Refusal, diagnosed at scale: CausalT5K, RAudit's blind process audits, and the Skepticism Trap and Scaling Paradox in frontier models. KDD 2026ACL 2026
The end of what-only feedback: epistemic regret critiques the why, temporal regret measures the when, and RLER turns accumulated interventional evidence into reward. arXiv 2026
SagaLLM and ALAS bring compensable, disruption-aware planning and execution; REALM-Bench supplies the real-world planning benchmark that stress-tests them; Mnemosyne and the ATP contract promote admission, compensation, and durable obligations into substrate guarantees. VLDB 2025KDD 2026arXiv 2026
TRW's world-model-as-materialized-view bridge between cognitive and physical AI, and Architectural Wisdom, governing what systems optimize. arXiv 2026
The field is converging on this program under new names. This table keeps the record straight: the term now in circulation, the name and date it carries here, and where it lives.
| Circulating term (2025–26) | This program's name | First stated |
|---|---|---|
| “Reasoning models,” test-time compute | System-2 coordination over pattern repositories | Vol. 1 (2024–25); ICML 2026 |
| Process reward, critique models | Epistemic regret: why-feedback, label-free critique | ERM (Feb 2026); Vol. 2 |
| Agent durability, workflow engines, “sagas” | Transactional agents: SagaLLM/ALAS; the generated→admitted break | VLDB 2025 |
| Agent benchmarks, agentic evaluation | REALM-Bench: real-world planning under disruption; the on-ramp that carried SagaLLM’s adoption | REALM (2025); KDD 202640+ citations |
| Agent memory (vector databases) | Governed memory: ATP admission, three times, durable obligations | Mnemosyne (Jun 2026); Vol. 3 |
| “Grounding,” does-it-really-understand | The semantic-anchoring invariant: S = ρd − dr − log k | UCCT (Jun 2025); ICML 2026 |
| World-model reliability | The world contract: materialized views, commitment reads, priced refresh | TRW (Jul 2026); Vol. 3 |
| Multi-agent “discussion,” LLM-as-judge, process reward | Behavior-controlled debate; and RCA: the first no-gold-label judge, grading process, not answers | SocraSynth (2023), EVINCE (2024); ACL 2026 |
| Refusal calibration, abstention, “knowing when not to answer” | Wise Refusal, scored on three axes: Utility, Safety, calibrated abstention (CausalT5k’s 5,147 expert cases) | CausalT3/T5k; KDD 2026 |
| Knowing-doing gap, faithfulness of stated reasoning | Detection is not correction: models name the flaw (77–91% detection), then endorse the claim anyway (only 26–42% aligned) | KDD 2026 |
| AI constitutions, oversight | Checks-and-balances: DIKE–ERIS three-branch governance | NeurIPS SafeAI 2024 ICML 2025 |
The transactions thread is inheritance, not analogy: high-performance transaction processing built with Jim Gray; workflow with Dieter Gawlick at DEC; a Stanford Ph.D. under Hector Garcia-Molina, co-author of the 1987 Sagas paper that SagaLLM extends four decades later. The data thread was earned in computer vision's feature-engineering years, where scale kept beating cleverness, and became the data-centric program and the ImageNet sponsorship. The applied thread ran a decade of healthcare AI to an XPRIZE. The pre-LLM virtual-assistant years showed exactly where traditional NLP ends, which made the 2023 commitment to LLMs a decision, not a fashion. And the wisdom thread rests on fifteen-plus philosophy courses, a student philosophy-journal editorship, and published poetry: Volume 3's Part III was five decades in preparation.
Full publication list at Google Scholar →
As Director of Google Research (Beijing), our team built web-scale annotated datasets and parallel learning infrastructure years before "data-centric AI" was coined: one of the first web-scale annotated image datasets, sponsorship of Fei-Fei Li's ImageNet project, and a series of parallel ML algorithms on MapReduce, consolidated in the Springer book Foundations of Large-Scale Multimedia Information Management and Retrieval (2011), whose Chapter 2 formulated the data-driven + model-based hybrid (DMD) a decade early.
Scalable support vector machines on distributed systems (NeurIPS 2007).
30K+ web images annotated via distributed LDA on MapReduce.
Distributed Gibbs sampling on MapReduce.
Large-scale graph partitioning on distributed infrastructure.
Scalable pattern mining for web-scale data.
Sponsored Fei-Fei Li's ImageNet project at Google Research.
Early distributed CNN training, anticipating the deep learning revolution.
DMD hybrid architecture; all parallel algorithms documented (Chs. 9–12).