A historical investigation from the semantic gap and SVM active query-concept learning to automatic annotation, data scaling, mutable datasets, Web-scale learning, ImageNet, and modern data-centric AI.
TL;DR: Large datasets, data-driven learning, and data-centric AI are related but not identical. Early image-semantic systems learned labels from examples; SVMActive used uncertainty to choose which human semantic label to acquire next; later work asked which categories, features, and training instances should change when the model failed. The article treats ImageNet as a defining scale transition without turning the history into a single-origin story.
A mother tidies a study but disrupts her daughter’s work. A robot faces shared laundry, a recently reused office cup, and a leak hidden behind a wall. These everyday cases expose the gap between modeling physical possibilities and establishing the current, person-specific context needed for action.
TL;DR: Better reconstruction, prediction, and motor control do not by themselves establish what is current, intended, and authorized. The essay connects these gaps to eight grounding questions from The Path to AGI, Volume III: freshness, reference, effect, ownership, provenance, interpretation, verification, and sufficiency. The goal is enough reliable context for the task—not continuous tracking of every shared object.
A critique and boundary demarcation of Richard Sutton's "Bitter Lesson." Tracing data-centric AI from its real origins at Google in 2006 to AlphaFold, this essay explains why data-driven learning solves System-1 associative pattern matching, but cannot bridge the four structural gaps required for AGI: grounding, causality, transactional memory, and meta-cognitive initiative.
TL;DR: Richard Sutton is not the original pioneer of data-centric AI, nor is data-driven scale the complete answer to AGI. While data and compute systematically beat human feature engineering in closed-world, noiseless settings like AlphaZero or crystallized protein structures, they break down in open-world reality. Because human life contains absorbing ruin states (you cannot trial-and-error Clorox on COVID patients), uncalibrated subjective metrics (pain scales have no universal ruler), and long-horizon operational constraints requiring undo/redo Sagas, general intelligence requires an external System-2 regulatory architecture—as explored throughout The Path to AGI trilogy.
A response to the polarized industry debate between raw scaling and world models. Introducing the 2,000-page trilogy (MACI, System-2 Reasoning, Beyond Intelligence) as the Fifth Paradigm: Verified Causal Collaborative Intelligence.
TL;DR: Foundation models are high-capacity System-1 statistical pattern repositories, not AGI. Scaling, alignment, process supervision, and world models all fail because they remain locked in Pearl's Level 1 (association). The trilogy delivers the solution: Operational AGI emerges when the System-1 substrate is governed by external System-2 regulatory controls (contextual, causal, temporal, and meta-cognitive); the ASI transition automates strategic initiative; and Wisdom governs what systems are permitted to optimize.
A critique of Jensen Huang’s suggestion that AI makes foundational mathematical drills obsolete, evaluated through Schrödinger’s habit-formation laws and the CoCoMo multi-level feedback queue.
TL;DR: Conflating mature cognitive tool offloading with developmental neural wiring is a category error. Foundational arithmetic drills do not prepare children to be manual calculators; they automate mechanics into unconscious reflex (shifting latency from 10s down to 100ms), which is the necessary prerequisite to liberate single-threaded conscious working memory for genuine System-2 reasoning and error auditing.
Analyzing the shift in AISTATS 2027 to mandate centralized factual AI reviews, comparative metrics against ICML, NeurIPS, and ICLR, alongside risks of platform scraping and semantic cryptomnesia.
TL;DR: Honor-system AI bans failed, turning peer review into a covert, token-level diff engine. AISTATS 2027 normalizes hybrid scientific workflows by introducing centralized, temporally blinded factual audits. In response, authors must adopt defensive submission engineering: pre-submission cross-examination sweeps, strict disambiguation of numeric claims, and inverse drafting (compression over padding).