Frontier Systems for the Physical World
Robot learning, autonomous science, and new human-machine interfaces are three faces of one emerging substrate for physical-world AI, maturing on five shared primitives whose binding constraint is reliability at scale.
One substrate, three faces
The dominant, production-ready AI paradigm today runs on language and code, where scaling laws are well characterized and the data-compute-algorithm loop is spinning. A newer paradigm has been maturing alongside it in what a16z's essay calls a gestation phase: robot learning, autonomous science, and new human-machine interfaces, argued as three instances of a single emerging substrate rather than three separate fields.1
Five primitives are maturing concurrently across all three. Learned representations of physical dynamics are being approached from three converging directions: vision-language-action models that extend pretrained vision-language systems with action decoders, world-action models built on video-diffusion transformers, and native embodied foundation models trained from scratch on physical interaction data. Spatial intelligence, the work coming out of World Labs and Fei-Fei Li, fills a gap none of the three explicitly address on their own: none of them model the three-dimensional structure of the scenes they operate in.
Architectures for embodied action have converged on a dual-system design, a slow reasoning model paired with a fast visuomotor policy, and the most consequential recent development is reinforcement learning layered onto pretrained action models: one system folds laundry across fifty distinct garment types, assembles boxes, and makes espresso for hours unattended, more than doubling throughput and cutting failure rates by half or more compared with imitation learning alone.
Simulation and synthetic data have become the scaling infrastructure of the field, because a robot cannot yet have a billion physical interactions the way a language model can have a billion documents; the bottleneck shifts from collecting real-world data to designing diverse virtual environments, and simulation scales with compute rather than with human labor or physical hardware. An expanding sensory manifold, touch, muscle signal, subvocal speech, brain-computer interfaces, digitized smell, means the physical world communicates through a far richer channel set than vision and language, and every new consumer device built to capture one of these signals is simultaneously a new data-generation platform. And closed-loop agentic systems address sustained autonomy over long horizons, since a physical agent that drops a beaker of reagent cannot simply undo the action the way a text-based one can, which forces long-horizon memory, provenance, safety, and recovery into the design from the start.
A structural flywheel
The three domains reinforce one another. Robotics stress-tests every primitive at once; autonomous science and self-driving laboratories draw on them most completely, acting as a data engine that converts physical reality into structured, causal, empirically verified training signal; and new interfaces form a spectrum of increasingly high-bandwidth channels feeding both of the other two. Robotics enables autonomous science and the reverse is also true, while interfaces feed both. Together they open new scaling axes that are complementary to, rather than competing with, the existing digital frontier.
The binding constraint
The honest limitation is stated precisely in the source essay: even a 95% per-step success rate yields only about 60% success on a ten-step task chain, so reliability at scale, addressed largely through reinforcement-learning post-training, is what stands between an impressive demo and something deployable. The five primitives are maturing concurrently on the essay's own account, but nothing guarantees their progress compounds rather than bottlenecking on whichever one turns out to be slowest, most likely real-world reliability or safe closed-loop autonomy.
The founder-supply side
Micky Malka of Ribbit Capital reports a preference shift among founders in their early twenties that complements the capital-side argument: many of them find working on software alone boring and unchallenging, and want to build physical objects instead, to touch materials and build things that reach people's hands directly.2 He reads this as an inversion of the prior fifteen to twenty years, when the default instinct was to stay capital-light and software-only to move through cycles as fast as possible, and offers a generational explanation: a cohort that grew up with a phone, a tablet, and a computer from birth does not find bits interesting enough to be the whole world. If the observation holds broadly rather than describing only the founders who already reached one investor, it loosens the constraint on hard-technology formation from a direction nobody had been modeling, founder supply rather than capital or policy.
Practiced by
Connections
Loading connections…
References
- 01
Frontier Systems for the Physical World
a16z · article · 2026
- 02
Lessons From Backing The Best Founders In Fintech
Micky Malka · podcast · 2026
Related