At the VB Transform 2026 conference held in late July 2026, Manasi Joshi, Director of Engineering for System Intelligence and Machine Learning at Waymo, shared the autonomous driving company's 'eval-centric development' methodology. For Waymo, an artificial intelligence (AI) project is not considered ready for operation simply because the model performs well. Instead, readiness must be measured by the maturity of the evaluation system built around it. This rigorous operational philosophy aims to minimize physical risks when deploying AI in autonomous vehicles.
Background & Drivers
Unlike generative text AI or office automation applications, Waymo's models directly control vehicles navigating unpredictable real-world streets. According to Waymo, its autonomous vehicles have completed over 220 million rider-only miles (without a safety driver), with a bodily injury claim rate 17 times lower than human drivers over the same distance. However, to achieve this safety benchmark, Waymo does not rely solely on final checks prior to deployment. Instead, quality evaluation has been shifted deep into the core of the product development lifecycle, turning evaluation into a continuous process spanning training, post-training, and both open-loop and closed-loop simulation scenarios.
Technical Analysis & Technology
On a technical level, Waymo's system is clearly divided into two parts: the real-time onboard system and the off-board development infrastructure for data processing and simulation. Having adopted the Transformer architecture in 2017, Waymo has continuously expanded to Large Language Models (LLMs), Vision-Language Models (VLMs), and Vision-Language-Action (VLA) models. Currently, the company is integrating generative multimodal models into its foundation model strategy. To optimize increasingly scarce hardware resources, Waymo focuses heavily on 'data efficiency'—selecting the most informative training data samples rather than blindly chasing quantity. At the same time, the company leverages internal AI agents to analyze data distribution and identify edge cases during testing.
Expert Insights & Perspectives
According to Manasi Joshi, evaluation is not a one-time task to greenlight a model. Waymo combines real-world datasets, standardized benchmarks, and a massive simulation infrastructure to reconstruct billions of virtual driving miles. However, she emphasized that Waymo's release gatekeeping process is by no means 100% automated. 'This is not an AI-led system without human oversight. Lives are at stake,' Joshi asserted. The involvement of internal safety leaders and expert manual approval teams serves as the final gatekeeper before any software update is widely deployed.
Impact & Future Outlook
Waymo's approach offers a major practical lesson for businesses in Vietnam and globally when deploying agentic AI systems. The key takeaway is that having a powerful foundation model is not enough; enterprises must build specialized evaluation datasets, implement continuous testing against rare but hazardous corner cases, and, most importantly, establish clear human accountability mechanisms. In the future, as AI penetrates deeper into high-risk industries like healthcare, finance, and logistics, this 'eval-centric' philosophy will shift from a best practice to an absolute necessity.