Waymo Requires AI Project Evaluations Before Release
Self-driving car company Waymo emphasizes the importance of AI evaluations before product launch. According to the company, an AI model is not ready until its evaluation processes are mature.

Waymo, the Alphabet-owned autonomous vehicle company, has implemented an "eval-centric development" approach for its AI models. The company asserts that AI projects are not released until their evaluation frameworks are sufficiently mature, irrespective of the model's standalone performance. This strategy prioritizes continuous assessment and measurement to mitigate risks associated with AI deployment.
Manasi Joshi, Waymo's Director of Engineering for Systems Intelligence and Machine Learning, explained at VB Transform 2026 that evaluation is an integral part of the development lifecycle, rather than a final check. "We gauge the maturity of our projects based on the maturity of their evals," Joshi stated. To date, Waymo has accumulated over 220 million fully autonomous miles, reporting significantly fewer serious crash injuries compared to human drivers over the same distance.
This methodology mandates that enterprises must be able to reliably measure the performance of their AI systems before deploying them into production. Furthermore, evaluations must continue post-launch, as business processes, user behaviors, and incoming data evolve. Waymo draws parallels to the development of customer service agents or other AI applications.
The company places a strong emphasis on safety, testing its systems through extensive simulations and real-world driving scenarios, including rare and hazardous situations. Ultimately, human oversight remains critical in release decisions, not automated systems, given the high stakes involved. Balancing efficiency with reliability is paramount, and Waymo focuses on optimizing both its onboard vehicle systems and the supporting infrastructure.
Waymo's experience suggests that agentic AI requires a clear objective, representative data, continuous testing, and defined human decision-makers who are accountable for deployment. "Earning trust is supremely important," Joshi concluded.