The Shift Toward VLA 2.0 Architecture
On September 8, 2026, AI researcher Jim Fan (@DrJimFan) stated on X that traditional Vision-Language-Action (VLA) models are no longer sufficient and are being superseded by the concept of 'VLA 2.0', with Astra serving as its core foundation. According to Fan, multimodal coding represents the precise action space required for the 'System 2' (slow thinking and deliberate reasoning) brain in robotic systems.
Fan's remarks signal a pivotal shift in AI control architectures for robotics. Conventional VLA models integrate vision, language, and low-level control action tokens to help robots interact with the physical world. However, Fan contends that transitioning to a multimodal programming space provides the essential framework for robots to execute complex logical reasoning and multi-step planning before taking physical action.
Multimodal Coding as the New Action Space
In a post on X, Fan wrote: 'Astra is the new VLA. VLA is dead. Long live VLA 2.0: multimodal coding is the right action space for the System 2 (slow) brain of robotics.' The 'System 2' paradigm referenced here reflects deliberative cognitive processes that demand deeper computation and reasoning, contrasting with the rapid, reflexive control loops typically executed on hardware.
At present, Fan's post outlines a strategic conceptual direction rather than a full release. Detailed technical specifications, comprehensive research papers, internal architectural designs of Astra, and empirical benchmark results demonstrating the efficacy of VLA 2.0 over existing VLA solutions have yet to be disclosed.