The humanoid robotics boom is no longer just hype; it is a capital-intensive reality. Startups NEURA Robotics and Figure AI raised $1.4 billion and $1 billion, respectively, while Apptronik raised $520 million. China’s Unitree went public, raising roughly $900 million, and Agility Robotics plans to go public via SPAC later this year. While hardware actuators get the headlines, the real intelligence driving these machines lies in a specific class of AI models: the Vision-Language-Action (VLA) model. As we see robots sorting packages and folding boxes, understanding the underlying architecture is critical for developers looking to build or integrate with this emerging infrastructure.

From Text to Action: The VLA Paradigm

Unlike Large Language Models (LLMs) that map text to text, VLAs map text, images, and robot sensor data to a series of physical robot actions. This architecture, employed by major players like Nvidia, Figure, and Physical Intelligence, leverages the same transformer components that made LLMs successful, specifically attention mechanisms. A prime example is Physical Intelligence’s open-weight π0.5 VLA, released in 2025, which has become widely used even if it isn’t the most advanced model available. The shift from pure language processing to action generation requires a fundamental change in how we view input-output mappings in neural networks.

The Math Behind the Motion

At its core, the AI controlling these robots relies on linear algebra operations that developers might recognize from basic machine learning courses. The source article breaks down how neural networks are effectively implemented as matrix multiplications, where input vectors are multiplied by weight matrices and added to bias vectors before passing through nonlinearities like ReLU or GELU. This structure allows the network to approximate complex functions, mapping high-dimensional sensor data to precise motor commands. Training involves a forward pass to generate an action, a loss function calculation (often cross-entropy or mean squared error), and backpropagation to adjust weights via gradient descent. The key difference in robotics is the complexity of the output space: instead of predicting the next token in a sentence, the model predicts the next set of joint angles or force vectors.

Attention Mechanisms in Physical Space

The article details how attention mechanisms, traditionally used for context in text, are adapted for physical tasks. Just as GPT-3 uses attention to understand semantic relationships between words, VLAs use similar mechanisms to correlate visual patches and textual instructions with motor outputs. The process involves converting images into patches—essentially visual tokens—and embedding them into vectors. Unlike text, where a fixed vocabulary exists, visual embeddings use learned functions to convert pixel data into high-dimensional vectors. These vectors are then modified by attention layers that calculate query, key, and value matrices, allowing the robot to focus on relevant parts of the environment. This enables the robot to understand that a 'cup' in a specific location requires a different grasping action than a 'cup' in a different orientation.

Key Takeaways

  • VLA models are the dominant AI architecture for current humanoid robots, used by Figure, Unitree, and Nvidia.
  • The π0.5 model from Physical Intelligence is a key open-weight example released in 2025, serving as a standard reference.
  • Robotic AI relies on the same transformer and attention mechanisms as LLMs but maps multimodal inputs to action outputs.
  • Major funding rounds for robotics startups indicate that software-defined motion is becoming as valuable as hardware design.

The Bottom Line

VLA models represent the critical software layer that transforms raw hardware into capable robots, making understanding their transformer-based architecture essential for developers entering the robotics space.