NewsStocksDyna Robotics Unveils DYNA-2, a Robot Foundation Model Trained on One Million Hours of Human Video

Dyna Robotics Unveils DYNA-2, a Robot Foundation Model Trained on One Million Hours of Human Video

Author: Cryptopolitan·

Key Takeaways

  • DYNA-2 was trained exclusively on egocentric human video, which Dyna says is equivalent to about 170 years of continuous experience.
  • Dyna reports that DYNA-2 improved success rates on high-precision manufacturing tasks from 20% to as high as 80% to 90%.
  • The model uses a World-Action architecture that predicts the next video frame and the next action during pre-training.
  • Dyna says the model can be adapted to new robotic platforms with only hours of fine-tuning, including a reported 13-minute example for twisting off a bottle cap.
  • Dyna already has its earlier DYNA-1 robots deployed in commercial settings such as hotels, restaurants, laundromats, and gyms.
Dyna Robotics Unveils DYNA-2, a Robot Foundation Model Trained on One Million Hours of Human Video

Dyna Robotics introduced DYNA-2 on Monday, a new robot foundation model that the company says was trained on over one million hours of human video and zero robot data. The US-based company believes this approach could address the cost and scalability challenges that have constrained general-purpose robotics — a hurdle that has kept physical AI behind other domains where large-scale data pre-training has already driven rapid capability gains.

The Training Data Bottleneck in Robotics

A persistent obstacle in developing general-purpose robots is acquiring sufficient training data. Most robots today learn through teleoperation, a process in which a human operator manually guides the machine through tasks over many hours. This method is slow, expensive, and difficult to scale, ultimately capping how capable such systems can become. The problem is widely recognized across the industry, with players including Google DeepMind, Tesla, and Physical Intelligence all investing heavily in alternative data strategies for their own robot foundation models.

DYNA-2 takes a different path by eliminating robot data from the training phase entirely. Instead, the model was trained exclusively on egocentric human video — footage recorded from a person's point of view as they interact with objects and their environment. According to the company, this dataset amounts to approximately 170 years of continuous experience.

Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws:

• world-action models exhibit scaling law on human data across four orders of magnitude, from 1000… pic.twitter.com/wZamR0axzS — Dyna Robotics (@DynaRobotics) August 10, 2026

Co-founder Jason Ma explained the rationale: "Action data is scarce, but video is everywhere." He argued that physical intuition "can be learned directly from human video" rather than requiring millions of hours of data collected on a robot arm.

On high-precision manufacturing tasks, DYNA-2 lifted success rates from 20% to between 80% and 90%. Dyna attributes this improvement to the sheer scale of video pre-training. Across 15 benchmark tasks, models trained on larger volumes of human video consistently outperformed those trained on smaller volumes. This finding echoes the scaling laws that transformed natural language processing, where model performance improved predictably with increases in training data and compute — though whether those principles translate fully to physical-world tasks remains an open question.

How DYNA-2's Architecture Differs

Dyna distinguishes itself through a novel architecture. Rather than relying on a vision-language model — the approach used by systems like Google's RT-2 — DYNA-2 is a World-Action Model built on video generation. During pre-training, the model simultaneously predicts both the next video frame and the next action. Dyna says this dual objective gives the model an internal understanding of contact physics and spatial reasoning.

The company expects that knowledge transferred from human videos will generalize to different robotic platforms — including stationary arms, humanoid prototypes, and dexterous hands — even though none of these devices were involved in pre-training. Adapting DYNA-2 to a new platform reportedly requires only hours of fine-tuning rather than weeks of fresh data collection. For instance, Dyna reports that just 13 minutes of fine-tuning data enabled robotic hands to twist off a bottle cap. If this cross-platform transfer holds up under independent evaluation, it could significantly lower the barrier to deploying capable robots across diverse hardware configurations.

Real-World Deployment and Future Plans

Dyna already operates robots in commercial environments, a factor that differentiates it from many peers in the robotics field. The company's DYNA-1 robots are currently deployed in hotels, restaurants, laundromats, and gyms, according to reports. This operational experience provides Dyna with deployment feedback that many research-focused robotics companies lack.

Dyna was founded by Lindon Gao, York Yang, and Jason Ma, who previously worked at DeepMind. The company's stated roadmap is to scale the same training methodology to 10 million hours of video, characterizing this as a data collection challenge rather than a fleet-building one. As the field watches whether video-based pre-training can deliver on its promise, key questions include whether DYNA-2's demonstrated gains will generalize beyond benchmark tasks to unpredictable real-world environments, and how quickly competitors pursuing alternative architectures can close any gap.