NewsMacroReward AI Unveils OM1, a Robot Foundation Model That Learns Manipulation Directly From Human Data

Reward AI Unveils OM1, a Robot Foundation Model That Learns Manipulation Directly From Human Data

Author: Metaverse Post·

Key Takeaways

  • Reward AI unveiled OM-1, a robot foundation model trained exclusively on human manipulation data, with no teleoperation or on-robot experience used during training.
  • The model is reported to generalize zero-shot across different robot bodies, including tabletop arms, industrial arms, and humanoids, without robot-specific fine-tuning.
  • Demonstrations are captured with the Omnibody Hand, a seven-degree-of-freedom wearable built on the team's prior Stanford DexCap research and designed around functional grasping capabilities rather than replicating hand joints.
  • The One Data Interface combines tactile feedback, proximity sensing, global-shutter in-hand cameras, and electromagnetic hand-pose tracking, and Reward AI says adding electromagnetic sensing cut mean overshoot error by 60% at high speed in controlled tests.
  • The company states OM-1 can learn brand-new tasks, including those with challenging dynamics and long horizons, from less than 30 minutes of data, though the reported figures come from its own testing.
Reward AI Unveils OM1, a Robot Foundation Model That Learns Manipulation Directly From Human Data

Reward AI has unveiled OM-1, a general-purpose robot foundation model designed to acquire physical intelligence directly from human manipulation data. According to the company's official announcement, the system can generalize zero-shot across different robot bodies — including tabletop arms, industrial arms, and humanoids — without teleoperation data, on-robot training, or robot-specific fine-tuning. That last point is where much of the significance sits: teleoperation, in which a human drives a robot to record demonstrations, has long been a standard data-collection method in manipulation research, and it ties every recording to a specific machine.

The hardware foundation of the system is the Omnibody Hand, a wearable device with seven degrees of freedom built on DexCap, the team's prior research at Stanford. Rather than replicating the human hand joint by joint, the device is designed around functional capabilities: selecting useful contact points, reorienting objects within the palm, and transitioning smoothly between precision and power grasps. An ergonomic fit accommodates variations in hand size and finger proportion, ensuring that recorded movements reflect natural, uncompensated human behavior.

On top of this hardware sits the "One Data Interface" layer, which combines high-frequency tactile feedback, proximity sensing, global-shutter in-hand cameras, and electromagnetic hand-pose tracking to capture demonstrations at full human pace. Reward AI reports that augmenting visual-inertial tracking with electromagnetic sensing reduced mean overshoot error by 60% at high speed in controlled tests. Force data is also recorded along trajectories, so demonstrations encode both the path of a movement and the effort behind it.

Learning From People, Not Robots

OM-1 is trained exclusively on human data. According to Reward AI, neither teleoperation nor any on-robot experience is used during training; the policy instead learns to generate robot actions directly from human motion, distilling what the team describes as subconscious physical intelligence into a single, unified model. Because all demonstrations arrive through the same data interface, pre-training and post-training are merged into a single stage — new data requires no re-collection and no separate training pipeline per robot, task, or deployment. Cross-embodiment transfer of this kind has been a long-standing goal in robot learning, where manipulation policies have typically been trained and tuned for one platform at a time.

Architecturally, OM-1 processes each sensory modality — images, tactile signals, inter-finger proximity, and hand-pose trajectories — at its native sampling rate, preserving high-frequency contact cues alongside slower visual context. A temporal history of these streams allows the policy to reason about how contact and task progress evolve over time. Beneath the model, a reinforcement learning-based control layer running at high frequency translates actions into actuation for any given body, compensating for dynamics, disturbances, and system delays while optimizing transitions between successive policy predictions to keep motion smooth at speed.

In its conclusion, the company states that OM-1 can pick up brand-new tasks — including those involving challenging dynamics and long horizons — from less than 30 minutes of data. If validated, the approach positions human demonstration capture, sensing, learning, inference, and control as a single integrated stack — one the company frames as future-proof for robot hardware that has yet to be designed. Consistent with that hedge, the reported figures — the 60% tracking improvement and sub-30-minute task learning — come from the company's own reporting and controlled tests; performance across unfamiliar robot bodies in real deployments is the natural next test of the single-stack framing.

Source: Metaverse Post