Observe once, then perform it. That’s the concept behind a new robotic foundation model that can acquire a physical task from a demonstration lasting just 3 to 12 seconds. The robot observes the example, then attempts the task immediately. No retraining. No gradient adjustments. No fine-tuning. The system relies on contextual prompting instead of the typical training process.
The model is named GEN-1.5. It originates from Generalist AI, a company that characterizes it as an embodied foundational model. The company states the system learns directly from physical demonstrations instead of depending on engineering retraining. It treats each demonstration as a “physical prompt” within its contextual window. The model interprets sensorimotor data and deduces what action the demonstration aims to convey.
Generalist AI evaluated the model across 10 physical tasks. With just one demonstration, the robot succeeded 59% of the time on average. After incorporating 10 training steps using five minutes of data, that figure increased to 83%. The company presents the one-shot setting as a tool for rapid deployment, not maximum accuracy. Still, the results indicate genuine generalization from a minimal amount of physical input. These figures come from the company’s own reporting and external coverage of it.
Physical skills in this situation encompass a broad spectrum of actions. They include manipulation tasks, like opening a jar. They encompass object interaction, like withdrawing money from a wallet. Broader robotics research also addresses whole-body grasps, ambulation, and trajectory learning. Many of these skills require both motion control and force adaptation. They often blend perception, planning, and low-level control. The model generates these action trajectories at a frequency of 100 Hz.
This method builds on imitation learning, a technique where robots acquire knowledge by observing examples. A human or teleoperated system typically offers the demonstration. Some systems modify a demonstrated motion to adapt to a new object or setting. Others extract latent plans from recorded human play to inform their own actions. The model can even chain several of these learned prompts together to produce compositional behaviors for more complex tasks.
Robotics has experienced milestones like this before. A Berkeley-led team once demonstrated a robot learning to walk in about 20 minutes without any previous model training. Other researchers have instructed robots in imitative motions by combining observation with exploration. Some teams have even investigated learning skills from a single video. GEN-1.5 contributes a new chapter to that ongoing narrative.
References
- https://interestingengineering.com/ai-robotics/gen-1-5-robot-learns-physical-tasks-one-demonstration
- https://generalistai.com/blog/gen-1.5
- https://www.foxnews.com/tech/robots-learn-1000-tasks-one-day-from-single-demo
- http://dspace.mit.edu/bitstream/handle/1721.1/30251/CS0?sequence=1
- https://dl.acm.org/doi/10.1016/j.robot.2023.104427
- https://www.humanoidsdaily.com/news/generalist-ai-unveils-gen-1-5-one-shot-robot-learning-and-the-end-of-heavy-fine-tuning
- https://engineering.berkeley.edu/news/2022/10/step-by-step/
- https://the-decoder.com/gen-1-5-generalist-ai-teaches-robots-new-tasks-from-a-single-demo/
- https://proceedings.mlr.press/v229/wang23a.html
- https://www.cs.umd.edu/~reggia/supplement/index.html