The next-gen embodied agent doubles its predecessor’s performance by reasoning, adapting, and self-learning in virtual environments — all powered by Gemini 2.5.
A New Leap for Embodied AI: Meet SIMA 2
Google DeepMind has unveiled a research preview of SIMA 2, the next iteration of its embodied AI agent that can observe, reason, act, and self-improve in 3D virtual environments. Powered by Gemini 2.5 flash-lite, SIMA 2 marks a major upgrade over the first SIMA, which launched in early 2024 with limited success in task completion.
Where SIMA 1 had a 31% success rate on complex tasks (vs. 71% for humans), SIMA 2 doubles performance, and takes a step closer to Artificial General Intelligence (AGI).
“It’s a more general agent,” said Joe Marino, senior research scientist at DeepMind. “It can self-improve based on its own experience — a key step toward general-purpose robots and AGI systems.”

What Makes SIMA 2 Different?
While SIMA 1 could follow instructions, SIMA 2 understands and reasons about its surroundings, forming high-level plans using natural language, emojis, and contextual clues.
Key upgrades include:
- Advanced reasoning: Uses Gemini to draw logical inferences (e.g., ripe tomatoes are red → walk to the red house)
- Emoji-based instructions: Try “🪓🌲” and SIMA 2 will chop down a tree
- Descriptive understanding: Recognizes and describes 3D surroundings in real-time
- Cross-environment transfer: Can perform well in unseen, photorealistic worlds
- Self-improvement: Learns from its own mistakes without human data
SIMA 2 isn’t just reacting — it’s thinking, learning, and planning.
Inside the Tech: How SIMA 2 Learns
SIMA 2’s architecture blends language models and interactive environments, making it a cognitive agent capable of:
- Understanding tasks through natural or symbolic inputs
- Interpreting environments by observing and labeling surroundings
- Planning actions with reasoning based on prior knowledge
- Acting in real time, like navigating worlds or manipulating objects
- Improving autonomously through trial-and-error reinforced by AI-generated feedback
Marino explained that in new virtual worlds, SIMA 2 receives self-generated tasks from one Gemini model and rewards from another. Over time, it builds its own playbook of strategies — learning without explicit human labels or gameplay footage.
The Power of Embodiment: Beyond Static AI
DeepMind defines embodied agents as systems that observe and act in an environment, be it physical or virtual — simulating how humans and animals interact with the world.
“If we want robots that operate in the real world, they need more than just sensor data. They need reasoning,” said Frederic Besse, senior staff research engineer at DeepMind.
Unlike static AI tools that only read documents or manage calendars, SIMA 2’s embodiment gives it real-world application potential — particularly for tasks that require spatial understanding, adaptability, and environmental interaction.
Examples include:
- Helping a user navigate virtual simulations
- Assisting robots in high-level task planning
- Running long-form simulations to train digital twins
- Interfacing with AI agents that manage or optimize complex environments
Real-World Demos: From No Man’s Sky to Genie
In the “No Man’s Sky” demo, SIMA 2:
- Recognizes a rocky terrain
- Identifies a distress beacon
- Chooses to investigate based on reasoning
It also worked within Genie, DeepMind’s photorealistic world model, accurately identifying butterflies, benches, and trees — demonstrating adaptability across realistic, dynamic simulations.
Not Ready for Robots… Yet
Although robotics is a key goal, DeepMind clarified that SIMA 2 is not yet deployed in physical robots, and is separate from its newly announced robotics foundation models.
However, the underlying capabilities — understanding instructions, planning actions, navigating space — are essential prerequisites for intelligent robots.
“This is about high-level cognition, not low-level motor control,” Besse noted.
The Road to AGI: What’s Next?
SIMA 2 isn’t a product release — it’s a research preview meant to explore collaborations, use cases, and next steps.
But it’s a clear marker of how language models like Gemini are evolving from text-based chat tools to full-fledged agents that understand, act, and learn inside interactive environments.
As DeepMind’s Jane Wang put it:
“We’re asking it to reason in a way that’s grounded in the environment — and that’s very hard.”
DeepMind’s SIMA 2 uses Gemini to power an AI agent that can understand, reason, act, and self-improve in virtual environments. Doubling its predecessor’s performance, SIMA 2 learns tasks through trial and error, understands emojis, and showcases the potential of embodied AI in future robotics and AGI systems.








