AI Models

NVIDIA Cosmos 3: The World's First Fully Open Omnimodel for Physical AI

NVIDIA has announced Cosmos 3 at GTC Taipei COMPUTEX — the world's first fully open omnimodel for Physical AI, combining native vision reasoning, world generation, and action simulation in a single model. Available now in Super (32B) and Nano (8B) variants.

A
AIDeveloper44 Team
June 1, 2026·5 min read
NVIDIA Cosmos 3: The World's First Fully Open Omnimodel for Physical AI

Cosmos 3: Physical AI's Unified Foundation

Announced at NVIDIA GTC Taipei at COMPUTEX, NVIDIA Cosmos 3 is the world's first fully open omnimodel for Physical AI. It unifies three capabilities that have previously required separate specialized models: native vision reasoning, world generation, and action simulation — all in a single foundation model released under an open license.

Cosmos 3 is available today in two variants: Super (32B) and Nano (8B).

What Makes It an "Omnimodel"

Cosmos 3's mixture-of-transformers architecture separates two distinct functions:

  • A reasoning block that interprets what is happening in a scene
  • A generation block that uses that context to produce physically grounded outputs

This means the model can reason over a scene and simultaneously generate outputs across text, video, images, ambient sound, and robot action data. It's not a language model with vision bolted on — it's a unified system designed from the ground up for physical world understanding and generation.

Robot Action Data Generation

One of Cosmos 3's core capabilities is native action generation — producing numerical action data such as joint angles, gripper positions, and trajectory points that describe how a robot should move to complete a task. This is distinct from generating video of a robot — Cosmos 3 generates the actual control signals.

Developers can fine-tune Cosmos 3 for a specific robot embodiment, camera layout, workspace, or task. Agile Robots is already using Cosmos 3 to generate action-conditioned robot data for its humanoids including Thor 3, creating diverse task trajectories at scale. NVIDIA's own GEAR team uses it to develop video action models for embodied agents across games, simulations, and real-world robotics.

Cosmos 3 Nano post-trained policy leads on RoboLab (language-guided simulation tasks) and RoboArena (real DROID robot comparisons).

Smart Cities and Traffic Intelligence

Cosmos 3 can reason across a video scene to identify moving objects, predict path intersections, and anticipate future states — then generate dense captions, predicted scene changes, or scenario variations. This connects understanding, prediction, and alerting for vision AI agents in industrial and infrastructure environments.

Linker Vision uses Cosmos' vision language reasoning capabilities to analyze live camera streams, understand spatial contexts, and perform root-cause analysis across thousands of feeds for smart city and industrial solutions. Cosmos 3 tops the VANTAGE-Bench leaderboard for smart-infrastructure scene understanding and the TAR challenge for traffic anomaly reasoning.

Long-Tail Scenario Generation for Autonomous Vehicles

Rare events — collisions, edge cases, unusual pedestrian behavior — are critical for training physical AI systems but nearly impossible to capture at scale in the real world. Cosmos 3 can generate physically plausible video sequences of these scenarios, supporting synthetic data workflows for humanoids, arm robots, surgical robots, and autonomous vehicles. Cosmos 3 variants rank first on open-weights leaderboards on Artificial Analysis and top the Physics-IQ, R-Bench, and PAI-Bench leaderboards for world generation.

Open License and Availability

Cosmos 3 is released under the OpenMDW 1.1 license from the Linux Foundation — a model-centric open license that covers weights, architecture, documentation, datasets, benchmarks, and code in a single agreement. This makes it straightforward for developers to train, modify, redistribute, and deploy Cosmos 3 across physical AI workflows.

Developers can:

Enjoyed this?

Get more posts like this delivered to your inbox.

🚀 Join the AI dev community — follow us everywhere

© 2026 MARKTECHPOST AI MEDIA INC. All rights reserved.Terms & ConditionsPrivacy Policy
Beta Mode