NVIDIA has launched Alpamayo 2 Tremendous, a 34B-parameter vision-language-action (VLA) model for autonomous driving, beneath an open industrial license. The said design goal is the long-tail occasions: uncommon, multi-agent conditions that typical detection-and-prediction stacks deal with poorly. The mannequin pairs a 32B VLM spine, constructed on NVIDIA Cosmos 3 Tremendous Reasoner and post-trained with reinforcement studying, with a 2.3B diffusion-based motion decoder. From one cross over full-surround digicam video it emits a deliberate trajectory, a causal clarification of that trajectory, and a meta-action.
Is it deployable
Sure, and for industrial use from day one. The weights are launched beneath OpenMDW-1.1, the Linux Basis’s permissive license for open mannequin distributions; supply code is Apache 2.0. The license covers fine-tuning, by-product fashions and industrial redistribution. NVIDIA is making use of OpenMDW throughout your entire Alpamayo household, so earlier releases launched for R&D at the moment are deployable commercially with out further permission.
Inputs, outputs and coaching information
Inputs are multi-camera RGB video, textual content, and egomotion historical past with timestamps. The validated public pocket book profiles use six cameras and 4 historic frames per digicam. Egomotion is 3D translation plus a 3×3 rotation matrix, multi-timestep.
The trajectory API returns 64 waypoints spanning 0.1 to six.4 seconds at 0.1-second intervals. Every waypoint carries ego-frame XYZ and a 3×3 rotation matrix.
Coaching information is roughly 115,000 hours of multi-camera driving video with egomotion and trajectory annotations. It consists of about 3,700,000 Chain-of-Causation (CoC) traces — structured, causally linked explanations of driving selections. Picture coaching information exceeds one billion pictures.
Benchmarks
On LingoQA, Alpamayo 2 Tremendous data a Lingo-Decide rating of 79.2 and ranks first amongst almost 40 fashions evaluated. In NVIDIA’s testing it beat Qwen2.5-VL 72B by 17.0 factors, Gemini 2.5 Professional by 15.1, and GPT-4o by 23.2.
Two extra numbers matter for planning work. Closed-loop analysis with AlpaSim on 910 situations from the PhysicalAI-AV-NuRec dataset offers an AlpaSim rating of 1.50 ± 0.13. Open-loop analysis on 937 difficult samples from the PhysicalAI-AV dataset offers minADE₆ at 6.4s of 0.911m.
5 outputs from one mannequin
For every driving scenario, the mannequin produces a trajectory, a CoC hint explaining the choice, a meta-action comparable to yield or lane change, reasoning auto-labels, and visible query answering with 2D grounding.
That mixture is what makes the discharge attention-grabbing operationally. Builders can tie what the mannequin noticed to the motion it selected. CoC traces combine with NVIDIA Halos safety-validation workflows and help AI security aligned with ISO/PAS 8800.
Used as an autolabeler on proprietary fleet information, NVIDIA says the mannequin compresses annotation cycles from months to days.
Interactive explainer
Key Takeaways
- 34B VLA mannequin — 32B Cosmos 3 Tremendous Reasoner spine plus a 2.3B diffusion motion knowledgeable.
- OpenMDW-1.1 weights and Apache 2.0 code; industrial use and redistribution allowed, no further permission wanted.
- LingoQA Lingo-Decide 79.2, first amongst almost 40 fashions; AlpaSim 1.50 ± 0.13; minADE₆ 0.911m at 6.4s.
- One cross yields trajectory, Chain-of-Causation hint, meta-action, auto-labels, and grounded VQA.
- Cloud-scale mannequin examined on 1× H100 80GB at 72,115 MiB peak; distill it for in-car inference.
Take a look at the NVIDIA weblog and Hugging Face mannequin card. Additionally, be at liberty to observe us on Twitter and don’t neglect to hitch our 150k+ML SubReddit and Subscribe to our Publication. Wait! are you on telegram? now you’ll be able to be a part of us on telegram as properly.
Have to associate with us for selling your GitHub Repo OR Hugging Face Web page OR Product Launch OR Webinar and many others.? Join with us
Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is dedicated to harnessing the potential of Synthetic Intelligence for social good. His most up-to-date endeavor is the launch of an Synthetic Intelligence Media Platform, Marktechpost, which stands out for its in-depth protection of machine studying and deep studying information that’s each technically sound and simply comprehensible by a large viewers. The platform boasts of over 2 million month-to-month views, illustrating its reputation amongst audiences.








