TAAFT
Free mode
100% free
Freemium
Free Trial
Prompts Deals

Cosmos3-Edge

By NVIDIA
Cosmos3-Edge is a 4 billion parameter world foundation model using a Mixture-of-Transformers architecture with two towers: an autoregressive tower for vision and text understanding and reasoning, and a diffusion tower for vision, audio, and action token prediction and generation. The towers share multimodal attention layers aligning language, video, audio, and action information. As a post-trained world action model, it operates at robot-control resolution of 640x360, generating 32 actions per inference with real-time control at 15 Hz on NVIDIA Jetson Thor. Actions across embodiments (cameras, vehicles, robot arms, humanoids) map to a shared representation encoding translation, rotation, and manipulation state. It can predict the visual outcome of an action or infer the action from its effect. A companion checkpoint, Cosmos3-Edge Policy (DROID), is post-trained for robot pick-and-place tasks.
New Multimodal Visit model
Released: July 20, 2026

Overview

Cosmos3-Edge is a 4 billion parameter open world model built to run on-device. It helps robots and vision AI agents understand their surroundings, reason in real time, and generate robot actions on edge hardware such as NVIDIA Jetson and RTX GPUs. It combines an autoregressive reasoning tower with a diffusion generation tower to unify scene understanding, video prediction, and action generation.

About NVIDIA

Industry: Computer Hardware Manufacturing
Company Size: 42000
Location: Santa Clara, California, US
Website: nvidia.com
View Company Profile
Last updated: July 21, 2026
0 AIs selected
Clear selection
#
Name
Task