TAAFT
Free mode
100% free
Freemium
Free Trial
Prompts Deals

FLUX-mimic

FLUX-mimic is a Video-Action Model (VAM) that generates robot manipulation actions by attending to the internal latent representations of a pretrained video generation backbone, rather than relying on low-dimensional action-only training data. A compact action decoder denoises action chunks directly from these latent video plans, so inference requires a single forward pass with no video rollout. The model is trained jointly on video and action prediction, letting it keep improving as a world model while learning to act. It uses techniques including post-training quantization, Real-Time Chunking, and adaptive action denoising with caching to run locally on a single NVIDIA RTX 5090 GPU. In benchmarks on a soft-body kitting task, it reached 95 percent success without task-specific fine-tuning. It is being tested and deployed on real production use cases such as parts kitting, component insertion, and soft-body assembly with manufacturing partners including Audi.
New Multimodal Visit model
Released: July 23, 2026

Overview

FLUX-mimic is a Video-Action Model built on the FLUX 3 video backbone, enabling general-purpose robot manipulation. It predicts how a manipulation scene will evolve and decodes robot action chunks directly from the video model's internal latent representations, requiring only a single forward pass at inference with no video rollout. Deployed on real factory tasks including soft-body kitting and assembly, running locally on a single NVIDIA RTX 5090 GPU.

About Mimic AI

Company Size: 30
View Company Profile
Last updated: July 27, 2026
0 AIs selected
Clear selection
#
Name
Task