Qwen models
Browse all models from this model family.
ID
Model
Company
Type
Primary task
Open source
-
By AlibabaQwen Image 3.0 is Alibaba's third generation image generation model, built to produce information dense images such as infographics, newspaper layouts, and documents with legible small text and mathematical notation. It accepts prompts of up to 4,500 tokens, renders text natively in 12 languages, and can incorporate live web data such as weather graphics. No benchmarks, parameter counts, or downloadable weights have been released for this version.NewImageReleased 2d ago
-
By AlibabaQwen Audio 3.0 TTS is a production oriented text to speech model built on a 12.5 Hz speech tokenizer and a five stage progressive training pipeline. It supports zero shot voice cloning, 16 languages, 20 Chinese dialect regions, and one pass synthesis up to 3 minutes, with strong robustness to noisy or degraded reference audio.NewAudioReleased 3d ago
-
By AlibabaQwen Audio 3.0 Realtime Flash is Alibaba's speed optimized real time duplex voice dialogue model. It listens and speaks simultaneously, supports millisecond level responses, autonomous tool calling, and empathetic tone and emotion adaptation, while resisting interruption from background noise or side conversations.NewAudioReleased 7d ago
-
By AlibabaQwen-Audio-3.0-Realtime-Plus is a real-time duplex speech interaction model from Alibaba's Qwen family. It accepts streaming text and audio input and produces streaming text and audio output over a WebSocket connection, with millisecond-level responses, barge-in interruption, and reasoning depth that adapts to task complexity.NewAudioReleased 8d ago
-
By PrismMLTernary Bonsai 27B AWQ 4bit is a 4 bit AWQ quantized build of PrismML's Ternary Bonsai 27B, a multimodal model based on Qwen3.6 27B. It accepts text and images, supports a 262K token context window, and is designed for reasoning, coding, and agentic tool use workflows served efficiently via vLLM or SGLang.NewMultimodalReleased 9d ago
-
By PrismMLTernary Bonsai 27B is a 27B parameter multimodal language model from PrismML, built on a hybrid linear and full attention (Qwen3.5 based) architecture with a 262,144 token context window. This repository provides the full size FP16 safetensors weights, an alternative to the natively packed 2 bit ternary format for use with standard model-serving tooling.NewMultimodalReleased 9d ago
-
By PrismMLBonsai 27B is a 27.8B-parameter multimodal language model built on Qwen3.6 27B, quantized end-to-end to 1-bit or 1.58-bit ternary weights across embeddings, attention, MLPs, and the LM head. It accepts text and image input with a 262,144-token context window, and supports reasoning, tool use, and coding, while running locally on phones, laptops, and edge GPUs.NewMultimodalReleased 9d ago
-
By AlibabaOvisOCR2 is a compact 0.8B parameter end-to-end model for page-level document parsing. Given a document page image, it generates a Markdown representation in natural reading order, covering text, formulas, tables, and visual regions. Built by post-training Qwen3.5-0.8B with a multi-stage SFT, RL, and OPD training recipe.NewMultimodalReleased 10d ago
-
By PrismMLA 27B parameter language model compressed to ternary (1.71 bit) weights, retaining 95% of full precision intelligence while shrinking deployed size to about 7.2 GB. Runs full 27B class reasoning, math, and coding on a standard laptop or single GPU with a 262K token context window.NewMultimodalReleased 22d ago
-
By PrismMLTernary Bonsai 27B GGUF is a 27B parameter reasoning language model compressed to true 2 bit ternary weights (1.71 bits per weight), derived from a Qwen3.6-27B hybrid attention backbone. It runs in a roughly 7.2GB footprint on laptops, single GPUs, and CPUs via llama.cpp, while retaining strong math, coding, tool use, and long context (262K token) reasoning performance.NewMultimodalReleased 22d ago
-
By PrismML1-bit quantized version of the 27B-parameter Qwen3.6-27B language model, using binary g128 weight representation (1.125 bits per weight) for a ~3.9GB deployed footprint, about 14.2x smaller than FP16. Supports a 262K token context, retains around 89.5% of FP16 accuracy across reasoning, math, coding, and tool-use benchmarks, and runs on-device on phones, laptops, and GPUs via Apple MLX and CUDA.NewMultimodalReleased 22d ago
-
By BitInfArko T is a 4B parameter model that converts natural language design requests into executable, parametric CAD programs. Unlike text to 3D systems that output static meshes, it preserves named features, parameters, and construction history so the result is an editable engineering design, not just a rendered shape.NewCodingReleased 23d ago
-
By AlibabaQwen-AgentWorld-35B-A3B is a 35B Mixture-of-Experts language world model with 3B active parameters, built on Qwen3.5-35B-A3B-Base. It natively simulates seven agent interaction domains: MCP tool calling, search, terminal, software engineering, Android, web, and game environments. Trained end-to-end via CPT to SFT to RL. Supports 262,144-token context and uses thinking mode by default.NewMultimodalReleased 1mo ago
-
By AlibabaLanguage-conditioned video world model for embodied AI that predicts physically grounded future visual trajectories from natural language instructions and current observations. Supports robotic manipulation, autonomous driving, indoor navigation, and human-to-robot transfer. Built on a 60-layer double-stream diffusion transformer coupled with Qwen2.5-VL semantics.NewVideoReleased 1mo ago
-
By AlibabaLOGOS-8B (Language Of Generative Objects in Science) is an 8B parameter autoregressive scientific foundation model. It encodes proteins, antibodies, small molecules, chemical reactions, and materials into a unified scientific grammar, enabling a single model to perform generation, prediction, and design tasks across the natural sciences without relying on 3D geometry.NewTextReleased 1mo ago
-
By AlibabaQwen3-Max is Alibaba Cloud’s closed-source trillion-parameter flagship LLM for coding, reasoning, enterprise agentic AI, and tool-use workflows.NewMultimodalReleased 2mo ago
-
By AlibabaQwen3.5-LiveTranslate-Flash-Realtime is Alibaba’s vision-enhanced real-time audio/video translation model for live multilingual interpretation across 60 languages.NewMultimodalReleased 2mo ago
-
By AlibabaQwen3.6-27B is Qwen’s open-weight multimodal model for coding, agent workflows, long-context reasoning, and vision-language tasks. It combines a 27B causal language model with a vision encoder, supports image-text-to-text use, and offers a native 262,144-token context window extendable to about 1.01M tokens.MultimodalReleased 3mo ago
-
By AlibabaQwen 3.6 Max Preview appears to be Alibaba’s newest preview flagship in the Qwen line, positioned above Qwen3.6-Plus and aimed at top-end coding, world knowledge, and instruction-following performance. I could not read the exact official blog page you linked directly, so I’m relying on current public reporting for this description.MultimodalReleased 3mo ago
-
By AlibabaQwen3.6-35B-A3B is Qwen’s open-weight multimodal MoE model for coding, agentic workflows, long-context reasoning, and vision-language tasks. It has 35B total parameters with 3B activated, supports image-text-to-text use, preserves reasoning context across turns, and natively handles 262,144 tokens with extension up to about 1.01M.MultimodalReleased 3mo ago
-
By AlibabaWan2.7-Video is Alibaba’s comprehensive video generation suite for director-style creation and editing. It includes text-to-video, image-to-video, reference-to-video, and video-editing models, and is built to support natural-language control over actions, dialogue, appearance, cinematography, and multi-shot narrative workflows.MultimodalReleased 3mo ago
