Robots Atlas>ROBOTS ATLAS
Qwen3-32B

Qwen3-32B

32B · Family: Qwen
Dense 32.8B-parameter LLM from Alibaba’s Qwen3 family with hybrid thinking modes and up to 128K context.
✓ Active✓ Public access⚖ Open sourceLLMReasoning model📁 Qwen
Context window
128K (natywnie 32 768, do 131 072 z YaRN)
tokens
Parameters
32.8B (31.2B non-embedding)
parameters
Release date
29 April 2025
Access:DownloadAPIHostedDeployment:💻 Local☁ Cloud

Overview

Qwen3-32B is the largest dense large language model in the Qwen3 family developed by the Qwen team at Alibaba Cloud, released on 29 April 2025 under the Apache 2.0 license.

Architecture

The model has 32.8B parameters (31.2B non-embedding) and 64 transformer decoder layers. It uses Grouped Query Attention with 64 query (Q) heads and 8 key/value (KV) heads, SwiGLU, Rotary Position Embeddings (RoPE), RMSNorm with pre-normalization and QK-Norm. The QKV bias used in Qwen2 was removed. The model natively supports a 32,768-token context, extendable to 131,072 tokens via YaRN.

Thinking modes

A defining feature of Qwen3 is its hybrid operation: a thinking mode for complex reasoning, math and coding, and a non-thinking mode for fast dialogue. The modes can be switched dynamically based on the query or the chat template.

Training

The model was pretrained on roughly 36 trillion tokens covering 119 languages and dialects, followed by a multi-stage post-training pipeline (long chain-of-thought SFT and reinforcement learning). Qwen3-32B has strong agentic and external tool-integration capabilities. Weights are released under the Apache 2.0 license.

Classification
LLMReasoning model
Family: Qwen
Access & deployment
DownloadAPIHosted
LocalCloud
Weights: Open source
Key parameters
📏 Context: 128K (natywnie 32 768, do 131 072 z YaRN)
🧩 Parameters: 32.8B (31.2B non-embedding)
Tools · ✓ Fine-tuning
📥 Input: text

Technical specification

Context window
128K (natywnie 32 768, do 131 072 z YaRN)
tokens
Parameters
32.8B (31.2B non-embedding)
parameters
License
Apache 2.0
Hardware requirements
BF16 weights ~65 GB; single-GPU inference typically needs quantization (INT4/AWQ/GPTQ) or 2+ A100/H100 80 GB GPUs.
Features:Tool useFine-tuning
Modalities
⬇ Input
text
⬆ Output
textcode

Capabilities and applications

Native model capabilities
Reasoning
The model's ability to perform multi-step logical inference, solve complex problems and decompose tasks into steps.
Category: reasoning
Coding
Generating, completing, explaining and debugging code across multiple programming languages.
Category: coding
Multilingual
Understanding and generating text in many languages and translating between them.
Category: language
Long context
Processing very long inputs (tens to hundreds of thousands of tokens) while maintaining coherence.
Category: language
Function Calling
Category: planning

Benchmark results

14 benchmarks
MMLU
accuracy · Qwen3-32B-Base
83.61%
📄 technical_report
MMLU-Redux
accuracy · Qwen3-32B-Base
83.41%
📄 technical_report
MMLU-Pro
accuracy · Qwen3-32B-Base
65.54%
📄 technical_report
SuperGPQA
accuracy · Qwen3-32B-Base
39.78%
📄 technical_report
BIG-Bench Hard (BBH)
accuracy · Qwen3-32B-Base
87.38%
📄 technical_report
GPQA
accuracy · Qwen3-32B-Base
49.49%
📄 technical_report
GSM8K
accuracy · Qwen3-32B-Base
93.40%
📄 technical_report
MATH
accuracy · Qwen3-32B-Base
61.62%
📄 technical_report
EvalPlus
accuracy · Qwen3-32B-Base
72.05%
📄 technical_report
MultiPL-E
accuracy · Qwen3-32B-Base
67.06%
📄 technical_report
MBPP
accuracy · Qwen3-32B-Base
78.20%
📄 technical_report
CRUX-O
accuracy · Qwen3-32B-Base
72.50%
📄 technical_report
MGSM
accuracy · Qwen3-32B-Base
83.06%
📄 technical_report
MMMLU
accuracy · Qwen3-32B-Base
83.83%
📄 technical_report

Pricing

Technical architecture