QwQ 32B Preview
Qwen's reasoning-focused 'thinking' model. Generates long chains-of-thought before answering, similar to OpenAI's o1 and DeepSeek R1 lineage. Optimised for math and competition-style problem solving. The Preview tag means Qwen is iterating quickly; later versions may obsolete this one. Useful today for math-heavy workloads where a slow, careful answer is preferred to a fast wrong one.
- Parameters
- 32B
- Context length
- 33K
- Modality
- text
- Released
- 2024-11-28
- Tokenizer
- Qwen2Tokenizer
Memory & hardware
- VRAM (fp16)
- 64 GB
- VRAM (Q4)
- 19.2 GB
- Recommended
- H100 80GB or RTX 4090 (Q4)
- Quantizations
- fp16, q8_0, q4_k_m
Benchmarks
Hosted inference pricing
No provider we track publishes a per-token price for this model today. What each one used to offer is listed below.
No longer listed
Providers that used to serve this model. We don't republish their old rates — the dates below are the provider's own.
- togetherNot in Together AI’s serverless catalogue when we checked on 20 September 2026.Source ↗
Run it yourself
Drop-in commands for the three most common open-source inference paths. The Ollama tag is a best-effort match against the registry; verify the size variant before pulling.
vllm serve Qwen/QwQ-32B-Preview
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("Qwen/QwQ-32B-Preview")
model = AutoModelForCausalLM.from_pretrained(
"Qwen/QwQ-32B-Preview", device_map="auto", torch_dtype="auto"
)Qwen/QwQ-32B-Preview Related models
Same family or similar size — useful when shopping around.
32B sweet-spot Qwen 3, Apache 2.0. Reasoning-mode toggle inherited from smaller siblings; strong on math, code and agentic tool use. Fits on a single H100 in fp16 and on a 4090 at Q4.
- Context
- 33K
- License
- apache-2-0
- VRAM Q4
- 19.2 GB
32B sweet-spot model: strong reasoning, fits on one H100 in fp16, on a 4090 at Q4. The 32B size in particular hits a quality/cost knee — quality scales with parameters faster than cost up to ~32B, and slower afterwards. Favoured for production chat where 7B isn't sharp enough and where 70B+ would over-spec the hardware budget. Apache 2.0 licence.
- Context
- 128K
- License
- apache-2-0
- VRAM Q4
- 19.2 GB
Coding-specialised Qwen2.5 32B fine-tune. GPT-4o-class on HumanEval and BigCodeBench at the time of release. Trained on additional code-heavy data with extended pre-training. Apache 2.0. Natural pick for self-hosted coding assistants, code-review automation, and any agent loop that primarily writes code.
- Context
- 128K
- License
- apache-2-0
- VRAM Q4
- 19.2 GB
Bilingual EN/中 34B chat model. Apache 2.0 licensed with strong Chinese-language performance and competitive English chat quality. Good default for bilingual production workloads.
- Context
- 33K
- License
- apache-2-0
- VRAM Q4
- 20.4 GB
Vision-language variant of Yi 34B. Image-text reasoning via an MLP adapter on a CLIP encoder. Useful for bilingual EN/中 multimodal workloads where the major Western vision-language models underperform on Chinese text in images.
- Context
- 4K
- License
- apache-2-0
- VRAM Q4
- 20.4 GB
Cohere's 35B model tuned for RAG and tool use. The open weights are released under CC-BY-NC (commercial use requires the Cohere API). Strong multilingual coverage and a fine-grained RAG-mode output format that makes downstream citation easier.
- Context
- 128K
- License
- mrl
- VRAM Q4
- 21 GB