Llama 3.2 1B
The smallest Llama 3 release, designed for on-device inference on phones and laptops. The 1B model runs comfortably in <2 GB of RAM at Q4 quantization and is fast enough for real-time chat on a modern smartphone. Useful for edge inference, on-device assistants where round-tripping to a server is undesirable, and as a draft model for speculative decoding in front of a larger Llama 3 variant.
- Parameters
- 1B
- Context length
- 128K
- Modality
- text
- Released
- 2024-09-25
Memory & hardware
- VRAM (fp16)
- 2 GB
- VRAM (Q4)
- 0.6 GB
- Recommended
- CPU or any GPU
- Quantizations
- fp16, q8_0, q4_k_m, gguf
License: Llama 3 Community License
- SPDX
- —
- Commercial use
- Yes
- Modification
- Yes
- Redistribution
- Yes
Benchmarks
Hosted inference pricing
No provider we track publishes a per-token price for this model today. What each one used to offer is listed below.
No longer listed
Providers that used to serve this model. We don't republish their old rates — the dates below are the provider's own.
- groqGroq shut down llama-3.2-1b-preview on 14 April 2025.Source ↗
Run it yourself
Drop-in commands for the three most common open-source inference paths. The Ollama tag is a best-effort match against the registry; verify the size variant before pulling.
ollama run llama3.2:1b
vllm serve meta-llama/Llama-3.2-1B
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.2-1B")
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.2-1B", device_map="auto", torch_dtype="auto"
)meta-llama/Llama-3.2-1B Related models
Same family or similar size — useful when shopping around.
Compact Gemma variant designed for on-device inference. Trained with knowledge distillation from larger Gemma 2 teachers. Runs comfortably on a phone at Q4.
- Context
- 8K
- License
- gemma
- VRAM Q4
- 1.6 GB
Pocket-sized Llama 3 variant for edge deployment. Surprising chat quality after instruction tuning makes it competitive with much larger models from a previous generation. At Q4 it fits in ~2 GB of VRAM and runs on consumer GPUs and recent Apple Silicon. A strong default for on-device chat, summarisation, and structured extraction tasks where the workload doesn't need frontier reasoning quality.
- Context
- 128K
- License
- llama-3
- VRAM Q4
- 1.8 GB
The original 7B RLHF chat model. Historically important — the first widely-adopted commercially-usable open-weights chat model. Still cited as a baseline in most 2024–25 papers.
- Context
- 4K
- License
- llama-2
- VRAM Q4
- 4.2 GB
The workhorse 8B instruction-tuned model. Excellent quality-to-cost ratio and the broadest ecosystem support of any open-weights model — every major inference engine, fine-tuning library, and quantization toolchain has a 3.1 8B preset. Fits in 24 GB of VRAM at fp16, ~6 GB at Q4. Strong default for production chat where 70B is overkill, for fine-tuning on a specialist task, and for any workload where you want a known-good baseline.
- Context
- 128K
- License
- llama-3
- VRAM Q4
- 4.8 GB
Llama 3's first vision-language model. Image understanding via a separately-trained ViT adapter bolted onto Llama 3 weights. Useful for OCR-adjacent workloads, document understanding, and image captioning at a permissive licence. The 11B size makes it cheap to host. Combined with the 128K text context, it handles long PDF-with-images workflows comfortably on a single 4090.
- Context
- 128K
- License
- llama-3
- VRAM Q4
- 6.6 GB
Mid-size Llama 2 chat model. Deprecated in most 2025 workloads by Llama 3.1 8B, but remains the baseline against which many post-2023 fine-tunes report.
- Context
- 4K
- License
- llama-2
- VRAM Q4
- 7.8 GB