Mid-tier Gemma. Strong general-purpose chat model at small scale. The Gemma Terms of Use permit commercial use subject to Google's prohibited-use policy.
- Parameters
- 9B
- Context length
- 8K
- Modality
- text
- Released
- 2024-06-27
Memory & hardware
- VRAM (fp16)
- 18 GB
- VRAM (Q4)
- 5.4 GB
- Recommended
- RTX 3090 24GB
- Quantizations
- fp16, q8_0, q5_k_m, q4_k_m
Benchmarks
Hosted inference pricing
No provider we track publishes a per-token price for this model today. What each one used to offer is listed below.
No longer listed
Providers that used to serve this model. We don't republish their old rates — the dates below are the provider's own.
- groqGroq shut down gemma2-9b-it on 8 October 2025.Source ↗
Run it yourself
Drop-in commands for the three most common open-source inference paths. The Ollama tag is a best-effort match against the registry; verify the size variant before pulling.
ollama run gemma2:9b
vllm serve google/gemma-2-9b-it
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("google/gemma-2-9b-it")
model = AutoModelForCausalLM.from_pretrained(
"google/gemma-2-9b-it", device_map="auto", torch_dtype="auto"
)google/gemma-2-9b-it Related models
Same family or similar size — useful when shopping around.
The workhorse 8B instruction-tuned model. Excellent quality-to-cost ratio and the broadest ecosystem support of any open-weights model — every major inference engine, fine-tuning library, and quantization toolchain has a 3.1 8B preset. Fits in 24 GB of VRAM at fp16, ~6 GB at Q4. Strong default for production chat where 70B is overkill, for fine-tuning on a specialist task, and for any workload where you want a known-good baseline.
- Context
- 128K
- License
- llama-3
- VRAM Q4
- 4.8 GB
The April 2025 refresh of Qwen at 8B. Native mixed-mode reasoning: the model can 'think' before answering when triggered, or answer directly for simple queries — configurable per request. Apache 2.0. A strong upgrade over Qwen 2.5 7B on math and code, with much better instruction following.
- Context
- 33K
- License
- apache-2-0
- VRAM Q4
- 4.8 GB
NousResearch's community-driven fine-tune on the Llama 3.1 8B base. Tuned for strong tool use, function calling and steerable persona behaviour. Inherits Llama 3's community licence and its 128K context.
- Context
- 128K
- License
- llama-3
- VRAM Q4
- 4.8 GB
Llama 3's first vision-language model. Image understanding via a separately-trained ViT adapter bolted onto Llama 3 weights. Useful for OCR-adjacent workloads, document understanding, and image captioning at a permissive licence. The 11B size makes it cheap to host. Combined with the 128K text context, it handles long PDF-with-images workflows comfortably on a single 4090.
- Context
- 128K
- License
- llama-3
- VRAM Q4
- 6.6 GB
The original 7B RLHF chat model. Historically important — the first widely-adopted commercially-usable open-weights chat model. Still cited as a baseline in most 2024–25 papers.
- Context
- 4K
- License
- llama-2
- VRAM Q4
- 4.2 GB
TII's latest dense 7B from December 2024. Strong scores on commonsense reasoning benchmarks. TII's Falcon licence permits royalty-free commercial use with attribution.
- Context
- 33K
- License
- falcon-2
- VRAM Q4
- 4.2 GB