Cost calculator
Estimate monthly hosted-inference cost across every model + provider we track. Numbers assume the pricing on our verified pricing table — always confirm on the provider's own page before committing.
| Provider | Input $/M | Output $/M | ||
|---|---|---|---|---|
| Mistral Nemo 12B | deepinfra | $0.019 | $0.030 | $15.90 |
| Mistral Nemo 12B | openrouter | $0.019 | $0.030 | $15.90 |
| Mistral Small 3 | openrouter | $0.050 | $0.080 | $42.00 |
| Phi-4 14B | openrouter | $0.070 | $0.140 | $63.00 |
| Qwen 3 32B | deepinfra | $0.080 | $0.280 | $90.00 |
| Qwen 3 32B | openrouter | $0.080 | $0.280 | $90.00 |
| Qwen2.5 7B Instruct | openrouter | $0.100 | $0.200 | $90.00 |
| Llama 4 Scout 17B (16E) | openrouter | $0.100 | $0.300 | $105.00 |
| Llama 3.3 70B Instruct | openrouter | $0.100 | $0.320 | $108.00 |
| Qwen 3 8B | openrouter | $0.117 | $0.455 | $138.45 |
| Llama 4 Maverick 17B (128E) | openrouter | $0.188 | $0.652 | $210.37 |
| Qwen2.5 72B Instruct | openrouter | $0.360 | $0.400 | $276.00 |
| DeepSeek V3 | deepinfra | $0.320 | $0.890 | $325.50 |
| DeepSeek V3 | openrouter | $0.320 | $0.890 | $325.50 |
| Gemma 2 27B | openrouter | $0.650 | $0.650 | $487.50 |
| Hermes 3 Llama 3.1 70B | openrouter | $0.700 | $0.700 | $525.00 |
| Qwen 3 235B (A22B) | openrouter | $0.455 | $1.820 | $546.00 |
| Qwen2.5 Coder 32B | openrouter | $0.660 | $1.000 | $546.00 |
| DeepSeek R1 Distill Llama 70B | openrouter | $0.800 | $0.800 | $600.00 |
| Kimi K2 Instruct | openrouter | $0.570 | $2.300 | $687.00 |
| Llama 3.3 70B Instruct | together | $1.040 | $1.040 | $780.00 |
| DeepSeek R1 | openrouter | $0.700 | $2.500 | $795.00 |
| Mixtral 8×22B Instruct | openrouter | $2.000 | $6.000 | $2100.00 |
Self-hosting comparison: at typical utilisation, a single RTX 4090 (~$400/mo amortised) breaks even against hosted 7B pricing around 2B tokens/month, and against hosted 70B pricing around 200M tokens/month. See /hardware.