DeepInfra
The cheapest way to run large open-weight models at volume.
11 models on DeepInfra
DeepInfra's own rate for each of these. Where another provider serves the same model it may charge differently — the model page shows every route side by side.
| Model | Price | Context |
|---|---|---|
| BGE-M3baai/bge-m3 | $0.010 inFree out | 8K |
| Llama 3.1 8B Instructmeta/llama-3.1-8b-instruct | $0.020 in$0.040 out | 131K |
| Qwen3 32Bqwen/qwen3-32b | $0.080 in$0.280 out | 131K |
| Llama 3.3 70B Instructmeta/llama-3.3-70b-instruct | $0.100 in$0.320 out | 131K |
| DeepSeek V3.2deepseek/deepseek-v3.2 | $0.260 in$0.380 out | 164K |
| Qwen3 235B A22B Instructqwen/qwen3-235b-a22b-instruct | $0.090 in$0.550 out | 262K |
| DeepSeek V3.1deepseek/deepseek-v3.1 | $0.250 in$0.950 out | 164K |
| DeepSeek R1 (0528)deepseek/deepseek-r1 | $0.500 in$2.15 out | 164K |
| DeepSeek V4 Prodeepseek/deepseek-v4-pro | $1.30 in$2.60 out | 512K |
| Claude Haiku 4.5anthropic/claude-haiku-4-5 | $1.00 in$5.00 out | 200K |
| Claude Sonnet 5anthropic/claude-sonnet-5 | $2.00 in$10.00 out | 1M |
Availability and latency
Multigrid does not measure DeepInfra's availability, and will not publish numbers it did not measure.
For live status, read DeepInfra’s own status page — it is the only authoritative source, and the one they update during an incident. Our own reachability probes against every provider are on the status page, and they answer “is anyone home”, not “is it fast”.
What Multigrid can tell you is what happened to your own requests: latency, errors and cost per provider are on your analytics page, measured from traffic you actually sent.
Reach DeepInfra on one balance
We hold a platform key for DeepInfra, so everything above is reachable on Multigrid credit — one balance across every vendor, at the prices listed. Signing up takes no card and commits you to nothing. If you already have a DeepInfra contract, bring that key instead and we take no percentage on the traffic.