Skip to content

Open-weight models served fast enough that latency stops being the thing you design around.

Lowest latencyCheap open weightsHigh throughput

4 models on Groq

Groq's own rate for each of these. Where another provider serves the same model it may charge differently — the model page shows every route side by side.

ModelPriceContext
Llama 3.1 8B Instructmeta/llama-3.1-8b-instruct$0.050 in$0.080 out131K
GPT-OSS 20Bopenai/gpt-oss-20b$0.075 in$0.300 out128K
GPT-OSS 120Bopenai/gpt-oss-120b$0.150 in$0.600 out131K
Llama 3.3 70B Instructmeta/llama-3.3-70b-instruct$0.590 in$0.790 out131K

Availability and latency

Multigrid does not measure Groq's availability, and will not publish numbers it did not measure.

For live status, read Groq’s own status page — it is the only authoritative source, and the one they update during an incident. Our own reachability probes against every provider are on the status page, and they answer “is anyone home”, not “is it fast”.

What Multigrid can tell you is what happened to your own requests: latency, errors and cost per provider are on your analytics page, measured from traffic you actually sent.

Reach Groq on one balance

We hold a platform key for Groq, so everything above is reachable on Multigrid credit — one balance across every vendor, at the prices listed. Signing up takes no card and commits you to nothing. If you already have a Groq contract, bring that key instead and we take no percentage on the traffic.

Groq — provider · Multigrid