Cerebras · OpenAI-Compatible Engines
GPT OSS 120B
Language
Tool use
OpenAI's open-weight 120B model on Cerebras wafer-scale hardware — ~3,000 tokens/second. Production tier.
Specifications
- Parameters
- 117B total · 5.1B active
- Architecture
- Open-weight MoE
- Token factory
- Cerebras (wafer-scale)
- Context window
- 128K tokens
- Max output
- 128K tokens
- Licence
- Apache-2.0
Best suited for
A capable all-rounder — the sensible default for everyday assistant and agent work.
General assistants
Drafting & summarising
Everyday agent tasks
Tool use & function calling
Pricing (per million tokens)
US$0.350 input / US$0.750 output
Price history (USD per 1M tokens)
▮ cheaper
▮ pricier
wick = input rate
Input rate has risen
+133.3%
since 5 Aug 2025.
Model id: gpt-oss-120b
Wafer-scale inference for open-weight models — the same OpenAI-compatible format as the rest of the gateway, at ~1,000–3,000 tokens/second.
View all model pricing →