← Model Library

Cerebras · OpenAI-Compatible Engines

GPT OSS 120B

Language Tool use

OpenAI's open-weight 120B model on Cerebras wafer-scale hardware — ~3,000 tokens/second. Production tier.

Specifications

Parameters
117B total · 5.1B active
Architecture
Open-weight MoE
Token factory
Cerebras (wafer-scale)
Context window
128K tokens
Max output
128K tokens
Licence
Apache-2.0

Best suited for

A capable all-rounder — the sensible default for everyday assistant and agent work.

General assistants Drafting & summarising Everyday agent tasks Tool use & function calling

Pricing (per million tokens)

US$0.350 input / US$0.750 output

Price history (USD per 1M tokens)

cheaper pricier wick = input rate
$0 $0.4 $0.8 5 Aug 2025: $0.15 input / $0.6 output 1 Feb 2026: $0.037 input / $0.1 output 4 Jun 2026: $0.3 input / $0.7 output 4 Aug 2026: $0.35 input / $0.75 output 5 Aug 2025 4 Aug 2026

Input rate has risen +133.3% since 5 Aug 2025.

Model id: gpt-oss-120b

Wafer-scale inference for open-weight models — the same OpenAI-compatible format as the rest of the gateway, at ~1,000–3,000 tokens/second.

View all model pricing →

One Gateway, Every Model

Get an API key and point your first request at ApiSpi in minutes