We signed a multi-year agreement to deploy Cerebras wafer-scale inference. Live Q1 2027.Read

Pricing

Priced per deployment, not per token.

We sell dedicated capacity on purpose-built inference silicon under one contract and one set of SLAs. Tell us the workload and we reply with a rate card and a deployment slot.

Workload shape
Prompt length, output length, streaming and concurrency all change the serving economics.
Silicon
Cerebras, SambaNova, Positron and d-Matrix sit at different points on the latency and cost curve. You pick against a target.
Model ownership
Hosted open models and private weights have different operational requirements.
Term
Reserved racks are priced on commitment length. Longer terms get lower rates.
ModeHumanAgent