Pricing
Priced per deployment, not per token.
We sell dedicated capacity on purpose-built inference silicon under one contract and one set of SLAs. Tell us the workload and we reply with a rate card and a deployment slot.
- Workload shape
- Prompt length, output length, streaming and concurrency all change the serving economics.
- Silicon
- Cerebras, SambaNova, Positron and d-Matrix sit at different points on the latency and cost curve. You pick against a target.
- Model ownership
- Hosted open models and private weights have different operational requirements.
- Term
- Reserved racks are priced on commitment length. Longer terms get lower rates.