We Signed a Multi-Year Agreement to Deploy Cerebras Wafer-Scale Inference
General Compute will bring Cerebras wafer-scale inference to its customers at scale. Capacity opens in Q1 2027, and agentic coding is the first workload. Here is what we are doing and why.
Today we are announcing a multi-year agreement with Cerebras Systems to deploy Cerebras wafer-scale inference at scale. General Compute will finance the systems and sell the result as inference. Capacity opens to customers in Q1 2027, and the first workload we are committing it to is agentic coding.
This is the largest single hardware commitment we have made, and it is the first deployment drawn against the $400M debt facility we secured with Upper90 in July. This post covers what we are deploying, why we picked agentic coding as the first workload, and why the financing structure matters as much as the silicon.
What we are deploying
Cerebras builds the largest chip in the world on purpose. The Wafer-Scale Engine spreads model weights across on-chip SRAM the size of an entire wafer, so decode never has to wait on an external memory bus. That is the reason Cerebras sits at or near the top of independent per-user speed leaderboards model after model. For a single user streaming tokens, there is nothing faster in production today.
We are adding that hardware to the General Compute fleet as the tier for work where per-token latency is the constraint that matters most.
The mechanics are the same as the rest of our fleet. We buy the systems, and customers get dedicated capacity under one contract with one set of SLAs. If you are already running with us, Cerebras shows up as a tier you can select against a latency target and a price.
Why agentic coding goes first
A chat product makes one inference call and shows you the answer. A coding agent makes hundreds or thousands of them in sequence: read the repository, plan, write a change, run the tests, read the failure, revise, and repeat until the task is done. Every call sits on the critical path. Latency does not average out across those steps. It adds up into wall-clock time the developer waits on.
The math is unforgiving. Take an agent that needs 400 model calls to finish a task and generates roughly 300 output tokens per call.
| Decode speed | Time per call | Time per task |
|---|---|---|
| 100 tok/s | 3.0 s | 20 min |
| 300 tok/s | 1.0 s | 6.7 min |
| 1,000 tok/s | 0.3 s | 2 min |
| 2,000 tok/s | 0.15 s | 1 min |
Same model, same agent, same task. The only variable is how fast the hardware decodes. At the bottom of that table the developer walks away and comes back. At the top they stay in the loop, and that changes how people use the tool. We wrote about this dynamic in more depth in how coding agents depend on inference speed and building a code agent: why each step needs sub-second inference.
The companies building coding agents and AI developer tools already know this. Most of them cannot put a wafer-scale system on their own balance sheet, and the ones who could still do not want to run one. That is exactly the gap we exist to close, so it is where we are starting.
Sean Lie, Cerebras CTO and co-founder, put it this way:
In AI, speed is productivity. An agent that takes hundreds of steps to finish a task is only as fast as its slowest step. Working with General Compute puts Cerebras speed in front of the developers building these agents, on a platform they already trust.
Why the financing matters
Chips do not win markets on architectural merit. They win on deployed capacity, and deployed capacity is a financing problem before it is an engineering problem.
The fastest-growing infrastructure companies in AI are, in practice, extensions of NVIDIA's balance sheet. They are funded with NVIDIA equity, collateralized by NVIDIA GPUs, and backstopped by NVIDIA purchase guarantees. Every alternative accelerator that beats a GPU on inference has faced the same problem: no equivalent channel existed to get it financed, deployed, and in front of customers at scale.
We built General Compute to be that channel. We buy the accelerators our customers want to run on but cannot or will not own, and we sell the output as inference. The Cerebras agreement is the proof that the model works at scale. Lenders will finance alternative accelerators. A neocloud can deploy them. Customers can consume the result without a capital commitment. That is the whole thesis, and you can read the long version in our whitepaper.
Here is how our co-founder and CEO Finn Puklowski framed it:
The chips that win inference are not going to come from one vendor, and most customers cannot put a wafer-scale system on their own balance sheet. That is the gap we exist to close. We buy the hardware, and our customers get Cerebras speed on a contract they can actually sign. Agentic coding is where that speed is worth the most right now, so that is where we are starting.
If you want to be on the first racks, reserve capacity and tell us about your workload. If you want to understand the broader picture of why we are betting on specialized silicon, the whitepaper walks through the full argument.