We secured a $400M debt facility with Upper90 to scale inference compute.Read

General Compute

Blog

Insights on AI inference, ASIC infrastructure, and building fast AI applications.

114 posts across 5 topics

Latest

Browse by topic

Why inference speed is the new moat, real-time AI guides, and benchmarks comparing latency, throughput, and cost.

Technical deep-dives on the building blocks of modern LLM inference: attention, quantization, decoding, and architectures.

Tool calling, multi-agent architectures, reasoning patterns, and the inference requirements behind production agents.

Speculative decoding, KV cache, tensor parallelism, batching strategies, and the systems that serve LLMs at scale.

Guides and benchmarks for the latest open-source and proprietary models, with practical tips for running them in production.

ModeHumanAgent