SemiAnalysis podcast episode discusses competitive landscape of AI inference chips, including Google TPU v7, Vera Rubin architecture, and vendors Groq, Cerebras, and SambaNova.
Open-source inference serving is currently extraordinarily profitable, according to InferenceX. By their own profit estimator, serving Kimi K3 on a GB300 NVL72 rack at 60% utilization—after accounting for Moonshot's 30% license fee—yields roughly a 40–50% profit margin. The team argues that even 30–40% utilization rates remain profitable. This assessment rests on the AgentX benchmark, which replays real agentic coding sessions captured from SemiAnalysis employees' own infrastructure, where the theoretical cache hit rate exceeds 98% and DRAM offloading recovers 20–30% of that at high-throughput operating points.
**N-gram layers and memory-bandwidth constraints**
N-gram layers represent a co-design response to HBM scarcity. Rather than having every layer of a transformer build up the meaning of a token sequence, 2-gram and 3-gram token embeddings are learned at training time, stored in a hash table, and looked up at inference. This mechanism effectively creates layers "for free"—the model stops spending parameters memorizing word meaning and moves toward next-token prediction earlier in the stack. DeepSeek's ablation shows loss is minimized at a middle point with some n-gram layers and some MoE layers rather than at either extreme, and the KL divergence between early and late hidden states shrinks with n-grams, indicating the model grounds meaning sooner.
At inference, n-gram inputs are just token IDs, enabling asynchronous retrieval overlapped with other computation. This creates a placement trade-off: n-gram layers help most early in the model, but placing them at layer one leaves nothing to overlap the retrieval against. Qwen 3.8 Flash Next placed engrams at layer two to give the DRAM or SSD retriever time to return, while V4.1 Flash placed them at layer one, leaving less overlap room. Offloading uses Unified Virtual Addressing to move n-grams into CPU-pinned host memory. The architectural divergence between American and Chinese labs reflects underlying hardware constraints: American labs optimize for HBM bandwidth over capacity because the constraint has not been forced on them, while Chinese labs have, particularly given Ascend 950's domestic high-bandwidth memory (lower bandwidth than export-available HBM). The LongCat team at Meituan claims to have been working on n-grams concurrently, suggesting this technique represents a convergent response to capacity limits.
**AgentX benchmark and KV cache offloading**
AgentX replaces earlier chip-level benchmarks with a system-level replay of agentic coding traffic. The workload shape follows this pattern: an agentic harness sends a message, the model calls tools, tool output is appended, and the growing message stream is resent each turn because the model is stateless. With million-token context models, KV cache storage becomes severe. The dataset is captured by a proxy that intercepts SemiAnalysis employees' agentic coding sessions, post-processed to remove bad requests, then replayed against real servers running open-source inference. The theoretical cache hit rate in the replayed traces exceeds 98%.
On B300 vLLM, offloading becomes highly relevant in the 40–60 tokens-per-second range at very high throughput—offline batch inference or reinforcement learning rollouts with hundreds of concurrent agentic sessions—where DRAM offloading recovers 20–30% additional cache hit rate over HBM alone. The stack uses Nixl for KV transfer, vLLM for KV offloading to CPU, DRAM-only offloading, and the Dynamo router. The P50 request in the dataset contains nearly 100,000 input tokens.
**Profit margins and TCO modeling**
InferenceX's TCO model breaks down hourly accelerator cost—datacenter, electricity, chip—and combines it with OpenRouter pricing data and InferenceX throughput and concurrency numbers to estimate profit. On a GB300 NVL72 with optimal serving configuration, Kimi K3 yields approximately a 40.70% profit margin, while DeepSeek V4.1 Flash shows configurations where no profit is made. At 60% utilization with the 30% Moonshot license fee, margins reach approximately 50%, with profitability sustained even at 30–40% utilization. Inference providers with private configurations should exceed these open-source-based figures. As compute pricing adjusts, these margins are expected to normalize.
**TPU v7 cost competitiveness**
TPU v7, which Google is now making purchasable through its own cloud, sits at the bottom of the cost-per-token Pareto curve against B200 and B300, with an external TCO of approximately $1.21 per million tokens. The recent TPU mega-kernel work produced a large performance jump, with further improvements expected. The structural argument concerns externalization: as Google shifts from allocating all TPUs to internal workloads toward selling a large share externally, this drives open-source ecosystem improvements, including a native torch backend for TPU replacing the current chain of intermediate representations. Historically, TPUs' fixed-size systolic arrays forced padding and waste on smaller models, and because TPUs were primarily internal, many models were built around this limitation while open source was not. TPUs being installed in neoclouds—hosting outside GCP for the first time—represents a likely price-reducing development.
**Vera Rubin performance**
On DeepSeek V4 Pro, Vera Rubin delivers approximately 3x better performance in the interactivity ranges providers actually serve, translating to 3x more concurrent users at the same latency and corresponding profit increases reflected in the profit estimator. The generation prior (HGX H100/H200 to GB200/GB300) introduced the ARM CPU, scale-up NVL switches, new 400G/800G networking, and the SM100 revision with new FP4 data types. Vera Rubin is more incremental physically and in software support. The 8-high configuration and reduced VRAM reflect a priority for memory bandwidth over per-GPU capacity, as the scale-up network allows spreading the model across more GPUs to use additional memory bandwidth. The genuinely new item is LUTB, a hardware-accelerated lookup-table quantization path confirmed in PTX and open-source Triton branches, using 3-bit keys to look up 8-bit table values.