Friday, August 7, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

GPU cloud provider cuts L40S prices in half, intensifying competition in inference-grade accelerator markets.

Signals pricing power erosion for older GPU generations, pressuring supply chain economics while reducing inference workload deployment costs.
Trade pressSlicast · August 16, 2024 · Global · Source: fly.io
importance 60

Fly.io, a public cloud provider with developer-friendly ergonomics, has lowered the prices on NVIDIA L40S GPUs to $1.25 per hour. The company offers four different NVIDIA GPU models in increasing order of performance: the A10, the L40S, the 40G PCI A100, and the 80G SXM A100. By a wide margin, the most popular GPU in their inventory is the A10, an older generation NVIDIA GPU with fewer, slower cores and less memory that is the least capable GPU they offer. However, it remains capable enough for random inference tasks and mid-sized generative AI work like Mistral Nemo or Stable Diffusion, such that there is not much benefit in getting a beefier GPU. As a result, Fly.io cannot get new A10s in fast enough for their users.

Looking back over four years of customer conversations, Fly.io has learned that what seemed like the obvious GPU strategy problems were not actually where customer demand lay. In 2023, the company thought the biggest problem to solve was selling fractional A100 slices and spent an entire quarter trying to get MIG or vGPUs working through IOMMU PCI passthrough on Fly Machines in a project so cursed that Thomas forswore ever programming again. They then went to market selling whole A100s, and for several more months believed the biggest problem was finding a secure way to expose NVLink-ganged A100 clusters to VMs for training work. Then came the H100s; could they find H100s anywhere, perhaps only in a black market in Shenzhen? A year later, looking at actual data, the least sexy, least interesting GPU in the catalog is where all the action is.

The data reveals that training workloads tend to look more like batch jobs, while inference tends to look more like transactions. Batch training jobs are not that sensitive to networking or reliability, but live inference jobs responding to end-user HTTP requests are highly sensitive. Given Fly.io's pricing, the A10s hit a sweet spot. The next step up in their lineup is the L40S, an AI-optimized version of the L40, which is the data center version of the GeForce RTX 4090. While the RTX 4090 is a gaming GPU used for ray-traced Witcher 3 and similar applications, NVIDIA's high-end gaming GPUs are reasonably good at AI workloads despite consuming excessive power and being hard to cool in a data center rack. The L40 offers much more memory, less energy consumption, and better rack density than gaming GPUs. A funny thing happened in the middle of 2023 when the market for ultra-high-end NVIDIA cards went absolutely batshit; the huge cards used for training jobs became impossible to find, and NVIDIA became one of the most valuable companies in the world. NVIDIA responded by launching the L40S, an L40 with AI workload compute performance comparable to that of the A100.

The L40S is positioned as an A100-performer that Fly.io can price for A10 customers—the Volkswagen GTI of their lineup. To make this happen, Fly.io is making it official: customers won't pay a dime extra to use an L40S instead of an A10, with pricing at $1.25 an hour. The company believes the combination of just-right-sized inference GPUs and Tigris object storage is particularly powerful, and customers should be able to use L40S cards without thinking hard about it. As Chicago's weather turns chilly in the coming month, Fly.io invites users to go light some cycles on fire.

Read the original
GPU cloud provider cuts L40S prices in half,… · Slicast