Clockwork.io raised $31M to eliminate wasted GPU-hours through AI infrastructure fault tolerance software, with production adoption by LinkedIn and Together AI.
Clockwork.io has raised $31 million in new funding, bringing its total capital to $73 million, as AI infrastructure providers confront an expensive problem: thousands of GPUs can sit idle when a single component fails. The round was co-led by Premji Invest, Wing Ventures and Seligman Ventures, with participation from existing investors NEA and e& Capital.
The company is addressing a challenge that grows more costly as AI models and clusters scale. Meta previously reported unexpected interruptions roughly every three hours during a 54-day Llama 3 training run involving 16,384 GPUs. Conventional checkpoint-and-restart systems force an entire training job to restart from its last checkpoint when a GPU, network link, NIC or server fails, potentially losing hours of completed work and leaving healthy GPUs idle.
Clockwork.io was founded in 2018 from Stanford research by Balaji Prabhakar, Yilong Geng and Deepak Merugu, with Mendel Rosenblum, the VMware co-founder, serving as chief scientist. Prabhakar is a Stanford professor whose research focuses on computer networks and data-centre systems, while Geng's Stanford research produced the Huygens clock-synchronisation technology that became foundational to Clockwork's platform. Merugu previously co-founded Urban Engines, which was acquired by Google. The company is now led by Suresh Vasudevan, who joined as CEO after leading Sysdig and Nimble Storage. At Nimble, he took the company through an IPO and its eventual acquisition by HPE, and he held senior leadership roles at NetApp.
Clockwork's software sits between infrastructure and workloads to keep jobs running through failures. LinkPass reroutes network traffic around failed links, while TorchPass moves workloads from failing GPUs to healthy ones without requiring a complete restart. TorchSnap adds another layer by capturing the state of distributed AI jobs across multiple nodes—for training, its platform snapshots can preserve a running job without requiring changes to the training code. Its faster application checkpoints can also help reinforcement-learning systems move updated model weights to inference replicas more quickly.
LinkedIn has deployed Clockwork's LinkPass across its AI infrastructure and reports it prevents tens of thousands of GPU-hours of downtime each month. Together AI is bringing TorchPass to its GPU clusters, while WhiteFiber is expanding Clockwork across its global GPU-as-a-service infrastructure.
Gartner forecasts global AI spending will reach $2.7 trillion in 2026, up 49.5% year over year, with AI infrastructure representing the largest spending category. The boom has created a wave of heavily funded players: Together AI raised $800 million at an $8.3 billion valuation in its latest round, Crusoe recently raised more than $3 billion at a $30 billion valuation, and Nscale has raised $3.36 billion in pre-IPO financing. Clockwork is attacking a different layer of the same infrastructure stack—ensuring that the enormous amount of compute being deployed actually gets used.
"Failures are inevitable at AI scale. Losing hours of useful work to them should not be," said Suresh Vasudevan, CEO of Clockwork.io. "Fault tolerance is a goodput multiplier: it keeps GPUs doing useful work instead of waiting for recovery or repeating work already done. We built our software alongside enterprises and cloud providers operating some of the largest GPU fleets, so it handles the failures they actually see. That protection belongs in the infrastructure enterprises and cloud providers rely on every day."