Clockwork.io, a Palo Alto-based company that builds software to keep large artificial intelligence workloads running when hardware fails, has raised $31 million in new funding, bringing its total capital raised to $73 million.

The round, announced on 5 October 2026, was co-led by Premji Invest, the investment firm backed by Wipro founder Azim Premji, together with Wing Venture Capital and Seligman Ventures. Existing investors NEA and e& Capital also participated.

The financing comes as companies operating large clusters of graphics processing units grapple with an increasingly expensive problem. As AI clusters grow to thousands or tens of thousands of GPUs, hardware and network failures become routine rather than exceptional, and every failure can halt or slow jobs that cost enormous sums to run.

The problem of wasted compute

Training a large AI model involves thousands of GPUs working in tight coordination over days or weeks. If a single GPU fails, or a network link between servers flaps, the entire job can stall. Traditionally, engineers respond by restarting the job from its last saved checkpoint, losing all progress made since then and leaving costly hardware idle while systems recover.

Clockwork's chief executive, Suresh Vasudevan, argues that this approach no longer makes sense at modern scale. In the company's announcement, he said failures are inevitable in AI infrastructure but that losing hours of useful work to them should not be. He described fault tolerance as a goodput multiplier, using the industry term for the share of computing time that produces useful results rather than being spent on recovery or repeated work.

The economics are stark. High-end GPUs rent for significant sums per hour, and large clusters can cost millions of dollars a month to operate. Even a modest improvement in the share of time that GPUs spend doing productive work can translate into substantial savings or faster model development.

How the software works

Clockwork's products sit as a software layer between the hardware and the AI workloads running on it. Its LinkPass technology provides network fault tolerance: when a network link, optical component, cable or network card fails, it automatically reroutes traffic onto healthy paths so that jobs continue uninterrupted while repairs are made.

Its TorchPass technology addresses GPU failures. Rather than forcing a training job to roll back to a checkpoint, it migrates work off a failing GPU onto a healthy one while the job is still running. A related capability takes snapshots of an entire running job without requiring changes to the underlying code, allowing faster recovery when problems do occur.

An example cited by one of the company's customers illustrates the stakes. Before adopting Clockwork's software, a single network card flap could remove an eight-GPU server from service, and a switch port issue could take down a second server, doubling the impact to 16 GPUs. With automatic rerouting, those jobs can continue running while the faulty components are fixed.

c3d5f971-d11e-481e-8870-35fabad0d214.png

Customers and partners

The company has secured production deployments at several notable organisations. LinkedIn has deployed LinkPass across its AI infrastructure fleet and says the software prevents tens of thousands of GPU-hours of downtime each month. Together AI, a cloud provider focused on AI workloads, is integrating TorchPass into its GPU Clusters offering, so its customers can access the technology as a service. WhiteFiber, which provides GPU capacity as a service, is expanding its use of Clockwork's software across its global infrastructure.

Clockwork also works with companies including Wells Fargo and Uber, and partners with AI cloud providers such as Nebius and Nscale. Denmark's national AI supercomputing centre has also cited the company's technology as helping it operate its Gefion system reliably. Together AI and Clockwork plan to demonstrate a live multi-node training job continuing through deliberately injected network and GPU failures at the PyTorch Conference.

Why investors are interested

The investment thesis is that, as the AI industry matures, attention is shifting from simply acquiring more computing capacity to using existing capacity more efficiently. Hyperscale cloud providers, specialist AI clouds and enterprises building their own GPU clusters all face the same pressure to extract more useful work from each expensive chip.

Clockwork's bet is that resilience will become as important as raw GPU performance. Faster chips matter, but if large clusters lose a meaningful fraction of their time to failures and recovery, improving reliability can deliver gains comparable to a hardware upgrade at a fraction of the cost.

Greg Papadopoulos, a venture partner at NEA, said the firm first backed Clockwork in 2021 and that the company's technology is already saving customers tens of thousands of GPU-hours a month. The company plans to use the new capital to accelerate the rollout of its fault-tolerance suite across training, inference and reinforcement learning, broaden enterprise adoption and scale distribution through cloud partners.

The Indian connection

The round has a notable Indian dimension. Premji Invest, one of the co-leads, manages capital for the family of Azim Premji and has become an active investor in technology companies in India and the United States. Vasudevan, who leads Clockwork, is a veteran Silicon Valley executive who previously served as chief executive of storage company Nimble Storage and later of cloud security company Sysdig. Indian capital backing a Silicon Valley infrastructure company of this kind reflects how closely the two technology ecosystems are now intertwined.

For India, which is building its own national AI computing capacity through government-backed programmes and private data centre investment, the problems Clockwork is solving are directly relevant. As Indian cloud providers and research institutions deploy larger GPU clusters, the efficiency of those clusters will matter as much as their size.

The bigger picture

The AI infrastructure market has been dominated by headlines about chip shortages, multi-billion-dollar data centre projects and soaring capital expenditure. Clockwork's funding is a reminder that a significant opportunity also lies in the less visible layer of software that keeps that infrastructure running efficiently.

As AI clusters continue to grow in size and cost, tolerance for wasted compute will shrink. Companies that can measurably increase goodput are likely to find willing customers among cloud providers and enterprises alike. With $73 million raised and deployments at some of the world's most demanding AI operators, Clockwork is positioning itself to become a standard part of that infrastructure stack. Its next challenge is to prove that its technology can scale as quickly as the clusters it is designed to protect.