Clockwork.io Raises $31M to Make AI Infrastructure Resilient

Clockwork.io Raises $31M for AI Fault Tolerance Clockwork.io Raises $31M for AI Fault Tolerance

Clockwork.io has raised $31 million to expand software designed to keep large-scale AI training, inference and reinforcement-learning workloads running when GPUs, network links or servers fail. The company is also adding new checkpointing capabilities to its TorchPass platform and says deployments at LinkedIn, Together AI and WhiteFiber are expanding as AI infrastructure operators focus increasingly on GPU utilization and workload “goodput.”

AI infrastructure is moving beyond simply adding more GPUs

The economics of large-scale AI increasingly depend on how much useful computation a GPU fleet can deliver—not simply how many accelerators an operator can deploy. As training clusters grow into thousands of GPUs, individual hardware and networking failures become increasingly difficult to treat as exceptional events.

Clockwork.io is positioning fault tolerance as an infrastructure layer for that problem. The company announced $31 million in new funding alongside expanded customer deployments and two additions to its TorchPass workload-resilience platform.

The company’s software is designed to keep distributed AI jobs operating through infrastructure failures rather than forcing workloads to restart from an earlier checkpoint.

That distinction matters because conventional recovery can discard completed computation. A distributed training job may have to reload a previously saved state, leaving healthy GPUs idle while the failed portion of the workload is restored. The larger the cluster and the longer the interval between checkpoints, the more compute can be lost.

Clockwork.io cites Meta’s experience training Llama 3 as an indication of the problem. Meta reported unexpected interruptions roughly once every three hours during a 54-day training period involving 16,384 GPUs.

The company’s approach divides fault tolerance across several layers. LinkPass reroutes traffic around failed network paths, while TorchPass can move work away from a failing GPU to another available GPU. Both capabilities are already in production, according to Clockwork.io.

The latest TorchPass release adds multi-node platform snapshots and fast asynchronous application checkpoints.

Multi-node snapshots capture the execution state of an entire distributed AI job across its nodes. Clockwork.io says the capability is designed to work without requiring changes to training code, giving infrastructure teams a way to protect workloads even when they do not control the underlying application.

That could be particularly relevant for cloud and GPU-as-a-service providers. Platform operators increasingly run workloads belonging to customers, making application-level modifications impractical. A platform-level recovery mechanism can instead sit below the workload.

The second capability targets reinforcement learning and reduces the time required to move updated model weights from a training process to inference replicas.

In reinforcement learning systems, inference replicas generate rollouts that become training data. The trainer then updates the model, with those new weights needing to reach the replicas so they can generate subsequent rollouts using the latest model. Delays in that exchange can leave inference workers operating on stale weights.

Clockwork.io says its asynchronous checkpoints transfer updated weights in the background while the workload continues running.

LinkedIn and GPU cloud providers push resilience further down the stack

Customer deployments illustrate why fault tolerance is becoming a concern for both enterprise AI teams and infrastructure providers.

LinkedIn has deployed Clockwork.io’s LinkPass across its AI infrastructure fleet. The company says the technology prevents tens of thousands of GPU-hours of downtime each month by allowing workloads to continue while network components such as NICs, optics, cables and links are repaired.

Together AI is taking a different route by bringing TorchPass to customers through its GPU Clusters service. The companies plan to demonstrate distributed training continuing through deliberately injected GPU and network failures without restarting the job.

For GPU cloud providers, the objective is closely tied to utilization. Every hour that a customer’s workload spends recovering rather than computing represents capacity that generates less value.

WhiteFiber is also expanding its deployment across its GPU-as-a-service infrastructure. The company uses Clockwork.io’s fleet auditing capabilities to identify faulty or degraded links, NICs and optical components before clusters enter production.

The broader market is therefore shifting from basic infrastructure availability toward AI workload resilience. A cluster can remain technically online while still losing substantial productive capacity because individual failures interrupt distributed jobs.

That makes “goodput”—the amount of GPU time that actually advances a workload—an increasingly important metric alongside conventional uptime.

Fault tolerance becomes part of the AI platform stack

Clockwork.io’s expansion reflects a broader evolution in AI infrastructure. Training, inference and reinforcement learning are becoming increasingly interconnected, while model sizes and distributed deployments continue to increase the operational cost of failures.

The competitive landscape includes GPU infrastructure providers, cloud platforms and software companies building orchestration, observability and workload-management layers around accelerators from NVIDIA and other vendors. As AI workloads become more distributed, resilience can become a differentiating feature for those platforms.

The challenge is not eliminating hardware failures. It is preventing an individual failure from becoming a workload-level outage.

Clockwork.io’s strategy is to make that protection part of the infrastructure layer, allowing platform teams to recover or route around failures without relying entirely on application developers.

If AI infrastructure continues toward larger clusters and increasingly autonomous workloads, the ability to preserve computation may become as important as raw accelerator availability.

Market Landscape

AI infrastructure is evolving from a capacity race into an efficiency and resilience race. Cloud and neocloud providers are competing to offer scarce GPU capacity, but customers increasingly care about how much purchased compute actually advances their models.

Fault tolerance sits at the intersection of AI infrastructure, machine learning infrastructure, AI cloud platforms and AI applications. Network resilience protects distributed inference and training, GPU migration limits lost computation, and faster checkpointing reduces the amount of work that must be repeated.

The trend also aligns with the growth of reinforcement learning and agentic AI. As AI agents increasingly depend on continuous inference, tool use and iterative model improvement, infrastructure interruptions can affect not only training duration but also the availability of production AI services.

Top Insights

  • Clockwork.io is adding platform-level snapshots and asynchronous checkpoints to protect distributed AI workloads from costly infrastructure failures.
  • LinkedIn says Clockwork.io prevents tens of thousands of GPU-hours of downtime each month across its AI infrastructure fleet.
  • Together AI plans to offer TorchPass through its GPU Clusters service, extending fault tolerance into the GPU cloud market.
  • Faster checkpoints are particularly relevant to reinforcement learning, where stale model weights can slow the training-inference feedback loop.
  • AI infrastructure providers are increasingly measuring goodput, not just uptime, as GPU fleets become larger and more expensive.

Power Tomorrow’s Intelligence — Build It with TechEdgeAI

Grow Your
Brand Visibility

Looking to publish a press release, guest article, interview or podcast? Connect with us.

GET FEATURED
Subscribe

Sign up today for exclusive insights and updates.

Newsletter Signup