Delos Data has raised more than $100 million to launch Delos Nonstop AI, a software, systems and silicon approach aimed at a growing problem in AI infrastructure: moving data efficiently between heterogeneous compute resources. The company says its new architecture is designed for persistent agentic AI inference, combining a Nonstop AI Data Interface, clusters and servers to keep workloads running across GPUs, XPUs, CPUs, memory and storage.
The next constraint in scaling AI may not be another generation of GPUs. It may be the network connecting them.
That is the thesis behind Delos Nonstop AI, a new AI infrastructure architecture launched by Delos Data alongside more than $100 million in new funding from Matrix, Playground, Socratic Partners, Capricorn’s Technology Impact Fund, Matter Venture Partners, IAG and industry investors.
The Palo Alto-based company says the capital will support additional software and hardware engineering, product development and sales as it targets the infrastructure requirements of large-scale AI inference.
Delos argues that the architecture of modern AI clusters was largely designed around workloads that do not resemble emerging agentic inference. Traditional inference can often be treated as a sequence of requests and responses. Agentic systems can execute longer, multistep workflows, repeatedly invoking models and moving information among different types of compute, memory and storage.
That changes the networking problem.
Instead of optimizing only for peak accelerator performance, infrastructure operators increasingly need to keep heterogeneous resources continuously supplied with data while maintaining performance when individual components or connections fail.
Delos calls this a data movement bottleneck and has built Nonstop AI around the idea that networking should become a first-class component of the AI architecture rather than an interconnect added around compute.
From GPU clusters to heterogeneous AI infrastructure
The company’s new Nonstop AI Reference Architecture is designed to combine different hardware, models, switches, links and network topologies within a common data domain.
The approach is aimed at what Delos describes as a “mixture of X” infrastructure, or MoXI: clusters increasingly composed of GPUs, XPUs, accelerators, CPUs, memory and storage rather than a homogeneous collection of identical processors.
That heterogeneity is becoming more important as AI infrastructure expands beyond conventional GPU clusters. Hyperscalers and AI infrastructure providers are deploying multiple generations of accelerators and specialized silicon, while enterprises increasingly use combinations of foundation models, smaller models and domain-specific inference systems.
The new Delos Nonstop AI Data Interface is the company’s hardware and systems answer to that problem. Delos claims the interface can deliver 10x lower latency and 10x higher efficiency, although those figures are company-reported performance claims rather than independent benchmarks. The interface is available in three proposed form factors: an I/O chiplet, near-packaged optics and a card.
It joins Delos Nonstop AI Clusters and the Nonstop AI Server, giving the company a portfolio spanning infrastructure software, systems and data-movement hardware.
The architecture is intended to allow customers to choose their own combination of hardware and models while treating the resulting environment as a unified domain.
Why inference changes the infrastructure equation
The timing of the launch reflects a broader change in AI workloads.
Gartner forecasts that worldwide spending on AI-optimized infrastructure will reach approximately $42.3 billion in 2026, up 96.4% from 2025. More significantly for companies such as Delos, Gartner expects global spending on inference to reach $23.3 billion in 2026, exceeding the $19 billion projected for training.
McKinsey similarly projects that AI inference will become the dominant AI workload in data centers by 2030, accounting for more than half of AI compute and potentially more than 40% of overall data-center demand. Its model projects inference demand growing at a 35% compound annual growth rate between 2025 and 2030.
That transition changes the economics of infrastructure.
Training can involve enormous but relatively discrete compute runs. Inference for production AI applications is continuous. Agentic applications can make multiple model calls to complete one task, increasing the amount of data moving through the system and placing sustained pressure on compute, memory and networking.
Gartner estimates that agentic models can require five to 30 times more tokens per task than a standard generative AI chatbot. It also forecasts that the cost of inference per agentic workflow will increase more than fivefold through 2028 as increasingly sophisticated workflows consume more tokens.
That creates an infrastructure paradox: individual tokens may become cheaper to process, but the amount of inference required by increasingly capable AI systems can rise faster.
Resilience becomes part of AI performance
Delos is also emphasizing resilience.
Agentic workloads are expected to operate continuously across large clusters, meaning infrastructure failures can translate directly into interrupted inference and idle accelerator capacity. If a GPU is available but waiting for data, the capital invested in that GPU is not producing useful output.
Delos Nonstop AI is designed around keeping workloads running despite failures by creating a more resilient data-movement layer. The company says its architecture can combine heterogeneous endpoints and maintain the workload through infrastructure failures.
That is a different proposition from simply building a faster network. It treats availability, scale and data movement as interconnected parts of inference performance.
For AI infrastructure operators, the objective is increasingly to maximize tokens per dollar, rather than simply tokens per accelerator.
Competing for the AI infrastructure layer
Delos enters a market where the major technology companies are simultaneously investing in compute, networking and custom silicon.
NVIDIA has built networking into its accelerated-computing platform through technologies including InfiniBand and Ethernet, while hyperscalers such as Google, Amazon and Microsoft have developed custom accelerators and networking architectures for their AI workloads.
Specialist infrastructure companies are targeting different layers of the same problem. Some focus on optical connectivity, others on Ethernet fabrics, switches, accelerators, memory or distributed computing software.
Delos is attempting to differentiate by spanning the data path from architecture and software to systems and interface silicon.
Its investors include former and current technology executives with experience across processors, networking silicon, data-center systems and infrastructure software. Playground Global General Partner Pat Gelsinger and other investors have described the network as an increasingly important determinant of inference performance, but those statements represent investor views rather than independent validation of Delos’s technology.
The company is targeting frontier AI labs, hyperscalers, neoclouds and sovereign AI programs—markets where infrastructure utilization and failure recovery can have significant economic consequences.
The larger question is whether AI networking architectures can evolve quickly enough as inference becomes more distributed and agentic systems generate increasingly complex workloads.
Delos is betting that the answer requires redesigning the data path itself.
Its Nonstop AI architecture therefore represents a broader trend in AI infrastructure: as models become more capable, performance increasingly depends not just on the silicon doing the computation, but on how efficiently and reliably information moves between every component involved in producing the result.
Market Landscape
AI infrastructure is moving from a compute-centric market toward a broader systems race involving accelerators, networking, memory, storage, power and orchestration.
Gartner expects AI-optimized IaaS spending to reach $42.3 billion in 2026, while its forecast puts inference spending above training spending this year. Gartner also says AI infrastructure demand will continue to exceed supply through at least 2030, highlighting the pressure on infrastructure providers to improve utilization rather than simply add capacity.
Meanwhile, Gartner predicts that neocloud providers could capture 20% of a $267 billion AI cloud market by 2030, indicating a growing market for specialized infrastructure outside the largest hyperscalers.
Delos is targeting this infrastructure layer by combining software, servers, clusters and networking silicon around persistent agentic inference. Its central proposition is that maximizing AI capacity increasingly requires optimizing the movement of data between compute resources, not only increasing the number or performance of those resources.
Top Insights
- Delos Data raised more than $100 million to develop Nonstop AI, targeting networking and data movement as an emerging AI inference bottleneck.
- Its Nonstop AI portfolio combines reference architecture, data interfaces, servers, clusters and software for heterogeneous AI infrastructure.
- Delos claims its new Data Interface delivers 10x lower latency and 10x higher efficiency, subject to independent benchmarking.
- Gartner forecasts AI inference spending will exceed training spending in 2026 as enterprises operationalize increasingly persistent AI workloads.
- McKinsey projects inference could account for more than 40% of total data-center demand by 2030.
Power Tomorrow’s Intelligence — Build It with TechEdgeAI












