Snorkel AI has raised $350 million at a $3.5 billion valuation to expand its agentic data factory, which develops specialized datasets, environments and evaluation infrastructure for frontier AI systems. The Series E round, led by Insight Partners and S32, comes as AI developers increasingly require expert-built data and testing environments rather than conventional annotation at scale.
AI data becomes a research problem
The next phase of AI development is putting pressure on one of the industry’s least visible infrastructure layers: the data used to train, evaluate and improve increasingly capable models and agents.
Snorkel AI is positioning itself around that problem. The company announced a $350 million Series E financing on September 22, bringing its valuation to $3.5 billion. Insight Partners and S32 co-led the round, with significant participation from existing investor Addition and new and existing investors including March Capital, Blumberg Capital, Allegis Capital, Frontline, Standard, Third Point Ventures, Greylock, Lightspeed and GV.
Snorkel says the funding will be used to expand its “agentic data factory,” increase capacity, invest in vertical and enterprise AI, and extend its research and technology into additional domains and modalities.
The company argues that AI data is entering a different phase. Early machine-learning workflows often relied on large volumes of relatively straightforward human labeling. Frontier models and autonomous AI agents require a more complicated combination of expert tasks, simulated or controlled environments, evaluation rubrics and domain-specific examples.
That shift could make data development a more technical part of the AI infrastructure stack.
From labeling to agentic environments
Snorkel describes the transition as moving from “Data 1.0” to “Data 2.0.”
The distinction is less about replacing annotation than expanding what qualifies as useful AI development data. An agent working on software engineering, legal analysis or healthcare tasks may need to interact with tools, follow multi-step procedures and produce outputs that can be assessed against detailed criteria.
That creates a different engineering problem from labeling images or categorizing text.
A useful training or evaluation environment may need realistic tasks, domain expertise, tool access, ground-truth information and a rubric capable of distinguishing partially correct behavior from successful completion.
Snorkel’s platform is designed around that type of data development. The company grew from research at the Stanford AI Lab and has focused on data-centric AI for roughly a decade. Snorkel says its research has produced more than 250 peer-reviewed papers with more than 25,000 citations.
Its newer Data-as-a-Service offering, launched in September 2025, extends that research orientation into commercial work with AI labs and enterprises.
Agentic AI creates a new evaluation layer
The development of autonomous AI agents makes evaluation more complicated because an agent’s behavior can change depending on context, available tools, data and the path it takes to complete a task.
Gartner’s 2026 research explicitly distinguishes agent evaluation from traditional output-focused evaluation, noting that the nondeterministic nature of agentic workflows requires organizations to evaluate agent behavior and performance across workflows.
Gartner has also identified simulation as an emerging capability for designing, evaluating and continuously improving agentic systems. Its research argues that simulated testing environments can help reduce deployment friction and support iterative improvement.
That makes environments themselves part of AI infrastructure.
Instead of simply asking whether a model generates a correct answer, developers increasingly need to test whether an agent can select the right tool, follow a process, recover from an error, respect constraints and reach an intended outcome.
Snorkel’s focus on expert-built tasks, environments and rubrics sits directly within this emerging layer.
Data quality becomes an AI systems issue
The development of agentic AI is also changing how enterprises think about data readiness.
Gartner published research in August 2026 arguing that AI-ready data practices need to expand toward “agent-ready” data, with organizations evaluating whether data is suitable for agents and their use of enterprise information.
McKinsey’s June 2026 research similarly identifies data readiness as a constraint on scaling enterprise AI. The firm argues that organizations need governed, reusable data foundations spanning structured and unstructured information.
For frontier AI labs, the problem extends beyond internal enterprise data. They need high-quality examples of complex tasks that can teach or measure increasingly capable models.
This is where specialized data providers can become part of the model-development pipeline rather than simply outsourced annotation vendors.
Snorkel targets high-complexity AI work
Snorkel says its customers include frontier AI labs, hyperscalers, AI companies, enterprises and U.S. government agencies. The company plans to increase its capacity for complex data development while expanding into areas including healthcare, law and software engineering.
These verticals are particularly demanding because the quality of an AI system cannot always be measured through simple factual accuracy.
A healthcare agent, for example, may need to reason through clinical information while following defined safety constraints. A software engineering agent may need to modify code, execute tests and recover from failures. A legal system may need to interpret documents while maintaining traceability to source material.
Each scenario requires a different definition of successful behavior.
Snorkel’s model therefore points toward a more specialized AI development ecosystem in which data scientists, domain experts, engineers and AI researchers jointly construct the environments used to develop and evaluate models.
The economics of frontier AI data
The company’s financing also reflects the growing commercial value assigned to AI development infrastructure.
Snorkel says its Data-as-a-Service business has grown more than 18-fold since launching in September 2025 and that it reached a $375 million annualized revenue run rate in September 2026. Those figures are company-reported and have not been independently verified.
The broader AI market is also moving toward larger-scale agent deployment. McKinsey’s 2026 global survey found that 40% of respondents at organizations with more than $1 billion in annual revenue reported scaling AI agents, up from 27% the previous year.
As deployment expands, the supporting infrastructure must address more than model inference. Data pipelines, evaluation systems, agent environments, governance and observability increasingly become part of the production architecture.
Building an infrastructure layer for AI labs
Snorkel’s strategy is consequently broader than supplying training examples.
The company is attempting to establish an infrastructure and research layer around the development of expert-grade data for frontier and agentic AI. Its Open Benchmarks Grants initiative is another component of that strategy, with Snorkel supporting open research into evaluation and AI-agent benchmarks.
The challenge will be scaling expert-generated data without sacrificing the rigor that makes the data valuable in the first place.
As AI models become more capable, the bottleneck may increasingly shift from generating more data to constructing the right data: realistic environments, difficult tasks, reliable rubrics and measurable outcomes.
Snorkel’s new funding is aimed at expanding precisely that layer of the AI stack.
Market Landscape
AI infrastructure is expanding beyond model training and inference into data preparation, evaluation, simulation and agent environments. Gartner’s 2026 research identifies agent-ready data and agent evaluation as emerging requirements, while McKinsey’s research highlights data readiness as a constraint on scaling enterprise AI.
The market is also shifting toward agentic systems that interact with tools and execute multi-step workflows. McKinsey’s 2026 survey found that 40% of respondents from organizations with more than $1 billion in annual revenue reported scaling AI agents.
This creates demand for infrastructure capable of testing not just model outputs but entire agent behaviors. Data providers that combine domain expertise, evaluation design, environments and research capabilities could consequently occupy a more strategic position in the AI development stack.
Top Insights
- Snorkel AI raised $350 million at a $3.5 billion valuation to expand its agentic data factory for frontier AI development.
- The company is moving beyond conventional labeling toward expert-built tasks, environments and evaluation rubrics for complex AI systems.
- Gartner identifies agent-ready data and behavior-focused agent evaluation as emerging requirements for reliable agentic AI deployment.
- Snorkel says its Data-as-a-Service business has grown more than 18-fold since its September 2025 launch.
- Funding will support expansion across enterprise AI, vertical applications, new modalities and open AI research.
Power Tomorrow’s Intelligence — Build It with TechEdgeAI












