As AI developers exhaust increasingly large portions of the internet as a source of training material, a new data frontier is emerging: people performing real tasks in physical workplaces. Realset AI and Flatkey have raised $10 million in Series A funding to expand a data-capture network built around real workers, expert demonstrators and operational environments for training and evaluating AI models and agents.
Realset AI is targeting one of the harder problems in AI development: generating reliable training and evaluation data that reflects what people actually do outside controlled digital environments.
The company says its $10 million Series A will fund expansion of its real-world capture network, increase its pool of domain experts and support open benchmarks designed to measure whether AI policies work in real environments.
Rather than relying primarily on web-scraped material, crowdsourced annotations or simulated environments, Realset builds datasets from people performing tasks in workplaces and physical settings. The approach targets applications ranging from embodied AI and robotics to enterprise agents handling commerce, customer support and logistics workflows.
From Web Data to Real-World Ground Truth
Realset divides its offering into three areas.
Realset Body captures skilled workers performing manipulation tasks in environments such as homes, kitchens, warehouses and light-assembly operations. Its datasets can combine egocentric and third-person video with stereo imagery, inertial measurements and action logs. Bimanual teleoperation episodes are also captured for vision-language-action (VLA) model training.
The emphasis is on recording not simply what an object looks like, but how a skilled person actually performs a task.
Realset Field takes a similar approach to software agents. The company creates reinforcement-learning environments based on real workflows, including e-commerce operations, customer support, logistics and manufacturing procedures. Domain experts generate trajectories, preference data and evaluation rubrics inside those environments.
That gives AI developers a way to train agents against business processes rather than abstract tasks.
Realset Judge focuses on evaluation after deployment. Experts who understand the underlying job can identify agent failures, create targeted corrective data and continuously re-score systems as models and prompts change.
A Data Pipeline Built Around Experts
The three services operate through the Realset Workspace. Domain experts across six languages—English, Simplified Chinese, Japanese, Korean, Spanish and Arabic—complete tasks, provide evidence and screen recordings, and undergo independent quality review.
Approved records are exported in JSONL with provenance information that customers can verify.
That workflow positions Realset closer to an AI data infrastructure provider than a conventional annotation vendor. The company is attempting to make the people performing a task part of the data-generation and evaluation loop.
Measuring Whether AI Works Outside the Lab
Realset is also developing open benchmarks around real-world performance.
Its planned Realset Household Manipulation Bench will evaluate VLA policies including π0, OpenVLA, GR00T and Octo on tasks such as folding, loading, sorting and wiping in real household environments. The company says results are expected in Q4 2026.
Additional benchmarks are planned for light assembly and computer-use agents operating real e-commerce seller workflows.
The broader shift is important for AI developers. As models become more capable, simply measuring performance on static datasets may provide less information about whether an agent can reliably complete a physical or business task.
Realset’s approach instead treats data capture, domain expertise and evaluation as one continuous system—a model-development layer increasingly relevant as AI moves from generating content to performing work.
Market Landscape
AI training has historically benefited from enormous quantities of publicly available digital information. Generative AI and agentic systems, however, increasingly require data that represents actions, workflows, outcomes and physical environments.
This has created demand for specialized datasets for robotics, computer-use agents, reinforcement learning and enterprise AI evaluation. Companies such as NVIDIA, Google DeepMind and Amazon are also investing heavily in embodied AI and robotics research, increasing the importance of real-world interaction data.
Realset’s model reflects another emerging trend: using domain experts not merely to label examples, but to generate demonstrations, establish evaluation criteria and diagnose failures.
For enterprise AI, that could become particularly important as organizations move from benchmark performance toward measuring whether agents can reliably complete specific operational processes.
Top Insights
- Realset AI and Flatkey have raised $10 million in Series A funding to expand real-world AI training and evaluation data infrastructure.
- Realset Body targets embodied AI with video, sensor and action data captured from skilled workers performing physical tasks.
- Realset Field creates agent-training environments around real business workflows, including commerce, support, logistics and manufacturing.
- Realset Judge uses domain experts to evaluate deployed agents, diagnose failures and generate targeted corrective data.
- Open benchmarks could provide a more practical measure of AI performance by testing models against real physical and business tasks.
Power Tomorrow’s Intelligence — Build It with TechEdgeAI











