The partnership between Genoria AI—a subsidiary of MGI Tech—and the Shanghai Artificial Intelligence Laboratory has produced two new offerings aimed at turning AI‑generated experimental designs into concrete actions on automated laboratory equipment. The duo introduced ProtoPilot, a self‑evolving multi‑agent platform, and BioLab Bench, the first end‑to‑end benchmark that evaluates an AI agent’s ability to move from intent to execution on real lab devices.
Physical AI moves beyond textual answers
The two products are positioned as the foundation of what Genoria and the Shanghai AI Lab call “Physical AI for life sciences.” Unlike most generative models that stop at producing a textual protocol, these systems strive to close the loop: an AI interprets a research goal, drafts a protocol, translates it into device‑specific code, runs the experiment on an automation platform, and then incorporates wet‑lab feedback to refine its next attempt.
The underlying research was released as a preprint on arXiv (arXiv:2606.31763) in June 2026, underscoring the academic rigor behind the commercial push.
ProtoPilot: a full‑cycle, self‑improving agent
ProtoPilot tackles the entire experimental workflow—design, protocol generation, code conversion, device execution, and feedback—through a chain of cooperating agents. Its most striking claim is the ability to learn from failure. In a recent pilot, a PCA assembly step faltered because of an antibiotic‑resistance screening error. ProtoPilot identified the root cause, automatically generated a corrected protocol, and reran the experiment, demonstrating a closed‑loop learning capability that many AI‑driven lab tools still lack.
Performance on the public ProtocolQA benchmark—a widely cited test for AI reasoning in experimental design—places ProtoPilot close to human expertise:
- GPT‑5.6‑sol: 43.5 %
- Human expert baseline: 54 %
- ProtoPilot: 52.38 %
The result suggests that a purpose‑built, domain‑specific agent can rival general‑purpose large language models on specialized tasks.
BioLab Bench: the first real‑world evaluation framework
Benchmarking AI for scientific workflows has traditionally focused on the correctness of generated text. BioLab Bench expands that scope by measuring whether an AI can actually drive a physical instrument to a successful outcome. Its key attributes include:
- Comprehensive task set – Scenarios range from simple pipetting to multi‑step, high‑complexity workflows, organized into three difficulty tiers (L1–L3).
- Full‑chain assessment – The framework scores each stage: intent parsing, protocol drafting, device‑agnostic SOP creation, translation to device‑specific SOPs, generation of executable code, and verification of successful execution.
- Cross‑device portability – Tests can be run on different automation platforms, allowing vendors to gauge how well an AI agent adapts to varying hardware constraints.
By demanding actual device interaction, BioLab Bench forces developers to address latency, error handling, and hardware abstraction—issues that are often glossed over in purely textual evaluations.
Toward 24 × 7 unattended laboratories
Both ProtoPilot and BioLab Bench are framed as stepping stones toward fully autonomous, round‑the‑clock laboratories. The vision is that AI agents will no longer rely solely on text‑based fine‑tuning. Instead, they will ingest a continuous stream of real‑world data: experimental tasks, automation logs, expert annotations, failure cases, and wet‑lab outcomes. Over time, this corpus is expected to endow agents with integrated reasoning, execution, and validation capabilities, enabling enterprises to run high‑throughput experiments without human supervision.
A decade of AI‑biology synergy at MGI
MGI’s foray into AI dates back to 2019. In 2025, Dr. Yang Meng—then Chief AI Officer—co‑authored a Nature Biomedical Engineering paper with Professor Nattiya Hirankarn (Chulalongkorn University) that introduced “PrimeGen,” a dry‑wet collaborative multi‑agent system that linked primer design, experimental validation, and automated workstation execution. The success of PrimeGen laid the groundwork for the 2026 launch of Genoria AI, a dedicated AI‑for‑Science subsidiary focused on building closed‑loop infrastructure for biotech research.
The latest Physical AI initiative leverages MGI’s deep integration with its own automation hardware and the experience gathered from more than 3,800 global users. Rather than chasing sheer compute scale, Genoria’s strategy emphasizes “agent scaling and closed‑loop data engineering,” as Dr. Yang Meng explained after assuming the role of CEO:
“It reflects a different path from the pure compute race. While leading AI companies rely on scale compute to push the capabilities of general‑purpose models, we take a different approach. Through agent scaling and closed‑loop data engineering, we organize real‑world tasks, device constraints, expert feedback, and wet‑lab results into a training ground where AI continuously evolves.”
From a strategic standpoint, the collaboration blends China’s state‑backed AI research capacity (Shanghai AI Lab) with MGI’s commercial hardware expertise. This hybrid model could give both parties a competitive edge against Western AI giants that are primarily focused on scaling large language models without direct hardware integration.
Market implications and competitive landscape
The emergence of Physical AI tools could reshape the enterprise biotech automation enterprise market, which has so far been dominated by hardware vendors offering point‑solution software layers. By providing an open, benchmark‑driven framework, BioLab Bench may encourage a more modular ecosystem where third‑party AI agents can be swapped in without redesigning the underlying hardware stack.
For developers, the shift toward real‑world execution metrics means that future AI models will need to be trained on multimodal data that includes device telemetry and experimental outcomes—not just text corpora. Enterprises looking to adopt these systems will have to invest in data pipelines capable of capturing, normalizing, and feeding back wet‑lab results at scale.
Outlook
If ProtoPilot and BioLab Bench deliver on their promises, the next wave of biotech R&D could see a significant reduction in manual protocol engineering, faster iteration cycles, and higher reproducibility across labs. The real test will be how quickly enterprises can integrate these agents into existing automation workflows and whether the closed‑loop data collection can sustain continuous model improvement without prohibitive overhead.
Power Tomorrow’s Intelligence — Build It with TechEdgeAI












