Sondera, a New York‑based startup focused on AI governance, announced that its latest research on automatically turning natural‑language policy into formally verified controls for autonomous agents has been selected for presentation at two prominent academic venues and a high‑profile security showcase.
Academic validation at ICML and FLoC
The paper titled “Autoformalization of Agent Instructions into Policy‑as‑Code”, authored by Adam Mondl, Matthew Maisel, and John Brock, received acceptance at the Agents in the Wild workshop during the 2026 International Conference on machine learning (ICML). The same work was also approved for the LLM‑Solve workshop at the 2026 Federated Logic Conference (FLoC). Both venues are known for rigorous peer review and attract leading researchers in machine learning, formal methods, and automated reasoning.
In the study, the authors used the MedAgentBench benchmark—a set of clinical AI agent tasks published in NEJM AI—to evaluate their pipeline. The results show that Sondera’s system automatically formalized 23 of 88 policy rules, surpassing prior hand‑coded attempts, and successfully blocked every adversarial unsafe write operation (99 out of 99). Those figures illustrate a tangible step toward scaling policy enforcement beyond manual rule authoring.
Black Hat Arsenal demonstration
A companion tool, “GolemHalt: A Deterministic Reference Monitor for AI Coding Agents,” will be showcased at Black Hat Arsenal, the conference’s curated exhibition of open‑source security tools. The demo is expected to illustrate how the monitor enforces deterministic decisions—allow, deny, or escalate—on each agent action, independent of the model’s prompt context.
How the pipeline works
Sondera’s approach begins by ingesting natural‑language policy documents—ranging from HIPAA manuals and FINRA guidelines to internal SOPs—and compiling them directly into Cedar policy‑as‑code. A built‑in theorem prover validates every rule, while an adversarial simulation suite stress‑tests the generated code before it reaches production. This dual verification aims to catch edge cases and ensure that legitimate workflows remain uninterrupted.
The architecture blends neural and symbolic techniques. Large language models act as “judges,” providing probabilistic assessments of an agent’s behavior, whereas the formally verified symbolic rules make the final, deterministic access decision. Because enforcement occurs outside the LLM’s context window, classic attack vectors such as prompt injection or model drift cannot bypass the policy layer. Moreover, the system maintains state across an agent’s entire execution trace, allowing decisions to adapt based on prior actions.
Business implications
“The agent incidents we see today aren’t from prompt injection and hijacking. They’re from authorized humans asking authorized agents to do legitimate tasks, like analyzing a financial file or configuring a server. Along the way, the agent reaches the goal with unintended behavior, like leaking or destroying data,” explained Josh Devon, co‑founder and CEO of Sondera. “Even beyond that, enterprises are struggling to apply business logic at scale to their agents, rules that aren’t security or compliance but just standard operating procedure, like a coding agent that should stand up a server on the company’s approved cloud account rather than an unapproved vendor. Our research is focused on letting organizations turn on the most powerful, long‑running agents possible while having the confidence that they will follow the rules.”
For enterprises that are rapidly deploying AI agents across finance, healthcare, and IT operations, the ability to enforce policy at scale without hand‑coding each rule could be a game changer. The deterministic nature of the enforcement layer also offers a clearer audit trail, which is increasingly important for regulatory compliance and internal governance.
Power Tomorrow’s Intelligence — Build It with TechEdgeAI












