Tools

Microsoft Introduces run-assert-eval for AI Agent Risk Management

Microsoft has released run-assert-eval, a new skill designed to automate risk discovery, policy generation, and evaluation for AI agents.

A
AIDeveloper44 Team
September 27, 2026·3 min read
Microsoft Introduces run-assert-eval for AI Agent Risk Management

The run-assert-eval tool automates the process of discovering, fixing, and verifying AI agent risks.

TL;DR
  • The new run-assert-eval skill automates the identification and mitigation of risks for AI agents through a unified workflow.
  • It integrates existing tools like Clarity, ASSERT, and Agent Control Specification to streamline the development cycle.
  • The process allows developers to discover failure modes, generate runtime policies, and verify fixes using a single prompt.

Automating AI Agent Governance

Microsoft has announced the release of run-assert-eval, a new skill designed to assist developers in governing AI agents at runtime. This tool aims to reduce the manual overhead typically associated with identifying agent failure modes, authoring runtime policies, and validating improvements. By automating these steps, the tool seeks to provide a more rigorous and evidence-based approach to agent security and reliability.

The governance of AI agents is a relatively new field, and the development of run-assert-eval follows previous open-source releases from Microsoft, including ASSERT and the Agent Control Specification (ACS). These tools were created to help teams move from theoretical requirements to structured, executable evaluations. While these tools were already available, the company noted that integrating them manually could be complex and prone to errors, such as losing context between test runs or failing to maintain the integrity of performance comparisons.

The Integrated Workflow

Run-assert-eval is designed to address the challenges of manual integration. According to Microsoft's official documentation, the tool functions by enabling developers to execute a complete diagnostic and remediation loop with a single prompt. This process involves four key stages:

  • Risk Discovery: The tool uses Clarity to identify potential failure modes and risks that may not have been previously accounted for in written requirements.
  • Measurement: Using the ASSERT framework, the tool assesses the frequency and nature of agent failures based on the identified risks.
  • Policy Generation: It generates runtime policy directly from the findings, utilizing the Agent Control Specification to ensure compatibility and portability.
  • Validation: Finally, the tool reruns the evaluation to provide evidence that the applied policy effectively mitigates the identified risks without negatively impacting the agent's performance.

This automated loop ensures that the behavior definition, the test cases, and the judge remain consistent throughout the remediation process. By removing the need for developers to manually translate failure modes into evaluation configurations or author individual Rego rules, the tool helps maintain the validity of the comparison between the original and the governed agent.

Addressing Foundational Assumptions

The development of this tool is rooted in a desire to correct two common assumptions in AI development. The first is that written requirements fully capture every possible risk; in practice, many significant failure modes are unanticipated. The second is that teams possess the resources and time to manually connect every step of the evaluation and governance pipeline. By integrating Clarity at the beginning of the evaluation loop, run-assert-eval allows the entire process to begin with discovery rather than relying solely on pre-defined requirements.

The tool is designed to be framework-agnostic and trace-aware, allowing it to evaluate a wide range of agentic systems, including those built with LangGraph, CrewAI, and other orchestration frameworks. By grounding judgments in OpenTelemetry traces, it provides clear evidence for why an agent failed, citing specific tool calls, routing decisions, and model responses.

Diagram: The run-assert-eval pipeline for AI agent risk management.

Enjoyed this?

Get more posts like this delivered to your inbox.