
Open SourceSep 11, 2026·4 min read
Microsoft Open-Sources ThinkingBox to Benchmark AI Agent Side Effects
Microsoft has released ThinkingBox, an open-source evaluation framework that verifies AI agents by inspecting tool side effects rather than transcripts.
Read more →ToolsAug 4, 2026·4 min read
OpenRouter Launches Ori Eval for Automated Model Benchmarking
OpenRouter has introduced Ori Eval, a framework for developers to evaluate LLMs using repository-specific prompts, tool assertions, and automated LLM judges.
Read more →