
Anthropic Adds Plugin Evaluation Framework to Claude Code
Claude Code now features built-in evaluation tools allowing developers to benchmark plugin performance against a no-plugin baseline and gate CI workflows.
Read more →
OpenAI Introduces Agents API Powered by Codex Harness
OpenAI has announced the launch of the Agents API, enabling developers to build cloud-based agents connected to sandbox environments and custom tools.
Read more →
Microsoft Open-Sources ThinkingBox to Benchmark AI Agent Side Effects
Microsoft has released ThinkingBox, an open-source evaluation framework that verifies AI agents by inspecting tool side effects rather than transcripts.
Read more →
Perplexity Search API Integrated into Hermes Agent
Perplexity has added its Search API to Hermes Agent, providing live web retrieval and page extraction across an index of over 400 billion URLs.
Read more →
GitHub Introduces Project HydraFusion Multi-Model Orchestration
GitHub has unveiled Project HydraFusion, a multi-model orchestration workflow designed to improve coding task quality while lowering inference costs.
Read more →
Anthropic Introduces 'ant apply' for Declarative AI Resource Management
Anthropic has released ant apply in the ant CLI, enabling developers to manage Claude agents, environments, and skills declaratively via code repositories.
Read more →
Google Releases Gemini 3.8 Flash and Gemini 3.8 Flash Cyber
Google has announced two new AI models, Gemini 3.8 Flash and Gemini 3.8 Flash Cyber, designed for agentic workflows and cybersecurity applications.
Read more →
Tencent AI Releases CubeSandbox v7 with Cross-Node State Persistence
Tencent Cloud's CubeSandbox v0.7.0 introduces cross-node pause and resume capabilities, optimized networking, and improved control plane separation for AI agents.
Read more →
OpenClaw Releases Version 2.0 After Significant Structural Rework
OpenClaw 2.0 introduces a rebuilt browser interface, simplified installation workflows, and new collaborative shared cloud sessions.
Read more →
Vercel Releases Run SDK for Secure Agentic Code Execution
Vercel has launched the Run SDK, a hardened QuickJS sandbox for executing untrusted JavaScript and TypeScript with support for human-in-the-loop approval.
Read more →