Tools

Automating Eval Design and Hillclimbing in Claude Code

New updates to the claude-api skill introduce automated evaluation building and hillclimbing to help developers improve application performance systematically.

A
AIDeveloper44 Team
September 30, 2026·3 min read
Automating Eval Design and Hillclimbing in Claude Code

Visualizing the iterative process of model performance optimization.

TL;DR
  • The new /claude-api build-eval command assists developers in creating codebase-integrated evaluations using production logs and synthetic data.
  • The /claude-api hillclimb command automates iterative performance improvements while utilizing held-out datasets to prevent overfitting.
  • Guidelines emphasize using mirror-production tasks, maintaining passable headroom at the frontier, and minimizing run-to-run variance for reliable results.

Overview of Evaluation and Optimization

Developing robust AI applications requires reliable signals regarding performance. However, designing meaningful evaluations and iteratively improving agent capabilities without introducing bias or overfitting is a technical challenge. Anthropic has addressed these needs by integrating specialized guidance and commands into the claude-api skill, designed to work within the Claude Code environment.

Principles of Effective Evaluation Design

Effective evaluations must provide actionable data. According to the guidance provided by the official documentation, a well-structured evaluation should mirror production environments. Developers are encouraged to select tasks that represent actual use cases rather than those that are merely easy to generate or grade. Furthermore, evaluation tasks should exhibit "passable headroom"—the most capable models should perform well but remain below 100% accuracy, allowing for meaningful measurement of iterative improvements.

High variance is a common obstacle in AI evaluation. It often stems from ambiguous task definitions or inconsistent graders. To combat this, developers should ensure that tasks are specific enough that independent domain experts would reach consistent verdicts. Additionally, the infrastructure used to run evaluations must be checked for potential noise, such as leftover environment state or inconsistent effort configurations.

Automated Workflow with /claude-api build-eval

The /claude-api build-eval command streamlines the creation of these evaluations. When triggered, the tool initiates a guided workflow where it interviews the developer to build tests directly within the codebase. The process typically follows a structured sequence for sampling inputs:

  • Extraction from production transcripts, with considerations for sensitive data.
  • Utilization of existing bug reports and support tickets.
  • Manual input of five to ten specific test cases.
  • Synthesis of cases based on the existing codebase.

Once inputs are collected, the tool assists in validating the grader. It can suggest programmatic verification for constrained outputs or implement an "LLM-as-judge" approach for more open-ended scenarios. In the latter case, a separate model is employed to analyze the input, output, and a rubric of checkable claims, providing a score and reasoning without relying on subjective 1-to-5 scales.

Iterative Improvement via Hillclimbing

Once a reliable evaluation baseline is established, developers can apply the /claude-api hillclimb command to improve their application’s performance. Hillclimbing involves making incremental changes to parameters like prompts or skill definitions to optimize for cost, latency, or quality. To ensure these improvements are genuine, the tool employs held-out sets of examples that are not part of the primary training or configuration data, which helps detect and prevent overfitting.

Successful hillclimbing relies on iteration paths that are inexpensive and attributable. Modifying text-based elements, such as prompts, is generally faster and easier to revert than making deep, structural changes to an agent's core architecture. By automating these feedback loops, the claude-api skill provides a framework for consistent model tuning while maintaining rigorous standards for testing and validation.

Diagram: The Claude Code evaluation pipeline showing the iterative hillclimbing process.

Enjoyed this?

Get more posts like this delivered to your inbox.