every ai dev release, the day it lands
find the tool
for the thing you
are building right now
latest releases, docs and repos for models, agents, inference stacks and clis — with the numbers that say whether anyone actually uses them.
one field, one button — and it is free.
latest releases
all tools →RRSI is a framework for the recursive self-improvement of LLM agent harnesses, designed to optimize prompts, control flow, tools, and memory. It addresses overfitting in agent evolution by regularizing both the proposal and selection processes, ensuring robust performance across diverse task domains.
issue-graph is a command-line interface tool designed to map and visualize the reference graph of GitHub issues, pull requests, and repository backlogs. It enables developers and AI coding agents to identify related work, competing changes, and unresolved follow-ups before beginning new tasks.
run-assert-eval is an automated skill designed to streamline the responsible AI lifecycle for agents by integrating threat modeling, evaluation, and policy enforcement. It enables developers to discover risks, measure agent failure rates, generate runtime policies, and verify fixes through a unified, repeatable loop.
Gitea Runner 4.0.0 is a major update to the Gitea Actions runner, introducing built-in actions, S3-compatible shared caching, and OpenTelemetry tracing. This release focuses on performance and operational flexibility, allowing runners to handle cache requests directly and support more complex workflow configurations.
OrcaSAQ-2 27B is a highly compressed, sensitivity-aware mixed-precision quantized version of the Qwen3.8-27B model. It reduces the storage footprint by over 77% while maintaining high fidelity to the original model, making it suitable for deployment on hardware with limited GPU memory.
LensVLM-9B is a 9B-parameter vision-language model designed to process compressed images of text by selectively expanding relevant regions to their uncompressed form. By utilizing a learned tool-based approach, it maintains high accuracy in document and code understanding tasks while significantly reducing the visual token count required for processing.
TypeLLM is a framework that enables type-safe generation for autoregressive large language models without requiring changes to their architecture or weights. It allows models to retain native thinking and free-form generation capabilities while ensuring outputs strictly adhere to defined JSON schemas.
The Docker Sandbox Kit Specification v3 is an open-source standard for packaging AI agent environments, including network rules, credentials, and volumes, into a single, reproducible OCI image. It enables developers to define agent authority as code, ensuring that security policies and environment configurations are versioned, diffable, and portable across different runtimes.
Cloud Sandboxes are microVM-based isolated environments that allow AI coding agents to run autonomously on Docker-managed compute. They enable developers to offload long-running tasks from their local machines to the cloud, ensuring security and persistence for tasks that take hours to complete.
Qwen-Planner-Agent is a closed-loop AI-for-AI framework designed for real-world mobile planner agents. It integrates a trained planner model with a stateful harness to manage complex mobile tasks, utilizing persistent memory, reusable skills, and sub-agent coordination to achieve high performance in tool-use and reasoning.
Worker Previews provides isolated, production-like environments for every Git branch or pull request within Cloudflare Workers. It allows developers and coding agents to test changes in parallel with dedicated URLs, configurations, and state, ensuring that updates behave as expected before merging to production.
GPT-6 Sol and Luna are the latest additions to the GPT-6 model family, designed to provide high-performance intelligence at improved cost-efficiency. These models leverage the advanced reasoning and alignment capabilities of GPT-6 Astra, optimized for professional workflows, coding, and computer-use tasks.
Claude Opus 5.5 is a high-performance AI model optimized for agentic coding, complex knowledge work, and large-scale data analysis. It delivers significant improvements in efficiency and communication, offering a 40% reduction in operating costs compared to its predecessor, Claude Opus 5.
The MiMo-V2.6 series is Xiaomi's latest frontier intelligence model lineup, designed to provide advanced multimodal capabilities and high-performance agentic workflows. The series includes specialized variants like the MiMo-V2.6-Pro and MiMo-V2.6-Flash, which have undergone extensive reinforcement learning training to optimize professional-grade tasks.
AX is a high-throughput, declarative orchestrator designed to run autonomous agent workloads at scale. It provides a sandboxed, stateful runtime that manages agent lifecycles, including sub-second suspension and resumption, to optimize compute density and cost.
Grok 4.7 is a frontier-level AI model optimized for complex coding, knowledge work, and agentic tasks. It features enhanced self-verification capabilities, a larger base model, and a new safeguard stack designed for high-stakes dual-use domains.
ZCode is an Agentic Development Environment (ADE) designed for long-horizon, multi-step software engineering tasks. It integrates a self-developed AI agent with desktop, web, and terminal interfaces to manage the entire development lifecycle, including planning, coding, debugging, and deployment.
jina-ocr-v1 is an efficient, end-to-end document parsing model designed to convert images and PDFs into structured Markdown in a single pass. It utilizes a 3.4B parameter mixture-of-experts architecture with 570M active parameters to deliver high-performance document transcription on low-budget hardware.
Qwen3.8-Omni-Flash is a next-generation native omnimodal model designed to enhance agentic capabilities through advanced audio, video, and text understanding. It features a 1M-token context window and provides significant performance improvements in multimodal reasoning and long-horizon task execution.
Parse Gateway is an intelligent document routing tool that analyzes PDF pages to determine their complexity. It automatically directs each page to the most cost-effective parsing tier—ranging from local LiteParse to advanced LlamaParse—ensuring optimal performance and cost efficiency for LLM pipelines.
what we wrote this week
all posts →Microsoft Introduces run-assert-eval for AI Agent Risk Management
Microsoft has released run-assert-eval, a new skill designed to automate risk discovery, policy generation, and evaluation for AI agents.
Gitea Releases Runner v4: Shared S3 Cache and Faster CI/CD with Built-in Actions
Gitea has released version 4.0.0 of its runner, introducing S3-compatible caching, built-in actions, and improved observability via OpenTelemetry.
Anthropic Introduces New Plugin Submission Portal for Claude
Anthropic has launched a new developer portal to streamline the submission, review, and analytics process for plugins in the Claude directory.
give us a follow
every release also goes out on x, linkedin and reddit the moment it lands.