every ai dev release, the day it lands
find the tool
for the thing you
are building right now
latest releases, docs and repos for models, agents, inference stacks and clis — with the numbers that say whether anyone actually uses them.
one field, one button — and it is free.
latest releases
all tools →WeKnora is an open-source, LLM-powered knowledge framework designed for enterprise-grade document understanding, semantic retrieval, and autonomous reasoning. It transforms raw documents into a queryable RAG system, an autonomous agent, and a self-maintaining, interlinked Wiki.
The A2A CLI is the official command-line interface for interacting with agents compatible with the Agent2Agent (A2A) protocol. It enables developers to discover, message, and manage autonomous AI agents directly from their terminal, providing a standardized way to handle agent communication across different frameworks and languages.
SpecShip is an autonomous engineering workflow designed for AI coding agents that orchestrates a complete pipeline from reconnaissance to shipping. It enforces rigorous quality standards through test-driven development, adversarial validation, and spec-driven quality gates to ensure AI-generated code is production-ready.
OpenCodeReview is an AI-powered CLI tool designed for automated code reviews by combining deterministic engineering pipelines with LLM agents. It provides precise, line-level feedback and includes a built-in ruleset for detecting common vulnerabilities like NPEs, thread-safety issues, XSS, and SQL injection.
Supabase Evals is an open-source benchmark and testing framework designed to evaluate how effectively AI coding agents can build, deploy, and troubleshoot projects using the Supabase platform. It provides a standardized way to measure agent performance across real-world developer tasks, such as schema design, debugging Edge Functions, and managing Row Level Security (RLS) policies.
Channels SDK is an open-source library that enables developers to deploy AI agents into popular chat platforms like Slack, Microsoft Teams, Discord, and Telegram. It allows agents to utilize native, interactive UI elements, tool calling, and file handling directly within the conversation flow.
SoL-Pi is a standalone extension for the Pi coding agent that introduces four efficiency mechanisms to reduce token consumption and inference costs. It optimizes agent performance by streamlining tool calls, observation handling, and context management without requiring modifications to the core Pi source code.
ThinkingBox is an open-source sandbox and benchmarking framework designed to evaluate the reliability of AI agents in stateful business workflows. It separates the execution framework from the benchmark package, allowing for independent updates to the harness and test scenarios.
ARTEMIS is an open-source framework that enables AI assistants to perform reliable Android automation using natural language instructions. It integrates with AI coding assistants via the Model Context Protocol (MCP) to execute end-to-end workflows, capture logs, and conduct autonomous testing on real devices.
The Agents API is an OpenAI managed cloud service that allows developers to build and execute long-running AI agents using the Codex harness. It provides context management with automated compaction, multi-agent coordination with subagent support, programmatic tool calling, and flexible deployment options across OpenAI-hosted sandboxes or third-party infrastructure.
DeepJIT is a lightweight, header-only C++20 runtime library designed for Just-In-Time (JIT) compilation of kernels on NVIDIA CUDA GPUs and HUAWEI Ascend NPUs. It provides a unified interface for compiling, caching, and launching device kernels, enabling developers to manage JIT infrastructure efficiently within their C++ or Python extensions.
GPT-Live-1 is a full-duplex voice model designed for the OpenAI API that enables natural, real-time conversational experiences. It allows developers to build voice agents capable of simultaneous listening and speaking, featuring advanced interruption handling and customizable tone, pace, and style.
Neki is a sharded PostgreSQL solution architected from first principles to provide horizontal scaling and high reliability for demanding database workloads. Developed by the team behind Vitess, it leverages deep operational expertise to manage distributed Postgres clusters without being a fork of existing technologies.
SWE-2 is an advanced agentic coding model designed to push the Pareto frontier of capability and cost-efficiency. It utilizes a novel reinforcement learning (RL) algorithm that trains across multiple reasoning-effort levels in a single run, enabling high-performance software engineering tasks at a fraction of the cost of previous frontier models.
Gemini API Skills is a library of specialized context-enhancing modules designed to improve the performance of AI coding agents when working with the Gemini API. By providing up-to-date best practices, SDK syntax, and model-specific guidance, these skills help bridge the knowledge gap between static LLM training data and the rapidly evolving Gemini ecosystem.
AuK is a 1.5B-parameter foundational model designed to unify speech generation and editing through a single natural-language instruction interface. It supports a wide range of tasks including zero-shot TTS, content editing, speech enhancement, and source separation, utilizing a hybrid rectified-flow Transformer architecture.
The Megakernel Serving Engine is a high-performance inference system designed specifically for the North Mini Code model. It utilizes a persistent CUDA megakernel to execute the complete decode forward pass, significantly reducing overhead by eliminating per-op kernel launches and grid synchronization.
Xiaomi-TabLDM is a tabular foundation model designed for classification and regression tasks using in-context learning. It eliminates the need for task-specific fine-tuning by leveraging large-scale synthetic pretraining on structural causal models to achieve high predictive accuracy and efficiency.
Skill Recorder is a desktop application that captures your on-screen work sessions, including clicks, window switches, and optional narration. It leverages the GitHub Copilot CLI to reconstruct these actions into clear, intent-based steps, enabling the creation of reusable skills or automations for AI agents like Microsoft Scout and Copilot Studio.
tgrep is a high-performance, trigram-indexed grep tool designed for rapid regex searching within massive codebases. By utilizing a client/server architecture with persistent indexing, it significantly reduces search latency compared to traditional tools like grep or ripgrep.
what we wrote this week
all posts →OpenAI Details Habitat Storage Scaling and Rust Rewrite
OpenAI has shared details on Habitat, the online storage platform supporting ChatGPT and Codex, which shifted from Python to Rust to handle massive throughput.
Anthropic Adds Plugin Evaluation Framework to Claude Code
Claude Code now features built-in evaluation tools allowing developers to benchmark plugin performance against a no-plugin baseline and gate CI workflows.
OpenAI Introduces Agents API Powered by Codex Harness
OpenAI has announced the launch of the Agents API, enabling developers to build cloud-based agents connected to sandbox environments and custom tools.
give us a follow
every release also goes out on x, linkedin and reddit the moment it lands.