Sat, Aug 15
(1)Ori Harness is a command-line interface tool that allows developers to run existing agent CLIs on OpenRouter without changing their workflow. It provides optimized configurations, unified billing, and organization-wide guardrails for various AI agent harnesses.
Fri, Aug 14
(4)GLM-5.3 is a frontier-grade large language model optimized for advanced coding and complex agentic workflows. It serves as the latest iteration in the GLM-5 series, delivering significant performance improvements in reasoning and code generation tasks.
The Perplexity Search SDK is an agents-first Python library that enables AI agents to orchestrate custom search pipelines using 'Search as Code' primitives. By allowing agents to write Python code to control retrieval, ranking, and filtering, it reduces token usage and improves performance for complex, multi-step research tasks.
pgbot is an open-source PostgreSQL intelligence tool that provides real-time database health monitoring, performance analysis, and AI-powered insights. It acts as an 'AI DBA' by analyzing database statistics to offer actionable recommendations, helping users identify and resolve performance bottlenecks without manual dashboard monitoring.
Qwen3.8-27B is a highly capable, open-weight multimodal model designed for advanced reasoning and agentic tasks. It features native vision and video understanding, a 256K context window, and a sophisticated thinking mechanism that can be adjusted for various performance needs.
Thu, Aug 13
(4)DeepSeek Harness (dsh) is an open-source agent runtime framework designed for building autonomous agents. It utilizes a modular architecture where every component—including models, tools, session management, and the agent loop—is implemented as a plugin powered by the Cordis framework.
Dots3-Note Preview is an open-weight, multimodal Mixture-of-Experts (MoE) model featuring 280B total parameters and 16B active parameters. Designed for long-horizon agentic tasks, it supports a 512K context window and integrates text, visual, and audio understanding to solve complex, real-world problems.
Gemini 3.7 Flash is an advanced, high-performance AI model optimized for coding, agentic workflows, and complex reasoning tasks. It delivers significant improvements in debugging, issue resolution, and web development while maintaining cost-efficiency for developers and enterprises.
Toast 1 is a specialized search agent designed for knowledge-intensive tasks, capable of decomposing queries, gathering evidence, and curating context. It matches or outperforms frontier models like Claude Opus 5 and GPT-5.6 Sol while offering significantly higher speed and lower operational costs.
Wed, Aug 12
(4)Delta is a multiplayer environment designed for coding with AI agents and reviewing their output. It utilizes DeltaDB to keep code and conversation threads synchronized in real-time, allowing developers and agents to collaborate with full context.
Docker VMM is a fully rebuilt, container-optimized virtualization layer for Docker Desktop that replaces third-party solutions. It provides enhanced performance, stability, and memory management by allowing Docker to own and tune the entire virtualization stack.
Grok 4.6 is a frontier AI model designed for long-running agentic tasks, complex coding, and multi-step knowledge work. It features improved reasoning capabilities and visual processing, optimized for turning broad ideas into polished applications.
LFM2.5-VL-3B is a 3.1-billion-parameter vision-language model designed for high-performance, on-device inference. It features significant improvements in screen understanding, grounding, and function calling, allowing it to outperform larger models while maintaining low latency on edge hardware.
Tue, Aug 11
(4)MAI-Code-1.1-Flash is a lightweight, agentic coding model designed to provide high-quality code generation with superior efficiency. It is specifically optimized for GitHub Copilot and VS Code, offering faster token streaming and reduced operational costs for engineering teams.
NVIDIA Nemotron 3.5 Lightning 30B-A3B-NVFP4 is a high-performance, latency-optimized large language model featuring a hybrid Mamba-2 and Mixture-of-Experts (MoE) architecture. Designed for efficient agentic workflows, it utilizes 30B total parameters with 3B active parameters per token to deliver fast, accurate execution for specialized tasks.
Nemotron-RL Agentic Terminal Pivot v1 is an open-source reinforcement learning dataset designed to train large language models for agentic command-line interface (CLI) tasks. It provides high-quality, expert-derived trajectories that enable models to perform complex software engineering operations, tool use, and reasoning within Linux environments.
Solid 2.0 is a major update to the declarative JavaScript library for building user interfaces, introducing first-class asynchronous reactivity. By integrating async directly into the reactive graph, it eliminates the need for specialized primitives like createResource, simplifying data fetching and state management.
Mon, Aug 10
(3)Muse Glimmer is a 30-billion-parameter open agentic model optimized for always-on local workflows on consumer hardware. It enables advanced capabilities like multi-step reasoning, reliable tool use, and multimodal understanding while running locally on devices like Macs and PCs.
North Micro Vision Instruct is a 2.4B-parameter open-weight vision-language model designed for efficient multimodal tasks. It features native-resolution image processing and is optimized for prototyping, task-specific fine-tuning, and specialized visual applications.
SIE is an open-source, self-hosted inference server designed to run a wide variety of AI models for agentic workflows, including embedding, reranking, entity extraction, and text generation. It provides a unified API that replaces fragmented model-specific servers, allowing developers to serve over 100 models from a single cluster with on-demand loading.
Sun, Aug 9
(1)LLaDA2.2-flash is an agent-oriented Mixture-of-Experts (MoE) diffusion language model designed for long-context agentic workloads. It introduces Levenshtein Editing with DELETE and INSERT control tokens to enable efficient parallel generation, error correction, and multi-turn tool use.
Fri, Aug 7
(1)Agent Orchestrator is an open-source development tool designed to manage fleets of AI coding agents in parallel. It provides a centralized dashboard to monitor agent sessions, handle git worktree isolation, and autonomously manage CI failures, code reviews, and merge conflicts.
Thu, Aug 6
(3)Adapt-1 Preview is a non-transformer, neuro-symbolic substrate designed for test-time learning. It enables applications to learn and adapt while operating by maintaining a persistent, inspectable state without requiring task-specific pretraining or an LLM in the decision loop.
Agent Plugins is an open, vendor-neutral standard for packaging reusable components into portable plugins for AI agents. It provides a consistent format for bundling Agent Skills and Model Context Protocol (MCP) servers, allowing them to be discovered and loaded across different compatible AI agent clients.
Kitesurf is a stateless, highly scalable web browser designed specifically for AI agents to perform tasks like HTML extraction and screenshot generation. It runs entirely on Cloudflare Workers, offering significantly higher efficiency in CPU and memory usage compared to traditional Chromium-based browsers.
Wed, Aug 5
(4)Cloudflare OS is an open-source AI productivity platform designed to function as an operating system for enterprise AI workloads. It enables employees to build, share, and run secure AI-powered applications (gadgets) while maintaining strict governance through capability-based security and Gatekeepers.
Muse Code is a terminal-based coding agent powered by the Muse Spark 1.2 model, designed to handle complex software engineering tasks across large repositories. It features persistent background agents, repository-scale execution, and built-in verification to automate planning, coding, and debugging workflows.
Prime Agent is an open-source, self-improving coding harness that utilizes Recursive Language Model (RLM) and Continual Harness abstractions to manage complex, long-horizon tasks. It enables agents to programmatically manage their own state, sub-agents, and memory through a persistent IPython kernel, allowing for autonomous, iterative improvement.
The Cursor SDK Bridge is an open-source protocol and local server that enables developers to drive Cursor agents from any programming language. By providing a stable sdk.v1 protobuf contract, it allows languages like Rust, Go, and Java to interact with Cursor's agent runtime without requiring direct dependency on the official TypeScript or Python SDKs.
Tue, Aug 4
(7)Anydoc is a high-performance Rust library designed to convert various office document formats, including Word, PowerPoint, Excel, and PDF, into clean, LLM-ready GitHub-Flavored Markdown. It provides consistent output across different file types and includes language bindings for Node.js and Python.
Kiro Crew is an open-source, persistent development workspace designed to transform AI coding agents into autonomous engineering teams. It enables agents to maintain context across sessions, schedule recurring tasks, and integrate with various developer tools for long-running workflows.
LFM2.5-2.6B is a compact, agentic foundation model designed for efficient on-device deployment. It features a 128K context window and is specifically optimized for agentic workflows, including planning, tool use, and multi-step task execution.
Maple-Preview is an open-source 20B-A1B ternary-weight reasoning large language model optimized for efficient on-device inference. It features a 24-layer, 256-expert architecture designed to deliver state-of-the-art reasoning capabilities while maintaining high performance on consumer hardware like the Mac mini M4.
Pokee-Isaac 28B is a 28-billion parameter agentic model designed for long-context tasks, featuring a 10-million-token context window. It is built on a proprietary non-decoder-only architecture that allows for efficient deployment on single consumer-grade GPUs like the RTX 4090.
Qwen-MM-Plugins is an open-source collection of multimodal skills and Model Context Protocol (MCP) servers designed to make various AI agent harnesses natively multimodal. It enables agents to perform complex tasks such as long-video analysis, 3D modeling in Blender, parametric CAD in FreeCAD, and media generation without requiring a custom runtime.
Remote Agent Browser is a tool that enables running the agent-browser automation CLI within isolated, cloud-based Vercel Sandboxes. It allows AI agents to perform browser tasks like navigation, interaction, and data extraction in a secure, ephemeral environment.
Mon, Aug 3
(4)The Computer by Cloudflare package introduces a new way to build and scale AI agents by providing them with their own dedicated 'computer' environment. Instead of relying solely on heavy, expensive containers, this runtime intelligently switches between lightweight isolates and container sandboxes based on the task requirements, such as file manipulation or native binary execution. This approach allows developers to scale agents to millions of concurrent instances while maintaining cost-efficiency and performance. By offering a durable, unified filesystem that stays in sync across different execution backends, the library simplifies the development of complex agentic systems. It provides developers with fine-grained control, auditability, and a clear paper trail of agent actions, making it a robust solution for production-grade autonomous workloads. The project is currently available as an open-source library, enabling developers to build agents that are secure, scalable, and capable of handling diverse tasks.
DeepCode is an open-source, multi-agent coding platform designed to automate the translation of high-level inputs—such as academic research papers, technical documentation, and natural language requirements—into production-ready software. By orchestrating specialized agents through a unified runtime, it enables complex workflows like Paper2Code reproduction, full-stack prototyping, and automated repository maintenance.
Ori Eval is an automated evaluation tool designed to help developers systematically identify the best AI models for their specific applications. It scans your codebase, generates test cases from your prompts, and uses an LLM-as-a-judge to grade performance across metrics like accuracy, latency, and cost.
Sandbox SDK is an open-source TypeScript library that provides a unified interface for managing isolated code execution environments. It allows developers to spin up sandboxes across various providers, including local environments, E2B, Daytona, Vercel, Upstash, Ascii Box, and Railway, using a consistent, clean syntax.
Sun, Aug 2
(1)Qwen 3.8 Max is a flagship large language model developed by Alibaba, featuring a massive 2.4 trillion parameter architecture. Currently available in a preview version, it is designed to deliver frontier-level performance for complex reasoning and coding tasks.
Sat, Aug 1
(3)OpenVuln is an AI-powered tool hosted on Hugging Face that scans software repositories to identify potential security vulnerabilities. It leverages the GLM (General Language Model) family of models to assist developers in safeguarding their codebases.
Sprites are persistent, hardware-isolated Linux environments designed for running arbitrary code with stateful checkpoint and restore capabilities. They function as dedicated microVMs that hibernate when idle and wake automatically, making them ideal for AI agents, development environments, and secure code execution.
The Windows App Development CLI (winapp CLI) is a unified command-line interface designed to streamline the development of Windows applications. It simplifies complex tasks such as managing Windows SDKs, handling MSIX packaging, generating app identity, and configuring manifests and certificates for various app frameworks.