
ResearchSep 11, 2026·4 min read
Specific Labs Introduces Real-SWE Benchmark Using Private Codebases
Specific Labs has launched Real-SWE, evaluating AI coding agents on licensed, private enterprise codebases across 640 scored rollouts.
Read more →
ResearchSep 2, 2026·4 min read
Google Research Releases TimesFM-3 for Multivariate Forecasting
Google Research has launched TimesFM-3, a 300M parameter zero-shot foundation model designed for complex multivariate time-series forecasting tasks.
Read more →
ResearchJul 3, 2026·5 min read
Mistral Releases Leanstral 1.5 for Formal Verification in Lean 4
Mistral AI has launched Leanstral 1.5, an open-source model with 6 billion active parameters designed specifically for automated theorem proving and code verification.
Read more →
ResearchMay 15, 2026·2 min read
Poetiq's Meta-System Hits SOTA on LiveCodeBench Pro With No Fine-Tuning
Poetiq just hit state-of-the-art performance on LiveCodeBench Pro — without fine-tuning a single model or using any privileged API access. Just standard APIs and a Meta-System that built its own coding harness from scratch.
Read more →