Poetiq
🔥 This release made it to our blog
Poetiq's Recursive Self-Improvement Tops LiveCodeBench Pro: Flash Model Beats Gemini Deep Think
Poetiq's Meta-System has set a new state-of-the-art on LiveCodeBench Pro (LCB Pro) by automatically constructing and optimizing a coding harness through recursive self-improvement. The system improved Gemini 3.1 Pro by 12.3%, pushed GPT-5.5 to 93.9%, and surpassed Google's own Gemini Deep Think — all without fine-tuning or privileged model access.
### TL;DR
Poetiq is a model-agnostic meta-system that recursively improves the reasoning and knowledge extraction of any LLM without fine-tuning. It establishes new state-of-the-art results on ARC-AGI-1 & 2, Humanity's Last Exam, SimpleQA, and LiveCodeBench Pro — making cheaper, smaller models outperform larger expensive ones by orchestrating smarter multi-step reasoning strategies.
Key Insights & Metrics
Key Features
- Recursive self-improvement harness — LLM-agnostic system that automatically builds optimized reasoning strategies for any model (Gemini, GPT, Claude, Grok); integrates new models within hours of release and boosts performance across every tested provider
- SOTA on ARC-AGI-1 & 2, HLE (55.0%), SimpleQA (77.3%), and LiveCodeBench Pro — makes Gemini 3.5 Flash match Gemini Deep Think at a fraction of the cost by extracting more knowledge from cheaper models
- Self-auditing iterative problem-solving loop — multi-step generation, feedback, and refinement with autonomous termination; open-sourced configurations available on GitHub
Discussion
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!