CL-bench
### TL;DR
CL-bench is a benchmark designed to evaluate language models' ability to learn from complex, context-dependent tasks. It comprises 500 complex contexts, 1,899 tasks, and 31,607 verification rubrics, all crafted by experienced domain experts.
Key Insights & Metrics
Key Features
- 500 complex contexts
- 1,899 tasks
- 31,607 verification rubrics
- Evaluates context learning in language models
- Developed by Tencent's Hunyuan Team
→ Related Releases
Protenix
Protenix is an open-source, trainable PyTorch implementation of AlphaFold 3, designed for high-accuracy biomolecular structure prediction. It aims to advance accessible and extensible research tools for the computational biology community. ([github.com](https://github.com/bytedance/Protenix?utm_source=openai))
AIBuildAI
AIBuildAI is an AI agent that autonomously constructs AI models. Given a specific task, it initiates an agent loop to analyze the problem, design models, and execute training processes, all without human intervention.
MiroThinker
MiroThinker is an open-source search agent model developed by MiroMindAI, designed for tool-augmented reasoning and real-world information seeking. It aims to match the deep research capabilities of leading AI models like OpenAI's Deep Research and Google's Gemini Deep Research.
Letta Code SDK
The Letta Code SDK is a software development kit that enables developers to build deeply personalized agents with persistent memory that learn over time. It serves as the interface to Letta Code, facilitating the creation of stateful agents capable of continuous learning and improvement.
Discussion
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!