Back to News
NemoStation (Shubham Sharma)•May 19, 2026
Marlin-2B
Paid
VLM
Video Understanding
Temporal Grounding
Video Captioning
Open Source
HuggingFace
2B Model
Structured Extraction
### TL;DR
Marlin-2B is a tiny 2B video vision-language model (VLM) fine-tuned for extracting structured information from videos. It produces second-precise Scene + Event captions and resolves natural-language temporal queries — the strongest open model in its weight class on dense captioning and temporal grounding benchmarks, competitive with Gemini-2.5-Flash.
Key Insights & Metrics
Pricing
Free (Open Source, Apache-2.0)
Cost structure
Version
Marlin-2B (SimPO checkpoint)
Current release version
Hardware
Single consumer GPU (NVIDIA recommended); transformers>=5.7.0, torch>=2.11.0, torchcodec
Compute requirements
Category
Paid
Licensing model
Region
Global
Primary region
Key Features
- marlin.caption() — returns structured Scene + Events JSON with second-precise timestamps; state-of-the-art at 2B on CaReBench and DREAM-1K dense video captioning benchmarks, beating all open models in its weight class
- marlin.find() — resolves natural-language event queries to span-grounded (start, end) time ranges; beats Qwen2.5-VL-7B by +6.4 mIoU on TimeLens-Bench and matches Gemini-2.0-Flash at only 2B params
- Deployable on a single consumer GPU — vLLM and swift-deploy compatible, standard HuggingFace transformers API, Gradio demo ready; fine-tuned from Qwen3.5-2B with SimPO preference optimization
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!