← all releases
Apple·Sep 25, 2026·v9B·open source

LensVLM

LensVLM-9B is a 9B-parameter vision-language model designed to process compressed images of text by selectively expanding relevant regions to their uncompressed form. By utilizing a learned tool-based approach, it maintains high accuracy in document and code understanding tasks while significantly reducing the visual token count required for processing.

vision-language-modellong-contextvisual-text-compressionconversational
overview

LensVLM-9B is a 9B-parameter vision-language model designed to process compressed images of text by selectively expanding relevant regions to their uncompressed form. By utilizing a learned tool-based approach, it maintains high accuracy in document and code understanding tasks while significantly reducing the visual token count required for processing.

key features
  • 01Selective context expansion for dynamic high-resolution zooming
  • 02Efficient processing of compressed visual text representations
  • 03Built on the robust Qwen3.5-9B-Base architecture
  • 04Supports multimodal document and code understanding
  • 05Optimized for high compression ratios up to 10.1x
related productsbrowse all →