Skip to content

North Micro Vision Instruct

CohereLabs/North-Micro-Vision-Instruct
Open Source · chat · open-weights
Open Alert me on changes
Context
8K
Max output
1K
Weights
Open
API $/1M
Modalities
text · image
Released
10 Aug 2026
Download image Share on X Share on LinkedIn
AI summary
● machine-written

North Micro Vision Instruct: 2.4B open-weight vision-language model

North Micro Vision Instruct is a 2.4B-parameter open-weight vision-language model from Cohere Labs combining a 400M native-resolution vision encoder with a 2B language model. It specializes in document understanding, OCR, and visual grounding while maintaining deployability on edge hardware. Released under Apache 2.0 license with no per-token API pricing.

What's new
  • Native-resolution vision encoder preserves up to A4 page detail (1654×2339 px at 200 dpi)
  • 92.1% DocVQA and 73.2% RefCOCO-avg grounding—strongest in size class
  • 2.4B parameters (400M vision + 2B language) designed for edge deployment and fine-tuning
  • Open-weights, self-hostable with no per-token API pricing
Best for
Document and chart understanding with layout preservationCompact edge-aware multimodal applicationsVisual grounding and object counting tasksMultilingual visual question answering
Sources

Source: https://huggingface.co/CohereLabs/North-Micro-Vision-Instruct