Skip to content

Vision-Language Models

About this topic

Vision-Language Models couple a visual encoder with a language model so a single system can perceive images (and video) and reason or converse about them. This topic tracks architectures, training recipes, benchmarks, grounding/interpretability, and the move toward unified multimodal foundation models.

๐Ÿ—“๏ธ By year

  • 2026 โ€” 718 papers ยท 15 in latest batch
  • 2025 โ€” 400 papers ยท year in review