Vision-Language Models¶
About this topic
Vision-Language Models couple a visual encoder with a language model so a single system can perceive images (and video) and reason or converse about them. This topic tracks architectures, training recipes, benchmarks, grounding/interpretability, and the move toward unified multimodal foundation models.