Skip to content

Vision-Language Models โ€” 2025

๐Ÿ—“๏ธ 2025 in Review

๐Ÿ•’ Timeline

  • 2021: CLIP establishes contrastive image-text pretraining; global-embedding alignment becomes the substrate for zero-shot transfer, retrieval, and OOD detection.
  • 2022: BLIP-2/Q-Former and Flamingo show frozen-encoder + LLM bridging; prompt tuning (CoOp) emerges as the lightweight adaptation paradigm.
  • 2023: LLaVA-style visual instruction tuning + GPT-4V make conversational VLMs mainstream; "bag-of-words" compositionality weaknesses of CLIP become widely documented.
  • 2024-early: Open VLM scaling (Qwen-VL, InternVL, SigLIP backbones) and any-resolution tiling; long-context and high-resolution become active frontiers.
  • 2024-mid: ฯ€0 and flow/diffusion action experts launch the modern VLA era, grafting continuous-action heads onto pretrained VLM backbones.
  • 2024-late: DeepSeek-R1 / o1 reasoning RL lands; the open question becomes whether verifiable-reward RL transfers to multimodal reasoning.
  • 2025-early: GRPO-based multimodal reasoning (OpenVLThinker, GRIT, ViCrit) and "thinking with images" (ViLaSR, VISER) demonstrate RL-driven visual reasoning at 3B-7B scale.
  • 2025-mid: Mechanistic critique matures (two-hop problem, narrow gate, binding problem); VLA engineering (knowledge insulation, real-time chunking, force/safety) and agentic VLMs (OpenCUA, VAGEN) consolidate.

๐Ÿ“ˆ Trend

Where the field stands

Vision-language models in 2025 have matured from contrastive dual-encoders (CLIP) and instruction-tuned captioners (LLaVA) into a sprawling ecosystem spanning reasoning, embodied control, generation, and safety. The defining shift this year is the migration of the LLM reasoning playbook โ€” R1-style reinforcement learning with verifiable rewards, chain-of-thought, test-time scaling โ€” into the multimodal setting, alongside the rise of Vision-Language-Action (VLA) models as the dominant interface between perception and robotics. Simultaneously, a maturing critical literature is diagnosing why VLMs fail: binding errors, late-emerging visual entities, modality bottlenecks, and compositional gaps that scaling alone does not close. The result is a field that is bifurcating between making VLMs bigger/longer-context and making them mechanistically understood, efficient, and safe. The NeurIPS 2025 corpus here captures this consolidation: fewer architectural revolutions, more targeted interventions on reasoning, grounding, efficiency, and reliability.

โžก๏ธ Read the full trend analysis

๐Ÿ“„ Papers (400)

โžก๏ธ Paper list โ€” 400 papers, grouped by month. ยท โญ 16 key papers (see Key papers in the left nav).