Skip to content

Vision-Language Models — 2026 · Paper list (718)

July 2026 (308)

Date Paper Venue Why selected
2026-07-23 K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
2026-07-23 Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
2026-07-22 Test-Time Training for Modality Order Consistency in Vision-Language Models Gandelsman (Berkeley); TTT fixes systematic VLM modality-order bias; reproducible across 3 models
2026-07-22 Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs ECCV 2026 ECCV 2026; joint token-compute pruning cuts MLLM inference cost; directly deployable
2026-07-22 Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning RLVR training environment for VLMs with taxonomy-guided verifiable visual rewards
2026-07-22 LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition Medical VLM fine-tuning for surgical instrument-tissue interaction; Intel Labs; reproducible
2026-07-22 ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models Taxonomic probe showing VLMs reliably succumb to irrelevant contextual entrainment
2026-07-22 SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments
2026-07-22 ReferTrack: Referring Then Tracking for Embodied Visual Tracking
2026-07-22 Robostral Navigate
2026-07-21 Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models Deepak Pathak + Zhiqiu Lin; LLM-generated CLIP descriptors carry no visual evidence — fundamental rethink
2026-07-21 Dual Adversarial Fine-tuning for Enhancing Robustness of Large Vision Language Model ICT CAS (Shan/Chen/Gao); dual adversarial fine-tuning closes LVLM visual robustness gap
2026-07-21 Stochastic Meta-Unlearning: Bridging Language Backbone and Multimodal Unlearning Sijia Liu + Tianlong Chen; first principled multimodal unlearning bridging language and vision components
2026-07-21 Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval ECCV 2026 ECCV 2026; zero-shot VMR resolving modality and language-style alignment gaps
2026-07-21 PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image PathAgentBench; evidence-seeking VLMs evaluated on multi-scale whole-slide pathology
2026-07-21 ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting ECCV 2026 ECCV 2026; multi-target referring segmentation extending 3DGS to language-guided scene understanding
2026-07-21 No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation Test-time scaling lifts VLM navigation performance without any additional training
2026-07-21 Now You See the Hate: Adaptive View Retrieval for Hidden Hateful Illusions VLM safety gap: state-of-the-art models under 25% accuracy on hateful optical illusions
2026-07-21 Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio Unified text+image+video+audio embedding; fills critical audio gap in VLM retrieval stacks
2026-07-21 DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization GRPO-trained medical VLM for CXR reports; directly replicable RL recipe for domain-specific VLMs
2026-07-21 MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings Multi-party ToM benchmark; measures social/epistemic reasoning gap in MLLMs for agentic settings
2026-07-21 ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis Expert-level visual reasoning benchmark; stress-tests knowledge-intensive VLM synthesis beyond commonsense
2026-07-21 IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer Streaming 4D instance-grounded geometry; spatial intelligence backbone for embodied agents
2026-07-21 OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation On-policy self-distillation adapts LVLM judgments to pixel-level IAD; novel spec-to-detector recipe
2026-07-21 HPD-Parsing: Hierarchical Parallel Document Parsing Hierarchical parallel document parsing with VLMs; speed+accuracy gain for doc-understanding pipelines
2026-07-21 ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
2026-07-21 Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
2026-07-21 DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking Yu Qiao (Shanghai AI Lab); reasoning-guided physics-aware video generation via spatial-temporal masking
2026-07-21 ChronoStitch: Training-Free Composition of Visual KV Memories for Long-Horizon Temporal Reasoning Training-free KV cache composition for long-video QA; avoids video reprocessing
2026-07-21 MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts ECCV 2026 ECCV 2026; benchmark exposing VLM failure on missing-part detection; honest hallucination audit
2026-07-20 Patch Policy: Efficient Embodied Control via Dense Visual Representations Yann LeCun co-author; dense ViT features for embodied control, underexplored direction
2026-07-20 The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Jun-Yan Zhu + Shechtman + Hertzmann; text-prompted multi-sense perceptual metric, novel
2026-07-20 Memory-Supported Synergistic Adaptation for Training-Free Test-Time Medical Image Segmentation ECCV 2026 ECCV 2026; memory-synergistic TTA for medical VLM segmentation, classification→seg bridge
2026-07-20 O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning ECCV 2026 ECCV 2026; object-centric VLM reasoning for industrial anomaly detection
2026-07-20 FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications Agent harness for real-time multimodal deployment; directly actionable for harness builders
2026-07-20 Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation Future-state-conditioned VLN; novel supervision signal beyond next-action cloning
2026-07-20 Reasoning as a Double-Edged Sword: Architecture and Cross-Stage Robustness in Vision-Language-Action Models Empirical robustness audit of reasoning in VLAs; honest baselines; safety-critical insight
2026-07-20 SEE: Structure-aware Exploring \& Exploiting for Long-horizon GUI Agent Trajectory Synthesis GUI agent trajectory synthesis with VLMs; long-horizon coverage; agentic scaffolding
2026-07-20 Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs Causal audit proves attention heatmaps don't reflect actual medical VLM evidence
2026-07-20 Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs Domain-generalized pixel-level tampering detection in VLMs; practical safety/integrity tooling
2026-07-20 FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry
2026-07-20 PRiSM: Prototype Regularization for Few-Shot VLMs Dolz + Ben Ayed; few-shot VLM adaptation under realistic class-imbalance; buildable recipe
2026-07-20 Measuring and Improving Complex-Atomic Answer Consistency in Endoscopic VQA Complex-atomic answer consistency metric for endoscopic VQA; fills medical eval gap
2026-07-19 TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Generalist video temporal grounding with MLLMs; fills key capability gap in video VLMs
2026-07-19 EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding Jiwen Lu (Tsinghua); EvoGUI isolates state-transition reasoning from end-to-end GUI success
2026-07-19 Searching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language Models Evolutionary block pruning finds task-specific vision paths; cuts VLM compute without retraining
2026-07-18 Look Clearly Before Answering: Mitigating Hallucinations in LVLMs via Saliency-Driven Perceptual Realignment Saliency-driven hallucination mitigation in LVLMs; core reliability problem, buildable method
2026-07-18 Dataset Distillation by Influence Matching CVPR 2026 CVPR 2026; outcome-centric dataset distillation; broadly applicable to VLM training data
2026-07-18 Can Experts Adapt Without Training? On Test-Time Modality Generalization in MVLMs MICCAI 2026 MICCAI 2026; test-time modality generalization for medical VLMs; directly buildable
2026-07-18 HTT-Net: Hierarchical Text-guided Transition Modeling for Surgical Video Phase Recognition Text-guided surgical video phase recognition; strong medical VLM baseline with temporal structure
2026-07-17 Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models Deva Ramanan (CMU); prompt-position paradox fix; universal VLM prompting impact
2026-07-17 Von Mises-Fisher Mixture Model with Dynamic Shrinkage for Realistic Test-Time Transduction ICML 2026 ICML 2026; vMF mixture TTA-transduction; rigorous test-time VLM adaptation theory
2026-07-17 More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe Luc Van Gool (ETH); simple recipe beats RS-specific archs; transferable scaling insight
2026-07-17 Model Merging for Medical LVLMs: A Benchmark and a Winner-Take-All Approach Medical LVLM model-merging benchmark + WTA method; practical LoRA-expert fusion
2026-07-17 How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA Vision-operation misalignment taxonomy; actionable failure map for compositional VQA
2026-07-17 ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning VL tool-use + RL for multimodal scientific verification; novel reward signal design
2026-07-17 When Can Test-Time Adaptation Help Zero-Shot CT Vision-Language Models? Leonid Sigal (UBC); systematic TTA scope study on zero-shot 3D CT VLMs
2026-07-17 Region-Grounded Vision-Language Learning for Detection-Guided Mammographic Lesion Classification Region-grounded contrastive VL for mammography; spatial alignment beats global CLIP
2026-07-17 Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs
2026-07-17 Vision-Language Assistant for Emotional Reactions to Risky Driving
2026-07-17 An Exam for Active Observers
2026-07-17 Attention-Guided Saliency Maps for Interpreting Visualization Literacy in VLMs
2026-07-17 Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting
2026-07-17 Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling
2026-07-17 One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models Cross-modal unlearning vulnerability in VLMs; important safety gap with concrete attack surface
2026-07-17 Foresight Residual RL for Long-Horizon Robot Manipulation with Vision-Language-Action Models VLA + residual RL for long-horizon manipulation; Boularias lab; addresses credit assignment gap
2026-07-17 JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models Multi-tenant post-training service for VLA; novel deployment angle; Wentao Zhang group
2026-07-17 PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation Reflective agentic physics control for video gen; VLM+agent+simulation pipeline
2026-07-16 Symbal: Detecting Systematic Misalignments in Model-Generated Captions ICML 2026 Stanford/Langlotz group; ICML 2026; systematic caption error detection for medical MLLMs
2026-07-16 Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology Multi-LLM VLM pipeline for 3D brain MRI reports; novel visual instruction tuning for oncology
2026-07-16 Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models Action supervision reshapes multimodal representations in VLA; structural insight for builders
2026-07-16 FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models Visual foresight + motion guidance in VLA models; explicit forward prediction addresses reactivity
2026-07-16 WorkDrive: Roadwork Chain of Causation for Autonomous Driving Chain-of-causation reasoning for roadwork zones; tackles hard OOD failure mode in driving VLMs
2026-07-16 HoloGeo: Mitigating Landmark Bias in Geo-localization via Evidence-Driven Reasoning Evidence-driven reasoning to mitigate landmark bias in VLM geo-localization
2026-07-16 RoboTTT: Context Scaling for Robot Policies
2026-07-16 SportD: Can VLMs Physically Strategize? Probes VLM physical strategy/reasoning in soccer; novel eval axis for agent decision-making
2026-07-16 U-shaped Multi-granularity Learning for Vision-Language Models Prompt learning granularity dilemma for VLMs; practical recipe for cross-task generalization
2026-07-16 GeoDetect: Geometric Adversarial Detection for VLPs ECCV 2026 Adversarial detection for multimodal VLPs; ECCV 2026; geometric approach is novel angle
2026-07-16 VTM-Nav: Hierarchical Visual-Topological Memory for Cross-Episode Object-Goal Navigation VLM-driven object-goal navigation with topological memory; training-free embodied agent system
2026-07-16 On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline Simplicity-first transferable attack on VLPMs; questions complexity of existing pipelines
2026-07-16 ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
2026-07-16 Knowing You at First Glance: Inferring Apparent Personality from Faces
2026-07-16 Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Xiaomi robotics; 100K hrs real VLA trajectories; massive-scale grounded VLM-action
2026-07-16 Video = World + Event Stream Wan-Streamer (Alibaba); world+event decomposition reframes video VLM architecture
2026-07-16 Training-Free Open-Vocabulary 3D Point-Cloud Segmentation on the Generalized Few-Shot Benchmark
2026-07-15 Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation ECCV 2026
2026-07-15 ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding ECCV 2026
2026-07-15 M\(^\text{4}\)World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming
2026-07-15 SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning
2026-07-15 Fine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT Understanding
2026-07-15 Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs
2026-07-15 Fine-grained CLIP fine-tuning with self-annotated region alignment
2026-07-15 Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models Attention-free token reduction; directly enables VLM deployment on edge/resource-constrained devices
2026-07-15 Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Representation anchoring + language-action alignment; fixes BC fine-tuning catastrophic forgetting in VLAs
2026-07-15 Semantic Anchoring for Robotic Action Representations Semantic anchoring preserves VLM priors in VLA fine-tuning; Yizhou Wang (PKU) group
2026-07-15 Self-Improving is Often Sudden: Enlightenment-style Finetuning for Large-Scale Models Phase-transition self-improvement phenomenon; Tianwei Zhang (NTU); actionable finetuning recipe
2026-07-15 FM\(^2\): Unified Federated Foundation Models for Heterogeneous Multimodal Medical Imaging Federated multimodal medical VLM; addresses privacy-constrained cross-institution medical AI deployment
2026-07-15 SD-MAR: Multi-image Analytical Reasoning via Synthetic Data and Reinforcement Learning RL + synthetic data for multi-image analytical reasoning; addresses core VLM reasoning gap
2026-07-15 MonteRET: AI Agent Enhancing Multimodal LLMs with Multi-granularity Knowledge Retrieval for Chest CT Report Generation Multi-granularity RAG agent for chest CT reports; retrospective clinical evaluation included
2026-07-15 DiMaS: Distribution Matching for Steering Vision-Language-Action Models Fine-grained behavioral steering for flow-matching VLA; enables controllable robot policy
2026-07-15 Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection Force injection during VLA post-training; addresses contact-state blindspot in vision-driven policies
2026-07-15 RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
2026-07-15 GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning Zero-shot transit video analytics with grounded hybrid VLM reasoning; practical deployment recipe
2026-07-15 Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making
2026-07-15 ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning
2026-07-14 VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression Visual token compression for VLMs; training-free LLM-as-encoder approach; addresses core inference efficiency bottleneck
2026-07-14 CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models CoRe framework for cross-image comparative reasoning; fine-grained attribute grounding; fills known VLM compositional gap
2026-07-14 Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding Gaussian mixture keyframe selection for long video VLMs; event-aware visual allocation; addresses uniform-sampling information loss
2026-07-14 Hy-Embodied-VLM-1.0: Efficient Physical-World Agents Hy-Embodied-VLM-1.0; end-to-end embodied agent with multimodal perception and agentic reasoning; buildable system report
2026-07-14 Auditing Data Leakage in Whole-Slide Image Multimodal Benchmarks Data leakage audit in WSI pathology VLM benchmarks; questions benchmark validity; critical for medical VLM evaluation
2026-07-14 A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism Controlled null result on GRPO failing for small VLMs; mechanistic analysis; essential calibration for agent RL fine-tuning
2026-07-14 What Does a Temporal Benchmark Score Measure? Decomposing Channel Use in Video VLM Evaluation Decomposes what temporal benchmark scores actually measure in video VLMs; channel-use analysis; methodological contribution
2026-07-14 Let RGB Be the Language of Vision Unified RGB formulation for all visual modalities (masks, depth, etc.); architectural simplification with broad VLM implications
2026-07-14 ReflectVLN: Training Vision-Language Navigation Agents with Reflective Reasoning ReflectVLN adds closed-loop reflective reasoning to VLN agents; explicit failure diagnosis mechanism; directly applicable to agent harnesses
2026-07-14 The Sound of Absence: Audio-Language Embedding Models Struggle with Negation Hung-yi Lee (NTU); exposes fundamental negation blindspot in CLAP-style audio-language models
2026-07-14 Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering
2026-07-14 MQAdapter: Multi-Modal Quantum Adapter for Coarse-to-Fine VLM Fine-tuning Quantum-inspired adapter for coarse-to-fine VLM fine-tuning; novel PEFT direction
2026-07-14 DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery Spatial cognition augmentation for VLMs on street-view; directly relevant to geo-VLM eval
2026-07-14 Lost in Visual Translation: A VLM-Assisted Perceptual-Semantic Coherence Framework for EEG-to-Image Reconstruction VLM-assisted perceptual-semantic coherence for EEG-to-image; novel eval framework
2026-07-14 Breaking Déjà Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning VLM reasoning for visual place recognition auditing; novel robustness eval for embodied nav
2026-07-14 Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation
2026-07-13 MED-DSLC: Multi-Expert-Domain Classification via Domain Supervision and Logit Calibration ECCV 2026 ECCV 2026; Nuno Vasconcelos (UCSD); fixes CLIP logit non-comparability across domains—principled zero-shot fix
2026-07-13 StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure Causal structure for long-horizon VLM digital agents; directly relevant to agent harness design
2026-07-13 Confidence Scores in Open-Vocabulary Detection Are a Biased Mixture of Scale and Semantics Open-vocab detector confidence = scale+semantics mixture; exposes fundamental reliability flaw in CLIP-based OVD
2026-07-13 TIGER: Text-Conditioned Visual Gated Routing with Acceptance Alignment for Multimodal Speculative Decoding Multimodal speculative decoding with gated routing; practical inference speedup for VLMs
2026-07-13 When Depth Is Better Told Than Shown: Depth-Ordinal Prompting for Vision-Language Spatial Reasoning Depth-ordinal prompting for VLM spatial reasoning; addresses a core 3D perception weakness in VLMs
2026-07-13 Think When It Matters: Conditional VLM Reasoning for Social Navigation with RL Policies Conditional VLM reasoning with RL for social navigation; conditional compute in embodied agents
2026-07-13 StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description ECCV
2026-07-13 DynEval: Holistic Evaluations of T2I Generative Models in the Wild ECCV 2026
2026-07-13 An Empirical Analysis of Continual Learning for Heterogeneous Medical Visual Question Answering Empirical continual-learning study for medical VQA; practical for multi-task clinical VLM deployment
2026-07-13 LaGuadia: Language-Guided Adaptive Distillation from Pathology Foundation Models Language-guided distillation from pathology FMs; Won-Ki Jeong; reduces WSI inference cost
2026-07-13 MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models
2026-07-13 The Emerging Paradigm of Geospatial Foundation Models: From Pre-Training to Agentic Reasoning
2026-07-13 SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning Self-verification RL for multimodal reasoning; extends R1-style training to VLMs
2026-07-13 See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models Robot-frame 3D pointmaps for VLA; fixes camera-frame mismatch, strong geometric inductive bias
2026-07-12 Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models ICML 2026 ICML 2026; spectral-heat-flow token condensation; directly cuts VLM inference cost without info loss
2026-07-12 Learning Anatomy-Grounded CT Vision-Language Representations with Organ-Hierarchical Report Knowledge Anatomy-hierarchical VLP from CT+reports; strong medical VLM pretraining recipe for organ-level grounding
2026-07-12 3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects Factorial study of VLM eval pipelines for 3D generation; rigorous methodology for automated judge reliability
2026-07-12 Mixture of Cognitive Experts in Large Vision-Language Models Mixture-of-cognitive-experts for LVLMs; novel framing of metacognition and diverse representations
2026-07-12 Detecting AI-Generated Video: A Vision-Language Dual-View Survey ACL 2026
2026-07-12 Traj-VLN: Learning Pixel-Space Interaction via Autoregressive Trajectory Generation Autoregressive pixel-space trajectory generation for VLN; novel grounding approach for navigation agents
2026-07-11 What Does Your Short-Answer VQA Score Actually Measure? Evaluator-Dependent Instability in Multimodal Short-Answer Benchmarks Demonstrates VQA benchmark scores are evaluator-dependent; critical for anyone running VLM evals
2026-07-11 Empowering Long-form Omni-modal Understanding with Robust Audio Perception Weidi Xie (SJTU/Oxford); long-form omni-modal understanding with explicit audio alignment; scarce labeled data angle
2026-07-11 SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding COLM 2026 COLM 2026; controlled benchmark for long-context VLM document understanding, fills real eval gap
2026-07-11 PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning Physical plausibility reasoning in video-LMs; targets a known VLM failure mode with retrieval+verification
2026-07-11 Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift Benchmarks 15 foundation model backbones for mammography under domain shift; directly relevant to medical VLM eval
2026-07-11 WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding Training-free UHR remote sensing VLM reasoning with structured evidence; novel efficiency angle
2026-07-11 ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Robotic agent OS with lifelong multimodal memory; directly relevant to agent harness architecture
2026-07-11 Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection Reversible adversarial examples against VLMs for privacy; novel threat model with practical reversibility
2026-07-11 ABot-N1: Toward a General Visual Language Navigation Foundation Model Foundation model for Visual Language Navigation; unifies spatial reasoning and embodied versatility
2026-07-10 ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts Pathology foundation model fusing vision, VL, and slide-level experts; builds on strong computational pathology trend
2026-07-10 Video Generation Models are General-Purpose Vision Learners ECCV 2026 Video generation as general vision learner; ECCV 2026; Jasper Uijlings (Google); strong foundational claim
2026-07-10 SigLIP-HD by Fine-to-Coarse Supervision ICLR 2026 SigLIP-HD fine-to-coarse supervision; ICLR 2026; Hengshuang Zhao group; directly improves MLLM visual encoding
2026-07-10 Super-Generalist: Towards Comprehensive and Accurate Medical Image Understanding via Generalist-Specialist Synergy Generalist-specialist synergy for medical image understanding; addresses known VLM accuracy gap in clinical settings
2026-07-10 MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models Heterogeneous VLM with linear attention interleaving; efficiency + performance; actionable architecture direction
2026-07-10 Robustifying Vision-Language Models via Test-Time Prompt Adaptation ICML 2026 ICML 2026; test-time prompt adaptation for adversarial robustness in VLMs; strong venue, practical defense method
2026-07-10 The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs Diagnoses counting failures as representation-verbalization gap; mechanistic insight actionable for VLM evaluation
2026-07-10 Test-Time Scaling for Small VLMs on Multilingual Visual MCQ Test-time scaling on small open VLMs—transfers TTS gains beyond frontier models; highly actionable
2026-07-10 Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference Reveals decode (LLM text gen) not vision tokens as edge energy bottleneck; inverts prevailing optimization target
2026-07-10 Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models Decade-long longitudinal analysis of VLM visual-cognitive errors; rare rigorous error taxonomy work
2026-07-10 TCLA: Training-Free Class-wise Logit Adaptation for Medical Vision-Language Models Training-free class-wise logit adaptation for OOD medical VLMs; lightweight, reproducible, practical
2026-07-09 Dive Into the Implicit Biases of Low-rank Vision-language Alignment ECCV 2026 ECCV 2026; challenges full-param alignment dogma with low-rank analysis — rethinks VLM training recipe
2026-07-09 Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Cognitive memory architecture for unified understand/generate/edit — directly improves multimodal agent design
2026-07-09 When Thinking Hurts: Epistemic Signals in the Reasoning Chains of Visual Language Models ICML 2026 ICML 2026; first empirical characterization of reasoning-chain uncertainty in thinking-mode VLMs — calibration insight
2026-07-09 APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts Stanford (Haber lab); interleaved vision-language thoughts for long-horizon robot planning — actionable VLA scaffold
2026-07-09 Rethinking Small VLM Quantization: From Component-Wise Analysis to Hardware-Aware Edge Deployment ICML 2026 ICML 2026; component-wise quantization analysis for sub-3B VLMs — direct edge deployment recipe
2026-07-09 Playing ZendoWorld: Challenging AI Agents on Active Visual Concept Induction ZendoWorld benchmark for active visual concept induction — novel evaluation of hypothesis-driven VLM reasoning
2026-07-09 When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities Sparse autoencoders for cross-modal mechanistic interpretability in VLMs — novel probe into internal representations
2026-07-09 VEGAS: Human-Aligned Video Caption Evaluation via Gaze Gaze-grounded VLM caption eval; training-free metric; novel human-attention alignment signal
2026-07-09 LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action Selective visual token routing in VLA; addresses core dynamic-scene limitation in robot VLMs
2026-07-09 AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding CVPR CVPR dashcam VQA benchmark; incident-centric eval fills gap in autonomous driving VLM testing
2026-07-09 Post-Training in End-to-End Autonomous Driving Comprehensive post-training survey for E2E autonomous driving VLAs; strong practical recipes
2026-07-09 WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving Dual-level world-cognitive VLA for autonomous driving; addresses reactive limitation with world foresight
2026-07-09 VocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database Retrieval Visual tokenization + vector DB for open-vocab detection/segmentation; novel retrieval-grounded approach
2026-07-09 Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data Adaptive benchmark sizing for model evaluation; directly actionable for rigorous VLM evaluation practice
2026-07-09 Open-ended Multi-agent Autocurricula via Visual Inspection of Policies with Multi-modal LLMs Multimodal LLMs as policy inspectors for RL curricula; novel VLM-as-evaluator application
2026-07-09 Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation Attribute-driven compositional referring segmentation for endoscopy; medical VLM with fine-grained control
2026-07-09 Write-Protected Discrete Bottlenecks for Language-Grounded World Models: A Structural Limitation and Sufficient Fix Structural analysis of language-grounded world models; critiques RT-2/PaLM-E paradigm with fix
2026-07-09 MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs Multi-view integration benchmark; targets untested allocentric 3D reasoning gap in VLMs; diagnostic benchmark value
2026-07-09 Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing Learning from privileged modalities via probing; addresses inference-time modality mismatch; practical for deployment
2026-07-09 Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Blind-Spots-Bench; systematic evaluation of multimodal model failure modes humans find trivial; benchmark for VLM robustness
2026-07-08 HIVE: Understanding Post-Hallucination Reasoning in Vision Language Models ECCV 2026 ECCV 2026; dissects post-hallucination reasoning chains in VLMs, not just detection
2026-07-08 AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning ECCV 2026 ECCV 2026; query-relevance-anchored visual token pruning cuts VLM inference cost
2026-07-08 Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
2026-07-08 Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering Comparative domain-adapted VLMs for DocVQA — honest baselines across architectures for document understanding
2026-07-08 When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs Graph-based attribute reasoning for VLM calibration — addresses overconfidence in prompt-tuned models
2026-07-08 UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma
2026-07-08 Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks ACL ACL survey; comprehensive multimodal unlearning across modalities; timely safety/privacy topic
2026-07-08 Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models ACL ACL; multilingual spatial deictic eval; exposes systematic VLM spatial reasoning failure mode
2026-07-08 On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces Spectral subspace lens on VLM adversarial vulnerability; novel mechanistic robustness analysis
2026-07-08 InfraQR: Edge-Placed QR-Inspired Structured Patch Attacks on Infrared Vision-Language Models Structured adversarial patch attacks on infrared VLMs; underexplored robustness axis with security relevance
2026-07-08 Generalist Vision-Language Models for Fast Radio Burst detection: a zero-shot benchmark against a specialized detector Zero-shot VLM benchmark vs. specialized FRB detector; honest generalist vs. specialist evaluation
2026-07-08 Heterogeneity-Adaptive Diffusion Schrodinger Bridge for PET-Guided Whole-Body MRI Translation MICCAI 2026
2026-07-08 Prototype-Anchored Generalized Manifold Regression for Unknown-Domain Object Detection
2026-07-08 BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning Brain-inspired backward prediction for VLM self-reflection; novel training-free reasoning improvement mechanism
2026-07-07 Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders ICML 2026 ICML 2026; Daniel Cohen-Or/Or Patashnik; novel mechanistic insight into VLM localization signals
2026-07-07 Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment ICML 26 ICML 2026; VLM physical reasoning failure modes directly impact VLA agent generalization
2026-07-07 SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs ECCV ECCV; SAM-based optimizer for VLM prompt tuning; directly addresses performance-generalization tradeoff
2026-07-07 What Images Cannot Say: Language-Guided Olfactory Representation Learning ECCV 2026
2026-07-07 Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation
2026-07-07 MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation
2026-07-07 AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models
2026-07-07 A VLM-Enhanced Framework for Comprehensive Traffic Sign Condition Assessment Integrating Daytime Visual Performance and Nighttime Retroreflectivity Evaluation
2026-07-07 VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery
2026-07-07 Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement
2026-07-07 TMF-RSE: Tri-Modal Fusion with Regional Semantics and Evidential Uncertainty for Lung Severity Scoring
2026-07-07 PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails Policy-adaptive image guardrails benchmark; critical for real deployment safety
2026-07-07 Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability Multicultural/multilingual VLM safety benchmark exposing Western-centric evaluation gaps
2026-07-07 Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding
2026-07-07 SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models
2026-07-07 RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures
2026-07-07 Structured-Condensed Prompt Tuning in Vision-Language Models for Fine-grained Image Recognition Structured prompt tuning for CLIP fine-grained recognition — practical, reproducible PEFT recipe for VLMs
2026-07-07 Reward Valuation in Vision Language Models: Causal Mechanisms Underlying Anhedonia Schrimpf lab; mechanistic VLM–cognition alignment; novel anhedonia/reward valuation probe
2026-07-07 Pelican-VLA 0.5: Attending Before Acting Benefits Generalization Unified VLA with attention-before-action; integrates future-frame gen + action in one model
2026-07-07 UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation UI-to-app generation benchmark; visual interaction inference; novel multimodal code-gen eval
2026-07-07 PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet 3D dense captioning with pseudo-cap + voxel net; concrete architectural advance for 3D VLMs
2026-07-07 Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review Review of VLA models across UAV and bimanual manipulation; useful scope survey for builders
2026-07-07 VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection
2026-07-06 TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models Interpretable concept-overlap token reduction cuts VLM compute with explainable mechanism
2026-07-06 Does It Fail to See or Fail to Know? Attributing Errors in Vision-Language Models Disentangles perception vs. knowledge failures in VLM errors — useful for evaluation design
2026-07-06 Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data? VLMs for medical data standardization — tackles missing preprocessing step in clinical AI pipelines
2026-07-06 QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding ECCV 2026 ECCV 2026; query-conditioned temporal retrieval tackles core long-video VLM degradation problem
2026-07-06 PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving ECCV 2026 ECCV 2026; end-to-end VLA for autonomous driving; scalable architecture with practical evaluation
2026-07-06 InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization InternVLA series (Shanghai AI Lab); unifies semantic priors and future prediction in one VLA
2026-07-06 Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation Bidirectional hierarchical VLA agent addressing Markovian failure in long-horizon manipulation
2026-07-06 Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Object-evidence token merging for dense retrieval; practical efficiency for deploying VLMs
2026-07-06 Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models TMLR
2026-07-06 StructuredEdit: Constraint-Aware Graphic Design Editing via Differentiable Parameter Propagation SIGGRAPH
2026-07-06 DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation Dynamic gating for reasoning segmentation bridges language queries to pixel masks
2026-07-06 Repurposing CLIP to Localize at Pixel Level Repurposes CLIP for pixel-level localization, addressing global-feature bias bottleneck
2026-07-06 Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis 3D scene graph + VLM open-vocab scene understanding beyond isolated object lifting
2026-07-06 HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better Tencent HunyuanOCR-1.5; lightweight OCR-VLM unifying doc parsing, extraction, translation
2026-07-05 DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics ICML 2026 ICML; schema-guided world modeling for hierarchical visual dynamics in multimodal LLMs
2026-07-05 SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering Visual token engineering to mitigate hallucinations in LVLMs — practical LVLM reliability fix
2026-07-05 RL Forgets! Towards Continual Policy Optimization Exposes catastrophic forgetting in RL-based VLM post-training — critical for continual adaptation
2026-07-05 ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
2026-07-05 AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in Memes
2026-07-04 Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs Counterfactual grounding + hard-negative contrastive training exposes visual shortcut problem in medical VLMs
2026-07-04 USE: A Unified Self-Ensembling Framework for Test-Time Prompt Tuning ICML 2026 ICML; unified self-ensembling for test-time prompt tuning advances CLIP-based TTA
2026-07-04 Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process RL over interleaved text-image generation; directly informs training unified multimodal agents
2026-07-04 Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Distills tree search into frozen VLA action evaluation; novel adaptation without fine-tuning
2026-07-03 RADIO1D: Elastic Representations for Condensed Vision Modeling ICML 2026 ICML; challenges fixed 2D patch assumption in VLMs with elastic 1D representations
2026-07-03 Present but Not Remembered: Auditing How Frozen VLAs Encode, Deploy, and Steer Visual History Mechanistic audit of how frozen VLAs encode visual history — informs VLA memory architecture
2026-07-03 MentalThink: Shaping Thoughts in Mental SVG World Novel think-with-SVG paradigm enabling executable visual reasoning in MLLMs
2026-07-03 Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment
2026-07-02 Show Me Examples: Inferring Visual Concepts from Image Sets Apple Research (Susskind/Bautista); fills fundamental VLM gap: concept inference from image sets without text
2026-07-02 Visually Grounded Self-Reflection for Vision-Language Models via Reinforcement Learning Durrett lab; RL trains visual self-reflection/CoT correction — key for VLM agent reasoning chains
2026-07-02 LASER: A Corrective Lens for LVLMs via Visual Attention Preservation and Sink Suppression ECCV 2026 ECCV 2026; LASER attention-sink suppression fixes visual forgetting in long LVLM decoding
2026-07-02 Towards Robustness against Typographic Attack with Training-free Concept Localization ECCV 2026 ECCV 2026; training-free CLIP typographic attack defense — critical for all CLIP-encoder VLMs
2026-07-02 Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs Task-agnostic VLA pretraining; novel approach to VLA data bottleneck via motion pre-training
2026-07-02 VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment VLAFlow unified VLA co-training + future latent alignment; clean ablation of pretraining paradigms
2026-07-02 MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models MMBench-Live; continuously evolving benchmark combats contamination — raises evaluation rigor bar
2026-07-02 Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning ECCV 2026 ECCV 2026; entropy-aware visual token pruning under dense/fine-grained instructions — practical efficiency
2026-07-02 Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editing Online recursive MLLM editing with bounded overhead — essential for keeping deployed VLMs current
2026-07-02 Towards Real-World Ultrasound Understanding: Large Vision-Language Models from Multi-Image Examinations with Long-Form Reports Multi-image ultrasound VLM with long-form reports; practical medical LVLM with realistic data regime
2026-07-02 AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models
2026-07-02 Teaching Vision-Language-Action Models What to See and Where to Look ECCV 2026 VLA training data design for autonomous driving; addresses vision-text imbalance in VLA supervision
2026-07-02 SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video Unified 3D spatial reasoning in video; directly relevant to embodied VLM spatial understanding
2026-07-02 DeepGaze3.5-VL: Modeling Scanpaths via Autoregressive Token Prediction Autoregressive scanpath modeling via VLM token prediction; novel attention modeling approach
2026-07-02 COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows Self-evolving skill harnesses for image workflows; directly relevant to agent memory and reusable skills
2026-07-02 MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding Graph-aware VLM for molecular images; novel architecture bridging chemistry and multimodal understanding
2026-07-02 MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding Time-aware medical video benchmark; proactive prediction timing critical for clinical AI systems
2026-07-02 Seek to Segment: Active Perception for Panoramic Referring Segmentation ECCV 2026 Active perception for panoramic segmentation; bridges VLMs and embodied AI agent perception
2026-07-02 Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias VLM reliability under corruption for medical imaging; rigorous evaluation of deployment robustness
2026-07-02 SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models Binarization compression for large VLMs; practical inference efficiency for deployment
2026-07-02 ProCal: Inference-Time Proposal Calibration for Open-Vocabulary Object Detection Inference-time calibration for open-vocab detection; plug-in improvement without retraining
2026-07-02 FlowCIR: Semantic Transport via Flow Matching for Zero-Shot Composed Image Retrieval Zero-shot composed image retrieval via flow matching; novel semantic transport approach for VLMs
2026-07-02 Rank-Then-Act: Reward-Free Control from Frame-Order Progress VLM as reward-free ordinal progress scorer for control; directly relevant to agent harness design
2026-07-02 The Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data Collection Hybrid dynamic data collection for VLA spatial generalization; actionable recipe for robot VLM training
2026-07-02 LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension ECCV 2026 ECCV egocentric video REC benchmark; tests long-form temporal grounding in VLMs
2026-07-02 Domain Generalization via Text-Anchored Information Bottleneck ECCV 2026 Text-anchored information bottleneck for domain generalization; VLM-grounded invariant representation method
2026-07-02 LIME: Learning Intent-aware Camera Motion from Egocentric Video Active vision for robotics; intent-aware camera motion from egocentric video, novel VLM application
2026-07-02 Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots Portable VLA/WAM inference runtime; directly addresses real-world embodied AI deployment gap
2026-07-02 Multimodal Fusion for Fine-Grained Classification of Breast Fibroadenoma and Phyllodes Tumors Medical VLM fusion for breast tumor classification; fine-grained multimodal clinical decision support
2026-07-02 Search-based Testing of Vision Language Models for In-Car Scene Understanding Search-based adversarial testing of VLMs for automotive safety; evaluation methodology contribution
2026-07-02 Efficient PEFT Methods with Adaptive Checkpointing for Vision Models and VLMs on Resource Constrained Consumer-GPUs PEFT comparison for VLMs on consumer GPUs; practical fine-tuning on constrained hardware
2026-07-02 SINA: A Fully Automated Circuit Schematic Image to Netlist Generator Using Artificial Intelligence VLM applied to circuit schematic-to-netlist; novel EDA domain grounding task with structured output
2026-07-02 EduArt: An educational-level benchmark for evaluating art history knowledge in large language models Domain-specific benchmark for art history in LLMs; benchmark methodology over synthetic baselines
2026-07-02 Evaluating Vision-Language Models as a Zero-Shot Learning Alternative to You Only Look Once and Optical Character Recognition for Nigerian License Plate Recognition Zero-shot VLM vs. YOLO+OCR for license plate recognition; narrow but concrete evaluation comparison
2026-07-02 RTE-FM-Dehazer: Radiative Transfer Equation Inspired Flow Matching for Real-World Image Dehazing
2026-07-02 VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon
2026-07-01 Information-Regularized Attention for Visual-Centric Reasoning ECCV 2026 ECCV; tackles hallucination, grounding, forgetting via attention regularization — core VLM reliability
2026-07-01 AdaBoosting Text Prompts for Vision-Language Models ECCV 2026 ECCV 2026; AdaBoosting text prompts improves zero-shot VLM classification, practical and reproducible
2026-07-01 LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives Non-contrastive VL pretraining without negatives — architectural shift away from CLIP-style objectives
2026-07-01 DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding ECCV ECCV; difficulty-adaptive routing for zero-shot video grounding, strong venue, novel routing idea
2026-07-01 Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning Decouples perception/reasoning for fine-grained VQA — addresses critical VLM failure mode
2026-07-01 Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts ECCV 2026 ECCV 2026; domain arithmetic for VLA one-shot adaptation, directly relevant to robot VLM deployment
2026-07-01 StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning ECCV 2026 ECCV 2026; stochastic turn depth for visual instruction tuning addresses training-inference discrepancy
2026-07-01 MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos Counterfactual spatial reasoning benchmark probes genuine VLM scene understanding beyond observation
2026-07-01 GenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language Models VLM for industrial anomaly detection+localization+explanation — practical grounded multimodal system
2026-07-01 Personalized Object Identification and Localization via In-Context Inference with Vision-Language Models In-context inference for personalized object localization — few-shot VLM grounding without finetuning
2026-07-01 Selective Test-Time Debiasing for CLIP via Reward Gating CLIP debiasing via reward gating; test-time, training-free; addresses VLM fairness at deployment
2026-07-01 Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs ECCV 2026 Tactile modality in MLLMs; mask-isolated alignment; ECCV 2026; fills physical-grounding gap
2026-07-01 What's Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models VLMs for occlusion-aware autonomous planning; novel selective attention over hidden agents
2026-07-01 DroneFINE: Domain-Aware Parameter-Efficient Fine-Tuning of Vision-Language Detectors for Drone Images Domain-aware PEFT for VLM detectors on UAV imagery; parameter-efficient adaptation recipe
2026-07-01 LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models Long-video quality understanding benchmark for LVLMs; temporal distortion evaluation gap
2026-07-01 Foundation Model-driven Key Anatomy Frame Selection for Blind-sweep Ultrasound Fetal Birth Weight Estimation MICCAI 2026 MICCAI 2026; foundation model for ultrasound frame selection; low-resource medical VLM
2026-07-01 Human-Centric Transferable Tactile Pre-Training for Dexterous Robotic Manipulation Tactile pre-training transferable to dexterous manipulation; cross-modal grounding for robotics
2026-07-01 Attribute-Prompted Kernel Hashing for Unsupervised Data-Efficient Cross-Modal Retrieval Unsupervised cross-modal hashing with attribute prompts; data-efficient retrieval baseline
2026-07-01 DroneIQA-VLE: Multi-Task Drone Image Quality Assessment via Vision-Language Ensemble VLM ensemble for drone image quality assessment; ICME 2026 challenge solution
2026-07-01 Discrete Diffusion Language Models for Interactive Radiology Report Drafting
2026-07-01 Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task ECCV 2026 Probes VLM vision vs. language bias on depth perception; diagnostic for building reliable VLMs
2026-07-01 CPG-PAD: Concept-Informed Prompts Guided Presentation Attack Detection Concept-informed prompting for face PAD; practical VLM prompting technique for security-critical vision
2026-07-01 ESC: Emotional Self-Correction for Reliable Vision-Language Models ECCV Emotional self-correction for VLMs without post-training; lightweight reliability fix applicable broadly
2026-07-01 GMO-E\(^2\)DIT: Grounded Multi-Operation Editing for E-Commerce Images Multi-operation grounded image editing for e-commerce; tests compositional VLM instruction following
2026-07-01 MIBE: Multi-subject Interaction Benchmark and Evaluator for Personalized Image Generation Multi-subject personalized image generation benchmark; evaluates identity-binding in generative VLMs
2026-07-01 Multi-modal Rail Crossing Safety Analysis Multimodal safety analysis fusing vision and structured accident reports; applied safety-critical domain
2026-07-01 Towards Developing a Multimodal Chat Assistant for University Stakeholders: RAG-based Approach RAG-based multimodal university assistant; practical retrieval-augmented VLM deployment case study

June 2026 (109)

Date Paper Venue Why selected
2026-06-30 One Video, One World: Turning Monocular Video into Physical 4D Scenes ECCV 2026 Saining Zhang; first training-free 4D mesh scene from single video; highly novel
2026-06-30 Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning Tat-Seng Chua; probes overestimate VLM grounding — critical eval result
2026-06-30 Harnessing Textual Refusal Directions for Multimodal Safety Nicu Sebe; refusal-direction safety for MLLMs without unsafe data; timely
2026-06-30 Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference Operator-level visual token skipping; directly cuts MLLM inference cost
2026-06-30 Evidence Triangulation for Multimodal Fact-Checking in the Wild Papadopoulos; multimodal fact-checking in the wild; real-world relevance
2026-06-30 HyperVLP: Enhancing Hierarchical Surgical Video-Language Pre-training in Hyperbolic Space Nassir Navab; surgical VLM pretraining in hyperbolic space; strong med contribution
2026-06-30 Language-Assisted Super-Resolution from Real-World Low-Resolution Patches Kyoung Mu Lee; language-guided real-world SR; creative cross-modal approach
2026-06-30 RCT: A Robot-Collected Touch-Vision-Language Dataset for Tactile Generalization Calandra; robot-collected tactile-vision-language dataset; enables tactile VLMs
2026-06-30 Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning Step-aware RL for medical VLM reasoning; fixes failure cascades in multimodal CoT
2026-06-30 Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning Dual-stream RL for token-sparse medical VLM reasoning; novel sparse-evidence approach
2026-06-30 Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity? Visual ambiguity detection for VLMs; essential for reliable deployed systems
2026-06-30 Technical Report of RoboSpatial Challenge at CVPR 2026: Selective Reasoning Activation and Reference-Frame Disambiguation for Embodied Spatial Reasoning CVPR 2026; embodied spatial reasoning benchmark + winning solution
2026-06-30 Xiaomi-GUI-0 Technical Report Industry-scale GUI agent from Xiaomi; directly applicable to agent building
2026-06-30 Reasoning-aware Speculative Decoding for Efficient Vision-Language-Action Models in Autonomous Driving Speculative decoding for VLA reasoning; practical latency reduction technique
2026-06-30 Think While You Map: Asynchronous Vision-Language Agents for Incremental 3D Scene Graphs ECCV 2026 ECCV 2026; asynchronous incremental 3D scene graphs for embodied agents
2026-06-30 Rethinking Foundation Model Collaboration: Enhancing Specialized Models through Proxy Task Reasoning FM + specialized model collaboration via proxy tasks; key agent architecture pattern
2026-06-30 Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue
2026-06-30 InstanceControl: Controllable Complex Image Generation without Instance Labeling
2026-06-30 3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance
2026-06-29 H-GRPO: Permutation-Invariant Reinforcement Learning for Grounded Visual Reasoning GRPO + grounded VLM reasoning; directly relevant to RL-based VLM improvement
2026-06-29 VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context ECCV 2026 Addresses the long-visual-context bottleneck that limits real-world VLM deployment
2026-06-29 MemLeak: Diagnosing Information Leaks in Multimodal Agent Memory Surprising finding on recoverable visual data after memory deletion; security implications
2026-06-29 See Only When Needed: Context-Aware Attention Intervention for Mitigating Hallucinations in LVLMs ECCV 2026 Snoek group; ECCV 2026; training-free hallucination mitigation without amplifying spurious regions
2026-06-29 Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision Dense embodied CoT supervision tackles the cross-embodiment bottleneck in VLA models
2026-06-29 From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Akata group; fundamental audit of whether VLMs truly use visual evidence
2026-06-29 Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders First sparse autoencoder analysis of cross-modal feature heterogeneity in VLMs
2026-06-29 GROW\(^2\): Grounding Which and Where for Robot Tool Use David Hsu; open-world affordance grounding for creative robot tool use with VLMs
2026-06-29 UnfoldArt: Zero-Shot Recovery of Full Articulated 3D Objects from Text or Image Laptev + Angela Dai; zero-shot articulated 3D recovery; crucial for embodied AI
2026-06-29 The Label Imitation Game: Turing Test Network for Zero-Shot Pseudo-Label Pruning ECCV 2026 Jason Corso; Turing-test method cleans zero-shot pseudo-labels; novel + practical
2026-06-29 Automating the Design of Embodied AgentArchitectures Automated embodied agent architecture design; changes how we build agent harnesses
2026-06-29 DOPD: Dual On-policy Distillation
2026-06-29 Unlocking the Visual Record of Materials Science: A Large-Scale Multimodal Dataset from Scientific Literature
2026-06-29 JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications
2026-06-28 ScAle: Attention Head Scaling as a Minimal Adapter for Spatial Reasoning in Vision Language Models ECCV 2026 Minimal efficient adapter for spatial reasoning, a known VLM weakness; Yanzhi Wang group
2026-06-28 MIRROR: Aligning Semantic Relations from Language to Image via Gromov--Wasserstein ECCV 2026 Fundamental structural alignment gap in MLLMs; novel Gromov-Wasserstein approach
2026-06-28 Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs Token merging for low-latency robotic VLM inference; practical efficiency contribution
2026-06-28 Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model Event camera fusion addresses RGB degradation in VLA; novel sensor modality for embodied AI
2026-06-28 Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation
2026-06-27 On Test-Time Scaling for Vision-Language Models ECCV 2026 Test-time scaling is the hottest topic; timely application to VLMs is high-impact
2026-06-27 HKVLM: Faithful Reasoning Grounding by Binding Language Queries to a Frozen Detector Practical binding of LM reasoning to frozen detector; addresses VLM faithfulness directly
2026-06-27 ViPSim: Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Models Consistent world models for evaluating VLA systems; critical embodied AI infrastructure
2026-06-26 Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models Mechanistic causal account of perception-knowledge conflict; VLM reliability foundation
2026-06-26 AirGroundBench: Probing Spatial Intelligence in Multimodal Large Models under Heterogeneous Multi-View Embodied Collaboration ECCV; spatial reasoning benchmark for MLLMs across heterogeneous multi-view embodied settings
2026-06-26 HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration ECCV 2026 ECCV; monocular→4D pipeline for scalable VLA training data at low cost
2026-06-26 MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving Dynamic token pruning for multi-view VLMs; practical efficiency for autonomous driving
2026-06-26 RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning ECCV 2026
2026-06-26 EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography
2026-06-26 Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots
2026-06-26 ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures
2026-06-26 SpatialUAV: Benchmarking Spatial Intelligence for Low-Altitude UAV Perception, Collaboration, and Motion
2026-06-26 Panoramic Scene Analysis: A Survey from Distortion-Aware Engineering to Sphere-Native Foundation Modeling
2026-06-26 Enhancing Numerical Prediction in LLMs via Smooth MMD Alignment
2026-06-26 ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion Retrieval
2026-06-26 VLM-Aware Meta-Optic Front-End Design for Frozen Vision-Language Models
2026-06-26 DataComp-VLM: Improved Open Datasets for Vision-Language Models DataComp series has proven impact; open VLM dataset benchmark is foundational for the whole field
2026-06-26 Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation Temporal visual reasoning evaluation gap; Alane Suhr; tests if VLMs understand dynamics
2026-06-26 Detecting Clinical Hallucinations in LVLMs via Counterfactual Visual Grounding Uncertainty Counterfactual visual grounding for clinical hallucination detection; important safety work
2026-06-26 Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models? Systematic VLA redundancy analysis; actionable pruning for efficient deployment
2026-06-25 From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP ECCV 2026 ECCV 2026; structural diagnostic separating language priors from true spatial reasoning in VLMs
2026-06-25 GAVEL: Grounded Caption Error Verification and Localization VLM hallucination detection+localization; directly useful for grounded VLM eval pipelines
2026-06-25 Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models ECCV 2026 ECCV 2026; self-evolving LMM training with visual token attention — unsupervised reasoning improvement
2026-06-25 Do Image Editing Models Understand Lighting? Nießner+Rother at ECCV; probes fundamental lighting understanding gap in generative VLMs
2026-06-25 Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision Structured stage+keyframe supervision for VLA fine-tuning; practical reproducible recipe
2026-06-25 Bridging Vision and Language Concepts through Optimal Transport Semantic Flow OT-based visual-textual concept alignment; directly improves interpretable VLM reasoning in CBMs
2026-06-25 HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models Systematic harmful-video benchmark for LMMs; fills critical safety evaluation gap for video VLMs
2026-06-25 Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models Comprehensive survey unifying adversarial attacks across text/vision/VLMs; useful robustness reference
2026-06-25 Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding ICML 2026 ICML; causal root-cause of LVLM hallucination beyond attention heuristics
2026-06-25 Large Language Model Teaches Visual Students: Cross-Modality Transfer of Fine-Grained Conceptual Knowledge ICML 2026 Yong Jae Lee, ICML; LLM→vision cross-modal fine-grained knowledge distillation
2026-06-25 Aloe-Vision: Robust Vision-Language Models for Healthcare MIDL 2026 MIDL; robustness-focused healthcare LVLM with scarce-data evaluation
2026-06-24 Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography ICML 2026 ICML; 3D CT VLP with hybrid encoding; directly buildable medical foundation recipe
2026-06-24 MRI2Rep: Autoregressive Structured Report Generation for 3D Liver MRI MICCAI 2026 MICCAI; Zongwei Zhou; autoregressive structured MRI reports end-to-end
2026-06-24 Beyond Visual Forensics: Auditing Multimodal Robustness for Synthetic Medical Image Detection MICCAI 2026 MICCAI; Pin-Yu Chen; audits VLM robustness on synthetic medical images
2026-06-24 Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation ECCV 2026 Bansal group (UNC, ECCV 2026); scene-graph eval for physical plausibility in T2V — new rigorous eval tool
2026-06-24 H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks ECCV 2026 ECCV 2026; attention-derived masks for pose-robust image editing with VLM-adjacent mechanisms
2026-06-24 Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity ECCV 2026 ECCV; benchmark exposing VLM retrieval failures under visual homogeneity at scale
2026-06-23 ActiveScope: Actively Seeking and Correcting Perception for MLLMs ICML 2026 ICML 2026; training-free fine-grained perception correction for MLLMs — drop-in applicable
2026-06-23 VisCritic: Visual State Comparison as Process Reward for GUI Agents ECCV 2026 ECCV 2026; visual-state process reward for GUI agents — directly useful for agent eval harnesses
2026-06-23 BioMedVR: Confusion-Aware Mixture-of-Prompt Experts for Biomedical Visual Reprogramming ECCV 2026 ECCV 2026; confusion-aware mixture-of-prompt experts for biomedical CLIP adaptation
2026-06-23 EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding ECCV 2026 Yin Li group (ECCV 2026); first streaming egocentric VLM benchmark — critical for real-time agent evaluation
2026-06-23 Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations ECCV 2026 ECCV 2026; first concept-annotation-grounded evaluation of SAE interpretability in VLMs
2026-06-23 Transition-Aware best-of-N sampling for Longitudinal Chest X-ray Reports
2026-06-22 CFPO: Counterfactual Policy Optimization for Multimodal Reasoning ICML 2026 ICML 2026; counterfactual causal RL for VLM reasoning—novel mechanism beyond standard GRPO/PPO
2026-06-22 Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views ECCV 2026 ECCV 2026; dense reward shaping for multi-view 3D VQA—novel RL signal for spatial chain-of-thought
2026-06-22 Brain-Adapter: A Dual-Stream Vision-Language MIL Framework for Comprehensive 3D CT Diagnosis of Acute Intracranial Pathologies MICCAI 2026 MICCAI 2026; dual-stream MIL VLM for 3D CT; directly applicable to medical foundation model work
2026-06-22 Mind the Heads: Topological Representation Alignment for Multimodal LLMs
2026-06-21 PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models ECCV 2026 ECCV 2026; intrinsic policy efficiency for VLA deployment; directly buildable
2026-06-18 Timage: A Generative Text-in-Image Paradigm for Fine-Tuning Vision-Language Models ECCV ECCV; text-in-image geometric anchoring directly addresses fine-grained spatial reasoning gap in MLLMs
2026-06-18 Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA MICCAI 2025 MICCAI 2025; confidence calibration mismatch in medical MLLMs creates direct misdiagnosis risk
2026-06-18 FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation ICML 2026 ICML 2026; systematic few-shot stress test of VLA adaptation with future-oriented conditioning
2026-06-18 Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology MICCAI 2026 MICCAI 2026; 1.2M CT/MR grounded pairs; scalable spatially grounded radiology VLM recipe
2026-06-17 Moving Beyond Diversity: Visual Token Pruning as Subspace Reconstruction for Efficient VLMs ECCV 2026 ECCV 2026; novel subspace-reconstruction reframing of token pruning; critical efficiency gain
2026-06-17 Semantic Robustness Certification for Vision-Language Models ICML ICML; first certified robustness framework for semantic distribution shifts in VLMs
2026-06-17 Hallucination Detection and Correction in Medical VLMs via Counter-Evidence Verification MICCAI 2026 MICCAI 2026; counter-evidence verification for hallucinations in high-stakes medical VLMs
2026-06-17 Bridging Creative Intent and Visual Quality: Creator-Driven Recurrent Video Generation with Agentic Feedback Loops ICML 2026 ICML 2026 acceptance; agentic feedback loop for coherent long-form video generation
2026-06-17 Language-Instructed Vision Embeddings for Controllable and Generalizable Perception ICLR 2026 ICLR 2026; novel paradigm using language to steer vision embeddings for controllable perception
2026-06-16 Visuals Lie, Consistency Speaks: Disentangling Spatial Attention from Reliability in Vision-Language Models ICLR 2026 ICLR 2026; challenges Attention-Confidence Assumption; exposes core VLM reliability blind spot
2026-06-16 MODE-RAG: Manifold Outlier Diagnosis and Energy-based Retrieval-Augmented Generation Evaluation ACL 2026 ACL 2026; multimodal RAG evaluation exposing cross-modal hallucination and sycophancy failures
2026-06-16 Evaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in Meetings INTERSPEECH 2026
2026-06-16 Domain Generalizable Adaptation of 3D Vision-Language Models via Regularized Fine-Tuning TMLR TMLR; regularized fine-tuning for domain-generalizable 3D vision-language model adaptation
2026-06-15 Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering EMNLP 2026
2026-06-15 XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models INTERSPEECH 2026
2026-06-14 Learning Directional Semantic Transitions for Longitudinal Chest X-ray Analysis MICCAI 2026
2026-06-14 The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages ICML 2026 ICML 2026; inherited truthful heads across model lineages offers new grounding tool against hallucination
2026-06-13 Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams ICML
2026-06-13 Think Less, Act Early: Reinforced Latent Reasoning with Early Exit in Vision-Language-Action Models ICML 2026 ICML 2026; reinforced latent reasoning with early exit cuts VLA compute and error propagation
2026-06-12 CORA: Analyzing and bridging thinking-answer gap in Multimodal RLVR via Consistency-Oriented Reasoning Alignment EMNLP 2026
2026-06-12 One Layer's Trash is Another Layer's Treasure: Adaptive Layer-wise Visual Token Selection in LVLMs CVPR 2026
2026-06-11 Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality ACL 2026
2026-06-11 OpenMedQ: Broad Open Pretraining for Medical Vision-Language Models MIDL

May 2026 (20)

Date Paper Venue Why selected
2026-05-30 Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding ICML 2026 ICML 2026; decomposes on-policy distillation gradients for VL visual grounding
2026-05-30 Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs CVPR 2026 CVPR 2026; PRISM: principle-aware multi-scale evaluation of visual designs
2026-05-29 Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization ICML 2026 ICML 2026; visual contrastive DPO for multimodal hallucination mitigation
2026-05-29 Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models ICML 2026 Mrinmaya Sachan (ETH); ICML 2026; systematic study of test-time compute in VLMs
2026-05-29 How can embedding models bind concepts? ICML 2026 Seong Joon Oh (Tübingen); ICML 2026; concept binding failure in CLIP-style models
2026-05-29 Immuno-VLM: Immunizing Large Vision-Language Models via Generative Semantic Antibodies for Open-World Trustworthiness ICML 2026 ICML 2026; immunizing VLMs against open-world adversarial semantic vulnerabilities
2026-05-28 Grounded 3D-Aware Spatial Vision-Language Modeling CVPR 2026 Ligeng Zhu (MIT/NVIDIA); CVPR 2026; unified 2D+3D grounding in spatial VLMs
2026-05-28 Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning CVPR 2026 Yi Ma (Berkeley) + Joseph Tighe (Meta); CVPR 2026; 3D spatial priors improve VLM geometry
2026-05-28 Unveiling the Visual Counting Bottleneck in Vision-Language Models ICML 2026 Mrinmaya Sachan (ETH); ICML 2026; reveals visual counting extrapolation bottleneck
2026-05-28 Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning ICML 2026 ICML 2026; breaks tail alignment in CLIP for cross-domain few-shot learning
2026-05-27 Self-Prophetic Decoding to Unlock Visual Search in LVLMs ICML 2026 ICML 2026; self-prophetic decoding enables visual search in LVLMs
2026-05-24 Interpretability Transfer from Language to Vision via Sparse Autoencoders ICML 2026 ICML 2026; sparse autoencoders transfer language interpretability to vision
2026-05-24 Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation ICML 2026 ICML 2026; in-depth analysis + simple mitigation of language bias in LVLMs
2026-05-22 CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception ICML 2026 ICML 2026; cognitive visual search for high-resolution image perception in MLLMs
2026-05-22 Multimodal Distribution Matching for Vision-Language Dataset Distillation CVPR 2026 CVPR 2026; multimodal distribution matching for vision-language dataset distillation
2026-05-22 Turning Adaptation into Assets: Cross-Domain Bridging for Online Vision-Language Navigation ICML 2026 ICML 2026; test-time adaptation for VLN under non-stationary distribution shifts
2026-05-22 SPACENUM: Revisiting Spatial Numerical Understanding in VLMs
2026-05-21 Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? CVPR 2026 CVPR 2026; reveals VLM benchmark scores don't require grounded visual evidence
2026-05-21 From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model ICML 2026 ICML 2026; learning generalized behavioral representations for VLA distribution shift
2026-05-21 AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation CVPR 2026 CVPR 2026; self-awareness in VLN agents for reasoning over visual environments

April 2026 (33)

Date Paper Venue Why selected
2026-04-30 Improving Calibration in Test-Time Prompt Tuning for Vision-Language Models via Data-Free Flatness-Aware Prompt Pretraining CVPR 2026 Flatness-aware pretraining improves calibration in VLM test-time prompt tuning
2026-04-29 Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning CVPR Qualitative reasoning mitigates optical illusion failures in frozen VLMs; CVPR
2026-04-28 Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models CVPR 2026 Novel prefill-time intervention eliminates LVLM hallucinations; inference-time method
2026-04-27 Improving Vision-language Models with Perception-centric Process Reward Models CVPR Process-level reward models for VLM RLVR; finer than outcome-only supervision
2026-04-27 LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models ICLR 2026 Attention-based token pruning rethought; attention scores don't indicate importance
2026-04-27 Jailbreaking Frontier Foundation Models Through Intention Deception CVPR 2026 Jailbreaks via intention deception bypass safety training; CVPR 2026 novel vector
2026-04-27 ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning ICML 2026 Rebuilds 3D spatial reasoning evaluation exposing VLM annotation biases; ICML 2026
2026-04-27 Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-System ACL 2026 Asynchronous coarse-to-fine dual system balances VLA learning equilibrium; ACL 2026
2026-04-27 SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs CVPR 2026 Soft modality routing in MoE-VLMs; first study of modality-specific expert allocation
2026-04-27 World-R1: Reinforcing 3D Constraints for Text-to-Video Generation ICML 2026 RL reinforces 3D geometric consistency in text-to-video generation; ICML 2026
2026-04-25 Exploring Hierarchical Consistency and Unbiased Objectness for Open-Vocabulary Object Detection CVPR 2026 Hierarchical consistency with unbiased objectness for CLIP open-vocab detection
2026-04-24 DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning CVPR 2026 Background/question-aware token pruning for efficient document VQA; CVPR 2026
2026-04-23 Ramen: Robust Test-Time Adaptation of Vision-Language Models with Active Sample Selection CVPR 2026 Active sample selection for robust CLIP test-time adaptation; CVPR 2026
2026-04-22 Building a Precise Video Language with Human-AI Oversight CVPR 2026 CVPR 2026 scalable video-language human-AI oversight; datasets and benchmarks
2026-04-22 Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation ACL 2026 Hallucination-free fine-tuning without performance degradation; ACL 2026
2026-04-22 Breaking the Illusion: When Positive Meets Negative in Multimodal Decoding CVPR 2026 Training-free positive-negative decoding reduces object hallucination; CVPR 2026
2026-04-22 OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model ACL 2026 Olympiad-level multi-image reasoning benchmark for frontier VLMs; ACL 2026
2026-04-22 Evian: Towards Explainable Visual Instruction-tuning Data Auditing ACL 2026 Explainable visual instruction data auditing for LVLM training quality
2026-04-21 CAPruner: Conceptual-Adjacent Scene Graph Pruner for Enhancing 3D Spatial Reasoning of Large Language Models ACL 2026 Scene graph pruning removes irrelevant anchors for LLM 3D spatial reasoning
2026-04-21 Environmental Understanding Vision-Language Model for Embodied Agent CVPR Environmental understanding VLM improves instruction-following embodied agent generalization
2026-04-20 Hierarchically Robust Zero-shot Vision-language Models CVPR 26 Hierarchically robust VLM fine-tuning preserves zero-shot under adversarial attack
2026-04-20 S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models ACL 2026 Hardness-aware DPO closes multi-image reasoning gap in VLM alignment
2026-04-20 Test-Time Perturbation Learning with Delayed Feedback for Vision-Language-Action Models CVPR 2026 Test-time perturbation learning corrects VLA trajectory overfitting; CVPR 2026
2026-04-20 Enhancing Continual Learning of Vision-Language Models via Dynamic Prefix Weighting CVPR 2026 Dynamic prefix weighting for domain-class incremental VLM learning; CVPR 2026
2026-04-20 From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models ACL 2026 Neuron-level causal attribution reveals cross-task VLM mechanisms; ACL 2026
2026-04-19 More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage ACL 2026 High visual fidelity hinders abstract VLM reasoning; semiotic gap study
2026-04-19 Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception ACL 2026 Cold-start agentic VLM grounding without expensive supervised trajectory tuning
2026-04-18 EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling CVPR 2026 Evolutionary semantic labeling for visual token compression; CVPR 2026 efficiency
2026-04-18 SIF: Semantically In-Distribution Fingerprints for Large Vision-Language Models CVPR 2026 In-distribution semantic fingerprints for VLM IP protection; CVPR 2026
2026-04-17 Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow CVPR 2026 Adaptive info flow fixes VLMs that see but don't perceive; CVPR 2026
2026-04-17 TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Models CVPR 2026 Test-time textual learning improves CLIP-based OOD detection without training
2026-04-17 Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization CVPR World-scale geolocalization analysis reveals systematic VLM perception failures; CVPR
2026-04-15 One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding NEURIPS 2025 Extreme compression to one token per frame enables long-video VLMs; NeurIPS 2025

March 2026 (43)

Date Paper Venue Why selected
2026-03-31 Robust Multimodal Safety via Conditional Decoding ACL 2026 ACL 2026; conditional decoding restores safety alignment in multimodal LLMs
2026-03-31 A Comprehensive Information-Decomposition Analysis of Large Vision-Language Models ICLR 2026 ICLR 2026; information-decomposition reveals if VLMs use true multimodal fusion or shortcuts
2026-03-31 Hierarchical Pre-Training of Vision Encoders with Large Language Models CVPR CVPR; hierarchical joint pre-training tightly couples vision encoders with LLMs
2026-03-31 AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Models CVPR 2026 CVPR 2026; alignment-guided fine-tuning preserves zero-shot adversarial robustness in VLMs
2026-03-30 Explaining CLIP Zero-shot Predictions Through Concepts CVPR 2026 CVPR 2026; concept bottleneck explanations make CLIP zero-shot predictions transparent
2026-03-30 Learning Multi-View Spatial Reasoning from Cross-View Relations CVPR 2026 CVPR 2026; cross-view relational learning for multi-view spatial reasoning in VLMs
2026-03-30 ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation CVPR 2026 CVPR 2026; comprehensive real-world benchmark for VLA manipulation evaluation
2026-03-30 FlowHijack: A Dynamics-Aware Backdoor Attack on Flow-Matching Vision-Language-Action Models CVPR 2026 CVPR 2026; first backdoor attack targeting flow-matching VLA models like π0
2026-03-30 \(AutoDrive\text{-}P^3\): Unified Chain of Perception-Prediction-Planning Thought via Reinforcement Fine-Tuning ICLR 2026 ICLR 2026; chain of perception-prediction-planning with RL for autonomous driving VLMs
2026-03-29 On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language Models CVPR 2026 CVPR 2026; drift-aware dynamic MoE token assignment for continual multimodal learning
2026-03-28 Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models CVPR 2026 CVPR 2026; diagnoses root cause and mitigates hallucination in multimodal chain-of-thought
2026-03-28 Diagnosing and Repairing Unsafe Channels in Vision-Language Models via Causal Discovery and Dual-Modal Safety Subspace Projection CVPR 2026 CVPR 2026; causal discovery identifies unsafe channels; dual-modal safety subspace repair
2026-03-28 Structural Graph Probing of Vision-Language Models CVPR CVPR; structural graph probing reveals neural topology organization in VLMs
2026-03-28 SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning CVPR 2026 CVPR 2026; layered geometry-language fusion for reliable 3D spatial VLM reasoning
2026-03-28 ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding CVPR 2026 CVPR 2026; million-scale high-quality chart dataset for geometry-numerical-language reasoning
2026-03-28 Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark CVPR 2026 CVPR 2026; scene-aware benchmark reveals catastrophic forgetting in long-video VLMs
2026-03-28 VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation CVPR 2026 CVPR 2026; VLM-driven spatiotemporal reasoning for referring video object segmentation
2026-03-28 Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving ECCV 2026 ECCV 2026; interleaved world modeling and planning unifies AD perception and action
2026-03-28 MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence CVPR 2026 CVPR 2026; medical VLM with lesion detection, tracking, and visual explainability
2026-03-27 A Provable Energy-Guided Test-Time Defense Boosting Adversarial Robustness of Large Vision-Language Models CVPR CVPR; provable energy-guided test-time defense for VLM adversarial robustness
2026-03-27 HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models CVPR 2026 CVPR 2026; benchmark diagnosing fine-grained hand spatial reasoning failures in VLMs
2026-03-27 Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining ICLR 2026 ICLR 2026; disentangled forward/inverse dynamics pretraining resolves VLA misalignment
2026-03-27 Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding CVPR 2026 CVPR 2026; diffusion VLMs for GUI grounding challenge autoregressive model dominance
2026-03-27 FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants CVPR 2026 CVPR 2026; fairness-aware PEFT for multimodal LLMs in clinical high-stakes settings
2026-03-26 No Hard Negatives Required: Concept Centric Learning Leads to Compositionality without Degrading Zero-shot Capabilities of Contrastive Models CVPR 2026 CVPR 2026; concept-centric learning enables compositionality without zero-shot degradation
2026-03-26 HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models CVPR 2026 CVPR 2026; hierarchical 3D spatial understanding with perception-to-reasoning pipeline
2026-03-26 Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized Driving CVPR 2026 CVPR 2026; preference alignment personalizes VLA driving behavior to individual styles
2026-03-26 Demographic Fairness in Multimodal LLMs: A Benchmark of Gender and Ethnicity Bias in Face Verification CVPR 2026 CVPR 2026; demographic fairness benchmark for MLLM-based face verification
2026-03-26 MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models CVPR 2026 CVPR 2026; GRPO-based reinforcement learning improves MoE routing in VLMs
2026-03-25 From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector Decomposition CVPR 2026 CVPR 2026; data-free CLIP interpretability via SVD; no activation data required
2026-03-25 Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification CVPR 2026 CVPR 2026; attention imbalance rectification directly mitigates object hallucination
2026-03-25 LensWalk: Agentic Video Understanding by Planning How You See in Videos CVPR 2026 CVPR 2026; agentic lens planning for dense temporal video understanding
2026-03-25 Unleashing Vision-Language Semantics for Deepfake Video Detection CVPR 2026 CVPR 2026; VLM semantic priors boost generalizable deepfake video detection
2026-03-24 VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions CVPR 2026 CVPR 2026; sparse dynamic VL interactions for LVLM efficiency without information bottleneck
2026-03-24 Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding CVPR 2026 CVPR 2026; identifies instruction-relevant image regions for information-dense VLM inputs
2026-03-23 Language Models Can Explain Visual Features via Steering CVPR 2026 CVPR 2026; LLMs explain SAE visual features autonomously via causal steering
2026-03-23 Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language Models CVPR 2026 CVPR 2026; null-space projection steering principled defense against visual jailbreaks
2026-03-23 Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models CVPR 2026 CVPR 2026; hyperbolic embeddings capture part-to-whole hierarchy in VL alignment
2026-03-23 Which Concepts to Forget and How to Refuse? Decomposing Concepts for Continual Unlearning in Large Vision-Language Models CVPR 2026 CVPR 2026; concept decomposition enables continual selective unlearning in LVLMs
2026-03-22 Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models CVPR 2026 CVPR 2026; uncertainty-aware distillation balances teacher vs. data guidance in MLLMs
2026-03-21 Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models CVPR 2026 CVPR 2026; predictive regularization prevents visual representation degradation during LLM training
2026-03-20 When Negation Is a Geometry Problem in Vision-Language Models CVPR CVPR; negation failure in CLIP reframed as geometry problem with novel fix
2026-03-20 PersonaVLM: Long-Term Personalized Multimodal LLMs CVPR 2026 CVPR 2026; multi-turn long-term personalized preference alignment in multimodal LLMs

January 2026 (205)

Date Paper Venue Why selected
2026-01-01 Visual symbolic mechanisms: Emergent symbol processing in Vision Language Models ICLR 2026 Yoshua Bengio (Mila); ICLR 2026; symbolic binding mechanisms in VLMs
2026-01-01 ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning ICLR 2026 Linjie Li (Microsoft); ICLR 2026; multimodal interleaved chain-of-thought emergent properties
2026-01-01 Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation ICLR 2026 Serge Belongie (Cornell Tech); ICLR 2026; comprehensive fine-grained LVLM evaluation
2026-01-01 Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning ICLR 2026 Serena Yeung-Levy (Stanford); ICLR 2026; evolving tool libraries for 3D spatial reasoning
2026-01-01 Scaling up Memory for Robotic Control via Experience Retrieval ICLR 2026 Chelsea Finn (Stanford); ICLR 2026; scaling memory for robot policies via retrieval
2026-01-01 WorldGym: World Model as An Environment for Policy Evaluation ICLR 2026 Percy Liang (Stanford); ICLR 2026; world-model as policy evaluation environment
2026-01-01 Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning ICLR 2026 Tsung-Yi Lin + Yen-Chen Lin; ICLR 2026; fine-tuning video models for visuomotor control
2026-01-01 HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model ICLR 2026 Renrui Zhang (CUHK); ICLR 2026; hybrid diffusion+autoregressive VLA model
2026-01-01 InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models ICLR 2026 InternVL team (Zhe Chen); ICLR 2026; large-scale spatial reasoning dataset for VLMs
2026-01-01 Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving ICLR 2026 Hang Zhao (Tsinghua); ICLR 2026; discrete diffusion VLA for end-to-end autonomous driving
2026-01-01 Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs ICLR 2026 Bohyung Han (SNU); ICLR 2026; mechanistic analysis of information flow in VideoLLMs
2026-01-01 Zebra-CoT: A Dataset for Interleaved Vision-Language Reasoning ICLR 2026 ICLR 2026; Zebra-CoT large dataset for interleaved vision-language reasoning
2026-01-01 Self-Improving Vision-Language-Action Models with Data Generation via Residual RL ICLR 2026 ICLR 2026; self-improving VLA via residual RL, reduces reliance on human demos
2026-01-01 Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models ICLR 2026 ICLR 2026; test-time matching unlocks compositional reasoning in frontier VLMs
2026-01-01 Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded Verification ICLR 2026 ICLR 2026; agreement bias mitigation via self-grounded verification in MLLMs
2026-01-01 Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization ICLR 2026 ICLR 2026; importance-sampling multi-negative multimodal DPO for VLMs
2026-01-01 Hallucination-aware Intermediate Representation Edit in Large Vision-Language Models ICLR 2026 ICLR 2026; intermediate representation editing for hallucination mitigation
2026-01-01 Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models ICLR 2026 ICLR 2026; dynamic activation steering identifies truthfulness/visual directions
2026-01-01 Matched Data, Better Models: Target Aligned Data Filtering with Sparse Autoencoders ICLR 2026 ICLR 2026; SAE-based target-aligned data filtering for VLM pretraining
2026-01-01 Label-Free Mitigation of Spurious Correlations in VLMs using Sparse Autoencoders ICLR 2026 ICLR 2026; label-free SAE-based spurious correlation removal in VLMs
2026-01-01 MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse ICLR 2026 ICLR 2026; first RL framework for 3D spatial reasoning in VLMs for metaverse
2026-01-01 Pursuing Minimal Sufficiency in Spatial Reasoning ICLR 2026 ICLR 2026; minimal sufficiency framework for spatial reasoning bottlenecks
2026-01-01 Do 3D Large Language Models Really Understand 3D Spatial Relationships? ICLR 2026 ICLR 2026; exposes 3D LLMs do not truly understand spatial relationships
2026-01-01 SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs ICLR 2026 ICLR 2026; SpinBench: perspective-taking diagnostic for spatial reasoning in VLMs
2026-01-01 ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis ICLR 2026 ICLR 2026; RLVR agentic data synthesis for complex video reasoning
2026-01-01 VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding ICLR 2026 Bhiksha Raj (CMU); ICLR 2026; bootstrapping scalable MLLM-as-judge for video
2026-01-01 VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning? ICLR 2026 ICLR 2026; VideoReasonBench: vision-centric complex video reasoning benchmark
2026-01-01 FOCUS: Efficient Keyframe Selection for Long Video Understanding ICLR 2026 ICLR 2026; FOCUS efficient keyframe selection for hour-long video MLLMs
2026-01-01 Memento: Toward an All-Day Proactive Assistant for Ultra-Long Streaming Video ICLR 2026 ICLR 2026; Memento: proactive assistant for ultra-long streaming video
2026-01-01 SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning ICLR 2026 ICLR 2026; SimpleVLA-RL: scaling VLA training via reinforcement learning
2026-01-01 X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model ICLR 2026 ICLR 2026; X-VLA: soft-prompted transformer for cross-embodiment VLA
2026-01-01 Interleave-VLA: Enhancing Robot Manipulation with Image-Text Interleaved Instructions ICLR 2026 ICLR 2026; interleaved image-text instructions improve robot manipulation
2026-01-01 PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model ICLR 2026 ICLR 2026; PixelVLA: pixel-level understanding in vision-language-action models
2026-01-01 HAMLET: Switch Your Vision-Language-Action Model into a History-Aware Policy ICLR 2026 ICLR 2026; HAMLET: history-aware policy for VLA manipulation models
2026-01-01 TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models ICLR 2026 ICLR 2026; TwinVLA: data-efficient bimanual manipulation via VLAs
2026-01-01 Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation ICLR 2026 ICLR 2026; Embodied-R1: unified pointing representation for generalist manipulation
2026-01-01 From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation ICLR 2026 ICLR 2026; bridging reasoning to action for novel robotic manipulation scenarios
2026-01-01 Hybrid Training for Vision-Language-Action Models ICLR 2026 ICLR 2026; hybrid VLA training combining CoT with diffusion-based actions
2026-01-01 Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denosing Diffusion Process ICLR 2026 ICLR 2026; unified discrete diffusion VLA with future image understanding
2026-01-01 Verifier-free Test-Time Sampling for Vision-Language-Action Models ICLR 2026 ICLR 2026; verifier-free test-time scaling for VLA precision tasks
2026-01-01 Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models ICLR 2026 ICLR 2026; identifies fundamental bottlenecks in VLM safety fine-tuning
2026-01-01 Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning ICLR 2026 ICLR 2026; spurious correlations undermine safety fine-tuning; machine unlearning fix
2026-01-01 ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks ICLR 2026 ICLR 2026; adaptive plug-and-play red-teaming agents against multimodal models
2026-01-01 Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots ICLR 2026 ICLR 2026; structural input perturbations reveal multimodal alignment blind spots
2026-01-01 AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference Optimization ICLR 2026 ICLR 2026; AdPO: preference optimization for adversarial robustness in LVLMs
2026-01-01 BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning ICLR 2026 ICLR 2026; BEAT: contrastive visual backdoor attacks on VLM embodied agents
2026-01-01 Do Vision-Language Models Respect Contextual Integrity in Location Disclosure? ICLR 2026 ICLR 2026; VLMs violate contextual integrity in location privacy disclosure
2026-01-01 TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models ICLR 2026 ICLR 2026; TrustGen: dynamic benchmarking platform for generative model trustworthiness
2026-01-01 Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation ICLR 2026 ICLR 2026; Kaleidoscope: massively multilingual in-language VLM evaluation
2026-01-01 RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding ICLR 2026 ICLR 2026; RAVENEA: retrieval-augmented visual culture understanding benchmark
2026-01-01 Vision Language Models are Biased ICLR 2026 ICLR 2026; systematic evidence that VLMs exhibit strong prior-knowledge biases
2026-01-01 MMReD: a Cross-Modal Benchmark for Dense Context Reasoning ICLR 2026 ICLR 2026; MMReD: dense multi-modal context reasoning over long inputs
2026-01-01 ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation ICLR 2026 ICLR 2026; ChartGalaxy: large-scale infographic chart understanding and generation
2026-01-01 GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation ICLR 2026 ICLR 2026; GeoBench: hierarchical geometric reasoning evaluation for VLMs
2026-01-01 Unified Vision–Language Modeling via Concept Space Alignment ICLR 2026 ICLR 2026; V-SONAR: extends multilingual text embeddings to 1500-language vision
2026-01-01 VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models ICLR 2026 ICLR 2026; VisCodex: unified multimodal code generation merging vision and code models
2026-01-01 DiVE-k: DIFFERENTIAL VISUAL REASONING FOR FINE-GRAINED IMAGE RECOGNITION ICLR 2026 Ram Nevatia (USC); ICLR 2026; differential visual reasoning for fine-grained recognition
2026-01-01 Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning ICLR 2026 ICLR 2026; Sparse CLIP co-optimizes interpretability and performance in CLIP
2026-01-01 Post-hoc Probabilistic Vision-Language Models ICLR 2026 ICLR 2026; post-hoc probabilistic VLMs with calibrated uncertainty in joint space
2026-01-01 Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation ICLR 2026 ICLR 2026; speculative drafting approach for information-intensive visual reasoning
2026-01-01 Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning ICLR 2026 ICLR 2026; game-based multimodal verifiable data boosts VLM general reasoning
2026-01-01 ProxyThinker: Test-Time Guidance through Small Visual Reasoners ICLR 2026 ICLR 2026; small proxy VLMs guide test-time reasoning of large VLMs cheaply
2026-01-01 Vision-SR1: Self-Rewarding Vision-Language Model via Reasoning Decomposition and Multi-Reward Policy Optimization ICLR 2026 ICLR 2026; Vision-SR1: self-rewarding VLM via multi-reward policy optimization
2026-01-01 ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Models ICLR 2026 ICLR 2026; ViPER: self-evolution for fine-grained visual perception in VLMs
2026-01-01 Vision-Zero: Scalable VLM Self-Evolution via Multi-Agent Self-Play ICLR 2026 ICLR 2026; Vision-Zero: multi-agent self-play for label-free VLM self-evolution
2026-01-01 CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs ICLR 2026 ICLR 2026; attention distillation from large to small MLLMs for compositional reasoning
2026-01-01 Imitating the Truth: Attention-aware Truth-Guided Enhancement for Hallucination Mitigation in Large Vision-Language Models ICLR 2026 ICLR 2026; attention-aware truth-guided enhancement for LVLM hallucination
2026-01-01 Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs ICLR 2026 ICLR 2026; probes disconnect between visual attention and answer correctness in VLMs
2026-01-01 AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models ICLR 2026 ICLR 2026; empirical study of attention+diversity for adaptive visual token pruning
2026-01-01 LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models ICLR 2026 ICLR 2026; rethinks attention-based token pruning in VLMs
2026-01-01 Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity ICLR 2026 ICLR 2026; synergistic importance-diversity token compression for VLMs
2026-01-01 WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models ICLR 2026 ICLR 2026; weighted SVD for fast low-precision VLM execution
2026-01-01 Embodied Navigation Foundation Model ICLR 2026 ICLR 2026; foundation model for embodied navigation leveraging VLM priors
2026-01-01 CitySeeker: How Do VLMs Explore Embodied Urban Navigation with Implicit Human Needs? ICLR 2026 ICLR 2026; VLMs interpreting implicit human needs in urban embodied navigation
2026-01-01 Talking Points: Describing and Localizing Pixels ICLR 2026 ICLR 2026; Talking Points: pixel-precise keypoint comprehension via natural language
2026-01-01 Decomposition of Concept-Level Rules in Visual Scenes ICLR 2026 ICLR 2026; decomposing concept-level visual rules for compositional scene parsing
2026-01-01 DaVinci: Reinforcing Visual-Structural Syntax in MLLMs for Generalized Scientific Diagram Parsing ICLR 2026 ICLR 2026; DaVinci: RL-driven visual-structural syntax for scientific diagram parsing
2026-01-01 Mordal: Automated Pretrained Model Selection for Vision Language Models ICLR 2026 ICLR 2026; Mordal: automated pretrained model selection for building VLMs
2026-01-01 EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing ICLR 2026 ICLR 2026; EditReward: human-aligned reward model for instruction-guided image editing
2026-01-01 VLMgineer: Vision-Language Models as Robotic Toolsmiths ICLR 2026 ICLR 2026; VLMs as robotic toolsmiths — tool design as measure of physical intelligence
2026-01-01 VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision–Language Models ICLR 2026 ICLR 2026; VITA: test-time adaptation of VLMs as zero-shot goal-conditioned value functions
2026-01-01 DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning ICLR 2026 RL incentivizes visual thinking in VLMs; novel image-based reasoning paradigm
2026-01-01 VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use ICLR 2026 VLMs learn multimodal tool use via RL; key agentic VLM capability
2026-01-01 Revisiting Multimodal Positional Encoding in Vision–Language Models ICLR 2026 Comprehensive multimodal RoPE analysis; foundational for VLM positional encoding
2026-01-01 To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models ICLR 2026 Mechanistic analysis of visual information pathways and attention sinks in LVLMs
2026-01-01 Generative Universal Verifier as Multimodal Meta-Reasoner ICLR 2026 Generative universal verifier enables VLM meta-reasoning and self-refinement
2026-01-01 Cross-Modal Redundancy and the Geometry of Vision–Language Embeddings ICLR 2026 Geometric probe of VL joint embedding space via cross-modal redundancy
2026-01-01 OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning ICLR 2026 OneTwoVLA: unified VLA with adaptive reasoning-acting switching for robotics
2026-01-01 Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting ICLR 2026 Fine-tuning VLMs into VLAs without catastrophic forgetting; solves core challenge
2026-01-01 Rethinking Causal Mask Attention for Vision-Language Inference ICLR 2026 Rethinks causal masking in autoregressive VLMs; non-trivial architectural insight
2026-01-01 PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models ICLR 2026 Positional information preserved during token merging prevents spatial degradation
2026-01-01 GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models ICLR 2026 Test-time safety alignment for LVLMs; input detection plus output steering
2026-01-01 Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models ICLR 2026 Universal transferable jailbreak attacks on VLMs; multimodal attack surface analysis
2026-01-01 From Sure" toSorry": Detecting Jailbreak in Large Vision Language Model via JailNeurons ICLR 2026 JailNeurons: fast neuron-based jailbreak detection for LVLMs; ICLR 2026
2026-01-01 Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP ICLR 2026 Mechanistic defense against typographic attacks in CLIP; ICLR 2026 analysis
2026-01-01 Identifying Robust Neural Pathways: Few-Shot Adversarial Mask Tuning for Vision-Language Models ICLR 2026 Adversarial mask tuning identifies robust neural pathways in VLMs
2026-01-01 Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness ICLR 2026 Inference-time compute improves OOD robustness in VLMs; ICLR 2026 analysis
2026-01-01 Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes ICLR 2026 Egocentric multi-view spatial reasoning benchmark; VLMs struggle with 3D
2026-01-01 SpatiaLab: Can Vision–Language Models Perform Spatial Reasoning in the Wild? ICLR 2026 In-the-wild spatial reasoning benchmark challenges current VLM spatial capabilities
2026-01-01 OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models ICLR 2026 Comprehensive 3D spatial reasoning benchmark with diverse relationship categories
2026-01-01 IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs ICLR 2026 South-Asian cultural VLM benchmark exposes Western-centric evaluation blindspots
2026-01-01 Contamination Detection for VLMs Using Multi‑Modal Semantic Perturbations ICLR 2026 Multi-modal semantic perturbations detect VLM benchmark data contamination
2026-01-01 IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs ICLR 2026 Image-grounded video benchmark tests VLMs on cross-modal video comprehension
2026-01-01 Spotlight on Token Perception for Multimodal Reinforcement Learning ICLR 2026 Token-level perceptual spotlight improves RLVR visual reasoning in VLMs
2026-01-01 SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start ICLR 2026 Self-distilled cold start decouples visual/language learning for RLVR VLMs
2026-01-01 P\(^2\)-DPO: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization ICLR 2026 Calibration DPO grounds hallucination in perceptual processing; preference learning
2026-01-01 Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation ICLR 2026 Action-aware dynamic pruning enables efficient VLA inference; ICLR 2026
2026-01-01 villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models ICLR 2026 villa-X: enhanced latent action modeling for VLA models; architectural advance
2026-01-01 Robust Fine-tuning of Vision-Language-Action Robot Policies via Parameter Merging ICLR 2026 Parameter merging preserves generalist VLA skills during task-specific fine-tuning
2026-01-01 Memory-Free Continual Learning with Null Space Adaptation for Zero-Shot Vision-Language Models ICLR 2026 Null space adaptation enables memory-free continual VLM learning; preserves zero-shot
2026-01-01 Enhanced Continual Learning of Vision-Language Models with Model Fusion ICLR 2026 Model fusion enhances continual VLM learning; mitigates catastrophic forgetting
2026-01-01 KeepLoRA: Continual Learning with Residual Gradient Adaptation ICLR 2026 KeepLoRA: residual gradient adaptation balances VLM plasticity and stability
2026-01-01 Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard Problems ICLR 2026 Real-world Bongard problems test abstract visual concept learning in VLMs
2026-01-01 PoSh: Using Scene Graphs to Guide LLMs-as-a-Judge for Detailed Image Descriptions ICLR 2026 Scene graph-guided LLM-as-judge improves detailed image description evaluation
2026-01-01 K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge ICLR 2026 K-Sort corrects VLM-as-judge bias; scalable visual generation preference evaluation
2026-01-01 Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory ICLR 2026 Multimodal Item Response Theory enables rigorous VLM reasoning benchmark design
2026-01-01 DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection ICLR 2026 DeCo-DETR decoupled cognition achieves efficient open-vocabulary detection; ICLR 2026
2026-01-01 Fantastic Tractor-Dogs and How Not to Find Them With Open-Vocabulary Detectors ICLR 2026 OVDs produce confident false positives on negative images; critical deployment flaw
2026-01-01 VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning ICLR 2026 RL temporal focusing on relevant frames enables efficient long-video VLM reasoning
2026-01-01 CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval ICLR 2026 Fine-grained video captioning and retrieval benchmark reveals VLM temporal limits
2026-01-01 STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning ICLR 2026 RL reduces hallucination in spatial-temporal video grounding; ICLR 2026
2026-01-01 Enhancing Multi-Image Understanding through Delimiter Token Scaling ICLR 2026 Delimiter token scaling resolves cross-image leakage in multi-image VLMs
2026-01-01 QLIP: A Dynamic Quadtree Vision Prior Enhances MLLM Performance Without Retraining ICLR 2026 Quadtree spatial prior boosts MLLM performance without retraining; ICLR 2026
2026-01-01 ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models ICLR 2026 ERGO efficient high-resolution VLM processing via selective token allocation
2026-01-01 MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning ICLR 2026 Iterative preference learning improves mobile VLM agent chain-of-action reasoning
2026-01-01 SMAN-Bench: A Cross-System Benchmark for Mobile Agents under Single- and Multi-path, Ambiguous, and Noisy Tasks ICLR 2026 Cross-system mobile agent benchmark with ambiguous and noisy real-world tasks
2026-01-01 Simulation to Rules: A Dual-VLM Framework for Formal Visual Planning ICLR 2026 Dual-VLM bridges formal PDDL planning with visual scene understanding
2026-01-01 DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking ICLR 2026 DriveAgent-R1: active perception and hybrid thinking for autonomous driving VLMs
2026-01-01 Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty Quantification ICLR 2026 Evidential uncertainty quantification detects VLM misbehaviors on shifted inputs
2026-01-01 Revisiting Confidence Calibration for Misclassification Detection in VLMs ICLR 2026 Standard calibration degrades misclassification detection in VLMs; key theoretical finding
2026-01-01 Asynchronous Matching with Dynamic Sampling for Multimodal Dataset Distillation ICLR 2026 Asynchronous matching for multimodal dataset distillation; efficient VLM training
2026-01-01 Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks ICLR 2026 Adversarial referring expression tasks expose gaps in VLM visual reasoning
2026-01-01 WIMFRIS: WIndow Mamba Fusion and Parameter Efficient Tuning for Referring Image Segmentation ICLR 2026 Window Mamba fusion with parameter-efficient tuning for referring segmentation
2026-01-01 LINK: Learning Instance-level Knowledge from Vision-Language Models for Human-Object Interaction Detection ICLR 2026 Instance-level VLM knowledge for HOI detection; balances generalization and specialization
2026-01-01 WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM ICLR 2026 WAVE: unified audio-visual VLM embeddings outperform dedicated audio-visual pretraining
2026-01-01 Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds ICLR 2026 Hyperbolic manifold alignment for hierarchical multimodal feature fusion in VLMs
2026-01-01 MoRA: Missing Modality Low-Rank Adaptation for Visual Recognition ICLR 2026 Low-rank adaptation handles missing modalities in VLM visual recognition tasks
2026-01-01 Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization ICLR 2026 Test-time query refinement bridges modality gap in multimodal retrieval with VLMs
2026-01-01 DAVE: A VLM Vision Encoder for Document Understanding and Web Agents ICLR 2026 DAVE encoder captures low-level spatial structure improving document VLM performance
2026-01-01 Adaptive Logit Adjustment for Debiasing Multimodal Language Models ICLR 2026 Adaptive logit adjustment debiases VLM captioning and VQA generation tasks
2026-01-01 WRING Out The Bias: A Rotation-Based Alternative To Projection Debiasing ICLR 2026 Rotation-based debiasing removes CLIP spurious correlations; outperforms projection
2026-01-01 Read the Room: Video Social Reasoning with Mental-Physical Causal Chains ICLR 2026 Video social reasoning benchmark with mental-physical causal chains; VLM ToM
2026-01-01 ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations ICLR 2026 Interleaved text-image reasoning chains mitigate relation hallucination in VLMs
2026-01-01 Hallucination Reduction with CASAL: Contrastive Activation Steering for Amortized Learning ICLR 2026 Contrastive activation steering reduces hallucination via linear knowledge representations
2026-01-01 Reasoning-Driven Multimodal LLM for Domain Generalization ICLR 2026 MLLM reasoning improves domain generalization beyond visual feature invariance
2026-01-01 Discrete Latent Features Ablate Adversarial Attack: A Robust Prompt Tuning Framework for VLMs ICLR 2026 Discrete latent features resist adversarial attacks in VLM prompt tuning
2026-01-01 Inducing Dyslexia in Vision Language Models ICLR 2026 Mechanistic VLM study reveals VWFA analog in CLIP vision encoders
2026-01-01 AQuA: Toward Strategic Response Generation for Ambiguous Visual Questions ICLR 2026 Strategic response generation for ambiguous VQA; addresses real-world uncertainty
2026-01-01 VL-JEPA: Joint Embedding Predictive Architecture for Vision-language ICLR 2026 Novel JEPA architecture for VLMs; continuous embedding prediction vs. autoregressive tokens
2026-01-01 No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers ICLR 2026 Georgia Gkioxari/Meta; label-free visual reasoning training via multimodal verifiers
2026-01-01 Unleashing Perception-Time Scaling to Multimodal Reasoning Models ICLR 2026 Extends inference-time compute scaling to VLM perception; frontier capability paper
2026-01-01 PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs ICLR 2026 Philip Torr/Oxford; multimodal grounding corrects MLLM fine-grained visual failures
2026-01-01 MaskInversion: Localized Embeddings via Optimization of Explainability Maps ICLR 2026 Ferrari/Rupprecht (Google/Oxford); data-free localized CLIP embeddings via mask optimization
2026-01-01 RegionReasoner: Region-Grounded Multi-Round Visual Reasoning ICLR 2026 Cees Snoek; iterative region-grounded multi-round visual reasoning beyond single-pass VLMs
2026-01-01 Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding ICLR 2026 ICLR 2026; chain-of-embedding contrast exposes language prior dominance in LVLMs
2026-01-01 LLaVA-FA: Learning Fourier Approximation for Compressing Large Multimodal Models ICLR 2026 ICLR 2026; Fourier approximation compresses large multimodal models with strong efficiency
2026-01-01 Mitigating Hallucination in Vision-Language Model with Depth and Spatial-aware Key-Value Refinement ICLR 2026 ICLR 2026; depth and spatial-aware KV cache refinement reduces VLM hallucination
2026-01-01 AFTER: Mitigating the Object Hallucination of LVLM via Adaptive Factual-Guided Activation Editing ICLR 2026 ICLR 2026; adaptive factual-guided activation editing targets object hallucination categories
2026-01-01 Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models ICLR 2026 ICLR 2026; query-adaptive entropy-based contrastive decoding mitigates LVLM hallucination
2026-01-01 Cat-PO: Cross-modal Adaptive Token-rewards for Preference Optimization in Truthful Multimodal LLMs ICLR 2026 ICLR 2026; cross-modal token-level reward optimization reduces MLLM hallucination
2026-01-01 GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs ICLR 2026 ICLR 2026; adversarial image synthesis stress-tests MLLM hallucination failure modes
2026-01-01 Transferable and Stealthy Adversarial Attacks on Large Vision-Language Models ICLR 2026 ICLR 2026; progressive semantic infusion yields transferable stealthy LVLM attacks
2026-01-01 Seeing What’s Not There: Negation Understanding Needs More Than Training ICLR 2026 ICLR 2026; shows negation understanding requires more than training data scaling
2026-01-01 NePTune: A Neuro-Pythonic Framework for Tunable Compositional Reasoning on Vision-Language ICLR 2026 ICLR 2026; neuro-symbolic tunable compositional reasoning framework for VLMs
2026-01-01 Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence ICLR 2026 ICLR 2026; Cauchy-Schwarz divergence for distributional VL alignment beyond pairwise InfoNCE
2026-01-01 Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models ICLR 2026 ICLR 2026; unified spatial reasoning benchmark spanning robotics, AR, and navigation
2026-01-01 TABLET: A Large-Scale Dataset for Robust Visual Table Understanding ICLR 2026 Mirella Lapata; large-scale real-world dataset for robust visual table understanding
2026-01-01 FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting ICLR 2026 ICLR 2026; multi-turn frame spotlighting for efficient long-video VLM reasoning
2026-01-01 V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model Interaction ICLR 2026 ICLR 2026; visual prompt-based video benchmark enables richer human-model interaction
2026-01-01 A Training-Free Framework for Long Video Understanding via Video-Query-Options Similarity ICLR 2026 ICLR 2026; training-free long-video understanding via video-query-options similarity
2026-01-01 WMPO: World Model-based Policy Optimization for Vision-Language-Action Models ICLR 2026 ICLR 2026; world-model RL policy optimization for VLA; goes beyond demonstrations
2026-01-01 On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations ICLR 2026 ICLR 2026; systematic study of VLA robustness to multi-modal perturbations
2026-01-01 Sim2Real VLA: Zero-Shot Generalization of Synthesized Skills to Realistic Manipulation ICLR 2026 ICLR 2026; zero-shot sim-to-real VLA transfer via synthesized skill generalization
2026-01-01 Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance ICLR 2026 ICLR 2026; unified latent guidance adapts VLA models to downstream manipulation tasks
2026-01-01 QVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization ICLR 2026 ICLR 2026; channel-aware quantization enables VLA deployment on resource-constrained robots
2026-01-01 EVLP: Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning ICLR 2026 ICLR 2026; reinforced supervised fine-tuning for unified embodied VL planner
2026-01-01 WebWatcher: Breaking New Frontiers of Vision-Language Deep Research Agent ICLR 2026 ICLR 2026; VL deep research agent navigates visual web for complex information-seeking
2026-01-01 GhostEI-Bench: Do Mobile Agent Resilience to Environmental Injection in Dynamic On-Device Environments? ICLR 2026 ICLR 2026; benchmark for VLM agent resilience to environmental injection on mobile
2026-01-01 Uncertainty-Aware Gaussian Map for Vision-Language Navigation ICLR 2026 ICLR 2026; uncertainty-aware Gaussian map improves VLN in partially observed 3D environments
2026-01-01 Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-Language Navigation ICLR 2026 ICLR 2026; dual-system slow-grounding/fast-moving VLM for generalizable VLN
2026-01-01 Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks ICLR 2026 ICLR 2026; argues current medical VLM benchmarks mask true clinical reasoning gaps
2026-01-01 TumorChain: Interleaved Multimodal Chain-of-Thought Reasoning for Traceable Clinical Tumor Analysis ICLR 2026 ICLR 2026; interleaved multimodal CoT enables traceable clinical tumor analysis
2026-01-01 MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning ICLR 2026 ICLR 2026; annotation-free medical VLM reasoning via agentic reinforcement learning
2026-01-01 OmniCT: Towards a Unified Slice-Volume LVLM for Comprehensive CT Analysis ICLR 2026 ICLR 2026; unified slice-volume LVLM covering full CT image analysis pipeline
2026-01-01 MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning ICLR 2026 ICLR 2026; multi-agent RL framework for specialist medical VLM collaboration
2026-01-01 Towards Text-Mask Consistency in Medical Image Segmentation ICLR 2026 ICLR 2026; text-mask consistency constraints fix VLM multi-lesion segmentation failures
2026-01-01 Reversible Primitive–Composition Alignment for Continual Vision–Language Learning ICLR 2026 ICLR 2026; reversible primitive-composition alignment for tight-budget continual VL learning
2026-01-01 Pi-CCA: Prompt-Invariant CCA Certificates for Replay-Free Continual Multimodal Learning ICLR 2026 ICLR 2026; prompt-invariant CCA certificates enable replay-free continual VL learning
2026-01-01 RLAP-CLIP: Continual Multimodal Learning with Prototype Adaptation and Difficulty-Aware Routing ICLR 2026 ICLR 2026; prototype adaptation with difficulty routing for class-incremental CLIP
2026-01-01 Naming to Learn: Class Incremental Learning for Vision-Language Model with Unlabeled Data ICLR 2026 ICLR 2026; naming paradigm enables class incremental VLM learning from unlabeled data
2026-01-01 A-TPT: Angular Diversity Calibration Properties for Test-Time Prompt Tuning of Vision-Language Models ICLR 2026 ICLR 2026; angular diversity calibration addresses prompt dispersion gap in VLM TTA
2026-01-01 Flatness Guided Test-Time Adaptation for Vision-Language Models ICLR 2026 ICLR 2026; flatness-guided TTA links loss landscape geometry to VLM distribution shift
2026-01-01 Bilateral Information-aware Test-time Adaptation for Vision-Language Models ICLR 2026 ICLR 2026; bilateral information-aware TTA handles covariate shifts without test distribution
2026-01-01 Adaptive Debiasing Tsallis Entropy for Test-Time Adaptation ICLR 2026 ICLR 2026; Tsallis entropy debiasing corrects CLIP's built-in bias during TTA
2026-01-01 Benchmarking Bias Mitigation Toward Fairness Without Harm from Vision to LVLMs ICLR 2026 ICLR 2026; first comprehensive benchmark comparing bias mitigation in VLMs
2026-01-01 Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images ICLR 2026 ICLR 2026; VLM-based semantic anomaly detection in AI-generated image outputs
2026-01-01 Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models ICLR 2026 ICLR 2026; perceptually grounded chain-of-thought for faithful VLM geospatial reasoning
2026-01-01 FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark ICLR 2026 ICLR 2026; million-scale T2I reasoning dataset and benchmark from top labs
2026-01-01 Object-Centric Refinement for Enhanced Zero-Shot Segmentation ICLR 2026 ICLR 2026; object-centric region refinement enhances CLIP zero-shot segmentation
2026-01-01 Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition ICLR 2026 ICLR 2026; MLLM as detector-agnostic interaction recognizer for zero-shot HOI
2026-01-01 lmgame-Bench: How Good are LLMs at Playing Games? ICLR 2026 Eric P. Xing/CMU; game-playing benchmark requiring perception, reasoning, planning
2026-01-01 pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language Models ICLR 2026 ICLR 2026; personalized federated fine-tuning for CLIP with multi-modal adapters
2026-01-01 Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking ICLR 2026 ICLR 2026; self-supervised voting and ranking evolves VLMs for image quality assessment
2026-01-01 VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation ICLR 2026 ICLR 2026; VL anchoring enables one-shot generalization for bimanual manipulation
2026-01-01 Seeing What’s Wrong: A Trajectory-Guided Approach to Caption Error Detection ICLR 2026 ICLR 2026; trajectory-guided multi-score method detects subtle image-caption errors