Vision-Language Models — 2026 · Paper list (718)¶
July 2026 (308)¶
| Date | Paper | Venue | Why selected |
|---|---|---|---|
| 2026-07-23 | K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs | — | — |
| 2026-07-23 | Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text | — | — |
| 2026-07-22 | Test-Time Training for Modality Order Consistency in Vision-Language Models | — | Gandelsman (Berkeley); TTT fixes systematic VLM modality-order bias; reproducible across 3 models |
| 2026-07-22 | Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs | ECCV 2026 | ECCV 2026; joint token-compute pruning cuts MLLM inference cost; directly deployable |
| 2026-07-22 | Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning | — | RLVR training environment for VLMs with taxonomy-guided verifiable visual rewards |
| 2026-07-22 | LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition | — | Medical VLM fine-tuning for surgical instrument-tissue interaction; Intel Labs; reproducible |
| 2026-07-22 | ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models | — | Taxonomic probe showing VLMs reliably succumb to irrelevant contextual entrainment |
| 2026-07-22 | SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments | — | — |
| 2026-07-22 | ReferTrack: Referring Then Tracking for Embodied Visual Tracking | — | — |
| 2026-07-22 | Robostral Navigate | — | — |
| 2026-07-21 | Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models | — | Deepak Pathak + Zhiqiu Lin; LLM-generated CLIP descriptors carry no visual evidence — fundamental rethink |
| 2026-07-21 | Dual Adversarial Fine-tuning for Enhancing Robustness of Large Vision Language Model | — | ICT CAS (Shan/Chen/Gao); dual adversarial fine-tuning closes LVLM visual robustness gap |
| 2026-07-21 | Stochastic Meta-Unlearning: Bridging Language Backbone and Multimodal Unlearning | — | Sijia Liu + Tianlong Chen; first principled multimodal unlearning bridging language and vision components |
| 2026-07-21 | Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval | ECCV 2026 | ECCV 2026; zero-shot VMR resolving modality and language-style alignment gaps |
| 2026-07-21 | PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image | — | PathAgentBench; evidence-seeking VLMs evaluated on multi-scale whole-slide pathology |
| 2026-07-21 | ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting | ECCV 2026 | ECCV 2026; multi-target referring segmentation extending 3DGS to language-guided scene understanding |
| 2026-07-21 | No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation | — | Test-time scaling lifts VLM navigation performance without any additional training |
| 2026-07-21 | Now You See the Hate: Adaptive View Retrieval for Hidden Hateful Illusions | — | VLM safety gap: state-of-the-art models under 25% accuracy on hateful optical illusions |
| 2026-07-21 | Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio | — | Unified text+image+video+audio embedding; fills critical audio gap in VLM retrieval stacks |
| 2026-07-21 | DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization | — | GRPO-trained medical VLM for CXR reports; directly replicable RL recipe for domain-specific VLMs |
| 2026-07-21 | MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings | — | Multi-party ToM benchmark; measures social/epistemic reasoning gap in MLLMs for agentic settings |
| 2026-07-21 | ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis | — | Expert-level visual reasoning benchmark; stress-tests knowledge-intensive VLM synthesis beyond commonsense |
| 2026-07-21 | IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer | — | Streaming 4D instance-grounded geometry; spatial intelligence backbone for embodied agents |
| 2026-07-21 | OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation | — | On-policy self-distillation adapts LVLM judgments to pixel-level IAD; novel spec-to-detector recipe |
| 2026-07-21 | HPD-Parsing: Hierarchical Parallel Document Parsing | — | Hierarchical parallel document parsing with VLMs; speed+accuracy gain for doc-understanding pipelines |
| 2026-07-21 | ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU | — | — |
| 2026-07-21 | Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges | — | — |
| 2026-07-21 | DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking | — | Yu Qiao (Shanghai AI Lab); reasoning-guided physics-aware video generation via spatial-temporal masking |
| 2026-07-21 | ChronoStitch: Training-Free Composition of Visual KV Memories for Long-Horizon Temporal Reasoning | — | Training-free KV cache composition for long-video QA; avoids video reprocessing |
| 2026-07-21 | MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts | ECCV 2026 | ECCV 2026; benchmark exposing VLM failure on missing-part detection; honest hallucination audit |
| 2026-07-20 | Patch Policy: Efficient Embodied Control via Dense Visual Representations | — | Yann LeCun co-author; dense ViT features for embodied control, underexplored direction |
| 2026-07-20 | The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric | — | Jun-Yan Zhu + Shechtman + Hertzmann; text-prompted multi-sense perceptual metric, novel |
| 2026-07-20 | Memory-Supported Synergistic Adaptation for Training-Free Test-Time Medical Image Segmentation | ECCV 2026 | ECCV 2026; memory-synergistic TTA for medical VLM segmentation, classification→seg bridge |
| 2026-07-20 | O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning | ECCV 2026 | ECCV 2026; object-centric VLM reasoning for industrial anomaly detection |
| 2026-07-20 | FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications | — | Agent harness for real-time multimodal deployment; directly actionable for harness builders |
| 2026-07-20 | Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation | — | Future-state-conditioned VLN; novel supervision signal beyond next-action cloning |
| 2026-07-20 | Reasoning as a Double-Edged Sword: Architecture and Cross-Stage Robustness in Vision-Language-Action Models | — | Empirical robustness audit of reasoning in VLAs; honest baselines; safety-critical insight |
| 2026-07-20 | SEE: Structure-aware Exploring \& Exploiting for Long-horizon GUI Agent Trajectory Synthesis | — | GUI agent trajectory synthesis with VLMs; long-horizon coverage; agentic scaffolding |
| 2026-07-20 | Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs | — | Causal audit proves attention heatmaps don't reflect actual medical VLM evidence |
| 2026-07-20 | Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs | — | Domain-generalized pixel-level tampering detection in VLMs; practical safety/integrity tooling |
| 2026-07-20 | FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry | — | — |
| 2026-07-20 | PRiSM: Prototype Regularization for Few-Shot VLMs | — | Dolz + Ben Ayed; few-shot VLM adaptation under realistic class-imbalance; buildable recipe |
| 2026-07-20 | Measuring and Improving Complex-Atomic Answer Consistency in Endoscopic VQA | — | Complex-atomic answer consistency metric for endoscopic VQA; fills medical eval gap |
| 2026-07-19 | TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs | — | Generalist video temporal grounding with MLLMs; fills key capability gap in video VLMs |
| 2026-07-19 | EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding | — | Jiwen Lu (Tsinghua); EvoGUI isolates state-transition reasoning from end-to-end GUI success |
| 2026-07-19 | Searching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language Models | — | Evolutionary block pruning finds task-specific vision paths; cuts VLM compute without retraining |
| 2026-07-18 | Look Clearly Before Answering: Mitigating Hallucinations in LVLMs via Saliency-Driven Perceptual Realignment | — | Saliency-driven hallucination mitigation in LVLMs; core reliability problem, buildable method |
| 2026-07-18 | Dataset Distillation by Influence Matching | CVPR 2026 | CVPR 2026; outcome-centric dataset distillation; broadly applicable to VLM training data |
| 2026-07-18 | Can Experts Adapt Without Training? On Test-Time Modality Generalization in MVLMs | MICCAI 2026 | MICCAI 2026; test-time modality generalization for medical VLMs; directly buildable |
| 2026-07-18 | HTT-Net: Hierarchical Text-guided Transition Modeling for Surgical Video Phase Recognition | — | Text-guided surgical video phase recognition; strong medical VLM baseline with temporal structure |
| 2026-07-17 | Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models | — | Deva Ramanan (CMU); prompt-position paradox fix; universal VLM prompting impact |
| 2026-07-17 | Von Mises-Fisher Mixture Model with Dynamic Shrinkage for Realistic Test-Time Transduction | ICML 2026 | ICML 2026; vMF mixture TTA-transduction; rigorous test-time VLM adaptation theory |
| 2026-07-17 | More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe | — | Luc Van Gool (ETH); simple recipe beats RS-specific archs; transferable scaling insight |
| 2026-07-17 | Model Merging for Medical LVLMs: A Benchmark and a Winner-Take-All Approach | — | Medical LVLM model-merging benchmark + WTA method; practical LoRA-expert fusion |
| 2026-07-17 | How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA | — | Vision-operation misalignment taxonomy; actionable failure map for compositional VQA |
| 2026-07-17 | ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning | — | VL tool-use + RL for multimodal scientific verification; novel reward signal design |
| 2026-07-17 | When Can Test-Time Adaptation Help Zero-Shot CT Vision-Language Models? | — | Leonid Sigal (UBC); systematic TTA scope study on zero-shot 3D CT VLMs |
| 2026-07-17 | Region-Grounded Vision-Language Learning for Detection-Guided Mammographic Lesion Classification | — | Region-grounded contrastive VL for mammography; spatial alignment beats global CLIP |
| 2026-07-17 | Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs | — | — |
| 2026-07-17 | Vision-Language Assistant for Emotional Reactions to Risky Driving | — | — |
| 2026-07-17 | An Exam for Active Observers | — | — |
| 2026-07-17 | Attention-Guided Saliency Maps for Interpreting Visualization Literacy in VLMs | — | — |
| 2026-07-17 | Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting | — | — |
| 2026-07-17 | Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling | — | — |
| 2026-07-17 | One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models | — | Cross-modal unlearning vulnerability in VLMs; important safety gap with concrete attack surface |
| 2026-07-17 | Foresight Residual RL for Long-Horizon Robot Manipulation with Vision-Language-Action Models | — | VLA + residual RL for long-horizon manipulation; Boularias lab; addresses credit assignment gap |
| 2026-07-17 | JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models | — | Multi-tenant post-training service for VLA; novel deployment angle; Wentao Zhang group |
| 2026-07-17 | PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation | — | Reflective agentic physics control for video gen; VLM+agent+simulation pipeline |
| 2026-07-16 | Symbal: Detecting Systematic Misalignments in Model-Generated Captions | ICML 2026 | Stanford/Langlotz group; ICML 2026; systematic caption error detection for medical MLLMs |
| 2026-07-16 | Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology | — | Multi-LLM VLM pipeline for 3D brain MRI reports; novel visual instruction tuning for oncology |
| 2026-07-16 | Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models | — | Action supervision reshapes multimodal representations in VLA; structural insight for builders |
| 2026-07-16 | FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models | — | Visual foresight + motion guidance in VLA models; explicit forward prediction addresses reactivity |
| 2026-07-16 | WorkDrive: Roadwork Chain of Causation for Autonomous Driving | — | Chain-of-causation reasoning for roadwork zones; tackles hard OOD failure mode in driving VLMs |
| 2026-07-16 | HoloGeo: Mitigating Landmark Bias in Geo-localization via Evidence-Driven Reasoning | — | Evidence-driven reasoning to mitigate landmark bias in VLM geo-localization |
| 2026-07-16 | RoboTTT: Context Scaling for Robot Policies | — | — |
| 2026-07-16 | SportD: Can VLMs Physically Strategize? | — | Probes VLM physical strategy/reasoning in soccer; novel eval axis for agent decision-making |
| 2026-07-16 | U-shaped Multi-granularity Learning for Vision-Language Models | — | Prompt learning granularity dilemma for VLMs; practical recipe for cross-task generalization |
| 2026-07-16 | GeoDetect: Geometric Adversarial Detection for VLPs | ECCV 2026 | Adversarial detection for multimodal VLPs; ECCV 2026; geometric approach is novel angle |
| 2026-07-16 | VTM-Nav: Hierarchical Visual-Topological Memory for Cross-Episode Object-Goal Navigation | — | VLM-driven object-goal navigation with topological memory; training-free embodied agent system |
| 2026-07-16 | On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline | — | Simplicity-first transferable attack on VLPMs; questions complexity of existing pipelines |
| 2026-07-16 | ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors | — | — |
| 2026-07-16 | Knowing You at First Glance: Inferring Apparent Personality from Faces | — | — |
| 2026-07-16 | Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories | — | Xiaomi robotics; 100K hrs real VLA trajectories; massive-scale grounded VLM-action |
| 2026-07-16 | Video = World + Event Stream | — | Wan-Streamer (Alibaba); world+event decomposition reframes video VLM architecture |
| 2026-07-16 | Training-Free Open-Vocabulary 3D Point-Cloud Segmentation on the Generalized Few-Shot Benchmark | — | — |
| 2026-07-15 | Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation | ECCV 2026 | — |
| 2026-07-15 | ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding | ECCV 2026 | — |
| 2026-07-15 | M\(^\text{4}\)World: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming | — | — |
| 2026-07-15 | SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning | — | — |
| 2026-07-15 | Fine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT Understanding | — | — |
| 2026-07-15 | Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs | — | — |
| 2026-07-15 | Fine-grained CLIP fine-tuning with self-annotated region alignment | — | — |
| 2026-07-15 | Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models | — | Attention-free token reduction; directly enables VLM deployment on edge/resource-constrained devices |
| 2026-07-15 | Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment | — | Representation anchoring + language-action alignment; fixes BC fine-tuning catastrophic forgetting in VLAs |
| 2026-07-15 | Semantic Anchoring for Robotic Action Representations | — | Semantic anchoring preserves VLM priors in VLA fine-tuning; Yizhou Wang (PKU) group |
| 2026-07-15 | Self-Improving is Often Sudden: Enlightenment-style Finetuning for Large-Scale Models | — | Phase-transition self-improvement phenomenon; Tianwei Zhang (NTU); actionable finetuning recipe |
| 2026-07-15 | FM\(^2\): Unified Federated Foundation Models for Heterogeneous Multimodal Medical Imaging | — | Federated multimodal medical VLM; addresses privacy-constrained cross-institution medical AI deployment |
| 2026-07-15 | SD-MAR: Multi-image Analytical Reasoning via Synthetic Data and Reinforcement Learning | — | RL + synthetic data for multi-image analytical reasoning; addresses core VLM reasoning gap |
| 2026-07-15 | MonteRET: AI Agent Enhancing Multimodal LLMs with Multi-granularity Knowledge Retrieval for Chest CT Report Generation | — | Multi-granularity RAG agent for chest CT reports; retrospective clinical evaluation included |
| 2026-07-15 | DiMaS: Distribution Matching for Steering Vision-Language-Action Models | — | Fine-grained behavioral steering for flow-matching VLA; enables controllable robot policy |
| 2026-07-15 | Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection | — | Force injection during VLA post-training; addresses contact-state blindspot in vision-driven policies |
| 2026-07-15 | RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination | — | — |
| 2026-07-15 | GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning | — | Zero-shot transit video analytics with grounded hybrid VLM reasoning; practical deployment recipe |
| 2026-07-15 | Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making | — | — |
| 2026-07-15 | ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning | — | — |
| 2026-07-14 | VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression | — | Visual token compression for VLMs; training-free LLM-as-encoder approach; addresses core inference efficiency bottleneck |
| 2026-07-14 | CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models | — | CoRe framework for cross-image comparative reasoning; fine-grained attribute grounding; fills known VLM compositional gap |
| 2026-07-14 | Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding | — | Gaussian mixture keyframe selection for long video VLMs; event-aware visual allocation; addresses uniform-sampling information loss |
| 2026-07-14 | Hy-Embodied-VLM-1.0: Efficient Physical-World Agents | — | Hy-Embodied-VLM-1.0; end-to-end embodied agent with multimodal perception and agentic reasoning; buildable system report |
| 2026-07-14 | Auditing Data Leakage in Whole-Slide Image Multimodal Benchmarks | — | Data leakage audit in WSI pathology VLM benchmarks; questions benchmark validity; critical for medical VLM evaluation |
| 2026-07-14 | A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism | — | Controlled null result on GRPO failing for small VLMs; mechanistic analysis; essential calibration for agent RL fine-tuning |
| 2026-07-14 | What Does a Temporal Benchmark Score Measure? Decomposing Channel Use in Video VLM Evaluation | — | Decomposes what temporal benchmark scores actually measure in video VLMs; channel-use analysis; methodological contribution |
| 2026-07-14 | Let RGB Be the Language of Vision | — | Unified RGB formulation for all visual modalities (masks, depth, etc.); architectural simplification with broad VLM implications |
| 2026-07-14 | ReflectVLN: Training Vision-Language Navigation Agents with Reflective Reasoning | — | ReflectVLN adds closed-loop reflective reasoning to VLN agents; explicit failure diagnosis mechanism; directly applicable to agent harnesses |
| 2026-07-14 | The Sound of Absence: Audio-Language Embedding Models Struggle with Negation | — | Hung-yi Lee (NTU); exposes fundamental negation blindspot in CLAP-style audio-language models |
| 2026-07-14 | Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering | — | — |
| 2026-07-14 | MQAdapter: Multi-Modal Quantum Adapter for Coarse-to-Fine VLM Fine-tuning | — | Quantum-inspired adapter for coarse-to-fine VLM fine-tuning; novel PEFT direction |
| 2026-07-14 | DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery | — | Spatial cognition augmentation for VLMs on street-view; directly relevant to geo-VLM eval |
| 2026-07-14 | Lost in Visual Translation: A VLM-Assisted Perceptual-Semantic Coherence Framework for EEG-to-Image Reconstruction | — | VLM-assisted perceptual-semantic coherence for EEG-to-image; novel eval framework |
| 2026-07-14 | Breaking Déjà Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning | — | VLM reasoning for visual place recognition auditing; novel robustness eval for embodied nav |
| 2026-07-14 | Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation | — | — |
| 2026-07-13 | MED-DSLC: Multi-Expert-Domain Classification via Domain Supervision and Logit Calibration | ECCV 2026 | ECCV 2026; Nuno Vasconcelos (UCSD); fixes CLIP logit non-comparability across domains—principled zero-shot fix |
| 2026-07-13 | StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure | — | Causal structure for long-horizon VLM digital agents; directly relevant to agent harness design |
| 2026-07-13 | Confidence Scores in Open-Vocabulary Detection Are a Biased Mixture of Scale and Semantics | — | Open-vocab detector confidence = scale+semantics mixture; exposes fundamental reliability flaw in CLIP-based OVD |
| 2026-07-13 | TIGER: Text-Conditioned Visual Gated Routing with Acceptance Alignment for Multimodal Speculative Decoding | — | Multimodal speculative decoding with gated routing; practical inference speedup for VLMs |
| 2026-07-13 | When Depth Is Better Told Than Shown: Depth-Ordinal Prompting for Vision-Language Spatial Reasoning | — | Depth-ordinal prompting for VLM spatial reasoning; addresses a core 3D perception weakness in VLMs |
| 2026-07-13 | Think When It Matters: Conditional VLM Reasoning for Social Navigation with RL Policies | — | Conditional VLM reasoning with RL for social navigation; conditional compute in embodied agents |
| 2026-07-13 | StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description | ECCV | — |
| 2026-07-13 | DynEval: Holistic Evaluations of T2I Generative Models in the Wild | ECCV 2026 | — |
| 2026-07-13 | An Empirical Analysis of Continual Learning for Heterogeneous Medical Visual Question Answering | — | Empirical continual-learning study for medical VQA; practical for multi-task clinical VLM deployment |
| 2026-07-13 | LaGuadia: Language-Guided Adaptive Distillation from Pathology Foundation Models | — | Language-guided distillation from pathology FMs; Won-Ki Jeong; reduces WSI inference cost |
| 2026-07-13 | MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models | — | — |
| 2026-07-13 | The Emerging Paradigm of Geospatial Foundation Models: From Pre-Training to Agentic Reasoning | — | — |
| 2026-07-13 | SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning | — | Self-verification RL for multimodal reasoning; extends R1-style training to VLMs |
| 2026-07-13 | See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models | — | Robot-frame 3D pointmaps for VLA; fixes camera-frame mismatch, strong geometric inductive bias |
| 2026-07-12 | Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models | ICML 2026 | ICML 2026; spectral-heat-flow token condensation; directly cuts VLM inference cost without info loss |
| 2026-07-12 | Learning Anatomy-Grounded CT Vision-Language Representations with Organ-Hierarchical Report Knowledge | — | Anatomy-hierarchical VLP from CT+reports; strong medical VLM pretraining recipe for organ-level grounding |
| 2026-07-12 | 3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects | — | Factorial study of VLM eval pipelines for 3D generation; rigorous methodology for automated judge reliability |
| 2026-07-12 | Mixture of Cognitive Experts in Large Vision-Language Models | — | Mixture-of-cognitive-experts for LVLMs; novel framing of metacognition and diverse representations |
| 2026-07-12 | Detecting AI-Generated Video: A Vision-Language Dual-View Survey | ACL 2026 | — |
| 2026-07-12 | Traj-VLN: Learning Pixel-Space Interaction via Autoregressive Trajectory Generation | — | Autoregressive pixel-space trajectory generation for VLN; novel grounding approach for navigation agents |
| 2026-07-11 | What Does Your Short-Answer VQA Score Actually Measure? Evaluator-Dependent Instability in Multimodal Short-Answer Benchmarks | — | Demonstrates VQA benchmark scores are evaluator-dependent; critical for anyone running VLM evals |
| 2026-07-11 | Empowering Long-form Omni-modal Understanding with Robust Audio Perception | — | Weidi Xie (SJTU/Oxford); long-form omni-modal understanding with explicit audio alignment; scarce labeled data angle |
| 2026-07-11 | SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding | COLM 2026 | COLM 2026; controlled benchmark for long-context VLM document understanding, fills real eval gap |
| 2026-07-11 | PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning | — | Physical plausibility reasoning in video-LMs; targets a known VLM failure mode with retrieval+verification |
| 2026-07-11 | Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift | — | Benchmarks 15 foundation model backbones for mammography under domain shift; directly relevant to medical VLM eval |
| 2026-07-11 | WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding | — | Training-free UHR remote sensing VLM reasoning with structured evidence; novel efficiency angle |
| 2026-07-11 | ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory | — | Robotic agent OS with lifelong multimodal memory; directly relevant to agent harness architecture |
| 2026-07-11 | Imperceptible and Reversible Adversarial Examples against Vision-Language Models for Privacy Protection | — | Reversible adversarial examples against VLMs for privacy; novel threat model with practical reversibility |
| 2026-07-11 | ABot-N1: Toward a General Visual Language Navigation Foundation Model | — | Foundation model for Visual Language Navigation; unifies spatial reasoning and embodied versatility |
| 2026-07-10 | ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts | — | Pathology foundation model fusing vision, VL, and slide-level experts; builds on strong computational pathology trend |
| 2026-07-10 | Video Generation Models are General-Purpose Vision Learners | ECCV 2026 | Video generation as general vision learner; ECCV 2026; Jasper Uijlings (Google); strong foundational claim |
| 2026-07-10 | SigLIP-HD by Fine-to-Coarse Supervision | ICLR 2026 | SigLIP-HD fine-to-coarse supervision; ICLR 2026; Hengshuang Zhao group; directly improves MLLM visual encoding |
| 2026-07-10 | Super-Generalist: Towards Comprehensive and Accurate Medical Image Understanding via Generalist-Specialist Synergy | — | Generalist-specialist synergy for medical image understanding; addresses known VLM accuracy gap in clinical settings |
| 2026-07-10 | MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models | — | Heterogeneous VLM with linear attention interleaving; efficiency + performance; actionable architecture direction |
| 2026-07-10 | Robustifying Vision-Language Models via Test-Time Prompt Adaptation | ICML 2026 | ICML 2026; test-time prompt adaptation for adversarial robustness in VLMs; strong venue, practical defense method |
| 2026-07-10 | The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs | — | Diagnoses counting failures as representation-verbalization gap; mechanistic insight actionable for VLM evaluation |
| 2026-07-10 | Test-Time Scaling for Small VLMs on Multilingual Visual MCQ | — | Test-time scaling on small open VLMs—transfers TTS gains beyond frontier models; highly actionable |
| 2026-07-10 | Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference | — | Reveals decode (LLM text gen) not vision tokens as edge energy bottleneck; inverts prevailing optimization target |
| 2026-07-10 | Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models | — | Decade-long longitudinal analysis of VLM visual-cognitive errors; rare rigorous error taxonomy work |
| 2026-07-10 | TCLA: Training-Free Class-wise Logit Adaptation for Medical Vision-Language Models | — | Training-free class-wise logit adaptation for OOD medical VLMs; lightweight, reproducible, practical |
| 2026-07-09 | Dive Into the Implicit Biases of Low-rank Vision-language Alignment | ECCV 2026 | ECCV 2026; challenges full-param alignment dogma with low-rank analysis — rethinks VLM training recipe |
| 2026-07-09 | Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing | — | Cognitive memory architecture for unified understand/generate/edit — directly improves multimodal agent design |
| 2026-07-09 | When Thinking Hurts: Epistemic Signals in the Reasoning Chains of Visual Language Models | ICML 2026 | ICML 2026; first empirical characterization of reasoning-chain uncertainty in thinking-mode VLMs — calibration insight |
| 2026-07-09 | APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts | — | Stanford (Haber lab); interleaved vision-language thoughts for long-horizon robot planning — actionable VLA scaffold |
| 2026-07-09 | Rethinking Small VLM Quantization: From Component-Wise Analysis to Hardware-Aware Edge Deployment | ICML 2026 | ICML 2026; component-wise quantization analysis for sub-3B VLMs — direct edge deployment recipe |
| 2026-07-09 | Playing ZendoWorld: Challenging AI Agents on Active Visual Concept Induction | — | ZendoWorld benchmark for active visual concept induction — novel evaluation of hypothesis-driven VLM reasoning |
| 2026-07-09 | When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities | — | Sparse autoencoders for cross-modal mechanistic interpretability in VLMs — novel probe into internal representations |
| 2026-07-09 | VEGAS: Human-Aligned Video Caption Evaluation via Gaze | — | Gaze-grounded VLM caption eval; training-free metric; novel human-attention alignment signal |
| 2026-07-09 | LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action | — | Selective visual token routing in VLA; addresses core dynamic-scene limitation in robot VLMs |
| 2026-07-09 | AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding | CVPR | CVPR dashcam VQA benchmark; incident-centric eval fills gap in autonomous driving VLM testing |
| 2026-07-09 | Post-Training in End-to-End Autonomous Driving | — | Comprehensive post-training survey for E2E autonomous driving VLAs; strong practical recipes |
| 2026-07-09 | WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving | — | Dual-level world-cognitive VLA for autonomous driving; addresses reactive limitation with world foresight |
| 2026-07-09 | VocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database Retrieval | — | Visual tokenization + vector DB for open-vocab detection/segmentation; novel retrieval-grounded approach |
| 2026-07-09 | Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data | — | Adaptive benchmark sizing for model evaluation; directly actionable for rigorous VLM evaluation practice |
| 2026-07-09 | Open-ended Multi-agent Autocurricula via Visual Inspection of Policies with Multi-modal LLMs | — | Multimodal LLMs as policy inspectors for RL curricula; novel VLM-as-evaluator application |
| 2026-07-09 | Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation | — | Attribute-driven compositional referring segmentation for endoscopy; medical VLM with fine-grained control |
| 2026-07-09 | Write-Protected Discrete Bottlenecks for Language-Grounded World Models: A Structural Limitation and Sufficient Fix | — | Structural analysis of language-grounded world models; critiques RT-2/PaLM-E paradigm with fix |
| 2026-07-09 | MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs | — | Multi-view integration benchmark; targets untested allocentric 3D reasoning gap in VLMs; diagnostic benchmark value |
| 2026-07-09 | Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing | — | Learning from privileged modalities via probing; addresses inference-time modality mismatch; practical for deployment |
| 2026-07-09 | Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models | — | Blind-Spots-Bench; systematic evaluation of multimodal model failure modes humans find trivial; benchmark for VLM robustness |
| 2026-07-08 | HIVE: Understanding Post-Hallucination Reasoning in Vision Language Models | ECCV 2026 | ECCV 2026; dissects post-hallucination reasoning chains in VLMs, not just detection |
| 2026-07-08 | AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning | ECCV 2026 | ECCV 2026; query-relevance-anchored visual token pruning cuts VLM inference cost |
| 2026-07-08 | Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation | — | — |
| 2026-07-08 | Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering | — | Comparative domain-adapted VLMs for DocVQA — honest baselines across architectures for document understanding |
| 2026-07-08 | When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs | — | Graph-based attribute reasoning for VLM calibration — addresses overconfidence in prompt-tuned models |
| 2026-07-08 | UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma | — | — |
| 2026-07-08 | Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks | ACL | ACL survey; comprehensive multimodal unlearning across modalities; timely safety/privacy topic |
| 2026-07-08 | Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models | ACL | ACL; multilingual spatial deictic eval; exposes systematic VLM spatial reasoning failure mode |
| 2026-07-08 | On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces | — | Spectral subspace lens on VLM adversarial vulnerability; novel mechanistic robustness analysis |
| 2026-07-08 | InfraQR: Edge-Placed QR-Inspired Structured Patch Attacks on Infrared Vision-Language Models | — | Structured adversarial patch attacks on infrared VLMs; underexplored robustness axis with security relevance |
| 2026-07-08 | Generalist Vision-Language Models for Fast Radio Burst detection: a zero-shot benchmark against a specialized detector | — | Zero-shot VLM benchmark vs. specialized FRB detector; honest generalist vs. specialist evaluation |
| 2026-07-08 | Heterogeneity-Adaptive Diffusion Schrodinger Bridge for PET-Guided Whole-Body MRI Translation | MICCAI 2026 | — |
| 2026-07-08 | Prototype-Anchored Generalized Manifold Regression for Unknown-Domain Object Detection | — | — |
| 2026-07-08 | BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning | — | Brain-inspired backward prediction for VLM self-reflection; novel training-free reasoning improvement mechanism |
| 2026-07-07 | Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders | ICML 2026 | ICML 2026; Daniel Cohen-Or/Or Patashnik; novel mechanistic insight into VLM localization signals |
| 2026-07-07 | Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment | ICML 26 | ICML 2026; VLM physical reasoning failure modes directly impact VLA agent generalization |
| 2026-07-07 | SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs | ECCV | ECCV; SAM-based optimizer for VLM prompt tuning; directly addresses performance-generalization tradeoff |
| 2026-07-07 | What Images Cannot Say: Language-Guided Olfactory Representation Learning | ECCV 2026 | — |
| 2026-07-07 | Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation | — | — |
| 2026-07-07 | MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation | — | — |
| 2026-07-07 | AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models | — | — |
| 2026-07-07 | A VLM-Enhanced Framework for Comprehensive Traffic Sign Condition Assessment Integrating Daytime Visual Performance and Nighttime Retroreflectivity Evaluation | — | — |
| 2026-07-07 | VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery | — | — |
| 2026-07-07 | Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement | — | — |
| 2026-07-07 | TMF-RSE: Tri-Modal Fusion with Regional Semantics and Evidential Uncertainty for Lung Severity Scoring | — | — |
| 2026-07-07 | PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails | — | Policy-adaptive image guardrails benchmark; critical for real deployment safety |
| 2026-07-07 | Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability | — | Multicultural/multilingual VLM safety benchmark exposing Western-centric evaluation gaps |
| 2026-07-07 | Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding | — | — |
| 2026-07-07 | SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models | — | — |
| 2026-07-07 | RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures | — | — |
| 2026-07-07 | Structured-Condensed Prompt Tuning in Vision-Language Models for Fine-grained Image Recognition | — | Structured prompt tuning for CLIP fine-grained recognition — practical, reproducible PEFT recipe for VLMs |
| 2026-07-07 | Reward Valuation in Vision Language Models: Causal Mechanisms Underlying Anhedonia | — | Schrimpf lab; mechanistic VLM–cognition alignment; novel anhedonia/reward valuation probe |
| 2026-07-07 | Pelican-VLA 0.5: Attending Before Acting Benefits Generalization | — | Unified VLA with attention-before-action; integrates future-frame gen + action in one model |
| 2026-07-07 | UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation | — | UI-to-app generation benchmark; visual interaction inference; novel multimodal code-gen eval |
| 2026-07-07 | PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet | — | 3D dense captioning with pseudo-cap + voxel net; concrete architectural advance for 3D VLMs |
| 2026-07-07 | Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review | — | Review of VLA models across UAV and bimanual manipulation; useful scope survey for builders |
| 2026-07-07 | VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection | — | — |
| 2026-07-06 | TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models | — | Interpretable concept-overlap token reduction cuts VLM compute with explainable mechanism |
| 2026-07-06 | Does It Fail to See or Fail to Know? Attributing Errors in Vision-Language Models | — | Disentangles perception vs. knowledge failures in VLM errors — useful for evaluation design |
| 2026-07-06 | Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data? | — | VLMs for medical data standardization — tackles missing preprocessing step in clinical AI pipelines |
| 2026-07-06 | QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding | ECCV 2026 | ECCV 2026; query-conditioned temporal retrieval tackles core long-video VLM degradation problem |
| 2026-07-06 | PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving | ECCV 2026 | ECCV 2026; end-to-end VLA for autonomous driving; scalable architecture with practical evaluation |
| 2026-07-06 | InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization | — | InternVLA series (Shanghai AI Lab); unifies semantic priors and future prediction in one VLA |
| 2026-07-06 | Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation | — | Bidirectional hierarchical VLA agent addressing Markovian failure in long-horizon manipulation |
| 2026-07-06 | Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval | — | Object-evidence token merging for dense retrieval; practical efficiency for deploying VLMs |
| 2026-07-06 | Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models | TMLR | — |
| 2026-07-06 | StructuredEdit: Constraint-Aware Graphic Design Editing via Differentiable Parameter Propagation | SIGGRAPH | — |
| 2026-07-06 | DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation | — | Dynamic gating for reasoning segmentation bridges language queries to pixel masks |
| 2026-07-06 | Repurposing CLIP to Localize at Pixel Level | — | Repurposes CLIP for pixel-level localization, addressing global-feature bias bottleneck |
| 2026-07-06 | Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis | — | 3D scene graph + VLM open-vocab scene understanding beyond isolated object lifting |
| 2026-07-06 | HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better | — | Tencent HunyuanOCR-1.5; lightweight OCR-VLM unifying doc parsing, extraction, translation |
| 2026-07-05 | DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics | ICML 2026 | ICML; schema-guided world modeling for hierarchical visual dynamics in multimodal LLMs |
| 2026-07-05 | SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering | — | Visual token engineering to mitigate hallucinations in LVLMs — practical LVLM reliability fix |
| 2026-07-05 | RL Forgets! Towards Continual Policy Optimization | — | Exposes catastrophic forgetting in RL-based VLM post-training — critical for continual adaptation |
| 2026-07-05 | ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog | — | — |
| 2026-07-05 | AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in Memes | — | — |
| 2026-07-04 | Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs | — | Counterfactual grounding + hard-negative contrastive training exposes visual shortcut problem in medical VLMs |
| 2026-07-04 | USE: A Unified Self-Ensembling Framework for Test-Time Prompt Tuning | ICML 2026 | ICML; unified self-ensembling for test-time prompt tuning advances CLIP-based TTA |
| 2026-07-04 | Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process | — | RL over interleaved text-image generation; directly informs training unified multimodal agents |
| 2026-07-04 | Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models | — | Distills tree search into frozen VLA action evaluation; novel adaptation without fine-tuning |
| 2026-07-03 | RADIO1D: Elastic Representations for Condensed Vision Modeling | ICML 2026 | ICML; challenges fixed 2D patch assumption in VLMs with elastic 1D representations |
| 2026-07-03 | Present but Not Remembered: Auditing How Frozen VLAs Encode, Deploy, and Steer Visual History | — | Mechanistic audit of how frozen VLAs encode visual history — informs VLA memory architecture |
| 2026-07-03 | MentalThink: Shaping Thoughts in Mental SVG World | — | Novel think-with-SVG paradigm enabling executable visual reasoning in MLLMs |
| 2026-07-03 | Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment | — | — |
| 2026-07-02 | Show Me Examples: Inferring Visual Concepts from Image Sets | — | Apple Research (Susskind/Bautista); fills fundamental VLM gap: concept inference from image sets without text |
| 2026-07-02 | Visually Grounded Self-Reflection for Vision-Language Models via Reinforcement Learning | — | Durrett lab; RL trains visual self-reflection/CoT correction — key for VLM agent reasoning chains |
| 2026-07-02 | LASER: A Corrective Lens for LVLMs via Visual Attention Preservation and Sink Suppression | ECCV 2026 | ECCV 2026; LASER attention-sink suppression fixes visual forgetting in long LVLM decoding |
| 2026-07-02 | Towards Robustness against Typographic Attack with Training-free Concept Localization | ECCV 2026 | ECCV 2026; training-free CLIP typographic attack defense — critical for all CLIP-encoder VLMs |
| 2026-07-02 | Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs | — | Task-agnostic VLA pretraining; novel approach to VLA data bottleneck via motion pre-training |
| 2026-07-02 | VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment | — | VLAFlow unified VLA co-training + future latent alignment; clean ablation of pretraining paradigms |
| 2026-07-02 | MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models | — | MMBench-Live; continuously evolving benchmark combats contamination — raises evaluation rigor bar |
| 2026-07-02 | Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning | ECCV 2026 | ECCV 2026; entropy-aware visual token pruning under dense/fine-grained instructions — practical efficiency |
| 2026-07-02 | Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editing | — | Online recursive MLLM editing with bounded overhead — essential for keeping deployed VLMs current |
| 2026-07-02 | Towards Real-World Ultrasound Understanding: Large Vision-Language Models from Multi-Image Examinations with Long-Form Reports | — | Multi-image ultrasound VLM with long-form reports; practical medical LVLM with realistic data regime |
| 2026-07-02 | AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models | — | — |
| 2026-07-02 | Teaching Vision-Language-Action Models What to See and Where to Look | ECCV 2026 | VLA training data design for autonomous driving; addresses vision-text imbalance in VLA supervision |
| 2026-07-02 | SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video | — | Unified 3D spatial reasoning in video; directly relevant to embodied VLM spatial understanding |
| 2026-07-02 | DeepGaze3.5-VL: Modeling Scanpaths via Autoregressive Token Prediction | — | Autoregressive scanpath modeling via VLM token prediction; novel attention modeling approach |
| 2026-07-02 | COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows | — | Self-evolving skill harnesses for image workflows; directly relevant to agent memory and reusable skills |
| 2026-07-02 | MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding | — | Graph-aware VLM for molecular images; novel architecture bridging chemistry and multimodal understanding |
| 2026-07-02 | MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding | — | Time-aware medical video benchmark; proactive prediction timing critical for clinical AI systems |
| 2026-07-02 | Seek to Segment: Active Perception for Panoramic Referring Segmentation | ECCV 2026 | Active perception for panoramic segmentation; bridges VLMs and embodied AI agent perception |
| 2026-07-02 | Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias | — | VLM reliability under corruption for medical imaging; rigorous evaluation of deployment robustness |
| 2026-07-02 | SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models | — | Binarization compression for large VLMs; practical inference efficiency for deployment |
| 2026-07-02 | ProCal: Inference-Time Proposal Calibration for Open-Vocabulary Object Detection | — | Inference-time calibration for open-vocab detection; plug-in improvement without retraining |
| 2026-07-02 | FlowCIR: Semantic Transport via Flow Matching for Zero-Shot Composed Image Retrieval | — | Zero-shot composed image retrieval via flow matching; novel semantic transport approach for VLMs |
| 2026-07-02 | Rank-Then-Act: Reward-Free Control from Frame-Order Progress | — | VLM as reward-free ordinal progress scorer for control; directly relevant to agent harness design |
| 2026-07-02 | The Moving Eye: Enhancing VLA Spatial Generalization via Hybrid Dynamic Data Collection | — | Hybrid dynamic data collection for VLA spatial generalization; actionable recipe for robot VLM training |
| 2026-07-02 | LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension | ECCV 2026 | ECCV egocentric video REC benchmark; tests long-form temporal grounding in VLMs |
| 2026-07-02 | Domain Generalization via Text-Anchored Information Bottleneck | ECCV 2026 | Text-anchored information bottleneck for domain generalization; VLM-grounded invariant representation method |
| 2026-07-02 | LIME: Learning Intent-aware Camera Motion from Egocentric Video | — | Active vision for robotics; intent-aware camera motion from egocentric video, novel VLM application |
| 2026-07-02 | Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots | — | Portable VLA/WAM inference runtime; directly addresses real-world embodied AI deployment gap |
| 2026-07-02 | Multimodal Fusion for Fine-Grained Classification of Breast Fibroadenoma and Phyllodes Tumors | — | Medical VLM fusion for breast tumor classification; fine-grained multimodal clinical decision support |
| 2026-07-02 | Search-based Testing of Vision Language Models for In-Car Scene Understanding | — | Search-based adversarial testing of VLMs for automotive safety; evaluation methodology contribution |
| 2026-07-02 | Efficient PEFT Methods with Adaptive Checkpointing for Vision Models and VLMs on Resource Constrained Consumer-GPUs | — | PEFT comparison for VLMs on consumer GPUs; practical fine-tuning on constrained hardware |
| 2026-07-02 | SINA: A Fully Automated Circuit Schematic Image to Netlist Generator Using Artificial Intelligence | — | VLM applied to circuit schematic-to-netlist; novel EDA domain grounding task with structured output |
| 2026-07-02 | EduArt: An educational-level benchmark for evaluating art history knowledge in large language models | — | Domain-specific benchmark for art history in LLMs; benchmark methodology over synthetic baselines |
| 2026-07-02 | Evaluating Vision-Language Models as a Zero-Shot Learning Alternative to You Only Look Once and Optical Character Recognition for Nigerian License Plate Recognition | — | Zero-shot VLM vs. YOLO+OCR for license plate recognition; narrow but concrete evaluation comparison |
| 2026-07-02 | RTE-FM-Dehazer: Radiative Transfer Equation Inspired Flow Matching for Real-World Image Dehazing | — | — |
| 2026-07-02 | VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon | — | — |
| 2026-07-01 | Information-Regularized Attention for Visual-Centric Reasoning | ECCV 2026 | ECCV; tackles hallucination, grounding, forgetting via attention regularization — core VLM reliability |
| 2026-07-01 | AdaBoosting Text Prompts for Vision-Language Models | ECCV 2026 | ECCV 2026; AdaBoosting text prompts improves zero-shot VLM classification, practical and reproducible |
| 2026-07-01 | LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives | — | Non-contrastive VL pretraining without negatives — architectural shift away from CLIP-style objectives |
| 2026-07-01 | DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding | ECCV | ECCV; difficulty-adaptive routing for zero-shot video grounding, strong venue, novel routing idea |
| 2026-07-01 | Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning | — | Decouples perception/reasoning for fine-grained VQA — addresses critical VLM failure mode |
| 2026-07-01 | Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts | ECCV 2026 | ECCV 2026; domain arithmetic for VLA one-shot adaptation, directly relevant to robot VLM deployment |
| 2026-07-01 | StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning | ECCV 2026 | ECCV 2026; stochastic turn depth for visual instruction tuning addresses training-inference discrepancy |
| 2026-07-01 | MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos | — | Counterfactual spatial reasoning benchmark probes genuine VLM scene understanding beyond observation |
| 2026-07-01 | GenAU: Language-Grounded Industrial Anomaly Understanding with Vision-Language Models | — | VLM for industrial anomaly detection+localization+explanation — practical grounded multimodal system |
| 2026-07-01 | Personalized Object Identification and Localization via In-Context Inference with Vision-Language Models | — | In-context inference for personalized object localization — few-shot VLM grounding without finetuning |
| 2026-07-01 | Selective Test-Time Debiasing for CLIP via Reward Gating | — | CLIP debiasing via reward gating; test-time, training-free; addresses VLM fairness at deployment |
| 2026-07-01 | Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs | ECCV 2026 | Tactile modality in MLLMs; mask-isolated alignment; ECCV 2026; fills physical-grounding gap |
| 2026-07-01 | What's Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models | — | VLMs for occlusion-aware autonomous planning; novel selective attention over hidden agents |
| 2026-07-01 | DroneFINE: Domain-Aware Parameter-Efficient Fine-Tuning of Vision-Language Detectors for Drone Images | — | Domain-aware PEFT for VLM detectors on UAV imagery; parameter-efficient adaptation recipe |
| 2026-07-01 | LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models | — | Long-video quality understanding benchmark for LVLMs; temporal distortion evaluation gap |
| 2026-07-01 | Foundation Model-driven Key Anatomy Frame Selection for Blind-sweep Ultrasound Fetal Birth Weight Estimation | MICCAI 2026 | MICCAI 2026; foundation model for ultrasound frame selection; low-resource medical VLM |
| 2026-07-01 | Human-Centric Transferable Tactile Pre-Training for Dexterous Robotic Manipulation | — | Tactile pre-training transferable to dexterous manipulation; cross-modal grounding for robotics |
| 2026-07-01 | Attribute-Prompted Kernel Hashing for Unsupervised Data-Efficient Cross-Modal Retrieval | — | Unsupervised cross-modal hashing with attribute prompts; data-efficient retrieval baseline |
| 2026-07-01 | DroneIQA-VLE: Multi-Task Drone Image Quality Assessment via Vision-Language Ensemble | — | VLM ensemble for drone image quality assessment; ICME 2026 challenge solution |
| 2026-07-01 | Discrete Diffusion Language Models for Interactive Radiology Report Drafting | — | — |
| 2026-07-01 | Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task | ECCV 2026 | Probes VLM vision vs. language bias on depth perception; diagnostic for building reliable VLMs |
| 2026-07-01 | CPG-PAD: Concept-Informed Prompts Guided Presentation Attack Detection | — | Concept-informed prompting for face PAD; practical VLM prompting technique for security-critical vision |
| 2026-07-01 | ESC: Emotional Self-Correction for Reliable Vision-Language Models | ECCV | Emotional self-correction for VLMs without post-training; lightweight reliability fix applicable broadly |
| 2026-07-01 | GMO-E\(^2\)DIT: Grounded Multi-Operation Editing for E-Commerce Images | — | Multi-operation grounded image editing for e-commerce; tests compositional VLM instruction following |
| 2026-07-01 | MIBE: Multi-subject Interaction Benchmark and Evaluator for Personalized Image Generation | — | Multi-subject personalized image generation benchmark; evaluates identity-binding in generative VLMs |
| 2026-07-01 | Multi-modal Rail Crossing Safety Analysis | — | Multimodal safety analysis fusing vision and structured accident reports; applied safety-critical domain |
| 2026-07-01 | Towards Developing a Multimodal Chat Assistant for University Stakeholders: RAG-based Approach | — | RAG-based multimodal university assistant; practical retrieval-augmented VLM deployment case study |
June 2026 (109)¶
May 2026 (20)¶
| Date | Paper | Venue | Why selected |
|---|---|---|---|
| 2026-05-30 | Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding | ICML 2026 | ICML 2026; decomposes on-policy distillation gradients for VL visual grounding |
| 2026-05-30 | Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs | CVPR 2026 | CVPR 2026; PRISM: principle-aware multi-scale evaluation of visual designs |
| 2026-05-29 | Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization | ICML 2026 | ICML 2026; visual contrastive DPO for multimodal hallucination mitigation |
| 2026-05-29 | Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models | ICML 2026 | Mrinmaya Sachan (ETH); ICML 2026; systematic study of test-time compute in VLMs |
| 2026-05-29 | How can embedding models bind concepts? | ICML 2026 | Seong Joon Oh (Tübingen); ICML 2026; concept binding failure in CLIP-style models |
| 2026-05-29 | Immuno-VLM: Immunizing Large Vision-Language Models via Generative Semantic Antibodies for Open-World Trustworthiness | ICML 2026 | ICML 2026; immunizing VLMs against open-world adversarial semantic vulnerabilities |
| 2026-05-28 | Grounded 3D-Aware Spatial Vision-Language Modeling | CVPR 2026 | Ligeng Zhu (MIT/NVIDIA); CVPR 2026; unified 2D+3D grounding in spatial VLMs |
| 2026-05-28 | Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning | CVPR 2026 | Yi Ma (Berkeley) + Joseph Tighe (Meta); CVPR 2026; 3D spatial priors improve VLM geometry |
| 2026-05-28 | Unveiling the Visual Counting Bottleneck in Vision-Language Models | ICML 2026 | Mrinmaya Sachan (ETH); ICML 2026; reveals visual counting extrapolation bottleneck |
| 2026-05-28 | Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning | ICML 2026 | ICML 2026; breaks tail alignment in CLIP for cross-domain few-shot learning |
| 2026-05-27 | Self-Prophetic Decoding to Unlock Visual Search in LVLMs | ICML 2026 | ICML 2026; self-prophetic decoding enables visual search in LVLMs |
| 2026-05-24 | Interpretability Transfer from Language to Vision via Sparse Autoencoders | ICML 2026 | ICML 2026; sparse autoencoders transfer language interpretability to vision |
| 2026-05-24 | Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation | ICML 2026 | ICML 2026; in-depth analysis + simple mitigation of language bias in LVLMs |
| 2026-05-22 | CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception | ICML 2026 | ICML 2026; cognitive visual search for high-resolution image perception in MLLMs |
| 2026-05-22 | Multimodal Distribution Matching for Vision-Language Dataset Distillation | CVPR 2026 | CVPR 2026; multimodal distribution matching for vision-language dataset distillation |
| 2026-05-22 | Turning Adaptation into Assets: Cross-Domain Bridging for Online Vision-Language Navigation | ICML 2026 | ICML 2026; test-time adaptation for VLN under non-stationary distribution shifts |
| 2026-05-22 | SPACENUM: Revisiting Spatial Numerical Understanding in VLMs | — | — |
| 2026-05-21 | Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? | CVPR 2026 | CVPR 2026; reveals VLM benchmark scores don't require grounded visual evidence |
| 2026-05-21 | From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model | ICML 2026 | ICML 2026; learning generalized behavioral representations for VLA distribution shift |
| 2026-05-21 | AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation | CVPR 2026 | CVPR 2026; self-awareness in VLN agents for reasoning over visual environments |
April 2026 (33)¶
| Date | Paper | Venue | Why selected |
|---|---|---|---|
| 2026-04-30 | Improving Calibration in Test-Time Prompt Tuning for Vision-Language Models via Data-Free Flatness-Aware Prompt Pretraining | CVPR 2026 | Flatness-aware pretraining improves calibration in VLM test-time prompt tuning |
| 2026-04-29 | Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning | CVPR | Qualitative reasoning mitigates optical illusion failures in frozen VLMs; CVPR |
| 2026-04-28 | Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models | CVPR 2026 | Novel prefill-time intervention eliminates LVLM hallucinations; inference-time method |
| 2026-04-27 | Improving Vision-language Models with Perception-centric Process Reward Models | CVPR | Process-level reward models for VLM RLVR; finer than outcome-only supervision |
| 2026-04-27 | LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models | ICLR 2026 | Attention-based token pruning rethought; attention scores don't indicate importance |
| 2026-04-27 | Jailbreaking Frontier Foundation Models Through Intention Deception | CVPR 2026 | Jailbreaks via intention deception bypass safety training; CVPR 2026 novel vector |
| 2026-04-27 | ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning | ICML 2026 | Rebuilds 3D spatial reasoning evaluation exposing VLM annotation biases; ICML 2026 |
| 2026-04-27 | Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-System | ACL 2026 | Asynchronous coarse-to-fine dual system balances VLA learning equilibrium; ACL 2026 |
| 2026-04-27 | SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs | CVPR 2026 | Soft modality routing in MoE-VLMs; first study of modality-specific expert allocation |
| 2026-04-27 | World-R1: Reinforcing 3D Constraints for Text-to-Video Generation | ICML 2026 | RL reinforces 3D geometric consistency in text-to-video generation; ICML 2026 |
| 2026-04-25 | Exploring Hierarchical Consistency and Unbiased Objectness for Open-Vocabulary Object Detection | CVPR 2026 | Hierarchical consistency with unbiased objectness for CLIP open-vocab detection |
| 2026-04-24 | DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning | CVPR 2026 | Background/question-aware token pruning for efficient document VQA; CVPR 2026 |
| 2026-04-23 | Ramen: Robust Test-Time Adaptation of Vision-Language Models with Active Sample Selection | CVPR 2026 | Active sample selection for robust CLIP test-time adaptation; CVPR 2026 |
| 2026-04-22 | Building a Precise Video Language with Human-AI Oversight | CVPR 2026 | CVPR 2026 scalable video-language human-AI oversight; datasets and benchmarks |
| 2026-04-22 | Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation | ACL 2026 | Hallucination-free fine-tuning without performance degradation; ACL 2026 |
| 2026-04-22 | Breaking the Illusion: When Positive Meets Negative in Multimodal Decoding | CVPR 2026 | Training-free positive-negative decoding reduces object hallucination; CVPR 2026 |
| 2026-04-22 | OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model | ACL 2026 | Olympiad-level multi-image reasoning benchmark for frontier VLMs; ACL 2026 |
| 2026-04-22 | Evian: Towards Explainable Visual Instruction-tuning Data Auditing | ACL 2026 | Explainable visual instruction data auditing for LVLM training quality |
| 2026-04-21 | CAPruner: Conceptual-Adjacent Scene Graph Pruner for Enhancing 3D Spatial Reasoning of Large Language Models | ACL 2026 | Scene graph pruning removes irrelevant anchors for LLM 3D spatial reasoning |
| 2026-04-21 | Environmental Understanding Vision-Language Model for Embodied Agent | CVPR | Environmental understanding VLM improves instruction-following embodied agent generalization |
| 2026-04-20 | Hierarchically Robust Zero-shot Vision-language Models | CVPR 26 | Hierarchically robust VLM fine-tuning preserves zero-shot under adversarial attack |
| 2026-04-20 | S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models | ACL 2026 | Hardness-aware DPO closes multi-image reasoning gap in VLM alignment |
| 2026-04-20 | Test-Time Perturbation Learning with Delayed Feedback for Vision-Language-Action Models | CVPR 2026 | Test-time perturbation learning corrects VLA trajectory overfitting; CVPR 2026 |
| 2026-04-20 | Enhancing Continual Learning of Vision-Language Models via Dynamic Prefix Weighting | CVPR 2026 | Dynamic prefix weighting for domain-class incremental VLM learning; CVPR 2026 |
| 2026-04-20 | From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models | ACL 2026 | Neuron-level causal attribution reveals cross-task VLM mechanisms; ACL 2026 |
| 2026-04-19 | More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage | ACL 2026 | High visual fidelity hinders abstract VLM reasoning; semiotic gap study |
| 2026-04-19 | Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception | ACL 2026 | Cold-start agentic VLM grounding without expensive supervised trajectory tuning |
| 2026-04-18 | EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling | CVPR 2026 | Evolutionary semantic labeling for visual token compression; CVPR 2026 efficiency |
| 2026-04-18 | SIF: Semantically In-Distribution Fingerprints for Large Vision-Language Models | CVPR 2026 | In-distribution semantic fingerprints for VLM IP protection; CVPR 2026 |
| 2026-04-17 | Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow | CVPR 2026 | Adaptive info flow fixes VLMs that see but don't perceive; CVPR 2026 |
| 2026-04-17 | TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Models | CVPR 2026 | Test-time textual learning improves CLIP-based OOD detection without training |
| 2026-04-17 | Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization | CVPR | World-scale geolocalization analysis reveals systematic VLM perception failures; CVPR |
| 2026-04-15 | One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding | NEURIPS 2025 | Extreme compression to one token per frame enables long-video VLMs; NeurIPS 2025 |
March 2026 (43)¶
| Date | Paper | Venue | Why selected |
|---|---|---|---|
| 2026-03-31 | Robust Multimodal Safety via Conditional Decoding | ACL 2026 | ACL 2026; conditional decoding restores safety alignment in multimodal LLMs |
| 2026-03-31 | A Comprehensive Information-Decomposition Analysis of Large Vision-Language Models | ICLR 2026 | ICLR 2026; information-decomposition reveals if VLMs use true multimodal fusion or shortcuts |
| 2026-03-31 | Hierarchical Pre-Training of Vision Encoders with Large Language Models | CVPR | CVPR; hierarchical joint pre-training tightly couples vision encoders with LLMs |
| 2026-03-31 | AGFT: Alignment-Guided Fine-Tuning for Zero-Shot Adversarial Robustness of Vision-Language Models | CVPR 2026 | CVPR 2026; alignment-guided fine-tuning preserves zero-shot adversarial robustness in VLMs |
| 2026-03-30 | Explaining CLIP Zero-shot Predictions Through Concepts | CVPR 2026 | CVPR 2026; concept bottleneck explanations make CLIP zero-shot predictions transparent |
| 2026-03-30 | Learning Multi-View Spatial Reasoning from Cross-View Relations | CVPR 2026 | CVPR 2026; cross-view relational learning for multi-view spatial reasoning in VLMs |
| 2026-03-30 | ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation | CVPR 2026 | CVPR 2026; comprehensive real-world benchmark for VLA manipulation evaluation |
| 2026-03-30 | FlowHijack: A Dynamics-Aware Backdoor Attack on Flow-Matching Vision-Language-Action Models | CVPR 2026 | CVPR 2026; first backdoor attack targeting flow-matching VLA models like π0 |
| 2026-03-30 | \(AutoDrive\text{-}P^3\): Unified Chain of Perception-Prediction-Planning Thought via Reinforcement Fine-Tuning | ICLR 2026 | ICLR 2026; chain of perception-prediction-planning with RL for autonomous driving VLMs |
| 2026-03-29 | On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language Models | CVPR 2026 | CVPR 2026; drift-aware dynamic MoE token assignment for continual multimodal learning |
| 2026-03-28 | Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models | CVPR 2026 | CVPR 2026; diagnoses root cause and mitigates hallucination in multimodal chain-of-thought |
| 2026-03-28 | Diagnosing and Repairing Unsafe Channels in Vision-Language Models via Causal Discovery and Dual-Modal Safety Subspace Projection | CVPR 2026 | CVPR 2026; causal discovery identifies unsafe channels; dual-modal safety subspace repair |
| 2026-03-28 | Structural Graph Probing of Vision-Language Models | CVPR | CVPR; structural graph probing reveals neural topology organization in VLMs |
| 2026-03-28 | SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning | CVPR 2026 | CVPR 2026; layered geometry-language fusion for reliable 3D spatial VLM reasoning |
| 2026-03-28 | ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding | CVPR 2026 | CVPR 2026; million-scale high-quality chart dataset for geometry-numerical-language reasoning |
| 2026-03-28 | Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark | CVPR 2026 | CVPR 2026; scene-aware benchmark reveals catastrophic forgetting in long-video VLMs |
| 2026-03-28 | VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation | CVPR 2026 | CVPR 2026; VLM-driven spatiotemporal reasoning for referring video object segmentation |
| 2026-03-28 | Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving | ECCV 2026 | ECCV 2026; interleaved world modeling and planning unifies AD perception and action |
| 2026-03-28 | MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence | CVPR 2026 | CVPR 2026; medical VLM with lesion detection, tracking, and visual explainability |
| 2026-03-27 | A Provable Energy-Guided Test-Time Defense Boosting Adversarial Robustness of Large Vision-Language Models | CVPR | CVPR; provable energy-guided test-time defense for VLM adversarial robustness |
| 2026-03-27 | HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models | CVPR 2026 | CVPR 2026; benchmark diagnosing fine-grained hand spatial reasoning failures in VLMs |
| 2026-03-27 | Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining | ICLR 2026 | ICLR 2026; disentangled forward/inverse dynamics pretraining resolves VLA misalignment |
| 2026-03-27 | Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding | CVPR 2026 | CVPR 2026; diffusion VLMs for GUI grounding challenge autoregressive model dominance |
| 2026-03-27 | FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants | CVPR 2026 | CVPR 2026; fairness-aware PEFT for multimodal LLMs in clinical high-stakes settings |
| 2026-03-26 | No Hard Negatives Required: Concept Centric Learning Leads to Compositionality without Degrading Zero-shot Capabilities of Contrastive Models | CVPR 2026 | CVPR 2026; concept-centric learning enables compositionality without zero-shot degradation |
| 2026-03-26 | HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models | CVPR 2026 | CVPR 2026; hierarchical 3D spatial understanding with perception-to-reasoning pipeline |
| 2026-03-26 | Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized Driving | CVPR 2026 | CVPR 2026; preference alignment personalizes VLA driving behavior to individual styles |
| 2026-03-26 | Demographic Fairness in Multimodal LLMs: A Benchmark of Gender and Ethnicity Bias in Face Verification | CVPR 2026 | CVPR 2026; demographic fairness benchmark for MLLM-based face verification |
| 2026-03-26 | MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models | CVPR 2026 | CVPR 2026; GRPO-based reinforcement learning improves MoE routing in VLMs |
| 2026-03-25 | From Weights to Concepts: Data-Free Interpretability of CLIP via Singular Vector Decomposition | CVPR 2026 | CVPR 2026; data-free CLIP interpretability via SVD; no activation data required |
| 2026-03-25 | Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification | CVPR 2026 | CVPR 2026; attention imbalance rectification directly mitigates object hallucination |
| 2026-03-25 | LensWalk: Agentic Video Understanding by Planning How You See in Videos | CVPR 2026 | CVPR 2026; agentic lens planning for dense temporal video understanding |
| 2026-03-25 | Unleashing Vision-Language Semantics for Deepfake Video Detection | CVPR 2026 | CVPR 2026; VLM semantic priors boost generalizable deepfake video detection |
| 2026-03-24 | VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions | CVPR 2026 | CVPR 2026; sparse dynamic VL interactions for LVLM efficiency without information bottleneck |
| 2026-03-24 | Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding | CVPR 2026 | CVPR 2026; identifies instruction-relevant image regions for information-dense VLM inputs |
| 2026-03-23 | Language Models Can Explain Visual Features via Steering | CVPR 2026 | CVPR 2026; LLMs explain SAE visual features autonomously via causal steering |
| 2026-03-23 | Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language Models | CVPR 2026 | CVPR 2026; null-space projection steering principled defense against visual jailbreaks |
| 2026-03-23 | Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models | CVPR 2026 | CVPR 2026; hyperbolic embeddings capture part-to-whole hierarchy in VL alignment |
| 2026-03-23 | Which Concepts to Forget and How to Refuse? Decomposing Concepts for Continual Unlearning in Large Vision-Language Models | CVPR 2026 | CVPR 2026; concept decomposition enables continual selective unlearning in LVLMs |
| 2026-03-22 | Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models | CVPR 2026 | CVPR 2026; uncertainty-aware distillation balances teacher vs. data guidance in MLLMs |
| 2026-03-21 | Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language Models | CVPR 2026 | CVPR 2026; predictive regularization prevents visual representation degradation during LLM training |
| 2026-03-20 | When Negation Is a Geometry Problem in Vision-Language Models | CVPR | CVPR; negation failure in CLIP reframed as geometry problem with novel fix |
| 2026-03-20 | PersonaVLM: Long-Term Personalized Multimodal LLMs | CVPR 2026 | CVPR 2026; multi-turn long-term personalized preference alignment in multimodal LLMs |
January 2026 (205)¶
| Date | Paper | Venue | Why selected |
|---|---|---|---|
| 2026-01-01 | Visual symbolic mechanisms: Emergent symbol processing in Vision Language Models | ICLR 2026 | Yoshua Bengio (Mila); ICLR 2026; symbolic binding mechanisms in VLMs |
| 2026-01-01 | ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning | ICLR 2026 | Linjie Li (Microsoft); ICLR 2026; multimodal interleaved chain-of-thought emergent properties |
| 2026-01-01 | Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation | ICLR 2026 | Serge Belongie (Cornell Tech); ICLR 2026; comprehensive fine-grained LVLM evaluation |
| 2026-01-01 | Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning | ICLR 2026 | Serena Yeung-Levy (Stanford); ICLR 2026; evolving tool libraries for 3D spatial reasoning |
| 2026-01-01 | Scaling up Memory for Robotic Control via Experience Retrieval | ICLR 2026 | Chelsea Finn (Stanford); ICLR 2026; scaling memory for robot policies via retrieval |
| 2026-01-01 | WorldGym: World Model as An Environment for Policy Evaluation | ICLR 2026 | Percy Liang (Stanford); ICLR 2026; world-model as policy evaluation environment |
| 2026-01-01 | Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning | ICLR 2026 | Tsung-Yi Lin + Yen-Chen Lin; ICLR 2026; fine-tuning video models for visuomotor control |
| 2026-01-01 | HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model | ICLR 2026 | Renrui Zhang (CUHK); ICLR 2026; hybrid diffusion+autoregressive VLA model |
| 2026-01-01 | InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models | ICLR 2026 | InternVL team (Zhe Chen); ICLR 2026; large-scale spatial reasoning dataset for VLMs |
| 2026-01-01 | Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving | ICLR 2026 | Hang Zhao (Tsinghua); ICLR 2026; discrete diffusion VLA for end-to-end autonomous driving |
| 2026-01-01 | Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs | ICLR 2026 | Bohyung Han (SNU); ICLR 2026; mechanistic analysis of information flow in VideoLLMs |
| 2026-01-01 | Zebra-CoT: A Dataset for Interleaved Vision-Language Reasoning | ICLR 2026 | ICLR 2026; Zebra-CoT large dataset for interleaved vision-language reasoning |
| 2026-01-01 | Self-Improving Vision-Language-Action Models with Data Generation via Residual RL | ICLR 2026 | ICLR 2026; self-improving VLA via residual RL, reduces reliance on human demos |
| 2026-01-01 | Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models | ICLR 2026 | ICLR 2026; test-time matching unlocks compositional reasoning in frontier VLMs |
| 2026-01-01 | Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded Verification | ICLR 2026 | ICLR 2026; agreement bias mitigation via self-grounded verification in MLLMs |
| 2026-01-01 | Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization | ICLR 2026 | ICLR 2026; importance-sampling multi-negative multimodal DPO for VLMs |
| 2026-01-01 | Hallucination-aware Intermediate Representation Edit in Large Vision-Language Models | ICLR 2026 | ICLR 2026; intermediate representation editing for hallucination mitigation |
| 2026-01-01 | Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models | ICLR 2026 | ICLR 2026; dynamic activation steering identifies truthfulness/visual directions |
| 2026-01-01 | Matched Data, Better Models: Target Aligned Data Filtering with Sparse Autoencoders | ICLR 2026 | ICLR 2026; SAE-based target-aligned data filtering for VLM pretraining |
| 2026-01-01 | Label-Free Mitigation of Spurious Correlations in VLMs using Sparse Autoencoders | ICLR 2026 | ICLR 2026; label-free SAE-based spurious correlation removal in VLMs |
| 2026-01-01 | MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse | ICLR 2026 | ICLR 2026; first RL framework for 3D spatial reasoning in VLMs for metaverse |
| 2026-01-01 | Pursuing Minimal Sufficiency in Spatial Reasoning | ICLR 2026 | ICLR 2026; minimal sufficiency framework for spatial reasoning bottlenecks |
| 2026-01-01 | Do 3D Large Language Models Really Understand 3D Spatial Relationships? | ICLR 2026 | ICLR 2026; exposes 3D LLMs do not truly understand spatial relationships |
| 2026-01-01 | SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs | ICLR 2026 | ICLR 2026; SpinBench: perspective-taking diagnostic for spatial reasoning in VLMs |
| 2026-01-01 | ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis | ICLR 2026 | ICLR 2026; RLVR agentic data synthesis for complex video reasoning |
| 2026-01-01 | VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding | ICLR 2026 | Bhiksha Raj (CMU); ICLR 2026; bootstrapping scalable MLLM-as-judge for video |
| 2026-01-01 | VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning? | ICLR 2026 | ICLR 2026; VideoReasonBench: vision-centric complex video reasoning benchmark |
| 2026-01-01 | FOCUS: Efficient Keyframe Selection for Long Video Understanding | ICLR 2026 | ICLR 2026; FOCUS efficient keyframe selection for hour-long video MLLMs |
| 2026-01-01 | Memento: Toward an All-Day Proactive Assistant for Ultra-Long Streaming Video | ICLR 2026 | ICLR 2026; Memento: proactive assistant for ultra-long streaming video |
| 2026-01-01 | SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning | ICLR 2026 | ICLR 2026; SimpleVLA-RL: scaling VLA training via reinforcement learning |
| 2026-01-01 | X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model | ICLR 2026 | ICLR 2026; X-VLA: soft-prompted transformer for cross-embodiment VLA |
| 2026-01-01 | Interleave-VLA: Enhancing Robot Manipulation with Image-Text Interleaved Instructions | ICLR 2026 | ICLR 2026; interleaved image-text instructions improve robot manipulation |
| 2026-01-01 | PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model | ICLR 2026 | ICLR 2026; PixelVLA: pixel-level understanding in vision-language-action models |
| 2026-01-01 | HAMLET: Switch Your Vision-Language-Action Model into a History-Aware Policy | ICLR 2026 | ICLR 2026; HAMLET: history-aware policy for VLA manipulation models |
| 2026-01-01 | TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models | ICLR 2026 | ICLR 2026; TwinVLA: data-efficient bimanual manipulation via VLAs |
| 2026-01-01 | Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation | ICLR 2026 | ICLR 2026; Embodied-R1: unified pointing representation for generalist manipulation |
| 2026-01-01 | From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation | ICLR 2026 | ICLR 2026; bridging reasoning to action for novel robotic manipulation scenarios |
| 2026-01-01 | Hybrid Training for Vision-Language-Action Models | ICLR 2026 | ICLR 2026; hybrid VLA training combining CoT with diffusion-based actions |
| 2026-01-01 | Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denosing Diffusion Process | ICLR 2026 | ICLR 2026; unified discrete diffusion VLA with future image understanding |
| 2026-01-01 | Verifier-free Test-Time Sampling for Vision-Language-Action Models | ICLR 2026 | ICLR 2026; verifier-free test-time scaling for VLA precision tasks |
| 2026-01-01 | Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models | ICLR 2026 | ICLR 2026; identifies fundamental bottlenecks in VLM safety fine-tuning |
| 2026-01-01 | Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning | ICLR 2026 | ICLR 2026; spurious correlations undermine safety fine-tuning; machine unlearning fix |
| 2026-01-01 | ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks | ICLR 2026 | ICLR 2026; adaptive plug-and-play red-teaming agents against multimodal models |
| 2026-01-01 | Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots | ICLR 2026 | ICLR 2026; structural input perturbations reveal multimodal alignment blind spots |
| 2026-01-01 | AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference Optimization | ICLR 2026 | ICLR 2026; AdPO: preference optimization for adversarial robustness in LVLMs |
| 2026-01-01 | BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning | ICLR 2026 | ICLR 2026; BEAT: contrastive visual backdoor attacks on VLM embodied agents |
| 2026-01-01 | Do Vision-Language Models Respect Contextual Integrity in Location Disclosure? | ICLR 2026 | ICLR 2026; VLMs violate contextual integrity in location privacy disclosure |
| 2026-01-01 | TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models | ICLR 2026 | ICLR 2026; TrustGen: dynamic benchmarking platform for generative model trustworthiness |
| 2026-01-01 | Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation | ICLR 2026 | ICLR 2026; Kaleidoscope: massively multilingual in-language VLM evaluation |
| 2026-01-01 | RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding | ICLR 2026 | ICLR 2026; RAVENEA: retrieval-augmented visual culture understanding benchmark |
| 2026-01-01 | Vision Language Models are Biased | ICLR 2026 | ICLR 2026; systematic evidence that VLMs exhibit strong prior-knowledge biases |
| 2026-01-01 | MMReD: a Cross-Modal Benchmark for Dense Context Reasoning | ICLR 2026 | ICLR 2026; MMReD: dense multi-modal context reasoning over long inputs |
| 2026-01-01 | ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation | ICLR 2026 | ICLR 2026; ChartGalaxy: large-scale infographic chart understanding and generation |
| 2026-01-01 | GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation | ICLR 2026 | ICLR 2026; GeoBench: hierarchical geometric reasoning evaluation for VLMs |
| 2026-01-01 | Unified Vision–Language Modeling via Concept Space Alignment | ICLR 2026 | ICLR 2026; V-SONAR: extends multilingual text embeddings to 1500-language vision |
| 2026-01-01 | VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models | ICLR 2026 | ICLR 2026; VisCodex: unified multimodal code generation merging vision and code models |
| 2026-01-01 | DiVE-k: DIFFERENTIAL VISUAL REASONING FOR FINE-GRAINED IMAGE RECOGNITION | ICLR 2026 | Ram Nevatia (USC); ICLR 2026; differential visual reasoning for fine-grained recognition |
| 2026-01-01 | Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning | ICLR 2026 | ICLR 2026; Sparse CLIP co-optimizes interpretability and performance in CLIP |
| 2026-01-01 | Post-hoc Probabilistic Vision-Language Models | ICLR 2026 | ICLR 2026; post-hoc probabilistic VLMs with calibrated uncertainty in joint space |
| 2026-01-01 | Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation | ICLR 2026 | ICLR 2026; speculative drafting approach for information-intensive visual reasoning |
| 2026-01-01 | Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning | ICLR 2026 | ICLR 2026; game-based multimodal verifiable data boosts VLM general reasoning |
| 2026-01-01 | ProxyThinker: Test-Time Guidance through Small Visual Reasoners | ICLR 2026 | ICLR 2026; small proxy VLMs guide test-time reasoning of large VLMs cheaply |
| 2026-01-01 | Vision-SR1: Self-Rewarding Vision-Language Model via Reasoning Decomposition and Multi-Reward Policy Optimization | ICLR 2026 | ICLR 2026; Vision-SR1: self-rewarding VLM via multi-reward policy optimization |
| 2026-01-01 | ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Models | ICLR 2026 | ICLR 2026; ViPER: self-evolution for fine-grained visual perception in VLMs |
| 2026-01-01 | Vision-Zero: Scalable VLM Self-Evolution via Multi-Agent Self-Play | ICLR 2026 | ICLR 2026; Vision-Zero: multi-agent self-play for label-free VLM self-evolution |
| 2026-01-01 | CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs | ICLR 2026 | ICLR 2026; attention distillation from large to small MLLMs for compositional reasoning |
| 2026-01-01 | Imitating the Truth: Attention-aware Truth-Guided Enhancement for Hallucination Mitigation in Large Vision-Language Models | ICLR 2026 | ICLR 2026; attention-aware truth-guided enhancement for LVLM hallucination |
| 2026-01-01 | Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs | ICLR 2026 | ICLR 2026; probes disconnect between visual attention and answer correctness in VLMs |
| 2026-01-01 | AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models | ICLR 2026 | ICLR 2026; empirical study of attention+diversity for adaptive visual token pruning |
| 2026-01-01 | LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models | ICLR 2026 | ICLR 2026; rethinks attention-based token pruning in VLMs |
| 2026-01-01 | Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity | ICLR 2026 | ICLR 2026; synergistic importance-diversity token compression for VLMs |
| 2026-01-01 | WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models | ICLR 2026 | ICLR 2026; weighted SVD for fast low-precision VLM execution |
| 2026-01-01 | Embodied Navigation Foundation Model | ICLR 2026 | ICLR 2026; foundation model for embodied navigation leveraging VLM priors |
| 2026-01-01 | CitySeeker: How Do VLMs Explore Embodied Urban Navigation with Implicit Human Needs? | ICLR 2026 | ICLR 2026; VLMs interpreting implicit human needs in urban embodied navigation |
| 2026-01-01 | Talking Points: Describing and Localizing Pixels | ICLR 2026 | ICLR 2026; Talking Points: pixel-precise keypoint comprehension via natural language |
| 2026-01-01 | Decomposition of Concept-Level Rules in Visual Scenes | ICLR 2026 | ICLR 2026; decomposing concept-level visual rules for compositional scene parsing |
| 2026-01-01 | DaVinci: Reinforcing Visual-Structural Syntax in MLLMs for Generalized Scientific Diagram Parsing | ICLR 2026 | ICLR 2026; DaVinci: RL-driven visual-structural syntax for scientific diagram parsing |
| 2026-01-01 | Mordal: Automated Pretrained Model Selection for Vision Language Models | ICLR 2026 | ICLR 2026; Mordal: automated pretrained model selection for building VLMs |
| 2026-01-01 | EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing | ICLR 2026 | ICLR 2026; EditReward: human-aligned reward model for instruction-guided image editing |
| 2026-01-01 | VLMgineer: Vision-Language Models as Robotic Toolsmiths | ICLR 2026 | ICLR 2026; VLMs as robotic toolsmiths — tool design as measure of physical intelligence |
| 2026-01-01 | VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision–Language Models | ICLR 2026 | ICLR 2026; VITA: test-time adaptation of VLMs as zero-shot goal-conditioned value functions |
| 2026-01-01 | DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning | ICLR 2026 | RL incentivizes visual thinking in VLMs; novel image-based reasoning paradigm |
| 2026-01-01 | VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use | ICLR 2026 | VLMs learn multimodal tool use via RL; key agentic VLM capability |
| 2026-01-01 | Revisiting Multimodal Positional Encoding in Vision–Language Models | ICLR 2026 | Comprehensive multimodal RoPE analysis; foundational for VLM positional encoding |
| 2026-01-01 | To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models | ICLR 2026 | Mechanistic analysis of visual information pathways and attention sinks in LVLMs |
| 2026-01-01 | Generative Universal Verifier as Multimodal Meta-Reasoner | ICLR 2026 | Generative universal verifier enables VLM meta-reasoning and self-refinement |
| 2026-01-01 | Cross-Modal Redundancy and the Geometry of Vision–Language Embeddings | ICLR 2026 | Geometric probe of VL joint embedding space via cross-modal redundancy |
| 2026-01-01 | OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning | ICLR 2026 | OneTwoVLA: unified VLA with adaptive reasoning-acting switching for robotics |
| 2026-01-01 | Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting | ICLR 2026 | Fine-tuning VLMs into VLAs without catastrophic forgetting; solves core challenge |
| 2026-01-01 | Rethinking Causal Mask Attention for Vision-Language Inference | ICLR 2026 | Rethinks causal masking in autoregressive VLMs; non-trivial architectural insight |
| 2026-01-01 | PPE: Positional Preservation Embedding for Token Compression in Multimodal Large Language Models | ICLR 2026 | Positional information preserved during token merging prevents spatial degradation |
| 2026-01-01 | GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models | ICLR 2026 | Test-time safety alignment for LVLMs; input detection plus output steering |
| 2026-01-01 | Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models | ICLR 2026 | Universal transferable jailbreak attacks on VLMs; multimodal attack surface analysis |
| 2026-01-01 | From Sure" toSorry": Detecting Jailbreak in Large Vision Language Model via JailNeurons |
ICLR 2026 | JailNeurons: fast neuron-based jailbreak detection for LVLMs; ICLR 2026 |
| 2026-01-01 | Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP | ICLR 2026 | Mechanistic defense against typographic attacks in CLIP; ICLR 2026 analysis |
| 2026-01-01 | Identifying Robust Neural Pathways: Few-Shot Adversarial Mask Tuning for Vision-Language Models | ICLR 2026 | Adversarial mask tuning identifies robust neural pathways in VLMs |
| 2026-01-01 | Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness | ICLR 2026 | Inference-time compute improves OOD robustness in VLMs; ICLR 2026 analysis |
| 2026-01-01 | Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes | ICLR 2026 | Egocentric multi-view spatial reasoning benchmark; VLMs struggle with 3D |
| 2026-01-01 | SpatiaLab: Can Vision–Language Models Perform Spatial Reasoning in the Wild? | ICLR 2026 | In-the-wild spatial reasoning benchmark challenges current VLM spatial capabilities |
| 2026-01-01 | OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models | ICLR 2026 | Comprehensive 3D spatial reasoning benchmark with diverse relationship categories |
| 2026-01-01 | IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs | ICLR 2026 | South-Asian cultural VLM benchmark exposes Western-centric evaluation blindspots |
| 2026-01-01 | Contamination Detection for VLMs Using Multi‑Modal Semantic Perturbations | ICLR 2026 | Multi-modal semantic perturbations detect VLM benchmark data contamination |
| 2026-01-01 | IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs | ICLR 2026 | Image-grounded video benchmark tests VLMs on cross-modal video comprehension |
| 2026-01-01 | Spotlight on Token Perception for Multimodal Reinforcement Learning | ICLR 2026 | Token-level perceptual spotlight improves RLVR visual reasoning in VLMs |
| 2026-01-01 | SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start | ICLR 2026 | Self-distilled cold start decouples visual/language learning for RLVR VLMs |
| 2026-01-01 | P\(^2\)-DPO: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization | ICLR 2026 | Calibration DPO grounds hallucination in perceptual processing; preference learning |
| 2026-01-01 | Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation | ICLR 2026 | Action-aware dynamic pruning enables efficient VLA inference; ICLR 2026 |
| 2026-01-01 | villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models | ICLR 2026 | villa-X: enhanced latent action modeling for VLA models; architectural advance |
| 2026-01-01 | Robust Fine-tuning of Vision-Language-Action Robot Policies via Parameter Merging | ICLR 2026 | Parameter merging preserves generalist VLA skills during task-specific fine-tuning |
| 2026-01-01 | Memory-Free Continual Learning with Null Space Adaptation for Zero-Shot Vision-Language Models | ICLR 2026 | Null space adaptation enables memory-free continual VLM learning; preserves zero-shot |
| 2026-01-01 | Enhanced Continual Learning of Vision-Language Models with Model Fusion | ICLR 2026 | Model fusion enhances continual VLM learning; mitigates catastrophic forgetting |
| 2026-01-01 | KeepLoRA: Continual Learning with Residual Gradient Adaptation | ICLR 2026 | KeepLoRA: residual gradient adaptation balances VLM plasticity and stability |
| 2026-01-01 | Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard Problems | ICLR 2026 | Real-world Bongard problems test abstract visual concept learning in VLMs |
| 2026-01-01 | PoSh: Using Scene Graphs to Guide LLMs-as-a-Judge for Detailed Image Descriptions | ICLR 2026 | Scene graph-guided LLM-as-judge improves detailed image description evaluation |
| 2026-01-01 | K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge | ICLR 2026 | K-Sort corrects VLM-as-judge bias; scalable visual generation preference evaluation |
| 2026-01-01 | Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory | ICLR 2026 | Multimodal Item Response Theory enables rigorous VLM reasoning benchmark design |
| 2026-01-01 | DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection | ICLR 2026 | DeCo-DETR decoupled cognition achieves efficient open-vocabulary detection; ICLR 2026 |
| 2026-01-01 | Fantastic Tractor-Dogs and How Not to Find Them With Open-Vocabulary Detectors | ICLR 2026 | OVDs produce confident false positives on negative images; critical deployment flaw |
| 2026-01-01 | VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning | ICLR 2026 | RL temporal focusing on relevant frames enables efficient long-video VLM reasoning |
| 2026-01-01 | CaReBench: A Fine-grained Benchmark for Video Captioning and Retrieval | ICLR 2026 | Fine-grained video captioning and retrieval benchmark reveals VLM temporal limits |
| 2026-01-01 | STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning | ICLR 2026 | RL reduces hallucination in spatial-temporal video grounding; ICLR 2026 |
| 2026-01-01 | Enhancing Multi-Image Understanding through Delimiter Token Scaling | ICLR 2026 | Delimiter token scaling resolves cross-image leakage in multi-image VLMs |
| 2026-01-01 | QLIP: A Dynamic Quadtree Vision Prior Enhances MLLM Performance Without Retraining | ICLR 2026 | Quadtree spatial prior boosts MLLM performance without retraining; ICLR 2026 |
| 2026-01-01 | ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models | ICLR 2026 | ERGO efficient high-resolution VLM processing via selective token allocation |
| 2026-01-01 | MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning | ICLR 2026 | Iterative preference learning improves mobile VLM agent chain-of-action reasoning |
| 2026-01-01 | SMAN-Bench: A Cross-System Benchmark for Mobile Agents under Single- and Multi-path, Ambiguous, and Noisy Tasks | ICLR 2026 | Cross-system mobile agent benchmark with ambiguous and noisy real-world tasks |
| 2026-01-01 | Simulation to Rules: A Dual-VLM Framework for Formal Visual Planning | ICLR 2026 | Dual-VLM bridges formal PDDL planning with visual scene understanding |
| 2026-01-01 | DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking | ICLR 2026 | DriveAgent-R1: active perception and hybrid thinking for autonomous driving VLMs |
| 2026-01-01 | Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty Quantification | ICLR 2026 | Evidential uncertainty quantification detects VLM misbehaviors on shifted inputs |
| 2026-01-01 | Revisiting Confidence Calibration for Misclassification Detection in VLMs | ICLR 2026 | Standard calibration degrades misclassification detection in VLMs; key theoretical finding |
| 2026-01-01 | Asynchronous Matching with Dynamic Sampling for Multimodal Dataset Distillation | ICLR 2026 | Asynchronous matching for multimodal dataset distillation; efficient VLM training |
| 2026-01-01 | Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks | ICLR 2026 | Adversarial referring expression tasks expose gaps in VLM visual reasoning |
| 2026-01-01 | WIMFRIS: WIndow Mamba Fusion and Parameter Efficient Tuning for Referring Image Segmentation | ICLR 2026 | Window Mamba fusion with parameter-efficient tuning for referring segmentation |
| 2026-01-01 | LINK: Learning Instance-level Knowledge from Vision-Language Models for Human-Object Interaction Detection | ICLR 2026 | Instance-level VLM knowledge for HOI detection; balances generalization and specialization |
| 2026-01-01 | WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM | ICLR 2026 | WAVE: unified audio-visual VLM embeddings outperform dedicated audio-visual pretraining |
| 2026-01-01 | Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds | ICLR 2026 | Hyperbolic manifold alignment for hierarchical multimodal feature fusion in VLMs |
| 2026-01-01 | MoRA: Missing Modality Low-Rank Adaptation for Visual Recognition | ICLR 2026 | Low-rank adaptation handles missing modalities in VLM visual recognition tasks |
| 2026-01-01 | Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization | ICLR 2026 | Test-time query refinement bridges modality gap in multimodal retrieval with VLMs |
| 2026-01-01 | DAVE: A VLM Vision Encoder for Document Understanding and Web Agents | ICLR 2026 | DAVE encoder captures low-level spatial structure improving document VLM performance |
| 2026-01-01 | Adaptive Logit Adjustment for Debiasing Multimodal Language Models | ICLR 2026 | Adaptive logit adjustment debiases VLM captioning and VQA generation tasks |
| 2026-01-01 | WRING Out The Bias: A Rotation-Based Alternative To Projection Debiasing | ICLR 2026 | Rotation-based debiasing removes CLIP spurious correlations; outperforms projection |
| 2026-01-01 | Read the Room: Video Social Reasoning with Mental-Physical Causal Chains | ICLR 2026 | Video social reasoning benchmark with mental-physical causal chains; VLM ToM |
| 2026-01-01 | ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations | ICLR 2026 | Interleaved text-image reasoning chains mitigate relation hallucination in VLMs |
| 2026-01-01 | Hallucination Reduction with CASAL: Contrastive Activation Steering for Amortized Learning | ICLR 2026 | Contrastive activation steering reduces hallucination via linear knowledge representations |
| 2026-01-01 | Reasoning-Driven Multimodal LLM for Domain Generalization | ICLR 2026 | MLLM reasoning improves domain generalization beyond visual feature invariance |
| 2026-01-01 | Discrete Latent Features Ablate Adversarial Attack: A Robust Prompt Tuning Framework for VLMs | ICLR 2026 | Discrete latent features resist adversarial attacks in VLM prompt tuning |
| 2026-01-01 | Inducing Dyslexia in Vision Language Models | ICLR 2026 | Mechanistic VLM study reveals VWFA analog in CLIP vision encoders |
| 2026-01-01 | AQuA: Toward Strategic Response Generation for Ambiguous Visual Questions | ICLR 2026 | Strategic response generation for ambiguous VQA; addresses real-world uncertainty |
| 2026-01-01 | VL-JEPA: Joint Embedding Predictive Architecture for Vision-language | ICLR 2026 | Novel JEPA architecture for VLMs; continuous embedding prediction vs. autoregressive tokens |
| 2026-01-01 | No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers | ICLR 2026 | Georgia Gkioxari/Meta; label-free visual reasoning training via multimodal verifiers |
| 2026-01-01 | Unleashing Perception-Time Scaling to Multimodal Reasoning Models | ICLR 2026 | Extends inference-time compute scaling to VLM perception; frontier capability paper |
| 2026-01-01 | PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs | ICLR 2026 | Philip Torr/Oxford; multimodal grounding corrects MLLM fine-grained visual failures |
| 2026-01-01 | MaskInversion: Localized Embeddings via Optimization of Explainability Maps | ICLR 2026 | Ferrari/Rupprecht (Google/Oxford); data-free localized CLIP embeddings via mask optimization |
| 2026-01-01 | RegionReasoner: Region-Grounded Multi-Round Visual Reasoning | ICLR 2026 | Cees Snoek; iterative region-grounded multi-round visual reasoning beyond single-pass VLMs |
| 2026-01-01 | Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding | ICLR 2026 | ICLR 2026; chain-of-embedding contrast exposes language prior dominance in LVLMs |
| 2026-01-01 | LLaVA-FA: Learning Fourier Approximation for Compressing Large Multimodal Models | ICLR 2026 | ICLR 2026; Fourier approximation compresses large multimodal models with strong efficiency |
| 2026-01-01 | Mitigating Hallucination in Vision-Language Model with Depth and Spatial-aware Key-Value Refinement | ICLR 2026 | ICLR 2026; depth and spatial-aware KV cache refinement reduces VLM hallucination |
| 2026-01-01 | AFTER: Mitigating the Object Hallucination of LVLM via Adaptive Factual-Guided Activation Editing | ICLR 2026 | ICLR 2026; adaptive factual-guided activation editing targets object hallucination categories |
| 2026-01-01 | Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models | ICLR 2026 | ICLR 2026; query-adaptive entropy-based contrastive decoding mitigates LVLM hallucination |
| 2026-01-01 | Cat-PO: Cross-modal Adaptive Token-rewards for Preference Optimization in Truthful Multimodal LLMs | ICLR 2026 | ICLR 2026; cross-modal token-level reward optimization reduces MLLM hallucination |
| 2026-01-01 | GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs | ICLR 2026 | ICLR 2026; adversarial image synthesis stress-tests MLLM hallucination failure modes |
| 2026-01-01 | Transferable and Stealthy Adversarial Attacks on Large Vision-Language Models | ICLR 2026 | ICLR 2026; progressive semantic infusion yields transferable stealthy LVLM attacks |
| 2026-01-01 | Seeing What’s Not There: Negation Understanding Needs More Than Training | ICLR 2026 | ICLR 2026; shows negation understanding requires more than training data scaling |
| 2026-01-01 | NePTune: A Neuro-Pythonic Framework for Tunable Compositional Reasoning on Vision-Language | ICLR 2026 | ICLR 2026; neuro-symbolic tunable compositional reasoning framework for VLMs |
| 2026-01-01 | Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence | ICLR 2026 | ICLR 2026; Cauchy-Schwarz divergence for distributional VL alignment beyond pairwise InfoNCE |
| 2026-01-01 | Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models | ICLR 2026 | ICLR 2026; unified spatial reasoning benchmark spanning robotics, AR, and navigation |
| 2026-01-01 | TABLET: A Large-Scale Dataset for Robust Visual Table Understanding | ICLR 2026 | Mirella Lapata; large-scale real-world dataset for robust visual table understanding |
| 2026-01-01 | FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting | ICLR 2026 | ICLR 2026; multi-turn frame spotlighting for efficient long-video VLM reasoning |
| 2026-01-01 | V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model Interaction | ICLR 2026 | ICLR 2026; visual prompt-based video benchmark enables richer human-model interaction |
| 2026-01-01 | A Training-Free Framework for Long Video Understanding via Video-Query-Options Similarity | ICLR 2026 | ICLR 2026; training-free long-video understanding via video-query-options similarity |
| 2026-01-01 | WMPO: World Model-based Policy Optimization for Vision-Language-Action Models | ICLR 2026 | ICLR 2026; world-model RL policy optimization for VLA; goes beyond demonstrations |
| 2026-01-01 | On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations | ICLR 2026 | ICLR 2026; systematic study of VLA robustness to multi-modal perturbations |
| 2026-01-01 | Sim2Real VLA: Zero-Shot Generalization of Synthesized Skills to Realistic Manipulation | ICLR 2026 | ICLR 2026; zero-shot sim-to-real VLA transfer via synthesized skill generalization |
| 2026-01-01 | Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance | ICLR 2026 | ICLR 2026; unified latent guidance adapts VLA models to downstream manipulation tasks |
| 2026-01-01 | QVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization | ICLR 2026 | ICLR 2026; channel-aware quantization enables VLA deployment on resource-constrained robots |
| 2026-01-01 | EVLP: Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning | ICLR 2026 | ICLR 2026; reinforced supervised fine-tuning for unified embodied VL planner |
| 2026-01-01 | WebWatcher: Breaking New Frontiers of Vision-Language Deep Research Agent | ICLR 2026 | ICLR 2026; VL deep research agent navigates visual web for complex information-seeking |
| 2026-01-01 | GhostEI-Bench: Do Mobile Agent Resilience to Environmental Injection in Dynamic On-Device Environments? | ICLR 2026 | ICLR 2026; benchmark for VLM agent resilience to environmental injection on mobile |
| 2026-01-01 | Uncertainty-Aware Gaussian Map for Vision-Language Navigation | ICLR 2026 | ICLR 2026; uncertainty-aware Gaussian map improves VLN in partially observed 3D environments |
| 2026-01-01 | Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-Language Navigation | ICLR 2026 | ICLR 2026; dual-system slow-grounding/fast-moving VLM for generalizable VLN |
| 2026-01-01 | Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks | ICLR 2026 | ICLR 2026; argues current medical VLM benchmarks mask true clinical reasoning gaps |
| 2026-01-01 | TumorChain: Interleaved Multimodal Chain-of-Thought Reasoning for Traceable Clinical Tumor Analysis | ICLR 2026 | ICLR 2026; interleaved multimodal CoT enables traceable clinical tumor analysis |
| 2026-01-01 | MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning | ICLR 2026 | ICLR 2026; annotation-free medical VLM reasoning via agentic reinforcement learning |
| 2026-01-01 | OmniCT: Towards a Unified Slice-Volume LVLM for Comprehensive CT Analysis | ICLR 2026 | ICLR 2026; unified slice-volume LVLM covering full CT image analysis pipeline |
| 2026-01-01 | MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning | ICLR 2026 | ICLR 2026; multi-agent RL framework for specialist medical VLM collaboration |
| 2026-01-01 | Towards Text-Mask Consistency in Medical Image Segmentation | ICLR 2026 | ICLR 2026; text-mask consistency constraints fix VLM multi-lesion segmentation failures |
| 2026-01-01 | Reversible Primitive–Composition Alignment for Continual Vision–Language Learning | ICLR 2026 | ICLR 2026; reversible primitive-composition alignment for tight-budget continual VL learning |
| 2026-01-01 | Pi-CCA: Prompt-Invariant CCA Certificates for Replay-Free Continual Multimodal Learning | ICLR 2026 | ICLR 2026; prompt-invariant CCA certificates enable replay-free continual VL learning |
| 2026-01-01 | RLAP-CLIP: Continual Multimodal Learning with Prototype Adaptation and Difficulty-Aware Routing | ICLR 2026 | ICLR 2026; prototype adaptation with difficulty routing for class-incremental CLIP |
| 2026-01-01 | Naming to Learn: Class Incremental Learning for Vision-Language Model with Unlabeled Data | ICLR 2026 | ICLR 2026; naming paradigm enables class incremental VLM learning from unlabeled data |
| 2026-01-01 | A-TPT: Angular Diversity Calibration Properties for Test-Time Prompt Tuning of Vision-Language Models | ICLR 2026 | ICLR 2026; angular diversity calibration addresses prompt dispersion gap in VLM TTA |
| 2026-01-01 | Flatness Guided Test-Time Adaptation for Vision-Language Models | ICLR 2026 | ICLR 2026; flatness-guided TTA links loss landscape geometry to VLM distribution shift |
| 2026-01-01 | Bilateral Information-aware Test-time Adaptation for Vision-Language Models | ICLR 2026 | ICLR 2026; bilateral information-aware TTA handles covariate shifts without test distribution |
| 2026-01-01 | Adaptive Debiasing Tsallis Entropy for Test-Time Adaptation | ICLR 2026 | ICLR 2026; Tsallis entropy debiasing corrects CLIP's built-in bias during TTA |
| 2026-01-01 | Benchmarking Bias Mitigation Toward Fairness Without Harm from Vision to LVLMs | ICLR 2026 | ICLR 2026; first comprehensive benchmark comparing bias mitigation in VLMs |
| 2026-01-01 | Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images | ICLR 2026 | ICLR 2026; VLM-based semantic anomaly detection in AI-generated image outputs |
| 2026-01-01 | Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models | ICLR 2026 | ICLR 2026; perceptually grounded chain-of-thought for faithful VLM geospatial reasoning |
| 2026-01-01 | FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark | ICLR 2026 | ICLR 2026; million-scale T2I reasoning dataset and benchmark from top labs |
| 2026-01-01 | Object-Centric Refinement for Enhanced Zero-Shot Segmentation | ICLR 2026 | ICLR 2026; object-centric region refinement enhances CLIP zero-shot segmentation |
| 2026-01-01 | Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition | ICLR 2026 | ICLR 2026; MLLM as detector-agnostic interaction recognizer for zero-shot HOI |
| 2026-01-01 | lmgame-Bench: How Good are LLMs at Playing Games? | ICLR 2026 | Eric P. Xing/CMU; game-playing benchmark requiring perception, reasoning, planning |
| 2026-01-01 | pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language Models | ICLR 2026 | ICLR 2026; personalized federated fine-tuning for CLIP with multi-modal adapters |
| 2026-01-01 | Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking | ICLR 2026 | ICLR 2026; self-supervised voting and ranking evolves VLMs for image quality assessment |
| 2026-01-01 | VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation | ICLR 2026 | ICLR 2026; VL anchoring enables one-shot generalization for bimanual manipulation |
| 2026-01-01 | Seeing What’s Wrong: A Trajectory-Guided Approach to Caption Error Detection | ICLR 2026 | ICLR 2026; trajectory-guided multi-score method detects subtle image-caption errors |