| 2025-11-30 |
Agentic Persona Control and Task State Tracking for Realistic User Simulation in Interactive Scenarios |
NEURIPS 2025 |
— |
| 2025-11-29 |
CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency |
— |
— |
| 2025-11-28 |
Towards Corpus-Grounded Agentic LLMs for Multilingual Grammatical Analysis |
— |
— |
| 2025-11-28 |
Towards Continuous Intelligence Growth: Self-Training, Continual Learning, and Dual-Scale Memory in SuperIntelliAgent |
— |
— |
| 2025-11-28 |
TIM-PRM: Verifying multimodal reasoning with Tool-Integrated PRM |
— |
— |
| 2025-11-27 |
Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI Automation |
CVPR 2026 |
— |
| 2025-11-27 |
Solving Context Window Overflow in AI Agents |
— |
— |
| 2025-11-27 |
Automated Design Optimization via Strategic Search with Large Language Models |
— |
— |
| 2025-11-27 |
NOMAD: A Multi-Agent LLM System for UML Class Diagram Generation from Natural Language Requirements |
— |
— |
| 2025-11-27 |
Co-Evolving Agents: Learning from Failures as Hard Negatives |
— |
— |
| 2025-11-27 |
A Safety and Security Framework for Real-World Agentic Systems |
— |
— |
| 2025-11-26 |
Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework |
— |
— |
| 2025-11-26 |
EWE: An Agentic Framework for Extreme Weather Analysis |
— |
— |
| 2025-11-26 |
AI Urban Scientist: Multi-Agent Collaborative Automation for Urban Research |
— |
— |
| 2025-11-25 |
MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology |
NEURIPS 2025 |
— |
| 2025-11-25 |
EnergyTwin: A Multi-Agent System for Simulating and Coordinating Energy Microgrids |
— |
— |
| 2025-11-25 |
VICoT-Agent: A Vision-Interleaved Chain-of-Thought Framework for Interpretable Multimodal Reasoning and Scalable Remote Sensing Analysis |
— |
— |
| 2025-11-25 |
Learning Multi-Access Point Coordination in Agentic AI Wi-Fi with Large Language Models |
— |
— |
| 2025-11-24 |
Agint: Agentic Graph Compilation for Software Engineering Agents |
NEURIPS 2025 |
— |
| 2025-11-24 |
A Multi-Agent LLM Framework for Multi-Domain Low-Resource In-Context NER via Knowledge Retrieval, Disambiguation and Reflective Analysis |
AAAI 2026 |
— |
| 2025-11-24 |
Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion |
CVPR 2026 |
— |
| 2025-11-24 |
HeaRT: A Hierarchical Circuit Reasoning Tree-Based Agentic Framework for AMS Design Optimization |
— |
— |
| 2025-11-24 |
IRSDA: An Agent-Orchestrated Framework for Enterprise Intrusion Response |
— |
— |
| 2025-11-24 |
Beyond Protein Language Models: An Agentic LLM Framework for Mechanistic Enzyme Design |
— |
— |
| 2025-11-24 |
Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration |
— |
— |
| 2025-11-24 |
AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents |
— |
— |
| 2025-11-24 |
Addressing Situated Teaching Needs: A Multi-Agent Framework for Automated Slide Adaptation |
— |
— |
| 2025-11-24 |
MAGMA-Edu: Multi-Agent Generative Multimodal Framework for Text-Diagram Educational Question Generation |
— |
— |
| 2025-11-23 |
\(A^2Flow:\) Automating Agentic Workflow Generation via Self-Adaptive Abstraction Operators |
AAAI |
— |
| 2025-11-23 |
FHE-Agent: Automating CKKS Configuration for Practical Encrypted Inference via an LLM-Guided Agentic Framework |
— |
— |
| 2025-11-23 |
End-to-End Automated Logging via Multi-Agent Framework |
— |
— |
| 2025-11-23 |
LLMs as Firmware Experts: A Runtime-Grown Tree-of-Agents Framework |
— |
— |
| 2025-11-23 |
A Multimodal Conversational Agent for Tabular Data Analysis |
— |
— |
| 2025-11-23 |
DiscoVerse: Multi-Agent Pharmaceutical Co-Scientist for Traceable Drug Discovery and Reverse Translation |
— |
— |
| 2025-11-23 |
Hybrid Agentic AI and Multi-Agent Systems in Smart Manufacturing |
— |
— |
| 2025-11-22 |
A superpersuasive autonomous policy debating system |
AAAI 2026 |
— |
| 2025-11-22 |
ARIAL: An Agentic Framework for Document VQA with Precise Answer Localization |
— |
— |
| 2025-11-21 |
Episodic Memory in Agentic Frameworks: Suggesting Next Tasks |
— |
— |
| 2025-11-21 |
AutoLink: Autonomous Schema Exploration and Expansion for Scalable Schema Linking in Text-to-SQL at Scale |
— |
— |
| 2025-11-20 |
Trustworthy AI in the Agentic Lakehouse: from Concurrency to Governance |
— |
— |
| 2025-11-20 |
Distributed Agent Reasoning Across Independent Systems With Strict Data Locality |
— |
— |
| 2025-11-20 |
ChemLabs on ChemO: A Multi-Agent System for Multimodal Reasoning on IChO 2025 |
— |
— |
| 2025-11-20 |
InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution |
— |
— |
| 2025-11-20 |
Hiding in the AI Traffic: Abusing MCP for LLM-Powered Agentic Red Teaming |
— |
— |
| 2025-11-20 |
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference |
— |
— |
| 2025-11-19 |
What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity |
— |
— |
| 2025-11-19 |
Know Your Intent: An Autonomous Multi-Perspective LLM Agent Framework for DeFi User Transaction Intent Mining |
— |
— |
| 2025-11-19 |
Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization |
— |
— |
| 2025-11-19 |
Multi-Agent LLM Orchestration Achieves Deterministic, High-Quality Decision Support for Incident Response |
— |
— |
| 2025-11-19 |
Knowledge-Informed Automatic Feature Extraction via Collaborative Large Language Model Agents |
— |
— |
| 2025-11-19 |
Beyond GeneGPT: A Multi-Agent Architecture with Open-Source LLMs for Enhanced Genomic Question Answering |
— |
— |
| 2025-11-18 |
AutoTool: Efficient Tool Selection for Large Language Model Agents |
AAAI 2026 |
— |
| 2025-11-18 |
Tell Me: An LLM-powered Mental Well-being Assistant with RAG, Synthetic Dialogue Generation, and Agentic Planning |
ACL |
— |
| 2025-11-18 |
Agentic AI Systems in Electrical Power Systems Engineering: Current State-of-the-Art and Challenges |
— |
— |
| 2025-11-18 |
MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents |
— |
— |
| 2025-11-18 |
DataSage: Multi-agent Collaboration for Insight Discovery with External Knowledge Retrieval, Multi-role Debating, and Multi-path Reasoning |
— |
— |
| 2025-11-18 |
APD-Agents: A Large Language Model-Driven Multi-Agents Collaborative Framework for Automated Page Design |
— |
— |
| 2025-11-17 |
Extracting Events Like Code: A Multi-Agent Programming Framework for Zero-Shot Event Extraction |
AAAI 2026 |
— |
| 2025-11-17 |
Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly? |
— |
— |
| 2025-11-17 |
P1: Mastering Physics Olympiads with Reinforcement Learning |
— |
— |
| 2025-11-17 |
Multi-Agent Multimodal Large Language Model Framework for Automated Interpretation of Fuel Efficiency Analytics in Public Transportation |
— |
— |
| 2025-11-17 |
MedDCR: Learning to Design Agentic Workflows for Medical Coding |
— |
— |
| 2025-11-16 |
Co-Layout: LLM-driven Co-optimization for Interior Layout |
AAAI 2026 |
— |
| 2025-11-16 |
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs |
— |
— |
| 2025-11-16 |
Multi-agent Self-triage System with Medical Flowcharts |
— |
— |
| 2025-11-15 |
Mobile-Agent-RAG: Driving Smart Multi-Agent Coordination with Contextual Knowledge Empowerment for Long-Horizon Mobile Automation |
— |
— |
| 2025-11-15 |
GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory |
— |
— |
| 2025-11-14 |
Refine and Align: Confidence Calibration through Multi-Agent Interaction in VQA |
AAAI 2026 |
— |
| 2025-11-14 |
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding |
— |
— |
| 2025-11-14 |
ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation |
— |
— |
| 2025-11-14 |
NOVA: An Agentic Framework for Automated Histopathology Analysis and Discovery |
— |
— |
| 2025-11-14 |
AIonopedia: an LLM agent orchestrating multimodal learning for ionic liquid discovery |
— |
— |
| 2025-11-14 |
Learning to Refine: An Agentic RL Approach for Iterative SPARQL Query Construction |
— |
— |
| 2025-11-14 |
Generative Caching for Structurally Similar Prompts and Responses |
— |
— |
| 2025-11-13 |
Towards an Agentic Workflow for Internet Measurement Research |
— |
— |
| 2025-11-13 |
Evaluating Prompting Strategies with MedGemma for Medical Order Extraction |
— |
— |
| 2025-11-13 |
Mastering Olympiad-Level Physics with Artificial Intelligence |
— |
— |
| 2025-11-13 |
SCARE: A Benchmark for SQL Correction and Question Answerability Classification for Reliable EHR Question Answering |
— |
— |
| 2025-11-13 |
EnvTrace: Simulation-Based Semantic Evaluation of LLM Code via Execution Trace Alignment -- Demonstrated at Synchrotron Beamlines |
— |
— |
| 2025-11-13 |
MINDS: A Cross-cultural Dialogue Corpus for Social Norm Classification and Adherence Detection |
— |
— |
| 2025-11-12 |
SlideBot: A Multi-Agent Framework for Generating Informative, Reliable, Multi-Modal Presentations |
— |
— |
| 2025-11-12 |
Evaluating Software Process Models for Multi-Agent Class-Level Code Generation |
— |
— |
| 2025-11-12 |
BarrierBench: Evaluating Large Language Models for Safety Verification in Dynamical Systems |
— |
— |
| 2025-11-12 |
ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset |
— |
— |
| 2025-11-12 |
Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation |
— |
— |
| 2025-11-12 |
BioVerge: A Comprehensive Benchmark and Study of Self-Evaluating Agents for Biomedical Hypothesis Generation |
— |
— |
| 2025-11-11 |
Adaptive Multi-Agent Response Refinement in Conversational Systems |
AAAI 2026 |
— |
| 2025-11-11 |
QLCoder: A Query Synthesizer For Static Analysis of Security Vulnerabilities |
— |
— |
| 2025-11-11 |
LLM-Powered Fully Automated Chaos Engineering: Towards Enabling Anyone to Build Resilient Software Systems at Low Cost |
— |
— |
| 2025-11-10 |
FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation |
AAAI 2026 |
— |
| 2025-11-10 |
Voice-Interactive Surgical Agent for Multimodal Patient Data Control |
— |
— |
| 2025-11-10 |
Discourse Graph Guided Document Translation with Large Language Models |
— |
— |
| 2025-11-10 |
AgentSUMO: An Agentic Framework for Interactive Simulation Scenario Generation in SUMO via Large Language Models |
— |
— |
| 2025-11-09 |
Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models |
— |
— |
| 2025-11-09 |
PRAGMA: A Profiling-Reasoned Multi-Agent Framework for Automatic Kernel Optimization |
— |
— |
| 2025-11-09 |
WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation |
— |
— |
| 2025-11-08 |
MTTR-A: Measuring Cognitive Recovery Latency in Multi-Agent Systems |
— |
— |
| 2025-11-08 |
EduAgentQG: A Multi-Agent Workflow Framework for Personalized Question Generation |
— |
— |
| 2025-11-08 |
Self-Abstraction from Grounded Experience for Plan-Guided Policy Refinement |
— |
— |
| 2025-11-07 |
AdvisingWise: Supporting Academic Advising in Higher Education Settings Through a Human-in-the-Loop Multi-Agent Framework |
— |
— |
| 2025-11-07 |
SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models |
— |
— |
| 2025-11-06 |
RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG |
— |
— |
| 2025-11-06 |
Multi-Agent Collaborative Framework For Math Problem Generation |
— |
— |
| 2025-11-06 |
PEFA-AI: Advancing Open-source LLMs for RTL generation using Progressive Error Feedback Agentic-AI |
— |
— |
| 2025-11-05 |
AnaFlow: Agentic LLM-based Workflow for Reasoning-Driven Explainable and Sample-Efficient Analog Circuit Sizing |
— |
— |
| 2025-11-05 |
U2F: Encouraging SWE-Agent to Seize Novelty without Losing Feasibility |
— |
— |
| 2025-11-05 |
Towards Realistic Project-Level Code Generation via Multi-Agent Collaboration and Semantic Architecture Modeling |
— |
— |
| 2025-11-05 |
Toward Autonomous Engineering Design: A Knowledge-Guided Multi-Agent Framework |
— |
— |
| 2025-11-04 |
No-Human in the Loop: Agentic Evaluation at Scale for Recommendation |
NEURIPS 2025 |
— |
| 2025-11-04 |
MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning |
ACL 2026 |
— |
| 2025-11-04 |
PublicAgent: Multi-Agent Design Principles From an LLM-Based Open Data Analysis Framework |
— |
— |
| 2025-11-04 |
PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts |
— |
— |
| 2025-11-04 |
VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation |
— |
— |
| 2025-11-04 |
Towards Iterative End-to-End Software Development: A Feature-Driven Multi-Agent Framework |
— |
— |