| 2026-05-31 |
SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories |
— |
SkillAdaptor self-adapts reusable skills from step-level trajectory feedback; training-free |
| 2026-05-31 |
Leyline: KV Cache Directives for Agentic Inference |
— |
KV cache directives for agentic inference; breaks chatbot-only append-only cache assumptions |
| 2026-05-31 |
Self-Healing Agentic Orchestrators for Reliable Tool-Augmented Large Language Model Systems |
— |
Self-healing orchestrators; failure taxonomy and recovery strategies for tool-augmented LLMs |
| 2026-05-31 |
Bridging Requirements and Architecture: Multi-Agent Orchestration with External Knowledge and Hierarchical Memory |
— |
Multi-agent orchestration with external knowledge and hierarchical memory for architecture |
| 2026-05-31 |
Reducing Token Usage of State-in-Context Agents using Minification |
— |
Minification reduces state-in-context agent token cost; independent SWE-bench replication |
| 2026-05-31 |
A New Framework for Cybersecurity Refusals in AI Agents |
— |
New framework for cybersecurity refusals in multi-step agentic scaffolds; beyond single-turn |
| 2026-05-31 |
Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing |
— |
Traceable multi-agent workflows preserving narrative evidence across video editing pipeline |
| 2026-05-31 |
Can LLM Agents Sustain Long-Horizon Organizational Dynamics? |
— |
Tests whether LLM agents sustain coherent behavior in structured organizational hierarchies |
| 2026-05-30 |
MOSAIC: Modular Orchestration for Structured Agentic Intelligence and Composition |
— |
MOSAIC: modular orchestration substrate for structured agentic composition |
| 2026-05-30 |
Dynamic Coordination Strategy Selection for Enterprise Multi-Agent Systems |
— |
Evidence-based guidance for choosing consensus, debate, or single-agent coordination |
| 2026-05-29 |
From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors |
— |
Trojan backdoor defense for agentic harnesses; persistent workspace attack model |
| 2026-05-29 |
Model-Native Computing Architecture: Envisioning Future System Architecture Through the Lens of Computer Architecture |
— |
Frames LLM agent scheduling as classical computer-architecture problem; novel systems lens |
| 2026-05-29 |
Learning to Construct Practical Agentic Systems |
— |
Empirical study exposing gap between research and production agentic system requirements |
| 2026-05-29 |
Stateful Online Monitoring Catches Distributed Agent Attacks |
— |
Stateful monitoring catches distributed agent attacks split across many sessions |
| 2026-05-29 |
How Generation Architecture Shapes Code Complexity in Multi-Agent LLM Systems: A Paired Study on HumanEval |
— |
Multi-agent code generation increases complexity vs single-shot; controlled paired study |
| 2026-05-28 |
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs |
— |
Multi-agent harness for editable scientific figure generation from diverse inputs |
| 2026-05-28 |
OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents |
— |
Automated auditing of the open skill ecosystem for LLM agents; lifecycle reliability lens |
| 2026-05-28 |
Scaling Laws for Agent Harnesses via Effective Feedback Compute |
— |
Scaling laws for agent harnesses via feedback compute; bridges test-time scaling to harness design |
| 2026-05-28 |
Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation |
— |
Multi-agent harness for verifiable interleaved multimodal deep research; directly named topic |
| 2026-05-28 |
Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Framework for Dynamic Scheduling |
— |
Asynchronous agent framework balancing reactive scheduling and long-horizon planning |
| 2026-05-28 |
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction |
— |
WorldMemArena: memory must track evolving world state and revise stale entries |
| 2026-05-28 |
Formalizing Mathematics at Scale |
— |
AutoformBot: multi-agent system formalizes mathematics at textbook scale in Lean 4 |
| 2026-05-28 |
Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents |
— |
Agora: LLM agents autonomously detect bugs in production-level consensus protocols |
| 2026-05-28 |
EvoRepair: Enhancing Vulnerability Repair Agents Through Experience-Based Self-Evolution |
— |
EvoRepair: cross-vulnerability experience accumulation for self-evolving repair agents |
| 2026-05-28 |
A Multi-AI-agent Framework Enabling End-to-end Finite Element Analysis for Solid Mechanics Problems |
— |
— |
| 2026-05-27 |
Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows |
— |
Harness-Bench directly benchmarks harness effects on model outputs; fills critical measurement gap |
| 2026-05-27 |
Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refinement |
— |
Prompt Codebooks: discrete compositional optimization for LLM instruction refinement |
| 2026-05-27 |
AI Research Agents Narrow Scientific Exploration |
— |
Finds AI research agents narrow scientific exploration diversity; critical warning for harness designers |
| 2026-05-27 |
SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents |
— |
SNARE: tests overeager behavior where benign-prompt agents exceed authorized scope |
| 2026-05-27 |
Multi-Agent LLM-based Metamorphic Testing for REST APIs |
— |
Multi-agent metamorphic testing for REST APIs; LLMs discover metamorphic relations |
| 2026-05-27 |
From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence |
— |
Agentic framework reproduces under-specified ML papers into benchmark-ready implementations |
| 2026-05-27 |
Adaptive Multimodal Agents-Based Framework for Automatic Workflow Execution |
— |
Adaptive multimodal agents for automatic workflow execution bridging metadata to perception |
| 2026-05-27 |
GUI-CIDER: Mid-training GUI Agents via Causal Internalization and Density-aware Exemplar Reselection |
— |
— |
| 2026-05-27 |
VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild |
— |
— |
| 2026-05-26 |
Governed Evolution of Agent Runtimes through Executable Operational Cognition |
— |
Governed harness evolution via executable operational cognition artifacts |
| 2026-05-26 |
A Policy-Driven Runtime Layer for Agentic LLM Serving |
— |
Policy-driven runtime layer separating agent framework identity from serving infrastructure |
| 2026-05-26 |
AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems |
— |
AgensFlow: coordination-policy substrate decoupling agent skills from runtime routing |
| 2026-05-26 |
UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems |
— |
Unified RL interface optimizing all roles in LLM multi-agent systems end-to-end |
| 2026-05-26 |
Testing Agentic Workflows with Structural Coverage Criteria |
— |
Structural coverage criteria for agentic workflow testing beyond task-success metrics |
| 2026-05-26 |
MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation |
— |
MUSE-Autoskill: self-evolving skill management with creation, memory, and evaluation |
| 2026-05-26 |
Lessons from Penetration Tests on Large-Scale Agent Systems |
— |
Penetration test taxonomy reveals recurring vulnerability classes across large agent systems |
| 2026-05-25 |
From Model Scaling to System Scaling: Scaling the Harness in Agentic AI |
— |
Theoretical treatment of system scaling vs model scaling; defines harness architecture requirements |
| 2026-05-25 |
Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems |
— |
Agent lifespan engineering; first systematic study of deployed agent degradation over time |
| 2026-05-25 |
Automated Benchmark Auditing for AI Agents and Large Language Models |
— |
Automated benchmark auditing finds latent bugs in expert-authored agent tasks at scale |
| 2026-05-25 |
KYA: A Framework-Agnostic Trust Layer for Autonomous Systems with Verifiable Provenance and Hierarchical Policy Composition |
— |
KYA: framework-agnostic trust layer with verifiable provenance and policy composition |
| 2026-05-25 |
Security of OpenClaw Agents: Fundamentals, Attacks, and Countermeasures |
— |
Security fundamentals, attacks, countermeasures for persistent OpenClaw agent frameworks |
| 2026-05-25 |
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly |
— |
CyberEvolver: self-evolving cybersecurity agent adapts scaffold on-the-fly to targets |
| 2026-05-24 |
Meta-Agent: From Task Descriptions to Verified Multi-Agent Systems |
— |
Auto-generates verified multi-agent systems from natural-language task descriptions |
| 2026-05-23 |
DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations |
— |
Harness evolution via demonstrations; sparse-feedback sample-efficient adaptation paradigm |
| 2026-05-22 |
When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems |
— |
Epistemic miscalibration: novel failure mode where agents misjudge plan feasibility |
| 2026-05-22 |
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills |
— |
SkillEvolBench: benchmarks distillation of episodic experience into reusable procedural skills |
| 2026-05-22 |
EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation |
— |
— |
| 2026-05-21 |
Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost |
— |
Compiles orchestration workflows into LLM weights; claims 100× cost reduction over frameworks |
| 2026-05-21 |
The Log is the Agent: Event-Sourced Reactive Graphs for Auditable, Forkable Agentic Systems |
— |
Event-sourced reactive graphs; log-first design for auditable forkable agentic systems |
| 2026-05-21 |
MOSS: Self-Evolution through Source-Level Rewriting in Autonomous Agent Systems |
— |
MOSS: runtime self-evolution via source-level agent code rewriting without weight updates |
| 2026-05-21 |
Benchmarking Autonomous Agents against Temporal, Spatial, and Semantic Evasions |
— |
Benchmarks agents against temporal, spatial, semantic evasions in persistent systems |
| 2026-05-21 |
ExComm: Exploration-Stage Communication for Error-Resilient Agentic Test-Time Scaling |
— |
ExComm: exploration-stage communication reduces error propagation in long-horizon agents |
| 2026-05-21 |
One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systems |
— |
— |
| 2026-05-20 |
Toward AI VIS Co-Scientists: A General and End-to-End Agent Harness for Solving Complex Data Visualization Tasks |
— |
End-to-end agent harness for scientific visualization; concrete harness architecture case study |
| 2026-05-20 |
Causal Past Logic for Runtime Verification of Distributed LLM Agent Workflows |
— |
Causal past logic for correct runtime monitoring of asynchronous distributed agent workflows |
| 2026-05-20 |
Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems |
— |
Goal-level energy accounting metric; new efficiency unit for multi-step agentic AI |
| 2026-05-20 |
RMA: an Agentic System for Research-Level Mathematical Problems |
— |
RMA: agentic framework targeting research-level mathematics beyond competition benchmarks |
| 2026-05-19 |
BOHM: Zero-Cost Hierarchical Attribution for Compound AI Systems |
— |
Zero-cost Shapley attribution for compound AI without expensive coalition re-evaluation |
| 2026-05-19 |
Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs |
— |
Mix-Quant: quantized prefilling with precise decoding cuts agentic LLM input overhead |
| 2026-05-19 |
AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows |
— |
Retrieval-based synthesis of interoperable multi-agent workflows for open-ended science |
| 2026-05-19 |
MuMuTestUp: Mutation-based Multi-Agent Test Case Update |
— |
MuMuTestUp: mutation-based multi-agent test update for CI/CD code evolution |
| 2026-05-19 |
EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design |
— |
EngiAI: multi-agent benchmark combining simulation, retrieval, and manufacturing preparation |
| 2026-05-19 |
STAR-PólyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision |
— |
STAR-PólyaMath: persistent meta-strategic supervision for long-horizon math reasoning |
| 2026-05-18 |
Code as Agent Harness |
— |
Establishes code as executable agent harness; foundational paradigm shift for agentic systems |
| 2026-05-18 |
PROTEA: Offline Evaluation and Iterative Refinement for Multi-Agent LLM Workflows |
ACL 2026 |
ACL; offline evaluation and iterative refinement pipeline for multi-agent LLM workflows |
| 2026-05-18 |
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows |
— |
DecisionBench: benchmarks emergent delegation across 11 models in long-horizon workflows |
| 2026-05-18 |
MINTEval: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems |
— |
MINTEval: evaluates memory under multi-target interference in long-horizon agent systems |
| 2026-05-18 |
Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search |
— |
Lean Refactor: plug-and-play harness for multi-objective controllable proof optimization |
| 2026-05-18 |
MMoA: An AI-Agent framework with recurrence for Memoried Mixure-of-Agent |
— |
MMoA: recurrent memoried mixture-of-agent with dynamic context-aware routing |
| 2026-05-18 |
STRIDE: A Self-Reflective Agent Framework for Reliable Automatic Equation Discovery |
— |
STRIDE: self-reflective agent avoids generation-only loops in symbolic equation discovery |
| 2026-05-18 |
It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs |
— |
— |
| 2026-05-17 |
Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security |
— |
Comprehensive survey of safety, robustness, privacy, security in agentic AI systems |
| 2026-05-17 |
MemRepair: Hierarchical Memory for Agentic Repository-Level Vulnerability Repair |
— |
MemRepair: hierarchical memory enabling repository-scale vulnerability repair agents |
| 2026-05-17 |
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration |
— |
— |
| 2026-05-16 |
S-Bus: Automatic Read-Set Reconstruction for Multi-Agent LLM State Coordination |
— |
S-Bus: automatic read-set reconstruction preventing multi-agent structural race conditions |
| 2026-05-16 |
Responsible Agentic AI Requires Explicit Provenance |
— |
Explicit provenance as foundation for responsible autonomous agentic AI deployment |
| 2026-05-16 |
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents |
— |
AgentKernelArena: generalization-aware benchmark for GPU kernel optimization agents |
| 2026-05-16 |
Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework |
— |
Systematic analysis of three agent interaction paradigms within unified production framework |
| 2026-05-15 |
TopoEvo: A Topology-Aware Self-Evolving Multi-Agent Framework for Root Cause Analysis in Microservices |
— |
TopoEvo: topology-aware self-evolving multi-agent framework for root cause analysis |
| 2026-05-15 |
BootstrapAgent: Distilling Repository Setup into Reusable Agent Knowledge |
— |
BootstrapAgent: distills costly repository setup knowledge into reusable agent memory |
| 2026-05-15 |
STAR: A Stage-attributed Triage and Repair framework for RCA Agents in Microservices |
— |
STAR: stage-attributed triage and repair improves reliability of RCA agents |
| 2026-05-15 |
RTL-BenchMT: Dynamic Maintenance of RTL Generation Benchmark Through Agent-Assisted Analysis and Revision |
— |
RTL-BenchMT: agentic dynamic maintenance prevents benchmark staleness in EDA research |
| 2026-05-15 |
An Agentic Retrieval Framework for Autonomous Context-Aware Data Quality Assessment |
— |
Agentic retrieval framework for autonomous context-aware data quality assessment |
| 2026-05-15 |
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? |
— |
CHI-Bench: end-to-end automation of policy-dense healthcare workflows via agents |
| 2026-05-14 |
Is Grep All You Need? How Agent Harnesses Reshape Agentic Search |
— |
Isolates harness contributions to agentic search; grep-sufficiency ablation study |
| 2026-05-14 |
LEMON: Learning Executable Multi-Agent Orchestration via Counterfactual Reinforcement Learning |
NEURIPS 2026 |
NEURIPS; RL learns multi-agent orchestration topology outperforming hand-designed workflows |
| 2026-05-14 |
APWA: A Distributed Architecture for Parallelizable Agentic Workflows |
— |
Distributed architecture enabling parallelizable agentic workflow execution at scale |
| 2026-05-14 |
Orchard: An Open-Source Agentic Modeling Framework |
— |
Orchard: open-source agentic modeling framework addressing open-research bottleneck |
| 2026-05-14 |
Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution |
— |
Solvita: stateful self-evolving agent accumulates experience for competitive programming |
| 2026-05-14 |
Veritas: A Semantically Grounded Agentic Framework for Memory Corruption Vulnerability Detection in Binaries |
— |
Veritas: semantically grounded agentic framework for binary memory-corruption detection |
| 2026-05-14 |
Auditing Agent Harness Safety |
— |
First audit of harness safety; unauthorized access via correct-looking trajectories |
| 2026-05-13 |
Position: Agentic AI System Is a Foreseeable Pathway to AGI |
ICML 26 |
ICML position; argues agentic AI systems are necessary pathway to AGI beyond scaling |
| 2026-05-13 |
Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context |
— |
— |
| 2026-05-11 |
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation |
— |
Real-world long-horizon CLI harness benchmark against synthetic sandboxes and mock servers |
| 2026-05-11 |
NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation |
— |
Co-evolving skills, memory, and policy for personalized multi-agent research automation |
| 2026-05-11 |
Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents |
— |
— |
| 2026-05-09 |
Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection |
— |
— |
| 2026-05-08 |
Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge |
— |
— |
| 2026-05-07 |
Retrieval-Conditioned Topology Selection with Provable Budget Conservation for Multi-Agent Code Generation |
NEURIPS 2026 |
NEURIPS; retrieval-conditioned topology selection for code-generation multi-agent routing |
| 2026-05-05 |
SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agents |
— |
Portable skill compilation across LLM agent frameworks; cross-framework interoperability |
| 2026-05-05 |
Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies |
— |
Workspace-Bench: AI agents on real heterogeneous file-dependency workspace tasks |