Skip to content

Harnesses / Meta-Harnesses — 2026 · Paper list (575)

July 2026 (157)

Date Paper Venue Why selected
2026-07-23 Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
2026-07-23 GS-Agent: Creating 4D Physical Worlds With Generative Simulation
2026-07-23 EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization
2026-07-22 IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
2026-07-22 OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills
2026-07-22 PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning
2026-07-22 EvoDRC: A Self-Evolving Agentic Framework for Automated DRC Violation Repair
2026-07-22 DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
2026-07-22 Bridging Behavior and Implementation: Automated Java Glue Code Generation for Behavior-Driven Development
2026-07-21 Comparative Study of Multi-Agent Actor-Critic Algorithms in Parameterized Action Reinforcement Learning
2026-07-21 TraceDev: A Traceability-Driven Multi-agent Framework for Requirement-to-Code Development
2026-07-21 PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents
2026-07-20 How Agent Skills Fail under Long Contexts: A White-Box Study in Code Auditing White-box study of skill/instruction degradation in long-context agent trajectories; directly actionable for harness design
2026-07-20 TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory Minimization Google team; agent trajectory minimization to cut codeslop; novel objective for coding agent harnesses
2026-07-20 Is Progressive Disclosure All You Need for Long-Context Agents? Progressive disclosure as context management primitive for long-context agents; empirical comparison vs retriever
2026-07-20 (Over)Reliance on Test Agents in AI-Assisted Software Testing Empirical study of over-reliance on test agents; calibration signal for agent harness quality gates
2026-07-20 O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning ECCV 2026
2026-07-20 Natural Language Access to Domain-Specific Metadata: A Reusable Framework for LLM Query Generation
2026-07-20 FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications
2026-07-20 MAGE: Human-Like Macro Placement via Agentic Multimodal Reasoning
2026-07-20 FinSAgent: Corpus-Aligned Multi-Agent RAG Framework for Evidence-Grounded SEC Filing Question Answering
2026-07-20 SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategy Refinement in E-Commerce Recommendation
2026-07-19 Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making POMDP routing + self-correction synthesis; principled framework for autonomous LLM agent decision-making
2026-07-19 Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions
2026-07-19 When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering
2026-07-18 CLOSER-Bench: Evaluating Budgeted Cross-Stage Design Closure for Hardware Agents CLOSER-Bench: long-horizon hardware agent eval with delayed/heterogeneous feedback; novel benchmark design for harness eval
2026-07-18 Just A Rather Very Intelligent Spoken Agent Voice-interactive long-horizon agent with richer user feedback channels beyond text; modality extension for harnesses
2026-07-17 Knowledge-Centric Agents for Workflow Generation ECCV 2026
2026-07-17 AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
2026-07-17 Understanding Agent-Reactive Bugs at the Model-Harness Boundary: An Empirical Study of LLM Agent Issue Reports
2026-07-17 EpiNarrate: Agentic Generation of Grounded Narratives from Epidemiological Scenario Projections
2026-07-17 Nonuniformity Principle in Human-AI Coworking Nonuniformity principle for human-AI coworking; formal framework for where to insert human checkpoints in agentic pipelines
2026-07-17 PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation
2026-07-16 Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution Specialist vs. generalist harness architectures for structured code workflows; practical design lessons
2026-07-16 TopoAgent: A Self-Evolving Topological Agent for Multimodal Scientific Reasoning Self-evolving topological planning agent; novel non-linear orchestration for multimodal reasoning
2026-07-16 Setup Complete, Now You Are Compromised: Weaponizing Setup Instructions Against AI Coding Agents Supply-chain attack vector targeting coding agent setup phase; critical harness security finding
2026-07-16 SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration Multi-agent collaboration for long-horizon search; addresses context/progress tracking in orchestration
2026-07-16 FirmPilot: Evidence-Guided Multi-Agent Environment Recovery for IoT Firmware Rehosting Evidence-guided multi-agent recovery harness; fault-tolerance patterns for fragile tool environments
2026-07-16 AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery
2026-07-16 Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
2026-07-16 BrainPilot: Automating Brain Discovery with Agentic Research
2026-07-15 Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling KDD 2026
2026-07-15 Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
2026-07-15 Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild
2026-07-15 Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity
2026-07-15 Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System
2026-07-15 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis
2026-07-15 Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents
2026-07-15 Set-shifting Behavioral Test for Harnessed Agents
2026-07-15 Copy-on-Write Scoring: Application-Specific Agent Evaluations ICML 2026 ICML 2026; copy-on-write evaluation granularity for application-specific agent workflows
2026-07-15 Towards Reliable AI-Assisted Analog Design: Template-Constrained LLM Agents for SAR ADC Generation
2026-07-15 Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents
2026-07-15 Stop Means Stop: Measuring and Repairing the Enforcement Gap in Agent-Framework Control Primitives
2026-07-14 PalmClaw: A Native On-Device Agent Framework for Mobile Phones
2026-07-14 Software Supply Chains are Dead: Use-Case-Oriented Regeneration
2026-07-14 Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions
2026-07-14 Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
2026-07-14 RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar
2026-07-14 KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill
2026-07-14 MemoHarness: Agent Harnesses That Learn from Experience Literally defines/studies harness layer; memory-augmented harness design; most on-topic
2026-07-14 Orchestrating Power Grid Studies with Multi-Agent AI and MCP Servers MCP + multi-agent orchestration pattern for simulation tools; concrete harness-integration blueprint
2026-07-14 Enhancing Small Language Models Reasoning through Knowledge Graph Grounding
2026-07-13 ToFu: A White-Box, Token-Efficient Agent Harness for Researchers White-box token-efficient harness design; directly targets researcher-buildable agent scaffolding
2026-07-13 A Formal Hierarchical Architecture for Agentic Orchestration with Stack-Based Execution and Lazy Discovery Formal hierarchical orchestration with stack-based execution; addresses flat-registry bottleneck architecturally
2026-07-13 NVAITC AI Scientist: A Governed End-to-End Research System -- A Hypertension GWAS Case Study Governed end-to-end research agent system in biomedical; institutional deployment constraints for harnesses
2026-07-13 Efficient Test-Time Optimization for Multi-Agent Proof Autoformalization Multi-agent proof autoformalization with test-time optimization; novel orchestration for formal reasoning
2026-07-13 Fine-Tuned Multi-Agent Framework for Detecting OCEAN in Life Narratives
2026-07-13 Phantom Guardrails: When Self-Improving Agent Harnesses Fix Failures That Never Happened
2026-07-12 SETA: Scaling Environments for Terminal Agents Scales terminal agent environments; practical benchmark infrastructure for agentic eval harnesses
2026-07-12 MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference Region-aware KV cache eviction for agent inference; efficiency critical for long-horizon harness runs
2026-07-12 Towards Autonomous and Auditable Medical Imaging Model Development Autonomous MLE agent for medical imaging; concrete harness applied to buildable medical AI pipeline
2026-07-12 UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp Data-to-agent framework for multimodal web browsing; compositional tool-use and perception integration
2026-07-12 Large language model agents accelerate inverse design of metal-organic frameworks for gas separation
2026-07-11 Temporary Authority, Permanent Effects: Commit-Time Authorization for LLM Agents Novel commit-time authorization model for LLM agents; directly addresses harness safety boundaries
2026-07-11 AgentAbstain: Do LLM Agents Know When Not to Act? Abstention/self-awareness in agents; critical eval dimension missing from most harness benchmarks
2026-07-11 A Stepwise Questioning Expert-Editor Multi-Agent Framework for Long-Document Summarization
2026-07-11 Falsifiable Release Gates for Self-Improving Systems
2026-07-10 Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift Directly theorizes harness-managed agentic context evolution and scoped verification under distribution shift
2026-07-10 Failure as a Process: An Anatomy of CLI Coding Agent Trajectories Empirical anatomy of CLI agent failure trajectories; actionable for harness designers improving reliability
2026-07-10 ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning Multi-agent synthesis framework with iterative perception-hypothesis-execution loop; ARC-AGI-2 scale
2026-07-10 ProofCouncil: An LLM Agent for Solving Open Mathematical Problems Agentic workflow tailored to open math problems; council-style multi-agent orchestration pattern
2026-07-10 ReProAgent: Tool-Augmented Multi-Stage Agentic Generation of Bug Reproduction Tests from Issue Reports Multi-stage agentic pipeline with tool augmentation for bug reproduction; concrete harness recipe
2026-07-10 Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents Hypothesis evolution protocol for auditable LLM-agent science; addresses harness-level auditability
2026-07-10 Malaika: Understanding Malware through Tri-Grounded Agentic Reasoning Tri-grounded agentic reasoning under partial observability; novel reasoning scaffold for tool-use agents
2026-07-10 VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents Autonomous vulnerability exploitation agents; structured multi-step agentic harness for security testing
2026-07-10 All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language Models
2026-07-10 A Knowledge-Based Multi-Agent Framework for Security Control Recommendation
2026-07-10 Exploring Agentic Workflows for Generating High Quality Math Visual Aids
2026-07-09 TTHE: Test-Time Harness Evolution Core topic: defines harness as executable scaffold; proposes test-time harness evolution
2026-07-09 From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents Harness engineering framing for enterprise: prompt-to-contract pipeline, auditable agent design
2026-07-09 WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search Recursive multi-agent orchestration for deep-wide search; canonical meta-harness decomposition pattern
2026-07-09 Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents Proactive memory agent decouples memory from action; critical for long-horizon harness design
2026-07-09 Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Cognitive scaffolding separates modalities across context windows; novel agent architecture pattern
2026-07-09 UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks Benchmark for proactive real-world agents; evaluates harness behavior across tool-use scenarios
2026-07-09 ProjAgent: Procedural Similarity Retrieval for Repository-Level Code Generation
2026-07-09 Game Theory Driven Multi-Agent Framework Mitigates Language Model Hallucination
2026-07-09 Out of Sight: Compression-Aware Content Protection against Agentic Crawlers
2026-07-09 ASMR: Agentic Schema Generation for Ship Maintenance Report Writing
2026-07-09 Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution SOP-enforcing agentic system for procedural task execution; harness compliance enforcement pattern
2026-07-09 MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs
2026-07-08 Rethinking Code Performance Benchmarks for LLMs
2026-07-08 SpaCellAgent: A Self-Evolving LLM-Based Multi-Agent Framework for Trajectory Analysis
2026-07-08 From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents
2026-07-08 Progressive Crystallization: Turning Agent Exploration into Deterministic, Lower-Cost Workflows in Production
2026-07-08 The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
2026-07-08 PERFOPT-Bench: Evaluating Coding Agents on Software Performance Optimization First benchmark targeting performance-optimization coding agents; expands harness evaluation scope
2026-07-08 From Triggers to Emotions: A CPM-Grounded Appraisal Multi-Agent for Dynamic Emotional Evolution in Persona-Based Dialogue
2026-07-08 Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE ACL 2026 Dynamic RoPE for long agentic context; directly addresses harness token-budget and context management
2026-07-08 PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language
2026-07-07 LogicHunter: Testing LLM Agent Frameworks with an Agentic Oracle Automated oracle for testing LLM agent frameworks (LangChain/LlamaIndex/CrewAI) — critical infra gap
2026-07-07 The Balkanization of Execution-Security Research for AI Coding Agents: Isolation, Access Control, and Time-of-Check-to-Time-of-Use Vulnerabilities Systematic survey of execution-security isolation for coding agents — TOCTOU vulns in harness layer
2026-07-07 Spider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL Workflows AI-native SQL workflows extending agentic tool use to cloud data platforms — real deployment setting
2026-07-07 Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities Real-world multi-tabular benchmark exposes evaluation gaps for data analysis agents
2026-07-07 Articulating Assumptions in AI-Generated Scientific Analyses through Task Decomposition Task decomposition to surface and articulate assumptions in LLM-generated scientific analysis pipelines
2026-07-07 Detecting Vulnerability-Inducing Commits via Multi-Stage Reasoning with LLM-Based Agents Multi-stage LLM agent reasoning for security-critical commit analysis — practical agentic pipeline
2026-07-07 Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1
2026-07-07 VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery
2026-07-06 Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation Hierarchical dual-system VLA harness; generalizable pattern for long-horizon agentic control
2026-07-06 On the risk of coding before testing: An empirical study on LLM-based test generation workflow Papadakis group; empirical study of test-first vs code-first ordering in agentic SE workflows
2026-07-06 OptiAgent: End-to-End Optimization Modeling via Multi-Agent Iterative Refinement End-to-end multi-agent iterative refinement pipeline; generalizable orchestration pattern
2026-07-06 An Exploration of Agentic Information Fusion for Test Maintenance Prediction Agentic information fusion for test maintenance; novel multi-signal agent coordination pattern
2026-07-06 PDEFlow: Autonomous Agentic PDE Pipelines for Neural Operator Learning and Solver-Free Inference Full autonomous pipeline from spec to trained neural operator; meta-harness for scientific AI
2026-07-06 PatchOptic for Shared-State LLM Workflows with Projected Views and Verified Structured Updates Projected views + verified structured updates for shared-state LLM workflows — core scaffolding primitive
2026-07-06 TACTIC-KG: Toward Small Agent Teams for Cyber Threat Intelligence Knowledge Graph Construction
2026-07-06 Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory
2026-07-06 AgenticPD: A Stage-Aware Agentic Framework for Physical Design QoR Optimization
2026-07-06 Collective Intelligence with Foundation Models Multi-model coordination framework with judge/solver separation; foundational meta-harness pattern
2026-07-05 Forethought: Verifiable Reasoning from Neurosymbolic Primitive Programming Neurosymbolic primitive programming makes tool-call sequences verifiable; novel workflow abstraction
2026-07-05 Learning to Control LLM Agent Harnesses with Offline Reinforcement Learning Frames agent harness as learnable via offline RL — directly novel for harness builders
2026-07-05 Prompt-to-Paper: Agentic AI System for Bioinformatics End-to-end agentic bioinformatics manuscript system with verifiable grounding — domain harness case study
2026-07-05 Agentic SABRE: An Uncertainty-Aware Neuro-Symbolic Multi-Agent Framework for Adaptive Ransomware Detection
2026-07-04 Don't Blame the Large Language Model: How Scaffolding Evolution Shapes Coding Agent Quality Hassan group; empirical dissection of scaffolding middleware as first-class quality variable
2026-07-04 The Remarkable Effectiveness of Providing AI Agents with Natural Language Tools: A Replication Study Validating NLT Performance Across 14 Models Replicated NLT across 14 models; directly answers structured-vs-NL tool calling for harness builders
2026-07-04 ProACT: Towards Breakdown-Aware Proactive Agent in Multi-User Collaboration Proactive breakdown detection in multi-user agent collaboration; novel harness intervention trigger
2026-07-04 AutoCedar: An Agentic Framework for Verifier-Guided Access Control Policy Synthesis Verifier-guided agentic synthesis; novel safety pattern for harnesses in high-stakes domains
2026-07-04 Context Graphs for Proactive Enterprise Agents Context graphs enable proactive enterprise agents; shifts harness from reactive to anticipatory
2026-07-03 Diagnosis-Driven Automatic Repair for Agentic Workflow via Symbolic Inference Symbolic inference auto-repairs platform-orchestrated workflows; addresses harness reliability gap
2026-07-03 VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning RL trains multi-tool agentic reasoning for video; multimodal + harness design jointly optimized
2026-07-03 MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents Medical calculator agents with implicit tool selection; practical benchmark for clinical agent harnesses
2026-07-03 APeB: Benchmarking Personalization Ability of Large Language Model Agents Benchmark for LLM agent personalization; evaluation frameworks drive field standards
2026-07-02 AgentFlow: Building Agent Dependency Graphs for Static Analysis of Agent Programs Static analysis of agent dependency graphs; framework-agnostic harness structure formalization
2026-07-02 Steerability via constraints: a substrate for scalable oversight of coding agents Constraint-based steerability for coding agents; directly addresses scalable human oversight bottleneck
2026-07-02 Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational study Observational study isolating reasoning effort vs tool access in agentic code generation reliability
2026-07-02 SimWorlds: A Multi-Agent System for Dynamic 3D Scene Creation Multi-agent orchestration for dynamic scene generation; novel 4D output from LLM agent pipelines
2026-07-02 CausalSteward: An Agentic Divide-Conquer-Combine Copilot for Causal Discovery Divide-conquer-combine agentic pattern for causal discovery; reusable scaffold design
2026-07-02 KRCA: An Efficient Root Cause Analysis System in Hyper-Scale Microservice Systems via Agentic AI Production-scale agentic RCA system in hyper-scale microservices; industrial deployment evidence
2026-07-02 Language Models as Measurement Apparatus for Culture ACL 2026
2026-07-02 HULAT2 at MER-TRANS 2026: Governed Multi-Agent Simplification for Spanish Easy-to-Read Generation
2026-07-02 Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification Scale-up safety-testing pipeline for tool-using agents; risk-discovery to evidence-grounded verification
2026-07-01 Auto-FL-Research: Agentic Search for Federated Learning Algorithms Agentic search over FL algorithm space; meta-harness pattern applied to automated research
2026-07-01 AGC-Bench: Measuring Artificial General Creativity
2026-07-01 Rise From The Ashes: LLM-based Static Analysis for Deep Learning Framework Bugs
2026-07-01 VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement ECCV 2026
2026-07-01 Agent4cs: A Multi-agent System for Code Summarization in Large Hierarchical Codebases Multi-agent hierarchical code summarization; scales agent teams to large structured codebases
2026-07-01 Knowledge-Centric Information Systems Data engineering governance principles applied to LLM-era systems; architectural framing
2026-07-01 CHARLIE: An On-Premise Multi-Agent Retrieval-Augmented Generation System for Evidential Reasoning in Forensic Science On-premise multi-agent RAG for forensic evidential reasoning — structured domain harness reference design
2026-07-01 Agentic generation of verifiable rules for deterministic, self-expanding reaction classification

June 2026 (113)

Date Paper Venue Why selected
2026-06-30 Design and Implementation of Agentic Orchestrations and Orchestration of Agents Rinderle-Ma (U Vienna, BPM authority); directly about designing agent orchestration — the core harness problem
2026-06-30 ACE: Pluggable Adaptive Context Elasticizer across Agents Addresses fundamental harness bottleneck: adaptive context management across agents
2026-06-30 Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents Cohere + LG CNS; multilingual tool-use agents at 111B scale — real enterprise harness
2026-06-30 Agentic-Ideation: Sample Efficient Agentic Trajectories Synthesis for Scientific Ideation Agents Key meta-harness concept: moving beyond fixed workflows via self-evolving trajectories
2026-06-30 DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation Self-evolving multi-agent data construction — meta-harness pattern for iterative refinement
2026-06-30 An Agentic AI Framework to Accelerate Scientific Discovery in Plant Phenotyping ORNL national lab; scientific discovery agent framework from a high-credibility group
2026-06-30 CSTrader: A Testbed for Language-Grounded Trading in a Community-Driven Virtual Asset Market Novel agent testbed with grounded language-trading environment; valuable for harness eval
2026-06-30 MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning Multi-UAV LLM platform/benchmark — directly useful for comparing multi-agent harnesses
2026-06-30 SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks ICML 2026 ICML 2026; cost-routing in multi-turn agentic SWE harnesses directly actionable
2026-06-30 Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics Agentic autoformalization pipeline with Lean 4 verification; novel harness for correctness guarantees
2026-06-30 A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols Self-evolving agentic system with SOP-to-code alignment and feedback loops; generalizable scaffold
2026-06-30 TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models
2026-06-30 ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping
2026-06-30 DDIAgents: Mechanism-Conditioned Context Flow for Drug-Drug Interaction Prediction
2026-06-30 Knowledge Distillation from Large Reasoning Models to Compact Student Models: A Case Study on the John O Bryan Mathematics Competition
2026-06-30 A Semantic-Layer-Mediated Agent for Natural Language to SQL over Heterogeneous Enterprise Databases
2026-06-29 Experience Graphs: The Data Foundation for Self-Improving Agents Database-community perspective on self-improving agents; fresh infrastructure paradigm
2026-06-29 SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows End-to-end agent benchmark; benchmarks are core harness infrastructure
2026-06-29 MicroAgent: Context-Augmented Multi-Agent Framework for Automatic Microservice Decomposition Multi-agent orchestration for software decomposition; relevant harness pattern
2026-06-29 AgRefactor: Self-Evolving Agentic Workflow for HLS Compatibility and Performance Jason Cong (UCLA, ACM/IEEE Fellow, EDA giant); self-evolving agentic workflow for HLS
2026-06-29 SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon CivRealm Strategy Planning Long-horizon multi-domain agent with scene-aware goal evolution; strong harness design patterns
2026-06-29 Investigating Multi-Agent Deliberation in Law
2026-06-29 TopoAgent: An Agentic Framework for Automated Topology Learning in Medical Imaging
2026-06-28 Agent Security Meets Regulatory Reality -- A Practitioner Systematization of Autonomous-Agent Threats and Controls in Regulated Financial Systems Bridges lab security research and regulated deployment; rare practitioner systematization
2026-06-28 Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting Practitioner-focused loop-design framework; reframes harness architecture around autonomous iteration
2026-06-27 Capability Gates Are Not Authorization: Confused-Deputy Failures in LLM Agent Frameworks Audits confused-deputy failures in agent frameworks; practical security contribution to harness design
2026-06-27 Agentic Abstention: Do Agents Know When to Stop Instead of Act? Novel abstention framing for agent eval; critical for building reliable harnesses
2026-06-27 Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem Empirical characterization of real-world LLM agent workflows; valuable for harness designers
2026-06-27 LAMP: Lean-based Agentic framework with MCP and Proof Repair Novel MCP + formal verification (Lean) synthesis; proof repair for agent tool-use
2026-06-27 Why Trust Your Agent? Empirical Security Gains from TRiSM-Guided Agentic Workflows in Healthcare Security evaluation framework for healthcare agents; TRiSM-guided methodology
2026-06-27 A Task-Driven and Quality-Assured Agent Framework for SAR Data Generation
2026-06-27 From Determinism to Delegation: AI-Native Software Engineering and the Evolution of the Agentic Engineer
2026-06-26 Agentic Hardware Design as Repository-Level Code Evolution
2026-06-26 HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration ECCV 2026
2026-06-26 LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior ICML 2026
2026-06-26 Digitizing Coaching Intelligence: An Agentic Framework for Holistic Athlete Profiling using VLM and RAG
2026-06-25 Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy Meta-harness orchestrating heterogeneous cyber+physical tools with autonomous recovery; core scaffolding architecture
2026-06-25 EGG: An Expert-Guided Agent Framework for Kernel Generation Expert-guided agentic framework for GPU kernel generation; novel human-in-loop harness design with measurable compute impact
2026-06-25 Content-Based Smart E-Mail Dispatcher Using Large Language Models
2026-06-25 Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments
2026-06-25 Glite ARF: Verifier-Driven Research with Parallel LLM Coding Agents
2026-06-25 NOVA: A Verification-Aware Agent Harness for Architecture Evolution in Industrial Recommender Systems
2026-06-24 Unlocking Model Potentials Through Adaptive Multi-Agent Scaffolding for Efficient Issue Resolution Adaptive scaffolding for ambiguous issue resolution; directly models harness-level workflow design
2026-06-24 IntentTester: Intent-Driven Multi-agent Framework for Cross-Library Test Migration Multi-agent test migration harness; concrete scaffold reuse pattern across library boundaries
2026-06-24 BrainAgent: A Large Language Model-Driven Multi-Agent Framework for Autonomous Brain Signal Understanding LLM multi-agent framework for brain signal tasks; shows harness generalization to specialized sensor domains
2026-06-24 Knowledge-augmented Agentic AI for Mental Health Medication Information Seeking Multi-source knowledge-augmented agent for drug safety; concrete harness integrating regulatory + patient-narrative retrieval
2026-06-23 SHERLOC: Structured Diagnostic Localization for Code Repair Agents NVIDIA+TU Darmstadt; structured localization cuts agent budget waste in half for repo repair
2026-06-23 MANGO: Automated Multi-Agent Test Oracle Generation for Vision-Language-Action Models Lionel Briand (top SE researcher); automated test oracle generation harness for VLA evaluation
2026-06-23 PixJail: Self-Evolving Paper-to-Pipeline Reproduction for Text-to-Image Jailbreak Evaluation Self-evolving paper-to-pipeline reproduction; meta-harness that auto-generates evaluation pipelines
2026-06-23 Agentic Generation of AST Transformation Rules for Fixing Breaking Updates Martin Monperrus (KTH, program repair); agentic workflow generates migration rules at scale
2026-06-23 DeepBD: A Grounded Agentic Workflow for Variant Prioritization and Diagnosis of Genetic Birth Defects Grounded agentic workflow for clinical genetics; concrete harness design with domain-expert loop
2026-06-23 SoK: AI Secure Code Generation: Progress, Pitfalls, and Paths Forward SoK systematization of coding agent security; high-citation survey covering what harnesses must enforce for safe code gen
2026-06-23 Privacy-Preserving RAG via Multi-Agent Semantic Rewriting: Achieving Confidentiality Without Compromising Contextual Fidelity Multi-agent semantic rewriting harness for privacy-preserving RAG; reusable coordination pattern for sensitive deployments
2026-06-23 OmniPath: A Multi-Modal Agentic Framework for Auditing Wheelchair Accessibility
2026-06-22 AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction OS-level agent harness; open-source; covers tool-call, memory, cross-app orchestration
2026-06-22 From Task-Guided Conversational Graphs to Goal-Oriented Dialogue Runtimes Graph-to-runtime orchestration for multi-goal LLM conversations; practical harness wiring pattern
2026-06-22 Emergent Relational Order in LLM Agent Societies: From Collective Affect to Authority Stratification ACL 2026 ACL 2026; emergent authority stratification in LLM agent societies informs multi-agent coordination design
2026-06-22 Semantic Browsing: Controllable Diversity for Image Generation
2026-06-22 The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Practitioner-oriented full-stack agentic AI reference; foundations to production, covers harness architecture
2026-06-22 RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation Multi-stage agentic framework coupling reasoning and search; reusable scaffold pattern for OOD generation
2026-06-22 VideoAgent: All-in-One Framework for Video Understanding and Editing All-in-one video understanding/editing agent; long-horizon multi-task orchestration architecture
2026-06-22 StatABench: Dataset and Framework for Evaluating Statistical Analysis Capabilities of LLMs StatABench evaluation framework for LLM tool-use in statistical analysis; benchmark harness design reusable across domains
2026-06-22 IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPO
2026-06-21 Code Isn't Memory: A Structural Codebase Index Inside a Coding Agent Controlled study of structural indexing inside a fixed harness; actionable retrieval finding
2026-06-21 Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation Agent-native full-stack inference acceleration; practical for deploying harness-wrapped generative pipelines
2026-06-21 SVGym (SciVerseGym): An Environment for Reinforcement Learning and Bayesian Optimization in Crystal Discovery SVGym: unified RL/BO environment harness for crystal discovery; shows how to scaffold closed-loop scientific agent pipelines
2026-06-20 Revelio: Cost-Efficient Agentic Memory Safety Vulnerability Detection For Repository-Scale Codebases Cost-efficient multi-agent harness for repo-scale memory-safety detection; addresses scalability directly
2026-06-20 CodeTeam: An LLM-Powered Multi-Agent Framework for Repository-Level Code Generation Multi-agent NL2Repo framework; repo-level orchestration is the key harness benchmark frontier
2026-06-19 Building Agent Harnesses for Scientific Curation from Multimodal Sources Title literally 'Agent Harnesses'; multimodal scientific curation; strong MSR-adjacent group
2026-06-19 AutoRAS: Learning Robust Agentic Systems with Primitive Representations Automated agentic system design with primitive representations; meta-harness self-improvement
2026-06-19 Composing Verifiable Conceptual Models via Building Blocks: Towards Design-Time Verification of Agentic AI Workflows Design-time formal verification of agentic workflows; fills major harness safety gap
2026-06-19 SwarmX: Agentic Scheduling for Low-Latency Agentic Systems GPU-CPU scheduling for multi-call agentic pipelines; latency-critical harness infrastructure
2026-06-19 Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents Supervised orchestrator training for PDDL planning; end-to-end harness-to-planner bridge
2026-06-19 Agentic Time Machine as an Infrastructure for Future-Event Forecasting Time-machine infrastructure for forecasting agents; harness-as-temporal-scaffold framing
2026-06-19 BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discovery Multi-agent orchestration for biomedical discovery; interactive evidence inspection matches medical+agent interests
2026-06-19 Dementia-Agents: A Multi-Modal Multi-Agent System for Dementia Staging and Phenotyping Multi-modal multi-agent orchestration for clinical staging; strong medical-AI+agent harness fit
2026-06-19 EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory Evolvable embeddings for agentic long-context memory; foundational retrieval infrastructure every persistent agent harness needs
2026-06-19 Who Checks the Citations? Benchmarking Legal Hallucination Detection
2026-06-18 AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning
2026-06-18 Dual-Agent Framework for Cross-Model Verified Translation of Natural-Language Protocols into Robotic Laboratory Platform
2026-06-18 Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale
2026-06-18 AI Economist Agent: An Agentic Framework for Model-Grounded Economic Analysis with RAG, Knowledge Graphs, and Large Language Models
2026-06-18 Human-on-the-Loop Orchestration for AI-Assisted Legal Discovery
2026-06-18 Whose Agent Are You? Multi-Layer Fingerprinting and Attribution of Autonomous Web Agents Agent fingerprinting and attribution; critical governance layer for deploying autonomous web agents
2026-06-18 Prompt, Plan, Extract: Zero-Shot Agentic LLMs Workflows for Lung Pathology Extraction from Clinical Narratives Zero-shot agentic workflow harness for clinical NLP; concrete Prompt→Plan→Extract pipeline pattern directly reusable
2026-06-18 GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling
2026-06-17 AdsMind: A Physics-Grounded Multi-Agent System for Self-Correcting Discovery of Adsorption Configurations on Heterogeneous Catalyst Surfaces
2026-06-17 LandslideAgent with Multimodal LandslideBench: A Domain-Rule-Augmented Agent for Autonomous Landslide Identification and Analysis
2026-06-17 StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns
2026-06-17 Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark Multi-agent deep research framework + comprehensive benchmark for physical sciences
2026-06-16 SEAGym: An Evaluation Environment for Self-Evolving LLM Agents Defines and benchmarks 'agent harness' as execution layer; Tsinghua; fills key evaluation gap
2026-06-16 Environment-Grounded Automated Prompt Optimization for LLM Game Agents Lindauer & Feurer (top AutoML); automated prompt optimization as harness tuning layer
2026-06-16 Trustworthy Self-Composable Big-Data-as-a-Service: An LLM-Orchestrated Multi-Agent Framework for Automated Data Engineering, AutoML, MLOps Deployment, and Drift-Aware Lifecycle Optimization LLM-orchestrated meta-harness spanning ingest→MLOps→drift; self-composable pipeline
2026-06-16 Agentic AI-based Framework for Mitigating Premature Diagnostic Handoff and Silent Hallucination in Healthcare Applications Buyya co-author; agentic healthcare harness addressing hallucination and diagnostic handoff
2026-06-16 WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning
2026-06-16 Divide, Deliberate, Decide: A Multi-Agent Framework for Fine-Grained Egocentric Action Recognition
2026-06-16 AUTOGATE: Automated Clock Gating via Toggling-Aware LLM-based RTL Rewriting
2026-06-16 Towards Scalable Customization and Deployment of Multi-Agent Systems for Enterprise Applications
2026-06-16 Guava: An Effective and Universal Harness for Embodied Manipulation
2026-06-16 Dissecting model behavior through agent trajectories
2026-06-16 Agentra: A Supervisable Multi-Agent Framework for Enterprise Intrusion Response
2026-06-16 FinAcumen: Financial Multimodal Reasoning via Self-Evolving Experience Memory Harness Self-evolving experience memory harness; multimodal financial reasoning; novel memory loop
2026-06-15 LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control
2026-06-15 MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
2026-06-15 AgentFairBench: Do LLM Agents Discriminate When They Act?
2026-06-15 ACCORD: Action-Conditioned Contextual Grounding for Language Agents
2026-06-15 LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching
2026-06-15 AdaSTORM: Scaling LLM Reasoning on Dynamic Graphs via Adaptive Spatio-Temporal Multi-Agent Collaboration
2026-06-15 Agentic Discovery of Non-Canonical Antimicrobial Peptides with AMPGAN v3 ICML 2026 ICML 2026 top venue; agentic discovery harness for non-canonical peptide generation
2026-06-14 LLM-as-Code Agentic Programming for Agent Harness KDD 2026
2026-06-13 EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning INTERSPEECH 2026
2026-06-11 The Containment Gap: How Deployed Agentic AI Frameworks Fail Public-Facing Safety Requirements ICML 2026 ICML 2026; systematic audit of agentic framework safety gaps in public-facing deployments
2026-06-11 PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue INTERSPEECH 2026 INTERSPEECH 2026; multi-agent prosody+semantic reasoning pipeline framework

May 2026 (106)

Date Paper Venue Why selected
2026-05-31 SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories SkillAdaptor self-adapts reusable skills from step-level trajectory feedback; training-free
2026-05-31 Leyline: KV Cache Directives for Agentic Inference KV cache directives for agentic inference; breaks chatbot-only append-only cache assumptions
2026-05-31 Self-Healing Agentic Orchestrators for Reliable Tool-Augmented Large Language Model Systems Self-healing orchestrators; failure taxonomy and recovery strategies for tool-augmented LLMs
2026-05-31 Bridging Requirements and Architecture: Multi-Agent Orchestration with External Knowledge and Hierarchical Memory Multi-agent orchestration with external knowledge and hierarchical memory for architecture
2026-05-31 Reducing Token Usage of State-in-Context Agents using Minification Minification reduces state-in-context agent token cost; independent SWE-bench replication
2026-05-31 A New Framework for Cybersecurity Refusals in AI Agents New framework for cybersecurity refusals in multi-step agentic scaffolds; beyond single-turn
2026-05-31 Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing Traceable multi-agent workflows preserving narrative evidence across video editing pipeline
2026-05-31 Can LLM Agents Sustain Long-Horizon Organizational Dynamics? Tests whether LLM agents sustain coherent behavior in structured organizational hierarchies
2026-05-30 MOSAIC: Modular Orchestration for Structured Agentic Intelligence and Composition MOSAIC: modular orchestration substrate for structured agentic composition
2026-05-30 Dynamic Coordination Strategy Selection for Enterprise Multi-Agent Systems Evidence-based guidance for choosing consensus, debate, or single-agent coordination
2026-05-29 From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors Trojan backdoor defense for agentic harnesses; persistent workspace attack model
2026-05-29 Model-Native Computing Architecture: Envisioning Future System Architecture Through the Lens of Computer Architecture Frames LLM agent scheduling as classical computer-architecture problem; novel systems lens
2026-05-29 Learning to Construct Practical Agentic Systems Empirical study exposing gap between research and production agentic system requirements
2026-05-29 Stateful Online Monitoring Catches Distributed Agent Attacks Stateful monitoring catches distributed agent attacks split across many sessions
2026-05-29 How Generation Architecture Shapes Code Complexity in Multi-Agent LLM Systems: A Paired Study on HumanEval Multi-agent code generation increases complexity vs single-shot; controlled paired study
2026-05-28 Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Multi-agent harness for editable scientific figure generation from diverse inputs
2026-05-28 OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Automated auditing of the open skill ecosystem for LLM agents; lifecycle reliability lens
2026-05-28 Scaling Laws for Agent Harnesses via Effective Feedback Compute Scaling laws for agent harnesses via feedback compute; bridges test-time scaling to harness design
2026-05-28 Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation Multi-agent harness for verifiable interleaved multimodal deep research; directly named topic
2026-05-28 Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Framework for Dynamic Scheduling Asynchronous agent framework balancing reactive scheduling and long-horizon planning
2026-05-28 WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction WorldMemArena: memory must track evolving world state and revise stale entries
2026-05-28 Formalizing Mathematics at Scale AutoformBot: multi-agent system formalizes mathematics at textbook scale in Lean 4
2026-05-28 Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents Agora: LLM agents autonomously detect bugs in production-level consensus protocols
2026-05-28 EvoRepair: Enhancing Vulnerability Repair Agents Through Experience-Based Self-Evolution EvoRepair: cross-vulnerability experience accumulation for self-evolving repair agents
2026-05-28 A Multi-AI-agent Framework Enabling End-to-end Finite Element Analysis for Solid Mechanics Problems
2026-05-27 Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows Harness-Bench directly benchmarks harness effects on model outputs; fills critical measurement gap
2026-05-27 Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refinement Prompt Codebooks: discrete compositional optimization for LLM instruction refinement
2026-05-27 AI Research Agents Narrow Scientific Exploration Finds AI research agents narrow scientific exploration diversity; critical warning for harness designers
2026-05-27 SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents SNARE: tests overeager behavior where benign-prompt agents exceed authorized scope
2026-05-27 Multi-Agent LLM-based Metamorphic Testing for REST APIs Multi-agent metamorphic testing for REST APIs; LLMs discover metamorphic relations
2026-05-27 From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence Agentic framework reproduces under-specified ML papers into benchmark-ready implementations
2026-05-27 Adaptive Multimodal Agents-Based Framework for Automatic Workflow Execution Adaptive multimodal agents for automatic workflow execution bridging metadata to perception
2026-05-27 GUI-CIDER: Mid-training GUI Agents via Causal Internalization and Density-aware Exemplar Reselection
2026-05-27 VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild
2026-05-26 Governed Evolution of Agent Runtimes through Executable Operational Cognition Governed harness evolution via executable operational cognition artifacts
2026-05-26 A Policy-Driven Runtime Layer for Agentic LLM Serving Policy-driven runtime layer separating agent framework identity from serving infrastructure
2026-05-26 AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems AgensFlow: coordination-policy substrate decoupling agent skills from runtime routing
2026-05-26 UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems Unified RL interface optimizing all roles in LLM multi-agent systems end-to-end
2026-05-26 Testing Agentic Workflows with Structural Coverage Criteria Structural coverage criteria for agentic workflow testing beyond task-success metrics
2026-05-26 MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation MUSE-Autoskill: self-evolving skill management with creation, memory, and evaluation
2026-05-26 Lessons from Penetration Tests on Large-Scale Agent Systems Penetration test taxonomy reveals recurring vulnerability classes across large agent systems
2026-05-25 From Model Scaling to System Scaling: Scaling the Harness in Agentic AI Theoretical treatment of system scaling vs model scaling; defines harness architecture requirements
2026-05-25 Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems Agent lifespan engineering; first systematic study of deployed agent degradation over time
2026-05-25 Automated Benchmark Auditing for AI Agents and Large Language Models Automated benchmark auditing finds latent bugs in expert-authored agent tasks at scale
2026-05-25 KYA: A Framework-Agnostic Trust Layer for Autonomous Systems with Verifiable Provenance and Hierarchical Policy Composition KYA: framework-agnostic trust layer with verifiable provenance and policy composition
2026-05-25 Security of OpenClaw Agents: Fundamentals, Attacks, and Countermeasures Security fundamentals, attacks, countermeasures for persistent OpenClaw agent frameworks
2026-05-25 CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly CyberEvolver: self-evolving cybersecurity agent adapts scaffold on-the-fly to targets
2026-05-24 Meta-Agent: From Task Descriptions to Verified Multi-Agent Systems Auto-generates verified multi-agent systems from natural-language task descriptions
2026-05-23 DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations Harness evolution via demonstrations; sparse-feedback sample-efficient adaptation paradigm
2026-05-22 When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems Epistemic miscalibration: novel failure mode where agents misjudge plan feasibility
2026-05-22 SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills SkillEvolBench: benchmarks distillation of episodic experience into reusable procedural skills
2026-05-22 EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
2026-05-21 Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost Compiles orchestration workflows into LLM weights; claims 100× cost reduction over frameworks
2026-05-21 The Log is the Agent: Event-Sourced Reactive Graphs for Auditable, Forkable Agentic Systems Event-sourced reactive graphs; log-first design for auditable forkable agentic systems
2026-05-21 MOSS: Self-Evolution through Source-Level Rewriting in Autonomous Agent Systems MOSS: runtime self-evolution via source-level agent code rewriting without weight updates
2026-05-21 Benchmarking Autonomous Agents against Temporal, Spatial, and Semantic Evasions Benchmarks agents against temporal, spatial, semantic evasions in persistent systems
2026-05-21 ExComm: Exploration-Stage Communication for Error-Resilient Agentic Test-Time Scaling ExComm: exploration-stage communication reduces error propagation in long-horizon agents
2026-05-21 One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systems
2026-05-20 Toward AI VIS Co-Scientists: A General and End-to-End Agent Harness for Solving Complex Data Visualization Tasks End-to-end agent harness for scientific visualization; concrete harness architecture case study
2026-05-20 Causal Past Logic for Runtime Verification of Distributed LLM Agent Workflows Causal past logic for correct runtime monitoring of asynchronous distributed agent workflows
2026-05-20 Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems Goal-level energy accounting metric; new efficiency unit for multi-step agentic AI
2026-05-20 RMA: an Agentic System for Research-Level Mathematical Problems RMA: agentic framework targeting research-level mathematics beyond competition benchmarks
2026-05-19 BOHM: Zero-Cost Hierarchical Attribution for Compound AI Systems Zero-cost Shapley attribution for compound AI without expensive coalition re-evaluation
2026-05-19 Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Mix-Quant: quantized prefilling with precise decoding cuts agentic LLM input overhead
2026-05-19 AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows Retrieval-based synthesis of interoperable multi-agent workflows for open-ended science
2026-05-19 MuMuTestUp: Mutation-based Multi-Agent Test Case Update MuMuTestUp: mutation-based multi-agent test update for CI/CD code evolution
2026-05-19 EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design EngiAI: multi-agent benchmark combining simulation, retrieval, and manufacturing preparation
2026-05-19 STAR-PólyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision STAR-PólyaMath: persistent meta-strategic supervision for long-horizon math reasoning
2026-05-18 Code as Agent Harness Establishes code as executable agent harness; foundational paradigm shift for agentic systems
2026-05-18 PROTEA: Offline Evaluation and Iterative Refinement for Multi-Agent LLM Workflows ACL 2026 ACL; offline evaluation and iterative refinement pipeline for multi-agent LLM workflows
2026-05-18 DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows DecisionBench: benchmarks emergent delegation across 11 models in long-horizon workflows
2026-05-18 MINTEval: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems MINTEval: evaluates memory under multi-target interference in long-horizon agent systems
2026-05-18 Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search Lean Refactor: plug-and-play harness for multi-objective controllable proof optimization
2026-05-18 MMoA: An AI-Agent framework with recurrence for Memoried Mixure-of-Agent MMoA: recurrent memoried mixture-of-agent with dynamic context-aware routing
2026-05-18 STRIDE: A Self-Reflective Agent Framework for Reliable Automatic Equation Discovery STRIDE: self-reflective agent avoids generation-only loops in symbolic equation discovery
2026-05-18 It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs
2026-05-17 Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security Comprehensive survey of safety, robustness, privacy, security in agentic AI systems
2026-05-17 MemRepair: Hierarchical Memory for Agentic Repository-Level Vulnerability Repair MemRepair: hierarchical memory enabling repository-scale vulnerability repair agents
2026-05-17 Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
2026-05-16 S-Bus: Automatic Read-Set Reconstruction for Multi-Agent LLM State Coordination S-Bus: automatic read-set reconstruction preventing multi-agent structural race conditions
2026-05-16 Responsible Agentic AI Requires Explicit Provenance Explicit provenance as foundation for responsible autonomous agentic AI deployment
2026-05-16 AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents AgentKernelArena: generalization-aware benchmark for GPU kernel optimization agents
2026-05-16 Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework Systematic analysis of three agent interaction paradigms within unified production framework
2026-05-15 TopoEvo: A Topology-Aware Self-Evolving Multi-Agent Framework for Root Cause Analysis in Microservices TopoEvo: topology-aware self-evolving multi-agent framework for root cause analysis
2026-05-15 BootstrapAgent: Distilling Repository Setup into Reusable Agent Knowledge BootstrapAgent: distills costly repository setup knowledge into reusable agent memory
2026-05-15 STAR: A Stage-attributed Triage and Repair framework for RCA Agents in Microservices STAR: stage-attributed triage and repair improves reliability of RCA agents
2026-05-15 RTL-BenchMT: Dynamic Maintenance of RTL Generation Benchmark Through Agent-Assisted Analysis and Revision RTL-BenchMT: agentic dynamic maintenance prevents benchmark staleness in EDA research
2026-05-15 An Agentic Retrieval Framework for Autonomous Context-Aware Data Quality Assessment Agentic retrieval framework for autonomous context-aware data quality assessment
2026-05-15 CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? CHI-Bench: end-to-end automation of policy-dense healthcare workflows via agents
2026-05-14 Is Grep All You Need? How Agent Harnesses Reshape Agentic Search Isolates harness contributions to agentic search; grep-sufficiency ablation study
2026-05-14 LEMON: Learning Executable Multi-Agent Orchestration via Counterfactual Reinforcement Learning NEURIPS 2026 NEURIPS; RL learns multi-agent orchestration topology outperforming hand-designed workflows
2026-05-14 APWA: A Distributed Architecture for Parallelizable Agentic Workflows Distributed architecture enabling parallelizable agentic workflow execution at scale
2026-05-14 Orchard: An Open-Source Agentic Modeling Framework Orchard: open-source agentic modeling framework addressing open-research bottleneck
2026-05-14 Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution Solvita: stateful self-evolving agent accumulates experience for competitive programming
2026-05-14 Veritas: A Semantically Grounded Agentic Framework for Memory Corruption Vulnerability Detection in Binaries Veritas: semantically grounded agentic framework for binary memory-corruption detection
2026-05-14 Auditing Agent Harness Safety First audit of harness safety; unauthorized access via correct-looking trajectories
2026-05-13 Position: Agentic AI System Is a Foreseeable Pathway to AGI ICML 26 ICML position; argues agentic AI systems are necessary pathway to AGI beyond scaling
2026-05-13 Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context
2026-05-11 WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation Real-world long-horizon CLI harness benchmark against synthetic sandboxes and mock servers
2026-05-11 NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation Co-evolving skills, memory, and policy for personalized multi-agent research automation
2026-05-11 Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents
2026-05-09 Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection
2026-05-08 Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge
2026-05-07 Retrieval-Conditioned Topology Selection with Provable Budget Conservation for Multi-Agent Code Generation NEURIPS 2026 NEURIPS; retrieval-conditioned topology selection for code-generation multi-agent routing
2026-05-05 SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agents Portable skill compilation across LLM agent frameworks; cross-framework interoperability
2026-05-05 Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies Workspace-Bench: AI agents on real heterogeneous file-dependency workspace tasks

April 2026 (84)

Date Paper Venue Why selected
2026-04-30 Crab: A Semantics-Aware Checkpoint/Restore Runtime for Agent Sandboxes Checkpoint/restore runtime for agent sandboxes; enables rollout branching and fault tolerance
2026-04-30 In-Context Prompting Obsoletes Agent Orchestration for Procedural Tasks Controlled study: in-context prompting often beats LangGraph/CrewAI orchestration
2026-04-30 Trace-Level Analysis of Information Contamination in Multi-Agent Systems Information contamination propagates through multi-agent workflows; trace-level analysis
2026-04-30 Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study Layered security attack/defense review for autonomous agent framework vulnerabilities
2026-04-30 Building Persona-Based Agents On Demand: Tailoring Multi-Agent Workflows to User Needs Persona-based agent workflow tailoring on demand; adaptive multi-agent customization
2026-04-30 Heterogeneous Scientific Foundation Model Collaboration Heterogeneous scientific foundation model collaboration beyond text-only agent interfaces
2026-04-30 CTM-AI: A Blueprint for General AI Inspired by a Model of Consciousness Consciousness-inspired blueprint for general-purpose multi-modal AI agent architecture
2026-04-30 WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments WindowsWorld: process-centric benchmark for GUI agents across professional cross-app tasks
2026-04-29 Agent Name Service (ANS): A Proof-of-Concept Trust Layer for Secure AI Agent Discovery, Identity, and Governance in Kubernetes Agent Name Service: cryptographic identity, discovery, governance for agent ecosystems
2026-04-29 Bian Que: An Agentic Framework with Flexible Skill Arrangement for Online System Operations Flexible skill arrangement in agentic framework for large-scale system operations
2026-04-29 Web2BigTable: A Bi-Level Multi-Agent LLM System for Internet-Scale Information Search and Extraction Bi-level multi-agent system for internet-scale structured information extraction
2026-04-29 DreamProver: Evolving Transferable Lemma Libraries via a Wake-Sleep Theorem-Proving Agent Wake-sleep program induction discovers reusable lemma libraries for theorem proving
2026-04-29 GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents GLM-5V-Turbo: native foundation model explicitly designed for multimodal agent deployment
2026-04-28 Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses Observability-driven automatic evolution of coding-agent harnesses; core harness engineering
2026-04-28 Recursive Multi-Agent Systems Extends recursive LM scaling to multi-agent systems; new meta-harness paradigm
2026-04-28 Scalable Inference Architectures for Compound AI Systems: A Production Deployment Study Production deployment study of compound AI systems; practical inference architectures
2026-04-28 Toward Scalable Terminal Task Synthesis via Skill Graphs Skill graphs enable scalable terminal task synthesis for training agent harnesses
2026-04-28 FAMA: Failure-Aware Meta-Agentic Framework for Open-Source LLMs in Interactive Tool Use Environments ACL 2026 Failure-aware meta-agentic framework for open-source LLMs in tool-use [ACL 2026]
2026-04-28 SAFEdit: Does Multi-Agent Decomposition Resolve the Reliability Challenges of Instructed Code Editing? Multi-agent decomposition evaluated for code editing reliability improvement
2026-04-28 From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling Memory-enhanced LLM agents with decentralized debate for optimization modeling
2026-04-28 Towards Agentic Investigation of Security Alerts Agentic multi-source log correlation automates early-stage security alert investigation
2026-04-27 SeaEvo: Advancing Algorithm Discovery with Strategy Space Evolution Strategy space evolution improves LLM-guided algorithm discovery beyond scalar fitness
2026-04-27 Constraint-Guided Multi-Agent Decompilation for Executable Binary Recovery Constraint-guided multi-agent decompilation improves binary recovery compilability
2026-04-27 Co-Director: Agentic Generative Video Storytelling
2026-04-26 AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking ACL 2026 DAG-structured step-level agentic evaluation with error propagation tracking [ACL 2026]
2026-04-26 Optimas: An Intelligent Analytics-Informed Generative AI Framework for Performance Optimization Performance-analytics-informed LLM framework for automated code optimization
2026-04-26 KISS Sorcar: A Stupidly-Simple General-Purpose and Software Engineering AI Assistant Minimal SE assistant addressing finite context windows and multi-turn session gaps
2026-04-25 RAT: RunAnyThing via Fully Automated Environment Configuration Automated repo environment configuration; foundational infrastructure for coding agent harnesses
2026-04-24 Beyond Single-Agent Alignment: Preventing Context-Fragmented Violations in Multi-Agent Systems Context-fragmented violations: individually safe agent actions collectively break policies
2026-04-24 Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning Behavioral canaries audit private retrieved context usage during RL fine-tuning
2026-04-24 A Decoupled Human-in-the-Loop System for Controlled Autonomy in Agentic Workflows Decoupled human-in-the-loop for controlled autonomy in agentic workflows
2026-04-23 Tool Attention Is All You Need: Dynamic Tool Gating and Lazy Schema Loading for Eliminating the MCP/Tools Tax in Scalable Agentic Workflows Dynamic tool gating and lazy schema loading eliminates MCP harness tax overhead
2026-04-23 Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework Taxonomy and evaluation framework for emergent strategic reasoning risks in agents
2026-04-23 Read the Paper, Write the Code: Agentic Reproduction of Social-Science Results LLM agents reproduce social science results from methods text alone; generalization test
2026-04-22 HARBOR: Automated Harness Optimization First paper automating harness optimization end-to-end; directly on topic
2026-04-22 The Last Harness You'll Ever Build Principles for building domain-specific long-horizon agentic harnesses
2026-04-22 Learning to Evolve: A Self-Improving Framework for Multi-Agent Systems via Textual Parameter Graph Optimization Textual parameter graph optimization enables self-improving multi-agent system design
2026-04-22 Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows Agents exploit public eval scores; benchmark-gaming risk in coding harnesses
2026-04-22 EvoAgent: An Evolvable Agent Framework with Skill Learning and Multi-Agent Delegation Structured skill learning with hierarchical sub-agent delegation in EvoAgent framework
2026-04-22 CHORUS: An Agentic Framework for Generating Realistic Deliberation Data Agentic framework generates large-scale realistic deliberation data at scale
2026-04-21 Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs Analyzes latency-reliability-cost tradeoffs in LLM-enabled agentic workflow design
2026-04-21 A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression Observational context compression for efficient long-horizon terminal agent harnesses
2026-04-21 Security Is Relative: Training-Free Vulnerability Detection via Multi-Agent Behavioral Contract Synthesis Multi-agent behavioral contract synthesis for vulnerability detection; deduplication-robust
2026-04-21 Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment ACL 2026 Dialectical alignment resolves actor-observer asymmetry in multi-agent frameworks [ACL 2026]
2026-04-21 AgenticRecTune: Multi-Agent with Self-Evolving Skillhub for Recommendation System Optimization Multi-agent with self-evolving skillhub optimizes multi-stage recommendation pipelines
2026-04-21 ChipCraftBrain: Validation-First RTL Generation via Multi-Agent Orchestration Validation-first RTL generation via multi-agent orchestration; correctness improvement
2026-04-21 ClawNet: Human-Symbiotic Agent Network for Cross-User Autonomous Cooperation
2026-04-20 Architectural Design Decisions in AI Agent Harnesses Systematic survey of architectural design decisions specific to AI agent harnesses
2026-04-20 SDOF: Taming the Alignment Tax in Multi-Agent Orchestration with State-Constrained Dispatch State-constrained dispatch fixes alignment tax in multi-agent orchestration frameworks
2026-04-20 Co-evolving Agent Architectures and Interpretable Reasoning for Automated Optimization Co-evolves agent architectures and interpretable reasoning for automated optimization
2026-04-20 AI scientists produce results without reasoning scientifically LLM-based scientific agents achieve results but violate epistemic reasoning norms
2026-04-20 Human-Guided Harm Recovery for Computer Use Agents Formalizes human-guided harm recovery for computer-use agents after prevention fails
2026-04-20 ClawEnvKit: Automatic Environment Generation for Claw-Like Agents Automatic environment generation pipeline for training and evaluating claw-like agents
2026-04-20 HiRAS: A Hierarchical Multi-Agent Framework for Paper-to-Code Generation and Execution Hierarchical multi-agent with adaptive pipelines for paper-to-code generation
2026-04-20 Poster: EdgeCitadel -- Hybrid NATS-MQTT Orchestration for Edge Multi-Agent Systems EdgeCitadel: hybrid NATS-MQTT orchestration for edge-resident multi-agent systems
2026-04-19 Compiling Deterministic Structure into SLM Harnesses Compiles deterministic structure into SLM harnesses via semantic gradient descent
2026-04-19 Clover: A Neural-Symbolic Agentic Harness with Stochastic Tree-of-Thoughts for Verified RTL Repair Neural-symbolic agentic harness with stochastic tree-of-thoughts for verified repair
2026-04-19 EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale Foundational evolving agent framework enabling iterative agentic science at scale
2026-04-19 Agents Explore but Agents Ignore: LLMs Lack Environmental Curiosity LLM agents fail to exploit unexpected discoveries; environmental curiosity gap exposed
2026-04-19 Safe and Policy-Compliant Multi-Agent Orchestration for Enterprise AI Policy-compliant multi-agent orchestration with SOX/HIPAA/GDPR enterprise constraints
2026-04-19 SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents Benchmark for lifelong skill discovery and evolution across autonomous agents
2026-04-19 Project Prometheus: Bridging the Intent Gap in Agentic Program Repair via Reverse-Engineered Executable Specifications Reverse-engineered executable specs bridge intent gap in agentic program repair
2026-04-18 Harness as an Asset: Enforcing Determinism via the Convergent AI Agent Framework (CAAF) Determinism enforcement in harnesses via CAAF; safety-critical deployment
2026-04-18 Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks Most comprehensive cross-model eval of LLM agents on offensive cyber tasks
2026-04-17 Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Multi-Agent Workflows Cyclic subtask graphs model flexibility vs coordination overhead in multi-agent workflows
2026-04-17 Weak-Link Optimization for Multi-Agent Reasoning and Collaboration Weak-link optimization prevents error amplification in multi-agent collaborative reasoning
2026-04-17 Agentic Frameworks for Reasoning Tasks: An Empirical Study Empirical comparison of agentic frameworks across reasoning tasks and efficiency
2026-04-17 Know When to Trust the Skill: Delayed Appraisal and Epistemic Vigilance for Single-Agent LLMs Delayed appraisal and epistemic vigilance for reliable tool use in LLM agents
2026-04-17 AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficient Kernel Generation Failure-driven adaptation with diversity preservation for LLM kernel generation agents
2026-04-17 The Amazing Agent Race: Strong Tool Users, Weak Navigators
2026-04-16 The Semi-Executable Stack: Agentic Software Engineering and the Expanding Scope of SE Agentic harnesses expand scope of SE; conceptual framing for the field
2026-04-16 Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines Serving agentic workflows via aggregate LLM pipelines; latency-throughput infrastructure
2026-04-16 Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex Trajectory-level safety evaluation and diagnosis benchmark for agent harnesses
2026-04-16 ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants Agentic GPU kernel optimization via data-flow invariants; closes performance gap
2026-04-16 VeriGraphi: A Multi-Agent Framework of Hierarchical RTL Generation for Large Hardware Designs Multi-agent hierarchical RTL generation for large hardware designs
2026-04-16 Dr. RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement Tool-grounded self-improvement cycle for autonomous agentic RTL optimization
2026-04-16 CAMO: An Agentic Framework for Automated Causal Discovery from Micro Behaviors to Macro Emergence in LLM Agent Simulations Agentic causal discovery from micro-behaviors in LLM agent simulations
2026-04-14 AgentSPEX: An Agent SPecification and EXecution Language Agent specification and execution language; foundational DSL for harness programming
2026-04-14 Numerical Instability and Chaos: Quantifying the Unpredictability of Large Language Models Numerical instability in LLMs undermines agentic harness reliability and determinism
2026-04-13 Mem\(^2\)Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation ACL 2026 Co-evolutionary capability expansion and experience distillation for self-evolving agents [ACL]
2026-04-09 SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents ACL 2026 Joint optimization of policy and tool graph memory for self-evolving agents [ACL 2026]
2026-04-07 Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework ACL Open-source multi-agent research discovery and analysis pipeline framework [ACL]
2026-04-04 Explainable Model Routing for Agentic Workflows Explainable model routing for agentic workflows; cost-quality tradeoff analysis
2026-04-03 OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments

March 2026 (67)

Date Paper Venue Why selected
2026-03-31 VeriAct: Beyond Verifiability -- Agentic Synthesis of Correct and Complete Formal Specifications
2026-03-31 KAIJU: An Executive Kernel for Intent-Gated Execution of LLM Agents
2026-03-31 An Artifact-based Agent Framework for Adaptive and Reproducible Medical Image Processing
2026-03-31 A Study on the Impact of Fault localization Granularity for Repository-Scale Code Repair Tasks
2026-03-31 One Panel Does Not Fit All: Case-Adaptive Multi-Agent Deliberation for Clinical Prediction
2026-03-31 SkillReducer: Optimizing LLM Agent Skills for Token Efficiency
2026-03-31 Near-Miss: Latent Policy Failure Detection in Agentic Workflows
2026-03-31 ASI-Evolve: AI Accelerates AI
2026-03-31 KPI2KVI: A Multi Agent Workflow for Calculating Key Value Indicators from Service Descriptions
2026-03-31 How and Why Agents Can Identify Bug-Introducing Commits
2026-03-31 AEC-Bench: A Multimodal Benchmark for Agentic Systems in Architecture, Engineering, and Construction
2026-03-31 SimMOF: AI agent for Automated MOF Simulations
2026-03-30 Mimosa Framework: Toward Evolving Multi-Agent Systems for Scientific Research
2026-03-30 Multi-Agent LLMs for Adaptive Acquisition in Bayesian Optimization
2026-03-30 BACE: LLM-based Code Generation through Bayesian Anchored Co-Evolution of Code and Test Populations
2026-03-30 Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering
2026-03-30 Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification
2026-03-30 Entropic Claim Resolution: Uncertainty-Driven Evidence Selection for RAG
2026-03-29 Towards Context-Aware Image Anonymization with Multi-Agent Reasoning CVPR 2026
2026-03-29 EffiSkill: Agent Skill Based Automated Code Efficiency Optimization
2026-03-29 A Security Analysis of the OpenClaw AI Agent Framework
2026-03-28 AutoMS: Multi-Agent Evolutionary Search for Cross-Physics Inverse Microstructure Design
2026-03-28 MediHive: A Decentralized Agent Collective for Medical Reasoning
2026-03-28 Autonomous Agent-Orchestrated Digital Twins (AADT): Leveraging the OpenClaw Framework for State Synchronization in Rare Genetic Disorders
2026-03-28 Story2Proposal: A Scaffold for Structured Scientific Paper Writing
2026-03-27 FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified? ICLR 2026
2026-03-27 ClinicalAgents: Multi-Agent Orchestration for Clinical Decision Making with Dual-Memory KDD 2026
2026-03-27 Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
2026-03-27 Agents on a Tree: Pathwise Coordination for Multi-Objective Molecular Optimization
2026-03-27 Knowdit: Agentic Smart Contract Vulnerability Detection with Auditing Knowledge Summarization
2026-03-27 Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents
2026-03-27 GISclaw: A Comprehensive Open-Source LLM Agent System for Realistic Multi-Step Geospatial Analysis
2026-03-27 AutoB2G: A Large Language Model-Driven Agentic Framework For Automated Building-Grid Co-Simulation
2026-03-26 AVDA: Autonomous Vibe Detection Authoring for Cybersecurity
2026-03-26 Natural-Language Agent Harnesses
2026-03-26 ReCUBE: Evaluating Repository-Level Context Utilization in Code Generation
2026-03-26 SEVerA: Verified Synthesis of Self-Evolving Agents
2026-03-26 From Logic Monopoly to Social Contract: Separation of Power and the Institutional Foundations for Autonomous Agent Economies
2026-03-26 TopoPilot: Reliable Conversational Workflow Automation for Topological Data Analysis and Visualization
2026-03-25 LensWalk: Agentic Video Understanding by Planning How You See in Videos CVPR 2026
2026-03-25 SentinelAI: A Multi-Agent Framework for Structuring and Linking NG9-1-1 Emergency Incident Data
2026-03-25 AutoSAM: an Agentic Framework for Automating Input File Generation for the SAM Code with Multi-Modal Retrieval-Augmented Generation
2026-03-25 Multi-Agent Reasoning with Consistency Verification Improves Uncertainty Calibration in Medical MCQA
2026-03-25 CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
2026-03-25 AI-Supervisor: Autonomous AI Research Supervision via a Persistent Research World Model
2026-03-25 Environment-Grounded Multi-Agent Workflow for Autonomous Penetration Testing
2026-03-25 ELITE: Experiential Learning and Intent-Aware Transfer for Self-improving Embodied Agents
2026-03-25 Language-Grounded Multi-Agent Planning for Personalized and Fair Participatory Urban Sensing
2026-03-25 AnalogAgent: Self-Improving Analog Circuit Design Automation with LLM Agents
2026-03-24 Efficient Benchmarking of AI Agents
2026-03-24 Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment
2026-03-24 Can LLM Agents Generate Real-World Evidence? Evaluating Observational Studies in Medical Databases
2026-03-24 Benchmarking Multi-Agent LLM Architectures for Financial Document Processing: A Comparative Study of Orchestration Patterns, Cost-Accuracy Tradeoffs and Production Scaling Strategies
2026-03-20 PersonaVLM: Long-Term Personalized Multimodal LLMs CVPR 2026
2026-03-17 DanceHA: A Multi-Agent Framework for Document-Level Aspect-Based Sentiment Analysis AAAI 2026
2026-03-17 RECOVER: Robust Entity Correction via agentic Orchestration of hypothesis Variants for Evidence-based Recovery INTERSPEECH 2026
2026-03-16 The PokeAgent Challenge: Competitive and Long-Context Learning at Scale NEURIPS 2025
2026-03-16 AgentTrace: Causal Graph Tracing for Root Cause Analysis in Deployed Multi-Agent Systems ICLR 2026
2026-03-14 LiveWeb-IE: A Benchmark For Online Web Information Extraction ICLR 2026
2026-03-13 EvoClaw: Evaluating AI Agents on Continuous Software Evolution ICML 2026
2026-03-12 Verified Multi-Agent Orchestration: A Plan-Execute-Verify-Replan Framework for Complex Query Resolution ICLR 2026
2026-03-11 SciNav: A General Agent Framework for Scientific Coding Tasks ICLR 2026
2026-03-10 A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation MICCAI 2026
2026-03-09 UIS-Digger: Towards Comprehensive Research Agent Systems for Real-world Unindexed Information Seeking ICLR 2026
2026-03-06 LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models AAAI 2026
2026-03-02 LexChronos: An Agentic Framework for Structured Event Timeline Extraction in Indian Jurisprudence AAAI 2026
2026-03-02 CARE: Towards Clinical Accountability in Multi-Modal Medical Reasoning with an Evidence-Grounded Agentic Framework ICLR 2026

January 2026 (48)

Date Paper Venue Why selected
2026-01-01 Textual Equilibrium Propagation for Deep Compound AI Systems ICLR 2026 ICLR; textual equilibrium propagation for compound AI; gradient-like credit through tool chains
2026-01-01 Risk-Sensitive Agent Compositions ICLR 2026 ICLR; formalizes agentic workflows as DAGs; risk-sensitive composition with safety guarantees
2026-01-01 ResiliBench: Evaluating Agentic Workflow Adaptation in Stochastic Environments ICLR 2026 ICLR ResiliBench; evaluates workflow adaptation under instruction variability and tool uncertainty
2026-01-01 Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents ICLR 2026 ICLR OSU; Agent Data Protocol unifying diverse datasets for fine-tuning LLM agents at scale
2026-01-01 Stop Wasting Your Tokens: Towards Efficient Runtime Multi-Agent Systems ICLR 2026 ICLR; targets token waste and misinformation in runtime multi-agent systems
2026-01-01 Multi-View Encoders for Performance Prediction in LLM-Based Agentic Workflows ICLR 2026 ICLR KAIST; multi-view performance prediction for agentic workflow configuration optimization
2026-01-01 Kimi-Dev: Agentless Training as Skill Prior for SWE-agents ICLR 2026 ICLR Kimi/Moonshot; agentless training as skill prior for SWE-agent harness design
2026-01-01 Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation ICLR 2026 ICLR; multi-agent deliberation loop overcomes lazy-agent failure in RL-trained reasoning
2026-01-01 FlowSearcher: Synthesizing Memory-Guided Agentic Workflows for Web Information Seeking ICLR 2026 ICLR; memory-guided agentic workflow synthesis for non-linear web information seeking
2026-01-01 WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research ICLR 2026 ICLR; dynamic outline structures mitigate static-pipeline failures in deep research agents
2026-01-01 Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents ICLR 2026 ICLR; process-level trajectory evaluation for SE agent environment configuration tasks
2026-01-01 WideSearch: Benchmarking Agentic Broad Info-Seeking ICLR 2026 ICLR; benchmarks wide-scale info-seeking workflows; exposes automation bottleneck
2026-01-01 MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use ICLR 2026 ICLR; stress-tests realistic MCP use with full read/write/deep interaction coverage
2026-01-01 PLAGUE: Plug-and-play Framework for Lifelong Adaptive Generation of Multi-turn Jailbreaks ICLR 2026 ICLR; PLAGUE: plug-and-play lifelong adaptive multi-turn jailbreak generation framework
2026-01-01 Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory ICLR 2026 ICLR; M3-Agent: multimodal agent with episodic and semantic long-term memory
2026-01-01 Demystifying Deep Search: A Holistic Evaluation with Hint-free Multi-Hop Questions and Factorised Metrics ICLR 2026 ICLR; hint-free multi-hop queries with factorised metrics expose deep search agent gaps
2026-01-01 Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing ICLR 2026 First AI-agent vs. human pentester comparison in live enterprise environment [ICLR 2026]
2026-01-01 Towards Multimodal Data-Driven Scientific Discovery Powered by LLM Agents ICLR 2026 LLM agents for multimodal data-driven scientific discovery; benchmark and pipeline [ICLR 2026]
2026-01-01 SciNav: A General Agent Framework for Scientific Coding Tasks ICLR 2026 General agent framework for scientific coding; reasoning-to-execution pipeline [ICLR 2026]
2026-01-01 AI-for-Science Low-code Platform with Bayesian Adversarial Multi-Agent Framework ICLR 2026 Bayesian adversarial multi-agent for AI-for-science low-code automation [ICLR 2026]
2026-01-01 Helmsman: Autonomous Synthesis of Federated Learning Systems via Collaborative LLM Agents ICLR 2026 Autonomous synthesis of federated learning systems via collaborative agents [ICLR 2026]
2026-01-01 M\(^2\)-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining ICLR 2026 Multi-agent MCTS for GUI agent trajectory data mining at scale [ICLR 2026]
2026-01-01 MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning ICLR 2026 RL-optimized multi-agent collaboration for multimodal medical reasoning [ICLR 2026]
2026-01-01 CellAgent: LLM-Driven Multi-Agent Framework for Natural Language-Based Single-Cell Analysis ICLR 2026 LLM-driven multi-agent framework for natural language single-cell analysis [ICLR 2026]
2026-01-01 AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent ICLR 2026 Tool-augmented math agent integrating LRMs with symbolic execution harness [ICLR 2026]
2026-01-01 Social Agents: Collective Intelligence Improves LLM Predictions ICLR 2026 Social multi-agent collective intelligence improves LLM prediction accuracy [ICLR 2026]
2026-01-01 CoDA: Agentic Systems for Collaborative Data Visualization ICLR 2026 Agentic systems for collaborative data visualization via deep research paradigm [ICLR 2026]
2026-01-01 MedAgent-Pro: Towards Evidence-based Multi-modal Medical Diagnosis via Reasoning Agentic Workflow ICLR 2026 Evidence-based multi-modal medical diagnosis via structured agentic workflow [ICLR 2026]
2026-01-01 Learning to Summarize by Learning to Quiz: Adversarial Agentic Collaboration for Long Document Summarization ICLR 2026 Adversarial agentic collaboration for long-document summarization [ICLR 2026]
2026-01-01 ATLAS: Constraints-Aware Multi-Agent Collaboration for Real-World Travel Planning ICLR 2026 Multi-agent travel planning under complex constraints; grounded reasoning [ICLR 2026]
2026-01-01 WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality ICLR 2026
2026-01-01 SMAN-Bench: A Cross-System Benchmark for Mobile Agents under Single- and Multi-path, Ambiguous, and Noisy Tasks ICLR 2026
2026-01-01 UIS-Digger: Towards Comprehensive Research Agent Systems for Real-world Unindexed Information Seeking ICLR 2026
2026-01-01 PerfGuard: A Performance-Aware Agent for Visual Content Generation ICLR 2026
2026-01-01 From Medical Records to Diagnostic Dialogues: A Clinical-Grounded Approach and Dataset for Psychiatric Comorbidity ICLR 2026
2026-01-01 SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents ICLR 2026
2026-01-01 MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents ICLR 2026
2026-01-01 IR-Agent: Expert-Inspired LLM Agents for Structure Elucidation from Infrared Spectra ICLR 2026
2026-01-01 Pursuing Minimal Sufficiency in Spatial Reasoning ICLR 2026
2026-01-01 From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization ICLR 2026
2026-01-01 P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark ICLR 2026
2026-01-01 CoLLMLight: Cooperative Large Language Model Agents for Network-Wide Traffic Signal Control ICLR 2026
2026-01-01 HAMLET: A Hierarchical and Adaptive Multi-Agent Framework for Live Embodied Theatrics ICLR 2026
2026-01-01 Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations ICLR 2026
2026-01-01 ContextNav: Towards Agentic Multimodal In-Context Learning ICLR 2026
2026-01-01 VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning ICLR 2026
2026-01-01 Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents ICLR 2026
2026-01-01 InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search ICLR 2026