Skip to content

Harnesses / Meta-Harnesses โ€” 2025

๐Ÿ—“๏ธ 2025 in Review

๐Ÿ•’ Timeline

  • 2025-07: Emergence of 30+ agent scientific systems (cmbagent) and hierarchical theatrical orchestration (HAMLET, HIMA); FSM-based auto-construction demonstrated (MetaAgent)
  • 2025-08: Survey consolidation of OS-agent field; cost-efficiency realism sets in โ€” observation masking beats LLM summarization; VistaWise shows low-parameter embodied agents can match API-scale competitors
  • 2025-09: Domain-specialization wave: medical (RadAgents, MedLA, AMANDA), financial (FinDebate, CodeGym), scientific (Re4, MOSAIC) harnesses proliferate; red-teaming meta-frameworks emerge (SafeSearch)
  • 2025-10: Meta-harness evaluation infrastructure matures โ€” self-evolving safety benchmarks (AgenticEval), no-human annotation at Walmart scale (ScalingEval), plug-and-play scientific knowledge (xKG, PLAGUE)
  • 2025-11: Automated workflow generation breaks through โ€” AยฒFlow eliminates hand-crafted operators via MCTS over auto-derived abstractions; Agint compiles natural-language specs to typed DAGs; enterprise production deployment (MAFA at JPMorgan, 1M-utterance backlog cleared)
  • 2025-12: RL-augmented harnesses set new performance ceilings (AgentMath-30B: 90.6% AIME24); empirical workflow comparison shows agentic MCP flows beat rigid expert pipelines at small model scales (Workflows vs Agents); context management hardened to O(1) (MemSearcher)

๐Ÿ“ˆ Trend

Where the field stands

The 2025 literature reveals a decisive shift from hand-crafted single-agent pipelines toward composable, self-constructing multi-agent harnesses โ€” meta-level frameworks that orchestrate agent populations rather than individual models. The dominant pattern is decomposition: tasks are split across specialized sub-agents (planners, coders, critics, verifiers) coordinated by a controller that manages routing, context, and feedback. Three concurrent pressures are driving this: the context-window ceiling on single-agent depth, the reliability gap in monolithic LLM calls, and the explosion of domain-specific deployment demands (medical, financial, scientific). By late 2025, the frontier has moved again toward automated harness construction โ€” systems that design, compile, and self-optimize their own agent topologies from task descriptions alone โ€” and toward RL-trained meta-controllers that replace human-designed orchestration logic.

โžก๏ธ Read the full trend analysis

๐Ÿ“„ Papers (400)

โžก๏ธ Paper list โ€” 400 papers, grouped by month. ยท โญ 15 key papers (see Key papers in the left nav).