Skip to content

Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation

🕒 Published (v1): 2026-07-06 17:55 UTC · Source: Arxiv · link

Why this paper was selected

Bidirectional hierarchical VLA agent addressing Markovian failure in long-horizon manipulation

Ask a follow-up

Open an assistant pre-loaded with this paper's context.

💬 Ask ChatGPT✦ Ask Claude

TL;DR

Cortex is a hierarchical dual-system embodied agent that bridges the gap between a high-level VLM planner and a low-level VLA executor via a bidirectionally aligned subtask interface. It standardizes manipulation into 32 canonical skill primitives and employs event-balanced sampling to handle planning ambiguity at subtask transitions. On standard benchmarks, Cortex achieves state-of-the-art long-horizon manipulation performance and enables zero-shot transfer to unseen real-world tasks like multi-stage chemistry experiments.

Problem

Monolithic VLA models are Markovian—they lack persistent memory and cannot track long-horizon task progress, causing repeated actions and compounding errors. Hierarchical dual-system approaches partially address this but suffer from a semantic-kinematic domain gap: high-level VLM planners generate unconstrained natural language plans that are kinematically ungrounded for the low-level executor. Conversely, existing VLAs are not designed to consume structured planner outputs, breaking bidirectional alignment.

Method

Cortex decomposes long-horizon manipulation into a two-system architecture:

  • Interface formulation: Subtasks \(s_t\) are instantiated as elements from 32 canonical skill primitives with strict language templates (e.g., [Pick] [object], [Unscrew] [object]). A recurrently updated semantic memory \(M^{(t)} = M^{(0)} \oplus \bigoplus_{i=1}^{L_{k-1}} \Phi(s_i)\) encodes completed milestones and is fed back into the VLM at each step.
  • Metadata construction: 4,000+ hours of open-source video are automatically re-annotated using Qwen3-VL-235B, with subtask boundaries inferred via dynamic programming over fused state-action and visual features \(x_t = [\phi_s(o_t, a_t), \phi_v(I_t)]\). Simulation data (~30 hours) is synthesized with explicit tractability constraints: object color/spatial attributes and kinematic reachability checks for embodiment-aware routing.
  • Event-balanced sampling: Trajectories are split into boundary transition phases (\(t \in [t_k - \epsilon_1, t_k + \epsilon_2]\), ~1 s total, with \(\epsilon_2 > \epsilon_1\)) and intra-task execution phases. Transition frames are densely oversampled to counterbalance the majority of steady-state execution frames, forcing the model to learn both "semantic patience" and discrete state advancement.
  • Harness engineering: At inference, a lightweight arbitration layer maps open-ended instructions into the unified skill interface, enforces skill constraints on VLM outputs, and uses sequence matching to canonicalize raw descriptions to the nearest primitive.
  • Asynchronous inference: System-2 (VLM) runs at ~2 Hz monitoring progress; System-1 (VLA, \(\pi_0^{0.5}\)) executes at ~10 Hz. A timeout-driven kinematic reset handles perceptual deadlocks.

Key Contributions

  • A bidirectionally aligned subtask interface with 32 canonical skill primitives that enforces both executability (standardized syntax the VLA can consume) and tractability (object attributes, reachability, spatial disambiguation).
  • A scalable automatic annotation pipeline covering 4,000+ hours of heterogeneous video data and 30 hours of procedurally generated simulation data.
  • An event-balanced sampling strategy with asymmetric temporal margins around subtask boundaries, explicitly balancing intra-task state maintenance versus discrete transition learning.
  • Harness engineering for inference-time grounding and closed-loop physical error recovery via timeout-driven kinematic resets.
  • Zero-shot generalization to unseen real-world long-horizon tasks by combining the generalist VLM with a task-specific fine-tuned VLA.

Results

  • LIBERO-Long: 95.5% success rate, +3.1% over best end-to-end baseline (Ď€0.5 at 92.4%) and +4.5% over best agentic baseline (Gemini-3.1-Pro at 91.0%).
  • RoboTwin 2.0: 86.8% overall success rate (+4.1% over Ď€0.5), with 88.0% on long-horizon splits where all monolithic VLAs degrade significantly.
  • Open-loop VLM evaluation (LLM-as-Judge): Cortex (full harness) scores 8.741 avg. total at step-level and 8.464 at episode-level, outperforming GPT-5 (7.376), Gemini (7.006), and Qwen3-VL-8B-Instruct (6.388) at episode-level.
  • Real-world chemistry experiments (zero-shot, 20 trials): 65% success rate and 11.0/14 avg. progress; end-to-end Ď€0 and Ď€0.5 achieve 0% success (2.5/14 and 4.1/14 progress).
  • Human + fine-tuned VLA upper bound on real-world tasks: 75% SR (chemistry), 70% SR (washing)—Cortex achieves 65%/55% respectively.

Limitations

  • The 32 canonical skill primitives form a closed vocabulary; tasks requiring novel manipulation primitives outside this set cannot be directly executed without extending the taxonomy.
  • The low-level VLA (\(\pi_{\text{mem}}^{\text{sub}}\)) still requires ~10 hours of fine-tuning on segmented subtask-to-action data per embodiment, limiting plug-and-play deployment.
  • Real-world evaluation covers only two task types (chemistry experiments, kitchen assembly) on a single robot platform (ARX ACONE); broader generalization is unverified.
  • The timeout-driven kinematic reset for error recovery is a heuristic mechanism; its behavior in more dynamic or cluttered environments is uncharacterized.
  • Subtask boundary inference via dynamic programming assumes relatively clean state-action signals; noisy or highly variable demonstrations may degrade annotation quality.

Relevance to Vision-Language Models

Cortex directly addresses how VLMs can be specialized as grounded cognitive orchestrators within robotic systems—a question central to embodied VLM research. It demonstrates that a VLM fine-tuned with structured skill-primitive supervision and event-balanced temporal sampling substantially outperforms frontier general-purpose VLMs (GPT-5, Gemini) at physically grounded long-horizon planning, suggesting that domain-specific alignment matters more than raw model scale for embodied tasks. The bidirectional interface design—where the VLM is constrained to outputs the VLA can kinematically execute—is a concrete instantiation of the broader challenge of grounding VLM outputs in physical feasibility. For researchers tracking VLMs, the automatic annotation pipeline scaling to 4,000+ hours of heterogeneous video data via a large VLM (Qwen3-VL-235B) also represents a practically significant methodology for building embodied training corpora.