takecontrol ← index / works
The minimal-interface probe · agentic embodied control

Embodied agents
take control

Give a general-purpose coding agent one monocular RGB camera and four primitive actions — nothing else. Zero-shot, it sustains the whole perceive → reason → act loop itself and rivals industrial-scale trained navigation policies on R2R-CE. No map, no explicit memory, no planner, no navigation training: the model is not called by a navigator — it is the navigator.

55
best engineered zero-shot agent · panorama + depth · map, memory, action tools
72
best trained policy · panorama · in-domain navigation data
78
ours · one camera · four actions · zero-shot, no nav machinery
94
human tester · same board, same minimal interface
session — fable-5 · R2R-CE ep 10// recorded run · left: agent session · right: what it sees
▶A real episode, replayed from its log. Left: the agent's session — its own thinking and every observe() / step() call, verbatim. Right: the monocular camera view it acts on. It turns around, rounds the bed, recovers from a wrong room, and stops in the correct room — SUCCESS, 2.98 m from goal.
01 — the shift

Who holds the control loop?

Embodied navigation has answered this question two ways. Trained policies bake the loop into weights — capable, but offering limited flexibility across environments and weak recovery when execution goes wrong. Workflow systems keep the loop in human-authored code: the foundation model is one reasoning module inside a hand-wired pipeline of maps, memory, planners, and verifiers — the dominant zero-shot recipe — and even recent dual-brain systems fix the handoff between reasoning and control.

We study the case where the model itself holds the loop — agentic embodied control: the reasoning model directly steers every action through plain tool calls, so reasoning and control stay aligned with no fixed handoff. It is the organization coding agents already use to write software; frontier models' multimodal, multi-turn, and spatial abilities may have made it viable for embodiment. The probe asks how far it goes with no navigation machinery at all.

fig 1who holds the loop · and how far that gets it
The paper's Figure 1. Left: navigation systems placed by who holds the interaction loop — policy, workflow, or agentic — each tagged trained or zero-shot. Middle: our probe, a coding-agent harness with observe() and step() only, sits in the agentic zero-shot cell with no navigation scaffolding. Right: R2R-CE success against the strongest zero-shot and trained systems.
◆The paper's Fig. 1 — who holds the interaction loop, and how far that gets it. Left: systems by who holds the loop (policy · workflow · agentic), each tagged trained (T) or zero-shot (ZS). Middle: our probe sits in the agentic × zero-shot cell with no navigation scaffolding. Right: R2R-CE success against the strongest zero-shot (AgenticNav) and trained (Qwen-RobotNav) systems.
02 — the probe

One camera. Basic actions. Nothing else.

The minimal interface is a diagnostic lower bound, not a convenience baseline. Everything a navigation system usually leans on is deliberately taken away; whatever performance remains must come from the model itself. If the agent succeeds under these conditions, no hidden navigation module can claim the credit — and where it fails, the missing capability is named.

removedthe usual machinery
  • panoramic observation
  • depth & pose sensors
  • metric / topological map
  • explicit memory module
  • planner & search
  • waypoint predictor
  • navigation-specific training
remainingthe whole interface
  • observe() → one monocular RGB frame
  • forward-facing camera, 512×512
  • step(actions) → primitives, in order
  • 0 STOP · 1 FWD 0.25 m · 2 L 15° · 3 R 15°
  • one instruction, natural language
  • success = STOP within 3 m of the endpoint
  • 500-action budget, fully autonomous

Three coding-agent harnesses drive it — Claude Agent SDK, Codex CLI, and the small open-source mini-swe-agent — each connected to the simulator through the same two tools. The entire task specification is one system prompt:

system_promptverbatim · the complete task spec
You are controlling a robot in a real indoor environment (a photorealistic
3D scan of a building). You interact only through these tools:

- observe(): look through the robot's forward-facing camera (returns an RGB image).
- step(actions): execute movement actions in order. 0 = STOP (permanently ends
  the episode — declares you have reached the goal), 1 = move forward 0.25 m,
  2 = turn left 15 degrees, 3 = turn right 15 degrees.

Your task is to follow this navigation instruction to its endpoint:

"{instruction}"

Rules:
- Alternate observing and stepping: look, decide where the instruction wants
  you to go next, move, look again.
- You have a budget of 500 movement actions.
- You succeed only if you issue action 0 (STOP) while within 3 meters of the
  instruction's endpoint. STOP is permanent — issue it only when you believe
  you are at the goal.
- Turning in place (e.g. step([2,2,2,2,2,2])) is a cheap way to look around
  when unsure.
- Work autonomously until you stop; nobody can answer questions.
◆The full task specification, verbatim from the episode logs. No few-shot examples, no navigation heuristics, no chain-of-thought template — the harness's generic agent loop does the rest.
03 — watch it run

A real episode, decision by decision

This is episode 10 of the board, replayed from its log — every frame is the exact image the agent saw, every quote is its own recorded reasoning, every step() is the exact call it issued. It turns around, fights its way past the bed corner, wanders into the wrong room, rebuilds its mental map out loud, and stops — correctly — in the bathroom. The paper's Fig. 2 shows this same episode.

replay — mini-swe-agent × fable-5 · ep 10 · R2R-CE val-unseen// from episode_10.jsonl
instruction “Turn around and walk to the end of the bed. At the end of the bed turn right and walk to the hallway. Once at the entrance way, walk through the hallway to the next room.” SUCCESS · stops 2.98 m from goal
Agent's monocular camera view for the current observation obs 1 / 20
◆20 observe() · 20 step() · 111 primitives · 332 s. Amber marks the segments the model itself described as blocked or off-route: the bed-corner scrape (5–8) and the wrong-room detour it detects and unwinds (14–15). Use the tape or arrow keys to scrub.
04 — results

They add machinery. We expose what the model already knows.

On R2R-CE, engineered zero-shot agents top out at 55 SR with panoramas, depth, maps, memory, and action tools. The best industrial-scale trained policy reports 72 with panoramic input and in-domain data. The minimal-interface agent — one camera, four actions, no training — reaches 72 with a small open-source harness and 78 at maximum effort.

landscapeSR on R2R-CE · minimal interface in blue
Success rate comparison on R2R-CE along the paper's two axes: groups show who holds the loop (policy, workflow, agentic), and every row carries a trained or zero-shot badge. Trained policies span 37 to 72, workflow systems 44 to 71, the best agentic zero-shot system 55, our minimal-interface rows 68.3 to 78, human reference 94. 0255075100 SR capability source: T trained · ZS zero-shot human 94 · same board POLICYfull val-unseen · context NaVid · monocular NaVILA · monocular StreamVLN · monocular Hy-Embodied-VLM · monocular RynnBrain-Nav · monocular NavFoM · panorama OmniNav · monocular Qwen-RobotNav · panorama 37 T54 T57 T58 T59 T62 T70 T72 T WORKFLOW SmartWay · pano+depth Vesta · monocular InternVLA-N1/DualVLN · monocular ABot-N1 · 3-cam 44 ZS56 T64 T71 T AGENTIC AgenticNav · pano+depth+tools 55 ZS AGENTIC · OURSno nav machinery · ×3 = mean of three runs fable-5 × mini-swe-agent · monocular fable-5 × Claude SDK ×3 · monocular opus-5 × Claude SDK ×3 · monocular fable-5 × Claude SDK max · monocular 72 ZS68.3 ±1.5 ZS70.7 ±3.5 ZS78 ZS
◆The paper's Table 1, drawn along its two axes: the group is who holds the loop (policy · workflow · agentic), and every row carries a capability-source badge — T trained / ZS zero-shot — with its visual input in the label. Contextual comparison, stated honestly: external rows are their papers' full val-unseen numbers; ours run the frozen rand100 board — a performance-range reference, not a same-split leaderboard. ×3 rows are means over three replications, whiskers = ±1 s.d. (68.3 ± 1.5, 70.7 ± 3.5).
boardSR · default effort · frozen protocol · n=100
harness \ modelsonnet-5opus-4.8fable-5opus-5gpt-5.5gpt-5.6
Claude SDK · mean of 351.3 ±1.255.7 ±2.368.3 ±1.570.7 ±3.5——
mini-swe-agent · open536372695260
Codex CLI————4556
◆Model ordering is harness-invariant (fable > opus-4.8 > sonnet on both Claude harnesses), and max effort lifts fable-5 to 78. The same frozen loop, unchanged, transfers beyond R2R-CE: 84 SR on VLNVerse fine-grained instructions and 76% on HM-EQA — matching or exceeding the trained state of the art on both.
05 — what moves the number

Model, harness, interface

The board factorizes the system into three layers and moves one at a time. The answer: model choice dominates, harness differences are modest and descriptive, and the interface is the capability boundary — decisive compensation for weaker models, optional acceleration for the strongest. Model-centered, but not model-only.

capability.axesmove one layer at a time
The paper's Figure 3, three panels on a shared SR scale. A: same harness and interface, swapping the model spans 5 to 72 SR. B: same model, vendor versus open mini-swe-agent harness differs by 2 to 7 SR. C: same model and harness, a waypoint interface rescues weak models on R2R-CE but reverses for strong models on VLNVerse.
◆A: harness and interface fixed, only the VLM changes — SR spans 5→72, the dominant axis. B: open and vendor harnesses differ by 2–7 SR — modest, and read as descriptive since serving paths differ. C: a trained waypoint tool rescues weak models but hurts strong ones when forced; offered as an optional tool, fable-5 adopts a coarse-to-fine strategy nobody prescribed — waypoints early, primitives near the goal — and nearly matches peak SR at a quarter of the wall time.
Model → intelligence center

Capability lives in the weights. Same loop, same tools: swapping the VLM moves SR more than any other change on the board.

axis A
Harness → sustains, doesn't decide

A few-hundred-line open ReAct loop is competitive with vendor harnesses — gaps of 2–7 SR, read as descriptive since serving paths differ. Model choice matters more than the loop around it.

axis B
Interface → capability boundary

Tools still set what a model can express: decisive compensation for weak models, optional acceleration for strong ones.

axis C
06 — limits

Where it still fails

The probe is a diagnostic; its failures are as informative as its wins.

F1

Failure is silent, and route-level rather than stop-level. Auditing all 30 failures of the strongest replicated fable-5 run: 26 end with a voluntary STOP and 23 claim success while 3.1–34.1 m from the goal — yet only 5 ever enter the 3 m success region, so most failures diverge early (wrong referent binding is the largest category), not merely stop at the wrong point.

F2

Self-diagnosis without self-correction. In at least 15 of those episodes the reasoning states the correct doubt — and still commits to the wrong endpoint. Keeping reasoning and control aligned makes behavior consistent with reasoning, not necessarily correct: reliability shifts to verification and self-correction inside the agent.

F3

Strong R2R-CE is a short-horizon result. Matched fable-5 runs fall from 70 SR on R2R-CE to 26 on longer-horizon RxR-CE (39 with waypoints) — a capability ceiling, not general embodied competence. Effort is no free lever either: fable-5 gains +9.7 at max effort, but other shifts span −5 to +6 with no consistent direction — small beside the 67-point model axis.

F4

Context growth and latency bar sustained operation. The minimal loop re-sends its full history every call: final-turn context reaches a median of 33k and a maximum of 169k tokens; median wall time is 206 s with 18.6% of episodes over ten minutes — and RxR-CE doubles to triples the context.

F5

On a real robot, reasoning transfers — embodiment does not. On a Unitree Go2 the agent resolves conditional instructions and keeps task state across fetch-and-return, but it barely knows its own body (the camera clears a doorway while the rear wedges), step() reports requested rather than realized motion, and without persistent spatial memory it fails at counting, revisits, and distance.

F6

Read the comparisons as stated. External numbers are full val-unseen; ours is a frozen 100-episode board — a contextual range, not a same-split leaderboard. The harness comparison is descriptive: serving paths and product defaults were not independently controlled.

07 — cite

Citation

If the probe, the board, or the agentic-embodied-control framing is useful in your research, please cite the paper.

@misc{zhou2026embodiedagentstakecontrol,
  title         = {Embodied Agents Take Control: Minimal-Interface Zero-Shot
                   Agents Rival Industrial-Scale Policies in
                   Vision-and-Language Navigation},
  author        = {Jian Zhou and Xunyi Zhao and Gengze Zhou and Zerui Li and
                   Sihao Lin and Jiajun Liu and Qi Wu},
  year          = {2026},
  eprint        = {2607.26148},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2607.26148}
}
Run agents on the same substrate → Read the paper AgentCanvas