Aerial navigation agents may know how to reach a destination yet still fail because people do not speak like their training data. Researchers from Zhejiang University and Zhejiang Lab report that a learned translation layer can convert short, intent-driven directions into commands a frozen drone-navigation model is more likely to execute. In their strongest test, the front end nearly tripled one agent’s success rate on human-written instructions, although the work remains a simulation study and the gains were much smaller on weaker navigation systems.

The paper, submitted to arXiv on October 7, focuses on an interface mismatch in vision-and-language navigation. These systems are often trained with detailed, trajectory-aligned descriptions, while a user might say something closer to “go to the building by the trees.” Holding routes, initial conditions and the OpenFly navigator fixed, the researchers evaluated 203 episodes with instructions collected from 16 annotators. Success fell from 31.03 percent with the original dataset instructions to 11.33 percent with the human instructions, which the authors interpret as evidence that wording alone can block capabilities the agent has already demonstrated.

Three rewritten instructions are tested as separate simulated drone routes and ranked by their outcomes.
TGIT learns from simulated flight outcomes, ranking alternative rewrites before updating only the translator.

Their proposed system, called the Trajectory-Grounded Instruction Translator, or TGIT, sits in front of an existing navigator. The navigation policy and controller stay frozen. During training, the translator receives shortened instructions and visual context, proposes several expanded commands, and lets the unchanged navigator execute each one in a simulator. Those trajectories are ranked primarily by whether they reach the goal, with additional signals for path efficiency and final distance. The resulting preferences update only a low-rank adapter on the translator, not the flight policy.

The training data is deliberately limited to cases in which the original detailed command succeeds but a weakened version fails under the same conditions. That filter is meant to distinguish a language-interface problem from missing perception or control skill. For each of three UAV agents, the team built a pool of 100 such hard cases, sampled 20 per training round over five rounds, and generated three candidate rewrites for each sampled instruction. At deployment, the feedback loop disappears: the translator gets the user instruction and available RGB frames, emits one command, and exits before the navigator acts.

Three abstract drone agents show strong, partial, and unsuccessful navigation outcomes in separate simulated environments.
Cross-agent tests showed relative gains, but weaker navigators still had low absolute success rates.

On the 203-episode OpenFly evaluation, direct use of the weakened instructions produced a 15.27 percent success rate. Prompt-only rewriting reached 19.70 percent, and supervised fine-tuning on target commands reached 29.55 percent. TGIT reached 37.93 percent. The same adapter, trained on weakened rather than human-written inputs, was then tested without human-specific training and lifted success on the human commands from 11.33 percent to 32.51 percent. The paper says these results suggest that execution feedback teaches the translator which plausible wording actually activates a particular agent, rather than merely teaching it to imitate a reference sentence.

The improvement was not uniform. On OpenFly scenes outside the translator’s selected pool, success rose from 14.11 percent to 25.25 percent in agent-seen environments and from 4.95 percent to 20.79 percent in environments unseen by the agent. CityNav also improved across its reported splits, but absolute success remained low. AirVLN rose from 0.83 percent to 7.50 percent on its selected-seen split and remained at zero on unseen scenes. The authors therefore frame TGIT as a way to recover skills already present in a navigation model, not as a substitute for better perception or control.

There are additional boundaries on the finding. Human-written instructions were evaluated only for the OpenFly selected-seen set because annotation was expensive, and the paper explicitly says a translator trained directly on human inputs was not measured. Path efficiency did not improve in lockstep with success, and the method currently rewrites an instruction once before navigation rather than revising it as the flight unfolds. The researchers identify real-time, closed-loop translation as future work. The broader result is narrower but useful: for embodied agents, a natural-language failure may sometimes be an interface failure, and simulator-grounded translation can be tested before engineers retrain the entire system.