Turning Depth into Text for Embodied Navigation
DepthJev converts every RGB frame into metric depth and target detections, then writes them as short text facts.
Monocular metric depth and open-vocabulary detection become a compact, metre-level description of the scene. Jev reasons over that description and returns one of eight navigation actions.

Depth Anything 3 estimates metric depth; OWLv2 finds the target in the frame.
Free space in five sectors, a 0.25 m collision check, target distance and recent actions become text facts.
Jev reads the facts, picks one action, and the loop repeats with the next observation.
A common_sense episode. DepthJev resolves the instruction to Pot and reaches it in 9 steps. Step through what the agent sees and what Jev reads.
step 0Success rate on all 300 EB-Navigation episodes. Every other agent passes the image to its decision model; DepthJev passes only text facts.
Baselines from Table 3 of the EmbodiedBench paper.