Plann3r: Predicting Planning Costs Grounded in 3D

Anonymous Authors
Accepted at CoRL 2026!

Abstract

Planning paths or trajectories for robot navigation requires understanding scene geometry and traversability. Classic approaches relied on accurate 3D maps to define occupancy based planning costs. Learning based alternatives predict planning costs either in terms of distance to goal or temporal distance between images. However, the former overfits to scene layout and the latter lacks geometric understanding. Moreover, most of these methods estimate image-level scalar costs, which are not sufficiently informative to effectively guide the robot. In this paper, we propose Plann3r as a 3D grounded method to predict pixel-level planning costs in terms of geodesic distances for any given set of images with an arbitrary subgoal pixel. We integrate Plann3r within a navigation pipeline, dubbed VGGT-Nav, in two unique ways: the offline mapping cum global planning phase iteratively uses Plann3r to generate subgoals and reference image global costmaps; and in the execution phase Plann3r performs simultaneous localization and local planning to generate planning costmaps to directly condition a learnt control policy. We show that the proposed Plann3r and VGGT-Nav largely outperform baselines respectively on the planning and navigation tasks on an existing prior-traversal navigation benchmark. We further show real world demonstrations that highlight sim-to-real deployment.

Walkthrough

VGGT-Nav, step by step

Two HM3D validation episodes, rendered from the logged runs. Poses, images, costmaps and paths come from the logs. The 3D mesh is only for visualization, since Plann3r works on the reference images in a sliding window.

Scene A

Goal: chair, alternate goal: refrigerator

Scene B

Goal: plant, alternate goal: refrigerator cabinet

Key ideas

What's new

  • Plann3r, a method to predict planning costs in the form of pixel-level geodesic distances for a query image to an arbitrary subgoal pixel given within a submap.
  • VGGT-Nav, a single-camera visual navigation pipeline that leverages 3D grounded backend for mapping, localization and planning, which enables improved control prediction and navigation success rates across multiple tasks of a prior-traversal (seen-environment) navigation benchmark.
  • Simultaneous localization and planning. Plann3r enables simultaneous mapping/localization and planning as the query costmap is implicitly based on query-to-submap cross-attention; this forgoes the need for explicit pairwise feature matching which is more prone to failures.

Method

Plann3r architecture

Plann3r architecture: VGGT backbone, goal token, and MLP cost decoder.

We build Plann3r on top of a pre-trained VGGT model, fine-tuning it to predict geometrically consistent planning costs in terms of geodesic distances from an image pixel to a (sub)goal pixel given within a submap of images.

The VGGT-Nav pipeline

VGGT-Nav pipeline: offline mapping/global planning and online localization/local planning.

We use the proposed Plann3r, based on VGGT, to predict planning costs in terms of geodesic distances from query image pixels to an arbitrary (sub)goal pixel given within a submap. Plann3r is used iteratively both during the offline mapping cum global planning phase and the online execution phase of localization cum local planning, where the output costmaps from the latter are used to predict egocentric trajectory.

Offline: mapping and global planning

Functional map of costs

Plann3r predicts local planning costs on a submap around the goal image. The lowest-cost patch of the image shared with the next submap is the next subgoal, and its cost is added as an offset to the next window. Repeating this back to the start gives every reference image a costmap of global planning costs to the goal, without a 3D occupancy grid or a topological graph.

Online: localization and local planning

Cost-conditioned control

The robot retrieves a submap of 8 reference images, and the subgoal pixel is the minimum cost in their precomputed costmaps. Plann3r takes the query image, the submap and the subgoal together, so its global attention localizes the query without explicit feature matching. A controller with the architecture and training of ObjectReact turns the query costmap into 5 future waypoints in the local BEV space.

Why geodesic

Dense geodesic costmaps

Recent 3D mapping methods do not directly tell whether a place can be reached. That needs reasoning about traversable regions in a geodesic space, not a Euclidean one. Plann3r gives each pixel the geodesic distance from its closest floor point to the subgoal.

Qualitative comparison of planning costmaps: Ground Truth, Plann3r, Euclidean, ObjectReact.

The goal is outside the room and we observe that the side of the wall inside the room gets a high cost (shades of red) in both the Ground Truth and Plann3r generated costmaps, whereas the same wall either gets a low cost (shades of blue) when directly using 3D Euclidean distances or misses a cost gradient in ObjectReact.

Try it

Demo

Pick a scene and goal, then scrub the traverse to see Plann3r's predicted geodesic costmap per frame. Blue is a low cost to the goal and red is a high cost. The second scene is from the video game Sekiro, which the model never saw in training, run with no fine-tuning.

Scene
Goal
Normalize
Cost to goal along the traverse

Loading demo data...

Each goal is a pixel selected in one reference frame (white-and-red marker). The costmaps come straight from Plann3r's offline propagation pass over this scene. Global normalizes color across the whole traverse, so it shows the absolute cost to the goal. Relative normalizes each frame on its own, which shows local structure even far from the goal.

Demonstrations

Mapping and execution in action

Real-world recordings of the VGGT-Nav pipeline.

Offline mapping

Mapping to an alternate goal

Costs propagate from the alternate goal through overlapping local submaps of the traverse.

Execution view

Query, submap, costmap, and rollout

The query image, retrieved submap, predicted costmap and controller waypoints, beside the robot's run to the alternate goal.

Offline mapping

Mapping to the final goal

The same traverse, with costs propagated from the final goal.

Real-world rollout

Reaching the alternate goal

The robot follows the controller's waypoints in closed loop to the alternate goal.

Results

Navigation performance

Method Imitate Alt Goal Shortcut Reverse Average
SPLSSPL SPLSSPL SPLSSPL SPLSSPL SPLSSPL
GNM 78.7982.95 8.7015.44 15.3831.74 3.336.11 26.5534.56
GNM (HM3D) 81.8286.38 0.0010.91 15.3824.57 13.2820.77 27.6235.66
PixNav 42.4246.75 26.0931.66 7.6922.29 16.1625.56 23.0931.57
RoboHop 57.5664.99 30.4338.23 30.7740.87 9.9816.92 32.1940.25
ObjectReact 60.6068.51 21.7426.68 23.0839.64 30.0042.01 33.3644.71
Plann3r + GT-trained Control (Ours) 61.3075.53 58.8869.42 45.3661.10 0.009.80 41.3953.96
Plann3r + Control (Ours) 78.4387.69 65.9277.78 55.6870.82 0.0011.99 50.0162.07

State-of-the-art comparison on four navigation tasks of the IIN-val benchmark (HM3D), in SPL and Soft-SPL (SSPL). Best scores in bold, and the shaded rows are ours. GT-trained Control denotes a controller trained on simulator ground-truth (GT) costmaps but deployed with Plann3r-predicted costmaps.

Costmap quality

Costmap Source All Pixels Below Threshold Within Cost Range
5%15%30%50%100% 0-5%5-15%15-30%30-50%50-100%
Euclidean Distance 2.2842.5332.6172.5612.232 2.2362.4032.4122.2962.645
ObjectReact 3.6503.6003.4833.2602.837 3.7313.3122.9752.6193.304
Plann3r (Ours) 2.3982.4242.3692.2991.986 2.3872.2622.0912.0432.318

Mean Absolute Rank Difference (MARD) between ground truth and predicted costs (sweeping over min-k% cost pixels or k-based bins). Lower is better. Left: cumulative MARD over all pixels below the given cost threshold. Right: MARD restricted to pixels within each cost bin. Euclidean distance is more accurate in the lowest 5% of costs around the subgoal, and Plann3r has the lowest error at every other threshold and bin.

Qualitative navigation overview

Qualitative overview of the navigation pipeline: submap and query costmaps, subgoal/goal pixels, predicted waypoints, and a top-down view.

Qualitative overview of the navigation pipeline. The query costmaps remain geodesically consistent with the submap costmaps and correctly direct the robot at both the timesteps, regardless of the visibility of the final goal in the current image at timestep t1.

BibTeX

@inproceedings{plann3r2026,
  title     = {Plann3r: Predicting Planning Costs Grounded in 3D},
  author    = {Anonymous Authors},
  booktitle = {Conference on Robot Learning (CoRL)},
  year      = {2026},
  note      = {Project page: https://plann3r.github.io/}
}