Planning paths or trajectories for robot navigation requires understanding scene geometry
and traversability. Classic approaches relied on accurate 3D maps to define occupancy based
planning costs. Learning based alternatives predict planning costs either in terms of
distance to goal or temporal distance between images. However, the former overfits to scene
layout and the latter lacks geometric understanding. Moreover, most of these methods estimate
image-level scalar costs, which are not sufficiently informative to effectively guide the
robot. In this paper, we propose Plann3r as a 3D grounded method to predict pixel-level
planning costs in terms of geodesic distances for any given set of images with an arbitrary
subgoal pixel. We integrate Plann3r within a navigation pipeline, dubbed VGGT-Nav, in two
unique ways: the offline mapping cum global planning phase iteratively uses Plann3r to
generate subgoals and reference image global costmaps; and in the execution phase Plann3r
performs simultaneous localization and local planning to generate planning costmaps to
directly condition a learnt control policy. We show that the proposed Plann3r and VGGT-Nav
largely outperform baselines respectively on the planning and navigation tasks on an
existing prior-traversal navigation benchmark. We further show real world demonstrations
that highlight sim-to-real deployment.