NVIDIA's unified 6D object pose estimation and tracking model — model-based (CAD) or model-free (reference images), no fine-tuning for novel objects. #1 on the BOP leaderboard at release. Research use only.
Input
RGB-D
Per frame: RGB image + aligned depth image (metres) + 3x3 camera intrinsics K. First frame also needs a binary object mask. Per object: a CAD mesh (model-based) or ~16 reference views to build a neural object field (model-free).
Output
6D object pose
4x4 object-to-camera pose per frame (register on frame 0, then track_one).
From the repo's run_demo.py. Inputs: RGB, depth in metres, intrinsics K, and an object mask on the first frame. Output: a 4x4 object-to-camera pose per frame.
# Model-based setup: CAD mesh of the object + an RGB-D stream + camera intrinsics
mesh = trimesh.load("object.obj")
est = FoundationPose(model_pts=mesh.vertices, model_normals=mesh.vertex_normals,
mesh=mesh, scorer=ScorePredictor(), refiner=PoseRefinePredictor(),
glctx=dr.RasterizeCudaContext())
# Frame 0: register (needs a binary mask of the object)
pose = est.register(K=K, rgb=rgb, depth=depth, ob_mask=mask, iteration=5)
# Frames 1..n: track (no mask needed)
pose = est.track_one(rgb=rgb, depth=depth, K=K, iteration=2)
# pose: 4x4 object-to-camera transform for that frameNVIDIA's unified 6D object pose estimation and tracking model — model-based (CAD) or model-free (reference images), no fine-tuning for novel objects. #1 on the BOP leaderboard at release. Research use only.
Save it, compare it or record that you have used it — with an account.
FoundationPose (CVPR 2024 Highlight) is a unified foundation model for 6D object pose estimation and tracking that works both model-based (with a CAD model) and model-free (from a few reference images via neural implicit novel-view synthesis). Trained on large-scale synthetic data with LLM-aided generation and contrastive learning, it generalises to novel objects with no test-time fine-tuning and ranked #1 on the BOP leaderboard for model-based novel-object pose estimation at release (Mar 2024). ROS integration is available through NVIDIA Isaac ROS Pose Estimation.
Derived from its documented connections — mechanical, electrical, software.
3D/Depth Camera (RGB-D) 1 · Pre-trained Vision Model 1
missing: consumes cuda-runtime · Find a cuda-runtime
missing: consumes bounding-boxes-2d · Find a bounding-boxes-2d
3 other Pre-trained Vision Model products, closest hero specs and price first. Suggestions, not recommendations.
FoundationPose
This product
YOLO11
Ultralytics
SAM 2
Meta AI
Grounding DINO
IDEA Research
This product is part of the Birdwave Atlas community.
Contribute evaluations, share integrations and help improve compatibility data.
Discussions, evaluations, integrations and updates
Discussions and reviews are for members
Sign in to read what members say about FoundationPose and to add your own evaluation, integration notes or spec fixes.
Sign in to join the discussion