Open-set object detector that finds objects from free-text prompts — no fine-tuning needed. 52.5 AP zero-shot on COCO; pairs with SAM 2 for text-prompted segmentation, the standard combo for flexible robotic perception.
Open-set object detector that finds objects from free-text prompts — no fine-tuning needed. 52.5 AP zero-shot on COCO; pairs with SAM 2 for text-prompted segmentation, the standard combo for flexible robotic perception.
Save it, compare it or record that you have used it — with an account.
Grounding DINO from IDEA Research marries a DINO transformer detector with grounded language pre-training, producing an open-set detector: give it a text prompt ("the red gear next to the housing") and it returns bounding boxes for novel categories without any fine-tuning — 52.5 AP zero-shot on COCO. In robotics pipelines it is the de-facto language-to-region front end: Grounding DINO proposes boxes from natural-language object descriptions, SAM turns them into precise masks ("Grounded-SAM"), and a pose estimator like FoundationPose lifts them to 6-DoF. Apache-2.0 licensed with checkpoints on Hugging Face; Grounding DINO 1.5/1.6 Pro editions (API-gated) push accuracy and edge speed further.
Derived from its documented connections — mechanical, electrical, software.
3D/Depth Camera (RGB-D) 1 · Pre-trained Vision Model 1
missing: consumes ros2-distro · Find a ros2-distro
missing: consumes cuda-runtime · Find a cuda-runtime
3 other Pre-trained Vision Model products, closest hero specs and price first. Suggestions, not recommendations.
Grounding DINO
This product
FoundationPose
NVIDIA
YOLO11
Ultralytics
SAM 2
Meta AI
This product is part of the Birdwave Atlas community.
Contribute evaluations, share integrations and help improve compatibility data.
Discussions, evaluations, integrations and updates
Discussions and reviews are for members
Sign in to read what members say about Grounding DINO and to add your own evaluation, integration notes or spec fixes.
Sign in to join the discussion