Birdwave Atlas.
Discover
OrganisationsProductsPeopleProjectsBeta
OpportunitiesEvents
Birdwave.
ImprintPrivacy PolicyTerms of UseCommunity RulesReport Illegal ContentContact

© 2026 Birdwave. Aachen, Germany. v0.1.4

Back to productsPre-trained Vision Model

Grounding DINO

IDEA Research

Community
Back to productsPre-trained Vision Model

Grounding DINO

IDEA Research

Community

Open-set object detector that finds objects from free-text prompts — no fine-tuning needed. 52.5 AP zero-shot on COCO; pairs with SAM 2 for text-prompted segmentation, the standard combo for flexible robotic perception.

Save it, compare it, request a quote or record that you have used it — with an account.

Sign in to save or compare
Updated39d ago

Tags

open-vocabularyobject-detectionvision-languagezero-shotopen-source

Open-set object detector that finds objects from free-text prompts — no fine-tuning needed. 52.5 AP zero-shot on COCO; pairs with SAM 2 for text-prompted segmentation, the standard combo for flexible robotic perception.

licenseApache-2.0frameworkPyTorchtaskdetection

Open-set object detector that finds objects from free-text prompts — no fine-tuning needed. 52.5 AP zero-shot on COCO; pairs with SAM 2 for text-prompted segmentation, the standard combo for flexible robotic perception.

Save it, compare it, request a quote or record that you have used it — with an account.

Sign in to save or compare
Updated39d ago

Tags

open-vocabularyobject-detectionvision-languagezero-shotopen-source

Overview

Grounding DINO from IDEA Research marries a DINO transformer detector with grounded language pre-training, producing an open-set detector: give it a text prompt ("the red gear next to the housing") and it returns bounding boxes for novel categories without any fine-tuning — 52.5 AP zero-shot on COCO. In robotics pipelines it is the de-facto language-to-region front end: Grounding DINO proposes boxes from natural-language object descriptions, SAM turns them into precise masks ("Grounded-SAM"), and a pose estimator like FoundationPose lifts them to 6-DoF. Apache-2.0 licensed with checkpoints on Hugging Face; Grounding DINO 1.5/1.6 Pro editions (API-gated) push accuracy and edge speed further.

open-vocabularyobject-detectionvision-languagezero-shotopen-source

Key features

  • Detect novel objects from text prompts — zero fine-tuning
  • 52.5 AP zero-shot COCO (Swin-L)
  • Standard front end of Grounded-SAM robotic perception stacks
  • Apache-2.0 with Hugging Face checkpoints

This product is part of the Birdwave Atlas community.

Contribute evaluations, share integrations and help improve compatibility data.

Community feed

Discussions, evaluations, integrations and updates

At a glance

Task
detection
Architecture
DETR-style transformer (DINO) with grounded language-image pre-training; Swin backbone; open-set via text prompts
Framework
PyTorch

Manufacturer

IR
IDEA Research
View organisation →

At a glance

Task
detection
Architecture
DETR-style transformer (DINO) with grounded language-image pre-training; Swin backbone; open-set via text prompts
Framework
PyTorch

Manufacturer

IR
IDEA Research
View organisation →