Pre-trained Vision ModelYOLO11by UltralyticsTaskdetectionArchitectureSingle-stage anchor-free CNN (C3k2 backbone, C2PSA attention), 5 sizes n→xFrameworkPyTorchSign in to save YOLO11
Pre-trained Vision ModelGrounding DINOby IDEA ResearchTaskdetectionArchitectureDETR-style transformer (DINO) with grounded language-image pre-training; Swin backbone; open-set via text promptsFrameworkPyTorchSign in to save Grounding DINO
Pre-trained Vision ModelSAM 2by Meta AITasksegmentationArchitectureTransformer with streaming memory (promptable, video-capable)FrameworkPyTorchSign in to save SAM 2
Pre-trained Vision ModelFoundationPoseby NVIDIATaskpose estimationArchitectureTransformer with contrastive learning + neural implicit representation (unified model-based/model-free)FrameworkPyTorchSign in to save FoundationPose