Meta's foundation model for promptable segmentation in images and video — point/box prompts, multi-object tracking, streaming memory for real-time video, Apache-2.0.
Meta's foundation model for promptable segmentation in images and video — point/box prompts, multi-object tracking, streaming memory for real-time video, Apache-2.0.
Save it, compare it or record that you have used it — with an account.
Segment Anything Model 2 (SAM 2) extends promptable visual segmentation from images to video by treating an image as a single-frame video. A transformer architecture with streaming memory enables real-time video processing and multi-object tracking from point or box prompts. SAM 2.1 ships four checkpoint sizes (Tiny 38.9M to Large 224.4M params) with 39.5–91.2 FPS on an A100, PyTorch training/fine-tuning code, and Hugging Face integration. Widely used in robotics perception pipelines for open-vocabulary object segmentation and tracking.
Derived from its documented connections — mechanical, electrical, software.
Pre-trained Vision Model 2 · 3D/Depth Camera (RGB-D) 1
missing: consumes bounding-boxes-2d · Find a bounding-boxes-2d
missing: consumes cuda-runtime · Find a cuda-runtime
missing: consumes bounding-boxes-2d · Find a bounding-boxes-2d
3 other Pre-trained Vision Model products, closest hero specs and price first. Suggestions, not recommendations.
SAM 2
This product
FoundationPose
NVIDIA
YOLO11
Ultralytics
Grounding DINO
IDEA Research
This product is part of the Birdwave Atlas community.
Contribute evaluations, share integrations and help improve compatibility data.
Discussions, evaluations, integrations and updates
Discussions and reviews are for members
Sign in to read what members say about SAM 2 and to add your own evaluation, integration notes or spec fixes.
Sign in to join the discussion