MobileNet SSD v2 vs. OpenAI CLIP: Compared and Contrasted

Models

MobileNet SSD v2

This architecture provides good realtime results on limited compute. It's designed to run in realtime (30 frames per second) even on mobile devices.

Learn more about MobileNet SSD v2

OpenAI CLIP

CLIP (Contrastive Language-Image Pre-Training) is an impressive multimodal zero-shot image classifier that achieves impressive results in a wide range of domains with no fine-tuning. It applies the recent advancements in large-scale transformers like GPT-3 to the vision arena.

Learn more about OpenAI CLIP

Model Type

Object Detection

Classification

Model Features

Item 1 Info

Item 2 Info

Architecture

Annotation Format

Instance Segmentation

Framework

TensorFlow 1.5