Vision-Language Model Benchmarks
Erik KokaljPublished Aug 4, 2026
A model can describe an image and still struggle to locate the exact objects an application needs. For inspection, annotation, or component identification, knowing what is in the scene is only part of the task. The model also needs to show where it is.
Vision-language model benchmarks help measure those capabilities against defined tasks and labeled examples. Roboflow’s RF100-VL benchmark focuses on object detection across diverse image domains, testing whether models can return useful bounding boxes for the objects developers actually need to find.
In this webinar, Matvei Popov, a Research Engineer at Roboflow, compares GPT-6 Astra, Gemini 3.8 Flash, Qwen3.8-Max, and Claude Fable 5.1. Learn how the evaluation works, where few-shot examples improve results, and why the strongest model overall may not lead in every category. You can watch it here or follow along below:
What Do Vision-Language Model Benchmarks Measure?
Vision-language models, or VLMs, process images and language together. Benchmarks can evaluate different abilities, from reading text to answering questions about a scene or locating individual objects.
For object detection, the task is specific: given an image and target classes, return bounding boxes around the matching objects. A useful prediction needs both the correct class and an accurate location.
That distinction matters when choosing a model. Strong image-description results do not establish that a model can identify a small industrial component or follow the labeling conventions of a specialized dataset. The benchmark should reflect the output your application requires.
RF100-VL was designed to test detection across real-world visual domains, including imagery that differs from common object detection datasets.
Test Across Domains with RF100-VL
The RF100-VL research brings together 100 datasets sourced from Roboflow Universe. The collection covers seven broad domains, including aerial imagery, industrial scenes, document processing, medical imagery, sports, and flora and fauna.
The benchmark supports evaluation of specialized detectors such as RF-DETR as well as vision-language models. It also includes annotation instructions and a few-shot setup with 10 labeled examples per class.
For the results in this webinar, Matvei uses RF20-VL-FSOD, a 20-dataset subset that makes evaluation more efficient. Scores are calculated for each dataset and then averaged. These results describe that subset rather than an evaluation across all 100 datasets.
The comparison also focuses on VLM prompting. The models receive class names and, depending on the setting, labeled examples.
Give the Model Examples of What You Mean
A class name does not always fully explain the intended detection. A volleyball serve, for example, requires more context than a familiar object such as a dog. A dataset may also use a specific convention for how tightly a box should enclose an object.
The webinar compares three prompting settings:
| Setting | Information supplied to the model |
|---|---|
| Zero-shot | Target class names without labeled reference examples. |
| One-shot | Class names and one labeled example per class. |
| 10-shot | Class names and 10 labeled examples per class. |
The examples pair an image with the expected bounding-box response. They show the model both what the target looks like and how to represent it.
Although RF100-VL includes textual annotation instructions, the comparison presented here isolates the zero-shot, one-shot, and 10-shot settings without those additional instructions.
One central finding is that newer models can use this context more effectively. In this comparison, few-shot examples generally improve results, particularly for the stronger models.
Choose Metrics That Fit the Output
Object detection benchmarks commonly use mean average precision. The webinar reports mAP50-95, which evaluates detections across intersection-over-union thresholds from 0.50 to 0.95. These thresholds measure how closely predicted boxes overlap the reference boxes.
Average precision also depends on ranking detections by confidence. Specialized detectors typically produce confidence scores, while the VLM responses discussed here return box coordinates without reliable confidence estimates.
To examine performance from another angle, Matvei also uses Rex F1@mIoU, a confidence-independent metric focused on detection and localization. The comparison looks at whether the model rankings hold across both metrics.
For your own evaluation, inspect the predicted boxes alongside the aggregate score. Missed objects, incorrect classes, and poorly placed boxes can impact an application differently, even when they contribute to the same headline metric.
What the Model Comparison Shows
In the webinar’s reported evaluation, GPT-6 Astra leads overall, followed by Gemini 3.8 Flash, Qwen3.8-Max, and Claude Fable 5.1. Claude was evaluated on 19 datasets because it declined requests for one dataset, so its aggregate covers a different set.
The domain breakdown adds useful detail. Gemini 3.8 Flash leads Astra on the aerial imagery category shown, while Astra leads across the other categories discussed. That makes the category results relevant for anyone selecting a model for drone imagery or another specific visual environment.
Few-shot gains also depend on the task. Datasets with ambiguous class names benefit more from reference examples.
More examples do not help every model on every dataset. Some Qwen and Claude results deteriorate when examples are added, particularly where the class name already describes the target clearly.
Comparing prompting settings is part of the evaluation, rather than assuming the longest prompt will perform best.
Apply the Findings to Your Own Images
VLM-assisted annotation and rapid task testing are worth exploring. A small set of labeled references can communicate a specialized visual concept before investing in a larger training dataset. Roboflow’s Astra auto-labeling guide shows how to apply that capability to dataset creation.
Use the benchmark to choose models to test, then compare them on representative images from your application. Include ambiguous classes and difficult examples, and evaluate whether adding references improves the output.
In addition, RF100-VL is public, so exposure during model pre-training cannot be ruled out. Fresh evaluation images from your own environment help establish whether the benchmark gains transfer to your task.
Watch the Benchmark Results Explained
Watch the full webinar for the RF100-VL overview at 1:21, prompting comparisons at 7:30, evaluation metrics at 9:26, model results at 14:22, and domain-specific findings at 22:41.
Then compare vision-language models in Roboflow Playground using your own images to find the model that best fits your detection task.
