Common computer vision questions
Short, practical answers to the questions we hear most: choosing models, training and deploying, dataset problems, and how vision fits with LLMs.
Best model for detecting multiple small objects in real time?
For detecting many small objects in real time, RF-DETR is a top choice. Built on the DETR architecture and fine-tuned by Roboflow, RF-DETR excels at identifying small and densely packed objects. It performs especially well when paired with high-resolution images and tiling, both of which you can easily set up in Roboflow. Unlike YOLO models (YOLOv8, YOLOv9, or YOLO-NAS), RF-DETR uses a transformer-based attention mechanism to pinpoint object locations with high precision, even when objects are overlapping or varied in size. Once trained, you can deploy RF-DETR seamlessly to edge devices using Roboflow Inference, with support for ONNX export and hardware acceleration to ensure real-time performance.
Want to try RF-DETR on your dataset? Deploy now with free GPU.
Cloud vs on-device inference for cv model
When deciding between cloud and on-device inference for your computer vision model, it comes down to speed, connectivity, and control. Cloud inference is great when you need easy scalability, automatic updates, and access to powerful GPUs - perfect for prototypes, dashboards, or apps with stable internet access. But when latency matters or you’re working in environments with limited connectivity (like factories, retail stores, or edge cameras), on-device inference is the way to go. With Roboflow Inference, you can deploy models to devices like Jetson, Raspberry Pi, or even a local CPU, no internet required. It’s the same model, just running right where the data is collected. You can even switch between cloud and edge without retraining.
Need help deciding? Check out our inference deployment guide.
What hardware is best for training YOLOv8 models?
For training YOLOv8 models, your best bet is a GPU with plenty of VRAM - the more, the better. A NVIDIA RTX 4080 or 4090 is recommended if you’re training locally; they offer fast training times and can handle large batch sizes, which improve stability and performance. If you're on a tighter budget, even a 3060 or 3070 will get the job done for smaller datasets. Training on a CPU is possible but extremely slow - not recommended unless you're just testing. Don’t want to worry about hardware at all? Roboflow Train handles the GPUs for you - just click "Train" and let us spin up optimized infrastructure under the hood.
Want to see how your dataset performs before committing to hardware? Try training with Roboflow first.
How to fine-tune a pretrained vision model on a custom dataset
- Step 1: Label your images
After labeling, upload them to Roboflow for object detection, classification, or segmentation model training. - Step 2: Generate a version
Choose your preprocessing steps (resize, format, augmentations) and generate a dataset version. - Step 3: Click “Train New Model”
Select a pretrained model like RF-DETR to fine-tune - no setup required. - Step 4: Adjust training settings (optional)
Tweak epochs, batch size, or enable advanced options like class balancing. - Step 5: Start training
Roboflow spins up optimized GPU infrastructure and begins fine-tuning your model. - Step 6: Evaluate and deploy
Review results, download weights, or deploy directly to cloud or edge with Roboflow Inference.
Want help choosing the right model? Compare training options here.
How do I visualize bounding boxes on my dataset?
Visualizing bounding boxes is a key part of reviewing and managing your computer vision dataset. Use Roboflow to see bounding boxes on your images. Here’s how:
1. How to View Bounding Boxes During Annotation
- In Roboflow Annotate, click on any image in your project to enter the labeling interface.
- Existing bounding boxes will automatically appear over the image.
- You can select, move, and resize bounding boxes, or draw new ones using the bounding box tool (found in the toolbar, usually represented by a rectangle icon).
- To edit a bounding box, simply click on it. Circular handles will appear that let you adjust the size, and you can drag the box to move it.
- You can also switch between images in your dataset to quickly review other bounding boxes.
2. How to Preview Bounding Boxes After Upload or Processing
- Once your images are annotated (either manually or with AI tools like Label Assist or Smart Polygon), Roboflow automatically overlays bounding boxes for each labeled object when viewing images in the Annotation or Dataset tabs.
- This visual preview lets you verify that your dataset annotations are correct and precise before you proceed with model training.
3. How to Visualize Bounding Boxes with Model Predictions
- Within the Roboflow dashboard, after training or uploading a model, run the model on a sample image or batch.
- Roboflow will display output images with predicted bounding boxes clearly drawn over detected objects, helping you visually compare model predictions to your ground truth annotations.
How to convert YOLO annotations to COCO format?
To convert YOLO annotations to COCO format, you'll need to create a Python script or use a library that can parse the YOLO format and generate the corresponding COCO JSON structure. This involves mapping YOLO's normalized bounding box coordinates (center x, center y, width, height) to COCO's [x, y, width, height] format and organizing the data into the required JSON structure with keys like "images", "annotations", and "categories". Simply upload your YOLO-labeled dataset to Roboflow, generate a new dataset version with any preprocessing or augmentations you need, then export that version selecting the COCO JSON format.
Why is my object detection model predicting all objects as the same class?
If your object detection model is predicting everything as the same class, it’s usually a dataset or labeling issue. Double check that your annotations are correctly labeled. Sometimes all objects accidentally get tagged with the same class name or ID. It’s also possible that your dataset is imbalanced, with one class dominating the training data. In Roboflow, you can quickly spot this by reviewing your annotation stats and class distribution in the Versions tab. Try rebalancing your dataset, reviewing your labelmap, and regenerating your version. Small fixes here can lead to big improvements in model performance.
How to fix class imbalance in object detection datasets?
- Add more images of underrepresented classes
Upload targeted examples to boost low-frequency classes. - Use class-specific augmentation
Apply stronger augmentations to underrepresented classes to help the model generalize. - Try tiling
Break large images into smaller sections so small or rare objects get more visibility. - Balance before training
Always review and rebalance your dataset before kicking off a training job - Roboflow makes it easy.
Can GPT-4 see images?
Yes, GPT-4 with vision (often referred to as GPT-4V or GPT-4o) can analyze images alongside text. That means you can show it a photo, diagram, or screenshot and ask questions about what’s in it. While it’s powerful for general understanding, it's not specialized for tasks like object detection or segmentation. That’s where dedicated vision models, like those you train in Roboflow, come in. Think of GPT-4 with vision as a broad visual assistant, and Roboflow as your tool for building production-ready computer vision systems.
What are multimodal AI models?
Multimodal AI models are designed to understand and generate insights across multiple types of data such as images, text, audio, and video all at once. Instead of focusing on just one input (like a photo or a sentence), these models combine signals from different modalities to make smarter, more context-aware predictions. For example, a multimodal model could look at a product image and read a description to answer a question or generate tags. While tools like GPT-4o bring general multimodal capabilities, Roboflow helps you build specialized vision modelsthat plug into those pipelines, giving you full control over the visual side of your AI stack.
How to use LLM with computer vision pipeline
- Step 1: Run your vision model
Use Roboflow to detect or segment objects in an image - generate structured JSON output. - Step 2: Pass results to an LLM
Feed your detection output (labels, counts, locations) into a large language model like GPT-4 to interpret or summarize. - Step 3: Add context or instructions
Prompt the LLM with domain-specific questions like “What does this quality defect mean?” or “Summarize what was found in this image.” - Step 4: Build the chain
Combine vision and language steps in a single pipeline using tools such as Roboflow Workflows and Inference + OpenAI API. - Step 5: Deploy your intelligent agent
Connect your pipeline to a dashboard, API, or automation system - now it sees and speaks.
Roboflow handles the vision side, LLMs handle the reasoning. Together, they power next-gen multimodal workflows.
How do I improve accuracy on a custom object detection model?
- Step 1: Clean your labels
Use Roboflow’s annotation tools to fix mislabeled or missing objects. - Step 2: Balance your classes
Check class distribution in the Versions tab and add more examples where needed. - Step 3: Apply smart augmentations
Use blur, mosaic, and rotation to help your model generalize to new conditions. - Step 4: Train longer or tune settings
Adjust epochs, learning rate, or batch size in advanced training settings. - Step 5: Upgrade your architecture
Switch to a higher-performing model such as RF-DETR for better accuracy. - Step 6: Add more data
Take data from your production environment and label more data.
Roboflow makes it easy to test and compare training runs so you can see what actually improves your model.
What’s the best way to deploy a computer vision model on a Raspberry Pi?
- Train your model using Roboflow Train or export from your preferred framework (YOLOv5, YOLOv8, etc.).
- Convert your model to a lightweight format (e.g. TensorFlow Lite, ONNX, or PyTorch with quantization).
- Use Roboflow’s Python SDK or Docker template optimized for Raspberry Pi.
- Leverage hardware acceleration (like the Pi Camera and Coral USB Accelerator if available).
- Run real-time inference directly on the Pi with minimal latency and power usage.
- Monitor, test, and iterate - all from the Roboflow dashboard.
How do I use image augmentation to improve model performance?
To improve model performance, use Roboflow's built-in image augmentation options during dataset generation (such as flip, rotate, blur, brightness). Simulate real-world variation to help your model generalize better. Apply augmentations like noise, cutout, and weather effects to boost robustness. Target rare classes with extra augmentation to fix class imbalance. Avoid over-augmenting, which can confuse the model; start light and increase gradually. Preview and version your augmented dataset in the Roboflow platform. Train with both original and augmented images for improved accuracy and reduced overfitting.
How can I detect multiple object classes in one model?
Detecting multiple object classes in one model is straightforward with Roboflow. When you upload and label your dataset, simply include all the classes you want your model to recognize: each with accurate annotations. Roboflow’s training pipeline automatically supports multiclass detection, so your model learns to differentiate and predict each class simultaneously. Make sure your dataset is balanced across classes to avoid bias, and consider using augmentations to help the model generalize. Once trained, your model will confidently identify and classify multiple object types in a single pass.
How do I evaluate an object detection model?
Evaluating an object detection model starts with measuring key metrics like precision, recall, and mean Average Precision (mAP) to understand how well your model identifies and localizes objects. In Roboflow, evaluation is made simple: after training, you get detailed performance reports with easy-to-read charts showing where your model excels and where it struggles. You can visually inspect predictions against ground truth in the web app or use the Python SDK to analyze errors and tune your model further. Regular evaluation helps you track progress and ensures your model is ready for real-world deployment.
What format should my dataset be in for training YOLO?
For training YOLO models, your dataset should be in the YOLO format, which uses text files with normalized bounding box coordinates (center x, center y, width, height) alongside your images. Roboflow supports seamless upload of images and labels in many formats and can automatically convert between YOLO, COCO, Pascal VOC, and more. Just upload your dataset to Roboflow, generate a version, and export it in the YOLO format ready for training.
How do I retrain a vision model after updating my dataset?
The easiest way to retrain a vision model after updating your dataset is to use Roboflow.
- Step 1: Upload new images or fix labels
Add more data or make annotation corrections directly in your Roboflow project. - Step 2: Generate a new dataset version
Click Generate to create a fresh version with your updated images, labels, and augmentations. This ensures all preprocessing and augmentation settings are consistent with your prior training run - Step 3: Start a new training run
Go to the Train tab and launch training on the new version - no need to start from scratch manually.- Select your model architecture (e.g., RF-DETR, Roboflow 3.0, YOLOv12) based on your project type.
- Choose a model size (Fast for quick iteration, Accurate for production).
- Select a checkpoint:
- Train from previous checkpoint to leverage existing model knowledge.
- Train from scratch for a fresh start (advanced use only).
- Step 4: Compare performance
Use Roboflow’s built-in training metrics and version history to see how the updated model performs.
With version control and one-click retraining, Roboflow makes it easy to iterate fast and improve your model over time.
Can I run a computer vision model without an Internet connection?
Yes, you can run a computer vision model completely offline using Roboflow Inference. Once you’ve trained or uploaded your model, you can export it and deploy it to edge devices like a Jetson, Raspberry Pi, or even a local CPU or GPU machine. Roboflow Inference supports fully offline execution, so once the model and runtime are set up, you don’t need an internet connection to process images or video streams. This enables secure, low-latency, and cost-effective predictions in air-gapped or remote environments.
What’s the fastest way to prototype a vision model for my business?
The fastest way to prototype a vision model for your business is to use Roboflow.
- Step 1: Collect and upload images
Start with just 20–50 labeled images - upload them to Roboflow in minutes. - Step 2: Label or verify annotations
Use Roboflow’s built-in annotation tool to label objects or fine-tune existing labels. - Step 3: Generate a dataset version
Apply preprocessing, set image size, and add augmentations to improve generalization. - Step 4: Train a model in one click
Use Roboflow Train to start training, no setup, no GPUs required on your end. - Step 5: Build your app
Build and test your workflow visually with Roboflow Workflows - drag in your model, set up logic, and run sample images in your browser. - Step 5: Test and deploy instantly
Run inference in the browser, via API, or on edge devices using Roboflow Inference. - Step 6: Iterate
Iterate fast: update data, retrain, and redeploy, from the same dashboard.
How do I monitor performance of a deployed vision model?
To monitor the performance of a deployed vision model, use Roboflow’s built-in tools to track predictions, confidence scores, and real-world edge cases over time. Whether you’re deploying via the Hosted API or Roboflow Inference on-device, you can log outputs, visualize results, and compare predictions against ground truth data. This helps you catch drift, identify failure modes, and continuously improve your model by feeding new edge cases back into your dataset.
Simply head to the “Monitoring” tab in your workspace to access real-time insights like total inference requests, average prediction confidence, and average inference time, all filterable by time period. You can drill down to see performance for individual models, review recent inferences - including actual images and detections - and even set up alerts for anomalies or key events.
How do I label images for object detection?
The best way to label images for object detection is with Roboflow.
- Step 1: Create a project in Roboflow and select “Object Detection” as the project type to start labeling with bounding boxes.
- Step 2: Add your images by dragging them into the project: JPG, PNG, and other formats are supported.
- Step 3: Click into the Annotate tab to launch Roboflow’s labeling workspace.
- Step 4: Draw bounding boxes around each object of interest in your images using the intuitive annotation tool.
- Step 5: Assign class labels to each box- type or select from your predefined list to keep labels consistent.
- Step 6: Use keyboard shortcuts to speed up annotation (e.g., copy/paste boxes, switch classes, zoom).
- Step 7: Leverage auto-labeling tools like Roboflow Annotate Assist to accelerate labeling with AI-powered suggestions.
- Step 8: Review and refine: double-check annotations for accuracy, fix any missed or incorrect labels, and use the “Review” mode for quality control.
- Step 9: Collaborate with your team - invite others to label, review, or manage your dataset in real time.
- Step 10: Export or version your labeled dataset when ready for training, ensuring reproducibility and easy rollback.
What’s the best image size for training a computer vision model?
The best image size for training a computer vision model depends on your use case. A great starting point is to use square images between 224x224 and 640x640 pixels - these sizes are widely supported by popular model architectures like RF-DETR, YOLO and ResNet. Larger images can capture more detail but require more compute and memory, while smaller images train faster but may lose important features. For most projects, resizing all images to a consistent size (ideally a multiple of 32) ensures smooth training and compatibility with modern models. If your task involves detecting small objects or fine details, opt for higher resolutions such as 1024x1024; otherwise, stick to standard sizes for efficiency. Roboflow’s preprocessing defaults make this easy - just select your target resolution when generating a dataset version, no manual resizing required - so you can focus on building and iterating, not wrangling image dimensions.
How do I train a model with a small dataset?
Training a model with a small dataset is totally doable when you use Roboflow. First, make sure your annotations are clean and accurate. Then, apply data augmentation like flips, blur, and rotation when generating your dataset version, this helps your model generalize by simulating real-world variation. Use a pretrained model (like Roboflow 3.0, RF-DETR or YOLOv8) to leverage transfer learning so you’re fine-tuning instead of starting from scratch. Roboflow Train handles this out of the box, so you get great performance even with limited data. The key: clean labels, smart augmentations, and transfer learning. Monitor your results with built-in evaluation tools, and iterate by refining labels or adding more data as you spot weaknesses.
Can I use multiple datasets to train one model?
Yes, you can absolutely use multiple datasets to train one model. Simply upload each dataset into the same project in Roboflow or use Roboflow’s dataset merging feature to combine them. Just select the datasets you want, click “Merge Datasets,” and Roboflow will create a unified dataset without duplicating your originals. As long as the class labels are consistent (after merging, review and standardize your class labels), your model will learn to detect objects across different environments, lighting conditions, or sources. This is a great way to boost generalization and improve performance on real-world data. Roboflow handles the formatting, class unification, and versioning.
How do I automatically label images with a pre-trained model?
Automatically labeling images with a pre-trained model is easy with Roboflow. Just upload your images to a project, head to the Annotate tab, and click Auto Label. Enter prompts describing the objects you want to detect (like “car” or “yellow object”). You can choose from a list of available models - including your own previously trained models - to generate predictions across your unlabeled images. Roboflow will apply bounding boxes and class labels automatically, giving you a huge head start on annotation. You can preview and adjust results on a sample batch, tweak confidence thresholds, and then run auto-labeling on your entire dataset with one click.
How can I improve model performance on edge cases?
To improve model performance on edge cases, start by identifying where your model struggles. Once you spot failure modes (such as missed detections or misclassifications), collect more examples of those specific scenarios and add them to your dataset. You can also use Roboflow’s versioning to fine-tune your model with targeted data, apply class-specific augmentations, or use tiling to boost visibility for small or rare objects.
How do I build a computer vision API without writing code?
Building a computer vision API without writing code is easy with Roboflow. After training your model, just drag and drop blocks in the Workflows editor to connect your model, set up logic, and define outputs like JSON responses or webhooks. Then click “Deploy” in the Workflow editor to host your API in the cloud or run it locally - Roboflow handles all the infrastructure for you. Roboflow automatically provides you with a ready-to-use REST endpoint where you can send images and get back predictions, complete with bounding boxes, class labels, and confidence scores. Just copy the curl command or use the Python snippet generated for you. Update your Workflow anytime to add new logic, models, or integrations - all changes go live with a click, keeping your API flexible and business-ready.
How can I deploy a vision model into a production app?
The best way to deploy a vision model into a production app is Roboflow.
- Step 1: Train or upload your model in Roboflow
Use Roboflow Train or upload custom weights to get started. - Step 2: Choose your deployment method
Decide between cloud (scalable APIs), edge (on-device for low latency), or hybrid. Head to the Deploy tab and select Hosted API for cloud apps or Roboflow Inference for edge/on-device use. - Step 3: Use the provided code snippet
Copy the prebuilt Python, JavaScript, or cURL snippet to send images and receive predictions instantly. - Step 4: Integrate into your app
Connect the Roboflow API to your frontend, backend, or automation, it's REST-based and easy to use.
- Step 5: Monitor and iterate
Log predictions, collect edge cases, and retrain your model with new data, all within Roboflow. - Step 6: Secure your deployment
Protect endpoints with API keys, HTTPS, and workspace-level access controls.
What’s the difference between classification, detection, and segmentation?
Classification, detection, and segmentation are three core types of computer vision tasks, each offering a different level of visual understanding. Classification tells you what’s in an image (e.g. “This is a cat”), but not where. Object detection adds location by drawing bounding boxes around each object (“There are two cats and one dog”). Instance segmentation goes even further by outlining each object with a pixel-perfect mask - ideal for overlapping or irregular shapes. You can build models for all three depending on how precise your application needs to be using Roboflow. Get started free.
What’s the difference between object detection and instance segmentation?
Object detection and instance segmentation both locate objects in images, but they differ in precision. Object detection draws bounding boxes around objects and labels them. For example, drawing a rectangle around each car in a street scene. Instance segmentation goes further by returning a pixel-perfect mask for each individual object, even when they overlap. It’s like moving from a rough outline to a detailed cutout. If you need exact shape, position, or surface measurements, such as detecting defects on circuit boards or counting overlapping cells - instance segmentation is the way to go. You can build models for both depending on how precise your application needs to be using Roboflow. Get started free.
Roboflow free vs paid
Roboflow offers both free and paid plans depending on your computer vision needs. The free tier requires no credit card and is great for getting started and personal projects: it includes unlimited image uploads, annotation tools, dataset versioning, and access to public models. Keep in mind all of your datasets and models are listed publicly on Roboflow Universe and you’re limited to one workspace.
For privacy, scale, or advanced capabilities for production, upgrading to a paid plan is the way to go. The paid plans unlock features such as private projects, more training credits, advanced augmentations, team collaboration, commercial usage rights, model monitoring and professional labeling, and priority support. Whether you're prototyping or deploying to production, Roboflow makes it easy to start free and upgrade as you grow. Learn more about pricing.