Visual Intelligence Summit: Oct 22 in San Francisco Get your ticket

Deploying Computer Vision Models as Mircroservices

Brad DwyerPublished Sep 27, 2021
2 min read
SUMMARY

Deploying a computer vision model as a microservice keeps vision logic isolated from application code, which matters because models often require specific versions of Python, CUDA, or specialized hardware that you do not want to force on the rest of your stack. The microservice pattern also handles bursty workloads more cheaply (idle compute costs less when decoupled) and lets the vision team ship new model versions independently without a full application redeploy. Roboflow Inference supports this pattern with a standardized API that runs identically on the hosted cloud or in a Docker container on devices like an NVIDIA Jetson or Luxonis OAK.

Roboflow's philosophy around MLOps revolves around treating your computer vision model as a microservice. The reasons for this are myriad; in this post we highlight the benefits of this approach and how it works in practice.

Why Microservices?

Separation of Concerns

The primary motivation behind deploying your model as a microservice is that it separates concerns. Computer vision models can be finicky, requiring specific dependencies and sometimes even specialized hardware. You don't want to have to build your entire application around these specialized constraints.

With a microservice, your vision model can run the specific version of Ubuntu, Python, CUDA, and Tensorflow it needs (for example), while your application code is running in node.js, Go, or C#.

Scalability and Cost

In many situations, usage of your model is going to be bursty. It will sit idle waiting for something to trigger it then have a flurry of activity all at once.

For example, let's say your model is monitoring 20 security camera feeds at a worksite. You could connect each camera to a powerful machine that runs your model. But a better approach would be to connect each camera to a cheap, low-power device that speaks to a microservice running your model when motion is detected.

Speed of Iteration

With a microservice, updating your model doesn't mean re-deploying your entire app (which, in the case of a mobile app could mean days spent waiting for review, or with software deployed on physical hardware could take weeks or months).

Additionally, if you're working on a team, the folks responsible for your computer vision model can iterate independently of the application's overall release cadence.

How Does it Work with Roboflow?

Roboflow supports a standardized inference API for your models that works across several different platforms. You can test against our autoscaling Hosted API, then deploy the exact same model via a Docker container to devices like the NVIDIA Jetson, a server on your private cloud, or the Luxonis OAK.

When you train a new model, updating just means changing a reference to the version number in a configuration file. Or you can migrate users over slowly via a staged rollout to ensure that your new model performs just as well in the real world.

More AboutModel Deployment

Get started

Build on the Platform

For developers, engineers, and technical founders who want to get hands on. Try the free tier; the docs are open.

Bring it into your operation

For heads of AI, operations leaders, and enterprise teams. Bring a known problem, or work with us to find the first one worth taking on.