Visual Intelligence Summit: Oct 22 in San Francisco Get your ticket

Multimodal JSONL

A JSONL format for multimodal datasets (i.e. VQA).

Overview

A JSONL format for multimodal datasets (i.e. VQA).

To use this format, you need a directory of images that contains JSONL files with the same name as each image (i.e. "image.png" should have a corresponding file called "image.png.jsonl"). You can then drag and drop the image and corresponding JSONL annotation file for use in a multimodal Roboflow project.

There should be one annotation per line in the JSONL file.

Format Description

Below, learn the structure of Multimodal JSONL.

JSON
{"image":"Beer-Can-Loading_mp4-21_jpg.rf.53ac14e905aa1f13529df4d1191c590e.jpg","prefix":"What's in this image?","suffix":"Beer cans on production line"} {"image":"Beer-Can-Loading_mp4-23_jpg.rf.b0ecce950abd8bce4d4ca7b979b40cb2.jpg","prefix":"What's in this image?","suffix":"Beer cans on production line"} 
Start where you are

Speak with an AI expert

Our team will help you start solving business problems on the first call.

  • Solution architecting
  • Live demonstration
  • Pricing and specifications
  • Feasibility assessment

Over 16,000 organizations build with Roboflow.

Rivian
Pella
Chobani
USG Corporation
BNSF Railway
American Woodmark
Outcomes
Patrick Industries
Peer Robotics