Multimodal JSONL
A JSONL format for multimodal datasets (i.e. VQA).
Overview
A JSONL format for multimodal datasets (i.e. VQA).
To use this format, you need a directory of images that contains JSONL files with the same name as each image (i.e. "image.png" should have a corresponding file called "image.png.jsonl"). You can then drag and drop the image and corresponding JSONL annotation file for use in a multimodal Roboflow project.
There should be one annotation per line in the JSONL file.
Format Description
Below, learn the structure of Multimodal JSONL.
JSON
{"image":"Beer-Can-Loading_mp4-21_jpg.rf.53ac14e905aa1f13529df4d1191c590e.jpg","prefix":"What's in this image?","suffix":"Beer cans on production line"} {"image":"Beer-Can-Loading_mp4-23_jpg.rf.b0ecce950abd8bce4d4ca7b979b40cb2.jpg","prefix":"What's in this image?","suffix":"Beer cans on production line"}