Page content last modified on 2026-07-30.
UVG-VCM: Benchmarking Dataset for Machine-Oriented Visual Data Compression
UVG‑VCM is an open video dataset specifically designed for machine‑oriented codec evaluation. The dataset comprises 20 uncompressed and annotated video sequences released under the CC-BY 4.0 license, most of which are provided in 4K resolution, 60 fps, 16‑bit YUV 4:4:4 format. The sequences cover diverse and realistic machine vision scenarios and support tasks including object detection, tracking, human pose estimation, depth estimation, license‑plate recognition, and instance, semantic, and panoptic segmentation. UVG‑VCM is intended to enable comprehensive and reproducible benchmarking of machine‑oriented video compression techniques.
Please cite the following paper for any usage of the dataset:
T. Partanen, M. Anttila, R. Kortelahti, G. Gautier, A. Mercat, and J. Vanne, “UVG-VCM: Benchmarking dataset for machine-oriented visual data compression,” in Proc. Int. Conf. Qual. Multimedia Exper., Cardiff, United Kingdom, Jun.—Jul. 2026.
About the annotations
Each sequence includes a YUV video file together with machine vision task annotations. Depth maps and
segmentation masks are provided as per‑frame PNG images, while all other annotations are stored in JSON
format. Some JSON annotation files additionally contain instance segmentation masks represented as
polygons. All bounding boxes (track_id in JSON) and instance masks (mask_color in JSON) are assigned
persistent identifiers across frames, enabling object and instance tracking. The depth annotations
correspond to the left view of the stereo camera used during acquisition; the corresponding right‑view
sequence is also provided.
Note: Annotation download links include the annotation version
number. Version v1.0 denotes fully reviewed and manually refined annotations. Annotations labeled v0.9
remain under active refinement and may contain minor inaccuracies, although they are suitable for
benchmarking and evaluation purposes.
Privacy and data protection
The UVG‑VCM dataset has been curated in accordance with established privacy and ethical principles. Personally identifiable information has been removed or anonymized wherever required. In particular, faces and vehicle license plates are masked unless explicit permission for their publication and research use was obtained from the corresponding individuals or vehicle owners. Only content for which the necessary permissions and rights have been secured is distributed in recognizable form.
COCO Categories — Object Detection / Tracking / Segmentation (80 classes)
COCO Pose Keypoints (17 points)
- nose
- left_eye
- right_eye
- left_ear
- right_ear
- left_shoulder
- right_shoulder
- left_elbow
- right_elbow
- left_wrist
- right_wrist
- left_hip
- right_hip
- left_knee
- right_knee
- left_ankle
- right_ankle
DOTA Categories — Oriented Object Detection (15 classes)
Disclaimer
All the information and any part thereof provided on this website are provided « AS IS » without warranty
of any kind either expressed or implied including, without limitation, warranties of merchantability,
fitness for a particular purpose or non infringement of intellectual property rights.
Tampere
University makes no representations or warranties as to the accuracy or completeness of any materials
and information incorporated thereto and contained on this website. Tampere University makes no
representations or warranties that access to this website will be uninterrupted or error-free, that this
website (the materials and/or any information incorporated thereto) will be secure and free of virus or
other harmful components.