Dataset

A multimodal hybrid dataset for validating perception systems in rare, safety‑critical, and underrepresented driving scenarios

Dataset

A multimodal hybrid dataset for validating perception systems in rare, safety‑critical, and underrepresented driving scenarios

The just better DATAset is a multimodal dataset designed to support robust validation and analysis of perception systems in safety‑critical driving scenarios. The dataset has been curated with a strong focus on underrepresented, rare, and safety‑relevant situations, addressing known gaps in conventional real‑world driving data.

Dataset Overview

The dataset consists of three data batches, each collected using three independent test vehicles from the following project partners:

  • AVL Deutschland GmbH (AVL)
  • b-plus technologies GmbH (b-plus)
  • FZI Forschungszentrum Informatik (FZI)

equipped with heterogeneous multimodal sensor setups, including:

  • LiDAR
  • Radar
  • Camera

This setup enables comprehensive multimodal perception research and provides realistic sensor diversity reflective of real‑world deployments.

Focus on Safety‑Critical and Long‑Tail Scenarios

Unlike generic driving datasets, the jbDATAset explicitly targets challenging and safety‑critical scenarios, including but not limited to:

  • Unknown or uncommon objects on the road
  • Diverse and adverse weather conditions
  • Situations that are typically underrepresented in large‑scale fleet data

This makes the dataset particularly well‑suited for validation, robustness testing, and failure analysis of perception models.

Hybrid Real & Synthetic Data Design

The dataset follows a hybrid data strategy, combining:

  • Real‑world sensor recordings
  • Carefully generated synthetic data

The synthetic components are not intended to replace real data, but to systematically complement it, enabling controlled variation, improved coverage of rare scenarios, and targeted stress testing. This hybrid approach is especially valuable for model validation and generalization assessment.

Ground Truth and Pseudolabels

To support different stages of development and evaluation, the dataset provides:

  • High‑quality ground truth annotations for selected samples
  • Pseudolabels for other parts of the dataset

This reflects realistic industrial workflows and allows users to study:

  • Model performance under varying label fidelity
  • The impact of pseudolabeling in validation and benchmarking pipelines
  • Strategies for combining ground truth and weak supervision

Intended Use

The jbDATAset is designed primarily for:

  • Validation and robustness evaluation
  • Safety‑oriented perception research
  • Analysis of long‑tail and rare scenarios
  • Benchmarking multimodal perception systems
  • Research on hybrid real‑synthetic data strategies

It is not intended as a generic training dataset, but as a high‑value asset for testing, analysis and validation, particularly in safety‑critical contexts.

Publication

The consortium of just better DATA will publish and present a paper about the jbDATAset within the Curated Data for Efficient Learning Workshop at the ECCV 2026 in Malmö (September 8 – 12th, 2026). Once the paper is published it will be linked here.

Key Information about the data batches

AVLB-PLUSFZI
NUMBER of SEQUENCES151635500
LENGTH of SEQUENCES20 sec10 sec20 sec
DATA SIZE~0.97TiB2.0TiB~1.7TiB
DATA TYPES- Camera images
- LiDAR pointclouds
- Camera images
- LiDAR pointclouds
- Radar pointclouds
- Camera images
- LiDAR pointclouds
- Radar pointclouds
ANNO-TATIONS- 3D Bounding Boxes
(Detection 9Hz,
Tracking 5Hz, LiDAR)
- 2D Bounding Boxes
(6Hz, Front Camera)
none- 3D bounding boxes (2 Hz, human annotated)
- 3D LiDAR semantic segmentation (2 Hz, human annotated)
- 2D bounding boxes (2 Hz, pseudolabels, front camera)
WEATHER- Sun
- Rain
- Snow
- Sun
- Rain
- Snow
- Fog
- Sun
- Rain
ROADTYPES- Urban
- Rural
- Highway
- Urban
- Rural
- Highway
- Urban
- Rural
- Highway
TRIGGERS- Criticality detection
- Anomaly Detection
(unusual acceleration
behaviour of Ego Vehicle)
nonenone

Preview of the data batches

AVL Dynamic Ground Truth (DGT) is a highly precise reference measurement system for automated and connected driving. It was developed to operate as an independent, objective environmental reference for sensor evaluation. The modular roof-mounted system is equipped with 3 lidar sensors, 6 cameras, and a dGPS (differential GPS) device for centimeter-precision dynamic positioning.

The dataset was recorded on a variety of scenes, including complex traffic situations, targeted operational design domains. The recorded dataset provides 3D LiDAR detections and tracking alongside 2D object detections from the front-facing camera, with every tenth annotated with 2D Bounding Boxes.

 

The b-plus research vehicle “NOVA” was developed as part of the jbDATA project and specifically designed to meet its requirements. It features a multimodal sensor suite comprising cameras, LiDAR, and radar. The sensors are arranged to closely resemble the configuration found in production vehicles. In addition, the setup is complemented by reference sensors.

The dataset was primarily collected in eastern Bavaria and includes urban, rural, and highway scenarios. Recordings were conducted across multiple seasons to capture a wide range of weather conditions, including heavy snow, rain, and fog, as well as varying lighting conditions.

 

CoCar NextGen is a multi purpose research platform for automated and connected driving. It was set up in-house and operates independently from industry manufactures and OEMs. The Audi A6 Avant plug-in hybrid is equipped with 12 state-of-the-art lidar sensors, 3 radars, 9 cameras, a Car2X onboard unit, and a high precision IMU unit with dual antenna GNSS. The modular design facilitates its use in various applications and research fields in new mobility concepts.
The dataset was recorded on a variety of scenes, including urban, cross country and highway driving in various weather conditions. A total of 500 sequences were selected for the jbDATAset; of these, every fifth frame contains 3D bounding boxes and a semantic segmentation of the point cloud, whilst every tenth frame contains 2D bounding boxes from the front camera.

The jbDATA Synthetic Dataset contains 10,000 automatically annotated synthetic driving images generated to complement real-world sensor recordings with scenarios that are difficult to acquire and balance through conventional data collection alone. Developed within the jbDATA Smart Data Loop framework, the dataset targets underrepresented operating conditions relevant to autonomous driving, with a particular focus on adverse weather, including rain and snow, challenging illumination, nighttime scenes, tunnels, and other situations that can impact perception performance and robustness. The synthetic data was created as part of a targeted data-enrichment process guided by coverage and data-gap analysis, supporting the development and evaluation of AI systems under challenging environmental conditions. The release is provided in the widely adopted MS COCO format with object-level annotations, enabling straightforward integration into existing computer vision validation and benchmarking workflows.

The following picture shows some examples of the AUMOVIO synthetic dataset:

Download

Please note that the data batches are published under the following license: CC-BY-SA 4.0.

Please refer to the respective partner’s download page for more information and the correct citation of the data batches.


Data batch #01

AVL


data batch #02

b-plus


data batch #03

FZI

Synthetic
data batch

Aumovio

Nach oben scrollen