123D: Unifying Multi-Modal Autonomous Driving Data at Scale
Publish this paper in The AIPR Journal
Are you an author? Turn this AI review into a permanent, citable journal entry with a cover, open comments, and Scholar metadata.
AIPR assessment
This is a hard, crowded infrastructure problem, not a niche benchmark, because the paper tries to unify many established datasets, sensor modalities, and downstream tasks that the field has optimized separately for years. The strengths reinforce each other: the broad format, open-source release, and multi-task demonstrations make the toolkit credible and useful. The weaknesses also interact: the novelty is limited by a dense prior ecosystem of partial unification efforts, and several claims rest
Abstract
The pursuit of autonomous driving has produced one of the richest sensor data collections in all of robotics. However, its scale and diversity remain largely untapped. Each dataset adopts different 2D and 3D modalities, such as cameras, lidar, ego states, annotations, traffic lights, and HD maps, with different rates and synchronization schemes. They come in fragmented formats requiring complex dependencies that cannot natively coexist in the same development environment. Further, major inconsistencies in annotation conventions prevent training or measuring generalization across multiple datasets. We present 123D, an open-source framework that unifies such multi-modal driving data through a single API. To handle synchronization, we store each modality as an independent timestamped event stream with no prescribed rate, enabling synchronous or asynchronous access across arbitrary datasets. Using 123D, we consolidate eight real-world driving datasets spanning 3,300 hours and 90,000 kilometers, together with a synthetic dataset with configurable collection scripts, and provide tools for data analysis and visualization. We conduct a systematic study comparing annotation statistics and assessing each dataset's pose and calibration accuracy. Further, we showcase two applications 123D enables: cross-dataset 3D object detection transfer and reinforcement learning for planning, and offer recommendations for future directions. Code and documentation are available at https://github.com/kesai-labs/py123d.
Score Breakdown
More from this week
- \(\mathit{SB}(3,n)\) has no Hamiltonian cycle when \(n\) is even: a sign-of-permutation proof, with extension to all odd \(m\equiv 3\pmod 4\)
- Space as a spectroscopic laboratory: High-resolution spectroscopy of the [\(^{13}\)C II] hyperfine structure with SOFIA/upGREAT
- Statistical inference with belief functions: A survey
- AERO-VIS: Asynchronous Event-based Real-time Onboard Visual-Inertial SLAM
- Deterministically finding an element of large order in \(\mathbb{Z}_N^*\)
More in Robotics
- Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis
- vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models
- Learning High-Frequency Continuous Action Chunks in Latent Space
- Dynamic Neural Koopman Distillation for Real-Time Robot Control Using Diffusion Models
- AcroRL: Learning Aggressive Quadrotor Inversion using Bidirectional Thrust