Empirical evaluation of simulation-based fuzz testing for autonomous driving systems
摘要
Autonomous driving systems (ADS) are an emerging technology with promising potential in areas such as intelligent cities and transportation. Testing ADS is of paramount importance before their deployment in real-world vehicles, where simulation-based testing provides a cost-effective way to assess the performance of ADS. Several simulation-based fuzz testing tools have been developed, but there lacks a systematic empirical study to evaluate and compare their performance. To address this, we propose the first comprehensive evaluation framework with unified initial driving scenarios and violation oracles to ensure fair comparisons. We conducted extensive experiments of over 500 hours. The results demonstrate that initial driving scenario datasets may impact the performance of fuzzers, and introduce potential evaluation biases. Furthermore, we consider two additional metrics, i.e., map waypoint coverage and code coverage. We find that the map waypoint coverage can be a complementary indicator to the number of unique violations to evaluate the performance of fuzz testing methods, whereas the code coverage fails to distinguish performance differences among fuzzers. These insights may provide guidance for comprehensively evaluating and comparing ADS simulation testing approaches in future studies.