<p>Autonomous driving systems (ADS) are an emerging technology with promising potential in areas such as intelligent cities and transportation. Testing ADS is of paramount importance before their deployment in real-world vehicles, where simulation-based testing provides a cost-effective way to assess the performance of ADS. Several simulation-based fuzz testing tools have been developed, but there lacks a systematic empirical study to evaluate and compare their performance. To address this, we propose the first comprehensive evaluation framework with unified initial driving scenarios and violation oracles to ensure fair comparisons. We conducted extensive experiments of over 500 hours. The results demonstrate that initial driving scenario datasets may impact the performance of fuzzers, and introduce potential evaluation biases. Furthermore, we consider two additional metrics, i.e., map waypoint coverage and code coverage. We find that the map waypoint coverage can be a complementary indicator to the number of unique violations to evaluate the performance of fuzz testing methods, whereas the code coverage fails to distinguish performance differences among fuzzers. These insights may provide guidance for comprehensively evaluating and comparing ADS simulation testing approaches in future studies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Empirical evaluation of simulation-based fuzz testing for autonomous driving systems

  • Huiwen Yang,
  • Yu Zhou,
  • Taolue Chen

摘要

Autonomous driving systems (ADS) are an emerging technology with promising potential in areas such as intelligent cities and transportation. Testing ADS is of paramount importance before their deployment in real-world vehicles, where simulation-based testing provides a cost-effective way to assess the performance of ADS. Several simulation-based fuzz testing tools have been developed, but there lacks a systematic empirical study to evaluate and compare their performance. To address this, we propose the first comprehensive evaluation framework with unified initial driving scenarios and violation oracles to ensure fair comparisons. We conducted extensive experiments of over 500 hours. The results demonstrate that initial driving scenario datasets may impact the performance of fuzzers, and introduce potential evaluation biases. Furthermore, we consider two additional metrics, i.e., map waypoint coverage and code coverage. We find that the map waypoint coverage can be a complementary indicator to the number of unique violations to evaluate the performance of fuzz testing methods, whereas the code coverage fails to distinguish performance differences among fuzzers. These insights may provide guidance for comprehensively evaluating and comparing ADS simulation testing approaches in future studies.