Improving Autonomous Driving Performance through Risk-Constrained Adaptive Reinforcement Learning
DOI:
https://doi.org/10.64972/jaat.2023v1.314p27e:361-374Keywords:
Autonomous Driving, Adaptive Control, Risk-Constrained Optimization, Context Estimation, Safety ShieldAbstract
Autonomous driving policies trained in a fixed traffic environment may not be efficient or safe when road friction, traffic density, sensing quality, or driver behaviour changes. This paper introduces a risk-constrained adaptive reinforcement learning framework that integrates online context estimation with a mixture-of-experts driving policy, uncertainty-aware constrained optimisation, and a control-barrier safety shield. The controller was trained in a procedurally varied simulator and evaluated over 12,000 episodes at urban intersections, multilane highways, under rain and sensor corruption, and in previously unseen behaviour profiles. Compared with the non-adaptive soft actor-critic baseline, the proposed method raised the route completion rate from 91.8% to 97.1%, reduced collision frequency from 2.84 to 0.91 events per 1,000 km, and lowered mean travel time by 11.6%. With the combination of weather and perception shift, it still has a 94.3% completion rate and limits the 95th-percentile lateral jerk to 2.31 m/s³. Ablation results show that context adaptation adds 3.4 per cent to the completion rate, and the safety shield removes 58.7 per cent of the remaining high-risk behaviours. Policy adaptation needs 6.8 ms per control cycle on an automotive graphics processor and still has enough margin for a 20 Hz planning loop. Therefore, the improvement in performance will be more consistently obtained if adaptation is not directly tied to explicit risk constraints, and uncertainty directly influences the size of the update.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2023 Juliusz Ignasiak

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.