Training Strategies for Reinforcement Learning-based Fixed-Wing UAV Flight Control System


Khanzada H. R., MAQSOOD A., Basit A., Ur Rehman S. S., Riaz T. B., Ali A.

JOURNAL OF CONTROL AUTOMATION AND ELECTRICAL SYSTEMS, 2026 (ESCI, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1007/s40313-026-01301-w
  • Dergi Adı: JOURNAL OF CONTROL AUTOMATION AND ELECTRICAL SYSTEMS
  • Derginin Tarandığı İndeksler: Emerging Sources Citation Index (ESCI), Scopus, Aerospace Database, Compendex, INSPEC, Materials Science & Engineering Collection (ProQuest), Technology Collection (ProQuest)
  • Orta Doğu Teknik Üniversitesi Kuzey Kıbrıs Kampüsü Adresli: Evet

Özet

Reinforcement learning (RL) is a promising approach for adaptive flight control of fixed-wing unmanned aerial vehicles (UAVs) operating under nonlinear dynamics and disturbances. However, the influence of RL training strategies on hierarchical flight control performance remains insufficiently explored. This study compares three RL-based outer-loop strategies for altitude and heading tracking, while conventional PID controllers stabilize inner-loop angular rates. The RL agent observes altitude and heading errors, their derivatives and integrals, and selected aircraft states, and outputs continuous roll, pitch, and yaw angle commands. The reward function penalizes tracking errors, oscillations, and excessive control effort to ensure smooth and stable responses. Three training paradigms are evaluated using DDPG: centralized training with decentralized execution (CTDE) using cooperative agents, fully decentralized task-specific agents, and a single-agent multi-task (SAMT) framework that learns a unified policy. Under identical disturbance conditions, the decentralized strategy achieves the highest steady-state accuracy (1.34 m altitude error, 0.59 degrees heading error) but requires similar to 30,000 episodes. The centralized strategy converges in similar to 5000 episodes with higher RMSE ( similar to 2.48 m). The SAMT approach converges in fewer than 1000 episodes while maintaining low RMSE ( similar to 1.92 m). These results demonstrate that training architecture significantly affects learning efficiency, robustness, and control smoothness.