Offline Reinforcement Learning on a Real-World Power Grid Control Problem
📋 Type
Project
⚡ Status
finished
📅 Duration
Mar 10, 2026 – Jun 10, 2026
👤
Primary supervisor
Julian Oelhaf
👥
Co-supervisors
Siming Bayer
Andreas Maier
🎓 Student
Alexander Luce
Artificial Intelligence, M.Sc.
Abstract
This project explores offline reinforcement learning for power grid protection in a simulated medium-voltage distribution network. Using precomputed fault and non-fault trajectories from the CIGRE MV benchmark system, a Conservative Q-Learning agent is trained to issue line-specific trip or wait decisions from voltage and current measurements. The study compares raw waveforms, phasor-based features, combined input representations, different observation windows, reward settings, and conservative regularization strengths. Results show that combining waveform and phasor information improves protection performance, while conservative regularization is essential for stable offline learning in this safety-critical setting.