Offline Reinforcement Learning on a Real-World Power Grid Control Problem

📋 Type Project
Status finished
📅 Duration Mar 10, 2026 – Jun 10, 2026
👤 Primary supervisor Julian Oelhaf
👥 Co-supervisors Siming Bayer Andreas Maier
🎓 Student Alexander Luce Artificial Intelligence, M.Sc.

Abstract

This project explores offline reinforcement learning for power grid protection in a simulated medium-voltage distribution network. Using precomputed fault and non-fault trajectories from the CIGRE MV benchmark system, a Conservative Q-Learning agent is trained to issue line-specific trip or wait decisions from voltage and current measurements. The study compares raw waveforms, phasor-based features, combined input representations, different observation windows, reward settings, and conservative regularization strengths. Results show that combining waveform and phasor information improves protection performance, while conservative regularization is essential for stable offline learning in this safety-critical setting.