高级检索

计及动态可行域约束的电动汽车集群深度强化学习调频控制

Deep reinforcement learning frequency regulation control of electric vehicle clusters considering dynamic feasible region constraints

  • 摘要: 高比例可再生能源并网对电力系统频率稳定构成严峻挑战,电动汽车(electric vehicle,EV)作为移动式储能资源参与调频极具潜力。为解决海量EV在满足用户出行与电池安全约束下高效跟踪二次调频指令的实时控制难题,提出了一种计及动态可行域约束的深度强化学习(deep reinforcement learning,DRL)双层调频控制策略。首先,构建从单体到集群的荷电状态(state of charge,SOC)可行域动态聚合模型,并建立调频容量优化问题,将用户行为与物理约束量化为集群实时的安全调节边界。其次,设计“集中决策-分层执行”架构,中央层采用深度确定性策略梯度(deep deterministic policy gradient,DDPG)算法,在连续动作空间中学习各功能区间的最优功率分配;功能区层则基于SOC规则将指令分解至单车。最后,引入动作投影层,将DDPG输出的原始动作实时映射至可行域内,以确保控制过程的安全性与可行性。仿真结果表明,所提策略在满足用户出行需求与电池安全约束的同时,实现了平均绝对误差0.084 MW和均方根误差0.129 MW的高精度调频跟踪,与无约束的强化学习算法相比将跟踪误差降低了97.8%,与模型预测控制(model predictive control,MPC)相比将计算效率提高了28倍,同时有效保障区域间调频资源利用的公平性。

     

    Abstract: The integration of high-penetration renewable energy sources poses significant challenges to power system frequency stability. As mobile energy storage resources, electric vehicles (EVs) have great potential to participate in frequency regulation. To address the real-time control difficulty that massive EVs efficiently track secondary frequency regulation commands while satisfying user travel demands and battery safety constraints, this paper proposes a deep reinforcement learning (DRL)-based two-layer frequency regulation control strategy considering dynamic feasible-region constraints. Firstly, a dynamic aggregation model for the state-of-charge (SOC) feasible regions is established from individual EVs to the vehicle clusters, and a frequency regulation capacity optimization problem is formulated to quantify user behaviors and physical constraints as real-time safe regulation boundaries of the EV clusters. Secondly, a "centralized decision-making - hierarchical execution" framework is developed. At the central layer, the deep deterministic policy gradient (DDPG) algorithm is employed to learn the optimal power allocation among different functional zones within the continuous action space. At the functional zone layer, commands are decomposed to individual EVs based on SOC rules. Finally, an action projection layer is introduced to map the raw actions output by DDPG into the feasible region in real time, so as to ensure the safety and feasibility of the control process. Simulation results demonstrate that the proposed strategy achieves high-precision frequency regulation tracking while satisfying user travel demands and battery safety constraints, with a mean absolute error (MAE) of 0.084 MW and a root mean square error (RMSE) of 0.129 MW. Compared with the unconstrained reinforcement learning algorithm, the proposed method reduces the tracking error by 97.8% and improves the computational efficiency by 28 times in contrast to model predictive control (MPC). Meanwhile, the fairness of frequency regulation resources utilization among different regions is effectively ensured.

     

/

返回文章
返回