Information on the structure of the conference

Poster-No.

P4-002

Author:

Other authors:

Institution/company:

Machine learning (ML) methods have become an integral part of engineering in recent years. Reinforcement learning (RL) is one popular example of ML and uses a software agent to learn a strategy that maximizes a reward defined by the user. The focus of this research is on enhancing the combined operational strategies of an electrolyzer and a battery energy storage system to maximize economic gains. Consequently, mathematical models developed within the scope of the research project “ZET Reallabor Wunsiedel” for the electrolyzer and the battery of the Wunsiedel energy park are used.
The primary objective of this project revolves around the reduction of electricity procurement costs (EPC), alongside the simultaneous reduction of aging effects resulting from suboptimal operational strategies. Additionally, a pre-defined amount of hydrogen needs to be produced consistently on a weekly basis.
The RL agent is systematically trained to minimize EPCs reacting to the fluctuating electricity prices throughout the year, while ensuring that the desired amount of hydrogen is produced by the electrolyzer. Furthermore, aging is quantified monetarily and considered by the optimization. To evaluate the performance of the optimization with RL, the resulting costs as well as two defined key performance indices, i.e. equivalent operation hours (EOH) of the electrolyzer and full equivalent cycles (FEC) of the battery, are compared to the global optimum found by dynamic programming based on the principle of optimality by Bellman. The reference optimization is carried out with the same constraints as the RL algorithm and specifically finds the best solution for each individual week, while the RL is trained to find the overall optimum for 80 % of all weeks of the year. The evaluation is carried out with the other 20 %.
Additionally, the RL algorithm is trained to account for uncertain electricity prices that lie in the future and therefore to simulate incorrect models of the electricity market and its price developments.
It can be shown that with the RL method results can be achieved that deviate from the global optimum by only 3.5%.