Date of Award
4-2026
Document Type
Dissertation
Degree Name
Doctor of Philosophy in Electrical Engineering
Department
Electrical and Communication Engineering
First Advisor
Hussain Shareef
Abstract
Microgrid technology is essential in facilitating the transition to smart energy grids in developed countries and mitigating energy poverty in developing countries, particularly in areas where grid extensions are not feasible. Recently, the concept of networked microgrids (NMGs) has garnered tremendous attention due to the plausibility of interactions among interconnected microgrids leading to power networks that are more resilient, reliable, and stable. However, because each microgrid has diverse distributed generation resources (renewables and controllable generators) and each microgrid operator (MO) has different objectives, coordinated energy management is required to satisfy local and system-wide goals under conditions with significant uncertainty. Existing mathematical formulations for energy management among NMGs are often intractable; when tractable, they tend to be non-scalable in large and dynamic environments due to uncertainties in load demand, electricity prices, solar irradiance, and wind speed.
This study focuses on grid-tied NMGs with multiple electricity retailers (ERs) and multiple microgrids (MGs). The problem is first formulated as a bi-level Stackelberg game in which the upper level maximizes ER profits and the network's available transfer capacity (ATC), while simultaneously minimizing the carbon emissions pertaining to power purchases from the main grid. The lower level on the other hand minimized the operating costs of the MGs. In making the bi-level formulation tractable, the lower-level problem was replaced by its Karush–Kuhn–Tucker (KKT) conditions and eventually appended to the upper-level formulation consequently yielding a tractable single-level mathematical program with equilibrium constraints (MPEC). In solving the resulting NMG energy management problem, a lean multi-agent deep reinforcement learning (L-MADRL)/deep reinforcement learning in the loop (DRL-ITL) framework was developed. The framework combines a multi-agent deep reinforcement learning algorithm with the single-level MPEC obtained after the KKT-based reformulation and adopts a modular architecture in which constraint handling is delegated to this derived analytically tractable optimization layer. This ensures that every action proposed by the DRL agents is feasible with respect to network and market constraints, allowing the agents to focus on learning optimal policies over stochastic variables such as solar irradiance, wind speed, and load demand. As a result, training is more stable and computationally efficient because the agents do not need to navigate complex feasibility regions since technical constraints such as power flow limits, generator capacities, and market rules are embedded in the MPEC. The DRL agents are thus free from constraint enforcement, consequently enhancing the reliability and accuracy of the policy from the agents.
The L-MADRL framework which utilized a deep q network (DQN) as the DRL agent was evaluated across three benchmark systems, namely PJM 5-bus, IEEE 14-bus, and IEEE 30-bus, and benchmarked against a deterministic, risk-neutral stochastic optimization (RNSO), and risk-averse stochastic optimization (RASO) with conditional value at risk (CVaR) approach. In the 5-bus case, the L-MADRL achieved an overall 10.3% cost reduction for the MGs and increased the profits of the ERs by up to 3.7% over the best baseline. In the 14-bus case, total MG costs decreased by 2.6% and ER profits rose by 11.4%. For the 30-bus system, the framework delivered savings of about 2.7% in MG costs and boosted ER profits by up to 9.1%. Notably, L-MADRL enhanced the network's ATC from a starting value of 78 MW to a peak of 94 MW, exceeding the best benchmark by 17.5%. In another example using the DRL-ITL framework strictly on a modified IEEE 14-bus system where a double deep q network (DDQN) was utilized as the DRL agent and carbon emissions were considered, compared to the RASO and a robust optimization approach (RO), the framework reduced the aggregate MG operating costs to €128,408 (which are 9.01% and 5.28% lower than RASO and RO) and increased the total ER profits to €34,466 (10.43% and 15.44% higher), respectively. It further lowered the total CO2 emissions to 305.47 kg, representing a 16.7% and 11.3% reduction relative to the RASO and RO baselines. Across the test systems studied, the DRL-ITL/L-MADRL framework performed better on both economic and technical metrics when compared with other conventional approaches, and at the same time maintained a runtime below 3 seconds, which was also significantly lower than the benchmark methods. This consequently demonstrates the computational efficiency and scalability of the proposed DRL approach.
Arabic Abstract
الإدارة الكفؤة للطاقة في الشبكات المصغّرة المترابطة باستخدام التعلم العميق المعزز متعدد الوكلاء في ظل حالات عدم اليقين
تُعد تقنية الشبكات المصغّرة (Microgrids) عنصرًا أساسيًا في تسهيل الانتقال إلى شبكات طاقة ذكية في الدول المتقدمة، وفي التخفيف من فقر الطاقة في الدول النامية، لا سيما في المناطق التي لا يكون فيها توسيع الشبكة الكهربائية الرئيسية ممكنًا. في الآونة الأخيرة، حظي مفهوم الشبكات المصغّرة المترابطة (Networked Microgrids) باهتمام كبير، نظرًا لإمكانية التفاعل بين الشبكات المصغّرة المترابطة بما يؤدي إلى شبكات طاقة أكثر مرونة وموثوقية واستقرارًا. ومع ذلك، وبسبب امتلاك كل شبكة مصغّرة لمصادر توليد موزعة متنوعة (متجددة وقابلة للتحكم)، إضافة إلى اختلاف أهداف مشغلي الشبكات المصغّرة، فإن إدارة الطاقة المنسقة تصبح ضرورية لتحقيق الأهداف المحلية وعلى مستوى النظام في ظل وجود درجات عالية من عدم اليقين. وغالبًا ما تكون الصيغ الرياضية الحالية لإدارة الطاقة بين الشبكات المصغّرة المترابطة غير قابلة للحل عمليًا، وإن كانت قابلة للحل في بعض الحالات، فإنها تفتقر إلى القابلية للتوسع في البيئات الكبيرة والديناميكية نتيجة عدم اليقين في الطلب على الأحمال، وأسعار الكهرباء، والإشعاع الشمسي، وسرعات الرياح.
تركز هذه الدراسة على الشبكات المصغّرة المترابطة المتصلة بالشبكة الرئيسية والتي تضم عدة بائعي كهرباء (ERs) وعدة شبكات مصغّرة (MGs). تم أولًا صياغة المشكلة كلعبة ستاكلبرغ ثنائية المستوى، حيث يهدف المستوى العلوي إلى تعظيم أرباح بائعي الكهرباء وسعة النقل المتاحة للشبكة (ATC)، مع تقليل انبعاثات الكربون الناتجة عن شراء الطاقة من الشبكة الرئيسية في الوقت نفسه، بينما يهدف المستوى السفلي إلى تقليل تكاليف التشغيل الخاصة بالشبكات المصغّرة. ولجعل الصياغة ثنائية المستوى قابلة للحل، تم استبدال مسألة المستوى السفلي بشروط كاروش–كون–تاكر (KKT) وإلحاقها بصياغة المستوى العلوي، مما أدى في النهاية إلى نموذج رياضي أحادي المستوى قابل للحل من نوع برنامج رياضي بقيود توازنية (MPEC).
ولحل مشكلة إدارة الطاقة في الشبكات المصغّرة المترابطة الناتجة، تم تطوير إطار عمل خفيف الوزن للتعلم العميق المعزز متعدد الوكلاء / (L-MADRL) التعلم العميق المعزز ضمن الحلقة (DRL-ITL). يجمع هذا الإطار بين خوارزمية تعلم عميق معزز متعدد الوكلاء والنموذج الأحادي المستوى (MPEC) الناتج عن إعادة الصياغة المعتمدة على شروط KKT، ويعتمد بنية معيارية يتم فيها تفويض معالجة القيود إلى طبقة تحسين تحليلية قابلة للحل. ويضمن ذلك أن تكون جميع الأفعال التي يقترحها وكلاء التعلم العميق المعزز متوافقة مع قيود الشبكة والسوق، مما يسمح للوكلاء بالتركيز على تعلم السياسات المثلى في ظل المتغيرات العشوائية مثل الإشعاع الشمسي وسرعة الرياح والطلب على الأحمال. ونتيجة لذلك، تصبح عملية التدريب أكثر استقرارًا وكفاءة حسابية، حيث لا يحتاج الوكلاء إلى التعامل مع مناطق جدوى معقدة، إذ إن القيود التقنية مثل حدود تدفق القدرة، وسعات المولدات، وقواعد السوق تكون
مدمجة ضمن نموذج الـ MPEC. وبالتالي، يتحرر وكلاء التعلم العميق المعزز من فرض القيود مباشرة، مما يعزز موثوقية ودقة السياسات المتعلمة.
تم تقييم إطار L-MADRL باستخدام شبكة Q العميقة (DQN) كوكلاء تعلم عميق معزز عبر ثلاثة أنظمة مرجعية، وهي نظام PJM ذي خمس عقد، ونظام IEEE ذي 14 عقدة، ونظام IEEE ذي 30 عقدة، كما تمت مقارنته بأساليب التحسين الحتمي المحايد للمخاطر، والتحسين العشوائي المتحفظ للمخاطر باستخدام مقياس القيمة المعرضة للخطر الشرطية (CVaR). في حالة نظام الخمس عقد، حقق إطار L-MADRL خفضًا إجماليًا في تكاليف تشغيل الشبكات المصغّرة بنسبة 10.3% وزيادة في أرباح بائعي الكهرباء تصل إلى 3.7% مقارنة بأفضل نموذج مرجعي. وفي نظام الأربع عشرة عقدة، انخفضت التكاليف الإجمالية للشبكات المصغّرة بنسبة 2.6% وارتفعت أرباح بائعي الكهرباء بنسبة 11.4%. أما في نظام الثلاثين عقدة، فقد حقق الإطار وفورات بنحو 2.7% في تكاليف الشبكات المصغّرة وزيادة في أرباح بائعي الكهرباء تصل إلى %9.1
والجدير بالذكر أن إطار L-MADRL حسّن سعة النقل المتاحة للشبكة من قيمة ابتدائية قدرها 78 ميغاواط إلى قيمة قصوى بلغت 94 ميغاواط، متجاوزًا أفضل نموذج مرجعي بنسبة 17.5%. وفي مثال آخر باستخدام إطار DRL-ITL على نظام IEEE معدل ذي 14 عقدة مع استخدام شبكة Q العميقة المزدوجة (DDQN) وأخذ انبعاثات الكربون بعين الاعتبار، وبالمقارنة مع أسلوبي التحسين العشوائي المتحفظ للمخاطر والتحسين القوي، نجح الإطار في خفض التكاليف الإجمالية لتشغيل الشبكات المصغّرة إلى 128,408 يورو (أي أقل بنسبة 9.01% و5.28% على التوالي)، وزيادة أرباح بائعي الكهرباء إلى 34,466 يورو (بزيادة قدرها 10.43% و15.44%). كما خفّض الإطار إجمالي انبعاثات ثاني أكسيد الكربون إلى 305.47 كغ، أي بانخفاض نسبته 16.7% و11.3% مقارنة بالأساليب المرجعية. وعبر جميع الأنظمة المدروسة، تفوق إطار L-MADRL / DRL-ITL باستمرار على الطرق التقليدية من حيث المؤشرات الاقتصادية والتقنية، مع الحفاظ على زمن تنفيذ يقل عن 3 ثوانٍ، مما يبرهن على قابليته العالية للتوسع وكفاءته الحسابية.
Recommended Citation
Chukwuyem, Ayodele Benjamin, "EFFICIENT ENERGY MANAGEMENT IN NETWORKED MICROGRIDS USING MULTI-AGENT DEEP REINFORCEMENT LEARNING IN THE PRESENCE OF UNCERTAINTIES" (2026). Dissertations. 417.
https://scholarworks.uaeu.ac.ae/all_dissertations/417