Date of Award

11-2025

Document Type

Dissertation

Degree Name

Doctor of Philosophy in Informatics and Computing

Department

Computer Science and Software Engineering

First Advisor

Dr. Farag M. Sallabi

Second Advisor

Prof. Mohamed Adel Serhani

Abstract

Federated Learning (FL) emerged as a significant advancement in the field of Artificial Intelligence (AI), enabling collaborative model training across distributed devices while maintaining data privacy. As the importance of FL and its application in various areas increased, addressing trustworthiness issues in its various aspects became crucial. In the FL process, clients contribute updates computed on their local datasets, which the server aggregates to iteratively refine the global model. However, not all client data may be relevant to the learning objective, and incorporating updates from irrelevant data can harm the model's performance. The selection of training samples significantly impacts model performance, as datasets with errors, skewed distributions, or low diversity can lead to inaccurate and unstable models. To address these issues, a data quality evaluation model has been introduced to assess the quality of datasets in FL systems. This model dynamically selects high-quality data samples for FL training by utilizing intrinsic and contextual data quality dimensions. Additionally, an importance-based interpretable feature selection model and a data quality-based dynamic client selection model employing Nash equilibrium and joint differential privacy (DP) have been designed. This approach encourages clients with high-quality data to participate in FL training, thereby improving the overall quality of the training process. Building upon this, it has been identified that FL faces a major challenge of high communication overhead due to frequent model updates between clients and the central server. Existing solutions, such as reducing communication frequency or compressing gradients, struggle to balance efficiency and accuracy, particularly in resource-constrained edge environments. There is a growing need for trustworthy, communication-efficient FL solutions that ensure privacy, fairness, and security. To address this, a lightweight novel Hierarchical FL (HFL) framework is proposed that integrates adaptive model pruning, quantization, and communication frequency optimization. First, a joint model pruning, and quantization approach is introduced that dynamically adjusts pruning ratios and quantization levels, reducing communication costs while maintaining high accuracy. Second, a fairness-aware Stackelberg game-based communication frequency optimization model has been developed, where clients, edge servers, and the central server collaboratively determine optimal update frequencies to balance overhead and convergence speed. Third, privacy protection is enhanced using Selective Homomorphic Encryption (SHE), and a verifiable model trust assessment is introduced to ensure secure participation of edge devices. However, communication efficiency and quality selection alone cannot guarantee trustworthy FL, as the security of both client models and aggregation processes is very important. FL enables distributed training while preserving data privacy, but it is vulnerable to poisoning attacks due to heterogeneous, non-IID client data and limited participation. In addition, malicious clients, either fake or compromised, can launch targeted or untargeted attacks such as backdoors, label-flipping, and adaptive model poisoning, leading to corrupted global models. To address this, a multi-layered defense is proposed that combines game-theoretic aggregation, incentive-aware client regularization, and model-side verification with important parameter selection and SHE. Extensive experiments have been conducted to validate the proposed approaches, demonstrating their effectiveness in improving data-quality-driven data sample and client selection, optimizing communication efficiency, and defending against adversarial poisoning in the Trustworthy FL paradigm. We are hopeful that the results and discussions in this dissertation will help researchers to further improve trustworthy and secure FL systems.

Arabic Abstract


إطار التعلم الموحد الموثوق لتوفير ذكاء اصطناعي موزع آمن، فعال، ومراعي لجودة البيانات

يعد التعلم الموحد تطوراً بارزاً في مجال الذكاء الاصطناعي إذ يتيح تدريب نماذج الذكاء الاصطناعي بشكل تعاوني عبر أجهزة موزعة مع الحفاظ على خصوصية البيانات. ومع ازدياد أهمية التعلم الموحد وتطبيقاته في مجالات مختلفة، أصبح من الضروري معالجة قضايا الموثوقية في جوانب متعددة. في عملية التعلم الموحد (FL)، تقوم الأجهزة المشاركة (العملاء) بحساب تحديثات نموذج الذكاء الاصطناعي استناداً إلى مجموعات البيانات المحلية الخاصة بهم، ثم يقوم الخادم بتجميع هذه التحديثات بشكل تكراري لتحسين النموذج العالمي. ومع ذلك، قد لا تكون جميع بيانات هذه الأجهزة ذات صلة بالهدف التعليمي، وإدراج تحديثات من بيانات غير ذات صلة قد يؤثر سلباً على أداء النموذج. إن الاختيار المناسب لعينات التدريب يؤثر بشكل كبير على أداء نموذج الذكاء الاصطناعي، إذ إن مجموعات البيانات التي تحتوي على أخطاء أو توزيعات غير متوازنة أو تنوع منخفض قد تؤدي إلى نماذج غير دقيقة وغير مستقرة. ولمعالجة هذه المشكلات، تم اقتراح نموذج لتقييم جودة البيانات في أنظمة التعلم الموحد (FL). يقوم هذا النموذج باختيار عينات بيانات عالية الجودة بشكل ديناميكي لتدريب FL من خلال الاستفادة من أبعاد الجودة الجوهرية والسياقية للبيانات. بالإضافة إلى ذلك، تم تصميم نموذج قابل للتفسير لاختيار السمات بناءً على الأهمية، وكذلك نموذج لاختيار العملاء بشكل ديناميكي وفقاً لجودة البيانات باستخدام توازن ناش (Nash Equilibrium) والخصوصية التفاضلية المشتركة (DP). يهدف هذا النهج إلى تشجيع العملاء الذين يمتلكون بيانات عالية الجودة على المشاركة في تدريب FL، مما يحسن من الجودة العامة لعملية التدريب. بناءً على ذلك، يواجه التعلم الموحد (FL) تحدياً رئيسياً يتمثل في ارتفاع تكلفة الاتصال الناتجة عن التحديثات المتكررة للنموذج بين الأجهزة والخادم المركزي. الحلول الحالية، مثل تقليل عدد مرات الاتصال أو ضغط التدرجات، لا تنجح غالباً في الموازنة بين الكفاءة والدقة، خاصةً في بيئات الحافة محدودة الموارد. لذلك، هناك حاجة متزايدة إلى حلول موثوقة وفعالة في الاتصال، مع ضمان الخصوصية والعدالة والأمان. وللتغلب على هذا التحدي، يقترح إطار عمل جديد وخفيف للتعلم الموحد الهرمي (HFL) يدمج بين تقليم النموذج بشكل تكيفي (Adaptive Pruning) والكمية (Quantization)، وتحسين معدل الاتصال. أولاً، يقدم أسلوباً مشتركاً للتقليم (Pruning) والكمية (Quantization) يعمل على ضبط نسب التقليم ومستويات الكمية بشكل ديناميكي، مما يقلل من تكاليف الاتصال مع الحفاظ على دقة عالية. ثانياً، تم تطوير نموذج لتحسين معدل الاتصال على لعبة ستاكلبرغ (Stackelberg Game)، حيث يتعاون العملاء والخادم المركزي في تحديد المعدل الأمثل الذي يحقق التوازن بين تقليل تكلفة الاتصال وسرعة استقرار النموذج. ثالثاً، تم تعزيز حماية الخصوصية باستخدام التشفير المتماثل الانتقائي (Selective Homomorphic Encryption (SHE))، كما تم إدخال آلية للتحقق من موثوقية النموذج لضمان أمان أجهزة الحافة. ومع ذلك، فإن تحسين كفاءة الاتصال وجودة البيانات وحده لا يكفي لضمان موثوقية التعلم الموحد (FL)، إذ تظل حماية النماذج لدى العملاء وتأمين عملية التجميع عند الخادم أمراً بالغ الأهمية. يمكن للتعلم الموحد (FL) من تنفيذ تدريب موزع مع الحفاظ على

 خصوصية البيانات، لكنه يبقى عرضة لهجمات التسميم (Poisoning Attacks) نتيجة لتباين البيانات بين العملاء (Heterogeneous) وعدم تجانسها أو استقلاليتها (Non-IID)، بالإضافة إلى محدودية مشاركة العملاء في عملية التدريب. علاوة على ذلك، يمكن للعملاء الخبيثين، سواء كانوا مهاجمين أو مخترقين، شن هجمات مستهدفة أو غير مستهدفة مثل زرع الأبواب الخلفية (backdoors)، أو قلب التسميات (Label Flipping)، أو تسميم النماذج التكيفي (Adaptive Model Poisoning)، مما يؤدي إلى إفساد النماذج العالمية. ولمعالجة ذلك، يُقترح نظام دفاع متعدد الطبقات يجمع بين تجميع قائم على نظرية الألعاب (Game Theoretic Aggregation)، وتنظيم مشاركة العملاء مع مراعاة الحوافز (Incentive Aware Client Regularization)، والتحقق الجانبي من النموذج عبر اختيار المعاملات المهمة، والتشفير المتماثل الانتقائي (SHE). وقد أُجريت تجارب مكثفة للتحقق من فعالية هذه الأساليب المقترحة، وأظهرت النتائج فعاليتها في تحسين جودة البيانات لاختيار العينات والعملاء، وتحسين كفاءة الاتصال، والتصدي لهجمات التسميم العدائية ضمن إطار التعلم الموحد الموثوق (Trustworthy FL). ونحن نأمل أن تسهم النتائج والمناقشات الواردة في هذه الرسالة في مساعدة الباحثين على تحسين أنظمة التعلم الموحد لتصبح أكثر موثوقية وأماناً.

Share

COinS