A HYBRID APPROACH TO PHISHING EMAIL DETECTION: LEVERAGING MACHINE LEARNING AND LARGE LANGUAGE MODELS
Date of Award
6-2026
Document Type
Thesis
Degree Name
Master of Science in Information Security
Department
Information Systems and Security
First Advisor
Prof. Khaled Shuaib
Abstract
Phishing attacks have reached a new level of sophistication through the deployment of large language models by attackers. The current AI-generated threats defeat existing detection systems which base their operation on past data. The thesis presents a hybrid system for phishing email detection which combines real email data with synthetic LLM-created samples to enhance traditional machine learning classifiers performance. The study created a hybrid dataset of 20,627 emails by combining 18,631 real messages from the Kaggle Email Classification Dataset with 1,996 synthetic emails. The synthetic emails were generated using four large language models LLaMA-3, Falcon, LLaVA, and Mistral to capture different writing styles found in modern AI-driven phishing. Each email used 121 features which included 21 manually created features that measured text statistics and character patterns and URL details and keyword signals and sentiment scores in addition to 100 TF-IDF unigram and bigram features.
The system trained and tested four classical classifiers Gradient Boosting, Random Forest, Logistic Regression, and SVM (Linear) on both Kaggle-only and hybrid datasets. Performance was assessed using stratified five-fold cross-validation and seven distinct metrics. Gradient Boosting outperformed other conventional classifiers, achieving an accuracy of 95.78%, a recall of 95.55%, and an AUC-ROC of 0.9915 on the hybrid dataset. The Sentence-BERT semantic embedding pipeline, combined with an SVM, yielded the highest overall accuracy of 96.40%, although this method necessitated significantly greater computational resources. Furthermore, the hybrid dataset enhanced accuracy for three of the four classifiers examined.
These findings show the value of adding diverse LLM-generated synthetic data. The system deployed the trained models as a lightweight, stateless Flask REST API. The system can classify emails in under 100 milliseconds. It meets GDPR guidelines by processing data in memory and keeping storage to a minimum. The results show that well-designed hand-crafted features, combined with synthetic data from several LLMs, perform almost as well as Sentence-BERT-based semantic models. The study demonstrates that a system was developed which matches the performance of lightweight transformer classifiers. The study intends to evaluate the method against BERT and RoBERTa in upcoming studies which will maintain practical usage for organizations facing financial constraints.
Arabic Abstract
نهج هجين للكشف عن رسائل البريد الإلكتروني الاحتيالية: الاستفادة من التعلم الآلي ونماذج اللغة الكبيرة
أصبحت هجمات التصيد الاحتيالي أكثر تطوراً مع استخدام المهاجمين للنماذج اللغوية الكبيرة(LLMs)، مما جعل أنظمة الكشف القديمة التي تعتمد فقط على البيانات السابقة أقل فعالية في مواجهة التهديدات الجديدة التي يتم إنشاؤها بواسطة الذكاء الاصطناعي. تقدم هذه الرسالة نظاماً هجيناً لاكتشاف رسائل التصيد الاحتيالي يجمع بين بيانات البريد الإلكتروني الحقيقية وعينات اصطناعية تم إنشاؤها باستخدام النماذج اللغوية الكبيرة، وذلك لجعل مصنفات التعلم الآلي التقليدية أكثر قوة وقدرة على التكيف.
قامت الدراسة ببناء مجموعة بيانات هجينة تتكون من 20,627 رسالة بريد إلكتروني من خلال دمج 18,631 رسالة حقيقية من مجموعة بيانات Kaggle لتصنيف البريد الإلكتروني مع 1,996 رسالة اصطناعية تم إنشاء هذه الرسائل الاصطناعية باستخدام أربعة نماذج لغوية كبيرة وهي LLaMA-3 و Falcon وLLaVA وMistral لتعكس أنماط الكتابة المختلفة التي تظهر في هجمات التصيد الاحتيالي الحديثة المدعومة بالذكاء الاصطناعي. تم تمثيل كل رسالة بريد إلكتروني بواسطة متجه خصائص مكون من 121 بُعداً، يتضمن 21 خاصية مصممة يدوياً مثل إحصائيات النص، وأنماط الأحرف، وتفاصيل الروابط، وإشارات الكلمات المفتاحية، ودرجات تحليل المشاعر، بالإضافة إلى 100 خاصية من نوع TF-IDF للأحادية والثنائية (bigram unigram).
تم تدريب واختبار أربعة مصنفات تقليدية وهي Gradient Boosting وRandom Forest وLogistic Regression وSVM (Linear) باستخدام كل من مجموعة بيانات Kaggle فقط ومجموعة البيانات الهجينة. اعتمد التقييم على التحقق المتقاطع الطبقي بخمس طيات (Stratified Five-Fold Cross-Validation) واستخدام سبعة مقاييس مختلفة للأداء. حقق Gradient Boosting أفضل أداء بين المصنفات التقليدية، حيث وصل إلى دقة 95.78% ومعدل استرجاع 95.55%، وقيمة AUC-ROC بلغت 0.9915 عند استخدام مجموعة البيانات الهجينة. كما أن استخدام تمثيلات دلالية عبر Sentence-BERT مع مصنف SVM حقق أعلى دقة إجمالية بلغت 96.40%، لكنه تطلب قدرة حاسوبية أكبر بكثير. وقد حسنت مجموعة البيانات الهجينة الدقة لثلاثة من أصل أربعة مصنفات، مما يوضح الفائدة الواضحة لإضافة بيانات اصطناعية متنوعة تم إنشاؤها بواسطة النماذج اللغوية الكبيرة.
تم نشر النماذج المدربة كنظام Flask REST API خفيف وعديم الحالة يمكنه تصنيف رسائل البريد الإلكتروني في أقل من 100 مللي ثانية. كما يلتزم النظام بإرشادات GDPR من خلال استخدام المعالجة في الذاكرة وتقليل تخزين البيانات إلى الحد الأدنى. وتشير النتائج إلى أن الخصائص المصممة يدوياً بشكل جيد، عند دمجها مع بيانات اصطناعية من عدة نماذج لغوية كبيرة، يمكن أن تقترب من أداء النماذج المعتمدة على المحولات (transformer-based models)، مع الحفاظ على قابلية التطبيق العملي في الشركات ذات الموارد المحدودة.
Recommended Citation
Biri, Hessa Shamal, "A HYBRID APPROACH TO PHISHING EMAIL DETECTION: LEVERAGING MACHINE LEARNING AND LARGE LANGUAGE MODELS" (2026). Theses. 1496.
https://scholarworks.uaeu.ac.ae/all_theses/1496