BERTPHIURL: A Teacher-Student Learning Approach Using DistilRoBERTa and RoBERTa for Detecting Phishing Cyber URLs

Authors

DOI:

https://doi.org/10.62411/faith.3048-3719-71

Keywords:

BERT, DistilRoBERTa, Natural Language Processing, NLP phishing detection, Phishing URL Detection, RoBERTa, Teacher-Student Learning, Transformer

Abstract

Phishing is a fraudulent activity wherein an attacker impersonates a trusted individual or organization to acquire sensitive information from an online user. Phishing websites have become a major cyber-security issue in the contemporary digital landscape. As online activities expand in e-commerce, banking, and social media, the hazards presented by these fraudulent websites have intensified. Deep learning-based Natural Language Processing (NLP) approaches offer an effective solution for detecting phishing URLs. However, deploying large models like BERT or RoBERTa for real-time detection poses computational challenges. This study proposes BERTPHIURL, a Teacher-Student Learning framework that leverages RoBERTa as the Teacher model and DistilRoBERTa as the Student model to improve phishing detection efficiency while reducing computational overhead. By applying knowledge distillation, the Student model learns from the Teacher, preserving high detection accuracy with significantly lower resource consumption. The proposed approach effectively captures contextual relevance and local features in malicious URL detection tasks. The experiments were conducted on a dataset exceeding 50,000 URLs to evaluate performance. Results indicate that BERTPHIURL achieves a 94.22% accuracy, outperforming existing phishing detection methods while maintaining efficiency suitable for real-time applications.

Downloads

Download data is not yet available.

Author Biographies

Payman Hussein Hussan, Al-Furat Al-Awsat Technical University

Department of Computer Networks and Software Techniques, Babylon Technical Institute, Al-Furat  Al-Awsat Technical University, Kufa, Iraq

Syefy Mohammed Mangj, Al-Furat  Al-Awsat Technical University

Department of Computer Networks and Software Techniques, Babylon Technical Institute, Al-Furat Al-Awsat Technical University, Kufa, Iraq

References

M. Somesha, A. R. Pais, R. S. Rao, and V. S. Rathour, “Efficient deep learning techniques for the detection of phishing websites,” Sādhanā, vol. 45, no. 1, p. 165, Dec. 2020, doi: 10.1007/s12046-020-01392-4.

A. Safi and S. Singh, “A systematic literature review on phishing website detection techniques,” J. King Saud Univ. - Comput. Inf. Sci., vol. 35, no. 2, pp. 590–611, Feb. 2023, doi: 10.1016/j.jksuci.2023.01.004.

W. Sarasjati, S. Rustad, Purwanto, H. A. Santoso, and D. R. I. M. Setiadi, “Phishing Detection Using Random Forest-Based Weighted Bootstrap Sampling and LASSO+ Feature Selection,” Int. J. Saf. Secur. Eng., vol. 14, no. 6, pp. 1783–1794, Dec. 2024, doi: 10.18280/ijsse.140613.

Y. Said, A. A. Alsheikhy, H. Lahza, and T. Shawly, “Detecting phishing websites through improving convolutional neural networks with Self-Attention mechanism,” Ain Shams Eng. J., vol. 15, no. 4, p. 102643, Apr. 2024, doi: 10.1016/j.asej.2024.102643.

S. Nagarajan, “An Investigation of AI-Enabled Comprehensive Survey of Phishing Attacks detection techniques,” Int. J. Innov. Res. Eng., vol. 5, no. 1, pp. 59–63, Sep. 2024, doi: 10.59256/ijire.20220301001.

D. R. I. M. Setiadi, S. Widiono, A. N. Safriandono, and S. Budi, “Phishing Website Detection Using Bidirectional Gated Recurrent Unit Model and Feature Selection,” J. Futur. Artif. Intell. Technol., vol. 1, no. 2, pp. 75–83, Jul. 2024, doi: 10.62411/faith.2024-15.

F. Heiding, B. Schneier, A. Vishwanath, J. Bernstein, and P. S. Park, “Devising and Detecting Phishing Emails Using Large Language Models,” IEEE Access, vol. 12, pp. 42131–42146, 2024, doi: 10.1109/ACCESS.2024.3375882.

G. Desolda, L. S. Ferro, A. Marrella, T. Catarci, and M. F. Costabile, “Human Factors in Phishing Attacks: A Systematic Literature Review,” ACM Comput. Surv., vol. 54, no. 8, pp. 1–35, Nov. 2022, doi: 10.1145/3469886.

T. Xu and P. Rajivan, “Determining psycholinguistic features of deception in phishing messages,” Inf. Comput. Secur., vol. 31, no. 2, pp. 199–220, May 2023, doi: 10.1108/ICS-11-2021-0185.

C.-S. Lin, C.-N. Tsai, J.-S. Jwo, C.-H. Lee, and X. Wang, “Heterogeneous Student Knowledge Distillation From BERT Using a Lightweight Ensemble Framework,” IEEE Access, vol. 12, pp. 33079–33088, 2024, doi: 10.1109/ACCESS.2024.3372568.

L. Wang and K.-J. Yoon, “Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 6, pp. 3048–3068, Jun. 2022, doi: 10.1109/TPAMI.2021.3055564.

J. K. S and A. B, “Phishing URL detection by leveraging RoBERTa for feature extraction and LSTM for classification,” in 2023 Second International Conference on Augmented Intelligence and Sustainable Systems (ICAISS), Aug. 2023, pp. 972–977. doi: 10.1109/ICAISS58487.2023.10250684.

M. Songailaitė, E. Kankevičiūtė, B. Zhyhun, and J. Mandravickaitė, “BERT-Based Models for Phishing Detection,” in CEUR Workshop Proceedings, 2023, vol. 3575, pp. 34–44.

K. Zhang, C. Zhang, S. Li, D. Zeng, and S. Ge, “Student Network Learning via Evolutionary Knowledge Distillation,” ArXiv. Mar. 22, 2021. [Online]. Available: http://arxiv.org/abs/2103.01381

A. Saleem Raja, R. Vinodini, and A. Kavitha, “Lexical features based malicious URL detection using machine learning techniques,” Mater. Today Proc., vol. 47, pp. 163–166, 2021, doi: 10.1016/j.matpr.2021.04.041.

L. Tang and Q. H. Mahmoud, “A Deep Learning-Based Framework for Phishing Website Detection,” IEEE Access, vol. 10, pp. 1509–1521, 2022, doi: 10.1109/ACCESS.2021.3137636.

A. Karim, M. Shahroz, K. Mustofa, S. B. Belhaouari, and S. R. K. Joga, “Phishing Detection System Through Hybrid Machine Learning Based on URL,” IEEE Access, vol. 11, pp. 36805–36822, 2023, doi: 10.1109/ACCESS.2023.3252366.

M. K. Prabakaran, P. Meenakshi Sundaram, and A. D. Chandrasekar, “An enhanced deep learning‐based phishing detection mechanism to effectively identify malicious URLs using variational autoencoders,” IET Inf. Secur., vol. 17, no. 3, pp. 423–440, May 2023, doi: 10.1049/ise2.12106.

M. Amanullah, V. Selvakumar, A. Jyot, N. Purohit, S. S, and M. Fahlevi, “CNN based Prediction Analysis for Web Phishing Prevention,” in 2022 International Conference on Edge Computing and Applications (ICECAA), Oct. 2022, pp. 1–7. doi: 10.1109/ICECAA55415.2022.9936112.

O. K. Sahingoz, E. BUBEr, and E. Kugu, “DEPHIDES: Deep Learning Based Phishing Detection System,” IEEE Access, vol. 12, no. December, pp. 8052–8070, 2024, doi: 10.1109/ACCESS.2024.3352629.

P. Vaitkevicius and V. Marcinkevicius, “Comparison of Classification Algorithms for Detection of Phishing Websites,” Informatica, vol. 31, no. 1, pp. 143–160, Mar. 2020, doi: 10.15388/20-INFOR404.

O. G. Nabila, H. R. Wicaksono, Girinoto, R. N. Yasa, and H. Setiawan, “Benchmarking Model URL Features and Image Based for Phishing URL Detection,” in 2023 International Conference on Informatics, Multimedia, Cyber and Informations System (ICIMCIS), Nov. 2023, pp. 177–182. doi: 10.1109/ICIMCIS60089.2023.10349059.

Z. Alshingiti, R. Alaqel, J. Al-Muhtadi, Q. E. U. Haq, K. Saleem, and M. H. Faheem, “A Deep Learning-Based Phishing Detection System Using CNN, LSTM, and LSTM-CNN,” Electronics, vol. 12, no. 1, p. 232, Jan. 2023, doi: 10.3390/electronics12010232.

T. Tiwari, “Phishing Site URLs,” Kaggle.com, 2017. https://www.kaggle.com/datasets/taruntiwarihp/phishing-site-urls

Y. Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” ArXiv. Jul. 26, 2019.

J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2019, pp. 4171–4186.

V. Sanh, L. Debut, J. Chaumond, and T. Wolf, “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,” ArXiv. pp. 2–6, Oct. 02, 2019. [Online]. Available: http://arxiv.org/abs/1910.01108

A. Çetin and S. Öztürk, “Comprehensive Exploration of Ensemble Machine Learning Techniques for IoT Cybersecurity Across Multi-Class and Binary Classification Tasks,” J. Futur. Artif. Intell. Technol., vol. 1, no. 4, pp. 371–384, Feb. 2025, doi: 10.62411/faith.3048-3719-51.

M. Korkmaz, O. K. Sahingoz, and B. Diri, “Detection of Phishing Websites by Using Machine Learning-Based URL Analysis,” in 2020 11th International Conference on Computing, Communication and Networking Technologies (ICCCNT), Jul. 2020, pp. 1–7. doi: 10.1109/ICCCNT49239.2020.9225561.

A. Subasi and E. Kremic, “Comparison of Adaboost with MultiBoosting for Phishing Website Detection,” Procedia Comput. Sci., vol. 168, pp. 272–278, 2020, doi: 10.1016/j.procs.2020.02.251.

M. S. Hossen, A. H. Jony, T. Tabassum, M. T. Islam, M. M. Rahman, and T. Khatun, “Hotel review analysis for the prediction of business using deep learning approach,” in 2021 International Conference on Artificial Intelligence and Smart Systems (ICAIS), Mar. 2021, pp. 1489–1494. doi: 10.1109/ICAIS50930.2021.9395757.

P. K. Jain, V. Saravanan, and R. Pamula, “A Hybrid CNN-LSTM: A Deep Learning Approach for Consumer Sentiment Analysis Using Qualitative User-Generated Contents,” ACM Trans. Asian Low-Resource Lang. Inf. Process., vol. 20, no. 5, pp. 1–15, Sep. 2021, doi: 10.1145/3457206.

S. He, B. Li, H. Peng, J. Xin, and E. Zhang, “An Effective Cost-Sensitive XGBoost Method for Malicious URLs Detection in Imbalanced Dataset,” IEEE Access, vol. 9, pp. 93089–93096, 2021, doi: 10.1109/ACCESS.2021.3093094.

K. L. Tan, C. P. Lee, K. S. M. Anbananthen, and K. M. Lim, “RoBERTa-LSTM: A Hybrid Model for Sentiment Analysis With Transformer and Recurrent Neural Network,” IEEE Access, vol. 10, pp. 21517–21525, 2022, doi: 10.1109/ACCESS.2022.3152828.

K. Omari, “Comparative Study of Machine Learning Algorithms for Phishing Website Detection,” Int. J. Adv. Comput. Sci. Appl., vol. 14, no. 9, pp. 417–425, 2023, doi: 10.14569/IJACSA.2023.0140945.

P. Dhanavanthini and S. S. Chakkravarthy, “Phish-armour: phishing detection using deep recurrent neural networks,” Soft Comput., Mar. 2023, doi: 10.1007/s00500-023-07962-y.

S. Raminedi, T. N. Pandey, V. A. Woonna, S. C. Mascarenhas, and A. Bharani, “Classification of Phishing Websites using Machine Learning Models,” in 2023 3rd International conference on Artificial Intelligence and Signal Processing (AISP), Mar. 2023, pp. 1–5. doi: 10.1109/AISP57993.2023.10134944.

E. Benavides-Astudillo, W. Fuertes, S. Sanchez-Gordon, G. Rodriguez-Galan, V. Martínez-Cepeda, and D. Nuñez-Agurto, “Comparative Study of Deep Learning Algorithms in the Detection of Phishing Attacks Based on HTML and Text Obtained from Web Pages,” in Applied Technologies, M. Botto-Tobar, M. Z. Vizuete, S. M. León, P. Torres-Carrión, and B. Durakovic, Eds. Springer Nature Switzerland, 2023, pp. 386–398. doi: 10.1007/978-3-031-24985-3_28.

T. Ige, C. Kiekintveld, A. Piplai, A. Waggler, O. Kolade, and B. H. Matti, “An investigation into the performances of the Current state-of-the-art Naive Bayes, Non-Bayesian and Deep Learning Based Classifier for Phishing Detection: A Survey,” ArXiv. Nov. 24, 2024. [Online]. Available: http://arxiv.org/abs/2411.16751

Downloads

Published

2025-02-24

How to Cite

[1]
P. H. Hussan and S. M. Mangj, “BERTPHIURL: A Teacher-Student Learning Approach Using DistilRoBERTa and RoBERTa for Detecting Phishing Cyber URLs”, J. Fut. Artif. Intell. Tech., vol. 1, no. 4, pp. 417–428, Feb. 2025.

Issue

Section

Articles

Similar Articles

1 2 3 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.