Low-latency Malicious URL Detection based on Lexical and Statistical Features

Authors

  • Siyu Huang Department of Computing, The Hong Kong Polytechnic University, Hong Kong, China

DOI:

https://doi.org/10.54097/6g69az84

Keywords:

Malicious URL Detection, Feature Engineering, Ensemble Learning, Random Forest, Lightweight Security

Abstract

Malicious URLs remain a persistent threat because attackers can generate short-lived links faster than static blacklists can be updated. This paper presents a lightweight stateless detector that uses only URL lexical and statistical features, without webpage downloading, DNS probing, WHOIS lookup, or HTML rendering. Experiments on approximately 50,000 cleaned URLs assembled from PhishTank, CIC-Phishing2021/CIC-Bell-DNS2021, and Tranco show that the proposed LightGBM model achieves 98.64% Accuracy, an F1-Score of 0.9863, and 0.4165 ms feature-extraction-plus-inference latency on an Apple M4 device with 16 GB RAM. Under a stress test with 5% malicious traffic, it maintains 99.59% Precision, 96.60% Recall, 0.02% FPR, and PR-AUC of 0.9793. These results suggest that URL-only mathematical and linguistic signals can support low-latency browser, mobile-gateway, and edge-security screening, while domain-granularity and source-composition effects still require time-shifted and source-balanced deployment validation.

Downloads

Download data is not yet available.

References

[1] Antonakakis, M., Perdisci, R., Dagon, D., Lee, W., & Feamster, N. (2012). From throw-away traffic to bots: Detecting the rise of DGA-based malware. In Proceedings of the 21st USENIX Security Symposium (pp. 491–506). Bellevue.

[2] Sahoo, D., Liu, C., & Hoi, S. C. H. (2017). Malicious URL detection using machine learning: A survey. arXiv preprint arXiv:1701.07179. https://doi.org/10.48550/arXiv.1701.07179.

[3] Ma, J., Saul, L. K., Savage, S., & Voelker, G. M. (2009). Beyond blacklists: Learning to detect malicious web sites from suspicious URLs. In Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 1245–1254). Paris. https://doi.org/10.1145/ 1557019. 1557153. DOI: https://doi.org/10.1145/1557019.1557153

[4] Garera, S., Provos, N., Chew, M., & Rubin, A. D. (2007). A framework for detection and measurement of phishing attacks. In Proceedings of the 2007 ACM Workshop on Recurring Malcode (pp. 1–8). Alexandria. https://doi.org/10.1145/ 1313848. 1313850. DOI: https://doi.org/10.1145/1314389.1314391

[5] Whittaker, C., Ryner, B., & Nazif, M. (2010). Large-scale automatic classification of phishing pages. In Proceedings of the Network and Distributed System Security Symposium (pp. 1–14). San Diego.

[6] Marchal, S., Francois, J., State, R., & Engel, T. (2014). PhishStorm: Detecting phishing with streaming analytics. IEEE Transactions on Network and Service Management, 11(4), 458–471. https://doi.org/10.1109/TNSM.2014.2377295. DOI: https://doi.org/10.1109/TNSM.2014.2377295

[7] Xiang, G., Hong, J., Rose, C. P., & Cranor, L. (2011). CANTINA+: A feature-rich machine learning framework for detecting phishing web sites. ACM Transactions on Information and System Security, 14(2), 21. https://doi.org/10. 1145/ 1987346.1987348. DOI: https://doi.org/10.1145/2019599.2019606

[8] Blum, A., Wardman, B., Solorio, T., & Warner, G. (2010). Lexical feature based phishing URL detection using online learning. In Proceedings of the 3rd ACM Workshop on Artificial Intelligence and Security (pp. 54–60). Chicago. https:// doi.org/10.1145/1866642.1866651. DOI: https://doi.org/10.1145/1866423.1866434

[9] Kan, M.-Y., & Thi, H. O. N. (2005). Fast webpage classification using URL features. In Proceedings of the 14th ACM International Conference on Information and Knowledge Management (pp. 325–326). Bremen. https://doi.org/10.1145/ 1099554. 1099631. DOI: https://doi.org/10.1145/1099554.1099649

[10] Shirazi, H., Bezawada, B., & Ray, I. (2018). Kn0w Thy Doma1n Name: Unbiased phishing detection using domain name based features. In Proceedings of the 23rd ACM Symposium on Access Control Models and Technologies (pp. 69–75). Indianapolis. https://doi.org/10.1145/ 3205977. 32059 81. DOI: https://doi.org/10.1145/3205977.3205992

[11] Sahingoz, O. K., Buber, E., Demir, O., & Diri, B. (2019). Machine learning based phishing detection from URLs. Expert Systems with Applications, 117, 345–357. https://doi.org/10. 1016/ j. eswa. 2018.09.029. DOI: https://doi.org/10.1016/j.eswa.2018.09.029

[12] Shannon, C. E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27, 379–423, 623–656. https://doi.org/10.1002/j.1538-7305.1948.tb01338.x. DOI: https://doi.org/10.1002/j.1538-7305.1948.tb00917.x

[13] Kumi, S., Lim, C., & Lee, S. G. (2021). Malicious URL detection based on associative classification. Entropy, 23(2), 182. https://doi.org/10.3390/e23020182. DOI: https://doi.org/10.3390/e23020182

[14] Le, H., Pham, Q., Sahoo, D., & Hoi, S. C. H. (2018). URLNet: Learning a URL representation with deep learning for malicious URL detection. arXiv preprint arXiv:1802.03162. https:// doi.org/10.48550/arXiv.1802.03162.

[15] Vinayakumar, R., Soman, K. P., & Poornachandran, P. (2018). Evaluating deep learning approaches to characterize and classify malicious URLs. Journal of Intelligent & Fuzzy Systems, 34(3), 1333–1343. https://doi.org/10.3233/JIFS-171 439. DOI: https://doi.org/10.3233/JIFS-169429

[16] Saxe, J., & Berlin, K. (2017). eXpose: A character-level convolutional neural network with embeddings for detecting malicious URLs, file paths and registry keys. arXiv preprint arXiv:1702.08568. https://doi.org/10.48550/arXiv.1702.08568.

[17] Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324. DOI: https://doi.org/10.1023/A:1010933404324

[18] Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T.-Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems 30. Long Beach.

[19] Alqahtani, H., & Abu-Khadrah, A. (2024). Enhance the accuracy of malicious uniform resource locator detection based on effective machine learning approach. Bulletin of Electrical Engineering and Informatics, 13(6), 4422–4429. https://doi. org/ 10. 11591/eei. v13i6.7121. DOI: https://doi.org/10.11591/eei.v13i6.7797

[20] Butnaru, A., Mylonas, A., & Pitropakis, N. (2021). Towards lightweight URL-based phishing detection. Future Internet, 13(7), 154. https://doi.org/10.3390/fi13070154. DOI: https://doi.org/10.3390/fi13060154

[21] Cisco Talos. (2026). PhishTank developer information. https:// phishtank. org/developer_info.php. Accessed 30 May 2026.

[22] Mahdavifar, S., Maleki, N., Lashkari, A. H., Broda, M., & Razavi, A. H. (2021). Classifying malicious domains using DNS traffic analysis. In Proceedings of the IEEE International Conference on Dependable, Autonomic and Secure Computing (pp. 60–67). Calgary. https://doi.org/10. 1109/DASC 52594. 2021. 9566527. DOI: https://doi.org/10.1109/DASC-PICom-CBDCom-CyberSciTech52372.2021.00024

[23] Le Pochat, V., Van Goethem, T., Tajalizadehkhoob, S., Korczynski, M., & Joosen, W. (2019). Tranco: A research-oriented top sites ranking hardened against manipulation. In Proceedings of the Network and Distributed System Security Symposium. San Diego. https://doi.org/10. 14722/ndss. 2019. 23106. DOI: https://doi.org/10.14722/ndss.2019.23386

[24] Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321–357. https://doi.org/10.1613/jair.953. DOI: https://doi.org/10.1613/jair.953

Downloads

Published

28-09-2026

Issue

Section

Articles