پدافند الکترونیکی و سایبری

پدافند الکترونیکی و سایبری

دسته بندی داده های وب تاریک به کمک مدل زبانی BERT

نوع مقاله : مقاله پژوهشی

نویسندگان
1 دانشجوی کارشناسی‌ارشد،دانشگاه صنعتی شاهرود ، شاهرود، ایران
2 استادیار، دانشگاه صنعتی شاهرود، شاهرود، ایران
چکیده
ماهیت پنهان و دسترسی محدود وب‌تاریک، موجب گسترش فعالیت‌های مجرمانه بسیاری ازجمله تهدیدات سایبری، فروش اسلحه، فروش مواد مخدر و فروش ابزارهای غیرقانونی شده است. ظهور مدل‌های زبانی بزرگ این امید را ایجاد نموده است که بتوان با دقت مناسبی به تحلیل مطالب موجود در وب تاریک پرداخت. در همین راستا استفاده از داده‌های انبوه سایبری موجود در وب‌تاریک برای جلوگیری از تهدیدات سایبری و آموزش مدل‌های زبانی بسیار مفید و مؤثر خواهد بود. فناوری مدل‌های زبانی بزرگ برای آموزش بهتر و رسیدن به‌دقت کافی، به داده زیاد و باکیفیت بالا نیاز دارند و این چالشی است که محققان حوزه امنیت سایبری با توجه ‌به آلوده بودن داده‌های موجود در وب‌تاریک روبرو هستند. اغلب تحقیقات در این زمینه، متمرکز بر روی تمام مشخصه‌های دادگان وب‌تاریک و داده‌های باکیفیت پایین صورت پذیرفته است و نتوانسته‌اند دقت بالایی را کسب کنند. در این پژوهش یک مدل ‌زبانی جدید بر پایه مدل زبانی پایه BERT که بر روی‌داده استخراج‌شده از وب‌تاریک آموزش‌دیده است، ارائه کردیم. مدل پیشنهادی یک مدل متنی مبتنی بر ترانسفورماتور است که از رمزگذار دوطرفه از ترانسفورماتورها برای رویکرد یادگیری استفاده می‌کند و آن را بر روی یک دادگان باکیفیت بالا، بدون داده تکراری، عاری از کلمات نامعلوم، تماماً به زبان انگلیسی و به‌طور مشخص بر روی‌داده‌های سایبری و امنیت ارزیابی نمودیم. درنهایت با تحلیل مقادیر ارزیابی‌شده مدل پیشنهادی با مدل‌های قبلی، مشخص شد که مدل پیشنهادی به علت تزریق داده‌های باکیفیت نسبت به مدل‌های قبلی، توانسته دقت بهتری در دسته‌بندی داده‌ها داشته باشد.
کلیدواژه‌ها
موضوعات

عنوان مقاله English

Dark Web Text Classification using BERT’s Language Model

نویسندگان English

baratali akhtariyan 1
Mohsen Rezvani 2
1 Master's student,, Shahrood University of Technology, Shahrood, Iran
2 Assistant Professor, , Shahrood University of Technology, Shahrood, Iran
چکیده English

The hidden nature and limited access of the dark web has led to the proliferation of many criminal activities, including cyber threats, arms sales, drug sales, and the sale of illegal tools. The emergence of large language models has created the hope that it will be possible to analyze the content on the dark web with proper accuracy. In this regard, the use of mass cyber data available in the dark web will be very useful and effective to prevent cyber threats and train language models. The technology of large language models requires a lot of high-quality data for better training and to achieve sufficient accuracy, and this is the challenge that researchers in the field of cyber security face due to the contamination of the data available on the dark web. Most of the researches in this field have been focused on all the characteristics of the dark web dataset and low-quality data and have not been able to achieve high accuracy. In this thesis, we presented a new language model based on the BERT-based language model, which was trained on the data extracted from the dark web. The proposed model is a transformer-based text model that uses a two-way encoder of transformers for a learning approach and we evaluated it on a high - quality dataset, without repetitive data, free of unknown words, all in English and specifically on hacking and security data. Finally, by analyzing the evaluated values ​​of the proposed model with the previous models, it was found that the proposed model was able to have better accuracy in data classification due to the injection of quality data compared to the previous models.

کلیدواژه‌ها English

Dark Web
Large Language Models
Transformers
BERT
[1]     Aghaei, Ehsan, and Ehab Al-Shaer. "ThreatZoom: Neural Network for Automated Vulnerability Mitigation." In Proceedings of the 6th Annual Symposium on Hot Topics in the Science of Security, 2019, pp. 1-3. https://doi.org/10.1145/3314058.3318167
[2]     Tann, Wesley, Yuancheng Liu, Jun Heng Sim, Choon Meng Seah, and Ee-Chien Chang. "Using Large Language Models for Cybersecurity capture-the-flag Challenges and Certification Questions." 2023,arXiv:2308.10443.
https://doi.org/10.48550/arXiv.2308.10443
[3]     Devlin, Jacob. "Bert: Pre-training of Deep Bidirectional Transformers for Language Understanding." 2018, arXiv:1810.04805.
[4]     Jin, Youngjin, Eugene Jang, Jian Cui, Jin-Woo Chung, Yongjae Lee, and Seungwon Shin. "DarkBERT: A Language Model for the Dark Side of the Internet." 2023, arXiv:2305.08596.
https://doi.org/10.48550/arXiv.2305.08596
[5]     Choshen, Leshem, Dan Eldad, Daniel Hershcovich, Elior Sulem, and Omri Abend. "The Language of Legal and Illegal Activity on the Darknet." 2019, arXiv:1905.05543.
https://doi.org/10.48550/arXiv.1905.05543
[6]     Al Nabki, Mhd Wesam, Eduardo Fidalgo, Enrique Alegre, and Ivan De Paz. "Classifying Illegal Activities on Tor Network Based on Web Textual Contents." In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers, 2017, pp. 35-43.
[7]     Liao, Xiaojing, Kan Yuan, XiaoFeng Wang, Zhou Li, Luyi Xing, and Raheem Beyah. "Acing the Ioc Game: Toward Automatic Discovery and Analysis of Open-source Cyber Threat Intelligence." In Proceedings of the 2016 ACM SIGSAC conference on computer and communications security,2016, pp. 755-766. https://doi.org/10.1145/2976749.2978315
[8]     Bradbury, Danny. "Unveiling the Dark web." Network security, 2014, no. 4: 14-17. https://doi.org/10.1016/S1353-4858(14)70042-X
[9]     Faizan, Mohd, and Raees Ahmad Khan. "Exploring and Analyzing the Dark web: a New Alchemy." First Monday, 2019. https://doi.org/10.5210/fm.v24i5.9473
[10]     Dingledine, Roger, Nick Mathewson, and Paul F. Syverson. "Tor: The Second-generation Onion Router." In USENIX security symposium, vol. 4,2004, pp. 303-320.
[11]     Jin, Youngjin, Eugene Jang, Yongjae Lee, Seungwon Shin, and Jin-Woo Chung. "Shedding New Light on the Language of the Dark web." 2022, arXiv:2204.06885.
https://doi.org/10.48550/arXiv.2204.06885
[12]     Samtani, Sagar, Weifeng Li, Victor Benjamin, and Hsinchun Chen. "Informing Cyber Threat Intelligence Through Dark Web Situational Awareness: The AZSecure Hacker Assets Portal." Digital Threats: Research and Practice (DTRAP) 2, 2021, no. 4: 1-10.
https://doi.org/10.1145/3450972
[13]     Törnberg, Petter. " How to Use Large-Language Models for Text Analysis", SAGE Publications Ltd, 2024.
[14]     Rajaraman, Nived, Jiantao Jiao, and Kannan Ramchandran. "Toward a Theory of Tokenization in LLMs."2024, arXiv:2404.08335.
https://doi.org/10.48550/arXiv.2404.08335
[15]     Chang, Yupeng, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen et al. "A Survey on Evaluation of Large Language Models." ACM Transactions on Intelligent Systems and Technology 15, 2024, no. 3: 1-45. https://doi.org/10.1145/3641289
[16]     Shanahan, Murray. "Talking About Large Language Models." Communications of the ACM 67,2024, no. 2: 68-79.
https://doi.org/10.1145/3624724
[17]     Koroteev, Mikhail V. "BERT: a Review of Applications in Natural Language Processing and Understanding." 2021, arXiv:2103.11943 (2021).
https://doi.org/10.48550/arXiv.2103.11943
[18]     Biryukov, Alex, Ivan Pustogarov, Fabrice Thill, and Ralf-Philipp Weinmann. "Content and Popularity Analysis of Tor Hidden Services." In 2014 IEEE 34th International Conference on Distributed Computing Systems Workshops (ICDCSW),2014, pp. 188-193.
[19]     Moore, Daniel, and Thomas Rid. "Cryptopolitik And the Darknet." Survival 58,2016, no. 1: 7-38.
[20]     Wolf, Thomas, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac et al. "Huggingface's Transformers: State-of-the-art Natural Language Processing."2019, arXiv:1910.03771.
[21]     Mohammed, Athar Hussein, and Ali H. Ali. "Survey of Bert (bidirectional encoder representation transformer) Types." In Journal of Physics: Conference Series, vol. 1963, no. 1,2021, pp. 012173.
[22]     Loshchilov, Ilya, and Frank Hutter. "Decoupled Weight Decay Regularization."2017, arXiv:1711.05101.
[23]     Townsend, James T. "Theoretical Analysis of an Alphabetic Confusion Matrix.",1971, Perception & Psychophysics 9: 40-50. https://doi.org/10.3758/BF03213026
[24]     Magnus, Amy L., and Mark E. Oxley. "Theory of confusion." In Applications and Science of Neural Networks, Fuzzy Systems, and Evolutionary Computation IV, vol. 4479, 2001, pp. 105-116. https://doi.org/10.1117/12.448337
[25]     Grandini, Margherita, Enrico Bagli, and Giorgio Visani. "Metrics for Multi-class Classification: an overview."2020, arXiv:2008.05756.
https://doi.org/10.48550/arXiv.2008.05756
[26]     Sokolova, Marina, Nathalie Japkowicz, and Stan Szpakowicz. "Beyond Accuracy, F-score and ROC: a family of Discriminant Measures for Performance Evaluation." In Australasian joint conference on artificial intelligence,2006, pp. 1015-1021.
[27]     Liu, Yang, Jiahuan Cao, Chongyu Liu, Kai Ding, and Lianwen Jin. "Datasets for Large Language Models: A Comprehensive Survey."2024, arXiv:2402.18041.
https://doi.org/10.48550/arXiv.2402.18041
[28]     Wang, Jiahao, Bolin Zhang, Qianlong Du, Jiajun Zhang, and Dianhui Chu. "A Survey on Data Selection for LLM Instruction Tuning."2024, arXiv:2402.05123.
https://doi.org/10.48550/arXiv.2402.05123
[29]     Ketkar, Nikhil, Jojo Moolayil, Nikhil Ketkar, and Jojo Moolayil. "Introduction to Pytorch."2021, Deep learning with python: learn best practices of deep learning models with PyTorch : 27-91.
[30]     Salton, Gerard, and Clement T. Yu. "On the Construction of Effective Vocabularies for Information Retrieval." Acm Sigplan Notices 10, 1973, no. 1: 48-60. https://doi.org/10.1145/951787.951766
[31]     Ali, Mehdi, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max Lübbering, Johannes Leveling, Katrin Klug et al. "Tokenizer Choice For LLM Training: Negligible or Crucial?."2023, arXiv:2310.08754.
https://doi.org/10.48550/arXiv.2310.08754
[32]     Jiang, Ming, Jennifer D’Souza, Sören Auer, and J. Stephen Downie. "Improving Scholarly Knowledge Representation: Evaluating bert-based Models for Scientific Relation Classification." In Digital Libraries at Times of Massive Societal Transition: 22nd International Conference on Asia-Pacific Digital Libraries, ICADL 2020, Kyoto, Japan, Proceedings 22, November 30–December 1,2020, pp. 3-19.
دوره 13، شماره 4 - شماره پیاپی 52
زمستان
زمستان 1404
صفحه 63-76

  • تاریخ دریافت 28 شهریور 1404
  • تاریخ بازنگری 17 آبان 1404
  • تاریخ پذیرش 07 آذر 1404
  • تاریخ انتشار 01 دی 1404