Ningxi Zhang

Pronouns: She/Her
Research Mentor(s): Michael Kallitsis
Research Mentor School/College/Department: Merit Network, Inc. / Other
Program: CG
Session: Session 5 (2:40pm – 3:30pm)
Authors:
Abstract
This research aims to improve the efficiency of detecting and preventing cyber attacks in network security by constructing a labeled machine learning dataset using real-time network data from Merit Network at University of Michigan. Past studies have conducted behavioral clustering of HTTP-based malware and generated signatures using malicious network traces, used packet sizes and inter-arrival times for classification, transformed basic flow data into an intuitive picture and used image classification with deep learning, generated vector representations for bytes and payloads, and developed a neural network architecture based on LSTM and CNN layers. This study considers both payloads and malicious scannings, like Mirai, Masscan and Zmap, as labels to build payload-driven features and improve the overall performance of the classification model. Unlike traditional methods that compare payloads to find the right label in a large data library, this research takes a more comprehensive approach by learning from the past labeled data to predict the right label for the new real-life network data. The study uses advanced AI/ML methods, such as the n-gram language model and different payload types like hex, and base64, to enhance the payload-driven machine learning model’s performance. The hypothesis is that the constructed payload-driven security model will be more efficient than traditional methods in detecting and preventing various types of cyberattacks. The ultimate goal is to assist security analysts in obtaining useful insights and annotating observed scanners with meaningful data and improve network security.



