CN118413402B - Malicious domain name detection method based on large language model - Google Patents
Malicious domain name detection method based on large language model Download PDFInfo
- Publication number
- CN118413402B CN118413402B CN202410875232.5A CN202410875232A CN118413402B CN 118413402 B CN118413402 B CN 118413402B CN 202410875232 A CN202410875232 A CN 202410875232A CN 118413402 B CN118413402 B CN 118413402B
- Authority
- CN
- China
- Prior art keywords
- layer
- domain name
- output
- model
- attention
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Active
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/214—Generating training patterns; Bootstrap methods, e.g. bagging or boosting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/90—Details of database functions independent of the retrieved data types
- G06F16/95—Retrieval from the web
- G06F16/955—Retrieval from the web using information identifiers, e.g. uniform resource locators [URL]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/24—Classification techniques
- G06F18/243—Classification techniques relating to the number of classes
- G06F18/2431—Multiple classes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/25—Fusion techniques
- G06F18/253—Fusion techniques of extracted features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L61/00—Network arrangements, protocols or services for addressing or naming
- H04L61/45—Network directories; Name-to-address mapping
- H04L61/4505—Network directories; Name-to-address mapping using standardised directories; using standardised directory access protocols
- H04L61/4511—Network directories; Name-to-address mapping using standardised directories; using standardised directory access protocols using domain name system [DNS]
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L63/00—Network architectures or network communication protocols for network security
- H04L63/14—Network architectures or network communication protocols for network security for detecting or protecting against malicious traffic
- H04L63/1408—Network architectures or network communication protocols for network security for detecting or protecting against malicious traffic by monitoring network traffic
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L9/00—Cryptographic mechanisms or cryptographic arrangements for secret or secure communications; Network security protocols
- H04L9/40—Network security protocols
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Data Mining & Analysis (AREA)
- General Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Computer Security & Cryptography (AREA)
- Life Sciences & Earth Sciences (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Computation (AREA)
- Computer Networks & Wireless Communication (AREA)
- Evolutionary Biology (AREA)
- Signal Processing (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Bioinformatics & Computational Biology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Databases & Information Systems (AREA)
- Computing Systems (AREA)
- Biophysics (AREA)
- Mathematical Physics (AREA)
- General Health & Medical Sciences (AREA)
- Computer Hardware Design (AREA)
- Molecular Biology (AREA)
- Software Systems (AREA)
- Computational Linguistics (AREA)
- Biomedical Technology (AREA)
- Health & Medical Sciences (AREA)
- Machine Translation (AREA)
Abstract
The invention relates to a malicious domain name detection method based on a large language model, which solves the defect that the malicious domain name is difficult to detect compared with the prior art. The invention comprises the following steps: constructing a pre-training data set and a fine-tuning training data set; setting a URL-BERT model; pre-training of URL-BERT model; fine tuning of URL-BERT model; obtaining a domain name to be detected; obtaining a malicious domain name detection result. The invention adopts the big language model BERT to process the malicious domain name, and utilizes the strong semantic understanding capability of the big language model to better capture the implicit information and the context in the domain name and improve the identification accuracy of the malicious domain name.
Description
Technical Field
The invention relates to the technical field of network space safety and artificial intelligence, in particular to a malicious domain name detection method based on a large language model.
Background
With the rapid development of the internet, network security is threatened, various malicious websites and malicious software are flooded into disaster, and many users and websites are easily threatened by security, wherein malicious domain names are one of main paths of network attacks. Tens of thousands of new domain names are presented every day, and besides a large number of benign domain names (Benign Domain) which are normally registered, a large number of malicious domain names (Malicious Domain) which are maliciously registered for illegal activities exist, and the malicious domain names become one of main attack paths of illegal activities of various networks.
Domain names serve as bridges connecting users with malicious agents (e.g., luxury software, malware, and trojans), which are commonly used to implement malicious activities such as illegally stealing and storing personally sensitive information, illegally controlling and managing infected hosts, and the like. Almost all network attacks rely on the use of malicious domain names and the domain name system is subject to abuse by attackers. Therefore, how to effectively detect malicious domain names is a hotspot and difficulty problem in the network security field.
The existing malicious domain name detection method is mainly divided into three types, namely rule-based, machine learning and deep learning. The legitimacy of the domain name is detected through a black-and-white list based method, and the method is simple and direct; extracting features (such as domain name length, number of different IP addresses, etc.) from the domain name based on a machine learning method, and then constructing a classifier based on the machine learning to distinguish benign domain names from malicious domain names; the method based on deep learning carries out coding modeling on the text of the domain name through a deep neural network.
The existing malicious domain name detection method has good effects to a certain extent. But rule-based approaches have limited effectiveness in facing explosively growing malicious domain names; the machine learning-based approach relies heavily on expert knowledge and is limited by feature expression capabilities, and an attacker can easily bypass model detection by adding minimal feature perturbations (changing the length of the domain or modifying network parameters, such as TTL); the deep learning-based method has the defect of neglecting importance of position embedding on domain name vectors, lack of interpretability and the like.
Domain names are typically made up of multiple parts, such as secondary domain names, top-level domain names, etc., that are connected by points (), as shown in fig. 2. Removing special characters such as "," - ", and the like in the Domain name, and separating the Domain name by using the", "as a separator to obtain the Token of the Domain name, wherein the deep learning method generally regards the Domain name as plain text data, ignores the sensitivity of the Token in the Domain name to the position, namely, the same Token should have different meanings and representation methods when the Token is positioned in Domain/Top Level Domain (TLD)/Path.
Then, how to design a malicious domain name detection method based on the special text information of the domain name has become a technical problem to be solved urgently.
Disclosure of Invention
The invention aims to solve the defect that the malicious domain name is difficult to detect in the prior art, and provides a malicious domain name detection method based on a large language model to solve the problems.
In order to achieve the above object, the technical scheme of the present invention is as follows:
a malicious domain name detection method based on a large language model comprises the following steps:
11 Construction of a pre-training dataset and a fine-tuning training dataset: constructing a pre-training data set and a fine-tuning training data set, and performing data preprocessing;
12 Setting a URL-BERT model;
13 Pre-training of URL-BERT model: pre-training the URL-BERT model by utilizing a pre-training data set;
14 Fine tuning of URL-BERT model: utilizing the fine tuning training data set to carry out fine tuning on the pre-trained URL-BERT model;
15 Obtaining a domain name to be detected): obtaining a domain name to be detected;
16 Obtaining malicious domain name detection results: and inputting the domain name to be detected into the URL-BERT model after fine adjustment to obtain a malicious domain name detection result.
The construction of the pre-training data set and the fine-tuning training data set comprises the following steps:
21 A pre-training data set is set, the pre-training data set is unmarked domain name data, and the unmarked domain name data is obtained from a public DNS query log and a WHOIS database source;
22 Setting a fine tuning training data set, wherein the fine tuning training data set is marked domain name data, the marked domain name data is obtained from a network security company or a public malicious domain name database, and the number of positive samples and negative samples is balanced;
23 Pre-processing the pre-training data set and the fine-tuning training data set by the steps of de-duplication, irrelevant characters removal, conversion into lower case letters and standardization;
24 Word segmentation is carried out on the pre-training data set and the fine training data set after pretreatment.
The URL-BERT model setting method comprises the following steps of:
31 The URL-BERT model is set to comprise an input layer, a URL-BERT pre-training layer, a character level embedded feature extraction layer, a mark level embedded feature extraction layer, a feature fusion layer and a fully-connected classification layer;
Setting an input layer: the input layer receives text after word segmentation as the input of a model;
32 Setting a URL-BERT pre-training layer: training a large amount of unmarked domain name data by setting an MLM task in a pre-training stage, and learning and capturing rich semantic information and structural features in domain name texts;
33 Setting a character-level embedded feature extraction layer: the character level embedding extraction layer extracts character level embedding of the domain name text, and bidirectional vector representation is generated by BiGRU, wherein complex association information among characters in the domain name is contained; for the fine tuning training data set, the bi-directional gating unit BiGRU is utilized to perform character level embedding and extraction on the marked domain name data, and BiGRU combines hidden layer states in the forward direction and the backward direction by utilizing GRUs in two different directions, so that the integration of bi-directional information is realized, and the embedded vector of the domain name character level is obtained ;
34 Setting a mark level embedding feature extraction layer to perform mark level embedding;
35 Setting a feature fusion layer):
The feature fusion layer splices the extracted character level embedded feature vector and the mark level embedded feature vector, and then introduces a plurality of Transformer layers to capture complex relations among features;
The heterogeneous interaction module between the transducer layers is used for embedding the character level and the marking level into each transducer layer and then combining and separating, the combination operation enriches the relevance between different representations, the separation operation reserves the independence of the character level and the marking level characteristics, and the representation of the model in the aspect of double-channel distinction is promoted;
36 Setting a full connection classification layer): the full-connection classification layer is used for fine-tuning the whole URL-BERT model to adapt to a specific malicious domain name detection task, and the loss function is minimized through a back propagation algorithm by inputting the feature vector output by the feature fusion layer into the full-connection layer.
The pre-training of the URL-BERT model comprises the following steps:
41 Inputting the text subjected to word segmentation processing of the pre-training data set into a URL-BERT model;
42 URL-BERT pre-training layer Masked Language Model pre-training, i.e., MLM pre-training:
Is provided with Is randomly replaced by [ MASK ]Word sequences of a small number of token,15% Of the token is randomly processed [ MASK ], whereIndicating that the ith token was randomly replaced, the goal of the MLM pre-training process is to predict the contents of the [ MASK ] token based on context;
Will be Is input into the URL-BERT model,Corresponding hidden layer vector, From the following componentsThe code is obtained by the encoder of BERT;
421 Encoder processing: the input data of the encoded input layer is an embedded vector ;
422 Self-attention mechanism):
Input device Obtaining a query matrix by linear transformationKey matrix K and value matrixWherein, , Is the super parameter of the encoder:
,
,
,
The attention weight is then calculated, the attention is scaled using the dot product, the output produced in this step contains a representation of global dependencies inside the input sequence,
Wherein the method comprises the steps ofIs the dimension of the key vector, is used for scaling dot products, prevents the gradient vanishing problem,Is a matrix of attention weights that,Is the output of one of the attention heads,
,
Output of multi-head attention:
the outputs of the multiple heads are spliced together and then transformed by another linear transformation, formulated as follows,
,
Wherein, Is the firstThe output of the individual head is provided with,Is a super parameter matrix;
423 Feedforward neural network):
The output for the multi-head attention is characterized by a feed-forward neural network, which consists of two linear layers and an activation function,
,
Wherein the method comprises the steps of、、、Is a learnable parameter;
and then carrying out residual connection and layer normalization for enhancing training stability and convergence rate of the model:
,
Here, the Representation layer normalization operations;
Encoder output layer: coding sequence of final output encoder ;
Next to this, the process is carried out,Obtaining corresponding token distribution through Sigmoid operation;
the formula of the MLM process is as follows,
,
Wherein the method comprises the steps ofRepresenting a target probability distribution, i.e. the whole sequenceIs a function of the probability distribution of (1),Representing the predicted probability of the ith token given all the tokens of the sequence,The parameters of the model are represented by the parameters,Indicating whether the ith token is [ MASK ], if soOtherwiseThe purpose is to weight the token of [ MASK ] which affects the probability of the whole sequence;
In MLM pre-training, the goal is to maximize the model over the entire sequence Based on this prediction, calculating a cross entropy loss function and updating model parameters by a back propagation algorithmFinally obtaining a pre-trained model,
,
Wherein, Is a parameter of the model and is a parameter of the model,Is a cross-entropy loss function,Is the true label of the ith token, if the token is [ MASK ] i.e. the true value of the token,Is the predictive probability of the ith token, i.eMinimizing loss function by back propagation algorithmUpdating model parametersFinally obtaining a pre-trained URL-BERT model。
The fine tuning of the URL-BERT model comprises the following steps:
51 Training the character level embedded feature extraction layer and the mark level embedded feature extraction layer to perform character level embedding;
the fine tuning data set forms word segmentation sequence after word segmentation In one aspect of the present invention,Through character level embedded feature extraction layer embedding, an embedded vector is obtained; On the other handEmbedded by a mark level embedded feature extraction layer, i.eEmbedding the pre-trained model, adding an attention mask vectorPerforming position embedding processing to obtain a mark level embedded vector with position embedding;
Finally, the marker level is embedded into the vectorWith character level embedded vectorsAre connected together, i.e. inVector end joinVectors forming a longer embedded vector,
;
52 Setting a feature fusion layer): setting the feature fusion layer as a modelModelModel comprising frozen parameterA layer, a plurality of transducer layers, a plurality of heterogeneous interaction layers, a fully connected layer and a Sigmoid layer; through interaction among the modules, the feature fusion layer can capture the complex relationship between the url character level and the mark level;
training the feature fusion layer:
521 Model for setting frozen parameters The layer is a pre-trained URL-BERT model,
522 Setting a transducer module:
the transducer consists of an encoder for processing the input data, and a decoder for generating an output,
The encoder comprises an encoding input layer, a self-attention mechanism, a multi-head attention, a feedforward neural network, residual connection and layer normalization, and an encoder output layer;
the decoder comprises a decoding input layer, a self-attention mechanism, an encoder-decoder attention mechanism, a feedforward neural network, a residual connection and layer normalization and a decoder output layer;
5221 Encoder processing: the input data of the encoded input layer is an embedded vector ;
Self-attention mechanism:
For each head Will inputObtaining a query matrix by linear transformationKey matrixSum matrixWherein, , ,Is the super parameter of the encoder:
,
,
,
next, for each head, an attention weight is calculated, and the scaled dot product attention is used, the output resulting from this step containing a representation of global dependencies inside the input sequence,
Wherein the method comprises the steps ofIs the dimension of the key vector, is used for scaling dot products, prevents the gradient vanishing problem,Is a matrix of attention weights that,Is the output of the attention mechanism of a head,
,
Output of multi-head attention:
the outputs of the multiple heads are spliced together and then transformed by another linear transformation, formulated as follows,
,
Wherein, Is the output of each of the attention heads,Is a super parameter matrix;
Feedforward neural network:
The output for the multi-head attention is characterized by a feed-forward neural network, which consists of two linear layers and an activation function,
,
Wherein the method comprises the steps of、、、Is a learnable parameter;
residual connection and layer normalization: for enhancing the training stability and convergence speed of the model,
,
Here, theRepresentation layer normalization operations;
Encoder output layer: coding sequence of final output encoder ;
5222 Decoder processing:
decoding input data of an input layer into embedded vectors
Self-attention mechanism: obtaining a query matrix by linear transformation Key matrixSum matrixWherein, , ,Is the super-parameter of the decoder and,
,
,
,
Calculating an attention weight for each head, using a scaled dot product attention;
the outputs of the multiple heads are spliced together, and the self-attention output of the decoder is obtained through residual connection and layer normalization ;
Encoder-decoder attention mechanism: the decoder performs attention calculation on the output of the encoder to capture the relationship between the input sequence and the output sequence, allowing the decoder to adjust its attention according to the information of the encoder to generate a target sequence;
For each position of the decoder Using the output of the decoderAs a query, the output of the encoder、As a key and value, an attention weight is calculated:
,
,
Output of multi-head attention: combining the self-attention output of the decoder with the encoder-decoder attention output to obtain a final attention output;
they are added by element weights, the expression of which is as follows:
,
wherein the final attention output is ,Is a learnable weight parameter:
Feedforward neural network: to the final attention output as Performing a feed forward neural network process to further extract and transform the feature representation:
;
Residual connection and layer normalization: residual connection and layer normalization are applied on the output of the feedforward neural network to enhance training stability and convergence speed of the model:
,
decoder output: the output sequence of the decoder is ;
523 Setting a heterogeneous interaction module: the marker-level and character-level representations are treated as two different types of feature representations in the heterogeneous interaction layer, and after each transducer conversion layer, the two representations are combined and separated;
The fully connected network is responsible for mapping the labels and character representations into the same feature space to ensure that they have consistent dimensions and semantic representations, then connecting and integrating through a CNN layer, and finally performing residual connection and layer normalization operation;
5231 Separation and combination treatment:
vector to be output by decoder Separated by length into mark-level representationsAnd character level representation, A mark/character representing the i-th position, respectively toAndConversion is performed by using a special full-connection network:
,
,
The converted token and character representations are then concatenated to form a fused representation:
;
5232 CNN convolution: for fused representations The CNN operation is applied, for integrating features,
,
Wherein, Representing slaveTo the point ofIs provided with an embedded splice of (a),A window size representing a jth filter;
5233 Separation: the convolved combined representation A characteristic representation separated into two different channels,
,
,
Wherein GELU is an activation function for separating the fused feature representations;
5234 Residual connection and layer normalization: the separated feature representation is added to the initial tag and character representations to reassemble the two channel representations:
,
,
finally, carrying out layer normalization operation on the recombined representation to ensure stability:
,
,
The separated features are connected to obtain the input of the next layer of transducer;
524 Setting a full connection and Sigmoid classification module:
training of the full connection layer:
Output of last layer of transducer Is transferred to the device withThe fully connected layers of the individual hidden units are mathematically expressed as follows:
,
Wherein, Is the input vector which is to be used for the input,AndIs the weight and bias of the layer,Is an activation function;
Sigmoid is used for mapping the output of the full connection layer to the category probability distribution, and the Sigmoid layer maps the output result of the full connection layer Passing to Sigmoiod function, the Sigmoid function maps the input real number value to probability value between 0 and1, and the output of the full connection layer isThe output after the Sigmoid function is expressed as:
;
53 URL-BERT model fine tuning process:
Embedding marker levels of a fine-tuning dataset into vectors With character level embedded vectorsProcessed embedded vectorUsing a pre-trained URL-BERT modelConstructed fine tuning modelBy inputting data into a modelComparing the output result with a real label, calculating a loss function, updating model parameters through a back propagation algorithm, and iteratively training the model until convergence, so that the URL-BERT model after fine adjustment carries out malicious domain name classification.
The word segmentation processing of the preprocessed pre-training data set and the fine-tuning training data set comprises the following steps:
61 For the domain name data in the pre-training data set and the fine-tuning training data set, splitting the whole domain name character string according to the structural characteristics, removing special characters in the domain name, and splitting the domain name by taking "/", "-" as separators:
inputting a domain name after pretreatment A word segmentation function f is defined to represent this process:
,
the word segmentation function f () input is a domain name in the form of a preprocessed character string, and the output is a word segmentation list composed of segments segmented by the domain name Wherein, Is thatNumber of medium word content; the expression is as follows:
,
after preliminary word segmentation, the BERT word segmentation device pair is utilized Performing embedding treatment;
62 Word embedding processing: in order to identify the beginning and end of an input sequence, in Adding [ CLS ] mark at the beginning, adding [ SEP ] mark at the end, and limiting the token sequence length to 128, for input sequences with a length of less than 128 marks, using a specific PAD mark [ PAD ] to extend the sequence length until it reaches the limit of 128 marks, thereby obtaining;
63 Segment embedding processing is performed: the segment embeddings are all set to 0, i.e.:
Added to the same length of [0, …,0 ];
64 Position embedding processing: position embedding is based on The different position information of the token encodes a position vector, the position relation of the token is introduced into the URL-BERT model, and the position embedding is set to 0, namely:
Adding the sequence subjected to segment embedding treatment with the sequence with the same length of 0;
65 Generating an attention mask vector: the attention mask vector is used to distinguish the actual markers in the input sequence from the filled markers,
For each real marker in the sequence, its corresponding attention is set to 1; for each PAD, its corresponding attention is set to 0; generating an attention mask directionWherein the number of 1 isNamely:
the sequence subjected to the position embedding processing is added to the sequence [1, …,0 ].
The character level embedding comprises the following steps:
71 Input sequence representation: first, for an input sequence Each of which is provided withIs the mark of sub-word segmentation,Is the total number of subwords in the sequence, each tagFrom charactersComposition of whichThe length of the subword is indicated,
The total length of the character input is expressed as;
72 Character vector generation: for each characterAll generate a character embedding vector,
By combining charactersMatrix for embedding charactersMultiplication is done, i.eThe process maps each character to a vector representation in a high-dimensional space;
73 BiGRU) treatment: generated character embedding vector Is fed into BiGRU for processing, the result of BiGRU operation is the generation of hidden state vectors for each character,
Hiding state vectorContextual information including characters;
74 Building character level embedding: for each mark The character-level embedded vectors thereof form a sequenceThis sequence is processed into a mark-level embedded vector;
For each markConnecting the hidden state vectors of the first character and the last character to obtain a mark level embedded vectorBy processing the characters of the whole input sequence, the embedded vector of the contextual characters of each mark, namely the embedded vector of the domain name character level, is obtainedThe expression is as follows:
。
Advantageous effects
Compared with the prior art, the malicious domain name detection method based on the large language model adopts the large language model BERT to process the malicious domain name, and utilizes the powerful semantic understanding capability of the large language model to better capture hidden information and context in the domain name and improve the identification accuracy of the malicious domain name. Meanwhile, a position embedding technology is introduced during pre-training, and the importance of the position information of the Token in the domain name to the representation is fully considered. The Token meaning at different positions can be expressed more accurately by combining the position information with the semantic vector of the Token, so that the detection performance of the model on malicious domain names is improved.
The invention also utilizes BiGRU to extract character level embedding, fully utilizes two-way information, and better captures context information with finer granularity, thereby more comprehensively understanding url data and improving embedding accuracy and expression capability.
Compared with the traditional feature extraction method based on machine learning, the method can avoid manual feature design and can automatically learn higher-level feature representation. The end-to-end training mode can be better adapted to the change and complexity of data, so that the generalization performance and the mobility of the model are improved, and the capability of coping with the attack resistance is improved. Compared with a deep learning method, the method introduces position embedding and improves the processing capacity of the model on the sequence data. Meanwhile, the generalization capability and the migration capability of the model are improved to the greatest extent through the pre-training task, and the model can be quickly adapted and has good effects even if facing a new malicious domain name classification task.
In conclusion, the malicious domain name detection method based on the large language model has obvious advantages in accuracy and universality.
Drawings
FIG. 1 is a process sequence diagram of the present invention;
FIG. 2 is a diagram of a URL structure in the prior art;
FIG. 3 is a block diagram of a BiGRU network;
FIG. 4 is a block diagram of a heterogeneous interaction layer;
fig. 5 is a diagram of the overall architecture of the model.
Detailed Description
For a further understanding and appreciation of the structural features and advantages achieved by the present invention, the following description is provided in connection with the accompanying drawings, which are presently preferred embodiments and are incorporated in the accompanying drawings, in which:
the traditional detection method based on the characteristic engineering is difficult to obtain good effect when facing to complex network change, and secondly, the same Token in the Domain name should have different meanings and representation methods when being positioned in Domain/TLD/Path, namely, the Token of the Domain name has position sensitivity. Existing research regards domain names as plain text sequences and applies a language model to characterize domain names. Many studies use a "bag of words" model to model domain name labels and extract lexical features, such as using n-gram features for domain name classification, ignoring the position sensitivity of Token altogether.
The invention pre-trains the BERT through unmarked domain name data, and then effectively combines character level embedded features and mark level embedded features of the domain name to finely tune the URL-BERT. By utilizing the context understanding capability and position sensitivity of the BERT model, the semantic information of the domain name deep layer is captured, and the serious dependence on manual definition rules or characteristics in a purely feature-based method is avoided.
As shown in fig. 1, the method for detecting malicious domain names based on a large language model according to the invention comprises the following steps:
First, constructing a pre-training data set and a fine-tuning training data set: and constructing a pre-training data set and a fine-tuning training data set, and performing data preprocessing.
(1) And setting a pre-training data set, wherein the pre-training data set is unmarked domain name data, and the unmarked domain name data is obtained from a public DNS query log and WHOIS database source.
(2) And setting a fine tuning training data set, wherein the fine tuning training data set is marked domain name data, the marked domain name data is obtained from a network security company or a public malicious domain name database, and the number of positive samples and negative samples is balanced.
(3) The pre-training data set and the fine-tuning training data set are subjected to preprocessing of steps of de-duplication, irrelevant characters removal, conversion into lower case letters and standardization.
(4) And performing word segmentation on the preprocessed pre-training data set and the fine-tuning training data set. Which comprises the following steps:
For the domain name data in the pre-training data set and the fine-tuning training data set, the word segmentation operation splits the whole domain name character string according to the structural characteristics of the domain name character string, removes special characters in the domain name, and segments the domain name by taking "/", "-" as separators:
inputting a domain name after pretreatment A word segmentation function f is defined to represent this process:
,
the word segmentation function f () input is a domain name in the form of a preprocessed character string, and the output is a word segmentation list composed of segments segmented by the domain name Wherein, Is thatNumber of medium word content; the expression is as follows:
,
after preliminary word segmentation, the BERT word segmentation device pair is utilized Performing embedding treatment;
Word embedding processing is carried out: in order to identify the beginning and end of an input sequence, in Adding [ CLS ] mark at the beginning, adding [ SEP ] mark at the end, and limiting the token sequence length to 128, for input sequences with a length of less than 128 marks, using a specific PAD mark [ PAD ] to extend the sequence length until it reaches the limit of 128 marks, thereby obtaining;
Segment embedding treatment is carried out: the segment embeddings are all set to 0, i.e.:
Added to the same length of [0, …,0 ];
And (3) performing position embedding processing: position embedding is based on The different position information of the token encodes a position vector, the position relation of the token is introduced into the URL-BERT model, and the position embedding is set to 0, namely:
Adding the sequence subjected to segment embedding treatment with the sequence with the same length of 0;
generating an attention mask vector: the attention mask vector is used to distinguish the actual markers in the input sequence from the filled markers,
For each real marker in the sequence, its corresponding attention is set to 1; for each PAD, its corresponding attention is set to 0; generating an attention mask directionWherein the number of 1 isNamely:
the sequence subjected to the position embedding processing is added to the sequence [1, …,0 ].
And secondly, setting a URL-BERT model.
The overall architecture of the model is designed, and the characteristics of character level and mark level feature extraction, feature fusion, classification and the like are comprehensively considered in consideration of the specificity of domain name text, so that a comprehensive model structure is designed.
In the pre-training stage, the URL-BERT model pre-trains large-scale unmarked domain name data through an MLM task, and learns and captures rich semantic information and structural features in domain name texts; in the feature extraction and fusion stage, a proper feature extraction method is designed to integrate character-level and mark-level features, and a transducer layer is utilized to capture complex relations between features so as to improve the representation capability and classification performance of the model.
The design has the advantage of comprehensively utilizing information of different levels, thereby improving the characterization capability and classification performance of the model. The character-level embedded features can capture fine-grained features of the domain name text, such as semantic information and sequential relationships of single characters; while the tag-level embedded features are able to capture higher level semantic information. Through feature fusion, the model can simultaneously consider the features of the two layers, so that the model can more comprehensively understand the meaning of the domain name text.
In the URL-BERT architecture design, feature fusion can comprehensively utilize embedded features of character level and mark level, so that a model can more comprehensively understand domain name text; while the Transformer can handle long distance dependencies in domain name text and better capture context information. The advantages are combined, and the accuracy and generalization capability of the URL-BERT model in malicious domain name detection tasks are improved.
The method comprises the following specific steps:
(1) As shown in fig. 5, the URL-BERT model is set to include an input layer, a URL-BERT pre-training layer, a character-level embedded feature extraction layer, a mark-level embedded feature extraction layer, a feature fusion layer and a full-connection classification layer;
Setting an input layer: the input layer receives text after word segmentation as input of a model.
(2) Setting a URL-BERT pre-training layer: in the pre-training stage, a large amount of unmarked domain name data are trained by setting an MLM task, and rich semantic information and structural features in domain name texts are learned and captured.
(3) Setting a character-level embedded feature extraction layer: the character level embedding extraction layer extracts character level embedding of the domain name text, and bidirectional vector representation is generated by BiGRU, wherein complex association information among characters in the domain name is contained; for the fine tuning training dataset, as shown in fig. 3, the bi-directional gating unit BiGRU is used to perform character-level embedding extraction on the tagged domain name data, and BiGRU combines hidden layer states in the forward direction and the backward direction by using two GRUs in different directions, so that the integration of bi-directional information is realized, and the embedded vector of the domain name character level is obtained。
(4) And setting a mark level embedding feature extraction layer, and performing mark level embedding.
(5) Setting a feature fusion layer:
The feature fusion layer splices the extracted character level embedded feature vector and the mark level embedded feature vector, and then introduces a plurality of Transformer layers to capture complex relations among features;
And the heterogeneous interaction module is used for embedding the character level and the mark level in each transducer layer and then combining and separating the character level and the mark level, the combination operation enriches the relevance between different representations, and the separation operation maintains the independence of the character level and the mark level characteristics and promotes the representation of the model in the aspect of double-channel distinction.
(6) Setting a full connection classification layer: the full-connection classification layer is used for fine-tuning the whole URL-BERT model to adapt to a specific malicious domain name detection task, and the loss function is minimized through a back propagation algorithm by inputting the feature vector output by the feature fusion layer into the full-connection layer.
Third, pre-training of URL-BERT model: the URL-BERT model is pre-trained using the pre-training dataset.
The URL-BERT model is obtained by pre-training BERT (Bidirectional Encoder Representations from Transformers) of a large number of unmarked domain name data, learns rich semantic information and structural features in the domain name data in a self-supervision learning mode, understands semantic relations, lexical rules and structural features in domain name texts, and provides deeper and comprehensive language understanding capability for subsequent malicious domain name detection tasks.
The pre-training process involves knowledge in multiple fields such as natural language processing, deep learning, etc. The method improves the effect and efficiency of natural language processing applied to url, and can greatly improve the accuracy of understanding url text and the downstream generalization capability of the model. Meanwhile, the training cost of the model is reduced, the training time is shortened, the existing url database can be effectively utilized, and the data utilization rate is improved.
(1) And inputting the text subjected to word segmentation processing of the pre-training data set into a URL-BERT model.
(2) The URL-BERT pre-training layer performs Masked Language Model pre-training, namely MLM pre-training:
Is provided with Is randomly replaced by [ MASK ]Word sequences of a small number of token,15% Of the token is randomly processed [ MASK ], whereIndicating that the ith token was randomly replaced, the goal of the MLM pre-training process is to predict the contents of the [ MASK ] token based on context;
Will be Is input into the URL-BERT model,Corresponding hidden layer vector,From the following componentsThe code is obtained by the encoder of BERT;
a1 Encoder processing: the input data of the encoded input layer is an embedded vector ;
A2 Self-attention mechanism):
Will input Obtaining a query matrix by linear transformationKey matrix K and value matrixWherein, , Is the super parameter of the encoder:
,
,
,
The attention weight is then calculated, the attention is scaled using the dot product, the output produced in this step contains a representation of global dependencies inside the input sequence,
Wherein the method comprises the steps ofIs the dimension of the key vector, is used for scaling dot products, prevents the gradient vanishing problem,Is a matrix of attention weights that,Is the output of the attention mechanism of a head,
,
Output of multi-head attention:
the outputs of the multiple heads are spliced together and then transformed by another linear transformation, formulated as follows,
,
Wherein, Is the firstThe output of the individual head is provided with,Is a super parameter matrix;
A3 Feedforward neural network):
The output for the multi-head attention is characterized by a feed-forward neural network, which consists of two linear layers and an activation function,
,
Wherein the method comprises the steps of、、、Is a learnable parameter;
residual connection and layer normalization: for enhancing the training stability and convergence speed of the model,
,
Here, theRepresentation layer normalization operations;
Encoder output layer: coding sequence of final output encoder ;
Next to this, the process is carried out,Obtaining corresponding token distribution through Sigmoid operation;
the formula of the MLM process is as follows,
,
Wherein the method comprises the steps ofRepresenting a target probability distribution, i.e. the whole sequenceIs a function of the probability distribution of (1),Representing the predicted probability of the ith token given all the tokens of the sequence,The parameters of the model are represented by the parameters,Indicating whether the ith token is [ MASK ], if soOtherwiseThe purpose is to weight the token of [ MASK ] which affects the probability of the whole sequence;
In MLM pre-training, the goal is to maximize the model over the entire sequence Based on this prediction, calculating a cross entropy loss function and updating model parameters by a back propagation algorithmFinally obtaining a pre-trained model,
,
Wherein, Is a parameter of the model and is a parameter of the model,Is a cross-entropy loss function,Is the true label of the ith token, if the token is [ MASK ] i.e. the true value of the token,Is the predictive probability of the ith token, i.eMinimizing loss function by back propagation algorithmUpdating model parametersFinally obtaining a pre-trained URL-BERT model。
Fourth, fine tuning of URL-BERT model: and performing fine tuning on the pre-trained URL-BERT model by utilizing the fine tuning training data set.
The fine tuning step needs to optimize and adjust the model, mainly trains aiming at each module in the feature fusion layer, selects proper super parameters and adjusts the learning force of the model so as to achieve the best fine tuning effect. The fine tuning process of the model further improves the classification accuracy of malicious domain name detection on the basis of pre-training, and the model has better feature expression and text modeling capability through feature fusion and application of a transducer, so that the classification efficiency of malicious domain name detection can be accelerated. Meanwhile, the fine tuning process can be suitable for multiple fields, such as the problems of malicious family classification applied to url, url multiple classification, url opposite relation judgment, countermeasure sample generation and the like, and has wide application prospects.
The method effectively combines character level features and mark level features of the domain name through fine tuning. Capturing sequential relation information of the domain name text at a single character level through character features, and intuitively reflecting important information of the domain name text while not depending on semantic understanding; by utilizing the context understanding capability and position sensitivity of the BERT model, the deep semantic information of the domain name is captured through the mark-level features, so that the serious dependence on manual definition rules or features in a purely feature-based method is avoided.
(1) Training the character level embedded feature extraction layer and the mark level embedded feature extraction layer to perform character level embedding;
the character level embedding includes the steps of:
Input sequence representation: first, for an input sequence Each of which is provided withIs the mark of sub-word segmentation,Is the total number of subwords in the sequence, each tagFrom charactersComposition of whichThe length of the subword is indicated,
The total length of the character input is expressed as;
Character vector generation: for each characterAll generate a character embedding vector,
By combining charactersMatrix for embedding charactersMultiplication is done, i.eThe process maps each character to a vector representation in a high-dimensional space;
BiGRU treatment: generated character embedding vector Is fed into BiGRU for processing, the result of BiGRU operation is the generation of hidden state vectors for each character,
Hiding state vectorContextual information including characters;
constructing character level embedding: for each mark The character-level embedded vectors thereof form a sequenceThis sequence is processed into a mark-level embedded vector;
For each markConnecting the hidden state vectors of the first character and the last character to obtain a mark level embedded vectorBy processing the characters of the whole input sequence, the embedded vector of the contextual characters of each mark, namely the embedded vector of the domain name character level, is obtainedThe expression is as follows:
。
the fine tuning data set forms word segmentation sequence after word segmentation In one aspect of the present invention,Through character level embedded feature extraction layer embedding, an embedded vector is obtained; On the other handEmbedded by a mark level embedded feature extraction layer, i.eEmbedding the pre-trained model, adding an attention mask vectorPerforming position embedding processing to obtain a mark level embedded vector with position embedding;
Finally, the marker level is embedded into the vectorWith character level embedded vectorsAre connected together, i.e. inVector end joinVectors forming a longer embedded vector,
。
(2) Setting a feature fusion layer: setting the feature fusion layer as a modelModelModel comprising frozen parameterA layer, a plurality of transducer layers, a plurality of heterogeneous interaction layers, a fully connected layer and a Sigmoid layer; through interaction among the modules, the feature fusion layer can capture the complex relationship between the url character level and the mark level;
training the feature fusion layer:
b1 Model for setting frozen parameters The layer is a pre-trained URL-BERT model,
B2 Setting a transducer module:
the transducer consists of an encoder for processing the input data, and a decoder for generating an output,
The encoder comprises an encoding input layer, a self-attention mechanism, a multi-head attention, a feedforward neural network, residual connection and layer normalization, and an encoder output layer;
the decoder comprises a decoding input layer, a self-attention mechanism, an encoder-decoder attention mechanism, a feedforward neural network, a residual connection and layer normalization and a decoder output layer;
B21 Encoder processing: the input data of the encoded input layer is an embedded vector ;
Self-attention mechanism:
For each head Will inputObtaining a query matrix by linear transformationKey matrixSum matrixWherein, , ,Is the super parameter of the encoder:
,
,
,
next, for each head, an attention weight is calculated, and the scaled dot product attention is used, the output resulting from this step containing a representation of global dependencies inside the input sequence,
Wherein the method comprises the steps ofIs the dimension of the key vector used to scale the dot product, preventing the gradient vanishing problem.Is a matrix of attention weights that,Is the output of the attention mechanism of a head,
Output of multi-head attention:
the outputs of the multiple heads are spliced together and then transformed by another linear transformation, formulated as follows,
,
Wherein, Is the output of each of the attention heads,Is a super parameter matrix;
Feedforward neural network:
The output for the multi-head attention is characterized by a feed-forward neural network, which consists of two linear layers and an activation function,
,
Wherein the method comprises the steps of、、、Is a learnable parameter;
residual connection and layer normalization: for enhancing the training stability and convergence speed of the model,
,
Here, theRepresentation layer normalization operations;
Encoder output layer: coding sequence of final output encoder ;
B22 Decoder processing:
decoding input data of an input layer into embedded vectors ;
Self-attention mechanism: obtaining a query matrix by linear transformation Key matrixSum matrixWherein, , ,Is the super-parameter of the decoder and,
,
,
,
Calculating an attention weight for each head, using a scaled dot product attention;
the outputs of the multiple heads are spliced together, and the self-attention output of the decoder is obtained through residual connection and layer normalization ;
Encoder-decoder attention mechanism: the decoder performs attention calculation on the output of the encoder to capture the relationship between the input sequence and the output sequence, allowing the decoder to adjust its attention according to the information of the encoder to generate a target sequence;
For each position of the decoder Using the output of the decoderAs a query, the output of the encoder、As a key and value, an attention weight is calculated:
,
,
Output of multi-head attention: combining the self-attention output of the decoder with the encoder-decoder attention output to obtain a final attention output;
they are added by element weights, the expression of which is as follows:
,
wherein the final attention output is ,Is a learnable weight parameter:
Feedforward neural network: to the final attention output as Performing a feed forward neural network process to further extract and transform the feature representation:
;
Residual connection and layer normalization: residual connection and layer normalization are applied on the output of the feedforward neural network to enhance training stability and convergence speed of the model:
,
decoder output: the output sequence of the decoder is ;
B3 Setting a heterogeneous interaction module: as shown in fig. 4, the marker-level and character-level representations are treated as two different types of feature representations at the heterogeneous interaction layer, which are combined and separated after each transducer conversion layer;
The fully connected network is responsible for mapping the labels and character representations into the same feature space to ensure that they have consistent dimensions and semantic representations, then connecting and integrating through a CNN layer, and finally performing residual connection and layer normalization operation;
b31 Separation and combination treatment:
vector to be output by decoder Separated by length into mark-level representationsAnd character level representation, A mark/character representing the i-th position, respectively toAndConversion is performed by using a special full connection layer:
,
,
The converted token and character representations are then concatenated to form a fused representation:
;
b32 CNN convolution: for fused representations The CNN operation is applied, for integrating features,
,
Wherein, Representing slaveTo the point ofIs provided with an embedded splice of (a),A window size representing a jth filter;
b33 Separation: the convolved combined representation A characteristic representation separated into two different channels,
,
,
Wherein GELU is an activation function for separating the fused feature representations;
b34 Residual connection and layer normalization: the separated feature representation is added to the initial tag and character representations to reassemble the two channel representations:
,
,
finally, carrying out layer normalization operation on the recombined representation to ensure stability:
,
,
The separated features are connected to obtain the input of the next layer of transducer;
B4 Setting a full connection and Sigmoid classification module:
training of the full connection layer:
Output of last layer of transducer Is transferred to the device withThe full link layer of the hidden units is expressed mathematically as follows
,
Wherein, Is the input vector which is to be used for the input,AndIs the weight and bias of the layer,Is an activation function;
Sigmoid is used for mapping the output of the full connection layer to the category probability distribution, and the Sigmoid layer maps the output result of the full connection layer Passing to Sigmoiod function, the Sigmoid function maps the input real number value to probability value between 0 and1, and the output of the full connection layer isThe output after the Sigmoid function is expressed as:
。
(3) URL-BERT model fine tuning process:
Embedding marker levels of a fine-tuning dataset into vectors With character level embedded vectorsProcessed embedded vectorUsing a pre-trained URL-BERT modelConstructed fine tuning modelBy inputting data into a modelComparing the output result with a real label, calculating a loss function, updating model parameters through a back propagation algorithm, and iteratively training the model until convergence, so that the URL-BERT model after fine adjustment carries out malicious domain name classification.
Fifthly, obtaining a domain name to be detected: obtaining the domain name to be detected.
Sixth, obtaining a malicious domain name detection result: and inputting the domain name to be detected into the URL-BERT model after fine adjustment to obtain a malicious domain name detection result.
The invention provides a malicious domain name detection method based on a BERT model. Compared with the traditional detection method based on the rule and the blacklist, the method can understand the semantic information and the context relation of the domain name in a deeper level, thereby greatly improving the identification accuracy of the malicious domain name. By combining character level embedding and mark level embedding, the method introduces position information, captures information relations among different dimensions, and further enhances the understanding capability of the model on domain name structures; second, the pre-trained BERT model, as a marker-level feature extractor, can take advantage of the rich linguistic knowledge learned in large-scale pre-training data. By introducing the feature fusion layer, the method and the device can flexibly adapt to the specific requirements of the malicious domain name classification task, and provide a high-efficiency and accurate malicious domain name detection scheme. Overall, the method of the invention has the advantages of high efficiency, accuracy, easy implementation and the like, and has positive effects on improving the network security protection level.
The method takes the domain name as a research object, takes the domain name as a special text, can capture the interrelationship between character level characteristics and mark level characteristics of the domain name by means of the strong representation capability of a large-scale language model, digs the characteristics of the domain name from different scales, and provides an innovative solution for malicious domain name detection.
The foregoing has shown and described the basic principles, principal features and advantages of the invention. It will be understood by those skilled in the art that the present invention is not limited to the embodiments described above, and that the above embodiments and descriptions are merely illustrative of the principles of the present invention, and various changes and modifications may be made therein without departing from the spirit and scope of the invention, which is defined by the appended claims. The scope of the invention is defined by the appended claims and equivalents thereof.
Claims (5)
1. The malicious domain name detection method based on the large language model is characterized by comprising the following steps of:
11 Construction of a pre-training dataset and a fine-tuning training dataset: constructing a pre-training data set and a fine-tuning training data set, and performing data preprocessing;
12 Setting a URL-BERT model;
The URL-BERT model setting method comprises the following steps of:
121 The URL-BERT model is set to comprise an input layer, a URL-BERT pre-training layer, a character level embedded feature extraction layer, a mark level embedded feature extraction layer, a feature fusion layer and a fully-connected classification layer;
Setting an input layer: the input layer receives text after word segmentation as the input of a model;
122 Setting a URL-BERT pre-training layer: training a large amount of unmarked domain name data by setting an MLM task in a pre-training stage, and learning and capturing rich semantic information and structural features in domain name texts;
123 Setting a character-level embedded feature extraction layer: the character level embedding extraction layer extracts character level embedding of the domain name text, and bidirectional vector representation is generated by BiGRU, wherein complex association information among characters in the domain name is contained; for the fine tuning training data set, the bi-directional gating unit BiGRU is utilized to perform character level embedding and extraction on the marked domain name data, and BiGRU combines hidden layer states in the forward direction and the backward direction by utilizing GRUs in two different directions, so that the integration of bi-directional information is realized, and the embedded vector of the domain name character level is obtained ;
124 Setting a mark level embedding feature extraction layer to perform mark level embedding;
125 Setting a feature fusion layer):
The feature fusion layer splices the extracted character level embedded feature vector and the mark level embedded feature vector, and then introduces a plurality of Transformer layers to capture complex relations among features;
The heterogeneous interaction module between the transducer layers is used for embedding the character level and the marking level into each transducer layer and then combining and separating, the combination operation enriches the relevance between different representations, the separation operation reserves the independence of the character level and the marking level characteristics, and the representation of the model in the aspect of double-channel distinction is promoted;
126 Setting a full connection classification layer): the full-connection classification layer is used for fine-tuning the whole URL-BERT model to adapt to a specific malicious domain name detection task, and minimizing a loss function through a back propagation algorithm by inputting the feature vector output by the feature fusion layer into the full-connection layer;
13 Pre-training of URL-BERT model: pre-training the URL-BERT model by utilizing a pre-training data set;
The pre-training of the URL-BERT model comprises the following steps:
131 Inputting the text subjected to word segmentation processing of the pre-training data set into a URL-BERT model;
132 URL-BERT pre-training layer Masked Language Model pre-training, i.e., MLM pre-training:
Is provided with Is randomly replaced by [ MASK ]Word sequences of a small number of token,15% Of the token is randomly processed [ MASK ], whereIndicating that the ith token was randomly replaced, the goal of the MLM pre-training process is to predict the contents of the [ MASK ] token based on context;
Will be Is input into the URL-BERT model,Corresponding hidden layer vector, From the following componentsThe code is obtained by the encoder of BERT;
1321 Encoder processing: the input data of the encoded input layer is an embedded vector ;
1322 Self-attention mechanism):
Input device Obtaining a query matrix by linear transformationKey matrix K and value matrixWherein, , Is the super parameter of the encoder:
,
,
,
The attention weight is then calculated, the attention is scaled using the dot product, the output produced in this step contains a representation of global dependencies inside the input sequence,
Wherein the method comprises the steps ofIs the dimension of the key vector, is used for scaling dot products, prevents the gradient vanishing problem,Is a matrix of attention weights that,Is the output of one of the attention heads,
,
Output of multi-head attention:
the outputs of the multiple heads are spliced together and then transformed by another linear transformation, formulated as follows,
,
Wherein, Is the firstThe output of the individual head is provided with,Is a super parameter matrix;
1323 Feedforward neural network):
The output for the multi-head attention is characterized by a feed-forward neural network, which consists of two linear layers and an activation function,
,
Wherein the method comprises the steps of、、、Is a learnable parameter;
and then carrying out residual connection and layer normalization for enhancing training stability and convergence rate of the model:
,
Here, the Representation layer normalization operations;
Encoder output layer: coding sequence of final output encoder ;
Next to this, the process is carried out,Obtaining corresponding token distribution through Sigmoid operation;
the formula of the MLM process is as follows,
,
Wherein the method comprises the steps ofRepresenting a target probability distribution, i.e. the whole sequenceIs a function of the probability distribution of (1),Representing the predicted probability of the ith token given all the tokens of the sequence,The parameters of the model are represented by the parameters,Indicating whether the ith token is [ MASK ], if soOtherwiseThe purpose is to weight the token of [ MASK ] which affects the probability of the whole sequence;
In MLM pre-training, the goal is to maximize the model over the entire sequence Based on this prediction, calculating a cross entropy loss function and updating model parameters by a back propagation algorithmFinally obtaining a pre-trained model,
,
Wherein, Is a parameter of the model and is a parameter of the model,Is a cross-entropy loss function,Is the true label of the ith token, if the token is [ MASK ] i.e. the true value of the token,Is the predictive probability of the ith token, i.eMinimizing loss function by back propagation algorithmUpdating model parametersFinally obtaining a pre-trained URL-BERT model M;
14 Fine tuning of URL-BERT model: utilizing the fine tuning training data set to carry out fine tuning on the pre-trained URL-BERT model;
15 Obtaining a domain name to be detected): obtaining a domain name to be detected;
16 Obtaining malicious domain name detection results: and inputting the domain name to be detected into the URL-BERT model after fine adjustment to obtain a malicious domain name detection result.
2. The method for detecting malicious domain names based on a large language model according to claim 1, wherein the construction of the pre-training data set and the fine-tuning training data set comprises the following steps:
21 A pre-training data set is set, the pre-training data set is unmarked domain name data, and the unmarked domain name data is obtained from a public DNS query log and a WHOIS database source;
22 Setting a fine tuning training data set, wherein the fine tuning training data set is marked domain name data, the marked domain name data is obtained from a network security company or a public malicious domain name database, and the number of positive samples and negative samples is balanced;
23 Pre-processing the pre-training data set and the fine-tuning training data set by the steps of de-duplication, irrelevant characters removal, conversion into lower case letters and standardization;
24 Word segmentation is carried out on the pre-training data set and the fine training data set after pretreatment.
3. The method for detecting malicious domain name based on large language model according to claim 1, wherein the fine tuning of URL-BERT model comprises the steps of:
31 Training the character level embedded feature extraction layer and the mark level embedded feature extraction layer to perform character level embedding;
the fine tuning data set forms word segmentation sequence after word segmentation In one aspect of the present invention,Through character level embedded feature extraction layer embedding, an embedded vector is obtained; On the other handEmbedded by a mark level embedded feature extraction layer, i.eEmbedding the pre-trained model, adding an attention mask vectorPerforming position embedding processing to obtain a mark level embedded vector with position embedding;
Finally, the marker level is embedded into the vectorWith character level embedded vectorsAre connected together, i.e. inVector end joinVectors forming a longer embedded vector,
;
32 Setting a feature fusion layer): setting the feature fusion layer as a modelModelModel comprising frozen parameterA layer, a plurality of transducer layers, a plurality of heterogeneous interaction layers, a fully connected layer and a Sigmoid layer; through interaction among the modules, the feature fusion layer can capture the complex relationship between the url character level and the mark level;
training the feature fusion layer:
321 Model for setting frozen parameters The layer is a pre-trained URL-BERT model,
322 Setting a transducer module:
the transducer consists of an encoder for processing the input data, and a decoder for generating an output,
The encoder comprises an encoding input layer, a self-attention mechanism, a multi-head attention, a feedforward neural network, residual connection and layer normalization, and an encoder output layer;
the decoder comprises a decoding input layer, a self-attention mechanism, an encoder-decoder attention mechanism, a feedforward neural network, a residual connection and layer normalization and a decoder output layer;
3221 Encoder processing: the input data of the encoded input layer is an embedded vector ;
Self-attention mechanism:
For each head Will inputObtaining a query matrix by linear transformationKey matrixSum matrixWherein, , ,Is the super parameter of the encoder:
,
,
,
next, for each head, an attention weight is calculated, and the scaled dot product attention is used, the output resulting from this step containing a representation of global dependencies inside the input sequence,
Wherein the method comprises the steps ofIs the dimension of the key vector, is used for scaling dot products, prevents the gradient vanishing problem,Is a matrix of attention weights that,Is the output of the attention mechanism of a head,
,
Output of multi-head attention:
the outputs of the multiple heads are spliced together and then transformed by another linear transformation, formulated as follows,
,
Wherein, Is the output of each of the attention heads,Is a super parameter matrix;
Feedforward neural network:
The output for the multi-head attention is characterized by a feed-forward neural network, which consists of two linear layers and an activation function,
,
Wherein the method comprises the steps of、、、Is a learnable parameter;
residual connection and layer normalization: for enhancing the training stability and convergence speed of the model,
,
Here, theRepresentation layer normalization operations;
Encoder output layer: coding sequence of final output encoder ;
3222 Decoder processing:
decoding input data of an input layer into embedded vectors ;
Self-attention mechanism: obtaining a query matrix by linear transformation Key matrixSum matrixWherein, , ,Is the super-parameter of the decoder and,
,
,
,
Calculating an attention weight for each head, using a scaled dot product attention;
the outputs of the multiple heads are spliced together, and the self-attention output of the decoder is obtained through residual connection and layer normalization ;
Encoder-decoder attention mechanism: the decoder performs attention calculation on the output of the encoder to capture the relationship between the input sequence and the output sequence, allowing the decoder to adjust its attention according to the information of the encoder to generate a target sequence;
For each position of the decoder Using the output of the decoderAs a query, the output of the encoder、As a key and value, an attention weight is calculated:
,
,
Output of multi-head attention: combining the self-attention output of the decoder with the encoder-decoder attention output to obtain a final attention output;
they are added by element weights, the expression of which is as follows:
,
wherein the final attention output is ,Is a learnable weight parameter:
Feedforward neural network: to the final attention output as Performing a feed forward neural network process to further extract and transform the feature representation:
;
Residual connection and layer normalization: residual connection and layer normalization are applied on the output of the feedforward neural network to enhance training stability and convergence speed of the model:
,
decoder output: the output sequence of the decoder is ;
323 Setting a heterogeneous interaction module: the marker-level and character-level representations are treated as two different types of feature representations in the heterogeneous interaction layer, and after each transducer conversion layer, the two representations are combined and separated;
The fully connected network is responsible for mapping the labels and character representations into the same feature space to ensure that they have consistent dimensions and semantic representations, then connecting and integrating through a CNN layer, and finally performing residual connection and layer normalization operation;
3231 Separation and combination treatment:
vector to be output by decoder Separated by length into mark-level representationsAnd character level representation, A mark/character representing the i-th position, respectively toAndConversion is performed by using a special full-connection network:
,
,
The converted token and character representations are then concatenated to form a fused representation:
;
3232 CNN convolution: for fused representations The CNN operation is applied, for integrating features,
,
Wherein, Representing slaveTo the point ofIs provided with an embedded splice of (a),A window size representing a jth filter;
3233 Separation: the convolved combined representation A characteristic representation separated into two different channels,
,
,
Wherein GELU is an activation function for separating the fused feature representations;
3234 Residual connection and layer normalization: the separated feature representation is added to the initial tag and character representations to reassemble the two channel representations:
,
,
finally, carrying out layer normalization operation on the recombined representation to ensure stability:
,
,
The separated features are connected to obtain the input of the next layer of transducer;
324 Setting a full connection and Sigmoid classification module:
training of the full connection layer:
Output of last layer of transducer Is transferred to the device withThe fully connected layers of the individual hidden units are mathematically expressed as follows:
,
Wherein, Is the input vector which is to be used for the input,AndIs the weight and bias of the layer,Is an activation function;
Sigmoid is used for mapping the output of the full connection layer to the category probability distribution, and the Sigmoid layer maps the output result of the full connection layer Passing to Sigmoiod function, the Sigmoid function maps the input real number value to probability value between 0 and1, and the output of the full connection layer isThe output after the Sigmoid function is expressed as:
;
33 URL-BERT model fine tuning process:
Embedding marker levels of a fine-tuning dataset into vectors With character level embedded vectorsProcessed embedded vectorUsing a pre-trained URL-BERT modelConstructed fine tuning modelBy inputting data into a modelComparing the output result with a real label, calculating a loss function, updating model parameters through a back propagation algorithm, and iteratively training the model until convergence, so that the URL-BERT model after fine adjustment carries out malicious domain name classification.
4. The method for detecting malicious domain names based on a large language model according to claim 2, wherein the word segmentation processing of the pre-processed pre-training data set and the fine training data set comprises the following steps:
41 For the domain name data in the pre-training data set and the fine-tuning training data set, splitting the whole domain name character string according to the structural characteristics, removing special characters in the domain name, and splitting the domain name by taking "/", "-" as separators:
inputting a domain name after pretreatment A word segmentation function f is defined to represent this process:
,
the word segmentation function f () input is a domain name in the form of a preprocessed character string, and the output is a word segmentation list composed of segments segmented by the domain name Wherein, Is thatNumber of medium word content; the expression is as follows:
,
after preliminary word segmentation, the BERT word segmentation device pair is utilized Performing embedding treatment;
42 Word embedding processing: in order to identify the beginning and end of an input sequence, in Adding [ CLS ] mark at the beginning, adding [ SEP ] mark at the end, and limiting the token sequence length to 128, for input sequences with a length of less than 128 marks, using a specific PAD mark [ PAD ] to extend the sequence length until it reaches the limit of 128 marks, thereby obtaining;
43 Segment embedding processing is performed: the segment embeddings are all set to 0, i.e.:
Added to the same length of [0, …,0 ];
44 Position embedding processing: position embedding is based on The different position information of the token encodes a position vector, the position relation of the token is introduced into the URL-BERT model, and the position embedding is set to 0, namely:
Adding the sequence subjected to segment embedding treatment with the sequence with the same length of 0;
45 Generating an attention mask vector: the attention mask vector is used to distinguish the actual markers in the input sequence from the filled markers,
For each real marker in the sequence, its corresponding attention is set to 1; for each PAD, its corresponding attention is set to 0; generating an attention mask directionWherein the number of 1 isNamely:
the sequence subjected to the position embedding processing is added to the sequence [1, …,0 ].
5. A method for detecting malicious domain names based on a large language model according to claim 3, wherein the character level embedding comprises the steps of:
51 Input sequence representation: first, for an input sequence Each of which is provided withIs the mark of sub-word segmentation,Is the total number of subwords in the sequence, each tagFrom charactersComposition of whichThe length of the subword is indicated,
The total length of the character input is expressed as;
52 Character vector generation: for each characterAll generate a character embedding vector,
By combining charactersMatrix for embedding charactersMultiplication is done, i.eThe process maps each character to a vector representation in a high-dimensional space;
53 BiGRU) treatment: generated character embedding vector Is fed into BiGRU for processing, the result of BiGRU operation is the generation of hidden state vectors for each character,
Hiding state vectorContextual information including characters;
54 Building character level embedding: for each mark The character-level embedded vectors thereof form a sequenceThis sequence is processed into a mark-level embedded vector;
For each markConnecting the hidden state vectors of the first character and the last character to obtain a mark level embedded vectorBy processing the characters of the whole input sequence, the embedded vector of the contextual characters of each mark, namely the embedded vector of the domain name character level, is obtainedThe expression is as follows:
。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202410875232.5A CN118413402B (en) | 2024-07-02 | 2024-07-02 | Malicious domain name detection method based on large language model |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202410875232.5A CN118413402B (en) | 2024-07-02 | 2024-07-02 | Malicious domain name detection method based on large language model |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| CN118413402A CN118413402A (en) | 2024-07-30 |
| CN118413402B true CN118413402B (en) | 2024-09-06 |
Family
ID=92003295
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| CN202410875232.5A Active CN118413402B (en) | 2024-07-02 | 2024-07-02 | Malicious domain name detection method based on large language model |
Country Status (1)
| Country | Link |
|---|---|
| CN (1) | CN118413402B (en) |
Families Citing this family (12)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN119071023B (en) * | 2024-08-02 | 2025-08-22 | 清华大学 | A method and device for detecting malicious traffic and countering attacks |
| CN118972123A (en) * | 2024-08-05 | 2024-11-15 | 威海蓝海银行股份有限公司 | Identification method of malicious domain name access based on large language model |
| CN118585632B (en) * | 2024-08-06 | 2024-12-03 | 浪潮云信息技术股份公司 | Large model question-answering method, device, equipment and storage medium for traffic industry |
| CN118657439A (en) * | 2024-08-19 | 2024-09-17 | 宁波宁途智能科技有限公司 | Industrial production data quality management method and device based on SPC detection |
| CN118675073B (en) * | 2024-08-23 | 2024-12-27 | 牡丹区公路事业发展中心 | A highway slope displacement prediction method |
| CN119276539B (en) * | 2024-09-09 | 2025-09-09 | 鹏城实验室 | Malicious domain name identification method, malicious domain name identification device, computer equipment and readable storage medium |
| CN119675890A (en) * | 2024-10-15 | 2025-03-21 | 北京亚鸿世纪科技发展有限公司 | A method for identifying malicious domain names based on search engines and generative language models |
| CN119720193A (en) * | 2024-11-25 | 2025-03-28 | 交通银行股份有限公司 | Program detection method, device, computer equipment and storage medium |
| CN119204232B (en) * | 2024-11-28 | 2025-04-22 | 齐鲁工业大学(山东省科学院) | A method to improve the accuracy of language processing models |
| CN119649599B (en) * | 2024-12-12 | 2025-12-09 | 南京邮电大学 | Urban traffic flow prediction method and system integrating space-time characteristics and large language model |
| CN119357965A (en) * | 2024-12-25 | 2025-01-24 | 浙江工业大学 | A malicious URL detection method based on rule and deep learning dual drive |
| CN120277500B (en) * | 2025-06-10 | 2025-08-12 | 苏州工学院 | Malicious URL detection method based on fusion of character-level language model and structural features |
Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN118054951A (en) * | 2024-02-26 | 2024-05-17 | 东北大学 | A malicious domain name detection method based on unsupervised learning |
Family Cites Families (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP4383686A1 (en) * | 2019-04-02 | 2024-06-12 | Bright Data Ltd. | System and method for managing non-direct url fetching service |
| EP4377817A4 (en) * | 2021-07-26 | 2025-05-28 | Bright Data Ltd. | WEB BROWSER EMULATION IN A DEDICATED MIDDLE BOX |
| CN117675238A (en) * | 2022-08-23 | 2024-03-08 | 腾讯云计算(北京)有限责任公司 | Data access methods, devices, electronic equipment and storage media |
| CN116167084B (en) * | 2023-02-24 | 2026-01-02 | 北京工业大学 | A privacy-preserving method and system for training federated learning models based on a hybrid strategy. |
| CN116633623B (en) * | 2023-05-24 | 2025-07-04 | 北京邮电大学 | A DGA domain name detection method for IoT botnet |
| CN116385808B (en) * | 2023-06-02 | 2023-08-01 | 合肥城市云数据中心股份有限公司 | Big data cross-domain image classification model training method, image classification method and system |
-
2024
- 2024-07-02 CN CN202410875232.5A patent/CN118413402B/en active Active
Patent Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN118054951A (en) * | 2024-02-26 | 2024-05-17 | 东北大学 | A malicious domain name detection method based on unsupervised learning |
Also Published As
| Publication number | Publication date |
|---|---|
| CN118413402A (en) | 2024-07-30 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN118413402B (en) | Malicious domain name detection method based on large language model | |
| CN113596007B (en) | A deep learning-based vulnerability attack detection method and device | |
| CN114444516B (en) | Cantonese rumor detection method based on deep semantic perception map convolutional network | |
| CN111459491B (en) | A code recommendation method based on tree neural network | |
| CN115033670A (en) | Cross-modal image-text retrieval method with multi-granularity feature fusion | |
| CN112183064B (en) | Text emotion reason recognition system based on multi-task joint learning | |
| CN115759092A (en) | Network threat information named entity identification method based on ALBERT | |
| CN115329776B (en) | Semantic analysis method for network security co-processing based on less-sample learning | |
| CN113705218A (en) | Event element gridding extraction method based on character embedding, storage medium and electronic device | |
| CN114416479A (en) | Log sequence anomaly detection method based on out-of-stream regularization | |
| CN115169285A (en) | A method and system for event extraction based on graph parsing | |
| CN112668013A (en) | Java source code-oriented vulnerability detection method for statement-level mode exploration | |
| CN117688560A (en) | An intelligent detection method for malware based on semantic analysis | |
| CN114169329A (en) | Named entity identification method oriented to information security field | |
| CN115422945A (en) | A rumor detection method and system integrating emotion mining | |
| CN120150993A (en) | SQL injection detection method, system and storage medium based on syntactic and semantic feature fusion network | |
| CN117332411B (en) | Abnormal login detection method based on transducer model | |
| CN117390189B (en) | Neutral text generation method based on pre-classifier | |
| CN120146051A (en) | Multimodal entity and relationship extraction method and system based on cross-modal alignment and fusion | |
| CN115242539A (en) | Network attack detection method and device for power grid information system based on feature fusion | |
| CN120105149A (en) | A network threat intelligence analysis method and system | |
| CN113886593A (en) | Method for improving relation extraction performance by using reference dependence | |
| CN113342982B (en) | Enterprise industry classification method integrating Roberta and external knowledge base | |
| CN120277500A (en) | Malicious URL detection method based on fusion of character-level language model and structural features | |
| CN117010382A (en) | A method and system for named entity recognition in the field of network security based on auxiliary vectors |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PB01 | Publication | ||
| PB01 | Publication | ||
| SE01 | Entry into force of request for substantive examination | ||
| SE01 | Entry into force of request for substantive examination | ||
| GR01 | Patent grant | ||
| GR01 | Patent grant |