WO2024164828A1 - 数据流识别方法及装置、电子设备 - Google Patents
数据流识别方法及装置、电子设备 Download PDFInfo
- Publication number
- WO2024164828A1 WO2024164828A1 PCT/CN2024/073409 CN2024073409W WO2024164828A1 WO 2024164828 A1 WO2024164828 A1 WO 2024164828A1 CN 2024073409 W CN2024073409 W CN 2024073409W WO 2024164828 A1 WO2024164828 A1 WO 2024164828A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- data flow
- data
- data stream
- flow
- nat
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L41/00—Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks
- H04L41/06—Management of faults, events, alarms or notifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L61/00—Network arrangements, protocols or services for addressing or naming
- H04L61/09—Mapping addresses
- H04L61/25—Mapping addresses of the same type
- H04L61/2503—Translation of Internet protocol [IP] addresses
- H04L61/2514—Translation of Internet protocol [IP] addresses between local and global IP addresses
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L43/00—Arrangements for monitoring or testing data switching networks
- H04L43/02—Capturing of monitoring data
- H04L43/026—Capturing of monitoring data using flow identification
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L61/00—Network arrangements, protocols or services for addressing or naming
- H04L61/09—Mapping addresses
- H04L61/25—Mapping addresses of the same type
- H04L61/2503—Translation of Internet protocol [IP] addresses
- H04L61/2557—Translation policies or rules
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L61/00—Network arrangements, protocols or services for addressing or naming
- H04L61/09—Mapping addresses
- H04L61/25—Mapping addresses of the same type
- H04L61/2503—Translation of Internet protocol [IP] addresses
- H04L61/256—NAT traversal
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L43/00—Arrangements for monitoring or testing data switching networks
- H04L43/10—Active monitoring, e.g. heartbeat, ping or trace-route
- H04L43/106—Active monitoring, e.g. heartbeat, ping or trace-route using time related information in packets, e.g. by adding timestamps
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L43/00—Arrangements for monitoring or testing data switching networks
- H04L43/12—Network monitoring probes
Definitions
- the present application relates to the field of network technology, and in particular to a data flow identification method and device, and an electronic device.
- NAT network address translation
- IP Internet protocol
- the relevant technology can read the NAT session table information of a device with NAT function enabled (hereinafter referred to as NAT device), obtain the correspondence between two data flows before and after NAT conversion according to the NAT session table information, determine that a data flow belongs to the same NAT session according to the correspondence between the data flows, and then consider the data flows belonging to the same NAT session at the same time when performing fault diagnosis.
- NAT device a device with NAT function enabled
- the present application provides a data flow identification method and device, and an electronic device, which can reduce the difficulty and cost of determining a data flow related to a first data flow NAT, and reduce the difficulty of network deployment.
- the present application provides a data flow identification method, which includes: obtaining a fingerprint feature and an establishment time of a first data flow; and determining at least one data flow related to the first data flow NAT according to the fingerprint feature and the establishment time of the first data flow.
- At least one data flow related to the first data flow NAT is determined by the fingerprint feature of the first data flow, without reading the session table from the NAT device, with low implementation difficulty and cost, and low network deployment difficulty.
- a data stream related to the first data stream NAT will be generated in the first-level NAT, and multiple data streams related to the first data stream NAT will be generated in the multi-level NAT. Therefore, one or more data streams related to the first data stream NAT can be determined.
- being related to the first data stream NAT may mean belonging to the same NAT session as the first data stream, or being related to the first data stream NAT may mean belonging to the same NAT session as the data stream after the first data stream undergoes N-level NAT conversion.
- the method can be executed by an analysis platform, network equipment, collector, storage platform, etc. in the network.
- Collectors i.e., probes
- the analysis platform executes the data flow identification method through the fingerprint features of the data flow collected by the collector.
- the collector is usually implemented by software. In other implementations, the collector may also be implemented by hardware.
- collectors can be deployed upstream and downstream of the node that enables the NAT function.
- other methods can also be used to deploy collectors. As an example, three methods of deploying collectors are listed below.
- the first method is to deploy a collector on each node to ensure that the fingerprint characteristics of the data flow of each NAT session can be collected.
- the second method is to deploy collectors on multiple core nodes of the network, such as multiple spine switches, multiple border nodes, and/or firewalls, so as to collect fingerprint features of the data flow of each NAT session as much as possible.
- the third method is to deploy a collector on a node with NAT function to collect fingerprint features of the data flow of each NAT session.
- the collector In the first and second deployment modes, the collector is usually deployed in the inbound direction of the node interface, and can also be deployed in the outbound direction of the interface. In the third deployment mode, the collector needs to be deployed in both the inbound and outbound directions of the interface.
- each node where a collector (software) is deployed or the collector (hardware) stores the information of the data stream collected by each node.
- the network may further include a storage platform for centrally storing information of data streams collected by various collectors.
- the analysis platform interacts with the collector or storage platform to obtain the data flow information and further determine at least one data flow related to the first data flow NAT.
- the storage platform determines at least one data flow related to the first data flow NAT through the locally stored data flow information.
- the first one is that the analysis platform implements this method by interacting with the collector or storage platform:
- obtaining the fingerprint feature and establishment time of the first data stream includes:
- a first query response sent by a first network node is received, where the first query response includes a fingerprint feature and an establishment time of the first data flow.
- the first network node here may be a collector, a network device with a collector deployed, or a storage platform. Furthermore, when sending a query request, if it is sent to a collector or a network device with a collector deployed, it may be sent to multiple collectors or network devices with collectors deployed at the same time.
- the first flow identifier of the data flow is a five-tuple of the data flow, including a source Internet Protocol (IP) address, a destination IP address, a source port, a destination port, and a protocol type, where the protocol type is a transmission control protocol (TCP).
- IP Internet Protocol
- TCP transmission control protocol
- the first network node after receiving the query request, uses the first flow identifier as an index to perform a query to obtain the fingerprint feature and establishment time of the first data flow.
- the first query request also includes a query time range, and the query time range is used to limit the time range within which the establishment time of the queried data stream falls.
- the query time range is expressed in an exact range, such as 0:00 to 24:00 on a certain day, or 0:00 to 60:00 on a certain hour on a certain day, etc.
- the query time range is expressed in time granularity, such as one day, then the corresponding range is the current day, or one hour, then the corresponding range is within the current hour, etc.
- the query time range may adopt a default value, such as querying the current day.
- the fingerprint feature of the data stream may include one or more features.
- the fingerprint feature includes at least one of the following features:
- IP identifier IP identifier
- IPID IP identifier
- SYN synchronize sequence numbers
- ISN initial sequence number
- the first data packet refers to the first data packet transmitted after the TCP three-way handshake establishes a connection.
- Different data streams need to transmit different content. Some data streams require multiple data packets to complete the content transmission, while some data streams only need to transmit one data packet to complete the content transmission. Therefore, the embodiment of the present application uses the information of the first data packet as a fingerprint feature, and can collect fingerprint features for each data stream that needs to be judged.
- the first data packet is the first data packet transmitted after TCP establishes the link, and its sequence number is equal to the ISN+1 carried by the SYN message. The collector can easily determine the first data packet. Therefore, the resources and The time is less than that of subsequent packets.
- the collector can determine whether it is the first data packet based on the sequence number of the data packet.
- the sequence number of the first data packet is the ISN+1 of the SYN message.
- the sequence number of the second data packet is the ISN+1 of the SYN message plus the length of the first data packet. Therefore, if you want to collect the features of the data packets after the first data packet, it is more difficult to determine the sequence number of the data packet, which takes up more resources and time. Therefore, the first data packet is preferred for fingerprint feature collection.
- the payload of the first data packet can be the entire payload of the first data packet, or a portion of the payload of the first data packet, for example, a portion of the payload intercepted from the packet header, such as a payload of 500 or 1000 bytes in length. Accordingly, the hash value of the first data packet can also be the hash value of the entire payload, or the hash value of a portion of the payload, which will not be described in detail here.
- each collector collects the same fingerprint features.
- the applicant has found through theoretical analysis and a large number of experiments that the fingerprint features listed above will not change before and after NAT conversion, and different data streams are unlikely to conflict, that is, although the two data streams before and after NAT conversion may have different IP addresses or ports, the above fingerprint features in the two data streams are the same, so the above fingerprint features can identify NAT-related data streams.
- the fingerprint feature collection process listed above is relatively easy to implement, occupies less resources, and is conducive to the implementation of the method provided in the embodiment of the present application.
- the fingerprint features of the data flow may also include other features, as long as they remain unchanged before and after NAT conversion and different data flows are unlikely to conflict, and this application does not impose any restrictions on this.
- the creation time of the data stream collected by the collector may be the time of receiving or sending a certain data packet or message.
- the collector can use one of the following times as the creation time of the data stream:
- SYN message receiving time SYN-ACK message receiving time
- ACK message receiving time The ACK message here is the ACK message in the TCP three-way handshake.
- the above is for the collector in the inbound direction of the interface. If it is a collector in the outbound direction of the interface, one of the following times is used as the creation time of the data flow:
- the collector may also use other times, such as the time when the first data packet is received or sent, when collecting the creation time of the data stream.
- the manner of determining at least one data flow related to the first data flow NAT includes multiple methods:
- the analysis platform uses the fingerprint feature of the first data flow to query the second data flow with the same fingerprint feature, and then determines whether NAT is related based on the creation time of the two.
- the steps are as follows:
- the second query response including a first flow identifier and an establishment time of a second data flow having the same fingerprint feature as the first data flow;
- the second network node and the first network node may be the same network node or different network nodes.
- the first network node and the second network node may be a plurality of distributed collectors or network devices deployed with collectors.
- the creation time of the second data stream with the same fingerprint feature is queried for the second time. Then, the NAT-related data streams are judged based on the same fingerprint feature and similar creation time as the judgment criteria. Since there is a time delay in NAT conversion and the transmission of data streams between different collectors, there is a time delay in the establishment time of the data stream collected before NAT conversion and the establishment time of the data stream collected after NAT conversion. By limiting the absolute value of the time difference of the establishment time of the two data streams to be less than a threshold, it can be determined that the two data streams with the same fingerprint feature are two data streams converted by NAT.
- the above threshold can be determined based on time synchronization accuracy and/or transmission delay.
- the first query response also includes the first Flow ID and creation time; at this time, the analysis platform directly determines whether NAT is related based on the creation time of the two data flows.
- the steps are as follows:
- the first network node when the analysis platform sends the first query request, can not only feed back the information of the first data flow to the analysis platform, but also simultaneously query the information of the second data flow with the same fingerprint as the first data flow, and feed it back to the analysis platform together. In this way, the analysis platform can directly determine whether the two data flows are NAT related based on the content in the first query response.
- the second data flow may be one or more. If there are more than one, it is determined whether each second data flow is related to the first data flow NAT.
- the first network device may be a storage platform.
- the second is the case where the storage platform implements this method based on the collected information stored locally:
- obtaining the fingerprint feature and establishment time of the first data stream includes:
- a first query request is received, wherein the first query request includes a first flow identifier of a first data flow; and a fingerprint feature and an establishment time of the first data flow are determined according to the first flow identifier of the first data flow.
- the method is executed by the storage platform, and the storage platform can query the fingerprint feature and establishment time of the first data flow according to the flow identifier in the query request.
- the step of determining at least one data flow related to the first data flow NAT according to the fingerprint feature and the establishment time of the first data flow includes:
- the establishment time of the second data stream may be obtained by searching the fingerprint feature of the first data stream to find the second data stream with the same fingerprint feature.
- fingerprint conflicts can also be excluded in the above process. Fingerprint conflicts refer to the same fingerprint features of non-NAT-related data flows.
- determining that the first data flow and the second data flow are NAT-related includes:
- determining that the fingerprints of the first data stream and the second data stream conflict may be performed through the following steps:
- the second flow identifier of the data flow may be a four-tuple, including a source IP address, a destination IP address, a destination port, and a protocol type.
- the principle used is to query whether a data flow with the same four-tuple as the data flow has been NAT-converted in the historical records. If a data flow with the same four-tuple has been NAT-converted, it means that this is also a NAT conversion, not a fingerprint conflict.
- the first data stream may be a data stream selected by a user, and when the user observes that the first data stream has a fault, the user determines the data stream related to the first data stream NAT.
- the first data stream may also be automatically selected by the diagnostic device, and the diagnostic device detects the fault of the first data stream and then determines the data stream related to the first data stream NAT.
- the method further includes:
- fault diagnosis is performed on the first data flow and at least one data flow related to the first data flow NAT.
- the method further includes:
- a fault diagnosis is performed on the first data stream.
- the data flow failure refers to the failure of the application based on the data flow, including but not limited to high data flow packet loss rate, large delay, etc.
- the present application provides a data flow identification method.
- the method includes: receiving a query request, the query request includes a first flow identifier of a first data flow; sending a query response, the query response includes a fingerprint feature and an establishment time of the first data flow, and the fingerprint feature and the establishment time of the first data flow are used to determine at least one data flow related to the first data flow NAT.
- the difference between the method provided in the second aspect and the method provided in the first aspect is that the method provided in the second aspect only queries and feedbacks the fingerprint features of the data flow, but does not judge at least one data flow related to the first data flow NAT.
- the method may be executed by a collector, a network device equipped with a collector, or a storage platform.
- the query response further includes a first flow identifier and an establishment time of the second data flow having the same fingerprint feature as the first data flow.
- the method further comprises:
- the establishment time, first stream identifier and fingerprint features of each data stream are stored.
- the collector is connected in a bypass mode. At this time, the establishment time, first stream identifier and fingerprint features of each data stream are collected, including:
- the message of the data stream is captured by the mirroring function, including:
- ACL access control lists
- ACL By configuring ACL to collect the characteristics of SYN packets and/or SYN-ACK packets, you can reduce the number of packets that need to be captured and reduce the collection burden.
- the collector is connected in a path-dependent manner.
- the data stream message is directly analyzed to obtain the establishment time, the first stream identifier and the fingerprint feature of the data stream.
- a data flow identification device comprising:
- An acquisition unit used to acquire the fingerprint feature and establishment time of the first data stream
- the determination unit is used to determine at least one data flow related to the first data flow NAT according to the fingerprint feature and establishment time of the first data flow.
- the acquisition unit is used to send a first query request to the first network node, the first query request including a first flow identifier of the first data flow; and receive a first query response sent by the first network node, the first query response including a fingerprint feature and an establishment time of the first data flow.
- the determination unit is used to send a second query request to the second network node, the second query request including the fingerprint feature of the first data flow; receive a second query response sent by the second network node, the second query response including the first flow identifier and establishment time of the second data flow that are the same as the fingerprint feature of the first data flow; and determine that the first data flow and the second data flow are NAT-related when the absolute value of the time difference between the establishment time of the second data flow and the establishment time of the first data flow is less than a threshold.
- the first query response further includes a first stream identifier and an establishment time of a second data stream having the same fingerprint feature as the first data stream;
- the determining unit is configured to determine that the first data flow and the second data flow are NAT-related when the absolute value of the time difference between the establishment time of the second data flow and the establishment time of the first data flow is less than a threshold.
- the acquisition unit is configured to receive a first query request, the first query request including a first stream identifier of the first data stream; and determine a fingerprint feature and an establishment time of the first data stream according to the first stream identifier of the first data stream.
- the determination unit is used to determine the second data stream based on the fingerprint characteristics of the first data stream, and the fingerprint characteristics of the second data stream are the same as the fingerprint characteristics of the first data stream; when the absolute value of the time difference between the establishment time of the second data stream and the establishment time of the first data stream is less than a threshold, it is determined that the first data stream and the second data stream are NAT related.
- the fingerprint feature includes at least one of the following features:
- TCP Transmission Control Protocol
- ISN Initial Sequence Code
- the first query request further includes a query time range, and the query time range is used to limit the time range within which the establishment time of the queried data flow falls.
- the determination unit is used to determine that the first data stream and the second data stream are NAT-related when the absolute value of the time difference between the establishment time of the second data stream and the establishment time of the first data stream is less than a threshold and there is no fingerprint conflict between the first data stream and the second data stream, and the fingerprint conflict means that the fingerprint features of non-NAT-related data streams are the same.
- the determination unit is also used to determine at least one third data stream having the same second stream identifier as the second data stream; determine at least one fourth data stream having the same fingerprint feature as each third data stream; and determine that there is no fingerprint conflict between the first data stream and the second data stream when, among the at least one fourth data stream, there is a fourth data stream whose establishment time and the corresponding third data stream have an absolute value of a time difference less than a threshold, and the second stream identifier of the fourth data stream is the same as the second stream identifier of the first data stream.
- the device further comprises:
- the diagnosis unit is used to perform fault diagnosis on the first data flow and at least one data flow related to the first data flow NAT when a fault occurs in the first data flow.
- a data flow identification device comprising:
- a receiving unit configured to receive a query request, wherein the query request includes a first stream identifier of a first data stream;
- the sending unit is used to send a query response, the query response includes the fingerprint feature and establishment time of the first data flow, the fingerprint feature and establishment time of the first data flow are used to determine at least one data flow related to the first data flow NAT.
- the query response further includes a first stream identifier and an establishment time of a second data stream having the same fingerprint feature as the first data stream.
- the device further comprises:
- a collection unit used to collect the establishment time, first stream identifier and fingerprint features of each data stream
- the storage unit is used to store the establishment time, first stream identifier and fingerprint characteristics of each data stream.
- the acquisition unit is used to capture the packets of the data stream through a mirroring function; analyze the captured packets of the data stream to obtain the establishment time, the first stream identifier and the fingerprint feature of the data stream.
- an electronic device in a fifth aspect, includes a processor and a memory.
- the memory is used to store software programs and modules.
- the processor implements the method in the first aspect or any possible implementation of the first aspect, or implements the method in the second aspect or any possible implementation of the second aspect by running or executing the software program and/or module stored in the memory.
- the number of the processors is one or more, and the number of the memories is one or more.
- the memory may be integrated with the processor, or the memory may be provided separately from the processor.
- the memory can be a non-transitory memory, such as a read-only memory (ROM), which can be integrated with the processor on the same chip or can be set on different chips.
- ROM read-only memory
- a data flow identification system comprising a collector and an analysis platform, the analysis platform being used to execute the method in the above-mentioned first aspect or any possible implementation of the first aspect, and the collector being used to execute the method in the above-mentioned second aspect or any possible implementation of the second aspect.
- the data flow identification system further includes a storage platform for storing information of the data flow collected by the collector.
- a computer program product includes a computer program code, and when the computer program code is executed by a computer, the computer executes the method in the first aspect or any possible implementation of the first aspect. Or execute the method in the above second aspect or any possible implementation manner of the second aspect.
- the present application provides a computer-readable storage medium, which is used to store program codes executed by a processor, wherein the program code includes a method for implementing any possible implementation of the above-mentioned first aspect, or a method for implementing the above-mentioned second aspect or any possible implementation of the second aspect.
- a chip comprising a processor, the processor being used to call and execute instructions stored in a memory from the memory, so that a communication device equipped with the chip executes a method in any possible implementation of the first aspect above, or executes a method in any possible implementation of the second aspect or the second aspect above.
- another chip in the tenth aspect, another chip is provided.
- the another chip includes an input interface, an output interface, a processor and a memory.
- the input interface, the output interface, the processor and the memory are connected via an internal connection path.
- the processor is used to execute the code in the memory.
- the processor is used to execute the method in any possible implementation of the first aspect above, or execute the second aspect or the method in any possible implementation of the second aspect.
- FIG1 is a schematic diagram of a structure of an application scenario provided by an embodiment of the present application.
- FIG2 is a schematic diagram of the structure of another application scenario provided by an embodiment of the present application.
- FIG3 is a schematic diagram of a network topology of a data center provided in an embodiment of the present application.
- FIG4 is a flow chart of a data flow identification method provided in an embodiment of the present application.
- FIG5 is a flow chart of a data flow identification method provided in an embodiment of the present application.
- FIG6 is a flow chart of a method for collecting data stream information provided in an embodiment of the present application.
- FIG7 is a flow chart of a data flow identification method provided in an embodiment of the present application.
- FIG8 is a flow chart of a data flow identification method provided in an embodiment of the present application.
- FIG9 is a flow chart of a data flow identification method provided in an embodiment of the present application.
- FIG10 is a block diagram of a data flow identification device provided in an embodiment of the present application.
- FIG11 is a block diagram of a data flow identification device provided in an embodiment of the present application.
- FIG. 12 shows a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
- FIG1 is a schematic diagram of a structure of an application scenario provided by an embodiment of the present application.
- the application scenario includes multiple network devices 11, and the multiple network devices 11 constitute a network, such as a data center network.
- the network device 11 may be a switch, a routing device, a firewall device, a server, etc.
- the NAT function is deployed and enabled on at least some of the network devices 11. As shown in Fig. 1, a collector 12 is deployed on at least some of the network devices 11 among the multiple network devices 11 to collect information of data flows.
- Fig. 2 is a schematic diagram of another application scenario provided by an embodiment of the present application. Referring to Fig. 2, the difference between this application scenario and Fig. 1 is that it also includes a storage platform 13, which is connected to each collector 12 at the same time and is used to centrally store the information of the data streams collected by each collector.
- the collector and the storage platform can both provide a data query interface to the outside.
- an analysis platform 14 may also be included, and the analysis platform 14 is connected to each collector 12 at the same time, or is connected to the storage platform 13, and determines the NAT-related data flow by analyzing the information of the data flow.
- the storage platform and the analysis platform may be integrated into the same device, for example, the analysis platform includes a storage module, or the storage platform includes an analysis module, etc.
- a network device, a collector, a storage platform or an analysis platform can determine NAT-related data flows through information about the data flows.
- Fig. 3 is a schematic diagram of a network topology of a data center provided in an embodiment of the present application.
- the data center includes a spine switch 14 and a virtualized server cluster 15, and the spine switch 14 is connected to the virtualized server cluster 15 respectively.
- the data center may further include an analyzer network 16, which includes a collector cluster 161, an analyzer cluster 162, and The Leaf switch 163 and the Spine switch 14 are respectively connected to the collector cluster 161 and the analyzer cluster 162 through the Leaf switch 163.
- the collector cluster 161 includes the aforementioned storage platform
- the analyzer cluster 162 includes the aforementioned analysis platform.
- the analyzer network 16 may also be independent of the data center, and accordingly, the Leaf switch 163 may be other network devices.
- the spine switch 14 and the virtualized server cluster 15 are both network devices 11, and the collectors 12 are deployed on the spine switch 14 and the network devices 11 between the spine switch 14 and the virtualized server cluster 15.
- the three collectors 12 collect the information of the data stream and then upload it to the storage platform of the collector cluster 161 through the path shown by the dotted line.
- the data center will deploy a firewall (firewall, FW) (not shown in the figure) at the border and enable the NAT function to protect the network of the data center.
- FW firewall
- collectors in order to ensure that the collected fingerprint features can be used to identify whether the data flow is NAT-related, collectors can be deployed both upstream and downstream of the node that enables the NAT function.
- other methods can also be used to deploy the collector. As an example, three methods of deploying the collector are listed below.
- the first method is to deploy a collector on each node to ensure that the fingerprint characteristics of the data flow of each NAT session can be collected.
- the second method is to deploy collectors on multiple core nodes of the network, such as multiple Spine switches, multiple border nodes, and/or firewalls, so as to collect the fingerprint characteristics of the data flow of each NAT session as much as possible.
- the above-mentioned core nodes are deployed on both sides of the device with NAT function enabled, so deploying collectors on these core nodes can collect the fingerprint characteristics of the data flow of each NAT session.
- the firewall with NAT function enabled is usually connected to two Spine switches. The data flow first passes through a Spine switch, then performs NAT conversion in the firewall, and then passes through another Spine switch. Therefore, deploying collectors on two Spine switches can collect the fingerprint characteristics of the data flow before and after NAT conversion.
- the third method is to deploy a collector on a node with NAT function to collect fingerprint features of the data flow of each NAT session.
- the collector In the first and second deployment modes, the collector is usually deployed in the inbound direction of the node interface, and can also be deployed in the outbound direction of the interface. In the third deployment mode, the collector needs to be deployed in both the inbound and outbound directions of the interface.
- each node where a collector (software) is deployed or the collector (hardware) stores the information of the data stream collected by each node.
- the collector sends the collected data stream information to the storage platform, which stores it centrally.
- Figure 4 is a flow chart of a data flow identification method provided by an embodiment of the present application.
- the method can be executed by an analysis platform in the application scenarios shown in Figures 1 to 3, and the analysis platform is connected to each collector, or connected to a storage platform to obtain data flow information to implement the method, or the analysis platform obtains the data flow information stored in itself to implement the method.
- the method can also be executed by a network device in the application scenarios shown in Figures 1 to 3. The following is an example of the method being executed by an analysis platform. As shown in Figure 4, the method includes the following steps.
- S11 Obtain fingerprint features and creation time of the first data stream.
- the fingerprint feature includes at least one of the following features:
- the fingerprint feature is any one of the above features, for example, the fingerprint feature is the IPID of the first data packet, or the hash value of the payload of the first data packet, or the ISN of the SYN message. Using a single field as the fingerprint feature reduces the computation and storage burden.
- the fingerprint feature includes multiple of the above features, for example, the fingerprint feature includes the ISN of the SYN message + the IPID of the first data packet, or the payload of the first data packet + the ISN of the SYN message, or the payload of the first data packet + the ISN of the SYN message + the IPID of the first data packet, etc.
- the fingerprint feature includes the ISN of the SYN message + the IPID of the first data packet, or the payload of the first data packet + the ISN of the SYN message, or the payload of the first data packet + the ISN of the SYN message + the IPID of the first data packet, etc.
- IPID is a field used to identify IP data packets in the IP layer. It is used to identify each IP message sent by the host. It is 16 bits long. ISN is the initial sequence code carried by the SYN message in the three-way handshake process of TCP link establishment and used to notify the other end. It is 32 bits long.
- the payload of the data packet is the application layer content carried in the data packet (that is, the content after the TCP header in the data packet).
- the last SYN message or data packet received during the retransmission process is used as the collection object to collect fingerprint features to ensure the consistency of fingerprint features collected by each collector.
- the collector obtains fingerprint features from the last received SYN message or the last received first packet.
- the first data packet refers to the first data packet transmitted after the TCP three-way handshake establishes a connection.
- some data streams require multiple data packets to complete the transmission of content, while some data streams only need to transmit one data packet to complete the transmission of content. Therefore, the embodiment of the present application uses the information of the first data packet as the fingerprint feature, which can ensure that the fingerprint feature can be collected for each data stream that needs to be judged.
- the first data packet is the first data packet transmitted after TCP establishes the link, and its sequence number is equal to the ISN+1 carried by the SYN message. The collector can easily determine the first data packet. Therefore, the resources and time required to collect the first data packet are less than those of subsequent data packets.
- the collector can determine whether it is the first data packet based on the sequence number of the data packet.
- the sequence number of the first data packet is the ISN+1 of the SYN message.
- the sequence number of the second data packet is the ISN+1 of the SYN message plus the length of the first data packet. Therefore, if you want to collect the features of the data packets after the first data packet, it is more difficult to determine the sequence number of the data packet, which takes up more resources and time. Therefore, the first data packet is preferred for fingerprint feature collection.
- the payload of the first data packet can be the entire payload of the first data packet, or a portion of the payload of the first data packet, for example, a portion of the payload intercepted from the packet header, such as a payload of 500 or 1000 bytes in length. Accordingly, the hash value of the first data packet can also be the hash value of the entire payload, or the hash value of a portion of the payload, which will not be described in detail here.
- the payload of the first data packet may include not only the payload content but also the payload length
- the hash value of the payload of the first data packet may include not only the hash value but also the payload length.
- the payload length may be used as an identifier to identify the length of the payload content or the length of the payload content corresponding to the hash value, and is optional.
- the payload length may be used to filter fingerprint features.
- the analysis platform when the analysis platform obtains the payload content or the hash value of the payload collected by each collector, it may be determined whether the lengths of the collected payloads configured by each collector are the same based on the payload length. If they are the same, the collected payload content or the hash value of the payload may be used as the fingerprint feature. Otherwise, the collected payload content or the hash value of the payload may not be used as the fingerprint feature. If the fingerprint feature is only for a portion of the payload, during the collection process, the payload of the required length may be directly intercepted through port mirroring for feature extraction.
- each collector collects the same fingerprint features.
- the applicant has found through theoretical analysis and a large number of experiments that the fingerprint features listed above will not change before and after NAT conversion, and different data streams are unlikely to conflict, that is, although the two data streams before and after NAT conversion may have different IP addresses or ports, the above fingerprint features in the two data streams are the same, so the above fingerprint features can identify NAT-related data streams.
- the fingerprint feature collection process listed above is relatively easy to implement, occupies less resources, and is conducive to the implementation of the method provided in the embodiment of the present application.
- the establishment time indicates the establishment time of the first data stream, for example, the time when the connection of the first data stream is started to be created, or the time when the connection of the first data stream is completed, or the time when the data packet of the first data stream is started to be sent.
- the establishment time can be the SYN message reception time, the SYN-ACK message reception time, the ACK message reception time, the SYN message sending time, the SYN-ACK message sending time, the ACK message sending time, or the reception or sending time of the first data packet.
- the SYN, SYN-ACK and ACK messages here are the messages in the TCP three-way handshake corresponding to the first data stream.
- the first data packet is the first data packet sent or received after the TCP connection corresponding to the first data stream is created.
- S12 Determine at least one data flow related to the first data flow NAT according to the fingerprint feature and establishment time of the first data flow.
- the NAT-related to the first data flow may be that the NAT-related to the first data flow belongs to the same NAT session as the first data flow.
- a NAT session includes a data stream before NAT conversion and a data stream after NAT conversion.
- the first data stream is the data stream before NAT device A converts
- the second data stream is the data stream obtained after NAT device A converts the first data stream by NAT, then the first data stream and the second data stream belong to the same NAT session.
- the data stream before NAT conversion and the data stream after NAT conversion have the same fingerprint features.
- the IPID of the SYN message of the first data stream is the same as the IPID of the SYN message of the second data stream
- the ISN of the SYN message of the first data stream is the same as the ISN of the SYN message of the second data stream
- the IPID of the first packet of the first data stream is the same as the IPID of the first packet of the second data stream
- the payload of the first packet of the first data stream is the same as the payload of the first packet of the second data stream
- the hash value of the payload of the first packet of the first data stream is the same as the hash value of the payload of the first packet of the second data stream
- the ISN of the SYN-ACK message of the first data stream is the ISN of the SYN-ACK message of the second data stream.
- the data flow related to the first data flow NAT may also be that the data flow after the first data flow is converted by NAT N times belongs to the same NAT session.
- N is a natural number greater than or equal to 1.
- a service data stream may undergo multiple NAT conversions from the service source to the service destination (for example, M NAT devices are deployed on the transmission path).
- the first data stream is a data stream sent from the service source
- the data stream becomes a second data stream after NAT conversion by the first NAT device.
- the second data stream and the first data stream belong to the same NAT session.
- the second data stream becomes a third data stream after NAT conversion by the second NAT device.
- the third data stream and the second data stream belong to the same NAT session.
- the third data stream is NAT-related to the first data stream
- the M-1th data stream becomes the Mth data stream after being converted by the M-1th NAT device
- the Mth data stream and the M-1th data stream belong to the same NAT session
- the Mth data stream is NAT-related to the first data stream
- the Mth data stream becomes the M+1th data stream after being converted by the Mth NAT device
- the M+1th data stream and the Mth data stream belong to the same NAT session
- the M+1th data stream is also NAT-related to the first data stream.
- the second data stream is the data stream after the first data stream is converted by NAT once
- the third data stream belongs to the same NAT session as the second data stream, that is, the third data stream and the data stream after the first data stream is converted by NAT once belong to the same NAT session.
- the M-1th data stream is the data stream after the first data stream is converted by NAT M-1 times
- the Mth data stream and the M-1th data stream belong to the same NAT session, that is, the Mth data stream and the data stream after the first data stream is converted by NAT M-1 times belong to the same NAT session.
- the Mth data stream is the data stream after the first data stream has been converted by NAT for M times, and the M+1th data stream and the Mth data stream belong to the same NAT session, that is, the M+1th data stream and the data stream after the first data stream has been converted by NAT for M times belong to the same NAT session.
- the fingerprint features of the data streams before and after NAT conversion are the same, that is, the fingerprint features of the data streams belonging to the same NAT session are the same, and every two data streams of the first data stream and at least one data stream related to the first data stream NAT belong to the same NAT session, therefore, the fingerprint features of the first data stream and the data stream related to the first data stream NAT are the same.
- the analysis platform may determine a data flow having the same fingerprint feature as the fingerprint feature of the first data flow and a data flow having an establishment time close to the establishment time of the first data flow as a data flow related to the first data flow NAT.
- the establishment time being close to the establishment time of the first data flow for example, means that the absolute value of the difference between the establishment time and the establishment time of the first data flow is less than a threshold.
- At least one data flow related to the first data flow NAT is determined by the fingerprint feature of the first data flow, without the need to read the session table from the NAT device, with low implementation difficulty and cost, and low network deployment difficulty.
- Figure 5 is a flow chart of a data flow identification method provided by an embodiment of the present application.
- the method can be executed by a collector, a network device or a storage platform deployed with a collector in the application scenarios shown in Figures 1 to 3. As shown in Figure 5, the method includes the following steps.
- S21 Receive a query request, where the query request includes a first stream identifier of a first data stream.
- S22 Send a query response, the query response includes the fingerprint feature and the establishment time of the first data flow, the fingerprint feature and the establishment time of the first data flow are used to determine at least one data flow related to the first data flow NAT.
- a query is performed using the first stream identifier as an index to obtain the fingerprint feature and establishment time of the first data stream, and then a query response is fed back.
- the query request may further include a query time range.
- the query request includes the query time range, the data stream established within the query time range is queried.
- the query time range is used to assist in finding the record conditions of the data stream. If it is not specified, the system can use a default value, such as today.
- the fingerprint features of multiple data streams can be returned as a response, or the fingerprint features of a data stream can be selected as a response from multiple data streams.
- the basis for selection can be the creation time of the stream.
- the collector or storage platform automatically selects the fingerprint features of the data stream with the most recent creation time as a response.
- the data stream with the most recent creation time is, for example, the data stream with the most recent record in the collector or storage platform.
- the data stream with the most recent creation time is, for example, the data stream with the most recent record in the query time range.
- the creation time of the multiple data streams is fed back to the analysis platform.
- the analysis platform provides a selection interface for the user to select from the multiple data streams according to the creation time, and then feeds back the selection results to the collector or storage platform.
- the collector or storage platform then feeds back the fingerprint features of the corresponding data stream based on the selection.
- Fig. 6 is a flow chart of a method for collecting data stream information provided in an embodiment of the present application. The method is executed by a collector. As shown in Fig. 6, the method includes the following steps.
- S31 Collect the establishment time, first stream identifier and fingerprint features of each data stream.
- the collector is connected in a bypass mode. At this time, the establishment time, first stream identifier and fingerprint features of each data stream are collected, including:
- the mirroring function can be message port mirroring or remote mirroring (such as encapsulated remote switch port analyzer (ERSPAN) remote mirroring).
- Message port mirroring is to copy the port message intact.
- Remote mirroring on the other hand, encapsulates port messages and then replicates them remotely.
- the message of the data stream is captured by the mirroring function, including:
- the mirroring function When configuring the mirroring function, you can configure multiple optional parameters of the mirroring function, which may include ACL.
- the ACL can be used to filter specified messages or specified message lengths to implement SYN message and/or SYN-ACK message collection.
- ACL By configuring ACL to capture only the three-way handshake messages during the TCP establishment process to collect the features of SYN messages and/or SYN-ACK messages, the number of mirror messages that need to be captured can be reduced, reducing the mirroring burden of network devices and the message processing burden of collectors, which is conducive to large-scale deployment. Compared with the features of the first packet, using the features of the SYN message or SYN-ACK message as the fingerprint feature has a smaller processing burden.
- the collector is connected in a path-by-path manner, for example, the forwarding device also includes a collection function. In this case, the collector directly analyzes the message of the data flow to obtain the establishment time, the first flow identifier and the fingerprint feature of the data flow.
- the collector when collecting the creation time of the data stream, the collector may use one of the following times as the creation time of the data stream:
- SYN message receiving time SYN-ACK message receiving time
- ACK message receiving time The ACK message here is the ACK message in the TCP three-way handshake.
- the collector when collecting the reception time of the ACK message in the TCP three-way handshake, can use the SYN message or SYN-ACK message as a benchmark to collect the reception time of the first ACK message after the SYN message or SYN-ACK message, that is, the reception time of the ACK message in the TCP three-way handshake.
- the above is for the collector in the inbound direction of the interface. If it is a collector in the outbound direction of the interface, one of the following times is used as the creation time of the data flow:
- S32 Store the creation time, first stream identifier and fingerprint feature of each data stream.
- the collector stores the collected data locally, or stores it in a network device where it is deployed.
- the collector sends the collected data to a storage platform for centralized storage.
- the first flow identifier of the data flow is a five-tuple of the data flow, including a source IP address, a destination IP address, a source port, a destination port, and a protocol type, wherein the protocol type is TCP.
- the information of the data stream may be stored in the format of the following Table 1:
- the recording time may be the acquisition time or the storage time. Since the recording time and the creation time are usually close, the query time range may be limited to the creation time or the recording time, which is not limited.
- the fingerprint feature of the data stream may consist of one feature or multiple features.
- Table 1 includes two features, which is only for example and is not intended to be a limitation of the present application.
- the collector can also collect and record information such as the number of packets, rate, and packet loss rate.
- FIG7 is a flow chart of a data flow identification method provided by an embodiment of the present application. The method is illustrated by taking the analysis platform, the first network node and the second network node as an example, where the first network node and the second network node are collectors or network devices with collectors deployed thereon. As shown in FIG7 , the method includes the following steps.
- the analysis platform sends a first query request to a first network node, where the first query request includes a first flow identifier of a first data flow; the first network node receives the first query request.
- the first network node here may be one or more.
- the first query request may also include a parameter field indicating the request.
- the first query request may carry different identifiers to indicate different parameters to be requested, such as identifier a indicating the establishment time of the requested data flow, and identifier b indicating the fingerprint of the requested data flow.
- identifier a indicating the establishment time of the requested data flow
- identifier b indicating the fingerprint of the requested data flow.
- the aforementioned method of indicating a request is only an example and is not intended to be a limitation of the present application.
- S42 The first network node performs a query using the first flow identifier as an index. When a result is found, the fingerprint feature and the creation time of the first data flow are obtained, and then S43 is executed, otherwise, subsequent steps are not executed.
- the network node here is a collector or a network device deployed with a collector, which stores the information of the data flow shown in Table 1, so the first flow identifier (quintuple) can be used as an index for query.
- each first network node executes step S42.
- the method may further include: the first network node feeds back a query failure message to the analysis platform to notify the analysis platform that no query result is found.
- the first network node sends a first query response to the analysis platform, where the first query response includes the fingerprint feature and establishment time of the first data flow; the analysis platform receives the first query response.
- the analysis platform sends a second query request to the second network node, where the second query request includes the fingerprint feature of the first data flow; the second network node receives the second query request.
- the second network node here may be one or more.
- the second network node and the first network node may be at least partially identical or completely different.
- S45 The second network node performs a query using the fingerprint feature of the first data flow as an index.
- a query result is obtained, the first flow identifier and establishment time of the second data flow having the same fingerprint feature as the first data flow are obtained, and then S46 is executed, otherwise the subsequent steps are not executed.
- each second network node executes step S45.
- the method may further include: the second network node feeds back a query failure message to the analysis platform to notify the analysis platform that no query result is found.
- the second network node sends a second query response to the analysis platform, where the second query response includes a first flow identifier of the second data flow having the same fingerprint feature as the first data flow and an establishment time; the analysis platform receives the second query response.
- the analysis platform may receive second query responses sent by different second network nodes.
- Each second query response may contain the same first stream identifier of the second data stream, which corresponds to a scenario in which there is only one NAT device on the transmission path between the source and destination of the first data stream.
- Different second query responses may contain different first stream identifiers of second data streams, which corresponds to a scenario in which there are multiple NAT devices on the transmission path between the source and destination of the first data stream. Accordingly, the analysis platform may obtain information on multiple second data streams.
- the analysis platform determines a data flow related to the first data flow NAT according to the establishment time of the first data flow and the establishment time of at least one second data flow.
- the analysis platform determines that the first data stream and the second data stream belong to the same NAT session, that is, the second data stream is a data stream related to the NAT of the first data stream. Otherwise, the analysis platform determines that the first data stream and the second data stream do not belong to the same NAT session.
- NAT sessions may have multiple levels, there may be multiple data flows related to the first data flow NAT. Therefore, when multiple data flows with the same fingerprint are queried, each data flow may be judged to determine whether it is related to the first data flow NAT.
- the conversion and transmission delay of multi-level NAT is small, and the absolute value of the time difference between the establishment time of the first data flow and the data flow after multi-level NAT conversion is less than the set threshold. Therefore, the above scheme can accurately determine whether the two data flows are NAT-related.
- the above threshold can be determined based on time synchronization accuracy and/or transmission delay.
- one or more NAT-related data flows travel a short distance and have a low transmission delay.
- the threshold can be determined based solely on the time synchronization accuracy of the collector. For example, none of the collectors belong to a high-speed data center, the transmission delay between the collectors is at the microsecond level, and the collectors use the network time protocol (NTP) for synchronization, and the synchronization accuracy is at the millisecond level. Relative to the synchronization accuracy, the transmission delay can be ignored. Therefore, the threshold can be determined directly based on the synchronization accuracy. For example, the threshold can be 10 milliseconds.
- the threshold when the synchronization accuracy between collectors is higher, the threshold can be designed to be smaller, such as 1 millisecond, 2 milliseconds, etc.
- the synchronization accuracy of a time synchronization protocol indicates the time error that may exist when devices in the network perform time synchronization based on the time synchronization protocol. For example, device A and device B perform time synchronization based on the NTP protocol. If the synchronization accuracy is at the millisecond level, it means that device A and device B The timing may differ by a few milliseconds.
- the transmission path between the collection points of NAT-related data streams is long and the transmission delay is high. Compared with the transmission delay, the time synchronization accuracy problem can be ignored.
- the above threshold can be determined based on the transmission delay alone. For example, two collection points span a wide area network, and the transmission delay between the collection points is about 50 to 60ms, which is far more than the synchronization accuracy of several milliseconds. In this case, the threshold can be determined directly based on the transmission delay, for example, the threshold is 100ms.
- the transmission delay between the collection points is close to the time synchronization accuracy.
- the transmission delay also needs to be considered, and the threshold can be determined based on the time synchronization accuracy and the transmission delay. For example, if the synchronization accuracy is in the millisecond level and the transmission delay is also in the millisecond level, the threshold can be 20ms.
- fingerprint conflicts refer to the same fingerprint features of non-NAT-related data flows.
- data flows with fingerprint conflicts can also be excluded.
- determining that the first data flow and the second data flow are NAT-related includes:
- the absolute value of the time difference between the establishment time of the second data stream and the establishment time of the first data stream is less than a threshold, and there is no fingerprint conflict between the first data stream and the second data stream, it is determined that the first data stream and the second data stream are NAT-related, and the fingerprint conflict means that the fingerprint features of non-NAT-related data streams are the same.
- two data flows are not data flows in the same NAT session, but the fingerprint features of the two data flows are the same. In this case, fingerprint conflicts exist between the two data flows.
- multiple data flows with the same fingerprint characteristics as the first data flow are not data flows after the first data flow is converted by N-level NAT. In this case, these multiple data flows conflict with the fingerprint of the first data flow.
- the analysis platform can perform fingerprint conflict judgment on each second data stream respectively.
- the fingerprint conflict is determined as follows:
- the second flow identifier of the data flow may be a four-tuple, including a source IP address, a destination IP address, a destination port, and a protocol type.
- This solution can not only determine the fingerprint conflict in the scenario where there is only a single-level NAT, but also determine the fingerprint conflict in the scenario where there are multiple levels of NAT. Because, in the above process of determining fingerprint conflicts, the principle used is:
- Devices with NAT function turned on will perform NAT conversion on data streams from the same host to the same destination and the same port (that is, data streams with the same four-tuple). Therefore, query the historical records to see whether the data streams with the same four-tuple have undergone NAT conversion. If it is found that there are data streams with the same four-tuple that have undergone NAT conversion, it is considered that this is also a NAT conversion, not a fingerprint conflict. Corresponding to the scenario with multiple levels of NAT, if there are data streams with the same four-tuple that have multiple levels of NAT conversion in the historical records, it is considered that this is also a multi-level NAT conversion. Therefore, the above judgment process can determine whether each second data stream has a fingerprint conflict with the first data stream or is NAT-related.
- the source port of the data stream is not limited because the second data stream and the third data stream can be data streams generated by different applications of the same host, but due to the same destination address and port, the same NAT conversion will be performed during transmission.
- data stream a is converted into data stream b after NAT conversion
- data stream b is converted into data stream C after NAT conversion
- data stream A having the same second stream identifier as data stream a is converted into data stream B after NAT conversion
- data stream B is converted into data stream C after NAT conversion.
- fingerprint conflict judgment is performed:
- the first data stream is data stream a
- the second data stream is data stream c
- the third data stream having the same second stream identifier as data stream c is determined to be data stream C
- the fourth data stream having the same fingerprint feature as data stream C is determined to be data stream B and data stream A.
- the absolute value of the time difference between the establishment time of data stream A and the establishment time of data stream C is less than the threshold, and the second stream identifier of data stream A is the same as data stream a, so it is determined that there is no fingerprint conflict.
- the first data flow input is flow11.
- the second data flows flow21 and flow31 can be found to be NAT-related data flows, as shown in Table 2 below:
- the third data flows flow22 and flow32 are queried according to the second data flows flow21 and flow31 respectively.
- At least one fourth data flow corresponding to the third data flow flow22 is flow12, and the third data flow flow32 has no corresponding fourth data flow.
- At least one fourth data flow corresponding to the third data flow flow22 is flow12, which meets the condition that the absolute value of the time difference of the establishment time is less than the threshold and has the same condition as the second flow identifier of the first data flow, it can be determined that there is no fingerprint conflict between the second data flow flow21 and the first data flow flow11. However, there is a fingerprint conflict between the second data flow flow31 and the first data flow flow11.
- the analysis platform may be a network diagnostic device. Accordingly, the method further includes:
- the fault diagnosis device first determines at least one second data stream related to the first data stream NAT, and then analyzes the reasons for the high packet loss rate or delay for the first data stream and at least one second data stream, determines the fault point on the link that causes the packet loss or delay, and completes the fault diagnosis.
- the network diagnostic device determines that there is packet loss in the first data stream based on the ACK message of the data packet, but the location of the packet loss is not in the link through which the first data stream passes. At this time, fault diagnosis is performed on the first data stream and the second data stream to determine that the location of the packet loss is in the link through which the second data stream passes, and the location of the fault point causing the packet loss is determined, thereby achieving more accurate positioning or demarcation of network faults.
- the network diagnostic equipment finds that the data packet transmission delay of the first data stream is large. In this case, fault diagnosis is performed on the first data stream and the second data stream to determine that the fault point causing the transmission delay is located on the link where the second data stream is located, thereby achieving more accurate positioning or demarcation of the network fault.
- the embodiment of the present application performs fault diagnosis at the granularity of data flow.
- the reason is that in some scenarios, the reason for packet loss or transmission delay is that the device on the link where the second data flow is located limits the burst traffic of the data flow. This situation does not involve other data flows on the link, and diagnosis at the granularity of data flow is more accurate.
- the data flow identification solution provided by the present application can determine NAT-related data flows, which enables the analysis platform to not only perform fault diagnosis on the original data flow, but also continue to perform fault diagnosis on the NAT-related data flow, and can more accurately locate or delimit network faults.
- the method further includes:
- the analysis platform When a fault occurs in the first data stream, the analysis platform performs fault diagnosis on the first data stream.
- Figure 8 is a flow chart of a data flow identification method provided by an embodiment of the present application. The method is illustrated by taking the analysis platform, the first network node and the second network node as an example, where the first network node is a storage platform. As shown in Figure 8, the method includes the following steps.
- the analysis platform sends a first query request to a first network node, where the first query request includes a first flow identifier of a first data flow; the first network node receives the first query request.
- S52 The first network node performs a query using the first flow identifier as an index. When a result is found, the fingerprint feature and the creation time of the first data flow are obtained, and then S53 is executed, otherwise, subsequent steps are not executed.
- the method may further include: the first network node feeds back a query failure message to the analysis platform to notify the analysis platform that no query result is found.
- S53 The first network node performs a query using the fingerprint feature of the first data flow as an index.
- a query result is obtained, the first flow identifier and establishment time of the second data flow having the same fingerprint feature as the first data flow are obtained, and then S54 is executed, otherwise the subsequent steps are not executed.
- the method may further include: the first network node feeds back a query failure message to the analysis platform to notify the analysis platform that no query result is found.
- S54 The first network node determines whether the first data flow and the second data flow are NAT-related according to the establishment time of the first data flow and the second data flow. If it is determined that the first data flow and the second data flow are NAT-related, S55 is executed, otherwise the subsequent steps are not executed.
- the absolute value of the time difference between the establishment time of the second data flow and the establishment time of the first data flow is less than the threshold, it is determined that the first data flow and the second data flow are NAT-related. Otherwise, it is determined that the first data flow and the second data flow are not NAT-related.
- determining that the first data flow and the second data flow are NAT-related includes:
- the absolute value of the time difference between the establishment time of the second data stream and the establishment time of the first data stream is less than a threshold, and there is no fingerprint conflict between the first data stream and the second data stream, it is determined that the first data stream and the second data stream are NAT-related, and the fingerprint conflict means that the fingerprint features of non-NAT-related data streams are the same.
- step S47 The detailed process of determining fingerprint conflict is shown in step S47 and will not be described in detail here.
- the method may further include: the first network node feeds back a notification message to the analysis platform to notify the analysis platform that there is no data flow related to the first data flow NAT.
- the first network node sends a third query response to the analysis platform, where the third query response includes a first flow identifier of a second data flow related to the first data flow NAT; the analysis platform receives the third query response.
- the analysis platform determines, based on the third query response, that the first data flow and the second data flow are NAT-related.
- step S57 For the detailed process of step S57, please refer to step S48.
- Figure 9 is a flow chart of a data flow identification method provided by an embodiment of the present application. The method is described by taking the analysis platform and the first network node as an example, where the first network node is a storage platform. As shown in Figure 9, the method includes the following steps.
- the analysis platform sends a first query request to a first network node, where the first query request includes a first flow identifier of a first data flow; the first network node receives the first query request.
- S62 The first network node performs a query using the first flow identifier as an index. When a result is found, the fingerprint feature and the creation time of the first data flow are obtained, and then S63 is executed, otherwise, subsequent steps are not executed.
- S63 The first network node performs a query using the fingerprint feature of the first data flow as an index.
- a query result is obtained, the first flow identifier and establishment time of the second data flow having the same fingerprint feature as the first data flow are obtained, and then S64 is executed, otherwise the subsequent steps are not executed.
- the method may further include: the first network node feeds back a query failure message to the analysis platform to notify the analysis platform that no query result is found.
- the first network node sends a first query response to the analysis platform, the first query response including the fingerprint characteristics and establishment time of the first data flow, and the first flow identifier and establishment time of the second data flow that is the same as the fingerprint characteristics of the first data flow; the analysis platform receives the first query response.
- the analysis platform determines whether the first data flow and the second data flow are NAT related according to the establishment time of the first data flow and the second data flow. close.
- the absolute value of the time difference between the establishment time of the second data flow and the establishment time of the first data flow is less than the threshold, it is determined that the first data flow and the second data flow are NAT-related. Otherwise, it is determined that the first data flow and the second data flow are not NAT-related.
- determining that the first data flow and the second data flow are NAT-related includes:
- the absolute value of the time difference between the establishment time of the second data stream and the establishment time of the first data stream is less than a threshold, and there is no fingerprint conflict between the first data stream and the second data stream, it is determined that the first data stream and the second data stream are NAT-related, and the fingerprint conflict means that the fingerprint features of non-NAT-related data streams are the same.
- step S47 The detailed process of determining fingerprint conflict is shown in step S47 and will not be described in detail here.
- step S66 For the detailed process of step S66, please refer to step S48.
- FIG10 is a block diagram of a data flow identification device provided in an embodiment of the present application.
- the data flow identification device can be implemented as all or part of an analysis platform, a network device, a collector or a storage platform through software, hardware or a combination of both.
- the data flow identification device may include: an acquisition unit 701 and a determination unit 702.
- the acquisition unit 701 is used to acquire the fingerprint feature and establishment time of the first data stream
- the determination unit 702 is configured to determine at least one data flow related to the first data flow NAT according to the fingerprint feature and establishment time of the first data flow.
- the acquisition unit 701 is used to send a first query request to the first network node, the first query request including a first flow identifier of the first data flow; receive a first query response sent by the first network node, the first query response including a fingerprint feature and an establishment time of the first data flow.
- the determination unit 702 is used to send a second query request to the second network node, the second query request including the fingerprint feature of the first data flow; receive a second query response sent by the second network node, the second query response including the first flow identifier and establishment time of the second data flow that are the same as the fingerprint feature of the first data flow; when the absolute value of the time difference between the establishment time of the second data flow and the establishment time of the first data flow is less than a threshold, determine that the first data flow and the second data flow are NAT related.
- the first query response further includes a first stream identifier and an establishment time of a second data stream having the same fingerprint feature as the first data stream;
- the determining unit 702 is configured to determine that the first data flow and the second data flow are NAT-related if the absolute value of the time difference between the establishment time of the second data flow and the establishment time of the first data flow is less than a threshold.
- the acquisition unit 701 is configured to receive a first query request, where the first query request includes a first flow identifier of a first data flow; and determine a fingerprint feature and an establishment time of the first data flow according to the first flow identifier of the first data flow.
- the determination unit 702 is used to determine the second data stream based on the fingerprint characteristics of the first data stream, and the fingerprint characteristics of the second data stream are the same as the fingerprint characteristics of the first data stream; when the absolute value of the time difference between the establishment time of the second data stream and the establishment time of the first data stream is less than a threshold, it is determined that the first data stream and the second data stream are NAT related.
- the fingerprint feature includes at least one of the following features:
- TCP Transmission Control Protocol
- ISN Initial Sequence Code
- the first query request further includes a query time range, and the query time range is used to limit the time range within which the establishment time of the queried data flow falls.
- the determination unit 702 is used to determine that the first data stream and the second data stream are NAT-related when the absolute value of the time difference between the establishment time of the second data stream and the establishment time of the first data stream is less than a threshold and there is no fingerprint conflict between the first data stream and the second data stream, and the fingerprint conflict means that the fingerprint features of non-NAT-related data streams are the same.
- the determination unit 702 is further used to determine at least one third data stream having the same second stream identifier as the second data stream; determine at least one fourth data stream having the same fingerprint feature as each third data stream; and determine that there is no fingerprint conflict between the first data stream and the second data stream when, in at least one fourth data stream, there is a fourth data stream whose establishment time and the corresponding third data stream have an absolute value of a time difference less than a threshold, and the second stream identifier of the fourth data stream is the same as the second stream identifier of the first data stream.
- the device further comprises:
- the diagnosis unit 703 is configured to perform fault diagnosis on the first data flow and at least one data flow related to the first data flow NAT when a fault occurs on the first data flow.
- the data flow identification device provided in the above embodiment only uses the division of the above functional units as an example when performing data flow identification.
- the above functions can be assigned to different functional units as needed, that is, the internal structure of the device is divided into different functional units to complete all or part of the functions described above.
- the data flow identification device provided in the above embodiment and the data flow identification method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
- FIG11 is a block diagram of a data flow identification device provided in an embodiment of the present application.
- the display device can be implemented as all or part of a collector, a network device deployed with a collector, or a storage platform through software, hardware, or a combination of both.
- the display device may include: a receiving unit 801 and a sending unit 802.
- the receiving unit 801 is configured to receive a query request, where the query request includes a first stream identifier of a first data stream;
- the sending unit 802 is used to send a query response, the query response includes the fingerprint feature and establishment time of the first data flow, the fingerprint feature and establishment time of the first data flow are used to determine at least one data flow related to the first data flow NAT.
- the query response further includes a first stream identifier and an establishment time of a second data stream having the same fingerprint feature as the first data stream.
- the device further comprises:
- the collection unit 803 is used to collect the establishment time, first stream identifier and fingerprint characteristics of each data stream;
- the storage unit 804 is used to store the establishment time, first flow identifier and fingerprint characteristics of each data flow.
- the collection unit 803 is used to capture the packets of the data flow through the mirroring function; analyze the captured packets of the data flow to obtain the establishment time, the first flow identifier and the fingerprint feature of the data flow.
- the data flow identification device provided in the above embodiment only uses the division of the above functional units as an example when performing data flow identification.
- the above functions can be assigned to different functional units as needed, that is, the internal structure of the device is divided into different functional units to complete all or part of the functions described above.
- the data flow identification device provided in the above embodiment and the data flow identification method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
- FIG12 shows a schematic diagram of the structure of an electronic device 150 provided in an embodiment of the present application.
- the electronic device may be an analysis platform, a network device, a collector or a storage platform.
- the electronic device 150 shown in FIG12 is used to perform the operations involved in the data flow identification method shown in any of FIG4 to FIG9 above.
- the electronic device 150 may be implemented by a general bus architecture.
- the electronic device 150 includes at least one processor 151 , a memory 153 , and at least one communication interface 154 .
- the processor 151 is, for example, a general-purpose central processing unit (CPU), a digital signal processor (DSP), a network processor (NP), a data processing unit (DPU), a microprocessor, or one or more integrated circuits for implementing the solution of the present application.
- the processor 151 includes an application-specific integrated circuit (ASIC), a programmable logic device (PLD) or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof.
- ASIC application-specific integrated circuit
- PLD programmable logic device
- the PLD is, for example, a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
- CPLD complex programmable logic device
- FPGA field-programmable gate array
- GAL generic array logic
- the processor may also be a combination that implements a computing function, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
- the electronic device 150 further includes a bus.
- the bus is used to transmit information between components of the electronic device 150.
- the bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus.
- PCI peripheral component interconnect
- EISA extended industry standard architecture
- the bus may be divided into an address bus, a data bus, a control bus, etc.
- FIG12 only uses a thick line, but does not mean that there is only one bus or one type of bus.
- the memory 153 is, for example, a read-only memory (ROM) or other type of static storage device that can store static information and instructions, or a random access memory (RAM) or other type of memory that can store information and instructions.
- the memory 153 may be a dynamic storage device such as an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
- the memory 153 for example, exists independently and is connected to the processor 151 via a bus.
- the memory 153 may also be integrated with the processor 151.
- the communication interface 154 uses any transceiver-like device to communicate with other devices or communication networks, and the communication network can be Ethernet, a radio access network (RAN) or a wireless local area network (WLAN), etc.
- the communication interface 154 may include a wired communication interface and may also include a wireless communication interface.
- the communication interface 154 may be an Ethernet interface, a Fast Ethernet (FE) interface, a Gigabit Ethernet (GE) interface, an Asynchronous Transfer Mode (ATM) interface, a wireless local area network (WLAN) interface, a cellular network communication interface or a combination thereof.
- the Ethernet interface may be an optical interface, an electrical interface or a combination thereof.
- the communication interface 154 may be used for the electronic device 150 to communicate with other devices.
- the processor 151 may include one or more CPUs, such as CPU0 and CPU1 shown in FIG12 . Each of these processors may be a single-CPU processor or a multi-CPU processor.
- the processor here may refer to one or more devices, circuits, and/or processing cores for processing data (e.g., computer program instructions).
- the electronic device 150 may include multiple processors, such as the processor 151 and the processor 155 shown in FIG12. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU).
- the processor here may refer to one or more devices, circuits, and/or processing cores for processing data (such as computer program instructions).
- the electronic device 150 may also include an output device and an input device.
- the output device communicates with the processor 151 and can display information in a variety of ways.
- the output device may be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector.
- the input device communicates with the processor 151 and can receive user input in a variety of ways.
- the input device may be a mouse, a keyboard, a touch screen device, or a sensor device.
- the memory 153 is used to store the program code 1510 for executing the solution of the present application, and the processor 151 can execute the program code 1510 stored in the memory 153. That is, the electronic device 150 can implement the data flow identification method provided by the method embodiment by executing the program code 1510 in the memory 153 through the processor 151.
- the program code 1510 may include one or more software modules.
- the processor 151 itself may also store the program code or instruction for executing the solution of the present application.
- the electronic device 150 of the embodiment of the present application may correspond to the controller in the above-mentioned method embodiments, and the processor 151 in the electronic device 150 reads the instructions in the memory 153, so that the electronic device 150 shown in Figure 12 can execute all or part of the operations performed by the controller.
- the processor 151 is used to obtain the fingerprint characteristics and establishment time of the first data flow; and determine at least one data flow related to the first data flow NAT according to the fingerprint characteristics and establishment time of the first data flow.
- each step of the data flow identification method shown in any of Figures 4 to 9 is completed by an integrated logic circuit of hardware in the processor of the electronic device 150 or an instruction in the form of software.
- the steps of the method disclosed in conjunction with the embodiment of the present application can be directly embodied as a hardware processor, or a combination of hardware and software modules in the processor.
- the software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc.
- the storage medium is located in the memory, and the processor reads the information in the memory, and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it is not described in detail here.
- the embodiment of the present application also provides a chip, including: an input interface, an output interface, a processor and a memory.
- the input interface, the output interface, the processor and the memory are connected through an internal connection path.
- the processor is used to execute the code in the memory.
- the processor is used to execute any of the above-mentioned data flow identification methods.
- the processor may be a CPU, or other general-purpose processors, DSP, ASIC, FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
- a general-purpose processor may be a microprocessor or any conventional It is worth noting that the processor may be a processor supporting ARM architecture.
- the processor is one or more and the memory is one or more.
- the memory can be integrated with the processor, or the memory is separately arranged with the processor.
- the memory can include a read-only memory and a random access memory, and provide instructions and data to the processor.
- the memory can also include a non-volatile random access memory.
- the memory can also store a reference block and a target block.
- the memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory.
- the non-volatile memory may be a ROM, a PROM, an EPROM, an EEPROM, or a flash memory.
- the volatile memory may be a RAM, which is used as an external cache.
- many forms of RAM are available. For example, SRAM, DRAM, SDRAM, DDR SDRAM, ESDRAM, SLDRAM, and DR RAM.
- a computer-readable storage medium is further provided.
- the computer-readable storage medium stores computer instructions.
- the electronic device executes the data flow identification method provided above.
- a computer program product including instructions is also provided.
- the electronic device executes the data flow identification method provided above.
- the computer program product includes one or more computer instructions.
- the computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.
- the computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.
- the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means.
- the computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated.
- the available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk), etc.
Landscapes
- Engineering & Computer Science (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Data Exchanges In Wide-Area Networks (AREA)
Abstract
公开了一种数据流识别方法及装置、电子设备。该方法包括:获取第一数据流的指纹特征和建立时间;根据所述第一数据流的指纹特征和建立时间,确定与所述第一数据流NAT相关的至少一个数据流。
Description
本申请要求于2023年02月07日提交的申请号202310129577.1、申请名称为“数据流识别方法及装置、电子设备”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请涉及网络技术领域,特别涉及一种数据流识别方法及装置、电子设备。
网络地址转换(network address translation,NAT)的作用是进行内部私有网络地址和公网互联网协议(internet protocol,IP)地址的转换。网络故障分析器通常基于数据流的五元组检测数据流在传输过程中的传输质量,以进行故障诊断。如果一个数据流的传输过程中涉及NAT设备,则NAT设备转换前后的数据流的五元组将发生变化,这将导致网络故障分析器无法准确精细的定界故障范围。例如,网络故障分析器仅能将NAT设备到目的端或源端到NAT设备之间的范围确定为故障范围。
相关技术可以读取开启了NAT功能的设备(后文简称NAT设备)的NAT会话表信息,根据NAT会话表信息获得NAT转换前后两条数据流的对应关系,根据数据流的对应关系确定一条数据流属于同一NAT会话的数据流,进而可以在进行故障诊断时同时考虑属于同一NAT会话的数据流。
但是,不同厂商和型号的NAT设备的NAT会话表的导出方式不一,部署上述方案需要逐个NAT设备对接,甚至部分NAT设备没有现成的管理接口来获取NAT会话表,实现难度及代价较大;并且,需要预先获取网络中的NAT设备的部署信息,从而能够从NAT设备读取NAT会话表信息,并且在NAT设备的部署变化时需要同步更新部署信息,部署难度较大。
发明内容
本申请提供了一种数据流识别方法及装置、电子设备,能够降低确定与第一数据流NAT相关的数据流的实现难度及代价,降低网络部署难度。
第一方面,本申请提供了一种数据流识别方法。该方法包括:获取第一数据流的指纹特征和建立时间;根据第一数据流的指纹特征和建立时间,确定与第一数据流NAT相关的至少一个数据流。
在该实现方式中,通过第一数据流的指纹特征确定与第一数据流NAT相关的至少一个数据流,不需要从NAT设备读取会话表,实现难度及代价小,并且网络部署难度小。
由于实际网络中可以是一级NAT,也可能存在多级NAT。一级NAT时会产生一个数据流与第一数据流NAT相关,多级NAT时会产生多个数据流与第一数据流NAT相关。因此,可以确定出一个或多个与第一数据流NAT相关的数据流。
在本申请的实现方式中,与第一数据流NAT相关可以是与第一数据流属于同一个NAT会话。或者,与第一数据流NAT相关还可以是与第一数据流经过N级NAT转换后的数据流属于同一个NAT会话。
该方法可以由网络中的分析平台、网络设备、采集器、存储平台等执行,网络中的节点上部署有采集器(也即探针),通过采集器采集数据流的信息,数据流的信息包括指纹特征。以分析平台为例,分析平台通过采集器采集到的数据流的指纹特征执行数据流识别方法。
在一些可能的实现方式中,采集器通常采用软件实现。在其他实现方式中,采集器也可以采用硬件实现。
为了保证采集到的指纹特征能够用来识别数据流是否NAT相关,可以在启用NAT功能的节点的上下游均部署采集器。可选地,也可以采用其他方式部署采集器。作为示例,下方列出了三种采集器的部署方式。
第一种,在每个节点上均部署采集器,从而保证能够采集到每个NAT会话的数据流的指纹特征。
第二种,在网络的多个核心节点部署采集器,例如多个Spine交换机、多个边界节点、和/或防火墙上部署采集器,从而能够尽量采集到每个NAT会话的数据流的指纹特征。
第三种,在具有NAT功能的节点上部署采集器,以采集到每个NAT会话的数据流的指纹特征。
在第一种和第二种部署方式中,采集器通常部署在节点的接口的入方向,当然也可以部署在接口的出方向。在第三种部署方式中,采集器需要同时在接口的入方向和接口的出方向部署。
在一些可能的实现方式中,部署采集器(软件)的各个节点或者采集器(硬件)自己存储各自采集到的数据流的信息。
在另一些可能的实现方式中,该网络还可以包括存储平台,用于集中存储各个采集器采集到的数据流的信息。
以分析平台为例,在数据流的信息采集完成后,分析平台通过与采集器或者存储平台交互,从而获取数据流的信息并进一步确定与第一数据流NAT相关的至少一个数据流。
以存储平台为例,在数据流的信息采集完成后,存储平台通过本地存储的数据流的信息确定与第一数据流NAT相关的至少一个数据流。
下面分别对这两种情况进行说明:
第一种,分析平台通过与采集器或者存储平台交互实现该方法的情况:
在这种实施方式中,获取第一数据流的指纹特征和建立时间,包括:
向第一网络节点发送第一查询请求,第一查询请求包括第一数据流的第一流标识;
接收第一网络节点发送的第一查询应答,第一查询应答包括第一数据流的指纹特征和建立时间。
通过从第一网络节点请求第一数据流的指纹特征和建立时间,为后续确定与第一数据流NAT相关的至少一个数据流做准备。
这里的第一网络节点可以是采集器、部署有采集器的网络设备或者存储平台。并且,在发送查询请求时,如果是发送给采集器或者部署有采集器的网络设备,可以同时发送给多个采集器或者部署有采集器的网络设备。
示例性地,数据流的第一流标识为数据流的五元组,包括源互联网协议(internet protocol,IP)地址、目的IP地址、源端口、目的端口和协议类型,其中协议类型为传输控制协议(transmission control protocol,TCP)。
在上述过程中,第一网络节点接收到查询请求后,采用第一流标识作为索引进行查询,得到第一数据流的指纹特征和建立时间。
由于一条数据流是双向传输的,因此查询时除了按照原本的五元组进行查询外,还可以将源、目的IP调换,将源、目的端口调换后进行查询。
可选地,第一查询请求除了包括第一流表示外,还包括查询时间范围,查询时间范围用于限制查询到的数据流的建立时间所处的时间范围。
例如,查询时间范围采用确切的范围表示,如某一天的0点至24点,或者某一天某个小时的0分至60分等。再例如,查询时间范围采用时间粒度表示,如一天,则对应的范围是当天,或者一个小时,则对应的范围是当前小时内等。
当然,在第一查询请求中不包括查询时间范围时,查询时间范围可以采用默认值,例如查询当天等。
在一些可能的实现方式中,数据流的指纹特征可以包括一个或多个特征。
示例性地,指纹特征包括如下特征中的至少一项:
首个数据包的互联网协议标识(IP identifier,IPID)、首个数据包的载荷、首个数据包的载荷的哈希值、TCP三次握手过程中同步序列编号(synchronize sequence numbers,SYN)报文的IPID、TCP三次握手过程中SYN报文的初始序列码(initial sequence number,ISN)、TCP三次握手过程中SYN-ACK报文的ISN。
其中,首个数据包(简称首包)是指TCP三次握手建立连接后传输的第一个数据包。不同数据流需要传输的内容不同,有的数据流需要多个数据包才能完成内容的传输,而有的数据流仅需要传输一个数据包就可以完成内容的传输。因此,本申请实施例采用首个数据包的信息作为指纹特征,可以为每个需要判断的数据流采集到指纹特征。另一方面,首个数据包即是TCP建链后第一个被传输的数据包,而且其序列号等于SYN报文携带的ISN+1,采集器可以很容易地确定首个数据包,因此,采集首个数据包所需的资源和
时间相比后续数据包更少。
采集器在采集时,可以根据数据包的序列号(sequence)来判断是否是首个数据包,通常首个数据包的序列号是SYN报文的ISN+1。而第二个数据包的序列号是SYN报文的ISN+1再加首个数据包的长度,因此,如果要采集首个数据包之后的数据包的特征,则存在判断数据包的序列号更加困难,占用资源和时间更多的问题。所以优先选用首个数据包进行指纹特征采集。
首个数据包的载荷可以是首个数据包的全部载荷,也可以是首个数据包的部分载荷,例如是从数据包包头开始截取的部分载荷,如500或1000字节长度的载荷。相应地,首个数据包的哈希值也可以是全部载荷的哈希值,也可以是部分载荷的哈希值,这里不再赘述。
在该实现方式中,各个采集器采集相同的指纹特征。申请人通过理论分析和大量的试验发现,上述所列举的指纹特征在NAT转换前后不会变化,并且不同数据流难以发生冲突,即,NAT转换前后的两个数据流尽管IP地址或端口可能不同,但两个数据流中的上述指纹特征相同,因此,上述指纹特征能够标识NAT相关的数据流。另外,上述所列举的指纹特征采集过程较为容易实现,占用资源少,有利于实现本申请实施例提供的方法。
在其他可能的实现方式中,数据流的指纹特征还可以包括其他特征,只要能够在NAT转换前后不变,并且不同数据流难以发生冲突即可,本申请对此不做限制。
在一些可能的实现方式中,采集器采集的数据流的创建时间可以是接收或发送某个数据包或者报文的时间。
采集器在采集数据流的创建时间时,可以采用如下时间中的一种作为数据流的创建时间:
SYN报文接收时间、SYN-ACK报文接收时间、ACK报文接收时间。此处的ACK报文为TCP三次握手中的ACK报文。
当然上面针对的是接口的入方向的采集器,如果是接口的出方向的采集器,则采用如下时间中的一种作为数据流的创建时间:
SYN报文发送时间、SYN-ACK报文发送时间、ACK报文发送时间。
在其他可能的实现方式中,采集器在采集数据流的创建时间时,也可以采用其他时间,例如首个数据包的接收或发送时间等。
在这种实施方式中,根据第一数据流的指纹特征和建立时间,确定与第一数据流NAT相关的至少一个数据流的方式包括多种:
在一种可能的实现方式中,分析平台利用第一数据流的指纹特征,查询指纹特征相同的第二数据流,然后根据二者的创建时间判断是否NAT相关,步骤如下:
向第二网络节点发送第二查询请求,第二查询请求包括第一数据流的指纹特征;
接收第二网络节点发送的第二查询应答,第二查询应答包括与第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间;
在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定第一数据流和第二数据流NAT相关。
在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值不小于阈值的情况下,确定第一数据流和第二数据流不NAT相关。
其中,第二网络节点和第一网络节点可以是相同的网络节点,也可以是不同的网络节点。
第一网络节点和第二网络节点可以是分布式的多个采集器或部署有采集器的网络设备。
在该实现方式中,通过第二次查询指纹特征相同的第二数据流的创建时间。然后以指纹特征相同、创建时间相近作为判断标准,判断NAT相关的数据流。由于NAT转换以及数据流在不同采集器之间传输存在时间延迟,因此,NAT转换前采集到的数据流的建立时间和NAT转换后采集到的数据流的建立时间有时间延迟,通过限定两个数据流的建立时间的时间差的绝对值小于阈值,能够确定两个相同指纹特征的数据流是经过NAT转换的两个数据流。
不同网络节点上的采集器采集时存在同步精度问题,不同的采集器之间还会存在传输时延,因此,可以基于时间同步精度和/或传输时延确定上述阈值。
在另一种可能的实现方式中,第一查询应答还包括与第一数据流的指纹特征相同的第二数据流的第一
流标识以及建立时间;此时,分析平台直接根据两个数据流的创建时间判断是否NAT相关,步骤如下:
在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定第一数据流和第二数据流NAT相关。
在这种情况下,分析平台发送第一查询请求时,第一网络节点除了能够将第一数据流的信息反馈给分析平台外,还可以同时查询与第一数据流指纹相同的第二数据流的信息,一起反馈给分析平台。这样,分析平台可以直接根据第一查询应答中的内容,确定两个数据流是否NAT相关。
在上述几种确定与第一数据流NAT相关的至少一个数据流的方式中,第二数据流可以是一个,也可以是多个,如果是多个,则分别确定每一个第二数据流是否与第一数据流NAT相关。
这种情况下,第一网络设备可以为存储平台。
第二种,存储平台基于本地存储的采集信息实现该方法的情况:
在这种实施方式中,获取第一数据流的指纹特征和建立时间,包括:
接收第一查询请求,第一查询请求包括第一数据流的第一流标识;根据第一数据流的第一流标识确定第一数据流的指纹特征和建立时间。
在这种情况下,方法由存储平台执行,存储平台可以根据查询请求中的流标识查询第一数据流的指纹特征和建立时间。
在这种实施方式中,根据第一数据流的指纹特征和建立时间,确定与第一数据流NAT相关的至少一个数据流的步骤包括:
根据第一数据流的指纹特征确定第二数据流,第二数据流的指纹特征与第一数据流的指纹特征相同;
在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定第一数据流和第二数据流NAT相关。
其中,第二数据流的建立时间可以是查询第一数据流的指纹特征找到相同指纹特征的第二数据流,然后得到的。
在本申请的实现方式中,由于指纹特征采用了冲突可能性小的特征来实现,但为了保证确定NAT相关的数据流的准确性,在上述过程中,还可以将指纹冲突的情况排除掉,指纹冲突是指非NAT相关的数据流的指纹特征相同。
在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定第一数据流和第二数据流NAT相关,包括:
在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值,第一数据流和第二数据流不存在指纹冲突的情况下,确定第一数据流和第二数据流NAT相关。
在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值,第一数据流和第二数据流存在指纹冲突的情况下,确定第一数据流和第二数据流非NAT相关,也即存在指纹冲突。
示例性地,确定第一数据流和第二数据流指纹冲突,可以通过如下步骤:
确定与第二数据流具有相同第二流标识的至少一条第三数据流;
确定与每条第三数据流具有相同指纹特征的至少一条第四数据流;
在至少一条第四数据流中存在一条第四数据流的建立时间和对应第三数据流的建立时间的时间差的绝对值小于阈值,且这一条第四数据流的第二流标识和第一数据流的第二流标识相同的情况下,确定第一数据流和第二数据流不存在指纹冲突。
其中,数据流的第二流标识可以是四元组,包括源IP地址、目的IP地址、目的端口和协议类型。
在上述确定指纹冲突的过程中,使用的原理是查询与该数据流具有相同四元组的数据流在历史记录中是否进行过NAT转换,如果存在具有相同四元组的数据流经过NAT转换,则说明本次也是NAT转换,而非指纹冲突。
实施时,先确定与第二数据流的第二流标识相同的第三数据流,然后查询与第三数据流指纹特征相同且建立时间差小于阈值的第四数据流,如果第四数据流与第一数据流的第二流标识相同,则说明是NAT转换,而非指纹冲突。其中,第二流标识相同说明,两条数据流是从相同的源地址出发去往相同的目的地址,并且在NAT转换的同一侧(转换前或后)。如果第三数据流和第四数据流存在NAT转换,那么分别与之对应的第二数据流和第一数据流也存在NAT转换。
在本申请的实现方式中,第一数据流可以是用户选定的数据流,用户在观测到第一数据流出现故障时,确定与第一数据流NAT相关的数据流。当然,在其他实现方式中,第一数据流也可以是诊断设备自动选取的,诊断设备检测到第一数据流的故障,然后确定与第一数据流NAT相关的数据流。
在确定存在与第一数据流NAT相关的数据流之后,该方法还包括:
在第一数据流发生故障的情况下,对第一数据流以及与第一数据流NAT相关的至少一个数据流进行故障诊断。
通过对NAT相关数据流同时进行故障诊断,能够保证故障诊断的准确性。
在确定不存在与第一数据流NAT相关的数据流之后,该方法还包括:
在第一数据流发生故障的情况下,对第一数据流进行故障诊断。
其中,数据流发生故障是指基于该数据流的应用发生故障,包括但不限于数据流丢包率高、延迟大等情况。
第二方面,本申请提供了一种数据流识别方法。该方法包括:接收查询请求,查询请求包括第一数据流的第一流标识;发送查询应答,查询应答包括第一数据流的指纹特征和建立时间,第一数据流的指纹特征和建立时间用于确定与第一数据流NAT相关的至少一个数据流。
第二方面提供的方法与第一方面提供的方法的区别在于:第二方面提供的方法仅进行数据流的指纹特征的查询和反馈,而不进行与第一数据流NAT相关的至少一个数据流的判断。
执行该方法的可以是采集器、部署有采集器的网络设备或者存储平台。
在一些可能的实现方式中,查询应答还包括与第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间。
可选地,该方法还包括:
采集各个数据流的建立时间、第一流标识和指纹特征;
存储各个数据流的建立时间、第一流标识和指纹特征。
在一些可能的实现方式中,采集器以旁路的方式接入。此时,采集各个数据流的建立时间、第一流标识和指纹特征,包括:
通过镜像功能抓取数据流的报文;
分析抓取到的数据流的报文,以获取数据流的建立时间、第一流标识和指纹特征。
在本申请的实现方式中,当指纹特征采用TCP三次握手中的SYN报文的特征时,通过镜像功能抓取数据流的报文,包括:
通过配置访问控制列表(access control lists,ACL)抓取SYN报文和/或SYN-ACK报文。
通过配置ACL实现SYN报文和/或SYN-ACK报文的特征的采集,可以减少需要抓取的报文数量,降低采集负担。
在另一些可能的实现方式中,采集器以随路的方式接入。此时,直接分析数据流的报文,获取数据流的建立时间、第一流标识和指纹特征即可。
第三方面,提供了一种数据流识别装置,该装置包括:
获取单元,用于获取第一数据流的指纹特征和建立时间;
确定单元,用于根据第一数据流的指纹特征和建立时间,确定与第一数据流NAT相关的至少一个数据流。
可选地,获取单元,用于向第一网络节点发送第一查询请求,第一查询请求包括第一数据流的第一流标识;接收第一网络节点发送的第一查询应答,第一查询应答包括第一数据流的指纹特征和建立时间。
可选地,确定单元,用于向第二网络节点发送第二查询请求,第二查询请求包括第一数据流的指纹特征;接收第二网络节点发送的第二查询应答,第二查询应答包括与第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间;在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定第一数据流和第二数据流NAT相关。
可选地,第一查询应答还包括与第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间;
确定单元,用于在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定第一数据流和第二数据流NAT相关。
可选地,获取单元,用于接收第一查询请求,第一查询请求包括第一数据流的第一流标识;根据第一数据流的第一流标识确定第一数据流的指纹特征和建立时间。
可选地,确定单元,用于根据第一数据流的指纹特征确定第二数据流,第二数据流的指纹特征与第一数据流的指纹特征相同;在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定第一数据流和第二数据流NAT相关。
可选地,指纹特征包括如下特征中的至少一项:
首个数据包的互联网协议标识IPID、首个数据包的载荷、首个数据包的载荷的哈希值、传输控制协议TCP三次握手过程中SYN报文的IPID、TCP三次握手过程中SYN报文的初始序列码ISN、TCP三次握手过程中SYN-ACK报文的ISN。
可选地,第一查询请求还包括查询时间范围,查询时间范围用于限制查询到的数据流的建立时间所处的时间范围。
可选地,确定单元,用于在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值,第一数据流和第二数据流不存在指纹冲突的情况下,确定第一数据流和第二数据流NAT相关,指纹冲突是指非NAT相关的数据流的指纹特征相同。
可选地,确定单元,还用于确定与第二数据流具有相同第二流标识的至少一条第三数据流;确定与每条第三数据流具有相同指纹特征的至少一条第四数据流;在至少一条第四数据流中存在一条第四数据流的建立时间和对应第三数据流的建立时间的时间差的绝对值小于阈值,且这一条第四数据流的第二流标识和第一数据流的第二流标识相同的情况下,确定第一数据流和第二数据流不存在指纹冲突。
可选地,该装置还包括:
诊断单元,用于在第一数据流发生故障的情况下,对第一数据流以及与第一数据流NAT相关的至少一个数据流进行故障诊断。
第四方面,提供了一种数据流识别装置,装置包括:
接收单元,用于接收查询请求,查询请求包括第一数据流的第一流标识;
发送单元,用于发送查询应答,查询应答包括第一数据流的指纹特征和建立时间,第一数据流的指纹特征和建立时间用于确定与第一数据流NAT相关的至少一个数据流。
可选地,查询应答还包括与第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间。
可选地,该装置还包括:
采集单元,用于采集各个数据流的建立时间、第一流标识和指纹特征;
存储单元,用于存储各个数据流的建立时间、第一流标识和指纹特征。
可选地,采集单元,用于通过镜像功能抓取数据流的报文;分析抓取到的数据流的报文,以获取数据流的建立时间、第一流标识和指纹特征。
第五方面,提供了一种电子设备。所述电子设备包括处理器和存储器。所述存储器用于存储软件程序以及模块。所述处理器通过运行或执行存储在所述存储器内的软件程序和/或模块实现上述第一方面或第一方面的任一种可能的实施方式中的方法,或者实现上述第二方面或第二方面的任一种可能的实施方式中的方法。
可选地,所述处理器为一个或多个,所述存储器为一个或多个。
可选地,所述存储器可以与所述处理器集成在一起,或者所述存储器与处理器分离设置。
在具体实现过程中,存储器可以为非瞬时性(non-transitory)存储器,例如只读存储器(read only memory,ROM),其可以与处理器集成在同一块芯片上,也可以分别设置在不同的芯片上,本申请对存储器的类型以及存储器与处理器的设置方式不做限定。
第六方面,提供了一种数据流识别系统,所述数据流识别系统包括采集器和分析平台,所述分析平台用于执行上述第一方面或第一方面的任一种可能的实施方式中的方法,所述采集器用于执行上述第二方面或第二方面的任一种可能的实施方式中的方法。
可选地,所述数据流识别系统还包括存储平台,用于存储所述采集器采集的数据流的信息。
第七方面,提供了一种计算机程序产品。所述计算机程序产品包括计算机程序代码,当所述计算机程序代码被计算机运行时,使得所述计算机执行上述第一方面或第一方面的任一种可能的实施方式中的方法,
或者执行上述第二方面或第二方面的任一种可能的实施方式中的方法。
第八方面,本申请提供了一种计算机可读存储介质,所述计算机可读存储介质用于存储处理器所执行的程序代码,所述程序代码包括用于实现上述第一方面任一种可能的实施方式中的方法,或者实现上述第二方面或第二方面的任一种可能的实施方式中的方法。
第九方面,提供了一种芯片,包括处理器,处理器用于从存储器中调用并运行所述存储器中存储的指令,使得安装有所述芯片的通信设备执行上述第一方面任一种可能的实施方式中的方法,或者执行上述第二方面或第二方面的任一种可能的实施方式中的方法。
第十方面,提供另一种芯片。所述另一种芯片包括输入接口、输出接口、处理器和存储器。所述输入接口、输出接口、所述处理器以及所述存储器之间通过内部连接通路相连。所述处理器用于执行所述存储器中的代码,当所述代码被执行时,所述处理器用于执行上述第一方面任一种可能的实施方式中的方法,或者执行第二方面或第二方面的任一种可能的实施方式中的方法。
图1是本申请实施例提供的一种应用场景的结构示意图;
图2是本申请实施例提供的另一种应用场景的结构示意图;
图3是本申请实施例提供的一种数据中心的网络拓扑示意图;
图4是本申请实施例提供的一种数据流识别方法的流程图;
图5是本申请实施例提供的一种数据流识别方法的流程图;
图6是本申请实施例提供的一种数据流信息的采集方法的流程图;
图7是本申请实施例提供的一种数据流识别方法的流程图;
图8是本申请实施例提供的一种数据流识别方法的流程图;
图9是本申请实施例提供的一种数据流识别方法的流程图;
图10是本申请实施例提供的一种数据流识别装置的框图;
图11是本申请实施例提供的一种数据流识别装置的框图;
图12示出了本申请实施例提供的一种电子设备的结构示意图。
为使本申请的目的、技术方案和优点更加清楚,下面将结合附图对本申请实施方式作进一步地详细描述。
图1是本申请实施例提供的一种应用场景的结构示意图。参见图1,该应用场景包括多个网络设备11,多个网络设备11构成网络,例如数据中心网络。网络设备11可以是交换机、路由设备、防火墙设备、服务器等。
多个网络设备11中的至少部分网络设备11上部署并启用NAT功能。如图1所示,多个网络设备11中的至少部分网络设备11上部署有采集器12,用于采集数据流的信息。
图2是本申请实施例提供的另一种应用场景的结构示意图。参见图2,该应用场景与图1的区别在于,还包括存储平台13,存储平台13同时与各个采集器12连接,用于集中存储各个采集器采集到的数据流的信息。
在本申请实施例中,采集器和存储平台可以均对外提供数据查询接口。
可选地,在图1和图2的场景中,还可以包括分析平台14,分析平台14同时与各个采集器12连接,或者与存储平台13连接,通过分析数据流的信息以确定NAT相关的数据流。可选地,存储平台可以和分析平台集成到同一个设备,例如,分析平台包括存储模块,或,存储平台包括分析模块等。
在本申请实施例中,网络设备、采集器、存储平台或分析平台可以通过数据流的信息确定NAT相关的数据流。
图3是本申请实施例提供的一种数据中心的网络拓扑示意图。参见图3,数据中心包括Spine交换机14和虚拟化服务器集群15,Spine交换机14分别连接虚拟化服务器集群15。
可选地,数据中心还可以包括分析器网络16,分析器网络16包括采集器集群161、分析器集群162和
Leaf交换机163,Spine交换机14通过Leaf交换机163分别连接采集器集群161和分析器集群162,采集器集群161包括前述存储平台,分析器集群162包括前述分析平台。在其他实现方式中,分析器网络16也可以独立于数据中心之外,相应地,Leaf交换机163可以是其他网络设备。
其中,Spine交换机14和虚拟化服务器集群15都属于前述网络设备11,在Spine交换机14上、以及Spine交换机14和虚拟化服务器集群15之间的网络设备11上部署采集器12。在数据流17传输过程中,三个采集器12采集数据流的信息,然后通过虚线所示的路径上传到采集器集群161的存储平台中。
数据中心在边界会部署防火墙(firewall,FW)(图中未示出),启用NAT功能,保护数据中心的网络。
在本申请实施例中,为了保证采集到的指纹特征能够用来识别数据流是否NAT相关,可以在启用NAT功能的节点的上下游均部署采集器。可选地,也可以采用其他方式部署采集器。作为示例,下方列出了三种采集器的部署方式。
第一种,在每个节点上均部署采集器,从而保证能够采集到每个NAT会话的数据流的指纹特征。
第二种,在网络的多个核心节点部署采集器,例如多个Spine交换机、多个边界节点、和/或防火墙上部署采集器,从而能够尽量采集到每个NAT会话的数据流的指纹特征。通常,开启NAT功能的设备两侧都部署有上述核心节点,因此在这些核心节点上部署采集器,能够采集到每个NAT会话的数据流的指纹特征。以图3所示的数据中心为例,开启NAT功能的防火墙通常连接两个Spine交换机,数据流先经过一个Spine交换机,然后在防火墙中进行NAT转换,然后经过另一个Spine交换机,因此,在2个Spine交换机上部署采集器,能够采集到该数据流NAT转换前后的指纹特征。
第三种,在具有NAT功能的节点上部署采集器,以采集到每个NAT会话的数据流的指纹特征。
在第一种和第二种部署方式中,采集器通常部署在节点的接口的入方向,当然也可以部署在接口的出方向。在第三种部署方式中,采集器需要同时在接口的入方向和接口的出方向部署。
在一些可能的实现方式中,部署采集器(软件)的各个节点或者采集器(硬件)自己存储各自采集到的数据流的信息。
在一些可能的实现方式中,采集器将采集到的数据流信息发送给存储平台,由存储平台集中存储。
图4是本申请实施例提供的一种数据流识别方法的流程图。在一些可能的实现方式中,该方法可以由图1~图3所示的应用场景中的分析平台执行,分析平台与各个采集器连接,或者与存储平台连接,以获取数据流信息以实现该方法,或者,分析平台获取自身存储的数据流信息以实现该方法。在另一些可能的实现方式中,该方法还可以由图1~图3所示的应用场景中的网络设备执行。下文以该方法由分析平台执行为例进行说明。如图4所示,该方法包括如下步骤。
S11:获取第一数据流的指纹特征和建立时间。
在本申请实施例中,指纹特征包括如下特征中的至少一项:
首个数据包的IPID、首个数据包的载荷、首个数据包的载荷的哈希值、TCP三次握手过程中SYN报文的IPID、TCP三次握手过程中SYN报文的ISN、TCP三次握手过程中SYN-ACK报文的ISN。
示例性地,指纹特征为上述特征中的任一个,例如,指纹特征为首个数据包的IPID,或者,首个数据包的载荷的哈希值,或者,SYN报文的ISN。采用单独字段作为指纹特征,减少了计算和存储负担。
示例性地,指纹特征包括上述特征中的多个,例如,指纹特征包括SYN报文的ISN+首个数据包的IPID,或者,首个数据包的载荷+SYN报文的ISN,或者,首个数据包的载荷+SYN报文的ISN+首个数据包的IPID等。采用多个字段的组合作为指纹特征,减小了指纹冲突的可能性。
其中,IPID是IP层中用来标识IP数据包的字段,用来标识主机发出的每一个IP报文,长度16bit。ISN是在TCP建链的3次握手过程的SYN报文携带的用于通知给对端的初始的序列码,长度32bit。数据包的载荷是数据包内携带的应用层内容(也即数据包中TCP头之后的内容)。
如果采集过程中SYN报文或者首个数据包(简称首包)发生重传,采用重传过程中最后一次收到的SYN报文或数据包作为采集对象,采集指纹特征,保证各个采集器采集的指纹特征的一致。也就是说,针对一个数据流,若采集器在一段时间内接收多个SYN报文或多个首包,则采集器从最后一个接收到的SYN报文或最后一个接收到的首包获取指纹特征。
其中,首个数据包是指TCP三次握手建立连接后传输的第一个数据包。不同数据流需要传输的内容不
同,有的数据流需要多个数据包才能完成内容的传输,而有的数据流仅需要传输一个数据包就可以完成内容的传输。因此,本申请实施例采用首个数据包的信息作为指纹特征,能够保证可以为每个需要判断的数据流采集到指纹特征。另一方面,首个数据包即是TCP建链后第一个被传输的数据包,而且其序列号等于SYN报文携带的ISN+1,采集器可以很容易地确定首个数据包,因此,采集首个数据包所需的资源和时间相比后续数据包更少。
采集器在采集时,可以根据数据包的序列号来判断是否是首个数据包,通常首个数据包的序列号是SYN报文的ISN+1。而第二个数据包的序列号是SYN报文的ISN+1再加首个数据包的长度,因此,如果要采集首个数据包之后的数据包的特征,则存在判断数据包的序列号更加困难,占用资源和时间更多的问题。所以优先选用首个数据包进行指纹特征采集。
首个数据包的载荷可以是首个数据包的全部载荷,也可以是首个数据包的部分载荷,例如是从数据包包头开始截取的部分载荷,如500或1000字节长度的载荷。相应地,首个数据包的哈希值也可以是全部载荷的哈希值,也可以是部分载荷的哈希值,这里不再赘述。
由于首个数据包的载荷或首个数据包的载荷的哈希值可能仅对应部分长度载荷,因此,首个数据包的载荷除了可以包括载荷内容,还可以包括载荷长度,首个数据包的载荷的哈希值除了可以包括哈希值,还可以包括载荷长度。其中,载荷长度可以作为一个标识,标识载荷内容的长度或者哈希值对应的载荷内容的长度,属于可选项。该载荷长度可以用于筛选指纹特征,例如,在分析平台获取到各个采集器采集的载荷内容或者载荷的哈希值时,可以根据载荷长度判断各个采集器采集配置的采集载荷的长度是否相同,如果相同则可以使用采集的载荷内容或者载荷的哈希值作为指纹特征,否则不使用采集的载荷内容或者载荷的哈希值作为指纹特征。如果指纹特征仅针对部分载荷,在采集过程中,通过端口镜像直接截取所需长度的载荷进行特征提取即可。
在该实现方式中,各个采集器采集相同的指纹特征。申请人通过理论分析和大量的试验发现,上述所列举的指纹特征在NAT转换前后不会变化,并且不同数据流难以发生冲突,即,NAT转换前后的两个数据流尽管IP地址或端口可能不同,但两个数据流中的上述指纹特征相同,因此,上述指纹特征能够标识NAT相关的数据流。另外,上述所列举的指纹特征采集过程较为容易实现,占用资源少,有利于实现本申请实施例提供的方法。
建立时间指示第一数据流的建立时刻,例如,开始创建第一数据流的连接的时刻,或者,第一数据流的连接建立完成的时刻,或者开始发送第一数据流的数据报文的时刻。例如,建立时间可以是SYN报文接收时间、SYN-ACK报文接收时间、ACK报文接收时间、SYN报文发送时间、SYN-ACK报文发送时间、ACK报文发送时间或首个数据包的接收或发送时间等。此处的SYN、SYN-ACK和ACK报文为第一数据流对应的TCP三次握手中的报文。首个数据包是第一数据流对应的TCP连接创建完成后发送或接收的第一个数据包。
S12:根据第一数据流的指纹特征和建立时间,确定与第一数据流NAT相关的至少一个数据流。
与第一数据流NAT相关可以是与第一数据流属于同一个NAT会话。
一个NAT会话包括NAT转换前的数据流和NAT转换后的数据流。例如,第一数据流为NAT设备A转换前的数据流,第二数据流为NAT设备A对第一数据流进行NAT转换后获得的数据流,则第一数据流和第二数据流属于同一个NAT会话。NAT转换前的数据流和NAT转换后的数据流具有相同的指纹特征。例如,第一数据流的SYN报文的IPID和第二数据流的SYN报文的IPID相同,第一数据流的SYN报文的ISN和第二数据流的SYN报文的ISN相同,第一数据流的首包的IPID和第二数据流的首包的IPID相同,第一数据流的首包的载荷和第二数据流的首包的载荷相同,第一数据流的首包载荷的哈希值和第二数据流的首包载荷的哈希值相同,第一数据流的SYN-ACK报文的ISN和第二数据流的SYN-ACK报文的ISN。
与第一数据流NAT相关还可以是与第一数据流经过N次NAT转换后的数据流属于同一个NAT会话。
其中,N为大于等于1的自然数。例如,业务数据流从业务源端到业务目的端可能会经过多次NAT(例如,传输路径上部署了M个NAT设备)转换,假设第一数据流为从业务源端发出的数据流,则该数据流经过第一个NAT设备的NAT转换后变为第二数据流,第二数据流与第一数据流属于同一个NAT会话,第二数据流经过第二个NAT设备的NAT转换后变成第三数据流,第三数据流与第二数据流属于同一
个NAT会话,第三数据流与第一数据流NAT相关,以此类推,第M-1数据流经过第M-1个NAT设备转换后变成第M数据流,第M数据流与第M-1数据流属于同一个NAT会话,第M数据流与第一数据流NAT相关,第M数据流经过第M个NAT设备转换后变成第M+1数据流,第M+1数据流与第M数据流属于同一个NAT会话,第M+1数据流也与第一数据流NAT相关。其中,第二数据流是第一数据流经过一次NAT转换后的数据流,第三数据流与第二数据流属于同一个NAT会话,即,第三数据流与第一数据流经过一次NAT转换后的数据流属于同一个NAT会话。其中,第M-1数据流是第一数据流经过M-1次NAT转换后的数据流,第M数据流与第M-1数据流属于同一个NAT会话,即,第M数据流与第一数据流经过M-1次NAT转换后的数据流属于同一个NAT会话。其中,第M数据流是第一数据流经过M次NAT转换后的数据流,第M+1数据流与第M数据流属于同一个NAT会话,即,第M+1数据流与第一数据流经过M次NAT转换后的数据流属于同一个NAT会话。NAT转换前后的数据流的指纹特征相同,即,属于同一个NAT会话的数据流的指纹特征相同,而与第一数据流以及与第一数据流NAT相关的至少一个数据流中的每两个数据流均属于同一个NAT会话,因此,第一数据流和与第一数据流NAT相关的数据流的指纹特征均相同。
分析平台可以将指纹特征与第一数据流的指纹特征相同,且建立时间与第一数据流的建立时间接近的数据流确定为与第一数据流NAT相关的数据流。其中,建立时间与第一数据流的建立时间接近例如为建立时间与第一数据流的建立时间的差值的绝对值小于阈值。
在本申请实施例中,通过第一数据流的指纹特征确定与第一数据流NAT相关的至少一个数据流,不需要从NAT设备读取会话表,实现难度及代价小,并且网络部署难度小。
图5是本申请实施例提供的一种数据流识别方法的流程图。该方法可以由图1~图3所示的应用场景中的采集器、部署有采集器的网络设备或存储平台执行。如图5所示,该方法包括如下步骤。
S21:接收查询请求,查询请求包括第一数据流的第一流标识。
S22:发送查询应答,查询应答包括第一数据流的指纹特征和建立时间,第一数据流的指纹特征和建立时间用于确定与第一数据流NAT相关的至少一个数据流。
在接收到第一数据流的第一流标识后,以第一流标识为索引进行查询,以获取第一数据流的指纹特征和建立时间,然后反馈查询应答。
可选地,查询请求还可以包括查询时间范围,当查询请求包括查询时间范围时,查询在该查询时间范围内建立的数据流。
查询时间范围是用来辅助查找数据流的记录条件,如果不指定,系统可以采用默认值,如当天。
如果查询时间范围较大,查询到多条数据流,可以返回多条数据流的指纹特征作为应答,也可以通过从多条数据流中选择一条数据流的指纹特征作为应答,选择的依据可以是流的创建时间。例如,由采集器或存储平台自动选择创建时间最近的数据流的指纹特征作为应答。创建时间最近的数据流例如为采集器或存储平台最新记录的数据流。当查询请求还包括查询时间范围时,创建时间最近的数据流例如为查询时间范围内最新记录的数据流。或者,当查找到多条数据流时,将多条数据流的创建时间反馈给分析平台,分析平台提供选择界面由用户根据创建时间从多条数据流中选择,然后将选择结果反馈给采集器或存储平台,采集器或存储平台再根据选择反馈对应数据流的指纹特征。
图6是本申请实施例提供的一种数据流信息的采集方法的流程图。该方法由采集器执行。如图6所示,该方法包括如下步骤。
S31:采集各个数据流的建立时间、第一流标识和指纹特征。
在一些可能的实现方式中,采集器以旁路的方式接入。此时,采集各个数据流的建立时间、第一流标识和指纹特征,包括:
通过镜像功能抓取数据流的报文;
分析抓取到的数据流的报文,以获取数据流的建立时间、第一流标识和指纹特征。
其中,镜像功能可以是报文端口镜像或者远程镜像(例如封装式远程交换端口分析仪技术(encapsulated remote switch port analyzer,ERSPAN)远程镜像)。其中,报文端口镜像是对端口报文原封不动进行复制,
而远程镜像则是将端口报文封装后进行远程复制。
在本申请的实现方式中,当指纹特征采用TCP三次握手中的SYN报文和/或SYN-ACK报文的特征时,通过镜像功能抓取数据流的报文,包括:
通过配置ACL抓取SYN报文和/或SYN-ACK报文。
在配置镜像功能时,可以配置镜像功能的多个可选参数,其中可以包括ACL,通过ACL能够对指定报文或者指定报文长度进行过滤,实现SYN报文和/或SYN-ACK报文采集。
通过配置ACL仅抓取TCP建立过程中的3次握手报文以采集SYN报文和/或SYN-ACK报文的特征,可以减少需要抓取的镜像报文数量,减轻了网络设备的镜像负担和采集器的报文处理负担,有利于大规模部署。相比于首包的特征,使用SYN报文或SYN-ACK报文的特征作为指纹特征,处理负担更小。
在另一些可能的实现方式中,采集器以随路的方式接入,例如,转发设备也包含采集功能。此时,采集器直接分析数据流的报文,以获取数据流的建立时间、第一流标识和指纹特征即可。
示例性地,采集器在采集数据流的创建时间时,可以采用如下时间中的一种作为数据流的创建时间:
SYN报文接收时间、SYN-ACK报文接收时间、ACK报文接收时间。此处的ACK报文为TCP三次握手中的ACK报文。
其中,采集器在采集TCP三次握手中的ACK报文的接收时间时,可以以SYN报文或SYN-ACK报文为基准,采集SYN报文或SYN-ACK报文后的第一个ACK报文的接收时间,即为TCP三次握手中的ACK报文的接收时间。
当然上面针对的是接口的入方向的采集器,如果是接口的出方向的采集器,则采用如下时间中的一种作为数据流的创建时间:
SYN报文发送时间、SYN-ACK报文发送时间、ACK报文发送时间。
S32:存储各个数据流的建立时间、第一流标识和指纹特征。
在一些可能的实现方式中,采集器本地存储采集到数据,或者存储到部署所在的网络设备中。
在另一些可能的实现方式中,采集器将采集到的数据发送到存储平台进行集中式存储。
示例性地,数据流的第一流标识为数据流的五元组,包括源IP地址、目的IP地址、源端口、目的端口和协议类型,其中协议类型为TCP。
在本申请的实施例中,数据流的信息可以按照下表1的格式进行存储:
表1
其中,记录时间可以是采集时间,也可以是存储时间。由于记录时间和建立时间通常较为接近,因此,查询时间范围除了可以限定建立时间,也可以限定的是记录时间,对此不做限定。
数据流的指纹特征可以由1个特征或多个特征组成,表1中包括两个特征,仅为举例,不作为本申请的限制。
当然,采集器除了采集和存储上述流的信息外,还可以采集和记录包数、速率、丢包率等信息。
图7是本申请实施例提供的一种数据流识别方法的流程图。该方法以分析平台、第一网络节点和第二网络节点共同执行为例进行说明,示例性地,第一网络节点和第二网络节点为采集器或部署有采集器的网络设备。如图7所示,该方法包括如下步骤。
S41:分析平台向第一网络节点发送第一查询请求,第一查询请求包括第一数据流的第一流标识;第一网络节点接收第一查询请求。
这里的第一网络节点可以是一个也可以是多个。
第一查询请求除了包括流标识外,还可以包括指示请求的参数字段,例如第一查询请求可以通过携带不同标识指示要请求的不同参数,比如标识a指示请求数据流的建立时间,标识b指示请求数据流的指纹
特征。当然,前述指示请求的方式仅为一种示例,不作为本申请的限制。
S42:第一网络节点以第一流标识为索引进行查询。当查询到结果时,获取第一数据流流的指纹特征和建立时间,然后执行S43,否则不执行后续步骤。
这里的网络节点是采集器或部署有采集器的网络设备,存储有如表1所示的数据流的信息,因此可以采用第一流标识(五元组)为索引进行查询。
若分析平台向多个第一网络节点发送第一查询请求,则每个第一网络节点均执行步骤S42。
可选地,在没有查询到结果时,该方法还可以包括:第一网络节点向分析平台反馈查询失败消息,以通知分析平台未查询到结果。
S43:第一网络节点向分析平台发送第一查询应答,第一查询应答包括第一数据流的指纹特征和建立时间;分析平台接收第一查询应答。
S44:分析平台向第二网络节点发送第二查询请求,第二查询请求包括第一数据流的指纹特征;第二网络节点接收第二查询请求。
这里的第二网络节点可以是一个也可以是多个。第二网络节点和第一网络节点可以至少部分相同,也可以完全不同。
S45:第二网络节点以第一数据流的指纹特征为索引进行查询。当查询到结果时,获取与第一数据流的指纹特征相同的第二数据流的第一流标识和建立时间,然后执行S46,否则不执行后续步骤。
若分析平台向多个第二网络节点发送第二查询请求,则每个第二网络节点均执行步骤S45。
可选地,在没有查询到结果时,该方法还可以包括:第二网络节点向分析平台反馈查询失败消息,以通知分析平台未查询到结果。
S46:第二网络节点向分析平台发送第二查询应答,第二查询应答包括与第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间;分析平台接收第二查询应答。
若分析平台向多个第二网络节点发送第二查询请求,则分析平台有可能接收到不同的第二网络节点发送的第二查询应答。每个第二查询应答可能包含相同的第二数据流的第一流标识,这对应于第一数据流的源端和目的端的传输路径上仅存在一个NAT设备的场景。不同的第二查询应答可能包含不同的第二数据流的第一流标识,这对应于第一数据流的源端和目的端的传输路径上存在多个NAT设备的场景,相应地,分析平台可能获取到多个第二数据流的信息。
S47:分析平台根据第一数据流的建立时间和至少一个第二数据流的建立时间,确定与第一数据流NAT相关的数据流。
当分析平台仅获取到一个第二数据流时,在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,分析平台确定第一数据流和第二数据流属于同一个NAT会话,也即,第二数据流是第一数据流NAT相关的数据流。否则,分析平台确定第一数据流和第二数据流不属于同一个NAT会话。
由于NAT会话可能有多级,因此可能存在多个数据流和第一数据流NAT相关,因此在查询到多个指纹相同的数据流时,可以分别针对每个数据流进行判断,确定是否与第一数据流NAT相关。
其中,对于多级NAT,多级NAT的转换和传输延迟较小,第一数据流和经过多级NAT转换后的数据流的建立时间的时间差的绝对值小于设定的阈值,因此,采用上述方案能够准确判断两个数据流是否NAT相关。
不同网络节点上的采集器采集时存在同步精度问题,不同的采集器之间还会存在传输时延,因此,可以基于时间同步精度和/或传输时延确定上述阈值。
在一些可能的实现方式中,NAT相关的一个或多个数据流经过的路程短,传输延迟低,这种情况下,可以仅基于采集器的时间同步精度确定阈值。例如,采集器均不属于高速数据中心中,采集器间的传输时延为微妙级,采集器之间采用网络时间协议(network time protocol,NTP)进行同步,同步精度在毫秒级,相对于同步精度,传输时延可以忽略不计,因此,可以直接基于同步精度确定阈值,例如,阈值可以为10毫秒等。再例如,采集器之间的同步精度更高时,阈值可以设计得更小,例如1毫秒,2毫秒等。其中,一个时间同步协议的同步精度指示网络中的设备基于该时间同步协议进行时间同步时可能存在的时间误差,例如,设备A和设备B基于NTP协议进行时间同步,若同步精度为毫秒级,其表示设备A和设备B
的时间可能会相差几毫秒。
在另一些可能的实现方式中,NAT相关的数据流在采集点间的传输路径长,传输延迟高,相对于传输延迟,时间同步精度问题可以忽略不计,这种情况下,可以仅基于传输时延确定上述阈值。例如,两个采集点跨广域网,采集点间的传输延迟约为50~60ms,远超同步精度的几毫秒,此时可以直接基于传输延迟确定阈值,例如,阈值为100ms。
在本申请另一些可能的实现方式中,采集点间的传输延迟与时间同步精度接近,这种情况下,除了考虑采集器的时间同步精度,还需要考虑传输延迟,可以基于时间同步精度和传输时延确定阈值。例如,同步精度在毫秒级,传输延迟也是毫秒级,则阈值可以为20ms。
随着网络规模的增大和网络中部署业务的增多,有可能出现指纹冲突的情况,指纹冲突是指非NAT相关的数据流的指纹特征相同。为进一步提高准确性,可选地,在确定NAT相关的数据流时,还可以排除指纹冲突的数据流。可选地,在本申请实施例中,在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定第一数据流和第二数据流NAT相关,包括:
在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值,第一数据流和第二数据流不存在指纹冲突的情况下,确定第一数据流和第二数据流NAT相关,指纹冲突是指非NAT相关的数据流的指纹特征相同。
例如,两个数据流不是同一NAT会话中的数据流,但两个数据流的指纹特征相同,此时这两个数据流存在指纹冲突。
再例如,与第一数据流指纹特征相同的多个数据流,不是第一数据流经过N级NAT转换后的数据流,此时这多个数据流与第一数据流指纹冲突。
分析平台可以分别对每个第二数据流进行指纹冲突判断。
示例性地,确定指纹冲突的方式如下:
确定与第二数据流具有相同第二流标识的至少一条第三数据流;
确定与每条第三数据流具有相同指纹特征的至少一条第四数据流;
在至少一条第四数据流中的存在一条第四数据流的建立时间和对应第三数据流的建立时间的时间差的绝对值小于阈值,且这一条第四数据流的第二流标识和第一数据流的第二流标识相同的情况下,确定第一数据流和第二数据流不存在指纹冲突。
其中,数据流的第二流标识可以是四元组,包括源IP地址、目的IP地址、目的端口和协议类型。
该方案不仅可以判断仅存在单级NAT的场景下的指纹冲突,还可以判断存在多级NAT的场景下的指纹冲突。因为,上述确定指纹冲突的过程中,使用的原理是:
开启NAT功能的设备会对从相同主机发出到相同目的地的相同端口的数据流(也即四元组相同的数据流)都进行NAT转换。故,查询历史记录中四元组相同的数据流是否经过NAT转换,如果查询到存在四元组相同的数据流经过NAT转换,则认为本次也是NAT转换,而非指纹冲突。对应到存在多级NAT的场景下,如果存在四元组相同的数据流在历史记录中存在过多级NAT转换,则认为本次也是多级NAT转换。故上述判断过程能够确定每条第二数据流与第一数据流是指纹冲突,还是NAT相关。
其中,确定第三数据流时,不限定数据流的源端口是因为第二数据流和第三数据流可以是同一主机的不同应用产生的数据流,但由于目的地址和端口相同,传输时会进行同样的NAT转换。
例如,数据流a经过NAT转换得到数据流b,数据流b经过NAT转换得到数据流c。在历史记录中,与数据流a具有相同第二流标识的数据流A经过NAT转换得到数据流B,数据流B经过NAT转换得到数据流C。
在数据流识别时,确定出数据流b和c的指纹特征和a相同,并且建立时间的时间差的绝对值也小于阈值。此时进行指纹冲突判断:
第一数据流为数据流a,第二数据流为数据流c,确定与数据流c具有相同第二流标识的第三数据流为数据流C,确定与数据流C具有相同指纹特征的第四数据流为数据流B和数据流A。
在第四数据流中,数据流A的建立时间和数据流C的建立时间的时间差的绝对值小于阈值,并且数据流A的第二流标识和数据流a相同,此时确定不存在指纹冲突。
下面结合表2和表3提供的数据流信息来对指纹冲突进行说明。
输入的第一数据流为flow11,根据指纹匹配能查询到第二数据流flow21和flow31可能是NAT相关的数据流,如下表2所示:
表2
分别查询具有相同第二流标识的其他流,可以得到3组流信息,如下表3所示:
表3
根据第二数据流flow21和flow31分别查询到第三数据流flow22和flow32。
其中,第三数据流flow22对应的至少一条第四数据流为flow12,第三数据流flow32没有对应的第四数据流。
由于第三数据流flow22对应的至少一条第四数据流为flow12符合建立时间的时间差的绝对值小于阈值,且与第一数据流的第二流标识相同的条件,因此可以判断第二数据流flow21和第一数据流flow11不存在指纹冲突。而第二数据流flow31和第一数据流flow11存在指纹冲突。
可选地,该分析平台可以是网络诊断设备,相应地,该方法还包括:
S48:分析平台在第一数据流发生故障的情况下,对第一数据流以及与第一数据流NAT相关的至少一个数据流进行故障诊断。
例如,故障诊断设备在确定第一数据流存在丢包率或延迟较高的故障后,先确定与第一数据流NAT相关的至少一个第二数据流,然后针对第一数据流和至少一个第二数据流,分析丢包率或延迟高的原因,确定链路上造成丢包或延迟的故障点,完成故障诊断。
在一些场景中,网络诊断设备在第一数据流根据数据包的ACK报文,确定存在丢包,但丢包的位置并非在第一数据流经过的链路中,此时针对第一数据流和第二数据流进行故障诊断,确定出丢包的位置在第二数据流经过的链路中,并且确定出造成丢包的故障点位置,即实现了更精确地定位或定界网络故障。
在另一些场景中,网络诊断设备发现第一数据流的数据包传输延迟较大,此时针对第一数据流和第二数据流进行故障诊断,确定出造成传输延迟的故障点位置在第二数据流所在链路上,即实现了更精确地定位或定界网络故障。
本申请实施例以数据流为粒度进行故障诊断,原因在于一些场景下,造成数据包丢包或者传输的延迟的原因是第二数据流所在链路上的设备对该数据流的突发流量进行了限速,这种情况不涉及链路上的其他数据流,以数据流为粒度诊断更准确。
可见,通过本申请提供的数据流识别方案能够确定出NAT相关的数据流,这使得分析平台不仅可以针对原数据流进行故障诊断,还可以继续针对NAT相关的数据流进行故障诊断,可以更精确地定位或定界网络故障。
可选地,在确定不存在与第一数据流NAT相关的数据流时,可以确定不存在与第一数据流属于同一个NAT会话的数据流。可选地,在确定不存在与第一数据流NAT相关的数据流时,该方法还包括:
分析平台在第一数据流发生故障的情况下,对第一数据流进行故障诊断。
图8是本申请实施例提供的一种数据流识别方法的流程图。该方法以分析平台、第一网络节点和第二网络节点共同执行为例进行说明,示例性地,第一网络节点为存储平台。如图8所示,该方法包括如下步骤。
S51:分析平台向第一网络节点发送第一查询请求,第一查询请求包括第一数据流的第一流标识;第一网络节点接收第一查询请求。
S52:第一网络节点以第一流标识为索引进行查询。当查询到结果时,获取第一数据流流的指纹特征和建立时间,然后执行S53,否则不执行后续步骤。
可选地,在没有查询到结果时,该方法还可以包括:第一网络节点向分析平台反馈查询失败消息,以通知分析平台未查询到结果。
S53:第一网络节点以第一数据流的指纹特征为索引进行查询。当查询到结果时,获取与第一数据流的指纹特征相同的第二数据流的第一流标识和建立时间,然后执行S54,否则不执行后续步骤。
可选地,在没有查询到结果时,该方法还可以包括:第一网络节点向分析平台反馈查询失败消息,以通知分析平台未查询到结果。
S54:第一网络节点根据第一数据流和第二数据流的建立时间,确定第一数据流和第二数据流是否NAT相关。当确定第一数据流和第二数据流NAT相关时,执行S55,否则不执行后续步骤。
在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定第一数据流和第二数据流NAT相关。否则,确定第一数据流和第二数据流非NAT相关。
可选地,在本申请实施例中,在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定第一数据流和第二数据流NAT相关,包括:
在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值,第一数据流和第二数据流不存在指纹冲突的情况下,确定第一数据流和第二数据流NAT相关,指纹冲突是指非NAT相关的数据流的指纹特征相同。
确定指纹冲突的详细过程参见步骤S47,在此不做赘述。
可选地,在确定第一数据流和第二数据流非NAT相关时,该方法还可以包括:第一网络节点向分析平台反馈通知消息,以通知分析平台没有与第一数据流NAT相关的数据流。
S55:第一网络节点向分析平台发送第三查询应答,第三查询应答包括与第一数据流NAT相关的第二数据流的第一流标识;分析平台接收第三查询应答。
S56:分析平台根据第三查询应答,确定第一数据流和第二数据流NAT相关。
S57:分析平台在第一数据流发生故障的情况下,对第一数据流以及与第一数据流NAT相关的至少一个数据流进行故障诊断。
步骤S57的详细过程参考步骤S48。
图9是本申请实施例提供的一种数据流识别方法的流程图。该方法以分析平台和第一网络节点共同执行为例进行说明,示例性地,第一网络节点为存储平台。如图9所示,该方法包括如下步骤。
S61:分析平台向第一网络节点发送第一查询请求,第一查询请求包括第一数据流的第一流标识;第一网络节点接收第一查询请求。
S62:第一网络节点以第一流标识为索引进行查询。当查询到结果时,获取第一数据流流的指纹特征和建立时间,然后执行S63,否则不执行后续步骤。
S63:第一网络节点以第一数据流的指纹特征为索引进行查询。当查询到结果时,获取与第一数据流的指纹特征相同的第二数据流的第一流标识和建立时间,然后执行S64,否则不执行后续步骤。
可选地,在没有查询到结果时,该方法还可以包括:第一网络节点向分析平台反馈查询失败消息,以通知分析平台未查询到结果。
S64:第一网络节点向分析平台发送第一查询应答,第一查询应答包括第一数据流的指纹特征和建立时间、与第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间;分析平台接收第一查询应答。
S65:分析平台根据第一数据流和第二数据流的建立时间,确定第一数据流和第二数据流是否NAT相
关。
在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定第一数据流和第二数据流NAT相关。否则,确定第一数据流和第二数据流非NAT相关。
可选地,在本申请实施例中,在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定第一数据流和第二数据流NAT相关,包括:
在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值,第一数据流和第二数据流不存在指纹冲突的情况下,确定第一数据流和第二数据流NAT相关,指纹冲突是指非NAT相关的数据流的指纹特征相同。
确定指纹冲突的详细过程参见步骤S47,在此不做赘述。
S66:分析平台在第一数据流发生故障的情况下,对第一数据流以及与第一数据流NAT相关的至少一个数据流进行故障诊断。
步骤S66的详细过程参考步骤S48。
图10是本申请实施例提供的一种数据流识别装置的框图。该数据流识别装置可以通过软件、硬件或者两者的结合实现成为分析平台、网络设备、采集器或存储平台的全部或者一部分。该数据流识别装置可以包括:获取单元701和确定单元702。
其中,获取单元701,用于获取第一数据流的指纹特征和建立时间;
确定单元702,用于根据第一数据流的指纹特征和建立时间,确定与第一数据流NAT相关的至少一个数据流。
可选地,获取单元701,用于向第一网络节点发送第一查询请求,第一查询请求包括第一数据流的第一流标识;接收第一网络节点发送的第一查询应答,第一查询应答包括第一数据流的指纹特征和建立时间。
可选地,确定单元702,用于向第二网络节点发送第二查询请求,第二查询请求包括第一数据流的指纹特征;接收第二网络节点发送的第二查询应答,第二查询应答包括与第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间;在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定第一数据流和第二数据流NAT相关。
可选地,第一查询应答还包括与第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间;
确定单元702,用于在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定第一数据流和第二数据流NAT相关。
可选地,获取单元701,用于接收第一查询请求,第一查询请求包括第一数据流的第一流标识;根据第一数据流的第一流标识确定第一数据流的指纹特征和建立时间。
可选地,确定单元702,用于根据第一数据流的指纹特征确定第二数据流,第二数据流的指纹特征与第一数据流的指纹特征相同;在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定第一数据流和第二数据流NAT相关。
可选地,指纹特征包括如下特征中的至少一项:
首个数据包的互联网协议标识IPID、首个数据包的载荷、首个数据包的载荷的哈希值、传输控制协议TCP三次握手过程中SYN报文的IPID、TCP三次握手过程中SYN报文的初始序列码ISN、TCP三次握手过程中SYN-ACK报文的ISN。
可选地,第一查询请求还包括查询时间范围,查询时间范围用于限制查询到的数据流的建立时间所处的时间范围。
可选地,确定单元702,用于在第二数据流的建立时间与第一数据流的建立时间的时间差的绝对值小于阈值,第一数据流和第二数据流不存在指纹冲突的情况下,确定第一数据流和第二数据流NAT相关,指纹冲突是指非NAT相关的数据流的指纹特征相同。
可选地,确定单元702,还用于确定与第二数据流具有相同第二流标识的至少一条第三数据流;确定与每条第三数据流具有相同指纹特征的至少一条第四数据流;在至少一条第四数据流中存在一条第四数据流的建立时间和对应第三数据流的建立时间的时间差的绝对值小于阈值,且这一条第四数据流的第二流标识和第一数据流的第二流标识相同的情况下,确定第一数据流和第二数据流不存在指纹冲突。
可选地,该装置还包括:
诊断单元703,用于在第一数据流发生故障的情况下,对第一数据流以及与第一数据流NAT相关的至少一个数据流进行故障诊断。
需要说明的是,上述实施例提供的数据流识别装置在进行数据流识别时,仅以上述各功能单元的划分进行举例说明,实际应用中,可以根据需要而将上述功能分配由不同的功能单元完成,即将设备的内部结构划分成不同的功能单元,以完成以上描述的全部或者部分功能。另外,上述实施例提供的数据流识别装置与数据流识别方法实施例属于同一构思,其具体实现过程详见方法实施例,这里不再赘述。
图11是本申请实施例提供的一种数据流识别装置的框图。该显示装置可以通过软件、硬件或者两者的结合实现成为采集器、部署有采集器的网络设备或存储平台的全部或者一部分。该显示装置可以包括:接收单元801和发送单元802。
其中,接收单元801,用于接收查询请求,查询请求包括第一数据流的第一流标识;
发送单元802,用于发送查询应答,查询应答包括第一数据流的指纹特征和建立时间,第一数据流的指纹特征和建立时间用于确定与第一数据流NAT相关的至少一个数据流。
可选地,查询应答还包括与第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间。
可选地,该装置还包括:
采集单元803,用于采集各个数据流的建立时间、第一流标识和指纹特征;
存储单元804,用于存储各个数据流的建立时间、第一流标识和指纹特征。
可选地,采集单元803,用于通过镜像功能抓取数据流的报文;分析抓取到的数据流的报文,以获取数据流的建立时间、第一流标识和指纹特征。
需要说明的是,上述实施例提供的数据流识别装置在进行数据流识别时,仅以上述各功能单元的划分进行举例说明,实际应用中,可以根据需要而将上述功能分配由不同的功能单元完成,即将设备的内部结构划分成不同的功能单元,以完成以上描述的全部或者部分功能。另外,上述实施例提供的数据流识别装置与数据流识别方法实施例属于同一构思,其具体实现过程详见方法实施例,这里不再赘述。
上述各个附图对应的流程的描述各有侧重,某个流程中没有详述的部分,可以参见其他流程的相关描述。
图12示出了本申请实施例提供的电子设备150的结构示意图。该电子设备可以为分析平台、网络设备、采集器或存储平台。图12所示的电子设备150用于执行上述图4至图9任一幅所示的数据流识别方法所涉及的操作。该电子设备150可以由一般性的总线体系结构来实现。
如图12所示,电子设备150包括至少一个处理器151、存储器153以及至少一个通信接口154。
处理器151例如是通用中央处理器(central processing unit,CPU)、数字信号处理器(digital signal processor,DSP)、网络处理器(network processer,NP)、数据处理单元(Data Processing Unit,DPU)、微处理器或者一个或多个用于实现本申请方案的集成电路。例如,处理器151包括专用集成电路(application-specific integrated circuit,ASIC),可编程逻辑器件(programmable logic device,PLD)或者其他可编程逻辑器件、晶体管逻辑器件、硬件部件或者其任意组合。PLD例如是复杂可编程逻辑器件(complex programmable logic device,CPLD)、现场可编程逻辑门阵列(field-programmable gate array,FPGA)、通用阵列逻辑(generic array logic,GAL)或其任意组合。其可以实现或执行结合本申请实施例公开内容所描述的各种逻辑方框、模块和电路。所述处理器也可以是实现计算功能的组合,例如包括一个或多个微处理器组合,DSP和微处理器的组合等等。
可选的,电子设备150还包括总线。总线用于在电子设备150的各组件之间传送信息。总线可以是外设部件互连标准(peripheral component interconnect,简称PCI)总线或扩展工业标准结构(extended industry standard architecture,简称EISA)总线等。总线可以分为地址总线、数据总线、控制总线等。为便于表示,图12中仅用一条粗线表示,但并不表示仅有一根总线或一种类型的总线。
存储器153例如是只读存储器(read-only memory,ROM)或可存储静态信息和指令的其它类型的静态存储设备,又如是随机存取存储器(random access memory,RAM)或者可存储信息和指令的其它类型
的动态存储设备,又如是电可擦可编程只读存储器(electrically erasable programmable read-only Memory,EEPROM)、只读光盘(compact disc read-only memory,CD-ROM)或其它光盘存储、光碟存储(包括压缩光碟、激光碟、光碟、数字通用光碟、蓝光光碟等)、磁盘存储介质或者其它磁存储设备,或者是能够用于携带或存储具有指令或数据结构形式的期望的程序代码并能够由计算机存取的任何其它介质,但不限于此。存储器153例如是独立存在,并通过总线与处理器151相连接。存储器153也可以和处理器151集成在一起。
通信接口154使用任何收发器一类的装置,用于与其它设备或通信网络通信,通信网络可以为以太网、无线接入网(RAN)或无线局域网(wireless local area networks,WLAN)等。通信接口154可以包括有线通信接口,还可以包括无线通信接口。具体的,通信接口154可以为以太(Ethernet)接口、快速以太(Fast Ethernet,FE)接口、千兆以太(Gigabit Ethernet,GE)接口,异步传输模式(Asynchronous Transfer Mode,ATM)接口,无线局域网(wireless local area networks,WLAN)接口,蜂窝网络通信接口或其组合。以太网接口可以是光接口,电接口或其组合。在本申请实施例中,通信接口154可以用于电子设备150与其他设备进行通信。
在具体实现中,作为一种实施例,处理器151可以包括一个或多个CPU,如图12中所示的CPU0和CPU1。这些处理器中的每一个可以是一个单核(single-CPU)处理器,也可以是一个多核(multi-CPU)处理器。这里的处理器可以指一个或多个设备、电路、和/或用于处理数据(例如计算机程序指令)的处理核。
在具体实现中,作为一种实施例,电子设备150可以包括多个处理器,如图12中所示的处理器151和处理器155。这些处理器中的每一个可以是一个单核处理器(single-CPU),也可以是一个多核处理器(multi-CPU)。这里的处理器可以指一个或多个设备、电路、和/或用于处理数据(如计算机程序指令)的处理核。
在具体实现中,作为一种实施例,电子设备150还可以包括输出设备和输入设备。输出设备和处理器151通信,可以以多种方式来显示信息。例如,输出设备可以是液晶显示器(liquid crystal display,LCD)、发光二级管(light emitting diode,LED)显示设备、阴极射线管(cathode ray tube,CRT)显示设备或投影仪(projector)等。输入设备和处理器151通信,可以以多种方式接收用户的输入。例如,输入设备可以是鼠标、键盘、触摸屏设备或传感设备等。
在一些实施例中,存储器153用于存储执行本申请方案的程序代码1510,处理器151可以执行存储器153中存储的程序代码1510。也即是,电子设备150可以通过处理器151执行存储器153中的程序代码1510,来实现方法实施例提供的数据流识别方法。程序代码1510中可以包括一个或多个软件模块。可选地,处理器151自身也可以存储执行本申请方案的程序代码或指令。
在具体实施例中,本申请实施例的电子设备150可对应于上述各个方法实施例中的控制器,电子设备150中的处理器151读取存储器153中的指令,使图12所示的电子设备150能够执行控制器所执行的全部或部分操作。
具体的,处理器151用于获取第一数据流的指纹特征和建立时间;根据所述第一数据流的指纹特征和建立时间,确定与所述第一数据流NAT相关的至少一个数据流。
其他可选的实施方式,为了简洁,在此不再赘述。
其中,图4至图9任一幅所示的数据流识别方法的各步骤通过电子设备150的处理器中的硬件的集成逻辑电路或者软件形式的指令完成。结合本申请实施例所公开的方法的步骤可以直接体现为硬件处理器执行完成,或者用处理器中的硬件及软件模块组合执行完成。软件模块可以位于随机存储器,闪存、只读存储器,可编程只读存储器或者电可擦写可编程存储器、寄存器等本领域成熟的存储介质中。该存储介质位于存储器,处理器读取存储器中的信息,结合其硬件完成上述方法的步骤,为避免重复,这里不再详细描述。
本申请实施例还提供了一种芯片,包括:输入接口、输出接口、处理器和存储器。输入接口、输出接口、处理器以及存储器之间通过内部连接通路相连。处理器用于执行存储器中的代码,当代码被执行时,处理器用于执行上述任一种的数据流识别方法。
应理解的是,上述处理器可以是CPU,还可以是其他通用处理器、DSP、ASIC、FPGA或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。通用处理器可以是微处理器或者是任何常规
的处理器等。值得说明的是,处理器可以是支持ARM架构的处理器。
进一步地,在一种可选的实施例中,上述处理器为一个或多个,存储器为一个或多个。可选地,存储器可以与处理器集成在一起,或者存储器与处理器分离设置。上述存储器可以包括只读存储器和随机存取存储器,并向处理器提供指令和数据。存储器还可以包括非易失性随机存取存储器。例如,存储器还可以存储参考块和目标块。
该存储器可以是易失性存储器或非易失性存储器,或可包括易失性和非易失性存储器两者。其中,非易失性存储器可以是ROM、PROM、EPROM、EEPROM或闪存。易失性存储器可以是RAM,其用作外部高速缓存。通过示例性但不是限制性说明,许多形式的RAM可用。例如,SRAM、DRAM、SDRAM、DDR SDRAM、ESDRAM、SLDRAM和DR RAM。
本申请实施例中,还提供了一种计算机可读存储介质,计算机可读存储介质存储有计算机指令,当计算机可读存储介质中存储的计算机指令被电子设备执行时,使得电子设备执行上述所提供的数据流识别方法。
本申请实施例中,还提供了一种包含指令的计算机程序产品,当其在电子设备上运行时,使得电子设备执行上述所提供的数据流识别方法。
在上述实施例中,可以全部或部分地通过软件、硬件、固件或者其任意组合来实现。当使用软件实现时,可以全部或部分地以计算机程序产品的形式实现。所述计算机程序产品包括一个或多个计算机指令。在计算机上加载和执行所述计算机程序指令时,全部或部分地产生按照本申请所述的流程或功能。所述计算机可以是通用计算机、专用计算机、计算机网络、或者其他可编程装置。所述计算机指令可以存储在计算机可读存储介质中,或者从一个计算机可读存储介质向另一个计算机可读存储介质传输,例如,所述计算机指令可以从一个网站站点、计算机、服务器或数据中心通过有线(例如同轴电缆、光纤、数字用户线)或无线(例如红外、无线、微波等)方式向另一个网站站点、计算机、服务器或数据中心进行传输。所述计算机可读存储介质可以是计算机能够存取的任何可用介质或者是包含一个或多个可用介质集成的服务器、数据中心等数据存储设备。所述可用介质可以是磁性介质,(例如,软盘、硬盘、磁带)、光介质(例如,DVD)、或者半导体介质(例如固态硬盘Solid State Disk)等。
本领域普通技术人员可以理解实现上述实施例的全部或部分步骤可以通过硬件来完成,也可以通过程序来指令相关的硬件完成,所述的程序可以存储于一种计算机可读存储介质中,上述提到的存储介质可以是只读存储器,磁盘或光盘等。
以上所述仅为本申请的可选实施例,但本申请的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本申请揭露的技术范围内,可轻易想到的变化或替换,都应涵盖在本申请的保护范围之内。因此,本申请的保护范围应该以权利要求的保护范围为准。
除非另作定义,此处使用的技术术语或者科学术语应当为本申请所属领域内具有一般技能的人士所理解的通常意义。本申请专利申请说明书以及权利要求书中使用的“第一”、“第二”、“第三”以及类似的词语并不表示任何顺序、数量或者重要性,而只是用来区分不同的组成部分。同样,“一个”或者“一”等类似词语也不表示数量限制,而是表示存在至少一个。“包括”或者“包含”等类似的词语意指出现在“包括”或者“包含”前面的元件或者物件涵盖出现在“包括”或者“包含”后面列举的元件或者物件及其等同,并不排除其他元件或者物件。
以上仅为本申请一个实施例,并不用以限制本申请,凡在本申请的精神和原则之内,所作的任何修改、等同替换、改进等,均应包含在本申请的保护范围之内。
Claims (35)
- 一种数据流识别方法,其特征在于,所述方法包括:获取第一数据流的指纹特征和建立时间;根据所述第一数据流的指纹特征和建立时间,确定与所述第一数据流网络地址转换NAT相关的至少一个数据流。
- 根据权利要求1所述的方法,其特征在于,所述获取第一数据流的指纹特征和建立时间,包括:向第一网络节点发送第一查询请求,所述第一查询请求包括所述第一数据流的第一流标识;接收所述第一网络节点发送的第一查询应答,所述第一查询应答包括所述第一数据流的指纹特征和建立时间。
- 根据权利要求1或2所述的方法,其特征在于,所述根据所述第一数据流的指纹特征和建立时间,确定与所述第一数据流NAT相关的至少一个数据流,包括:向第二网络节点发送第二查询请求,所述第二查询请求包括所述第一数据流的指纹特征;接收所述第二网络节点发送的第二查询应答,所述第二查询应答包括与所述第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间;在所述第二数据流的建立时间与所述第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定所述第一数据流和所述第二数据流NAT相关。
- 根据权利要求2所述的方法,其特征在于,所述第一查询应答还包括与所述第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间;所述根据所述第一数据流的指纹特征和建立时间,确定与所述第一数据流NAT相关的至少一个数据流,包括:在所述第二数据流的建立时间与所述第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定所述第一数据流和所述第二数据流NAT相关。
- 根据权利要求1所述的方法,其特征在于,所述获取第一数据流的指纹特征和建立时间,包括:接收第一查询请求,所述第一查询请求包括所述第一数据流的第一流标识;根据所述第一数据流的第一流标识确定所述第一数据流的指纹特征和建立时间。
- 根据权利要求1或5所述的方法,其特征在于,所述根据所述第一数据流的指纹特征和建立时间,确定与所述第一数据流NAT相关的至少一个数据流,包括:根据所述第一数据流的指纹特征确定第二数据流,所述第二数据流的指纹特征与所述第一数据流的指纹特征相同;在所述第二数据流的建立时间与所述第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定所述第一数据流和所述第二数据流NAT相关。
- 根据权利要求1至6任一项所述的方法,其特征在于,所述指纹特征包括如下特征中的至少一项:首个数据包的互联网协议标识IPID、首个数据包的载荷、首个数据包的载荷的哈希值、传输控制协议TCP三次握手过程中同步序列号SYN报文的IPID、TCP三次握手过程中SYN报文的初始序列码ISN、TCP三次握手过程中SYN-ACK报文的ISN。
- 根据权利要求2、4或5所述的方法,其特征在于,所述第一查询请求还包括查询时间范围,所述查询时间范围用于限制查询到的数据流的建立时间所处的时间范围。
- 根据权利要求3、4或6所述的方法,其特征在于,所述在所述第二数据流的建立时间与所述第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定所述第一数据流和所述第二数据流NAT相关,包括:在所述第二数据流的建立时间与所述第一数据流的建立时间的时间差的绝对值小于阈值,所述第一数据流和所述第二数据流不存在指纹冲突的情况下,确定所述第一数据流和所述第二数据流NAT相关,所述指纹冲突是指非NAT相关的数据流的指纹特征相同。
- 根据权利要求9所述的方法,其特征在于,所述方法还包括:确定与所述第二数据流具有相同第二流标识的至少一条第三数据流;确定与每条所述第三数据流具有相同指纹特征的至少一条第四数据流;在所述至少一条第四数据流中存在一条第四数据流的建立时间和对应所述第三数据流的建立时间的时间差的绝对值小于阈值,且所述一条第四数据流的第二流标识和所述第一数据流的第二流标识相同的情况下,确定所述第一数据流和所述第二数据流不存在指纹冲突。
- 根据权利要求1至10任一项所述的方法,其特征在于,所述方法还包括:在所述第一数据流发生故障的情况下,对所述第一数据流以及与所述第一数据流NAT相关的至少一个数据流进行故障诊断。
- 一种数据流识别方法,其特征在于,所述方法包括:接收查询请求,所述查询请求包括第一数据流的第一流标识;发送查询应答,所述查询应答包括所述第一数据流的指纹特征和建立时间,所述第一数据流的指纹特征和建立时间用于确定与所述第一数据流NAT相关的至少一个数据流。
- 根据权利要求12所述的方法,其特征在于,所述查询应答还包括与所述第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间。
- 根据权利要求12或13所述的方法,其特征在于,所述方法还包括:采集各个数据流的建立时间、第一流标识和指纹特征;存储所述各个数据流的建立时间、第一流标识和指纹特征。
- 根据权利要求14所述的方法,其特征在于,所述采集各个数据流的建立时间、第一流标识和指纹特征,包括:通过镜像功能抓取数据流的报文;分析抓取到的所述数据流的报文,以获取所述数据流的建立时间、第一流标识和指纹特征。
- 一种数据流识别装置,其特征在于,所述装置包括:获取单元,用于获取第一数据流的指纹特征和建立时间;确定单元,用于根据所述第一数据流的指纹特征和建立时间,确定与所述第一数据流NAT相关的至少一个数据流。
- 根据权利要求16所述的装置,其特征在于,所述获取单元,用于向第一网络节点发送第一查询请求,所述第一查询请求包括所述第一数据流的第一流标识;接收所述第一网络节点发送的第一查询应答,所述第一查询应答包括所述第一数据流的指纹特征和建立时间。
- 根据权利要求16或17所述的装置,其特征在于,所述确定单元,用于向第二网络节点发送第二查询请求,所述第二查询请求包括所述第一数据流的指纹特征;接收所述第二网络节点发送的第二查询应答,所述第二查询应答包括与所述第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间;在所述第二数据流的建立时间与所述第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定所述第一数据流和所述第二数据流NAT相关。
- 根据权利要求17所述的装置,其特征在于,所述第一查询应答还包括与所述第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间;所述确定单元,用于在所述第二数据流的建立时间与所述第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定所述第一数据流和所述第二数据流NAT相关。
- 根据权利要求16所述的装置,其特征在于,所述获取单元,用于接收第一查询请求,所述第一查询请求包括所述第一数据流的第一流标识;根据所述第一数据流的第一流标识确定所述第一数据流的指纹特征和建立时间。
- 根据权利要求16或20所述的装置,其特征在于,所述确定单元,用于根据所述第一数据流的指纹特征确定第二数据流,所述第二数据流的指纹特征与所述第一数据流的指纹特征相同;在所述第二数据流的建立时间与所述第一数据流的建立时间的时间差的绝对值小于阈值的情况下,确定所述第一数据流和所述第二数据流NAT相关。
- 根据权利要求16至21任一项所述的装置,其特征在于,所述指纹特征包括如下特征中的至少一项:首个数据包的互联网协议标识IPID、首个数据包的载荷、首个数据包的载荷的哈希值、传输控制协议TCP三次握手过程中SYN报文的IPID、TCP三次握手过程中SYN报文的初始序列码ISN、TCP三次握手过程中SYN-ACK报文的ISN。
- 根据权利要求17、19或20所述的装置,其特征在于,所述第一查询请求还包括查询时间范围,所述查询时间范围用于限制查询到的数据流的建立时间所处的时间范围。
- 根据权利要求18、19或21所述的装置,其特征在于,所述确定单元,用于在所述第二数据流的建立时间与所述第一数据流的建立时间的时间差的绝对值小于阈值,所述第一数据流和所述第二数据流不存在指纹冲突的情况下,确定所述第一数据流和所述第二数据流NAT相关,所述指纹冲突是指非NAT相关的数据流的指纹特征相同。
- 根据权利要求24所述的装置,其特征在于,所述确定单元,还用于确定与所述第二数据流具有相同第二流标识的至少一条第三数据流;确定与每条所述第三数据流具有相同指纹特征的至少一条第四数据流;在所述至少一条第四数据流中存在一条第四数据流的建立时间和对应所述第三数据流的建立时间的时间差的绝对值小于阈值,且所述一条第四数据流的第二流标识和所述第一数据流的第二流标识相同的情况下,确定所述第一数据流和所述第二数据流不存在指纹冲突。
- 根据权利要求16至25任一项所述的装置,其特征在于,所述装置还包括:诊断单元,用于在所述第一数据流发生故障的情况下,对所述第一数据流以及与所述第一数据流NAT相关的至少一个数据流进行故障诊断。
- 一种数据流识别装置,其特征在于,所述装置包括:接收单元,用于接收查询请求,所述查询请求包括第一数据流的第一流标识;发送单元,用于发送查询应答,所述查询应答包括所述第一数据流的指纹特征和建立时间,所述第一数据流的指纹特征和建立时间用于确定与所述第一数据流NAT相关的至少一个数据流。
- 根据权利要求27所述的装置,其特征在于,所述查询应答还包括与所述第一数据流的指纹特征相同的第二数据流的第一流标识以及建立时间。
- 根据权利要求27或28所述的装置,其特征在于,所述装置还包括:采集单元,用于采集各个数据流的建立时间、第一流标识和指纹特征;存储单元,用于存储所述各个数据流的建立时间、第一流标识和指纹特征。
- 根据权利要求29所述的装置,其特征在于,所述采集单元,用于通过镜像功能抓取数据流的报文;分析抓取到的所述数据流的报文,以获取所述数据流的建立时间、第一流标识和指纹特征。
- 一种电子设备,其特征在于,所述电子设备包括处理器和存储器,所述存储器用于存储软件程序,所述处理器通过运行或执行存储在所述存储器内的软件程序,以使所述电子设备实现如权利要求1至15任一项所述的方法。
- 一种数据流识别系统,其特征在于,所述数据流识别系统包括采集器和分析平台,所述分析平台用于执行如权利要求1至11任一项所述的方法,所述采集器用于执行如权利要求12至15任一项所述的方法。
- 根据权利要求32所述的系统,其特征在于,所述数据流识别系统还包括存储平台,用于存储所述采集器采集的数据流的信息。
- 一种计算机可读存储介质,其特征在于,所述计算机可读存储介质用于存储处理器所执行的程序代码,所述程序代码包括用于实现如权利要求1至15任一项所述的方法的指令。
- 一种计算机程序产品,其特征在于,包括程序代码,当计算机运行所述计算机程序产品时,使得所述计算机执行如权利要求1至15任一项所述的方法。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24752685.8A EP4654540A4 (en) | 2023-02-07 | 2024-01-22 | METHOD, APPARATUS AND ELECTRONIC DEVICE FOR DATA FLOW IDENTIFICATION |
| US19/290,640 US20250358256A1 (en) | 2023-02-07 | 2025-08-05 | Data flow identification method and apparatus and electronic device |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310129577.1 | 2023-02-07 | ||
| CN202310129577.1A CN118473894A (zh) | 2023-02-07 | 2023-02-07 | 数据流识别方法及装置、电子设备 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US19/290,640 Continuation US20250358256A1 (en) | 2023-02-07 | 2025-08-05 | Data flow identification method and apparatus and electronic device |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024164828A1 true WO2024164828A1 (zh) | 2024-08-15 |
Family
ID=92154563
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/073409 Ceased WO2024164828A1 (zh) | 2023-02-07 | 2024-01-22 | 数据流识别方法及装置、电子设备 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US20250358256A1 (zh) |
| EP (1) | EP4654540A4 (zh) |
| CN (1) | CN118473894A (zh) |
| WO (1) | WO2024164828A1 (zh) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US12155691B2 (en) * | 2022-04-25 | 2024-11-26 | At&T Intellectual Property I, L.P. | Detecting and mitigating denial of service attacks over home gateway network address translation |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101150505A (zh) * | 2007-07-31 | 2008-03-26 | 杭州华三通信技术有限公司 | 通过网络地址转换转发数据流的方法和装置 |
| US20190114350A1 (en) * | 2017-10-17 | 2019-04-18 | Salesforce.Com, Inc. | Systems, Methods, and Apparatuses for Implementing Concurrent Dataflow Execution with Write Conflict Protection Within a Cloud Based Computing Environment |
| CN112468373A (zh) * | 2020-12-08 | 2021-03-09 | 武汉蜘易科技有限公司 | 一种指纹设备网络流量精准定位分析系统和方法 |
| CN114915566A (zh) * | 2021-01-28 | 2022-08-16 | 腾讯科技(深圳)有限公司 | 应用识别方法、装置、设备及计算机可读存储介质 |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US8219675B2 (en) * | 2009-12-11 | 2012-07-10 | Tektronix, Inc. | System and method for correlating IP flows across network address translation firewalls |
| GB201101723D0 (en) * | 2011-02-01 | 2011-03-16 | Roke Manor Research | A method and apparatus for identifier correlation |
| GB201211323D0 (en) * | 2012-06-26 | 2012-08-08 | Bae Systems Plc | Resolution of address translations |
| US9800542B2 (en) * | 2013-03-14 | 2017-10-24 | International Business Machines Corporation | Identifying network flows under network address translation |
-
2023
- 2023-02-07 CN CN202310129577.1A patent/CN118473894A/zh active Pending
-
2024
- 2024-01-22 WO PCT/CN2024/073409 patent/WO2024164828A1/zh not_active Ceased
- 2024-01-22 EP EP24752685.8A patent/EP4654540A4/en active Pending
-
2025
- 2025-08-05 US US19/290,640 patent/US20250358256A1/en active Pending
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101150505A (zh) * | 2007-07-31 | 2008-03-26 | 杭州华三通信技术有限公司 | 通过网络地址转换转发数据流的方法和装置 |
| US20190114350A1 (en) * | 2017-10-17 | 2019-04-18 | Salesforce.Com, Inc. | Systems, Methods, and Apparatuses for Implementing Concurrent Dataflow Execution with Write Conflict Protection Within a Cloud Based Computing Environment |
| CN112468373A (zh) * | 2020-12-08 | 2021-03-09 | 武汉蜘易科技有限公司 | 一种指纹设备网络流量精准定位分析系统和方法 |
| CN114915566A (zh) * | 2021-01-28 | 2022-08-16 | 腾讯科技(深圳)有限公司 | 应用识别方法、装置、设备及计算机可读存储介质 |
Non-Patent Citations (2)
| Title |
|---|
| GONG, BO; WANG, RUCHUAN: "Research on Method of Session-Based P2P Network Traffic Identification", JISUANJI JISHU YU FAZHAN - COMPUTER TECHNOLOGY AND DEVELOPMENT, JISUANJI JISHU YU FAZHAN BIANJIBU - SHAANXI COMPUTER SOCIETY, CN, vol. 20, no. 3, 10 March 2010 (2010-03-10), CN , pages 5 - 8, XP009556777, ISSN: 1673-629X * |
| See also references of EP4654540A1 * |
Also Published As
| Publication number | Publication date |
|---|---|
| EP4654540A1 (en) | 2025-11-26 |
| CN118473894A (zh) | 2024-08-09 |
| US20250358256A1 (en) | 2025-11-20 |
| EP4654540A4 (en) | 2026-04-29 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2021203623A1 (zh) | 一种物联网资源接入系统及资源接入方法 | |
| US10917322B2 (en) | Network traffic tracking using encapsulation protocol | |
| CN112491717B (zh) | 一种服务路由方法及装置 | |
| EP3905590A1 (en) | System and method for obtaining network topology, and server | |
| US10033602B1 (en) | Network health management using metrics from encapsulation protocol endpoints | |
| US20130304915A1 (en) | Network system, controller, switch and traffic monitoring method | |
| US9137305B2 (en) | Information processing device, computer-readable recording medium, and control method | |
| WO2017206841A1 (zh) | 一种网络设备的服务质量检测方法和装置 | |
| JP7456502B2 (ja) | トラフィック監視装置、トラフィック監視方法およびトラフィック監視プログラム | |
| US20250358256A1 (en) | Data flow identification method and apparatus and electronic device | |
| US12621249B2 (en) | Data processing method, apparatus, network device and storage medium | |
| CN114598636A (zh) | 流量调度方法、设备及系统 | |
| WO2025066530A1 (zh) | 应用程序的性能测量方法、装置、设备、系统及存储介质 | |
| CN110708209B (zh) | 虚拟机流量采集方法、装置、电子设备及存储介质 | |
| US11888959B2 (en) | Data transmission method, system, device, and storage medium | |
| US10805206B1 (en) | Method for rerouting traffic in software defined networking network and switch thereof | |
| WO2024199223A1 (zh) | 设备接入位置的获取方法及装置 | |
| JP7085638B2 (ja) | データ伝送方法および関連装置 | |
| CN115473948A (zh) | 一种数据包分析方法、装置、计算机设备和存储介质 | |
| CN121000642B (zh) | 一种流量采集方法和装置 | |
| EP4128667B1 (en) | Distributed network flow record | |
| CN110545196A (zh) | 一种数据传输方法及相关网络设备 | |
| CN118921337A (zh) | 流量镜像系统、流量镜像方法及装置、电子设备 | |
| CN121842042A (zh) | 一种基于sdn的网络诊断方法、装置、电子设备 | |
| WO2025025720A1 (zh) | 网络传输质量的检测方法、装置、设备及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24752685 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| WWP | Wipo information: published in national office |
Ref document number: 2024752685 Country of ref document: EP |