WO2019178714A1 - 一种故障检测的方法、装置及系统 - Google Patents
一种故障检测的方法、装置及系统 Download PDFInfo
- Publication number
- WO2019178714A1 WO2019178714A1 PCT/CN2018/079422 CN2018079422W WO2019178714A1 WO 2019178714 A1 WO2019178714 A1 WO 2019178714A1 CN 2018079422 W CN2018079422 W CN 2018079422W WO 2019178714 A1 WO2019178714 A1 WO 2019178714A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- node
- delay data
- nodes
- heartbeat
- heartbeat delay
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L43/00—Arrangements for monitoring or testing data switching networks
- H04L43/10—Active monitoring, e.g. heartbeat, ping or trace-route
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
- G06F11/0703—Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation
- G06F11/0706—Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation the processing taking place on a specific hardware platform or in a specific software environment
- G06F11/0709—Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation the processing taking place on a specific hardware platform or in a specific software environment in a distributed system consisting of a plurality of standalone computer nodes, e.g. clusters, client-server systems
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/07—Responding to the occurrence of a fault, e.g. fault tolerance
- G06F11/0703—Error or fault processing not based on redundancy, i.e. by taking additional measures to deal with the error or fault not making use of redundancy in operation, in hardware, or in data representation
- G06F11/0751—Error or fault detection not based on redundancy
- G06F11/0754—Error or fault detection not based on redundancy by exceeding limits
- G06F11/0757—Error or fault detection not based on redundancy by exceeding limits by exceeding a time limit, i.e. time-out, e.g. watchdogs
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L1/00—Arrangements for detecting or preventing errors in the information received
- H04L1/20—Arrangements for detecting or preventing errors in the information received using signal quality detector
- H04L1/205—Arrangements for detecting or preventing errors in the information received using signal quality detector jitter monitoring
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L43/00—Arrangements for monitoring or testing data switching networks
- H04L43/08—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters
- H04L43/0805—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters by checking availability
- H04L43/0817—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters by checking availability by checking functioning
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L43/00—Arrangements for monitoring or testing data switching networks
- H04L43/08—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters
- H04L43/0852—Delays
- H04L43/0864—Round trip delays
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/30—Monitoring
- G06F11/3003—Monitoring arrangements specially adapted to the computing system or computing system component being monitored
- G06F11/3006—Monitoring arrangements specially adapted to the computing system or computing system component being monitored where the computing system is distributed, e.g. networked systems, clusters, multiprocessor systems
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F2201/00—Indexing scheme relating to error detection, to error correction, and to monitoring
- G06F2201/805—Real-time
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F2201/00—Indexing scheme relating to error detection, to error correction, and to monitoring
- G06F2201/81—Threshold
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L43/00—Arrangements for monitoring or testing data switching networks
- H04L43/08—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters
- H04L43/0823—Errors, e.g. transmission errors
- H04L43/0829—Packet loss
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L43/00—Arrangements for monitoring or testing data switching networks
- H04L43/08—Monitoring or testing based on specific metrics, e.g. QoS, energy consumption or environmental parameters
- H04L43/0852—Delays
- H04L43/087—Jitter
Definitions
- the present application relates to the field of communications, and in particular, to a method, device, and system for fault detection.
- each node appears in a peer-to-peer manner and generally has similar functionality. In such a cluster, when any single node fails or performance degradation occurs, the performance of the entire cluster is not greatly affected. In order to deal with node or link failures under the characteristics of avoiding single-point bottlenecks, each node joining the cluster should have network fault detection capability and a certain degree of cluster management capability.
- each node in the cluster is monitored by the failure detection mechanism of other parts of the node.
- a node in the cluster fails, its monitor will discover the failure and broadcast the failure message to the cluster.
- the management node (Leader) in the cluster receives the fault message, it will time the node. If the node still does not return to normal before the timer expires, the fault node will be determined by the management node at the end of the timer. , such as moving the failed node out of the cluster.
- the decision of the program is based only on time, and in some cases an erroneous decision will occur. If the fault of the faulty node appears as a packet loss with a certain probability, in this case the faulty node still has the opportunity to communicate with the normal node. However, based on the prior art solution, the normal node may be marked as a fault due to the problem of packet loss and the fault information is broadcasted to the cluster. If the management node of the cluster receives the fault report, the management node may move the normal node out. The situation of the cluster. Therefore, the failure detection rate of the failure detection scheme in the prior art solution is low.
- the embodiment of the present application provides a method, device, and system for fault detection, which are used to improve the accuracy of fault detection.
- the embodiment of the present application provides a method for fault detection, where the method is applied to a distributed node cluster, where the node cluster includes multiple nodes, and the method is performed by any one of the multiple nodes, and any node is the first A node, the method includes:
- the first node When the first node determines that the node health assessment trigger condition is met, the first node evaluates the health of other nodes in the node cluster according to the heartbeat delay data between the first node and other nodes in the node cluster. And get the assessment of the health of other nodes in the cluster.
- the first node acquires heartbeat delay data of multiple groups and all other nodes in the node cluster, and then evaluates the health status of other nodes according to the heartbeat delay data, thereby being able to be based on the health status of other nodes.
- the evaluation results are used for fault detection. In this way, the accuracy of fault detection is improved.
- the first node collects N sets of heartbeat delay data from all nodes in the node cluster except the first node, wherein each group of heartbeat delay data includes M heartbeat delay data, and N and M are An integer greater than 1, the node cluster containing the first node.
- the first node further calculates M evaluation values of the communication situation of the first node to the M nodes according to the N sets of heartbeat delay data. The lower the evaluation value, the better the communication situation, and the evaluation value is greater than or equal to the preset health value.
- the corresponding node is a faulty node, and the first node is not included in the M nodes, and the total number of the node clusters is M+1.
- the first node compares each of the calculated M evaluation values with a preset health value. If the preset health value is less than the preset health value, the evaluation value is a health value, and the corresponding node is healthy. If the node or the normal node is greater than or equal to the preset health value, it indicates that the evaluation value is an unhealthy value. Correspondingly, the node corresponding to the evaluation value is an unhealthy node or a faulty node.
- the first node may further evaluate whether the first node itself is healthy according to the number of the evaluation values indicating health or unhealth among the foregoing M evaluation values.
- the specific evaluation may be: if the evaluation value indicating that the dismissal represents the health value exceeds the preset proportion in the M evaluation values, for example, more than 50%, the first node is determined to be the fault node. Further, after the first node determines that it is a faulty node, the first node may perform corresponding processing according to the processing manner specified by the fault decision, for example, idle or turn itself off.
- the first node acquires heartbeat delay data of multiple groups and all other nodes in the node cluster, and then evaluates communication status of the first node to each node according to the heartbeat delay data, thereby, according to the communication situation.
- the proportion of health or not determines whether the first node is a faulty node. Therefore, each node has the self-evaluation ability of faults, without relying on other nodes in the node cluster for fault assessment, and autonomously performing idle or equal processing after judging that the self is a faulty node. In this way, the faulty node can be prevented to make fault decision and the multiple fault handling that may be caused to some extent, so that the fault detection mode is more reasonable and efficient, and the accuracy and efficiency of fault detection are improved.
- the first node calculates M evaluation values according to the N sets of heartbeat delay data, which may be:
- the first node calculates M evaluation values according to the jitter condition of the M heartbeat delay data in the N sets of heartbeat delay data, wherein the jitter condition of the M heartbeat delay data is from the first node to each of the M nodes.
- the jitter condition of a heartbeat delay data refers to the jitter situation exhibited by the N heartbeat delay data of the first node to the same node.
- the evaluation value is calculated by the jitter condition of the heartbeat delay data, so that the calculation of the evaluation value is more reasonable, and thus more accurate in the subsequent failure determination.
- the first node may further calculate M evaluation values according to the jitter condition of the M heartbeat delay data in the N sets of heartbeat delay data and the delay level of the M heartbeat delay data, where M heartbeats
- the delay level of the delay data is the delay level of the heartbeat delay data of each of the first node to the M nodes.
- the delay level situation represents the delay level of the heartbeat round trip of one of the first node to the M nodes.
- the delay level is the average delay value, and the delay level can reflect the first node to a certain extent.
- the communication status of other nodes, therefore, the evaluation value can be calculated more reasonably in combination with the delay level and the jitter condition.
- the first node may also delete the invalid data in the N sets of heartbeat data. Due to the collection of invalid data, such as noise data, due to network conditions during heartbeat delay data collection, the difference between these invalid data and the actual heartbeat delay data is generally large, so it can be filtered or denoised first. The way to delete the invalid data.
- the first node may further normalize the M evaluation values, so that the first node is in different delay network environments or different window size heartbeat delay data. A uniform comparison of evaluation values can be performed.
- the method further includes:
- the first node determines that the first node is a fault node, and if the number of the M evaluation values greater than the preset health value does not exceed the preset ratio The first node determines that the first node is a normal node.
- the first node If the number of unhealthy evaluation values is large, it means that the communication between the first node and most other nodes is poor, then the first node considers itself a faulty node. If the number of unhealthy evaluation values is small, it means that the communication between the first node and most other nodes is better, then the first node considers itself to be a normal node.
- the first node may also idle or close the first node.
- the idling means that the first node does not perform any processing on the fault condition, and still performs the current normal processing operation, waiting for the management node to remove it; the closing means that the first node closes all the ongoing processes in the first node. .
- the method further includes:
- the first node determines the management node in the node cluster according to the M evaluation values.
- the first node Since the faulty node cannot become the management node, if the first node is a faulty node, it is not necessary to determine whether the first node is a management node. If the first node is a normal node, the first node also needs to determine whether the first node is a management node.
- the way to determine the management node can be:
- the first node determines, according to the M evaluation values, that the node whose evaluation value is less than or equal to the preset health value in the node cluster is a normal node, and then obtains the determined sequence number of the normal node.
- the first node compares the sequence number of the first node with the obtained sequence number. If the sequence number of the first node is smaller than any one of the obtained sequence numbers, the first node is determined to be the management node.
- the node with the smallest sequence number among all the normal nodes in the node cluster is determined as the management node.
- the manner of determining the management node may also be other manners, for example, determining that the node with the highest sequence number among all the normal nodes in the node cluster is the management node. This application does not limit this.
- each node can determine whether it is a management node in the node cluster through self-evaluation, and avoids multiple nodes as management nodes, thereby generating fault handling conflicts that may be caused by multiple fault decisions.
- the node self-evaluates as a faulty node it does not participate in the confirmation of the management node, so that the faulty node does not become a management node, thereby improving the stability of the node cluster.
- the method further includes:
- the first node removes or isolates K nodes of the M nodes whose evaluation values are greater than the preset health value according to the M evaluation values, where K is an integer greater than or equal to 0.
- the first node needs to perform the management function of the management node, that is, the nodes in the node cluster that are faulty nodes need to be removed or isolated. In this way, the impact of the faulty node on the cluster of nodes is reduced.
- the node health triggering condition is that the first node detects an abnormal node existing in the node cluster, or the first node receives a message that the abnormal node exists broadcasted by another node; Before calculating the M evaluation values of the N-group heartbeat delay data, the method further includes:
- the first node determines whether the abnormal node returns to normal within a preset duration. If the abnormal node does not return to normal within the preset duration, the first node starts to calculate M evaluation values according to the collected N sets of heartbeat delay data. If the abnormal node returns to normal within the preset time, the first node does not need to calculate the evaluation value.
- each node has one or more other nodes to monitor it, that is, to detect whether the node is abnormal in real time.
- the first node monitors the second node in real time, and the first node sends a measurement signal to the second node. If the second node does not reply to receive the confirmation information or does not reply for a long time, it indicates that the second node may be abnormal. Thus, the first node can start collecting heartbeat delay data.
- the abnormality detection mode is firstly used to detect whether there is a node that may be abnormal in the node cluster. If the detected abnormal node returns to normal in a short time, no subsequent evaluation and failure determination are performed; If the abnormal node that arrives does not return to normal within a short period of time, it means that a faulty node appears in the node cluster, and then subsequent evaluation and determination of the faulty node are performed. In this way, the utilization of system resources is improved.
- an embodiment of the present application provides a device for detecting a fault, the device having the function of implementing the implementation described in any one of the foregoing first aspects.
- This function can be implemented in hardware or in hardware by executing the corresponding software.
- the hardware or software includes one or more modules corresponding to the functions described above.
- an embodiment of the present application provides a device for detecting a fault, the device comprising: a processor, a memory, a processor, a memory connected by a bus, a memory storing a computer instruction, and the processor is configured to implement the computer instruction Any of the implementations described on the one hand.
- an embodiment of the present application provides a readable storage medium storing computer instructions for implementing any of the implementations described in the first aspect.
- the embodiment of the present application provides a computer program product, where the computer program product includes computer software instructions, which can be loaded by a processor to implement any implementation manner as described in the first aspect. Process.
- an embodiment of the present application provides a chip device, where the chip system includes a processor and a memory, the processor is coupled to the memory, and the processor can execute a memory storage instruction to enable the chip device to perform the first aspect. Any of the implementations described.
- the embodiment of the present application provides a fault detection system, where the system is a distributed cluster system, where the node cluster system includes multiple nodes, wherein each node implements any implementation manner of the first aspect. One node.
- FIG. 1 is a schematic diagram of an embodiment of a system architecture applied to a method for detecting faults in an embodiment of the present application
- FIG. 2 is a schematic diagram of an embodiment of node module division in an embodiment of the present application.
- FIG. 3 is a schematic diagram of an embodiment of unit division of a node in an embodiment of the present application.
- FIG. 4 is a schematic diagram of an embodiment of a method for detecting faults in an embodiment of the present application
- FIG. 5 is a schematic flowchart of triggering heartbeat collection and evaluation according to an embodiment of the present application.
- FIG. 6 is a schematic diagram of a framework of an evaluation model in an embodiment of the present application.
- FIG. 7 is a schematic diagram of an application scenario applied to a method for detecting faults according to an embodiment of the present disclosure
- FIG. 8 is a schematic diagram of another application scenario applied to the method for detecting faults in the embodiment of the present application.
- FIG. 9 is a schematic diagram of another application scenario applied to the method for detecting faults in the embodiment of the present application.
- FIG. 10 is a schematic diagram of an embodiment of an apparatus for detecting faults according to an embodiment of the present application.
- FIG. 11 is a schematic diagram of another embodiment of a fault detecting apparatus according to an embodiment of the present application.
- the embodiment of the present application provides a method, device, and system for fault detection, which are used to improve the accuracy of fault detection.
- FIG. 1 is a schematic diagram of a system architecture applied to a method for detecting faults according to an embodiment of the present application.
- Figure 1 shows a node cluster, which includes multiple nodes.
- the node cluster in Figure 1 is illustrated by four nodes, including node 1, node 2, node 3, and node 4. .
- Four nodes are connected to each other, and each node is monitored by one or more other nodes.
- Each node has an anomaly detection module, a heartbeat acquisition module, and an evaluation decision module, as shown in Figure 2.
- the fault detection method provided by the embodiment of the present application can be applied to any decentralized peer-to-peer (P2P)-based cluster environment (such as Akka). -Cluster cluster organization). In this regard, the application is not limited.
- P2P peer-to-peer
- the abnormality detecting module is configured to detect abnormal conditions of other nodes monitored by the node in real time.
- the node 1 detects that the node 2 may be abnormal through the abnormality detecting module, and the node 1 broadcasts the abnormality information, notifies the node 2, the node 3, and the node 4, and agrees to be within a preset time length, for example, within 10 seconds. 2 Failure to return to normal conditions, then all nodes make assessment decisions. If the node 2 returns to the normal situation within the preset time period, no evaluation decision is needed. For details, refer to the description of step 402 in the example in FIG. 4 below.
- the heartbeat acquisition module is configured to collect heartbeat delay data of all nodes in the node cluster, where the heartbeat delay data is a heartbeat round trip delay from one node to another. For example, when node 1 detects that node 2 may be abnormal, node 1 broadcasts the abnormal information, and all nodes in the node cluster start collecting heartbeat delay data through the heartbeat collection module.
- the heartbeat delay data collected by the heartbeat module is the heartbeat delay data collected by the node transmitting the heartbeat measurement information to the other nodes in the node cluster after the node receives the broadcast of the abnormal information.
- the heartbeat delay data collected by the implementation manner can more accurately reflect the network communication status of the entire node cluster.
- the node collects a certain amount of heartbeat delay data by sending heartbeat measurement information to other nodes in the node cluster, and the node may also collect a certain amount of the latest history.
- Heartbeat delay data historical heartbeat delay data refers to the heartbeat delay data collected by the node by sending heartbeat measurement information before performing the current heartbeat measurement information.
- Each node needs to collect N sets of heartbeat delay data. The N is an integer greater than 1. For example, each node collects 10 sets of heartbeat delay data, and the heartbeat delay data collected by each node is used to evaluate the decision module. Perform failure analysis.
- the evaluation decision module is configured to evaluate the health of the communication link between the nodes according to the N sets of heartbeat delay data collected by the heartbeat acquisition module, and determine whether the node is faulty according to the evaluation result.
- the node 1 calculates an evaluation value of the node 1 to other nodes according to the heartbeat jitter of the collected N sets of heartbeat delay data.
- the evaluation value of node 1 to node 2 is calculated to be 2.3
- the evaluation value of node 1 to node 3 is 3.1
- the evaluation value of node 1 to node 4 is 1.4. Node 1 then determines which link's evaluation values are unhealthy based on the preset health values.
- the preset health value is 2.5
- the node 1 determines whether the node 1 is a normal node according to the statistics of the multiple links, and if the number of normal links exceeds a preset ratio, for example, more than fifty percent, the node is determined to be the node.
- a normal node determines that the node is a faulty node if the number of normal links does not exceed the preset ratio. For example, in the above example, it can be determined that node 1 is a normal node.
- the evaluation decision module is also used by the node to determine whether the node can become a management node (Leader) and has the ability to remove or isolate the failed node when the node is a management node. For example, when node 1 is a normal node, the node determines whether node 1 is a management node according to its serial number and other serial numbers that are normal nodes.
- the sequence number of the first node is the first node according to its IP address and port number. After the Greek operation, each of the calculated sequence numbers is broadcast to all nodes when joining the node cluster. Therefore, the first node stores the sequence numbers of all other nodes.
- the determining method of the management node is: determining that the node with the smallest serial number in the normal node is the management node, and the normal node herein refers to the above-mentioned node 1 determining that the node 1 is a normal node after evaluation, and communicating with the node 1
- the nodes that are healthy are all normal nodes (such as node 2 and node 4 above).
- node 1 determines that node 1 is a management node, node 1 also needs to remove or isolate the failed node 3 in the node cluster.
- the node described in this application may be a server, a terminal, or another device capable of the foregoing modules.
- the node 100 includes a processor 110, a memory 120, a network controller 130, and a network interface 131.
- the processor 110, the memory 120, the network controller 130, and the network interface 131 are respectively connected by a bus.
- the processor 110 is configured to control the network interface 131 to collect heartbeat delay data of other nodes, and calculate a plurality of evaluation values according to the heartbeat delay data, and determine whether the node 100 is faulty according to the plurality of evaluation values.
- it is determined that the node 100 is a normal node it is also necessary to determine whether the node 100 is a management node. If it is determined that the node 100 is a management node, the processor 110 further determines a faulty node to be removed or isolated according to the evaluation value.
- the memory 120 is configured to store a time for sending a message (for example, sending a heartbeat measurement message), and according to the time of returning the message received by the network controller 130, calculated by the central processing unit 110, obtained and stored to and from the message. Delay.
- the network controller 130 transmits data to the destination node via the network interface 131 according to an instruction of the central processing unit 110; correspondingly, at the destination node end, data is sent to the central processing unit 110 via the network interface 131 via the network control unit 130.
- the fault detection method embodiment is applied to the system architecture described above in Fig. 1, wherein the first node may be the node in the embodiment shown in Fig. 2 or 3 above.
- the fault detection method may include the following processing:
- the first node collects N sets of heartbeat delay data.
- Each group of heartbeat delay data in the N sets of heartbeat delay data includes M heartbeat delay data, and the M heartbeat delay data is the heartbeat delay data of the M nodes in the first node to the node cluster, and the M nodes are All other nodes except the first node in the node cluster, N and M are integers greater than 1, and the number of nodes in the node cluster is M+1.
- the N sets of heartbeat delay data collected by the first node are used for subsequent health assessment of other nodes. Before the first node collects N sets of heartbeat delay data, the first node first determines whether the node health assessment is satisfied. The condition is judged in the following way:
- each node in a cluster of nodes has an adjacent one or more nodes that monitor it.
- the node cluster includes four nodes, which are a first node, a second node, a third node, and a fourth node, and the fourth node monitors the first node, the first node monitors the second node, and the third node monitors the fourth node. node.
- the monitoring mode is that the measurement signal is sent in real time between the nodes.
- the first node sends the measurement signal to the second node in real time. If the feedback confirmation of the second node is not received or the feedback delay is too high, the first node confirms the second. The node may have an exception. At this time, the first node can collect the heartbeat delay data.
- the first node when the first node does not monitor the second node, an abnormality occurs, but the third node monitors the fourth node abnormality, and the third node is in the node cluster. After the abnormal information of the fourth node is broadcast, the first node also triggers the step of collecting the heartbeat delay data by the first node.
- all nodes in the node cluster pre-agreed and collect heartbeat delay data according to a preset fixed period. For example, every 10 minutes, all nodes perform heartbeat delay data collection, collect 10 sets of heartbeat delay data, and then end the acquisition, or 10 seconds after the acquisition.
- heartbeat delay data collection is performed on all nodes in the node cluster.
- the manner of collecting heartbeat delay data is as shown in the embodiment of FIG. 1 by means of transmitting heartbeat measurement information or collecting historical heartbeat delay data.
- the first node collects N sets of heartbeat delay data, wherein each group of heartbeat delay data in the N sets of heartbeat delay data includes M heartbeat delay data.
- the heartbeat delay data of the N group is represented as the heartbeat delay data collected by the first node at different N time points, and the M heartbeat delay data indicates that the first node to all other nodes in each group of heartbeat delay data M heartbeat delay data, the number of all other nodes is M.
- the first node collects two sets of heartbeat delay data, including the first group and the second group of heartbeat delay data, respectively: [1, 2, 1], [3, 4, 5], then the first The heartbeat delay data of the group indicates that the heartbeat round-trip delay of the first node to the second node is 1 millisecond, and the heartbeat round-trip delay of the first node to the third node is 2 milliseconds, and the heartbeat of the first node to the fourth node The round trip is 1 millisecond; the second set of heartbeat delay data indicates that the heartbeat round trip delay of the first node to the second node is 3 milliseconds, and the heartbeat round trip delay of the first node to the third node is 4 milliseconds, first The heartbeat round trip from node to node is 5 milliseconds.
- the unit of the delay of the center hop back and forth in the embodiment of the present application may be the foregoing milliseconds, or may be microseconds or seconds. In this regard, the application is not limited.
- the first node calculates M evaluation values according to the N sets of heartbeat delay data.
- the M evaluation values are used to indicate the communication between the first node and the M nodes, wherein the lower the evaluation value, the better the communication situation.
- each set of heartbeat delay data includes M heartbeat delay data
- the M heartbeat delay data refers to the heartbeat delay collected by the first node and each of the other nodes in the node cluster. data.
- the first node After collecting the N sets of heartbeat delay data, the first node indicates that the first node performs N times of heartbeat delay data collection for all other nodes in the node cluster.
- the first node has N heartbeat delay data collected at different time points for each other node. Therefore, the first node can perform N heartbeats collected by N at different time points for any other node.
- the extended data calculates an evaluation value. Since the number of other nodes in the node cluster is M, M evaluation values are obtained.
- the M evaluation values represent the communication status of the first node to the other M nodes in the node cluster calculated by the first node. For example, if the evaluation value of the first node to the second node is lower, it means that the communication between the first node and the second node is better.
- the first node described in step 401 monitors that another node has an abnormality, or the first node receives the abnormality information broadcasted by other nodes, the first node is configured according to the N sets of heartbeat delay data.
- the embodiment of the present application may further include the following steps:
- Determining whether the abnormal node returns to normal within a preset duration if not, performing the step of the first node calculating M evaluation values according to the N sets of heartbeat delay data; if yes, performing the first node according to the N sets of heartbeat delay data The step of calculating M evaluation values.
- the abnormality detection under the monitoring mechanism since the abnormality detection under the monitoring mechanism is not particularly perfect, it may be caused by a link problem or a very small packet loss situation. For example, if the actual situation of the second node is a normal node, after the first node detects the abnormality of the second node, the first node needs to continue to measure the second node, and if the second node detects that the second node returns to normal within a certain time. After that, the calculation of the evaluation value is not required. Conversely, if the second node has not returned to normal within a certain period of time, the calculation of the evaluation value can be performed (the calculation method of the evaluation value is described in the subsequent part of this step).
- the first node After the first node detects the abnormal node, or the first node receives the abnormal information broadcasted by other nodes, the first node starts collecting the heartbeat delay data. During the acquisition phase, the first node continues to detect the monitored abnormal node, or continues to receive information broadcast by other nodes. If the abnormal node does not return to normal within a certain period of time, the first node performs the calculation of the evaluation value, if After the abnormal node returns to normal for a period of time, the first node does not need to calculate the evaluation value.
- the first node may evaluate the health of the link communication between the first node and the other nodes according to the collected heartbeat delay data, thereby obtaining M evaluation values. .
- the embodiment of the present application further provides a method for calculating M evaluation values according to the N sets of heartbeat delay data, where the first node is based on the jitter of M heartbeat delay data in the N sets of heartbeat delay data.
- the first node collects three sets of heartbeat delay data, which are: [1, 2, 1], [1, 100, 2], [1, 1, 1], according to the three groups.
- the heartbeat delay data shows that the heartbeat delay data collected by the first node from the second node is 1, and there is no heartbeat jitter, thereby indicating that the communication between the first node and the second node is stable;
- the heartbeat delay data collected by the three nodes is 2, 100, and 1, respectively. If the heartbeat jitter is large, the communication link between the first node and the third node may be poor; the first node is from the fourth node. If the heartbeat delay data collected by the node is 1, 2, and 1, respectively, the heartbeat jitter is small, thereby indicating that the communication link between the first node and the fourth node is better.
- FIG. 6 is a general framework diagram of the evaluation model.
- the evaluation model is mainly composed of a noise reduction filter module, a delay level evaluation module, a cumulative jitter module, a standardized module, and a packet loss impact module, wherein the noise reduction filter module, the delay level evaluation module, the standardization module, and the packet loss impact module are Optional module.
- the noise reduction filter module is used to filter out some invalid heartbeat delay data; the delay level evaluation module is used to evaluate the round-trip delay of the heartbeat; and the cumulative jitter module is used to collect the heartbeat delay data collected at different times of the same node.
- the jitter amplitude is calculated; the normalization module is used to standardize the jitter value output by the accumulated jitter module, so that the network with different delay levels and the number of different heartbeat delay data can be adapted; Calculate the magnitude of the impact on the evaluation value when a packet loss occurs.
- the steps for evaluation are:
- the object is based on the N heartbeat delay data of the N-group heartbeat delay data in a single node dimension.
- the data is input to the evaluation model, it is input in the manner of (1, 1), (2, 100) and (1, 2), so that the three evaluation values calculated are the communication evaluations for the three nodes respectively. .
- T0 refers to multiple heartbeat delay data of one of the other nodes initially collected by the first node
- Ln represents a heartbeat delay lost during the heartbeat delay data collection of the first node by the first node.
- the number of data, Dw0 represents the window size of the heartbeat delay data T0, and the window size is the number of heartbeat delay data of a node initially collected by the first node.
- T represents a plurality of heartbeat delay data after filtering
- Dw represents a window size of the filtered heartbeat delay data T.
- S represents the number of invalid data that needs to be filtered out in the plurality of heartbeat delay data, and S is a constant set by a preset.
- the implementation of the noise reduction filtering is performed by performing the culling operation on the largest S data in the heartbeat delay data, and obtaining the filtered heartbeat delay data T and the number of heartbeat delay data.
- Dw len(T)+Ln, where “len” is the vector length operator, which is used to calculate the number of filtered heartbeat delay data T, Ln represents the number of lost heartbeat delay data, and Dw represents filtering.
- T0 is [2, 3, 10]
- the heartbeat delay data after filtering and noise reduction is calculated by the delay level evaluation module, the jitter accumulation module, and the packet loss impact module, respectively, to calculate the inherent delay level of the network, the jitter of the delay, and the impact of the heartbeat loss. .
- Delay level evaluation module In the packet loss network, the observation delay of the TCP ping-pong message at the originating application layer may be greater than the inherent delay (caused by the packet loss and TCP retransmission mechanism), if the mean value of the delay is used. As an evaluation criterion for the inherent delay of the link, there is a case where the evaluation value is too large, which adversely affects the standardization step. In order to estimate the inherent delay of the network from the delay history of the ping-pong message as accurately as possible, the scheme assumes that the inherent delay of the target network does not fluctuate drastically within the heartbeat history window, which makes the heartbeat delay history. Large fluctuations are mainly caused by the TCP retransmission mechanism.
- the input of the module is the filtered heartbeat delay data T
- the output is the estimated delay level l
- the time delay is evaluated using the quantile parameter p
- the evaluation algorithm is as follows:
- mean represents the mean of the elements in the vector
- sort represents the sort
- ASC represents the ascending order
- len(T) represents the vector length of the filtered heartbeat delay data T, that is, the number of filtered heartbeat delay data T is calculated.
- Jitter accumulation module This module is used to calculate the jitter of the heartbeat delay.
- the cumulative method is to differentiate the heartbeat history curve to obtain the heartbeat change vector, quantify the heartbeat change vector using the absolute value, and then perform the quantified heartbeat change rate.
- the integral gets the accumulated jitter value.
- the input of the module is the filtered heartbeat data T, and the output is the accumulated jitter value A. Calculate the cumulative jitter value using the following formula:
- ⁇ T represents a first-order difference to the vector T, and the result is a difference vector
- represents an absolute value of an element-by-element of the vector ⁇ T, and constitutes a new vector
- ) represents an element of the vector
- the result is the jitter cumulative value A.
- Packet loss impact module This module is used to calculate the impact of packet loss on the evaluation value.
- a specific implementation method of the packet loss impact module in this solution is as follows: the product of the packet loss impact coefficient Lf and the packet loss ratio is lost.
- the jitter accumulation module that the jitter of the heartbeat delay data does not change when the packet loss rate is constant, indicating that the network condition is relatively stable, but with the vector length of the heartbeat data T Increase, the jitter accumulation value A will also increase; on the other hand, in networks with different delays, such as a network with a delay of one hundred microseconds and a network with a delay of one hundred milliseconds, the packet loss rate is not In the case of change, as the delay increases, the uncertainty caused by factors such as queuing will also bring additional jitter increase.
- the model needs to standardize A.
- the normalized model normalizes the jitter cumulative value A using the link average delay level l, the vector length len(T) of the heartbeat delay data T, and superimposes the influence L of the lost heartbeat in the cumulative result.
- the input to this module is A, L, len(T), l, and the output is Use the following formula to normalize A and superimpose it with the impact of packet loss:
- a set of heartbeat delay data collected by the first node to the second node is (1, 2, 1, 5, 3, 260, -1, 5, 4, 4), where -1 represents a heartbeat loss.
- T0 (1,2,1,5,3,260,5,4,4)
- the heartbeat delay data obtained after the heartbeat history data is denoised by the noise reduction module is:
- the filtered heartbeat data passes through the jitter accumulation module, and the calculation process and result A are:
- the above formula is expressed as the difference between the two in the heartbeat delay data T, that is, the last heartbeat delay data in the heartbeat delay data T is subtracted from the previous heartbeat delay data, and the following results are obtained:
- the filtered heartbeat history data passes through the packet loss impact module, and the result L is:
- the first node determines whether the number of the M evaluation values greater than the preset health value exceeds a preset proportion. If yes, step 404 is performed, and if no, step 405 is performed.
- each evaluation value is healthy. For example, by comparing each evaluation value with a preset health value, if the evaluation value is less than or equal to the preset health value, it indicates that the corresponding link is normal, and if the evaluation value is greater than the preset evaluation value, the corresponding chain is indicated. Road failure.
- the first node determines whether the first node is a faulty node according to the quantity of the unhealthy evaluation value. For example, the preset ratio is 50%. The method is determined.
- the link from the first node to the most nodes is unhealthy.
- the first node may be determined to be a fault. If the number of unhealthy evaluation values does not exceed the preset ratio, the link between the first node and the majority of the nodes is normal, and the first node may be determined to be a normal node. And other nodes that can determine the evaluation value of the small preset health value are also normal nodes.
- the first node determines that the first node is a faulty node.
- the first node may be determined to be faulty. After determining that the first node is a faulty node, the first node may perform self-shutdown to reduce the impact on the node cluster, and the first node may also output fault prompt information for prompting the user to the first node fault.
- the first node may also idle or close itself.
- the idling means that the first node does not perform any processing on the fault condition, and still performs the current normal processing operation, waiting for the management node to remove it; the closing means that the first node closes all the ongoing processes in the first node. .
- the first node determines that the first node is a normal node.
- the first node may be determined to be a normal node.
- the first node determines, according to the M evaluation values, a management node in the node cluster.
- the first node may determine, according to the M evaluation values, the same normal node in the node cluster. The method may be determined by comparing the M evaluation values with the preset health values. If the value is less than the preset health value, the corresponding node is a normal node. The first node then determines a unique management node in the cluster of nodes based on the determined normal node.
- the method for determining the management node may be:
- the first node acquires the sequence number of the normal node among the M nodes.
- the serial number of the node is the serial number obtained by each node according to its inherent IP address or port number, and the serial number of each node is different from each other in the node cluster.
- the new node or the management node broadcasts the sequence number of the new node to all the nodes in the node cluster. Therefore, the first node stores the sequence numbers of all the nodes in the node cluster. The first node searches for the sequence number of the normal node at this time from the set of sequence numbers of all stored nodes.
- the first node compares the sequence number of the first node with the obtained sequence number.
- sequence number of the first node is smaller than any one of the obtained sequence numbers, it is determined that the first node is a management node, and vice versa, it is determined that the first node is not a management node.
- the manner in which the first node determines the management node according to the sequence number may also be multiple, and may be determined according to the agreement of all nodes in the node cluster.
- the node with the largest sequence number may be used as the management node, or a node may be randomly determined.
- the application is not limited.
- the first node removes or isolates K nodes of the M nodes according to the M evaluation values, where the K nodes are nodes whose evaluation value is greater than a preset health value.
- the first node After determining that the first node is a management node, in order to reduce the impact of the faulty node on the node cluster in the node cluster, the first node needs to remove or isolate the K nodes that are faulty in the node node cluster.
- the method of removing is that the first node broadcasts the information of the K faulty nodes to all the nodes in the node cluster, and after receiving the receiving confirmation information of all the nodes, disconnects the communication connection of the K faulty nodes.
- the method of isolating is that the first node broadcasts the information of the K faulty nodes to all the nodes in the node cluster, and notifies all the nodes to pull the K nodes into the blacklist, so that all the nodes temporarily do not communicate with the K nodes. .
- each node collects heartbeat delay data of all the nodes in the node cluster, and calculates corresponding evaluation values according to the heartbeat delay data, and determines the node according to the proportion of the healthy quantity in the evaluation value. Whether the fault is caused, and then the management node in the node cluster is determined again, thus improving the accuracy of fault detection.
- each node has the self-evaluation capability of the fault, without relying on other nodes in the node cluster for fault assessment, and independently idling or waiting after determining that the self is a faulty node. deal with.
- the faulty node can be prevented to make fault decision and the multiple fault handling that may be caused to some extent, so that the fault detection mode is more reasonable and efficient, and the accuracy and efficiency of fault detection are improved.
- each node can determine whether it is a management node in the node cluster through self-evaluation, and avoid multiple nodes as management nodes, thereby generating fault handling conflicts that may be caused by multiple fault decisions. happening.
- the node self-evaluates as a faulty node, it does not participate in the confirmation of the management node, so that the faulty node does not become a management node, thereby improving the stability of the node cluster.
- the node can overcome other problems of the prior art by relying on the heartbeat delay data as the fault detection basis.
- the disadvantage of the existing TCP-based seq and ack serial numbers for determining the health of the target node is that the application logic of the nodes in the cluster is generally located in the user space of the operating system, and the application logic layer cannot be directly read in the space.
- the content of the transport layer in the internet protocol (IP) protocol stack is located in the kernel state of the system. If you take a further mechanism to read this content, it will increase the complexity of the system, as well as the dependence on the operating system, and increase the maintenance cost of the node.
- IP internet protocol
- a node network health assessment model based on the round-trip delay history data of the heartbeat as an input is proposed on the basis of the TCP protocol, so that the assessment based on the heartbeat round-trip delay of the transport layer can be implemented.
- the model does not need to rely on the underlying system, which reduces the complexity of the system and reduces the cost of node maintenance.
- the application scenario shown in FIG. 7 is a triggering and decision process of the present solution in the case where a node node composed of three nodes has a faulty node of a non-management node.
- the node 1 is a leader node, and the node 3 is a fault node.
- Node1 loses more heartbeat messages from Node3, Node1 considers Node3 to be faulty and informs Node2 of the message.
- Node3 also lost more messages from Node1 and thought that Node1 is faulty, but Node3 still informs Node2 of this message that is not correct from a global perspective.
- These messages trigger their respective heartbeat delay acquisition modules and evaluation decision plans in each node, so that these nodes determine whether they are fault nodes in the decision period (ie, calculate the evaluation value described in the above embodiment and determine the fault value according to the evaluation value). Evaluation and decision making will take place after the end. If the decision is made directly without evaluation, Node2 will consider itself a Leader node and remove Node1 and Node3 from the cluster.
- each node After the decision cycle, each node enters its own evaluation decision process. Each node performs a self-test based on the evaluation results. Node1 considers itself to be a faulty node because both evaluation results indicate that the health is poor. Node1 and Node2 find that the evaluation result of at least one other node is good, and the number of good nodes reaches half of the estimated value. They think they are in a normal working state. Since the sequence number of Node1 is the smallest among Node1 and Node2, Node1 considers itself to be the leader node of the cluster, and the sequence number of Node2 is larger than the sequence number of Node1, and considers that it is not the leader of the cluster. Node1 will move Node3 out of the cluster at the decision time, and Node3 will idle or close itself according to the processing method specified by the fault decision.
- FIG. 8 is a triggering and decision-making process of the present solution when a management node failure and a non-management node failure condition exist in a node cluster composed of five nodes.
- each node After the decision cycle, each node enters its own decision process. Among them, Node1 and Node4 consider themselves to be faulty nodes because they indicate that their status to all other nodes is poor. Node2, Node3, and Node5 are better because the evaluation result of the existing node is good, and the number of such nodes that are evaluated as good is greater than half of the number of evaluation values, so they all consider themselves to be normal nodes, and because of the sequence number of Node2 in these normal nodes. The smallest, so Node2 is determined to be the new leader. Node2 will remove Node1 and Node4 from the cluster, and Node1 and Node4 will also idle or close themselves according to the processing method specified by the fault decision.
- FIG. 9 is a triggering and decision-making process of the present application when a switching device fails in a node cluster composed of five nodes.
- the failure of the core switch in which the failure occurs is represented by a certain probability of packet loss.
- the packet loss caused by the failure of the core switch will result in the loss of the round-trip delay jitter or heartbeat message of the heartbeat message between the left cluster and the right cluster between the five nodes. This phenomenon will eventually be discovered by the fault detection mechanism of each node and trigger their respective evaluation decision plans.
- each node After the decision cycle, each node enters its own evaluation decision process. Node4 and Node5, because their self-test results indicate that more than half of the cluster nodes are unhealthy, they consider themselves to be faulty nodes (essentially in an unhealthy cluster); Node1, Node2, and Node3 have more than half of the nodes evaluated. Ok, so these nodes think they are normal nodes. And because the number of Node1 in these nodes is the smallest, Node1 is the leader of the cluster. Node1 removes Node4 and Node5 from the cluster according to the evaluation result; Node4 and Node5 also shut themselves down or idle according to the rules defined by the cluster failure.
- each node when each node performs the evaluation value calculation, it is required to collect the heartbeat delay data of all other nodes in the node cluster, thereby calculating the evaluation of the communication link of the node to each node. Value, then determine if you are faulty, and then determine whether you are a management node.
- the embodiment of the present application further provides another implementation manner, which is as follows:
- the node first self-tests through coarse-grained heartbeat delay data (ie, as little heartbeat delay data as possible), and first tries to exclude the possibility that it is a cluster manager. Then, determine your own fault or when you are a normal node but your own serial number is greater than the serial number of other normal nodes. If this may not be ruled out, then a more granular heartbeat delay data (ie, collecting heartbeat delay data for all nodes of the entire network) is self-checked to confirm that it is the leader for decision making.
- coarse-grained heartbeat delay data ie, as little heartbeat delay data as possible
- the first node first collects heartbeat delay data of the other 20 nodes, and calculates 20 evaluation values. If more than half of the evaluation values are unhealthy values, it can be determined that the first node is The faulty node determines that the first node does not become a management node, so the first node does not need to collect heartbeat delay data.
- the first node If more than half of the evaluation values are health values, it indicates that the first node is a normal node, and the first node compares the sequence numbers of the normal nodes with the 20 nodes, and if the sequence number of the first node is not the minimum sequence number, Then it is determined that the first node will not become a management node, so that it is no longer necessary to collect heartbeat delay data with other nodes. If the first node determines that the first node is a normal node by using 20 evaluation values and the sequence number of the first node is the smallest among the sequence numbers of the normal nodes of the 20 nodes, the first node may be a management node.
- the first node needs to collect the heartbeat delay data of the other 79 nodes, and calculate 79 evaluation values, and determine the normal nodes among the 79 nodes, so that the serial number of the first node and the serial number of the normal node in the 79 nodes are For comparison, if the sequence number of the first node is still the smallest, it is determined that the first node is a management node, and if the sequence number of the first node is not the smallest, it is determined that the first node is not a management node.
- the network In a node cluster with a large number of nodes, the network is generally stable, and there are few cases where large-area node faults occur. Generally, a very small number of fault nodes occur. Therefore, through the implementation manner, the self-failure detection of the node can generally be accurately performed, and the overhead of the network is reduced.
- FIG. 10 is a schematic diagram of an embodiment of an apparatus for detecting faults according to an embodiment of the present application.
- the apparatus 600 is applied to a distributed node cluster, where the node cluster includes multiple nodes, and the method is described by Any one of the plurality of nodes is executed, the node is the first node, the device is the first node, and the device 600 includes: a determining unit 601 and an evaluating unit 602;
- the determining unit 601 is configured to determine whether the node health assessment trigger condition is met, and the evaluating unit 602 is configured to use, according to the first node, the other node in the node cluster when the node health assessment trigger condition is met.
- the heartbeat delay data between the nodes respectively evaluates the health of other nodes in the node cluster, and obtains the evaluation results of the health of other nodes in the cluster.
- the determining unit 601 is configured to perform the implementation of the three node health assessment triggers described in step 401 in the embodiment of FIG. 4;
- the evaluation unit 602 is configured to perform steps 401 to 403 in the embodiment of FIG. 4.
- the evaluating unit 602 includes:
- the collecting unit 6021 is configured to execute the content described in step 401 in the embodiment of FIG. 4;
- the calculating unit 6022 is configured to perform step 402 in the embodiment of FIG. 4.
- the device 600 further includes:
- the deleting unit 605 is configured to delete the invalid data in the N sets of heartbeat delay data before the evaluation unit 602 calculates the M evaluation values according to the N sets of heartbeat delay data; and delete the N sets of heartbeats after the invalid data is deleted.
- the delay data is used to calculate the M evaluation values.
- the device 600 further includes:
- the determining unit 603 is configured to perform step 404 in the embodiment of FIG. 4.
- the determining unit 603 is further configured to perform step 405 in the embodiment of FIG. 4.
- the determining unit 603 is further configured to: perform step 406 in the embodiment of FIG. 4 .
- the device 600 further includes:
- the processing unit 604 is configured to perform step 407 in the embodiment of FIG. 4.
- the execution of the first node described in the embodiment of FIG. 4 is performed at the time of execution of the embodiment of the present invention. For details, refer to the embodiment of FIG. 4, and details are not described herein.
- the apparatus of the embodiment of Figure 6 has yet another form of embodiment.
- the apparatus 700 includes: a processor 701, a memory 702, a transceiver 703, the processor 701, the memory 702, and
- the transceiver 703 is coupled by a bus 704, which may include a transmitter and a receiver, the memory 702 storing computer instructions.
- the transceiver 703 is configured to perform heartbeat delay data collection on other nodes
- the memory 702 is configured to store heartbeat delay data collected by the transceiver 703
- the processor 701 is configured to call the heartbeat delay data in the memory 702 for evaluation. The health of other nodes and determine if the node is a management node.
- the processor 701 is configured to determine, at least, whether a node health assessment trigger condition is met;
- the processor 701 is further configured to: when the node health assessment trigger condition is met, according to the heartbeat delay data between the first node and other nodes in the node cluster, respectively, in the node cluster The other node health is evaluated and the results of the assessment of the health of other nodes in the cluster are obtained.
- the transceiver 703 is configured to perform step 401 in the embodiment of FIG. 4;
- the memory 702 is configured to store the collected heartbeat delay data
- the processor 701 is configured to perform steps 402 to 407 in the embodiment of FIG.
- the chip comprises: a processing unit and a communication unit
- the processing unit may be, for example, a processor
- the communication unit may be, for example, an input/output interface, Pin or circuit, etc.
- the processing unit may execute a computer-executed instruction stored by the storage unit to cause a chip within the device to perform a method of resource scheduling in any of the above embodiments.
- the storage unit is a storage unit in the chip, such as a register, a cache, etc., and the storage unit may also be a storage unit located outside the chip in the terminal, such as a read-only memory (read) -only memory, ROM) or other types of static storage devices, random access memory (RAM), etc. that can store static information and instructions.
- the processor mentioned in any of the above may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more for controlling the above.
- CPU central processing unit
- ASIC application-specific integrated circuit
- the integrated circuit of the program execution of the first aspect wireless communication method may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more for controlling the above.
- CPU central processing unit
- ASIC application-specific integrated circuit
- the computer program product includes one or more computer instructions.
- the computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable device.
- the computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be from a website site, computer, server or data center Transfer to another website site, computer, server, or data center by wire (eg, coaxial cable, fiber optic, digital subscriber line (DSL), or wireless (eg, infrared, wireless, microwave, etc.).
- wire eg, coaxial cable, fiber optic, digital subscriber line (DSL), or wireless (eg, infrared, wireless, microwave, etc.).
- the computer readable storage medium can be any available media that can be stored by a computer or a data storage device such as a server, data center, or the like that includes one or more available media.
- the usable medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
- the disclosed system, apparatus, and method may be implemented in other manners.
- the device embodiments described above are merely illustrative.
- the division of the unit is only a logical function division.
- there may be another division manner for example, multiple units or components may be combined or Can be integrated into another system, or some features can be ignored or not executed.
- the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interface, device or unit, and may be in an electrical, mechanical or other form.
- the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of the embodiment.
- each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
- the above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
- the integrated unit if implemented in the form of a software functional unit and sold or used as a standalone product, may be stored in a computer readable storage medium.
- a computer readable storage medium A number of instructions are included to cause a computer device (which may be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present application.
- the foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, and the like. .
Landscapes
- Engineering & Computer Science (AREA)
- Signal Processing (AREA)
- Computer Networks & Wireless Communication (AREA)
- Theoretical Computer Science (AREA)
- Environmental & Geological Engineering (AREA)
- Quality & Reliability (AREA)
- General Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Cardiology (AREA)
- General Health & Medical Sciences (AREA)
- General Physics & Mathematics (AREA)
- Physics & Mathematics (AREA)
- Computer Hardware Design (AREA)
- Debugging And Monitoring (AREA)
Abstract
一种故障检测的方法,所述方法应用于分布式的节点集群,所述节点集群包括多个节点,所述方法由所述多个节点中的任一节点执行,所述任一节点为第一节点,所述方法包括:所述第一节点判断是否满足健康度评估触发条件,当满足所述健康度评估触发条件时,所述第一节点根据所述第一节点与所述节点集群中的其它节点之间的心跳时延数据分别对所述节点集群中的其它节点健康度进行评估,并获得所述集群中的其它节点的健康度的评估结果。
Description
本申请涉及通信领域,尤其涉及一种故障检测的方法、装置及系统。
分布式系统由于具备高并发性、自治性、容错性、可靠性等先进特性,已被广泛的应用于工业界的各类控制系统中。在去中心化的集群中,每个节点以对等的方式出现且一般具备相似的功能。在这类集群中,任何一个单一节点出现故障或出现性能下降时,整个集群的性能不会受到较大的影响。为了在避免单点瓶颈的特质下应对节点或链路故障的出现,每个加入集群的节点都应该具备网络故障的检测能力和一定限度的集群管理能力。
在现有技术中,集群中每个节点都受其他一部分节点的故障检测机制的监视。当集群中某节点故障时,其监视者会发现这种故障,并向集群广播该故障消息。当集群中的管理节点(Leader)收到该条故障消息时将会对该节点进行计时,若计时结束前该节点仍然未恢复正常,则该故障节点会在计时结束时被该管理节点执行决策,如将该故障节点移出集群。
该方案的判决仅以时间作为依据,在某些情况下还会发生错误判决。如故障节点的故障表现为以一定概率丢包的情况,在该情况下故障节点仍然有机会与正常节点通信。但基于现有技术方案,可能因丢包的问题将正常的节点标记为故障并将该故障信息广播至集群中,若集群的管理节点收到了该故障报告,则会出现管理节点将正常节点移出集群的情况。因此,现有技术方案中的故障检测方案的故障检测的准确率较低。
发明内容
本申请实施例提供了一种故障检测的方法、装置及系统,用于提高故障检测的准确性。
第一方面,本申请实施例提供一种故障检测的方法,该方法应用于分布式的节点集群,节点集群包括多个节点,方法由多个节点中的任一节点执行,任一节点为第一节点,该方法包括:
第一节点在确定满足节点健康度评估触发条件的情况下,该第一节点根据第一节点与节点集群中的其它节点之间的心跳时延数据分别对节点集群中的其它节点健康度进行评估,并获得集群中的其它节点的健康度的评估结果。
在该实现方式中,第一节点获取了多组与节点集群中其它所有的节点的心跳时延数据,再根据这些心跳时延数据评估其它节点的健康状况,从而可以根据其它节点的健康状况的评估结果进行故障检测。这样,提高了故障检测的准确性。
一种可能的实现方式中,第一节点向节点集群中除第一节点意外的所有节点采集N组心跳时延数据其中,每组心跳时延数据包含M个心跳时延数据,N和M为大于1的整数,节点集群包含该第一节点。第一节点再根据N组心跳时延数据计算出第一节点至M个节点的通信情况的M个评估值,评估值越低,则通信情况越好,评估值大于或等于预设健康值所对应的节点为故障节点,M个节点中不包括该第一节点,节点集群的总数为M+1。该第一节点再将所计算得到的M个评估值中每个评估值与预设的健康值进行比较,若小于该预设健康值,则表示该评估值为健康值,对应的节点为健康节点或正常节点,若大于或者等于 该预设健康值,则表示该评估值为不健康值,相应的,该评估值多对应的节点为不健康节点或为故障节点。
进一步的,第一节点还可以根据前述M个评估值中表示健康或者不健康的评估值的数量作为依据评估该第一节点自身是否健康。具体评估可以为:若为M个评估值中评估解雇表示健康值的评估值超过预设占比,比如超过百分之五十,则确定第一节点为故障节点。进一步的,当该第一节点确定自身为故障节点后,该第一节点可以根据故障决策规定的处理方式进行相应的处理,例如使自己空转或关闭。
在该实现方式中,第一节点获取了多组与节点集群中其它所有的节点的心跳时延数据,再根据这些心跳时延数据评估第一节点至每个节点的通信情况,从而根据通信情况的健康与否的占比确定第一节点是否为故障节点。从而使得每个节点都具备故障的自我评估能力,而无需依赖于节点集群中的其它节点进行故障评估,而且在判断自我为故障节点后自主进行空转或等处理。这样,可以一定程度上防止故障节点进行故障决策以及可能引起的多重故障处理等情况,从而使得故障检测的方式更为合理和高效,提高了故障检测的准确性和效率。
另一种可能的实现方式中,第一节点根据N组心跳时延数据计算M个评估值,可以为:
第一节点根据N组心跳时延数据中M个心跳时延数据的抖动情况计算M个评估值,其中,M个心跳时延数据的抖动情况为第一节点至所述M个节点中每个节点的心跳时延数据的抖动情况,心跳时延数据的抖动幅度越大,评估值越大。
一个心跳时延数据的抖动情况指的是第一节点至同一个节点的N个心跳时延数据所展示的抖动情况。抖动幅度越大,则表示通信情况越差,则评估值越大,若抖动越小,则表示通信情况越好,则评估值越小。该实现方式中,通过心跳时延数据的抖动情况计算评估值,使得评估值的计算更为合理,从而在后续的故障确定中更为准确。
可选的,第一节点还可以根据N组心跳时延数据中M个心跳时延数据的抖动情况、以及M个心跳时延数据的时延水平情况计算M个评估值,其中,M个心跳时延数据的时延水平情况为第一节点至M个节点中每个节点的心跳时延数据的时延水平情况。
时延水平情况表示第一节点至M个节点中某一节点的心跳往返的时延水平,比如,该时延水平为平均时延值,通过时延水平能够在一定程度上反应第一节点与其它节点的通信情况,因此,结合时延水平以及抖动情况可以更合理计算该评估值。
可选的,第一节点在根据N组心跳数据计算M个评估值之前,第一节点还可以先将N组心跳数据中的无效数据删除。由于在心跳时延数据采集时,可能会由于网络情况采集到一些无效数据,比如噪声数据,这些无效数据与实际的心跳时延数据区别一般是很大的,因此可以先通过滤波或降噪的方式删除掉其中的无效数据。
可选的,第一节点在计算得到M个评估值后,还可以对该M个评估值进行标准化处理,使得第一节点在不同时延网络环境下或者不同窗口大小的心跳时延数据的情况下均能进行统一的评估值的比较。
另一种可能的实现方式中,该方法还包括:
若M个评估值中大于预设健康值的数量超过预设占比,第一节点确定第一节点为故障 节点,以及若M个评估值中大于预设健康值的数量未超过预设占比,第一节点则确定第一节点为正常节点。
若不健康的评估值的数量较多,则表示第一节点与多数的其它节点的通信情况都较差,那么第一节点则认为自己为故障节点。若不健康的评估值的数量较少,则表示第一节点与多数的其它节点的通信情况较好,那么第一节点则认为自己为正常节点。
可选的,当第一节点确定第一节点为故障节点之后,第一节点还可以将第一节点空转或者关闭。
其中,空转指的是第一节点不对该故障情况做任何处理,仍然进行当前正常的处理作业,等待管理节点将其移除;关闭指的是第一节点将第一节点中正在进行进程全部关闭。
可选的,在第一节点确定第一节点为正常节点之后,该方法还包括:
第一节点根据M个评估值确定节点集群中的管理节点。
由于故障节点不能成为管理节点,因此,若第一节点为故障节点,则无需确定第一节点是否为管理节点。若第一节点为正常节点,则第一节点还需确定第一节点是否为管理节点。
确定管理节点的方式可以为:
第一节点根据M个评估值确定节点集群中评估值小于或等于所述预设健康值所对应的节点为正常节点,再获取所确定的正常节点的序号。第一节点将第一节点的序号与所获取的序号进行比较,若第一节点的序号小于所获取的序号的任意一个序号,则确定第一节点为管理节点。
在该实现方式中,是通过确定节点集群中所有正常节点中序号最小的节点为管理节点。可选的,确定管理节点的方式还可以为其它方式,比如,确定节点集群中所有正常节点中序号最大的节点为管理节点。本申请对此不做限定。该实现方式中,每个节点均能通过自我评估从而确定自己是否为节点集群中的管理节点,避免了多个节点为管理节点,从而产生多重故障决策所可能引起的故障处理冲突等情况。当节点自我评估为故障节点后,则不会参与管理节点的确认,从而故障节点不会成为管理节点,从而也提高了节点集群的稳定性。
另一种可能的实现方式中,在确定第一节点为管理节点之后,该方法还包括:
第一节点根据所述M个评估值移除或隔离所述M个节点中评估值大于所述预设健康值的节点的K个节点,K为大于或等于0的整数。
在该实现方式中,当第一节点确定为管理节点之后,第一节点则需要执行管理节点的管理功能,即将节点集群中为故障节点的节点均需要移除或者隔离。这样,减少了故障节点对节点集群的影响。
另一种可能的实现方式中,节点健康度触发条件为所述第一节点检测节点集群中存在的异常节点,或者第一节点接收到其它节点广播的存在异常节点的消息;在第一节点根据N组心跳时延数据计算M个评估值之前,该方法还包括:
第一节点判断异常节点在预设时长内是否恢复正常,若在预设时长内异常节点未恢复正常,则第一节点开始根据所采集到的N组心跳时延数据计算M个评估值。若在预设的时 长内异常节点恢复正常,则第一节点无需进行评估值的计算。
在节点集群中,每个节点都有一个或多个其它节点对其进行监视,即实时检测该节点是否异常。比如,第一节点实时监视着第二节点,第一节点向第二节点发送测量信号,若第二节点未回复接收确认信息或者长时间未回复,则表示第二节点可能异常。从而第一节点则可以开始进行心跳时延数据的采集。
在实现方式中,先通过异常检测方式检测节点集群中是否存在可能异常的节点,若检测到的异常节点在较短的时间内恢复为正常,则无需进行后续的评估和故障的确定;若检测到的异常节点在较短的时间内未恢复正常,则表示在节点集群中出现了故障节点,则进行后续的评估以及故障节点的确定等步骤。这样,提高了系统资源的利用率。
第二方面,本申请实施例提供一种故障检测的装置,该装置具有实现上述第一方面中任意一种所描述实现方式中功能。该功能可以通过硬件实现,也可以通过硬件执行相应的软件实现。该硬件或软件包括一个或多个与上述功能相对应的模块。
第三方面,本申请实施例提供一种故障检测的装置,该装置包括:处理器、存储器,处理器、存储器通过总线连接,存储器存储有计算机指令,处理器通过执行计算机指令用于实现如第一方面所描述的任意一种实现方式。
第四方面,本申请实施例提供一种可读存储介质,该存储介质存储有用于实现如第一方面所描述的任意一种实现方式的计算机指令。
第五方面,本申请实施例提供一种计算机程序产品,该计算机程序产品包括计算机软件指令,该计算机软件指令可通过处理器进行加载来实现如第一方面所描述的任意一种实现方式中的流程。
第六方面,本申请实施例提供一种芯片装置,该芯片系统包括处理器和存储器,处理器连接到存储器,该处理器可以运行存储器存储的指令,以使该芯片装置执行上述第一方面所描述的任意一种实现方式。
第七方面,本申请实施例提供一种故障检测的系统,该系统为分布式的集群系统,该节点集群系统中包括多个节点,其中,每个节点如实现第一方面任意实现方式的第一节点。
图1为本申请实施例中故障检测的方法所应用的系统架构的一个实施例的示意图;
图2为本申请实施例中节点模块划分的一个实施例示意图;
图3为本申请实施例中节点的单元划分的一个实施例示意图;
图4为本申请实施例中故障检测的方法的一个实施例示意图;
图5为本申请实施例中触发心跳采集和评估的一个流程示意图;
图6为本申请实施例中评估模型的框架示意图;
图7为本申请实施例中故障检测的方法所应用的一个应用场景示意图;
图8为本申请实施例中故障检测的方法所应用的另一应用场景示意图;
图9为本申请实施例中故障检测的方法所应用的另一应用场景示意图;
图10为本申请实施例中故障检测的装置的一个实施例示意图;
图11为本申请实施例中故障检测装置的另一实施例示意图。
本申请实施例提供了一种故障检测的方法、装置及系统,用于提高故障检测的准确率。
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本申请一部分实施例,而不是全部的实施例。本申请的说明书和权利要求书及上述附图中的术语“第一”、“第二”、“第三”、“第四”等(如果存在)是用于区别类似的对象,而不必用于描述特定的顺序或先后次序。应该理解这样使用的数据在适当情况下可以互换,以便这里描述的实施例能够以除了在这里图示或描述的内容以外的顺序实施。此外,术语“包括”和“具有”以及他们的任何变形,意图在于覆盖不排他的包含,例如,包含了一系列步骤或单元的过程、方法、系统、产品或设备不必限于清楚地列出的那些步骤或单元,而是可包括没有清楚地列出的或对于这些过程、方法、产品或设备固有的其它步骤或单元。
参照图1所示,图1为本申请实施例中故障检测的方法所应用的系统架构示意图。图1所展示的是一个节点集群,该节点集群中包括多个节点,为了方便描述,图1中的节点集群以四个节点为例进行说明,包括节点1、节点2、节点3以及节点4。四个节点相互连接,每个节点都被一个或多个其它节点监视。每个节点都具备异常检测模块、心跳采集模块以及评估决策模块,如图2所示。需要说明的是,本申请实施例所提供的故障检测方法可以适用于任何去中心化的、无专用故障检测节点的基于对等网络(peer to peer,P2P)方式组织的集群环境中(如Akka-Cluster的集群组织方式)。对此,本申请不做限定。
在本申请中,异常检测模块,用于实时检测节点所监视的其它节点的异常情况。比如,节点1通过异常检测模块检测到节点2可能异常,节点1则会将该异常信息进行广播,通知节点2、节点3以及节点4,并约定在预设时长内,比如10秒内若节点2未能恢复到正常情况,则所有节点进行评估决策。若在预设时长内节点2恢复到正常情况,则无需进行评估决策,具体可参考下面图4实例中步骤402的描述内容。
心跳采集模块,用于对节点集群中所有的节点进行心跳时延数据的采集,该心跳时延数据为一个节点至另一个节点的心跳往返时延。比如,当节点1检测到节点2可能异常时,节点1将该异常信息进行广播,节点集群中所有的节点则开始通过心跳采集模块对心跳时延数据进行采集。在一种实现方式中,心跳模块所采集的心跳时延数据为节点接收到异常信息的广播后,向节点集群中的其它节点发送心跳测量信息所采集到的心跳时延数据。通过该实现方式所采集的心跳时延数据能够较为准确地体现出整个节点集群的网络通信状况。另一种实现方式中,由于通过发送心跳测量信息采集心跳时延数据会耗费较大的采集时间,且在较短时间内采集到的心跳时延数据的数量不会太多。而采集到的心跳时延数据要用作后续的评估决策,越多数量的心跳时延数据评估决策的准确性越高。因此,在该实现方式中,节点在接收到异常信息的广播后,既通过向节点集群中的其它节点发送心跳测量信息采集一定数量的心跳时延数据,节点还可以采集一定数量的最新的历史心跳时延数据,历史心跳时延数据指的是节点在进行当前进行发送心跳测量信息之前,通过发送心跳测量信息所采集到的心跳时延数据。每个节点均需要采集N组心跳时延数据,该N为大于1的整数,比如,每个节点采集10组心跳时延数据,每个节点所采集到的心跳时延数据用 于评估决策模块进行故障分析。
评估决策模块,用于根据心跳采集模块所采集到的N组心跳时延数据对节点之间的通信链路健康情况进行评估,并根据评估结果确定节点是否故障。比如,节点1根据所采集到的N组心跳时延数据的心跳抖动情况计算出节点1至其它节点的评估值。如,计算出节点1至节点2的评估值为2.3,节点1至节点3的评估值为3.1,节点1至节点4的评估值为1.4。节点1再根据预设的健康值确定哪些链路的评估值不健康。比如,预设的健康值为2.5,则可以确定节点1至节点2的链路正常,节点1至节点3的链路不正常,节点1至节点4的链路正常。节点1再通过评估决策模块根据所统计的多条链路的情况确定节点1是否为正常节点,若正常链路的数量超过预设占比,比如超过百分之五十,则确定该节点为正常节点,若正常链路的数量未超过预设占比,则确定该节点为故障节点。比如,上述举例中,可以确定节点1为正常节点。评估决策模块还用于节点确定该节点是否能够成为管理节点(Leader),并具备当该节点为管理节点时移除或隔离故障节点的能力。比如,当节点1为正常节点时,该节点则根据其序号与其它为正常节点的序号确定节点1是否为管理节点,其中第一节点的序号为第一节点根据其IP地址以及端口号通过哈希运算后得到的,每个在加入该节点集群时均会将其所计算得到的序号所广播给所有节点,因此,第一节点中存储有所有的其它节点的序号。比如,该管理节点的确定方法为:确定为正常节点中序号最小的节点为管理节点,此处的正常节点指的是上述节点1通过评估后确定节点1为正常节点,则与节点1通信链路为健康的节点均为正常节点(比如上述的节点2和节点4)。当节点1确定节点1为管理节点后,节点1还需将节点集群中故障的节点3进行移除或者隔离。
需要说明的是,本申请中所描述的节点可以为服务器,也可以为终端,或者其它具备上述模块能力的设备,本申请对此不做限定。
本申请实施例中所描述的节点还可以以另一种形式体现,如图3所示,在该实施例中,该节点100包括处理器110、存储器120、网络控制器130以及网络接口131,处理器110、存储器120、网络控制器130以及网络接口131分别通过总线连接。其中,处理器110,用于控制网络接口131采集其它节点的心跳时延数据,并根据心跳时延数据计算多个评估值,并根据多个评估值决策节点100是否故障。当确定节点100为正常节点时,还需确定该节点100是否为管理节点。若确定该节点100为管理节点,处理器110还需根据评估值确定待移除或者隔离的故障节点。
存储器120,用于存储发送消息(例如,发送心跳测量消息)的时间,并根据网络控制器130收到的该消息的返回消息的时间,经由中央处理器110计算,得到并存储到该消息往返时延。
网络控制器130:依据中央处理器110的指令向目的节点经由网络接口131发送数据;相应的,在目的节点端,数据经由网络接口131,通过网络控制130器发送至中央处理器110。
参照图4所示,对本申请实施例中故障检测的方法进行示例性描述。该故障检测方法实施例应用在上述图1所述的系统架构中,其中的第一节点可以是上述图2或3所示实施 例中的节点。该故障检测方法可以包括如下处理:
401、第一节点采集N组心跳时延数据。
其中,N组心跳时延数据中每组心跳时延数据包含M个心跳时延数据,M个心跳时延数据为第一节点至节点集群中M个节点的心跳时延数据,M个节点为节点集群中除第一节点以外的所有其它节点,N和M为大于1的整数,节点集群中的节点数量为M+1。
第一节点所采集的N组心跳时延数据是用于后续对其它节点进行健康度评估,在第一节点采集N组心跳时延数据之前,第一节点先判断是否满足节点健康度评估的触发条件,其判断方式为如下实现方式:
在一种实现方式中,节点集群中每个节点都有相邻的一个或多个节点对其监视。比如,节点集群中包括四个节点,分别为第一节点、第二节点、第三节点以及第四节点,第四节点监视第一节点,第一节点监视第二节点,第三节点监视第四节点。监视的方式为节点之间实时发送测量信号,比如,第一节点实时向第二节点发送测量信号,若未接收到第二节点的反馈确认或者反馈时延过高,则第一节点确认第二节点可能存在异常。此时,第一节点则可以进行心跳时延数据的采集。
在另一种实现方式中,基于上一种实现方式的监视机制,当第一节点未监视到第二节点出现异常,但第三节点监视到第四节点异常,且第三节点在节点集群中广播了第四节点异常的信息,第一节点接收到该异常信息后,也会触发第一节点进行心跳时延数据采集的步骤。
在另一种实现方式中,节点集群中所有的节点预先约定,按照预设的固定周期进行心跳时延数据的采集。比如,每隔10分钟,所有节点则进行心跳时延数据采集,采集10组心跳时延数据后结束采集,或者采集10秒后结束采集。
第一节点进行心跳时延数据采集时,需对节点集群中所有的节点均进行心跳时延数据采集。心跳时延数据采集的方式如图1实施例中所描述的通过发送心跳测量信息的方式或采集历史心跳时延数据的方式。第一节点采集N组心跳时延数据,其中,N组心跳时延数据中每组心跳时延数据包含M个心跳时延数据。N组心跳时延数据表示为第一节点在不同的N个时间点所采集到的心跳时延数据,M个心跳时延数据表示每组心跳时延数据中都有第一节点至其它所有节点的M个心跳时延数据,其它所有节点的数量为M。比如,第一节点所采集了两组心跳时延数据,包括第一组和第二组心跳时延数据,分别为:[1,2,1]、[3,4,5],那么第一组心跳时延数据中则表示了第一节点至第二节点的心跳往返时延为1毫秒,第一节点至第三节点的心跳往返时延为2毫秒,第一节点至第四节点的心跳往返为1毫秒;第二组心跳时延数据中则表示了第一节点至第二节点的心跳往返时延为3毫秒,第一节点至第三节点的心跳往返时延为4毫秒,第一节点至第四节点的心跳往返为5毫秒。基于前述举例,那么N=2,分别为前述的第一组和第二组;M=3,表示其它节点的数量为3个,因此每组所采集到的心跳时延数据的数量也为3个。需要说明的是,本申请实施例中心跳往返的时延的单位可以为前述的毫秒,也可以为微秒或者秒。对此,本申请不做限定。
402、第一节点根据N组心跳时延数据计算M个评估值。
具体的实施例中,M个评估值用于指示第一节点与M个节点的通信情况,其中,评估 值越低,通信情况越好。
在步骤401中已经描述了每组心跳时延数据都包含M个心跳时延数据,且该M个心跳时延数据指代第一节点分别与节点集群中其它每个节点所采集的心跳时延数据。第一节点在采集到N组心跳时延数据后,则表示第一节点分别对节点集群中所有的其它节点均进行了N次的心跳时延数据的采集。第一节点针对每个其它节点,都有N个在不同时间点所采集到的心跳时延数据,因此,第一节点可以对任何一个其它节点根据N个在不同时间点所采集到的心跳时延数据计算一个评估值,由于节点集群中其它节点的数量为M个,从而得到M个评估值。该M个评估值表示第一节点所计算的第一节点至节点集群中的其它M个节点的通信情况。比如,第一节点至第二节点的评估值较低,则表示第一节点与第二节点的通信情况较好。
可选的,若为步骤401所描述的第一节点监视到其它节点出现异常,或者第一节点接收到其它节点广播的异常信息的实现方式,则,在第一节点根据N组心跳时延数据计算M个评估值之前,本申请实施例还可以包括如下步骤:
判断异常节点在预设时长内是否恢复正常,若否,则执行第一节点根据N组心跳时延数据计算M个评估值的步骤;若是,则可不执行第一节点根据N组心跳时延数据计算M个评估值的步骤。
结合图5所示,由于该监视机制下的异常检测并非特别完善,有可能是链路问题或者极少的丢包情况导致的。比如,第二节点的实际情况为正常节点,那么第一节点在检测到第二节点异常后,第一节点还需继续对第二节点进行测量,若在一定时间内检测到第二节点恢复正常后,则无需进行评估值的计算,反之,若在一定时间内第二节点仍未恢复正常,则可以进行评估值的计算(评估值的计算方式参考本步骤后续部分的描述)。需要说明的是,第一节点在检测到异常节点,或者第一节点接收到其它节点广播的异常信息后,第一节点则开始进行心跳时延数据的采集。在采集阶段,第一节点继续对所监视的异常节点继续检测,或者继续接收其它节点广播的信息,若在一段时间内,异常节点未恢复正常,则第一节点进行评估值的计算,若在一段时间内,异常节点恢复正常,则第一节点无需进行评估值的计算。
第一节点在采集到N组心跳时延数据后,则可以根据所采集到的心跳时延数据对第一节点至其它节点之间的链路通信的健康情况进行评估,从而得到M个评估值。
可选的,本申请实施例还提供了一种根据N组心跳时延数据计算M个评估值的方法,为:第一节点根据N组心跳时延数据中M个心跳时延数据的抖动情况计算M个评估值,其中,所述M个心跳时延数据的抖动情况为第一节点至M个节点中每个节点的心跳时延数据的抖动情况,心跳时延数据的抖动幅度越大,评估值越大。下面进行举例说明:
比如,N等于3,第一节点所采集到3组心跳时延数据,分别为:[1,2,1]、[1,100,2]、[1,1,1],根据这三组心跳时延数据可知,第一节点从第二节点所采集到的心跳时延数据均为1,无心跳抖动情况,从而表示第一节点至第二节点之间的通信平稳;第一节点从第三节点所采集到的心跳时延数据分别为2、100、1,该心跳抖动幅度较大,则表示第一节点至第三节点之间的通信链路可能较差;第一节点从第四节点所采集到的心跳时延数 据分别1、2、1,则该心跳抖动较小,从而表示第一节点至第四节点之间的通信链路较好。
由于可靠性的要求,节点集群内节点间的通信常采用传输控制协议(transmission control protocol,TCP)来保证通信的可靠性。本申请提出了一套在TCP协议的基础上,依靠心跳的往返时延历史数据作为输入的节点网络健康度评估模型。图6为该评估模型的总体框架图。该评估模型主要由降噪滤波模块、时延水平评估模块、累计抖动模块、标准化模块、丢包影响模块构成,其中,降噪滤波模块、时延水平评估模块、标准化模块、丢包影响模块为可选模块。降噪滤波模块用于过滤掉一些无效的心跳时延数据;时延水平评估模块用于对心跳的往返时延的评估;累计抖动模块用于对同一节点不同时刻所采集到的心跳时延数据的抖动幅度大小进行计算;标准化模块用于对累计抖动模块所输出的抖动值进行标准化处理,使得不同时延水平的网络和不同的心跳时延数据的数量都能适配;丢包影响模块用于计算当出现丢包时对评估值的影响的大小。评估的步骤为:
1)对采集到的N组心跳时延数据通过降噪滤波模块进行一次降噪处理。
需要说明的是,第一节点在根据N组心跳时延数据进行通过评估模型进行评估计算时,所依据的对象是N组心跳时延数据中以单节点为维度的N个心跳时延数据进行逐一评估。比如,N=2,M=3,第一组和第二组心跳时延数据分别为[1,2,1]、[1,100,2],第一节点在将这两组心跳时延数据输入评估模型时,是以(1,1)、(2,100)以及(1,2)的方式输入的,这样所计算得到的3个评估值则是分别针对3个节点的通信评估情况。
为了避免用户线程调度等因素导致心跳消息的处理被小概率拖延导致评估结果受到不良影响,本方案考虑对心跳时延数据进行滤波后再进行进一步评估。该模块的输入为采集的心跳时延数据T0,Ln,Dw0,输出为滤波后的心跳时延数据T,Ln,Dw,使用滤波强度参数:S。其中,T0指代第一节点初步采集的其它节点中某一节点的多个心跳时延数据,Ln表示在第一节点在对某一节点进行心跳时延数据采集过程中所丢失的心跳时延数据的数量,Dw0表示心跳时延数据T0的窗口大小,该窗口大小为第一节点初步采集的某个节点的心跳时延数据的数量。T表示进行滤波后的多个心跳时延数据,Dw表示滤波后的心跳时延数据T的窗口大小。S表示该多个心跳时延数据中需要滤波掉的无效数据的数量,S为预设设置的常量。
在一种具体的实施方式中,一种降噪滤波的实现方式为,对心跳时延数据内最大的S个数据进行剔除操作,得到滤波后的心跳时延数据T及心跳时延数据的数量Dw=len(T)+Ln,其中“len”为取向量长度运算符,即用于计算该滤波后的心跳时延数据T的数量,Ln表示丢失的心跳时延数据的数量,Dw表示滤波后心跳时延数据T的窗口大小。比如,S等于1,第一节点初步采集的某一节点的一组心跳时延数据为[-1,2,3,10],“-1”表示丢失一个心跳时延数据,则Ln=1,则T0为[2,3,10],Dw0=len(2,3,10)+1,即3+1=4个。根据该滤波方式,则需要剔除掉其中的最大值10,剔除后得到的T为[2,3],Ln为1,根据窗口大小Dw的计算公式得到Dw=len(2,3)+1,即2+1=3个。
2)对滤波降噪后的心跳时延数据通过时延水平评估模块、抖动累积模块、丢包影响模块,分别计算出网络的固有时延水平、时延的抖动情况、心跳丢失带来的影响。
时延水平评估模块:在丢包网络中,TCP ping-pong消息在发端应用层的观测时延有 可能大于固有时延(丢包与TCP的重传机制引发),若使用该时延的均值作为链路的固有时延的评估标准会存在评估值偏大的情况,进而对标准化步骤产生不良影响。为了尽可能准确的从ping-pong消息的时延历史中评估网络的固有时延,本方案假设目标网络的固有时延不会在心跳历史窗口内发生剧烈波动,这就使得心跳时延历史的大幅度波动主要由TCP重传机制引起。而由于概率性丢包的假设,我们认为在窗口内仍然有数量占比为p的包未受重传机制的影响,也就是没有丢包的完成了往返,使用此部分心跳包的心跳往返时延可以反映链路的固有时延。所以我们从心跳时延数据中选取最小的P*100%的部分,以该部分的均值作为衡量固有时延的标准。
该模块的输入为滤波后的心跳时延数据T,输出为评估的时延水平l,使用分位参量p进行时延评估,分位参量p为经验值,其含义是在该时延水平评估时取多少的心跳时延数据作为链路时延的评估依据。比如,p=0.5,取百分比值为50%,则表示把一半的高时延的心跳时延数据剔除,留下一半低时延的心跳时延数据。评估算法为如下公式:
l=mean(left(sort(T,ASC),p*len(T)))
其中,mean代表向量内元素的均值,sort代表排序,ASC代表升序,len(T)代表对滤波后的心跳时延数据T取向量长度,即计算滤波后的心跳时延数据T的数量。
极端情况下,可以直接从历史中选取最小的历时时延作为输出结果,即
l=min(T)
抖动累积模块:该模块用于计算心跳时延的抖动情况,累积的方法为将心跳历史曲线进行微分得到心跳变化矢量,使用绝对值将心跳变化矢量数量化,再对数量化的心跳变化率进行积分得到累计抖动值。
模块的输入为滤波后的心跳数据T,输出为累计抖动值A。计算累计抖动值采用如下公式:
A=∑(|ΔT|)
其中ΔT表示对向量T进行一阶差分,结果为差分向量;|ΔT|表示对向量ΔT的逐个元素取绝对值,构成的新向量;∑(|ΔT|)表示对向量|ΔT|的元素求和,结果为抖动累积值A。
丢包影响模块:该模块用于计算丢包对评估值的影响,本方案中丢包影响模块的一种具体的实施方式为:将丢包影响系数Lf与丢包所占比例的乘积作为丢包影响模块的输出L并传递给标准化模块,其中Lf为经验值,比如Lf=5。
3)将上述三个模块的输出会作为标准化模块的输入,通过标准化模块使用这些计算结果计算出标准化的心跳时延抖动情况作为评估模型的输出。
从抖动累积模块的实现中可以看出,在丢包率不变的情况下,心跳时延数据的抖动情况也不会发生变化,表示网络情况比较稳定,但是随着心跳数据T的向量长度的增大,抖 动累积值A也会随之增大;另一方面,在不同时延的网络中,例如百微秒级别时延的网络和百毫秒级别时延的网络中,在丢包率不变的情况下,随着时延增大,由于排队等因素引起的不确定性也会带来额外的抖动增加。
为了使抖动累积值A对不同时延水平的网络和不同的心跳时延数据T的向量长度的尺寸都能适配,以便决策者进行判决,模型需要对A进行标准化。
标准化模型使用链路平均时延水平l、心跳时延数据T的向量长度len(T)对抖动累计值A进行标准化,并将丢失心跳带来的影响L叠加在累计结果中。
为了更清楚的描述上述各个模块的功能与使用,下面进行举例说明:
假设第一节点对第二节点所采集的的一组心跳时延数据为(1,2,1,5,3,260,-1,5,4,4),其中-1代表心跳丢失。去掉丢失后的心跳后,则
T0=(1,2,1,5,3,260,5,4,4)
Ln=1
Dw0=10
设降噪模块的强度S=1,则心跳历史数据经过降噪模块后得到降噪后的心跳时延数据为:
T=(1,2,1,5,3,5,4,4)
Ln=1
Dw=9
设时延水平评估模块的分位参数p=0.5,则滤波后的心跳历史数据经过时延水平模块,其计算过程及结果l为:
sort(T,ASC)=(1,1,2,3,4,4,5,5)
p*len(T)=4
left(sort(T,ASC),p*len(T0))=(1,1,2,3)
l=mean(left(sort(T,ASC),p*len(T)))=2.75
滤波后的心跳数据经过抖动累积模块,其计算过程及结果A为:
ΔT={Ti-Ti-1(0<i<len(T))}
上述公式表示为计算心跳时延数据T中两两之间的差值,即将心跳时延数据T中后一个心跳时延数据减去前一个心跳时延数据,得到如下结果:
ΔT=(2-1,1-2,5-1,3-5,…,4-4)=(1,-1,4,-2,2,-1,0)
|ΔT|=(1,1,4,2,2,1,0)
A=∑(|ΔT|)=11
设丢包影响模块的丢包影响系数Lf=5,则滤波后的心跳历史数据经过丢包影响模块, 其结果L为:
L=5*1/9=0.56
403、第一节点判断M个评估值中大于预设健康值的数量是否超过预设占比,若是,则执行步骤404,若否,则执行步骤405。
根据上述步骤402所描述的内容,第一节点根据N组心跳时延数据计算得到M个评估值后,则判断每个评估值是否健康。比如通过将每个评估值与预设的健康值进行比较,若评估值小于或等于预设健康值,则表示所对应的链路正常,若评估值大于预设评估值,则表示对应的链路故障。第一节点在计算得到每个评估值的健康情况后,再根据其中不健康评估值的数量确定第一节点是否为故障节点。比如,该预设占比为百分之五十,该确定的方式为,若不健康评估值的数量超过预设占比,则表示该第一节点至大多数的节点的链路都不健康,因此,则可以确定该第一节点故障;若不健康评估值的数量未超过预设占比,则表示第一节点至大部分的节点之间的链路为正常,则可以确定第一节点为正常节点,并且可以确定评估值小预设健康值的其它节点也是正常节点。
参照图6所示,在计算得到M个评估值
后,则将每个
与预设健康值thres进行比较,若小于该预设健康值thres,则表示该
为健康值,否则,则表示该
为不健康值。比如,根据步骤402中的举例中所计算的
为1.1,若预设健康值thres=2.5,1.1<2.5,则表示该
为健康值。
404、第一节点确定所述第一节点为故障节点。
若不健康评估值的数量超过预设占比,则表示该第一节点至大多数的节点的链路都不健康,因此,则可以确定该第一节点故障。在确定第一节点为故障节点后,第一节点可以进行自我关闭,以减少对节点集群的影响,第一节点还可以输出故障提示信息,用于提示用户第一节点故障。
可选的,第一节点在确认自己为故障节点之后,第一节点还可以将自己进行空转或关闭。其中,空转指的是第一节点不对该故障情况做任何处理,仍然进行当前正常的处理作业,等待管理节点将其移除;关闭指的是第一节点将第一节点中正在进行进程全部关闭。
405、第一节点确定所述第一节点为正常节点。
若不健康评估值的数量未超过预设占比,则表示第一节点至大部分的节点之间的链路为正常,则可以确定第一节点为正常节点。
406、第一节点根据M个评估值确定节点集群中的管理节点。
第一节点为在确定第一节点为正常节点后,第一节点则可以根据该M个评估值确定节点集群中的同样为正常的节点。确定的方式可以为将M个评估值与预设的健康值进行比较,若小于预设健康值,则表示所对应的节点为正常节点。从而第一节点再根据所确定的正常节点确定节点集群中的唯一管理节点。
可选的,确定管理节点的方式可以为:
第一节点获取M个节点中正常节点的序号。
节点的序号是每个节点根据其固有的IP地址或者端口号通过哈希计算所得到的序号,每个节点的序号在节点集群中互不相同。当有新的节点加入该节点集群中,该新的节点或者管理节点向节点集群中所有节点广播该新的节点的序号,因此,第一节点中存储有节点集群中所有节点的序号。第一节点则从所存储的所有节点的序号的集合中查找出此时为正常节点的序号。
第一节点将第一节点的序号与所获取的序号进行比较。
若第一节点的序号小于所获取的序号的任意一个序号,则确定第一节点为管理节点,反之,则确定第一节点不是管理节点。
可选的,第一节点根据序号确定管理节点的方式也可以有多种,具体可以根据节点集群中所有节点的约定而定,比如,可以根据序号最大的节点作为管理节点,或者随机确定一个节点为管理节点,并广播给其它节点等等。对此,本申请不做限定。
407、若确定第一节点为管理节点,则第一节点根据M个评估值移除或隔离所述M个节点中的K个节点,K个节点为评估值大于预设健康值的节点。
在确定第一节点为管理节点后,为了减少节点集群中故障节点对节点集群的影响,第一节点则需要将节点节点集群中故障的K个节点进行移除或者隔离。其中,移除的方式为第一节点将该K个故障节点的信息广播至节点集群中的所有节点,再收到所有节点的接收确认信息后,则断开该K个故障节点的通信连接。隔离的方式为第一节点将该K个故障节点的信息广播至节点集群中的所有节点,并通知所有节点将该K个节点拉入黑名单,使得所有节点暂时不与该K个节点进行通信。
本申请实施例中,每个节点通过对节点集群中所有的节点进行心跳时延数据的采集,并根据心跳时延数据计算出相应的评估值,根据评估值中为健康数量的占比确定节点是否故障,从而再确定节点集群中的管理节点,这样,提高了故障检测的准确率。
此外,通过本申请实施例所示的方案,每个节点都具备故障的自我评估能力,而无需依赖于节点集群中的其它节点进行故障评估,而且在判断自我为故障节点后自主进行空转或等处理。这样,可以一定程度上防止故障节点进行故障决策以及可能引起的多重故障处理等情况,从而使得故障检测的方式更为合理和高效,提高了故障检测的准确性和效率。
另外,本申请实施例中,每个节点均能通过自我评估从而确定自己是否为节点集群中的管理节点,避免了多个节点为管理节点,从而产生多重故障决策所可能引起的故障处理冲突等情况。当节点自我评估为故障节点后,则不会参与管理节点的确认,从而故障节点不会成为管理节点,从而也提高了节点集群的稳定性。
另外,本申请实施例中节点依靠心跳时延数据作为故障检测依据还可以克服现有技术的其他问题。现有的基于TCP协议的seq、ack序号判断目标节点的健康情况的方案的缺点是:集群中的节点的应用逻辑一般位于操作系统的用户空间中,该空间下应用逻层无法直接读取到位于系统底层内核态下网络互连协议(internet protocol,IP)协议栈中传输层的相关内容。如果采取进更进一步的机制来读取这些内容,则会增加系统的复杂性,以及对操作系统的依赖性,增加对节点的维护成本。本申请方案中,提出了在TCP协议的基础 上,依靠心跳的往返时延历史数据作为输入的节点网络健康度评估模型,这样,基于传输层的心跳往返时延的采集,则能实现该评估模型,无需依赖于底层系统,减少了系统的复杂度,从而减少了节点维护的成本。
为了更清楚的阐述本申请的方案,下面结合图7所示应用场景,对本申请实施例中故障检测的方法的进行举例描述。图7所示的应用场景为由三个节点组成的节点集群内存在一个非管理(Leader)节点出现故障节点的情况下,本申请方案的触发及决策流程。
如图7所示,其中节点(Node)1为Leader节点,Node3为故障节点。在该实例中,因Node1丢失了较多来自Node3的心跳消息,故Node1认为Node3故障,并将该消息告知Node2。Node3同样因为丢失了较多来自Node1的消息,认为Node1故障,但Node3仍然将这条从全局角度看不正确的消息告知了Node2。这些消息触在各个节点内都触发了各自的心跳时延采集模块与评估决策计划,使得这些节点在决策周期(即上述实施例所描述的计算评估值并根据评估值确定自己是否为故障节点)结束后将进行评估与决策。若不进行评估而直接进行决策,则Node2将会认为自己是Leader节点,并将Node1、Node3从集群中移除。
经过决策周期,各节点分别进入各自的评估决策流程中。各节点依据评估结果进行自检。其中Node3由于两个评估结果均表示健康度为差,则其认为自己属于故障节点;Node1、Node2则发现存在至少一个其他节点的评估结果为良好,良好的节点数量达到评估值数量的一半,则他们认为自己处于正常工作的状态。又由于Node1的序号在Node1和Node2中最小,故Node1认为自己是集群的Leader节点,Node2的序号大于Node1的序号,认为自己不是集群的Leader。则Node1会在决策时刻时将Node3移出集群,Node3也会根据故障决策所规定的处理方式空转或将自己关闭。
参照图8所示的应用场景,图8为以五节点所组成的节点集群内存在一个管理节点故障和一个非管理节点故障情况时,本申请方案的触发及决策流程。
如图8所示,其中Node1为故障前的Leader节点,现在发生了故障,Node4为另一非Leader故障节点。其中图8中还展示了决策前各节点故障检测机制发现的节点故障情况。
经过决策周期,各节点分别进入各自的决策流程中。其中Node1、Node4因评估结果表明其到所有其他节点的状态都为差,则认为自己为故障节点。Node2、Node3、Node5则因为存在节点的评估结果为好,且这种评估为好的节点的数量大于评估值数量的半数,所以他们都认为自己是正常节点,又因这些正常节点中Node2的序号最小,故Node2被确定为新的Leader。Node2会将Node1与Node4从集群中移除,Node1与Node4也会根据故障决策所规定的处理方式空转或将自己关闭。
参照图9所示的应用场景,图9为以五节点所组成的节点集群内存在一个交换设备故障的情况时本申请方案的触发及决策流程。
如图9所示,假设其中故障的核心交换机的故障表现为以一定概率丢包。则因核心交换机故障导致的丢包将会导致五个节点间左集群与右集群的心跳消息往返时延的抖动或心跳消息的丢失。这种现象将最终被各节点的故障检测机制发现并触发各自的评估决策计划。
经过决策周期,各节点分别进入自己的评估决策流程。其中Node4、Node5因其自检结 果表明超过半数的集群节点为不健康状态,则他们认为自己为故障节点(实质为处于不健康集群中);Node1、Node2、Node3则因超过半数的节点的评估状态为好,所以这些节点认为自己为正常节点。又因为这些节点中Node1的序号最小,故Node1为集群的Leader。Node1根据评估结果将Node4、Node5从集群中移除;Node4、Node5也会根据集群故障所定义的规则将自己关闭或空转。
上述实施例所描述的实现方式中,为每个节点在进行评估值计算时,需要采集节点集群中所有的其它节点的心跳时延数据,从而计算该节点至每个节点的通信链路的评估值,再确定自己是否故障,进而判断自己是否为管理节点。可选的,本申请实施例还提供了另一种实现方式,如下描述:
在针对节点集群内节点较少的情况,认为单位时间内全网范围内心跳测量消息的发送是可以容忍的。然而,对于节点集群内节点数量较多的网络环境下,心跳测量消息可能带来较大的网络开销。因此,在本实现方式中,可以设计更进一步的方案:节点先通过粗粒度的心跳时延数据(即,尽可能少的心跳时延数据)自检,首先尝试排除自己是集群管理者的可能,则确定自己故障或者当自己为正常节点但自己的序号大于其它正常节点的序号。如果这种可能不能排除,则进行更细粒度的心跳时延数据(即,采集全网所有节点的心跳时延数据)进行自检,用来确认自己是领导者以进行决策。
比如,节点集群中有100个节点,第一节点首先采集其它20个节点的心跳时延数据,计算出20个评估值,若超过半数的评估值为不健康值,则可以确定该第一节点为故障节点,从而确定了第一节点不会成为管理节点,因此该第一节点无需再做心跳时延数据的采集。若超过半数的评估值为健康值,则表示该第一节点为正常节点,该第一节点再比较自己与这20个节点中正常节点的序号,若确定该第一节点的序号不是最小序号,那么确定该第一节点也不会成为管理节点,从而无需再采集与其它节点的心跳时延数据。若该第一节点通过20个评估值确定为第一节点为正常节点且第一节点的序号与20个节点中正常节点的序号中比较是最小的,那么第一节点则有可能为管理节点。从而第一节点还需采集其它79个节点的心跳时延数据,并计算79个评估值,确定79个节点中的正常节点,从而在将第一节点的序号与79个节点中正常节点的序号进行比较,若第一节点的序号仍是最小的,则确定第一节点为管理节点,若第一节点的序号不是最小的,则确定第一节点不是管理节点。
由于在节点数量较多的节点集群中,一般情况下网络是较稳定的,出现大面积节点故障的情况较少,一般情况是出现极少数的故障节点。因此,通过本实现方式,一般也能准确地进行节点的自我故障检测,并且减少了网络的开销。
参照图10所示,图10为本申请实施例中故障检测的装置的一个实施例示意图,该装置600应用于分布式的节点集群,所述节点集群包括多个节点,所述方法由所述多个节点中的任一节点执行,所述任一节点为第一节点,所述装置为所述第一节点,该装置600包括:判断单元601以及评估单元602;
其中,判断单元601用于判断是否满足节点健康度评估触发条件,评估单元602用于用于当满足所述节点健康度评估触发条件时,根据所述第一节点与所述节点集群中的其它节点之间的心跳时延数据分别对所述节点集群中的其它节点健康度进行评估,并获得所述 集群中的其它节点的健康度的评估结果。
具体的,判断单元601,用于执行图4实施例中步骤401中所描述的三种节点健康度评估触发的实现方式;
评估单元602,用于执行图4实施例中的步骤401至403。
可选的,所述评估单元602包括:
采集单元6021,用于执行图4实施例中的步骤401中所描述的内容;
计算单元6022,用于执行图4实施例中的步骤402。
可选的,所述装置600还包括:
删除单元605,用于在所述评估单元602根据所述N组心跳时延数据计算M个评估值之前,删除所述N组心跳时延数据中的无效数据;删除无效数据后的N组心跳时延数据用于计算所述M个评估值。
可选的,所述装置600还包括:
确定单元603,用于执行图4实施例中的步骤404。
可选的,确定单元603还用于执行图4实施例中的步骤405。
可选的,所述确定单元603还用于:执行图4实施例中的步骤406。
可选的,所述装置600还包括:
处理单元604,用于执行图4实施例中的步骤407。
图6实施例所描述的各个单元在运行时执行图4实施例中所描述的第一节点的执行步骤,详细内容可参照图4实施例,此处不做赘述。
图6实施例所述的装置还有另一个形式的实施例。
参照图11所示,对本申请实施例提供的一种故障检测的装置进行示例性介绍,该装置700包括:处理器701、存储器702、收发器703,所述处理器701、所述存储器702以及所述收发器703通过总线704连接,收发器703可以包括发送器与接收器,所述存储器702存储有计算机指令。其中,收发器703用于对其它节点进行心跳时延数据采集,存储器702用于存储收发器703所采集到的心跳时延数据,处理器701用于调用存储器702中的心跳时延数据进行评估其它节点的健康状况,并确定本节点是否为管理节点。
该处理器701至少用于判断是否满足节点健康度评估触发条件;
该处理器701至少还用于:当满足所述节点健康度评估触发条件时,根据所述第一节点与所述节点集群中的其它节点之间的心跳时延数据分别对所述节点集群中的其它节点健康度进行评估,并获得所述集群中的其它节点的健康度的评估结果。
具体的:
收发器703用于执行图4实施例中的步骤401;
存储器702用于存储所采集到的心跳时延数据;
处理器701用于执行图4实施例中的步骤402至407。
所属领域的技术人员可以清楚地了解到,为描述的方便和简洁,上述描述的系统,装置和单元的具体工作过程,可以参考前述方法实施例中的对应过程,在此不再赘述。
在另一种可能的设计中,当上述装置为设备内的芯片时,芯片包括:处理单元和通信 单元,所述处理单元例如可以是处理器,所述通信单元例如可以是输入/输出接口、管脚或电路等。该处理单元可执行存储单元存储的计算机执行指令,以使该设备内的芯片执行上述实施例中任意一种资源调度的方法。可选地,所述存储单元为所述芯片内的存储单元,如寄存器、缓存等,所述存储单元还可以是所述终端内的位于所述芯片外部的存储单元,如只读存储器(read-only memory,ROM)或可存储静态信息和指令的其他类型的静态存储设备,随机存取存储器(random access memory,RAM)等。
其中,上述任一处提到的处理器,可以是一个通用中央处理器(CPU),微处理器,特定应用集成电路(application-specific integrated circuit,ASIC),或一个或多个用于控制上述第一方面无线通信方法的程序执行的集成电路。
在上述实施例中,可以全部或部分地通过软件、硬件、固件或者其任意组合来实现。当使用软件实现时,可以全部或部分地以计算机程序产品的形式实现。
所述计算机程序产品包括一个或多个计算机指令。在计算机上加载和执行所述计算机程序指令时,全部或部分地产生按照本发明实施例所述的流程或功能。所述计算机可以是通用计算机、专用计算机、计算机网络、或者其他可编程装置。所述计算机指令可以存储在计算机可读存储介质中,或者从一个计算机可读存储介质向另一计算机可读存储介质传输,例如,所述计算机指令可以从一个网站站点、计算机、服务器或数据中心通过有线(例如同轴电缆、光纤、数字用户线(DSL))或无线(例如红外、无线、微波等)方式向另一个网站站点、计算机、服务器或数据中心进行传输。所述计算机可读存储介质可以是计算机能够存储的任何可用介质或者是包含一个或多个可用介质集成的服务器、数据中心等数据存储设备。所述可用介质可以是磁性介质,(例如,软盘、硬盘、磁带)、光介质(例如,DVD)、或者半导体介质(例如固态硬盘Solid State Disk(SSD))等。
在本申请所提供的几个实施例中,应该理解到,所揭露的系统,装置和方法,可以通过其它的方式实现。例如,以上所描述的装置实施例仅仅是示意性的,例如,所述单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口,装置或单元的间接耦合或通信连接,可以是电性,机械或其它的形式。
所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本申请各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能单元的形式实现。
所述集成的单元如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的全部或部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可 以是个人计算机,服务器,或者网络设备等)执行本申请各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、磁碟或者光盘等各种可以存储程序代码的介质。
以上所述,以上实施例仅用以说明本申请的技术方案,而非对其限制;尽管参照前述实施例对本申请进行了详细的说明,本领域的普通技术人员应当理解:其依然可以对前述各实施例所记载的技术方案进行修改,或者对其中部分技术特征进行等同替换;而这些修改或者替换,并不使相应技术方案的本质脱离本申请各实施例技术方案的范围。
Claims (34)
- 一种故障检测的方法,所述方法应用于分布式的节点集群,所述节点集群包括多个节点,所述方法由所述多个节点中的任一节点执行,所述任一节点为第一节点,其特征在于,所述方法包括:所述第一节点判断是否满足节点健康度评估触发条件;当满足所述节点健康度评估触发条件时,所述第一节点根据所述第一节点与所述节点集群中的其它节点之间的心跳时延数据分别对所述节点集群中的其它节点的健康度进行评估,并获得所述集群中的其它节点的健康度的评估结果。
- 根据权利要求1所述的方法,其特征在于,所述第一节点根据所述第一节点与所述节点集群中的其它节点之间的心跳时延数据分别对所述节点集群中的其它节点的健康度进行评估,并获得所述集群中的其它节点的健康度的评估结果,包括:所述第一节点采集N组心跳时延数据,所述N组心跳时延数据中每组心跳时延数据包含M个心跳时延数据,所述M个心跳时延数据为所述第一节点至节点集群中M个节点的心跳时延数据,所述M个节点为所述节点集群中的其它节点,其中,所述N和所述M为大于1的整数;所述第一节点根据所述N组心跳时延数据计算M个评估值,所述M个评估值用于指示所述第一节点与所述M个节点的通信情况,其中,评估值大于预设健康值所对应的节点为故障节点。
- 根据权利要求2所述的方法,其特征在于,所述第一节点根据所述N组心跳时延数据计算M个评估值,包括:所述第一节点根据所述N组心跳时延数据中M个心跳时延数据的抖动情况计算M个评估值,其中,所述M个心跳时延数据的抖动情况为所述第一节点至所述M个节点中每个节点的心跳时延数据的抖动情况,心跳时延数据的抖动幅度越大,评估值越大。
- 根据权利要求3所述的方法,其特征在于,所述第一节点根据所述N组心跳时延数据中M个心跳时延数据的抖动情况计算M个评估值,包括:所述第一节点根据所述N组心跳时延数据中M个心跳时延数据的抖动情况、以及M个心跳时延数据的时延水平情况计算M个评估值,其中,所述M个心跳时延数据的时延水平情况为所述第一节点至所述M个节点中每个节点的心跳时延数据的时延水平情况。
- 根据权利要求4所述的方法,其特征在于,所述第一节点根据所述N组心跳时延数据中M个心跳时延数据的抖动情况、以及M个心跳时延数据的时延水平情况计算M个评估值,包括:所述第一节点根据所述N组心跳时延数据中M个心跳时延数据的抖动情况、M个心跳时延数据的时延水平情况以及M个心跳时延数据的丢包情况计算M个评估值,其中,所述M个心跳时延数据的丢包情况为所述第一节点至所述M个节点中每个节点的心跳时延数据的丢包情况,丢包数量越多,评估值越大。
- 根据权利要求3至5其中任意一项所述的方法,其特征在于,在所述第一节点根据所述N组心跳时延数据计算M个评估值之前,所述方法还包括:所述第一节点删除所述N组心跳时延数据中的无效数据;删除无效数据后的N组心跳时延数据用于计算所述M个评估值。
- 根据权利要求2至6其中任一项所述的方法,其特征在于,在所述第一节点根据所述N组心跳时延数据计算M个评估值之后,所述方法还包括:若所述M个评估值中大于所述预设健康值的数量超过预设占比,所述第一节点确定所述第一节点为故障节点;以及若所述M个评估值中大于预设健康值的数量未超过预设占比,所述第一节点确定所述第一节点为正常节点。
- 根据权利要求7所述的方法,其特征在于,在所述第一节点确定所述第一节点为故障节点之后,所述方法还包括:所述第一节点将所述第一节点空转或关闭。
- 根据权利要求7所述的方法,其特征在于,在所述第一节点确定所述第一节点为正常节点之后,所述方法还包括:所述第一节点根据所述M个评估值确定所述节点集群中的管理节点。
- 根据权利要求9所述的方法,其特征在于,所述第一节点根据所述M个评估值确定所述节点集群中的管理节点,包括:所述第一节点根据所述M个评估值确定所述节点集群中的所有正常节点,其中,评估值小于或等于所述预设健康值,则表示对应的节点为正常节点;所述第一节点获取所确定的正常节点的序号;所述第一节点将所述第一节点的序号与所获取的序号进行比较,若所述第一节点的序号小于所述所获取的序号的任意一个序号,则确定所述第一节点为管理节点。
- 根据权利要求10所述的方法,其特征在于,在确定所述第一节点为管理节点之后,所述方法还包括:所述第一节点根据所述M个评估值移除或隔离所述M个节点中的K个节点,所述K个节点为评估值大于所述预设健康值的节点。
- 根据权利要求1至11其中任意一项所述的方法,其特征在于,所述节点健康度评估触发条件为所述第一节点检测所述节点集群中存在的异常节点,或者所述第一节点接收到其它节点广播的存在异常节点的消息。
- 根据权利要求12所述的方法,其特征在于,在所述第一节点根据所述N组心跳时延数据计算M个评估值之前,所述方法还包括:所述第一节点判断所述异常节点在预设时长内是否恢复正常;所述第一节点在所述异常节点在预设时长内未恢复正常的情况下执行所述根据所述N组心跳时延数据计算M个评估值的步骤。
- 根据权利要求1至11其中任意一项所述的方法,其特征在于,所述节点健康度评估触发条件为所述第一节点检测到当前时刻为预设周期时刻。
- 根据权利要求10或11所述的方法,其特征在于,所述序号为节点根据其网络协议IP地址以及端口号按照预设算法得到的数值。
- 一种故障检测的装置,所述装置应用于分布式的节点集群,所述节点集群包括多个节点,所述装置由所述多个节点中的任一节点,所述任一节点为第一节点,所述装置为所述第一节点,其特征在于,所述装置包括:判断单元,用于判断是否满足节点健康度评估触发条件;评估单元,用于当满足所述节点健康度评估触发条件时,根据所述第一节点与所述节点集群中的其它节点之间的心跳时延数据分别对所述节点集群中的其它节点健康度进行评估,并获得所述集群中的其它节点的健康度的评估结果。
- 根据权利要求16所述的装置,其特征在于,所述评估单元包括:采集单元,用于采集N组心跳时延数据,所述N组心跳时延数据中每组心跳时延数据包含M个心跳时延数据,所述M个心跳时延数据为所述第一节点至节点集群中M个节点的心跳时延数据,所述M个节点为所述节点集群中除第一节点以外的所有其它节点,其中,所述N和所述M为大于1的整数;计算单元,用于根据所述N组心跳时延数据计算M个评估值,所述M个评估值用于指示所述第一节点与所述M个节点的通信情况,其中,评估值大于预设健康值所对应的节点为故障节点。
- 根据权利要求17所述的装置,其特征在于,所述计算单元具体用于:根据所述N组心跳时延数据中M个心跳时延数据的抖动情况计算M个评估值,其中,所述M个心跳时延数据的抖动情况为所述第一节点至所述M个节点中每个节点的心跳时延数据的抖动情况,心跳时延数据的抖动幅度越大,评估值越大。
- 根据权利要求18所述的装置,其特征在于,所述计算单元具体用于:根据所述N组心跳时延数据中M个心跳时延数据的抖动情况、以及M个心跳时延数据的时延水平情况计算M个评估值,其中,所述M个心跳时延数据的时延水平情况为所述第一节点至所述M个节点中每个节点的心跳时延数据的时延水平情况。
- 根据权利要求19所述的装置,其特征在于,所述计算单元具体用于:根据所述N组心跳时延数据中M个心跳时延数据的抖动情况、M个心跳时延数据的时延水平情况以及M个心跳时延数据的丢包情况计算M个评估值,其中,所述M个心跳时延数据的丢包情况为所述第一节点至所述M个节点中每个节点的心跳时延数据的丢包情况,丢包数量越多,评估值越大。
- 根据权利要求18至20其中任意一项所述的装置,其特征在于,所述装置还包括:删除单元,用于在所述评估单元根据所述N组心跳时延数据计算M个评估值之前,删除所述N组心跳时延数据中的无效数据;删除无效数据后的N组心跳时延数据用于计算所述M个评估值。
- 根据权利要求17至21其中任意一项所述的装置,其特征在于,所述装置还包括:确定单元,用于在所述计算单元根据所述N组心跳时延数据计算M个评估值之后,若所述M个评估值中大于所述预设健康值的数量超过预设占比,确定所述第一节点为故障节点;以及若所述M个评估值中大于预设健康值的数量未超过预设占比,确定所述第一节点为正 常节点。
- 根据权利要求22所述的装置,其特征在于,所述装置还包括:处理单元,用于在所述确定单元确定所述第一节点为故障节点之后,将所述第一节点空转或关闭。
- 根据权利要求22所述的装置,其特征在于,所述确定单元还用于:在确定所述第一节点为正常节点之后,根据所述M个评估值确定所述节点集群中的管理节点。
- 根据权利要求24所述的装置,其特征在于,所述确定单元具体用于:据所述M个评估值确定所述节点集群中的所有正常节点,其中,评估值小于或等于所述预设健康值,则表示对应的节点为正常节点;获取所确定的正常节点的序号;将所述第一节点的序号与所获取的序号进行比较,若所述第一节点的序号小于所述所获取的序号的任意一个序号,则确定所述第一节点为管理节点。
- 根据权利要求25所述的装置,其特征在于,所述装置还包括:处理单元,用于在所述确定单元确定所述第一节点为管理节点之后,根据所述M个评估值移除或隔离所述M个节点中的K个节点,所述K个节点为评估值大于所述预设健康值的节点。
- 根据权利要求16至26其中任意一项所述的装置,其特征在于,所述节点健康度评估触发条件为所述第一节点检测所述节点集群中存在的异常节点,或者所述第一节点接收到其它节点广播的存在异常节点的消息。
- 根据权利要求27其中任意一项所述的装置,其特征在于,所述判断单元还用于:在所述计算单元根据所述N组心跳时延数据计算M个评估值之前,判断所述异常节点在预设时长内是否恢复正常;所述计算单元在所述异常节点在预设时长内未恢复正常的情况下根据所述N组心跳时延数据计算M个评估值。
- 根据权利要求16至26其中任意一项所述的装置,其特征在于,所述节点健康度评估触发条件为所述第一节点检测到当前时刻为预设周期时刻。
- 根据权利要求25或26所述的装置,其特征在于,所述序号为节点根据其网络协议IP地址以及端口号按照预设算法得到的数值。
- 一种故障检测的系统,所述系统为分布式的节点集群系统,所述节点集群系统中包括多个节点,所述多个节点中每个节点如权利要求16至30中任一项所述的第一节点。
- 一种故障检测的装置,包括:处理器、收发器以及存储器,其中,所述处理器、所述收发器以及所述存储器通过总线连接,所述存储器存储有计算机指令,所述处理器通过执行计算机指令用于实现如权利要求1至15中任意一项所述的方法。
- 一种计算机可读存储介质,包括指令,当其在计算机上运行时,使得计算机执行如权利要求1至15中任意一项所述的方法。
- 一种计算机程序产品,该计算机程序产品包括计算机软件指令,该计算机软件指 令可通过处理器进行加载来实现如权利要求1至15中任意一项所述的方法。
Priority Applications (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2018/079422 WO2019178714A1 (zh) | 2018-03-19 | 2018-03-19 | 一种故障检测的方法、装置及系统 |
| CN201880091411.2A CN111869163B (zh) | 2018-03-19 | 2018-03-19 | 一种故障检测的方法、装置及系统 |
| EP18910654.5A EP3761559A4 (en) | 2018-03-19 | 2018-03-19 | ERROR DETECTION METHOD, DEVICE AND SYSTEM |
| US17/025,805 US20210006484A1 (en) | 2018-03-19 | 2020-09-18 | Fault detection method, apparatus, and system |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2018/079422 WO2019178714A1 (zh) | 2018-03-19 | 2018-03-19 | 一种故障检测的方法、装置及系统 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/025,805 Continuation US20210006484A1 (en) | 2018-03-19 | 2020-09-18 | Fault detection method, apparatus, and system |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2019178714A1 true WO2019178714A1 (zh) | 2019-09-26 |
Family
ID=67988268
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2018/079422 Ceased WO2019178714A1 (zh) | 2018-03-19 | 2018-03-19 | 一种故障检测的方法、装置及系统 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US20210006484A1 (zh) |
| EP (1) | EP3761559A4 (zh) |
| CN (1) | CN111869163B (zh) |
| WO (1) | WO2019178714A1 (zh) |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111901421A (zh) * | 2020-07-28 | 2020-11-06 | 腾讯科技(深圳)有限公司 | 一种数据处理方法及相关设备 |
| CN112804113A (zh) * | 2021-04-15 | 2021-05-14 | 北京全路通信信号研究设计院集团有限公司 | 一种故障判断方法及系统 |
| CN115643432A (zh) * | 2022-09-08 | 2023-01-24 | 湖南快乐阳光互动娱乐传媒有限公司 | p2p分享率异常检测方法及装置 |
| CN116127149A (zh) * | 2023-04-14 | 2023-05-16 | 杭州悦数科技有限公司 | 图数据库集群健康度的量化方法和系统 |
| CN119210992A (zh) * | 2024-09-23 | 2024-12-27 | 慧友科技股份有限公司 | 多节点智能调度的业务分段协同控制方法及系统 |
Families Citing this family (19)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| IT201900010362A1 (it) * | 2019-06-28 | 2020-12-28 | Telecom Italia Spa | Abilitazione della misura di perdita di pacchetti round-trip in una rete di comunicazioni a commutazione di pacchetto |
| CN111556345B (zh) * | 2020-03-19 | 2023-08-29 | 视联动力信息技术股份有限公司 | 一种网络质量检测的方法、装置、电子设备及存储介质 |
| US11811641B1 (en) * | 2020-03-20 | 2023-11-07 | Juniper Networks, Inc. | Secure network topology |
| WO2022085260A1 (ja) * | 2020-10-22 | 2022-04-28 | パナソニックIpマネジメント株式会社 | 異常検知装置、異常検知方法及びプログラム |
| US11584382B2 (en) * | 2021-02-12 | 2023-02-21 | Fca Us Llc | System and method for malfuncton operation machine stability determination |
| CN112988463B (zh) * | 2021-02-23 | 2022-08-30 | 新华三大数据技术有限公司 | 一种故障节点隔离方法及装置 |
| CN115348157B (zh) * | 2021-05-14 | 2023-09-05 | 中国移动通信集团浙江有限公司 | 分布式存储集群的故障定位方法、装置、设备及存储介质 |
| CN113312234B (zh) * | 2021-05-18 | 2022-07-26 | 福建天泉教育科技有限公司 | 一种健康检测的优化方法及终端 |
| CN115484268B (zh) * | 2021-05-31 | 2025-07-01 | 耀灵人工智能(浙江)有限公司 | 协同计算的对等网络与非特定特征识别的对等计算网络 |
| CN113760592B (zh) * | 2021-07-30 | 2024-02-27 | 郑州云海信息技术有限公司 | 一种节点内核检测方法和相关装置 |
| CN114285602B (zh) * | 2021-11-26 | 2024-02-02 | 成都安恒信息技术有限公司 | 一种分布式业务安全检测方法 |
| CN114979188A (zh) * | 2022-05-30 | 2022-08-30 | 阿里云计算有限公司 | 边缘设备的自愈方法、装置、电子设备及存储介质 |
| CN115225775B (zh) * | 2022-09-19 | 2022-12-09 | 苏州华兴源创科技股份有限公司 | 多通道的延迟修正方法、装置、计算机设备 |
| CN115550144B (zh) * | 2022-11-30 | 2023-03-24 | 季华实验室 | 分布式故障节点预测方法、装置、电子设备及存储介质 |
| US20240256357A1 (en) * | 2023-01-27 | 2024-08-01 | VMware LLC | Methods, systems and apparatus to provide a highly available cluster network |
| CN116668335B (zh) * | 2023-05-19 | 2026-04-10 | 超聚变数字技术股份有限公司 | 一种集群业务处理方法、服务器及系统 |
| CN118467112B (zh) * | 2024-07-12 | 2024-09-24 | 济南浪潮数据技术有限公司 | 一种故障处理方法、装置及设备、介质和计算机程序产品 |
| CN120223560A (zh) * | 2025-04-03 | 2025-06-27 | 吉林农业科技学院 | 一种温室大棚气候监控方法及系统 |
| CN120448178B (zh) * | 2025-07-10 | 2025-09-12 | 苏州元脑智能科技有限公司 | 故障诊断方法、系统、装置、电子设备及存储介质 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20050050398A1 (en) * | 2003-08-27 | 2005-03-03 | International Business Machines Corporation | Reliable fault resolution in a cluster |
| CN101795234A (zh) * | 2010-03-10 | 2010-08-04 | 北京航空航天大学 | 一种基于应用层组播算法的流媒体传输方案 |
| CN102355369A (zh) * | 2011-09-27 | 2012-02-15 | 华为技术有限公司 | 虚拟化集群系统及其处理方法和设备 |
| CN107204879A (zh) * | 2017-06-05 | 2017-09-26 | 浙江大学 | 一种基于指数移动平均的分布式系统自适应故障检测方法 |
Family Cites Families (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| PL1627316T3 (pl) * | 2003-05-27 | 2018-10-31 | Vringo Infrastructure Inc. | Zbieranie danych w klastrze komputerowym |
| JP2012173996A (ja) * | 2011-02-22 | 2012-09-10 | Nec Corp | クラスタシステム、クラスタ管理方法、およびクラスタ管理プログラム |
| CN103023716A (zh) * | 2012-11-26 | 2013-04-03 | 中怡(苏州)科技有限公司 | 一种零流量消耗的网络质量监控系统及监控方法 |
| JP6089884B2 (ja) * | 2013-03-29 | 2017-03-08 | 富士通株式会社 | 情報処理システム,情報処理装置,情報処理装置の制御プログラム,及び情報処理システムの制御方法 |
| WO2017008698A1 (zh) * | 2015-07-10 | 2017-01-19 | 努比亚技术有限公司 | 多通道路由方法及装置 |
| CN106998302B (zh) * | 2016-01-26 | 2020-04-14 | 华为技术有限公司 | 一种业务流量的分配方法及装置 |
| JP6709689B2 (ja) * | 2016-06-16 | 2020-06-17 | 株式会社日立製作所 | 計算機システム及び計算機システムの制御方法 |
-
2018
- 2018-03-19 EP EP18910654.5A patent/EP3761559A4/en not_active Withdrawn
- 2018-03-19 WO PCT/CN2018/079422 patent/WO2019178714A1/zh not_active Ceased
- 2018-03-19 CN CN201880091411.2A patent/CN111869163B/zh active Active
-
2020
- 2020-09-18 US US17/025,805 patent/US20210006484A1/en not_active Abandoned
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20050050398A1 (en) * | 2003-08-27 | 2005-03-03 | International Business Machines Corporation | Reliable fault resolution in a cluster |
| CN101795234A (zh) * | 2010-03-10 | 2010-08-04 | 北京航空航天大学 | 一种基于应用层组播算法的流媒体传输方案 |
| CN102355369A (zh) * | 2011-09-27 | 2012-02-15 | 华为技术有限公司 | 虚拟化集群系统及其处理方法和设备 |
| CN107204879A (zh) * | 2017-06-05 | 2017-09-26 | 浙江大学 | 一种基于指数移动平均的分布式系统自适应故障检测方法 |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP3761559A4 * |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111901421A (zh) * | 2020-07-28 | 2020-11-06 | 腾讯科技(深圳)有限公司 | 一种数据处理方法及相关设备 |
| CN112804113A (zh) * | 2021-04-15 | 2021-05-14 | 北京全路通信信号研究设计院集团有限公司 | 一种故障判断方法及系统 |
| CN115643432A (zh) * | 2022-09-08 | 2023-01-24 | 湖南快乐阳光互动娱乐传媒有限公司 | p2p分享率异常检测方法及装置 |
| CN116127149A (zh) * | 2023-04-14 | 2023-05-16 | 杭州悦数科技有限公司 | 图数据库集群健康度的量化方法和系统 |
| CN119210992A (zh) * | 2024-09-23 | 2024-12-27 | 慧友科技股份有限公司 | 多节点智能调度的业务分段协同控制方法及系统 |
Also Published As
| Publication number | Publication date |
|---|---|
| EP3761559A1 (en) | 2021-01-06 |
| EP3761559A4 (en) | 2021-03-17 |
| CN111869163A (zh) | 2020-10-30 |
| CN111869163B (zh) | 2022-05-24 |
| US20210006484A1 (en) | 2021-01-07 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN111869163B (zh) | 一种故障检测的方法、装置及系统 | |
| US11575559B1 (en) | Monitoring and detecting causes of failures of network paths | |
| KR101881409B1 (ko) | 소프트웨어 정의 네트워크에서 멀티-마스터 선택 | |
| CA2973991C (en) | Determining link conditions of a client lan/wan from measurement point to client devices and application servers of interest | |
| US8661295B1 (en) | Monitoring and detecting causes of failures of network paths | |
| CN104506482B (zh) | 网络攻击检测方法及装置 | |
| US9009305B1 (en) | Network host inference system | |
| CN101483547A (zh) | 一种网络突发事件度量评估方法及系统 | |
| US8971871B2 (en) | Radio base station, control apparatus, and abnormality detection method | |
| EP3682595A1 (en) | Obtaining local area network diagnostic test results | |
| US20140198660A1 (en) | Communication monitor, occurrence prediction method, and recording medium | |
| US12074807B2 (en) | Detecting shortfalls in an agreement between a publisher and a subscriber | |
| WO2012012986A1 (zh) | 告警防抖动的处理方法及装置 | |
| JP6531755B2 (ja) | ネットワーク制御方法およびシステム | |
| US11652682B2 (en) | Operations management apparatus, operations management system, and operations management method | |
| JP4299210B2 (ja) | ネットワーク監視方法及び装置 | |
| KR20240174667A (ko) | Nf 장치, nf에서 수행되는 시그널링 제어 방법 | |
| CN109688031B (zh) | 一种网络监控方法及相关设备 | |
| CN115733726A (zh) | 网络群障确定方法、装置、存储介质及电子装置 | |
| CN107018016B (zh) | 一种用于监控联网设备的方法和装置 | |
| KR100921335B1 (ko) | 인터넷 트래픽 특성을 이용한 회선의 안정성 진단 장치 및그 방법 | |
| JP2016025653A (ja) | 通信システムおよび通信システムの通信方法 | |
| JP5537692B1 (ja) | 品質劣化原因推定装置、品質劣化原因推定方法、品質劣化原因推定プログラム | |
| JP6920835B2 (ja) | 設備監視装置 | |
| KR101231600B1 (ko) | 단말의 접속 유형 판별 방법 및 시스템 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 18910654 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2018910654 Country of ref document: EP Effective date: 20200929 |

