WO2021232969A1 - 动作识别方法、装置、设备及存储介质 - Google Patents

动作识别方法、装置、设备及存储介质 Download PDF

Info

Publication number
WO2021232969A1
WO2021232969A1 PCT/CN2021/085386 CN2021085386W WO2021232969A1 WO 2021232969 A1 WO2021232969 A1 WO 2021232969A1 CN 2021085386 W CN2021085386 W CN 2021085386W WO 2021232969 A1 WO2021232969 A1 WO 2021232969A1
Authority
WO
WIPO (PCT)
Prior art keywords
video data
layer
frame
feature map
feature
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2021/085386
Other languages
English (en)
French (fr)
Inventor
李斌泉
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Bigo Technology Pte Ltd
Original Assignee
Bigo Technology Pte Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Bigo Technology Pte Ltd filed Critical Bigo Technology Pte Ltd
Priority to EP21808837.5A priority Critical patent/EP4156017B1/en
Priority to US17/999,284 priority patent/US12412426B2/en
Publication of WO2021232969A1 publication Critical patent/WO2021232969A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/764Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/7715Feature extraction, e.g. by transforming the feature space, e.g. multi-dimensional scaling [MDS]; Mappings, e.g. subspace methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/80Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
    • G06V10/806Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level of extracted features
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/40Scenes; Scene-specific elements in video content
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/40Scenes; Scene-specific elements in video content
    • G06V20/41Higher-level, semantic clustering, classification or understanding of video scenes, e.g. detection, labelling or Markovian modelling of sport events or news items
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/40Scenes; Scene-specific elements in video content
    • G06V20/44Event detection
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/40Scenes; Scene-specific elements in video content
    • G06V20/46Extracting features or characteristics from the video content, e.g. video fingerprints, representative shots or key frames
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/40Scenes; Scene-specific elements in video content
    • G06V20/49Segmenting video sequences, i.e. computational techniques such as parsing or cutting the sequence, low-level clustering or determining units such as shots or scenes
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/20Movements or behaviour, e.g. gesture recognition
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/17Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
    • H04N19/172Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a picture, frame or field
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/50Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
    • H04N19/503Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
    • H04N19/51Motion estimation or motion compensation
    • H04N19/513Processing of motion vectors
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/50Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
    • H04N19/503Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
    • H04N19/51Motion estimation or motion compensation
    • H04N19/513Processing of motion vectors
    • H04N19/517Processing of motion vectors by encoding
    • H04N19/52Processing of motion vectors by encoding by predictive encoding

Definitions

  • the embodiments of the present application relate to the technical field of computer vision applications, such as motion recognition methods, devices, equipment, and storage media.
  • Video-based action recognition has always been an important field of computer vision research.
  • the realization of video action recognition mainly includes two parts: feature extraction and representation, and feature classification.
  • Classical methods such as density trajectory tracking are generally the method of manually designing features.
  • people have found that deep learning has powerful feature representation capabilities, and neural networks have gradually become the mainstream method in the field of action recognition.
  • the feature method greatly improves the performance of action recognition.
  • most of the neural network action recognition schemes are based on the sequence of pictures obtained from the video to construct the timing relationship, so as to judge the action. For example: based on the recurrent neural network to construct the timing between pictures, based on the 3D convolution to extract the timing information of multiple pictures, and the picture-based deep learning technology to superimpose the optical flow information of the action change, etc.
  • the above scheme has the situation that the calculation amount and the recognition accuracy cannot be taken into account, and it needs to be improved.
  • the embodiments of the present application provide an action recognition method, device, equipment, and storage medium, which can optimize an action recognition solution for videos in related technologies.
  • an action recognition method which includes:
  • the packet video data to be identified is input into a second preset model, and the action type included in the packet video data to be identified is determined according to the output result of the second preset model.
  • an action recognition device which includes:
  • the video grouping module is configured to perform grouping processing on the original compressed video data to obtain grouped video data
  • a target grouped video determining module configured to input the grouped video data into a first preset model, and determine the target grouped video data containing actions according to the output result of the first preset model;
  • the video decoding module is configured to decode the target packet video data to obtain the packet video data to be identified;
  • the action type identification module is configured to input the packet video data to be identified into a second preset model, and determine the action type contained in the packet video data to be identified according to the output result of the second preset model.
  • embodiments of the present application provide a computer device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor.
  • the processor executes the computer program
  • the computer program is The action recognition method provided by the embodiment.
  • an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the action recognition method as provided in the embodiment of the present application is implemented.
  • FIG. 1 is a schematic flowchart of an action recognition method provided by an embodiment of this application.
  • FIG. 2 is a schematic diagram of a frame arrangement in a compressed video provided by an embodiment of the application
  • FIG. 3 is a schematic diagram of a feature transformation operation provided by an embodiment of this application.
  • FIG. 5 is a schematic diagram of an action recognition process based on a compressed video stream provided by an embodiment of this application;
  • FIG. 6 is a schematic structural diagram of a first 2D residual network provided by an embodiment of this application.
  • FIG. 7 is a schematic diagram of an action label provided by an embodiment of the application.
  • FIG. 8 is a schematic flowchart of another action recognition method provided by an embodiment of this application.
  • FIG. 9 is a schematic diagram of an application scenario of short video-based action recognition provided by an embodiment of the application.
  • FIG. 10 is a schematic diagram of an action recognition process based on sequence pictures according to an embodiment of this application.
  • FIG. 11 is a schematic diagram of a second 2D residual network structure provided by an embodiment of this application.
  • FIG. 12 is a schematic diagram of a 3D residual network structure provided by an embodiment of this application.
  • FIG. 13 is a schematic diagram of a calculation process of a two-level multi-receptive field pooling operation provided by an embodiment of the application;
  • FIG. 14 is a structural block diagram of an action recognition device provided by an embodiment of this application.
  • FIG. 15 is a structural block diagram of a computer device provided by an embodiment of this application.
  • the action recognition solution in the embodiments of the present application can be applied to various video-oriented action recognition scenarios, such as short video review scenarios, video surveillance scenarios, real-time call recognition scenarios, robot visual recognition scenarios, and so on.
  • the video can be a video file or a video stream.
  • most of the neural network action recognition schemes are based on the sequence of pictures obtained from the video to construct the timing relationship, so as to judge the action. For example: based on Recurrent Neural Network (RNN) or Long Short-Term Memory (LSTM, Long Short-Term Memory) to construct the timing between pictures, extract the timing information of multiple pictures based on 3D convolution, and based on pictures
  • RNN Recurrent Neural Network
  • LSTM Long Short-Term Memory
  • the deep learning technology superimposes the optical flow information of the action changes.
  • the method of action recognition based on sequence images has at least the following two shortcomings: First, these technical solutions require a lot of computing resources to make judgments on a short video, and rely heavily on the machine central processing unit ( Central Processing Unit (CPU) and graphics processor (Graphics Processing Unit, GPU) computing power, and extracting pictures from compressed short videos needs to be decoded, such as decode decoding (the technology to decode compressed videos into pictures), and The full-time and long-segment decoding of short-term video itself requires a large amount of CPU and GPU computing power. Therefore, the image-based motion recognition scheme requires high machine computing power.
  • CPU Central Processing Unit
  • GPU Graphics Processing Unit
  • Fig. 1 is a schematic flowchart of an action recognition method provided by an embodiment of the application.
  • the method can be executed by an action recognition device, where the device can be implemented by software and/or hardware, and generally can be integrated in a computer device.
  • the computer device may be, for example, a server, or a device such as a mobile phone, or two or more devices may perform part of the steps respectively.
  • the method includes:
  • Step 101 Perform grouping processing on the original compressed video data to obtain grouped video data.
  • video data contains a large amount of image and sound information.
  • the coding standard and coding parameters are not limited, and may be H264, H265, MPEG-4, etc., for example.
  • the original compressed video data may be divided into intervals based on preset grouping rules to obtain interval compressed video (segment, hereinafter may be referred to as segment), the interval compressed video is used as the grouped video data, or part of the data is selected from the interval compressed video As packet video data.
  • the preset grouping rule may include, for example, interval division time intervals, that is, the duration corresponding to each interval compressed video after being divided by intervals.
  • the interval division time interval may be constant or variable, and is not limited. Taking a short video review scenario as an example, the interval division time interval may be, for example, 5 seconds.
  • the preset grouping rule for example, may also include the number of intervals, and the value is not limited.
  • the packet processing and interval division can be completed by using time stamps, that is, the interval range of the grouped video data or interval compressed video can be defined by the start and end timestamps. This step can be understood as extracting data of different time periods from the original compressed video data and inputting them to the first preset model respectively.
  • the data corresponding to 0 ⁇ 5s in the original compressed video is a piece of packet video data
  • the data corresponding to 5 ⁇ 10s is a piece of packet video data
  • the two pieces of data are respectively entered into the first preset model of.
  • Step 102 Input the grouped video data into a first preset model, and determine the target grouped video data containing the action according to the output result of the first preset model.
  • each piece of grouped video data can be input into the first preset model to improve calculation efficiency; or after all grouping is completed, each segment can be grouped sequentially or in parallel
  • the video data is input into the first preset model to ensure that coarse-grained action recognition is performed when the grouping processing is accurately completed.
  • the first preset model may be a pre-trained neural network model, which is directly loaded when needed.
  • the model is mainly set to identify whether the grouped video data contains an action, and does not care which action it is. You can set the action
  • the label is two-category, such as "yes” and "no", the two-category label can be marked in the training sample of the first preset model. In this way, according to the output result of the first preset model, it is possible to filter out which segments contain actions, and determine the corresponding grouped video data as the target grouped video data.
  • the calculation amount of the first preset model is small, and the recognition is performed without decoding, which can save a lot of decoding power and eliminate a large number of errors.
  • the video segment containing the action is also guaranteed to be retained for subsequent identification.
  • the network structure and related parameters of the first preset model are not limited in the embodiment of the present application, and can be set according to actual needs, for example, it can be a lightweight model.
  • Step 103 Decode the target packet video data to obtain the packet video data to be identified.
  • an appropriate decoding method can be selected with reference to factors such as the encoding standard of the original compressed video, which is not limited.
  • the obtained video image can be used as the grouped video data to be identified (the decoded video images are generally arranged in chronological order, that is, sequence images.
  • the grouped video data to be identified can contain video images Time sequence information), other information can also be extracted on this basis and used as the grouped video data to be identified.
  • the other information here can be, for example, frequency domain information.
  • Step 104 Input the packet video data to be identified into a second preset model, and determine the action type included in the packet video data to be identified according to the output result of the second preset model.
  • the second preset model may be a pre-trained neural network model, which is directly loaded when needed.
  • the model is mainly set to identify the types of actions contained in the grouped video data to be recognized, that is, perform fine-grained For recognition, multi-category labels can be marked in the training samples of the second preset model, so that the final action recognition result can be determined according to the output result of the second preset model.
  • the data that can enter the second preset model has been greatly reduced compared to the original compressed video data, and the data purity is much higher than that of the original compressed video data.
  • the magnitude of the number of video clips that need to be identified is not large. , So you can use the decoded sequence image for identification.
  • the 3D convolutional network structure with more neural network parameters can be used to extract the time series features, and the tags are multi-classifications with finer granularity.
  • the number of tags is not limited. For example, it can be 50. The recognition accuracy is adjusted.
  • the original compressed video data is grouped to obtain the grouped video data, the grouped video data is input into the first preset model, and the output result of the first preset model is determined to include
  • the target packet video data of the action is decoded to obtain the packet video data to be identified, and the packet video data to be identified is input into the second preset model, and the to be identified is determined according to the output result of the second preset model
  • the type of action contained in the packet video data is used before decompressing the compressed video, the first preset model is used to roughly filter out the video clips that contain actions, and then the second preset model is used to accurately identify the types of actions contained, which can ensure recognition accuracy. Under the premise of effectively reducing the amount of calculation and improving the efficiency of action recognition.
  • the client can perform preliminary screening of the compressed video to be uploaded, and upload the target grouped video data containing the action to the server for identification and review.
  • the entire recognition process can also be completed locally on the client, that is, the related operations from step 101 to step 104, to realize the control of whether to allow the upload of the video according to the finally recognized action type. .
  • the grouping processing of the original compressed video data to obtain the grouped video data includes: dividing the original compressed video data into intervals based on a preset grouping rule to obtain the interval compressed video; and extracting the compressed video based on a preset extraction strategy.
  • the I frame data and P frame data in the interval compression video are used to obtain packetized video data, where the P frame data includes motion vector information and/or color element change residual information corresponding to the P frame.
  • I frame also known as key frame
  • P frame also known as forward predictive coding frame
  • I frame generally includes the change information of the reference I frame in the compressed video, and may include motion vector (MV) information
  • RGBR residual
  • the distribution or content of the I frame and the P frame may be different.
  • the preset grouping rule may be as described above, for example, may include the interval division time interval or the number of segments, etc.
  • the preset extraction strategy may include, for example, an extraction strategy for I frames and an extraction strategy for P frames.
  • the extraction strategy of I frames may include the time interval for acquiring I frames, or the number of I frames acquired in a unit time, for example, acquiring 1 I frame every second.
  • the P frame extraction strategy may include acquiring the number of P frames following an I frame and the time interval between every two acquired P frames, for example, the number is two. Then, taking an interval compressed video of 5 seconds as an example, data corresponding to 5 I frames and 10 P frames can be obtained. If the P frame data includes both MV and RGBR, 10 MVs and 10 RGBRs can be obtained.
  • the first preset model includes a first 2D residual network, a first splicing layer, and a first fully connected layer; after the packet video data is input into the first preset model , Obtaining corresponding feature maps of the same dimension through the first 2D residual network; obtaining the splicing feature maps after performing splicing operations in the order of the frames via the first splicing layer; the splicing feature maps A classification result of whether an action is included is obtained through the first fully connected layer.
  • the technical solutions provided by the embodiments of the present application use relatively simplified network results to obtain higher recognition efficiency and ensure a higher recall rate of video clips containing actions.
  • the first 2D residual network may adopt a lightweight ResNet18 model. Since I frame, MV, and RGBR cannot be directly concatenated in the data features, three ResNet18 can be used to process the I frame, MV and RGBR separately to obtain feature maps with the same dimensions, which can be recorded separately I frame feature maps, MV feature maps, and RGBR feature maps.
  • MV feature maps and RGBR feature maps are collectively referred to as P frame feature maps.
  • the same dimension can mean that C*H*W is consistent, where C represents channel, H represents height, W represents width, and * can also be represented as ⁇ .
  • feature maps with the same dimension can be spliced through the first splicing layer.
  • extracting I frame data and P frame data in the interval compressed video based on a preset extraction strategy to obtain packetized video data includes: extracting I frame data in the interval compressed video based on a preset extraction strategy And P frame data; accumulatively transform the P frame data, so that the transformed P frame data depends on the forward adjacent I frame; determine the packet video data according to the I frame data and the transformed P frame data .
  • the first preset model further includes an addition layer located before the splicing layer, the feature map corresponding to the P frame data in the feature map is denoted as the P frame feature map, and the feature map corresponds to I
  • the feature map of the frame data is marked as an I frame feature map;
  • the P frame feature map and the I frame feature map obtain the P frame feature after the addition operation on the basis of the I frame feature map through the addition layer Figure;
  • the I frame feature map and the P frame feature map after the addition operation are obtained through the first splicing layer after the splicing operation is performed in the order of the frames.
  • Figure 2 is a schematic diagram of frame arrangement in a compressed video provided by an embodiment of this application.
  • the MV and RGBR of each P frame depend on the previous one.
  • the P frame of the frame (such as the P2 frame depends on the P1 frame).
  • the MV and RGBR of the P frame can be cumulatively transformed to obtain the MV and the P frame of the input neural network.
  • RGBR is relative to the previous I frame (for example, after accumulative transformation, the P2 frame becomes dependent on the previous I frame), rather than relative to the previous P frame.
  • the above-mentioned addition operation can refer to the residual addition method in ResNet to directly output the MV feature map and RGBR feature map according to each element (pixel) and I frame (I frame after the first 2D residual network processing Feature maps) are added, and then the 3 feature maps are spliced in the order of the frames.
  • a feature transformation layer is included before the residual structure of the first 2D residual network; before the packet video data enters the residual structure, an upward feature transformation is obtained through the feature transformation layer. And/or down-characterized grouped video data.
  • the feature map before entering the residual structure, that is, before performing the convolution operation, the feature map may be subjected to a feature shift (shift) operation through Feature Shift (FS), Make a feature map contain some of the features in the feature map at different time points, so that when the convolution operation is performed, the feature map contains timing information, which can handle the collection of timing information without increasing the amount of calculation. And fusion, enrich the information of the feature map to be recognized, and improve the recognition accuracy. Compared with the related technology based on 3D convolution or optical flow information, it can effectively reduce the amount of calculation.
  • FIG. 3 is a schematic diagram of a feature transformation operation provided by an embodiment of this application.
  • Figure 3 taking the MV feature map as an example, for ease of description, Figure 3 only shows the process of performing feature transformation operations on three MV feature maps.
  • the three MV feature maps correspond to different time points.
  • Divide the MV feature map into 4 parts. Assuming that the first part is subjected to downward feature transformation (Down Shift), the second part is subjected to upward feature transformation (Up Shift), and the second MV feature map after transformation is at the same time Contains some features of the MV feature map at 3 time points.
  • the division rule and the area of the upward feature transformation and/or the downward feature transformation can be set according to the actual situation.
  • extracting P frame data in the interval compressed video based on a preset extraction strategy includes: extracting a preset number of P frame data in the interval compressed video in an equally spaced manner; wherein, in the In the training phase of the first preset model, a preset number of P-frame data in the interval compressed video is extracted in a random interval manner.
  • FIG. 4 is a schematic flowchart of another action recognition method provided by an embodiment of the application, which is refined on the basis of the foregoing exemplary embodiments.
  • the method includes:
  • Step 401 Perform interval division on the original compressed video data based on a preset grouping rule to obtain an interval compressed video.
  • the original compressed video can be enhanced with a preset video and image enhancement strategy, and the enhancement strategy can be selected according to the data requirements of the service and the configuration conditions of the driver.
  • the same enhancement strategy can be used in model training and model application.
  • Step 402 Extract I frame data and P frame data in the interval compressed video based on a preset extraction strategy.
  • the P frame data includes the motion vector information and the pigment change residual information corresponding to the P frame.
  • Step 403 Perform cumulative transformation on the P frame data, so that the transformed P frame data depends on the forward adjacent I frame, and determine the grouped video data according to the I frame data and the transformed P frame data.
  • Step 404 Input the grouped video data into the first preset model, and determine the target grouped video data containing the action according to the output result of the first preset model.
  • FIG. 5 is a schematic diagram of an action recognition process based on a compressed video stream provided by an embodiment of the application.
  • the compressed video is divided into n segments.
  • the extracted I frame data, MV The data and RGBR data are input into the first preset model.
  • the first preset model includes the first 2D residual network, the addition layer, the first splicing layer, and the first fully connected layer (Fully connected layer, FC layer).
  • the first 2D residual network is 2D Res18.
  • the feature transformation layer (FS) is included before the residual structure.
  • the training method in the training phase of the first preset model, can be selected according to actual needs, including loss function (loss), etc., for example, the loss function can use cross entropy, and other auxiliary loss functions can also be used to improve the model. Effect.
  • some non-heuristic optimization algorithms can be used to improve the convergence speed and optimization performance of stochastic gradient descent.
  • Fig. 6 is a schematic diagram of a first 2D residual network structure provided by an embodiment of this application.
  • 2D Res18 is an 18-layer residual neural network, which consists of 4 stages and a total of 8 uses 2D Convolutional residual block (block) composition, the network structure is relatively shallow, in order to use the convolutional layer as much as possible to fully extract features, 8 2D residual blocks can all be performed by 3*3 convolution, that is, the convolution kernel parameters It is 3*3.
  • the convolutional layer generally refers to a network layer used to complete the weighted summation of local pixel values and nonlinear activation.
  • each 2D residual block can use a bottleneck.
  • the design concept of (bottleneck) is that each residual block is composed of 3 convolutional layers (convolution kernel parameters are 1*1, 3*3, and 1*1), and the first and last layers are used for compression. And restore the image channel.
  • the first 2D residual network structure and various parameters can also be adjusted according to actual needs.
  • the I frame data, MV data and RGBR data are respectively obtained through the FS layer through the upward feature transformation and the downward feature transformation of the I frame data.
  • MV data and RGBR data and then obtain a C*H*W consistent feature map through the residual structure respectively, and directly compare the MV feature map and RGBR feature map with the residual addition method in ResNet through the addition layer
  • Add each element to the output of the I frame and then concate the three feature maps in the order of the frame to get the spliced feature map.
  • the spliced feature map is passed through the first fully connected layer (FC) to get whether it contains action The classification results.
  • FC fully connected layer
  • Figure 7 is a schematic diagram of an action label provided by an embodiment of the application.
  • the last two circles in Figure 5 indicate that the output is two-category.
  • Setting the action label to two-category is to ensure that the recall of fragments containing actions is improved.
  • the label granularity is whether to include The action is designed to have a two-category level of "Yes” or "No", that is, it does not care which sub-category action is included.
  • the short video used for training can be cut into 6 segments, where segment S2 contains action A1, segment S4 contains action A2, and A1 and A2 are two different actions, but The labels corresponding to these two positions are the same, and they are both set to 1, which means that the actions of A1 and A2 are not distinguished. Therefore, in actual application, the output result of the first preset model is also "Yes” or "No", which realizes coarse-grained recognition of actions in the compressed video.
  • Step 405 Decode the target packet video data to obtain the packet video data to be identified.
  • Step 406 Input the packet video data to be identified into the second preset model, and determine the action type included in the packet video data to be identified according to the output result of the second preset model.
  • the action recognition method provided by the embodiments of this application first performs action recognition based on compressed video. Without decompressing the video, extracts the MV and RGBR information of the I frame and the P frame, and changes the MV and RGBR to increase the I frame.
  • the information dependence of realizes the processing of a large number of videos of variable duration with less computational power requirements, and uses FS without computational power requirements to increase the model's extraction of timing information.
  • the enhancement of model capabilities does not lead to a decrease in computational efficiency. Redesigned the label granularity of actions to ensure the goal of recall. Let the model deal with simple binary classification problems, which can improve the recall rate, and then accurately identify the preliminarily screened video clips containing actions, which is effective under the premise of ensuring recognition accuracy. Reduce the amount of calculation and improve the efficiency of action recognition.
  • the decoding the target packet video data to obtain the to-be-identified packet video data includes: decoding the target packet video data to obtain the to-be-identified segmented video image; and obtaining the to-be-identified segmented video image
  • the frequency domain information in the segmented video image is generated according to the frequency domain information, and the corresponding frequency domain map is generated; the segmented video image to be identified and the corresponding frequency domain map are used as the packetized video data to be identified.
  • the second preset model includes a model based on an Efficient Convolutional Network for Online Video (ECO) architecture for online video understanding.
  • ECO Efficient Convolutional Network for Online Video
  • the ECO architecture can be understood as a video feature extractor. It provides an architecture design for video feature acquisition, which includes a 2D feature extraction network and a 3D feature extraction network. This architecture can achieve better performance while increasing speed.
  • This application is implemented The example can be improved and designed on the basis of the ECO architecture to obtain the second preset model.
  • the second preset model includes a second splicing layer, a second 2D residual network, a 3D residual network, a third splicing layer, and a second fully connected layer;
  • the packet video data to be identified is After being input into the second preset model, the spliced image data obtained by splicing the segmented video image to be identified and the corresponding frequency domain map is obtained through the second splicing layer;
  • the spliced image data passes through the first
  • the second 2D residual network obtains a 2D feature map;
  • the middle layer output result of the second 2D residual network is used as the input of the 3D residual network, and a 3D feature map is obtained through the 3D residual network;
  • the 2D The feature map and the 3D feature map obtain a spliced feature map to be identified via the third splicing layer;
  • the feature map to be identified obtains a corresponding action type label via the second fully connected layer.
  • the technical solution provided by the embodiment of the present application
  • the grouped video data to be identified first enters the 2D convolution part of the model to extract the features of each image, then enters the 3D convolution part to extract the timing information of the action, and finally outputs the result of the multi-classification.
  • the network structure and related parameters of the second 2D residual network and the 3D residual network can be set according to actual requirements.
  • the second 2D residual network is 2D Res50
  • the 3D residual network is 3D Res10.
  • the second preset model further includes a first pooling layer and a second pooling layer; the 2D feature map passes through the first pooling layer before being input to the third stitching layer Obtain the corresponding one-dimensional 2D feature vector containing the first element quantity; before the 3D feature map is input to the third stitching layer, obtain the corresponding one containing the second element quantity through the second pooling layer Dimensional 3D feature vector.
  • the 2D feature map and the 3D feature map obtain the spliced feature map to be recognized through the third splicing layer, including: the one-dimensional 2D feature vector and the one-dimensional 3D feature vector are passed through the The third splicing layer obtains the spliced vector to be recognized.
  • the feature map still has a larger size after feature extraction through the convolutional layer.
  • the feature map is directly flattened, the dimension of the feature vector may be too high. Therefore, the 2D and After the 3D residual network completes a series of feature extraction operations, the pooling operation can be used to directly aggregate the feature maps into one-dimensional feature vectors to reduce the dimensionality of the feature vectors.
  • the first pooling layer and the second pooling layer can be designed according to actual needs, for example, it can be global average pooling (GAP) or maximum pooling.
  • GAP global average pooling
  • the number of the first element and the number of the second element can also be set freely, for example, according to the number of channels of the feature map.
  • the first pooling layer includes a multi-receptive field pooling layer.
  • the receptive field generally refers to the image or video range covered by the feature value, which is used to indicate the size of the receptive range of the original image by different neurons in the network, or in other words, the pixels on the feature map output by each layer of the convolutional neural network The size of the area mapped on the original image.
  • the larger the value of the receptive field the larger the range of the original image it can touch, and it also means that it may contain more global and higher semantic features; on the contrary, the smaller the value, the more the features it contains. Part and detail. Therefore, the value of the receptive field can be used to roughly judge the abstraction level of each layer.
  • the advantage of adopting the multi-sensory field pooling layer is that the features can be sensitive to targets of different scales, and the recognition range of the action category is broadened.
  • the multi-receptive field realization method can be set according to actual needs.
  • the first pooling layer includes a first-level local pooling layer, a second-level global pooling layer, and a vector merging layer
  • the first-level local pooling layer includes at least two pools with different sizes.
  • the core; the 2D feature map obtains corresponding at least two sets of 2D pooling feature maps of different scales through the first-level local pooling layer; the at least two sets of 2D pooling feature maps of different scales are passed through the second level
  • the global pooling layer obtains at least two sets of feature vectors; the at least two sets of feature vectors obtain corresponding one-dimensional 2D feature vectors containing the first number of elements through the vector merging layer.
  • the technical solution provided by the embodiments of this application adopts two-level multi-receptive field pooling to summarize the total feature map, and uses different size pooling cores for multi-scale pooling, so that the summarized features have different receptive fields, and at the same time, the feature map
  • the size is greatly reduced, the recognition efficiency is improved, and then the two-level global pooling is used to summarize the features, and the vector merging layer is used to obtain a one-dimensional 2D feature vector.
  • Figure 8 is a schematic flow chart of another action recognition method provided by an embodiment of this application. It is refined on the basis of the above-mentioned exemplary embodiments.
  • Figure 9 is a short-based A schematic diagram of the application scenario of video action recognition, as shown in Figure 9, after the user uploads a short video, first extract the compressed video stream information, and use the pre-trained and constructed action video model based on the compressed video (the first preset model) to identify that it contains The target segment of the action. Other segments are screened out because they do not contain actions.
  • the target segment is decoded and the time domain and frequency domain information of the picture is extracted, and the action recognition model based on the decoded picture constructed in advance is used (the second preset The model) performs finer-grained action recognition based on the picture sequence, and obtains the action type corresponding to each target segment, and then determines whether it is a target action. If it is not a target action, it is also filtered out.
  • the method may include:
  • Step 801 Perform interval division on the original compressed video data based on a preset grouping rule to obtain an interval compressed video.
  • the original compressed video data is the video stream data of the short video uploaded by the user, and the length of each section of the compressed video may be 5 seconds.
  • Step 802 Extract I frame data and P frame data in the interval compressed video based on a preset extraction strategy.
  • the P frame data contains the MV and RGBR corresponding to the P frame, and 1 I frame and 2 P frames thereafter are obtained every second, that is, a 5 second segment can obtain 5 I frames, 10 MVs and 10 RGBR.
  • Step 803 Perform cumulative transformation on the P frame data, so that the transformed P frame data depends on the forward adjacent I frame, and determine the grouped video data according to the I frame data and the transformed P frame data.
  • Step 804 Input the grouped video data into the first preset model, and determine the target grouped video data containing the action according to the output result of the first preset model.
  • Step 805 Decode the target grouped video data to obtain the segmented video image to be identified.
  • the images are decoded from the target grouped video data, and they are arranged in a sequence in chronological order to obtain the sequence pictures.
  • a set number of images can be acquired from the sequence of pictures according to a preset acquisition strategy as the segmented video images to be identified.
  • the preset acquisition strategy is, for example, acquisition at equal intervals, and the set number is, for example, 15.
  • Step 806 Obtain frequency domain information in the segmented video image to be identified, and generate a corresponding frequency domain map according to the frequency domain information.
  • the frequency domain information collection method is not limited, and the frequency domain information can be collected for each picture in the sequence picture, and a frequency domain (FD) map corresponding to the sequence picture one-to-one can be generated.
  • FD frequency domain
  • Step 807 Use the segmented video image to be identified and the corresponding frequency domain map as the grouped video data to be identified, input them into the second preset model, and determine according to the output result of the second preset model that the grouped video data to be identified contains The type of action.
  • the sequence picture and the corresponding frequency domain map in one segment will correspond to one label in the multi-category label, that is, in the output result of the second preset model, each segment corresponds to one label.
  • the model will use the corresponding type of the main action as the label of the segment. For example, in a segment A of 5 seconds, there is an action in 4 seconds, the action in 3 seconds is a hand wave, and the action in 1 second is a kick, then the action tag corresponding to the segment A is a hand wave.
  • FIG. 10 is a schematic diagram of an action recognition process based on a sequence of pictures provided by an embodiment of the present application, where the value of n is generally different from the value of n in FIG. Picture (Image), after extracting the frequency domain information, generate the corresponding frequency domain map (FD).
  • the second preset model can be designed based on the ECO architecture, which can include the second splicing layer, the second 2D residual network, the 3D residual network, the first pooling layer, the second pooling layer, the third splicing layer, and the second Fully connected layer.
  • training methods including loss functions, can be selected according to actual needs.
  • the loss function can use cross entropy, and other auxiliary loss functions can also be used to improve the model effect.
  • some non-heuristic optimization algorithms can be used to improve the convergence speed and optimization performance of stochastic gradient descent.
  • the decoded sequence picture and the corresponding frequency domain map pass through the second splicing layer in the second preset model to obtain spliced image data.
  • the subsequent spliced image data may also pass through the convolutional layer. After processing (conv) and the maximum pooling layer (maxpool), they are then input into the second 2D residual network (2D Res50).
  • the output of 2D Res50 passes through the first pooling layer (multi-receptive field pooling layer) to obtain a 1024-dimensional 2D feature vector (which can be understood as a one-dimensional row vector or column vector containing 1024 elements).
  • the intermediate output result of the 2D Res50 will be used as the input of the 3D residual network (3D Res10), and the 512-dimensional 3D feature vector will be obtained after the second pooling layer (GAP).
  • the feature map to be recognized that is, the feature vector to be recognized, is obtained, and finally the final action type label is obtained through the second fully connected layer.
  • FIG 11 is a schematic diagram of a second 2D residual network structure provided by an embodiment of the application.
  • the second 2D residual network is 2D Res50, which is a 50-layer residual neural network, consisting of 4 stages in total It is composed of 16 residual blocks using 2D convolution.
  • each residual block can use the bottleneck design concept, that is, each residual block is composed of 3 convolutional layers, of which the import and export two 1* 1 Convolution is used to compress and restore the number of channels of the feature map.
  • 2D projection residual blocks can be used at the entrance of each stage.
  • This residual block is A 1*1 convolutional layer is added to the bypass to ensure that the size of the feature map and the number of channels remain the same when the pixel-by-pixel addition operation is performed.
  • only using 2D projection residual blocks at the entrance of each stage can also reduce network parameters.
  • FIG. 12 is a schematic diagram of a 3D residual network structure provided by an embodiment of this application.
  • Each segment takes N video frames (frame pictures obtained after the splicing operation of the 15 images obtained from the sequence of images and the corresponding frequency domain images as described above) the middle layer feature map group obtained through 2D Res50, for example
  • the feature map group from stage2-block4 can be assembled into a three-dimensional video tensor, the tensor shape is (c, f, h, w), where c is the number of channels of the frame picture, f is the number of video frames, h, w refers to the height and width of the frame picture respectively, and the video tensor is input to 3D Res10 to extract the spatiotemporal features of the entire video.
  • 3D Res10 consists of only 3 stages and a total of 5 residual blocks. All convolutional layers use three-dimensional convolution kernels. During the convolution process, the information in the time dimension will also participate in the calculation. Also in order to reduce network parameters, 3D Res10 can use residual blocks with fewer convolutional layers, and can eliminate the channel number expansion technology used in bottleneck residual blocks.
  • global average pooling can be used to average the pixel values in the space-time range to obtain a 512-dimensional video space-time feature vector.
  • the kernel size of the global average pooling can be, for example, It is 2*7*7.
  • the 2D Res50 can be different from the general convolutional neural network. It does not use a simple global pooling operation to summarize the total feature map, but can use two-level multi-receptive field pooling to summarize the total feature map.
  • FIG. 13 is a schematic diagram of the operation process of the two-level multi-receptive field pooling operation provided by an embodiment of the application.
  • the first-level local pooling operation three maximum pooling with different totalization ranges (pooling cores) are used to totalize the pixels on the 2D feature map (three
  • the maximum pooling core size can be, for example, 7*7, 4*4, and 2*2, respectively. Pooling cores of different sizes make the aggregated features have different receptive fields, making the features sensitive to targets of different scales. Broaden the scope of recognition of target categories.
  • three sets of 2D pooling feature maps with greatly reduced sizes are obtained. These 2D pooling feature maps contain video features of different scales. Then, the three sets of feature maps are respectively subjected to the second level of global sum-pooling operation.
  • the feature maps on each channel obtain a feature by summing the pixel values, and the three sets of feature maps are summed into Three feature vectors containing video spatial information of different scales. Finally, perform a pixel-by-pixel addition operation on the three feature vectors to obtain a 1024-dimensional 2D feature vector (as shown by the bottom circle in Figure 13).
  • N 15 as mentioned above
  • frames of the video will obtain N 1024-dimensional 2D feature vectors, and the sum of the feature vectors of these frame pictures can be averaged to obtain the 2D feature vector representing the spatial information of the entire video. (That is, the 1024-dimensional 2D feature vector shown in Figure 10).
  • the action recognition method extracts frequency domain information from the decoded image in the action recognition part based on the sequence of pictures, and splices it into the image, increasing the richness of information extracted by the feature network
  • the information is pooled in different ways and merged into a one-dimensional vector to be recognized, using the ECO network structure, and classifying the network with more tags, providing a higher-precision recognition .
  • the action recognition part based on the compressed video stream has filtered out a large number of non-target action videos, the amount of video input to the action recognition part based on sequence pictures, which requires a higher computing power, is small. Combining the two parts, it is guaranteed In the case of motion recognition accuracy, all video detection tasks are completed with very small computing power requirements, which effectively improves recognition efficiency, and achieves a good balance of computing power requirements, recall rate and recognition accuracy.
  • FIG. 14 is a structural block diagram of an action recognition device provided by an embodiment of the application.
  • the device can be implemented by software and/or hardware, and generally can be integrated in a computer device, and can perform action recognition by executing an action recognition method. As shown in Figure 14, the device includes:
  • the video grouping module 1401 is configured to perform grouping processing on the original compressed video data to obtain grouped video data;
  • the target grouped video determining module 1402 is configured to input the grouped video data into a first preset model, and determine the target grouped video data containing actions according to the output result of the first preset model;
  • the video decoding module 1403 is configured to decode the target packet video data to obtain the packet video data to be identified;
  • the action type identification module 1404 is configured to input the packet video data to be identified into a second preset model, and determine the action type contained in the packet video data to be identified according to the output result of the second preset model .
  • the action recognition device Before decompressing the compressed video, the action recognition device provided in the embodiment of the present application first uses the first preset model to roughly filter out the video clips that contain actions, and then uses the second preset model to accurately identify the types of actions contained. On the premise of ensuring recognition accuracy, the amount of calculation is effectively reduced and the efficiency of action recognition is improved.
  • FIG. 15 is a structural block diagram of a computer device provided by an embodiment of this application.
  • the computer device 1500 includes a memory 1501, a processor 1502, and a computer program that is stored on the memory 1501 and can run on the processor 1502.
  • the processor 1502 implements the action recognition method provided in the embodiment of the present application when the computer program is executed.
  • the embodiments of the present application also provide a storage medium containing computer-executable instructions, which are used to execute the action recognition method provided in the embodiments of the present application when the computer-executable instructions are executed by a computer processor.
  • the action recognition apparatus, equipment, and storage medium provided in the foregoing embodiments can execute the action recognition method provided in any embodiment of the present application, and have corresponding functional modules for executing the method.
  • the action recognition method provided in any embodiment of this application can execute the action recognition method provided in any embodiment of this application, and have corresponding functional modules for executing the method.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • General Physics & Mathematics (AREA)
  • Computing Systems (AREA)
  • Evolutionary Computation (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Software Systems (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Artificial Intelligence (AREA)
  • Databases & Information Systems (AREA)
  • Medical Informatics (AREA)
  • Signal Processing (AREA)
  • Computational Linguistics (AREA)
  • Biophysics (AREA)
  • General Engineering & Computer Science (AREA)
  • Mathematical Physics (AREA)
  • Biomedical Technology (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Data Mining & Analysis (AREA)
  • Molecular Biology (AREA)
  • Psychiatry (AREA)
  • Social Psychology (AREA)
  • Human Computer Interaction (AREA)
  • Image Analysis (AREA)
  • Television Signal Processing For Recording (AREA)

Abstract

一种动作识别方法、装置、设备及存储介质。其中,该方法包括:对原始压缩视频数据进行分组处理,得到分组视频数据(101),将分组视频数据输入至第一预设模型中,并根据第一预设模型的输出结果确定包含动作的目标分组视频数据(102),对目标分组视频数据进行解码,得到待识别分组视频数据(103),将待识别分组视频数据输入至第二预设模型中,并根据第二预设模型的输出结果确定待识别分组视频数据中包含的动作类型(104)。

Description

动作识别方法、装置、设备及存储介质
本申请要求在2020年5月20日提交中国专利局、申请号为202010431706.9的中国专利申请的优先权,该申请的全部内容通过引用结合在本申请中。
技术领域
本申请实施例涉及计算机视觉应用技术领域,例如涉及动作识别方法、装置、设备及存储介质。
背景技术
基于视频的动作识别,一直是计算机视觉研究的重要领域。视频动作识别的实现主要包括特征抽取与表示,以及特征分类两大部分。经典的如密度轨迹追跟踪等方法,一般为手动设计特征的方法,而近些年来,人们发现深度学习具备强大的特征表示能力,神经网络便逐渐成为动作识别领域的主流方法,相对于手动设计特征的方法,大大提升了动作识别的性能。
目前,神经网络动作识别方案大部分基于从视频中获取的序列图片构建出时序关系,从而对动作进行判断。例如:基于循环神经网络构建图片之间的时序、基于3D卷积提取多个图片时间的时序信息以及基于图片的深度学习技术叠加动作变化的光流信息等。上述方案存在计算量与识别精度不能兼顾的情况,需要改进。
发明内容
以下是对本文详细描述的主题的概述。本概述并非是为了限制权利要求的保护范围。
本申请实施例提供了动作识别方法、装置、设备及存储介质,可以优化相关技术中的针对视频的动作识别方案。
第一方面,本申请实施例提供了一种动作识别方法,该方法包括:
对原始压缩视频数据进行分组处理,得到分组视频数据;
将所述分组视频数据输入至第一预设模型中,并根据所述第一预设模型的输出结果确定包含动作的目标分组视频数据;
对所述目标分组视频数据进行解码,得到待识别分组视频数据;
将所述待识别分组视频数据输入至第二预设模型中,并根据所述第二预设模型的输出结果确定所述待识别分组视频数据中包含的动作类型。
第二方面,本申请实施例提供了一种动作识别装置,该装置包括:
视频分组模块,设置为对原始压缩视频数据进行分组处理,得到分组视频数据;
目标分组视频确定模块,设置为将所述分组视频数据输入至第一预设模型中,并根据所述第一预设模型的输出结果确定包含动作的目标分组视频数据;
视频解码模块,设置为对所述目标分组视频数据进行解码,得到待识别分组视频数据;
动作类型识别模块,设置为将所述待识别分组视频数据输入至第二预设模型中,并根据所述第二预设模型的输出结果确定所述待识别分组视频数据中包含的动作类型。
第三方面,本申请实施例提供了一种计算机设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,所述处理器执行所述计算机程序时实现如本申请实施例提供的动作识别方法。
第四方面,本申请实施例提供了一种计算机可读存储介质,其上存储有计算机程序,该程序被处理器执行时实现如本申请实施例提供的动作识别方法。
附图说明
图1为本申请实施例提供的一种动作识别方法的流程示意图;
图2为本申请实施例提供的一种压缩视频中帧排列示意图;
图3为本申请实施例提供的一种特征变换操作示意图;
图4为本申请实施例提供的另一种动作识别方法的流程示意图;
图5为本申请实施例提供的一种基于压缩视频流的动作识别过程示意图;
图6为本申请实施例提供的一种第一2D残差网络结构示意图;
图7为本申请实施例提供的一种动作标签示意图;
图8为本申请实施例提供的又一种动作识别方法的流程示意图;
图9为本申请实施例提供的基于短视频的动作识别应用场景示意图;
图10为本申请实施例提供的一种基于序列图片的动作识别过程示意图;
图11为本申请实施例提供的一种第二2D残差网络结构示意图;
图12为本申请实施例提供的一种3D残差网络结构示意图;
图13为本申请实施例提供的两级多感受野池化操作的运算过程示意图;
图14为本申请实施例提供的一种动作识别装置的结构框图;
图15为本申请实施例提供的一种计算机设备的结构框图。
具体实施方式
下面结合附图和实施例对本申请作详细说明。可以理解的是,此处所描述的示例实施例仅仅用于解释本申请,而非对本申请的限定。另外还需要说明的是,为了便于描述,附图中仅示出了与本申请相关的部分而非全部结构。此外,在不冲突的情况下,本申请中的实施例及实施例中的特征可以相互组合。
本申请实施例中的动作识别方案可应用于各种针对视频的动作识别场景,如短视频审核场景、视频监控场景、实时通话识别场景以及机器人视觉识别场景等等。其中,视频可以是视频文件,也可以是视频流。
目前,神经网络动作识别方案大部分基于从视频中获取的序列图片构建出时序关系,从而对动作进行判断。例如:基于循环神经网络(Recurrent Neural Network,RNN)或长短期记忆网络(LSTM,Long Short-Term Memory)等构建图片之间的时序、基于3D卷积提取多个图片时间的时序信息以及基于图片的深度学习技术叠加动作变化的光流信息。
以短视频审核场景为例,基于序列图片进行动作识别的方法至少存在以下两个不足:第一,这些技术方案对一个短视频做出判断都需要大量的计算资源,重度依赖机器中央处理器(Central Processing Unit,CPU)以及图像处理器(Graphics Processing Unit,GPU)的计算力,而且从压缩的短视频中提取出图片需要经过解码,如decode解码(将压缩视频解码成图片的技术),而对短时视频进行全时长段解码本身又需要大量的CPU和GPU算力,因此基于图片的动作识别方案对机器算力需求高,且随着从短视频中取帧的间隔的减小,计算资源的需求与短视频时长成线性增长关系;第二,基于循环神经网络和光流的技术方案,机器审核精度较低,为保证对目标动作的召回率,势必增加机审推送比,导致人审阶段对人力需求变大,从而增大了审核成本。由此可见,上述方案存在计算量与识别精度不能兼顾的情况,需要改进,且除短视频审核场景外的其他类似应用场景同理也存在上述情况。
图1为本申请实施例提供的一种动作识别方法的流程示意图,该方法可以 由动作识别装置执行,其中该装置可由软件和/或硬件实现,一般可集成在计算机设备中。其中,计算机设备例如可以是服务器,也可以是手机等设备,也可由两种或两种以上的设备分别执行部分步骤。如图1所示,该方法包括:
步骤101、对原始压缩视频数据进行分组处理,得到分组视频数据。
示例性的,视频数据中包含了大量的图像以及声音信息,在传输或存储时,通常需要对视频数据进行压缩编码,得到压缩视频数据,本申请实施例这里称为原始压缩视频数据。编码标准以及编码参数等不做限定,例如可以是H264、H265以及MPEG-4等等。
示例性的,可以基于预设分组规则对原始压缩视频数据进行区间划分,得到区间压缩视频(segment,以下可简称片段),将区间压缩视频作为分组视频数据,或从区间压缩视频中选取部分数据作为分组视频数据。其中,预设分组规则中例如可包含区间划分时间间隔,即被区间划分后的每个区间压缩视频对应的时长,区间划分时间间隔可以是恒定的,也可以是变化的,不做限定。以短视频审核场景为例,区间划分时间间隔例如可以是5秒钟。另外,预设分组规则中例如也可以包含区间数量,数值不做限定。
需要说明的是,虽然本文将区间压缩视频简称为片段,但仅为了描述方便,上述分组处理及区间划分可以不涉及对原始压缩视频进行切分或切割的操作,避免引入额外的计算力和存储,保证工程开发效率。分组处理及区间划分可以利用时间戳来完成,也即可以由起止时间戳来限定分组视频数据或区间压缩视频的区间范围。本步骤可理解为从原始压缩视频数据中提取到不同时间段的数据分别输入至第一预设模型。以区间划分时间间隔为5秒为例,原始压缩视频中0~5s对应的数据为一段分组视频数据,5~10s对应的数据为一段分组视频数据,2段数据是分别进入第一预设模型的。
步骤102、将分组视频数据输入至第一预设模型中,并根据所述第一预设模型的输出结果确定包含动作的目标分组视频数据。
示例性的,对原始压缩视频数据进行分组处理时,可以每得到一段分组视频数据就输入至第一预设模型中,提高计算效率;也可以全部分组完成后,再依次或并行将各段分组视频数据输入至第一预设模型中,确保在分组处理准确完成的情况下再进行粗粒度的动作识别。
示例性的,第一预设模型可以是预先训练的神经网络模型,在需要使用时直接进行加载,该模型主要设置为识别分组视频数据中是否包含动作,并不关 心是哪个动作,可以设置动作标签为二分类,如“是”和“否”,可在第一预设模型的训练样本中标记二分类标签。这样,根据第一预设模型的输出结果就可以筛选出哪些片段中包含动作,将对应的分组视频数据确定为目标分组视频数据。由于不需要识别动作类型,也即仅进行粗粒度的识别,因此第一预设模型的计算量较小,且在未解码的情况下进行识别,可以节省大量的解码算力,在排除大量不包含动作的视频片段的同时保证包含动作的视频片段被保留,用于后续的识别。
其中,第一预设模型的网络结构以及相关参数等本申请实施例不做限定,可根据实际需求设置,例如可以为轻量级的模型。
步骤103、对所述目标分组视频数据进行解码,得到待识别分组视频数据。
示例性的,可参考原始压缩视频的编码标准等因素选择适当的解码方式,不做限定。在对目标分组视频数据进行解码后,可以将得到的视频图像作为待识别分组视频数据(解码后的视频图像一般按照时间顺序排列,即序列图像,此时待识别分组视频数据中可包含视频图像的时序信息),也可以在此基础上提取其他信息一并作为待识别分组视频数据,这里的其他信息例如可以是频域信息等。
步骤104、将所述待识别分组视频数据输入至第二预设模型中,并根据所述第二预设模型的输出结果确定所述待识别分组视频数据中包含的动作类型。
示例性的,第二预设模型可以是预先训练的神经网络模型,在需要使用时直接进行加载,该模型主要设置为识别待识别分组视频数据中包含的动作的类型,也即进行细粒度的识别,可在第二预设模型的训练样本中标记多分类标签,这样,根据第二预设模型的输出结果就可以确定出最终的动作识别结果。经过前述步骤的初筛,能够进入到第二预设模型的数据相比原始压缩视频数据已经大大减少,且数据纯度也远高于原始压缩视频数据,需要识别的视频片段数量的量级不大,所以可以采用解码之后的序列图像进行识别。在一些实施例中,可采用神经网络参数较多的3D卷积网络结构进行时序特征的提取,标签则是采用粒度更细的多分类,标签数量不做限定,例如可以是50个,可根据的识别精度进行调整。
本申请实施例中提供的动作识别方法,对原始压缩视频数据进行分组处理,得到分组视频数据,将分组视频数据输入至第一预设模型中,并根据第一预设模型的输出结果确定包含动作的目标分组视频数据,对目标分组视频数据进行 解码,得到待识别分组视频数据,将待识别分组视频数据输入至第二预设模型中,并根据第二预设模型的输出结果确定待识别分组视频数据中包含的动作类型。通过采用上述技术方案,在对压缩视频进行解压前,先利用第一预设模型粗略筛选出包含动作的视频片段,再利用第二预设模型精确识别包含的动作的类型,可以在保证识别精度的前提下有效减少计算量,提高动作识别效率。
需要说明的是,对于一些如视频审核的应用场景,一般可包括用于视频上传的客户端和视频审核的服务器,由于上述步骤101和步骤102的相关操作计算量较小,可以在客户端完成,也即客户端可以针对即将上传的压缩视频进行初步筛选,将包含动作的目标分组视频数据上传至服务器进行识别和审核。另外,对于一些配置较高的设备来说,也可在客户端本地完成整个识别流程,即步骤101至步骤104的相关操作,实现根据最终识别出来的动作类型来确定是否允许视频的上传等控制。
在一些实施例中,所述对原始压缩视频数据进行分组处理,得到分组视频数据,包括:基于预设分组规则对原始压缩视频数据进行区间划分,得到区间压缩视频;基于预设提取策略提取所述区间压缩视频中的I帧数据和P帧数据,得到分组视频数据,其中,所述P帧数据包含P帧对应的运动矢量信息和/或色素变化残差信息。本申请实施例提供的技术方案,可以快速提取压缩视频数据中可以用于动作识别的特征数据,提高识别效率。
其中,I帧又称关键帧,指压缩视频中包含的图像;P帧又称前向预测编码帧,一般包括压缩视频中参照I帧的变化信息,可包括运动矢量(motion vector,MV)信息和RGB色素变化的残差(RGB Residual Frame,RGBR)信息。一个I帧后面通常存在多个P帧,不同编码方式中,I帧和P帧的分布或包含的内容可能存在差异。示例性的,预设分组规则可如前文所述,例如可包括区间划分时间间隔或分段数量等。预设提取策略例如可包括I帧的提取策略和P帧的提取策略。I帧的提取策略可包括获取I帧的时间间隔,或单位时间内获取的I帧数量,例如,每秒钟获取1个I帧。P帧的提取策略可包括获取一个I帧后面的P帧的数量以及每两个被获取的P帧之间的时间间隔,例如数量为2。那么以一个区间压缩视频为5秒为例,可以获取5个I帧和10个P帧对应的数据,若P帧数据同时包括MV和RGBR,则可获取10个MV和10个RGBR。
在一些实施例中,所述第一预设模型中包含第一2D残差网络、第一拼接层和第一全连接层;所述分组视频数据被输入至所述第一预设模型中后,经由所 述第一2D残差网络得到对应的维度相同的特征图;所述特征图经由所述第一拼接层得到按照帧的先后顺序进行拼接操作后的拼接特征图;所述拼接特征图经由所述第一全连接层得到是否包含动作的分类结果。本申请实施例提供的技术方案,采用比较精简的网络结果来获取较高的识别效率,并保证包含动作的视频片段有较高的召回率。
示例性的,第一2D残差网络可以采用轻量级的ResNet18模型。由于I帧、MV和RGBR在数据特征上并不能直接拼接(concate)在一起,因此可以使用3个ResNet18对I帧、MV和RGBR分别进行单独处理,得出维度相同的特征图,可分别记为I帧特征图、MV特征图和RGBR特征图,MV特征图和RGBR特征图统称P帧特征图。其中,维度相同可以指C*H*W一致,其中,C表示通道(channel)、H表示高度(height)以及W表示宽度(width),*又可表示为×。经过处理后,维度相同的特征图便可经过第一拼接层实现拼接。
在一些实施例中,基于预设提取策略提取所述区间压缩视频中的I帧数据和P帧数据,得到分组视频数据,包括:基于预设提取策略提取所述区间压缩视频中的I帧数据和P帧数据;对所述P帧数据进行累加变换,以使得变换后的P帧数据依赖于前向相邻的I帧;根据所述I帧数据和变换后的P帧数据确定分组视频数据。相应的,所述第一预设模型中还包括位于所述拼接层之前的相加层,所述特征图中对应P帧数据的特征图记为P帧特征图,所述特征图中对应I帧数据的特征图记为I帧特征图;所述P帧特征图和所述I帧特征图经由所述相加层得到在所述I帧特征图基础上经过相加操作后的P帧特征图;所述I帧特征图和所述经过相加操作后的P帧特征图经由所述第一拼接层得到按照帧的先后顺序进行拼接操作后的拼接特征图。本申请实施例提供的技术方案通过累加变换提供与I帧更紧密的信息关联,经过神经网络中的Add(相加)计算,提供更加全面的待识别信息。
以H264为例,图2为本申请实施例提供的一种压缩视频中帧排列示意图,如图2所示,从第二个P帧开始,每一个P帧的MV和RGBR都依赖于前面一帧的P帧(如P2帧依赖于P1帧),为了使P帧与I帧的关联更加紧密,可以对P帧的MV和RGBR做累加的变换,获取到输入神经网络的P帧的MV和RGBR是相对于前面的I帧(如累加变换后,P2帧变为依赖前面的I帧),而不是相对于前一个P帧。上述相加操作可以是参照ResNet中的残差相加的方式直接对MV特征图和RGBR特征图按每个元素(像素)与I帧的输出(经过第一2D残 差网络处理后的I帧特征图)进行相加,之后再按帧的先后顺序对3个特征图进行拼接操作。
在一些实施例中,在所述第一2D残差网络的残差结构前包括特征变换层;所述分组视频数据在进入所述残差结构前,经由所述特征变换层得到经过向上特征变换和/或向下特征变换的分组视频数据。本申请实施例提供的技术方案,在进入残差(residual)结构之前,也即进行卷积操作前,可以通过特征变换(Feature Shift,FS)对特征图进行部分特征的变换(shift)操作,使得一个特征图中包含不同时间点的特征图中的部分特征,这样在进行卷积操作时,特征图中包含了时序信息,可以在基本不增加计算量的前提下有能力处理时序信息的采集和融合,丰富待识别特征图的信息,提升识别准确度,相比于相关技术中的基于3D卷积或光流信息的方案来说,可有效降低计算量。
示例性的,图3为本申请实施例提供的一种特征变换操作示意图。如图3所示,以MV特征图为例,为了便于说明,图3中仅示出了针对3个MV特征图进行特征变换操作的过程,3个MV特征图分别对应不同的时间点,假设将MV特征图划分为4个部分,假设针对第1部分进行向下特征变换(Down Shift),针对第2部分进行向上特征变换(Up Shift),经过变换后的第2个MV特征图则同时包含了3个时间点的MV特征图的部分特征。在实际应用时,划分规则,以及向上特征变换和/或向下特征变换作用区域可根据实际情况进行设置。
在一些实施例中,基于预设提取策略提取所述区间压缩视频中的P帧数据,包括:采用等间隔方式提取所述区间压缩视频中的预设数量的P帧数据;其中,在所述第一预设模型的训练阶段,采用随机间隔方式提取区间压缩视频中的预设数量的P帧数据。本申请实施例提供的技术方案,可以增强第一预设模型的鲁棒性。
图4为本申请实施例提供的另一种动作识别方法的流程示意图,在上述各示例实施例基础上进行细化,例如,该方法包括:
步骤401、基于预设分组规则对原始压缩视频数据进行区间划分,得到区间压缩视频。
在一些实施例中,原始压缩视频可采用预设的视频和图像增强策略进行增强处理,可根据业务的数据需求以及驱动等配置情况来选择增强策略。在模型训练和模型应用时可采用相同的增强策略。
步骤402、基于预设提取策略提取区间压缩视频中的I帧数据和P帧数据。
其中,P帧数据包含P帧对应的运动矢量信息和色素变化残差信息。
步骤403、对P帧数据进行累加变换,以使得变换后的P帧数据依赖于前向相邻的I帧,根据I帧数据和变换后的P帧数据确定分组视频数据。
步骤404、将分组视频数据输入至第一预设模型中,并根据第一预设模型的输出结果确定包含动作的目标分组视频数据。
示例性的,图5为本申请实施例提供的一种基于压缩视频流的动作识别过程示意图,压缩视频被划分为n个片段,以其中的S2为例,将提取到的I帧数据、MV数据和RGBR数据输入至第一预设模型中。第一预设模型中包含第一2D残差网络、相加层、第一拼接层和第一全连接层(Fully connected layer,FC layer),其中,第一2D残差网络为2D Res18,其残差结构前包括特征变换层(FS)。在一些实施例中,在第一预设模型的训练阶段,可根据实际需求选择训练方式,包括损失函数(loss)等,例如损失函数可采用交叉熵,还可以采用其他辅助损失函数来提高模型效果。另外,可使用一些非启发式的优化算法来提高随机梯度下降的收敛速度以及优化性能。
图6为本申请实施例提供的一种第一2D残差网络结构示意图,如图6所示,2D Res18是一个18层的残差神经网络,由4个阶段(stage)共8个使用2D卷积的残差块(block)组成,网络结构较浅,为尽可能使用卷积层充分提取特征,8个2D残差块可均采用3*3的卷积进行,也即卷积核参数为3*3。其中,卷积层一般指一个用于完成局部像素值的加权求和以及非线性激活的网络层,为了减少卷积层运算的通道数进而减少参数量,每个2D残差块都可使用瓶颈(bottleneck)的设计理念,即每个残差块都由3个卷积层组成(卷积核参数分别为1*1、3*3和1*1),首层和尾层分别用于压缩和恢复图像通道。当然,第一2D残差网络结构以及各参数也可根据实际需求进行调整。
如图5所示,I帧数据、MV数据和RGBR数据在进入2D Res18残差结构(也即2D残差块)前,分别经由FS层得到经过向上特征变换和向下特征变换的I帧数据、MV数据和RGBR数据,然后再分别经由残差结构得到得出C*H*W一致的特征图,经由相加层参照ResNet中的残差相加的方式直接对MV特征图和RGBR特征图按每个元素与I帧的输出进行相加,之后再按帧的先后顺序对3个特征图进行concate操作,得到拼接特征图,拼接特征图经由第一全连接层(FC)得到是否包含动作的分类结果。
图7为本申请实施例提供的一种动作标签示意图,图5中最后2个圆圈表 示输出是二分类,设置动作标签为二分类是为了保证提高包含动作的片段的召回,标签粒度为是否包含动作,设计为只有“是”或者“否”的二分类级别,即不关心包含的是哪个细分类动作。如图7所示,在模型训练过程中,可将训练用的短视频切割分为6个segments,其中segment S2包含动作A1,segment S4包含动作A2,A1和A2为2个不同的动作,但这2个位置对应的标签是一样的,都设置为1,也即不区分A1和A2动作的不同。因而,在实际应用时,第一预设模型输出的结果也同样为“是”或“否”,实现压缩视频中动作的粗粒度识别。
步骤405、对目标分组视频数据进行解码,得到待识别分组视频数据。
步骤406、将待识别分组视频数据输入至第二预设模型中,并根据第二预设模型的输出结果确定待识别分组视频数据中包含的动作类型。
本申请实施例提供的动作识别方法,先基于压缩视频进行动作识别,在无需解压视频的情况下,提取I帧、P帧的MV和RGBR信息,并对MV、RGBR做变化以增加对I帧的信息依赖,实现了以较少计算力需求处理大量的时长不定的视频,且使用无需计算力要求的FS来增加模型对时序信息的提取,增强模型能力的情况下没有导致计算效率的降低,重新设计动作的标签粒度,以保证召回的目标,让模型处理简单的二分类问题,可以提高召回率,随后对初步筛选出来的包含动作的视频片段做精确识别,在保证识别精度的前提下有效减少计算量,提高动作识别效率。
在一些实施例中,所述对所述目标分组视频数据进行解码,得到待识别分组视频数据,包括:对所述目标分组视频数据进行解码,得到待识别分段视频图像;获取所述待识别分段视频图像中的频域信息,根据所述频域信息生成对应的频域图;将所述待识别分段视频图像和对应的频域图作为待识别分组视频数据。本申请实施例提供的技术方案,可以丰富图像信息,提高第二预设模型的准确度。
在一些实施例中,所述第二预设模型包括基于用于在线视频理解的高效卷积网络(Efficient Convolutional Network for Online Video,ECO)架构的模型。ECO架构可理解为视频特征提取器,提供了一种视频特征获取的架构设计,里面包含了2D特征提取网络和3D特征提取网络,该架构可以在得到较好性能的同时提高速度,本申请实施例可以在ECO架构基础上进行改进和设计,得到第二预设模型。
在一些实施例中,所述第二预设模型包括第二拼接层、第二2D残差网络、 3D残差网络、第三拼接层和第二全连接层;所述待识别分组视频数据被输入至第二预设模型中后,经由所述第二拼接层得到对所述待识别分段视频图像和对应的频域图进行拼接后的拼接图像数据;所述拼接图像数据经由所述第二2D残差网络得到2D特征图;将所述第二2D残差网络的中间层输出结果作为所述3D残差网络的输入,并经由所述3D残差网络得到3D特征图;所述2D特征图和所述3D特征图经由所述第三拼接层得到拼接后的待识别特征图;所述待识别特征图经由所述第二全连接层得到对应的动作类型标签。本申请实施例提供的技术方案,在ECO架构基础上,将第二2D残差网络的中间层输出结果作为3D残差网络的输入,实现网络结构的复用,提升模型速度。
示例性的,待识别分组视频数据首先进入模型的2D卷积部分提取每张图像的特征,然后进入3D卷积部分提取动作的时序信息,最后输出多分类的结果。其中,第二2D残差网络和3D残差网络的网络结构和相关参数可根据实际需求设置,在一些实施例中,第二2D残差网络为2D Res50,3D残差网络为3D Res10。
示例性的,所述第二预设模型还包括第一池化层和第二池化层;所述2D特征图在被输入至所述第三拼接层之前,经由所述第一池化层得到对应的包含第一元素数量的一维2D特征向量;所述3D特征图在被输入至所述第三拼接层之前,经由所述第二池化层得到对应的包含第二元素数量的一维3D特征向量。相应的,所述2D特征图和所述3D特征图经由所述第三拼接层得到拼接后的待识别特征图,包括:所述一维2D特征向量和所述一维3D特征向量经由所述第三拼接层得到拼接后的待识别向量。本申请实施例提供的技术方案,特征图在经过卷积层进行特征提取以后仍然有较大的尺寸,此时若直接把特征图展平,可能会使得特征向量维度过高,因此,2D和3D残差网络在完成一系列的特征提取操作以后,可使用池化操作把特征图直接归总为一维特征向量,以降低特征向量的维度。其中,第一池化层和第二池化层可根据实际需求设计,例如可以是全局平均池化(Global average Pooling,GAP)或最大池化等。第一元素数量和第二元素数量也可以自由设置,例如可根据特征图的通道数来设置。
示例性的,所述第一池化层包括多感受野池化层。感受野一般指特征值所覆盖的图像或视频范围,用来表示网络内部的不同神经元对原图像的感受范围的大小,或者说,卷积神经网络每一层输出的特征图上的像素点在原始图像上映射的区域大小。感受野的值越大表示其能接触到的原始图像范围就越大,也意味着它可能蕴含更为全局、语义层次更高的特征;相反,值越小则表示其所 包含的特征越趋向局部和细节。因此感受野的值可以用来大致判断每一层的抽象层次。采用多感受野池化层的好处在于,使得特征可以对不同尺度的目标敏感,拓宽了对动作类别的识别范围。多感受野实现方式可根据实际需求设置。
在一些实施例中,所述第一池化层包括一级局部池化层、二级全局池化层和向量合并层,所述一级局部池化层中包含至少两个尺寸不相同的池化核;所述2D特征图经由所述一级局部池化层得到对应的至少两组不同尺度的2D池化特征图;所述至少两组不同尺度的2D池化特征图经由所述二级全局池化层得到至少两组特征向量;所述至少两组特征向量经由所述向量合并层得到对应的包含第一元素数量的一维2D特征向量。本申请实施例提供的技术方案,采用两级多感受野池化来归总特征图,利用不同尺寸的池化核进行多尺度池化,使得归总的特征拥有不同的感受野,同时也使得特征图尺寸大大减少,提高识别效率,随后利用二级全局池化来对特征进行归总,并利用向量合并层得到一维2D特征向量。
图8为本申请实施例提供的又一种动作识别方法的流程示意图,在上述各示例实施例基础上进行细化,以短视频审核场景为例,图9为本申请实施例提供的基于短视频的动作识别应用场景示意图,如图9所示,用户上传短视频后,先提取压缩视频流信息,并利用预先训练构造的基于压缩视频的动作视频模型(第一预设模型)识别出包含动作的目标片段,其他片段因不包含动作而被大量筛除,随后,对目标片段进行解码并提取图片时域频域信息,利用预先训练构造的基于解码图片的动作识别模型(第二预设模型)进行基于图片序列的更细粒度的动作识别,得到每个目标片段对应的动作类型,进而判断是否为目标动作,若不是目标动作,则也被筛除。
示例性的,该方法可包括:
步骤801、基于预设分组规则对原始压缩视频数据进行区间划分,得到区间压缩视频。
其中,原始压缩视频数据为用户上传的短视频的视频流数据,每个区间压缩视频的长度可以是5秒。
步骤802、基于预设提取策略提取区间压缩视频中的I帧数据和P帧数据。
其中,P帧数据包含P帧对应的MV和RGBR,每秒获取1个I帧以及其后对应的2个P帧,也即一个5秒的片段可以获取到5个I帧、10个MV和10个RGBR。
步骤803、对P帧数据进行累加变换,以使得变换后的P帧数据依赖于前向相邻的I帧,根据I帧数据和变换后的P帧数据确定分组视频数据。
步骤804、将分组视频数据输入至第一预设模型中,并根据第一预设模型的输出结果确定包含动作的目标分组视频数据。
步骤805、对目标分组视频数据进行解码,得到待识别分段视频图像。
示例性的,从目标分组视频数据中解码出图像(图片),按时间顺序排成序列,得到序列图片。在一些实施例中,为了减少待识别数据量,可以按照预设获取策略从序列图片中获取设定数量的图像,作为待识别分段视频图像。预设获取策略例如等间隔获取,设定数量例如为15。
步骤806、获取待识别分段视频图像中的频域信息,根据频域信息生成对应的频域图。
其中,频域信息的采集方式不做限定,可针对序列图片中的每个图片采集频域信息,并生成与序列图片一一对应的频域(Frequency Domain,FD)图。
步骤807、将待识别分段视频图像和对应的频域图作为待识别分组视频数据,输入至第二预设模型中,并根据第二预设模型的输出结果确定待识别分组视频数据中包含的动作类型。
示例性的,一个片段中的序列图片和对应的频域图将对应多分类标签中的一个标签,也即,第二预设模型的输出结果中,每个片段对应一个标签,若一个片段中包含多个动作,则模型会将主要动作对应的类型作为该片段的标签。例如,5秒钟的A片段中有4秒钟存在动作,3秒钟的动作为挥手,1秒钟的动作为踢腿,则该片段A对应的动作标签为挥手。
图10本申请实施例提供的一种基于序列图片的动作识别过程示意图,其中n的数值一般与图5中的n的数值不相同,以片段S2为例,经过解码(Decode)后,得到序列图片(Image),提取频域信息后,生成对应的频域图(FD)。第二预设模型可基于ECO架构设计,其中可包含第二拼接层、第二2D残差网络、3D残差网络、第一池化层、第二池化层、第三拼接层和第二全连接层。在一些实施例中,在第二预设模型的训练阶段,可根据实际需求选择训练方式,包括损失函数等,例如损失函数可采用交叉熵,还可以采用其他辅助损失函数来提高模型效果。另外,可使用一些非启发式的优化算法来提高随机梯度下降的收敛速度以及优化性能。如图10所示,解码后的序列图片和对应的频域图经过第二预设模型中第二拼接层后得到拼接图像数据,在一些实施例中,随后拼接图 像数据还可经过卷积层(conv)和最大池化层(maxpool)的处理后,再输入到第二2D残差网络(2D Res50)中。2D Res50的输出经过第一池化层(多感受野池化层)后得到1024维2D特征向量(可理解为包含1024个元素的一维的行向量或列向量)。2D Res50的中间输出结果将作为3D残差网络(3D Res10)的输入,经过第二池化层(GAP)后得到512维3D特征向量。1024维2D特征向量和512维3D特征向量经过第三拼接层后,得到待识别特征图,也即待识别特征向量,最后经由第二全连接层得到最终的动作类型标签。
图11为本申请实施例提供的一种第二2D残差网络结构示意图,如图所示,第二2D残差网络为2D Res50,是一个50层的残差神经网络,由4个stage共16个使用2D卷积的残差块组成。为了减少卷积层运算的通道数进而减少参数量,每个残差块都可使用bottleneck的设计理念,即每个残差块都由3个卷积层组成,其中进出口的两个1*1卷积分别用来压缩和还原特征图的通道数。另外,因为每经过一个stage,需要把特征图的尺寸缩小至四分之一、通道扩大为两倍,所以在每个stage的入口处都可使用2D投影残差块,这种残差块在旁路中增加了一个1*1卷积层,用来保证做逐像素相加操作时,特征图的尺寸和通道数保持一致。同理,只在每个stage的入口处使用2D投影残差块也可以减少网络参数。
图12为本申请实施例提供的一种3D残差网络结构示意图。每个片段取N个视频帧(如前文所述的从序列图像中获取的15个图像与对应的频域图经过拼接操作后得到的帧图片)通过2D Res50获得的中间层特征图组,例如来自stage2-block4的特征图组可被组装为三维的视频张量,张量形状为(c,f,h,w),其中,c是帧图片的通道数,f是视频帧数,h、w分别指帧图片的高和宽,视频张量被输入至3D Res10,以进行整个视频时空特征的提取。如图12所示,3D Res10仅由3个stage共5个残差块组成,所有卷积层都使用三维的卷积核,卷积过程中,时间维度上的信息也将一起参与计算。同样为了减少网络参数,3D Res10可使用卷积层数更少的残差块,并且可去除在bottleneck残差块中使用的通道数扩张技术。
示例性的,3D Res10在完成卷积操作后,可通过全局平均池化来对时空范围内的像素值求平均值,得到一个512维的视频时空特征向量,全局平均池化的核尺寸例如可以是2*7*7。而2D Res50可与一般的卷积神经网络不同,并不使用简单的全局池化操作来归总特征图,而可采用两级多感受野池化来归总特 征图。图13为本申请实施例提供的两级多感受野池化操作的运算过程示意图。如图13所示,首先,在第一级的局部池化操作中,三个归总范围(池化核)不同的最大池化被用于对2D特征图上的像素进行归总(三个最大池化的核尺寸例如可以分别为7*7、4*4和2*2),不同大小的池化核使得归总的特征拥有不同的感受野,使得特征可以对不同尺度的目标敏感,拓宽了对目标类别的识别范围。经过了第一级的多尺度局部池化操作以后,得到三组尺寸大大减少的2D池化特征图,这些2D池化特征图包含了不同尺度的视频特征。然后,三组特征图分别进行第二级的求和全局池化(global sum-pooling)操作,每个通道上的特征图都通过求和像素值得到一个特征,三组特征图被归总为三个包含不同尺度视频空间信息的特征向量。最后对三个特征向量执行逐像素相加操作,获得一个1024维的2D特征向量(如图13中最底部的圆圈所示)。经过前述过程,视频的N(如前文所述15)个帧将获得N个1024维的2D特征向量,对这些帧图片的特征向量求和取平均即可获得代表整个视频空间信息的2D特征向量(即图10所示的1024维2D特征向量)。
本申请实施例提供的动作识别方法,在上述实施例基础上,在基于序列图片的动作识别部分,从解码之后的图像中提取出频域信息拼接到图像中,增加特征网络提取到的信息丰富度,采用了ECO网络结构,提取2D和3D信息后,针对信息进行不同方式的池化,并融合为一维的待识别向量,对网络进行更多标签的分类,提供了精度更高的识别。因为基于压缩视频流的动作识别部分已经过滤掉大量的非目标动作视频,因此输入到对计算力需求较高的基于序列图片的动作识别部分的视频量较小,综合两部分来说,在保证动作识别精度的情况下,以非常小的计算力需求完成了所有视频的检测任务,有效提高识别效率,做到了很好地平衡计算力需求、召回率和识别精度。
图14为本申请实施例提供的一种动作识别装置的结构框图,该装置可由软件和/或硬件实现,一般可集成在计算机设备中,可通过执行动作识别方法来进行动作识别。如图14所示,该装置包括:
视频分组模块1401,设置为对原始压缩视频数据进行分组处理,得到分组视频数据;
目标分组视频确定模块1402,设置为将所述分组视频数据输入至第一预设模型中,并根据所述第一预设模型的输出结果确定包含动作的目标分组视频数据;
视频解码模块1403,设置为对所述目标分组视频数据进行解码,得到待识别分组视频数据;
动作类型识别模块1404,设置为将所述待识别分组视频数据输入至第二预设模型中,并根据所述第二预设模型的输出结果确定所述待识别分组视频数据中包含的动作类型。
本申请实施例提供的动作识别装置,在对压缩视频进行解压前,先利用第一预设模型粗略筛选出包含动作的视频片段,再利用第二预设模型精确识别包含的动作的类型,可以在保证识别精度的前提下有效减少计算量,提高动作识别效率。
本申请实施例提供了一种计算机设备,该计算机设备中可集成本申请实施例提供的动作识别装置。图15为本申请实施例提供的一种计算机设备的结构框图。计算机设备1500包括存储器1501、处理器1502及存储在存储器1501上并可在处理器1502上运行的计算机程序,所述处理器1502执行所述计算机程序时实现本申请实施例提供的动作识别方法。
本申请实施例还提供一种包含计算机可执行指令的存储介质,所述计算机可执行指令在由计算机处理器执行时用于执行本申请实施例提供的动作识别方法。
上述实施例中提供的动作识别装置、设备以及存储介质可执行本申请任意实施例所提供的动作识别方法,具备执行该方法相应的功能模块。未在上述实施例中详尽描述的技术细节,可参见本申请任意实施例所提供的动作识别方法。
注意,上述仅为本申请的实例实施例。本领域技术人员会理解,本申请不限于这里所述的特定实施例,对本领域技术人员来说能够进行各种明显的变化、重新调整和替代而不会脱离本申请的保护范围。因此,虽然通过以上实施例对本申请进行了较为详细的说明,但是本申请不仅仅限于以上实施例,在不脱离本申请构思的情况下,还可以包括更多其他等效实施例,而本申请的范围由权利要求范围决定。

Claims (15)

  1. 一种动作识别方法,包括:
    对原始压缩视频数据进行分组处理,得到分组视频数据;
    将所述分组视频数据输入至第一预设模型中,并根据所述第一预设模型的输出结果确定包含动作的目标分组视频数据;
    对所述目标分组视频数据进行解码,得到待识别分组视频数据;
    将所述待识别分组视频数据输入至第二预设模型中,并根据所述第二预设模型的输出结果确定所述待识别分组视频数据中包含的动作类型。
  2. 根据权利要求1所述的方法,其中,所述对原始压缩视频数据进行分组处理,得到分组视频数据,包括:
    基于预设分组规则对原始压缩视频数据进行区间划分,得到区间压缩视频;
    基于预设提取策略提取所述区间压缩视频中的关键帧I帧数据和前向预测编码帧P帧数据,得到分组视频数据,其中,所述P帧数据包含P帧对应的以下至少一种信息:运动矢量信息,色素变化残差信息。
  3. 根据权利要求2所述的方法,其中,所述第一预设模型中包含第一2D残差网络、第一拼接层和第一全连接层;
    所述分组视频数据被输入至所述第一预设模型中后,经由所述第一2D残差网络得到对应的维度相同的特征图;
    所述特征图经由所述第一拼接层得到按照帧的先后顺序进行拼接操作后的拼接特征图;
    所述拼接特征图经由所述第一全连接层得到是否包含动作的分类结果。
  4. 根据权利要求3所述的方法,其中,基于预设提取策略提取所述区间压缩视频中的I帧数据和P帧数据,得到分组视频数据,包括:
    基于预设提取策略提取所述区间压缩视频中的I帧数据和P帧数据;
    对所述P帧数据进行累加变换,以使得变换后的P帧数据依赖于前向相邻的所述I帧数据;
    根据所述I帧数据和变换后的P帧数据确定分组视频数据;
    所述第一预设模型中还包括位于所述拼接层之前的相加层,所述特征图中对应P帧数据的特征图记为P帧特征图,所述特征图中对应I帧数据的特征图记为I帧特征图;
    所述P帧特征图和所述I帧特征图经由所述相加层得到在所述I帧特征图基 础上经过相加操作后的P帧特征图;
    所述I帧特征图和所述经过相加操作后的P帧特征图经由所述第一拼接层得到按照帧的先后顺序进行拼接操作后的拼接特征图。
  5. 根据权利要求3所述的方法,其中,在所述第一2D残差网络的残差结构前包括特征变换层;
    所述分组视频数据在进入所述残差结构前,经由所述特征变换层得到以下至少一种分组视频数据:经过向上特征变换的分组视频数据,经过向下特征变换的分组视频数据。
  6. 根据权利要求2所述的方法,其中,基于预设提取策略提取所述区间压缩视频中的P帧数据,包括:
    采用等间隔方式提取所述区间压缩视频中的预设数量的P帧数据;
    其中,在所述第一预设模型的训练阶段,采用随机间隔方式提取区间压缩视频中的预设数量的P帧数据。
  7. 根据权利要求1所述的方法,其中,所述对所述目标分组视频数据进行解码,得到待识别分组视频数据,包括:
    对所述目标分组视频数据进行解码,得到待识别分段视频图像;
    获取所述待识别分段视频图像中的频域信息,根据所述频域信息生成对应的频域图;
    将所述待识别分段视频图像和对应的频域图作为待识别分组视频数据。
  8. 根据权利要求7所述的方法,其中,所述第二预设模型包括基于用于在线视频理解的高效卷积网络ECO架构的模型。
  9. 根据权利要求8所述的方法,其中,所述第二预设模型包括第二拼接层、第二2D残差网络、3D残差网络、第三拼接层和第二全连接层;
    所述待识别分组视频数据被输入至第二预设模型中后,经由所述第二拼接层得到对所述待识别分段视频图像和对应的频域图进行拼接后的拼接图像数据;
    所述拼接图像数据经由所述第二2D残差网络得到2D特征图;
    将所述第二2D残差网络的中间层输出结果作为所述3D残差网络的输入,并经由所述3D残差网络得到3D特征图;
    所述2D特征图和所述3D特征图经由所述第三拼接层得到拼接后的待识别特征图;
    所述待识别特征图经由所述第二全连接层得到对应的动作类型标签。
  10. 根据权利要求9所述的方法,其中,所述第二预设模型还包括第一池化层和第二池化层;
    所述2D特征图在被输入至所述第三拼接层之前,经由所述第一池化层得到对应的包含第一元素数量的一维2D特征向量;
    所述3D特征图在被输入至所述第三拼接层之前,经由所述第二池化层得到对应的包含第二元素数量的一维3D特征向量;
    所述2D特征图和所述3D特征图经由所述第三拼接层得到拼接后的待识别特征图,包括:
    所述一维2D特征向量和所述一维3D特征向量经由所述第三拼接层得到拼接后的待识别向量。
  11. 根据权利要求10所述的方法,其中,所述第一池化层包括多感受野池化层。
  12. 根据权利要求11所述的方法,其中,所述第一池化层包括一级局部池化层、二级全局池化层和向量合并层,所述一级局部池化层中包含至少两个尺寸不相同的池化核;
    所述2D特征图经由所述一级局部池化层得到对应的至少两组不同尺度的2D池化特征图;
    所述至少两组不同尺度的2D池化特征图经由所述二级全局池化层得到至少两组特征向量;
    所述至少两组特征向量经由所述向量合并层得到对应的包含第一元素数量的一维2D特征向量。
  13. 一种动作识别装置,包括:
    视频分组模块,设置为对原始压缩视频数据进行分组处理,得到分组视频数据;
    目标分组视频确定模块,设置为将所述分组视频数据输入至第一预设模型中,并根据所述第一预设模型的输出结果确定包含动作的目标分组视频数据;
    视频解码模块,设置为对所述目标分组视频数据进行解码,得到待识别分组视频数据;
    动作类型识别模块,设置为将所述待识别分组视频数据输入至第二预设模型中,并根据所述第二预设模型的输出结果确定所述待识别分组视频数据中包含的动作类型。
  14. 一种计算机设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,所述处理器执行所述计算机程序时实现如权利要求1-12任一项所述的方法。
  15. 一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现如权利要求1-12中任一所述的方法。
PCT/CN2021/085386 2020-05-20 2021-04-02 动作识别方法、装置、设备及存储介质 Ceased WO2021232969A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
EP21808837.5A EP4156017B1 (en) 2020-05-20 2021-04-02 Action recognition method and apparatus, and device and storage medium
US17/999,284 US12412426B2 (en) 2020-05-20 2021-04-02 Action recognition method and apparatus, and device and storage medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202010431706.9 2020-05-20
CN202010431706.9A CN111598026B (zh) 2020-05-20 2020-05-20 动作识别方法、装置、设备及存储介质

Publications (1)

Publication Number Publication Date
WO2021232969A1 true WO2021232969A1 (zh) 2021-11-25

Family

ID=72183951

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2021/085386 Ceased WO2021232969A1 (zh) 2020-05-20 2021-04-02 动作识别方法、装置、设备及存储介质

Country Status (4)

Country Link
US (1) US12412426B2 (zh)
EP (1) EP4156017B1 (zh)
CN (1) CN111598026B (zh)
WO (1) WO2021232969A1 (zh)

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114827739A (zh) * 2022-06-06 2022-07-29 百果园技术(新加坡)有限公司 一种直播回放视频生成方法、装置、设备及存储介质
CN115909479A (zh) * 2022-10-20 2023-04-04 中国科学院自动化研究所 人体行为识别方法、装置、电子设备及可读存储介质
CN116129936A (zh) * 2023-02-02 2023-05-16 百果园技术(新加坡)有限公司 一种动作识别方法、装置、设备、存储介质及产品
CN116246259A (zh) * 2022-12-31 2023-06-09 南京行者易智能交通科技有限公司 一种基于3d卷积神经网络的公交安全驾驶行为识别方法
CN117036789A (zh) * 2023-07-24 2023-11-10 西安电子科技大学 基于动态推理的全局与局部特征融合的骨架行为识别方法
CN117975376A (zh) * 2024-04-02 2024-05-03 湖南大学 基于深度分级融合残差网络的矿山作业安全检测方法

Families Citing this family (20)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111598026B (zh) 2020-05-20 2023-05-30 广州市百果园信息技术有限公司 动作识别方法、装置、设备及存储介质
CN112183359B (zh) * 2020-09-29 2024-05-14 中国科学院深圳先进技术研究院 视频中的暴力内容检测方法、装置及设备
CN112337082B (zh) * 2020-10-20 2021-09-10 深圳市杰尔斯展示股份有限公司 一种ar沉浸式虚拟视觉感知交互系统和方法
CN112396161B (zh) * 2020-11-11 2022-09-06 中国科学技术大学 基于卷积神经网络的岩性剖面图构建方法、系统及设备
CN112507920B (zh) * 2020-12-16 2023-01-24 重庆交通大学 一种基于时间位移和注意力机制的考试异常行为识别方法
CN112712042B (zh) * 2021-01-04 2022-04-29 电子科技大学 嵌入关键帧提取的行人重识别端到端网络架构
CN112686193B (zh) * 2021-01-06 2024-02-06 东北大学 基于压缩视频的动作识别方法、装置及计算机设备
CN112749666B (zh) * 2021-01-15 2024-06-04 百果园技术(新加坡)有限公司 一种动作识别模型的训练及动作识别方法与相关装置
CN116263618A (zh) * 2021-12-13 2023-06-16 成都拟合未来科技有限公司 一种健身动作识别方法及系统
CN114333065B (zh) * 2021-12-31 2025-03-04 济南博观智能科技有限公司 一种应用于监控视频的行为识别方法、系统及相关装置
CN114627154B (zh) * 2022-03-18 2023-08-01 中国电子科技集团公司第十研究所 一种在频域部署的目标跟踪方法、电子设备及存储介质
CN114758271A (zh) * 2022-03-24 2022-07-15 京东科技信息技术有限公司 视频处理方法、装置、计算机设备及存储介质
CN115019390A (zh) * 2022-05-26 2022-09-06 北京百度网讯科技有限公司 视频数据处理方法、装置以及电子设备
CN115171080B (zh) * 2022-06-14 2025-09-09 安徽工程大学 面向车载设备的压缩视频驾驶员行为识别方法
CN117033954B (zh) * 2022-07-25 2025-01-28 腾讯科技(深圳)有限公司 一种数据处理方法及相关产品
CN115588235B (zh) * 2022-09-30 2023-06-06 河南灵锻创生生物科技有限公司 一种宠物幼崽行为识别方法及系统
US12608936B2 (en) * 2023-05-30 2026-04-21 Microsoft Technology Licensing, Llc Prior-driven supervision for weakly-supervised temporal action localization
CN116719420B (zh) * 2023-08-09 2023-11-21 世优(北京)科技有限公司 一种基于虚拟现实的用户动作识别方法及系统
CN118227984B (zh) * 2024-02-08 2025-05-13 深圳酷源数联科技有限公司 基于ai模型的工业环境数据分析方法、装置和系统
CN119207462B (zh) * 2024-09-30 2025-11-18 平安科技(深圳)有限公司 基于频带分割的声码器音频生成方法、装置、设备及介质

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20110182469A1 (en) * 2010-01-28 2011-07-28 Nec Laboratories America, Inc. 3d convolutional neural networks for automatic human action recognition
CN110414335A (zh) * 2019-06-20 2019-11-05 北京奇艺世纪科技有限公司 视频识别方法、装置及计算机可读存储介质
CN110826545A (zh) * 2020-01-09 2020-02-21 腾讯科技(深圳)有限公司 一种视频类别识别的方法及相关装置
CN111598026A (zh) * 2020-05-20 2020-08-28 广州市百果园信息技术有限公司 动作识别方法、装置、设备及存储介质

Family Cites Families (24)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5612735A (en) * 1995-05-26 1997-03-18 Luncent Technologies Inc. Digital 3D/stereoscopic video compression technique utilizing two disparity estimates
US5619256A (en) * 1995-05-26 1997-04-08 Lucent Technologies Inc. Digital 3D/stereoscopic video compression technique utilizing disparity and motion compensated predictions
US6055012A (en) * 1995-12-29 2000-04-25 Lucent Technologies Inc. Digital multi-view video compression with complexity and compatibility constraints
US7266150B2 (en) * 2001-07-11 2007-09-04 Dolby Laboratories, Inc. Interpolation of video compression frames
KR101454025B1 (ko) * 2008-03-31 2014-11-03 엘지전자 주식회사 영상표시기기에서 녹화정보를 이용한 영상 재생 장치 및 방법
ATE547775T1 (de) * 2009-08-21 2012-03-15 Ericsson Telefon Ab L M Verfahren und vorrichtung zur schätzung von interframe-bewegungsfeldern
CN106407889B (zh) * 2016-08-26 2020-08-04 上海交通大学 基于光流图深度学习模型在视频中人体交互动作识别方法
EP3321844B1 (en) * 2016-11-14 2021-04-14 Axis AB Action recognition in a video sequence
US11568545B2 (en) * 2017-11-20 2023-01-31 A9.Com, Inc. Compressed content object and action detection
US10528819B1 (en) * 2017-11-20 2020-01-07 A9.Com, Inc. Compressed content object and action detection
CN107808150A (zh) * 2017-11-20 2018-03-16 珠海习悦信息技术有限公司 人体视频动作识别方法、装置、存储介质及处理器
CN108280436A (zh) * 2018-01-29 2018-07-13 深圳市唯特视科技有限公司 一种基于堆叠递归单元的多级残差网络的动作识别方法
CN110163052B (zh) * 2018-08-01 2022-09-09 腾讯科技(深圳)有限公司 视频动作识别方法、装置和机器设备
US10909424B2 (en) * 2018-10-13 2021-02-02 Applied Research, LLC Method and system for object tracking and recognition using low power compressive sensing camera in real-time applications
CN109522867A (zh) * 2018-11-30 2019-03-26 国信优易数据有限公司 一种视频分类方法、装置、设备和介质
US11501532B2 (en) * 2019-04-25 2022-11-15 International Business Machines Corporation Audiovisual source separation and localization using generative adversarial networks
CN110490055A (zh) * 2019-07-08 2019-11-22 中国科学院信息工程研究所 一种基于三重编码的弱监督行为识别定位方法和装置
CN110490078B (zh) * 2019-07-18 2024-05-03 平安科技(深圳)有限公司 监控视频处理方法、装置、计算机设备和存储介质
CN110472531B (zh) * 2019-07-29 2023-09-01 腾讯科技(深圳)有限公司 视频处理方法、装置、电子设备及存储介质
CN112422863B (zh) * 2019-08-22 2022-04-12 华为技术有限公司 一种视频拍摄方法、电子设备和存储介质
CN110569814B (zh) * 2019-09-12 2023-10-13 广州酷狗计算机科技有限公司 视频类别识别方法、装置、计算机设备及计算机存储介质
CN110765860B (zh) * 2019-09-16 2023-06-23 平安科技(深圳)有限公司 摔倒判定方法、装置、计算机设备及存储介质
CN111080699B (zh) * 2019-12-11 2023-10-20 中国科学院自动化研究所 基于深度学习的单目视觉里程计方法及系统
WO2021140483A1 (en) * 2020-01-10 2021-07-15 Everseen Limited System and method for detecting scan and non-scan events in a self check out process

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20110182469A1 (en) * 2010-01-28 2011-07-28 Nec Laboratories America, Inc. 3d convolutional neural networks for automatic human action recognition
CN110414335A (zh) * 2019-06-20 2019-11-05 北京奇艺世纪科技有限公司 视频识别方法、装置及计算机可读存储介质
CN110826545A (zh) * 2020-01-09 2020-02-21 腾讯科技(深圳)有限公司 一种视频类别识别的方法及相关装置
CN111598026A (zh) * 2020-05-20 2020-08-28 广州市百果园信息技术有限公司 动作识别方法、装置、设备及存储介质

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See also references of EP4156017A4 *

Cited By (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114827739A (zh) * 2022-06-06 2022-07-29 百果园技术(新加坡)有限公司 一种直播回放视频生成方法、装置、设备及存储介质
CN115909479A (zh) * 2022-10-20 2023-04-04 中国科学院自动化研究所 人体行为识别方法、装置、电子设备及可读存储介质
CN116246259A (zh) * 2022-12-31 2023-06-09 南京行者易智能交通科技有限公司 一种基于3d卷积神经网络的公交安全驾驶行为识别方法
CN116129936A (zh) * 2023-02-02 2023-05-16 百果园技术(新加坡)有限公司 一种动作识别方法、装置、设备、存储介质及产品
CN117036789A (zh) * 2023-07-24 2023-11-10 西安电子科技大学 基于动态推理的全局与局部特征融合的骨架行为识别方法
CN117975376A (zh) * 2024-04-02 2024-05-03 湖南大学 基于深度分级融合残差网络的矿山作业安全检测方法
CN117975376B (zh) * 2024-04-02 2024-06-07 湖南大学 基于深度分级融合残差网络的矿山作业安全检测方法

Also Published As

Publication number Publication date
US12412426B2 (en) 2025-09-09
EP4156017B1 (en) 2025-10-08
EP4156017A4 (en) 2024-09-25
CN111598026B (zh) 2023-05-30
CN111598026A (zh) 2020-08-28
EP4156017A1 (en) 2023-03-29
US20230196837A1 (en) 2023-06-22

Similar Documents

Publication Publication Date Title
CN111598026B (zh) 动作识别方法、装置、设备及存储介质
Liao et al. Video-based person re-identification via 3d convolutional networks and non-local attention
Liu et al. Robust video super-resolution with learned temporal dynamics
CN111062314B (zh) 图像选取方法、装置、计算机可读存储介质及电子设备
US20200329233A1 (en) Hyperdata Compression: Accelerating Encoding for Improved Communication, Distribution & Delivery of Personalized Content
Ma et al. Joint feature and texture coding: Toward smart video representation via front-end intelligence
CN114973049B (zh) 一种统一卷积与自注意力的轻量视频分类方法
CN115171014B (zh) 视频处理方法、装置、电子设备及计算机可读存储介质
CN111444370A (zh) 图像检索方法、装置、设备及其存储介质
Li et al. Cross-level parallel network for crowd counting
Wang et al. Exploring hybrid spatio-temporal convolutional networks for human action recognition
CN112991239A (zh) 一种基于深度学习的图像反向恢复方法
CN111402118A (zh) 图像替换方法、装置、计算机设备和存储介质
Qi et al. Long-term temporal context gathering for neural video compression
CN112200816A (zh) 视频图像的区域分割及头发替换方法、装置及设备
Peng et al. Motion boundary emphasised optical flow method for human action recognition
Shen et al. Codedvision: Towards joint image understanding and compression via end-to-end learning
CN111680618A (zh) 基于视频数据特性的动态手势识别方法、存储介质和设备
CN115909479A (zh) 人体行为识别方法、装置、电子设备及可读存储介质
CN116543339B (zh) 一种基于多尺度注意力融合的短视频事件检测方法及装置
Dai et al. Slowfast diversity-aware prototype learning for egocentric action recognition
Zhang et al. Local compressed video stream learning for generic event boundary detection
CN113392269A (zh) 一种视频分类方法、装置、服务器及计算机可读存储介质
Li et al. MFMENet: multi-scale features mutual enhancement network for change detection in remote sensing images
CN116485807A (zh) 一种数字图像分割方法及装置

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 21808837

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

ENP Entry into the national phase

Ref document number: 2021808837

Country of ref document: EP

Effective date: 20221220

WWG Wipo information: grant in national office

Ref document number: 17999284

Country of ref document: US

WWG Wipo information: grant in national office

Ref document number: 2021808837

Country of ref document: EP