WO2021232969A1 - 动作识别方法、装置、设备及存储介质 - Google Patents
动作识别方法、装置、设备及存储介质 Download PDFInfo
- Publication number
- WO2021232969A1 WO2021232969A1 PCT/CN2021/085386 CN2021085386W WO2021232969A1 WO 2021232969 A1 WO2021232969 A1 WO 2021232969A1 CN 2021085386 W CN2021085386 W CN 2021085386W WO 2021232969 A1 WO2021232969 A1 WO 2021232969A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- video data
- layer
- frame
- feature map
- feature
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/764—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/7715—Feature extraction, e.g. by transforming the feature space, e.g. multi-dimensional scaling [MDS]; Mappings, e.g. subspace methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/80—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
- G06V10/806—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level of extracted features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
- G06V20/41—Higher-level, semantic clustering, classification or understanding of video scenes, e.g. detection, labelling or Markovian modelling of sport events or news items
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
- G06V20/44—Event detection
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
- G06V20/46—Extracting features or characteristics from the video content, e.g. video fingerprints, representative shots or key frames
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
- G06V20/49—Segmenting video sequences, i.e. computational techniques such as parsing or cutting the sequence, low-level clustering or determining units such as shots or scenes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/20—Movements or behaviour, e.g. gesture recognition
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/172—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a picture, frame or field
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/513—Processing of motion vectors
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/513—Processing of motion vectors
- H04N19/517—Processing of motion vectors by encoding
- H04N19/52—Processing of motion vectors by encoding by predictive encoding
Definitions
- the embodiments of the present application relate to the technical field of computer vision applications, such as motion recognition methods, devices, equipment, and storage media.
- Video-based action recognition has always been an important field of computer vision research.
- the realization of video action recognition mainly includes two parts: feature extraction and representation, and feature classification.
- Classical methods such as density trajectory tracking are generally the method of manually designing features.
- people have found that deep learning has powerful feature representation capabilities, and neural networks have gradually become the mainstream method in the field of action recognition.
- the feature method greatly improves the performance of action recognition.
- most of the neural network action recognition schemes are based on the sequence of pictures obtained from the video to construct the timing relationship, so as to judge the action. For example: based on the recurrent neural network to construct the timing between pictures, based on the 3D convolution to extract the timing information of multiple pictures, and the picture-based deep learning technology to superimpose the optical flow information of the action change, etc.
- the above scheme has the situation that the calculation amount and the recognition accuracy cannot be taken into account, and it needs to be improved.
- the embodiments of the present application provide an action recognition method, device, equipment, and storage medium, which can optimize an action recognition solution for videos in related technologies.
- an action recognition method which includes:
- the packet video data to be identified is input into a second preset model, and the action type included in the packet video data to be identified is determined according to the output result of the second preset model.
- an action recognition device which includes:
- the video grouping module is configured to perform grouping processing on the original compressed video data to obtain grouped video data
- a target grouped video determining module configured to input the grouped video data into a first preset model, and determine the target grouped video data containing actions according to the output result of the first preset model;
- the video decoding module is configured to decode the target packet video data to obtain the packet video data to be identified;
- the action type identification module is configured to input the packet video data to be identified into a second preset model, and determine the action type contained in the packet video data to be identified according to the output result of the second preset model.
- embodiments of the present application provide a computer device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor.
- the processor executes the computer program
- the computer program is The action recognition method provided by the embodiment.
- an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the action recognition method as provided in the embodiment of the present application is implemented.
- FIG. 1 is a schematic flowchart of an action recognition method provided by an embodiment of this application.
- FIG. 2 is a schematic diagram of a frame arrangement in a compressed video provided by an embodiment of the application
- FIG. 3 is a schematic diagram of a feature transformation operation provided by an embodiment of this application.
- FIG. 5 is a schematic diagram of an action recognition process based on a compressed video stream provided by an embodiment of this application;
- FIG. 6 is a schematic structural diagram of a first 2D residual network provided by an embodiment of this application.
- FIG. 7 is a schematic diagram of an action label provided by an embodiment of the application.
- FIG. 8 is a schematic flowchart of another action recognition method provided by an embodiment of this application.
- FIG. 9 is a schematic diagram of an application scenario of short video-based action recognition provided by an embodiment of the application.
- FIG. 10 is a schematic diagram of an action recognition process based on sequence pictures according to an embodiment of this application.
- FIG. 11 is a schematic diagram of a second 2D residual network structure provided by an embodiment of this application.
- FIG. 12 is a schematic diagram of a 3D residual network structure provided by an embodiment of this application.
- FIG. 13 is a schematic diagram of a calculation process of a two-level multi-receptive field pooling operation provided by an embodiment of the application;
- FIG. 14 is a structural block diagram of an action recognition device provided by an embodiment of this application.
- FIG. 15 is a structural block diagram of a computer device provided by an embodiment of this application.
- the action recognition solution in the embodiments of the present application can be applied to various video-oriented action recognition scenarios, such as short video review scenarios, video surveillance scenarios, real-time call recognition scenarios, robot visual recognition scenarios, and so on.
- the video can be a video file or a video stream.
- most of the neural network action recognition schemes are based on the sequence of pictures obtained from the video to construct the timing relationship, so as to judge the action. For example: based on Recurrent Neural Network (RNN) or Long Short-Term Memory (LSTM, Long Short-Term Memory) to construct the timing between pictures, extract the timing information of multiple pictures based on 3D convolution, and based on pictures
- RNN Recurrent Neural Network
- LSTM Long Short-Term Memory
- the deep learning technology superimposes the optical flow information of the action changes.
- the method of action recognition based on sequence images has at least the following two shortcomings: First, these technical solutions require a lot of computing resources to make judgments on a short video, and rely heavily on the machine central processing unit ( Central Processing Unit (CPU) and graphics processor (Graphics Processing Unit, GPU) computing power, and extracting pictures from compressed short videos needs to be decoded, such as decode decoding (the technology to decode compressed videos into pictures), and The full-time and long-segment decoding of short-term video itself requires a large amount of CPU and GPU computing power. Therefore, the image-based motion recognition scheme requires high machine computing power.
- CPU Central Processing Unit
- GPU Graphics Processing Unit
- Fig. 1 is a schematic flowchart of an action recognition method provided by an embodiment of the application.
- the method can be executed by an action recognition device, where the device can be implemented by software and/or hardware, and generally can be integrated in a computer device.
- the computer device may be, for example, a server, or a device such as a mobile phone, or two or more devices may perform part of the steps respectively.
- the method includes:
- Step 101 Perform grouping processing on the original compressed video data to obtain grouped video data.
- video data contains a large amount of image and sound information.
- the coding standard and coding parameters are not limited, and may be H264, H265, MPEG-4, etc., for example.
- the original compressed video data may be divided into intervals based on preset grouping rules to obtain interval compressed video (segment, hereinafter may be referred to as segment), the interval compressed video is used as the grouped video data, or part of the data is selected from the interval compressed video As packet video data.
- the preset grouping rule may include, for example, interval division time intervals, that is, the duration corresponding to each interval compressed video after being divided by intervals.
- the interval division time interval may be constant or variable, and is not limited. Taking a short video review scenario as an example, the interval division time interval may be, for example, 5 seconds.
- the preset grouping rule for example, may also include the number of intervals, and the value is not limited.
- the packet processing and interval division can be completed by using time stamps, that is, the interval range of the grouped video data or interval compressed video can be defined by the start and end timestamps. This step can be understood as extracting data of different time periods from the original compressed video data and inputting them to the first preset model respectively.
- the data corresponding to 0 ⁇ 5s in the original compressed video is a piece of packet video data
- the data corresponding to 5 ⁇ 10s is a piece of packet video data
- the two pieces of data are respectively entered into the first preset model of.
- Step 102 Input the grouped video data into a first preset model, and determine the target grouped video data containing the action according to the output result of the first preset model.
- each piece of grouped video data can be input into the first preset model to improve calculation efficiency; or after all grouping is completed, each segment can be grouped sequentially or in parallel
- the video data is input into the first preset model to ensure that coarse-grained action recognition is performed when the grouping processing is accurately completed.
- the first preset model may be a pre-trained neural network model, which is directly loaded when needed.
- the model is mainly set to identify whether the grouped video data contains an action, and does not care which action it is. You can set the action
- the label is two-category, such as "yes” and "no", the two-category label can be marked in the training sample of the first preset model. In this way, according to the output result of the first preset model, it is possible to filter out which segments contain actions, and determine the corresponding grouped video data as the target grouped video data.
- the calculation amount of the first preset model is small, and the recognition is performed without decoding, which can save a lot of decoding power and eliminate a large number of errors.
- the video segment containing the action is also guaranteed to be retained for subsequent identification.
- the network structure and related parameters of the first preset model are not limited in the embodiment of the present application, and can be set according to actual needs, for example, it can be a lightweight model.
- Step 103 Decode the target packet video data to obtain the packet video data to be identified.
- an appropriate decoding method can be selected with reference to factors such as the encoding standard of the original compressed video, which is not limited.
- the obtained video image can be used as the grouped video data to be identified (the decoded video images are generally arranged in chronological order, that is, sequence images.
- the grouped video data to be identified can contain video images Time sequence information), other information can also be extracted on this basis and used as the grouped video data to be identified.
- the other information here can be, for example, frequency domain information.
- Step 104 Input the packet video data to be identified into a second preset model, and determine the action type included in the packet video data to be identified according to the output result of the second preset model.
- the second preset model may be a pre-trained neural network model, which is directly loaded when needed.
- the model is mainly set to identify the types of actions contained in the grouped video data to be recognized, that is, perform fine-grained For recognition, multi-category labels can be marked in the training samples of the second preset model, so that the final action recognition result can be determined according to the output result of the second preset model.
- the data that can enter the second preset model has been greatly reduced compared to the original compressed video data, and the data purity is much higher than that of the original compressed video data.
- the magnitude of the number of video clips that need to be identified is not large. , So you can use the decoded sequence image for identification.
- the 3D convolutional network structure with more neural network parameters can be used to extract the time series features, and the tags are multi-classifications with finer granularity.
- the number of tags is not limited. For example, it can be 50. The recognition accuracy is adjusted.
- the original compressed video data is grouped to obtain the grouped video data, the grouped video data is input into the first preset model, and the output result of the first preset model is determined to include
- the target packet video data of the action is decoded to obtain the packet video data to be identified, and the packet video data to be identified is input into the second preset model, and the to be identified is determined according to the output result of the second preset model
- the type of action contained in the packet video data is used before decompressing the compressed video, the first preset model is used to roughly filter out the video clips that contain actions, and then the second preset model is used to accurately identify the types of actions contained, which can ensure recognition accuracy. Under the premise of effectively reducing the amount of calculation and improving the efficiency of action recognition.
- the client can perform preliminary screening of the compressed video to be uploaded, and upload the target grouped video data containing the action to the server for identification and review.
- the entire recognition process can also be completed locally on the client, that is, the related operations from step 101 to step 104, to realize the control of whether to allow the upload of the video according to the finally recognized action type. .
- the grouping processing of the original compressed video data to obtain the grouped video data includes: dividing the original compressed video data into intervals based on a preset grouping rule to obtain the interval compressed video; and extracting the compressed video based on a preset extraction strategy.
- the I frame data and P frame data in the interval compression video are used to obtain packetized video data, where the P frame data includes motion vector information and/or color element change residual information corresponding to the P frame.
- I frame also known as key frame
- P frame also known as forward predictive coding frame
- I frame generally includes the change information of the reference I frame in the compressed video, and may include motion vector (MV) information
- RGBR residual
- the distribution or content of the I frame and the P frame may be different.
- the preset grouping rule may be as described above, for example, may include the interval division time interval or the number of segments, etc.
- the preset extraction strategy may include, for example, an extraction strategy for I frames and an extraction strategy for P frames.
- the extraction strategy of I frames may include the time interval for acquiring I frames, or the number of I frames acquired in a unit time, for example, acquiring 1 I frame every second.
- the P frame extraction strategy may include acquiring the number of P frames following an I frame and the time interval between every two acquired P frames, for example, the number is two. Then, taking an interval compressed video of 5 seconds as an example, data corresponding to 5 I frames and 10 P frames can be obtained. If the P frame data includes both MV and RGBR, 10 MVs and 10 RGBRs can be obtained.
- the first preset model includes a first 2D residual network, a first splicing layer, and a first fully connected layer; after the packet video data is input into the first preset model , Obtaining corresponding feature maps of the same dimension through the first 2D residual network; obtaining the splicing feature maps after performing splicing operations in the order of the frames via the first splicing layer; the splicing feature maps A classification result of whether an action is included is obtained through the first fully connected layer.
- the technical solutions provided by the embodiments of the present application use relatively simplified network results to obtain higher recognition efficiency and ensure a higher recall rate of video clips containing actions.
- the first 2D residual network may adopt a lightweight ResNet18 model. Since I frame, MV, and RGBR cannot be directly concatenated in the data features, three ResNet18 can be used to process the I frame, MV and RGBR separately to obtain feature maps with the same dimensions, which can be recorded separately I frame feature maps, MV feature maps, and RGBR feature maps.
- MV feature maps and RGBR feature maps are collectively referred to as P frame feature maps.
- the same dimension can mean that C*H*W is consistent, where C represents channel, H represents height, W represents width, and * can also be represented as ⁇ .
- feature maps with the same dimension can be spliced through the first splicing layer.
- extracting I frame data and P frame data in the interval compressed video based on a preset extraction strategy to obtain packetized video data includes: extracting I frame data in the interval compressed video based on a preset extraction strategy And P frame data; accumulatively transform the P frame data, so that the transformed P frame data depends on the forward adjacent I frame; determine the packet video data according to the I frame data and the transformed P frame data .
- the first preset model further includes an addition layer located before the splicing layer, the feature map corresponding to the P frame data in the feature map is denoted as the P frame feature map, and the feature map corresponds to I
- the feature map of the frame data is marked as an I frame feature map;
- the P frame feature map and the I frame feature map obtain the P frame feature after the addition operation on the basis of the I frame feature map through the addition layer Figure;
- the I frame feature map and the P frame feature map after the addition operation are obtained through the first splicing layer after the splicing operation is performed in the order of the frames.
- Figure 2 is a schematic diagram of frame arrangement in a compressed video provided by an embodiment of this application.
- the MV and RGBR of each P frame depend on the previous one.
- the P frame of the frame (such as the P2 frame depends on the P1 frame).
- the MV and RGBR of the P frame can be cumulatively transformed to obtain the MV and the P frame of the input neural network.
- RGBR is relative to the previous I frame (for example, after accumulative transformation, the P2 frame becomes dependent on the previous I frame), rather than relative to the previous P frame.
- the above-mentioned addition operation can refer to the residual addition method in ResNet to directly output the MV feature map and RGBR feature map according to each element (pixel) and I frame (I frame after the first 2D residual network processing Feature maps) are added, and then the 3 feature maps are spliced in the order of the frames.
- a feature transformation layer is included before the residual structure of the first 2D residual network; before the packet video data enters the residual structure, an upward feature transformation is obtained through the feature transformation layer. And/or down-characterized grouped video data.
- the feature map before entering the residual structure, that is, before performing the convolution operation, the feature map may be subjected to a feature shift (shift) operation through Feature Shift (FS), Make a feature map contain some of the features in the feature map at different time points, so that when the convolution operation is performed, the feature map contains timing information, which can handle the collection of timing information without increasing the amount of calculation. And fusion, enrich the information of the feature map to be recognized, and improve the recognition accuracy. Compared with the related technology based on 3D convolution or optical flow information, it can effectively reduce the amount of calculation.
- FIG. 3 is a schematic diagram of a feature transformation operation provided by an embodiment of this application.
- Figure 3 taking the MV feature map as an example, for ease of description, Figure 3 only shows the process of performing feature transformation operations on three MV feature maps.
- the three MV feature maps correspond to different time points.
- Divide the MV feature map into 4 parts. Assuming that the first part is subjected to downward feature transformation (Down Shift), the second part is subjected to upward feature transformation (Up Shift), and the second MV feature map after transformation is at the same time Contains some features of the MV feature map at 3 time points.
- the division rule and the area of the upward feature transformation and/or the downward feature transformation can be set according to the actual situation.
- extracting P frame data in the interval compressed video based on a preset extraction strategy includes: extracting a preset number of P frame data in the interval compressed video in an equally spaced manner; wherein, in the In the training phase of the first preset model, a preset number of P-frame data in the interval compressed video is extracted in a random interval manner.
- FIG. 4 is a schematic flowchart of another action recognition method provided by an embodiment of the application, which is refined on the basis of the foregoing exemplary embodiments.
- the method includes:
- Step 401 Perform interval division on the original compressed video data based on a preset grouping rule to obtain an interval compressed video.
- the original compressed video can be enhanced with a preset video and image enhancement strategy, and the enhancement strategy can be selected according to the data requirements of the service and the configuration conditions of the driver.
- the same enhancement strategy can be used in model training and model application.
- Step 402 Extract I frame data and P frame data in the interval compressed video based on a preset extraction strategy.
- the P frame data includes the motion vector information and the pigment change residual information corresponding to the P frame.
- Step 403 Perform cumulative transformation on the P frame data, so that the transformed P frame data depends on the forward adjacent I frame, and determine the grouped video data according to the I frame data and the transformed P frame data.
- Step 404 Input the grouped video data into the first preset model, and determine the target grouped video data containing the action according to the output result of the first preset model.
- FIG. 5 is a schematic diagram of an action recognition process based on a compressed video stream provided by an embodiment of the application.
- the compressed video is divided into n segments.
- the extracted I frame data, MV The data and RGBR data are input into the first preset model.
- the first preset model includes the first 2D residual network, the addition layer, the first splicing layer, and the first fully connected layer (Fully connected layer, FC layer).
- the first 2D residual network is 2D Res18.
- the feature transformation layer (FS) is included before the residual structure.
- the training method in the training phase of the first preset model, can be selected according to actual needs, including loss function (loss), etc., for example, the loss function can use cross entropy, and other auxiliary loss functions can also be used to improve the model. Effect.
- some non-heuristic optimization algorithms can be used to improve the convergence speed and optimization performance of stochastic gradient descent.
- Fig. 6 is a schematic diagram of a first 2D residual network structure provided by an embodiment of this application.
- 2D Res18 is an 18-layer residual neural network, which consists of 4 stages and a total of 8 uses 2D Convolutional residual block (block) composition, the network structure is relatively shallow, in order to use the convolutional layer as much as possible to fully extract features, 8 2D residual blocks can all be performed by 3*3 convolution, that is, the convolution kernel parameters It is 3*3.
- the convolutional layer generally refers to a network layer used to complete the weighted summation of local pixel values and nonlinear activation.
- each 2D residual block can use a bottleneck.
- the design concept of (bottleneck) is that each residual block is composed of 3 convolutional layers (convolution kernel parameters are 1*1, 3*3, and 1*1), and the first and last layers are used for compression. And restore the image channel.
- the first 2D residual network structure and various parameters can also be adjusted according to actual needs.
- the I frame data, MV data and RGBR data are respectively obtained through the FS layer through the upward feature transformation and the downward feature transformation of the I frame data.
- MV data and RGBR data and then obtain a C*H*W consistent feature map through the residual structure respectively, and directly compare the MV feature map and RGBR feature map with the residual addition method in ResNet through the addition layer
- Add each element to the output of the I frame and then concate the three feature maps in the order of the frame to get the spliced feature map.
- the spliced feature map is passed through the first fully connected layer (FC) to get whether it contains action The classification results.
- FC fully connected layer
- Figure 7 is a schematic diagram of an action label provided by an embodiment of the application.
- the last two circles in Figure 5 indicate that the output is two-category.
- Setting the action label to two-category is to ensure that the recall of fragments containing actions is improved.
- the label granularity is whether to include The action is designed to have a two-category level of "Yes” or "No", that is, it does not care which sub-category action is included.
- the short video used for training can be cut into 6 segments, where segment S2 contains action A1, segment S4 contains action A2, and A1 and A2 are two different actions, but The labels corresponding to these two positions are the same, and they are both set to 1, which means that the actions of A1 and A2 are not distinguished. Therefore, in actual application, the output result of the first preset model is also "Yes” or "No", which realizes coarse-grained recognition of actions in the compressed video.
- Step 405 Decode the target packet video data to obtain the packet video data to be identified.
- Step 406 Input the packet video data to be identified into the second preset model, and determine the action type included in the packet video data to be identified according to the output result of the second preset model.
- the action recognition method provided by the embodiments of this application first performs action recognition based on compressed video. Without decompressing the video, extracts the MV and RGBR information of the I frame and the P frame, and changes the MV and RGBR to increase the I frame.
- the information dependence of realizes the processing of a large number of videos of variable duration with less computational power requirements, and uses FS without computational power requirements to increase the model's extraction of timing information.
- the enhancement of model capabilities does not lead to a decrease in computational efficiency. Redesigned the label granularity of actions to ensure the goal of recall. Let the model deal with simple binary classification problems, which can improve the recall rate, and then accurately identify the preliminarily screened video clips containing actions, which is effective under the premise of ensuring recognition accuracy. Reduce the amount of calculation and improve the efficiency of action recognition.
- the decoding the target packet video data to obtain the to-be-identified packet video data includes: decoding the target packet video data to obtain the to-be-identified segmented video image; and obtaining the to-be-identified segmented video image
- the frequency domain information in the segmented video image is generated according to the frequency domain information, and the corresponding frequency domain map is generated; the segmented video image to be identified and the corresponding frequency domain map are used as the packetized video data to be identified.
- the second preset model includes a model based on an Efficient Convolutional Network for Online Video (ECO) architecture for online video understanding.
- ECO Efficient Convolutional Network for Online Video
- the ECO architecture can be understood as a video feature extractor. It provides an architecture design for video feature acquisition, which includes a 2D feature extraction network and a 3D feature extraction network. This architecture can achieve better performance while increasing speed.
- This application is implemented The example can be improved and designed on the basis of the ECO architecture to obtain the second preset model.
- the second preset model includes a second splicing layer, a second 2D residual network, a 3D residual network, a third splicing layer, and a second fully connected layer;
- the packet video data to be identified is After being input into the second preset model, the spliced image data obtained by splicing the segmented video image to be identified and the corresponding frequency domain map is obtained through the second splicing layer;
- the spliced image data passes through the first
- the second 2D residual network obtains a 2D feature map;
- the middle layer output result of the second 2D residual network is used as the input of the 3D residual network, and a 3D feature map is obtained through the 3D residual network;
- the 2D The feature map and the 3D feature map obtain a spliced feature map to be identified via the third splicing layer;
- the feature map to be identified obtains a corresponding action type label via the second fully connected layer.
- the technical solution provided by the embodiment of the present application
- the grouped video data to be identified first enters the 2D convolution part of the model to extract the features of each image, then enters the 3D convolution part to extract the timing information of the action, and finally outputs the result of the multi-classification.
- the network structure and related parameters of the second 2D residual network and the 3D residual network can be set according to actual requirements.
- the second 2D residual network is 2D Res50
- the 3D residual network is 3D Res10.
- the second preset model further includes a first pooling layer and a second pooling layer; the 2D feature map passes through the first pooling layer before being input to the third stitching layer Obtain the corresponding one-dimensional 2D feature vector containing the first element quantity; before the 3D feature map is input to the third stitching layer, obtain the corresponding one containing the second element quantity through the second pooling layer Dimensional 3D feature vector.
- the 2D feature map and the 3D feature map obtain the spliced feature map to be recognized through the third splicing layer, including: the one-dimensional 2D feature vector and the one-dimensional 3D feature vector are passed through the The third splicing layer obtains the spliced vector to be recognized.
- the feature map still has a larger size after feature extraction through the convolutional layer.
- the feature map is directly flattened, the dimension of the feature vector may be too high. Therefore, the 2D and After the 3D residual network completes a series of feature extraction operations, the pooling operation can be used to directly aggregate the feature maps into one-dimensional feature vectors to reduce the dimensionality of the feature vectors.
- the first pooling layer and the second pooling layer can be designed according to actual needs, for example, it can be global average pooling (GAP) or maximum pooling.
- GAP global average pooling
- the number of the first element and the number of the second element can also be set freely, for example, according to the number of channels of the feature map.
- the first pooling layer includes a multi-receptive field pooling layer.
- the receptive field generally refers to the image or video range covered by the feature value, which is used to indicate the size of the receptive range of the original image by different neurons in the network, or in other words, the pixels on the feature map output by each layer of the convolutional neural network The size of the area mapped on the original image.
- the larger the value of the receptive field the larger the range of the original image it can touch, and it also means that it may contain more global and higher semantic features; on the contrary, the smaller the value, the more the features it contains. Part and detail. Therefore, the value of the receptive field can be used to roughly judge the abstraction level of each layer.
- the advantage of adopting the multi-sensory field pooling layer is that the features can be sensitive to targets of different scales, and the recognition range of the action category is broadened.
- the multi-receptive field realization method can be set according to actual needs.
- the first pooling layer includes a first-level local pooling layer, a second-level global pooling layer, and a vector merging layer
- the first-level local pooling layer includes at least two pools with different sizes.
- the core; the 2D feature map obtains corresponding at least two sets of 2D pooling feature maps of different scales through the first-level local pooling layer; the at least two sets of 2D pooling feature maps of different scales are passed through the second level
- the global pooling layer obtains at least two sets of feature vectors; the at least two sets of feature vectors obtain corresponding one-dimensional 2D feature vectors containing the first number of elements through the vector merging layer.
- the technical solution provided by the embodiments of this application adopts two-level multi-receptive field pooling to summarize the total feature map, and uses different size pooling cores for multi-scale pooling, so that the summarized features have different receptive fields, and at the same time, the feature map
- the size is greatly reduced, the recognition efficiency is improved, and then the two-level global pooling is used to summarize the features, and the vector merging layer is used to obtain a one-dimensional 2D feature vector.
- Figure 8 is a schematic flow chart of another action recognition method provided by an embodiment of this application. It is refined on the basis of the above-mentioned exemplary embodiments.
- Figure 9 is a short-based A schematic diagram of the application scenario of video action recognition, as shown in Figure 9, after the user uploads a short video, first extract the compressed video stream information, and use the pre-trained and constructed action video model based on the compressed video (the first preset model) to identify that it contains The target segment of the action. Other segments are screened out because they do not contain actions.
- the target segment is decoded and the time domain and frequency domain information of the picture is extracted, and the action recognition model based on the decoded picture constructed in advance is used (the second preset The model) performs finer-grained action recognition based on the picture sequence, and obtains the action type corresponding to each target segment, and then determines whether it is a target action. If it is not a target action, it is also filtered out.
- the method may include:
- Step 801 Perform interval division on the original compressed video data based on a preset grouping rule to obtain an interval compressed video.
- the original compressed video data is the video stream data of the short video uploaded by the user, and the length of each section of the compressed video may be 5 seconds.
- Step 802 Extract I frame data and P frame data in the interval compressed video based on a preset extraction strategy.
- the P frame data contains the MV and RGBR corresponding to the P frame, and 1 I frame and 2 P frames thereafter are obtained every second, that is, a 5 second segment can obtain 5 I frames, 10 MVs and 10 RGBR.
- Step 803 Perform cumulative transformation on the P frame data, so that the transformed P frame data depends on the forward adjacent I frame, and determine the grouped video data according to the I frame data and the transformed P frame data.
- Step 804 Input the grouped video data into the first preset model, and determine the target grouped video data containing the action according to the output result of the first preset model.
- Step 805 Decode the target grouped video data to obtain the segmented video image to be identified.
- the images are decoded from the target grouped video data, and they are arranged in a sequence in chronological order to obtain the sequence pictures.
- a set number of images can be acquired from the sequence of pictures according to a preset acquisition strategy as the segmented video images to be identified.
- the preset acquisition strategy is, for example, acquisition at equal intervals, and the set number is, for example, 15.
- Step 806 Obtain frequency domain information in the segmented video image to be identified, and generate a corresponding frequency domain map according to the frequency domain information.
- the frequency domain information collection method is not limited, and the frequency domain information can be collected for each picture in the sequence picture, and a frequency domain (FD) map corresponding to the sequence picture one-to-one can be generated.
- FD frequency domain
- Step 807 Use the segmented video image to be identified and the corresponding frequency domain map as the grouped video data to be identified, input them into the second preset model, and determine according to the output result of the second preset model that the grouped video data to be identified contains The type of action.
- the sequence picture and the corresponding frequency domain map in one segment will correspond to one label in the multi-category label, that is, in the output result of the second preset model, each segment corresponds to one label.
- the model will use the corresponding type of the main action as the label of the segment. For example, in a segment A of 5 seconds, there is an action in 4 seconds, the action in 3 seconds is a hand wave, and the action in 1 second is a kick, then the action tag corresponding to the segment A is a hand wave.
- FIG. 10 is a schematic diagram of an action recognition process based on a sequence of pictures provided by an embodiment of the present application, where the value of n is generally different from the value of n in FIG. Picture (Image), after extracting the frequency domain information, generate the corresponding frequency domain map (FD).
- the second preset model can be designed based on the ECO architecture, which can include the second splicing layer, the second 2D residual network, the 3D residual network, the first pooling layer, the second pooling layer, the third splicing layer, and the second Fully connected layer.
- training methods including loss functions, can be selected according to actual needs.
- the loss function can use cross entropy, and other auxiliary loss functions can also be used to improve the model effect.
- some non-heuristic optimization algorithms can be used to improve the convergence speed and optimization performance of stochastic gradient descent.
- the decoded sequence picture and the corresponding frequency domain map pass through the second splicing layer in the second preset model to obtain spliced image data.
- the subsequent spliced image data may also pass through the convolutional layer. After processing (conv) and the maximum pooling layer (maxpool), they are then input into the second 2D residual network (2D Res50).
- the output of 2D Res50 passes through the first pooling layer (multi-receptive field pooling layer) to obtain a 1024-dimensional 2D feature vector (which can be understood as a one-dimensional row vector or column vector containing 1024 elements).
- the intermediate output result of the 2D Res50 will be used as the input of the 3D residual network (3D Res10), and the 512-dimensional 3D feature vector will be obtained after the second pooling layer (GAP).
- the feature map to be recognized that is, the feature vector to be recognized, is obtained, and finally the final action type label is obtained through the second fully connected layer.
- FIG 11 is a schematic diagram of a second 2D residual network structure provided by an embodiment of the application.
- the second 2D residual network is 2D Res50, which is a 50-layer residual neural network, consisting of 4 stages in total It is composed of 16 residual blocks using 2D convolution.
- each residual block can use the bottleneck design concept, that is, each residual block is composed of 3 convolutional layers, of which the import and export two 1* 1 Convolution is used to compress and restore the number of channels of the feature map.
- 2D projection residual blocks can be used at the entrance of each stage.
- This residual block is A 1*1 convolutional layer is added to the bypass to ensure that the size of the feature map and the number of channels remain the same when the pixel-by-pixel addition operation is performed.
- only using 2D projection residual blocks at the entrance of each stage can also reduce network parameters.
- FIG. 12 is a schematic diagram of a 3D residual network structure provided by an embodiment of this application.
- Each segment takes N video frames (frame pictures obtained after the splicing operation of the 15 images obtained from the sequence of images and the corresponding frequency domain images as described above) the middle layer feature map group obtained through 2D Res50, for example
- the feature map group from stage2-block4 can be assembled into a three-dimensional video tensor, the tensor shape is (c, f, h, w), where c is the number of channels of the frame picture, f is the number of video frames, h, w refers to the height and width of the frame picture respectively, and the video tensor is input to 3D Res10 to extract the spatiotemporal features of the entire video.
- 3D Res10 consists of only 3 stages and a total of 5 residual blocks. All convolutional layers use three-dimensional convolution kernels. During the convolution process, the information in the time dimension will also participate in the calculation. Also in order to reduce network parameters, 3D Res10 can use residual blocks with fewer convolutional layers, and can eliminate the channel number expansion technology used in bottleneck residual blocks.
- global average pooling can be used to average the pixel values in the space-time range to obtain a 512-dimensional video space-time feature vector.
- the kernel size of the global average pooling can be, for example, It is 2*7*7.
- the 2D Res50 can be different from the general convolutional neural network. It does not use a simple global pooling operation to summarize the total feature map, but can use two-level multi-receptive field pooling to summarize the total feature map.
- FIG. 13 is a schematic diagram of the operation process of the two-level multi-receptive field pooling operation provided by an embodiment of the application.
- the first-level local pooling operation three maximum pooling with different totalization ranges (pooling cores) are used to totalize the pixels on the 2D feature map (three
- the maximum pooling core size can be, for example, 7*7, 4*4, and 2*2, respectively. Pooling cores of different sizes make the aggregated features have different receptive fields, making the features sensitive to targets of different scales. Broaden the scope of recognition of target categories.
- three sets of 2D pooling feature maps with greatly reduced sizes are obtained. These 2D pooling feature maps contain video features of different scales. Then, the three sets of feature maps are respectively subjected to the second level of global sum-pooling operation.
- the feature maps on each channel obtain a feature by summing the pixel values, and the three sets of feature maps are summed into Three feature vectors containing video spatial information of different scales. Finally, perform a pixel-by-pixel addition operation on the three feature vectors to obtain a 1024-dimensional 2D feature vector (as shown by the bottom circle in Figure 13).
- N 15 as mentioned above
- frames of the video will obtain N 1024-dimensional 2D feature vectors, and the sum of the feature vectors of these frame pictures can be averaged to obtain the 2D feature vector representing the spatial information of the entire video. (That is, the 1024-dimensional 2D feature vector shown in Figure 10).
- the action recognition method extracts frequency domain information from the decoded image in the action recognition part based on the sequence of pictures, and splices it into the image, increasing the richness of information extracted by the feature network
- the information is pooled in different ways and merged into a one-dimensional vector to be recognized, using the ECO network structure, and classifying the network with more tags, providing a higher-precision recognition .
- the action recognition part based on the compressed video stream has filtered out a large number of non-target action videos, the amount of video input to the action recognition part based on sequence pictures, which requires a higher computing power, is small. Combining the two parts, it is guaranteed In the case of motion recognition accuracy, all video detection tasks are completed with very small computing power requirements, which effectively improves recognition efficiency, and achieves a good balance of computing power requirements, recall rate and recognition accuracy.
- FIG. 14 is a structural block diagram of an action recognition device provided by an embodiment of the application.
- the device can be implemented by software and/or hardware, and generally can be integrated in a computer device, and can perform action recognition by executing an action recognition method. As shown in Figure 14, the device includes:
- the video grouping module 1401 is configured to perform grouping processing on the original compressed video data to obtain grouped video data;
- the target grouped video determining module 1402 is configured to input the grouped video data into a first preset model, and determine the target grouped video data containing actions according to the output result of the first preset model;
- the video decoding module 1403 is configured to decode the target packet video data to obtain the packet video data to be identified;
- the action type identification module 1404 is configured to input the packet video data to be identified into a second preset model, and determine the action type contained in the packet video data to be identified according to the output result of the second preset model .
- the action recognition device Before decompressing the compressed video, the action recognition device provided in the embodiment of the present application first uses the first preset model to roughly filter out the video clips that contain actions, and then uses the second preset model to accurately identify the types of actions contained. On the premise of ensuring recognition accuracy, the amount of calculation is effectively reduced and the efficiency of action recognition is improved.
- FIG. 15 is a structural block diagram of a computer device provided by an embodiment of this application.
- the computer device 1500 includes a memory 1501, a processor 1502, and a computer program that is stored on the memory 1501 and can run on the processor 1502.
- the processor 1502 implements the action recognition method provided in the embodiment of the present application when the computer program is executed.
- the embodiments of the present application also provide a storage medium containing computer-executable instructions, which are used to execute the action recognition method provided in the embodiments of the present application when the computer-executable instructions are executed by a computer processor.
- the action recognition apparatus, equipment, and storage medium provided in the foregoing embodiments can execute the action recognition method provided in any embodiment of the present application, and have corresponding functional modules for executing the method.
- the action recognition method provided in any embodiment of this application can execute the action recognition method provided in any embodiment of this application, and have corresponding functional modules for executing the method.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- Multimedia (AREA)
- General Physics & Mathematics (AREA)
- Computing Systems (AREA)
- Evolutionary Computation (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Software Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Artificial Intelligence (AREA)
- Databases & Information Systems (AREA)
- Medical Informatics (AREA)
- Signal Processing (AREA)
- Computational Linguistics (AREA)
- Biophysics (AREA)
- General Engineering & Computer Science (AREA)
- Mathematical Physics (AREA)
- Biomedical Technology (AREA)
- Life Sciences & Earth Sciences (AREA)
- Data Mining & Analysis (AREA)
- Molecular Biology (AREA)
- Psychiatry (AREA)
- Social Psychology (AREA)
- Human Computer Interaction (AREA)
- Image Analysis (AREA)
- Television Signal Processing For Recording (AREA)
Abstract
Description
Claims (15)
- 一种动作识别方法,包括:对原始压缩视频数据进行分组处理,得到分组视频数据;将所述分组视频数据输入至第一预设模型中,并根据所述第一预设模型的输出结果确定包含动作的目标分组视频数据;对所述目标分组视频数据进行解码,得到待识别分组视频数据;将所述待识别分组视频数据输入至第二预设模型中,并根据所述第二预设模型的输出结果确定所述待识别分组视频数据中包含的动作类型。
- 根据权利要求1所述的方法,其中,所述对原始压缩视频数据进行分组处理,得到分组视频数据,包括:基于预设分组规则对原始压缩视频数据进行区间划分,得到区间压缩视频;基于预设提取策略提取所述区间压缩视频中的关键帧I帧数据和前向预测编码帧P帧数据,得到分组视频数据,其中,所述P帧数据包含P帧对应的以下至少一种信息:运动矢量信息,色素变化残差信息。
- 根据权利要求2所述的方法,其中,所述第一预设模型中包含第一2D残差网络、第一拼接层和第一全连接层;所述分组视频数据被输入至所述第一预设模型中后,经由所述第一2D残差网络得到对应的维度相同的特征图;所述特征图经由所述第一拼接层得到按照帧的先后顺序进行拼接操作后的拼接特征图;所述拼接特征图经由所述第一全连接层得到是否包含动作的分类结果。
- 根据权利要求3所述的方法,其中,基于预设提取策略提取所述区间压缩视频中的I帧数据和P帧数据,得到分组视频数据,包括:基于预设提取策略提取所述区间压缩视频中的I帧数据和P帧数据;对所述P帧数据进行累加变换,以使得变换后的P帧数据依赖于前向相邻的所述I帧数据;根据所述I帧数据和变换后的P帧数据确定分组视频数据;所述第一预设模型中还包括位于所述拼接层之前的相加层,所述特征图中对应P帧数据的特征图记为P帧特征图,所述特征图中对应I帧数据的特征图记为I帧特征图;所述P帧特征图和所述I帧特征图经由所述相加层得到在所述I帧特征图基 础上经过相加操作后的P帧特征图;所述I帧特征图和所述经过相加操作后的P帧特征图经由所述第一拼接层得到按照帧的先后顺序进行拼接操作后的拼接特征图。
- 根据权利要求3所述的方法,其中,在所述第一2D残差网络的残差结构前包括特征变换层;所述分组视频数据在进入所述残差结构前,经由所述特征变换层得到以下至少一种分组视频数据:经过向上特征变换的分组视频数据,经过向下特征变换的分组视频数据。
- 根据权利要求2所述的方法,其中,基于预设提取策略提取所述区间压缩视频中的P帧数据,包括:采用等间隔方式提取所述区间压缩视频中的预设数量的P帧数据;其中,在所述第一预设模型的训练阶段,采用随机间隔方式提取区间压缩视频中的预设数量的P帧数据。
- 根据权利要求1所述的方法,其中,所述对所述目标分组视频数据进行解码,得到待识别分组视频数据,包括:对所述目标分组视频数据进行解码,得到待识别分段视频图像;获取所述待识别分段视频图像中的频域信息,根据所述频域信息生成对应的频域图;将所述待识别分段视频图像和对应的频域图作为待识别分组视频数据。
- 根据权利要求7所述的方法,其中,所述第二预设模型包括基于用于在线视频理解的高效卷积网络ECO架构的模型。
- 根据权利要求8所述的方法,其中,所述第二预设模型包括第二拼接层、第二2D残差网络、3D残差网络、第三拼接层和第二全连接层;所述待识别分组视频数据被输入至第二预设模型中后,经由所述第二拼接层得到对所述待识别分段视频图像和对应的频域图进行拼接后的拼接图像数据;所述拼接图像数据经由所述第二2D残差网络得到2D特征图;将所述第二2D残差网络的中间层输出结果作为所述3D残差网络的输入,并经由所述3D残差网络得到3D特征图;所述2D特征图和所述3D特征图经由所述第三拼接层得到拼接后的待识别特征图;所述待识别特征图经由所述第二全连接层得到对应的动作类型标签。
- 根据权利要求9所述的方法,其中,所述第二预设模型还包括第一池化层和第二池化层;所述2D特征图在被输入至所述第三拼接层之前,经由所述第一池化层得到对应的包含第一元素数量的一维2D特征向量;所述3D特征图在被输入至所述第三拼接层之前,经由所述第二池化层得到对应的包含第二元素数量的一维3D特征向量;所述2D特征图和所述3D特征图经由所述第三拼接层得到拼接后的待识别特征图,包括:所述一维2D特征向量和所述一维3D特征向量经由所述第三拼接层得到拼接后的待识别向量。
- 根据权利要求10所述的方法,其中,所述第一池化层包括多感受野池化层。
- 根据权利要求11所述的方法,其中,所述第一池化层包括一级局部池化层、二级全局池化层和向量合并层,所述一级局部池化层中包含至少两个尺寸不相同的池化核;所述2D特征图经由所述一级局部池化层得到对应的至少两组不同尺度的2D池化特征图;所述至少两组不同尺度的2D池化特征图经由所述二级全局池化层得到至少两组特征向量;所述至少两组特征向量经由所述向量合并层得到对应的包含第一元素数量的一维2D特征向量。
- 一种动作识别装置,包括:视频分组模块,设置为对原始压缩视频数据进行分组处理,得到分组视频数据;目标分组视频确定模块,设置为将所述分组视频数据输入至第一预设模型中,并根据所述第一预设模型的输出结果确定包含动作的目标分组视频数据;视频解码模块,设置为对所述目标分组视频数据进行解码,得到待识别分组视频数据;动作类型识别模块,设置为将所述待识别分组视频数据输入至第二预设模型中,并根据所述第二预设模型的输出结果确定所述待识别分组视频数据中包含的动作类型。
- 一种计算机设备,包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,所述处理器执行所述计算机程序时实现如权利要求1-12任一项所述的方法。
- 一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现如权利要求1-12中任一所述的方法。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP21808837.5A EP4156017B1 (en) | 2020-05-20 | 2021-04-02 | Action recognition method and apparatus, and device and storage medium |
| US17/999,284 US12412426B2 (en) | 2020-05-20 | 2021-04-02 | Action recognition method and apparatus, and device and storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202010431706.9 | 2020-05-20 | ||
| CN202010431706.9A CN111598026B (zh) | 2020-05-20 | 2020-05-20 | 动作识别方法、装置、设备及存储介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2021232969A1 true WO2021232969A1 (zh) | 2021-11-25 |
Family
ID=72183951
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2021/085386 Ceased WO2021232969A1 (zh) | 2020-05-20 | 2021-04-02 | 动作识别方法、装置、设备及存储介质 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US12412426B2 (zh) |
| EP (1) | EP4156017B1 (zh) |
| CN (1) | CN111598026B (zh) |
| WO (1) | WO2021232969A1 (zh) |
Cited By (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114827739A (zh) * | 2022-06-06 | 2022-07-29 | 百果园技术(新加坡)有限公司 | 一种直播回放视频生成方法、装置、设备及存储介质 |
| CN115909479A (zh) * | 2022-10-20 | 2023-04-04 | 中国科学院自动化研究所 | 人体行为识别方法、装置、电子设备及可读存储介质 |
| CN116129936A (zh) * | 2023-02-02 | 2023-05-16 | 百果园技术(新加坡)有限公司 | 一种动作识别方法、装置、设备、存储介质及产品 |
| CN116246259A (zh) * | 2022-12-31 | 2023-06-09 | 南京行者易智能交通科技有限公司 | 一种基于3d卷积神经网络的公交安全驾驶行为识别方法 |
| CN117036789A (zh) * | 2023-07-24 | 2023-11-10 | 西安电子科技大学 | 基于动态推理的全局与局部特征融合的骨架行为识别方法 |
| CN117975376A (zh) * | 2024-04-02 | 2024-05-03 | 湖南大学 | 基于深度分级融合残差网络的矿山作业安全检测方法 |
Families Citing this family (20)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111598026B (zh) | 2020-05-20 | 2023-05-30 | 广州市百果园信息技术有限公司 | 动作识别方法、装置、设备及存储介质 |
| CN112183359B (zh) * | 2020-09-29 | 2024-05-14 | 中国科学院深圳先进技术研究院 | 视频中的暴力内容检测方法、装置及设备 |
| CN112337082B (zh) * | 2020-10-20 | 2021-09-10 | 深圳市杰尔斯展示股份有限公司 | 一种ar沉浸式虚拟视觉感知交互系统和方法 |
| CN112396161B (zh) * | 2020-11-11 | 2022-09-06 | 中国科学技术大学 | 基于卷积神经网络的岩性剖面图构建方法、系统及设备 |
| CN112507920B (zh) * | 2020-12-16 | 2023-01-24 | 重庆交通大学 | 一种基于时间位移和注意力机制的考试异常行为识别方法 |
| CN112712042B (zh) * | 2021-01-04 | 2022-04-29 | 电子科技大学 | 嵌入关键帧提取的行人重识别端到端网络架构 |
| CN112686193B (zh) * | 2021-01-06 | 2024-02-06 | 东北大学 | 基于压缩视频的动作识别方法、装置及计算机设备 |
| CN112749666B (zh) * | 2021-01-15 | 2024-06-04 | 百果园技术(新加坡)有限公司 | 一种动作识别模型的训练及动作识别方法与相关装置 |
| CN116263618A (zh) * | 2021-12-13 | 2023-06-16 | 成都拟合未来科技有限公司 | 一种健身动作识别方法及系统 |
| CN114333065B (zh) * | 2021-12-31 | 2025-03-04 | 济南博观智能科技有限公司 | 一种应用于监控视频的行为识别方法、系统及相关装置 |
| CN114627154B (zh) * | 2022-03-18 | 2023-08-01 | 中国电子科技集团公司第十研究所 | 一种在频域部署的目标跟踪方法、电子设备及存储介质 |
| CN114758271A (zh) * | 2022-03-24 | 2022-07-15 | 京东科技信息技术有限公司 | 视频处理方法、装置、计算机设备及存储介质 |
| CN115019390A (zh) * | 2022-05-26 | 2022-09-06 | 北京百度网讯科技有限公司 | 视频数据处理方法、装置以及电子设备 |
| CN115171080B (zh) * | 2022-06-14 | 2025-09-09 | 安徽工程大学 | 面向车载设备的压缩视频驾驶员行为识别方法 |
| CN117033954B (zh) * | 2022-07-25 | 2025-01-28 | 腾讯科技(深圳)有限公司 | 一种数据处理方法及相关产品 |
| CN115588235B (zh) * | 2022-09-30 | 2023-06-06 | 河南灵锻创生生物科技有限公司 | 一种宠物幼崽行为识别方法及系统 |
| US12608936B2 (en) * | 2023-05-30 | 2026-04-21 | Microsoft Technology Licensing, Llc | Prior-driven supervision for weakly-supervised temporal action localization |
| CN116719420B (zh) * | 2023-08-09 | 2023-11-21 | 世优(北京)科技有限公司 | 一种基于虚拟现实的用户动作识别方法及系统 |
| CN118227984B (zh) * | 2024-02-08 | 2025-05-13 | 深圳酷源数联科技有限公司 | 基于ai模型的工业环境数据分析方法、装置和系统 |
| CN119207462B (zh) * | 2024-09-30 | 2025-11-18 | 平安科技(深圳)有限公司 | 基于频带分割的声码器音频生成方法、装置、设备及介质 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20110182469A1 (en) * | 2010-01-28 | 2011-07-28 | Nec Laboratories America, Inc. | 3d convolutional neural networks for automatic human action recognition |
| CN110414335A (zh) * | 2019-06-20 | 2019-11-05 | 北京奇艺世纪科技有限公司 | 视频识别方法、装置及计算机可读存储介质 |
| CN110826545A (zh) * | 2020-01-09 | 2020-02-21 | 腾讯科技(深圳)有限公司 | 一种视频类别识别的方法及相关装置 |
| CN111598026A (zh) * | 2020-05-20 | 2020-08-28 | 广州市百果园信息技术有限公司 | 动作识别方法、装置、设备及存储介质 |
Family Cites Families (24)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5612735A (en) * | 1995-05-26 | 1997-03-18 | Luncent Technologies Inc. | Digital 3D/stereoscopic video compression technique utilizing two disparity estimates |
| US5619256A (en) * | 1995-05-26 | 1997-04-08 | Lucent Technologies Inc. | Digital 3D/stereoscopic video compression technique utilizing disparity and motion compensated predictions |
| US6055012A (en) * | 1995-12-29 | 2000-04-25 | Lucent Technologies Inc. | Digital multi-view video compression with complexity and compatibility constraints |
| US7266150B2 (en) * | 2001-07-11 | 2007-09-04 | Dolby Laboratories, Inc. | Interpolation of video compression frames |
| KR101454025B1 (ko) * | 2008-03-31 | 2014-11-03 | 엘지전자 주식회사 | 영상표시기기에서 녹화정보를 이용한 영상 재생 장치 및 방법 |
| ATE547775T1 (de) * | 2009-08-21 | 2012-03-15 | Ericsson Telefon Ab L M | Verfahren und vorrichtung zur schätzung von interframe-bewegungsfeldern |
| CN106407889B (zh) * | 2016-08-26 | 2020-08-04 | 上海交通大学 | 基于光流图深度学习模型在视频中人体交互动作识别方法 |
| EP3321844B1 (en) * | 2016-11-14 | 2021-04-14 | Axis AB | Action recognition in a video sequence |
| US11568545B2 (en) * | 2017-11-20 | 2023-01-31 | A9.Com, Inc. | Compressed content object and action detection |
| US10528819B1 (en) * | 2017-11-20 | 2020-01-07 | A9.Com, Inc. | Compressed content object and action detection |
| CN107808150A (zh) * | 2017-11-20 | 2018-03-16 | 珠海习悦信息技术有限公司 | 人体视频动作识别方法、装置、存储介质及处理器 |
| CN108280436A (zh) * | 2018-01-29 | 2018-07-13 | 深圳市唯特视科技有限公司 | 一种基于堆叠递归单元的多级残差网络的动作识别方法 |
| CN110163052B (zh) * | 2018-08-01 | 2022-09-09 | 腾讯科技(深圳)有限公司 | 视频动作识别方法、装置和机器设备 |
| US10909424B2 (en) * | 2018-10-13 | 2021-02-02 | Applied Research, LLC | Method and system for object tracking and recognition using low power compressive sensing camera in real-time applications |
| CN109522867A (zh) * | 2018-11-30 | 2019-03-26 | 国信优易数据有限公司 | 一种视频分类方法、装置、设备和介质 |
| US11501532B2 (en) * | 2019-04-25 | 2022-11-15 | International Business Machines Corporation | Audiovisual source separation and localization using generative adversarial networks |
| CN110490055A (zh) * | 2019-07-08 | 2019-11-22 | 中国科学院信息工程研究所 | 一种基于三重编码的弱监督行为识别定位方法和装置 |
| CN110490078B (zh) * | 2019-07-18 | 2024-05-03 | 平安科技(深圳)有限公司 | 监控视频处理方法、装置、计算机设备和存储介质 |
| CN110472531B (zh) * | 2019-07-29 | 2023-09-01 | 腾讯科技(深圳)有限公司 | 视频处理方法、装置、电子设备及存储介质 |
| CN112422863B (zh) * | 2019-08-22 | 2022-04-12 | 华为技术有限公司 | 一种视频拍摄方法、电子设备和存储介质 |
| CN110569814B (zh) * | 2019-09-12 | 2023-10-13 | 广州酷狗计算机科技有限公司 | 视频类别识别方法、装置、计算机设备及计算机存储介质 |
| CN110765860B (zh) * | 2019-09-16 | 2023-06-23 | 平安科技(深圳)有限公司 | 摔倒判定方法、装置、计算机设备及存储介质 |
| CN111080699B (zh) * | 2019-12-11 | 2023-10-20 | 中国科学院自动化研究所 | 基于深度学习的单目视觉里程计方法及系统 |
| WO2021140483A1 (en) * | 2020-01-10 | 2021-07-15 | Everseen Limited | System and method for detecting scan and non-scan events in a self check out process |
-
2020
- 2020-05-20 CN CN202010431706.9A patent/CN111598026B/zh active Active
-
2021
- 2021-04-02 US US17/999,284 patent/US12412426B2/en active Active
- 2021-04-02 EP EP21808837.5A patent/EP4156017B1/en active Active
- 2021-04-02 WO PCT/CN2021/085386 patent/WO2021232969A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20110182469A1 (en) * | 2010-01-28 | 2011-07-28 | Nec Laboratories America, Inc. | 3d convolutional neural networks for automatic human action recognition |
| CN110414335A (zh) * | 2019-06-20 | 2019-11-05 | 北京奇艺世纪科技有限公司 | 视频识别方法、装置及计算机可读存储介质 |
| CN110826545A (zh) * | 2020-01-09 | 2020-02-21 | 腾讯科技(深圳)有限公司 | 一种视频类别识别的方法及相关装置 |
| CN111598026A (zh) * | 2020-05-20 | 2020-08-28 | 广州市百果园信息技术有限公司 | 动作识别方法、装置、设备及存储介质 |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP4156017A4 * |
Cited By (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114827739A (zh) * | 2022-06-06 | 2022-07-29 | 百果园技术(新加坡)有限公司 | 一种直播回放视频生成方法、装置、设备及存储介质 |
| CN115909479A (zh) * | 2022-10-20 | 2023-04-04 | 中国科学院自动化研究所 | 人体行为识别方法、装置、电子设备及可读存储介质 |
| CN116246259A (zh) * | 2022-12-31 | 2023-06-09 | 南京行者易智能交通科技有限公司 | 一种基于3d卷积神经网络的公交安全驾驶行为识别方法 |
| CN116129936A (zh) * | 2023-02-02 | 2023-05-16 | 百果园技术(新加坡)有限公司 | 一种动作识别方法、装置、设备、存储介质及产品 |
| CN117036789A (zh) * | 2023-07-24 | 2023-11-10 | 西安电子科技大学 | 基于动态推理的全局与局部特征融合的骨架行为识别方法 |
| CN117975376A (zh) * | 2024-04-02 | 2024-05-03 | 湖南大学 | 基于深度分级融合残差网络的矿山作业安全检测方法 |
| CN117975376B (zh) * | 2024-04-02 | 2024-06-07 | 湖南大学 | 基于深度分级融合残差网络的矿山作业安全检测方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| US12412426B2 (en) | 2025-09-09 |
| EP4156017B1 (en) | 2025-10-08 |
| EP4156017A4 (en) | 2024-09-25 |
| CN111598026B (zh) | 2023-05-30 |
| CN111598026A (zh) | 2020-08-28 |
| EP4156017A1 (en) | 2023-03-29 |
| US20230196837A1 (en) | 2023-06-22 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN111598026B (zh) | 动作识别方法、装置、设备及存储介质 | |
| Liao et al. | Video-based person re-identification via 3d convolutional networks and non-local attention | |
| Liu et al. | Robust video super-resolution with learned temporal dynamics | |
| CN111062314B (zh) | 图像选取方法、装置、计算机可读存储介质及电子设备 | |
| US20200329233A1 (en) | Hyperdata Compression: Accelerating Encoding for Improved Communication, Distribution & Delivery of Personalized Content | |
| Ma et al. | Joint feature and texture coding: Toward smart video representation via front-end intelligence | |
| CN114973049B (zh) | 一种统一卷积与自注意力的轻量视频分类方法 | |
| CN115171014B (zh) | 视频处理方法、装置、电子设备及计算机可读存储介质 | |
| CN111444370A (zh) | 图像检索方法、装置、设备及其存储介质 | |
| Li et al. | Cross-level parallel network for crowd counting | |
| Wang et al. | Exploring hybrid spatio-temporal convolutional networks for human action recognition | |
| CN112991239A (zh) | 一种基于深度学习的图像反向恢复方法 | |
| CN111402118A (zh) | 图像替换方法、装置、计算机设备和存储介质 | |
| Qi et al. | Long-term temporal context gathering for neural video compression | |
| CN112200816A (zh) | 视频图像的区域分割及头发替换方法、装置及设备 | |
| Peng et al. | Motion boundary emphasised optical flow method for human action recognition | |
| Shen et al. | Codedvision: Towards joint image understanding and compression via end-to-end learning | |
| CN111680618A (zh) | 基于视频数据特性的动态手势识别方法、存储介质和设备 | |
| CN115909479A (zh) | 人体行为识别方法、装置、电子设备及可读存储介质 | |
| CN116543339B (zh) | 一种基于多尺度注意力融合的短视频事件检测方法及装置 | |
| Dai et al. | Slowfast diversity-aware prototype learning for egocentric action recognition | |
| Zhang et al. | Local compressed video stream learning for generic event boundary detection | |
| CN113392269A (zh) | 一种视频分类方法、装置、服务器及计算机可读存储介质 | |
| Li et al. | MFMENet: multi-scale features mutual enhancement network for change detection in remote sensing images | |
| CN116485807A (zh) | 一种数字图像分割方法及装置 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21808837 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2021808837 Country of ref document: EP Effective date: 20221220 |
|
| WWG | Wipo information: grant in national office |
Ref document number: 17999284 Country of ref document: US |
|
| WWG | Wipo information: grant in national office |
Ref document number: 2021808837 Country of ref document: EP |