WO2024082933A1 - 视频处理方法、装置、电子设备及存储介质 - Google Patents
视频处理方法、装置、电子设备及存储介质 Download PDFInfo
- Publication number
- WO2024082933A1 WO2024082933A1 PCT/CN2023/121354 CN2023121354W WO2024082933A1 WO 2024082933 A1 WO2024082933 A1 WO 2024082933A1 CN 2023121354 W CN2023121354 W CN 2023121354W WO 2024082933 A1 WO2024082933 A1 WO 2024082933A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- processed
- frames
- video
- interlaced
- model
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N5/00—Details of television systems
- H04N5/14—Picture signal circuitry for video frequency region
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T5/00—Image enhancement or restoration
- G06T5/50—Image enhancement or restoration using two or more images, e.g. averaging or subtraction
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T5/00—Image enhancement or restoration
- G06T5/60—Image enhancement or restoration using machine learning, e.g. neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T5/00—Image enhancement or restoration
- G06T5/77—Retouching; Inpainting; Scratch removal
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/20—Analysis of motion
- G06T7/246—Analysis of motion using feature-based methods, e.g. the tracking of corners or segments
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N5/00—Details of television systems
- H04N5/222—Studio circuitry; Studio devices; Studio equipment
- H04N5/262—Studio circuits, e.g. for mixing, switching-over, change of character of image, other special effects ; Cameras specially adapted for the electronic generation of special effects
- H04N5/265—Mixing
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N7/00—Television systems
- H04N7/01—Conversion of standards, e.g. involving analogue television standards or digital television standards processed at pixel level
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/10—Image acquisition modality
- G06T2207/10016—Video; Image sequence
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20081—Training; Learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20084—Artificial neural networks [ANN]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/20—Special algorithmic details
- G06T2207/20212—Image combination
- G06T2207/20221—Image fusion; Image merging
Definitions
- the embodiments of the present disclosure relate to the field of video processing technology, for example, to a video processing method, device, electronic device and storage medium.
- interlaced video When displaying interlaced video on an existing display interface, it needs to be de-interlaced to display the complete video.
- de-interlacing is usually performed on interlaced videos to remove the brushing effect in interlaced videos.
- the de-interlacing effect of this method is not good. For example, in scenes with moving objects, the brushed areas are relatively blurred, and it is easy to lose details and have brushed images.
- the present disclosure provides a video processing method, device, electronic device and storage medium to achieve an effect of effectively restoring video images. For example, for video images of motion scenes, a more significant restoration effect can be achieved.
- an embodiment of the present disclosure provides a video processing method, the method comprising:
- a target video is determined.
- an embodiment of the present disclosure further provides a video processing device, the device comprising:
- a module for acquiring interlaced frames to be processed configured to acquire at least three interlaced frames to be processed; wherein the interlaced frames to be processed are determined based on two adjacent video frames to be processed;
- a target video frame determination module configured to input the at least three interlaced frames to be processed into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three interlaced frames to be processed; wherein the image fusion model includes a feature processing sub-model and a motion perception sub-model;
- the target video determination module is configured to determine the target video based on the at least two target video frames.
- an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:
- processors one or more processors
- a storage device configured to store one or more programs
- the one or more processors When the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method as described in any one of the embodiments of the present disclosure.
- the embodiments of the present disclosure further provide a storage medium comprising computer executable instructions, which, when executed by a computer processor, are used to execute the video processing method as described in any one of the embodiments of the present disclosure.
- FIG1 is a schematic flow chart of a video processing method provided by an embodiment of the present disclosure.
- FIG2 is a schematic diagram of an image fusion model provided by an embodiment of the present disclosure.
- FIG3 is a schematic diagram of a motion perception model provided by an embodiment of the present disclosure.
- FIG4 is a schematic flow chart of a video processing method provided by an embodiment of the present disclosure.
- FIG5 is a schematic diagram of a video frame to be processed provided by an embodiment of the present disclosure.
- FIG6 is a schematic diagram of the structure of a video processing device provided by an embodiment of the present disclosure.
- FIG. 7 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure.
- a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information.
- the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
- the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form.
- the pop-up window may also carry a selection control for the user to select "agree” or “disagree” to provide personal information to the electronic device.
- the first implementation method is to perform deinterlacing on the interlaced frames to be processed based on the deinterlacing algorithm (YADIF) to remove the brushed effect of the original video.
- YADIF deinterlacing algorithm
- the second implementation method is to input multiple interlaced frames to be processed into the ST-Deint deep learning neural network model that combines time domain and spatial domain information prediction, and process the interlaced frames to be processed based on the deep learning algorithm.
- the third implementation method is to process the interlaced frames to be processed based on the deep learning model DIN, so that the interlaced frames to be processed first fill in the missing information and then fuse the interfield content to obtain the processed video. Similar to the second implementation method, this method can only achieve a rough restoration effect for motion scenes, and the detail restoration effect is not good. For example, the motion brushed area is prone to blurring and loss of details. Based on the above, it can be seen that the video processing method in the related art still has the problem of poor output video display effect.
- the interlaced frames to be processed can be processed based on the image fusion model including multiple sub-models, thereby avoiding the output target video from missing details, screen brushing, and blurred images.
- FIG1 is a flow chart of a video processing method provided by an embodiment of the present disclosure.
- the embodiment of the present disclosure is applicable to the situation where feature information is supplemented for an original video using interlaced scanning so that the obtained target video can be fully displayed on an existing display device.
- the method can be executed by a video processing device, which can be implemented in the form of software and/or hardware, for example, by an electronic device, which can be a mobile terminal, a PC or a server, etc.
- the technical solution provided by the embodiment of the present disclosure can be executed based on a client, can be executed based on a server, or can be executed based on the cooperation of a client and a server.
- the method comprises:
- the interlaced frame to be processed is determined based on two adjacent video frames to be processed.
- the device for executing the video processing method can be integrated into an application software that supports the video processing function, and the software can be installed in an electronic device, for example, the electronic device can be a mobile terminal or a PC.
- the application software can be a type of software for image/video processing, and its specific application software will not be described one by one here, as long as the image/video processing can be achieved. It can also be a specially developed application program to implement video processing and display the output video in the software, or it can be integrated in the corresponding page, and the user can realize the processing of special effect videos through the page integrated in the PC.
- the user can shoot a video in real time based on the camera device of the mobile terminal, or actively upload a video based on the pre-developed control in the application software. Therefore, it can be understood that the real-time video captured by the application or the video actively uploaded by the user is the video to be processed. For example, based on the pre-written program, a plurality of video frames to be processed can be obtained.
- early video display methods usually adopt an interlaced scanning method, that is, first scan the odd rows to obtain a video frame in which only the odd rows of pixels have rendered pixel values, and then scan the even rows to obtain a video frame in which only the even rows of pixels have rendered pixel values, and combine the two video frames to obtain a complete video frame.
- This display method will cause a large time interval between the display of two adjacent video frames, resulting in large flickering of the video frame, jagged lines, false images and other image quality problems.
- the current video display usually adopts a line-by-line scanning method
- the interlaced video frames can be used as the video frames to be processed, that is, only odd-numbered rows of pixels have rendered pixel values or even-numbered rows of pixels have rendered pixel values.
- the video frame combined by two adjacent frames of video frames to be processed is the interlaced frame to be processed.
- de-interlacing refers to filling the missing half-field information of the odd and even fields of the interlaced frames of two adjacent frames of images to restore the original frame size, and finally obtaining odd and even frames.
- the video frame to be processed is a video frame in which only odd-numbered rows of pixels have rendered pixel values or even-numbered rows of pixels have rendered pixel values
- two adjacent frames of the video frame to be processed can be combined to obtain an interlaced frame to be processed, so that the interlaced frame to be processed can be de-interlaced.
- I 1 , I 2 , I 3 , I 4 , I 5 , and I 6 can be used as video frames to be processed.
- I 1 and I 2 can be combined to obtain a frame of interlaced frame to be processed D 1
- I 3 and I 4 can be combined to obtain a frame of interlaced frame to be processed D 2
- I 5 and I 6 can be combined to obtain a frame of video frame to be processed D 3 .
- the principle of making the interlaced frame to be processed can be expressed based on the following formula:
- the number of interlaced frames to be processed may be three or more than three, and this is not specifically limited in the embodiments of the present disclosure.
- the number of interlaced frames to be processed corresponds to the number of video frames of the original video
- the number of interlaced frames to be processed input into the model can be three frames or more than three frames.
- the interlaced frames to be processed can be input into a pre-trained image fusion model.
- the image fusion model can be a deep learning neural network model including multiple sub-models.
- the image fusion model includes a feature processing sub-model and a motion perception sub-model.
- the number of interlaced frames to be processed is at least three frames.
- the feature processing submodel can be a neural network model including multiple convolution modules.
- the feature processing submodel can be used to extract, fuse and perform other processing on the features in the interlaced frames to be processed.
- the feature processing submodel can include multiple 3D convolution layers, so that the feature processing submodel can not only process the time domain feature information of multiple frames, but also process the spatial feature information, thereby strengthening the information interaction between the interlaced frames to be processed.
- the motion perception submodel may be a neural network model for perceiving inter-frame motion.
- the motion perception submodel may be composed of at least one convolutional network, a network including a backward warping function, and a residual network.
- the backward warping function may realize the mapping between images.
- the motion perception submodel may be used to process the feature information between frames, thereby making the inter-frame content more continuous and also achieving the effect of complementing each other's details.
- the video frame to be processed can be processed based on multiple sub-models in the image fusion model, so as to obtain at least two target video frames corresponding to the interlaced frame to be processed.
- the image fusion model includes multiple sub-models
- the multiple sub-models in the model can be used in turn to perform corresponding processing on the video frames to be processed, thereby outputting at least two target video frames corresponding to the interlaced frames to be processed.
- the image fusion model includes multiple sub-models, and the arrangement order of the multiple sub-models can be arranged according to the data input and output order.
- the image fusion model includes a feature processing sub-model, a motion perception sub-model, and a 2D convolution layer, wherein the 2D convolution layer may be a neural network layer that performs feature processing on only the height and width of the data.
- determining the arrangement order of multiple sub-models in the image fusion model based on the data input and output order can enable the image fusion model to not only process the feature information of the interlaced frame to be processed, but also perceive the motion between multiple interlaced frames to be processed, so as to make the content between frames more continuous and achieve the effect of detail supplementation.
- the solution adopted by the related technology is to split the interlaced frame to be processed into odd and even rows, that is, to halve the H dimension of the interlaced frame to be processed. For example, if the matrix of the interlaced frame to be processed is (H ⁇ W ⁇ C), then the matrix after the odd and even row split is (2/H ⁇ W ⁇ C). Such a solution may cause the objects in the interlaced frame to be processed to be deformed in structure, thereby affecting the visual effect of the target video frame.
- the processing process of the embodiment of the present disclosure can be understood as, on the basis of splitting the interlaced frame to be processed into odd and even rows, it is also split into odd and even columns, that is, a dual feature processing branch is adopted, so as to ensure that when the interlaced frame to be processed is processed based on the image fusion model, it can not only process the overall structural feature information of the interlaced frame to be processed, but also process the high-frequency detail feature information of the interlaced frame to be processed.
- the feature processing sub-model includes a first feature extraction branch and a second feature extraction branch; the output of the first feature extraction branch is the input of the first motion perception sub-model in the motion perception sub-model, and the output of the second feature extraction branch is the input of the second motion perception sub-model in the motion perception sub-model; the output of the first motion perception sub-model and the output of the second motion perception sub-model are the input of the 2D convolution layer, so that the 2D convolution layer outputs the target Label the video frame.
- the first feature extraction branch can be a neural network model for processing the structural features of the interlaced frame to be processed.
- the first feature extraction branch includes a structural feature extraction network and a structural feature fusion network.
- the structural feature extraction network can be composed of at least one convolutional network, so that at least one convolutional network can process the interlaced frame to be processed according to a preset structural splitting ratio to obtain structural features corresponding to the interlaced frame to be processed.
- the structural feature fusion network can be a neural network of a U-Net structure stacked by at least one 3D convolutional layer. It should be noted that the convolution kernels of at least one 3D convolutional layer can be the same value or different values, and the present embodiment does not specifically limit this.
- the structural feature fusion network can be used to strengthen the information interaction between frames, so that not only the spatial features of the interlaced frame to be processed can be processed, but also the time domain features between multiple frames can be strengthened.
- the second feature extraction branch may be a neural network model for processing detail features of the interlaced frame to be processed.
- the second feature extraction branch includes a detail feature extraction network and a detail feature fusion network.
- the detail feature extraction network may be composed of at least one convolutional layer, so that at least one convolutional layer can process the interlaced frame to be processed according to a preset detail splitting ratio to obtain detail features corresponding to the interlaced frame to be processed.
- the detail feature fusion network may be a neural network of a U-Net structure stacked by at least one 3D convolutional layer. It should be noted that the convolution kernels of at least one 3D convolutional layer may be the same value or different values, and the disclosed embodiment does not specifically limit this.
- the interlaced frames to be processed are respectively input into the first feature extraction branch and the second feature extraction branch.
- the interlaced frames to be processed are processed by the structural feature extraction network and the structural feature fusion network in the first feature extraction branch, they can be input into the first motion perception sub-model.
- the interlaced frames to be processed are processed by the detail feature extraction network and the detail feature fusion network in the second feature extraction branch, they can be input into the second motion perception sub-model.
- the model input is processed by the first motion perception sub-model, it can be input into the 2D convolution layer.
- the model input is processed by the second motion perception sub-model, it is input into the 2D convolution layer so that the 2D convolution layer can output the target video frame.
- the image fusion model can process the feature information of the interlaced frames to be processed and perceive the motion between the interlaced frames to be processed, so that the content between frames is more continuous and the effect of detail supplementation is achieved.
- the interlaced frame to be processed is input into the image fusion model, and it can be processed based on multiple sub-models in the model, so as to obtain the target video frame corresponding to the interlaced frame to be processed.
- the following is a detailed description of the process of the image fusion model processing the interlaced frame to be processed in conjunction with Figure 2.
- At least three interlaced frames to be processed are input into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three interlaced frames to be processed, including: performing equal-proportion feature extraction on the at least three interlaced frames to be processed based on a structural feature extraction network to obtain structural features corresponding to the interlaced frames to be processed; and performing even-odd field feature extraction on the at least three interlaced frames to be processed based on a detail feature extraction network to obtain detail features corresponding to the interlaced frames to be processed; processing the structural features based on a structural feature fusion network to obtain a first inter-frame feature map between two adjacent interlaced frames to be processed; and processing the detail features based on a detail feature fusion network to obtain a second inter-frame feature map between two adjacent interlaced frames to be processed; processing the first inter-frame feature map based on a first motion perception sub-model to obtain a first fused feature map; and processing the second inter-frame feature map
- the structural feature may be a feature for reflecting the overall structural information of the interlaced frame to be processed.
- the detail feature may be a feature for reflecting the detail information of the interlaced frame to be processed.
- the detail feature may be a high-frequency feature, which is a higher-order feature than the structural feature.
- the interlaced frame to be processed is input into the image fusion model, and the interlaced frame to be processed can be subjected to proportional dimensionality reduction processing based on the structural feature extraction network to obtain the structural features corresponding to the interlaced frame to be processed.
- the detail feature extraction network performs an odd-even field splitting process on the interlaced frame to be processed to obtain detail features corresponding to the interlaced frame to be processed; for example, for structural features, feature fusion processing can be performed on the structural features based on the structural feature fusion network, and the structural features of two adjacent interlaced frames to be processed are fused, so that a fusion feature map between the two adjacent interlaced frames to be processed can be obtained, that is, the first inter-frame feature map; at the same time, for detail features, detail features can be fused based on the detail feature fusion network, and the detail features of two adjacent interlaced frames to be processed are fused, so that a fusion feature map between the two adjacent interlaced frames to be processed can be obtained.
- the first inter-frame feature map is input into the first motion perception sub-model, and the first inter-frame feature map is processed based on the first motion perception sub-model to obtain the first fused feature map.
- the second inter-frame feature map is input into the second motion perception sub-model, and the second inter-frame feature map is processed based on the second motion perception sub-model to obtain the second fused feature map.
- the first fused feature map and the second fused feature map are input into the 2D convolution layer, and the fused feature map is processed based on the 2D convolution layer to obtain at least two target video frames corresponding to the interlaced frame to be processed.
- the dual feature processing branch enables the image fusion model to process both the overall structural feature information of the interlaced frame to be processed and the detailed feature information of the interlaced frame to be processed.
- the 3D convolution layer can be used to strengthen the information interaction between frames.
- the motion perception sub-model can be used to perceive the motion between frames and perform feature alignment, so that the content between frames is more continuous, thereby improving the display effect of the target video frame.
- the first inter-frame feature map may include a first feature map and a second feature map.
- the first inter-frame feature map is processed based on the first motion perception sub-model, and the first feature map and the second feature map may be processed separately based on the first motion perception sub-model to obtain a first fused feature map.
- the processing process of the first inter-frame feature map by the first motion perception sub-model may be specifically described below in conjunction with FIG. 3 .
- the first inter-frame feature map is processed based on the first motion perception sub-model to obtain a first fused feature map, including: based on the convolutional network in the first motion perception sub-model, the first feature map and the second feature map are processed respectively to obtain a first optical flow map and a second optical flow map; based on the distortion network in the first motion perception sub-model, the first optical flow map and the second optical flow map are mapped and processed to obtain an offset; based on the first optical flow map, the second optical flow map and the offset, the first fused feature map is determined.
- an optical flow graph can represent the speed and direction of movement of each pixel in two adjacent frames of an image.
- Optical flow is the instantaneous speed of the pixel movement of a moving object in space on the observation imaging plane. It is a method of finding the correspondence between the previous frame and the current frame by using the change of pixels in the time domain in an image sequence and the correlation between connected frames, thereby calculating the motion information of the object between adjacent frames.
- the distortion network can be a network containing a backward warping function, which can realize the mapping between images.
- the offset can be data obtained after mapping based on the optical flow graph, which is used to represent the feature displacement offset.
- the first inter-frame feature map is input into the first motion perception sub-model, and the first feature map and the second feature map can be processed based on the convolutional network, respectively, so as to obtain a first optical flow map for characterizing the motion speed and motion direction of the pixels in the two adjacent interlaced frames to be processed corresponding to the first feature map, and a second optical flow map for characterizing the motion speed and motion direction of the pixels in the two adjacent interlaced frames to be processed corresponding to the second feature map.
- the first optical flow map and the second optical flow map are mapped based on the distortion network to obtain the offset corresponding to the first optical flow map and the offset corresponding to the second optical flow map.
- the first optical flow map, the second optical flow map and the offsets corresponding to the two optical flow maps are fused to obtain the first fused feature map.
- a first fused feature map is determined, including: residual processing of the first optical flow map and the offset to obtain a first feature map to be spliced; residual processing of the second optical flow map and the offset to obtain a second feature map to be spliced; and splicing of the first feature map to be spliced and the second feature map to be spliced to obtain a first fused feature map.
- residual processing may be performed on the first optical flow map and the offset to align multiple optical flow features in the first optical flow map, thereby obtaining feature alignment.
- the first feature map to be spliced is obtained, and at the same time, the second optical flow map and the offset are subjected to residual processing to align multiple optical flow features in the second optical flow map, thereby obtaining the second feature map to be spliced after feature alignment.
- the first feature map to be spliced and the second feature map to be spliced are spliced to obtain the first fused feature map.
- processing process of the second inter-frame feature map based on the second motion perception sub-model is the same as the processing process of the first inter-frame feature map based on the first motion perception sub-model, and the embodiment of the present disclosure will not be described in detail here.
- the process of processing the first inter-frame feature map by the first motion perception sub-model is described by taking three interlaced frames to be processed as an example.
- D 1 , D 2 , and D 3 can be used as interlaced frames to be processed, and these three interlaced frames to be processed are input into the first feature extraction branch to obtain the first feature map and the second feature map, which can be represented by F 1 and F 2.
- F 1 and F 2 are input into the first motion perception sub-model, and F 1 and F 2 are processed based on the convolution layer to obtain the first optical flow map IF 1 and the second optical flow map IF 2 .
- IF 1 and IF 2 are mapped based on the distortion network to obtain the offset, and IF 1 and the offset are subjected to residual processing to obtain the first feature map to be spliced.
- Perform residual processing on IF 2 and the offset to obtain the second feature map to be spliced Finally, and After the splicing process is performed, the first fused feature map F full can be obtained.
- S130 Determine a target video based on at least two target video frames.
- the target video frames can be spliced, so as to obtain a target video composed of multiple continuous target video frames.
- determining the target video based on at least two target video frames includes: splicing the at least two target video frames in the time domain to obtain the target video.
- the application can splice multiple video frames according to the timestamps corresponding to the target video frames to obtain the target video. It can be understood that by splicing multiple frames and generating a target video, the processed images can be displayed in a clear and coherent form.
- the video can be played directly to display the processed video screen on the display interface, or the target video can be stored in a specific space according to a preset path.
- the embodiments of the present disclosure do not specifically limit this.
- the at least three interlaced frames to be processed can be input into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three interlaced frames to be processed, and finally, based on the at least two target video frames, the target video is determined.
- the restoration effect of the video picture can be effectively improved. For example, for the video picture of a motion scene, a more significant restoration effect can also be achieved.
- the picture brushing and detail loss are avoided, the picture quality and clarity of the video picture are improved, and the user experience is enhanced.
- FIG4 is a flow chart of a video processing method provided by an embodiment of the present disclosure.
- a plurality of video frames to be processed in the original video can be processed to obtain an interlaced frame to be processed.
- the exemplary implementation method thereof can refer to the technical solution of this embodiment.
- the technical terms identical or corresponding to the above-mentioned embodiment are not repeated here.
- the method comprises the following steps:
- S210 Acquire a plurality of to-be-processed video frames corresponding to the original video.
- the two to-be-processed video frames include an odd-numbered video frame and an even-numbered video frame, and the odd-numbered video frame and the even-numbered video frame are determined based on the order of the to-be-processed video frames in the original video.
- the original video may be a video spliced from interlaced scanned video frames.
- the original video may be a video captured in real time by a terminal device, or a video pre-stored in a storage space by an application software, or a video uploaded to a server or client by a user based on a pre-set video upload control, etc., which is not specifically limited in this embodiment of the present disclosure.
- the original video may be an early image video.
- the odd-numbered video frames may be the original video.
- the number corresponding to the arrangement order in the original video is an odd number, and there are rendering pixel values for the odd-numbered rows of pixels, which can be rendered and displayed on the display interface, and the pixel values of the even-numbered rows of pixels may be preset values, and the video frames are displayed in the form of black holes in the display interface.
- the even-numbered video frames may be the number corresponding to the arrangement order in the original video is an even number, and there are rendering pixel values for the even-numbered rows of pixels, which can be rendered and displayed on the display interface, and the pixel values of the odd-numbered rows of pixels may be preset values, and the video frames are displayed in the form of black holes in the display interface.
- Figure 5a can be an odd video frame
- Figure 5b can be an even video frame.
- the odd video frame only scans and samples the odd rows. Therefore, in Figure 5a, only the pixel values of the pixels in the odd rows are rendering pixel values, which can be blue, and are rendered and displayed in the display interface, while the pixel values of the pixels in the even rows can be preset values, which can be black. At this time, when the pixels in the even rows are displayed in the display interface, they may be displayed in the display interface as a black hole; similarly, the even video frame only scans and samples the even rows.
- the original video can be parsed based on a pre-written program to obtain multiple video frames to be processed. For example, starting from the first video frame to be processed, two adjacent video frames to be processed are fused to obtain an interlaced frame to be processed.
- fusing two adjacent video frames to be processed to obtain an interlaced frame to be processed includes: extracting odd-numbered line data from an odd-numbered video frame and even-numbered line data from an even-numbered video frame; and fusing the odd-numbered line data and the even-numbered line data to obtain an interlaced frame to be processed.
- the odd-numbered row data may be pixel point information in the odd-numbered row.
- the even-numbered row data may be pixel point information in the even-numbered row.
- the pixel point information of the odd-numbered row may be sampled first to obtain an odd-numbered video frame, and then the pixel point information of the even-numbered row may be sampled to obtain an even-numbered video frame, and the pixel point sampling information of the odd-numbered row in the odd-numbered video frame may be used as the odd-numbered row data, and the pixel point sampling information of the even-numbered row in the even-numbered video frame may be used as the even-numbered row data.
- the odd-numbered row data of the odd-numbered video frame can be extracted, and the even-numbered row data of the even-numbered video frame can be extracted.
- the odd-numbered row data and the even-numbered row data are fused to obtain the interlaced frame to be processed.
- the interlaced frame to be processed containing both the pixel information of the odd-numbered row pixels and the pixel information of the even-numbered row pixels can be obtained, so that the target video frame that meets the user's needs can be obtained by processing the interlaced frame to be processed.
- S240 Input at least three interlaced frames to be processed into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three interlaced frames to be processed.
- S250 Determine a target video based on at least two target video frames.
- the technical solution of the disclosed embodiment obtains multiple to-be-processed video frames corresponding to the original video, fuses two adjacent to-be-processed video frames to obtain to-be-processed interlaced frames, then obtains at least three to-be-processed interlaced frames, and inputs the at least three to-be-processed interlaced frames into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three to-be-processed interlaced frames, and finally, determines the target video based on the at least two target video frames.
- the restoration effect of the video picture can be effectively improved. For example, for the video picture of a motion scene, a more significant restoration effect can also be achieved.
- the picture brushing and detail loss are avoided, the picture quality and clarity of the video picture are improved, and the user experience is improved.
- FIG6 is a schematic diagram of the structure of a video processing device provided by an embodiment of the present disclosure. As shown in FIG6 , the device includes: a to-be-processed interlaced frame acquisition module 310 , a target video frame determination module 320 , and a target video determination module 330 .
- the to-be-processed interlaced frame acquisition module 310 is configured to acquire at least three to-be-processed interlaced frames; wherein the to-be-processed interlaced frames are determined based on two adjacent to-be-processed video frames;
- the target video frame determination module 320 is configured to input the at least three interlaced frames to be processed into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three interlaced frames to be processed; wherein the image fusion model includes a feature processing sub-model and a motion perception sub-model;
- the target video determination module 330 is configured to determine the target video based on the at least two target video frames.
- the device further includes: a to-be-processed video frame acquisition module and a to-be-processed video frame processing module.
- a module for acquiring video frames to be processed configured to acquire a plurality of video frames to be processed corresponding to the original video before acquiring at least three interlaced frames to be processed;
- the processing module of the video frames to be processed is configured to fuse two adjacent video frames to be processed to obtain the interlaced frame to be processed; wherein the two video frames to be processed include an odd video frame and an even video frame, and the odd video frame and the even video frame are determined based on the order of the video frames to be processed in the original video.
- the to-be-processed video frame processing module includes: a data extraction unit and a data processing unit.
- a data extraction unit configured to extract odd-numbered line data in the odd-numbered video frame and even-numbered line data in the even-numbered video frame;
- the data processing unit is configured to obtain the to-be-processed interlaced frame by fusing the odd-numbered line data and the even-numbered line data.
- the image fusion model includes a feature processing sub-model, a motion perception sub-model and a 2D convolution layer.
- the feature processing sub-model includes a first feature extraction branch and a second feature extraction branch;
- the output of the first feature extraction branch is the input of the first motion perception sub-model in the motion perception sub-model
- the output of the second feature extraction branch is the input of the second motion perception sub-model in the motion perception sub-model
- the output of the first motion perception sub-model and the output of the second motion perception sub-model are inputs to the 2D convolutional layer, so that the 2D convolutional layer outputs a target video frame.
- the first feature extraction branch includes a structural feature extraction network and a structural feature fusion network
- the second feature extraction branch includes a detail feature extraction network and a detail feature fusion network
- the target video frame determination module 320 includes: an equal-proportional feature extraction submodule, an odd-even field feature extraction submodule, a structural feature processing submodule, a detail feature processing submodule, a first fusion feature map determination submodule, a second fusion feature map determination submodule and a target video frame determination submodule.
- an equal-proportional feature extraction submodule configured to perform equal-proportional feature extraction on the at least three interlaced frames to be processed based on the structural feature extraction network to obtain structural features corresponding to the interlaced frames to be processed;
- an odd-even field feature extraction submodule configured to extract odd-even field features from the at least three interlaced frames to be processed based on the detail feature extraction network, to obtain detail features corresponding to the interlaced frames to be processed;
- a structural feature processing submodule configured to process the structural feature based on the structural feature fusion network to obtain a first inter-frame feature map between two adjacent interlaced frames to be processed;
- a detail feature processing submodule configured to process the detail feature based on the detail feature fusion network to obtain a second inter-frame feature map between two adjacent interlaced frames to be processed;
- a first fused feature map determining submodule is configured to process the first inter-frame feature map based on a first motion perception submodel to obtain a first fused feature map
- a second fused feature map determining submodule configured to process the second inter-frame feature map based on a second motion perception submodel to obtain a second fused feature map
- the target video frame determination submodule is configured to perform a 2D convolutional layer on the first fusion feature map and the first The two fused feature maps are processed to obtain the at least two target video frames.
- the first fused feature map determination submodule includes: a feature map processing unit, an optical flow map mapping processing unit and a first fused feature map determination unit.
- a feature map processing unit configured to process the first feature map and the second feature map respectively based on the convolutional network in the first motion perception sub-model to obtain a first optical flow map and a second optical flow map;
- an optical flow map mapping processing unit configured to map and process the first optical flow map and the second optical flow map based on the distortion network in the first motion perception sub-model to obtain an offset
- the first fused feature map determining unit is configured to determine the first fused feature map based on the first optical flow map, the second optical flow map and the offset.
- the first fused feature map determining unit is configured to process the first optical flow map and the offset residual to obtain a first feature map to be spliced; process the second optical flow map and the offset residual to obtain a second feature map to be spliced; and obtain the first fused feature map by splicing the first feature map to be spliced and the second feature map to be spliced.
- the target video determination module 330 is configured to splice the at least two target video frames in the time domain to obtain the target video.
- the at least three interlaced frames to be processed can be input into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three interlaced frames to be processed, and finally, based on the at least two target video frames, the target video is determined.
- the restoration effect of the video picture can be effectively improved. For example, for the video picture of a motion scene, a more significant restoration effect can also be achieved.
- the picture brushing and detail loss are avoided, the picture quality and clarity of the video picture are improved, and the user experience is enhanced.
- the video processing device provided in the embodiments of the present disclosure can execute the video processing method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
- FIG7 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure.
- a schematic diagram of the structure of an electronic device e.g., a terminal device or server in FIG7
- the terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
- the electronic device shown in FIG7 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.
- the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503.
- a processing device e.g., a central processing unit, a graphics processing unit, etc.
- RAM random access memory
- various programs and data required for the operation of the electronic device 500 are also stored.
- the processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504.
- An edit/output (I/O) interface 505 is also connected to the bus 504.
- the following devices may be connected to the I/O interface 505: input devices 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 508 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 509.
- the communication device 509 may allow the electronic device 500 to communicate wirelessly or wired with other devices to exchange data.
- FIG. 7 shows an electronic device 500 with a variety of devices, it should be understood that it is not required to implement or have all the devices shown. More or fewer devices may be implemented or have alternatively.
- an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart.
- the computer program product includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart.
- the computer program may be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502.
- the electronic device provided by the embodiment of the present disclosure and the video processing method provided by the above embodiment belong to the same concept.
- the technical details not fully described in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
- the embodiments of the present disclosure provide a computer storage medium on which a computer program is stored.
- the program is executed by a processor, the video processing method provided by the above embodiments is implemented.
- the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two.
- the computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above.
- Computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
- a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device.
- a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried.
- This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above.
- the computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device.
- the program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
- the client and server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network).
- HTTP HyperText Transfer Protocol
- Examples of communication networks include a local area network ("LAN”), a wide area network ("WAN”), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
- the computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
- the computer-readable medium carries one or more programs.
- the electronic device When the one or more programs are executed by the electronic device, the electronic device:
- the computer-readable medium carries one or more programs.
- the electronic device When the one or more programs are executed by the electronic device, the electronic device:
- a target video is determined.
- Computer program code for performing operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages.
- the program code may be executed entirely on a user's computer, partially on a user's computer, as a stand-alone software package, or partially on a computer.
- the program on the user's computer may be executed partially on the remote computer or completely on the remote computer or server.
- the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
- LAN local area network
- WAN wide area network
- each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function.
- the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved.
- each square box in the block diagram and/or flow chart, and the combination of the square boxes in the block diagram and/or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
- the units involved in the embodiments described in the present disclosure may be implemented by software or hardware.
- the name of a unit does not limit the unit itself in some cases.
- the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses".
- exemplary types of hardware logic components include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
- FPGAs field programmable gate arrays
- ASICs application specific integrated circuits
- ASSPs application specific standard products
- SOCs systems on chip
- CPLDs complex programmable logic devices
- a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment.
- a machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
- a machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing.
- a more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
- RAM random access memory
- ROM read-only memory
- EPROM or flash memory erasable programmable read-only memory
- CD-ROM portable compact disk read-only memory
- CD-ROM compact disk read-only memory
- magnetic storage device or any suitable combination of the foregoing.
- Example 1 provides a video processing method, the method comprising:
- a target video is determined.
- Example 2 provides a video processing method, and before obtaining at least three interlaced frames to be processed, the method further includes:
- the two video frames to be processed include an odd video frame and an even video frame, and the odd video frame and the even video frame are determined based on the order of the video frames to be processed in the original video.
- Example 3 provides a video processing method, wherein the method of fusing two adjacent video frames to be processed to obtain the interlaced frame to be processed includes:
- the interlaced frame to be processed is obtained by fusing the odd-numbered line data and the even-numbered line data.
- Example 4 provides a video processing method, wherein the image fusion model includes a feature processing sub-model, a motion perception sub-model and a 2D convolution layer.
- Example 5 provides a video processing method, further comprising:
- the feature processing sub-model includes a first feature extraction branch and a second feature extraction branch;
- the output of the first feature extraction branch is the input of the first motion perception sub-model in the motion perception sub-model
- the output of the second feature extraction branch is the input of the second motion perception sub-model in the motion perception sub-model
- the output of the first motion perception sub-model and the output of the second motion perception sub-model are inputs to the 2D convolutional layer, so that the 2D convolutional layer outputs a target video frame.
- Example Six provides a video processing method, wherein the first feature extraction branch includes a structural feature extraction network and a structural feature fusion network, and the second feature extraction branch includes a detail feature extraction network and a detail feature fusion network.
- Example 7 provides a video processing method, wherein the step of inputting the at least three interlaced frames to be processed into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three interlaced frames to be processed includes:
- the first fused feature map and the second fused feature map are processed based on the 2D convolution layer to obtain the at least two target video frames.
- Example 8 provides a video processing method, wherein the first inter-frame feature map includes a first feature map and a second feature map, and the first inter-frame feature map is processed based on the first motion perception sub-model to obtain a first fused feature map, including:
- the first feature map and the second feature map are processed respectively to obtain a first optical flow map and a second optical flow map;
- the first fusion feature map is determined based on the first optical flow map, the second optical flow map, and the offset.
- Example 9 provides a video processing method, wherein the determining the first fusion feature map based on the first optical flow map, the second optical flow map, and the offset includes:
- the first fused feature map is obtained by splicing the first feature map to be spliced and the second feature map to be spliced.
- Example 10 provides a video processing method, wherein determining a target video based on the at least two target video frames includes:
- the at least two target video frames are spliced in the time domain to obtain the target video.
- Example 11 provides a video processing device, the device comprising:
- a to-be-processed interlaced frame acquisition module is configured to acquire at least three to-be-processed interlaced frames; wherein the to-be-processed interlaced frames are determined based on two adjacent to-be-processed video frames;
- the target video frame determination module is configured to input the at least three interlaced frames to be processed into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three interlaced frames to be processed; wherein the The image fusion model includes a feature processing sub-model and a motion perception sub-model;
- the target video determination module is configured to determine the target video based on the at least two target video frames.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Image Analysis (AREA)
- Television Systems (AREA)
Abstract
Description
Claims (13)
- 一种视频处理方法,包括:获取至少三个待处理交错帧;其中,所述待处理交错帧是基于相邻两个待处理视频帧确定的;将所述至少三个待处理交错帧输入至预先训练得到的图像融合模型中,得到与所述至少三个待处理交错帧所对应的至少两个目标视频帧;其中,所述图像融合模型中包括特征处理子模型以及运动感知子模型;基于所述至少两个目标视频帧,确定目标视频。
- 根据权利要求1所述的方法,在所述获取至少三个待处理交错帧之前,还包括:获取与原始视频相对应的多个待处理视频帧;对相邻两个待处理视频帧进行融合处理,得到所述待处理交错帧;其中,两个所述待处理视频帧中包括一个奇数视频帧和一个偶数视频帧,所述奇数视频帧和所述偶数视频帧是基于所述待处理视频帧在所述原始视频中的顺序确定的。
- 根据权利要求2所述的方法,其中,所述对相邻两个待处理视频帧进行融合处理,得到所述待处理交错帧,包括:提取所述奇数视频帧中的奇数行数据以及所述偶数视频帧中的偶数行数据;通过对所述奇数行数据以及所述偶数行数据进行融合处理,得到所述待处理交错帧。
- 根据权利要求1所述的方法,其中,所述图像融合模型还包括2D卷积层。
- 根据权利要求4所述的方法,还包括:所述特征处理子模型中包括第一特征提取分支和第二特征提取分支;所述第一特征提取分支的输出为所述运动感知子模型中第一运动感知子模型的输入,所述第二特征提取分支的输出为所述运动感知子模型中第二运动感知子模型的输入;所述第一运动感知子模型的输出和所述第二运动感知子模型的输出为所述2D卷积层的输入,以使所述2D卷积层输出目标视频帧。
- 根据权利要求5所述的方法,其中,所述第一特征提取分支包括结构特征提取网络和结构特征融合网络,所述第二特征提取分支包括细节特征提取网络和细节特征融合网络。
- 根据权利要求6所述的方法,其中,所述将所述至少三个待处理交错帧输入至预先训练得到的图像融合模型中,得到与所述至少三个待处理交错帧所对应的至少两个目标视频帧,包括:基于所述结构特征提取网络对所述至少三个待处理交错帧进行等比例特征提取,得到与所述待处理交错帧所对应的结构特征;基于所述细节特征提取网络对所述至少三个待处理交错帧进行奇偶场特征提取,得到与所述待处理交错帧所对应的细节特征;基于所述结构特征融合网络对所述结构特征进行处理,得到相邻两个待处理交错帧之间的第一帧间特征图;以及,基于所述细节特征融合网络对所述细节特征进行处理,得到相邻两个待处理交错帧之间的第二帧间特征图;基于所述第一运动感知子模型对所述第一帧间特征图进行处理,得到第一融合特征图;以及,基于所述第二运动感知子模型对所述第二帧间特征图进行处理,得到第二融合特征图;基于所述2D卷积层对所述第一融合特征图以及所述第二融合特征图进行处理,得到所述至少两个目标视频帧。
- 根据权利要求7所述的方法,其中,所述第一帧间特征图包括第一特征图和第二特征图,所述基于第一运动感知子模型对所述第一帧间特征图进行处理,得到第一融合特征图,包括:基于所述第一运动感知子模型中的卷积网络分别对所述第一特征图和所述第二特征图进行处理,得到第一光流图和第二光流图;基于所述第一运动感知子模型中的畸变网络对所述第一光流图和所述第二光流图进行映射处理,得到偏移量;基于所述第一光流图、第二光流图以及所述偏移量,确定所述第一融合特征图。
- 根据权利要求8所述的方法,其中,所述基于所述第一光流图、第二光流图以及所述偏移量,确定所述第一融合特征图,包括:对所述第一光流图和所述偏移量进行残差处理,得到第一待拼接特征图;对所述第二光流图和所述偏移量进行残差处理,得到第二待拼接特征图;通过对所述第一待拼接特征图和所述第二待拼接特征图进行拼接处理,得到所述第一融合特征图。
- 根据权利要求1所述的方法,其中,所述基于所述至少两个目标视频帧,确定目标视频,包括:对所述至少两个目标视频帧在时域上进行拼接处理,得到所述目标视频。
- 一种视频处理装置,包括:待处理交错帧获取模块,设置为获取至少三个待处理交错帧;其中,所述待处理交错帧是基于相邻两个待处理视频帧确定的;目标视频帧确定模块,设置为将所述至少三个待处理交错帧输入至预先训练得到的图像融合模型中,得到与所述至少三个待处理交错帧所对应的至少两个目标视频帧;其中,所述图像融合模型中包括特征处理子模型以及运动感知子模型;目标视频确定模块,设置为基于所述至少两个目标视频帧,确定目标视频。
- 一种电子设备,包括:一个或多个处理器;存储装置,设置为存储一个或多个程序,当所述一个或多个程序被所述一个或多个处理器执行,使得所述一个或多个处理器实现如权利要求1-10中任一所述的视频处理方法。
- 一种包含计算机可执行指令的存储介质,所述计算机可执行指令在由计算机处理器执行时用于执行如权利要求1-10中任一所述的视频处理方法。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP23878926.7A EP4529154A4 (en) | 2022-10-21 | 2023-09-26 | VIDEO PROCESSING METHOD AND APPARATUS, AND ELECTRONIC DEVICE AND RECORDING MEDIUM |
| US18/867,268 US20250356462A1 (en) | 2022-10-21 | 2023-09-26 | Video processing method and apparatus, electronic device, and storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202211294643.2A CN115633144B (zh) | 2022-10-21 | 2022-10-21 | 视频处理方法、装置、电子设备及存储介质 |
| CN202211294643.2 | 2022-10-21 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024082933A1 true WO2024082933A1 (zh) | 2024-04-25 |
Family
ID=84907105
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2023/121354 Ceased WO2024082933A1 (zh) | 2022-10-21 | 2023-09-26 | 视频处理方法、装置、电子设备及存储介质 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US20250356462A1 (zh) |
| EP (1) | EP4529154A4 (zh) |
| CN (1) | CN115633144B (zh) |
| WO (1) | WO2024082933A1 (zh) |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN115633144B (zh) * | 2022-10-21 | 2025-07-29 | 抖音视界有限公司 | 视频处理方法、装置、电子设备及存储介质 |
| CN116668750A (zh) * | 2023-05-05 | 2023-08-29 | 北京达佳互联信息技术有限公司 | 视频处理方法、装置、设备及存储介质 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108134938A (zh) * | 2016-12-01 | 2018-06-08 | 中兴通讯股份有限公司 | 视频扫描方式检测、纠正方法、及视频播放方法和装置 |
| KR101979584B1 (ko) * | 2017-11-21 | 2019-05-17 | 에스케이 텔레콤주식회사 | 디인터레이싱 방법 및 장치 |
| CN112218081A (zh) * | 2020-09-03 | 2021-01-12 | 深圳市捷视飞通科技股份有限公司 | 视频图像去隔行的方法和装置、电子设备及存储介质 |
| CN112750094A (zh) * | 2020-12-30 | 2021-05-04 | 合肥工业大学 | 一种视频处理方法及系统 |
| US20220014708A1 (en) * | 2020-07-10 | 2022-01-13 | Disney Enterprises, Inc. | Deinterlacing via deep learning |
| CN115633144A (zh) * | 2022-10-21 | 2023-01-20 | 抖音视界有限公司 | 视频处理方法、装置、电子设备及存储介质 |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114299105B (zh) * | 2021-08-04 | 2025-07-18 | 腾讯科技(深圳)有限公司 | 图像处理方法、装置、计算机设备及存储介质 |
| CN113344794B (zh) * | 2021-08-04 | 2021-10-29 | 腾讯科技(深圳)有限公司 | 一种图像处理方法、装置、计算机设备及存储介质 |
-
2022
- 2022-10-21 CN CN202211294643.2A patent/CN115633144B/zh active Active
-
2023
- 2023-09-26 EP EP23878926.7A patent/EP4529154A4/en active Pending
- 2023-09-26 WO PCT/CN2023/121354 patent/WO2024082933A1/zh not_active Ceased
- 2023-09-26 US US18/867,268 patent/US20250356462A1/en active Pending
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108134938A (zh) * | 2016-12-01 | 2018-06-08 | 中兴通讯股份有限公司 | 视频扫描方式检测、纠正方法、及视频播放方法和装置 |
| KR101979584B1 (ko) * | 2017-11-21 | 2019-05-17 | 에스케이 텔레콤주식회사 | 디인터레이싱 방법 및 장치 |
| US20220014708A1 (en) * | 2020-07-10 | 2022-01-13 | Disney Enterprises, Inc. | Deinterlacing via deep learning |
| CN112218081A (zh) * | 2020-09-03 | 2021-01-12 | 深圳市捷视飞通科技股份有限公司 | 视频图像去隔行的方法和装置、电子设备及存储介质 |
| CN112750094A (zh) * | 2020-12-30 | 2021-05-04 | 合肥工业大学 | 一种视频处理方法及系统 |
| CN115633144A (zh) * | 2022-10-21 | 2023-01-20 | 抖音视界有限公司 | 视频处理方法、装置、电子设备及存储介质 |
Non-Patent Citations (2)
| Title |
|---|
| See also references of EP4529154A4 |
| WANG CHONG: "Video Deinterlacing Method Based on Optical Flow Method", INFORMATIZATION RESEARCH, vol. 39, no. 1, 20 March 2013 (2013-03-20), pages 52 - 57, XP093160126 * |
Also Published As
| Publication number | Publication date |
|---|---|
| US20250356462A1 (en) | 2025-11-20 |
| EP4529154A4 (en) | 2025-12-24 |
| EP4529154A1 (en) | 2025-03-26 |
| CN115633144B (zh) | 2025-07-29 |
| CN115633144A (zh) | 2023-01-20 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP4053784A1 (en) | Image processing method and apparatus, electronic device, and storage medium | |
| EP4529154A1 (en) | Video processing method and apparatus, and electronic device and storage medium | |
| CN116527748B (zh) | 一种云渲染交互方法、装置、电子设备及存储介质 | |
| US11893770B2 (en) | Method for converting a picture into a video, device, and storage medium | |
| US20250329086A1 (en) | Image processing method and apparatus, electronic device and storage medium | |
| CN111860363A (zh) | 一种视频图像的处理方法及装置、电子设备、存储介质 | |
| WO2024037556A1 (zh) | 图像处理方法、装置、设备及存储介质 | |
| CN116302268A (zh) | 媒体内容的展示方法、装置、电子设备和存储介质 | |
| CN114245028A (zh) | 图像展示方法、装置、电子设备及存储介质 | |
| CN115113838A (zh) | 投屏画面的显示方法、装置、设备及介质 | |
| CN111402133A (zh) | 图像处理方法、装置、电子设备及计算机可读介质 | |
| CN116939130A (zh) | 一种视频生成方法、装置、电子设备和存储介质 | |
| US20250024011A1 (en) | Video processing method and apparatus, electronic device, and storage medium | |
| CN115272061A (zh) | 特效视频的生成方法、装置、设备及存储介质 | |
| CN114745545B (zh) | 一种视频插帧方法、装置、设备和介质 | |
| CN112218081A (zh) | 视频图像去隔行的方法和装置、电子设备及存储介质 | |
| CN102769732A (zh) | 一种实现视频场的转换方法 | |
| CN106658095A (zh) | 一种直播视频传输的方法、服务器和用户设备 | |
| CN113535645A (zh) | 共享文档的展示方法、装置、电子设备及存储介质 | |
| CN113706385A (zh) | 一种视频超分辨率方法、装置、电子设备及存储介质 | |
| WO2025113388A1 (zh) | 一种三维数据的生成方法、装置、电子设备及存储介质 | |
| WO2025161719A1 (zh) | 图像处理方法、装置、可读介质及电子设备 | |
| CN115802105B (zh) | 视频注入方法及其设备、信息处理系统 | |
| CN115767126B (zh) | 一种yuv纹理生成方法、装置及电子设备 | |
| CN117057993A (zh) | 一种轻量化图像超分方法、装置、设备、存储介质及产品 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23878926 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18867268 Country of ref document: US |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2023878926 Country of ref document: EP |
|
| ENP | Entry into the national phase |
Ref document number: 2023878926 Country of ref document: EP Effective date: 20241217 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| WWP | Wipo information: published in national office |
Ref document number: 18867268 Country of ref document: US |