WO2023246888A1 - 多媒体数据处理方法、装置和计算机可读存储介质 - Google Patents
多媒体数据处理方法、装置和计算机可读存储介质 Download PDFInfo
- Publication number
- WO2023246888A1 WO2023246888A1 PCT/CN2023/101786 CN2023101786W WO2023246888A1 WO 2023246888 A1 WO2023246888 A1 WO 2023246888A1 CN 2023101786 W CN2023101786 W CN 2023101786W WO 2023246888 A1 WO2023246888 A1 WO 2023246888A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- feature data
- text
- preset
- mapping relationship
- expression
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/233—Processing of audio elementary streams
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T11/00—Two-dimensional [2D] image generation
- G06T11/60—Creating or editing images; Combining images with text
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
- H04N21/4402—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving reformatting operations of video signals for household redistribution, storage or real-time display
- H04N21/440236—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving reformatting operations of video signals for household redistribution, storage or real-time display by media transcoding, e.g. video is transformed into a slideshow of still pictures, audio is converted into text
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/279—Recognition of textual entities
- G06F40/289—Phrasal analysis, e.g. finite state techniques or chunking
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/20—Image preprocessing
- G06V10/25—Determination of region of interest [ROI] or a volume of interest [VOI]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/20—Image preprocessing
- G06V10/26—Segmentation of patterns in the image field; Cutting or merging of image elements to establish the pattern region, e.g. clustering-based techniques; Detection of occlusion
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/74—Image or video pattern matching; Proximity measures in feature spaces
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
- G06V40/16—Human faces, e.g. facial parts, sketches or expressions
- G06V40/174—Facial expression recognition
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
- G06V40/16—Human faces, e.g. facial parts, sketches or expressions
- G06V40/174—Facial expression recognition
- G06V40/176—Dynamic expression
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/18—Speech classification or search using natural language modelling
- G10L15/1822—Parsing for meaning understanding
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/26—Speech to text systems
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
- G10L25/57—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for processing of video signals
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
- H04N21/23412—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs for generating or manipulating the scene composition of objects, e.g. MPEG-4 objects
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
- H04N21/23418—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
- H04N21/2343—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving reformatting operations of video signals for distribution or compliance with end-user requests or end-user device requirements
- H04N21/234336—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving reformatting operations of video signals for distribution or compliance with end-user requests or end-user device requirements by media transcoding, e.g. video is transformed into a slideshow of still pictures or audio is converted into text
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/235—Processing of additional data, e.g. scrambling of additional data or processing content descriptors
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/435—Processing of additional data, e.g. decrypting of additional data, reconstructing software from modules extracted from the transport stream
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/439—Processing of audio elementary streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/439—Processing of audio elementary streams
- H04N21/4394—Processing of audio elementary streams involving operations for analysing the audio stream, e.g. detecting features or characteristics in audio streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
- H04N21/44008—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics in the video stream
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
- H04N21/44012—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving rendering scenes according to scene graphs, e.g. MPEG-4 scene graphs
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/81—Monomedia components thereof
- H04N21/8106—Monomedia components thereof involving special audio data, e.g. different tracks for different languages
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
- G10L2015/221—Announcement of recognition results
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
- G10L25/63—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for estimating an emotional state
Definitions
- the embodiments of the present application relate to but are not limited to the field of data processing technology, and in particular, to a multimedia data processing method, device and computer-readable storage medium.
- Embodiments of the present application provide a multimedia data processing method, device and computer-readable storage medium.
- embodiments of the present application provide a multimedia data processing method, which includes: obtaining audio streams and video streams of multimedia data; parsing the audio streams to obtain text feature data, and matching the described data according to a preset mapping relationship Text feature data to determine topic feature data; parse the video stream to obtain expression feature data, and compare and match the expression feature data according to a preset mapping relationship to determine the emotion index; based on the text feature data, the emotion index and The topic feature data renders the multimedia data.
- embodiments of the present application provide a multimedia data processing device, including: an audio processing module configured to receive and parse the audio stream to obtain text feature data; a video processing module configured to receive and parse the video stream, Obtain the expression feature data; the mapping relationship module is configured to process the text feature data to obtain topic feature data, and process the expression feature data to obtain the emotional index; the rendering module is configured to combine the text feature data and the topic feature The data and the sentiment index are rendered with the video stream.
- embodiments of the present application provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above mentioned steps are implemented.
- embodiments of the present application provide a computer-readable storage medium, which stores A computer executable program is stored, and the computer executable program is used to cause the computer to execute the multimedia data processing method described in the first aspect above.
- Figure 1 is a main flow chart of a multimedia data processing method provided by an embodiment of the present application.
- Figure 2A is a sub-flow chart of a multimedia data processing method provided by an embodiment of the present application.
- Figure 2B is a sub-flow chart for parsing an audio stream to obtain text feature data provided by an embodiment of the present application
- Figure 3A is a sub-flow chart of a multimedia data processing method provided by an embodiment of the present application.
- Figure 3B is a schematic diagram of a micro-expression provided by an embodiment of the present application.
- Figure 3C is a sub-flow chart for parsing a video stream to obtain expression feature data provided by an embodiment of the present application
- Figure 3D is a schematic diagram of the mapping relationship between preset micro-expression combinations and emotional indexes provided by an embodiment of the present application
- Figure 4A is a sub-flow chart of a multimedia data processing method provided by an embodiment of the present application.
- Figure 4B is a schematic diagram of the mapping relationship between preset key phrases and scene templates provided by an embodiment of the present application.
- Figure 4C is a sub-flow chart of a multimedia data processing method provided by an embodiment of the present application.
- Figure 4D is a schematic diagram of the mapping relationship between preset declaration statements and command sequences provided by an embodiment of the present application.
- Figure 4E is a schematic diagram of a command sequence type provided by an embodiment of the present application.
- Figure 5A is a sub-flow chart of a multimedia data processing method provided by an embodiment of the present application.
- Figure 5B is a schematic diagram of the mapping relationship between preset sensitive statements and rendering sets provided by an embodiment of the present application.
- Figure 6 is a schematic diagram of the structure of a multimedia data processing device provided by an embodiment of the present application.
- Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present application.
- embodiments of the present application provide a multimedia data processing method, device and computer-readable storage medium.
- obtain the audio stream and video stream of the multimedia data analyze the audio stream , obtain the text feature data, and compare and match the text feature data according to the preset mapping relationship to determine the topic feature data; parse the video stream to obtain the expression feature data, and compare and match the expression feature data according to the preset mapping relationship to determine the emotional index; based on the text Feature data, sentiment index and topic feature data render multimedia data.
- this application analyzes and extracts, intelligently analyzes and renders the multimedia information of multi-person discussion calls, and digs out multi-dimensional value information based on the voice and video of existing call products, including voice text, emotion Index, and the content of the voice text is further processed according to the preset mapping relationship, so that the call product can present richer and smarter information.
- this application can flexibly handle different business scenarios and talks.
- Content expand the application fields of technical solutions, can effectively tap the potential value of call products, bring a new experience to call negotiations, increase the negotiation competitiveness of negotiation business users, improve decision-making efficiency, and win more voice. Seize business opportunities effectively.
- FIG. 1 is a flow chart of a multimedia data processing method provided by an embodiment of the present application.
- Multimedia data processing methods include but are not limited to the following steps:
- Step S100 obtain the audio stream and video stream of multimedia data
- Step S200 parse the audio stream to obtain text feature data, and compare and match the text feature data according to the preset mapping relationship to determine the topic feature data;
- Step S300 parse the video stream to obtain expression feature data, and compare and match the expression feature data according to the preset mapping relationship to determine the emotional index;
- Step S400 Render multimedia data based on text feature data, emotion index and topic feature data.
- the multimedia transmitted in the network includes audio and video.
- the audio stream and video of the multimedia data are first obtained.
- the sounds in the audio stream are parsed and extracted to obtain text feature data.
- the sounds in the audio stream can be human voices, and the human voices are converted into text, that is, language characters, through speech recognition technology; similarly, it can also be Whether it is a non-human voice, based on the frequency spectrum, identify its corresponding sound type, name, characteristics, etc., and convert the name or action corresponding to the sound into a text description.
- the sound of an instrument is recognized as piano music, text features
- the data can include piano music, the name of the piano music, related information of the piano music, etc.; then the video stream in the multimedia data is parsed to obtain expression feature data.
- the expression feature data includes the expressions of the participants in the video conversation, for example, through the video Recognize consecutive frames in the stream, analyze the faces in the video, extract the dynamic changes in facial expressions, and generate different expression feature data for different expressions. For example, identify the angry expression on the face through the eyebrows. , the combination of eyes, face, nose, lips and other postures generates expression feature data.
- the text feature data is matched according to the preset mapping relationship to determine the topic feature data.
- the preset mapping relationship is to set some text feature data and topic feature data into a mapping relationship, that is, some text short sentences are followed by words.
- the question keyword is set as a mapping relationship.
- the text phrase and some action commands can also be set as a mapping relationship.
- the text phrase is matched as a search action, that is, if a match is achieved, the matched phrase is retrieved. , use the search results as topic feature data.
- the text feature data is obtained after multi-level mapping transformation and action processing. Therefore, according to The preset mapping relationship matches the text phrases in the text feature data with the preset text phrases in the mapping relationship. If the matching degree reaches the preset threshold, the text phrases in the text feature data are matched according to the preset mapping relationship.
- the keywords mapped to the preset text phrases matched by the text phrases are determined as topic feature data.
- the text feature data may contain an entire text.
- the text in the text feature data is dispersed into short sentences or short words, and the scattered short sentences or short words are combined with the preset
- the preset threshold can be understood as the proportion of short sentences or short words scattered in the text feature data that match the preset short text sentences. You can gradually test by setting values with different granularities. A suitable preset threshold range can also be used to determine whether the threshold setting is appropriate through the recognition effect.
- the preset mapping relationship is to set some expression feature data and the emotion index into a mapping relationship, that is, to set the expression extracted from the video stream and the emotion index.
- the emotional index can be expressed in the form of color, numerical value, graphics, etc.
- the emotion is set from angry to happy to 0-100 points, 0 means very angry, and 100 means very happy.
- the expression extracted from the video stream can be matched to the corresponding score.
- the emotional index can be represented by color. Different colors correspond to different emotions. Through color psychology, warm colors represent boldness. , sunshine, enthusiasm, enthusiasm, and liveliness. Cold colors represent gracefulness, femininity, calmness, and elegance.
- Expression feature data can include expression changes over a period of time, and may include multiple expression combinations.
- the preset threshold can be understood as comparing and matching these multiple expression combinations with preset expressions, and what proportion of expressions match the preset expressions. To achieve a matching expression, you can gradually test the appropriate preset threshold range by setting values at different granularities, or you can reversely judge whether the threshold setting is appropriate through the recognition effect.
- multimedia data is rendered based on text feature data, sentiment index and topic feature data.
- the speech text obtained by speech conversion in the text feature data audio stream has an emotional index of the interviewer's emotional index.
- the topic feature data is compared and matched based on the speech text and the preset text.
- the matched preset text is Mapping topic feature data, or topic feature data obtained after matching the preset text through multi-level mapping and executing actions. Therefore, the multimedia data is rendered based on the text feature data, emotion index and topic feature data, and the text feature data is Render into the video and become the subtitles of the conference conversation. Render the emotional index into the video to display the emotional state of the participants in real time.
- the text feature data obtained from the audio stream will
- the topic feature data of the secondary processing content obtained through mapping or multi-level mapping of the text feature data, and the emotional index obtained from the video stream are rendered with the multimedia data respectively, so that the video stream not only contains the picture content of the meeting, but also includes Multi-dimensional auxiliary information is returned to the media plane routing, and sent to the conference site through the transmission system.
- the multimedia data is parsed and presented through XR equipment at the conference site. According to 3Gpp specifications, this technical solution can be deployed on the routing or media processing network elements, which can also be deployed independently. Participants can see the original video information and hear the original audio information.
- the meeting discusses the topic of financial security, a number of financial-related short sentences are scattered from the text feature data.
- the preset financial-related short sentences For example, the stock market, names of listed companies, financial reports of listed companies, insurance companies, China Securities Regulatory Commission, etc., match a number of financial-related short sentences dispersed in the text feature data with the preset financial short sentences.
- the preset threshold For example, when 10 short sentences or keywords are reached within a certain period of time, it is considered to have reached the preset threshold, that is, the successfully matched short sentences or keywords are mapped.
- the mapping content can be the corresponding topic keywords, or it can be The content obtained after multi-level actions.
- the successfully matched keywords are the names and financial data of listed companies.
- the preset short sentences map to the actions.
- the action content is "retrieve its income data in the past 2 years.” Therefore, based on matching If the keywords listed company name and financial data are found, the mapping content is to retrieve the revenue data of the corresponding listed company in the past two years. At the meeting site, the revenue data of the corresponding listed company in the past two years will be immediately displayed in the space, allowing participants to Participants can grasp the topic content more efficiently.
- mapping relationships for various scenarios and topics can be preset.
- the mapping relationships can also be multi-level mappings.
- mapping relationships can be keywords and can also include execution actions. , also contains multi-level actions. After a lot of training and accumulation, the mapping relationship can be continuously optimized to become more accurate and intelligent. The accuracy and intelligence of the mapping relationship can also be judged by the length of time the participants visually pay attention to the content on site. , to achieve automated optimization of mapping relationships and automatic entry and update of mapping relationships.
- mapping relationship includes at least one of: a mapping relationship between a preset micro-expression combination and an emotional index; a mapping relationship between a preset key phrase and a scene template; a mapping relationship between a preset statement statement and a command sequence; a preset sensitivity
- mapping relationship between statements and rendering sets Setting up mapping relationships in multiple dimensions enables data processing to be performed based on the information in different dimensions and using mapping relationships in different dimensions.
- step S200 may include but is not limited to the following sub-steps:
- Step S2001 Perform frequency feature learning on the audio stream to obtain the frequency feature band
- Step S2002 filter the audio stream using frequency characteristic bands to obtain several characteristic audio streams
- Step S2003 Perform speech recognition on the characteristic audio stream to obtain subtitle text
- Step S2004 Perform sound intensity analysis on the characteristic audio stream to obtain a sound intensity value.
- the sound intensity value reaches a preset threshold, additional subtitle text is output.
- frequency feature learning is performed on the audio stream, that is, the frequency features of different voices are learned, and the frequency feature band is obtained, that is, the frequency band where the frequency features of different voices are located.
- the audio stream is Filter to obtain several characteristic audio streams, that is, separate the human voice from the audio stream, and further separate the voices of different people in the human voice to obtain the audio information of several channels of human voices.
- Perform speech recognition on the characteristic audio stream to obtain For several subtitle texts corresponding to human voices, perform sound intensity analysis on the audio stream to obtain the sound intensity value.
- additional subtitle text corresponding to several human voices is output, that is, short sentences whose pitch or volume exceeds the preset value. or keywords.
- step S300 may include but is not limited to the following sub-steps:
- Step S3001 Input the video stream to the preset deep learning model to identify and obtain facial area coordinates
- Step S3002 Based on the coordinates of the facial area, perform micro-expression recognition and segmentation on the continuous frames of the facial area in the video stream to obtain a micro-expression combination.
- facial expressions involve movements of multiple parts, among which "micro-expressions" naturally correspond to "macro-expressions". Macro expressions are easy to observe in our daily life. The duration of facial expressions is within 0.5-4 seconds, and the facial muscles involved in facial expressions contract or relax to a large extent. Micro-expressions are not only not controlled by thinking consciousness, but also last for a very short time. People's consciousness has been exposed before they have time to control them, so they reflect true emotions to a certain extent. It is precisely because of this characteristic of micro-expressions that it has very important applications in criminal investigation, security, justice, negotiation and other fields. Microexpression is a special facial expression. Compared with ordinary expressions, microexpressions mainly have the following characteristics:
- the duration is short, usually only 1/25s ⁇ 1/3s;
- the action intensity is low and difficult to detect
- micro-expressions usually needs to be done in videos, while ordinary expressions can be analyzed in images.
- micro-expressions The recognition of micro-expressions is initially trained manually. After about an hour and a half of training, the accuracy can be improved to 30%-40%, but psychologists have demonstrated that manual recognition will not exceed 47%. Later, as psychological experiments gradually evolved into computer applications, micro-expressions had AU (Action Units) combinations. For example, if it is happy, the happy macro-expression is AU6+AU12, and the micro-expression is AU6 or AU12, or AU6 +AU12. The second is micro-expression movement, which is a local muscle movement, rather than two muscle groups moving at the same time. For example, if you are happy, the movement is AU6 or AU12. But if the intensity is relatively high, AU6 and AU12 may occur at the same time.
- AU Application Units
- micro expressions can be extracted from videos through the micro expression recognition technology MER (Micro expression recognition).
- MER Micro expression recognition
- the expression changes of the participants can be accurately described, and the emotional changes of the participants can be reflected laterally, and the interest of the participants in the topic of discussion can be further analyzed. Therefore, as shown in Figure 3C, the video stream is input to the preset deep learning model.
- the convolutional neural network (Convolutional Neural Networks, CNN) model can be used to first perform learning and training to obtain a mature CNN Model, input the video stream into the CNN model, perform face recognition, obtain the facial area coordinates through recognition, obtain the area coordinates of the faces of each interviewer in the video, use the MER blocker to perform MER block on the continuous frames of the video face area , further analyze the MER sequence, output the changes in the continuous time series pictures, and obtain the combination sequence of micro-expression AU, that is, the micro-expression combination.
- CNN convolutional Neural Networks
- the mapping relationship between the preset micro-expression combination and the emotional index is that the micro-expression combination consists of several micro-expressions, for example, micro-expression 1, micro-expression 2 and micro-expression 3, where micro-expression Expression 1 is AU6, micro-expression 2 is AU7, micro-expression 3 is AU12, that is, the micro-expression combination is AU6+AU7+AU12.
- the micro-expressions can be arranged and combined to obtain a set of micro-expression combinations. According to the micro-expression set The corresponding expressed emotion is mapped to the corresponding emotion index.
- the emotion index can be expressed in color, numerical, graphical and other forms. For example, set the emotion from angry to happy to 0-100 points, 0 means very angry, 100 means very happy. According to different expressions, set the corresponding emotion index.
- the expressions extracted from the video stream can be matched to the corresponding scores.
- the emotional index can be represented by color. Different colors correspond to different emotions, through color psychology, because warm colors represent boldness, sunshine, enthusiasm, and enthusiasm. , lively, and cold colors represent graceful, feminine, calm, and elegant.
- step S200 also includes but is not limited to the following sub-steps:
- Step S2005 process the subtitle text through natural language processing to obtain a sequence of short sentences
- Step S2006 use the mapping relationship to compare and match the short sentence sequence with the preset key short sentences to obtain the corresponding scene template
- Step S2007 Perform in-depth processing on the subtitle text and subtitle additional text according to the corresponding scene template to obtain topic feature data.
- the subtitle text is semantically processed.
- the subtitle text is semantically segmented into several short sentence sequences through natural language processing (Natural language processing), and a mapping relationship is used to combine the short sentence sequence with the predetermined sentence sequence.
- Natural language processing Natural language processing
- Set key phrases for comparison and matching When the matching degree reaches the preset threshold, the corresponding scene template is obtained.
- the mapping relationship between preset key phrases and scene templates is that at least one preset key phrase is mapped to a scene.
- a topic scene in the financial field is preset
- the preset key phrase is the Shanghai Composite Index. , secondary market, Shanghai Stock Exchange, Shenzhen Stock Exchange, capital inflow and outflow in the two cities, A-shares, 6 preset key phrases mapped to financial scenarios
- the subtitle text extracted from the video conference audio stream is divided into semantics Several short sentence sequences are compared and matched with 6 preset key short sentences, and the similarity is calculated. When the similarity reaches the preset threshold, it is judged that the current meeting topic is a financial topic, and then the financial scene is entered. .
- the subtitle text and subtitle additional text are deeply processed to obtain topic feature data.
- the subtitle text and the subtitle additional text are further semantically segmented according to the preset mapping relationship, and compared with the preset sentences to obtain action instructions for data processing.
- the mapped Action process subtitle text and subtitle attached text for secondary processing, and output and render the secondary processing result.
- the threshold can be set according to the strictness of scene matching and the number of preset key phrases. For example, in the same scene, the more the number of preset key phrases, the higher the matching degree required, then the threshold The higher it is, the more difficult it is to match the scene, but the accuracy will be higher.
- this solution is preset with multiple scenarios, and multiple preset key phrases map to one scenario. Scenarios of different dimensions and levels can be set to apply to different types of meeting themes.
- step S2007 also includes but is not limited to the following sub-steps:
- Step S20071 use the mapping relationship to compare and match the short sentence sequence and subtitle additional text with the preset statement statement to obtain the corresponding command sequence;
- Step S20072 executes the command sequence to obtain topic feature data.
- mapping relationship that is, the mapping relationship between the preset declaration statement and the command sequence
- the short sentence sequence and the subtitle additional text are compared and matched with the preset declaration statement to obtain the corresponding command sequence.
- the mapping relationship between preset declaration statements and command sequences is that at least one preset declaration statement is mapped to a command sequence.
- the command sequence contains an action group, that is, several data processing actions, for example, presetting a 5G communication field Topic scenarios, the preset key phrases are 5G networking, non-standalone networking NSA, independent networking SA, control bearer separation, mobile edge computing, network slicing, 6 preset key phrases are mapped to 5G communication scenarios, Segment the subtitle text extracted from the video conference audio stream into several short sentence sequences according to semantics, compare and match these short sentence sequences with 6 preset key short sentences, and calculate the similarity.
- the mapping relationship between the preset declaration statement and the command sequence is used to attach text to the short sentence sequence and subtitles and the preset declaration.
- the statements are compared and matched to obtain the corresponding command sequence.
- the preset statement statements are the number of 5G patents, 5G standard essential patents, 5G patent holdings, 5G technology companies, 5G market size, and 5 preset statement statements.
- the command sequence mapped by the five preset declaration statements can be set as:
- Action 1 The search engine retrieves the preset declaration statement
- Action 2 Extract the content introduction or keywords of the first five URLs in the search results
- Action 3 Search the patent database for keywords extracted from the web page
- Action 4 Remove duplicates from the search results and output them.
- the command sequence contains 4 actions. Through this mapping relationship, several short sentence sequences are compared and matched with 5 preset statement sentences in 6 5G communication scenarios. When the similarity reaches the preset value, the match is considered successful. Therefore, the command sequence mapped by the preset declaration statement is executed, that is, 4 actions, and the action result is finally output.
- the threshold can be set according to the strictness of the match and the number of preset declaration statements. For example, under the same command sequence, the more the number of preset declaration statements, the higher the matching degree required, the higher the threshold. , the more difficult it is to match the command sequence, but the accuracy will be higher. Therefore, a suitable threshold range can be found through a large amount of data training, taking into account both accuracy and practicality. It should be noted that multiple command sequences are preset in this solution. Multiple command sequences map to one command sequence. Command sequences of different dimensions and levels can be set to apply to different types of secondary data processing scenarios.
- step S20071 also includes but is not limited to at least one of the following actions:
- the command sequence includes an online search instruction
- the online database is searched and the online search data set is obtained
- the offline database is searched to obtain the offline retrieval data set
- the command sequence contains data processing instructions
- the data is processed twice to obtain secondary processing data.
- search engine To search the database online, you can use one or more combinations of Google search engine, Baidu search engine, Bing search engine, Sogou search engine, Haosou search engine, Shenma search engine, and Yahoo search engine.
- Search engines in subdivided professional fields such as special databases in the field of patent retrieval, special databases in the field of literature retrieval, etc.
- offline databases are local databases, which can be internal databases in the local area network, targeting the fields used by this technical solution or the user's own
- There is a database which can be set flexibly; when the command sequence contains data processing instructions, the data will be processed twice.
- the secondary processing method can still use mapping relationships to perform multiple mappings and process the data multiple times and in multiple dimensions.
- step S400 also includes but is not limited to the following sub-steps:
- Step S4001 use the subtitle text to obtain the corresponding sound time sequence when the characteristic audio stream is started;
- Step S4002 use the micro-expression combination to obtain the corresponding mouth shape time sequence when the mouth expression is activated;
- Step S4003 determine the consistency of the sound time series and the mouth shape time series, and align the subtitle text, emotion index and facial area coordinates in time series;
- Step S4004 use the mapping relationship to compare and match the short sentence sequence and the preset sensitive sentences to obtain the corresponding rendering set;
- Step S4005 Render the subtitle text, emotion index, and topic feature data with the video stream according to the configuration of the rendering set.
- the subtitle text is used to obtain the sound time sequence corresponding to the start of the characteristic audio stream.
- the subtitle text is extracted with a time mark, and the time corresponding to the text can be determined; the micro-expression combination is used to obtain the corresponding sound time sequence when the mouth expression is started.
- the micro-expression combination When extracting micro-expressions from the video stream, they are extracted sequentially according to the time series, so the micro-expression combination also has a time stamp.
- the video stream and audio stream are the same multimedia source, they have consistent timing, so , during data processing, due to the separate processing of audio streams and video streams and secondary data processing, data timing may occur.
- Asynchronous therefore, determine the consistency of the sound time series and the mouth shape time series, and align the subtitle text, emotion index and facial area coordinates in time series.
- the subtitle text corresponds to the sound time series
- the emotion index corresponds to the mouth shape time.
- sequence, the facial area coordinates correspond to the mouth shape time series.
- the short sentence sequence and the preset sensitive statement are compared and matched using the mapping relationship to obtain the corresponding rendering set.
- the preset sensitive statement and the rendering set setting have a mapping relationship.
- At least one preset sensitive statement and the rendering set are Mapping, according to the needs of the usage scenario, preset sensitive statements of different dimensions can be preset as matching references.
- the configuration of the rendering set includes at least one of rendering coordinates, rendering time, and rendering color theme.
- Different rendering sets will present different visual effects when rendering subtitle text, emotion index, and topic feature data with the video stream.
- the rendering set after the subtitle text, emotion index and topic feature data are rendered with the video stream, meeting participants will see different colors of subtitle text and different styles near different faces. Sentiment index, topic feature data at different display times. Participants can see the nearby areas of different interviewers on the XR device. There are display boards or barrage pop-ups with subtitles for their respective talks. They can also see the display of the emotional index of the participants. The key information during the meeting will be processed after secondary processing.
- the display duration, font color, and display position of different levels of key information are different.
- the presentation method can be icons or numbers, or other designs.
- different key information can be displayed through the management interface. Words and sentences are grouped, and the grouped collection can be set with corresponding processing actions and different rendering effect settings.
- the embodiment of the present application also provides a multimedia data processing device, which can parse, extract, intelligently analyze and render the multimedia information of multi-person discussion calls, thereby improving the voice and video quality of existing call products.
- multi-dimensional value information is mined, including voice text and emotional index, and the content of the voice text is further processed according to the preset mapping relationship, so that the call product can present richer and smarter information.
- it can flexibly handle different business scenarios and meeting contents, expand the application areas of technical solutions, and effectively tap the potential value of call products, bringing a new experience to call discussions and increasing the number of negotiation business users. Improve negotiation competitiveness, improve decision-making efficiency, win more voice, and effectively seize business opportunities.
- the multimedia data processing device includes:
- the audio processing module 501 is configured to receive audio data, analyze and extract sounds in the audio stream, and obtain text feature data;
- the video processing module 502 is configured to receive video data, analyze the video stream, and obtain expression feature data.
- the expression feature data includes the expressions of the participants in the video session, and analyzes the people in the video through the recognition of consecutive frames in the video stream. face, extract the dynamic changes in facial expressions, and generate different expression feature data for different expressions;
- the mapping relationship module 503 converts the text feature data into topic feature data according to the preset mapping relationship; converts the expression feature data into emotional index;
- the rendering module 504 renders and outputs text feature data, sentiment index and topic feature data.
- the audio processing module receives the audio data, analyzes and extracts the sounds in the audio stream, and obtains text feature data;
- the video processing module receives the video data, analyzes the video stream, and obtains the expression feature data.
- the expression feature data includes The expressions of the participants in the video conversation are analyzed through the recognition of consecutive frames in the video stream, and the dynamic changes in facial expressions are extracted to generate different expression feature data for different expressions; mapping relationship The module converts text feature data into topic feature data and expression feature data into emotion index according to the preset mapping relationship; the rendering module renders and outputs text feature data, emotion index and topic feature data. Based on this, this application is approved Analyze, extract, intelligently analyze and render the multimedia information of multi-person discussion calls.
- an embodiment of the present application also provides an electronic device.
- multi-dimensional value information can be mined, including voice text, emotional index, and The content of the voice text is further processed according to the preset mapping relationship, so that the call product can present richer and smarter information.
- the application fields of the solution can effectively tap the potential value of call products, bring a new experience to call negotiations, increase the negotiation competitiveness of negotiation business users, improve decision-making efficiency, win more voice, and effectively seize business opportunities. .
- the electronic equipment includes:
- the processor can be implemented by a general-purpose CPU (Central Processing Unit, central processing unit), a microprocessor, an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is configured to execute Relevant procedures to implement the technical solutions provided by the embodiments of this application;
- a general-purpose CPU Central Processing Unit, central processing unit
- a microprocessor central processing unit
- an application specific integrated circuit Application Specific Integrated Circuit, ASIC
- ASIC Application Specific Integrated Circuit
- Memory can be implemented in the form of read-only memory (Read Only Memory, ROM), static storage device, dynamic storage device, or random access memory (Random Access Memory, RAM).
- ROM read-only memory
- RAM random access memory
- the memory can store operating systems and other application programs.
- the relevant program codes are stored in the memory and called by the processor to execute the goals of the embodiments of this application. How to sort data;
- Input/output interface configured to implement information input and output
- the communication interface is configured to realize communication interaction between this device and other devices. Communication can be achieved through wired methods (such as USB, network cables, etc.) or wireless methods (such as mobile network, WIFI, Bluetooth, etc.);
- Buses which transmit information between various components of a device (such as processors, memory, input/output interfaces, and communication interfaces);
- the processor, memory, input/output interface and communication interface realize the communication connection between each other within the device through the bus.
- the electronic device includes: one or more processors and memories.
- one processor and memory are taken as an example.
- the processor and memory can be connected through a bus or other means.
- Figure 7 takes the connection through a bus as an example.
- the memory can be used to store non-transitory software programs and non-transitory computer executable programs, such as the multimedia data processing method in the above embodiments of the present application.
- the processor implements the multimedia data processing method in the above embodiments of the present application by running non-transient software programs and programs stored in the memory.
- the memory may include a program storage area and a data storage area, where the program storage area may store an operating system and an application program required for at least one function; the storage data area may store information required to execute the multimedia data processing method in the embodiments of the present application. Data etc.
- the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, flash memory device, or other non-transitory solid-state storage device.
- the memory includes memory located remotely relative to the processor, and these remote memories may be connected to the multimedia media via a network. Data processing device. Examples of the above-mentioned networks include but are not limited to the Internet, intranets, local area networks, mobile communication networks and combinations thereof.
- the non-transitory software programs and programs required to implement the multimedia data processing method in the above embodiments of the present application are stored in the memory.
- the multimedia data processing method in the above embodiments of the present application is executed. , for example, perform the above-described method steps S100 to step S400 in Figure 1, method steps S2001 to step S2004 in Figure 2A, method steps S3001 to step S3002 in Figure 3A, and method steps S2005 to step S2007 in Figure 4A. .
- this application analyzes and extracts, intelligently analyzes and renders the multimedia information of multi-person discussion calls, and digs out multi-dimensional value information based on the voice and video of existing call products, including voice text, emotion Index, and the content of the voice text is further processed according to the preset mapping relationship, so that the call product can present richer and smarter information.
- mapping relationships can flexibly handle different business scenarios and talks.
- Content expand the application fields of technical solutions, can effectively tap the potential value of call products, bring a new experience to call negotiations, increase the negotiation competitiveness of negotiation business users, improve decision-making efficiency, and win more voice. Seize business opportunities effectively.
- embodiments of the present application also provide a computer-readable storage medium that stores a computer-executable program, and the computer-executable program is executed by one or more control processors, for example, as shown in FIG. 7
- Execution by one of the processors may cause the one or more processors to execute the multimedia data processing method in the embodiment of the present application, for example, execute the above-described method steps S100 to S400 in FIG. 1, and execute steps S400 in FIG. 2A.
- this application analyzes and extracts, intelligently analyzes and renders the multimedia information of multi-person discussion calls, and digs out multi-dimensional value information based on the voice and video of existing call products, including voice text, emotion Index, and the content of the voice text is further processed according to the preset mapping relationship, so that the call product can present richer and smarter information.
- mapping relationships can flexibly handle different business scenarios and talks.
- Content expand the application fields of technical solutions, can effectively tap the potential value of call products, bring a new experience to call negotiations, increase the negotiation competitiveness of negotiation business users, improve decision-making efficiency, and win more voice. Seize business opportunities effectively.
- Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, Digital Versatile Disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, disk storage or other magnetic storage device, or any other medium that can be used to store the desired information and can be accessed by a computer.
- communication media typically embodies a computer-readable program, data structure, program module or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media .
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Signal Processing (AREA)
- General Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Human Computer Interaction (AREA)
- General Health & Medical Sciences (AREA)
- Computational Linguistics (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Artificial Intelligence (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Acoustics & Sound (AREA)
- Oral & Maxillofacial Surgery (AREA)
- Evolutionary Computation (AREA)
- General Engineering & Computer Science (AREA)
- Software Systems (AREA)
- Medical Informatics (AREA)
- Databases & Information Systems (AREA)
- Computing Systems (AREA)
- Life Sciences & Earth Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Data Mining & Analysis (AREA)
- Evolutionary Biology (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
Abstract
Description
Claims (13)
- 一种多媒体数据处理方法,包括:获取多媒体数据的音频流和视频流;解析所述音频流,得到文本特征数据,根据预设的映射关系对照匹配所述文本特征数据以确定话题特征数据;解析所述视频流,得到表情特征数据,根据预设的映射关系对照匹配所述表情特征数据以确定情感指数;基于所述文本特征数据、所述情感指数和所述话题特征数据对所述多媒体数据进行渲染。
- 根据权利要求1所述的方法,其中,所述映射关系包括如下至少之一:预置微表情组合与情感指数的映射关系;或预置关键短句与场景模板的映射关系;或预置声明语句与命令序列的映射关系;或预置敏感语句与渲染集的映射关系。
- 根据权利要求2所述的方法,其中,所述解析所述音频流,得到文本特征数据,包括:对所述音频流进行频率特征学习,得到频率特征频段;利用所述频率特征频段对所述音频流过滤,得到若干特征音频流;对所述特征音频流进行语音识别,得到字幕文本;对所述特征音频流进行音强分析,得到音强值,当所述音强值达到预设阈值,则输出字幕附加文本。
- 根据权利要求3所述的方法,其中,所述解析所述视频流,得到表情特征数据,包括:将所述视频流输入到预置的深度学习模型,识别获得面部区域坐标;根据所述面部区域坐标,对所述视频流中面部区域的连续帧进行微表情识别分块,得到微表情组合。
- 根据权利要求4所述的方法,其中,所述根据预设的映射关系对照匹配所述表情特征数据以确定情感指数,包括:利用所述映射关系,将所述微表情组合与所述预置微表情组合进行对照匹配,得到对应的情感指数。
- 根据权利要求3所述的方法,其中,所述根据预设的映射关系对照匹配所述文本特征数据以确定话题特征数据,包括:将所述字幕文本通过自然语言处理,获得短句序列;利用所述映射关系,将所述短句序列与所述预置关键短句进行对照匹配,得到对应的所述场景模板;依据对应的所述场景模板,对所述字幕文本、所述字幕附加文本进行深层处理获得话题特征数据。
- 根据权利要求6所述的方法,其中,所述依据对应的所述场景模板,对所述字幕文本、所述字幕附加文本进行深层处理获得话题特征数据,包括:利用所述映射关系,将所述短句序列和所述字幕附加文本与所述预置声明语句进行对照匹配,得到对应的所述命令序列;执行所述命令序列,获得话题特征数据。
- 根据权利要求7所述的方法,其中,所述执行所述命令序列,获得话题特征数据,包括如下至少之一:当所述命令序列包含在线检索指令,则检索在线数据库,得到在线检索数据集;或当所述命令序列包含离线检索指令,则检索离线数据库,得到离线检索数据集;或当所述命令序列包含数据加工指令,则对数据进行二次加工,得到二次加工数据。
- 根据权利要求4至8任一所述的方法,其中,所述基于所述文本特征数据、所述情感指数和所述话题特征数据对所述多媒体数据进行渲染,包括:利用所述字幕文本,获得所述特征音频流启动时对应的声音时间序列;利用所述微表情组合,获得嘴型表情启动时对应的嘴型时间序列;判断所述声音时间序列和所述嘴型时间序列的一致性,将所述字幕文本、所述情感指数和所述面部区域坐标在时序上对齐;利用所述映射关系,将所述短句序列和所述预置敏感语句进行对照匹配,得到对应的所述渲染集;根据所述渲染集的配置,将所述字幕文本、所述情感指数和话题特征数据与所述视频流渲染。
- 根据权利要求9所述的方法,其中,所述渲染集的配置,包括如下至少之一:渲染坐标;或渲染时间;或渲染颜色主题。
- 一种多媒体数据处理装置,包括:音频处理模块,被配置为接收并解析音频流,获得文本特征数据;视频处理模块,被配置为接收并解析视频流,得到表情特征数据;映射关系模块,被配置为处理所述文本特征数据得到话题特征数据,处理所述表情特征数据得到情感指数;渲染模块,被配置为将所述文本特征数据、所述话题特征数据和所述情感指数与所述视频流渲染。
- 一种电子设备,包括:存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,所述处理器执行所述计算机程序时实现如权利要求1至10中任意一项所述的多媒体数据处理方法。
- 一种计算机可读存储介质,存储有计算机可执行程序,所述计算机可执行程序用于使计算机执行如权利要求1至10任意一项所述的多媒体数据处理方法。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP23826533.4A EP4543012A4 (en) | 2022-06-24 | 2023-06-21 | METHOD AND APPARATUS FOR PROCESSING MULTIMEDIA DATA, AND COMPUTER-READABLE STORAGE MEDIUM |
| US18/878,212 US20250384605A1 (en) | 2022-06-24 | 2023-06-21 | Multimedia data processing method and apparatus, and computer-readable storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202210722924.7A CN117319701A (zh) | 2022-06-24 | 2022-06-24 | 多媒体数据处理方法、装置和计算机可读存储介质 |
| CN202210722924.7 | 2022-06-24 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023246888A1 true WO2023246888A1 (zh) | 2023-12-28 |
Family
ID=89235961
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2023/101786 Ceased WO2023246888A1 (zh) | 2022-06-24 | 2023-06-21 | 多媒体数据处理方法、装置和计算机可读存储介质 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US20250384605A1 (zh) |
| EP (1) | EP4543012A4 (zh) |
| CN (1) | CN117319701A (zh) |
| WO (1) | WO2023246888A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN119110023A (zh) * | 2024-11-08 | 2024-12-10 | 广州兴趣岛信息科技有限公司 | 一种ai自动拨打通话处理方法及系统 |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20250008290A1 (en) * | 2023-06-27 | 2025-01-02 | University Of Central Florida Research Foundation, Inc. | Spatially Explicit Auditory Cues for Enhanced Situational Awareness |
| CN119402699A (zh) * | 2024-10-31 | 2025-02-07 | 北京字跳网络技术有限公司 | 视频处理方法、装置、设备及介质 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103365959A (zh) * | 2013-06-03 | 2013-10-23 | 深圳市爱渡飞科技有限公司 | 一种语音搜索方法及装置 |
| CN106161873A (zh) * | 2015-04-28 | 2016-11-23 | 天脉聚源(北京)科技有限公司 | 一种视频信息提取推送方法及系统 |
| CN109886258A (zh) * | 2019-02-19 | 2019-06-14 | 新华网(北京)科技有限公司 | 提供多媒体信息的关联信息的方法、装置及电子设备 |
| CN110365933A (zh) * | 2019-05-21 | 2019-10-22 | 武汉兴图新科电子股份有限公司 | 一种基于ai的视频会议会议纪要在线生成装置及方法 |
| CN110414465A (zh) * | 2019-08-05 | 2019-11-05 | 北京深醒科技有限公司 | 一种视频通讯的情感分析方法 |
| CN114095782A (zh) * | 2021-11-12 | 2022-02-25 | 广州博冠信息科技有限公司 | 一种视频处理方法、装置、计算机设备及存储介质 |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10049263B2 (en) * | 2016-06-15 | 2018-08-14 | Stephan Hau | Computer-based micro-expression analysis |
| US10791078B2 (en) * | 2017-07-30 | 2020-09-29 | Google Llc | Assistance during audio and video calls |
| US20210076002A1 (en) * | 2017-09-11 | 2021-03-11 | Michael H Peters | Enhanced video conference management |
| CN110798636B (zh) * | 2019-10-18 | 2022-10-11 | 腾讯数码(天津)有限公司 | 字幕生成方法及装置、电子设备 |
-
2022
- 2022-06-24 CN CN202210722924.7A patent/CN117319701A/zh active Pending
-
2023
- 2023-06-21 WO PCT/CN2023/101786 patent/WO2023246888A1/zh not_active Ceased
- 2023-06-21 EP EP23826533.4A patent/EP4543012A4/en active Pending
- 2023-06-21 US US18/878,212 patent/US20250384605A1/en active Pending
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103365959A (zh) * | 2013-06-03 | 2013-10-23 | 深圳市爱渡飞科技有限公司 | 一种语音搜索方法及装置 |
| CN106161873A (zh) * | 2015-04-28 | 2016-11-23 | 天脉聚源(北京)科技有限公司 | 一种视频信息提取推送方法及系统 |
| CN109886258A (zh) * | 2019-02-19 | 2019-06-14 | 新华网(北京)科技有限公司 | 提供多媒体信息的关联信息的方法、装置及电子设备 |
| CN110365933A (zh) * | 2019-05-21 | 2019-10-22 | 武汉兴图新科电子股份有限公司 | 一种基于ai的视频会议会议纪要在线生成装置及方法 |
| CN110414465A (zh) * | 2019-08-05 | 2019-11-05 | 北京深醒科技有限公司 | 一种视频通讯的情感分析方法 |
| CN114095782A (zh) * | 2021-11-12 | 2022-02-25 | 广州博冠信息科技有限公司 | 一种视频处理方法、装置、计算机设备及存储介质 |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP4543012A4 * |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN119110023A (zh) * | 2024-11-08 | 2024-12-10 | 广州兴趣岛信息科技有限公司 | 一种ai自动拨打通话处理方法及系统 |
Also Published As
| Publication number | Publication date |
|---|---|
| US20250384605A1 (en) | 2025-12-18 |
| EP4543012A1 (en) | 2025-04-23 |
| CN117319701A (zh) | 2023-12-29 |
| EP4543012A4 (en) | 2025-09-10 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Gleason et al. | Making memes accessible | |
| Ou et al. | Multimodal local-global attention network for affective video content analysis | |
| EP4543012A1 (en) | Multimedia data processing method and apparatus, and computer-readable storage medium | |
| CN117252259A (zh) | 基于深度学习的自然语言理解方法及ai助教系统 | |
| Baur et al. | eXplainable cooperative machine learning with NOVA | |
| US10062384B1 (en) | Analysis of content written on a board | |
| CN119904901A (zh) | 基于大模型的情绪识别方法以及相关装置 | |
| US12288379B1 (en) | Common sense reasoning for deepfake detection | |
| CN116977992A (zh) | 文本信息识别方法、装置、计算机设备和存储介质 | |
| Wang et al. | Automatic Chinese meme generation using deep neural networks | |
| CN119831057A (zh) | 基于情感识别的ai对话系统 | |
| CN119402728A (zh) | 基于文本的视频生成方法和装置、电子设备、存储介质 | |
| CN112084788A (zh) | 一种影像字幕隐式情感倾向自动标注方法及系统 | |
| CN114741472B (zh) | 辅助绘本阅读的方法、装置、计算机设备及存储介质 | |
| Diamantini et al. | Automatic annotation of corpora for emotion recognition through facial expressions analysis | |
| CN119049051B (zh) | 一种基于文本数据的图像生成方法、系统及装置 | |
| CN118916707A (zh) | 多模态人格分析方法、分析模型、可读存储介质及装置 | |
| CN118428340A (zh) | 作文批改方法、装置、电子设备及存储介质 | |
| CN116842218A (zh) | 全自动生成视频会议纪要的方法以及存储介质 | |
| Nivethika et al. | Roberta-Powered Sentiment Detection in Natural Language Processing | |
| Letaifa et al. | The CG-MER dyadic multimodal dataset for spontaneous french conversations: annotation, analysis and assessment benchmark | |
| Lakshmi et al. | Virtual Caricature Assistant using Voice Analysis and | |
| Poornima et al. | Police sketch generator using generative adversarial networks | |
| CN120950901B (zh) | 基于多模态数据的员工情绪指数评测方法以及装置 | |
| Punj et al. | Detection of emotions with deep learning |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23826533 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18878212 Country of ref document: US |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2023826533 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2023826533 Country of ref document: EP Effective date: 20250120 |
|
| WWP | Wipo information: published in national office |
Ref document number: 2023826533 Country of ref document: EP |
|
| WWP | Wipo information: published in national office |
Ref document number: 18878212 Country of ref document: US |