WO2021043053A1 - 一种基于人工智能的动画形象驱动方法和相关装置 - Google Patents
一种基于人工智能的动画形象驱动方法和相关装置 Download PDFInfo
- Publication number
- WO2021043053A1 WO2021043053A1 PCT/CN2020/111615 CN2020111615W WO2021043053A1 WO 2021043053 A1 WO2021043053 A1 WO 2021043053A1 CN 2020111615 W CN2020111615 W CN 2020111615W WO 2021043053 A1 WO2021043053 A1 WO 2021043053A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- expression
- base
- target
- expression base
- speaker
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T13/00—Animation
- G06T13/20—Three-dimensional [3D] animation
- G06T13/40—Three-dimensional [3D] animation of characters, e.g. humans, animals or virtual beings
-
- A—HUMAN NECESSITIES
- A63—SPORTS; GAMES; AMUSEMENTS
- A63F—CARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
- A63F13/00—Video games, i.e. games using an electronically generated display having two or more dimensions
- A63F13/50—Controlling the output signals based on the game progress
- A63F13/54—Controlling the output signals based on the game progress involving acoustic signals, e.g. for simulating revolutions per minute [RPM] dependent engine sounds in a driving game or reverberation against a virtual wall
-
- A—HUMAN NECESSITIES
- A63—SPORTS; GAMES; AMUSEMENTS
- A63F—CARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
- A63F13/00—Video games, i.e. games using an electronically generated display having two or more dimensions
- A63F13/55—Controlling game characters or game objects based on the game progress
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/011—Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/017—Gesture based interaction, e.g. based on a set of recognized hand gestures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/16—Sound input; Sound output
- G06F3/167—Audio in a user interface, e.g. using voice commands for navigating, audio feedback
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T13/00—Animation
- G06T13/20—Three-dimensional [3D] animation
- G06T13/205—Three-dimensional [3D] animation driven by audio data
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
- G06V20/46—Extracting features or characteristics from the video content, e.g. video fingerprints, representative shots or key frames
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
- G06V40/16—Human faces, e.g. facial parts, sketches or expressions
- G06V40/174—Facial expression recognition
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/02—Feature extraction for speech recognition; Selection of recognition unit
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/06—Transformation of speech into a non-audible representation, e.g. speech visualisation or speech processing for tactile aids
- G10L21/10—Transforming into visible information
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/02—Feature extraction for speech recognition; Selection of recognition unit
- G10L2015/025—Phonemes, fenemes or fenones being the recognition units
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/06—Transformation of speech into a non-audible representation, e.g. speech visualisation or speech processing for tactile aids
- G10L21/10—Transforming into visible information
- G10L2021/105—Synthesis of the lips movements from speech, e.g. for talking heads
Definitions
- This application relates to the field of data processing, especially to the driving of animated images.
- the interactive device can determine the reply content according to the text or voice input by the user, and play a virtual sound synthesized according to the reply content.
- the present application provides an artificial intelligence-based animation image driving method and device, which brings a realistic sense of substitution and immersion to the user, and improves the user's experience of interacting with the animation image.
- an embodiment of the present application provides a method for driving an animated character, the method is executed by an audio and video processing device, and the method includes:
- the acoustic features and target expression parameters corresponding to the target text information are determined; the acoustic features are used to identify and simulate the speaker uttering the target text
- the voice of the information, the target expression parameter is used to identify the degree of change of the facial expression that simulates the speaker uttering the target text information with respect to the first expression base;
- a second animation character with a second expression base is driven.
- an embodiment of the present application provides an animation image driving device, the device is deployed on an audio and video processing device, and the device includes an acquiring unit, a first determining unit, a second determining unit, and a driving unit:
- the acquiring unit is configured to acquire media data containing the facial expression of the speaker and the corresponding voice
- the first determining unit is configured to determine a first expression base of the first animated image corresponding to the speaker according to the facial expression, and the first expression base is used to identify the expression of the first animated image;
- the second determining unit is configured to determine the acoustic characteristics and target expression parameters corresponding to the target text information according to the target text information, the media data, and the first expression base; the acoustic characteristics are used to identify the simulation location The voice of the speaker uttering the target text information, and the target expression parameter is used to identify the degree of change of the facial expression that simulates the speaker uttering the target text information with respect to the first expression base;
- the driving unit is configured to drive a second animation character with a second expression base according to the acoustic characteristics and the target expression parameters.
- an embodiment of the present application provides a method for driving an animated character, the method is executed by an audio and video processing device, and the method includes:
- the first expression base of the first animated image corresponding to the speaker is determined according to the facial expression, the first expression base is used to identify the expression of the first animated image; the dimension of the first expression base Is the first dimension, and the vertex topology is the first vertex topology;
- the target expression base is determined;
- the dimension of the second expression base is the second dimension, and the vertex topology is the second vertex topology, so
- the target expression base is the expression base corresponding to the first animation image with the second vertex topology, and the dimension of the target expression base is the second dimension;
- the target expression parameters are used to identify that the speaker uttered the voice The degree of change of the facial expression relative to the target expression base;
- the second animation character with the second expression base is driven.
- an embodiment of the present application provides an animation image driving device, the device is deployed on an audio and video processing device, and the device includes an acquisition unit, a first determination unit, a second determination unit, a third determination unit, and a driver unit:
- the acquiring unit is configured to acquire first media data including the facial expression of the speaker and the corresponding voice;
- the first determining unit is configured to determine a first expression base of the first animated image corresponding to the speaker according to the facial expression, and the first expression base is used to identify the expression of the first animated image;
- the dimension of the first expression base is the first dimension, and the vertex topology is the first vertex topology;
- the second determining unit is configured to determine a target expression base according to the first expression base and the second expression base of the second animation image to be driven;
- the dimension of the second expression base is the second dimension
- the vertex topology is the second vertex topology
- the target expression base is the expression base corresponding to the first animation image with the second vertex topology
- the dimension of the target expression base is the second dimension
- the third determining unit is configured to determine target expression parameters and acoustic features according to the second media data containing the facial expression and corresponding voice of the speaker and the target expression base; the target expression parameters are used to identify The degree of change of the facial expression of the speaker speaking the voice relative to the target expression base;
- the driving unit is configured to drive the second animation character with the second expression base according to the target expression parameters and acoustic characteristics.
- an embodiment of the present application provides a device for driving animated characters, the device including a processor and a memory:
- the memory is used to store program code and transmit the program code to the processor
- the processor is configured to execute the method described in the first aspect or the third aspect according to instructions in the program code.
- an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium is used to store program code, and the program code is used to execute the method described in the first aspect or the third aspect.
- embodiments of the present application provide a computer program product, which is used to execute the method described in the first aspect or the third aspect when the computer program product is executed.
- the first expression base of the first animation image corresponding to the speaker can be determined, and the first expression base can reflect the first animation image.
- Different expressions After determining the target text information used to drive the second animated image, the acoustic features and target expression parameters corresponding to the target text information can be determined according to the target text information, the aforementioned collected media data, and the first expression base.
- the acoustic characteristics can be Identifies the voice that simulates the speaker uttering the target text information, and the target expression parameter may identify the degree of change of the facial expression simulating the speaker uttering the target text information with respect to the first expression base.
- the second animation image with the second expression base can be driven, so that the second animation image can simulate the voice of the speaker uttering the target text information through the acoustic characteristics, and make conformance during the vocalization process.
- the speaker should have an expressive facial expression, which brings a realistic sense of substitution and immersion to the user, and improves the user's experience of interacting with the animated image.
- FIG. 1 is a schematic diagram of an application scenario of an artificial intelligence-based animation image driving method provided by an embodiment of the application
- FIG. 2 is a flowchart of an artificial intelligence-based animation image driving method provided by an embodiment of the application
- FIG. 3 is a structural flow of an animation image driving system provided by an embodiment of the application.
- FIG. 4 is an example diagram of a scene for collecting media data provided by an embodiment of the application
- FIG. 5 is an example diagram of each dimension distribution and meaning of the 3DMM library M provided by an embodiment of the application.
- FIG. 6 is a schematic diagram of an application scenario of an animation image driving method based on determining face pinching parameters provided by an embodiment of the application;
- FIG. 7 is a schematic diagram of an application scenario of an animation image driving method based on determining a mapping relationship provided by an embodiment of the application;
- FIG. 8 is an exemplary diagram of the correspondence between time intervals and phonemes provided by an embodiment of the application.
- FIG. 9 is a flowchart of an artificial intelligence-based animation image driving method provided by an embodiment of the application.
- Fig. 10a is a flowchart of an artificial intelligence-based animation image driving method provided by an embodiment of the application.
- Fig. 10b is a structural diagram of an animated character driving device provided by an embodiment of the application.
- FIG. 11 is a structural diagram of an animated character driving device provided by an embodiment of the application.
- FIG. 12 is a structural diagram of a device for driving animated characters provided by an embodiment of the application.
- FIG. 13 is a structural diagram of a server provided by an embodiment of this application.
- the main research direction of human-computer interaction is to use animated images with the ability to change expressions as the interactive objects of interaction with users.
- a game character (animated image) with the same face shape as the user's own face can be constructed.
- the game character can make a voice and make a corresponding expression (such as mouth shape, etc.);
- a game character with the same face shape as the user's own.
- the opponent inputs text or voice
- the game character can respond to the voice according to the opponent's input and make a corresponding expression.
- an embodiment of the present application provides an artificial intelligence-based method for driving an animated image.
- This method can determine the first expression base of the first animated image corresponding to the speaker by collecting the media data of the facial expression changes when the speaker speaks the voice.
- the acoustic characteristics and target expression parameters corresponding to the target text information can be determined according to the target text information, the aforementioned collected media data and the first expression base, so as to drive the second animation image with the second expression base through the acoustic characteristics and the target expression parameters.
- the second animation image is made to simulate the voice of the speaker uttering the target text information through the acoustic feature, and a facial expression conforming to the expected expression of the speaker is made during the vocalization process, so that the second animation image is driven based on the text information.
- AI Artificial Intelligence
- theory, methods, technologies, and application systems that perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
- artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a similar way to human intelligence.
- Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making.
- Artificial intelligence technology is a comprehensive discipline, covering a wide range of fields, including both hardware-level technology and software-level technology.
- Basic artificial intelligence technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation/interaction systems, and mechatronics.
- Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning/deep learning.
- the artificial intelligence technologies mainly involved include speech processing technology, machine learning, and computer vision (image).
- Speech recognition technology includes speech signal preprocessing (Speech signal preprocessing), speech signal frequency domain analysis (Speech signal frequency analyzing), speech signal feature extraction (Speech signal feature extraction), speech signal feature matching/recognition (Speech signal feature matching/ recognition), speech training, etc.
- Speech synthesis includes text analysis and speech generation.
- Machine learning is a multi-field interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other subjects. Specializing in the study of how computers simulate or realize human learning behaviors in order to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve its own performance.
- Machine learning is the core of artificial intelligence, the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence.
- Machine learning usually includes deep learning (Deep Learning) and other technologies. Deep learning includes artificial neural networks, such as Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), and deep learning. Deep neural network (DNN), etc.
- CNN Convolutional Neural Network
- RNN Recurrent Neural Network
- DNN Deep neural network
- Video processing includes target recognition, target detection and localization, etc.
- face recognition includes face 3D reconstruction (Face 3D Reconstruction), face detection (Face Detection), and face tracking (Face Tracking) etc.
- the artificial intelligence-based animation image driving method provided in the embodiments of the present application can be applied to an audio and video processing device with the ability to drive an animation image, and the audio and video processing device may be a terminal device or a server.
- the audio and video processing equipment can have the ability to implement voice technology, so that the audio and video processing equipment can listen, see, and feel, which is the future development direction of human-computer interaction, and voice becomes one of the most promising human-computer interaction methods in the future .
- the audio and video processing device can determine the first expression base of the first animated image corresponding to the speaker in the media data by implementing the above-mentioned computer vision technology, and can be based on the target text information and media data through voice technology and machine learning. , Determine the acoustic characteristics and target expression parameters corresponding to the target text information, and then use the acoustic characteristics and target expression parameters to drive the second animation image with the second expression base.
- the audio and video processing device is a terminal device
- the terminal device may be a smart terminal, a computer, a personal digital assistant (Personal Digital Assistant, PDA for short), a tablet computer, etc.
- PDA Personal Digital Assistant
- the server may be an independent server or a cluster server.
- the terminal device can upload media data containing the facial expression of the speaker and the corresponding voice to the server.
- the server determines the acoustic characteristics and target expression parameters, and uses the acoustic characteristics and target expression parameters to drive the terminal device The second animated image on top.
- the artificial intelligence-based animation image driving method provided by the embodiments of the present application can be applied to various application scenarios applicable to animation images, such as news broadcast, weather forecast, game commentary, and game scenes that are allowed to be used for construction and users.
- Game characters with the same face shape can also be used in scenes that use animated images to undertake personal services, such as one-to-one services for individuals such as psychologists and virtual assistants.
- the method provided in the embodiments of the present application can be used to drive the animated image.
- FIG. 1 is a schematic diagram of an application scenario of an artificial intelligence-based animation image driving method provided by an embodiment of the application.
- This application scenario is introduced by taking an audio and video processing device as a terminal device as an example.
- the application scenario includes a terminal device 101, and the terminal device 101 can obtain media data containing the facial expression of the speaker and the corresponding voice.
- the media data can be one or multiple.
- the media data can be video, video and audio.
- the languages corresponding to the characters included in the voice in the media data can be Chinese, English, Korean and other languages.
- Facial expressions can be actions made by the face when the speaker speaks the voice, for example, it can include mouth shape, eye movements, eyebrow movements, etc.
- the video viewer can feel the voice in the media data through the face expression of the speaker That's what the speaker said.
- the terminal device 101 can determine the first expression base of the first animation image corresponding to the speaker according to the facial expressions, and the first expression base is used to identify different expressions of the first animation image.
- an expression form of the expression parameter and the face pinching parameter that may be involved later may be a coefficient, for example, a vector having a certain dimension.
- the acoustic features and target expression parameters are obtained based on the media data, and they correspond to the same time axis. Therefore, the sound identified by the acoustic features and the target expression parameters are identified 'S expression changes synchronously on the same timeline.
- the generated acoustic feature is a sequence related to the time axis
- the target expression parameter is a sequence related to the same time axis, and the two can be adjusted accordingly as the text information changes.
- the acoustic feature is used to identify the voice that simulates the target text information of the speaker in the above-mentioned media data
- the target expression parameter is used to identify the facial expression that simulates the target text information of the speaker in the above-mentioned media data. State the degree of change of the first expression base.
- the terminal device 101 can drive the second animation image with the second expression base through the acoustic characteristics and the target expression parameters, so that the second animation image can simulate the sound of the speaker uttering the target text information through the acoustic characteristics, and is uttering
- a facial expression that conforms to the expression of the speaker is made.
- the second animation image may be the same animation image as the first animation image, or it may be an animation image different from the first animation image, which is not limited in the embodiment of the application.
- the method includes:
- S201 Acquire media data including the facial expression of the speaker and the corresponding voice.
- Media data containing facial expressions and corresponding voices can be obtained by recording the voice spoken by the speaker in a recording environment with a camera, and recording the facial expressions of the speaker corresponding to the speaker through the camera.
- the media data is the video; if the video captured by the camera includes the facial expression of the speaker, and the voice is through other devices such as
- the media data collected by the recording device includes video and audio. At this time, the video and audio are collected simultaneously.
- the video includes the facial expression of the speaker, and the audio includes the speaker's voice.
- the method provided in the embodiments of the present application can be implemented by an animation image driving system.
- the system can be seen in Figure 3 and mainly includes four parts, namely a data acquisition module, a face modeling module, an acoustic feature, and Expression parameter determination module and animation drive module.
- the data acquisition module is used to perform S201
- the face modeling module is used to perform S202
- the acoustic feature and expression parameter determination module is used to perform S203
- the animation drive module is used to perform S204.
- the media data including the facial expression of the speaker and the corresponding voice can be obtained through the data collection module.
- the data collection module can have more choices.
- the data collection module can collect media data including the speaker's voice and facial expressions through professional equipment, such as motion capture systems, facial expression capture systems and other professional equipment to capture speech Human facial expressions, facial expressions can be, for example, facial movements, expressions, mouth shapes, etc., use professional recording equipment to record the speaker’s voice, and synchronize the data between the voice and facial expressions by triggering synchronization signals between different devices and many more.
- professional equipment is not limited to the use of expensive capture systems. It can also be a multi-view ultra-high-definition device.
- the multi-view ultra-high-definition device collects videos including the speaker's voice and facial expressions.
- the data collection module can also collect media data including the speaker's voice and facial expressions in a multi-camera surround mode.
- the collection environment needs to have stable ambient lighting, and the speaker is not required to wear specific clothes.
- Figure 4 takes three ultra-high-definition cameras as an example.
- the upper dashed arrow represents stable lighting, and the three arrows on the left represent the relationship between the viewing angle of the ultra-high-definition camera and the speaker, so as to collect the voice and face of the speaker.
- Media data of the facial expressions At this time, the video captured by the ultra-high-definition camera can include both voice and facial expressions, that is, the media data is a video.
- the form of expression of the collected media data may be different according to the different sensors used to collect facial expressions.
- the establishment of a face model can be achieved by shooting a speaker with a Red Green Blue Deep (RGBD) sensor. Since the RGBD sensor can collect depth information and obtain the speaker's three-dimensional reconstruction result, the media data includes the static modeling of the speaker's corresponding face, that is, 3 dimensions (3D) data. In other cases, there may be no RGBD sensor but a two-dimensional sensor to take pictures of the speaker. At this time, there is no 3D reconstruction result of the speaker.
- the media data includes the video frame corresponding to the speaker, that is, 2 Dimensions (2 Dimensions). , 2D) data.
- S202 Determine the first expression base of the first animation image corresponding to the speaker according to the facial expression.
- the face modeling module in FIG. 3 can be used to model the face of the speaker, thereby obtaining the first expression base of the first animation image corresponding to the speaker, and the first expression base is used To identify the expression of the first animated image.
- the purpose of modeling the face is to make the collected objects, such as the aforementioned speaker, be understood and stored by the computer, including the shape and texture of the collected objects.
- face modeling mainly from three perspectives: hardware, manual, and software.
- the hardware perspective can be realized by using professional equipment to perform high-precision scanning of the speaker, such as a 3D scanning instrument, and the obtained face model can be selected to manually/automatically clean up the data;
- the artificial perspective can be realized by manually designing the data by the art designer , Clean up the data, adjust the data;
- the software perspective can be realized by using the parameterized face pinching algorithm to automatically generate the speaker's face model.
- a professional face scanning device can be used to scan a speaker with an expression, and a parameterized description of the current expression will be automatically given. This description is related to the custom expression description in the scanning device.
- the expression parameters manually adjusted by the art designer it is generally necessary to predefine the expression type and the corresponding face parameterization, such as the opening and closing degree of the mouth, the movement range of the facial muscles, and so on.
- the mathematical description of the face in different expressions For example, by decomposing a large amount of real face data, the principal component analysis (PCA) method can be used to obtain the best expression.
- PCA principal component analysis
- the software-based face modeling and expression parameterization are mainly introduced.
- the mathematical description of the face in different expressions can be defined through the model library.
- the animation images in the embodiments of the present application may be models in the model library, or may be obtained by linear combination of the models in the model library.
- the model library may be a human face 3D deformable model (3DMM) library or other model libraries, which is not limited in this implementation.
- the animated image can be a 3D grid.
- the 3DMM library is obtained from a large amount of high-precision face data through the principal component analysis method, and describes the main changes of high-dimensional face shape and expression relative to the average face, and can also describe texture information.
- the 3DMM library when the 3DMM library describes an expressionless face, it can be obtained by mu+ ⁇ (Pface i -mu)* ⁇ i .
- mu is the average face under natural expression
- Pface i is the i-th face principal component component
- ⁇ i is the weight of each face principal component component, that is, the face pinch parameter.
- the grid corresponding to the animation image in the 3DMM library can be represented by M, that is, the relationship between the face shape, expression and vertices in the 3DMM library is represented by M, and M is a three-dimensional matrix of [m ⁇ n ⁇ d], where each One dimension is the vertex coordinates of the grid (m), the principal component of face shape (n), and the principal component of expression (d).
- M is a three-dimensional matrix of [m ⁇ n ⁇ d], where each One dimension is the vertex coordinates of the grid (m), the principal component of face shape (n), and the principal component of expression (d).
- the distribution and meaning of each dimension of the 3DMM library M is shown in Figure 5, and each coordinate axis represents the vertex coordinates (m), the principal component of face (n), and the principal component of expression (d).
- M can be a two-dimensional matrix.
- the texture dimension in the 3DMM library is not considered, and assuming that the driving of the animated image is F, then:
- M is the grid of the animated image
- ⁇ is the face pinching parameter
- ⁇ is the expression parameter
- n is the number of face pinching grids in the face pinching base
- d is the number of expression grids in the expression base
- M k, j,i is the k-th grid with the i-th expression grid and the j-th face pinching grid
- ⁇ j is the j-th dimension in a set of face pinching parameters, representing the weight of the j-th face principal component component
- ⁇ i is the i-th dimension in a set of expression parameters, and represents the weight of the i-th expression principal component.
- the process of determining the face pinching parameters is the face pinching algorithm
- the process of determining the expression parameters is the face pinching algorithm.
- the pinching parameters are used to make a linear combination with the pinching base to obtain the corresponding face shape.
- a pinching base including 50 pinching grids (belonging to deformable grids, such as blendshape), and the pinching face corresponding to the pinching base
- the parameter is a 50-dimensional vector, and each dimension can identify the degree of correlation between the face shape corresponding to the face pinching parameter and a face pinching grid.
- the face pinch grids included in the face pinch base represent different face shapes, and each face pinch grid is a face image with a relatively large change from the average face. It is a face shape master of different dimensions obtained after a large number of faces are decomposed by PCA. Component, and the number of vertices corresponding to different face-pinching meshes in the same face-pinching base remains the same.
- the expression parameters are used to linearly combine with the expression base to obtain the corresponding expression.
- an expression base including 50 (equivalent to 50) expression grids (belonging to deformable grids, such as blendshape), which corresponds to
- the expression parameter of is a 50-dimensional vector, and each dimension can identify the degree of correlation between the expression corresponding to the expression parameter and an expression grid.
- the expression grids included in the expression base represent different expressions. Each expression grid is formed by changing the same 3D model under different expressions. The number of vertices corresponding to different expression grids in the same expression base remains the same.
- a single grid can be deformed through a predefined shape to obtain any number of grids.
- the first expression base of the first animated image corresponding to the speaker can be obtained, which can be used to drive the subsequent second animated image.
- S203 Determine acoustic features and target expression parameters corresponding to the target text information according to the target text information, the media data, and the first expression base.
- the acoustic feature and expression parameter determination module in FIG. 3 can determine the acoustic characteristics and target expression parameters corresponding to the target text information. Among them, the acoustic feature is used to identify the voice of the simulated speaker uttering the target text information, and the target expression parameters are used to identify the degree of change of the facial expression of the simulated speaker uttering the target text information with respect to the first expression base.
- the target text information may be obtained in multiple ways.
- the target text information may be input by the user through the terminal device, or may be obtained by converting the voice input to the terminal device.
- the expressions identified by the target expression parameters are combined with the voices identified by the acoustic features to be displayed in a way that humans can intuitively understand, using multiple senses.
- the target expression parameter represents the weight of each expression grid in the second expression base, and the corresponding expression can be obtained through the weighted linear combination of the second expression base.
- the second animation image that makes the expression corresponding to the voice is rendered through the rendering method, so as to realize the driving of the second animation image.
- the first expression base of the first animation image corresponding to the speaker can be determined, and the first expression base can reflect the image of the first animation image.
- Different expressions After determining the target text information used to drive the second animated image, the acoustic features and target expression parameters corresponding to the target text information can be determined according to the target text information, the aforementioned collected media data, and the first expression base.
- the acoustic characteristics can be Identifies the voice that simulates the speaker uttering the target text information, and the target expression parameter may identify the degree of change of the facial expression simulating the speaker uttering the target text information with respect to the first expression base.
- the second animation image with the second expression base can be driven, so that the second animation image can simulate the voice of the speaker uttering the target text information through the acoustic characteristics, and make conformance during the vocalization process.
- the speaker should have an expressive facial expression, which brings a realistic sense of substitution and immersion to the user, and improves the user's experience of interacting with the animated image.
- S203 may include multiple implementations, and the embodiment of the present application focuses on one implementation.
- the implementation manner of S203 may be to determine the acoustic feature and expression feature corresponding to the target text information according to the target text information and media data.
- the acoustic feature is used to identify the voice of the simulated speaker uttering the target text information
- the expression feature is used to identify the facial expression of the simulated speaker uttering the target text information. Then, the target expression parameters are determined according to the first expression base and expression characteristics.
- the facial expression and voice of the speaker have been synchronously recorded in the media data, that is, the facial expression and voice of the speaker in the media data correspond to the same time axis. Therefore, a large amount of media data can be pre-collected offline as training data, and text features, acoustic features, and expression features can be extracted from these media data, and the duration model, acoustic model, and expression model can be trained based on these features.
- the duration model can be used to determine the duration corresponding to the target text information, and then the duration and the corresponding text characteristics of the target text information are determined by the acoustic model and the expression model respectively Corresponding acoustic features and expression features. Since the acoustic features and expression features are based on the time length obtained by the same time length model, it is easy to synchronize the voice and expression, so that the second animation image simulates the speech while simulating the speaker to say the target text information corresponding to the voice. People make corresponding expressions.
- the second animation image may be the same animation image as the first animation image, or it may be an animation image different from the first animation image. In these two cases, the implementation of S204 may be different.
- the first case the first animation image and the second animation image are the same animation image.
- the animation image to be driven is the first animation image.
- the first animation image in order to drive the first animation image, in addition to determining the first expression base, it is also necessary to determine the pinch face parameters of the first animation image to obtain the face shape of the first animation image. Therefore, in S202, the first expression base of the first animated image and the pinching parameters of the first animated image can be determined according to the facial expressions.
- the pinching parameters are used to identify that the face of the first animated image is relative to the first animated image. Corresponding to the degree of change of the pinch face base.
- the first expression base of the first animation image and the pinching parameters of the first animation image There are many ways to determine the first expression base of the first animation image and the pinching parameters of the first animation image.
- the collected media data often has low accuracy and high noise, which makes the quality of the established face model not high and has many uncertainties, making it difficult to build a face model.
- an embodiment of the present application provides a method for determining face pinching parameters, as shown in FIG. 6.
- the initial pinching parameters can be determined based on the first vertex data among them and the target vertex data used to identify the target facial model in the 3DMM library .
- the expression parameters are determined based on the initial pinching parameters and the target vertex data. After that, the expression parameters are fixed, and the pinching parameters are reversed or reversed. Push how to change the face shape to obtain the face image of the speaker under the expression parameters, that is, to correct the initial pinch parameters by fixing the expression to reverse the face shape to obtain the target pinch parameters, and then use the target pinch parameters as the first animation Image pinch face parameters.
- the target pinching parameters corrected by the second vertex data can offset the noise in the first vertex data to a certain extent, and the face model corresponding to the speaker determined by the target pinching parameters The accuracy is relatively higher.
- the determined target expression parameters can directly drive the second animation image. Therefore, the second animation is driven in S204
- the way of the image may be to drive the second animation image with the second expression base according to the acoustic characteristics, target expression parameters and face pinch parameters.
- the second case the first animation image and the second animation image are different animation images.
- the first expression base is different from the second expression base, that is, the dimensions of the two and the semantic information of each dimension are different, so it is difficult to directly use the target expression parameters to drive the second animation with the second expression base Image.
- the expression parameters corresponding to the first animation image and the expression parameters corresponding to the second animation image should have a mapping relationship, the mapping relationship between the expression parameters corresponding to the first animation image and the expression parameters corresponding to the second animation image can be through the function f() ,
- the formula for calculating the expression parameters corresponding to the second animation image from the expression parameters corresponding to the first animation image is as follows:
- ⁇ b is the expression parameter corresponding to the second animation image
- ⁇ a is the expression parameter corresponding to the first animation image
- f() represents the mapping between the expression parameters corresponding to the first animation image and the expression parameters corresponding to the second animation image. relationship.
- the mapping relationship may be a linear mapping relationship or a non-linear mapping relationship.
- mapping relationship In order to drive the second animation image with the second expression base according to the target expression parameters, the mapping relationship needs to be determined. There may be multiple ways to determine the mapping relationship, and this embodiment mainly introduces two ways of determining.
- the first determination method may be to determine the mapping relationship between the expression parameters based on the first expression base corresponding to the first animation image and the second expression base corresponding to the second animation image.
- the actual expression parameters can reflect the degree of correlation between the actual expression and its expression base in different dimensions, that is, the first
- the actual expression parameters corresponding to the second animation image can also reflect the degree of correlation between the actual expression of the second animation image and its expression base in different dimensions. Therefore, based on the above-mentioned correlation between the expression parameters and the expression base, it can be based on the corresponding relationship of the first animation image.
- the first expression base and the second expression base corresponding to the second animation image determine the mapping relationship between the expression parameters. Then, according to the acoustic characteristics, the target expression parameters and the mapping relationship, the second animation image with the second expression base is driven.
- the second determination method may be to determine the mapping relationship between the expression parameters based on the preset relationship between the phoneme and the second expression base.
- Phoneme is the smallest phonetic unit divided according to the natural attributes of the speech. It is analyzed according to the pronunciation actions in the syllable.
- An action (such as mouth shape) constitutes a phoneme.
- the phoneme has nothing to do with the speaker, no matter who the speaker is, whether the speech is English or Chinese, whether the text corresponding to the phoneme is the same, as long as the phoneme in a time interval in the speech is the same, then the corresponding expression is for example
- the mouth shape is consistent. Refer to Figure 8, which shows the correspondence between time intervals and phonemes, and describes which phoneme corresponds to which time interval in a speech.
- the corresponding video frames can be easily divided by voice, that is, the phoneme identified by the voice, the time interval corresponding to the phoneme, and the media data are determined according to the media data.
- the video frame of the time interval Then, the first expression parameter corresponding to the phoneme is determined according to the video frame, and the first expression parameter is used to identify the degree of change of the facial expression of the speaker relative to the first expression base when the phoneme is emitted.
- the corresponding time interval is 5.65 seconds to 6.3 seconds.
- the first expression parameter If the first animation image is an animation image a, the first expression parameter can be represented by ⁇ a .
- the dimension of the first group is n a expression
- the resulting parameter ⁇ a first expression of a set of vectors of length n a.
- the premise of the method of determining the mapping relationship is that the expression bases of other animated images, such as the second expression base corresponding to the second animated image, are generated according to the preset relationship with the phoneme, and the preset relationship indicates that one phoneme corresponds to one expression network.
- the phoneme "u" in the preset relationship corresponds to the first expression grid
- the phoneme "i" corresponds to the second expression grid...
- the second expression base including n b expression grids can be determined.
- the second expression parameter corresponding to the phoneme can be determined according to the preset relationship and the second expression base.
- the mapping relationship is determined according to the first expression parameter and the second expression parameter.
- the phoneme identified by the voice is "u”
- ⁇ b includes n b elements, except for the first element which is 1, the other n b -1 elements are all 0.
- the formula for determining the mapping relationship can be:
- f is the mapping relationship
- ⁇ A is the first matrix
- ⁇ B is the second matrix
- inv is the matrix inversion operation.
- the foregoing embodiment mainly introduced how to drive an animated image based on text information.
- the first animation image corresponding to the speaker in the media data has a first expression base
- the dimension of the first expression base is the first dimension
- the vertex topology is the first vertex topology.
- the first expression base can be represented by Ea
- One dimension can be represented by Na
- the first vertex topology can be represented by Ta
- the first expression base Ea looks like Fa
- the second animation image to be driven has a second expression base
- the dimension of the second expression base is second Dimension
- the vertex topology is the second vertex topology
- the second expression base can be represented by Eb
- the second dimension can be represented by Nb
- the second vertex topology can be represented by Tb.
- the second expression base Eb looks like Fb.
- the media data including the facial expression and voice of the speaker drives the second animated image.
- an embodiment of the present application also provides an artificial intelligence-based animation image driving method. As shown in FIG. 9, the method includes:
- S902 Determine the first expression base of the first animation image corresponding to the speaker according to the facial expression.
- S903 Determine the target expression base according to the first expression base and the second expression base of the second animation image to be driven.
- the dimension of the first expression base is different from the dimension of the second expression base, in order to use the facial expression and voice of the speaker in the media data to drive the second animation image, a new animation image can be constructed.
- the expression base of is, for example, the target expression base, so that the target expression base has the characteristics of the first expression base and the second expression base at the same time.
- the implementation manner of S903 may be: determine from the first expression base the corresponding expressionless grid when the first animation image is in the absence of expression, and determine from the second expression base that the second animation image is in the absence of expression.
- the expressionless grid corresponding to the expression According to the expressionless grid corresponding to the first image and the expressionless grid corresponding to the second image, an adjustment grid is determined, and the adjustment grid has a second vertex topology and is used to identify the first animated image in an expressionless state. According to the grid deformation relationship in the adjusted grid and the second expression base, the target expression base is generated.
- the flowchart of the method can also be referred to as shown in FIG. 10a.
- the target expression base Eb' is determined based on the first expression base Ea and the second expression base Eb. Wherein, the method of determining the target expression base Eb' may be to extract the expressionless grid of the second expression base Eb and the expressionless grid of the first expression base Ea.
- the expressionless mesh of Eb is pasted to the expressionless mesh of Ea, so that the expressionless mesh of Eb changes its appearance and becomes the appearance of Ea while maintaining the vertex topology Fb.
- Get the adjustment grid which can be expressed as Newb.
- the mesh deformation relationship of the expressions in each dimension in Newb and the second expression base Eb relative to the natural expression (no expression) is known, the mesh deformation relationship in Newb and the second expression base Eb can be obtained from The target expression base Eb' is transformed in Newb.
- the appearance of the target expression base Eb' is Fa
- the dimension is Nb
- the vertex topology is Tb.
- S904 Determine target expression parameters and acoustic features according to the second media data including the facial expression and corresponding voice of the speaker and the target expression base.
- the target expression parameter Bb is used to identify the degree of change of the facial expression of the speaker speaking the voice relative to the target expression base.
- target expression parameters and acoustic features obtained by this method can be used to retrain the aforementioned acoustic models and expression models.
- the terminal device can obtain the media data including the facial expression of the speaker and the corresponding voice, and determine the first expression base of the first animation image corresponding to the speaker according to the facial expression.
- the acoustic characteristics and target expression parameters corresponding to the target text information are determined, so as to drive the second animation with the second expression base according to the acoustic characteristics and target expression parameters
- the image makes the second animated image emit the voice corresponding to the target text information and make the corresponding expression.
- the user can see that the game character imitates the speaker's voice and makes a corresponding expression, which brings a realistic sense of substitution and immersion to the user, and improves the user's experience of interacting with the animated image.
- this embodiment also provides an animated character driving device 1000, which is deployed on audio and video processing equipment.
- the device 1000 includes an acquiring unit 1001, a first determining unit 1002, a second determining unit 1003, and a driving unit 1004:
- the acquiring unit 1001 is configured to acquire media data containing the facial expression of the speaker and the corresponding voice;
- the first determining unit 1002 is configured to determine a first expression base of the first animated image corresponding to the speaker according to the facial expression, and the first expression base is used to identify the expression of the first animated image ;
- the second determining unit 1003 is configured to determine acoustic features and target expression parameters corresponding to the target text information according to the target text information, the media data, and the first expression base; the acoustic characteristics are used to identify simulations The speaker utters the voice of the target text information, and the target expression parameter is used to identify the degree of change of the facial expression that simulates the speaker uttering the target text information with respect to the first expression base;
- the driving unit 1004 is configured to drive a second animation character with a second expression base according to the acoustic characteristics and the target expression parameters.
- the first animation image and the second animation image are the same animation image
- the first expression base is the same as the second expression base
- the first determining unit 1002 For:
- the first expression base of the first animated avatar and the pinch face parameters of the first animated avatar are determined according to the facial expressions, and the face pinch parameters are used to identify that the face of the first animated avatar is relative to the face shape of the first animated avatar.
- the driving unit 1004 is used for:
- the second animated image is driven according to the acoustic feature, the target expression parameter, and the face pinching parameter.
- the first animation image and the second animation image are different animation images
- the first expression base is different from the second expression base
- the driving unit 1004 is configured to :
- the second animated image is driven.
- the second expression base is generated according to a preset relationship between the second expression base and phonemes, and the driving unit 1004 is further configured to:
- the mapping relationship is determined according to the first expression parameter and the second expression parameter.
- the second determining unit 1003 is configured to:
- the target text information and the media data determine the acoustic characteristics and the expression characteristics corresponding to the target text information; the expression characteristics are used to identify facial expressions that simulate the speaker uttering the target text information;
- the target expression parameter is determined according to the first expression base and the expression feature.
- the device 1100 includes an acquiring unit 1101, a first determining unit 1102, a second determining unit 1103, a third determining unit 1104, and a driving unit 1105:
- the acquiring unit 1101 is configured to acquire first media data including the facial expression of the speaker and the corresponding voice;
- the first determining unit 1102 is configured to determine a first expression base of the first animated image corresponding to the speaker according to the facial expression, and the first expression base is used to identify the expression of the first animated image ;
- the dimension of the first expression base is the first dimension, and the vertex topology is the first vertex topology;
- the second determining unit 1103 is configured to determine a target expression base according to the first expression base and the second expression base of the second animation image to be driven; the dimension of the second expression base is the second dimension , The vertex topology is the second vertex topology, the target expression base is the expression base corresponding to the first animation image with the second vertex topology, and the dimension of the target expression base is the second dimension;
- the third determining unit 1104 is configured to determine target expression parameters and acoustic features according to the second media data containing the facial expression and corresponding voice of the speaker and the target expression base; the target expression parameters are used for Identifying the degree of change in the facial expression of the speaker uttering the voice relative to the target expression base;
- the driving unit 1105 is configured to drive the second animation character with the second expression base according to the target expression parameters and acoustic characteristics.
- the second determining unit 1103 is configured to determine from the first expression base the expressionless grid corresponding to the first animation image in the expressionless state, and obtain the expression from the first expression base. In the second expression base, determine the expressionless grid corresponding to the second animation image when it is expressionless;
- an adjustment grid is determined.
- the adjustment grid has a second vertex topology and is used to identify when in an expressionless state.
- the target expression base is generated.
- the embodiment of the present application also provides a device for driving an animation image.
- the device can drive the animation through voice, and the device can be an audio and video processing device.
- the equipment will be introduced below in conjunction with the drawings.
- an embodiment of the present application provides a device for driving animated characters.
- the device may also be a terminal device.
- the terminal device may include a mobile phone, a tablet computer, and a personal digital assistant (Personal Digital Assistant, Any smart terminal such as PDA), Point of Sales (POS), on-board computer, etc. Take the terminal device as a mobile phone as an example:
- FIG. 12 shows a block diagram of a part of the structure of a mobile phone related to a terminal device provided in an embodiment of the present application.
- the mobile phone includes: Radio Frequency (RF) circuit 1210, memory 1220, input unit 1230, display unit 1240, sensor 1250, audio circuit 1260, wireless fidelity (wireless fidelity, WiFi for short) module 1270, processing 1280, and power supply 1290 and other components.
- RF Radio Frequency
- the RF circuit 1210 can be used for receiving and sending signals during the process of sending and receiving information or talking. In particular, after receiving the downlink information of the base station, it is processed by the processor 1280; in addition, the designed uplink data is sent to the base station.
- the RF circuit 1210 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA for short), a duplexer, and the like.
- the RF circuit 1210 can also communicate with the network and other devices through wireless communication.
- the above-mentioned wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access ( Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), Email, Short Message Service (Short Messaging Service, SMS) Wait.
- GSM Global System of Mobile Communication
- GPRS General Packet Radio Service
- CDMA Code Division Multiple Access
- WCDMA Wideband Code Division Multiple Access
- LTE Long Term Evolution
- Email Short Message Service
- SMS Short Messaging Service
- the memory 1220 may be used to store software programs and modules.
- the processor 1280 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1220.
- the memory 1220 may mainly include a program storage area and a data storage area.
- the program storage area may store an operating system, an application program required by at least one function (such as a sound playback function, an image playback function, etc.), etc.; Data created by the use of mobile phones (such as audio data, phone book, etc.), etc.
- the memory 1220 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
- the input unit 1230 can be used to receive input digital or character information, and generate key signal input related to the user settings and function control of the mobile phone.
- the input unit 1230 may include a touch panel 1231 and other input devices 1232.
- the touch panel 1231 also called a touch screen, can collect the user's touch operations on or near it (for example, the user uses any suitable objects or accessories such as fingers, stylus, etc.) on the touch panel 1231 or near the touch panel 1231. Operation), and drive the corresponding connection device according to the preset program.
- the touch panel 1231 may include two parts: a touch detection device and a touch controller.
- the touch detection device detects the user's touch position, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it To the processor 1280, and can receive and execute the commands sent by the processor 1280.
- the touch panel 1231 can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave.
- the input unit 1230 may also include other input devices 1232.
- the other input device 1232 may include, but is not limited to, one or more of a physical keyboard, function keys (such as volume control buttons, switch buttons, etc.), trackball, mouse, and joystick.
- the display unit 1240 can be used to display information input by the user or information provided to the user and various menus of the mobile phone.
- the display unit 1240 may include a display panel 1241.
- the display panel 1241 may be configured in the form of a liquid crystal display (Liquid Crystal Display, LCD for short), Organic Light-Emitting Diode (OLED for short), etc.
- the touch panel 1231 can cover the display panel 1241. When the touch panel 1231 detects a touch operation on or near it, it is transmitted to the processor 1280 to determine the type of the touch event, and then the processor 1280 determines the type of the touch event. The type provides corresponding visual output on the display panel 1241.
- the touch panel 1231 and the display panel 1241 are used as two independent components to implement the input and input functions of the mobile phone, in some embodiments, the touch panel 1231 and the display panel 1241 can be integrated. Realize the input and output functions of the mobile phone.
- the mobile phone may also include at least one sensor 1250, such as a light sensor, a motion sensor, and other sensors.
- the light sensor can include an ambient light sensor and a proximity sensor.
- the ambient light sensor can adjust the brightness of the display panel 1241 according to the brightness of the ambient light.
- the proximity sensor can close the display panel 1241 and/or when the mobile phone is moved to the ear. Or backlight.
- the accelerometer sensor can detect the magnitude of acceleration in various directions (usually three-axis), and can detect the magnitude and direction of gravity when it is stationary.
- the audio circuit 1260, the speaker 1261, and the microphone 1262 can provide an audio interface between the user and the mobile phone.
- the audio circuit 1260 can transmit the electrical signal converted from the received audio data to the speaker 1261, which is converted into a sound signal for output by the speaker 1261; on the other hand, the microphone 1262 converts the collected sound signal into an electrical signal, which is then output by the audio circuit 1260. After being received, it is converted into audio data, and then processed by the audio data output processor 1280, and sent to, for example, another mobile phone via the RF circuit 1210, or the audio data is output to the memory 1220 for further processing.
- WiFi is a short-distance wireless transmission technology.
- the mobile phone can help users send and receive emails, browse web pages, and access streaming media through the WiFi module 1270. It provides users with wireless broadband Internet access.
- FIG. 12 shows the WiFi module 1270, it is understandable that it is not a necessary component of the mobile phone, and can be omitted as needed without changing the essence of the invention.
- the processor 1280 is the control center of the mobile phone. It uses various interfaces and lines to connect various parts of the entire mobile phone. It executes by running or executing software programs and/or modules stored in the memory 1220, and calling data stored in the memory 1220. Various functions and processing data of the mobile phone can be used to monitor the mobile phone as a whole.
- the processor 1280 may include one or more processing units; preferably, the processor 1280 may integrate an application processor and a modem processor, where the application processor mainly processes the operating system, user interface, application programs, etc. , The modem processor mainly deals with wireless communication. It can be understood that the foregoing modem processor may not be integrated into the processor 1280.
- the mobile phone also includes a power supply 1290 (such as a battery) for supplying power to various components.
- a power supply 1290 (such as a battery) for supplying power to various components.
- the power supply can be logically connected to the processor 1280 through a power management system, so that functions such as charging, discharging, and power management can be managed through the power management system.
- the mobile phone may also include a camera, a Bluetooth module, etc., which will not be repeated here.
- the processor 1280 included in the terminal device also has the following functions:
- the acoustic features and target expression parameters corresponding to the target text information are determined; the acoustic features are used to identify and simulate the speaker uttering the target text
- the voice of the information, the target expression parameter is used to identify the degree of change of the facial expression that simulates the speaker uttering the target text information with respect to the first expression base;
- a second animation character with a second expression base is driven.
- the first expression base of the first animated image corresponding to the speaker is determined according to the facial expression, the first expression base is used to identify the expression of the first animated image; the dimension of the first expression base Is the first dimension, and the vertex topology is the first vertex topology;
- the target expression base is determined;
- the dimension of the second expression base is the second dimension, and the vertex topology is the second vertex topology, so
- the target expression base is the expression base corresponding to the first animation image with the second vertex topology, and the dimension of the target expression base is the second dimension;
- the target expression parameters are used to identify that the speaker uttered the voice The degree of change of the facial expression relative to the target expression base;
- the second animation character with the second expression base is driven.
- FIG. 13 is a structural diagram of the server 1300 provided by the embodiment of the present application.
- the server 1300 may have relatively large differences due to different configurations or performance, and may include one or one The above central processing unit (Central Processing Units, CPU for short) 1322 (for example, one or more processors) and memory 1332, one or more storage media 1330 for storing application programs 1342 or data 1344 (for example, one or more storage equipment).
- the memory 1332 and the storage medium 1330 may be short-term storage or persistent storage.
- the program stored in the storage medium 1330 may include one or more modules (not shown in the figure), and each module may include a series of command operations on the server.
- the central processing unit 1322 may be configured to communicate with the storage medium 1330, and execute a series of instruction operations in the storage medium 1330 on the server 1300.
- the server 1300 may also include one or more power supplies 1326, one or more wired or wireless network interfaces 1350, one or more input and output interfaces 1358, and/or one or more operating systems 1341, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
- operating systems 1341 such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
- the steps performed by the server in the foregoing embodiment may be based on the server structure shown in FIG. 13.
- the embodiments of the present application also provide a computer-readable storage medium, where the computer-readable storage medium is used to store program code, and the program code is used to execute the animation image driving method described in each of the foregoing embodiments.
- the embodiments of the present application also provide a computer program product including instructions, which when run on a computer, cause the computer to execute the animation image driving method described in each of the foregoing embodiments.
- At least one (item) refers to one or more, and “multiple” refers to two or more.
- “And/or” is used to describe the association relationship of associated objects, indicating that there can be three types of relationships, for example, “A and/or B” can mean: only A, only B, and both A and B , Where A and B can be singular or plural.
- the character “/” generally indicates that the associated objects before and after are in an “or” relationship.
- the following at least one item (a) or similar expressions refers to any combination of these items, including any combination of a single item (a) or a plurality of items (a).
- At least one of a, b, or c can mean: a, b, c, "a and b", “a and c", “b and c", or "a and b and c" ", where a, b, and c can be single or multiple.
- the disclosed system, device, and method may be implemented in other ways.
- the device embodiments described above are merely illustrative, for example, the division of the units is only a logical function division, and there may be other divisions in actual implementation, for example, multiple units or components may be combined or It can be integrated into another system, or some features can be ignored or not implemented.
- the displayed or discussed mutual coupling or direct coupling or communication connection may be indirect coupling or communication connection through some interfaces, devices or units, and may be in electrical, mechanical or other forms.
- the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the objectives of the solutions of the embodiments.
- the functional units in the various embodiments of the present application may be integrated into one processing unit, or each unit may exist alone physically, or two or more units may be integrated into one unit.
- the above-mentioned integrated unit can be implemented in the form of hardware or software functional unit.
- the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium.
- the technical solution of the present application essentially or the part that contributes to the existing technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium , Including several instructions to make a computer device (which may be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in the various embodiments of the present application.
- the aforementioned storage media include: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), magnetic disks or optical disks, etc., which can store program codes. Medium.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Multimedia (AREA)
- Human Computer Interaction (AREA)
- General Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- General Health & Medical Sciences (AREA)
- Acoustics & Sound (AREA)
- Computational Linguistics (AREA)
- Oral & Maxillofacial Surgery (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Data Mining & Analysis (AREA)
- Quality & Reliability (AREA)
- Signal Processing (AREA)
- Processing Or Creating Images (AREA)
Abstract
Description
Claims (16)
- 一种动画形象驱动方法,所述方法由音视频处理设备执行,所述方法包括:获取包含说话人的脸部表情和对应语音的媒体数据;根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;根据目标文本信息、所述媒体数据和所述第一表情基,确定对应所述目标文本信息的声学特征和目标表情参数;所述声学特征用于标识模拟所述说话人说出所述目标文本信息的声音,所述目标表情参数用于标识模拟所述说话人说出所述目标文本信息的脸部表情相对于所述第一表情基的变化程度;根据所述声学特征和所述目标表情参数,驱动具有第二表情基的第二动画形象。
- 根据权利要求1所述的方法,所述第一动画形象和所述第二动画形象为同一个动画形象,所述第一表情基与所述第二表情基相同,所述根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,包括:根据所述脸部表情确定所述第一动画形象的第一表情基和所述第一动画形象的捏脸参数,所述捏脸参数用于标识所述第一动画形象的脸型相对于所述第一动画形象所对应捏脸基的变化程度;所述根据所述声学特征和所述目标表情参数,驱动具有第二表情基的第二动画形象,包括:根据所述声学特征、所述目标表情参数和所述捏脸参数,驱动所述第二动画形象。
- 根据权利要求1所述的方法,所述第一动画形象和所述第二动画形象为不同动画形象,所述第一表情基与所述第二表情基不同,所述根据所述声学特征和所述目标表情参数,驱动具有第二表情基的第二动画形象,包括:确定所述第一表情基所对应表情参数与所述第二表情基所对应表情参数间的映射关系;根据所述声学特征、所述目标表情参数和所述映射关系,驱动所述第二动画形象。
- 根据权利要求3所述的方法,所述第二表情基是根据所述第二表情基与音素的预设关系生成的,所述确定所述第一表情基所对应表情参数与所述第二表情基所对应表情参数间的映射关系,包括:根据所述媒体数据确定所述语音所标识音素、所述音素对应的时间区间和所述媒体数据处于所述时间区间的视频帧;根据所述视频帧确定所述音素对应的第一表情参数,所述第一表情参数用于标识发出所述音素时所述说话人的脸部表情相对于所述第一表情基的变化程度;根据所述预设关系和所述第二表情基,确定所述音素对应的第二表情参数;根据所述第一表情参数和所述第二表情参数,确定所述映射关系。
- 根据权利要求1所述的方法,所述根据目标文本信息、所述媒体数据和所述第一表情基,确定对应所述目标文本信息的声学特征和目标表情参数,包括:根据所述目标文本信息和所述媒体数据,确定对应所述目标文本信息的声学特征和表情特征;所述表情特征用于标识模拟所述说话人说出所述目标文本信息的脸部表情;根据所述第一表情基和所述表情特征确定所述目标表情参数。
- 一种动画形象驱动装置,所述装置部署在音视频处理设备上,所述装置包括获取单元、第一确定单元、第二确定单元和驱动单元:所述获取单元,用于获取包含说话人的脸部表情和对应语音的媒体数据;所述第一确定单元,用于根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;所述第二确定单元,用于根据目标文本信息、所述媒体数据和所述第一表情基,确定对应所述目标文本信息的声学特征和目标表情参数;所述声学特征用于标识模拟所述说话人说出所述目标文本信息的声音,所述目标表情参数用于标识模拟所述说话人说出所述目标文本信息的脸部表情相对于所述第一表情基的变化程度;所述驱动单元,用于根据所述声学特征和所述目标表情参数,驱动具有第二表情基的第二动画形象。
- 根据权利要求6所述的装置,所述第一动画形象和所述第二动画形象为同一个动画形象,所述第一表情基与所述第二表情基相同,所述第一确定单元,用于:根据所述脸部表情确定所述第一动画形象的第一表情基和所述第一动画形象的捏脸参数,所述捏脸参数用于标识所述第一动画形象的脸型相对于所述第一动画形象所对应捏脸基的变化程度;所述驱动单元,用于:根据所述声学特征、所述目标表情参数和所述捏脸参数,驱动所述第二动画形象。
- 根据权利要求6所述的装置,所述第一动画形象和所述第二动画形象为不同动画形象,所述第一表情基与所述第二表情基不同,所述驱动单元,用于:确定所述第一表情基所对应表情参数与所述第二表情基所对应表情参数间的映射关系;根据所述声学特征、所述目标表情参数和所述映射关系,驱动所述第二动画形象。
- 根据权利要求8所述的装置,所述第二表情基是根据所述第二表情基与音素的预设关系生成的,所述驱动单元,还用于:根据所述媒体数据确定所述语音所标识音素、所述音素对应的时间区间和所述媒体数据处于所述时间区间的视频帧;根据所述视频帧确定所述音素对应的第一表情参数,所述第一表情参数用于标识发出所述音素时所述说话人的脸部表情相对于所述第一表情基的变化程度;根据所述预设关系和所述第二表情基,确定所述音素对应的第二表情参数;根据所述第一表情参数和所述第二表情参数,确定所述映射关系。
- 根据权利要求6所述的装置,所述第二确定单元,用于:根据所述目标文本信息和所述媒体数据,确定对应所述目标文本信息的声学特征和表情特征;所述表情特征用于标识模拟所述说话人说出所述目标文本信息的脸部表情;根据所述第一表情基和所述表情特征确定所述目标表情参数。
- 一种动画形象驱动方法,所述方法由音视频处理设备执行,所述方法包括:获取包含说话人的脸部表情和对应语音的第一媒体数据;根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;所述第一表情基的维数为第一维数,顶点拓扑为第一顶点拓扑;根据所述第一表情基和待驱动的第二动画形象的第二表情基,确定目标表情基;所述第二表情基的维数为第二维数,顶点拓扑为第二顶点拓扑,所述目标表情基为具有第二顶点拓扑的第一动画形象对应的表情基,所述目标表情基的维数为第二维数;根据包含所述说话人的脸部表情和对应语音的第二媒体数据和所述目标表情基,确定目标表情参数和声学特征;所述目标表情参数用于标识所述说话人说出所述语音的脸部表情相对于所述目标表情基的变化程度;根据所述目标表情参数和声学特征,驱动具有所述第二表情基的所述第二动画形象。
- 根据权利要求11所述的方法,所述根据所述第一表情基和待驱动的第二动画形象的第二表情基,确定目标表情基,包括:从所述第一表情基中确定所述第一动画形象处于无表情时对应的无表情网格,并从所述第二表情基中确所述第二动画形象处于无表情时对应的无表情网格;根据所述第一动画形象对应的无表情网格和所述第二动画形象对应的无表情网格,确定调整网格,所述调整网格具有第二顶点拓扑,用于标识处于无表情时的第一动画形象;根据所述调整网格和所述第二表情基中的网格形变关系,生成所述目标表情基。
- 一种动画形象驱动装置,所述装置部署在音视频处理设备上,所述装置包括获取单元、第一确定单元、第二确定单元、第三确定单元和驱动单元:所述获取单元,用于获取包含说话人的脸部表情和对应语音的第一媒体数据;所述第一确定单元,用于根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;所述第一表情基的维数为第一维数,顶点拓扑为第一顶点拓扑;所述第二确定单元,用于根据所述第一表情基和待驱动的第二动画形象的第二表情基,确定目标表情基;所述第二表情基的维数为第二维数,顶点拓扑为第二顶点拓扑,所述目标表情基为具有第二顶点拓扑的第一动画形象对应的表情基,所述目标表情基的维数为第二维数;所述第三确定单元,用于根据包含所述说话人的脸部表情和对应语音的第二媒体数据和所述目标表情基,确定目标表情参数和声学特征;所述目标表情参数用于标识所述说话人说出所述语音的脸部表情相对于所述目标表情基的变化程度;所述驱动单元,用于根据所述目标表情参数和声学特征,驱动具有所述第二表情基的所述第二动画形象。
- 一种用于动画形象驱动的设备,所述设备包括处理器以及存储器:所述存储器用于存储程序代码,并将所述程序代码传输给所述处理器;所述处理器用于根据所述程序代码中的指令执行权利要求1-5或11-12任一项所述的方法。
- 一种计算机可读存储介质,所述计算机可读存储介质用于存储程序代码,所述程序代码用于执行权利要求1-5或11-12任一项所述的方法。
- 一种计算机程序产品,当所述计算机程序产品被执行时,用于执行权利要求1-5或11-12任一项所述的方法。
Priority Applications (5)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2021557135A JP7408048B2 (ja) | 2019-09-02 | 2020-08-27 | 人工知能に基づくアニメキャラクター駆動方法及び関連装置 |
| EP20860658.2A EP3929703B1 (en) | 2019-09-02 | 2020-08-27 | Animation image driving method based on artificial intelligence, and related device |
| KR1020217029221A KR102694330B1 (ko) | 2019-09-02 | 2020-08-27 | 인공 지능에 기초한 애니메이션 이미지 구동 방법, 및 관련 디바이스 |
| US17/405,965 US11605193B2 (en) | 2019-09-02 | 2021-08-18 | Artificial intelligence-based animation character drive method and related apparatus |
| US18/080,655 US12112417B2 (en) | 2019-09-02 | 2022-12-13 | Artificial intelligence-based animation character drive method and related apparatus |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201910824770.0 | 2019-09-02 | ||
| CN201910824770.0A CN110531860B (zh) | 2019-09-02 | 2019-09-02 | 一种基于人工智能的动画形象驱动方法和装置 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/405,965 Continuation US11605193B2 (en) | 2019-09-02 | 2021-08-18 | Artificial intelligence-based animation character drive method and related apparatus |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2021043053A1 true WO2021043053A1 (zh) | 2021-03-11 |
Family
ID=68666304
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2020/111615 Ceased WO2021043053A1 (zh) | 2019-09-02 | 2020-08-27 | 一种基于人工智能的动画形象驱动方法和相关装置 |
Country Status (6)
| Country | Link |
|---|---|
| US (2) | US11605193B2 (zh) |
| EP (1) | EP3929703B1 (zh) |
| JP (1) | JP7408048B2 (zh) |
| KR (1) | KR102694330B1 (zh) |
| CN (1) | CN110531860B (zh) |
| WO (1) | WO2021043053A1 (zh) |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114420088A (zh) * | 2022-01-20 | 2022-04-29 | 安徽淘云科技股份有限公司 | 一种展示方法及其相关设备 |
| CN115617169A (zh) * | 2022-10-11 | 2023-01-17 | 深圳琪乐科技有限公司 | 一种语音控制机器人及基于角色关系的机器人控制方法 |
| US12136159B2 (en) * | 2022-12-31 | 2024-11-05 | Theai, Inc. | Contextually oriented behavior of artificial intelligence characters |
Families Citing this family (104)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9318108B2 (en) | 2010-01-18 | 2016-04-19 | Apple Inc. | Intelligent automated assistant |
| US8977255B2 (en) | 2007-04-03 | 2015-03-10 | Apple Inc. | Method and system for operating a multi-function portable electronic device using voice-activation |
| US8676904B2 (en) | 2008-10-02 | 2014-03-18 | Apple Inc. | Electronic devices with voice command and contextual data processing capabilities |
| US10706373B2 (en) | 2011-06-03 | 2020-07-07 | Apple Inc. | Performing actions associated with task items that represent tasks to perform |
| US10276170B2 (en) | 2010-01-18 | 2019-04-30 | Apple Inc. | Intelligent automated assistant |
| US10057736B2 (en) | 2011-06-03 | 2018-08-21 | Apple Inc. | Active transport based notifications |
| US10417037B2 (en) | 2012-05-15 | 2019-09-17 | Apple Inc. | Systems and methods for integrating third party services with a digital assistant |
| EP4560630A3 (en) | 2013-02-07 | 2025-08-06 | Apple Inc. | Voice trigger for a digital assistant |
| US10652394B2 (en) | 2013-03-14 | 2020-05-12 | Apple Inc. | System and method for processing voicemail |
| US10748529B1 (en) | 2013-03-15 | 2020-08-18 | Apple Inc. | Voice activated device for use with a voice-based digital assistant |
| EP3008641A1 (en) | 2013-06-09 | 2016-04-20 | Apple Inc. | Device, method, and graphical user interface for enabling conversation persistence across two or more instances of a digital assistant |
| US10176167B2 (en) | 2013-06-09 | 2019-01-08 | Apple Inc. | System and method for inferring user intent from speech inputs |
| US10791216B2 (en) | 2013-08-06 | 2020-09-29 | Apple Inc. | Auto-activating smart responses based on activities from remote devices |
| US10170123B2 (en) | 2014-05-30 | 2019-01-01 | Apple Inc. | Intelligent assistant for home automation |
| US9966065B2 (en) | 2014-05-30 | 2018-05-08 | Apple Inc. | Multi-command single utterance input method |
| US9715875B2 (en) | 2014-05-30 | 2017-07-25 | Apple Inc. | Reducing the need for manual start/end-pointing and trigger phrases |
| US9338493B2 (en) | 2014-06-30 | 2016-05-10 | Apple Inc. | Intelligent automated assistant for TV user interactions |
| US9886953B2 (en) | 2015-03-08 | 2018-02-06 | Apple Inc. | Virtual assistant activation |
| US10460227B2 (en) | 2015-05-15 | 2019-10-29 | Apple Inc. | Virtual assistant in a communication session |
| US10200824B2 (en) | 2015-05-27 | 2019-02-05 | Apple Inc. | Systems and methods for proactively identifying and surfacing relevant content on a touch-sensitive device |
| US20160378747A1 (en) | 2015-06-29 | 2016-12-29 | Apple Inc. | Virtual assistant for media playback |
| US10671428B2 (en) | 2015-09-08 | 2020-06-02 | Apple Inc. | Distributed personal assistant |
| US10331312B2 (en) | 2015-09-08 | 2019-06-25 | Apple Inc. | Intelligent automated assistant in a media environment |
| US10747498B2 (en) | 2015-09-08 | 2020-08-18 | Apple Inc. | Zero latency digital assistant |
| US10740384B2 (en) | 2015-09-08 | 2020-08-11 | Apple Inc. | Intelligent automated assistant for media search and playback |
| US11587559B2 (en) | 2015-09-30 | 2023-02-21 | Apple Inc. | Intelligent device identification |
| US10691473B2 (en) | 2015-11-06 | 2020-06-23 | Apple Inc. | Intelligent automated assistant in a messaging environment |
| US10956666B2 (en) | 2015-11-09 | 2021-03-23 | Apple Inc. | Unconventional virtual assistant interactions |
| US10223066B2 (en) | 2015-12-23 | 2019-03-05 | Apple Inc. | Proactive assistance based on dialog communication between devices |
| US12223282B2 (en) | 2016-06-09 | 2025-02-11 | Apple Inc. | Intelligent automated assistant in a home environment |
| US10586535B2 (en) | 2016-06-10 | 2020-03-10 | Apple Inc. | Intelligent digital assistant in a multi-tasking environment |
| DK201670540A1 (en) | 2016-06-11 | 2018-01-08 | Apple Inc | Application integration with a digital assistant |
| DK179415B1 (en) | 2016-06-11 | 2018-06-14 | Apple Inc | Intelligent device arbitration and control |
| US12197817B2 (en) | 2016-06-11 | 2025-01-14 | Apple Inc. | Intelligent device arbitration and control |
| US11204787B2 (en) | 2017-01-09 | 2021-12-21 | Apple Inc. | Application integration with a digital assistant |
| US10726832B2 (en) | 2017-05-11 | 2020-07-28 | Apple Inc. | Maintaining privacy of personal information |
| DK180048B1 (en) | 2017-05-11 | 2020-02-04 | Apple Inc. | MAINTAINING THE DATA PROTECTION OF PERSONAL INFORMATION |
| DK179496B1 (en) | 2017-05-12 | 2019-01-15 | Apple Inc. | USER-SPECIFIC Acoustic Models |
| DK201770428A1 (en) | 2017-05-12 | 2019-02-18 | Apple Inc. | LOW-LATENCY INTELLIGENT AUTOMATED ASSISTANT |
| DK179745B1 (en) | 2017-05-12 | 2019-05-01 | Apple Inc. | SYNCHRONIZATION AND TASK DELEGATION OF A DIGITAL ASSISTANT |
| DK201770411A1 (en) | 2017-05-15 | 2018-12-20 | Apple Inc. | Multi-modal interfaces |
| DK179549B1 (en) | 2017-05-16 | 2019-02-12 | Apple Inc. | FAR-FIELD EXTENSION FOR DIGITAL ASSISTANT SERVICES |
| US10303715B2 (en) | 2017-05-16 | 2019-05-28 | Apple Inc. | Intelligent automated assistant for media exploration |
| US20180336892A1 (en) | 2017-05-16 | 2018-11-22 | Apple Inc. | Detecting a trigger of a digital assistant |
| US10818288B2 (en) | 2018-03-26 | 2020-10-27 | Apple Inc. | Natural assistant interaction |
| US10928918B2 (en) | 2018-05-07 | 2021-02-23 | Apple Inc. | Raise to speak |
| US11145294B2 (en) | 2018-05-07 | 2021-10-12 | Apple Inc. | Intelligent automated assistant for delivering content from user experiences |
| EP3815050B1 (en) * | 2018-05-24 | 2024-01-24 | Warner Bros. Entertainment Inc. | Matching mouth shape and movement in digital video to alternative audio |
| DK179822B1 (da) | 2018-06-01 | 2019-07-12 | Apple Inc. | Voice interaction at a primary device to access call functionality of a companion device |
| US10892996B2 (en) | 2018-06-01 | 2021-01-12 | Apple Inc. | Variable latency device coordination |
| DK180639B1 (en) | 2018-06-01 | 2021-11-04 | Apple Inc | DISABILITY OF ATTENTION-ATTENTIVE VIRTUAL ASSISTANT |
| DK201870355A1 (en) | 2018-06-01 | 2019-12-16 | Apple Inc. | VIRTUAL ASSISTANT OPERATION IN MULTI-DEVICE ENVIRONMENTS |
| US11462215B2 (en) | 2018-09-28 | 2022-10-04 | Apple Inc. | Multi-modal inputs for voice commands |
| CN111627095B (zh) * | 2019-02-28 | 2023-10-24 | 北京小米移动软件有限公司 | 表情生成方法及装置 |
| US11348573B2 (en) | 2019-03-18 | 2022-05-31 | Apple Inc. | Multimodality in digital assistant systems |
| US11307752B2 (en) | 2019-05-06 | 2022-04-19 | Apple Inc. | User configurable task triggers |
| DK201970509A1 (en) | 2019-05-06 | 2021-01-15 | Apple Inc | Spoken notifications |
| US11140099B2 (en) | 2019-05-21 | 2021-10-05 | Apple Inc. | Providing message response suggestions |
| DK201970510A1 (en) | 2019-05-31 | 2021-02-11 | Apple Inc | Voice identification in digital assistant systems |
| DK180129B1 (en) | 2019-05-31 | 2020-06-02 | Apple Inc. | User activity shortcut suggestions |
| US11227599B2 (en) | 2019-06-01 | 2022-01-18 | Apple Inc. | Methods and user interfaces for voice-based control of electronic devices |
| CN110531860B (zh) | 2019-09-02 | 2020-07-24 | 腾讯科技(深圳)有限公司 | 一种基于人工智能的动画形象驱动方法和装置 |
| CN111145777A (zh) * | 2019-12-31 | 2020-05-12 | 苏州思必驰信息科技有限公司 | 一种虚拟形象展示方法、装置、电子设备及存储介质 |
| US11593984B2 (en) | 2020-02-07 | 2023-02-28 | Apple Inc. | Using text for avatar animation |
| CN111294665B (zh) * | 2020-02-12 | 2021-07-20 | 百度在线网络技术(北京)有限公司 | 视频的生成方法、装置、电子设备及可读存储介质 |
| CN111311712B (zh) * | 2020-02-24 | 2023-06-16 | 北京百度网讯科技有限公司 | 视频帧处理方法和装置 |
| CN111372113B (zh) * | 2020-03-05 | 2021-12-21 | 成都威爱新经济技术研究院有限公司 | 基于数字人表情、嘴型及声音同步的用户跨平台交流方法 |
| CN111736700B (zh) * | 2020-06-23 | 2025-01-07 | 上海商汤临港智能科技有限公司 | 基于数字人的车舱交互方法、装置及车辆 |
| KR20220004156A (ko) * | 2020-03-30 | 2022-01-11 | 상하이 센스타임 린강 인텔리전트 테크놀로지 컴퍼니 리미티드 | 디지털 휴먼에 기반한 자동차 캐빈 인터랙션 방법, 장치 및 차량 |
| CN111459450A (zh) * | 2020-03-31 | 2020-07-28 | 北京市商汤科技开发有限公司 | 交互对象的驱动方法、装置、设备以及存储介质 |
| US12301635B2 (en) | 2020-05-11 | 2025-05-13 | Apple Inc. | Digital assistant hardware abstraction |
| US11043220B1 (en) | 2020-05-11 | 2021-06-22 | Apple Inc. | Digital assistant hardware abstraction |
| US11061543B1 (en) | 2020-05-11 | 2021-07-13 | Apple Inc. | Providing relevant data items based on context |
| US11755276B2 (en) | 2020-05-12 | 2023-09-12 | Apple Inc. | Reducing description length based on confidence |
| US11490204B2 (en) | 2020-07-20 | 2022-11-01 | Apple Inc. | Multi-device audio adjustment coordination |
| US11438683B2 (en) | 2020-07-21 | 2022-09-06 | Apple Inc. | User identification using headphones |
| CN111988658B (zh) * | 2020-08-28 | 2022-12-06 | 网易(杭州)网络有限公司 | 视频生成方法及装置 |
| CN118981266A (zh) * | 2020-10-14 | 2024-11-19 | 住友电气工业株式会社 | 计算机可读取的存储介质 |
| US12128322B2 (en) * | 2020-11-12 | 2024-10-29 | Tencent Technology (Shenzhen) Company Limited | Method and apparatus for driving vehicle in virtual environment, terminal, and storage medium |
| CN112527115B (zh) * | 2020-12-15 | 2023-08-04 | 北京百度网讯科技有限公司 | 用户形象生成方法、相关装置及计算机程序产品 |
| CN112669424B (zh) * | 2020-12-24 | 2024-05-31 | 科大讯飞股份有限公司 | 一种表情动画生成方法、装置、设备及存储介质 |
| CN112286366B (zh) * | 2020-12-30 | 2022-02-22 | 北京百度网讯科技有限公司 | 用于人机交互的方法、装置、设备和介质 |
| CN112927712B (zh) * | 2021-01-25 | 2024-06-04 | 网易(杭州)网络有限公司 | 视频生成方法、装置和电子设备 |
| CN115205917A (zh) * | 2021-04-12 | 2022-10-18 | 上海擎感智能科技有限公司 | 一种人机交互的方法及电子设备 |
| CN113066156B (zh) * | 2021-04-16 | 2025-06-03 | 广州虎牙科技有限公司 | 表情重定向方法、装置、设备和介质 |
| CN113256821B (zh) * | 2021-06-02 | 2022-02-01 | 北京世纪好未来教育科技有限公司 | 一种三维虚拟形象唇形生成方法、装置及电子设备 |
| CN115984452A (zh) * | 2021-10-14 | 2023-04-18 | 聚好看科技股份有限公司 | 一种头部三维重建方法及设备 |
| KR20230100205A (ko) * | 2021-12-28 | 2023-07-05 | 삼성전자주식회사 | 영상 처리 방법 및 장치 |
| CN114898019A (zh) * | 2022-02-08 | 2022-08-12 | 武汉路特斯汽车有限公司 | 一种动画融合方法和装置 |
| CN114612600B (zh) * | 2022-03-11 | 2023-02-17 | 北京百度网讯科技有限公司 | 虚拟形象生成方法、装置、电子设备和存储介质 |
| CN116778107B (zh) | 2022-03-11 | 2024-10-15 | 腾讯科技(深圳)有限公司 | 表情模型的生成方法、装置、设备及介质 |
| CN114758040A (zh) * | 2022-03-24 | 2022-07-15 | 努比亚技术有限公司 | 一种虚拟人物生成方法、设备及计算机可读存储介质 |
| CN114708636A (zh) * | 2022-04-01 | 2022-07-05 | 成都市谛视科技有限公司 | 一种密集人脸网格表情驱动方法、装置及介质 |
| CN115050067B (zh) * | 2022-05-25 | 2024-07-02 | 中国科学院半导体研究所 | 人脸表情构建方法、装置、电子设备、存储介质及产品 |
| CN115311394A (zh) * | 2022-07-27 | 2022-11-08 | 湖南芒果无际科技有限公司 | 一种驱动数字人面部动画方法、系统、设备及介质 |
| KR102838168B1 (ko) * | 2022-09-02 | 2025-07-25 | 동서대학교 산학협력단 | 딥페이크 기술을 활용한 동화 미디어 셀프 제작방법 |
| KR102652652B1 (ko) * | 2022-11-29 | 2024-03-29 | 주식회사 일루니 | 아바타 생성 장치 및 방법 |
| US20240265605A1 (en) * | 2023-02-07 | 2024-08-08 | Google Llc | Generating an avatar expression |
| CN116188649B (zh) * | 2023-04-27 | 2023-10-13 | 科大讯飞股份有限公司 | 基于语音的三维人脸模型驱动方法及相关装置 |
| CN116452709A (zh) * | 2023-06-13 | 2023-07-18 | 北京好心情互联网医院有限公司 | 动画生成方法、装置、设备及存储介质 |
| CN116778043B (zh) * | 2023-06-19 | 2024-02-09 | 广州怪力视效网络科技有限公司 | 一种表情捕捉及动画自动生成系统和方法 |
| US12045639B1 (en) * | 2023-08-23 | 2024-07-23 | Bithuman Inc | System providing visual assistants with artificial intelligence |
| CN120935387A (zh) * | 2024-05-08 | 2025-11-11 | 北京字跳网络技术有限公司 | 生成媒体内容的方法、装置、设备和存储介质 |
| CN118331431B (zh) * | 2024-06-13 | 2024-08-27 | 海马云(天津)信息技术有限公司 | 虚拟数字人驱动方法与装置、电子设备及存储介质 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US8725507B2 (en) * | 2009-11-27 | 2014-05-13 | Samsung Eletronica Da Amazonia Ltda. | Systems and methods for synthesis of motion for animation of virtual heads/characters via voice processing in portable devices |
| CN104217454A (zh) * | 2014-08-21 | 2014-12-17 | 中国科学院计算技术研究所 | 一种视频驱动的人脸动画生成方法 |
| CN108875633A (zh) * | 2018-06-19 | 2018-11-23 | 北京旷视科技有限公司 | 表情检测与表情驱动方法、装置和系统及存储介质 |
| CN109447234A (zh) * | 2018-11-14 | 2019-03-08 | 腾讯科技(深圳)有限公司 | 一种模型训练方法、合成说话表情的方法和相关装置 |
| CN110531860A (zh) * | 2019-09-02 | 2019-12-03 | 腾讯科技(深圳)有限公司 | 一种基于人工智能的动画形象驱动方法和装置 |
Family Cites Families (21)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2003141564A (ja) | 2001-10-31 | 2003-05-16 | Minolta Co Ltd | アニメーション生成装置およびアニメーション生成方法 |
| US8555164B2 (en) * | 2001-11-27 | 2013-10-08 | Ding Huang | Method for customizing avatars and heightening online safety |
| US8224652B2 (en) * | 2008-09-26 | 2012-07-17 | Microsoft Corporation | Speech and text driven HMM-based body animation synthesis |
| US8803889B2 (en) | 2009-05-29 | 2014-08-12 | Microsoft Corporation | Systems and methods for applying animations or motions to a character |
| US20130257877A1 (en) * | 2012-03-30 | 2013-10-03 | Videx, Inc. | Systems and Methods for Generating an Interactive Avatar Model |
| WO2015145219A1 (en) * | 2014-03-28 | 2015-10-01 | Navaratnam Ratnakumar | Systems for remote service of customers using virtual and physical mannequins |
| JP2015210739A (ja) | 2014-04-28 | 2015-11-24 | 株式会社コロプラ | キャラクタ画像生成方法及びキャラクタ画像生成プログラム |
| CN107004287B (zh) * | 2014-11-05 | 2020-10-23 | 英特尔公司 | 化身视频装置和方法 |
| US9911218B2 (en) * | 2015-12-01 | 2018-03-06 | Disney Enterprises, Inc. | Systems and methods for speech animation using visemes with phonetic boundary context |
| CN105551071B (zh) * | 2015-12-02 | 2018-08-10 | 中国科学院计算技术研究所 | 一种文本语音驱动的人脸动画生成方法及系统 |
| US10528801B2 (en) * | 2016-12-07 | 2020-01-07 | Keyterra LLC | Method and system for incorporating contextual and emotional visualization into electronic communications |
| WO2018175892A1 (en) * | 2017-03-23 | 2018-09-27 | D&M Holdings, Inc. | System providing expressive and emotive text-to-speech |
| US10586368B2 (en) * | 2017-10-26 | 2020-03-10 | Snap Inc. | Joint audio-video facial animation system |
| US10657695B2 (en) * | 2017-10-30 | 2020-05-19 | Snap Inc. | Animated chat presence |
| KR20190078015A (ko) * | 2017-12-26 | 2019-07-04 | 주식회사 글로브포인트 | 3d 아바타를 이용한 게시판 관리 서버 및 방법 |
| CN108763190B (zh) * | 2018-04-12 | 2019-04-02 | 平安科技(深圳)有限公司 | 基于语音的口型动画合成装置、方法及可读存储介质 |
| CN109377540B (zh) * | 2018-09-30 | 2023-12-19 | 网易(杭州)网络有限公司 | 面部动画的合成方法、装置、存储介质、处理器及终端 |
| CN109961496B (zh) * | 2019-02-22 | 2022-10-28 | 厦门美图宜肤科技有限公司 | 表情驱动方法及表情驱动装置 |
| US11202131B2 (en) * | 2019-03-10 | 2021-12-14 | Vidubly Ltd | Maintaining original volume changes of a character in revoiced media stream |
| CN110288682B (zh) * | 2019-06-28 | 2023-09-26 | 北京百度网讯科技有限公司 | 用于控制三维虚拟人像口型变化的方法和装置 |
| US10949715B1 (en) * | 2019-08-19 | 2021-03-16 | Neon Evolution Inc. | Methods and systems for image and voice processing |
-
2019
- 2019-09-02 CN CN201910824770.0A patent/CN110531860B/zh active Active
-
2020
- 2020-08-27 WO PCT/CN2020/111615 patent/WO2021043053A1/zh not_active Ceased
- 2020-08-27 JP JP2021557135A patent/JP7408048B2/ja active Active
- 2020-08-27 KR KR1020217029221A patent/KR102694330B1/ko active Active
- 2020-08-27 EP EP20860658.2A patent/EP3929703B1/en active Active
-
2021
- 2021-08-18 US US17/405,965 patent/US11605193B2/en active Active
-
2022
- 2022-12-13 US US18/080,655 patent/US12112417B2/en active Active
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US8725507B2 (en) * | 2009-11-27 | 2014-05-13 | Samsung Eletronica Da Amazonia Ltda. | Systems and methods for synthesis of motion for animation of virtual heads/characters via voice processing in portable devices |
| CN104217454A (zh) * | 2014-08-21 | 2014-12-17 | 中国科学院计算技术研究所 | 一种视频驱动的人脸动画生成方法 |
| CN108875633A (zh) * | 2018-06-19 | 2018-11-23 | 北京旷视科技有限公司 | 表情检测与表情驱动方法、装置和系统及存储介质 |
| CN109447234A (zh) * | 2018-11-14 | 2019-03-08 | 腾讯科技(深圳)有限公司 | 一种模型训练方法、合成说话表情的方法和相关装置 |
| CN110531860A (zh) * | 2019-09-02 | 2019-12-03 | 腾讯科技(深圳)有限公司 | 一种基于人工智能的动画形象驱动方法和装置 |
Non-Patent Citations (2)
| Title |
|---|
| CAO, CHEN ET AL.: "FaceWarehouse: A 3D Facial Expression Database for Visual Computing", IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, vol. 20, no. 3, 31 March 2014 (2014-03-31), XP011543570, DOI: 20201116102735A * |
| See also references of EP3929703A4 * |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114420088A (zh) * | 2022-01-20 | 2022-04-29 | 安徽淘云科技股份有限公司 | 一种展示方法及其相关设备 |
| CN115617169A (zh) * | 2022-10-11 | 2023-01-17 | 深圳琪乐科技有限公司 | 一种语音控制机器人及基于角色关系的机器人控制方法 |
| US12136159B2 (en) * | 2022-12-31 | 2024-11-05 | Theai, Inc. | Contextually oriented behavior of artificial intelligence characters |
Also Published As
| Publication number | Publication date |
|---|---|
| US11605193B2 (en) | 2023-03-14 |
| CN110531860B (zh) | 2020-07-24 |
| EP3929703B1 (en) | 2025-04-30 |
| US12112417B2 (en) | 2024-10-08 |
| KR20210123399A (ko) | 2021-10-13 |
| KR102694330B1 (ko) | 2024-08-13 |
| EP3929703A4 (en) | 2022-10-05 |
| US20230123433A1 (en) | 2023-04-20 |
| JP7408048B2 (ja) | 2024-01-05 |
| EP3929703A1 (en) | 2021-12-29 |
| JP2022527155A (ja) | 2022-05-31 |
| US20210383586A1 (en) | 2021-12-09 |
| CN110531860A (zh) | 2019-12-03 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN110531860B (zh) | 一种基于人工智能的动画形象驱动方法和装置 | |
| CN112379812B (zh) | 仿真3d数字人交互方法、装置、电子设备及存储介质 | |
| CN115909015B (zh) | 一种可形变神经辐射场网络的构建方法和装置 | |
| US11858118B2 (en) | Robot, server, and human-machine interaction method | |
| CN110288077B (zh) | 一种基于人工智能的合成说话表情的方法和相关装置 | |
| WO2021036644A1 (zh) | 一种基于人工智能的语音驱动动画方法和装置 | |
| CN110517340B (zh) | 一种基于人工智能的脸部模型确定方法和装置 | |
| CN110286756A (zh) | 视频处理方法、装置、系统、终端设备及存储介质 | |
| WO2023246163A1 (zh) | 一种虚拟数字人驱动方法、装置、设备和介质 | |
| CN113421547A (zh) | 一种语音处理方法及相关设备 | |
| WO2021098338A1 (zh) | 一种模型训练的方法、媒体信息合成的方法及相关装置 | |
| CN110517339A (zh) | 一种基于人工智能的动画形象驱动方法和装置 | |
| CN117370605A (zh) | 一种虚拟数字人驱动方法、装置、设备和介质 | |
| KR20190126906A (ko) | 돌봄 로봇을 위한 데이터 처리 방법 및 장치 | |
| CN116229311B (zh) | 视频处理方法、装置及存储介质 | |
| CN112767520A (zh) | 数字人生成方法、装置、电子设备及存储介质 | |
| CN109343695A (zh) | 基于虚拟人行为标准的交互方法及系统 | |
| WO2025209111A1 (zh) | 视频生成模型的训练方法、装置、设备、存储介质和产品 | |
| WO2025232361A1 (zh) | 一种视频生成方法和相关装置 | |
| CN109558853B (zh) | 一种音频合成方法及终端设备 | |
| CN118250523A (zh) | 数字人视频生成方法、装置、存储介质及电子设备 | |
| CN116248811A (zh) | 视频处理方法、装置及存储介质 | |
| CN115526772A (zh) | 视频处理方法、装置、设备和存储介质 | |
| KR20260013532A (ko) | 생성형 인공지능을 이용한 3차원 캐릭터의 발화 얼굴 영상 생성 방법 | |
| CN116074577A (zh) | 视频处理方法、相关装置及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20860658 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 20217029221 Country of ref document: KR Kind code of ref document: A |
|
| ENP | Entry into the national phase |
Ref document number: 2021557135 Country of ref document: JP Kind code of ref document: A Ref document number: 2020860658 Country of ref document: EP Effective date: 20210922 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| WWG | Wipo information: grant in national office |
Ref document number: 2020860658 Country of ref document: EP |