WO2021043053A1 - 一种基于人工智能的动画形象驱动方法和相关装置 - Google Patents

一种基于人工智能的动画形象驱动方法和相关装置 Download PDF

Info

Publication number
WO2021043053A1
WO2021043053A1 PCT/CN2020/111615 CN2020111615W WO2021043053A1 WO 2021043053 A1 WO2021043053 A1 WO 2021043053A1 CN 2020111615 W CN2020111615 W CN 2020111615W WO 2021043053 A1 WO2021043053 A1 WO 2021043053A1
Authority
WO
WIPO (PCT)
Prior art keywords
expression
base
target
expression base
speaker
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2020/111615
Other languages
English (en)
French (fr)
Inventor
暴林超
康世胤
王盛
林祥凯
季兴
朱展图
李广之
陀得意
刘朋
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Tencent Technology (Shenzhen) Co Ltd
Original Assignee
Tencent Technology (Shenzhen) Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Tencent Technology (Shenzhen) Co Ltd filed Critical Tencent Technology (Shenzhen) Co Ltd
Priority to JP2021557135A priority Critical patent/JP7408048B2/ja
Priority to EP20860658.2A priority patent/EP3929703B1/en
Priority to KR1020217029221A priority patent/KR102694330B1/ko
Publication of WO2021043053A1 publication Critical patent/WO2021043053A1/zh
Priority to US17/405,965 priority patent/US11605193B2/en
Anticipated expiration legal-status Critical
Priority to US18/080,655 priority patent/US12112417B2/en
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T13/00Animation
    • G06T13/20Three-dimensional [3D] animation
    • G06T13/40Three-dimensional [3D] animation of characters, e.g. humans, animals or virtual beings
    • AHUMAN NECESSITIES
    • A63SPORTS; GAMES; AMUSEMENTS
    • A63FCARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
    • A63F13/00Video games, i.e. games using an electronically generated display having two or more dimensions
    • A63F13/50Controlling the output signals based on the game progress
    • A63F13/54Controlling the output signals based on the game progress involving acoustic signals, e.g. for simulating revolutions per minute [RPM] dependent engine sounds in a driving game or reverberation against a virtual wall
    • AHUMAN NECESSITIES
    • A63SPORTS; GAMES; AMUSEMENTS
    • A63FCARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
    • A63F13/00Video games, i.e. games using an electronically generated display having two or more dimensions
    • A63F13/55Controlling game characters or game objects based on the game progress
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/011Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/017Gesture based interaction, e.g. based on a set of recognized hand gestures
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/16Sound input; Sound output
    • G06F3/167Audio in a user interface, e.g. using voice commands for navigating, audio feedback
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T13/00Animation
    • G06T13/20Three-dimensional [3D] animation
    • G06T13/205Three-dimensional [3D] animation driven by audio data
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/40Scenes; Scene-specific elements in video content
    • G06V20/46Extracting features or characteristics from the video content, e.g. video fingerprints, representative shots or key frames
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16Human faces, e.g. facial parts, sketches or expressions
    • G06V40/174Facial expression recognition
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/02Feature extraction for speech recognition; Selection of recognition unit
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/06Transformation of speech into a non-audible representation, e.g. speech visualisation or speech processing for tactile aids
    • G10L21/10Transforming into visible information
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L13/00Speech synthesis; Text to speech systems
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/02Feature extraction for speech recognition; Selection of recognition unit
    • G10L2015/025Phonemes, fenemes or fenones being the recognition units
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/06Transformation of speech into a non-audible representation, e.g. speech visualisation or speech processing for tactile aids
    • G10L21/10Transforming into visible information
    • G10L2021/105Synthesis of the lips movements from speech, e.g. for talking heads

Definitions

  • This application relates to the field of data processing, especially to the driving of animated images.
  • the interactive device can determine the reply content according to the text or voice input by the user, and play a virtual sound synthesized according to the reply content.
  • the present application provides an artificial intelligence-based animation image driving method and device, which brings a realistic sense of substitution and immersion to the user, and improves the user's experience of interacting with the animation image.
  • an embodiment of the present application provides a method for driving an animated character, the method is executed by an audio and video processing device, and the method includes:
  • the acoustic features and target expression parameters corresponding to the target text information are determined; the acoustic features are used to identify and simulate the speaker uttering the target text
  • the voice of the information, the target expression parameter is used to identify the degree of change of the facial expression that simulates the speaker uttering the target text information with respect to the first expression base;
  • a second animation character with a second expression base is driven.
  • an embodiment of the present application provides an animation image driving device, the device is deployed on an audio and video processing device, and the device includes an acquiring unit, a first determining unit, a second determining unit, and a driving unit:
  • the acquiring unit is configured to acquire media data containing the facial expression of the speaker and the corresponding voice
  • the first determining unit is configured to determine a first expression base of the first animated image corresponding to the speaker according to the facial expression, and the first expression base is used to identify the expression of the first animated image;
  • the second determining unit is configured to determine the acoustic characteristics and target expression parameters corresponding to the target text information according to the target text information, the media data, and the first expression base; the acoustic characteristics are used to identify the simulation location The voice of the speaker uttering the target text information, and the target expression parameter is used to identify the degree of change of the facial expression that simulates the speaker uttering the target text information with respect to the first expression base;
  • the driving unit is configured to drive a second animation character with a second expression base according to the acoustic characteristics and the target expression parameters.
  • an embodiment of the present application provides a method for driving an animated character, the method is executed by an audio and video processing device, and the method includes:
  • the first expression base of the first animated image corresponding to the speaker is determined according to the facial expression, the first expression base is used to identify the expression of the first animated image; the dimension of the first expression base Is the first dimension, and the vertex topology is the first vertex topology;
  • the target expression base is determined;
  • the dimension of the second expression base is the second dimension, and the vertex topology is the second vertex topology, so
  • the target expression base is the expression base corresponding to the first animation image with the second vertex topology, and the dimension of the target expression base is the second dimension;
  • the target expression parameters are used to identify that the speaker uttered the voice The degree of change of the facial expression relative to the target expression base;
  • the second animation character with the second expression base is driven.
  • an embodiment of the present application provides an animation image driving device, the device is deployed on an audio and video processing device, and the device includes an acquisition unit, a first determination unit, a second determination unit, a third determination unit, and a driver unit:
  • the acquiring unit is configured to acquire first media data including the facial expression of the speaker and the corresponding voice;
  • the first determining unit is configured to determine a first expression base of the first animated image corresponding to the speaker according to the facial expression, and the first expression base is used to identify the expression of the first animated image;
  • the dimension of the first expression base is the first dimension, and the vertex topology is the first vertex topology;
  • the second determining unit is configured to determine a target expression base according to the first expression base and the second expression base of the second animation image to be driven;
  • the dimension of the second expression base is the second dimension
  • the vertex topology is the second vertex topology
  • the target expression base is the expression base corresponding to the first animation image with the second vertex topology
  • the dimension of the target expression base is the second dimension
  • the third determining unit is configured to determine target expression parameters and acoustic features according to the second media data containing the facial expression and corresponding voice of the speaker and the target expression base; the target expression parameters are used to identify The degree of change of the facial expression of the speaker speaking the voice relative to the target expression base;
  • the driving unit is configured to drive the second animation character with the second expression base according to the target expression parameters and acoustic characteristics.
  • an embodiment of the present application provides a device for driving animated characters, the device including a processor and a memory:
  • the memory is used to store program code and transmit the program code to the processor
  • the processor is configured to execute the method described in the first aspect or the third aspect according to instructions in the program code.
  • an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium is used to store program code, and the program code is used to execute the method described in the first aspect or the third aspect.
  • embodiments of the present application provide a computer program product, which is used to execute the method described in the first aspect or the third aspect when the computer program product is executed.
  • the first expression base of the first animation image corresponding to the speaker can be determined, and the first expression base can reflect the first animation image.
  • Different expressions After determining the target text information used to drive the second animated image, the acoustic features and target expression parameters corresponding to the target text information can be determined according to the target text information, the aforementioned collected media data, and the first expression base.
  • the acoustic characteristics can be Identifies the voice that simulates the speaker uttering the target text information, and the target expression parameter may identify the degree of change of the facial expression simulating the speaker uttering the target text information with respect to the first expression base.
  • the second animation image with the second expression base can be driven, so that the second animation image can simulate the voice of the speaker uttering the target text information through the acoustic characteristics, and make conformance during the vocalization process.
  • the speaker should have an expressive facial expression, which brings a realistic sense of substitution and immersion to the user, and improves the user's experience of interacting with the animated image.
  • FIG. 1 is a schematic diagram of an application scenario of an artificial intelligence-based animation image driving method provided by an embodiment of the application
  • FIG. 2 is a flowchart of an artificial intelligence-based animation image driving method provided by an embodiment of the application
  • FIG. 3 is a structural flow of an animation image driving system provided by an embodiment of the application.
  • FIG. 4 is an example diagram of a scene for collecting media data provided by an embodiment of the application
  • FIG. 5 is an example diagram of each dimension distribution and meaning of the 3DMM library M provided by an embodiment of the application.
  • FIG. 6 is a schematic diagram of an application scenario of an animation image driving method based on determining face pinching parameters provided by an embodiment of the application;
  • FIG. 7 is a schematic diagram of an application scenario of an animation image driving method based on determining a mapping relationship provided by an embodiment of the application;
  • FIG. 8 is an exemplary diagram of the correspondence between time intervals and phonemes provided by an embodiment of the application.
  • FIG. 9 is a flowchart of an artificial intelligence-based animation image driving method provided by an embodiment of the application.
  • Fig. 10a is a flowchart of an artificial intelligence-based animation image driving method provided by an embodiment of the application.
  • Fig. 10b is a structural diagram of an animated character driving device provided by an embodiment of the application.
  • FIG. 11 is a structural diagram of an animated character driving device provided by an embodiment of the application.
  • FIG. 12 is a structural diagram of a device for driving animated characters provided by an embodiment of the application.
  • FIG. 13 is a structural diagram of a server provided by an embodiment of this application.
  • the main research direction of human-computer interaction is to use animated images with the ability to change expressions as the interactive objects of interaction with users.
  • a game character (animated image) with the same face shape as the user's own face can be constructed.
  • the game character can make a voice and make a corresponding expression (such as mouth shape, etc.);
  • a game character with the same face shape as the user's own.
  • the opponent inputs text or voice
  • the game character can respond to the voice according to the opponent's input and make a corresponding expression.
  • an embodiment of the present application provides an artificial intelligence-based method for driving an animated image.
  • This method can determine the first expression base of the first animated image corresponding to the speaker by collecting the media data of the facial expression changes when the speaker speaks the voice.
  • the acoustic characteristics and target expression parameters corresponding to the target text information can be determined according to the target text information, the aforementioned collected media data and the first expression base, so as to drive the second animation image with the second expression base through the acoustic characteristics and the target expression parameters.
  • the second animation image is made to simulate the voice of the speaker uttering the target text information through the acoustic feature, and a facial expression conforming to the expected expression of the speaker is made during the vocalization process, so that the second animation image is driven based on the text information.
  • AI Artificial Intelligence
  • theory, methods, technologies, and application systems that perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
  • artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a similar way to human intelligence.
  • Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making.
  • Artificial intelligence technology is a comprehensive discipline, covering a wide range of fields, including both hardware-level technology and software-level technology.
  • Basic artificial intelligence technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation/interaction systems, and mechatronics.
  • Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning/deep learning.
  • the artificial intelligence technologies mainly involved include speech processing technology, machine learning, and computer vision (image).
  • Speech recognition technology includes speech signal preprocessing (Speech signal preprocessing), speech signal frequency domain analysis (Speech signal frequency analyzing), speech signal feature extraction (Speech signal feature extraction), speech signal feature matching/recognition (Speech signal feature matching/ recognition), speech training, etc.
  • Speech synthesis includes text analysis and speech generation.
  • Machine learning is a multi-field interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other subjects. Specializing in the study of how computers simulate or realize human learning behaviors in order to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve its own performance.
  • Machine learning is the core of artificial intelligence, the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence.
  • Machine learning usually includes deep learning (Deep Learning) and other technologies. Deep learning includes artificial neural networks, such as Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), and deep learning. Deep neural network (DNN), etc.
  • CNN Convolutional Neural Network
  • RNN Recurrent Neural Network
  • DNN Deep neural network
  • Video processing includes target recognition, target detection and localization, etc.
  • face recognition includes face 3D reconstruction (Face 3D Reconstruction), face detection (Face Detection), and face tracking (Face Tracking) etc.
  • the artificial intelligence-based animation image driving method provided in the embodiments of the present application can be applied to an audio and video processing device with the ability to drive an animation image, and the audio and video processing device may be a terminal device or a server.
  • the audio and video processing equipment can have the ability to implement voice technology, so that the audio and video processing equipment can listen, see, and feel, which is the future development direction of human-computer interaction, and voice becomes one of the most promising human-computer interaction methods in the future .
  • the audio and video processing device can determine the first expression base of the first animated image corresponding to the speaker in the media data by implementing the above-mentioned computer vision technology, and can be based on the target text information and media data through voice technology and machine learning. , Determine the acoustic characteristics and target expression parameters corresponding to the target text information, and then use the acoustic characteristics and target expression parameters to drive the second animation image with the second expression base.
  • the audio and video processing device is a terminal device
  • the terminal device may be a smart terminal, a computer, a personal digital assistant (Personal Digital Assistant, PDA for short), a tablet computer, etc.
  • PDA Personal Digital Assistant
  • the server may be an independent server or a cluster server.
  • the terminal device can upload media data containing the facial expression of the speaker and the corresponding voice to the server.
  • the server determines the acoustic characteristics and target expression parameters, and uses the acoustic characteristics and target expression parameters to drive the terminal device The second animated image on top.
  • the artificial intelligence-based animation image driving method provided by the embodiments of the present application can be applied to various application scenarios applicable to animation images, such as news broadcast, weather forecast, game commentary, and game scenes that are allowed to be used for construction and users.
  • Game characters with the same face shape can also be used in scenes that use animated images to undertake personal services, such as one-to-one services for individuals such as psychologists and virtual assistants.
  • the method provided in the embodiments of the present application can be used to drive the animated image.
  • FIG. 1 is a schematic diagram of an application scenario of an artificial intelligence-based animation image driving method provided by an embodiment of the application.
  • This application scenario is introduced by taking an audio and video processing device as a terminal device as an example.
  • the application scenario includes a terminal device 101, and the terminal device 101 can obtain media data containing the facial expression of the speaker and the corresponding voice.
  • the media data can be one or multiple.
  • the media data can be video, video and audio.
  • the languages corresponding to the characters included in the voice in the media data can be Chinese, English, Korean and other languages.
  • Facial expressions can be actions made by the face when the speaker speaks the voice, for example, it can include mouth shape, eye movements, eyebrow movements, etc.
  • the video viewer can feel the voice in the media data through the face expression of the speaker That's what the speaker said.
  • the terminal device 101 can determine the first expression base of the first animation image corresponding to the speaker according to the facial expressions, and the first expression base is used to identify different expressions of the first animation image.
  • an expression form of the expression parameter and the face pinching parameter that may be involved later may be a coefficient, for example, a vector having a certain dimension.
  • the acoustic features and target expression parameters are obtained based on the media data, and they correspond to the same time axis. Therefore, the sound identified by the acoustic features and the target expression parameters are identified 'S expression changes synchronously on the same timeline.
  • the generated acoustic feature is a sequence related to the time axis
  • the target expression parameter is a sequence related to the same time axis, and the two can be adjusted accordingly as the text information changes.
  • the acoustic feature is used to identify the voice that simulates the target text information of the speaker in the above-mentioned media data
  • the target expression parameter is used to identify the facial expression that simulates the target text information of the speaker in the above-mentioned media data. State the degree of change of the first expression base.
  • the terminal device 101 can drive the second animation image with the second expression base through the acoustic characteristics and the target expression parameters, so that the second animation image can simulate the sound of the speaker uttering the target text information through the acoustic characteristics, and is uttering
  • a facial expression that conforms to the expression of the speaker is made.
  • the second animation image may be the same animation image as the first animation image, or it may be an animation image different from the first animation image, which is not limited in the embodiment of the application.
  • the method includes:
  • S201 Acquire media data including the facial expression of the speaker and the corresponding voice.
  • Media data containing facial expressions and corresponding voices can be obtained by recording the voice spoken by the speaker in a recording environment with a camera, and recording the facial expressions of the speaker corresponding to the speaker through the camera.
  • the media data is the video; if the video captured by the camera includes the facial expression of the speaker, and the voice is through other devices such as
  • the media data collected by the recording device includes video and audio. At this time, the video and audio are collected simultaneously.
  • the video includes the facial expression of the speaker, and the audio includes the speaker's voice.
  • the method provided in the embodiments of the present application can be implemented by an animation image driving system.
  • the system can be seen in Figure 3 and mainly includes four parts, namely a data acquisition module, a face modeling module, an acoustic feature, and Expression parameter determination module and animation drive module.
  • the data acquisition module is used to perform S201
  • the face modeling module is used to perform S202
  • the acoustic feature and expression parameter determination module is used to perform S203
  • the animation drive module is used to perform S204.
  • the media data including the facial expression of the speaker and the corresponding voice can be obtained through the data collection module.
  • the data collection module can have more choices.
  • the data collection module can collect media data including the speaker's voice and facial expressions through professional equipment, such as motion capture systems, facial expression capture systems and other professional equipment to capture speech Human facial expressions, facial expressions can be, for example, facial movements, expressions, mouth shapes, etc., use professional recording equipment to record the speaker’s voice, and synchronize the data between the voice and facial expressions by triggering synchronization signals between different devices and many more.
  • professional equipment is not limited to the use of expensive capture systems. It can also be a multi-view ultra-high-definition device.
  • the multi-view ultra-high-definition device collects videos including the speaker's voice and facial expressions.
  • the data collection module can also collect media data including the speaker's voice and facial expressions in a multi-camera surround mode.
  • the collection environment needs to have stable ambient lighting, and the speaker is not required to wear specific clothes.
  • Figure 4 takes three ultra-high-definition cameras as an example.
  • the upper dashed arrow represents stable lighting, and the three arrows on the left represent the relationship between the viewing angle of the ultra-high-definition camera and the speaker, so as to collect the voice and face of the speaker.
  • Media data of the facial expressions At this time, the video captured by the ultra-high-definition camera can include both voice and facial expressions, that is, the media data is a video.
  • the form of expression of the collected media data may be different according to the different sensors used to collect facial expressions.
  • the establishment of a face model can be achieved by shooting a speaker with a Red Green Blue Deep (RGBD) sensor. Since the RGBD sensor can collect depth information and obtain the speaker's three-dimensional reconstruction result, the media data includes the static modeling of the speaker's corresponding face, that is, 3 dimensions (3D) data. In other cases, there may be no RGBD sensor but a two-dimensional sensor to take pictures of the speaker. At this time, there is no 3D reconstruction result of the speaker.
  • the media data includes the video frame corresponding to the speaker, that is, 2 Dimensions (2 Dimensions). , 2D) data.
  • S202 Determine the first expression base of the first animation image corresponding to the speaker according to the facial expression.
  • the face modeling module in FIG. 3 can be used to model the face of the speaker, thereby obtaining the first expression base of the first animation image corresponding to the speaker, and the first expression base is used To identify the expression of the first animated image.
  • the purpose of modeling the face is to make the collected objects, such as the aforementioned speaker, be understood and stored by the computer, including the shape and texture of the collected objects.
  • face modeling mainly from three perspectives: hardware, manual, and software.
  • the hardware perspective can be realized by using professional equipment to perform high-precision scanning of the speaker, such as a 3D scanning instrument, and the obtained face model can be selected to manually/automatically clean up the data;
  • the artificial perspective can be realized by manually designing the data by the art designer , Clean up the data, adjust the data;
  • the software perspective can be realized by using the parameterized face pinching algorithm to automatically generate the speaker's face model.
  • a professional face scanning device can be used to scan a speaker with an expression, and a parameterized description of the current expression will be automatically given. This description is related to the custom expression description in the scanning device.
  • the expression parameters manually adjusted by the art designer it is generally necessary to predefine the expression type and the corresponding face parameterization, such as the opening and closing degree of the mouth, the movement range of the facial muscles, and so on.
  • the mathematical description of the face in different expressions For example, by decomposing a large amount of real face data, the principal component analysis (PCA) method can be used to obtain the best expression.
  • PCA principal component analysis
  • the software-based face modeling and expression parameterization are mainly introduced.
  • the mathematical description of the face in different expressions can be defined through the model library.
  • the animation images in the embodiments of the present application may be models in the model library, or may be obtained by linear combination of the models in the model library.
  • the model library may be a human face 3D deformable model (3DMM) library or other model libraries, which is not limited in this implementation.
  • the animated image can be a 3D grid.
  • the 3DMM library is obtained from a large amount of high-precision face data through the principal component analysis method, and describes the main changes of high-dimensional face shape and expression relative to the average face, and can also describe texture information.
  • the 3DMM library when the 3DMM library describes an expressionless face, it can be obtained by mu+ ⁇ (Pface i -mu)* ⁇ i .
  • mu is the average face under natural expression
  • Pface i is the i-th face principal component component
  • ⁇ i is the weight of each face principal component component, that is, the face pinch parameter.
  • the grid corresponding to the animation image in the 3DMM library can be represented by M, that is, the relationship between the face shape, expression and vertices in the 3DMM library is represented by M, and M is a three-dimensional matrix of [m ⁇ n ⁇ d], where each One dimension is the vertex coordinates of the grid (m), the principal component of face shape (n), and the principal component of expression (d).
  • M is a three-dimensional matrix of [m ⁇ n ⁇ d], where each One dimension is the vertex coordinates of the grid (m), the principal component of face shape (n), and the principal component of expression (d).
  • the distribution and meaning of each dimension of the 3DMM library M is shown in Figure 5, and each coordinate axis represents the vertex coordinates (m), the principal component of face (n), and the principal component of expression (d).
  • M can be a two-dimensional matrix.
  • the texture dimension in the 3DMM library is not considered, and assuming that the driving of the animated image is F, then:
  • M is the grid of the animated image
  • is the face pinching parameter
  • is the expression parameter
  • n is the number of face pinching grids in the face pinching base
  • d is the number of expression grids in the expression base
  • M k, j,i is the k-th grid with the i-th expression grid and the j-th face pinching grid
  • ⁇ j is the j-th dimension in a set of face pinching parameters, representing the weight of the j-th face principal component component
  • ⁇ i is the i-th dimension in a set of expression parameters, and represents the weight of the i-th expression principal component.
  • the process of determining the face pinching parameters is the face pinching algorithm
  • the process of determining the expression parameters is the face pinching algorithm.
  • the pinching parameters are used to make a linear combination with the pinching base to obtain the corresponding face shape.
  • a pinching base including 50 pinching grids (belonging to deformable grids, such as blendshape), and the pinching face corresponding to the pinching base
  • the parameter is a 50-dimensional vector, and each dimension can identify the degree of correlation between the face shape corresponding to the face pinching parameter and a face pinching grid.
  • the face pinch grids included in the face pinch base represent different face shapes, and each face pinch grid is a face image with a relatively large change from the average face. It is a face shape master of different dimensions obtained after a large number of faces are decomposed by PCA. Component, and the number of vertices corresponding to different face-pinching meshes in the same face-pinching base remains the same.
  • the expression parameters are used to linearly combine with the expression base to obtain the corresponding expression.
  • an expression base including 50 (equivalent to 50) expression grids (belonging to deformable grids, such as blendshape), which corresponds to
  • the expression parameter of is a 50-dimensional vector, and each dimension can identify the degree of correlation between the expression corresponding to the expression parameter and an expression grid.
  • the expression grids included in the expression base represent different expressions. Each expression grid is formed by changing the same 3D model under different expressions. The number of vertices corresponding to different expression grids in the same expression base remains the same.
  • a single grid can be deformed through a predefined shape to obtain any number of grids.
  • the first expression base of the first animated image corresponding to the speaker can be obtained, which can be used to drive the subsequent second animated image.
  • S203 Determine acoustic features and target expression parameters corresponding to the target text information according to the target text information, the media data, and the first expression base.
  • the acoustic feature and expression parameter determination module in FIG. 3 can determine the acoustic characteristics and target expression parameters corresponding to the target text information. Among them, the acoustic feature is used to identify the voice of the simulated speaker uttering the target text information, and the target expression parameters are used to identify the degree of change of the facial expression of the simulated speaker uttering the target text information with respect to the first expression base.
  • the target text information may be obtained in multiple ways.
  • the target text information may be input by the user through the terminal device, or may be obtained by converting the voice input to the terminal device.
  • the expressions identified by the target expression parameters are combined with the voices identified by the acoustic features to be displayed in a way that humans can intuitively understand, using multiple senses.
  • the target expression parameter represents the weight of each expression grid in the second expression base, and the corresponding expression can be obtained through the weighted linear combination of the second expression base.
  • the second animation image that makes the expression corresponding to the voice is rendered through the rendering method, so as to realize the driving of the second animation image.
  • the first expression base of the first animation image corresponding to the speaker can be determined, and the first expression base can reflect the image of the first animation image.
  • Different expressions After determining the target text information used to drive the second animated image, the acoustic features and target expression parameters corresponding to the target text information can be determined according to the target text information, the aforementioned collected media data, and the first expression base.
  • the acoustic characteristics can be Identifies the voice that simulates the speaker uttering the target text information, and the target expression parameter may identify the degree of change of the facial expression simulating the speaker uttering the target text information with respect to the first expression base.
  • the second animation image with the second expression base can be driven, so that the second animation image can simulate the voice of the speaker uttering the target text information through the acoustic characteristics, and make conformance during the vocalization process.
  • the speaker should have an expressive facial expression, which brings a realistic sense of substitution and immersion to the user, and improves the user's experience of interacting with the animated image.
  • S203 may include multiple implementations, and the embodiment of the present application focuses on one implementation.
  • the implementation manner of S203 may be to determine the acoustic feature and expression feature corresponding to the target text information according to the target text information and media data.
  • the acoustic feature is used to identify the voice of the simulated speaker uttering the target text information
  • the expression feature is used to identify the facial expression of the simulated speaker uttering the target text information. Then, the target expression parameters are determined according to the first expression base and expression characteristics.
  • the facial expression and voice of the speaker have been synchronously recorded in the media data, that is, the facial expression and voice of the speaker in the media data correspond to the same time axis. Therefore, a large amount of media data can be pre-collected offline as training data, and text features, acoustic features, and expression features can be extracted from these media data, and the duration model, acoustic model, and expression model can be trained based on these features.
  • the duration model can be used to determine the duration corresponding to the target text information, and then the duration and the corresponding text characteristics of the target text information are determined by the acoustic model and the expression model respectively Corresponding acoustic features and expression features. Since the acoustic features and expression features are based on the time length obtained by the same time length model, it is easy to synchronize the voice and expression, so that the second animation image simulates the speech while simulating the speaker to say the target text information corresponding to the voice. People make corresponding expressions.
  • the second animation image may be the same animation image as the first animation image, or it may be an animation image different from the first animation image. In these two cases, the implementation of S204 may be different.
  • the first case the first animation image and the second animation image are the same animation image.
  • the animation image to be driven is the first animation image.
  • the first animation image in order to drive the first animation image, in addition to determining the first expression base, it is also necessary to determine the pinch face parameters of the first animation image to obtain the face shape of the first animation image. Therefore, in S202, the first expression base of the first animated image and the pinching parameters of the first animated image can be determined according to the facial expressions.
  • the pinching parameters are used to identify that the face of the first animated image is relative to the first animated image. Corresponding to the degree of change of the pinch face base.
  • the first expression base of the first animation image and the pinching parameters of the first animation image There are many ways to determine the first expression base of the first animation image and the pinching parameters of the first animation image.
  • the collected media data often has low accuracy and high noise, which makes the quality of the established face model not high and has many uncertainties, making it difficult to build a face model.
  • an embodiment of the present application provides a method for determining face pinching parameters, as shown in FIG. 6.
  • the initial pinching parameters can be determined based on the first vertex data among them and the target vertex data used to identify the target facial model in the 3DMM library .
  • the expression parameters are determined based on the initial pinching parameters and the target vertex data. After that, the expression parameters are fixed, and the pinching parameters are reversed or reversed. Push how to change the face shape to obtain the face image of the speaker under the expression parameters, that is, to correct the initial pinch parameters by fixing the expression to reverse the face shape to obtain the target pinch parameters, and then use the target pinch parameters as the first animation Image pinch face parameters.
  • the target pinching parameters corrected by the second vertex data can offset the noise in the first vertex data to a certain extent, and the face model corresponding to the speaker determined by the target pinching parameters The accuracy is relatively higher.
  • the determined target expression parameters can directly drive the second animation image. Therefore, the second animation is driven in S204
  • the way of the image may be to drive the second animation image with the second expression base according to the acoustic characteristics, target expression parameters and face pinch parameters.
  • the second case the first animation image and the second animation image are different animation images.
  • the first expression base is different from the second expression base, that is, the dimensions of the two and the semantic information of each dimension are different, so it is difficult to directly use the target expression parameters to drive the second animation with the second expression base Image.
  • the expression parameters corresponding to the first animation image and the expression parameters corresponding to the second animation image should have a mapping relationship, the mapping relationship between the expression parameters corresponding to the first animation image and the expression parameters corresponding to the second animation image can be through the function f() ,
  • the formula for calculating the expression parameters corresponding to the second animation image from the expression parameters corresponding to the first animation image is as follows:
  • ⁇ b is the expression parameter corresponding to the second animation image
  • ⁇ a is the expression parameter corresponding to the first animation image
  • f() represents the mapping between the expression parameters corresponding to the first animation image and the expression parameters corresponding to the second animation image. relationship.
  • the mapping relationship may be a linear mapping relationship or a non-linear mapping relationship.
  • mapping relationship In order to drive the second animation image with the second expression base according to the target expression parameters, the mapping relationship needs to be determined. There may be multiple ways to determine the mapping relationship, and this embodiment mainly introduces two ways of determining.
  • the first determination method may be to determine the mapping relationship between the expression parameters based on the first expression base corresponding to the first animation image and the second expression base corresponding to the second animation image.
  • the actual expression parameters can reflect the degree of correlation between the actual expression and its expression base in different dimensions, that is, the first
  • the actual expression parameters corresponding to the second animation image can also reflect the degree of correlation between the actual expression of the second animation image and its expression base in different dimensions. Therefore, based on the above-mentioned correlation between the expression parameters and the expression base, it can be based on the corresponding relationship of the first animation image.
  • the first expression base and the second expression base corresponding to the second animation image determine the mapping relationship between the expression parameters. Then, according to the acoustic characteristics, the target expression parameters and the mapping relationship, the second animation image with the second expression base is driven.
  • the second determination method may be to determine the mapping relationship between the expression parameters based on the preset relationship between the phoneme and the second expression base.
  • Phoneme is the smallest phonetic unit divided according to the natural attributes of the speech. It is analyzed according to the pronunciation actions in the syllable.
  • An action (such as mouth shape) constitutes a phoneme.
  • the phoneme has nothing to do with the speaker, no matter who the speaker is, whether the speech is English or Chinese, whether the text corresponding to the phoneme is the same, as long as the phoneme in a time interval in the speech is the same, then the corresponding expression is for example
  • the mouth shape is consistent. Refer to Figure 8, which shows the correspondence between time intervals and phonemes, and describes which phoneme corresponds to which time interval in a speech.
  • the corresponding video frames can be easily divided by voice, that is, the phoneme identified by the voice, the time interval corresponding to the phoneme, and the media data are determined according to the media data.
  • the video frame of the time interval Then, the first expression parameter corresponding to the phoneme is determined according to the video frame, and the first expression parameter is used to identify the degree of change of the facial expression of the speaker relative to the first expression base when the phoneme is emitted.
  • the corresponding time interval is 5.65 seconds to 6.3 seconds.
  • the first expression parameter If the first animation image is an animation image a, the first expression parameter can be represented by ⁇ a .
  • the dimension of the first group is n a expression
  • the resulting parameter ⁇ a first expression of a set of vectors of length n a.
  • the premise of the method of determining the mapping relationship is that the expression bases of other animated images, such as the second expression base corresponding to the second animated image, are generated according to the preset relationship with the phoneme, and the preset relationship indicates that one phoneme corresponds to one expression network.
  • the phoneme "u" in the preset relationship corresponds to the first expression grid
  • the phoneme "i" corresponds to the second expression grid...
  • the second expression base including n b expression grids can be determined.
  • the second expression parameter corresponding to the phoneme can be determined according to the preset relationship and the second expression base.
  • the mapping relationship is determined according to the first expression parameter and the second expression parameter.
  • the phoneme identified by the voice is "u”
  • ⁇ b includes n b elements, except for the first element which is 1, the other n b -1 elements are all 0.
  • the formula for determining the mapping relationship can be:
  • f is the mapping relationship
  • ⁇ A is the first matrix
  • ⁇ B is the second matrix
  • inv is the matrix inversion operation.
  • the foregoing embodiment mainly introduced how to drive an animated image based on text information.
  • the first animation image corresponding to the speaker in the media data has a first expression base
  • the dimension of the first expression base is the first dimension
  • the vertex topology is the first vertex topology.
  • the first expression base can be represented by Ea
  • One dimension can be represented by Na
  • the first vertex topology can be represented by Ta
  • the first expression base Ea looks like Fa
  • the second animation image to be driven has a second expression base
  • the dimension of the second expression base is second Dimension
  • the vertex topology is the second vertex topology
  • the second expression base can be represented by Eb
  • the second dimension can be represented by Nb
  • the second vertex topology can be represented by Tb.
  • the second expression base Eb looks like Fb.
  • the media data including the facial expression and voice of the speaker drives the second animated image.
  • an embodiment of the present application also provides an artificial intelligence-based animation image driving method. As shown in FIG. 9, the method includes:
  • S902 Determine the first expression base of the first animation image corresponding to the speaker according to the facial expression.
  • S903 Determine the target expression base according to the first expression base and the second expression base of the second animation image to be driven.
  • the dimension of the first expression base is different from the dimension of the second expression base, in order to use the facial expression and voice of the speaker in the media data to drive the second animation image, a new animation image can be constructed.
  • the expression base of is, for example, the target expression base, so that the target expression base has the characteristics of the first expression base and the second expression base at the same time.
  • the implementation manner of S903 may be: determine from the first expression base the corresponding expressionless grid when the first animation image is in the absence of expression, and determine from the second expression base that the second animation image is in the absence of expression.
  • the expressionless grid corresponding to the expression According to the expressionless grid corresponding to the first image and the expressionless grid corresponding to the second image, an adjustment grid is determined, and the adjustment grid has a second vertex topology and is used to identify the first animated image in an expressionless state. According to the grid deformation relationship in the adjusted grid and the second expression base, the target expression base is generated.
  • the flowchart of the method can also be referred to as shown in FIG. 10a.
  • the target expression base Eb' is determined based on the first expression base Ea and the second expression base Eb. Wherein, the method of determining the target expression base Eb' may be to extract the expressionless grid of the second expression base Eb and the expressionless grid of the first expression base Ea.
  • the expressionless mesh of Eb is pasted to the expressionless mesh of Ea, so that the expressionless mesh of Eb changes its appearance and becomes the appearance of Ea while maintaining the vertex topology Fb.
  • Get the adjustment grid which can be expressed as Newb.
  • the mesh deformation relationship of the expressions in each dimension in Newb and the second expression base Eb relative to the natural expression (no expression) is known, the mesh deformation relationship in Newb and the second expression base Eb can be obtained from The target expression base Eb' is transformed in Newb.
  • the appearance of the target expression base Eb' is Fa
  • the dimension is Nb
  • the vertex topology is Tb.
  • S904 Determine target expression parameters and acoustic features according to the second media data including the facial expression and corresponding voice of the speaker and the target expression base.
  • the target expression parameter Bb is used to identify the degree of change of the facial expression of the speaker speaking the voice relative to the target expression base.
  • target expression parameters and acoustic features obtained by this method can be used to retrain the aforementioned acoustic models and expression models.
  • the terminal device can obtain the media data including the facial expression of the speaker and the corresponding voice, and determine the first expression base of the first animation image corresponding to the speaker according to the facial expression.
  • the acoustic characteristics and target expression parameters corresponding to the target text information are determined, so as to drive the second animation with the second expression base according to the acoustic characteristics and target expression parameters
  • the image makes the second animated image emit the voice corresponding to the target text information and make the corresponding expression.
  • the user can see that the game character imitates the speaker's voice and makes a corresponding expression, which brings a realistic sense of substitution and immersion to the user, and improves the user's experience of interacting with the animated image.
  • this embodiment also provides an animated character driving device 1000, which is deployed on audio and video processing equipment.
  • the device 1000 includes an acquiring unit 1001, a first determining unit 1002, a second determining unit 1003, and a driving unit 1004:
  • the acquiring unit 1001 is configured to acquire media data containing the facial expression of the speaker and the corresponding voice;
  • the first determining unit 1002 is configured to determine a first expression base of the first animated image corresponding to the speaker according to the facial expression, and the first expression base is used to identify the expression of the first animated image ;
  • the second determining unit 1003 is configured to determine acoustic features and target expression parameters corresponding to the target text information according to the target text information, the media data, and the first expression base; the acoustic characteristics are used to identify simulations The speaker utters the voice of the target text information, and the target expression parameter is used to identify the degree of change of the facial expression that simulates the speaker uttering the target text information with respect to the first expression base;
  • the driving unit 1004 is configured to drive a second animation character with a second expression base according to the acoustic characteristics and the target expression parameters.
  • the first animation image and the second animation image are the same animation image
  • the first expression base is the same as the second expression base
  • the first determining unit 1002 For:
  • the first expression base of the first animated avatar and the pinch face parameters of the first animated avatar are determined according to the facial expressions, and the face pinch parameters are used to identify that the face of the first animated avatar is relative to the face shape of the first animated avatar.
  • the driving unit 1004 is used for:
  • the second animated image is driven according to the acoustic feature, the target expression parameter, and the face pinching parameter.
  • the first animation image and the second animation image are different animation images
  • the first expression base is different from the second expression base
  • the driving unit 1004 is configured to :
  • the second animated image is driven.
  • the second expression base is generated according to a preset relationship between the second expression base and phonemes, and the driving unit 1004 is further configured to:
  • the mapping relationship is determined according to the first expression parameter and the second expression parameter.
  • the second determining unit 1003 is configured to:
  • the target text information and the media data determine the acoustic characteristics and the expression characteristics corresponding to the target text information; the expression characteristics are used to identify facial expressions that simulate the speaker uttering the target text information;
  • the target expression parameter is determined according to the first expression base and the expression feature.
  • the device 1100 includes an acquiring unit 1101, a first determining unit 1102, a second determining unit 1103, a third determining unit 1104, and a driving unit 1105:
  • the acquiring unit 1101 is configured to acquire first media data including the facial expression of the speaker and the corresponding voice;
  • the first determining unit 1102 is configured to determine a first expression base of the first animated image corresponding to the speaker according to the facial expression, and the first expression base is used to identify the expression of the first animated image ;
  • the dimension of the first expression base is the first dimension, and the vertex topology is the first vertex topology;
  • the second determining unit 1103 is configured to determine a target expression base according to the first expression base and the second expression base of the second animation image to be driven; the dimension of the second expression base is the second dimension , The vertex topology is the second vertex topology, the target expression base is the expression base corresponding to the first animation image with the second vertex topology, and the dimension of the target expression base is the second dimension;
  • the third determining unit 1104 is configured to determine target expression parameters and acoustic features according to the second media data containing the facial expression and corresponding voice of the speaker and the target expression base; the target expression parameters are used for Identifying the degree of change in the facial expression of the speaker uttering the voice relative to the target expression base;
  • the driving unit 1105 is configured to drive the second animation character with the second expression base according to the target expression parameters and acoustic characteristics.
  • the second determining unit 1103 is configured to determine from the first expression base the expressionless grid corresponding to the first animation image in the expressionless state, and obtain the expression from the first expression base. In the second expression base, determine the expressionless grid corresponding to the second animation image when it is expressionless;
  • an adjustment grid is determined.
  • the adjustment grid has a second vertex topology and is used to identify when in an expressionless state.
  • the target expression base is generated.
  • the embodiment of the present application also provides a device for driving an animation image.
  • the device can drive the animation through voice, and the device can be an audio and video processing device.
  • the equipment will be introduced below in conjunction with the drawings.
  • an embodiment of the present application provides a device for driving animated characters.
  • the device may also be a terminal device.
  • the terminal device may include a mobile phone, a tablet computer, and a personal digital assistant (Personal Digital Assistant, Any smart terminal such as PDA), Point of Sales (POS), on-board computer, etc. Take the terminal device as a mobile phone as an example:
  • FIG. 12 shows a block diagram of a part of the structure of a mobile phone related to a terminal device provided in an embodiment of the present application.
  • the mobile phone includes: Radio Frequency (RF) circuit 1210, memory 1220, input unit 1230, display unit 1240, sensor 1250, audio circuit 1260, wireless fidelity (wireless fidelity, WiFi for short) module 1270, processing 1280, and power supply 1290 and other components.
  • RF Radio Frequency
  • the RF circuit 1210 can be used for receiving and sending signals during the process of sending and receiving information or talking. In particular, after receiving the downlink information of the base station, it is processed by the processor 1280; in addition, the designed uplink data is sent to the base station.
  • the RF circuit 1210 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA for short), a duplexer, and the like.
  • the RF circuit 1210 can also communicate with the network and other devices through wireless communication.
  • the above-mentioned wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access ( Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), Email, Short Message Service (Short Messaging Service, SMS) Wait.
  • GSM Global System of Mobile Communication
  • GPRS General Packet Radio Service
  • CDMA Code Division Multiple Access
  • WCDMA Wideband Code Division Multiple Access
  • LTE Long Term Evolution
  • Email Short Message Service
  • SMS Short Messaging Service
  • the memory 1220 may be used to store software programs and modules.
  • the processor 1280 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1220.
  • the memory 1220 may mainly include a program storage area and a data storage area.
  • the program storage area may store an operating system, an application program required by at least one function (such as a sound playback function, an image playback function, etc.), etc.; Data created by the use of mobile phones (such as audio data, phone book, etc.), etc.
  • the memory 1220 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
  • the input unit 1230 can be used to receive input digital or character information, and generate key signal input related to the user settings and function control of the mobile phone.
  • the input unit 1230 may include a touch panel 1231 and other input devices 1232.
  • the touch panel 1231 also called a touch screen, can collect the user's touch operations on or near it (for example, the user uses any suitable objects or accessories such as fingers, stylus, etc.) on the touch panel 1231 or near the touch panel 1231. Operation), and drive the corresponding connection device according to the preset program.
  • the touch panel 1231 may include two parts: a touch detection device and a touch controller.
  • the touch detection device detects the user's touch position, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it To the processor 1280, and can receive and execute the commands sent by the processor 1280.
  • the touch panel 1231 can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave.
  • the input unit 1230 may also include other input devices 1232.
  • the other input device 1232 may include, but is not limited to, one or more of a physical keyboard, function keys (such as volume control buttons, switch buttons, etc.), trackball, mouse, and joystick.
  • the display unit 1240 can be used to display information input by the user or information provided to the user and various menus of the mobile phone.
  • the display unit 1240 may include a display panel 1241.
  • the display panel 1241 may be configured in the form of a liquid crystal display (Liquid Crystal Display, LCD for short), Organic Light-Emitting Diode (OLED for short), etc.
  • the touch panel 1231 can cover the display panel 1241. When the touch panel 1231 detects a touch operation on or near it, it is transmitted to the processor 1280 to determine the type of the touch event, and then the processor 1280 determines the type of the touch event. The type provides corresponding visual output on the display panel 1241.
  • the touch panel 1231 and the display panel 1241 are used as two independent components to implement the input and input functions of the mobile phone, in some embodiments, the touch panel 1231 and the display panel 1241 can be integrated. Realize the input and output functions of the mobile phone.
  • the mobile phone may also include at least one sensor 1250, such as a light sensor, a motion sensor, and other sensors.
  • the light sensor can include an ambient light sensor and a proximity sensor.
  • the ambient light sensor can adjust the brightness of the display panel 1241 according to the brightness of the ambient light.
  • the proximity sensor can close the display panel 1241 and/or when the mobile phone is moved to the ear. Or backlight.
  • the accelerometer sensor can detect the magnitude of acceleration in various directions (usually three-axis), and can detect the magnitude and direction of gravity when it is stationary.
  • the audio circuit 1260, the speaker 1261, and the microphone 1262 can provide an audio interface between the user and the mobile phone.
  • the audio circuit 1260 can transmit the electrical signal converted from the received audio data to the speaker 1261, which is converted into a sound signal for output by the speaker 1261; on the other hand, the microphone 1262 converts the collected sound signal into an electrical signal, which is then output by the audio circuit 1260. After being received, it is converted into audio data, and then processed by the audio data output processor 1280, and sent to, for example, another mobile phone via the RF circuit 1210, or the audio data is output to the memory 1220 for further processing.
  • WiFi is a short-distance wireless transmission technology.
  • the mobile phone can help users send and receive emails, browse web pages, and access streaming media through the WiFi module 1270. It provides users with wireless broadband Internet access.
  • FIG. 12 shows the WiFi module 1270, it is understandable that it is not a necessary component of the mobile phone, and can be omitted as needed without changing the essence of the invention.
  • the processor 1280 is the control center of the mobile phone. It uses various interfaces and lines to connect various parts of the entire mobile phone. It executes by running or executing software programs and/or modules stored in the memory 1220, and calling data stored in the memory 1220. Various functions and processing data of the mobile phone can be used to monitor the mobile phone as a whole.
  • the processor 1280 may include one or more processing units; preferably, the processor 1280 may integrate an application processor and a modem processor, where the application processor mainly processes the operating system, user interface, application programs, etc. , The modem processor mainly deals with wireless communication. It can be understood that the foregoing modem processor may not be integrated into the processor 1280.
  • the mobile phone also includes a power supply 1290 (such as a battery) for supplying power to various components.
  • a power supply 1290 (such as a battery) for supplying power to various components.
  • the power supply can be logically connected to the processor 1280 through a power management system, so that functions such as charging, discharging, and power management can be managed through the power management system.
  • the mobile phone may also include a camera, a Bluetooth module, etc., which will not be repeated here.
  • the processor 1280 included in the terminal device also has the following functions:
  • the acoustic features and target expression parameters corresponding to the target text information are determined; the acoustic features are used to identify and simulate the speaker uttering the target text
  • the voice of the information, the target expression parameter is used to identify the degree of change of the facial expression that simulates the speaker uttering the target text information with respect to the first expression base;
  • a second animation character with a second expression base is driven.
  • the first expression base of the first animated image corresponding to the speaker is determined according to the facial expression, the first expression base is used to identify the expression of the first animated image; the dimension of the first expression base Is the first dimension, and the vertex topology is the first vertex topology;
  • the target expression base is determined;
  • the dimension of the second expression base is the second dimension, and the vertex topology is the second vertex topology, so
  • the target expression base is the expression base corresponding to the first animation image with the second vertex topology, and the dimension of the target expression base is the second dimension;
  • the target expression parameters are used to identify that the speaker uttered the voice The degree of change of the facial expression relative to the target expression base;
  • the second animation character with the second expression base is driven.
  • FIG. 13 is a structural diagram of the server 1300 provided by the embodiment of the present application.
  • the server 1300 may have relatively large differences due to different configurations or performance, and may include one or one The above central processing unit (Central Processing Units, CPU for short) 1322 (for example, one or more processors) and memory 1332, one or more storage media 1330 for storing application programs 1342 or data 1344 (for example, one or more storage equipment).
  • the memory 1332 and the storage medium 1330 may be short-term storage or persistent storage.
  • the program stored in the storage medium 1330 may include one or more modules (not shown in the figure), and each module may include a series of command operations on the server.
  • the central processing unit 1322 may be configured to communicate with the storage medium 1330, and execute a series of instruction operations in the storage medium 1330 on the server 1300.
  • the server 1300 may also include one or more power supplies 1326, one or more wired or wireless network interfaces 1350, one or more input and output interfaces 1358, and/or one or more operating systems 1341, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
  • operating systems 1341 such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
  • the steps performed by the server in the foregoing embodiment may be based on the server structure shown in FIG. 13.
  • the embodiments of the present application also provide a computer-readable storage medium, where the computer-readable storage medium is used to store program code, and the program code is used to execute the animation image driving method described in each of the foregoing embodiments.
  • the embodiments of the present application also provide a computer program product including instructions, which when run on a computer, cause the computer to execute the animation image driving method described in each of the foregoing embodiments.
  • At least one (item) refers to one or more, and “multiple” refers to two or more.
  • “And/or” is used to describe the association relationship of associated objects, indicating that there can be three types of relationships, for example, “A and/or B” can mean: only A, only B, and both A and B , Where A and B can be singular or plural.
  • the character “/” generally indicates that the associated objects before and after are in an “or” relationship.
  • the following at least one item (a) or similar expressions refers to any combination of these items, including any combination of a single item (a) or a plurality of items (a).
  • At least one of a, b, or c can mean: a, b, c, "a and b", “a and c", “b and c", or "a and b and c" ", where a, b, and c can be single or multiple.
  • the disclosed system, device, and method may be implemented in other ways.
  • the device embodiments described above are merely illustrative, for example, the division of the units is only a logical function division, and there may be other divisions in actual implementation, for example, multiple units or components may be combined or It can be integrated into another system, or some features can be ignored or not implemented.
  • the displayed or discussed mutual coupling or direct coupling or communication connection may be indirect coupling or communication connection through some interfaces, devices or units, and may be in electrical, mechanical or other forms.
  • the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the objectives of the solutions of the embodiments.
  • the functional units in the various embodiments of the present application may be integrated into one processing unit, or each unit may exist alone physically, or two or more units may be integrated into one unit.
  • the above-mentioned integrated unit can be implemented in the form of hardware or software functional unit.
  • the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium.
  • the technical solution of the present application essentially or the part that contributes to the existing technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium , Including several instructions to make a computer device (which may be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in the various embodiments of the present application.
  • the aforementioned storage media include: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), magnetic disks or optical disks, etc., which can store program codes. Medium.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Human Computer Interaction (AREA)
  • General Engineering & Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • General Health & Medical Sciences (AREA)
  • Acoustics & Sound (AREA)
  • Computational Linguistics (AREA)
  • Oral & Maxillofacial Surgery (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Data Mining & Analysis (AREA)
  • Quality & Reliability (AREA)
  • Signal Processing (AREA)
  • Processing Or Creating Images (AREA)

Abstract

一种基于人工智能的动画形象驱动方法和相关装置,采集说话人说出语音时脸部表情变化的媒体数据,确定说话人所对应第一动画形象的第一表情基,第一表情基可以体现第一动画形象的不同表情。在确定出用于驱动第二动画形象的目标文本信息后,根据目标文本信息、前述采集的媒体数据和第一表情基,确定对应目标文本信息的声学特征和目标表情参数。通过声学特征和目标表情参数,可以驱动具有第二表情基的第二动画形象,使得第二动画形象可以通过声学特征模拟发出说话人说出目标文本信息的声音,并且在发声过程中做出符合该说话人应有表情的脸部表情,给用户带来逼真的代入感和沉浸感,提高了用户与动画形象进行交互的体验。

Description

一种基于人工智能的动画形象驱动方法和相关装置
本申请要求于2019年9月2日提交中国专利局、申请号201910824770.0、申请名称为“一种基于人工智能的动画形象驱动方法和装置”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及数据处理领域,特别是涉及动画形象驱动。
背景技术
随着计算机技术的发展,人机交互已经比较常见,但多为单纯的语音交互,例如,交互设备可以根据用户输入的文字或语音确定回复内容,并播放根据回复内容合成的虚拟声音。
这种类型的人机交互带来的用户沉浸感难以满足目前用户的交互需求,为了提高用户沉浸感,具有表情变化能力例如可以口型变化的动画形象作为与用户进行交互的交互对象属于目前的研发方向。
然而,目前并没有完善的动画形象驱动方式。
发明内容
为了解决上述技术问题,本申请提供了一种基于人工智能的动画形象驱动方法和装置,给用户带来逼真的代入感和沉浸感,提高了用户与动画形象进行交互的体验。
本申请实施例公开了如下技术方案:
第一方面,本申请实施例提供一种动画形象驱动方法,所述方法由音视频处理设备执行,所述方法包括:
获取包含说话人的脸部表情和对应语音的媒体数据;
根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;
根据目标文本信息、所述媒体数据和所述第一表情基,确定对应所述目标文本信息的声学特征和目标表情参数;所述声学特征用于标识模拟所述说话人说出所述目标文本信息的声音,所述目标表情参数用于标识模拟所述说话人说出所述目标文本信息的脸部表情相对于所述第一表情基的变化程度;
根据所述声学特征和所述目标表情参数,驱动具有第二表情基的第二动画形象。
第二方面,本申请实施例提供一种动画形象驱动装置,所述装置部署在音视频处理设备上,所述装置包括获取单元、第一确定单元、第二确定单元和驱动单元:
所述获取单元,用于获取包含说话人的脸部表情和对应语音的媒体数据;
所述第一确定单元,用于根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;
所述第二确定单元,用于根据目标文本信息、所述媒体数据和所述第一表情基,确定对应所述目标文本信息的声学特征和目标表情参数;所述声学特征用于标识模拟所述说话人说出所述目标文本信息的声音,所述目标表情参数用于标识模拟所述说话人说出所述目标文本信息的脸部表情相对于所述第一表情基的变化程度;
所述驱动单元,用于根据所述声学特征和所述目标表情参数,驱动具有第二表情基的第二动画形象。
第三方面,本申请实施例提供一种动画形象驱动方法,所述方法由音视频处理设备执行,所述方法包括:
获取包含说话人的脸部表情和对应语音的第一媒体数据;
根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;所述第一表情基的维数为第一维数,顶点拓扑为第一顶点拓扑;
根据所述第一表情基和待驱动的第二动画形象的第二表情基,确定目标表情基;所述第二表情基的维数为第二维数,顶点拓扑为第二顶点拓扑,所述目标表情基为具有第二顶点拓扑的第一动画形象对应的表情基,所述目标表情基的维数为第二维数;
根据包含所述说话人的脸部表情和对应语音的第二媒体数据和所述目标表情基,确定目标表情参数和声学特征;所述目标表情参数用于标识所述说话人说出所述语音的脸部表情相对于所述目标表情基的变化程度;
根据所述目标表情参数和声学特征,驱动具有所述第二表情基的所述第二动画形象。
第四方面,本申请实施例提供一种动画形象驱动装置,所述装置部署在音视频处理设备上,所述装置包括获取单元、第一确定单元、第二确定单元、第三确定单元和驱动单元:
所述获取单元,用于获取包含说话人的脸部表情和对应语音的第一媒体数据;
所述第一确定单元,用于根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;所述第一表情基的维数为第一维数,顶点拓扑为第一顶点拓扑;
所述第二确定单元,用于根据所述第一表情基和待驱动的第二动画形象的第二表情基,确定目标表情基;所述第二表情基的维数为第二维数,顶点拓扑为第二顶点拓扑,所述目标表情基为具有第二顶点拓扑的第一动画形象对应的表情基,所述目标表情基的维数为第二维数;
所述第三确定单元,用于根据包含所述说话人的脸部表情和对应语音的第二媒体数据和所述目标表情基,确定目标表情参数和声学特征;所述目标表情参数用于标识所述说话人说出所述语音的脸部表情相对于所述目标表情基的变化程度;
所述驱动单元,用于根据所述目标表情参数和声学特征,驱动具有所述第二表情基的所述第二动画形象。
第五方面,本申请实施例提供一种用于动画形象驱动的设备,所述设备包括处理器以及存储器:
所述存储器用于存储程序代码,并将所述程序代码传输给所述处理器;
所述处理器用于根据所述程序代码中的指令执行第一方面或第三方面所述的方法。
第六方面,本申请实施例提供一种计算机可读存储介质,所述计算机可读存储介质用于存储程序代码,所述程序代码用于执行第一方面或第三方面所述的方法。
第七方面,本申请实施例提供一种计算机程序产品,当所述计算机程序产品被执行时,用于执行第一方面或第三方面所述的方法。
由上述技术方案可以看出,通过采集说话人说出语音时脸部表情变化的媒体数据,可以确定说话人所对应第一动画形象的第一表情基,第一表情基可以体现第一动画形象的不同表情。在确定出用于驱动第二动画形象的目标文本信息后,可以根据目标文本信息、前述采集的媒体数据和第一表情基,确定对应目标文本信息的声学特征和目标表情参数,该声学特征可以标识模拟所述说话人说出所述目标文本信息的声音,该目标表情参数可以标识模拟所述说话人说出所述目标文本信息的脸部表情相对于所述第一表情基的变化程度。通过声学特征和目标表情参数,可以驱动具有第二表情基的第二动画形象,使得第二动画形象可以通过声学特征模拟发出说话人说出目标文本信息的声音,并且在发声过程中做出符合该说话人应有表情的脸部表情,给用户带来逼真的代入感和沉浸感,提高了用户与动画形象进行交互的体验。
附图说明
为了更清楚地说明本申请实施例或现有技术中的技术方案,下面将对实施例或现有技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1为本申请实施例提供的一种基于人工智能的动画形象驱动方法的应用场景示意图;
图2为本申请实施例提供的一种基于人工智能的动画形象驱动方法的流程图;
图3为本申请实施例提供的一种动画形象驱动系统的结构流程;
图4为本申请实施例提供的采集媒体数据的场景示例图;
图5为本申请实施例提供的3DMM库M的各个维度分布和意义示例图;
图6为本申请实施例提供的一种基于确定捏脸参数的动画形象驱动方法的应用场景示意图;
图7为本申请实施例提供的一种基于确定映射关系的动画形象驱动方法的应用场景示意图;
图8为本申请实施例提供的时间区间与音素的对应关系的示例图;
图9为本申请实施例提供的一种基于人工智能的动画形象驱动方法的流程图;
图10a为本申请实施例提供的一种基于人工智能的动画形象驱动方法的流程图;
图10b为本申请实施例提供的一种动画形象驱动装置的结构图;
图11为本申请实施例提供的一种动画形象驱动装置的结构图;
图12为本申请实施例提供的一种用于动画形象驱动的设备的结构图;
图13为本申请实施例提供的一种服务器的结构图。
具体实施方式
下面结合附图,对本申请的实施例进行描述。
目前,将具有表情变化能力的动画形象作为与用户交互的交互对象是人机交互的主要研究方向。
例如,在游戏场景中,可以构建和用户自己脸型一样的游戏人物(动画形象)等,当用户输入文字或语音时,该游戏人物可以发出语音并做出对应的表情(例如口型等);或者,在游戏场景中,构建和用户自己脸型一样的游戏人物等,当对方输入文字或语音时,该游戏人物可以根据对方的输入回复语音并做出对应的表情。
为了可以更好的驱动动画形象,例如驱动动画形象发出语音并做出对应的表情,本申请实施例提供一种基于人工智能的动画形象驱动方法。该方法通过采集说话人说出语音时脸部表情变化的媒体数据,可以确定说话人所对应第一动画形象的第一表情基,在确定出用于驱动第二动画形象的目标文本信息后,可以根据目标文本信息、前述采集的媒体数据和第一表情基,确定对应目标文本信息的声学特征和目标表情参数,从而通过声学特征和 目标表情参数驱动具有第二表情基的第二动画形象,使得第二动画形象通过声学特征模拟发出说话人说出目标文本信息的声音,并且在发声过程中做出符合该说话人应有表情的脸部表情,实现基于文本信息驱动第二动画形象。
需要强调的是,本申请实施例所提供的动画形象驱动方法是基于人工智能实现的,人工智能(Artificial Intelligence,AI)是利用数字计算机或者数字计算机控制的机器模拟、延伸和扩展人的智能,感知环境、获取知识并使用知识获得最佳结果的理论、方法、技术及应用系统。换句话说,人工智能是计算机科学的一个综合技术,它企图了解智能的实质,并生产出一种新的能以人类智能相似的方式做出反应的智能机器。人工智能也就是研究各种智能机器的设计原理与实现方法,使机器具有感知、推理与决策的功能。
人工智能技术是一门综合学科,涉及领域广泛,既有硬件层面的技术也有软件层面的技术。人工智能基础技术一般包括如传感器、专用人工智能芯片、云计算、分布式存储、大数据处理技术、操作/交互系统、机电一体化等技术。人工智能软件技术主要包括计算机视觉技术、语音处理技术、自然语言处理技术以及机器学习/深度学习等几大方向。
在本申请实施例中,主要涉及的人工智能技术包括语音处理技术、机器学习和计算机视觉(图像)等方向。
例如可以涉及语音技术(Speech Technology)中的语音识别技术(Automatic Speech Recognition,ASR)、语音合成(Text To Speech,TTS)和声纹识别。语音识别技术中包括语音信号预处理(Speech signal preprocessing)、语音信号频域分析(Speech signal frequency analyzing)、语音信号特征提取(Speech signal feature extraction)、语音信号特征匹配/识别(Speech signal feature matching/recognition)、语音的训练(Speech training)等。语音合成中包括文本分析(Text analyzing)、语音生成(Speech generation)等。
例如可以涉及机器学习(Machine learning,ML),机器学习是一门多领域交叉学科,涉及概率论、统计学、逼近论、凸分析、算法复杂度理论等多门学科。专门研究计算机怎样模拟或实现人类的学习行为,以获取新的知识或技能,重新组织已有的知识结构使之不断改善自身的性能。机器学习是人工智能的核心,是使计算机具有智能的根本途径,其应用遍及人工智能的各个领域。机器学习通常包括深度学习(Deep Learning)等技术,深度学习包括人工神经网络(artificial neural network),例如卷积神经网络(Convolutional Neural Network,CNN)、循环神经网络(Recurrent Neural Network,RNN)、深度神经网络(Deep neural network,DNN)等。
例如可以涉及计算机视觉(Computer Vision)中的视频处理(video processing)、视频语义理解(video semantic understanding,VSU)、人脸识别(face recognition)等。视频语义理解中包括目标识别(target recognition)、目标检测与定位(target detection/localization)等;人脸识别中包括人脸3D重建(Face 3D Reconstruction)、人脸检测(Face Detection)、人脸跟踪(Face Tracking)等。
本申请实施例提供的基于人工智能的动画形象驱动方法可以应用于具有驱动动画形象能力的音视频处理设备上,该音视频处理设备可以是终端设备,也可以是服务器。
该音视频处理设备可以具有实施语音技术的能力,让音视频处理设备能听、能看、能感觉,是未来人机交互的发展方向,其中语音成为未来最被看好的人机交互方式之一。
在本申请实施例中,音视频处理设备通过实施上述计算机视觉技术可以确定媒体数据中说话人所对应第一动画形象的第一表情基,通过语音技术和机器学习可以根据目标文本信息和媒体数据,确定对应目标文本信息的声学特征和目标表情参数,进而利用声学特征和目标表情参数,驱动具有第二表情基的第二动画形象。
其中,若音视频处理设备是终端设备,则终端设备可以是智能终端、计算机、个人数字助理(Personal Digital Assistant,简称PDA)、平板电脑等。
若该音视频处理设备是服务器,则服务器可以为独立服务器,也可以为集群服务器。当服务器实施该方法时,终端设备可以将包含说话人的脸部表情和对应语音的媒体数据上传给服务器,服务器确定出声学特征和目标表情参数,利用该声学特征和目标表情参数驱动终端设备上的第二动画形象。
可以理解的是,本申请实施例提供的基于人工智能的动画形象驱动方法可以应用到各种适用动画形象的应用场景,例如新闻播报、天气预报、游戏解说以及游戏场景中允许用于构建和用户自己脸型一样的游戏人物等,还能用于利用动画形象承担私人化的服务的场景,例如心理医生,虚拟助手等面向个人的一对一服务。在这些场景下,利用本申请实施例提供的方法可以实现动画形象的驱动。
为了便于理解本申请的技术方案,下面结合实际应用场景对本申请实施例提供的基于人工智能的动画形象驱动方法进行介绍。
参见图1,图1为本申请实施例提供的基于人工智能的动画形象驱动方法的应用场景示意图。该应用场景以音视频处理设备为终端设备为例进行介绍,该应用场景中包括终端设备101,终端设备101可以获取包含说话人的脸部表情和对应语音的媒体数据。该媒体数据可以是一个,也可以是多个。媒体数据可以是视频,也可以是视频和音频。媒体数据中语音包括的字符所对应的语种可以是汉语、英语、韩语等各种语种。
脸部表情可以是说话人说出语音时脸部所做出的动作,例如可以包括口型、眼睛动作、眉毛动作等,视频观看者通过说话人的脸部表情可以感受到媒体数据中的语音就是该说话人说出的。
终端设备101根据脸部表情可以确定说话人所对应第一动画形象的第一表情基,第一表情基用于标识第一动画形象的不同表情。
终端设备101在确定出用于驱动第二动画形象的目标文本信息后,可以根据目标文本信息、前述采集的媒体数据和第一表情基,确定对应目标文本信息的声学特征和目标表情 参数。其中,表情参数以及后续可能涉及到的捏脸参数的一种表现形式可以是系数,例如可以是具有某一维数的向量。
由于媒体数据中语音和脸部表情是同步的,声学特征和目标表情参数均是根据媒体数据得到的,所对应的是同一个时间轴,故,声学特征所标识的声音和目标表情参数所标识的表情在同一时间轴上同步变化。生成的声学特征是与时间轴相关的一个序列,目标表情参数是与同一时间轴相关的序列,二者可以随着文本信息的变化而有相应的调整。但无论如何调整,声学特征用于标识模拟上述媒体数据中说话人说出目标文本信息的声音,目标表情参数用于标识模拟上述媒体数据中说话人说出目标文本信息的脸部表情相对于所述第一表情基的变化程度。
之后,终端设备101可以通过声学特征和目标表情参数,驱动具有第二表情基的第二动画形象,使得第二动画形象可以通过声学特征模拟发出说话人说出目标文本信息的声音,并且在发声过程中做出符合该说话人应有表情的脸部表情。其中,第二动画形象可以是与第一动画形象相同的动画形象,也可以是与第一动画形象不同的动画形象,本申请实施例对此不做限定。
接下来,将结合附图对本申请实施例提供的基于人工智能的动画形象驱动方法进行详细介绍。参见图2,所述方法包括:
S201、获取包含说话人的脸部表情和对应语音的媒体数据。
包含了脸部表情和对应语音的媒体数据可以是在有摄像头的录音环境下,录制说话人说出的语音,以及通过摄像头录制说话人对应的脸部表情得到的。
若通过摄像头采集到的视频中同时包括说话人的脸部表情和对应语音,则媒体数据为该视频;若通过摄像头采集到的视频中包括说话人的脸部表情,而语音是通过其他设备例如录音设备采集的,则媒体数据包括视频和音频,此时,该视频和音频是同步采集的,视频中包括说话人的脸部表情,音频中包括说话人的语音。
需要说明的是,本申请实施例提供的方法可以通过动画形象驱动系统实现,该系统可以参见图3所示,主要包括四个部分,分别是数据采集模块、脸部建模模块、声学特征和表情参数确定模块以及动画驱动模块。其中,数据采集模块用于执行S201,脸部建模模块用于执行S202,声学特征和表情参数确定模块用于执行S203,动画驱动模块用于执行S204。
包含说话人的脸部表情和对应语音的媒体数据可以是通过数据采集模块得到的。该数据采集模块可以有较多的选择,该数据采集模块可以通过专业设备采集包括说话人的语音和脸部表情的媒体数据,比如使用动作捕捉系统、脸部表情捕捉系统等专业设备来捕捉说话人的脸部表情,脸部表情例如可以是脸部动作、表情、口型等等,使用专业录音设备录制说话人的语音,不同设备之间通过同步信号触发实现语音和脸部表情的数据同步等等。
当然,专业设备并不局限于使用昂贵的捕捉系统,也可以是多视角超高清设备,通过多视角超高清设备采集包括说话人的语音和脸部表情的视频。
该数据采集模块还可以利用多相机环绕的方式采集包括说话人的语音和脸部表情的媒体数据。在一种可能的实现方式中,可以选择3个,5个,乃至更多的超高清相机,正面围绕说话人拍摄。采集环境中需要有稳定的环境光照,不要求说话人穿特定的衣服。参见图4所示,图4以3个超高清相机为例,上方虚线箭头表示稳定光照,左侧三个箭头表示超高清相机的视角和说话人的关系,从而采集包括说话人的语音和脸部表情的媒体数据。此时,通过超高清相机采集的视频中可以同时包括语音和脸部表情,即媒体数据为视频。
需要说明的是,在采集媒体数据时,根据采集脸部表情所使用传感器的不同,采集的媒体数据的表现形式可以有所不同。在一些情况下,可以通过具有红绿蓝深度(Red Green Blue Deep,RGBD)传感器的对说话人进行拍摄,实现对脸部模型的建立。由于RGBD传感器可以采集到深度信息,得到说话人的三维重建结果,因此,媒体数据中包括说话人对应的脸部静态建模,即3维(3 Dimensions,3D)数据。在另外一些情况下,可能没有RGBD传感器而是使用二维传感器对说话人进行拍摄,此时,没有说话人的三维重建结果,媒体数据中包括说话人对应的视频帧,即2维(2 Dimensions,2D)数据。
S202、根据脸部表情确定该说话人所对应第一动画形象的第一表情基。
在获取到上述媒体数据后,通过图3中的脸部建模模块可以对说话人进行脸部建模,从而得到说话人所对应第一动画形象的第一表情基,该第一表情基用于标识所述第一动画形象的表情。
进行脸部建模的目的在于使得被采集的对象例如前述提到的说话人可以被计算机理解并存储,包括被采集对象的形状、纹理等。进行脸部建模的方式可以包括多种,主要从硬件、人工、软件三个角度来实现。其中,硬件角度实现可以是采用专业设备对说话人进行高精度的扫描,如3D扫描仪器,对得到的脸部模型可以选择手动/自动清理数据;人工角度实现可以是由美术设计师手工设计数据、清理数据、调节数据;软件角度实现可以是采用参数化捏脸算法自动生成说话人脸部模型。
在表情参数化时,同样可以从硬件、人工、软件三个角度来实现。比如可以使用专业人脸扫描设备扫描带有表情的说话人之后,会自动给出对当前表情的参数化描述,这种描述与扫描设备中自定义的表情描述相关。而对于美术设计师手工调节的表情参数,一般需要预先定义表情类型和对应的人脸参数化,比如嘴巴的张合程度,脸部肌肉的运动幅度等等。而对于软件实现表情参数化,一般需要定义脸部在不同表情中的数学描述,比如通过对大量真实脸部数据,进行主成分分析方法(Principal Component Analysis,PCA)分解之后,得到最能体现各个表情相对平均脸的变化程度的数字描述。
在本实施例中,主要对基于软件的脸部建模和表情参数化进行介绍。在这种情况下,脸部在不同表情中的数学描述可以通过模型库定义。本申请实施例中的动画形象(例如第一动画形象和后续的第二动画形象)可以为模型库中的模型,也可以是通过模型库中模型的线性组合得到的。该模型库可以是人脸3D可变形模型(3DMM)库,也可以是其他模型库,本实施对此不做限定。动画形象可以是一个3D网格。
以3DMM库为例,3DMM库由大量高精度脸部数据通过主成分分析方法得到,描述了高维脸型和表情相对平均脸的主要变化,也可以描述纹理信息。
一般来说,3DMM库描述一个无表情的脸型时,可以通过mu+∑(Pface i-mu)*α i得到。其中,mu是自然表情下的平均脸,Pface i是第i个脸型主成分分量,α i就是各个脸型主成分分量的权重,也就是捏脸参数。
假设3DMM库中的动画形象对应的网格可以通过M表示,即通过M表示3DMM库中的脸型、表情和顶点之间的关系,M是一个[m×n×d]的三维矩阵,其中每一维分别为网格的顶点坐标(m)、脸型主成分(n)、表情主成分(d)。3DMM库M的各个维度分布和意义如图5所示,每个坐标轴分别表示顶点坐标(m)、脸型主成分(n)、表情主成分(d)。由于m表示xyz三个坐标的值,所以网格的顶点数为m/3,记作v。如果确定了动画形象的脸型或者表情,那么M可以是一个二维矩阵。
在本申请实施例中,不考虑3DMM库中的纹理维度,假设动画形象的驱动为F,则:
Figure PCTCN2020111615-appb-000001
其中,M为动画形象的网格,α为捏脸参数,β为表情参数;n为捏脸基中捏脸网格的个数,d为表情基中表情网格的个数,M k,j,i为具有第i个表情网格、第j个捏脸网格的第k个网格,α j为一组捏脸参数中的第j维,表示第j个脸型主成分分量的权重,β i为一组表情参数中的第i维,表示第i个表情主成分分量的权重。
其中,确定捏脸参数的过程为捏脸算法,确定表情参数的过程为捏表情算法。捏脸参数用于与捏脸基做线性组合得到对应的脸型,例如存在一个包括50个捏脸网格(属于可变形网格,例如blendshape)的捏脸基,该捏脸基对应的捏脸参数为一个50维的向量,每一维可以标识该捏脸参数所对应脸型与一个捏脸网格的相关程度。捏脸基所包括的捏脸网格分别代表不同脸型,每一个捏脸网格均为相对平均脸变化较大的脸部形象,是大量的脸通过PCA分解之后的得到的不同维度的脸型主成分,且同一个捏脸基中不同捏脸网格对应的顶点序号保持一致。
表情参数用于与表情基做线性组合得到对应的表情,例如存在一个包括50个(相当于维数为50)表情网格(属于可变形网格,例如blendshape)的表情基,该表情基对应的表情参数为一个50维的向量,每一维可以标识该表情参数所对应表情与一个表情网格的相关程度。表情基所包括的表情网格分别代表不同表情,每一个表情网格均由同一个3D模型在不同表情下变化而成,同一个表情基中不同表情网格对应的顶点序号保持一致。
针对前述的可变形网格,单个网格可以通过预定义形状变形,得到任意数量网格。
结合上述公式(1),可以得到说话人所对应第一动画形象的第一表情基,从而用于后续第二动画形象的驱动。
S203、根据目标文本信息、该媒体数据和第一表情基确定对应目标文本信息的声学特征和目标表情参数。
通过图3中的声学特征和表情参数确定模块可以确定对应目标文本信息的声学特征和目标表情参数。其中,声学特征用于标识模拟说话人说出目标文本信息的声音,目标表情参数用于标识模拟说话人说出目标文本信息的脸部表情相对于第一表情基的变化程度。
可以理解的是,目标文本信息的获取方式可以包括多种,例如,目标文本信息可以是用户通过终端设备输入的,也可以是根据输入至终端设备的语音转换得到的。
S204、根据声学特征和目标表情参数,驱动具有第二表情基的第二动画形象。
通过图3中的动画驱动模块将目标表情参数所标识的表情,配合声学特征所标识的语音,通过人类能直观理解的方式,利用多种感官来展示。其中一种可行的方式是,假设目标表情参数表示了第二表情基中各个表情网格的权重,通过第二表情基加权线性组合可以得到对应的表情。在发出语音的同时,通过渲染方法将做出与该语音对应表情的第二动画形象渲染出来,从而实现第二动画形象的驱动。
由上述技术方案可以看出,通过采集说话人说出语音时脸部表情变化的视频,可以确定说话人所对应第一动画形象的第一表情基,第一表情基可以体现第一动画形象的不同表情。在确定出用于驱动第二动画形象的目标文本信息后,可以根据目标文本信息、前述采集的媒体数据和第一表情基,确定对应目标文本信息的声学特征和目标表情参数,该声学特征可以标识模拟所述说话人说出所述目标文本信息的声音,该目标表情参数可以标识模拟所述说话人说出所述目标文本信息的脸部表情相对于所述第一表情基的变化程度。通过声学特征和目标表情参数,可以驱动具有第二表情基的第二动画形象,使得第二动画形象可以通过声学特征模拟发出说话人说出目标文本信息的声音,并且在发声过程中做出符合该说话人应有表情的脸部表情,给用户带来逼真的代入感和沉浸感,提高了用户与动画形象进行交互的体验。
需要说明的是,S203的实现方式可以包括多种,本申请实施例着重对一种实现方式进行介绍。
在一种可能的实现方式中,S203的实现方式可以是根据目标文本信息和媒体数据,确定对应目标文本信息的声学特征和表情特征。该声学特征用于标识模拟说话人说出目标文本信息的声音,该表情特征用于标识模拟说话人说出所述目标文本信息的脸部表情。然后,根据第一表情基和表情特征确定目标表情参数。
由于媒体数据中已经同步记录了说话人的脸部表情和语音,即媒体数据中说话人的脸部表情和语音对应同一时间轴。故,可以在线下预先收集大量的媒体数据作为训练数据,从这些媒体数据中提取文本特征、声学特征和表情特征,根据这些特征训练得到时长模型、声学模型、表情模型。当线上获取包含说话人的脸部表情和对应语音的媒体数据时,可以使用时长模型确定目标文本信息对应的时长,再将时长结合目标文本信息对应的文本特征 分别通过声学模型和表情模型确定对应的声学特征和表情特征。由于声学特征和表情特征都是基于同一个时长模型得到的时长,因此很容易做到语音以及表情的同步,从而使得第二动画形象在模拟说话人说出目标文本信息对应语音的同时,模拟说话人做出相应的表情。
接下来,将对S204的可能实现方式进行介绍。应理解,在本实施例中,第二动画形象可以是与第一动画形象相同的动画形象,也可以是与第一动画形象不同的动画形象。在这两种情况下,S204的实现方式可能有所不同。
第一种情况:第一动画形象和第二动画形象为同一个动画形象。
在这种情况下,所需驱动的动画形象即第一动画形象。那么,为了驱动第一动画形象,除了需要确定第一表情基,还需要确定第一动画形象的捏脸参数,得到第一动画形象的脸型。因此,在S202中,可以根据脸部表情确定第一动画形象的第一表情基和第一动画形象的捏脸参数,该捏脸参数用于标识第一动画形象的脸型相对于第一动画形象所对应捏脸基的变化程度。
确定第一动画形象的第一表情基和第一动画形象的捏脸参数的方式有很多种。在一些情况下,基于媒体数据确定捏脸参数以建立脸部模型时,所采集的媒体数据往往精度不高,噪声较大,使得建立的脸部模型质量不高,具有很多不确定性,难以准确体现待建对象的实际外形。例如,由于采集不规范导致建模质量低;重建过程容易受到环境光照、用户化妆等影响;重建的脸部模型中含有表情,并非自然状态;建立的脸部模型无法适应之后将要提取表情参数的视频等。为了解决这一问题,本申请实施例提供一种捏脸参数的确定方法,参见图6所示。
在图6中,若获取的媒体数据中可以包括多组脸部顶点数据,可以基于其中的第一顶点数据,以及3DMM库中用于标识目标脸部模型的目标顶点数据,确定初始捏脸参数。在确定出初始捏脸参数的基础上,通过获取媒体数据中的第二顶点数据,基于初始捏脸参数和目标顶点数据确定表情参数,之后,固定该表情参数,反推捏脸参数或者说反推如何变化脸型得到在该表情参数下的说话人的脸部形象,即通过固定表情反推脸型的方式修正初始捏脸参数,得到目标捏脸参数,从而将该目标捏脸参数作为第一动画形象的捏脸参数。
由于第二顶点数据和第一顶点数据分别标识待建对象的不同脸部形象,故第二顶点数据和第一顶点数据受到完全相同的不确定性影响的几率较小,在通过第一顶点数据确定出初始捏脸参数的基础上,通过第二顶点数据修正出的目标捏脸参数可以一定程度上抵消第一顶点数据中的噪声,以目标捏脸参数确定出的说话人对应的脸部模型精确度相对更高。
由于第一表情基与第二表情基相同,即二者的维数以及各个维数的语义信息相同,确定出的目标表情参数可以直接驱动第二动画形象,故,在S204中驱动第二动画形象的方式可以是根据声学特征、目标表情参数和捏脸参数,驱动具有第二表情基的第二动画形象。
第二种情况:第一动画形象和第二动画形象为不同动画形象。
在这种情况下,第一表情基与第二表情基不同,即二者的维数以及各个维数的语义信息存在不同,故难以直接利用目标表情参数驱动具有第二表情基的第二动画形象。由于第一动画形象对应的表情参数与第二动画形象对应的表情参数应具有映射关系,第一动画形象对应的表情参数与第二动画形象对应的表情参数间的映射关系可以通过函数f()表示,则通过第一动画形象对应的表情参数计算第二动画形象对应的表情参数的公式如下:
β b=f(β a)   (2)
其中,β b为第二动画形象对应的表情参数,β a为第一动画形象对应的表情参数,f()表示第一动画形象对应的表情参数与第二动画形象对应的表情参数间的映射关系。
故,若确定出该映射关系,便可以利用第一动画形象(例如动画形象a)对应的表情参数直接驱动第二动画形象(例如动画形象b)。其中,映射关系可以是线性映射关系,也可以是非线性映射关系。
为了实现根据目标表情参数驱动具有第二表情基的第二动画形象,需要确定出映射关系。确定映射关系的方式可以包括多种,本实施例主要对两种确定方式进行介绍。
第一种确定方式可以是基于第一动画形象对应的第一表情基和第二动画形象对应的第二表情基,确定表情参数间的映射关系。参见图7所示,由于第一动画形象对应的实际表情参数可以驱动被第一动画形象做出实际表情,该实际表情参数可以体现该实际表情与其表情基的不同维度下的相关程度,即第二动画形象对应的实际表情参数也可以体现第二动画形象的实际表情与其表情基的不同维度下的相关程度,故基于上述表情参数与表情基间的关联关系,可以根据第一动画形象对应的第一表情基和第二动画形象对应的第二表情基,确定出表情参数间的映射关系。然后,根据声学特征、目标表情参数和该映射关系,驱动具有第二表情基的第二动画形象。
第二种确定方式可以是基于音素和第二表情基之间的预设关系确定表情参数间的映射关系。
音素是根据语音的自然属性划分出来的最小语音单位,依据音节里的发音动作来分析,一个动作(例如口型)构成一个音素。也就是说,音素与说话人无关,无论说话人是谁、无论语音是英语还是汉语、无论发出音素所对应的文本是否相同,只要语音中一个时间区间内的音素相同,那么,对应的表情例如口型具有一致性。参见图8所示,图8示出了时间区间与音素的对应关系,描述了在一个语音中,哪个时间区间对应了哪个音素。例如,第二行中“5650000”和“6300000”代表时间戳,表示5.65秒至6.3秒这一时间区间,在该时间区间内说话人发出的音素是“u”。音素的统计方法并不唯一,本实施例以33个中文音素为例。
由于媒体数据中,面部表情和语音是同步采集的,因此可以方便的通过语音的划分,得到对应的视频帧,即根据媒体数据确定语音所标识音素、该音素对应的时间区间和媒体 数据处于该时间区间的视频帧。然后,根据该视频帧确定音素对应的第一表情参数,第一表情参数用于标识发出该音素时说话人的脸部表情相对于第一表情基的变化程度。
例如图8中第二行,对于音素“u”,其所对应的时间区间是5.65秒至6.3秒,确定处于时间区间5.65秒至6.3秒的视频帧,根据该视频帧提取音素“u”对应的第一表情参数。若第一动画形象为动画形象a,第一表情参数可以用β a表示。若第一表情基的维数是n a,则得到的第一表情参数β a为一组n a长度的向量。
由于该确定映射关系的方式的前提是其他动画形象的表情基例如第二动画形象对应的第二表情基是根据与音素的预设关系生成的,预设关系表示的是一个音素对应一个表情网格,比如对于第二动画形象b而言,预设关系中音素“u”对应第1个表情网格,音素“i”对应第2个表情网格……,若音素的个数为n b个,则根据预设关系可以确定出包括n b个表情网格的第二表情基。那么,当确定出语音所标识的音素后,便可以根据预设关系和第二表情基,确定该音素对应的第二表情参数。然后,根据第一表情参数和第二表情参数,确定映射关系。
例如,语音所标识的音素为“u”,通过第二表情基和预设关系可知音素“u”对应第1个表情网格,则可以确定出第二表情参数为β b=[1 0…0],β b中包括n b个元素,除了第一个元素为1,其余n b-1个元素均为0。
由此,一组β b和β a的映射关系就建立了。当得到大量第一表情参数β a时,可以产生大量对应的第二表情参数β b。假设第一表情参数β a和第二表情参数β b的个数分别是L个,L个第一表情参数β a构成第一矩阵,L个第二表情参数β b构成第二矩阵,分别记作β A和β B。有:
β A=[L×n a],β B=[L×n b]  (3)
本方案以第一表情参数和第二表情参数之间满足线性映射关系为例,则上述公式(2)可以变形为:
β b=f*β a  (4)
根据公式(3)和(4)所示,则映射关系的确定公式可以为:
f=β B*inv(β A)  (5)
其中,f为映射关系,β A为第一矩阵,β B为第二矩阵,inv为矩阵求逆运算。
在得到映射关系f后,对于任意一组第一表情参数β a,可以得到对应的β b=f*β a,从而根据第一表情参数得到第二表情参数,以便驱动第二动画形象,例如动画形象b。
前述实施例主要介绍了如何基于文本信息驱动动画形象。在一些情况下,还可以基于媒体数据直接驱动动画形象。例如,媒体数据中说话人所对应第一动画形象具有第一表情基,第一表情基的维数为第一维数,顶点拓扑为第一顶点拓扑,第一表情基可以用Ea表示, 第一维数可以用Na表示,第一顶点拓扑可以用Ta表示,第一表情基Ea的样子是Fa;待驱动的第二动画形象具有第二表情基,第二表情基的维数为第二维数,顶点拓扑为第二顶点拓扑,第二表情基可以用Eb表示,第二维数可以用Nb表示,第二顶点拓扑可以用Tb表示,第二表情基Eb的样子是Fb,希望通过包括该说话人脸部表情和语音的媒体数据来驱动第二动画形象。
为此,本申请实施例还提供一种基于人工智能的动画形象驱动方法,参见图9所示,所述方法包括:
S901、获取包含说话人的脸部表情和对应语音的第一媒体数据。
S902、根据脸部表情确定所述说话人所对应第一动画形象的第一表情基。
S903、根据第一表情基和待驱动的第二动画形象的第二表情基,确定目标表情基。
在本实施例中,由于第一表情基的维数与第二表情基的维数不同,为了可以利用媒体数据中该说话人的脸部表情和语音驱动第二动画形象,可以构造出一个新的表情基例如目标表情基,使得该目标表情基同时具有第一表情基和第二表情基的特点。
在一种实现方式中,S903的实现方式可以是:从第一表情基中确定第一动画形象处于无表情时对应的无表情网格,并从第二表情基中确第二动画形象处于无表情时对应的无表情网格。根据第一形象对应的无表情网格和第二形象对应的无表情网格,确定调整网格,该调整网格具有第二顶点拓扑,用于标识处于无表情时的第一动画形象。根据调整网格和第二表情基中的网格形变关系,生成目标表情基。
若第一表情基为Ea,第一维数为Na,第一顶点拓扑为Ta,第一表情基Ea的样子是Fa;第二表情基为Eb,第二维数为Nb,第二顶点拓扑为Tb,第二表情基Eb的样子是Fb,则该方法的流程图还可以参见图10a所示。基于第一表情基Ea和第二表情基Eb确定目标表情基Eb’。其中,确定目标表情基Eb’的方式可以是提取第二表情基Eb的无表情网格和第一表情基Ea的无表情网格。通过捏脸算法例如nricp算法,将Eb的无表情网格贴到Ea的无表情网格上,使得Eb的无表情网格在保持顶点拓扑Fb的前提下,改变样子,变成Ea的样子,得到调整网格,该调整网格可以表示为Newb。随后,由于Newb和第二表情基Eb中各个维度的表情相对自然表情(无表情)的网格形变关系是已知的,故,可以根据Newb和第二表情基Eb中的网格形变关系从Newb中形变出目标表情基Eb’。目标表情基Eb’的样子是Fa,维数是Nb,顶点拓扑是Tb。
S904、根据包含所述说话人的脸部表情和对应语音的第二媒体数据和所述目标表情基,确定目标表情参数和声学特征。
在得到目标表情基后,参见图10a所示,基于根据包含该说话人的脸部表情和对应语音的第二媒体数据和该目标表情基Eb’,提取出声学特征并通过捏表情算法得到目标表情参数Bb。其中,目标表情参数用于标识说话人说出所述语音的脸部表情相对于目标表情基的变化程度。
可以理解的是,利用该方法得到的目标表情参数和声学特征可以用于重新训练前述所提到的声学模型、表情模型。
S905、根据目标表情参数和声学特征,驱动具有第二表情基的所述第二动画形象。
S901、S902和S905的具体实现方式分别可以参见前述S201、S202和S204的实现方式,此处不再赘述。
接下来,将结合实际应用场景对本申请实施例提供的基于人工智能的动画形象驱动方法进行介绍。
在该应用场景中,在该应用场景中,假设第一动画形象为仿照说话人的形象构建的,第二动画形象为在游戏中与用户进行交互的游戏角色的形象。当该游戏角色通过输入的目标文本信息与用户进行交流时,希望通过该目标文本信息驱动该游戏角色模仿说话人发出目标文本信息对应的语音,并做出对应的表情。故,终端设备可以获取包含说话人的脸部表情和对应语音的媒体数据,根据脸部表情确定该说话人所对应第一动画形象的第一表情基。接着,根据目标文本信息、媒体数据和第一表情基,确定对应所述目标文本信息的声学特征和目标表情参数,从而根据该声学特征和目标表情参数,驱动具有第二表情基的第二动画形象,使得第二动画形象发出目标文本信息对应的语音,并做出对应的表情。这样,用户可以看到该游戏角色模仿说话人说出语音,并做出对应的表情,为用户带来逼真的代入感和沉浸感,提高了用户与动画形象进行交互的体验。
基于前述实施例提供的方法,本实施例还提供一种动画形象驱动装置1000,所述装置1000部署在音视频处理设备上。参见图10b,所述装置1000包括获取单元1001、第一确定单元1002、第二确定单元1003和驱动单元1004:
所述获取单元1001,用于获取包含说话人的脸部表情和对应语音的媒体数据;
所述第一确定单元1002,用于根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;
所述第二确定单元1003,用于根据目标文本信息、所述媒体数据和所述第一表情基,确定对应所述目标文本信息的声学特征和目标表情参数;所述声学特征用于标识模拟所述说话人说出所述目标文本信息的声音,所述目标表情参数用于标识模拟所述说话人说出所述目标文本信息的脸部表情相对于所述第一表情基的变化程度;
所述驱动单元1004,用于根据所述声学特征和所述目标表情参数,驱动具有第二表情基的第二动画形象。
在一种可能的实现方式中,所述第一动画形象和所述第二动画形象为同一个动画形象,所述第一表情基与所述第二表情基相同,所述第一确定单元1002,用于:
根据所述脸部表情确定所述第一动画形象的第一表情基和所述第一动画形象的捏脸参数,所述捏脸参数用于标识所述第一动画形象的脸型相对于所述第一动画形象所对应捏脸基的变化程度;
所述驱动单元1004,用于:
根据所述声学特征、所述目标表情参数和所述捏脸参数,驱动所述第二动画形象。
在一种可能的实现方式中,所述第一动画形象和所述第二动画形象为不同动画形象,所述第一表情基与所述第二表情基不同,所述驱动单元1004,用于:
确定所述第一表情基所对应表情参数与所述第二表情基所对应表情参数间的映射关系;
根据所述声学特征、所述目标表情参数和所述映射关系,驱动所述第二动画形象。
在一种可能的实现方式中,所述第二表情基是根据所述第二表情基与音素的预设关系生成的,所述驱动单元1004,还用于:
根据所述媒体数据确定所述语音所标识音素、所述音素对应的时间区间和所述媒体数据处于所述时间区间的视频帧;
根据所述视频帧确定所述音素对应的第一表情参数,所述第一表情参数用于标识发出所述音素时所述说话人的脸部表情相对于所述第一表情基的变化程度;
根据所述预设关系和所述第二表情基,确定所述音素对应的第二表情参数;
根据所述第一表情参数和所述第二表情参数,确定所述映射关系。
在一种可能的实现方式中,所述第二确定单元1003,用于:
根据所述目标文本信息和所述媒体数据,确定对应所述目标文本信息的声学特征和表情特征;所述表情特征用于标识模拟所述说话人说出所述目标文本信息的脸部表情;
根据所述第一表情基和所述表情特征确定所述目标表情参数。
本实施例还提供一种动画形象驱动装置1100,所述装置1100部署在音视频处理设备上。参见图11,所述装置1100包括获取单元1101、第一确定单元1102、第二确定单元1103、第三确定单元1104和驱动单元1105:
所述获取单元1101,用于获取包含说话人的脸部表情和对应语音的第一媒体数据;
所述第一确定单元1102,用于根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;所述第一表情基的维数为第一维数,顶点拓扑为第一顶点拓扑;
所述第二确定单元1103,用于根据所述第一表情基和待驱动的第二动画形象的第二表情基,确定目标表情基;所述第二表情基的维数为第二维数,顶点拓扑为第二顶点拓扑,所述目标表情基为具有第二顶点拓扑的第一动画形象对应的表情基,所述目标表情基的维数为第二维数;
所述第三确定单元1104,用于根据包含所述说话人的脸部表情和对应语音的第二媒体数据和所述目标表情基,确定目标表情参数和声学特征;所述目标表情参数用于标识所述说话人说出所述语音的脸部表情相对于所述目标表情基的变化程度;
所述驱动单元1105,用于根据所述目标表情参数和声学特征,驱动具有所述第二表情基的所述第二动画形象。
在一种可能的实现方式中,所述第二确定单元1103,用于从所述第一表情基中确定所述第一动画形象处于无表情时对应的无表情网格,并从所述第二表情基中确所述第二动画形象处于无表情时对应的无表情网格;
根据所述第一动画形象对应的无表情网格和所述第二动画形象对应的无表情网格,确定调整网格,所述调整网格具有第二顶点拓扑,用于标识处于无表情时的第一动画形象;
根据所述调整网格和所述第二表情基中的网格形变关系,生成所述目标表情基。
本申请实施例还提供了一种用于动画形象驱动的设备,该设备可以通过语音驱动动画,该设备可以为音视频处理设备。下面结合附图对该设备进行介绍。请参见图12所示,本申请实施例提供了一种用于动画形象驱动的设备,该设还可以是终端设备,该终端设备可以为包括手机、平板电脑、个人数字助理(Personal Digital Assistant,简称PDA)、销售终端(Point of Sales,简称POS)、车载电脑等任意智能终端,以终端设备为手机为例:
图12示出的是与本申请实施例提供的终端设备相关的手机的部分结构的框图。参考图12,手机包括:射频(Radio Frequency,简称RF)电路1210、存储器1220、输入单元1230、显示单元1240、传感器1250、音频电路1260、无线保真(wireless fidelity,简称WiFi)模块1270、处理器1280、以及电源1290等部件。本领域技术人员可以理解,图12中示出的手机结构并不构成对手机的限定,可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件布置。
下面结合图12对手机的各个构成部件进行具体的介绍:
RF电路1210可用于收发信息或通话过程中,信号的接收和发送,特别地,将基站的下行信息接收后,给处理器1280处理;另外,将设计上行的数据发送给基站。通常,RF电路1210包括但不限于天线、至少一个放大器、收发信机、耦合器、低噪声放大器(Low Noise Amplifier,简称LNA)、双工器等。此外,RF电路1210还可以通过无线通信与网络和其他设备通信。上述无线通信可以使用任一通信标准或协议,包括但不限于全球移动通讯系统(Global System of Mobile communication,简称GSM)、通用分组无线服务(General Packet Radio Service,简称GPRS)、码分多址(Code Division Multiple Access,简称CDMA)、 宽带码分多址(Wideband Code Division Multiple Access,简称WCDMA)、长期演进(Long Term Evolution,简称LTE)、电子邮件、短消息服务(Short Messaging Service,简称SMS)等。
存储器1220可用于存储软件程序以及模块,处理器1280通过运行存储在存储器1220的软件程序以及模块,从而执行手机的各种功能应用以及数据处理。存储器1220可主要包括存储程序区和存储数据区,其中,存储程序区可存储操作系统、至少一个功能所需的应用程序(比如声音播放功能、图像播放功能等)等;存储数据区可存储根据手机的使用所创建的数据(比如音频数据、电话本等)等。此外,存储器1220可以包括高速随机存取存储器,还可以包括非易失性存储器,例如至少一个磁盘存储器件、闪存器件、或其他易失性固态存储器件。
输入单元1230可用于接收输入的数字或字符信息,以及产生与手机的用户设置以及功能控制有关的键信号输入。具体地,输入单元1230可包括触控面板1231以及其他输入设备1232。触控面板1231,也称为触摸屏,可收集用户在其上或附近的触摸操作(比如用户使用手指、触笔等任何适合的物体或附件在触控面板1231上或在触控面板1231附近的操作),并根据预先设定的程式驱动相应的连接装置。可选的,触控面板1231可包括触摸检测装置和触摸控制器两个部分。其中,触摸检测装置检测用户的触摸方位,并检测触摸操作带来的信号,将信号传送给触摸控制器;触摸控制器从触摸检测装置上接收触摸信息,并将它转换成触点坐标,再送给处理器1280,并能接收处理器1280发来的命令并加以执行。此外,可以采用电阻式、电容式、红外线以及表面声波等多种类型实现触控面板1231。除了触控面板1231,输入单元1230还可以包括其他输入设备1232。具体地,其他输入设备1232可以包括但不限于物理键盘、功能键(比如音量控制按键、开关按键等)、轨迹球、鼠标、操作杆等中的一种或多种。
显示单元1240可用于显示由用户输入的信息或提供给用户的信息以及手机的各种菜单。显示单元1240可包括显示面板1241,可选的,可以采用液晶显示器(Liquid Crystal Display,简称LCD)、有机发光二极管(Organic Light-Emitting Diode,简称OLED)等形式来配置显示面板1241。进一步的,触控面板1231可覆盖显示面板1241,当触控面板1231检测到在其上或附近的触摸操作后,传送给处理器1280以确定触摸事件的类型,随后处理器1280根据触摸事件的类型在显示面板1241上提供相应的视觉输出。虽然在图12中,触控面板1231与显示面板1241是作为两个独立的部件来实现手机的输入和输入功能,但是在某些实施例中,可以将触控面板1231与显示面板1241集成而实现手机的输入和输出功能。
手机还可包括至少一种传感器1250,比如光传感器、运动传感器以及其他传感器。具体地,光传感器可包括环境光传感器及接近传感器,其中,环境光传感器可根据环境光线的明暗来调节显示面板1241的亮度,接近传感器可在手机移动到耳边时,关闭显示面板1241和/或背光。作为运动传感器的一种,加速计传感器可检测各个方向上(一般为三轴)加速度的大小,静止时可检测出重力的大小及方向,可用于识别手机姿态的应用(比如横 竖屏切换、相关游戏、磁力计姿态校准)、振动识别相关功能(比如计步器、敲击)等;至于手机还可配置的陀螺仪、气压计、湿度计、温度计、红外线传感器等其他传感器,在此不再赘述。
音频电路1260、扬声器1261,传声器1262可提供用户与手机之间的音频接口。音频电路1260可将接收到的音频数据转换后的电信号,传输到扬声器1261,由扬声器1261转换为声音信号输出;另一方面,传声器1262将收集的声音信号转换为电信号,由音频电路1260接收后转换为音频数据,再将音频数据输出处理器1280处理后,经RF电路1210以发送给比如另一手机,或者将音频数据输出至存储器1220以便进一步处理。
WiFi属于短距离无线传输技术,手机通过WiFi模块1270可以帮助用户收发电子邮件、浏览网页和访问流式媒体等,它为用户提供了无线的宽带互联网访问。虽然图12示出了WiFi模块1270,但是可以理解的是,其并不属于手机的必须构成,完全可以根据需要在不改变发明的本质的范围内而省略。
处理器1280是手机的控制中心,利用各种接口和线路连接整个手机的各个部分,通过运行或执行存储在存储器1220内的软件程序和/或模块,以及调用存储在存储器1220内的数据,执行手机的各种功能和处理数据,从而对手机进行整体监控。可选的,处理器1280可包括一个或多个处理单元;优选的,处理器1280可集成应用处理器和调制解调处理器,其中,应用处理器主要处理操作系统、用户界面和应用程序等,调制解调处理器主要处理无线通信。可以理解的是,上述调制解调处理器也可以不集成到处理器1280中。
手机还包括给各个部件供电的电源1290(比如电池),优选的,电源可以通过电源管理系统与处理器1280逻辑相连,从而通过电源管理系统实现管理充电、放电、以及功耗管理等功能。
尽管未示出,手机还可以包括摄像头、蓝牙模块等,在此不再赘述。
在本实施例中,该终端设备所包括的处理器1280还具有以下功能:
获取包含说话人的脸部表情和对应语音的媒体数据;
根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;
根据目标文本信息、所述媒体数据和所述第一表情基,确定对应所述目标文本信息的声学特征和目标表情参数;所述声学特征用于标识模拟所述说话人说出所述目标文本信息的声音,所述目标表情参数用于标识模拟所述说话人说出所述目标文本信息的脸部表情相对于所述第一表情基的变化程度;
根据所述声学特征和所述目标表情参数,驱动具有第二表情基的第二动画形象。
或,
获取包含说话人的脸部表情和对应语音的第一媒体数据;
根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;所述第一表情基的维数为第一维数,顶点拓扑为第一顶点拓扑;
根据所述第一表情基和待驱动的第二动画形象的第二表情基,确定目标表情基;所述第二表情基的维数为第二维数,顶点拓扑为第二顶点拓扑,所述目标表情基为具有第二顶点拓扑的第一动画形象对应的表情基,所述目标表情基的维数为第二维数;
根据包含所述说话人的脸部表情和对应语音的第二媒体数据和所述目标表情基,确定目标表情参数和声学特征;所述目标表情参数用于标识所述说话人说出所述语音的脸部表情相对于所述目标表情基的变化程度;
根据所述目标表情参数和声学特征,驱动具有所述第二表情基的所述第二动画形象。
本申请实施例还提供服务器,请参见图13所示,图13为本申请实施例提供的服务器1300的结构图,服务器1300可因配置或性能不同而产生比较大的差异,可以包括一个或一个以上中央处理器(Central Processing Units,简称CPU)1322(例如,一个或一个以上处理器)和存储器1332,一个或一个以上存储应用程序1342或数据1344的存储介质1330(例如一个或一个以上海量存储设备)。其中,存储器1332和存储介质1330可以是短暂存储或持久存储。存储在存储介质1330的程序可以包括一个或一个以上模块(图示没标出),每个模块可以包括对服务器中的一系列指令操作。更进一步地,中央处理器1322可以设置为与存储介质1330通信,在服务器1300上执行存储介质1330中的一系列指令操作。
服务器1300还可以包括一个或一个以上电源1326,一个或一个以上有线或无线网络接口1350,一个或一个以上输入输出接口1358,和/或,一个或一个以上操作系统1341,例如Windows ServerTM,Mac OS XTM,UnixTM,LinuxTM,FreeBSDTM等等。
上述实施例中由服务器所执行的步骤可以基于该图13所示的服务器结构。
本申请实施例还提供一种计算机可读存储介质,所述计算机可读存储介质用于存储程序代码,所述程序代码用于执行前述各个实施例所述的动画形象驱动方法。
本申请实施例还提供一种包括指令的计算机程序产品,当其在计算机上运行时,使得计算机执行前述各个实施例所述的动画形象驱动方法。
本申请的说明书及上述附图中的术语“第一”、“第二”、“第三”、“第四”等(如果存在)是用于区别类似的对象,而不必用于描述特定的顺序或先后次序。应该理解这样使用的数据在适当情况下可以互换,以便这里描述的本申请的实施例例如能够以除了在这里图示或描述的那些以外的顺序实施。此外,术语“包括”和“具有”以及他们的任何变形,意图在于覆盖不排他的包含,例如,包含了一系列步骤或单元的过程、方法、系统、 产品或设备不必限于清楚地列出的那些步骤或单元,而是可包括没有清楚地列出的或对于这些过程、方法、产品或设备固有的其它步骤或单元。
应当理解,在本申请中,“至少一个(项)”是指一个或者多个,“多个”是指两个或两个以上。“和/或”,用于描述关联对象的关联关系,表示可以存在三种关系,例如,“A和/或B”可以表示:只存在A,只存在B以及同时存在A和B三种情况,其中A,B可以是单数或者复数。字符“/”一般表示前后关联对象是一种“或”的关系。“以下至少一项(个)”或其类似表达,是指这些项中的任意组合,包括单项(个)或复数项(个)的任意组合。例如,a,b或c中的至少一项(个),可以表示:a,b,c,“a和b”,“a和c”,“b和c”,或“a和b和c”,其中a,b,c可以是单个,也可以是多个。
在本申请所提供的几个实施例中,应该理解到,所揭露的系统,装置和方法,可以通过其它的方式实现。例如,以上所描述的装置实施例仅仅是示意性的,例如,所述单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口,装置或单元的间接耦合或通信连接,可以是电性,机械或其它的形式。
所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本申请各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能单元的形式实现。
所述集成的单元如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的全部或部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行本申请各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(Read-Only Memory,简称ROM)、随机存取存储器(Random Access Memory,简称RAM)、磁碟或者光盘等各种可以存储程序代码的介质。
以上所述,以上实施例仅用以说明本申请的技术方案,而非对其限制;尽管参照前述实施例对本申请进行了详细的说明,本领域的普通技术人员应当理解:其依然可以对前述各实施例所记载的技术方案进行修改,或者对其中部分技术特征进行等同替换;而这些修改或者替换,并不使相应技术方案的本质脱离本申请各实施例技术方案的精神和范围。

Claims (16)

  1. 一种动画形象驱动方法,所述方法由音视频处理设备执行,所述方法包括:
    获取包含说话人的脸部表情和对应语音的媒体数据;
    根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;
    根据目标文本信息、所述媒体数据和所述第一表情基,确定对应所述目标文本信息的声学特征和目标表情参数;所述声学特征用于标识模拟所述说话人说出所述目标文本信息的声音,所述目标表情参数用于标识模拟所述说话人说出所述目标文本信息的脸部表情相对于所述第一表情基的变化程度;
    根据所述声学特征和所述目标表情参数,驱动具有第二表情基的第二动画形象。
  2. 根据权利要求1所述的方法,所述第一动画形象和所述第二动画形象为同一个动画形象,所述第一表情基与所述第二表情基相同,所述根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,包括:
    根据所述脸部表情确定所述第一动画形象的第一表情基和所述第一动画形象的捏脸参数,所述捏脸参数用于标识所述第一动画形象的脸型相对于所述第一动画形象所对应捏脸基的变化程度;
    所述根据所述声学特征和所述目标表情参数,驱动具有第二表情基的第二动画形象,包括:
    根据所述声学特征、所述目标表情参数和所述捏脸参数,驱动所述第二动画形象。
  3. 根据权利要求1所述的方法,所述第一动画形象和所述第二动画形象为不同动画形象,所述第一表情基与所述第二表情基不同,所述根据所述声学特征和所述目标表情参数,驱动具有第二表情基的第二动画形象,包括:
    确定所述第一表情基所对应表情参数与所述第二表情基所对应表情参数间的映射关系;
    根据所述声学特征、所述目标表情参数和所述映射关系,驱动所述第二动画形象。
  4. 根据权利要求3所述的方法,所述第二表情基是根据所述第二表情基与音素的预设关系生成的,所述确定所述第一表情基所对应表情参数与所述第二表情基所对应表情参数间的映射关系,包括:
    根据所述媒体数据确定所述语音所标识音素、所述音素对应的时间区间和所述媒体数据处于所述时间区间的视频帧;
    根据所述视频帧确定所述音素对应的第一表情参数,所述第一表情参数用于标识发出所述音素时所述说话人的脸部表情相对于所述第一表情基的变化程度;
    根据所述预设关系和所述第二表情基,确定所述音素对应的第二表情参数;
    根据所述第一表情参数和所述第二表情参数,确定所述映射关系。
  5. 根据权利要求1所述的方法,所述根据目标文本信息、所述媒体数据和所述第一表情基,确定对应所述目标文本信息的声学特征和目标表情参数,包括:
    根据所述目标文本信息和所述媒体数据,确定对应所述目标文本信息的声学特征和表情特征;所述表情特征用于标识模拟所述说话人说出所述目标文本信息的脸部表情;
    根据所述第一表情基和所述表情特征确定所述目标表情参数。
  6. 一种动画形象驱动装置,所述装置部署在音视频处理设备上,所述装置包括获取单元、第一确定单元、第二确定单元和驱动单元:
    所述获取单元,用于获取包含说话人的脸部表情和对应语音的媒体数据;
    所述第一确定单元,用于根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;
    所述第二确定单元,用于根据目标文本信息、所述媒体数据和所述第一表情基,确定对应所述目标文本信息的声学特征和目标表情参数;所述声学特征用于标识模拟所述说话人说出所述目标文本信息的声音,所述目标表情参数用于标识模拟所述说话人说出所述目标文本信息的脸部表情相对于所述第一表情基的变化程度;
    所述驱动单元,用于根据所述声学特征和所述目标表情参数,驱动具有第二表情基的第二动画形象。
  7. 根据权利要求6所述的装置,所述第一动画形象和所述第二动画形象为同一个动画形象,所述第一表情基与所述第二表情基相同,所述第一确定单元,用于:
    根据所述脸部表情确定所述第一动画形象的第一表情基和所述第一动画形象的捏脸参数,所述捏脸参数用于标识所述第一动画形象的脸型相对于所述第一动画形象所对应捏脸基的变化程度;
    所述驱动单元,用于:
    根据所述声学特征、所述目标表情参数和所述捏脸参数,驱动所述第二动画形象。
  8. 根据权利要求6所述的装置,所述第一动画形象和所述第二动画形象为不同动画形象,所述第一表情基与所述第二表情基不同,所述驱动单元,用于:
    确定所述第一表情基所对应表情参数与所述第二表情基所对应表情参数间的映射关系;
    根据所述声学特征、所述目标表情参数和所述映射关系,驱动所述第二动画形象。
  9. 根据权利要求8所述的装置,所述第二表情基是根据所述第二表情基与音素的预设关系生成的,所述驱动单元,还用于:
    根据所述媒体数据确定所述语音所标识音素、所述音素对应的时间区间和所述媒体数据处于所述时间区间的视频帧;
    根据所述视频帧确定所述音素对应的第一表情参数,所述第一表情参数用于标识发出所述音素时所述说话人的脸部表情相对于所述第一表情基的变化程度;
    根据所述预设关系和所述第二表情基,确定所述音素对应的第二表情参数;
    根据所述第一表情参数和所述第二表情参数,确定所述映射关系。
  10. 根据权利要求6所述的装置,所述第二确定单元,用于:
    根据所述目标文本信息和所述媒体数据,确定对应所述目标文本信息的声学特征和表情特征;所述表情特征用于标识模拟所述说话人说出所述目标文本信息的脸部表情;
    根据所述第一表情基和所述表情特征确定所述目标表情参数。
  11. 一种动画形象驱动方法,所述方法由音视频处理设备执行,所述方法包括:
    获取包含说话人的脸部表情和对应语音的第一媒体数据;
    根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;所述第一表情基的维数为第一维数,顶点拓扑为第一顶点拓扑;
    根据所述第一表情基和待驱动的第二动画形象的第二表情基,确定目标表情基;所述第二表情基的维数为第二维数,顶点拓扑为第二顶点拓扑,所述目标表情基为具有第二顶点拓扑的第一动画形象对应的表情基,所述目标表情基的维数为第二维数;
    根据包含所述说话人的脸部表情和对应语音的第二媒体数据和所述目标表情基,确定目标表情参数和声学特征;所述目标表情参数用于标识所述说话人说出所述语音的脸部表情相对于所述目标表情基的变化程度;
    根据所述目标表情参数和声学特征,驱动具有所述第二表情基的所述第二动画形象。
  12. 根据权利要求11所述的方法,所述根据所述第一表情基和待驱动的第二动画形象的第二表情基,确定目标表情基,包括:
    从所述第一表情基中确定所述第一动画形象处于无表情时对应的无表情网格,并从所述第二表情基中确所述第二动画形象处于无表情时对应的无表情网格;
    根据所述第一动画形象对应的无表情网格和所述第二动画形象对应的无表情网格,确定调整网格,所述调整网格具有第二顶点拓扑,用于标识处于无表情时的第一动画形象;
    根据所述调整网格和所述第二表情基中的网格形变关系,生成所述目标表情基。
  13. 一种动画形象驱动装置,所述装置部署在音视频处理设备上,所述装置包括获取单元、第一确定单元、第二确定单元、第三确定单元和驱动单元:
    所述获取单元,用于获取包含说话人的脸部表情和对应语音的第一媒体数据;
    所述第一确定单元,用于根据所述脸部表情确定所述说话人所对应第一动画形象的第一表情基,所述第一表情基用于标识所述第一动画形象的表情;所述第一表情基的维数为第一维数,顶点拓扑为第一顶点拓扑;
    所述第二确定单元,用于根据所述第一表情基和待驱动的第二动画形象的第二表情基,确定目标表情基;所述第二表情基的维数为第二维数,顶点拓扑为第二顶点拓扑,所述目标表情基为具有第二顶点拓扑的第一动画形象对应的表情基,所述目标表情基的维数为第二维数;
    所述第三确定单元,用于根据包含所述说话人的脸部表情和对应语音的第二媒体数据和所述目标表情基,确定目标表情参数和声学特征;所述目标表情参数用于标识所述说话人说出所述语音的脸部表情相对于所述目标表情基的变化程度;
    所述驱动单元,用于根据所述目标表情参数和声学特征,驱动具有所述第二表情基的所述第二动画形象。
  14. 一种用于动画形象驱动的设备,所述设备包括处理器以及存储器:
    所述存储器用于存储程序代码,并将所述程序代码传输给所述处理器;
    所述处理器用于根据所述程序代码中的指令执行权利要求1-5或11-12任一项所述的方法。
  15. 一种计算机可读存储介质,所述计算机可读存储介质用于存储程序代码,所述程序代码用于执行权利要求1-5或11-12任一项所述的方法。
  16. 一种计算机程序产品,当所述计算机程序产品被执行时,用于执行权利要求1-5或11-12任一项所述的方法。
PCT/CN2020/111615 2019-09-02 2020-08-27 一种基于人工智能的动画形象驱动方法和相关装置 Ceased WO2021043053A1 (zh)

Priority Applications (5)

Application Number Priority Date Filing Date Title
JP2021557135A JP7408048B2 (ja) 2019-09-02 2020-08-27 人工知能に基づくアニメキャラクター駆動方法及び関連装置
EP20860658.2A EP3929703B1 (en) 2019-09-02 2020-08-27 Animation image driving method based on artificial intelligence, and related device
KR1020217029221A KR102694330B1 (ko) 2019-09-02 2020-08-27 인공 지능에 기초한 애니메이션 이미지 구동 방법, 및 관련 디바이스
US17/405,965 US11605193B2 (en) 2019-09-02 2021-08-18 Artificial intelligence-based animation character drive method and related apparatus
US18/080,655 US12112417B2 (en) 2019-09-02 2022-12-13 Artificial intelligence-based animation character drive method and related apparatus

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201910824770.0 2019-09-02
CN201910824770.0A CN110531860B (zh) 2019-09-02 2019-09-02 一种基于人工智能的动画形象驱动方法和装置

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US17/405,965 Continuation US11605193B2 (en) 2019-09-02 2021-08-18 Artificial intelligence-based animation character drive method and related apparatus

Publications (1)

Publication Number Publication Date
WO2021043053A1 true WO2021043053A1 (zh) 2021-03-11

Family

ID=68666304

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2020/111615 Ceased WO2021043053A1 (zh) 2019-09-02 2020-08-27 一种基于人工智能的动画形象驱动方法和相关装置

Country Status (6)

Country Link
US (2) US11605193B2 (zh)
EP (1) EP3929703B1 (zh)
JP (1) JP7408048B2 (zh)
KR (1) KR102694330B1 (zh)
CN (1) CN110531860B (zh)
WO (1) WO2021043053A1 (zh)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114420088A (zh) * 2022-01-20 2022-04-29 安徽淘云科技股份有限公司 一种展示方法及其相关设备
CN115617169A (zh) * 2022-10-11 2023-01-17 深圳琪乐科技有限公司 一种语音控制机器人及基于角色关系的机器人控制方法
US12136159B2 (en) * 2022-12-31 2024-11-05 Theai, Inc. Contextually oriented behavior of artificial intelligence characters

Families Citing this family (104)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9318108B2 (en) 2010-01-18 2016-04-19 Apple Inc. Intelligent automated assistant
US8977255B2 (en) 2007-04-03 2015-03-10 Apple Inc. Method and system for operating a multi-function portable electronic device using voice-activation
US8676904B2 (en) 2008-10-02 2014-03-18 Apple Inc. Electronic devices with voice command and contextual data processing capabilities
US10706373B2 (en) 2011-06-03 2020-07-07 Apple Inc. Performing actions associated with task items that represent tasks to perform
US10276170B2 (en) 2010-01-18 2019-04-30 Apple Inc. Intelligent automated assistant
US10057736B2 (en) 2011-06-03 2018-08-21 Apple Inc. Active transport based notifications
US10417037B2 (en) 2012-05-15 2019-09-17 Apple Inc. Systems and methods for integrating third party services with a digital assistant
EP4560630A3 (en) 2013-02-07 2025-08-06 Apple Inc. Voice trigger for a digital assistant
US10652394B2 (en) 2013-03-14 2020-05-12 Apple Inc. System and method for processing voicemail
US10748529B1 (en) 2013-03-15 2020-08-18 Apple Inc. Voice activated device for use with a voice-based digital assistant
EP3008641A1 (en) 2013-06-09 2016-04-20 Apple Inc. Device, method, and graphical user interface for enabling conversation persistence across two or more instances of a digital assistant
US10176167B2 (en) 2013-06-09 2019-01-08 Apple Inc. System and method for inferring user intent from speech inputs
US10791216B2 (en) 2013-08-06 2020-09-29 Apple Inc. Auto-activating smart responses based on activities from remote devices
US10170123B2 (en) 2014-05-30 2019-01-01 Apple Inc. Intelligent assistant for home automation
US9966065B2 (en) 2014-05-30 2018-05-08 Apple Inc. Multi-command single utterance input method
US9715875B2 (en) 2014-05-30 2017-07-25 Apple Inc. Reducing the need for manual start/end-pointing and trigger phrases
US9338493B2 (en) 2014-06-30 2016-05-10 Apple Inc. Intelligent automated assistant for TV user interactions
US9886953B2 (en) 2015-03-08 2018-02-06 Apple Inc. Virtual assistant activation
US10460227B2 (en) 2015-05-15 2019-10-29 Apple Inc. Virtual assistant in a communication session
US10200824B2 (en) 2015-05-27 2019-02-05 Apple Inc. Systems and methods for proactively identifying and surfacing relevant content on a touch-sensitive device
US20160378747A1 (en) 2015-06-29 2016-12-29 Apple Inc. Virtual assistant for media playback
US10671428B2 (en) 2015-09-08 2020-06-02 Apple Inc. Distributed personal assistant
US10331312B2 (en) 2015-09-08 2019-06-25 Apple Inc. Intelligent automated assistant in a media environment
US10747498B2 (en) 2015-09-08 2020-08-18 Apple Inc. Zero latency digital assistant
US10740384B2 (en) 2015-09-08 2020-08-11 Apple Inc. Intelligent automated assistant for media search and playback
US11587559B2 (en) 2015-09-30 2023-02-21 Apple Inc. Intelligent device identification
US10691473B2 (en) 2015-11-06 2020-06-23 Apple Inc. Intelligent automated assistant in a messaging environment
US10956666B2 (en) 2015-11-09 2021-03-23 Apple Inc. Unconventional virtual assistant interactions
US10223066B2 (en) 2015-12-23 2019-03-05 Apple Inc. Proactive assistance based on dialog communication between devices
US12223282B2 (en) 2016-06-09 2025-02-11 Apple Inc. Intelligent automated assistant in a home environment
US10586535B2 (en) 2016-06-10 2020-03-10 Apple Inc. Intelligent digital assistant in a multi-tasking environment
DK201670540A1 (en) 2016-06-11 2018-01-08 Apple Inc Application integration with a digital assistant
DK179415B1 (en) 2016-06-11 2018-06-14 Apple Inc Intelligent device arbitration and control
US12197817B2 (en) 2016-06-11 2025-01-14 Apple Inc. Intelligent device arbitration and control
US11204787B2 (en) 2017-01-09 2021-12-21 Apple Inc. Application integration with a digital assistant
US10726832B2 (en) 2017-05-11 2020-07-28 Apple Inc. Maintaining privacy of personal information
DK180048B1 (en) 2017-05-11 2020-02-04 Apple Inc. MAINTAINING THE DATA PROTECTION OF PERSONAL INFORMATION
DK179496B1 (en) 2017-05-12 2019-01-15 Apple Inc. USER-SPECIFIC Acoustic Models
DK201770428A1 (en) 2017-05-12 2019-02-18 Apple Inc. LOW-LATENCY INTELLIGENT AUTOMATED ASSISTANT
DK179745B1 (en) 2017-05-12 2019-05-01 Apple Inc. SYNCHRONIZATION AND TASK DELEGATION OF A DIGITAL ASSISTANT
DK201770411A1 (en) 2017-05-15 2018-12-20 Apple Inc. Multi-modal interfaces
DK179549B1 (en) 2017-05-16 2019-02-12 Apple Inc. FAR-FIELD EXTENSION FOR DIGITAL ASSISTANT SERVICES
US10303715B2 (en) 2017-05-16 2019-05-28 Apple Inc. Intelligent automated assistant for media exploration
US20180336892A1 (en) 2017-05-16 2018-11-22 Apple Inc. Detecting a trigger of a digital assistant
US10818288B2 (en) 2018-03-26 2020-10-27 Apple Inc. Natural assistant interaction
US10928918B2 (en) 2018-05-07 2021-02-23 Apple Inc. Raise to speak
US11145294B2 (en) 2018-05-07 2021-10-12 Apple Inc. Intelligent automated assistant for delivering content from user experiences
EP3815050B1 (en) * 2018-05-24 2024-01-24 Warner Bros. Entertainment Inc. Matching mouth shape and movement in digital video to alternative audio
DK179822B1 (da) 2018-06-01 2019-07-12 Apple Inc. Voice interaction at a primary device to access call functionality of a companion device
US10892996B2 (en) 2018-06-01 2021-01-12 Apple Inc. Variable latency device coordination
DK180639B1 (en) 2018-06-01 2021-11-04 Apple Inc DISABILITY OF ATTENTION-ATTENTIVE VIRTUAL ASSISTANT
DK201870355A1 (en) 2018-06-01 2019-12-16 Apple Inc. VIRTUAL ASSISTANT OPERATION IN MULTI-DEVICE ENVIRONMENTS
US11462215B2 (en) 2018-09-28 2022-10-04 Apple Inc. Multi-modal inputs for voice commands
CN111627095B (zh) * 2019-02-28 2023-10-24 北京小米移动软件有限公司 表情生成方法及装置
US11348573B2 (en) 2019-03-18 2022-05-31 Apple Inc. Multimodality in digital assistant systems
US11307752B2 (en) 2019-05-06 2022-04-19 Apple Inc. User configurable task triggers
DK201970509A1 (en) 2019-05-06 2021-01-15 Apple Inc Spoken notifications
US11140099B2 (en) 2019-05-21 2021-10-05 Apple Inc. Providing message response suggestions
DK201970510A1 (en) 2019-05-31 2021-02-11 Apple Inc Voice identification in digital assistant systems
DK180129B1 (en) 2019-05-31 2020-06-02 Apple Inc. User activity shortcut suggestions
US11227599B2 (en) 2019-06-01 2022-01-18 Apple Inc. Methods and user interfaces for voice-based control of electronic devices
CN110531860B (zh) 2019-09-02 2020-07-24 腾讯科技(深圳)有限公司 一种基于人工智能的动画形象驱动方法和装置
CN111145777A (zh) * 2019-12-31 2020-05-12 苏州思必驰信息科技有限公司 一种虚拟形象展示方法、装置、电子设备及存储介质
US11593984B2 (en) 2020-02-07 2023-02-28 Apple Inc. Using text for avatar animation
CN111294665B (zh) * 2020-02-12 2021-07-20 百度在线网络技术(北京)有限公司 视频的生成方法、装置、电子设备及可读存储介质
CN111311712B (zh) * 2020-02-24 2023-06-16 北京百度网讯科技有限公司 视频帧处理方法和装置
CN111372113B (zh) * 2020-03-05 2021-12-21 成都威爱新经济技术研究院有限公司 基于数字人表情、嘴型及声音同步的用户跨平台交流方法
CN111736700B (zh) * 2020-06-23 2025-01-07 上海商汤临港智能科技有限公司 基于数字人的车舱交互方法、装置及车辆
KR20220004156A (ko) * 2020-03-30 2022-01-11 상하이 센스타임 린강 인텔리전트 테크놀로지 컴퍼니 리미티드 디지털 휴먼에 기반한 자동차 캐빈 인터랙션 방법, 장치 및 차량
CN111459450A (zh) * 2020-03-31 2020-07-28 北京市商汤科技开发有限公司 交互对象的驱动方法、装置、设备以及存储介质
US12301635B2 (en) 2020-05-11 2025-05-13 Apple Inc. Digital assistant hardware abstraction
US11043220B1 (en) 2020-05-11 2021-06-22 Apple Inc. Digital assistant hardware abstraction
US11061543B1 (en) 2020-05-11 2021-07-13 Apple Inc. Providing relevant data items based on context
US11755276B2 (en) 2020-05-12 2023-09-12 Apple Inc. Reducing description length based on confidence
US11490204B2 (en) 2020-07-20 2022-11-01 Apple Inc. Multi-device audio adjustment coordination
US11438683B2 (en) 2020-07-21 2022-09-06 Apple Inc. User identification using headphones
CN111988658B (zh) * 2020-08-28 2022-12-06 网易(杭州)网络有限公司 视频生成方法及装置
CN118981266A (zh) * 2020-10-14 2024-11-19 住友电气工业株式会社 计算机可读取的存储介质
US12128322B2 (en) * 2020-11-12 2024-10-29 Tencent Technology (Shenzhen) Company Limited Method and apparatus for driving vehicle in virtual environment, terminal, and storage medium
CN112527115B (zh) * 2020-12-15 2023-08-04 北京百度网讯科技有限公司 用户形象生成方法、相关装置及计算机程序产品
CN112669424B (zh) * 2020-12-24 2024-05-31 科大讯飞股份有限公司 一种表情动画生成方法、装置、设备及存储介质
CN112286366B (zh) * 2020-12-30 2022-02-22 北京百度网讯科技有限公司 用于人机交互的方法、装置、设备和介质
CN112927712B (zh) * 2021-01-25 2024-06-04 网易(杭州)网络有限公司 视频生成方法、装置和电子设备
CN115205917A (zh) * 2021-04-12 2022-10-18 上海擎感智能科技有限公司 一种人机交互的方法及电子设备
CN113066156B (zh) * 2021-04-16 2025-06-03 广州虎牙科技有限公司 表情重定向方法、装置、设备和介质
CN113256821B (zh) * 2021-06-02 2022-02-01 北京世纪好未来教育科技有限公司 一种三维虚拟形象唇形生成方法、装置及电子设备
CN115984452A (zh) * 2021-10-14 2023-04-18 聚好看科技股份有限公司 一种头部三维重建方法及设备
KR20230100205A (ko) * 2021-12-28 2023-07-05 삼성전자주식회사 영상 처리 방법 및 장치
CN114898019A (zh) * 2022-02-08 2022-08-12 武汉路特斯汽车有限公司 一种动画融合方法和装置
CN114612600B (zh) * 2022-03-11 2023-02-17 北京百度网讯科技有限公司 虚拟形象生成方法、装置、电子设备和存储介质
CN116778107B (zh) 2022-03-11 2024-10-15 腾讯科技(深圳)有限公司 表情模型的生成方法、装置、设备及介质
CN114758040A (zh) * 2022-03-24 2022-07-15 努比亚技术有限公司 一种虚拟人物生成方法、设备及计算机可读存储介质
CN114708636A (zh) * 2022-04-01 2022-07-05 成都市谛视科技有限公司 一种密集人脸网格表情驱动方法、装置及介质
CN115050067B (zh) * 2022-05-25 2024-07-02 中国科学院半导体研究所 人脸表情构建方法、装置、电子设备、存储介质及产品
CN115311394A (zh) * 2022-07-27 2022-11-08 湖南芒果无际科技有限公司 一种驱动数字人面部动画方法、系统、设备及介质
KR102838168B1 (ko) * 2022-09-02 2025-07-25 동서대학교 산학협력단 딥페이크 기술을 활용한 동화 미디어 셀프 제작방법
KR102652652B1 (ko) * 2022-11-29 2024-03-29 주식회사 일루니 아바타 생성 장치 및 방법
US20240265605A1 (en) * 2023-02-07 2024-08-08 Google Llc Generating an avatar expression
CN116188649B (zh) * 2023-04-27 2023-10-13 科大讯飞股份有限公司 基于语音的三维人脸模型驱动方法及相关装置
CN116452709A (zh) * 2023-06-13 2023-07-18 北京好心情互联网医院有限公司 动画生成方法、装置、设备及存储介质
CN116778043B (zh) * 2023-06-19 2024-02-09 广州怪力视效网络科技有限公司 一种表情捕捉及动画自动生成系统和方法
US12045639B1 (en) * 2023-08-23 2024-07-23 Bithuman Inc System providing visual assistants with artificial intelligence
CN120935387A (zh) * 2024-05-08 2025-11-11 北京字跳网络技术有限公司 生成媒体内容的方法、装置、设备和存储介质
CN118331431B (zh) * 2024-06-13 2024-08-27 海马云(天津)信息技术有限公司 虚拟数字人驱动方法与装置、电子设备及存储介质

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8725507B2 (en) * 2009-11-27 2014-05-13 Samsung Eletronica Da Amazonia Ltda. Systems and methods for synthesis of motion for animation of virtual heads/characters via voice processing in portable devices
CN104217454A (zh) * 2014-08-21 2014-12-17 中国科学院计算技术研究所 一种视频驱动的人脸动画生成方法
CN108875633A (zh) * 2018-06-19 2018-11-23 北京旷视科技有限公司 表情检测与表情驱动方法、装置和系统及存储介质
CN109447234A (zh) * 2018-11-14 2019-03-08 腾讯科技(深圳)有限公司 一种模型训练方法、合成说话表情的方法和相关装置
CN110531860A (zh) * 2019-09-02 2019-12-03 腾讯科技(深圳)有限公司 一种基于人工智能的动画形象驱动方法和装置

Family Cites Families (21)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2003141564A (ja) 2001-10-31 2003-05-16 Minolta Co Ltd アニメーション生成装置およびアニメーション生成方法
US8555164B2 (en) * 2001-11-27 2013-10-08 Ding Huang Method for customizing avatars and heightening online safety
US8224652B2 (en) * 2008-09-26 2012-07-17 Microsoft Corporation Speech and text driven HMM-based body animation synthesis
US8803889B2 (en) 2009-05-29 2014-08-12 Microsoft Corporation Systems and methods for applying animations or motions to a character
US20130257877A1 (en) * 2012-03-30 2013-10-03 Videx, Inc. Systems and Methods for Generating an Interactive Avatar Model
WO2015145219A1 (en) * 2014-03-28 2015-10-01 Navaratnam Ratnakumar Systems for remote service of customers using virtual and physical mannequins
JP2015210739A (ja) 2014-04-28 2015-11-24 株式会社コロプラ キャラクタ画像生成方法及びキャラクタ画像生成プログラム
CN107004287B (zh) * 2014-11-05 2020-10-23 英特尔公司 化身视频装置和方法
US9911218B2 (en) * 2015-12-01 2018-03-06 Disney Enterprises, Inc. Systems and methods for speech animation using visemes with phonetic boundary context
CN105551071B (zh) * 2015-12-02 2018-08-10 中国科学院计算技术研究所 一种文本语音驱动的人脸动画生成方法及系统
US10528801B2 (en) * 2016-12-07 2020-01-07 Keyterra LLC Method and system for incorporating contextual and emotional visualization into electronic communications
WO2018175892A1 (en) * 2017-03-23 2018-09-27 D&M Holdings, Inc. System providing expressive and emotive text-to-speech
US10586368B2 (en) * 2017-10-26 2020-03-10 Snap Inc. Joint audio-video facial animation system
US10657695B2 (en) * 2017-10-30 2020-05-19 Snap Inc. Animated chat presence
KR20190078015A (ko) * 2017-12-26 2019-07-04 주식회사 글로브포인트 3d 아바타를 이용한 게시판 관리 서버 및 방법
CN108763190B (zh) * 2018-04-12 2019-04-02 平安科技(深圳)有限公司 基于语音的口型动画合成装置、方法及可读存储介质
CN109377540B (zh) * 2018-09-30 2023-12-19 网易(杭州)网络有限公司 面部动画的合成方法、装置、存储介质、处理器及终端
CN109961496B (zh) * 2019-02-22 2022-10-28 厦门美图宜肤科技有限公司 表情驱动方法及表情驱动装置
US11202131B2 (en) * 2019-03-10 2021-12-14 Vidubly Ltd Maintaining original volume changes of a character in revoiced media stream
CN110288682B (zh) * 2019-06-28 2023-09-26 北京百度网讯科技有限公司 用于控制三维虚拟人像口型变化的方法和装置
US10949715B1 (en) * 2019-08-19 2021-03-16 Neon Evolution Inc. Methods and systems for image and voice processing

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8725507B2 (en) * 2009-11-27 2014-05-13 Samsung Eletronica Da Amazonia Ltda. Systems and methods for synthesis of motion for animation of virtual heads/characters via voice processing in portable devices
CN104217454A (zh) * 2014-08-21 2014-12-17 中国科学院计算技术研究所 一种视频驱动的人脸动画生成方法
CN108875633A (zh) * 2018-06-19 2018-11-23 北京旷视科技有限公司 表情检测与表情驱动方法、装置和系统及存储介质
CN109447234A (zh) * 2018-11-14 2019-03-08 腾讯科技(深圳)有限公司 一种模型训练方法、合成说话表情的方法和相关装置
CN110531860A (zh) * 2019-09-02 2019-12-03 腾讯科技(深圳)有限公司 一种基于人工智能的动画形象驱动方法和装置

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
CAO, CHEN ET AL.: "FaceWarehouse: A 3D Facial Expression Database for Visual Computing", IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, vol. 20, no. 3, 31 March 2014 (2014-03-31), XP011543570, DOI: 20201116102735A *
See also references of EP3929703A4 *

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114420088A (zh) * 2022-01-20 2022-04-29 安徽淘云科技股份有限公司 一种展示方法及其相关设备
CN115617169A (zh) * 2022-10-11 2023-01-17 深圳琪乐科技有限公司 一种语音控制机器人及基于角色关系的机器人控制方法
US12136159B2 (en) * 2022-12-31 2024-11-05 Theai, Inc. Contextually oriented behavior of artificial intelligence characters

Also Published As

Publication number Publication date
US11605193B2 (en) 2023-03-14
CN110531860B (zh) 2020-07-24
EP3929703B1 (en) 2025-04-30
US12112417B2 (en) 2024-10-08
KR20210123399A (ko) 2021-10-13
KR102694330B1 (ko) 2024-08-13
EP3929703A4 (en) 2022-10-05
US20230123433A1 (en) 2023-04-20
JP7408048B2 (ja) 2024-01-05
EP3929703A1 (en) 2021-12-29
JP2022527155A (ja) 2022-05-31
US20210383586A1 (en) 2021-12-09
CN110531860A (zh) 2019-12-03

Similar Documents

Publication Publication Date Title
CN110531860B (zh) 一种基于人工智能的动画形象驱动方法和装置
CN112379812B (zh) 仿真3d数字人交互方法、装置、电子设备及存储介质
CN115909015B (zh) 一种可形变神经辐射场网络的构建方法和装置
US11858118B2 (en) Robot, server, and human-machine interaction method
CN110288077B (zh) 一种基于人工智能的合成说话表情的方法和相关装置
WO2021036644A1 (zh) 一种基于人工智能的语音驱动动画方法和装置
CN110517340B (zh) 一种基于人工智能的脸部模型确定方法和装置
CN110286756A (zh) 视频处理方法、装置、系统、终端设备及存储介质
WO2023246163A1 (zh) 一种虚拟数字人驱动方法、装置、设备和介质
CN113421547A (zh) 一种语音处理方法及相关设备
WO2021098338A1 (zh) 一种模型训练的方法、媒体信息合成的方法及相关装置
CN110517339A (zh) 一种基于人工智能的动画形象驱动方法和装置
CN117370605A (zh) 一种虚拟数字人驱动方法、装置、设备和介质
KR20190126906A (ko) 돌봄 로봇을 위한 데이터 처리 방법 및 장치
CN116229311B (zh) 视频处理方法、装置及存储介质
CN112767520A (zh) 数字人生成方法、装置、电子设备及存储介质
CN109343695A (zh) 基于虚拟人行为标准的交互方法及系统
WO2025209111A1 (zh) 视频生成模型的训练方法、装置、设备、存储介质和产品
WO2025232361A1 (zh) 一种视频生成方法和相关装置
CN109558853B (zh) 一种音频合成方法及终端设备
CN118250523A (zh) 数字人视频生成方法、装置、存储介质及电子设备
CN116248811A (zh) 视频处理方法、装置及存储介质
CN115526772A (zh) 视频处理方法、装置、设备和存储介质
KR20260013532A (ko) 생성형 인공지능을 이용한 3차원 캐릭터의 발화 얼굴 영상 생성 방법
CN116074577A (zh) 视频处理方法、相关装置及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20860658

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 20217029221

Country of ref document: KR

Kind code of ref document: A

ENP Entry into the national phase

Ref document number: 2021557135

Country of ref document: JP

Kind code of ref document: A

Ref document number: 2020860658

Country of ref document: EP

Effective date: 20210922

NENP Non-entry into the national phase

Ref country code: DE

WWG Wipo information: grant in national office

Ref document number: 2020860658

Country of ref document: EP