JPH0473698A - Shape control method based on audio signal - Google Patents
Shape control method based on audio signalInfo
- Publication number
- JPH0473698A JPH0473698A JP2185557A JP18555790A JPH0473698A JP H0473698 A JPH0473698 A JP H0473698A JP 2185557 A JP2185557 A JP 2185557A JP 18555790 A JP18555790 A JP 18555790A JP H0473698 A JPH0473698 A JP H0473698A
- Authority
- JP
- Japan
- Prior art keywords
- lips
- formant
- audio signal
- tongue
- center frequency
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
Abstract
Description
【発明の詳細な説明】
(産業上の利用分野)
本発明は、音声信号に基づいて映像或いは人形等の顔の
顎と口唇の形状を制御する音声信号に基づく形状制御方
法に関するものである。DETAILED DESCRIPTION OF THE INVENTION (Field of Industrial Application) The present invention relates to a shape control method based on an audio signal, which controls the shape of the chin and lips of a face of an image or a doll based on the audio signal.
本発明は、入力音声信号のホルマント中心周波数を線形
変換と非線形変換して下顎と口唇の開大度を得るように
したことにより、顎と口唇の形状をよりリアルに制御す
ることができる音声信号に基づく形状制御方法を提供す
るものである。The present invention provides an audio signal that allows the shapes of the jaw and lips to be controlled more realistically by linearly and non-linearly transforming the formant center frequency of the input audio signal to obtain the degree of opening of the lower jaw and lips. The present invention provides a shape control method based on the following.
〔従来の技術]
従来の例えばいわゆるアニメーションにおいて、そのア
ニメーション中の人物が会話等を行う際の口唇及び顎等
の動きは、当該アニメーション画像作成者が、該会話に
合わせた口唇等の動きを例えば従来の経験に照らし合わ
せて推測することで決めるようにしている。また、例え
ば大型のいわゆるロボット或いは人形等の口唇及び顎等
を会話に合わせて動かす場合も同様であった。[Prior Art] In conventional animations, for example, the movements of the lips, jaws, etc. when a person in the animation performs a conversation are made by the creator of the animation image, for example, by adjusting the movements of the lips, etc., to match the conversation. I try to make decisions by making guesses based on past experience. The same applies to cases in which the lips, jaws, etc. of large so-called robots or dolls are moved in accordance with conversation.
ところで、近年、上記アニメーション或いはロボット、
人形等においては、会話に合わせて、よりリアルに口唇
及び顎等を動かすことができるようになることが求めら
れている。By the way, in recent years, the above animations or robots,
There is a need for dolls and the like to be able to move their lips, jaws, etc. more realistically in accordance with conversation.
しかし、上述したように、従来は会話に合わせた口唇等
の動きを経験等に基づいて推測するようにしているため
、到底リアルな動きとは言い難いものとなっている。ま
た、例えばコンピュータ等を用いて口唇等の動き演算す
るものも考えられているが、膨大な演算量が必要で、簡
単に、よりリアルな口唇等の動きを得ることはできない
のが実情である。更に、従来は、リアルタイムで口唇等
を動かすこともできない。However, as described above, in the past, movements of the lips, etc. in accordance with conversation have been estimated based on experience, so it is difficult to say that the movements are realistic. In addition, for example, methods have been considered to calculate the movements of the lips, etc. using computers, etc., but the reality is that it requires a huge amount of calculation and it is not possible to easily obtain more realistic movements of the lips, etc. . Furthermore, conventionally, it is not possible to move the lips or the like in real time.
そこで、本発明は、上述のような実情に鑑みて提案され
たものであり、映像、或いは、人形等の顔の顎と口唇の
形状を、よりリアルに制御することができ、更にリアル
タイムでの制御も可能な音声信号に基づく形状制御方法
を提供することを目的とするものである。Therefore, the present invention was proposed in view of the above-mentioned circumstances, and it is possible to more realistically control the shape of the jaw and lips of the face of an image or a doll, and furthermore, it can be controlled in real time. It is an object of the present invention to provide a shape control method based on an audio signal that can also be controlled.
本発明の音声信号に基づく形状制御方法は、上述の目的
を達成するために提案されたものであって、入力音声信
号から、当該人力音声信号のスペクトルエンベロープの
ピークを示すホルマント間波数の中心周波数を求め、こ
のホルマント中心周波数を線形変換及び非線形変換する
ことにより、例えば映像3人形(ロボノ日等の顔の形状
の下顎の開大度と口唇の横方向の開大度を得るようにし
たものである。すなわち、ホルマント中心周波数を線形
変換することで顎の動きと舌の動きを求め、この舌の動
きを非線形変換するこ出で下請の開大度と口唇の横方向
の開大度を得るよう二ニしでいる。The shape control method based on an audio signal of the present invention has been proposed to achieve the above-mentioned object, and is capable of determining the center frequency of the interformant wave number indicating the peak of the spectral envelope of the human input audio signal from the input audio signal. By performing linear and non-linear transformation on this formant center frequency, for example, the degree of opening of the lower jaw and the degree of opening of the lips in the lateral direction of the face shape of three video dolls (such as Robono Hi) are obtained. In other words, the movement of the jaw and the tongue are obtained by linearly transforming the formant center frequency, and the degree of opening of the subcontractor and the degree of lateral opening of the lips are determined by nonlinearly transforming this tongue movement. I'm trying to get what I want.
本発明によ7′1.ば、実際の音声に基づく入力音声信
号のホルマント中心周波数に、簡単な線形変換演算を施
し、更に、簡単な非線形変換演算を行って、下顎と口唇
の開大度を求めるようにしているため、簡単にリアルタ
イムで口唇と下麗の動きを再現できるようになる。According to the present invention 7'1. For example, a simple linear transformation operation is performed on the formant center frequency of an input audio signal based on actual speech, and a simple nonlinear transformation operation is further performed to obtain the degree of opening of the lower jaw and lips. You can easily reproduce the movements of your lips and lower jaw in real time.
:実施例]
以下、本発明を適用した実施例について図面を参暉しな
がら説明する。Embodiments] Hereinafter, embodiments to which the present invention is applied will be described with reference to the drawings.
第1図に本発明実施例の音声信号に基づく形状制御方法
が適用される例えばアニメーションの顔画像を示す。FIG. 1 shows, for example, an animated face image to which the shape control method based on an audio signal according to an embodiment of the present invention is applied.
この第1図に示す本実施例の顔画像においては、入力音
声信号から、例えば第2図〜第4図に示すような入力音
声信号のスペクトルエンベロープのピークを示すホルマ
ント(例えば第1.第2ホルマントH,、Hl)周波数
の中心周波数を求め、このホルマント中心周波数を線形
変換及び非線形変換することにより、顔の下顎の開大度
D (cm)と口唇の横方向の開大度L (cm)を得
るようにしたものである。すなわち、本実施例では、上
記ホルマント中心周波数を、後述する(1)弐を用いて
線形変換することで、顎の動きすなわち上記下顎の開大
度りと、第8図に示す舌の各動作位買P1〜P、での動
き(第6図)及び/又は第7図に示す舌の先端形状の動
きとを求め、この舌の動きを後述する(2)式を用いて
非線形変換することで上記口唇の横方向の開大度りを得
るようにしている。In the face image of this embodiment shown in FIG. 1, formants (for example, formants 1 and 2) indicating the peaks of the spectral envelope of the input audio signal as shown in FIGS. By finding the center frequency of the formant H,, Hl) frequency and performing linear and nonlinear transformation on this formant center frequency, we can calculate the lower jaw opening D (cm) and the lateral lip opening L (cm). ). That is, in this example, by linearly converting the formant center frequency using (1) 2, which will be described later, the movement of the jaw, that is, the degree of opening of the lower jaw, and each movement of the tongue shown in FIG. The movements at positions P1 to P (Fig. 6) and/or the movements of the tip shape of the tongue shown in Fig. 7 are determined, and the movements of the tongue are nonlinearly transformed using equation (2) described later. This is to obtain the degree of lateral opening of the lips.
ここで、第2図には例えば「ア2の音を発音した場合の
音声信号のスペクトルエンヘローブを示し、第3図には
例えば1゛イ」の音声信号のスペクトルエンベロープを
、第4図には例えば「つ・のスペクトルエンベロープを
示している。これら第2図〜第4図に示すスペクトルエ
ンヘローブのピーク部分を通常ホルマントと呼び、この
音声信号のホルマントは、−船ムこ、声道の音響的イン
パルス応答の減衰正弦波成分と定義されるものである。Here, FIG. 2 shows the spectral envelope of the audio signal when the sound "A2" is produced, for example, FIG. For example, the spectral envelope of ``Tsu'' is shown.The peak part of the spectral enchelobe shown in Figures 2 to 4 is usually called a formant, and the formant of this audio signal is It is defined as the damped sinusoidal component of the acoustic impulse response of a road.
このホルマントは、長さが約17cmの平均的声道に対
しては、一般に3kHz以内に3〜4個のホルマントが
あり、5kHz以内では4〜5個のホルマントがある。There are typically 3-4 formants within 3 kHz and 4-5 formants within 5 kHz for an average vocal tract of approximately 17 cm in length.
有声音では最初の3個のホルマントが最も重要であり、
一般に、周波数の最も低い所に現れるピークを第1ホル
マント(Hl)と呼び、この第1ホルマントの次に現れ
るピークを第2ホルマント(Hl)と、以後、第3ホル
マント、第4ホルマント、・・・と続いている。これら
ホルマントは、例えばいわゆるケプストル分析或いは線
形予測分析に基づいて求めることができるものであり、
例えば、該線形予測分析を用いることによって、少ない
演算量で求めることができる。In voiced sounds, the first three formants are the most important;
Generally, the peak that appears at the lowest frequency is called the first formant (Hl), and the peak that appears after this first formant is called the second formant (Hl), and henceforth, the third formant, the fourth formant, etc.・It continues. These formants can be obtained based on, for example, so-called cepstral analysis or linear predictive analysis.
For example, by using the linear predictive analysis, it can be determined with a small amount of calculation.
上述のようにして例えば各母音「ア2.′イ。As mentioned above, for example, each vowel ``A2.'I.
「つ」、′工1.・−オ、のホルマント周波数を求める
。第5圀に咳各母音のホルマント周波数の例えば第1ホ
ルマントH1と第2ホルマントH2の中心周波数の位置
を平面にプロットした時の位置関係を示す。この第5図
において、各母音の位置関係は、175と1つ、の間に
「オづが位置し、「イ」と「ア」の間に1ニーが位置す
るような位置関係となっていることが確認できる。``tsu'', 'technique 1. - Find the formant frequency of -o. The fifth field shows the positional relationship when the center frequencies of the first formant H1 and the second formant H2 of the formant frequencies of each cough vowel are plotted on a plane. In this Figure 5, the positional relationship of each vowel is such that ``ozu'' is located between 175 and 1, and 1 knee is located between ``i'' and ``a.'' I can confirm that there is.
ところで、音声と舌の動きとは、例えば、第6図、第7
図のような関係を有していることが知られている。該第
6図には各母音「ア4.1イ」1つJ+ r工、・、
[オ」に対応する第8図に示す舌の各動作位置P1〜P
、での曲率関数(舌形状曲率) C(crrr’)を示
し、第7図には各母音に対応する舌の先端形状を示して
いる。ここで、第8図において、各動作位置P1〜P、
は、舌の表側の中心線上の位置であって、上記動作位置
Pは舌の先端から例えば10mmの位置であり、動作位
置P2は上記動作位置P、から更に51奥の位置で、以
下動作位置P x、 P 4. P sの順に5mmず
つ奥の位置を示している。すなわち、該第8圀に基づき
、第6図には、上記各動作位置P1〜P、における各母
音の発声時の、これら各動作位置P〜P、上の5mmの
範囲における舌形状の曲率関数(舌形状曲率)Cを示し
、第7図には、各母音に対してこの第6図のような舌形
状曲率Cを、舌の先端の形状に変換したものを示してい
る。By the way, speech and tongue movements are, for example, shown in Figures 6 and 7.
It is known that there is a relationship as shown in the figure. In Figure 6, there is one vowel for each vowel ``a4.1i'' J + r ,...
Each operating position P1 to P of the tongue shown in FIG. 8 corresponding to [O]
, the curvature function (tongue shape curvature) C(crrr') is shown, and FIG. 7 shows the tip shape of the tongue corresponding to each vowel. Here, in FIG. 8, each operating position P1 to P,
is a position on the center line on the front side of the tongue, the operating position P is, for example, 10 mm from the tip of the tongue, and the operating position P2 is a position further 51 mm from the above operating position P, hereinafter referred to as operating position. P x, P 4. In the order of P s , the positions are shown in 5 mm increments. That is, based on the eighth area, FIG. 6 shows the curvature function of the tongue shape in a range of 5 mm above each of the operating positions P1 to P when each vowel is uttered at each of the operating positions P1 to P. (tongue shape curvature) C, and FIG. 7 shows the tongue shape curvature C as shown in FIG. 6 converted into the shape of the tip of the tongue for each vowel.
また、各母音における各動作位置P、−P、での舌形状
曲率Cと、顔の下顎の開大度りとは、第9図〜第13図
に示すような関係となっている。Further, the tongue shape curvature C at each action position P, -P for each vowel and the degree of opening of the lower jaw of the face have a relationship as shown in FIGS. 9 to 13.
すなわち、第9図は各母音発声時の顔の下顎の開大度り
と上記動作位置P1での各母音発声時の舌形状曲率Cと
の関係を示し、第10図は各母音発声時の顔の下顎の開
大度りと上記動作位置P2での各母音発声時の舌形状曲
率Cとの関係を、以下同様に、第11図は下顎開大度り
と動作位置P3、第12図は下顎開犬度りと動作位置P
4、第13図は下顎開大度りと動作位置P5の舌形状曲
率Cとの関係を示している。これら第9回〜第13図か
ら、下顎開大度りと舌形状曲率Cとの関係は、「オヨは
1ア:、と:ウーとの間↓こ位置し、工は「ア」とコイ
」の間に位置していることがわかる。これらの各母音に
おける位置関係は、上記第9図〜第13図で全て共通し
ていることが確認できる。That is, FIG. 9 shows the relationship between the degree of opening of the lower jaw of the face when each vowel is uttered and the tongue shape curvature C when each vowel is uttered at the above-mentioned operating position P1, and FIG. Similarly, FIG. 11 shows the relationship between the degree of opening of the mandible of the face and the tongue shape curvature C when each vowel is uttered at the above operating position P2. The degree of mandibular opening and the operating position P
4. FIG. 13 shows the relationship between the degree of mandibular opening and the tongue shape curvature C at the operating position P5. From these figures 9 to 13, the relationship between the mandibular opening degree and the tongue shape curvature C is as follows. It can be seen that it is located between ``. It can be confirmed that the positional relationships among these vowels are all the same in FIGS. 9 to 13 above.
上述の第9回〜第13図と前述の第5図とから、上記下
顎開大度り及び舌形状曲率Cと、上記第1第2ホルマン
)H,、H2における各母音の位置関係が、上述同様に
、−オーが−アーと−ウ、との間に位置し、−工、がニ
ア−とコイ の間に位置するような関係を存しでいるこ
とが確認できる。From the above-mentioned Figures 9 to 13 and the above-mentioned Figure 5, the positional relationship between the mandibular opening degree and tongue shape curvature C and each vowel in the first and second Holmans) H, H2 is as follows. Similarly to the above, it can be confirmed that -o is located between -a and -u, and -work is located between near and carp.
すなわち、ホルマント周波数と、舌形状曲率C及び下顎
開大度りの位置関係とは、略一致していると確認できる
。That is, it can be confirmed that the formant frequency, the positional relationship between the tongue shape curvature C and the mandibular opening degree substantially match.
このようなことから、上記第5図に示したような第1.
第2ホルマントH1,Hzのホルマント周波数から、第
9図〜第13図に示したような上記舌形状曲率C及び下
顎開大度りを近似的に写像する様な関数を比較的容易に
導くことができるようになる。For this reason, as shown in FIG.
From the formant frequency of the second formant H1, Hz, it is relatively easy to derive a function that approximately maps the tongue shape curvature C and the degree of mandibular opening as shown in FIGS. 9 to 13. You will be able to do this.
本実施例では、当該近似的に写像する関数を線形として
いる。この場合、その線形変換は、で表すことができる
。ただし、咳(1)式中、F。In this embodiment, the approximate mapping function is linear. In this case, the linear transformation can be expressed as. However, in formula (1), F.
は第1ホルマントH4のホルマント周波数(Hz)であ
り、F2は第2ホルマントH2のホルマント周波数(H
z)である。また、
である。is the formant frequency (Hz) of the first formant H4, and F2 is the formant frequency (Hz) of the second formant H2.
z). Also, .
この時、(1)式中、A及びBは、例えば第14図のよ
うなホルマント周波数の「ア」、「イ」「つ」のそれぞ
れの位置を示す点rx、P工、qえを、第16図に示す
ような舌形状曲率C及び下顎開大度りでの「アシ、「イ
ー1.「つ」のそれぞれの点rll+ Py+ Q
yに線形変換するようにして求められる。At this time, in formula (1), A and B are the points rx, P, and q that indicate the respective positions of "a", "i", and "tsu" of the formant frequency as shown in FIG. 14, for example, At the tongue shape curvature C and the degree of mandibular opening as shown in Fig. 16, each point rll+ Py+ Q
It is obtained by performing a linear transformation to y.
なお、上記ホルマント周波数は、発声する人によって個
人差があるため、この個人差を正規化によって取り除く
。この正規化としては、例えば、この第14図の「ア」
、「イコ、「つ」を頂点とする三角形を、第15図のよ
うな正三角形に変換することにより行う、これにより、
各母音の正規化が可能となる。この正三角形への変換に
ついては後述する。Note that since the formant frequency has individual differences depending on the person speaking, this individual difference is removed by normalization. As for this normalization, for example, "A" in Fig.
, by converting the triangle whose vertices are ``ico'' and ``tsu'' into an equilateral triangle as shown in Figure 15.
It becomes possible to normalize each vowel. This conversion to an equilateral triangle will be described later.
更に、本発明実施例では、上記舌形状曲率Cから口唇の
横方向の開大度りへの変換を行うようにしている。この
時の変換は、非線形変換を用いることでなされる。当該
非線形変換としては、例えば、
L = (C+d +)”” = d t
(2)を用いる。ただし、(2)式中、d 、、 d
2は定数である。この非線形変換により口の動きの自
然性を高めることができる。Furthermore, in the embodiment of the present invention, the tongue shape curvature C is converted to the degree of lateral opening of the lips. The transformation at this time is performed using nonlinear transformation. As the nonlinear transformation, for example, L = (C + d +)"" = d t
Use (2). However, in formula (2), d , d
2 is a constant. This nonlinear transformation can enhance the naturalness of mouth movements.
本実施例においては、上述したようなホルマント周波数
の線形変換による舌形状曲率C8下顎開大度りへの変換
、及び、該舌形状曲率Cの非線形変換による口唇の横方
向の開大度りへの変換の操作を行うことで、音声信号か
ら下顎の開大度り及び口唇の横方向の開大度りを推定す
ることができるようになる。したがって、本実施例の形
状制御方法を用いれば、アニメーション等の映像に限ら
ず、人形等の顔の顎と口唇の形状をよりリアルに制御す
ることができるようになる。更に、発声の個人差を正規
化することで取り除いているため、より正確な形状制御
が可能となる。In this example, the tongue shape curvature C8 is converted to the mandibular opening degree by linear transformation of the formant frequency as described above, and the tongue shape curvature C is converted to the lateral opening degree of the lips by nonlinear transformation. By performing the conversion operation, it becomes possible to estimate the degree of opening of the lower jaw and the degree of lateral opening of the lips from the audio signal. Therefore, by using the shape control method of this embodiment, it is possible to more realistically control the shape of the chin and lips of the face of a doll or the like, not just videos such as animations. Furthermore, since individual differences in vocalization are removed by normalization, more accurate shape control is possible.
ここで、上記正規化に用いられる任意の三角形を正三角
形に変換或いは逆変換する手法について説明する。すな
わち、第17図に示すように、χY平面内の三角形pq
rを考え、以下の手順によって、第23図に示すような
一辺の長さが1の正三角形p(“) qi61 r
(61、、変換する。先ず、第17図の三角形pqrに
おいて点p(χ+、 )’ 1)が原点に移るように平
行移動する(第18図)。Here, a method of converting an arbitrary triangle used in the above normalization into an equilateral triangle or inversely converting it will be described. That is, as shown in FIG. 17, the triangle pq in the χY plane
Considering r, and using the following procedure, we can create an equilateral triangle p(“) qi61 r with side length 1 as shown in Figure 23.
(61, Convert. First, in the triangle pqr of FIG. 17, point p(χ+, )' 1) is translated in parallel so that it moves to the origin (FIG. 18).
第18図の三角形p(1)q(1)r(1)の点q(1
)のX座標(xi (1))、及び、点r(1)のX座
標(y3(1))が1となるように、X、X座標をスケ
ール変換する(第19図)、当該第19図の点q′tl
のX座標(y、 (Z))がOとなるように角度θだけ
三角形pftl q121 rItT を回転させ
る(第20図)。当該20図の三角形pf31 9fi
lr(3) の点q(3) のX座標(x2(ff+
)及び、点r′3〉 のX座標(y、(ff+)が1と
なるように、x、X座標をスケール変換する(第21圀
)。該21図の三角形p(419(41r(41の点r
(4) のX座標(x、 +41 )とX = 0.5
との差をaとし、直線Y = X / aを利用して点
rf41のX座標をスケール変換する(第22図)。該
第22図の三角形pf5) q(51rTS)のX座
標が3””/2となるようにX座標のスケール変換を行
う(第23図)。Point q(1) of triangle p(1)q(1)r(1) in FIG.
), and the X coordinate (y3(1)) of point r(1) becomes 1 (Figure 19). Point q′tl in Figure 19
The triangle pftl q121 rItT is rotated by the angle θ so that the X coordinate (y, (Z)) of is O (FIG. 20). Triangle pf31 9fi in Figure 20
X coordinate of point q(3) of lr(3) (x2(ff+
) and the x, X coordinates are scale-transformed so that the point r
(4) X coordinate (x, +41) and X = 0.5
The difference between the two points is set as a, and the X coordinate of the point rf41 is scale-transformed using the straight line Y=X/a (FIG. 22). The scale of the X coordinate is converted so that the X coordinate of the triangle pf5)q(51rTS) in FIG. 22 becomes 3""/2 (FIG. 23).
・以上の手順tこより任意の三角形pqrは正三角形に
変換できる。また、この千1117!を逆にたどること
により逆変換も可能である。・From the above procedure t, any triangle pqr can be converted into an equilateral triangle. Also, this 11117! Inverse transformation is also possible by tracing backwards.
本発明の音声信号に基づく形状制御方法においては、入
力音声信号のホルマント中心周波数を線形変換と非線形
変換して下顎と口唇の開大度を得るようにしたことによ
り、例えばアニメーション等の映像、或いは、人形、ロ
ボット等の顔の顎と口唇の形状を簡単で、よりリアルに
制御可能とし、更に、リアルタイムでも制御することが
可能となった。In the shape control method based on the audio signal of the present invention, the formant center frequency of the input audio signal is linearly and non-linearly transformed to obtain the degree of opening of the lower jaw and lips. It has become possible to easily and more realistically control the shape of the jaws and lips of the faces of dolls, robots, etc., and it has also become possible to control them in real time.
第1図は本発明実施例の顔画像を示す図、第2図〜第4
図は音声信号のスペクトルエンベロープを示す特性図、
第5図は音声信号のホルマント周波数を説明するための
図、第6図は舌形状曲率を示す図、第7図は舌の先端の
形状を示す図、第8図は動作位置を示す図、第9図〜第
13図は舌の各動作位置での舌形状曲率と下7量大度を
説明するための図、第14図〜第16区はホルマント周
波数から舌形状曲率、下顎開大度への変換を説明するた
めの図、第17図〜第23圓は任意の三角形から正三角
形への変換方法を説明するためのメである。
D・・・・・・・下顎の開大度FIG. 1 is a diagram showing a face image according to an embodiment of the present invention, and FIGS.
The figure is a characteristic diagram showing the spectral envelope of the audio signal.
Fig. 5 is a diagram for explaining the formant frequency of the audio signal, Fig. 6 is a diagram showing the tongue shape curvature, Fig. 7 is a diagram showing the shape of the tip of the tongue, and Fig. 8 is a diagram showing the operating position. Figures 9 to 13 are diagrams for explaining the tongue shape curvature and lower jaw opening degree at each operating position of the tongue. Figures 17 to 23 are diagrams for explaining the conversion from an arbitrary triangle to an equilateral triangle. D・・・・・・Degree of mandibular opening
Claims (1)
ベロープのピークを示すホルマント周波数の中心周波数
を求め、このホルマント中心周波数を線形変換及び非線
形変換することにより、顔の形状の下顎の開大度と口唇
の横方向の開大度を得るようにしたことを特徴とする音
声信号に基づく形状制御方法。From the input audio signal, find the center frequency of the formant frequency that indicates the peak of the spectral envelope of the input audio signal, and linearly and non-linearly transform this formant center frequency to determine the opening degree of the lower jaw and the lips of the face shape. A shape control method based on an audio signal, characterized in that the degree of expansion in the lateral direction is obtained.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2185557A JP3070073B2 (en) | 1990-07-13 | 1990-07-13 | Shape control method based on audio signal |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2185557A JP3070073B2 (en) | 1990-07-13 | 1990-07-13 | Shape control method based on audio signal |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPH0473698A true JPH0473698A (en) | 1992-03-09 |
| JP3070073B2 JP3070073B2 (en) | 2000-07-24 |
Family
ID=16172894
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP2185557A Expired - Fee Related JP3070073B2 (en) | 1990-07-13 | 1990-07-13 | Shape control method based on audio signal |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JP3070073B2 (en) |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2009180958A (en) * | 2008-01-31 | 2009-08-13 | Yamaha Corp | Parameter setting device, sound generating device and program |
| CN101894566A (en) * | 2010-07-23 | 2010-11-24 | 北京理工大学 | A Visualization Method of Mandarin Chinese Compound Finals Based on Formant Frequency |
| JP2012173389A (en) * | 2011-02-18 | 2012-09-10 | Advanced Telecommunication Research Institute International | Lip action parameter generation device and computer program |
| CN116749174A (en) * | 2023-05-18 | 2023-09-15 | 清华大学 | Method and system for controlling actions of webcast lecture assistant robot based on voice content |
| JP2025517080A (en) * | 2022-04-25 | 2025-06-03 | グレース ジョンウン シン | AI-based disease diagnosis method and device using voice data |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2644789B2 (en) | 1987-12-18 | 1997-08-25 | 富士通株式会社 | Image transmission method |
| JP2667455B2 (en) | 1988-07-27 | 1997-10-27 | 富士通株式会社 | Facial video synthesis system |
| JP2518683B2 (en) | 1989-03-08 | 1996-07-24 | 国際電信電話株式会社 | Image combining method and apparatus thereof |
-
1990
- 1990-07-13 JP JP2185557A patent/JP3070073B2/en not_active Expired - Fee Related
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2009180958A (en) * | 2008-01-31 | 2009-08-13 | Yamaha Corp | Parameter setting device, sound generating device and program |
| CN101894566A (en) * | 2010-07-23 | 2010-11-24 | 北京理工大学 | A Visualization Method of Mandarin Chinese Compound Finals Based on Formant Frequency |
| JP2012173389A (en) * | 2011-02-18 | 2012-09-10 | Advanced Telecommunication Research Institute International | Lip action parameter generation device and computer program |
| JP2025517080A (en) * | 2022-04-25 | 2025-06-03 | グレース ジョンウン シン | AI-based disease diagnosis method and device using voice data |
| CN116749174A (en) * | 2023-05-18 | 2023-09-15 | 清华大学 | Method and system for controlling actions of webcast lecture assistant robot based on voice content |
Also Published As
| Publication number | Publication date |
|---|---|
| JP3070073B2 (en) | 2000-07-24 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Waters et al. | DECface: An automatic lip-synchronization algorithm for synthetic faces | |
| US20020024519A1 (en) | System and method for producing three-dimensional moving picture authoring tool supporting synthesis of motion, facial expression, lip synchronizing and lip synchronized voice of three-dimensional character | |
| US9431027B2 (en) | Synchronized gesture and speech production for humanoid robots using random numbers | |
| De Martino et al. | Facial animation based on context-dependent visemes | |
| Waters et al. | An automatic lip-synchronization algorithm for synthetic faces | |
| JPH02234285A (en) | Method and device for synthesizing picture | |
| Chu et al. | Corrtalk: Correlation between hierarchical speech and facial activity variances for 3d animation | |
| CN119782828A (en) | AIGC-based multi-mode digital person generation method, system and storage medium | |
| Petersen et al. | Musical-based interaction system for the Waseda Flutist Robot: Implementation of the visual tracking interaction module | |
| JP2974655B1 (en) | Animation system | |
| Pausch et al. | Tailor: creating custom user interfaces based on gesture | |
| Beskow | Talking heads-communication, articulation and animation | |
| Waters et al. | DECface: A system for synthetic face applications | |
| JP3070073B2 (en) | Shape control method based on audio signal | |
| Carré et al. | Vowel–vowel trajectories and region modeling | |
| CN113160366A (en) | 3D face animation synthesis method and system | |
| Lin et al. | A face robot for autonomous simplified musical notation reading and singing | |
| JP3298076B2 (en) | Image creation device | |
| JP2012150278A (en) | Automatic generation system of sound effect accommodated to visual change in virtual space | |
| Dabbaghchian et al. | Synthesis of VV utterances from muscle activation to sound with a 3D model | |
| JPS63184875A (en) | Sound/image conversion device | |
| JP4459415B2 (en) | Image processing apparatus, image processing method, and computer-readable information storage medium | |
| US6408274B2 (en) | Method and apparatus for synchronizing a computer-animated model with an audio wave output | |
| Williamson et al. | Audio feedback for gesture recognition | |
| Mochida et al. | Control system for talking robot to replicate articulatory movement of natural speech. |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| FPAY | Renewal fee payment (event date is renewal date of database) |
Free format text: PAYMENT UNTIL: 20090526 Year of fee payment: 9 |
|
| LAPS | Cancellation because of no payment of annual fees |