JPH04293097A - Speaker identification device - Google Patents

Speaker identification device

Info

Publication number
JPH04293097A
JPH04293097A JP3058961A JP5896191A JPH04293097A JP H04293097 A JPH04293097 A JP H04293097A JP 3058961 A JP3058961 A JP 3058961A JP 5896191 A JP5896191 A JP 5896191A JP H04293097 A JPH04293097 A JP H04293097A
Authority
JP
Japan
Prior art keywords
speaker
codebook
quantization
speaker identification
vector
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP3058961A
Other languages
Japanese (ja)
Inventor
Toshio Akaha
俊夫 赤羽
Satoru Nakamura
哲 中村
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sharp Corp
Original Assignee
Sharp Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sharp Corp filed Critical Sharp Corp
Priority to JP3058961A priority Critical patent/JPH04293097A/en
Publication of JPH04293097A publication Critical patent/JPH04293097A/en
Pending legal-status Critical Current

Links

Abstract

PURPOSE:To offer the speaker identification device which identifies a speaker corresponding to an inputted voice by vector quantization. CONSTITUTION:This speaker identification device is equipped with code books 12 for plural speakers to be recognized, a vector quantization part 11 which performs the vector quantization of voice parameters of respective frames of the input voice according to the code books 12, a code book selection part 13 which selects a code book minimizing the quantization distortion of the quantized voice parameter among the code books 12, and a decision part 15 which decides that the recognized speaker corresponding to the code book selected most among the code books 12 in a specific time is the speaker of the input voice.

Description

【発明の詳細な説明】[Detailed description of the invention]

【0001】0001

【産業上の利用分野】本発明は入力された音声により話
者を識別する話者識別装置に関する。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a speaker identification device for identifying a speaker based on input voice.

【0002】0002

【従来の技術】第1の従来の話者識別方法としては、ベ
クトル量子化を用いたテキスト独立型の話者識別方法が
知られており、各フレームの音声パラメータをベクトル
量子化して量子化歪みを時間方向に積分した値の最も小
さくなるものを認識結果とする話者識別方法がロ−ゼン
ベルグ (Rosenberg)、ス−ン(Soong
) らにより研究されている(エフ.ケイ.ス−ン他、
共著、「話者認識に対するベクトル量子化方法」,IC
ASSP論文集,米国電気電子学会,頁387−390
 , 1985 (F.K.Soong, et al
., ”A Vector Quantaizatio
n Approach to Speaker Rec
ognition”, Proc. ICASSP, 
IEEE, pp.387−390, 1985) )
[Prior Art] As a first conventional speaker identification method, a text-independent speaker identification method using vector quantization is known. A speaker identification method that uses the smallest integrated value in the time direction as the recognition result is the one proposed by Rosenberg and Soong.
) et al. (F.K. Soon et al.
Co-author, “Vector quantization method for speaker recognition”, IC
ASSP Proceedings, Institute of Electrical and Electronics Engineers, pp. 387-390
, 1985 (F.K. Soong, et al.
.. , ”A Vector Quantaizatio
n Approach to Speaker Rec
ignition”, Proc. ICASSP,
IEEE, pp. 387-390, 1985)
.

【0003】第2の従来のベクトル量子化を用いたテキ
スト独立型の話者識別方法としては、ロ−ゼンベルグ 
(Rosenberg)、ス−ン(Soong) らに
より研究されており(エフ.ケイ.ス−ン他、共著、「
話者認識における瞬時及び過渡スペクトル情報の使用に
関して」,ICASSP´86論文集,米国電気電子学
会,頁877−880 ,1986 (F.K.Soo
ng, et al.,”On the Use of
 Instantaneous and Transi
tional Spectral Informati
on in Speaker Recognition
”, Proc. ICASSP´86, TOKYO
, IEEE,pp.877−880, 1986)、
各フレームの音声パラメータに加えて音声パラメータの
時間微分を求めて、それぞれをベクトル量子化し、量子
化歪の重み付き和を時間方向に積分した値の最も小さく
なるものを認識結果とする話者識別方法が知られている
A second conventional text-independent speaker identification method using vector quantization is Rosenberg's
(Rosenberg), Soong et al. (Co-authored by F.K. Soong et al.,
"On the Use of Instantaneous and Transient Spectral Information in Speaker Recognition", Proceedings of ICASSP'86, Institute of Electrical and Electronics Engineers, pp. 877-880, 1986 (F.K. Soo
ng, et al. ,”On the Use of
Instant and Transi
tional Spectral Information
On-in Speaker Recognition
”, Proc. ICASSP´86, TOKYO
, IEEE, pp. 877-880, 1986),
In addition to the audio parameters of each frame, time differentials of the audio parameters are calculated, vector quantized, and the recognition result is the smallest value obtained by integrating the weighted sum of quantization distortion in the time direction.Speaker identification method is known.

【0004】また、第3の従来のベクトル量子化を用い
たテキスト独立型の話者識別方法としては、各フレーム
の音声パラメータをベクトル量子化した後に、セグメン
ト量子化して、量子化歪みを時間方向に積分した値の最
も小さくなるものを認識結果とする話者識別方法が、杉
山により研究されている(杉山雅英、”音声セグメント
を用いるテキスト独立話者認識”,日本音響学会講演論
文集,昭和63年3月,pp.75−76)。
[0004] A third conventional text-independent speaker identification method using vector quantization is to vector quantize the audio parameters of each frame and then perform segment quantization to remove quantization distortion in the temporal direction. Sugiyama has researched a speaker identification method in which the recognition result is the smallest integrated value of March 1963, pp. 75-76).

【0005】[0005]

【本発明が解決しようとする課題】しかしながら、上述
した従来の第1のベクトル量子化を用いたテキスト独立
型の話者識別方法では、ベクトル量子化による量子化歪
みの大きさは、話者の違い以外にも発声された音素によ
って異なるため、量子化歪を時間方向に積分した場合に
その分布が話者同士で重なって、話者識別率が低下して
いまうという問題点がある。
[Problems to be Solved by the Invention] However, in the conventional text-independent speaker identification method using the first vector quantization described above, the magnitude of quantization distortion due to vector quantization is In addition to the difference, since it differs depending on the phoneme uttered, there is a problem that when the quantization distortion is integrated in the time direction, the distribution overlaps between speakers, reducing the speaker identification rate.

【0006】また、上述した従来の第2及び第3のベク
トル量子化を用いたテキスト独立型の話者識別方法では
、音声の動的特徴を捕えようとするときに、パターンの
持つバリエーションが爆発的に大きくなることによって
コードブックを作成するための音声データが不足して、
話者識別率が低下してしまうという問題点がある。
[0006] Furthermore, in the conventional text-independent speaker identification method using the second and third vector quantization described above, variations in patterns explode when attempting to capture dynamic features of speech. As the data size increases, there is a shortage of audio data to create a codebook.
There is a problem that the speaker identification rate decreases.

【0007】本発明は、上記従来のベクトル量子化を用
いたテキスト独立型の話者識別方法の問題点に鑑み、ベ
クトル量子化による量子化歪の分布が話者同士で重なら
ず、また、音声の動的特徴を捕えるときにコードブック
を作成するための十分な音声データを供給できるように
構成された高い話者識別率を有する話者識別装置を提供
する。
In view of the above problems of the conventional text-independent speaker identification method using vector quantization, the present invention provides that the distribution of quantization distortion due to vector quantization does not overlap between speakers, and A speaker identification device having a high speaker identification rate is provided, which is configured to supply sufficient speech data for creating a codebook when capturing dynamic features of speech.

【0008】[0008]

【課題を解決するための手段】本発明は、複数の認識対
象話者のコ−ドブックと、複数のコ−ドブックに基づい
て入力音声の各フレームの音声パラメータをベクトル量
子化する量子化手段と、量子化された音声パラメ−タの
量子化歪を最小にするコードブックを複数のコードブッ
クから選択する選択手段と、所定の時間内に複数のコー
ドブックから最も多く選択されたコードブックの認識対
象話者を入力音声の話者と判定する判定手段とを備えて
いる話者識別装置によって達成される。
[Means for Solving the Problems] The present invention provides codebooks of a plurality of speakers to be recognized, and quantization means for vector quantizing the speech parameters of each frame of input speech based on the plurality of codebooks. , selection means for selecting a codebook from a plurality of codebooks that minimizes quantization distortion of quantized speech parameters, and recognition of the codebook most frequently selected from the plurality of codebooks within a predetermined time. This is achieved by a speaker identification device including a determination means for determining the target speaker as the speaker of the input voice.

【0009】[0009]

【作用】本発明の話者識別装置によれば、量子化手段は
複数の認識対象話者のコ−ドブックに基づいて入力音声
の各フレームの音声パラメータをベクトル量子化し、選
択手段は量子化された音声パラメ−タの量子化歪を最小
にするコードブックを複数のコードブックから選択し、
判定手段は複数のコードブックのうち所定の時間内に最
も多く選ばれたコードブックに基づいて話者を判定して
最も多く選ばれたコードブックの認識対象話者を入力音
声の話者とする。
[Operation] According to the speaker identification device of the present invention, the quantization means vector quantizes the speech parameters of each frame of the input speech based on the codebooks of a plurality of speakers to be recognized, and the selection means vector quantizes the speech parameters of each frame of the input speech. A codebook is selected from multiple codebooks that minimizes the quantization distortion of the voice parameters.
The determining means determines the speaker based on the codebook selected most frequently within a predetermined time from among the plurality of codebooks, and determines the speaker to be recognized in the codebook selected most frequently as the speaker of the input speech. .

【0010】本発明の話者識別装置は、学習ベクトル量
子化を用いて複数の認識対象話者のコードブックを修正
する修正手段を備えてもよい。
The speaker identification device of the present invention may include a modification means for modifying the codebooks of a plurality of speakers to be recognized using learning vector quantization.

【0011】[0011]

【実施例】以下、図面を参照して本発明の話者識別装置
における実施例を詳述する。
DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, embodiments of a speaker identification device according to the present invention will be described in detail with reference to the drawings.

【0012】図1は、本発明の話者識別装置における一
実施例の構成を示す。
FIG. 1 shows the configuration of an embodiment of a speaker identification device according to the present invention.

【0013】図1の話者識別装置は、音声分析部10、
音声分析部10に接続されたベクトル量子化部11、ベ
クトル量子化部11に接続されたコ−ドブック12、各
ベクトル量子化部11に接続されたコ−ドブック選択部
13、コ−ドブック選択部13に接続された複数のフレ
−ムカウンタ部14及び各フレ−ムカウンタ部14に接
続された判定部15により構成されている。
The speaker identification device shown in FIG. 1 includes a speech analysis section 10,
A vector quantization unit 11 connected to the speech analysis unit 10, a codebook 12 connected to the vector quantization unit 11, a codebook selection unit 13 connected to each vector quantization unit 11, and a codebook selection unit 13 and a determining section 15 connected to each frame counter section 14.

【0014】入力音声は、音響分析部10によりフレー
ム毎に音声パラメータに変換される。音声の個人性を捉
える音声パラメータとしては、スペクトル分析、声道特
性分析または声帯特性のうちいずれか1つでよい。
The input speech is converted into speech parameters frame by frame by the acoustic analysis section 10. The voice parameter that captures the individuality of voice may be any one of spectrum analysis, vocal tract characteristic analysis, and vocal cord characteristics.

【0015】また、スペクトル分析、声道特性分析、声
帯特性を組み合せることにより、更に良い音声パラメー
タを構築し得る。
Furthermore, by combining spectrum analysis, vocal tract characteristic analysis, and vocal fold characteristics, even better speech parameters can be constructed.

【0016】次に、図1に示す話者識別装置による話者
識別時の動作を説明する。
Next, the operation of speaker identification by the speaker identification device shown in FIG. 1 will be explained.

【0017】話者識別時には、発声音素による量子化歪
の大きさの違いによる量子化歪の分布の重なりの影響を
軽減するために、量子化の直後に量子化歪によるコード
ブック選択の工程を設けてフレーム毎にコードブック1
2を選択して、量子化歪の総和を取る代りに選択された
フレーム数をカウントする。
When identifying a speaker, in order to reduce the influence of overlapping distributions of quantization distortion due to differences in the magnitude of quantization distortion depending on the uttered phoneme, a codebook selection process based on quantization distortion is performed immediately after quantization. 1 codebook for each frame.
2 is selected, and the number of selected frames is counted instead of calculating the sum of quantization distortion.

【0018】抽出された音声パラメータを、ベクトル量
子化部11でコードブック12に基づいて量子化し、そ
の時の歪を式(1)に従って算出する。
The extracted audio parameters are quantized by the vector quantizer 11 based on the codebook 12, and the distortion at that time is calculated according to equation (1).

【0019】話者Sのコードブック12に含まれるi番
めのコードをCs(i)とする。
Let the i-th code included in the codebook 12 of speaker S be Cs(i).

【0020】入力音声パラメータ時系列をX,フレーム
tでのベクトルをX(t)とする。フレームtでのベク
トルX(t)をコードブックCsで量子化した時の歪を
d(X(t),Cs)とすると、 で表わされる。
Let the input audio parameter time series be X and the vector at frame t be X(t). Letting d(X(t), Cs) be the distortion when the vector X(t) at frame t is quantized using the codebook Cs, it is expressed as follows.

【0021】 )の中で、ベクトルF(x)との差を最小にするG(y
)を与えるyの値を示す。
), G(y
).

【0022】次に、算出した歪をコードブック選択部1
3で比較して、式(2)に従ってコードブックを選択す
る。
Next, the calculated distortion is sent to the codebook selection section 1.
3 and select a codebook according to equation (2).

【0023】コードブック選択部13は、ベクトル量子
化装置で計算した各話者に対する入力X(t)の量子化
歪d(X(t),Cs)を受け取り、歪の最小になるコ
ードブックの話者s(t)を選ぶ 選択したコードブック12に対して、フレームカウント
部14で式(3)に従ってフレーム数をカウントする。
The codebook selection unit 13 receives the quantization distortion d(X(t), Cs) of the input X(t) for each speaker calculated by the vector quantizer, and selects the codebook that minimizes the distortion. For the selected codebook 12 in which speaker s(t) is selected, the frame count unit 14 counts the number of frames according to equation (3).

【0024】フレームカウント部14は、各フレームt
毎にs(t)を受け取り、各コードブックが選ばれたフ
レーム数N(X,s)をカウントする。
The frame counting section 14 counts each frame t.
s(t) for each codebook, and count the number of frames N(X,s) in which each codebook was selected.

【0025】   N(X,s)=N(X,s)+1  (s=s(X
(t))のとき)  N(X,s)=N(X,s)  
    (s≠s(X(t))のとき)…(3)カウン
トされたフレーム数を判定部15で比較して、式(4)
に従って識別結果を出力する。
N(X,s)=N(X,s)+1 (s=s(X
(t))) N(X, s) = N(X, s)
(When s≠s(X(t)))...(3) The number of counted frames is compared in the determination unit 15, and the formula (4)
Output the identification results according to the following.

【0026】即ち、判定部15は所定のフレーム数の音
声入力の後半定結果r(X)を出力する。
That is, the determination unit 15 outputs the second half of the predetermined result r(X) of audio input for a predetermined number of frames.

【0027】   以下、図2を参照して図1の話者識別装置を用いて
コードブック学習時を説明する。
Hereinafter, with reference to FIG. 2, a description will be given of codebook learning using the speaker identification device of FIG. 1.

【0028】なお、図2では、図1の構成部分と共通な
部分には同じ参照番号を付してある。更に、入力音声に
対する学習デ−タベ−ス16、各コ−ドブック12に接
続された学習部17、学習デ−タベ−ス16及び学習部
17に接続された学習制御部18が付け加えられて構成
されている。
In FIG. 2, the same reference numerals are given to the parts common to those in FIG. 1. Furthermore, a learning database 16 for input speech, a learning section 17 connected to each codebook 12, and a learning control section 18 connected to the learning database 16 and the learning section 17 are added. has been done.

【0029】まず、コードブック作成時における音声デ
ータの不足に対処するために、コードブック作成後、学
習ベクトル量子化による修正を行う。
First, in order to deal with the lack of audio data when creating a codebook, correction is performed by learning vector quantization after the codebook is created.

【0030】始めのコードブックは、従来の話者識別方
法と同様にk−means法やLBGアルゴリズム等を
用いて作成される。
The initial codebook is created using the k-means method, LBG algorithm, etc., similar to conventional speaker identification methods.

【0031】次に、学習ベクトル量子化を説明する。Next, learning vector quantization will be explained.

【0032】一般に、学習ベクトル量子化は、コードブ
ックを用いた識別を行う時に、識別誤りを少なくするよ
う少しずつコードブック内の参照ベクトルを繰り返して
移動し、正確な判別境界をコードブックにより形成する
In general, in learning vector quantization, when performing classification using a codebook, reference vectors in the codebook are repeatedly moved little by little in order to reduce identification errors, and accurate discrimination boundaries are formed using the codebook. do.

【0033】まず、抽出された音声パラメータを、ベク
トル量子化部11によりコードブック12に基づいて量
子化し,その時の歪を式(5)及び式(6)により計算
する。
First, the extracted audio parameters are quantized by the vector quantizer 11 based on the codebook 12, and the distortion at that time is calculated using equations (5) and (6).

【0034】各話者s(s=1,2,…,S)のコード
ブックをBsとし、コードブックBsに含まれるi番め
のコードをBs(i)とする。また、識別対象話者r(
r=1,2,…,R)の学習用音声時系列をXr(t)
(t=1,2,…,Tr)とすると、一般にS=Rが成
り立つ。
Let Bs be the codebook of each speaker s (s=1, 2, . . . , S), and let Bs(i) be the i-th code included in the codebook Bs. Also, the speaker to be identified r(
r = 1, 2, ..., R) training audio time series as Xr(t)
If (t=1, 2, . . . , Tr), then S=R generally holds.

【0035】各話者の各フレームの音声パラメータを各
コードブックCsでベクトル量子化したときの歪みをd
(Xr(t),Cs)とし、そのときに使われたコード
ブックCs内の参照ベクトルをKs(Xr(t))とす
る。
The distortion when the speech parameters of each frame of each speaker are vector quantized using each codebook Cs is d
(Xr(t), Cs), and the reference vector in the codebook Cs used at that time is Ks(Xr(t)).

【0036】 コードブック選択部13は、式(7)、式(8)を計算
する。
The codebook selection unit 13 calculates equations (7) and (8).

【0037】歪が最小になるコードブックを2つ選んで
、最初に選ばれたコードブックCsの話者をs1、次に
選ばれたコードブックCsの話者をs2とすると、学習
制御部18は、学習データベース16から入力に与える
学習データを指定すると同時に学習部17に正解を示す
If two codebooks with minimum distortion are selected and the speaker of the first selected codebook Cs is s1 and the speaker of the second selected codebook Cs is s2, then the learning control unit 18 specifies the learning data to be inputted from the learning database 16 and at the same time indicates the correct answer to the learning section 17.

【0038】この時、   s1(Xr(t))≠r  かつ  s2(Xr(
t))=r        …(9)ならば、Cs1(
Ks1(Xr(t)))を  Xr(t)から遠ざけ   Cs2(Ks2(Xr(t)))を  Xr(t)
に近付ける      …(10)全ての認識対象話者
の学習データについて上記の手続きを行って、充分学習
が収束し、参照ベクトルの移動が所定の量より少なくな
った時に学習を終了する。
At this time, s1(Xr(t))≠r and s2(Xr(
t))=r...(9), then Cs1(
Move Ks1(Xr(t))) away from Xr(t) and move Cs2(Ks2(Xr(t))) to Xr(t)
(10) Perform the above procedure on the learning data of all speakers to be recognized, and end the learning when the learning has converged sufficiently and the movement of the reference vector becomes less than a predetermined amount.

【0039】学習部17は、式(9)で示された条件の
ときに、式(10)に示された操作を行う。学習ベクト
ル量子化には式(9)、式(10)以外にもいろいろな
変形が考えられるので、特に式(9)、式(10)に限
る必要はない。
The learning section 17 performs the operation shown in equation (10) under the condition shown in equation (9). Since various modifications other than equations (9) and (10) can be considered for learning vector quantization, there is no need to limit the equations to equations (9) and (10).

【0040】上記実施例では、フレーム毎の静的なパラ
メータベクトルに対してベクトル量子化を行っているが
、2つのフレーム以上にわたる動的なパターンに対して
も同じ方法を用いることが可能である。この場合、学習
ベクトル量子化の効果が一層大きくなる。
In the above embodiment, vector quantization is performed on a static parameter vector for each frame, but the same method can be used for dynamic patterns spanning two or more frames. . In this case, the effect of learning vector quantization becomes even greater.

【0041】[0041]

【発明の効果】複数の認識対象話者のコ−ドブックと、
複数のコ−ドブックに基づいて入力音声の各フレームの
音声パラメータをベクトル量子化する量子化手段と、量
子化された音声パラメ−タの量子化歪を最小にするコー
ドブックを複数のコードブックから選択する選択手段と
、所定の時間内に複数のコードブックから最も多く選択
されたコードブックの認識対象話者を入力音声の話者と
判定する判定手段とを備えているので、発声された音素
による量子化歪の大きさの違いの影響を軽減でき、その
結果、より正確に話者を識別することができる。
[Effect of the invention] A codebook of multiple recognition target speakers,
A quantization means vector quantizes the audio parameters of each frame of input audio based on multiple codebooks, and a codebook that minimizes quantization distortion of the quantized audio parameters from the multiple codebooks. Since it is equipped with a selection means for selecting, and a determination means for determining the recognition target speaker of the codebook selected most frequently from a plurality of codebooks within a predetermined time as the speaker of the input speech, the uttered phoneme It is possible to reduce the influence of differences in the magnitude of quantization distortion, and as a result, speakers can be identified more accurately.

【図面の簡単な説明】[Brief explanation of drawings]

【図1】本発明の話者識別装置における一実施例の構成
を示す。
FIG. 1 shows the configuration of an embodiment of a speaker identification device of the present invention.

【図2】本発明により話者識別装置のコードブック学習
時の説明図である。
FIG. 2 is an explanatory diagram during codebook learning of the speaker identification device according to the present invention.

【符号の説明】[Explanation of symbols]

10  音響分析部 11  ベクトル量子化部 12  コードブック 13  コードブック選択部 14  フレームカウント部 15  判定部 16  学習データベース 17  学習部 18  学習制御部 10 Acoustic analysis department 11 Vector quantization section 12 Codebook 13 Codebook selection section 14 Frame count section 15 Judgment section 16 Learning database 17 Learning Department 18 Learning control section

Claims (2)

【特許請求の範囲】[Claims] 【請求項1】複数の認識対象話者のコ−ドブックと、前
記複数のコ−ドブックに基づいて入力音声の各フレーム
の音声パラメータをベクトル量子化する量子化手段と、
前記量子化された音声パラメ−タの量子化歪を最小にす
るコードブックを前記複数のコードブックから選択する
選択手段と、所定の時間内に前記複数のコードブックか
ら最も多く選択されたコードブックの認識対象話者を前
記入力音声の話者と判定する判定手段とを備えているこ
とを特徴とする話者識別装置。
1. Codebooks of a plurality of speakers to be recognized, and quantization means for vector quantizing speech parameters of each frame of input speech based on the plurality of codebooks.
selection means for selecting a codebook from the plurality of codebooks that minimizes quantization distortion of the quantized speech parameters; and a codebook selected most often from the plurality of codebooks within a predetermined time. a speaker identification device, comprising determining means for determining a speaker to be recognized as the speaker of the input voice.
【請求項2】学習ベクトル量子化を用いて前記複数の認
識対象話者のコードブックを修正する修正手段を備えて
いることを特徴とする請求項1に記載の話者識別装置。
2. The speaker identification device according to claim 1, further comprising a modification means for modifying the codebooks of the plurality of speakers to be recognized using learning vector quantization.
JP3058961A 1991-03-22 1991-03-22 Speaker identification device Pending JPH04293097A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP3058961A JPH04293097A (en) 1991-03-22 1991-03-22 Speaker identification device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP3058961A JPH04293097A (en) 1991-03-22 1991-03-22 Speaker identification device

Publications (1)

Publication Number Publication Date
JPH04293097A true JPH04293097A (en) 1992-10-16

Family

ID=13099440

Family Applications (1)

Application Number Title Priority Date Filing Date
JP3058961A Pending JPH04293097A (en) 1991-03-22 1991-03-22 Speaker identification device

Country Status (1)

Country Link
JP (1) JPH04293097A (en)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS593491A (en) * 1982-06-29 1984-01-10 富士通株式会社 Voice recognition equipment
JPS59111699A (en) * 1982-12-17 1984-06-27 富士通株式会社 Speaker recognition system

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS593491A (en) * 1982-06-29 1984-01-10 富士通株式会社 Voice recognition equipment
JPS59111699A (en) * 1982-12-17 1984-06-27 富士通株式会社 Speaker recognition system

Similar Documents

Publication Publication Date Title
EP4028953B1 (en) Convolutional neural network with phonetic attention for speaker verification
Haeb-Umbach et al. Linear discriminant analysis for improved large vocabulary continuous speech recognition.
US5167004A (en) Temporal decorrelation method for robust speaker verification
US4773093A (en) Text-independent speaker recognition system and method based on acoustic segment matching
US5278942A (en) Speech coding apparatus having speaker dependent prototypes generated from nonuser reference data
Cheung et al. Feature selection via dynamic programming for text-independent speaker identification
Zhu et al. Serialized multi-layer multi-head attention for neural speaker embedding
EP0685835B1 (en) Speech recognition based on HMMs
Korshunov et al. On the use of convolutional neural networks for speech presentation attack detection
Schuller et al. Comparing one and two-stage acoustic modeling in the recognition of emotion in speech
US6502070B1 (en) Method and apparatus for normalizing channel specific speech feature elements
US6961703B1 (en) Method for speech processing involving whole-utterance modeling
KR20190135916A (en) Apparatus and method for determining user stress using speech signal
Naik et al. Evaluation of a high performance speaker verification system for access Control
Merlin et al. Non directly acoustic process for costless speaker recognition and indexation
JPH0766734A (en) Equipment and method for voice coding
JPH02232696A (en) Voice recognition device
Lohrenz et al. On temporal context information for hybrid BLSTM-based phoneme recognition
JPH0197997A (en) Voice quality conversion system
JPH07160287A (en) Standard pattern making device
KR101838947B1 (en) Voice authentication method and apparatus applicable to defective utterance
Koniaris et al. Selecting static and dynamic features using an advanced auditory model for speech recognition
Dean et al. Comparing audio and visual information for speech processing
WO2004064040A1 (en) A method for processing speech
Ma et al. Further feature extraction for speaker recognition