JPH03157698A - Speaker recognizing system - Google Patents

Speaker recognizing system

Info

Publication number
JPH03157698A
JPH03157698A JP1298503A JP29850389A JPH03157698A JP H03157698 A JPH03157698 A JP H03157698A JP 1298503 A JP1298503 A JP 1298503A JP 29850389 A JP29850389 A JP 29850389A JP H03157698 A JPH03157698 A JP H03157698A
Authority
JP
Japan
Prior art keywords
neural network
speaker
threshold
network
time
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP1298503A
Other languages
Japanese (ja)
Other versions
JP2510301B2 (en
Inventor
Kazuhiko Okashita
和彦 岡下
Shingo Nishimura
新吾 西村
Masashi Miyagawa
宮川 正志
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sekisui Chemical Co Ltd
Original Assignee
Sekisui Chemical Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sekisui Chemical Co Ltd filed Critical Sekisui Chemical Co Ltd
Priority to JP1298503A priority Critical patent/JP2510301B2/en
Publication of JPH03157698A publication Critical patent/JPH03157698A/en
Application granted granted Critical
Publication of JP2510301B2 publication Critical patent/JP2510301B2/en
Anticipated expiration legal-status Critical
Expired - Lifetime legal-status Critical Current

Links

Abstract

PURPOSE:To lessen the deterioration in a recognition rate with lapse of time and to allow real time processing with the recognizing system using a neutral network by making decision on a speaker and decision on the execution of additional learning in accordance with the thresholds for recognition of registered speakers and for additional learning. CONSTITUTION:The digital voice information from a voice processing section 12 is supplied to the neural network 13 and is stored into a memory 15. The recognition output value by this network 13 is compared with the threshold for deciding the speaker in a deciding section 14. Decision is made that the speaker is the registered speaker when the value is above the threshold. Further, the value is compared with the threshold for learning higher than this threshold. The learning is executed in accordance of the stored value of this time of the memory 15 via a network control section 16 when the value is below the threshold. The variable and coefft. of the network 13 are then changed and the network 13 is corrected to the network before the deterioration with lapse of time. The speaker recognizing system which lessens the deterioration in the recognition rate with lapse of time and allows the real time processing by using the neutral network is obtd. in this way.

Description

【発明の詳細な説明】 [産業上の利用分野] 本発明は、電子錠等において入力音声からその話者を照
合するに好適な話者認識システムに関する。
DETAILED DESCRIPTION OF THE INVENTION [Field of Industrial Application] The present invention relates to a speaker recognition system suitable for verifying a speaker from an input voice in an electronic lock or the like.

[従来の技術] 従来の話者認識システムは、例えば特公昭56−139
56に記載される如く、以下の手順による。
[Prior art] A conventional speaker recognition system is, for example, the Japanese Patent Publication No. 56-139
56, according to the following procedure.

■入力音声に含まれる話者に関する特徴量を抽出する。■Extract features related to the speaker included in the input voice.

■予め上記■と同様にして抽出しておいた標準パターン
と上記■で抽出した特徴量との距離を計算する。
(2) Calculate the distance between the standard pattern extracted in advance in the same manner as (2) above and the feature amount extracted in (2) above.

■上記■で計算した距離が、予め設定しであるしきい値
よりも小なることを条件に、今回の入力話者をその標準
パターンの登録話者であるものと判定する。
(2) On the condition that the distance calculated in (2) above is smaller than a preset threshold, the current input speaker is determined to be the registered speaker of the standard pattern.

[発明か解決しようとする課題] 然しなから、上記従来の話者認識システムでは、下記■
、■の問題点がある。
[Problem to be solved by the invention] However, in the conventional speaker recognition system described above, the following
, ■There are problems.

■標準パターン作成時から時間が経過するにつれ、認m
率が劣化する。例えば、3ケ月経過により、認識率は1
00.0%から85.0%に劣化する。
■As time passes from the time of standard pattern creation,
rate deteriorates. For example, after 3 months, the recognition rate is 1.
It deteriorates from 00.0% to 85.0%.

■実時間処理か困難である。即ち、−室以上の認識率を
確保するためには複雑な特徴量を用いる必要かあるが、
複雑な特徴量を抽出するためには複雑な処理装置が必要
であり、処理時間も多大となる。
■Real-time processing is difficult. In other words, in order to secure a recognition rate higher than -room, it is necessary to use complex features;
In order to extract complex feature quantities, a complex processing device is required, and the processing time is also large.

本発明は、経時的な認識率の劣化が極めて少なく、容易
に実時間処理できる話者認識システムを得ることを目的
とする。
SUMMARY OF THE INVENTION An object of the present invention is to provide a speaker recognition system that exhibits extremely little deterioration in recognition rate over time and that can easily perform real-time processing.

[課題を解決するための手段] 請求項1に記載の本発明は、ニューラルネットワークを
用いた話者認識システムであって、登録話者に対応する
出力ユニットの出力値に対し、登録話者認識用しきい値
と追加学習用しきい値とを設定し、上記出力値が登録話
者認識用しきい値より大なることを条件に、今回の入力
話者を登録話者と判定し、上記出力値が登録話者認識用
しきい値より大、かつ追加学習用しきい値より小なるこ
とを条件に、今回の入力音声データを用いてニューラル
ネットワークの追加学習を行なうようにしたものである
[Means for Solving the Problems] The present invention according to claim 1 is a speaker recognition system using a neural network, which performs registered speaker recognition based on an output value of an output unit corresponding to a registered speaker. The current input speaker is determined to be a registered speaker, and the above output value is determined to be a registered speaker under the condition that the above output value is greater than the registered speaker recognition threshold. The current input voice data is used to perform additional learning of the neural network on the condition that the output value is greater than the registered speaker recognition threshold and smaller than the additional learning threshold. .

請求項2に記載の本発明は、前記ニューラルネットワー
クへの入力として、 ■音声の周波数特性の時間的変化、 ■音声の平均的な線形予測係数、 ■音声の平均的なPARCOR係数、 ■音声の平均的な周波数特性、及びピッチ周波数、 ■高域強調を施された音声波形の平均的な周波数特性、
並びに ■音声の平均的な周波数特性 のうちの1つ以上を使用するようにしたものである。
The present invention according to claim 2 provides, as inputs to the neural network, (1) temporal changes in the frequency characteristics of audio, (2) average linear prediction coefficients of audio, (2) average PARCOR coefficients of audio, and (2) the average PARCOR coefficient of audio. Average frequency characteristics and pitch frequency, ■Average frequency characteristics of high-frequency emphasized audio waveforms,
and (1) one or more of the average frequency characteristics of voice is used.

請求項3に記載の本発明は、前記ニューラルネットワー
クが階層的なニューラルネットワークであるようにした
ものである。
According to a third aspect of the present invention, the neural network is a hierarchical neural network.

[作用] (1)経時的な認識率の劣化が極めて少ない。このこと
は、後述する実験結果により確認されていることである
か、ニューラルネットワークが音声の時期差による変動
の影響を受けにくい構造をとることか可能なためと推定
される。
[Effect] (1) Deterioration of recognition rate over time is extremely small. This has been confirmed by the experimental results described below, or it is presumed that the neural network can have a structure that is less susceptible to fluctuations due to differences in audio timing.

(2)ニューラルネットワークを構成する、登録話者に
対応する出力ユニットの出力値に対し、登録話者認識用
しきい値の他に、追加学習用しきい値を設けた。即ち、
上記出力値が登録話者認識用しきい値を超えて大なるも
のであり、入力話者を登録話者と判定できるものであっ
ても、該出力値が該登録話者認識用しきい値より大なる
追加学習用しきい値を超えるものでない場合には、今回
の入力音声データを用いてニューラルネットワークの追
加学習を行なう、これにより、話者の特徴が経時変化し
ても認識率が劣化する前にニューラルネットワークを更
新でき、結果として、音声の経時変化に強い話者認識シ
ステムを構成できる。
(2) In addition to the registered speaker recognition threshold, an additional learning threshold is provided for the output value of the output unit corresponding to the registered speaker, which constitutes the neural network. That is,
Even if the above output value exceeds the registered speaker recognition threshold and the input speaker can be determined to be a registered speaker, the output value exceeds the registered speaker recognition threshold. If it does not exceed a larger additional training threshold, the neural network performs additional training using the current input speech data.This prevents the recognition rate from deteriorating even if the speaker's characteristics change over time. As a result, it is possible to construct a speaker recognition system that is resistant to changes in speech over time.

(3)ニューラルネットワークは、原理的に、ネットワ
ーク全体の演算処理が単純且つ迅速である。
(3) In principle, in a neural network, the calculation processing of the entire network is simple and quick.

(4)ニューラルネットワークは、原理的に、それを構
成している各ユニットが独立に動作しており、並列的な
演算処理が可能である。従って、演算処理が迅速である
(4) In principle, each unit constituting a neural network operates independently, and parallel arithmetic processing is possible. Therefore, calculation processing is quick.

(5)上記(3)〜(4)により、話者認識システムを
複雑な処理装置によることなく容易に実時間処理できる
(5) With the above (3) and (4), the speaker recognition system can be easily processed in real time without using a complicated processing device.

又、請求項2に記載の本発明によれば上記(1)〜(5
)の作用効果に加えて、下記(6)の作用効果がある。
Further, according to the present invention as set forth in claim 2, the above (1) to (5)
), there is the following effect (6).

(6)ニューラルネットワークへの入力として、請求項
2に記載の■〜■の各要素のうちの1つ以上を用いるか
ら、入力を得るための前処理が、従来の複雑な特徴量抽
出に対して、単純となり、この前処理に要する時間が短
くて足りる。
(6) Since one or more of each of the elements (■ to ■) described in claim 2 is used as input to the neural network, the preprocessing for obtaining the input is different from conventional complex feature extraction. Therefore, the process is simple, and the time required for this preprocessing is short.

又、請求項3に記載の本発明によれば上記(1)〜(6
)の作用効果に加えて、下記(7)の作用効果かある。
Moreover, according to the present invention according to claim 3, the above (1) to (6)
In addition to the effects of ), there is also the effect of (7) below.

(7)階層的なニューラルネットワークにあっては、現
在、後述する如くの簡単な学習アルゴリズム(パックプ
ロパゲーション)が確立されており、高い認識率を実現
できるニューラルネットワークを容易に形成できる。
(7) Regarding hierarchical neural networks, a simple learning algorithm (pack propagation) as described later has been established, and a neural network that can achieve a high recognition rate can be easily formed.

[実施例] 第1図は本発明が適用された話者認識システムの一例を
示す模式図、第2図は音声処理部とニューラルネットワ
ークの一例を示す模式図、第3図は入力音声を示す模式
図、第4図はバンドパスフィルタの出力を示す模式図、
第5図はニューラルネットワークを示す模式図、第6図
は階層的なニューラルネットワークを示す模式図、第7
図はユニットの構造を示す模式図である。
[Example] Fig. 1 is a schematic diagram showing an example of a speaker recognition system to which the present invention is applied, Fig. 2 is a schematic diagram showing an example of a speech processing unit and a neural network, and Fig. 3 shows input speech. Schematic diagram, Figure 4 is a schematic diagram showing the output of the bandpass filter,
Figure 5 is a schematic diagram showing a neural network, Figure 6 is a schematic diagram showing a hierarchical neural network, and Figure 7 is a schematic diagram showing a hierarchical neural network.
The figure is a schematic diagram showing the structure of the unit.

本発明の具体的実施例の説明に先立ち、ニューラルネッ
トワークの構成、学習アルゴリズムについて説明する。
Prior to describing specific embodiments of the present invention, the configuration of the neural network and the learning algorithm will be described.

(1)ニューラルネットワークは、その構造から、第5
図(A)に示す階層的ネットワークと第5図(B)に示
す相互結合ネットワークの2種に大別できる0本発明は
、両ネットワークのいずれを用いて構成するものであっ
ても良いが、階層的ネットワークは後述する如くの簡単
な学習アルゴリズムが確立されているためより有用であ
る。
(1) Due to its structure, neural networks are
The present invention can be roughly divided into two types: the hierarchical network shown in FIG. 5(A) and the interconnected network shown in FIG. 5(B). Hierarchical networks are more useful because simple learning algorithms have been established as described below.

(2)ネットワークの構造 階層的ネットワークは、第6図に示す如く、入力層、中
間層、出力層からなる階層構造をとる。
(2) Network Structure A hierarchical network has a hierarchical structure consisting of an input layer, an intermediate layer, and an output layer, as shown in FIG.

各層は1以上のユニットから構成される。結合は、入力
層→中間層→出力層という前向きの結合だけで、各層内
での結合はない。
Each layer is composed of one or more units. The connections are only forward connections such as input layer → middle layer → output layer, and there are no connections within each layer.

(3)ユニットの構造 ユニットは第7図に示す如く脳のニューロンのモデル化
であり構造は簡単である。他のユニットから入力を受け
、その総和をとり一定の規則(変換関数)で変換し、結
果を出力する。他のユニットとの結合には、それぞれ結
合の強さを表わす可変の重みを付ける。
(3) Structure of the unit The unit is a model of a neuron in the brain and has a simple structure as shown in FIG. It receives input from other units, sums it up, transforms it using a certain rule (conversion function), and outputs the result. Each connection with another unit is given a variable weight that represents the strength of the connection.

(4)学習(パックプロパゲーション)ネットワークの
学習とは、実際の出力を目標値(望ましい出力)に近づ
けることであり、−a的には第7図に示した各ユニット
の変換関数及び重みを変化させて学習を行なう。
(4) Learning (pack propagation) Learning of a network is to bring the actual output closer to the target value (desired output). Learn by making changes.

又、学習のアルゴリズムとしては、例えば、Rumel
hart、 D、E、、McClelland、 J、
L、 and thePDP Re5earch Gr
oup、 PARALLEL DISTRIBUTED
PRO(:ESSING、 the MIT Pres
s、 1986.に記載されているパックプロパゲーシ
ョンを用いることができる。
Further, as a learning algorithm, for example, Rumel
hart, D.E., McClelland, J.
L, and thePDP Re5earch Gr.
oup, PARALLEL DISTRIBUTED
PRO(:ESSING, the MIT Pres
s, 1986. Pack propagation as described in .

以下、本発明の具体的な実施例について説明する。Hereinafter, specific examples of the present invention will be described.

話者認識システム10は、第1図に示す如く、音声入力
部11、音声処理部12、ニューラルネットワーク13
、判定部14、メモリ部15、ネットワーク制御部16
、機器制御部17を有して構成される。
As shown in FIG. 1, the speaker recognition system 10 includes a voice input section 11, a voice processing section 12, and a neural network 13.
, determination unit 14, memory unit 15, network control unit 16
, and a device control section 17.

(1)音声入力部11に登録音声を入力する。(1) Input the registered voice to the voice input section 11.

この時、学習単語を「タダイマ」、入力単語を「タダイ
マ」とする。
At this time, the learning word is "Tadaima" and the input word is "Tadaima".

又、登録話者を9名、詐称者を27名とする。Also, there are 9 registered speakers and 27 impostors.

(2)音声処理部12で、上記(1)の入力音声に簡単
な前処理を施す。
(2) The audio processing unit 12 performs simple preprocessing on the input audio in (1) above.

前処理結果は、今回の話者認識のためにニューラルネッ
トワーク13に転送されるとともに、追加学習の可能性
に備えて、メモリ部15に転送される。
The preprocessing results are transferred to the neural network 13 for the current speaker recognition, and are also transferred to the memory unit 15 in preparation for the possibility of additional learning.

(3)ニューラルネットワーク13は、下記■の学習動
作と下記■の評価動作を行なう。
(3) The neural network 13 performs the following learning operation (2) and the following evaluation operation (2).

■学習 目標値(出力層を構成する各出力ユニットの目標出力値
)を、登録話者については(1,0)、詐称者について
は(0,1)とする。
(2) The learning target value (target output value of each output unit constituting the output layer) is set to (1, 0) for the registered speaker and (0, 1) for the impostor.

登録話者の入力音声「タダイマ」に、音声処理部12に
よる前処理を施し、この前処理結果なニューラルネット
ワーク13に入力する。そして、ニューラルネットワー
ク13の出力値(出力層を構成する各出力ユニットの出
力値)が上記目標値に近づくように、ニューラルネット
ワーク13の各ユニットの変換関数及び重みを修正する
The input voice "Tadaima" of the registered speaker is preprocessed by the voice processing unit 12, and the preprocessing result is input to the neural network 13. Then, the conversion function and weight of each unit of the neural network 13 are corrected so that the output value of the neural network 13 (the output value of each output unit constituting the output layer) approaches the target value.

この学習動作を例えば3万回くり返す。This learning operation is repeated, for example, 30,000 times.

■評価 今回話者の入力音声に前処理を施し、この前処理を施し
た音声をニューラルネットワーク13に入力し、ニュー
ラルネットワークの出力値(X、Y)を得る。
■Evaluation This time, the input voice of the speaker is preprocessed, and the preprocessed voice is input to the neural network 13 to obtain the output values (X, Y) of the neural network.

そして、ニューラルネットワーク13の上記出力値(X
、Y)は判定部14に転送される。
Then, the above output value (X
, Y) are transferred to the determination unit 14.

(4)判定部14は、ニューラルネットワーク13の出
力値(X、Y)に対し、しきい値θ1.02.03 (
θ1〉θ2)を設ける。
(4) The determination unit 14 determines the threshold value θ1.02.03 (
θ1>θ2).

Olは追加学習用しきい値、θ2は登録話者認識用しき
い値、θ3は詐称者認識用しきい値である。
Ol is a threshold for additional learning, θ2 is a threshold for registered speaker recognition, and θ3 is a threshold for impostor recognition.

判定部14は、上記しきい値を用いて、下記■〜■の判
定動作を行なう。
The determination unit 14 performs the following determination operations (1) to (2) using the above threshold value.

■[X>02かつY〈θ3] であることを条件に、判定部14は、今回の入力話者を
登録話者と判定し、この登録話者判定信号を機器制御部
17に出力する。
(2) On the condition that [X>02 and Y<θ3], the determination unit 14 determines the current input speaker as a registered speaker, and outputs this registered speaker determination signal to the device control unit 17.

■[x〉θ2かつY〉θ3コ又は[X<02かつY〉θ
3]又は[X<02かつY〈θ3]であることを条件に
、判定部14は、今回の入力話者を詐称者と判定し、こ
の詐称者判定信号を機器制御部17に出力する。
■[x>θ2 and Y>θ3 or [X<02 and Y>θ
3] or [X<02 and Y<θ3], the determination unit 14 determines the current input speaker to be an impostor, and outputs this impostor determination signal to the device control unit 17.

■上記■の登録話者判定時に限り、判定部14は、更に
次の(a)  (b)の処理を行なう。
(2) Only when determining the registered speaker in (2) above, the determination unit 14 further performs the following processes (a) and (b).

(a)[X<θ1] であることを条件に、判定部14は、今回の入力音声デ
ータを用いてニューラルネットワーク13の追加学習を
行なうべく、ネットワーク制御部16に追加学習実行信
号を出力する。
(a) On the condition that [X<θ1], the determination unit 14 outputs an additional learning execution signal to the network control unit 16 in order to perform additional learning of the neural network 13 using the current input audio data. .

(b)[X>θl] である時、判定部14は何もしない。(b) [X>θl] When this is the case, the determination unit 14 does nothing.

(5)機器制御部17は、判定部14による上記■の判
定結果に基づく登録話者判定信号により、機器を制御す
る。
(5) The device control section 17 controls the device using the registered speaker determination signal based on the determination result of the above-mentioned (2) by the determination section 14.

この機器は、例えば電子錠であり、上記登録話者判定信
号に基づいて開錠制御を行なう。
This device is, for example, an electronic lock, and performs unlocking control based on the registered speaker determination signal.

(6)ネットワーク制御部16は、判定部14による上
記■の判定結果に基づく追加学習実行信号により、ニュ
ーラルネットワーク13の追加学習を行なうことを判断
する。この時、ネットワーク制御部16は、メモリ部1
5より、今回の入力音声データを取出し、この入力音声
データをニューラルネットワーク13に再入力し、この
入力に対するニューラルネットワーク13の出力値(X
、Y)か口述(3)■の登録話者についての目標値(1
、0)に近づくように、ニューラルネットワーク13の
各ユニットの変換関数及び重みを修正する。ネットワー
ク制御部16は、この追加学習動作を例えば3万回くり
返す。
(6) The network control unit 16 determines to perform additional learning of the neural network 13 based on the additional learning execution signal based on the determination result of the above-mentioned (2) by the determining unit 14. At this time, the network control unit 16 controls the memory unit 1
5, extract the current input audio data, re-input this input audio data to the neural network 13, and calculate the output value of the neural network 13 for this input (X
, Y) or oral (3)■ Target value (1
, 0), the conversion function and weight of each unit of the neural network 13 are corrected. The network control unit 16 repeats this additional learning operation, for example, 30,000 times.

以下、第2図に示す如く、階層的なニューラルネットワ
ーク13を用い、ニューラルネットワーク13の入力と
して音声の一定時間内における平均的な周波数特性の時
間的変化を用いた場合の具体的実施例について説明する
Hereinafter, as shown in FIG. 2, a specific example will be described in which a hierarchical neural network 13 is used and temporal changes in the average frequency characteristics of audio within a certain period of time are used as input to the neural network 13. do.

尚、音声処理部12は、第2図に示す如く、ローパスフ
ィルタ21、バンドパスフィルタ22、平均化回路23
の結合にて構成される。
Note that the audio processing section 12 includes a low-pass filter 21, a band-pass filter 22, and an averaging circuit 23, as shown in FIG.
It is composed of the combination of

■入力音声の音声信号の高域成分を、ローパスフィルタ
21にてカットする。そして、この入力音声を第3図に
示す如く、4つのブロックに時間的に等分割する。
(2) The high-frequency components of the audio signal of the input audio are cut by the low-pass filter 21. Then, this input audio is temporally equally divided into four blocks as shown in FIG.

■音声波形を、第2図に示す如く、複数(n個)チャン
ネルのバンドパスフィルタ22に通し、各ブロック即ち
各一定時間毎に第4図(A)〜(D)のそれぞれに示す
如くの周波数特性を得る。
■The audio waveform is passed through a plurality (n) channel band pass filter 22 as shown in FIG. Obtain frequency characteristics.

この時、バンドパスフィルタ22の出力信号は、平均化
回路23にて、各ブロック毎、即ち一定時間で平均化さ
れる。
At this time, the output signal of the bandpass filter 22 is averaged by the averaging circuit 23 for each block, that is, for a certain period of time.

以上の前処理により、「音声の一定時間内における平均
的な周波数特性の時間的変化」が得られた。
Through the above preprocessing, the "temporal change in the average frequency characteristics of audio within a certain period of time" was obtained.

平均化回路23の出力は、直接的にニューラルネットワ
ーク13に転送され、或いはメモリ部15を経由して間
接的にニューラルネットワーク13に転送される。
The output of the averaging circuit 23 is transferred directly to the neural network 13, or indirectly transferred to the neural network 13 via the memory section 15.

■ニューラルネットワーク13は、3層の階層的なニュ
ーラルネットワークにて構成される。入力層3】は、前
処理の4ブロツク、nチャンネルに対応する4Xnユニ
ツトにて構成される。出力層32は、登録話者群と詐称
者群との2ユニツトにて構成される。
■The neural network 13 is composed of a three-layer hierarchical neural network. The input layer 3 consists of 4×n units corresponding to 4 blocks of preprocessing and n channels. The output layer 32 is composed of two units: a group of registered speakers and a group of impostors.

出力層32の目標値は、登録話者については(1、Q 
)詐称者については((1、1)である。
The target value of the output layer 32 is (1, Q
) for the imposter ((1, 1).

実験 上記の如く、追加学習用しきい値θ1を設けて追加学習
したネットワークの認識率と、追加学習しないネットワ
ークの認識率とを比較した結果、表1を得た0本発明力
式により、時期差による認識率劣化を防止できることが
認められる。
ExperimentAs mentioned above, as a result of comparing the recognition rate of the network that underwent additional learning by setting the threshold value θ1 for additional learning and the recognition rate of the network that did not perform additional learning, Table 1 was obtained. It is recognized that deterioration in recognition rate due to differences can be prevented.

次に、上記実施例の作用について説明する。Next, the operation of the above embodiment will be explained.

(1)経時的な認識率の劣化が極めて少ない。このこと
は、後述する実験結果により確認されていることである
が、ニューラルネットワーク13が音声の時期差による
変動の影響を受けにくい構造をとることが可能なためと
推定される。
(1) Deterioration of recognition rate over time is extremely small. This has been confirmed by experimental results described below, and is presumed to be because the neural network 13 can have a structure that is less susceptible to fluctuations due to differences in audio timing.

(2)ニューラルネットワーク13を構成する、登録話
者に対応する出力ユニットの出力値に対し、登録話者認
識用しきい値θ2の他に追加学習用しきい値θlを設け
た。即ち、上記出力値が登録話者認識用しきい値θ2を
超えて大なるものであり、入力話者をU録話者と判定で
きるものであっても、該出力値が該登録話者認識用しき
い値θ2より大なる追加学習用しきい値θ1を超えるも
のでない場合には、今回の入力音声データを用いてニュ
ーラルネットワーク13の追加学習を行なう、これによ
り、話者の特徴が経時変化しても認識率が劣化する前に
ニューラルネットワーク13を更新でき、結果として、
音声の経時変化に強い話者認識システムを構成できる。
(2) In addition to the registered speaker recognition threshold value θ2, an additional learning threshold value θl is provided for the output value of the output unit corresponding to the registered speaker that constitutes the neural network 13. In other words, even if the output value is greater than the registered speaker recognition threshold value θ2 and the input speaker can be determined to be the U-recorded speaker, the output value does not exceed the registered speaker recognition threshold value θ2. If the additional learning threshold θ1, which is larger than the additional learning threshold θ2, is not exceeded, the neural network 13 performs additional learning using the current input voice data.This allows the speaker's characteristics to change over time. However, the neural network 13 can be updated before the recognition rate deteriorates, and as a result,
It is possible to construct a speaker recognition system that is resistant to changes in speech over time.

(3)ニューラルネットワーク13は、原理的に、ネッ
トワーク全体の演算処理が単純且つ迅速である。
(3) In principle, the neural network 13 has simple and quick arithmetic processing for the entire network.

(4)ニューラルネットワーク13は、原理的に、それ
を構成している各ユニットが独立に動作しており、並列
的な演算処理が可能である。従って、演算処理が迅速で
ある。
(4) In principle, each unit constituting the neural network 13 operates independently, and parallel arithmetic processing is possible. Therefore, calculation processing is quick.

(5)上記(3)〜(4)により、話者認識システム1
0を抜雑な処理装置によることなく容易に実時間処理で
きる。
(5) According to (3) to (4) above, the speaker recognition system 1
0 can be easily processed in real time without using complicated processing equipment.

(6)ニューラルネットワーク13への入力として、「
音声の周波数特性の時間的変化」を用いたから、入力を
得るための前処理が従来の複雑な特徴社抽出に比して、
単純となりこの前処理に要する時間か短くて足りる。
(6) As an input to the neural network 13, “
Because it uses "temporal changes in the frequency characteristics of the voice," the preprocessing to obtain input is complicated compared to conventional feature extraction.
It is simple and the time required for this preprocessing is short.

この時、上記ニューラルネットワークへの入力として、
更に、「音声の一定時間内における平均的な周波数特性
の時間的変化」を用いたから、ニューラルネットワーク
13における処理が単純となり、この処理に要する時間
がより短くて足りる。
At this time, as an input to the above neural network,
Furthermore, since the "temporal change in the average frequency characteristic within a certain period of time" is used, the processing in the neural network 13 is simple, and the time required for this processing is shorter.

(7)階層的なニューラルネットワーク13を用いたか
ら、現在、既に確立している簡単な学習アルゴリズム(
パックプロパゲーション)を用いて、高い認識率を達成
できる。
(7) Since the hierarchical neural network 13 is used, a simple learning algorithm that has already been established (
Pack propagation) can be used to achieve high recognition rates.

尚、本発明の実施においては、ニューラルネットワーク
への入力として、 ■音声の周波数特性の時間的変化、 ■音声の平均的な線形予測係数、 ■音声の平均的なPARCOR係数、 ■音声の平均的な周波数特性、及びピッチ周波数、 ■高域強調を施された音声波形の平均的な周波数特性、
並びに ■音声の平均的な周波数特性 のうちの1つ以上を使用できる。
In the implementation of the present invention, as inputs to the neural network, ■temporal changes in the frequency characteristics of audio, ■average linear prediction coefficients of audio, ■average PARCOR coefficients of audio, ■average audio frequency characteristics and pitch frequency, ■Average frequency characteristics of high-frequency emphasized audio waveforms,
and ■ one or more of the average frequency characteristics of voice can be used.

そして、上記■の要素が更に「音声の一定時間内におけ
る平均的な周波数特性の時間的変化」として用いられた
ように、上記■の要素は「音声の一定時間内における平
均的な線形予測係数の時間豹変化」、上記■の要素は「
音声の一定時間内における平均的なPARCOR係数の
時間的変化」、上記■の要素は「音声の一定時間内にお
ける平均的な周波数特性、及びピッチ周波数の時間的変
化」、上記■の要素は、[高域強調を施された音声波形
の一定時間内における平均的な周波数特性の時間的変化
」として用いることができる。
Then, just as the element (■) above was further used as "temporal change in the average frequency characteristics within a certain period of time", the element (■) above is also used as "the average linear prediction coefficient within a certain period of time". ``time leopard change'', the element of ■ above is ``
``Temporal change in the average PARCOR coefficient within a certain time period of audio'', the above element (■) is ``the average frequency characteristic and temporal change in pitch frequency within a certain time period of audio'', and the above element (■) is: It can be used as [temporal change in the average frequency characteristics within a certain period of time of a high-frequency emphasized audio waveform."

尚、上記■の線形予測係数は、以下の如く定義される。Incidentally, the linear prediction coefficient of (2) above is defined as follows.

即ち、音声波形のサンプル値(χ。)の間には、−mに
高い近接相関があることが知られている。
That is, it is known that -m has a high proximity correlation between sample values (χ) of audio waveforms.

そこで次のような線形予測か可能であると仮定する。Therefore, assume that the following linear prediction is possible.

線形予測値  χ、=−Σα1χt−1・・・(1)線
形予測誤差 ε、=χ、−χL  ・・・(2)ここで
、χ、:時刻tにおける音声波形のサンプル値、(α+
)(1=1.・・・、p): (1次の)線形予測係数 さて、本発明の実施においては、線形予測誤差εtの2
乗平均値が最小となるように線形予測係数(α邂)を求
める。
Linear predicted value χ, = -Σα1χt-1...(1) Linear prediction error ε, =χ, -χL...(2) Here, χ: sample value of the audio waveform at time t, (α+
) (1=1..., p): (first-order) linear prediction coefficient Now, in the implementation of the present invention, 2 of the linear prediction error εt
Find the linear prediction coefficient (α) so that the root mean value is the minimum.

具体的には (εt)2を求め、その時間平均を(t 
t)”と表わして、θ(t t)” / aa (=O
r z=1.2.・・・、pとおくことによって、次の
式から(a五)が求められる。
Specifically, (εt)2 is calculated and its time average is (t
t)” and θ(t t)” / aa (=O
rz=1.2. By setting . . . , p, (a5) can be obtained from the following equation.

又、上記■のPARCOR係数は以下の如く定義される
Further, the PARCOR coefficient of (2) above is defined as follows.

即ち、[k、](n=1.・・・、p)を(1次の)P
AR(:OR係数(偏自己相関係数)とする時、PAR
COR係数k。、、は、線形予測による前向き残差ε 
<11と後向き残差ε’−(n * I ) (b 1
間の正規化相関係数として、次の式によって定義される
That is, [k,] (n=1...,p) is expressed as (first-order) P
When AR (:OR coefficient (partial autocorrelation coefficient)), PAR
COR coefficient k. , , is the forward residual ε due to linear prediction
<11 and backward residual ε′−(n*I)(b 1
The normalized correlation coefficient between is defined by the following formula.

・・・(4) ココテ、εt<1ゝ=χを一Σ α直χt−1゜(α弧
) :前向き予測係数、 εt−(n+II Lb’=χt−in+Il  −F
、II J−χt−J 。
...(4) Here, εt<1ゝ=χ is one Σ α direct χt−1゜(α arc): Forward prediction coefficient, εt−(n+II Lb'=χt−in+Il −F
, II J-χt-J.

(βj):後向き予測係数 又、上記■の音声のピッチ周波数とは、声帯波の繰り返
し周期(ピッチ周期)の逆数である。
(βj): Backward prediction coefficient Also, the pitch frequency of the voice mentioned in (2) above is the reciprocal of the repetition period (pitch period) of the vocal cord wave.

尚、ニューラルネットワークへの入力として、個人差か
ある声帯の基本的なパラメータであるピッチ周波数を付
加したから、特に大人/小人、男性/女性間の話者の認
識率を向上することができる。
Furthermore, since the pitch frequency, which is a basic parameter of the vocal cords that varies from person to person, was added as an input to the neural network, it is possible to improve the recognition rate of speakers, especially between adults/dwarfs and male/female. .

又、上記■の高域強調とは、音声波形のスペクトルの平
均的な傾きを補償して、低域にエネルギが集中すること
を防止することである。然るに、音声波形のスペクトル
の平均的な傾きは話者に共通のものであり、話者の認識
には無関係である。
Furthermore, the above-mentioned high frequency enhancement (2) is to compensate for the average slope of the spectrum of the audio waveform to prevent concentration of energy in the low frequency range. However, the average slope of the spectrum of the speech waveform is common to all speakers and is unrelated to the speaker's recognition.

ところが、このスペクトルの平均的な傾きが補償されて
いない音声波形をそのままニューラルネットワークへ入
力する場合には、ニューラルネットワークが学習する時
にスペクトルの平均的な傾きの特徴の方を抽出してしま
い、話者の認識に必要なスペクトルの山と谷を抽出する
のに時間がかかる。これに対し、ニューラルネットワー
クへの入力を高域強調する場合には、話者に共通で、認
識には無関係でありながら、学習に影響を及ぼすスペク
トルの平均的な傾きを補償できるため、学習速度が速く
なるのである。
However, when inputting a speech waveform that has not been compensated for the average slope of the spectrum to a neural network as is, the neural network extracts the feature of the average slope of the spectrum during learning, and the speech becomes distorted. It takes time to extract the peaks and valleys of the spectrum necessary for human recognition. On the other hand, when high-frequency emphasis is applied to the input to a neural network, it is possible to compensate for the average slope of the spectrum that is common to all speakers and is unrelated to recognition, but that affects learning, which speeds up the learning process. becomes faster.

[発明の効果] 以上のように本発明によれば、経時的な認識率の劣化が
極めて少なく、容易に実時間処理できる話者認識システ
ムを得ることができる。
[Effects of the Invention] As described above, according to the present invention, it is possible to obtain a speaker recognition system that exhibits extremely little deterioration in recognition rate over time and that can easily perform real-time processing.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は本発明が適用された話者認識システムの一例を
示す模式図、第2図は音声処理部とニューラルネットワ
ークの一例を示す模式図、第3図は入力音声を示す模式
図、第4図はバンドパスフィルタの出力を示す模式図、
第5図はニューラルネットワークを示す模式図、第6図
は階層的なニューラルネットワークを示す模式図、第7
図はユニットの構造を示す模式図である。 ○・・・話者認識システム、 1・・・音声入力部、 2・・・音声処理部、 3・・・ニューラルネッ 4・・・判定部、 5・・・メモリ部、 6・・・ネットワーク制御部、 7・・・機器制御部。 トワーク、
FIG. 1 is a schematic diagram showing an example of a speaker recognition system to which the present invention is applied, FIG. 2 is a schematic diagram showing an example of a speech processing unit and a neural network, FIG. 3 is a schematic diagram showing an input voice, and FIG. Figure 4 is a schematic diagram showing the output of the bandpass filter.
Figure 5 is a schematic diagram showing a neural network, Figure 6 is a schematic diagram showing a hierarchical neural network, and Figure 7 is a schematic diagram showing a hierarchical neural network.
The figure is a schematic diagram showing the structure of the unit. ○...Speaker recognition system, 1...Speech input section, 2...Speech processing section, 3...Neural network 4...Judgment section, 5...Memory section, 6...Network Control unit, 7... equipment control unit. twerk,

Claims (3)

【特許請求の範囲】[Claims] (1)ニューラルネットワークを用いた話者認識システ
ムであって、登録話者に対応する出力ユニットの出力値
に対し、登録話者認識用しきい値と追加学習用しきい値
とを設定し、上記出力値が登録話者認識用しきい値より
大なることを条件に、今回の入力話者を登録話者と判定
し、上記出力値が登録話者認識用しきい値より大、かつ
追加学習用しきい値より小なることを条件に、今回の入
力音声データを用いてニューラルネットワークの追加学
習を行なう話者認識システム。
(1) A speaker recognition system using a neural network, which sets a registered speaker recognition threshold and an additional learning threshold for the output value of an output unit corresponding to a registered speaker, On the condition that the above output value is greater than the threshold for registered speaker recognition, the current input speaker is determined to be a registered speaker, and the above output value is greater than the threshold for registered speaker recognition, and additional information is added. A speaker recognition system that performs additional training of a neural network using the current input voice data, provided that it is smaller than a learning threshold.
(2)前記ニューラルネットワークへの入力として、 [1]音声の周波数特性の時間的変化、 [2]音声の平均的な線形予測係数、 [3]音声の平均的なPARCOR係数、 [4]音声の平均的な周波数特性、及びピッチ周波数、 [5]高域強調を施された音声波形の平均的な周波数特
性、並びに [6]音声の平均的な周波数特性 のうちの1つ以上を使用する請求項1記載の話者認識シ
ステム。
(2) As inputs to the neural network, [1] Temporal changes in the frequency characteristics of speech, [2] Average linear prediction coefficients of speech, [3] Average PARCOR coefficients of speech, [4] speech using one or more of the following: average frequency characteristics and pitch frequency; [5] average frequency characteristics of high-frequency emphasized audio waveform; and [6] average frequency characteristics of audio. The speaker recognition system according to claim 1.
(3)前記ニューラルネットワークが階層的なニューラ
ルネットワークである請求項1又は2記載の話者認識シ
ステム。
(3) The speaker recognition system according to claim 1 or 2, wherein the neural network is a hierarchical neural network.
JP1298503A 1989-11-16 1989-11-16 Speaker recognition system Expired - Lifetime JP2510301B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP1298503A JP2510301B2 (en) 1989-11-16 1989-11-16 Speaker recognition system

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP1298503A JP2510301B2 (en) 1989-11-16 1989-11-16 Speaker recognition system

Publications (2)

Publication Number Publication Date
JPH03157698A true JPH03157698A (en) 1991-07-05
JP2510301B2 JP2510301B2 (en) 1996-06-26

Family

ID=17860556

Family Applications (1)

Application Number Title Priority Date Filing Date
JP1298503A Expired - Lifetime JP2510301B2 (en) 1989-11-16 1989-11-16 Speaker recognition system

Country Status (1)

Country Link
JP (1) JP2510301B2 (en)

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2002269047A (en) * 2001-03-07 2002-09-20 Nec Eng Ltd Sound user authentication system
JPWO2006109515A1 (en) * 2005-03-31 2008-10-23 パイオニア株式会社 Operator recognition device, operator recognition method, and operator recognition program
JP2012181280A (en) * 2011-02-28 2012-09-20 Sogo Keibi Hosho Co Ltd Sound processing device and sound processing method
CN111883106A (en) * 2020-07-27 2020-11-03 腾讯音乐娱乐科技(深圳)有限公司 Audio processing method and device

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2002269047A (en) * 2001-03-07 2002-09-20 Nec Eng Ltd Sound user authentication system
JPWO2006109515A1 (en) * 2005-03-31 2008-10-23 パイオニア株式会社 Operator recognition device, operator recognition method, and operator recognition program
JP4588069B2 (en) * 2005-03-31 2010-11-24 パイオニア株式会社 Operator recognition device, operator recognition method, and operator recognition program
JP2012181280A (en) * 2011-02-28 2012-09-20 Sogo Keibi Hosho Co Ltd Sound processing device and sound processing method
CN111883106A (en) * 2020-07-27 2020-11-03 腾讯音乐娱乐科技(深圳)有限公司 Audio processing method and device
CN111883106B (en) * 2020-07-27 2024-04-19 腾讯音乐娱乐科技(深圳)有限公司 Audio processing method and device

Also Published As

Publication number Publication date
JP2510301B2 (en) 1996-06-26

Similar Documents

Publication Publication Date Title
JPH06161496A (en) Voice recognition system for recognizing remote control command words for home appliances
JP2510301B2 (en) Speaker recognition system
JPH03157697A (en) Word recognizing system
JPH03230200A (en) Voice recognizing method
CN116434758B (en) Voiceprint recognition model training method and device, electronic equipment and storage medium
JPH03111899A (en) Voice lock device
EP0369485B1 (en) Speaker recognition system
JP2518939B2 (en) Speaker verification system
JPH02273798A (en) Speaker recognition system
JP2559506B2 (en) Speaker verification system
JP2518940B2 (en) Speaker verification system
JPH03144176A (en) Voice-controlled hot water supply device
JPH02275996A (en) Word recognition system
JPH05143094A (en) Speaker recognition system
JPH02273799A (en) Speaker recognition system
JPH0415700A (en) Speaker recognition system
CN115862636B (en) A method of Internet man-machine verification based on speech recognition technology
JPH05257496A (en) Word recognizing system
JPH02273796A (en) Speaker recognition method
JPH02304498A (en) Word recognition system
JPH03230255A (en) Sound recognizing method
JPH0415694A (en) Word recognition system
JPH05181500A (en) Word recognition system
JPH02273800A (en) Speaker recognition system
JPH03114345A (en) Caller recognition telephone system