JPH0415700A - Speaker recognition system - Google Patents

Speaker recognition system

Info

Publication number
JPH0415700A
JPH0415700A JP2120867A JP12086790A JPH0415700A JP H0415700 A JPH0415700 A JP H0415700A JP 2120867 A JP2120867 A JP 2120867A JP 12086790 A JP12086790 A JP 12086790A JP H0415700 A JPH0415700 A JP H0415700A
Authority
JP
Japan
Prior art keywords
neural network
speakers
speaker
group
output
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP2120867A
Other languages
Japanese (ja)
Inventor
Hidekazu Tsuda
津田 英一
Shingo Nishimura
新吾 西村
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sekisui Chemical Co Ltd
Original Assignee
Sekisui Chemical Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sekisui Chemical Co Ltd filed Critical Sekisui Chemical Co Ltd
Priority to JP2120867A priority Critical patent/JPH0415700A/en
Publication of JPH0415700A publication Critical patent/JPH0415700A/en
Pending legal-status Critical Current

Links

Abstract

(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。
(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.

Description

【発明の詳細な説明】 [産業上の利用分野] 本発明は話者認識システムに関する。[Detailed description of the invention] [Industrial application field] The present invention relates to speaker recognition systems.

[従来の技術] 本出願人は、ニューラルネットワークを用いた話者認識
システムを提案している。この話者認識システムにあっ
ては、ニューラルネットワークの出カバターンに対し一
定のしきい値θを設け、例えばニューラルネットワーク
の1つの出力ユニットの出力値が0以上の値をとり、他
の出力ユニットの出力値の全てか(1−θ)以下の値を
とる場合に、今回の入力話者は出力値が0以上である出
力ユニットに対応する登録話者と同一話者であるものと
認識する。
[Prior Art] The present applicant has proposed a speaker recognition system using a neural network. In this speaker recognition system, a certain threshold value θ is set for the output pattern of the neural network. For example, when the output value of one output unit of the neural network takes a value of 0 or more, If all of the output values are less than or equal to (1-θ), it is recognized that the current input speaker is the same speaker as the registered speaker corresponding to the output unit whose output value is 0 or more.

[発明か解決しようとする課題] 然しなから、ニューラルネットワークの出カバターンは
必ずしも上述の如くにならず、■θ以−Fの出力値を示
す出力ユニットか複数個あるパターン、或いは■全ての
出力ユニットの出力値が(l−θ)以下の値をとるパタ
ーン等の出現をみることかある。そして、このような出
カバターンについては、話者認識か困難ないし不能とな
るのである。
[Invention or problem to be solved] However, the output pattern of a neural network is not necessarily as described above, and there is a pattern in which there are two or more output units that exhibit an output value of θ or more, or ■ all outputs. Sometimes we see the appearance of a pattern in which the output value of the unit takes a value less than or equal to (l-θ). Such output patterns make it difficult or impossible to recognize the speaker.

本発明は、全登録話者を学習対象として構築されたニュ
ーラルネットワークの出力パターンが入力話者を特定て
きないパターンである場合にも、正確に話者認識を行な
うことを目的とする。
An object of the present invention is to accurately perform speaker recognition even when the output pattern of a neural network constructed using all registered speakers as learning targets is a pattern that does not identify the input speaker.

[課題を解決するための手段] 請求項1に記載の本発明は、ニューラルネットワークを
用いた話者認識システムにおいて、全登録話者を学習対
象として構築した全話者用ニューラルネットワークと、
全話者用ニューラルネットワークを構成する各出力ユニ
ットの出力値か互いに一定の類似関係にありしきい値判
定できないことに基づいて、全登録話者をグループ分け
した各グループ毎の登録話者を学習対象として構築した
各グループ用ニューラルネットワークとを用い、入力話
者について、全話者用ニューラルネットワークによる認
識を行ない、該全話者用ニューラルネットワークの出力
パターンが入力話者を特定できないパターンである時、
更に、該全話者用ニューラルネットワークにおける最大
出力値の出力ユニットに対応する登録話者を含むグルー
プの、グループ用ニューラルネットワークによる認識を
行なうようにしたものである 請求項2に記載の本発明は、前記ニューラルネットワー
クへの入力として、 ■音声の周波数特性の時間的変化、 ■音声の平均的な線形予測係数、 ■音声の平均的なPARCOR係数、 ■音声の平均的な周波数特性、及びピッチ周波数、 ■高域強調を施された音声波形の平均的な周波数特性、
並びに ■音声の平均的な周波数特性 のうちの1つ以上を使用するようにしたものである。
[Means for Solving the Problems] The present invention according to claim 1 is a speaker recognition system using a neural network, which includes a neural network for all speakers constructed with all registered speakers as learning targets;
All registered speakers are divided into groups based on the fact that the output values of the output units that make up the neural network for all speakers have a certain similarity relationship with each other and cannot be judged by a threshold value.Learn the registered speakers for each group by dividing all registered speakers into groups. When the input speaker is recognized by the neural network for all speakers using the neural network for each group constructed as a target, and the output pattern of the neural network for all speakers is a pattern in which the input speaker cannot be identified. ,
Further, the present invention according to claim 2 is characterized in that a group including registered speakers corresponding to the output unit of the maximum output value in the neural network for all speakers is recognized by the group neural network. , As inputs to the neural network, ■temporal changes in the frequency characteristics of the voice, ■average linear prediction coefficients of the voice, ■average PARCOR coefficients of the voice, ■average frequency characteristics of the voice, and pitch frequency. , ■Average frequency characteristics of the audio waveform with high-frequency emphasis,
and (1) one or more of the average frequency characteristics of voice is used.

[作用] 請求項1に記載の本発明によれば、下記■の作用効果か
ある。
[Function] According to the present invention as set forth in claim 1, there is the following function and effect.

■全登録話者を学習対象として構築した全話者用ニュー
ラルネットワークと、類似話者にてグループ化された登
録話者を学習対象として構築した各グループ用ニューラ
ルネットワークとを用いることにより、話者認識率を向
上てきる。これにより、全登録話者を学習対象として構
築されたニューラルネットワークの出カバターンが入力
話者を特定できないパターンである場合にも、正確に話
者認識を行なうことができる。
■By using a neural network for all speakers constructed using all registered speakers as learning targets and a neural network for each group constructed using registered speakers grouped by similar speakers as learning targets, It can improve the recognition rate. As a result, even if the output pattern of a neural network constructed using all registered speakers as learning targets is a pattern that cannot identify the input speaker, accurate speaker recognition can be performed.

請求項2に記載の本発明によれば、下記■の作用がある
According to the present invention as set forth in claim 2, there is the following effect (2).

■ニューラルネットワークへの入力として、請求項1に
記載の■〜■の各要素のうちの1つ以上を用いるから、
入力を得るための前処理か単純となり、この前処理に要
する時間か短くて足りるため、話者認識システムを複雑
な処理装置によることなく容易に実時間処理できる。
■As input to the neural network, one or more of the elements of ■ to ■ according to claim 1 are used,
Since the preprocessing for obtaining input is simple and the time required for this preprocessing is short, the speaker recognition system can be easily processed in real time without using a complicated processing device.

[実施例] 第1図は全話者用ニューラルネットワークの学習系統を
示すブロック図、第2図はグループ用ニューラルネット
ワークの学習系統を示すブロック図、第3図は話者認識
系統を示すブロック図、第4図は話者認識系統を示す流
れ図である。
[Example] Fig. 1 is a block diagram showing the learning system of the neural network for all speakers, Fig. 2 is a block diagram showing the learning system of the neural network for groups, and Fig. 3 is a block diagram showing the speaker recognition system. , FIG. 4 is a flowchart showing the speaker recognition system.

fA)先ず、全話者用ニューラルネットワークの学習系
統について説明する(第1図参照)。
fA) First, the learning system of the neural network for all speakers will be explained (see Fig. 1).

この系統は、音声入力部11、前処理部12、全話者用
ニューラルネットワーク13にて構成される。
This system includes a voice input section 11, a preprocessing section 12, and a neural network 13 for all speakers.

以下、前処理部12、全話者用ニューラルネットワーク
13の構成について説明する。
The configurations of the preprocessing unit 12 and the neural network for all speakers 13 will be described below.

(1)前処理部 前処理部12は、入力音声に簡単な前処理を施し、上記
全話者用ニューラルネットワーク13、後述するグルー
プ用ニューラルネットワーク14への入力データを作成
する。
(1) Pre-processing unit The pre-processing unit 12 performs simple pre-processing on input speech to create input data to the all-speaker neural network 13 and the group neural network 14 described below.

前処理部12の具体的構成を例示すれば以下の如くであ
る。
A specific example of the configuration of the preprocessing section 12 is as follows.

即ち、前処理部12としては、ローパスフィルタ、バン
ドパスフィルタ、平均化回路の結合からなるものを用い
ることかできる。
That is, the preprocessing section 12 may be a combination of a low-pass filter, a band-pass filter, and an averaging circuit.

■入力音声の音声信号の高域の雑音成分を、ローパスフ
ィルタにてカットする。そして、この入力音声を4つの
ブロックに時間的に等分割する。
■Cut the high-frequency noise components of the input audio signal using a low-pass filter. Then, this input audio is temporally equally divided into four blocks.

■音声波形を、複数(n個)チャンネルのバンドパスフ
ィルタに通し、各ブロック即ち各一定時間毎の周波数特
性を得る。
(2) Pass the audio waveform through a bandpass filter of multiple (n) channels to obtain frequency characteristics for each block, that is, for each fixed time period.

この時、バントパスフィルタの出力信号は、平均化回路
にて、各ブロック毎、即ち各一定時間で平均化される。
At this time, the output signal of the band pass filter is averaged for each block, that is, for each fixed period of time, in an averaging circuit.

以上の前処理により、「音声の一定時間内における平均
的な周波数特性の時間的変化」か得られる。
Through the above pre-processing, the "temporal change in the average frequency characteristics of audio within a certain period of time" can be obtained.

(2)全話者用ニューラルネットワーク全話者用ニュー
ラルネットワーク13は、入力音声か全登録話者のいず
れであるかを判定する。
(2) Neural network for all speakers The neural network for all speakers 13 determines whether the input voice is the input voice or all registered speakers.

全話者用ニューラルネットワーク13の具体的構成を例
示すれば、以下の如くである。
A specific configuration of the neural network 13 for all speakers is as follows.

■構造 全話者用ニューラルネットワーク13は例えば3層バー
セプトロン型てあり、入カニニット数は前処理部12の
4ブロツク、nチャンネルに対応する4n個、出力ユニ
ット数は登録話者と同数個である。
■Structure The neural network 13 for all speakers is, for example, a three-layer berceptron type, and the number of input units is 4 blocks of the preprocessing section 12, 4n units corresponding to n channels, and the number of output units is the same number as the number of registered speakers. .

■学習 目標値は、登録話者について対応する出力ユニットの出
力値を 1、その他の出力値を 0とする。
■The learning target value is 1 for the output value of the corresponding output unit for the registered speaker, and 0 for the other output values.

(a)登録話者の音声に前処理部12による前処理を施
し、全話者用ニューラルネットワーク13に入力する。
(a) The voice of the registered speaker is subjected to preprocessing by the preprocessing unit 12 and is input to the neural network 13 for all speakers.

目標値に近づくように全話者用ニューラルネットワーク
】3の重みと変換関数を修正する。
Modify the weights and conversion functions in [Neural Network for All Speakers] 3 so that they approach the target values.

(a)を目標値と出力ユニットの出力値の誤差か、十分
に小さな値(例えば、I X 10−’)になるまて繰
り返す。
(a) is repeated until the error between the target value and the output value of the output unit becomes a sufficiently small value (for example, I x 10-').

(B1次に、グループ用ニューラルネットワークの学習
系統について説明する(第2図参照)。
(B1 Next, the learning system of the group neural network will be explained (see Fig. 2).

この系統は、音声入力部11、前処理部12、全話者用
ニューラルネットワーク13、グループ用ニューラルネ
ットワーク14、判定部15、学習パターン記憶部16
にて構成される。
This system includes a voice input section 11, a preprocessing section 12, a neural network for all speakers 13, a neural network for groups 14, a determination section 15, and a learning pattern storage section 16.
Consists of.

以下、グループ用ニューラルネットワーク14、判定部
15、学習パターン記憶部16の構成について説明する
。前処理部12は、前述(A)の前処理部12と同一で
あり、全話者用ニューラルネットワーク13は前述(A
)にて構築済のものを用いる。
The configurations of the group neural network 14, determination section 15, and learning pattern storage section 16 will be described below. The preprocessing unit 12 is the same as the preprocessing unit 12 described above (A), and the neural network 13 for all speakers is the same as the preprocessing unit 12 described above (A).
) is used.

各グループ用ニューラルネットワーク14は、予めグル
ープ分けした各グループ毎に対応して設けられ、入力音
声か各グループ内の登録話者のいずれであるかを判定す
る。尚、各グループは、[全話者用ニューラルネットワ
ーツク13を構成する各出力ユニットの出力値か互いに
一定の類似関係にありしきい値判定てきない話者の組」
毎にグループ化して構成したものである。そして、各グ
ループを構成することとなる話者の入力音声は、判定部
15の判定結果に基づき上述の如くにグループ化されて
学習パターン記憶部16に記憶され、グループ用ニュー
ラルネットワーク14の学習のために供される。尚、学
習パターン記憶部16はグループ記憶部16Aを付帯的
に備えている。
Each group neural network 14 is provided corresponding to each group that has been divided into groups in advance, and determines whether it is an input voice or a registered speaker in each group. Each group is defined as a group of speakers whose output values of the output units constituting the neural network 13 for all speakers have a certain similarity to each other and whose threshold value cannot be determined.
It is organized by grouping. Then, the input voices of the speakers forming each group are grouped as described above based on the determination result of the determination unit 15 and stored in the learning pattern storage unit 16, and the learning of the group neural network 14 is performed. provided for. Note that the learning pattern storage section 16 additionally includes a group storage section 16A.

ここで、上述の「全話者用ニューラルネットワーク13
の各出力ユニットの出力値か互いに一定の類似関係にあ
りしきい値判定できない話者の組」とは、■出カバター
ンの中てしきい値0以上の値をとるユニットか複数個あ
り、或いは全てのユニットか(1−θ)以下の値をとる
ためにしきい値判定できない状況下て、■例えば最大出
力値の出力ユニットに対応する話者、及び、該最大出力
値から一定の差をなす範囲内に出力値かある出力ユニッ
トに対応する話者(又は該最大出力値と一定の比率をな
す範囲内に出力値がある出力ユニットに対応する話者)
からなる話者の組をいう。
Here, the above-mentioned "neural network 13 for all speakers"
A set of speakers whose output values of each output unit have a certain similarity relationship with each other and whose threshold value cannot be determined is defined as: ■ There are multiple units whose output values are equal to or higher than the threshold value of 0 in the output turn, or In a situation where threshold judgment cannot be made because all units take a value less than (1-θ), for example, the speaker corresponding to the output unit with the maximum output value and the speaker whose output unit has a certain difference from the maximum output value A speaker corresponding to an output unit whose output value is within a range (or a speaker corresponding to an output unit whose output value is within a range that is a certain ratio to the maximum output value)
A group of speakers consisting of

グループ用ニューラルネットワーク14の具体的構成を
例示すれば、以下の如くである。
A specific example of the configuration of the group neural network 14 is as follows.

■構造 グループ用ニューラルネットワーク14は例えば3層パ
ーセプトロン型てあり、入カニニット数は前処理部】2
の4ブロツク、nチャンネルに対応する4n個、出力ユ
ニット数は当該グループを構成する登録話者数と同数個
である。
■The structural group neural network 14 is, for example, a three-layer perceptron type, and the number of input units is 2 in the preprocessing section.
There are 4 blocks, 4n units corresponding to n channels, and the number of output units is the same as the number of registered speakers constituting the group.

■学習 目標値は、当該グループを構成する登録話者について対
応する出力ユニットの出力値を1、その他の出力値を0
とする。
■The learning target value is the output value of the corresponding output unit for the registered speakers composing the group, and the other output values are 0.
shall be.

(a)当該グループを構成する登録話者の音声に前処理
部12による前処理を施し、グループ用ニューラルネッ
トワーク14に入力する。目標値に近づくようにグルー
プ用ニューラルネットワーク14の重みと変換関数を修
正する。
(a) The preprocessing unit 12 performs preprocessing on the voices of the registered speakers constituting the group, and inputs the preprocessed voices to the group neural network 14. The weights and conversion function of the group neural network 14 are corrected so as to approach the target values.

(a)を目標値と出力ユニットの出力値の誤差が、十分
に小さな値(例えば、I X 10”’)になるまで繰
り返す。
(a) is repeated until the error between the target value and the output value of the output unit becomes a sufficiently small value (for example, I x 10'').

(C)次に、本発明による話者認識系統について説明す
る(第3図、第4図参照)。
(C) Next, the speaker recognition system according to the present invention will be explained (see FIGS. 3 and 4).

話者認識システム10は、音声入力部11、前処理部1
2、全話者用ニューラルネットワーク13、グループ用
ニューラルネットワーク14、及び、第1判定部21、
グループ記憶部22、ネットワーク選択部23、最終判
定部24にて構成される。この時、全話者用ニューラル
ネットワーク13は前述(A)にて構築され、グループ
用ニューラルネットワーク14は前述(B)にて構築さ
れたものである。
The speaker recognition system 10 includes a voice input section 11 and a preprocessing section 1.
2. Neural network for all speakers 13, neural network for group 14, and first determination unit 21,
It is composed of a group storage section 22, a network selection section 23, and a final determination section 24. At this time, the all-speaker neural network 13 was constructed in the above (A), and the group neural network 14 was constructed in the above (B).

話者認識システム1oは、下記(1)〜(3)のアルゴ
リズムにより話者認識する。
The speaker recognition system 1o recognizes speakers using the following algorithms (1) to (3).

(1)入力音声に対し、全話者用ニューラルネットワー
ク13による認識を行なう。第1判定部21により、全
話者用ニューラルネットワーク13の出カバターンをし
きい値判定して話者認識し、入力話者を特定する。
(1) The input speech is recognized by the neural network 13 for all speakers. The first determination unit 21 performs speaker recognition by thresholding the output patterns of the neural network 13 for all speakers, and identifies the input speaker.

(2)上記(1)において、全話者用ニューラルネット
ワーク13の出カバターンが入力話者を特定てきないパ
ターンである時、グループ記憶部22、ネットワーク選
択部23を用いて、該全話者用ニューラルネットワーク
13における最大出力値の出力ユニットに対応する登録
話者を含むグループのグループ用ニューラルネットワー
ク14を選択する。
(2) In (1) above, when the output pattern of the neural network 13 for all speakers is a pattern that does not specify the input speaker, the group storage unit 22 and network selection unit 23 are used to The group neural network 14 of the group including the registered speaker corresponding to the output unit with the maximum output value in the neural network 13 is selected.

(3)前記(1)の入力音声に対し、上記(2)により
選択されたグループ用ニューラルネットワーク14によ
る認識を行なう。最終判定部24により、グループ用ニ
ューラルネットワーク14の出カバターンをしきい値判
定して話者認識し、入力話者を特定する。
(3) The input voice in (1) is recognized by the group neural network 14 selected in (2) above. The final determination unit 24 performs speaker recognition by thresholding the output pattern of the group neural network 14 and identifies the input speaker.

以下、上記話者認識システム10の具体的実施結果につ
いて説明する。
Hereinafter, specific implementation results of the speaker recognition system 10 will be explained.

■登録話者として、5名を用い、各話者音声試料に前処
理を施し、64次元(4ブロツク×16チヤンネル)の
特徴ベクトルを得る。これをニューラルネットワークの
入力として、全話者用ニューラルネットワーク13を構
築する。学習パターンとして、登録話者5名、非登録話
者25名を用いた。
(5) Using 5 people as registered speakers, preprocessing is performed on each speaker's voice sample to obtain a 64-dimensional (4 blocks x 16 channels) feature vector. Using this as input to the neural network, a neural network 13 for all speakers is constructed. As learning patterns, 5 registered speakers and 25 non-registered speakers were used.

■全話者用ニューラルネットワーク13を構成する各出
力ユニットの出力値か互いに一定の類似関係にあり、し
きい値判定できないことに基づいてグループ分けした各
グループの登録話者毎にグループ化したグループ用ニュ
ーラルネットワーク14を構築する。
■Groups for each registered speaker grouped based on the fact that the output values of the output units constituting the neural network 13 for all speakers have a certain similar relationship to each other and cannot be judged by a threshold value. A neural network 14 is constructed for this purpose.

グループ1 話者Aと話者B グループ2 話者Cと話者り等 ■全話者用ニューラルネットワーク13、及びグループ
用ニューラルネットワーク14を用い、前記アルゴリズ
ムに従って話者認識を行なう。
Group 1 Speaker A and Speaker B Group 2 Speaker C and Speaker R etc. ② Using the neural network 13 for all speakers and the neural network 14 for group, speaker recognition is performed according to the above algorithm.

従来の全登録話者を対象としたニューラルネットワーク
のみの場合では、出力ユニット値か口〜1の範囲て、 ■入力話者Aに対し、出力ユニットAの出力値か0.1
7、出力ユニットBの出力値は0.19■入力話者Cに
対し、出力ユニットCの出力値は0.19 、出力ユニ
ットDの出力値は0.20等となり、話者の認識が困難
であった。
In the conventional case of only a neural network targeting all registered speakers, the output unit value is in the range of 1 to 1. ■The output value of output unit A is 0.1 for input speaker A.
7. The output value of output unit B is 0.19 ■For input speaker C, the output value of output unit C is 0.19, the output value of output unit D is 0.20, etc., making it difficult to recognize the speaker. Met.

これに対し、本発明の話者認識システム10によれば、
全話者用ニューラルネットワーク13に加え、グループ
用ニューラルネットワーク14を用いることにより、上
述の各登録話者を正確に認識することか可能となった。
In contrast, according to the speaker recognition system 10 of the present invention,
By using the group neural network 14 in addition to the all-speaker neural network 13, it has become possible to accurately recognize each of the registered speakers described above.

又、前述の前処理部12により、入力音声を前処理して
作成されるニューラルネットワークへの入力としては、 ■音声の周波数特性の時間的変化、 ■音声の平均的な線形予測係数、 ■音声の平均的なPARCOR係数、 ■音声の平均的な周波数特性、及びピッチ周波数、 ■高域強調を施された音声波形の平均的な周波数特性、
並びに ■音声の平均的な周波数特性 のうちの1つ以上を使用できる。
In addition, the inputs to the neural network created by preprocessing the input audio by the preprocessing unit 12 described above include: ■ Temporal changes in the frequency characteristics of the audio, ■ Average linear prediction coefficients of the audio, and ■ Audio. The average PARCOR coefficient of, ■The average frequency characteristic of the voice and the pitch frequency, ■The average frequency characteristic of the voice waveform with high-frequency emphasis,
and ■ one or more of the average frequency characteristics of voice can be used.

そして、上記■の要素は「音声の一定時間内における平
均的な周波数特性の時間的変化」、上記■の要素は「音
声の一定時間内における平均的な線形予測係数の時間的
変化」、上記■の要素は「音声の一定時間内における平
均的なPARCOR係数の時間的変化」、上記■の要素
は「音声の一定時間内における平均的な周波数特性、及
びピッチ周波数の時間的変化」、上記■の要素は、「高
域強調を施された音声波形の一定時間内における平均的
な周波数特性の時間的変化」として用いることがてきる
The element ■ above is the "temporal change in the average frequency characteristics of the audio within a certain time", the element ■ above is the "temporal change in the average linear prediction coefficient within the certain time of the audio", and the element The element (■) is "temporal change in the average PARCOR coefficient within a certain period of time", the element (■) above is "the average frequency characteristic and temporal change in pitch frequency within a certain period of time" (above). The element (2) can be used as a "temporal change in the average frequency characteristic within a certain period of time of a high-frequency emphasized audio waveform."

尚、上記■の線形予測係数は、以下の如く定義される。Incidentally, the linear prediction coefficient of (2) above is defined as follows.

即ち、音声波形のサンプル値(χ。)の間には、一般に
高い近接相関かあることか知られている。
That is, it is known that there is generally a high proximity correlation between sample values (χ) of audio waveforms.

そこで次のような線形予測か可能であると仮定する。Therefore, assume that the following linear prediction is possible.

線形予測値  χt=−Σα1χ、−1  ・・・(1
)線形予測誤差 ε、=χ、−χ、  ・・・(2)こ
こで、χt=時刻tにおける音声波形のサンプル値、(
α+) (1= L”・・、p): (9次の)線形予
測係数 さて、本発明の実施においては、線形予測誤差ε、の2
乗平均値か最小となるように線形予測係数(α口を求め
る。
Linear predicted value χt=-Σα1χ, -1...(1
) Linear prediction error ε, = χ, -χ, ... (2) Here, χt = sample value of the audio waveform at time t, (
α+) (1=L”..., p): (9th-order) linear prediction coefficient Now, in the implementation of the present invention, the linear prediction error ε, 2
Find the linear prediction coefficient (α) so that the root mean value is the minimum.

具体的には (εt)2を求め、その時間平均を(ε、
)2と表わして、θ(εt)2/δα、:o、i=1.
2.・・・、pとおくことによって、次の式から(α1
)か求められる。
Specifically, (εt)2 is calculated and its time average is (ε,
)2, θ(εt)2/δα, :o, i=1.
2. ..., p, from the following equation (α1
) is required.

ΣCI Ivll−Jl  =L J= ’ +  2
+ ””+  p”” (3)又、上記■のPAR(1
:OR係数は以下の如く定義される。
ΣCI Ivll-Jl = L J = ' + 2
+ “”+ p”” (3) Also, PAR (1
:The OR coefficient is defined as follows.

即ち、[k、1)(n=1.・・・、p)を(9次の)
PARCOR係数(偏自己相関係数)とする時、PAR
COR係数k nilは、線形予測による前向き残差ε
、<1)と後向き残差εt−fn+ll (b’間の正
規化相関係数として、次の式によって定義される。
That is, [k, 1) (n=1..., p) (9th order)
When using PARCOR coefficient (partial autocorrelation coefficient), PAR
The COR coefficient k nil is the forward residual ε due to linear prediction
, <1) and the backward residual εt-fn+ll (b') is defined by the following equation.

ε (f)・ εt−+n++1 ・・・ (4) ここて、ε 11 :χt−Σ α・χ゛−”(α、)
:前向き予測係数、 (βj):後向き予測係数 又、上記■の音声のピッチ周波数とは、声帯波の繰り返
し周期(ピッチ周期)の逆数である。
ε (f)・εt−+n++1 ... (4) Here, ε 11: χt−Σ α・χ゛−”(α,)
: Forward prediction coefficient, (βj) : Backward prediction coefficient Further, the pitch frequency of the voice mentioned above (■) is the reciprocal of the repetition period (pitch period) of the vocal cord wave.

尚、ニューラルネットワークへの入力として、個人差か
ある声帯の基本的なパラメータであるピッチ周波数を付
加したから、特に大人/小人、男性/女性間の話者の認
識率を向上することかてきる。
Furthermore, since we added pitch frequency, which is a basic parameter of the vocal cords that varies from person to person, as an input to the neural network, it is possible to improve the recognition rate of speakers, especially between adults/dwarfs and male/female. Ru.

又、上記■の高域強調とは、音声波形のスペクトルの平
均的な傾きを補償して、低域にエネルギか集中すること
を防止することである。然るに、音声波形のスペクトル
の平均的な傾きは話者に共通のものてあり、話者の認識
には無関係である。
Furthermore, the above-mentioned high frequency enhancement (2) is to compensate for the average slope of the spectrum of the audio waveform to prevent energy from concentrating in the low frequency range. However, the average slope of the spectrum of the speech waveform is common to all speakers and is unrelated to speaker recognition.

ところが、このスペクトルの平均的な傾きか補償されて
いない音声波形をそのままニューラルネットワークへ入
力する場合には、ニューラルネットワークが学習する時
にスペクトルの平均的な傾きの特徴の方を抽出してしま
い、話者の認識に必要なスペクトルの山と谷を抽出する
のに時間がかかる。これに対し、ニューラルネットワー
クへの入力を高域強調する場合には、話者に共通て、認
識には無関係てありながら、学習に影響を及ぼすスペク
トルの平均的な傾きを補償できるため、学習速度が速く
なるのである。
However, if the average slope of the spectrum is not compensated for and the audio waveform is directly input to the neural network, when the neural network learns, it will extract the feature of the average slope of the spectrum, and the speech will be distorted. It takes time to extract the peaks and valleys of the spectrum necessary for human recognition. On the other hand, when emphasizing the high frequencies of the input to a neural network, it is possible to compensate for the average slope of the spectrum, which is common to all speakers and has nothing to do with recognition, but which affects learning, which speeds up the learning process. becomes faster.

上記実施例によれば、下記■、■の作用効果かある。According to the above embodiment, there are the following effects (1) and (2).

■全登録話者を学習対象として構築した全話者用ニュー
ラルネットワーク13と、類似話者にてグループ化され
た登録話者を学習対象として構築した各グループ用ニュ
ーラルネットワーク14とを用いることにより1話者認
識率を向上できる。
■ By using a neural network 13 for all speakers constructed with all registered speakers as learning targets, and a neural network 14 for each group constructed with registered speakers grouped by similar speakers as learning targets. Speaker recognition rate can be improved.

これにより、全登録話者を学習対象として構築されたニ
ューラルネットワークの出力パターンが入力話者を特定
できないパターンである場合にも、正確に話者認識を行
なうことかできる。
As a result, even if the output pattern of a neural network constructed using all registered speakers as learning targets is a pattern in which the input speaker cannot be identified, accurate speaker recognition can be performed.

■ニューラルネットワーク13.14への入力として、
「音声の一定時間内における平均的な周波数特性の時間
的変化」等、前述■〜■の各要素のうちの1つ以上を用
いるから、入力を得るための前処理が単純となり、この
前処理に要する時間が短くて足りるため、話者認識シス
テム10を複雑な処理装置によることなく容易に実時間
処理できる。
■As an input to neural network 13.14,
Since one or more of the above-mentioned elements (■ to ■), such as "temporal changes in the average frequency characteristics within a certain period of time", are used, the preprocessing to obtain input is simple; Since the time required for processing is short, the speaker recognition system 10 can be easily processed in real time without using a complicated processing device.

[発明の効果] 以上のように本発明によれば、全登録話者を学習対象と
して構築されたニューラルネットワークの出力パターン
が入力話者を特定できないパターンである場合にも、正
確に話者認識を行なうことかできる。
[Effects of the Invention] As described above, according to the present invention, even when the output pattern of a neural network constructed using all registered speakers as learning targets is a pattern in which the input speaker cannot be identified, speaker recognition can be performed accurately. It is possible to do this.

【図面の簡単な説明】[Brief explanation of drawings]

第1図は全話者用ニューラルネットワークの学習系統を
示すブロック図、第2図はグループ用ニューラルネット
ワークの学習系統を示すブロック図、第3図は話者認識
系統を示すブロック図、第4図は話者認識系統を示す流
れ図である。 10・・・話者認識システム、 11・・・音声入力部、 12・・・前処理部、 13・・・全話者用ニューラルネットワーク、14・・
・グループ用ニューラルネットワーク、21・・・第1
判定部、 22・・・グループ記憶部、 23・・・ネットワーク選択部、 24・・・最終判定部。 特許出願人 積水化学工業株式会社 代表者 廣 1) 馨
Figure 1 is a block diagram showing the learning system of the neural network for all speakers, Figure 2 is a block diagram showing the learning system of the neural network for groups, Figure 3 is a block diagram showing the speaker recognition system, and Figure 4 is a block diagram showing the learning system of the neural network for all speakers. is a flowchart showing a speaker recognition system. DESCRIPTION OF SYMBOLS 10...Speaker recognition system, 11...Voice input unit, 12...Preprocessing unit, 13...Neural network for all speakers, 14...
・Group neural network, 21...1st
Judgment unit, 22...Group storage unit, 23...Network selection unit, 24...Final judgment unit. Patent applicant: Sekisui Chemical Co., Ltd. Representative Hiroshi 1) Kaoru

Claims (2)

【特許請求の範囲】[Claims] (1)ニューラルネットワークを用いた話者認識システ
ムにおいて、全登録話者を学習対象として構築した全話
者用ニューラルネットワークと、全話者用ニューラルネ
ットワークを構成する各出力ユニットの出力値が互いに
一定の類似関係にありしきい値判定できないことに基づ
いて、全登録話者をグループ分けした各グループ毎の登
録話者を学習対象として構築した各グループ用ニューラ
ルネットワークとを用い、入力話者について、全話者用
ニューラルネットワークによる認識を行ない、該全話者
用ニューラルネットワークの出力パターンが入力話者を
特定できないパターンである時、更に、該全話者用ニュ
ーラルネットワークにおける最大出力値の出力ユニット
に対応する登録話者を含むグループの、グループ用ニュ
ーラルネットワークによる認識を行なうことを特徴とす
る話者認識システム。
(1) In a speaker recognition system using a neural network, the output values of the all-speaker neural network built with all registered speakers as learning targets and the output units that make up the all-speakers neural network are mutually constant. Based on the fact that the threshold value cannot be determined due to the similarity relationship, all registered speakers are divided into groups, and a neural network for each group is constructed using the registered speakers of each group as learning targets. When recognition is performed using a neural network for all speakers, and when the output pattern of the neural network for all speakers is a pattern that cannot identify the input speaker, the output unit of the maximum output value in the neural network for all speakers is A speaker recognition system characterized in that a group including corresponding registered speakers is recognized by a group neural network.
(2)前記ニューラルネットワークへの入力として、 [1]音声の周波数特性の時間的変化、 [2]音声の平均的な線形予測係数、 [3]音声の平均的なPARCOR係数、 [4]音声の平均的な周波数特性、及びピッチ周波数、 [5]高域強調を施された音声波形の平均的な周波数特
性、並びに [6]音声の平均的な周波数特性 のうちの1つ以上を使用する請求項1に記載の話者認識
システム。
(2) As inputs to the neural network, [1] Temporal changes in the frequency characteristics of speech, [2] Average linear prediction coefficients of speech, [3] Average PARCOR coefficients of speech, [4] speech using one or more of the following: average frequency characteristics and pitch frequency; [5] average frequency characteristics of high-frequency emphasized audio waveform; and [6] average frequency characteristics of audio. The speaker recognition system according to claim 1.
JP2120867A 1990-05-09 1990-05-09 Speaker recognition system Pending JPH0415700A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2120867A JPH0415700A (en) 1990-05-09 1990-05-09 Speaker recognition system

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP2120867A JPH0415700A (en) 1990-05-09 1990-05-09 Speaker recognition system

Publications (1)

Publication Number Publication Date
JPH0415700A true JPH0415700A (en) 1992-01-21

Family

ID=14796922

Family Applications (1)

Application Number Title Priority Date Filing Date
JP2120867A Pending JPH0415700A (en) 1990-05-09 1990-05-09 Speaker recognition system

Country Status (1)

Country Link
JP (1) JPH0415700A (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6182037B1 (en) 1997-05-06 2001-01-30 International Business Machines Corporation Speaker recognition over large population with fast and detailed matches

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6182037B1 (en) 1997-05-06 2001-01-30 International Business Machines Corporation Speaker recognition over large population with fast and detailed matches

Similar Documents

Publication Publication Date Title
Su et al. Performance analysis of multiple aggregated acoustic features for environment sound classification
CN103106903B (en) Single channel blind source separation method
CN110728360A (en) Micro-energy device energy identification method based on BP neural network
US5963904A (en) Phoneme dividing method using multilevel neural network
JPH03273722A (en) Sound/modem signal identifying circuit
WO2006000103A1 (en) Spiking neural network and use thereof
CN116682444A (en) Single-channel voice enhancement method based on waveform spectrum fusion network
JPH06161496A (en) Voice recognition system for recognizing remote control command words for home appliances
CN120388575A (en) A speech recognition method and system based on artificial intelligence
JP3887028B2 (en) Signal source characterization system
Suh et al. Phoneme segmentation of continuous speech using multi-layer perceptron
TWI749547B (en) Speech enhancement system based on deep learning
Krishnakumar et al. A comparison of boosted deep neural networks for voice activity detection
Shaltaf Neural-network-based time-delay estimation
JPH0415700A (en) Speaker recognition system
Ryu et al. Microphone conversion: Mitigating device variability in sound event classification
CA2045612A1 (en) Time series association learning
CN119537877A (en) Radar emitter individual recognition method based on incremental neural network
JPH0415695A (en) Word recognition system
JPH0415694A (en) Word recognition system
JPH03157697A (en) Word recognizing system
Khan et al. Hybrid BiLSTM-HMM based event detection and classification system for food intake recognition
JPH0462599A (en) Noise removing device
JPH03175498A (en) Speaker collating system
Park et al. Advancing Temporal Spike Encoding for Efficient Speech Recognition