JPH0632003B2 - Word registration method - Google Patents

Word registration method

Info

Publication number
JPH0632003B2
JPH0632003B2 JP62202546A JP20254687A JPH0632003B2 JP H0632003 B2 JPH0632003 B2 JP H0632003B2 JP 62202546 A JP62202546 A JP 62202546A JP 20254687 A JP20254687 A JP 20254687A JP H0632003 B2 JPH0632003 B2 JP H0632003B2
Authority
JP
Japan
Prior art keywords
standard pattern
pattern
word
phoneme class
voice
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Expired - Lifetime
Application number
JP62202546A
Other languages
Japanese (ja)
Other versions
JPS6444493A (en
Inventor
隆夫 渡辺
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NEC Corp
Original Assignee
Nippon Electric Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Electric Co Ltd filed Critical Nippon Electric Co Ltd
Priority to JP62202546A priority Critical patent/JPH0632003B2/en
Publication of JPS6444493A publication Critical patent/JPS6444493A/en
Publication of JPH0632003B2 publication Critical patent/JPH0632003B2/en
Anticipated expiration legal-status Critical
Expired - Lifetime legal-status Critical Current

Links

Description

【発明の詳細な説明】 (産業上の利用分野) 本願発明は特定話者の単語音声認識装置に適用する単語
登録方法の改良に関する。
Description: TECHNICAL FIELD The present invention relates to an improvement in a word registration method applied to a specific speaker's word voice recognition device.

(従来の技術とその問題点) 特定話者の単語音声認識装置では、あらかじめ登録され
た標準パタンの各々と入力された未知音声のパタンを比
較して最大の類似度を与える単語を選択することによっ
て認識が行われる。標準パタンの登録においては、単語
音声の区間を切り出すことが必要であり、音声信号の振
幅パラメータを用いた端点検出が行われる。しかしなが
ら、使用者の舌打ち音や息音や周囲雑音のため、単語音
声の区間を誤って登録することがあり、誤認識の原因と
なっている。
(Prior art and its problems) In a specific speaker's word voice recognition device, each standard pattern registered in advance is compared with an input unknown voice pattern to select a word giving the maximum similarity. Recognition is performed by. In registering a standard pattern, it is necessary to cut out a word voice section, and end point detection is performed using an amplitude parameter of a voice signal. However, because of the user's tongue-clicking sound, breathing sound, and ambient noise, the word voice section may be erroneously registered, which causes erroneous recognition.

本願発明は、登録する単語の音声学的な構造を考慮して
単語音声の端点を検出することによって、より正確な単
語検出を実現し、これによって認識精度の高い音声認識
を実現することを目的とする。
An object of the present invention is to realize more accurate word detection by detecting the end points of word speech in consideration of the phonetic structure of a word to be registered, and thereby to realize speech recognition with high recognition accuracy. And

(問題点を解決するための手段) 前述の問題点を解決するために本願の第1の発明が提供
する単語登録方法は、使用者の発声した登録音声を標準
パタンとして登録するに際して、環境雑音パタン及び不
特定話者の音素クラス標準パタンを記憶する手段を有
し、音素クラス系列として表現された登録単語に対し
て、前記音素クラス標準パタンを連結することによって
結合標準パタンとして仮の単語標準パタンを生成し、前
記結合標準パタンの前後に前記環境雑音パタンを連結し
たパタンと登録音声との時間軸対応付けによるマッチン
グを行って、結合標準パタンと環境雑音パタンとの境界
位置に対応する登録音声上の位置を抽出することによっ
て登録音声の端点を検出することを特徴とする。
(Means for Solving Problems) A word registration method provided by the first invention of the present application in order to solve the above problems is such that when registering a registered voice uttered by a user as a standard pattern, environmental noise is generated. A temporary word standard as a combined standard pattern by connecting the phoneme class standard pattern to a registered word expressed as a phoneme class sequence, having means for storing the pattern and the phoneme class standard pattern of the unspecified speaker. A pattern is generated, registration is performed corresponding to the boundary position between the combined standard pattern and the environmental noise pattern by performing matching based on the time axis correspondence between the pattern in which the environmental noise pattern is connected before and after the combined standard pattern and the registered voice. The feature is that the end point of the registered voice is detected by extracting the position on the voice.

また、本願の第2の発明が提供する単語登録方法は、使
用者の発声した登録音声を標準パタンとして登録するに
際して、環境雑音パタン及び音素クラス位置の情報を付
加されている不特定話者の基本単語標準パタンを記憶す
る手段を有し、あらかじめ使用者の発声した基本単語音
声と前記基本単語標準パタンの前後に前記環境雑音パタ
ンを連結したパタンとの時間軸対応付けによるマッチン
グを行って、基本単語標準パタンの音素クラス位置に対
応する基本単語音声の音素クラス位置を抽出することに
より音素クラス標準パタンを求め、音素クラス系列とし
て表現された登録単語に対して、前記音素クラス標準パ
タンを連結することによって結合標準パタンとして仮の
標準パタンを生成し、結合標準パタンの前後に前記環境
雑音パタンを連結したパタンと登録音声との時間軸対応
付けによるマッチングを行って、結合標準パタンと環境
雑音標準パタンとの境界位置に対応する登録音声上の位
置を抽出することによって登録音声の端点を検出するこ
とを特徴とする。
In addition, the word registration method provided by the second invention of the present application, when registering the registered voice uttered by the user as a standard pattern, an unspecified speaker to which the environmental noise pattern and the phoneme class position information are added. Having a means for storing the basic word standard pattern, by performing matching by time axis correspondence between the basic word voice uttered by the user in advance and the pattern in which the environmental noise pattern is connected before and after the basic word standard pattern, The phoneme class standard pattern is obtained by extracting the phoneme class position of the basic word speech corresponding to the phoneme class position of the basic word standard pattern, and the phoneme class standard pattern is concatenated with the registered word expressed as a phoneme class sequence. By generating a temporary standard pattern as a combined standard pattern, and connecting the environmental noise pattern before and after the combined standard pattern. And detecting the end point of the registered voice by extracting the position on the registered voice corresponding to the boundary position between the combined standard pattern and the environmental noise standard pattern by performing matching based on the time axis correspondence between the pattern and the registered voice. Is characterized by.

(作用) 本願発明の基本的な原理を以下に説明する。(Operation) The basic principle of the present invention will be described below.

単語を認識装置に登録するためには、発声された単語音
声の始点と終点(端点)を検出することが必要であり、
このため、一般に、音声信号の振幅パラメータが用いら
れるが、発声された単語音声の前後に舌打ち音、息音、
背景雑音が付加されている場合には、端点の検出を誤る
ことがある。本願発明では、このような雑音の存在する
場合にも正確に単語の端点を検出するため、単語の音声
学的な構造を利用する。ここでは、単語の音声学的な表
現として、単語音素クラスの系列として表現する方法を
用いる。音素クラスとは、一般に、/a/,/i/,/u/,/e/,/o
/,/p/,/k/,/t/,/s/,/z/,/n/,/m/等の音素を指すが、複
数の音素を1つのクラスにまとめたものであってもよい
(例えば、/p,k,t/は無声破裂音、/s,z/は
摩擦音として1つの音素クラスにまとめることが可能で
ある)。
In order to register a word in the recognition device, it is necessary to detect the start point and end point (end point) of the spoken word voice,
For this reason, generally, the amplitude parameter of the voice signal is used, but before and after the uttered word voice, a fluttering sound, a breath sound,
If background noise is added, the end points may be erroneously detected. The present invention utilizes the phonetic structure of a word in order to accurately detect the end points of the word even in the presence of such noise. Here, as a phonetic expression of a word, a method of expressing it as a series of word phoneme classes is used. Phoneme classes are generally / a /, / i /, / u /, / e /, / o
It refers to phonemes such as /, / p /, / k /, / t /, / s /, / z /, / n /, / m / etc., but it is a group of multiple phonemes. (For example, / p, k, t / can be combined into one phoneme class as unvoiced plosives and / s, z / as fricatives).

発声される単語の音素クラス表現が既知である場合に
は、音素クラス標準パタンを用いて次のようにして単語
区間を検出することができる。発声される単語の音素ク
ラス表現(即ち、音素クラス系列)を {s1,s2,…,sk,…,sk} とする。また、各音素クラスsについて特徴ベクトルと
して表される音素クラス標準パタンX(s)が与えられてい
るものとする。このとき、音素クラス標準パタンを連結
して単語の結合標準パタンXを生成する。
If the phoneme class representation of the spoken word is known, the phoneme class standard pattern can be used to detect word intervals as follows. Let the phoneme class representation (ie, the phoneme class sequence) of the spoken word be {s 1 , s 2 , ..., S k , ..., S k }. Further, it is assumed that a phoneme class standard pattern X (s) represented as a feature vector is given for each phoneme class s. At this time, the phoneme class standard patterns are connected to generate a combined standard pattern X of words.

X={X(s1),…,X(s1),X(s2),…,X(s2,…, X(sk),…,X(sk)} ここで、各音素クラス標準パタンX(sk)はそれぞれ固有
の個数が並べられる。さらに、環境雑音パタンを特徴ベ
クトルNで表すものとし、結合標準パタンの前後に環境
雑音パタンを付加した拡張結合標準パタンX*を考える。
X = {X (s 1 ), ..., X (s 1 ), X (s 2 ), ..., X (s 2 , ..., X (s k ), ..., X (s k )} where each A specific number of phoneme class standard patterns X (s k ) are arranged, and the environmental noise pattern is represented by a feature vector N, and an extended joint standard pattern X * in which the ambient noise pattern is added before and after the joint standard pattern . think of.

X*={N,X(s1),…,X(s1),X(s2),…,X(S2),…, X(sk),…,X(sk),N} ={X*(1),X*(2),…,X
*(i),…,X*(1)} 環境雑音パタンは、環境雑音の平均的なパタンであり、
発声する前に一定時間の入力を観測、平均化することに
よって得ることができる。一方、入力される単語音声を
特徴ベクトルの系列Y={Y(1),Y(2),…,Y(J)}とす
る。2つの特徴ベクトル系列パタンX*,Yはよく知られ
ている動的計画法を利用したDPマッチングにより時間
軸を非線形的に対応付けることができる(DPマッチン
グの原理は日本音響学会誌Vol.27,No.9,p.483参照)。
即ち、ベクトル間距離 d(i,j)=distance(X*(i),Y(j)) に関する累積量g(i,j)について次の漸化式を解けばよ
い。
X * = {N, X (s 1 ), ..., X (s 1 ), X (s 2 ), ..., X (S 2 ), ..., X (s k ), ..., X (s k ), N} = {X * (1), X * (2), ..., X
* (i), ..., X * (1)} The environmental noise pattern is an average pattern of environmental noise,
It can be obtained by observing and averaging inputs for a certain period of time before uttering. On the other hand, the input word speech is set as a feature vector sequence Y = {Y (1), Y (2), ..., Y (J)}. The two feature vector series patterns X * and Y can be associated with the time axis non-linearly by DP matching using the well-known dynamic programming (The principle of DP matching is the Acoustical Society of Japan Vol.27, No. 9, p. 483).
That is, the following recurrence formula may be solved for the cumulative amount g (i, j) regarding the inter-vector distance d (i, j) = distance (X * (i), Y (j)).

初期条件:g(0,0)=0,g(i,0)=∞(i=1,…,I), g(0,j)=∞(j=1,…,J) 漸化式: この漸化式計算は入力音声の時間軸に沿って実行するこ
とができ、時間軸の対応付けを求めるには仮の終点(i,
j)からマッチング経路をh(i,j)を用いてトレースバック
すればよい。入力音声の仮の終点Jは、振幅を用いて音
声の存在する可能性のある区間を求める、発声を終えた
時点を使用者がマニュアルで指定する等の方法により決
定することができる。第1図にDPマッチングによる時
間軸対応付け結果の例を示す。図においてIS,IEは
それぞれ結合標準パタンの始点、終点であり、これらに
対応する位置JS,JEが単語音声の始点、終点とな
る。このようにして決定された端点は、入力音声を単語
に先行する雑音部分、これに続く音素クラスの系列から
構成される単語部分、単語に後続する雑音部分として解
釈した結果として得られたものであり、単に振幅の大小
によって決定された端点に比べて正確なものである。
Initial condition: g (0,0) = 0, g (i, 0) = ∞ (i = 1, ..., I), g (0, j) = ∞ (j = 1, ..., J) Recurrence formula : This recurrence formula calculation can be executed along the time axis of the input voice, and the temporary end point (i,
The matching path from j) can be traced back using h (i, j). The tentative end point J of the input voice can be determined by methods such as obtaining a section in which the voice may exist using the amplitude and manually designating the time when the utterance is finished by the user. FIG. 1 shows an example of the time axis correspondence result by DP matching. In the figure, I S and I E are the start point and end point of the combined standard pattern, and the positions J S and J E corresponding to these are the start point and end point of the word speech. The end points determined in this way are obtained as a result of interpreting the input speech as a noise part preceding a word, a word part consisting of a sequence of phoneme classes following the word, and a noise part following the word. Yes, it is more accurate than the end point determined simply by the magnitude of the amplitude.

音素クラス標準パタンについては不特定話者の標準パタ
ンを用いる方法と、特定話者の標準パタンを用いる方法
の2通りが考えられる。不特定話者標準パタンを用いる
場合には、多数の話者の音声サンプルから抽出した音素
クラスパタンを平均化する方法、クラスタリングによっ
て複数個の標準パタンを求める方法などによって、あら
かじめ音素クラス標準パタンを求めておくことができ
る。特定話者標準パタンを用いる場合には、不特定話者
標準パタンを用いる方法より更に精度の良い端点の検出
を行うことができる。この場合には、あらかじめ音素ク
ラス標準パタンを求めておくことができないため、単語
の登録に先だって、話者の音素クラス標準パタンの作成
を以下の手順によって行う。
There are two possible phoneme class standard patterns: a method using a standard pattern of an unspecified speaker and a method using a standard pattern of a specific speaker. When using the unspecified speaker standard pattern, the phoneme class standard pattern is previously determined by a method of averaging phoneme class patterns extracted from a large number of speaker voice samples or a method of obtaining a plurality of standard patterns by clustering. You can ask for it. When the specific speaker standard pattern is used, it is possible to detect the end points with higher accuracy than the method using the unspecified speaker standard pattern. In this case, since the phoneme class standard pattern cannot be obtained in advance, the speaker phoneme class standard pattern is created by the following procedure prior to the word registration.

まず、あらかじめ基本となる基本単語セットを定義して
おく。この基本単語セットはすべての音素クラスを含む
ようにし、これに対応する不特定話者の単語標準パタン
を用意しておく。基本単語の標準パタンはあらかじめ多
数話者の基本単語の音声サンプルから平均化、クラスタ
リングなどの手法によって作成することが可能である。
ここで基本単語の標準パタン上で音素クラスの位置を指
定しておく。第2図にその一例を示す。話者の発声した
基本単語音声と不特定話者の基本単語標準パタンとの時
間軸対応付けをDPマッチングにより実行する。このと
き、既に説明したように、環境雑音パタンを標準パタン
の前後に付加しておくことによって単語区間の検出が確
実なものとなる。DPマッチングの結果から、話者の発
声した音声上での音素クラス位置が決定される。これら
の位置にある特徴ベクトルをその話者の音素クラス標準
パタンとすることができる。(1つの音素クラスに対し
て複数個の位置が存在する場合には、ここでも、平均
化、クラスタリング等の手法を利用できる) 本説明では、音素クラス標準パタン、環境雑音パタンは
共に単一のベクトルとしたが、一般には複数のベクトル
の集合とすることが可能である。即ち、ベクトル間距離
の計算において、ある音素クラスに対して標準パタンの
集合が X(s)={X* 1,X* 2,…,X* m,…,X* M} として与えられている場合には としてベクトル間距離を計算すれば良い。ただし、ここ
でyは入力音声の特徴ベクトルである。
First, a basic basic word set is defined in advance. This basic word set should include all phoneme classes, and the standard word patterns of the unspecified speaker corresponding to this should be prepared. The standard pattern of basic words can be created in advance from speech samples of the basic words of many speakers by means of averaging, clustering, or the like.
Here, the position of the phoneme class is specified on the standard pattern of the basic word. FIG. 2 shows an example thereof. The time base correspondence between the basic word voice uttered by the speaker and the standard word standard pattern of the unspecified speaker is executed by DP matching. At this time, as already described, by adding the environmental noise pattern before and after the standard pattern, the detection of the word section becomes reliable. The phoneme class position on the voice uttered by the speaker is determined from the result of the DP matching. The feature vectors at these positions can be the phoneme class standard pattern of the speaker. (When there are a plurality of positions for one phoneme class, methods such as averaging and clustering can be used here as well.) In this description, the phoneme class standard pattern and the environmental noise pattern are both single. Although it is a vector, it can be generally a set of a plurality of vectors. That is, in the calculation of the distance between vectors, a set of standard patterns is given as X (s) = {X * 1 , X * 2 , ,, X * m , ..., X * M } for a certain phoneme class. If there is The vector distance may be calculated as Here, y is the feature vector of the input voice.

(実施例) 第3図は本願の第1の発明の一実施例を示すブロック図
である。参照数字1は音声分析部であり、入力された音
声信号が特徴ベクトルの系列に変換され出力される。こ
こで、特徴ベクトルの表現としてスペクトラム、ケプス
トラム、バンドパスフィルタ等任意のものが可能であ
る。2は切り替えスイッチであり、環境雑音パタンを登
録する時は端子Aに、標準パタン作成のため単語登録を
行う時は端子Cに、認識時には端子Dに接続される。
(Embodiment) FIG. 3 is a block diagram showing an embodiment of the first invention of the present application. Reference numeral 1 is a voice analysis unit, which converts an input voice signal into a series of feature vectors and outputs the sequence. Here, an arbitrary expression such as a spectrum, a cepstrum, or a bandpass filter is possible as the expression of the feature vector. Reference numeral 2 denotes a changeover switch, which is connected to a terminal A when registering an environmental noise pattern, a terminal C when registering a word for creating a standard pattern, and a terminal D when recognizing.

まず最初に、スイッチ2においてAが選択され、環境雑
音パタンの登録が行われる。無音状態において音声分析
部により分析された結果である特徴ベクトルが環境雑音
パタン記憶部3に格納される。6は音素クラス標準パタ
ン記憶部で、あらかじめ多数話者の音声サンプルから作
成された音素クラス標準パタンが格納されている。7は
登録単語リスト記憶部で、認識対象語いの各々について
の音素クラス系列記述が格納される。単語の登録時に
は、スイッチ2においてCが選択される。ここでは、ま
ず、7に格納されている登録単語の音素クラス系列に基
づいて、記憶部6に格納されている音素クラス標準パタ
ンを連結することによって結合標準パタンを作成し、更
に、記憶部3から読みだされた環境雑音パタンを結合標
準パタンの前後に付加し拡張結合標準パタンが作成さ
れ、バッファ8に格納される。処理部9はバッファ8か
ら読みだされた拡張結合標準パタンと、スイッチ2Cを
経由して音声分析部1から送出されてきた登録単語音声
とのDPマッチングを実行し、時間軸対応を求め、入力
音声の始点、終点を決定し区間内の特徴ベクトル系列を
標準パタンとして記憶部10に格納する。
First, A is selected in the switch 2, and the environmental noise pattern is registered. The feature vector, which is the result of the analysis performed by the voice analysis unit in the silent state, is stored in the environmental noise pattern storage unit 3. A phoneme class standard pattern storage unit 6 stores phoneme class standard patterns created in advance from voice samples of many speakers. A registered word list storage unit 7 stores phoneme class sequence descriptions for each recognition target word. When registering a word, the switch 2 selects C. Here, first, based on the phoneme class sequence of the registered word stored in 7, the phoneme class standard patterns stored in the storage unit 6 are concatenated to create a combined standard pattern, and further, the storage unit 3 The ambient noise pattern read from the above is added before and after the combined standard pattern to create an expanded combined standard pattern, which is stored in the buffer 8. The processing unit 9 executes DP matching between the extended combined standard pattern read from the buffer 8 and the registered word voice sent from the voice analysis unit 1 via the switch 2C, obtains the time axis correspondence, and inputs. The start point and the end point of the voice are determined, and the feature vector series in the section is stored in the storage unit 10 as a standard pattern.

認識時には、スイッチ2においてDが選択され、記憶部
10から読みだされた標準パタンを用いて認識部11におい
て認識処理が行われる。
At the time of recognition, D is selected by the switch 2 and the storage unit
The recognition processing is performed in the recognition unit 11 using the standard pattern read from 10.

第4図は本願の第2の発明の一実施例を示すブロック図
である。参照数字21は音声分析部であり、入力された音
声信号が特徴ベクトルの系列に変換され出力される。こ
こで、特徴ベクトルの表現としてスペクトラム、ケプス
トラム、バンドパスフィルタ等任意のものが可能であ
る。22は切り替えスイッチであり、環境雑音パタンを登
録する時は端子Aに、音素標準パタン生成のため基本単
語を入力する時は端子Bに、標準パタン作成のため単語
登録を行う時は端子Cに、認識時には端子Dに接続され
る。
FIG. 4 is a block diagram showing an embodiment of the second invention of the present application. Reference numeral 21 is a voice analysis unit, which converts the input voice signal into a series of feature vectors and outputs the sequence. Here, an arbitrary expression such as a spectrum, a cepstrum, or a bandpass filter is possible as the expression of the feature vector. Reference numeral 22 denotes a changeover switch. When registering an environmental noise pattern, terminal A is used. When a basic word is input to generate a phoneme standard pattern, terminal B is used. When a standard pattern is created, terminal C is used. , At the time of recognition, it is connected to the terminal D.

まず最初に、スイッチ22においてAが選択され、環境雑
音パタンの登録が行われる。無音状態で入力され、音声
分析部21により分析された結果である特徴ベクトルが環
境雑音パタン記憶部23に格納される。24は不特定話者の
基本単語の標準パタンを保持する記憶部で、この標準パ
タンはあらかじめ多数話者のサンプルから作成されたも
のを搭載しておく。続いて、スイッチ22においてBが選
択され音素クラス標準パタンの生成が行われる。25は音
素クラス標準パタン抽出部であり、記憶部24から読みだ
された不特定話者の基本単語標準パタンの前後に記憶部
23から読みだされた環境雑音パタンを付加したパタン
と、スイッチ22Bを経由して音声分析部21から送出され
てきた基本単語音声とのDPマツチングが実行される。
この結果得られた2つのパタンの時間軸の対応関係から
入力音声上の音素クラス位置が決定され、その位置上の
特徴ベクトルが音素クラス標準パタンとして記憶部26に
格納される。27は登録単語リスト記憶部で、認識対象語
いの各々についての音素クラス系列記述が格納される。
単語の登録の時には、スイッチ22においてCが選択され
る。ここでは、まず、27に格納されている登録単語の音
素クラス系列に基づいて、記憶部26に格納されている音
素クラス標準パタンを連結することによって結合標準パ
タンを作成し、更に、記憶部23から読みだされた環境雑
音パタンを結合標準パタンの前後に付加し拡張結合標準
パタンが作成され、バッファ28に格納される。処理部29
はバッファ28から読みだされた拡張結合標準パタンと、
スイッチ22Cを経由して音声分析部21から送出されてき
た登録単語音声とのDPマッチングを実行し、時間軸対
応を求め、入力音声の始点、終点を決定し区間内の特徴
ベクトル系列を標準パタンとして記憶部30に格納する。
First, A is selected by the switch 22 and the environmental noise pattern is registered. A feature vector that is input in a silent state and analyzed by the voice analysis unit 21 is stored in the environmental noise pattern storage unit 23. Reference numeral 24 is a storage unit that holds a standard pattern of basic words of unspecified speakers, and this standard pattern is prepared in advance from a sample of many speakers. Then, B is selected by the switch 22 and a phoneme class standard pattern is generated. Reference numeral 25 denotes a phoneme class standard pattern extraction unit, which is provided before and after the basic word standard pattern of the unspecified speaker read from the storage unit 24.
DP matching of the pattern added with the environmental noise pattern read from 23 and the basic word voice sent from the voice analysis unit 21 via the switch 22B is executed.
The phoneme class position on the input speech is determined from the correspondence of the time axes of the two patterns obtained as a result, and the feature vector on the position is stored in the storage unit 26 as a phoneme class standard pattern. A registered word list storage unit 27 stores a phoneme class sequence description for each recognition target vocabulary.
When registering a word, the switch 22 selects C. Here, first, based on the phoneme class sequence of the registered word stored in 27, a phoneme class standard pattern stored in the storage unit 26 is connected to create a combined standard pattern, and further, the storage unit 23 The environmental noise pattern read from the above is added before and after the combined standard pattern to create an expanded combined standard pattern, which is stored in the buffer 28. Processing unit 29
Is the extended combined standard pattern read from buffer 28,
DP matching is performed with the registered word voice sent from the voice analysis unit 21 via the switch 22C, the time axis correspondence is obtained, the start point and end point of the input voice are determined, and the feature vector series in the section is standardized. Is stored in the storage unit 30 as

認識時には、スイッチ22においてDが選択され、記憶部
30から読みだされた標準パタンを用いて認識部31におい
て認識処理が行われる。
At the time of recognition, D is selected by the switch 22 and the storage unit
The recognition processing is performed in the recognition unit 31 using the standard pattern read from 30.

以上に於て、DPマッチングにおける漸化式は本発明で
述べたものに限定されるものではない。また、認識部1
1,31はDPマッチング以外にも任意の方式が利用可能
である。
In the above, the recurrence formula in DP matching is not limited to the one described in the present invention. Also, the recognition unit 1
As for 1 and 31, any method other than DP matching can be used.

(発明の効果) 以上に述べたように本願発明によれば、登録する単語の
音声学的な構造を考慮した単語音声の端点の検出が行わ
れるので、より正確な単語検出を実現し、これによって
認識精度の高い音声認識を実現することが可能となる。
(Effect of the invention) As described above, according to the present invention, since the end point of the word voice is detected in consideration of the phonetic structure of the word to be registered, more accurate word detection is realized. Thus, it becomes possible to realize voice recognition with high recognition accuracy.

【図面の簡単な説明】[Brief description of drawings]

第1図はDPマッチングによる時間軸の対応付けの一例
を示す概念図、第2図は不特定話者の単語標準パタンの
一例を示す概念図、第3図および第4図は本願発明の実
施例を示すブロック図である。 これら図において、参照数字1,21は音声分析部、2,
22は切り替えスイッチ、3,23は環境雑音パタン記憶
部、24は不特定話者単語標準パタン記憶部、25は音素ク
ラス標準パタン抽出部、6,26は音素クラス標準パタン
記憶部、7,27は登録単語リスト記憶部、8,28は結合
標準パタンバッファ、9,29は処理部(単語標準パタン
登録部)、10,30は単語標準パタン記憶部、11,31は認
識部である。
FIG. 1 is a conceptual diagram showing an example of time axis correspondence by DP matching, FIG. 2 is a conceptual diagram showing an example of a word standard pattern of an unspecified speaker, and FIGS. 3 and 4 are implementations of the present invention. It is a block diagram which shows an example. In these figures, reference numerals 1 and 21 are voice analysis units and 2,
22 is a changeover switch, 3 and 23 are environmental noise pattern storage units, 24 is an unspecified speaker word standard pattern storage unit, 25 is a phoneme class standard pattern extraction unit, 6 and 26 are phoneme class standard pattern storage units, 7 and 27. Is a registered word list storage unit, 8 and 28 are combined standard pattern buffers, 9 and 29 are processing units (word standard pattern registration units), 10 and 30 are word standard pattern storage units, and 11 and 31 are recognition units.

Claims (2)

【特許請求の範囲】[Claims] 【請求項1】使用者の発声した登録音声を標準パタンと
して登録するに際して、環境雑音パタン及び不特定話者
の音素クラス標準パタンを記憶する手段を有し、音素ク
ラス系列として表現された登録単語に対して、前記音素
クラス標準パタンを連結することによって結合標準パタ
ンとして仮の単語標準パタンを生成し、前記結合標準パ
タンの前後に前記環境雑音パタンを連結したパタンと登
録音声との時間軸対応付けによるマッチングを行って、
結合標準パタンと環境雑音パタンとの境界位置に対応す
る登録音声上の位置を抽出することによって登録音声の
端点を検出することを特徴とする単語音声認識における
単語登録方法。
1. A registered word expressed as a phoneme class sequence, having means for storing an environmental noise pattern and a phoneme class standard pattern of an unspecified speaker when registering a registered voice uttered by a user as a standard pattern. On the other hand, a temporary word standard pattern is generated as a combined standard pattern by connecting the phoneme class standard patterns, and the time axis correspondence between the pattern in which the environmental noise pattern is connected before and after the combined standard pattern and the registered voice is generated. Match by attaching,
A word registration method in word voice recognition, characterized in that an end point of a registered voice is detected by extracting a position on the registered voice corresponding to a boundary position between a combined standard pattern and an environmental noise pattern.
【請求項2】使用者の発声した登録音声を標準パタンと
して登録するに際して、環境雑音パタン及び音素クラス
位置の情報を付加されている不特定話者の基本単語標準
パタンを記憶する手段を有し、あらかじめ使用者の発声
した基本単語音声と前記基本単語標準パタンの前後に前
記環境雑音パタンを連結したパタンとの時間軸対応付け
によるマッチングを行って、基本単語標準パタンの音素
クラス位置に対応する基本単語音声の音素クラス位置を
抽出することにより音素クラス標準パタンを求め、音素
クラス系列として表現された登録単語に対して、前記音
素クラス標準パタンを連結することによって結合標準パ
タンとして仮の標準パタンを生成し、結合標準パタンの
前後に前記環境雑音パタンを連結したパタンと登録音声
との時間軸対応付けによるマッチングを行って、結合標
準パタンと環境雑音標準パタンとの境界位置に対応する
登録音声上の位置を抽出することによって登録音声の端
点を検出することを特徴とする単語音声認識における単
語登録方法。
2. When registering a registered voice uttered by a user as a standard pattern, a means is provided for storing a standard pattern of a basic word of an unspecified speaker, to which information on environmental noise patterns and phoneme class positions is added. , By matching the basic word voice uttered by the user in advance and the pattern in which the environmental noise pattern is connected before and after the basic word standard pattern by time-axis correspondence to correspond to the phoneme class position of the basic word standard pattern. A phoneme class standard pattern is obtained by extracting the phoneme class position of the basic word speech, and the phoneme class standard pattern is concatenated to the registered words expressed as a phoneme class sequence to form a temporary standard pattern as a combined standard pattern. Is generated, and the time axis correspondence between the pattern in which the environmental noise pattern is connected before and after the combined standard pattern and the registered voice is generated. A word registration method in word voice recognition characterized by detecting the end points of the registered voice by extracting the position on the registered voice corresponding to the boundary position between the combined standard pattern and the environmental noise standard pattern by performing matching by .
JP62202546A 1987-08-12 1987-08-12 Word registration method Expired - Lifetime JPH0632003B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP62202546A JPH0632003B2 (en) 1987-08-12 1987-08-12 Word registration method

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP62202546A JPH0632003B2 (en) 1987-08-12 1987-08-12 Word registration method

Publications (2)

Publication Number Publication Date
JPS6444493A JPS6444493A (en) 1989-02-16
JPH0632003B2 true JPH0632003B2 (en) 1994-04-27

Family

ID=16459293

Family Applications (1)

Application Number Title Priority Date Filing Date
JP62202546A Expired - Lifetime JPH0632003B2 (en) 1987-08-12 1987-08-12 Word registration method

Country Status (1)

Country Link
JP (1) JPH0632003B2 (en)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH082015A (en) * 1994-06-27 1996-01-09 Nec Corp Printer device

Also Published As

Publication number Publication date
JPS6444493A (en) 1989-02-16

Similar Documents

Publication Publication Date Title
JP3050934B2 (en) Voice recognition method
JP3337233B2 (en) Audio encoding method and apparatus
EP2048655A1 (en) Context sensitive multi-stage speech recognition
JPH0968994A (en) Method of recognizing words by pattern matching and apparatus for implementing the method
Rivlin et al. A phone-dependent confidence measure for utterance rejection
JPS6247320B2 (en)
US20040199385A1 (en) Methods and apparatus for reducing spurious insertions in speech recognition
JP2001166789A (en) Chinese speech recognition method and apparatus using initial / final phoneme similarity vector
Kuamr et al. Implementation and performance evaluation of continuous Hindi speech recognition
JP2996019B2 (en) Voice recognition device
JP2745562B2 (en) Noise adaptive speech recognizer
WO2007114346A1 (en) Speech recognition device
JP2813209B2 (en) Large vocabulary speech recognition device
JPS58108590A (en) Voice recognition equipment
JP3277522B2 (en) Voice recognition method
JP3457578B2 (en) Speech recognition apparatus and method using speech synthesis
KR100677224B1 (en) Speech Recognition Using Anti-Word Model
JP2882791B2 (en) Pattern comparison method
JP3029654B2 (en) Voice recognition device
JP2943473B2 (en) Voice recognition method
JPH0997095A (en) Voice recognition device
JP2004309654A (en) Voice recognition device
JP3357752B2 (en) Pattern matching device
JP2862306B2 (en) Voice recognition device
JPS63217399A (en) Voice section detecting system