JPH0424717B2 - - Google Patents

Info

Publication number
JPH0424717B2
JPH0424717B2 JP59108668A JP10866884A JPH0424717B2 JP H0424717 B2 JPH0424717 B2 JP H0424717B2 JP 59108668 A JP59108668 A JP 59108668A JP 10866884 A JP10866884 A JP 10866884A JP H0424717 B2 JPH0424717 B2 JP H0424717B2
Authority
JP
Japan
Prior art keywords
speech
section
voice
blocks
block
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Expired
Application number
JP59108668A
Other languages
Japanese (ja)
Other versions
JPS60254100A (en
Inventor
Atsuko Hirota
Yutaka Iizuka
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Oki Electric Industry Co Ltd
Original Assignee
Oki Electric Industry Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Oki Electric Industry Co Ltd filed Critical Oki Electric Industry Co Ltd
Priority to JP59108668A priority Critical patent/JPS60254100A/en
Publication of JPS60254100A publication Critical patent/JPS60254100A/en
Publication of JPH0424717B2 publication Critical patent/JPH0424717B2/ja
Granted legal-status Critical Current

Links

Description

【発明の詳細な説明】 (技術分野) 本発明は、音声認識装置に関し、特に精度良く
音声区間の検出を行う音声区間検出に関するもの
である。
DETAILED DESCRIPTION OF THE INVENTION (Technical Field) The present invention relates to a speech recognition device, and particularly to speech section detection for detecting speech sections with high accuracy.

(背景技術) 従来の音声認識装置のブロツク図を第1図に示
す。第1図において、1は信号入力端子、2は周
波数分析部、3は音声取込制御部、4は取込開始
信号、5は音声区間検出部、6は取込終了信号、
7は始端・終端情報、8は認識部、9は出力端子
の如く構成されており、以下各部の説明をする。
(Background Art) A block diagram of a conventional speech recognition device is shown in FIG. In FIG. 1, 1 is a signal input terminal, 2 is a frequency analysis section, 3 is an audio capture control section, 4 is a capture start signal, 5 is a voice section detection section, 6 is a capture end signal,
7 is a starting end/end end information, 8 is a recognition section, and 9 is an output terminal. Each section will be explained below.

周波数分析部2は、第2図に示す如く構成され
ており、入力音声信号11は前置増幅器12によ
り適当なレベルに増幅され、約200Hzから6000Hz
までを対数尺度で等間隔に分割されたN個のバン
ドパスフイルタ群13、全波整流器群14、およ
びローパスフイルタ群15により分析され、さら
に、あらかじめ定められた時間周期(以後サンプ
ル周期と記す)毎にマルチプレクサ16を順次切
り換えながらAD変換器17によつて量子化さ
れ、サンプル周期毎にN個の分析結果18を出力
する。
The frequency analysis unit 2 is configured as shown in FIG.
is analyzed by N band-pass filter groups 13, full-wave rectifier groups 14, and low-pass filter groups 15, which are divided at equal intervals on a logarithmic scale. The data are quantized by the AD converter 17 while sequentially switching the multiplexer 16 at each sample period, and N analysis results 18 are output at each sample period.

音声取込制御部3は、取込開始信号4を受信し
たのち、周波数分析部2の分析結果18を音声区
間検出部5および認識部8へ一定時間、または確
かに音声の入力が終了したと判断するまで出力す
る。音声の入力終了の判断法としては、たとえ
ば、各サンプル周期毎のN個のデータの平均値
(以後フレームパワーと記す)を利用して、、フレ
ームパワーがあらかじめ設定された閾値を越える
ものが、ある一定数存在したのち、閾値を越えな
いものが連続一定数続いたとき音声の入力が終了
したと判断する方法がある。
After receiving the capture start signal 4, the voice capture control unit 3 transmits the analysis result 18 of the frequency analysis unit 2 to the voice section detection unit 5 and the recognition unit 8 for a certain period of time, or when the voice input has certainly finished. Output until a decision is made. As a method for determining the end of audio input, for example, the average value of N pieces of data for each sample period (hereinafter referred to as frame power) is used, and if the frame power exceeds a preset threshold, There is a method of determining that the voice input has ended when a certain number of values do not exceed the threshold value and a certain number of consecutive values do not exceed the threshold value.

音声区間検出部5におけるブロツク図を第3図
に示す。第3図において、18は分析結果、21
はパラメータ演算部、6は取込終了信号、22は
ブロツク化部、23は音声区間判定部、7は始端
終端情報の如く構成され、以下詳細に説明する。
A block diagram of the voice section detection section 5 is shown in FIG. In Figure 3, 18 is the analysis result, 21
Reference numeral numeral 6 indicates a parameter calculation unit, numeral 6 indicates a capture end signal, numeral 22 indicates a blocking unit, numeral 23 indicates a voice section determination unit, and numeral 7 indicates start/end information, which will be described in detail below.

パラメータ演算部21は、分析結果18から音
声区間検出に使用する(1)式で定義されるパラメー
タを求める部分である。
The parameter calculation unit 21 is a part that calculates parameters defined by equation (1) used for voice section detection from the analysis result 18.

Pj=ajj …(1) ただしaj;第j番目の分析結果のスペクタル傾
j;第j番目の分析結果の平均値 また、スペクトル傾斜ajすなわち最少2乗近似
直線の傾きは、第j番目のN個の分析結果をxij
とすると(i;N分割されたバンドパスフイルタ
群の周波数の低いものから順に付けられた番号)、
ajは(2)式によつて求められる。
P j = a jj …(1) where a j ; Spectral slope j of the jth analysis result; Average value of the jth analysis result Also, the spectral slope a j, that is, the slope of the least squares approximation straight line is , the j-th N analysis results are x ij
Assuming that (i; the number assigned in order from the lowest frequency to the group of N-divided bandpass filters),
a j is determined by equation (2).

(2)式においてNを固定すれば、Ni=1 i及びNi=1 i2
定数となり、C1Ni=1 i及び C2=N・Ni=1 i2−{Ni=1 i}2 と置き換えることができ、(2)式は(3)式に変形され
る。
If N is fixed in equation (2), Ni=1 i and Ni=1 i 2 become constants, and C 1 = Ni=1 i and C 2 = N・Ni=1 i 2 −{ Ni=1 i} 2 can be replaced, and equation (2) is transformed into equation (3).

従つて、Ni=1 i・xijNi=1 xijを求めればajを求め
ることができる。
Therefore, a j can be found by finding Ni=1 i・x ij and Ni=1 x ij .

また、jNi=1 xijをNで除すことによつて得ら
れる。第4図は、Pjを演算するブロツク図であ
り、以下図に従つて説明する。
Also, j can be obtained by dividing Ni=1 x ij by N. FIG. 4 is a block diagram for calculating P j , which will be explained below with reference to the figure.

第j番目のN個の分析結果xij(i=1,2,…
N)が順番に出力されるものとすると、加算器1
01およびレジスタ102によつてxijの累積Ni=1
xijをレジスタ102にセツトすることができ、
その結果を乗算器103と除算器106に出力さ
れる。乗算器103でNi=1 xijとC1(=Ni=1 i)との乗
算を行ない、さらに補数器104によつて−
C1Ni=1 xijの値を求め、加算器105の一方に入
力される。また、xijのデータ出力と同期して働
くカウンタ107の出力と、xijとの積i・xij
乗算器108によつて求め、乗算器108の出力
に接続されている加算器109と、さらにそれに
接続されているレジスタ110によつてNi=1 i・
xijを求めることができる。レジスタ110の出
Ni=1 i・xijは乗算器111の一方の入力に接続
されており、乗算器111の他方の入力にはNが
セツトされていて、乗算器111ではN・Ni=1
i・xijが演算され、加算器105のもう一方に
入力される。加算器105では、N・Ni=1 i・xij
が演算され、除算器112に接続されている。除
算器112では、N・Ni=1 i・xij−C1Ni=1 i・xij
をC2で除すことによつて、第j番目のサンプル
データのスペクトル傾斜ajを求められ、その結果
は乗算器113の一方の入力となる。また除算器
106では、Ni=1 xijをNで除すことによつてj
求められ、その結果は乗算器113の他方の入力
となり、乗算器113によつてPj(=ajj)を
求めることができる。以上の演算をサンプル周期
毎に行なつて、各サンプル時のPjの値を全て演算
することができる。
j-th N analysis results x ij (i=1, 2,...
N) are output in order, adder 1
01 and register 102, the cumulative value of x ij is Ni=1
x ij can be set in register 102,
The result is output to multiplier 103 and divider 106. The multiplier 103 multiplies Ni=1 x ij by C 1 (= Ni=1 i), and the complementer 104 multiplies -
The value of C 1 · Ni=1 x ij is determined and input to one side of the adder 105 . Further, the product i x ij of the output of the counter 107 that operates in synchronization with the data output of x ij and x ij is obtained by the multiplier 108, and the product is calculated by the adder 109 connected to the output of the multiplier 108. , furthermore, by the register 110 connected to it, Ni=1 i・
x ij can be found. The output Ni=1 i・x ij of the register 110 is connected to one input of the multiplier 111, and the other input of the multiplier 111 is set to N.i=1
i·x ij is calculated and input to the other side of the adder 105. In the adder 105, N・Ni=1 i・x ij
is calculated and connected to the divider 112. In the divider 112, N・Ni=1 i・x ij −C 1Ni=1 i・x ij
By dividing C 2 by C 2 , the spectral slope a j of the j-th sample data is obtained, and the result becomes one input of the multiplier 113 . Further, in the divider 106, j is obtained by dividing Ni=1 x ij by N, and the result becomes the other input of the multiplier 113 . jj ) can be obtained. By performing the above calculation for each sampling period, all the values of P j at each sampling time can be calculated.

ブロツク化部22は、パラメータ演算部21の
結果Pjを取込終了信号6を検出するまで受け取
り、取込終了信号6を検出後、音声のブロツク化
(音声であると思われる部分のかたまりの検出)
を行なう部分で、第5図にブロツク図を示し、第
5図に従つて説明する。
The blocking unit 22 receives the result P j of the parameter calculation unit 21 until it detects the capture end signal 6, and after detecting the capture end signal 6, blocks the audio (blocks the parts that are considered to be audio). detection)
FIG. 5 shows a block diagram of the portion where this is carried out, and the explanation will be given with reference to FIG.

パラメータ演算部21の各サンプル周期毎のPj
は、順次Pパラメータメモリ200に格納されて
いるので、それを順番に読取し絶対値回路201
によつて絶対値化され、|Pj|を比較器202の
一方に入力する。比較器202の他方の入力に
は、|Pj|の閾値PTHがセツトされている。比較器
202では、|Pj|≧PTHのときにはα出力に、|
Pj|<PTHのときにはβ出力にそれぞれ有意信号
を出力する。カウンタ203は、|Pj|≧PTHのと
きカウントアツプし、|Pj|<PTHのときクリアさ
れるようになつており、|Pj|≧PTHとなる連続量
をカウントする。また、カウンタ203の出力
は、常にレジスタ204にセツトされている。レ
ジスタ204にセツトされている値(|Pj|≧
PTHである連続数)は、比較器205に入力され、
比較器205の他方の入力にはKがセツトされて
おり、|Pj|≧PTHである連続量(以下ブロツク長
と記す)がK以上のとき、比較器205の出力C
に有意信号が出力される。
P j for each sample period of the parameter calculation unit 21
are sequentially stored in the P parameter memory 200, so they are read in order and the absolute value circuit 201
|P j | is input to one side of the comparator 202. At the other input of the comparator 202, a threshold value P TH of |P j | is set. The comparator 202 outputs α when |P j |≧P TH ;
When P j |<P TH , a significant signal is output to each β output. The counter 203 counts up when |P j |≧P TH , and is cleared when |P j |<P TH , and counts the continuous amount that satisfies |P j |≧P TH . Further, the output of the counter 203 is always set in the register 204. The value set in the register 204 (|P j |≧
P TH (consecutive number) is input to the comparator 205,
K is set to the other input of the comparator 205, and when the continuous amount (hereinafter referred to as block length) where |P j |≧P TH is greater than or equal to K, the output C of the comparator 205
A significant signal is output.

ブロツク長がK(K≧2の自然数)以上(C信
号出力時)で、かつ、比較器202のβ出力(|
Pj|<PTH)が表われたタイミングをAND回路2
06によつて捕える。カウンタ207は、AND
回路206の出力から出力までのPjを読み出した
量を数えるもので、減算器208によつてカウン
タ7の出力からレジスタ204の結果(ブロツク
長)を差し引くことにより、ブロツク間の距離
(時間)を求めることができる。またカウンタ2
09は、Pjの読出しと同期してカウントしてお
り、減算器210によつてカウンタ209の結果
からレジスタ204の出力(ブロツク長)を引く
ことによつて、当該ブロツクの先頭を求められ
る。加算器211とレジスタ212により|Pj
≧PTHの部分の累積を求め、ブロツクの大きさを
表わすSBなるものを求め、AND回路206の信
号を検出したとき、レジスタ213にセツトする
と同時に、レジスタ213の出力(以下ブロツク
量と記す)、減算器210の出力(ブロツク先頭
情報)、レジスタ204の出力(ブロツク長)、お
よび減算器208の出力(ブロツク間距離)をブ
ロツクテーブル214に登録する。このようにし
て取込んだ量全てについてブロツク化が行なうこ
とができる。
The block length is K (a natural number of K≧2) or more (when outputting C signal), and the β output of the comparator 202 (|
The timing at which P j |<P TH ) appears is determined by AND circuit 2.
Captured by 06. The counter 207 is AND
It counts the amount of P j read from the output of the circuit 206, and by subtracting the result of the register 204 (block length) from the output of the counter 7 by the subtracter 208, the distance (time) between blocks is calculated. can be found. Also counter 2
09 is counted in synchronization with the reading of Pj , and by subtracting the output (block length) of the register 204 from the result of the counter 209 by the subtracter 210, the head of the block can be determined. By the adder 211 and the register 212, |P j |
≧P TH is calculated, S B representing the block size is calculated, and when the signal of the AND circuit 206 is detected, it is set in the register 213 and at the same time the output of the register 213 (hereinafter referred to as block amount) is calculated. ), the output of the subtracter 210 (block head information), the output of the register 204 (block length), and the output of the subtractor 208 (interblock distance) are registered in the block table 214. In this way, the entire amount taken in can be blocked.

音声区間判定部23は、ブロツク化部22で得
れたブロツクテーブル214から、次のようにし
て音声区間の判定を行なつていた。すなわち、ブ
ロツク量の最大値となるブロツクを検出し、それ
を音声区間の中心として前後のブロツクについ
て、ブロツク間距離が一定値以下であれば当該ブ
ロツクも音声区間に含めるという方法で、音声区
間の判定を行なつていた。
The speech section determining section 23 judges the speech section from the block table 214 obtained by the blocking section 22 in the following manner. In other words, the block with the maximum block amount is detected, and if the distance between the blocks before and after the detected block is the center of the speech section, and the distance between the blocks is less than a certain value, the block is included in the speech section. was making a judgment.

認識部8は、音声取込制御部3に取込開始信号
を送るとともに、音声取込制御部3からの分析結
果を格納しておき、さらに音声区間検出部5から
の始端終端情報7を受けると、あらかじめ用意さ
れている内容既知の標準パターンとの類似度演算
を行ない、最も類似度の高い標準パターンと同一
内容の音声が入力されたと判断し、その結果を出
力する。
The recognition unit 8 sends a capture start signal to the voice capture control unit 3, stores the analysis results from the voice capture control unit 3, and further receives start and end information 7 from the voice section detection unit 5. , and a standard pattern whose content is known and has been prepared in advance, and determines that a voice with the same content as the standard pattern with the highest degree of similarity has been input, and outputs the result.

しかしながら、上記従来の技術における音声区
間検出では、 (1) 入力音声の強弱によりスペクトル傾斜ajが変
化するため、不安定なパラメータすなわち、Pj
が不安定なパラメータである。
However, in the speech interval detection using the above-mentioned conventional technology, (1) the spectral slope a j changes depending on the strength of the input speech, so an unstable parameter, that is, P j
is an unstable parameter.

(2) スペクトル傾斜ajは、音韻、話者による変化
とともにマイクの特性等によつて往往にして、
音声部においても0に近い値を取り、結果とし
てPjも0に近い値となり、ブロツク化を誤ま
る。
(2) The spectral slope a j varies depending on the phonology, speaker characteristics, microphone characteristics, etc.
The audio portion also takes a value close to 0, and as a result, P j also takes a value close to 0, resulting in incorrect blocking.

(3) ノイズが大きい場合、ノイズとの区別(特に
子音)がつけにくい。
(3) When the noise is loud, it is difficult to distinguish it from the noise (especially consonants).

という欠点があつた。There was a drawback.

(発明の課題) この発明の目的は誤認識をなくして認識率の向
上をはかることの出来る音声認識装置を提供する
ことにあり、その特徴は、音声区間検出時に、音
声パターンからノイズパターンを差し引くことに
より、音声区間検出をより精度よく行ない、認識
率を上げる手段を提供するもので、以下詳細に説
明する。
(Problem to be solved by the invention) An object of the present invention is to provide a speech recognition device that can eliminate misrecognition and improve the recognition rate.The feature is that a noise pattern is subtracted from a speech pattern when detecting a speech section. This provides a means for detecting voice segments with higher accuracy and increasing the recognition rate, which will be described in detail below.

(発明の構成および作用) 第6図は、本発明のブロツク図であり、100
は入力端子、200は周波数分析部、300は対
数変換部、400はスペクトル変換部、500は
音声区間決定部であり、対数変換済データ部50
1、ノイズパターン検出部502、減算回路50
3、乗算回路504、加算回路505、除算回路
506、Pパラメータメモリ507、比較器1
508、FLAG509、スムージング1 51
0、スムージング2 511、ブロツク化51
2、比較器2 513、ブロツク決定514、音
声区間決定515、MAXBLKテーブル516か
ら成る、600は再サンプル部、700は距離演
算部、800は標準パターンメモリ、900は判
定部、1000は認識結果出力端子である。
(Structure and operation of the invention) FIG. 6 is a block diagram of the invention.
200 is an input terminal, 200 is a frequency analysis section, 300 is a logarithmic conversion section, 400 is a spectrum conversion section, 500 is a speech interval determination section, and a logarithmically transformed data section 50
1. Noise pattern detection section 502, subtraction circuit 50
3. Multiplication circuit 504, addition circuit 505, division circuit 506, P parameter memory 507, comparator 1
508, FLAG509, smoothing 1 51
0, smoothing 2 511, blocking 51
2. Comparator 2 Consists of 513, block determination 514, speech interval determination 515, and MAXBLK table 516, 600 is a resampling section, 700 is a distance calculation section, 800 is a standard pattern memory, 900 is a judgment section, and 1000 is a recognition result output It is a terminal.

このような構成において、入力端子100から
入力される入力音声信号は、周波数分析部200
に入力され、複数の周波数帯域に対応した量子化
信号U(i,j)として周波数分析され、対数変換部30
0に送られる。
In such a configuration, the input audio signal input from the input terminal 100 is processed by the frequency analysis section 200.
is input into the quantized signal U (i,j) corresponding to multiple frequency bands, and is frequency-analyzed as a quantized signal U (i,j) corresponding to a plurality of frequency bands.
Sent to 0.

対数変換部300に送られたデータは、スペク
トル情報と、パワー情報等となり、スペクトル変
換器400へはスペクトル情報、音声区間決定部
500へはスペクトル情報及びパワー情報が送ら
れる。
The data sent to the logarithmic conversion section 300 becomes spectral information, power information, etc., the spectral information is sent to the spectral converter 400, and the spectral information and power information are sent to the speech interval determination section 500.

対数変換部300では第(4)式の計算が行なわれ
る。周波数分析データをU(i,j)とする。
The logarithmic conversion unit 300 calculates equation (4). Let the frequency analysis data be U (i,j) .

U(i,j) i=1〜19 j=1〜∞ 0≦U(i,j)≦2047 対数変換データをV(i,j)とする。 U (i,j) i=1~19 j=1~∞ 0≦U (i,j) ≦2047 Let the logarithmically transformed data be V (i,j) .

V(i,j) i=1〜19 j=1〜∞ ここでiは周波数(1ch〜19ch)を示し、jは
時間(1フレーム〜∞フレーム)を示す。また前
処理部からの入力データをU(i,j)とする。U(i,j)i=
1〜19 j=1〜∞ 0≦U(i,j)≦2047対数変換ビ
ツト数をNBとする。ここではNB=8である。
V (i,j) i=1 to 19 j=1 to ∞ Here, i indicates frequency (1ch to 19ch), and j indicates time (1 frame to ∞ frame). Also, input data from the preprocessing section is assumed to be U (i,j) . U (i,j) i=
1 to 19 j=1 to ∞ 0≦U (i,j) ≦2047 Let NB be the number of logarithmic conversion bits. Here, NB=8.

ここで入力パターンのパワーPOW(j)及び入
力パターンの10フレームパワーの計算式を第(5)
式,第(6)式で定義する。
Here, the calculation formula for the input pattern power POW (j) and the input pattern 10 frame power is expressed as (5)
It is defined by Equation (6).

POW(j)=1/1919i=1 V(i,j) j=1〜∞ (5) POW10(k)=10l=1 POW(j+l-1) (6) k=(j−1)/10+1 但し、j=(k−1)*10+1とする。POW(j)=1/19 19i=1 V(i,j) j=1~∞ (5) POW10(k)= 10l=1 POW (j+l-1) (6) k= (j-1)/10+1 However, j=(k-1)*10+1.

ノイズパターンは第(7)式で定義する。The noise pattern is defined by equation (7).

ノイズパターン測定区間をk=k1〜k2とした時、 NLEVEL=1/k2−k1+1k2k=k1 POW10(k) …(7) 但し、k2=k1+2とする ここで切り出しスライスレベルL1を L1=NLEVEL+LO として、はじめてPOW10(k3)がL1よりも大き
くPOW10(k3+1)がL1よりも大きい点k3から
40フレーム逆のぼつたフレームj1を j1=(k3−1)*10+1−40 として、仮の音声始端フレームSTFR1を STFR1=MAX(j,1) とする。
When the noise pattern measurement interval is k=k 1 to k 2 , NLEVEL=1/k 2 −k 1 +1 k2k=k1 POW10(k) …(7) However, here, k 2 = k 1 + 2. When the slice level L 1 is set to L 1 = NLEVEL + LO, POW10 (k 3 ) is greater than L1 and POW10 (k 3 +1) is greater than L1 from the point k 3 .
Let the frame j 1 with 40 frames reversed be j 1 = (k 3 -1)*10+1-40, and the temporary voice start frame STFR1 be STFR1 = MAX (j, 1).

終端検出はk4がk2+1よりも大きく、かつ
POW10(k4)がL1よりも小さいか等しくなつた
時に、仮の音声終端フレームEDFR1を EDFR1=(k4−1)*10−1+9 とする。
Termination detection is performed if k 4 is greater than k 2 +1, and
When POW10 (k 4 ) becomes smaller than or equal to L1, a temporary audio end frame EDFR1 is set as EDFR1=(k 4 -1)*10-1+9.

さて、対数変換部300より計算された対数変換
データV(i,j)は、対数変換済データ部50
1へ送られた後、ノイズパターンNPAT(i)を
求めるためノイズパターン検出部502にて、ノ
イズパターンNPAT(i)を計算する。但し、ノ
イズレベル測定区間をk=k1〜k2とした時、j2
びj3の値を第(8)式において計算する。
Now, the logarithmically transformed data V(i,j) calculated by the logarithmically transformed section 300 is transferred to the logarithmically transformed data section 50.
1, the noise pattern detection unit 502 calculates the noise pattern NPAT(i) in order to obtain the noise pattern NPAT(i). However, when the noise level measurement interval is set to k= k1 to k2 , the values of j2 and j3 are calculated using equation (8).

j2=(k1−1)*10+1 j3=(k2−1)*10+1+9 …(8) ノイズパターンNPAT(i)を求める式を第(9)
式に示す。
j 2 = (k 1 -1) *10 + 1 j 3 = (k 2 -1) *10 + 1 + 9 ... (8) The formula for determining the noise pattern NPAT (i) is expressed as (9)
As shown in the formula.

J=STFR1〜EDFR1NPAT(i)=1/j3−j2+1 (j3j=j2 (i,j)+j3−j2+1/2) …(9) 次に、減算回路503、乗算回路504、加算
回路505、除算回路506、において、対数変
換済データ部501に格納されているV(i,j)及びノ
イズパターン検出部502において、第(9)式より
求まつたNPAT(i)を用い、ノイズパターンを
差し引いたパワーの計算を第(10)式により行なう。
J=STFR1~EDFR1NPAT(i)=1/ j3 - j2 +1 ( j3j=j2 (i,j)+ j3 -j2 +1 /2)...(9) Next, the subtraction circuit 503 and the multiplication circuit 504, addition circuit 505, and division circuit 506, V (i, j) stored in logarithmically transformed data section 501 and noise pattern detection section 502, NPAT(i) found from equation (9). The power after subtracting the noise pattern is calculated using equation (10).

P(j)=1/1919i=1 ((V(i,j)−NPAT(i)/4)2+9…(10) 第(10)式より求まつたP(j)はPパラメータメモ
リ507へ格納され、比較器1508により次の
第(11)式の比較を行なう。
P(j)=1/19 19i=1 ((V(i,j)−NPAT(i)/4) 2 +9…(10) P(j) found from equation (10) is P The signal is stored in the parameter memory 507, and the comparator 1508 performs a comparison according to the following equation (11).

FLAG(j)=0 Pp(j)<L2 1 P(j)≧L2 …(11) 第(11)式において、スライスレベルL2がP(j)よ
りも大きい場合は、FLAG(j)=0とする。また
L2がP(j)よりも等しいか小さい場合はFLAG
(j)=1とする。第(11)式において決定された
FLAG(j)の値は、FLAG509へ格納され、
FLAG(j)の値に応じて、スムージング1 5
10あるいはスムージング2 511へ送られ
る。スムージング1 510ではFLAG(j)=0
の場合の操作を行ないFLAG(j−1)=0であ
り、FLAG(j+1)=0である時は、FLAG(j)
=0とする。また、スムージング2 511では
FLAG(j)=1の場合の操作を行ないFLAG(j
−1)=1であり、FLAG(j+1)=1である時
は、FLAG(j)=1とする。
FLAG(j)=0 Pp(j)<L2 1 P(j)≧L2 …(11) In equation (11), if slice level L2 is greater than P(j), FLAG(j)=0 shall be. Also
FLAG if L2 is less than or equal to P(j)
Let (j)=1. Determined in equation (11)
The value of FLAG(j) is stored in FLAG509,
Smoothing 1 5 depending on the value of FLAG(j)
10 or smoothing 2 511. Smoothing 1 For 510, FLAG(j) = 0
When FLAG (j-1) = 0 and FLAG (j + 1) = 0, FLAG (j)
=0. Also, in smoothing 2 511
Perform the operation when FLAG(j) = 1 and set FLAG(j
-1)=1 and FLAG(j+1)=1, then FLAG(j)=1.

次にブロツク化512においてFLAG(j)=1
が4フレーム以上連続し、その区間の、POW1
(l)=1/8〓P(j)がPOW1(l)≧L3、すなわ ちPOW1(l)がスライスレベルL3よりも大きい
か等しい場合のものをブロツクとする。
Next, in blocking 512, FLAG(j)=1
is continuous for 4 or more frames, POW1 of that section
(l)=1/8〓P(j) is defined as a block if POW1(l)≧L3, that is, POW1(l) is greater than or equal to the slice level L3.

ブロツク数をBLKSとし、ブロツクlの先頭フ
レームをS(l)、ブロツクlの最終フレームをE
(l)とする。ブロツクlのノイズパターンを差
し引いたパワーP(j)の加算値は第(12)式により
求められる。
The number of blocks is BLKS, the first frame of block l is S(l), and the last frame of block l is E.
Let it be (l). The added value of the power P(j) after subtracting the noise pattern of block 1 is obtained by equation (12).

POW1(l)=1/8E(l)j=s(l) P(j) …(12) ブロツクlのフレーム数は第(13)式により求め
られる。
POW1(l)=1/8 E(l)j=s(l) P(j)...(12) The number of frames of block 1 is determined by equation (13).

FR1(l)=E(l)−S(l)+1 …(13) また、前ブロツク(l−1)との間隔は第(14)
式により求められる。
FR1(l)=E(l)-S(l)+1...(13) Also, the distance from the previous block (l-1) is the (14th)
It is determined by the formula.

FR2(l)=S(l)−E(l−1) …(14) ここでl1を音声先頭ブロツク、l2を音声最終ブロ
ツクとして比較器2 513において、音声先頭
ブロツクl1については、第(15)式の条件を満た
している限りl1=l1−1とする。
FR2(l)=S(l)-E(l-1)...(14) Here, l1 is the first audio block and l2 is the final audio block.In the comparator 2 513, regarding the audio first block l1 , As long as the condition of equation (15) is satisfied, l 1 =l 1 -1.

FR2(l1)≦MIN(POW1(l1−1)/SC1+SC2,
SC3) …(15) また音声最終ブロツクl2については、第(16)式
の条件を満たしている限りl2=l2+1とする。
FR2(l 1 )≦MIN(POW1(l 1 −1)/SC1+SC2,
SC3) ...(15) Also, regarding the final audio block l2 , as long as the condition of equation (16) is satisfied, l2 = l2 +1.

FR(l2+1)≦MIN(POW1(l2+1)/SC1+
SC2,SC3) …(16) ここでSC1〜SC3は定数でありSC1=16,SC2=
8,SC3=30である。
FR(l 2 +1)≦MIN(POW1(l 2 +1)/SC1+
SC2, SC3) …(16) Here, SC1 to SC3 are constants, SC1=16, SC2=
8, SC3=30.

以上の式より、最大ブロツクを中心に前後のブ
ロツクを音声区間のブロツクとして取り込むかど
うかの判定を行ない、音声区間ブロツク候補とし
て採用する。
Based on the above equation, it is determined whether or not the blocks before and after the largest block are to be taken in as speech section blocks, and are adopted as speech section block candidates.

一般に、日本語の50音におけるカ行、タ行、パ
行、バ行、ダ行及びガ行の音は、破裂音と呼ばれ
ているものである。このような破裂音は、一旦息
を止めた後、一気に声帯を開放して振動させるこ
とにより、発声される。一般に、破裂音を含む単
語は、その単語の発声期間中に10msから30ms程
度の無音声期間を生じることがあり、その長さに
は個人差がある。
In general, the sounds in the Japanese 50 sounds such as ka, ta, pa, ba, da, and ga are called plosives. Such plosive sounds are produced by once holding the breath and then opening the vocal cords all at once to vibrate them. In general, a word containing a plosive may have a silent period of about 10 to 30 ms during the utterance of the word, and the length varies from person to person.

従つて、この発明では、前述のように、単語の
発声期間中において、そのパワーが所定値以上、
かつ所定フレーム以上連続した場合に、その音声
部分をブロツクとして定義するものである。
Therefore, in this invention, as mentioned above, during the utterance period of a word, the power is equal to or greater than a predetermined value,
If the audio portion continues for a predetermined number of frames or more, the audio portion is defined as a block.

一つの単語におけるブロツク数は、破裂音をい
くつ含むかによつて異なるが、通常は1以上であ
る。しかし、電話の音声を認識する場合に、受信
された音声に重畳するノイズレベルによつては、
ノイズを音声のブロツクであると語認識してしま
う可能性が存在する。従つて、連続する2以上の
ブロツクには、音声ブロツクだけでなく、ノイズ
ブロツクも含まれ得るので、このようなブロツク
を音声区間ブロツク候補と呼んでいる。換言すれ
ば、音声区間ブロツク候補は、そのまま音声ブロ
ツクの場合もあるし、そうでないノイズブロツク
の場合もある。このようにして決定された音声区
間ブロツク候補である音声先頭ブロツクl1及び音
声最終ブロツクl2の値はブロツク決定514に送
られる。
The number of blocks in one word varies depending on how many plosives it contains, but is usually 1 or more. However, when recognizing telephone voices, depending on the noise level superimposed on the received voice,
There is a possibility that the noise will be recognized as a block of speech. Therefore, since two or more consecutive blocks can include not only speech blocks but also noise blocks, such blocks are called speech section block candidates. In other words, a speech section block candidate may be a speech block as it is, or it may be a noise block. The values of the voice section block candidates L1 and L2 , which are voice section block candidates determined in this way, are sent to block determination 514.

次に、音声区間決定515に用いる認識語の最
大ブロツク数のMAXBLK516を説明する。
Next, MAXBLK 516, which is the maximum number of blocks of recognition words used in speech section determination 515, will be explained.

一般に、一つの単語は、これに破裂音が含まれ
ていれば、複数のブロツに分割され得る。しか
し、発声の個人差、及びフレームサンプリングの
タイミング(1フレームは、10ms程度である。)
により、その単語に破裂音が含まれていたとし
て、、常に複数のブロツクに分割されるとは限ら
ない。しかし、単語には、その先頭の破裂音を無
視しても、その中に含まれる破裂音の数+1を超
えることはない。
Generally, one word can be divided into multiple blots if it contains a plosive. However, there are individual differences in vocalization and the timing of frame sampling (one frame is about 10ms).
Therefore, even if the word contains a plosive, it is not always divided into multiple blocks. However, even if you ignore the plosive at the beginning of a word, the number of plosives in the word will not exceed the number of plosives + 1.

一つの単語に含まれるブロツク数はその単語に
含まれる破裂音の数に基づいて決定され、その最
大値を最大ブロツク数という。実際のブロツ数
は、最大ブロツク数以下となることがあつても、
これを超えることはない。例えば、「イチ」の最
大ブロツク数は、2である。しかし、実際には、
前述の理由により、「イチ」のブロツク数が1に
なることもある。
The number of blocks included in one word is determined based on the number of plosives included in the word, and the maximum value is called the maximum number of blocks. Even if the actual number of blocks may be less than the maximum number of blocks,
It will never exceed this. For example, the maximum number of blocks for "1" is 2. However, in reality,
For the reasons mentioned above, the number of "first" blocks may be one.

このようなブロツク数を複数の単語についてそ
れぞれ対応付けしてテーブルにしたものが、最大
ブロツク(MAXBLK)テーブルである。最大ブ
ロツク(MAXBLK)テーブルーをメモリに記憶
したものがMAXBLKテーブル516である。
A maximum block (MAXBLK) table is a table in which such block numbers are associated with each of a plurality of words. The MAXBLK table 516 is a maximum block (MAXBLK) table stored in memory.

最大ブロツク数MAXBLKの例を第8図に示
す。左側がカテゴリ(16語)を示し、右側は、予
め発声データから求めた各カテゴリの最大ブロツ
ク数を示す。これらの認識語セツトの中で最大の
MAXBLKを選ぶ。例えば認識語の中に「モーイ
チド」を含むならMAXBLK=3とする。一般化
すると、最大ブロツク数(MAXBLK)は、原則
として単語中に含まれる破裂音の数+1である。
ただし、先頭の破裂音は含まれないものとする。
また、例外として、例えば「オワリ」なる語にお
ける「リ」は、破裂音ではないが、発声者によつ
てブロツク数が2となることが経験されるので、
そのブロツク数は2とする。
An example of the maximum block number MAXBLK is shown in FIG. The left side shows the categories (16 words), and the right side shows the maximum number of blocks for each category calculated in advance from the utterance data. The largest of these recognition word sets is
Select MAXBLK. For example, if the recognized words include "Moichido", MAXBLK=3. Generalizing, the maximum number of blocks (MAXBLK) is, in principle, the number of plosives included in a word + 1.
However, the initial plosive sound shall not be included.
As an exception, for example, the ``ri'' in the word ``owari'' is not a plosive, but the number of blocks is 2, which is experienced by the utterer.
The number of blocks is 2.

音声区間決定部515において、 BLKS≦MAXBLK とする時、すなわちブロツク数BLKSが最大ブロ
ツク数MAXBLKよりも小さいか等しい場合であ
ればすべてのブロツクを音声区間とする。逆に BLKS>MAXBLK とする時、すなわちブロツク数BLKが最大ブロ
ツク数MAXBLKよりも大きい場合、例えば第7
図においてブロツク数BLKS=3で最大ブロツク
数MAXBLK=2であればまたはの組み合わ
せが考えられ、及びのブロツクの組み合わせ
の各々のパワーPP(l)を求めた後PPの比較を
行ないブロツクのパワーPP(l)が最大となるブ
ロツクの組合せを音声区間とする。ブロツクのパ
ワーPP(l)は第(17)式により求められる。
In the voice section determining section 515, when BLKS≦MAXBLK, that is, when the number of blocks BLKS is smaller than or equal to the maximum number of blocks MAXBLK, all blocks are determined as voice sections. Conversely, when BLKS>MAXBLK, that is, when the number of blocks BLK is larger than the maximum number of blocks MAXBLK, for example, the seventh
In the figure, if the number of blocks BLKS = 3 and the maximum number of blocks MAXBLK = 2, a combination of or is considered, and after calculating the power PP (l) of each combination of blocks, the power PP of the block is calculated by comparing the PP. The combination of blocks for which (l) is the maximum is defined as a voice section. The power PP(l) of the block is obtained from equation (17).

PP(l)=1/E(l+MAXBLK−1)−S(l)+1 E(l+MAXBLK-4)j=s(l) P(j) …(17) l=1〜BLKS−MAXBLK+1 第(17)式より求められたS(l1)は音声先頭ブ
ロツクであり、E(l2)は音声最終ブロツクとな
り、音声始端フレームSTFRは STFR=S(l1) また音声終端フレームEDFRは EDFR=E(l2) となる。また、入力パターンフレーム数IFRは次
の第(18)式で表わされる。
PP(l)=1/E(l+MAXBLK-1)-S(l)+1 E(l+MAXBLK-4)j=s(l) P(j)...(17) l=1~BLKS-MAXBLK+1 th S(l 1 ) obtained from equation (17) is the audio start block, E(l 2 ) is the audio final block, audio start frame STFR is STFR=S(l 1 ), and audio end frame EDFR is EDFR =E(l 2 ). Further, the number of input pattern frames IFR is expressed by the following equation (18).

IFR=EDFR−STFR+1 …(18) 処理終了の判定は、音声最終ブロツクl2が以下
の第(19)式の条件を全て満たした時、処理を終
了とする。
IFR=EDFR-STFR+1 (18) The process is determined to be finished when the final audio block l2 satisfies all of the conditions in equation (19) below.

POW10(K4)≦L1 POW10(k4+1)≦L1 POW10(K4+2)≦L1 POW10(k4+3)≦L1 POW10(k4+4)≦L1 …(19) すなわち、L1がk4,k4+1,k4+2,k4+3,
k4+4,のいずれに対しても大きいか等しい場合
は、終理終了となる。
POW10(K 4 )≦L1 POW10(k 4 +1)≦L1 POW10(K 4 +2)≦L1 POW10(k 4 +3 )≦L1 POW10(k 4 +4)≦L1 …(19) That is, L1 is k 4 , k 4 +1, k 4 +2, k 4 +3,
If it is greater than or equal to any of k 4 +4, the process is terminated.

また第(19)式の条件が満たされなかつた場合
は、認識を打ち切り POW10(k4)≦L1 すなわちL1が大きいか等しくなる次のk4の値を
求める。
If the condition of equation (19) is not satisfied, recognition is stopped and the next value of k 4 where POW10(k 4 )≦L1, that is, L1 is greater than or equal to, is determined.

このように決定された音声区間STFR及び
EDFRは、スペクトル変換部400から送られる
W(i,j)と同時に再サンプル部600に送ら
れる。再サンプル部600では、音声の時間軸の
正規化を行われる。時間軸の正規化の方法は従来
公知の技術であり、リニアマツチング方法では、
音声区間を認識装置の条件によつて定められた一
定数に、時間的に等間隔に分割、再サンプする方
法である。そして、距離演算部700において、
同様に作成された標準パターンメモリ800の出
力との距離演算を行ないその結果を判定部900
へ送る。
The voice interval STFR determined in this way and
EDFR is sent to the resampling unit 600 at the same time as W(i,j) sent from the spectrum conversion unit 400. The resampling unit 600 normalizes the time axis of the audio. The time axis normalization method is a conventionally known technique, and the linear matching method
This is a method in which a voice section is divided into a fixed number determined by the conditions of the recognition device at equal intervals in time and resampled. Then, in the distance calculation section 700,
The determination unit 900 calculates the distance from the output of the standard pattern memory 800 created in the same way and uses the result.
send to

判定部900では、トータル距離との距離値の
比較を行ない、最も小さいトータル距離のカテゴ
リ名を認識結果として、認識結果出力端子100
0から出力する。
The determination unit 900 compares the distance value with the total distance, and outputs the category name with the smallest total distance as the recognition result to the recognition result output terminal 100.
Output from 0.

以上説明したように、本発明では、音声区間検
出時に音声パターンからノイズパターンを差し引
くことにより、音声区間検出をより精度よく行な
い、認識率を上げることができる。
As described above, in the present invention, by subtracting the noise pattern from the speech pattern when detecting the speech section, the speech section can be detected more accurately and the recognition rate can be increased.

(発明の効果) 本発明は、音声区間検出の際に、音声のノイズ
パターンの情報を音声パターン情報から差し引く
ことにより、音声区間検出をより精度よく行なう
ことができ、音声認識装置の認識性能を向上する
のに効果がある。
(Effects of the Invention) The present invention makes it possible to detect speech segments more accurately by subtracting information about speech noise patterns from speech pattern information when detecting speech segments, thereby improving the recognition performance of the speech recognition device. It is effective in improving.

【図面の簡単な説明】[Brief explanation of drawings]

第1図は従来の音声認識装置のブロツク図、第
2図は第1図の周波数分析部の詳細ブロツク図、
第3図は第1図の音声区間検出部のブロツク図、
第4図は第3図のパラメータ演算部の詳細ブロツ
ク図、第5図は第3図のブロツク化部の詳細図、
第6図は本発明の音声認識装置のブロツク図、第
7図は音声区間のブロツクの組合せを示す図、第
8図は音声の最大ブロツク数を示す図である。 1…入力端子、2…周波数分析部、3…音声取
込制御部、4…取込開始信号、5…音声区間検出
部、6…取込終了信号、7…始端・終端情報、8
…認識部、9…出力端子、11…入力音声信号、
12…前置増幅器、13…バンドパスフイルタ
群、14…全波整流器群、15ローパスフイルタ
群、16…マルチプレクサ、17…AD変換器、
18…分析結果、21…パラメータ演算部、22
…ブロツク化部、23…音声区間判定部、10
1,105,109…加算器、102,110…
レジスタ、103,108,111,113…乗
算器、104…補数器、106,112…除算
器、107…カウンタ、200…Pパラメータメ
モリ、201…絶対値回路、202,205…比
較器、203,207,209…カウンタ、20
4,212,213…レジスタ、206…AND
回路、208,210…減算器、211…加算
器、214…ブロツクテーブル、100…入力端
子、200…周波数分析部、300…対数変換
部、400…スペクトル変換部、500…音声区
間決定部、501…対数変換部、502…ノイズ
パターン検出部、503…減算回路、504…乗
算回路、505…加算回路、506…除算回路、
507…Pパラメータメモリ、508…比較器
1、509…FLAG、510スムージング1、5
11…スムージング2、512…ブロツク化、5
13…比較器2、512,514…ブロツク決
定、515…音声区間決定、516…
MAXBLK、600…再サンプル部、700…距
離演算部、800…標準パタンメモリ、900…
判定部、1000…認識結果出力端子。
Figure 1 is a block diagram of a conventional speech recognition device, Figure 2 is a detailed block diagram of the frequency analysis section in Figure 1,
FIG. 3 is a block diagram of the voice section detection section in FIG.
4 is a detailed block diagram of the parameter calculation section in FIG. 3, FIG. 5 is a detailed diagram of the blocking section in FIG. 3,
FIG. 6 is a block diagram of the speech recognition apparatus of the present invention, FIG. 7 is a diagram showing combinations of blocks in speech sections, and FIG. 8 is a diagram showing the maximum number of speech blocks. DESCRIPTION OF SYMBOLS 1... Input terminal, 2... Frequency analysis section, 3... Audio capture control section, 4... Capture start signal, 5... Audio section detection section, 6... Capture end signal, 7... Start/end information, 8
... recognition unit, 9 ... output terminal, 11 ... input audio signal,
12... Preamplifier, 13... Band pass filter group, 14... Full wave rectifier group, 15... Low pass filter group, 16... Multiplexer, 17... AD converter,
18... Analysis result, 21... Parameter calculation section, 22
...Blocking unit, 23...Voice section determining unit, 10
1,105,109...adder, 102,110...
Register, 103, 108, 111, 113... Multiplier, 104... Complementer, 106, 112... Divider, 107... Counter, 200... P parameter memory, 201... Absolute value circuit, 202, 205... Comparator, 203, 207, 209...Counter, 20
4,212,213...Register, 206...AND
Circuit, 208, 210... Subtractor, 211... Adder, 214... Block table, 100... Input terminal, 200... Frequency analysis section, 300... Logarithmic conversion section, 400... Spectrum conversion section, 500... Speech interval determination section, 501 ... Logarithmic conversion section, 502 ... Noise pattern detection section, 503 ... Subtraction circuit, 504 ... Multiplication circuit, 505 ... Addition circuit, 506 ... Division circuit,
507...P parameter memory, 508...Comparator 1, 509...FLAG, 510 Smoothing 1, 5
11...Smoothing 2, 512...Blocking, 5
13... Comparator 2, 512, 514... Block determination, 515... Voice section determination, 516...
MAXBLK, 600...Resample section, 700...Distance calculation section, 800...Standard pattern memory, 900...
Judgment unit, 1000... recognition result output terminal.

Claims (1)

【特許請求の範囲】 1 入力された音声信号の音声区間を判定する音
声認識装置において、 前記音声信号を周波数分析し、その結果を対数
変換して得られた対数変換データに含まれるノイ
ズのノイズパターンを演算する手段と、 前記対数変換データから前記ノイズパターンを
差し引いて前記音声パターンに含まれるパワー情
報を演算する手段と、 前記パワー情報を所定の第1基準レベルと比較
し、両者間の大小関係を示す理論レベルの音声区
間フラグを求める手段と、 特定の音声区間フラグの論理レベルがその前後
の音声フラグの論理レベルと同一であつたときに
前記特定の音声区間フラグの論理レベルを確定す
ることによりスムージングを行なう手段と、 前記スムージングを行なう手段によりスムージ
ングされた音声区間フラグが所定の期間について
連続し、かつ前記パワー情報が前記第1基準レベ
ルより大きい第2基準レベルを超えているとき
は、当該パワー情報を音声ブロツク候補と判断す
る手段と、 前記音声ブロツク候補の最大ブロツク数と、前
記音声ブロツク候補に対応した最大ブロツクテー
ブルにおける最大ブロツク数とを比較した結果に
基づいて前記音声信号の音声区間を決定する手段
と を有し、 前記最大ブロツクテーブルは、複数の語とそれ
ら語の発声データについて予め求めた最大ブロツ
ク数とを対応させ、前記音声区間を決定する手段
により続み出し可能に記憶されていることを特徴
とする音声認識装置。
[Claims] 1. In a speech recognition device that determines the speech interval of an input speech signal, noise contained in logarithmically transformed data obtained by frequency-analyzing the speech signal and logarithmically transforming the result. means for calculating a pattern; means for calculating power information included in the voice pattern by subtracting the noise pattern from the logarithmically transformed data; and comparing the power information with a predetermined first reference level to determine the magnitude between the two. means for determining a theoretical level speech interval flag indicating a relationship; and determining the logic level of the specific speech interval flag when the logic level of the specific speech interval flag is the same as the logic level of the speech flags before and after it; means for performing smoothing by performing smoothing; and when the voice section flags smoothed by the means for performing smoothing are continuous for a predetermined period and the power information exceeds a second reference level that is greater than the first reference level; , means for determining the power information as an audio block candidate; and means for determining the audio signal based on the result of comparing the maximum number of blocks of the audio block candidate with the maximum number of blocks in the maximum block table corresponding to the audio block candidate. and means for determining a voice interval, and the maximum block table can be continued by the means for determining a voice interval by associating a plurality of words with a maximum number of blocks determined in advance for the utterance data of those words. A voice recognition device characterized by being stored in a voice recognition device.
JP59108668A 1984-05-30 1984-05-30 Voice recognition system Granted JPS60254100A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP59108668A JPS60254100A (en) 1984-05-30 1984-05-30 Voice recognition system

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP59108668A JPS60254100A (en) 1984-05-30 1984-05-30 Voice recognition system

Publications (2)

Publication Number Publication Date
JPS60254100A JPS60254100A (en) 1985-12-14
JPH0424717B2 true JPH0424717B2 (en) 1992-04-27

Family

ID=14490648

Family Applications (1)

Application Number Title Priority Date Filing Date
JP59108668A Granted JPS60254100A (en) 1984-05-30 1984-05-30 Voice recognition system

Country Status (1)

Country Link
JP (1) JPS60254100A (en)

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
GB8613327D0 (en) * 1986-06-02 1986-07-09 British Telecomm Speech processor
JP2701431B2 (en) * 1989-03-06 1998-01-21 株式会社デンソー Voice recognition device
JPH03212697A (en) * 1990-01-18 1991-09-18 Matsushita Electric Ind Co Ltd signal processing device
WO2020218597A1 (en) 2019-04-26 2020-10-29 株式会社Preferred Networks Interval detection device, signal processing system, model generation method, interval detection method, and program

Also Published As

Publication number Publication date
JPS60254100A (en) 1985-12-14

Similar Documents

Publication Publication Date Title
CA1227286A (en) Speech recognition method and apparatus thereof
EP1083542B1 (en) A method and apparatus for speech detection
US5583961A (en) Speaker recognition using spectral coefficients normalized with respect to unequal frequency bands
EP1393300B1 (en) Segmenting audio signals into auditory events
US20060053003A1 (en) Acoustic interval detection method and device
EP0411290A2 (en) Method and apparatus for extracting information-bearing portions of a signal for recognizing varying instances of similar patterns
AU2002252143A1 (en) Segmenting audio signals into auditory events
JPS6128998B2 (en)
US5159637A (en) Speech word recognizing apparatus using information indicative of the relative significance of speech features
EP0474496B1 (en) Speech recognition apparatus
US5522013A (en) Method for speaker recognition using a lossless tube model of the speaker&#39;s
JPS60200300A (en) Voice head/end detector
JPS60254100A (en) Voice recognition system
EP0537316B1 (en) Speaker recognition method
JPH0556520B2 (en)
JP2606211B2 (en) Sound source normalization method
JPS62113197A (en) Voice recognition equipment
JPS6131880B2 (en)
JP2658104B2 (en) Voice recognition device
JPS61256399A (en) Voice recognition system
JP2744622B2 (en) Plosive consonant identification method
JPS6255798B2 (en)
JPH0752355B2 (en) Voice recognizer
JPS61203497A (en) Voice recognition system
JPH0221598B2 (en)