JPH0570837B2 - - Google Patents

Info

Publication number
JPH0570837B2
JPH0570837B2 JP20020984A JP20020984A JPH0570837B2 JP H0570837 B2 JPH0570837 B2 JP H0570837B2 JP 20020984 A JP20020984 A JP 20020984A JP 20020984 A JP20020984 A JP 20020984A JP H0570837 B2 JPH0570837 B2 JP H0570837B2
Authority
JP
Japan
Prior art keywords
voice
section
pattern
threshold value
registered
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Expired - Lifetime
Application number
JP20020984A
Other languages
Japanese (ja)
Other versions
JPS6177900A (en
Inventor
Hiromi Fujii
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NEC Corp
Original Assignee
Nippon Electric Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Electric Co Ltd filed Critical Nippon Electric Co Ltd
Priority to JP59200209A priority Critical patent/JPS6177900A/en
Publication of JPS6177900A publication Critical patent/JPS6177900A/en
Publication of JPH0570837B2 publication Critical patent/JPH0570837B2/ja
Granted legal-status Critical Current

Links

Classifications

    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02EREDUCTION OF GREENHOUSE GAS [GHG] EMISSIONS, RELATED TO ENERGY GENERATION, TRANSMISSION OR DISTRIBUTION
    • Y02E60/00Enabling technologies; Technologies with a potential or indirect contribution to GHG emissions mitigation
    • Y02E60/10Energy storage using batteries
    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02PCLIMATE CHANGE MITIGATION TECHNOLOGIES IN THE PRODUCTION OR PROCESSING OF GOODS
    • Y02P10/00Technologies related to metal processing
    • Y02P10/20Recycling

Description

【発明の詳細な説明】 (産業上の利用分野) 本発明は、音声認識技術などで用いられる入力
音声の存在範囲を検出する音声区間検出装置に関
するものである。
DETAILED DESCRIPTION OF THE INVENTION (Field of Industrial Application) The present invention relates to a speech section detection device for detecting the range of input speech used in speech recognition technology.

(従来技術) 従来の音声区間検出装置としては、種々の方式
によるものが試みられている。それらのうち代表
的な方法として、パワーや零交差数などの特徴パ
ラメータの閾値をあらかじめ設定し、入力信号の
パワーや零交差数、あるいは、単位時間内のパワ
ーから求められるエネルギーなどがある閾値を超
えるかどうかで音声区間の検出を行うものがあ
る。
(Prior Art) Various methods have been tried as conventional voice segment detection devices. Among these methods, a typical method is to set thresholds for characteristic parameters such as power and number of zero crossings in advance, and then set a threshold value such as the power of the input signal, the number of zero crossings, or the energy calculated from the power within a unit time. There is a method that detects a voice section based on whether or not it exceeds the range.

以下に図面を用いて従来の音声検出装置の原理
を説明する。
The principle of a conventional voice detection device will be explained below with reference to the drawings.

第1図は従来の音声検出装置の原理を示すブロ
ツク図である。
FIG. 1 is a block diagram showing the principle of a conventional voice detection device.

マイク1から入力された音声は音声パラメータ
抽出部2においてパワーを含む特徴パラメータの
時系列に変換される。
The voice input from the microphone 1 is converted into a time series of characteristic parameters including power in the voice parameter extraction section 2.

次に、あらかじめ閾値記憶部3に設定された閾
値に従い、音声区間検出部4において検出が行わ
れる。この方法は固定、あるいは、ノイズレベル
に適応した閾値と入力音声の特徴パラメータを時
間ごとに比較することにより実行される。音声検
出部4における処理はたとえば特願昭58−156098
号明細書「可変閾値型音声検出器」に記載されて
いる方式が知られている。
Next, detection is performed in the voice section detection section 4 according to a threshold value set in the threshold value storage section 3 in advance. This method is performed by comparing the characteristic parameters of the input speech with a fixed threshold or a threshold value adapted to the noise level at each time. The processing in the voice detection section 4 is described in, for example, Japanese Patent Application No. 58-156098.
A method described in the specification "Variable Threshold Type Speech Detector" is known.

この方式は、入力信号中の雑音信号の平均電力
を求め、その値からパワー、パワーの累積などの
閾値を設定し、入力信号との比較により求めた、
それぞれの閾値に対する結果から始端、終端を決
定するものである。
This method calculates the average power of the noise signal in the input signal, sets a threshold value for power, power accumulation, etc. from that value, and then compares it with the input signal.
The starting end and ending end are determined from the results for each threshold value.

音声区間検出後、入力音声が登録パタンの場合
はS1をAに切り換えた後、登録パタンバツフア
メモリ5に、認識パタンの場合はスイツチS1を
Bに切り換えた後、認識パタンバツフアメモリ6
にそれぞれのパタンが記憶される。
After detecting the voice section, if the input voice is a registered pattern, switch S1 to A and then store it in the registered pattern buffer memory 5. If the input voice is a recognized pattern, switch S1 to B and then store it in the recognition pattern buffer memory 6.
Each pattern is memorized.

(従来技術の問題点) しかし、以上説明してきたような音声区間検出
装置では、音声区間検出のための閾値が、音声の
個人差に適応しないため、音声者によつては、語
頭、語尾の誤検出が起こる場合がある。このこと
は認識時に認識エラーが発生することを意味す
る。特に登録時にこのような誤検出が起こると、
どのようにすぐれた認識方式に用いたとしても認
識エラーが大幅に増加してしまうという問題が発
生することになる。
(Problems with the prior art) However, in the speech segment detection device as described above, the threshold value for detecting speech segments does not adapt to individual differences in speech. False positives may occur. This means that a recognition error occurs during recognition. Especially when such false positives occur during registration,
No matter how good the recognition method is, a problem will arise in that recognition errors will increase significantly.

(発明の目的) 本発明の目的は、登録パタンの音声区間検出が
正確であり、しかも、発声者ごとに、最適な閾値
を学習する機能を備えた音声区間検出装置を提供
することにある。
(Object of the Invention) An object of the present invention is to provide a speech section detection device that accurately detects speech sections of registered patterns and has a function of learning an optimal threshold value for each speaker.

(発明の構成) 本発明による音声区間検出装置は次のような各
部を必要とする。すなわち、入力された音声から
パワーを含むパラメータ時系列を抽出する音声パ
ラメータ抽出部と、前記音声パラメータ抽出部で
得られた音声パラメータを記憶する、パタンバツ
フアメモリと、入力された登録パタンごとにあら
かじめ定められたルールを記録するルール記憶部
と、前記パターンバツフアメモリ内の登録パタン
とルール記憶部のルールにより、最適の閾値を学
習する最適閾値学習部と、そこで得られた最適閾
値を用いて、前記パタンバツフアメモリ内の登録
パタンおよび音声パラメータ抽出部で得られた認
識パタンの音声区間の検出を行う音声区間検出部
と、前記音声区間検出部により音声検出された登
録パタンを記憶する登録パタンバツフアメモリ
と、同じく前記音声区間検出部により音声検出さ
れた認識パタンを記憶する認識パタンバツフアメ
モリの各部である。
(Structure of the Invention) The speech section detection device according to the present invention requires the following parts. That is, a voice parameter extraction unit extracts a parameter time series including power from input voice, a pattern buffer memory stores the voice parameters obtained by the voice parameter extraction unit, and a pattern buffer memory for each input registered pattern. A rule storage unit that records predetermined rules; an optimal threshold learning unit that learns an optimal threshold based on registered patterns in the pattern buffer memory and rules in the rule storage unit; and an optimal threshold learning unit that uses the optimal threshold obtained therein. a voice section detecting section that detects the registered pattern in the pattern buffer memory and the voice section of the recognition pattern obtained by the voice parameter extracting section; and storing the registered pattern voice detected by the voice section detecting section. These sections include a registered pattern buffer memory and a recognition pattern buffer memory that stores recognition patterns voice-detected by the voice section detecting section.

(本発明の作用・原理) 本発明の原理は、以下の3つのステツプに分け
て考える事ができる。まず第1ステツプでは、複
数個の登録パタンに対して求められた複数個の閾
値から最適な閾値を学習する。
(Operation/Principle of the Present Invention) The principle of the present invention can be considered in the following three steps. First, in the first step, an optimal threshold value is learned from a plurality of threshold values obtained for a plurality of registered patterns.

すなわち登録パタンごとに音声区間検出が正確
に求められる条件をルールとして与え、そのルー
ルを満足する閾値を各登録パタンに対してそれぞ
れ求める。このように求められた複数個の閾値は
登録パタンに対して求められたものであり、これ
らの中から最適の閾値を得ることにより、音声者
の音声に適した閾値を得ることができる。第2の
ステツプは、第1ステツプで得られた最適閾値を
用いて、登録パタンの音声検出を行う。ここでは
最適閾値が登録パタンから得られたものであるた
め音声検出の確実性が期待できる。第3のステツ
プは登録パタンの音声検出後、認識パタンに対し
て第2ステツプと同様に音声検出を行う。
That is, a rule is given to the conditions under which voice segment detection can be accurately obtained for each registered pattern, and a threshold value that satisfies the rule is determined for each registered pattern. The plurality of threshold values obtained in this way are obtained for the registered pattern, and by obtaining the optimal threshold value from among these, it is possible to obtain a threshold value suitable for the voice of the speaker. In the second step, audio detection of the registered pattern is performed using the optimal threshold obtained in the first step. Here, since the optimal threshold value is obtained from the registered pattern, reliability of voice detection can be expected. In the third step, after detecting the voice of the registered pattern, voice detection is performed on the recognized pattern in the same manner as in the second step.

(実施例) 以下に本発明の実施例について図面を参照しな
がら詳細に説明する。
(Example) Examples of the present invention will be described in detail below with reference to the drawings.

第3図は本発明の音声区間検出装置の一実施例
を示すブロツク図であり、マイク1、音声パラメ
ータ抽出部2、パタンバツフアメモリ7、ルール
記憶部8、最適閾値学習部9、音声区間検出部
4、登録パタンバツフアメモリ5、認識パタンバ
ツフアメモリ6とからなる。
FIG. 3 is a block diagram showing an embodiment of the speech section detection device of the present invention, which includes a microphone 1, a speech parameter extraction section 2, a pattern buffer memory 7, a rule storage section 8, an optimum threshold value learning section 9, and a speech section detection device. It consists of a detection section 4, a registered pattern buffer memory 5, and a recognized pattern buffer memory 6.

ここでは入力音声を数字に限り、登録パタンと
して「ゼロ」「イチ」「ニイ」……「キユー」を1
回ずつ計10パタン、認識パタンとして「ニイサ
ン」の1パタンを用いた場合を例にとつて説明す
る。
Here, the input voice is limited to numbers, and the registered patterns are ``zero'', ``ichi'', ``nii''... ``kyu''.
An example will be explained in which a total of 10 patterns are used each time, and one pattern of "Niisan" is used as a recognition pattern.

第1ステツプでは、まずスイツチS2をAに切
り換えた後マイク1から入力されたデータ「ゼ
ロ」〜「キユー」が音声パラメータ抽出部2にお
いてパワーを含む特徴パラメータの時系列に変換
される。
In the first step, after switching the switch S2 to A, the data "ZERO" to "QUE" inputted from the microphone 1 are converted into a time series of characteristic parameters including power in the audio parameter extraction section 2.

そこで求められたパラメータはパタンバツフア
メモリ7に収納される。ルール記憶部8には「ゼ
ロ」〜「キユー」それぞれのデータごとに音声区
間が確実に検出できるような条件が定めてある。
The parameters thus determined are stored in the pattern buffer memory 7. In the rule storage unit 8, conditions are defined so that the voice section can be reliably detected for each data of "Zero" to "Kyuu".

たとえば「ロク」に対しては EROKU<Eio ……(1) NP=2 ……(2) (P′(t)=P(t)−P(t−1) (t=2,
……T)とし、 P′(x)×P′(x+1)<0かつP′(x)>0な
らば
時刻xで極大値をとるとする) が考えられる。EROKUはパタン「ロク」に対して
あらかじめ定めたエネルギー閾値であり、Eio
入力音声に対して仮に検出された音声区間に対す
るパワーの積分である。また、NPは極大値をと
るxの数、P(t)は時刻tにおけるパワー、T
は入力音声の時間長である。
For example, for "Roku", E ROKU <E io ...(1) NP=2 ...(2) (P'(t)=P(t)-P(t-1) (t=2,
...T), and if P'(x)×P'(x+1)<0 and P'(x)>0, it takes the maximum value at time x). E ROKU is an energy threshold predetermined for the pattern "ROKU", and E io is an integral of power for a speech section temporarily detected for the input speech. Also, NP is the number of x that takes the maximum value, P(t) is the power at time t, T
is the time length of the input audio.

すなわち、ここでの条件とは「パワーの積分が
EROKUより大きく、かつ極大値を2つ持つ」と言
い換えることができる。
In other words, the condition here is ``if the integral of power is
It can be rephrased as "is larger than E ROKU and has two maximum values."

「ロク」の後半部分「ク」は、無声化すると、
パワーのレベルがかなり下がるため、標準的な音
声検出レベルでは「ク」は検出できない。そこで
そのような場合を防ぐために“極大値を2つ持
つ”という条件をルールとして与えるわけであ
る。
When the second half of “Roku”, “ku”, is devoiced,
Because the power level is so low, the "ku" cannot be detected using standard audio detection levels. Therefore, in order to prevent such cases, the condition "having two maximum values" is given as a rule.

次の処理はこのような全登録パタンに対するル
ールと登録パタンとから最適閾値学習部9におい
て最適な閾値を求めることである。最適閾値学習
部9の構成は第2図1に示される通りである。登
録パタン「ロク」を例にとつて以下にその処理を
説明する。仮閾値記憶部91に記憶されている閾
値を初期値とし、入力された登録データ「ロク」
に対するルールを用いて、そのルールを満足する
閾値を閾値決定部92において求める。第2図2
は閾値決定部92における「ロク」の音声検出の
様子を示している。LSは仮閾値LFはルールを満
足する閾値であり、閾値は極の数NP=2になる
まで徐々に下げられる。
The next process is to find an optimal threshold value in the optimal threshold value learning section 9 from the rules for all registered patterns and the registered patterns. The configuration of the optimal threshold value learning section 9 is as shown in FIG. 2. The process will be explained below using the registered pattern "Roku" as an example. The threshold value stored in the temporary threshold value storage unit 91 is set as the initial value, and the input registered data "ROKU" is set as the initial value.
Using a rule for , the threshold determining unit 92 determines a threshold that satisfies the rule. Figure 2 2
2 shows how the threshold value determination unit 92 detects the sound of "Roku". L S is a provisional threshold, L F is a threshold that satisfies the rule, and the threshold is gradually lowered until the number of poles NP = 2.

閾値決定部92の構成は第2図3に示す通りで
ある。
The configuration of the threshold determining section 92 is as shown in FIG. 2.

マイクロプロセツサ921は第2図4に示すフ
ローチヤートに従つて動作し、最終的に「ロク」
に対する閾値LFが決定される。
The microprocessor 921 operates according to the flowchart shown in FIG.
A threshold value for L F is determined.

第2図4中のブロツクAでは、従来の方法、た
とえば、前記の特願昭58−156098号明細書に記載
されている方法、によつて音声検出を行い始端
(TS)と終端(TF)を求める処理を行う。またLV
は一度に下げられる閾値の値である。このように
して求められた閾値は閾値記憶部93に記憶され
る。同様にして全登録パタンについてそれぞれの
ルールを満足する閾値が閾値記憶部93に記憶さ
れる。
In block A in FIG. 2, voice detection is performed by a conventional method, for example, the method described in the above-mentioned Japanese Patent Application No. 156098/1985, and detects the start end (T S ) and the end end (T S ). F ). Also L V
is the threshold value that is lowered at once. The threshold value determined in this manner is stored in the threshold value storage section 93. Similarly, thresholds that satisfy the respective rules for all registered patterns are stored in the threshold storage section 93.

全登録データに対して閾値が決定すると、次に
最適閾値決定部94において最適の閾値を決定す
る。
Once the threshold values have been determined for all registered data, the optimal threshold value determination unit 94 then determines the optimal threshold value.

ここでは、たとえば、閾値記憶部93における
閾値の最小値をとることが考えられる。最適閾値
決定部94は最小値検出回路により構成され、こ
こで得られた最適閾値は音声区間検出部4に記憶
される。
Here, for example, it is possible to take the minimum value of the threshold values in the threshold value storage section 93. The optimum threshold value determination section 94 is constituted by a minimum value detection circuit, and the optimum threshold value obtained here is stored in the speech section detection section 4.

以上説明したように、第1ステツプでは入力さ
れた登録データに対して閾値を適応させる機能を
持ち、従来方法に比べ、より話者に適した閾値が
得られる。
As explained above, the first step has a function of adapting the threshold value to the input registered data, and as compared to the conventional method, a threshold value more suitable for the speaker can be obtained.

第2ステツプでは、スイツチS3,S4をAに
切り換えた後、第1ステツプで求められた最適閾
値を用いて登録パタンの音声検出を音声区間検出
部4において行う。登録パタンバツフアメモリ7
には各登録データのパラメータが既に記憶されて
いるためそれを利用することができる。
In the second step, after switching the switches S3 and S4 to A, the voice section detecting section 4 detects the voice of the registered pattern using the optimum threshold value determined in the first step. Registered pattern buffer memory 7
Since the parameters of each registered data are already stored in the , it is possible to use them.

検出部4における処理は前述と同様に特願昭58
−156098号明細書に記載されている音声検出器を
用いる事ができる。
The processing in the detection unit 4 is similar to that described above.
It is possible to use the voice detector described in -156098.

第3ステツプでは、S2,S3,S4をBに切
り換えた後、認識パタンの音声区間検出を行う。
マイク1より発声入力された「ニイサン」は登録
時と同様に音声パラメータ抽出部2で特徴パラメ
ータが抽出される。
In the third step, after switching S2, S3, and S4 to B, the speech section of the recognition pattern is detected.
The characteristic parameters of "Nii-san" inputted through the microphone 1 are extracted by the voice parameter extraction unit 2 in the same manner as at the time of registration.

次に第1ステツプど求められた最適閾値を用い
て、音声区間検出部4において第2ステツプと同
様に検出を行う。
Next, using the optimum threshold value determined in the first step, detection is performed in the voice section detecting section 4 in the same manner as in the second step.

以上、本発明による一実施例を説明しかが、扱
うデータは、数字に限る必要はなく、音声ならば
何でも適用できることは自明である。また、説明
中でルールの一例をあげたが、ルールは閾値を決
定するための条件であるという意味において、ル
ールの記述内容、記述法は本発明の本質を何らか
えるものではない。したがつて、本発明に含まれ
る。
Although one embodiment of the present invention has been described above, it is obvious that the data to be handled need not be limited to numbers and can be applied to any voice data. Furthermore, although an example of a rule has been given in the explanation, the description content and method of writing the rule do not change the essence of the present invention in any way in the sense that the rule is a condition for determining a threshold value. Therefore, it is included in the present invention.

また、最適閾値は全登録データの閾値すべてを
用いて決める必要はなく、そのうちの1個、ある
いは複数個の登録データに対する閾値のみを用い
て決める事も考えられる。
Furthermore, it is not necessary to determine the optimal threshold value using all the threshold values of all the registered data, and it is also possible to determine the optimal threshold value using only the threshold value for one or more of the registered data.

さらに最適閾値決定部94において複数個の閾
値から最適閾値を求める方法は最小値の他に平均
値をとる方法、重み付け平均値をとる方法などが
考えられる。
Further, as a method for determining the optimal threshold value from a plurality of threshold values in the optimal threshold value determination unit 94, there are a method of taking an average value in addition to the minimum value, a method of taking a weighted average value, and the like.

(発明の効果) 今まで述べてきたように、本発明による音声区
間検出装置では、音声の回数が従来とまつたく変
わらないにもかかわらず、以下の利点を生じる。
(Effects of the Invention) As described above, the voice section detection device according to the present invention has the following advantages, even though the number of voices is not significantly different from that of the conventional method.

まず第1に、発声の個人性に対応できる。すな
わち登録パタンから最適の閾値を学習するため、
固定閾値型音声検出装置に比べると、より発声者
の音声に適応した閾値を検出に使うことができ
る。
First of all, it can accommodate the individuality of vocalizations. In other words, in order to learn the optimal threshold from the registered pattern,
Compared to a fixed threshold type voice detection device, a threshold value that is more suited to the voice of the speaker can be used for detection.

第2に、登録パタンに対する閾値から求めた最
適閾値によつて再び登録パタンの音声区間検出を
行うという点で、従来の方法より登録パタンの誤
検出が減るといえる。このことは登録パタンの質
の悪さに起因する認識エラーを大幅に減少させる
ことを示している。
Second, since the voice section detection of the registered pattern is performed again using the optimal threshold value determined from the threshold value for the registered pattern, it can be said that false detection of the registered pattern is reduced compared to the conventional method. This shows that recognition errors caused by poor quality registered patterns can be significantly reduced.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は従来の音声区間検出装置を示すブロツ
ク図、第2図1は最適閾値学習部9を示すブロツ
ク図、第2図2は登録パタン「6」に対する閾値
決定法を示す図、第2図3は閾値決定部92の回
路図、第2図4は閾値決定部921のフローチヤ
ート、第3図は、本発明による音声区間検出装置
を示すブロツク図である。 図において、1……マイク、2……音声パラメ
ータ抽出部、3……閾値記憶部、4……音声検出
部、5……登録パタンバツフアメモリ、6……認
識パタンバツフアメモリ、7……パタンバツフア
メモリ、8……ルール記憶部、9……最適閾値学
習部、91……仮閾値記憶部、92……閾値決定
部、93……閾値記憶部、94……最適閾値決定
部、921……マイクロプロセツサ、922……
I/Oポート、923……メモリ、S1,S2,
S3,S4……スイツチ。
FIG. 1 is a block diagram showing a conventional speech interval detection device, FIG. 2 1 is a block diagram showing an optimal threshold value learning section 9, FIG. FIG. 3 is a circuit diagram of the threshold determining section 92, FIG. 2 is a flowchart of the threshold determining section 921, and FIG. 3 is a block diagram showing a voice section detection apparatus according to the present invention. In the figure, 1...Microphone, 2...Audio parameter extraction section, 3...Threshold value storage section, 4...Audio detection section, 5...Registered pattern buffer memory, 6...Recognition pattern buffer memory, 7... . . . Pattern buffer memory, 8 . , 921...Microprocessor, 922...
I/O port, 923...Memory, S1, S2,
S3, S4...Switch.

Claims (1)

【特許請求の範囲】[Claims] 1 入力された音声からパワーを含むパラメータ
時系列を抽出する音声パラメータ抽出部と、前記
音声パラメータ抽出部で得られた音声パラメータ
を記憶するパタンバツフアメモリと、登録パタン
ごとにあらかじめ定められた判定条件を記憶する
判定条件記憶部と、前記パタンバツフアメモリ内
の登録パタンと判定条件記憶部出力とに適合する
閾値を定める閾値決定部と、前記閾値決定部にお
いて得られた閾値を用いて、前記パタンバツフア
メモリ内の登録パタンおよび音声パラメータ抽出
部で得られた認識パタンの音声区間の検出を行う
音声区間検出部と、前記音声区間検出部により音
声検出された登録パタンを記憶する登録パタンバ
ツフアメモリと、同じく前記音声区間検出部によ
り音声検出された認識パタンを記憶する認識パタ
ンバツフアメモリとを有することを特徴とする音
声区間検出装置。
1. A voice parameter extraction unit that extracts a parameter time series including power from input voice, a pattern buffer memory that stores the voice parameters obtained by the voice parameter extraction unit, and a predetermined judgment for each registered pattern. using a judgment condition storage unit that stores conditions, a threshold determination unit that determines a threshold that matches the registered pattern in the pattern buffer memory and the output of the judgment condition storage unit, and the threshold obtained in the threshold determination unit, a registered pattern in the pattern buffer memory and a voice section detecting section that detects the voice section of the recognition pattern obtained by the voice parameter extracting section; and a registered pattern that stores the registered pattern voice detected by the voice section detecting section. A speech segment detection device comprising: a buffer memory; and a recognition pattern buffer memory that also stores a recognition pattern detected by the speech segment detection section.
JP59200209A 1984-09-25 1984-09-25 Voice section detector Granted JPS6177900A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP59200209A JPS6177900A (en) 1984-09-25 1984-09-25 Voice section detector

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP59200209A JPS6177900A (en) 1984-09-25 1984-09-25 Voice section detector

Publications (2)

Publication Number Publication Date
JPS6177900A JPS6177900A (en) 1986-04-21
JPH0570837B2 true JPH0570837B2 (en) 1993-10-05

Family

ID=16420620

Family Applications (1)

Application Number Title Priority Date Filing Date
JP59200209A Granted JPS6177900A (en) 1984-09-25 1984-09-25 Voice section detector

Country Status (1)

Country Link
JP (1) JPS6177900A (en)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
DE19939102C1 (en) * 1999-08-18 2000-10-26 Siemens Ag Method and arrangement for recognizing speech

Also Published As

Publication number Publication date
JPS6177900A (en) 1986-04-21

Similar Documents

Publication Publication Date Title
EP0614169B1 (en) Voice signal processing device
JP2000330587A (en) Method and device for recognizing speech
JPH0570837B2 (en)
EP0255529A4 (en) METHOD FOR COMPARING SEQUENCES FOR THE RECOGNITION OF WORDS IN HIGH ENVIRONMENTAL NOISE ENVIRONMENTS.
JP2754960B2 (en) Voice recognition device
JPS62141595A (en) Voice detection system
JPH02293797A (en) voice recognition device
JPH0376471B2 (en)
JP3096564B2 (en) Voice detection device
JPS6326879Y2 (en)
JP2901976B2 (en) Pattern matching preliminary selection method
JPS61260299A (en) Voice recognition equipment
JPH0443277B2 (en)
JPH05210763A (en) Automatic learning type character recognizing device
JP2891259B2 (en) Voice section detection device
JPS62237498A (en) Voice section detecting method
JPS61259296A (en) Voice section detection system
JPS58159599A (en) Monosyllabic voice recognition system
JPS61113099A (en) Voice section detecting system for voice recognition equipment
JPS6285393A (en) Pattern recognizing device with rejecting function
JPS60205600A (en) Voice recognition equipment
JPH10336314A (en) Method and apparatus for detecting a tone, such as DTMF, and a telephone device including such a detection apparatus
JPS6199196A (en) Voice recognition processor
JPH11119793A (en) Voice recognition device
JPS58152299A (en) Voice input controller