JPS6334479B2 - - Google Patents

Info

Publication number
JPS6334479B2
JPS6334479B2 JP57230980A JP23098082A JPS6334479B2 JP S6334479 B2 JPS6334479 B2 JP S6334479B2 JP 57230980 A JP57230980 A JP 57230980A JP 23098082 A JP23098082 A JP 23098082A JP S6334479 B2 JPS6334479 B2 JP S6334479B2
Authority
JP
Japan
Prior art keywords
formant
circuit
filter
vowels
output
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Expired
Application number
JP57230980A
Other languages
Japanese (ja)
Other versions
JPS59123899A (en
Inventor
Megumi Motomya
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sharp Corp
Original Assignee
Sharp Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sharp Corp filed Critical Sharp Corp
Priority to JP23098082A priority Critical patent/JPS59123899A/en
Publication of JPS59123899A publication Critical patent/JPS59123899A/en
Publication of JPS6334479B2 publication Critical patent/JPS6334479B2/ja
Granted legal-status Critical Current

Links

Description

【発明の詳細な説明】[Detailed description of the invention]

<技術分野> 本発明は簡易型の音声認識装置に関するもので
ある。 <従来技術> 日本語の音声の特徴はすべての音節に母音が付
加されているので、限定された文章や単語では、
子音を無視して母音の系列としたときでもかなり
の部分が解読できる。母音はエネルギーも大き
く、持続時間も長いため騒音に対しても影響が少
ない。そのため日本語音声中から母音を正確に認
識することは極めて重要である。ところが、母音
の認識で重要なホルマントは性別、年令により異
なり又連続音声中の母音は前後の音韻や発声速度
によつても異なり、5母音の範囲がかなり広がつ
ている。 これを考慮に入れると次の2つの方法を考える
必要がある。 (i) フイルターを多数使用して周波数分解能を上
げる。その上でホルマント周波数の正規化を行
なう。この方法は後の処理が複雑になる欠点が
あるが、ばらつきの範囲がかなり広くても適用
が可能である。 (ii) フイルターの周波数範囲を広げて、ばらつき
を吸収する様に考える。 上記(ii)の方法において統計的な実験の結果、5
母音の識別には5個のフイルターで充分であるこ
とがわかつた。1例として下記に示す。
<Technical Field> The present invention relates to a simple speech recognition device. <Prior art> A characteristic of Japanese speech is that a vowel is added to every syllable, so in a limited number of sentences and words,
Even when consonants are ignored and the vowel sequence is used, a large portion can still be deciphered. Vowels have a lot of energy and last a long time, so they are less affected by noise. Therefore, it is extremely important to accurately recognize vowels in Japanese speech. However, the formant, which is important for vowel recognition, differs depending on gender and age, and the vowels in continuous speech also differ depending on the preceding and following phonemes and the rate of speech, and the range of five vowels has expanded considerably. Taking this into consideration, it is necessary to consider the following two methods. (i) Increase frequency resolution by using multiple filters. Then, the formant frequency is normalized. Although this method has the disadvantage of complicating subsequent processing, it can be applied even if the range of variation is quite wide. (ii) Consider expanding the frequency range of the filter to absorb variations. Results of statistical experiments in method (ii) above, 5
Five filters were found to be sufficient for vowel identification. An example is shown below.

【表】 フイルターNo.1、No.2は第1ホルマント領域、
フイルターNo.3は第1又は第2ホルマントの共存
領域であり母音により異なる。フイルターNo.4、
No.5は第2ホルマント領域である。 これら5個のフイルターの出力に適当な重み付
を行なつた後、各フイルター相互を比較すること
により母音を識別する。これはかなり幅広い範囲
にわたり5母音を識別することができる方式であ
る。しかしながら連続音声中ではこのような5母
音だけでは不充分であり、他の情報が必要であ
る。 <発明の目的> 本発明はこのような点に鑑み、上記(ii)の方式に
おけるフイルターを用いて効果的に実現できる構
成簡単な簡易型の音声認識装置を提供する。 <実施例> 以下図面に従つて本発明の一実施例を説明す
る。 第1図は本案音声認識装置の一実施例を示すブ
ロツク図である。 マイクロホン1により音声信号が入力され、ア
ンプ2に入る。アンプ2は自動音量調整回路を付
加されており、必要に応じて使用できる。又この
アンプ2の周波数特性は約1KHzから高域強調
(+6dB/oct)となつている。アンプ2の出力は
フイルター群3に入力される。これは上記5個の
フイルターで構成されている。各フイルターNo.1
〜No.5の出力は包絡線検出回路4に入れられる。
この他、包絡線検出回路4にはプリアンプ2から
直接入力されるものがあり合計6個で構成されて
いる。この回路は低域濾波回路を有しており、20
〜50Hzの遮断周波数に設定されている。包絡線検
出回路4の出力は母音・無声子音・無音識別回路
5及びフイルター選択回路6に入力される。母
音・無声子音・無音識別回路5は包絡線検出回路
4の各出力を相互比較することにより5母音及び
無声子音、無音(結果的に有声子音も含む)を識
別する回路であり、識別状態は3ビツトのデイジ
タル信号として出力される。 一方フイルター選択回路6は上記の様に5個の
フイルター出力に対応した包絡線検出回路4の
各々の出力を二つのグループ即ちフイルターNo.
1、No.2、No.3とフイルター(No.3)、No.4、No.
5(後者グループのフイルターNo.3は、前者グル
ープの選択により含めるか又は含めないかが決定
される)に分離して、二つのグループ内での出力
最大のフイルターを選択し、スイツチ回路7を制
御することにより2つの入力としてホルマント周
波数遷移部分検出回路8に入力される。 即ちフイルターNo.1、No.2、No.3を第1ホルマ
ント領域、フイルターNo.3、No.4、No.5を第2ホ
ルマント領域として2つのグループ分けを行な
い、第1ホルマント領域のフイルター3個の中で
最大出力を示すフイルターを選択する。もし最大
出力がフイルターNo.3でなければNo.3を含む第2
ホルマント領域のフイルター3個の中で最大出力
を出すフイルターを選択する。もし第1ホルマン
ト領域のフイルターでNo.3が最大値を示す場合
は、第2ホルマント領域はNo.4、No.5から出力の
大きい方を選択する。以上の方法により有効に第
1及び第2ホルマントが存在するフイルターを自
動的に選択できる。 このフイルターの出力は零交差計数回路に入力
され、その出力の時間的変化が観測される。この
出力が第1及び第2ホルマント周波数を推定して
いることも実験的にも知られている。従つてこの
零交差数の時間的変化はホルマントのトランジエ
ント部分を表現しているため定常母音からの時間
変化の特徴をとり、定常母音に入る前のホルマン
トのトランジエントを「入りわたり」、定常母音
の後のホルマントトランジエントの「出わたり」
であり、音節や子音・母音・子音連鎖として特徴
が表われる。 第2図はホルマント周波数遷移部分検出回路8
を示すもの(第1ホルマント部、第2ホルマント
部も同じ)であり、選択されたフイルター出力は
まず零交差計数回路8aに入力される。これは一
定時間(約10ms)の時間内で音声波形が零を何
度横切るかを計数する回路である。その次段にそ
れらの計数結果を一時記憶する回路8bがあり
(約200ms)、これに接続された判定回路8cに
よつて、その範囲内での計数の変化度合により定
常母音と遷移部分の判定を行ない、さらに変化の
度合、方向、変化時間等の特徴を抽出することに
より子音部と母音部の「わたり」の特徴を抽出す
る。 このホルマント周波数遷移部分検出回路8の出
力と母音・無声子音・無音識別回路5の出力は識
別単位判定回路9に入力される。これにより母
音、無声子音、無音と有声子音又はホルマント遷
移部分等から識別単位を選び決定する。ここで、
識別単位とは日本語の種々の母音、子音、及び母
音と子音間に存在する過渡的な部分(わたり)等
を示す音声を構成する一種の単位である。例え
ば、「た」という日本語の場合、「ta」という発音
がなされるわけであるが、「t」の破裂音部と
「a」の母音定常部との間にスペクトル(特にホ
ルマント)の急激に変化する非定常な区間が存在
するがこの部分も1つの識別単位である。識別単
位標準パターン記憶回路10はこれら識別単位の
特徴パターンを予じめ記憶し、ホルマント周波数
遷移部分検出回路8の出力及び母音・無声子音・
無音識別回路5の出力と照合され識別単位を選択
する。 この選択によつて、識別単位に該当する出力に
判別される。 識別単位判定回路9の出力は識別系照合回路1
1に入力される。照合される識別単位別の標準パ
ターンは多数のサンプルから統計的に決定された
ものをメモリー回路12に記憶している。上記識
別系列照合回路11では識別単位判定回路9から
の識別単位に該当するデータとメモリー回路から
の識別単位別の標準パターンとを比較して、その
時のデータがどの識別単位に相当するかを判定す
る。そして判別回路13によつて誤つた識別対等
の認識情報は除去され、認識すべき音声の識別系
列の相互比較により最も類似した識別系列を選択
し、連続音声中の5母音を判定認識する。 なお、以上の装置は汎用の8ビツトマイクロコ
ンピユータ等により構成することも可能である。 <発明の効果> 以上のように本発明は、5個のフイルター回路
を用いて有効に連続音声中の5母音を認識できる
ものであり、装置としても簡単である。また、フ
イルター出力の相互比較により予じめ第1及び第
2ホルマントが存在するフイルターを選択し、そ
れに零交差計数回路を接続することにより、耐雑
音性のあるホルマント推定が可能であり、又時間
分解能の優れた方式が複合され有用である。
[Table] Filter No. 1 and No. 2 are in the first formant region,
Filter No. 3 is the coexistence region of the first or second formant, which differs depending on the vowel. Filter No.4,
No. 5 is the second formant region. After appropriately weighting the outputs of these five filters, the vowels are identified by comparing the filters with each other. This is a method that can identify five vowels over a fairly wide range. However, in continuous speech, such five vowels alone are insufficient, and other information is required. <Object of the Invention> In view of the above points, the present invention provides a simple speech recognition device with a simple configuration that can be effectively realized using the filter in the method (ii) above. <Example> An example of the present invention will be described below with reference to the drawings. FIG. 1 is a block diagram showing an embodiment of the speech recognition device of the present invention. An audio signal is input through a microphone 1 and enters an amplifier 2. Amplifier 2 is equipped with an automatic volume adjustment circuit and can be used as required. Also, the frequency characteristics of this amplifier 2 are high frequency emphasized (+6dB/octave) starting from about 1KHz. The output of amplifier 2 is input to filter group 3. This is composed of the five filters mentioned above. Each filter No.1
The output of No. 5 is input to the envelope detection circuit 4.
In addition, the envelope detection circuit 4 includes a circuit that receives direct input from the preamplifier 2, and consists of six circuits in total. This circuit has a low pass filter circuit and has a 20
The cutoff frequency is set to ~50Hz. The output of the envelope detection circuit 4 is input to a vowel/voiceless consonant/silence discrimination circuit 5 and a filter selection circuit 6. The vowel/voiceless consonant/silence discrimination circuit 5 is a circuit that discriminates five vowels, voiceless consonants, and silence (including voiced consonants as a result) by mutually comparing each output of the envelope detection circuit 4, and the discrimination state is as follows. It is output as a 3-bit digital signal. On the other hand, the filter selection circuit 6 divides each output of the envelope detection circuit 4 corresponding to the five filter outputs into two groups, that is, filter numbers, as described above.
1, No.2, No.3 and filter (No.3), No.4, No.
5 (filter No. 3 of the latter group is included or excluded depending on the selection of the former group), selects the filter with the maximum output in the two groups, and controls the switch circuit 7. As a result, the two input signals are inputted to the formant frequency transition portion detection circuit 8 as two inputs. That is, filters No. 1, No. 2, and No. 3 are divided into two groups, with filters No. 1, No. 2, and No. 3 as a first formant region, and filters No. 3, No. 4, and No. 5 as a second formant region. Select the filter with the maximum output among the three. If the maximum output is not filter No. 3, then the second
Select the filter that produces the maximum output among the three filters in the formant region. If No. 3 of the filters in the first formant region shows the maximum value, the filter with the larger output is selected from No. 4 and No. 5 for the second formant region. By the above method, it is possible to automatically select a filter in which the first and second formants are effectively present. The output of this filter is input to a zero-crossing counting circuit, and the temporal change in the output is observed. It is also known experimentally that this output estimates the first and second formant frequencies. Therefore, since this temporal change in the number of zero crossings expresses the transient part of the formant, it takes the characteristics of the temporal change from the stationary vowel, and the transient part of the formant before entering the stationary vowel is ``crossed'' and the stationary vowel is The “emergence” of formant transients after vowels
It is characterized by syllables, consonants, vowels, and consonant chains. Figure 2 shows the formant frequency transition part detection circuit 8.
(the same applies to the first formant part and the second formant part), and the selected filter output is first input to the zero crossing counting circuit 8a. This is a circuit that counts how many times the audio waveform crosses zero within a certain period of time (approximately 10ms). At the next stage, there is a circuit 8b that temporarily stores the counting results (approximately 200ms), and a judgment circuit 8c connected to this judges whether the vowel is a stationary vowel or a transitional part based on the degree of change in the counting within that range. Then, by extracting features such as the degree of change, direction, and time of change, the characteristics of the "crossing" between the consonant and vowel parts are extracted. The output of the formant frequency transition portion detection circuit 8 and the output of the vowel/voiceless consonant/silence discrimination circuit 5 are input to a discrimination unit determination circuit 9. In this way, identification units are selected and determined from vowels, voiceless consonants, voiceless and voiced consonants, formant transition parts, etc. here,
Discrimination units are a type of unit that constitutes sounds that indicate various Japanese vowels, consonants, and transitional parts (watari) that exist between vowels and consonants. For example, in the case of the Japanese word "ta", it is pronounced as "ta", but there is an abrupt change in the spectrum (particularly formant) between the plosive part of "t" and the vowel stationary part of "a". Although there is an unsteady section where the value changes, this section is also one identification unit. The identification unit standard pattern storage circuit 10 stores the characteristic patterns of these identification units in advance, and stores the output of the formant frequency transition part detection circuit 8 and vowels, voiceless consonants,
The identification unit is selected by comparing with the output of the silence identification circuit 5. This selection determines the output that corresponds to the identification unit. The output of the identification unit determination circuit 9 is sent to the identification system matching circuit 1.
1 is input. Standard patterns for each identification unit to be compared are statistically determined from a large number of samples and are stored in the memory circuit 12. The identification series matching circuit 11 compares the data corresponding to the identification unit from the identification unit determination circuit 9 with the standard pattern for each identification unit from the memory circuit, and determines which identification unit the current data corresponds to. do. Then, the discriminating circuit 13 removes the recognition information of the erroneously identified equals, selects the most similar discriminating series by comparing the discriminating series of the speech to be recognized, and recognizes the five vowels in the continuous speech. Incidentally, the above-mentioned device can also be constructed from a general-purpose 8-bit microcomputer or the like. <Effects of the Invention> As described above, the present invention can effectively recognize five vowels in continuous speech using five filter circuits, and is a simple device. In addition, by mutually comparing filter outputs and selecting in advance a filter in which the first and second formants exist, and connecting it to a zero-crossing counting circuit, it is possible to estimate formants with noise resistance, and also to It is a useful combination of methods with excellent resolution.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は本発明の一実施例を示すブロツク図、
第2図は第1図の要部を詳細に示すブロツク図で
ある。 3……フイルター群、4……包絡線検出回路、
5……母音・無声子音・無音識別回路、6……フ
イルター選択回路、7……スイツチ回路、8……
ホルマント周波数遷移部分検出回路、9……識別
単位判定回路、10……識別単位標準パターン記
憶回路、11……識別系列照合回路、12……メ
モリー回路、13……判定回路。
FIG. 1 is a block diagram showing one embodiment of the present invention;
FIG. 2 is a block diagram showing the main parts of FIG. 1 in detail. 3... Filter group, 4... Envelope detection circuit,
5... Vowel/vowel/voiceless consonant/silence identification circuit, 6... Filter selection circuit, 7... Switch circuit, 8...
Formant frequency transition portion detection circuit, 9...Identification unit determination circuit, 10...Identification unit standard pattern storage circuit, 11...Identification series matching circuit, 12...Memory circuit, 13...Identification circuit.

Claims (1)

【特許請求の範囲】 1 第1ホルマント領域及び第2ホルマント領域
を含む5個のフイルター回路を備え、 該5個のフイルター回路の出力の相互比較によ
り、5母音、無声子音、無音等を識別する手段
と、 前記5個のフイルター回路の中から第1ホルマ
ントと第2ホルマントが存在する出力最大の2個
のフイルター回路の出力を選択する手段と、 該選択された2個のフイルター回路の出力に基
づいて零交差計数によりホルマント周波数遷移部
分を検出する手段と、 前記5母音、無声子音、無音等の識別出力と前
記ホルマント周波数遷移部分の検出出力により識
別単位を判定する手段と、 該判定した識別単位の時系列から連続音声の判
定認識を行う手段と、 を設けてなることを特徴とする音声認識装置。
[Claims] 1. Five filter circuits including a first formant region and a second formant region are provided, and five vowels, voiceless consonants, silence, etc. are identified by mutual comparison of the outputs of the five filter circuits. means for selecting the outputs of the two filter circuits with the maximum outputs in which the first formant and the second formant are present from among the five filter circuits; and the outputs of the two selected filter circuits. means for detecting a formant frequency transition part by zero crossing counting based on the discrimination output of the five vowels, voiceless consonants, silence, etc. and means for determining a discrimination unit based on the detection output of the formant frequency transition part; and the determined discrimination. A speech recognition device comprising: means for performing judgment recognition of continuous speech from a time series of units;
JP23098082A 1982-12-29 1982-12-29 Voice recognition equipment Granted JPS59123899A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP23098082A JPS59123899A (en) 1982-12-29 1982-12-29 Voice recognition equipment

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP23098082A JPS59123899A (en) 1982-12-29 1982-12-29 Voice recognition equipment

Publications (2)

Publication Number Publication Date
JPS59123899A JPS59123899A (en) 1984-07-17
JPS6334479B2 true JPS6334479B2 (en) 1988-07-11

Family

ID=16916333

Family Applications (1)

Application Number Title Priority Date Filing Date
JP23098082A Granted JPS59123899A (en) 1982-12-29 1982-12-29 Voice recognition equipment

Country Status (1)

Country Link
JP (1) JPS59123899A (en)

Also Published As

Publication number Publication date
JPS59123899A (en) 1984-07-17

Similar Documents

Publication Publication Date Title
JP2891259B2 (en) Voice section detection device
JPS59123899A (en) Voice recognition equipment
JP2886879B2 (en) Voice recognition method
JPS5936759B2 (en) Voice recognition method
JPS60138599A (en) Voice section detector
JPH01158499A (en) Standing noise eliminaton system
JPH026078B2 (en)
JPS61246800A (en) Voice response switch
JPH0311478B2 (en)
JPS63191199A (en) Voiced plosive consonant identifier
JPS62289898A (en) Voiced plosive consonant identification system
JPS59170894A (en) Voice section starting system
JPS61246799A (en) Voice response switch
JPS5936299A (en) Voice recognition equipment
JPH0652479B2 (en) Speech analysis method
JPS6375800A (en) Voice recognition equipment
JPS6070497A (en) voice recognition device
JPH026079B2 (en)
Reitboeck et al. Speaker-identification with real time formant extraction
JPS61177000A (en) Audio pattern registration method
JPS60163099A (en) Voice sorter
JPS63155197A (en) Voiceless sound detection
JPS6236699A (en) voice identification device
JPH01159699A (en) Voice recognition apparatus
JPS6240497A (en) Voice pattern sorting system