JPS5882300A - Voice recognition system - Google Patents

Voice recognition system

Info

Publication number
JPS5882300A
JPS5882300A JP56180850A JP18085081A JPS5882300A JP S5882300 A JPS5882300 A JP S5882300A JP 56180850 A JP56180850 A JP 56180850A JP 18085081 A JP18085081 A JP 18085081A JP S5882300 A JPS5882300 A JP S5882300A
Authority
JP
Japan
Prior art keywords
feature
feature pattern
pattern
section
extracted
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP56180850A
Other languages
Japanese (ja)
Inventor
大岡 明裕
豊 和田
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sumitomo Electric Industries Ltd
Original Assignee
Sumitomo Electric Industries Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sumitomo Electric Industries Ltd filed Critical Sumitomo Electric Industries Ltd
Priority to JP56180850A priority Critical patent/JPS5882300A/en
Publication of JPS5882300A publication Critical patent/JPS5882300A/en
Pending legal-status Critical Current

Links

Abstract

(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。
(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.

Description

【発明の詳細な説明】 本発明は音声認識方式に係る。[Detailed description of the invention] The present invention relates to a voice recognition method.

従来の音声g織方式においては、音声特徴抽出パターン
のメモリ容量が多くなシ、これらのメモリ内容で識別す
るための処理時間が長くなシ、処理速度が上らない欠点
があった。第1図に従来の音声認識方式を示す構成図を
示す。
The conventional voice weaving method has disadvantages such as a large memory capacity for voice feature extraction patterns, a long processing time for identification based on the memory contents, and an inability to increase processing speed. FIG. 1 shows a block diagram showing a conventional speech recognition system.

111図において、音声入力が前処理部1に入力され、
前処理部1は入力信号を一定レベルに増巾すると共に音
声区間の切出しを行ない、音声特徴抽出部2へ出力する
。音声特徴抽出部2は音声入力信号から、(1)$9形
予測係数(LPC)法、偉)周波数帯域別パワー法、(
3)偏自己相関係数法、(4)自己相関係数法醇によっ
て音声特徴パターンの抽出をタイミング制御部3で指定
されるqIf徴パターン抽出周期ごとに行なう。そして
音声i!識を行なうためには、始めに登録モードとして
音声特徴抽出部2の出力を標準パターンメモリ部5へ点
線で示す如く入力し、使用用途に−よって特定される用
語の標準パターンを標準パターンメモリ部5に予め記録
する。次いで使用モードに切替え、音声特徴抽出部2の
出力を実線で示す如く入力音声特徴パターンメモリ部4
へ入力し、音声識別をするために記憶する。識別処理部
6は入力音声特徴パターンメモリ部4に蓄えられた特徴
抽出パターンと標準ハターンメモリ部5に蓄えられてい
る標準パターンとを類似度計算法あるいはパターン!ツ
チング法によシ比較し、一致した標準パターンに基づい
て入力された音声を識別出力する。
In FIG. 111, voice input is input to the preprocessing unit 1,
The preprocessing section 1 amplifies the input signal to a certain level, cuts out a speech section, and outputs it to the speech feature extraction section 2. The audio feature extraction unit 2 extracts the audio input signal from the audio input signal using (1) the $9 type prediction coefficient (LPC) method, (1) the frequency band power method, (
3) Partial autocorrelation coefficient method and (4) Autocorrelation coefficient method, audio feature patterns are extracted every qIf feature pattern extraction period designated by the timing control unit 3. And audio i! In order to perform this recognition, first input the output of the voice feature extraction section 2 into the standard pattern memory section 5 as shown by the dotted line in the registration mode, and then input the standard pattern of the term specified by the purpose of use into the standard pattern memory section. 5 in advance. Next, the mode is switched to the use mode, and the output of the audio feature extracting unit 2 is stored in the input audio feature pattern memory unit 4 as shown by the solid line.
and store it for voice identification. The identification processing unit 6 uses a similarity calculation method or a pattern! A comparison is made using the matching method, and the input voice is identified and output based on the matched standard pattern.

処で、音声特徴抽出部2では、特徴パターン抽出周期を
例えば10 ms@cとし、線形予測係数法を用い次数
10次で予測を行ない且つ次数当り8ビツトで特徴パタ
ーンを表示すると、単@gg*の場合通常、1@は0.
5〜1.Omecテあるから0.5〜1.0 kbyt
e/語になる。これらを記憶する入力音声特徴パターン
メモリ部4及び標準パターンメそり郁Sの記憶容量はほ
う大になる。特に多くの語の標準パターンを記憶する標
準パターンメモリ部50記憶容量はぼり大になる。
Here, in the audio feature extraction section 2, if the feature pattern extraction period is set to 10 ms@c, prediction is performed at the 10th order using the linear prediction coefficient method, and the feature pattern is displayed at 8 bits per order, then a single @gg In the case of *, 1@ is usually 0.
5-1. 0.5 to 1.0 kbyte because there is Omec
e/becomes a word. The storage capacity of the input voice feature pattern memory section 4 and the standard pattern memory S for storing these becomes larger. In particular, the storage capacity of the standard pattern memory section 50, which stores standard patterns of many words, becomes enormous.

一般に音声パターンは先行する子音区間と、これKlE
行する母音区間とからなるが、母音区間については特徴
パターンの変化は少ないにもかかわらず特徴パターン抽
出周期毎に近似した特徴パターンが入力音声特徴パター
ンメモリ部4及び標準パターンメモリ部Sへ全て配憶さ
れねばならないととKなシ、これらを記憶するメモリの
数はほう大となった。従って識別処理に時間がか\多処
理速度があがらないという欠点があった。
In general, the phonetic pattern consists of the preceding consonant section and this KlE
Although there are few changes in the characteristic patterns for vowel intervals, the approximated characteristic patterns are all allocated to the input speech characteristic pattern memory section 4 and the standard pattern memory section S at each feature pattern extraction cycle. The number of memories needed to store these data has become larger and larger. Therefore, there is a drawback that the identification process takes a long time and the multi-processing speed cannot be increased.

本発明は以上述べた欠点を除き、メモリ部の容量が著し
く少なくて済み且つ処理速度が上がる音声l!職方式を
提供することを目的とする。斯かる目的を達成する零発
W14の$ll!itは、音声入力を一定レベルに増巾
する前処理部と、前処理部の出力信号から周期的に特徴
パターンを抽出する音声特徴抽出部と、該音声特徴抽出
部によって考次抽出される特徴パターンと既に抽出され
かつ選択された特徴パターンとの変化が特定の値よ〕大
きいか否かを計算し、大きいときに上記抽出された特徴
パターンな配憶させるためのクロックを出す距離計算部
と、骸距離計算部がらのクロックによって上記抽出され
た特徴パターンを上記の既に抽出されかつ選択された特
徴パターンの代F)K記憶する選択特徴パターンレジス
タと、一つの特徴パターンが選択された後次の特徴パタ
ーンが選択されるまでの特徴パターン抽出周期の回数即
ち圧縮回数を計数する圧縮数カウンタと、前記距離計算
部からのクロックによシ上配の選択された特徴パターン
と上記圧縮数カウンタの内容とを記憶する久方音声特徴
パターンメモリ部と、各用語の標準パターン及び圧縮回
数を記憶する標準パターンメモリ部と、上記入力音声特
徴パターンメモリ部に記憶された特徴パターン及び圧縮
回数と標準パターンメ毫す部に記憶された標準パターン
及び圧縮回数とを比較して音声入力を認識する識別処理
部とを備えたことを特徴とする。
The present invention eliminates the above-mentioned drawbacks, requires significantly less memory capacity, and improves processing speed. The purpose is to provide a job format. Zero-shot W14 that achieves this purpose is $ll! It consists of a preprocessing unit that amplifies the audio input to a certain level, an audio feature extraction unit that periodically extracts feature patterns from the output signal of the preprocessing unit, and features that are sequentially extracted by the audio feature extraction unit. a distance calculation unit that calculates whether or not the change between the pattern and the already extracted and selected feature pattern is greater than a specific value, and outputs a clock for storing the extracted feature pattern when the change is greater than a specific value; , a selected feature pattern register that stores the extracted feature pattern in place of the already extracted and selected feature pattern by the clock of the skeleton distance calculation unit; a compression number counter that counts the number of feature pattern extraction cycles, that is, the number of compressions until the feature pattern is selected; a standard pattern memory section that stores the standard pattern of each term and the number of times of compression; and a standard pattern memory section that stores the standard pattern of each term and the number of times of compression, and the characteristic pattern, the number of times of compression, and the standard pattern stored in the input voice feature pattern memory section. The present invention is characterized by comprising an identification processing section that recognizes voice input by comparing the standard pattern stored in the printing section and the number of times of compression.

本発明による音声il!識方式の一実施例を第2図に示
す。第2図において、前処理部1と音声特徴抽出部2で
の処理は第1因の場合とと〈K変らない。従って前処理
部1では入力音声を一定レベルに増巾し且つ音声区間を
検出する。音声特徴抽出部2では特徴抽出が行なわれ、
若し線形予測係数法で特徴抽出が行なわれ特徴パターン
抽出周期を10m5ec、線形予測係数の次数を10次
とし、次数当り8ビツトとすると、特徴パターンのビッ
ト数は上記で説明した通シ0.5〜1.0 kbyte
/語となる。
Audio il according to the invention! An example of the identification method is shown in FIG. In FIG. 2, the processing in the preprocessing unit 1 and the audio feature extraction unit 2 is the same as in the case of the first factor. Therefore, the preprocessing section 1 amplifies the input voice to a certain level and detects the voice section. The audio feature extraction unit 2 performs feature extraction,
If feature extraction is performed using the linear prediction coefficient method, the feature pattern extraction period is 10m5ec, the order of the linear prediction coefficient is 10th, and each order is 8 bits, the number of bits of the feature pattern will be 0.5m as explained above. 5 to 1.0 kbytes
/becomes a word.

本発明の音声I!識方式では、特徴パターン抽出周期毎
に抽出された特徴パターンの内炭化の少ない特徴パター
ンは捨てて記憶せず、直前に選択して一時蓄えておいた
特徴パターンと比較し、両者の変化が設定値より大きい
か否かを距離計算部Tが計算し、大きい場合の特徴パタ
ーンだけを記憶するものでおる。即ち、第2図の距離計
算部7は、音声特徴抽出部2の出力を既に選択されて選
択特徴レジスタ8に記憶されている特徴パターンと比較
し、変化があらかじめ設定した値よシ大きい時だけ、上
記出力の特徴パターンが選択特徴レジスタ8及び入力音
声特徴パターンメモリ部4あるいは標準パターンメモリ
部5へ記憶されるためのクロックCKを出力する。
Audio I of the present invention! In the identification method, feature patterns with little internal carbonization extracted at each feature pattern extraction cycle are discarded and not stored, but are compared with the previously selected feature pattern and temporarily stored, and changes in both are set. The distance calculation unit T calculates whether the distance is larger than the value, and stores only the feature pattern when the distance is larger than the value. That is, the distance calculating section 7 in FIG. 2 compares the output of the audio feature extracting section 2 with the feature pattern that has already been selected and stored in the selected feature register 8, and only when the change is larger than a preset value. , outputs a clock CK for storing the output feature pattern into the selected feature register 8 and the input voice feature pattern memory section 4 or the standard pattern memory section 5.

尚、入力音声特徴パターンメモリ部4に一つの特徴パタ
ーンが選択された後火の特徴パターンが選択されるまで
特徴に変化がないとして選択されなかったものの数を圧
縮数カウンタ9がカウントする。圧縮数カウンタ8の内
容は選択された特徴パターンと共に入力音声特徴パター
ンメモリ部4あるいは標準パターンメそり部囮記憶され
、圧縮数カウンタ9の内容は距離計算部7の出力でクリ
アされる。識別処理部6は類似度計算あるいはパターン
マツチング法によシ、入力音声特徴パターンメモリ部4
の特徴パターンと標準パターンメモリ部5の標準パター
ンと比較し、圧縮された音声特徴パターンに対して圧縮
カウンタ9の圧縮数を重み係数として利用し、音声入力
を認識する。
After one feature pattern is selected in the input voice feature pattern memory section 4, a compression number counter 9 counts the number of unselected feature patterns, assuming that there is no change in the feature until a second feature pattern is selected. The contents of the compression number counter 8 are stored together with the selected feature pattern in the input voice feature pattern memory section 4 or the standard pattern mesori section decoy, and the contents of the compression number counter 9 are cleared by the output of the distance calculation section 7. The identification processing unit 6 uses similarity calculation or pattern matching method, and the input voice feature pattern memory unit 4
The feature pattern is compared with the standard pattern in the standard pattern memory section 5, and the compression number of the compression counter 9 is used as a weighting factor for the compressed voice feature pattern to recognize the voice input.

第3図はGAKKO(学校)と発音した場合の音声波形
を示している。第3図に示されるように、母音ム、0の
定常部分は一つの発生区間のはv14を占めている。母
音区間での特徴パターンは殆んど変化がないため、本発
明の方式では、距離計算部Tによって圧縮され、母音区
間で選択される特徴パターンは大巾に削減される。した
がって、メモリ部4及びSK記憶される特徴パターンは
従来の場合にくらべ、1音声区間ではソン2に削減され
た。
FIG. 3 shows the speech waveform when GAKKO (school) is pronounced. As shown in FIG. 3, the constant part of the vowel m, 0 occupies v14 of one occurrence interval. Since the feature patterns in the vowel section hardly change, in the method of the present invention, the feature patterns selected in the vowel section are compressed by the distance calculation unit T and are greatly reduced. Therefore, compared to the conventional case, the number of characteristic patterns stored in the memory unit 4 and SK has been reduced to 2 songs in one voice section.

以上説明した如く本発明によれば、音声の冗長性を除去
でき、特に多くの標準パターンを記憶している標準パタ
ーンメやり部5のメモリ容量が大巾に削減される。また
、従来のもののように特徴パターンに変化の表い場合で
もいちいち処理しなければならなかったものKくらべて
、本発明忙よるt−序認識方式では選択された特徴パタ
ーンについてのみ処理して識別すればよくその処理速度
も著しく迅速に々つた。
As explained above, according to the present invention, it is possible to eliminate redundancy in audio, and in particular, the memory capacity of the standard pattern mailing section 5, which stores a large number of standard patterns, can be greatly reduced. In addition, compared to the conventional method that requires processing every time there is a change in the feature pattern, the t-order recognition method of the present invention processes and identifies only selected feature patterns. The processing speed increased significantly.

なお、実施例では音声区間の検出を前処理部1に行わせ
ているが、これは音声特徴抽出部2に行わせてもかまわ
ない。
In the embodiment, the preprocessing section 1 detects the voice section, but the voice feature extraction section 2 may also detect the voice section.

【図面の簡単な説明】[Brief explanation of drawings]

第1図は従来の音声認識方式の構成図、第2図は本発明
による音声認識方式の構成図、第3図社音声入力波形の
一つの例を示した因である。 図面中、 1は前処理部、2は音声特徴抽出部、 3はタイミング制御部、 4は入力音声特徴パターンメモリ部、 5は標準パターンメモリ部、 6は識別処理部、 Tは距離計算部、 8は選択特徴レジスタ、 9は圧縮数カウンタである。 特許出願人  住友電気工業株式会社
FIG. 1 is a block diagram of a conventional speech recognition method, FIG. 2 is a block diagram of a speech recognition method according to the present invention, and FIG. 3 shows an example of a speech input waveform. In the drawings, 1 is a preprocessing unit, 2 is a voice feature extraction unit, 3 is a timing control unit, 4 is an input voice feature pattern memory unit, 5 is a standard pattern memory unit, 6 is a recognition processing unit, T is a distance calculation unit, 8 is a selection feature register, and 9 is a compression number counter. Patent applicant: Sumitomo Electric Industries, Ltd.

Claims (1)

【特許請求の範囲】[Claims] 音声入力を一定レベルに増巾する前処理部と、前処理部
の出力信号から周期的に特徴パターンを抽出する音声特
徴抽出部と、該音声特徴抽出部によって造次抽出される
特徴パターンと既に抽出されかつ選択された特徴パター
ンとの変化が特定の値よ〕大きいか否かを計算し、大き
いときに上記抽出された特徴パターンを記憶させるため
Oクロックを出す距離計算部と、該距離計算部からのク
ロックによって上記抽出された特徴パターンを上記の既
に抽出されかつ選択された特徴パターン0代シに記憶す
る選択特徴パターンレジスタと、一つの特徴パターンが
選択された後次の特徴パターンが選択されるまでの特徴
パターン抽出周期の回数即ち圧縮回数を計数する圧縮数
カウンタと、前記距離計算部からのクロックによシ上記
の選択された特徴パターンと上記圧縮数カウンタの内容
とを記憶する入力音声特徴パターンメモリ部と、各用語
の標準パター、ン及び圧縮回数を記憶する標準パターン
メモリ部と上記入力音声特徴パターンメモリ部に記憶さ
れた特徴パターン及び圧縮回数と標準パターンメモリ部
に記憶された標準パターン及び圧縮回数とを比較して音
声入力を認識する識別処理部とを健えたことを特徴とす
る音声認識方式。
a pre-processing section that amplifies the audio input to a certain level; an audio feature extraction section that periodically extracts feature patterns from the output signal of the pre-processing section; and a feature pattern that is sequentially extracted by the audio feature extraction section. a distance calculation unit that calculates whether or not a change from the extracted and selected feature pattern is larger than a specific value, and outputs an O clock to store the extracted feature pattern when the change is larger than a specific value; A selection feature pattern register that stores the extracted feature pattern in the already extracted and selected feature pattern 0 by a clock from the section, and after one feature pattern is selected, the next feature pattern is selected. a compression number counter that counts the number of feature pattern extraction cycles, that is, the number of compressions until the feature pattern is extracted; and an input that stores the selected feature pattern and the contents of the compression number counter according to a clock from the distance calculation section. an audio feature pattern memory section, a standard pattern memory section that stores the standard pattern of each term, the number of compressions, and the number of compressions stored in the input audio feature pattern memory section; A speech recognition method characterized by having an identification processing section that recognizes speech input by comparing it with a standard pattern and the number of times of compression.
JP56180850A 1981-11-11 1981-11-11 Voice recognition system Pending JPS5882300A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP56180850A JPS5882300A (en) 1981-11-11 1981-11-11 Voice recognition system

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP56180850A JPS5882300A (en) 1981-11-11 1981-11-11 Voice recognition system

Publications (1)

Publication Number Publication Date
JPS5882300A true JPS5882300A (en) 1983-05-17

Family

ID=16090448

Family Applications (1)

Application Number Title Priority Date Filing Date
JP56180850A Pending JPS5882300A (en) 1981-11-11 1981-11-11 Voice recognition system

Country Status (1)

Country Link
JP (1) JPS5882300A (en)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS60158498A (en) * 1984-01-27 1985-08-19 株式会社リコー pattern matching device
JPS6227798A (en) * 1985-07-29 1987-02-05 株式会社日立製作所 Voice recognition equipment

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS60158498A (en) * 1984-01-27 1985-08-19 株式会社リコー pattern matching device
JPS6227798A (en) * 1985-07-29 1987-02-05 株式会社日立製作所 Voice recognition equipment

Similar Documents

Publication Publication Date Title
US4813074A (en) Method of and device for segmenting an electric signal derived from an acoustic signal
TW347619B (en) A communication system and method using a speaker dependent time-scaling technique a method for time-scale modification of speech using a modified version of the Waveform Similarity based Overlap-Add technique (WSOLA).
US3943295A (en) Apparatus and method for recognizing words from among continuous speech
JPS6360919B2 (en)
EP0810583A3 (en) Speech recognition system
US5095508A (en) Identification of voice pattern
US4790017A (en) Speech processing feature generation arrangement
JPS6131880B2 (en)
JPS6069695A (en) Initial consonant segmentation method
JP2655637B2 (en) Voice pattern matching method
JPS63223696A (en) Audio pattern creation method
JPS6363919B2 (en)
JPS63318600A (en) Voice recognition system
JPS61143800A (en) Voice recognition equipment
JPH04198999A (en) Method for searching minimum value of matching distance in speech recognition
JPS59204099A (en) Voice recognition system
JPS636599A (en) Word preselection system
JPS6195399A (en) Voice pattern matching method
JPS61137197A (en) Continuous word voice recognition equipment
PFEIFER Isolated word phoneme recognition using features derived from wavefunction parameters(Isolated-word phoneme recognition using features derived from wave function parameters)
JPS6069694A (en) Segmentation of head consonant
JPS61252599A (en) Voice recognition method
JPS61200596A (en) Continuous voice recognition equipment
JPS6293000A (en) Voice recognition method
JPS5915999A (en) Monosyllable recognition equipment