JPH06295196A - Speech recognition device and signal recognition device - Google Patents
Speech recognition device and signal recognition deviceInfo
- Publication number
- JPH06295196A JPH06295196A JP5107709A JP10770993A JPH06295196A JP H06295196 A JPH06295196 A JP H06295196A JP 5107709 A JP5107709 A JP 5107709A JP 10770993 A JP10770993 A JP 10770993A JP H06295196 A JPH06295196 A JP H06295196A
- Authority
- JP
- Japan
- Prior art keywords
- voice
- unit
- signal
- recognition
- voice signal
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Abstract
(57)【要約】 (修正有)
【目的】 演算量が少なく正確な音声信号の検出ができ
る音声認識装置及び信号認識装置。
【構成】 音声認識装置10は、音声信号検出用の特徴
抽出部14と、予め何種類かに分類されたノイズ信号と
音声信号によって学習がなされるとともに、特徴抽出部
14から送られた特徴パラメータを基にニューラルネッ
ト計算によりノイズと音声信号を判別する音声信号検出
部15と、特徴抽出部14から全てのフレームの特徴パ
ラメータが送られてきた特徴パラメータに所定の演算を
加え認識用の特徴パラメータを算出する特徴抽出部16
と、特徴抽出部16から送られた特徴パラメータを基に
演算を行なって音声を認識し認識結果を出力する認識部
17と、認識部17からの認識結果を表示する表示部1
8と、各部の動作を制御する制御部19とを設け、ニュ
ーラルネットワークを用いた音声信号検出部15によっ
てノイズ信号と音声信号のニューラルネットにより音声
区間を検出する。
(57) [Summary] (Correction) [Purpose] A voice recognition device and a signal recognition device capable of accurately detecting a voice signal with a small amount of calculation. [Structure] The speech recognition apparatus 10 learns from a feature extraction unit 14 for detecting a voice signal, a noise signal and a voice signal that are classified into several types in advance, and features parameters sent from the feature extraction unit 14. A voice signal detection unit 15 for discriminating noise and a voice signal by a neural network calculation based on the above, and a feature parameter for all frames sent from the feature extraction unit 14 is subjected to a predetermined calculation to perform a feature parameter for recognition. Feature extraction unit 16 for calculating
And a recognition unit 17 that performs a calculation based on the characteristic parameter sent from the characteristic extraction unit 16 to recognize a voice and output a recognition result, and a display unit 1 that displays the recognition result from the recognition unit 17.
8 and a control unit 19 for controlling the operation of each unit, and a voice signal detection unit 15 using a neural network detects a voice section by a neural network of noise signals and voice signals.
Description
【0001】[0001]
【産業上の利用分野】本発明は、音声等を認識する音声
認識装置及び信号認識装置に係り、詳細には、音声信号
検出機能を備え、正確な音声信号の判定が可能な音声認
識装置及び信号認識装置に関する。BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a voice recognition device and a signal recognition device for recognizing a voice or the like, and more particularly to a voice recognition device having a voice signal detection function and capable of accurately determining a voice signal. The present invention relates to a signal recognition device.
【0002】[0002]
【従来の技術】音声認識装置は、利用する前に本人の音
声の登録を必要とする特定話者方式と誰でもが登録せず
に利用できる不特定話者方式がある。特定話者音声認識
では、音声認識を困難にしている個人差の要因が無視で
きるので、容易に高性能の音声認識装置を実現すること
ができる。2. Description of the Related Art A voice recognition device is classified into a specific speaker system that requires registration of a person's voice before use and an unspecified speaker system that anyone can use without registration. In the specific speaker voice recognition, the factors of individual differences that make voice recognition difficult can be ignored, so that a high-performance voice recognition device can be easily realized.
【0003】従来のこの種の音声認識装置として、単語
単位に区切って発声した音声を認識する離散単語認識装
置が知られている。この離散単語認識装置では、マイク
ロホンから入力された音声は、音声分析部でスペクトラ
ムパラメータの時系列に変換される。先ず、登録モード
では、標準パターンとして発声された音声の分析結果は
音声検出により単語単位に切り出され、カテゴリーコー
ドと共に単語毎に標準パターンメモリに記憶される。次
いで、認識モードでは入力音声の分析結果は同様に単語
単位に切り出され、照合部で各標準パターン(テンプレ
ート)とのマッチング処理が行なわれ、マッチング量
(距離)が計算される。判定部では、最も距離の小さい
標準パターンのカテゴリーコードが認識結果として表示
・出力される。As a conventional voice recognition device of this type, a discrete word recognition device for recognizing a voice uttered by dividing it into word units is known. In this discrete word recognition device, the voice input from the microphone is converted into a time series of spectrum parameters by the voice analysis unit. First, in the registration mode, the analysis result of the voice uttered as the standard pattern is cut out in units of words by voice detection, and is stored in the standard pattern memory for each word together with the category code. Next, in the recognition mode, the analysis result of the input voice is similarly cut out for each word, and the matching unit performs matching processing with each standard pattern (template) to calculate the matching amount (distance). The determination unit displays / outputs the category code of the standard pattern with the smallest distance as the recognition result.
【0004】音声分析としては、12〜16チャネルの
対数的に配置された帯域フィルタによるスペクトル分析
が用いられる場合が多い。分析結果は10〜20msの
フレームレートでサンプリングされ、フレーム毎に振幅
で正規化され、8ビット程度の精度で表現される。その
他の分析方法としては、ケプストラム分析、メルケプス
トラム分析、LPCケプストラム分析、LPC分析など
が用いられる。例えば、音声を分析してパラメータを抽
出する際、発声の長さの影響をなくすため図6に示すよ
うに、フレーム数を決めてフレームの重なりを調整して
分析を行なっている。As the voice analysis, a spectrum analysis by band-pass filters of 12 to 16 channels is often used. The analysis result is sampled at a frame rate of 10 to 20 ms, normalized by amplitude for each frame, and expressed with an accuracy of about 8 bits. Other analysis methods include cepstrum analysis, mel cepstrum analysis, LPC cepstrum analysis, and LPC analysis. For example, when a voice is analyzed and parameters are extracted, in order to eliminate the influence of the utterance length, as shown in FIG. 6, the number of frames is determined and the overlap of the frames is adjusted for analysis.
【0005】[0005]
【発明が解決しようとする課題】しかしながら、このよ
うな従来の音声認識装置にあっては、音声区間の検出は
音声レベルや電力レベルで行なう構成となっていたた
め、ノイズのレベルが大きかったり、音声のレベルが小
さかったりすると音声区間が検出できなかったり、検出
した位置が不正確なるという問題点があった。認識と同
じ分析を行なっているものでは、演算量が多すぎてハー
ドウェアが大きくなったり、または実時間で処理できな
いという問題点があった。However, in such a conventional voice recognition device, since the detection of the voice section is performed by the voice level or the power level, the noise level is high or the voice level is high. If the level is low, the voice section cannot be detected, or the detected position is inaccurate. With the same analysis as recognition, there are problems that the amount of calculation is too large and the hardware becomes large, or that it cannot be processed in real time.
【0006】そこで本発明は、演算量が少なくてかつ正
確な音声信号の検出ができる音声認識装置及び信号認識
装置を提供することを目的としている。Therefore, an object of the present invention is to provide a voice recognition device and a signal recognition device which can accurately detect a voice signal with a small amount of calculation.
【0007】[0007]
【課題を解決するための手段】請求項1記載の発明は、
上記目的達成のため、音声データを記憶する音声データ
記憶手段と、前記音声データ記憶手段に記憶された音声
データのノイズ信号と音声信号を検出する音声信号検出
手段と、前記音声信号検出手段により検出された音声信
号の音声区間をフレームに分割するフレーム分割手段
と、前記フレーム分割手段により分割されたフレームデ
ータから特徴パラメータを抽出する特徴パラメータ抽出
手段と、前記特徴パラメータ抽出手段により抽出された
特徴パラメータを用いて音声データを認識する認識手段
とを備えている。The invention according to claim 1 is
To achieve the above object, a voice data storage means for storing voice data, a voice signal detection means for detecting a noise signal and a voice signal of the voice data stored in the voice data storage means, and a detection by the voice signal detection means Frame dividing means for dividing the voice section of the audio signal thus obtained into frames, characteristic parameter extracting means for extracting characteristic parameters from the frame data divided by the frame dividing means, and characteristic parameters extracted by the characteristic parameter extracting means And a recognition means for recognizing voice data.
【0008】請求項2記載の発明は、音声データを記憶
する音声データ記憶手段と、前記音声データ記憶手段に
記憶された音声データから音声信号検出手段で使用する
特徴パラメータを抽出する第1の特徴パラメータ抽出手
段と、前記第1の特徴パラメータ抽出手段から出力され
た特徴パラメータに基づいてノイズ信号と音声信号を検
出する音声信号検出手段と、前記音声信号検出手段によ
り検出された音声信号の音声区間をフレームに分割する
フレーム分割手段と、前記フレーム分割手段により分割
されたフレームデータから音声認識用の特徴パラメータ
を抽出する第2の特徴パラメータ抽出手段と、前記第2
の特徴パラメータ抽出手段により抽出された特徴パラメ
ータを用いて音声データを認識する認識手段とを備えて
いる。According to a second aspect of the present invention, there is provided a voice data storage means for storing voice data, and a first feature for extracting a characteristic parameter used by the voice signal detection means from the voice data stored in the voice data storage means. Parameter extracting means, audio signal detecting means for detecting a noise signal and an audio signal based on the characteristic parameter output from the first characteristic parameter extracting means, and an audio section of the audio signal detected by the audio signal detecting means. And a second characteristic parameter extracting means for extracting a characteristic parameter for voice recognition from the frame data divided by the frame dividing means,
And recognition means for recognizing voice data using the characteristic parameters extracted by the characteristic parameter extracting means.
【0009】前記音声信号検出手段は、例えば請求項3
に記載されているように、ニューラルネットワークによ
り構成されるものであってもよく、前記音声信号検出手
段は、例えば請求項4に記載されているように、ニュー
ラルネットワークにより構成され、予め所定種類に分類
されたノイズ信号と音声信号によって学習がなされると
ともに、ニューラルネット計算によりノイズ信号と音声
信号を判別するものであってもよい。The voice signal detecting means may be, for example, the device of claim 3.
Alternatively, the audio signal detecting means may be configured by a neural network as described in claim 4, and may be of a predetermined type. Learning may be performed by the classified noise signal and voice signal, and the noise signal and the voice signal may be discriminated by neural network calculation.
【0010】請求項5記載の発明は、時系列に入力され
る信号をデータとして記憶するデータ記憶手段と、前記
データ記憶手段に記憶されたデータのノイズと信号を検
出する信号検出手段と、前記信号検出手段により検出さ
れた信号の区切り区間をフレームに分割するフレーム分
割手段と、前記フレーム分割手段により分割されたフレ
ームデータから特徴パラメータを抽出する特徴パラメー
タ抽出手段と、前記特徴パラメータ抽出手段により抽出
された特徴パラメータを用いて時系列データを認識する
認識手段とを備えている。According to a fifth aspect of the present invention, data storage means for storing signals inputted in time series as data, signal detection means for detecting noise and signals of the data stored in the data storage means, and Frame dividing means for dividing the delimiter section of the signal detected by the signal detecting means into frames, characteristic parameter extracting means for extracting characteristic parameters from the frame data divided by the frame dividing means, and extraction by the characteristic parameter extracting means And a recognition means for recognizing the time-series data by using the generated characteristic parameter.
【0011】前記音声信号検出手段は、例えば請求項6
に記載されているように、ニューラルネットワークによ
り構成されるものであっもよい。The audio signal detecting means may be, for example, the device of claim 6.
, It may be configured by a neural network.
【0012】[0012]
【作用】請求項1、2、3及び4記載の発明では、音声
取込手段により音声がデジタル化して取り込まれ、取り
込まれた音声データは音声データ記憶手段に記憶され
る。音声データ記憶手段に記憶された音声データは第1
の特徴パラメータ抽出手段によって音声信号検出用の特
徴パラメータが抽出され、この特徴パラメータを基にニ
ューラルネットワークを用いた音声信号検出手段によっ
てノイズ信号と音声信号が検出される。ノイズ信号と音
声信号が検出されるとフレーム分割手段により検出され
た音声信号の音声区間がフレームに分割される。この場
合、音声データを分析し語頭及び語尾を判定して音声区
間が決定される。According to the invention described in claims 1, 2, 3 and 4, the voice is digitized and captured by the voice capturing means, and the captured voice data is stored in the voice data storage means. The voice data stored in the voice data storage means is the first
The characteristic parameter extracting means extracts the characteristic parameter for voice signal detection, and the noise signal and the voice signal are detected by the voice signal detecting means using the neural network based on the characteristic parameter. When the noise signal and the voice signal are detected, the voice section of the voice signal detected by the frame dividing means is divided into frames. In this case, the voice section is determined by analyzing the voice data and determining the beginning and end of the word.
【0013】このようにして検出された音声区間を基に
フレーム分割手段により音声信号の音声区間がフレーム
に分割され、分割された所定フレームのフレームデータ
から第2の特徴パラメータ抽出手段により音声認識用の
特徴パラメータが抽出され、抽出された特徴パラメータ
を用いて認識手段が音声データを認識する。The voice division of the voice signal is divided into frames by the frame division means based on the voice division thus detected, and the second feature parameter extraction means for voice recognition is used from the divided frame data of a predetermined frame. Feature parameters are extracted, and the recognition unit recognizes the voice data using the extracted feature parameters.
【0014】従って、演算量が少なくてかつ正確な音声
信号の検出ができる。Therefore, it is possible to accurately detect a voice signal with a small amount of calculation.
【0015】請求項5及び6記載の発明では、時系列に
入力される信号がデータとして取り込まれ、データ記憶
手段に記憶される。データ記憶手段に記憶されたデータ
は信号検出手段によってノイズと信号が検出され、フレ
ーム分割手段によって検出された信号の区切り区間がフ
レームに分割される。According to the fifth and sixth aspects of the invention, the signals input in time series are fetched as data and stored in the data storage means. Noise and a signal are detected by the signal detection means in the data stored in the data storage means, and the delimiter section of the signal detected by the frame division means is divided into frames.
【0016】そして、分割された所定フレームのフレー
ムデータから特徴パラメータ抽出手段により特徴パラメ
ータが抽出され、抽出された特徴パラメータを用いて認
識手段が時系列に入力される信号のデータを認識する。Then, the characteristic parameter is extracted from the divided frame data of the predetermined frame by the characteristic parameter extracting means, and the recognizing means recognizes the data of the signal input in time series using the extracted characteristic parameter.
【0017】従って、時系列に入力される信号を正確に
認識することができる。Therefore, the signals input in time series can be accurately recognized.
【0018】[0018]
【実施例】以下、本発明を図面に基づいて説明する。DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described below with reference to the drawings.
【0019】図1〜図5は本発明に係る音声認識装置及
び信号認識装置の一実施例を示す図であり、本実施例は
発声した音声を認識する音声認識装置に適用した例であ
る。1 to 5 are diagrams showing an embodiment of a voice recognition device and a signal recognition device according to the present invention, and this embodiment is an example applied to a voice recognition device for recognizing a uttered voice.
【0020】先ず、構成を説明する。図1は、音声認識
装置10のブロック構成図であり、図中、白抜き矢印は
データの流れを示す。この図において、音声認識装置1
0は、音声信号や音楽信号が入力される入力部11と、
入力部11に入力される音声信号や音楽信号を所定のサ
ンプリング数でサンプリングして音情報パターンを出力
するサンプリング部12と、A/D変換された音声信号
を記憶するメモリ部13と、後述する音声信号検出部1
5に使用する特徴パラメータを、後述する特徴抽出部1
6(特徴抽出部2)の前段処理をも兼ねた簡易な処理に
より抽出する特徴抽出部14(特徴抽出部1)と、ニュ
ーラルネットワークにより構成され、予め何種類かに分
類されたノイズ信号と音声信号によって学習がなされる
とともに、特徴抽出部14(特徴抽出部1)から送られ
た特徴パラメータを基にニューラルネット計算によりノ
イズと音声信号を判別する音声信号検出部15と、特徴
抽出部14(特徴抽出部1)から全てのフレームの特徴
パラメータが送られてきたとき、これら送られてきた特
徴パラメータに所定の演算を加えて認識用の特徴パラメ
ータを算出する特徴抽出部16(特徴抽出部2)と、特
徴抽出部16(特徴抽出部2)から送られた特徴パラメ
ータを基に所定の演算を行なって音声を認識し認識結果
を出力する認識部17と、認識部17からの認識結果を
表示するLCD等からなる表示部18と、上記各部から
の信号を受け各部に制御信号を出力して各部の動作を制
御する制御部19とから構成されている。First, the structure will be described. FIG. 1 is a block diagram of the voice recognition apparatus 10, in which white arrows indicate the flow of data. In this figure, the voice recognition device 1
0 is an input unit 11 to which a voice signal or a music signal is input,
A sampling unit 12 that outputs a sound information pattern by sampling a voice signal or a music signal input to the input unit 11 at a predetermined sampling number, a memory unit 13 that stores the A / D converted voice signal, and will be described later. Audio signal detector 1
The feature parameter used for 5 is the feature extraction unit 1 described later.
6 (feature extraction unit 2), a feature extraction unit 14 (feature extraction unit 1) for extracting by a simple process that also functions as a pre-stage process, and a noise signal and voice that are configured by a neural network and are classified into several types in advance. The learning is performed by the signal, and the voice signal detection unit 15 that discriminates noise from a voice signal by neural network calculation based on the feature parameter sent from the feature extraction unit 14 (feature extraction unit 1), and the feature extraction unit 14 ( When the feature parameters of all the frames are sent from the feature extracting unit 1), the feature extracting unit 16 (feature extracting unit 2) that calculates the feature parameters for recognition by adding a predetermined calculation to the sent feature parameters ) And a feature extraction unit 16 (feature extraction unit 2), a recognition unit for performing a predetermined calculation based on the feature parameter to recognize a voice and output a recognition result. 7, a display unit 18 such as an LCD that displays the recognition result from the recognition unit 17, and a control unit 19 that receives signals from the above units and outputs a control signal to each unit to control the operation of each unit. ing.
【0021】上記入力部11は、音声を入力するマイク
21、入力された音声を増幅する増幅器22及び入力さ
れた音声信号のノイズを除去するローパスフィルタ23
からなり、入力される音声情報や音楽情報を所定のアナ
ログ信号に変換してサンプリング部12に出力する。The input section 11 includes a microphone 21 for inputting voice, an amplifier 22 for amplifying the input voice, and a low-pass filter 23 for removing noise of the input voice signal.
The input voice information and music information are converted into a predetermined analog signal and output to the sampling unit 12.
【0022】上記サンプリング部12は、A/Dコンバ
ータ等から構成され、所定のサンプリング間隔(サンプ
リング周波数)でサンプリングして標本化し、その標本
化データをメモリ部13に出力する。The sampling section 12 is composed of an A / D converter or the like, and samples and samples at a predetermined sampling interval (sampling frequency), and outputs the sampled data to the memory section 13.
【0023】上記メモリ部13は、リング・バッファ構
造のRAM(Random Access Memory)等から構成され、
制御部19からの制御信号に従ってデジタル信号に変換
された音声信号を常時更新しながら保存する。The memory unit 13 is composed of a ring buffer structure RAM (Random Access Memory) or the like,
The audio signal converted into the digital signal according to the control signal from the control unit 19 is constantly updated and stored.
【0024】上記特徴抽出部14(特徴抽出部1)は、
音声信号検出部15に使用する特徴パラメータを抽出す
るもので、特徴抽出部16(特徴抽出部2)の前段処理
をも兼ねた比較的簡単な処理を行なう。この特徴抽出部
14(特徴抽出部1)では、上記メモリ部13に蓄えら
れたデータ数が一定値(1フレーム分)に達したときデ
ータを取り込んで特徴パラメータを抽出する。The feature extraction unit 14 (feature extraction unit 1) is
A feature parameter used in the audio signal detection unit 15 is extracted, and a relatively simple process that doubles as a preceding stage process of the feature extraction unit 16 (feature extraction unit 2) is performed. The feature extracting unit 14 (feature extracting unit 1) takes in data when the number of data stored in the memory unit 13 reaches a constant value (one frame), and extracts a feature parameter.
【0025】上記音声信号検出部15は、図2に示すよ
うなニューラルネットワーク構造により構成されてお
り、予め何種類かに分類されたノイズ信号と音声信号に
よって学習がなされている。図2において、ニューラル
ネットワーク30は、一層の相互結合された連合層とも
呼ぶべきニューロン(○印参照)と、全ての連合層と結
合している入力シナプスから構成されるネットワーク構
造となっており、ニューラルネットワーク30を用いた
音声信号検出部15は、特徴抽出部14(特徴抽出部
1)から送られた特徴パラメータ(Pk+1〜Pl)を前回
送られた特徴パラメータ(P0〜Pk)に追加して特徴パ
ラメータ(P0〜Pl)とし、この特徴パラメータ(P0
〜Pl)を入力として所定の演算を行なって出力のうち
のO0〜Omが発火した時はノイズ信号と、またOm+1〜
Onが発火した時は音声信号と判定し判定結果を出力す
る構成となっている。The voice signal detecting section 15 has a neural network structure as shown in FIG. 2, and is trained by noise signals and voice signals classified into several types in advance. In FIG. 2, the neural network 30 has a network structure composed of one layer of neurons (see circles) that should also be called as an interconnected associative layer and input synapses connected to all the associative layers, The voice signal detection unit 15 using the neural network 30 adds the feature parameters (Pk + 1 to Pl) sent from the feature extraction unit 14 (feature extraction unit 1) to the previously sent feature parameters (P0 to Pk). The characteristic parameters (P0 to Pl), and the characteristic parameters (P0
~ Pl) is used as an input and a predetermined operation is performed to generate a noise signal when O0 to Om of the outputs are ignited, and Om + 1 to
When On fires, it is determined that it is a voice signal and outputs the determination result.
【0026】図3は上記ニューラルネットの入力ニュー
ロンに入力される前段処理を示す図であり、前段処理と
して2次元メルケプストラム変換を用いた例である。FIG. 3 is a diagram showing a pre-stage process input to the input neuron of the above-mentioned neural network, and is an example using a two-dimensional mel-cepstrum transform as the pre-stage process.
【0027】入力ニューロンに入力される前段処理とし
て1フレーム毎に、以下のような処理が行なわれる。す
なわち、1フレーム毎に音声信号の入力→フーリエ変換
→メル変換(対数変換)→逆フーリエ変換→ケプストラ
ムが行なわれ、上記前段処理によって得られたケプスト
ラムを32フレーム分毎にフーリエ変換して図2のニュ
ーラルネットの入力ニューロンに入力する。The following processing is performed for each frame as the pre-stage processing input to the input neuron. That is, the input of the voice signal → Fourier transform → Mel transform (logarithmic conversion) → inverse Fourier transform → cepstrum is performed for each frame, and the cepstrum obtained by the above-described preprocessing is Fourier transformed for every 32 frames to obtain the result shown in FIG. Input to the input neuron of the neural network.
【0028】次に、本実施例の動作を説明する。Next, the operation of this embodiment will be described.
【0029】ところで、このような音声認識装置の音声
認識手法としてニューラルネットワーク技術からのアプ
ローチも試みられている。By the way, as a voice recognition method for such a voice recognition device, an approach based on a neural network technique has been attempted.
【0030】ニューラルネットワークは、それまでの逐
次処理型コンピュータシステムでは手に負えない問題に
対して、脳の神経細胞をモデルとして複数のプロッセッ
サを相互に密に結合させて並列処理を行う並列処理型コ
ンピュータシステムにより解決しようとするものであ
り、特に、パターンマッピング、パターン完全化、パタ
ーン認識等のパターンに関連した問題において優れた能
力を発揮する。これらのパターン認識に関連した問題を
多く含む領域の例として、音声合成と音声認識があり、
音声認識手法としてニューラルネットワーク技術におい
て提案されている各種手法が適用可能である。ニューラ
ルネットワークの各種手法は、パターン分類器として分
類した場合、入力信号が2値パターンか連続値をとる
か、結合効率を教師つきで学習するか教師なしで学習す
るか、などによって分類できる。音声認識にニューラル
ネットワークの手法を適用する場合は、音声が時系列で
連続的に変化するため入力パターンに連続値をとる手法
のものを適用することになる。連続値をとるニューラル
ネットワークのうち教師ありのものとしては、パーセプ
トロン及び多層パーセプトロン等があり、教師なしのも
のとしては、自己組織化特徴写像がある。これらのニュ
ーラルネットワーク手法のうちベクトル量子化手法にお
けるボロノイ分割を実現するものが、自己組織化特徴写
像マップである。The neural network is a parallel processing type in which a plurality of processors are closely coupled to each other as a model by using nerve cells of the brain as a model to solve a problem that the conventional sequential processing type computer system cannot handle. It is intended to be solved by a computer system, and exerts an excellent ability particularly in pattern-related problems such as pattern mapping, pattern perfection, and pattern recognition. Speech synthesis and speech recognition are examples of areas that contain many problems related to pattern recognition.
As the speech recognition method, various methods proposed in the neural network technology can be applied. When classified as a pattern classifier, various methods of neural networks can be classified according to whether an input signal takes a binary pattern or a continuous value, whether coupling efficiency is learned with a teacher or without a teacher, and the like. When the neural network method is applied to the voice recognition, since the voice changes continuously in time series, the method of taking a continuous value in the input pattern is applied. Among the neural networks that take continuous values, there are perceptrons and multilayer perceptrons as supervised ones, and there is a self-organizing feature map as unsupervised ones. Among these neural network methods, the one that realizes Voronoi division in the vector quantization method is a self-organizing feature map.
【0031】自己組織化特徴写像マップは、ヘルシンキ
大学のKohonen(コホーネン)によって提案され
たニューラルネットワークのマッピング手法であり、教
師なしで学習を行うため、そのネットワーク構造は単純
なものであり、一層の相互結合された連合層とも呼ぶべ
きニューロンと、全ての連合層と結合している入力シナ
プスから構成される。自己組織化特徴写像マップでは、
入力信号はシナプス荷重により積和計算され、連合層の
どれか特定のニューロンを出力状態にする。出力状態に
なった特定のニューロン及びその近傍のニューロンは、
同様の入力信号に最も敏感に反応するようにシナプス荷
重が変更される。このとき、最初に最も高い出力状態に
なったニューロンは同一の入力信号に対して他のニュー
ロンと競い合って自己の出力を最大にするように見える
ため、この学習過程は協合学習とも呼ばれている。この
協合学習により、入力信号に対して特定のニューロンの
活性度が上がると、距離的に近い位置関係にあるニュー
ロンは引きずられて活性度が上がり、さらにその外側に
あるニューロンは逆に活性度が下がるという現象が発生
する。この現象により活性度が上がるニューロンをまと
めて「泡」と呼んでいる。すなわち、自己組織化特徴写
像マップにおける基本動作は、入力信号に従ってニュー
ロンの2次元平面上に、ある特定の「泡」が発生するこ
とであり、その「泡」は徐々に成長しながら、特定の入
力信号(刺激)に対して選択的に反応するようニューラ
ルネットワークのシナプス荷重が自動的に修正されてい
くものである。この動作の結果、最終的には特定の入力
信号と「泡」が対応づけられ、入力信号の自動的な分類
が行なわれる。The self-organizing feature map is a neural network mapping method proposed by Kohonen of the University of Helsinki. Since the learning is performed without supervision, the network structure is simple and further It is composed of neurons that should be called interconnected layers, and input synapses that are connected to all connected layers. In the self-organizing feature map,
The input signals are sum-of-products calculated by the synaptic weight, and any specific neuron in the association layer is put into the output state. The specific neuron that is in the output state and the neighboring neurons are
The synaptic weights are modified to be most sensitive to similar input signals. At this time, it seems that the neuron with the highest output state first competes with other neurons for the same input signal and maximizes its own output, so this learning process is also called cooperative learning. There is. With this joint learning, when the activity of a specific neuron increases with respect to the input signal, the neuron in a close distance relationship is dragged to increase the activity, and the neurons outside the neuron have the opposite activity. The phenomenon that the value goes down occurs. The neurons whose activity is increased by this phenomenon are collectively called "bubbles". That is, the basic operation in the self-organizing feature map is that a certain "bubble" is generated on the two-dimensional plane of the neuron according to the input signal, and the "bubble" gradually grows and becomes a specific bubble. The synaptic weight of the neural network is automatically corrected so as to selectively respond to the input signal (stimulus). As a result of this operation, a specific input signal is finally associated with a "bubble", and the input signal is automatically classified.
【0032】本実施例では、上記ニューラルネットワー
クを用いた音声信号検出部15を設けることにより、音
声区間の検出を単に音声レベルや電力レベルで行なうの
ではなく、予めノイズ信号と音声信号により学習された
ニューラルネットによって音声区間を検出しているの
で、正確な音声信号の検出ができるようになる。In the present embodiment, by providing the voice signal detecting section 15 using the above neural network, the voice section is not detected simply by the voice level or the power level, but is learned in advance by the noise signal and the voice signal. Since the voice section is detected by the neural network, it becomes possible to accurately detect the voice signal.
【0033】図4は音声認識装置10の音声認識処理を
示すフローチャートであり、同図中、符号Sn(n=
1,2,…)はフローの各ステップを示している。FIG. 4 is a flow chart showing the voice recognition processing of the voice recognition apparatus 10, and in the figure, reference numeral Sn (n = n =
1, 2, ...) Indicate each step of the flow.
【0034】先ず、ステップS1で音声データの1デー
タを取込み、取込んだ音声データをメモリ部13に書き
込み、ステップS2でメモリ部13に蓄えられたデータ
が1フレーム分になったか否かを判別する。次いで、ス
テップS3で特徴抽出部14(特徴抽出部1)により1
フレーム分のデータを取込んで特徴パラメータを抽出
し、ステップS4で音声信号検出部15により特徴抽出
部14(特徴抽出部1)から送られた特徴パラメータを
入力として音声信号検出用ニューラルネットを計算す
る。すなわち、音声信号検出部15では、特徴抽出部1
4(特徴抽出部1)から送られた特徴パラメータ(Pk+
1〜Pl)を前回送られた特徴パラメータ(P0〜Pk)に
追加して特徴パラメータ(P0〜Pl)とし、この特徴パ
ラメータ(P0〜Pl)を入力として所定の演算を行なっ
て出力のうちのO0〜Omが発火した時はノイズ信号と、
またOm+1〜Onが発火した時は音声信号と判定する音声
信号検出用ニューラルネット計算を行なう。First, in step S1, one piece of voice data is fetched, the fetched voice data is written in the memory section 13, and in step S2 it is judged whether or not the data stored in the memory section 13 is one frame. To do. Next, in step S3, the feature extraction unit 14 (feature extraction unit 1) sets 1
The data for the frame is fetched to extract the characteristic parameter, and in step S4, the speech signal detecting section 15 calculates the neural network for detecting the speech signal by using the characteristic parameter sent from the characteristic extracting section 14 (feature extracting section 1) as an input. To do. That is, in the voice signal detection unit 15, the feature extraction unit 1
4 (feature extraction unit 1) sends the feature parameter (Pk +
1 to Pl) is added to the previously sent characteristic parameter (P0 to Pk) to make it a characteristic parameter (P0 to Pl), and this characteristic parameter (P0 to Pl) is used as an input to perform a predetermined calculation and output When O0-Om ignites, a noise signal
Further, when Om + 1 to On are fired, a voice signal detecting neural network calculation for determining a voice signal is performed.
【0035】次いで、ステップS5で上記ニューラルネ
ット計算の結果、音声信号検出部15からの信号が音声
信号か否かを判別し、音声信号検出部15からの信号が
音声信号でないときはノイズ信号であると判断してステ
ップS6で前回の音声信号検出部15からの信号が音声
信号か否かを判別する。前回の音声信号検出部15から
の信号が音声信号でないときはノイズ信号であると判断
してステップS7でフレームの頭のアドレスを1/2フ
レーム分ずらしてステップS1に戻り1データの取込み
を行なう。すなわち、特徴抽出部14(特徴抽出部1)
はフレームの頭のアドレス(ポインタ)を1/2フレー
ム分ずらしてステップS1に戻って1フレーム分のデー
タが蓄えられるのを待ち、特徴パラメータを抽出する。
音声信号検出部15は、特徴抽出部14(特徴抽出部
1)からの特徴パラメータを更新しながらニューラルネ
ット計算を行なう。上記ステップS1〜ステップS7の
処理はステップS5またはステップS6で音声信号が検
出されるまで繰り返される。上記ステップS5で音声信
号が検出されるとステップS8で音声信号が初検出か否
かを検出し、音声信号が初検出のときはステップS9で
検出した音声信号の語頭アドレスを設定する。次いで、
ステップS10でメモリオーバーになったか否かを判別
し、メモリオーバーでなければ音声信号検出部15から
の音声信号が所定の長さに達していないと判断してメモ
リオーバーになるまでデータをメモリに書き込むためス
テップS7に移行して上記ステップS1〜ステップS7
の処理を繰り返す。Next, in step S5, as a result of the neural network calculation, it is determined whether or not the signal from the voice signal detecting section 15 is a voice signal. If the signal from the voice signal detecting section 15 is not a voice signal, a noise signal is used. If it is determined that the signal is present, it is determined in step S6 whether or not the previous signal from the audio signal detection unit 15 is an audio signal. If the previous signal from the voice signal detecting unit 15 is not a voice signal, it is determined that it is a noise signal, and in step S7 the address of the head of the frame is shifted by 1/2 frame and the process returns to step S1 to fetch one data. . That is, the feature extraction unit 14 (feature extraction unit 1)
Shifts the address (pointer) at the head of the frame by 1/2 frame, returns to step S1, waits for the data for one frame to be stored, and extracts the characteristic parameter.
The voice signal detection unit 15 performs the neural network calculation while updating the feature parameter from the feature extraction unit 14 (feature extraction unit 1). The processes of steps S1 to S7 are repeated until a voice signal is detected in step S5 or step S6. When the voice signal is detected in step S5, whether or not the voice signal is initially detected is detected in step S8. When the voice signal is initially detected, the word address of the voice signal detected in step S9 is set. Then
In step S10, it is determined whether the memory is over, and if the memory is not over, it is determined that the audio signal from the audio signal detector 15 has not reached a predetermined length, and the data is stored in the memory until the memory is over. In order to write, the process proceeds to step S7 and the above steps S1 to S7.
The process of is repeated.
【0036】上記ステップS10でメモリオーバーにな
ったときあるいは上記ステップS6で前回の音声信号が
検出されたときは音声信号検出部15からの音声信号が
所定の長さに達したと判断してステップS11で語尾ア
ドレスを設定し、ステップS12で語頭及び語尾の判定
結果を基に語長(音声区間)を計算して音声区間のフレ
ームを分割し、ステップS13で特徴抽出部16(特徴
抽出部2)により1フレームの特徴パラメータを抽出す
る。次いで、ステップS14で全てのフレームについて
特徴パラメータの抽出が終了したか否かを判別し、全フ
レームの特徴パラメータの抽出が終了していないときは
ステップS13に戻って全てのフレームについて特徴パ
ラメータの抽出が終了するまで特徴パラメータの抽出を
繰り返す。全フレームの特徴パラメータの抽出が終了す
るとステップS15で認識部17により特徴抽出部16
(特徴抽出部2)から送られた全フレームの特徴パラメ
ータを入力として認識用ニューラルネットを計算し、ス
テップS16で認識結果を表示部18に表示して本フロ
ーの処理を終える。When the memory is over in step S10 or when the previous voice signal is detected in step S6, it is determined that the voice signal from the voice signal detecting section 15 has reached a predetermined length, and step In step S11, the ending address is set, in step S12 the word length (speech section) is calculated based on the result of the beginning and ending judgment, and the frame of the speech section is divided. In step S13, the feature extraction unit 16 (feature extraction unit 2 ), The characteristic parameter of one frame is extracted. Next, in step S14, it is determined whether or not the extraction of the characteristic parameters has been completed for all the frames. If the extraction of the characteristic parameters of all the frames has not been completed, the process returns to step S13 to extract the characteristic parameters of all the frames. The extraction of the characteristic parameter is repeated until. When the extraction of the feature parameters of all frames is completed, the recognition unit 17 causes the feature extraction unit 16 to perform the extraction in step S15.
The feature neural network for recognition is calculated using the feature parameters of all the frames sent from the (feature extraction unit 2), and the recognition result is displayed on the display unit 18 in step S16, and the processing of this flow ends.
【0037】以上の処理を実行することにより具体的に
は音声認識装置10の各部で以下のような音声認識動作
が行われる。By executing the above processing, specifically, the following voice recognition operation is performed in each unit of the voice recognition apparatus 10.
【0038】 マイク21から入力された音声は、増
幅器22で増幅された後、ローパスフィルタ23で入力
された音声信号のノイズが除去され、A/Dコンバータ
12でデジタル信号に変換され、制御部19からの制御
信号に従ってリング・バッファ構造のメモリ部13に常
時更新されながら保存される。After the voice input from the microphone 21 is amplified by the amplifier 22, the noise of the voice signal input by the low-pass filter 23 is removed, converted into a digital signal by the A / D converter 12, and the control unit 19 In accordance with the control signal from the above, it is constantly updated and stored in the memory unit 13 of the ring buffer structure.
【0039】 特徴抽出部14(特徴抽出部1)は、
音声信号検出部15に使用する特徴パラメータを抽出す
るもので、特徴抽出部16(特徴抽出部2)の前段処理
をも兼ねた比較的簡単な処理を行なう。この特徴抽出部
14(特徴抽出部1)では、メモリ部13に蓄えられた
データ数が一定値(1フレーム分)に達したときデータ
を取り込んで特徴パラメータを抽出して音声信号検出部
15に送る。The feature extraction unit 14 (feature extraction unit 1)
A feature parameter used in the audio signal detection unit 15 is extracted, and a relatively simple process that doubles as a preceding stage process of the feature extraction unit 16 (feature extraction unit 2) is performed. In the feature extraction unit 14 (feature extraction unit 1), when the number of data stored in the memory unit 13 reaches a fixed value (for one frame), the data is taken in to extract the feature parameter and the voice signal detection unit 15 is extracted. send.
【0040】 音声信号検出部15は、図2に示すよ
うなニューラルネットワークで構成され、予め何種類か
に分類されたノイズ信号と音声信号によって学習がなさ
れている。ニューラルネットワーク30を用いた音声信
号検出部15は、特徴抽出部14(特徴抽出部1)から
送られた特徴パラメータ(Pk+1〜Pl)を前回送られた
特徴パラメータ(P0〜Pk)に追加して特徴パラメータ
(P0〜Pl)とし、この特徴パラメータ(P0〜Pl)を
入力として所定の演算を行なって出力のうちのO0〜Om
が発火した時はノイズ信号と、またOm+1〜Onが発火し
た時は音声信号と判定し制御部19に判定結果を出力す
る。The voice signal detection unit 15 is composed of a neural network as shown in FIG. 2, and is trained by noise signals and voice signals that are classified into several types in advance. The voice signal detection unit 15 using the neural network 30 adds the feature parameters (Pk + 1 to Pl) sent from the feature extraction unit 14 (feature extraction unit 1) to the previously sent feature parameters (P0 to Pk). Then, the characteristic parameters (P0 to Pl) are set, and a predetermined calculation is performed using the characteristic parameters (P0 to Pl) as an input to output O0 to Om of the outputs.
When it is fired, it is determined to be a noise signal, and when Om + 1 to On are fired, it is determined to be a voice signal and the determination result is output to the control unit 19.
【0041】 音声信号検出部15からの信号がノイ
ズ信号と判定したときは、制御部19は特徴抽出部14
(特徴抽出部1)に音声信号状態を続けるように指示を
出す。When the signal from the audio signal detection unit 15 is determined to be a noise signal, the control unit 19 causes the feature extraction unit 14
The (feature extraction unit 1) is instructed to continue the audio signal state.
【0042】 すると、特徴抽出部14(特徴抽出部
1)は、データのポイントを1/2フレームずらしてま
た1フレーム分のデータが蓄えられるのを待ち、特徴パ
ラメータを抽出して音声信号検出部15に送る。音声信
号検出部15は、特徴パラメータを更新しながら演算を
行なう。Then, the feature extraction unit 14 (feature extraction unit 1) shifts the data points by ½ frame and waits for one frame of data to be stored again, and extracts the feature parameter to extract the voice signal detection unit. Send to 15. The voice signal detection unit 15 performs calculation while updating the characteristic parameter.
【0043】上記〜の動作を、音声信号検出部15
が音声信号を検出するまで繰返し行なう。The above-described operations (1) to (5) are performed by the audio signal detecting section 15
Repeat until the voice signal is detected by.
【0044】 音声信号検出部15が音声信号を検出
すると、制御部19は音声信号検出部15から音声信号
検出信号のきている間若しくはメモリ分の所定時間メモ
リにデータを書き込んだ後、メモリへのデータの書き込
みを禁止する。When the audio signal detection unit 15 detects an audio signal, the control unit 19 writes the data to the memory while the audio signal detection signal is being received from the audio signal detection unit 15 or after writing data in the memory for a predetermined time. Prohibit writing of data.
【0045】 制御部19は、音声区間のフレーム分
割を行なった後に、各フレームのスタート・ポイントを
特徴抽出部14(特徴抽出部1)に送って、改めてフレ
ーム毎の特徴抽出を行なわせその結果を特徴抽出部16
(特徴抽出部2)に取り込ませる。After performing the frame division of the voice section, the control unit 19 sends the start point of each frame to the feature extraction unit 14 (feature extraction unit 1) so that the feature extraction is performed again for each frame. The feature extraction unit 16
(Feature extraction unit 2).
【0046】 特徴抽出部14(特徴抽出部1)から
全てのフレームのパラメータが送られてきたら特徴抽出
部16(特徴抽出部2)はこれらに演算を加えて認識用
のパラメータを算出して認識部17に送る。When the parameters of all the frames are sent from the feature extraction unit 14 (feature extraction unit 1), the feature extraction unit 16 (feature extraction unit 2) performs calculation on these to calculate the recognition parameters and recognizes them. Send to section 17.
【0047】 認識部17は、特徴抽出部16(特徴
抽出部2)から送られてきた特徴パラメータを基に演算
を行なって認識結果を表示部18に送り、表示部18
は、制御部19の指示に従って認識部17からの音声認
識結果を表示して一連の音声認識動作が終了する。The recognition unit 17 performs calculation based on the characteristic parameters sent from the characteristic extraction unit 16 (feature extraction unit 2) and sends the recognition result to the display unit 18, and the display unit 18
Displays the voice recognition result from the recognition unit 17 according to the instruction of the control unit 19 and ends the series of voice recognition operations.
【0048】制御部19は、各部に制御信号を出力して
音声認識装置10の動作状態を上記の状態に戻して次
の認識を開始する。The control unit 19 outputs a control signal to each unit to return the operation state of the voice recognition apparatus 10 to the above state and start the next recognition.
【0049】以上説明したように、本実施例の音声認識
装置10は、音声信号検出用の特徴パラメータを抽出す
る特徴抽出部14(特徴抽出部1)と、ニューラルネッ
トワークにより構成され、予め何種類かに分類されたノ
イズ信号と音声信号によって学習がなされるとともに、
特徴抽出部14(特徴抽出部1)から送られた特徴パラ
メータを基にニューラルネット計算によりノイズと音声
信号を判別する音声信号検出部15と、特徴抽出部14
(特徴抽出部1)から全てのフレームの特徴パラメータ
が送られてきたとき、これら送られてきた特徴パラメー
タに所定の演算を加えて認識用の特徴パラメータを算出
する特徴抽出部16(特徴抽出部2)と、特徴抽出部1
6(特徴抽出部2)から送られた特徴パラメータを基に
所定の演算を行なって音声を認識し認識結果を出力する
認識部17と、認識部17からの認識結果を表示する表
示部18と、制御信号を出力して各部の動作を制御する
制御部19とを設け、ニューラルネットワークを用いた
音声信号検出部15によって予めノイズ信号と音声信号
により学習されたニューラルネットにより音声区間を検
出しているので、実時間で途切れることなく正確な音声
信号の検出が行なえ、ノイズや変化に強い音声認識装置
が実現できる。すなわち、従来の音声認識装置では、音
声区間の検出を単に音声レベルや電力レベルで行なって
いたため、音声のレベルが小さかったり、ノイズのレベ
ルが大きかったりしたたときに音声区間が検出できな
い、あるいは検出した位置が不正確なるという問題点が
あったが、本実施例の音声認識装置では、ニューラルネ
ットワークを用いた音声信号検出部15を設けることに
より、ノイズを単に信号に対するノイズ(S/N比)と
して捉えるのではなく、ノイズ成分も信号に対するノイ
ズ信号としてニューラルネットの入力として入力し、ノ
イズ信号と音声信号のニューラルネット計算により音声
区間を検出しているので、信号レベルの大小の影響を受
けない正確な音声信号の検出ができるようになる。As described above, the voice recognition device 10 of the present embodiment is composed of the feature extraction unit 14 (feature extraction unit 1) for extracting the feature parameters for voice signal detection and the neural network, and has various types in advance. Learning is performed by the noise signal and the voice signal classified into
A voice signal detection unit 15 that discriminates noise from a voice signal by neural network calculation based on the feature parameters sent from the feature extraction unit 14 (feature extraction unit 1), and the feature extraction unit 14
When the feature parameters of all the frames are sent from the (feature extracting unit 1), the feature extracting unit 16 (feature extracting unit) that calculates a feature parameter for recognition by performing a predetermined calculation on the sent feature parameters 2) and the feature extraction unit 1
6 (feature extracting unit 2), a recognition unit 17 that performs a predetermined calculation based on the feature parameters sent from the feature recognition unit to recognize a voice and output a recognition result, and a display unit 18 that displays the recognition result from the recognition unit 17. A control section 19 for outputting a control signal to control the operation of each section is provided, and a voice signal detecting section 15 using a neural network detects a voice section by a neural network previously learned from a noise signal and a voice signal. As a result, a voice signal can be accurately detected without interruption in real time, and a voice recognition device that is resistant to noise and changes can be realized. That is, in the conventional voice recognition device, since the detection of the voice section is simply performed by the voice level or the power level, the voice section cannot be detected or detected when the voice level is low or the noise level is high. However, the voice recognition apparatus of this embodiment is provided with the voice signal detection unit 15 using a neural network, so that the noise is simply noise to the signal (S / N ratio). The noise component is also input as a noise signal to the signal as the input of the neural network, and the voice section is detected by the neural network calculation of the noise signal and the voice signal, so it is not affected by the magnitude of the signal level. It becomes possible to detect an audio signal accurately.
【0050】また、音声信号検出部15で使用する音声
信号検出用の特徴パラメータを演算する特徴抽出部14
(特徴抽出部1)は、本来認識に用いる特徴パラメータ
抽出の一部であるので、この特徴抽出部14のため新た
にハードやソフトを追加することなく実現できる。Further, the feature extraction unit 14 for calculating the feature parameter for voice signal detection used by the voice signal detection unit 15
Since the (feature extraction unit 1) is a part of the feature parameter extraction originally used for the recognition, the feature extraction unit 14 can be realized without newly adding hardware or software.
【0051】また、本実施例では、認識部17にニュー
ラルネットワークを用いているので実時間で正確な音声
認識結果を得ることが可能になる。Further, in this embodiment, since the recognition unit 17 uses the neural network, it is possible to obtain the accurate voice recognition result in real time.
【0052】なお、本実施例では、発声した音声を認識
する音声認識装置に適用した例であるが、勿論これには
限定されず、音声データのノイズ信号と音声信号を検出
する音声信号検出部を備えた装置であれば全ての装置に
適用可能であることは言うまでもない。例えば、時系列
に入力される信号(音声信号もこれに含まれるが、音声
信号に限定されない)に対して同様の認識処理を行なう
装置でもよい。Although the present embodiment is an example applied to a voice recognition device for recognizing a uttered voice, of course, the present invention is not limited to this, and a voice signal detecting section for detecting a noise signal and a voice signal of voice data. It goes without saying that any device provided with can be applied to all devices. For example, a device that performs similar recognition processing on a signal input in time series (including a voice signal, but not limited to a voice signal) may be used.
【0053】また、本実施例では、特徴抽出部14(特
徴抽出部1)を設け、特徴抽出部16(特徴抽出部2)
の前段処理をも兼ねた簡易な処理により音声信号検出用
の特徴パラメータを抽出するようにしているが、ノイズ
信号と音声信号を判別する音声信号検出部にデータを供
給できるものであればこの特徴抽出部14はどのような
構成のものであってもよく、かかる特徴抽出部を用いな
い構成であってもよい。Further, in this embodiment, the feature extracting unit 14 (feature extracting unit 1) is provided, and the feature extracting unit 16 (feature extracting unit 2) is provided.
The characteristic parameters for audio signal detection are extracted by a simple process that doubles as the preceding stage processing, but if the data can be supplied to the audio signal detection unit that distinguishes between a noise signal and an audio signal, this characteristic The extraction unit 14 may have any configuration, and may not have such a feature extraction unit.
【0054】また、本実施例では、音声信号検出部15
及び認識部17にニューラルネットワークを用いること
によって正確な音声認識結果を得ることが可能にしてい
るが、このニューラルネットワークの態様はどのような
ものでもよく、また、認識部17にニューラルネットワ
ークを用いないものであってもよいことは言うまでもな
い。Further, in the present embodiment, the voice signal detecting section 15
By using a neural network for the recognition unit 17 and the recognition unit 17, it is possible to obtain an accurate speech recognition result. However, any form of this neural network may be used, and the recognition unit 17 does not use a neural network. It goes without saying that it may be one.
【0055】さらに、上記音声認識装置を構成する回路
や部材の数、種類などは前述した実施例に限られないこ
とは言うまでもなく、ソフトウェアにより実現するよう
にしてもよい。Further, it is needless to say that the number and types of the circuits and members constituting the above speech recognition apparatus are not limited to those in the above-mentioned embodiment, but may be realized by software.
【0056】[0056]
【発明の効果】請求項1、2、3、4、5及び6記載の
発明によれば、音声データのノイズ信号と音声信号を検
出する音声信号検出手段を設けているので、演算量が少
なくてかつ正確な音声信号の検出ができ、ノイズや変化
に強い音声認識装置が実現できる。According to the first, second, third, fourth, fifth and sixth aspects of the present invention, since the voice signal detecting means for detecting the noise signal and the voice signal of the voice data is provided, the calculation amount is small. In addition, it is possible to realize a voice recognition device that can detect a voice signal accurately and accurately and that is resistant to noise and changes.
【0057】請求項7、8及び9記載の発明によれば、
時系列に入力される信号をデータとして取り込み、この
データのノイズ信号と音声信号を検出しているので、時
系列に入力される信号に対し演算量が少なくかつ正確な
時系列信号の検出ができ、ノイズや変化に強い信号認識
装置が実現できる。According to the invention described in claims 7, 8 and 9,
Since the signal input in time series is captured as data and the noise signal and voice signal of this data are detected, the amount of calculation is small and the time series signal can be detected accurately with respect to the signal input in time series. A signal recognition device that is resistant to noise and changes can be realized.
【図1】音声認識装置の全体構成図である。FIG. 1 is an overall configuration diagram of a voice recognition device.
【図2】音声認識装置の音声信号検出部のニューラルネ
ットワーク構造を示す図である。FIG. 2 is a diagram showing a neural network structure of a voice signal detection unit of the voice recognition device.
【図3】音声認識装置のニューラルネットワークの前段
処理を示す図である。FIG. 3 is a diagram showing a pre-stage process of a neural network of the voice recognition device.
【図4】音声認識装置のの音声認識処理を示すフローチ
ャートである。FIG. 4 is a flowchart showing a voice recognition process of the voice recognition device.
【図5】音声認識装置のフレームの重なりを示す図であ
る。FIG. 5 is a diagram showing overlapping of frames of the voice recognition device.
10 音声認識装置 11 入力部 12 サンプリング部(A/Dコンバータ) 13 メモリ部 14 特徴抽出部(特徴抽出部1) 15 音声信号検出部 16 特徴抽出部(特徴抽出部2) 17 認識部 18 表示部 19 制御部 21 マイク 22 増幅器 23 ローパスフィルタ 30 ニューラルネットワーク 10 voice recognition device 11 input unit 12 sampling unit (A / D converter) 13 memory unit 14 feature extraction unit (feature extraction unit 1) 15 voice signal detection unit 16 feature extraction unit (feature extraction unit 2) 17 recognition unit 18 display unit 19 control unit 21 microphone 22 amplifier 23 low-pass filter 30 neural network
Claims (6)
段と、 前記音声データ記憶手段に記憶された音声データのノイ
ズ信号と音声信号を検出する音声信号検出手段と、 前記音声信号検出手段により検出された音声信号の音声
区間をフレームに分割するフレーム分割手段と、 前記フレーム分割手段により分割されたフレームデータ
から特徴パラメータを抽出する特徴パラメータ抽出手段
と、 前記特徴パラメータ抽出手段により抽出された特徴パラ
メータを用いて音声データを認識する認識手段と、 を具備したことを特徴とする音声認識装置。1. A voice data storage unit for storing voice data; a voice signal detection unit for detecting a noise signal and a voice signal of the voice data stored in the voice data storage unit; and a voice signal detection unit for detecting the noise signal and the voice signal. A frame dividing unit that divides the voice section of the audio signal into frames; a characteristic parameter extracting unit that extracts a characteristic parameter from the frame data divided by the frame dividing unit; and a characteristic parameter that is extracted by the characteristic parameter extracting unit. A voice recognition device comprising: a recognition unit that recognizes voice data by using the recognition unit.
段と、 前記音声データ記憶手段に記憶された音声データから音
声信号検出手段で使用する特徴パラメータを抽出する第
1の特徴パラメータ抽出手段と、 前記第1の特徴パラメータ抽出手段から出力された特徴
パラメータに基づいてノイズ信号と音声信号を検出する
音声信号検出手段と、 前記音声信号検出手段により検出された音声信号の音声
区間をフレームに分割するフレーム分割手段と、 前記フレーム分割手段により分割されたフレームデータ
から音声認識用の特徴パラメータを抽出する第2の特徴
パラメータ抽出手段と、 前記第2の特徴パラメータ抽出手段により抽出された特
徴パラメータを用いて音声データを認識する認識手段
と、 を具備したことを特徴とする音声認識装置。2. A voice data storage unit for storing voice data, a first feature parameter extraction unit for extracting a feature parameter used by a voice signal detection unit from the voice data stored in the voice data storage unit, A voice signal detection unit that detects a noise signal and a voice signal based on the feature parameter output from the first feature parameter extraction unit, and a frame that divides the voice section of the voice signal detected by the voice signal detection unit into frames. By using the dividing means, the second characteristic parameter extracting means for extracting the characteristic parameter for voice recognition from the frame data divided by the frame dividing means, and the characteristic parameter extracted by the second characteristic parameter extracting means. A voice recognition device comprising: a recognition unit for recognizing voice data.
ットワークにより構成されていることを特徴とする請求
項1又は請求項2の何れかに記載の音声認識装置。3. The voice recognition device according to claim 1, wherein the voice signal detection means is configured by a neural network.
ットワークにより構成され、予め所定種類に分類された
ノイズ信号と音声信号によって学習がなされるととも
に、ニューラルネット計算によりノイズ信号と音声信号
を判別するようにしたことを特徴とする請求項1又は請
求項2の何れかに記載の音声認識装置。4. The voice signal detecting means is composed of a neural network, and learning is performed by a noise signal and a voice signal which are classified into a predetermined type in advance, and the noise signal and the voice signal are discriminated by a neural network calculation. The voice recognition device according to claim 1 or 2, wherein
記憶するデータ記憶手段と、 前記データ記憶手段に記憶されたデータのノイズと信号
を検出する信号検出手段と、 前記信号検出手段により検出された信号の区切り区間を
フレームに分割するフレーム分割手段と、 前記フレーム分割手段により分割されたフレームデータ
から特徴パラメータを抽出する特徴パラメータ抽出手段
と、 前記特徴パラメータ抽出手段により抽出された特徴パラ
メータを用いて時系列データを認識する認識手段と、 を具備したことを特徴とする信号認識装置。5. A data storage means for storing signals inputted in time series as data, a signal detection means for detecting noise and a signal of the data stored in the data storage means, and a signal detection means for detecting the noise and the signal. A frame dividing unit that divides a signal delimiter section into frames, a characteristic parameter extracting unit that extracts a characteristic parameter from the frame data divided by the frame dividing unit, and a characteristic parameter that is extracted by the characteristic parameter extracting unit are used. A signal recognition device comprising: a recognition unit that recognizes time series data.
ットワークにより構成されていることを特徴とする請求
項5に記載の信号認識装置。6. The signal recognition device according to claim 5, wherein the voice signal detection means is configured by a neural network.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP5107709A JPH06295196A (en) | 1993-04-08 | 1993-04-08 | Speech recognition device and signal recognition device |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP5107709A JPH06295196A (en) | 1993-04-08 | 1993-04-08 | Speech recognition device and signal recognition device |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH06295196A true JPH06295196A (en) | 1994-10-21 |
Family
ID=14465964
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP5107709A Pending JPH06295196A (en) | 1993-04-08 | 1993-04-08 | Speech recognition device and signal recognition device |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH06295196A (en) |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109523994A (en) * | 2018-11-13 | 2019-03-26 | 四川大学 | A kind of multitask method of speech classification based on capsule neural network |
| CN115641852A (en) * | 2022-10-18 | 2023-01-24 | 中国电信股份有限公司 | Voiceprint recognition method and device, electronic equipment and computer readable storage medium |
| JP2024532786A (en) * | 2021-08-12 | 2024-09-10 | ドルビー ラボラトリーズ ライセンシング コーポレイション | Reverberation and noise robust voice activity detection based on modulation domain attention |
-
1993
- 1993-04-08 JP JP5107709A patent/JPH06295196A/en active Pending
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109523994A (en) * | 2018-11-13 | 2019-03-26 | 四川大学 | A kind of multitask method of speech classification based on capsule neural network |
| JP2024532786A (en) * | 2021-08-12 | 2024-09-10 | ドルビー ラボラトリーズ ライセンシング コーポレイション | Reverberation and noise robust voice activity detection based on modulation domain attention |
| CN115641852A (en) * | 2022-10-18 | 2023-01-24 | 中国电信股份有限公司 | Voiceprint recognition method and device, electronic equipment and computer readable storage medium |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US10679612B2 (en) | Speech recognizing method and apparatus | |
| JP3168779B2 (en) | Speech recognition device and method | |
| CN109785857A (en) | Abnormal sound event recognition method based on MFCC+MP fusion feature | |
| US20240374219A1 (en) | Body action detection, identification and/or characterization using a machine learning model | |
| CN110600014A (en) | Model training method and device, storage medium and electronic equipment | |
| Chinmayi et al. | Emotion classification using deep learning | |
| EP3847646B1 (en) | An audio processing apparatus and method for audio scene classification | |
| KR20210098083A (en) | Method and Apparatus for Determining Stress in Speech Signal Using Weight | |
| CN113724692A (en) | Voice print feature-based phone scene audio acquisition and anti-interference processing method | |
| CN111785302A (en) | Speaker separation method, device and electronic device | |
| JPH06295196A (en) | Speech recognition device and signal recognition device | |
| CN120356488A (en) | Voice emotion recognition method, device, equipment and readable storage medium | |
| EP1489597B1 (en) | Vowel recognition device | |
| CN115376494B (en) | Voice detection method, device, equipment and medium | |
| CN112712823A (en) | Detection method, device and equipment of trailing sound and storage medium | |
| Estrebou et al. | Voice recognition based on probabilistic SOM | |
| CN113671031B (en) | Wall hollowing detection method and device | |
| JPH06295197A (en) | Speech recognition device and signal recognition device | |
| KR100202424B1 (en) | Real time speech recognition method | |
| JPH0466999A (en) | Device for detecting clause boundary | |
| Pagidirayi et al. | Speech Feature Extraction and Emotion Recognition using Deep Learning Techniques | |
| JPH0442299A (en) | Sound block detector | |
| CN119339711B (en) | Voice recognition awakening method, equipment and system for preventing recording detection | |
| Radha et al. | Raw-Waveform Based Bark Scale Initialized SincNet Model in Child Speaker Identification | |
| Gupta et al. | Application of Multilayer Perceptron in Speech Emotion Recognition Naman Gupta and Nikunj Agarwal |