JPH03230200A - Voice recognizing method - Google Patents
Voice recognizing methodInfo
- Publication number
- JPH03230200A JPH03230200A JP2026670A JP2667090A JPH03230200A JP H03230200 A JPH03230200 A JP H03230200A JP 2026670 A JP2026670 A JP 2026670A JP 2667090 A JP2667090 A JP 2667090A JP H03230200 A JPH03230200 A JP H03230200A
- Authority
- JP
- Japan
- Prior art keywords
- neural network
- input
- band
- speech
- block
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Landscapes
- Lock And Its Accessories (AREA)
Abstract
(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.
Description
【発明の詳細な説明】
[産業上の利用分野]
本発明は、電気錠、ICカード等のオンライン端末等で
入力音声からその単語を認識するに好適な音声認識方法
に関する。DETAILED DESCRIPTION OF THE INVENTION [Field of Industrial Application] The present invention relates to a speech recognition method suitable for recognizing words from input speech using online terminals such as electric locks and IC cards.
[従来の技術]
本出願人は、容易に実時間処理できる音声認識方法とし
て、特願平1−98376号により、ニューラルネット
ワークを用いて入力音声からその単語を認識するものを
提案している。この音声認識方法にあっては、ニューラ
ルネットワークへの入力として、入力音声の周波数特性
を算出し、各帯域のそれぞれにおいて時間的に等分割し
た音声区間のそれぞれを1つのブロックとして、各ブロ
ックの中で周波数特性の平均を算出し、それらの平均を
単語のパワー全体で正規化したものを用いることとして
いる。[Prior Art] The present applicant has proposed, in Japanese Patent Application No. 1-98376, a method for recognizing words from input speech using a neural network as a speech recognition method that can be easily processed in real time. In this speech recognition method, the frequency characteristics of the input speech are calculated as input to the neural network, and each speech interval divided equally in time in each band is treated as one block. The average of the frequency characteristics is calculated, and the average is normalized by the power of the word as a whole.
[発明が解決しようとする課題]
然しながら、上述の従来技術による場合には、ニューラ
ルネットワークを構築するために標準入カバターン(学
習入カバターン)を作製する時と、構築されたニューラ
ルネットワークを使用して音声認識するために評価入カ
バターンを作製する時との間で、定常雑音の混入や回線
等の入力系の相違等によってそれらの作製条件が異なる
と、認識率の低下か見られることとなる。[Problem to be solved by the invention] However, in the case of the above-mentioned conventional technology, it is difficult to create a standard input cover pattern (learning input cover pattern) in order to construct a neural network, and to use the constructed neural network. If the conditions for creating evaluation cover patterns for speech recognition differ due to the introduction of stationary noise or differences in input systems such as lines, the recognition rate will drop.
この認識率の低下は、以下に解析する如く、単語のパワ
ー全体で正規化するために、スペクトル歪を消去できな
いことによる。即ち、iをブロック番号、kを帯域番号
、Akをに帯域の周波数伝送特性、S mikを学習段
階でのに帯域iブロックの音声信号、S tikを評価
段階で電話回線を通した後における如く、定常的な周波
数伝送特性Akの影響によりスペクトルか歪んだ、k帯
域iブロックの音声信号とする時、
5tik=Ak−3■i
である、そして、評価段階での各音声信号S tikを
単語のパワー全体で正規化したものは、S tik
A k S *ikΣΣS tik Σ
ΣA k S aikてあって、右辺の周波数伝送特性
Akを消去できない、即ち、スペクトル歪を消去てきな
いのである。This reduction in recognition rate is due to the inability to eliminate spectral distortion due to normalization across word powers, as analyzed below. That is, i is the block number, k is the band number, Ak is the frequency transmission characteristic of the band, Smik is the audio signal of the band i block in the learning stage, and Stik is the sound signal after passing through the telephone line in the evaluation stage. , when the speech signal of k-band i-block is spectrum-distorted due to the influence of the stationary frequency transmission characteristic Ak, 5tik=Ak-3■i, and each speech signal S tik at the evaluation stage is expressed as a word. Normalized across the power of S tik
A k S *ikΣΣS tik Σ
Since ΣA k S aik, the frequency transmission characteristic Ak on the right side cannot be eliminated, that is, the spectral distortion cannot be eliminated.
本発明は、容易に実時間処理てき、かつ高い認識率を確
保てきる音声認識方法を提供することを目的とする。SUMMARY OF THE INVENTION An object of the present invention is to provide a speech recognition method that can easily perform real-time processing and ensure a high recognition rate.
[課題を解決するための手段]
請求項1に記載の本発明は、ニューラルネットワークを
用いて入力音声からその単語を認識する単語認識方法て
ありて、入力音声の周波数特性を算出し、各帯域のそれ
ぞれにおいて時間的に等分割した音声区間のそれぞれを
1つのブロックとして、各ブロックの中で周波数特性の
平均を算出し、それらの平均を対応する帯域毎に正規化
したものを、ニューラルネットワークへの入力として用
いるようにしたものである。[Means for Solving the Problems] The present invention according to claim 1 provides a word recognition method for recognizing a word from an input voice using a neural network, calculating the frequency characteristics of the input voice, and Each of the speech intervals divided equally in time is treated as one block, and the average frequency characteristics are calculated within each block, and the average is normalized for each corresponding band and sent to the neural network. It is designed to be used as an input.
請求項2に記載の本発明は、前記ニューラルネットワー
クか階層的なニューラルネットワークであるようにした
ものである。According to a second aspect of the present invention, the neural network is a hierarchical neural network.
[作用]
請求項1に記載の本発明によれば、下記■〜■の作用効
果かある。[Action] According to the present invention as set forth in claim 1, there are the following effects (1) to (2).
■ニューラルネットワークへ入力する特徴パラメータと
して「周波数特性」を用いたから、入力を得るための前
処理か、LPC相関やLPCケプストラムの如くの複雑
な特徴量抽出に比して単純で並列的に周波数分析でき、
その前処理に要する時間か短くて足りる。■Since "frequency characteristics" are used as the feature parameters input to the neural network, frequency analysis is simpler and more parallel than complex feature extraction such as LPC correlation or LPC cepstrum. I can do it,
The time required for the preprocessing is short.
■ニューラルネットワークは、原理的に、ネットワーク
全体の演算処理か単純かつ迅速である。■Neural networks are, in principle, capable of simple and quick arithmetic processing for the entire network.
■ニューラルネットワークは、原理的に、それを構成し
ている各ユニットが独立に動作しており、並列的な演算
処理が可能である。従って、演算処理が迅速である。■In principle, each unit that makes up a neural network operates independently, and parallel arithmetic processing is possible. Therefore, calculation processing is quick.
■上記■〜■により、音声認識処理を複雑な処理装置に
よることなく容易に実時間処理できる。(2) With the above (2) to (4), voice recognition processing can be easily performed in real time without using a complicated processing device.
■定常的なスペクトル歪に強く、高い認識率を維持でき
る。これは、以下に解析する如く、入力音声の各ブロッ
クでの周波数特性の平均を同一帯域内で正規化するもの
であるため、スペクトル歪を消去できることによる。即
ち、前述の如く、lをブロック番号、kを帯域番号、A
kをに帯域の周波数伝送特性、S■ikを学習段階での
に帯域iブロックの音声信号、S tikを評価段階で
電話回線を通した後における如く、定常的な周波数伝送
特性Akの影響によりスペクトルが歪んだ、k帯域iブ
ロックの音声信号とする時、
S tik = A k−3mi
・・・(1)である。そして、評価段階での各音声信号
S tikを帯域毎に正規化したものは、
Σ 5tik Ak Σ S■ik Σ
S膳ikてあって、周波数伝送特性Akを消去できる
、即ち、スペクトル歪を消去てきるのである。■It is resistant to constant spectral distortion and can maintain a high recognition rate. This is because, as will be analyzed below, the average of the frequency characteristics of each block of input audio is normalized within the same band, so that spectral distortion can be eliminated. That is, as mentioned above, l is the block number, k is the band number, and A
k is the frequency transmission characteristic of the band, Sik is the audio signal of the band i block in the learning stage, and Stik is the voice signal of the band i block in the evaluation stage due to the influence of the steady frequency transmission characteristic Ak, such as after passing through a telephone line. When the spectrum is distorted and it is a k-band i-block audio signal, S tik = A k-3mi
...(1). Then, each audio signal S tik at the evaluation stage is normalized for each band as follows: Σ 5tik Ak Σ S■ik Σ
The frequency transmission characteristic Ak can be eliminated, that is, the spectral distortion can be eliminated.
請求項2に記載の本発明によれば、下記■の作用効果が
ある。According to the present invention as set forth in claim 2, there is the following effect (2).
0階層的なニューラルネットワークにあっては、現在、
後述する如くの簡単な学習アルゴリズム(パックプロパ
ゲーション)が確立されており、高い認識率を実現でき
るニューラルネットワークを容易に形成できる。Currently, in a zero-layer neural network,
A simple learning algorithm (pack propagation) as described below has been established, and a neural network that can achieve a high recognition rate can be easily formed.
[実施例]
第1図は本発明が適用された音声認識システムの一例を
示す模式図、第2図はニューラルネ・ソトヮークを示す
模式図、第3図は階層的なニューラルネットワークを示
す模式図、第4図はユニ・ソトの構造を示す模式図であ
る。[Example] Fig. 1 is a schematic diagram showing an example of a speech recognition system to which the present invention is applied, Fig. 2 is a schematic diagram showing a neural network, and Fig. 3 is a schematic diagram showing a hierarchical neural network. , FIG. 4 is a schematic diagram showing the structure of Uni-Soto.
本発明の具体的実施例の説明に先立ち、ニューラルネッ
トワークの構成、学習アルゴリズムについて説明する。Prior to describing specific embodiments of the present invention, the configuration of the neural network and the learning algorithm will be described.
(1)ニューラルネットワークは、その構造から、第2
図(A)に示す階層的ネ・ソトワークと第2図(B)に
示す相互結合ネットワークの2種に大別できる。本発明
は、両ネットワークのいずれを用いて構成するものてあ
っても良いか、階層的ネ・ソトワークは後述する如くの
簡単な学習アルゴリズムか確立されているためより有用
である。(1) Due to its structure, neural networks
It can be roughly divided into two types: a hierarchical network shown in FIG. 2(A) and an interconnected network shown in FIG. 2(B). Although the present invention may be constructed using either of these networks, the hierarchical network is more useful because a simple learning algorithm as described below has been established.
(2)ネットワークの構造
階層的ネットワークは、第3図に示す如く、入力層、中
間層、出力層からなる階層構造をとる。(2) Network Structure A hierarchical network has a hierarchical structure consisting of an input layer, an intermediate layer, and an output layer, as shown in FIG.
各層は1以上のユニットから構成される。結合は、入力
層→中間層→出力層という前向きの結合たけで、各層内
での結合はない。Each layer is composed of one or more units. The connections are forward connections from the input layer to the middle layer to the output layer, and there are no connections within each layer.
(3)ユニットの構造
ユニットは第4図に示す如く脳のニューロンのモデル化
であり構造は簡単である。他のユニットから入力を受け
、その総和をとり一定の規則(変換関数)で変換し、結
果を出力する。他のユニットとの結合には、それぞれ結
合の強さを表わす可変の重みを付ける。(3) Structure of the unit The unit is a model of a neuron in the brain and has a simple structure as shown in FIG. It receives input from other units, sums it up, transforms it using a certain rule (conversion function), and outputs the result. Each connection with another unit is given a variable weight that represents the strength of the connection.
(4)学習(パックプロパゲーション)ネットワークの
学習とは、実際の出力を目標値(望ましい出力)に近づ
けることであり、一般的には第4図に示した各ユニット
の変換関数及び重みを変化させて学習を行なう。(4) Learning (pack propagation) Learning of a network is to bring the actual output closer to the target value (desired output), and generally changes the conversion function and weight of each unit shown in Figure 4. Let them learn.
又、学習のアルゴリズムとしては、例えば、Ru+5e
lhart、 D、E、、McClelland、 J
、L、 and thePDP Re5earch G
roup、 PARALLEL DISTRIBLIT
EDPROCESSING、 the MIT Pre
ss、 1986.に記載されているパックプロパゲー
ションを用いることができる。Also, as a learning algorithm, for example, Ru+5e
lhart, D.E., McClelland, J.
, L, and thePDP Re5earch G
roup, PARALLEL DISTRIBLIT
EDPROCESSING, the MIT Pre
ss, 1986. Pack propagation as described in .
以下、本発明の具体的な実施例について説明する。Hereinafter, specific examples of the present invention will be described.
認識システム1は、16チヤンネルのバンドパスフィル
タ11、平均化回路12、正規化回路13、ニューラル
ネットワーク20、判定回路30の結合にて構成される
(第1図参照)。The recognition system 1 is composed of a 16-channel bandpass filter 11, an averaging circuit 12, a normalization circuit 13, a neural network 20, and a determination circuit 30 (see FIG. 1).
この認識システム1にあつては、認識単語を47都道府
県名、特定話者を1名とした。以下、認識システム1の
学習動作と評価動作について詳述する。In this recognition system 1, the recognized words were the names of 47 prefectures, and the specific speaker was one person. The learning operation and evaluation operation of the recognition system 1 will be described in detail below.
(学習)
1、入力作成
■各認識単語の既知入力音声波・形を16チヤンネルの
バンドパスフィルタ11に通し、入力音声の周波数特性
を算出する。(Learning) 1. Input creation ■ Pass the known input speech wave/form of each recognition word through the 16-channel bandpass filter 11 to calculate the frequency characteristics of the input speech.
■バントパスフィルタ11の各帯域のそれぞれにおいて
音声波形を時間的に8等分割した音声区間のそれぞれを
1つのブロックとして、平均化回路12により、各ブロ
ックの中で、上記■て求めた周波数特性の平均を算出す
る。この学習段階における音声信号のに帯域1ブロツク
での周波数特性の平均を、S■ikとする。■In each band of the band pass filter 11, each of the voice sections obtained by temporally dividing the voice waveform into 8 equal parts is treated as one block, and the frequency characteristics obtained in the above (■) are determined in each block by the averaging circuit 12. Calculate the average of Let Sik be the average frequency characteristic of the audio signal in one band block at this learning stage.
■上記■で各帯域にて求めた各ブロックの周波数特性の
平均を、対応する帯域の全ブロックのレベルの和ΣS
mikで除算し、対応する帯域毎に、
として正規化する。■The average of the frequency characteristics of each block obtained in each band in the above ■is the sum of the levels of all blocks in the corresponding band ΣS
Divide by mik and normalize as follows for each corresponding band.
■上記■で求めた値をニューラルネットワーク20への
入力とする。入力個数は16チヤンネル×8ブロック=
128個となる。(2) The value obtained in (2) above is input to the neural network 20. Number of inputs is 16 channels x 8 blocks =
There will be 128 pieces.
2、学習
■ 128個の入力層と48個の出力層をもつニューラ
ルネットワーク20を用いる。2. Learning ■ A neural network 20 with 128 input layers and 48 output layers is used.
047個の認識単語のそれぞれに番号付けし、47個の
出力層と対応させ、各認識単語について上記1の■で求
めた入力に対し、その単語に対応した出力層が「1」、
その他の出力層がrOJという値(目標値)になるよう
に、パックプロパゲーションにより5000回学習する
。これにより、一定認識率を保証し得るニューラルネッ
トワーク20を構築する。Each of the 047 recognized words is numbered and made to correspond to the 47 output layers, and for each recognized word, the output layer corresponding to that word is "1" for the input obtained in step 1 above.
Learning is performed 5000 times by pack propagation so that the other output layers have the value rOJ (target value). In this way, a neural network 20 that can guarantee a constant recognition rate is constructed.
(評価)
1、入力作成
■各認識単語の未知入力音声波形を16チヤンネルのバ
ンドパスフィルタ11に通し、入力音声の周波数特性を
算出する。(Evaluation) 1. Input Creation - Pass the unknown input speech waveform of each recognized word through a 16-channel bandpass filter 11 to calculate the frequency characteristics of the input speech.
■バンドパスフィルタ11の各帯域のそれぞれにおいて
音声波形を時間的に8等分割した音声区間のそれぞれを
1つのブロックとして、平均化回路12により、各ブロ
ックの中で、上記■で求めた周波数特性の平均を算出す
る。この評価段階における音声信号のに帯域1ブロツク
での周波数特性の平均を、S tikとする。■In each band of the band pass filter 11, the audio waveform is temporally divided into eight equal parts, each of which is treated as one block, and the averaging circuit 12 calculates the frequency characteristics obtained in the above (■) within each block. Calculate the average of The average frequency characteristic of the audio signal in one band block at this evaluation stage is defined as Stik.
■上記■で各帯域にて求めた各ブロックの周波数特性の
平均を、対応する帯域の全ブロックのレベルの和ΣS
tikで除算し、対応する帯域毎に、
tik
・・・(4)
Σ S tik
として正規化する。■The average of the frequency characteristics of each block obtained in each band in the above ■is the sum of the levels of all blocks in the corresponding band ΣS
Divide by tik and normalize as tik (4) Σ S tik for each corresponding band.
2、学習
■上記■で求めた値をニューラルネットワーク20へ入
力する。2. Learning ■ Input the values obtained in (■) above to the neural network 20.
■ニューラルネットワーク20の出力層の値より判定回
路30にて入力単語を判定する。(2) The determination circuit 30 determines the input word based on the value of the output layer of the neural network 20.
以下、本発明の実験結果について説明する。Hereinafter, experimental results of the present invention will be explained.
(実験1)
本発明例として、周波数特性の平均を帯域毎に正規化し
たものをニューラルネットワーク20への入力とした。(Experiment 1) As an example of the present invention, the average of frequency characteristics normalized for each band was input to the neural network 20.
認識単語を47都道府県名、特定話者を1名とした。The words to be recognized were the names of 47 prefectures, and the specific speaker was one person.
結果、認識率は94.4%であった。As a result, the recognition rate was 94.4%.
(実験2)
比較例として、周波数特性の平均を算出し、単語のパワ
ー全体で正規化したものをニューラルネットワーク20
への入力とした。認識単語を47都道府県名、特定話者
を1名とした。(Experiment 2) As a comparative example, the average frequency characteristics were calculated and normalized by the entire word power, and then the neural network 20
It was used as an input. The words to be recognized were the names of 47 prefectures, and the specific speaker was one person.
結果、認識率は57.0%であった。As a result, the recognition rate was 57.0%.
以下、上記実施例の作用について説明する。Hereinafter, the operation of the above embodiment will be explained.
■ニューラルネットワーク20へ入力する特徴パラメー
タとして「周波数特性」を用いたから、入力を得るため
の前処理が、LPG相関やLPCケプストラムの如くの
複雑な特徴量抽出に比して単純で並列的に周波数分析で
き、その前処理に要する時間が短くて足りる。■Since "frequency characteristics" are used as the feature parameters input to the neural network 20, the preprocessing to obtain the input is simpler and parallel to the frequency characteristics than complex feature extraction such as LPG correlation or LPC cepstrum. can be analyzed, and the time required for preprocessing is short.
■ニューラルネットワーク20は、原理的に、ネットワ
ーク全体の演算処理が単純かつ迅速である。(2) In principle, in the neural network 20, the calculation processing of the entire network is simple and quick.
■ニューラルネットワーク20は、原理的に、それを構
成している各ユニットが独立に動作しており、並列的な
演算処理が可能である。従って、演算処理が迅速である
。(2) In principle, each unit constituting the neural network 20 operates independently, and parallel arithmetic processing is possible. Therefore, calculation processing is quick.
■上記■〜■により、音声認識処理を複雑な処理袋!に
よることなく容易に実時間処理できる。■With the above ■~■, voice recognition processing becomes complicated! It can be easily processed in real time without having to rely on
■定常的なスペクトル歪に強く、高い認識率を維持てき
る。これは、[作用]の■にて前述の如く、評価段階で
正規化された(4)式の如くの値が、(2)式にて解析
された如くに周波数伝送特性Akを消去されて、学習段
階で正規化された(3)式の如くの値と同等となり、雑
音の影響や回線等の入力系の相違に起因するスペクトル
歪を消去できるからである。■It is resistant to constant spectral distortion and maintains a high recognition rate. This is due to the fact that the value of equation (4) normalized at the evaluation stage is canceled out by the frequency transmission characteristic Ak as analyzed using equation (2), as mentioned above in [effect]. This is because it becomes equivalent to the value of equation (3) normalized in the learning stage, and spectral distortion caused by the influence of noise or differences in input systems such as lines can be eliminated.
■階層的なニューラルネットワーク20にあっては、現
在、前述の如くの簡単な学習アルゴリズム(パックプロ
パゲーション)が確立されており、高い認識率を実現で
きるニューラルネットワーク20を容易に形成できる。(2) Regarding the hierarchical neural network 20, a simple learning algorithm (pack propagation) as described above has been established at present, and a neural network 20 that can achieve a high recognition rate can be easily formed.
[発明の効果]
以上のように本発明によれば、容易に実時間処理でき、
かつ高い認識率を確保できる音声認識方法を得ることが
できる。[Effects of the Invention] As described above, according to the present invention, real-time processing is easily possible.
Moreover, it is possible to obtain a speech recognition method that can ensure a high recognition rate.
第1図は本発明が適用された音声認識システムの一例を
示す模式図、第2図はニューラルネットワークを示す模
式図、第3図は階層的なニューラルネットワークを示す
模式図、第4図はユニットの構造を示す模式図である。
1・・・認識システム、
10・・・バントパスフィルタ、
12・・・平均化回路、
13・・・正規化回路、
20・・・ニューラルネットワーク、
30・・・判定回路。Fig. 1 is a schematic diagram showing an example of a speech recognition system to which the present invention is applied, Fig. 2 is a schematic diagram showing a neural network, Fig. 3 is a schematic diagram showing a hierarchical neural network, and Fig. 4 is a schematic diagram showing a unit. FIG. DESCRIPTION OF SYMBOLS 1... Recognition system, 10... Band pass filter, 12... Averaging circuit, 13... Normalization circuit, 20... Neural network, 30... Judgment circuit.
Claims (2)
の単語を認識する単語認識方法であって、入力音声の周
波数特性を算出し、各帯域のそれぞれにおいて時間的に
等分割した音声区間のそれぞれを1つのブロックとして
、各ブロックの中で周波数特性の平均を算出し、それら
の平均を対応する帯域毎に正規化したものを、ニューラ
ルネットワークへの入力として用いる音声認識方法。(1) A word recognition method that uses a neural network to recognize words from input speech, in which the frequency characteristics of the input speech are calculated, and each speech interval divided equally in time in each band is divided into one A speech recognition method in which the average frequency characteristics of each block are calculated, and the averages are normalized for each corresponding band and used as input to a neural network.
ルネットワークである請求項1記載の音声認識方法。(2) The speech recognition method according to claim 1, wherein the neural network is a hierarchical neural network.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2026670A JPH03230200A (en) | 1990-02-05 | 1990-02-05 | Voice recognizing method |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2026670A JPH03230200A (en) | 1990-02-05 | 1990-02-05 | Voice recognizing method |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH03230200A true JPH03230200A (en) | 1991-10-14 |
Family
ID=12199837
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP2026670A Pending JPH03230200A (en) | 1990-02-05 | 1990-02-05 | Voice recognizing method |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH03230200A (en) |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2005020212A1 (en) * | 2003-08-22 | 2005-03-03 | Sharp Kabushiki Kaisha | Signal analysis device, signal processing device, speech recognition device, signal analysis program, signal processing program, speech recognition program, recording medium, and electronic device |
| JP2006343544A (en) * | 2005-06-09 | 2006-12-21 | Miyazaki Prefecture | Voice recognition method |
| US7343288B2 (en) | 2002-05-08 | 2008-03-11 | Sap Ag | Method and system for the processing and storing of voice information and corresponding timeline information |
| US7406413B2 (en) | 2002-05-08 | 2008-07-29 | Sap Aktiengesellschaft | Method and system for the processing of voice data and for the recognition of a language |
-
1990
- 1990-02-05 JP JP2026670A patent/JPH03230200A/en active Pending
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7343288B2 (en) | 2002-05-08 | 2008-03-11 | Sap Ag | Method and system for the processing and storing of voice information and corresponding timeline information |
| US7406413B2 (en) | 2002-05-08 | 2008-07-29 | Sap Aktiengesellschaft | Method and system for the processing of voice data and for the recognition of a language |
| WO2005020212A1 (en) * | 2003-08-22 | 2005-03-03 | Sharp Kabushiki Kaisha | Signal analysis device, signal processing device, speech recognition device, signal analysis program, signal processing program, speech recognition program, recording medium, and electronic device |
| CN1839427B (en) | 2003-08-22 | 2010-04-28 | 夏普株式会社 | Signal analysis device, signal processing device, speech recognition device, and electronic apparatus |
| JP2006343544A (en) * | 2005-06-09 | 2006-12-21 | Miyazaki Prefecture | Voice recognition method |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US5461697A (en) | Speaker recognition system using neural network | |
| JPH03273722A (en) | Sound/modem signal identifying circuit | |
| WO2006000103A1 (en) | Spiking neural network and use thereof | |
| JPH06161496A (en) | Voice recognition system for recognizing remote control command words for home appliances | |
| WO1998037543A1 (en) | Method and apparatus for training a speaker recognition system | |
| Gaikwad et al. | Classification of Indian classical instruments using spectral and principal component analysis based cepstrum features | |
| JPH03230200A (en) | Voice recognizing method | |
| Sunny et al. | Feature extraction methods based on linear predictive coding and wavelet packet decomposition for recognizing spoken words in malayalam | |
| EP0369485B1 (en) | Speaker recognition system | |
| JPH0462599A (en) | Noise removing device | |
| El-Wakdy et al. | Speech Recognition Using a Wavelet Transform to Establish Fuzzy Inference System Through Subtractive Clustering and Neural Network(ANFIS) | |
| JPH03230255A (en) | Sound recognizing method | |
| JPH02273798A (en) | Speaker recognition system | |
| Sunny et al. | Development of a speech recognition system for speaker independent isolated Malayalam words | |
| JPH05143094A (en) | Speaker recognition system | |
| JPH02275996A (en) | Word recognition system | |
| JPH03157697A (en) | Word recognizing system | |
| JPH03157698A (en) | Speaker recognizing system | |
| Czyżewski | Soft processing of audio signals | |
| JPH03230256A (en) | Voice recognizing method | |
| JPH02273799A (en) | Speaker recognition system | |
| George et al. | Speaker recognition using dynamic synapse based neural networks with wavelet preprocessing | |
| JP2518939B2 (en) | Speaker verification system | |
| JPH02135500A (en) | Talker recognizing system | |
| JPH05181500A (en) | Word recognition system |