JPS59201100A - Voice standard pattern registration system - Google Patents

Voice standard pattern registration system

Info

Publication number
JPS59201100A
JPS59201100A JP58076562A JP7656283A JPS59201100A JP S59201100 A JPS59201100 A JP S59201100A JP 58076562 A JP58076562 A JP 58076562A JP 7656283 A JP7656283 A JP 7656283A JP S59201100 A JPS59201100 A JP S59201100A
Authority
JP
Japan
Prior art keywords
voice
standard
button
registration
dictionary
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP58076562A
Other languages
Japanese (ja)
Other versions
JPH037960B2 (en
Inventor
加世田 光子
佐藤 泰雄
教幸 藤本
一成 畑中
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Fujitsu Ltd
Original Assignee
Fujitsu Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Fujitsu Ltd filed Critical Fujitsu Ltd
Priority to JP58076562A priority Critical patent/JPS59201100A/en
Publication of JPS59201100A publication Critical patent/JPS59201100A/en
Publication of JPH037960B2 publication Critical patent/JPH037960B2/ja
Granted legal-status Critical Current

Links

Abstract

(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。
(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.

Description

【発明の詳細な説明】 (al  発明の技術分野 本発明は単語または単音節を認識対象とする音?認識に
おける音声標準バタンの登録方式に関する。
DETAILED DESCRIPTION OF THE INVENTION Technical Field of the Invention The present invention relates to a method for registering standard speech sounds in sound recognition in which words or monosyllables are recognized.

(bl  技術の背景 近年データ処理技術の発達と普及に伴いデータ処j18
システムにおけるデータ入出力手段の一端として、自初
は音声制御による仕分け、電話回線における案内サービ
ス程度にとソまっていた音声認識−・合成技術も半導体
特に集積化技術と回路設計技術の進展に支えられ、演算
処理の高速あるいは大容量記憶を要する実現手段の小形
且低コスト化が得られるようになって、日本語による音
声入出力手段が分散処理および対話形式に適し操作者に
特別の習熟を必要とすることのない操作が容易な入力音
声〜テンクルデー2間の変換機能を生かしたデータ処理
装置として普及するようになった。
(bl Technological Background With the development and spread of data processing technology in recent years, data processing j18
As part of the data input/output means in the system, voice recognition/synthesis technology, which was originally limited to voice-controlled sorting and guidance services on telephone lines, has also been supported by advances in semiconductors, especially integration technology and circuit design technology. As a result, implementation means that require high-speed arithmetic processing or large-capacity storage can be made smaller and lower in cost, and voice input/output means in Japanese are suitable for distributed processing and interactive formats and require special skill for operators. It has become popular as a data processing device that takes advantage of the conversion function between input voice and Tenkle Day 2, which is easy to operate and does not require any input.

(C1従来技術と問題点 従来より音声認識装置は通常特定話者のため認識すべき
入力音声における複数の単語または/および単音節を設
定して、先行入力する各単語または/および単音節を予
め帯域フィルタ群に印加して得るスペクトル出力毎に標
本化して得た特徴パラメータをデータとして蓄積し、こ
れを各単語ま尼は/および単音節に対応する音声標準バ
タンとし、その后は該話者の入力音声による音声バタン
を該標準バタンと比較することによって未知音声を入力
する都度対応するディジタルデータに変換する機能を備
えている。従って音声認識装置では入力音声による音声
パターンを認識するため、単語寸たは/および単音節に
対応する音声標準バタンを登録・する都度n回例えば4
〜8回ずつ、複数P ’jMの単音節例えば68個lた
は101個では総計68/l0IX(4〜8)の発声を
必要とする他のデータ入力装置にはない煩わしさが存在
する。この音声・漂卓バタン登録時の発声繰返しは話者
の負担だけではなく例えばRA Mによる記憶容量およ
び装置におけるデータ処理量が増大するのでコスト上か
らも少い方が望ましいが単純に発声回数を減少すること
(’:L m a= m能の信頼性を低下する欠点があ
った。
(C1 Prior Art and Problems Conventionally, speech recognition devices usually set a plurality of words and/or monosyllables in the input speech to be recognized for a specific speaker, and pre-input each word or/and monosyllable. The characteristic parameters obtained by sampling each spectrum output obtained by applying it to a group of bandpass filters are stored as data, and this is used as a speech standard button corresponding to each word and/or monosyllable. The device has a function of converting an unknown voice into corresponding digital data each time it is input by comparing the voice bang of the input voice with the standard bang.Therefore, in order to recognize the voice pattern of the input voice, the voice recognition device For example, 4 times each time you register a phonetic standard button corresponding to a single syllable.
For example, 68 l or 101 monosyllables of P'jM each time ~8 times each require a total of 68/10 IX (4 to 8), which is a hassle not found in other data input devices. Repetition of utterances when registering voice/speech buttons not only burdens the speaker, but also increases the storage capacity of RAM and the amount of data processing in the device, so it is desirable to reduce the number of utterances from a cost standpoint. This has the disadvantage of reducing the reliability of the ability (': L m a = m).

tdl  発明の目的 本発明の目的は上記の欠点を除去するため、よりlp数
回可能な限り例えば単語または/および単廿節毎に1回
、心安な対象には2〜3回レベルによって音声81準パ
ターンを登録して寧ろ従来の複u、回ずりの)6戸」入
力によるレベルに匹敵する音声標準パターンを確保せし
めるところの、発声回数の削減とイ己順性の確保を両立
させる音声標準バタン登録の手段を提供しようとするも
のである。
tdl OBJECTS OF THE INVENTION The object of the present invention is to eliminate the above-mentioned drawbacks, in order to eliminate the above-mentioned drawbacks by repeating the voice 81 as many times as possible, for example once per word or/and single clause, and for safe subjects 2-3 times. A voice standard that achieves both a reduction in the number of utterances and a guarantee of self-sequence by registering a quasi-pattern to ensure a voice standard pattern that is comparable to the level achieved by the conventional multiple u, rotation) input. This is intended to provide a means of registration.

fel  発明のイ′1イ成 この目的は、未知入力音声の認識を予め辞書中に登録さ
れた複数の音声標準バタンと入力音声バタンとの照合に
よって行う音声認識システムにおいて、音声標準バタン
候補の多数からなりたつ音声標準バタン候補辞書を有し
、登録時制御部は登録する特定話者の音声を音声処理部
に入力して得られる音声バタンと該候補辞書中の音声標
1社バタン候補とを比較してその類似度を求め、登録す
べき音声毎に設定した!Jii似度の閾値および登録数
に応じて、該閾値以上かつ登録数以下の音声捺蘂バタン
候補を選択し、記憶部に記憶登録せしめ特定話者め音声
標準バタンからなる辞書とすることを特徴とする音声標
準バタン登録方式を(J供することによって達成するこ
とが出来る。
fel A'1 Achievement of the Invention The purpose of this invention is to recognize a large number of speech standard slam candidates in a speech recognition system that recognizes an unknown input speech by comparing input speech stamps with a plurality of speech standard stamps registered in advance in a dictionary. It has a voice standard slam candidate dictionary consisting of , and at the time of registration, the control unit compares the voice slam obtained by inputting the voice of the specific speaker to be registered into the voice processing unit with the phonetic standard one company slam candidates in the candidate dictionary. The similarity was determined and set for each voice to be registered! According to a threshold of Jii similarity and the number of registrations, candidates for voice stamps that are equal to or greater than the threshold and less than or equal to the number of registrations are selected and stored and registered in a storage unit to form a dictionary consisting of standard speech stamps for a specific speaker. This can be achieved by providing an audio standard bang registration method (J).

+f+  発明の実施例 以下図面を参照しつ\本発明の一実施夕1」について説
明する。第1図は本発明の一実施例における音声標準バ
タン登録方式のプロ、り図、8g2図は音声標準バタン
候補辞書における音声標準バタン候補、標準バタンおよ
び音声バタンの相関を示す模式図および第3図は本発明
の一実施例における音声標準バタン登録方式における処
理手順を示すフローチャートである。図において1は制
御部、2は記憶部、21は制御プログラム、22は制御
データ、23は音声標準バタン候補辞書、23a。
+f+ Embodiment of the Invention An embodiment 1 of the present invention will be described below with reference to the drawings. FIG. 1 is a professional diagram of the audio standard button registration method in an embodiment of the present invention, FIG. The figure is a flowchart showing the processing procedure in the audio standard button registration method according to an embodiment of the present invention. In the figure, 1 is a control unit, 2 is a storage unit, 21 is a control program, 22 is control data, 23 is a speech standard slam candidate dictionary, and 23a.

b・・・p、は音声標準バタン候補群、23aa 、 
ab、 ac 、・・・8g・・・は音声標準バタン候
補、24は音p5登録標準バタン辞曹、24a 、 b
・・・pは音声標準バタンgl、24aa 、ab・・
・ah・・・は音声標準バタンである。制御部1は記憶
部2の記憶領域に”蓄積する制御プログラム21および
制御データ22に従って構成各部を制御して音声入力信
号に伴いその音声標準バタンを選択して特定話者に対応
する音声標準バタン辞書24を作成する。記憶部2はそ
の記憶領域に制御プログラム21および各標準バタン候
補と比較対象となる音声バタンとの類似既における回位
あるいは各単語または/および単音節に対応する憚準バ
タン候補群23a 、 b・・・p毎から選択して標準
バタン枇24a 、 b・・・p毎に登録する標準バタ
ンの認(以下登録数)値を設定する。
b...p is a voice standard slam candidate group, 23aa,
ab, ac,...8g... are voice standard bang candidates, 24 is sound p5 registered standard bang dictionary, 24a, b
...p is the audio standard button gl, 24aa, ab...
・ah... is a standard voice button. The control unit 1 controls each component according to the control program 21 and control data 22 stored in the storage area of the storage unit 2, selects the standard audio button in response to the audio input signal, and selects the standard audio button corresponding to the specific speaker. A dictionary 24 is created.The storage unit 2 stores in its storage area the control program 21 and the similarity level of each standard bang candidate with the voice bang to be compared, or the standard bang corresponding to each word and/or monosyllable. The recognition (hereinafter referred to as the number of registrations) value of the standard button to be selected from each candidate group 23a, b...p and registered for each standard button group 24a, b...p is set.

尚単音節例えばパア″に対応する多数話者(a+b+C
・・・g)の音声バタンは標準バタン候補群23aにお
ける標準バタン族+nj 23 aa −agにd己1
]ξ(されており、未知話者の“ア″のための標準バタ
ン24aa。
It should be noted that multiple speakers (a+b+C
...G)'s audio bang is the standard bang group +nj 23 aa -ag in the standard bang candidate group 23a.
] ξ (standard slam 24aa for "a" of an unknown speaker.

ab・・・abは憬争バクン群24aに収容されるもの
とする。但し標準バタン群袖23aaは標準バタン24
aaとは直裁対応するものではない。従って標準バタン
候補群の23 a −1) 、材2竿バクン荏Iの24
a〜pの数は青しく且各単mlまたは/および単行11
1jの琲位総4又に対応する。例え(+:(68−6:
たば101である才だ各標渠バタン@イIjイ拌23a
〜pに共通する標準バタン候補の蚊a−gは予め畜積し
た多V、話者gの叡に対応し、標準バタン釧−24a−
pに共通する標準バタンa −hは金U叡((対応し・
ン[jえは6でおる。こ\で本発明の一実施例において
は図示省略したが通常特定話者の発hレリえば゛ア″を
マイクロフォンに入力して得られるアナログ′電気信号
による入力16号を音声処理m3に入力してその特徴パ
ラメータを抽出して音声バタンを作成する。
It is assumed that ab...ab is accommodated in the conflict Bakun group 24a. However, the standard batan group sleeve 23aa is the standard batan 24
AA does not correspond to direct judgment. Therefore, 23 a -1) of the standard batan candidate group, 24 of the material 2 rod Bakun Ei I
The numbers a to p are in blue and each single ml or/and single row 11
It corresponds to the total four prongs of 1j. Example (+: (68-6:
Taba 101 is a talented person.
The standard slam candidates a-g that are common to ~p correspond to the pre-accumulated multi-V and speaker g's words, and are standard slam candidates -24a-
The standard batons a - h common to p are Kim Ue ((corresponding to
N [j is 6. Although not shown in this embodiment of the present invention, input No. 16 in the form of an analog electrical signal obtained by inputting a particular speaker's utterance "A" into a microphone is input to the audio processing m3. and extracts its characteristic parameters to create a voice button.

即ち入力信号を音声周波数200〜5400Hzをmチ
ャンネル例えば16の帯域フィルタと時間的変化をn個
例えば16または32個に標本化する手段によって得ら
れるスペクトルの特徴を256または512個のデータ
に表現する音声バタンXaに変換して送出する、音声バ
タンXaを印加された比較部4け制御部1の制御に従い
第2図に示す標準バタン候補群23aの○印に対応する
標準バタン候補23aa−agのデータと比較して予め
RrlJ 御データ22に設定された閾値の範囲で最も
類似度の高い即ちデータとの距離が近い標準バタン候補
から順に同じく制御データ22の登m数だけ例えば6個
、二ハ択して標準バタン群24aに標準バタンaa〜a
hこ\では6個の標準バタンaa−afを登録する。同
様に他の音声バタンXpは標準バタン候補群23pの0
印に対応する標準バタン候補23pa〜pgのデータと
比較して標準バタンpa〜pfを標準バタン群24pに
登録する。第2図における標準バタン候補群23aを示
す変形楕円は領域を囲む外部線ではなく最外分布部に存
在する○印の標A% /(タン候補を結ぶ表示線であり
、同様にセ、(準](タン群24aを示す円形もX印に
より示した音声)<タンXaから近い距離に選択した標
準ノ(タンを結んだ表示線である。このように特定Hf
fi者の1回発声による音声バタンによっても消去にa
KAした多数話者の音声バタンにおけるデータによって
構成されるd弗バタン候補辞書23の中から選択して標
準バタン辞■24を作成すれは従来標漁)々タン群24
a 、 b・・・pを登録するのに複数回ずつ発声を必
要としていた煩しさを谷単語または/単音節毎に1回ず
つ計p回の発声だけで特定話者に対応する標準バタン辞
書24が発録出来るので有用である。
That is, the spectral characteristics obtained from the input signal by means of sampling the audio frequency of 200 to 5400 Hz into m channels, e.g., 16 bandpass filters, and the temporal changes into n, e.g., 16 or 32, are expressed as 256 or 512 pieces of data. Comparison unit 4-digit comparison unit 1 to which the audio bang Xa is applied, which is converted into an audio slam Xa and sent out, selects the standard slam candidates 23aa-ag corresponding to the circle marks in the standard baton candidate group 23a shown in FIG. RrlJ is compared with the control data 22 in advance, and the standard slam candidates with the highest degree of similarity within the threshold range set in the control data 22, i.e., those with the closest distance to the data, are selected, for example, by the number of meters of the control data 22, for example, Select the standard button group 24a to select the standard button aa~a.
In hko\, six standard buttons aa-af are registered. Similarly, the other voice button Xp is 0 in the standard button candidate group 23p.
The data of the standard slam candidates 23pa to 23pg corresponding to the marks are compared and the standard batons pa to pf are registered in the standard slam group 24p. The deformed ellipse showing the standard button candidate group 23a in FIG. 2 is not an external line surrounding the area, but a display line connecting the mark A% / standard] (The circle indicating the tongue group 24a is also indicated by an X mark)
It can also be erased by a single voice click from the fi person.
A standard batan dictionary 24 is created by selecting it from the d弗 batan candidate dictionary 23, which is composed of the data of the voice batans of many speakers who have KA.
We have replaced the trouble of having to pronounce each word multiple times to register a, b...p with a standard slam dictionary that can be used by a specific speaker with just one utterance for each word or monosyllable, p times in total. It is useful because 24 can be recorded.

尚上り己の説明では発声に伴う音声データ例えばXa第
2図の×印点即ち音声ノ(クンXaについては標準バタ
ン枇24aの標準)くタンa a −a fには採用し
なかったが音声バタンXa自牙についても例えはbv<
準バタンagとして標準ノくタン群24aの構成とすれ
はより高い密度のデータとして期待できる。
In addition, in my own explanation, the sound data accompanying the vocalization, for example, the X mark in Figure 2 of Xa, that is, the sound (for Kun Xa, the standard of the standard button 24a) is not adopted for aa - a f, but the sound The analogy for Batan Xa self-fang is bv<
The structure of the standard notch group 24a as a quasi-button ag can be expected to provide higher density data.

梃にある入力信号による音声ノ(タンZa12図に示す
■印点のように従来の音声バタン候補23aa−agと
は著しく具なる)1イリ度として分布−から逸脱して得
られたときは、これを誤り入力信号または音声処11部
3の誤動作として制御部1が判定して取次のデータ処理
を抑止し標準バタンaa〜ahを設廼しないように制御
すれば誤った標準バタン辞書24が登録されること(1
才ない。この時は必要にJニジ図示省陥し、化がその旨
例えば注意表示をして1!+度話者に同−音こ\では例
えば17′″を発声させるようにする。また本実施例で
は標準バタン候補として多数話者の音声バタンを用いた
が、多数の該音声バタンから平均化等の手法により合成
するバタンないしはその両方を用いても同様に実現出来
ることはいう迄もない1、+g+  発明の詳細 な説明したように本発明によれば従来特定話者の未知音
声を認トにするためp個の巣語または/および単音節に
対する標準バタン辞書を登録するのに各複数回ずつを発
声させて得た煩しさとそのデータ処理に対して1回また
はより少数回によりて今録出来るので話者の発声におけ
る煩しさとその処理工数を大幅に減殺出来るので有用で
ある。
When the voice sound due to the input signal on the lever (as shown in the ■ mark in Figure 12, which is significantly different from the conventional voice button candidate 23aa-ag) is obtained with a deviation from the distribution as 1 degree, If the control unit 1 determines this as an erroneous input signal or a malfunction of the audio processor 11 section 3, suppresses the intermediary data processing, and controls so as not to set the standard buttons aa to ah, an incorrect standard button dictionary 24 will be registered. to be done (1
I'm not talented. At this time, it is necessary to omit the illustrations, and the system should display a warning to that effect, for example, 1! For example, the speaker is asked to utter 17''' for the same sound.Also, in this embodiment, the voice bangs of many speakers are used as standard bang candidates, but the average number of voice bangs from a large number of voice bangs is It goes without saying that the same effect can be achieved by using a synthesized button or both of the methods described above1, To register a standard slam dictionary for p nest words and/or monosyllables, we now have to deal with the hassle of having to utter each word multiple times and process the data once or fewer times. This is useful because it can greatly reduce the troublesomeness of the speaker's utterances and the amount of processing time required.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は本発明の一実施例における音声標準ノくタン登
録方式のプロ、り図、第2図は音声標準ノ(邦 1 図 第 2 図 第′3 図
Figure 1 is a professional diagram of the voice standard registration method in one embodiment of the present invention, and Figure 2 is a diagram of the voice standard registration method (Japanese).

Claims (3)

【特許請求の範囲】[Claims] (1)未知入力音声の認識を予め辞芸中に登録された被
数の音声標準バタンと入力音1−バタンとの照合によっ
て行う音声認滝システムにおいて、音声標準バタン候補
の多数からなりたつ音声標準バタン候補辞書を有し、登
録時制御部は特定品者の音声を音声処理部に入力して得
られる音声バタンと該候補辞書中の音声elバタン候補
とを比軟してその類似度を求め、登録すべき音声毎に設
定した類似度の閾値2よび登録数に応じて、該閾値以上
かつ登録数以下の音声椋単バタン候補を選択し、記憶部
に記憶登録せしめ特定話者の音声4ぶ準バクンからなる
辞書とすること全特徴とする音声標準バタン登録方式。
(1) In a voice recognition system that recognizes an unknown input voice by comparing the number of voice standard bangs registered in advance during a speech with the input sound 1-bang, a voice standard consisting of a large number of voice standard bang candidates is used. It has a slam candidate dictionary, and the control unit at the time of registration compares the voice button obtained by inputting the voice of a specific product to the voice processing unit and the voice el button candidates in the candidate dictionary and calculates the degree of similarity between them. , according to the similarity threshold 2 set for each voice to be registered and the number of registrations, select voice single-bang candidates that are above the threshold and below the number of registrations, and store and register them in the storage unit. The dictionary is made up of quasi-bakuns, and the speech standard baptan registration method is fully characterized.
(2)上記制御部は選択した音声標準バタン候補と共に
、登録時の入力音声により得られた該音声バタンを併せ
て記憶部に記憶登録せしめ特定品名−の音声標準バタン
からなる辞書とすることを特徴とする特許請求の範囲第
1項記載の音声標準バタン登録方式。
(2) The control unit stores and registers the selected audio standard button candidate and the audio button obtained from the input voice at the time of registration in the storage unit to form a dictionary consisting of the audio standard button of the specific product name. A voice standard button registration method according to claim 1, which is characterized by:
(3)  上記制御部は登録時音声処理部に入力して得
られる音声バタンか音声標準バタン候補辞書の音声標準
バタン候補による分布より逸脱することを検出したとき
は、該音声バタンによって行う音声標準バタンの記憶登
録を抑止することを特徴とする特許請求の範囲第1項記
載の音声標準バタン登録方式。
(3) When the control unit detects that the voice beat obtained by inputting it to the voice processing unit at the time of registration deviates from the distribution according to the voice standard slam candidates in the voice standard beat candidate dictionary, the control unit performs the voice standard based on the voice beat. 2. The audio standard button registration method according to claim 1, wherein storage registration of a button is suppressed.
JP58076562A 1983-04-30 1983-04-30 Voice standard pattern registration system Granted JPS59201100A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP58076562A JPS59201100A (en) 1983-04-30 1983-04-30 Voice standard pattern registration system

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP58076562A JPS59201100A (en) 1983-04-30 1983-04-30 Voice standard pattern registration system

Publications (2)

Publication Number Publication Date
JPS59201100A true JPS59201100A (en) 1984-11-14
JPH037960B2 JPH037960B2 (en) 1991-02-04

Family

ID=13608679

Family Applications (1)

Application Number Title Priority Date Filing Date
JP58076562A Granted JPS59201100A (en) 1983-04-30 1983-04-30 Voice standard pattern registration system

Country Status (1)

Country Link
JP (1) JPS59201100A (en)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS602998A (en) * 1983-06-20 1985-01-09 富士通株式会社 Method of composing voice dictionary for voice recognition system
JPH05143093A (en) * 1990-10-23 1993-06-11 Internatl Business Mach Corp <Ibm> Method and apparatus for forming model of uttered word

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS57133495A (en) * 1981-02-12 1982-08-18 Oki Electric Ind Co Ltd Voice registering method for voice typewriter

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS57133495A (en) * 1981-02-12 1982-08-18 Oki Electric Ind Co Ltd Voice registering method for voice typewriter

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS602998A (en) * 1983-06-20 1985-01-09 富士通株式会社 Method of composing voice dictionary for voice recognition system
JPH05143093A (en) * 1990-10-23 1993-06-11 Internatl Business Mach Corp <Ibm> Method and apparatus for forming model of uttered word

Also Published As

Publication number Publication date
JPH037960B2 (en) 1991-02-04

Similar Documents

Publication Publication Date Title
Schuller et al. The INTERSPEECH 2021 computational paralinguistics challenge: COVID-19 cough, COVID-19 speech, escalation & primates
Potamianos et al. Robust recognition of children's speech
US20040148161A1 (en) Normalization of speech accent
JPS62239231A (en) Speech recognition method by inputting lip picture
WO2000058943A1 (en) Speech synthesizing system and speech synthesizing method
US12512102B2 (en) Providing prompts in speech recognition results in real time
JPH0756594A (en) Device and method for recognizing unspecified speaker&#39;s voice
CN117894294B (en) Personification auxiliary language voice synthesis method and system
Liao et al. Formosa speech recognition challenge 2020 and taiwanese across taiwan corpus
Shahin Gender-dependent emotion recognition based on HMMs and SPHMMs
Shahin Employing both gender and emotion cues to enhance speaker identification performance in emotional talking environments
JP6849977B2 (en) Synchronous information generator and method for text display and voice recognition device and method
Trinh et al. Directly comparing the listening strategies of humans and machines
US7844459B2 (en) Method for creating a speech database for a target vocabulary in order to train a speech recognition system
US20040006469A1 (en) Apparatus and method for updating lexicon
JPWO2020136948A1 (en) Speech rhythm converters, model learning devices, their methods, and programs
Kurian et al. Connected digit speech recognition system for Malayalam language
JP3576066B2 (en) Speech synthesis system and speech synthesis method
Sasmal et al. Robust automatic continuous speech recognition for'Adi', a zero-resource indigenous language of Arunachal Pradesh
KR102457822B1 (en) apparatus and method for automatic speech interpretation
JPS5941226B2 (en) voice translation device
Mittal et al. Speaker-independent automatic speech recognition system for mobile phone applications in Punjabi
Sudhakar et al. Development of Concatenative Syllable-Based Text to Speech Synthesis System for Tamil
JPH037960B2 (en)
JPS6126678B2 (en)