JPH0361958B2 - - Google Patents

Info

Publication number
JPH0361958B2
JPH0361958B2 JP57112923A JP11292382A JPH0361958B2 JP H0361958 B2 JPH0361958 B2 JP H0361958B2 JP 57112923 A JP57112923 A JP 57112923A JP 11292382 A JP11292382 A JP 11292382A JP H0361958 B2 JPH0361958 B2 JP H0361958B2
Authority
JP
Japan
Prior art keywords
signal
pushphone
recognition
signals
voice
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Expired - Lifetime
Application number
JP57112923A
Other languages
Japanese (ja)
Other versions
JPS593498A (en
Inventor
Yasuo Takahashi
Toshishige Sakai
Haruo Asada
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Toshiba Corp
Original Assignee
Tokyo Shibaura Electric Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Tokyo Shibaura Electric Co Ltd filed Critical Tokyo Shibaura Electric Co Ltd
Priority to JP57112923A priority Critical patent/JPS593498A/en
Publication of JPS593498A publication Critical patent/JPS593498A/en
Publication of JPH0361958B2 publication Critical patent/JPH0361958B2/ja
Granted legal-status Critical Current

Links

Description

【発明の詳細な説明】 〔発明の技術分野〕 本発明は電話回線を通じて入力される音声信号
とプツシユホン信号とをそれぞれ確実に認識する
ことのできる音声認識装置に関する。
DETAILED DESCRIPTION OF THE INVENTION [Technical Field of the Invention] The present invention relates to a voice recognition device that can reliably recognize voice signals and pushphone signals input through a telephone line.

〔発明の技術的背景とその問題点〕[Technical background of the invention and its problems]

近時、情報化社会の発達に伴つて電話回線を介
して接続された系において、音声信号や電話機か
ら発せられるプツシユホン信号をそれぞれ認識し
てデータ処理を行うことが考えられている。とこ
ろがこのような音声信号とプツシユホン信号と云
う明らかに性質の異なる信号を1つのアルゴリズ
ムに従つて認識処理することは甚だ困難であり、
またその認識精度の向上も望めない。そこで従来
では第1図に示すようにプツシユホン信号を認識
する為の専用のアルゴリズムを備えたプツシユホ
ン信号認識部1と音声信号を認識する為の専用の
アルゴリズムを備えた音声認識部2と、これらの
認識部1,2による認識結果を総合判定する総合
判定部3とにより音声認識装置を構成することが
行われている。
BACKGROUND ART Recently, with the development of the information society, it has been considered to perform data processing by recognizing voice signals and pushphone signals emitted from telephones in systems connected via telephone lines. However, it is extremely difficult to recognize and process such signals, which have clearly different characteristics, such as a voice signal and a pushphone signal, according to a single algorithm.
Furthermore, improvement in recognition accuracy cannot be expected. Therefore, conventionally, as shown in Fig. 1, a pushphone signal recognition section 1 equipped with a dedicated algorithm for recognizing pushphone signals, a voice recognition section 2 equipped with a dedicated algorithm for recognizing voice signals, and a A speech recognition device is constituted by a comprehensive judgment section 3 that comprehensively judges the recognition results obtained by the recognition sections 1 and 2.

このようにすれば音声認識部2における認識対
象語数を整理することができるので、或る程度信
頼性の高い認識処理を行うことが可能となる。然
し乍ら、例えば曖昧な信号が入力された場合等、
プツシユホン信号認識部1はこれを音声信号であ
るとして確実にリジエクトすることが困難であ
り、また音声信号認識部2にあつても同様にこれ
をプツシユホン信号であると認定して確実にリジ
エクトすることが困難である為、結局総合判定部
3においても上記入力信号がプツシユホン信号で
あるが、或いは音声信号であるかを確実に識別す
ることができないと云う問題があつた。またこの
ような不具合を解消する為には各認識部1,2の
リジエクト能力を高めなければならず、結局装置
構成が複雑化すると云う問題があつた。またこの
ような複雑化に見合う効果がさほど期待されない
と云う問題もあつた。
In this way, the number of words to be recognized by the speech recognition unit 2 can be arranged, so that recognition processing can be performed with a certain degree of reliability. However, if, for example, an ambiguous signal is input,
It is difficult for the pushphone signal recognition unit 1 to reliably reject this as a voice signal, and it is also difficult for the voice signal recognition unit 2 to similarly recognize this as a pushphone signal and reliably reject it. Since it is difficult to determine whether the input signal is a pushphone signal or an audio signal, the overall determination section 3 has a problem in that it cannot reliably identify whether the input signal is a pushphone signal or an audio signal. Furthermore, in order to eliminate such problems, it is necessary to improve the reject ability of each of the recognition sections 1 and 2, resulting in a problem that the device configuration becomes complicated. There was also the problem that the effects commensurate with such increased complexity were not expected to be that great.

〔発明の目的〕[Purpose of the invention]

本発明はこのような事情を考慮してなされたも
ので、その目的とするところは、簡易に且つ確実
に音声信号とプツシユホン信号とを識別すること
のできる実用性の高い音声認識装置を提供するこ
とにある。
The present invention has been made in consideration of these circumstances, and its purpose is to provide a highly practical voice recognition device that can easily and reliably distinguish between voice signals and pushphone signals. There is a particular thing.

〔発明の概要〕[Summary of the invention]

本発明は、音声信号を認識する専用のアルゴリ
ズムを備えた音声信号認識手段とプツシユホン信
号を認識する専用のアルゴリズムを備えたプツシ
ユホン信号認識手段に対して、プツシユホン信号
のみを認識処理対象とし入力信号がプツシユホン
信号か否かを判定する判定手段を設け、この判定
手段での判定結果に応じてプツシユホン信号認識
手段での認識結果または音声信号認識手段での認
識結果について所定の処理を行うようにしたもの
である。
The present invention provides an audio signal recognition means equipped with an algorithm dedicated to recognizing audio signals and a pushphone signal recognition means equipped with an algorithm dedicated to recognizing pushphone signals. A determination means for determining whether or not it is a pushphone signal is provided, and a predetermined process is performed on the recognition result by the pushphone signal recognition means or the recognition result by the voice signal recognition means in accordance with the determination result by the determination means. It is.

〔発明の効果〕〔Effect of the invention〕

従つて本発明によれば、所定の時間軸フレーム
においてトーンが安定であると云うプツシユホン
信号特有の特徴を利用して入力信号を識別したの
ち、この識別結果に従つて音声信号およびプツシ
ユホン信号をそれぞれ別個に認識処理するので、
その認識精度は非常に高いものとなる。しかも処
理形式が簡単であり、装置構成も簡易であるか
ら、容易にその信頼性の向上を図ることができ、
実用的利点が多大である。
Therefore, according to the present invention, the input signal is identified by utilizing the unique feature of the pushphone signal that the tone is stable in a predetermined time axis frame, and then the audio signal and the pushphone signal are respectively differentiated according to the identification results. Since recognition processing is performed separately,
The recognition accuracy is extremely high. Moreover, since the processing format is simple and the device configuration is simple, the reliability can be easily improved.
The practical advantages are enormous.

〔発明の実施例〕[Embodiments of the invention]

以下、図面を参照して本発明の一実施例につき
説明する。
Hereinafter, one embodiment of the present invention will be described with reference to the drawings.

第2図は実施例装置の概略構成図であり、第1
図に示す従来装置と同一構成部分には同一符号を
付して示してある。この実施例装置が特徴とする
ところは、判定部4にて入力信号をプツシユホン
信号の音響特徴辞書を用いて類似度計算処理し、
これによつて上記入力信号がプツシユホン信号で
あるか否かを判定するようにしたところにある。
上記類似度計算処理は、入力信号に対して所定の
時間軸フレーム毎に行われる。そして、類似度値
に基づく判定は、例えば音声信号/プツシユホン
信号の2値として、あるいはこれに判定不能なる
信号を加えた3値によつて行われる。このような
判定結果が認識部1,2および総合判定部3に送
られる。
FIG. 2 is a schematic configuration diagram of the embodiment device, and the first
Components that are the same as those of the conventional device shown in the figure are designated by the same reference numerals. The feature of this embodiment device is that the determination unit 4 processes the input signal by calculating the similarity using the acoustic feature dictionary of the pushphone signal.
Based on this, it is determined whether the input signal is a pushphone signal or not.
The above similarity calculation process is performed on the input signal every predetermined time axis frame. The determination based on the similarity value is performed, for example, as a binary value of an audio signal/pushphone signal, or as a ternary value obtained by adding an undeterminable signal to the binary value. Such determination results are sent to the recognition units 1 and 2 and the comprehensive determination unit 3.

判定結果がプツシユホン信号であるとして識別
したとき、その判定信号によつてプツシユホン信
号認識部1が駆動されて入力信号の認識が行われ
る。そしてその認識結果は総合判定部3を介して
出力される。また判定結果が音声信号であるとし
て識別したとき、その判定信号によつて音声認識
部2が駆動される。これにより入力信号は音声認
識され、その認識結果が総合判定部3を介して出
力されることになる。そして判定不能なる判定結
果が得られた場合には、認識部1,2がそれぞれ
駆動され、その各々において認識結果が求められ
る。このとき総合判定部3は所定のアルゴリズム
に従つて上記両認識結果を総合判定し、その判定
結果を入力信号に対する最終的な認識結果として
出力することになる。
When the judgment result is that the pushphone signal is identified, the pushphone signal recognition section 1 is driven by the judgment signal to recognize the input signal. Then, the recognition result is outputted via the comprehensive determination section 3. Further, when the determination result is identified as a voice signal, the voice recognition unit 2 is driven by the determination signal. As a result, the input signal is voice recognized, and the recognition result is outputted via the comprehensive determination section 3. If an undeterminable determination result is obtained, the recognition units 1 and 2 are each driven, and a recognition result is determined in each of them. At this time, the comprehensive judgment section 3 performs a comprehensive judgment on both of the above recognition results according to a predetermined algorithm, and outputs the judgment result as the final recognition result for the input signal.

かくして上記の如く構成された装置によれば明
らかに性質の異なる音声信号とプツシユホン信号
とを簡易に且つ精度良く識別したのち、その各々
の場合に応じて適切なアルゴリズムに従つて信号
認識することができる。これ故、従来非常に複雑
であつた音声信号およびプツシユホン信号に対す
る認識処理プロセスを系統別に分けることによつ
て、簡易にすることができ、またその認識精度の
向上を図ることができる。つまり簡易に装置の高
性能化を図ることが可能となる。
Thus, with the device configured as described above, it is possible to easily and accurately distinguish between an audio signal and a pushphone signal, which have clearly different properties, and then recognize the signal according to an appropriate algorithm for each case. can. Therefore, by dividing the recognition processing process for voice signals and pushphone signals, which has conventionally been extremely complicated, into different systems, it is possible to simplify the recognition process and improve the recognition accuracy. In other words, it is possible to easily improve the performance of the device.

ところで、前記の如く入力信号の識別を行う判
定部4は、例えば第3図に示す如く構成すること
ができる。即ち、入力信号を前処理部11に導び
き、例えば数10msecの所定時間軸フレーム毎に
上記入力信号を分析し、例えばそのバンドパスフ
イルタ出力Aと、低域または全帯域フイルタ出力
Bとを得る。上記バンドパスフイルタ出力Aを類
似度計算部12に導びき、プツシユホン信号・音
響信号特辞辞書13に格納されたプツシユホン信
号のカテゴリ毎の特徴データとの類似度計算を行
わしめる。
By the way, the determination section 4 that identifies input signals as described above can be configured as shown in FIG. 3, for example. That is, the input signal is led to the preprocessing unit 11, and the input signal is analyzed every predetermined time axis frame of, for example, several tens of milliseconds, to obtain, for example, the bandpass filter output A and the low-frequency or full-band filter output B. . The bandpass filter output A is led to the similarity calculation section 12, and similarity calculation is performed with feature data for each category of the pushphone signal stored in the pushphone signal/acoustic signal dictionary 13.

一方、分析区間決定部14では前記全帯域フイ
ルタ出力Bを用い、例えばその信号レベルの大な
る区間を検出する等して分析処理区間を求めてい
る。そして、その分析開始点と分析終了点におい
て計数処理部15に制御信号を与えている。この
計数処理部15は、上記の如く設定される区間内
において、前記類似度計算部12が所定値θs以上
の類似度値を得る回数を計数するものである。こ
の所定値θsを越える類似度値の判定は、全てのカ
テゴリについて行われる。そして、この計数され
た回数kは、前記分析区間の情報lと共に認識判
定部16に与えられるようになつている。
On the other hand, the analysis section determining section 14 uses the full-band filter output B to determine an analysis processing section, for example, by detecting sections where the signal level is high. A control signal is given to the counting processing section 15 at the analysis start point and analysis end point. This counting processing unit 15 counts the number of times the similarity calculation unit 12 obtains a similarity value equal to or greater than a predetermined value θs within the interval set as described above. This determination of similarity values exceeding the predetermined value θs is performed for all categories. The counted number of times k is then given to the recognition determining section 16 together with the information l of the analysis section.

認識判定部16は、上記分析区間lの値に応じ
て2つの閾値θks(i)、θkw(i)を持つており、これ
らの閾値と前記計数値kとを比較して入力信号の
判定を行つている。但し、iは1、2、3…lな
る値をとる。そして、 kθks(l)…(1) θks(l)>kθkw(l) …(2) θkw(l)>k …(3) なる3通りの判定を行い、上記条件が(1)なる場合
にはこれを入力信号がプツシユホン信号であると
判定結果を得ている。また上記条件が(2)なる場合
には入力信号の判定が不能であり、また条件が(3)
なる場合には前記入力信号が音声信号であるとの
判定結果をそれぞれ得ている。
The recognition determination unit 16 has two thresholds θks(i) and θkw(i) depending on the value of the analysis interval l, and compares these thresholds with the count value k to determine the input signal. I'm going. However, i takes a value of 1, 2, 3...l. Then, we make three judgments: kθks(l)…(1) θks(l)>kθkw(l)…(2) θkw(l)>k…(3) If the above condition becomes (1), then has determined that the input signal is a pushphone signal. In addition, if the above condition (2) is satisfied, it is impossible to judge the input signal, and if the condition (3) is satisfied, the input signal cannot be determined.
In each case, a determination result is obtained that the input signal is an audio signal.

このようにして求められる判定結果に応じて前
述した認識部1,2、および総合判定部3におけ
る認識・判定処理がそれぞれ行われることにな
る。
According to the determination results obtained in this manner, the recognition and determination processes described above are performed in the recognition units 1 and 2 and the comprehensive determination unit 3, respectively.

以上のように本装置によれば、認識処理の中心
となる音声認識部2に、プツシユホン信号に対す
る辞書を設けることが必要でなくなるので、従来
装置に比して処理速度の大幅な向上と、辞書分離
度の上昇による認識率の著しい改善、更には辞書
記憶領域の減少による装置構成の簡素化を図るこ
とが可能となる。またこの音声認識部2がプツシ
ユホン信号に対するリジエクト能力が低い場合で
も、プツシユホン信号認識部1では音声信号に対
するリジエクト能力を考慮することなしに、その
処理を簡易に行い得る。つまり特徴の変動が激し
い音声信号に比べて、特徴変動の小さいプツシユ
ホン信号のみを処理対象とし得るので、極めて簡
単な構成を採用して信頼性の高いプツシユホン信
号の認識を行い得る。また判定部4の構成につい
ても、第3図に示すように簡易に実現できる。ま
た分析区間判定処理を装置の前処理結果をそのま
ま利用して、つまり判定部4として格別に前処理
部11等を設けることなしに行うことも可能であ
り、装置全体として、その構成の簡易化を図り得
る。故に、辞書処理を始めとするその他関連した
処理の簡易化を図り、処理速度の向上を図り得る
等、実用上多大なる効果が奏せられる。
As described above, according to the present device, it is no longer necessary to provide a dictionary for pushphone signals in the speech recognition unit 2, which is the center of recognition processing, so that the processing speed is greatly improved compared to conventional devices, and the dictionary It is possible to significantly improve the recognition rate by increasing the degree of separation, and to simplify the device configuration by reducing the dictionary storage area. Further, even if the voice recognition section 2 has a low reject ability for pushphone signals, the pushphone signal recognition section 1 can easily perform the processing without considering the reject ability for voice signals. In other words, compared to audio signals with large feature variations, only pushphone signals with small feature variations can be processed, so that highly reliable pushphone signal recognition can be achieved using an extremely simple configuration. Furthermore, the configuration of the determining section 4 can be easily realized as shown in FIG. In addition, it is also possible to carry out the analysis interval determination process by directly using the preprocessing results of the device, that is, without providing a special preprocessing section 11 or the like as the determination section 4, which simplifies the configuration of the device as a whole. can be achieved. Therefore, it is possible to simplify dictionary processing and other related processing, and to improve the processing speed, which brings about great practical effects.

尚、本発明は上記実施例に限定されるものでは
ない。例えばプツシユホン信号は、音響結合器等
を用いた擬似プツシユホン信号をも含むことは云
うまでもない。またプツシユホン信号に対する辞
書を全てのカテゴリに対して持つことなく、カテ
ゴリを相互にクラスタリングして少数にまとめて
辞書として与えることも有効である。また音声信
号とプツシユホン信号との特徴の分離度が比較的
大きい場合には、所定の識別性能をそのまま維持
した状態で上述した処理を行うようにしてもよ
い。このようにすれば辞書とのマツチング処理に
要する時間を短くすることができ、更に辞書とし
ての記憶領域を軽減できる等の利点が生まれる。
このように本発明は、その要旨を逸脱しない範囲
で種々変形して実施することができる。
Note that the present invention is not limited to the above embodiments. For example, it goes without saying that the pushphone signal includes a pseudo pushphone signal using an acoustic coupler or the like. Furthermore, instead of having a dictionary for pushphone signals for all categories, it is also effective to mutually cluster the categories and provide a small number of them as a dictionary. Furthermore, if the degree of separation of the features between the audio signal and the pushphone signal is relatively large, the above-described processing may be performed while maintaining the predetermined identification performance. In this way, the time required for matching with the dictionary can be shortened, and the storage area for the dictionary can also be reduced, among other advantages.
As described above, the present invention can be implemented with various modifications without departing from the gist thereof.

【図面の簡単な説明】[Brief explanation of drawings]

第1図は従来装置の一例を示す構成図、第2図
は本発明の一実施例装置の概略構成図、第3図は
実施例装置における判定部の構成図である。 1……プツシユホン信号認識部、2……音声信
号認識部、3……総合判定部、4……判定部、1
1……前処理部、12……類似度計算部、13…
…特徴辞書、14……分析区間決定部、15……
特徴区間計数部、16……認識判定部。
FIG. 1 is a configuration diagram showing an example of a conventional device, FIG. 2 is a schematic configuration diagram of an embodiment of the device of the present invention, and FIG. 3 is a configuration diagram of a determining section in the embodiment device. DESCRIPTION OF SYMBOLS 1... Pushphone signal recognition section, 2... Audio signal recognition section, 3... Comprehensive judgment section, 4... Judgment section, 1
1... Preprocessing section, 12... Similarity calculation section, 13...
...Feature dictionary, 14... Analysis interval determination section, 15...
Feature section counting section, 16... Recognition determining section.

Claims (1)

【特許請求の範囲】[Claims] 1 音声信号とプツシユホン信号が入力される電
話回線に接続される音声認識装置において、音声
信号を認識する専用のアルゴリズムを備えた音声
信号認識手段と、プツシユホン信号を認識する専
用のアルゴリズムを備えたプツシユホン信号認識
手段と、プツシユホン信号のみを認識処理対象と
し上記電話回線を介して入力される信号がプツシ
ユホン信号か否かを判定する判定手段とを具備
し、この判定手段の判定結果に応じて上記音声信
号認識手段またはプツシユホン信号認識手段の認
識結果に対する処理を行うことを特徴とする音声
認識装置。
1. In a voice recognition device connected to a telephone line into which voice signals and pushphone signals are input, a voice signal recognition means equipped with an algorithm dedicated to recognizing voice signals, and a pushphone equipped with an algorithm dedicated to recognizing pushphone signals. It is equipped with a signal recognition means and a determination means for determining whether or not a signal inputted through the telephone line is a pushphone signal by treating only the pushphone signal as a recognition processing target, and according to the determination result of the determination means, the voice A speech recognition device characterized in that it processes a recognition result of a signal recognition means or a pushphone signal recognition means.
JP57112923A 1982-06-30 1982-06-30 Voice recognition equipment Granted JPS593498A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP57112923A JPS593498A (en) 1982-06-30 1982-06-30 Voice recognition equipment

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP57112923A JPS593498A (en) 1982-06-30 1982-06-30 Voice recognition equipment

Publications (2)

Publication Number Publication Date
JPS593498A JPS593498A (en) 1984-01-10
JPH0361958B2 true JPH0361958B2 (en) 1991-09-24

Family

ID=14598869

Family Applications (1)

Application Number Title Priority Date Filing Date
JP57112923A Granted JPS593498A (en) 1982-06-30 1982-06-30 Voice recognition equipment

Country Status (1)

Country Link
JP (1) JPS593498A (en)

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS5768965A (en) * 1980-10-17 1982-04-27 Fujitsu Ltd Telephone line information service system
JPS58193595A (en) * 1982-05-07 1983-11-11 株式会社日立製作所 Telephone information input unit

Also Published As

Publication number Publication date
JPS593498A (en) 1984-01-10

Similar Documents

Publication Publication Date Title
CA1246228A (en) Endpoint detector
JPS62217295A (en) Voice recognition system
JPS5972496A (en) Single sound identifier
US5101434A (en) Voice recognition using segmented time encoded speech
Beigi et al. Speaker, channel and environment change detection
CN113590873A (en) Processing method and device for white list voiceprint feature library and electronic equipment
CN111341295A (en) Offline real-time multilingual broadcast sensitive word monitoring method
JPH0361958B2 (en)
US3647978A (en) Speech recognition apparatus
EP0177854B1 (en) Keyword recognition system using template-concatenation model
CN114155840B (en) Voice initiator distinguishing method and device
Han et al. Robust speaker clustering strategies to data source variation for improved speaker diarization
JPS5953897A (en) Voice recognition equipment
JPS5953899A (en) Voice recognition equipment
JPH0119597B2 (en)
JPS5953898A (en) Voice recognition equipment
KR100262564B1 (en) Car Speech Recognition Device
JPS5952388A (en) Dictionary collating system
JPH0376471B2 (en)
JP2602271B2 (en) Consonant identification method in continuous speech
JPS60115996A (en) Voice recognition equipment
JPS59219798A (en) Voice recognition equipment
JPH01253799A (en) Recognizing method for speech
GB1602036A (en) Signal pattern encoder and classifier
JPS6070496A (en) Voice recognition processing system