JPS5961897A - Recognition equipment - Google Patents

Recognition equipment

Info

Publication number
JPS5961897A
JPS5961897A JP57172786A JP17278682A JPS5961897A JP S5961897 A JPS5961897 A JP S5961897A JP 57172786 A JP57172786 A JP 57172786A JP 17278682 A JP17278682 A JP 17278682A JP S5961897 A JPS5961897 A JP S5961897A
Authority
JP
Japan
Prior art keywords
transition
syllable
candidate
transition matrix
string
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP57172786A
Other languages
Japanese (ja)
Other versions
JPH0652478B2 (en
Inventor
外川 文雄
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Computer Basic Technology Research Association Corp
Original Assignee
Computer Basic Technology Research Association Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Computer Basic Technology Research Association Corp filed Critical Computer Basic Technology Research Association Corp
Priority to JP57172786A priority Critical patent/JPH0652478B2/en
Publication of JPS5961897A publication Critical patent/JPS5961897A/en
Publication of JPH0652478B2 publication Critical patent/JPH0652478B2/en
Anticipated expiration legal-status Critical
Expired - Lifetime legal-status Critical Current

Links

Abstract

(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。
(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.

Description

【発明の詳細な説明】 く技術分野〉 本発明は認識装置の改良に関し、更に詳細には例えば文
節等の一区切りの音声等の一区切りの認識すべき情報を
音韻、かな、音節2文節等のより細分化された単位要素
で認識する認識装置の改良に関するものである。
[Detailed Description of the Invention] [Technical Field] The present invention relates to an improvement in a recognition device, and more specifically, the present invention relates to an improvement in a recognition device, and more specifically, for example, a recognition device that recognizes a section of information such as a section of speech such as a phrase, etc. This invention relates to the improvement of a recognition device that recognizes subdivided unit elements.

〈従来技術〉 文節等の一区切りの音声等を音韻、かな、音節等のより
細分化された単位で認識する場合、従来一般的には入力
された認識すべき一区切りの音声情報等を例えば音響処
理して音韻、音節等の単位毎の特徴ベクトル入カバター
ンを得ると共に、この入カバターンと予め記憶されてい
る標準パターンとのマツチングを行って、入力された情
報を候補単位列として類似度の高いものから出力し、こ
の出力された候補単位列と文節等の辞書の内容とを照合
して入力された情報に対する文節等の一区切りの情報を
認識している。
<Prior art> When recognizing one segment of speech, such as a phrase, in more subdivided units such as phonemes, kana, syllables, etc., conventionally, the input speech information of one segment to be recognized is generally processed through acoustic processing, for example. Then, a feature vector input cover pattern is obtained for each unit such as phoneme, syllable, etc., and this input cover pattern is matched with a pre-stored standard pattern, and the input information is used as a candidate unit sequence to select a pattern with high similarity. The output candidate unit string is compared with the contents of a dictionary of phrases, etc., to recognize one section of information, such as a phrase, for the input information.

しかし、このような従来の方法によれば、全ての音韻、
音節等の標準パターンと入カバターンとのマツチングを
行なって類似度を算出し、類似度と の高いものから順に候補音節i(て出力している。
However, according to such conventional methods, all phonemes,
A standard pattern such as a syllable and an input cover pattern are matched to calculate the degree of similarity, and candidate syllables i are output in order of similarity.

したがって1例えば拗音を含む単音節単位で認識する場
合、各音節単位全てについて100種以上の単音節の標
準パターンと入カバターンとの間でマツチングを行う必
要があり、その処理時間か多大なものとなっていた。
Therefore, 1. For example, when recognizing monosyllables including syllables, it is necessary to match over 100 types of monosyllable standard patterns and input kabata patterns for each syllable, which takes a lot of processing time. It had become.

また、その後に類似度の高いものから出力される候補単
位列の全てについて辞書照合処理を行なう必要があり、
その処理時間か長くなり、正しい文節等を認識する確度
が向上せず、結果的に全体の認識に要する処理量が膨大
なものになっていた。
In addition, after that, it is necessary to perform dictionary matching processing on all candidate unit sequences output from those with high similarity.
The processing time becomes long, the accuracy of recognizing correct phrases, etc. does not improve, and as a result, the amount of processing required for overall recognition becomes enormous.

〈目的〉 本発明は、上記従来の欠点を除去した認識装置を提供す
ることを目的とし、正しい文節等の一区切りの認識すべ
き情報を認識する確度を向上させると共に、異なる話題
2分野等の異なる種類の認識すべき情報に応じた処理を
指定することが出来、結果的に全体の認識に要する処理
量を減少させることのできる認識装置を提供するもので
ある。
<Purpose> The present invention aims to provide a recognition device that eliminates the above-mentioned conventional drawbacks, and improves the accuracy of recognizing a section of information such as a correct phrase, and also improves the accuracy of recognizing a section of information such as a correct phrase. The present invention provides a recognition device that can specify processing according to the type of information to be recognized, and as a result can reduce the amount of processing required for overall recognition.

〈実施例〉 以下、本発明の認識装置を文節等の一区切りの音声入力
を音節等のより細分化された単位要素で認識する場合の
例を実施例として説明する。
<Example> Hereinafter, an example in which the recognition device of the present invention recognizes one segment of audio input such as a phrase using more subdivided unit elements such as syllables will be described as an example.

本発明の実施例によれば、文節等の一区切りの音声等の
認識すべき情報を音韻、かな、音節等のより細分化され
たN個の単位要素で認識する認識装置において、認識対
象となる文節あるいは文章等の文字(単位要素)列につ
いて、その話題あるいは分野毎に、(N+1)個の文字
(単位要素)間の接続関係である遷移関係を記述した異
なる遷移行列を複数種類記憶した遷移行列記憶手段と、
この遷移行列記憶手段に記憶された複数種類の遷移行列
より、認識すべき文字列の内容(種類1分野等)に応じ
て所望の遷移行列を指定する遷移行列指定手段と、この
遷移行列記憶手段番こより指定された遷移行列にもとす
いて、音節(単位要素〕ラティス生成時に、−音節(単
位要素)前のと゛のイ侯補音節(単位要素)からも遷移
しなし)音節(単位要素)群は認識対照から除外し、ま
た(ま及び候ン甫列作成時に各候補列に対して遷移行列
を参照し、遷移しない音節c単位要素)の組合せを含む
候7m列は除外する等の認識処理を行う処理手段とを備
えて、次の高次の辞書照合の際の処理量の削減を図るよ
うに構成されている。
According to an embodiment of the present invention, in a recognition device that recognizes information to be recognized, such as a segment of speech such as a phrase, using N unit elements that are further divided into phonemes, kana, syllables, etc., the recognition target A transition that stores multiple types of different transition matrices that describe transition relations, which are connections between (N+1) characters (unit elements), for character (unit element) strings such as clauses or sentences, for each topic or field. matrix storage means;
A transition matrix specifying means for specifying a desired transition matrix according to the content of the character string to be recognized (type 1 field, etc.) from among the plurality of types of transition matrices stored in the transition matrix storage means, and the transition matrix storage means When the syllable (unit element) lattice is generated based on the transition matrix specified from the number, there is no transition from the i-complementary syllable (unit element) of the previous - syllable (unit element). Groups are excluded from the recognition control, and (when creating a matrix and a candidate column, the transition matrix is referred to for each candidate column, and candidate 7m columns that include a combination of syllable c unit elements that do not transition) are excluded. It is configured to include processing means for performing processing, and to reduce the amount of processing in the next high-order dictionary matching.

まず、本発明の詳細な説明に先立ち、本発明の認識装置
に用いられる単位要素間の接続関係である遷移関係を示
した遷移付列番こつし1て説明する。
First, prior to a detailed explanation of the present invention, a series number with transition 1 indicating a transition relationship, which is a connection relationship between unit elements used in the recognition device of the present invention, will be explained.

一般に日本語文章は、全てかな文字で表現した場合、か
な文字列に対応しソこ音節列で表現できる。
In general, when a Japanese sentence is expressed entirely in kana characters, it can be expressed as a string of syllables corresponding to the kana character string.

例えば文節「地球の」は“ち゛“きゆ”“う゛“°の゛
という4個の単音節といわれる単位要素力)ら成り立っ
ている。2つの音節間の接続関係(“ち力)ら “きゅ
、″きゆ“から “う、“′う゛から ゛の゛)を、日
本語全て、あるいは特定の分野2話題番こおける文章等
について調べると接続(遷移;以下遷移ということばを
使う)しない音節対かある。例えばば行の音節の前には
1“ん++ 、 Itつ”5以外はこない。またパにや
′は語頭にはこないし、“へパ(へと発声するもの)は
語尾にこない。
For example, the phrase ``Earth's'' is made up of four monosyllable unit elements called ``chi゛``kiyu'' and ``u゛''゛゛.The connection relationship between the two syllables (``chiriki'') kyu, from ``kiyu'' to ``u, ``u゛ to ゛ no ゛) can be connected if you look up sentences in all Japanese or two topics in a specific field (transition; the word transition will be used below) There are some syllable pairs that do not come before the syllables of the line. For example, nothing other than 1 "n++, Ittsu" 5 comes before the syllable in the line. Also, paniya' does not come at the beginning of a word, and "hepa (something pronounced as he) does not come at the end of the word.

このような文節を構成する音節の1次の遷移関係を以下
に示す式(1)に従って記述して、第1図に示すような
遷移行列M(X、Y)を作成する。
The first-order transition relationship of the syllables constituting such a bunsetsu is described according to the following equation (1), and a transition matrix M(X, Y) as shown in FIG. 1 is created.

第1図において遷移行列M(M、Y)は単位要素列であ
る文字列の文字Xから次の文字Yへの遷移を記述したも
のであり、単位要素(音節)がN個の場合、(N+])
x(N+1)の行列であり、ノ1−ド的にはROM等に
記憶される。また70列には各単位要素(1〜N)が節
類に来るか否かを表わし、X0行には各単位要素(1〜
N)が節尾に来るか否かを表わすデータが書込まれる。
In FIG. 1, the transition matrix M(M, Y) describes the transition from character X to the next character Y in a character string, which is a unit element string, and when there are N unit elements (syllables), ( N+])
It is a matrix of x(N+1), and is stored in a ROM or the like in terms of a node. In addition, column 70 indicates whether each unit element (1 to N) comes in a clause, and row X0 indicates each unit element (1 to N).
Data indicating whether or not N) comes at the end of the knot is written.

例えば°“赤い゛という文字列の遷移を遷移行列に書込
んだ例を第2図に示す。遷移行列の要素は0(遷移不可
能)か1(遷移可能)の2値のとちらかで表現され、1
ビツトで記憶される。なお、第2図においては表記++
 + ++以外の行列要素は全て“0′”であり、その
表示を省略している。
For example, Figure 2 shows an example in which the transition of the character string °“red゛” is written in a transition matrix.The elements of the transition matrix can be either 0 (transition not possible) or 1 (transition possible). expressed, 1
It is stored in bits. In addition, in Figure 2, the notation ++
All matrix elements other than +++ are "0'" and their display is omitted.

次に遷移行列の作成について、今少し詳細に説明する。Next, the creation of the transition matrix will be explained in a little more detail.

まず遷移行列の作成にあたって遷移行列メモリを0”に
初期セットCMCX、Y)−〇]する。
First, when creating a transition matrix, the transition matrix memory is initialized to 0'' (CMCX,Y)-0].

次に文字列バー(al 、a2 +aB +・・・、a
l)但し、■=列の文字数 とした場合、次式(1) に従って、文字列A\の文字遷移関係を遷移行列M(X
 、 Y)に書込む。同様に認識対象となる文字列の全
てについて遷移関係を書込む遷移行列(1次〕の作成を
完了する。
Next, the string bar (al, a2 +aB +..., a
l) However, when ■=number of characters in a string, the character transition relationship of the character string A\ is expressed as a transition matrix M(X
, Y). Similarly, the creation of a transition matrix (primary) in which transition relationships are written for all character strings to be recognized is completed.

このようにして作成された具体的な遷移行列(1次)M
(X、Y)の例を第3図に示している。
The concrete transition matrix (first order) M created in this way
An example of (X, Y) is shown in FIG.

この第3図より明らかなように例えば(X、Y)=(え
、<)のビット位置が“′ピであるため、゛え″から“
く”への遷移が存在し、また(X 、 Y)=(え。
As is clear from FIG. 3, for example, the bit position of (X, Y) = (E, <) is "'pi," so from "E" to "
There is a transition to ``, and (X, Y) = (Eh.

け)のビット位置が0“′であるため、“え”から′け
”への遷移が存在しないことを表わしている。
Since the bit position of ke) is 0 "', it means that there is no transition from "e" to "ke".

」1記は1次の遷移であるが、2次遷移、更には一般に
M次へ拡張したM次遷移行列も同様に次式(2)に従っ
て作成することが出来る。
1 is a first-order transition, but a second-order transition, and generally an M-order transition matrix expanded to an M-order, can be similarly created according to the following equation (2).

M次遷移行列:Mcxl、x2.x3.・・・+XM+
、Y)+CN刊)M+1次元M(a + y + a 
+−CM−1) 置・+ ai) −1+ (+ =I
 −I +1 )−(2)本発明の実施例は、この遷移
行列を認識対象の種類7話題2公野等毎に複数個備え、
必要に応じて特定の遷移行列を選択して認識処理を実行
し得るようにしたものである。
M-order transition matrix: Mcxl, x2. x3. ...+XM+
, Y) + CN publication) M + 1-dimensional M (a + y + a
+-CM-1) Place・+ ai) −1+ (+ =I
-I+1)-(2) The embodiment of the present invention provides a plurality of transition matrices for each type of recognition target, 7 topics, 2 public areas, etc.
This allows recognition processing to be performed by selecting a specific transition matrix as necessary.

次に本発明の実施例を図面を参照して説明する。Next, embodiments of the present invention will be described with reference to the drawings.

第4図は本発明の一実施例装置の構成を示すブロック図
である。
FIG. 4 is a block diagram showing the configuration of an apparatus according to an embodiment of the present invention.

第4図において、lは遷移行列指定手段であり、該指定
手段1は中央処理装置(CPU)に接続されており、操
作面に設けた選択キーあるいは音声による選択人力手段
により構成される。また8は認識すべき音声情報の入力
される入力部、4は増幅部、5は音響処理部、61,6
2.・・、6にはそれぞれ異なった種類の遷移行列を記
憶する遷移行列記憶手段、7は認識処理部である。
In FIG. 4, reference numeral 1 denotes a transition matrix designation means, and the designation means 1 is connected to a central processing unit (CPU) and is constituted by a selection key provided on an operation surface or a manual selection means by voice. Further, 8 is an input section into which voice information to be recognized is input, 4 is an amplification section, 5 is an acoustic processing section, 61, 6
2. . . , 6 is a transition matrix storage means for storing different types of transition matrices, and 7 is a recognition processing section.

上記の如き構成において遷移行列記憶手段61゜62、
・・・、6Kにはそれぞれ異なる分野(例えば科学2文
学、経済等)の文章等から作成された異なる種類の遷移
行列が記憶されており、今入力部3に入力される音声情
報か例えば科学関係のものであれば、遷移行列指定手段
lを操作して科学関係の文章等から作成された遷移行列
記憶手段(例えはM + )を選択指定し、この選択指
定してメモリM1に記憶している遷移行列を用いて認識
処理部7で認識処理動作が行なわれる。
In the above configuration, transition matrix storage means 61, 62,
. . , 6K stores different types of transition matrices created from texts in different fields (for example, science, literature, economics, etc.). If it is related, operate the transition matrix specifying means 1 to select and specify a transition matrix storage means (for example, M + ) created from science-related texts, etc., and store this selection and specification in the memory M1. A recognition processing operation is performed in the recognition processing section 7 using the transition matrix.

次に」1記のようにして認識すべき情報の種類C分野)
等に応じて選択指定された遷移行列を用いた認識動作に
ついて説明する。
Next, the type of information that should be recognized as described in 1. Field C)
A recognition operation using a transition matrix selected and specified according to the following will be explained.

第5図は上記第4図に示した音響処理部5及び認識処理
部7の詳細ブロック図である。
FIG. 5 is a detailed block diagram of the acoustic processing section 5 and the recognition processing section 7 shown in FIG. 4 above.

第5図において、文節音声入力部21に入力された音声
情報は次段の音響処理・比較部22に入力される。この
音響処理・比較部22は遷移行列メモリ26を用いた処
理部分を除いて従来公知のものであり、例えば文節音声
入力部21に入力された文節音声信号が音響処理部22
により単音節毎に特徴抽出処理が行なわれ、各単音節毎
の特徴パターンが同処理部22内のバッファに一時記憶
される。一方記憶装置23には各単音節毎の標準パター
ンPH(i=I〜N)か記憶されており、この標準パタ
ーンP1が順次読出されて処理・比較部22において該
処理部内のバッファに記憶された入力音声の入力特徴パ
ターンとのマツチング計算が行なわれる。
In FIG. 5, the speech information input to the phrase speech input section 21 is input to the next stage acoustic processing/comparison section 22. This acoustic processing/comparison section 22 is of a conventionally known type except for a processing section using a transition matrix memory 26. For example, the phrase speech signal input to the phrase speech input section 21 is processed by the acoustic processing section 22.
Feature extraction processing is performed for each single syllable, and the feature pattern for each single syllable is temporarily stored in a buffer within the processing unit 22. On the other hand, the storage device 23 stores a standard pattern PH (i=I to N) for each monosyllable, and this standard pattern P1 is sequentially read out and stored in a buffer in the processing/comparison section 22. A matching calculation is performed with the input feature pattern of the input voice.

従来技術によれば、この標準パターンと人力特徴パター
ンとのマツチング計算処理は全ての標準パターンについ
て行なわれていたが、本実施例によれば、後述するよう
に遷移行列メモリ26に記憶法れた情報にもとずいて前
に候補として認識した音節に接続可能な音節(最初の場
合は先頭に来る可能性のある音節)の標準パターンとの
マツチングが計算され、最も近似したものが第1候補と
して、また順次近似したものか次候補として選出され、
その結果か候補音節メモリ24に記憶される。即ち、音
節ラティス生成時に、−音節前のどの候補音節からも遷
移しない音節群は認識対照から除外するように処理され
る。
According to the prior art, this matching calculation process between the standard pattern and the human feature pattern was performed for all standard patterns, but according to the present embodiment, as will be described later, the matching calculation process is performed in the transition matrix memory 26. Based on the information, the matching with the standard pattern of syllables that can be connected to the syllable previously recognized as a candidate (in the first case, a syllable that may come at the beginning) is calculated, and the most similar one is the first candidate. , and the successive approximations are selected as the next candidate,
The result is stored in the candidate syllable memory 24. That is, when generating a syllable lattice, syllable groups that do not transition from any candidate syllable before the -syllable are excluded from recognition comparison.

なお、遷移行列メモリ26は遷移行列指定手段1によっ
て指定された遷移行列記憶手段6]、62゜・・、6に
の一つのメモ’J(Mi)に対応したものである。
The transition matrix memory 26 corresponds to one memo 'J (Mi) in the transition matrix storage means 6], 62°, . . . , 6 designated by the transition matrix designation means 1.

上記候補音節ラティスメモリ24に記憶された複数個の
x*N音節の時系列は候補列作成部25及び遷移行列メ
モリ26より成る候補列出力部27に入力され、該候補
列出力部27において、特定の話題1分野等に対応した
遷移行列メモリ26の内容を参照して遷移不可能な音節
遷移を含む候補列は除外して、遷移可能な候補列のみ、
信頼度の4.4.>’組合せ順に作成され、この候補列
と辞書28に記憶された文節とが辞書照合部29により
照合され、一致すればその結果が文節出力部30に出力
されるように構成されている。
The time series of a plurality of x*N syllables stored in the candidate syllable lattice memory 24 is input to a candidate string output section 27 consisting of a candidate string creation section 25 and a transition matrix memory 26, and in the candidate string output section 27, By referring to the contents of the transition matrix memory 26 corresponding to a specific topic field, etc., candidate strings that include syllable transitions that cannot be transitioned are excluded, and only candidate strings that are transitionable are selected.
Reliability: 4.4. >' The candidate string is created in the order of combination, and the dictionary collation unit 29 collates this candidate string with the clauses stored in the dictionary 28, and if they match, the result is output to the clause output unit 30.

次に遷移行列M(X、Y)を用いた音節認識処理につい
て第6図に示す遷移行列を用いた候補音節作成処理ブロ
ック図を参照して説明する。
Next, syllable recognition processing using the transition matrix M(X, Y) will be explained with reference to a block diagram of candidate syllable creation processing using the transition matrix shown in FIG.

本実施例においては、結果として得る候補音節を時系列
順に候補音節ラティスバッファ24に一次記憶する。ま
た上記した遷移行列情報はメモリ26に記憶されており
、音節標準パターンはメモリ23に記憶されている。
In this embodiment, the resulting candidate syllables are temporarily stored in the candidate syllable lattice buffer 24 in chronological order. Further, the above-mentioned transition matrix information is stored in the memory 26, and the syllable standard pattern is stored in the memory 23.

候補音節ラティス24には認識結果が次表の如く記憶さ
れていくが今、第i音節を認識する場合には、以下の如
く処理が実行される。
The recognition results are stored in the candidate syllable lattice 24 as shown in the table below, and when the i-th syllable is to be recognized, the following processing is executed.

但 J(i)’第i音節候補数 Sl、:第j音節■候補音節番号 令、前音節候補を X=(si 、、j)j−1〜J(i−1)組合せ数:
 J(i−])  C1−0のとき Sl、j=o)と
した場合、次式(3)に従って直前の複数個(J(i−
1)個)の候補音節について遷移行列の和をとり、得ら
れた行m(Y)がOである音節は遷移不可能であると指
定する。
However, J(i)' Number of i-th syllable candidates Sl,: j-th syllable ■ Candidate syllable number order, previous syllable candidates X = (si ,, j) j-1 to J (i-1) Number of combinations:
J(i-]) When C1-0, when Sl, j=o, the immediately preceding multiple (J(i-)
1) The sum of the transition matrices is calculated for the candidate syllables, and the syllables whose row m(Y) obtained is O are designated as non-transitionable.

m(Y)−VM(S−Y)        −・−−−
−−・−t311  ’+3+ = M (S    Y)→−M(S    Y)+・
→i−1、I、     i −1,2。
m(Y)-VM(S-Y) ---
−−・−t311′+3+=M(SY)→−M(SY)+・
→i-1, I, i-1,2.

M(Si I、Jci−’)+ y) この(3)式においてm(Y)−oとなり、遷移不可能
と指定された音節群は 除外して、次の類似比較の処理
を行い、第i音節の候補音節を出力し、候補音節ラティ
ス7に書込む。但し、1−1(節類の音節)のときは第
0行M(0,Y)によって遷移不可能と指定された音節
群を除外して類似比較の処理を行なう。
M(Si I, Jci-') + y) In this equation (3), m(Y)-o, excluding the syllable group designated as non-transitionable, performs the next similarity comparison process, and then The candidate syllable of the i syllable is output and written in the candidate syllable lattice 7. However, in the case of 1-1 (syllables of clause class), the syllable group designated as non-transitionable by the 0th row M(0, Y) is excluded and the similarity comparison process is performed.

以上を繰返して、−文節音声の候補音節ラティスの作成
を完了する。
By repeating the above steps, the creation of the candidate syllable lattice of the -phrasal speech is completed.

今、−文節音声として「国民は」を入力した場合、音響
処理部22により音節毎に特徴抽出が行なわれ、その音
節毎の特徴パターン肩 が入カバターン時系列バッファ
31に記憶される。次に遷移行列を用いた候補音節作成
処理に移り、最初に第1音節の特徴パターンが次1が入
カバターンバッファ32に読み込まれ、次にステップn
3に移行して前候補音節群により式(3)にしたがって
遷移行列の行を指定する。最初の場合はステップn4に
おいて第0行のM(0、Y)が指定されその内容がバッ
ファ33に一時記憶され、ステップn5の生起音節の指
定が成される。
Now, when "Kokuminwa" is input as the -syllable speech, the acoustic processing unit 22 extracts features for each syllable, and the feature pattern for each syllable is stored in the input pattern time series buffer 31. Next, the process moves to candidate syllable creation using the transition matrix, and first, the feature pattern of the first syllable is read into the input cover turn buffer 32, and then step n
3, the rows of the transition matrix are specified using the previous candidate syllable group according to equation (3). In the first case, M(0, Y) in the 0th line is specified in step n4, its contents are temporarily stored in the buffer 33, and the occurring syllable is specified in step n5.

次にステップn6に移行して入カバターンバッファ32
に記憶された第1音節×1の特徴パターンかロードされ
、この特徴パターン次、と音節標塾パターンメモリ23
に記憶された標準パターンノ内バッファ33によって生
起音節と指定されて順次標準パターンバッファ34に読
出される標準パターンとの間で類似比較が行なわれ(ス
テップn7)、その結果にもとずいて候補音節が出力さ
れ(ステップn8〕、その結果か候補音節ラティス24
に書かれる。この実施例においては第1音節候補として
“KO”′、“+ G OII 、 I“BO”が記憶
される。
Next, proceeding to step n6, the input cover turn buffer 32
The characteristic pattern of the first syllable x 1 stored in
A similarity comparison is made between the standard patterns designated as occurring syllables by the standard pattern internal buffer 33 stored in the standard pattern buffer 33 and sequentially read out to the standard pattern buffer 34 (step n7), and based on the result, candidates are selected. The syllables are output (step n8), and the result is a candidate syllable lattice 24.
written in. In this embodiment, "KO"', "+G OII, and I"BO" are stored as first syllable candidates.

次にステップn2に戻り、第2音節特徴パターン×2か
バッファ32に入力され、ステップn3に移行して、候
補音節ラティス24の第1候補音節にもとずいて+IK
O!1 、 l“GO°゛、BO’”に対応した各行の
M(S、、1〜B+y)が指定され、ステップn4にお
いて、その遷移行列の和(OR)が作成されてその結果
がバッファ33に一時記憶され、ステ九プn5の生起音
節の指定が成される。
Next, the process returns to step n2, where the second syllable feature pattern x2 is input to the buffer 32, and the process proceeds to step n3, where +IK is added based on the first candidate syllable of the candidate syllable lattice 24.
O! 1, M(S,, 1 to B+y) of each row corresponding to l "GO°゛, BO'" is specified, and in step n4, the sum (OR) of the transition matrices is created and the result is stored in the buffer 33. The syllable of step n5 is specified.

次にステップn6に移行し、以下同様のステップn6〜
n9を実行して第2候補音節“’KU’”、“G U”
をメモリ24に記憶する。
Next, the process moves to step n6, and similar steps n6 to
Execute n9 to select the second candidate syllable "'KU'", "G U"
is stored in the memory 24.

以上の動作を繰返して一文節の候補音節ラティスの作成
を完了する。
By repeating the above operations, the creation of a candidate syllable lattice for one phrase is completed.

以上のようにして候補音節ラティス24に候補例が記憶
されることになるが、遷移行列を用いない場合の従来方
式の場合と木刀式の場合の実例を入力音声「国民は」に
ついて次表に示す。
Candidate examples are stored in the candidate syllable lattice 24 as described above, but examples for the conventional method and the wooden sword method when no transition matrix is used are shown in the table below for the input voice "Kokumin wa". show.

」−記の例から明らかなように1木刀式による方が正し
い文字列が候補列の上位に上がっている様子がわかる。
” - As is clear from the example below, it can be seen that the correct character strings are higher in the candidate string when using the 1-bokuto method.

以」−の遷移行列は1次遷移であるが、2次遷移、史に
は一般的なM次遷移まで同じ手法で拡張することができ
る。
Although the transition matrix ``-'' is a first-order transition, it can be extended to a second-order transition, or even a general M-order transition, using the same method.

なおM次の遷移行列の作成は上述の式(2)に従い、n
iJ候補音節(M音節前まで)からの音節指定は次に示
す式(4)によって行なうことか出来る。
Note that the M-order transition matrix is created according to the above equation (2), and n
Syllable designation from the iJ candidate syllable (up to the M syllable) can be performed using the following equation (4).

即ちM次遷移行列M(X、 、X2 、=・、’XM、
Y ) ヘの拡張の場合、前音節候補列を +x、  、x 2.− ツXM)= (Si−M、j
l  Si −(M−1)、j2 °” Si −1、
jM)jl−1〜J(i −M) j2−1−J(i−(M−1)) jM−1−J(+  1) 組合せの数:J(i−M)・J(i−CM−1))・・
・J(i−1)(l!<0のとき  S、、、−0) とした場合、 音節指定は +−M、J 1.1(M−1)、j2.・・・+ Sl
−+、jM、 Y)−(i)m(Y)=VM(S   
−S。
That is, the M-order transition matrix M(X, ,X2 ,=・,'XM,
In the case of expansion to Y ), the previous syllable candidate sequence is +x, , x 2. - TSXM) = (Si-M,j
l Si −(M−1), j2 °” Si −1,
jM) jl-1~J(i-M) j2-1-J(i-(M-1)) jM-1-J(+ 1) Number of combinations: J(i-M)・J(i- CM-1))・・
・If J(i-1) (S, ,, -0 when l!<0), the syllable specification is +-M, J 1.1(M-1), j2.・・・+ Sl
−+,jM, Y)−(i)m(Y)=VM(S
-S.

j、=l〜J(i −M) j2−1〜J (i −(M−1)) jM−1〜J(i−1) によって行なうことになる。j, =l~J(i-M) j2-1~J (i-(M-1)) jM-1~J(i-1) This will be done by.

なお、Mの次数を大きくとれは、生成音節の限定が強く
なり効果(1より大きくなる。
Note that when the order of M is increased, the syllables to be generated become more limited, and the effect becomes larger than 1.

次に上記候補列出力部27で実行されている遷移行列を
用いた候補音節列作成動作について、第7図に示す遷移
行列を用いた候補列作成の処理ブロック図を参照して説
明する。
Next, the operation of creating a candidate syllable string using a transition matrix, which is executed by the candidate string output section 27, will be explained with reference to the process block diagram of creating a candidate string using a transition matrix shown in FIG.

上記第5図に示した音響処理・比較部22から出力され
た複数個の候補音節の時系列を記憶する候補音節ラティ
スメモリ24の内容をもとに、候補音節列作成部41に
おいて信頼度の高い順に候補列が作成され、その結果か
候補音節列バッファ42に一次記憶される。この候補音
節列バッファ42に記憶された候補音節列は遷移行列参
照部43においてメモリ26に記憶された遷移行列:M
(X。
Based on the contents of the candidate syllable lattice memory 24 that stores the time series of a plurality of candidate syllables output from the acoustic processing/comparison section 22 shown in FIG. Candidate strings are created in ascending order, and the results are temporarily stored in the candidate syllable string buffer 42. The candidate syllable string stored in the candidate syllable string buffer 42 is converted to the transition matrix M stored in the memory 26 by the transition matrix reference unit 43.
(X.

Y〕を参照して、遷移可能か不可能かを次式(5)によ
って判定部44において判定し、可能な候補列のみ候補
音節列書込み部45を介して候補音節列出力バッファ4
6に記憶していく。
Y], the determination section 44 determines whether the transition is possible or not using the following equation (5), and only possible candidate strings are sent to the candidate syllable string output buffer 4 via the candidate syllable string writing section 45.
I will remember it in 6.

令弟J番目の候補音節列を バー(al+a2+・・・、al) 但し、a、:第1番目の音節番号 ■ 1列の音節数 とした場合、判定部44による遷移行列M(X 、 Y
)を用いた候補列否定は のいずれか一つが成立した場合に成される。
If the candidate syllable string of the J-th younger brother is bar (al+a2+..., al), where a,: the first syllable number ■ the number of syllables in one column, the transition matrix M(X, Y
) is used to negate a candidate sequence if any one of the following is true.

この(5)式において、いずれか一つが成立した遷移不
可能な音節列を含んだ候補音節列は除外され、次の候補
音節列について同様の判定を行ない、遷移可能な候補音
節列のみが出力バッファ46に記憶される。
In this formula (5), candidate syllable strings containing non-transitionable syllable strings in which any one of them is true are excluded, the same judgment is made for the next candidate syllable string, and only transitional candidate syllable strings are output. The data is stored in buffer 46.

今、−文節音声として「国民は」を入力した場合、音響
処理・比較部2の処理により候補音節ラティスメモリ4
に次表の如き候補音節が時系列に記憶される。
Now, when "Kokuminwa" is input as the -syllable sound, the candidate syllable lattice memory 4 is processed by the acoustic processing/comparison unit 2.
Candidate syllables as shown in the following table are stored in chronological order.

このメモリ24に記憶された音節ラティスを基に、信頼
度の高い順に候補列が作成され、遷移行列:MCX、Y
) を参照して作成された候補列が遷移可能なもののみ
か出力され、この例の場合には候補音節列か次の如く出
力される。
Based on the syllable lattice stored in the memory 24, candidate columns are created in descending order of reliability, and transition matrices: MCX, Y
), only transitionable candidate strings are output, and in this example, candidate syllable strings are output as follows.

遷移行列を参照しない従来方式によれば信頼度の最も高
い候補列としてrGOKUI) INWAJが出力され
ることになるが、本方式によれば、この候補列の音節の
遷移例えば’KU“から“PI”が遷移不1丁能である
と遷移行列:M(X、Y)を用いて判断され、以後の辞
書照合処理から除外される。
According to the conventional method that does not refer to the transition matrix, rGOKUI) INWAJ would be output as the candidate string with the highest reliability, but according to this method, the syllable transition of this candidate string, for example from 'KU' to 'PI ” is determined to be transition-invalid using the transition matrix: M(X, Y), and is excluded from subsequent dictionary matching processing.

以−にの遷移行列は1次遷移であるが、2次遷移、更に
は一般的なM次遷移まで同じ手法で拡張することができ
る。
Although the transition matrix described above is a first-order transition, it can be extended to a second-order transition or even a general M-order transition using the same method.

なおM次の遷移行列の作成は上述の式(2)に従い、候
補音節列の否定は次に示す式(6)によって行うことが
出来る。
Note that the M-order transition matrix can be created according to the above equation (2), and the candidate syllable string can be negated using the following equation (6).

即ち、M次遷移行列’M(XI + X2 +”’+ 
xMI Y )への拡張の場合、第j候補列をA j 
−(a 1 + a 2 +・・・。
That is, the M-order transition matrix 'M(XI + X2 +'''+
xMI Y ), the j-th candidate column is A j
-(a 1 + a 2 +...

al)とすると M(a;1.aH(Ml)、、、、al)=OC1川〜
I + 1 )−113+(但し l≦0.l>1のと
きa i −0)のいずれか一つが成立した場合に否定
が成される。
al) then M(a; 1.aH(Ml), ,, al) = OC1 river ~
Negation is performed when any one of I + 1 )-113+ (a i -0 when l≦0.l>1) holds true.

なお、Mの次数を大きくとれば、候補音節列の限定が強
くなり、効果はより大きくなる。
Note that if the degree of M is increased, the candidate syllable string becomes more limited, and the effect becomes greater.

秩フのようにして、候補列作成時に、各候補列に対して
行列Mを参照し、遷移しない音節の組合せを含む候補列
は除外されることになる。
As in Chichifu, when creating a candidate string, the matrix M is referred to for each candidate string, and candidate strings that include combinations of syllables that do not transition are excluded.

上記した認識装置の認識対象は文節に限らず、音節、単
語1文章でもよく、また細分化された単位は音節に限ら
ず、音韻、単語でもよい。
The recognition target of the recognition device described above is not limited to phrases, but may also be syllables or single-word sentences, and the subdivided units are not limited to syllables, but may also be phonemes or words.

またアルファベット等の文字列あるいはFORTRAN
言語等のプログラム言語の文字列でもよい。
Also, character strings such as alphabets or FORTRAN
It may be a character string of a programming language such as a language.

一般に認識対象語を構成する細分化した単位の遷移関係
の存在する文字列であれば、本発明を適用することが出
来る。
In general, the present invention can be applied to any character string in which there is a transition relationship between subdivided units constituting a recognition target word.

く効果〉 以上の如く、本発明によれば、確度高く正しい候補列を
抽出することが出来るため、正しい文節等を認識する確
度が高くなり、結果的に高次の辞書照合等の処理量を減
少させることが出来ると共に、認識すべき情報の種類、
内容9話題2公野等に応じて、その都度必要に応じて話
題9分野別等の遷移行列を任意に選択指定して用いるこ
とが出来るため、遷移行列を用いた認識処理の効果をよ
り大きくすることが可能である。
As described above, according to the present invention, it is possible to extract correct candidate sequences with high accuracy, so the accuracy of recognizing correct phrases, etc. is increased, and as a result, the amount of processing such as high-level dictionary matching can be reduced. types of information that can be reduced and should be recognized;
Depending on the content, 9 topics, 2 public fields, etc., it is possible to arbitrarily select and use transition matrices for each topic, 9 fields, etc., as necessary each time, which increases the effectiveness of recognition processing using transition matrices. It is possible to do so.

なお、本発明において、話題毎の文章や文節について作
成したような同次数の異なる種類の遷移行列;M、、M
、から、それ等の和をとって合成することにより、簡単
に新しい遷移行列;M(M=MiUMj  )を作成す
ることが出来る。
In addition, in the present invention, transition matrices of different types with the same degree, such as those created for sentences and clauses for each topic;
, a new transition matrix; M (M=MiUMj) can be easily created by summing and composing them from .

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は1次遷移行列を示す図、第2図は文字列の遷移
を書込んだ遷移行列例を示す図、第3図は文節文字列の
遷移行列例を示す図、第4図は本発明を実施した認識装
置の一実施例の構成を示すブロック図、第5図は遷移行
列を用いた認識処理部の詳細ブロック図、第6図は遷移
行列を用いた候補音節作成の処理フロー図、第7図は遷
移行列を用いた候補列作成の処理ブロック図である。 1・・遷移行列指定手段、2・・・中央処理装置(CP
U)、61 、62 、・・・、6K・・・遷移行列記
憶手段、7・・・認識処理部。 師(財) 第1図 話尾 話頭 ■ 0″″″                基、第3図
Figure 1 shows a linear transition matrix, Figure 2 shows an example of a transition matrix in which character string transitions are written, Figure 3 shows an example of a transition matrix for bunsetsu character strings, and Figure 4 shows an example of a transition matrix in which character string transitions are written. A block diagram showing the configuration of an embodiment of a recognition device implementing the present invention, FIG. 5 is a detailed block diagram of a recognition processing unit using a transition matrix, and FIG. 6 is a processing flow for creating candidate syllables using a transition matrix. 7 are processing block diagrams for creating candidate columns using a transition matrix. 1...Transition matrix designation means, 2...Central processing unit (CP)
U), 61 , 62 , . . . , 6K . . . transition matrix storage means, 7 . . . recognition processing unit. Master (Treasury) Figure 1 End of story Beginning of story ■ 0″″″ Base, Figure 3

Claims (1)

【特許請求の範囲】 1、一区切りの認識すべき情報をより細分化されたN個
の単位要素で認識する認識装置において、認識すべき所
定の単位要素列について(N+1)個の単位要素間の接
続関係である遷移関係を記述した異なる遷移行列を複数
種類記憶した遷移行列記憶手段と、 上記遷移行列記憶手段に記憶された複数種類の異なる遷
移行列の所定の遷移行列を指定する遷移行列指定手段と
、 上記遷移行列指定手段により指定された遷移行列にもと
ずいて認識処理する処理手段と、を備えたことを特徴と
する認識装置。 2 一区切りの認識すべき情報は単語あるいは文節単位
の音声情報であり、単位要素列は単語あるいは文節単位
の文字列であるところの特許請求の範囲第1項記載の認
識装置。 3、複数種類の異なる遷移行列は、それぞれ異なる分野
の文章から作成された複数個の遷移行列であるところの
特許請求の範囲第1項記載の認識装置。
[Scope of Claims] 1. In a recognition device that recognizes one section of information to be recognized using N unit elements that are further subdivided, the information between (N+1) unit elements for a predetermined unit element string to be recognized is Transition matrix storage means for storing a plurality of different transition matrices that describe transition relationships that are connection relationships; and transition matrix designation means for specifying a predetermined transition matrix among the plurality of different transition matrices stored in the transition matrix storage means. and processing means for performing recognition processing based on the transition matrix designated by the transition matrix designation means. 2. The recognition device according to claim 1, wherein the information to be recognized in one section is speech information in units of words or phrases, and the unit element string is character strings in units of words or phrases. 3. The recognition device according to claim 1, wherein the plurality of different transition matrices are a plurality of transition matrices created from texts in different fields.
JP57172786A 1982-09-30 1982-09-30 Recognition device Expired - Lifetime JPH0652478B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP57172786A JPH0652478B2 (en) 1982-09-30 1982-09-30 Recognition device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP57172786A JPH0652478B2 (en) 1982-09-30 1982-09-30 Recognition device

Publications (2)

Publication Number Publication Date
JPS5961897A true JPS5961897A (en) 1984-04-09
JPH0652478B2 JPH0652478B2 (en) 1994-07-06

Family

ID=15948322

Family Applications (1)

Application Number Title Priority Date Filing Date
JP57172786A Expired - Lifetime JPH0652478B2 (en) 1982-09-30 1982-09-30 Recognition device

Country Status (1)

Country Link
JP (1) JPH0652478B2 (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS5991499A (en) * 1982-11-18 1984-05-26 伊福部 達 Voice recognition system

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS5629299A (en) * 1979-07-16 1981-03-24 Western Electric Co Voice identifier

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS5629299A (en) * 1979-07-16 1981-03-24 Western Electric Co Voice identifier

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS5991499A (en) * 1982-11-18 1984-05-26 伊福部 達 Voice recognition system

Also Published As

Publication number Publication date
JPH0652478B2 (en) 1994-07-06

Similar Documents

Publication Publication Date Title
Lee Voice dictation of mandarin chinese
JPH0855122A (en) Context tagger
Lee et al. Golden Mandarin (I)-A real-time Mandarin speech dictation machine for Chinese language with very large vocabulary
US20080270115A1 (en) System and method for diacritization of text
Laporte Rational transductions for phonetic conversion and phonology
JP4738847B2 (en) Data retrieval apparatus and method
JP2002278579A (en) Voice data search device
JP2974621B2 (en) Speech recognition word dictionary creation device and continuous speech recognition device
Dolatian et al. Reduplication with finite-state technology
Chowdhury et al. Bangla grapheme to phoneme conversion using conditional random fields
JPH0552506B2 (en)
Hoste et al. Meta-learning for phonemic annotation of corpora
JPH0552507B2 (en)
JPS5855995A (en) Voice recognition system
Sarikaya et al. Maximum entropy modeling for diacritization of Arabic text.
JPS6342279B2 (en)
JPH0652478B2 (en) Recognition device
JPS61122781A (en) Speech word processor
Rabiner Speech recognition based on pattern recognition approaches
JPS58123126A (en) Dictionary retrieving device
Gorman et al. Rewrite rules
JPH02308194A (en) Foreign language learning device
JPS60164800A (en) Voice recognition equipment
Quan et al. A robust method for the Vietnamese handwritten and speech recognition
Hoste et al. A rule induction approach to modeling regional pronunciation variation