JPH02193280A - Character recognizing device - Google Patents

Character recognizing device

Info

Publication number
JPH02193280A
JPH02193280A JP1012720A JP1272089A JPH02193280A JP H02193280 A JPH02193280 A JP H02193280A JP 1012720 A JP1012720 A JP 1012720A JP 1272089 A JP1272089 A JP 1272089A JP H02193280 A JPH02193280 A JP H02193280A
Authority
JP
Japan
Prior art keywords
character
characters
calculation accuracy
calculation
candidate
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP1012720A
Other languages
Japanese (ja)
Inventor
Naoki Maeda
直樹 前田
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sumitomo Electric Industries Ltd
Original Assignee
Sumitomo Electric Industries Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sumitomo Electric Industries Ltd filed Critical Sumitomo Electric Industries Ltd
Priority to JP1012720A priority Critical patent/JPH02193280A/en
Publication of JPH02193280A publication Critical patent/JPH02193280A/en
Pending legal-status Critical Current

Links

Landscapes

  • Character Discrimination (AREA)

Abstract

(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。
(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.

Description

【発明の詳細な説明】 〈産業上の利用分野〉 本発明は文字認識装置に関し、さらに詳細にいえば、画
像入力装置、またはファクシミリ等の通信媒体を通して
文字、記号等(以下代表して「文字」という)を表わす
画像信号を取得し、その特徴量を抽出し、上記特徴量を
基に演算を行って候補文字をひとまず選定し、上記候補
文字の中から被読取対象である文字に最も近い文字を詳
細識別して当該文字を表わす信号を出力することのでき
る文字認識装置に関するものである。
[Detailed Description of the Invention] <Industrial Field of Application> The present invention relates to a character recognition device, and more specifically, the present invention relates to a character recognition device, and more specifically, to recognize characters, symbols, etc. (hereinafter representatively referred to as “characters”) through an image input device or a communication medium such as a facsimile. ”), extract its feature amount, perform calculations based on the feature amount, select a candidate character, and select the candidate character that is closest to the character to be read from among the candidate characters. The present invention relates to a character recognition device that can identify characters in detail and output signals representing the characters.

〈従来の技術〉 従来の文字認識装置による文字認識は、入力された文字
信号から入力文字の特徴を抽出し、自ら持っている認識
用辞書に蓄えられている特徴との相違度を比較し、相違
度の最も小さな文字を選択することにより行う。
<Conventional technology> Character recognition using conventional character recognition devices extracts the characteristics of input characters from input character signals, compares the degree of difference with the characteristics stored in the recognition dictionary owned by the user, This is done by selecting the character with the smallest degree of difference.

例えば、第4図に示すように、スキャナ等の画像入力手
段(11)で文字を含む画像を入力し、文字切り出し手
段(12)で1つ1つの単位文字を切り出す。そして、
特徴抽出手段(13)で、切り出された文字画像に基づ
いて特徴量を抽出する。詳細識別手段(15)は、この
特徴量と、認識用辞書(1B) (認識用辞書(16)
は、特徴量の外、例えば特徴量の平均値や分布の状態、
各特徴量が認識に影響を与える順位等を記憶している。
For example, as shown in FIG. 4, an image containing characters is input using an image input means (11) such as a scanner, and each unit character is cut out using a character cutting means (12). and,
A feature extraction means (13) extracts feature amounts based on the cut out character image. The detailed identification means (15) uses this feature amount and the recognition dictionary (1B) (recognition dictionary (16)
In addition to the features, for example, the average value of the features, the state of the distribution,
The ranking of each feature amount influencing recognition is stored.

)に蓄えられた文字数分の特徴量に関する情報とを逐一
比較し、相違度あるいは類似度(以下代表して「相違度
」という)をそれぞれ計算しその結果を出力する。順位
決定手段(17)は、相違度の小さな順に並べ変え、例
えば最も小さな相違度を有する文字を認識結果として出
力する。
) is compared point by point with the information on the feature quantities for the number of characters stored in , the degree of dissimilarity or degree of similarity (hereinafter representatively referred to as "degree of dissimilarity") is calculated, and the results are output. The ranking determining means (17) rearranges the characters in descending order of dissimilarity, and outputs, for example, the character having the smallest dissimilarity as a recognition result.

また、第5図は他の従来例を示すブロック図であり、特
徴抽出手段(13)で、切り出された文字信号に基づい
て特徴量を抽出した後、大分類識別手段(14)で、簡
単な手法を用いて複数個の候補文字を絞る。詳細識別手
段(15)では大分類識別手段(14)から送られてき
た各候補文字について、認識用辞書(16)に蓄えられ
た当該候補文字の特徴量に関する情報を得、特徴抽出手
段(18)で得られた特徴量と比較し、それぞれの相違
度を計算してその結果を出力する。順位決定手段(17
)は、上記と同様の手法により相違度の小さな順に並べ
変え、例えば最も小さな相違度を出した文字を認識結果
として出力する。
FIG. 5 is a block diagram showing another conventional example. After the feature extracting means (13) extracts the feature amount based on the cut out character signal, the major classification discriminating means (14) Narrow down multiple candidate characters using a method. The detailed identification means (15) obtains information regarding the feature amount of the candidate character stored in the recognition dictionary (16) for each candidate character sent from the major classification identification means (14). ), calculate the degree of difference for each, and output the results. Ranking determining means (17
) rearranges the characters in order of decreasing degree of difference using the same method as above, and outputs, for example, the character with the smallest degree of difference as the recognition result.

特に、この第5図の文字認識装置の一具体例として、特
公昭83−28915号公報記載の文字認識装置があげ
られる。この文字認識装置では、大分類識別手段で、同
じような偏を持つカテゴリの辞書を選択し、詳細識別手
段では、選択された候補カテゴリの辞書に入っている文
字に対して偏部分はマスクし、労 部分のみを詳細識別
する。そして、同じような偏を持つ文字群の認識率を向
上させている。
In particular, a specific example of the character recognition device shown in FIG. 5 is the character recognition device described in Japanese Patent Publication No. 83-28915. In this character recognition device, the major classification identification means selects dictionaries of categories with similar biases, and the detailed identification means masks the biased parts of the characters in the dictionary of the selected candidate category. , only the labor part is identified in detail. This also improves the recognition rate for groups of characters with similar biases.

〈発明が解決しようとする課題〉 ところが、上記第4図の構成では、詳細識別手段(15
)は、認識用辞書(16)に蓄えられた全ての文字の特
徴量を比較の対象とするため、認識字種が多い文学認識
装置では計算量が大きくなり、認識時間が増大すること
になる。例えば、JIs第一水準3,000字種を認識
する漢字OCRでは、−文字を認識するのに相違度をa
、ooo回計算しなければならない。また、相違度を計
算する精度(例えば相違度を計算するのに必要な級数の
項数)を、その文字認識装置で最も認識することが困難
な文字を正続(正しく認識すること)できるように、字
に対しても同一の認識処理を行うことになる。
<Problem to be solved by the invention> However, in the configuration shown in FIG.
) compares the features of all the characters stored in the recognition dictionary (16), which increases the amount of calculations and increases the recognition time for literature recognition devices that recognize many types of characters. . For example, in Kanji OCR, which recognizes 3,000 character types at the first level of JIs, the degree of dissimilarity is a to recognize - characters.
, ooo times. In addition, the accuracy of calculating the degree of dissimilarity (for example, the number of terms in the series required to calculate the degree of dissimilarity) can be adjusted so that characters that are most difficult to recognize with the character recognition device can be concatenated (correctly recognized). In addition, the same recognition process is performed for characters.

したがって、認識が簡単な入力文字に対しても上記の精
度で計算するため、時間の無駄が生じている。
Therefore, even input characters that are easy to recognize are calculated with the above accuracy, which results in wasted time.

第5図の文字認識装置では、上記の欠点はある程度解決
されでいる。第5図の構成では、大分類識別手段(14
)で、認識しようとする文字と類似の候補文字を複数個
選定する。ただし、この場合も認識字種が多い文字認識
装置では、全ての字種について候補選定のための計算し
ければならないが、1回の計算に要する時間が少なくて
済むので、候補文字を選定するのに要する時間は全体と
して少なくなる。そして、その次に、候補文字を詳細識
別手段(15)に送り込み、その候補文字について、第
4図の文字認識装置と同様にして相違度を計算する。
In the character recognition device shown in FIG. 5, the above-mentioned drawbacks have been solved to some extent. In the configuration shown in FIG. 5, the major classification identification means (14
) to select multiple candidate characters similar to the character you want to recognize. However, in this case too, in a character recognition device that recognizes many character types, it is necessary to perform calculations for selecting candidates for all character types, but since the time required for one calculation is short, it is necessary to select candidate characters. The overall time required will be reduced. Then, the candidate characters are sent to the detailed identification means (15), and the degree of dissimilarity is calculated for the candidate characters in the same manner as the character recognition device shown in FIG.

ここで、nを文字認識装置b(認識可能な字種数、nl
を大分類識別手段(14)が出力する候補文字数、tl
を大分類文字識別に要する1文字当たりの時間、t2を
詳細分類文字識別に要する1文字当たりの時間とすると
、第4図の文字認識装置では、t2  Xn の時間がかかっていたのを、第5図の文字認識装置では
、 tl  Xn+t2  Xnl の時間がかかることとになる。もし、tlがt2より遥
かに小さく、nlがnよりも遥かに小さければ、演算時
間を短縮することができる。
Here, n is character recognition device b (number of recognizable character types, nl
The number of candidate characters output by the major classification identification means (14), tl
Assuming that t2 is the time per character required for major classification character identification and t2 is the time per character required for detailed classification character identification, the character recognition device shown in Fig. 4 takes t2 In the character recognition device shown in FIG. 5, it takes tl Xn+t2 Xnl time. If tl is much smaller than t2 and nl is much smaller than n, the calculation time can be reduced.

しかし、第5図の文字認識装置では、計算の精度を、詳
細分類文字識別が最も困難な文字に合わせておく点では
、第4図の文字認識装置と変わらないので、詳細分類文
字識別に要する1文字当たりの時間t2が短縮する訳で
はない。例えば、上記した特公昭63−26915号公
報記載の文字認識装置では、労部分の特徴量相違度の計
算をするときに、同数の次元までの計算(同公報中(1
)式のΣをに−1からnまでとる計算)を行っており、
実際の計算には時間がかかるものである。
However, the character recognition device shown in Fig. 5 is the same as the character recognition device shown in Fig. 4 in that the calculation accuracy is adjusted to the character for which detailed classification characters are most difficult to identify. This does not mean that the time t2 per character is shortened. For example, in the character recognition device described in Japanese Patent Publication No. 63-26915 mentioned above, when calculating the degree of feature difference in the labor part, calculations up to the same number of dimensions (in the same publication (1)
) is calculated by taking Σ of the formula from −1 to n.
Actual calculation takes time.

したがって、従来の文字認識装置では、演算時間の短縮
にも一定の限度が存在していたということができる。
Therefore, it can be said that in conventional character recognition devices, there is a certain limit to the reduction in calculation time.

本発明の目的は、簡単な手法を採用することにより、文
字の認識時間を全体として短縮することのできる文字認
識装置を提供することにある。
An object of the present invention is to provide a character recognition device that can shorten the overall character recognition time by employing a simple method.

く課題を解決するための手段〉 上記の目的を達成するための本発明の文字認識装置は、
第1図に示すように、文字を含む被読取対象を表わす画
像信号を取得する画像信号取得手段(1)と、上記画像
信号に基づき画像中の認識しようとする文字の特徴量を
抽出する特徴量抽出手段(2)と、上記特徴量を基に演
算を行い1つ以上の候補文字を選定する大分類識別手段
(3)と、各文字の認識に必要な特徴量に関する情報を
記憶した認識用辞書(6)と、上記認識しようとする文
字の特徴量を認識用辞書(6)に記憶された候補文字の
特徴量その他の情報と比較し特徴量比較計算を行う詳細
識別手段(4)と、詳細識別手段(4)で識別された文
字の中から一定の基準で文字を選択して当該文字を表わ
す信号を出力する認識文字出力手段(5)と、各文字に
対応して、詳細識別手段(4)が特徴量比較計算をする
時の計算精度を指定する定数を記憶している計算精度対
応表(7)と、上記候補文字の中から所定の基準で1つ
または複数の候補文字を選びだし、この選び出した候補
文字に対応する計算精度を計算精度対応表から検索し、
検索した計算精度を基に1つの計算精度を決定して詳細
識別手段に送り出す計算精度決定手段(8)とを具備し
ている。そして、上記詳細識別手段(4)は、計算精度
決定手段(8)から指定される計算精度で特徴量比較計
算をする機能を有するものである。
Means for Solving the Problems> The character recognition device of the present invention for achieving the above objects has the following features:
As shown in FIG. 1, an image signal acquisition means (1) that acquires an image signal representing an object to be read including characters, and a feature that extracts the feature amount of the character to be recognized in the image based on the image signal. a quantity extraction means (2), a major classification identification means (3) that performs calculations based on the feature quantities and selects one or more candidate characters, and a recognition device that stores information regarding the feature quantities necessary for recognizing each character. detailed identification means (4) that compares the feature amounts of the character to be recognized with the feature amounts and other information of the candidate characters stored in the recognition dictionary (6) and performs a feature amount comparison calculation; , a recognized character output means (5) which selects a character according to a certain standard from among the characters identified by the detailed identification means (4) and outputs a signal representing the character; A calculation accuracy correspondence table (7) that stores constants specifying calculation accuracy when the identification means (4) performs feature value comparison calculations, and one or more candidates from the above candidate characters based on a predetermined standard. Select a character, search the calculation accuracy corresponding to the selected candidate character from the calculation accuracy correspondence table,
The calculation accuracy determining means (8) determines one calculation accuracy based on the retrieved calculation accuracy and sends it to the detailed identification means. The detailed identification means (4) has a function of performing feature value comparison calculation with the calculation accuracy specified by the calculation accuracy determination means (8).

く作用〉 上記の構成の文字認識装置によれば、画像信号取得手段
(1)により取り込まれた画像信号に含まれる文字信号
に対して、特徴量抽出手段(2)によって特徴量(幾つ
かの要素を含むものであってもよい)が抽出される。大
分類識別手段(3)は、この特徴量を用いて簡単な手法
で候補文字を1つ以上選定する。上記候補文字は、認識
しようとする文字と類似の文字群から構成され、これら
の候補文字の中には正続文字が存在していると考えられ
る。
Effects> According to the character recognition device configured as described above, the feature amount extracting means (2) extracts the feature amount (several elements) are extracted. The major classification identification means (3) selects one or more candidate characters by a simple method using this feature amount. The above candidate characters are composed of a group of characters similar to the character to be recognized, and it is thought that regular continuation characters exist among these candidate characters.

ところで、認識しようとする文字が複雑な場合、候補文
字もこれに類似した複雑な文字からなり、認識しようと
する文字が簡単な場合、候補文字もこれに類似した簡単
な文字からなることは容易に推測できる。ここで、「複
雑」 「簡単」というのは、認識する上で他の文字を誤
読してしまう可能性が高いか低いかをいい、例えば画数
の少ない文字でも類似した文字が多くあれば「複雑」な
文字であり、画数の多い文字でもそれが特異な特徴を持
っていて類似した文字がほとんどない場合「簡単」な文
字といえる。
By the way, if the character you are trying to recognize is complex, the candidate characters will also consist of similar complex characters, and if the character you are trying to recognize is simple, the candidate characters will also easily consist of similar simple characters. It can be inferred that Here, "complex" and "easy" refer to whether there is a high or low possibility of misreading other characters during recognition. Even if a character has a large number of strokes, it can be said to be an "easy" character if it has unique characteristics and there are almost no similar characters.

そこで、「複雑」な文字を認識しようとすれば、詳細識
別手段(4)において、認識用辞書(6)に記憶された
特徴量その他の情報と詳細に比較しなければ、目的とす
る文字認識率を達成することはできない。
Therefore, in order to recognize "complex" characters, the detailed identification means (4) must compare them in detail with the feature values and other information stored in the recognition dictionary (6). rate cannot be achieved.

しかし、「簡単」な文字を認識しようとする時には、必
ずしも上記「複雑」な文字と同様の比較を行わなくても
よい。
However, when attempting to recognize "simple" characters, it is not necessary to perform the same comparison as for the "complex" characters described above.

本発明では、詳細識別手段(4)に、指定された精度で
特徴量の比較計算をする機能を持たせるとともに、各候
補文字の「簡単」さ「複雑」さに対応して、特徴量比較
計算をする時の計算精度を指定する定数を計算精度対応
表(7)に記憶させている。
In the present invention, the detailed identification means (4) is provided with a function to compare and calculate feature quantities with specified precision, and also perform feature quantity comparisons according to the "simpleness" and "complexity" of each candidate character. Constants that specify calculation accuracy when performing calculations are stored in calculation accuracy correspondence table (7).

大分類識別手段(3)で選定された候補文字は計算精度
決定手段(8)に通知され、計算精度決定手段(8)は
通知された候補文字の中から所定の基準で1つまたは複
数の候補文字を選ぶ。そして、計算精度対応表(7)を
参照して、上記選ばれた候補文字に対応する計算精度を
基にして、1つの計算精度を決定して詳細識別手段(4
)に送る。詳細識別手段(4)は、認識用辞書(6)を
参照しながら、上記指定された計算精度で各候補文字ご
とに比較計算をする。認識文字出力手段(5)は、詳細
識別手段(4)で識別された文字の中から特定の文字を
選定して当該文字を表わす信号を出力する。
The candidate characters selected by the major classification identification means (3) are notified to the calculation accuracy determination means (8), and the calculation accuracy determination means (8) selects one or more candidate characters from among the notified candidate characters based on predetermined criteria. Select candidate characters. Then, with reference to the calculation accuracy correspondence table (7), one calculation accuracy is determined based on the calculation accuracy corresponding to the selected candidate character, and the detailed identification means (4) is determined.
). The detailed identification means (4) performs comparative calculations for each candidate character with the specified calculation accuracy while referring to the recognition dictionary (6). The recognized character output means (5) selects a specific character from among the characters identified by the detailed identification means (4) and outputs a signal representing the character.

このように、候補文字の「簡単」さ、「複雑」さに対応
して計算精度を決定し、この計算精度で詳細計算を行う
ようにした。
In this way, the calculation accuracy is determined according to the "simpleness" or "complexity" of the candidate character, and detailed calculations are performed using this calculation accuracy.

〈実施例〉 以下実施例を示す添付図面によって詳細に説明する。<Example> Embodiments will be described in detail below with reference to the accompanying drawings showing embodiments.

第2図は、本発明の文字認識装置の一構成を示すブロッ
ク図であり、(la)は原稿全体を写し出すビジコン等
のイメージカメラ、およびその出力信号を二値化して整
形された信号を得る二値化回路からなる画像入力部を表
わす。画像信号は、文字切り出し手段(1b)によって
1字1字ごとに細分される。特徴抽出手段(2)は、各
文字の特徴量を抽出する。例えば、文字輪郭線の方向ベ
クトルのヒストグラムや空白部領域量の分布である。こ
の特徴量は、大分類識別手段(3)に入力され、大分類
識別手段(3)は、演算時間の短い簡単な識別関数、例
えばシティ・ブロック関数や線形−次間数等で全字種と
の相違度を計算する。そして、複数の候補文字を選定す
る。
FIG. 2 is a block diagram showing the configuration of the character recognition device of the present invention, in which (la) shows an image camera such as a bidicon that captures the entire document, and the output signal thereof is binarized to obtain a formatted signal. It represents an image input section consisting of a binarization circuit. The image signal is subdivided character by character by character cutting means (1b). The feature extraction means (2) extracts the feature amount of each character. For example, it is a histogram of directional vectors of character outlines or a distribution of blank area amounts. This feature quantity is input to the major classification discriminating means (3), which uses a simple discriminating function that takes a short calculation time, such as a city block function or a linear-dimensional number, for all character types. Calculate the degree of difference between Then, a plurality of candidate characters are selected.

各候補文字のデータは計算精度決定手段(8)に送られ
、計算精度決定手段B)は、計算精度対応表(7)を用
いて、候補文字のうちの最上位の文字(相違度が最少で
ある文字)、または上位数位の文字について計算精度対
応表(7)から計算精度を検索し、これらの計算精度に
基づいて1つの計算精度を決定し、詳細識別手段(4)
に送り出す。計算精度対応表ωの内容は、計算精度対応
表(7)の作成時に予め決定しておけばよいが、実際に
使用した結果により内容を更新していくことが好ましい
。例えば、文字認識装置に学習機能を付け、学習により
更新させるようにすればよい。また、学習という方法を
とらず人為的に更新するようにしてもよい。
The data of each candidate character is sent to the calculation accuracy determination means (8), and the calculation accuracy determination means (B) uses the calculation accuracy correspondence table (7) to determine the highest character among the candidate characters (the character with the lowest degree of difference). Search the calculation accuracy from the calculation accuracy correspondence table (7) for the character with the highest number) or the character in the top number, determine one calculation accuracy based on these calculation accuracies, and use the detailed identification means (4)
send to. The contents of the calculation accuracy correspondence table ω may be determined in advance when creating the calculation accuracy correspondence table (7), but it is preferable to update the contents according to the results of actual use. For example, a learning function may be added to the character recognition device, and the character recognition device may be updated through learning. Alternatively, the information may be updated artificially without using the learning method.

詳細識別手段(4)では、上記決定された計算精度に基
づき、各候補文字について相違度を計算する。
The detailed identification means (4) calculates the degree of dissimilarity for each candidate character based on the calculation accuracy determined above.

例えば、認識用辞書(6)に入っている候補文字の特徴
量と、特徴抽出手段(2)で求められた特徴量との距離
を求める。これらの特徴量は、一般にベクトル(次元を
nとする)であるが、上記計算精度が低いときは、n次
元まで求める必要はなく、その計算精度に見合った低い
次元での距離を求めればよいのである。逆に、計算精度
が高いときは、それに見合った高い次元での距離を求め
る必要がある。
For example, the distance between the feature amount of the candidate character included in the recognition dictionary (6) and the feature amount obtained by the feature extraction means (2) is determined. These feature quantities are generally vectors (dimension is n), but when the above calculation accuracy is low, there is no need to calculate up to n dimensions, and it is sufficient to calculate distances in a lower dimension commensurate with the calculation accuracy. It is. On the other hand, when the calculation accuracy is high, it is necessary to find the distance in a correspondingly high dimension.

このようにして、各候補文字につき相違度が求まるので
、順位決定手段(51)は相違度の小さいものから順に
並べ変え、第一順位の文字コードを出力する。
In this way, since the degree of dissimilarity is determined for each candidate character, the ranking determining means (51) rearranges the candidate characters in descending order of degree of dissimilarity and outputs the character code of the first rank.

次に、以上の文字認識装置の動作を「本日は晴天なり」
と書かれた原稿の最初の文字「本」を認識する場合を例
にとって説明する。
Next, we will change the operation of the above character recognition device to "It's sunny today".
An example of recognizing the first character "hon" in a manuscript written as "hon" will be explained.

画像入力部(1a)が「本日は晴天なり」と書かれた原
稿を読み取ると、文字切り出し手段(1b)は「本」 
「日」 「は」 「晴」 「天」 「な」 「す」とい
った単文字に画像信号を変換する。そして特徴抽出手段
(2)ハ、「本」の特徴量X  (Xi、X2.−、X
n)を抽出する。大分類識別手段(3)は、前述したよ
うな手法で複数の候補文字「本」 「木」 「水」・・
・を選定する。計算精度決定手段(8)は、例えば候補
文字のうちの最上位の文字「本」について計算精度を検
索し、計算精度対応表(7)から精度3との結果を得る
。詳細識別手段(4)は、上記精度3を基に距離計算を
行う。例えば認識用辞書(6)に入っている候補文字「
本」の特徴量をA (Al 、 A2、−An)、「木
」の特徴量をB (B1.B2.・・・Bn) 、・・
・とすると、下記のシティ・ブロック関数により距離X
−A  I  −[(Xi−Al)2 +  (X2−
A2)2+   (X3−A3) 2 ]  112X
 −B l −[(Xi−BL)2+(X2−82)2
十  (X3−B3)2 ]  ”2 等を求めていく。
When the image input unit (1a) reads a manuscript that says "Today is sunny", the character extraction means (1b) reads "Book".
Image signals are converted into single characters such as ``Sun'', ``Ha'', ``Sunny'', ``Ten'', ``Na'', and ``Su''. And feature extraction means (2) c. Feature amount of "book" X (Xi, X2.-, X
Extract n). The major classification identification means (3) uses the method described above to select multiple candidate characters such as ``book,''``tree,''``water,'' and so on.
・Select. The calculation accuracy determining means (8) searches for the calculation accuracy of, for example, the highest character "hon" among the candidate characters, and obtains a result of accuracy 3 from the calculation accuracy correspondence table (7). The detailed identification means (4) performs distance calculation based on the accuracy 3 described above. For example, candidate characters in the recognition dictionary (6)
The feature amount of "book" is A (Al, A2, -An), the feature amount of "tree" is B (B1.B2...Bn),...
・Then, the distance X is calculated by the city block function below.
-A I -[(Xi-Al)2 + (X2-
A2) 2+ (X3-A3) 2 ] 112X
-B l -[(Xi-BL)2+(X2-82)2
10 (X3-B3)2 ] ” 2 etc. will be found.

なお、前述した大分類識別の際でもこのシティ・ブロッ
ク関数が使われることがあるが、大分類識別の際は、例
えば最初の1項のみについて計算する等、非常に大雑把
な計算をする点で詳細分類の計算手法と異なっている。
Note that this city-block function is sometimes used in the case of the above-mentioned major classification identification, but in the case of major classification identification, it is difficult to perform very rough calculations, such as calculating only the first term. This is different from the calculation method for detailed classification.

順位決定手段(51)は距離の小さいものから順に並べ
変え、第一順位の文字コードを出力する。
A ranking determining means (51) rearranges the characters in descending order of distance and outputs the character code of the first ranking.

以上のようにして、「本日は晴天なり」全ての文字の認
識に要する時間を第3図に示す。
As described above, FIG. 3 shows the time required to recognize all the characters "It's sunny today."

第3図(a)は、従来の文字認識装置で認識した時間を
示すものであり、各文字について同じ計算時間tがかか
っているので、全ての文字に対しては、 X7 の時間がかかる。一方、第3図(b)は本実施例の文字
認識装置を用いて認識した時間を示し、各文字の「複雑
」さの相違に応じて計算時間t 1. t 2゜・・・
X7が異なっている。この結果、全体としての認識時間
は、 1−Σtn となり、第3図(a)の場合よりも短縮できることが分
かる。
FIG. 3(a) shows the time required for recognition by a conventional character recognition device. Since the same calculation time t is required for each character, it takes X7 time for all characters. On the other hand, FIG. 3(b) shows the time required for recognition using the character recognition device of this embodiment, and the calculation time t1. t 2゜...
X7 is different. As a result, it can be seen that the overall recognition time is 1-Σtn, which is shorter than in the case of FIG. 3(a).

なお、上記の実施例では、計算精度決定手段(8)が候
補文字のうちの最上位の文字「本」に基づいて計算精度
を決定する例を示したが、これに限定されるものではな
く、計算精度決定手段(8)は上位数文字についてそれ
ぞれ精度を求め、求めた複数の精度に基づき1つの計算
精度を決定するようにしてもよい。例えば、上位3文字
「本」 「木」「水」につき精度がそれぞれ3,4,5
であったとすると、これらの精度の平均値をとったり、
最大値、最小値をとったりして1つの計算精度を決定し
てもよい。なぜなら、大分類の結実現れる候補文字群に
含まれる各候補文字は、計算精度対応表(7)の中では
、おおむね等しい計算精度を伴っていると考えられるの
で、最上位の文字のみについて計算精度を決定しても、
上位幾つかの文字について計算精度を決定しても、結果
はあまり異ならないからである。
In addition, in the above embodiment, an example was shown in which the calculation accuracy determining means (8) determines the calculation accuracy based on the highest character "hon" among the candidate characters, but the invention is not limited to this. The calculation precision determining means (8) may determine the precision of each of the top several characters, and determine one calculation precision based on the plurality of precisions found. For example, the accuracy is 3, 4, and 5 for the top three characters ``book,''``tree,'' and ``water,'' respectively.
If so, take the average value of these accuracies, or
One calculation accuracy may be determined by taking the maximum value or the minimum value. This is because each candidate character included in the candidate character group realized as a result of the major classification is considered to have approximately the same calculation accuracy in the calculation accuracy correspondence table (7), so only the highest character has a calculation accuracy of Even if you decide
This is because even if the calculation accuracy is determined for the top few characters, the results will not differ much.

また、上記の実施例では、詳細識別手段(4)は、単純
に距離を求めていたが、認識率を上げるため、もっと複
雑な計算式を用いることも可能であるのは勿論である。
Further, in the above embodiment, the detailed identification means (4) simply calculates the distance, but it is of course possible to use a more complicated calculation formula in order to increase the recognition rate.

このときでも、候補文字の「簡単」さに応じて計算する
項数を限定することができるので、本発明は計算時間の
短縮に非常に効果的である。
Even in this case, the number of terms to be calculated can be limited depending on the "simpleness" of the candidate character, so the present invention is very effective in shortening calculation time.

また、上記実施例では、イメージカメラを用いて原稿画
像を入力していたが、これに限定されるものではなく、
ファクシミリ等通信回線を通して画像上方を入力するも
のであってもよい。
Further, in the above embodiment, the image camera was used to input the document image, but the invention is not limited to this.
The upper part of the image may be input through a communication line such as a facsimile.

その池水発明の要旨を変更しない範囲内において、種々
の設計変更を施すことが可能である。
Various design changes can be made without changing the gist of the invention.

〈発明の効果〉 以上のように、本発明の文字認識装置によれば、文字ご
とに比較計算の精度を文字の「複雑j 「簡単」の度合
いに応じて予め決定して計算精度対応表に記憶させてお
き、計算精度決定手段により大分類識別された候補文字
群に適応する計算精度を決定することとした。
<Effects of the Invention> As described above, according to the character recognition device of the present invention, the accuracy of comparison calculation for each character is determined in advance according to the degree of "complexity" and "simpleness" of the character, and the calculation accuracy correspondence table is created. The calculation accuracy is then stored and the calculation accuracy determined by the calculation accuracy determination means to be applied to the candidate character group classified into major categories.

したがって、従来どの文字に対しても最大限の精度で計
算していたため計算時間に無駄が生じていたところ、本
発明では、決定された計算精度により、文字ごとに最低
限の精度で比較計算できるため、「複雑」 「簡単」な
文字の混在する被読取対象を認識する際に、1文字当た
りの平均認識時間を短縮することができる。言い換えれ
ば、同じ認識時間を許されるならば文字認識の精度を高
めることができる。よって、文字認識装置としての認識
性能の向上を実現することができる。
Therefore, whereas conventional calculations were performed with maximum accuracy for each character, resulting in wasted calculation time, the present invention enables comparative calculations for each character with the minimum accuracy based on the determined calculation accuracy. Therefore, when recognizing an object to be read containing a mixture of "complex" and "simple" characters, the average recognition time per character can be shortened. In other words, if the same recognition time is allowed, the accuracy of character recognition can be increased. Therefore, it is possible to improve the recognition performance of the character recognition device.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は本発明の文字認識装置の構成を示すブロック図
、 第2図は文字認識装置一実施例を示すブロック構成図、 第3図は文字認識時間の従来例との比較表、第4図、第
5図は従来の文字認識装置の構成を示すブロック図であ
る。 (1)・・・画像信号取得手段、(2)・・・特徴量抽
出手段、(3)・・・大分類識別手段、(4)・・・詳
細識別手段、(5)・・・認識文字出力手段、(6)・
・・認識用辞書、(7)・・・計算精度対応表、(8)
・・・計算精度決定手段時 許 出 願 人 住友電気工業株式会社
FIG. 1 is a block diagram showing the configuration of the character recognition device of the present invention, FIG. 2 is a block diagram showing an embodiment of the character recognition device, FIG. 3 is a comparison table of character recognition time with the conventional example, and FIG. 5 are block diagrams showing the configuration of a conventional character recognition device. (1)...Image signal acquisition means, (2)...Feature amount extraction means, (3)...Major classification identification means, (4)...Detailed identification means, (5)...Recognition Character output means, (6)・
...Recognition dictionary, (7) ...Calculation accuracy correspondence table, (8)
... Calculation accuracy determining means Applicant: Sumitomo Electric Industries, Ltd.

Claims (1)

【特許請求の範囲】 1、文字を含む被読取対象を表わす画像信 号を取得する画像信号取得手段と、上記 画像信号に基づき画像中の認識しようと する文字の特徴量を抽出する特徴量抽出 手段と、上記特徴量を基に演算を行い1 つ以上の候補文字を選定する大分類識別 手段と、各文字の認識に必要な特徴量に 関する情報を記憶した認識用辞書と、上 記認識しようとする文字の特徴量を認識 用辞書に記憶された候補文字の情報と比 較し特徴量比較計算を行う詳細識別手段 と、詳細識別手段で識別された文字の中 から一定の基準で文字を選択して当該文 字を表わす信号を出力する認識文字出力 手段とを有する文字認識装置において、 特徴量比較計算をする時の計算精度を 指定する定数を各文字に対応して記憶し ている計算精度対応表と、上記候補文字 の中から所定の基準で1つまたは複数の 候補文字を選び出し、この選び出した候 補文字に対応する計算精度を計算精度対 応表から検索し、検索した計算精度を基 に1つの計算精度を決定して詳細識別手 段に送り出す計算精度決定手段とを具備 し、かつ、上記詳細識別手段が計算精度 決定手段から指定された計算精度に応じ て特徴量比較計算をすることを特徴とす る文字認識装置。[Claims] 1. Image signal representing the object to be read including characters an image signal acquisition means for acquiring the image signal; Trying to recognize images based on image signals Feature extraction to extract the features of characters Perform calculations based on the means and the above feature amounts 1 Broad classification identification that selects two or more candidate characters methods and features necessary for recognizing each character. A recognition dictionary that stores information about Recognizes the features of the characters to be recognized Candidate character information stored in the dictionary Detailed identification means for calculating feature value comparison and among the characters identified by the detailed identification means. Select characters based on certain criteria and write the relevant sentence. Recognized character output that outputs a signal representing a character In a character recognition device having means, Calculation accuracy when performing feature comparison calculations Memorizes the specified constant corresponding to each character. Calculation accuracy correspondence table and the above candidate characters one or more based on predetermined criteria from among Select a candidate character and use this selected candidate character. Calculate the calculation precision corresponding to the complementary character vs. calculation precision Search from the table and based on the searched calculation accuracy. Determine the calculation accuracy and perform detailed identification Equipped with means for determining the calculation accuracy to be sent to the stage and the above detailed identification means has calculation accuracy. Depending on the calculation accuracy specified by the determining means The feature is that feature comparison calculation is performed using character recognition device.
JP1012720A 1989-01-20 1989-01-20 Character recognizing device Pending JPH02193280A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP1012720A JPH02193280A (en) 1989-01-20 1989-01-20 Character recognizing device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP1012720A JPH02193280A (en) 1989-01-20 1989-01-20 Character recognizing device

Publications (1)

Publication Number Publication Date
JPH02193280A true JPH02193280A (en) 1990-07-30

Family

ID=11813264

Family Applications (1)

Application Number Title Priority Date Filing Date
JP1012720A Pending JPH02193280A (en) 1989-01-20 1989-01-20 Character recognizing device

Country Status (1)

Country Link
JP (1) JPH02193280A (en)

Similar Documents

Publication Publication Date Title
EP0844583B1 (en) Method and apparatus for character recognition
JPH0664631B2 (en) Character recognition device
KR19980018029A (en) Character recognition device
US20010051965A1 (en) Apparatus for rough classification of words, method for rough classification of words, and record medium recording a control program thereof
JPH087033A (en) Information processing method and device
JP3917349B2 (en) Retrieval device and method for retrieving information using character recognition result
JPH0682403B2 (en) Optical character reader
JP3589007B2 (en) Document filing system and document filing method
JP2000181931A (en) Automatic authoring device and recording medium
JP2586372B2 (en) Information retrieval apparatus and information retrieval method
KR19990016894A (en) How to search video database
JPH02193281A (en) character recognition device
JPH05314320A (en) Evaluation method of recognition result using difference of recognition distance and candidate order
JPH0528324A (en) English character recognition device
JP2996823B2 (en) Character recognition device
JPH0766423B2 (en) Character recognition device
JPH08180064A (en) Document retrieval method and document filing device
JP2728117B2 (en) Character recognition device
JP2851865B2 (en) Character recognition device
JP3720405B2 (en) Region identification apparatus and method
JP2622004B2 (en) Character recognition device
JP2977244B2 (en) Character recognition method and character recognition device
JPH0322094A (en) Character recognizing system
JPH0528323A (en) Character recognition device
JPS59188783A (en) Character discriminating and processing system