JPH02219190A - Character recognizing method - Google Patents

Character recognizing method

Info

Publication number
JPH02219190A
JPH02219190A JP1039308A JP3930889A JPH02219190A JP H02219190 A JPH02219190 A JP H02219190A JP 1039308 A JP1039308 A JP 1039308A JP 3930889 A JP3930889 A JP 3930889A JP H02219190 A JPH02219190 A JP H02219190A
Authority
JP
Japan
Prior art keywords
character
characters
recognition
shape feature
long
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP1039308A
Other languages
Japanese (ja)
Other versions
JPH07101441B2 (en
Inventor
Kazuyuki Yoshida
吉田 収志
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Fuji Electric Co Ltd
Fuji Facom Corp
Original Assignee
Fuji Electric Co Ltd
Fuji Facom Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Fuji Electric Co Ltd, Fuji Facom Corp filed Critical Fuji Electric Co Ltd
Priority to JP1039308A priority Critical patent/JPH07101441B2/en
Publication of JPH02219190A publication Critical patent/JPH02219190A/en
Publication of JPH07101441B2 publication Critical patent/JPH07101441B2/en
Anticipated expiration legal-status Critical
Expired - Lifetime legal-status Critical Current

Links

Landscapes

  • Character Discrimination (AREA)

Abstract

PURPOSE:To contrive the improvement of the recognition rate and the segmenting performance of a character, etc., by comparing the shape feature quantity of an input character with the shape feature quantity stored in advance of a candidate character corresponding to the result of its recognition and determining whether the result of recognition based on the similarity is adopted or not. CONSTITUTION:For instance, it is assumed that an input image is (u) (Japanese), and a shape feature obtained from the input image is 'long longitudinally, and separated into two in the longitudinal direction', and on the other hand, a first place in a first - a tenth places of candidate characters of a result of recognition is (ro) (Japanese). In this case, the shape feature of (ro) stored in advance in a dictionary is 'not long longitudinally nor long laterally and not separated longitudinally nor laterally'. Accordingly, since its result becomes 'NO' by a longitudinally long/laterally long check, the check of the candidate character of a second place is executed. The candidate character of a second place is (u) (Japanese), and the shape feature quantity coincides, therefore, it is adopted as the result of recognition. In such a way, the improvement of the recognition rate and the segmenting performance of a character, etc., can be contrived.

Description

【発明の詳細な説明】 〔産業上の利用分野〕 この発明は、文字、マーク(文字等とも云う)を認識す
るための文字認識方法、特にその改良に関する。
DETAILED DESCRIPTION OF THE INVENTION [Field of Industrial Application] The present invention relates to a character recognition method for recognizing characters and marks (also referred to as characters, etc.), and particularly to improvements thereof.

〔従来の技術〕[Conventional technology]

従来、この種の方法として、例えば文書画像から文字行
または文字列(以下、文字行と云う。)を切り出した後
、文字らしきものを仮文字として抽出し、この仮文字か
ら所定の文字サイズを基準として全角文字(漢字等)を
切り出し、文字サイズから全角文字と確定できないもの
は、半角文字として有効に成立し得る文字か否かを調べ
、成立すれば半角文字として認識する方法がある。
Conventionally, as a method of this kind, for example, after cutting out a character line or character string (hereinafter referred to as a character line) from a document image, what looks like characters is extracted as temporary characters, and a predetermined character size is extracted from these temporary characters. There is a method of cutting out full-width characters (such as Kanji characters) as a standard, and if characters cannot be determined to be full-width characters based on the character size, check whether the characters can be effectively established as half-width characters or not, and if so, recognize them as half-width characters.

しかしながら、このような方法では文字によって全角文
字が2つ以上の半角文字として誤認識されるおそれがあ
る。そこで、出願人は次のような方法を提案している(
特願昭63−292445号;提案漬方法とも云う)。
However, in such a method, there is a risk that a full-width character may be mistakenly recognized as two or more half-width characters. Therefore, the applicant has proposed the following method (
Japanese Patent Application No. 63-292445; also referred to as the proposed pickling method).

第4図は提案漬方法を説明するためのフローチャートで
ある。
FIG. 4 is a flow chart for explaining the proposed dipping method.

これは、図示されない画像処理装置の処理手順を示すも
ので、まず文書画像データを入力しく■参照)、その水
平方向の投影値をとることにより、各文字行を切り出す
(■参照)、これにより、行の幅寸法を求め、全角文字
の大きさに相当する量(文字サイズ)を得る。なお、こ
こでは横書きの場合を想定しているが、縦書きの場合も
同様である。
This shows the processing procedure of an image processing device (not shown). First, document image data is input (see ■), and each character line is cut out by taking its horizontal projection value (see ■). , determine the line width dimension and obtain the amount equivalent to the size of a full-width character (character size). Note that although horizontal writing is assumed here, the same applies to vertical writing.

次に、各行に垂直な方向の投影値を調べ、文字サイズを
考慮することにより、各文字行から文字らしきもの、す
なわち仮文字群を切り出しく■参照)、シかる後この仮
文字群の中から上記文字サイズを利用して全角文字を選
出する(■参照)。
Next, by examining the projection value in the direction perpendicular to each line and taking into account the character size, cut out what appears to be a character, that is, a group of temporary characters, from each line of characters. Select full-width characters using the above font size (see ■).

全角文字として選出する条件は次のとおりである。The conditions for selecting characters as full-width characters are as follows.

イ)それ単独で文字サイズが全角サイズのもの、つまり
他の仮文字と結合する余地の全くないもの。
b) The font size is full-width by itself, that is, there is no room to combine it with other temporary characters.

口)句読点 ハ)それ単独では半角サイズであるが、隣り合う他の半
角サイズの仮文字と結合させてみると全角サイズとなる
もの。
(mouth) Punctuation marks (c) Punctuation marks are half-width by themselves, but become full-width when combined with other half-width adjacent characters.

二)それ単独ではサイズが全角サイズよりも小さいが、
隣り合う他の半角サイズの仮文字との間に距離があり過
ぎ、これらを無理に結合させると全角文字サイズをこえ
るもの。
2) Although the size alone is smaller than the full-width size,
There is too much distance between adjacent half-width temporary characters, and if you forcefully combine them, the size will exceed the full-width character size.

以上の如き条件に従って全角文字を全て選出した後、あ
とに残った仮文字について、これを統合または分離して
統合文字2分離文字を作成しく■参照)、シかる後これ
らの統合文字9分離文字をOCR(文字読取装置)によ
り、辞書パターンとの類似度を利用して認識する(■参
照)。
After selecting all full-width characters according to the above conditions, combine or separate the remaining temporary characters to create 2 integrated characters and 2 separated characters (see ■), and after selecting these integrated characters and 9 separated characters. is recognized by an OCR (character reading device) using the degree of similarity with a dictionary pattern (see ■).

次に、その認識結果に対して次のような矛盾処理を実行
する(■参照)。
Next, the following contradiction processing is performed on the recognition result (see ■).

a)例えば認識すべき対象が分離文字であるにもかかわ
らず、OCRによる認識結果が全角サイズの漢字を示す
ものとすれば互いに矛盾するので、かかる認識結果は採
用しない。
a) For example, even though the object to be recognized is a separate character, if the recognition result by OCR indicates a full-width kanji character, it would be inconsistent with each other, so such a recognition result would not be adopted.

b)上記とは逆に、認識すべき対象が統合文字であるに
もかかわらず、OCRによる認識結果が英字、数字等の
半角サイズ文字を示す場合。
b) Contrary to the above case, even though the object to be recognized is an integrated character, the recognition result by OCR shows half-width size characters such as alphabetic characters and numbers.

そして、最後に残された文字につき、これを統合文字と
すべきか分離文字とすべきかを、OCRにより相対類似
度を用いて判別する(■参照)。
Then, regarding the last remaining character, whether it should be used as an integrated character or a separated character is determined by OCR using relative similarity (see ■).

なお、相対類似度X、は類似度Xと類似度の平均値mと
の比に成る定数(例えば、1024)を掛けたものとし
て定義する。すなわち、 x、=x/mx定数(1024) である。
Note that the relative similarity X is defined as the product of the similarity X and a constant (for example, 1024) that is the ratio of the average value m of the similarities. That is, x,=x/mx constant (1024).

〔発明が解決しようとする課題〕[Problem to be solved by the invention]

提案済方法では類似度から認識結果を決定するようにし
ているので、例えば「う」と「ろ」、「テ」と「チ」の
如く非常に似ているものは、文字画像の僅かな違い、例
えば文字等の太さや字体の相違1文字等のつぶれやかす
れ等により誤認識するおそれがある。
In the proposed method, recognition results are determined based on the degree of similarity, so for example, very similar words such as ``u'' and ``ro'' or ``te'' and ``chi'' may be recognized by slight differences in the character images. For example, there is a risk of erroneous recognition due to differences in the thickness or font of characters, such as one character being crushed or blurred.

また、同様の理由から、「は」の相対類似度よりも「(
」と「よ」の部分の相対類似度の方が高くなり、誤認識
してしまうことがある。この点について、もう少し具体
的に説明する。
Also, for the same reason, the relative similarity of ``(
” and the “yo” part have a higher relative similarity, which may lead to misrecognition. This point will be explained in more detail.

第5図は入力文字「は」について、これを統合文字1と
して処理した場合と、分離文字2,3として処理した場
合のI!(11度X、相対類僚度X、を示すものである
。すなわち、類似度x、 [似度の平均値mとが図示の
如く得られるものとすると、類似度とその平均値との比
に一定数(1024)を乗じて得られる相対類似度は、
統合文字lの場合はr954J 、分離文字2.3の場
合はそれぞれr1028J、r898Jでその平均値は
「963」となり、 954<963 であることから、入力文字「は」は「(」(前括弧)と
「は」からなるものとして誤認識されることになる。な
お、定数はr1024Jに限らないことは勿論である。
FIG. 5 shows I! of the input character "wa" when it is processed as integrated character 1 and when it is processed as separated characters 2 and 3. (11 degrees X, relative degree of assimilation The relative similarity obtained by multiplying by a constant number (1024) is
For the integrated character l, r954J is used, and for the separator character 2.3, r1028J and r898J are used, respectively, and the average value is "963". Since 954<963, the input character "wa" is "(" ) and "wa".The constant is of course not limited to r1024J.

〔課題を解決するための手段〕[Means to solve the problem]

少なくとも辞書パターンとの類似度を調べて文字を認識
するにあたり、認識すべき文字の縦横比と縦方向および
横方向にそれぞれいくつに分離できるかを示す形状特徴
量も辞書として予め記憶しておき、入力文字の形状特徴
量をその認識結果と対応する候補文字の予め記憶されて
いる形状特徴量と比較して類似度にもとづく認識結果を
採用するか否かを決定する。
When recognizing a character by checking at least its similarity with a dictionary pattern, the aspect ratio of the character to be recognized and shape feature values indicating how many parts can be separated in the vertical and horizontal directions are also stored in advance as a dictionary, The shape feature amount of the input character is compared with the shape feature amount stored in advance of the candidate character corresponding to the recognition result, and it is determined whether or not to adopt the recognition result based on the degree of similarity.

【作用〕[Effect]

類似度だけでなく、形状特徴量をも抽出するようにし、
「う」と「ろ」、「チ」と「テ」等のように、形状特徴
に差のある類似文字の認識率を向上させる。また、「は
」を「(」と「は」の如く誤認識しないようにし、文字
の切り出し性能の向上を図る。
In addition to the similarity, we also extract shape features,
Improves the recognition rate of similar characters with different shape characteristics, such as "u" and "ro", "chi" and "te", etc. Furthermore, the character extraction performance is improved by preventing ``wa'' from being misrecognized as ``('' and ``wa'').

〔実施例〕〔Example〕

第1図はこの発明の実施例を示すフローチャートである
。同図からも明らかなように、この実施例は形状特徴照
合ステップ■を付加した点が特徴である。
FIG. 1 is a flowchart showing an embodiment of the invention. As is clear from the figure, this embodiment is characterized by the addition of a shape feature matching step (2).

ステップ■は同図の如く[相]〜[相]に細分化される
Step (2) is subdivided into [phase] to [phase] as shown in the figure.

すなわち、ステップ[相]では入力文字が縦長か横長か
がチエツクされる。なお、縦長か横長かはそれが正体文
字の場合は縦横比(高さ7幅)が例えば2以上ならば縦
長とし、1/2以下ならば横長とする。また、長体文字
や平棒文字の場合は正体文字に直して判定することとす
る。
That is, in step [phase], it is checked whether the input character is vertically long or horizontally long. Regarding whether the character is vertically long or horizontally long, if the character is a regular character, if the aspect ratio (height 7 width) is, for example, 2 or more, it is considered vertically long, and if it is 1/2 or less, it is considered horizontally long. Furthermore, in the case of long characters or flat bar characters, the characters are converted to the normal characters for determination.

ステップ■、@ではそれぞれ縦、横をいくつに分離する
かを調べ、分離の態様が入力文字とその認識結果の候補
文字との間で一致するか否かを調ぺる。この操作を候補
文字(第1〜第10位)で適合するものが見つかる迄行
い、いずれの候補文字も適合しない場合はステップ[相
]でリジェクト出力を出す。
In steps {circle over (2)} and @, it is determined how many vertical and horizontal divisions are to be made, respectively, and whether or not the manner of separation matches the input character and the candidate character resulting from its recognition is checked. This operation is performed until a matching candidate character (1st to 10th) is found. If none of the candidate characters matches, a reject output is output in step [phase].

なお、以上の他は第3図と同様なので説明は省略する。Note that the other parts are the same as those in FIG. 3, so the explanation will be omitted.

次に、具体的な例について説明する。Next, a specific example will be explained.

第2図は入力画像(または入力文字)が「う」で、入力
画像から得られる形状特徴は“縦長で、縦方向に2つに
分離する”であり、これに対して認識結果の候補文字の
第1〜第10位のうちの第1位が「ろ」となった例であ
る。この場合、辞書に予め記憶されている「ろ」の形状
特徴は“縦長でも横長でもなく、縦横に分離しない”で
ある。
In Figure 2, the input image (or input character) is ``u'', and the shape feature obtained from the input image is ``vertically long and separated into two vertically'', and the recognition result candidate character This is an example in which the first place among the first to tenth places is "ro". In this case, the shape characteristics of "ro" stored in the dictionary in advance are "neither vertically nor horizontally long, nor separated vertically and horizontally."

したがって、縦長・横長チエツクのステップ(第1図[
相]参照)で“NO”となるため、第2位の候補文字の
チエツクが行われる。第2位の候補文字は「う」であり
、形状特徴量が一致するので、これを認識結果として採
用する。
Therefore, the step of vertical/horizontal check (Fig. 1 [
Since the answer is "NO", the second candidate character is checked. The second candidate character is "u", and since the shape feature values match, this is adopted as the recognition result.

第3図は入力文字「は」を統合文字として処理する場合
と、分離文字r(J、rよ」として処理する場合の例で
ある。すなわち、同図(イ)の如く統合文字として処理
する場合は、認識結果の第1候補「は」が形状特徴量の
点からも適合するので「は」が採用される。同様に、分
離文字「(」は同図(ロ)の如く第1候補文字「(」と
適合するが、分離文字「よ」は同図(ハ)の如く、その
形状特徴量から第1候補文字「は」とは適合せず、結局
第2位の候補文字「ま」と適合する。そして、同図(ニ
)に示す如く、「は」として認識したときの相対類似度
はr954J、「(」と「ま」に分離されるものとした
ときの相対類似度は「946」となり、 946<954 から、「は」が採用されることになる。つまり、入力文
字は「は」と認識され、「(」と「ま」に誤i!識され
るおそれをなくすことができる。
Figure 3 shows an example where the input character "wa" is processed as an integrated character and when it is processed as a separate character r (J, ryo). In other words, it is processed as an integrated character as shown in (a) of the same figure. In this case, the first candidate "ha" in the recognition result is suitable from the point of view of the shape feature, so "ha" is adopted.Similarly, the separated character "(" is the first candidate as shown in (b) in the same figure. Although it matches the character "(", the separated character "yo" does not match the first candidate character "ha" due to its shape features as shown in the same figure (c), and in the end it is matched with the second candidate character "ma". As shown in the same figure (d), the relative similarity when recognized as "wa" is r954J, and when it is separated into "(" and "ma"), the relative similarity is r954J. "946", and since 946<954, "ha" will be adopted.In other words, the input character will be recognized as "ha", eliminating the possibility that it will be mistakenly recognized as "(" and "ma"). be able to.

〔発明の効果〕〔Effect of the invention〕

この発明によれば、類似度による認識の他に、形状時f
!kit (縦長か横長か、縦方向にいくつに分離する
か、横方向にいくつに分離するか)も考慮するようにし
たので、類似度のみによる認識結果の誤りを修正するこ
とができ、文字等の認識率。
According to this invention, in addition to recognition based on similarity, shape time f
! Since the kit (vertically or horizontally, how many parts to separate vertically, how many parts to separate horizontally) can be taken into account, it is possible to correct errors in recognition results based only on similarity, and it is possible to correct characters, etc. recognition rate.

切出し性能を向上させることができる。Cutting performance can be improved.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図はこの発明の実施例を示すフローチャート、第2
図および第3図はいずれもこの発明を具体的に説明する
ための説明図、第4図は提案済方法を説明するためのフ
ローチャート、第5図は誤認識例を説明するための説明
図である。 1・・・統合文字、2,3・・・分離文字。 代理人弁理士  並 木 昭 夫
FIG. 1 is a flowchart showing an embodiment of the invention, and FIG.
3 and 3 are explanatory diagrams for specifically explaining the present invention, FIG. 4 is a flowchart for explaining the proposed method, and FIG. 5 is an explanatory diagram for explaining an example of erroneous recognition. be. 1... integrated character, 2, 3... separation character. Representative Patent Attorney Akio Namiki

Claims (1)

【特許請求の範囲】 1)文書を画像処理して個々の文字を切り出し、各文字
毎に辞書パターンとの類似度を求めて文字を認識するに
あたり、 認識すべき文字の縦横比と縦方向および横方向にそれぞ
れいくつに分離できるかを示す形状特徴量も辞書データ
として予め記憶しておき、入力文字の形状特徴量をその
認識結果と対応する候補文字の予め記憶されている形状
特徴量と比較して類似度にもとづく認識結果を採用する
か否かを決定することを特徴とする文字認識方法。
[Claims] 1) In recognizing characters by image processing a document, cutting out individual characters, and determining the degree of similarity with a dictionary pattern for each character, the aspect ratio and vertical direction of the characters to be recognized are determined. Shape features indicating how many parts each character can be separated in the horizontal direction are also stored in advance as dictionary data, and the recognition results are compared with the shape features of the corresponding candidate character stored in advance. A character recognition method characterized by determining whether or not to adopt a recognition result based on similarity.
JP1039308A 1989-02-21 1989-02-21 Character recognition method Expired - Lifetime JPH07101441B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP1039308A JPH07101441B2 (en) 1989-02-21 1989-02-21 Character recognition method

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP1039308A JPH07101441B2 (en) 1989-02-21 1989-02-21 Character recognition method

Publications (2)

Publication Number Publication Date
JPH02219190A true JPH02219190A (en) 1990-08-31
JPH07101441B2 JPH07101441B2 (en) 1995-11-01

Family

ID=12549486

Family Applications (1)

Application Number Title Priority Date Filing Date
JP1039308A Expired - Lifetime JPH07101441B2 (en) 1989-02-21 1989-02-21 Character recognition method

Country Status (1)

Country Link
JP (1) JPH07101441B2 (en)

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS6274181A (en) * 1985-09-27 1987-04-04 Sony Corp Character recognizing device

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS6274181A (en) * 1985-09-27 1987-04-04 Sony Corp Character recognizing device

Also Published As

Publication number Publication date
JPH07101441B2 (en) 1995-11-01

Similar Documents

Publication Publication Date Title
JP3452774B2 (en) Character recognition method
JP2001283152A (en) Device and method for discrimination of forms and computer readable recording medium stored with program for allowing computer to execute the same method
JPS60217477A (en) Handwritten character recognizing device
US5561720A (en) Method for extracting individual characters from raster images of a read-in handwritten or typed character sequence having a free pitch
JP2000315247A (en) Character recognition device
Baird Global-to-local layout analysis
JP2002063548A (en) Handwritten character recognizing method
JP4194020B2 (en) Character recognition method, program used for executing the method, and character recognition apparatus
JPH02219190A (en) Character recognizing method
JP3374762B2 (en) Character recognition method and apparatus
JP2963474B2 (en) Similar character identification method
JPH028348B2 (en)
JP3151866B2 (en) English character recognition method
JPH02230484A (en) Character recognizing device
JP2977244B2 (en) Character recognition method and character recognition device
JP2752499B2 (en) Character reader
Hwang et al. Segmentation of a text printed in Korean and English using structure information and character recognizers
JPH01231187A (en) Character recognizing device
JP2000207491A (en) Character string reading method and apparatus
JPH11120294A (en) Character recognition device and medium
JP2570311B2 (en) String recognition device
JP3725944B2 (en) Character recognition device
JPH07160831A (en) Rejection method for handwritten character recognition results
JP3033904B2 (en) Character recognition post-processing method
JPH08171608A (en) Form style identification method and device