JPH0746371B2 - Character reader - Google Patents

Character reader

Info

Publication number
JPH0746371B2
JPH0746371B2 JP62330959A JP33095987A JPH0746371B2 JP H0746371 B2 JPH0746371 B2 JP H0746371B2 JP 62330959 A JP62330959 A JP 62330959A JP 33095987 A JP33095987 A JP 33095987A JP H0746371 B2 JPH0746371 B2 JP H0746371B2
Authority
JP
Japan
Prior art keywords
character
pattern
missing
area
character pattern
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Expired - Lifetime
Application number
JP62330959A
Other languages
Japanese (ja)
Other versions
JPH01171082A (en
Inventor
一巳 松浦
啓二 小林
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Mitsubishi Electric Corp
Original Assignee
Mitsubishi Electric Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Mitsubishi Electric Corp filed Critical Mitsubishi Electric Corp
Priority to JP62330959A priority Critical patent/JPH0746371B2/en
Publication of JPH01171082A publication Critical patent/JPH01171082A/en
Publication of JPH0746371B2 publication Critical patent/JPH0746371B2/en
Anticipated expiration legal-status Critical
Expired - Lifetime legal-status Critical Current

Links

Landscapes

  • Character Input (AREA)
  • Character Discrimination (AREA)

Description

【発明の詳細な説明】 〔産業上の利用分野〕 この発明は、文字読取装置に関するものであり、特に、
工業用カメラ等で撮像した,欠けや隠れが生じた低品質
の刻印文字,印刷文字,手書き文字等を読取る文字読取
装置に関するものである。
TECHNICAL FIELD The present invention relates to a character reading device, and in particular,
The present invention relates to a character reading device for reading low-quality engraved characters, printed characters, handwritten characters, etc., which are imaged with an industrial camera or the like, and which are missing or hidden.

〔従来の技術〕[Conventional technology]

画像中の文字を読取る文字読取装置は、画像から個々の
文字を切り出す文字パターン切り出し装置と、切り出さ
れた文字パターンを認識する文字認識装置とを組合わせ
ることにより実現することが出来る。
A character reading device that reads characters in an image can be realized by combining a character pattern cutout device that cuts out individual characters from an image and a character recognition device that recognizes the cut out character pattern.

第3図は、例えば、特公昭55−42434号公報に示された
従来の文字パターン切り出し装置と、特公昭53−46062
号公報に示された従来の文字認識装置とを組合わせた従
来の文字読取装置の構成図である。
FIG. 3 shows, for example, a conventional character pattern cutting device disclosed in Japanese Patent Publication No. 55-42434 and Japanese Patent Publication No. 53-46062.
It is a block diagram of the conventional character reading device which combined with the conventional character recognition device shown in the publication.

図において、1は入力された帳票等のディジタル画像を
記憶する画像メモリ、2は帳票等の媒体上に設定した文
字枠或いは印刷又は刻印する文字の位置や大きさ等の既
知の文字配列情報を格納した文字配列情報格納メモリ、
3は文字配列情報格納メモリ2の内容を参照して画像メ
モリ1から文字パターンを切り出す文字切り出し手段、
4は認識対象のカテゴリに対応した基準パターンを格納
した認識辞書を記憶する辞書メモリ、5は画像メモリ1
から切り出した文字パターンと辞書メモリ4に記憶した
各カテゴリの基準パターンとを重ね合わせることにより
入力文字パターンのカテゴリを決定する文字認識手段で
ある。なお、上記文字切り出し手段3及び文字認識手段
5はマイクロプロセッサ等により実現されるものであ
る。
In the figure, 1 is an image memory for storing a digital image such as an input form, and 2 is a character frame set on a medium such as a form or known character arrangement information such as the position and size of characters to be printed or stamped. Stored character array information storage memory,
Reference numeral 3 denotes a character cutting means for cutting out a character pattern from the image memory 1 by referring to the contents of the character array information storage memory 2.
4 is a dictionary memory that stores a recognition dictionary that stores reference patterns corresponding to categories to be recognized, and 5 is an image memory 1.
It is a character recognition means for determining the category of the input character pattern by superimposing the character pattern cut out from the reference pattern of each category stored in the dictionary memory 4. The character cutout means 3 and the character recognition means 5 are realized by a microprocessor or the like.

第4図(a)〜(i)は、第3図に示した従来の文字読
取装置の動作例を示した図である。
FIGS. 4 (a) to 4 (i) are diagrams showing an operation example of the conventional character reading device shown in FIG.

同図(a)に示した6は画像メモリ1に記憶した入力画
像で、6aはカテゴリ「3」の文字パターン,6bはカテゴ
リ「4」の文字パターン,6cはカテゴリ「8」の文字パ
ターン,6dはカテゴリ「7」の文字パターン、同図
(b)に示した7は文字配列情報格納メモリ2の内容で
ある文字の位置及び大きさを矩形で表現したものであ
り、7a,7b,7c,7dは左からそれぞれ1文字目,2文字目,3
文字目,4文字目の位置及び大きさを表現した矩形、同図
(c)に示した8は文字配列情報格納メモリ2の内容を
数値で表現したテーブルであり、8a,8b,8c,8dはそれぞ
れ上記矩形7a,7b,7c,7dの位置番号(NO),左上端のX
座標(SX),左上端のY座標(SY),幅(W),高さ
(H)を数値により表現したものである。一方、同図
(d)に示した9は入力画像6から作成した周辺分布
で、9aはY方向の周辺分布,9bはX方向の周辺分布、同
図(e)に示した10は入力画像6から切り出した文字パ
ターンの矩形で、10a,10b,10c,10dはそれぞれ矩形7a,7
b,7c,7dに対応させて切り出した文字パターンの矩形、
同図(f)に示した11は辞書メモリ4に記憶した認識辞
書で、11a〜11jはそれぞれカテゴリ「0」〜「9」の基
準パターンである。また、同図(g)に示した12は文字
パターンの矩形10aに対応する入力画像6の文字パター
ン、同図(h)に示した13は文字パターン12と辞書メモ
リ4に記憶した認識辞書11の各カテゴリ「0」〜「9」
の基準パターン11a〜11jとを重ね合わせて求めた第5位
までの候補カテゴリとその類似度で、13a,13b,13c,13d,
13eはそれぞれ基準パターン11d,11i,11j,11a,11fと文字
パターン12を重ね合わせた場合の類似度、同図(i)に
示した14は候補カテゴリとその類似度13から文字パター
ン12に対して決定した認識カテゴリ「3」である。
6A shown in FIG. 3A is an input image stored in the image memory 1, 6a is a character pattern of category "3", 6b is a character pattern of category "4", 6c is a character pattern of category "8", 6d is a character pattern of category "7", and 7 shown in FIG. 7 (b) is a rectangular representation of the position and size of the character which is the contents of the character arrangement information storage memory 2, and 7a, 7b, 7c. , 7d are the 1st, 2nd and 3rd characters respectively from the left.
A rectangle representing the position and size of the fourth character and the fourth character, and 8 shown in (c) of the figure is a table in which the contents of the character array information storage memory 2 are represented by numerical values, 8a, 8b, 8c, 8d. Is the position number (NO) of the rectangle 7a, 7b, 7c, 7d, and X at the upper left corner, respectively.
The coordinates (SX), the Y coordinate (SY) at the upper left corner, the width (W), and the height (H) are expressed by numerical values. On the other hand, 9 shown in (d) of the figure is a peripheral distribution created from the input image 6, 9a is a peripheral distribution in the Y direction, 9b is a peripheral distribution in the X direction, and 10 shown in (e) is the input image. Character pattern rectangles cut out from 6, 10a, 10b, 10c, 10d are rectangles 7a, 7 respectively.
The rectangle of the character pattern cut out corresponding to b, 7c, 7d,
Reference numeral 11 shown in FIG. 6F is a recognition dictionary stored in the dictionary memory 4, and reference numerals 11a to 11j are reference patterns of categories "0" to "9", respectively. Further, 12 shown in FIG. 7G is a character pattern of the input image 6 corresponding to the rectangle 10a of the character pattern, and 13 shown in FIG. 7H is a character pattern 12 and the recognition dictionary 11 stored in the dictionary memory 4. Each category “0” to “9”
The candidate categories up to the fifth rank obtained by superimposing the reference patterns 11a to 11j of
13e is the similarity when the reference patterns 11d, 11i, 11j, 11a, 11f and the character pattern 12 are overlapped, and 14 shown in FIG. It is the recognition category “3” determined by the above.

次に、動作について説明する。入力画像6は画像メモリ
1に記憶する。
Next, the operation will be described. The input image 6 is stored in the image memory 1.

文字切り出し手段3では、入力画像6をX方向及びY方
向に走査し、各カラム及び各ライン毎に黒画素を計数し
てX方向の周辺分布9b及びY方向の周辺分布9aを作成す
る。次に周辺分布9a及び9bをそれぞれ2値化し、X方向
に連続する黒部分の両端のX座標とY方向に連続する黒
部分の両端のY座標から文字パターンに外装する矩形10
a,10b,10c,10dの位置及び大きさを求め、予め設定して
おいた文字配列情報格納メモリ2の内容であるテーブル
8の数値(位置及び大きさ)と比較して(比率を求め
て)矩形7a,7b,7c,7dに文字パターンの矩形10a,10b,10
c,10dをそれぞれ対応付けることにより文字パターンの
矩形10を決定し、文字パターンの矩形10に対応した文字
パターンを入力画像6から切り出す。例えば、文字パタ
ーン12は文字パターンの矩形10aに対応する入力画像6
のパターンである。
The character clipping means 3 scans the input image 6 in the X and Y directions, counts black pixels for each column and each line, and creates a peripheral distribution 9b in the X direction and a peripheral distribution 9a in the Y direction. Next, the marginal distributions 9a and 9b are each binarized, and a rectangle 10 to be externally attached to the character pattern from the X coordinates of both ends of the black portion continuous in the X direction and the Y coordinates of both ends of the black portion continuous in the Y direction.
The positions and sizes of a, 10b, 10c, and 10d are calculated, and compared with the numerical values (position and size) of the table 8 which is the contents of the character array information storage memory 2 set in advance (the ratio is calculated. ) Rectangle 7a, 7b, 7c, 7d in the character pattern rectangle 10a, 10b, 10
The character pattern rectangle 10 is determined by associating c and 10d with each other, and the character pattern corresponding to the character pattern rectangle 10 is cut out from the input image 6. For example, the character pattern 12 is the input image 6 corresponding to the rectangle 10a of the character pattern.
Pattern.

次に文字認識手段5では、文字切り出し手段3で切り出
した文字パターン12と辞書メモリ4に記憶した認識辞書
11に格納された認識対象のカテゴリの基準パターン11a
〜11jとをそれぞれ重ね合わせ、両者の整合の度合(類
似度)を求める。
Next, in the character recognition means 5, the character pattern 12 cut out by the character cutting means 3 and the recognition dictionary stored in the dictionary memory 4 are used.
Reference pattern 11a of the recognition target category stored in 11
~ 11j are overlapped with each other and the degree of matching (similarity) between them is obtained.

この場合、例えば、入力文字パターンと基準文字パター
ンについて白画素を−1,黒画素を+1で表す2値行列を としたとき、類似度は次に示すSで表されるものであ
る。
In this case, for example, for input character patterns and reference character patterns, a binary matrix in which white pixels are −1 and black pixels are +1 Then, the similarity is represented by S shown below.

ここで、分子は2つの行列の要素の一致数から不一致数
を引いた値,分母Nは行列の全要素数である。尚、 はベクトル、P,Qはスカラーを表す。
Here, the numerator is a value obtained by subtracting the number of mismatches from the number of matches of the elements of the two matrices, and the denominator N is the total number of elements of the matrix. still, Is a vector and P and Q are scalars.

例えば、類似度13a,13b,13c,13d,13eは、それぞれ基準
文字パターン11d,11i,11j,11a,11fと文字パターン12を
重ね合わせた場合に上記式により算出した類似度Sの値
である。
For example, the similarities 13a, 13b, 13c, 13d, 13e are the values of the similarity S calculated by the above formula when the reference character patterns 11d, 11i, 11j, 11a, 11f and the character pattern 12 are superposed, respectively. .

次に、認識対象の全カテゴリ「0」〜「9」の中から類
似度が高いカテゴリを候補カテゴリとし、候補カテゴリ
とその類似度13から認識方式で定められた条件を満たす
もの,例えば類似度が最も高いカテゴリ「3」を認識カ
テゴリ14と決定する。
Next, a category having a high degree of similarity among all the recognition target categories “0” to “9” is set as a candidate category, and a category satisfying a condition defined by the recognition method from the candidate category and the degree of similarity 13, for example, the degree of similarity The highest category "3" is determined as the recognition category 14.

〔発明が解決しようとする問題点〕[Problems to be solved by the invention]

従来の文字読取装置は以上のように構成されているの
で、工業用カメラなどで撮像する方向や位置により、文
字が欠けたり,障害物に一部分が隠れたりする場合に
は、正確に文字を切り出せないという問題点があった。
又,文字を切り出した場合でも、欠けや隠れにより正し
く文字を認識することが困難であるといった問題点があ
った。
Since the conventional character reading device is configured as described above, if a character is missing or partially obscured by an obstacle depending on the direction and position of the image taken by an industrial camera, etc., the character can be accurately cut out. There was a problem that it did not exist.
Further, even when the character is cut out, there is a problem that it is difficult to correctly recognize the character due to lack or occlusion.

この発明は、上記のような問題点を解消する為になされ
たもので、文字の一部分が欠けたり隠れていても精度良
く文字を切り出して正しく文字を認識することが出来る
文字読取装置を得ることを目的とする。
The present invention has been made in order to solve the above problems, and provides a character reading device that can accurately cut out a character and correctly recognize the character even if a part of the character is missing or hidden. With the goal.

〔問題点を解決するための手段〕[Means for solving problems]

この発明に係る文字読取装置は、入力画像の画素の値に
基づいて明らかに文字である文字パターンを優先して切
り出し優先文字切り出し手段と、既知の文字配列情報と
優先文字切り出し手段で切り出した文字パターンの位置
及び大きさ情報に基づいて残りの文字の位置及び大きさ
を推定して文字パターンを切り出し推定文字切り出し手
段と、上記文字パターンから成る領域の画素の値に基づ
いて文字が欠けた又は隠れた領域を検出する欠け隠れ領
域検出手段と、欠け隠れ領域検出手段で検出した欠け隠
れ領域を除く領域において上記文字パターンと認識の対
象となる文字の基準パターンとを比較することにより文
字を認識する文字認識手段とを備えたものである。
The character reading device according to the present invention preferentially cuts out a character pattern which is obviously a character based on the value of a pixel of an input image, a priority character cutting means, and a character cut out by known character arrangement information and priority character cutting means. The position and size of the remaining characters are estimated based on the position and size information of the pattern to cut out the character pattern, and the estimated character cutting means and the character is missing based on the value of the pixel in the area formed by the character pattern, or Character recognition is performed by comparing the above-mentioned character pattern with the reference pattern of the character to be recognized in the area except for the missing hidden area detected by the missing hidden area detecting means and the missing hidden area detecting means for detecting the hidden area. And a character recognizing means for doing so.

〔作用〕[Action]

この発明における文字読取装置は、事前に予想される文
字の位置及び大きさ等の文字配列情報と既に優先文字切
り出し手段で切り出された文字パターンの情報から推定
して文字に欠け又は隠れがある文字を推定文字切り出し
手段で切り出し、且つ,欠け隠れの領域を検出するにこ
とより、文字認識手段は欠け隠れ領域を除いたパターン
を用いて文字を認識する。従って、撮像方向や位置の変
動等によって文字に欠けや隠れがある場合でも精度良く
文字を切り出すことが出来る。又、文字に欠けや隠れが
ある場合でも正確に文字を認識することが出来る。
The character reading device according to the present invention is a character in which a character is missing or hidden, which is estimated from character arrangement information such as a predicted position and size of a character in advance and information of a character pattern already cut out by the priority character cutting means. Is cut out by the estimated character cutting-out means and the area of the lacking occlusion is detected, so that the character recognizing means recognizes the character using the pattern excluding the lacking occlusion area. Therefore, it is possible to accurately cut out a character even when the character is missing or hidden due to a change in the imaging direction or the position. Further, even if the character is missing or hidden, the character can be accurately recognized.

〔実施例〕〔Example〕

以下、この発明の一実施例を図について説明する。 An embodiment of the present invention will be described below with reference to the drawings.

第1図は本実施例による文字読取装置の構成図である。
図において、1は認識対象の文字が記録された媒体をテ
レビカメラ等の撮像装置で撮像して得たディジタル画像
を記憶する画像メモリ、3aは画像メモリ1を走査して求
めたX方向及びY方向の周辺分布値に基づいて明らかに
文字である文字パターンを優先して切り出す優先文字切
り出し手段、3bは優先切り出し手段3aで切り出すことが
出来なかった文字パターンを既に切り出した文字パター
ンの位置及び大きさ情報と文字配列情報格納メモリ2の
内容とから推定して切り出す推定文字切り出し手段、15
は推定文字切り出し手段3bで切り出した文字パターンか
ら成る領域を走査して求めたY方向の周辺分布値に基づ
いて文字が欠けた又は隠れた領域を検出する欠け隠れ領
域検出手段、5aは欠け隠れ領域検出手段15で検出した領
域を除外した領域において文字パターンと認識の対象と
なる文字の基準パターンとを比較して文字を認識する文
字認識手段である。尚、上記各手段3a,3b,15,5aはマイ
クロプロセッサ等により実現されるものであり、又、各
メモリ2及び4は第3図に示した従来の文字読取装置と
同一のものである。
FIG. 1 is a block diagram of a character reading device according to this embodiment.
In the figure, 1 is an image memory for storing a digital image obtained by capturing an image of a medium on which a character to be recognized is recorded with an image capturing device such as a television camera, and 3a is an X direction and Y obtained by scanning the image memory 1. Priority character cutting means for preferentially cutting out a character pattern that is clearly a character based on the peripheral distribution value in the direction, 3b is the position and size of the character pattern already cut out of the character pattern that could not be cut out by the priority cutting means 3a. Estimated character cut-out means for cutting out by estimating from the size information and the contents of the character array information storage memory 2, 15
Is a missing hidden area detecting means for detecting an area where a character is missing or hidden based on a peripheral distribution value in the Y direction obtained by scanning an area formed by the character pattern cut out by the estimated character cutting means 3b, and 5a is a missing hidden area. It is a character recognition means for recognizing a character by comparing a character pattern with a reference pattern of a character to be recognized in an area excluding the area detected by the area detection means 15. The respective means 3a, 3b, 15, 5a are realized by a microprocessor or the like, and the memories 2 and 4 are the same as the conventional character reading device shown in FIG.

第2図(a)〜(n)は、第1図に示した実施例による
文字読取装置を動作例を示した図である。
FIGS. 2A to 2N are diagrams showing an operation example of the character reading device according to the embodiment shown in FIG.

同図(a)に示した16は画像メモリ1に記憶した入力画
像で、16a〜16hはそれぞれカテゴリ「3」,「4」,
「8」,「7」,「6」,「2」,「3」,「4」の文
字パターンであり、文字パターン16a及び16bは文字の上
部が隠れており、文字パターン16g,16hは文字の上部が
欠けている。同図(b)に示した17は文字配列情報格納
メモリ2の内容である文字の位置及び大きさを矩形で表
現したものであり、17a〜17dは1行目の左からそれぞれ
1〜4文字目の位置及び大きさを表現した矩形、17e〜1
7hは2行目の左からそれぞれ1〜4文字目の位置及び大
きさを表現した矩形、同図(c)に示した18は文字配列
情報格納メモリ2の内容を数値で表現したテーブルであ
り、18a〜18hはそれぞれ矩形17a〜17hの位置番号(N
O),左上端のX座標(SX)及びY座標(SY),幅
(W),高さ(H)を数値により表現したものである。
一方、同図(d)に示した19は入力画像6から作成した
周辺分布で、19a及び19bはY方向の周辺分布,19c及び19
dはX方向の周辺分布であり、Y方向の周辺分布19aとX
方向の周辺分布19cが1行目の文字パターンに対応し、
Y方向の周辺分布19bとX方向の周辺分布19dが2行目の
文字パターンに対応する。同図(e)に示した20は優先
文字切り出し手段3aで切り出した文字パターンの矩形
で、20a〜20dはそれぞれ矩形17c〜17fに対応する文字パ
ターンの矩形、同図(f)に示した21は推定文字切り出
し手段3bで切り出した文字パターンの矩形で、21a〜21d
はそれぞれ矩形17a,17b,17g,17hに対応させて切り出し
た文字パターンの矩形、同図(g)に示した22は推定文
字切り出し手段3bで切り出された文字パターンの矩形21
a〜21dから成る領域の入力画像16を走査して求めたY方
向の周辺分布で、22aは1行目の文字パターンの矩形21a
及び21bから成る領域を走査して求めたY方向の周辺分
布,22bは2行目の文字パターンの矩形21c及び21dから成
る領域を走査して求めたY方向の周辺分布、同図(h)
に示した23は周辺分布22の値に基づいて検出した欠け隠
れ領域で、23aは周辺分布22aの値に基づいて検出した隠
れ領域,23bは周辺分布22bの値に基づいて検出した欠け
領域である。
Reference numeral 16 shown in FIG. 3A is an input image stored in the image memory 1, and 16a to 16h are categories "3", "4",
The character patterns are "8", "7", "6", "2", "3", and "4". The character patterns 16a and 16b hide the upper part of the characters, and the character patterns 16g and 16h are characters. Is missing the top of the. Reference numeral 17 shown in FIG. 2B represents the position and size of the character, which is the content of the character array information storage memory 2, in a rectangle, and 17a to 17d are 1 to 4 characters from the left of the first line, respectively. Rectangle expressing eye position and size, 17e-1
7h is a rectangle representing the positions and sizes of the first to fourth characters from the left of the second line, and 18 shown in FIG. 7C is a table representing the contents of the character array information storage memory 2 by numerical values. , 18a-18h are the position numbers (N
O), the X coordinate (SX) and the Y coordinate (SY) of the upper left corner, the width (W), and the height (H) are expressed by numerical values.
On the other hand, 19 shown in (d) of the figure is a marginal distribution created from the input image 6, 19a and 19b are marginal distributions in the Y direction, 19c and 19b.
d is the marginal distribution in the X direction, and the marginal distribution 19a and X in the Y direction
The peripheral distribution 19c in the direction corresponds to the character pattern on the first line,
The peripheral distribution 19b in the Y direction and the peripheral distribution 19d in the X direction correspond to the character pattern on the second line. Reference numeral 20 shown in FIG. 7E is a rectangle of the character pattern cut out by the priority character cutting means 3a, reference numerals 20a to 20d are rectangles of character patterns corresponding to the rectangles 17c to 17f, and 21 shown in FIG. Is a rectangle of the character pattern cut out by the estimated character cutting-out means 3b.
Is a rectangle of a character pattern cut out corresponding to each of the rectangles 17a, 17b, 17g, and 17h, and 22 shown in FIG. 7G is a rectangle 21 of the character pattern cut out by the estimated character cut-out means 3b.
22a is the peripheral distribution in the Y direction obtained by scanning the input image 16 in the area consisting of a to 21d, where 22a is the rectangle 21a of the character pattern in the first line.
And 21b are the peripheral distributions in the Y direction obtained by scanning the area, and 22b is the peripheral distribution in the Y direction obtained by scanning the area composed of the rectangles 21c and 21d of the second line character pattern, FIG.
23 is a missing hidden area detected based on the value of the peripheral distribution 22, 23a is a hidden area detected based on the value of the peripheral distribution 22a, and 23b is a missing area detected based on the value of the peripheral distribution 22b. is there.

又、同図(i)に示した24は文字パターンの矩形21に対
応させて入力画像16から切り出した文字パターンであ
り、24a,24bはそれぞれ文字パターンの矩形21b,21dに対
応する。同図(j)に示した25は整合画素を値「1」,,
非整合画素を値「0」として文字パターン24毎に作成し
たマスクパターン(図中黒で示した部分が整合画素,白
で示した部分が非整合画素)で、25a,25bはそれぞれ文
字パターン24a,24bのマスクパターン、同図(k),
(l)に示した26,27は文字パターン24とマスクパター
ン25と辞書メモリ4に記憶された第4図(f)に示す認
識辞書11のカテゴリ「0」〜「9」の基準パターン11a
〜11jとを重ね合わせることにより求めた第5位までの
候補カテゴリとその類似度で、26は文字パターン24aと
マスクパターン25aに対する候補カテゴリとその類似度
であり、26a〜26eはそれぞれ基準パターン11e,11a,11g,
11c,11iとの類似度、27は文字パターン24bとマスクパタ
ーン25bに対する候補カテゴリとその類似度であり、27a
〜27eはそれぞれ基準パターン11e,11g,11c,11i,11aに対
する類似度である。そして、同図(m)に示した28は候
補カテゴリとその類似度26から文字パターン24aに対し
て決定した認識カテゴリ「4」、同図(n)に示した29
は候補カテゴリとその類似度27から文字パターン24bに
対して決定した認識カテゴリ「4」である。尚、辞書メ
モリ4に格納する認識辞書は第4図に示した従来の文字
読取装置の動作例中の認識辞書11と同一のものである。
Further, 24 shown in FIG. 9I is a character pattern cut out from the input image 16 in correspondence with the character pattern rectangle 21, and 24a and 24b correspond to the character pattern rectangles 21b and 21d, respectively. In FIG. 25 (j), 25 is the matching pixel with the value "1",
A mask pattern is created for each character pattern 24 with the non-matching pixel value "0" (the black part in the figure is the matching pixel and the white part is the non-matching pixel), and 25a and 25b are the character patterns 24a, respectively. , 24b mask pattern, FIG.
26 and 27 shown in (l) are reference patterns 11a of the character patterns 24, the mask patterns 25 and the categories "0" to "9" of the recognition dictionary 11 shown in FIG.
~ 11j are the candidate categories up to the fifth rank and their similarity obtained by superimposing them with each other, 26 is the candidate category and its similarity to the character pattern 24a and the mask pattern 25a, and 26a to 26e are reference patterns 11e, respectively. , 11a, 11g,
11c, 11i similarity, 27 is a candidate category and its similarity to the character pattern 24b and the mask pattern 25b, 27a
27e are the similarities to the reference patterns 11e, 11g, 11c, 11i, and 11a, respectively. 28 is a recognition category “4” determined for the character pattern 24a from the candidate category and its similarity 26, and 29 shown in FIG.
Is the recognition category "4" determined for the character pattern 24b from the candidate category and its similarity 27. The recognition dictionary stored in the dictionary memory 4 is the same as the recognition dictionary 11 in the operation example of the conventional character reading apparatus shown in FIG.

次に、動作について説明する。Next, the operation will be described.

テレビカメラ等の撮像装置で撮像し、光電変換して得ら
れた入力画像16を画像メモリ1に記憶する。
An input image 16 obtained by picking up an image with an image pickup device such as a television camera and performing photoelectric conversion is stored in the image memory 1.

優先文字切り出し手段3aでは、入力画像16を走査し,各
ライン毎の黒画素を計数してY方向の周辺分布19a,19b
を作成する。そして、周辺分布19a,19bを2値化し、Y
方向に連続する黒部分の両端のY座標を求める。図の例
では連続する黒部分が2箇所あり、Y座標の対が2組抽
出される。次に連続する黒部分毎にY座標の対で定まる
範囲の入力画像16を走査し、各カテゴリ毎の黒画素を計
数してX方向の周辺分布19c,19dを求める。X方向の周
辺分布19cはY方向の周辺分布19aから求まるY座標の対
に対する周辺分布であり、X方向の周辺分布19dはY方
向の周辺分布19bから求まるY座標の対に対する周辺分
布である。そして、次にX方向の周辺分布19c,19dを2
値化し、連続する黒部分の両端のX座標を求める。図中
破線が2値化閾値を表すものとすると、周辺分布19cに
対しては3組のX座標の対が抽出され、周辺分布19dに
対しては2組のX座標の対が抽出される。次に、Y方向
の周辺分布19aに対するY座標の対とX方向の周辺分布1
9cに対するX座標の対を組合わせてパターンに外装する
矩形を求める。この矩形の位置及び大きさが所定の範囲
にある文字パターンの矩形20a,20bを優先して切り出
す。Y方向の周辺分布19bとX方向の周辺分布19dに対し
ても同様にして文字パターンの矩形20c,20dを切り出
す。
The priority character cutting means 3a scans the input image 16, counts the black pixels in each line, and distributes the margins 19a and 19b in the Y direction.
To create. Then, the marginal distributions 19a and 19b are binarized, and Y
The Y coordinates of both ends of the black portion continuous in the direction are obtained. In the illustrated example, there are two continuous black portions, and two pairs of Y coordinates are extracted. Next, the input image 16 in the range defined by the pair of Y coordinates is scanned for each continuous black portion, and the black pixels of each category are counted to obtain the peripheral distributions 19c and 19d in the X direction. The X-direction marginal distribution 19c is a marginal distribution for the pair of Y coordinates obtained from the Y-direction marginal distribution 19a, and the X-direction marginal distribution 19d is a marginal distribution for the Y-coordinate pair obtained from the Y-direction marginal distribution 19b. Then, the peripheral distributions 19c and 19d in the X direction are set to 2
It is binarized and the X coordinates of both ends of the continuous black portion are obtained. If the broken line in the figure represents a binarization threshold, three pairs of X coordinates are extracted for the marginal distribution 19c, and two pairs of X coordinates are extracted for the marginal distribution 19d. . Next, the pair of Y coordinates for the peripheral distribution 19a in the Y direction and the peripheral distribution 1 in the X direction 1
A pair of X coordinates for 9c is combined to find a rectangle to be attached to the pattern. The rectangles 20a and 20b of the character pattern whose position and size are within a predetermined range are preferentially cut out. The rectangles 20c and 20d of the character pattern are similarly cut out for the peripheral distribution 19b in the Y direction and the peripheral distribution 19d in the X direction.

推定文字切り出し手段3bでは、優先文字切り出し手段3a
で切り出した文字パターンの矩形20a〜20dの位置及び大
きさと文字配列情報格納メモリ2の内容である矩形17a
〜17hの位置及び大きさ18a〜18hとの比率を比較するこ
とにより文字パターンの矩形20a〜20dをそれぞれ矩形17
c〜17fに対応させ、残りの矩形17a,17b,17g,17hに対し
ては、それぞれ位置的に最も近い文字パターンの矩形20
c,20a,20d,20bの位置及び大きさに対する上記比率を掛
けることにより文字パターンの矩形21a〜21dの位置及び
大きさを推定して切り出す。
The estimated character cutout means 3b has a priority character cutout means 3a.
Positions and sizes of the rectangles 20a to 20d of the character pattern cut out in step 17 and the rectangle 17a which is the contents of the character array information storage memory 2
~ 17h and size 18a ~ 18h by comparing the ratio with the size of the rectangles 20a ~ 20d of the character pattern rectangle 17a.
For the remaining rectangles 17a, 17b, 17g, and 17h corresponding to c to 17f, the rectangle 20 of the character pattern that is closest in position respectively
The positions and sizes of the rectangles 21a to 21d of the character pattern are estimated and cut out by multiplying the positions and sizes of c, 20a, 20d, and 20b by the above ratios.

この矩形21a〜21dも、文字配列情報格納メモリ2の内容
である矩形17a〜17hとの比率は優先文字切り出し手段3a
で切り出された矩形20a〜20dと同じであるので、文字パ
ターンにほぼ外装するものとなる。
The ratio of the rectangles 21a to 21d to the rectangles 17a to 17h, which are the contents of the character array information storage memory 2, is the priority character cutting means 3a.
Since it is the same as the rectangles 20a to 20d cut out in, it is almost put on the character pattern.

欠け隠れ領域検出手段15では、推定文字切り出し手段3b
で求めた文字パターンの矩形21a〜21dに対して、同一行
にある矩形21a,21bから成る領域を走査し、各ライン毎
に黒画素を計数してY方向の周辺分布22aを求める。矩
形21c,21dに対しても同様にしてY方向の周辺分布22bを
求める。次に周辺分布22a,22bをそれぞれ所定の閾値で
2値化し、欠け領域23b,隠れ領域23aの境界線のY座標
を求める。上記閾値には欠け検出用と隠れ検出用の2種
類存在するが、周辺分布19の値に基づいて求めた外接矩
形の形状により何れか一方を用いる。図の例では、周辺
分布22aに対しては隠れ検出用が,周辺分布22bに対して
は欠れ検出用がそれぞれ採用される。
In the missing part detection means 15, the estimated character cutting means 3b
With respect to the rectangles 21a to 21d of the character pattern obtained in step 1, the area consisting of the rectangles 21a and 21b in the same row is scanned, and black pixels are counted for each line to obtain the peripheral distribution 22a in the Y direction. The peripheral distribution 22b in the Y direction is similarly obtained for the rectangles 21c and 21d. Next, the marginal distributions 22a and 22b are binarized with predetermined threshold values, and the Y coordinate of the boundary line between the missing area 23b and the hidden area 23a is obtained. There are two types of threshold values, one for missing detection and the other for hidden detection, but either one is used depending on the shape of the circumscribed rectangle obtained based on the value of the peripheral distribution 19. In the example shown in the figure, the hidden detection is used for the peripheral distribution 22a, and the missing detection is used for the peripheral distribution 22b.

文字認識手段5aでは、優先文字切り出し手段3a及び推定
文字切り出し手段3bで切り出した矩形20及び21に対する
入力画像16の文字パターンに対して,欠け隠れ領域23と
重ならないものは全ての画素の値が「1」であるマスク
パターンを作成し、それ以外のものは上記境界線のY座
標に基づいて欠け隠れ領域と重なる画素の値を「0」,
それ以外の画素の値を「1」としたマスクパターンを作
成する。例えば、文字パターンの矩形21b,21dに関して
は、それぞれ対応する文字パターン24a,24bに対してマ
スクパターン25a,25bが作成される。マスクパターン25a
は隠れ領域をマスクしたものであり、マスクパターン25
bは欠け領域をマスクしたものである。次に、文字パタ
ーン24aとマスクパターン25aと辞書メモリ4に記憶され
た第4図に示す認識辞書11の各カテゴリの基準パターン
11a〜11jとをそれぞれ重ね合わせることにより類似度を
求める。文字パターン24bとマスクパターン25bに対して
も同様にして類似度を求める。
In the character recognition means 5a, with respect to the character patterns of the input image 16 for the rectangles 20 and 21 cut out by the priority character cutout means 3a and the estimated character cutout means 3b, the values of all the pixels that do not overlap with the lacking hidden area 23 are A mask pattern of "1" is created, and other than that, the value of the pixel overlapping with the missing hidden area is set to "0", based on the Y coordinate of the boundary line.
A mask pattern in which the values of the other pixels are "1" is created. For example, for rectangles 21b and 21d of character patterns, mask patterns 25a and 25b are created for the corresponding character patterns 24a and 24b, respectively. Mask pattern 25a
Is a mask of the hidden area. Mask pattern 25
b is a mask of the missing area. Next, the character pattern 24a, the mask pattern 25a, and the reference pattern of each category of the recognition dictionary 11 shown in FIG.
The degree of similarity is obtained by superimposing 11a to 11j on each other. Similarity is similarly obtained for the character pattern 24b and the mask pattern 25b.

この場合、例えば、入力文字パターンについて整合画素
を1,非整合画素を0で表す2値行列を は従来の文字読取装置の類似度計算式と同一のものとし
たとき、類似度は次に示すSで表されるものである。
In this case, for example, for the input character pattern, a binary matrix in which matching pixels are 1 and non-matching pixels are 0 Is the same as the similarity calculation formula of the conventional character reading device, the similarity is represented by S shown below.

ここで、分子は の要素が1である位置での の要素の一致数から不一致数を引いた値、分母は の値1の要素数である。即ち、分子はマスクパターンが
1である画素に関して、入力文字パターンと基準文字パ
ターンが一致する画素数から不一致となる画素数を引い
た値となり、分母はマスクパターンの値が1である画素
のみがカウントされる。
Where the numerator is At the position where the element of is 1 The value obtained by subtracting the number of mismatches from the number of matches of the elements of, the denominator is Is the number of elements with a value of 1. That is, the numerator is a value obtained by subtracting the number of pixels that do not match from the number of pixels that the input character pattern and the reference character pattern match with respect to the pixel whose mask pattern is 1, and the denominator is only the pixels that have a mask pattern value of 1. Is counted.

例えば、類似度26a〜26eはそれぞれ基準パターン11e,11
a,11g,11c,11iと文字パターン24aとマスクパターン25a
とを重ね合わせた場合に上記式により算出した類似度S
の値である。同様に、類似度27a〜27eはそれぞれ基準パ
ターン11e,11g,11c,11i,11aと文字パターン24bとマスク
パターン25bを重ね合わせた場合の類似度である。
For example, the similarities 26a to 26e are the reference patterns 11e and 11e, respectively.
a, 11g, 11c, 11i and character pattern 24a and mask pattern 25a
Similarity S calculated by the above equation when and are superposed
Is the value of. Similarly, the similarities 27a to 27e are the similarities when the reference patterns 11e, 11g, 11c, 11i, 11a, the character pattern 24b, and the mask pattern 25b are superimposed.

次に、認識対象の全カテゴリ「0」〜「9」の中から類
似度の高いカテゴリを候補カテゴリとし、候補カテゴリ
とその類似度26から認識方式で定められた条件を満たす
もの,例えば類似度が最も高いカテゴリ「4」を認識カ
テゴリ28と決定する。同様に、候補カテゴリと類似度27
に対してはカテゴリ「4」を認識カテゴリ29と決定す
る。従って、文字パターン24aの認識結果も文字パター
ン24bの認識結果もカテゴリ「4」となる。
Next, a category with a high degree of similarity is selected as a candidate category from all the recognition target categories “0” to “9”, and a category satisfying the condition defined by the recognition method from the candidate category and the degree of similarity 26, for example, the degree of similarity The highest category “4” is determined as the recognition category 28. Similarly, the candidate category and similarity 27
For the above, the category “4” is determined as the recognition category 29. Therefore, the recognition result of the character pattern 24a and the recognition result of the character pattern 24b are in category "4".

尚、上記実施例では、欠け隠れ領域をY方向についての
み求めたが、X方向及びY方向について求めても良い。
Incidentally, in the above-mentioned embodiment, the chipped hidden area is obtained only in the Y direction, but it may be obtained in the X direction and the Y direction.

又、欠け隠れ領域を検出する走査範囲は推定文字切り出
し手段3bで切り出された文字パターンから成る領域とし
て説明したが、切り出された全ての文字パターンから成
る領域としても良いし、推定文字切り出し手段3bで切り
出された個々の文字パターンを走査範囲としても良い。
Further, the scanning range for detecting the lacking hidden area has been described as an area formed of the character patterns cut out by the estimated character cutting means 3b, but it may be an area formed of all the cut out character patterns, or the estimated character cutting means 3b. The individual character patterns cut out in step may be used as the scanning range.

更に、文字認識方法はパターンを重ね合わせるパターン
マッチング法を用いた場合について説明したが、パター
ンから抽出した特徴を用いる方法を用いても良い。
Further, as the character recognition method, the case of using the pattern matching method of overlapping patterns has been described, but a method using a feature extracted from the pattern may be used.

〔発明の効果〕〔The invention's effect〕

以上の様に、この発明によれば、優先文字切り出し手段
と推定文字切り出し手段とを設けたことにより、撮像方
向及び位置の変動等により文字に欠けや隠れがある場合
でも精度良く文字を切り出すことが出来る。又、欠け隠
れ領域検出手段と欠け隠れ領域を除外して認識する文字
認識手段を設けたことにより、文字に欠けや隠れがある
場合でも正確に文字認識が出来るといった効果がある。
As described above, according to the present invention, by providing the priority character cutting-out means and the estimated character cutting-out means, it is possible to accurately cut out a character even if the character is missing or hidden due to a change in the imaging direction and position. Can be done. Further, by providing the missing and hidden area detecting means and the character recognizing means for recognizing the missing hidden area, there is an effect that the character can be accurately recognized even when there is a missing or hidden character.

【図面の簡単な説明】[Brief description of drawings]

第1図はこの発明の一実施例による文字読取装置の構成
図、第2図(a)〜(n)は第1図に示す装置の動作例
を示す図、第3図は従来の文字読取装置の構成図、第4
図(a)〜(i)は第3図に示す装置の動作例を示す図
である。 1は画像メモリ、2は文字配列情報格納メモリ、3aは優
先文字切り出し手段、3bは推定文字切り出し手段、5aは
文字認識手段、15は欠け隠れ領域検出手段。 尚、図中、同一符号は同一,又は相当部分を示す。
FIG. 1 is a block diagram of a character reading device according to an embodiment of the present invention, FIGS. 2 (a) to (n) are diagrams showing an operation example of the device shown in FIG. 1, and FIG. 3 is a conventional character reading device. Device configuration diagram, 4th
(A)-(i) is a figure which shows the operation example of the apparatus shown in FIG. Reference numeral 1 is an image memory, 2 is a character array information storage memory, 3a is a priority character cutting means, 3b is an estimated character cutting means, 5a is a character recognition means, and 15 is a missing area detection means. In the drawings, the same reference numerals indicate the same or corresponding parts.

Claims (1)

【特許請求の範囲】[Claims] 【請求項1】文字配列が既知である個々の文字を入力画
像から切り出して認識することにより文字を読取る文字
読取装置において、入力画像の画素の値に基づいて明ら
かに文字である文字パターンを優先して切り出す優先文
字切り出し手段と、上記文字配列の情報と優先文字切り
出し手段で切り出した文字パターンの位置及び大きさ情
報に基づいて残りの文字の位置及び大きさを推定して文
字パターンを切り出す推定文字切り出し手段と、上記文
字パターンから成る領域の画素の値に基づいて文字が欠
けた又は隠れた領域を検出する欠け隠れ領域検出手段
と、欠け隠れ領域検出手段で検出した欠け隠れ領域を除
く領域において上記文字パターンと認識の対象となる文
字の基準パターンとを比較することにより文字を認識す
る文字認識手段とを備えたことを特徴とする文字読取装
置。
1. In a character reading device for reading a character by cutting out and recognizing individual characters whose character arrangement is known from an input image, a character pattern which is obviously a character is prioritized based on a pixel value of the input image. Estimating the character pattern by estimating the position and size of the remaining characters based on the information of the character array and the position and size information of the character pattern cut out by the priority character cutting device. Character cut-out means, a missing hidden area detecting means for detecting an area where a character is missing or hidden based on the pixel value of an area consisting of the character pattern, and an area excluding the missing hidden area detected by the missing hidden area detecting means A character recognition means for recognizing a character by comparing the above-mentioned character pattern with a reference pattern of a character to be recognized. Character reader, characterized in that was e.
JP62330959A 1987-12-26 1987-12-26 Character reader Expired - Lifetime JPH0746371B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP62330959A JPH0746371B2 (en) 1987-12-26 1987-12-26 Character reader

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP62330959A JPH0746371B2 (en) 1987-12-26 1987-12-26 Character reader

Publications (2)

Publication Number Publication Date
JPH01171082A JPH01171082A (en) 1989-07-06
JPH0746371B2 true JPH0746371B2 (en) 1995-05-17

Family

ID=18238302

Family Applications (1)

Application Number Title Priority Date Filing Date
JP62330959A Expired - Lifetime JPH0746371B2 (en) 1987-12-26 1987-12-26 Character reader

Country Status (1)

Country Link
JP (1) JPH0746371B2 (en)

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP5561102B2 (en) * 2009-12-14 2014-07-30 富士通株式会社 Character recognition device, character recognition program, and character recognition method
JP5908825B2 (en) * 2012-11-01 2016-04-26 日本電信電話株式会社 Character recognition device and computer-readable recording medium on which character recognition program is recorded

Also Published As

Publication number Publication date
JPH01171082A (en) 1989-07-06

Similar Documents

Publication Publication Date Title
JP2940936B2 (en) Tablespace identification method
US5018216A (en) Method of extracting a feature of a character
JP2000242798A (en) Extraction of feature quantity of binarty image
JPH07168910A (en) Document layout analysis device and document format identification device
JPH01171082A (en) Character reader
JP2795860B2 (en) Character recognition device
JPH07120392B2 (en) Character pattern cutting device
JP2803224B2 (en) Pattern matching method
JP2708604B2 (en) Character recognition method
JP3566738B2 (en) Shaded area processing method and shaded area processing apparatus
JPS622382A (en) Image processing method
JP4439054B2 (en) Character recognition device and character frame line detection method
JP3127413B2 (en) Character recognition device
JPS6324482A (en) Hole extraction method
JPH04311283A (en) Line direction discriminating device
JP2963532B2 (en) Line direction determination device
JPH03172983A (en) Table processing method
JPH09167205A (en) Method for picking-up character pattern
HK40011374A (en) Two-dimensional code identification method, device and equipment
HK40011374B (en) Two-dimensional code identification method, device and equipment
JPH06274692A (en) Character extractor
JPS62278686A (en) Extracting method for hole
JPH0354384B2 (en)
JPH06274690A (en) Character recognition device and optical character reader
JPH0199189A (en) Character segmenting system

Legal Events

Date Code Title Description
EXPY Cancellation because of completion of term
FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20080517

Year of fee payment: 13