JPH0517598B2 - - Google Patents
Info
- Publication number
- JPH0517598B2 JPH0517598B2 JP61285877A JP28587786A JPH0517598B2 JP H0517598 B2 JPH0517598 B2 JP H0517598B2 JP 61285877 A JP61285877 A JP 61285877A JP 28587786 A JP28587786 A JP 28587786A JP H0517598 B2 JPH0517598 B2 JP H0517598B2
- Authority
- JP
- Japan
- Prior art keywords
- character
- characters
- binary data
- frame
- compression ratio
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired - Lifetime
Links
Landscapes
- Character Input (AREA)
Description
【発明の詳細な説明】
(産業上の利用分野)
本発明は光学式文字読取装置等の文字読取装置
における文字切出し方法に関するものである。DETAILED DESCRIPTION OF THE INVENTION (Field of Industrial Application) The present invention relates to a method for cutting out characters in a character reading device such as an optical character reading device.
(従来の技術)
一般的な光学式文字読取装置(OCR)のブロ
ツク図を第5図aに示す。OCRは、光電変換部
11、多値データバツフア12、フイルタ回路1
3及び認識部14を備え、認識部14は2値デー
タバツフア14a及び認識制御回路14bから構
成される。帳票上の文字部分は光電変換部11に
よつて、光電変換されて複数階調のデイジタル信
号に変換された後、多値データとして多値データ
バツフア12に格納される。格納された多値デー
タはフイルタ回路13により「0」と「1」に2
値化して平滑化された後、2値データとして認識
部14の2値データバツフア14aに格納され
る。このとき、2値データバツフア14aには、
後述するように、文字の大きさの正規化のため、
予め定められた圧縮率に応じて間引きした2値デ
ータが格納される。2値データバツフア14aに
格納された2値データを基に、認識制御部14b
で文字の認識が行われる。なお、以下、2値デー
タの「0」を白点、「1」を黒点という。(Prior Art) A block diagram of a general optical character reader (OCR) is shown in FIG. 5a. OCR includes a photoelectric conversion section 11, a multi-value data buffer 12, and a filter circuit 1.
3 and a recognition section 14, and the recognition section 14 is composed of a binary data buffer 14a and a recognition control circuit 14b. The character portion on the form is photoelectrically converted into a multi-gradation digital signal by the photoelectric converter 11, and then stored in the multi-value data buffer 12 as multi-value data. The stored multi-value data is divided into "0" and "1" by the filter circuit 13.
After being digitized and smoothed, it is stored in the binary data buffer 14a of the recognition unit 14 as binary data. At this time, in the binary data buffer 14a,
As explained later, in order to normalize the font size,
Binary data thinned out according to a predetermined compression rate is stored. Based on the binary data stored in the binary data buffer 14a, the recognition control unit 14b
Character recognition is performed. Note that, hereinafter, "0" in the binary data is referred to as a white point, and "1" is referred to as a black point.
文字の認識に先立ち、2値データバツフア14
a内の2値データに対し、認識の対象となる範囲
である認識対象範囲を決める作業があり、これを
文字切出しという。 Prior to character recognition, binary data buffer 14
For the binary data in a, there is a task of determining the recognition target range, which is the range to be recognized, and this is called character extraction.
2値データバツフア14aの内部構成を第5図
bに示す。2値データバツフア14aは、2値デ
ータを格納するデータバツフア141、データバ
ツフア141に格納された2値データのうち文字
パタン部分(即ち黒点の集まり)をx方向に投影
した結果(投影像)を格納する投影バツフア14
2、及び同様にy方向に投影した結果を格納する
投影バツフア143から構成される。 The internal configuration of the binary data buffer 14a is shown in FIG. 5b. The binary data buffer 14a is a data buffer 141 that stores binary data, and a projection that stores the result (projection image) of projecting a character pattern portion (i.e., a collection of black dots) in the x direction of the binary data stored in the data buffer 141. Batsuhua 14
2, and a projection buffer 143 that similarly stores the results of projection in the y direction.
また、2値データバツフア14aに格納される
文字パタン例を第6図a,bに示す。 Further, examples of character patterns stored in the binary data buffer 14a are shown in FIGS. 6a and 6b.
次に、第6図を参照して認識部14の認識制御
回路14bの制御による文字切出し手順を説明す
る。第6図aにおいて、文字パタン20,21,
22がデータバツフア141の2文字分に相当す
る領域23に格納されており、これらをy方向に
投影した投影像24,25,26は投影バツフア
143に格納される。これらの投影像24,2
5,26の切れ目から投影像25に該当する範囲
を1文字の巾として見なすことができる。更に、
第6図bに示すように、文字パタン21をx方向
に投影した投影像27が投影バツフア142に格
納される。この投影像27に該当する範囲を文字
の高さとしてみなされる。従つて、投影像25,
27で囲まれる矩形領域28を文字枠と称する。
この文字枠は認識対象範囲内の1文字パタンの外
接枠となつており、文字の大きさを示すものであ
る(以下、文字の大きさとは文字枠の高さ、巾を
合わせたものとして述べる)。 Next, a character extraction procedure under the control of the recognition control circuit 14b of the recognition section 14 will be described with reference to FIG. In FIG. 6a, character patterns 20, 21,
22 is stored in an area 23 corresponding to two characters in the data buffer 141, and projected images 24, 25, and 26 obtained by projecting these in the y direction are stored in the projection buffer 143. These projected images 24,2
The range corresponding to the projected image 25 from the breaks 5 and 26 can be regarded as the width of one character. Furthermore,
As shown in FIG. 6b, a projected image 27 of the character pattern 21 projected in the x direction is stored in the projection buffer 142. The range corresponding to this projected image 27 is regarded as the height of the character. Therefore, the projected image 25,
A rectangular area 28 surrounded by 27 is called a character frame.
This character frame serves as a circumscribing frame for a single character pattern within the recognition target range, and indicates the size of the character (hereinafter, character size is defined as the sum of the height and width of the character frame. ).
次に、前述の文字切出しの際の文字の大きさの
正規化について説明する。 Next, the normalization of the character size during the above-mentioned character extraction will be explained.
まず、2値データの間引き例を第7図に示す。
同図において、多値データバツフア12内の1
1,12,…17,…,67はアドレス番号を示
し、各アドレスには多値データが格納されてい
る。また、2値データバツフア14a内の11,
13,15,17,…,57は2値データバツフ
ア14aに格納された2値データに対応する多値
データの多値データバツフア12内の格納位置を
示すものである。同図はフイルタ回路13で2値
化した2値データを1つ置きに間引いて2値デー
タバツフア14aに格納した状態を示す。 First, an example of thinning out binary data is shown in FIG.
In the figure, 1 in the multilevel data buffer 12
1, 12, . . . 17, . . . , 67 indicate address numbers, and multi-value data is stored in each address. 11 in the binary data buffer 14a,
13, 15, 17, . . . , 57 indicate storage positions in the multi-value data buffer 12 of multi-value data corresponding to the binary data stored in the binary data buffer 14a. This figure shows a state in which the binary data that has been binarized by the filter circuit 13 is thinned out every other data and stored in the binary data buffer 14a.
このように、1つ置きに間引く、あるいは2つ
置きに間引くといつた間引きの割合は予め帳票フ
オーマツトで決められており、その決定方法は以
下の方法で行う。 In this way, the rate of thinning, such as thinning out every other item or every two items, is determined in advance by the form format, and the method for determining it is as follows.
帳票上の文字を記入する位置は、通常文字を1
つずつ書き込む記入枠で決められている。記入枠
の大きさは、複数種類用意されており、書かれる
文字の大きさは、記入枠の大きさに応じて変化す
ることを想定し、正規化する必要がある。具体的
には、正規化は記入枠を文字の大きさが一定の大
きさに書かれるべきとして定めた特定の大きさ
(以下、標準パタンと称する)の枠に変換する圧
縮・拡大の処理であり、標準パタンの高さ、巾に
記入枠の高さ、巾を合わせる為に、どのぐらい圧
縮するかという圧縮率を決め、この圧縮率を間引
きの割合として、フオーマツトに入れる。 The position where you write the characters on the form is usually 1 character.
It is determined by the entry frame for each entry. Multiple types of entry frame sizes are available, and it is necessary to normalize the size of written characters assuming that they will change depending on the size of the entry frame. Specifically, normalization is a compression/enlargement process that converts a writing frame into a frame of a specific size (hereinafter referred to as a standard pattern), which is determined so that the font size should be written at a constant size. Yes, in order to match the height and width of the standard pattern with the height and width of the entry frame, determine the compression rate of how much to compress, and enter this compression rate as the thinning rate in the format.
第8図a,bは、記入枠による正規化を説明す
る図であつて、31,32は記入枠、33は高さ
33a、巾33bを持つ標準パタンである。記入
枠31,32の各々に対してx方向及びy方向に
圧縮率を持たせる。 FIGS. 8a and 8b are diagrams for explaining normalization using a writing frame, and 31 and 32 are writing frames, and 33 is a standard pattern having a height 33a and a width 33b. A compression ratio is given to each of the entry frames 31 and 32 in the x direction and the y direction.
(発明が解決しようとする問題点)
しかしながら、以上述べた従来の文字切出し方
法における帳票の記入枠による正規化では次のよ
うな問題点がある。(Problems to be Solved by the Invention) However, the normalization using the form entry frame in the conventional character extraction method described above has the following problems.
記入枠の大きさによつて、圧縮率が定められる
ため、実際に書かれた文字の大きさの正規化に寄
与するものではない。例えば、第8図bに示すよ
うに、実際に書かれた文字34が記入枠31に標
準パタン33より小さく(即ち、記入枠31に対
して非常に小さく)書かれ、記入枠31を標準パ
タン33に合わせるように圧縮すると、文字34
自体の大きさが小さくなり、従つて、切出される
文字が小さくなるので、文字を正確に認識するこ
とが困難となる。 Since the compression rate is determined by the size of the entry frame, it does not contribute to normalizing the size of the actually written characters. For example, as shown in FIG. 8b, an actually written character 34 is written in the writing frame 31 smaller than the standard pattern 33 (that is, very small compared to the writing frame 31), and the writing frame 31 is written in the standard pattern. When compressed to fit 33, the character 34
Since the size of the cutout itself becomes smaller and the cut out characters become smaller, it becomes difficult to accurately recognize the characters.
本発明は、以上述べた問題点を除去し、実際に
書かれた文字の大きさに応じた正規化により正確
に文字認識することが可能な文字切出し方法を提
供することを目的とする。 SUMMARY OF THE INVENTION An object of the present invention is to provide a character extraction method that eliminates the above-mentioned problems and allows accurate character recognition by normalization according to the size of the actually written character.
(問題点を解決するための手段)
本発明は前記問題点を解決するために、帳票上
の文字を読取つて得られた2値データを正規化
し、正規化した2値データに対し文字パタンの外
接矩形領域である文字枠を求めて文字の切出しを
行う文字読取装置の文字切出し方法において、帳
票上の文字を読取つて得られた2値データを、帳
票のフオーマツトの記入枠と標準パタンにより第
1の圧縮率を求めるとともに、前記第1の圧縮率
で正規化された2値データに対して文字パタンの
文字枠を求める第1のステツプと、前記第1のス
テツプで求めた文字枠の大きさと標準パタンの大
きさに基づいて第2の圧縮率を求める第2のステ
ツプと、前記文字枠内の黒点数と文字枠面積との
比及び前記文字パタンの包括線内の黒点数と包括
線内面積との比に基づいて、第2の圧縮率を補正
するか否かを決定する第3のステツプと、第3の
ステツプの決定に基づいて、高さ方向又は横方向
の圧縮率のいずれか小さい方に揃えるように補正
した第2の圧縮率又は補正しない第2の圧縮率に
応じて第1の圧縮率を補正して第3の圧縮率を決
定する第4のステツプと、帳票上の文字を読取つ
て得られた前記2値データを第3の圧縮率に応じ
た間引きにより正規化して文字の切出しを再度行
う第5のステツプとを有するものである。(Means for Solving the Problems) In order to solve the above-mentioned problems, the present invention normalizes binary data obtained by reading characters on a form, and adds character patterns to the normalized binary data. In the character extraction method of a character reading device, which extracts characters by finding a character frame, which is a circumscribed rectangular area, the binary data obtained by reading characters on a form is extracted using the entry frame of the form format and a standard pattern. A first step in which a compression rate of 1 is determined, and a character frame of a character pattern is determined for the binary data normalized by the first compression rate, and the size of the character frame determined in the first step is a second step of calculating a second compression ratio based on the size of the character pattern and the size of the standard pattern; A third step of determining whether or not to correct the second compression ratio based on the ratio to the inner area; and determining whether the compression ratio in the height direction or the lateral direction is corrected based on the determination in the third step. a fourth step of determining a third compression ratio by correcting the first compression ratio according to the second compression ratio corrected to be equal to the smaller one or the second compression ratio that is not corrected; and a fifth step of normalizing the binary data obtained by reading the characters by thinning out according to a third compression ratio and cutting out the characters again.
(作用)
本発明の文字読取装置の文字切出し方法は次の
ように作用する。第1のステツプでは、従来と同
様に、帳票フオーマツトで定められる第1の圧縮
率(例えば帳票上の記入枠の大きさと標準パタン
の大きさによる)に応じて、2値データを正規化
した2値データ(間引きによる2値データ)に対
して文字枠を得る。第2のステツプでは、このよ
うにして得られた文字枠の大きさと標準パタンの
大きさに基づいて第3の圧縮率が求められる。第
2のステツプでは、文字枠内の黒点数と文字枠面
積との比及び包括線(凸開包)内の黒点数と包括
線内との比に基づいて、第2の圧縮率を補正する
か否かが決定される。例えば、文字枠内及び包括
線内にほとんど白点がないような文字パタンは、
第2の圧縮率は補正するように決定される。第4
のステツプでは第3のステツプの決定に基づいて
補正した第2の圧縮率(例えば、縦長、横長の文
字パタンは高さ、巾の圧縮率とも小さい方の圧縮
率に揃えられる)、又は第2の圧縮率に応じて第
1の圧縮率を補正して第3の圧縮率が決定される
(例えば、補正した第2の圧縮率又は第2の圧縮
率と第1の圧縮率との掛算により決定される)。
第5のステツプでは第3の圧縮率に応じて帳票上
の文字を読取つて得られた2値データを間引きす
る正規化が行われ、再度の文字切出しが行われ
る。このように、文字枠の大きさによる第2の圧
縮率だけでなく、更に文字の形状(抱括線によ
る)に応じて第2の圧縮率を補正して従来の帳票
フオーマツトによる第1の圧縮率を補正すること
により、実際に書かれた文字に忠実な第3の圧縮
率を決定し、この第3の圧縮率により正規化を行
つているので、最適な文字切出しを行うことがで
き、かつ文字を正確に認識することが可能とな
る。従つて、前記従来技術の問題点を解決できる
のである。(Operation) The character extraction method of the character reading device of the present invention operates as follows. In the first step, as in the past, binary data is normalized according to the first compression rate determined by the form format (for example, depending on the size of the entry frame on the form and the size of the standard pattern). Obtain a character frame for value data (binary data by thinning). In the second step, a third compression ratio is determined based on the size of the character frame and the size of the standard pattern obtained in this way. In the second step, the second compression ratio is corrected based on the ratio between the number of black dots within the character frame and the area of the character frame, and the ratio between the number of black dots within the inclusive line (convex hull) and the area within the inclusive line. It is determined whether or not. For example, a character pattern with almost no white dots inside the character frame or within the comprehensive line,
A second compression ratio is determined to compensate. Fourth
In this step, the second compression rate corrected based on the determination in the third step (for example, vertically long and horizontally long character patterns are aligned to the smaller compression rate for both height and width), or The third compression rate is determined by correcting the first compression rate according to the compression rate of (for example, by multiplying the corrected second compression rate or the second compression rate and the first compression rate) It is determined).
In the fifth step, normalization is performed to thin out the binary data obtained by reading the characters on the form according to the third compression ratio, and character extraction is performed again. In this way, in addition to the second compression rate based on the size of the character frame, the second compression rate is also corrected based on the shape of the characters (by enclosing lines) to achieve the first compression rate based on the conventional form format. By correcting the ratio, a third compression ratio that is faithful to the actually written characters is determined, and normalization is performed using this third compression ratio, so optimal character extraction can be performed. In addition, it becomes possible to accurately recognize characters. Therefore, the problems of the prior art described above can be solved.
(実施例)
第1図は本発明の一実施例における光学式文字
読取装置の文字切出し方法の処理手順を示すフロ
ーチヤートであつて、第5図で述べた認識部14
の認識制御回路14bの制御により行われる。第
2図a,b,cは本実施例の文字切出しにおける
文字の大きさの正規化を具体的に説明するための
文字パタン例を示す図である。(Embodiment) FIG. 1 is a flowchart showing the processing procedure of a character segmentation method of an optical character reading device in an embodiment of the present invention, and shows the recognition unit 14 described in FIG.
This is performed under the control of the recognition control circuit 14b. FIGS. 2a, 2b, and 2c are diagrams showing examples of character patterns for specifically explaining the normalization of character size in character extraction in this embodiment.
まず、第1図を参照して文字切出し手順を説明
する。 First, the character extraction procedure will be explained with reference to FIG.
従来と同様に、予め帳票フオーマツトで定めら
れる圧縮率、即ち記入枠と標準パタンにより求め
られる圧縮率(y方向の圧縮率PY1、x方向の圧
縮率PX1)で間引きされた2値データに対して文
字切出しが行なわれ、文字枠が決定される。この
文字枠の高さYと巾Xとを標準パタンの高さY0
と巾X0に対して比較を行い、y方向の圧縮率PY
=Y0/Y、x方向の圧縮率文字PX=X0/Xを求
める(ステツプ1)。次に文字枠内の黒点数と文
字枠の面積との比較を行い、ある定数C1に対し
てその比較結果の大小の判断をする(ステツプ
2)。 As in the past, binary data is thinned out at the compression rate determined in advance by the form format, that is, the compression rate determined by the entry frame and standard pattern (y-direction compression rate PY 1 , x-direction compression rate PX 1 ). Character cutting is performed on the character, and a character frame is determined. The height Y and width X of this character frame are the standard pattern height Y 0
and the width X 0 , and the compression ratio in the y direction PY
=Y 0 /Y, the compression ratio in the x direction PX = X 0 /X is determined (step 1). Next, the number of black dots in the character frame is compared with the area of the character frame, and the magnitude of the comparison result is determined with respect to a certain constant C1 (step 2).
ステツプ2で(黒点数)/(文字枠面積)<C1
のときは、文字パタンの包括線(凸開包)を求
め、包括線内の黒点数と包括線内の面積との比較
を行い、ある定数C2に対してその比較結果の大
小の判断をする(ステツプ3)。 In step 2, (number of black dots) / (character frame area) < C 1
In this case, find the comprehensive line (convex hull) of the character pattern, compare the number of black dots within the comprehensive line and the area within the comprehensive line, and judge the magnitude of the comparison result for a certain constant C 2 . (Step 3).
ステツプ3で(黒点数)/(包括線内面積)<
C2のとき、ステツプ1で求めた圧縮率PY,PX
をそれぞれ記入枠より求められた圧縮率PY1,
PX1に掛けて最終の圧縮率PY0,PX0を求める
(ステツプ9)。求めた圧縮率PY0,PX0を基に間
引された2値データに対して文字切出しを再度行
う(ステツプ10)。 In step 3, (number of black dots) / (area within the comprehensive line) <
When C 2 , the compression ratio PY, PX obtained in step 1
The compression ratio PY 1 obtained from the respective entry boxes is PY 1 ,
Multiply by PX 1 to find the final compression ratio PY 0 and PX 0 (step 9). Character extraction is performed again on the thinned out binary data based on the obtained compression ratios PY 0 and PX 0 (step 10).
一方、ステツプ2で(黒点数)/(文字枠面
積)≧C1のときは、文字枠の高さYと巾Xとの比
較を行い、ある定数C3,C4に対して比較結果の
大小の判断をする(ステツプ4,5)。ステツプ
4でY/X>C3、又はステツプ5でX/Y>C4
の場合、又はステツプ3で(黒点数)/(包括線
内面積)≧C2の場合には、PXとPYの比較を行う
(ステツプ7)。ステツプ7でPX<PYのときは
PYにPXの値を入れ(ステツプ8a)、PX≧PYの
ときはPXにPYの値を入れる(ステツプ8b)。即
ち、ステツプ7,8a,8bではPY,PXは圧縮率
の小さい方の値に揃えられる。このy方向の圧縮
率PYを記入枠より求められた圧縮率PY1に掛け
てy方向の最終の圧縮率PY0を求め、同様にx方
向の圧縮率PXを圧縮率PX1に掛けてx方向の最
終の圧縮率PX0を求める(ステツプ9)。ステツ
プ9で求めた圧縮率PY0,PX0を基に間引きされ
た2値データに対して文字切出しを再度行う(ス
テツプ10)。 On the other hand, in step 2, if (number of black points)/(character frame area)≧C 1 , the height Y and width X of the character frame are compared, and the comparison result is calculated for certain constants C 3 and C 4 . Judging the size (steps 4 and 5). Y/X>C 3 in step 4, or X/Y>C 4 in step 5
, or if (number of sunspots)/(area within the comprehensive line)≧C 2 in step 3, PX and PY are compared (step 7). If PX<PY in step 7
Put the value of PX into PY (step 8a), and if PX≧PY, put the value of PY into PX (step 8b). That is, in steps 7, 8a, and 8b, PY and PX are adjusted to the value of the smaller compression ratio. Multiply this y-direction compression ratio PY by the compression ratio PY 1 obtained from the entry frame to obtain the final y-direction compression ratio PY 0 , and similarly multiply the x-direction compression ratio PX by the compression ratio PX 1 to obtain x The final compression ratio PX 0 in the direction is determined (step 9). Character extraction is performed again on the thinned out binary data based on the compression ratios PY 0 and PX 0 determined in step 9 (step 10).
また、ステツプ4でY/X≦C3かつステツプ
5でX/Y≦C4の場合には圧縮率は変えないと
いうことで、PX,PY共に1を入れ(ステツプ
6)、ステツプ9,10を行う。 Also, if Y/X≦C 3 in step 4 and X/Y≦C 4 in step 5, the compression ratio is not changed, so 1 is entered in both PX and PY (step 6), and steps 9 and 10 are set. I do.
このように、文字枠及び包括線を用いて実際に
書かれた文字の大きさに応じて正規化を行つた
後、文字切出しを再度行つている。 In this way, after normalization is performed according to the size of the actually written characters using the character frame and comprehensive line, character segmentation is performed again.
次に、第2図a,b,cの文字パタン例を参照
して文字切出し手順を具体的に説明する。 Next, a character cutting procedure will be specifically explained with reference to character pattern examples shown in FIGS. 2a, b, and c.
まず、文字パタンが、第2図aに示す3種類の
「1」の文字パタン101,102,103とす
ると、ステツプ1でこれらの文字パタンの文字枠
101a,102b,103bの高さYと巾Xよ
り、各各圧縮率PY,PXを求める。ステツプ2に
おいて、定数C1がC1≒1とすると、文字パタン
101は文字枠101a内に白点がほとんどない
ため、ステツプ4以降の処理が行われ、文字パタ
ン102,103は文字枠内の白点が多いので、
ステツプ3の処理が行なわれる。従つて、文字パ
タン102,103は、第2図bに示すように、
各々の包括線102b,103bが求められ、定
数C2がC2≒1とすると、文字パタン103は包
括線内に白点がほとんどないため、ステツプ7以
降の処理が行われる。 First, if the character patterns are three types of character patterns 101, 102, 103 of "1" shown in FIG. From X, find each compression ratio PY, PX. In step 2, if the constant C 1 is C 1 ≒ 1, character pattern 101 has almost no white dots within the character frame 101a, so the processing from step 4 onwards is performed, and character patterns 102 and 103 are Because there are many white spots,
Processing in step 3 is performed. Therefore, the character patterns 102 and 103 are as shown in FIG. 2b.
If the respective comprehensive lines 102b and 103b are obtained and the constant C 2 is C 2 ≈1, then the character pattern 103 has almost no white dots within the comprehensive line, so the processing from step 7 onwards is performed.
第2図cはステツプ2の判断で、ステツプ4の
処理が行われる文字パタン例である。ステツプ4
で定数C3がC3≒1とすると、文字パタン101
は縦長であるため、ステツプ7以降の処理が行わ
れる。また、文字パタン104と105はステツ
プ5の処理が行われる。ステツプ5の定数C4が
C4≒1とすると、ほぼ正方形の文字パタン10
4はステツプ6以降の処理が行われ、横長の文字
パタン105はステツプ7以降の処理が行われ
る。 FIG. 2c is an example of a character pattern for which the processing in step 4 is performed based on the judgment in step 2. Step 4
If the constant C 3 is C 3 ≒ 1, then the character pattern 101
Since it is vertically long, the processing from step 7 onwards is performed. Further, the character patterns 104 and 105 are processed in step 5. The constant C 4 in step 5 is
If C 4 ≒ 1, almost square character pattern 10
4 undergoes the processing from step 6 onwards, and the horizontally long character pattern 105 undergoes the processing from step 7 onwards.
次に包括線を決定する手順の一例を説明する。
ここで用いる文字パタンを第3図に示す。 Next, an example of a procedure for determining a comprehensive line will be explained.
The character pattern used here is shown in FIG.
まず、x方向を主走査、y方向を副走査とし
て、文字パタン106の文字枠107の左端から
黒点に当たるまでの白点数をとり、y方向を横軸
として、縦軸に(文字巾)−(求めた白点数)をと
つたヒストグラムを第4図aに示す。この時、ヒ
ストグラムの最大値イは、文字巾を示すが、この
最大値に達した点を最大ピークと称する。このヒ
ストグラムの各々の先端の点に対して包括線を求
める。まず、最大ピークから右側であるy方向の
正方向を見て、各々の先端への変化率を第4図b
に示す。最大ピークイから各先端までのy方向に
離れた距離を分母に、x方向に離れた距離を分子
として表わす。すなわち(縦座標、横座標)と表
わすと、最大ピークがm,nのとき、m−M,n
−Nの点を見た場合は、その変化率をM/Nとして
表わす。こうして求めた変化率に対して、変化率
が最も小さい点を第1の包括接点とする(第4図
のロ)。次に、この第1の包括接点からの変化率
をヒストグラム上の第1の包括接点より右側の各
先端に対して取り、求めた変化率が最も小さい点
を第2の包括接点とする(第4図のハ)。以下、
y方向の右端ニが、包括接点となるまで続ける。
このようにして、最大ピークより右側における包
括接点を全て求め、次に、最大ピークより左側の
y方向の負方向における包括接点を求める。 First, with the x direction as the main scan and the y direction as the sub scan, take the number of white points from the left end of the character frame 107 of the character pattern 106 to the black point, and take the y direction as the horizontal axis and the vertical axis as (character width) - ( A histogram of the obtained white point numbers is shown in FIG. 4a. At this time, the maximum value A of the histogram indicates the character width, and the point at which this maximum value is reached is called the maximum peak. A comprehensive line is found for each tip point of this histogram. First, look at the positive direction of the y direction, which is the right side from the maximum peak, and calculate the rate of change to each tip as shown in Figure 4b.
Shown below. The distance from the maximum peak point to each tip in the y direction is expressed as the denominator, and the distance in the x direction is expressed as the numerator. In other words, when expressed as (ordinate, abscissa), when the maximum peak is m, n, m-M, n
When looking at the point -N, the rate of change is expressed as M/N. With respect to the rate of change thus determined, the point where the rate of change is the smallest is defined as the first global contact point (FIG. 4, b). Next, the rate of change from this first comprehensive contact point is taken for each tip on the right side of the first comprehensive contact point on the histogram, and the point where the calculated rate of change is the smallest is set as the second comprehensive contact point ( (c in Figure 4). below,
Continue until the right end d in the y direction becomes a comprehensive contact point.
In this way, all comprehensive contacts on the right side of the maximum peak are found, and then comprehensive contacts in the negative direction of the y direction on the left side of the maximum peak are found.
以上により、ヒストグラム上の包括接点を全て
求め、次に、包括接点をつなぐ包括線を求める。
第4図cは、最大ピークと包括接点、又は、包括
接点間の包括線を求める方法を示すものである。
PpとPxは包括接点で、Ppを基にしたPxの変化率
はM/Nとし、Pp,Px間には包括接点がないとする
と、Ppからの変化率がM/N以下で、かつPpよりy
方向に等しく離れた点のうち変化率の最大の点
(この場合は白点)を包括線上の点とする。第4
図cにおいては、PpからP2の変化率をM−1/N−1と
すると、M−1/N−1>M/Nの場合は、P2は包括線
上
になく、N−1/N−1≦M/Nの場合はP2は包括線上
に
ある。P2が包括線上である場合は、Ppからのy
方向の距離がP2と等しく、Ppからの変化率が
M−2/N−1となるP1は包括線上にはない。 Through the above steps, all comprehensive contact points on the histogram are found, and then a comprehensive line connecting the comprehensive contact points is found.
FIG. 4c shows a method for determining the maximum peak and the comprehensive contact point or the comprehensive line between the comprehensive contact points.
P p and P x are comprehensive contacts, and the rate of change of P x based on P p is M/N. Assuming that there is no comprehensive contact between P p and P x , the rate of change from P p is M /N or less and which is equally distant from P p in the y direction, the point with the largest rate of change (in this case, the white point) is defined as the point on the comprehensive line. Fourth
In Figure c, if the rate of change from P p to P 2 is M-1/N-1, if M-1/N-1>M/N, P 2 is not on the comprehensive line and N-1 When /N-1≦M/N, P 2 is on the comprehensive line. If P 2 is on the comprehensive line, then y from P p
P 1 whose distance in the direction is equal to P 2 and whose rate of change from P p is M-2/N- 1 is not on the comprehensive line.
以上より、求めた包括線上の点の集まりと最大
ピークと包括接点を求めることで、第3図に示す
文字パタン106の左半分の包括線108が求ま
る。 From the above, the comprehensive line 108 of the left half of the character pattern 106 shown in FIG. 3 can be found by finding the collection of points on the obtained comprehensive line, the maximum peak, and the comprehensive contact point.
一方、第3図の文字枠107の右端から、x方
向の負方向を主走査とし、y方向を副走査として
黒点に当たるまでの白点数をとり、以下、左端か
らの場合と同様に行つて、第3図に示す文字パタ
ン106の右半分の包括線109が求まる。以上
の結果、第3図の文字パタン106の包括線11
0が求まる。 On the other hand, from the right end of the character frame 107 in FIG. 3, the negative direction in the x direction is used as the main scan, and the direction in the y direction is used as the sub scan, and the number of white points is taken until it hits a black point. A comprehensive line 109 for the right half of the character pattern 106 shown in FIG. 3 is determined. As a result of the above, the inclusive line 11 of the character pattern 106 in FIG.
Find 0.
本実施例では光学式文字読取装置の文字切出し
方法について説明したが本発明の方法は光学式以
外の文字読取装置にも適用可能である。 Although the present embodiment describes a character cutting method for an optical character reading device, the method of the present invention is also applicable to character reading devices other than optical type.
(発明の効果)
以上詳細に説明したように本発明によれば、文
字枠及び包括線により実際に書かれた文字の大き
さ及び文字の形状に忠実に2値データを正規化し
ているので、最適な文字の切出しを行うことがで
きる。従つて、正確に文字を認識することが可能
となる。(Effects of the Invention) As explained in detail above, according to the present invention, binary data is normalized faithfully to the size and shape of the character actually written using the character frame and comprehensive line. Optimal character extraction can be performed. Therefore, it becomes possible to accurately recognize characters.
第1図は本発明の実施例を示すフローチヤー
ト、第2図a,b,cは本実施例の正規化の説明
図、第3図は文字パタン例を示す図、第4図a,
b,cは包括線の決定の説明図、第5図a,bは
一般的なOCRのブロツク図、第6図a,bは文
字パタン例を示す図、第7図は2値データの間引
き例を示す図、第8図a,bは記入枠による正規
化の説明図である。
11……光電変換部、12……多値データバツ
フア、13……フイルタ回路、14……認識部、
14a……2値データバツフア、14b……認識
制御回路。
Fig. 1 is a flowchart showing an embodiment of the present invention, Fig. 2 a, b, and c are explanatory diagrams of normalization in this embodiment, Fig. 3 is a diagram showing an example of a character pattern, and Fig. 4 a,
Figures b and c are diagrams explaining the determination of comprehensive lines, Figures 5 a and b are block diagrams of general OCR, Figures 6 a and b are diagrams showing examples of character patterns, and Figure 7 is thinning of binary data. Figures 8a and 8b, which illustrate examples, are explanatory diagrams of normalization using entry boxes. 11...Photoelectric conversion unit, 12...Multi-value data buffer, 13...Filter circuit, 14...Recognition unit,
14a...Binary data buffer, 14b...Recognition control circuit.
Claims (1)
を正規化し、正規化した2値データに対し文字パ
ターンの外接矩形領域である文字枠を求めて文字
の切出しを行う文字読取装置の文字切出し方法に
おいて、 帳票上の文字を読取つて得られた2値データ
を、帳票のフオーマツトの記入枠と標準パタンに
より第1の圧縮率を求めるとともに、前記第1の
圧縮率で正規化された2値データに対して文字パ
タンの文字枠を求める第1のステツプと、 前記第1のステツプで求めた文字枠の大きさと
標準パタンの大きさに基づいて第2の圧縮率を求
める第2のステツプと、 前記文字枠内の黒点数と文字枠面積との比及び
前記文字パタンの包括線内の黒点数と包括線内面
積との比に基づいて、第2の圧縮率を補正するか
否かを決定する第3のステツプと、 第3のステツプの決定に基づいて、高さ方向又
は横方向の圧縮率のいずれか小さい方に揃えるよ
うに補正した第2の圧縮率又は補正しない第2の
圧縮率に応じて第1の圧縮率を補正して第3の圧
縮率を決定する第4のステツプと、 帳票上の文字を読取つて得られた前記2値デー
タを第3の圧縮率に応じた間引きにより正規化し
て文字の切出しを再度行う第5のステツプとを有
することを特徴とする文字読取装置の文字切出し
方法。[Claims] 1. Normalize binary data obtained by reading characters on a form, and extract characters by finding a character frame, which is a circumscribed rectangular area of a character pattern, from the normalized binary data. In a character extraction method for a character reading device, a first compression rate is determined for binary data obtained by reading characters on a form using a form entry frame and a standard pattern, and the first compression rate is A first step of determining a character frame of a character pattern for the normalized binary data, and a second compression ratio based on the size of the character frame determined in the first step and the size of the standard pattern. A second compression ratio is determined based on the second step to be determined, the ratio of the number of black dots in the character frame to the area of the character frame, and the ratio of the number of black dots in the comprehensive line of the character pattern to the area within the comprehensive line. a third step of determining whether or not to correct; and a second compression ratio corrected to be equal to the smaller of the compression ratio in the height direction or the lateral direction based on the determination in the third step; a fourth step of determining a third compression rate by correcting the first compression rate according to the uncorrected second compression rate; 1. A method for cutting out characters for a character reading device, comprising: a fifth step of normalizing the characters by thinning out according to a compression ratio of the characters and cutting out the characters again.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP61285877A JPS63140387A (en) | 1986-12-02 | 1986-12-02 | Character segmenting method for character reader |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP61285877A JPS63140387A (en) | 1986-12-02 | 1986-12-02 | Character segmenting method for character reader |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPS63140387A JPS63140387A (en) | 1988-06-11 |
| JPH0517598B2 true JPH0517598B2 (en) | 1993-03-09 |
Family
ID=17697186
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP61285877A Granted JPS63140387A (en) | 1986-12-02 | 1986-12-02 | Character segmenting method for character reader |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPS63140387A (en) |
-
1986
- 1986-12-02 JP JP61285877A patent/JPS63140387A/en active Granted
Also Published As
| Publication number | Publication date |
|---|---|
| JPS63140387A (en) | 1988-06-11 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US5075895A (en) | Method and apparatus for recognizing table area formed in binary image of document | |
| US6226402B1 (en) | Ruled line extracting apparatus for extracting ruled line from normal document image and method thereof | |
| EP0287995B1 (en) | Method and apparatus for recognizing pattern of halftone image | |
| US4556985A (en) | Pattern recognition apparatus | |
| US5243668A (en) | Method and unit for binary processing in image processing unit and method and unit for recognizing characters | |
| US6983071B2 (en) | Character segmentation device, character segmentation method used thereby, and program therefor | |
| JP2868134B2 (en) | Image processing method and apparatus | |
| JP3095470B2 (en) | Character recognition device | |
| JP3850488B2 (en) | Character extractor | |
| JP3936039B2 (en) | Screened area extraction device | |
| JP2795860B2 (en) | Character recognition device | |
| JP2708604B2 (en) | Character recognition method | |
| JP3162414B2 (en) | Ruled line recognition method and table processing method | |
| JPH08194776A (en) | Form processing method and device | |
| JP3084833B2 (en) | Feature extraction device | |
| JP3104355B2 (en) | Feature extraction device | |
| CN121907963A (en) | Scanner image paper fold detection methods, systems, media and products | |
| JPS622382A (en) | Image processing method | |
| JP2980636B2 (en) | Character recognition device | |
| JPS63140387A (en) | Character segmenting method for character reader | |
| JP3127413B2 (en) | Character recognition device | |
| JPH08129612A (en) | Character recognition method and character recognition device | |
| JPH02231690A (en) | Linear picture recognizing method | |
| JPH04372085A (en) | Character reading method | |
| JPH0425987A (en) | Character recognizing device |