JPS6330666B2 - - Google Patents
Info
- Publication number
- JPS6330666B2 JPS6330666B2 JP55187608A JP18760880A JPS6330666B2 JP S6330666 B2 JPS6330666 B2 JP S6330666B2 JP 55187608 A JP55187608 A JP 55187608A JP 18760880 A JP18760880 A JP 18760880A JP S6330666 B2 JPS6330666 B2 JP S6330666B2
- Authority
- JP
- Japan
- Prior art keywords
- character
- buffer
- image data
- scanning
- memory
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/10—Character recognition
- G06V30/14—Image acquisition
- G06V30/148—Segmentation of character regions
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/10—Character recognition
Landscapes
- Engineering & Computer Science (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Multimedia (AREA)
- Theoretical Computer Science (AREA)
- Character Input (AREA)
Description
【発明の詳細な説明】
本発明は文字分離方式に関し、特に複数の文字
が一定の形状のマスクでは個々の文字に分離でき
ないように互に接近して手書き原稿や図面等に書
込まれられている場合でも、これらの各文字を
個々の文字に分離できるようにした文字分離方式
に関する。DETAILED DESCRIPTION OF THE INVENTION The present invention relates to a character separation method, and in particular, when a plurality of characters are written in handwritten manuscripts, drawings, etc. in close proximity to each other so that they cannot be separated into individual characters with a mask of a certain shape. This invention relates to a character separation method that allows each character to be separated into individual characters even when the characters are different from each other.
例えば自由に手書きされた手書き文字を文字読
取装置で認識する場合、まずその第1段階として
文字を一字毎に分離して抽出し、それからこの抽
出した文字を識別している。このように文字を分
離して抽出する場合、例えば第1図イに示す如き
矩形状のマスクMを使用し、これを同ロの如く文
字列上に走査して個々の文字を抽出していた。こ
の場合、第1図ロに示すように、複数の文字が互
に離れて、しかもある大きさの文字で揃つて書か
れている場合には、これをマスクMのような一定
形状のマスクにより各文字毎に分離することがで
きるので、これらをそれぞれ互に分離して抽出す
ることが可能である。 For example, when freely handwritten characters are recognized by a character reading device, the first step is to separate and extract each character, and then to identify the extracted characters. When separating and extracting characters in this way, for example, a rectangular mask M as shown in Figure 1A was used, and this was scanned over the character string as shown in Figure 1B to extract individual characters. . In this case, as shown in Figure 1 B, if multiple characters are written apart from each other and in a certain size, they can be written using a mask of a certain shape such as mask M. Since each character can be separated, it is possible to separate and extract these characters from each other.
しかしながら、第1図ハに示すように、複数の
文字が互に接近して記載されている場合、もはや
第1図イに示す如きマスクを単に走査する方法で
はこれらの各文字を正確に分離することが不可能
であり、したがつてこのような場合には文字読取
装置で文字を正確に認識することが困難である。 However, when multiple characters are written close to each other as shown in Figure 1C, it is no longer possible to accurately separate these characters using the method of simply scanning a mask as shown in Figure 1B. Therefore, in such a case, it is difficult for a character reading device to accurately recognize the characters.
したがつて本発明の目的は、このように複数の
文字が近接して記載されている場合でも各文字を
正確に分離することができる文字分離方式を提供
することを目的とするものである。そしてこのた
めに本発明における文字分離方式では、入力手段
から入力された文字の画像情報を保持する情報保
持手段と、上記文字情報を第1の方向に走査して
その白黒変化点を求め、該変化点の間の領域であ
る第1領域を検出する第1特徴抽出手段と、上記
画像情報を、前記第1の方向と略直交する第2の
方向に走査してその白黒変化点を求め、該変化点
の間の領域である第2領域を検出する第2特徴抽
出手段と、上記第1領域と第2領域とが重畳しな
い領域である中間領域を前記第1の方向に順次走
査し走査線分を得る走査手段と、上記中間領域の
各走査線分における中間点を求める中間点設定手
段と、上記文字の画像情報に近接して前記中間点
から第1の方向に水平線を発生する水平線発生手
段を有し、上記中間点列および上記水平線により
個々の文字情報の存在する文字領域を区別できる
ようにしたことを特徴とする。 Therefore, an object of the present invention is to provide a character separation method that can accurately separate each character even when a plurality of characters are written close to each other. To this end, the character separation method of the present invention includes an information holding means that holds image information of characters input from an input means, and an information holding means that scans the character information in a first direction to find black and white change points. a first feature extracting means for detecting a first region that is an area between the changing points; and scanning the image information in a second direction substantially orthogonal to the first direction to obtain the black and white changing points; a second feature extracting means for detecting a second region that is a region between the change points; and a second feature extraction means that sequentially scans in the first direction an intermediate region that is a region where the first region and the second region do not overlap. scanning means for obtaining a line segment; midpoint setting means for determining a midpoint in each scanned line segment of the intermediate area; and a horizontal line that is close to the image information of the character and generates a horizontal line in a first direction from the midpoint. The present invention is characterized in that it has a generating means, and character areas in which individual character information exists can be distinguished by the intermediate point string and the horizontal line.
以下本発明を具体的に説明するに先立ち、本発
明の原理を第2図〜第14図にもとづき説明す
る。 Before explaining the present invention in detail below, the principle of the present invention will be explained based on FIGS. 2 to 14.
(a) まず手書き文字の記入された原稿あるいは図
面を読出すことにより得られた画像データm0
の記入されたメモリm1,m2を用意する。そし
てメモリm1を第2図イに実線で示すようにx
方向に走査する。このとき文字が画かれている
黒点を「1」、文字の画かれていないところを
「0」として画像データm0を得る。そして第2
図ロに示す如く、変化点すなわち「0」→
「1」および「1」→「0」に変化する点を求
める。例えば第2図ロに示すように、ラインl
上を走査するとき、x1では「0」→「1」に変
化し、x2では「1」→「0」に変化し、x3では
「0」→「1」に変化し、x4では「1」→「0」
に変化するので、これらのx1〜x4はいずれも変
化点である。この場合変化点P1(「1」→
「0」)およびP2(「0」→「1」)の対を検出
し、メモリm2上のこの変化点P1〜P2間の領域
をすべて「1」とする。(a) First, image data m 0 obtained by reading a manuscript or drawing with handwritten characters
Prepare memories m 1 and m 2 in which . Then, the memory m 1 is
Scan in the direction. At this time, image data m0 is obtained by setting black dots where characters are drawn as "1" and places where no characters are drawn as "0". and the second
As shown in Figure B, the point of change, i.e. "0" →
Find "1" and the point where "1" changes to "0". For example, as shown in Figure 2 (b), the line l
When scanning above, x 1 changes from "0" to "1", x 2 changes from "1" to "0", x 3 changes from "0" to "1", x 4 Then "1" → "0"
Therefore, all of these x 1 to x 4 are changing points. In this case, the change point P 1 (“1” →
0) and P 2 (“0”→“1”), and all areas between the change points P 1 and P 2 on the memory m 2 are set to “1”.
(b) 次にメモリm1上の画像データm0を、第2図
イに鎖線で示すように、y方向に走査して、同
様に変化点Q1(「1」→「0」ただしy方向)
およびQ2(「0」→「1」)の対を検出し、これ
により、該Q1〜Q2間のメモリm2上の領域を、
もし「0」ならば「1」に、
もし「1」ならば「0」に、
反転させる。これにより第3図に示す如く、二
重線領域が「0」になり、1線領域が「1」と
なり、メモリm2には第4図に示すデータが記
入されることになる。このとき文字の部分は通
常「1」が連続しているのでそのまま残る。か
くして第4図の斜線部が「1」という状態のデ
ータがメモリm2に保持される。(b) Next, the image data m 0 on the memory m 1 is scanned in the y direction as shown by the chain line in Fig. 2 direction)
and Q 2 (“0” → “1”), and thereby the area on memory m 2 between Q 1 and Q 2 is changed to “1” if “0” and “1” if “ If it is "1", it is reversed to "0". As a result, as shown in FIG. 3, the double line area becomes "0", the single line area becomes "1", and the data shown in FIG. 4 is written in the memory m2. At this time, the character part is usually a series of "1"s, so it remains as is. In this way, the data in the state where the shaded area in FIG. 4 is "1" is held in the memory m2 .
(c) 今度は、この第4図に示す状態のデータが記
入されてにる領域のクラスタリングを行なつ
て、第5図に示すように、各領域にラベル記号
1〜8を付与する。(c) Next, clustering is performed on the areas in which the data in the state shown in FIG. 4 has been entered, and label symbols 1 to 8 are assigned to each area as shown in FIG.
次にこのクラスタリングについて説明する。 Next, this clustering will be explained.
(c‐1) 第6図イに示すような3×2のマスクM2
を、同ロに示すようにx方向に走査する。そ
して第4図に示す斜線部の領域について、こ
のマスクM2の走査を行なうとき、このマス
クM2の区分c〜fにラベル記号の付与され
た点があるか否かを調べる。もし存在すれば
その区分aにそのラベル記号を付与するが、
ラベル記号付けされた点がなければ区分fに
新らしいラベル記号例えばラベル信号1を付
与し、これにもとづき区分aにラベル信号1
を付与する。そして次にマスクM2がx方向
に移動したとき、今度は区分fのその前の区
分aが位置することになり、この新しい区分
fにはラベル記号1が付与されているので、
その新しい区分aにラベル記号1を付与す
る。このようにしてメモリmをこのマスク
M2で全部走査してラベル記号付けを行なう。(c-1) 3×2 mask M 2 as shown in Figure 6 A
is scanned in the x direction as shown in the same figure. When the mask M 2 is scanned for the shaded area shown in FIG. 4, it is checked whether there is a point given a label symbol in the sections c to f of the mask M 2 . If it exists, give that label symbol to the category a, but
If there is no point with a label symbol, a new label symbol, for example, label signal 1, is assigned to the segment f, and based on this, a label signal 1 is assigned to the segment a.
Grant. Then, when the mask M 2 moves in the x direction, the previous section a of the section f will now be located, and this new section f is given the label symbol 1, so
Label symbol 1 is assigned to the new category a. In this way you can store memory m using this mask.
Scan everything with M 2 and label it.
(c‐2) しかしながらこの場合、次のような問題が
生ずる。例えば第7図イに示すように、同一
の区分を走査しても、その形状によつては別
のラベル記号が付与されることが生ずる。例
えば、第7図イの如き形状の領域αを走査す
るとき、まず区域α1を走査し、次にある間隔
を置いて全域α2を走査する。このとき区域α1
とα2が同一のものであるということは、最初
判断できないので、x方向に走査するときは
区域α1に例えばラベル信号11を付与し、区域
α2にラベル記号12を付与することになる。こ
のため、一応マスクM2による走査が終了後、
第7図ロに示すように3×3のマスクM3を
使用して再度走査する。この場合、区分αの
周囲の区分b〜iに異なるラベル記号が付与
されている場合、これを例えば数の少ないラ
ベル信号に統一する。これにより、第7図イ
の点線の近くの部分がラベル記号1に訂正さ
れる。このような操作をラベル記号の変化が
なくなるまで繰返す。これにより、第7図ハ
に示すように、同一の領域には同一のラベル
記号が付与されることになる。(c-2) However, in this case, the following problems arise. For example, as shown in FIG. 7A, even if the same section is scanned, different label symbols may be given depending on its shape. For example, when scanning an area α having a shape as shown in FIG. 7A, first the area α 1 is scanned, and then the entire area α 2 is scanned at a certain interval. In this case, the area α 1
Since it cannot be determined at first that and α 2 are the same, when scanning in the x direction, for example, label signal 11 is given to area α 1 , and label symbol 12 is given to area α 2 . . For this reason, after the scanning by mask M2 is completed,
Scanning is performed again using a 3×3 mask M 3 as shown in FIG. 7B. In this case, if different label symbols are given to the sections b to i surrounding the section α, these are unified into a small number of label signals, for example. As a result, the part near the dotted line in FIG. 7A is corrected to label symbol 1. This operation is repeated until there are no more changes in the label symbol. As a result, as shown in FIG. 7C, the same area is given the same label symbol.
(c‐3) このようにして同一領域におけるラベル記
号の統一が行なわれたときに、各ラベル記号
はシーケンシヤルに付与されていないことが
ほとんどである。例えばラベル記号が1、
4、9、12はと付与されることが多いので、
これをシーケンシヤルに訂正し、ラベル記号
の整列を行なう。かくして上記の1、4、
9、12はのものは、1、2、3、4はと訂正
され、ラベル記号が整列されることになる。(c-3) When label symbols in the same area are unified in this way, in most cases the label symbols are not assigned sequentially. For example, the label symbol is 1,
4, 9, and 12 are often given as .
This is corrected sequentially and the label symbols are aligned. Thus, 1, 4 above,
The numbers 9 and 12 will be corrected to 1, 2, 3, and 4, and the label symbols will be aligned.
(d) このようにしてラベル記号が付与された第5
図の画面を、第2図イの実線で示すようにx方
向に走査する。そして同ラベル記号(Lm)毎
に変曲点R1(「0」→「Lm」)はR2(「Lm」→
「0」)の対を検出する。ここで「Lm」はラベ
ル記号Lmの付与された領域である。そして上
記変化点R1、R2の対の中点Mの軌跡を記憶す
る。この場合、R1の座標を(x1、y1)とし、
R2の座標を(x2、y2)とし、中点Mの座標を
(xm、ym)とすれば、xm=(x1+x2)/2、
ym=y1=y2である。このようにして求めた中
点M1(xm1、ym1)、M2(xm2、ym2)はの位置
を、例えば第8図ロの如きテーブルに記憶する
ことにより、この中点Mの軌跡を記憶すること
ができる。このような中点の軌跡を全ラベル記
号領域毎に求める。これにより、第9図に示す
中点を得ることができる。(d) No. 5 labeled in this way.
The screen shown in the figure is scanned in the x direction as shown by the solid line in Figure 2A. And for each label symbol (Lm), the inflection point R 1 ("0" → "Lm") is R 2 ("Lm" →
"0") pairs are detected. Here, "Lm" is the area to which the label symbol Lm is assigned. Then, the locus of the midpoint M of the pair of change points R 1 and R 2 is stored. In this case, let the coordinates of R 1 be (x 1 , y 1 ),
If the coordinates of R 2 are (x 2 , y 2 ) and the coordinates of the midpoint M are (xm, ym), then xm = (x 1 + x 2 )/2,
ym= y1 = y2 . By storing the positions of the midpoints M 1 (xm 1 , ym 1 ) and M 2 (xm 2 , ym 2 ) obtained in this way in a table such as the one shown in FIG. The trajectory of can be memorized. The trajectory of such a midpoint is obtained for each label symbol area. As a result, the midpoint shown in FIG. 9 can be obtained.
(e) このようにして求めた各ラベル記号領域の中
点Mの軌跡に対して、第10図イに示すよう
に、互にオーバラツプしている2点間にx方向
の直線(水平線)を引く。そしてその直線上を
メモリm1にセツトしている原画像データm0で
検索し、黒点すなわち文字情報がなければ第1
0図ロに示すように、その各々の2点のテーブ
ル上に互のラベル記号を付与する。一度相手の
ラベル記号が付与されると、その2つのラベル
記号領域間でのこの処理を終了する。そしてこ
のような処理を全ラベル記号領域間で行なう。(e) For the trajectory of the midpoint M of each label symbol area obtained in this way, draw a straight line (horizontal line) in the x direction between two points that overlap each other, as shown in Figure 10A. Pull. Then, the original image data m0 set in the memory m1 is searched on the straight line, and if there is no black point, that is, character information, the first
As shown in Figure 0B, mutual label symbols are given to each of the two points on the table. Once the other party's label symbol is assigned, this process between the two label symbol areas is completed. Then, such processing is performed for all label symbol areas.
(f) 次に上記(d)で求めた中点Mの軌跡が黒点に接
触した位置からεだけ手前の距離を求める。こ
のεの大きさはシミユレーシヨンにより適宜定
めるものである。したがつて第11図イに示す
ように中点Mの軌跡が黒点に接触した位置から
εだけ離れた位置を同ロ,ハのようにテーブル
より求めることができる。第11図ロは例えば
ラベル記号1の領域の場合を示し、同ハは例え
ばラベル記号5の領域の場合を示す。(f) Next, find the distance ε before the position where the trajectory of the midpoint M found in (d) above touches the black dot. The magnitude of ε is determined appropriately by simulation. Therefore, as shown in FIG. 11A, the position ε apart from the position where the locus of the midpoint M contacts the black dot can be determined from the table as shown in FIGS. 11A and 11B. FIG. 11B shows, for example, the case of an area labeled with label symbol 1, and FIG.
(g) このようにεだけ手前の距離の位置からx方
向に水平線を引く。この水平線が黒点と接触し
たとき、その水平線の左端を該水平線のスター
ト点x3とし右端をエンド点xEとする。そしてこ
れを第12図イの如きテーブルに記入する。(g) In this way, draw a horizontal line in the x direction from a position ε in front of you. When this horizontal line touches the black dot, the left end of the horizontal line is the starting point x3 , and the right end is the end point xE . Then, enter this into a table as shown in Figure 12A.
(h) このようにして求めた各水平線において他の
水平線とx座標がオーバラツプする部分につい
て垂直線を引き、その垂直線がメモリm1にセ
ツトしている原画像データm0でその軌跡を検
索する。そして黒点と接触せず、しかも両ラベ
ル記号領域間で相互ラベルの付与がない場合
(中点テーブルより検出)ならば、その垂直線
を境界線の1つとして採用することになる。か
くして第13図に示す如き線引きで行なうこと
ができる。(h) For each horizontal line obtained in this way, draw a vertical line where the x-coordinate overlaps with another horizontal line, and search for the trajectory of that vertical line in the original image data m0 set in memory m1 . do. If there is no contact with the black dot and there is no mutual labeling between both label symbol areas (detected from the midpoint table), that vertical line will be used as one of the boundaries. In this way, it can be done by drawing lines as shown in FIG.
(i) かくして得られた、第14図に示した中点軌
跡および水平線、垂直線等で構成された境界線
により、文字情報を完全に個々に分離すること
ができる。それ故、上記境界線情報によりメモ
リm1に保持されている原画像データを抽出す
れば個々の文字情報を得ることができる。(i) Character information can be completely separated into individual pieces using the midpoint locus and the boundary line composed of horizontal lines, vertical lines, etc., as shown in FIG. 14 thus obtained. Therefore, individual character information can be obtained by extracting the original image data held in the memory m1 using the boundary line information.
次に本発明の一実施例を第15図にもとづき説
明する。 Next, one embodiment of the present invention will be described based on FIG. 15.
図中、1は入力部、2は第1画像メモリ、3は
バツフア、4は第1特徴抽出部、5は第1アドレ
ス・テーブル、6は第1アドレス発生部、7は第
2アドレス発生部、8はバツフア、9は第2特徴
抽出部、10は第2アドレス・テーブル、11は
エクスクルシーブ・オア回路、12は第2画像メ
モリ、13は制御部、14は第3アドレス発生
部、15はバツフア、16は第1ラベル付処理
部、17は第3画像メモリ、18は第4アドレス
発生部、19はバツフア、20は第2ラベル付処
理部、21は第4画像メモリ、22はラベル整列
回路、23は第5画像メモリ、24はバツフア、
25は第5アドレス発生部、26は中点抽出回
路、27は第1中点アドレス・テーブル、28は
境界線処理部、29はバツフア、30は第2中点
アドレス・テーブル、31は水平線発生処理部、
32は水平線テーブル・メモリ、33,34はバ
ツフア、35は境界線抽出部、36は境界線テー
ブル、37はバツフア、38は出力メモリであ
る。 In the figure, 1 is an input section, 2 is a first image memory, 3 is a buffer, 4 is a first feature extraction section, 5 is a first address table, 6 is a first address generation section, and 7 is a second address generation section. , 8 is a buffer, 9 is a second feature extraction section, 10 is a second address table, 11 is an exclusive OR circuit, 12 is a second image memory, 13 is a control section, 14 is a third address generation section, 15 is a buffer, 16 is a first label processing section, 17 is a third image memory, 18 is a fourth address generation section, 19 is a buffer, 20 is a second label processing section, 21 is a fourth image memory, and 22 is a buffer. a label alignment circuit, 23 a fifth image memory, 24 a buffer,
25 is a fifth address generator, 26 is a midpoint extraction circuit, 27 is a first midpoint address table, 28 is a boundary line processing unit, 29 is a buffer, 30 is a second midpoint address table, and 31 is a horizontal line generator. processing section,
32 is a horizontal line table memory, 33 and 34 are buffers, 35 is a boundary line extractor, 36 is a boundary line table, 37 is a buffer, and 38 is an output memory.
入力部1は手書き原稿等を例えば光電変換部で
変換する電気信号発生部である。第1画像メモリ
2は、入力部1から入力された画像データが保持
されるメモリである。バツフア3は、該第1画像
メモリ4に保持された画像データが送出されこれ
にセツトされるものであつて、上記(a)の如き処理
を行なうための作業用のバツフア・メモリであ
る。 The input unit 1 is an electrical signal generation unit that converts a handwritten manuscript or the like using, for example, a photoelectric conversion unit. The first image memory 2 is a memory in which image data input from the input section 1 is held. The buffer 3 is a working buffer memory to which the image data held in the first image memory 4 is sent and set, and is used to perform the processing as described in (a) above.
第1特徴抽出部4は、バツフア3に保持された
画像データから、上記(a)の如く、変化点P1およ
びP2の対を求め、そのP1〜P2の領域を「1」と
して読出すものであり、第1アドレス・テーブル
5は上記変化点P1〜P2の間の「1」の領域のア
ドレスが保持されるテーブルである。 The first feature extraction unit 4 obtains a pair of change points P 1 and P 2 from the image data held in the buffer 3, as shown in (a) above, and sets the region of P 1 to P 2 as "1". The first address table 5 is a table in which the addresses of the areas of "1" between the change points P1 and P2 are held.
第1アドレス発生部6は、バツフア3に保持さ
れた画像データを、上記(a)に示す如く、x方向に
走査するためのアドレスを発生するものである。 The first address generating section 6 generates an address for scanning the image data held in the buffer 3 in the x direction as shown in (a) above.
第2アドレス発生部7はバツフア8に保持され
た第1画像メモリ2から送出された画像データ
を、上記(b)に示すように、y方向に走査するため
のアドレスを発生するものである。 The second address generating section 7 generates an address for scanning the image data sent from the first image memory 2 held in the buffer 8 in the y direction as shown in (b) above.
第2特徴抽出部9は、バツフア8に保持された
画像データから、上記(b)の如く、変化点Q1,Q2
の対を検出し、そのQ1,Q2の領域を「1」とし
て読出すものであり、第2アドレス・テーブル1
0は上記変化点Q1〜Q2間の「1」の領域のアド
レスが保持されるテーブルである。 The second feature extraction unit 9 extracts the change points Q 1 and Q 2 from the image data held in the buffer 8 as shown in (b) above.
, and reads out the Q 1 and Q 2 areas as "1".
0 is a table in which addresses of the area of "1" between the change points Q 1 and Q 2 are held.
第2画像メモリ12は、上記(a)、(b)の結果得ら
れた第4図に示す画像データがセツトされるメモ
リである。 The second image memory 12 is a memory in which the image data shown in FIG. 4 obtained as a result of the above (a) and (b) is set.
制御部13は、第1画像メモリ2に入力された
画像データを、上記(a)〜(i)の手順にしたがつて処
理し、個々の文字領域を作成抽出するための各種
制御を行なうものであつて、例えばバツフア3を
走査するための第1アドレス発生部6を制御した
り、第1特徴抽出部4を制御するものである。 The control unit 13 processes the image data input to the first image memory 2 according to the steps (a) to (i) above, and performs various controls for creating and extracting individual character areas. For example, it controls the first address generation section 6 for scanning the buffer 3 or the first feature extraction section 4.
第3アドレス発生部14は、バツフア15に保
持された画像データを第6図イに示すマスクを使
用して、同ロの如く走査させ、上記(c−1)の
如き処理を行なうためのアドレスを発生させるも
のである。バツフア15は第2画像メモリ12か
ら伝達された、第4図に示す如き画像データが保
持されるものである。 The third address generation unit 14 scans the image data held in the buffer 15 using the mask shown in FIG. 6A, as shown in FIG. It is something that generates. The buffer 15 is for holding image data as shown in FIG. 4 transmitted from the second image memory 12.
第1ラベル付処理部16は、バツフア15にセ
ツトされた第4図に示す画像データに第6図イに
示すマスクを同ロの如く走査して得られたデータ
にもとづき、上記(c−1)の如き処理を行なつ
て各領域にラベル記号を付与する処理を行うもの
であり、この結果ラベル記号が付与された画像デ
ータが第3画像メモリ17にセツトされることに
なる。 The first labeling processing unit 16 scans the image data shown in FIG. 4 set in the buffer 15 with the mask shown in FIG. ) to add a label symbol to each area, and as a result, the image data to which the label symbol is assigned is set in the third image memory 17.
アドレス発生部18は、上記(c−2)に説明
したように、第7図ロに示すマスクM3をバツフ
ア19に走査させるためのアドレスを発生するも
のである。バツフア19は上記第3画像メモリ1
7から伝達された画像データがセツトされるもの
である。 As explained in (c-2) above, the address generating section 18 generates an address for causing the buffer 19 to scan the mask M3 shown in FIG. 7B. The buffer 19 is the third image memory 1
The image data transmitted from 7 is set here.
第2ラベル付処理部20は、上記バツフア19
に上記マスクM3を走査させた結果得られたデー
タにもとづき、上記(c−2)に説明したような
処理を行なうものである。 The second labeling processing section 20
Based on the data obtained as a result of scanning the mask M3 , the processing described in (c-2) above is performed.
第4画像メモリ21は、上記第2ラベル付処理
部20により処理された、同一領域におけるラベ
ル記号が統一された画像データがセツトされるメ
モリである。 The fourth image memory 21 is a memory in which image data processed by the second labeling processing section 20 and having unified label symbols in the same area is set.
ラベル整列回路22は、上記(c−3)に記載
した如く、ラベル記号をシーケンシヤルに整列す
るものであり、この結果得られた、第5図に示す
如きラベル記号の付与された画像データが第5画
像メモリ23にセツトされる。そしてこの第5画
像メモリ23にセツトされた画像データはバツフ
ア24に送出される。第5アドレス発生部25
は、バツフア24にセツトされた。第5図に示さ
れる画像データをx方向に走査するためのアドレ
スを発生するものである。 The label alignment circuit 22 sequentially aligns the label symbols as described in (c-3) above, and the resulting image data with label symbols as shown in FIG. 5 is set in the image memory 23. The image data set in the fifth image memory 23 is then sent to the buffer 24. Fifth address generation section 25
was set to buffer 24. This generates an address for scanning the image data shown in FIG. 5 in the x direction.
中点抽出回路26は、上記(d)の如き処理を行な
いその中点アドレスを得るものであつて、得られ
た中点アドレスは、第8図ロのような状態で、第
1中点アドレス・テーブル27にセツトされる。 The midpoint extraction circuit 26 performs the process as described in (d) above to obtain the midpoint address, and the midpoint address obtained is the first midpoint address in a state as shown in FIG. - Set on table 27.
境界線連絡処理部28は、上記(e)の如き処理を
行なうものであつて、そのために必要な原画像デ
ータは第1画像メモリ2からバツフア29に送出
されるものである。そしてこの結果得られた、第
10図ロに示すラベル記号の付与されたテーブル
が第2中点アドレス・テーブル30にセツトされ
る。 The boundary line communication processing section 28 performs the processing as described in (e) above, and the original image data necessary for this purpose is sent from the first image memory 2 to the buffer 29. The table obtained as a result and labeled with the label symbol shown in FIG. 10B is set in the second midpoint address table 30.
水平線発生処理部31は、上記(f)および(g)の処
理を行なうものであり、そのために必要な原画像
データは、第1画像メモリ2からバツフア33に
送出されるものである。そしてこの結果得られた
第12図イの如きテーブルが、水平線テーブル,
メモリ32にセツトされる。 The horizontal line generation processing unit 31 performs the processes (f) and (g) above, and the original image data necessary for this purpose is sent from the first image memory 2 to the buffer 33. The resulting table as shown in Figure 12A is the horizontal line table,
It is set in memory 32.
バツフア34は、第1画像メモリ2から送出さ
れた原画像データがセツトされるものであり、水
平線テーブル・メモリ32に記入された水平線の
オーバラツプ領域に黒点が存在するか否かを検出
するため等に使用されるものである。 The buffer 34 is set with the original image data sent from the first image memory 2, and is used to detect whether a black point exists in the overlap area of the horizontal line written in the horizontal line table memory 32, etc. It is used for.
境界線抽出部35は上記(h)(i)の如き処理を行な
い、第13図に示す境界線用の垂直線を引き、こ
れにより第14図に示す如き文字毎の境界線を作
成する。そしてこの各境界線の座標データ境界線
テーブル36に記入される。 The boundary line extracting unit 35 performs the processing as in (h) and (i) above, draws a vertical line for the boundary line shown in FIG. 13, and thereby creates a boundary line for each character as shown in FIG. 14. Then, the coordinate data of each boundary line is entered in the boundary line table 36.
バツフア37は、第1画像メモリ2から原画像
データが送出されてこれにセツトされるものであ
る。そしてこの原画像データは境界線テーブル3
6に記入された境界線を示す座標データにもとづ
き抽出され、出力メモリ38に文字毎の画像デー
タとして出力される。 The buffer 37 is to which original image data is sent from the first image memory 2 and set therein. And this original image data is border table 3
6 is extracted based on the coordinate data indicating the boundary line written in 6, and outputted to the output memory 38 as image data for each character.
次に第15図の動作について簡単に説明する。 Next, the operation shown in FIG. 15 will be briefly explained.
(1) まず手書き原稿等の画像情報が入力部1で電
気信号に変換され、「1」、「0」の画像データ
となり、第1画像メモリ2にセツトされる。そ
してこの画像データはバツフア3およびバツフ
ア8に送出され、これらにもセツトされる。(1) First, image information such as a handwritten manuscript is converted into an electrical signal by the input unit 1, resulting in image data of "1" and "0", and is set in the first image memory 2. This image data is then sent to buffers 3 and 8 and set there as well.
(2) このようにバツフア3およびバツフア8に画
像データがセツトされた後、制御部13は第1
アドレス発生部6および第2アドレス発生部7
に対して制御信号を送出し、第1アドレス発生
部6に対してはバツフア3をx方向に走査する
ように、また第2アドレス発生部7に対しては
バツフア8をy方向に走査するように、それぞ
れアドレスを発生するよう制御する。これによ
り第1特徴抽出部4は上記(a)に示した変化点対
P1,P2を検出してこの変化点対P1〜P2の領域
を「1」となし、この「1」としたアドレス領
域を第1アドレス・テーブル5にセツトする。
一方第2特徴抽出部9は上記(b)に示した変化点
Q1,Q2の対を検出してこの変化点対Q1〜Q2の
領域を「1」となし、この「1」としたアドレ
ス領域を第2アドレス・テーブル10にセツト
する。(2) After the image data is set in the buffers 3 and 8 in this way, the control section 13
Address generation section 6 and second address generation section 7
A control signal is sent to the first address generating section 6 to scan the buffer 3 in the x direction, and to the second address generating section 7 to scan the buffer 8 in the y direction. control to generate an address for each. As a result, the first feature extraction unit 4 extracts the change point pair shown in (a) above.
P 1 and P 2 are detected and the area of this change point pair P 1 to P 2 is set to ``1'', and this address area set to ``1'' is set in the first address table 5.
On the other hand, the second feature extraction unit 9 extracts the change points shown in (b) above.
The pair of Q 1 and Q 2 is detected and the area of this change point pair Q 1 to Q 2 is set to "1", and this address area set to "1" is set in the second address table 10.
(3) このようにして変化点P1〜P2およびQ1〜Q2
の間の領域を「1」にした後、エクスクルシー
ブ・オア回路14で第1アドレス・テーブル5
および第2アドレス・テーブル10の「1」の
領域のエクスクルシーブ・オアをとり、これに
より、第4図に示す如き画像データが得られ、
これが第2画像メモリ12にセツトされる。(3) In this way, the change points P 1 ~ P 2 and Q 1 ~ Q 2
After setting the area between
and performs an exclusive OR on the area "1" in the second address table 10, thereby obtaining image data as shown in FIG.
This is set in the second image memory 12.
(4) このようにして得られた第4図の画像データ
は第2画像メモリ12からバツフア15にセツ
トされる。それから制御部13は第3アドレス
発生部14を制御してバツフア15をx方向に
走査し、これにより読出されたデータにもとづ
き、第1ラベル付処理部16では、上記(c−
1)のように、第6図イに示すマスクM2によ
りデータを読出したマスク処理を行なう。そし
てこの結果ラベル記憶が付与された画像データ
が第3画像メモリ17にセツトされる。(4) The image data of FIG. 4 obtained in this way is set from the second image memory 12 to the buffer 15. Then, the control section 13 controls the third address generation section 14 to scan the buffer 15 in the x direction, and based on the data thus read out, the first label processing section 16 performs the above (c-
As shown in 1), mask processing is performed in which data is read out using the mask M2 shown in FIG. 6A. As a result, the image data to which label storage has been added is set in the third image memory 17.
(5) 次に第3画像メモリ17にセツトされた画像
データはバツフア19に送出されてこれにセツ
トされる。このとき制御部13は第4アドレス
発生部18を制御してバツフア19をx方向に
走査し、これにより読出されたデータにもとづ
き、第2ラベル付処理部20では、上記(c−
2)のように、第7図ロに示すマスクM3によ
りデータを読出したマスク処理を行なう。そし
てこの結果同一領域内に同一のラベル記号が付
与されることになる。(5) Next, the image data set in the third image memory 17 is sent to the buffer 19 and set therein. At this time, the control section 13 controls the fourth address generation section 18 to scan the buffer 19 in the x direction, and based on the data thus read out, the second labeling processing section 20 performs the above (c-
2), a masking process is performed in which the data is read using the mask M3 shown in FIG. 7B. As a result, the same label symbol is given within the same area.
(6) この第2ラベル付処理部20から出力された
画像データは、第4画像メモリ4にセツトされ
たのちラベル整列回路22により、上記(c−
3)の処理が行なわれて、第5図に示す如き状
態にラベル記号がシーケンシヤルなものに整列
された後、第5画像メモリ23にセツトされ
る。(6) The image data output from the second labeling processing unit 20 is set in the fourth image memory 4 and then processed by the label alignment circuit 22 (c-
After the process 3) is performed and the label symbols are sequentially arranged as shown in FIG. 5, they are set in the fifth image memory 23.
(7) 第5画像メモリ23にセツトされた第5図に
示す如くラベル記号の付与された画像データ
は、バツフア24に送出される。このとき制御
部13は第5アドレス発生部25に対して制御
信号を送出し、これをx方向に走査する。そし
てこの結果出力されたデータにもとづき中点抽
出回路26が上記dの如き処理を行ない、かく
して得られた第8図ロの如き中点の軌跡が第1
中点アドレス・テーブル27に記入された第9
図に示す中点の軌跡が得られる。(7) The image data set in the fifth image memory 23 and provided with a label symbol as shown in FIG. 5 is sent to the buffer 24. At this time, the control section 13 sends a control signal to the fifth address generation section 25 to scan it in the x direction. Then, based on the data outputted as a result, the midpoint extraction circuit 26 performs the process as described in d above, and the trajectory of the midpoint as shown in FIG.
The ninth address entered in the midpoint address table 27
The trajectory of the midpoint shown in the figure is obtained.
(8) この中点の軌跡データは境界線連絡処理部2
8に伝達される。またバツフア29には第1画
像メモリ2から送出された原画像データがセツ
トされておりこの原画像データを参照しつつ境
界線連絡処理部28では上記(e)の処理が行なわ
れる。そしてこの結果第2中点アドレス・テー
ブル30に第10図ロに示す如きラベル記号の
付与されたテーブルが記入される。(8) The trajectory data of this midpoint is the boundary line communication processing unit 2.
8. Further, the original image data sent from the first image memory 2 is set in the buffer 29, and the boundary line communication processing section 28 performs the process (e) above while referring to this original image data. As a result, a table with label symbols as shown in FIG. 10B is entered in the second midpoint address table 30.
(9) この第2中点アドレス・テーブル30にセツ
トされたデータは水平線発生処理部31に伝達
され、上記(f)、(g)の処理が行なわれる。このと
き必要な原画像データはバツフア33にセツト
されており、水平線発生処理部31はこのバツ
フア33にセツトされた原画像データを参照し
つつ上記(f)、(g)の如き処理を行なう。そしてこ
の結果得られたデータが、第12図イに示すよ
うな状態で水平線テーブル・メモリ32にセツ
トされることになる。(9) The data set in the second midpoint address table 30 is transmitted to the horizontal line generation processing section 31, where the above processes (f) and (g) are performed. At this time, the necessary original image data is set in the buffer 33, and the horizontal line generation processing section 31 performs the processes (f) and (g) above while referring to the original image data set in the buffer 33. The data obtained as a result is then set in the horizontal line table memory 32 in the state shown in FIG. 12A.
(10) 次に境界線抽出部35は、上記水平線テーブ
ル・メモリ32から伝達されたデータと、バツ
フア34にセツトされている原画像データにも
とづき、上記(h)(i)の如き処理を行なう。かくし
て第13図に示す如き境界線用の垂直線が引か
れ、これにもとづき、第14図に示すような文
字毎の境界線を作成することができる。そして
この第14図に太線で示した各境界線の座標デ
ータが境界線テーブル36にセツトされる。バ
ツフア37には第1画像メモリ2から原画像デ
ータが送出されているので、この境界線座標デ
ータにもとづき各文字毎にその原画像データが
抽出することができ、出力メモリ38にセツト
することができる。かくして出力メモリ38か
ら必要とする文字情報が文字毎に個別に得るこ
とができるので、これを文字読取装置等に入力
することにより手書き文字を認識することがで
きる。(10) Next, the boundary line extraction unit 35 performs the processing as in (h) and (i) above based on the data transmitted from the horizontal line table memory 32 and the original image data set in the buffer 34. . In this way, vertical lines for boundaries as shown in FIG. 13 are drawn, and based on these, boundaries for each character as shown in FIG. 14 can be created. The coordinate data of each boundary line shown in thick lines in FIG. 14 is then set in the boundary line table 36. Since the original image data is sent from the first image memory 2 to the buffer 37, the original image data can be extracted for each character based on this boundary line coordinate data and set in the output memory 38. can. In this way, the required character information can be obtained individually for each character from the output memory 38, so that handwritten characters can be recognized by inputting this information into a character reading device or the like.
以上説明の如く、結局本発明によれば、複数の
文字が文字枠等の記入位置に無関係に且つランダ
ムに互いに近接して手書きされているような場合
でも、その文字間の中間に境界線を引くことがで
き、各文字を正確に分離することができる。それ
故、文字情報を個別に抽出することができるの
で、手書き文字の認識等に非常に有効な文字分離
手段を提供することができる。 As explained above, according to the present invention, even when a plurality of characters are handwritten randomly and close to each other, regardless of the writing position in the character frame etc., a border line can be created in the middle between the characters. and can accurately separate each letter. Therefore, since character information can be extracted individually, it is possible to provide a character separation means that is very effective in recognizing handwritten characters.
第1図は文字読取用のマスクおよび、該マスク
により文字分離が可能な場合および不可能な場合
の説明図、第2図は走査状態説明図、第3図〜第
14図は本発明の動作状態説明図、第15図は本
発明の一実施例構成図である。
図中、1は入力部、2は第1画像メモリ、3は
バツフア、4は第1特徴抽出部、5は第1アドレ
ス・テーブル、6は第1アドレス発生部、7は第
2アドレス発生部、8はバツフア、9は第2特徴
抽出部、10は第2アドレス・テーブル、11は
エクスクルシーブ・オア回路、12は第2画像メ
モリ、13は制御部、14は第3アドレス発生
部、15はバツフア、16は第1ラベル付処理
部、17は第3画像メモリ、18は第4アドレス
発生部、19はバツフア、20は第2ラベル付処
理部、21は第4画像メモリ、22はラベル整列
回路、23は第5画像メモリ、24はバツフア、
25は第5アドレス発生部、26は中点抽出回
路、27は第1中点アドレス・テーブル、28は
境界線処理部、29はバツフア、30は第2中点
アドレス・テーブル、31は水平線発生処理部、
32は水平線テーブル・メモリ、33,34はバ
ツフア、35は境界線抽出部、36は境界線テー
ブル、37はバツフア、38は出力メモリをそれ
ぞれ示す。
FIG. 1 is an explanatory diagram of a mask for character reading and cases in which character separation is possible and impossible using the mask, FIG. 2 is an explanatory diagram of the scanning state, and FIGS. 3 to 14 are diagrams illustrating the operation of the present invention. A state explanatory diagram, FIG. 15, is a configuration diagram of an embodiment of the present invention. In the figure, 1 is an input section, 2 is a first image memory, 3 is a buffer, 4 is a first feature extraction section, 5 is a first address table, 6 is a first address generation section, and 7 is a second address generation section. , 8 is a buffer, 9 is a second feature extraction section, 10 is a second address table, 11 is an exclusive OR circuit, 12 is a second image memory, 13 is a control section, 14 is a third address generation section, 15 is a buffer, 16 is a first label processing section, 17 is a third image memory, 18 is a fourth address generation section, 19 is a buffer, 20 is a second label processing section, 21 is a fourth image memory, and 22 is a buffer. a label alignment circuit, 23 a fifth image memory, 24 a buffer,
25 is a fifth address generator, 26 is a midpoint extraction circuit, 27 is a first midpoint address table, 28 is a boundary line processing unit, 29 is a buffer, 30 is a second midpoint address table, and 31 is a horizontal line generator. processing section,
32 is a horizontal line table memory, 33 and 34 are buffers, 35 is a boundary line extractor, 36 is a boundary line table, 37 is a buffer, and 38 is an output memory.
Claims (1)
持する情報保持手段と、上記文字情報を第1の方
向に走査してその白黒変化点を求め、該変化点の
間の領域である第1領域を検出する第1特徴抽出
手段と、上記画像情報を、前記第1の方向と略直
交する第2の方向に走査してその白黒変化点を求
め、該変化点の間の領域である第2領域を検出す
る第2特徴抽出手段と、上記第1領域と第2領域
とが重畳しない領域である中間領域を前記第1の
方向に順次走査し走査線分を得る走査手段と、上
記中間領域の各走査線分における中間点を求める
中間点設定手段と、上記文字の画像情報に近接し
て前記中間点から第1の方向に水平線を発生する
水平線発生手段を有し、上記中間点列および上記
水平線により個々の文字情報の存在する文字領域
を区別できるようにしたことを特徴とする文字分
離方式。1. An information holding means for holding image information of characters inputted from an input means, and a first area that scans the character information in a first direction to obtain points of black and white change, and is an area between the points of change. a first feature extraction means for detecting a black and white change point by scanning the image information in a second direction substantially perpendicular to the first direction; a second feature extraction means for detecting a region; a scanning means for sequentially scanning in the first direction an intermediate region where the first region and the second region do not overlap to obtain scanning line segments; and a scanning means for obtaining scanning line segments; intermediate point setting means for determining the intermediate point in each scanning line segment; horizontal line generating means for generating a horizontal line in a first direction from the intermediate point in proximity to the image information of the character; A character separation method characterized in that character areas in which individual character information exists can be distinguished by the horizontal lines.
Priority Applications (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP55187608A JPS57111784A (en) | 1980-12-29 | 1980-12-29 | Character separation and pickup system |
| DE8282900151T DE3177075D1 (en) | 1980-12-29 | 1981-12-28 | Character and figure isolating and extracting system |
| PCT/JP1981/000424 WO1982002268A1 (en) | 1980-12-29 | 1981-12-28 | Character and figure isolating and extracting system |
| EP82900151A EP0067236B1 (en) | 1980-12-29 | 1981-12-28 | Character and figure isolating and extracting system |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP55187608A JPS57111784A (en) | 1980-12-29 | 1980-12-29 | Character separation and pickup system |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPS57111784A JPS57111784A (en) | 1982-07-12 |
| JPS6330666B2 true JPS6330666B2 (en) | 1988-06-20 |
Family
ID=16209081
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP55187608A Granted JPS57111784A (en) | 1980-12-29 | 1980-12-29 | Character separation and pickup system |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPS57111784A (en) |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS56166587A (en) * | 1980-05-28 | 1981-12-21 | Toshiba Corp | Character segmenting system |
-
1980
- 1980-12-29 JP JP55187608A patent/JPS57111784A/en active Granted
Also Published As
| Publication number | Publication date |
|---|---|
| JPS57111784A (en) | 1982-07-12 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| KR20010102224A (en) | Apparatus and system for reproduction of handwritten input | |
| JP2846486B2 (en) | Image input device | |
| JP2000113111A (en) | Image recognition feature value extraction method and apparatus, storage medium for storing image analysis program | |
| JPH03142691A (en) | Table format document recognizing system | |
| JPS6330665B2 (en) | ||
| JPH02138674A (en) | Document processing method and device | |
| JP3124854B2 (en) | Character string direction detector | |
| JPS61150081A (en) | Character recognizing device | |
| JP3594625B2 (en) | Character input device | |
| JP3903540B2 (en) | Image extraction method and apparatus, recording medium on which image extraction program is recorded, information input / output / selection method and apparatus, and recording medium on which information input / output / selection processing program is recorded | |
| JP2978801B2 (en) | Character input method for handwritten character recognition | |
| JP3502130B2 (en) | Table recognition device and table recognition method | |
| JPH02176973A (en) | Drawing read processing method | |
| JPS6327752B2 (en) | ||
| JP2762476B2 (en) | Copy-writing device | |
| JPH08297718A (en) | Character segmentation device and character recognition device | |
| JPH06150056A (en) | Table recognition device | |
| JPH053631B2 (en) | ||
| JPH0119189B2 (en) | ||
| JPH08185475A (en) | Image recognition device | |
| JPS6097480A (en) | Apex extracting system of closed area graphic | |
| JPH0160870B2 (en) | ||
| JPH058670U (en) | Optical character reader | |
| JPS6292080A (en) | Character pattern recognition correction device | |
| JPS6316786B2 (en) |