JPH04264993A - One character segmenting method - Google Patents

One character segmenting method

Info

Publication number
JPH04264993A
JPH04264993A JP3026020A JP2602091A JPH04264993A JP H04264993 A JPH04264993 A JP H04264993A JP 3026020 A JP3026020 A JP 3026020A JP 2602091 A JP2602091 A JP 2602091A JP H04264993 A JPH04264993 A JP H04264993A
Authority
JP
Japan
Prior art keywords
character
characters
storage area
stored
product
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP3026020A
Other languages
Japanese (ja)
Inventor
Yasuhiko Murayama
靖彦 村山
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Seiko Epson Corp
Original Assignee
Seiko Epson Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Seiko Epson Corp filed Critical Seiko Epson Corp
Priority to JP3026020A priority Critical patent/JPH04264993A/en
Publication of JPH04264993A publication Critical patent/JPH04264993A/en
Pending legal-status Critical Current

Links

Landscapes

  • Character Input (AREA)

Abstract

PURPOSE:To obtain a character separate position deciding method to deal with connected characters with different character width and character intervals in a document and operated at high processing speed by using simultaneously the peripheral distribution and line density of a part of the connected characters. CONSTITUTION:The peripheral distribution and the line density are found from a character image, and are stored in the storage area of a RAM 13, respectively. Thence, the product of the peripheral distribution stored in a peripheral distribution storage area 13e and the line density stored in a line density storage area 13f is found, and a result is stored in a product storage area, however, the product can be found by multiplying the value of the peripheral distribution by that of the line density at the same scanning position. The separate position of the connected characters can be decided by using a numeric value stored in the product storage area 13g. A deciding method is to set the scanning position where the minimum product value can be obtained as the separate position of the character.

Description

【発明の詳細な説明】[Detailed description of the invention]

【0001】0001

【産業上の利用分野】本発明は、文字認識装置において
、文書のイメージ情報中で連結した文字を分離する一文
字切り出し方法に関する。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a single character extraction method for separating connected characters in image information of a document in a character recognition device.

【0002】0002

【従来の技術】文字認識装置において、文書自体に文字
同士の黒画素の連結がある、もしくはスキャナーの解像
度が十分でないことにより文字同士の黒画素の連結を生
じることがある。このように連結した文字は強制的に分
離をする必要がある。
2. Description of the Related Art In character recognition devices, black pixels may be connected between characters in the document itself, or because the resolution of the scanner is insufficient. Characters concatenated in this way must be forcibly separated.

【0003】そこで、文字が連結している場合、従来は
一文字に分離するために、例えば横書き文書の場合、連
結した文字の垂直方向の射影(水平軸に対する射影)を
求め、射影した値が特定の数値以下の部分で強制的に分
離する方法や、連結した文字を含む文字列全体の垂直方
向の射影を求め、射影した値がゼロとなる位置の周期性
から連結した文字の分離位置を決定した。
Therefore, when characters are connected, conventionally, in order to separate them into single characters, for example, in the case of a horizontally written document, the vertical projection (projection on the horizontal axis) of the connected characters is calculated, and the projected value is specified. A method of forcibly separating characters that is less than or equal to the value of , or finding the vertical projection of the entire string containing connected characters, and determining the separation position of the connected characters from the periodicity of the position where the projected value is zero. did.

【0004】0004

【発明が解決しようとする課題】しかしながら、上記の
連結した文字の分離方法によると、射影した値が特定の
数値以下となる部分が複数個存在し、分離位置が決定で
きなかったり、文字幅、文字間隔が等しくないために、
射影した値がゼロとなる位置のの周期性がなく、連結し
た文字の分離位置を誤ったりすることがあった。
[Problems to be Solved by the Invention] However, according to the method for separating connected characters described above, there are multiple parts where the projected value is less than a specific value, and the separation position cannot be determined, or the character width or Due to unequal letter spacing,
There was no periodicity in the position where the projected value was zero, and the separation position of connected characters was sometimes incorrect.

【0005】そこで本発明はこのような問題点を解決す
るもので、その目的とするところは連結した文字の部分
の周辺分布と線密度を併用することにより、文字幅や文
字間隔が等しくない文書中にある連結した文字に対応で
き、しかも文字の分離位置を探索する範囲のみを処理す
るので処理速度が速く、文字分離位置を決定できる方法
を提供するところにある。
[0005]The present invention is intended to solve these problems, and its purpose is to solve the problems of documents with unequal character widths and character spacing by using the peripheral distribution and line density of connected character parts together. The object of the present invention is to provide a method that can deal with connected characters in the text, has a high processing speed, and can determine character separation positions because it processes only the range in which character separation positions are to be searched.

【0006】[0006]

【課題を解決するための手段】本発明による文字分離方
法は、処理対象の各行に連結した文字がある場合、文字
の分離位置を探索する範囲を設定し、設定した範囲内で
連結した文字のイメージから周辺分布と線密度を求め、
求めた周辺分布と線密度の各走査位置での積を求め、求
めた積から連結した文字の分離位置を決定することを特
徴とする。
[Means for Solving the Problems] In the character separation method according to the present invention, when there are connected characters in each line to be processed, a range is set to search for the character separation position, and the connected characters are searched within the set range. Find the marginal distribution and line density from the image,
The method is characterized in that the product of the obtained marginal distribution and line density at each scanning position is obtained, and the separation position of connected characters is determined from the obtained product.

【0007】[0007]

【実施例】(実施例1)以下本発明の実施例につき図面
を用いて詳細に説明する。
EXAMPLES (Example 1) Examples of the present invention will be described below in detail with reference to the drawings.

【0008】図1は本発明の実施に必要な装置構成を示
すブロック図である。10は文書画像を入力するための
スキャナー、11は処理を実行するためのCPU、12
は各処理のプログラムを格納したROM、13は画像や
処理に関連したデータ、処理結果を格納するためのRA
Mである。図1では各処理のプログラムをROM12に
格納したが、ROM12の代りにRAMに各処理のプロ
グラムをロードしてから処理を始めてもかまわない。
FIG. 1 is a block diagram showing the equipment configuration necessary for implementing the present invention. 10 is a scanner for inputting document images; 11 is a CPU for executing processing; 12
13 is a ROM that stores programs for each process, and RA is for storing data related to images and processing, and processing results.
It is M. Although the program for each process is stored in the ROM 12 in FIG. 1, the program for each process may be loaded in the RAM instead of the ROM 12 before starting the process.

【0009】以上の装置の構成例による処理内容につい
て説明する。
[0009] The processing contents of the above-mentioned apparatus configuration example will be explained.

【0010】スキャナー10によって入力した文書の画
像はRAM13内の画像イメージ格納領域13aに蓄え
られる。なお、ここで扱う分書の画像は2値、すなわち
0(白画素)と1(黒画素)で構成されたものを対象と
する。画像イメージ格納領域13aに蓄えた文書画像に
対し、CUP11は行切り出しプログラム12aに従っ
て行の切り出しを行い、切り出した行のイメージをRA
M13内の行イメージ格納領域13bに格納する。この
行切り出しは、例えば文書(ここでは横書き文書とする
)の水平方向の射影(垂直軸に対する射影)を測定し、
射影の値が特定の数値を越える範囲を行の範囲として切
り出す方法などによって行う。なお、行を複数のブロッ
クに分割し、ブロック毎の水平射影によって行切り出し
を行う方法など、行切り出しの方法自体は任意である。
The image of the document inputted by the scanner 10 is stored in an image storage area 13a in the RAM 13. Note that the image of the separate book handled here is a binary image, that is, one composed of 0 (white pixel) and 1 (black pixel). The CUP 11 cuts out lines from the document image stored in the image storage area 13a according to the line cutting program 12a, and converts the image of the cut lines into RA.
It is stored in the row image storage area 13b in M13. This line cutting is performed by, for example, measuring the horizontal projection (projection on the vertical axis) of a document (in this case, a horizontally written document).
This is done by cutting out the range where the projection value exceeds a specific numerical value as a row range. Note that the line extraction method itself may be arbitrary, such as a method of dividing a line into a plurality of blocks and performing line extraction by horizontal projection for each block.

【0011】切り出した行イメージから、CPU11は
文字切り出しプログラム12bに従い文字の切り出しを
行い、切り出した文字のイメージをRAM13内の文字
イメージ格納領域13cに格納する。それと同時に文字
幅を求め、RAM13内の文字幅格納領域13dに格納
する。この文字の切り出しは、例えば文字例(ここでは
横書き文書とする)の垂直方向の射影(水平軸に対する
射影)を測定し、射影の数値がゼロより大きい数値とな
る部分を文字の範囲として切り出す方法などによって行
う。文字幅を求める方法は、切り出した文字の範囲の平
均値等を用いることによって行う。また、文字幅は予め
与えられてRAM13内の文字幅格納領域13dに格納
されていてもかまわない。なお、文字の切り出し、文字
幅の計算の方法自体は任意である。
From the cut out line image, the CPU 11 cuts out characters according to the character cutting program 12b, and stores the cut out character images in the character image storage area 13c in the RAM 13. At the same time, the character width is determined and stored in the character width storage area 13d in the RAM 13. This method of cutting out characters is, for example, by measuring the vertical projection (projection on the horizontal axis) of a character example (in this case, a horizontally written document), and then cutting out the part where the projection value is greater than zero as the range of characters. etc. The character width is determined by using the average value of the range of cut out characters. Further, the character width may be given in advance and stored in the character width storage area 13d in the RAM 13. Note that the method of cutting out characters and calculating the character width is arbitrary.

【0012】一文字切り出しにおいて、文書自体に文字
同士の黒画素の連結がある、もしくはスキャナー10の
解像度が十分でないことにより文字同士の黒画素の連結
を生じることがある。そこでCPU11は文字分離プロ
グラム12cに従い、文字同士が連結したまま文字イメ
ージの格納領域13cに登録されたものがないかをチェ
クし、文字同士が連結したまま文字イメージの格納領域
13cに登録されたと判断した文字イメージについては
文字の分離を行い、再登録を行う処理をする。図2を参
照し、文字分離プログラム12cの処理の流れを説明す
る。
When cutting out a single character, there may be a connection of black pixels between characters in the document itself, or there may be a connection of black pixels between characters due to insufficient resolution of the scanner 10. Therefore, the CPU 11 checks whether any characters are registered in the character image storage area 13c while being connected to each other according to the character separation program 12c, and determines that the characters are registered in the character image storage area 13c while being connected to each other. For character images that have been registered, the characters are separated and re-registered. The flow of processing of the character separation program 12c will be explained with reference to FIG.

【0013】CPU11が文字切り出しプログラム12
bに従い、登録した文字イメージの幅と、求めた文字幅
をα倍した値を比較する(ステップ202)。ここでα
は文字同士が連結した場合の幅を示す数値であり、例え
ばα=2とする。
[0013] The CPU 11 executes the character cutting program 12.
According to b, the width of the registered character image is compared with the obtained character width multiplied by α (step 202). Here α
is a numerical value indicating the width when characters are connected to each other; for example, α=2.

【0014】比較した結果、文字イメージの幅が文字幅
をα倍した数値以上であったならば、2文字以上が連結
していると判断し、文字分離の位置を探索する範囲を設
定する(ステップ203)。文字分離の位置を探索する
範囲を文字イメージの先頭から文字幅進んだ位置を中心
に、文字幅のβ倍左右に進んだ範囲とする。ここでβは
文字分離の位置を探索する範囲を設定するときに用いる
数値で、例えばβ=0.25とする。これは印刷文字に
おいて同じポイント数の文字の場合、文字の仮想ボディ
(図3、301)は同じであるが、文字に外接する矩形
で示される字面(3図、302)は文字によって異なり
、仮想ボディの幅に対して、字面の幅は平均5%から2
5%小さいことにより設定した数値である。
As a result of the comparison, if the width of the character image is equal to or larger than the character width multiplied by α, it is determined that two or more characters are connected, and a range to search for character separation positions is set ( Step 203). The range in which the character separation position is searched is centered on the position leading by the character width from the beginning of the character image, and is set to the range β times the character width to the left and right. Here, β is a numerical value used when setting the range to search for character separation positions, and is set to β=0.25, for example. This is because when printed characters have the same number of points, the virtual bodies of the characters (Fig. 3, 301) are the same, but the face (Fig. 3, 302) shown by the rectangle circumscribing the characters differs depending on the character, and the virtual body of the characters (Fig. 3, 301) is the same. The width of the font is on average 5% to 2% of the width of the body.
This value was set based on the fact that it is 5% smaller.

【0015】ステップ203で設定した範囲内で、文字
イメージから周辺分布を求め、RAM13内の周辺分布
格納領域13eに格納する(ステップ204)。横書き
文書の場合、周辺分布は垂直方向に走査し、画素値1が
(黒画素)の画素を各走査位置毎に計数することにより
求める。
The peripheral distribution is determined from the character image within the range set in step 203, and is stored in the peripheral distribution storage area 13e in the RAM 13 (step 204). In the case of a horizontally written document, the peripheral distribution is obtained by scanning in the vertical direction and counting pixels with a pixel value of 1 (black pixels) at each scanning position.

【0016】次にステップ203で設定した範囲内で、
文字イメージから線密度を求め、RAM13内の線密度
各領域13fに格納する(ステップ205)。横書き文
書の場合、線密度は垂直方向に走査し、画素値が0から
1に反転する場所を各走査位置毎に計数することにより
求める。
Next, within the range set in step 203,
The line density is determined from the character image and stored in each line density area 13f in the RAM 13 (step 205). In the case of a horizontally written document, the line density is determined by scanning in the vertical direction and counting the locations where the pixel value is reversed from 0 to 1 at each scanning position.

【0017】周辺分布格納領域13eに格納された周辺
分布と、線密度格納領域13fに格納された線密度の積
を求め。RAM内の積格納領域13gに格納する(ステ
ップ206)。積は同じ走査位置の周辺分布の値と線密
度の値を掛けることにより求める。
The product of the marginal distribution stored in the marginal distribution storage area 13e and the linear density stored in the linear density storage area 13f is calculated. It is stored in the product storage area 13g in the RAM (step 206). The product is obtained by multiplying the value of the marginal distribution and the value of the linear density at the same scanning position.

【0018】積格納領域13gに格納された数値を用い
て、連結した文字の分離位置の決定を行う(ステップ2
07)。位置の決定の方法は、積の値が最小となる走査
位置を連結した文字の分離位置とする。最小値が範囲内
に複数個存在する場合には、文字イメージの先頭から、
文字幅進んだ位置にいちばん近い最小値の走査位置を文
字分離の位置とする。
Using the numerical value stored in the product storage area 13g, the separation position of the connected characters is determined (step 2
07). The position is determined by determining the scanning position where the product value is the minimum as the separation position of the connected characters. If there are multiple minimum values within the range, from the beginning of the character image,
The scanning position of the minimum value closest to the position advanced by the character width is set as the character separation position.

【0019】処理した文字イメージを文字イメージ格納
領域13cから削除し、ステップ205で、決定した文
字の分離位置で、文字イメージを分離し、改めて文字イ
メージ格納領域13cに再登録する(ステップ208)
。
The processed character image is deleted from the character image storage area 13c, separated at the character separation position determined in step 205, and re-registered in the character image storage area 13c (step 208).
.

【0020】次に、文字イメージ格納領域13cに、ス
テップ202で用いた条件を満たす文字イメージが残っ
ているかどうかを判断し、存在しない場合にはCPU1
1は文字分離プログラム12cの処理を終了する。存在
する場合には同様の処理を続ける(ステップ201)。
Next, it is determined whether a character image satisfying the conditions used in step 202 remains in the character image storage area 13c, and if there is no character image, the CPU 1
1 ends the processing of the character separation program 12c. If it exists, similar processing is continued (step 201).

【0021】図4にCPU11が文字分離プログラム1
2cに従い、連結した文字イメージから分離位置を決定
する例を示す。図4(a)は文字イメージ格納領域13
cに連結したまま登録した文字イメージの例である。 「煙」という文字と「道」という文字が連結しているの
がわかる。また図4(a)に示す範囲rは、ステップ2
03で設定された文字の分離位置を探索する範囲である
。図4(b)は、設定した範囲内で図4(a)から求め
た周辺分布である。周辺分布の最小値はポイントp11
からp12まで2ポイントみられる。図4(c)は、設
定した範囲内で、図4(a)から求めた線密度である。 線密度の最小値はポイントp21からp22まで2ポイ
ントみられる。図4(d)は、図4(b)と図4(c)
から求めた積である。最小値はポイントp31のみで、
文字の分離位置はポイントp31と決定する。
FIG. 4 shows that the CPU 11 runs the character separation program 1.
An example of determining a separation position from connected character images according to 2c is shown below. FIG. 4(a) shows the character image storage area 13.
This is an example of a character image registered while being connected to c. You can see that the characters ``smoke'' and ``road'' are connected. Moreover, the range r shown in FIG. 4(a) is
This is the range to search for the character separation position set in step 03. FIG. 4(b) shows the marginal distribution obtained from FIG. 4(a) within the set range. The minimum value of the marginal distribution is point p11
2 points can be seen from to p12. FIG. 4(c) shows the linear density determined from FIG. 4(a) within the set range. The minimum value of the linear density is found at two points from point p21 to p22. Figure 4(d) is similar to Figure 4(b) and Figure 4(c).
This is the product obtained from . The minimum value is only at point p31,
The character separation position is determined to be point p31.

【0022】図4に示すように、周辺分布や線密度のみ
では最小値となる走査位置が複数あり、連結した文字の
分離位置を決定しがたい。しかし、文字同士が連結する
とき、連結する部分が1カ所である確率が高いため、周
辺分布に線密度を掛け、重み付けをすることにより文字
同士が連結した位置を決定するのが容易になる。
As shown in FIG. 4, there are a plurality of scanning positions where the minimum value is obtained based only on the marginal distribution and linear density, and it is difficult to determine the separation position of connected characters. However, when characters are connected, there is a high probability that there is only one connected part, so by multiplying the marginal distribution by the linear density and weighting it, it becomes easier to determine the position where the characters are connected.

【0023】[0023]

【発明の効果】以上説明したように本発明によれば、連
結した文字の分離において周辺分布と線密度を併用して
文字の分離位置の決定を行うことにより、文字の間隔が
不定の文字列中にある連結した文字の分離が可能であり
、文字の分離位置の候補を絞ることができ、しかも文字
の分離位置を探索する範囲を設定するので処理速度が速
く、分離位置の決定ができるという効果を有する。
Effects of the Invention As explained above, according to the present invention, character separation positions are determined by using both marginal distribution and linear density in separating connected characters, thereby improving character strings with irregular character intervals. It is possible to separate connected characters inside, narrow down candidates for character separation positions, and set the search range for character separation positions, so processing speed is fast and separation positions can be determined. have an effect.

【図面の簡単な説明】[Brief explanation of the drawing]

【図1】本発明の実施に必要な装置の構成例を示すブロ
ック図である。
FIG. 1 is a block diagram showing an example of the configuration of a device necessary for implementing the present invention.

【図2】本発明の一文字切り出しの処理の流れを示す流
れ図である。
FIG. 2 is a flowchart showing the flow of a single character extraction process according to the present invention.

【図3】文字の大きさについて説明するための図である
。
FIG. 3 is a diagram for explaining the size of characters.

【図4】図2で示す処理により、連結した文字から分離
位置を決定する例を示す図である。
FIG. 4 is a diagram showing an example of determining a separation position from connected characters by the process shown in FIG. 2;

【符号の説明】[Explanation of symbols]

10  スキャナー 11  CPU 12  ROM 13  RAM 10 Scanner 11 CPU 12 ROM 13 RAM

Claims (1)

【特許請求の範囲】[Claims] 【請求項1】処理対象の文書の各行において、行中に連
結した文字が存在する場合、文字の分離位置を探索する
範囲を設定し、前記設定範囲内で、連結した文字イメー
ジから周辺分布、線密度を求め、前記周辺分布、前記線
密度の対応する各走査位置での積を求め、前記積が最小
値となる走査位置を連結した文字の分離位置とすること
を特徴とする一文字切り出し方法。
1. In each line of a document to be processed, if there are connected characters in the line, a range is set to search for character separation positions, and within the set range, peripheral distribution is determined from the connected character images. A method for cutting out one character, characterized in that the linear density is determined, the product of the peripheral distribution and the linear density at each corresponding scanning position is determined, and the scanning position where the product has a minimum value is set as the separation position of the connected characters. .
JP3026020A 1991-02-20 1991-02-20 One character segmenting method Pending JPH04264993A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP3026020A JPH04264993A (en) 1991-02-20 1991-02-20 One character segmenting method

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP3026020A JPH04264993A (en) 1991-02-20 1991-02-20 One character segmenting method

Publications (1)

Publication Number Publication Date
JPH04264993A true JPH04264993A (en) 1992-09-21

Family

ID=12182017

Family Applications (1)

Application Number Title Priority Date Filing Date
JP3026020A Pending JPH04264993A (en) 1991-02-20 1991-02-20 One character segmenting method

Country Status (1)

Country Link
JP (1) JPH04264993A (en)

Similar Documents

Publication Publication Date Title
JP7026165B2 (en) Text recognition method and text recognition device, electronic equipment, storage medium
US5075895A (en) Method and apparatus for recognizing table area formed in binary image of document
JP2940936B2 (en) Tablespace identification method
US5889885A (en) Method and apparatus for separating foreground from background in images containing text
CN111461133B (en) Express delivery surface single item name identification method, device, equipment and storage medium
US4556985A (en) Pattern recognition apparatus
CN115223172A (en) Text extraction method, device and equipment
JPH07160812A (en) Image processing apparatus and method
JPS6325391B2 (en)
JP3548234B2 (en) Character recognition method and device
JPH04248688A (en) Method for segmenting one character
CN111414919A (en) Method, device and equipment for extracting characters from printed pictures with forms and storage medium
JPS6254380A (en) character recognition device
JP3000480B2 (en) Character area break detection method
JP2982221B2 (en) Character reader
CN115731250A (en) Text segmentation method, device, equipment and storage medium
JPH04264687A (en) Character recognition processing method
JPH05114047A (en) Device for segmenting character
JPH05298487A (en) Alphabet recognizing device
CN115188009A (en) Character recognition method and device, electronic equipment and computer readable medium
JPH0343879A (en) Character area separating system for character recognizing device
JP2784059B2 (en) Method and apparatus for removing noise from binary image
JPH0573718A (en) Area attribute identification method
JP2974167B2 (en) Large Classification Recognition Method for Characters
JPH0344788A (en) Area extracting method for document picture