JPH06259601A - Character recognition device - Google Patents

Character recognition device

Info

Publication number
JPH06259601A
JPH06259601A JP5077693A JP7769393A JPH06259601A JP H06259601 A JPH06259601 A JP H06259601A JP 5077693 A JP5077693 A JP 5077693A JP 7769393 A JP7769393 A JP 7769393A JP H06259601 A JPH06259601 A JP H06259601A
Authority
JP
Japan
Prior art keywords
character
pattern
character pattern
standard
input
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP5077693A
Other languages
Japanese (ja)
Inventor
Yasumasa Araki
保昌 荒木
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Dainippon Screen Manufacturing Co Ltd
Original Assignee
Dainippon Screen Manufacturing Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Dainippon Screen Manufacturing Co Ltd filed Critical Dainippon Screen Manufacturing Co Ltd
Priority to JP5077693A priority Critical patent/JPH06259601A/en
Publication of JPH06259601A publication Critical patent/JPH06259601A/en
Pending legal-status Critical Current

Links

Landscapes

  • Character Discrimination (AREA)

Abstract

PURPOSE:To raise the probability that characters are correctly recognized by determining the character pattern, where the number of picture elements of line parts is largest, among standard character patterns included in a character picture as a recognized character. CONSTITUTION:The character picture of each character is segmented and is converted to a dot pattern (binary picture) One input character pattern PI as the object of character recognition is selected, and the number Nb of black picture elements of this pattern is calculated, and one standard character pattern PR whose number of black picture elements is smaller than Nb is selected. When AND between the standard character pattern PR and the input character pattern PI is equal to the standard character pattern PR, it is judged that the standard character pattern PR is included in the input, character pattern PI. In this case, the character code of the standard character pattern PR is determined as that of the input character pattern. Thus, the number of standard character patterns used for recognition of character pictures is reduced.

Description

【発明の詳細な説明】Detailed Description of the Invention

【0001】[0001]

【産業上の利用分野】この発明は、画像入力装置で読取
られた文字画像から文字を認識する装置に関する。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a device for recognizing characters from a character image read by an image input device.

【0002】[0002]

【従来の技術】文字画像から文字を認識する装置として
は、特開昭52−42025号公報や特開昭60−23
0279号公報に記載されているものが知られている。
これらの装置では、標準の文字パターンを太らせた第1
のパターンと細らせた第2のパターンとを準備してお
き、第1と第2のパターンと文字画像とを比較すること
によって文字を認識するようにしていた。
2. Description of the Related Art As a device for recognizing a character from a character image, there are Japanese Patent Laid-Open Nos. 52-42025 and 60-23.
The one described in Japanese Patent No. 0279 is known.
In these devices, the standard character pattern is thickened first
The pattern and the thinned second pattern are prepared, and the character is recognized by comparing the first and second patterns with the character image.

【0003】[0003]

【発明が解決しようとする課題】上述の従来の技術で
は、1つの文字に対して2つの文字パターンを準備しな
ければならず、また、同じ文字コードの文字であっても
書体が違えばそれぞれ別個の文字パターンを準備する必
要があった。例えば、日本語の1書体は約6000個の
文字を含むので、5書体の文字に対して文字パターンを
準備するとすれば、6000×5×2=60000個と
いう膨大な数の文字パターンを準備する必要があった。
In the above-mentioned conventional technique, two character patterns must be prepared for one character, and even if the characters have the same character code but different typefaces, It was necessary to prepare a separate character pattern. For example, one typeface in Japanese contains about 6000 characters, so if a character pattern is prepared for five typeface characters, an enormous number of character patterns of 6000 × 5 × 2 = 60,000 are prepared. There was a need.

【0004】なお、特開昭52−42025号公報の装
置では、標準パターンから細らせ処理と太らせ処理を行
なうことによって2つの文字パターンを作成しているの
で、予め記憶しておくべき標準パターンの数は減少する
が、太らせ処理や細らせ処理が必要になるという問題が
あった。また、結局1つの文字に対して2つの文字パタ
ーンを作成して文字画像と比較するという点では特開昭
60−230279号公報と同じであった。
In the apparatus disclosed in Japanese Patent Laid-Open No. 52-42025, two character patterns are created by performing a thinning process and a thickening process from the standard pattern, so that the standard pattern should be stored in advance. Although the number of patterns is reduced, there is a problem that thickening processing and thinning processing are required. Further, it is the same as in Japanese Patent Laid-Open No. 60-230279 in that two character patterns are created for one character and compared with a character image.

【0005】この発明は、従来技術における上述の課題
を解決するためになされたものであり、文字画像を認識
する際に用いる標準文字パターンの数を低減することの
できる文字認識装置を提供することを目的とする。
The present invention has been made to solve the above-mentioned problems in the prior art, and provides a character recognition device capable of reducing the number of standard character patterns used when recognizing a character image. With the goal.

【0006】[0006]

【課題を解決するための手段】上述の課題を解決するた
め、この発明による装置は、画像入力装置で読取られた
文字画像から文字を認識する装置であって、(A)複数
の文字のそれぞれに対する標準文字パターンを記憶する
記憶手段と、(B)画像入力装置で読取られた文字画像
と、前記記憶装置に記憶された標準文字パターンとを比
較し、前記文字画像の線部分に前記標準文字パターンが
包含される文字の中で線部分の画素数が最も多い文字を
認識された文字と決定する識別手段と、を備えることを
特徴とする。
In order to solve the above-mentioned problems, an apparatus according to the present invention is an apparatus for recognizing a character from a character image read by an image input apparatus, wherein (A) each of a plurality of characters And (B) comparing the character image read by the image input device with the standard character pattern stored in the storage device, and storing the standard character pattern in the line portion of the character image. Among the characters included in the pattern, a character having the largest number of pixels in the line portion is identified as a recognized character.

【0007】[0007]

【作用】文字画像に包含される標準文字パターンの中
で、線部分の画素数が最も多い文字を認識された文字と
決定することによって、文字を正しく認識する確率を高
めることができる。
In the standard character pattern included in the character image, the probability that the character is correctly recognized can be increased by determining the character having the largest number of pixels in the line portion as the recognized character.

【0008】[0008]

【実施例】【Example】

A.装置の構成:図1は、この発明の一実施例としての
文字認識システムを示すブロック図である。この文字認
識システムは、文字識別装置10と画像入力装置20と
で構成されている。文字識別装置10はワークステーシ
ョンやパーソナルコンピュータであり、画像入力装置2
0はスキャナである。
A. Device Configuration: FIG. 1 is a block diagram showing a character recognition system as an embodiment of the present invention. The character recognition system includes a character identification device 10 and an image input device 20. The character identification device 10 is a workstation or a personal computer, and the image input device 2
0 is a scanner.

【0009】文字識別装置10は、CPU11と、RO
M12と、RAM13と、キーボード14と、表示制御
回路15と、CRT16と、プリンタ17とを備えてい
る。画像入力装置20で読取られた文字画像は、RAM
13に記憶される。RAM13には、標準文字パターン
も記憶されている。ただし、標準文字パターンは、磁気
ディスクや光磁気ディスクなどの他の記憶手段に記憶し
ておくようにしてもよい。CPU11はROM12に記
憶されたアプリケーションプログラムを実行することに
より、文字画像に含まれている文字を認識する文字認識
部としての機能を発揮する。
The character identification device 10 includes a CPU 11 and an RO.
An M12, a RAM 13, a keyboard 14, a display control circuit 15, a CRT 16 and a printer 17 are provided. The character image read by the image input device 20 is stored in the RAM.
13 is stored. A standard character pattern is also stored in the RAM 13. However, the standard character pattern may be stored in another storage means such as a magnetic disk or a magneto-optical disk. By executing the application program stored in the ROM 12, the CPU 11 exerts a function as a character recognition unit that recognizes a character included in a character image.

【0010】B.標準文字パターンの準備:図2は、標
準文字パターンを作成する方法を示す説明図である。図
2(A)、(B)は、同じ字種で書体が異なる2つの文
字の2値画像(以下、「文字パターン」と呼ぶ)P1,
P2を示している。ここで、「字種」とは同じ文字コー
ド(JISコードや区点コードなど)で表わされる文字
を言う。図2(A),(B)の2つの文字パターンP
1,P2は、文字コードと書体とで規定されている。ま
た、文字パターン内の線の部分には2値データの値
「1」が割り当てられ、背景には値「0」が割り当てら
れている。
B. Preparation of Standard Character Pattern: FIG. 2 is an explanatory diagram showing a method of creating a standard character pattern. 2A and 2B are binary images (hereinafter, referred to as "character pattern") P1 of two characters having the same character type but different typefaces.
P2 is shown. Here, the "character type" refers to a character represented by the same character code (JIS code, division mark code, etc.). Two character patterns P in FIGS. 2A and 2B
1 and P2 are defined by a character code and a typeface. In addition, the value "1" of the binary data is assigned to the line portion in the character pattern, and the value "0" is assigned to the background.

【0011】標準文字パターンを作成する際には、ま
ず、同じ字種で書体が異なる複数の文字のパターンP
1,P2の論理積をとり、図2(C)に示すようなパタ
ーンP3を得る。論理積を取る文字パターンは2つに限
らず3つ以上であってもよい。
When creating a standard character pattern, first, a pattern P of a plurality of characters having the same character type but different typefaces is used.
The logical product of 1 and P2 is taken to obtain a pattern P3 as shown in FIG. The number of character patterns for which the logical product is obtained is not limited to two and may be three or more.

【0012】次に、パターンP3を細線化することによ
り、図2(D)に示す1画素の幅の標準文字パターンP
Rを作成する。この標準文字パターンPRは、その元と
なった2つの文字パターンP1,P2に共通する標準文
字パターンである。
Next, by thinning the pattern P3, the standard character pattern P having a width of one pixel shown in FIG.
Create R. This standard character pattern PR is a standard character pattern common to the two original character patterns P1 and P2.

【0013】なお、文字によっては複数の文字パターン
の論理積を取らずに、その文字パターンそのものを芯線
化することによって標準文字パターンを作成してもよ
い。
Depending on the character, the standard character pattern may be created by skeletonizing the character pattern itself without taking the logical product of a plurality of character patterns.

【0014】ところで、書体には、明朝系統(細明朝
体、中明朝体、太明朝体)やゴシック系統(細ゴシッ
ク、中ゴシック、太ゴシック)というように、形状の骨
格がほぼ同じで線の太さが異なる書体の系統がいくつか
存在する。このように形状が近似している複数の書体に
対して図2の方法で標準文字パターンを作成すれば、複
数の書体に対して1組の標準文字パターンが作成され
る。この結果、標準文字パターンの量を大幅に減少させ
ることができる。
By the way, in the typeface, there is almost a skeleton of a shape, such as the Mincho system (Hosyo-Mincho, Chu-Mincho, Tai-Mincho) or the Gothic system (Thin Gothic, Middle Gothic, Thick Gothic). There are several typeface systems with the same line weight but different line thickness. If standard character patterns are created by the method of FIG. 2 for a plurality of fonts whose shapes are similar to each other, one set of standard character patterns is created for a plurality of fonts. As a result, the amount of standard character patterns can be significantly reduced.

【0015】図3は、こうして作成された標準文字パタ
ーンを含む標準文字パターンファイルの構成を示す概念
図である。標準文字パターンファイルには、文字コード
と、標準文字パターン(2値画像)と、パターン内の黒
画素数とが含まれている。図3に示すように、同じ
「カ」という文字であっても複数の標準文字パターンが
必要になる場合には、同じ文字に対して複数の標準文字
パターンが登録される。また、標準文字パターンファイ
ルにおいては、黒画素数(2値データの値が1の画素
数)が多い順に標準文字パターンが登録されている。こ
の理由については後述する。なお、この標準文字パター
ンファイルはRAM13に記憶される。
FIG. 3 is a conceptual diagram showing the structure of a standard character pattern file containing the standard character pattern thus created. The standard character pattern file includes a character code, a standard character pattern (binary image), and the number of black pixels in the pattern. As shown in FIG. 3, when a plurality of standard character patterns are required even for the same character "Ka", a plurality of standard character patterns are registered for the same character. Further, in the standard character pattern file, standard character patterns are registered in the order of increasing number of black pixels (the number of pixels having a binary data value of 1). The reason for this will be described later. The standard character pattern file is stored in the RAM 13.

【0016】C.文字認識の手順:図4は、図1の文字
認識システムを用いて文字を認識する手順を示すフロー
チャートである。ステップS1では、画像入力装置20
を用いて、文書に印刷された文字画像を読取る。
C. Character Recognition Procedure: FIG. 4 is a flowchart showing a procedure for recognizing a character using the character recognition system of FIG. In step S1, the image input device 20
To read the character image printed on the document.

【0017】ステップS2では、文字画像を1文字ずつ
切り出すとともに、各文字の文字画像をドットパターン
化(2値画像化)する。この際、文字画像に含まれてい
るゴミや傷などの雑音成分を除去する処理なども実行さ
れる。図5(A)は、こうして得られた入力文字パター
ンPIを示している。
In step S2, the character images are cut out one by one, and the character image of each character is formed into a dot pattern (binary image formation). At this time, a process of removing noise components such as dust and scratches included in the character image is also executed. FIG. 5 (A) shows the input character pattern PI thus obtained.

【0018】ステップS3では、文字認識の対象とする
入力文字パターンPIを1つ選び、その黒画素数Nbを
算出する。ステップS4では、黒画素数がNb以下の標
準文字パターンPR(図5(B))を1つ選択する。図
3に示したように、標準文字パターンファイルには、黒
画素数の多い順に標準文字パターンが並べられており、
ステップS4では黒画素数の多い順に(すなわち上から
順に)標準文字パターンPRが選択される。例えば、入
力文字パターンの黒画素数Nbが31の場合には、まず
ゴシック体の「カ」が選択され、次にステップS4が実
行されるときには明朝体の「カ」が選択される。
In step S3, one input character pattern PI which is the object of character recognition is selected, and the number Nb of black pixels thereof is calculated. In step S4, one standard character pattern PR (FIG. 5B) in which the number of black pixels is Nb or less is selected. As shown in FIG. 3, in the standard character pattern file, standard character patterns are arranged in descending order of the number of black pixels.
In step S4, the standard character patterns PR are selected in descending order of the number of black pixels (that is, from the top). For example, when the number Nb of black pixels of the input character pattern is 31, the Gothic type “ka” is first selected, and when step S4 is executed next, the Mincho type “ka” is selected.

【0019】ステップS5では、選択された標準文字パ
ターンPRの線部分(黒画素の部分)が入力文字パター
ンPIの線部分に包含されるか否かが判断される。具体
的には、標準文字パターンPRと入力文字パターンPI
との論理積が標準文字パターンPRに等しい場合に標準
文字パターンPRが入力文字パターンPIに包含される
と判断される。あるいは、標準文字パターンPRと入力
文字パターンPIとの論理和が入力文字パターンPIに
等しい場合に標準文字パターンPRが入力文字パターン
PIに包含されると判断してもよい。
In step S5, it is determined whether the line portion (black pixel portion) of the selected standard character pattern PR is included in the line portion of the input character pattern PI. Specifically, the standard character pattern PR and the input character pattern PI
When the logical product of and is equal to the standard character pattern PR, it is determined that the standard character pattern PR is included in the input character pattern PI. Alternatively, it may be determined that the standard character pattern PR is included in the input character pattern PI when the logical sum of the standard character pattern PR and the input character pattern PI is equal to the input character pattern PI.

【0020】標準文字パターンPRが入力文字パターン
PIに包含されると判断された場合には、ステップS6
においてその標準文字パターンPRの文字コードが入力
文字パターンの文字であると決定される。
If it is determined that the standard character pattern PR is included in the input character pattern PI, step S6.
In, the character code of the standard character pattern PR is determined to be the character of the input character pattern.

【0021】一方、ステップS5において、標準文字パ
ターンPRが入力文字パターンPIに包含されない場合
には、ステップS7に移行し、他の標準文字パターンが
標準文字パターンファイルに残っている場合には、ステ
ップS4に戻って次の標準文字パターンが選択されてス
テップS5が再度実行される。ステップS7において黒
画素数Nb以下の全ての標準文字パターンとの比較が終
了した場合には文字の識別が不能なので、ステップS8
において、その入力文字パターンに対して所定の識別不
能コードを設定する。
On the other hand, in step S5, when the standard character pattern PR is not included in the input character pattern PI, the process proceeds to step S7, and when another standard character pattern remains in the standard character pattern file, the step is executed. Returning to S4, the next standard character pattern is selected and step S5 is executed again. When the comparison with all the standard character patterns having the number of black pixels Nb or less is completed in step S7, the character cannot be identified, so that the step S8 is performed.
At, a predetermined unidentifiable code is set for the input character pattern.

【0022】ステップS6またはステップS8によって
1つの入力文字パターンPIの文字コードが決定される
とステップS9に移行し、次の入力文字パターンが有る
場合にはステップS3に戻ってステップS4〜S8の処
理を繰り返す。こうして、画像入力装置20で読取られ
たすべての入力文字パターンについて文字コードを決定
する。
When the character code of one input character pattern PI is determined in step S6 or step S8, the process proceeds to step S9. If there is the next input character pattern, the process returns to step S3 and the processes of steps S4 to S8 are performed. repeat. In this way, character codes are determined for all input character patterns read by the image input device 20.

【0023】以上のように、入力文字パターンPIに包
含される標準文字パターンPRが認識文字であると決定
されるので、認識される標準文字パターンの黒画素数は
入力文字パターンPIの黒画素数Nb以下である。この
ため、ステップS4において、入力文字パターンPIの
黒画素数Nb以下の黒画素数の標準文字パターンを選択
しているのである。
As described above, since the standard character pattern PR included in the input character pattern PI is determined to be the recognized character, the number of black pixels of the recognized standard character pattern is the number of black pixels of the input character pattern PI. It is Nb or less. Therefore, in step S4, the standard character pattern having the number of black pixels equal to or less than the number Nb of black pixels of the input character pattern PI is selected.

【0024】さらに、図3の標準文字パターンファイル
は黒画素数の多い順に並べられているので、認識される
文字は入力文字パターンPIに包含される標準文字パタ
ーンの中で最も黒画素数の多いものである。例えば、入
力文字パターン「五」を認識しようとした場合に、2つ
の標準文字パターン「五」,「三」のいずれも入力文字
パターン「五」に包含されるが、このうちで黒画素数の
多い文字「五」が認識文字となる。同様に、入力文字パ
ターン「が」を認識しようとした場合に、2つの標準文
字パターン「が」,「か」のいずれも入力文字パターン
「が」に包含されるが、このうちで黒画素数の多い文字
「が」が認識文字となる。一方、入力文字パターン
「か」を認識しようとした場合には、標準文字パターン
「が」は入力文字パターン「か」に包含されず、標準文
字パターン「か」は包含されるので、文字「か」が認識
文字となる。このように、入力文字パターンPIに包含
される標準文字パターンの中で最も黒画素数の多いもの
を認識文字としているので、文字を正しく認識すること
が可能である。
Furthermore, since the standard character pattern files of FIG. 3 are arranged in the order of the number of black pixels, the recognized character has the largest number of black pixels among the standard character patterns included in the input character pattern PI. It is a thing. For example, when trying to recognize the input character pattern "5", both of the two standard character patterns "5" and "3" are included in the input character pattern "5". Many characters "5" are recognized characters. Similarly, when trying to recognize the input character pattern "ga", both of the two standard character patterns "ga" and "ka" are included in the input character pattern "ga". The character "ga" with many characters is the recognized character. On the other hand, when trying to recognize the input character pattern "ka", the standard character pattern "ga" is not included in the input character pattern "ka" and the standard character pattern "ka" is included. Is the recognition character. As described above, since the character having the largest number of black pixels among the standard character patterns included in the input character pattern PI is the recognized character, it is possible to correctly recognize the character.

【0025】また、標準文字パターンを1文字の幅の線
で構成するようにしているので、入力文字パターンの線
が多少ずれている場合にも、その入力文字パターンの線
部分に標準文字パターンの線部分が含まれることにな
る。従って、正しい文字認識を行なえる確率が高くなっ
ている。
Further, since the standard character pattern is composed of lines having a width of one character, even if the lines of the input character pattern are slightly deviated, the line part of the input character pattern is The line part will be included. Therefore, the probability of correct character recognition is high.

【0026】D.変形例:なお、この発明は上記実施例
に限られるものではなく、その要旨を逸脱しない範囲に
おいて種々の態様において実施することが可能であり、
例えば次のような変形も可能である。
D. Modification: The present invention is not limited to the above-described embodiment, and can be implemented in various modes without departing from the scope of the invention.
For example, the following modifications are possible.

【0027】(1)上記実施例では、入力文字パターン
に包含される標準文字パターンが存在しない場合におい
て、認識不能コードを設定する代わりに、入力文字パタ
ーンの包含度が最も高い標準文字パターンを認識文字と
して決定するようにしてもよい。入力文字パターンの包
含度としては種々のものが考えられる。例えば、入力文
字パターンと標準文字パターンとを画素の行ごと(また
は列ごと)に比較し、入力文字パターンに包含される行
数(または列数)を包含度としてもよい。この際、入力
文字パターンに包含されない行数が所定の数以上になっ
た文字は、認識文字候補になり得ないものとして、行ご
との比較を途中で終了すれば、処理時間を短縮すること
が可能である。また、入力文字パターンに包含される黒
画素数そのものを包含度として使用してもよい。
(1) In the above embodiment, when the standard character pattern included in the input character pattern does not exist, the standard character pattern having the highest degree of inclusion of the input character pattern is recognized instead of setting the unrecognizable code. You may make it determine as a character. There are various possible inclusion degrees of the input character pattern. For example, the input character pattern and the standard character pattern may be compared for each row (or each column) of pixels, and the number of rows (or the number of columns) included in the input character pattern may be used as the inclusion degree. At this time, if the number of lines that is not included in the input character pattern exceeds a predetermined number, it is considered that the characters cannot be recognized character candidates, and the processing time can be shortened if the line-by-line comparison is terminated halfway. It is possible. Alternatively, the number of black pixels included in the input character pattern itself may be used as the inclusion degree.

【0028】なお、入力文字パターンに完全に包含され
る標準文字パターンが存在しない場合に、包含度の高い
順にCRT等に表示して、オペレータに正しい文字を選
択させるようにしてもよい。
When there is no standard character pattern completely included in the input character pattern, the characters may be displayed in the CRT or the like in descending order of inclusion degree so that the operator can select the correct character.

【0029】(2)書体の中には、明朝体のように横線
が縦線に比べてかなり細いものがある。このような書体
の文字を認識する場合には、入力文字パターンの横線の
みを太らせた後に、標準文字パターンと比較するように
してもよい。横線を太らせるか否かは、オペレータが指
定することができる。
(2) In some typefaces, horizontal lines are much thinner than vertical lines, such as Mincho typeface. When recognizing characters in such a typeface, only the horizontal lines of the input character pattern may be thickened and then compared with the standard character pattern. The operator can specify whether to thicken the horizontal line.

【0030】例えば、図6(C)の入力文字パターンP
Iaを図6(A),(B)の標準文字パターンPR1,
PR2と比較する場合を考える。この場合、図6(A)
の標準文字パターンPR1は入力文字パターンPIaに
包含されるが、図6(B)の標準文字パターンPR2は
包含されないので、図6(A)の標準文字パターンPR
1の文字「±」が認識文字と決定される。しかし、図6
(C)の入力文字パターンPIaに対しては、むしろ図
6(B)の標準文字パターンPR2の文字「土」が認識
文字となるのが好ましい。図6(C)のように、横線が
細い入力文字パターンPIaについては、その入力文字
パターンPIaの横線部分を、図6(C)の白抜きの四
角形で示す部分まで太らせる。こうすれば、2つの標準
文字パターンPR1,PR2の両方が入力文字パターン
PIaに包含されるので、黒画素数がより多い標準文字
パターンPR2が認識文字と決定され、正しい認識が行
なわれる。
For example, the input character pattern P of FIG.
Ia is the standard character pattern PR1 shown in FIGS.
Consider the case of comparison with PR2. In this case, FIG. 6 (A)
6A is included in the input character pattern PIa, but is not included in the standard character pattern PR2 of FIG. 6B, the standard character pattern PR of FIG.
The character "1" of 1 is determined as the recognized character. However, FIG.
For the input character pattern PIa of (C), it is preferable that the character “Sat” of the standard character pattern PR2 of FIG. 6B be the recognized character. As shown in FIG. 6C, for an input character pattern PIa with a thin horizontal line, the horizontal line portion of the input character pattern PIa is thickened to the portion shown by the white square in FIG. 6C. In this way, since both of the two standard character patterns PR1 and PR2 are included in the input character pattern PIa, the standard character pattern PR2 having a larger number of black pixels is determined as a recognized character, and correct recognition is performed.

【0031】別の例として、図7(C)の入力文字パタ
ーンPIbを図7(A),(B)の標準文字パターンP
R1,PR2と比較する場合を考える。この時にも図6
の場合と同様に、図7(C)の入力文字パターンPIb
の横線部分を、図7(C)の白抜きの四角形で示す部分
まで太らせる。この場合、図7(A)の標準文字パター
ンPR1は太らせた入力文字パターンPIbに包含され
るが、図7(B)の標準文字パターンPR2は包含され
ないので、図7(A)の標準文字パターンPR1の文字
「±」が認識文字と決定される。こうして、図7(A)
の標準文字パターンPR1の文字であることが正しく認
識される。
As another example, the input character pattern PIb of FIG. 7C is replaced by the standard character pattern P of FIGS. 7A and 7B.
Consider the case of comparison with R1 and PR2. Also at this time
7C, the input character pattern PIb of FIG.
7C is thickened up to the portion indicated by the white square in FIG. 7C. In this case, the standard character pattern PR1 of FIG. 7 (A) is included in the thickened input character pattern PIb, but the standard character pattern PR2 of FIG. 7 (B) is not included, so the standard character pattern of FIG. 7 (A) is included. The character “±” of the pattern PR1 is determined as the recognized character. Thus, FIG. 7 (A)
It is correctly recognized that the character has the standard character pattern PR1.

【0032】以上のように、横線が細い書体の文字を認
識する場合には、入力文字パターンの横線部分のみを太
らせることによって正しく認識できる確率を高めること
ができる。なお、入力文字パターンの横線部分を太らせ
る方法としては、画像入力装置20において光学的に行
なう第1の方法と、文字識別装置10において画像デー
タを修正する第2の方法とがある。
As described above, in the case of recognizing characters in a typeface with thin horizontal lines, it is possible to increase the probability of correct recognition by thickening only the horizontal line portion of the input character pattern. As a method of thickening the horizontal line portion of the input character pattern, there are a first method optically performed by the image input device 20 and a second method of correcting image data by the character identification device 10.

【0033】第1の太らせ方法では、線を太らせる方向
(図7の場合は上下方向)に光源を走査し、光源の走査
方向と直角な方向に設けられたスリットの幅を広くす
る。これによって、光源の走査方向にボケが生じ、文字
パターンの横線の上下が太ることになる。太らせる幅
は、スリットの幅によって調整することができる。
In the first thickening method, the light source is scanned in the direction in which the line is thickened (vertical direction in FIG. 7), and the width of the slit provided in the direction perpendicular to the scanning direction of the light source is widened. As a result, blurring occurs in the scanning direction of the light source, and the upper and lower horizontal lines of the character pattern are thickened. The width to be thickened can be adjusted by the width of the slit.

【0034】第2の太らせ方法では、線を太らせる方向
と直角な方法(すなわち水平方向)に並ぶ黒画素の数を
調べ、所定の数(例えば7個)以上黒画素が並んでいる
ものを横線と認識して、その上下に黒画素を追加するよ
うにすればよい。なお、この際、図8に示すように、横
線同士の間隔が1画素しか無い場合には、横線同士に挟
まれた画素(図8で×印が記入されている画素)は白画
素のまま残すようにする。こうすれば、横線同士の間の
空間がつぶれることを防止できる。
In the second thickening method, the number of black pixels arranged in a direction perpendicular to the direction of thickening the line (that is, in the horizontal direction) is checked, and a predetermined number (eg, 7) or more of black pixels are arranged. Should be recognized as a horizontal line, and black pixels should be added above and below it. At this time, as shown in FIG. 8, when there is only one pixel between the horizontal lines, the pixels sandwiched between the horizontal lines (pixels marked with an X in FIG. 8) remain white pixels. Try to leave. By doing so, it is possible to prevent the space between the horizontal lines from being collapsed.

【0035】[0035]

【発明の効果】以上説明したように、本発明の文字認識
装置によれば、文字画像を認識する際に用いる標準文字
パターンの数を低減することができるという効果があ
る。
As described above, according to the character recognition device of the present invention, it is possible to reduce the number of standard character patterns used when recognizing a character image.

【図面の簡単な説明】[Brief description of drawings]

【図1】この発明の一実施例を適用して文字を認識する
文字認識システムを示すブロック図。
FIG. 1 is a block diagram showing a character recognition system for recognizing characters by applying an embodiment of the present invention.

【図2】標準文字パターンを作成する方法を示す説明
図。
FIG. 2 is an explanatory diagram showing a method of creating a standard character pattern.

【図3】標準文字パターンファイルの構成を示す概念
図。
FIG. 3 is a conceptual diagram showing the configuration of a standard character pattern file.

【図4】文字認識の手順を示すフローチャート。FIG. 4 is a flowchart showing a procedure of character recognition.

【図5】入力文字パターンPIと標準文字パターンとを
比較して示す図。
FIG. 5 is a diagram showing a comparison between an input character pattern PI and a standard character pattern.

【図6】横線を太らせる場合における入力文字パターン
と標準文字パターンの第1の例を示す図。
FIG. 6 is a diagram showing a first example of an input character pattern and a standard character pattern when a horizontal line is thickened.

【図7】横線を太らせる場合における入力文字パターン
と標準文字パターンの第2の例を示す図。
FIG. 7 is a diagram showing a second example of an input character pattern and a standard character pattern when a horizontal line is thickened.

【図8】横線を太らせる一例を示す図。FIG. 8 is a diagram showing an example of thickening a horizontal line.

【符号の説明】[Explanation of symbols]

10…文字識別装置 11…CPU 12…ROM 13…RAM 14…キーボード 15…表示制御回路 16…CRT 17…プリンタ 20…画像入力装置 Nb…入力文字パターンの黒画素数 PI…入力文字パターン PR…標準文字パターン 10 ... Character identification device 11 ... CPU 12 ... ROM 13 ... RAM 14 ... Keyboard 15 ... Display control circuit 16 ... CRT 17 ... Printer 20 ... Image input device Nb ... Number of black pixels of input character pattern PI ... Input character pattern PR ... Standard Character pattern

Claims (1)

【特許請求の範囲】[Claims] 【請求項1】 画像入力装置で読取られた文字画像から
文字を認識する装置であって、(A)複数の文字のそれ
ぞれに対する標準文字パターンを記憶する記憶手段と、
(B)画像入力装置で読取られた文字画像と、前記記憶
装置に記憶された標準文字パターンとを比較し、前記文
字画像の線部分に前記標準文字パターンが包含される文
字の中で線部分の画素数が最も多い文字を認識された文
字と決定する識別手段と、を備えることを特徴とする文
字認識装置。
1. A device for recognizing a character from a character image read by an image input device, comprising: (A) storage means for storing a standard character pattern for each of a plurality of characters;
(B) A character image read by the image input device is compared with a standard character pattern stored in the storage device, and a line portion of the character image includes the line portion of the standard character pattern. A character recognizing device that determines the character having the largest number of pixels as the recognized character.
JP5077693A 1993-03-10 1993-03-10 Character recognition device Pending JPH06259601A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP5077693A JPH06259601A (en) 1993-03-10 1993-03-10 Character recognition device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP5077693A JPH06259601A (en) 1993-03-10 1993-03-10 Character recognition device

Publications (1)

Publication Number Publication Date
JPH06259601A true JPH06259601A (en) 1994-09-16

Family

ID=13640979

Family Applications (1)

Application Number Title Priority Date Filing Date
JP5077693A Pending JPH06259601A (en) 1993-03-10 1993-03-10 Character recognition device

Country Status (1)

Country Link
JP (1) JPH06259601A (en)

Similar Documents

Publication Publication Date Title
US5075895A (en) Method and apparatus for recognizing table area formed in binary image of document
EP0343786A2 (en) Method and apparatus for reading and recording text in digital form
JP3737177B2 (en) Document generation method and apparatus
US5724455A (en) Automated template design method for print enhancement
US6195473B1 (en) Non-integer scaling of raster images with image quality enhancement
JP5294798B2 (en) Image processing apparatus and image processing method
EP1093078B1 (en) Reducing apprearance differences between coded and noncoded units of text
JP3172498B2 (en) Image recognition feature value extraction method and apparatus, storage medium for storing image analysis program
JP2957729B2 (en) Line direction determination device
JP2003317107A (en) Ruled line extraction method and apparatus
JP2977230B2 (en) Character extraction method
JP3162414B2 (en) Ruled line recognition method and table processing method
KR20250154355A (en) Apparatus and method for generating document images used in machine-learning of text detection and recognition
JP3961730B2 (en) Form processing apparatus, form identification method, and recording medium
JP3486246B2 (en) Character recognition device
JP2801730B2 (en) Data compression method
JP2771045B2 (en) Document image segmentation method
JPH0535872A (en) Contour tracing system for binary image
JP2918363B2 (en) Character classification method and character recognition device
JP2857260B2 (en) Judgment method of rectangular area
JPH09121279A (en) Picture processor
JPH01296385A (en) Method for improving picture quality of binary picture data
JPH04330584A (en) Line direction decision device
JPH08194772A (en) Optical character reader
JPH04360294A (en) Device and method for recognizing table

Legal Events

Date Code Title Description
LAPS Cancellation because of no payment of annual fees