JPH04268685A - Method for discriminating type of slips - Google Patents

Method for discriminating type of slips

Info

Publication number
JPH04268685A
JPH04268685A JP3050497A JP5049791A JPH04268685A JP H04268685 A JPH04268685 A JP H04268685A JP 3050497 A JP3050497 A JP 3050497A JP 5049791 A JP5049791 A JP 5049791A JP H04268685 A JPH04268685 A JP H04268685A
Authority
JP
Japan
Prior art keywords
pattern
ruled line
area
extracted
recognized
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP3050497A
Other languages
Japanese (ja)
Other versions
JP3096481B2 (en
Inventor
Yasuo Fujita
藤田 泰生
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Glory Ltd
Original Assignee
Glory Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Glory Ltd filed Critical Glory Ltd
Priority to JP03050497A priority Critical patent/JP3096481B2/en
Publication of JPH04268685A publication Critical patent/JPH04268685A/en
Application granted granted Critical
Publication of JP3096481B2 publication Critical patent/JP3096481B2/en
Anticipated expiration legal-status Critical
Expired - Fee Related legal-status Critical Current

Links

Landscapes

  • Character Input (AREA)

Abstract

PURPOSE:To provide the method stably discriminating the type of slips based on ruled line information forming the table of slips. CONSTITUTION:The horizontal and vertical line segments of a table constituting a character reading frame are extracted from inputted slip picture data. Vector patterning is performed by using the direction, length, and position of the line segments extracted for each area so as to comparatively collated with the feature vector of the reference pattern.

Description

【発明の詳細な説明】[Detailed description of the invention]

【0001】0001

【産業上の利用分野】本発明は、表形式の複数の帳票類
を読取る場合において、帳票の種類判別を行なう帳票類
の種類判別方法に関する。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a method for determining the type of a document when reading a plurality of tabular documents.

【0002】0002

【従来の技術】従来、表を含む帳票類を対象とする光学
式文字読取装置において文字読取りを行なう場合、上記
帳票類には文字記入位置を示す読取り枠が縦、横の罫線
等で印刷されている。光学式文字読取装置は、帳票を光
学的に走査して得られる帳票画像から読取り枠位置にお
いて文字パターンの検出、切出しを行ない、切出した文
字パターンについて文字認識処理を行ない、帳票に記入
された文字を読取るように構成されている。このとき、
読取り枠の位置情報はそれぞれの帳票について予め光学
式読取装置内に設定される。そして、読取り枠位置が異
なる複数種類の帳票が読取りの対象となる場合、帳票を
同定するためのバーコード又は番号(帳票識別マーク)
領域を帳票に設け、この領域内のバーコード又は番号を
読取って帳票の種類を判別し、帳票の種類に応じた読取
り枠位置情報を用いて文字読取処理を行なっている。
[Prior Art] Conventionally, when reading characters with an optical character reading device for documents including tables, a reading frame indicating the position of writing characters is printed on the documents with vertical and horizontal ruled lines. ing. An optical character reader detects and cuts out a character pattern at the reading frame position from a form image obtained by optically scanning a form, performs character recognition processing on the cut out character pattern, and recognizes the characters written on the form. is configured to read. At this time,
The position information of the reading frame is set in advance in the optical reading device for each form. If multiple types of documents with different reading frame positions are to be scanned, a barcode or number (document identification mark) to identify the documents is used.
An area is provided on a form, a bar code or number within this area is read to determine the type of form, and character reading processing is performed using reading frame position information corresponding to the type of form.

【0003】0003

【発明が解決しようとする課題】しかしながら従来の方
法では、文字読取り枠の構造が異なる複数種類の帳票を
混在させて文字読取処理を行なう場合には、各帳票に帳
票の種類を判別するためのバーコード又は番号領域を帳
票に設ける必要があった。そして、複数種類の帳票を混
在させて読取処理を行なう場合は、この領域が設けてあ
る帳票のみが対象であり、この領域の位置が帳票毎に異
なっている場合や、バーコード又は番号領域を設けてい
ない帳票を読取り対象にする場合には種類判別の処理が
困難であった。
[Problem to be Solved by the Invention] However, in the conventional method, when performing character reading processing on a mixture of multiple types of forms with different character reading frame structures, it is necessary to provide a method for determining the type of form for each form. It was necessary to provide a barcode or number area on the form. When reading a mixture of multiple types of forms, only the forms that have this area are targeted; if the position of this area is different for each form, or if the barcode or number area is When a document that is not provided is to be read, it is difficult to determine the type.

【0004】本発明はこのような点に着目して成された
ものであり、このような領域の異なる帳票が混在してい
る場合やこのような領域の設けられていない帳票であっ
ても、安定して帳票の種類判別を行ない得るように帳票
類の表を形成する罫線情報を基に帳票類の種類を判別す
る方法を提供することを目的としている。
[0004] The present invention has been made with attention to these points, and even when forms with different areas are mixed together, or even when forms do not have such areas, It is an object of the present invention to provide a method for discriminating types of forms based on ruled line information forming a table of forms so as to stably discriminate the types of forms.

【0005】[0005]

【課題を解決するための手段】本発明は帳票類の種類判
別方法に関するもので、本発明の上記目的は、被認識帳
票類の画像情報を入力し、この入力された画像データか
ら水平、垂直方向の線分を抽出し、抽出された各線分の
位置、長さ、方向の情報を検出し、前記入力された画像
を複数エリアに区分し、この区分されたエリア毎に存在
する前記抽出された線分の方向毎の長さ毎の重みである
特徴量を被認識対象のパターン特徴量として求め、認識
すべきパターンの基準となる標準パターンの特徴量を予
め記憶しておき、前記抽出された被認識パターンの特徴
量と前記記憶されている標準パターンの特徴量とを比較
照合し、その比較結果に基づいて前記被認識パターンが
どの標準パターンであるかを判断して帳票の種類を判別
することによって達成される。
[Means for Solving the Problems] The present invention relates to a method for determining the type of documents, and the above object of the present invention is to input image information of documents to be recognized, and to use the input image data horizontally and vertically. Line segments in the direction are extracted, information on the position, length, and direction of each extracted line segment is detected, the input image is divided into a plurality of areas, and the extracted line segments that exist in each divided area are The feature quantity, which is the weight for each length in each direction of the line segment, is determined as the pattern feature quantity of the object to be recognized. The feature amount of the recognized pattern and the feature amount of the stored standard pattern are compared and matched, and based on the comparison result, it is determined which standard pattern the recognized pattern is, and the type of the form is determined. This is achieved by

【0006】また、被認識帳票類の画像情報を入力し、
この入力された画像データから水平、垂直方向の線分を
抽出し、抽出された各線分の位置、長さ、方向の情報を
検出し、前記入力された画像を複数エリアに区分し、こ
の区分されたエリア毎に存在する前記抽出された線分の
方向毎の長さ毎の重みである特徴量を被認識対象のパタ
ーン特徴量として求め、認識すべきパターンの基準とな
る標準パターンの特徴量を予め記憶しておき、前記抽出
された被認識パターンの特徴量と前記記憶されている標
準パターンの特徴量とを比較照合し、所定値以上類似す
るものを類似する順に候補として抽出し、候補として抽
出された標準パターンの予め定められた特定領域に対応
する入力画像情報を検出して、前記標準パターンのもの
と一致するかを前記候補順に確認することによって帳票
の種類を判別することによって達成される。
[0006] Furthermore, inputting image information of documents to be recognized,
Line segments in the horizontal and vertical directions are extracted from this input image data, information on the position, length, and direction of each extracted line segment is detected, and the input image is divided into multiple areas. The feature quantity, which is the weight for each length of each direction of the extracted line segment existing in each area, is determined as the pattern feature quantity of the recognition target, and the feature quantity of the standard pattern, which is the standard of the pattern to be recognized, is determined. are stored in advance, the extracted feature values of the recognized pattern are compared and matched with the feature values of the stored standard pattern, and those that are similar by more than a predetermined value are extracted as candidates in the order of similarity, and the candidates are This is achieved by detecting input image information corresponding to a predetermined specific area of a standard pattern extracted as a standard pattern, and determining the type of form by checking whether it matches the standard pattern in the order of the candidates. be done.

【0007】[0007]

【作用】本発明は、バーコード又は番号(帳票識別マー
ク)領域の設けられていない帳票や、各々異なる位置に
設けられている帳票等の複数種類の帳票を混在させて文
字読取処理を行なう場合の、帳票の種類判別を行なうた
めの表形式パターン認識方法に関するもので、入力され
た帳票画像データから文字読取り枠を構成している表の
構成要素である水平、垂直の線分(罫線)を抽出して複
数エリアに区分し、区分されたエリア毎に抽出された線
分の方向、長さ、位置を用いてベクトルパターン化して
標準パターンの特徴ベクトルと比較照合するようにして
いる。また、請求項2の発明では、区分されたエリア毎
に抽出された線分の方向、長さ、位置を用いて、ベクト
ルパターン化したパターンベクトルと標準パターンの特
徴ベクトルとを比較照合して一旦候補を抽出し、候補順
に詳細確認を行なうようにしている。
[Operation] The present invention is applicable when performing character reading processing on a mixture of multiple types of forms, such as forms that do not have a bar code or number (form identification mark) area, or forms that are provided in different positions. This method relates to a tabular pattern recognition method for distinguishing the type of form, and uses input form image data to identify the horizontal and vertical line segments (ruled lines) that are the constituent elements of the table that make up the character reading frame. It is extracted and divided into a plurality of areas, and the direction, length, and position of the line segment extracted for each divided area are used to create a vector pattern, which is then compared and verified with the feature vector of a standard pattern. In addition, in the invention of claim 2, the direction, length, and position of the line segment extracted for each divided area are used to compare and match the pattern vector converted into a vector pattern and the feature vector of the standard pattern. Candidates are extracted and detailed confirmation is performed in the order of the candidates.

【0008】[0008]

【実施例】図1は本発明の認識の対象である帳票の一例
を示しており、図2は本発明の認識アルゴリズムのフロ
ーチャートであり、このフローチャートを参照して認識
動作を説明する。
DESCRIPTION OF THE PREFERRED EMBODIMENTS FIG. 1 shows an example of a form to be recognized by the present invention, and FIG. 2 is a flowchart of the recognition algorithm of the present invention, and the recognition operation will be explained with reference to this flowchart.

【0009】先ずイメージセンサ等の画像入力手段で帳
票の画像入力を行ない(ステップS1)、帳票のサイズ
を検出する(ステップS2)。帳票のサイズ検出は、入
力画像の濃淡レベルを用いて帳票のエッジを検出するこ
とによって行なう。次に画像データの前処理を行なうが
、この前処理は2値化及びスムージングを行なうことに
よって実行する(ステップS3)。2値画像のスムージ
ング処理は図3に示すように、注目画素2に対する8近
傍領域3での穴埋め処理又は孤立点除去処理を行なうこ
とによって実行する。すなわち、黒画素を“1”、白画
素を“0”としたとき、穴埋め処理は、注目画素2が“
0”でかつ8近傍画素3の全てが“1”のとき、注目画
素2の値を“0”→“1”に変更する。また、孤立点除
去処理は、注目画素2が“1”でかつ8近傍画素3の全
てが“0”のとき、注目画素2の値を“1”→“0”に
変更する。
First, an image of a form is input using image input means such as an image sensor (step S1), and the size of the form is detected (step S2). The size of the form is detected by detecting the edges of the form using the gray level of the input image. Next, the image data is preprocessed, and this preprocessing is performed by binarizing and smoothing (step S3). As shown in FIG. 3, the smoothing process of the binary image is executed by performing hole filling process or isolated point removal process in the 8-neighborhood area 3 for the pixel of interest 2. In other words, when a black pixel is set to "1" and a white pixel is set to "0", the hole-filling process is performed so that the pixel of interest 2 is set to "0".
0" and all of the 8 neighboring pixels 3 are "1", the value of the pixel of interest 2 is changed from "0" to "1". In addition, the isolated point removal process is performed when the pixel of interest 2 is "1". And when all of the eight neighboring pixels 3 are "0", the value of the pixel of interest 2 is changed from "1" to "0".

【0010】次に、前処理(2値化、スムージング)さ
れた画像に対して、黒画素が水平又は垂直方向に連続す
る数をカウントし、各画素にこの水平、垂直に黒画素が
連続する値(以下、ランレングスという)を求め、この
ランレングスを用いた罫線の抽出を行なうが(ステップ
S4)、この動作フローは図4に示すようになっており
、図4を参照してその動作を説明する。
Next, for the preprocessed (binarized, smoothed) image, count the number of consecutive black pixels in the horizontal or vertical direction, and calculate the number of consecutive black pixels in the horizontal or vertical direction for each pixel. The value (hereinafter referred to as run length) is determined, and the ruled line is extracted using this run length (step S4). This operation flow is shown in FIG. 4. Explain.

【0011】先ず横方向の罫線についての抽出について
説明すると、前処理された画像データから水平方向に連
続する直線を抽出して横ランレングスを検出する(ステ
ップS30)。そして、抽出された横ランレングスのう
ち極端に短いもの、例えば5mm未満のランレングスを
持つ画素はノイズとして無効データとし削除し、5mm
以上のランレングスを持つ画素は有効として取扱う横ラ
ンレングス画像の2値化を行なう(ステップS31)。 この画像データの様子を示すのが図5の(A)である。 次にこの画像データに基づいてつなぎ処理を行なう。つ
なぎ処理を具体的に説明すると、図5の(A)の画像デ
ータの横方向のヒストグラムである横方向周辺分布特徴
を図5の(B)のように求め(ステップS32)、周辺
分布特徴の高さがしきい値TH1を超える部分を罫線候
補位置とする(ステップS33)。図5の(A)、(B
)においては、罫線候補位置X、Y、Zの3つの候補が
抽出されている。そして、各々の罫線候補位置毎にその
候補領域内の縦方向のヒストグラムである縦方向周辺分
布を求める(ステップS34)。図5の(A)の罫線候
補位置Yについて縦方向周辺分布を求めた様子を示すの
が図5の(C)である。図5の(D)に示すように、こ
の縦方向周辺分布をしきい値TH2(例えばTH2=1
画素)で比較し、しきい値TH2を超える部分を有効と
する。この有効線分の間隔が所定値TH3(例えばTH
3=5画素)以下の場合はつなぎ処理を施し、つながっ
た一つの直線として取扱って画像入力データの罫線の補
正を行なう。図1の帳票について、この処理を行なった
様子を示すのが図6の(A)である。なお、ここで図5
の(D)に示すような各横罫線区間での縦周辺分布特徴
の値の平均値を、その罫線の幅としておく。線幅は文字
読取領域を画像より切出すときに、線で囲まれた領域の
内側のみを切出す時の情報として用いる。縦方向の罫線
についても同様な処理により縦罫線を求めることができ
る。図4のステップS20〜S25が縦罫線の抽出を示
している。図1の帳票について縦罫線を抽出した様子を
示すのが図6の(B)であり、罫線の抽出を終了したと
きの様子を示すのが図6の(C)である。
First, to explain the extraction of horizontal ruled lines, horizontally continuous straight lines are extracted from the preprocessed image data to detect the horizontal run length (step S30). Then, among the extracted horizontal run lengths, pixels with extremely short run lengths, for example, less than 5 mm, are treated as noise and invalid data and are deleted.
Pixels having the above run length are treated as valid and the horizontal run length image is binarized (step S31). FIG. 5A shows the state of this image data. Next, connection processing is performed based on this image data. To explain the connection process specifically, the horizontal peripheral distribution feature, which is a horizontal histogram of the image data in FIG. 5(A), is obtained as shown in FIG. 5(B) (step S32), and the peripheral distribution feature is A portion whose height exceeds the threshold value TH1 is set as a ruled line candidate position (step S33). (A) and (B) in Figure 5
), three candidates of ruled line candidate positions X, Y, and Z are extracted. Then, for each ruled line candidate position, a vertical peripheral distribution, which is a vertical histogram within the candidate area, is determined (step S34). FIG. 5C shows how the vertical peripheral distribution for the ruled line candidate position Y in FIG. 5A is obtained. As shown in (D) of FIG.
(pixel), and the portion exceeding the threshold value TH2 is considered valid. The interval between these effective line segments is a predetermined value TH3 (for example, TH
(3=5 pixels) or less, a connection process is performed and the ruled lines of the image input data are corrected by treating them as one connected straight line. FIG. 6A shows how this process is performed for the form shown in FIG. In addition, here, Figure 5
The average value of the vertical peripheral distribution feature values in each horizontal ruled line section as shown in (D) is set as the width of that ruled line. The line width is used as information when cutting out only the inside of the area surrounded by lines when cutting out a character reading area from an image. Vertical ruled lines can also be obtained by similar processing. Steps S20 to S25 in FIG. 4 show extraction of vertical ruled lines. FIG. 6(B) shows how vertical ruled lines are extracted for the form in FIG. 1, and FIG. 6(C) shows how the ruled line extraction is completed.

【0012】上述のようなランレングスを用いた罫線の
抽出後、画像より抽出された罫線の各々について縦横罫
線情報(位置(始終点座標)、長さ、方向(縦、横))
を求める(ステップS40)。以上の段階で画像入力デ
ータに修正が加えられ、罫線の抽出及び抽出された各罫
線についての情報が確実に検出されたことになる。以下
、この縦横罫線情報を基に処理を実行する。
After the ruled lines are extracted using the run length as described above, vertical and horizontal ruled line information (position (starting and ending point coordinates), length, direction (vertical, horizontal)) is obtained for each ruled line extracted from the image.
(Step S40). In the above steps, the image input data has been corrected, and the ruled lines have been extracted and information about each extracted ruled line has been reliably detected. Hereafter, processing is executed based on this vertical and horizontal ruled line information.

【0013】次に上記縦横罫線情報に基づいて帳票のパ
ターンベクトルを求める(図2のステップS5)。パタ
ーンベクトルについては後に詳述するが、ここでその概
略を述べておくと、パターンベクトルは図7に示されて
いるようなもので、P[m][n][t][w](但し
、mは縦方向のエリア番号(1〜M)、nは横方向のエ
リア番号(1〜N)、wは罫線の長さのレベル(1〜W
)、tは罫線の方向でt=1が横方向、t=2が縦方向
)で表わされ、M×N×W×2次元ベクトルである。 つまり、M×Nの各エリアについて、縦、横それぞれの
罫線が、罫線の長さのレベル(短いものから長いものま
でをW段階で量子化し表現している)毎に含まれている
度合い(以下、重みという)を表現している。図8がこ
のパターンベクトルPを求める動作を示すフローチャー
トであり、先ず初期化されたパターンベクトルに、各罫
線ki(但し、iは罫線番号)毎に、罫線の始点(Xs
i,Ysi)、終点(Xei,Yei)、長さLiに基
づいて、罫線kiが含まれているそれぞれのエリアの罫
線kiの長さのレベルにおいて、罫線kiによる重みを
求め、順次加算することによりパターンベクトルPを求
める。
Next, a pattern vector of the form is determined based on the vertical and horizontal ruled line information (step S5 in FIG. 2). The pattern vector will be explained in detail later, but to give an overview here, the pattern vector is as shown in Figure 7, and P[m][n][t][w] (however, , m is the vertical area number (1 to M), n is the horizontal area number (1 to N), and w is the ruled line length level (1 to W
), t is the direction of the ruled line, t=1 is the horizontal direction, t=2 is the vertical direction), and is an M×N×W×2-dimensional vector. In other words, for each M×N area, the degree to which vertical and horizontal ruled lines are included for each ruled line length level (expressed by quantizing in W steps from short to long) ( (hereinafter referred to as weight). FIG. 8 is a flowchart showing the operation to obtain this pattern vector P. First, for each ruled line ki (where i is the ruled line number), the initial point of the ruled line (Xs
i, Ysi), the end point (Xei, Yei), and the length Li, find the weight of the ruled line ki at the level of the length of the ruled line ki in each area in which the ruled line ki is included, and add it sequentially. The pattern vector P is determined by

【0014】以下、パターンベクトルを求める動作を図
8に従って具体的に説明する。
The operation for obtaining a pattern vector will be explained in detail below with reference to FIG.

【0015】先ず図9のように入力帳票4をM×Nのエ
リアに区分する(ステップS50)。次にパターンベク
トルを初期化し(ステップS51)、全ての罫線が処理
済か否かを判定する(ステップS52)。処理済でない
ときは未処理の罫線kiについて、その罫線kiの長さ
量子化レベルを求める(ステップS53)。長さ量子化
レベルとは、罫線の長さの程度をW段階で量子化して表
現するものである。量子化レベル数はレベル1〜WのW
段階で表わし、帳票の両端を結ぶ罫線のレベルがWにな
るように、横罫線は帳票幅のL0で、縦罫線は帳票高さ
h0で規格化を行なっている。そして、罫線kiの長さ
をLiとすると、その罫線kiの長さ量子化レベルWi
は、次式により求まる
First, the input form 4 is divided into M×N areas as shown in FIG. 9 (step S50). Next, the pattern vector is initialized (step S51), and it is determined whether all ruled lines have been processed (step S52). If the ruled line ki has not been processed, the length quantization level of the unprocessed ruled line ki is determined (step S53). The length quantization level represents the length of a ruled line by quantizing it in W steps. The number of quantization levels is W from level 1 to W.
It is expressed in stages, and so that the level of the ruled line connecting both ends of the form is W, horizontal ruled lines are normalized by the form width L0, and vertical ruled lines are normalized by the form height h0. If the length of the ruled line ki is Li, then the length quantization level Wi of the ruled line ki
is determined by the following formula

【0016】[0016]

【数1】 但し、横罫線のときa0=L0、縦罫線のときa0=h
0である。ceil[X]はXを下まわらない最小の整
数を示す。ここで、a0/Wは長さ量子化レベルの1レ
ベル分に相当する長さとなる。
[Equation 1] However, for horizontal ruled lines, a0=L0, for vertical ruled lines, a0=h
It is 0. ceil[X] indicates the smallest integer not less than X. Here, a0/W is a length corresponding to one length quantization level.

【0017】次に、罫線kiの存在するエリアを検出す
るために罫線の始点エリア、終点エリアを判定する(ス
テップS54、S55)。始点エリア、終点エリアの判
定は罫線情報として既に求められている始点(Xsi,
Ysi)、終点(Xei,Yei)の座標を用いて行な
う。先ず罫線の始点エリアの縦方向位置は、次式により
求まる。
Next, in order to detect the area where the ruled line ki exists, the starting point area and end point area of the ruled line are determined (steps S54 and S55). The starting point area and ending point area are determined based on the starting point (Xsi,
Ysi) and the coordinates of the end point (Xei, Yei). First, the vertical position of the starting point area of the ruled line is determined by the following equation.

【数2】 但し、横罫線のときy=(Ysi+Yei)/2、縦罫
線のときy=Ysi、h0は帳票の高さ、Mは帳票の縦
方向のエリア区分数である。
##EQU00002## where y=(Ysi+Yei)/2 for horizontal ruled lines, y=Ysi for vertical ruled lines, h0 is the height of the form, and M is the number of area divisions in the vertical direction of the form.

【0018】ここで、h0/Mは1エリアの高さに相当
する。また、横方向位置は次式により求まる。
Here, h0/M corresponds to the height of one area. Further, the lateral position is determined by the following equation.

【0019】[0019]

【数3】 但し、横罫線のときx=Xsi、縦罫線のときx=(X
si+Xei)/2、L0は帳票の幅、Nは帳票の横方
向のエリア区分数である。
[Equation 3] However, when it is a horizontal ruled line, x=Xsi, and when it is a vertical ruled line, x=(X
si+Xei)/2, L0 is the width of the form, and N is the number of area divisions in the horizontal direction of the form.

【0020】ここで、L0/Nは1エリアの幅に相当す
る。図10に示す罫線k1,k2の例では、罫線k1で
は始点エリアは(Ms1,Ns1)=(1,2)として
、罫線k2では、始点エリアは(Ms2,Ns2)=(
2,1)として求まる。
[0020] Here, L0/N corresponds to the width of one area. In the example of ruled lines k1 and k2 shown in FIG. 10, the starting point area for ruled line k1 is (Ms1, Ns1) = (1, 2), and the starting point area for ruled line k2 is (Ms2, Ns2) = (
2,1).

【0021】次に、罫線kiの終点エリアの縦方向位置
は次式により求まる。
Next, the vertical position of the end point area of the ruled line ki is determined by the following equation.

【0022】[0022]

【数4】 但し、横罫線のときy=(Ysi+Yei)/2、縦罫
線のときy=Yeiまた、横方向位置は次式により求ま
る。
[Equation 4] However, for horizontal ruled lines, y=(Ysi+Yei)/2, and for vertical ruled lines, y=Yei. Also, the horizontal position is determined by the following equation.

【0023】[0023]

【数5】 但し、横罫線のときx=Xei、縦罫線のときx=(X
si+Xei)/2図10に示す罫線k1,k2の例で
は、罫線k1では終点エリアは(Me1,Ne1)=(
1,2)として、罫線k2では終点エリアは(Me2,
Ne2)=(2,4)として求まる。
[Equation 5] However, when it is a horizontal ruled line, x = Xei, and when it is a vertical ruled line, x = (X
si+Xei)/2 In the example of ruled lines k1 and k2 shown in FIG. 10, the end point area of ruled line k1 is (Me1, Ne1)=(
1, 2), the end point area for ruled line k2 is (Me2,
Ne2)=(2,4).

【0024】この段階で、罫線kiによりパターンベク
トルPが更新される要素の範囲が求まったことになる。 すなわち、罫線kiによりパターンベクトルPが更新さ
れる範囲はエリア(Msi,Nsi)から(Mei,N
ei)までの区間のエリアの、罫線kiの方向(横罫線
か縦罫線の種類)の、罫線kiの長さレベルWiに対応
する要素となる。例えば図10の罫線k1では、始点と
終点が同じエリア(1,2)で横方向罫線であるので、
パターンベクトルPの更新される要素は、P[Ms1]
[Ns1][1][W1](=P[1][2][1][
W1])の1個である。図10の罫線k2の例では、始
点エリアと終点エリアが異なり複数のエリアにまたがっ
ており横罫線であるので、パターンベクトルPの更新さ
れる要素は、P[Ms2][Ns2][1][W2]か
らP[Me2][Ne2][1][W2]までの要素、
すなわちP[2][1][1][W2],P[2][2
][1][W2],P[2][3][1][W2],P
[2][4][1][W2]の4個となる。なお、パタ
ーンベクトルPが更新される重みについては後述する。
At this stage, the range of elements in which the pattern vector P is updated by the ruled line ki has been determined. That is, the range in which the pattern vector P is updated by the ruled line ki is from the area (Msi, Nsi) to (Mei, Nsi).
This element corresponds to the length level Wi of the ruled line ki in the direction of the ruled line ki (horizontal ruled line or vertical ruled line type) in the area of the section up to ei). For example, in the ruled line k1 in FIG. 10, the starting point and ending point are in the same area (1, 2) and is a horizontal ruled line, so
The updated element of pattern vector P is P[Ms1]
[Ns1][1][W1](=P[1][2][1][
W1]). In the example of ruled line k2 in FIG. 10, the start point area and end point area are different and it spans multiple areas and is a horizontal ruled line, so the updated elements of pattern vector P are P[Ms2][Ns2][1][ W2] to P[Me2][Ne2][1][W2],
That is, P[2][1][1][W2], P[2][2
][1][W2],P[2][3][1][W2],P
There are four items: [2], [4], [1], and [W2]. Note that the weight with which the pattern vector P is updated will be described later.

【0025】次に罫線kiによりパターンベクトルPが
更新される重みを計算し、パターンベクトルPを更新す
る。罫線が1つのエリア内に入るとき(図10の罫線k
1の場合)と、複数のエリアにまたがるとき(図10の
罫線k2の場合)では、重みの更新の方法が異なるので
、先ず罫線kiについて求めた始点エリアと終点エリア
が一致するかを判定する(ステップS56)。同じエリ
アであるときは、罫線kiは1つのエリア内に入ると判
定して重み加算処理1(ステップS57)を行ない、一
致しないときは、罫線kiは複数のエリアにまたがると
判定して重み加算処理2(ステップS58)を行なうこ
とにより、パターンベクトルPの更新を行なう。重み加
算処理1及び重み加算処理2については後述する。この
ようにして罫線kiによりパターンベクトルPを更新し
、順次全ての罫線により(ステップS59,S52)パ
ターンベクトルPを更新することにより、入力帳票の持
つ罫線情報をパターンベクトル化することができる。
Next, the weight by which the pattern vector P is updated by the ruled line ki is calculated, and the pattern vector P is updated. When a ruled line falls within one area (ruled line k in Figure 10)
1) and when spanning multiple areas (in the case of ruled line k2 in FIG. 10), the weight update method is different, so first, it is determined whether the starting point area and end point area found for ruled line ki match. (Step S56). If they are in the same area, it is determined that the ruled line ki falls within one area and weight addition processing 1 (step S57) is performed; if they do not match, it is determined that the ruled line ki spans multiple areas and weighted addition is performed. By performing process 2 (step S58), the pattern vector P is updated. Weight addition processing 1 and weight addition processing 2 will be described later. In this way, by updating the pattern vector P using the ruled line ki and sequentially updating the pattern vector P using all the ruled lines (steps S59, S52), the ruled line information of the input form can be converted into a pattern vector.

【0026】次に重み加算処理1(図8のステップS5
7)の動作を説明する。この処理は図10の罫線k1の
例のように、1つのエリア内に存在している罫線による
パターンベクトルPの更新の処理を行うものである。1
本の罫線によりパターンベクトルPが更新される重みの
総和は常に“1”として規格化する。このことにより、
重み加算処理1で更新されるパターンベクトルの要素は
1個のみであるので、このときの重み更新量はa0=1
となる。従って、罫線kiの始点エリア(Msi,Ns
i)と終点エリア(Mei,Nei)が等しく、罫線k
iが1つのエリア内に存在していると判定されたときの
パターンベクトルPの更新は、次の数6で行なわれる。
Next, weight addition processing 1 (step S5 in FIG. 8)
The operation of 7) will be explained. This process updates the pattern vector P based on ruled lines existing in one area, as in the example of ruled line k1 in FIG. 10. 1
The sum of the weights by which the pattern vector P is updated by the ruled lines of the book is always normalized as "1". Due to this,
Since only one element of the pattern vector is updated in weight addition process 1, the weight update amount at this time is a0=1
becomes. Therefore, the starting point area (Msi, Ns
i) and the end point area (Mei, Nei) are equal, and the ruled line k
When it is determined that i exists within one area, the pattern vector P is updated using the following equation 6.

【0027】[0027]

【数6】P[Msi][Nsi][t][Wi]=P[
Msi][Nsi][t][Wi]+a0但し、横罫線
のときt=1、縦罫線のときt=2、a0=1 次に重み加算処理2の動作を説明する。この処理は、図
10の罫線k2の例のように複数のエリアにまたがって
存在している罫線によるパターンベクトルPの更新の処
理を行なうものである。重み加算処理1でも述べたよう
に、1本の罫線によりパターンベクトルPが更新される
重みの総和を常に“1”として規格化する。このことよ
り、重み加算処理2では、複数のエリアにまたがる罫線
のそれぞれのエリアに含まれている長さの罫線全体に対
する比率を、それぞれのエリアの重みの更新量とする。 これにより、パターンベクトルPが1本の罫線により更
新される重みの総和を全て“1”とすることができる。 図11がこの重み加算処理2の動作を示すフローチャー
トであり、以下、図11に従って具体的に説明する。
[Formula 6] P[Msi][Nsi][t][Wi]=P[
Msi][Nsi][t][Wi]+a0 However, for horizontal ruled lines, t=1; for vertical ruled lines, t=2, a0=1 Next, the operation of weight addition processing 2 will be explained. This process updates the pattern vector P based on a ruled line that spans a plurality of areas, such as the example of ruled line k2 in FIG. As described in the weight addition process 1, the sum of the weights by which the pattern vector P is updated by one ruled line is always standardized as "1". Therefore, in the weight addition process 2, the ratio of the length of a ruled line that spans a plurality of areas included in each area to the entire ruled line is used as the update amount of the weight of each area. Thereby, the total sum of the weights of the pattern vector P updated by one ruled line can all be set to "1". FIG. 11 is a flowchart showing the operation of this weight addition process 2, and will be specifically explained below with reference to FIG.

【0028】この重み加算処理2で扱う罫線は、図10
の罫線k2のように複数のエリアにまたがるものである
。先ず、始点エリアについて重みasを求める(ステッ
プS60)。罫線kiの始点エリアに含まれている部分
の長さの、罫線ki全体の長さに対する比率を重みとし
て求める。具体的には罫線kiが横罫線の場合は数7で
、罫線kiが縦罫線の場合は数8で求まる。
The ruled lines handled in this weight addition process 2 are as shown in FIG.
It spans multiple areas like the ruled line k2. First, the weight as is determined for the starting point area (step S60). The ratio of the length of the portion included in the starting point area of the ruled line ki to the entire length of the ruled line ki is determined as a weight. Specifically, if the ruled line ki is a horizontal ruled line, it can be determined by Equation 7, and if the ruled line ki is a vertical ruled line, it can be determined by Equation 8.

【0029】[0029]

【数7】[Math 7]

【0030】[0030]

【数8】 数7及び数8共に分子第1項は始点から終点をみたとき
の始点エリアの境界座標であり、従って分子は罫線ki
の始点エリアに含まれている部分の長さであり、分母の
罫線kiの長さLiで規格化することにより、始点エリ
アに含まれている部分の長さの全体に対する比率が求め
られる。数7及び数8により求められた始点重みasi
を用いて、パターンベクトルの始点エリア部分を数9に
より更新する(ステップS61)。
[Equation 8] In both Equations 7 and 8, the first term of the numerator is the boundary coordinate of the starting point area when looking from the starting point to the end point, so the numerator is
is the length of the part included in the starting point area, and by normalizing it by the length Li of the ruled line ki in the denominator, the ratio of the length of the part included in the starting point area to the whole can be found. Starting point weight asi determined by Equations 7 and 8
The starting point area of the pattern vector is updated using Equation 9 (step S61).

【0031】[0031]

【数9】   P[Msi][Nsi][t][wi]=P[Ms
i][Nsi][t][wi]+asi 但し、横罫線のとき  t=1 縦罫線のとき  t=2 次に終点エリアについて重みaeiを求める(ステップ
S62)。始点エリア重みasiと同様に、罫線kiの
終点エリアに含まれている部分の長さの、罫線ki全体
の長さに対する比率を重みとして求める。具体的には罫
線kiが横罫線の場合は数10で、罫線kiが縦罫線の
場合は数11で求まる。
[Formula 9] P[Msi][Nsi][t][wi]=P[Ms
i][Nsi][t][wi]+asi However, for horizontal ruled lines, t=1; for vertical ruled lines, t=2 Next, the weight aei is determined for the end point area (step S62). Similar to the starting point area weight asi, the ratio of the length of the portion of the ruled line ki included in the end point area to the entire length of the ruled line ki is determined as the weight. Specifically, if the ruled line ki is a horizontal ruled line, it can be found by Equation 10, and if the ruled line ki is a vertical ruled line, it can be found by Equation 11.

【0032】[0032]

【数10】[Math. 10]

【0033】[0033]

【数11】 数10及び数11共に分子第2項は終点から始点をみた
ときの終点エリアの境界座標であり、従って分子は罫線
kiの終点エリアに含まれている部分の長さであり、分
母の罫線kiの長さLiで規格化することにより、終点
エリアに含まれている部分の長さの全体に対する比率が
求められる。数10及び数11により求められた終点重
みaeiを用いて、パターンベクトルの終点エリア部分
を数12により更新する(ステップ63)。
[Equation 11] In both Equations 10 and 11, the second term in the numerator is the boundary coordinate of the end point area when looking from the end point to the start point, and therefore the numerator is the length of the portion of the ruled line ki included in the end point area, By normalizing by the length Li of the ruled line ki in the denominator, the ratio of the length of the portion included in the end point area to the entire length is determined. Using the end point weight aei obtained from Equations 10 and 11, the end point area portion of the pattern vector is updated using Equation 12 (step 63).

【0034】[0034]

【数12】   P[Mei][Nei][t][wi]=P[Me
i][Nei][t][wi]+aei 但し、横罫線のとき  t=1 縦罫線のとき  t=2 次に、始点エリアと終点エリアに挟まれた部分のエリア
(以下、中間エリアという)についての重み更新を行な
う。先ず始点エリアと終点エリアが隣り合わせのエリア
かどうかを判定する(ステップS64)。始点エリアと
終点エリアが隣り合わせのエリアかどうかの判定は、横
罫線の場合は(Nsi+1)とNeiが等しいかどうか
、縦罫線の場合は(Msi+1)とMeiが等しいかど
うかで判定される。この判定の結果、始点エリアと終点
エリアが隣り合ったエリアと判定された場合は、中間エ
リアはないとして罫線kiについての重み加算処理2を
終了する。また、隣り合っていないと判定された場合は
、中間エリアについて重みamiを求める(ステップS
65)。始点エリア重みasi,終点エリア重みaei
と同様に、罫線kiの中間エリア1つに含まれている部
分の長さの、罫線ki全体の長さに対する比率を重みと
して求める。具体的には罫線kiが横罫線の場合は数1
3で罫線kiが横罫線の場合は数14で求まる。ここで
中間エリア1つに含まれる罫線kiの長さは、そのエリ
アの大きさと等しくなるので、これで置きかえると、
[Formula 12] P[Mei][Nei][t][wi]=P[Me
i] [Nei] [t] [wi] + aei However, for horizontal ruled lines t=1 For vertical ruled lines t=2 Next, the area sandwiched between the start point area and end point area (hereinafter referred to as the intermediate area) Update the weights for . First, it is determined whether the starting point area and the ending point area are adjacent areas (step S64). Whether the starting point area and the ending point area are adjacent areas is determined by whether (Nsi+1) and Nei are equal in the case of horizontal ruled lines, and whether (Msi+1) and Mei are equal in the case of vertical ruled lines. As a result of this determination, if it is determined that the start point area and the end point area are adjacent areas, it is assumed that there is no intermediate area and the weight addition process 2 for the ruled line ki is ended. Furthermore, if it is determined that they are not adjacent, the weight ami is calculated for the intermediate area (step S
65). Starting point area weight asi, ending point area weight aei
Similarly, the ratio of the length of the portion included in one intermediate area of the ruled line ki to the length of the entire ruled line ki is determined as a weight. Specifically, if the ruled line ki is a horizontal ruled line, the formula is 1
If the ruled line ki is a horizontal ruled line in 3, it can be found by Equation 14. Here, the length of the ruled line ki included in one intermediate area is equal to the size of that area, so if you replace it with this,

【0035】[0035]

【数13】[Math. 13]

【0036】[0036]

【数14】 である。数13及び数14共に分子は中間エリア1つを
横切るときの長さで、それは横罫線のときはエリアの幅
と等しく、縦罫線のときはエリアの高さと等しくなる。 これを罫線kiの長さLiで規格化することにより、中
間エリア1つに含まれている部分の長さの全体に対する
比率が求められる。数13及び数14により求められた
中間エリア重みamiを用いて、パターンベクトルの全
ての中間エリアの部分を横罫線の場合は数15により、
また縦罫線の場合は数16により更新する(ステップS
66)。横罫線の場合
[Formula 14]. In both Equations 13 and 14, the numerator is the length when crossing one intermediate area, which is equal to the width of the area in the case of a horizontal ruled line, and equal to the height of the area in the case of a vertical ruled line. By normalizing this with the length Li of the ruled line ki, the ratio of the length of the portion included in one intermediate area to the entire length is determined. Using the intermediate area weight ami obtained by Equations 13 and 14, if all intermediate areas of the pattern vector are horizontal ruled lines, use Equation 15 as follows:
In addition, in the case of vertical ruled lines, it is updated according to equation 16 (step S
66). For horizontal ruled lines

【0037】[0037]

【数15】   P[Msi][n][1][wi]=P[Msi]
[n][1][wi]+ami 但し、nは(Nsi+1)〜(Nei−1)の区間全て
縦罫線の場合
[Formula 15] P[Msi][n][1][wi]=P[Msi]
[n][1][wi]+ami However, when n is all vertical ruled lines in the interval from (Nsi+1) to (Nei-1)

【0038】[0038]

【数16】   P[m][Nsi][2][wi]=P[m][N
si][2][wi]+ami 但し、mは(Msi+1)〜(Mei−1)の区間全て
このようにして罫線kiによりパターンベクトルPの更
新が行なわれる。例えば図10に示す罫線k1によるパ
ターンベクトルPの更新は1要素のみで、P[1][2
][1][w1]=P[1][2][1][w1]+1
となる。また罫線k2によるパターンベクトルPの更新
は4要素で
[Formula 16] P[m][Nsi][2][wi]=P[m][N
si][2][wi]+ami However, the pattern vector P is updated in this manner using the ruled line ki throughout the interval from (Msi+1) to (Mei-1). For example, the pattern vector P based on the ruled line k1 shown in FIG. 10 is updated with only one element, P[1][2
][1][w1]=P[1][2][1][w1]+1
becomes. Also, the pattern vector P is updated by ruled line k2 with 4 elements.

【0039】[0039]

【数17】   P[2][1][1][w2]=P[2][1][
1][w2]+as2  P[2][2][1][w2
]=P[2][2][1][w2]+am2  P[2
][3][1][w2]=P[2][3][1][w2
]+am2  P[2][4][1][w2]=P[2
][4][1][w2]+ae2となる。以上の処理動
作(ステップS52〜59)を全ての横罫線について行
ない、その後全ての縦罫線について行なうことにより入
力帳票の罫線情報をパターンベクトル化し、パターンベ
クトルPが求められる。
[Formula 17] P[2][1][1][w2]=P[2][1][
1][w2]+as2 P[2][2][1][w2
]=P[2][2][1][w2]+am2 P[2
][3][1][w2]=P[2][3][1][w2
]+am2 P[2][4][1][w2]=P[2
][4][1][w2]+ae2. By performing the above processing operations (steps S52 to S59) for all horizontal ruled lines and then for all vertical ruled lines, the ruled line information of the input form is converted into a pattern vector, and a pattern vector P is obtained.

【0040】次に上述のようにして求めた認識対象帳票
のパターンベクトルPに基ずいて、その帳票の種類の識
別に移る。帳票の種類の識別は、先ず帳票のパターンベ
クトルPと、判別すべき複数種類の基本帳票の予め定め
られた標準パターンベクトルSx(但し、xは標準パタ
ーン番号)とを比較して距離計算をし、候補を決定する
(ステップS6)。距離計算は、パターンベクトルPと
標準パターンSとの間でシティブロック距離を用いた距
離計算を行なう。その距離dxは次式で求まる。
Next, based on the pattern vector P of the form to be recognized obtained as described above, the process moves on to identifying the type of the form. To identify the type of form, first, the distance is calculated by comparing the pattern vector P of the form with a predetermined standard pattern vector Sx (where x is the standard pattern number) of the plurality of basic forms to be discriminated. , a candidate is determined (step S6). Distance calculation is performed between pattern vector P and standard pattern S using city block distance. The distance dx is determined by the following formula.

【0041】[0041]

【数18】 そして、距離dxが予め設定したしきい値以下になるも
のを全て帳票の候補とする。このとき距離の小さいもの
から順に並べておいて、詳細確認での優先順位とする。 詳細確認においては、距離判定で候補として選択された
帳票についてそれぞれ特定領域の解析と検定を行なう(
ステップS7)。このとき、距離判定での優先順位に従
った順番に解析、検定を行ない、満足する結果が得られ
た時点で入力された帳票の種類を決定し、未処理のまま
残された候補は破棄することにより、処理の効率化を計
る。
##EQU00001## Then, all documents whose distance dx is less than or equal to a preset threshold are determined as document candidates. At this time, the distances are arranged in descending order of distance, and this is used as the priority order for detailed confirmation. In detailed confirmation, each specific area of the form selected as a candidate in the distance judgment is analyzed and verified (
Step S7). At this time, analysis and verification are performed in the order of priority in distance determination, and when a satisfactory result is obtained, the type of input form is determined, and candidates that remain unprocessed are discarded. By doing so, we aim to improve processing efficiency.

【0042】詳細確認ができれば最終候補を決定し(ス
テップS9)、帳票の種類が決定したことになる。詳細
確認ができなければ次の候補が有るか否かを判定し(ス
テップS10)、次の候補が有ればステップS7にリタ
ーンし、無ければリジェクトとなる。
[0042] Once the details have been confirmed, the final candidates are determined (step S9), and the type of form has been determined. If detailed confirmation is not possible, it is determined whether or not there is a next candidate (step S10). If there is a next candidate, the process returns to step S7; if there is not, the process is rejected.

【0043】ここで詳細確認を説明すると、詳細確認は
、標準帳票毎に特徴のある特定領域を予め定めておき、
候補として挙がった標準帳票のその特定領域に相当する
領域を入力帳票画像から切出し、その領域域内の文字や
番号を読取って確認したり、又、切り出された領域につ
いて濃淡パターンを読取って一致するかどうかを確認す
る処理である。例えば図1に示す帳票の場合では詳細確
認用特定領域1内の文字を読取り、「振替伝票」と書か
れてあるか否かで詳細確認を行なう。そして、最終的に
帳票の種類が決定されると、その帳票の記載形式に従っ
たその帳票の情報記入欄を切出し、その欄の内容を読取
り(ステップS11)、帳票の種類の判別及び読取処理
を終了する。
[0043] Detailed confirmation will be explained here. Detailed confirmation involves predetermining specific areas with characteristics for each standard form.
An area corresponding to the specific area of the standard form that has been raised as a candidate is cut out from the input form image, and the characters and numbers within that area are read and confirmed, and the grayscale pattern of the cut out area is read to see if they match. This is a process to check whether the For example, in the case of the form shown in FIG. 1, the characters in the specific area 1 for checking details are read and the details are checked to see if it says "transfer slip". When the type of the form is finally determined, the information entry field of the form is cut out according to the writing format of the form, and the contents of that field are read (step S11), and the form type is determined and the reading process is performed. end.

【0044】なお、上記実施例では入力画像の罫線の情
報により類似する帳票の候補を抽出して、その候補とし
て挙がった標準帳票の特定領域で詳細確認を行なってい
るが、この詳細確認を省略して候補の一番目のものと決
定してもよい。特に、特定領域として設けていない帳票
については詳細確認は不要である。
[0044] In the above embodiment, similar form candidates are extracted based on the information on the ruled lines of the input image, and detailed confirmation is performed in a specific area of the standard form selected as the candidate, but this detailed confirmation is omitted. The first candidate may be determined as the first candidate. In particular, detailed confirmation is not necessary for forms that are not provided as specific areas.

【0045】[0045]

【発明の効果】以上のように本願発明によれば帳票を形
成する罫線の情報を基にベクトルパターン化したパター
ンベクトルを求めて標準帳票のものと比較照合するよう
にしているので、帳票に種類番号やバーコードを付加し
ていないものであっても帳票の種類の判別を可能とする
ことができる。
Effects of the Invention As described above, according to the present invention, a pattern vector is obtained by converting it into a vector pattern based on the information of ruled lines forming a form, and is compared with that of a standard form. It is possible to distinguish the type of form even if it does not have a number or bar code added to it.

【0046】又、一旦パターンベクトルと標準帳票のも
のと比較照合して候補を抽出し、候補順に候補として挙
がった帳票の一部を確認するようにしたので、確実に帳
票の種類を判別できる。しかも、このようにしたことに
より帳票毎に特徴領域や番号、バーコードの付されてい
る位置が異なっても、最終的に詳細確認の段階で行なう
のみで良く、通常最後に一回特定の場所を見るのみで良
く、種類判定可能とする帳票の特定領域の全てについて
確認する必要がなく、判別処理の高速化を実現できる。
Further, candidates are extracted by comparing the pattern vector with that of a standard form, and some of the forms listed as candidates are checked in the order of the candidates, so that the type of form can be reliably determined. Moreover, by doing this, even if the characteristic areas, numbers, and barcode positions differ for each document, it only needs to be done at the final detailed confirmation stage; It is not necessary to check all the specific areas of the form for which the type can be determined, and the speed of the determination process can be increased.

【図面の簡単な説明】[Brief explanation of the drawing]

【図1】本発明が対象とする帳票の一例を示す図である
FIG. 1 is a diagram showing an example of a form targeted by the present invention.

【図2】本発明の動作例を示すフローチャートである。FIG. 2 is a flowchart showing an example of the operation of the present invention.

【図3】スムージング処理を説明するための図である。FIG. 3 is a diagram for explaining smoothing processing.

【図4】罫線抽出の動作例を示すフローチャートである
FIG. 4 is a flowchart illustrating an operation example of ruled line extraction.

【図5】罫線抽出の動作を説明するための図である。FIG. 5 is a diagram for explaining the ruled line extraction operation.

【図6】図1に示す帳票の処理結果を示す図である。FIG. 6 is a diagram showing a processing result of the form shown in FIG. 1;

【図7】パターンベクトルを説明するための図である。FIG. 7 is a diagram for explaining pattern vectors.

【図8】パターンベクトルを求める動作例を示すフロー
チャートである。
FIG. 8 is a flowchart showing an example of operation for obtaining a pattern vector.

【図9】入力帳票をM×Nのエリアに区分する様子を示
す図である。
FIG. 9 is a diagram showing how an input form is divided into M×N areas.

【図10】帳票上の罫線の例を示す図である。FIG. 10 is a diagram showing an example of ruled lines on a form.

【図11】重み加算処理の動作例を示すフローチャート
である。
FIG. 11 is a flowchart showing an example of the operation of weight addition processing.

【符号の説明】[Explanation of symbols]

1  特定領域 2  注目画素 3  8近傍画素 4  入力帳票 5  始点エリア 6  終点エリア 1 Specific area 2 Pixel of interest 3 8 neighboring pixels 4 Input form 5 Starting point area 6 End area

Claims (2)

【特許請求の範囲】[Claims] 【請求項1】  被認識帳票類の画像情報を入力し、こ
の入力された画像データから水平、垂直方向の線分を抽
出し、抽出された各線分の位置、長さ、方向の情報を検
出し、前記入力された画像を複数エリアに区分し、この
区分されたエリア毎に存在する前記抽出された線分の方
向毎の長さ毎の重みである特徴量を被認識対象のパター
ン特徴量として求め、認識すべきパターンの基準となる
標準パターンの特徴量を予め記憶しておき、前記抽出さ
れた被認識パターンの特徴量と前記記憶されている標準
パターンの特徴量とを比較照合し、その比較結果に基づ
いて前記被認識パターンがどの標準パターンであるかを
判断して帳票の種類を判別することを特徴とする帳票類
の種類判別方法。
[Claim 1] Image information of documents to be recognized is input, horizontal and vertical line segments are extracted from the input image data, and information on the position, length, and direction of each extracted line segment is detected. Then, the input image is divided into a plurality of areas, and the feature quantity, which is the weight for each direction of the extracted line segment existing in each divided area, is used as the pattern feature quantity of the target to be recognized. and store in advance the feature quantities of a standard pattern that serve as a reference for the pattern to be recognized, and compare and match the extracted feature quantities of the recognized pattern with the stored feature quantities of the standard pattern; A method for determining types of forms, characterized in that the types of forms are determined by determining which standard pattern the recognized pattern is based on the comparison result.
【請求項2】  被認識帳票類の画像情報を入力し、こ
の入力された画像データから水平、垂直方向の線分を抽
出し、抽出された各線分の位置、長さ、方向の情報を検
出し、前記入力された画像を複数エリアに区分し、この
区分されたエリア毎に存在する前記抽出された線分の方
向毎の長さ毎の重みである特徴量を被認識対象のパター
ン特徴量として求め、認識すべきパターンの基準となる
標準パターンの特徴量を予め記憶しておき、前記抽出さ
れた被認識パターンの特徴量と前記記憶されている標準
パターンの特徴量とを比較照合し、所定値以上類似する
ものを類似する順に候補として抽出し、候補として抽出
された標準パターンの予め定められた特定領域に対応す
る入力画像情報を検出して、前記標準パターンのものと
一致するかを前記候補順に確認することによって帳票類
の種類を判別することを特徴とする帳票類の種類判別方
法。
[Claim 2] Input image information of documents to be recognized, extract horizontal and vertical line segments from the input image data, and detect information on the position, length, and direction of each extracted line segment. Then, the input image is divided into a plurality of areas, and the feature quantity, which is the weight for each direction of the extracted line segment existing in each divided area, is used as the pattern feature quantity of the target to be recognized. and store in advance the feature quantities of a standard pattern that serve as a reference for the pattern to be recognized, and compare and match the extracted feature quantities of the recognized pattern with the stored feature quantities of the standard pattern; Items that are similar by more than a predetermined value are extracted as candidates in order of similarity, and input image information corresponding to a predetermined specific area of the standard pattern extracted as a candidate is detected to determine whether it matches that of the standard pattern. A method for determining types of forms, characterized in that the types of forms are determined by checking the candidates in the order of the candidates.
JP03050497A 1991-02-22 1991-02-22 How to determine the type of form Expired - Fee Related JP3096481B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP03050497A JP3096481B2 (en) 1991-02-22 1991-02-22 How to determine the type of form

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP03050497A JP3096481B2 (en) 1991-02-22 1991-02-22 How to determine the type of form

Publications (2)

Publication Number Publication Date
JPH04268685A true JPH04268685A (en) 1992-09-24
JP3096481B2 JP3096481B2 (en) 2000-10-10

Family

ID=12860579

Family Applications (1)

Application Number Title Priority Date Filing Date
JP03050497A Expired - Fee Related JP3096481B2 (en) 1991-02-22 1991-02-22 How to determine the type of form

Country Status (1)

Country Link
JP (1) JP3096481B2 (en)

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH07114616A (en) * 1993-10-20 1995-05-02 Hitachi Ltd Slip document information system
JP2003030583A (en) * 2001-07-11 2003-01-31 Oki Electric Ind Co Ltd Method and device for identifying chart classification, and method and device for identifying format classification
US6813381B2 (en) 2000-03-30 2004-11-02 Glory Ltd. Method and apparatus for identification of documents, and computer product
JP2015528960A (en) * 2012-07-24 2015-10-01 アリババ・グループ・ホールディング・リミテッドAlibaba Group Holding Limited Form recognition method and form recognition apparatus
JP2017199086A (en) * 2016-04-25 2017-11-02 富士通株式会社 Method, device, program, and dictionary data for recognizing business form
JP2019040585A (en) * 2017-06-30 2019-03-14 コニカ ミノルタ ラボラトリー ユー.エス.エー.,インコーポレイテッド Typesetness score for table

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP3853943B2 (en) 1997-11-12 2006-12-06 株式会社ジェイテクト Vehicle steering device

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH07114616A (en) * 1993-10-20 1995-05-02 Hitachi Ltd Slip document information system
US6813381B2 (en) 2000-03-30 2004-11-02 Glory Ltd. Method and apparatus for identification of documents, and computer product
JP2003030583A (en) * 2001-07-11 2003-01-31 Oki Electric Ind Co Ltd Method and device for identifying chart classification, and method and device for identifying format classification
JP2015528960A (en) * 2012-07-24 2015-10-01 アリババ・グループ・ホールディング・リミテッドAlibaba Group Holding Limited Form recognition method and form recognition apparatus
JP2017199086A (en) * 2016-04-25 2017-11-02 富士通株式会社 Method, device, program, and dictionary data for recognizing business form
JP2019040585A (en) * 2017-06-30 2019-03-14 コニカ ミノルタ ラボラトリー ユー.エス.エー.,インコーポレイテッド Typesetness score for table

Also Published As

Publication number Publication date
JP3096481B2 (en) 2000-10-10

Similar Documents

Publication Publication Date Title
US6687401B2 (en) Pattern recognizing apparatus and method
JP6080259B2 (en) Character cutting device and character cutting method
JP6900164B2 (en) Information processing equipment, information processing methods and programs
JP3096481B2 (en) How to determine the type of form
EP0651337A1 (en) Object recognizing method, its apparatus, and image processing method and its apparatus
JP2898562B2 (en) License plate determination method
CN110210467A (en) A kind of formula localization method, image processing apparatus, the storage medium of text image
JP4853313B2 (en) Character recognition device
JP3622347B2 (en) Form recognition device
US4607387A (en) Pattern check device
CN102682308A (en) Imaging processing method and device
JPH04111085A (en) Pattern recognizing device
JP2899383B2 (en) Character extraction device
JPH10222602A (en) Optical character reading device
JP2965165B2 (en) Pattern recognition method and recognition dictionary creation method
JP4810853B2 (en) Character image cutting device, character image cutting method and program
KR101001693B1 (en) Character recognition method of giro ticket holder
JP4221960B2 (en) Form identification device and identification method thereof
JP7343115B1 (en) information processing system
CN121505641B (en) Intelligent verification method and system for financial document
JP2902097B2 (en) Information processing device and character recognition device
JPS5850078A (en) Character recognizing device
JP2832035B2 (en) Character recognition device
JPH08115428A (en) Area discriminator
JP2959054B2 (en) Line type discrimination method in pattern recognition device

Legal Events

Date Code Title Description
LAPS Cancellation because of no payment of annual fees