JPH01283671A - information processing equipment - Google Patents

information processing equipment

Info

Publication number
JPH01283671A
JPH01283671A JP63114173A JP11417388A JPH01283671A JP H01283671 A JPH01283671 A JP H01283671A JP 63114173 A JP63114173 A JP 63114173A JP 11417388 A JP11417388 A JP 11417388A JP H01283671 A JPH01283671 A JP H01283671A
Authority
JP
Japan
Prior art keywords
keyword
character
string
information
ocr
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP63114173A
Other languages
Japanese (ja)
Inventor
Shoji Ihara
正二 井原
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Canon Inc
Original Assignee
Canon Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Canon Inc filed Critical Canon Inc
Priority to JP63114173A priority Critical patent/JPH01283671A/en
Publication of JPH01283671A publication Critical patent/JPH01283671A/en
Pending legal-status Critical Current

Links

Landscapes

  • Character Discrimination (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

PURPOSE:To usefully and efficiently use the surface of a sheet by providing a means for managing a discriminated specific character-string as attached information of document information. CONSTITUTION:An input means 5 inputs document information, a discriminating means 7 discriminates a specific character-string arranged regularly from the document information, and a managing means 11 manages the specific character- string as attached information of the document information. From the character- string which has been brought to character recognition from an OCR input sheet, a delimiting symbol, a continuance symbol and an end symbol are discriminated, and in accordance therewith, the character-string is registered in each item of a key word. In such a way, at the time of entering the character-string to be registered into each key word item on the OCR input sheet, only a necessary character-string is put close and entered, and it is unnecessary to leave a surplus space, therefore, the space of the OCR input sheet can be used effectively.

Description

【発明の詳細な説明】 [産業上の利用分野] 本発明は情報処理装置に関し、例えば文書情報を特定な
文字列から成る付属情報と対応づけて保存する情報処理
装置に関するものである。
DETAILED DESCRIPTION OF THE INVENTION [Industrial Field of Application] The present invention relates to an information processing device, and more particularly, to an information processing device that stores document information in association with attached information consisting of a specific character string.

[従来の技術] 従来、この種の装置においては、スキャナから文書をイ
メージデータとして読込み保存する機能を有した情報処
理装置があり、例えばOCR(光学式文字読取装置)入
力シートに書き込まれたそれぞれの文書情報(以下、「
文書画像」という)に付属情報のキーワードを付加し、
そのキーワードを検索する事により、保存した幾つかの
文書画像の中から目的の文書画像を見つけ出している。
[Prior Art] Conventionally, in this type of device, there is an information processing device that has a function of reading and saving a document as image data from a scanner. document information (hereinafter referred to as “
Add keywords of attached information to the document image),
By searching for the keyword, the desired document image is found from among several saved document images.

このときのキーワードは、予め幾つかの項目が定められ
、それぞれの項目については該当事項を入力するように
して用いられている。
Several items are predetermined as keywords at this time, and each item is used by inputting the relevant information.

このように、文書毎にキーワードを入力する場合には、
OCR入力シートに、手書で、各項目のキーワードを書
き込み、それをスキャナで読込ませ、手書き文字をOC
R認識させる事により、キーワード入力させる方法を用
いている。従ってOCR入力シートとしては、予めキー
ワードの各項目の入力欄の位置が、用紙上で規定されて
いた。
In this way, when entering keywords for each document,
Write the keywords for each item by hand on the OCR input sheet, read it with a scanner, and OC the handwritten characters.
A method is used in which keywords are input by having R recognized. Therefore, in the OCR input sheet, the positions of the input fields for each keyword item are predefined on the paper.

[発明が解決しようとする課題] しかしながら上記従来例では、OCR入力シートにおい
て、キーワードの各項目を記入する欄が予め規定されて
いるため、以下に述べるような欠点がある。
[Problems to be Solved by the Invention] However, in the conventional example described above, the OCR input sheet has predetermined fields for entering each keyword item, and therefore has the following drawbacks.

(1)キーワードの各項目の記入欄は、その項目で入力
が許されている最大文字数分のスペースが必要である。
(1) The entry field for each keyword item requires space equal to the maximum number of characters that are allowed to be entered in that item.

ところが、項目によっては、一部の例外的データや将来
の拡張性を考えて、この入力最大文字数を通常のデータ
に比べて大きめに設定する場合がある。このような項目
では、通常のキーワードを記入すると、常に、欄に余白
が生じ、その分だけ、OCR入力シートの紙面が無駄と
なり、その結果多数のキーワードを入力する場合、OC
R入力シートを余分に必要としていた。
However, depending on the item, the maximum number of input characters may be set to be larger than normal data in consideration of some exceptional data or future expandability. In such items, when entering regular keywords, there is always a blank space in the field, which wastes space on the OCR input sheet.As a result, when entering a large number of keywords, the OC
An extra R input sheet was required.

(2)OCR入力シートに、キーワードに対応する文書
画像を貼り付はキーワードと同時に、文書画像もスキャ
ナから読込ませる場合、キーワードの各項目の記入欄が
、固定位置にあると、それに従って文書画像の貼り付け
られる場所も制限され、その場所から、キーワード記入
欄に少しでもはみ出る画像は処理することができなかっ
た。
(2) Paste the document image corresponding to the keyword on the OCR input sheet When reading the document image from the scanner at the same time as the keyword, if the entry fields for each keyword item are in a fixed position, the document image will be There are also restrictions on where images can be pasted, and images that extend even slightly into the keyword entry field cannot be processed.

本発明は上述した従来例の欠点に鑑みてなされたもので
あり、その目的とするところは、文字認識シート、例え
ばOCR入力シート上に文字を有効に効率的に記入する
と共に、特定な文字列、例えばキーワードの文字数を調
整し、何枚かのOCR入力シートに分割して書き込める
ような融通のきく情報処理装置を提供する点にある。
The present invention has been made in view of the drawbacks of the conventional examples described above, and its purpose is to write characters effectively and efficiently on a character recognition sheet, for example, an OCR input sheet, and to write a specific character string. The object of the present invention is to provide a flexible information processing device that can, for example, adjust the number of characters of a keyword and write it on several OCR input sheets separately.

[課題を解決するための手段] 上述した問題点を解決し、目的を達成するため、この発
明に係わる情報処理装置は、文書情報を特定な文字列か
ら成る付属情報と対応づけて保存する情報処理装置にお
いて、前記文書情報を入力する入力手段と、該入力手段
の入力した文書情報より規則的に並ぶ特定な文字列を識
別する識別手段と、前記識別した特定な文字列を前記文
書情報の付属情報として管理する管理手段とを備えるこ
とを特徴とする。
[Means for Solving the Problems] In order to solve the above-mentioned problems and achieve the purpose, an information processing device according to the present invention stores information in which document information is associated with attached information consisting of a specific character string. In the processing device, an input means for inputting the document information, an identification means for identifying a specific character string arranged regularly from the document information inputted by the input means, and an identification means for identifying a specific character string arranged regularly from the document information inputted by the input means; The information processing apparatus is characterized by comprising a management means for managing the attached information.

[作用] 以上の構成によれば、入力手段は文書情報を入力し、識
別手段は文書情報より規則的に並ぶ特定な文字列を識別
し、管理手段は特定な文字列を前記文書情報の付属情報
として管理する。
[Operation] According to the above configuration, the input means inputs document information, the identification means identifies a specific character string arranged regularly from the document information, and the management means identifies the specific character string as an attachment of the document information. Manage as information.

[実施例] 以下添付図面を参照して本発明に係る好適な実施例を詳
細に説明する。
[Embodiments] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings.

く情報処理装置の構成の説明(第1図)〉第1図は本発
明の一実施例の構成を示す概略ブロック図である。図に
おいて、1は本実施例の0CR(光学式文字読取装置)
機能を有した情報処理装置であり、2は本装置全体を制
御するCPUである。3は制御プログラム、エラー処理
用プログラム、後述の第4図(a)、(b)に示すフロ
ーチャートに従った処理プログラム等を格納したROM
、4は各種プログラムのワークエリア及びエラー処理時
の一時退避エリアとして用いるRAMである。
Description of Configuration of Information Processing Apparatus (FIG. 1)> FIG. 1 is a schematic block diagram showing the configuration of an embodiment of the present invention. In the figure, 1 is the OCR (optical character reader) of this embodiment.
It is an information processing device with functions, and 2 is a CPU that controls the entire device. 3 is a ROM that stores a control program, an error processing program, a processing program according to the flowcharts shown in FIGS. 4(a) and (b) described later, etc.
, 4 is a RAM used as a work area for various programs and a temporary save area during error processing.

そして、5は光学的にOCR入力シートを読み取って画
像データを入力するスキャナ、6はスキャナ5で読み取
った人力画像データを格納する画像メモリ、7は画像メ
モリ6に格納されている入力画像データよりの文字パタ
ーンから文字を認識して文字コードに変換する文字認識
部である。この文字認識部7の認識結果である文字コー
ドは画像メモリ6に格納される。また8は画像メモリ6
中に格納されている文字コード化された文書情報から文
書の種類等を示す付属情報のキーワードを識別するキー
ワード識別部、9はキーワード識別部8で識別されたキ
ーワード或は通常の文字列を編集する文字列編集部であ
る。
5 is a scanner that optically reads the OCR input sheet and inputs the image data; 6 is an image memory that stores the manual image data read by the scanner 5; and 7 is the input image data stored in the image memory 6. This is a character recognition unit that recognizes characters from character patterns and converts them into character codes. The character code that is the recognition result of the character recognition section 7 is stored in the image memory 6. 8 is image memory 6
A keyword identification section 9 identifies a keyword of attached information indicating the type of document etc. from the character-coded document information stored therein; 9 edits the keyword or normal character string identified by the keyword identification section 8; This is the string editing section.

また、lOは文字列編集部9で編集された文字列を文書
毎に最新の文書画像データとして記憶する光ディスク、
11は文字列編集部9で編集されたキーワード、即ち光
ディスクlo内の文書画像データを検索するために必要
となる管理情報(後述するキーワード管理テーブル)を
格納した磁気ディスクである。12は種々の情報をオペ
レータが入力する各種キーを備えたキーボード、13は
オペレータによるキーボード12の入力手順、画像デー
タ等を表示する表示部、14はアドレス、データ、制御
信号等の伝送路となるシステムバスである。
Further, IO is an optical disk that stores the character strings edited by the character string editing section 9 as the latest document image data for each document;
Reference numeral 11 denotes a magnetic disk that stores keywords edited by the character string editing section 9, that is, management information (keyword management table to be described later) necessary for searching document image data in the optical disk LO. 12 is a keyboard equipped with various keys for the operator to input various information; 13 is a display section for displaying input procedures of the keyboard 12 by the operator; image data; and 14 is a transmission path for addresses, data, control signals, etc. It is a system bus.

〈キーワード管理テーブルの説明(第2図)〉次に、上
述した情報処理装置lによるキーワードの管理方法につ
いて説明する。
<Explanation of Keyword Management Table (FIG. 2)> Next, a method of managing keywords by the information processing device 1 described above will be described.

第2図は本実施例による磁気ディスク11のキーワード
管理テーブルを示す図である。図中の磁気ディスク11
において、100は本実施例のキーワード管理テーブル
であり、画像のアドレスに複数のキーワード項目を対応
させてテーブルを形成している。101は各種の文書画
像データが光デイスク10内のどの位置に格納されてい
るのかを示すアドレス(番地)情報(Ai;i=o。
FIG. 2 is a diagram showing a keyword management table of the magnetic disk 11 according to this embodiment. Magnetic disk 11 in the diagram
100 is a keyword management table of this embodiment, which is formed by associating a plurality of keyword items with image addresses. Reference numeral 101 denotes address information (Ai; i=o) indicating where in the optical disk 10 various document image data are stored.

1.2・・・)が格納されている画像アドレス、102
〜105は、図示の如く、キーワードの各項目(Xi 
; i=o、1.2・・・、j=1.2.3・・・)が
、その項目順に格納されている第1キーワード項目〜第
4キーワード項目である。このとき画像アドレスAiに
対応するキーワードの各項目がX1Nl  、  )(
if21  、  )(ill  、  )(i(41
トなる。尚、本実施例では第2図の如く、キーワード項
目の数を4つとしたが、この項目数は本発明の主旨を逸
脱しない範囲であれば、数を自由に設定することができ
る。
1.2...) is stored at the image address 102
~105 are each keyword item (Xi
; i=o, 1.2..., j=1.2.3...) are the first to fourth keyword items stored in the order of the items. At this time, each item of the keyword corresponding to the image address Ai is X1Nl, )(
if21 , )(ill , )(i(41
It will be. In this embodiment, the number of keyword items is four as shown in FIG. 2, but the number of keyword items can be set freely as long as it does not depart from the gist of the present invention.

ここで、光ディスクに格納されている画像の中から目的
の画像を表示するためには、まず、オペレータがキーボ
ード部12から目的の文書画像データに付加されたキー
ワードの各項目内容を入力する。次に、入力された内容
と各項目内容が一致する行をキーワード管理テーブル1
00の中から探し出し、それと同一行の画像アドレスが
示している文書画像データを光ディスク10から読み出
して、表示部13に表示する。
Here, in order to display a target image from among the images stored on the optical disk, the operator first inputs the contents of each item of the keyword added to the target document image data from the keyboard section 12. Next, select the rows that match the entered content and the content of each item in the keyword management table 1.
00, and the document image data indicated by the image address on the same line is read from the optical disc 10 and displayed on the display section 13.

<OCR記入方法の説明 (第3図(a)、(b)、(c))> 次に、本実施例によるOCR入力シートへのキーワード
記入方法について第3図(a)。
<Explanation of OCR entry method (FIGS. 3(a), (b), (c))> Next, FIG. 3(a) shows the method of entering keywords into the OCR input sheet according to this embodiment.

(b)、(C)を用いて説明する。This will be explained using (b) and (C).

第3図(a)、(b)、(c)は本実施例のキーワード
の記入例を示す図である。尚、本実施例においては区切
り記号を“、”、継続記号を“((”、終了記号を“〃
”とした。またキーワードは4つの項目に分かれている
FIGS. 3(a), 3(b), and 3(c) are diagrams showing examples of entering keywords in this embodiment. In this example, the delimiter is ",", the continuation symbol is "(", and the end symbol is "〃
The keywords are divided into four categories.

第3図(a)は1枚目のOCR入力シートを示し、この
図中の31は地図を示す画像データである。第3図(b
)は2枚目のOCR入力シートを示す。また第3図(c
)は第3図(a)、(b)に示すOCR入力シートより
文字認識した結果を示す。それぞれのシートの上部に記
入された文字列が最上行から順にOCR認識される。1
枚目のOCR入力シートの3行目までが文字認識された
所で、以下の文字列は2枚目のOCR入力シートから読
み込まれ、2枚目のOCR入力シートの1行目で読込み
は終了する。読込まれた文字列は、区切り記号“、”に
従って、第3図(C)で示されるようにキーワード項目
N011〜4のキーワードの各項目に分けられ、これを
磁気ディスク11のキーワード項目テーブルに格納され
る。
FIG. 3(a) shows the first OCR input sheet, and numeral 31 in this figure is image data representing a map. Figure 3 (b
) indicates the second OCR input sheet. Also, Figure 3 (c
) shows the results of character recognition from the OCR input sheets shown in FIGS. 3(a) and 3(b). The character strings written at the top of each sheet are sequentially recognized by OCR starting from the top row. 1
When characters are recognized up to the third line of the second OCR input sheet, the following character strings are read from the second OCR input sheet, and reading ends at the first line of the second OCR input sheet. do. The read character string is divided into keyword items N011 to N04 according to the delimiter "," as shown in FIG. 3(C), and these are stored in the keyword item table of the magnetic disk 11. be done.

また、OCR認識するためにスキャナから読込んだ第3
図(a)の画像データは、そのOCR入力シートに張り
付けられている画像データ31を光ディスクに格納する
時のデータとしても使用する。
In addition, the third image read from the scanner for OCR recognition
The image data in Figure (a) is also used when the image data 31 pasted on the OCR input sheet is stored on the optical disc.

く文字認識処理手順(第4図(a)、(b))>本実施
例における文字認識処理手順を第4図のフローチャート
に従って説明する。尚、ここでは表示部13への文字列
の表示処理やキーボード部12からの信号に基づく処理
においては通常の情報処理装置と同様なため省略し、こ
こではOCR入力シートの文字認識処理のみを説明する
Character Recognition Processing Procedure (FIGS. 4(a), (b))>The character recognition processing procedure in this embodiment will be explained with reference to the flowchart in FIG. 4. Note that the process of displaying character strings on the display unit 13 and the process based on signals from the keyboard unit 12 are omitted here because they are similar to those of a normal information processing device, and only the character recognition process of the OCR input sheet will be explained here. do.

第4図(a)、(b)は本実施例の文字認識処理を示す
フローチャートである。
FIGS. 4(a) and 4(b) are flowcharts showing the character recognition process of this embodiment.

まず、スキャナ5でOCR入力シートから画像データを
読み込み、画像メモリ6に格納する(ステップSl)。
First, the scanner 5 reads image data from an OCR input sheet and stores it in the image memory 6 (step Sl).

ここで、画像データにキーワード以外に、光ディスク1
0にも格納しておく必要のある、例えば第3図(a)、
(b)、(c)で説明した画像データ31が含まれてい
るような場合には、ここで読み取った入力画像データを
格納のためにも利用する事ができる。
Here, in addition to the keyword in the image data, the optical disc 1
For example, in Figure 3(a),
If the image data 31 described in (b) and (c) is included, the input image data read here can also be used for storage.

次に、キーワード項目No、を示す変数にの値を“0”
にクリアし、更にキーワード項目内容を一時的に保存す
る変数χ、(1=0.1.・・・。
Next, set the value of the variable indicating the keyword item No. to “0”.
, and a variable χ, (1=0.1...) that temporarily stores the keyword item contents.

L−1)の内容を“O”にクリアする(ステップS2)
。ここで、各項目の文字数は最高り文字とし、それぞれ
のχ皿には1文字分のデータが保存できるものとする。
Clear the contents of L-1) to “O” (step S2)
. Here, it is assumed that the number of characters in each item is the maximum number of characters, and each χ plate can store data for one character.

そしてOCR入力シート上のどの行をOCR機能により
文字認識するかを示す行の変数iの値を“O”にクリア
する(ステップS3)。ここで、OCR入力シートには
最大1行まで記入できるものとし、行の変数iと最大値
Iとの値を比較し、もしi≧IならばステップS1へ戻
り、もしi<Iならば次の処理に進む(ステップS4)
Then, the value of the variable i of the line indicating which line on the OCR input sheet is to be character-recognized by the OCR function is cleared to "O" (step S3). Here, it is assumed that a maximum of one line can be entered in the OCR input sheet, and the value of the variable i in the line is compared with the maximum value I. If i≧I, the process returns to step S1, and if i<I, the next step is Proceed to the process of (step S4)
.

上記のi<1と判定した後に、OCR入力シートのi行
目を文字認識し、その結果得られたi行目の文字列を変
数U。、ul、・・・、uJ−1とする(ステップS5
)、ここで、−行は最大J文字からなるものとする。そ
して処理する文字の列を示す列の変数jを“O”にクリ
アしくステップS6)、列の変数jと最大値Jとの値を
比較し、もしj≧Jならば行の変数iの値を1つインク
リメントさせ、再びステップS4へ戻る(ステップ86
〜ステツプS8)。これは次の行に処理を移すために行
われる。またステップS7でj<Jならばi行j列目の
変数Ujの値が終了記号“〃”かどうかをキーワード識
別部8で調べ、終了記号と識別したならば全文字認識処
理を終了する(ステップS9)。また変数UJが終了記
号として識別できなかったときには、次に継続記号“〜
(”か否かを更にキーワード識別部8で調べ、この結果
、継続記号と識別したときには再びステップS1へ戻り
(ステップ5IO)、2枚目のOCR入力シートから画
像データを読み込んでステップS2の処理へと続く。こ
のように継続記号“((”を識別したときには、現時点
で文字認識している1枚のOCR入力シート上のキーワ
ードはこれで終了であることを意味し、続くキーワード
は次に読み込むOCR入力シートより調べる。例えば第
3図(a)、(b)、(c)の場合には、第3図(a)
の第1枚目のOCR入力シートでは、継続記号”が記入
されており、キーワード項目No、1〜No、4の“1
987年”までのキーワードを各種記号を識別すること
により認識し、以降のキーワードは第2枚目のOCRシ
ートから調べ、そこで“8月20日“を認識するのであ
る。
After determining that i<1 above, the character in the i-th line of the OCR input sheet is recognized, and the resulting character string in the i-th line is set as a variable U. , ul, ..., uJ-1 (step S5
), where the - line consists of a maximum of J characters. Then, clear the variable j in the column indicating the string of characters to be processed to "O" (step S6), compare the value of the variable j in the column with the maximum value J, and if j≧J, the value of the variable i in the row is incremented by one and returns to step S4 again (step 86
~Step S8). This is done to move processing to the next line. Further, in step S7, if j<J, the keyword identification unit 8 checks whether the value of the variable Uj in the i-th row and the j-th column is the end symbol "〃", and if it is identified as the end symbol, the all character recognition process ends ( Step S9). Also, if variable UJ cannot be identified as a termination symbol, then the continuation symbol “~
The keyword identification unit 8 further checks whether the symbol is a continuation symbol, and as a result, when it is identified as a continuation symbol, the process returns to step S1 (step 5IO), reads the image data from the second OCR input sheet, and processes the step S2. When you identify the continuation symbol “((” in this way, it means that the keyword on the one OCR input sheet that is currently being recognized is the end, and the following keyword is Check from the OCR input sheet to be read. For example, in the case of Figures 3 (a), (b), and (c), Figure 3 (a)
In the first OCR input sheet, the continuation symbol" is entered, and the keyword item No. 1 to No. 4 "1" is entered.
Keywords up to ``987'' are recognized by identifying various symbols, keywords after that are checked from the second OCR sheet, and ``August 20th'' is recognized.

そして、ステップS10において継続記号と識別できな
かったときには、更に変数U、が区切り記号“、”かど
うかを調べる(ステップ511)。もし変数U、が区切
り記号と識別できなかったときには、変数U、の値を変
数χ1に代入しくステップ5L2)、変数χ!の1の値
を1つだけインクリメントさせる(ステップ513)。
If the continuation symbol cannot be identified in step S10, it is further checked whether the variable U is a delimiter "," (step 511). If variable U cannot be identified as a delimiter, the value of variable U is assigned to variable χ1 (step 5L2), and variable χ! The value of 1 is incremented by one (step 513).

さらに変数1と最大値りとの値を比較しくステップ51
4)、この結果がl≧Lのときには、OCR入力シート
上で記入されている文字数がそれぞれのキーワード項目
の文字数として許されている最大文字数りよりも多い誤
ったデータが記入されている場合であり、誤データが発
見されたとして全文字認識処理を終了する。このときエ
ラー表示を行う。またステップS14においてl<Lと
判定されたときには、さらに次のステップに進む。
Furthermore, compare the values of variable 1 and the maximum value in step 51.
4) If the result is l≧L, it means that incorrect data has been entered in which the number of characters entered on the OCR input sheet is greater than the maximum number of characters allowed for each keyword item. Yes, all character recognition processing is terminated as erroneous data has been found. At this time, an error is displayed. Further, when it is determined in step S14 that l<L, the process proceeds to the next step.

ところで、ステップSllにおいて、変数U。By the way, in step Sll, the variable U.

が区切り記号と識別されたときには、一つのキーワード
項目の情報を取得できたことになり、変数χ0.χ8.
・・・χl−1をキーワードの変数節に+1番目(k=
o、1、・・・k−1)の項目としてキーワード管理テ
ーブル100に登録する(ステップ515)。例えば上
述した第3図(a)。
is identified as a delimiter, it means that information for one keyword item has been obtained, and the variable χ0. χ8.
...Add χl-1 to the variable clause of the keyword +1st (k=
o, 1, . . . k-1) in the keyword management table 100 (step 515). For example, FIG. 3(a) mentioned above.

(b)、(Cンの例を用いると、第3図(a)。(b), (Using the example of C, FIG. 3(a).

(b)に示す2枚のOCR入力シートからは区切り記号
“、”を3回文字認識するので、キーワード項目No、
1−No、4は4つに分けてキーワード管理テーブル1
00に登録される。
From the two OCR input sheets shown in (b), the delimiter “,” is recognized three times, so the keyword item No.
1-No, 4 is divided into four keyword management table 1
Registered as 00.

次に、変数にの値を1つだけインクリメントしくステッ
プ516)、変数にとキーワード項目の最大許容設定数
にとの値を比較しくステップ517)、もしに≧にであ
ればOCR入力シートで記入されているキーワード項目
数がキーワード管理テーブル100のキーワード項目の
最大許容設定数により多い、矛盾したデータが入力され
たことになり、以降の処理を中断してエラー表示を行う
。またステップS17においてk<Kであれば(ステッ
プS 17) 、変数χ0.χ1.・・・。
Next, increment the value of the variable by one (Step 516), compare the value of the variable with the maximum allowable setting number of keyword items (Step 517), and if ≧, enter it in the OCR input sheet. This means that contradictory data has been input in which the number of keyword items is greater than the maximum permissible number of keyword items in the keyword management table 100, and subsequent processing is interrupted and an error message is displayed. Further, if k<K in step S17 (step S17), the variable χ0. χ1. ....

χし−1をクリアする(ステップ518)。Clear χ and -1 (step 518).

このようにして、ステップS14或はステップS18の
処理を終えると、次に変数UJの列jの値を1つだけイ
ンクリメントしくステップ519)、再びステップS7
に戻り上記処理を繰り返す。
In this way, when the process of step S14 or step S18 is finished, the value of column j of variable UJ is incremented by one (step 519), and step S7 is performed again.
Return to and repeat the above process.

以上の説明により本実施例によれば、OCR入力シート
から文字認識した文字列より、区切り記号、継続記号、
終了記号を識別し、それに従って文字列をキーワードの
各項目に登録することで以下の効果がある。
As described above, according to this embodiment, delimiters, continuation symbols,
By identifying the ending symbol and registering the character string in each keyword item accordingly, the following effects can be achieved.

(1)OCR入力シート上に各キーワード項目に登録す
る文字列を記入する際、必要な文字列のみを詰めて記入
し、余分な余白をあける必要がないため、OCR入力シ
ートの紙面を有効に使うことができる。
(1) When entering the character strings to be registered in each keyword item on the OCR input sheet, only the necessary character strings are filled in and there is no need to leave extra space, making the space of the OCR input sheet effective. You can use it.

(2)必要に応じて何枚かのOCR入力シートに分けて
キーワードを記入することができるので、OCR入力シ
ートに他のデータが記入されて(例えば第3図(a)の
如く、そのキーワードに対応する画像データがはりつけ
られている)キーワードが記入できる行がキーワードの
全文字を記入するだけのスペースを有しない場合でも、
記入できなかった文字を別のOCR入力シートに分けて
記入することができるので、OCR入力シートの用途を
拡張することができる。
(2) Keywords can be entered on several OCR input sheets as needed, so other data can be entered on the OCR input sheet (for example, as shown in Figure 3 (a), the keyword (the image data corresponding to the keyword is pasted) Even if the line where you can write the keyword does not have enough space to write all the characters of the keyword,
Since characters that could not be entered can be entered separately on another OCR input sheet, the uses of the OCR input sheet can be expanded.

(3)キーワードの項目の構造に応じた特定の形式のO
CR入力シートを用意する必要がなく、単に文字列が記
入できるOCR入力シートを用いるだけで、任意の項目
構造をもつキーワードの入力データが記入できる。
(3) O in a specific format depending on the structure of the keyword item
There is no need to prepare a CR input sheet, and input data of keywords having an arbitrary item structure can be entered simply by using an OCR input sheet in which character strings can be entered.

ところで、本実施例では、OCR入力シートをスキャナ
から読み込んだが、本発明を実施する上ではOCR入力
シートの読込画像データが得られれば充分であり、入力
デバイスは必ずしもスキャナに限定する必要はない。こ
の場合、例えば、通信媒体のファクシミリで受信した画
像であっても本発明は適用されることは十分に可能であ
り、ここでも上述の実施例と同様にOCR入力シートに
記入されたキーワードの文字認識許容範囲を拡張させる
ことができる。
By the way, in this embodiment, the OCR input sheet was read from a scanner, but in carrying out the present invention, it is sufficient to obtain the read image data of the OCR input sheet, and the input device does not necessarily need to be limited to a scanner. In this case, for example, the present invention can be applied even to images received by facsimile communication medium, and here as well, the keyword characters entered on the OCR input sheet are The recognition tolerance range can be expanded.

また、本実施例において、OCR認識により入力される
文字列は、必ずしもキーワードである必要はない。例え
ば、コメント等の文字列を、適当に幾枚かのシートに分
けて記入し、それを入力する場合でも、本発明は有効に
適用できる。
Furthermore, in this embodiment, the character string input by OCR recognition does not necessarily have to be a keyword. For example, the present invention can be effectively applied even when a character string such as a comment is appropriately divided into several sheets and inputted.

または、本実施例における区切り記号、継続記号、終了
記号の識別は、必ずしもそれに対応する記号を定義する
必要はなく、スキャナより文字列を読込んだとき、確実
に区切り状態、継続状態。
Alternatively, in order to identify the delimiter, continuation symbol, and end symbol in this embodiment, it is not necessary to define the corresponding symbol, and when a character string is read by the scanner, the delimiter state and continuation state are reliably identified.

終了状態の3つの状態が識別できれば良い。この3つの
状態を文字で示す以外の方法としては、例えば、各キー
ワード項目は、1行に書く事にし、行末を区切りにする
方法や1つの用紙には3行しか記入しない事にし、3行
目まで読込んでも、まだ文字列がセットされていないキ
ーワード項目が残っている場合は、継続状態とする方法
があり、さらには1行が総て空白であるとき、その行末
を終了状態とする方法等がある。
It is sufficient if the three end states can be identified. Other ways to indicate these three states with letters include, for example, writing each keyword item on one line and using the end of the line as a separator, or writing only three lines on one sheet of paper, and writing three lines. If there are still keyword items for which no strings have been set even after reading up to the first line, there is a way to make it a continuous state.Furthermore, if a line is entirely blank, the end of that line is set to an end state. There are methods etc.

[発明の効果] 以上の説明により本発明によれば、文字認識シートに特
定な文字列の記入規則を与えることで、シート面を無駄
なく、効率的に使用できることは勿論、文字認識シート
の用途を拡張した情報処理装置を提供してくれる。
[Effects of the Invention] As described above, according to the present invention, by giving a character recognition sheet a specific character string entry rule, the sheet surface can be used efficiently without wasting it, and the use of the character recognition sheet can be improved. It provides an information processing device that has expanded.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は本発明の一実施例の構成を示す概略ブロック図
、 第2図は本実施例による磁気ディスク11のキーワード
管理テーブルを示す図、 第3図(a)、(b)、(c)は本実施例のキーワード
の記入例を示す図、 第4図(a)、(b)は本実施例の文字認識処理を示す
フローチャートである。 図中、1・・・情報処理装置、2・・・CPU、3・・
・ROM、4・・・RAM、5・・・スキャナ、6・・
・画像メモリ、7・・・文字認識部、8・・・キーワー
ド識別部、9・・・文字列編集部、10・・・光ディス
ク、11・・・磁気ディスク、12・・・キーボード部
、13・・・表示部、14・・・システムバス、31・
・・画像データ、100・・・キーワード項目テーブル
、101・・・画像アドレス、102・・・第1キーワ
ード項目、103・・・第2キーワード項目、104・
・・第3キーワード項目、105・・・第4キーワード
項目である。 、11 / 第2図 (C) 第3図
FIG. 1 is a schematic block diagram showing the configuration of an embodiment of the present invention, FIG. 2 is a diagram showing a keyword management table of the magnetic disk 11 according to the embodiment, and FIGS. ) is a diagram showing an example of entering keywords in this embodiment, and FIGS. 4(a) and 4(b) are flowcharts showing character recognition processing in this embodiment. In the figure, 1... information processing device, 2... CPU, 3...
・ROM, 4...RAM, 5...Scanner, 6...
- Image memory, 7... Character recognition section, 8... Keyword identification section, 9... Character string editing section, 10... Optical disk, 11... Magnetic disk, 12... Keyboard section, 13 ...Display section, 14...System bus, 31.
...Image data, 100...Keyword item table, 101...Image address, 102...First keyword item, 103...Second keyword item, 104...
...Third keyword item, 105...Fourth keyword item. , 11 / Figure 2 (C) Figure 3

Claims (1)

【特許請求の範囲】[Claims] 文書情報を特定な文字列から成る付属情報と対応づけて
保存する情報処理装置において、前記文書情報を入力す
る入力手段と、該入力手段の入力した文書情報より規則
的に並ぶ特定な文字列を識別する識別手段と、前記識別
した特定な文字列を前記文書情報の付属情報として管理
する管理手段とを備えることを特徴とする情報処理装置
An information processing device that stores document information in association with attached information consisting of a specific character string, which includes an input means for inputting the document information, and a specific character string arranged regularly from the document information inputted by the input means. An information processing apparatus comprising: an identification means for identifying; and a management means for managing the identified specific character string as attached information of the document information.
JP63114173A 1988-05-11 1988-05-11 information processing equipment Pending JPH01283671A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP63114173A JPH01283671A (en) 1988-05-11 1988-05-11 information processing equipment

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP63114173A JPH01283671A (en) 1988-05-11 1988-05-11 information processing equipment

Publications (1)

Publication Number Publication Date
JPH01283671A true JPH01283671A (en) 1989-11-15

Family

ID=14631006

Family Applications (1)

Application Number Title Priority Date Filing Date
JP63114173A Pending JPH01283671A (en) 1988-05-11 1988-05-11 information processing equipment

Country Status (1)

Country Link
JP (1) JPH01283671A (en)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH0470967A (en) * 1990-07-05 1992-03-05 Canon Inc Picture retrieving device
JPH0498363A (en) * 1990-08-10 1992-03-31 Pfu Ltd Knowledge processing system for continuous field

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH0470967A (en) * 1990-07-05 1992-03-05 Canon Inc Picture retrieving device
JPH0498363A (en) * 1990-08-10 1992-03-31 Pfu Ltd Knowledge processing system for continuous field

Similar Documents

Publication Publication Date Title
EP0285449B1 (en) Document processing system
JPH03161873A (en) Electronic filing device with database construction function
JP2005018678A (en) Form data input processing device, form data input processing method and program
CN114692603A (en) Sensitive data identification method, system, device and medium based on CRF
US5680630A (en) Computer-aided data input system
JP2002024761A (en) Image processing apparatus, image processing method, and storage medium
JP2932667B2 (en) Information retrieval method and information storage device
JPS62257568A (en) Line boundary character selecting system for text processor
JP2784004B2 (en) Character recognition device
JPS6074094A (en) Character recognizing device
JPH03263182A (en) Electronic filing system
JP3355289B2 (en) Automatic proofing method and apparatus for character data
JP4462508B2 (en) Information processing apparatus and definition information generation method
JP2634926B2 (en) Kana-Kanji conversion device
JPH06251187A (en) Method and device for correcting character recognition error
JP2623292B2 (en) How to create dictionary data
JPH02114390A (en) Symbol input system for drawing
JP2624124B2 (en) Character recognition device and character recognition method
JPH0554180A (en) Slop format defining system for optical character reader
JPH01283670A (en) Information processor
JPH03142663A (en) Document file device
JPH02206893A (en) Character reader
JPS6326789A (en) Character recognizing device
JPS62160534A (en) String matching method
JPS5851390A (en) Font character recognizing device