JPH04111186A - Character recognition result correction method for address character string - Google Patents

Character recognition result correction method for address character string

Info

Publication number
JPH04111186A
JPH04111186A JP2230927A JP23092790A JPH04111186A JP H04111186 A JPH04111186 A JP H04111186A JP 2230927 A JP2230927 A JP 2230927A JP 23092790 A JP23092790 A JP 23092790A JP H04111186 A JPH04111186 A JP H04111186A
Authority
JP
Japan
Prior art keywords
place name
character
word
words
address
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP2230927A
Other languages
Japanese (ja)
Inventor
Hideyuki Isoyama
磯山 秀幸
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
N T T DATA TSUSHIN KK
NTT Data Group Corp
Original Assignee
N T T DATA TSUSHIN KK
NTT Data Communications Systems Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by N T T DATA TSUSHIN KK, NTT Data Communications Systems Corp filed Critical N T T DATA TSUSHIN KK
Priority to JP2230927A priority Critical patent/JPH04111186A/en
Publication of JPH04111186A publication Critical patent/JPH04111186A/en
Pending legal-status Critical Current

Links

Landscapes

  • Character Discrimination (AREA)

Abstract

(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。
(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.

Description

【発明の詳細な説明】 [産業上の利用分野] 本発明は、光学式文字読み取り装置(OCR:opii
cal character reader)等を用い
た文字認識方法に関し、特に売上伝票、配送伝票等に記
入される住所の文字認識結果について認識誤りを修正す
る住所文字列に対する文字認識結果修正方法に関する。
[Detailed Description of the Invention] [Industrial Application Field] The present invention is directed to an optical character reader (OCR: opii
The present invention relates to a character recognition method using a computer (cal character reader), etc., and particularly to a method for correcting character recognition results for address character strings to correct recognition errors in character recognition results for addresses written on sales slips, delivery slips, etc.

[従来の技術] 従来、OCRを用いた文字認識システムでは、紙の上の
光学像を電気信号に変換し、整形処理を行った後、特徴
抽出を行い、構造解析法やパターンマツチングによって
個々の文字パターンを認識している。
[Conventional technology] Conventionally, character recognition systems using OCR convert an optical image on paper into an electrical signal, perform shaping processing, extract features, and identify individual characters using structural analysis methods or pattern matching. Recognizes character patterns.

また、住所文字列を含む売上伝票、配送伝票等の帳票を
文字認識する際、文字認識結果と地名単語辞書との照合
を行い、確からしい単語を選び出す処理では、文字認識
結果の候補文字を組み合わせて文字列を作成し、これに
一致する単語を地名単語辞書から検索している。
In addition, when character recognition is performed on sales slips, delivery slips, and other forms that contain address character strings, the character recognition results are checked against a place name word dictionary, and in the process of selecting likely words, candidate characters from the character recognition results are combined. A character string is created using a string, and a word matching this string is searched from a place name word dictionary.

なお、完全に一致する単語がない場合には、部分的に一
致した候補文字にその候補順位に対して点数をつけて最
高点をあげたものを正解単語とするという方法がとられ
ている。
Note that if there is no completely matching word, a method is used in which points are assigned to partially matching candidate characters according to their candidate rankings, and the one with the highest score is determined as the correct word.

この種の方法については、例えば″自由記載住所文字列
に対する知識処理、清野他、電子情報通信学会春季全国
大会D−465,pp、6−185.1989”、OC
Rにおける住所データ読取りについて、情報処理学会第
34目金国大会4E−5,pp、 1841.1987
″等において論じられている。
For this type of method, see, for example, "Knowledge Processing for Free-Text Address Strings," Seino et al., IEICE Spring National Conference D-465, pp. 6-185.1989, OC
About reading address data in R, Information Processing Society of Japan 34th Gold Country Conference 4E-5, pp. 1841.1987
”, etc.

[発明が解決しようとする課題] 上記従来技術では、文字認識装置の認識結果に対して、
地名辞書との照合を行い、認識誤りを修正することは可
能であった。
[Problem to be solved by the invention] In the above-mentioned conventional technology, the recognition result of the character recognition device is
It was possible to correct recognition errors by checking with a place name dictionary.

しかし、】つの文字パターンに対して複数の候補文字を
持つ認識結果を組合せるため、組合せ数が増大し、さら
ζ二番々の生成文字列に対して文字列ごとの一致文字数
を調べるので、文字単位でマツチングをとる回数が非常
に多くなり、処理時間が増大する。逆に、処理時間を短
縮するため、候補文字を減らして組合せ数を少くすると
、正解候補文字が処理対象から除外されてしまうことが
ある。
However, since recognition results with multiple candidate characters are combined for one character pattern, the number of combinations increases, and the number of matching characters for each character string is checked for ζ second generated character strings. The number of times that character-by-character matching is performed becomes extremely large, increasing processing time. Conversely, if the number of candidate characters is reduced to reduce the number of combinations in order to shorten processing time, correct candidate characters may be excluded from the processing target.

また、類似度合の判定に候補順位を用いているため、同
一順位にある候補文字は全て一律に同じ確からしさであ
るとされていた。すなわち、文字認識結果の確からしぎ
を表す値である距lIi値が異っていても、この相違を
後処理に反映することができなかった。
In addition, since the candidate ranking is used to determine the degree of similarity, all candidate characters in the same ranking are uniformly considered to have the same probability. That is, even if the distance lIi value, which is a value representing the certainty of the character recognition result, differs, this difference cannot be reflected in post-processing.

本発明の目的は、このような問題点を改善し、処理時間
を短縮させることができ、また従来の方法では活用され
ていなかった文字認識過程での情報を有効に利用して、
処理精度を向上させることが可能な住所文字列に対する
文字認識結果修正方法を提供することにある。
The purpose of the present invention is to improve such problems, shorten processing time, and effectively utilize information in the character recognition process that was not utilized in conventional methods.
An object of the present invention is to provide a method for correcting character recognition results for address character strings, which can improve processing accuracy.

[課題を解決するための手段] 上記目的を達成するため、本発明の住所文字列に対する
文字認識結果修正方法は、記入された住所文字列の認識
結果の判定および誤り訂正を行う文字認識システムにお
いて、住所文字列を構成する地名単語の階層関係が表現
された地名単語テーブルと、各階層ごとに各々の単語の
文字位置における文字コードでソートされたインデクス
テーブルとを有し、文字認識装置から出力された住所文
字列の認識結果に対して、そのインデクステーブルを用
い、カラムごとに一致する文字を検索して、その文字が
存在する地名単語の地名単語テーブルにおけるレコード
番号を求めることにより、地名単語テーブルの中から類
似性のある地名単語を抽出する処理を高速化して、抽出
された地名単語群の中から、地名単語テーブルにより、
先に決定された上位階層の地名単語に正しく接続できる
地名単語を選出して、選出された地名単語に対し、その
文字ごとの距離値を基に計算された評価値によって、最
も類似度の高い地名単語を決定することに特徴がある。
[Means for Solving the Problems] In order to achieve the above object, the method for correcting character recognition results for address character strings of the present invention is provided in a character recognition system that determines recognition results of written address character strings and corrects errors. , has a place name word table that expresses the hierarchical relationship of place name words that make up an address string, and an index table that is sorted by character code at the character position of each word for each hierarchy, and is output from a character recognition device. Using the index table for the recognition result of the address string, search for matching characters in each column and find the record number in the place name word table of the place name word where that character exists. By speeding up the process of extracting similar place name words from the table, from the extracted place name word group, using the place name word table,
We select place name words that can be correctly connected to the place name words in the upper hierarchy determined earlier, and then select the word with the highest degree of similarity to the selected place name word based on the evaluation value calculated based on the distance value for each character. It is distinctive in determining place name words.

〔作用3 本発明においては、文字認識結果の各候補文字について
、第1カラムから順にインデクステーブルを検索し、そ
の文字が存在する地名単語の地名単語テーブルにおける
レコード番号を求める。この際、求められたレコード番
号の中で数多く現われたものが類似性のある地名単語を
表わすレコード番号である。この処理では、レコード番
号を抽出するのは1文字だけの比較であり、また、一致
した文字数を求めて類似性のある地名単語を選ぶときに
は、レコード番号のみを比較すればよい。
[Operation 3] In the present invention, for each candidate character in the character recognition result, the index table is searched in order from the first column, and the record number in the place name word table of the place name word in which that character exists is determined. At this time, the record numbers that appear many times among the obtained record numbers are record numbers representing similar place name words. In this process, record numbers are extracted by comparing only one character, and when finding the number of matching characters and selecting similar place name words, only record numbers need to be compared.

また、地名単語テーブルには、各地名単語の上位レベル
となる地名単語のレコード番号が格納されているため、
これを用いて先に決定された上位レベルの地名に正しく
接続することができるかを判定することができる。
In addition, the place name word table stores the record numbers of place name words, which are the upper level of each place name word.
Using this, it is possible to determine whether it is possible to correctly connect to the previously determined higher level place name.

さらに、上位レベルとの接続判定が終了した後の評価値
算出処理では、候補単語の確がらしさを表わす評価値と
して1文字認識結果の各々の候補文字の確からしさを表
わす値である距離値(文字パターンの適合度合い逆数)
を用いる。
Furthermore, in the evaluation value calculation process after the connection determination with the upper level is completed, a distance value ( Reciprocal of degree of conformance of character pattern)
Use.

これによって、処理時間を短縮し、かつ文字認識結果を
修正する精度を高めることができる。
This makes it possible to shorten the processing time and increase the precision with which character recognition results are corrected.

f実施例〕 以下、本発明の一実施例を図面により説明する。f Example] An embodiment of the present invention will be described below with reference to the drawings.

第2図は、本発明の一実施例における地名単語テーブル
の説明図、第3図は本発明の一実施例における都道府県
レベルの第1カラムおよび第2カラムに対するインデク
ステーブルの説明図である。
FIG. 2 is an explanatory diagram of a place name word table in one embodiment of the present invention, and FIG. 3 is an explanatory diagram of an index table for the first and second columns at the prefecture level in one embodiment of the present invention.

本実施例の文字認識システムは、○CR等の文字読み取
り装置、システム全体を制御するCPU、各種プログラ
ムおよび必要データを格納するためのメモリ、CRT等
の表示装置、キーボード等の入力装置、およびインデク
ステーブルや地名単語テーブル等を格納するための外部
記憶装置から構成される。
The character recognition system of this embodiment includes a character reading device such as ○CR, a CPU that controls the entire system, a memory for storing various programs and necessary data, a display device such as a CRT, an input device such as a keyboard, and an index. It consists of an external storage device for storing tables, place name word tables, etc.

この文字読み取り装置は、住所文字列を含む売上伝票や
配送伝票等の帳票が挿入されると、各文字を認識し、電
気信号に変換してCPIJに送る。
When a sales slip, delivery slip, or other form containing an address string is inserted, this character reading device recognizes each character, converts it into an electrical signal, and sends it to CPIJ.

また、CPUは、認識結果をもとに、メモリ内のインデ
クステーブルを検索し、これによって得られたレコード
番号によって、メモリ内の地名単語テーブルから地名単
語を抽出し、住所文字列認識結果を修正する。
In addition, the CPU searches the index table in memory based on the recognition results, extracts place name words from the place name word table in memory using the record number obtained thereby, and corrects the address string recognition results. do.

また、地名単語テーブルは、住所文字列を構成する地名
単語の階層関係を表現したものである。
Further, the place name word table expresses the hierarchical relationship of place name words that make up the address character string.

具体的には第2図のように示される。第2図の都道府県
テーブルにおいて、テーブルの左側の数字はそのレコー
ド自身のレコード番号であり、地名の右側に格納されて
いるのは、上位レコードを示すレコード番号である。例
えば、″京都布′”や“舞鶴布”は“京都府″″に含ま
れ(“京都府′”の下層にあり)、それらの上位レベル
の地名革語が格納されたレコードを示すレコード番号は
同じく“10″である。
Specifically, it is shown as shown in FIG. In the prefecture table shown in FIG. 2, the number on the left side of the table is the record number of the record itself, and the record number stored on the right side of the place name is the record number indicating the upper record. For example, ``Kyotofu'' and ``Maizurufu'' are included in ``Kyotofu'' (located in the lower level of ``Kyotofu''), and the record number indicates the record in which these higher-level place name revolution words are stored. is also "10".

また、インデクステーブルは、階層関係を持つ地名単語
の各階層ごとに、各々の単語の文字位置における文字コ
ードでソートされたものである。
In addition, the index table is sorted by the character code at the character position of each word for each layer of place name words that have a hierarchical relationship.

具体的には第3図のように示され、各文字位置ごとに、
候補文字、および候補文字と地名単語との関係を示すレ
コード番号から構成される。
Specifically, it is shown in Figure 3, and for each character position,
It consists of candidate characters and record numbers indicating the relationship between candidate characters and place name words.

次に、本実施例における住所文字列認識結果の修正方法
について述べる。
Next, a method for correcting address character string recognition results in this embodiment will be described.

第1図は、本発明の一実施例における文字認識システム
の動作フローチャート、第4図は本発明の一実施流側に
おける文字認識結果の候補文字と距離値を示す説明図、
第5図は本発明の一実施例において抽出されたレコード
番号を示す説明図である。
FIG. 1 is an operation flowchart of a character recognition system according to an embodiment of the present invention, and FIG. 4 is an explanatory diagram showing candidate characters and distance values of character recognition results in an embodiment of the present invention.
FIG. 5 is an explanatory diagram showing record numbers extracted in one embodiment of the present invention.

本実施例では、帳票に「京都府京都市下京区葛籠屋町」
と記入した場合について述べる。
In this example, the form is "Katsukagoyacho, Shimogyo Ward, Kyoto City, Kyoto Prefecture".
Let's talk about the case where this is entered.

第1図のように、文字読み取り装置に入力された帳票に
記されである住所「京都府京都市下京区葛籠屋町Jの各
文字を認識し、第4図に示す認識結果を得ると(101
)、これに対する修正を行う。
As shown in Figure 1, each character of the address "Katsukagoya-cho J, Shimogyo-ku, Kyoto City, Kyoto Prefecture" written on the form entered into the character reading device is recognized, and the recognition result shown in Figure 4 is obtained. 101
), make a fix for this.

まず、住所文字列の文字認識結果から、第3図に示した
インデクステーブルを参照して各カラム毎に一致する文
字を検索する(102)。
First, based on the character recognition result of the address character string, matching characters are searched for each column by referring to the index table shown in FIG. 3 (102).

すなわち、第4図の第1カラム目に対する認識結果の第
1位候補である“東”をキーとして、都道府県レベルの
地名単語の第1文字目でソートされたインデクステーブ
ルを、バイナリサーチアルゴリズムを用いて検索し、第
1文字目がパ東”である地名単語のレコード番号を得る
。次に第2位候補である゛京″′で同様に第1文字目が
゛京″である地名単語のレコード番号を得る。以下、候
補順位の最後までこの処理を行い、第1カラム目に対す
る処理を終える。第2カラム百以後についても上記の処
理を行い、都道府県レベルの最大文字列長となる第4カ
ラム目まで続ける。なお、第3図において、(a)は第
1カラムを示し、(b)は第2カラムを示す。
In other words, the index table sorted by the first character of place name words at the prefecture level is run through a binary search algorithm, using "East", which is the first candidate in the recognition results for the first column in Figure 4, as a key. Search using ``Pato'' as the first character to obtain the record number of a place name word whose first character is ``Pato''.Next, search for the second place candidate ゛ky''', a place name word whose first character is ``kyo'' Obtain the record number. From here on, this process is performed until the end of the candidate ranking, and the process for the first column is completed. The above process is also performed for the second column 100 and beyond, resulting in the maximum character string length at the prefecture level. Continue up to the fourth column. In Fig. 3, (a) shows the first column, and (b) shows the second column.

こうして、第5図に示すように地名単語のレコード番号
群を得る。
In this way, a record number group of place name words is obtained as shown in FIG.

次に、得られたレコード番号をソートして(103)、
出現頻度の高いレコード番号を選び(104)、地名単
語テーブルを参照して上位レベル単語との接続を調べる
(l○5)。
Next, the obtained record numbers are sorted (103),
A record number with a high frequency of appearance is selected (104), and the connection with higher level words is checked by referring to the place name word table (l○5).

すなわち、第5図のレコード番号群をマージソートし、
単語長の半数回以上現われたレコード番号は類似性があ
ると判断して、これらを抽出する。
That is, merge-sort the record numbers in Figure 5,
Record numbers that appear more than half of the word length are judged to have similarity and are extracted.

この際、抽出されるレコード番号は、3文字一致したr
l OJ、r40Jと2文字一致1.たr31Jである
At this time, the extracted record number is r
l OJ, r40J and 2 characters match 1. It is r31J.

なお、処理中の住所文字列において、先に決定された上
位レベルの地名単語がある場合、住所文字列における上
位レベルと下位レベルの地名の階層構造を利用して、候
補数を絞り込む。すなわち、第2図に示したように、地
名単語テーブルに格納された上位レベルのレコード番号
と、先に決定された地名単語のレコード番号が一致する
かを調べることで、上位レベルとの接続判定を行う。但
し、上位レベルの地名単語がない場合は、接続判定処理
は省略されるため、都道府県レベル(京都府)の場合は
行われない。
Note that if there is a previously determined upper-level place name word in the address string being processed, the number of candidates is narrowed down using the hierarchical structure of the upper-level and lower-level place names in the address string. In other words, as shown in Figure 2, the connection with the upper level is determined by checking whether the record number of the upper level stored in the place name word table matches the record number of the previously determined place name word. I do. However, if there is no higher-level place name word, the connection determination process is omitted, so it is not performed in the case of the prefecture level (Kyoto Prefecture).

次に、こうして得られた地名単語の中で、何れが最も確
からしいかを判定するための評価値計算を行い、認識結
果と地名単語テーブルとの照合において、「一致した文
字数が最大のものの中で評価値が最小であるもの」を、
最も確からしいと判断する(106)。
Next, among the place name words obtained in this way, an evaluation value is calculated to determine which one is the most likely, and when comparing the recognition results with the place name word table, it is determined that the The one with the lowest evaluation value in
It is determined that it is the most probable (106).

この評価値計算では、各候補文字の確からしさを表わす
値(距離値)を用い、次に示すように評価値を算出する
In this evaluation value calculation, a value (distance value) representing the probability of each candidate character is used to calculate the evaluation value as shown below.

(イン距NIM=第n位候補の距離値−第1位候補の距
離値 (ロ)評価値=一致した文字の距離差の和但し、距離値
とは、文字パターンの適合度合い逆数であり、値の大き
くなるほど、確からしさは低くなる。すなわち、文字認
識の候補文字は、距離値に関して昇順に並んでいる。ま
た、第1位候補の距離差はOである。
(In-distance NIM = distance value of the nth candidate - distance value of the first candidate (b)) Evaluation value = sum of distance differences between matched characters. However, the distance value is the reciprocal of the degree of matching of the character pattern, The larger the value, the lower the certainty. That is, the candidate characters for character recognition are arranged in ascending order with respect to the distance value. Also, the distance difference of the first candidate is O.

本実施例では、ステップ104において、3文字一致し
た単語として「101京都府」、「40:奈良系」、2
文字一致した単語として「31:大阪府」が抽出されて
いる。
In this embodiment, in step 104, the words that match three characters are "101 Kyoto Prefecture", "40: Nara-kei", and 2.
“31: Osaka Prefecture” is extracted as a word with character matching.

各々の評価値を計算すると以下のようになる。The calculation of each evaluation value is as follows.

京       都       府 (2994−1454)+(1632−1632)+(
3861−1495)・3906奈       良 
      県 (6534−1454)+(7962−1632)+(
6305−1495)=1622゜大       阪
       府 (6509−1454)         +(386
1−1495)=7421これらの中で、「一致した文
字数が最大のものの中で評価値が最小であるもの」の順
に正解候補の順位をつけると、確からしいほうから「京
都府」、「奈良系」、「大阪府」の順になる。よって、
本実施例の都道府県レベルでは、「京都府」が第1位の
正解地名単語となる。なお、求められる評価値がシステ
ム規定の閾値を超えてしまった場合は、候補単語からは
除外する。例えば、仮りに、「京都府」、「奈良系」の
評価値が閾値を超えてしまい、「大阪府」の評価値だけ
が閾値内にあれば、「大阪府」が正解単語となる。以上
で都道府県レベルの照合が終了したことになる。また、
本実施例では、都道府県レベルの地名は3文字で決定さ
れたため、次の市区郡レベルの照合は認識結果の第4カ
ラム目から行う。
Kyoto Prefecture (2994-1454) + (1632-1632) + (
3861-1495)・3906 Nara
Prefecture (6534-1454) + (7962-1632) + (
6305-1495) = 1622° Osaka Prefecture (6509-1454) + (386
1-1495) = 7421 Among these, if we rank the correct answer candidates in the order of ``those with the largest number of matching characters and the smallest evaluation value'', we will choose ``Kyoto Prefecture'', ``Nara Prefecture'' from the most likely. ``Kei'', then ``Osaka Prefecture''. Therefore,
At the prefecture level in this embodiment, "Kyoto Prefecture" is the first correct place name word. Note that if the required evaluation value exceeds a system-defined threshold, it is excluded from the candidate words. For example, if the evaluation values of "Kyoto Prefecture" and "Nara-kei" exceed the threshold value, and only the evaluation value of "Osaka Prefecture" is within the threshold value, "Osaka Prefecture" becomes the correct word. This concludes the prefecture-level verification. Also,
In this embodiment, since place names at the prefecture level are determined using three characters, the next comparison at the city, ward, and county level is performed starting from the fourth column of the recognition results.

さらに、上記の方法を繰り返すことにより、正解住所文
字列を得て、住所文字列の文字認識結果の修正処理を完
了する(107)。
Furthermore, by repeating the above method, a correct address string is obtained, and the correction process for the character recognition result of the address string is completed (107).

[発明の効果] 本発明によれば、住所文字列の文字認識結果の誤りを修
正する方法として、文字位置ごとのインデクステーブル
を用いた文字列照合を行うことにより、処理時間を短縮
することができ、また文字認識装置から出力された住所
文字列の認識結果を、階層構造からなる地名単語辞書を
用いて、上位レベル単語との接続関係を利用した判定を
行い、さらに距離値を用いて、候補単語の確からしさを
計算しているため、従来の方法では活用されていなかっ
た文字認識過程での情報を有効に利用し、処理精度を向
上させることができる。
[Effects of the Invention] According to the present invention, as a method for correcting errors in character recognition results of address character strings, processing time can be shortened by performing character string matching using an index table for each character position. In addition, the recognition result of the address character string output from the character recognition device is judged using the connection relationship with higher level words using a hierarchically structured place name word dictionary, and further using the distance value, Since the probability of candidate words is calculated, it is possible to effectively utilize information from the character recognition process, which is not used in conventional methods, and improve processing accuracy.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は本発明の一実施例における文字認識システムの
動作フローチャート、第2図は本発明の一実施例におけ
る地名単語テーブルの説明図、第3図は本発明の一実施
例における都道府県レベルの第1カラムおよび第2カラ
ムに対するインデクステーブルの説明図、第4図は本発
明の一実施流側における文字認識結果の候補文字と距離
値を示す説明図、第5図は本発明の一実施例において抽
出されたレコード番号を示す説明図である。 第 図 第 図 脈 荘 城 脈 鉦 薇 賊 脈 綜 脈
Fig. 1 is an operation flowchart of a character recognition system in an embodiment of the present invention, Fig. 2 is an explanatory diagram of a place name word table in an embodiment of the present invention, and Fig. 3 is a prefecture-level diagram in an embodiment of the present invention. FIG. 4 is an explanatory diagram showing candidate characters and distance values of character recognition results in one embodiment of the present invention, and FIG. 5 is an explanatory diagram of the index table for the first and second columns of the present invention. It is an explanatory diagram showing record numbers extracted in an example. Diagram Diagram Diagram Diagram Zhuang Castle Vein Dragon Bandit Vein

Claims (1)

【特許請求の範囲】[Claims] (1)記入された住所文字列の認識結果の判定および誤
り訂正を行う文字認識システムの文字認識結果修正方法
において、住所文字列を構成する地名単語の階層関係が
表現された地名単語テーブルと、階層ごとに各単語の文
字位置における文字コードでソートされたインデクステ
ーブルとを有し、該インデクステーブルを参照して、カ
ラムごとに一致する文字を検索し、該文字が存在する地
名単語の該地名単語テーブルにおけるレコード番号を求
め、該地名単語テーブルを参照して、該レコード番号に
より、類似性のある地名単語を抽出し、抽出された地名
単語群の中から、先に決定された上位階層の地名単語に
正しく接続できる地名単語を選出し、選出された地名単
語に対して、文字ごとの距離値を基に計算した評価値に
より、最も類似度の高い地名単語を決定することを特徴
とする住所文字列に対する文字認識結果修正方法。
(1) In a character recognition result correction method for a character recognition system that determines the recognition result of a written address character string and corrects errors, a place name word table that expresses the hierarchical relationship of place name words that make up the address character string; Each hierarchy has an index table sorted by character code at the character position of each word, and by referring to the index table, searching for matching characters in each column, and searching for the place name of the place name word where the character exists. Find the record number in the word table, refer to the place name word table, extract similar place name words based on the record number, and from the extracted place name word group, select the place name words in the upper hierarchy determined earlier. It is characterized by selecting place name words that can be correctly connected to place name words, and determining the place name word with the highest degree of similarity to the selected place name words based on an evaluation value calculated based on the distance value for each character. How to correct character recognition results for address strings.
JP2230927A 1990-08-31 1990-08-31 Character recognition result correction method for address character string Pending JPH04111186A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2230927A JPH04111186A (en) 1990-08-31 1990-08-31 Character recognition result correction method for address character string

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP2230927A JPH04111186A (en) 1990-08-31 1990-08-31 Character recognition result correction method for address character string

Publications (1)

Publication Number Publication Date
JPH04111186A true JPH04111186A (en) 1992-04-13

Family

ID=16915466

Family Applications (1)

Application Number Title Priority Date Filing Date
JP2230927A Pending JPH04111186A (en) 1990-08-31 1990-08-31 Character recognition result correction method for address character string

Country Status (1)

Country Link
JP (1) JPH04111186A (en)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH08167007A (en) * 1994-12-14 1996-06-25 Nec Corp Symbol string reader
US5754671A (en) * 1995-04-12 1998-05-19 Lockheed Martin Corporation Method for improving cursive address recognition in mail pieces using adaptive data base management

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH08167007A (en) * 1994-12-14 1996-06-25 Nec Corp Symbol string reader
US5754671A (en) * 1995-04-12 1998-05-19 Lockheed Martin Corporation Method for improving cursive address recognition in mail pieces using adaptive data base management

Similar Documents

Publication Publication Date Title
US7142716B2 (en) Apparatus for searching document images using a result of character recognition
JP7149976B2 (en) Error correction method and apparatus, computer readable medium
EP0564827B1 (en) A post-processing error correction scheme using a dictionary for on-line handwriting recognition
CN110674396B (en) Text information processing method and device, electronic equipment and readable storage medium
CN109344387B (en) Method and device for generating shape near word dictionary and method and device for correcting shape near word error
JPH11328317A (en) Japanese character recognition error correction method and apparatus, and recording medium recording error correction program
US20110229036A1 (en) Method and apparatus for text and error profiling of historical documents
Lehal et al. A shape based post processor for Gurmukhi OCR
CN111782892B (en) Similar character recognition method, device, apparatus and storage medium based on prefix tree
CN111814781B (en) Method, device and storage medium for correcting image block recognition results
JP3975825B2 (en) Character recognition error correction method, apparatus and program
JP2998054B2 (en) Character recognition method and character recognition device
JPS6262388B2 (en)
JP2012141742A (en) Character string retrieval device, character string retrieval method and character string retrieval program
JP3071745B2 (en) Post-processing method of character recognition result
JP3924899B2 (en) Text search apparatus and text search method
JP2827066B2 (en) Post-processing method for character recognition of documents with mixed digit strings
JP2000251017A (en) Word dictionary creation device and word recognition device
JPH03257693A (en) Character recognized result correcting system
KR101663521B1 (en) Method and program for proofreading word spacing
KR101629726B1 (en) Method and program for proofreading word spacing
JP2839515B2 (en) Character reading system
JPS63268082A (en) Pattern recognizing device
JPH0757059A (en) Character recognition device
JP2000036008A (en) Character recognition device and storage medium