JPH0337781A - Post-processing method for character recognition results - Google Patents

Post-processing method for character recognition results

Info

Publication number
JPH0337781A
JPH0337781A JP1172722A JP17272289A JPH0337781A JP H0337781 A JPH0337781 A JP H0337781A JP 1172722 A JP1172722 A JP 1172722A JP 17272289 A JP17272289 A JP 17272289A JP H0337781 A JPH0337781 A JP H0337781A
Authority
JP
Japan
Prior art keywords
line
recognition
recognition result
character
processing
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP1172722A
Other languages
Japanese (ja)
Other versions
JP2891368B2 (en
Inventor
Takakuni Minewaki
隆邦 嶺脇
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ricoh Co Ltd
Original Assignee
Ricoh Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ricoh Co Ltd filed Critical Ricoh Co Ltd
Priority to JP1172722A priority Critical patent/JP2891368B2/en
Publication of JPH0337781A publication Critical patent/JPH0337781A/en
Application granted granted Critical
Publication of JP2891368B2 publication Critical patent/JP2891368B2/en
Anticipated expiration legal-status Critical
Expired - Lifetime legal-status Critical Current

Links

Landscapes

  • Character Discrimination (AREA)

Abstract

PURPOSE:To curtail the capacity of a storage memory by setting a unit of the processing to one line, and also, shifting a delimiter part of the last word and sentence clause, etc., of a line being an object of the processing to the head of the next line. CONSTITUTION:When the character recognition of a one-line portion is ended, a recognition result reading-in part 21 reads in recognition of a one-line portion from a recognition result memory 8, and stores it in a line recognition result memory 22. Subsequently, a temporary sentence clause segmenting part 23 delimits a recognition result character-string of one line in the line recognition result memory 22 into plural parts being analytic units by a position in which the kind of a character is varied, etc., and an analytic range determining part 24 considers a position of a temporary sentence clause, and sets a head character position of the line recognition result memory 22 to a character position of this side of the temporary sentence clause to an analytic object range. Next, a sentence analyzing/correcting part 25 rewrites correctly a result of recognition in the line recognition result memory 22, and a final sentence clause writing part 26 executes an operation for shifting a result of recognition (the part excluded from an object to be analyzed and corrected) of the temporary sentence clause of the line end in the line recognition result memory 22, to the head. In such a way, the storage capacity can be curtailed.

Description

【発明の詳細な説明】 〔産業上の利用分野〕 本発明は1文字認識に係り、特に文字認識結果の修正の
ための後処理に関する。
DETAILED DESCRIPTION OF THE INVENTION [Field of Industrial Application] The present invention relates to single character recognition, and particularly to post-processing for modifying character recognition results.

〔従来の技術〕[Conventional technology]

王文字単位の文字認識によって誤認識を完全に排除する
ことは極めて困難であるため、漢字OCRにおいては、
単語照合や形態素解析等によって文字認識結果の誤り修
正(後処理)を行うことが多い。
It is extremely difficult to completely eliminate misrecognition by character recognition in units of royal characters, so in kanji OCR,
Error correction (post-processing) of character recognition results is often performed by word matching, morphological analysis, etc.

従来、一般文書を対象とした後処理の方式は、認識対象
となる領域(例えば段落)の認識が終了した後で、その
領域に改行はないものとして、領域の全部の文字列につ
いて後処理を行うようになっている。あるいは、文を後
処理の単位とした方式もあり、これにおいては、文の先
頭から句点または読点までの文字列を、改行を無視して
読み込み後処理を行う。
Conventionally, the post-processing method for general documents is to post-process all character strings in the area after the recognition of the area to be recognized (for example, a paragraph) is completed, assuming that there are no line breaks in that area. It is supposed to be done. Alternatively, there is also a method in which a sentence is the unit of post-processing, in which the character string from the beginning of the sentence to a period or comma is read and post-processed, ignoring line breaks.

なお、単語照合や形態素解析による日本語文書の文字認
識結果の後処理に関する公知資料としては、例えば「西
野ほか:゛日本語文書リーダ後処理の実現、自然言語処
理64−6 (1987゜11.20)、PP、45−
52Jがある。
Publicly known materials regarding post-processing of character recognition results of Japanese documents using word matching and morphological analysis include, for example, "Nishino et al.: Realization of post-processing for Japanese document readers, Natural Language Processing 64-6 (1987, November 2013). 20), PP, 45-
There is 52J.

〔発明が解決しようとする課題〕[Problem to be solved by the invention]

しかし、段落や文を処理単位とした後処理方式によれば
、段落や文が長い場合、その認識結果が終了するまで後
処理の開始が待たされ処理効率が悪く、また認識結果の
格納のために大きなメモリが必要となるという問題があ
る。
However, with post-processing methods that treat paragraphs or sentences as processing units, if the paragraphs or sentences are long, the post-processing has to wait until the recognition results are finished, resulting in poor processing efficiency. The problem is that it requires a large amount of memory.

よって本発明の目的は、このような問題を解決できる文
字認識結果の後処理方式を提供することにある。
Therefore, an object of the present invention is to provide a post-processing method for character recognition results that can solve such problems.

〔課題を解決するための手段〕[Means to solve the problem]

本発明は、文字認識結果に対し単語照合や形態素解析等
によって誤りを修正する後処理において、処理の単位を
1行とするとともに、処理の対象となっている行末の単
語や文節等の区切り部分を次行の洗頭へ移し、あるいは
解析不能となった行末の部分を次行の先頭へ移すことを
特徴とするものである。
In the post-processing of character recognition results to correct errors by word matching, morphological analysis, etc., the unit of processing is one line, and the delimiter of words, phrases, etc. at the end of the line to be processed is It is characterized by moving the line to the next line, or moving the part at the end of the line that cannot be parsed to the beginning of the next line.

〔作 用〕[For production]

上行の文字列は、単語や文節等に区切られて解析される
が、一つの単語が行末と次の行頭に分裂していることが
ある。このような改行によって分裂した単語は、それを
考慮せずに扱ったのでは正しい解析ができず、その部分
に誤認識があると修正に失敗する。
The character string in the upper line is divided into words, phrases, etc. and analyzed, but one word may be split into the end of a line and the beginning of the next line. Words split by line breaks like this cannot be analyzed correctly if they are handled without taking this into consideration, and if there is a misrecognition in that part, correction will fail.

しかし1本発明によれば、1行を単位として後処理を行
うけれども、行末の単語や文節の区切り部分、あるいは
、行末の解析不能な部分を次の行の先頭へ移し、次行の
後処理で扱うため、改行により分裂した単語についても
、その誤認識を正しく修正することができる。
However, according to the present invention, post-processing is performed on a line-by-line basis, but the word or clause separation part at the end of a line, or the unanalyzable part at the end of a line, is moved to the beginning of the next line, and the next line is post-processed. , it is possible to correctly correct misrecognitions of words that are split due to line breaks.

また、一般に文字認識装置においては″領域抽出または
領域指定→″′行切り出し″→″文字切り出し″→″文
字認識″→゛′後処理″という手順で処理が進む。した
がって、後処理の単位が1行であると、ある行の後処理
と次の行の文字認識処理とを並列的に実行することによ
り、後処理の待ち時間を少なくして処理全体を高速化す
ることができる。さらに、上行に印刷または記入される
文字数はほぼ決まっているので、認識結果の記憶に必要
なメモリ量を減らすことができる。
Further, in general, in a character recognition device, processing proceeds in the following order: ``area extraction or area specification → ``line extraction'' → ``character extraction'' → ``character recognition'' → ``post-processing.'' Therefore, if the unit of post-processing is one line, the post-processing of one line and the character recognition process of the next line can be executed in parallel, reducing the waiting time for post-processing and speeding up the entire process. can do. Furthermore, since the number of characters printed or written in the upper line is approximately fixed, the amount of memory required to store recognition results can be reduced.

〔実施例〕〔Example〕

以下、図面を用い本発明の実施例について説明する。 Embodiments of the present invention will be described below with reference to the drawings.

第1図に本発明に係る漢字OCRのブロック図を示す、
この漢字OCRの全体的な構成及び処理の流れは従来と
同様である。すなわち1画像入力部工がスキャナ(ある
いは他の入力装置り2より対象画像を読み込み、画像メ
モリ3に格納する。
FIG. 1 shows a block diagram of Kanji OCR according to the present invention.
The overall structure and processing flow of this Kanji OCR is the same as the conventional one. That is, one image input section reads a target image from a scanner (or other input device 2) and stores it in the image memory 3.

範囲指定部4が、画像メモリ3に入力された画像より認
識する範囲を自動的に指定し、あるいは指定手段によっ
て指定する。切り出し部5が範囲指定された領域から行
イメージを切り出し、この行イメージより個々の文字の
イメージを切り出す。
The range specifying unit 4 automatically specifies the range to be recognized from the image input to the image memory 3, or specifies it by a specifying means. A cutting unit 5 cuts out a line image from the specified range, and cuts out individual character images from this line image.

認識部6が、文字辞書メモリ7に格納された辞書を月い
、切り出された各文字イメージの文字認識を行い、認識
結果を認識結果メモリ8に格納する。
The recognition unit 6 uses the dictionary stored in the character dictionary memory 7, performs character recognition on each cut-out character image, and stores the recognition results in the recognition result memory 8.

後処理部9が、認識結果文字列を認識結果メモリ7より
読み込み、単語辞書・文法辞書メモリ10に格納されて
いる単語辞書や文法辞書を参照して解析し、文章として
不適当な部分を修正する。この後処理部9が本発明に直
接係わる部分であり、その詳細については実施例別に後
述する。出力部11は、後処理の結果をCRTまたは出
力ファイル12へ出力する。
The post-processing unit 9 reads the recognition result character string from the recognition result memory 7, analyzes it with reference to the word dictionary and grammar dictionary stored in the word dictionary/grammar dictionary memory 10, and corrects portions that are inappropriate as sentences. do. This post-processing section 9 is a section directly related to the present invention, and its details will be described later for each embodiment. The output unit 11 outputs the post-processing results to a CRT or an output file 12.

去」奥4L 本実施例における後処理部9のブロック図を第2@に示
す、上行分の文字認識が終了すると、認識結果読み込み
部21が認識結果メモリ8より1行分の認識結果を読み
込み1行認識結果メモリ22に格納する。なお、前の行
の行末の末解析部分が行認識結果メモリ22の先頭に書
き移されている場合は、その次の文字位置から書き込む
4L The block diagram of the post-processing unit 9 in this embodiment is shown in the second @. When the character recognition for the upper line is completed, the recognition result reading unit 21 reads the recognition result for one line from the recognition result memory 8. One line recognition result memory 22 is stored. Note that if the end analysis part of the end of the previous line has been written to the beginning of the line recognition result memory 22, it is written from the next character position.

次に仮文節切り部23が、行認識結果メモリ22内のt
行の認識結果文字列を、文字種の変化(例えば、ひらが
な→非ひらがなの切りかわり)する位置等によって、解
析単位としての複数の部分に区切る。このようにして区
切られた部分を“仮文節”と呼ぶことにする。
Next, the provisional phrase cutting unit 23 selects the t in the line recognition result memory 22.
The line recognition result character string is divided into a plurality of parts as units of analysis, depending on the position where the character type changes (for example, switching from hiragana to non-hiragana). The parts separated in this way will be called "kari bunsetsu".

解析範囲決定部24が、仮文節の位置を考慮し、行認識
結果メモリ22の先頭文字位置から最後の仮文節の手前
の文字位置までを解析対象範囲に設定する。
The analysis range determining unit 24 takes into account the position of the pseudo clause and sets the range from the first character position in the line recognition result memory 22 to the character position before the last pseudo clause as the analysis target range.

文章解析・修正部25が、設定された解析対象範囲の各
仮文節について、単語辞書・文法辞書メモリ10内の単
語辞書および文法辞書を利用し。
The sentence analysis/correction unit 25 uses the word dictionary and grammar dictionary in the word dictionary/grammar dictionary memory 10 for each temporary clause in the set analysis target range.

単語照合や形態素解析を行って誤認識の修正を行い1行
認識結果メモリ22内の認識結果を正しく書き換える。
Word matching and morphological analysis are performed to correct erroneous recognition, and the recognition results in the one-line recognition result memory 22 are rewritten correctly.

このように修正された結果が出力部1工により出力され
る。
The thus corrected results are output by the output unit 1.

解析対象範囲の最後の仮文節まで解析・修正が終わると
、最終文節書き込み部26が、行認識結果メモリ22内
の行末の仮文節の認識結果(解析修正の対象から外され
た部分)を、先頭へ移す操作を行う。すなわち、行末の
仮文節は改行により分断、されている可能性があるが、
これは次の行の行頭の仮文節として、次行の後処理の対
象とされることになる。
When the analysis and correction are completed up to the last pseudo clause in the analysis target range, the final clause writing unit 26 writes the recognition result of the pseudo clause at the end of the line in the line recognition result memory 22 (the part excluded from the analysis and correction target) to Perform the operation to move to the top. In other words, the pseudo clause at the end of the line may be separated by a line break, but
This will be treated as a provisional phrase at the beginning of the next line, and will be subject to post-processing on the next line.

なお、このような上行についての後処理の実行中に1次
行以降の文字認識処理が並列的に実行される。
Note that while the post-processing for the upper line is being executed, character recognition processing for the primary line and subsequent lines is executed in parallel.

ここで1次のような対象文章を例にとって、後まず、第
I行目に対し次の認識結果が得られたとする。
Let's take the following target sentence as an example, and assume that the following recognition result is obtained for the I-th line.

“文字認識技術は現代の情報処″ これは文字種変化(ひらがな→非ひらがな)によって次
のように仮文節に区切られる。
“Character recognition technology is modern information processing.” This is divided into temporary clauses as follows by character type change (hiragana → non-hiragana).

″文字認識技術は/現代の/情報処″ そして、解析対象範囲は次のようになる6“文字認識技
術は/現代の″ これで、改行により分裂していた単語″情報処(理)”
は第I行目の解析対象から外される。
``Character recognition technology is / modern / information processing'' And the scope of analysis is as follows 6 ``Character recognition technology is / modern / information processing'' Now, the word that was split by line breaks ``information processing (processing)''
is excluded from the analysis target in the Ith line.

この第1行目の解析対象範囲に対する解析・修正が終わ
り、結果が出力されると、行末の処理されなかった仮文
節部分が行認識結果メモリ22の先頭に書き移され、そ
れ以降の文字位置はクリアされる。
When the analysis and correction of the analysis target range of the first line is finished and the results are output, the unprocessed pseudo clause part at the end of the line is transferred to the beginning of the line recognition result memory 22, and the subsequent character positions are is cleared.

第2行目の処理に移り、その認識結果が行認識メモリ2
2の前行から移された部分に続けて書き込まれるので、
第2行目の認識結果は次のようになる。
The process moves to the second line, and the recognition result is stored in the line recognition memory 2.
It will be written following the part moved from the previous line of 2, so
The recognition result in the second line is as follows.

″′情報処理の花形技術である。″ これで、2行にまたがって分裂していた単語“情報処理
が一つにまとまり、解析・修正が可能となる。
``This is the star technology of information processing.'' With this, the word ``information processing'', which was split across two lines, has been unified into one, making it possible to analyze and modify it.

そして、同じように仮文節に区切り、対象範囲設定、解
析・修正を行う、ただし、この行は、最後の文字が句点
であるので、行末で単語の分裂は起こっていないと判断
し1行の最後の仮文節も解析対象範囲とする1行の最後
の文字が読点の場合も同様の扱いをする。
Then, divide it into pseudo clauses in the same way, set the target range, analyze and modify. However, since the last character of this line is a period, it is determined that there is no word splitting at the end of the line. The last pseudo clause is treated in the same way even if the last character of the line to be analyzed is a comma.

失凰量主 本実施例における後処理部9のブロック図を第3図に示
す、1行分の文字認識が終了すると、認識結果読み込み
部31が認識結果メモリ8より上行分の認識結果を読み
込み、行認識結果メモリ32に格納する。なお、前の行
の行末の解析不能部分が行認識結果メモリ32の先頭に
書き移されている場合は、その次の文字位置から書き込
む。
The block diagram of the post-processing unit 9 in this embodiment is shown in FIG. 3. When character recognition for one line is completed, the recognition result reading unit 31 reads the recognition results for the upper line from the recognition result memory 8. , are stored in the line recognition result memory 32. Note that if the unanalyzable part at the end of the previous line has been written to the beginning of the line recognition result memory 32, the data is written from the next character position.

次に文章解析・修正部33が、行認識結果メモリ32の
内容について、単語辞書・文法辞書メモリ10内の単語
辞書および文法辞書を利用し、単語照合や形態素解析に
よる修正処理を行い、行認識結果メモリ32内の認識結
果を正しく書き換える。このように修正された結果が出
力部11により出力される。ただし、行末に解析不能と
なった部分(該当単語が見つからない文字列)があれば
、その行末部分に″次行もちこし″の印を付ける。
Next, the sentence analysis/correction unit 33 uses the word dictionary and grammar dictionary in the word dictionary/grammar dictionary memory 10 to perform correction processing on the contents of the line recognition result memory 32 through word matching and morphological analysis, and performs line recognition. The recognition results in the result memory 32 are rewritten correctly. The result corrected in this way is outputted by the output unit 11. However, if there is a part at the end of a line that cannot be parsed (a character string for which the corresponding word cannot be found), the end of the line is marked as ``continued to the next line.''

次行もちこし部分書き込み部34が、行認識結果メモリ
32の先頭へ、7次行もちこし印がつけられた行末部分
を、先頭へ書き移し、それ以降の文字位置をクリアする
。これで、現在行の行末の解析不能となった部分は、次
行の先頭へ移され、次行の後処理の対象とされる。
The next line mochikoshi part writing unit 34 writes the end of the line marked with the 7th line mochikoshi mark to the beginning of the line recognition result memory 32, and clears subsequent character positions. Now, the unanalyzable portion at the end of the current line is moved to the beginning of the next line, and is subjected to post-processing of the next line.

なお、このような1行についての後処理の実行中に1次
行以降の文字L&識処理が並列的に実行される。
It should be noted that while the post-processing for one line is being executed, the character L& recognition process for the primary line and subsequent lines is executed in parallel.

ここで、次のような対象文章を例にとって、後まず、第
1行目に対し次の認識結果が得られたとする。
Let's take the following target sentence as an example and assume that the following recognition result is obtained for the first line.

゛′文字認識技術は現代の情報部” この文字列について、先頭から単語照合あるいは形態素
解析等によって解析を進めていくと、次のようになる(
/は単語の切れ目)。
``Character recognition technology is the modern information department.'' If this character string is analyzed from the beginning by word matching or morphological analysis, it will look like this (
/ is a word break).

゛′文字/認識/技術/は/現代/の/情報/処理この
最後の文字″処理は該当する単語がなく解析不能となる
ので、この“処理が次行もちこし部分となり、修正結果
出力は次のようになる。
``Character/Recognition/Technology/is/Modern/'s/Information/Processing This last character'' process cannot be parsed because there is no corresponding word, so this ``processing is carried over to the next line, and the corrected result output is It will look like this:

″文字認識技術は現代の情報′″ つぎに、″処理が次行の先頭に移されるので。``Character recognition technology is modern information'' Next, the ``processing is moved to the beginning of the next line.

第2行の認識結果は ″処理の花形技術である。The recognition result of the second line is ``It is a star processing technology.

となり、2行にまたがって分裂していた単語がつながる
。この行の最後の文字は句点であるので、行末まで解析
不能部分がなく、最後まで処理されることになる。
The words that were split across two lines are now connected. Since the last character of this line is a period, there is no part that cannot be parsed until the end of the line, and the line is processed until the end.

〔発明の効果〕〔Effect of the invention〕

以上説明した如く、本発明によれば、改行により分裂し
た単語を含めた認識結果を修正することが可能であり、
後処理と認識処理を行単位で並列的に実行して処理全体
の効率向上が可能であり、また、後処理の単位は1行で
あるので、後処理のための認識結果の記憶メモリの容量
削減が可能である。
As explained above, according to the present invention, it is possible to correct recognition results including words separated by line breaks,
It is possible to improve the efficiency of the entire process by performing post-processing and recognition processing in parallel row by row, and since the unit of post-processing is one row, the memory capacity for storing recognition results for post-processing is reduced. reduction is possible.

【図面の簡単な説明】[Brief explanation of drawings]

第1図は本発明に係る漢字OCRのブロック図、第2図
は本発明の一実施例における後処理部のブロック図、第
3図は本発明の他の実施例における後処理部のブロック
図である。 6・・・認識部、 8・・・認識結果メモリ、9・・・
後処理部、 10・・・単語辞書・文法辞書メモリ、 
 21・・・認識結果読み込み部、22・・・行認識結
果メモリ、  23・・・仮文節切り部、 24・・・
解析範囲決定部、 25・・・文章解析・修正部、 2
6・・・最終文節書き込み部、31・・・認識結果読み
込み部、 32・・・行認識結果メモリ、 33・・・
文章解析・修正部、34・・・次行もちこし部分書き込
み部。 ハ゛ス 第2図
FIG. 1 is a block diagram of a Kanji OCR according to the present invention, FIG. 2 is a block diagram of a post-processing unit in one embodiment of the present invention, and FIG. 3 is a block diagram of a post-processing unit in another embodiment of the present invention. It is. 6... Recognition unit, 8... Recognition result memory, 9...
Post-processing unit, 10... word dictionary/grammar dictionary memory,
21... Recognition result reading unit, 22... Line recognition result memory, 23... Provisional phrase cutting unit, 24...
Analysis range determination unit, 25...Text analysis/correction unit, 2
6... Final phrase writing unit, 31... Recognition result reading unit, 32... Line recognition result memory, 33...
Sentence analysis/correction section, 34... Next line sticky part writing section. Bus Figure 2

Claims (2)

【特許請求の範囲】[Claims] (1)文字認識結果の誤りを修正する後処理において、
処理の単位を1行とするとともに、処理の対象となって
いる行の最後の単語や文節等の区切り部分を次行の先頭
へ移すことを特徴とする文字認識結果の後処理方式。
(1) In post-processing to correct errors in character recognition results,
A post-processing method for character recognition results characterized in that the unit of processing is one line, and a delimiter such as the last word or clause of the line being processed is moved to the beginning of the next line.
(2)文字認識結果の誤りを修正する後処理において、
処理の単位を1行とするとともに、解析不能となった行
末の部分を次行の先頭へ移すことを特徴とする文字認識
結果の後処理方式。
(2) In post-processing to correct errors in character recognition results,
A post-processing method for character recognition results characterized in that the unit of processing is one line, and the part at the end of a line that cannot be analyzed is moved to the beginning of the next line.
JP1172722A 1989-07-04 1989-07-04 Post-processing method of character recognition result Expired - Lifetime JP2891368B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP1172722A JP2891368B2 (en) 1989-07-04 1989-07-04 Post-processing method of character recognition result

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP1172722A JP2891368B2 (en) 1989-07-04 1989-07-04 Post-processing method of character recognition result

Publications (2)

Publication Number Publication Date
JPH0337781A true JPH0337781A (en) 1991-02-19
JP2891368B2 JP2891368B2 (en) 1999-05-17

Family

ID=15947118

Family Applications (1)

Application Number Title Priority Date Filing Date
JP1172722A Expired - Lifetime JP2891368B2 (en) 1989-07-04 1989-07-04 Post-processing method of character recognition result

Country Status (1)

Country Link
JP (1) JP2891368B2 (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2010157241A (en) * 2008-12-30 2010-07-15 Nhn Corp Method and system for correcting ocr result, and computer-readable recording medium

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2010157241A (en) * 2008-12-30 2010-07-15 Nhn Corp Method and system for correcting ocr result, and computer-readable recording medium

Also Published As

Publication number Publication date
JP2891368B2 (en) 1999-05-17

Similar Documents

Publication Publication Date Title
CN117649670A (en) Document layout analysis model training method, application method, computer device and computer readable storage medium
CN115311666A (en) Image-text recognition method and device, computer equipment and storage medium
JPH08320914A (en) Table recognition method and device
CN120745560A (en) Outline level extraction method based on vision-language multimode algorithm
JPH0337781A (en) Post-processing method for character recognition results
JPH0619962A (en) Text dividing device
JPH04252390A (en) Post processing method for character recognition result
JPH0991371A (en) Character display device
JP3932912B2 (en) Character string shaping device, method and program
JP2918666B2 (en) Text image extraction method
JP2746345B2 (en) Post-processing method for character recognition
Alhubaiti et al. Typefaces and ligatures in printed Arabic text: a deep learning-based OCR perspective
JPH05143351A (en) Source program comparing system
JPH028348B2 (en)
Bumbu et al. Automation of PostOCR error correction in the digitization of historical texts
EP2804131A2 (en) System and methods for arabic text recognition and arabic corpus building
Konstantinidou Text Line Detection in Greek Polytonic Documents: A Comparative Analysis of CRAFT, EAST, PaddleOCR and YOLO
JPH10222612A (en) Document recognition device
KR100200666B1 (en) High speed character recognition device and method
CN119129918A (en) Structuring method of assembly process specification based on natural language processing
CN120747300A (en) Click-to-read configuration file generation method, device, equipment and medium
JPS6198487A (en) Dictionary selection method
JP2000123116A (en) Character recognition result correction method
JPH09167206A (en) Space detection method for Japanese-English mixed document, pitch format determination method, space detection method for constant pitch alphanumeric character string, and space detection method for proportional pitch alphanumeric character string
JP2723462B2 (en) Syntax signal analysis method and apparatus

Legal Events

Date Code Title Description
FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20080226

Year of fee payment: 9

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20090226

Year of fee payment: 10

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20100226

Year of fee payment: 11

EXPY Cancellation because of completion of term
FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20100226

Year of fee payment: 11