JPH0477965A - digital copying machine - Google Patents
digital copying machineInfo
- Publication number
- JPH0477965A JPH0477965A JP2192068A JP19206890A JPH0477965A JP H0477965 A JPH0477965 A JP H0477965A JP 2192068 A JP2192068 A JP 2192068A JP 19206890 A JP19206890 A JP 19206890A JP H0477965 A JPH0477965 A JP H0477965A
- Authority
- JP
- Japan
- Prior art keywords
- character
- image
- word
- translation
- information
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Landscapes
- Control Or Security For Electrophotography (AREA)
- Character Discrimination (AREA)
- Machine Translation (AREA)
- Document Processing Apparatus (AREA)
- Accessory Devices And Overall Control Thereof (AREA)
- Dot-Matrix Printers And Others (AREA)
Abstract
(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.
Description
【発明の詳細な説明】
[産業上の利用分野]
本発明は、デジタル複写装置に関し、特に画像中の文字
情報を単語単位で例えば英語から日本語に翻訳する機能
を備えるデジタル複写装置に関する。DETAILED DESCRIPTION OF THE INVENTION [Field of Industrial Application] The present invention relates to a digital copying apparatus, and more particularly to a digital copying apparatus having a function of translating character information in an image from English to Japanese in units of words.
[従来の技術]
言語間の翻訳作業は、一般に非常に高度な知識を要し、
しかも翻訳には時間がかかる。そこで従来より、翻訳作
業を支援もしくは自動化する技術について様々な提案が
され、商品化されている技術も存在する。[Conventional technology] Translation work between languages generally requires very advanced knowledge.
Moreover, translation takes time. Therefore, various proposals have been made for technologies to support or automate translation work, and some technologies have even been commercialized.
例えば、特開昭54−154955号公報及び特開昭5
5=15562号公報には、単語単位で翻訳する装置が
開示されており、特開昭511−4480号公報には文
法や意味の解析を実行して文章の翻訳をする装置が開示
されている。For example, JP-A No. 54-154955 and JP-A No. 5
5=15562 discloses a device that translates word by word, and Japanese Patent Application Laid-Open No. 511-4480 discloses a device that performs grammar and meaning analysis to translate sentences. .
なお、文章全体の翻訳については現在ではまだ誤りの発
生頻度が高いので、翻訳結果の修正作業が不可欠であり
、実用的な段階にはない。In addition, since errors still occur frequently when translating entire sentences, it is necessary to correct the translation results, and the method is not yet at a practical stage.
[発明が解決しようとする課題]
ところで従来の翻訳装置においては、例えば単語琳位で
、その文字情報をキーボード等を利用してオペレータが
入力する必要がある。従って、例えば英文の書類がある
場合には、オペレータがその書類上から順次に単語を読
み取って、その文字情報をキー人力する必要がある。こ
の種の入力作業は非常に大きな労力を伴ない、特にキー
操作に不慣れなオペレータの場合には、入力のために非
常に長い時間を必要とする。[Problems to be Solved by the Invention] However, in conventional translation devices, it is necessary for an operator to input character information using a keyboard or the like, for example, in the form of a word. Therefore, if there is a document in English, for example, it is necessary for the operator to read the words sequentially from the document and enter the character information manually. This type of input work requires a very large amount of effort and requires a very long time for inputting, especially in the case of an operator who is not familiar with key operations.
本発明は、翻訳すべき文字情報の翻訳装置への入力作業
を簡単にし、翻訳作業を容易にすることを課題とする。An object of the present invention is to simplify the input work of character information to be translated into a translation device, thereby facilitating the translation work.
[課題を解決するための手段]
上記課題を解決するために、本発明においては、原稿画
像を微小画素情報の集合として読取る画像読取手段、該
手段の読取った情報を蓄積するメモリ手段、及び該メモ
リ手段に蓄積された画像情報を出力する出力手段、を含
むデジタル複写装置において:前記メモリ手段に蓄積さ
れた入力画像情報から文字情報を各文字毎に切り出す文
字切り出し手段:切り出された文字の間隔に基づいて、
1つの文字もしくは複数文字の集合でなる単語毎に文字
情報を区分する単語抽出手段;文字毎の特徴量が予め登
録された文字パターン辞書を含み、文字毎に切り出され
た画像情報から、それが示す文字を認識する文字認識手
段:単語毎の翻訳結果が登録された単語辞書を含み、前
記単語抽出手段によって抽出された単語毎に、それを構
成する文字の構成に応じて単語辞書を参照し、各単語の
翻訳結果を認識する、翻訳手段;及び翻訳手段の翻訳結
果として得られる文字情報を、前記メモリ手段の出力画
像領域に画像情報として書込む、出力画像作成手段:を
設ける。[Means for Solving the Problems] In order to solve the above problems, the present invention includes an image reading means for reading a document image as a set of minute pixel information, a memory means for storing the information read by the means, and a memory means for storing the information read by the means. In a digital copying apparatus including an output means for outputting the image information stored in the memory means: a character cutting means for cutting out character information for each character from the input image information stored in the memory means: an interval between the cut out characters; On the basis of the,
Word extraction means for classifying character information into words consisting of one character or a set of multiple characters; includes a character pattern dictionary in which feature amounts for each character are registered in advance, and extracts character information from image information extracted for each character. Character recognition means for recognizing the characters indicated: includes a word dictionary in which translation results for each word are registered, and refers to the word dictionary for each word extracted by the word extraction means according to the structure of the characters constituting it. , a translation means for recognizing a translation result of each word; and an output image creation means for writing character information obtained as a translation result of the translation means into an output image area of the memory means as image information.
[作用]
本発明の装置によれば、原稿画像上から自動的に文字及
び単語を認識し、認識した単語毎に翻訳を実行してその
結果を出力画像の形で出力するので、翻訳作業は普通に
複写機を使うのと同様に、翻訳対象の原稿をイメージス
キャナにセットしてコピーボタンを押す、という極めて
簡単な作業になる。[Function] According to the device of the present invention, characters and words are automatically recognized from the original image, translation is executed for each recognized word, and the result is output in the form of an output image, so that the translation work is easy. Just like using a regular copy machine, it's an extremely simple process of placing the manuscript to be translated into the image scanner and pressing the copy button.
ところで、画像情報中から正しく文字を認識するのは非
常に難しいので、この種の装置においては、文字の認識
率を高めることが極めて重要である。By the way, since it is very difficult to correctly recognize characters from image information, it is extremely important to increase the character recognition rate in this type of device.
そこで本発明の好ましい実施例においては、文字の大き
さ及びフォントの違いに応じた複数種類の文字パターン
辞書を設けて、文字を切り出す時に検出される行ピッチ
と文字ピッチとに基づいて認識対象文字の大きさを識別
しその結果により使用する文字パターン辞書を限定する
とともに、原稿画像中の一部分の文字列に対して複数種
類の文字パターン辞書を利用して認識を実施し、その結
果において文字認識率が高い文字パターン辞書を自動的
に選択して以降の認識を実行する。Therefore, in a preferred embodiment of the present invention, a plurality of types of character pattern dictionaries are provided according to differences in character size and font, and characters to be recognized are recognized based on the line pitch and character pitch detected when cutting out characters. In addition to identifying the size of the character pattern and limiting the character pattern dictionary to be used based on the result, recognition is performed using multiple types of character pattern dictionaries for a part of the character string in the original image, and character recognition is performed based on the results. To automatically select a character pattern dictionary with a high rate and perform subsequent recognition.
原稿画像中の文字が例えば活字である時には、それのフ
ォントと一致する辞書を使うことによって文字の認識率
を大幅に向上しうる。しかし、般の人が画像上の活字を
見てそのフォントを認識することは非常に離しいし、複
数のフォントの辞書を認識処理中に常時参照するのでは
認識処理の所要時間が長くなってしまう。そこで上記の
ように一部分の文字列についてだけ複数フォントの辞書
を参照し、各辞書に対する認識率の大きいものだけを選
択して以後の処理で使用するようにすれば、適切なフォ
ントの辞書を自動的に選択できるし、認識率が向上し、
しかも短い時間で文字を認識できる。When the characters in the original image are printed characters, for example, the recognition rate of the characters can be greatly improved by using a dictionary that matches the font of the characters. However, it is very difficult for the general public to recognize the font by looking at the printed text on an image, and the time required for the recognition process increases if a dictionary of multiple fonts is constantly referenced during the recognition process. . Therefore, if you refer to multiple font dictionaries for only part of the character string as described above, select only the one with the highest recognition rate for each dictionary, and use it in subsequent processing, the appropriate font dictionary will be automatically selected. can be selectively selected, the recognition rate is improved,
Moreover, characters can be recognized in a short time.
また本発明の好ましい実施例においては、翻訳手段が翻
訳結果を検出できない時に、自動的に修正モードの処理
に移行し、翻訳対象の単語文字列毎に、その一部分の文
字を予め定められた規則に基づいて修正し、再び翻訳を
実行する。Further, in a preferred embodiment of the present invention, when the translation means cannot detect a translation result, it automatically shifts to the correction mode processing, and for each word string to be translated, a part of the characters is changed according to a predetermined rule. Make corrections based on this and run the translation again.
本発明の装置においては1画像中から単語の抽出を自動
的に行なってそれを翻訳するので、抽出された単語と単
語辞書の構成との相性が悪くなる場合があり、そのよう
な場合には翻訳の成功率が低下する。例えば、文の途中
にカンマ(1)があると、抽出された単語の最後にカン
マが付加される場合があるが、このような単語は通常は
辞書に登録されないので、翻訳不可能になる。しかしこ
の種の特別な単語の発生は、ある種の規則に従っている
ので、例えば単語の最後の文字がカンマの場合にはその
文字を削除する、というように予め定めた規則に基づい
て単語を自動的に修正することが可能である。このよう
な自動修正を実行すれば、オペレータの修正操作を伴な
わずに、翻訳の成功率を向上することができる。Since the device of the present invention automatically extracts words from one image and translates them, the extracted words may not be compatible with the structure of the word dictionary. Translation success rate decreases. For example, if there is a comma (1) in the middle of a sentence, a comma may be added to the end of the extracted word, but such words are usually not registered in dictionaries and cannot be translated. However, the generation of this kind of special word follows certain rules; for example, if the last character of a word is a comma, that character is deleted. It is possible to modify the If such automatic correction is performed, the success rate of translation can be improved without requiring correction operations by an operator.
本発明の他の目的及び特徴は、以下の5図面を参照した
実施例説明により明らかになろう。Other objects and features of the present invention will become clear from the description of the embodiments with reference to the following five drawings.
[実施例コ
第1図に、本発明を実施する一形式のデジタル複写機の
機構部の構成を示す。第1図を参照すると、このデジタ
ル複写機は、大きく分けて上部のイメージスキャナ10
0とその下に配置されたレーザプリンタ200で構成さ
れている。[Embodiment] FIG. 1 shows the structure of a mechanical section of a digital copying machine of one type in which the present invention is implemented. Referring to FIG. 1, this digital copying machine is roughly divided into an upper image scanner 10.
0 and a laser printer 200 placed below it.
イメージスキャナ100の最上部に、原稿を載置するコ
ンタクトガラスが配置されており、その下方に光学走査
系が設けられている。原稿は、光学走査系の露光ランプ
1によって露光され、その反射光、つまり両像光が光学
走査系に備わった各種ミラー及びレンズ2を通って受光
部3に結像される。この受光部3には、後述する一次元
CCDイ°メージセンサが設けられている。光学走査系
は、機械的な駆動系によって図面の左右方向に駆動され
るので、原稿面の各部の露光によって得られる画像光が
順次に、つまり1ライン毎にイメージセンナに読取られ
る。A contact glass on which a document is placed is placed at the top of the image scanner 100, and an optical scanning system is provided below it. A document is exposed to light by an exposure lamp 1 of an optical scanning system, and its reflected light, that is, both image lights, pass through various mirrors and lenses 2 provided in the optical scanning system and are imaged on a light receiving section 3. This light receiving section 3 is provided with a one-dimensional CCD image sensor, which will be described later. Since the optical scanning system is driven in the horizontal direction of the drawing by a mechanical drive system, the image light obtained by exposing each part of the document surface is sequentially read by the image sensor, line by line.
イメージセンナによって読取られた画像情報は、後述す
る処理によって出力両像に変換され、レーザプリンタ2
00の書込装置4から出力されるレーザ光を変調する。The image information read by the image sensor is converted into output images by the processing described later, and then sent to the laser printer 2.
The laser beam output from the writing device 4 of 00 is modulated.
画像情報によって変調されるレーザ光は、書込用の光学
系を通って、感光体ドラム5の表面に結像される。感光
体ドラム5の表面は、予めメインチャージャ6によって
全面が均一に所定の高電位に帯電しており、画像光の照
射を受けると、光強度に応じて電位が変化し1画像に対
応する電位分布、つまり静電潜像が形成される。The laser beam modulated by the image information passes through a writing optical system and forms an image on the surface of the photoreceptor drum 5. The entire surface of the photoreceptor drum 5 is uniformly charged to a predetermined high potential by the main charger 6 in advance, and when it is irradiated with image light, the potential changes depending on the light intensity and becomes a potential corresponding to one image. distribution, that is, an electrostatic latent image is formed.
感光体ドラム5に形成された静電潜像は、それが現像ユ
ニット7を通過する時にトナーの吸着によって可視化さ
れ、トナー像を形成する。The electrostatic latent image formed on the photosensitive drum 5 is made visible by adsorption of toner when it passes through the developing unit 7, forming a toner image.
一方、給紙カセット12又は13のうち選択されたもの
から記録紙が繰り出され、その記録紙は感光体ドラム5
上のトナー像の形成タイミングに同期して感光体ドラム
5の表面に重なるように送り込まれる。続いて、転写チ
ャージャの付勢により、感光体ドラム5上のトナー像は
記録紙に転写される。更に、分離チャージャ9の付勢に
よって、トナー像が転写された記録紙は感光体ドラム5
から分離して定着ユニット14に向かう。記録紙上のト
ナー像は、定着ユニット14によって記録紙に定着され
、その後、記録紙は複写機の外に排出される。On the other hand, recording paper is fed out from the paper feed cassette 12 or 13 selected, and the recording paper is fed to the photoreceptor drum 5.
The toner image is fed so as to overlap the surface of the photosensitive drum 5 in synchronization with the formation timing of the upper toner image. Subsequently, the toner image on the photosensitive drum 5 is transferred onto the recording paper by the urging of the transfer charger. Further, due to the biasing force of the separation charger 9, the recording paper onto which the toner image has been transferred is transferred to the photoreceptor drum 5.
The image is separated from the camera and goes to the fixing unit 14. The toner image on the recording paper is fixed onto the recording paper by the fixing unit 14, and then the recording paper is ejected from the copying machine.
画像の転写及び記録紙の分離が終了した後、感光体ドラ
ム5の表面は、クリーニングユニット10によってクリ
ーニングされ1次回の画像形成に備える。After the image transfer and recording paper separation are completed, the surface of the photosensitive drum 5 is cleaned by the cleaning unit 10 in preparation for the first image formation.
第2図に、第1図のデジタル複写機の電装部の構成を示
す。第2図を参照して説明する。イメージスキャナ10
0においては、CCDイメージセンナ110によって読
取られたビットマツプ形式の原稿画像の信号は、A/D
変換器120によってデジタル信号(この例では8ビツ
ト)に変換された後、シェーディング補正ユニット13
0によって濃度レベルのばらつきに関する補正を受け、
メモリユニット330に記憶される。メモリユニット3
30は、画像の1フレ一ム分の記憶領域の他に、文字認
識や翻訳処理の際に使用する記憶領域や出力画像用の記
憶領域を備えている。翻訳ユニット340は、後述する
ように、メモリユニット330上の入力画像情報を処理
し、文字認識や翻訳処理を実行した後で出力画像情報を
メモリユニット330上に生成する。文字認識のために
、文字パターン辞書350が設けられており、また翻訳
のために英和辞書が設けられている。文字パターン辞書
350には、サイズ及びフォントの異なる様々な文字種
に対応付けられた数組の辞書が備わっている。英和辞書
360には、様々な英単語とそれの各々に対応付けられ
た日本語の情報が対になって登録されている。FIG. 2 shows the configuration of the electrical equipment of the digital copying machine shown in FIG. 1. This will be explained with reference to FIG. image scanner 10
0, the bitmap format document image signal read by the CCD image sensor 110 is sent to the A/D
After being converted into a digital signal (8 bits in this example) by the converter 120, the shading correction unit 13
0 to correct for density level variations,
It is stored in memory unit 330. Memory unit 3
30 includes a storage area for one frame of an image, a storage area for use in character recognition and translation processing, and a storage area for output images. The translation unit 340 processes the input image information on the memory unit 330 and generates output image information on the memory unit 330 after performing character recognition and translation processing, as will be described later. A character pattern dictionary 350 is provided for character recognition, and an English-Japanese dictionary is provided for translation. The character pattern dictionary 350 includes several sets of dictionaries associated with various character types having different sizes and fonts. In the English-Japanese dictionary 360, various English words and Japanese information associated with each word are registered in pairs.
メモリユニット330上で生成された出力画像情報は、
各画素の黒/白に対応する二値情報の形でレーザプリン
タ200に印加され、バッファ220を通り、LDドラ
イバ230を通ってレーザダイオード240に付勢信号
として印加される。The output image information generated on the memory unit 330 is
The signal is applied to the laser printer 200 in the form of binary information corresponding to black/white of each pixel, passes through the buffer 220, passes through the LD driver 230, and is applied as an energizing signal to the laser diode 240.
従って、出力画像情報に応じて変調されたレーザ光をレ
ーザダイオード240が出力する。このレーザ光が書込
装置4から出力され、書込用の光学走査系を介して感光
体ドラム5の表面に照射される。Therefore, the laser diode 240 outputs laser light modulated according to the output image information. This laser light is output from the writing device 4 and is irradiated onto the surface of the photoreceptor drum 5 via an optical scanning system for writing.
オペレータからの指示は、この複写機の上面に配置され
た操作ボード310からのキー人力によって実施される
。メイン制御ユニット320は、操作ボード310上の
各種表示を制御するとともに、操作ボード310からの
キー人力を読取って、読取の開始、動作モード切換1画
像のハードコピー開始などを各部に指示する。Instructions from the operator are carried out by keystrokes from an operation board 310 arranged on the top of the copying machine. The main control unit 320 controls various displays on the operation board 310, reads key input from the operation board 310, and instructs each section to start reading, start hard copying one image of operation mode switching, and the like.
第1図の複写機の処理の概略を第3図に示す。FIG. 3 shows an outline of the processing of the copying machine shown in FIG.
第3図を参照して説明する。This will be explained with reference to FIG.
ステップ51では、イメージスキャナにセットされた原
稿の画像を読取り、画像データを画素単位で黒/白を示
す二値データの形で取込みそれをメモリユニット330
上の入力画像領域に蓄積する。この結果、例えば第4図
に示す画像IMGのような情報がメモリユニット330
上にビットマツプ形式で保持される。In step 51, the image of the document set in the image scanner is read, the image data is captured in the form of binary data indicating black/white in units of pixels, and the data is stored in the memory unit 330.
Accumulate in the upper input image area. As a result, information such as the image IMG shown in FIG. 4 is stored in the memory unit 330.
stored in bitmap format.
ステップ52では、入力画像における文章の行ピッチを
検出するために、縦方向における無画素の分布状態を第
4図に示すヒストグラムHGIのような形で検出する。In step 52, in order to detect the line pitch of sentences in the input image, the distribution state of non-pixels in the vertical direction is detected in the form of a histogram HGI shown in FIG.
このヒストグラムを利用して、文章の1行のピッチを検
出する。This histogram is used to detect the pitch of one line of text.
ステップ53では、検出された行ピッチに基づいて、入
力画像情報の中から1行毎に文字列情報を切り出す。In step 53, character string information is extracted line by line from the input image information based on the detected line pitch.
ステップ54では、切り出された1行毎の文字列情報に
対して、第5図に示すように2横方向における黒画素の
分布状態をヒストグラムHG2のような形で検出する。In step 54, the distribution state of black pixels in two horizontal directions is detected in the form of a histogram HG2, as shown in FIG. 5, for each line of extracted character string information.
そしてこのヒストグラムに基づいて文字間のピッチを検
出する。The pitch between characters is then detected based on this histogram.
ステップ55では、検出された文字ピッチに基づいて、
第6図に示すように、1行毎の文字列情報の中から、1
文字毎に文字情報を切り出す。In step 55, based on the detected character pitch,
As shown in Figure 6, from the character string information for each line, 1
Extract character information for each character.
ステップ56では、ステップ52で検出された行ピッチ
とステップ54で検出された文字ピッチとに基づいて、
処理対象の文章に含まれる文字の平均的な大きさを検出
する。In step 56, based on the line pitch detected in step 52 and the character pitch detected in step 54,
Detects the average size of characters included in the text to be processed.
ステップ57では、文字パターン辞書350に備わった
多数の辞書の中から、ステップ56で検出された文字サ
イズに最も適した一部の辞書を選択する。文字パターン
辞書350には様々なフォントの辞書が文字サイズ毎に
設けられているので、ここでは複数のフォントの辞書が
同時に選択される。In step 57, some dictionaries most suitable for the character size detected in step 56 are selected from among the large number of dictionaries provided in the character pattern dictionary 350. Since the character pattern dictionary 350 includes dictionaries of various fonts for each character size, dictionaries of a plurality of fonts are selected at the same time.
ステップ58では、ステップ55で切り出された文字情
報を処理して、第7図に示すように単語を抽出する。具
体的には、互いに隣り合う文字情報について、その間隔
が1文字の幅以上ある場合には、それらの間の位置を単
語の切れ目とみなし、その切れ目から切れ目までの文字
情報は同一の単語に属するものとしてグループ化する。In step 58, the character information cut out in step 55 is processed to extract words as shown in FIG. Specifically, if the distance between adjacent character information is one character or more, the position between them is considered to be a word break, and the character information from that break to the break is considered to be the same word. Group as belonging.
ステップ59では、抽出された単語に属する各々の文字
情報について、文字パターンの認識を実行する。つまり
、認識対象の文字情報の特徴パラメータを検出し、それ
を文字パターン辞書に登録された各種の文字の特徴パラ
メータと比較し、特徴が近い文字を、その文字情報の認
識結果として得る。なお、ステップ57では、複数のフ
ォントの辞書が選択されているので、それらの全てにつ
いて特徴パラメータの比較を実行する。また各々のフォ
ントの辞書について、文字認識が成功したか否かを示す
情報を保存する。In step 59, character pattern recognition is performed for each piece of character information belonging to the extracted word. That is, the feature parameters of the character information to be recognized are detected, compared with the feature parameters of various characters registered in the character pattern dictionary, and characters with similar characteristics are obtained as the recognition result of the character information. Note that in step 57, since a plurality of font dictionaries are selected, feature parameters are compared for all of them. Information indicating whether character recognition was successful or not is also stored for each font dictionary.
ステップ60では、ig*された文字の集合で構成され
る抽出された単語について、英和辞書を参照し、一致す
る英単語をみつける。一致する英単語が辞書上に存在す
る時には、それに対応付けられた日本語の情報を取り出
し、翻訳結果有とし、そうでなければ結果なしとする。In step 60, an English-Japanese dictionary is referenced to find a matching English word for the extracted word consisting of a set of ig* characters. When a matching English word exists in the dictionary, the Japanese information associated with it is retrieved and a translation result is determined; otherwise, no result is determined.
翻訳結果有の場合には、ステップ61を通って64に進
み、その結果をメモリの出力単語記憶領域にストアする
。翻訳結果なしの場合には、ステップ62において、翻
訳すべき英単語の修正の可否を判定する。つまり、第9
図に示すように、修正前の英単語に大文字が含まれる場
合にはそれを小文字に変換する。英単語の最後にカンマ
が付いている場合にはそれを削除する一単語の最後にS
が付いている場合にはそれを削除する1行末と行頭との
継続関係を示すハイフン「−」が含まれる場合にはそれ
を削除する。単語がハイフンで接続された複合語である
場合にはそれをハイフンの部分で複数に分割する等々の
予め定めた規則に基づいて、修正すべき単語か否かを調
べる。単語を修正可能な場合には、ステップ63におい
てその修正を実行し、修正された単語について再びステ
ップ60で辞書を検索し翻訳を実行する。修正が不可能
な場合には、翻訳エラーとみなし、ステップ65でエラ
ーコードを出力する。If there is a translation result, the process passes through step 61 and proceeds to 64, where the result is stored in the output word storage area of the memory. If there is no translation result, it is determined in step 62 whether or not the English word to be translated can be modified. In other words, the 9th
As shown in the figure, if the English word before correction contains uppercase letters, they are converted to lowercase letters. If there is a comma at the end of an English word, remove it. S at the end of a word
If a hyphen "-" is included, delete it. If the word is a compound word connected by a hyphen, it is checked whether the word should be corrected or not based on predetermined rules, such as dividing the word into multiple parts at the hyphen. If the word can be modified, the modification is performed in step 63, and the dictionary is searched for the modified word again in step 60 to perform translation. If correction is not possible, it is regarded as a translation error and an error code is output in step 65.
ステップ66では、翻訳対象の文章の中の最初の1行に
ついて全文字の認識が終了したか否かを調べる。1行目
を処理中であれば、認識が終了していないので、ステッ
プ67に進む。In step 66, it is checked whether recognition of all characters in the first line of the sentence to be translated has been completed. If the first line is being processed, the process proceeds to step 67 because recognition has not yet been completed.
ステップ67では、処理中の単語に含まれる文字の数と
その中で認識ができた文字の数を、使用した文字パター
ン辞書のフォント毎に計算し、その結果を記憶する。In step 67, the number of characters included in the word being processed and the number of recognized characters are calculated for each font in the character pattern dictionary used, and the results are stored.
翻訳対象の文章の中の最初の1行について全文字の認識
が終了した時には、ステップ66から68を通って69
に進む。ステップ69では、ステップ67の結果を保持
するメモリの内容を調べ1打金体についての文字認識率
をフォント間で比較し、最も認識率の高い結果が得られ
たフォントを検出する。そして、そのフォントの辞書を
選択し、以後はこの1つの辞書だけを文字認識(ステッ
プ59)で使用する。使用する辞書のフォントが決定さ
れると、以後はステップ68から70に進むのでステッ
プ69は処理されない。When recognition of all characters for the first line of the sentence to be translated is completed, the process passes through steps 66 to 68 and returns to step 69.
Proceed to. In step 69, the contents of the memory holding the results of step 67 are checked, and the character recognition rate for one font is compared between fonts, and the font with the highest recognition rate is detected. Then, a dictionary for that font is selected, and from then on only this one dictionary is used for character recognition (step 59). Once the dictionary font to be used is determined, the process proceeds from step 68 to step 70, so step 69 is not processed.
ステップ70では、画像全体について文字の認識及び単
語の翻訳が終了したか否かを判定し、終了していなけれ
ば、ステップ58に戻って上記の処理を繰り返す。In step 70, it is determined whether or not character recognition and word translation have been completed for the entire image. If not, the process returns to step 58 and the above process is repeated.
画像全体について文字の認識及び単語の翻訳が終了する
と、ステップ70から71に進む。ステップ71では、
メモリユニット330上の出力画像領域に、例えば第8
a図に示すような出力画像情報をビットマツプ形式で作
成する。つまり、ステップ64で出力された翻訳結果が
記憶されたメモリ領域を参照して翻訳結果を順番に取り
出し、各々の翻訳結果を文字パターンを示す画像情報に
展開して順次にメモリに書込む。なお第8a図に示され
る「?」は、認識不可能であった文字又は英和辞書に存
在しなかった単語を示している。When character recognition and word translation for the entire image are completed, the process proceeds from step 70 to step 71. In step 71,
In the output image area on the memory unit 330, for example, the eighth
Output image information as shown in Figure a is created in bitmap format. That is, the memory area in which the translation results outputted in step 64 are stored is referred to, and the translation results are taken out in order, and each translation result is developed into image information representing a character pattern and sequentially written into the memory. Note that "?" shown in FIG. 8a indicates a character that could not be recognized or a word that did not exist in the English-Japanese dictionary.
ステップ72では、ステップ71で作成された出力画像
情報を、プリンタに出力して出力画像のハードコピーを
作成する。In step 72, the output image information created in step 71 is output to a printer to create a hard copy of the output image.
この実施例では、操作ボード310からのキー操作によ
って、出力される画像の内容を切換えることができる。In this embodiment, the contents of the image to be output can be switched by key operations from the operation board 310.
この切換えによって、出ツノ画像として、第8a図、第
8b図、第8c図及び第8d図に示すような4種類の中
の1つの出力形式を選択することができる。By this switching, it is possible to select one of the four output formats as shown in FIGS. 8a, 8b, 8c, and 8d as the horn image.
第8b図の出力画像においては、文章の1行毎に、入力
画像のコピー像とそれの翻訳結果とが位置が対応するよ
うに上下に配置されている。また第8c図の出力画像に
おいては、文章の1行毎に、入力画像のコピー像、それ
を文字認識した結果。In the output image of FIG. 8b, the copy image of the input image and its translation result are arranged one above the other so that their positions correspond to each other for each line of text. The output image shown in FIG. 8c is a copy of the input image and the result of character recognition for each line of text.
及びそれを翻訳した結果が互いに位置が対応するように
上下に配置されている。同様に第8d図では、認識され
た文字の画像と翻訳の結果とが1行毎に配置されている
。and the results of its translation are arranged one above the other so that their positions correspond to each other. Similarly, in FIG. 8d, images of recognized characters and translation results are arranged line by line.
[効果]
以上のとおり本発明によれば、次のような効果が得られ
る。[Effects] As described above, according to the present invention, the following effects can be obtained.
請求項1:
原稿画像上から自動的に文字及び単語を認識し、認識し
た単語毎に翻訳を実行してその結果を出力画像の形で出
力するので、文字入力のためにキーボードを操作する必
要がなく、翻訳作業は普通に複写機を使うのと同様に、
翻訳対象の原稿をイメージスキャナにセットしてコピー
ボタンを押す、という極めて簡単な作業になる。Claim 1: Automatically recognizes characters and words from the original image, executes translation for each recognized word, and outputs the results in the form of an output image, so there is no need to operate the keyboard to enter characters. There is no need for translation, and the translation work is just like using a copy machine.
The process is extremely simple: place the manuscript to be translated into the image scanner and press the copy button.
請求項2
認識対象の文字に適した大きさ及びフォントの辞書を自
動的に選択でき、認識率が向上し、しかも文字認識の所
要時間も短い。Claim 2: A dictionary of size and font suitable for the characters to be recognized can be automatically selected, the recognition rate is improved, and the time required for character recognition is shortened.
請求項3コ
オペレータの修正操作を必要とすることなしに、翻訳の
成功率を向上することができる。Claim 3: The success rate of translation can be improved without requiring correction operations by co-operators.
第1図は、本発明を実施する一形式のデジタル複写機の
機構部の構成を示す正面図である。
第2図は、第1図の複写機の電装部の主要部の構成を示
すブロック図である。
第3図は、主として第2図の翻訳ユニット340の処理
の内容を示すフローチャートである。
第4図は入力画像と縦方向の黒画素分布のヒストグラム
を示す平面図、第5図は1行の入力画像とその中の横方
向の黒画素分布のヒストグラムを示す平面図、第6図は
文字毎に区分された画像を示す平面図、第7図は単語毎
にグループ化された文字の画像を示す平面図である。
第8a図、第8b図、第8c図及び第8d図は、各々1
作成された出力画像の例を示す平面図である。
第9図は、修正前の単語と修正後の単語との対応を示す
ブロック図である。
1:露光ランプ 2:レンズ
3コ受光部 4:書込装置
5:感光体ドラム 6:メインチャージャ7:現像
ユニット 8:分離チャージャ9:転写チャージャ
10:クリーニングユニット
12.13:給紙カセット
100:イメージスキャナ
110:CCDイメージセンナ
120:A/D変換器 200:レーザプリンタ310
:操作ボード 330:メモリユニット340:翻訳
ユニット
350:文字パターン辞書
360:英和辞書FIG. 1 is a front view showing the structure of a mechanical section of a digital copying machine of one type that implements the present invention. FIG. 2 is a block diagram showing the configuration of the main parts of the electrical equipment section of the copying machine shown in FIG. 1. FIG. 3 is a flowchart mainly showing the contents of the processing of the translation unit 340 of FIG. 2. Figure 4 is a plan view showing the input image and a histogram of the black pixel distribution in the vertical direction, Figure 5 is a plan view showing the input image in one row and the histogram of the black pixel distribution in the horizontal direction, and Figure 6 is a plan view showing the histogram of the black pixel distribution in the horizontal direction. FIG. 7 is a plan view showing an image of characters grouped by word. FIG. Figures 8a, 8b, 8c and 8d are each 1
FIG. 3 is a plan view showing an example of a created output image. FIG. 9 is a block diagram showing the correspondence between words before correction and words after correction. 1: Exposure lamp 2: Lens 3 light receiving section 4: Writing device 5: Photosensitive drum 6: Main charger 7: Developing unit 8: Separation charger 9: Transfer charger 10: Cleaning unit 12.13: Paper feed cassette 100: Image scanner 110: CCD image sensor 120: A/D converter 200: Laser printer 310
: Operation board 330: Memory unit 340: Translation unit 350: Character pattern dictionary 360: English-Japanese dictionary
Claims (3)
読取手段、該手段の読取った情報を蓄積するメモリ手段
、及び該メモリ手段に蓄積された画像情報を出力する出
力手段、を含むデジタル複写装置において: 前記メモリ手段に蓄積された入力画像情報から文字情報
を各文字毎に切り出す文字切り出し手段; 切り出された文字の間隔に基づいて、1つの文字もしく
は複数文字の集合でなる単語毎に文字情報を区分する単
語抽出手段; 文字毎の特徴量が予め登録された文字パターン辞書を含
み、文字毎に切り出された画像情報から、それが示す文
字を認識する文字認識手段;単語毎の翻訳結果が登録さ
れた単語辞書を 含み、前記単語抽出手段によって抽出された単語毎に、
それを構成する文字の構成に応じて単語辞書を参照し、
各単語の翻訳結果を認識する、翻訳手段;及び 翻訳手段の翻訳結果として得られる文字情報を、前記メ
モリ手段の出力画像領域に画像情報として書込む、出力
画像作成手段; を設けたことを特徴とする、デジタル複写装置。(1) A digital copying device that includes an image reading means for reading a document image as a set of minute pixel information, a memory means for storing the information read by the means, and an output means for outputting the image information stored in the memory means. In: character cutting means for cutting out character information for each character from the input image information stored in the memory means; character information for each word consisting of one character or a set of multiple characters based on the interval between the cut out characters; word extraction means for classifying; character recognition means that includes a character pattern dictionary in which feature amounts for each character are registered in advance; character recognition means for recognizing characters indicated by image information cut out for each character; Including a registered word dictionary, for each word extracted by the word extraction means,
Refer to the word dictionary according to the composition of the letters that make it up,
Translation means for recognizing the translation result of each word; and output image creation means for writing character information obtained as a translation result of the translation means into the output image area of the memory means as image information. A digital copying device.
の違いに応じた複数種類の文字パターン辞書を含み、文
字を切り出す時に検出される行ピッチと文字ピッチとに
基づいて認識対象文字の大きさを識別しその結果により
使用する文字パターン辞書を限定するとともに、原稿画
像中の一部分の文字列に対して複数種類の文字パターン
辞書を利用して認識を実施し、その結果において文字認
識率が高い文字パターン辞書を自動的に選択して以降の
認識を実行する、前記請求項1記載のデジタル複写装置
。(2) The character recognition means includes a plurality of character pattern dictionaries corresponding to differences in character size and font, and determines the size of the character to be recognized based on the line pitch and character pitch detected when cutting out the character. In addition to identifying the character strings and limiting the character pattern dictionary to be used based on the results, recognition is performed using multiple types of character pattern dictionaries for a part of the character string in the original image, and the character recognition rate is 2. The digital copying apparatus according to claim 1, wherein a high character pattern dictionary is automatically selected for subsequent recognition.
に修正モードの処理に移行し、翻訳対象の単語文字列毎
に、その一部分の文字を予め定められた規則に基づいて
修正し、再び翻訳を実行する、前記請求項1記載のデジ
タル複写装置。(3) When the translation means cannot detect a translation result, it automatically shifts to correction mode processing, corrects some characters for each word string to be translated based on predetermined rules, and then repeats the process again. 2. The digital reproduction apparatus of claim 1, wherein the digital reproduction apparatus performs translation.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2192068A JPH0477965A (en) | 1990-07-20 | 1990-07-20 | digital copying machine |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2192068A JPH0477965A (en) | 1990-07-20 | 1990-07-20 | digital copying machine |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH0477965A true JPH0477965A (en) | 1992-03-12 |
Family
ID=16285095
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP2192068A Pending JPH0477965A (en) | 1990-07-20 | 1990-07-20 | digital copying machine |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH0477965A (en) |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0756924A (en) * | 1993-06-30 | 1995-03-03 | Ricoh Co Ltd | Bilingual device |
| US5748805A (en) * | 1991-11-19 | 1998-05-05 | Xerox Corporation | Method and apparatus for supplementing significant portions of a document selected without document image decoding with retrieved information |
| JP2005014237A (en) * | 2003-06-23 | 2005-01-20 | Toshiba Corp | Copying machine having translation method, program, and external translation function in copying machine |
-
1990
- 1990-07-20 JP JP2192068A patent/JPH0477965A/en active Pending
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5748805A (en) * | 1991-11-19 | 1998-05-05 | Xerox Corporation | Method and apparatus for supplementing significant portions of a document selected without document image decoding with retrieved information |
| JPH0756924A (en) * | 1993-06-30 | 1995-03-03 | Ricoh Co Ltd | Bilingual device |
| JP2005014237A (en) * | 2003-06-23 | 2005-01-20 | Toshiba Corp | Copying machine having translation method, program, and external translation function in copying machine |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US5517409A (en) | Image forming apparatus and method having efficient translation function | |
| JP3729017B2 (en) | Image processing device | |
| JP3251959B2 (en) | Image forming device | |
| US8126270B2 (en) | Image processing apparatus and image processing method for performing region segmentation processing | |
| US5729618A (en) | Image forming apparatus for outputting equivalents of words spelled in a foreign language | |
| JP3050007B2 (en) | Image reading apparatus and image forming apparatus having the same | |
| JPH0844827A (en) | Digital copier | |
| JP5594269B2 (en) | File name creation device, image forming device, and file name creation program | |
| JPH09270902A (en) | Image filing method and image filing device | |
| JP2000050061A (en) | Image processor | |
| KR100306063B1 (en) | Image processing method and apparatus | |
| JPH0477965A (en) | digital copying machine | |
| JP3269842B2 (en) | Bilingual image forming device | |
| JPH05266074A (en) | Bilingual image forming device | |
| JPH11213089A (en) | Image processing apparatus and method | |
| JPH05324709A (en) | Bilingual image forming device | |
| JP3629959B2 (en) | Image recognition device | |
| JPH11220557A (en) | Image processing apparatus and method | |
| JPH0554069A (en) | Digital translator | |
| JPH0554074A (en) | Digital copying machine | |
| JP2001312697A (en) | Image direction determination method and apparatus | |
| JPH0554188A (en) | Picture processor | |
| JP3675181B2 (en) | Image recognition device | |
| JPH0546659A (en) | Digital translation / copying device | |
| JPH01144181A (en) | Optical character reader |