JPH0432960A - Reading processor - Google Patents
Reading processorInfo
- Publication number
- JPH0432960A JPH0432960A JP2132791A JP13279190A JPH0432960A JP H0432960 A JPH0432960 A JP H0432960A JP 2132791 A JP2132791 A JP 2132791A JP 13279190 A JP13279190 A JP 13279190A JP H0432960 A JPH0432960 A JP H0432960A
- Authority
- JP
- Japan
- Prior art keywords
- reading
- section
- color
- image
- document
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
- 238000000605 extraction Methods 0.000 claims abstract description 32
- 239000000284 extract Substances 0.000 claims abstract description 5
- 230000015572 biosynthetic process Effects 0.000 claims description 9
- 238000003786 synthesis reaction Methods 0.000 claims description 9
- 230000003993 interaction Effects 0.000 claims description 5
- 239000003086 colorant Substances 0.000 abstract description 2
- 238000000034 method Methods 0.000 description 18
- 238000010586 diagram Methods 0.000 description 15
- 230000008569 process Effects 0.000 description 8
- 230000008859 change Effects 0.000 description 3
- 230000000694 effects Effects 0.000 description 3
- 238000006243 chemical reaction Methods 0.000 description 2
- 101100136092 Drosophila melanogaster peng gene Proteins 0.000 description 1
- 230000004913 activation Effects 0.000 description 1
- 230000001174 ascending effect Effects 0.000 description 1
- 238000012790 confirmation Methods 0.000 description 1
- 210000004709 eyebrow Anatomy 0.000 description 1
- 230000006870 function Effects 0.000 description 1
- 230000003287 optical effect Effects 0.000 description 1
Landscapes
- Document Processing Apparatus (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
Abstract
Description
【発明の詳細な説明】
〔産業上の利用分野〕
本発明は文書、書籍上の文字を認識し、認識結果に基づ
いて朗読音声を出力する読書処理装置に関するものであ
る。DETAILED DESCRIPTION OF THE INVENTION [Field of Industrial Application] The present invention relates to a reading processing device that recognizes characters on a document or book and outputs a reading voice based on the recognition result.
従来、この種の装置としては、電子情報通信学会技術研
究報告 信学技報Vo1.87N0.334 1988
年1月22日発行 P79〜86に開示されたものがあ
った。この装置は第2図に示すような構成であり、画像
入力部1で文書、書籍をページ単位で入力し、この入力
された画像に対して画像処理部2で傾き補正などの画像
処理を行ない、該画像処理部2から出力され文書画像文
字の切出し、認識を文字認識処理部3で行ない、その後
言語処理部4で言語処理を施し、音声合成部5で音声に
変換し、音声出力部6から音声を出力するようになって
いる。Conventionally, this type of device was published in IEICE Technical Report Vol. 1.87 No. 334 1988
Published on January 22, 2016, there was something disclosed on pages 79-86. This device has a configuration as shown in Fig. 2, in which an image input section 1 inputs a document or book page by page, and an image processing section 2 performs image processing such as tilt correction on the input image. , the document image characters outputted from the image processing section 2 are cut out and recognized by the character recognition processing section 3, then subjected to language processing by the language processing section 4, converted into speech by the speech synthesis section 5, and then processed by the speech output section 6. It is designed to output audio from.
しかしながら上記構成の従来の読書処理装置では、健常
者が文書、書籍を読む時、日常的に行なっている下記の
ような読み方の選択が不可能であった。例えば、読書対
象の文書、書籍が論文などの場合、
(1)11名や執筆者などの書誌的事項を読む(2)要
約を証む
(3)章2節等の題名を読む
(4)参考文献を読む
(5)上記を知った上で全文を読むか読まないかを判断
する。However, with the conventional reading processing device having the above-mentioned configuration, it has been impossible for healthy people to select the reading method described below, which they do on a daily basis when reading documents and books. For example, if the document or book to be read is an essay, (1) Read the bibliographic information such as 11 names and authors (2) Prove the summary (3) Read the title of chapter 2, etc. (4) Read the references (5) After knowing the above, decide whether to read the entire text or not.
また、盲人や肢体障害者などの障害者にとって良い読書
環境であるとは言えなかった。In addition, it could not be said that it was a good reading environment for people with disabilities such as blind people and people with physical disabilities.
本発明は上述の点に鑑みてなされたもので上記問題点を
除去するため、文書、書籍のカラーマーク部分の文章を
自由に選択して読み取ることができるようにし、晴眼者
が日常的に行なっている文書書籍の拾い読みに近い形の
読み方が出来、障害者の読書環境の向上に役立つ読書処
理装置を提供することを目的とする。The present invention has been made in view of the above-mentioned problems, and in order to eliminate the above-mentioned problems, it is possible to freely select and read the text in the color mark part of documents and books, so that sighted people can read it on a daily basis. To provide a reading processing device capable of reading in a manner similar to the browsing method of reading documents and books, which is useful for improving the reading environment for handicapped people.
上記課題を解決するため本発明は 画像入力部と画像処
理部と文字認識部と言語処理部と音声合成部を有し、文
字認識結果に基づいて朗読音声を出力する読書処理装置
において、
画像入力部として、用紙上に記載されている文字0図形
、写真及び/又はマーク等の画像の濃淡及び色の一方又
は双方に応じた情報を出力する画像入力部を用い、
画像処理部として、所定の色特性と読書領域名称とを一
対の読書情報として設定するカラーコマンド設定部と、
音声対話処理によって読書領域名称を指定する読書モー
ド指定部と、該読書モード指定部で指定された読書領域
名称に基づいて読書情報の色特性を読取色特性として出
力する抽出モード切換部と、該読取色特性に基づいて文
書。In order to solve the above problems, the present invention provides a reading processing device that includes an image input section, an image processing section, a character recognition section, a language processing section, and a speech synthesis section, and outputs a reading voice based on the result of character recognition. As an image processing section, an image input section is used that outputs information corresponding to one or both of the shading and color of an image such as a character 0 figure, photograph, and/or mark written on a sheet of paper.As an image processing section, a predetermined a color command setting section that sets a color characteristic and a reading area name as a pair of reading information;
a reading mode specifying section that specifies a reading area name through voice interaction processing; an extraction mode switching section that outputs a color characteristic of reading information as a reading color characteristic based on the reading area name specified by the reading mode specifying section; Read documents based on color characteristics.
書籍上に記載されているマーク領域内の文書画像を抽出
する画像抽出部とを具備する画像処理部を用いたことを
特徴とする。The present invention is characterized by using an image processing section that includes an image extraction section that extracts a document image within a mark area written on a book.
本発明の読書処理装置は上記の如く画像入力部及び画像
処理部を構成するので、文書、書籍の所望の部分にカラ
ーマークを追記又は印刷し千おき、該文書、書籍のカラ
ーマーク部分の文章或いは全文を音声ガイダンス及び音
声認識で自由に選択できる読書処理装置となる。従って
、本読書処理装置を使用すると、製本業者が印刷を工夫
するか、健常者が一度追記マークを施した文書、書籍で
あれば、健常者が日常的に行なっている文書。Since the reading processing device of the present invention has an image input unit and an image processing unit as described above, it adds or prints a color mark to a desired part of a document or book, and prints the text in the color mark part of the document or book. Alternatively, it becomes a reading processing device that can freely select the entire text using voice guidance and voice recognition. Therefore, if you use this book reading processing device, you can print documents or books that have been printed by a bookbinder or once added by an able-bodied person, and which are printed on a daily basis by a able-bodied person.
書籍の拾い読みに近い形の読み方ができるから、盲人や
肢体障害者等の障害者の読書環境の向上に役立つ、また
健常者において聴覚による読書環境の提供が可能となる
。Since it is possible to read in a manner similar to browsing a book, it is useful for improving the reading environment for people with disabilities such as blind people and physically disabled people, and it is also possible to provide an auditory reading environment for healthy people.
以下、本発明の一実施例を図面に基づいて説明する。 Hereinafter, one embodiment of the present invention will be described based on the drawings.
第1図は本発明の実施例である読書処理装置の構成を示
すブロック図である。図示するように、本読書処理装置
は、画像入力部1、画像処理部2、文字認識部3、言語
処理部4、音声合成部5及び音声出力部6から構成され
る。FIG. 1 is a block diagram showing the configuration of a reading processing device that is an embodiment of the present invention. As shown in the figure, the reading processing device includes an image input section 1, an image processing section 2, a character recognition section 3, a language processing section 4, a speech synthesis section 5, and a speech output section 6.
画像処理部2はカラーコマンド設定部21、読書モード
指定部22、抽出モード切換部23及び画像抽出部24
から構成される。The image processing section 2 includes a color command setting section 21, a reading mode specifying section 22, an extraction mode switching section 23, and an image extraction section 24.
It consists of
また、カラーコマンド設定部21は学習/読書モード切
換部211及びカラーコマンド学習部212を具備し、
読書モード指定部22は音声入力部221、音声認識部
222及び読書モード設定部223を具備する。Further, the color command setting section 21 includes a learning/reading mode switching section 211 and a color command learning section 212,
The reading mode specifying section 22 includes a voice input section 221, a voice recognition section 222, and a reading mode setting section 223.
画像入力部1は、第4図、第5図に示すような用紙を主
走査を水平方向(XX)、副走査を垂直方向(YY)と
して走査し、用紙上に記載されているマーク及び/又は
文書9図面の濃淡及び色の一方又は双方に応じた色情報
、例えば濃度多値レベル信号又はR(赤)、G(緑)、
B(青)の多値レベル信号を色情報としてカラーコマン
ド設定部21へ出力する。The image input unit 1 scans a sheet of paper as shown in FIGS. 4 and 5 with main scanning in the horizontal direction (XX) and sub-scanning in the vertical direction (YY), and records marks and/or marks written on the paper. or Document 9 Color information corresponding to one or both of the shading and color of the drawing, such as a density multilevel signal or R (red), G (green),
A B (blue) multilevel signal is output to the color command setting section 21 as color information.
カラーコマンド設定部21の学習/読書モード切換部2
11は、図示しないスイッチ等によって学習モードと読
書モードとを切換える切換部であって、該学習/読書モ
ード切換部211の切換えが学習モードにある場合には
画像入力部1の出力をカラーコマンド学習部212に出
力すると共に、学習起動信号を出力する。また、切換え
が読書モードにある場合には、画像入力部1の出力を画
像抽出部24に出力すると共に、画像抽出起動信号を出
力する。Learning/reading mode switching section 2 of color command setting section 21
Reference numeral 11 denotes a switching unit that switches between a learning mode and a reading mode using a switch (not shown) or the like, and when the learning/reading mode switching unit 211 is in the learning mode, the output of the image input unit 1 is switched to the color command learning mode. 212, and also outputs a learning start signal. Furthermore, when the mode is switched to reading mode, the output of the image input section 1 is outputted to the image extraction section 24, and an image extraction activation signal is outputted.
カラーコマンド学習部212は、第4図に示すようなカ
ラーコマンド学習用紙の所定領域の色特性を学習し、当
該所定の領域に定義されている読書領域名称或いは前記
所定の領域に対応して記載きれている文字列から決まる
読書領域名称に対応する文字コード列と前記色特性とを
一対の読書情報として設定する。The color command learning unit 212 learns the color characteristics of a predetermined area of the color command study paper as shown in FIG. A character code string corresponding to a reading area name determined from a blank character string and the color characteristics are set as a pair of reading information.
前述のカラーコマンド設定部21は画像入力部1から出
力されるカラーコマンド学習用紙上の色特性及び読書領
域名称を学習して読書情報を設定するように構成したが
、画像入力部1から出力きれる色情報を用いるのではな
く、予め定められた色特性と読書領域名称の中から選択
して読書情報を設定するように構成してもよい。The color command setting section 21 described above is configured to set reading information by learning the color characteristics and reading area name on the color command learning paper output from the image input section 1, but the color command setting section 21 is configured to set reading information by learning the color characteristics and reading area name on the color command learning paper output from the image input section 1. Instead of using color information, reading information may be set by selecting from predetermined color characteristics and reading area names.
読書モード指定部22は上記のように、音声入力部22
1と音声認識部222で構成され、音声対話処理によっ
て所望の読書領域名称に対応する文字フード列を出力す
る。音声入力部221は読書希望者が発声する読書領域
名称、音声対話処理のための音声などを音声信号に変換
する。音声認識部222は前記音声信号を認識し、文字
コード列を出力する。As mentioned above, the reading mode designation section 22 is connected to the voice input section 22.
1 and a voice recognition unit 222, and outputs a character string corresponding to a desired reading area name through voice interaction processing. The voice input unit 221 converts the reading area name, voice for voice dialogue processing, etc. uttered by the person who wishes to read into voice signals. The voice recognition unit 222 recognizes the voice signal and outputs a character code string.
なお、音声認識部222は読書領域名称及び音声対話処
理に用いられる言葉を誤りなく認識するために音声認識
用パラメータを保有する。Note that the speech recognition unit 222 retains speech recognition parameters in order to accurately recognize reading area names and words used in speech dialogue processing.
読書モード設定部223は音声対話処理によって読書領
域名称を設定するために、下記の処理及び制御を行なう
。The reading mode setting unit 223 performs the following processing and control in order to set the reading area name by voice interaction processing.
■音声認識部222から出力される文字コード列と、対
話処理用の文字コード列との照合、■上記■の照合結果
に基づいて音声ガイダンスのための文字フード列を音声
合成部5へ出力、■音声認識部222から出力される文
字コード列と、カラーコマンド設定部21に設定されて
いる読書領域名称に対応する文字コード列との照合、
■上記■の照合結果に基づいて読書領域名称に対応する
文字コード列を出力する。■Comparing the character code string output from the speech recognition unit 222 with the character code string for dialogue processing, ■Outputting a character food string for voice guidance to the speech synthesis unit 5 based on the matching result of (■) above, ■ Matching the character code string output from the voice recognition unit 222 with the character code string corresponding to the reading area name set in the color command setting unit 21, ■ Based on the matching result of above ■, reading area name Outputs the corresponding character code string.
抽出モード切換部23は、読書モード指定部22から出
力される読書領域名称に対応する文字コード列に基づい
て前記読書情報の色特性を読取色特性として画像抽出部
24に出力する。The extraction mode switching unit 23 outputs the color characteristics of the reading information to the image extraction unit 24 as reading color characteristics based on the character code string corresponding to the reading area name output from the reading mode designation unit 22.
画像抽出部24は、前記読取色特性に基づいて文書、書
籍上に記載されているマーク領域内の文書画像を読書領
域画像として抽出する。The image extraction unit 24 extracts a document image within a mark area written on a document or book as a reading area image based on the reading color characteristics.
抽出された読書領域画像は従来技術と同様に傾き補正な
どの処理を経て1文字毎、文字画像を切り出す。The extracted reading area image is subjected to processes such as tilt correction as in the prior art, and character images are cut out character by character.
文字認識処理部3は切出きれた文字画像を順次認識し、
文字コード列を作成する。The character recognition processing unit 3 sequentially recognizes the cut out character images,
Create a character code string.
音声合成部5は、前記文字コード列及び読書モード設定
部から出力される音声ガイダンス用の文字コード列を音
声波形信号に変換する。The speech synthesis section 5 converts the character code string and the character code string for voice guidance outputted from the reading mode setting section into an audio waveform signal.
音声出力部6は、前記音声合成部5からの音声波形信号
を音声に変換し、朗読音声及びガイダンス音声として出
力する。The audio output unit 6 converts the audio waveform signal from the audio synthesis unit 5 into audio and outputs it as reading audio and guidance audio.
第3図は文書、書籍上のマークを色毎に分類するための
原理を示す色特性の説明図であって、同図(a)はマー
クの色座標の範囲を示す色度図、同図(b)はマークの
濃度範囲の説明図である。Fig. 3 is an explanatory diagram of color characteristics showing the principle for classifying marks on documents and books by color, and Fig. 3 (a) is a chromaticity diagram showing the range of color coordinates of marks; (b) is an explanatory diagram of the density range of the mark.
第3図(a)に示す原理を用いる場合、画像入力部1に
は、例えばRGB系の濃度多値レベル信号を出力できる
カラースキャナが必要である。When using the principle shown in FIG. 3(a), the image input section 1 requires a color scanner capable of outputting, for example, an RGB density multilevel signal.
第3図(b)に示す原理を用いる場合、画像入力部1は
濃度多値レベル信号を出力できるモノクロスキャナ又は
上記カラースキャナが必要である。When using the principle shown in FIG. 3(b), the image input section 1 requires a monochrome scanner or the above-mentioned color scanner capable of outputting a multilevel density signal.
第3図(a)の色度図は公知の方法により作成したもの
を概略的に示したものであって、RGB系の濃度信号を
座標変換式(1)を用いてXYZ系に変換し、
X −2,7689R+ 1.7517 G + 1.
1302 BY = R+ 4.5907 G
+ 0.0601 BZ= 0.056
5G+5.5943B (1)求められた、X、
Y、Zをもとに式(2)を用いて色度座標x、yを計算
し、
X雪X/(X+Y+Z)
y−y/(x+y+z) (2)求め
られた色座標x + Vによる直交座標を用いたもので
ある。前記X、yは色相と彩度を表すものであり、Yは
明度を表わすものである。The chromaticity diagram in FIG. 3(a) is a schematic diagram created by a known method, in which the RGB system density signal is converted to the XYZ system using coordinate conversion formula (1), X −2,7689R+ 1.7517 G + 1.
1302 BY = R+ 4.5907 G
+ 0.0601 BZ= 0.056
5G+5.5943B (1) Obtained,
Calculate the chromaticity coordinates x and y using formula (2) based on Y and Z, and calculate the chromaticity coordinates x and y as follows: It uses orthogonal coordinates. The X and y represent hue and saturation, and Y represents lightness.
上述のRGB系からXYZ系への変換C<(1)を用い
る〕色相と彩度を表わすx、yを求める計算〔式(2)
を用いる〕はそれぞれR,G、Bを入力し、x、y、z
を出力するROMを用いて、及びX、Y、Zを入力し、
x t 3’を整数形式で出力するROMを用いて変換
することによって、計算時間が不要となり、処理の高速
化がはかれる。Conversion C from the above RGB system to the XYZ system <Use (1)] Calculation to obtain x and y representing hue and saturation [Equation (2)
], input R, G, B, x, y, z
using a ROM that outputs and inputs X, Y, Z,
By converting x t 3' using a ROM that outputs it in an integer format, no calculation time is required and processing speed can be increased.
第3図(b)の場合、領域抽出対象の文書、書籍が白、
黒で表現されたものであって、文書、書籍上のマークの
濃度レベルが前記白、黒の濃度と重ならない濃度レベル
であって、前記マークの色が色毎に互いに重ならない濃
度レベルを持つ色を選択することによって色の分類が可
使となる。また、前記Yを同様に表現してもよい。In the case of Fig. 3(b), the document or book to be extracted is white;
It is expressed in black, and the density level of the mark on the document or book is a density level that does not overlap with the density of the white or black, and the color of the mark has a density level that does not overlap with each other for each color. Color classification becomes available by selecting a color. Moreover, the above Y may be expressed similarly.
第4図はカラーコマンド学習用紙の例を示す図であって
、同図(a)は読書領域名称を記入するための文字記入
枠m、1.と読書領域名称に対応するカラーマークを記
入するためのカラーコマンド記入枠m3+とで構成され
るカラーコマンドを垂直方向に並べたものである。FIG. 4 is a diagram showing an example of a color command learning sheet, in which (a) is a character entry frame m, 1. and a color command entry frame m3+ for entering a color mark corresponding to the reading area name, which are vertically arranged.
また、第4図(b)は読書領域名称を記入するための文
字記入枠m*、、と読書領域名称に対応するカラーマー
クを記入するためのカラーコマンド記入枠を一体化した
カラーコマンドを垂直方向に並べたものである。In addition, Fig. 4(b) shows a vertical color command that integrates a character entry frame m* for entering the reading area name and a color command entry frame for entering the color mark corresponding to the reading area name. They are arranged in the direction.
また、第4図(C)はカラーコマンド記入枠mHの位置
によっ読書領域名称が一意的に決定されるものである。Further, in FIG. 4(C), the reading area name is uniquely determined by the position of the color command entry frame mH.
第4図において、記入枠は黒で印刷されているものとす
る。In FIG. 4, it is assumed that the entry frame is printed in black.
なお、前記記入枠の色、大きさ、形状、数及び位置関係
は第4図に限定きれるものではなく、読書領域名称は略
称であってもよい。Note that the color, size, shape, number, and positional relationship of the entry frames are not limited to those shown in FIG. 4, and the reading area name may be an abbreviation.
第4図において、カラーコマンド学習用紙に記載したマ
ークg1は赤、g!は緑、g、は青、g4は黄とする。In Figure 4, the mark g1 written on the color command learning sheet is red, g! is green, g is blue, and g4 is yellow.
第3図(a)には前記マークの色度図上での位置を概略
的に示している。FIG. 3(a) schematically shows the position of the mark on the chromaticity diagram.
第5図は論文形式の文書の所望の読書領域に健常者が追
記マークを施した入力文書の例であって、m、は書誌的
事項の領域に追記した赤(g、)のカラーマーク、m、
は要約の領域に追記した緑(g、)のカラーマーク、m
st 、 m、、は章、節の題名の領域に追記した青(
g、)のカラーマークを示し、第4図のカラーコマンド
学習用紙の記入内容に対応して追記したものである。Figure 5 is an example of an input document in which an able-bodied person has added additional marks in the desired reading area of a paper-style document, where m is a red (g) color mark added in the area of bibliographic items; m,
is the green (g,) color mark added in the summary area, m
st, m,, are added in blue (
g,), which have been added to correspond to the contents filled in on the color command study form in Fig. 4.
次にカラーコマンド学習部212の処理について説明す
る。画像入力部1が第4図に示すカラーコマンド学習用
紙を主走査を水平方向(XX)、副走査方向(YY)と
して走査し、走査点がカラーコマンド記入枠m、1或い
は第4図(b)の文字記入枠m、1.の領域内にあって
、画像入力部1の出力が有彩色と判定された場合には、
画像入力部1から出力されるR、G、Hの多値レベル信
号を前述の式(1) 、 (2)を用いて色座標x、y
に変換し、当該X + 3’を整数化したX%、y%を
アドレスとするカラーマツプ上に読書領域名称番号を書
き込む。第4図(a)、(b)に示すように読書領域名
称番号が決まっていない場合は、カラーコマンド記入枠
の位置順位、例えばYY座標の昇順に付与した番号を用
いる。Next, the processing of the color command learning section 212 will be explained. The image input unit 1 scans the color command learning sheet shown in FIG. 4 with the main scanning in the horizontal direction (XX) and the sub-scanning direction (YY), and the scanning point is in the color command entry frame m, 1 or in the color command entry frame in FIG. 4 (b). ) character entry frame m, 1. If the output of the image input unit 1 is determined to be a chromatic color within the area of
The R, G, and H multi-level signals output from the image input unit 1 are converted into color coordinates x, y using the above equations (1) and (2).
Then, the reading area name number is written on the color map using X% and y%, which are obtained by converting X+3' into integers, as addresses. As shown in FIGS. 4(a) and 4(b), if the reading area name number has not been determined, the number assigned in the position order of the color command entry frame, for example, in ascending order of YY coordinates, is used.
前述の有彩色の判定は公知の方法、例えばR2O,Hの
信号レベルが各々異なっていることによって行なう。The aforementioned chromatic color determination is performed by a known method, for example, by making the signal levels of R2O and H different.
次に、第4図(a)、(b)に示すカラーコマンド学習
用紙のように読書領域名称を学習する場合には、文字記
入枠内の画像を抽出して、公知の文字認識手段を用いて
認識し、認識された文字列を読書領域名称とし、当該文
字記入枠に対応するカラーコマンド記入枠に対して付与
されている読書領域名称番号と読書領域名称と前述のカ
ラーマツプを読書情報として、カラーコマンド設定部2
1のカラーコマンド学習部212に登録する。Next, when learning reading area names as shown in the color command learning sheets shown in Figures 4(a) and (b), the image within the character entry frame is extracted and a known character recognition method is used to extract the image within the character entry frame. The recognized character string is used as a reading area name, and the reading area name number and reading area name assigned to the color command entry frame corresponding to the character entry frame, and the above-mentioned color map are used as reading information. Color command setting section 2
1 in the color command learning section 212.
次に、読書モード指定部22及び本読書処理装置を動作
させるための処理について第6図を用いて説明する。Next, a process for operating the reading mode specifying section 22 and the present reading processing device will be described using FIG. 6.
なお、第6図において、「 」内のひらがなで示す部分
は読書希望者が発声する音声認識の対象の音声コマンド
であって、対話処理用の言葉と読書領域名称を示す。′
」内のカタカナで示す部分はあらかじめ読書モード指
定部22の読書モード設定部223に登録されている対
話処理用の文字コード列及びカラーコマンド設定部21
のカラーコマンド学習部212に設定されている読書領
域名称に対応する文字コード列が音声合成部5、音声出
力部6を経て出力される合成音声を示す。先ずどのよう
な読書モードが可能かを知るため音声ガイダンスを要求
するために1領域」という音声を音声入力部221より
入力し、音声認識部222で「りよういき」と認識され
ると(ステップSTI 、5T2)、カラーコマンド設
定部21のカラーコマンド学習部212に設定きれてい
る読書領域名称に対応する文字コード列、例えば「書誌
的事項」、′要約」、1章2節の題名ヨ、・・・・「全
文」などが出力される。In FIG. 6, the part shown in hiragana in parentheses is the voice command to be voice recognized by the person who wishes to read, and indicates the words for dialogue processing and the name of the reading area. ′
The part indicated in katakana in "" is the character code string for interaction processing registered in advance in the reading mode setting section 223 of the reading mode specifying section 22 and the color command setting section 21.
The character code string corresponding to the reading area name set in the color command learning section 212 indicates the synthesized speech outputted via the speech synthesis section 5 and the speech output section 6. First, to request voice guidance to know what kind of reading mode is possible, input the voice "1 area" from the voice input section 221, and when the voice recognition section 222 recognizes it as "Ready to go" (step STI, 5T2), the character code string corresponding to the reading area name that has been set in the color command learning section 212 of the color command setting section 21, such as "Bibliographical matters", 'Summary', the title of Chapter 1, Section 2, etc. ..."Full text" etc. is output.
次に読書モードを指定するために「書誌的事項、と発声
し、音声認識部222で1しよしてきしこう」と認識さ
れると、前述の読書モード学習処理で登録されているパ
ラメータを用いて音声認識部222で音声認識され、間
違いなく認識きれると(ステップ5T4)、認識結果の
文字コード列に対応する音声をエコーバック(ステップ
5T11)した後、指定した読書モードであるか否かの
確認のために合成音声が出力される(ステップ12)。Next, in order to specify the reading mode, when the voice recognition unit 222 recognizes ``Bibliographical matters'' and ``Let's go ahead'', the parameters registered in the reading mode learning process described above are used. The voice is recognized by the voice recognition unit 222, and if it is recognized correctly (step 5T4), the voice corresponding to the character code string of the recognition result is echoed back (step 5T11), and then a check is made as to whether or not it is in the specified reading mode. Synthesized speech is output for confirmation (step 12).
指定された読書モードが所望のモードである場合には「
OK、の音声を入力すると(ステップ5T13)、音声
認識結果に基づいて読書領域名称を抽出モード切換部2
3に出力する(ステップ5T14)。他の読書モードを
指定する場合にも前述と同様な手順で行なう。If the specified reading mode is the desired mode, "
When the voice “OK” is input (step 5T13), the mode switching unit 2 extracts the reading area name based on the voice recognition result.
3 (step 5T14). When specifying other reading modes, follow the same procedure as described above.
前記ステップ13において、所望の読書モードでない場
合は10K」以外の音声を入力すると「もう−度」とい
う合成音声が出力され(ステップ5T8)、次の音声コ
マンドの入力を待つ(ステップ5T13,8.1)。In step 13, if a voice other than ``10K'' is input if the reading mode is not the desired one, a synthesized voice saying ``More'' is output (step 5T8), and the process waits for the next voice command to be input (steps 5T13, 8. 1).
読書モードの指定が終了すると、読書を開始するために
1開始」と発声し、音声認識部222で1かいし」と認
識されると(ステップSTI。When the designation of the reading mode is completed, ``1 start'' is uttered to start reading, and when the voice recognition unit 222 recognizes ``1'' (step STI).
2.3)、読書動作を開始するために画像入力部1がセ
ットされた文書、書籍の走査を開始する。2.3) In order to start the reading operation, the image input unit 1 starts scanning the set document or book.
抽出モード切換部23は読書モード指定部22の読書モ
ード設定部223から出力されている読書領域名称に基
づいてカラーコマンド設定部21のカラーコマンド学習
部212に登録されている読書情報の読書領域名称を選
択し、当該読書領域名称番号以外の色特性をマスク(0
に)したカラーマツプを画像抽出部24に出力する。な
お、前文読取モードの場合、カラーマツプは出力せず全
文読取信号を画像抽出部24に出力する。The extraction mode switching unit 23 selects the reading area name of the reading information registered in the color command learning unit 212 of the color command setting unit 21 based on the reading area name output from the reading mode setting unit 223 of the reading mode specifying unit 22. , and mask the color characteristics other than the reading area name number (0
The resulting color map is output to the image extraction section 24. In the case of the preamble reading mode, a full text reading signal is output to the image extracting section 24 without outputting a color map.
次に、画像抽出部24のマーク識別処理、マーク領域内
の画像抽出処理について説明する。Next, the mark identification process and the image extraction process in the mark area by the image extraction section 24 will be described.
マーク識別処理は次のようにして行なう。画像入力部1
が第5図に示すような文書を主走査を水平方向(XX)
、副走査を垂直方向(YY)として走査して出力される
RGBの多値レベル信号を前述の式(1)、(2)を用
いて色度座標x、yに変換し、該x、yを整数化したX
%、y%を用いて抽出モード切換部23から出力されて
でくるカラーマツプを参照し、該カラへマツプの出力値
がrO」以外の場合、走査点が所望の読書領域を示すマ
ーク上にあると識別する。Mark identification processing is performed as follows. Image input section 1
scans a document like the one shown in Figure 5 in the horizontal direction (XX)
, the RGB multi-level signal outputted by scanning with the sub-scanning in the vertical direction (YY) is converted into chromaticity coordinates x, y using the above equations (1) and (2), and the x, y X, which is an integer
%, y% to refer to the color map output from the extraction mode switching section 23, and if the output value of the color map is other than "rO", the scanning point is on the mark indicating the desired reading area. identify.
マーク領域内の画像抽出処理について第7図を用いて説
明する。The image extraction process within the mark area will be explained using FIG. 7.
第7図において、副走査YYI行ではマーク画像は検出
されないが、YYZ行では色特性(読取領域名称番号)
jのマーク画像が前述のマーク識別処理によって識別さ
れ、更にYYS行では前記マーク画像に連結し、且つ挾
まれた画像(x x s t〜X X = −b )が
存在するため、当該挾まれた画像を読書対象の画像とし
て画像抽出部24から出力する。In Fig. 7, no mark image is detected in the sub-scanning YYI row, but the color characteristics (reading area name number) are detected in the YYZ row.
The mark image of j is identified by the mark identification process described above, and furthermore, in the YYS row, there is an image (x x s t ~ X The extracted image is outputted from the image extraction unit 24 as an image to be read.
前述の手順で順次マーク画像に挾まれた画像を読書対象
の画像として出力していくが、副走査YYn行に達する
と連結するマーク画像が存在しなくなるので、読書対象
の画像の抽出は終了する。In the above-mentioned procedure, the images sandwiched between the mark images are sequentially output as images to be read, but when the sub-scanning line YYn is reached, there are no more connected mark images, so the extraction of images to be read ends. .
なお、マーク画像が連結しているか否かの判定は公知の
技術例えば特公昭60−55868号公報に詳述しであ
るように、下式(3)を用いて行なう。It should be noted that the determination as to whether or not the mark images are connected is made using the following equation (3) using a known technique, for example, as detailed in Japanese Patent Publication No. 60-55868.
XX<+−r)kb≦XX+*−
X X (H−’1 ) kv≦XX1k−(3)但し
、XX<l−5)*b + XX<+−r〕*wは各々
走査行YYi行の前走査行YYi−1の色特性「0.か
ら色特性rj」への変化点及び色特性「j」から色特性
「0」への変化点ペアのX座標を示す、xxlkmIX
Xlkkは各々走査行Yi(7)色特性「0」から前記
色特性と同一の色特性1j」への変化点及び前記色特性
rj」から色特性rOJへの変化点ペアのX座標を示す
。但し、抽出モード切換部23に全文読取モードが設定
されている場合、前述の処理は行なわず、画像入力部1
から出力され全画像を読書対象の画像として画像抽出部
24から出力する。XX<+-r) kb≦XX+*- X xxlkmIX, which indicates the X coordinate of a pair of points of change from color characteristic "0. to color characteristic rj" and change points from color characteristic "j" to color characteristic "0" in the previous scanning line YYi-1;
Xlkk indicates the X coordinate of a pair of points of change from the color characteristic "0" of the scanning line Yi(7) to the same color characteristic "1j" and from the color characteristic "rj" to the color characteristic rOJ, respectively. However, if the full text reading mode is set in the extraction mode switching unit 23, the above process is not performed and the image input unit 1
All the images outputted from the image extraction section 24 are outputted as images to be read.
なお、読書対象の文書、書籍が無彩色の場合には、画像
抽出部24に無彩色画像抽出手段を設けておき、画像抽
出部24から出力される画像は無彩色画像のみとしてお
くことにより、マーク画像による影響を除去することが
可能である。Note that when the document or book to be read is achromatic, the image extraction section 24 is provided with an achromatic image extraction means, and the images output from the image extraction section 24 are only achromatic images. It is possible to remove the influence of mark images.
前記無彩色画像抽出手段は、公知の方法、例えばR,G
、Bの信号が全て同一に近いレベルにあることによって
実現する。The achromatic image extraction means uses a known method, for example, R, G
, B are all at nearly the same level.
また、前記無彩色抽出手段には2値化手段を接続し、2
値化きれた画像を読書対象の画像とする。なお、2値化
のための閾値は公知の如何なる方法によってもよい。Further, a binarization means is connected to the achromatic color extraction means, and
The digitized image is used as the image to be read. Note that the threshold value for binarization may be determined using any known method.
読書対象に「書誌的事項、を指定し、第5図に示す文書
を入力した場合、マークmtに囲まれた画像が読書対象
の画像とし℃出力きれ、当該画像から文字の切出し、認
識を行なった後、音声合成されて朗読音声が出力される
。If you specify "Bibliographical matters" as the reading target and input the document shown in Figure 5, the image surrounded by the mark mt will be output as the reading target image, and characters will be cut out and recognized from the image. After that, the voice is synthesized and the reading voice is output.
読書対象に「要約、や1章1節の題名」が指定された場
合も同様な処理が行なわれる。Similar processing is performed when a "summary or title of chapter 1 section 1" is specified as the reading target.
なお、カラーコマンド学習用紙及び文書、書籍上に追記
又は印刷するマークの位置、大きさ、色数及び形状は前
述のものに限定されるものではない。Note that the position, size, number of colors, and shape of the mark to be added or printed on the color command learning paper, document, or book are not limited to those described above.
また、文書、書籍かカラーコマンド学習用紙かを判別で
きるマークを用紙上に付加し、前記マークの有無を識別
できる機能をカラーコマンド設定部21の学習/読書モ
ード切換部211に付加しておくことによって、モード
切換の動作を自動化することができる。Further, a mark is added to the paper to determine whether it is a document, a book, or a color command learning paper, and a function for identifying the presence or absence of the mark is added to the learning/reading mode switching unit 211 of the color command setting unit 21. This allows the mode switching operation to be automated.
なお、前記マークは例えば用紙の上辺に近い領域に所定
長、所定巾、所定数の黒の線分を記載したものである。Note that the mark is, for example, a black line segment of a predetermined length, a predetermined width, and a predetermined number written in an area near the top side of the paper.
本実施例では、カラーコマンド設定部21の学習/読書
モード切換部211及びカラーコマンド学習部212を
用いて、カラーコマンド用紙上のマーク及び文字を学習
し、色特性及び読書領域を設定したが、前述の説明では
読書領域は文書、書籍に追記又は印刷きれたマークを用
いたが、文書、書籍の文字を色分けしたものを用いるよ
うにし、前述の画像抽出部24のマーク識別処理を色分
けした文字画像を抽出するように構成し℃もよい。In this embodiment, the learning/reading mode switching section 211 and the color command learning section 212 of the color command setting section 21 are used to learn the marks and characters on the color command paper and set the color characteristics and reading area. In the above explanation, marks added or printed on documents and books were used for the reading area, but color-coded characters of documents and books were used, and the mark identification process of the image extraction unit 24 described above was performed using color-coded characters. Configured to extract images is also good.
なお、読書領域の画像抽出以降の処理については、前述
及び第1図に限定されるものではない。Note that the processing after image extraction of the reading area is not limited to that described above and to that shown in FIG.
以上詳細に説明したように本発明によれば、文書、書籍
の所望の部分にカラーマークを追記又は印刷しておき、
該文書、書籍のカラーマーク部分の文章或いは全文を音
声ガイダンス及び音声認識で自由に選択できる読書処理
装置としたので、本読書処理装置を使用すると、製本業
者が印刷を工夫するか、健常者が一度追記マークを施し
た文書、書籍であれば、健常者が日常的に行なっている
文書、書籍の拾い読みに近い形の読み方ができるので、
盲人や肢体障害者等の障害者の読書環境の向上に役立つ
。また健常者において聴覚による読書環境と同じ環境の
読書処理装置の提供が可能になるという優れた効果が得
られる。As explained in detail above, according to the present invention, a color mark is added or printed on a desired part of a document or book,
This reading processing device allows you to freely select the text or the full text of the color marked part of the document or book using voice guidance and voice recognition. Once a document or book has been marked with addendum marks, it can be read in a way similar to the way healthy people read documents and books on a daily basis.
Helps improve the reading environment for people with disabilities such as blind people and physically disabled people. Further, an excellent effect can be obtained in that it is possible to provide a reading processing device with the same environment as an auditory reading environment for a healthy person.
理装置の構成を示すブロック図、第3図は色特性の説明
図で、同図(a)はマークの色度座標の範囲の説明図、
同図(b)はマークの濃度範囲の説明図、第4図(a)
、cb)、(c)はそれぞれカラーコマンド学習用紙の
例を示す図、第5図は入力文書の例を示す図、第6図は
音声コマンド。
音声ガイダンスによる処理例を示す図、第7図は画像抽
出の説明図である。
図中、1・・・・画像入力部、2・・・・画像処理部、
3・・・・文字認識処理部、4・・・・言語処理部、5
・・・・音声合成部1.6・・・・音声出力部、21・
・・・カラーコマンド設定部、211・・・・学習/読
書モード切換部、212・・・・カラーコマンド学習部
、22・・・・読書モード指定部、221・・・・音声
入力部、222・・・・音声認識部、223・・・・読
書モード設定部、23・・・・抽出モード切換部、24
・・・・画像抽出部。
1LLs神事C理裏夏のλ鵬θく≦ルずフ゛O−Iフ目
第2図
色お眉−故萌目
第3図
−一一→主11″(XX)
入−jJj、ルイ枦Jf、iミ♂[〔ン]第5図
fpコづ ント′7 g戸1τイy″/ズ(・J32−
理f′j行、引コ第6図FIG. 3 is a block diagram showing the configuration of the optical system; FIG. 3 is an explanatory diagram of color characteristics; FIG.
Figure 4(b) is an explanatory diagram of the mark density range, and Figure 4(a)
, cb) and (c) are diagrams each showing an example of a color command learning sheet, FIG. 5 is a diagram showing an example of an input document, and FIG. 6 is a voice command. FIG. 7 is a diagram showing an example of processing using voice guidance, and is an explanatory diagram of image extraction. In the figure, 1... image input section, 2... image processing section,
3...Character recognition processing unit, 4...Language processing unit, 5
...Speech synthesis section 1.6...Speech output section, 21.
...Color command setting section, 211... Learning/reading mode switching section, 212... Color command learning section, 22... Reading mode specifying section, 221... Audio input section, 222 ...Speech recognition unit, 223...Reading mode setting unit, 23...Extraction mode switching unit, 24
...Image extraction section. 1LLs Shinto ritual C Riura summer λ Peng θ Ku ≦ Ruzufu O-I Fu eyes 2nd figure Color eyebrows - Late Moe eyes 3rd figure - 11 → Lord 11'' (XX) Enter-j Jj, Louis 枦 Jf , i min ♂ [[n] Fig. 5 fp comment '7 g door 1
ri f'j row, pull column Fig. 6
Claims (1)
声合成部等を具備し、文字認識結果に基づいて朗読音声
を出力する読書処理装置において、 前記画像入力部として、用紙上に記載されている文字、
図形、写真及び/又はマーク等の画像の濃淡及び色の一
方又は双方に応じた情報を出力する画像入力部を用い、 前記画像処理部として、所定の色特性と読書領域名称と
を一対の読書情報として設定するカラーコマンド設定部
と、音声対話処理によって読書領域名称を指定する読書
モード指定部と、該読書モード指定部で指定された読書
領域名称に基づいて読書情報の色特性を読取色特性とし
て出力する抽出モード切換部と、該読取色特性に基づい
て文書、書籍上に記載されているマーク領域内の文書画
像を抽出する画像抽出部とを具備する画像処理部を用い
たことを特徴とする読書処理装置。[Scope of Claims] A reading processing device that includes an image input section, an image processing section, a character recognition section, a language processing section, a speech synthesis section, etc., and outputs a reading voice based on a result of character recognition, comprising: the image input section. as written on the paper,
Using an image input unit that outputs information corresponding to one or both of the shading and color of images such as figures, photographs, and/or marks, the image processing unit inputs predetermined color characteristics and reading area names into a pair of reading images. A color command setting section for setting information, a reading mode specifying section for specifying a reading area name through voice interaction processing, and a color characteristic for reading color characteristics of reading information based on the reading area name specified in the reading mode specifying section. and an image extraction section that extracts a document image within a mark area written on a document or book based on the read color characteristics. Reading processing device.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2132791A JPH0432960A (en) | 1990-05-23 | 1990-05-23 | Reading processor |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2132791A JPH0432960A (en) | 1990-05-23 | 1990-05-23 | Reading processor |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH0432960A true JPH0432960A (en) | 1992-02-04 |
Family
ID=15089637
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP2132791A Pending JPH0432960A (en) | 1990-05-23 | 1990-05-23 | Reading processor |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH0432960A (en) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0973461A (en) * | 1995-09-06 | 1997-03-18 | Shinano Kenshi Co Ltd | Sentence information reproducing device using voice |
-
1990
- 1990-05-23 JP JP2132791A patent/JPH0432960A/en active Pending
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0973461A (en) * | 1995-09-06 | 1997-03-18 | Shinano Kenshi Co Ltd | Sentence information reproducing device using voice |
| US5822284A (en) * | 1995-09-06 | 1998-10-13 | Shinano Kenshi Kabushiki Kaisha | Audio player which allows written data to be easily searched and accessed |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US8458583B2 (en) | Document image generating apparatus, document image generating method and computer program | |
| EP1591954A1 (en) | Image processing device | |
| JP4337251B2 (en) | Image processing apparatus, image processing method, and computer-readable recording medium storing image processing program | |
| US9277094B2 (en) | Image processing apparatus and recording medium | |
| JP3211488B2 (en) | Document processing device | |
| KR102336051B1 (en) | Color-Tactile_Pattern Transformation System for Color Recognition and Method Using the Same | |
| JPH0432960A (en) | Reading processor | |
| JP2009038737A (en) | Image processing device | |
| EP3489859B1 (en) | Image processing apparatus | |
| KR100560314B1 (en) | the method of marking the elements of a foreign sentence by Using the Color chart of the Digital Color system | |
| JPH0424885A (en) | Reading processor | |
| JPH06208357A (en) | Document processor | |
| JP2002109542A (en) | Image processing system and data processing apparatus and method | |
| JPH0466998A (en) | Information processor | |
| JP2896919B2 (en) | Image processing device | |
| JP3030126B2 (en) | Image processing method | |
| CN111612007A (en) | An English secondary braille conversion system based on image acquisition and correction | |
| JPH07203230A (en) | Color image forming device | |
| Ina | Presentation of images for the blind | |
| JP3162575B2 (en) | Character recognition device | |
| JPH08202824A (en) | Document image recognition device | |
| JPH03225477A (en) | Image processor | |
| WO2024018553A1 (en) | Translated data creation device, translated data creation method, and translated data creation program | |
| Suzuki et al. | Comfortable color conversion using image segmentation | |
| JPH03248279A (en) | Picture processor |