JPH05224687A - Japanese sentence reading word conversion edit processing method - Google Patents

Japanese sentence reading word conversion edit processing method

Info

Publication number
JPH05224687A
JPH05224687A JP4030232A JP3023292A JPH05224687A JP H05224687 A JPH05224687 A JP H05224687A JP 4030232 A JP4030232 A JP 4030232A JP 3023292 A JP3023292 A JP 3023292A JP H05224687 A JPH05224687 A JP H05224687A
Authority
JP
Japan
Prior art keywords
word
conversion
polysemous
dictionary
japanese
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP4030232A
Other languages
Japanese (ja)
Inventor
Shinichiro Takagi
伸一郎 高木
Hisashi Nakada
寿 中田
Masashi Katsumata
雅司 勝俣
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Inc
Original Assignee
Nippon Telegraph and Telephone Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Telegraph and Telephone Corp filed Critical Nippon Telegraph and Telephone Corp
Priority to JP4030232A priority Critical patent/JPH05224687A/en
Publication of JPH05224687A publication Critical patent/JPH05224687A/en
Pending legal-status Critical Current

Links

Landscapes

  • Document Processing Apparatus (AREA)

Abstract

PURPOSE:To reduce the throughput of document editing by an editor by previously setting information flags which indicate homonym and polysemy in word information in a Japanese word dictionary and providing a conversion word dictionary which contains substitute conversion words. CONSTITUTION:The system is equipped with the conversion word dictionary 90 which contains the conversion word for homonym substitution for each combination of a polysemy index and a conversion word for polysemy substitution, and a homonym index and the meaning attribute of a case element word. A morphemic analysis process is performed for a Japanese document to be pronounced and a polysemy extracting process part 50 extracts polysemy from a certified word string. Further, a polysemy deciding process part 60 narrows down the polysemy according to the co-occurrence relation between a document decomposition attribute attached to a document file and the meaning attribute of the polysemy when there are plural word candidates meeting grammatical connection requirements to decide the word candidate. Further, a word converting and editing process part 100 retrieves the conversion word dictionary 90 to extract a conversion word and substitute the character string in a source document with it, and writes the character string in a document file 110.

Description

【発明の詳細な説明】Detailed Description of the Invention

【0001】[0001]

【産業上の利用分野】本発明は、日本文のニュース速報
や各種情報案内などを読み上げて通信手段で配信する際
に、日本文文章の単語を聞き取りやすく変換編集する日
本文読み上げ単語変換編集処理方式に関するものであ
る。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention is a Japanese sentence reading word conversion / editing process for converting and editing a Japanese sentence sentence word in an easy-to-understand manner when reading a Japanese sentence breaking news or various information guides and delivering them by communication means. It is about the method.

【0002】[0002]

【従来の技術】日本文のニュース速報や各種情報案内な
どの即時性の必要な情報データが通信手段で伝送され電
話やファクシミリなどのメディアで利用者に配信するサ
ービスが急増している。特に日本文文章を読み上げて電
話などの音声データで配信する場合には読み上げ処理が
一過性であるため、読み上げ内容の聞き取りやすさが重
要である。しかし、日本語には、同音語や同形で多義を
有する単語があり、これが混在することによっては聞き
誤りや聞きにくさなどの了解度の低下が発生する。
2. Description of the Related Art There is a rapid increase in services in which information data that requires immediacy such as breaking news in Japanese and various information guides are transmitted by communication means and distributed to users through media such as telephones and facsimiles. In particular, when a Japanese sentence is read aloud and delivered by voice data such as a telephone, the reading process is transient, so that it is important to make the read contents easy to hear. However, in Japanese, there are homophones and isomorphic words that have polysemy, and if they are mixed, the intelligibility such as erroneous listening and difficulty in hearing occurs.

【0003】これに対して、情報の了解度を上げるため
に、 利用者は繰り返し読み上げなどの操作を行う、 情報配信元で単語読み分け用の特殊読み付与や単語置
換を行う、 情報配信元で単語表現の変換の編集作業を行う、 などの対応が可能である。
On the other hand, in order to increase the intelligibility of information, the user repeatedly performs operations such as reading aloud, adding special reading for word reading distinction and word replacement at the information distribution source, and the word at the information distribution source. It is possible to take actions such as editing the expression conversion.

【0004】[0004]

【発明が解決しようとする課題】しかし、では、利用
者の操作が増えたりスピーカなどで複数利用者に音声デ
ータを流す場合は本質的に繰り返し読み上げなどは不可
能であるなどの問題点がある。
However, there is a problem that it is essentially impossible to repeat reading when the number of operations by the user is increased or when voice data is sent to a plurality of users through a speaker or the like. ..

【0005】また、では、文章校正処理の例で行われ
ているように、原文文字をイメージしやすいような特殊
読みを行う。例えば、「追及」「追究」「追求」は、
『ツイキュウオヨビ』『ツイキュウキワメ』『ツイキュ
ウモトメ』となる。しかし、一般の情報利用者には特殊
読みは聞きにくく自然性を損なってわかりにくくなる。
また出現する単語をわかりやすい単語に文字的に置換す
る手段では、「日中間で経済問題を日中協議する」の場
合で、『日中』→『日本と中国』とすると、「日本と中
国間で経済問題を日本と中国協議する」となる。さら
に、「汚染土壌を採る法的手段を採る」の場合で、『採
る』→『採取する』とすると、「汚染土壌を採取する法
的手段を採取する」となり、いずれも不適当な表現に置
き換えられてしまう問題がある。
In addition, as in the example of the grammar proofreading process, special reading is performed so that the original characters can be easily imaged. For example, "pursuit""pursuit""pursuit"
It will be "Tsuikyu Oyobi", "Tsuikyu Kiwame", and "Tsuikyu Motome". However, it is difficult for general information users to hear the special reading, and it becomes difficult to understand because the naturalness is impaired.
In addition, as a means to replace the appearing words with words that are easy to understand, in the case of "Japan-China talks about economic problems between Japan and China", if "Japan-China" → "Japan and China" Will discuss economic issues with Japan and China. " Furthermore, in the case of “taking legal means to take contaminated soil”, “take” → “collect” means “to take legal means to take contaminated soil”, and both are inappropriate expressions. There is a problem that they will be replaced.

【0006】また、では、日本文文章の編集作業段階
で多数の同音語や多義語に対する変換や編集を必要な範
囲で行わねばならないので熟練を要する作業であった。
このように、日本文文章の読み上げ処理において聞き取
りにくい同音語や多義語を抽出して聞き誤りが少なくわ
かりやすい単語表現に変換するなどの編集を必要な範囲
で行うことは、編集者の能力を有するので短時間で処理
することは困難であるという問題点があった。
[0006] In addition, since it is necessary to convert and edit a large number of homophones and polysemous words within a necessary range at the stage of editing a Japanese sentence sentence, it is a work requiring skill.
In this way, it is possible for the editor to perform editing within a necessary range, such as extracting homophones and polysemous words that are difficult to hear in the reading process of Japanese sentence sentences and converting them into easy-to-understand word expressions with few listening errors. Therefore, it is difficult to process in a short time.

【0007】本発明は、上記の従来手段における問題点
を解決するために、日本文文章を形態素解析処理の結果
から変換すべき条件に合う同音語や多義語について単語
変換辞書を検索して該当する変換単語に変換し編集する
ようにすることを目的としている。
In order to solve the above-mentioned problems in the conventional means, the present invention searches a word conversion dictionary for a homophone or a polysemous word that matches a condition for converting a Japanese sentence sentence from the result of morphological analysis processing. It is intended to be converted into a converted word and edited.

【0008】[0008]

【課題を解決するための手段】上記の目的を実現するた
めに、本発明では、読み上げ用の日本文文章ファイルに
含まれる同音語や多義語を聞き取りやすい単語に変換し
編集する処理において、読み上げ用の日本文原文章を日
本語単語辞書と文法辞書とを用いて形態素解析処理して
単語の認定と単語の言語情報とを取得する手段と、予め
日本語単語辞書の単語情報に多義語を示す情報フラグを
設定する手段と、該情報フラグによって多義語を抽出す
る手段と、多義語で文法的な接続条件を満たす単語候補
(意味的多義)が複数存在する場合に、文章ファイルに
付随する文章分野属性と多義語の意味属性(単語の意味
を一般的な用語で示したもの)との共起性(分野依存
度)で多義を絞り込む手段と、予め多義語の各多義毎に
多義語の単語見出しと品詞に対して置換する変換単語と
を収録した変換単語辞書と、該変換単語辞書を検索して
多義語を変換編集する手段と、をそなえることを特徴と
し、また予め日本語単語辞書の単語情報に動詞の同音語
を示す情報フラグを設定する手段と、該情報フラグによ
って同音語を抽出する手段と、同音語の前方に存在する
格構造関係の名詞(以下、格要素単語という)の意味属
性を抽出する手段と、予め同音語見出しと格要素単語の
意味属性との組み合わせ毎に置換する変換単語を収録し
た変換単語辞書と、該変換単語辞書を検索して同音語を
変換編集する手段とを備えることを特徴とする。
In order to achieve the above object, according to the present invention, in a process of converting a homonym or a polysemous word included in a Japanese sentence sentence file for reading into a word that is easy to hear and editing, A method for morphological analysis of original Japanese sentences using a Japanese word dictionary and a grammar dictionary to obtain word recognition and word linguistic information, and to add polysemous words to the word information of the Japanese word dictionary in advance. When a plurality of word candidates (semantic polysemy) satisfying a grammatical connection condition with a polysemous word are present, a means for setting an information flag shown, a means for extracting a polysemous word by the information flag, and the word file are attached to the text file. A means for narrowing down polysemy by co-occurrence (field dependency) of the text field attribute and the meaning attribute of polysemy (meaning of a word is shown in general terms), and polysemous word for each polysemous word in advance Word headlines The present invention is characterized by having a conversion word dictionary containing a conversion word to be replaced for a part of speech and a means for searching the conversion word dictionary to convert and edit a polysemous word, and also to previously prepare word information of the Japanese word dictionary. A means for setting an information flag indicating a homonym of a verb, a means for extracting the homonym with the information flag, and a semantic attribute of a noun (hereinafter referred to as a case element word) having a case structure relationship existing in front of the homonym. And a conversion word dictionary containing conversion words to be replaced in advance for each combination of the homophone heading and the meaning attribute of the case element word, and a means for searching the conversion word dictionary and converting and editing the homophone. It is characterized by including.

【0009】[0009]

【作用】本発明においては、読み上げ用の日本文文章フ
ァイルに含まれる同音語や多義語を聞き取りやすい単語
に変換し編集する処理において、日本語単語辞書の単語
情報に同音語や多義語を示す情報フラグを予め設定し、
置換する変換単語を収録した変換単語辞書を備えてあ
り、これにより、留意すべき同音語や多義語を抽出する
処理や前方の格要素単語の意味属性による聞き取りやす
い単語の置換処理を行うことが可能となるので、編集者
の文章編集の処理量を軽減できる。
In the present invention, in a process of converting a homonym or a polysemous word included in a Japanese sentence sentence file for reading into a word that is easy to hear and editing it, the homonym or polysemous word is shown in the word information of the Japanese word dictionary. Preset information flag,
It is equipped with a conversion word dictionary that contains the conversion words to be replaced, which enables the processing of extracting homonyms and polysemous words that should be noted and the processing of replacing easily recognizable words by the semantic attributes of case element words in the front. This enables the editor to reduce the amount of text editing processing.

【0010】[0010]

【実施例】以下、本発明の実施例を図面により詳細に説
明する。図1から図8は、本発明の一実施例を示す図で
ある。そして、図4から図6は図1のブロック構成にお
ける本発明の請求項1の実施例を示す図であり、図7か
ら図8は図1のブロック構成における本発明の請求項2
の実施例を示す図である。
Embodiments of the present invention will now be described in detail with reference to the drawings. 1 to 8 are diagrams showing an embodiment of the present invention. 4 to 6 are views showing an embodiment of claim 1 of the present invention in the block configuration of FIG. 1, and FIGS. 7 to 8 are claims 2 of the present invention in the block configuration of FIG.
It is a figure which shows the Example of.

【0011】図1は処理ブロック構成例、図2は日本語
単語辞書の構成例、図3は変換単語辞書の構成例、図4
と図7は処理概略フロー、図5、図6、図8は単語変換
編集処理の実施例を説明する図である。
FIG. 1 is a processing block configuration example, FIG. 2 is a Japanese word dictionary configuration example, FIG. 3 is a conversion word dictionary configuration example, and FIG.
7 and FIG. 7, and FIG. 5, FIG. 6, and FIG. 8 are diagrams for explaining an example of the word conversion editing process.

【0012】図1において、処理装置120はCPUお
よびメモリからなる処理装置で以下の機能部を有する。
すなわち、読み上げ用の日本文原文章ファイル10の日
本文文章を日本語単語辞書20、文法辞書30を用いて
単語の認定や単語の言語情報の取得を行う形態素解析処
理部40;単語情報から抽出した情報フラグのうち『同
形多義フラグ』がオンの単語を抽出する多義語抽出処理
部50;文法的な接続条件を満たす単語候補(意味的多
義)が存在する場合に多義を絞り込む多義語判定処理部
60;単語情報から抽出した情報フラグのうち『同音フ
ラグ』がオンの単語を抽出する同音語抽出処理部70;
同音語の前方に存在する格要素単語の意味属性を抽出す
る格要素意味属性抽出処理部80;多義語単語見出しと
多義語置換用の変換単語ならびに同音語単語見出しと格
要素単語の意味属性との組み合わせ毎に同音語置換用の
変換単語を予め収録した変換単語辞書90;該変換単語
辞書を検索して多義語や同音語を変換編集する単語変換
編集処理部100;編集済みの文章ファイル110から
なる。
In FIG. 1, a processing device 120 is a processing device including a CPU and a memory and has the following functional units.
That is, the Japanese sentence sentence of the Japanese sentence original sentence file 10 for reading out is extracted from the morphological analysis processing unit 40 that authenticates the word and acquires the language information of the word using the Japanese word dictionary 20 and the grammar dictionary 30; Of a plurality of information flags, a polysemous word extraction processing unit 50 for extracting a word for which the “isomorphic polysemous flag” is turned on; a polysemous word determination process for narrowing down polysemous words when there is a word candidate (semantic polysemous) that satisfies a grammatical connection condition Part 60; Homophone extraction processing unit 70 that extracts words for which the “homophone flag” is ON from the information flags extracted from the word information;
Case element semantic attribute extraction processing unit 80 for extracting the semantic attribute of the case element word existing in front of the homophone; the polysemous word headline and the conversion word for polysemous word replacement, and the homonym word headline and the semantic attribute of the case element word. A conversion word dictionary 90 in which conversion words for homophone substitution are previously recorded for each combination of; a word conversion edit processing unit 100 for searching the conversion word dictionary to convert and edit a polysemous word or a homophone; edited text file 110 Consists of.

【0013】この処理装置120では、読み上げ用の日
本文原文章の任意の文章について先頭から形態素解析処
理を行い、単語の認定と単語の言語情報として品詞、読
み、同音フラグ、同形多義フラグ、意味属性などの認定
とを行う。次に、認定された単語列の中から多義語
(『同形多義フラグ』がオンの単語)を抽出する(多義
語抽出処理部50)。文法的な接続条件を満たす単語候
補(意味的多義)が複数存在する場合に文章ファイルに
付随する文章分野属性と多義語の意味属性との共起性
(分野依存度)で多義を絞り込み単語候補を決定する
(多義語判定処理部60)。さらに、多義語の単語見出
しと認定品詞で変換単語辞書90を検索して変換単語を
抽出し原文章ファイルの文字列に置換して編集済みの文
章ファイル110に書き込む(単語変換編集処理部10
0)。
In this processing device 120, morphological analysis processing is performed from the beginning for an arbitrary sentence of a Japanese original sentence for reading, and the part of speech, reading, homophone flag, isomorphic flag, meaning as word recognition and word language information. Authenticate with attributes. Next, a polysemous word (a word for which the “isomorphic polysemous flag” is ON) is extracted from the recognized word string (polysemous word extraction processing unit 50). When there are multiple word candidates (semantic ambiguity) that satisfy the grammatical connection condition, the word sense is narrowed down by the co-occurrence (field dependency) between the sentence field attribute attached to the text file and the meaning attribute of the polysemous word. Is determined (the polysemous word determination processing unit 60). Furthermore, the converted word dictionary 90 is searched with the word headwords of the polysemous words and the certified part of speech, the converted words are extracted, replaced with the character strings of the original sentence file, and written in the edited sentence file 110 (the word conversion edit processing unit 10).
0).

【0014】また、認定された単語列の中から品詞が動
詞の同音語(『同音フラグ』がオンの単語)を抽出し
(同音語抽出処理部70)、同音語の前方に存在する格
要素単語の意味属性を抽出する(格要素意味属性抽出処
理部80)。
Further, a homonym of which the part of speech is a verb (a word whose "homophone flag" is ON) is extracted from the recognized word string (homophone extraction processing section 70), and the case element existing in front of the homophone is extracted. The semantic attribute of the word is extracted (case element semantic attribute extraction processing unit 80).

【0015】予め単語見出しと格要素単語の意味属性と
の組み合わせ毎に置換する変換単語を収録した変換単語
辞書90を、同音語の単語見出しと抽出した格要素の意
味属性で検索してマッチするパターンの変換単語を抽出
し、原文章ファイルの文字列に置換して編集済みの文章
ファイル110に書き込む(単語変換編集処理部10
0)。
The converted word dictionary 90 containing the converted words to be replaced in advance for each combination of the word heading and the meaning attribute of the case element word is searched and matched with the homologous word heading and the extracted meaning attribute of the case element. The conversion word of the pattern is extracted, replaced with the character string of the original text file, and written in the edited text file 110 (word conversion editing processing unit 10
0).

【0016】図2は、図1のブロック図の構成要素であ
る日本語単語辞書20の構成例を示す図である。「日
中」では、時詞と固有名詞との2つの多義を有する単語
候補があるので、同形多義フラグを付与されている。さ
らに、単語候補の多義を判定するために意味属性(単語
の意味を一般的な用語で示したもの)として、『時間』
と『国、地域』がある。また、動詞の同音語の「取る」
「採る」「捕る」「撮る」にはいずれも同音フラグが付
与されている。
FIG. 2 is a diagram showing a configuration example of the Japanese word dictionary 20 which is a component of the block diagram of FIG. In “Chinese”, there are word candidates having two ambiguous meanings such as a toki and a proper noun, so the isomorphic ambiguous flag is given. In addition, as a semantic attribute (indicating the meaning of a word in general terms) to determine the meaning of a word candidate, "time"
There is a "country, region". Also, the verb homophone “Take”
The same-sound flag is attached to each of “take”, “capture”, and “shoot”.

【0017】図3は、図1のブロック図の構成要素であ
る変換単語辞書90の構成例を示す図である。ここで、
130は同音語の前方の格要素単語の意味属性、140
は変換単語である。「日中」では、時詞と固有名詞との
2つの多義を有する単語候補があるので、それぞれの変
換単語は、時詞の場合に『昼間の間』、固有名詞の場合
に『日本と中国』となり、格納されている。また、動詞
の同音語の場合には格要素単語の意味属性で変換単語が
異なる。例えば、「採る」では、格要素単語の意味属性
が「制度」の時には『採用する』、「生物」の時には
『採取する』となる。
FIG. 3 is a diagram showing an example of the structure of the converted word dictionary 90 which is a component of the block diagram of FIG. here,
130 is the semantic attribute of the case element word preceding the homophone, 140
Is a conversion word. Since there are two ambiguous word candidates for "Japanese and Chinese", that is, a tongue and a proper noun, each converted word is "daytime" when it is a tongue, and "Japan and China" when it is a proper noun. ] And is stored. Further, in the case of a homonym of a verb, the converted word differs depending on the semantic attribute of the case element word. For example, in the case of "collect", when the semantic attribute of the case element word is "system", it is "adopted", and when it is "biological", it is "collected".

【0018】図4は図1に示した処理ブロック構成例に
おいて、日本文原文章ファイル10に含まれる多義語を
変換する処理の概略フローを示す図であり、概略フロー
に従って、動作の説明を行う。
FIG. 4 is a diagram showing a schematic flow of a process of converting a polysemous word included in the Japanese sentence original sentence file 10 in the processing block configuration example shown in FIG. 1. The operation will be described according to the schematic flow. ..

【0019】 日本文原文章ファイル10より単語変換編集処理を施す処理対象文章を読み込 む (ステップ100) 読み込んだ全文章について形態素解析処理部40において、日本語単語辞書2 0、文法辞書30を用いて、形態素解析処理を行い、単語の認定と、単語の言語 情報として品詞、読み、同音語フラグ、同形多義フラグ、意味属性などの認定と を行う (ステップ110) 多義語抽出処理部50において、認定された単語列の中から多義語を示す情報 フラグによって多義語を抽出し、多義語でない場合にはステップ180に分岐す る (ステップ120) 抽出した多義語において、文法的な接続条件を満たす単語候補(意味的多義) が複数存在するかを判定して、意味的多義がなければステップ160に分岐する (ステップ130) 多義語判定処理部60において、意味的多義を有する多義語の場合(ステップ 130の判定がYESの場合)には、文章分野属性と多義語の意味属性との共起 性(分野依存度)を調べる (ステップ140) 共起性(分野依存度)の高い方の単語候補を認定する (ステップ150) 認定した多義語の単語見出しと認定品詞で、変換単語辞書90を検索して変換 単語を抽出し原文章ファイルの文字列を置換する (ステップ160) 編集済みの文章ファイル110に書き込む (ステップ170) 全単語について処理を終了したかを判定して判定がYESの場合は処理を終え る。また、判定がNOの場合、次単語を読み込み(ステップ190)、ステップ 120へ移行し処理を継続する (ステップ180) 図5は図4の多義語を変換する処理の概略フローにおけ
る多義判定処理の実施例である。
The processing target sentence to which the word conversion editing process is applied is read from the Japanese sentence original sentence file 10 (step 100). In the morphological analysis processing unit 40 for all the read sentences, the Japanese word dictionary 20 and the grammar dictionary 30 are read. Using the morphological analysis process, the word is identified and the part of speech, reading, homophone flag, homomorphic flag, and semantic attribute are identified as the language information of the word (step 110). , A polysemous word is extracted from the recognized word string by an information flag indicating a polysemous word, and if it is not a polysemous word, the process branches to step 180 (step 120). It is determined whether there are a plurality of word candidates (semantic ambiguity) to be satisfied, and if there is no semantic ambiguity, the process branches to step 160 (step 1 0) In the polysemous word determination processing unit 60, in the case of a polysemous word having semantic polysemy (when the determination in step 130 is YES), the co-occurrence of the sentence field attribute and the meaning attribute of the polysemous word (field dependency degree) ) Is confirmed (step 140) The word candidate with the higher co-occurrence (field dependency) is recognized (step 150) The converted word dictionary 90 is searched with the recognized polysemous word heading and certified part of speech to convert the word Is extracted and the character string in the original sentence file is replaced (step 160) It is written in the edited sentence file 110 (step 170) It is determined whether or not the processing has been completed for all words, and if the determination is YES, the processing is terminated. .. If the determination is NO, the next word is read (step 190) and the process proceeds to step 120 to continue the process (step 180). FIG. 5 shows the polysemous determination process in the schematic flow of the process for converting polysemous words in FIG. This is an example.

【0020】ここで、150は文章分野属性、160は
原文章文字列、170は多義語、180は文法的単語接
続条件、190は分野依存度判定条件、195は多義判
定後の認定品詞である。
Here, 150 is a text field attribute, 160 is an original text character string, 170 is a polysemous word, 180 is a grammatical word connection condition, 190 is a field dependency determination condition, and 195 is a certified part of speech after polysemous judgment. ..

【0021】抽出した多義語「日中」において、文法的
な接続条件を満たす単語候補(意味的多義)が複数存在
するかを判定する。「日中開かれた」では、後続単語の
動詞「開かれ」との文法的接続条件で「日中(時詞)」
のみが認定されるため意味的多義はない。しかし、「日
中間の懸案」では、後続単語の接尾辞「間」との文法的
接続条件で「日中(時詞)」と「日中(固有名詞)」と
の複数の単語候補が認定されるため意味的多義が発生す
る。意味的多義が残留すると変換単語の検索処理で支障
があるため、文章分野属性150を用いて多義語の意味
属性との共起性(分野依存度)を調べ多義を絞る。ここ
では、文章分野属性が『国際政治』であるので、意味属
性が「国、地域」の「日中(固有名詞)」の方が認定さ
れる(認定後の品詞195)。
In the extracted polysemic word "Japanese-Chinese", it is determined whether or not there are a plurality of word candidates (semantic polysemous) that satisfy the grammatical connection condition. In "Japanese-Chinese opened", "Japanese (Chinese)" is defined as a grammatical connection condition with the verb "Opened" of the following word.
There is no semantic ambiguity because only is certified. However, in "Japan-China Suspicion", multiple word candidates "Chinese (Hijiki)" and "Chinese (Proper noun)" are identified based on the grammatical connection conditions with the suffix "Ma" of the subsequent words. As a result, semantic ambiguity occurs. If the semantic polysemy remains, it may hinder the conversion word search process. Therefore, the coexistence with the semantic attribute of the polysemous word (field dependency) is checked using the sentence field attribute 150 to narrow down the polysemy. Here, since the text field attribute is "international politics", "Japanese and Chinese (proper noun)" whose semantic attribute is "country, region" is recognized (part of speech 195 after recognition).

【0022】図6は図4の多義語を変換する処理の概略
フローの実施例である。ここで、200は変換単語、2
10は変換後文字列である。意味的多義を絞った多義語
に対して、認定された単語見出しと品詞で変換単語辞書
90を検索して変換単語200を抽出し、原文章ファイ
ルの文字列と置換して変換後文字列210を作成する。
実施例では、「日中(時詞)」に対して「昼間の間」を
変換し、「日中(固有名詞)」に対しては「日本と中
国」を変換している。
FIG. 6 is an embodiment of the schematic flow of the processing for converting the polysemous words in FIG. Here, 200 is a conversion word, 2
10 is a character string after conversion. For a polysemous word with a narrowed meaning of meaning, the converted word dictionary 90 is searched with the recognized word headings and parts of speech to extract the converted word 200, and the converted character string 210 is replaced with the character string of the original text file. To create.
In the embodiment, "daytime" is converted to "daytime (timer)", and "Japan and China" is converted to "daytime (proper noun)".

【0023】このように、予め多義語の各多義毎に多義
語の単語見出しと品詞と置換する変換単語を収録した変
換単語辞書として作成しておき、多義語が抽出された場
合に、意味的多義を絞って認定した多義語の単語見出し
と品詞とで該変換単語辞書を検索して聞き取りやすい変
換単語への変換を行うのであるから、編集者の文章編集
の処理量を軽減できる。
As described above, a conversion word dictionary is prepared in advance for each polysemous word of the polysemous words, which contains the conversion words for replacing the word headings of the polysemous words and the parts of speech. Since the converted word dictionary is searched with the word headings and the parts of speech of the polysemous words that have been identified by narrowing down the meaning of the polysemous word, conversion into a conversion word that is easy to hear is performed, so that the editor can reduce the amount of text editing processing.

【0024】図7は図1に示した処理ブロック構成例に
おいて、日本文原文章ファイル10に含まれる動詞の同
音語を変換する処理の概略フローを示す図であり、概略
フローに従って、動作の説明を行う。
FIG. 7 is a diagram showing a schematic flow of a process for converting a homonym of a verb included in the Japanese sentence original sentence file 10 in the processing block configuration example shown in FIG. I do.

【0025】ここで、ステップ200、ステップ21
0、ステップ280、ステップ290はそれぞれ、図4
のステップ100、ステップ110、ステップ180、
ステップ190に等しい。
Here, step 200 and step 21
0, step 280, and step 290 are respectively shown in FIG.
Step 100, Step 110, Step 180,
Equivalent to step 190.

【0026】 同音語抽出処理部70において、認定された単語列の中から動詞の同音語を示 す情報フラグによって同音語を抽出し、同音語でない場合にはステップ280に 分岐する (ステップ220) 格要素意味属性抽出処理部80において、抽出した同音語について同音語の前 方の単語を探索して格要素単語を抽出し、格要素単語の意味属性を抽出する (ステップ230) 同音語の単語見出しと抽出した格要素単語の意味属性で、変換単語辞書90を 検索する (ステップ240) 変換単語辞書でマッチするパターンがあるかを判定する。マッチするパターン がない場合にはステップ280に分岐する (ステップ250) 変換単語を抽出し原文章ファイルの文字列を置換する (ステップ260) 編集済みの文章ファイル110に書き込む (ステップ270) 図8は図7の同音語を変換する処理の概略フローの実施
例である。
In the homophone extraction processing unit 70, a homophone is extracted from the recognized word string by an information flag indicating a homophone of a verb, and if it is not a homophone, the process branches to step 280 (step 220). In the case element semantic attribute extraction processing unit 80, a word preceding the homophone is searched for in the extracted homophone, the case element word is extracted, and the semantic attribute of the case element word is extracted (step 230). The converted word dictionary 90 is searched using the headline and the extracted semantic attribute of the case element word (step 240). It is determined whether there is a matching pattern in the converted word dictionary. If there is no matching pattern, the process branches to step 280 (step 250). The converted word is extracted and the character string of the original sentence file is replaced (step 260). The edited sentence file 110 is written (step 270). 9 is an example of a schematic flow of a process of converting the same phoneme in FIG. 7.

【0027】ここで、220は原文章文字列、230は
同音語(動詞)、240は格要素単語との格構造関係、
250は変換単語、260は変換後文字列である。抽出
した同音語「捕る」「採る」について、同音語の前方の
単語を探索して格要素単語を抽出し、格要素単語の意味
属性を抽出する。実施例では、「捕る」に対して『サケ
(生物)』、「採る」に対して『病原菌(生物)』と
『手段(制度)』がそれぞれ抽出される。次に、格要素
単語を有する同音語の単語見出しと抽出した格要素単語
の意味属性で、変換単語辞書90を検索し、変換単語辞
書でマッチするパターンがあるかを判定する。マッチす
るパターンがあれば変換単語250を抽出し、原文章フ
ァイルの文字列と置換して変換後文字列260を作成す
る。実施例では、「捕る」に対して『サケ(生物)』が
格構造関係にあるので、「捕る」を「捕獲する」に変換
している。また、「採る」に対しては、格要素単語が
『病原菌(生物)』に対して「採る」を「採取する」に
変換し、格要素単語が『手段(制度)』に対しては「採
る」を「採用する」に変換している。
Here, 220 is an original sentence character string, 230 is a homophone (verb), 240 is a case structure relationship with a case element word,
Reference numeral 250 is a converted word, and 260 is a character string after conversion. With respect to the extracted homophones “capture” and “take”, a word in front of the homonym is searched to extract a case element word, and a semantic attribute of the case element word is extracted. In the embodiment, “salmon (biological)” is extracted for “catching”, and “pathogen (biological)” and “means (system)” are extracted for “collecting”. Next, the converted word dictionary 90 is searched by the homonym word heading having the case element word and the meaning attribute of the extracted case element word, and it is determined whether there is a matching pattern in the converted word dictionary. If there is a matching pattern, the converted word 250 is extracted and replaced with the character string of the original text file to create the converted character string 260. In the embodiment, since "salmon (biological)" has a case structure relationship with "capture", "capture" is converted to "capture". In addition, for "take", the case element word is converted from "take" to "collect" for "pathogen (biological)", and the case element word is "for means (system)". "Take" is converted to "Take".

【0028】このように、予め同音語見出しと格要素単
語の意味属性との組み合わせ毎に置換する変換単語を収
録した変換単語辞書を作成しておき、同音語が抽出され
た場合に、格要素単語を有する同音語の単語見出しと抽
出した格要素単語の意味属性で、該変換単語辞書を検索
して聞き取りやすい変換単語への変換を行う。
In this way, a conversion word dictionary containing conversion words to be replaced for each combination of the homophone heading and the meaning attribute of the case element word is prepared in advance, and when the homophone is extracted, the case element is extracted. The converted word dictionary is searched using the word heading of the homophone having the word and the meaning attribute of the extracted case element word, and converted into a converted word that is easy to hear.

【0029】すなわち、本発明の日本文読み上げ単語変
換編集処理方式では、読み上げ用の日本文文章ファイル
に含まれる同音語や多義語を聞き取りやすい単語に変換
し編集する処理において、既に記述した手段によって、
留意すべき同音語や多義語を抽出する処理や前方の格要
素単語の意味属性による聞き取りやすい変換単語への置
換処理を行うことが可能となる。
That is, in the Japanese sentence reading word conversion / editing processing method of the present invention, in the process of converting and editing the homonyms and polysemous words contained in the Japanese sentence sentence file for reading into easily audible words, the means already described is used. ,
It is possible to perform processing for extracting homonyms and polysemous words that should be noted, and processing for replacement with a conversion word that is easy to hear based on the semantic attribute of the preceding case element word.

【0030】[0030]

【発明の効果】以上説明したように、本発明によれば、
特に日本文文章を読み上げて電話などの音声データで速
報記事などを配信する際には、読み上げ内容の聞き誤り
や聞きにくさなどを排除して了解度の高い日本文文章を
作成するために、一般に日本文文章の編集作業を行うに
当って、編集者の文章編集の処理量を軽減できる。
As described above, according to the present invention,
In particular, when reading out Japanese sentences and delivering breaking news articles using voice data such as telephones, in order to eliminate misunderstanding and difficulty of listening to the read contents and create Japanese sentences with high intelligibility, Generally, when editing Japanese sentences, it is possible to reduce the editor's processing amount of sentence editing.

【図面の簡単な説明】[Brief description of drawings]

【図1】本発明の処理ブロック構成例である。FIG. 1 is an example of a processing block configuration of the present invention.

【図2】日本語単語辞書の構成例である。FIG. 2 is a configuration example of a Japanese word dictionary.

【図3】変換単語辞書の構成例である。FIG. 3 is a configuration example of a conversion word dictionary.

【図4】処理概略フローである。FIG. 4 is a schematic process flow.

【図5】単語変換編集処理の実施例を説明する図であ
る。
FIG. 5 is a diagram illustrating an example of a word conversion editing process.

【図6】単語変換編集処理の実施例を説明する図であ
る。
FIG. 6 is a diagram illustrating an example of a word conversion editing process.

【図7】処理概略フローである。FIG. 7 is a schematic process flow.

【図8】単語変換編集処理の実施例を説明する図であ
る。
FIG. 8 is a diagram illustrating an example of a word conversion editing process.

【符号の説明】[Explanation of symbols]

10 日本文原文章ファイル 20 日本語単語辞書 30 文法辞書 40 形態素解析処理部 50 多義語抽出処理部 60 多義語判定処理部 70 同音語抽出処理部 80 格要素意味属性抽出処理部 90 変換単語辞書 100 単語変換編集処理部 110 編集済みの文章ファイル 120 処理装置 130 格要素単語の意味属性 140 変換単語 150 文章分野属性 160 原文章文字列 170 多義語 180 文法的単語接続条件 190 分野依存度判定条件 195 認定品詞 200 変換単語 210 変換後文字列 220 原文章文字列 230 同音語 240 格構造関係 250 変換単語 260 変換後文字列 10 Japanese sentence original sentence file 20 Japanese word dictionary 30 Grammar dictionary 40 Morphological analysis processing unit 50 Polysemous word extraction processing unit 60 Polysemous word determination processing unit 70 Homophone extraction processing unit 80 Case element meaning attribute extraction processing unit 90 Converted word dictionary 100 Word conversion edit processing unit 110 Edited sentence file 120 Processing device 130 Case element Semantic attribute of word 140 Converted word 150 Text field attribute 160 Original sentence character string 170 Polysemous word 180 Sentence word connection condition 190 Field dependency judgment condition 195 Certification Part-of-speech 200 Converted word 210 Converted character string 220 Original sentence character string 230 Homophone 240 Case structure relation 250 Converted word 260 Converted character string

Claims (2)

【特許請求の範囲】[Claims] 【請求項1】 読み上げ用の日本文文章ファイルに含ま
れる同音語や多義語を聞き取りやすい単語に変換し編集
する処理において、 読み上げ用の日本文原文章を日本語単語辞書と文法辞書
とを用いて形態素解析処理して単語の認定と単語の言語
情報とを取得する手段と、 予め日本語単語辞書の単語情報に多義語を示す情報フラ
グを設定する手段と、 該情報フラグによって多義語を抽出する手段と、 多義語で文法的な接続条件を満たす単語候補が複数存在
する場合に、文章ファイルに付随する文章分野属性と多
義語の意味属性との共起性で多義を絞り込む手段と、 予め多義語の各多義毎に多義語の単語見出しと品詞に対
して置換する変換単語とを収録した変換単語辞書と、 該変換単語辞書を検索して多義語を変換編集する手段と
を備えることを特徴とする日本文読み上げ単語変換編集
処理方式。
1. In a process of converting a homonym or a polysemous word included in a Japanese sentence text file for reading into a word that is easy to hear and editing the original Japanese sentence for reading, a Japanese word dictionary and a grammar dictionary are used. Morphological analysis processing to obtain word recognition and word linguistic information, means for setting an information flag indicating a polysemous word in the word information of the Japanese word dictionary in advance, and polysemous word extraction by the information flag And a means for narrowing down the polysemous word by the co-occurrence of the sentence field attribute attached to the text file and the semantic attribute of the polysemous word when there are multiple word candidates that satisfy the grammatical connection condition for the polysemous word, A conversion word dictionary in which a polysemous word heading and a conversion word to be replaced with respect to a part of speech are recorded for each polysemy, and means for converting and editing a polysemous word by searching the conversion word dictionary Japanese sentence reading word conversion editing processing method characterized by.
【請求項2】 予め日本語単語辞書の単語情報に動詞の
同音語を示す情報フラグを設定する手段と、 該情報フラグによって同音語を抽出する手段と、 同音語の前方に存在する格構造関係の名詞の意味属性を
抽出する手段と、 予め同音語見出しと格構造関係の名詞の意味属性との組
み合わせ毎に置換する変換単語を収録した変換単語辞書
と、 該変換単語辞書を検索して同音語を変換編集する手段と
を備えることを特徴とする請求項1記載の日本文読み上
げ単語変換編集処理方式。
2. A means for setting an information flag indicating a homonym of a verb in the word information of a Japanese word dictionary in advance, a means for extracting a homonym by the information flag, and a case structure relationship existing in front of the homonym. Means for extracting the semantic attributes of the nouns, a conversion word dictionary containing the conversion words to be replaced in advance for each combination of the homonym headings and the meaning attributes of the nouns having a case structure relationship, and the same word by searching the conversion word dictionary. 2. A Japanese sentence reading word conversion / editing method according to claim 1, further comprising means for converting and editing words.
JP4030232A 1992-02-18 1992-02-18 Japanese sentence reading word conversion edit processing method Pending JPH05224687A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP4030232A JPH05224687A (en) 1992-02-18 1992-02-18 Japanese sentence reading word conversion edit processing method

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP4030232A JPH05224687A (en) 1992-02-18 1992-02-18 Japanese sentence reading word conversion edit processing method

Publications (1)

Publication Number Publication Date
JPH05224687A true JPH05224687A (en) 1993-09-03

Family

ID=12297970

Family Applications (1)

Application Number Title Priority Date Filing Date
JP4030232A Pending JPH05224687A (en) 1992-02-18 1992-02-18 Japanese sentence reading word conversion edit processing method

Country Status (1)

Country Link
JP (1) JPH05224687A (en)

Cited By (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH07244672A (en) * 1994-03-04 1995-09-19 Sony Corp Electronic dictionary and natural language processor
US6389386B1 (en) 1998-12-15 2002-05-14 International Business Machines Corporation Method, system and computer program product for sorting text strings
US6411948B1 (en) 1998-12-15 2002-06-25 International Business Machines Corporation Method, system and computer program product for automatically capturing language translation and sorting information in a text class
US6460015B1 (en) 1998-12-15 2002-10-01 International Business Machines Corporation Method, system and computer program product for automatic character transliteration in a text string object
KR100377475B1 (en) * 1999-07-23 2003-03-26 한국전자통신연구원 Device for Extracting Korean Speech Act using Mood Information
US7099876B1 (en) 1998-12-15 2006-08-29 International Business Machines Corporation Method, system and computer program product for storing transliteration and/or phonetic spelling information in a text string class
WO2008117432A1 (en) * 2007-03-27 2008-10-02 Fujitsu Limited Electronic document anonymizing program

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH0335296A (en) * 1989-06-30 1991-02-15 Sharp Corp Text voice synthesizing device
JPH0420998A (en) * 1990-05-16 1992-01-24 Ricoh Co Ltd Voice synthesizing device

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH0335296A (en) * 1989-06-30 1991-02-15 Sharp Corp Text voice synthesizing device
JPH0420998A (en) * 1990-05-16 1992-01-24 Ricoh Co Ltd Voice synthesizing device

Cited By (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH07244672A (en) * 1994-03-04 1995-09-19 Sony Corp Electronic dictionary and natural language processor
US6389386B1 (en) 1998-12-15 2002-05-14 International Business Machines Corporation Method, system and computer program product for sorting text strings
US6411948B1 (en) 1998-12-15 2002-06-25 International Business Machines Corporation Method, system and computer program product for automatically capturing language translation and sorting information in a text class
US6460015B1 (en) 1998-12-15 2002-10-01 International Business Machines Corporation Method, system and computer program product for automatic character transliteration in a text string object
US7099876B1 (en) 1998-12-15 2006-08-29 International Business Machines Corporation Method, system and computer program product for storing transliteration and/or phonetic spelling information in a text string class
KR100377475B1 (en) * 1999-07-23 2003-03-26 한국전자통신연구원 Device for Extracting Korean Speech Act using Mood Information
WO2008117432A1 (en) * 2007-03-27 2008-10-02 Fujitsu Limited Electronic document anonymizing program
JPWO2008117432A1 (en) * 2007-03-27 2010-07-08 富士通株式会社 Electronic document concealment program
JP5337020B2 (en) * 2007-03-27 2013-11-06 富士通株式会社 Electronic document concealment program

Similar Documents

Publication Publication Date Title
CN109840331B (en) A Neural Machine Translation Method Based on User Dictionary
KR100453227B1 (en) Similar sentence retrieval method for translation aid
CN1954315B (en) System and method for translating Chinese pinyin to Chinese characters
Cussens Part-of-speech tagging using Progol
Zechner Automatic generation of concise summaries of spoken dialogues in unrestricted domains
GB2407657A (en) Automatic grammar generator comprising phase chunking and morphological variation
KR101279707B1 (en) Definition extraction
JP3992348B2 (en) Morphological analysis method and apparatus, and Japanese morphological analysis method and apparatus
US20100185438A1 (en) Method of creating a dictionary
Brown et al. Capitalization recovery for text
Wang et al. Evaluation of spoken language grammar learning in the ATIS domain
Sankaravelayuthan et al. A Comprehensive Study of Shallow Parsing and Machine Translation in Malaylam
JPH0877196A (en) Document information extraction device
JPH03105465A (en) Compound word extraction device
JPH11338863A (en) Unknown noun and katakana spelling automatic collection / authorization device, and recording medium recording processing procedure for it
JPS62271057A (en) Dictionary register system for translation device
Hong et al. A Korean morphological analyzer for speech translation system
Sang et al. Reduction of Dutch Sentences for Automatic Subtitling.
JP2008225744A (en) Machine translation apparatus and program
JP2897942B2 (en) Japanese morphological analysis system and morphological analysis method
Lihemo et al. The Syntax of Head-Marked Phrases and Head-Marking Morphemes in Lunyore
JPH05250403A (en) Japanese sentence word analyzing system
JPS63109572A (en) Derivative processing system
JPS6389976A (en) Language analyzer
JPH05233689A (en) Automatic document summarization method