JPH10207889A - Document proofing device - Google Patents
Document proofing deviceInfo
- Publication number
- JPH10207889A JPH10207889A JP9006588A JP658897A JPH10207889A JP H10207889 A JPH10207889 A JP H10207889A JP 9006588 A JP9006588 A JP 9006588A JP 658897 A JP658897 A JP 658897A JP H10207889 A JPH10207889 A JP H10207889A
- Authority
- JP
- Japan
- Prior art keywords
- error
- word
- unit
- candidate
- correct answer
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
Landscapes
- Document Processing Apparatus (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Machine Translation (AREA)
Abstract
(57)【要約】
【課題】 文書校正装置において、誤り部分(誤りの可
能性のある部分)に対する正解候補の過剰指摘を減らす
ようにすること。
【解決手段】 形態素解析部100はテキストを単語列
に変換する。誤り部分検出部200は、単語列の中から
誤り部分を検出する。正解候補展開部300は、誤り部
分に対応する正解候補群を生成する。正解候補検証部4
00は、正解候補群の中から本当に正解候補らしいもの
を選び出し、選び出された正解候補から成る検証ずみ正
解候補群を出力する。正解候補検証部400は、符号4
10〜440の部分から構成されている。正解確率付与
部410は、単語生起確率データベース440を参照し
て、各正確候補に対して生起確率を付与する。誤り確率
計算部420は、生起確率に基づいて正解候補の誤り確
率を計算する。誤り候補選択部430は、各正解候補の
誤り確率に基づいて、正解候補の絞り込みを行う。
(57) [Summary] To provide a document proofreading apparatus that reduces excessive indication of correct candidates for an erroneous part (a part that may have an error). A morphological analysis unit converts a text into a word string. The error part detector 200 detects an error part from the word string. The correct answer candidate developing unit 300 generates a correct answer candidate group corresponding to the error part. Correct answer candidate verification unit 4
00 selects a correct answer candidate from the correct answer candidate group and outputs a verified correct answer candidate group including the selected correct answer candidates. Correct answer candidate verifying section 400 uses code 4
10 to 440. The correct answer probability giving unit 410 gives the occurrence probability to each correct candidate with reference to the word occurrence probability database 440. The error probability calculation unit 420 calculates the error probability of the correct answer candidate based on the occurrence probability. The error candidate selection unit 430 narrows down the correct candidates based on the error probability of each correct candidate.
Description
【0001】[0001]
【発明の属する技術分野】本発明は、文章処理装置にお
いてユーザが入力した又は電子的な媒体として獲得した
文書データに対して、ユーザが文書を校正する作業を軽
減し、文書校正の効率を大幅に向上させる文書校正装置
に関するものである。BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention reduces the work of a user to proofread a document with respect to document data input by a user or obtained as an electronic medium in a text processing apparatus, and greatly improves the efficiency of document proofreading. The present invention relates to a document proofreading device that can be improved.
【0002】[0002]
【従来の技術】従来の誤り指摘技術としては、 形態素解析をして結果中の未登録語部分を指摘する
もの。 同音異義語のある単語を指摘するもの。 などが先ず挙げられる。2. Description of the Related Art As a conventional error indication technique, a morphological analysis is performed and an unregistered word part in a result is indicated. Points out words with homonyms. And the like.
【0003】未登録語を指摘する機能の場合、誤った綴
の単語があれば、未登録語となる確率が高いため、未登
録語部分の周辺に誤った綴の単語がある可能性がある。
同様に、同音異義語の存在する単語箇所は、仮名漢字変
換のときに操作誤りをし易い箇所として指摘される。ユ
ーザは、その中で自分で正誤の判断を一つ一つのケース
に対して下すことになる。In the case of a function for pointing out an unregistered word, if there is a word with an incorrect spelling, it is highly likely that the word will be an unregistered word. .
Similarly, a word location where a homonym exists is pointed out as a location where an operation error is likely to occur during kana-kanji conversion. The user makes the right or wrong judgment for each case in the user.
【0004】他の手段としては、形態素解析の後に、特
定の単語列が検出された場合に誤りと認定して指摘する
ものがある。例えば、名詞+動詞と言う品詞列をチェッ
クする又は一文字の漢字単語があった場合に誤りとする
等である。他にも片仮名/漢字文字列を発音順に並べ、
同じ単語の僅かな表記の揺れのある単語が隣に来るよう
にして、表記の揺れを検出し易くしたものがある。As another means, when a specific word string is detected after morphological analysis, it is recognized as an error and pointed out. For example, a part-of-speech sequence called noun + verb is checked, or an error occurs when there is a one-character kanji word. In addition, katakana / kanji character strings are arranged in phonetic order,
In some cases, a word with a slight spelling of the same word is located next to the spelling to make it easier to detect the sway of the spelling.
【0005】更に、新たに誤りの候補を検出した後で、
誤りの内容を推定した仮説を複数作り出し、複合語等と
のマッチング等の手段で仮説の検定を行い、生き残った
尤もらしい仮説のみを提示するシステムも存在する。Further, after a new error candidate is detected,
There is also a system that creates a plurality of hypotheses that estimate the content of an error, tests the hypotheses by means such as matching with a compound word or the like, and presents only the likely hypotheses that have survived.
【0006】[0006]
【発明が解決しようとする課題】未登録語,同音異義語
の存在する単語の指摘機能等は誤りと断定できないが、
誤りが存在する可能性がある所を指摘するわけである。
しかし、未登録語の指摘に関して言うと、未登録語の生
まれる原因としては、綴誤り以外にも固有名詞などが辞
書中に存在しないと言った本来の未登録語の存在も挙げ
られる。同音異義語の存在する単語の指摘についても、
誤りが多いと言うだけでは必ず誤っている箇所と言うわ
けではない。このため、上記の方法については、指摘さ
れたものが全て本当の誤りではない(過剰指摘が多い)
ということが一番問題になる。The function of pointing out an unregistered word or a word having a homonymous word cannot be determined to be incorrect.
It points out where errors may exist.
However, regarding the indication of an unregistered word, the cause of the unregistered word may be the existence of an original unregistered word that a proper noun or the like does not exist in the dictionary other than the spelling error. Regarding the indication of words with homonyms,
Just because there are many mistakes does not necessarily mean that there are mistakes. For this reason, in the above method, all the points pointed out are not true errors (there are many over points)
That is the most problematic.
【0007】特定の品詞列によって誤りを発見する方法
では、扱う誤りの対象が非常に限定されたものとなり、
文章中の誤りの多くは検出されないと言う問題を持つ。
また、片仮名語句や漢字語句をソートしてユーザに示す
方法は、ユーザ自身でするべき作業が大きく、校正作業
の能率が余り改善されないと言う問題点があった。In the method of detecting an error using a specific part-of-speech sequence, the target of the error to be handled is very limited.
The problem is that many errors in the text are not detected.
In addition, the method of sorting and showing katakana words and kanji words to the user has a problem that the work to be performed by the user is large and the efficiency of the proofreading work is not improved much.
【0008】さらに、仮説を生成して検定によって確か
らしいものだけを残す方法においては、生成された各々
の仮説に対して正しい評価を与えることが重要になる。
この場合は、本来の未登録語が辞書に載っていないと言
うだけで指摘されると言う問題はないが、評価の揺れが
問題になる。例えば、テキスト中の原表記に対応する単
語が辞書中に無かった場合は他の仮説に比べて相対的な
評価が低くなり、対象部分が正しい場合にも指摘してし
まう可能性がある。Further, in a method of generating hypotheses and leaving only probable ones by a test, it is important to give a correct evaluation to each generated hypothesis.
In this case, there is no problem that the original unregistered word is simply pointed out that it is not listed in the dictionary, but there is a problem of fluctuation in evaluation. For example, when the word corresponding to the original notation in the text is not in the dictionary, the relative evaluation is lower than other hypotheses, and there is a possibility that the word may be indicated even when the target portion is correct.
【0009】一般の文書校正支援システムでは、誤り指
摘の精度を高くしようとすれば対象とする誤りの種類を
絞らざるを得ず、また可能な限り多くの誤りを指摘しよ
うとすれば指摘中に本来の誤りでない部分に対する指摘
(過剰指摘)が多く混じってしまう。これに対応するた
めに、入力テキストに存在する表記誤りの可能性を広く
考慮して多くのもとの正しい綴りの候補を生成する部分
(正解候補展開)と,それを辞書の内容とのマッチング
によって検証する部分(正解語探索)を独立させた文書
校正支援システムを本出願人は既に提案したが、検証能
力が弱く、未だに多くの過剰指摘が残っている。本発明
は、これらの点に鑑みて創作されたものであって、統計
的なデータや辞書情報を利用して、正解候補の展開時に
生成される正解候補の誤り確率(正解候補が誤って誤り
部分の単語または単語列になる確率)を求めるようにな
った文書校正装置を提供することを目的としている。In a general document proofreading support system, in order to increase the accuracy of pointing out errors, it is necessary to narrow down the types of errors to be targeted. A lot of indications (excessive indications) for parts that are not original errors are mixed. In order to cope with this, a part that generates many original spelling candidates in consideration of the possibility of typographical errors existing in the input text (correct candidate expansion) and matching it with the contents of the dictionary The applicant of the present invention has already proposed a document proofreading support system in which the part to be verified (correct word search) is made independent, but the verification ability is weak, and there are still many points of excess. The present invention has been made in view of these points, and uses statistical data and dictionary information to provide an error probability of a correct answer candidate generated at the time of developing the correct answer candidate (correct correct candidate It is an object of the present invention to provide a document proofreading apparatus for obtaining a partial word or a word string).
【0010】[0010]
【課題を解決するための手段】請求項1の文書校正装置
は、入力されたテキストを単語列に変換する形態素解析
部と、形態素解析の結果得られた単語列の中から誤り可
能性部分を抽出する誤り部分検出部と、誤り部分抽出部
によって抽出された誤り可能性部分に対して正解候補を
生成する正解候補展開部と、正解候補展開部の展開の結
果得られた1個または複数個の正解候補のそれぞれに対
して検証を行って確からしい正解候補のみに絞り込む正
解候補検証部とを具備する文書校正装置であって、正解
候補検証部が、単語又は単語列の生起確率に関するデー
タベースと、上記データベースを参照して、正解候補の
誤り確率を計算するために必要とされる単語又は単語列
の生起確率を出力する生起確率付与部と、生起確率付与
部から出力される単語又は単語列の生起確率に基づい
て、各正解候補の確からしさの検定を行い、各正解候補
に対して誤り確率を付与する誤り確率計算部と、誤り確
率計算部によって各正解候補に付与された誤り確率を参
照して、所定の閾値以上の正解候補を選択する誤り候補
選択部とを具備することを特徴とするものである。According to a first aspect of the present invention, there is provided a document proofreading apparatus for converting a text to be input into a word string, and a morphological analysis unit for extracting a possible error portion from the word string obtained as a result of the morphological analysis. An error part detection unit to be extracted, a correct candidate expansion unit for generating a correct candidate for the error-possible part extracted by the error part extraction unit, and one or more obtained as a result of expansion of the correct candidate expansion unit A document proofreading apparatus comprising a correct answer candidate verifying unit that verifies each correct candidate and narrows down to only correct correct candidates, wherein the correct candidate verifying unit includes a database regarding the occurrence probability of a word or a word string. Referring to the database, an occurrence probability assigning unit that outputs an occurrence probability of a word or a word string required to calculate an error probability of a correct answer candidate, and an output that is output from the occurrence probability assigning unit. Based on the occurrence probability of a word or word string, a test of the probability of each correct answer candidate is performed, and an error probability calculation unit that gives an error probability to each correct candidate, and an error probability calculation unit assigns each correct candidate to each correct candidate. An error candidate selecting unit for selecting a correct answer candidate having a predetermined threshold or more with reference to the error probability.
【0011】請求項2の文書校正装置は、請求項1の文
書校正装置において、誤り確率計算部が、テキスト中に
存在する誤り可能性部分の生起確率と正解候補の生起確
率との比によって誤り確率を計算することを特徴とする
ものである。According to a second aspect of the present invention, there is provided the document proofreading apparatus according to the first aspect, wherein the error probability calculation unit calculates an error based on a ratio between an occurrence probability of an error-possible portion existing in the text and an occurrence probability of a correct candidate. It is characterized by calculating a probability.
【0012】請求項3の文書校正装置は、請求項1の文
書校正装置において、誤り確率計算部が、各正解候補の
単語が単独に生起する生起確率とテキスト中の文脈にお
ける単語列としての生起確率との比を参照して、各正解
候補に対する誤り確率を計算することを特徴とするもの
である。According to a third aspect of the present invention, there is provided the document proofreading apparatus according to the first aspect, wherein the error probability calculation unit includes an occurrence probability that each correct candidate word independently occurs and an occurrence probability as a word string in a context in the text. The error probability is calculated for each correct candidate with reference to the ratio to the probability.
【0013】請求項4の文書校正装置は、請求項1の文
書校正装置において、誤り確率計算部が、各正解候補が
テスト対象の単語群と共起する共起確率を計算し、計算
の結果得られた共起確率のパターンと,テキスト中の誤
り可能性部分が上記テスト対象の単語群と共起する共起
確率のパターンとの類似度によって誤り確率を計算する
ことを特徴とするものである。According to a fourth aspect of the present invention, in the document proofreading apparatus of the first aspect, the error probability calculation unit calculates a co-occurrence probability that each correct candidate co-occurs with a word group to be tested. The error probability is calculated based on the similarity between the obtained co-occurrence probability pattern and the co-occurrence probability pattern in which the error-probable part in the text co-occurs with the word group to be tested. is there.
【0014】請求項5の文書校正装置は、請求項1の文
書校正装置において、生起確率付与部が、展開される群
内での優先度情報を持つ展開群内優先度情報付き単語辞
書と、入力される単語又は単語列に対応する上記単語辞
書の群内における優先度情報に基づいて、上記単語又は
単語列に対する生起確率を計算する相対生起確率計算部
とを具備することを特徴とするものである。According to a fifth aspect of the present invention, there is provided the document proofreading apparatus according to the first aspect, wherein the occurrence probability assigning unit includes a word dictionary with priority information in the expanded group having priority information in the expanded group; A relative occurrence probability calculation unit that calculates an occurrence probability for the word or the word string based on the priority information in the group of the word dictionaries corresponding to the input word or word string. It is.
【0015】請求項6の文書校正装置は、請求項1,請
求項2,請求項3,請求項4または請求項5の文書校正
装置において、正解候補展開部が、読み付き単語辞書
と、読み付き単語辞書を参照して、誤り可能性部分の単
語の読みを抽出する読み抽出部と、読み抽出部によって
抽出された単語の読みと同一の読みを持つ他の単語を読
み付き単語辞書から抽出し、抽出した単語を正解候補と
して出力する同音語抽出部とを具備することを特徴とす
るものである。According to a sixth aspect of the present invention, there is provided the document proofreading apparatus according to the first, second, third, fourth, or fifth aspect, wherein the correct answer candidate developing unit includes a read word dictionary and a read word dictionary. A reading extraction unit that extracts the reading of the word in the error-possible part by referring to the attached word dictionary, and other words that have the same reading as the reading of the word extracted by the reading extraction unit are extracted from the reading word dictionary. And a homophone extraction unit that outputs the extracted word as a correct answer candidate.
【0016】請求項7の文書校正装置は、請求項1,請
求項2,請求項3,請求項4または請求項5の文書校正
装置において、正解候補展開部が、誤り表記,これに対
応する正解候補および制約条件を持つ展開データが複数
個記述された展開データベースと、誤り可能性部分に適
合する展開データベース中の展開データを用いて、誤り
可能性部分を正解候補に展開する展開部と、展開部から
出力される正解候補が当該正解候補に対する制約条件を
満たしているか否かを調査し、制約条件に合致する正解
候補だけを残す条件検査部とを具備することを特徴とす
るものである。According to a seventh aspect of the present invention, there is provided the document proofreading apparatus according to the first, second, third, fourth, or fifth aspect, wherein the correct candidate expanding section corresponds to an error notation. An expansion database that describes a plurality of expansion data having correct candidates and constraints, and an expansion unit that expands the error-possible portion into a correct answer candidate using the expansion data in the expansion database that matches the error-possible portion; A condition checking unit that checks whether the correct answer candidate output from the developing unit satisfies the constraint condition for the correct answer candidate and leaves only the correct answer candidate that matches the constraint condition. .
【0017】請求項8の文書校正装置は、請求項1,請
求項2,請求項3,請求項4または請求項5の文書校正
装置において、正解候補展開部が、複数の日本語入力手
段のそれぞれに対応する,誤り可能性部分を正解候補に
展開するための展開データベースの複数個と、テキスト
を作成した際の日本語入力手段を特定する情報に基づい
て、参照先の展開データベースを選択する参照先制御部
と、選択された参照先の展開データベースを参照して、
誤り可能性部分を正解候補に展開する展開処理部とを具
備することを特徴とするものである。According to an eighth aspect of the present invention, in the document proofreading apparatus according to any one of the first, second, third, fourth, and fifth aspects, the correct candidate developing unit includes a plurality of Japanese input means. Select a reference development database based on a plurality of development databases corresponding to each of the plurality of development databases for developing possible error portions into correct answers and information for specifying a Japanese input unit when the text is created. With reference to the reference destination control unit and the expanded database of the selected reference destination,
And a development processing unit that develops the error-possible part into a correct answer candidate.
【0018】請求項9の文書校正装置は、複数の展開デ
ータを持ち、最初に校正対象のテキストの一部に対して
各々の展開データによる訂正を行い、その結果最も評価
値が高い展開データを利用してテキスト全体の校正を行
うことを特徴とするものである。According to a ninth aspect of the present invention, there is provided a document proofreading apparatus which has a plurality of expanded data, first corrects a part of a text to be proofed with each expanded data, and as a result, develops expanded data having the highest evaluation value. It is characterized in that proofreading of the entire text is performed using the proofreading.
【0019】請求項1ないし請求項7の文書校正装置に
よれば、正解候補を過剰に指摘すると言うことを無くす
ことが出来る。また、請求項8および請求項9の文書校
正装置によれば、校正対象のテキストに適合した展開デ
ータを使用することが可能になる。According to the document proofreading devices of the first to seventh aspects, it is possible to eliminate the possibility that excessively pointing out correct candidates. Further, according to the document proofreading devices of the eighth and ninth aspects, it is possible to use expanded data suitable for the text to be proofread.
【0020】[0020]
【発明の実施の形態】図1は本発明の文書校正装置の構
成例を示す図である。同図においては、100は形態素
解析部、200は誤り検出部、300は正解候補展開
部、400は正解候補検証部、410は生起確率付与
部、420は誤り確率計算部、430は誤り候補選択部
をそれぞれ示している。FIG. 1 is a diagram showing a configuration example of a document proofreading apparatus according to the present invention. In the figure, 100 is a morphological analysis unit, 200 is an error detection unit, 300 is a correct answer candidate developing unit, 400 is a correct answer candidate verification unit, 410 is an occurrence probability assignment unit, 420 is an error probability calculation unit, and 430 is an error candidate selection unit. Each part is shown.
【0021】図1(a) は本発明の文書校正装置の概要を
示す図である。形態素解析部100は、入力テキストを
単語列に分解し、得られた単語列を誤り部分検出部20
0に渡す。誤り部分検出部200は、受け取った単語列
から誤り部分(誤りの可能性のある部分)を検出し、誤
り部分を正解候補展開部300に渡す。正解候補展開部
300では、誤りの種類を推定して、誤り部分に対応す
る正しい単語又は単語列の候補(正解候補)を生成す
る。正解候補検証部400は、各正解候補を検証して、
正解度の高い正解候補を選択する。なお、本発明の文書
校正装置は、実際には計算機とソフトウェアによって実
現されている。FIG. 1A is a diagram showing an outline of a document proofreading apparatus according to the present invention. The morphological analysis unit 100 decomposes the input text into word strings, and converts the obtained word strings into error part detection units 20.
Pass to 0. The erroneous part detection unit 200 detects an erroneous part (a part that may have an error) from the received word string, and passes the erroneous part to the correct answer candidate developing unit 300. The correct answer candidate developing unit 300 estimates the type of error and generates a correct word or word string candidate (correct answer candidate) corresponding to the error part. The correct candidate verification unit 400 verifies each correct candidate,
A correct answer candidate with a high degree of correctness is selected. Note that the document proofreading apparatus of the present invention is actually realized by a computer and software.
【0022】図1(b) は正解候補検証部の構成例を示す
図である。正解候補検証部400は、生起確率付与部4
10,誤り確率計算部420,誤り候補選択部430,
単語生起確率データベース440を有している。生起確
率付与部410は、単語単体や単語列の生起確率に関す
るデータベース440(単語生起確率データベース)を
参照して、正解候補の誤り確率を計算するために必要と
なる単語または単語列(正解候補や誤り部分の単語等)
の生起確率を出力する。単語や単語列の生起確率とは、
テキストやコーパス(文例集)の中で、単語または単語
列を任意に選択した場合に、それが指定された単語又は
単語列である確率を意味している。単語生起確率データ
ベースとは、 単語 生起確率 安全 0.001 保証 0.002 保障 0.001 歩しょう 0.0005 アーク 0.001 のように、単語又は単語列と生起確率の対を複数個記憶
するものである。FIG. 1B is a diagram showing an example of the configuration of the correct answer candidate verifying unit. The correct answer candidate verification unit 400 includes the occurrence probability assignment unit 4
10, error probability calculation section 420, error candidate selection section 430,
It has a word occurrence probability database 440. The occurrence probability giving unit 410 refers to the database 440 (word occurrence probability database) relating to the occurrence probabilities of a single word or a word string, and calculates a word or a word string (correct candidate or Erroneous words, etc.)
Output the probability of occurrence of What is the probability of occurrence of a word or word string?
When a word or word string is arbitrarily selected in a text or corpus (sentence collection), it means the probability that it is a specified word or word string. The word occurrence probability database is a database that stores multiple pairs of words or word strings and occurrence probabilities, such as word occurrence probability safety 0.001 guarantee 0.002 guarantee 0.001 walking 0.0005 arc 0.001 It is.
【0023】誤り確率計算部420は、生起確率付与部
410から出力される単語または単語列の生起確率をも
とにして、正解候補の誤り確率を計算する。誤り確率と
は、正解候補の単語又は単語列が誤って誤り部分の単語
又は単語列になる確率を意味している。誤り候補選択部
430は、誤り確率計算部420から渡された誤り確率
に基づいて、正解候補展開部300から出力される正解
候補群の中から正解候補に相応しいものを選び出す。The error probability calculation unit 420 calculates the error probability of the correct answer candidate based on the occurrence probability of the word or word string output from the occurrence probability giving unit 410. The error probability means a probability that a word or word string of a correct answer candidate is mistakenly changed to a word or word string of an error part. The error candidate selection unit 430 selects a candidate suitable for the correct candidate from the correct candidate group output from the correct candidate expanding unit 300 based on the error probability passed from the error probability calculation unit 420.
【0024】図2は誤り確率計算部における誤り確率計
算の第1の例を説明するための図である。図示の例で
は、原テキストが「松本斎藤両名の努力が実を結ぶ」と
なっている。誤り検出部200によって、誤り部分とし
て「松本」と「斎藤」が検出されたと仮定する。正解候
補展開部300は、同音異義語誤りと推定して、誤り部
分「松本」に対応して正解候補「松元」を生成し、誤り
部分「斎藤」に対応して正解候補「斉藤」を生成する。
生起確率付与部410は、単語生起確率データベース4
40を参照して、誤り部分「松本」に対して同音グルー
プ内での生起確率=0.1を付与し、正解候補「松元」
に対して同音グループ内での生起確率=0.02を付与
すると共に、誤り部分「斎藤」に対して同音グループ内
での生起確率=0.2を付与し、正解候補「斉藤」に対
して同音グループ内での生起確率=0.2を付与する。FIG. 2 is a diagram for explaining a first example of error probability calculation in the error probability calculation section. In the illustrated example, the original text is "Mr. Saito Matsumoto's efforts bear fruit". It is assumed that “Matsumoto” and “Saito” have been detected as error portions by the error detection unit 200. The correct answer candidate developing unit 300 estimates a homonym error, generates a correct answer candidate “Matsumoto” corresponding to the error part “Matsumoto”, and generates a correct answer candidate “Saito” corresponding to the error part “Saito”. I do.
The occurrence probability assigning unit 410 stores the word occurrence probability database 4
With reference to 40, the probability of occurrence in the same sound group = 0.1 is assigned to the erroneous part “Matsumoto”, and the correct answer candidate “Matsumoto”
And the probability of occurrence in the same sound group = 0.2 for the error part "Saito", and the correct answer candidate "Saito" for the correct part. The probability of occurrence in the same sound group = 0.2 is given.
【0025】誤り確率計算部420は、例えば 誤り確率=0.01×誤り先の生起確率/誤り元生起確率 …… (1) なる式によって正解候補の誤り確率を計算する。(1) 式
に誤り部分「松本」の生起確率=0.1,正解候補「松
元」の生起確率=0.02を代入すると、「松元」の誤
り確率=0.5となる。同様に、上式に誤り部分「斎
藤」の生起確率=0.2,正解候補「斉藤」の生起確率
=0.2を代入すると、「斉藤」の誤り確率=0.1と
なる。The error probability calculation unit 420 calculates the error probability of the correct answer candidate by using the following formula: error probability = 0.01 × error destination occurrence probability / error source occurrence probability, for example. When the occurrence probability of the error part “Matsumoto” = 0.1 and the occurrence probability of the correct answer candidate “Matsumoto” = 0.02 are substituted into the equation (1), the error probability of “Matsumoto” becomes 0.5. Similarly, substituting the probability of occurrence of the erroneous part “Saito” = 0.2 and the probability of occurrence of the correct answer candidate “Saito” = 0.2 into the above equation gives the error probability of “Saito” = 0.1.
【0026】図3は誤り確率計算部における誤り確率計
算の第2の例を説明するための図である。図示の例で
は、原テキストが「安全保障に関する話題」となってい
る。誤り検出部200によって、誤り部分として「保
証」が検出されたと仮定する。正解候補展開部300
は、同音異義語誤りと推定して、誤り部分「保証」に対
応して正解候補「保障」,「補償」を生成する。生起確
率付与部410は、単語生起確率データベース440を
参照して、誤り部分「保証」に対して同音グループ内で
の生起確率=0.2を付与し、正解候補「保障」に対し
て同音グループ内での生起確率=0.1を付与し、正解
候補「補償」に対して同音グループ内での生起確率=
0.1を付与する。また、生起確率付与部410は、文
脈における単語列「安全保障」に対して生起確率=0.
02を付与し、「安全保証」に対して生起確率=0.0
01を付与し、「安全補償」に対して生起確率=0.0
01を付与する。FIG. 3 is a diagram for explaining a second example of the error probability calculation in the error probability calculator. In the illustrated example, the original text is “topic on security”. It is assumed that “guarantee” is detected by the error detection unit 200 as an error part. Correct answer candidate expansion unit 300
Predicts a homonym error and generates correct answer candidates “guarantee” and “compensation” corresponding to the error part “guarantee”. The occurrence probability assigning unit 410 refers to the word occurrence probability database 440, assigns an occurrence probability of 0.2 in the same sound group to the error part “guarantee”, and assigns the same sound group to the correct answer candidate “guarantee”. Within the same sound group for the correct candidate "compensation" =
0.1 is given. Further, the occurrence probability giving unit 410 generates the occurrence probability = 0.0 for the word string “security” in the context.
02, and the probability of occurrence = 0.0
01, and the probability of occurrence = 0.0 for “safety compensation”
01 is assigned.
【0027】誤り確率計算部420は、 正解候補の誤り確率=文脈内生起確率/単独生起確率 …… (2) なる式によって、正解候補の誤り確率を計算する。(2)
式に「保証」,「保障」,「補償」,「安全保障」,
「安全保証」,「安全補償」の生起確率を代入すると、 「保障」の誤り確率=0.02/0.1=0.2 「保証」の誤り確率=0.001/0.2=0.005 「補償」の誤り確率=0.001/0.1=0.01 誤り候補選択部430は、誤り確率が最も大きい「保
障」を検証済み正解候補として出力する。The error probability calculation unit 420 calculates the error probability of the correct answer candidate according to the following formula: error probability of the correct answer = occurrence probability in context / single occurrence probability. (2)
In the formula, “guarantee”, “guarantee”, “compensation”, “security”,
When the occurrence probabilities of “security guarantee” and “security compensation” are substituted, the error probability of “guarantee” = 0.02 / 0.1 = 0.2 The error probability of “guarantee” = 0.001 / 0.2 = 0 0.005 Error probability of “compensation” = 0.001 / 0.1 = 0.01 The error candidate selection unit 430 outputs “guarantee” having the largest error probability as a verified correct answer candidate.
【0028】図4は誤り確率計算部における誤り確率計
算の第3の例を説明するための図である。図示の例で
は、原テキストが「服を換える」となっている。誤り検
出部200によって、誤り部分として「換える」が検出
されたと仮定する。正解候補展開部300は、同音異義
語誤りと推定して、誤り部分「換える」に対応して正解
候補「替える」,「買える」を生成する。FIG. 4 is a diagram for explaining a third example of the error probability calculation in the error probability calculator. In the illustrated example, the original text is “change clothes”. It is assumed that “change” is detected by the error detection unit 200 as an error part. The correct answer candidate developing section 300 estimates the homonymous error and generates correct answer candidates “change” and “buy” corresponding to the error part “change”.
【0029】生起確率付与部410は、単語生起確率デ
ータベース440から誤り部分「換える」と助詞
「に」,「が」の共起パターンを取出し、正解候補「替
える」と助詞「に」,「が」の共起パターンを取出し、
正解候補「買える」と助詞「に」,「が」の共起パター
ンを取り出す。図示の例では、共起パターンは、 共起パターン に が 換える ○ ○ 替える ○ ○ 買える × ○ となっている。The occurrence probability assigning unit 410 extracts the co-occurrence pattern of the error part “change” and the particles “ni” and “ga” from the word occurrence probability database 440, and outputs the correct answer candidate “change” and the particles “ni” and “ga”. ”
The co-occurrence pattern of the correct answer candidate “buy” and the particles “ni” and “ga” is extracted. In the illustrated example, the co-occurrence pattern is changed to a co-occurrence pattern.
【0030】誤り確率計算部420は、誤り部分の単語
の共起パターンと,正解候補の単語の共起パターンとを
比較し、比較結果に基づいて正解候補の誤り確率を算出
する。図示の例においては、誤り部分の単語「換える」
の共起パターンと正解候補の単語「替える」の共起パタ
ーンは同じであるので、「替える」の誤り確率は高くさ
れる。また、誤り部分の単語「換える」の共起パターン
と正解候補の単語「買える」の共起パターンは異なるの
で、「買える」の誤り確率は低くされる。The error probability calculator 420 compares the co-occurrence pattern of the word in the error part with the co-occurrence pattern of the correct candidate word, and calculates the error probability of the correct candidate based on the comparison result. In the illustrated example, the word “change” in the error part
Is the same as the co-occurrence pattern of the correct answer word "replace", the error probability of "replace" is increased. Further, since the co-occurrence pattern of the word “change” in the error part and the co-occurrence pattern of the correct candidate word “buy” are different, the error probability of “buy” is reduced.
【0031】図5は本発明の生起確率付与部の構成例を
示す図である。同図において、411は相対生起確率計
算部、412は生起確率書込み部、441は展開群内優
先度情報付き単語辞書をそれぞれ示している。FIG. 5 is a diagram showing an example of the configuration of the occurrence probability giving section according to the present invention. In the figure, reference numeral 411 denotes a relative occurrence probability calculation unit, 412 denotes an occurrence probability writing unit, and 441 denotes a word dictionary with priority information in the expanded group.
【0032】展開群内優先度情報付き単語辞書441と
は、ワープロの仮名漢字辞書のように、同音の群(これ
を展開群とする)の中で変換キーを押した時に最初に選
択される単語から単語が順に並べてあるものである。例
えば、「ほしょう」と言う展開群には、「保証」,「保
障」,「補償」,「歩しょう」と言う単語が記述されて
いる。この例であると、「保証」の生起確率>「保障」
の生起確率>「補償」の生起確率>「歩しょう」の生起
確率となる。例えば、展開群内の第n番目の単語と第n
−1番目の単語との間に0.001の生起確率の差があ
ると仮定すれば、相対的な生起確率が判る。The word dictionary 441 with priority information in the expanded group is first selected when a conversion key is pressed in a group of the same sound (this is an expanded group) like a kana-kanji dictionary of a word processor. Words are arranged in order from word. For example, in a development group called “hosho”, words “guarantee”, “guarantee”, “compensation”, and “walk” are described. In this example, the probability of occurrence of “guarantee”> “guarantee”
Occurrence probability> “compensation” occurrence probability> “walking” occurrence probability. For example, the nth word and the nth word in the expansion group
Assuming that there is a 0.001 occurrence probability difference from the first word, the relative occurrence probability is known.
【0033】相対生起確率計算部411には正解候補や
正解候補の誤り確率に関係する単語(又は単語列)が入
力される。相対生起確率計算部411は、展開群内優先
度情報付き単語辞書441を参照しながら、入力された
単語又は単語列の相対的な生起確率を計算する。生起確
率書込み部412は、相対生起確率計算部411に入力
された単語又は単語列に対して、相対的な生起確率を付
加するものである。The relative occurrence probability calculation unit 411 receives a word (or word string) related to the correct answer candidate and the error probability of the correct answer candidate. The relative occurrence probability calculation unit 411 calculates the relative occurrence probability of the input word or word string, with reference to the word dictionary 441 with the priority information in the expanded group. The occurrence probability writing unit 412 adds a relative occurrence probability to the word or word string input to the relative occurrence probability calculation unit 411.
【0034】図6は本発明の正解候補展開部の第1の構
成例を示す図である。同図において、311は読み抽出
部、312は同音語抽出部、313は読み付き単語表記
辞書をそれぞれ示している。FIG. 6 is a diagram showing a first example of the configuration of the correct answer candidate developing section according to the present invention. In the figure, reference numeral 311 denotes a reading extraction unit, reference numeral 312 denotes a homophone extraction unit, and reference numeral 313 denotes a word dictionary with reading.
【0035】読み付き単語表記辞書313には、 安全 あんぜん 保証 ほしょう 候補 こうほ というように、単語(又は単語列)と読みの対が複数個
格納されている。The read word notation dictionary 313 stores a plurality of pairs of words (or word strings) and readings, such as safety guarantees.
【0036】読み抽出部311には、誤り部分が入力さ
れる。読み抽出部311は、入力された誤り部分の表記
をキーとして読み付き単語表記辞書313を検索し、誤
り部分の読みを抽出する。抽出された読みは、同音語抽
出部312に渡される。同音語抽出部312は、渡され
た読みをキーとして読み付き単語表記辞書313を検索
し、同音異義語を抽出する。抽出された同音異義語は正
解候補として出力される。An error part is input to the reading extraction unit 311. The reading extraction unit 311 searches the word dictionary with reading 313 using the input notation of the error part as a key, and extracts the reading of the error part. The extracted reading is passed to the homophone extraction unit 312. The homophone extraction unit 312 searches the word dictionary with reading 313 using the passed pronunciation as a key, and extracts homonyms. The extracted homonyms are output as correct answer candidates.
【0037】図7は本発明の正解候補展開部の第2の構
成例を示す図である。同図において、321は展開部、
322は条件検査部、323は展開データベースをそれ
ぞれ示している。FIG. 7 is a diagram showing a second example of the configuration of the correct candidate expanding section according to the present invention. In the figure, reference numeral 321 denotes a developing unit,
Reference numeral 322 denotes a condition inspection unit, and 323 denotes a development database.
【0038】展開データベースとは、或る表記があり、
それが誤りだと仮定したときに元の正しい表記の候補
(正解候補)が書かれたものである。展開データベース
は おう→おお ず→づ づ→ず 保証→保障,補償 エイ→ エー というような展開データを格納している。例えば、「お
う→おお」という展開データの中で左側が誤り部分に対
応し、右側が正解候補に対応する。その他の展開データ
についても同じである。例えば、「むづかしい」という
単語があれば、「づ→ず」と言う展開データを利用し
て、「むずかしい」という正解候補を生成することが出
来る。The expansion database has a certain notation,
It is the original correct notation candidate (correct answer candidate) written assuming that it is incorrect. The development database stores development data such as → → → → づ づ → → ず ず 保証 → 保証 → → 保証 →. For example, the left side of the expanded data “Oh → Oh” corresponds to an error portion, and the right side corresponds to a correct answer candidate. The same applies to other expanded data. For example, if there is a word "difficult", a correct answer candidate "difficult" can be generated using expanded data "zu-zu".
【0039】展開データ中の正解候補は、自分自身,前
後の品詞,表記に関する制約条件を記述できるフォーマ
ットを持っている。例えば、展開データが 生→性(単語列の最後に来たときのみ有効) と言うものであれば、誤り部分「有効生」に対応して
「有効性」と言う正解候補を生成することが出来る。The correct answer candidate in the expanded data has a format that can describe its own, preceding and following parts of speech, and constraints on notation. For example, if the expanded data is raw → gender (valid only when it comes to the end of a word string), it is possible to generate a correct answer candidate called “validity” corresponding to the error part “valid raw”. I can do it.
【0040】展開部321には、誤り部分が入力され
る。展開部321は、展開データベース323を参照し
て、入力された誤り部分に対応する正解候補群を生成
し、この正解候補群を第1の正解候補群として出力す
る。第1の正解候補群は、条件検査部322に入力され
る。条件検査部322は、第1の正解候補群に属する正
解候補のそれぞれに付加されている制約条件を検査し、
制約条件に合致した正解候補の集まりのみを第2の正解
候補群として出力する。An error part is input to the developing unit 321. The expansion unit 321 generates a correct answer candidate group corresponding to the input error part with reference to the expansion database 323, and outputs the correct answer candidate group as a first correct answer candidate group. The first correct answer candidate group is input to the condition checking unit 322. The condition checking unit 322 checks the constraint added to each of the correct answer candidates belonging to the first correct answer candidate group,
Only a group of correct answer candidates meeting the constraint conditions is output as a second correct answer candidate group.
【0041】図8は本発明の正解候補展開部の第3の構
成例を示す図である。同図において、331は展開処理
部、332は参照先制御部、333ないし335は展開
データベースをそれぞれ示している。FIG. 8 is a diagram showing a third example of the configuration of the correct answer expanding section according to the present invention. In the figure, reference numeral 331 denotes an expansion processing unit, 332 denotes a reference destination control unit, and 333 to 335 denote expansion databases.
【0042】日本語入力手段としては、例えばOAKと
か,ATOKとか,MS−IMEとかが知られている。
例えば、展開データベース333はOAKに対応してお
り、展開データベース334はATOKに対応してお
り、展開データベース335はMS−IMEに対応して
いる。For example, OAK, ATOK, and MS-IME are known as Japanese input means.
For example, the development database 333 corresponds to OAK, the development database 334 corresponds to ATOK, and the development database 335 corresponds to MS-IME.
【0043】参照先制御部332は、日本語入力手段に
関する設定情報を計算機のオペレーティング・システム
又は文書の付加情報から収集して、それに最も適切な展
開データベースを選択する。展開処理部331は、選択
した展開データベースを参照して、入力された誤り部分
に対応する正解候補を生成する。The reference destination control section 332 collects setting information relating to the Japanese input means from the operating system of the computer or the additional information of the document, and selects the most appropriate development database. The expansion processing unit 331 refers to the selected expansion database and generates a correct answer candidate corresponding to the input error part.
【0044】図9は本発明の文書校正装置の他の構成例
を示す図である。同図において、501ないし503は
誤り訂正部、504は訂正性能比較評価部、505は選
択部、506はテキスト全体に対する訂正処理部をそれ
ぞれ示している。FIG. 9 is a diagram showing another example of the configuration of the document proofreading apparatus of the present invention. In the figure, reference numerals 501 to 503 denote an error correction unit, 504 denotes a correction performance comparison and evaluation unit, 505 denotes a selection unit, and 506 denotes a correction processing unit for the entire text.
【0045】誤り訂正部501〜503のそれぞれは、
図1(a) に示すような構成を有している。しかし、各誤
り訂正部で使用される展開データや制約条件などは、互
いに相違している。第1の誤り訂正部501,第2の誤
り訂正部502,第3の誤り訂正部503には、テキス
トの一部が入力される。訂正性能比較評価部504は、
自動的に又はユーザとの対話によって、各誤り訂正部に
よる訂正結果の相違部分を検出し、何が正しいかを評価
する。選択部505は、訂正性能比較評価部504の評
価結果に基づいて、最も訂正性能の良好な誤り訂正部を
選択する。選択された誤り訂正部を使用して、テキスト
全体に対する訂正処理が行われる。Each of the error correction units 501 to 503
It has a configuration as shown in FIG. However, expanded data and constraint conditions used in each error correction unit are different from each other. A part of the text is input to the first error correction unit 501, the second error correction unit 502, and the third error correction unit 503. The correction performance comparison and evaluation unit 504
Automatically or through interaction with the user, the difference between the correction results by the error correction units is detected, and what is correct is evaluated. The selection unit 505 selects an error correction unit having the best correction performance based on the evaluation result of the correction performance comparison evaluation unit 504. Correction processing is performed on the entire text using the selected error correction unit.
【0046】[0046]
【発明の効果】以上説明したように、本発明によれば、
正解候補をユーザに提示する又は次の検証のための仮説
として利用する際にも、全てを提示するのではなく、誤
り確率の高いものだけを示す又は誤り確率の高いものか
ら低いものへソートして順に提示する等の手段によっ
て、訂正率の改善やユーザの行う校正作業をより効率化
することが可能である。また、入力手段やユーザの癖な
どによる生起確率のバリエーションに対して、仮名漢字
変換辞書からのデータ抽出,展開種別の調整によって常
に最適な誤りの適合率と再現率を実現することが可能と
なる。As described above, according to the present invention,
When presenting the correct answer candidate to the user or using it as a hypothesis for the next verification, instead of presenting all, only show those with high error probability or sort from high error probability to low error probability For example, it is possible to improve the correction rate and to make the calibration work performed by the user more efficient by means such as presenting in order. In addition, it is possible to always achieve optimal error relevance and reproducibility for variations in the occurrence probability due to input means or user habits by extracting data from the kana-kanji conversion dictionary and adjusting the expansion type. .
【図1】本発明の文書校正装置の構成例を示す図であ
る。FIG. 1 is a diagram illustrating a configuration example of a document proofreading device of the present invention.
【図2】誤り確率計算部における誤り確率計算の第1の
例を示す図である。FIG. 2 is a diagram illustrating a first example of error probability calculation in an error probability calculation unit.
【図3】誤り確率計算部における誤り確率計算の第2の
例を示す図である。FIG. 3 is a diagram illustrating a second example of error probability calculation in the error probability calculation unit.
【図4】誤り確率計算部における誤り確率計算の第3の
例を示す図である。FIG. 4 is a diagram illustrating a third example of error probability calculation in the error probability calculation unit.
【図5】本発明の生起確率付与部の構成例を示す図であ
る。FIG. 5 is a diagram illustrating a configuration example of an occurrence probability providing unit according to the present invention.
【図6】本発明の正解候補展開部の第1の構成例を示す
図である。FIG. 6 is a diagram illustrating a first configuration example of a correct answer candidate developing unit according to the present invention.
【図7】本発明の正解候補展開部の第2の構成例を示す
図である。FIG. 7 is a diagram illustrating a second configuration example of the correct answer candidate developing unit according to the present invention.
【図8】本発明の正解候補展開部の第3の構成例を示す
図である。FIG. 8 is a diagram showing a third example of the configuration of the correct answer candidate developing unit according to the present invention.
【図9】本発明の文書構成装置の他の構成例を示す図で
ある。FIG. 9 is a diagram showing another configuration example of the document composition device of the present invention.
100 形態素解析部 200 誤り部分検出部 300 正解候補展開部 311 読み抽出部 312 同音語抽出部 313 読み付き単語表記辞書 321 展開部 322 条件検査部 323 展開データベース 331 展開処理部 332 参照先制御部 333 展開データベース 334 展開データベース 335 展開データベース 400 正解候補検証部 410 生起確率付与部 420 誤り確率計算部 430 誤り候補選択部 440 単語生起確率データベース 411 相対生起確率計算部 412 生起確率書込み部 441 展開群内優先度情報付き単語辞書 501 第1の誤り訂正部 502 第2の誤り訂正部 503 第3の誤り訂正部 504 訂正性能比較評価部 505 選択部 506 テキスト全体に対する訂正処理部 Reference Signs List 100 Morphological analysis unit 200 Error part detection unit 300 Correct answer candidate expansion unit 311 Reading extraction unit 312 Homophone extraction unit 313 Reading word notation dictionary 321 Expansion unit 322 Condition inspection unit 323 Expansion database 331 Expansion processing unit 332 Reference control unit 333 Expansion Database 334 Expansion database 335 Expansion database 400 Correct answer candidate verification unit 410 Occurrence probability assignment unit 420 Error probability calculation unit 430 Error candidate selection unit 440 Word occurrence probability database 411 Relative occurrence probability calculation unit 412 Occurrence probability writing unit 441 Expansion group priority information Attached word dictionary 501 first error correction unit 502 second error correction unit 503 third error correction unit 504 correction performance comparison evaluation unit 505 selection unit 506 correction processing unit for the entire text
───────────────────────────────────────────────────── フロントページの続き (51)Int.Cl.6 識別記号 FI G06F 15/40 370J ──────────────────────────────────────────────────の Continued on the front page (51) Int.Cl. 6 Identification code FIG06F 15/40 370J
Claims (9)
形態素解析部と、 形態素解析の結果得られた単語列の中から誤り可能性部
分を抽出する誤り部分検出部と、 誤り部分抽出部によって抽出された誤り可能性部分に対
して正解候補を生成する正解候補展開部と、 正解候補展開部の展開の結果得られた1個または複数個
の正解候補のそれぞれに対して検証を行って確からしい
正解候補のみに絞り込む正解候補検証部とを具備する文
書校正装置であって、 正解候補検証部が、 単語又は単語列の生起確率に関するデータベースと、 上記データベースを参照して、正解候補の誤り確率を計
算するために必要とされる単語又は単語列の生起確率を
出力する生起確率付与部と、 生起確率付与部から出力される単語又は単語列の生起確
率に基づいて、各正解候補の確からしさの検定を行い、
各正解候補に対して誤り確率を付与する誤り確率計算部
と、 誤り確率計算部によって各正解候補に付与された誤り確
率を参照して、所定の閾値以上の正解候補を選択する誤
り候補選択部とを具備することを特徴とする文書校正装
置。A morphological analysis unit that converts an input text into a word string; an error part detection unit that extracts an error-possible part from a word string obtained as a result of morphological analysis; and an error part extraction unit. A correct answer candidate developing unit that generates a correct answer candidate for the extracted error-possible part, and one or more correct answer candidates obtained as a result of the expansion of the correct candidate expanding unit are verified and verified. A document proofreading apparatus comprising: a correct answer candidate verifying unit that narrows down only possible correct answer candidates, wherein the correct candidate verifying unit refers to a database relating to the occurrence probability of a word or a word string; Based on the occurrence probability of the word or word string output from the occurrence probability giving unit, and the occurrence probability giving unit that outputs the occurrence probability of the word or word string required to calculate It makes the likelihood of the assay of the correct candidate,
An error probability calculation unit that assigns an error probability to each of the correct answer candidates; and an error candidate selection unit that selects an answer candidate having a predetermined threshold or more by referring to the error probability assigned to each of the correct answer candidates by the error probability calculation unit. A document proofreading device comprising:
る誤り可能性部分の生起確率と正解候補の生起確率との
比によって誤り確率を計算することを特徴とする請求項
1の文書校正装置。2. The document proofreading apparatus according to claim 1, wherein the error probability calculation unit calculates the error probability based on a ratio between an occurrence probability of an error-probable part existing in the text and an occurrence probability of a correct answer candidate. .
単独に生起する生起確率とテキスト中の文脈における単
語列としての生起確率との比を参照して、各正解候補に
対する誤り確率を計算することを特徴とする請求項1の
文書構成装置。3. An error probability calculation unit calculates an error probability for each correct candidate by referring to a ratio between an occurrence probability of each correct candidate word independently and an occurrence probability as a word string in a context in a text. 2. The document composition device according to claim 1, wherein the calculation is performed.
対象の単語群と共起する共起確率を計算し、計算の結果
得られた共起確率のパターンと,テキスト中の誤り可能
性部分が上記テスト対象の単語群と共起する共起確率の
パターンとの類似度によって誤り確率を計算することを
特徴とする請求項1の文書校正装置。4. An error probability calculation unit calculates a co-occurrence probability that each correct answer co-occurs with a word group to be tested, and a pattern of the co-occurrence probability obtained as a result of the calculation and a possibility of error in the text. 2. The document proofreading apparatus according to claim 1, wherein an error probability is calculated based on a degree of similarity with a co-occurrence probability pattern in which a part co-occurs with the test target word group.
報付き単語辞書と、 入力される単語又は単語列に対応する上記単語辞書の群
内における優先度情報に基づいて、上記単語又は単語列
に対する生起確率を計算する相対生起確率計算部とを具
備することを特徴とする請求項1の文書校正装置。5. An occurrence probability assigning unit, comprising: a word dictionary with priority information in an expansion group having priority information in the group to be expanded; and a word dictionary in the word dictionary corresponding to an input word or word string. 2. The document proofreading apparatus according to claim 1, further comprising: a relative occurrence probability calculation unit that calculates an occurrence probability for the word or the word string based on the priority information in the above.
読みを抽出する読み抽出部と、 読み抽出部によって抽出された単語の読みと同一の読み
を持つ他の単語を読み付き単語辞書から抽出し、抽出し
た単語を正解候補として出力する同音語抽出部とを具備
することを特徴とする請求項1,請求項2,請求項3,
請求項4または請求項5の文書校正装置。6. A correct answer candidate developing unit comprising: a reading word dictionary; a reading extraction unit for referring to the reading word dictionary to extract a reading of a word in an error-prone portion; and a word extracted by the reading extraction unit. 3. A homophone extraction unit for extracting another word having the same reading as the reading of the word from the word dictionary with reading, and outputting the extracted word as a correct answer candidate. Claim 3,
The document proofreading device according to claim 4 or 5.
つ展開データが複数個記述された展開データベースと、 誤り可能性部分に適合する展開データベース中の展開デ
ータを用いて、誤り可能性部分を正解候補に展開する展
開部と、 展開部から出力される正解候補が当該正解候補に対する
制約条件を満たしているか否かを調査し、制約条件に合
致する正解候補だけを残す条件検査部とを具備すること
を特徴とする請求項1,請求項2,請求項3,請求項4
または請求項5の文書校正装置。7. An expansion database in which a correct answer candidate expansion unit describes a plurality of expansion data having an error notation, a correct answer candidate corresponding thereto, and constraint conditions, and expansion data in the expansion database adapted to an error-possible portion. And a developing unit that expands the error-possible part into a correct candidate, and investigates whether the correct candidate output from the developing unit satisfies a constraint condition for the correct candidate, and determines whether the correct candidate matches the constraint condition. And a condition inspection unit that leaves only the condition.
Or the document proofreading device according to claim 5.
性部分を正解候補に展開するための展開データベースの
複数個と、 テキストを作成した際の日本語入力手段を特定する情報
に基づいて、参照先の展開データベースを選択する参照
先制御部と、 選択された参照先の展開データベースを参照して、誤り
可能性部分を正解候補に展開する展開処理部とを具備す
ることを特徴とする請求項1,請求項2,請求項3,請
求項4または請求項5の文書校正装置。8. A correct answer candidate developing unit, comprising: a plurality of expanded databases corresponding to each of the plurality of Japanese input means for expanding an error-possible portion into a correct answer candidate; A reference control unit for selecting a reference development database based on the information specifying the input means; and a development processing unit for referencing the selected reference development database and developing the error-prone part as a correct answer candidate 6. The document proofreading device according to claim 1, wherein the document proofreading device comprises:
象のテキストの一部に対して各々の展開データによる訂
正を行い、その結果最も評価値が高い展開データを利用
してテキスト全体の校正を行うことを特徴とする文書校
正装置。9. A text having a plurality of decompressed data, a part of the text to be proofread is first corrected by each decompressed data, and as a result, the decompressed data having the highest evaluation value is used for the entire text. A document proofreading device for performing proofreading.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP00658897A JP3856515B2 (en) | 1997-01-17 | 1997-01-17 | Document proofing device |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP00658897A JP3856515B2 (en) | 1997-01-17 | 1997-01-17 | Document proofing device |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPH10207889A true JPH10207889A (en) | 1998-08-07 |
| JP3856515B2 JP3856515B2 (en) | 2006-12-13 |
Family
ID=11642499
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP00658897A Expired - Fee Related JP3856515B2 (en) | 1997-01-17 | 1997-01-17 | Document proofing device |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JP3856515B2 (en) |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US8782000B2 (en) | 2010-03-19 | 2014-07-15 | Fujitsu Limited | Management device, correction candidate output method, and computer product |
| JP2019016140A (en) * | 2017-07-06 | 2019-01-31 | 株式会社朝日新聞社 | Calibration support device, calibration support method and calibration support program |
| CN114677694A (en) * | 2022-03-30 | 2022-06-28 | 深圳市福流网络信息科技有限公司 | A customs clearance method with intelligent identification technology |
-
1997
- 1997-01-17 JP JP00658897A patent/JP3856515B2/en not_active Expired - Fee Related
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US8782000B2 (en) | 2010-03-19 | 2014-07-15 | Fujitsu Limited | Management device, correction candidate output method, and computer product |
| JP2019016140A (en) * | 2017-07-06 | 2019-01-31 | 株式会社朝日新聞社 | Calibration support device, calibration support method and calibration support program |
| CN114677694A (en) * | 2022-03-30 | 2022-06-28 | 深圳市福流网络信息科技有限公司 | A customs clearance method with intelligent identification technology |
Also Published As
| Publication number | Publication date |
|---|---|
| JP3856515B2 (en) | 2006-12-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP4568774B2 (en) | How to generate templates used in handwriting recognition | |
| US5485372A (en) | System for underlying spelling recovery | |
| US9411800B2 (en) | Adaptive generation of out-of-dictionary personalized long words | |
| US5535121A (en) | System for correcting auxiliary verb sequences | |
| US7983903B2 (en) | Mining bilingual dictionaries from monolingual web pages | |
| US20060015320A1 (en) | Selection and use of nonstatistical translation components in a statistical machine translation framework | |
| JPH07325828A (en) | Grammar check system | |
| CN110147546B (en) | Grammar correction method and device for spoken English | |
| US7536296B2 (en) | Automatic segmentation of texts comprising chunks without separators | |
| Tufiş et al. | DIAC+: A professional diacritics recovering system | |
| JP2003099426A (en) | Natural language processor, its control method and program | |
| Uthayamoorthy et al. | Ddspell-a data driven spell checker and suggestion generator for the tamil language | |
| US20110229036A1 (en) | Method and apparatus for text and error profiling of historical documents | |
| JP5097802B2 (en) | Japanese automatic recommendation system and method using romaji conversion | |
| US20070179779A1 (en) | Language information translating device and method | |
| JP3309174B2 (en) | Character recognition method and device | |
| JP3856515B2 (en) | Document proofing device | |
| JP6303508B2 (en) | Document analysis apparatus, document analysis system, document analysis method, and program | |
| Huang et al. | Large scale experiments on correction of confused words | |
| Sharma et al. | Improving existing punjabi grammar checker | |
| JPH07325825A (en) | English grammar check system device | |
| JP4047895B2 (en) | Document proofing apparatus and program storage medium | |
| JP4318223B2 (en) | Document proofing apparatus and program storage medium | |
| JP3907106B2 (en) | Translation rule creation device and program | |
| JP2003132059A (en) | Search device, search system, search method, program, and recording medium using language sentence |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| A977 | Report on retrieval |
Free format text: JAPANESE INTERMEDIATE CODE: A971007 Effective date: 20050708 |
|
| A131 | Notification of reasons for refusal |
Free format text: JAPANESE INTERMEDIATE CODE: A131 Effective date: 20051004 |
|
| A521 | Written amendment |
Free format text: JAPANESE INTERMEDIATE CODE: A523 Effective date: 20051201 |
|
| A131 | Notification of reasons for refusal |
Free format text: JAPANESE INTERMEDIATE CODE: A131 Effective date: 20060523 |
|
| A521 | Written amendment |
Free format text: JAPANESE INTERMEDIATE CODE: A523 Effective date: 20060721 |
|
| TRDD | Decision of grant or rejection written | ||
| A01 | Written decision to grant a patent or to grant a registration (utility model) |
Free format text: JAPANESE INTERMEDIATE CODE: A01 Effective date: 20060912 |
|
| A61 | First payment of annual fees (during grant procedure) |
Free format text: JAPANESE INTERMEDIATE CODE: A61 Effective date: 20060912 |
|
| R150 | Certificate of patent or registration of utility model |
Free format text: JAPANESE INTERMEDIATE CODE: R150 |
|
| FPAY | Renewal fee payment (event date is renewal date of database) |
Free format text: PAYMENT UNTIL: 20090922 Year of fee payment: 3 |
|
| FPAY | Renewal fee payment (event date is renewal date of database) |
Free format text: PAYMENT UNTIL: 20100922 Year of fee payment: 4 |
|
| FPAY | Renewal fee payment (event date is renewal date of database) |
Free format text: PAYMENT UNTIL: 20100922 Year of fee payment: 4 |
|
| FPAY | Renewal fee payment (event date is renewal date of database) |
Free format text: PAYMENT UNTIL: 20110922 Year of fee payment: 5 |
|
| FPAY | Renewal fee payment (event date is renewal date of database) |
Free format text: PAYMENT UNTIL: 20120922 Year of fee payment: 6 |
|
| FPAY | Renewal fee payment (event date is renewal date of database) |
Free format text: PAYMENT UNTIL: 20120922 Year of fee payment: 6 |
|
| FPAY | Renewal fee payment (event date is renewal date of database) |
Free format text: PAYMENT UNTIL: 20130922 Year of fee payment: 7 |
|
| LAPS | Cancellation because of no payment of annual fees |