JPH0246976B2 - - Google Patents
Info
- Publication number
- JPH0246976B2 JPH0246976B2 JP54172474A JP17247479A JPH0246976B2 JP H0246976 B2 JPH0246976 B2 JP H0246976B2 JP 54172474 A JP54172474 A JP 54172474A JP 17247479 A JP17247479 A JP 17247479A JP H0246976 B2 JPH0246976 B2 JP H0246976B2
- Authority
- JP
- Japan
- Prior art keywords
- kana
- candidate word
- character string
- kanji
- candidate
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired - Lifetime
Links
Landscapes
- Document Processing Apparatus (AREA)
Description
【発明の詳細な説明】
本発明は、漢字仮名混じりの日本語文章をその
読みであるカナ文字列に変換する漢字仮名変換装
置に関するものである。DETAILED DESCRIPTION OF THE INVENTION The present invention relates to a kanji-kana conversion device that converts a Japanese text containing kanji and kana into a kana character string that is the reading thereof.
従来から仮名文字列を漢字仮名混じりの日本語
文章に変換する仮名漢字変換装置については種々
の型式のものが提案されているが、漢字仮名混じ
りの日本語文章をその読みである仮名文字列に変
換することは行われていない。漢字仮名混じりの
日本語文章をその読みである仮名文字列に変換す
る漢字仮名変換装置は、例えば本を朗読する機械
などに必要なものであり、また、仮名漢字変換装
置における語変換を修正する場合などにおいても
有効な手段となるものである。 Various types of kana-kanji conversion devices have been proposed for converting kana character strings into Japanese text containing kanji and kana, but there is a method that converts Japanese text containing kanji and kana into the kana character string that is its reading. No conversion has been done. A kanji-kana conversion device that converts Japanese sentences containing kanji and kana into kana character strings that are the pronunciation is necessary for, for example, a machine that reads books aloud, and also corrects the word conversion in the kana-kanji conversion device. It is also an effective means in various situations.
本発明は、上記の要求に応えるものであつて、
漢字仮名混じりの日本語文章をその読みである仮
名文字列に変換できるようになつた漢字仮名変換
装置を提供することを目的としている。そしてそ
のため、本発明の漢字仮名変換装置は、
漢字仮名混じり文をその読みである仮名文字列
に変換する漢字仮名変換装置であつて、
漢字仮名混じり文を入力できる入力装置1と、
漢字表記もしくは仮名表記もしくは漢字仮名ま
じり表記の単語からなるキー、単語の読み、文法
情報および使用頻度が書き込まれた複数のエント
リを有する単語辞書5と、
品詞間の接続可否情報が書き込まれた文法辞書
6と、
解析装置3と、
候補単語抽出装置4と
を具備し、
上記候補単語抽出装置4は、上記解析装置3か
ら与えられる解析位置より右側の文字列をキーと
して上記単語辞書5を検索し、上記文法辞書6を
参照して検索の結果得られる単語の中から直前の
最尤候補単語と文法的に接続可能な候補単語を選
択し、選択の結果得られた候補単語および候補単
語の読みを含む関連情報を上記解析装置3に通知
するように構成され、
上記解析装置3は、最初は解析位置を入力文字
列の左端に設定して解析位置の右に続く入力文字
列を上記候補単語抽出装置4に与えると共に、上
記候補単語抽出装置4から候補単語および関連情
報が通知されたときには、候補単語の中から、単
語の文字数と使用頻度を評価して最尤候補単語を
選択し、最尤候補単語の文字数だけ解析位置を右
にずらし、解析位置が入力文字列の右端でないこ
とを条件に入力文字列における解析位置の右に続
く文字列を上記候補単語抽出装置4に与えるよう
に構成されている
ことを特徴とするものである。 The present invention meets the above requirements, and includes:
The purpose of the present invention is to provide a kanji-kana conversion device capable of converting a Japanese sentence containing kanji and kana into a kana character string that is its reading. Therefore, the kanji-kana conversion device of the present invention is a kanji-kana conversion device that converts a kanji-kana mixed sentence into a kana character string that is its reading, and includes an input device 1 capable of inputting a kanji-kana mixed sentence, and a kanji notation or A word dictionary 5 having a plurality of entries in which are written keys consisting of words written in kana or written in kanji and kana, word pronunciations, grammatical information, and frequency of use; and a grammar dictionary 6 in which information on connectivity between parts of speech is written. , an analysis device 3, and a candidate word extraction device 4, the candidate word extraction device 4 searches the word dictionary 5 using the character string on the right side of the analysis position given from the analysis device 3 as a key, and A candidate word that can be grammatically connected to the previous most likely candidate word is selected from among the words obtained as a result of the search with reference to the grammar dictionary 6, and the candidate word obtained as a result of the selection and the pronunciation of the candidate word are included. The analysis device 3 is configured to notify the analysis device 3 of related information, and the analysis device 3 initially sets the analysis position to the left end of the input character string, and uses the input character string continuing to the right of the analysis position to the candidate word extraction device. 4, and when candidate words and related information are notified from the candidate word extraction device 4, the maximum likelihood candidate word is selected from among the candidate words by evaluating the number of characters and frequency of use of the word, and the maximum likelihood candidate word is selected. It is configured to shift the analysis position to the right by the number of characters in the word, and provide the candidate word extraction device 4 with a character string that continues to the right of the analysis position in the input character string, provided that the analysis position is not at the right end of the input character string. It is characterized by the presence of
以下、本発明を図面を参照しつつ説明する。 Hereinafter, the present invention will be explained with reference to the drawings.
第1図は本発明の1実施例のブロツク図、第2
図は第1図の実施例の動作説明図である。 FIG. 1 is a block diagram of one embodiment of the present invention, and FIG.
The figure is an explanatory diagram of the operation of the embodiment of FIG. 1.
第1図において、1は入力装置、2は出力装
置、3は解析装置、4は候補単語抽出装置、5は
単語辞書、6は文法辞書をそれぞれ示している。
入力装置1は例えば光学文字読取装置や実時間手
書文字入力装置などである。出力装置2は、漢字
仮名変換装置によつて得られた仮名文字列を他装
置へ送出する部分である。解析装置3は、全体を
制御するものと考えて良く、例えばキーとなる漢
字仮名混じり文を候補単語抽出装置4へ渡す機能
や文の解析位置を移動させる機能、最尤候補を選
択する機能、求まつた読みである仮名を連結して
出力装置2へ渡す機能などを有している。候補単
語抽出装置4は、解析装置3から送られて来た漢
字仮名混じり文をキーとして単語辞書を索引して
一致する単語を取出し、文法辞書を参照して上記
単群の中から候補単語を取出す機能を有している
単語辞書5は複数のエントリを有しており、各ン
トリにはキーとなる漢字(平仮名を含む)、読み
(片仮名)、文法情報および使用頻度などがJIS漢
字コードで書込まれている。上記の文法情報と
は、例えば品詞の別や活用形に関する情報を意味
している。文法辞書6は、どのような種類の単語
とどのような単語とが接続可能であるかを示すも
のであつて、行列の第1行および第1列に各種の
品詞が書込まれ、その交点に接続可能であるか否
かを示す情報が書込まれている。 In FIG. 1, 1 is an input device, 2 is an output device, 3 is an analysis device, 4 is a candidate word extraction device, 5 is a word dictionary, and 6 is a grammar dictionary.
The input device 1 is, for example, an optical character reader or a real-time handwritten character input device. The output device 2 is a part that sends the kana character string obtained by the kanji-kana conversion device to another device. The analysis device 3 can be thought of as controlling the whole, and includes, for example, a function of passing a key sentence containing kanji and kana to the candidate word extraction device 4, a function of moving the analysis position of a sentence, a function of selecting a maximum likelihood candidate, It has a function of concatenating the kana with the readings found and passing them to the output device 2. The candidate word extraction device 4 indexes the word dictionary using the kanji/kana mixed sentence sent from the analysis device 3 as a key, extracts matching words, and refers to the grammar dictionary to select candidate words from the single group. The word dictionary 5, which has a retrieval function, has multiple entries, and each entry contains key kanji (including hiragana), reading (katakana), grammatical information, usage frequency, etc. in JIS kanji code. It is written. The above-mentioned grammatical information means, for example, information regarding parts of speech and conjugations. The grammar dictionary 6 shows what types of words can be connected with what words, and various parts of speech are written in the first row and first column of the matrix, and their intersection points are written in the first row and first column of the matrix. Information indicating whether connection is possible is written.
次に第1図の実施例の動作を第2図を参照しつ
つ説明する。漢字仮名混じり文を仮名文字列に変
換する場合、解析装置3は、iを1にセツトする
と共に解析位置を文の左端にセツトする。 Next, the operation of the embodiment shown in FIG. 1 will be explained with reference to FIG. 2. When converting a sentence containing kanji and kana into a kana character string, the analysis device 3 sets i to 1 and sets the analysis position to the left end of the sentence.
解析装置3は、解析位置より右の文を候補単語
抽出装置4へ渡す。候補単語抽出装置4は、与え
られた文をキーとして、単語辞書5から一致する
単語の全てを取出す。この取出しの方法について
説明する。入力文字列をa1…aoとし、また単語列
に変換されていない部分入力文字列の左端の入力
文字をajとする。この場合、候補単語抽出装置4
は、
aj…aj+k(1≦j≦n,1≦j+k≦n)
なる見出しを持つた単語を全て単語辞書5から取
出す。 The analysis device 3 passes the sentence to the right of the analysis position to the candidate word extraction device 4. The candidate word extraction device 4 extracts all matching words from the word dictionary 5 using the given sentence as a key. This extraction method will be explained. Let the input character string be a 1 ...a o , and let the leftmost input character of the partial input character string that has not been converted into a word string be a j . In this case, candidate word extraction device 4
extracts from the word dictionary 5 all words with the heading a j ...a j+k (1≦j≦n, 1≦j+k≦n).
例えば、解析位置の右に「日本語文の」という
漢字仮名文字列が存在する場合、「日本」、「日本
語」、「日本語文」という単語並びにその関連情報
(即ち、読み、文法情報、使用頻度など)が単語
辞書5から取出される。一致する全ての単語が取
出された後、候補単語抽出装置4は文法辞書6を
参照して最尤候補(i−1)と文法的に接続可能
なもののみの集合を{候補}iとする。なお、文
頭の単語に対しては、文頭文法規則が適用され
る。候補単語抽出装置4は、候補単語および関連
情報(即ち、読み、文法情報、使用頻度など)を
解析装置3に渡す。解析装置3は、{候補}iが
存在する場合には、単語の文字数(この場合には
漢字数)と使用頻度を評価して最尤候補(i)を
選択し、最尤候補(i)の文字数だけ解析位置を
右にずらし、解析位置が文の右端にあるか否かを
チエツクする。解析位置が文の右端に存在しない
場合には、解析装置3はiをi+1に更新し、解
析位置より右の文字列を候補単語抽出装置4に渡
す。これにより、上記の動作が繰返される。解析
位置が文の右端に存在する場合には終了1とされ
る。終了1において、解析装置3は最尤候補
(1)ないし最尤候補(i)の読みをつなぎ、こ
れを出力装置2に渡す。 For example, if there is a kanji-kana character string "Japanese sentence" to the right of the analysis position, the words "Japan", "Japanese", "Japanese sentence" and their related information (i.e., reading, grammatical information, usage frequency, etc.) are taken out from the word dictionary 5. After all matching words have been extracted, the candidate word extraction device 4 refers to the grammar dictionary 6 and sets {candidate}i the set of only words that can be grammatically connected to the maximum likelihood candidate (i-1). . Note that sentence-initial grammar rules are applied to words at the beginning of sentences. The candidate word extraction device 4 passes the candidate words and related information (ie, pronunciation, grammatical information, usage frequency, etc.) to the analysis device 3. If {candidate} i exists, the analysis device 3 evaluates the number of characters in the word (in this case, the number of kanji characters) and the frequency of use, selects the maximum likelihood candidate (i), and selects the maximum likelihood candidate (i). Shift the analysis position to the right by the number of characters, and check whether the analysis position is at the right end of the sentence. If the analysis position does not exist at the right end of the sentence, the analysis device 3 updates i to i+1 and passes the character string to the right of the analysis position to the candidate word extraction device 4. This causes the above operation to be repeated. If the analysis position is at the right end of the sentence, it is considered to be the end 1. At end 1, the analysis device 3 connects the readings of the maximum likelihood candidate (1) to the maximum likelihood candidate (i) and passes this to the output device 2.
{候補}iが存在しないことが通知されたとき
には、解析装置3は、iが1であるか否かをチエ
ツクし、iが1でなければiをi−1として解析
位置を最尤候補(i)の文字数だけ左へずらし、
{候補}iから最尤候補(i)を除去して、得ら
れた候補単語の集合に対して図示説明した処理を
繰返す。終了2に到達した場合には仮名文字列が
求まらなかつたことになる。 {Candidate} When notified that i does not exist, the analysis device 3 checks whether or not i is 1, and if i is not 1, sets i to i-1 and sets the analysis position to the maximum likelihood candidate ( Shift it to the left by the number of characters in i),
{Candidate} The maximum likelihood candidate (i) is removed from i, and the process illustrated and explained is repeated for the obtained set of candidate words. If end 2 is reached, it means that the kana character string has not been found.
以上の説明から明らかなように、本発明によれ
ば、漢字仮名混じり文を正しい読みの漢字文字列
に変換する漢字仮名変換装置を得ることが出来
る。 As is clear from the above description, according to the present invention, it is possible to obtain a kanji-kana conversion device that converts a sentence containing kanji and kana into a kanji character string with the correct reading.
第1図は本発明の1実施例のブロツク図、第2
図は第1図の実施例の説明図である。
1…入力装置、2…出力装置、3…解析装置、
4…候補単語抽出装置、5…単語辞書、6…文法
辞書。
FIG. 1 is a block diagram of one embodiment of the present invention, and FIG.
The figure is an explanatory diagram of the embodiment of FIG. 1. 1... Input device, 2... Output device, 3... Analysis device,
4... Candidate word extraction device, 5... Word dictionary, 6... Grammar dictionary.
Claims (1)
列に変換する漢字仮名変換装置であつて、 漢字仮名混じり文を入力できる入力装置1と、 漢字表記もしくは仮名表記もしくは漢字仮名ま
じり表記の単語からなるキー、単語の読み、文法
情報および使用頻度が書き込まれた複数のエント
リを有する単語辞書5と、 品詞間の接続可否情報が書き込まれた文法辞書
6と、 解析装置3と、 候補単語抽出装置4と を具備し、 上記候補単語抽出装置4は、上記解析装置3か
ら与えられる解析位置より右側の文字列をキーと
して上記単語辞書5を検索し、上記文法辞書6を
参照して検索の結果得られる単語の中から直前の
最尤候補単語と文法的に接続可能な候補単語を選
択し、選択の結果得られた候補単語および候補単
語の読みを含む関連情報を上記解析装置3に通知
するように構成され、 上記解析装置3は、最初は解析位置を入力文字
列の左端に設定して解析位置の右に続く入力文字
列を上記候補単語抽出装置4に与えると共に、上
記候補単語抽出装置4から候補単語および関連情
報が通知されたときには、候補単語の中から、単
語の文字数と使用頻度を評価して最尤候補単語を
選択し、最尤候補単語の文字数だけ解析位置を右
にずらし、解析位置が入力文字列の右端でないこ
とを条件に入力文字列における解析位置の右に続
く文字列を上記候補単語抽出装置4に与えるよう
に構成されている ことを特徴とする漢字仮名変換装置。[Scope of Claims] 1. A kanji-kana conversion device that converts a kanji-kana mixed sentence into a kana character string that is its reading, comprising an input device 1 capable of inputting a kanji-kana mixed sentence, and a kanji notation, a kana notation, or a kanji-kana character string. A word dictionary 5 that has a plurality of entries in which keys consisting of words written in mixed spellings, word pronunciations, grammatical information, and frequency of use are written; a grammar dictionary 6 in which information on connectivity between parts of speech is written; and an analysis device 3. , a candidate word extraction device 4, the candidate word extraction device 4 searches the word dictionary 5 using the character string on the right side of the analysis position given from the analysis device 3 as a key, and refers to the grammar dictionary 6. From among the words obtained as a result of the search, a candidate word that can be grammatically connected to the previous most likely candidate word is selected, and the candidate word obtained as a result of the selection and related information including the candidate word's pronunciation are analyzed as described above. The analysis device 3 initially sets the analysis position to the left end of the input character string and provides the input character string continuing to the right of the analysis position to the candidate word extraction device 4, and When candidate words and related information are notified from the candidate word extraction device 4, the most likely candidate word is selected by evaluating the number of characters and frequency of use of the word from among the candidate words, and only the number of characters of the most likely candidate word is analyzed. The candidate word extraction device 4 is configured to shift the position to the right and provide the candidate word extraction device 4 with a character string that continues to the right of the analysis position in the input character string on the condition that the analysis position is not at the right end of the input character string. Kanji-kana conversion device.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP17247479A JPS5692677A (en) | 1979-12-26 | 1979-12-26 | Kanji (chinese character)/kana (japanese syllabary) converter |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP17247479A JPS5692677A (en) | 1979-12-26 | 1979-12-26 | Kanji (chinese character)/kana (japanese syllabary) converter |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPS5692677A JPS5692677A (en) | 1981-07-27 |
| JPH0246976B2 true JPH0246976B2 (en) | 1990-10-18 |
Family
ID=15942650
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP17247479A Granted JPS5692677A (en) | 1979-12-26 | 1979-12-26 | Kanji (chinese character)/kana (japanese syllabary) converter |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPS5692677A (en) |
Families Citing this family (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0612539B2 (en) * | 1982-10-06 | 1994-02-16 | 株式会社日立製作所 | Kanji / Kana conversion device |
| JPH087747B2 (en) * | 1986-01-20 | 1996-01-29 | カシオ計算機株式会社 | Kana-Kanji mutual conversion device |
| JPH0731674B2 (en) * | 1986-01-28 | 1995-04-10 | カシオ計算機株式会社 | Kana-Kanji mutual conversion device |
| JPH0731675B2 (en) * | 1986-01-28 | 1995-04-10 | カシオ計算機株式会社 | Kana-Kanji mutual conversion device |
| JPS62189568A (en) * | 1986-02-15 | 1987-08-19 | Casio Comput Co Ltd | Kana-kanji interchange device |
-
1979
- 1979-12-26 JP JP17247479A patent/JPS5692677A/en active Granted
Also Published As
| Publication number | Publication date |
|---|---|
| JPS5692677A (en) | 1981-07-27 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| KR100656736B1 (en) | System and method for disambiguating phonetic input | |
| JPH03185561A (en) | Method for inputting european word | |
| Abd Alhadi et al. | Automatic Identification of Rhetorical Elements in classical Arabic Poetry. | |
| JPH10269204A (en) | Automatic Chinese document proofing method and device | |
| JPS584424A (en) | Japanese word input device | |
| JP2659700B2 (en) | Kana-Kanji conversion method | |
| JPS634206B2 (en) | ||
| JPS58123129A (en) | Converting device of japanese syllabary to chinese character | |
| JP2004206659A (en) | Reading information determination method and apparatus and program | |
| JPS62251986A (en) | Misread character correction processor | |
| JPS6229796B2 (en) | ||
| JP2002535768A (en) | Method and apparatus for Kanji input | |
| JP4232957B2 (en) | English and other languages bilingual dictionary of phoneme index multi-element matrix structure | |
| JP2821143B2 (en) | Morphological decomposition device | |
| JPS63316162A (en) | Document preparing device | |
| JPH0441399Y2 (en) | ||
| JPH0146895B2 (en) | ||
| JPH0475162A (en) | Japanese syllabary/chinese character conversion device | |
| JPS62224859A (en) | Japanese language processing system | |
| JPS6133569A (en) | "kana"/"kanji" converter | |
| JP3048793B2 (en) | Character converter | |
| JPH0441398Y2 (en) | ||
| JPH0991278A (en) | Document preparation device | |
| JPS6298456A (en) | Japanese language input device | |
| JPS6175467A (en) | Kana-kanji conversion method |