JPH0760433B2 - Kanji converter - Google Patents
Kanji converterInfo
- Publication number
- JPH0760433B2 JPH0760433B2 JP61287031A JP28703186A JPH0760433B2 JP H0760433 B2 JPH0760433 B2 JP H0760433B2 JP 61287031 A JP61287031 A JP 61287031A JP 28703186 A JP28703186 A JP 28703186A JP H0760433 B2 JPH0760433 B2 JP H0760433B2
- Authority
- JP
- Japan
- Prior art keywords
- syllable
- dictionary
- kanji
- mother
- input
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired - Fee Related
Links
- 238000006243 chemical reaction Methods 0.000 claims description 36
- 230000000873 masking effect Effects 0.000 claims description 9
- 239000000470 constituent Substances 0.000 claims 1
- 238000010586 diagram Methods 0.000 description 4
- 238000000034 method Methods 0.000 description 4
- 238000010276 construction Methods 0.000 description 1
- 230000000694 effects Effects 0.000 description 1
- 238000005516 engineering process Methods 0.000 description 1
- 239000000284 extract Substances 0.000 description 1
- 238000000926 separation method Methods 0.000 description 1
Landscapes
- Document Processing Apparatus (AREA)
Description
【発明の詳細な説明】 産業上の利用分野 本発明は中国語等の表音文字列を漢字列に変換する漢字
変換装置に関する。BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a kanji conversion device for converting a phonetic character string such as Chinese into a kanji character string.
従来の技術 中国語は原則として、一つの漢字が一つの音節に対応し
ている。音節は次に示すように聲母,韻母,聲調の順で
構成されている。Conventional technology In Chinese, in principle, one kanji corresponds to one syllable. The syllable is composed of a mother, a mother and a tone as shown below.
聲母+韻母+聲調 第1表は中国語の音韻要素をそれぞれ中国大陸(へい音
1)、台湾(へい音2、注音)で使われている表音方式
で表わしたものである(以下へい音1の表音方式で説明
することとする)。表中欄(1)〜(21)を聲母、欄
(22)〜(59)を韻母と呼ぶ。その内、特に欄(22)〜
(24)の韻母を介音、欄(25)〜(37)の韻母を主韻
母、欄(38)〜(59)の韻母を結合韻母と呼ぶことにす
る。結合韻母は介音と主韻母の組合わせによって構成さ
れる。つまり、結合韻母の表記記号として、必ずしも介
音の表記記号i,u,yuと主韻母の表記記号a,o,e,ai,……,
engとの組み合わせから成るとは限らないが、それらの
構成する音韻要素を分析すると、実は介音の音韻要素と
主韻母の音韻要素を含 んでいる。例えば、「ui」は「u」と「ei」の音韻要素
を含んでいる。第2表には各結合韻母に含まれる介音と
主韻母の音韻要素を示す。したがって、中国語の音節は 聲母+介音+主韻母+聲調 の順で構成されるとも言える。例えば。「時」「依」
「愛」「取」「打」「忘」「中」などの漢字の読みに含
まれる音韻要素は次の通りである。Table 1 shows the phonological elements of Chinese in the phonetic system used in the Chinese mainland (Hai-on 1) and Taiwan (Hai-on 2, note), respectively. The phonetic method of No. 1 will be used for explanation). Columns (1) to (21) in the table are called a mother, and columns (22) to (59) are called a mother. Among them, especially column (22) ~
The melody of (24) will be called as a phoneme, the melody of columns (25) to (37) will be called a main melody, and the melody of columns (38) to (59) will be called a combined melody. The combined syllable is composed of a combination of the consonant and the main syllable. In other words, the notation symbols i, u, yu of the consonants and the notation symbols a, o, e, ai, ...
Although it does not necessarily consist of a combination with eng, the analysis of the phonological elements that compose them actually reveals that the phonological elements of the consonant and the main vowel are included. I'm out. For example, “ui” includes phoneme elements “u” and “ei”. Table 2 shows the phoneme elements of the consonant and main vowel included in each combined vowel. Therefore, it can be said that the Chinese syllables are composed in the order of mother tongue + consonant + main rhyme + tone. For example. "Time""Yi"
The phonological elements included in the reading of kanji such as "love,""take,""hatsu,""forget," and "medium" are as follows.
この例からも分かるように、各音節に聲母,介音,主韻
母が全て含まれるとは限らない。例えば、「時」の読み
「shi′」には聲母の音韻要素しか含まれていない。
又、聲調は各音節の構成に不可欠の要素である。それに
よって、音節の調子の高さが分かるわけである。 As can be seen from this example, each syllable does not necessarily include all of the mother, the consonant, and the main syllable. For example, the reading "shi '" of "time" includes only the phoneme element of the mother of voice.
Also, the tone is an essential element in the construction of each syllable. By this, the height of the syllable tone can be known.
第1表に点線で区切られている音韻要素はそれらの発音
が最も類似し合ったもので、よく間違えられる。例え
ば、「ch」と「c」はいずれも破擦音で、発音する時の
舌の位置だけによって区別されるので、中国語を習う人
にとって非常に難しいことである。特に、表音文字を入
力手段とする漢字変換装置では、使用者が入力したい漢
字列に対応する音節を指定された表音文字列で表わして
入力することによって漢字列に変換されるので、上述類
似した音韻要素を正しく区別して入力しないと、誤変換
又は変換不能に陥り、漢字変換装置としては、致命的な
欠点となる。The phoneme elements separated by dotted lines in Table 1 are those whose pronunciations are most similar to each other and are often mistaken. For example, both “ch” and “c” are affricate sounds, which are distinguished only by the position of the tongue when pronounced, which is very difficult for a person who learns Chinese. In particular, in a kanji conversion device that uses phonetic characters as input means, the user can convert the kanji string corresponding to the kanji string that the user wants to input into the kanji string by inputting the syllable corresponding to the specified phonetic character string. If similar phonological elements are not correctly distinguished and input, erroneous conversion or non-conversion becomes possible, which is a fatal drawback for a kanji conversion device.
従来のへい音漢字変換装置としては、例えば特開昭59−
121425号公報に示されている。第2図はこの従来のへい
音漢字変換装置のブロック図を示すものである。21は入
力されたデータをローマ字データと聲調データに分離す
る分離手段である。23は下に示す要領で各単語について
ローマ字列,漢字列,聲調及び使用頻度の各項目を記憶
している辞書である。An example of a conventional kanji-to-kanji conversion device is Japanese Patent Laid-Open No. 59-
No. 121425 is disclosed. FIG. 2 shows a block diagram of this conventional syllable-kanji conversion device. Reference numeral 21 is a separating means for separating the input data into Roman character data and tone data. Reference numeral 23 is a dictionary that stores the Roman character string, the Chinese character string, the tone and the frequency of use for each word as shown below.
22は上記分離手段21より与えられるローマ字列データに
該当する全ての同音異義語を上記辞書23より取り出す参
照手段である。24は参照手段22より得られた漢字列と分
離手段21の聲調データを比較し所定の漢字列を出力する
と共に上記聲調データのない場合は該当する漢字列の使
用頻度を利用して頻度の高い順に出力し所望の漢字列を
選択可能とする比較手段である。 Reference numeral 22 is a reference means for extracting from the dictionary 23 all homonyms corresponding to the Roman character string data given by the separating means 21. Reference numeral 24 compares the kanji string obtained from the reference means 22 with the tone data of the separating means 21 and outputs a predetermined kanji string, and if there is no tone data, the frequency of use of the corresponding kanji string is used frequently. It is a comparison means that outputs in sequence and can select a desired kanji string.
以上のように構成された従来のへい音漢字変換装置にお
いては、例えば、「中国」を入力したい場合、先ずキー
ボードからその読みである「zhong1 guo2」を入力す
る。すると、分離手段21で(zhongguo)のローマ字列デ
ータと(1,2)の聲調データに分離される。参照手段24
で(zhongguo)を検索のキーとして、辞書23から単語を
逐次に検索する。辞書23に(zhongguo)で登録される単
語は「中国」と とがあるが、聲調データが(1,2)となるのが「中国」
であるので、比較手段24で「中国」を出力と判断する。In the conventional syllabary-kanji conversion device configured as described above, for example, when "Chinese" is desired to be input, the reading "zhong1 guo2" is first input from the keyboard. Then, the separation means 21 separates the Roman character string data of (zhongguo) and the tone data of (1,2). Reference means 24
With (zhongguo) as a search key, words are sequentially searched from the dictionary 23. The word registered in (zhongguo) in the dictionary 23 is "China" However, it is "China" that the voice data is (1,2).
Therefore, the comparison means 24 determines that “China” is output.
発明が解決しようとする問題点 しかし、上記のような構成には次の問題点がある。Problems to be Solved by the Invention However, the above configuration has the following problems.
(1)使用者が第1表に示すような類似した音韻要素を
区別することができない場合、例えば、「学生」を入力
したい場合、その正しい読みが と見当が付かない場合、間違った読みを入力すると、
「学生」と正しく変換することができない。このような
場合、試行錯誤をするように、全ての可能な組な合わせ
を一つ一つ試すより仕方がない。(1) When the user cannot distinguish similar phoneme elements as shown in Table 1, for example, when he wants to input "student", the correct reading is If you don't have a clue, enter the wrong reading,
Can't be converted correctly as "student". In such cases, there is no choice but to try all the possible combinations one by one, as if by trial and error.
(2)中国語の漢字の読みの種類は約1260があり、それ
を符号化すれば、せいぜい2bytes(byteを単位とする場
合)で済むが、辞書に各単語の読みを対応するローマ字
のままで登録すると、一つの漢字当たり2〜6bytesを要
し、無駄なメモリ空間を占めると共に、ローマ字列を辞
書検索時の比較対象とするので、必要の倍以上の時間が
かかり、また各単語に対応するローマ字列が固定長でな
いため、辞書構造に規則性がなく、検索が容易でない。
又、上記類似した音韻要素を同一視しようとすれば、ま
ず各ローマ字列を対応する音韻要素単位で分離し、対応
表の参照によって類似したものであるかどうかを判断す
るような複雑な処理を行なわなければならないので、非
効率的で実用的ではない。(2) There are about 1260 types of Chinese kanji readings, and if you code them, you can use at most 2bytes (when using bytes as a unit), but the dictionary will read each word as the corresponding Roman character. If you register with, it will take 2 to 6 bytes for each kanji, occupy a wasted memory space, and since the Roman character strings will be the comparison target when searching the dictionary, it will take more than double the time required and it will correspond to each word Since the Roman character string is not fixed length, there is no regularity in the dictionary structure and it is not easy to search.
Further, if it is attempted to identify the above-mentioned similar phoneme elements, first, each Roman character string is separated into corresponding phoneme element units, and complicated processing such as determining whether they are similar by referring to the correspondence table is performed. It must be done, so it is inefficient and impractical.
本発明はかかる点に鑑み、コンパクトな辞書を可能とす
ると共に、あいまいな入力に対しても高い確率で正しく
漢字列に変換できる漢字変換装置を提供することを目的
とする。In view of the above point, the present invention has an object to provide a kanji conversion device that enables a compact dictionary and can correctly convert an ambiguous input into a kanji string with a high probability.
問題点を解決するための手段 本発明は聲母,韻母の音韻要素に対し、それぞれ類似し
たものをグループに分け、各グループの音韻要素間に距
離が1であるようなビットパターンを割り当て、上記聲
母,韻母のビットパターン及び聲調を表わすビットパタ
ーンとの組み合わせにより一つの漢字の音節を示す音節
符号を用いて表わされた中国語の単語の読みと該当する
漢字コードとの組を格納した辞書と、あいまいな発音表
記に対して該当する音節符号の不明確なビット位置をマ
スクして上記辞書の検索を行う辞書検索手段とを備えた
漢字変換装置である。Means for Solving the Problems The present invention divides the phoneme elements of a mother and a phoneme element, which are similar to each other, into groups, and assigns a bit pattern having a distance of 1 between the phoneme elements of each group. , A dictionary that stores a set of readings of Chinese words represented by using a syllable code that indicates a syllable of one Chinese character by combining with a bit pattern of the vowel and a bit pattern that represents a tone and the corresponding Chinese character code A kanji conversion device provided with dictionary search means for searching the dictionary by masking unclear bit positions of the corresponding syllable code for ambiguous phonetic notations.
作用 本発明は前記した構成により、辞書がコンパク化され、
且つ規則的な構造をもつことにより、あいまいな入力に
対しても、辞書の読みの部分の特定の情報をマスクし、
類似した音韻要素を同一視することにより、所要の漢字
列に変換することができる。Action The present invention makes the dictionary compact by the above configuration,
And by having a regular structure, even for ambiguous input, mask certain information in the reading part of the dictionary,
By equating similar phoneme elements, it is possible to convert into a required kanji string.
実施例 第1図は本発明の実施例における漢字変換装置のブロッ
ク図を示すものである。第3表は本発明の実施例におけ
る漢字変換装置の内部処理に使われる音節符号におい
て、それぞれ聲母,介音,主韻母,結合韻母,聲調に割
り当てられるビットパターンであり、これらのビットパ
ターンは下記の構成で2バイトて定義されており、同時
に上記類似した音韻要素同士のビットパターン間の距離
は1(相違ビットは最下位のビット)となっている。Embodiment FIG. 1 is a block diagram of a kanji conversion device according to an embodiment of the present invention. Table 3 is a bit pattern assigned to a vowel, a consonant, a main vowel, a combined vowel, and a voice in the syllable code used for the internal processing of the kanji conversion device in the embodiment of the present invention. These bit patterns are as follows. Is defined by 2 bytes, and at the same time, the distance between the bit patterns of the phoneme elements similar to each other is 1 (the difference bit is the least significant bit).
第1,2byteのbitoは0で、結合韻母のビットパターンは
介音と主韻母とのビットパターンの組合せで表わされ
る。本実施例の漢字変換装置の内部処理に使われる音節
符号はASCII CODEのgraphic characterに対応し、つま
り、本実施例の音節符号によると、任意の中国語の音節
は二つのASCII CODEのgraphic characterで表わすこと
ができる。又、第1byteのbit5,bit7,第2byteのbit7をマ
スクするとそれぞれ類似した聲母,介音,主韻母を同一
視することができる。 The bito of the 1st and 2nd bytes is 0, and the bit pattern of the combined vowel is represented by the combination of the bit patterns of the consonant and the main vowel. The syllable code used in the internal processing of the Kanji conversion device of this embodiment corresponds to the ASCII CODE graphic character, that is, according to the syllable code of this embodiment, an arbitrary Chinese syllable has two ASCII CODE graphic characters. Can be expressed as Further, by masking bit5, bit7 of the first byte and bit7 of the second byte, it is possible to identify the same voice mother, consonant, and main rhyme.
第1図において、10は少なくとも表音文字、及び辞書検
索モードを指定する辞書検索モードキーを有する入力手
段、11は上記入力手段から送られてきた表音文字列を上
記音節符号に変換する音節変換手段、14は上記音節符号
を用いて表わされた中国語の単語の読みと上記単語に対
応する漢字コードとの組を格納した辞書、13は上記入力
手段から送られてきた辞書検索モードの指定によって、
上記辞書を検索する時、辞書に登録される単語の読みの
各音節の第1byteのbit5,bit7、第2byteのbit7をマスク
し、それぞれ対応する入力された表音文字列と類似した
聲母,介音,主韻母を同一視したり、第2byteのbit1〜b
it3をマスクし、聲調を無視したり、或いは各音韻要素
の対応するビットパターンをマスクし、その音韻要素を
無視したりして、該当する全ての単語候補を辞書から取
り出す辞書検索手段である。12は上記音節符号変換手段
11から送られてきた音節符号を変換単位毎に辞書検索手
段13に送ると共に、辞書検索手段13から送られてきた単
語候補を使用者の選択によって、対応する単語候補を出
力手段15に送る漢字変換手段である。In FIG. 1, 10 is at least a phonetic character and an input means having a dictionary search mode key for designating a dictionary search mode, and 11 is a syllable for converting a phonetic character string sent from the input means into the syllable code. A conversion means, 14 is a dictionary storing a set of readings of Chinese words represented by using the syllable code and Kanji codes corresponding to the words, and 13 is a dictionary search mode sent from the input means. By the specification of
When searching the above dictionary, mask the 1st byte bit5, bit7, the 2nd byte bit7 of each syllable of the word reading registered in the dictionary, and use the same mother and interface as the corresponding input phonetic character string. Identifies sounds and main syllables, bit 1 to b of the 2nd byte
It is a dictionary search means for extracting all relevant word candidates from the dictionary by masking it3 and ignoring the tone, or masking the corresponding bit pattern of each phoneme element and ignoring the phoneme element. 12 is the syllable code conversion means
The syllable code sent from 11 is sent to the dictionary search means 13 for each conversion unit, and the word candidates sent from the dictionary search means 13 are selected by the user and the corresponding word candidates are sent to the output means 15. It is a conversion means.
以上のように構成された本実施例の漢字変換装置につい
て、以下その動作を説明する。The operation of the Kanji conversion apparatus of this embodiment configured as described above will be described below.
入力手段10から入力された表音文字が先ず、音節符号変
換手段11で第3表に従って、音節毎に音節符号に変換さ
れる。例えば、 が入力されると、 のような音節符号が得られる。ASCII CODEのgraphic ch
aracterで表わすと、「v″h+」となる。音節符号変
換手段11では、入力された表音 文字列の各音節に対して、それに含まれる音韻要素と聲
調によって、第3表に示す対応するビットパターンを割
り当てるだけで良いので、変換は非常に簡単である。入
力された表音文字列が音節符号に変換された後、次に辞
書検索手段13で、入力手段10から指定された辞書検索モ
ードによって、該当する全ての単語候補を辞書14から取
り出す。上記の例では、辞書検索のキーは「v″h+」
の4文字だけで、従来の漢字漢字変換装置での の10文字に比べて、検索に必要な比較文字数が半分以下
となり、検索速度が従来より速い。それに、辞書14に上
記音節符号を用いるので、各単語の読みを表わすのに必
要なメモリ量はその単語を構成する文字数と正比例し、
次に示すように、単語を構成する文字数によって、分類
すると、従来の辞書に比べて、小メモリ量で、規則正し
い構造を持つことができる。First, the phonetic character input from the input means 10 is converted by the syllable code conversion means 11 into syllable codes for each syllable according to the third table. For example, Is entered, A syllable code such as ASCII CODE graphic ch
When expressed by aracter, it becomes “v ″ h +”. In the syllable code conversion means 11, the input phonetic For each syllable of a character string, it is only necessary to assign the corresponding bit pattern shown in Table 3 according to the phonological element and the tone contained in it, so the conversion is very simple. After the input phonetic character string is converted into a syllable code, the dictionary search means 13 then extracts all applicable word candidates from the dictionary 14 in the dictionary search mode specified by the input means 10. In the above example, the dictionary search key is “v ″ h +”
Only 4 characters of the conventional kanji-kanji conversion device Compared with 10 characters, the number of comparison characters required for search is less than half, and the search speed is faster than before. Besides, since the syllable code is used in the dictionary 14, the amount of memory required to represent the reading of each word is directly proportional to the number of characters that make up that word,
As shown below, when classified according to the number of characters that make up a word, it is possible to have a regular structure with a smaller amount of memory than a conventional dictionary.
以上の例で、例えば、使用者が「学生」を入力したいと
き、その読みが 見当がつかない場合、 のような間違った読みを入力すると、「学生」と正しく
変換できない。ところが、本実施例の音節符号による
と、 の音節は次に示すように表わされる。 In the above example, when the user wants to input "student", the reading is If you have no idea, If you enter the wrong reading like "," it cannot be converted correctly as "student". However, according to the syllable code of this embodiment, The syllable of is represented as follows.
したがって、類似した音韻要素を同一視するために入力
手段10に設けられるキーを押すだけで、辞書検索時、各
音節の第1byteのbit5、bit7、第2byteのbit7がマスクさ
れ、「学生」が検出される。更に、聲調を無視するため
に設けられるキーを押すと、 が検出される。この時、漢字変換手段12で、使用者の選
択によって、所要の単語に変換する。勿論、あるルーチ
ンによって、辞書14に登録した音節符号を指定された表
音方式に変換して、使用者に読みを知らせるのも簡単に
できる。音節符号に聲母,介音,主韻母,結合韻母,聲
調に対応するビットの位置は固定しているので、対応表
の参照だけで容易に変換できるからである。 Therefore, by simply pressing the key provided on the input means 10 to identify similar phonological elements, the first byte bit5, bit7, and second byte bit7 of each syllable are masked during the dictionary search, and the "student" To be detected. Furthermore, if you press the key provided to ignore the tone, Is detected. At this time, the kanji conversion means 12 converts the word into a desired word according to the user's selection. Of course, it is possible to easily notify the user of the reading by converting the syllable code registered in the dictionary 14 into the designated phonetic system by a certain routine. This is because the positions of the bits corresponding to the syllable code, the syllable, the consonant, the main syllable, the combined syllable, and the syllable are fixed, so that the conversion can be easily performed only by referring to the correspondence table.
なお、本発明は上記実施例にのみ限らず、要旨を変更し
ない範囲で適宜変形して、実施できる。例えば、入力手
段10はキーボードによる表音文字列の入力だけでなく、
音声信号の入力を音声認識によって対応する表音文字列
を生成する入力手段に変えても良い。音節符号変換手段
11で使う音韻要素−ビットパターン対応表における音韻
要素の表音方式としては、第3表に示すように単に上記
へい音1の表音方式だけでなく、同時に、音韻要素の各
種の表音方式に対応する表音文字列を用意することによ
って、入力の表音方式の切り換えだけで、上記実施例は
辞書の拡張などの変更をする必要がなく、同時に多種の
表音方式による入力文字列に対応することができる。い
ずれも、上記音節符号を内部処理に使うので、簡単な音
節符号変換手段と、辞書の単語の読みを音節符号で表わ
すことによって、容易に多種の表音方式による入力文字
列に対応できる。The present invention is not limited to the above-described embodiments, but can be modified and implemented as appropriate without departing from the scope of the invention. For example, the input means 10 is not limited to the input of phonetic character strings by the keyboard,
The input of the voice signal may be changed to an input means for generating a corresponding phonetic character string by voice recognition. Syllable code conversion means
As shown in Table 3, the phonetic system of the phoneme elements in the phonological element-bit pattern correspondence table used in 11 is not only the phonetic system of the above-mentioned vowel sound 1 but also various phonetic systems of the phoneme elements at the same time. By preparing a phonetic character string corresponding to, it is not necessary to change the expansion of the dictionary or the like in the above embodiment by simply switching the input phonetic method, and at the same time, input character strings of various phonetic methods can be used. Can respond. In both cases, since the syllable code is used for internal processing, a simple syllable code converting means and the reading of words in the dictionary are represented by syllable codes, so that input character strings of various phonetic systems can be easily dealt with.
なお、辞書検索手段13の辞書検索モードは、上記特定の
ビットをマスクして、類似した音韻要素を同一視するこ
とによって、該当する全ての単語を単語辞書から取り出
すような検索モードだけでなく、特別の指定によって、
辞書検索時、入力された読みのある音節の特定の部分を
マスクして検索することもできる。例えば、「*」を入
力列にある音節の不明な部分を表わす記号とする。する
と、 が入力された時、一番目の音節に対応する音節符号の介
音と主韻母に対応するビットがマスクされ、その条件を
満たした全ての単語候補が検索される。音節符号の各音
韻要素に対応するビットが固定であるので、このような
検索は簡単である。The dictionary search mode of the dictionary search means 13 is not limited to the search mode in which all the corresponding words are extracted from the word dictionary by masking the specific bits and identifying similar phoneme elements. By special designation,
At the time of dictionary search, it is also possible to search by masking a specific part of the input syllable with reading. For example, let “*” be a symbol representing an unknown part of the syllable in the input string. Then, Is input, the bits of the syllable code corresponding to the first syllable and the bits corresponding to the main syllables are masked, and all word candidates satisfying the condition are searched. Such a search is easy because the bits corresponding to each phoneme element of the syllable code are fixed.
なお、漢字変換手段12は単語単位の変換だけでなく、自
由文を変換単位とする漢字変換装置においても、本発明
に示す音節符号を用いた辞書の読みの部分の特定の位置
をマスクすることによって、あいまいな発音表記に対し
て、該当する音節符号の不明確なビット位置に対しての
照合を行なわない様マスクすることによって、変換すべ
き単語候補を選び出すことも応用できる。Note that the kanji conversion means 12 can mask not only a word-based conversion, but also a kanji conversion device that uses a free sentence as a conversion unit to mask a specific position of a reading part of a dictionary using a syllable code according to the present invention. It is also applicable to select a word candidate to be converted by masking an ambiguous phonetic notation so as not to perform matching on an uncertain bit position of a corresponding syllable code.
なお、音節符号の適用範囲は漢字変換装置だけでなく、
中国語の読みに対する任意の処理、メモリでの蓄積、コ
ンピュータ間の転送にも利用できる。The range of application of syllable codes is not limited to Kanji conversion devices,
It can also be used for arbitrary processing of Chinese reading, storage in memory, and transfer between computers.
発明の効果 以上説明したように、本発明によれば、音節符号を利用
することによって、辞書のメモリ量が従来の80%以下に
減り、同時に多種の表音方式による入力に対応できると
共に、特定の位置のビットをマスクすることによって、
あいまいな入力に対しても所要の単語候補を選び出し
て、変換することができる。それに、各単語の読みとそ
の読みに対応する漢字コードを登録するに必要なメモリ
量がその単語を構成する文字数に正比例するので、辞書
のROM化も簡単にでき、経済性の面も、辞書検索速度と
漢字変換速度の向上も図られ、その実用的効果は大き
い。As described above, according to the present invention, by using the syllable code, the memory capacity of the dictionary is reduced to 80% or less of the conventional one, and at the same time, it is possible to support input by various phonetic methods and By masking the bits in the position of
Even for ambiguous inputs, you can select the necessary word candidates and convert them. In addition, the amount of memory required to register the reading of each word and the kanji code corresponding to that reading is directly proportional to the number of characters that make up the word, so the dictionary can be easily ROMized and the dictionary is economical. The search speed and kanji conversion speed have also been improved, and their practical effects are great.
第1図は本発明における一実施例の漢字変換装置のブロ
ック図、第2図は従来の漢字変換装置のブロック図であ
る。 10……入力手段、11……音節符号変換手段、12……漢字
変換手段、13……辞書検索手段、14……辞書、15……出
力手段。FIG. 1 is a block diagram of a Kanji conversion apparatus according to an embodiment of the present invention, and FIG. 2 is a block diagram of a conventional Kanji conversion apparatus. 10 ... input means, 11 ... syllable code conversion means, 12 ... kanji conversion means, 13 ... dictionary search means, 14 ... dictionary, 15 ... output means.
Claims (1)
それぞれ類似するもの同士に分類して作られた複数グル
ープに対し、各グループ内の構成要素間に距離が1であ
るようなビットパターンを割り当て、上記聲母,韻母の
ビットパターン及び聲調を表わすビットパターンの組み
合わせにより一つの漢字の音節を示すように作成された
音節符号を用いて表わされた中国語の単語の読みと上記
単語に対応する漢字コードとの組を格納する辞書と、あ
いまいな発音表記に対して該当する音節符号の不明確な
ビット位置をマスクして上記辞書の検索を行う辞書検索
手段とを備えたことを特徴とする漢字変換装置。1. For a plurality of groups formed by classifying Chinese phonological elements, that is, a mother and a mother, which are similar to each other, a bit having a distance of 1 between constituent elements in each group. Allocating a pattern, reading the Chinese word represented by using a syllable code created so as to show one syllable of a Chinese character by a combination of the bit pattern of the above-mentioned mother, the mother and the bit pattern representing the tone and the above word And a dictionary storing means for storing a set of kanji codes corresponding to, and a dictionary search means for masking an unclear bit position of the corresponding syllable code for an ambiguous phonetic notation. Characteristic Kanji conversion device.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP61287031A JPH0760433B2 (en) | 1986-12-02 | 1986-12-02 | Kanji converter |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP61287031A JPH0760433B2 (en) | 1986-12-02 | 1986-12-02 | Kanji converter |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPS63140366A JPS63140366A (en) | 1988-06-11 |
| JPH0760433B2 true JPH0760433B2 (en) | 1995-06-28 |
Family
ID=17712148
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP61287031A Expired - Fee Related JPH0760433B2 (en) | 1986-12-02 | 1986-12-02 | Kanji converter |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH0760433B2 (en) |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS56152033A (en) * | 1980-04-23 | 1981-11-25 | Matsushita Electric Ind Co Ltd | Language input device |
| FR2482747B1 (en) * | 1980-05-19 | 1986-10-31 | Barouch Eleazar | IDEOGRAPHIC CHARACTER ENCODING DEVICE |
-
1986
- 1986-12-02 JP JP61287031A patent/JPH0760433B2/en not_active Expired - Fee Related
Also Published As
| Publication number | Publication date |
|---|---|
| JPS63140366A (en) | 1988-06-11 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US5995934A (en) | Method for recognizing alpha-numeric strings in a Chinese speech recognition system | |
| JPH10269204A (en) | Automatic Chinese document proofing method and device | |
| CN114822491A (en) | Text processing, model training and speech synthesis method, device, system and medium | |
| KR101777141B1 (en) | Apparatus and method for inputting chinese and foreign languages based on hun min jeong eum using korean input keyboard | |
| JP2002278579A (en) | Voice data search device | |
| JP2019095603A (en) | Information generation program, word extraction program, information processing device, information generation method and word extraction method | |
| JPS58123129A (en) | Converting device of japanese syllabary to chinese character | |
| JP3983313B2 (en) | Speech synthesis apparatus and speech synthesis method | |
| KR100564742B1 (en) | Text-to-speech device and method | |
| JPS63140366A (en) | Kanji converting device | |
| CN1323004A (en) | Automatic conversion method from Chinese braille to Chinese character | |
| JPS62251986A (en) | Misread character correction processor | |
| JPS61184683A (en) | Recognition-result selecting system | |
| JP2002189490A (en) | Method of pinyin speech input | |
| JPS58123126A (en) | Dictionary retrieving device | |
| JPH0227423A (en) | Method for rearranging japanese character data | |
| Phaiboon et al. | Isarn Dharma Alphabets lexicon for natural language processing | |
| JP3888701B2 (en) | Character converter | |
| CN1048341C (en) | Fuzzy character transtormer | |
| JP2997151B2 (en) | Kanji conversion device | |
| JPS6275757A (en) | Chinese input unit | |
| JPS61177575A (en) | Japanese sentence creation device | |
| CN118333010A (en) | Novel spelling Chinese character digital code | |
| JP2976682B2 (en) | Language playback device | |
| JPS61177574A (en) | Forming device of japanese document |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| LAPS | Cancellation because of no payment of annual fees |