JPS588379A - Kana (japanese syllabary)-kanji (chinese character) converting system - Google Patents

Kana (japanese syllabary)-kanji (chinese character) converting system

Info

Publication number
JPS588379A
JPS588379A JP56106806A JP10680681A JPS588379A JP S588379 A JPS588379 A JP S588379A JP 56106806 A JP56106806 A JP 56106806A JP 10680681 A JP10680681 A JP 10680681A JP S588379 A JPS588379 A JP S588379A
Authority
JP
Japan
Prior art keywords
character
kana
conversion
word
connection
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP56106806A
Other languages
Japanese (ja)
Inventor
Takahiko Chuma
中馬 高彦
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Panasonic Holdings Corp
Original Assignee
Matsushita Electric Industrial Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Matsushita Electric Industrial Co Ltd filed Critical Matsushita Electric Industrial Co Ltd
Priority to JP56106806A priority Critical patent/JPS588379A/en
Publication of JPS588379A publication Critical patent/JPS588379A/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/40Processing or translation of natural language
    • G06F40/53Processing of non-Latin text

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Computational Linguistics (AREA)
  • General Health & Medical Sciences (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Document Processing Apparatus (AREA)

Abstract

PURPOSE:To quickly obtain the result of desired connection and to perform the converting processing at high speed, by designating a conversion candidate even for a word short in the length of coincidence with an inputted Kana train, when the margin of connection with a character or a word connected in front is high. CONSTITUTION:A front connection character is read out to a character confirming section 3 from an input character buffer 1 with a conversion control section 2 of a Kana character conversion system, and the type of character of front-connected character is designated. In this case, when the type of character is other than Kana, the character type code and when Kane, the word of conversion condidate, corresponding to the part of end of the Kana are read out from conversion condidate buffer 6, the part of speech is recognized at a part of speech recognizing section 4 and the type code is inputted to a connection margin recognizing section 5. Words are sequentially given to the section 4 from word groups with different length of coincidence corresponding to the Kana train of the position of the front connected character from the control section 2 to recognize the part of the speech. The conversion candidate is stored in the buffer 6, and Kana next to the candidate is retrieved at a sentence retrieval section 7, and the result of connection desired is applied to the character output buffer 8 quickly.

Description

【発明の詳細な説明】 本発明は日本語文書作成装置等に有用な仮名漢字変換方
式に関するものであり、変換を高速に行なうことを目的
とする。
DETAILED DESCRIPTION OF THE INVENTION The present invention relates to a kana-kanji conversion method useful for Japanese document creation devices, etc., and aims to perform conversion at high speed.

従来文単位の仮名列を、漢字に変換する場合においては
、最長一致法に見られるように、被変換仮名列との一致
長さという点を最重要視して行なっていた。このため、
一致長さの短い単語が正しい場合には、まず、一致長さ
の長い単語が変換候補とされ、文法チェック等によって
排除されるという段階を経なければ正解が得られないと
いう欠点があった。
Conventionally, when converting a sentence-based kana string into a kanji character, the most important consideration was the length of the match with the converted kana string, as seen in the longest match method. For this reason,
If a word with a short matching length is correct, the word with a long matching length is first set as a conversion candidate, and the correct answer cannot be obtained unless the word is eliminated through a grammar check or the like.

第4図はその例を示すもので、Aは漢字に変換すべき入
力文字列、Bは学語検索結果、Cは変換例であり、正し
い漢字「回」が一致長さが短いため変換候補としての順
位が低くなる。
Figure 4 shows an example. A is the input character string to be converted to kanji, B is the academic language search result, and C is the conversion example. The correct kanji "kai" is a conversion candidate because the match length is short. As a result, the ranking will be lower.

本発明はこのような欠点全除去せんとするもので以下、
本発明の一実施例について図を用いて説明する。
The present invention aims to completely eliminate such defects, and the following will be explained below.
An embodiment of the present invention will be described with reference to the drawings.

第1図は、本発明の仮名漢字変換方式の一実施例を示す
ブロック図である。第1図において、1は入力文字バッ
ファ、2は変換制御部、3は文字種認定部、4は単語の
品詞認定部、6は文字種と品詞及び品詞同士の接続の尤
度認定部、6は変換候補バッファ、7は文節検索部、8
は出力文字バッファである。
FIG. 1 is a block diagram showing an embodiment of the kana-kanji conversion method of the present invention. In FIG. 1, 1 is an input character buffer, 2 is a conversion control unit, 3 is a character type recognition unit, 4 is a word part-of-speech recognition unit, 6 is a likelihood recognition unit for character types, parts of speech, and connections between parts of speech, and 6 is a conversion unit. Candidate buffer, 7 is clause search section, 8
is the output character buffer.

第2図は、本発明の説明のための入力文字列の例とこれ
に対応して検索された単語群の例、及び変換結果の例を
示す図であり、図において、Aは入力文字列、Bは「力
」の位置がらの単語検索結果、Cは変換結果である。
FIG. 2 is a diagram showing an example of an input character string, an example of a corresponding word group, and an example of a conversion result for explaining the present invention. In the figure, A is an input character string. , B is the word search result based on the position of "force", and C is the conversion result.

第3図は、接続尤度認定部6が尤度認定に用いる接続尤
度表の例であり、前側、後側の単語の品詞コード及び文
字種コードがそれぞれ行と列に配され、行と列が交差す
る位置には、該当する品詞コード及び文字種コード間の
接続尤度が設定されている。この例では尤度の高い箇所
には大きな数値を与えである。
FIG. 3 is an example of a connection likelihood table used by the connection likelihood recognition unit 6 for likelihood recognition, in which the part-of-speech code and character type code of the front and rear words are arranged in rows and columns, respectively. The likelihood of connection between the corresponding part-of-speech code and character type code is set at the intersection. In this example, large numbers are given to locations with high likelihood.

第2図、第3図の例を使って第1図の各部の動作を説明
する。
The operation of each part shown in FIG. 1 will be explained using the examples shown in FIGS. 2 and 3.

第2図に示すように、「力」の前接文字は数字の「8」
であり、「力」の位置がらの単語検索の結果として第2
図のB部の単語が得られたとする。
As shown in Figure 2, the prefix of "power" is the number "8".
, and as a result of a word search based on the position of "force", the second
Assume that the words in part B of the diagram are obtained.

変換制御部2は、「力」ア前接文字「8」を入力文字バ
ッファ1から文字種認定部3に読み出して文字種を認宥
させ、今の例のように文字種が仮名その仮名で終わる部
分に対応する変換候補の単語を変換候補バッファ6から
読み出して品詞認定部4に品詞を認定させてその品詞コ
ードを、接続尤度認定部6に渡す。一方、変換制御部2
は、「力」の位置からの仮名列に対応しうる一致長さの
異なるものを含む単語群から単語を順次に品詞認定部4
に与えて品詞を認定させ、品詞コードを接続尤度認定部
6に渡す。接続尤度認定部6はJ例えば第3図に示すよ
うな、前側品詞コードもしくは文字種と後側品詞コード
もしくは文字種との接続尤度表をもとに、変換制御部2
の制御を受けて順次、〜単語の品詞コードと「8」の・
文字種コードとの接続尤度を認定し、変換制御部2に渡
す。変換制御部2は、接続尤度に基いて、第2図のよう
に単語群から、「回」1下数作」2.「快作、」、「貝
」。
The conversion control unit 2 reads the character prefix “8” from the input character buffer 1 to the character type recognition unit 3 to recognize the character type, and converts the character type to kana, which ends with that kana, as in the current example. The corresponding conversion candidate word is read from the conversion candidate buffer 6, the part of speech is recognized by the part of speech recognition section 4, and the part of speech code is passed to the connection likelihood recognition section 6. On the other hand, conversion control section 2
The part-of-speech recognition unit 4 sequentially selects words from a group of words that include words with different matching lengths that can correspond to the kana sequence starting from the position of “power”.
The part-of-speech code is passed to the connection likelihood recognition unit 6. The connection likelihood recognition unit 6 uses the conversion control unit 2 based on a connection likelihood table between the front part-of-speech code or character type and the rear part-of-speech code or character type, as shown in FIG. 3, for example.
Under the control of
The likelihood of connection with the character type code is recognized and passed to the conversion control unit 2. Based on the connection likelihood, the conversion control unit 2 converts the word group from the word group as shown in FIG. ``Kaisaku,''``Kai.''

・・・・・−の順に1一つづつ、従って、まず「回」を
変換候補として変換候補バッファ6に格納し、次に「回
」が対応する仮名部の次の仮名「す」からの単語検索を
行なう。変換制御部2は、さらに、例えば単語検索の結
果、一致する単語が無いとか、前接単語又は前接文字と
接続する単順が無い場合に、前接単語を変換候補バッフ
ァ6がら削除し、新たに次順位単語を変換候補バッファ
6に格納するという動作を行なう。このような動作を、
変換候補群が・、該当する仮名郡全体に対応するまで、
繰り返し行なう。第2図(8)の下線部に対応して第2
図(qの下線部のような変換候補群が検索されている。
・・・・One by one in the order of −, therefore, first store “Kai” as a conversion candidate in the conversion candidate buffer 6, and then store it from the next kana “su” of the kana part to which “Kai” corresponds. Perform a word search. The conversion control unit 2 further deletes the prefix word from the conversion candidate buffer 6 if, for example, there is no matching word as a result of the word search or there is no single order connected to the prefix word or prefix character, An operation of newly storing the next ranking word in the conversion candidate buffer 6 is performed. This kind of behavior
Until the conversion candidate group corresponds to the entire corresponding kana county,
Do it repeatedly. The second line corresponds to the underlined part in Figure 2 (8).
A group of conversion candidates as shown in the underlined part of the figure (q) are being searched.

この時点で変換制御部2は、変換候補群を出力文字バッ
フ77に転送して動作を終了する。
At this point, the conversion control unit 2 transfers the conversion candidate group to the output character buffer 77 and ends the operation.

このように、接続尤度の高い単語を取り上げて変換が成
功すると、より接続尤度の低い単語につぃ7ての無駄な
処理を省くことができる。
In this way, if a word with a high connection likelihood is selected and the conversion is successful, wasteful processing for words with a lower connection likelihood can be omitted.

以上の説明から明らかなように、本発明によれば、入力
仮名列との一致長さが比較的短C単語でも、前接する文
字又は単語との接続尤度が高ければ先に変換候補とされ
て、所望の変換結果を速く得ることができ、従来性なわ
れていた、一致長さの長い方の単語から順次、変換候補
として扱っていく方法と比べて、変換時間お短縮化が図
れる。
As is clear from the above explanation, according to the present invention, even if the length of the match with the input kana string is relatively short, if the likelihood of connection with the preceding character or word is high, it is first selected as a conversion candidate. Therefore, the desired conversion result can be obtained quickly, and the conversion time can be shortened compared to the conventional method of sequentially treating words with the longest match length as conversion candidates.

なお、第3図に示した接続尤度は単に一例であるにすぎ
ず、品詞及び文字種の区分、並びに、尤度として与える
数値は、この限9ではなく、適宜設定されるものである
ことは言うまでもない。
Note that the connection likelihood shown in Figure 3 is merely an example, and the classification of parts of speech and character types, as well as the numerical value given as the likelihood, are not limited to 9 and may be set as appropriate. Needless to say.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は本発明の一実施例における仮名漢字変換方式を
示すブロック図、第2デ^、 (B) 、 (Qは、そ
の変換過程を説明する図、第3図は、説明に用いた接続
尤度の例を示す図、第4図(8)、 (B) 、 (Q
は従来の変換方式における変換例を示す図である。 1・・・・・・入力文字バッ/ファ、2・・・・・・変
換制御部、3・・・・・・文字認定部、4・・・・・・
品詞認、定部、6・・・・・・接続尤度認定部、6・・
・・・・変換候補バッファ、7・・−・・・文節検索部
、8・・・・・・出力文字バッファ。 代理人の氏名 弁理士 中 尾 敏 男 ほか1名11
図 (A)     タ゛イ8矛=チニtじこラーにΣ4J
!ど(c)  第8酬線だ 821」 (A)  ・タ゛イ8114サクラマツリか(8)  
 困    助数詞 区I月    名iq 囲司  名詞 0日     名詞 同   助詞 【口    接尾語 (C)  第8回4チ祭〆
Figure 1 is a block diagram showing a kana-kanji conversion method in an embodiment of the present invention, Figure 2 is a diagram explaining the conversion process, Figure 3 is a diagram used for explanation. A diagram showing an example of connection likelihood, Figure 4 (8), (B), (Q
1 is a diagram showing an example of conversion in a conventional conversion method. 1... Input character buffer/buffer, 2... Conversion control section, 3... Character recognition section, 4...
Part of speech recognition, fixed part, 6... Connection likelihood recognition part, 6...
... Conversion candidate buffer, 7 ... Clause search unit, 8 ... Output character buffer. Name of agent: Patent attorney Toshio Nakao and 1 other person11
Diagram (A) Type 8 swords = Σ4J to chini tjikora
! (c) It's the 8th line 821" (A) ・Tai 8114 Sakura Matsuri? (8)
Difficulty Particle Ward I Month Name iq Keishi Noun 0 Day Noun Same Particle [Mouth Suffix (C) 8th 4th Festival〆

Claims (1)

【特許請求の範囲】[Claims] ′ 文字種を認定する手段と、単語の品詞を認定する手
段と、文字種と単語の品詞及び単語の品詞同士の接続の
尤度を認定する手段と全具備し、仮名列の一部庄に対応
しつる単語群から、該単語群の中の個々の単語と、該部
分に仮名部分が前接する場合は該仮名部分が変換された
単語、また該部分に非仮名文字が前接する場合は該文字
との接続の尤度を、単語の品詞と非仮名文字の文字種に
基い、 て認定し、該尤度の高い順に単語を選択して変
換候補とすることを特徴とする仮名漢字変換方式。
′ It is fully equipped with a means to recognize the character type, a means to recognize the part of speech of a word, and a means to recognize the likelihood of connection between the character type, the part of speech of the word, and the parts of speech of the word, and corresponds to a part of the kana sequence. From a vine word group, each word in the word group, the word to which the kana part is converted if the part is preceded by a kana part, and the character if the part is preceded by a non-kana character. A kana-kanji conversion method characterized in that the likelihood of a connection is determined based on the part of speech of a word and the character type of a non-kana character, and words with the highest likelihood are selected as conversion candidates.
JP56106806A 1981-07-07 1981-07-07 Kana (japanese syllabary)-kanji (chinese character) converting system Pending JPS588379A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP56106806A JPS588379A (en) 1981-07-07 1981-07-07 Kana (japanese syllabary)-kanji (chinese character) converting system

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP56106806A JPS588379A (en) 1981-07-07 1981-07-07 Kana (japanese syllabary)-kanji (chinese character) converting system

Publications (1)

Publication Number Publication Date
JPS588379A true JPS588379A (en) 1983-01-18

Family

ID=14443093

Family Applications (1)

Application Number Title Priority Date Filing Date
JP56106806A Pending JPS588379A (en) 1981-07-07 1981-07-07 Kana (japanese syllabary)-kanji (chinese character) converting system

Country Status (1)

Country Link
JP (1) JPS588379A (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH03206551A (en) * 1990-08-31 1991-09-09 Hitachi Ltd Method and device for displaying conversion candidates for reading input character strings

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH03206551A (en) * 1990-08-31 1991-09-09 Hitachi Ltd Method and device for displaying conversion candidates for reading input character strings

Similar Documents

Publication Publication Date Title
JPS6140671A (en) Word division processing method
JPS588379A (en) Kana (japanese syllabary)-kanji (chinese character) converting system
JPS6229796B2 (en)
JPH0350669A (en) information processing equipment
JPS58123126A (en) Dictionary retrieving device
JPS62274366A (en) Dictionary retrieving device
JPS60207983A (en) Production system of dictionary for recognizing character
JPH01114976A (en) Dictionary structure for document processor
JPS61177575A (en) Japanese sentence creation device
JPH0695330B2 (en) Document creation device
JP3353769B2 (en) Character recognition device, character recognition method, and character recognition program recording medium
JP3344793B2 (en) Kana-Kanji conversion device
JPH04372047A (en) Kana/kanji converter
JPS62298869A (en) Sentence tail conversion method
JPS61156464A (en) Document preparing device
JPH0444301B2 (en)
JPH0916575A (en) Pronunciation dictionary device
JPS62296268A (en) Kana-kanji conversion device
JPH04167051A (en) Document editing device
JPS61156465A (en) Document preparing system
JPS592124A (en) "kana" (japanese syllabary) "kanji" (chinese character) converting system
JPS61128364A (en) dictionary search device
JPS6319068A (en) Kana/kanji converting device
JPH02155073A (en) Unknown word qualifying device
JPH06259413A (en) Japanese language input system