JPH0438018B2 - - Google Patents

Info

Publication number
JPH0438018B2
JPH0438018B2 JP60096324A JP9632485A JPH0438018B2 JP H0438018 B2 JPH0438018 B2 JP H0438018B2 JP 60096324 A JP60096324 A JP 60096324A JP 9632485 A JP9632485 A JP 9632485A JP H0438018 B2 JPH0438018 B2 JP H0438018B2
Authority
JP
Japan
Prior art keywords
information
kanji
idiom
stored
address
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Expired - Lifetime
Application number
JP60096324A
Other languages
Japanese (ja)
Other versions
JPS61255465A (en
Inventor
Tsutomu Kawada
Kimito Takeda
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Toshiba Corp
Original Assignee
Tokyo Shibaura Electric Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Tokyo Shibaura Electric Co Ltd filed Critical Tokyo Shibaura Electric Co Ltd
Priority to JP60096324A priority Critical patent/JPS61255465A/en
Publication of JPS61255465A publication Critical patent/JPS61255465A/en
Publication of JPH0438018B2 publication Critical patent/JPH0438018B2/ja
Granted legal-status Critical Current

Links

Landscapes

  • Document Processing Apparatus (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Description

【発明の詳細な説明】 〔発明の技術分野〕 本発明は辞書容量の削減を図つた言語処理装置
に関する。
DETAILED DESCRIPTION OF THE INVENTION [Technical Field of the Invention] The present invention relates to a language processing device that aims to reduce dictionary capacity.

〔発明の技術的背景とその問題点〕[Technical background of the invention and its problems]

近時、情報処理技術を利用した各種の言語処理
装置、例えば日本語ワードプロセツサや自動翻訳
装置等が広く普及している。
In recent years, various language processing devices using information processing technology, such as Japanese word processors and automatic translation devices, have become widespread.

この種の言語処理装置に組込まれる辞書は、一
般に語の表記とこれを検索する為のキー情報とを
相互に対応付けて格納し、このキー情報に従つて
上記語の表記を検索するものとなつている。例え
ば日本語ワードプロセツサでは、仮名漢字変換処
理の為の単語辞書が準備され、仮名見出し語を検
索キーとしてその漢字見出し語を検索するものと
なつている。
Dictionaries built into this type of language processing device generally store the notation of a word and key information for searching the word in correspondence with each other, and search for the notation of the word according to this key information. It's summery. For example, in a Japanese word processor, a word dictionary is prepared for kana-kanji conversion processing, and the kanji headword is searched using the kana headword as a search key.

ところで上記漢字情報はJISで定められるよう
にそれぞれ2バイトの情報で表現される。また単
語の多くは複数の漢字の組合せとして表現され
る。この為、前記日本語ワードプロセツサの単語
辞書を構成する場合、大容量の記憶装置を必要と
し、その小形化を図る上での大きな課題となつて
いる。
By the way, each of the above kanji information is expressed as 2-byte information as specified by JIS. Also, many words are expressed as combinations of multiple kanji. For this reason, when constructing a word dictionary for the Japanese word processor, a large capacity storage device is required, which poses a major problem in reducing its size.

〔発明の目的〕[Purpose of the invention]

本発明はこのような事情を考慮してなされたも
ので、その目的とするところは、辞書容量の削減
を図り、その小形化を容易に可能ならしめる言語
処理装置を提供することにある。
The present invention has been made in consideration of these circumstances, and its purpose is to provide a language processing device that reduces the dictionary capacity and can easily be miniaturized.

〔発明の概要〕[Summary of the invention]

本発明は、仮名文字列等のキー情報を入力する
ための入力部と、個々の漢字を表わす漢字情報と
複数の漢字により構成される熟語を表わす熟語情
報を見出し語としてそれぞれ所定のアドレスに格
納した変換辞書と、前記入力部から入力されたキ
ー情報を読みとして前記変換辞書を検索し、該仮
名文字列を漢字情報または熟語情報に変換する変
換部とを有する言語処理装置において、前記変換
辞書に格納された熟語情報は、該熟語情報を構成
する複数の漢字の漢字情報が格納されたアドレス
に関するアドレス情報と、該熟語情報の読みの区
切りを示す区切り情報とにより表現されているこ
とを特徴とする。
The present invention has an input section for inputting key information such as a kana character string, kanji information representing individual kanji, and idiom information representing an idiom composed of multiple kanji, each stored as a headword at a predetermined address. and a conversion unit that searches the conversion dictionary using the key information input from the input unit as a reading and converts the kana character string into kanji information or compound word information, the conversion dictionary The idiom information stored in the idiom information is expressed by address information regarding an address where kanji information of a plurality of kanji constituting the idiom information is stored, and delimiter information indicating the division of pronunciations of the idiom information. shall be.

〔発明の効果〕〔Effect of the invention〕

本発明によれば、変換辞書に格納された熟語情
報は、該熟語情報を構成する複数の漢字の漢字情
報が格納されたアドレスに関するアドレス情報
と、該熟語情報の読みの区切り情報とにより表現
されるので、熟語情報をこれを構成する複数の漢
字情報の組み合わせで単純に格納する場合に比較
して、情報量を少なくすることができる。
According to the present invention, the phrase information stored in the conversion dictionary is expressed by the address information regarding the address where the kanji information of the plurality of kanji constituting the phrase information is stored, and the reading delimiter information of the phrase information. Therefore, the amount of information can be reduced compared to the case where the compound word information is simply stored as a combination of the plurality of kanji information that makes up the compound word information.

具体的には、例えば漢字1文字の漢字情報は、
JISコードの場合で2バイトであるから、漢字2
文字で構成される熟語をJISコードで表現すると、
4バイトのデータ量を必要とするが、これを本発
明のように漢字2文字のJISコードが格納された
アドレスのアドレス情報と、読みの区切りを示す
区切り情報とで表現すると、例えば2バイト程度
以下のデータで表現できる。
Specifically, for example, kanji information for one kanji character is
In the case of JIS code, it is 2 bytes, so kanji 2
When expressions consisting of letters are expressed in JIS code,
It requires 4 bytes of data, but if this is expressed as the address information of the address where the JIS code of 2 kanji characters is stored and the delimiter information indicating the division of readings as in the present invention, for example, about 2 bytes. It can be expressed using the following data.

一般に、日本語文章で使用される単語の多く
は、複数の漢字の組み合わせで構成される熟語で
あるから、本発明のように変換辞書を構成するこ
とにより単語全体として見た場合、情報量を大幅
に少なくするでき、それだけ変換辞書の容量を削
減することが可能であるため、言語処理装置の小
形化に大きく貢献する。
In general, many of the words used in Japanese sentences are compound words made up of combinations of multiple kanji, so by configuring a conversion dictionary like the present invention, the amount of information can be reduced when looking at the word as a whole. Since it is possible to significantly reduce the size of the conversion dictionary and reduce the capacity of the conversion dictionary accordingly, it greatly contributes to downsizing of the language processing device.

〔発明の実施例〕[Embodiments of the invention]

以下、図面を参照して本発明の一実施例につき
説明する。
Hereinafter, one embodiment of the present invention will be described with reference to the drawings.

第1図は実施例装置の要部概略構成図であり、
1は検索キー情報の入力部、2は上記検索キー情
報に従つて変換辞書3を検索し、該検索キー情報
に該当した見出し語を求める変換部、4は変換部
2で検索された見出し語を出力する出力部であ
る。
FIG. 1 is a schematic diagram of the main parts of the embodiment device,
1 is an input section for search key information; 2 is a conversion section that searches the conversion dictionary 3 according to the search key information to find a headword corresponding to the search key information; 4 is a headword searched by the conversion section 2; This is an output section that outputs .

この言語処理装置が日本語ワードプロセツサと
して実現される場合には、前記入力部1に与えら
れる検索キー情報は、例えば仮名キーボードから
入力された仮名文字列であり、変換辞書3はその
仮名文字列を読み情報として漢字表記される単語
をそれぞれ格納した単語辞書として実現される。
そして変換部2は、この単語辞書(変換辞書3)
を用いて前記入力仮名文字列を仮名漢字変換処理
して出力することになる。
When this language processing device is implemented as a Japanese word processor, the search key information given to the input unit 1 is, for example, a kana character string input from a kana keyboard, and the conversion dictionary 3 It is realized as a word dictionary that stores words written in Kanji characters as reading information in columns.
Then, the conversion unit 2 converts this word dictionary (conversion dictionary 3)
is used to perform kana-kanji conversion processing on the input kana character string and output it.

第2図は、このような言語処理装置における変
換辞書の一構成例を示すもので、入力部1からの
キー情報すなわち熟語文字列を読みとし、この読
み(仮名文字列)に対応した漢字文字(文字列)
を漢字見出し語として格納して構成される。尚、
この変換辞書3には上記見出し語に対する品詞情
報や、さらに後述する熟語の見出しの区切り情報
などの属性情報も格納されている。
FIG. 2 shows an example of the configuration of a conversion dictionary in such a language processing device. The key information from the input unit 1, that is, the idiom character string is used as a reading, and the kanji character corresponding to this reading (kana character string) is (string)
is stored as a kanji headword. still,
This conversion dictionary 3 also stores attribute information such as part-of-speech information for the headwords and delimiter information for the headwords of compound words, which will be described later.

しかしてこの変換辞書3においては、漢字1文
字の見出し語を得る仮名文字列、例えば『あか』
なる仮名文字列に対する漢字見出し語は、その同
音語を含めて『赤,朱,垢,丹,緋…』等として
格納されている。また同様に『じ』なる仮名文字
列に対する漢字見出し語は、その同音語を含めて
『示,仕,字,地,自,寺…』等として格納され
ている。これらの各漢字見出し語は従来装置の場
合と同様に、そのJIS漢字コード等としてそれぞ
れ表現される。
However, in this conversion dictionary 3, a kana character string to obtain a one-character kanji entry word, for example, ``Aka'' is used.
The kanji headwords for the kana character strings, including their homophones, are stored as ``red, vermilion, dirt, tan, scarlet...'' and the like. Similarly, the kanji headword for the kana character string ``ji'', including its homophones, is stored as ``show, shi, ji, ji, ji, temple...'' and so on. Each of these kanji headwords is expressed as its JIS kanji code, etc., as in the case of the conventional device.

一方、変換辞書3において、漢字2文字のよう
に複数の漢字で構成される単語(すなわち熟語)
である見出し語は、これを構成する複数の漢字が
格納された辞書3の格納アドレスのアドレス情報
と、この見出し語についての読みの区切りを示す
区切り情報とにより表現されている。
On the other hand, in the conversion dictionary 3, words composed of multiple kanji such as two kanji characters (i.e. compound words)
The entry word is expressed by address information of the storage address of the dictionary 3 in which the plurality of kanji constituting the entry word are stored, and delimiter information indicating the division of readings for this entry word.

即ち、例えば『あかじ』なる読みに対しては、
その読みが『あか』『じ』の2つに分解できるこ
とを利用し、『赤字』の『赤』が読み『あか』の
1番目の見出し語『赤』として格納されており、
また上記『赤字』の『字』が読み『じ』の3番目
の見出し語『字』として格納されていることか
ら、その見出し語を『1,3』としてアドレス情
報として表現している。
That is, for example, for the reading ``Akaji'',
Taking advantage of the fact that the reading can be broken down into two parts, ``aka'' and ``ji'', ``red'' in ``red'' is stored as the first headword ``red'' in the reading ``aka''.
Furthermore, since the ``character'' of the above ``red'' is stored as the third headword ``character'' of the reading ``ji'', the headword is expressed as ``1, 3'' as address information.

尚、読み『あかじ』に対する属性情報として与
えられている『2』なる情報は、その読みが『先
頭から2文字目に区切りを有する』ことを示して
いる。また読みを示す文字列が分解できないよう
な場合には、上記区切りを示す情報は『0』とし
て与えられ、読みを示す仮名文字列が3つ以上に
分解できる場合には、区切りを示す情報は例えば
『2,4』等として与えられる。
Note that the information "2" given as attribute information for the reading "Akaji" indicates that the reading "has a break at the second character from the beginning." In addition, if the character string indicating the reading cannot be broken down, the information indicating the break is given as ``0'', and if the kana string indicating the reading can be broken down into three or more, the information indicating the break is given as ``0''. For example, it is given as "2, 4", etc.

例えば読み『あかしんごう』に対しては、区切
りの情報『2,4』によつて上記読みが『あか/
しん/ごう』に分解され、その『あか』に対して
『赤』が格納されたアドレス情報『1』、『しん』
に対して『信』が格納されたアドレス情報『5』、
『ごう』に対して『号』が格納されたアドレス情
報『4』により上記読み『あかしんごう』の見出
し語が『1,5,4(赤信号)』として与えられる
ことになる。
For example, for the reading ``Akashingo'', the above reading is changed to ``Aka/
Address information "1" and "Shin" are decomposed into "Shin/Go" and "Red" is stored for that "Aka".
Address information “5” where “trust” is stored for,
The address information ``4'' in which the ``go'' is stored for ``go'' gives the headword for the reading ``Akashingo'' as ``1, 5, 4 (red light)''.

尚、区切り情報を区切り文字数として、つまり
上述した例では『2,2』等として表現するよう
にしても良い。また該当語(漢字)が辞書3の他
の読みに対する見出し語として格納されていない
場合には、その見出し語をそのまま漢字情報とし
て格納するようにすれば良い。
Note that the delimiter information may be expressed as the number of delimiter characters, such as "2, 2" in the above example. Furthermore, if the corresponding word (kanji) is not stored as a headword for other readings in the dictionary 3, the headword may be stored as is as kanji information.

かくしてこのような変換辞書3を備えた本装置
によれば、読みに対する属性情報によつて区切り
が指定された場合、変換部2はその区切り情報に
従つて上記読みを分解し、分解された読みに対応
して格納された見出し語を前記アドレス情報に従
つてそれぞれ求めてその見出し語(漢字文字列)
を得ることになる。つまり、見出し語を構成する
漢字表記をアドレス情報に従つてそれぞれ求め
て、その見出し語を得ることが可能となる。
Thus, according to this device equipped with such a conversion dictionary 3, when a break is specified by the attribute information for the reading, the conversion unit 2 decomposes the reading according to the break information, and converts the decomposed reading into the decomposed reading. Find the headwords stored corresponding to each according to the address information and retrieve the headwords (kanji character strings).
You will get . In other words, it is possible to obtain the headword by finding the kanji notation that constitutes the headword according to the address information.

また前述した変換辞書3の構成によれば、複数
の漢字で表現される見出し語が、各漢字のアドレ
ス情報と、その見出し語(熟語)の読みの区切り
を示す区切り情報で表現されるので、漢字を直接
JISコードで表現した場合には1文字当り2バイ
トのデータ量が必要であつたところを、例えば1
文字当り1バイト以下のデータ量で表現すること
が可能となる。従つて変換辞書3に必要な容量を
大幅に少なくすることができ、例えば1チツプ半
導体ROMに辞書データの全てを収納することが
可能となる等の効果が奏せられる。
Furthermore, according to the configuration of the conversion dictionary 3 described above, a headword expressed by multiple kanji is expressed by the address information of each kanji and the delimiter information indicating the division of the pronunciation of the headword (idiom). Kanji directly
For example, instead of 2 bytes of data per character when expressed in JIS code,
It becomes possible to express the amount of data with less than 1 byte per character. Therefore, the capacity required for the conversion dictionary 3 can be significantly reduced, and for example, it is possible to store all the dictionary data in one chip semiconductor ROM.

尚、本発明はその要旨を逸脱しない範囲で種々
変形して実施することができる。
Note that the present invention can be implemented with various modifications without departing from the gist thereof.

【図面の簡単な説明】[Brief explanation of drawings]

第1図は本発明の一実施例装置の要部概略構成
図、第2図は実施例装置における変換辞書の構成
例を示す図である。 1…入力部、2…変換部、3…変換辞書、4…
出力部。
FIG. 1 is a schematic diagram of the main part of an apparatus according to an embodiment of the present invention, and FIG. 2 is a diagram showing an example of the configuration of a conversion dictionary in the apparatus according to the embodiment. 1... Input section, 2... Conversion section, 3... Conversion dictionary, 4...
Output section.

Claims (1)

【特許請求の範囲】 1 仮名文字列を入力するための入力部と、 個々の漢字を表わす漢字情報と複数の漢字によ
り構成される熟語を表わす熟語情報を見出し語と
してそれぞれ所定のアドレスに格納した変換辞書
と、 前記入力部から入力された仮名文字列に従つて
前記変換辞書を検索し、該仮名文字列を漢字情報
または熟語情報に変換する変換部とを有する言語
処理装置において、 前記変換辞書に格納された熟語情報は、該熟語
情報を構成する複数の漢字の漢字情報が格納され
たアドレスに関するアドレス情報と、該熟語情報
の読みの区切りを示す区切り情報とにより表現さ
れていることを特徴とする言語処理装置。
[Scope of Claims] 1. An input section for inputting a kana character string, kanji information representing individual kanji, and idiom information representing an idiom composed of a plurality of kanji, each stored as a headword at a predetermined address. A language processing device comprising: a conversion dictionary; and a conversion unit that searches the conversion dictionary according to a kana character string input from the input unit and converts the kana character string into kanji information or compound word information, The idiom information stored in the idiom information is expressed by address information regarding an address where kanji information of a plurality of kanji constituting the idiom information is stored, and delimiter information indicating the division of pronunciations of the idiom information. language processing device.
JP60096324A 1985-05-07 1985-05-07 Language processing device Granted JPS61255465A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP60096324A JPS61255465A (en) 1985-05-07 1985-05-07 Language processing device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP60096324A JPS61255465A (en) 1985-05-07 1985-05-07 Language processing device

Publications (2)

Publication Number Publication Date
JPS61255465A JPS61255465A (en) 1986-11-13
JPH0438018B2 true JPH0438018B2 (en) 1992-06-23

Family

ID=14161826

Family Applications (1)

Application Number Title Priority Date Filing Date
JP60096324A Granted JPS61255465A (en) 1985-05-07 1985-05-07 Language processing device

Country Status (1)

Country Link
JP (1) JPS61255465A (en)

Also Published As

Publication number Publication date
JPS61255465A (en) 1986-11-13

Similar Documents

Publication Publication Date Title
JP3196868B2 (en) Relevant word form restricted state transducer for indexing and searching text
JPS592125A (en) "kana" (japanese syllabary) "kanji" (chinese character) converting system
JPS62251876A (en) Language processing system
JPH0122948B2 (en)
JPH0438018B2 (en)
JPH03116375A (en) Information retriever
JPS58123129A (en) Converting device of japanese syllabary to chinese character
JPS6057421A (en) Documentation device
JPH0140372B2 (en)
JPS62212877A (en) Kanji-kana conversion device
JPH01229369A (en) character processing device
JPH0581945B2 (en)
JPS62144269A (en) information retrieval device
KR860000681B1 (en) Hangul/hanja(korean character/chinese character)word processor
JPS6389976A (en) Language analyzer
JPS6162970A (en) Kana-kanji conversion device
JPS59103136A (en) Kana-kanji conversion processing device
JP3273778B2 (en) Kana-kanji conversion device and kana-kanji conversion method
JPS61285571A (en) Compound word dictionary
JPS6180449A (en) Kana-Kanji conversion device
JPS59106029A (en) Kana (japanese syllabary) kanji (chinese character) converter
JPS603017A (en) Kana-kanji conversion processing device
JPS6198475A (en) Japanese text input device
JP2001318595A (en) Braille conversion system
JPH02289039A (en) Document processor with proper noun converting function