JPH0555912B2 - - Google Patents

Info

Publication number
JPH0555912B2
JPH0555912B2 JP61055683A JP5568386A JPH0555912B2 JP H0555912 B2 JPH0555912 B2 JP H0555912B2 JP 61055683 A JP61055683 A JP 61055683A JP 5568386 A JP5568386 A JP 5568386A JP H0555912 B2 JPH0555912 B2 JP H0555912B2
Authority
JP
Japan
Prior art keywords
data
bit
search
information
characters
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Expired - Lifetime
Application number
JP61055683A
Other languages
Japanese (ja)
Other versions
JPS62211728A (en
Inventor
Mikizo Kasugai
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NEC Corp
Tokai Television Broadcasting Co Ltd
Original Assignee
Tokai Television Broadcasting Co Ltd
Nippon Electric Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Tokai Television Broadcasting Co Ltd, Nippon Electric Co Ltd filed Critical Tokai Television Broadcasting Co Ltd
Priority to JP61055683A priority Critical patent/JPS62211728A/en
Publication of JPS62211728A publication Critical patent/JPS62211728A/en
Publication of JPH0555912B2 publication Critical patent/JPH0555912B2/ja
Granted legal-status Critical Current

Links

Landscapes

  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Description

【発明の詳細な説明】 (産業上の利用分野) 本発明は日本語による多量の文字情報から成る
データベースの中から検索条件に適合するデータ
を検索する日本語情報検索システムに関する。
DETAILED DESCRIPTION OF THE INVENTION (Field of Industrial Application) The present invention relates to a Japanese information retrieval system for searching data matching search conditions from a database consisting of a large amount of character information in Japanese.

(従来の技術) 従来のこの種の日本語情報検索システムは、日
本語によるデータを記憶するためのデータ部と、
予め定めた基本の日本語を起点としこの日本語を
つなぎ配列化したリストとを備え、データを登録
するときにはこのデータを構成する日本語をリス
トに登録しておき、情報検索時には検索条件に基
づきキーワードによつてリストを走査しこの結果
によつて指定されるデータをデータ部から読み出
すようにしている。
(Prior Art) A conventional Japanese information retrieval system of this type has a data section for storing data in Japanese;
It is equipped with a list that connects and arranges these Japanese words starting from a predetermined basic Japanese language, and when registering data, the Japanese words that make up this data are registered in the list, and when searching for information, it is possible to search based on search conditions. The list is scanned based on the keyword, and data specified by the result is read from the data section.

(発明が解決しようとする問題点) このような従来方式においては、階層構造のリ
ストを検索に使用するため、データベースのデー
タ量が尨大になるにつれて検索速度が劣化し、ま
た語によるキーワード検索が行なわれるため検索
に使用すべき文字列が適切なものであるか否か不
確かな場合にはヒツトしないケースがあるため、
データの登録時には1データに付帯して複数個の
日本語をリストに登録しておく必要がある場合が
多いため登録作業が厄介であるという問題点があ
る。
(Problems to be Solved by the Invention) In such conventional methods, since a hierarchically structured list is used for searching, the search speed deteriorates as the amount of data in the database increases, and keyword searches using words are difficult. is performed, so if it is uncertain whether the string to be used for the search is appropriate, there may be cases where the string is not hit.
When registering data, it is often necessary to register multiple Japanese words in a list along with one data, so there is a problem in that the registration work is cumbersome.

(問題点を解決するための手段) 本願発明の日本語情報検索システムは、日本語
によるデータを記憶するためのデータ部と、予め
定めた文字毎にすべての前記データの番号を指定
するためのビツト列からそれぞれ成る多重ビツト
テーブルとを設け、前記データを前記データ部に
登録する時には前記データ番号を発生させ該デー
タ内のすべての前記文字それぞれに対する前記多
重ビツトテーブル内の該データ番号指定ビツトを
オンしておき、指定された文字列をキーとして情
報を検索するときには該文字列内のすべての前記
文字に対応する前記ビツト列の同一位置ビツト同
士について検索条件に基づく論理演算を行いこの
結果によつて指定されるデータ番号のデータを前
記データ部から読み出すようにしたことを特徴と
する。
(Means for Solving the Problems) The Japanese information retrieval system of the present invention includes a data section for storing data in Japanese, and a section for specifying numbers of all the data for each predetermined character. A multiple bit table each consisting of a bit string is provided, and when registering the data in the data section, the data number is generated and the data number designation bit in the multiple bit table is assigned to each of all the characters in the data. When turned on, when searching for information using a specified character string as a key, a logical operation is performed based on the search condition on the bits at the same position in the bit string that correspond to all the characters in the string. Accordingly, the data of the specified data number is read out from the data section.

(実施例) 次に本発明の実施例について図面を参照して説
明する。
(Example) Next, an example of the present invention will be described with reference to the drawings.

ここで、具体的に表現するために多重ビツトテ
ーブルの多重度を2とし、それぞれ第1次ビツト
テーブル(あるいは上位索引)および第2次ビツ
トテーブル(あるいは下位索引)と表し、それぞ
れのビツトテーブルの大きさを1000ビツトと仮定
する。
Here, for concrete expression, the multiplicity of the multiple bit table is assumed to be 2, and each is expressed as a primary bit table (or upper index) and a secondary bit table (or lower index). Assume the size is 1000 bits.

第1図は本発明の一実施例を示すブロツク図で
ある。
FIG. 1 is a block diagram showing one embodiment of the present invention.

第1図は図面の繁雑化を避け説明を簡単化する
ために2つの第1次テーブルU1およびU2と、2
つの第2次テーブルL1およびL2と、データ部DA
が示されているが、第1次テーブルおよび第2次
テーブルは文字ごとに設けられる。すなわち、第
1図は2つの文字だけに対応するようになつてい
るが、検索のキーとして使用する文字をJIS第1
水準・第2水準の漢字(約6400字)、ひらがな、
カタカナおよび数字とすれば、これらの文字数と
同数の組数だけの第1ビツトテーブルと第2ビツ
トテーブルとが必要になる。
In order to avoid complicating the drawing and simplify the explanation, Figure 1 shows two primary tables U 1 and U 2 ,
two secondary tables L 1 and L 2 and a data section DA
is shown, a primary table and a secondary table are provided for each character. In other words, although Figure 1 corresponds to only two characters, the characters used as search keys are JIS 1.
Level/2nd level kanji (approximately 6400 characters), hiragana,
In the case of katakana and numbers, the same number of sets of first bit tables and second bit tables as the number of these characters are required.

データ部DAは100万個のエントリを有し、各
エントリに1つのデータを記憶する。データに固
有に付されるデータ番号をNとすると、エントリ
は(1)式を満足する座標(X,Y,Z)で指定され
る。
The data section DA has one million entries, and each entry stores one piece of data. Assuming that the data number uniquely assigned to data is N, an entry is designated by coordinates (X, Y, Z) that satisfy equation (1).

N=1000000(X-1)+1000(Y-1)+Z……(1) 第1次ビツトテーブルU1,U2と第2次ビツト
テーブルL1,L2はデータ部DAに対する索引部を
構成する。第1次ビツトテーブルU1,U2のアド
レスは上位索引の番号となる座標値×(本実施例
では1)で指定されそのビツト長は1000である。
また、第2次ビツトテーブルL1,L2のアドレス
は、上位索引のビツト、すなわち第1次ビツトテ
ーブルU1,U2のビツト位置に対応し下位索引の
番号となる座標値Y(001〜1000)で指定され、そ
のビツト長は1000である。第2次ビツトテーブル
L1,L2のビツト位置は下位索引のビツトである
座標値Z(001〜1000)に対応する。
N=1000000 (X -1 ) + 1000 (Y -1 ) + Z... (1) The primary bit tables U 1 and U 2 and the secondary bit tables L 1 and L 2 constitute an index section for the data section DA. do. The addresses of the primary bit tables U 1 and U 2 are specified by the coordinate value x (1 in this embodiment) which is the number of the upper index, and the bit length thereof is 1000.
Further, the addresses of the secondary bit tables L 1 and L 2 correspond to the bits of the upper index, that is, the bit positions of the primary bit tables U 1 and U 2 , and are the coordinate values Y (001 to 1000), and its bit length is 1000. Secondary bit table
The bit positions of L 1 and L 2 correspond to the coordinate value Z (001 to 1000), which is the bit of the lower index.

以上の説明から明らかなように、データ部DA
のデータは、第1次ビツトテーブルU1,U2と第
2次ビツトテーブルL1,L2との二重索引によつ
て索位されるようになつている。
As is clear from the above explanation, the data part DA
The data is indexed by a double index of the primary bit tables U 1 , U 2 and the secondary bit tables L 1 , L 2 .

次に本実施例の動作をデータの登録と情報検索
とに分けて説明する。
Next, the operation of this embodiment will be explained separately into data registration and information retrieval.

第2図は、例としてデータ「情報の蓄積と検
索」をとりあげて本実施例におけるデータ登録時
の動作を説明するためのものである。第1図にお
いて理解を容易化するために文字単位に個別に示
した第1次ビツトテーブルU1,U2と第2次ビツ
トテーブルL1,L2は、第2図および第3図にお
いては実際の姿に戻して一体化され、それぞれ
UXとLXとして表わされている。
FIG. 2 is for explaining the operation at the time of data registration in this embodiment, taking data "information storage and retrieval" as an example. The primary bit tables U 1 and U 2 and the secondary bit tables L 1 and L 2 , which are shown individually for each character in FIG. 1 to facilitate understanding, are not shown in FIGS. 2 and 3. Returned to their actual form and integrated, each
Represented as UX and LX.

いま、データ「情報の蓄積と検索」が入力され
ると、本システムによつてデータ番 号(本例で
は123002とする)を付与し、前述の式に基づいて
データ部の座標(1124002)にデータ番号123002
と共に格納する。
Now, when the data "Information storage and retrieval" is input, this system assigns a data number (123002 in this example) and sets the coordinates of the data section (1124002) based on the above formula. Data number 123002
Store with.

次いでデータ「情報の蓄積と検索」に含まれて
いるすべての文字を抽出し、1文字ごとに以下の
ようにして第1次ビツトテーブルUXのブロツク
アドレス001(以下UX001と記す)の第
124ビツトと第2次ビツトテーブルLX124の各
文字対応アドレスの第002ビツトをそれぞれ論理
“1”にする。なお、第2図においてはかな文字
「の」と「と」については図示を省略したが、か
な文字についても漢字と同様に登録される。
Next, extract all the characters included in the data "Information storage and retrieval", and write each character to the number of block address 001 (hereinafter referred to as UX001) of the primary bit table UX as follows.
The 124 bits and the 002nd bit of each character corresponding address in the secondary bit table LX124 are set to logic "1". Although the kana characters "no" and "to" are not shown in FIG. 2, the kana characters are also registered in the same way as the kanji characters.

第2図から明らかなように、文字「情」と文字
「積」に対しては、第1次ビツトテーブルUX0
01および第2次ビツトテーブルLX124の他
のビツトはすべて論理“0”であるため、これら
の文字はそれまで未登録であつたことがわかる。
文字「報」に対しては第1次ビツトテーブルUX
001の第001ビツトが“1”であるため第2次
ビツトテーブルLX001のいずれかのヒツト位
置(Zとする)が“1”である筈であり、データ
番号Zのデータがデータ部の座標(1001Z)に既
に格納されていることがわかる。
As is clear from Figure 2, for the characters ``jo'' and ``product'', the first bit table UX0
01 and all other bits of the secondary bit table LX 124 are logic "0", indicating that these characters were previously unregistered.
The first bit table UX is used for the character “information”.
Since the 001st bit of 001 is "1", any hit position (referred to as Z) in the secondary bit table LX001 should be "1", and the data of data number Z is the coordinate ( 1001Z).

第3図は、例として文字列「情報検索」をとり
あげて本実施例における情報検索時の動作を説明
するためのものである。いま、検索条件として文
字列「情報検索」を含むデータをすべて検索する
ように指示されているものとする。第3図におけ
る第1次ビツトテーブルUXおよび第2ビツトテ
ーブルLXの内容が第2図における内容と異なつ
ているのは、データ「情報の蓄積と検索」の登録
時から文字列「情報検索」による情報検索時まで
の間にデータ「東海テレビ情報検索システム」等
他のデータが登録されて“1”のビツト数が増え
たことによる。
FIG. 3 is for explaining the operation at the time of information retrieval in this embodiment, taking the character string "information search" as an example. Assume that you are instructed to search for all data containing the character string "information search" as a search condition. The reason why the contents of the primary bit table UX and the second bit table LX in Fig. 3 are different from those in Fig. 2 is that the character string ``information retrieval'' was created from the time of registration of the data ``information storage and retrieval''. This is because other data such as the data "Tokai Television Information Search System" was registered before the information search, and the number of "1" bits increased.

文字列「情報検索」が入力されると、第1次ビ
ツトテーブルUX001の文字「情」、「報」、
「検」および「索」に対応する各アドレスの内容
が同一ビツト位置同士で論理積演算される。
When the character string "information search" is input, the characters "information", "information",
The contents of each address corresponding to "search" and "search" are ANDed with the same bit positions.

この第1回目の論理積演算の結果により、第1
次ビツトテーブルUX001の第003ビツトと第
124ビツトが“1”になつたものすると、次に第
2次ビツトテーブルLX003とLX124それぞ
れについて、文字「情」、「報」、「検」および
「索」に対応する各アドレスの内容が同一ビツト
位置同士で論理積演算される。
Based on the result of this first logical AND operation, the first
The 003rd bit and the 003rd bit of the next bit table UX001
Assuming that the 124 bit becomes "1", the contents of each address corresponding to the characters "information", "information", "search", and "search" are the same for the second bit tables LX003 and LX124, respectively. A logical AND operation is performed between bit positions.

この第2回目の論理積演算の結果により、第2
次ビツトテーブルLX003については全ビツト
とも“0となり、第2次ビツトテーブルLX12
4については第002ビツトと第456ビツトが“1”
になつたとする。このことから、第2次ビツトテ
ーブルLX003に対応するデータ群には「情」、
「報」、「検」および「索」の4文字を含んでいる
データが一つも無いことがわかり、第2次ビツト
テーブルLX124に対応するデータ群のうちの
データ番号123002と123456で指定されるデータは
これら4文字を含んでいることがわかる。
Based on the result of this second logical AND operation, the second
Regarding the next bit table LX003, all bits become "0", and the second bit table LX12
For 4, the 002nd bit and 456th bit are “1”
Suppose that it becomes From this, the data group corresponding to the secondary bit table LX003 includes "emotion",
It is found that there is no data that contains the four characters ``report'', ``search'', and ``search'', which are specified by data numbers 123002 and 123456 of the data group corresponding to the secondary bit table LX124. It can be seen that the data contains these four characters.

したがつて、データ部DAからデータ番号
123002と123456のデータを読み出すが、データ番
号123002のデータは「情報の蓄積と検索」であつ
て検索キーとして使用した「情報検索」の構成文
字をすべて含むが「情報検索」という語は含んで
いないが、データ番号123456のデータは「東海テ
レビ情報検索システム」であつて「情報検索」と
いう語を含んでいることがわかるので、後者が検
索条件に適合するデータであることになる。
Therefore, data number from data part DA
Data 123002 and 123456 are read, but the data number 123002 is "information storage and retrieval" and includes all the constituent characters of "information retrieval" used as a search key, but does not include the word "information retrieval". However, it can be seen that the data with data number 123456 is "Tokai Television Information Search System" and includes the word "information search," so the latter is the data that meets the search conditions.

この場合、検索者の記憶が不確かで「情報検
索」を一応検索キーとしたものの、たとえば情報
検索端末のCRT画面に表示された検索結果のデ
ータを見て検索条件は「情報」および「検索」を
含むデータであつたことを思い出すことも度々あ
り得るものである。そのような場合には検索条件
を上記のように変更して検索をやり直すことがで
きる。
In this case, the searcher's memory is uncertain and the search key is "information search", but after looking at the search result data displayed on the CRT screen of the information search terminal, the search conditions are "information" and "search". It is possible that we often recall that the data contained . In such a case, you can change the search conditions as described above and start the search again.

以上に述べた実施例においては、索引の多重度
を2としているが本発明はこれに限定されること
はなく、たとえば(1)式の代りに(2)式を使用するよ
うにした三重索引方式を採用した実施例は容易に
実現できる。
In the embodiment described above, the multiplicity of the index is set to 2, but the present invention is not limited to this. For example, a triple index in which formula (2) is used instead of formula (1) is used. An embodiment employing this method can be easily realized.

N=1000000000(X-1)+1000000(Y-1)+1000
(Z-1)+W ……(2) さらに、以上に述べた実施例においては、検索
条件として検索キーのすべての文字列を含むデー
タを検索することにしているが、他の検索条件、
たとえば「情報」と「検索」のうちの一つのよう
に、複数個の文字列のうちのいずれかを含むデー
タを検索するような検索条件に対応できるように
した実施例も、ビツトテーブルの同一ビツト位置
同士の論理演算内容を変更することによつて容易
に実現できる。
N=1000000000(X -1 )+1000000(Y -1 )+1000
(Z -1 )+W ...(2) Furthermore, in the embodiment described above, data including all the character strings of the search key is searched as a search condition, but other search conditions,
For example, an embodiment that can support a search condition that searches for data that includes one of a plurality of character strings, such as one of "information" and "search", is also possible. This can be easily realized by changing the contents of logical operations between bit positions.

なお、以上の説明は日本語による情報検索につ
いて行なつているが、本発明は、漢字構成を採る
中国語による情報検索の場合にはよりいつそう適
性である。
Note that although the above explanation has been made regarding information retrieval in Japanese, the present invention is even more suitable for information retrieval in Chinese, which employs a kanji structure.

また、本発明における多重ビツトテーブルは、
情報検索システムに限らず一般の索引(存在チエ
ツク用等)にも適用可能である。
Furthermore, the multiple bit table in the present invention is
It is applicable not only to information retrieval systems but also to general indexes (for checking existence, etc.).

(発明の効果) 本発明によれば、従来方式におけるように、検
索キーワードとして使用される可能性のある語を
索引部に階層的にリスト化しておき、検索キーワ
ードでこのリストを走査して検索条件に適合する
データを見出す代りに、データ登録時には登録さ
れるデータに含まれる全文字について多重ビツト
テーブル内の該当するデータ番号指定ビツトをオ
ンにしておくようにした“字典方式”の採用によ
つて、文字が有限であるために、登録されるデー
タの量が尨大化しても索引部の規模は自ずから収
束し、さらに情報検索時には文字単位に多重ビツ
トテーブルを読み出して同一ビツト位置同士につ
いて検索条件に基づく論理演算を行なうようにし
た“フリーワード検索”の採用によつて、検索の
ステツプ数が少なくなるために検索速度の低下が
なく、また検索キーワードはいつたん分解されて
中間的には文字単位に検索が進行するため、検索
キーワードの不確かさにも対処できるように従来
方式で行なつているような1データに付帯した複
数個の日本語登録を不要化するという効果があ
る。
(Effects of the Invention) According to the present invention, unlike the conventional method, words that may be used as search keywords are hierarchically listed in the index section, and the list is scanned with the search keyword to perform the search. Instead of finding data that meets the conditions, the system uses a ``dictionary method'' that turns on the corresponding data number designation bit in the multiple bit table for all characters included in the registered data at the time of data registration. Since the number of characters is finite, even if the amount of registered data increases, the scale of the index section will naturally converge.Furthermore, when searching for information, multiple bit tables can be read out character by character and searches can be made for the same bit position. By adopting a "free word search" that performs logical operations based on conditions, the number of search steps is reduced, so there is no decrease in search speed, and the search keywords are broken down at any time and intermediately Since the search proceeds character by character, it has the effect of eliminating the need for multiple Japanese registrations associated with one piece of data, which is done in the conventional method, in order to deal with the uncertainty of search keywords.

【図面の簡単な説明】[Brief explanation of drawings]

第1図は本発明の一実施例を説明するための図
を示し、第2図および第3図は本実施例の動作を
説明するための図を示す。 U1,U2,UX……第1次ビツトテーブル、L1,
L2……第2次ビツトテーブル、DA……データ
部。
FIG. 1 shows a diagram for explaining one embodiment of the present invention, and FIGS. 2 and 3 show diagrams for explaining the operation of this embodiment. U 1 , U 2 , UX...primary bit table, L 1 ,
L2 ...Second bit table, DA...Data section.

Claims (1)

【特許請求の範囲】 1 日本語によるデータを記憶するためのデータ
部と、 予め定めた文字毎にすべての前記データの番号
を指定するためのビツト列からそれぞれ成る多重
ビツトテーブルとを設け、 前記データを前記データ部に登録する時には前
記データ番号を発生させ該データ内のすべての前
記文字それぞれに対する前記多重ビツトテーブル
内の該データ番号指定ビツトをオンにしておき、 指定された文字列をキーとして情報を検索する
ときには該文字列内のすべての前記文字に対応す
る前記ビツト列の同一位置ビツト同士について検
索条件に基づく論理演算を行いこの結果によつて
指定されるデータ番号のデータを前記データ部か
ら読み出すようにしたことを特徴とする日本語情
報検索システム。
[Scope of Claims] 1. A data section for storing data in Japanese, and a multiple bit table each consisting of a bit string for specifying the number of all the data for each predetermined character, When registering data in the data section, generate the data number, turn on the data number designation bit in the multiple bit table for each of the characters in the data, and use the specified character string as a key. When searching for information, a logical operation is performed based on the search condition on bits at the same position in the bit string that correspond to all the characters in the character string, and the data of the data number specified by the result is stored in the data section. A Japanese information retrieval system characterized by reading information from the .
JP61055683A 1986-03-12 1986-03-12 Japanese information retrieving system Granted JPS62211728A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP61055683A JPS62211728A (en) 1986-03-12 1986-03-12 Japanese information retrieving system

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP61055683A JPS62211728A (en) 1986-03-12 1986-03-12 Japanese information retrieving system

Publications (2)

Publication Number Publication Date
JPS62211728A JPS62211728A (en) 1987-09-17
JPH0555912B2 true JPH0555912B2 (en) 1993-08-18

Family

ID=13005698

Family Applications (1)

Application Number Title Priority Date Filing Date
JP61055683A Granted JPS62211728A (en) 1986-03-12 1986-03-12 Japanese information retrieving system

Country Status (1)

Country Link
JP (1) JPS62211728A (en)

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP3220865B2 (en) * 1991-02-28 2001-10-22 株式会社日立製作所 Full text search method
JP2986865B2 (en) * 1989-07-24 1999-12-06 株式会社日立製作所 Data search method and device

Also Published As

Publication number Publication date
JPS62211728A (en) 1987-09-17

Similar Documents

Publication Publication Date Title
US5551049A (en) Thesaurus with compactly stored word groups
JPH06325092A (en) Customer information retrieval system
JPS62211728A (en) Japanese information retrieving system
JPH01149127A (en) Information retrieving device
JPH0353378A (en) Name search method to search for surnames with homophones and homonyms
EP0649106B1 (en) Compactly stored word groups
JPH03118661A (en) Word retrieving device
JPH1021262A (en) Information retrieval device
JPH1021252A (en) Information retrieval device
JPH06348688A (en) Kana (japanese syllabary)/kanji (chinese character) conversion system
JPS6118071A (en) Dictionary retrieving system
JPH02148174A (en) Data retrieving device
JPH0991304A (en) Information retrieval method, information retrieval system, and information retrieval storage medium
JPH0531190B2 (en)
JPS6162163A (en) Japanese language word processor device
JPH0236988B2 (en)
KR100675161B1 (en) Search and save method using mobile terminal
JPH0746353B2 (en) Japanese text input device
JPH08180060A (en) Electronic dictionary display
JPH0748218B2 (en) Information processing equipment
JPH0375960A (en) Character processing device frequency change method
JPS5922255B2 (en) Kanji input method
JPH04322361A (en) Kanji retrieval system for word processor
JPS60252949A (en) Information retrieving method
JPS63314672A (en) Kana/kanji conversion processor

Legal Events

Date Code Title Description
EXPY Cancellation because of completion of term