JPH0765030A - Text search method and apparatus - Google Patents

Text search method and apparatus

Info

Publication number
JPH0765030A
JPH0765030A JP5212515A JP21251593A JPH0765030A JP H0765030 A JPH0765030 A JP H0765030A JP 5212515 A JP5212515 A JP 5212515A JP 21251593 A JP21251593 A JP 21251593A JP H0765030 A JPH0765030 A JP H0765030A
Authority
JP
Japan
Prior art keywords
sentence
similarity
search
word
phrase
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
JP5212515A
Other languages
Japanese (ja)
Inventor
Masato Yajima
真人 矢島
Noriko Koyama
紀子 小山
Yuuji Shimizu
勇詞 清水
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Toshiba Corp
Original Assignee
Toshiba Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Toshiba Corp filed Critical Toshiba Corp
Priority to JP5212515A priority Critical patent/JPH0765030A/en
Publication of JPH0765030A publication Critical patent/JPH0765030A/en
Withdrawn legal-status Critical Current

Links

Landscapes

  • Machine Translation (AREA)
  • Document Processing Apparatus (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

(57)【要約】 【目的】 1文での単語の使われ方を考慮し、文全体と
して検索対象の入力文との類似程度を判断して検索を行
い、広範囲かつ高精度の検索を出来るようにする。 【構成】 入力部1から入力した文字列を処理して表示
部3で表示する。入力部1からの1文を単語分割部4で
単語単位に分割し、この分割された単語から文節を文節
合成部5で作成する。この文節と文節の類似度を文節類
似度計算部6で計算して、文全体の類似度を類似文判断
部7で判断し、この判断で1文全体として類似している
と判断された検索対象文のみを表示部3に出力する。
(57) [Summary] [Purpose] Considering how words are used in one sentence, the degree of similarity with the input sentence that is the search target is determined for the entire sentence, and the search is performed, enabling a wide range and highly accurate search. To do so. [Structure] A character string input from the input unit 1 is processed and displayed on the display unit 3. One sentence from the input unit 1 is divided into word units by the word division unit 4, and a phrase synthesis unit 5 creates a phrase from the divided words. The similarity between the phrase and the phrase is calculated by the phrase similarity calculation unit 6, and the similarity of the entire sentence is determined by the similar sentence determination unit 7. By this determination, it is determined that the entire sentence is similar. Only the target sentence is output to the display unit 3.

Description

【発明の詳細な説明】Detailed Description of the Invention

【0001】[0001]

【産業上の利用分野】本発明は、日本語ワードプロセッ
サ、光学的文字読取装置(OCR)などに利用し、多量
の格納文書中から類似する文章を検索する文章検索方法
及びその装置に関する。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a sentence retrieval method and a device for retrieving a similar sentence from a large amount of stored documents, which is used in a Japanese word processor, an optical character reader (OCR) and the like.

【0002】[0002]

【従来の技術】近時の文書保管では、日本語ワードプロ
セッサ、光学的文字読取装置等が多用されている。これ
らの装置では、従来の記録紙又はマイクロフィルムなど
に代えて磁気ディスクや光磁気ディスクなどの記憶媒体
を用いて文書を保存している。この場合、多量の保存文
書から所望の文書を検索するテキスト検索技術が採用さ
れている。この検索では入力された検索文と文字列の一
致を判断し、又は使用している単語の一致を判断し、そ
の一致する文章を検索している。
2. Description of the Related Art Recently, Japanese word processors, optical character readers, etc. are widely used for document storage. In these devices, documents are stored using a storage medium such as a magnetic disk or a magneto-optical disk instead of a conventional recording paper or microfilm. In this case, a text search technique for searching a desired document from a large amount of stored documents is adopted. In this search, a match between the input search sentence and the character string is determined, or a match between the used words is determined, and the matching sentence is searched.

【0003】[0003]

【発明が解決しようとする課題】しかしながら上記のよ
うな従来例における文章検索では、文全体における文字
列を含んだ文を検索している。又は単語で切り離す単語
切りを行った後に、この単語と一致する文を検索してい
る。この場合、文字列の検索では、この文字列の一字で
も相違すると検索できない。さらに、単語での検索で
は、単語の使われ方を考慮しないで文中を検索してしま
う。したがって、必要な文章を検索できなかったり、無
用な文章を検索してしまうことがあり、広範囲かつ高精
度の検索が出来ない欠点がある。
However, in the sentence retrieval in the conventional example as described above, the sentence including the character string in the entire sentence is retrieved. Alternatively, after cutting the word into words, a sentence matching this word is searched. In this case, in the search of the character string, even if one character of this character string is different, it cannot be searched. Furthermore, in the word search, the sentence is searched without considering the usage of the word. Therefore, it may not be possible to search for a necessary sentence or may search for a useless sentence, and there is a drawback that a wide range and high precision cannot be searched.

【0004】本発明は、このような従来の技術における
欠点を解決するものであり、1文での単語の使われ方を
考慮し、また文全体として、検索対象の入力文との類似
程度を判断した検索が可能になり、広範囲かつ高精度の
検索が出来る文章検索方法及びその装置の提供を目的と
する。
The present invention solves the above-mentioned drawbacks of the prior art by taking into account the usage of words in one sentence, and by considering the degree of similarity between the entire sentence and the input sentence to be searched. It is an object of the present invention to provide a text search method and a device thereof that enable a judged search and can perform a wide range and highly accurate search.

【0005】[0005]

【課題を解決するための手段】上記目的を達成するため
に、本発明の文章検索方法は検索入力文及び検索対象文
を単語単位に分割し、次に文節単位にまとめて、検索入
力文と検索対象文とを1文節ごとに比較し、さらに文節
どうしの類似度を計算し、この計算で1文全体として類
似していると判断された検索対象文のみを出力を行う。
In order to achieve the above object, the sentence search method of the present invention divides a search input sentence and a search target sentence into word units, and then combines them into phrase units to obtain a search input sentence. The search target sentence is compared for each phrase, and the similarity between the phrases is calculated, and only the search target sentence determined to be similar as a whole sentence by this calculation is output.

【0006】この場合、文節どうしの類似度は、文節内
の自立語どうしの類似程度を計算した自立語類似度と、
文節内の付属語どうしの類似程度を計算した付属語類似
度との組み合わせ又はいずれか一方としている。
In this case, the similarity between the bunsetsus is the independent word similarity calculated by calculating the similarity between the independent words in the bunsetsu,
The degree of similarity between adjunct words in a clause is combined with the calculated adjunct word similarity or either.

【0007】また、本発明の文章検索装置は、文字列を
入力する入力手段と、文字列を処理して表示する表示手
段と、入力手段からの1文を単語単位に分割する単語分
割手段と、単語分割手段で分割された単語から文節を合
成して作成する文節合成手段と、文節合成手段で得られ
た文節どうしの類似度を計算する文節類似度計算手段
と、文全体の類似度を判断する類似文判断手段とを備え
ている。
Further, the sentence retrieval device of the present invention comprises an input means for inputting a character string, a display means for processing the character string and displaying it, and a word dividing means for dividing one sentence from the input means into word units. , A bunsetsu synthesizing unit that synthesizes bunsetsu from words divided by the word slicing unit, a bunsetsu similarity calculating unit that calculates the similarity between bunsetsu obtained by the bunsetsu synthesizing unit, and a similarity of the entire sentence. And a similar sentence judging means for judging.

【0008】[0008]

【作用】このように本発明の文章検索方法及びその装置
では、単語単位に分割して文節単位にまとめた検索入力
文と検索対象文とを1文節ごとに比較し、文節どうしの
類似度を計算して1文全体が類似している場合に検索対
象文のみを出力している。したがって、従来の単語の一
致のみで検索する場合に比較して、1文内での使われ方
が相違する文を検索することなく精度の高い検索が行わ
れる。また、自立語類似度及び付属語類似度を個別に評
価しているので、文全体の類似程度を判断しながら、検
索を行なうことができ、幅広い高度な検索が行われる。
As described above, in the sentence search method and apparatus according to the present invention, the search input sentence divided into word units and collected in phrase units is compared with the sentence to be searched for each phrase, and the similarity between the phrases is determined. Only the search target sentence is output when the whole sentence is similar after the calculation. Therefore, as compared with the conventional case of searching only by matching words, a highly accurate search can be performed without searching for sentences that are used differently in one sentence. In addition, since the independent word similarity and the adjunct word similarity are individually evaluated, it is possible to perform a search while judging the degree of similarity of the entire sentence, and a wide range of advanced searches are performed.

【0009】[0009]

【実施例】次に、本発明の文章検索方法及びその装置の
実施例を図面を参照して詳細に説明する。図1は本発明
の文章検索方法及びその装置における実施例の構成を示
すブロック図である。図1において、この文章検索方法
及びその装置は文字列や制御指示を行うための入力部1
と、この装置の各部の制御を行う制御部2と、入力部1
から入力された文字列などを表示する表示部3と、入力
された文字列を解析して単語単位に分割する単語分割部
4とを有している。さらに、単語分割部4で分割した単
語をまとめて、文節に合成する文節合成部5と、文節と
他の文節とを比較して、その類似度を計算する文節類似
度計算部6と、この文節類似度計算部6の計算結果か
ら、文全体の類似の程度を判断する類似文判断部7と、
検索対象文を保存記憶する検索対象文記憶部8とを有し
ている。
DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Next, an embodiment of a text search method and apparatus according to the present invention will be described in detail with reference to the drawings. FIG. 1 is a block diagram showing the configuration of an embodiment of a text search method and apparatus according to the present invention. In FIG. 1, this text search method and apparatus are provided with an input unit 1 for inputting character strings and control instructions.
And a control unit 2 for controlling each unit of the apparatus and an input unit 1.
It has a display unit 3 for displaying a character string and the like input from, and a word dividing unit 4 for analyzing the input character string and dividing it into word units. Furthermore, the words divided by the word dividing unit 4 are put together, and a bunsetsu synthesizing unit 5 for synthesizing them into bunsetsu, a bunsetsu similarity calculating unit 6 for comparing the bunsetsu with other bunsetsu, and calculating their similarity, A similar sentence determination unit 7 that determines the degree of similarity of the entire sentence from the calculation result of the phrase similarity calculation unit 6;
The search target sentence storage unit 8 stores and stores the search target sentence.

【0010】次に、この実施例の構成における動作につ
いて説明する。入力部1から、検索入力文となる文字列
が入力される。この検索入力文は、制御部2を通じて、
単語分割部4に送られ、ここで単語単位に分割される。
Next, the operation of the configuration of this embodiment will be described. A character string to be a search input sentence is input from the input unit 1. This search input sentence is
It is sent to the word division unit 4 and is divided here in word units.

【0011】図2は、この単語分割部4で文字列を解析
して単語単位に分割した状態を示す図である。図2にお
いて、単語分割部4では「太郎の父は、医者である」の
文字列を「太郎(名詞)」、「の(助詞)」、「父(名
詞)」、「は(助詞)」、「医者(名詞)」、「である
(助動詞終止形)」のように、分割単語に品詞、活用形
などの文法用語を付加している。このように分割された
単語は、文節合成部5によって文節単位に合成される。
この文節は原則として自立語と、これに続く付属語から
構成される。
FIG. 2 is a diagram showing a state in which the word dividing unit 4 analyzes a character string and divides it into word units. In FIG. 2, in the word division unit 4, the character string "Taro's father is a doctor" is changed to "Taro (noun)", "no (particle)", "father (noun)", "ha (particle)". , "Doctor (noun)", "is de (auxiliary verb end form)", grammatical terms such as part-of-speech and inflected forms are added to the divided words. The words thus divided are combined by the phrase combining unit 5 in units of phrases.
As a general rule, this bunsetsu consists of an independent word followed by an annex.

【0012】図3は、図2で解析した単語から文節を合
成した場合を示す図である。図3において、図2中に示
す単語が「太郎の」、「父は」、「医者である」の文節
として合成されている。
FIG. 3 is a diagram showing a case where a phrase is synthesized from the words analyzed in FIG. In FIG. 3, the words shown in FIG. 2 are synthesized as the clauses of “Taro's”, “father's”, and “is a doctor”.

【0013】次に、検索対象文記憶部8から、検索対象
文の1文を取り出し、検索入力文と同様に処理して単語
に分割し、かつ、文節の合成を順次行なう。文節類似度
計算部6では、検索入力文における1文節と検索対象文
の1文節とを比較して、類似度を計算する。この計算は
「(検索入力文の文節数)×(検索対象文の文節数)」
の総当りで行なう。この計算結果に基づいて類似文判断
部7は文全体としての類似度を評価する。この評価で類
似している場合は、この検索対象文を表示部3に出力す
る。検索対象文記憶部8に未検索の検索対象文がある場
合、記憶されている次の1文をとり出して同様の処理を
繰り返す。
Next, one sentence of the retrieval target sentence is fetched from the retrieval target sentence storage unit 8, processed in the same manner as the retrieval input sentence, divided into words, and the synthesizing of clauses is sequentially performed. The phrase similarity calculator 6 compares one phrase in the search input sentence with one phrase in the search target sentence to calculate the similarity. This calculation is “(number of phrases in the search input sentence) x (number of phrases in the search target sentence)”
Will be brute force. Based on this calculation result, the similar sentence judgment unit 7 evaluates the similarity of the whole sentence. If they are similar in this evaluation, this search target sentence is output to the display unit 3. When there is an unsearched search target sentence in the search target sentence storage unit 8, the next one stored sentence is taken out and the same processing is repeated.

【0014】次に文節類似度計算部6から類似文判断部
7における処理手順を説明する。図4は文節類似度計算
部6から類似文判断部7における動作の処理手順を示す
フローチャートである。ここでは検索入力文として図2
及び図3に示す「太郎の父は医者である」が入力され、
単語分割及び文節合成が行なわれている場合とする。ま
た、検索対象文として「次郎の父は医師である」が取り
出された場合を説明する。
Next, the processing procedure from the phrase similarity calculation unit 6 to the similar sentence determination unit 7 will be described. FIG. 4 is a flowchart showing the processing procedure of the operation from the phrase similarity calculation unit 6 to the similar sentence determination unit 7. Here, as a search input sentence,
And “Taro's father is a doctor” shown in FIG. 3 is entered,
It is assumed that word division and phrase synthesis are performed. Further, a case where "Jiro's father is a doctor" is retrieved as the search target sentence will be described.

【0015】図5は、この「次郎の父は医師である」の
文章を単語分割部4で解析して単語単位に分割した状態
を示す図である。図5において、単語分割部4では、
「次郎の父は、医者である」の文字列を「次郎(名
詞)」、「の(助詞)」、「父(名詞)」、「は(助
詞)」、「医者(名詞)」、「である(助動詞終止
形)」に示すように、分割した単語に品詞、活用形など
の文法用語を付加している。このように分割された単語
は、文節合成部5によって文節単位に合成される。原則
として、文節は自立語と後続する付属語から構成され
る。
FIG. 5 is a diagram showing a state in which the sentence "Jiro's father is a doctor" is analyzed by the word dividing unit 4 and divided into word units. In FIG. 5, in the word division unit 4,
The character string "Jiro's father is a doctor" is "Jiro (noun)", "no (particle)", "father (noun)", "ha (particle)", "doctor (noun)", " As shown in “(Auxiliary verb end form)”, grammatical terms such as parts of speech and inflectional forms are added to the divided words. The words thus divided are combined by the phrase combining unit 5 in units of phrases. In principle, a bunsetsu consists of an independent word followed by an adjunct.

【0016】図6は、図5で解析した単語から文節を合
成した状態を示す図である。図6において、図5中の単
語が「次郎の」、「父は」、「医者である」の文節に合
成されている。
FIG. 6 is a diagram showing a state in which a phrase is synthesized from the words analyzed in FIG. In FIG. 6, the words in FIG. 5 are synthesized into the clauses of “Jiro no”, “father is”, and “is a doctor”.

【0017】文節類似度計算部6は検索入力文の文節
を、先頭から順次取り出す。この際、未処理の文節があ
れば(図4中のステップ(S)10)、次に検索対象文
の文節を先頭から順次とり出して、未処理の文節を調べ
る(ステップ11)。この調べで未処理の文節がない場
合は、検索入力文における、次の文節の処理に進めて
(ステップ12)、さらに、検索対象文の先頭の文節処
理を開始する(ステップ13)。
The phrase similarity calculation unit 6 sequentially extracts the phrases of the search input sentence from the beginning. At this time, if there is an unprocessed phrase (step (S) 10 in FIG. 4), the phrases of the search target sentence are sequentially taken out from the beginning and the unprocessed phrase is examined (step 11). If there is no unprocessed phrase in this check, the process proceeds to the process of the next phrase in the search input sentence (step 12), and the phrase process at the beginning of the search target sentence is started (step 13).

【0018】ここで未処理の文節があれば、自立語の類
似度の計算を行なう(ステップ14)。例えば、自立語
が完全に一致している場合は10点、類語の場合は7点
等とする評価点を定める。検索入力文の第1文節の「太
郎は」と検索対象文の第1文節の「次郎は」とでは、自
立語である「太郎」と「次郎」とが一致しないため0点
となる。この計算を(検索入力文の全文節)×(検索対
象文の全文節)の総当りで行う。
If there is any unprocessed phrase, the similarity of the independent word is calculated (step 14). For example, an evaluation score is set to 10 points when the independent words are completely the same and 7 points when the synonyms are similar. Since the independent words “Taro” and “Jiro” do not match between the first phrase “Taroha” of the search input sentence and the first phrase “Jiroha” of the search target sentence, the score is 0. This calculation is performed by a brute force of (all phrases of search input sentence) × (all phrases of search target sentence).

【0019】図7は、この自立語の類似度の計算結果を
示す図である。図7において、検索入力文と検索対象文
の第2文節どうしの「父」と、第4文節どうしでの「医
者」とが、それぞれ自立語として一致しているので10
点が配点されている。続いて付属語の類似度の計算を行
なう(ステップ15)。付属語の類似度の計算の場合
も、完全に一致している場合は10点を評価点を定め、
また類似している場合は7点というように評価点を定め
て総当りの計算を行う。
FIG. 7 is a diagram showing the calculation result of the similarity of the independent word. In FIG. 7, since the “father” between the second bunsetsu and the “doctor” between the fourth bunsetsu of the search input sentence and the search target sentence match each other as independent words, 10
Points are assigned. Then, the similarity of the attached words is calculated (step 15). Also in the case of the calculation of the degree of similarity of the adjunct words, if they are completely in agreement, the evaluation point is set to 10 points,
If they are similar to each other, an evaluation score is set to 7 points and a brute force calculation is performed.

【0020】図8は付属語の類似度の計算結果を示す図
である。図8において、検索入力文と検索対象文におけ
る第1文節どうし、第2文節どうし、第3文節どうしに
おける「の(助詞)」、「は(助詞)」、「である(助
動詞終止形)」がそれぞれ一致しており、各10点づつ
が配点されている。そして、検索対象文の文節処理を次
の文節に進める(ステップ16)。さらに、この処理を
繰り返して検索入力文の未処理の文節がなくなった場合
に(ステップ10)、類似文判断部7が文全体の類似度
を計算する(ステップ17)。文全体の評価として、図
7、図8で示した自立語類似度と付属語類似度の配点を
加算して評価する。この評価の結果から検索入力文の文
節と検索対象文の文節とが1対1に対応する形で、配点
が最大になる組み合わせを求める。
FIG. 8 is a diagram showing the result of calculating the degree of similarity of attached words. In FIG. 8, "no (particle)", "wa (particle)", and "is (auxiliary verb end form)" in the first input phrase, the second output phrase, and the third output phrase in the search input sentence and the search target sentence. Are coincident with each other, and 10 points are assigned to each. Then, the phrase processing of the search target sentence is advanced to the next phrase (step 16). Further, when this processing is repeated and there is no unprocessed phrase in the search input sentence (step 10), the similar sentence judgment unit 7 calculates the similarity of the entire sentence (step 17). For the evaluation of the entire sentence, the scores of the independent word similarity and the adjunct word similarity shown in FIGS. 7 and 8 are added and evaluated. From the result of this evaluation, the combination of the phrase of the search input sentence and the phrase of the search target sentence is in one-to-one correspondence, and the combination with the maximum score is obtained.

【0021】図9は、この自立語類似度と付属語類似度
の配点を加算した図である。図9において、検索入力文
と検索対象文とにおける第1文節どうしの「太郎の」と
「次郎の」との10点と、第2文節どうしの「父は」と
「父は」との20点と、第3文節のどうしの「医者であ
る」と「医者である」との20点とを合計した50点の
評価点が得られる。この50点の評価点から文全体が類
似していると判断して(ステップ18)、検索対象文の
「次郎の父は医者である」を表示部3に出力する(ステ
ップ19)。なお、ここで得られた評価点が0点の場合
には検索対象文の「次郎の父は医者である」を表示部3
に出力しないで表示しないまま終了する。
FIG. 9 is a diagram in which the points of the independent word similarity and the adjunct word similarity are added. In FIG. 9, in the search input sentence and the search target sentence, 10 points of “Taro no” and “Jiro no” of the first phrase, and 20 points of “Father wa” and “Father wa” of the second phrase A total of 50 points, which is the sum of the points and 20 points of “being a doctor” and “being a doctor” between the third clauses, are obtained. Based on these 50 evaluation points, it is determined that the sentences are similar to each other (step 18), and the search target sentence "Jiro's father is a doctor" is output to the display unit 3 (step 19). If the evaluation score obtained here is 0, the search target sentence "Jiro's father is a doctor" is displayed on the display unit 3
It ends without displaying it without outputting to.

【0022】なお、この実施例では文節を自立語と後続
する付属語とで構成して説明したが、自立語及び付属語
は本要旨を逸脱しない範囲で自由に構成できる。また、
類似文判断部7が文全体の類似度を判断する際に、この
実施例のように単に自立語類似度と付属語類似度とを加
算するのではなく、自立語類似度のみで判断し、また、
付属語類似度のみで判断したり、自立語類似度と付属語
類似度の両方で評価された文節の類似度を判断するよう
に変更しても良い。
In this embodiment, the bunsetsu is composed of independent words and the following adjuncts, but the independent words and adjuncts can be freely constructed without departing from the scope of the present invention. Also,
When the similar sentence judgment unit 7 judges the similarity of the whole sentence, it does not simply add the independent word similarity and the adjunct word similarity as in this embodiment, but judges only by the independent word similarity, Also,
The determination may be made based only on the adjunct word similarity, or the similarity of the bunsetsu evaluated by both the independent word similarity and the adjunct word similarity may be changed.

【0023】[0023]

【発明の効果】以上の説明から明らかなように、本発明
の文章検索方法及びその装置は、単語単位に分割して文
節単位にまとめた検索入力文と検索対象文とを1文節ご
とに比較し、文節どうしの類似度を計算して1文全体が
類似している場合に検索対象文のみを出力しているの
で、従来の単語の一致のみで検索する場合に比較して、
1文内での使われ方が相違する文を検索することなく精
度の高い検索が出来るという効果を有する。さらに、自
立語類似度及び付属語類似度を個別に評価しているの
で、文全体の類似程度を判断しながら、検索を行なうこ
とができ、幅広い高度な検索が出来るという効果を有す
る。
As is clear from the above description, the sentence search method and apparatus according to the present invention compare the search input sentence divided into word units and the sentence to be searched for each phrase. However, since the similarity between the clauses is calculated and only the search target sentence is output when the entire sentence is similar, compared with the conventional case of searching only by matching the words,
This has the effect that a highly accurate search can be performed without searching for sentences that differ in how they are used within one sentence. Furthermore, since the independent word similarity and the adjunct word similarity are individually evaluated, it is possible to perform a search while judging the degree of similarity of the entire sentence, and it is possible to perform a wide range of advanced searches.

【図面の簡単な説明】[Brief description of drawings]

【図1】本発明の文章検索方法及びその装置の実施例に
おける構成を示すブロック図である。
FIG. 1 is a block diagram showing the configuration of an embodiment of a text search method and apparatus according to the present invention.

【図2】図1中の単語分割部で文字列を単語単位に分割
した状態を示す図である。
FIG. 2 is a diagram showing a state in which a character string is divided into word units by a word dividing unit in FIG.

【図3】実施例にあって解析した単語から合成した文節
を示す図である。
FIG. 3 is a diagram showing clauses synthesized from the analyzed words in the embodiment.

【図4】実施例にあって文節類似度計算部から類似文判
断部における動作の処理手順を示すフローチャートであ
る。
FIG. 4 is a flowchart showing a processing procedure of an operation from a phrase similarity calculation unit to a similar sentence determination unit in the embodiment.

【図5】実施例にあって単語分割部で文字列を解析して
単語単位に分割した状態を示す図である。
FIG. 5 is a diagram showing a state in which a character string is analyzed by a word dividing unit and divided into word units in the embodiment.

【図6】実施例にあって解析した単語から合成した文節
を示す図である。
FIG. 6 is a diagram showing clauses synthesized from the analyzed words in the embodiment.

【図7】実施例にあって自立語の類似度の計算結果を示
す図である。
FIG. 7 is a diagram showing a calculation result of the similarity of independent words in the embodiment.

【図8】実施例にあって付属語の類似度の計算結果を示
す図である。
FIG. 8 is a diagram showing calculation results of similarity of attached words in the embodiment.

【図9】実施例にあって自立語類似度と付属語類似度の
配点を加算した図である。
FIG. 9 is a diagram in which points of independent word similarity and adjunct word similarity are added in the embodiment.

【符号の説明】[Explanation of symbols]

1…入力部 2…制御部 3…表示部 4…単語分割部 5…文節合成部 6…文節類似度
計算部 7…類似文判断部 8…検索対象文
記憶部
DESCRIPTION OF SYMBOLS 1 ... Input part 2 ... Control part 3 ... Display part 4 ... Word division part 5 ... Phrase synthesis part 6 ... Phrase similarity degree calculation part 7 ... Similar sentence judgment part 8 ... Search target sentence storage part

───────────────────────────────────────────────────── フロントページの続き (51)Int.Cl.6 識別記号 庁内整理番号 FI 技術表示箇所 9194−5L 15/403 350 C ─────────────────────────────────────────────────── ─── Continuation of the front page (51) Int.Cl. 6 Identification code Internal reference number FI technical display location 9194-5L 15/403 350 C

Claims (3)

【特許請求の範囲】[Claims] 【請求項1】 検索入力文及び検索対象文を単語単位に
分割し、次に文節単位にまとめて、前記検索入力文と前
記検索対象文とを1文節ごとに比較し、さらに文節どう
しの類似度を計算し、この計算で1文全体として類似し
ていると判断された検索対象文のみを出力することを特
徴とする文章検索方法。
1. A search input sentence and a search target sentence are divided into word units, then grouped into phrase units, the search input sentence and the search target sentence are compared for each phrase, and the similarities between the clauses A sentence search method characterized by calculating a degree and outputting only a search target sentence determined to be similar as a whole sentence by this calculation.
【請求項2】 文節どうしの類似度は、文節内の自立語
どうしの類似程度を計算した自立語類似度と、文節内の
付属語どうしの類似程度を計算した付属語類似度との組
み合わせ又はいずれか一方であることを特徴とする請求
項1記載の文章検索方法。
2. The degree of similarity between bunsetsu is a combination of an independent word similarity calculated by calculating the degree of similarity between independent words within a bunsetsu and an adjunct word similarity calculated by calculating the degree of similarity between auxiliary words within a bunsetsu, or The text search method according to claim 1, wherein the text search method is one of the two.
【請求項3】 文字列を入力する入力手段と、文字列を
処理して表示する表示手段と、前記入力手段からの1文
を単語単位に分割する単語分割手段と、前記単語分割手
段で分割された単語から文節を合成して作成する文節合
成手段と、前記文節合成手段で得られた文節どうしの類
似度を計算する文節類似度計算手段と、文全体の類似度
を判断する類似文判断手段とを備える文章検索装置。
3. Input means for inputting a character string, display means for processing and displaying the character string, word dividing means for dividing one sentence from the input means into word units, and division by the word dividing means. A bunsetsu synthesizing means for synthesizing and creating bunsetsu from the selected words, a bunsetsu similarity calculating means for calculating the similarity between the bunsetsu obtained by the bunsetsu synthesizing means, and a similar sentence judgment for judging the similarity of the whole sentence. A text search device comprising means.
JP5212515A 1993-08-27 1993-08-27 Text search method and apparatus Withdrawn JPH0765030A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP5212515A JPH0765030A (en) 1993-08-27 1993-08-27 Text search method and apparatus

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP5212515A JPH0765030A (en) 1993-08-27 1993-08-27 Text search method and apparatus

Publications (1)

Publication Number Publication Date
JPH0765030A true JPH0765030A (en) 1995-03-10

Family

ID=16623952

Family Applications (1)

Application Number Title Priority Date Filing Date
JP5212515A Withdrawn JPH0765030A (en) 1993-08-27 1993-08-27 Text search method and apparatus

Country Status (1)

Country Link
JP (1) JPH0765030A (en)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH05311191A (en) * 1990-05-09 1993-11-22 Sanyo Chem Ind Ltd Coagulant for edible oil
JP2001243245A (en) * 2000-03-01 2001-09-07 Nippon Telegr & Teleph Corp <Ntt> Similar sentence search method and apparatus, and recording medium storing similar sentence search program
JP2011039639A (en) * 2009-08-07 2011-02-24 Yahoo Japan Corp Retrieval device and method

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH05311191A (en) * 1990-05-09 1993-11-22 Sanyo Chem Ind Ltd Coagulant for edible oil
JP2001243245A (en) * 2000-03-01 2001-09-07 Nippon Telegr & Teleph Corp <Ntt> Similar sentence search method and apparatus, and recording medium storing similar sentence search program
JP2011039639A (en) * 2009-08-07 2011-02-24 Yahoo Japan Corp Retrieval device and method

Similar Documents

Publication Publication Date Title
US6876998B2 (en) Method for cross-linguistic document retrieval
US6345253B1 (en) Method and apparatus for retrieving audio information using primary and supplemental indexes
US7805303B2 (en) Question answering system, data search method, and computer program
US6904429B2 (en) Information retrieval apparatus and information retrieval method
US10191892B2 (en) Method and apparatus for establishing sentence editing model, sentence editing method and apparatus
US5794177A (en) Method and apparatus for morphological analysis and generation of natural language text
JP3691844B2 (en) Document processing method
US7376634B2 (en) Method and apparatus for implementing Q&amp;A function and computer-aided authoring
JP3040945B2 (en) Document search device
JPH03172966A (en) Similar document retrieving device
JPH1145241A (en) Kana-kanji conversion system and computer-readable recording medium storing a program for causing a computer to function as each means of the system
JP2000200281A (en) Information retrieval apparatus, information retrieval method, and recording medium recording information retrieval program
JP2000222427A (en) Related word extraction device, related word extraction method, and storage medium storing related word extraction program
JP3198932B2 (en) Document search device
JPH05282367A (en) Related keyword automatic generator
JPH0765030A (en) Text search method and apparatus
JP2529418B2 (en) Document search device
JP2002108888A (en) Digital content keyword extraction apparatus and method, and computer-readable recording medium
JPH08339376A (en) Foreign language search device and information search system
JP3006526B2 (en) Similar document search method and similar document search device
JP2002215672A (en) Search expression expansion method, search system, and search expression expansion computer program
KR100657016B1 (en) Search method by combining source for recognition of relevant passages in texts
JPH09160928A (en) Document search method and apparatus
JP3436109B2 (en) Related search formula search device and computer-readable recording medium storing related search formula search program
JPH1145270A (en) Abstract sentence creation support system and computer-readable recording medium recording a program for causing a computer to function as the system

Legal Events

Date Code Title Description
A300 Application deemed to be withdrawn because no request for examination was validly filed

Free format text: JAPANESE INTERMEDIATE CODE: A300

Effective date: 20001031