JPH1031677A - Document search device - Google Patents

Document search device

Info

Publication number
JPH1031677A
JPH1031677A JP8185018A JP18501896A JPH1031677A JP H1031677 A JPH1031677 A JP H1031677A JP 8185018 A JP8185018 A JP 8185018A JP 18501896 A JP18501896 A JP 18501896A JP H1031677 A JPH1031677 A JP H1031677A
Authority
JP
Japan
Prior art keywords
document
data
document data
unit
search
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
JP8185018A
Other languages
Japanese (ja)
Inventor
Hiroshi Tanano
裕氏 棚野
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sharp Corp
Original Assignee
Sharp Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sharp Corp filed Critical Sharp Corp
Priority to JP8185018A priority Critical patent/JPH1031677A/en
Publication of JPH1031677A publication Critical patent/JPH1031677A/en
Withdrawn legal-status Critical Current

Links

Landscapes

  • Document Processing Apparatus (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
  • Machine Translation (AREA)

Abstract

(57)【要約】 【課題】 それぞれが異なる言語で記述された複数の文
書データから所定言語で記述された検索要求に意味的に
近似の文書データを検索できる文書検索装置を提供す
る。 【解決手段】 それぞれが異なる言語で記述された複数
の文書データと、複数文書データのそれぞれに対応して
意味的特徴を示す文書特徴ベクトルとを格納した蓄積文
書データベース2と、所定言語で記述された検索要求を
入力するための検索要求入力部3と、入力された検索要
求の意味的特徴を示す検索要求特徴ベクトルを生成する
検索要求特徴ベクトル生成部4と、検索部5とを備え、
検索部5は生成された検索要求特徴ベクトルと各文書特
徴ベクトルとの内積値に基づいて、入力された検索要求
に対して意味的に近似する文書データを蓄積文書データ
ベース2中の複数文書データから検索し出力する。
(57) [Summary] [PROBLEMS] To provide a document retrieval apparatus capable of retrieving document data semantically similar to a retrieval request described in a predetermined language from a plurality of document data each described in a different language. SOLUTION: A stored document database 2 storing a plurality of document data, each described in a different language, and a document feature vector indicating a semantic feature corresponding to each of the plurality of document data, A search request input unit 3 for inputting the search request, a search request feature vector generation unit 4 for generating a search request feature vector indicating a semantic feature of the input search request, and a search unit 5.
The search unit 5 extracts, based on the inner product value of the generated search request feature vector and each document feature vector, document data semantically similar to the input search request from the plurality of document data in the stored document database 2. Search and output.

Description

【発明の詳細な説明】DETAILED DESCRIPTION OF THE INVENTION

【0001】[0001]

【発明の属する技術分野】この発明は自然言語により記
述された文書データの検索を行なう文書検索装置に関
し、特に、異種言語で記述された複数の文書データが混
在する中から検索要求を満たす文書データの情報を獲得
するのに有効な文書検索装置に関する。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a document retrieval apparatus for retrieving document data described in a natural language, and more particularly, to a document data satisfying a retrieval request from a mixture of a plurality of document data described in different languages. The present invention relates to a document search device that is effective for obtaining information of a document.

【0002】[0002]

【従来の技術】文書データを検索する場合に、キーワー
ド検索のような表記の一致を前提とした硬直な検索方式
を補う方法として、単語の特徴ベクトルを用いて文書デ
ータを検索する方式が既に提案されている。
2. Description of the Related Art As a method of supplementing a rigid retrieval method that presupposes notation matching such as keyword retrieval when retrieving document data, a method of retrieving document data using a feature vector of a word has already been proposed. Have been.

【0003】この提案された文書検索方式においては、
いくつかの特徴単語(通常、数十から数百)によって特
徴空間を定義し、各単語に関して特徴単語との関係づけ
を数値化した特徴ベクトルを作成し、文書データに含ま
れる単語の特徴ベクトルの和をもって該文書データの特
徴ベクトルとする。そして、操作者からの検索要求に対
しても同様に特徴ベクトルが計算されて、これより文書
データと検索要求の双方における互いに正規化された
(ベクトル長を等しくした)特徴ベクトル間の近似度
(内積値)が計算され、この近似度が大きいほど検索要
求に近い文書データであると判断される。
[0003] In the proposed document retrieval system,
A feature space is defined by several feature words (usually several tens to several hundreds), a feature vector is created by quantifying the relationship between each word and the feature word, and a feature vector of the word included in the document data is created. The sum is used as the feature vector of the document data. Then, a feature vector is similarly calculated for a search request from the operator, and the degree of approximation between mutually normalized (equalized vector length) feature vectors in both the document data and the search request is calculated. The inner product value) is calculated, and it is determined that the larger the degree of approximation is, the closer the document data is to the search request.

【0004】この文書検索方式による場合、実用上高い
検索精度を実現するためには相当数の単語に対して特徴
ベクトルを作成する必要があり、このことは文書検索装
置の構築者にとっては多大な負担となるが、この負担を
軽減させるための単語の特徴ベクトルの自動付与に関す
る技術が特開平6−195388号公報において開示さ
れている。
In the case of this document search system, it is necessary to create feature vectors for a considerable number of words in order to achieve practically high search accuracy, which is a great deal for the builder of a document search apparatus. Japanese Patent Application Laid-Open No. 6-195388 discloses a technique for automatically assigning a feature vector of a word to reduce the burden.

【0005】[0005]

【発明が解決しようとする課題】ネットワーク技術の発
達により、必要な情報を入手するにあたって国境の存在
を意識する必要は皆無に等しく、海外に溢れている膨大
な情報を入手し利用したいという要求が高まってきてい
るが、依然として言語の違いという壁が存在する。従来
の特徴ベクトルを用いた文書検索方式は、特徴ベクトル
を付与している単語辞書が、ある言語Aに関するもので
あれば言語Aで記述されている文書データのみが検索対
象となり、自ずと検索対象とする情報の範囲が限定され
てしまい、実用性に優れないという問題があった。
[Problems to be Solved by the Invention] With the development of network technology, there is almost no need to be aware of the existence of borders when obtaining necessary information, and there is a demand to obtain and use a vast amount of information overflowing overseas. Although growing, there is still the barrier of language differences. In a conventional document search method using a feature vector, if a word dictionary to which a feature vector is assigned relates to a certain language A, only document data described in the language A is to be searched. However, there is a problem that the range of information to be obtained is limited, and the practicability is not excellent.

【0006】それゆえにこの発明の目的は、それぞれが
異なる言語で記述された複数の文書データから所定の言
語で記述された検索要求に意味的に近似の文書データを
検索できる文書検索装置を提供することである。
Therefore, an object of the present invention is to provide a document retrieval apparatus capable of retrieving document data semantically similar to a retrieval request described in a predetermined language from a plurality of document data each described in a different language. That is.

【0007】この発明の他の目的は、それぞれが異なる
言語で記述された複数の文書データから所定の言語で記
述された検索要求に意味的に近似の文書データを検索す
る場合に、検索の対象となる文書データを容易に追加で
きる文書検索装置を提供することである。
Another object of the present invention is to search for document data that is semantically similar to a search request described in a predetermined language from a plurality of document data each described in a different language. It is an object of the present invention to provide a document search device that can easily add document data to be used.

【0008】この発明のさらなる他の目的は、それぞれ
が異なる言語で記述された複数の文書データから所定の
言語で記述された検索要求に意味的に近似の文書データ
を検索する場合に、検索して得られた意味的に近似の文
書データを検索要求の記述言語で翻訳して提示できる文
書検索装置を提供することである。
Still another object of the present invention is to provide a method for retrieving document data semantically similar to a retrieval request described in a predetermined language from a plurality of document data each described in a different language. It is an object of the present invention to provide a document search apparatus capable of translating and presenting semantically approximated document data obtained by using a search request description language.

【0009】[0009]

【課題を解決するための手段】請求項1に記載の文書検
索装置は、それぞれが異なる言語で記述された複数の文
書データと、複数文書データのそれぞれに対応してその
意味的特徴を示す文書特徴データとが格納された文書デ
ータ蓄積部と、所定の言語で記述された検索要求を入力
するための要求入力部と、要求入力部により入力された
検索要求の意味的特徴を示す要求特徴データを検出する
要求特徴データ検出部と、文書データ蓄積部の複数文書
データから、要求特徴データ検出部により検出された要
求特徴データに基づいて、要求入力部により入力された
検索要求に対して意味的に近似する意味的近似文書デー
タを検索し出力する検索部とを備えて構成される。
According to a first aspect of the present invention, there is provided a document search apparatus comprising: a plurality of document data written in different languages; and a document showing a semantic characteristic corresponding to each of the plurality of document data. A document data storage unit storing the characteristic data, a request input unit for inputting a search request described in a predetermined language, and request characteristic data indicating semantic characteristics of the search request input by the request input unit A request feature data detection unit for detecting a search request; and a search request input by the request input unit based on the request feature data detected by the request feature data detection unit from a plurality of document data in the document data storage unit. And a retrieval unit that retrieves and outputs semantically approximated document data that approximates.

【0010】請求項1に係る文書検索装置によれば、入
力された検索要求の記述言語と検索対象となる蓄積文書
データの記述言語が一致しなくとも、意味的に近似した
文書データを検索し出力できる。したがって、文書デー
タ蓄積部中の複数文書データのそれぞれの記述言語にか
かわらず、操作者は要求する文書データを得ることがで
きるので、操作者に対して言語の壁を超えて検索対象と
なる情報源の範囲が広く設定されて実用性が向上する。
According to the first aspect of the present invention, even if the description language of the input search request does not match the description language of the stored document data to be searched, the document search apparatus searches for document data that is semantically similar. Can output. Therefore, the operator can obtain the requested document data regardless of the description language of each of the plurality of document data in the document data storage unit. The range of sources is set wider to improve practicability.

【0011】請求項2に記載の文書検索装置は、請求項
1に記載の装置がさらに、異なる言語のそれぞれについ
ての複数の単語と、各単語に対応してその意味的特徴を
示す単語特徴ベクトルデータとを格納した単語辞書部を
備え、文書特徴データおよび要求特徴データのそれぞれ
は、文書データおよび検索要求のそれぞれを構成する各
単語に対応の単語辞書部中の単語特徴ベクトルデータの
総和を正規化した単位ベクトルデータであり、意味的近
似文書データは、文書データ蓄積部の複数文書データの
うち、対応する単位ベクトルデータと検索要求に対応の
単位ベクトルデータとの内積値が大きい文書データであ
るよう構成される。
According to a second aspect of the present invention, there is provided the document search apparatus, wherein the apparatus according to the first aspect further comprises a plurality of words for each of different languages, and a word feature vector indicating a semantic feature corresponding to each word. And a word dictionary unit storing data and the document feature data and the requested feature data, respectively, and normalize the sum of the word feature vector data in the word dictionary unit corresponding to each word constituting the document data and the search request. And the semantically approximated document data is document data having a large inner product value between the corresponding unit vector data and the unit vector data corresponding to the search request among the plurality of document data in the document data storage unit. It is configured as follows.

【0012】請求項2に係る文書検索装置によれば、蓄
積文書データ中の複数の文書データ中から入力された検
索要求に意味的に近似する文書データをベクトルデータ
の内積値計算処理により求めることができるので、容易
に、かつ速やかに操作者が要求する文書データを検索
(特定)することができる。
According to the second aspect of the present invention, the document data which is semantically similar to the search request input from the plurality of document data in the stored document data is obtained by the inner product value calculation processing of the vector data. Therefore, the document data requested by the operator can be searched (specified) easily and promptly.

【0013】請求項3に記載の文書検索装置は、請求項
2の文書検索装置の単語辞書部が、異なる言語間の同義
である単語のすべては、同一の単語特徴ベクトルデータ
に対応づけられるよう構成される。
According to a third aspect of the present invention, the word dictionary section of the second aspect of the present invention is arranged such that all words having the same meaning in different languages are associated with the same word feature vector data. Be composed.

【0014】請求項3に係る文書検索装置によれば、単
語辞書部は異なる言語間の同義の単語は全て1つの単語
特徴ベクトルデータに対応づけられるように構成される
ので、該検索装置において単語辞書部に関する消費記憶
容量は抑制されて、該検索装置のメモリ有効利用が図ら
れる。
According to the third aspect of the present invention, the word dictionary section is configured such that all synonymous words between different languages are associated with one word feature vector data. The memory capacity consumed by the dictionary unit is suppressed, and the memory of the search device is effectively used.

【0015】また、いずれの言語の単語であっても同義
語であれば一意に単語特徴ベクトルデータが得られるの
で、記述言語にかかわらず検索要求に意味的に近似した
文書データを精度よく検索し出力することができる。
Further, word feature vector data can be uniquely obtained for words in any language as long as they are synonyms, so that document data that is semantically similar to the search request can be searched accurately regardless of the description language. Can be output.

【0016】請求項4に記載の文書検索装置は、請求項
1ないし3のいずれかに記載の文書検索装置がさらに、
新規の文書データを入力するための文書データ入力部
と、文書データ入力部により入力された文書データの文
書特徴データを生成する文書特徴データ生成部とを備
え、文書データ入力部により入力された文書データは文
書特徴データ生成部により生成された文書特徴データと
対応づけられて文書データ蓄積部に格納されるよう構成
される。
According to a fourth aspect of the present invention, there is provided the document search apparatus according to any one of the first to third aspects, further comprising:
A document data input unit for inputting new document data, a document characteristic data generation unit for generating document characteristic data of the document data input by the document data input unit, and a document input by the document data input unit The data is stored in the document data storage unit in association with the document feature data generated by the document feature data generation unit.

【0017】請求項4に係る文書検索装置によれば、入
力された新規文書データの記述言語にかかわらず、検索
のために必要な文書特徴データを生成して、検索対象で
ある文書データをその文書特徴データとともに文書デー
タ蓄積部に追加格納できる。したがって、検索対象とな
る文書データのための文書データ蓄積部の構築が容易に
可能となる。
According to the fourth aspect of the present invention, the document characteristic data required for the search is generated regardless of the description language of the new document data input, and the document data to be searched is converted to the document data. It can be additionally stored in the document data storage together with the document characteristic data. Therefore, it is possible to easily construct a document data storage unit for document data to be searched.

【0018】請求項5に記載の文書検索装置は、請求項
1ないし4のいずれかに記載の文書検索装置がさらに、
異なる言語間の翻訳に必要なデータを保持する翻訳辞書
部と、検索部による検索出力時、意味的近似文書データ
の記述言語が検索要求を記述する所定言語に一致しない
場合に、翻訳辞書部の内容を参照して意味的近似文書デ
ータを所定言語に翻訳する翻訳処理部とを備えて構成さ
れる。
According to a fifth aspect of the present invention, there is provided a document search apparatus according to any one of the first to fourth aspects, further comprising:
A translation dictionary unit that holds data necessary for translation between different languages; and a search unit that outputs a search when the description language of the semantically approximated document data does not match a predetermined language that describes the search request. A translation processing unit for translating the semantically approximate document data into a predetermined language with reference to the contents.

【0019】請求項5に係る文書検索装置によれば、操
作者が入力した検索要求の記述言語と検索対象となる文
書データの記述言語が一致しない場合でも、検索して得
られた文書データの持つ情報を検索要求の記述言語、す
なわち操作者が要求する言語表記に翻訳して提示するこ
とができる。したがって、操作者は要求する文書データ
のもつ情報が読解不可能な言語で記述されている場合で
も、言語の壁なく要求情報を獲得することができる。
According to the document search device of the present invention, even if the description language of the search request input by the operator does not match the description language of the document data to be searched, The information possessed can be translated and presented in the description language of the search request, that is, the language notation requested by the operator. Therefore, even when the information of the requested document data is described in a language that cannot be read, the operator can acquire the requested information without a language barrier.

【0020】[0020]

【発明の実施の形態】以下、この発明の実施の形態1〜
3について図面を参照し説明する。
BEST MODE FOR CARRYING OUT THE INVENTION Hereinafter, embodiments 1 to 1 of the present invention will be described.
3 will be described with reference to the drawings.

【0021】図1は、この発明の実施の形態1〜3に適
用される文書検索装置のブロック構成図である。図にお
いて文書検索装置は単語辞書データベース1、蓄積文書
データベース2、所望の自然言語で記述された検索要求
を入力するために、たとえばキーボードなどからなる検
索要求入力部3、検索要求入力部3から入力された検索
要求の特徴ベクトルを生成するための検索要求特徴ベク
トル生成部4、検索部5、処理結果などのデータを表示
する表示部6、キー入力部などからなり、外部から該装
置に新規の(未登録の)文書データを入力するための新
規文書入力部7、与えられるデータの記述言語を判定す
る言語判定部8、文書特徴ベクトル生成部9、翻訳辞書
データベース10および翻訳処理部11を含む。
FIG. 1 is a block diagram of a document search apparatus applied to the first to third embodiments of the present invention. In the figure, a document search device is inputted from a word dictionary database 1, a stored document database 2, a search request input unit 3 composed of, for example, a keyboard, and the like, for inputting a search request described in a desired natural language. A search request feature vector generation unit 4, a search unit 5, a display unit 6 for displaying data such as processing results, a key input unit, etc., for generating a feature vector of the searched search request. It includes a new document input unit 7 for inputting (unregistered) document data, a language determination unit 8 for determining the description language of given data, a document feature vector generation unit 9, a translation dictionary database 10, and a translation processing unit 11. .

【0022】図2は図1の単語辞書データベース1の構
成例を示す図である。図2において単語辞書データベー
ス1は複数の異なる自然言語A、B、C、D、…のそれ
ぞれについての単語辞書DA、DB、DC、DD、…お
よび複数の単語特徴ベクトルViからなり、各単語辞書
は複数の単語データWi(i=1、2、3、…)を含
み、各単語データWiはその意味的特徴を示す単語特徴
ベクトルViが対応づけられる。
FIG. 2 is a diagram showing a configuration example of the word dictionary database 1 of FIG. In FIG. 2, the word dictionary database 1 includes word dictionaries DA, DB, DC, DD,... For a plurality of different natural languages A, B, C, D,. Includes a plurality of word data Wi (i = 1, 2, 3,...), And each word data Wi is associated with a word feature vector Vi indicating its semantic feature.

【0023】図3は、図1の蓄積文書データベース2の
構成例を示す図である。図3において蓄積文書データベ
ース2は該装置に登録されてそれぞれが異なる自然言語
で記述された複数の文書データSDiと、各文書データ
SDiに対応して文書特徴ベクトルVDiとを含む。な
お、文書データSDiの文書特徴ベクトルVDiは文書
特徴ベクトル生成部9により求めることができるが、そ
の詳細は後述する。
FIG. 3 is a diagram showing a configuration example of the stored document database 2 of FIG. In FIG. 3, the stored document database 2 includes a plurality of document data SDi registered in the device and described in different natural languages, and a document feature vector VDi corresponding to each document data SDi. Note that the document feature vector VDi of the document data SDi can be obtained by the document feature vector generation unit 9, the details of which will be described later.

【0024】翻訳辞書データベース10はある自然言語
で記述された文書データSDiを他の自然言語に翻訳す
る際に、翻訳処理部11により参照されるデータベース
であり、たとえばある自然言語から他の自然言語への翻
訳における条件付けが定義されたものである。
The translation dictionary database 10 is a database that is referred to by the translation processing unit 11 when translating the document data SDi described in a certain natural language into another natural language. This defines the conditioning in the translation into.

【0025】(実施の形態1)図4は、この発明の実施
の形態1の文書データ検索処理動作に必要な図1の文書
検索装置の部分構成図であり、図5はこの発明の実施の
形態1の文書データ検索処理動作のフローチャートであ
る。
(Embodiment 1) FIG. 4 is a partial configuration diagram of the document search apparatus of FIG. 1 necessary for the document data search processing operation of Embodiment 1 of the present invention, and FIG. 5 is an embodiment of the present invention. 9 is a flowchart of a document data search processing operation according to the first embodiment.

【0026】図6はこの発明の実施の形態1の文書デー
タ検索処理動作において検索要求から検索要求特徴ベク
トルを得る手順を説明する図である。図7はこの発明の
実施の形態1の文書データ検索処理動作における検索要
求に近似の文書データを得る手順を説明する図である。
FIG. 6 is a diagram for explaining a procedure for obtaining a search request feature vector from a search request in the document data search processing operation according to the first embodiment of the present invention. FIG. 7 is a diagram for explaining a procedure for obtaining document data similar to a search request in the document data search processing operation according to the first embodiment of the present invention.

【0027】次に、この発明の実施の形態1として、図
1の文書検索装置において、蓄積文書データベース2か
ら検索要求に意味的に近似の文書データSDiを検索し
て出力する処理動作について図5のフローチャートに従
い説明する。
Next, as a first embodiment of the present invention, a processing operation of searching and outputting document data SDi that is semantically similar to a search request from the stored document database 2 in the document search apparatus of FIG. 1 will be described with reference to FIG. This will be described according to the flowchart of FIG.

【0028】まず、検索要求入力部3において入力され
た検索要求は、検索要求特徴ベクトル生成部4へ送られ
る(S301)。次に、検索要求特徴ベクトル生成部4
は、単語辞書データベース1を参照しながら入力された
検索要求を形態素解析して、検索要求に含まれる各単語
の単語特徴ベクトルを抽出し、これらの単語特徴ベクト
ルの総和を正規化した単位ベクトルを検索要求特徴ベク
トルとして検索部5へ送る(S302)。
First, the search request input in the search request input unit 3 is sent to the search request feature vector generation unit 4 (S301). Next, the search request feature vector generation unit 4
Performs a morphological analysis of a search request input with reference to the word dictionary database 1, extracts word feature vectors of each word included in the search request, and generates a unit vector obtained by normalizing the sum of these word feature vectors. The search request feature vector is sent to the search unit 5 (S302).

【0029】検索要求特徴ベクトル生成部4による検索
要求から検索要求特徴ベクトルが得られるまでの詳細手
順は図6に示される。図6では「パソコン通信の将来」
という検索要求SRが入力された例を示しているが、検
索要求SRは「パソコン」のように単語であっても構わ
ないし、複数の単語からなる文であってもよいし、複数
の文よりなる文書であってもよい。図6において検索要
求特徴ベクトル生成部4は「パソコン通信の将来」とい
う検索要求SRが入力されると、これを形態素(単語)
解析して「パソコン」「通信(の)」「将来」に分解
し、各形態素(単語データWi)に対応する単語特徴ベ
クトルViを該検索要求の記述言語に対応の単語辞書デ
ータベース1から抽出して、それぞれベクトル長を揃え
た(たとえば長さ1)ものをV1、V2、V3とする。
この例では、各単語特徴ベクトルViの大きさはすべて
同じになるようにしたが、状況によって単語ごとにベク
トル長を変えて重み付けをしてもよい。たとえば、専門
的分野に関する検索要求SRなので、検索要求SR中の
専門用語については重み付けを変更して検索効率を上げ
ることができる。
FIG. 6 shows the detailed procedure from the search request by the search request feature vector generation unit 4 to the obtaining of the search request feature vector. Figure 6 shows the “future of personal computer communication”
Is shown, the search request SR may be a word such as “PC”, a sentence composed of a plurality of words, or a plurality of sentences. Document. In FIG. 6, when a search request SR “future of personal computer communication” is input, the search request feature vector generation unit 4 converts this into a morpheme (word).
It is analyzed and decomposed into “PC”, “communication (no)”, and “future”, and a word feature vector Vi corresponding to each morpheme (word data Wi) is extracted from the word dictionary database 1 corresponding to the description language of the search request. In this case, V1, V2, and V3 have the same vector length (for example, length 1).
In this example, the size of each word feature vector Vi is all the same, but the vector length may be changed for each word and weighted depending on the situation. For example, since the search request SR is related to a specialized field, the weight of the technical terms in the search request SR can be changed to improve the search efficiency.

【0030】最後に、V1、V2、V3の総和を正規化
して検索要求特徴ベクトルVSとする。
Finally, the sum of V1, V2, and V3 is normalized to obtain a search request feature vector VS.

【0031】ここで、単語辞書データベース1は図2に
示されるように、複数の自然言語対応にしておくことに
より、その範囲内においては任意の自然言語による検索
要求SRを受付けることができる。すなわち図2におい
て、自然言語A、B、C、D、…間の同義語の単語デー
タWi群と1つの単語特徴ベクトルViを関係づけてお
くことにより、いずれの自然言語の単語データWiから
も全く同様に単語特徴ベクトルViが得られるので、検
索要求SRの意味的特徴を示す検索要求特徴ベクトルV
Sを容易に得ることができる。たとえば、図2の例で
は、検索要求SRが言語Aによる「家」であっても、言
語Bによる「house」であっても、同じ単語特徴ベ
クトルV1が抽出できることが示されている。
Here, as shown in FIG. 2, the word dictionary database 1 can accept a search request SR in an arbitrary natural language within the range by supporting a plurality of natural languages. That is, in FIG. 2, by associating the word data Wi group of synonyms between natural languages A, B, C, D,... With one word feature vector Vi, the word data Wi of any natural language can be obtained. Since the word feature vector Vi is obtained in exactly the same way, the search request feature vector V indicating the semantic feature of the search request SR is obtained.
S can be easily obtained. For example, the example in FIG. 2 shows that the same word feature vector V1 can be extracted whether the search request SR is “house” in language A or “house” in language B.

【0032】図5に戻って、検索部5は図7に示される
ように、蓄積文書データベース2に保持されている各文
書データSDiに対する各文書の意味的特徴を示す文書
特徴ベクトルVDiと検索要求特徴ベクトルVSの内積
VS・VDiを計算し、この内積値を検索要求SRと文
書データSDiとの意味的近似度Aiと定義する。そこ
で、意味的近似度Aiが大きくなる文書データSDi、
すなわち検索要求SRに対して意味的に近いと考えられ
る文書データSDiから順に表示部6に送られる(S3
03)。この際、意味的に最も近いもの1件だけを送っ
てもよいし、上位n件を送ってもよいし、一定のしきい
値を満たす意味的近似度Aiを有するものを送ってもよ
い。最後に、表示部6は送られてきた文書データSDi
を表示して(S304)、一連の処理が終了する。
Returning to FIG. 5, as shown in FIG. 7, the retrieval unit 5 transmits a document request vector VDi indicating the semantic characteristics of each document to each document data SDi held in the stored document database 2 and a retrieval request. The inner product VS · VDi of the feature vector VS is calculated, and the inner product value is defined as the semantic similarity Ai between the search request SR and the document data SDi. Therefore, the document data SDi in which the semantic similarity Ai increases,
That is, the document data SDi that is considered to be semantically close to the search request SR are sent to the display unit 6 in order (S3).
03). At this time, only the one that is semantically closest may be sent, the top n items may be sent, or one having a semantic approximation Ai that satisfies a certain threshold may be sent. Finally, the display unit 6 displays the sent document data SDi.
Is displayed (S304), and the series of processing ends.

【0033】この文書データの検索方法によれば、検索
要求と各文書データ間の意味的近似度は互いの記述言語
に全く依存しない形で計算されるので、検索要求と異な
る言語で記述された文書データも検索対象として検索範
囲を容易に拡張することができる。
According to this document data search method, the degree of semantic similarity between the search request and each piece of document data is calculated in a manner that does not depend on the description language at all. The search range can be easily expanded for document data as a search target.

【0034】(実施の形態2)図8は、この発明の実施
の形態2の文書データ登録処理動作に必要な図1の文書
検索装置の部分構成図であり、図9はこの発明の実施の
形態2の文書データ登録処理動作のフローチャートであ
る。
(Embodiment 2) FIG. 8 is a partial configuration diagram of the document search apparatus of FIG. 1 necessary for the document data registration processing operation of Embodiment 2 of the present invention, and FIG. 9 is an embodiment of the present invention. 13 is a flowchart of a document data registration processing operation according to the second embodiment.

【0035】図10はこの発明の実施の形態2の文書デ
ータ登録処理動作において言語判定部で作成されるデー
タの説明図である。図11はこの発明の実施の形態2の
文書データ登録処理動作において新規文書データを文書
特徴ベクトルを付与して蓄積文書データベースに格納す
る動作の説明図である。
FIG. 10 is an explanatory diagram of data created by the language determination section in the document data registration processing operation according to the second embodiment of the present invention. FIG. 11 is an explanatory diagram of an operation of adding new document data to a document feature vector and storing the new document data in a stored document database in the document data registration processing operation according to the second embodiment of the present invention.

【0036】次にこの発明の実施の形態2として、図1
の文書検索装置において蓄積文書データベース2に新規
の文書データを登録する処理動作について図9のフロー
チャートに従い説明する。
Next, as a second embodiment of the present invention, FIG.
The processing operation for registering new document data in the stored document database 2 in the document search device will be described with reference to the flowchart of FIG.

【0037】まず、新規文書入力部7から新規の文書デ
ータを入力し、言語判定部8へ送る(S401)。言語
判定部8では、図10に示されるように、入力された新
規の文書データを構成する文字の特徴や構文特徴を利用
して該文書データの記述言語を判定し、図10に示され
るように入力された新規の文書データに記述言語の判定
結果を添えたデータを文書特徴ベクトル生成部9へ送る
(S402)。
First, new document data is input from the new document input unit 7 and sent to the language determination unit 8 (S401). As shown in FIG. 10, the language determination unit 8 determines the description language of the input document data by using the character features and syntax features of the new document data, as shown in FIG. Is sent to the document feature vector generation unit 9 (S402).

【0038】言語判定部8による言語判定の方法は、単
純に文字だけで判断する方法でもよいし、言語A、B、
C、D、…について単語辞書データベース1を用いた解
析処理を入力された新規文書データに対して行ない、そ
の解析可能性の度合いによって判断する方法であっても
よいし、各言語で記述された文書データにおいて最も頻
繁に出現する特徴的ないくつかの単語を該入力文書デー
タにおいて検索して判断する方法であってもよい。
The method of language determination by the language determination unit 8 may be a method of simply determining only characters, or a method of determining languages A, B,
For C, D,..., An analysis process using the word dictionary database 1 may be performed on the input new document data, and a determination may be made based on the degree of the analysis possibility. A method may be used in which some characteristic words appearing most frequently in the document data are searched for and determined in the input document data.

【0039】図9に戻って、文書特徴ベクトル生成部9
は、図11に示されるように単語辞書データベース1に
保持する単語データWiのうち、該入力文書データSD
の記述言語に相当する単語データWiを参照しながら、
該入力文書データSDを形態素解析して、該入力文書デ
ータSDに含まれる各単語Wiの単語特徴ベクトルVi
を抽出し、これらの単語特徴ベクトルViの総和を正規
化した単位ベクトルを該入力文書データSDの特徴ベク
トルVDとして該文書データSDと対応づけ(S40
3)、蓄積文書データベース2に格納し(S404)、
一連の処理を終了する。
Returning to FIG. 9, the document feature vector generation unit 9
Among the word data Wi held in the word dictionary database 1 as shown in FIG.
While referring to word data Wi corresponding to the description language of
The input document data SD is subjected to morphological analysis to obtain a word feature vector Vi of each word Wi included in the input document data SD.
Is extracted, and a unit vector obtained by normalizing the sum of the word feature vectors Vi is associated with the document data SD as a feature vector VD of the input document data SD (S40).
3), stored in the stored document database 2 (S404),
A series of processing ends.

【0040】上述した文書データの登録方法によれば、
記述言語の異なる文書データを手動で分類する必要もな
く、蓄積文書データベース2に蓄積文書データとして一
元的に登録し保存することができる。また、文書データ
に対して自動的に文書特徴ベクトルが計算できることか
ら、文書データ入力と上述した検索処理の実行を連続し
て行なうことも可能であり、検索の対象を予め蓄積され
ていなかった文書データにまで拡大することが容易に可
能となる。
According to the document data registration method described above,
There is no need to manually classify document data having different description languages, and the data can be registered and stored as stored document data in the stored document database 2 centrally. Further, since the document feature vector can be automatically calculated for the document data, it is possible to continuously perform the input of the document data and the execution of the above-described search processing. It can be easily expanded to data.

【0041】(実施の形態3)図12はこの発明の実施
の形態3の検索して得られた文書データの翻訳処理動作
に必要な図1の文書検索装置の部分構成図であり、図1
3はこの発明の実施の形態3の検索して得られた文書デ
ータの翻訳処理動作のフローチャートである。
(Embodiment 3) FIG. 12 is a partial block diagram of the document search apparatus of FIG. 1 required for the operation of translating the document data obtained by the search according to Embodiment 3 of the present invention.
FIG. 3 is a flowchart of a translation processing operation of document data obtained by the search according to the third embodiment of the present invention.

【0042】次に、この発明の実施の形態3として図1
の文書検索装置において検索して得られた文書データを
翻訳して提示する処理動作を図13のフローチャートに
従い説明する。
Next, a third embodiment of the present invention will be described with reference to FIG.
A processing operation of translating and presenting document data obtained by searching in the document search device of the present embodiment will be described with reference to the flowchart of FIG.

【0043】まず、図13のステップS501からS5
03までの手順は、上述した実施の形態1における図5
のステップS301からS303までの手順に準ずるの
で、詳細説明を省略する。なお、図5のステップS30
3においては、検索して得られた文書データは直接表示
部6に送られたが、図13においては一旦、言語判定部
8に送られる。
First, steps S501 to S5 in FIG.
The procedures up to 03 are the same as those in FIG.
Since the procedure from step S301 to step S303 is followed, detailed description is omitted. Note that step S30 in FIG.
In FIG. 3, the document data obtained by the search is sent directly to the display unit 6, but in FIG.

【0044】言語判定部8では、上述した実施の形態2
の図9のステップS402で説明したのと全く同様の方
法により、検索部5から送られてきた文書データを構成
する文字の特徴や構文特徴を利用して文書データの記述
言語を判定し、該文書データとともに判定結果を翻訳処
理部11に送る(S504)。このとき、翻訳処理部1
1に送られるデータの形式は、図10に示されている新
規文書データとその記述言語が併記されているデータの
それと同じである。
In the language determination section 8, the second embodiment is used.
In exactly the same manner as described in step S402 of FIG. 9, the description language of the document data is determined using the character features and syntax features of the document data sent from the search unit 5 and The determination result is sent to the translation processing unit 11 together with the document data (S504). At this time, the translation processing unit 1
The format of the data sent to 1 is the same as the format of the new document data shown in FIG.

【0045】翻訳処理部11では、検索要求SRが形態
素解析されて言語Aで記述されていることがわかってい
る場合、まず文書データの記述言語を確認し、文書デー
タが言語A以外の言語で記述されていれば、言語Aの翻
訳処理が必要と判断し(S505)、翻訳辞書データベ
ース10を参照しながら翻訳処理を行ない(S50
6)、翻訳結果を表示部6に送って表示し(S50
7)、一連の処理を終了する。
When it is known that the search request SR has been morphologically analyzed and described in the language A, the translation processing unit 11 first checks the description language of the document data, and converts the document data into a language other than the language A. If it is described, it is determined that the translation processing of the language A is necessary (S505), and the translation processing is performed with reference to the translation dictionary database 10 (S50).
6) The translation result is sent to the display unit 6 and displayed (S50).
7), a series of processing ends.

【0046】一方、文書データが検索要求SRの記述言
語Aで記述されていれば翻訳処理は不要と判断され(S
505)、言語判定部8から送られてきた文書データを
そのまま表示部6に送って表示し(S508)、一連の
処理を終了する。
On the other hand, if the document data is described in the description language A of the search request SR, it is determined that the translation process is unnecessary (S
505), the document data sent from the language determination unit 8 is sent to the display unit 6 as it is and displayed (S508), and a series of processing ends.

【0047】上述した文書データの翻訳処理の方法によ
れば、必要な情報を含む文書データがいかなる言語で記
述されていても、操作者は理解できる単一の言語、すな
わち検索要求SRの記述言語による表現によって必要な
情報を入手することができる。
According to the above-described document data translation method, even if the document data containing necessary information is described in any language, a single language that the operator can understand, that is, the description language of the search request SR The necessary information can be obtained by the expression of.

【図面の簡単な説明】[Brief description of the drawings]

【図1】この発明の実施の形態1〜3に適用される文書
検索装置のブロック構成図である。
FIG. 1 is a block diagram of a document search device applied to Embodiments 1 to 3 of the present invention.

【図2】図1の単語辞書データベース1の構成例を示す
図である。
FIG. 2 is a diagram showing a configuration example of a word dictionary database 1 of FIG.

【図3】図1の蓄積文書データベース2の構成例を示す
図である。
FIG. 3 is a diagram showing a configuration example of a stored document database 2 of FIG. 1;

【図4】この発明の実施の形態1の文書データ検索処理
動作に必要な図1の文書検索装置の部分構成図である。
FIG. 4 is a partial configuration diagram of the document search device of FIG. 1 necessary for a document data search processing operation according to the first embodiment of the present invention;

【図5】この発明の実施の形態1の文書データ検索処理
動作のフローチャートである。
FIG. 5 is a flowchart of a document data search processing operation according to the first embodiment of the present invention.

【図6】この発明の実施の形態1の文書データ検索処理
動作において検索要求から検索要求特徴ベクトルを得る
手順を説明する図である。
FIG. 6 is a diagram illustrating a procedure for obtaining a search request feature vector from a search request in the document data search processing operation according to the first embodiment of the present invention.

【図7】この発明の実施の形態1の文書データ検索処理
動作における検索要求に近似の文書データを得る手順を
説明する図である。
FIG. 7 is a diagram illustrating a procedure for obtaining document data approximate to a search request in the document data search processing operation according to the first embodiment of the present invention.

【図8】この発明の実施の形態2の文書データ登録処理
動作に必要な図1の文書検索装置の部分構成図である。
FIG. 8 is a partial configuration diagram of the document search device of FIG. 1 required for a document data registration processing operation according to the second embodiment of the present invention.

【図9】この発明の実施の形態2の文書データ登録処理
動作のフローチャートである。
FIG. 9 is a flowchart of a document data registration processing operation according to the second embodiment of the present invention.

【図10】この発明の実施の形態2の文書データ登録処
理動作において言語判定部で作成されるデータの説明図
である。
FIG. 10 is an explanatory diagram of data created by a language determination unit in a document data registration processing operation according to the second embodiment of the present invention.

【図11】この発明の実施の形態2の文書データ登録処
理動作において新規文書データを文書特徴ベクトルを付
与して蓄積文書データベースに格納する動作の説明図で
ある。
FIG. 11 is an explanatory diagram of an operation of adding new document data to a new document data and storing it in a stored document database in a document data registration operation according to the second embodiment of the present invention;

【図12】この発明の実施の形態3の検索して得られた
文書データの翻訳処理動作に必要な図1の文書検索装置
の部分構成図である。
FIG. 12 is a partial configuration diagram of the document search device in FIG. 1 required for a translation operation of document data obtained by searching according to the third embodiment of the present invention.

【図13】この発明の実施の形態3の検索して得られた
文書データの翻訳処理動作のフローチャートである。
FIG. 13 is a flowchart of a translation processing operation of document data obtained by searching according to the third embodiment of the present invention.

【符号の説明】[Explanation of symbols]

1 単語辞書データベース 2 蓄積文書データベース 3 検索要求入力部 4 検索要求特徴ベクトル生成部 5 検索部 7 新規文書入力部 8 言語判定部 9 文書特徴ベクトル生成部 10 翻訳辞書データベース 11 翻訳処理部 Vi 単語特徴ベクトル Wi 単語データ SDi 文書データ VDi 文書特徴ベクトル SR 検索要求 VS 検索要求特徴ベクトル Ai 検索要求SRと文書データSDiとの間の意味的
近似度 ただし、i=1、2、3、… なお、各図中同一符号は同一または相当部分を示す。
REFERENCE SIGNS LIST 1 word dictionary database 2 stored document database 3 search request input unit 4 search request feature vector generation unit 5 search unit 7 new document input unit 8 language determination unit 9 document feature vector generation unit 10 translation dictionary database 11 translation processing unit Vi word feature vector Wi word data SDi document data VDi document feature vector SR search request VS search request feature vector Ai Semantic similarity between search request SR and document data SDi where i = 1, 2, 3,... The same reference numerals indicate the same or corresponding parts.

───────────────────────────────────────────────────── フロントページの続き (51)Int.Cl.6 識別記号 庁内整理番号 FI 技術表示箇所 G06F 15/403 330C ──────────────────────────────────────────────────続 き Continued on the front page (51) Int.Cl. 6 Identification code Agency reference number FI Technical display location G06F 15/403 330C

Claims (5)

【特許請求の範囲】[Claims] 【請求項1】 それぞれが異なる言語で記述された複数
の文書データと、前記複数文書データのそれぞれに対応
してその意味的特徴を示す文書特徴データとが格納され
た文書データ蓄積部と、 所定の言語で記述された検索要求を入力するための要求
入力部と、 前記要求入力部により入力された前記検索要求の意味的
特徴を示す要求特徴データを検出する要求特徴データ検
出部と、 前記要求特徴データ検出部により検出された前記要求特
徴データに基づいて、 前記文書データ蓄積部の前記複数文書データから前記要
求入力部により入力された前記検出要求に対して意味的
に近似する意味的近似文書データを検索し出力する検索
部とを備えた、文書検索装置。
A document data storage unit that stores a plurality of document data, each of which is described in a different language, and document feature data corresponding to each of the plurality of document data and indicating a semantic feature thereof; A request input unit for inputting a search request described in a language of the following; a request characteristic data detection unit for detecting required characteristic data indicating a semantic characteristic of the search request input by the request input unit; A semantic approximation document that semantically approximates the detection request input by the request input unit from the plurality of document data in the document data storage unit based on the required characteristic data detected by the characteristic data detection unit. A document search device comprising: a search unit for searching and outputting data.
【請求項2】 異なる言語のそれぞれについての複数の
単語と、各単語に対応してその意味的特徴を示す単語特
徴ベクトルデータとを格納した単語辞書部をさらに備
え、 前記文書特徴データおよび前記要求特徴データのそれぞ
れは、前記文書データおよび前記検索要求のそれぞれを
構成する各単語に対応の前記単語辞書部中の前記単語特
徴ベクトルデータの総和を正規化した単位ベクトルデー
タであり、 前記意味的近似文書データは、 前記文書データ蓄積部の複数文書データのうち、対応す
る前記単位ベクトルデータと前記検索要求に対応の前記
単位ベクトルデータとの内積値が大きい文書データであ
ることを特徴とする、請求項1に記載の文書検索装置。
2. The apparatus according to claim 1, further comprising: a word dictionary unit storing a plurality of words for each of different languages and word feature vector data indicating a semantic feature corresponding to each word; Each of the feature data is unit vector data obtained by normalizing a sum of the word feature vector data in the word dictionary unit corresponding to each word constituting each of the document data and the search request, and the semantic approximation. The document data is document data having a large inner product value between the corresponding unit vector data and the unit vector data corresponding to the search request among the plurality of document data in the document data storage unit. Item 2. The document search device according to Item 1.
【請求項3】 前記単語辞書部では、異なる言語間の同
義である単語のすべては、同一の前記単語特徴ベクトル
データに対応づけられることを特徴とする、請求項2に
記載の文書検索装置。
3. The document search device according to claim 2, wherein in the word dictionary section, all of the words having the same meaning in different languages are associated with the same word feature vector data.
【請求項4】 新規の文書データを入力するための文書
データ入力部と、 前記文書データ入力部により入力された前記文書データ
の前記文書特徴データを生成する文書特徴データ生成部
とをさらに備え、 前記文書データ入力部により入力された前記文書データ
は前記文書特徴データ生成部により生成された前記文書
特徴データと対応づけられて前記文書データ蓄積部に格
納されることを特徴とする、請求項1ないし3のいずれ
かに記載の文書検索装置。
4. A document data input unit for inputting new document data, and a document characteristic data generating unit for generating the document characteristic data of the document data input by the document data input unit, The document data input by the document data input unit is stored in the document data storage unit in association with the document feature data generated by the document feature data generation unit. 4. The document search device according to any one of claims 1 to 3.
【請求項5】 異なる言語間の翻訳に必要なデータを保
持する翻訳辞書部と、 前記検索部による検索出力時、前記意味的近似文書デー
タの記述言語が前記検索要求を記述する所定言語に一致
しない場合に、前記翻訳辞書部の内容を参照して前記意
味的近似文書データを前記所定言語に翻訳する翻訳処理
部とをさらに備えた、請求項1ないし4のいずれかに記
載の文書検索装置。
5. A translation dictionary unit for holding data necessary for translation between different languages, and a description language of the semantically approximate document data matches a predetermined language describing the search request when the search unit outputs a search. 5. The document search device according to claim 1, further comprising: a translation processing unit that translates the semantically approximated document data into the predetermined language by referring to the contents of the translation dictionary unit when not. .
JP8185018A 1996-07-15 1996-07-15 Document search device Withdrawn JPH1031677A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP8185018A JPH1031677A (en) 1996-07-15 1996-07-15 Document search device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP8185018A JPH1031677A (en) 1996-07-15 1996-07-15 Document search device

Publications (1)

Publication Number Publication Date
JPH1031677A true JPH1031677A (en) 1998-02-03

Family

ID=16163338

Family Applications (1)

Application Number Title Priority Date Filing Date
JP8185018A Withdrawn JPH1031677A (en) 1996-07-15 1996-07-15 Document search device

Country Status (1)

Country Link
JP (1) JPH1031677A (en)

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR100532585B1 (en) * 2000-12-30 2005-12-02 한국전자통신연구원 Construction of Knowledge Base for Question/Answering on Internet
KR100903599B1 (en) 2007-11-22 2009-06-18 한국전자통신연구원 Encrypted Data Retrieval Method using Inner Product and Terminal Device and Server for It
KR101442719B1 (en) * 2013-04-16 2014-09-19 한양대학교 에리카산학협력단 Apparatus and method for recommendation of academic paper
JP2018010482A (en) * 2016-07-13 2018-01-18 日本電信電話株式会社 Document concept base generation device, document concept search device, method, and program
CN111339261A (en) * 2020-03-17 2020-06-26 北京香侬慧语科技有限责任公司 Document extraction method and system based on pre-training model
JP2021157363A (en) * 2020-03-26 2021-10-07 株式会社野村総合研究所 Needs matching device and program

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR100532585B1 (en) * 2000-12-30 2005-12-02 한국전자통신연구원 Construction of Knowledge Base for Question/Answering on Internet
KR100903599B1 (en) 2007-11-22 2009-06-18 한국전자통신연구원 Encrypted Data Retrieval Method using Inner Product and Terminal Device and Server for It
KR101442719B1 (en) * 2013-04-16 2014-09-19 한양대학교 에리카산학협력단 Apparatus and method for recommendation of academic paper
JP2018010482A (en) * 2016-07-13 2018-01-18 日本電信電話株式会社 Document concept base generation device, document concept search device, method, and program
CN111339261A (en) * 2020-03-17 2020-06-26 北京香侬慧语科技有限责任公司 Document extraction method and system based on pre-training model
JP2021157363A (en) * 2020-03-26 2021-10-07 株式会社野村総合研究所 Needs matching device and program

Similar Documents

Publication Publication Date Title
US8185372B2 (en) Apparatus, method and computer program product for translating speech input using example
JP2002278964A (en) Translation support apparatus, method and translation support program
CN111259262A (en) Information retrieval method, device, equipment and medium
JP2021144348A (en) Information processing device and information processing method
CN114141384A (en) Method, apparatus and medium for retrieving medical data
JPH1031677A (en) Document search device
US11842165B2 (en) Context-based image tag translation
JP2002342361A (en) Information retrieval device
JPH05324719A (en) Document retrieval system
CN119621944A (en) Data retrieval method, device, electronic device and medium
US20230409620A1 (en) Non-transitory computer-readable recording medium storing information processing program, information processing method, information processing device, and information processing system
JP2000148754A (en) Multilingual system, multilingual processing method, and medium storing multilingual processing program
JP3315221B2 (en) Conversation sentence translator
JP4024137B2 (en) Quantity expression search device
JP4007630B2 (en) Bilingual example sentence registration device
JP2010009237A (en) Multi-language similar document retrieval device, method and program, and computer-readable recording medium
JPH08339376A (en) Foreign language search device and information search system
JPH0950435A (en) Translation device
JPH09245051A (en) Natural language case retrieval device and natural language case retrieval method
US20240265202A1 (en) Auto-suggestion with rich objects
JP2004334690A (en) Character data input / output device, character data input / output method, character data input / output program, and computer-readable recording medium
KR20110044697A (en) Transliteration methods and devices
JP2004280467A (en) TRANSLATION DEVICE, TRANSLATION METHOD, AND ITS PROGRAM
JP4054353B2 (en) Machine translation apparatus and machine translation program
CN114861664A (en) Chinese named entity recognition method, device, equipment and storage medium

Legal Events

Date Code Title Description
A300 Application deemed to be withdrawn because no request for examination was validly filed

Free format text: JAPANESE INTERMEDIATE CODE: A300

Effective date: 20031007