JPH07129588A - Pattern-guided document content analyzer - Google Patents
Pattern-guided document content analyzerInfo
- Publication number
- JPH07129588A JPH07129588A JP5279154A JP27915493A JPH07129588A JP H07129588 A JPH07129588 A JP H07129588A JP 5279154 A JP5279154 A JP 5279154A JP 27915493 A JP27915493 A JP 27915493A JP H07129588 A JPH07129588 A JP H07129588A
- Authority
- JP
- Japan
- Prior art keywords
- document content
- pattern
- document
- content pattern
- analysis
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Landscapes
- Machine Translation (AREA)
Abstract
Description
【0001】[0001]
【産業上の利用分野】本発明は、オフィスにおける文書
の内容を解析する文書内容解析装置に関する。BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a document content analysis device for analyzing the content of a document in an office.
【0002】[0002]
【従来の技術】従来は、文書内容を解析する場合、解析
に必要な辞書をあらかじめ用意し、その辞書の保持する
解析用の知識に応じて文書内容を解析していた。そのた
め、特定の文書に出現しない内容を辞書に知識として登
録していたり、必要な知識が辞書に記述されていない場
合があった。2. Description of the Related Art Conventionally, when analyzing document contents, a dictionary required for the analysis is prepared in advance, and the document contents are analyzed according to the analysis knowledge held by the dictionary. Therefore, there are cases where contents that do not appear in a specific document are registered as knowledge in the dictionary, or necessary knowledge is not described in the dictionary.
【0003】[0003]
【発明が解決しようとする課題】文書内容を解析する場
合、解析に必要な辞書をあらかじめ用意するのは、労力
が必要であり、さらに、必要十分な辞書知識を構築する
のが非常に難しい。When analyzing document contents, it is laborious to prepare a dictionary required for the analysis in advance, and it is very difficult to construct necessary and sufficient dictionary knowledge.
【0004】本発明の目的は、自然言語処理用の辞書を
用意しなくとも文書内容を解析することができる文書内
容解析装置を提供することにある。It is an object of the present invention to provide a document content analysis apparatus capable of analyzing document content without preparing a dictionary for natural language processing.
【0005】[0005]
【課題を解決するための手段】本発明の文書内容解析装
置は、対象とする文書内容の特徴を示す文書内容パター
ンを、該文書と同様の目的で記述された文書の内容から
抽出する文書内容パターン抽出手段と、抽出された文書
内容パターンの中で類似するパターンを抽出し、文書内
容パターンを圧縮する文書内容パターン圧縮手段と、該
文書内容パターン圧縮手段により圧縮された文書内容パ
ターンを用いて、新たに入力された文書内容の解析を行
なう文書内容解析手段を有する。A document content analyzing apparatus of the present invention extracts a document content pattern indicating a characteristic of a target document content from the content of a document described for the same purpose as the document content. Using a pattern extracting means, a document content pattern compressing means for extracting a similar pattern from the extracted document content patterns and compressing the document content pattern, and a document content pattern compressed by the document content pattern compressing means , And has document content analysis means for analyzing the newly input document content.
【0006】[0006]
【作用】文書内容を解析する場合、文書の解析に必要な
情報を適宜、対象となる文書と同じ目的で作成された文
書の内容から抽出し、抽出した対象文書特有の文章表現
を抽出化して保持し、その情報を利用して文書内容を解
析する。したがって、自然言語処理用の辞書を用意しな
くとも文書内容を解析することができる。When the document contents are analyzed, the information necessary for the document analysis is appropriately extracted from the contents of the document created for the same purpose as the target document, and the extracted text expression peculiar to the target document is extracted. Hold and use the information to analyze the document contents. Therefore, the document contents can be analyzed without preparing a dictionary for natural language processing.
【0007】[0007]
【実施例】次に、本発明の実施例について図面を参照し
て説明する。Embodiments of the present invention will now be described with reference to the drawings.
【0008】図1は本発明の一実施例のパターン誘導型
文書内容解析装置のブロック図、図2は文書内容パター
ン抽出手段1の動作を示すフローチャート、図3は文書
内容パターン圧縮手段2の動作を示すフローチャート、
図4は文書内容解析手段3の動作を示すフローチャート
である。FIG. 1 is a block diagram of a pattern-guided document content analyzing apparatus according to an embodiment of the present invention, FIG. 2 is a flow chart showing the operation of the document content pattern extracting means 1, and FIG. 3 is an operation of the document content pattern compressing means 2. A flow chart showing
FIG. 4 is a flowchart showing the operation of the document content analysis means 3.
【0009】文書内容パターン抽出手段1は、対象とな
る文書と同様の目的で記述された文書の内容4から文書
内容パターン5を抽出する。文書内容パターン圧縮手段
2は、抽出された文書内容パターン5の中で類似するパ
ターンを抽出し、文書内容パターン5を圧縮する。文書
内容解析手段3は、文書内容パターン圧縮手段2により
圧縮された文書内容パターン6を用いて、新たに入力さ
れた文書内容7の解析を行ない、解析結果8を出力す
る。The document content pattern extraction means 1 extracts the document content pattern 5 from the content 4 of the document described for the same purpose as the target document. The document content pattern compression unit 2 extracts a similar pattern from the extracted document content patterns 5 and compresses the document content pattern 5. The document content analysis unit 3 analyzes the newly input document content 7 using the document content pattern 6 compressed by the document content pattern compression unit 2 and outputs an analysis result 8.
【0010】次に、本実施例の動作を具体例により説明
する。Next, the operation of this embodiment will be described with reference to a concrete example.
【0011】ここでは、対象とする文書として、新聞記
事を考える。新聞記事は、一般対象向けに、企業で起こ
った出来事について編集・出版している。この場合、一
般大衆が知っているべき出来事を伝達するために記述さ
れているという点では、新聞記事の目的は、絶えず一定
であると考えてよい。Here, a newspaper article is considered as a target document. Newspaper articles are edited and published for the general public on what happened in the company. In this case, the purpose of the newspaper article may be considered to be constantly constant in that it is written to convey an event that the general public should know.
【0012】文書内容パターン抽出手段1では、まず、
文4を入力し(ステップ11)、入力された文4を文節
単位に分割して(ステップ12)、文書中に出現する文
節を固有名詞であるかそれ以外であるかに分けて処理が
行なわれる(ステップ13〜17)。文書中に出現する
文節が固有名詞の場合、その固有名詞の働きに応じて、
agent (主体)であるか、object(対象)であるかを区
別する(ステップ14)。通常、新聞記事などでは、ag
ent の働きをする場合、その固有名詞の後に括弧書き
で、主体に関係する情報を付加している。たとえば、企
業名であれば、社長が誰であり、本社の住所がどこで、
代表の電話番号が何番であるかの情報が括弧書きされて
いる。これらの情報より、括弧書きされた前の文節が固
有名詞であり、それがagent の働きをするということが
自動的に判断される。また、objectであれば、鍵括弧を
用いて固有名詞を修辞する。これらの修辞により、固有
名詞を自動的に判断する。固有名詞以外の場合には、各
文節の意味を最も適切に表現している単語を抽出する
(ステップ15)。In the document content pattern extraction means 1, first,
Sentence 4 is input (step 11), the input sentence 4 is divided into clauses (step 12), and the clauses appearing in the document are processed according to whether they are proper nouns or other proper nouns. (Steps 13 to 17). If the phrase appearing in the document is a proper noun, depending on the function of the proper noun,
A distinction is made between an agent (main body) and an object (target) (step 14). Usually, in newspaper articles, etc., ag
When acting as an ent, information related to the subject is added in brackets after the proper noun. For example, if it is a company name, who is the president, where is the address of the head office,
Information about what the representative telephone number is is written in brackets. From this information, it is automatically determined that the preceding clause in parentheses is a proper noun, which acts as an agent. If it is an object, the proper noun is rhetorical using brackets. Proper nouns are automatically judged by these rhetoric. If it is not a proper noun, the word that most appropriately expresses the meaning of each clause is extracted (step 15).
【0013】入力文4:A社(社長X、東京都Y市、電
話番号Z)が「B」の豊橋工場から前橋工場に移管す
る。Input sentence 4: Company A (President X, Y city, Tokyo, telephone number Z) transfers from "B" Toyohashi factory to Maebashi factory.
【0014】では、以下のような文節に分割される(文
節の区切りは“/”で表わす)。[0014] In the following, it is divided into the following clauses (separation of clauses is represented by "/").
【0015】A社(社長X、東京都Y市、電話番号Z)
が/「B」を/豊橋工場から/前橋工場に/移管する。Company A (President X, Y city, Tokyo, telephone number Z)
/ Transfer "B" from Toyohashi Plant to Maebashi Plant.
【0016】文節を区切るためには、形態素解析を用い
ることもできるが、文字種が平仮名から漢字又は記号に
変化した場所を文節区切りとすることもできる。文節区
切りが終了した後、個々の文節が固有名詞かどうかを判
断する(ステップ13)。この例では、“A社(社長
X、東京都Y市、電話番号Z)が”、“「B」を”が固
有名詞で、それぞれagent ,objectが付与される(ステ
ップ16,17)。一方、固有名詞以外は、“豊橋工場
から”、“前橋工場に”、“移管する。”の3文節であ
る。これらから、各文節の意味を最も適切に表わす単語
を抽出し、文節に付与する(ステップ15)。通常、文
節中の一番最後の単語と付属語が文節の意味を適切に表
わす単語であることが多い。この仮定を利用して、固有
名詞以外の文節も自動的に処理することができる。先の
入力文4であると、以下のような文書内容パターン5が
抽出される。Morphological analysis can be used to divide the bunsetsu, but the place where the character type is changed from hiragana to kanji or symbol can be used as the bunsetsu delimiter. After the phrase segmentation is completed, it is judged whether each segment is a proper noun (step 13). In this example, "A company (President X, Y city, Tokyo, telephone number Z)" is a proper noun, and "B" is a proper noun, and agent and object are respectively added (steps 16 and 17). , Except for proper nouns, "Toyohashi Factory", "Maebashi Factory", "Transfer". The three most appropriate words are extracted from these, and the word that most appropriately expresses the meaning of each word is extracted and assigned to the word (step 15). Usually, the last word in the word and the adjunct word mean the word. This assumption can be used to automatically process clauses other than proper nouns. With the input sentence 4 above, the following document content pattern 5 can be obtained. Is extracted.
【0017】抽出された文書内容パターン5:agent /
object/工場から/工場に/移管同様にして文書内容パ
ターン5を同様の新聞記事から抽出して、記憶する(ス
テップ18,19)。Extracted document content pattern 5: agent /
object / from factory / to factory / transfer In the same manner, the document content pattern 5 is extracted from a similar newspaper article and stored (steps 18 and 19).
【0018】文書内容パターン圧縮手段2は、文書内容
パターン抽出手段1で抽出された文書内容パターン5を
圧縮する(ステップ21)。圧縮処理では、agent ,ob
ject等が複数あった場合、それを1つとして処理する。
さらに、日本語特有の語順の自由度に対処するために、
出現する文節の順序に関係なく、出現した文節が同じで
あれば同一のパターンと判断し、圧縮する。さらに、省
略された文節がある場合には、パターンの他の部分が一
致していたら同一のパターンとして圧縮する。圧縮され
た文書内容パターン6は装置内に登録する(ステップ2
2)。ここで、同一と判断されるパターン例は次の3つ
である。The document content pattern compression means 2 compresses the document content pattern 5 extracted by the document content pattern extraction means 1 (step 21). In compression processing, agent, ob
If there are multiple jects, etc., they are treated as one.
Furthermore, in order to deal with the degree of freedom of word order peculiar to Japanese,
Regardless of the order of the appearing clauses, if the appearing clauses are the same, it is judged as the same pattern and compressed. Further, when there is an omitted clause, if the other parts of the pattern match, they are compressed as the same pattern. The compressed document content pattern 6 is registered in the device (step 2).
2). Here, the following three pattern examples are determined to be the same.
【0019】agent /agent /object/object/工場か
ら/工場に/移管 工場から/agent /object/工場に/移管 agent /object/工場に/移管 文書内容解析手段3では、文書内容パターン圧縮手段2
で圧縮された文書内容パターン6を用いて、文書内容を
解析する。まず、文7を入力し(ステップ31)、入力
された文7を文書内容パターン抽出手段1の処理と同じ
ように文節単位に分割し(ステップ32)、各文節を文
書内容パターン6とマッチングを取る(ステップ3
3)。マッチングでは、圧縮時と同様に出現する文節の
順序に関係なく、出現した文節が同じであれば一致した
と判断する。さらに、省略された文節がある場合には、
パターンの他の部分が一致していたら一致と判断する。Agent / agent / object / object / from factory / to factory / transfer from factory / agent / object / to factory / transfer agent / object / to factory / transfer In the document content analysis means 3, the document content pattern compression means 2
The document content is analyzed using the document content pattern 6 compressed in. First, the sentence 7 is input (step 31), and the input sentence 7 is divided into sentence units in the same manner as the processing of the document content pattern extraction means 1 (step 32), and each sentence is matched with the document content pattern 6. Take (Step 3
3). In the matching, as in the case of compression, regardless of the order of the appearing phrases, if the appearing phrases are the same, it is determined that they match. In addition, if there are omitted clauses,
If the other parts of the pattern match, it is determined that they match.
【0020】入力文7:大型橋梁用の三橋工場に移管し
た。Input sentence 7: Transferred to the Mihashi factory for large bridges.
【0021】入力文7を文節単位に分割:大型橋梁用の
/三橋工場に/移管した。The input sentence 7 is divided into clause units: for a large bridge / to / transferred to the Mitsuhashi factory.
【0022】マッチングした文書内容パターン6:agen
t /object/工場から/工場に/移管 マッチング部分:三橋工場に=工場に 移管した=移管 文書内容解析手段3は、文書内容パターン6にマッチン
グした部分を文書内容の解析結果8として最終的に出力
する(ステップ34)。Matched document content pattern 6: agen
t / object / from factory / to / factory / matching part: to Mitsuhashi factory = transferred to factory = transferred The document content analysis means 3 finally sets the part matching the document content pattern 6 as the analysis result 8 of the document content. Output (step 34).
【0023】出力(解析結果8):三橋工場に移管し
た。Output (analysis result 8): Transferred to the Mitsuhashi factory.
【0024】[0024]
【発明の効果】以上説明したように、本発明は、対象と
なる文書内容の特徴を示す文書内容パターンを、対象と
なる文書と同様の目的で記述された文書の内容から抽出
し、類似する文書内容パターンを圧縮して装置内に登録
し、該登録された文書内容パターンを用いて新たに入力
された文書内容の解析を行なうことにより、大規模な辞
書を人手で作成することなく文書内容を自動的に解析す
ることが可能となる効果がある。As described above, according to the present invention, the document content pattern indicating the characteristics of the target document content is extracted from the content of the document described for the same purpose as the target document and is similar. By compressing the document content pattern and registering it in the device, and analyzing the newly input document content using the registered document content pattern, the document content can be stored without manually creating a large-scale dictionary. There is an effect that it is possible to analyze automatically.
【図1】本発明の一実施例のパターン誘導型文書内容解
析装置のブロック図である。FIG. 1 is a block diagram of a pattern-guided document content analysis apparatus according to an embodiment of the present invention.
【図2】文書内容パターン抽出手段1の動作を示すフロ
ーチャートである。FIG. 2 is a flowchart showing the operation of the document content pattern extraction means 1.
【図3】文書内容パターン圧縮手段2の動作を示すフロ
ーチャートである。FIG. 3 is a flowchart showing the operation of the document content pattern compression unit 2.
【図4】文書内容解析手段3の動作を示すフローチャー
トである。FIG. 4 is a flowchart showing the operation of the document content analysis means 3.
1 文書内容パターン抽出手段 2 文書内容パターン圧縮手段 3 文書内容解析手段 4 入力文 5 文書内容パターン 6 圧縮された文書内容パターン 7 入力文 8 解析結果 11〜19,21,22,31〜34 ステップ 1 Document Content Pattern Extracting Means 2 Document Content Pattern Compressing Means 3 Document Content Analyzing Means 4 Input Statements 5 Document Content Patterns 6 Compressed Document Content Patterns 7 Input Statements 8 Analysis Results 11-19, 21, 22, 31-34 Steps
Claims (2)
容パターンを、該文書と同様の目的で記述された文書の
内容から抽出する文書内容パターン抽出手段と、 抽出された文書内容パターンの中で類似するパターンを
抽出し、文書内容パターンを圧縮する文書内容パターン
圧縮手段と、 該文書内容パターン圧縮手段により圧縮された文書内容
パターンを用いて、新たに入力された文書内容の解析を
行なう文書内容解析手段を有するパターン誘導型文書内
容解析装置。1. A document content pattern extracting means for extracting a document content pattern indicating a characteristic of a target document content from the content of a document described for the same purpose as the document, and among the extracted document content patterns. A document for analyzing a newly input document content by using a document content pattern compression unit for extracting a similar pattern by using the document content pattern compression unit and a document content pattern compressed by the document content pattern compression unit A pattern-guided document content analysis device having content analysis means.
された文を文節単位に分割し、分割された各文節につい
て、当該文節の働きに応じて当該文節から当該文節の意
味を表わす単語を抽出して当該文節に付与し、前記文書
内容解析手段は、入力された文を文節単位に分割し、分
割された各文節を前記文書内容パターン圧縮手段で圧縮
された文書内容パターンとマッチングを取り、マッチン
グした部分を解析結果として出力する、請求項1記載の
文書内容解析装置。2. The document content pattern extraction means divides an input sentence into phrases, and for each of the divided phrases, extracts a word representing the meaning of the phrase from the phrase according to the function of the phrase. Then, the document content analysis unit divides the input sentence into units of phrases, and matches each divided phrase with the document content pattern compressed by the document content pattern compression unit, The document content analysis apparatus according to claim 1, wherein the matched portion is output as an analysis result.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP5279154A JPH07129588A (en) | 1993-11-09 | 1993-11-09 | Pattern-guided document content analyzer |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP5279154A JPH07129588A (en) | 1993-11-09 | 1993-11-09 | Pattern-guided document content analyzer |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH07129588A true JPH07129588A (en) | 1995-05-19 |
Family
ID=17607209
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP5279154A Pending JPH07129588A (en) | 1993-11-09 | 1993-11-09 | Pattern-guided document content analyzer |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH07129588A (en) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP3276507A1 (en) | 2016-07-25 | 2018-01-31 | Fujitsu Limited | Encoding device, encoding method and search method |
-
1993
- 1993-11-09 JP JP5279154A patent/JPH07129588A/en active Pending
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP3276507A1 (en) | 2016-07-25 | 2018-01-31 | Fujitsu Limited | Encoding device, encoding method and search method |
| US9906238B2 (en) | 2016-07-25 | 2018-02-27 | Fujitsu Limited | Encoding device, encoding method and search method |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO1997004405A9 (en) | Method and apparatus for automated search and retrieval processing | |
| EP0839357A1 (en) | Method and apparatus for automated search and retrieval processing | |
| JPH07129588A (en) | Pattern-guided document content analyzer | |
| JPH0682376B2 (en) | Emotion information extraction device | |
| JPH0619968A (en) | Automatic extraction device for technical term | |
| JPH03105465A (en) | Compound word extraction device | |
| JP2812511B2 (en) | Keyword extraction device | |
| JPS6368972A (en) | Unregistered word processing system | |
| JP2003223441A (en) | Character string shaping method, device, and program | |
| JP3216725B2 (en) | Sentence structure analyzer | |
| JPH05233689A (en) | Automatic document summarization method | |
| JPH0244463A (en) | Original text input method in machine translation system | |
| JP3466669B2 (en) | Character processing method | |
| JPH05250403A (en) | Japanese sentence word analyzing system | |
| CN114398880A (en) | System and method for optimizing Chinese word segmentation | |
| JP2874378B2 (en) | Sentence analyzer | |
| JPH0262659A (en) | Extracting device for correction candidate character of japanese sentence | |
| JP2001022752A (en) | Method and device for character group extraction, and recording medium for character group extraction | |
| Cowie | CRL’s Approach to MET | |
| JPH06301715A (en) | Morpheme analyzing device | |
| JPH05233686A (en) | Japanese language processor | |
| JPH03127173A (en) | Japanese morpheme analyzing method | |
| JP2002269084A (en) | Morpheme conversion rule generating device and morpheme string converting device | |
| JPH01144162A (en) | System and device for morpheme analysis using key word | |
| JPS6316370A (en) | Word extracting system |