EP0968478A1 - Procede de generation automatique d'un resume d'un texte par un ordinateur - Google Patents
Procede de generation automatique d'un resume d'un texte par un ordinateurInfo
- Publication number
- EP0968478A1 EP0968478A1 EP98914784A EP98914784A EP0968478A1 EP 0968478 A1 EP0968478 A1 EP 0968478A1 EP 98914784 A EP98914784 A EP 98914784A EP 98914784 A EP98914784 A EP 98914784A EP 0968478 A1 EP0968478 A1 EP 0968478A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- sentence
- text
- word
- words
- probability
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Withdrawn
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/34—Browsing; Visualisation therefor
- G06F16/345—Summarisation for human users
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/205—Parsing
- G06F40/216—Parsing using statistical methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/279—Recognition of textual entities
- G06F40/284—Lexical analysis, e.g. tokenisation or collocates
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
- G06F40/35—Discourse or dialogue representation
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y10—TECHNICAL SUBJECTS COVERED BY FORMER USPC
- Y10S—TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y10S707/00—Data processing: database and file management or data structures
- Y10S707/99931—Database or file accessing
- Y10S707/99933—Query processing, i.e. searching
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y10—TECHNICAL SUBJECTS COVERED BY FORMER USPC
- Y10S—TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y10S707/00—Data processing: database and file management or data structures
- Y10S707/99931—Database or file accessing
- Y10S707/99933—Query processing, i.e. searching
- Y10S707/99934—Query formulation, input preparation, or translation
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y10—TECHNICAL SUBJECTS COVERED BY FORMER USPC
- Y10S—TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y10S707/00—Data processing: database and file management or data structures
- Y10S707/99931—Database or file accessing
- Y10S707/99933—Query processing, i.e. searching
- Y10S707/99935—Query augmenting and refining, e.g. inexact access
Definitions
- the invention relates to a method for the automatic generation of a summary of a text by a computer.
- a special type of information reduction consists in the merging of texts.
- a method for summarizing texts is known from [1] which uses heuristic features with a discrete range of values.
- the probability that a sentence from the text belongs to the summary on the condition that a heuristic feature has a certain value is estimated from a training set of summaries.
- the object of the invention is to automatically generate a summary from a given text, which summary is intended to represent the essential contents of the text in short form.
- the method according to the invention enables a text to be summarized by determining for each sentence of this text a probability that the sentence belongs to the summary.
- the relevance measure is determined for each word m in the sentence from a lexicon which contains all relevant words with a predefined relevance measure for each of these words.
- the accumulation of all relevance measures gives the probability of the sentence belonging to the summary. All records are then sorted according to their probability.
- a predeterminable reduction measure which indicates what percentage of the original text is shown in the summary, serves for the selection of the number of sentences given by this reduction measure from the sorted representation. If the most important x-percent sentences are selected, they are displayed as a summary of the text in its original order given by this text.
- An advantageous further development of the method according to the invention consists in introducing an frequency of Emzelworth in addition to the relevance measure. This level of detail indicates how often the word in question appears in the entire text to be summarized. Taking into account the relevance measure and this newly introduced
- N is the total number of words in the
- a further development of the method according to the invention consists in using an application-specific lexicon.
- a lexicon specified for sports contributions will rate sports-related words with a higher relevance for a text to be summarized than a lexicon that specializes in summaries of economic contributions. It is therefore advantageously possible to provide specific knowledge about predefinable categories by means of lexica corresponding to the respective categories.
- a text is also advantageous to assign a text to one or more categories. This can be done automatically by using specific, predefinable words in the subject-related lexica as a selection criterion for an assignment to the respective subject area. If several categories (subject areas), i.e. different perspectives or filters, are possible for the summary of a text, different summaries, one for each category, can be created automatically.
- FIG. 1 is a sketch illustrating a system for automatically generating a summary
- Fig. 2 is a block diagram illustrating the steps of the method according to the invention.
- FIG. 1 shows a system with which an automatic generation of a summary of text by a
- a text to be summarized can either be written TXT, e.g. on paper, or in digital form DIGTXT, e.g. as the result of a database query.
- the text TXT is read in by the scanner SC and stored as an image file BD.
- a text recognition software OCR converts the text TXT m present as an image file BD into a machine-readable format, e.g. ASCII format to.
- the digital text DIGTXT is already available in machine-readable format.
- the summary according to the invention is created using the corresponding lexicon (in the KatSel block).
- step 2a the first sentence is selected at the beginning of the method according to the invention and the probability that this sentence belongs to the summary is set to 0.
- step 2b the first word of this sentence is selected. Since the probability that this sentence belongs to the summary is derived from the
- step 2e If the probabilities of the individual words are put together, for each word in the sentence in the loop from step 2c to step 2e, the respective probability is cumulated to the overall probability for the entire sentence. Once all the words in the sentence have been processed, the probability for the individual sentence is normalized by the number of words. The steps described are carried out for all sentences in the text (step 2g, 2h, 2 ⁇ ). If the last sentence in the text has been processed, the sentences are after their
- step 2j Probability sorted (step 2j). According to a predeterminable reduction measure, the n best sentences corresponding to the reduction measure are selected in step 2k and then their original sequence is displayed in step 2m.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- Artificial Intelligence (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Computational Linguistics (AREA)
- General Health & Medical Sciences (AREA)
- Health & Medical Sciences (AREA)
- Probability & Statistics with Applications (AREA)
- Data Mining & Analysis (AREA)
- Databases & Information Systems (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Document Processing Apparatus (AREA)
Abstract
Le procédé selon l'invention permet la rédaction automatique, fondée sur les phrases, d'un texte, sur un ordinateur. A cet effet, sont utilisés des lexiques thématiques qui confèrent à chacun des mots qu'ils contiennent une grandeur de pertinence. Chaque phrase ou texte à résumer est traité mot par mot et, pour chaque mot, une fréquence de mot individuel est cumulée tout en étant pondérée avec la grandeur de pertinence. Pour la rédaction du résumé, sont assemblées les n phrases présentant la plus grande probabilité d'appartenir au résumé, n étant une grandeur de réduction prédéfinissable.
Applications Claiming Priority (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| DE19711284 | 1997-03-18 | ||
| DE19711284 | 1997-03-18 | ||
| PCT/DE1998/000485 WO1998041930A1 (fr) | 1997-03-18 | 1998-02-18 | Procede de generation automatique d'un resume d'un texte par un ordinateur |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP0968478A1 true EP0968478A1 (fr) | 2000-01-05 |
Family
ID=7823794
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP98914784A Withdrawn EP0968478A1 (fr) | 1997-03-18 | 1998-02-18 | Procede de generation automatique d'un resume d'un texte par un ordinateur |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US6401086B1 (fr) |
| EP (1) | EP0968478A1 (fr) |
| JP (1) | JP2001515623A (fr) |
| WO (1) | WO1998041930A1 (fr) |
Families Citing this family (46)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6789230B2 (en) * | 1998-10-09 | 2004-09-07 | Microsoft Corporation | Creating a summary having sentences with the highest weight, and lowest length |
| US7475334B1 (en) * | 2000-01-19 | 2009-01-06 | Alcatel-Lucent Usa Inc. | Method and system for abstracting electronic documents |
| JP2004514350A (ja) * | 2000-11-14 | 2004-05-13 | コーニンクレッカ フィリップス エレクトロニクス エヌ ヴィ | 番組の要約と索引付け |
| WO2003012661A1 (fr) * | 2001-07-31 | 2003-02-13 | Invention Machine Corporation | Recapitulation informatique de documents en langage naturel |
| US8799776B2 (en) * | 2001-07-31 | 2014-08-05 | Invention Machine Corporation | Semantic processor for recognition of whole-part relations in natural language documents |
| US9009590B2 (en) * | 2001-07-31 | 2015-04-14 | Invention Machines Corporation | Semantic processor for recognition of cause-effect relations in natural language documents |
| US6904564B1 (en) * | 2002-01-14 | 2005-06-07 | The United States Of America As Represented By The National Security Agency | Method of summarizing text using just the text |
| US7549114B2 (en) * | 2002-02-21 | 2009-06-16 | Xerox Corporation | Methods and systems for incrementally changing text representation |
| US7650562B2 (en) * | 2002-02-21 | 2010-01-19 | Xerox Corporation | Methods and systems for incrementally changing text representation |
| US20040199408A1 (en) * | 2003-04-01 | 2004-10-07 | Johnson Tolbert R. | Medical information card |
| US9275052B2 (en) | 2005-01-19 | 2016-03-01 | Amazon Technologies, Inc. | Providing annotations of a digital work |
| US8131647B2 (en) | 2005-01-19 | 2012-03-06 | Amazon Technologies, Inc. | Method and system for providing annotations of a digital work |
| US8234279B2 (en) * | 2005-10-11 | 2012-07-31 | The Boeing Company | Streaming text data mining method and apparatus using multidimensional subspaces |
| US7752204B2 (en) * | 2005-11-18 | 2010-07-06 | The Boeing Company | Query-based text summarization |
| US7831597B2 (en) * | 2005-11-18 | 2010-11-09 | The Boeing Company | Text summarization method and apparatus using a multidimensional subspace |
| US8352449B1 (en) | 2006-03-29 | 2013-01-08 | Amazon Technologies, Inc. | Reader device content indexing |
| US20080005284A1 (en) * | 2006-06-29 | 2008-01-03 | The Trustees Of The University Of Pennsylvania | Method and Apparatus For Publishing Textual Information To A Web Page |
| US8725565B1 (en) | 2006-09-29 | 2014-05-13 | Amazon Technologies, Inc. | Expedited acquisition of a digital item following a sample presentation of the item |
| US9672533B1 (en) | 2006-09-29 | 2017-06-06 | Amazon Technologies, Inc. | Acquisition of an item based on a catalog presentation of items |
| US7865817B2 (en) | 2006-12-29 | 2011-01-04 | Amazon Technologies, Inc. | Invariant referencing in digital works |
| US8024400B2 (en) | 2007-09-26 | 2011-09-20 | Oomble, Inc. | Method and system for transferring content from the web to mobile devices |
| US7751807B2 (en) | 2007-02-12 | 2010-07-06 | Oomble, Inc. | Method and system for a hosted mobile management service architecture |
| US9031947B2 (en) * | 2007-03-27 | 2015-05-12 | Invention Machine Corporation | System and method for model element identification |
| US9665529B1 (en) | 2007-03-29 | 2017-05-30 | Amazon Technologies, Inc. | Relative progress and event indicators |
| US7716224B2 (en) | 2007-03-29 | 2010-05-11 | Amazon Technologies, Inc. | Search and indexing on a user device |
| US20080288488A1 (en) * | 2007-05-15 | 2008-11-20 | Iprm Intellectual Property Rights Management Ag C/O Dr. Hans Durrer | Method and system for determining trend potentials |
| US20080293450A1 (en) | 2007-05-21 | 2008-11-27 | Ryan Thomas A | Consumption of Items via a User Device |
| US20080301579A1 (en) * | 2007-06-04 | 2008-12-04 | Yahoo! Inc. | Interactive interface for navigating, previewing, and accessing multimedia content |
| US9087032B1 (en) * | 2009-01-26 | 2015-07-21 | Amazon Technologies, Inc. | Aggregation of highlights |
| US8378979B2 (en) | 2009-01-27 | 2013-02-19 | Amazon Technologies, Inc. | Electronic device with haptic feedback |
| JP2012520529A (ja) * | 2009-03-13 | 2012-09-06 | インベンション マシーン コーポレーション | 知識調査のためのシステム及び方法 |
| CN102439590A (zh) * | 2009-03-13 | 2012-05-02 | 发明机器公司 | 用于自然语言文本的自动语义标注的系统和方法 |
| US8832584B1 (en) | 2009-03-31 | 2014-09-09 | Amazon Technologies, Inc. | Questions on highlighted passages |
| US8692763B1 (en) | 2009-09-28 | 2014-04-08 | John T. Kim | Last screen rendering for electronic book reader |
| US9495322B1 (en) | 2010-09-21 | 2016-11-15 | Amazon Technologies, Inc. | Cover display |
| US9298287B2 (en) | 2011-03-31 | 2016-03-29 | Microsoft Technology Licensing, Llc | Combined activation for natural user interface systems |
| US9244984B2 (en) | 2011-03-31 | 2016-01-26 | Microsoft Technology Licensing, Llc | Location based conversational understanding |
| US9858343B2 (en) | 2011-03-31 | 2018-01-02 | Microsoft Technology Licensing Llc | Personalization of queries, conversations, and searches |
| US9842168B2 (en) | 2011-03-31 | 2017-12-12 | Microsoft Technology Licensing, Llc | Task driven user intents |
| US9760566B2 (en) | 2011-03-31 | 2017-09-12 | Microsoft Technology Licensing, Llc | Augmented conversational understanding agent to identify conversation context between two humans and taking an agent action thereof |
| US10642934B2 (en) | 2011-03-31 | 2020-05-05 | Microsoft Technology Licensing, Llc | Augmented conversational understanding architecture |
| US9454962B2 (en) * | 2011-05-12 | 2016-09-27 | Microsoft Technology Licensing, Llc | Sentence simplification for spoken language understanding |
| US9064006B2 (en) | 2012-08-23 | 2015-06-23 | Microsoft Technology Licensing, Llc | Translating natural language utterances to keyword search queries |
| US9158741B1 (en) | 2011-10-28 | 2015-10-13 | Amazon Technologies, Inc. | Indicators for navigating digital works |
| CN108090094A (zh) * | 2016-11-23 | 2018-05-29 | 北京国双科技有限公司 | 一种文本信息分类方法及系统 |
| CN110162778B (zh) * | 2019-04-02 | 2023-05-26 | 创新先进技术有限公司 | 文本摘要的生成方法及装置 |
Family Cites Families (16)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US4965763A (en) * | 1987-03-03 | 1990-10-23 | International Business Machines Corporation | Computer method for automatic extraction of commonly specified information from business correspondence |
| US4930077A (en) * | 1987-04-06 | 1990-05-29 | Fan David P | Information processing expert system for text analysis and predicting public opinion based information available to the public |
| JP2783558B2 (ja) * | 1988-09-30 | 1998-08-06 | 株式会社東芝 | 要約生成方法および要約生成装置 |
| JP2790466B2 (ja) * | 1988-10-18 | 1998-08-27 | 株式会社日立製作所 | 文字列検索方法及び装置 |
| JPH03278270A (ja) * | 1990-03-28 | 1991-12-09 | Ricoh Co Ltd | 抄録文作成装置 |
| US5317507A (en) * | 1990-11-07 | 1994-05-31 | Gallant Stephen I | Method for document retrieval and for word sense disambiguation using neural networks |
| US5325298A (en) * | 1990-11-07 | 1994-06-28 | Hnc, Inc. | Methods for generating or revising context vectors for a plurality of word stems |
| EP0601666B1 (fr) | 1992-12-09 | 1999-04-21 | Koninklijke Philips Electronics N.V. | Dispositif d'affichage séquentiel couleur à valve optique |
| US5799299A (en) * | 1994-09-14 | 1998-08-25 | Kabushiki Kaisha Toshiba | Data processing system, data retrieval system, data processing method and data retrieval method |
| JPH08305695A (ja) * | 1995-04-28 | 1996-11-22 | Fujitsu Ltd | 文書処理装置 |
| US5887120A (en) * | 1995-05-31 | 1999-03-23 | Oracle Corporation | Method and apparatus for determining theme for discourse |
| US5778397A (en) | 1995-06-28 | 1998-07-07 | Xerox Corporation | Automatic method of generating feature probabilities for automatic extracting summarization |
| US6006221A (en) * | 1995-08-16 | 1999-12-21 | Syracuse University | Multilingual document retrieval system and method using semantic vector matching |
| US5963940A (en) * | 1995-08-16 | 1999-10-05 | Syracuse University | Natural language information retrieval system and method |
| US6026388A (en) * | 1995-08-16 | 2000-02-15 | Textwise, Llc | User interface and other enhancements for natural language information retrieval system and method |
| US5963893A (en) * | 1996-06-28 | 1999-10-05 | Microsoft Corporation | Identification of words in Japanese text by a computer system |
-
1998
- 1998-02-18 EP EP98914784A patent/EP0968478A1/fr not_active Withdrawn
- 1998-02-18 JP JP54000698A patent/JP2001515623A/ja active Pending
- 1998-02-18 WO PCT/DE1998/000485 patent/WO1998041930A1/fr not_active Ceased
-
1999
- 1999-09-16 US US09/381,180 patent/US6401086B1/en not_active Expired - Fee Related
Non-Patent Citations (1)
| Title |
|---|
| See references of WO9841930A1 * |
Also Published As
| Publication number | Publication date |
|---|---|
| US6401086B1 (en) | 2002-06-04 |
| WO1998041930A1 (fr) | 1998-09-24 |
| JP2001515623A (ja) | 2001-09-18 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO1998041930A1 (fr) | Procede de generation automatique d'un resume d'un texte par un ordinateur | |
| DE69804495T2 (de) | Informationsmanagement und wiedergewinnung von schlüsselbegriffen | |
| DE69617515T2 (de) | Automatisches Verfahren zur Erzeugung von thematischen Zusammenfassungen | |
| DE69726339T2 (de) | Verfahren und Apparat zur Sprachübersetzung | |
| DE68923981T2 (de) | Verfahren zur Bestimmung von Textteilen und Verwendung. | |
| DE19952769B4 (de) | Suchmaschine und Verfahren zum Abrufen von Informationen mit Abfragen in natürlicher Sprache | |
| DE69811066T2 (de) | Datenzusammenfassungsgerät. | |
| DE68928230T2 (de) | System zur grammatikalischen Verarbeitung eines aus natürlicher Sprache zusammengesetzten Satzes | |
| DE602004003361T2 (de) | System und verfahren zur erzeugung von verfeinerungskategorien für eine gruppe von suchergebnissen | |
| DE3901485C2 (de) | Verfahren und Vorrichtung zur Durchführung des Verfahrens zur Wiedergewinnung von Dokumenten | |
| DE69623082T2 (de) | Automatische Methode zur Extraktionszusammenfassung durch Gebrauch von Merkmal-Wahrscheinlichkeiten | |
| DE4015905C2 (de) | Sprachanalyseeinrichtung, -verfahren und -programm | |
| DE69423137T2 (de) | Verfahren zur Verarbeitung mehrerer elektronisch gespeicherte Dokumente | |
| DE69618089T2 (de) | Automatische Methode zur Erzeugung von Merkmalwahrscheinlichkeiten für automatische Extraktionszusammenfassung | |
| DE69423254T2 (de) | Verfahren und Gerät zur automatischen Spracherkennung von Dokumenten | |
| DE69731142T2 (de) | System zum Wiederauffinden von Dokumenten | |
| DE69032712T2 (de) | Hierarchischer vorsuch-typ dokument suchverfahren, vorrichtung dazu, sowie eine magnetische plattenanordnung für diese vorrichtung | |
| DE68928775T2 (de) | Verfahren und Vorrichtung zur Herstellung einer Zusammenfassung eines Dokumentes | |
| DE102004003878A1 (de) | System und Verfahren zum Identifizieren eines speziellen Wortgebrauchs in einem Dokument | |
| DE68922870T2 (de) | Textverarbeitungseinrichtung für europäische Sprachen mit Rechtschreibungs-Korrekturfunktion. | |
| DE10308550A1 (de) | System und Verfahren zur automatischen Daten-Prüfung und -Korrektur | |
| DE69710309T2 (de) | System für betriebliche veröffentlichung und speicherung | |
| DE69733294T2 (de) | Einrichtung und Verfahren zum Zugriff auf eine Datenbank | |
| DE69229583T2 (de) | Verfahren zur Flektieren von Wörtern und Datenverarbeitungseinheit zur Durchführung des Verfahrens | |
| DE19849855C1 (de) | Verfahren zur automatischen Generierung einer textlichen Äußerung aus einer Bedeutungsrepräsentation durch ein Computersystem |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| 17P | Request for examination filed |
Effective date: 19990903 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): DE FR GB |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN |
|
| 18D | Application deemed to be withdrawn |
Effective date: 20030902 |