EP3100176A1 - Verfahren zur semantischen analyse eines textes - Google Patents

Verfahren zur semantischen analyse eines textes

Info

Publication number
EP3100176A1
EP3100176A1 EP15703746.6A EP15703746A EP3100176A1 EP 3100176 A1 EP3100176 A1 EP 3100176A1 EP 15703746 A EP15703746 A EP 15703746A EP 3100176 A1 EP3100176 A1 EP 3100176A1
Authority
EP
European Patent Office
Prior art keywords
theme
coefficient
text
subset
words
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
EP15703746.6A
Other languages
English (en)
French (fr)
Inventor
Jean-Pierre Malle
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Deadia
Original Assignee
Deadia
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Deadia filed Critical Deadia
Publication of EP3100176A1 publication Critical patent/EP3100176A1/de
Withdrawn legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/35Clustering; Classification
    • G06F16/355Creation or modification of classes or clusters
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/33Querying
    • G06F16/3331Query processing
    • G06F16/334Query execution
    • G06F16/3344Query execution using natural language analysis
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • G06F40/205Parsing
    • G06F40/211Syntactic parsing, e.g. based on context-free grammar [CFG] or unification grammars
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/30Semantic analysis
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/40Processing or translation of natural language

Definitions

  • the present invention relates to the field of computer semantic understanding.
  • the semantic analysis of a text in natural language aims to establish the meaning by using the meaning of the words that constitute it, following a lexical analysis that allows to break down this text with the help of a lexicon or a grammar.
  • the human unconsciously realizes it to understand the texts it reads, and recent developments aim to confer machine-like capabilities.
  • Automated semantic analysis algorithms are known for the moment, designed for a computer to classify a text into several predetermined categories, for example general topics such as "nature”, “economy”, “literature”, etc.
  • the present invention proposes a semantic analysis method of a text in natural language received by a piece of equipment from input means, the method being characterized in that it comprises the implementation by data processing means of the device. equipment of steps of:
  • a coverage coefficient of a theme is calculated in step (d) as the number N of reference words associated with the theme included in said subset;
  • a coefficient of relevance of a theme is calculated in step (d) by the formula N * (1 + ln (#)), where N is the number of reference words associated with the theme included in the sub-group. together and R the total number of occurrences in said part of the text of reference words associated with the theme;
  • step (c) • two thematic orientation coefficients are calculated in step (c), including a theme certainty coefficient and a thematic shade coefficient;
  • a certainty coefficient of a theme is calculated in step (d) as:
  • a shade coefficient of a theme is a positive scalar greater than 1 when the words that are not part of the subset are representative of an amplification of the theme, and a positive scalar less than 1 when the words that do not part of the subset are representative of an attenuation of the theme;
  • the method comprises a step (aO) preceding the cutting of the text into a plurality of propositions, each being a part of the text for which the steps (a) to (d) of the method are repeated so as to obtain for each proposition a set of the coverage, relevance, and / or orientation coefficients associated with the proposal, the method comprising, before step (e), a step (eO) of calculation for each of said subsets and for each theme identified for minus a proposal of the text of a global coefficient of coverage of the theme and / or of an overall coefficient of relevance of the theme, and of at least one overall orientation coefficient of the theme according to all of said coefficients associated with a proposition;
  • An overall coefficient of coverage of a theme is calculated in step (eO) as the sum of the coverage coefficients of the theme associated with a proposition minus the number of reference words of the theme present in at least two propositions;
  • step (eO) an overall coefficient of relevance of a theme is calculated in step (eO) as the sum of the relevance coefficients of the theme associated with a proposition;
  • An overall orientation coefficient of a theme is calculated in step (eO) as the average of the orientation coefficients of the theme associated with a proposition weighted by the associated thematic coverage coefficients;
  • the step (eO) includes for each of said subsets and for each theme the calculation of an overall coefficient of divergence of the theme corresponding to the standard deviation of the product distribution of the products. orientation coefficients by the coverage coefficients associated with each proposal;
  • the subset / thematic pairs selected in step (f) are those such that for any partition of the subset into a plurality of parts of said subset, the semantic coefficient of the subset for the theme is greater than the sum of the semantic coefficients of the sub-parts of the subset for the theme;
  • step (g) comprising determining the group or groups comprising at least one subset / thematic pair selected at the step (f);
  • Step (g) comprises the creation of a new group if no subset / thematic group of couples of reference contains at least one subset / thematic pair selected for the text;
  • Each reference subset / thematic pair is associated with a score stored on the data storage means, the score of a reference subset / thematic couple decreasing with time but increasing each time this couple under -set / thematic is selected for a text;
  • the method comprises a step (h) of deleting a subassembly / thematic reference pair of a group if the score of said pair falls below a first threshold, or of modifying the data storage means (12) said plurality of topic-related lists if the score of said pair exceeds a second threshold;
  • Step (g) comprises for each group of couples subset / thematic reference the calculation of a dilution coefficient representing the number of occurrences in said part of the text of words references associated with themes of subset / thematic pairs of reference present in the text related to the total number of reference words associated with said themes;
  • the invention relates to an equipment comprising data processing means configured to implement following the reception of a text in natural language a method according to the first aspect of the invention semantic text analysis .
  • FIG. 1 is a diagram of a network architecture embodying the invention
  • FIG. 2 is a diagram showing schematically the steps of the semantic analysis method according to the invention.
  • the present method is implemented by data processing means 11 (which typically consist of one or more processors) of a piece of equipment 1.
  • the latter can be for example one or more servers connected to a network 4, typically the Internet, via which it is connected to clients 2 (for example personal computers).
  • the equipment 1 further comprises data storage means 12 (typically one or more hard disks).
  • a text here is any message in natural and meaningful language.
  • the text is received in electronic form, that is to say in a format directly processable by the processing means 1 1, for example XML (Extensible Markup Language).
  • XML Extensible Markup Language
  • input means 14 is meant a wide variety of origins.
  • the term input means means any means, hardware and / or software, for recovering the text and send it to the data processing means 11 in a readable format.
  • the text can be directly typed by a user, and the input means 14 refer for example to a keyboard and word processing software.
  • the text may be a paper text scanned and recognized by OCR (Optical Character Recognition), and the input means 14 then designate a scanner and a digital data processing software, or the text can be dictated and the means input 14 then designate a microphone and voice recognition software.
  • OCR Optical Character Recognition
  • the text can be received for example from a server of the Internet network, possibly directly in a readable format.
  • the present method is not limited to any type of text.
  • the input means are typically those of a client 2 or another server 1.
  • the text is structured in sections. Sections can be separated by paragraphs or simply chained. The sections are distinguished from each other by the fact that the concepts presented are substantially different. Detection of unmarked sections by the author is a complex operation. A section is composed of sentences separated by punctuation (colon, period, exclamation mark, question mark, hyphen, dots, etc.).
  • a sentence is composed of propositions separated by a punctuation (comma, semi-colon).
  • a proposition is a sequence of words separated by spaces.
  • a word is an ordered set of letters and special signs (accents, hyphens, etc.).
  • a first step (a), called "parsing” at least a portion of the text is cut syntactically into a plurality of words.
  • this part of the sentence is a proposition
  • the text is first cut proposal by proposition in a step (aO) before each proposal is in turn cut into words.
  • aO a step
  • the breakdown by propositions can be done following a division by sentences, itself after a division by sections.
  • the identification of words is done through spaces.
  • a parser (the parsing engine) using punctuation and formatting as a delimiter of the propositions may suffice if the punctuation is respected.
  • a text is classified in one or more "categories” according to the meaning it bears. Categories are here moving sets.
  • the categories are defined as groups of "rings" and can be induced by the appearance of a text under a new meaning.
  • a category is represented by a list of themes.
  • the theme is the meaning is attached to a set of words (so-called reference words) entering the composition of a proposal, present in a list called thematic.
  • the theme is attached to one or more categories.
  • the list of associated reference words is stored on the storage means 12 of the equipment 1.
  • a "motorization” theme can include reference words ⁇ motor, piston, cylinder, crankshaft, shaft, connecting rod, pedal, power, etc. ⁇
  • a “geometry” theme can include the reference words ⁇ right, angle, degree, star, rectangle, sphere, cylinder, pyramid, etc. ⁇ .
  • the word “cylinder” has many meanings and is thus related to both themes although they are distant.
  • the engine comprises three pistons connected to a crankshaft by star-shaped rods forming an angle of 120 ° two by two which reacts to the slightest pressure on the accelerator pedal ", or slight variation of this proposal.
  • step (b) at least one theme is identified among the plurality of themes each associated with a list of reference words of the stored theme.
  • a reference word associated with the theme is present for the theme to be associated.
  • at least two (or more) words are required.
  • the set of words of the part of the analyzed text associated with at least one theme is also identified. This is ⁇ engine, piston, crankshaft, connecting rod, pedal, angle, 120 °, star ⁇
  • V be a vocabulary of Nv words (in particular the set of reference words of at least one theme).
  • T be a subset of V of Nt words (in particular all of the reference words present in at least one theme), Nt ⁇ Nv.
  • P (P) and P (Q) are unitary commutative rings provided with two operators:
  • a symmetric difference operator denoted ⁇ (relative to two sets A and B, the symmetrical difference of A and B being the set containing the elements contained in A but not in B, and the elements contained in B and not in A) ; and - an intersection operator noted &.
  • P (P) is isomorphic to Z / NpZ and P (Q) is isomorphic to Z / NqZ
  • V A e P (P), P (A) is included in P (P) and A is also a unitary commutative ring.
  • A contains all the complete or partial combinations of a group of words.
  • A a "semantic ring". From the set of words of a proposition belonging to a theme, a semantic ring is defined by a subset of this set.
  • each ring is not the simple list of words that compose it, but rather the set of sets including i G HO, Kj of these words (which are other semantic rings).
  • the ring defined by vehicle and large actually corresponds to the set ⁇ ; ⁇ vehicle ⁇ ; ⁇ big ⁇ ⁇ vehicle, large ⁇ .
  • a ring is said centered if it does not exist two words that it contains belonging to two different thematics (but it can contain words not belonging to any thematic).
  • a ring is said to be regular if it also belongs to P (Q), that is to say that all the words it contains belong to one of the themes.
  • the method comprises constructing a plurality of subsets of the set of words of said part of the text associated with at least one theme, in other words the rings semantics, and advantageously the method comprises the construction of all of these rings.
  • step (d) a representation of the "meaning” of the semantic rings of a part of the text (which as explained is typically a proposition) is determined by the data processing means 1 1 of the equipment 1.
  • This representation takes the form of a matrix of vectors attached to the themes and comprising several dimensions and stored in the data storage means 12 of the equipment. This matrix is called “semantic matrix" (or sense matrix).
  • semantic matrix or sense matrix.
  • a sequence of semantic matrices is determined, and in a step (eO) a global semantic matrix of the text is determined according to the semantic matrices of the rings of the propositions.
  • a semantic matrix comprises at least two dimensions, advantageously three or even four: the coverage, the relevance (at least one of these two is required), the certainty, the nuance (the last two can be grouped into one dimension, the orientation).
  • the overall matrix of a text may include a fifth dimension (divergence).
  • the method comprises for each subgroup (ie semantic ring) and each identified theme, the calculation in a coverage coefficient of the theme and / or a relevance coefficient of the theme (advantageously both), depending of occurrences in the ring of reference words associated with the theme.
  • the coverage coefficient of a theme materializes the proximity between the ring and the theme, and is represented by an integer, typically the number N of words of the theme included in the ring. It is possible to add weightings (for example to certain "essential" words of the theme).
  • the coefficient of relevance is calculated by the data processing means 1 1 as the coverage coefficient but taking into account the total number of occurrences of the words of the theme.
  • N is the number of words of the theme contained in the ring, or each word counts only once (in other words the coverage coefficient of the theme) and R is the number of words of the theme contained in the ring, where each word counts as many times as it appears in the proposition (number of total occurrence, which increases with the length of the proposition)
  • the coefficient of relevance is for example given by the formula N * (1 + ln (#)), with In the natural logarithm.
  • Coefficient of certainty of a theme also comprises calculating, for each subgroup (ie semantic ring) and each identified theme, at least one orientation coefficient of the theme from the words of said part of the text that is not part of the ring (especially those not belonging to any ring).
  • two orientation coefficients of the theme are calculated in step (d), including a certainty coefficient of the theme and a coefficient of nuance of the theme.
  • Certainty is conveyed by a set of words whose order and nature can radically change the meaning of the proposition. These are typically words such as negations, punctuation, interrogative / negative words, a list of which can be stored on the data storage means 12. The position of these words relative to each other (typical of certain turns) gives besides indices on the certainty.
  • proximity can be affirmative, negative or uncertain.
  • the proximity is affirmative (for lack of words modifying the certainty).
  • the nuance is conveyed by a set of words whose order and nature can alter the meaning of the proposition.
  • This alteration can be a reinforcement or a weakening of the proximity with the theme, for example thanks to adverbs such as "certainly”, “assuredly”, “probably”, “possibly”.
  • adverbs such as "certainly”, “assuredly”, “probably”, “possibly”.
  • the data processing means 1 1 compare the words not associated with the theme with this list and deduce the value of the hue coefficient, which is in particular a positive scalar (greater than 1 for a reinforcement and less than 1 for a weakening )
  • each word representative of a shade can be stored associated with a coefficient, the coefficient of shade for the proposition being for example the product of the coefficients of the words found in the proposal.
  • the shade coefficient for the proposition can be the sum of the coefficients of the words found in the proposition.
  • the shading and certainty coefficients can constitute two distinct dimensions of the semantic matrix, or can be treated together as an orientation coefficient ("the orientator").
  • the orientation coefficient is thus typically a real number:
  • the semantic matrix obtained preferably has a structure of the type
  • Guidance 1 Guidance 2 Guidance 3 Guidance i Composition of semantic matrices As explained above, a text is formed of several sentences formed themselves of several propositions. A semantic matrix is advantageously generated for a ring for each proposition.
  • a step (eO) the semantic matrices of a ring are combined into a global matrix: is computed by the data processing means 1 1 for each ring and each theme identified for at least one proposition of the text a global coefficient of coverage of the theme and / or of an overall coefficient of relevance of the theme, and of at least one overall orientation coefficient of the theme as a function of all of said coefficients associated with a proposition.
  • the matrices of two propositions are complementary if they relate to different themes.
  • the meaning matrix of the set of two propositions consists of the juxtaposition of the two matrices (since no theme is common).
  • the matrices of two propositions are coherent if they relate to common themes with similar orientators.
  • the matrices of two propositions are opposed if they relate to common themes with opposite orientators (of different signs, i.e. the difference relates to the coefficient of certainty of the theme).
  • an overall coverage coefficient of a theme is calculated as the sum of the coverage coefficients of the theme associated with a proposal minus the number of reference words of the theme present in at least two propositions (in other words, it is necessary to count only once each word, the coverage of the sum is thus included between the largest of the covers (where all the reference words of the theme found in one proposal are also in the other), and the sum (case where no reference word is common to both thematic covers).
  • overall coverage coefficient can be easily recalculated as the number Nmax ôe theme words contained in the set of proposals);
  • an overall coefficient of relevance of a theme is calculated as the sum of the relevance coefficients of the theme associated with a proposition (since multiple occurrences are taken into account);
  • an overall orientation coefficient of a theme is calculated as the average of the orientation coefficients of the theme associated with a proposition weighted by the associated thematic coverage coefficients.
  • OS (OA * CA + OB * CB) / CS
  • thematic divergence is defined as representing the variations of meaning for a theme in a text.
  • the step (eO) thus comprises for each theme the calculation of an overall coefficient of divergence of the theme. It is calculated, for example, as the standard deviation of the distribution of the products of the advisers by the covers of the relevant proposals, brought back to the holistic product of the supplier by the coverage of the global text.
  • a text with strong divergence is a text in which the subject carried by the theme is approached with questions, comparisons, confrontations.
  • a low-divergence text is a text that constantly presents the same angle of view.
  • semantic coefficient representative of a degree of meaning carried by the subgroup according to said coverage coefficients, relevance and / or orientation of the theme, in particular the global coefficients.
  • This coefficient is calculated by the data processing means in step (e) of the method.
  • VAGP (P), with TGP (V), M (A, T) relevance (A, T) * orientator (A, T) * [1 + divergence (A, T) 2 ]
  • M (A, T) is the semantic coefficient of the ring A of the proposition P with respect to the theme T according to the vocabulary V.
  • M (A) is the semantic coefficient of the ring A of the proposition P with respect to all the themes according to the vocabulary V.
  • VAGP (P), with TGP (V), M (A, T) [relevance (A, T)] 2 * orientator (A, T), or
  • VAGP (P), with TGP (V), M (A, T) relevance (A, T) * neck under re (A, T)
  • the semantic coefficient makes it possible to select rings / thematic couples most meaningful in a step (f). In particular, it may be those for which the coefficient is the highest, but alternatively we can use the criterion of "growth" semantic rings.
  • a semantic ring increasing according to M any element A of P (Q) for which:
  • a growing semantic ring is a ring carrying a meaning greater than the sum of the meanings of its parts.
  • the sum of the semantic coefficients of the parts of the ring partition with respect to this theme is less than the semantic coefficient of the entire ring with respect to this theme.
  • the subset / thematic pairs selected in step (f) are those for which the ring is increasing for this theme.
  • this vehicle is big and blue
  • the rings ⁇ vehicle, big ⁇ and ⁇ vehicle, blue ⁇ carry less meaning than the global ring ⁇ vehicle, big, blue ⁇ . The latter is growing.
  • the union of two decreasing semantic rings is a descending semantic ring.
  • the union of a descending semantic ring and a growing semantic ring is a descending semantic ring.
  • the union of two increasing semantic rings is a semantic ring either increasing or decreasing.
  • the growing character is recessive to the union.
  • An expressive semantic ring is a set of words with a cultural meaning superior to that of the union of its parts.
  • this vehicle is a real bomb
  • the expressive ring ⁇ vehicle, bomb ⁇ associated with a reinforcing shade carries an expressive meaning not present in singletons rings ⁇ vehicle ⁇ and ⁇ bomb ⁇ and not present in the descending ring ⁇ vehicle, bomb ⁇ .
  • An expressive ring A is a descending ring that has become larger by a shade enhancement (ie, because of a high degree of shade due to the presence of "true” leading to a high orientation). Morphism M then has a discontinuity in the vicinity of A.
  • step (f) certain filters can eliminate certain rings according to a parameterization of the engine.
  • the first part which corresponds to the steps (a) to (f) already described, is implemented by a block called the analyzer for selecting rings / thematic couples representative of the meaning of the text.
  • a classifier associates the categories with the texts using the selected rings.
  • the categories corresponding groups of couples subset / thematic reference are stored on the data storage means 12, and the categories in which the text is classified are those comprising at least a subset / thematic couple selected at the step (f).
  • Step (g) may thus comprise the calculation of a so-called dilution coefficient, which represents the number of occurrences of terms of the themes related to the determined category or categories (in other words the themes of the pairs of the associated groups). to the categories), present in the text compared to the total number of terms of the said themes. It is said that the text is of category X according to dilution D.
  • the categories are not fixed and can evolve. In particular new categories can be generated and others segmented.
  • a new category can be generated with a new meaning: a new group is created if no subset / thematic group of couples of reference contains at least a subset / thematic couple selected for the text. The subset / thematic couples become the reference ones of this group.
  • each reference subset / thematic pair may be associated with a score stored on the data storage means 12, the score of a reference subassembly / thematic pair diminishing over time (for example, by damping hyperbolic) but increasing each time this subset / thematic couple is selected for a text.
  • the method may then comprise a step (h) of deleting a subassembly / thematic reference pair of a group if the score of said pair falls below a first threshold, or of modification on the storage means 12 of said plurality of lists associated with the themes if the score of said pair passes above a second threshold.
  • the connectivity between a ring and a theme can indeed be represented by a coefficient representing for each theme the frequency of appearance of this theme among the topics such that the pair ring / thematic associated has already been selected.
  • the connectivity between a ring and a theme is for example given as the score of this ring / thematic couple on the sum of the scores associated with pairs of this ring with a reference theme.
  • a strongly eroded ring (score passing below the first threshold) disappears from the stack.
  • Both thresholds can be set manually according to the "sensitivity", that is, the desired level of scalability of the system. Close thresholds (first high threshold and / or second low threshold) lead to a strong renewal of themes and categories.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Computational Linguistics (AREA)
  • Artificial Intelligence (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • General Health & Medical Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • Machine Translation (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
EP15703746.6A 2014-01-28 2015-01-28 Verfahren zur semantischen analyse eines textes Withdrawn EP3100176A1 (de)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
FR1400201A FR3016981A1 (fr) 2014-01-28 2014-01-28 Procede d'analyse semantique d'un texte
PCT/EP2015/051722 WO2015114014A1 (fr) 2014-01-28 2015-01-28 Procédé d'analyse sémantique d'un texte

Publications (1)

Publication Number Publication Date
EP3100176A1 true EP3100176A1 (de) 2016-12-07

Family

ID=51417300

Family Applications (1)

Application Number Title Priority Date Filing Date
EP15703746.6A Withdrawn EP3100176A1 (de) 2014-01-28 2015-01-28 Verfahren zur semantischen analyse eines textes

Country Status (5)

Country Link
US (1) US10289676B2 (de)
EP (1) EP3100176A1 (de)
CA (1) CA2937930A1 (de)
FR (1) FR3016981A1 (de)
WO (1) WO2015114014A1 (de)

Families Citing this family (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10565291B2 (en) * 2017-10-23 2020-02-18 International Business Machines Corporation Automatic generation of personalized visually isolated text
US11409749B2 (en) * 2017-11-09 2022-08-09 Microsoft Technology Licensing, Llc Machine reading comprehension system for answering queries related to a document
US11593561B2 (en) * 2018-11-29 2023-02-28 International Business Machines Corporation Contextual span framework
CN110134957B (zh) * 2019-05-14 2023-06-13 云南电网有限责任公司电力科学研究院 一种基于语义分析的科技成果入库方法及系统
US10978053B1 (en) * 2020-03-03 2021-04-13 Sas Institute Inc. System for determining user intent from text
CN117521669B (zh) * 2023-11-10 2024-07-02 北京博大网信股份有限公司 一种数学应用题语义分析自动求解系统

Family Cites Families (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6192360B1 (en) * 1998-06-23 2001-02-20 Microsoft Corporation Methods and apparatus for classifying text and for building a text classifier
US6510406B1 (en) * 1999-03-23 2003-01-21 Mathsoft, Inc. Inverse inference engine for high performance web search
US6990496B1 (en) * 2000-07-26 2006-01-24 Koninklijke Philips Electronics N.V. System and method for automated classification of text by time slicing
US7493253B1 (en) * 2002-07-12 2009-02-17 Language And Computing, Inc. Conceptual world representation natural language understanding system and method
EP1462950B1 (de) * 2003-03-27 2007-08-29 Sony Deutschland GmbH Verfahren zur Sprachmodellierung
JP4713870B2 (ja) * 2004-10-13 2011-06-29 ヒューレット−パッカード デベロップメント カンパニー エル.ピー. 文書分類装置、方法、プログラム

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
None *
See also references of WO2015114014A1 *

Also Published As

Publication number Publication date
WO2015114014A1 (fr) 2015-08-06
CA2937930A1 (fr) 2015-08-06
US20160350277A1 (en) 2016-12-01
US10289676B2 (en) 2019-05-14
FR3016981A1 (fr) 2015-07-31

Similar Documents

Publication Publication Date Title
US11734329B2 (en) System and method for text categorization and sentiment analysis
EP3100176A1 (de) Verfahren zur semantischen analyse eines textes
FR2694984A1 (fr) Procédé d'identification, de récupération et de classement de documents.
EP1977343A1 (de) Verfahren und einrichtung zum abrufen von daten und zum transformieren dieser in qualitative daten eines auf text basierenden dokuments
FR2975201A1 (fr) Analyse de texte utilisant des proprietes de listes linguistiques et non-linguistiques
Hofmann et al. The reddit politosphere: a large-scale text and network resource of online political discourse
US9633008B1 (en) Cognitive presentation advisor
WO2006120352A1 (fr) Dispositif et procede d'analyse semantique de documents par constitution d'arbres n-aire et semantique
CN113051374A (zh) 一种文本匹配优化方法及装置
CA3131157A1 (en) System and method for text categorization and sentiment analysis
Papegnies et al. Impact of content features for automatic online abuse detection
Beleveslis et al. A Hybrid Method for Sentiment Analysis of Election Related Tweets.
CN114372461A (zh) 一种隐性关键词提取方法、终端设备及存储介质
US11990131B2 (en) Method for processing a video file comprising audio content and visual content comprising text content
EP2013776A1 (de) Verfahren zur schnellen neuduplikation einer menge von dokumenten oder einer menge von in einer datei enthaltenen daten
US9569538B1 (en) Generating content based on a work of authorship
WO2017088126A1 (zh) 获取未登录词的方法与装置
EP4300326A1 (de) Verfahren zum abgleich einer zu untersuchenden anordnung und einer referenzliste, entsprechende paarungsmaschine und computerprogramm
Kaldeli et al. Combining Automatic Annotation with Human Validation for the Semantic Enrichment of Cultural Heritage Metadata.
WO2013117872A1 (fr) Procede d'identification d'un ensemble de phrases d'un document numerique, procede de generation d'un document numerique, dispositif associe
FR3030809A1 (fr) Procede d'analyse automatique de la qualite litteraire d'un texte
FR2970795A1 (fr) Procede de filtrage de synonymes.
FR3163752A1 (fr) Procédé et module de traitement de chaîne de caractères
FR3156945A1 (fr) Procédé et dispositif de sécurisation d'un modèle de langage de grande taille
Danezis Detecting Hate Speech Online using Machine Learning

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20160824

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR

AX Request for extension of the european patent

Extension state: BA ME

DAX Request for extension of the european patent (deleted)
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: EXAMINATION IS IN PROGRESS

17Q First examination report despatched

Effective date: 20200127

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN

18D Application deemed to be withdrawn

Effective date: 20220802