WO2021044475A1 - 文章解析システムおよびこれを用いたメッセージ交換における特徴評価システム - Google Patents
文章解析システムおよびこれを用いたメッセージ交換における特徴評価システム Download PDFInfo
- Publication number
- WO2021044475A1 WO2021044475A1 PCT/JP2019/034402 JP2019034402W WO2021044475A1 WO 2021044475 A1 WO2021044475 A1 WO 2021044475A1 JP 2019034402 W JP2019034402 W JP 2019034402W WO 2021044475 A1 WO2021044475 A1 WO 2021044475A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- sentence
- feature
- sentence analysis
- feature information
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/12—Use of codes for handling textual entities
- G06F40/126—Character encoding
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/205—Parsing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/33—Querying
- G06F16/3331—Query processing
- G06F16/334—Query execution
- G06F16/3344—Query execution using natural language analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/194—Calculation of difference between files
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/205—Parsing
- G06F40/216—Parsing using statistical methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/205—Parsing
- G06F40/226—Validation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F21/00—Security arrangements for protecting computers, components thereof, programs or data against unauthorised activity
- G06F21/50—Monitoring users, programs or devices to maintain the integrity of platforms, e.g. of processors, firmware or operating systems
Definitions
- the present invention relates to a text analysis system and a feature evaluation system for message exchange using the text analysis system.
- sample data in which a character string or the like is described is signalized as n-valued sample data (n is a natural number of 2 or more), and n-valued sample data and n-valued input are obtained. It discloses a technique for calculating the similarity with data and identifying whether or not the input data is spam mail based on the calculated similarity.
- the purpose of processing a message written in natural language is not only to understand the content but also to acquire the characteristics of the message creator.
- the characteristics of message creators are also utilized in the field of information security.
- Information leakage due to obstruction of operation of computer devices and electronic devices using messages, information fraud, fraudulent acts against users, etc. is a big problem, and there is a high demand for information leakage prevention by message analysis, and in addition, it is high speed. Processing is also required.
- the other is erroneous transmission by the user. For example, you can send a message to an unknown destination, use a topic or term that you don't normally deal with, or attach a file that you don't normally attach. A common feature of these is that they are accompanied by unusual behavior. Therefore, it is possible to prevent information leakage due to message exchange by detecting the peculiarity existing in the message at high speed and paying attention before transmission.
- an object of the present invention is to provide a sentence analysis system capable of detecting a sentence having peculiar expressive features and structural features at a lower cost and at a higher speed than before. Furthermore, an object of the present invention is to provide a message feature evaluation system that detects the peculiarity of the text in message exchange.
- the present invention realizes a system capable of processing a wide variety of languages by a single algorithm.
- the sentence analysis system of the present invention can be applied to the detection of features and exceptions of spoken words and sentences. According to the present invention, it is possible to detect extraordinary ideas buried in mediocre ideas and a small number of intentions in a large number, including the discovery of differences in meaning, misunderstandings, injustices and their signs caused by wording errors and irregularities. It becomes.
- the sentence analysis system of the present invention can be used in a wide variety of ways. Cross.
- the sentence analysis system for analyzing a sentence is converted into an acquisition means for acquiring sentence data and a conversion means for converting sentence data into a time-series signal by digitizing the characters of the acquired sentence data. It has a feature extraction means for extracting feature information from a time-series signal and storing the extracted feature information, and a determination means for determining the identity of newly acquired text data using the feature information.
- the sentence analysis system further includes a detection means for detecting a peculiar sentence different from the feature information based on the determination result of the determination means.
- the conversion means converts characters into numerical data based on a pre-prepared conversion table.
- the conversion means normalizes the time series signal so that it falls within the range of a minimum value of 0 and a maximum value of 1.
- the conversion means attenuates the value of the time series signal that exceeds a set threshold and normalizes the attenuated time series signal.
- the feature extraction means extracts features from a normalized time-series signal of text data described in normal expressive or structural features and uses the extracted features to use the time-series signal. Learn the features so that an output waveform that reproduces the input waveform of.
- the feature extraction means encodes the feature information with an autoencoder.
- the feature extraction means learns the feature information by a neural network.
- the feature evaluation system in the message exchange includes the sentence analysis system described above, and the detection means detects the specificity of the message based on the determination result of the determination means.
- the feature evaluation system in the message exchange includes transmission control means for stopping the transmission of the outgoing mail when the peculiarity of the outgoing mail is detected.
- the feature evaluation system in the message exchange further includes a notification means for notifying the transmission of the outgoing mail when the transmission of the outgoing mail is stopped by the transmission control means.
- the sentence analysis program executed by the computer terminal according to the present invention has been converted into a step of acquiring sentence data and a step of converting the sentence data into a time-series signal by digitizing the characters of the acquired sentence data. It has a step of extracting feature information from a time-series signal and storing the extracted feature information, and a step of determining the identity of newly acquired sentence data using the feature information.
- the step of determining identity identifies an outgoing email described with a unique expressive or structural feature that is different from the feature information.
- the sentence analysis method in the computer terminal includes a step of acquiring sentence data, a step of converting sentence data into a time-series signal by digitizing the characters of the acquired sentence data, and a converted time-series. It has a step of extracting feature information from a signal and storing the extracted feature information, and a step of determining the identity of newly acquired text data using the feature information.
- the step of determining identity identifies an outgoing email described with expressive or structural features that differ from the feature information.
- the text data is converted into a time-series signal, it is possible to reduce the cost without requiring morphological analysis of the text and dictionary data for that purpose. Furthermore, by determining the identity of the text data based on the feature information extracted from the time-series signal, it is possible to easily determine whether or not the text is the text of the person himself / herself. Further, according to the present invention, by detecting the peculiarity of the sent mail, it is possible to prevent information leakage by stopping the transmission of the abnormal sent mail.
- FIG. 1st Example of this invention It is a block diagram which shows the structure of the sentence analysis system which concerns on 1st Example of this invention. It is a block diagram which shows the internal structure of the feature extraction part shown in FIG. It is an example of a part of Unicode. It is a figure which shows the example which the e-mail is acquired as the text data, and the time series signal of the e-mail is normalized. It is a flowchart explaining the operation example of the signal normalization by the Example of this invention. It is a figure explaining the feature extraction from the input by the signal classification part by the Example of this invention. It is a figure explaining the outline of the autoencoder by the Example of this invention. It is a figure which shows the example of the classification by the threshold value by a signal classification part.
- the sentence analysis system can be applied to any electronic device (for example, a computer device, a mail server, a client terminal, a smartphone, etc.) having a function of electronically processing a sentence.
- electronic device for example, a computer device, a mail server, a client terminal, a smartphone, etc.
- FIG. 1 is a diagram showing a configuration example of a sentence analysis system according to an embodiment of the present invention.
- the sentence analysis system 100 is extracted by a sentence acquisition unit 110 that acquires sentence data, a feature extraction unit 120 that extracts features of the sentence data acquired by the sentence acquisition unit 110, and a feature extraction unit 120. It is configured to include a feature storage unit 130 that stores the features, and a peculiar sentence detection unit 140 that detects a peculiar sentence based on the features of the feature extraction unit 120 or the feature storage unit 130.
- the sentence analysis system 100 is implemented by software, hardware such as a mail server or a client terminal, or a combination of software and hardware.
- the sentence acquisition unit 110 acquires sentence data (for example, e-mail) created by the user.
- sentence data for example, e-mail
- the text data is an e-mail, for example, an HTML format e-mail created by the mail software installed in the client terminal, or an e-mail sent from the client terminal to the mail server via the Internet, or a message.
- the email in the exchange system is retrieved.
- the text acquisition unit 110 can acquire text data created by a plurality of users. Further, in order to give the sentence analysis system 100 a learning function in advance, the sentence data acquired by the sentence acquisition unit 110 is a normal sentence created by the user with normal behavior, that is, normal expressional features or structural features. It is data, and the feature extraction unit 120 extracts features included in normal text data created by the user's normal expressive features or structural features, and learns the features of the user's text. After training the sentence analysis system 100, the sentence acquisition unit 110 acquires arbitrary sentence data, and the sentence analysis system 100 creates the features of the arbitrary sentence data by ordinary expressive features or structural features. Identify whether or not it matches the characteristics of the text. For example, even if the text is created by the person himself / herself, it is possible to identify whether or not the text is created by ordinary expressive or structural features, or whether the text is created by a person other than the person himself / herself. To identify.
- FIG. 2 shows the internal configuration of the feature extraction unit 120.
- the feature extraction unit 120 receives the sentence data acquired by the sentence acquisition unit 110, and signals the characters described in the sentence into a time-series signal and a character signalization unit 122, and the character signalization unit 122. It has a normalization unit 124 that normalizes time-series signals and a signal classification unit 126 that classifies the normalized signals.
- the character signal conversion unit 122 converts a series of characters described in a sentence into a one-dimensional time series signal.
- the character signaling unit 122 converts one character of a sentence into numerical data based on Unicode.
- Unicode is one of the international standards for character codes, and characters, numbers, symbols, etc. in various languages around the world are assigned to the codes.
- FIG. 3 illustrates a partial excerpt of Unicode.
- Unicode encodes ASCII, Kanji, Arabic, Greek symbols, etc. into binary data with 16 bits or more.
- the character signaling unit 122 has a data amount of the number of bits per numerical value obtained by converting one character ⁇ the number of characters. Further, the character signalizing unit 122 may convert the fixed-length data into one continuous data without a break, or may convert the fixed-length data into variable-length data.
- a conversion table that uniquely defines the relationship between characters, idioms, phrases, etc. and numerical data is prepared in advance, and the character signaling unit 122 uses such a conversion table to describe each sentence. Characters, idioms, etc. may be converted into numerical data.
- the character signaling unit 122 converts from the first character to the last character of the sentence into numerical data. For example, in the case of a sentence having a size of P rows ⁇ Q columns (P and Q are arbitrary integers), a time series signal including binary data corresponding to the number of characters of P ⁇ Q is generated. Characters here are concepts that include natural language characters, numbers, symbols, figures, and spaces that do not represent such characters. For example, in the case of horizontal writing, characters are scanned sequentially from left to right or right to left from the first line to the last line, or in the case of vertical writing, from top to bottom or from top to last line. Characters are sequentially scanned from bottom to top and converted into numerical data from the first character to the last character. The scanning direction can be arbitrarily determined. If the page information (number of lines, number of characters in one line, etc.) that composes the text data is required, the page information is acquired at the same time, and the page information is referred to to identify the first character to the last character. May be good.
- the time-series signal of the sentence generated by the character signalizing unit 122 in this way can be regarded as an aperiodic waveform created by the characters of the sentence, and the words and idioms contained in the sentence appear as a waveform pattern there.
- the time-series signal will include a waveform pattern corresponding to " ⁇ ".
- a waveform pattern representing the usual expressive or structural features is also included when the user writes sentences in polite language, makes heavy use of punctuation marks, makes heavy use of specific conjunctions, and so on. It will be.
- Such a waveform pattern is one feature for identifying a user.
- the character signalizing unit 122 Since the character signalizing unit 122 according to this embodiment signals characters based on Unicode or a conversion table, it does not depend on a specific language and can be applied to multiple languages. Can be expressed by the difference between. Further, since the character signalizing unit 122 does not perform morphological analysis or syntactic analysis of sentences, a dictionary such as a corpus is unnecessary, and the cost can be reduced.
- the signal normalization unit 124 normalizes the time series signal generated by the character signalization unit 122.
- each number that produces a time-series signal represents a discrete value, and the range of that value can be very large. Therefore, the signal normalization unit 124 performs a process of suppressing outliers of the time-series signal and a process of normalizing the range.
- the outlier suppression process attenuates a numerical value that exceeds a set threshold value.
- processing is performed by the following equation.
- Avg is the average
- std is the standard deviation
- x is the target value (here, the numerical value of the time series signal)
- rate is the attenuation factor
- d is the purpose of raising the overall value. It is a coefficient to be multiplied by the numerical value to be added.
- the threshold value is set inside by a minute amount d from a point ⁇ away from the average value as described above (
- the range is normalized for the signal that has been processed to suppress outliers.
- (variance (std) 1 and mean (avg) 0 are normalized, and then the minimum value is 0 and the maximum value is 1 again, and the time series signal is kept in the range of 0 to 1.
- No. 4 is an example in which when an e-mail is acquired as text data, the characters in the e-mail body of the e-mail are converted into a time-series signal, and the time-series signal is normalized so as to converge in the range of 0 to 1. Shown.
- each character of the sentence acquired by the character signalizing unit 122 is digitized based on Unicode (S100).
- the signal normalization unit 124 multiplies the numerical value of the time series signal by an integer to expand the waveform (S102). This is corrected because the characters are adjacent to each other depending on the language.
- the signal normalization unit 124 performs the outlier suppression process as described above (S104). In the outlier suppression process, the numerical value exceeding the threshold value is attenuated, but this attenuation may be performed in a plurality of times (S106). Further, the number of attenuations may be adjusted according to the data.
- the signal normalization unit 124 normalizes the variance and the average, and then normalizes them to the minimum value 0 and the maximum value 1. If the variance value is not below a certain threshold, the processes of steps S104 to S108 are repeated. An upper limit may be set for the number of times of this repetition.
- the signal classification unit 126 receives the normalized time series signal from the signal normalization unit 124, and extracts the features included in the time series signal.
- the extracted feature can reproduce the input, and the signal classification unit 126 learns this feature.
- sentence data described with ordinary expressive features or structural features is learned. For example, a feature is extracted from the normalized input waveform as shown in FIG. 6, and the feature is learned by using the extracted feature so that an output waveform that substantially reproduces the input waveform can be obtained.
- the signal classification unit 216 reduces the dimension of features and suppresses the amount of information by an autoencoder using a neural network.
- FIG. 7 shows the concept of an autoencoder using a neural network.
- the autoencoder is configured with only fully coupled layers, includes four encoder layers and four decoder layers, and the width of each layer of the neural network is variable according to the length of the signal converted from the string. is there.
- the encoder compresses the feature by reducing the unwanted dimensions of the input, and the decoder reproduces the input from the compressed feature.
- the neural network adjusts the weights of the encoder and the decoder by the learning function. In this example, the neural network reproduces the input in a symmetric configuration and the input is fixed length.
- the signal classification unit 126 has a function of inspecting the reproducibility of the output waveform. Specifically, the distance between each point in the two time series of the input waveform and the output waveform as shown in FIG. 6 is brute-forced, and the path in which the distance between the two time series is the shortest is detected. This path becomes the DTW distance (Dynamic Time Warping). Although there are some errors in the reproduced waveform, this inspection is resistant to phase shifts. This DTW distance is used to measure the reproducibility of new data after the training model is determined. The new data here is new sentence data, and the sentence analysis system 100 determines whether or not the sentence is unique.
- DTW distance Dynamic Time Warping
- the signal classification unit 216 calculates a threshold value for classifying waveforms.
- the evaluation data that is, the features compressed by the autoencoder extracted from the sentences described by the usual expressive features and structural features (this is the weight of the autoencoder, for example, the neuron 1).
- Each one appears as a coefficient of the mathematical formula that it has inside) to calculate the identity, obtain the median and standard deviation of the identity, and calculate the threshold value from the following equation.
- This threshold means that when the waveform has a substantially normalized distribution, approximately 95% of the waveform is included in the range of the median value to the standard change ⁇ 2.
- Threshold median-standard deviation x 2
- FIG. 8 shows an example of classification by threshold value.
- the broken line graph is the text of the trained user, and the solid line is the text of another person.
- the threshold value of the feature is 5.8, and a sentence having more features is detected as a sentence of another person.
- the feature storage unit 130 stores features and their thresholds by the feature extraction unit 120. Whenever sentence data is learned, the features and thresholds are updated.
- the peculiar sentence detection unit 140 detects a peculiar sentence by using the learning result after the pre-learning by the feature extraction unit 120 is completed. That is, an arbitrary sentence A is acquired by the sentence acquisition unit 110, and the feature extraction unit 120 extracts the feature of the sentence A.
- the signal classification unit 126 compares the feature extracted from the sentence A with the threshold value stored in the feature storage unit 130, and if the feature is equal to or more than the threshold value, determines the sentence A as a peculiar sentence.
- This determination result is provided to the peculiar sentence detection unit 140, and the peculiar sentence detection unit 140 detects the sentence A determined to be a peculiar sentence as not being a sentence described by ordinary expressive features or structural features. .. For example, it is presumed that the sentence is written by a user other than the person himself / herself, or the sentence is written by a peculiar expressive feature or structural feature by the person himself / herself.
- FIG. 9 shows an example in which the text analysis system of this embodiment is applied to an outgoing mail monitoring system.
- the outgoing mail monitoring system 200 is realized, for example, in a mail server or a client terminal (computer device, mobile device, etc.) having a mail sending / receiving function.
- the outgoing mail monitoring system 200 includes an outgoing mail acquisition unit 210 that acquires an outgoing mail created by a user, a feature extraction unit 220 that extracts the characteristics of the outgoing mail acquired by the outgoing mail acquisition unit 210, and the extracted features.
- the sent mail acquisition unit 110 acquires the HTML format e-mail created by the mail software installed in the client terminal or the e-mail for sending uploaded from the client terminal to the mail server.
- the feature extraction unit 220 operates in the same manner as the feature extraction unit 120 of the sentence analysis system.
- the feature extraction unit 220 has learned in advance the features when the user X describes the e-mail with the usual expressive features and structural features. Therefore, if the outgoing mail acquired from the outgoing mail acquisition unit 210 is described by the user X, the characteristics of the outgoing mail have the same as the learned characteristics, so that the user X is a normal expression. It is identified as an outgoing mail described by features and structural features, but if User X is described by a unique expressive feature or structural feature, or by another person, the characteristics of the outgoing mail. Is not identical to the learned features, so it is identified as being described by User X by a unique expressive or structural feature, or by someone else. Whether or not they have the sameness is determined by whether or not the threshold value is exceeded, as described with reference to FIG.
- the abnormal mail detection unit 240 detects the transmitted mail as an abnormal mail and provides the detection result to the transmission control unit 250.
- the transmission control unit 250 causes, for example, the client terminal or the mail server to stop or suspend the transmission of the transmitted mail, and notifies the user of a warning that the transmitted mail cannot be transmitted.
- the display of the client terminal may be displayed to stop transmission, or voice guidance may be provided.
- the client terminal or mail server is made to send the sent mail.
- FIG. 10 is a flowchart illustrating an operation example of the outgoing mail monitoring system.
- the sent mail is acquired by the sent mail acquisition unit 210 (S200), each character in the body of the sent mail is signalized by the feature extraction unit 220, and a one-dimensional time series signal is generated (S202). Is normalized (S206), and then features are extracted from the time series signal.
- S208 the presence or absence of identity between the extracted features and the learned features was determined (S208), and if there was identity, it was described by the person's usual expressive features and structural features. It is determined to be a sent mail (S210), and the sent mail is sent to the sending address (S212). On the other hand, if there is no identity, it is determined that the sent mail is described by the person's peculiar expressive or structural features or the sent mail described by another person other than the person (S220), and the sent mail is sent. It is stopped (S222).
- the sent mail is described by the usual expressive features and structural features, and the person himself / herself describes the transmission according to the unique expressive features and structural features.
- the transmission of the sent e-mail is stopped, so that information leakage due to an unauthorized sent e-mail can be prevented.
- FIG. 11 shows the probability of being judged as another person in each language.
- the e-mail newsletters B and C are identified with fairly good accuracy, but the e-mail newsletter D has some variations between languages. This is a difference in the characteristics of each language. For example, the number of characters in Japanese is 50 + 50 + lowercase + Chinese characters, English is 26 characters + lower rank, and Chinese and Taiwanese are 87,000 (Unicode11).
- the next experiment evaluates the emails of three employees. Users A and B are sales occupations, respectively, and user C is a quality control engineer occupation.
- the graph of FIG. 12 shows the ratio of whether or not the user A is the person who trained and the users B and C could be detected as others.
- the ratio of detecting user A as another person is 5.95%, and users B and C are described with others (expressive features and structural features).
- the percentages detected as (emails received) were 62.00% and 51.00%, respectively.
- Sentence analysis system 110 Sentence acquisition unit 120: Feature extraction unit 130: Feature storage unit 140: Singular sentence detection unit 200: Sent mail monitoring system 210: Sent mail acquisition unit 220: Feature extraction unit 230: Feature storage unit 240: Abnormal mail detection unit 250: Output control unit
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Artificial Intelligence (AREA)
- Computational Linguistics (AREA)
- General Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- General Health & Medical Sciences (AREA)
- Databases & Information Systems (AREA)
- Data Mining & Analysis (AREA)
- Probability & Statistics with Applications (AREA)
- Information Transfer Between Computers (AREA)
- Machine Translation (AREA)
Abstract
Description
さらに本発明は、メッセージ交換における本文の特異性を検出するメッセージの特徴評価システムの提供を目的とする。
(数2)
閾値=中央値-標準偏差×2
110:文章取得部
120:特徴抽出部
130:特徴記憶部
140:特異文章検出部
200:送信メール監視システム
210:送信メール取得部
220:特徴抽出部
230:特徴記憶部
240:異常メール検出部
250:出力制御部
Claims (15)
- 文章を解析する文章解析システムであって、
文章データを取得する取得手段と、
取得された文章データの文字を数値化することにより文章データを時系列信号に変換する変換手段と、
変換された時系列信号から特徴情報を抽出し、抽出した特徴情報を格納する特徴抽出手段と、
前記特徴情報を用いて新たに取得された文章データの同一性を判定する判定手段と、
を有する文章解析システム。 - 文章解析システムはさらに、前記判定手段の判定結果に基づき前記特徴情報と異なる特異文章を検出する検出手段を有する、請求項1に記載の文章解析システム。
- 前記変換手段は、予め用意された変換テーブルに基づき文字を数値データに変換する、請求項1に記載の文章解析システム。
- 前記変換手段は、前記時系列信号を最小値0と最大値1の範囲内に収まるように正規化する、請求項1または3に記載の文章解析システム。
- 前記変換手段は、設定された閾値を超える前記時系列信号の値を減衰し、減衰した時系列信号を正規化する、請求項1または4に記載の文章解析システム。
- 前記特徴抽出手段は、通常の表現的特徴や構造的特徴で記載された文章データの正規化された時系列信号から特徴を抽出し、抽出した特徴を用いて前記時系列信号の入力波形を再現する出力波形が得られるように特徴を学習する、請求項1または4に記載の文章解析システム。
- 前記特徴抽出手段は、オートエンコーダにより前記特徴情報を符号化する、請求項6に記載の文章解析システム。
- 前記特徴抽出手段は、ニューラルネットワークにより前記特徴情報を学習する、請求項7に記載の文章解析システム。
- 請求項1ないし8に記載の文章解析システムを含むメッセージ交換における特徴評価システムであって、
前記検出手段は、前記判定手段の判定結果に基づき送信メールの異常を検出する、特徴評価システム。 - 特徴評価システムはさらに、送信メールの異常が検出された場合、当該送信メールの送信を停止する送信制御手段を含む、請求項9に記載の特徴評価システム。
- 特徴評価システムはさらに、前記送信制御手段により送信メールの送信が停止されたとき、送信メールの送信停止を通知する通知手段を含む、請求項10に記載の特徴評価システム。
- コンピュータ端末が実行する文章解析プログラムであって、
文章データを取得するステップと、
取得された文章データの文字を数値化することにより文章データを時系列信号に変換するステップと、
変換された時系列信号から特徴情報を抽出し、抽出した特徴情報を格納するステップと、
前記特徴情報を用いて新たに取得された文章データの同一性を判定するステップと、
を有する文章解析プログラム。 - 前記同一性を判定するステップは、前記特徴情報と異なる特異な表現的特徴または構造的特徴で記載された送信メールを識別する、請求項12に記載の文章解析プログラム。
- コンピュータ端末における文章解析方法であって、
文章データを取得するステップと、
取得された文章データの文字を数値化することにより文章データを時系列信号に変換するステップと、
変換された時系列信号から特徴情報を抽出し、抽出した特徴情報を格納するステップと、
前記特徴情報を用いて新たに取得された文章データの同一性を判定するステップと、
を有する文章解析方法。 - 前記同一性を判定するステップは、前記特徴情報と異なる表現的特徴および/または構造的特徴で記載された送信メールを識別する、請求項14に記載の文章解析方法。
Priority Applications (6)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201980099692.0A CN114341822B (zh) | 2019-09-02 | 2019-09-02 | 文章解析系统及使用其的消息交换的特征评价系统 |
| JP2021517726A JP7007693B2 (ja) | 2019-09-02 | 2019-09-02 | 文章解析システムおよびこれを用いたメッセージ交換における特徴評価システム |
| PCT/JP2019/034402 WO2021044475A1 (ja) | 2019-09-02 | 2019-09-02 | 文章解析システムおよびこれを用いたメッセージ交換における特徴評価システム |
| EP19944297.1A EP4027247A4 (en) | 2019-09-02 | 2019-09-02 | TEXT ANALYSIS SYSTEM AND EVALUATION SYSTEM OF THE CHARACTERISTICS FOR MESSAGE EXCHANGE WITH THIS SYSTEM |
| US17/639,866 US20220343067A1 (en) | 2019-09-02 | 2019-09-02 | Text Analysis System, and Characteristic Evaluation System for Message Exchange Using the Same |
| US18/189,819 US12602542B2 (en) | 2019-09-02 | 2023-03-24 | Text analysis system, and characteristic evaluation system for message exchange using the same |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2019/034402 WO2021044475A1 (ja) | 2019-09-02 | 2019-09-02 | 文章解析システムおよびこれを用いたメッセージ交換における特徴評価システム |
Related Child Applications (2)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/639,866 A-371-Of-International US20220343067A1 (en) | 2019-09-02 | 2019-09-02 | Text Analysis System, and Characteristic Evaluation System for Message Exchange Using the Same |
| US18/189,819 Continuation US12602542B2 (en) | 2019-09-02 | 2023-03-24 | Text analysis system, and characteristic evaluation system for message exchange using the same |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2021044475A1 true WO2021044475A1 (ja) | 2021-03-11 |
Family
ID=74852600
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2019/034402 Ceased WO2021044475A1 (ja) | 2019-09-02 | 2019-09-02 | 文章解析システムおよびこれを用いたメッセージ交換における特徴評価システム |
Country Status (5)
| Country | Link |
|---|---|
| US (2) | US20220343067A1 (ja) |
| EP (1) | EP4027247A4 (ja) |
| JP (1) | JP7007693B2 (ja) |
| CN (1) | CN114341822B (ja) |
| WO (1) | WO2021044475A1 (ja) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20250053755A1 (en) * | 2023-08-10 | 2025-02-13 | Dissanayake Gunasekera | Polyglot GPT Engine: Generative Pre-Trained Transformer Based Multilingual Chatbot |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH10227820A (ja) * | 1997-02-12 | 1998-08-25 | Nippon Telegr & Teleph Corp <Ntt> | センサ時間応答補正方法およびセンサ時間応答補正装置 |
| JP2006235949A (ja) * | 2005-02-24 | 2006-09-07 | Nec Corp | 電子メール誤送信監視方法及びシステム |
| JP2011081627A (ja) * | 2009-10-07 | 2011-04-21 | Kddi R & D Laboratories Inc | 特徴量算出装置、品詞推定装置およびプログラム |
| WO2017094202A1 (ja) * | 2015-12-01 | 2017-06-08 | アイマトリックス株式会社 | 画像処理を応用した文書構造解析装置 |
| US10104029B1 (en) * | 2011-11-09 | 2018-10-16 | Proofpoint, Inc. | Email security architecture |
| JP2019105979A (ja) * | 2017-12-12 | 2019-06-27 | 株式会社Ihi | 予測システム、予測方法、および予測プログラム |
Family Cites Families (55)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7092496B1 (en) * | 2000-09-18 | 2006-08-15 | International Business Machines Corporation | Method and apparatus for processing information signals based on content |
| US7421418B2 (en) * | 2003-02-19 | 2008-09-02 | Nahava Inc. | Method and apparatus for fundamental operations on token sequences: computing similarity, extracting term values, and searching efficiently |
| US7483947B2 (en) * | 2003-05-02 | 2009-01-27 | Microsoft Corporation | Message rendering for identification of content features |
| JP2006092346A (ja) * | 2004-09-24 | 2006-04-06 | Fuji Xerox Co Ltd | 文字認識装置、文字認識方法および文字認識プログラム |
| US9529984B2 (en) * | 2005-10-26 | 2016-12-27 | Cortica, Ltd. | System and method for verification of user identification based on multimedia content elements |
| US20080084972A1 (en) * | 2006-09-27 | 2008-04-10 | Michael Robert Burke | Verifying that a message was authored by a user by utilizing a user profile generated for the user |
| CN101500028A (zh) * | 2008-01-28 | 2009-08-05 | 英华达(上海)电子有限公司 | 采用读写模式的通信终端以及实现读写模式通信的方法 |
| WO2010037163A1 (en) * | 2008-09-30 | 2010-04-08 | National Ict Australia Limited | Measuring cognitive load |
| DE112011102691B4 (de) * | 2010-08-09 | 2019-05-09 | Charter Ip Llc | Verfahren und System zur Umwandlung von Zeitsteuerungsberichten in Zeitsteuerungswellenformen |
| KR101060639B1 (ko) * | 2010-12-21 | 2011-08-31 | 한국인터넷진흥원 | 자바스크립트 난독화 강도 분석을 통한 악성 의심 웹사이트 탐지 시스템 및 그 탐지방법 |
| WO2013008778A1 (ja) * | 2011-07-11 | 2013-01-17 | Mizunuma Takeshi | 識別名管理方法およびシステム |
| US11195057B2 (en) * | 2014-03-18 | 2021-12-07 | Z Advanced Computing, Inc. | System and method for extremely efficient image and pattern recognition and artificial intelligence platform |
| US11074495B2 (en) * | 2013-02-28 | 2021-07-27 | Z Advanced Computing, Inc. (Zac) | System and method for extremely efficient image and pattern recognition and artificial intelligence platform |
| JP5596649B2 (ja) * | 2011-09-26 | 2014-09-24 | 株式会社東芝 | 文書マークアップ支援装置、方法、及びプログラム |
| US20130091266A1 (en) * | 2011-10-05 | 2013-04-11 | Ajit Bhave | System for organizing and fast searching of massive amounts of data |
| US9354725B2 (en) * | 2012-06-01 | 2016-05-31 | New York University | Tracking movement of a writing instrument on a general surface |
| US8762302B1 (en) * | 2013-02-22 | 2014-06-24 | Bottlenose, Inc. | System and method for revealing correlations between data streams |
| US9497023B1 (en) * | 2013-03-14 | 2016-11-15 | Amazon Technologies, Inc. | Multiply-encrypted message for filtering |
| US10599697B2 (en) * | 2013-03-15 | 2020-03-24 | Uda, Llc | Automatic topic discovery in streams of unstructured data |
| US10430111B2 (en) * | 2013-03-15 | 2019-10-01 | Uda, Llc | Optimization for real-time, parallel execution of models for extracting high-value information from data streams |
| US10698935B2 (en) * | 2013-03-15 | 2020-06-30 | Uda, Llc | Optimization for real-time, parallel execution of models for extracting high-value information from data streams |
| EP2973042A4 (en) * | 2013-03-15 | 2016-11-09 | Uda Llc | HIERARCHICAL PARALLEL MODELS FOR REAL-TIME EXTRACTION OF HIGH-QUALITY INFORMATION FROM DATA STREAMS AND SYSTEM AND METHOD FOR MANUFACTURING THE SAME |
| US10204026B2 (en) * | 2013-03-15 | 2019-02-12 | Uda, Llc | Realtime data stream cluster summarization and labeling system |
| US20150242856A1 (en) * | 2014-02-21 | 2015-08-27 | International Business Machines Corporation | System and Method for Identifying Procurement Fraud/Risk |
| US9485209B2 (en) * | 2014-03-17 | 2016-11-01 | International Business Machines Corporation | Marking of unfamiliar or ambiguous expressions in electronic messages |
| US20150356571A1 (en) * | 2014-06-05 | 2015-12-10 | Adobe Systems Incorporated | Trending Topics Tracking |
| JP6762963B2 (ja) * | 2015-06-02 | 2020-09-30 | ライブパーソン, インコーポレイテッド | 一貫性重み付けおよびルーティング規則に基づく動的通信ルーティング |
| US10110531B2 (en) * | 2015-06-11 | 2018-10-23 | International Business Machines Corporation | Electronic rumor cascade management in computer network communications |
| US20170083817A1 (en) * | 2015-09-23 | 2017-03-23 | Isentium, Llc | Topic detection in a social media sentiment extraction system |
| JP6453202B2 (ja) * | 2015-10-30 | 2019-01-16 | 日本電産サンキョー株式会社 | 相互認証装置及び相互認証方法 |
| EP3403187A4 (en) * | 2016-01-14 | 2019-07-31 | Sumo Logic | SINGLE CLICK DELTA ANALYSIS |
| US20170250939A1 (en) * | 2016-02-28 | 2017-08-31 | Liquid Grids | Social dialogue listening, analytics, and engagement system and method |
| US20180025303A1 (en) * | 2016-07-20 | 2018-01-25 | Plenarium Inc. | System and method for computerized predictive performance analysis of natural language |
| US10796217B2 (en) * | 2016-11-30 | 2020-10-06 | Microsoft Technology Licensing, Llc | Systems and methods for performing automated interviews |
| US10133865B1 (en) * | 2016-12-15 | 2018-11-20 | Symantec Corporation | Systems and methods for detecting malware |
| US11580350B2 (en) * | 2016-12-21 | 2023-02-14 | Microsoft Technology Licensing, Llc | Systems and methods for an emotionally intelligent chat bot |
| US20180203851A1 (en) * | 2017-01-13 | 2018-07-19 | Microsoft Technology Licensing, Llc | Systems and methods for automated haiku chatting |
| US10403287B2 (en) * | 2017-01-19 | 2019-09-03 | International Business Machines Corporation | Managing users within a group that share a single teleconferencing device |
| US11216895B1 (en) * | 2017-03-30 | 2022-01-04 | Dividex Analytics, LLC | Securities claims identification, optimization and recovery system and methods |
| US10311454B2 (en) * | 2017-06-22 | 2019-06-04 | NewVoiceMedia Ltd. | Customer interaction and experience system using emotional-semantic computing |
| US20190028509A1 (en) * | 2017-07-20 | 2019-01-24 | Barracuda Networks, Inc. | System and method for ai-based real-time communication fraud detection and prevention |
| US11157831B2 (en) * | 2017-10-02 | 2021-10-26 | International Business Machines Corporation | Empathy fostering based on behavioral pattern mismatch |
| EP3788512A4 (en) * | 2017-12-30 | 2022-03-09 | Target Brands, Inc. | Hierarchical, parallel models for extracting in real time high-value information from data streams and system and method for creation of same |
| US11164239B2 (en) * | 2018-03-12 | 2021-11-02 | Ebay Inc. | Method, system, and computer-readable storage medium for heterogeneous data stream processing for a smart cart |
| US10817604B1 (en) * | 2018-06-19 | 2020-10-27 | Architecture Technology Corporation | Systems and methods for processing source codes to detect non-malicious faults |
| US11025649B1 (en) * | 2018-06-26 | 2021-06-01 | NortonLifeLock Inc. | Systems and methods for malware classification |
| CN108932220A (zh) * | 2018-06-29 | 2018-12-04 | 北京百度网讯科技有限公司 | 文章生成方法和装置 |
| US20220059083A1 (en) * | 2018-12-10 | 2022-02-24 | Interactive-Ai, Llc | Neural modulation codes for multilingual and style dependent speech and language processing |
| US11178170B2 (en) * | 2018-12-14 | 2021-11-16 | Ca, Inc. | Systems and methods for detecting anomalous behavior within computing sessions |
| US20200202280A1 (en) * | 2018-12-24 | 2020-06-25 | Level35 Pty Ltd | System and method for using natural language processing in data analytics |
| US11373243B2 (en) * | 2019-02-11 | 2022-06-28 | TD Ameritrade IP Compnay, Inc. | Time-series pattern matching system |
| US20240232539A1 (en) * | 2019-02-21 | 2024-07-11 | Charlee.Ai. Inc. | Systems and Methods for Insights Extraction Using Semantic Search |
| WO2020231453A1 (en) * | 2019-05-13 | 2020-11-19 | Google Llc | Automatic evaluation of natural language text generated based on structured data |
| CN114026557A (zh) * | 2019-07-04 | 2022-02-08 | 松下知识产权经营株式会社 | 说话解析装置、说话解析方法以及程序 |
| US11556992B2 (en) * | 2019-08-14 | 2023-01-17 | Royal Bank Of Canada | System and method for machine learning architecture for enterprise capitalization |
-
2019
- 2019-09-02 EP EP19944297.1A patent/EP4027247A4/en not_active Ceased
- 2019-09-02 US US17/639,866 patent/US20220343067A1/en not_active Abandoned
- 2019-09-02 CN CN201980099692.0A patent/CN114341822B/zh active Active
- 2019-09-02 JP JP2021517726A patent/JP7007693B2/ja active Active
- 2019-09-02 WO PCT/JP2019/034402 patent/WO2021044475A1/ja not_active Ceased
-
2023
- 2023-03-24 US US18/189,819 patent/US12602542B2/en active Active
Patent Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH10227820A (ja) * | 1997-02-12 | 1998-08-25 | Nippon Telegr & Teleph Corp <Ntt> | センサ時間応答補正方法およびセンサ時間応答補正装置 |
| JP2006235949A (ja) * | 2005-02-24 | 2006-09-07 | Nec Corp | 電子メール誤送信監視方法及びシステム |
| JP2011081627A (ja) * | 2009-10-07 | 2011-04-21 | Kddi R & D Laboratories Inc | 特徴量算出装置、品詞推定装置およびプログラム |
| US10104029B1 (en) * | 2011-11-09 | 2018-10-16 | Proofpoint, Inc. | Email security architecture |
| WO2017094202A1 (ja) * | 2015-12-01 | 2017-06-08 | アイマトリックス株式会社 | 画像処理を応用した文書構造解析装置 |
| JP6267830B2 (ja) | 2015-12-01 | 2018-01-24 | アイマトリックス株式会社 | 画像処理を応用した文書構造解析装置 |
| JP2019105979A (ja) * | 2017-12-12 | 2019-06-27 | 株式会社Ihi | 予測システム、予測方法、および予測プログラム |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP4027247A4 |
Also Published As
| Publication number | Publication date |
|---|---|
| JP7007693B2 (ja) | 2022-01-25 |
| US20230237258A1 (en) | 2023-07-27 |
| EP4027247A1 (en) | 2022-07-13 |
| CN114341822A (zh) | 2022-04-12 |
| EP4027247A4 (en) | 2023-05-10 |
| JPWO2021044475A1 (ja) | 2021-09-27 |
| US20220343067A1 (en) | 2022-10-27 |
| US12602542B2 (en) | 2026-04-14 |
| CN114341822B (zh) | 2022-12-02 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Ismail et al. | Efficient E‐Mail Spam Detection Strategy Using Genetic Decision Tree Processing with NLP Features | |
| CN105577660B (zh) | 基于随机森林的dga域名检测方法 | |
| US8489689B1 (en) | Apparatus and method for obfuscation detection within a spam filtering model | |
| WO2020230053A1 (en) | Detection of phishing campaigns | |
| US8112484B1 (en) | Apparatus and method for auxiliary classification for generating features for a spam filtering model | |
| CN111031026A (zh) | 一种dga恶意软件感染主机检测方法 | |
| CN117725161A (zh) | 文本中变种词的识别及提取敏感词的方法和系统 | |
| CN112948725A (zh) | 基于机器学习的钓鱼网站url检测方法及系统 | |
| CN113905016A (zh) | 一种dga域名检测方法、检测装置及计算机存储介质 | |
| CN112528682B (zh) | 语种检测方法、装置、电子设备和存储介质 | |
| Li et al. | Unbalanced network attack traffic detection based on feature extraction and GFDA-WGAN | |
| CN114372461B (zh) | 一种隐性关键词提取方法、终端设备及存储介质 | |
| JP2012088803A (ja) | 悪性ウェブコード判別システム、悪性ウェブコード判別方法および悪性ウェブコード判別用プログラム | |
| CN117376307B (zh) | 域名处理方法、装置及设备 | |
| CN115422926A (zh) | 模型训练方法、文本分类方法、装置、介质及电子设备 | |
| US12602542B2 (en) | Text analysis system, and characteristic evaluation system for message exchange using the same | |
| CN109284465B (zh) | 一种基于url的网页分类器构建方法及其分类方法 | |
| US11936686B2 (en) | System, device and method for detecting social engineering attacks in digital communications | |
| CN115766212B (zh) | 一种基于url多角度特征的钓鱼网站检测方法 | |
| CN101329668A (zh) | 一种信息规则生成方法及装置、信息类型判断方法及系统 | |
| CN120567440A (zh) | 一种多模态检测模型的训练方法、检测方法及装置 | |
| CN117749496B (zh) | 一种基于邮件的风险提示信息生成方法、装置及介质 | |
| Gaurav et al. | Defense-optimized BERT architecture against digital arrest attacks in intelligent systems: A. Gaurav et al. | |
| CN118761031A (zh) | 一种ai自动审核文案及图片的语义识别方法及系统 | |
| CN112771524A (zh) | 基于模糊包含的伪装检测 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| ENP | Entry into the national phase |
Ref document number: 2021517726 Country of ref document: JP Kind code of ref document: A |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19944297 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2019944297 Country of ref document: EP Effective date: 20220404 |
|
| WWW | Wipo information: withdrawn in national office |
Ref document number: 2019944297 Country of ref document: EP |
