JPH11242679A - Method and apparatus for classifying information based on user's interest, and recording medium storing program for classifying information based on user's interest - Google Patents

Method and apparatus for classifying information based on user's interest, and recording medium storing program for classifying information based on user's interest

Info

Publication number
JPH11242679A
JPH11242679A JP10043620A JP4362098A JPH11242679A JP H11242679 A JPH11242679 A JP H11242679A JP 10043620 A JP10043620 A JP 10043620A JP 4362098 A JP4362098 A JP 4362098A JP H11242679 A JPH11242679 A JP H11242679A
Authority
JP
Japan
Prior art keywords
information
classification
keywords
user
keyword
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP10043620A
Other languages
Japanese (ja)
Inventor
Sachiko Iori
祥子 庵
Hideaki Suzuki
英明 鈴木
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Inc
Original Assignee
Nippon Telegraph and Telephone Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Telegraph and Telephone Corp filed Critical Nippon Telegraph and Telephone Corp
Priority to JP10043620A priority Critical patent/JPH11242679A/en
Publication of JPH11242679A publication Critical patent/JPH11242679A/en
Pending legal-status Critical Current

Links

Landscapes

  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

(57)【要約】 【課題】 各利用者の興味に基づいて情報の検索を行う
とともに、利用者の興味が動的に変化していってもその
ときの興味を反映して情報の分類を自動的に行う。 【解決手段】 全体量が把握できる全ての情報から抽出
したキーワードに対応する情報の数と、そのキーワード
に対応する利用者がその範囲内で閲覧した情報の数を比
較し、全てのキーワードについて興味度合いを計算し、
興味度合いが大きいキーワードから順に上位N個のキー
ワードについて分類のカテゴリを作成する。キーワード
の全ての2つの組み合わせについてもキーワードの関連
度合いを計算し、関連度合いが高かった2つのキーワー
ドをグループ化して、分類カテゴリの上位階層を作成す
る。
(57) [Summary] [Problem] In addition to searching for information based on each user's interest, even if the user's interest is dynamically changing, information classification is performed by reflecting the interest at that time. Do it automatically. SOLUTION: The number of pieces of information corresponding to a keyword extracted from all pieces of information whose overall amount can be grasped is compared with the number of pieces of information browsed by a user corresponding to the keyword, and the user is interested in all the keywords. Calculate the degree,
Classification categories are created for the top N keywords in descending order of interest. The degree of relevance of keywords is calculated for all two combinations of keywords, and the two keywords having high degrees of relevance are grouped to create an upper layer of the classification category.

Description

【発明の詳細な説明】DETAILED DESCRIPTION OF THE INVENTION

【0001】[0001]

【発明の属する技術分野】本発明は、利用者の興味に基
づいて情報を分類する方法に関する。
[0001] The present invention relates to a method for classifying information based on a user's interest.

【0002】[0002]

【従来の技術】システムによる情報分類は、個々の利用
者にとってわかりやすい分類のカテゴリを生成するので
はなく、むしろどの利用者にも共通するような分類カテ
ゴリをシステムが初めから与え、その分類のカテゴリに
合わせて情報を分類するというアプローチであった。こ
れは個々の利用者の興味を反映していないので、分類さ
れた情報を利用するため、利用者は分類されている情報
を検索しなければならないなどの労力を必要とした。
2. Description of the Related Art Information classification by a system does not generate classification categories that are easy to understand for individual users, but rather, the system gives classification categories that are common to all users from the beginning, and the classification categories are classified. The approach was to classify the information according to. Since this does not reflect the individual user's interest, the user needs to search for the classified information in order to use the classified information.

【0003】これに対し、あらかじめ利用者が情報分類
をしてくれるシステムに対して自分の思い込みの興味を
提示して、それに基づいてシステムが分類のカテゴリを
生成し、情報を分類するというアプローチもあった。
On the other hand, there is also an approach in which a user presents his / her belief interest to a system that classifies information in advance, and the system generates a category of classification based on the interest and classifies the information. there were.

【0004】[0004]

【発明が解決しようとする課題】このアプローチでは、
利用者が最初に自分の思い込みの興味を提示するために
労力が必要となることや、利用者の興味が変化してもシ
ステム側が利用している分類カテゴリが固定しているの
で興味の変化に追従できないなどの問題点があった。
With this approach,
Since the user needs effort to present his / her own belief at first, and even if the interest of the user changes, the classification category used by the system is fixed, so the interest may change. There were problems such as inability to follow.

【0005】本発明の目的は、各利用者の興味に基づい
て個々に適した情報分類を行うことができ、また利用者
の興味が動的に変化していってもそのときの興味を反映
した情報分類を自動的に行うことができる、利用者の興
味に基づいて情報を分類する方法、装置、および情報分
類プログラムを記録した記録媒体を提供することであ
る。
SUMMARY OF THE INVENTION It is an object of the present invention to classify information individually based on the interests of each user, and to reflect the interests at the time even if the interests of the users change dynamically. It is an object of the present invention to provide a method and apparatus for classifying information based on a user's interest, and a recording medium on which an information classification program is recorded, which can automatically perform the classified information.

【0006】[0006]

【課題を解決するための手段】本発明の利用者の興味に
基づいて情報を分類する方法は、全体量が把握できるあ
る情報群のうちで利用者が閲覧した情報から抽出され
た、当該情報の特徴を表す各キーワードに対応する情報
の数を、前記情報群から抽出された、当該情報の特徴を
表す当該キーワードに対応する情報の数と比較し、後者
の情報の数に対する前者の情報の数の割合が大きいキー
ワードに1対1に対応させて、情報の分類に利用する入
れ物である分類カテゴリを生成して情報の分類を行う分
類カテゴリ生成段階と、それらの分類カテゴリのうちで
関連度合いが高い分類カテゴリをグループ化し、上位階
層を作成する上位階層作成段階を有する。
SUMMARY OF THE INVENTION According to the present invention, there is provided a method for classifying information based on a user's interest, wherein the information extracted from information browsed by a user in a group of information whose total amount can be grasped. The number of information corresponding to each keyword representing the characteristic of the information is compared with the number of information corresponding to the keyword extracted from the information group and representing the characteristic of the information. A classification category generation step of generating a classification category, which is a container used for information classification, in a one-to-one correspondence with a keyword having a large number, and classifying information, and a degree of association among the classification categories The method includes a higher-level creation step of grouping the classification categories having higher ranks and creating an upper-level hierarchy.

【0007】また、本発明の利用者の興味に基づいて情
報を分類する装置は、全体量が把握できるある情報群の
うち利用者が閲覧した情報から抽出された、当該情報の
特徴を表す各キーワードに対応する情報の数を、前記情
報群から抽出された、当該情報の特徴を表す当該キーワ
ードに対応する情報の数と比較し、後者の情報の数に対
する前者の情報の数の割合が大きいキーワードに1対1
に対応させて、情報の分類に利用する入れ物である分類
カテゴリを生成して情報の分類を行う分類カテゴリ生成
手段と、それらの分類カテゴリのうちで関連度合いが高
い分類カテゴリをグループ化し、上位階層を作成する上
位階層作成手段を有する。
Further, the apparatus for classifying information based on the user's interest according to the present invention is a device for extracting characteristics of the information extracted from information viewed by the user from a certain information group whose total amount can be grasped. The number of information corresponding to the keyword is compared with the number of information corresponding to the keyword extracted from the information group and representing the characteristic of the information, and the ratio of the number of the former information to the number of the latter information is large. One to one for keywords
A classification category generating means for generating a classification category, which is a container used for classifying information, and classifying information; and grouping classification categories having a high degree of association among the classification categories to form a higher hierarchy Is created.

【0008】また、本発明の、利用者の興味に基づいて
情報を分類するプログラムを記録した記録媒体は、全体
量が把握できるある情報群のうち利用者が閲覧した情報
から抽出された、当該情報の特徴を表す各キーワードに
対応する情報の数を、前記情報群から抽出された、当該
情報の特徴を表す当該キーワードに対応する情報の数と
比較し、後者の情報の数に対する前者の情報の数の割合
が大きいキーワードに1対1に対応させて情報の分類に
利用する入れ物である分類カテゴリを生成して情報の分
類を行う分類カテゴリ生成処理と、それらの分類カテゴ
リの関連度合いが高い場合には、さらに分類カテゴリを
グループ化し、上位階層を作成する上位階層生成処理を
コンピュータに実行させるための、利用者の興味に基づ
いて情報を分類するプログラムを格納している。
[0008] A recording medium according to the present invention, in which a program for classifying information based on the user's interest is recorded, the information extracted from the information viewed by the user from a certain information group whose overall amount can be grasped. The number of information corresponding to each keyword representing the feature of the information is compared with the number of information corresponding to the keyword representing the feature of the information extracted from the information group, and the former information is compared with the number of the latter information. A category that is a container used for information classification in a one-to-one correspondence with keywords having a large ratio of the number of keywords, and classifies the information, and the degree of relevance of the classification categories is high. In such a case, the classification categories are further grouped, and the information is classified based on the user's interest in order for the computer to execute an upper layer generation process of creating an upper layer. That contains the program.

【0009】ある一定範囲内の全ての情報から抽出した
キーワードと、その範囲内で利用者が閲覧した情報から
抽出したキーワードを比較し、利用者が高い割合で閲覧
している情報に対応するキーワードに基づいて、そのキ
ーワードに1対1で対応するような分類カテゴリの生成
を自動的に行う。この際分類カテゴリに対応している高
い割合で閲覧している情報の複数のキーワードの関連度
合いを計算して、関連度合いが高かった場合はそれらの
キーワードに対応する分類カテゴリをグループ化するよ
うな上位階層の分類カテゴリを自動的に生成し、分類の
階層化を行う。
[0009] A keyword extracted from all information within a certain range is compared with a keyword extracted from information viewed by a user within the certain range, and a keyword corresponding to information viewed by the user at a high rate is determined. Automatically generates a classification category that corresponds to the keyword on a one-to-one basis. At this time, the degree of relevance of a plurality of keywords of the information viewed at a high rate corresponding to the classification category is calculated, and if the degree of relevance is high, the classification categories corresponding to those keywords are grouped. The category of the upper hierarchy is automatically generated and the classification is hierarchized.

【0010】したがって、全体量が把握できる情報の中
から利用者が情報を閲覧することによって、利用者があ
らかじめ分類カテゴリを用意しなくても、利用者ごとの
興味に合わせた分類カテゴリを自動的に生成して情報を
分類することができる。また、時間の経過に伴って利用
者が新しい情報を閲覧すると、利用者の興味の変化に合
わせて自動的に新しい分類カテゴリを生成したり、不要
な分類カテゴリを消去することができる。
[0010] Therefore, when a user browses information from among information whose total amount can be grasped, a classification category according to the interest of each user can be automatically set without the user preparing a classification category in advance. And the information can be classified. Further, when the user browses new information as time elapses, a new classification category can be automatically generated according to a change in the user's interest, or unnecessary classification categories can be deleted.

【0011】[0011]

【発明の実施の形態】次に、本発明の実施の形態につい
て図面を参照して説明する。
Next, embodiments of the present invention will be described with reference to the drawings.

【0012】図1は本発明の一実施形態のシステムの概
略構成を示すもので、情報が格納されている情報サーバ
1と、利用者が情報を閲覧する際に利用する、パーソナ
ルコンピュータなどの端末2と、これらを任意に接続す
るコンピュータネットワーク3からなる。
FIG. 1 shows a schematic configuration of a system according to an embodiment of the present invention. An information server 1 in which information is stored, and a terminal such as a personal computer used when a user browses the information. 2 and a computer network 3 for arbitrarily connecting them.

【0013】本実施形態の利用者の興味に基づいて情報
を分類する方法は情報サーバ1と利用者端末2で行なわ
れる前処理と、情報サーバ1で行なわれる本処理からな
る。
The method for classifying information based on the user's interest according to the present embodiment includes a pre-process performed by the information server 1 and the user terminal 2 and a main process performed by the information server 1.

【0014】まず、前処理を図3と図4により説明す
る。
First, the preprocessing will be described with reference to FIGS.

【0015】図3(1)は情報サーバ1側の処理の流れ
を示すものである。まず、情報サーバ1内に情報(図4
(1))を登録する際に各情報に情報ID(図4
(2))をつける(ステップ11)。次に、情報サーバ
1内の各情報についてその情報の特徴を表すようなキー
ワード群(図4(3))を抽出する(ステップ12)。
また、このキーワード群をそれぞれの情報に添付してお
く(ステップ13、図4(4))。最後に、情報サーバ
1内の全ての情報に対して「あるキーワードとそのキー
ワードを含む情報の数とその情報ID」(図4(5))
を表形式で情報サーバ1に保存する(ステップ14)。
FIG. 3A shows the flow of processing on the information server 1 side. First, information (FIG. 4) is stored in the information server 1.
When registering (1)), information ID (FIG. 4) is assigned to each information.
(2)) is attached (step 11). Next, for each piece of information in the information server 1, a keyword group (FIG. 4 (3)) that represents the feature of the information is extracted (step 12).
Also, this keyword group is attached to each piece of information (step 13, FIG. 4 (4)). Finally, for all information in the information server 1, "a certain keyword, the number of information including the keyword, and the information ID" (FIG. 4 (5))
Is stored in the information server 1 in a table format (step 14).

【0016】図3(2)は端末2側の処理の流れを示す
ものである。利用者が端末2を利用し、コンピュータネ
ットワーク3を介して情報サーバ1内の情報を閲覧す
る。このとき、どの情報を閲覧したかという情報を端末
2に保存する(ステップ21)。この際、閲覧された各
情報に添付されているキーワード群についても保存する
(図4(6))。次に、情報サーバ1内で利用者に閲覧
された全ての情報に対して「あるキーワードとそのキー
ワードを含む情報の数とその情報ID」(図4(7))
を表形式で端末2に保存する(ステップ22)。
FIG. 3B shows the flow of processing on the terminal 2 side. A user browses information in the information server 1 via the computer network 3 using the terminal 2. At this time, information indicating which information has been browsed is stored in the terminal 2 (step 21). At this time, a keyword group attached to each piece of information viewed is also stored (FIG. 4 (6)). Next, "a certain keyword, the number of information including the keyword, and the information ID" are displayed for all the information browsed by the user in the information server 1 (FIG. 4 (7)).
Is stored in the terminal 2 in a table format (step 22).

【0017】次に、本処理について図2と図3(3)に
より説明する。本処理では、(I)キーワードに1対1
で対応する分類のカテゴリの生成、(II)そのカテゴ
リをグループ化することによる上位階層の生成、の2つ
の処理をそれぞれ分類カテゴリ生成部5、上位階層作成
部6で行う。
Next, this processing will be described with reference to FIGS. 2 and 3 (3). In this processing, (I) one-to-one for the keyword
, And (II) generation of an upper hierarchy by grouping the categories, respectively, are performed by the classification category generation unit 5 and the upper hierarchy creation unit 6, respectively.

【0018】(I)キーワードに1対1に対応する分類
カテゴリの生成 分類カテゴリの生成の基準は、あるキーワードに対して
利用者がどれだけ興味を持っているかという興味度合い
とする。興味度合いの計算には先に述べた「情報サーバ
1内の全ての情報に対する『あるキーワードとそのキー
ワードを含む情報の数とその情報ID』」と、「情報サ
ーバ1内で利用者に閲覧された全ての情報に対する『あ
るキーワードとそのキーワードを含む情報の数とその情
報ID』」を利用する。「情報サーバ1内で利用者に閲
覧された全ての情報に対する『あるキーワードを含む情
報の数』」を「情報サーバ1内の全ての情報に対する
『そのキーワードを含む情報の数』」で除算した結果を
興味度合いE(0≦E≦1)とする。全てのキーワード
について興味度合いEの計算を行い(ステップ31)、
値が大きいキーワードから順に上位N個(Nは可変で整
数)のキーワードについて、キーワードと1対1で対応
する分類のカテゴリを生成する(ステップ32)。
(I) Generation of Classification Categories Corresponding to Keywords One-to-one is based on the degree of interest indicating how much a user is interested in a certain keyword. The calculation of the degree of interest includes “a certain keyword and the number of pieces of information including the keyword and the information ID” for all pieces of information in the information server 1 described above, and “the information browsed by the user in the information server 1”. "A keyword and the number of information including the keyword and its information ID" for all the information. "The number of information including a certain keyword for all information viewed by the user in the information server 1" is divided by the "number of information including the keyword for all information in the information server 1". Let the result be the degree of interest E (0 ≦ E ≦ 1). The degree of interest E is calculated for all keywords (step 31),
For the top N keywords (N is a variable and an integer) in order from the keyword having the largest value, the category of the classification corresponding to the keyword in one-to-one correspondence is generated (step 32).

【0019】(II)(I)で生成されたカテゴリを利用
してグループ化を行うことによる上位階層の生成 先の操作で生成された分類のカテゴリをもとに、分類の
カテゴリの階層化を行う。分類のカテゴリの階層化の基
準は、利用者がキーワードAとキーワードBがどれだけ
関連があると考えているかをはかる関連度合いとする。
関連度合いの計算には先に述べた「情報サーバ内で利用
者に閲覧された全ての情報に対する『あるキーワードと
そのキーワードを含む情報の情報ID』」のキーワード
AおよびキーワードBに関する情報を利用する。「情報
サーバ内で利用者に閲覧された全ての情報に対する『キ
ーワードAを含む情報』の集合」と「情報サーバ内で利
用者に閲覧された全ての情報に対する『キーワードBを
含む情報』の集合」の重なりが大きければ、関連度合い
の重みづけが高くなるようになっている。2つのキーワ
ードA,Bの「情報サーバ内で利用者に閲覧された全て
の情報に対する『そのキーワードを含む情報』」の集合
の重なり(数)を「情報サーバ内で利用者に閲覧された
全ての情報に対する『キーワードAを含む情報の数』、
キーワードBを含む情報の数』」でそれぞれ除算した結
果を関連度合いとする(ステップ33)。この関連度合
いの少なくとも一方が閾値Xよりも高かった2つのキー
ワードはグループ化し、分類カテゴリの上位階層を生成
する(ステップ34)。3つ以上のキーワードについ
て、例えばキーワードA、キーワードB、キーワードC
について、キーワードAとキーワードB、キーワードB
とキーワードC、キーワードAとキーワードCのそれぞ
れの関連度合いが閾値Xよりも高い場合はこれら3つの
キーワードから生成された分類のカテゴリをグループ化
し、上位階層を生成する。
(II) Generation of an upper hierarchy by performing grouping using the categories generated in (I) Based on the classification categories generated by the previous operation, classification of the classification categories is performed. Do. The criterion for classifying the categories into categories is a degree of association that measures how much the user considers the keywords A and B to be related.
For the calculation of the degree of relevance, the information regarding the keywords A and B of the above-mentioned "information ID of a certain keyword and information including the keyword for all information browsed by the user in the information server" is used. . A set of "information including keyword A" for all information viewed by the user in the information server and a set of "information including keyword B" for all information viewed by the user in the information server Is greater, the degree of association is weighted higher. The overlap (number) of the set of “information including the keyword” for all the information browsed by the user in the information server of the two keywords A and B is set to “all the information browsed by the user in the information server”. "Number of information containing keyword A" for the information of
The number of pieces of information including the keyword B "is determined as the degree of association (step 33). Two keywords for which at least one of the degrees of relevance is higher than the threshold value X are grouped to generate an upper layer of the classification category (step 34). For three or more keywords, for example, keyword A, keyword B, keyword C
About keyword A and keyword B, keyword B
If the degree of association between the keyword A and the keyword A is higher than the threshold X, the categories of the classifications generated from these three keywords are grouped to generate an upper layer.

【0020】利用者が閲覧した情報の増加に併せて、分
類カテゴリの生成と分類カテゴリの階層化を定期的、か
つ自動的に行う。この際、以前に生成された分類のカテ
ゴリであっても興味度合いが相対的に低くなり、例えば
上位20個に入らなくなったキーワードに対応する分類
のカテゴリは消去する。また、以前は生成されなかった
分類のカテゴリであっても興味度合いが相対的に高くな
り、上位20個に入ったキーワードに対応する分類カテ
ゴリは新しく生成する。カテゴリの階層化についても同
じように消去、生成を行う。
In accordance with an increase in the information browsed by the user, the generation of the classification categories and the hierarchical classification of the classification categories are periodically and automatically performed. At this time, even if the category is a category that has been generated before, the degree of interest is relatively low, and, for example, the category of the category corresponding to the keyword that is no longer included in the top 20 categories is deleted. In addition, even if the category is a category that has not been generated before, the degree of interest is relatively high, and a category that corresponds to the top 20 keywords is newly generated. The deletion and generation of categories are performed in the same manner.

【0021】この方法を利用することによって、生成さ
れた分類のカテゴリおよび階層を利用者に提示する際
は、生成した分類のカテゴリごと分けて提示することが
できる。このときそのカテゴリに分類されている情報を
利用者が既に閲覧した情報とまだ閲覧していない情報に
分け、閲覧していない情報を推薦する形で提供する。
By using this method, when the categories and hierarchies of the generated classifications are presented to the user, they can be presented separately for each category of the generated classifications. At this time, the information classified into the category is divided into information that the user has already browsed and information that has not been browsed, and the information that has not been browsed is provided in a recommended form.

【0022】また、この方法を利用することによって利
用者の興味構造を把握することができるので、興味構造
の似た利用者同士を引き合わせてコミュニティを生成す
ることができる。
Further, since the user's interest structure can be grasped by using this method, a community can be created by drawing together users having similar interest structures.

【0023】なお、図2の処理部は端末2に設けてもよ
い。また、以上説明した分類カテゴリの生成と消去、分
類カテゴリの階層化と消去の処理は情報分類プログラム
として、FD、CD−ROM、半導体メモリなどの記録
媒体に記録しておき、コンピュータにより読み込んで実
行することもできる。
The processing section shown in FIG. 2 may be provided in the terminal 2. The above-described processes of generating and erasing the classification categories and hierarchizing and erasing the classification categories are recorded as information classification programs on a recording medium such as an FD, a CD-ROM, or a semiconductor memory, and are read and executed by a computer. You can also.

【0024】[0024]

【発明の効果】以上説明したように、本発明によれば、
同じ情報サーバから複数の利用者が情報の閲覧を行って
いるものとすると、各利用者の興味に基づいて個々に適
した情報分類を行うことができる。また、利用者の興味
が動的に変化していってもそのときの興味を反映した情
報分類を自動的に行うことが可能になる。
As described above, according to the present invention,
Assuming that a plurality of users are browsing information from the same information server, it is possible to perform information classification individually suitable for each user based on their interests. Further, even if the user's interest is dynamically changing, it is possible to automatically perform information classification reflecting the interest at that time.

【図面の簡単な説明】[Brief description of the drawings]

【図1】本発明が適用されるシステムの構成図である。FIG. 1 is a configuration diagram of a system to which the present invention is applied.

【図2】本処理を行う部分の構成図である。FIG. 2 is a configuration diagram of a part that performs this processing.

【図3】本発明の、利用者の興味に基づいて情報の分類
を行う方法を示す流れ図である。
FIG. 3 is a flowchart illustrating a method of classifying information based on a user's interest according to the present invention.

【図4】図3の方法にしたがって生成される情報の説明
図である。
FIG. 4 is an explanatory diagram of information generated according to the method of FIG. 3;

【符号の説明】[Explanation of symbols]

1 情報サーバ 2 端末 3 ネットワーク 5 分類カテゴリ生成部 6 上位階層作成部 11〜14、21,22,31〜34 ステップ DESCRIPTION OF SYMBOLS 1 Information server 2 Terminal 3 Network 5 Classification category generation part 6 Upper layer creation part 11-14, 21,22,31-34 Step

Claims (15)

【特許請求の範囲】[Claims] 【請求項1】 利用者の興味に基づいて情報を分類する
方法であって、全体量が把握できるある情報群のうちで
利用者が閲覧した情報から抽出された、当該情報の特徴
を表す各キーワードに対応する情報の数を、前記情報群
から抽出された、当該情報の特徴を表す当該キーワード
に対応する情報の数と比較し、後者の情報の数に対する
前者の情報の数の割合が大きいキーワードに1対1に対
応させて、情報の分類に利用する入れ物である分類カテ
ゴリを生成して情報の分類を行う分類カテゴリ生成段階
と、それらの分類カテゴリのうちで関連度合いが高い分
類カテゴリをグループ化し、上位階層を作成する上位階
層作成段階を有する、利用者の興味に基づいて情報を分
類する方法。
1. A method of classifying information based on a user's interest, wherein each of information representing a feature of the information extracted from information viewed by the user from a certain information group whose total amount can be grasped. The number of information corresponding to the keyword is compared with the number of information corresponding to the keyword extracted from the information group and representing the characteristic of the information, and the ratio of the number of the former information to the number of the latter information is large. A classification category generation step of generating a classification category, which is a container used for information classification, for one-to-one correspondence with keywords and classifying information, and a classification category having a high degree of association among those classification categories. A method of classifying information based on a user's interest, comprising an upper layer creation step of grouping and creating an upper layer.
【請求項2】 前記分類カテゴリ生成段階が、全体量が
把握できるある情報群のうちで利用者が閲覧した情報か
ら抽出された、当該情報の特徴を表す各キーワードに対
応する情報の数を、前記情報群から抽出された、当該情
報の特徴を表す当該キーワードに対応する情報の数で除
算し、値が大きいものから上位N個のキーワードについ
て1対1に対応する分類カテゴリを生成する、請求項1
記載の方法。
2. The method according to claim 1, wherein the classification category generation step includes the step of determining the number of pieces of information corresponding to each keyword representing a feature of the information extracted from the information browsed by the user in a certain information group whose total amount can be grasped. Dividing by the number of pieces of information corresponding to the keyword extracted from the information group and representing characteristics of the information, and generating a classification category corresponding to one-to-one with respect to the top N keywords from the largest value. Item 1
The described method.
【請求項3】 前記上位階層生成段階が、生成された分
類カテゴリに対応するキーワードの全ての2つの組み合
せについて、両方のキーワードを含む、利用者が閲覧し
た情報の数を各キーワードを含む、利用者が閲覧した情
報の数で除算し、この値の少なくとも一方があらかじめ
設定されている閾値Xを超えるキーワードに対応する分
類カテゴリをグループ化する、請求項1記載の方法。
3. The method according to claim 1, wherein the upper hierarchy generation step includes, for all two combinations of keywords corresponding to the generated classification categories, the number of pieces of information viewed by the user including both keywords. The method according to claim 1, wherein the category is divided by the number of information viewed by the user, and the classification categories corresponding to keywords whose at least one of them exceeds a preset threshold X are grouped.
【請求項4】 前記分類カテゴリ生成段階は、前記除算
を再度行なった結果、現在分類カテゴリを生成している
キーワードのうち、上位N個のキーワードに入らなくな
ったキーワードに対応する分類カテゴリを消去する、請
求項2記載の方法。
4. The classification category generating step, as a result of performing the division again, deletes a classification category corresponding to a keyword that is not included in the top N keywords among keywords that are currently generating a classification category. 3. The method of claim 2.
【請求項5】 前記上位階層生成段階は、前記除算を再
度行なった結果、現在上位階層の分類カテゴリを生成し
ているキーワードのうち、値が前記閾値X以下になった
キーワードに対応する上位階層の分類カテゴリを消去
し、個々の分類カテゴリとして扱う、請求項3記載の方
法。
5. The upper layer generation step, wherein, as a result of performing the division again, the upper layer corresponding to the keyword whose value is equal to or smaller than the threshold value X among the keywords that are currently generating the classification category of the upper layer. 4. The method according to claim 3, wherein the classification categories are deleted and treated as individual classification categories.
【請求項6】 利用者の興味に基づいて情報を分類する
装置であって、全体量が把握できるある情報群のうちで
利用者が閲覧した情報から抽出された、当該情報の特徴
を表す各キーワードに対応する情報の数を、前記情報群
から抽出された、当該情報の特徴を表す当該キーワード
に対応する情報の数と比較し、後者の情報の数に対する
前者の情報の数の割合が大きいキーワードに1対1に対
応させて、情報の分類に利用する入れ物である分類カテ
ゴリを生成して情報の分類を行う分類カテゴリ生成手段
と、それらの分類カテゴリのうちで関連度合いが高い分
類カテゴリをグループ化し、上位階層を作成する上位階
層作成手段を有する、利用者の興味に基づいて情報を分
類する装置。
6. An apparatus for classifying information based on an interest of a user, wherein each apparatus is extracted from information viewed by the user from a certain group of information whose total amount can be grasped, and represents characteristics of the information. The number of information corresponding to the keyword is compared with the number of information corresponding to the keyword extracted from the information group and representing the characteristic of the information, and the ratio of the number of the former information to the number of the latter information is large. Classification category generating means for generating a classification category, which is a container used for information classification, in one-to-one correspondence with keywords, and classifying information, and a classification category having a high degree of association among the classification categories. An apparatus for classifying information based on a user's interest, comprising an upper layer creating means for grouping and creating an upper layer.
【請求項7】 前記分類カテゴリ生成手段が、全体量が
把握できるある情報群のうちで利用者が閲覧した情報か
ら抽出された、当該情報の特徴を表す各キーワードに対
応する情報の数を、前記情報群から抽出された、当該情
報の特徴を表す当該キーワードに対応する情報の数で除
算し、値が大きいものから上位N個のキーワードについ
て1対1に対応する分類カテゴリを生成する、請求項6
記載の装置。
7. The classification category generation unit determines the number of pieces of information corresponding to each keyword representing a feature of the information extracted from information viewed by a user in a certain information group whose total amount can be grasped. Dividing by the number of pieces of information corresponding to the keyword extracted from the information group and representing characteristics of the information, and generating a classification category corresponding to one-to-one with respect to the top N keywords from the largest value. Item 6
The described device.
【請求項8】 前記上位階層生成手段が、生成された分
類カテゴリに対応するキーワードの全ての2つの組み合
せについて、両方のキーワードを含む、利用者が閲覧し
た情報の数を各キーワードを含む、利用者が閲覧した情
報の数で除算し、この値の少なくとも一方があらかじめ
設定されている閾値Xを超えるキーワードに対応する分
類カテゴリをグループ化する、請求項6記載の装置。
8. The method according to claim 1, wherein the upper hierarchy generation means uses, for each of the two combinations of keywords corresponding to the generated classification category, the number of pieces of information browsed by the user including both keywords. The apparatus according to claim 6, wherein the category is divided by the number of pieces of information viewed by the user, and the classification categories corresponding to the keywords whose at least one of them exceeds a preset threshold X are grouped.
【請求項9】 前記分類カテゴリ生成手段は、前記除算
を再度行なった結果、現在分類カテゴリを生成している
キーワードのうち、上位N個のキーワードに入らなくな
ったキーワードに対応する分類カテゴリを消去する、請
求項7記載の装置。
9. The classification category generating means, as a result of performing the division again, erases a classification category corresponding to a keyword that is not included in the top N keywords among keywords that are currently generating a classification category. An apparatus according to claim 7.
【請求項10】 前記上位階層生成手段は、前記除算を
再度行なった結果、現在上位階層の分類カテゴリを生成
しているキーワードのうち、値が前記閾値X以下になっ
たキーワードに対応する上位階層の分類カテゴリを消去
し、個々の分類カテゴリとして扱う、請求項8記載の装
置。
10. The upper hierarchy generating means, as a result of performing the division again, as a result of the upper hierarchy corresponding to the keyword whose value is equal to or less than the threshold X among the keywords for which the classification category of the upper hierarchy is currently generated. 9. The apparatus according to claim 8, wherein the classification categories are deleted and treated as individual classification categories.
【請求項11】 利用者の興味に基づいて情報を分類す
るプログラムであって、全体量が把握できるある情報群
のうちで利用者が閲覧した情報から抽出された、当該情
報の特徴を表す各キーワードに対応する情報の数を、前
記情報群から抽出された、当該情報の特徴を表す当該キ
ーワードに対応する情報の数と比較し、後者の情報の数
に対する前者の情報の数の割合が大きいキーワードに1
対1に対応させて情報の分類に利用する入れ物である分
類カテゴリを生成して情報の分類を行う分類カテゴリ生
成処理と、それらの分類カテゴリのうちで関連度合いが
高い分類カテゴリをグループ化し、上位階層を作成する
上位階層生成処理をコンピュータに実行させるための、
利用者の興味に基づいて情報を分類するプログラムを格
納した記録媒体。
11. A program for classifying information based on a user's interest, wherein each program is extracted from information viewed by a user from a group of information whose total amount can be grasped, and represents a characteristic of the information. The number of information corresponding to the keyword is compared with the number of information corresponding to the keyword extracted from the information group and representing the characteristic of the information, and the ratio of the number of the former information to the number of the latter information is large. 1 for keywords
Classification category generation processing for generating a classification category, which is a container used for information classification in correspondence with one-to-one, to classify information, and grouping the classification categories having a high degree of association among the classification categories, and For causing a computer to execute a higher hierarchy generation process for creating a hierarchy,
A recording medium storing a program for classifying information based on a user's interest.
【請求項12】 前記分類カテゴリ生成処理が、全体量
が把握できるある情報群のうちで利用者が閲覧した情報
から抽出された、当該情報の特徴を表す各キーワードに
対応する情報の数を、前記情報群から抽出された、当該
情報の特徴を表す当該キーワードに対応する情報の数で
除算し、値が大きいものから上位N個のキーワードにつ
いて1対1に対応する分類カテゴリを生成する、請求項
11記載の記録媒体。
12. The classification category generation processing includes the step of determining the number of pieces of information corresponding to each keyword representing a characteristic of the information extracted from information viewed by a user from a certain information group whose total amount can be grasped. Dividing by the number of pieces of information corresponding to the keyword extracted from the information group and representing characteristics of the information, and generating a classification category corresponding to one-to-one with respect to the top N keywords from the largest value. Item 12. The recording medium according to Item 11.
【請求項13】 前記上位階層生成処理が、生成された
分類カテゴリに対応するキーワードの全ての2つの組み
合せについて、両方のキーワード含む、利用者が閲覧し
た情報の数を各キーワードを含む、利用者が閲覧した情
報の数で除算し、この値の少なくとも一方があらかじめ
設定されている閾値Xを超えるキーワードに対応する分
類カテゴリをグループ化する、請求項11記載の記録媒
体。
13. The method according to claim 13, wherein the upper layer generation processing includes, for each of the two combinations of keywords corresponding to the generated classification category, the number of pieces of information browsed by the user including both keywords. 12. The recording medium according to claim 11, wherein the category is divided by the number of pieces of information browsed, and the classification categories corresponding to the keywords in which at least one of the values exceeds a preset threshold X are grouped.
【請求項14】 前記分類カテゴリ生成段階は、前記除
算を再度行なった結果、現在分類カテゴリを生成してい
るキーワードのうち上位N個のキーワードに入らなくな
ったキーワードに対応する分類カテゴリを消去する、請
求項12記載の記録媒体。
14. The classification category generating step, as a result of performing the division again, deletes a classification category corresponding to a keyword that is not included in the top N keywords among keywords that are currently generating a classification category, The recording medium according to claim 12.
【請求項15】 前記上位階層生成処理は、前記除算を
再度行なった結果、現在上位階層の分類カテゴリを生成
しているキーワードのうち、値が前記閾値X以下になっ
たキーワードに対応する上位階層の分類カテゴリを消去
し、個々の分類カテゴリとして扱う、請求項13記載の
記録媒体。
15. The upper layer generation processing includes, as a result of performing the division again, an upper layer corresponding to a keyword whose value is equal to or smaller than the threshold value X among keywords that are currently generating a classification category of an upper layer. 14. The recording medium according to claim 13, wherein said category is deleted and treated as an individual category.
JP10043620A 1998-02-25 1998-02-25 Method and apparatus for classifying information based on user's interest, and recording medium storing program for classifying information based on user's interest Pending JPH11242679A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP10043620A JPH11242679A (en) 1998-02-25 1998-02-25 Method and apparatus for classifying information based on user's interest, and recording medium storing program for classifying information based on user's interest

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP10043620A JPH11242679A (en) 1998-02-25 1998-02-25 Method and apparatus for classifying information based on user's interest, and recording medium storing program for classifying information based on user's interest

Publications (1)

Publication Number Publication Date
JPH11242679A true JPH11242679A (en) 1999-09-07

Family

ID=12668898

Family Applications (1)

Application Number Title Priority Date Filing Date
JP10043620A Pending JPH11242679A (en) 1998-02-25 1998-02-25 Method and apparatus for classifying information based on user's interest, and recording medium storing program for classifying information based on user's interest

Country Status (1)

Country Link
JP (1) JPH11242679A (en)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2002245212A (en) * 2000-11-22 2002-08-30 Matsushita Electric Ind Co Ltd Group forming system, group forming apparatus, group forming method, program, and medium
JP2008204374A (en) * 2007-02-22 2008-09-04 Fuji Xerox Co Ltd Cluster generating device and program
CN109992724A (en) * 2019-04-03 2019-07-09 西咸新区心灯软件科技有限公司 A kind of calculation method and device of user's compatible degree based on personal characteristic information

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH08153121A (en) * 1994-09-30 1996-06-11 Hitachi Ltd Document information classification method and document information classification device

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH08153121A (en) * 1994-09-30 1996-06-11 Hitachi Ltd Document information classification method and document information classification device

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2002245212A (en) * 2000-11-22 2002-08-30 Matsushita Electric Ind Co Ltd Group forming system, group forming apparatus, group forming method, program, and medium
JP2008204374A (en) * 2007-02-22 2008-09-04 Fuji Xerox Co Ltd Cluster generating device and program
CN109992724A (en) * 2019-04-03 2019-07-09 西咸新区心灯软件科技有限公司 A kind of calculation method and device of user's compatible degree based on personal characteristic information
CN109992724B (en) * 2019-04-03 2024-05-31 西咸新区心灯软件科技有限公司 A method and device for calculating user compatibility based on personal characteristic information

Similar Documents

Publication Publication Date Title
US6654742B1 (en) Method and system for document collection final search result by arithmetical operations between search results sorted by multiple ranking metrics
US6912550B2 (en) File classification management system and method used in operating systems
US7707201B2 (en) Systems and methods for managing and using multiple concept networks for assisted search processing
EP1678635B1 (en) Method and apparatus for automatic file clustering into a data-driven, user-specific taxonomy
JP5621773B2 (en) Classification hierarchy re-creation system, classification hierarchy re-creation method, and classification hierarchy re-creation program
EP1510938B1 (en) A method of providing a visualisation graph on a computer and a computer for providing a visualisation graph
CN102214208B (en) Method and equipment for generating structured information entity based on non-structured text
WO2001031502A1 (en) Multimedia information classifying/arranging device and method
JP5160312B2 (en) Document classification device
Pan et al. Content-based visual summarization for image collections
JP2004213626A (en) Storage and retrieval of information
CN108573408A (en) Popular commodity list making method for maximizing benefits
JP2005107688A (en) Information display method and system, and information display program
Macdonald et al. Blog track research at TREC
JP2004362451A (en) Search keyword information display method and system and search keyword information display program
US20060004809A1 (en) Method and system for calculating document importance using document classifications
JP4359075B2 (en) Concept extraction system, concept extraction method, concept extraction program, and storage medium
JP4407272B2 (en) Document classification method, document classification apparatus, and document classification program
JP4219122B2 (en) Feature word extraction system
JP5787924B2 (en) Cluster forming apparatus, cluster forming method, and cluster forming program
JPH11242679A (en) Method and apparatus for classifying information based on user's interest, and recording medium storing program for classifying information based on user's interest
JP2009211406A (en) Web browse history display device, method and program and computer readable recording medium
JP2004086262A (en) Visual information classification method, visual information classification device, visual information classification program, and recording medium recording the program
JP3692416B2 (en) Information filtering method and apparatus
JP2003323454A (en) Method, apparatus, and computer program for mapping content having meta information

Legal Events

Date Code Title Description
A131 Notification of reasons for refusal

Free format text: JAPANESE INTERMEDIATE CODE: A131

Effective date: 20040225

A02 Decision of refusal

Free format text: JAPANESE INTERMEDIATE CODE: A02

Effective date: 20040908