JPH01214967A - Character processing device and method - Google Patents

Character processing device and method

Info

Publication number
JPH01214967A
JPH01214967A JP63041594A JP4159488A JPH01214967A JP H01214967 A JPH01214967 A JP H01214967A JP 63041594 A JP63041594 A JP 63041594A JP 4159488 A JP4159488 A JP 4159488A JP H01214967 A JPH01214967 A JP H01214967A
Authority
JP
Japan
Prior art keywords
learning data
delimiter
clause
break
pronunciation
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP63041594A
Other languages
Japanese (ja)
Other versions
JP3029109B2 (en
Inventor
Eiichiro Toshima
英一朗 戸島
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Canon Inc
Original Assignee
Canon Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Canon Inc filed Critical Canon Inc
Priority to JP63041594A priority Critical patent/JP3029109B2/en
Publication of JPH01214967A publication Critical patent/JPH01214967A/en
Application granted granted Critical
Publication of JP3029109B2 publication Critical patent/JP3029109B2/en
Anticipated expiration legal-status Critical
Expired - Fee Related legal-status Critical Current

Links

Landscapes

  • Document Processing Apparatus (AREA)

Abstract

(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。
(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.

Description

【発明の詳細な説明】 [産業上の利用分野] 本発明はかな漢字変換を行ないながら漢字等の文字を入
力し、日本語の文書等のドキュメントを作成編集する文
字処理装置において、変換された漢字仮名混り文の文節
区切りを変更した際に、次回の入力からは正しく変換で
きるようにすることができる文字処理装置に関する。
[Detailed Description of the Invention] [Field of Industrial Application] The present invention is a character processing device that inputs characters such as kanji while performing kana-kanji conversion and creates and edits documents such as Japanese documents. The present invention relates to a character processing device that can perform correct conversion from the next input when the clause break of a sentence containing kana is changed.

[従来の技術] 日本語ワードプロセッサなどの日本文を入力する文字処
理装置においては、漢字入力の手段として、キーボード
等より仮名列を入力し、入力された仮名列を仮名漢字変
換することが一般に行なわれている。特に近年はオペレ
ータが文節の切れ目を意識せずに連続的に仮名人力が可
能な「ベタ書き変換」などを提供している機種もある。
[Prior Art] In a character processing device such as a Japanese word processor that inputs Japanese sentences, the method of inputting kanji is generally to input a kana string from a keyboard or the like, and then convert the input kana string into kana-kanji characters. It is. Particularly in recent years, some models offer ``solid writing conversion,'' which allows the operator to manually write kana continuously without being aware of the breaks in phrases.

しかし、このようなベタ書き変換は、システムの解析能
力がまだ不十分であるため、システムの提示する文節の
切れ目がオペレータの望む変換と異なる場合(誤変換、
あるいは誤分割)が発生する。このようなとき、オペレ
ータはいわゆる区切り変更キー(区切り縮小、区切り伸
長)を操作し、望む文節分割を指定するということが一
般に行なわれている。
However, since the analysis ability of the system is still insufficient for this kind of solid writing conversion, if the phrase break presented by the system is different from the conversion desired by the operator (erroneous conversion,
or incorrect division). In such cases, the operator generally operates so-called break change keys (break reduction, break expansion) to specify the desired bunsetsu division.

例えば、入力読み列として「そのもんだいにかんしてい
かのごういがえられた」と人力したとき、システムが誤
って「その/問題に/関し/定価の/合意が/得られた
J(’/Jは文節の切れ目を示す)と変換したとする。
For example, when inputting the input sequence ``Something has been said about that item,'' the system incorrectly reads ``A consensus has been reached on the /problem/at a fixed price.'''/J indicates a break between clauses).

このようなとき、オペレータ、は文節「関し」に対して
「区切り伸長」を指示する。その結果、「その問題に関
して以下の合意が得られた」と正しい変換となる。
In such a case, the operator instructs ``separation expansion'' for the phrase ``seki''. As a result, the correct translation is ``The following agreement was reached regarding the issue.''

このような区切り変更操作に対しては、学習が行なわれ
ないと非常な不便をオペレータに強いることになる0例
えば、上記の例では、次回にもう一度「そのもんだいに
かんしていかのごういかえられた」と入力すると、再び
、「その/問題に/関し/定価の/合意が/得られた」
と変換されると、オペレータはもう一度区切り変更を行
なわなければならない。
For such delimiter changing operations, if the operator is not trained, it will be very inconvenient for the operator.For example, in the above example, the next time you should ask the operator again, "A fixed price/agreement was reached/regarding/the issue."
, the operator must perform another delimiter change.

このため、機種によってはオペレータが行なった区切り
変更を次回に反映させるためのいわゆる区切り学習機能
を具備しているものもある0区切り学習機能は個々の単
語の頻度を上下することにより行なわれる。
For this reason, some models are equipped with a so-called break learning function to reflect the break changes made by the operator in the next time.The zero break learning function is performed by increasing or decreasing the frequency of individual words.

先の例で説明すると「かんしていかの」の変換について
は「関し」 「関して」の頻度!2、「定価」の頻度−
4、r以下」の頻度=3とすると、「関し/定価の」の
頻度−2+4−6、「関して/以下の」の頻度−2+3
−5であるので、第1候補としては「関し/定価の」が
変換される。ここで区切り変更を行なって「「関して/
以下の」に変更したとき、頻度を、「定価」の頻度−3
、r以下」の頻度間4と逆転させれば、次回に再び「か
んしていかの」を入力したとき、「関し/定価の」の頻
度−2+3!5、「関して/以下の」の頻度−2+4W
6となるので、正しい変換「関して/以下の」が得られ
、区切り学習が行なわれたことになる。
To explain using the previous example, the conversion of ``kanshiikano'' is ``seki shi'' and ``seki shite'' frequency! 2. Frequency of "list price" -
4. If the frequency of ``r or less'' = 3, the frequency of ``regarding/list price'' is -2+4-6, and the frequency of ``regarding/less than'' is -2+3.
-5, so "regarding/list price" is converted as the first candidate. At this point, change the delimiter to “Regarding/
When changing the frequency to "List price" frequency - 3
If you reverse the frequency between 4 and 4 for ``, r or less'', the next time you enter ``How about something'' again, the frequency for ``Regarding/List price'' will be -2 + 3!5, and the frequency for ``Regarding/below''. -2+4W
6, the correct conversion "with respect to/below" is obtained, and segmentation learning has been performed.

[発明が解決しようとしている問題点]しかしながら、
この従来方式の区切り学習では、関係のない文脈で思わ
ぬ誤変換が発生し、かえって変換率が悪くなる可能性が
ある。
[Problem that the invention seeks to solve] However,
In this conventional method of segmented learning, unexpected erroneous conversions may occur in unrelated contexts, which may actually worsen the conversion rate.

例えば、上述の例ではr以下」の頻度が向上しているの
で、引き続き「それにかんしたいかのいけんをきく」と
入力し変換すると、「それに/関した/以下の/意見を
7間く」などと変換され、オペレータの望む「それに/
関し/大家の/、!見を7間く」が変換されない可能性
がある。
For example, in the example above, the frequency of ``r or less'' has increased, so if you continue to input ``I would like to ask you what you think about it'' and convert it, you will get ``I would like to give 7 opinions regarding / below / about it''. etc., and the operator desires “and/
Seki/landlord/,! There is a possibility that ``Mise wo 7 kan'' will not be converted.

すなわち、「関する」という単語は「関し」「関して」
と使用されることはよくあるが、「関した」と使用され
ることは通常ないのであるが、従来の頻度に基く区切り
学習ではこの現象に対応することができない。
In other words, the word "regarding" is "regarding" and "regarding".
Although it is often used as ``related'', it is not usually used as ``related'', and conventional segmented learning based on frequency cannot deal with this phenomenon.

上記の問題を解決するために、区切り変更前後の1文節
分の局部的な読みを記憶する方式も考えられる0例えば
、「かルして」 「かんした」などの読みに対してどこ
で文節を分割すれば良いかを記憶し、「かんして」のと
きは「関して/Jと変換し、「かんした」のときは「関
し/た」と変換する方式である。しかし、その場合にも
問題は残る0例えば、「もんだいにかんしていげんをう
けいれる」という入力の場合、r問題に/関し/提言を
/受は入れる」と変換されるのが自然であり、r問題に
/関して/威厳を/受は入れる」は不自然である。これ
は「関して/以下の」の場合と異なる。すなわち、同じ
「かんして」の読みに対して「間し/て」と変換したほ
うが良い場合と「関して/」と変換したほうが良い場合
が存在する。
In order to solve the above problem, it is possible to memorize the local pronunciations of one bunsetsu before and after the delimiter is changed. This method memorizes the appropriate division and converts ``kanshita'' into ``sekishi/J'' and converts ``kanshita'' into ``sekishi/ta.'' However, even in that case, the problem still remains.For example, if the input is ``I will accept your concerns,'' it would naturally be converted to ``I will make a proposal/regarding the r problem.'' , ``I will accept/respect/regarding/regarding/regarding the issue.'' is unnatural. This is different from the case of "with respect to/below". In other words, there are cases where it is better to convert the same pronunciation of ``kanshite'' to ``mashi/te'' and cases where it is better to convert it to ``kishite/''.

以上をまとめると「かんしていかの」→「関して以下の
」、「かんしたいかの」叫「関し大家の」、「かルして
いげんを」→「関し提言を」と変換されるべきなのであ
り、従来方式のような個々の単語の頻度で対応する区切
り学習や、局部的な1文節分の読みを記憶する方式では
、区切り学習の副作用が生じ、使い勝手の悪い仮名漢字
変換となってしまう。
To summarize the above, it should be translated as ``How do you think about it'' → ``The following is the following'', ``How do you think about it'' and ``The landlord's concern'', and ``How do you think about it'' → ``Suggestions regarding the matter''. Therefore, the conventional method of segment learning that corresponds to the frequency of individual words, or the method that memorizes local pronunciations for one sentence segment, has the side effect of segment learning, resulting in inconvenient kana-kanji conversion. Put it away.

[問題点を解決するための手段(及び作用)]本発明は
、区切り変更が起動されたときに、区切り変更を行なう
とともに、区切り変更前後の2文節分の読みと、区切り
変更の結果生じた文節分割位置を記憶することにより、
副作用の生じにくい区切り学習手段を提案し、それによ
り、操作性の高いベタ書き仮名漢字変換を備える文字処
理装置を提供するものである。
[Means for Solving the Problems (and Effects)] The present invention performs a break change when a break change is activated, reads the two sentence segments before and after the break change, and reads the result of the break change. By memorizing the bunsetsu division positions,
The present invention proposes a delimiter learning means that is less likely to cause side effects, thereby providing a character processing device equipped with solid kana-kanji conversion that is highly operable.

[実施例1 以下図面を参照しながら本発明の詳細な説明する。[Example 1 The present invention will be described in detail below with reference to the drawings.

第1図は本発明の全体構成の一例である。FIG. 1 is an example of the overall configuration of the present invention.

図示の構成において、cpuは、マイクロプロセッサで
あり、文字処理のための演算、論理判断等を行ない、ア
ドレスバスAB、コントロールバスCB、データバスD
Bを介して、それらのバスに接続された各構成要素を制
御する。
In the illustrated configuration, the CPU is a microprocessor that performs calculations, logical judgments, etc. for character processing, and includes an address bus AB, a control bus CB, and a data bus D.
B to control each component connected to those buses.

アドレスバスABはマイクロプロセッサcPUの制御の
対象とする構成要素を指示するアドレス信号を転送する
。コントロールバスCBはマイクロプロセッサCPUの
制御の対象とする各構成要素のコントロール信号を転送
して印加する。データバスDBは各構成機器相互間のデ
ータの転送を行なう。
The address bus AB transfers address signals indicating the components to be controlled by the microprocessor cPU. The control bus CB transfers and applies control signals for each component to be controlled by the microprocessor CPU. The data bus DB transfers data between each component device.

つぎにROMは、読出し専用の・固定メモリであり、第
9図〜第14図につき後述するマイクロプロセッサCP
Uによる制御の手順、及び、単語辞書、文法辞書等の固
定データを記憶させておく。
Next, the ROM is a read-only fixed memory, and the microprocessor CP, which will be described later with reference to FIGS. 9 to 14,
The control procedure by U and fixed data such as a word dictionary and a grammar dictionary are stored.

単語辞書は読み、表記、文法情報等が対応して記憶され
たものであり、仮名漢字変換等で参照される0文法辞書
は形態素解析、構文解析等で必要となる単語間の接続規
則等が記憶されたものである。
A word dictionary stores information such as pronunciation, notation, and grammar in correspondence, and a zero-grammar dictionary that is referenced in kana-kanji conversion, etc. contains connection rules between words, etc. that are required for morphological analysis, syntactic analysis, etc. It is memorized.

また、RAMは、1ワード16ビツトの構成の書込み可
能のランダムアクセスメモリであって、各構成要素から
の各種データの一時記憶に用いる。にULDTは第5図
に詳述される区切り学習データである。HENTBLは
第7図に詳述される変換候補テーブルである。KBBU
Fは入力された読み列を蓄えるためのキーボードバッフ
ァであり、第8図に示すように構成される。
Further, the RAM is a writable random access memory having a configuration of 1 word and 16 bits, and is used for temporary storage of various data from each component. The ULDT is the delimited learning data detailed in FIG. HENTBL is a conversion candidate table detailed in FIG. KBBU
F is a keyboard buffer for storing input reading sequences, and is configured as shown in FIG.

KBはキーボードであって、アルファベットキー、ひら
かなキー、カタカナキー等の文字記号入カキ−1及び、
カーソル移動キー、仮名漢字変換キー、区切り縮小キー
、区切り伸長キー等の本文字処理装置に対する各種機能
を指示するための各種のファンクションキーを備えてい
る。
KB is a keyboard, and has character symbol keys such as alphabet keys, hirakana keys, katakana keys, etc.
It is equipped with various function keys for instructing various functions to the character processing device, such as a cursor movement key, a kana-kanji conversion key, a delimiter reduction key, and a delimiter expansion key.

D!Sには文書データを記憶するための外部記憶であり
、作成された文書の保管を行ない、保管された文書はキ
ーボードの指示により、必要な時呼び出される。
D! S is an external storage for storing document data, and stores created documents, and the stored documents can be called up when necessary by instructions from the keyboard.

CRはカーソルレジスタである。cpuにより、カーソ
ルレジスタの内容を読み書きできる。
CR is a cursor register. The CPU can read and write the contents of the cursor register.

後述するCRTコントローラCRTCは、ここに蓄えら
れたアドレスに対応する表示装fJCRT上の位置にカ
ーソルを表示する。
A CRT controller CRTC, which will be described later, displays a cursor at a position on the display device fJCRT corresponding to the address stored here.

DB[JFは表示用バッファメモリで、表示すべきデー
タのパターンを蓄える1文書データの内容の表示を行な
うときは、DBUF上にパターンを展開することにより
行なわれる。
DB[JF is a display buffer memory, which stores a pattern of data to be displayed.When displaying the contents of one document data, the pattern is developed on DBUF.

CRTCはカーソルレジスタOR及びバッファDBυF
に蓄えられた内容を表示器CRTに表示する役割を担う
CRTC is cursor register OR and buffer DBυF
It plays the role of displaying the contents stored in the CRT on the display device CRT.

またCRTは陰極線管等を用いた表示装置であり、その
表示装置CRTにおけるドツト構成の表示パターンおよ
びカーソルの表示をCRTコントローラで制御する。 
さらに、CGはキャラクタジェネレータであって、表示
装置CRTに表示する文字、記号のパターンを記憶する
ものである。
Further, a CRT is a display device using a cathode ray tube or the like, and a CRT controller controls the dot-configured display pattern and cursor display on the display device CRT.
Furthermore, CG is a character generator that stores patterns of characters and symbols to be displayed on the display device CRT.

かかる各構成要素からなる本発明文字処理装置において
は、キーボードにBからの各種の入力に応じて作動する
ものであって、キーボードK B hlらの入力が供給
されると、まず、インタラブド信号がマイクロプロセッ
サCPUに送られ、そのマイクロプロセッサCPUがR
OM内に記憶しである各種の制御信号を読出し、それら
の制御信号に従って各種の制御が行なわれる。
In the character processing device of the present invention, which is composed of each of these components, the keyboard operates in response to various inputs from B, and when inputs from the keyboard K B hl, etc. are supplied, the interwoven signal is first output. is sent to the microprocessor CPU, and the microprocessor CPU
Various control signals stored in the OM are read out, and various controls are performed in accordance with these control signals.

第2図は本発明装置の画面構成を示した図である0図中
、CRTは表示画面を意味する。CMはカーソルであり
、次にキー°入力を行なったとき、文字が入フていく位
置を示すものである。TSはテキスト画面であり、記憶
されている文書の内容が表示される。MSはモニタライ
ンであり、人力されるキーデータが逐一表示されるエリ
アである。仮名浅学変換等を行なうときは一旦MSに読
み列が表示され、変換後、変換結果がTS上に転送され
る。なお、仮名浅学変換キーは「/」で示される。
FIG. 2 is a diagram showing the screen configuration of the apparatus of the present invention. In FIG. 2, CRT means a display screen. CM is a cursor that indicates the position where a character will be entered when the next key input is performed. TS is a text screen on which the contents of stored documents are displayed. MS is a monitor line, which is an area where key data entered manually is displayed one by one. When performing kana/asagaku conversion, etc., the reading sequence is displayed on the MS, and after conversion, the conversion result is transferred to the TS. Note that the kana/asagaku conversion key is indicated by "/".

第3図は本発明における区切り伸長の操作方法を示した
図である。まず操作者は、「その」を文書中に入力し、
更にr問題に関して以下の合意が得られた」を入力しよ
うとして、読み「もんだいにかんしていかのごういがえ
られたJをキー人力し、次いで、仮名漢字変換キー「/
」を入力する。(a) 「もんだいにかルしていかのごういがえられた/」が仮
名浅学変換されr問題に関し定価の合意が得られた」と
なりテキスト画面に入フてぃ<、(b) この時操作者は「関し定価の」の部分が自分の望む文節
分割ではなかったことに気付き、カーソルを「関」の位
置に移動する。(C) 次いで区切り伸長キーを入力して文節「関し」の伸長を
指示する。「間して以下の」と再変換され、表示される
。(d) 第4図は本発明における区切り縮小の操作方法を示した
図である。まず操作者は、「その」を文書中に入力し、
更に「提案に対し対価を払った」を入力しようとして、
読み「ていあんにたいしたいかをはらった」をキー人力
し、次いで、仮名漢字変換キー「/」を入力する。(a
) 「ていあんにたいしたいかをはらった/」が仮名浅学変
換され「提案に対した医科を払りた」となりテキスト画
面に入ってい<、(b)この時操作者は「対した医科を
」の部分が自分の望む文節分割ではなかったことに気付
き、カーソルをr対」の位置に移動する。(C)次いで
区切り縮小キーを入力して文節「対した」の縮小を指示
する。「対し対価を」と再変換され、表示される。(d
) 第5図は区切り学習データKULDTの構成を示した図
である。
FIG. 3 is a diagram showing an operation method for delimiting and expanding according to the present invention. First, the operator inputs "so" into the document,
In addition, when trying to enter "The following agreement has been reached regarding problem
”. (a) ``In some way, I was able to find something/'' was converted into kana and sagaku and became ``We reached an agreement on the list price for the r problem'', which appears on the text screen. (b) At this time, the operator realizes that the part ``Seki Shigai no'' is not the phrase division he desires, and moves the cursor to the position of ``Seki.'' (C) Next, input the delimiter expansion key to instruct expansion of the phrase "Seki". It will be reconverted and displayed as "Below or less". (d) FIG. 4 is a diagram showing an operation method for dividing and reducing in the present invention. First, the operator inputs "so" into the document,
Furthermore, when trying to enter "I paid for the proposal,"
Enter the pronunciation ``Teian ni taitai wo wo hattata'' and then enter the kana-kanji conversion key ``/''. (a
) ``I paid the doctor for the proposal/'' is converted into kana and becomes ``I paid the doctor for the proposal'' and enters the text screen <, (b) At this time, the operator says ``The doctor for the proposal'' Realizing that part 2 was not the phrase division he wanted, he moved the cursor to ``r pair''. (C) Next, input the delimiter reduction key to instruct reduction of the clause ``tita''. It will be reconverted and displayed as "for consideration." (d
) FIG. 5 is a diagram showing the structure of the delimited learning data KULDT.

読みは区切り変更に関係する2文節分の読みを記憶する
エリアである。第1文節読み数は、前記読みのうち区切
り変更後の第1文節になる部分の読み数を格納する。第
2文節読み数は、前記読み−のりち区切り変更後の第2
文節になる部分の読み数を格納する。
The reading is an area that stores the readings of two clauses related to the break change. The number of readings of the first clause stores the number of readings of the part of the reading that becomes the first clause after the delimiter is changed. The number of readings for the second clause is the number of readings for the second clause after changing the reading-norichi division.
Stores the reading count of the part that becomes a clause.

例えば、入力「かんしていかのごういが」に対して「関
し/定価の/合意が」と変換され、区切り変更で「関し
て/以下の/合意が」に修正した場合を考える。この時
、2文節分の読みは「かんしていかの」であり、区切り
変更後の第1文節の読みは「かんして」第2文節の読み
は「いかの」である、従って、区切り学習データへの登
録は、読み−「かんしていかの」、第1文節読み数=4
、第2文節読み数−3となる。
For example, consider a case where the input ``About the situation'' is converted to ``Regarding/the following/agreement'', and is modified to ``Regarding/the following/agreement'' by changing the delimiter. At this time, the readings for the two clauses are "Kansei Ikano", and the readings of the first clause after the break change are "Kansete" and the readings of the second clause are "Ika no". Therefore, the break learning data To register, reading - "Kanseiikano", number of readings of the first clause = 4
, the number of readings of the second clause is -3.

また別の例として、入力「このへやはたいへんあつい」
に対して「古野へ/八幡/異変/暑い」と変換され、区
切り変更で「この/部屋は/大変/暑い」に修正した場
合を考える。この時、区切り変更前の2文節の読みは「
このへやはた」であリ、変更後の2文節の読みは「この
へやは」であるので、長いほうの「このへやけた」が区
切り学習データの読みとなる0区切り変更後の第1文節
の読みは「この」、第2文節の読みは「へやは」である
、従って、区切り学習データへの登録は、読み=「この
へやけた」、第1文節読み数=2、第2文節読み数諺3
となる。
As another example, input ``It's very hot in this room.''
Consider a case where the text is converted to "To Furuno/Yawata/It's strange/hot" and is modified to "This/room is/very/hot" by changing the delimiter. At this time, the reading of the two clauses before the break change is “
The reading of the two clauses after the change is "Konoheyaha", so the longer one "Konoheyaketa" is the reading of the delimiter learning data.After changing the 0 delimiter The reading of the first clause is "Kono" and the reading of the second clause is "Heyaha". Therefore, the reading of the first clause is "Kono Heyaketa" and the number of readings of the first clause is 2. , 2nd clause reading proverb 3
becomes.

第6図は仮名漢字変換を行なう際に、処理途中で解析さ
れる文節候補の例を示した図である。
FIG. 6 is a diagram showing an example of clause candidates that are analyzed during the process when performing kana-kanji conversion.

入力「かんしていかのごうい」に対して可能な文節構造
を図示している。左端は第1文節の候補であり、「官」
 「関し」 「関して」の可能性があることを意味する
。実際には例えば「官」に対して「缶」 「漢」 「環
」などの同音語があるが、煩雑になるので図では省略し
ている。
Possible phrase structures for the input ``Kansei Ika no Goi'' are illustrated. The leftmost candidate is the first clause, “kan”
``Seki'' means that there is a possibility of ``regarding''. In reality, for example, there are homophones for ``kan'' such as ``can'', ``kan'', and ``kan'', but these are omitted in the diagram to avoid complication.

第2文節の候補としては、第1文節「官」に対してr死
」 「仕手」 「指定」があり、第1文節「関し」に対
して1手」 「帝」 「定価」 「定価の」があり、第
1文節「関して」に対して「胃」「医科」 「医科の」
の可能性があることを意味する。
Candidates for the second clause include ``r death'', ``shite'', and ``designation'' for the first clause ``kan'', and ``tei'', ``list price'', and ``list price'' for the first clause ``seki''. ”, and the first clause “regarding” is replaced with “stomach”, “medical”, and “medical”.
This means that there is a possibility that

以下同様に文節の構造が表現される。The structure of the clause is expressed in the same manner below.

第7図は前述の文節候補を内部的に記述するための変換
候補テーブルHENTBLの構成を示した図である。
FIG. 7 is a diagram showing the structure of a conversion candidate table HENTBL for internally describing the above-mentioned bunsetsu candidates.

「読み」の欄は各文節の読みを記述する。「辞書」の欄
は、その文節の自立部の単語が単語辞書上のどの部分に
存在するかアドレスを記述する。
The "Yomi" column describes the reading of each clause. The "Dictionary" column describes the address of the part of the word dictionary in which the word in the independent part of the phrase exists.

図中では鍵括弧で括ってアドレスを意味して°いる。「
送り仮名数」は「読み」で記述した読みのうち自立部の
読みを除いた送り仮名部分の読み数を記述する0例えば
、「関して」であれば送り仮名「して」の読み数を記述
する。「次候補」の欄はその文節と交代しうる次候補の
文節を指し示すポインタである。第6図の文節候補の図
で言うと、縦の位置に並ぶ文節をリンクするものである
0例えば、「官」は「関し」 「関して」にリンクして
いる0次候補が存在しないときは「−1」を格納してそ
れ以上リンクが続かないことが分かるようになっている
。1次文節」の欄は、その文節に引き続く文節へのリン
クである。第6図の文節候補の図は右の位置になら部文
節をリンクするものである0例えば、r死」は「仕手」
 「指定」にリンクしている0次文節が存在しないとき
はr−1」を格納してそれ以上リンクが続かないことが
分かるようになっている。
In the figure, addresses are enclosed in square brackets. "
``Number of okurigana'' is the number of readings of the okurikana part of the reading written in ``yomi'' excluding the reading of the independent part 0 For example, for ``Kishite'', the number of readings of the okurikana ``shite'' is written. Describe. The "Next Candidate" column is a pointer that indicates the next candidate phrase that can replace the phrase. In the diagram of clause candidates in Figure 6, it links clauses arranged vertically. For example, "kan" is "seki" and there is no zero-order candidate linked to "seki". stores "-1" to indicate that the link will not continue any further. The ``primary clause'' column is a link to the clause following that clause. The phrase candidate diagram in Figure 6 links the part phrase in the right position.0For example, ``r death'' is ``shite''.
When there is no zero-order clause linked to "designation,""r-1" is stored to indicate that the link will not continue any further.

第8図はキーボードバッファにBBUFの構成を示した
図である。キーボードから入力されたキーデータは一旦
このKBBUFに蓄えられる。
FIG. 8 is a diagram showing the structure of the BBUF in the keyboard buffer. Key data input from the keyboard is temporarily stored in this KBBUF.

例えば、仮名漢字変換のときは変換される条件が整うま
で、読み列がこのKBBUFに蓄積され、変換条件が整
った段階で漢字に変換され、バッファがクリアされる。
For example, in the case of kana-kanji conversion, the reading sequence is stored in this KBBUF until the conversion conditions are met, and when the conversion conditions are met, it is converted to kanji and the buffer is cleared.

文字は例えばJIS X 0208コードコードを使用
して1文字2バイトで格納される0図中「/」は仮名漢
字変換キーを意味し、「/」までに格納されている読み
列を漢字に変換するという意味である。
Characters are stored as 2 bytes per character using, for example, the JIS It means to do.

上述の実施例の動作をフローに従って説明する。The operation of the above embodiment will be explained according to the flow.

第9図はキー人力を取り込み、処理を行なう部分のフロ
ーチャートである。
FIG. 9 is a flowchart of the part that takes in key human power and performs processing.

ステップ9−1はキーボードからのデータを大力バッフ
ァにBBUFに取り込む処理である。
Step 9-1 is a process of taking data from the keyboard into the BBUF in a large buffer.

ステップ9−2において人力バッファにBBUFのキー
内容をチエツクし、キーの種類に応じて各処理に分岐す
る。KBBUF内に仮名漢字変換キーのデータが含まれ
ていたときは仮名漢字変換を行なわなければならずステ
ップ9−3に分岐する。にBBUF先頭が区切り縮小キ
ーまたは区切り伸長キーであれば区切り変更を行なわな
ければならずステップ9−4に分岐する。上記以外のキ
ー内容であれば、通常のWA集処理を行なうので、ステ
ップ9−5に分岐し、カーソル移動、挿入、削除等の一
般のワードプロセッサにおいて見られるその他の処理を
行なう。
In step 9-2, the key content of BBUF is checked in the manual buffer, and the process branches to each process depending on the type of key. If KBBUF contains data of a kana-kanji conversion key, kana-kanji conversion must be performed and the process branches to step 9-3. If the beginning of the BBUF is a delimiter reduction key or a delimiter expansion key, the delimiter must be changed and the process branches to step 9-4. If the key content is other than the above, normal WA collection processing is performed, and the process branches to step 9-5 to perform other processing found in general word processors such as cursor movement, insertion, deletion, etc.

ステップ9−3において第10図に詳述するようにKB
BUF上の入力読み列を仮名漢字変換し、文書に出力し
、表示する。
In step 9-3, the KB is
The input reading sequence on the BUF is converted into kana-kanji, output to a document, and displayed.

ステップ9−4において第14図に詳述するように区切
り変更を行ない、文書に出力し、表示する。
In step 9-4, the delimiter is changed as detailed in FIG. 14, and the document is output and displayed.

第10図はステップ9−2「仮名漢字変換」を詳細化し
たフローチャートである。 ステップ10−1において
KBBUF上にある入力読み列を単語辞書、文法辞書等
を参照して形態素解析、構文解析等を行ない、変換候補
テーブルWENTBLを作成する。
FIG. 10 is a detailed flowchart of step 9-2 "kana-kanji conversion". In step 10-1, the input pronunciation sequence on KBBUF is subjected to morphological analysis, syntactic analysis, etc. with reference to a word dictionary, grammar dictionary, etc., and a conversion candidate table WENTBL is created.

ステップ10−2において、第11図に詳述するように
、上記変換候補のうちどの変換候補を変換すべきか、第
1候補を決定する。
In step 10-2, as detailed in FIG. 11, a first candidate is determined which of the conversion candidates should be converted.

ステップ1O−3において、決定された第1候補(採用
文節)を1文節ずつ漢字仮名混り文に変換していく。
In step 1O-3, the determined first candidates (adopted clauses) are converted clause by clause into sentences containing kanji and kana.

ステップ10−4において作成された変換文字列を文書
にセットする。更に文書の変更内容が分かるように変換
文字列がセットされた付近を表示する。
The converted character string created in step 10-4 is set in the document. Furthermore, the area near where the converted character string is set is displayed so that the content of changes to the document can be seen.

第11図はステップ1O−2r第1候補決定」を詳細化
したフローチャートである。
FIG. 11 is a detailed flowchart of Step 1O-2r, ``Determination of first candidate''.

ステップ11−1は変数の初期化処理である。Step 11-1 is variable initialization processing.

変換候補テーブルHETBL中の文節を指し示すポイン
タである現文節ポインタを1に初期化し、先頭の文節の
mt候補を、指し示すようにする0次に、採用された文
節が何番目の文節であるかを示すカウンタiを1に初期
化する。最後に入力読みバッファKBBUF上の読みの
うち、どの部分を現在処理中であるかを示すポインタ入
力読みポインタを1に初期化し、読みの先頭を示すよう
にする。
Initialize the current clause pointer, which is a pointer that points to a clause in the conversion candidate table HETBL, to 1, and point it to the mt candidate of the first clause.0 Next, determine what clause number the adopted clause is. The counter i shown is initialized to 1. Finally, the input reading pointer, which is a pointer indicating which part of the reading on the input reading buffer KBBUF is currently being processed, is initialized to 1 to indicate the beginning of the reading.

ステップ11−2においてそれまでに全文節の処理が終
了したかを判定する。具体的には現文節ポインタが−1
であるかどうかで判定する。全文節の処理が終了してい
るとき(現文節ポインタクー1のとき)はリターンする
In step 11-2, it is determined whether all clauses have been processed by then. Specifically, the current clause pointer is -1
Determine whether or not. When the processing of all the clauses has been completed (when the current clause pointer is 1), the process returns.

ステジブ11−3において第12図に詳述するように区
切り学習サーチを行ない、現在処理している読みが区切
り学習データに登録されているかどうかをサーチする。
In stage 11-3, a break learning search is performed as detailed in FIG. 12, and it is searched whether the pronunciation currently being processed is registered in the break learning data.

もしあれば、採用区切り学習として出力される。If there is, it will be output as adoption break learning.

ステップ11−4において採用区切り学習が出力された
かどうかを判定し、もしあればステップ1オー5に進む
が、無ければステップ11−8に分岐する。
In step 11-4, it is determined whether or not the adoption delimiter learning has been output, and if so, the process proceeds to step 1-5, but if not, the process branches to step 11-8.

ステップ11−5において第13図に詳述するように変
換候補サーチを行ない、ステップ11−3で出力された
採用区切り学習に整合する変換候補が存在するかどうか
サーチする。もしあれば、第1採用変換候補、第2変換
候補が出力される。
In step 11-5, a conversion candidate search is performed as detailed in FIG. 13, and a search is made to see if there is a conversion candidate matching the adopted delimiter learning output in step 11-3. If there are any, the first adopted conversion candidate and the second conversion candidate are output.

またこの時採用文節数雪2が設定される。Also, at this time, the adopted bunsetsu number 2 is set.

ステップ11−6において整合する変換候補が見つかっ
たかどうかを判定し、見つかったときはステップ1−7
に分岐し、見つからなかったときはステップ11−85
分岐する。
It is determined whether a matching conversion candidate is found in step 11-6, and if found, step 1-7
Branch to step 11-85 if not found
Branch out.

ステップ11−7において、ステップ11−5で出力さ
れた第1採用変換候補、第2変換候補を第i文節、第i
+1文節として採用し、ステップ11−10に分岐する
In step 11-7, the first adopted conversion candidate and the second conversion candidate output in step 11-5 are used as the i-th clause and the i-th
It is adopted as a +1 clause and branches to step 11-10.

ステップ11−8において、整合する区切り学習がなか
ったわけであるから、通常の2文節最長一致法にしたが
って採用文節を決定する。
In step 11-8, since there is no matching segment learning, the adopted clause is determined according to the usual two-clause longest matching method.

ステップ11−9において採用文節数を1に設定する。In step 11-9, the number of adopted clauses is set to 1.

ステップ11−10に、おいて、iに採用文節数の値を
加える。
In step 11-10, the value of the number of adopted clauses is added to i.

ステップ11−11において、現文節ポインタを更新し
、採用文節の次の文節を指すようにする。具体的には変
換候補テーブルHENTBLの1次文節」の欄の値を代
入することになる。ステップ11−5で変換候補が採用
されたときは第2採用変換候補の1次文節」を代入する
ようにする。また、入力読みポインタを採用された文節
の読み数分だけ加算して更新する。ステップ11−5で
変換候補が採用されたときは第1採用変換候補、第2採
用変換候補の読み数の和を加算するようにする。
In step 11-11, the current clause pointer is updated to point to the next clause of the adopted clause. Specifically, the value in the column "Primary Clause" of the conversion candidate table HENTBL is substituted. When the conversion candidate is adopted in step 11-5, the "primary phrase of the second adopted conversion candidate" is substituted. Also, the input reading pointer is updated by adding the number of readings of the adopted clause. When a conversion candidate is adopted in step 11-5, the sum of the reading numbers of the first adopted conversion candidate and the second adopted conversion candidate is added.

第12図はステップ1l−3r区切り学習サーチ」を詳
細化したフローチャートである。
FIG. 12 is a detailed flowchart of the step 11-3r section learning search.

ステップ12−1において、変数を初期化する。すなわ
ち、現在処理中の区切り学習データを指し示すぼんたで
あるr医学ポインタ」の値を1に初期化する。また、一
致した区切り学習データのうち最大の読み数のものを意
味する変数「最大読み数」の値をOに初期化する。
In step 12-1, variables are initialized. That is, the value of the "r medical pointer" which is a blank pointing to the delimited learning data currently being processed is initialized to 1. Further, the value of the variable "maximum number of readings", which means the one with the largest number of readings among the matching delimiter learning data, is initialized to O.

ステップ11−2において処理される区切り学習データ
がなくなったかどうかを判定する。具体的には医学ポイ
ンタの値が−1になったかどうかで判定する。もし、区
切り学習データが終りのとき(医学ポインタ=−1のと
き)はステップ12−8に分岐する。
In step 11-2, it is determined whether there is no more delimited learning data to be processed. Specifically, the determination is made based on whether the value of the medical pointer has become -1. If the delimited learning data ends (medical pointer=-1), the process branches to step 12-8.

ステップ12−3において、現在処理中の区切り学習デ
ータ(医学ポインタが指し示す)の読みと、現在の入力
読み(KBBUF上の読みで、入力読みポインタの示す
位置以降)とを比較する。
In step 12-3, the reading of the delimited learning data currently being processed (pointed to by the medical pointer) is compared with the current input reading (reading on the KBBUF, from the position indicated by the input reading pointer).

ステップ12−4において、ステップ12−3の比較が
一致するかどうか判定する。もし一致−しなければステ
ップ12−7に分岐する。
In step 12-4, it is determined whether the comparison in step 12-3 matches. If they do not match, the process branches to step 12-7.

ステップ12−5において一致した区切り学習データの
読み数が最大読み数を超えているかどうかをチエツクす
る。もし、超えていなければステップ12−7に分岐す
乞。
In step 12-5, it is checked whether the number of readings of the matching delimiter learning data exceeds the maximum number of readings. If not, proceed to step 12-7.

ステップ12−6において、見つかった区切り学習デー
タを当面の採用区切り学習データとする。
In step 12-6, the found delimiter learning data is set as the currently employed delimiter learning data.

ステップ12−7において、次の区切り学習データの処
理を行なうために、医学ポインタを次の区切り学習を示
すように+1して更新する。
In step 12-7, in order to process the next segment learning data, the medical pointer is updated by +1 to indicate the next segment learning.

第13図はステップ1l−5r変換候補サーチ」を詳細
化したフローチャートである。
FIG. 13 is a detailed flowchart of steps 11-5r conversion candidate search.

ステップ13−1において、変換候補テーブル)IEN
TBL上の文節のうち現在処理中の文節を示す変数であ
る量を現文節ポインタの値に初期設、定する。
In step 13-1, the conversion candidate table)IEN
A quantity, which is a variable indicating the clause currently being processed among the clauses on the TBL, is initialized and set to the value of the current clause pointer.

ステップ13−2において、現在の文節の次候補が既に
終りであるかどうかを判定する。具体的にはiの値が−
1であるかどうかで判定する。もし、次候補が終り(i
−−1)のときはステップ3−14に分岐する。
In step 13-2, it is determined whether the next candidate for the current phrase has already ended. Specifically, the value of i is -
Determine whether it is 1 or not. If the next candidate ends (i
--1), the process branches to step 3-14.

ステップ13−3において、iの指し示す変換候補の読
み数と、採用区切り学習の第1文節読み数とが一致する
かどうか比較する。
In step 13-3, it is compared whether the number of readings of the conversion candidate pointed to by i and the number of readings of the first clause in the adopted break learning match.

ステップ13−4において、読み数の一致がどうであっ
たか判定し、もし、一致すれば、ステップ13−6に進
んで、第2文節目のチエツクに入る。一致しなければ、
ステップ13−5に分岐し、iの値をiの次候補を示す
ように更新(HENTBLの「次候補」の欄を代入)し
、更にステップ13−2に戻る。
In step 13-4, it is determined whether the reading numbers match, and if they match, the process proceeds to step 13-6 to check the second bunsetsu. If it doesn't match,
The process branches to step 13-5, and the value of i is updated to indicate the next candidate for i (by substituting the "next candidate" column of HENTBL), and the process returns to step 13-2.

ステップ13−6においてiの示す変換候補を第1採用
変換候補とする。
In step 13-6, the conversion candidate indicated by i is set as the first adopted conversion candidate.

ステップ13−7において、jの値をlの変換候補の次
文節を示すように初期設定する。  (HENTBLの
1次文節」の欄を代入する。)ステップ13−8におい
て、jの次候補が既に終りであるかどうかを判定する。
In step 13-7, the value of j is initialized to indicate the next clause of the conversion candidate of l. (Substitute the column "Primary Clause of HENTBL.") In step 13-8, it is determined whether the next candidate for j has already ended.

具体的にはjの値が−lであるかどうかで判定する。も
し、次候補が終り(j=−1)のときはステップ12−
14に分岐する。
Specifically, the determination is made based on whether the value of j is -l. If the next candidate is the end (j=-1), step 12-
Branches into 14.

ステップ13−9においてjの指し示す変換候補の読み
数と、採用区切り学習の第2文節読み数とが一致するか
どうか比較する。
In step 13-9, the number of readings of the conversion candidate pointed to by j is compared with the number of readings of the second clause in the adopted break study.

ステップ13−10において読み数の一致がどうであっ
たか判定し、もし、一致すれば、ステップ13−12に
進む、一致しなければ、ステップ13−1tに分岐し、
jの値をjの次候補を示すように更新(I(ENTBL
の「次候補」の欄を代入)し、更にステップ13−8に
戻る。
In step 13-10, it is determined whether the reading numbers match, and if they match, proceed to step 13-12; if not, proceed to step 13-1t,
Update the value of j to indicate the next candidate for j (I(ENTBL
), and the process returns to step 13-8.

ステップ13−12においてjの示す変換候補を第2採
用変換候補とする。
In step 13-12, the conversion candidate indicated by j is set as the second adopted conversion candidate.

ステップ13−13において採用文節数を2に代入して
リターンする。
In step 13-13, the number of adopted clauses is substituted to 2 and the process returns.

ステップ13−’14では採用文節数を0に設定してリ
ターンする。
In step 13-'14, the number of adopted phrases is set to 0 and the process returns.

第14図はステップ9−4「区切り変更」を詳細化した
フローチャートである。
FIG. 14 is a detailed flowchart of step 9-4 "change delimiter".

ステップ14−1において、入カキ−が区切り縮小キー
であるか区切り伸長キーであるかに応じて、区切り縮小
または区切り伸長の処理を実行する。この処理は現実に
日本語ワードプロセッサ等において実現されており、公
知の技術であるの特に記述しない。
In step 14-1, a process of delimiting or decompressing is executed depending on whether the input key is a delimiting key or an expanding key. This processing is actually realized in a Japanese word processor, etc., and is a well-known technique, so it will not be specifically described.

ステップ14−2において、区切り変更前の2文節分の
読みの読み数を取り出しり、に代入する。
In step 14-2, the number of readings for the two clauses before the delimiter is changed is taken out and substituted into .

ステップ14−3において、区切り変更後の2文節分の
読みの読み数を取り出しL2に代入する。
In step 14-3, the number of readings for the two phrases after the delimiter change is extracted and substituted into L2.

ステップ14−4において、先に求めたLl とL2の
値を比較し、もし、L t > L 2であれば、すな
わち、区切り変更前の2文節分の読み数が長ければ、ス
テップ14−5に分岐する。LI≦L2であれば、すな
わち、区切り変更後の2文節分の読み数が長ければ、ス
テップ14−6に分岐する。
In step 14-4, the values of Ll and L2 obtained earlier are compared, and if L t > L 2, that is, if the number of readings for two clauses before changing the break is long, step 14-5 Branch into. If LI≦L2, that is, if the number of readings for two clauses after the delimiter change is long, the process branches to step 14-6.

ステップ14−5において区切り変更前の2文節分の読
みを区切り学習データ登録のための「読み」とする。
In step 14-5, the readings for the two sentence segments before the break change are set as the "readings" for registering the break learning data.

ステップ14−6において区切り変更後の2文節分の読
みを区切り学習データ登録のための「読み」とする。
In step 14-6, the readings for the two phrases after changing the delimiter are set as "readings" for registering the delimiter learning data.

ステップ14−7において、区切り変更後の第1文節の
読み数を区切り学習データ登録のための「第1文節読み
数」とす、る。
In step 14-7, the number of readings of the first clause after changing the break is set as the "number of first clause readings" for registering the break learning data.

ステップ14−8において、区切り変更後の第2文節の
読み数を区切り学習データ登録のための「第2文節読み
数」とする。
In step 14-8, the number of readings of the second clause after the break is changed is set as the "number of second clause readings" for registering the break learning data.

ステップ重4−9において、上記設定された通りに区切
り学習データを登録する。
In step 4-9, the delimited learning data is registered as set above.

ステップ14−10において、上記登録した区切り学習
データと矛盾する区切り学習データをサーチし、もし矛
盾する区切り学習データが見つかればそれを削除する。
In step 14-10, a search is made for delimiter learning data that is inconsistent with the registered delimiter learning data, and if contradictory delimiter learning data is found, it is deleted.

矛盾する区切り学習データとは、以下の学習データのこ
とである。
The contradictory delimited learning data refers to the following learning data.

■登録区切り学習データr読み」と一致する「読み」を
もつ、または、登録区切り学習データ「読み」よりも長
く入力読みと一致する区切り学習 かつ ■登録区切り学習データと文節分割が一致しないとぎ、
すなわち、「第1文節読み数」または「第2文節読み数
」が一致しない区切り学習 この矛盾区切り学習データ削除処理により、学習効果の
得られなくなる場合が発生するのを防いでいる。
■Have a reading that matches the registered break learning data r reading, or have a break learning that is longer than the registered break learning data ``reading'' and match the input reading;
In other words, this inconsistent segmentation learning data deletion process prevents the occurrence of a case where the learning effect cannot be obtained.

[他の実施例] 以上の説明においては、区切り学習データが登録される
タイミングとして、区切り変更をオペレータが明に指定
した場合、すなわち、区切り縮小キー、区切り伸長キー
を操作した場合について説明した。しかしオペレータが
区切り変更の方法を知らず、誤分割の変換結果が提示さ
れた時、入力読み列を次回は細かく分割して変換する可
能性も考えられ、この場合にも区切り学習を行なう必要
性がある。上記の場合にも区切り学習を行なうように構
成した実施例を以下に述べる。
[Other Embodiments] In the above description, the timing at which the delimiter learning data is registered is when the operator explicitly specifies a delimiter change, that is, when the delimiter reduction key or the delimiter expand key is operated. However, if the operator does not know how to change the delimiters and is presented with an incorrectly divided conversion result, there is a possibility that the input reading sequence will be divided into smaller parts the next time. be. An embodiment configured to perform segmentation learning also in the above case will be described below.

第15図は、上記に説明したオペレータの入力の例であ
る。(a)はテキスト画面上に既に「その」が入力され
ているの状態であり、オペレータは更にr問題に関して
以下の合意が得られた」と入力しようとしている。とこ
ろが、前回読みを続けて入力し誤変換されてしまったの
で、今回は人力単位を細かく分割することにし、「もん
だいにかんして/」と入力して変換を起動した。すると
、(b)の様に、正しくr問題に関して」と変換される
0次に(C)のように「いか9ごういがえられた/」と
入力する。その結果、(d)に示すよう゛に「以下の合
意が得られた」と変換される。この一連の動作を行なう
ことにより、「関して/以下の」の部分について区切り
学習が行なわれ、次回に、もし、「もんだいにかんして
いかのごついがえられた」と入力すれば正しく「問題に
関して以下の合意が得られた」と変換されるようになる
FIG. 15 is an example of the operator's input described above. In (a), ``that'' has already been input on the text screen, and the operator is about to input ``The following agreement has been reached regarding problem r''. However, last time I entered the reading in succession and it was converted incorrectly, so this time I decided to break up the manual unit into smaller units and started the conversion by typing ``Mondai Nikante/''. Then, as shown in (b), the user enters ``I was able to understand 9 squid/'' as shown in 0th order (C), which is converted to ``correctly regarding the r problem''. As a result, as shown in (d), it is converted to ``The following agreement has been obtained.'' By performing this series of operations, the section ``Regarding/Below'' will be learned in sections, and the next time, if you enter ``I was able to find out about a lot of things'', you will be able to correctly enter `` The following agreement has been reached regarding the issue.''

第10図で示した「仮名漢字変換」のフローチャートは
第16図に示す様に変更される。
The flow chart of "kana-kanji conversion" shown in FIG. 10 is changed as shown in FIG. 16.

ステップ16−1においてKBBUF上にある入力読み
列を単語辞書、文法辞書等を参照して形態素解析、構文
解析等を行ない、変換候補テープルHENTBLを作成
する。
In step 16-1, the input pronunciation sequence on KBBUF is subjected to morphological analysis, syntactic analysis, etc. with reference to a word dictionary, a grammar dictionary, etc., and a conversion candidate table HENTBL is created.

ステップ16−2において、第11図に詳述するように
、上記変換候補のうちどの変換候補を変換すべきか、第
1候補を決定する。
In step 16-2, as detailed in FIG. 11, a first candidate is determined which of the conversion candidates should be converted.

ステップ16−3において、今回の変換は前文節の変換
直後であるかどうかを判定する。変換直後であれば、ス
テップ16−4に進み、区切り学習データの登録処理を
行なうことになる。直後でなければステップ16−9に
分岐し、区切り学習データの登録処理をスキップする。
In step 16-3, it is determined whether the current conversion is immediately after the conversion of the previous clause. If it is immediately after conversion, the process proceeds to step 16-4, and registration processing of delimited learning data is performed. If it is not immediately after, the process branches to step 16-9 and the process of registering the delimited learning data is skipped.

ステップ16−4において、記憶されていた前回変換の
文節の読みを取り出し、今回の変換される先頭文節の読
みとマージし、区切り学習データ登録用の「読み」とす
る。
In step 16-4, the stored pronunciation of the previous converted clause is retrieved, merged with the pronunciation of the first clause to be converted this time, and used as the "reading" for registering the delimiter learning data.

ステップ16−5において、前回変換された文節の読み
数を取り出し、区切り学習データ登録用の「第1文節読
み数」とする。
In step 16-5, the number of pronunciations of the previously converted clause is taken out and set as the "first pronunciation number of clauses" for registration of delimited learning data.

ステップ16−6において今回変換される先頭文節の読
み数を取り出し、区切り学習データ登録用の「第2文節
読み数」とする。
In step 16-6, the number of pronunciations of the first clause to be converted this time is extracted and set as the "second pronunciation number" for registering the delimiter learning data.

ステップ16−7において上記設定された通りに区切り
学習データを¥IJ&する。
In step 16-7, the delimited learning data is ¥IJ& as set above.

ステップ16−8において、上記登録された区切り学習
データと矛盾する区切り学習データをサーチし、もし矛
盾する区切り学習データが見つかればそれを削除する。
In step 16-8, a search is made for delimiter learning data that is inconsistent with the registered delimiter learning data, and if contradictory delimiter learning data is found, it is deleted.

ステップ16−10において、決定された第1候補(採
用文節)を1文節ずつ浅学仮名混り文に変換していく。
In step 16-10, the determined first candidates (adopted clauses) are converted clause by clause into sentences containing shallow academic kana.

ステップ16−11において作成された変換文字列を文
書にセットする。更に文書の変更内容が分かるように変
換文字列がセットされた付近を表示する。
The converted character string created in step 16-11 is set in the document. Furthermore, the area near where the converted character string is set is displayed so that the content of changes to the document can be seen.

[発明の効果] 以上の説明から明らかなように本発明によれば、システ
ム側が間違った文節分割に基く変換を行なった場合であ
っても、オペレータが区切り変更を行なえば、区切り変
更前後の読みと、区切り変更の結果生じた文節分割位置
をシステムが区切り学習データとして記憶し、次回に同
じ読みを入力したときにその区切り学習データを参照し
て正しい変換を行なうことができ、毎回区切り変更キー
を操作するという煩わしさがなくなっている。しかも、
区切り学習データを起動する読みとして2文節分の読み
を記憶するので、区切り学習を行なったために別の状況
で誤変換を発生し、かえって不便になるという副作用も
無くしている。
[Effects of the Invention] As is clear from the above description, according to the present invention, even if the system side performs conversion based on wrong phrase division, if the operator changes the delimiter, the reading before and after the change is correct. Then, the system memorizes the bunsetsu division position resulting from the break change as break learning data, and the next time you enter the same reading, you can refer to that break learning data and perform the correct conversion. The hassle of operating the . Moreover,
Since the pronunciation of two clauses is stored as the pronunciation for activating the segment learning data, the side effect of erroneous conversion occurring in another situation due to segment learning and causing inconvenience is eliminated.

以上のように、操作性の高い文字処理装置を実現するこ
とができる。
As described above, it is possible to realize a character processing device with high operability.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は本発明の全体構成のブロック図第2図はCRT
の画面構成を示した口 笛3図は本発明における区切り伸長の操作方法を示した
図 第4図は本発明における区切り縮小の操作方法を示した
図 第5図は区切り学習データKULDTの構成を示した図 第6図は仮名漢字変換を行なう際に、処理途中で解析さ
れる文節候補の例を示した口笛7図は変換候補テーブル
HENTBLの構成を示した図 第8図はキーボードバッファKBBUFの構成を示した
図 第9図〜第14図は本発明文字処理装置の動作を示すフ
ローチャート 第15図は本発明の表示例を示す図 Zta図は本発明の動作の制御手順を示すフローチャー
ト DISK   ・・・外部記憶 CPU   ・・・マイクロプロセッサROM    
・・・誘出し専用メモリRAM    ・・・ランダム
アクセスメモリKBBUF  ・・・キーボードバッフ
ァKULDT  ・・・区切り学習データHENTBL
・・・変換候補テーブル
Figure 1 is a block diagram of the overall configuration of the present invention Figure 2 is a CRT
Whistle 3 shows the screen configuration of the screen. FIG. 4 shows the operation method of delimiter expansion in the present invention. FIG. 5 shows the configuration of the delimiter learning data KULDT. Fig. 6 shows an example of phrase candidates that are analyzed during the process when performing kana-kanji conversion. Whistle 7 shows the structure of the conversion candidate table HENTBL. Fig. 8 shows the structure of the keyboard buffer KBBUF. FIGS. 9 to 14 are flow charts showing the operation of the character processing device of the present invention. FIG. 15 is a flow chart showing display examples of the present invention.・External memory CPU...Microprocessor ROM
...Memory for extraction only RAM ...Random access memory KBBUF ...Keyboard buffer KULDT ...Separated learning data HENTBL
...conversion candidate table

Claims (4)

【特許請求の範囲】[Claims] (1)日本語の読み列を入力する入力手段と、読み列を
漢字仮名混り文に変換する変換手段 と、変換された漢字仮名混り文の文節の区切り方を変更
するための区切り変更手段と、区切りの変更対象となっ
た第1の文節及び第1の文節に引き続く第2の文節の2
文節分に対して区切り変更前の読みまたは区切り変更後
の読みを区切り変更後の文節の区切り方とともに記憶す
る区切り学習データと、前記区切り変更手段により文節
の区切り方を変更した際に起動され前記区切り学習デー
タを作成し登録する区切り学習データ登録手段とを設 け、前記変換手段は前記区切り学習データを参照し入力
読み列と一致する読みをもつ区切り学習データが存在し
た場合、その文節区切りに従って変換を行なう様に構成
され、オペレータが区切り変更手段を操作して文節の区
切り方を訂正した後、もう一度同じ読みで入力を行なっ
た時、オペレータの望むように変換されることを特徴と
する文字処理装置。
(1) An input means for inputting a Japanese pronunciation sequence, a conversion means for converting the pronunciation sequence into a sentence containing kanji and kana, and a delimiter change for changing the way in which clauses are separated in the converted sentence containing kanji and kana. the means, the first clause whose delimiter is subject to change, and the second clause following the first clause.
Separation learning data that stores the pronunciation before or after the separator change for the bunsetsu along with the way the bunsetsu is separated after the separator change, and the a break learning data registration means for creating and registering break learning data, and the converting means refers to the break learning data and, if there is break learning data having a pronunciation that matches the input pronunciation sequence, converts it according to the bunsetsu break. Character processing characterized in that when an operator operates a delimiter changing means to correct the way in which phrases are delimited and then inputs the same reading again, the character processing is converted as desired by the operator. Device.
(2)前記変換手段が前記区切り学習データを参照した
際、入力読みと一致する読みをもつ区切り学習データが
複数個あった場合、最長一致する区切り学習データの指
示に従うことを特徴とする特許請求の範囲第1項記載の
文字処理装置。
(2) A patent claim characterized in that when the converting means refers to the delimiter learning data and there is a plurality of delimiter learning data having pronunciations that match the input pronunciation, the instructions of the longest matching delimiter learning data are followed. The character processing device according to item 1.
(3)前記区切り学習データに記憶する読みは区切り変
更前の2文節分の読みと区切り変更後の2文節分の読み
のうち、長いほうの読みであることを特徴とする特許請
求の範囲第1項記載の文字処理装置。
(3) The pronunciation stored in the break learning data is the longer one of the readings for two clauses before changing the breaks and the readings for two clauses after changing the breaks. The character processing device according to item 1.
(4)前記区切り学習データ登録手段は区切り学習デー
タを登録後、それと矛盾する他の区切り学習データを削
除することを特徴とする特許請求の範囲第1項記載の文
字処理装置。
(4) The character processing device according to claim 1, wherein the delimiter learning data registration means deletes other delimiter learning data that contradicts the delimiter learning data after registering the delimiter learning data.
JP63041594A 1988-02-23 1988-02-23 Character processing apparatus and method Expired - Fee Related JP3029109B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP63041594A JP3029109B2 (en) 1988-02-23 1988-02-23 Character processing apparatus and method

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP63041594A JP3029109B2 (en) 1988-02-23 1988-02-23 Character processing apparatus and method

Publications (2)

Publication Number Publication Date
JPH01214967A true JPH01214967A (en) 1989-08-29
JP3029109B2 JP3029109B2 (en) 2000-04-04

Family

ID=12612730

Family Applications (1)

Application Number Title Priority Date Filing Date
JP63041594A Expired - Fee Related JP3029109B2 (en) 1988-02-23 1988-02-23 Character processing apparatus and method

Country Status (1)

Country Link
JP (1) JP3029109B2 (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH08272792A (en) * 1995-03-31 1996-10-18 Canon Inc Character processing apparatus and method

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS61156467A (en) * 1984-12-28 1986-07-16 Ricoh Co Ltd Word extraction method
JPS61173377A (en) * 1985-01-29 1986-08-05 Matsushita Electric Ind Co Ltd Forming device of japanese sentence
JPS62145463A (en) * 1985-12-20 1987-06-29 Ricoh Co Ltd Kana-kanji conversion method

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS61156467A (en) * 1984-12-28 1986-07-16 Ricoh Co Ltd Word extraction method
JPS61173377A (en) * 1985-01-29 1986-08-05 Matsushita Electric Ind Co Ltd Forming device of japanese sentence
JPS62145463A (en) * 1985-12-20 1987-06-29 Ricoh Co Ltd Kana-kanji conversion method

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH08272792A (en) * 1995-03-31 1996-10-18 Canon Inc Character processing apparatus and method

Also Published As

Publication number Publication date
JP3029109B2 (en) 2000-04-04

Similar Documents

Publication Publication Date Title
US5349368A (en) Machine translation method and apparatus
JPH07114568A (en) Data retrieval device
JPH01214967A (en) Character processing device and method
JP3278148B2 (en) Character processing apparatus and method
JP2862236B2 (en) Character processor
JPH0576065B2 (en)
JPH01214966A (en) character processing device
JP2771020B2 (en) Character processor
JP2744241B2 (en) Character processor
JPH04191966A (en) Character processor
JPH01204174A (en) Character processor
JP2688651B2 (en) String converter
JPH10187705A (en) Document processing method and apparatus
JPH03116369A (en) character processing device
JPH03116367A (en) Character processor
JPH0488550A (en) character processing device
JPH09223139A (en) Kana-Kanji conversion device
JPS62198956A (en) Character processor
JPH0638260B2 (en) Character processing apparatus and method
JPS62198953A (en) character processing device
JPH0638261B2 (en) Character processing apparatus and method
JPH03116368A (en) character processing device
JPH11175528A (en) Electronic dictionary
JPH0576064B2 (en)
JPH056360A (en) Character processor

Legal Events

Date Code Title Description
LAPS Cancellation because of no payment of annual fees