JPS6037954B2 - Feature extraction processing method - Google Patents

Feature extraction processing method

Info

Publication number
JPS6037954B2
JPS6037954B2 JP55012208A JP1220880A JPS6037954B2 JP S6037954 B2 JPS6037954 B2 JP S6037954B2 JP 55012208 A JP55012208 A JP 55012208A JP 1220880 A JP1220880 A JP 1220880A JP S6037954 B2 JPS6037954 B2 JP S6037954B2
Authority
JP
Japan
Prior art keywords
code
feature
point
white
character
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Expired
Application number
JP55012208A
Other languages
Japanese (ja)
Other versions
JPS56110188A (en
Inventor
茂 赤松
和昭 小森
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Inc
Original Assignee
Nippon Telegraph and Telephone Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Telegraph and Telephone Corp filed Critical Nippon Telegraph and Telephone Corp
Priority to JP55012208A priority Critical patent/JPS6037954B2/en
Publication of JPS56110188A publication Critical patent/JPS56110188A/en
Publication of JPS6037954B2 publication Critical patent/JPS6037954B2/en
Expired legal-status Critical Current

Links

Classifications

    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
    • G06V30/10—Character recognition
    • G06V30/18—Extraction of features or characteristics of the image
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
    • G06V30/10—Character recognition

Landscapes

  • Engineering & Computer Science (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Theoretical Computer Science (AREA)
  • Character Discrimination (AREA)

Description

【発明の詳細な説明】 本発明は、特徴抽出処理方式、特に位相幾何的にみて複
数な構造をもつ漢字等のパタンから識別に有効な特徴を
効果的に抽出する特徴抽出処理方式に関するものである
。
DETAILED DESCRIPTION OF THE INVENTION The present invention relates to a feature extraction processing method, and particularly to a feature extraction processing method that effectively extracts features useful for identification from patterns such as Chinese characters that have multiple structures from a topological perspective. be.

黒白2値パタン1個を含むパタン領域の各点から予め定
めた方向をみた時に見つかる位相幾何的特徴をして方向
情報を保存して各点に集積し、この結果得られる高次特
徴を用いて識別する文字論取方式として、特豚昭51一
85708号(特開昭53一10931号)がある。
The topological features found when looking in a predetermined direction from each point in a pattern area containing one black-and-white binary pattern are stored and accumulated at each point, and the resulting high-order features are used. Tokubuta No. 51-85708 (Japanese Unexamined Patent Publication No. 53-10931) is a method for character recognition.

この方式の基本部分である特徴抽出処理の原理を第1図
に示す。同図aにおいて、文字部分(黒)6を含むパタ
ン領域7の任意の白点5から、上下左右4方向に走査線
を出し(右方向1、上方向2、左方向3、下方向4)、
各方向に文字部分が存在するか杏かを調べる。各白点に
対応して4ビットの特徴ベクトル10が与えられ、上記
文字部分の存在状況に応じて存在すれば1を、存在しな
ければ0を特徴ベクトル10の所定の位置にセットする
。白点5が本例の位置にあるとき、この点の特徴ベクト
ル1川ま右方向走査線1に対応するビット位置11は1
、上方向走査線2に対応するビット位置12は0、左方
向走査線3に対応するビット位置13は1、そして下方
向走査線4に対応するビット位置14は1である。以上
の処理を1次位相特徴抽出と呼ぶ。この処理によってパ
タン領域7の白点領域は各白点が如何なる特徴ベクトル
をもつかに従っていくつかに分割される。本例では同図
bに図形的に示すように文字部分6の閉じ状況に応じて
図示21,22,23,24,25,26,27の如き
特徴をもつ特徴ベクトルが抽出される。例えば、特徴ベ
クトル21は右方向1および下方向4のみに文字部分が
あること、特徴ベクトル22は上方向2のみに空きがあ
ること、特徴ベクトル24は空き方向がないこと等を図
的に表わす。さて次に、1次位相特徴抽出にパタン領域
7の各白点に形成された位相幾何的特徴すなわち特徴ベ
クトル10を基に高次特徴を形成するようにする。
The principle of feature extraction processing, which is the basic part of this method, is shown in FIG. In the same figure a, scanning lines are drawn in four directions (rightward direction 1, upward direction 2, leftward direction 3, downward direction 4) from an arbitrary white point 5 in the pattern area 7 including the character part (black) 6. ,
Check whether a character part exists in each direction. A 4-bit feature vector 10 is given corresponding to each white point, and 1 is set in a predetermined position of the feature vector 10 if the character part exists, and 0 is set if it does not exist, depending on the existence status of the character part. When the white point 5 is at the position of this example, the bit position 11 corresponding to the feature vector 1 of this point or the rightward scanning line 1 is 1.
, bit position 12 corresponding to upper scan line 2 is 0, bit position 13 corresponding to left scan line 3 is 1, and bit position 14 corresponding to lower scan line 4 is 1. The above processing is called primary phase feature extraction. By this process, the white point area of the pattern area 7 is divided into several parts according to what kind of feature vector each white point has. In this example, feature vectors having features 21, 22, 23, 24, 25, 26, and 27 are extracted according to the closed state of the character portion 6, as graphically shown in FIG. For example, the feature vector 21 graphically represents that there is a character part only in the right direction 1 and the bottom direction 4, the feature vector 22 graphically represents that there is space only in the top direction 2, and the feature vector 24 graphically represents that there is no space in the direction 4. . Next, in primary phase feature extraction, high-order features are formed based on the topological features formed at each white point of the pattern area 7, that is, the feature vectors 10.

このため各白点については自分自身の特徴ベクトル10
の他に、各方向に走査線を出したときに文字部分6を越
えて初めて出合う白点の特徴ベクトルを方向情報を保存
して追加する。白点5が本体の位置にあるときは、その
高次特徴は図的に表わせば図示高次特徴30のようにな
る。該特徴3川こおいて小さな丸31は白点5の上方向
2に文字部分がないことを示している。また各黒V点‘
こついては、その点から各方向に走査線を出した時に初
めて出合う白点の特徴ベクトルを方向情報を保存して追
加する。黒点8が本例の位置にあるときはその高次特徴
は図的に表わせば図示高次特徴40のようになる。なお
特顔昭51一85708号(特開昭53一10931号
)では1次位相特徴抽出された特徴ベクトル10の中で
文字変形に対し不安定なものを安定なものに変換する統
合処理を行なう。例えば、特徴ベクトルのうちの特徴ベ
クトル24はループと区別がつかないので該ベクトル2
4と接している白点特徴ベクトル26に変換される。ま
た黒V点に対してその点の隣接点の白黒の状態をみて局
所懐き情報を与え、これを黒V点の1次位相特徴とする
特徴抽出法、持豚昭51一121792号(特関昭53
−47239号)がある。これら上記万式は手書きある
いは印字の英数字、カタカナ等の比較的単純な位相幾何
的構造をもったパタンの特徴抽出法に適している。そし
て入力文字の識別は上記高次特徴の存在状況とカテゴリ
毎に予め用意されている高次特徴のリストと比較して行
なうようにされる。上記従来方式では、背景白部の点に
与えられる位相幾何的特徴は当該白点を中心に上下左右
方向に文字枠まで走査線を伸ばした時の文字線の有無に
よって閉じ状態を記述する一次位相特徴を基本として構
成されている。
Therefore, for each white point, its own feature vector 10
In addition, when a scanning line is drawn in each direction, feature vectors of white points that are encountered for the first time beyond the character portion 6 are added while preserving direction information. When the white point 5 is located at the position of the main body, its higher-order feature becomes the illustrated higher-order feature 30 if expressed graphically. A small circle 31 in the feature 3 indicates that there is no character part in the upper direction 2 of the white point 5. Also, each black V point'
The trick is to save the direction information and add the feature vector of the white point that is encountered for the first time when scanning lines are drawn in each direction from that point. When the black point 8 is located at the position of this example, its higher-order feature is represented graphically as the illustrated higher-order feature 40. In addition, in Tokugan No. 51-85708 (Japanese Patent Application Laid-Open No. 53-10931), an integration process is performed to convert those that are unstable against character deformation into those that are stable among the feature vectors 10 extracted from the primary phase features. . For example, feature vector 24 of the feature vectors is indistinguishable from a loop, so
It is converted into a white point feature vector 26 that is in contact with 4. In addition, a feature extraction method that gives local familiarity information to a black V point by looking at the black and white state of the points adjacent to that point, and uses this as the primary phase feature of the black V point, Mochibuta Sho 51-1121792 (Special Showa 53
-47239). These above-mentioned formulas are suitable for extracting features of patterns with relatively simple topological structures such as handwritten or printed alphanumeric characters and katakana characters. The input character is identified by comparing the presence status of the higher-order features with a list of higher-order features prepared in advance for each category. In the conventional method described above, the topological feature given to a point in the white background is a primary phase that describes the closed state depending on the presence or absence of a character line when the scanning line is extended vertically and horizontally from the white point to the character frame. It is structured based on characteristics.

しかし、上記特徴では、漢字のように文字線の縦・横・
斜めの方向性の区別が比較的明確なパタンであっても、
対象文字の図形的変形・雑音によって生じた特徴と、安
定なストロークの構造を記述した特徴とを、特徴抽出の
段階で区別することが困難であること、並びに、漢字の
部首に代表されるような部分パタンの存在を文字全体の
中から区別して特徴づける際に必要となる文字線の分離
状態を表わす情報が充分でないこと等、特徴のもつ位相
幾何的情報の不足を原因とする欠点をなお含んでいる。
本発明は、これらの欠点を解決するために、パタン面上
の白点における位相幾何的特徴の基本となる1次特徴を
、予め定められた方向に存在する文字線の額きを表現し
ている傾斜コードを該白点上の集積させることによって
求め、併せて該白点の周辺の文字線の連結性をも表現し
うるようにすることを目的としており、以下図面につい
て詳細に説明する。
However, with the above characteristics, the vertical, horizontal, and
Even if the diagonal direction is relatively clear in the pattern,
It is difficult to distinguish at the feature extraction stage between features caused by graphical deformation/noise of the target character and features that describe stable stroke structure, and The disadvantages caused by the lack of topological information of the characteristics, such as the insufficient information representing the separation state of character lines, which is necessary to distinguish and characterize the existence of partial patterns from the entire character, It is included.
In order to solve these drawbacks, the present invention expresses the primary feature, which is the basic topological feature of the white point on the pattern surface, as the frame of the character line existing in a predetermined direction. The purpose of this invention is to obtain the slanted code by accumulating it on the white point, and also to express the connectivity of the character lines around the white point.The drawings will be described in detail below.

第2図は本発明の実施例であって、5川ま走査部、6川
まパタンメモリ部、61はパタンメモリ、7川ま制御部
、80は識別部、9川ま特徴抽出部を表わす。
FIG. 2 shows an embodiment of the present invention, which shows a 5-way scanning section, a 6-way pattern memory section, 61 a pattern memory, a 7-way control section, 80 an identification section, and a 9-way feature extraction section. .

該特徴抽出部90の構成要部として、910川ま傾斜コ
ード抽出部、9200は特徴集積部、9300は特徴コ
ード変換部、9400は特徴点抽出部、9500は背景
分割線抽出部、9600は周辺閉じ状態コード抽出部で
ある。これを動作するには、制御部70の指令により走
査部50が紙面を走査し、用紙上の1文字を含む領域の
黒及び白を表わす2値信号がパタンメモリ61の所定の
アドレスに順次格納される。
The main components of the feature extraction unit 90 include a river slope code extraction unit 910, a feature accumulation unit 9200, a feature code conversion unit 9300, a feature point extraction unit 9400, a background dividing line extraction unit 9500, and a surrounding area 9600. This is a closed state code extraction part. To operate this, the scanning unit 50 scans the paper surface according to a command from the control unit 70, and binary signals representing black and white of an area including one character on the paper are sequentially stored at a predetermined address in the pattern memory 61. be done.

パタンメモリ61上では、パタン領域内の1メッシュに
対して予め定められたワード数から成るメモリアドレス
が割りあてられている。パタン領域内の信号がパタンメ
モリ61に格納され終ると、制御部70の指令を受け、
特徴抽出部9川ま予め定められた手順でパタンメモリ6
1の内容を読み出し、格納された文字パタン上の各点に
その位相幾何的特徴を表わす特徴コードを求め、パタン
メモリ上の対応するアドレスに格納する。特徴抽出部9
0を用いた一連の特徴コード作成処理が終了すると、パ
タンメモリ61の内容は識別部8川こ順次転送される。
識別部80では、パタン領域の各メッシュに対応するパ
タンメモリ61上のアドレスに記憶された特徴コードを
読み出し、その種類と各々の個数とを求めて、この情報
に基づいて入力文字が何であったかを決定し、その結果
を出力する。特徴抽出部90では、まず額斜コード抽出
部9100でパタンメモリ61上に格納された文字パタ
ン上の黒V点に対して、その文字線の煩きを表わす額斜
コードを形成する。
On the pattern memory 61, a memory address consisting of a predetermined number of words is assigned to one mesh within the pattern area. When the signals in the pattern area have been stored in the pattern memory 61, receiving a command from the control unit 70,
The feature extraction unit 9 extracts the pattern from the pattern memory 6 according to a predetermined procedure.
1 is read out, a feature code representing the topological feature of each point on the stored character pattern is obtained, and is stored in the corresponding address on the pattern memory. Feature extraction section 9
When a series of feature code creation processes using 0 are completed, the contents of the pattern memory 61 are sequentially transferred to the identification unit 8.
The identification unit 80 reads out the feature codes stored in the addresses on the pattern memory 61 corresponding to each mesh in the pattern area, determines the type and number of each, and determines what the input character was based on this information. Decide and output the result. In the feature extraction unit 90, first, a forehead slant code extraction unit 9100 forms a forehead slant code representing the harshness of the character line for the black V point on the character pattern stored on the pattern memory 61.

頚斜コード抽出部9100による動作が終了すると、制
御部70の指令により、予め定められた手順に従ってパ
タンメモリ部60はパタンメモリ61を上下左右に走査
しつつその内容を読み出し、特徴集積部9200に転送
する。特徴集積部9200ではパタン領域の白点に対し
て該白点の上下左右方向に存在する文字線上の黒点から
煩斜コードを選択して抽出しこれに基づいて該白点より
ながめた時の該文字線の煩きを表わす頃斜コードを求め
、該コードを予め定められたビッド構成に従って配分す
ることにより集積コードを作成し、これをパタンメモリ
61上の該白点‘こ対応するアドレスの転送する。特徴
集積部9200の動作が終了した後特徴コード変換部9
300では、制御部70の指令によりパタンメモリ61
を上下左右に走査する過程を読み出された前記額斜コー
ドを含む特徴コードを変換する。即ち該走査線上にあっ
てパタンメモリ61上の隣接する点から読み出された特
徴コードを参照することにより、必要ならば予め定めら
れた規則に従って該額斜コードを変換し、パタンのより
安定な構造を記述するよう特徴コードを作成し、これを
パタンメモリ61上の対応するアドレスに格納する。以
上に示す特徴集積部9200及び特徴コード変換部93
00の動作とは独立に特徴点抽出部9400は、予め定
められた手順でパタンメモリ61の内容を読み出し、格
納された文字パタン上の端点、屈曲点などの位相幾何的
特徴点を検出する。特徴点抽出部9400動作が終了し
た後、背景分割線抽出部9500では、予め定められた
手順に従ってパタンメモリ61を走査する過程で、前記
特徴点抽出部94001こよって抽出された位相幾何的
特徴点が検出された時に、該特徴点の種別に応じて予め
定められた方向に該走査方向が一致するとき該特徴点を
起点として背景白部を分割する線分を延ばし、パタンメ
モリ上の該線分に対応する点のアドレスに該背景分割線
を表わすコードを格納する。周辺閉じ状態コード抽出部
9600では、前記特徴集積部9200及び背景分割線
抽出部9500‘こよる動作が終了した後、制御部70
の指令によりパタンメモリ61を上下左右に走査し、文
字線部分或いは背景分割線によって仕切られた任意の白
部領域に含まれる該走査線上の白点‘こ対して、該走査
線が文字線部分を通過してから該白熱こ至るまでの間に
交叉した背景分割線の回数とその方向、ならびにその間
に白点より読み出された特徴コード上に集積された方向
別の文字線傾斜コードの種類別出現個数などに応じて、
該白部領域の周囲の文字線閉じ状態を表現する特徴コー
ドを作成し、該白点に対応するアドレスに定められたビ
ット構成で格納する。以上に示す各部の動作によって、
特徴抽出部90はパタンメモリ61に格納された文字パ
タン上の各点にその位相幾何的特徴を表わす特徴コード
を格納する。次に特徴抽出部90を構成する各処理部に
ついて詳細に説明する。傾斜コード抽出部9100は、
文字線上の任意の黒V点から予め定められた各方向別の
黒V点の連続数を計数し方向別距離を求め、該方向別距
離を比較することにより文字線の方向成分を決定するな
どの方法によって、文字パタンを構成するストロークの
額斜を表わす煩斜コードを形成する。
When the operation by the cervical oblique code extraction section 9100 is completed, the pattern memory section 60 reads out the contents while scanning the pattern memory 61 vertically and horizontally according to a predetermined procedure according to a command from the control section 70, and stores the contents in the feature accumulation section 9200. Forward. The feature accumulation unit 9200 selects and extracts a slanting code from the black dots on the character lines that exist in the vertical and horizontal directions of the white dot in the pattern area, and based on this selects and extracts the oblique code from the black dots when viewed from the white dot. An integrated code is created by obtaining a diagonal code that represents the irregularity of character lines and allocating the code according to a predetermined bit configuration, and transfers this code to the address corresponding to the white dot on the pattern memory 61. do. After the operation of the feature accumulation unit 9200 is completed, the feature code conversion unit 9
300, the pattern memory 61 is
The characteristic code including the forehead oblique code read out is converted by scanning the screen vertically, horizontally, and horizontally. That is, by referring to the feature codes read from adjacent points on the pattern memory 61 on the scanning line, the forehead diagonal code is converted according to a predetermined rule, if necessary, to make the pattern more stable. A feature code is created to describe the structure and is stored at a corresponding address on the pattern memory 61. Feature accumulation unit 9200 and feature code conversion unit 93 shown above
Independently of the operation of 00, the feature point extraction unit 9400 reads the contents of the pattern memory 61 in a predetermined procedure and detects topological feature points such as end points and bending points on the stored character patterns. After the operation of the feature point extraction unit 9400 is completed, the background dividing line extraction unit 9500 extracts the topological feature points extracted by the feature point extraction unit 94001 in the process of scanning the pattern memory 61 according to a predetermined procedure. is detected, and when the scanning direction matches a predetermined direction according to the type of the feature point, a line segment that divides the white part of the background is extended starting from the feature point, and the line segment on the pattern memory is A code representing the background dividing line is stored at the address of the point corresponding to the minute. In the peripheral closed state code extraction section 9600, after the operations of the feature accumulation section 9200 and the background dividing line extraction section 9500' are completed, the control section 70
The pattern memory 61 is scanned vertically and horizontally according to the command, and when a white point on the scanning line included in a character line part or an arbitrary white area partitioned by a background dividing line is detected, the pattern memory 61 is scanned vertically and horizontally. The number of times the background dividing line intersects and its direction from passing through to reaching the incandescent point, as well as the type of character line slope code by direction accumulated on the feature code read from the white point during that time Depending on the number of occurrences etc.
A feature code expressing a character line closed state around the white area is created and stored in a predetermined bit configuration at an address corresponding to the white point. By the operation of each part shown above,
The feature extraction unit 90 stores a feature code representing the topological feature of each point on the character pattern stored in the pattern memory 61. Next, each processing unit making up the feature extraction unit 90 will be explained in detail. The slope code extraction unit 9100
The directional component of the character line is determined by counting the number of consecutive black V points in each predetermined direction from an arbitrary black V point on the character line to obtain the distance in each direction, and comparing the distances in each direction. By the method described above, a slant code representing the slant of strokes forming a character pattern is formed.

この種のストローク額斜抽出法は既に幾つか知られてお
り、公知の技術を利用することができる。第3図は、文
字パタン上の黒V点に対応するアドレスにおいて、傾斜
コード抽出部9100の出力として得られる頭斜コード
等を表現するビット構成の一実施例と、該傾斜コードに
よって記述されるストロークの方向成分の定義を示す。
本実施例ではストロークの方向を、右方向51と左方向
52、上方向53と下方向54、右上45o方向55と
左下45o方向56、左上45o方向57と右下45o
方向58の4対計8方向に童子化するものとする。61
0はパタンメモリ61上のメッシュに対応するアドレス
に用意された1語以上のメモリから構成される特徴コー
ドであり、該メッシュが文字線上に位置するとき黒白情
報部611は「1」となつて、額斜コード部620‘ま
該黒V点上に検出される頃斜コードを格納するビット領
域となる。
Several stroke amount oblique extraction methods of this type are already known, and known techniques can be used. FIG. 3 shows an example of a bit configuration expressing a prefix code obtained as an output of the slant code extracting section 9100 at an address corresponding to a black V point on a character pattern, and the bit structure described by the slant code. This shows the definition of the directional component of the stroke.
In this embodiment, the stroke directions are rightward 51, leftward 52, upward 53, downward 54, upper right 45o direction 55, lower left 45o direction 56, upper left 45o direction 57, lower right 45o direction.
It is assumed that the image is transformed into a doji in a total of 8 directions (4 pairs of directions 58). 61
0 is a feature code consisting of one or more words prepared at the address corresponding to the mesh on the pattern memory 61, and when the mesh is located on the character line, the black and white information section 611 becomes "1". , the forehead diagonal code section 620' is a bit area for storing the forehead diagonal code detected on the black V point.

ここでビット位置621,622,623,624,6
25,626,627,628は各々ストローク方向5
1,52,53,54,55,56,57,58に対応
している。たとえば、該黒V点を中心として51方向に
ストロークが検出された場合にはビット621を「1」
とし、逆に51方向にストロークが検出されないときビ
ット621を「0」とすることにより、該黒点における
頃斜コード620を8方向のストロークの有無を表わす
8ビットの領域でコード化する。第4図は、パタンメモ
リ61上の2値パタン500と、煩斜コード抽出部91
001こよって形成された該パタンの文字線上黒V点の
懐斜コードの一例を幾つかの点501,502,503
によって図示したものである。
Here bit positions 621, 622, 623, 624, 6
25, 626, 627, 628 are each stroke direction 5
1, 52, 53, 54, 55, 56, 57, and 58. For example, if a stroke is detected in 51 directions centering on the black point V, bit 621 is set to "1".
On the other hand, when no stroke is detected in the 51 direction, the bit 621 is set to "0", thereby encoding the rotational oblique code 620 at the sunspot with an 8-bit area representing the presence or absence of a stroke in the 8 directions. FIG. 4 shows the binary pattern 500 on the pattern memory 61 and the complex code extraction section 91.
001 An example of the slant code of the black V point on the character line of the pattern thus formed is shown as several points 501, 502, 503.
It is illustrated by.

ビット列506は図的表現511に示すように点501
が53,54方向に連なる垂直方向のストローク上にあ
ることを示す頃斜コードであり、ビット列507は図的
表現512に示すように点502が55方向にのみ連な
る右上り方向のストローク上にあることを示す煩斜コー
ドである。またビット列508は図的表現513に示す
様に点503においては54,55,56方向のみなら
ず52,53方向についての文字線の連なりをも検出さ
れたことを示す煩斜コードであり、該文字パタン500
のスケルトンのもつ構造とは矛盾する雑音とみなすべき
コードであるが、文字線が変動する太さをもっているこ
とから、このような雑音の発生を抑えることは困難であ
る。この種の雑音に対しては特徴集積部9200におて
対処することができる(後述)。第5図は、特徴集積部
9200をはじめとして特徴コード変換部9300背景
分割線抽出部9500、周辺閉じ状態コード抽出部96
00等によって、パタン領域上の任意の白点に形成され
る特徴コード630のビット構成一実施例と、該特徴コ
ードを構成し該白点の周囲4方向の文字線懐斜情報を表
現する集積コードの一実施例を示す。ビット列651,
652,653,654はそれぞれ任意の白点から予め
定められた方向に走査線を出した時に水平方向、垂直方
向、右上り方向、左上り方向の煩斜をもつ文字線が存在
することを表わす集積コード、ビット列655は該走査
線延長上に文字線の存在しないことを表わす集積コード
、ビット列656は該走査線上に最初に存在する文字線
上の黒点より抽出される頚斜コードが額斜コード判定回
路9211によっていずれも不安定と判断され該文字線
の傾斜を一意にコード化できなかったことを表わす集積
コードであり、これらに対応する記号661,662,
663,664,665,666は、第8図に一例を示
す白点上の集積コードの図的表現に用いるシンボルであ
る。白点の特徴コード63川ま、黒V点と白点とを弁別
するビット631と、文字線の煩斜を表わす前記集積コ
ードを該白点を起点として該文字線の存在する方向に対
応して右方向51を641、左方向52を642、上方
向53を643、下方向54を644のビット位置に配
置した集積コ−ド部640と、後述の特徴コード変換部
9300において該集積コードをそれぞれコード変換し
た結果を格納する変換済み集積コード部645と、後述
の背景分割線抽出部9500及び周辺閉じ状態コード抽
出部960川こよって形成されるその他の集積情報をコ
ード化し豚部分635とから成る。第6図は特徴集積部
9200の一実施例をブロック図で示したものである。
The bit string 506 is connected to the point 501 as shown in the graphical representation 511.
is a diagonal code indicating that the point 502 is on a vertical stroke extending in the 53 and 54 directions, and the bit string 507 is on a stroke in the upward right direction where the point 502 is extending only in the 55 direction, as shown in the graphical representation 512. This is a complicated code indicating that. Further, the bit string 508 is a slanting code indicating that, as shown in the graphical representation 513, at the point 503, a series of character lines not only in the 54th, 55th, and 56th directions but also in the 52nd and 53rd directions was detected. 500 character patterns
This code should be considered as noise that is inconsistent with the structure of the skeleton, but since the character lines have varying thicknesses, it is difficult to suppress the occurrence of such noise. This type of noise can be dealt with in the feature accumulation unit 9200 (described later). FIG. 5 shows a feature accumulation section 9200, a feature code conversion section 9300, a background dividing line extraction section 9500, and a peripheral closed state code extraction section 96.
An example of the bit configuration of a feature code 630 formed at an arbitrary white point on a pattern area by 00 etc., and an accumulation that constitutes the feature code and expresses character line nostalgic information in four directions around the white point. An example of code is shown. bit string 651,
652, 653, and 654 indicate that when a scanning line is drawn in a predetermined direction from an arbitrary white point, there are character lines with horizontal, vertical, upward-right, and upward-left directions. The bit string 655 is an integrated code indicating that there is no character line on the extension of the scanning line, and the bit string 656 is the cervical oblique code extracted from the black dot on the first character line on the scanning line. This is an integrated code indicating that the slope of the character line could not be uniquely coded because it was judged as unstable by the circuit 9211, and the corresponding symbols 661, 662,
663, 664, 665, and 666 are symbols used to graphically represent the integrated code on the white point, an example of which is shown in FIG. The feature code 63 of the white point, the bit 631 for distinguishing between the black V point and the white point, and the integrated code representing the slant of the character line correspond to the direction in which the character line exists starting from the white point. The integrated code section 640 arranges the right direction 51 at 641, the left direction 52 at 642, the upward direction 53 at 643, and the downward direction 54 at bit positions 644, and the feature code conversion section 9300 described later converts the integrated code into the integrated code section 640. The converted accumulated code section 645 stores the results of code conversion, and the pig section 635 encodes other accumulated information formed by the background dividing line extraction section 9500 and surrounding closed state code extraction section 960, which will be described later. Become. FIG. 6 is a block diagram showing one embodiment of the feature accumulation section 9200.

制御部70の指令にもとず〈文字パタン上の走査によっ
て、パタンメモリ61より読み出された黒V点の特徴コ
ード及び白点が結線303を通じて入力レジスタ921
0に絡納される。なおこの時点では白点に対応する特徴
は未だ形成されていない。黒点から転送された特徴コー
ド‘こ対しては懐斜コード判定回路9211が動作し第
3図に示す該特徴コード610上の傾斜コード620が
、後述の判定論理により、該文字線の煩斜情報を安定に
表現しているものと判定ざれた場合にのみ、集積コード
選択部9212が動作し、第5図で示したビット列65
1,652,653,654,655,656によつて
表わされる集積コードのうち、該黒点の煩斜情報を該走
査方向の白点の集積させる場合に対応する集積コードを
選択して出力し、3ビットから成る集積コードレジスタ
9213の内容を更新する。また、入力レジスタ921
0上のデータが白点から転送された場合には集積コード
合成部9215が動作し、集積コードレジスタ9213
の内容が入力レジスタ9210上の白点特徴コード63
0の頃斜コード集積部分640内の第5図で述べたよう
に該走査方向によって定まる集積方向に対応した所定ビ
ット位置に格納される。この処理の結果は走査方向と相
対する方向の文字線煩斜コードが該白点の特徴コード上
に集積されることを意味する。つづいて制御部70より
の指令により集積コード合成部9215で作られた1語
が結線304を通してパタンメモリ61上の対応するア
ドレスに格納される。煩斜コード判定回路9211によ
る額斜情報の安定性の判定論理の一実施例を以下に示す
。
Based on the command from the control unit 70, the characteristic code of the black V point and the white point read out from the pattern memory 61 by scanning the character pattern are input to the input register 921 through the connection 303.
Consolidated into 0. Note that at this point, features corresponding to the white dots have not yet been formed. In response to the characteristic code transferred from the black dot, the oblique code determination circuit 9211 operates, and the oblique code 620 on the characteristic code 610 shown in FIG. The integrated code selection unit 9212 operates only when it is determined that the bit string 65 shown in FIG.
Selecting and outputting an accumulation code among the accumulation codes represented by 1,652, 653, 654, 655, and 656 that corresponds to the case where the oblique information of the black point is accumulated of the white point in the scanning direction, The contents of the integrated code register 9213 consisting of 3 bits are updated. In addition, the input register 921
If the data above 0 is transferred from the white point, the integrated code synthesis unit 9215 operates, and the integrated code register 9213
The contents of the white point feature code 63 on the input register 9210
The value around 0 is stored at a predetermined bit position in the oblique code accumulation portion 640 corresponding to the accumulation direction determined by the scanning direction as described in FIG. The result of this processing means that the character line slanting code in the direction opposite to the scanning direction is accumulated on the characteristic code of the white point. Subsequently, one word created by the integrated code synthesis section 9215 according to a command from the control section 70 is stored at the corresponding address on the pattern memory 61 through the connection 304. An example of the logic for determining the stability of forehead oblique information by the oblique code determination circuit 9211 will be described below.

第3図に示したように、黒点より抽出される特徴コ‐ド
610上の鏡斜コード部620において、ビット対62
1,622、ビット対623,624、ビット対625
,626、ビット対627,628等はいずれか一方ま
たは双方のビットが「1」となることによって、それぞ
れ水平、垂直、右上り、左上りのストローク成分の存在
をあらわす。そこで上記4つのビット対のうちいずれか
1つだけ「1」なるビットが存在し、他はすべて「0」
となる銭斜コードを、文字線の煩斜情報を安定に表現し
ているものとみなす。そして、複数のビット対に「1」
なるビットが存在するときには、該コード‘こよって文
字線の鏡斜を一意に定めることは困難である為、これを
不安定な煩斜情報とみなすことにする。また、例えば右
水平方向51に走査する過程では、ビット対621,6
22のうち左水平方向52に対応するビット622のみ
「1」となる傾斜コードを不安定な情報コードとして捨
て去るなど、走査方向の情報を考慮した安定性の判定も
可能である。
As shown in FIG. 3, bit pairs 62
1,622, bit pair 623,624, bit pair 625
, 626, bit pairs 627, 628, etc., when one or both bits are set to "1", they represent the presence of horizontal, vertical, upward right, upward left stroke components, respectively. Therefore, of the four bit pairs mentioned above, only one bit is "1", and all others are "0".
We regard the zeni-shaku code as stably expressing the character line's angular information. And "1" for multiple bit pairs
When a bit exists, it is difficult to uniquely determine the mirror slant of a character line based on the code ', so this is regarded as unstable slant information. For example, in the process of scanning in the right horizontal direction 51, bit pairs 621, 6
It is also possible to determine stability in consideration of information in the scanning direction, such as by discarding a slope code in which only bit 622 corresponding to the left horizontal direction 52 is "1" as an unstable information code.

第7図はその一例であって、右方向51に走査するとき
、文字パタン520の文字線上の点521における、図
的表現526によって示される額斜コード‘ま、垂直方
向に安定と判断されて集積コードレジスタ9213に転
送されるが、点522における図的表現527で示され
る煩斜コードはビット622のみ「1」となることから
上記理由により不安定とみなされ、集積コード・レジス
タに書き込まれない。このことによって、走査線延長上
の白点523では、点522近傍の突起部分の影響が除
去され左方向52の文字線情報として、文字線の3分岐
点に相当する点521近傍の情報を保存して集積するこ
とができる。以上実例を示したような榎斜情報の安定性
に関する判定にはケース・バィ・ケースで種々の方法が
可能であるが、詳細については省略する。懐斜コード判
定回路9211の動作により、集積コード・レジスタ9
213の内容は文字線上の黒点を走査する過程で次々と
更新され、白点に対しては、走査線上に位置する安定な
額斜情報をもつ黒点のうちで最も該白則こ近い点の煩斜
情報がコード化されて集積されることになる。
FIG. 7 is an example of this, in which when scanning in the right direction 51, the forehead oblique code indicated by the graphical representation 526 at the point 521 on the character line of the character pattern 520 is determined to be stable in the vertical direction. It is transferred to the integrated code register 9213, but since only bit 622 of the confusing code shown by the graphical representation 527 at point 522 is "1", it is considered unstable for the above reason and is written to the integrated code register. do not have. As a result, at the white point 523 on the extension of the scanning line, the influence of the protrusion near the point 522 is removed, and the information near the point 521, which corresponds to the three branch points of the character line, is saved as character line information in the left direction 52. and can be accumulated. Various methods are possible on a case-by-case basis for determining the stability of the Enoki oblique information as shown in the example above, but the details will be omitted. By the operation of the nostalgic code determination circuit 9211, the integrated code register 9
The contents of 213 are updated one after another in the process of scanning the black dots on the character line, and for white dots, the content of the point closest to the white dot among the black dots with stable forehead slant information located on the scanning line is updated. The oblique information will be encoded and accumulated.

イニシャル・セット回路9214は制御部70の指令に
もとづき、集積コード・レジスタ9213上に予め定め
られた集積コードを書き込んで該レジスタの内容を初期
化する。
Initial set circuit 9214 writes a predetermined integrated code into integrated code register 9213 based on a command from control unit 70 to initialize the contents of the register.

初期化の内容は、周辺文字枠からパタン領域内に走査が
開始される時′点では該走査線が初めて文字線を横切る
までの間の白点に書き込む為の集積コード655(第5
図)が、また該走査線が白点から黒V真に移る時点では
該文字線上の黒点から安定な懐斜コードが1つも抽出さ
れない場合があることの為に集積コード656(第5図
)が、それぞれ集積コードレジスタ9213に転送され
る。以上述べた処理によって、パタン領域内のすべての
白点に対応するアドレスに、該白点から上下左右方向に
走査線を延ばした時初めて横切る文字線の傾斜をコード
化して集積した特徴コード630が形成される。
The content of the initialization is that when scanning starts from the peripheral character frame into the pattern area, an integrated code 655 (5th
However, at the time when the scanning line moves from the white point to the black line, there may be cases where no stable slant code is extracted from the black point on the character line, so the accumulated code 656 (Fig. 5) is used. are transferred to the integrated code register 9213, respectively. Through the above-described processing, feature codes 630 are accumulated at the addresses corresponding to all the white points in the pattern area, encoding the slope of the character line that first crosses when the scanning line is extended in the vertical and horizontal directions from the white point. It is formed.

第8図に一例としてその図的表現を示す。図中白点50
4には、図的表現531に示すように、上方向には右上
りの文字線、右方向には垂直の文字線がそれぞれ存在し
、さらに左方向及び下方向には文字枠まで文字線が存在
しないことを表わす特徴コードが形成される。一方、白
点505には、図的表現532に示すように、上方向及
び左方向には文字線が存在せず、下方向には右上りの文
字線が存在し、さらに右方向には文字線は存在するもの
の点505より右方向に延ばした走査線上に初めて出現
する黒V点列からは例えば黒点503に形成される煩斜
コード508(第4図)に代表されるように「不安定」
と判定される頃斜コードしか抽出されなかったことを表
わす特徴コードが形成される。白点505に形成された
特徴にみられるようにある方向についての文字線の懐き
に関して不確定集積コード(図的表現532における*
印)をもつ不確定特徴コードは、そのまま放置して識別
処理の段階で該特繊コードをdontcareとするな
どの対策も不可能ではないが、多くの漢字のように文字
パタンが複雑になるにつれてこの種の不確定特徴の出現
頻度が高まり、その結果パタンの詳細の構造を反映した
特徴の抽出が困難となるなど問題が多い。この問題は、
複数の文字線傾斜情報抽出手段の組み合せによる集積コ
ード決定法により解決可能である。第6図に示す輪郭形
状判定部922川まその一実施例をブロック図で示した
ものである。判定回路9221は入力レジスタ9210
の内容を調べることによって走査が文字線上の黒点から
白点‘こ移った時点を検出し、この時点の集積コードレ
ジス夕9213の内容を読み出す。該集積コードレジス
夕中に格納された集積コードが前記不確定集積コードで
あると判定されると、文字線輪郭追跡部9222は制御
部70に結線701を介して信号を出す。制御部70は
これを受けてパタンメモリ部6川こ現時点のパタンメモ
リレジスタを退避保存させるとともに、一時点前のアド
レス(不安定集積コードをもつ黒点に対応する)を起点
として一定範囲のパタン輪郭を追跡させ各点のアドレス
を結線305を介して逐次文字線輪郭追跡部9222に
転送する。この時、額斜判定部9223は、例えば文字
線輪郭追跡部9222により発せられる議論郭線上点の
アドレス情報の縦及び藤方向成分の変化分等により、該
輪郭線の懐斜を判定し、該判定結果を集積コード選択部
9212に与え所定の煩斜コードの集積コード・レジス
タ9213にセットする。この一連の処理が終了すると
、制御部70はパタンメモリ部60に指令を出し退避し
ていたパタンメモリレジス夕から再び走査をつづけさせ
る。第8図において図的表現532に示した下確定特徴
をもつ点505の場合について輪郭形状判定部9220
の動作の効果の一例を第9図に示す。
FIG. 8 shows a diagrammatic representation thereof as an example. White point 50 in the diagram
4, as shown in the graphical representation 531, there are character lines upward to the right, vertical character lines to the right, and further character lines to the left and down to the character frame. A feature code is created indicating its absence. On the other hand, in the white point 505, as shown in the graphical representation 532, there are no character lines in the upper and left directions, there are character lines in the upper right direction in the lower direction, and there are characters in the right direction. Although the line exists, from the black V dot series that first appears on the scanning line extending rightward from the point 505, it becomes unstable, as typified by the disturbing code 508 (FIG. 4) formed at the black dot 503. ”
When it is determined that , a characteristic code indicating that only oblique codes have been extracted is formed. An uncertain accumulation code (* in the graphical representation 532
It is not impossible to take measures such as leaving an uncertain feature code with a mark (mark) as it is and setting the special feature code as dontcare at the stage of identification processing, but as character patterns become more complex like many kanji, This type of indeterminate feature appears more frequently, resulting in many problems such as difficulty in extracting features that reflect the detailed structure of the pattern. This problem,
This problem can be solved by an integrated code determination method that combines a plurality of character line slope information extraction means. This is a block diagram showing an embodiment of the contour shape determining section 922 shown in FIG. 6. The judgment circuit 9221 is an input register 9210
By checking the contents of the character line, the point in time when scanning shifts from the black point to the white point on the character line is detected, and the contents of the integrated code register 9213 at this point are read out. When it is determined that the integrated code stored in the integrated code register is the uncertain integrated code, the character line contour tracing section 9222 outputs a signal to the control section 70 via the connection 701. In response to this, the control unit 70 saves and saves the current pattern memory register in the pattern memory unit 6, and stores the pattern outline in a certain range starting from the previous address (corresponding to the black dot with the unstable accumulation code). is traced and the address of each point is sequentially transferred to the character line contour tracing unit 9222 via the connection line 305. At this time, the forehead slant determination unit 9223 determines the slant of the contour line based on, for example, changes in the vertical and vertical components of the address information of the point on the argument line issued by the character line contour tracing unit 9222, and The determination result is given to the integrated code selection section 9212 and set in the integrated code register 9213 of a predetermined complicated code. When this series of processing is completed, the control unit 70 issues a command to the pattern memory unit 60 to continue scanning from the pattern memory register 60 that has been saved. In the case of the point 505 having the lower definite feature shown in the graphical representation 532 in FIG.
An example of the effect of the operation is shown in FIG.

走査線534を右より左に走査するとき、黒V点503
を含む文字線部分の黒V点列535より抽出される特徴
は図的表現によって示すとおりいずれも不安定懐斜コー
ドをもつ為に、走査が黒点から白点に切りかわる点53
6に達した時点において、該走査方向からの集積コード
を与える集積コード・レジスタ9213上にはビット列
537に示すように上記イニシャル・セット回路921
4によって予め設定された不確定集積コード656が格
納されている。そこで輪郭形状判定部9220が動作し
、例えば文字線輪郭上の黒点列538,539を一定メ
ッシュ数追跡し、この時の水平方向のパタンメモリ上の
距離変化分540,541及び垂直方向のパタンメモI
J上の距離変化分542,543を用いるなどして、文
字線輪郭部538,539の傾斜を検出し、集積コード
レジスタ9213の内容を544に示すとおり右上り頭
斜を表わすコード653(第5図)の形に更新する。こ
れによって、走査線534上に位置する白点505には
当該コードが集積されることになり、最終的には図的表
現533によって示される特徴コードが形成される。第
10図は特徴コード変換部9300の一実施例をブロッ
ク図で示したものである。
When scanning the scanning line 534 from the right to the left, the black V point 503
The features extracted from the black V point sequence 535 of the character line portion containing
6, the initial set circuit 921 is placed on the integrated code register 9213 that provides the integrated code from the scanning direction as shown in the bit string 537.
An uncertain accumulation code 656 preset by 4 is stored. Therefore, the contour shape determination unit 9220 operates, for example, traces the black dot rows 538 and 539 on the character line contour by a certain number of meshes, and at this time, the distance changes 540 and 541 on the horizontal pattern memory and the vertical pattern memo I
The inclinations of the character line contours 538 and 539 are detected by using distance changes 542 and 543 on J, and the contents of the integrated code register 9213 are converted to a code 653 (fifth Update to the form shown in Figure). As a result, the code is accumulated at the white point 505 located on the scanning line 534, and finally the feature code shown by the graphical representation 533 is formed. FIG. 10 shows a block diagram of an embodiment of the feature code converter 9300.

制御部70の指令によりパタンメモリ61上の上下左右
方向の走査に凝ってパタンメモリから読み出された白点
の特徴コードは、結線306を通じて現在レジスタ93
01に格納される。テーブル検索回路9303は、該走
査線上の該白点に至る直前までの走査でパタンメモリに
最終に書き込まれた白点の特徴コードを格納している過
去レジスタ9302の内容と、前記現在レジスタ930
1の内容とから、制御部70よりの指令によって与えら
れる走査方向に対して例えば該走査方向と直交する2方
向等のように予め定められた特定方向に関する鏡斜コー
ドをそれぞれ抽出する。そして両煩斜コ−Nこよって定
まる変換処理後の該白点の額斜コード内容が格納されて
いるコード変換テーブル9304上の該当アドレスを指
定する。変換コード合成回路9305では、コード変換
テーブル9304から読み出された変換済み煩斜コード
を、該白点の特徴コード630(第5図)上のビット位
置646,647,648,649のうちで方向の対応
する1つのビット位置に格納し、残り3つのビット位置
には該特徴コード上の集積コード部640の、それぞれ
方向の対応するビット位置の内容をそのまま格納する。
変換コード合成回路9305の動作によって一部修正さ
れた現在レジス夕9301の内容は、過去レジスタ93
02に格納されるとともに結線307を通じてパタンメ
モリ上の該当するアドレスに書き込まれる。特徴コード
変換部9300の動作による効果は、コード変換テ−ブ
ル9304にあらかじめ格納されるコード変換則の定義
に依存し、ノイズ除去や大域的特徴の抽出などが可能と
なる。
The characteristic code of the white point read from the pattern memory 61 by scanning the pattern memory 61 in the vertical and horizontal directions according to the command from the control unit 70 is stored in the current register 93 through the connection 306.
It is stored in 01. The table search circuit 9303 searches the contents of the past register 9302 storing the feature code of the white point finally written in the pattern memory during scanning up to the point immediately before reaching the white point on the scanning line, and the contents of the current register 9302.
1, mirror skew codes relating to specific directions predetermined, such as two directions perpendicular to the scanning direction given by the command from the control unit 70, are extracted. Then, the corresponding address on the code conversion table 9304 in which the contents of the forehead oblique code of the white spot after the conversion process determined by the double oblique code N is stored is specified. The converted code synthesis circuit 9305 converts the converted complicated code read from the code conversion table 9304 into a direction among the bit positions 646, 647, 648, 649 on the feature code 630 (FIG. 5) of the white point. The contents of the corresponding bit positions in each direction of the integrated code section 640 on the feature code are stored as they are in the remaining three bit positions.
The contents of the current register 9301 partially modified by the operation of the conversion code synthesis circuit 9305 are stored in the past register 93.
02 and written to the corresponding address on the pattern memory through the connection 307. The effect of the operation of the feature code conversion unit 9300 depends on the definition of the code conversion rule stored in advance in the code conversion table 9304, and enables noise removal, global feature extraction, and the like.

第11図はコード変換処理の−実施例を第4図の文字線
パタン500上の白点509の近傍について示したもの
である。
FIG. 11 shows an embodiment of the code conversion process in the vicinity of a white point 509 on the character line pattern 500 of FIG.

注目白点552(第4図図示の白点509に相当すると
する)を含む走査線550上をまず53方向に、次に5
4方向に走査する場合について、簡単の為に各々の白点
551,552,553における右方向51の頃斜コー
ドの推移にのみ注目することにする。図的表現554,
555,556は特徴コード変換部9300が動作する
以前の白点551,552,553における各方向の集
積コードを表わしている。557は上方向53又は下方
向54に走査するとき、任意の白点の右方向51の集積
コードの変換テーブルの一例である。
On the scanning line 550 including the white point 552 of interest (corresponding to the white point 509 shown in FIG. 4), first in the 53 direction, then in the 53 direction.
Regarding the case of scanning in four directions, for the sake of simplicity, we will focus only on the transition of the diagonal code in the right direction 51 at each of the white points 551, 552, and 553. Graphical representation 554,
555 and 556 represent the accumulated codes in each direction at the white points 551, 552, and 553 before the feature code converter 9300 operates. 557 is an example of a conversion table of accumulated codes in the right direction 51 of an arbitrary white point when scanning in the upward direction 53 or downward direction 54.

558に示す記号の列は現在レジスタ上の対応する集積
コードの内容を、また559に示す記号の行は過去レジ
スタ上の対応する集積コードの内容を示し、テーブル上
の記号は該コード変換テーブルによって書きかえられた
集積コードの出力を表わす。
The column of symbols shown at 558 shows the contents of the corresponding accumulated code on the current register, and the row of symbols shown at 559 shows the contents of the corresponding accumulated code on the past register, and the symbols on the table are changed according to the code conversion table. Represents the output of the rewritten integration code.

図的表現563,562,561は、走査線550上を
53方向に白点553,552,551と処理を進めた
時に、図的表現556,555,554に与えられた各
点の集積コードがコード変換された結果を示す。ここで
564,565,566は現在レジスタ、567,56
8,569は過去レジスタの、それぞれ右方向51の集
積コード格納部分の内容である。例えば点552におい
ては、図的表現555に示す特徴コードの右方向51の
集積コードを与える現在レジスタ565の内容と、1つ
前の点553において形成された過去レジスタ569の
内容とから、テーブル557を検索し新しい右方向51
の集積コードを過去レジスタ568に与えることにより
、新たに図的表現562に示す各方向の集積コードを得
る。同様に図的表現571,572,573は走査線5
50上を下方向54に白点551,552,553と処
理を進めた時に図的表現561,562,563に与え
られた各点の集積コードがさらにコード変換された結果
を示す。ここで574,575,576は現在レジスタ
、577,578,579は過去レジスタの、それぞれ
右方向51の集積コード格納部分の内容である。この結
果、557に示すようなコード変換テーブルを与えると
き、特徴コード変換部9300の動作により白点551
,552,553の特徴コード630上の変換済み集積
コード部645(第5図)にはいずれも図的表現571
,572,573に示す安定かつ大域的な図形構造を示
す集積コードが形成される。特徴点抽出部94001こ
おいて、文字パタン上の端点、屈曲点などの位相幾何的
特徴点を抽出する方法は既に幾つか知られており公知の
技術を利用することができる。
The graphical representations 563, 562, and 561 indicate that when processing is performed on the scanning line 550 in the 53 direction with the white points 553, 552, and 551, the integrated code of each point given to the graphical representations 556, 555, and 554 is Shows the result of code conversion. Here, 564, 565, 566 are current registers, 567, 56
8,569 are the contents of the accumulated code storage portions in the right direction 51 of the past registers. For example, at point 552, the table 557 Search for new right direction 51
By applying the accumulated code of 2 to the past register 568, the accumulated codes of each direction shown in the graphical representation 562 are newly obtained. Similarly, the graphical representations 571, 572, 573 are scan line 5
50 and downward 54 to white points 551, 552, and 553, the accumulated codes of each point given to graphical representations 561, 562, and 563 show the results of further code conversion. Here, 574, 575, and 576 are the contents of the current register, and 577, 578, and 579 are the contents of the past register, respectively, in the integrated code storage portion in the right direction 51. As a result, when a code conversion table as shown in 557 is provided, the white point 551
.
, 572, 573, an integrated code exhibiting a stable global graphical structure is formed. In the feature point extraction unit 94001, several methods are already known for extracting topological feature points such as end points and bending points on a character pattern, and known techniques can be used.

とくに前記実施例に示したように、傾斜コード抽出部9
100において、文字線上の任意の黒点から予め定めら
れた各方向別の黒点の連続数を計数し方向別距離を求め
、該方向別距離を比較することにより文字線の方向成分
を第3図図示6201こ示したようにコード化する方法
を用いた場合には、該煩斜コードをそのまま位相幾何的
特徴点の表現として用いることができ、その場合特徴点
抽出部9400は傾斜コード抽出部9100に包含する
ことができる。一例として第4図で、図的表現512な
る特徴コード507をもつ黒点502は唯一右方向にの
み文字線が一定の長さ以上連結していることを表現して
おり、このような特徴コードをもつ黒点を端点として定
義することにするようにする。背景分割線抽出部950
0の一実施例ブロック図を第12図に示す。
In particular, as shown in the above embodiment, the slope code extractor 9
100, the number of consecutive black dots in each predetermined direction from an arbitrary black point on the character line is counted to obtain the distance in each direction, and by comparing the distances in each direction, the directional component of the character line is determined as shown in FIG. 6201 When the encoding method shown above is used, the oblique code can be used as it is as an expression of topological feature points, and in that case, the feature point extracting unit 9400 uses the oblique code extracting unit 9100. can be included. As an example, in FIG. 4, a black dot 502 with a feature code 507, which is a graphical expression 512, represents that character lines are connected for a certain length or more only in the right direction, and such a feature code is Let us define the black point with the end point as the end point. Background dividing line extraction unit 950
FIG. 12 shows a block diagram of one embodiment of 0.

制御部70の指令により上下左右に走査してパタンメモ
リ61から読み出された黒点の特徴コード610及び白
点の特徴コード630は結線310を通して入力レジス
夕9501に格納される。黒点の集積コードの場合は位
相幾何的特徴点抽出部9502が動作し、該特徴コード
610(第3図)上のビット位置620に示す傾斜コー
ドの内容と、制御部70よりの指令によって与えられる
走査方向とから背景分割線の発生の有無を判定し、背景
分割線コードレジスタ9503上に該走査方向によって
定まる背景分割線コードをセットする。第13図図示の
67川ま背景分割線コードの一実施例である。ビット6
71,672,673,674はそれぞれ第3図51,
52,53,54方向の背景分割線が存在する時「1」
となり、いずれの方向の背景分割線も該白点上に存在し
ない時にはビット671,672,673,674はい
ずれも「0」とある。背景分割線発生の有無の判定では
、第3図62川こ示す黒V点の煩斜コードが、該全8ビ
ットのうち1つ又は互いに対となる方向に対応しない2
つのビットのみが「1」となる条件を満足する時に、「
1」となっている該ビットに対応する方向以外の走査方
向に背景分割線を発生せることととし、その為に背景分
割線コード670上の該走査方向に対応するビット位置
に1がセットされる。前記条件を満足しない黒V点の倭
斜コードが検出された時はビット671,672,67
3,674はすべて01こリセットされる。第13図に
は、例として文字パタン500上の黒点584,585
,586の煩斜コードを図的表現684,685.68
6によって示し、各々の黒点を起点として発生する背景
分割線を、黒V点584を起点とする851,852,
853,854及び黒点585を起点とする855,8
56さらに黒点586を起点とする857,858,8
59によって示している。入力レジスタ9501の内容
が白点の特徴コードである場合、背景分割線コード合成
部9504が動作し、入力レジスタ9501より転送さ
れた該特徴コード630(第5図)の中で背景分割線コ
ードを表わすビット位置680の内容を、背景分割線コ
ード・レジスタ9503の内容とビット対応でORをと
ることによって更新し、このデータを結線311を介し
てパタンメモリ61上の対応するアドレスに格納する。
The black dot characteristic code 610 and the white dot characteristic code 630 read out from the pattern memory 61 by scanning vertically and horizontally according to a command from the control unit 70 are stored in the input register 9501 through the connection 310. In the case of a sunspot accumulation code, the topological feature point extraction unit 9502 operates, and the information is given by the content of the slope code shown at the bit position 620 on the feature code 610 (FIG. 3) and the command from the control unit 70. The presence or absence of a background dividing line is determined based on the scanning direction, and a background dividing line code determined by the scanning direction is set in the background dividing line code register 9503. FIG. 13 is an example of the 67 river background dividing line code shown in FIG. bit 6
71, 672, 673, 674 are respectively 51,
"1" when there are background dividing lines in the 52, 53, and 54 directions
Therefore, when no background dividing line in any direction exists on the white point, bits 671, 672, 673, and 674 are all "0". In determining the presence or absence of a background dividing line, it is determined that the slanted code of the black V point shown in FIG.
When the condition that only one bit is "1" is satisfied, "
It is assumed that a background dividing line is generated in a scanning direction other than the direction corresponding to the bit that is "1", and for this purpose, 1 is set in the bit position corresponding to the scanning direction on the background dividing line code 670. Ru. Bits 671, 672, and 67 are set when a Japanese oblique code with a black V point that does not satisfy the above conditions is detected.
3,674 are all reset to 01. FIG. 13 shows black dots 584, 585 on the character pattern 500 as an example.
, 586 is a graphical representation of the 684, 685.68
6, and the background dividing lines generated starting from each black point are 851, 852, 852, and 852 starting from the black V point 584, respectively.
853, 854 and 855, 8 starting from black point 585
56 Furthermore, 857, 858, 8 starting from sunspot 586
59. When the content of the input register 9501 is a white dot feature code, the background dividing line code synthesis unit 9504 operates and combines the background dividing line code into the characteristic code 630 (FIG. 5) transferred from the input register 9501. The contents of the represented bit position 680 are updated by performing a bitwise OR with the contents of the background dividing line code register 9503, and this data is stored in the corresponding address on the pattern memory 61 via the connection 311.

第13図のビット表現681,682,683は白点5
81,582,583における特徴コード630上の背
景分割線コード680の内容である。例えば白点581
では黒点584,685,586からの背景分割線85
4,855,857が交叉していることが表わされてい
る。周辺閉じ状態コード抽出部9600の−実施例フロ
ック図を第14図に示す。
Bit representations 681, 682, 683 in Figure 13 are white dots 5
This is the content of the background dividing line code 680 on the feature code 630 in 81, 582, and 583. For example, white point 581
Then, the background dividing line 85 from the black points 584, 685, 586
It is shown that 4,855,857 intersect. A block diagram of an embodiment of the peripheral closed state code extraction unit 9600 is shown in FIG.

制御部70の指令により上下左右に走査してパタンメモ
リ61から読み出された黒点及び白点の特徴コードは結
線312を通して入力レジスタ961に格納される。黒
V点の特徴コードの場合は、イニシャルセット回路96
03が動作し、周辺閉じ状態コード・レジスタ部960
5上の走査方向に相対する方向に対応するビット位置に
「IJがセットされる。第15図図示695は周辺閉じ
状態コ−ドレジスタ部9605のレジスタの一実施例で
あり、ビット位贋696,697,698,699はそ
れぞれ方向51,52,53,54の周辺閉じ状態を表
わしており、議しジス外ま白点の特徴(第5図)630
上の周辺閉じ状態コード記述部分6901こ対応してい
る。例えば、右方向51に走査する過程で入力レジスタ
9601に黒点の特徴コードが格納された場合には、該
黒点にひきつづいて走査線上から読み出される各々の白
点に対して、左方向52が後述する意味で文字線によっ
て閉じていることを表現する為に、左方向52に対応す
るビット位置697が「1」をセットされる。入力レジ
ス夕9601の内容が白点の特徴コードである場合には
、背景分割線交叉検出部9604において、背景分割線
抽出部95001こよって求められた背景分割線コード
680を解析し、該走査線と直交する方向の背景分割線
との交叉の有無と判定する。該背景分割線が検出されな
かった時には、該特徴コード630上に集積コードを表
わすビット部分640のうちで、該走査線に直交する2
方向に対応する集積コードが、第5図に示した不確定集
積コード656以外の確定集積コードである時、コード
カウンタ部9606において、該方向に対応するコード
・カウンタ9611又は9612の内容を1加算する。
たとえば、右方向51に走査する時、上方向53の確定
集積コードの数はカウン夕9611によって、また下方
向54の確定集積コードの数はカウンタ9612によっ
て計数する。一方、背景分割線が検出された場合には、
周辺閉じ状態コード反転部9607が動作し、該背景分
割線が起点とする方向に対応するコードカウンタ961
1又は9612の内容が予め定められた閥値を越え、か
つ周辺閉じ状態コード・レジスタ部9605のレジスタ
695の内容のうち該走査方向と相対する方向に対応す
るビットが「1」となる場合に限り、該ビット内容をr
o」に変換して周辺閉じ状態コード・レジスタ部960
5内レジスタ695に格納する。コード・カウンタ96
11,9612の内容は、それぞれの対応する方向に起
点をもつ背景分割線が検出されて上記の処理が終了した
時点で、0にリセットされる。周辺閉じ状態コード合成
部9602は、周辺閉じ状態コード反転部9607にお
けるビット反転を必要ならば終了した周辺閉じ状態コー
ド・レジスタ部9605のレジスタ695の内客と、入
力レジスタ9601上の白点の特徴コード630のうち
対応するビット位置690の内容とのORをとって、該
白点の周辺閉じ状態コード記述部分690(第5図)の
内容を更新する。つづいて制御部70はこの内容を結線
313を通してパタンメモリ61上の対応するアドレス
に格納する。第16図に以上説明した周辺閉じ状態コー
ド抽出部9600の動作結果を例によって示す。
The characteristic codes of the black dots and white dots read out from the pattern memory 61 by scanning vertically and horizontally according to a command from the control unit 70 are stored in the input register 961 through the connection line 312. In the case of a feature code with a black V point, the initial set circuit 96
03 is activated, and the peripheral closed status code register section 960
"IJ" is set in the bit position corresponding to the direction opposite to the scanning direction on 5. Reference numeral 695 shown in FIG. 697, 698, and 699 represent peripheral closed states in the directions 51, 52, 53, and 54, respectively, and the characteristics of the white spot outside the frame (Fig. 5) 630
This corresponds to the peripheral closed state code description part 6901 above. For example, if the characteristic code of a black point is stored in the input register 9601 during the process of scanning in the right direction 51, the left direction 52 will be described later for each white point read out from the scanning line following the black point. In order to express that the character line is closed in meaning, the bit position 697 corresponding to the left direction 52 is set to "1". When the content of the input register 9601 is a feature code of a white point, the background dividing line intersection detection unit 9604 analyzes the background dividing line code 680 obtained by the background dividing line extracting unit 95001, and extracts the scanning line It is determined whether or not there is an intersection with the background dividing line in the direction orthogonal to the background dividing line. When the background dividing line is not detected, two of the bit portions 640 representing the integrated code on the feature code 630 are orthogonal to the scan line.
When the accumulation code corresponding to a direction is a definite accumulation code other than the uncertain accumulation code 656 shown in FIG. do.
For example, when scanning in the right direction 51, the number of confirmed accumulated codes in the upward direction 53 is counted by a counter 9611, and the number of confirmed accumulated codes in the downward direction 54 is counted by a counter 9612. On the other hand, if a background dividing line is detected,
The peripheral closed state code reversing unit 9607 operates, and the code counter 961 corresponding to the direction from which the background dividing line starts
When the content of 1 or 9612 exceeds a predetermined threshold value, and the bit corresponding to the direction opposite to the scanning direction among the contents of register 695 of peripheral closed state code register section 9605 becomes "1". If the bit contents are r
peripheral closed status code register section 960
5 internal register 695. code counter 96
The contents of 11 and 9612 are reset to 0 when the background dividing lines having their starting points in the respective corresponding directions are detected and the above processing is completed. The peripheral closed state code synthesis unit 9602 extracts the characteristics of the white dot on the input register 9601 and the inner part of the register 695 of the peripheral closed state code register unit 9605 after bit inversion in the peripheral closed state code inversion unit 9607 has been completed if necessary. By performing an OR with the contents of the corresponding bit position 690 in the code 630, the contents of the surrounding closed state code description portion 690 (FIG. 5) of the white point are updated. Subsequently, the control section 70 stores this content at the corresponding address on the pattern memory 61 through the connection 313. FIG. 16 shows an example of the operation results of the peripheral closed state code extraction unit 9600 described above.

文字パタン500内の同一走査線810上にある白点8
11,812における特徴コード630のビット位置6
401こは、特徴集積部9200‘こおける処理によっ
て図的表現814,815に示す額斜コードを集積され
ている。周辺閉じ状態コード抽出部9600の動作によ
り、該白点811,812の特徴コード630上のビッ
ト位置690には、それぞれ816,817に示す周辺
閉じ状態コードが抽出される。該周辺閉じ状態コードは
、走査線810上を走査するとき背景分割線854上の
白点813の前後において右方向.51に走査する場合
には左方向52の周辺閉じ状態を表わすビット位置69
7(第15図)が、一方左方向52に走査する場合には
右方向51の周辺閉じ状態を表わすビット位置696(
第15図)が、それぞれ「1」から「0」に変化するこ
とにより抽出される。特徴集積部920川こよって求め
られる図的表現814の表わす煩斜コードと周辺閉じ状
態コード抽出部9600によって求められる周辺閉じ状
態コード816とを、それぞれ方向51,52,53,
54に対応づけて組み合わせることにより、図的表現8
25に示す特徴コードが得られる。該図的表現において
「V′」はコード816においてビット821が「0」
となることに対応し、「V」との違いは、白点811よ
り右方向51には左方向52と同様に垂直ストロークが
存在するが、該白点と該ストロークとの間には図中84
0に示す文字線分機部分の影響によって背景分割線85
4を境として右方向51に直交する上方向53、下方向
54のいずれかの方向の文字線状態が異なることを示し
ている。同様に、白点812に関して図的表現815と
コード817とを組み合わせることにより、図的表現8
35に示す特徴コードが得られる。以上説明したように
、本発明によれば、パタン構造を記述する特徴を抽出で
きて識別処理においても特豚昭51−85708号のも
のと整合性をもち、従来法では白点の周囲上下左右方向
の文字線の存在のみしか表現できなかったのに比べて、
各方向に存在する文字線の煩斜情報を集積してコード化
した特徴を各白点より抽出することができる。
White dot 8 on the same scanning line 810 in the character pattern 500
Bit position 6 of feature code 630 at 11,812
401, the forehead oblique codes shown in graphical representations 814 and 815 are accumulated through processing in the feature accumulation unit 9200'. By the operation of the peripheral closed state code extraction unit 9600, peripheral closed state codes shown at 816 and 817 are extracted at bit positions 690 on the feature code 630 of the white points 811 and 812, respectively. The peripheral closed state code is written in the right direction before and after the white point 813 on the background dividing line 854 when scanning the scanning line 810. 51, the bit position 69 represents the peripheral closed state in the left direction 52.
7 (FIG. 15), while when scanning in the left direction 52, bit position 696 (representing the peripheral closed state in the right direction 51)
(Fig. 15) are extracted by changing from "1" to "0", respectively. The oblique code represented by the graphical representation 814 obtained by the feature accumulation section 920 and the surrounding closed state code 816 obtained by the surrounding closed state code extraction section 9600 are moved in directions 51, 52, 53,
By matching and combining 54, graphical representation 8
A characteristic code shown in 25 is obtained. In the graphical representation, "V'" means that bit 821 is "0" in code 816.
Corresponding to this, the difference from "V" is that there is a vertical stroke in the right direction 51 from the white point 811 as well as in the left direction 52, but there is a vertical stroke in the figure between the white point and the stroke. 84
Background dividing line 85 due to the influence of the character line segment machine part shown in 0
It is shown that the character line state is different in either an upward direction 53 or a downward direction 54 that is orthogonal to the right direction 51 with 4 as a boundary. Similarly, by combining the graphical representation 815 and the code 817 regarding the white point 812, the graphical representation 8
A characteristic code shown in 35 is obtained. As explained above, according to the present invention, features that describe the pattern structure can be extracted, and the identification process is consistent with that of Tokubuta No. 51-85708. Compared to the case where only the existence of directional character lines could be expressed,
Features coded by accumulating slant information of character lines existing in each direction can be extracted from each white point.

このため、漢字などのように直線ストロークが多くその
方向性が明確なパタンについてはノイズに対して安定な
特徴を抽出できる利点がある。また、上記の文字線傾斜
情報を集積した特徴と、該特徴を用いて抽出される前記
周辺閉じ状態コードとを組み合わせた特徴は、文字パタ
ンの一部だけに注目した特徴抽出に利用でき、漢字など
のように幾つかの部分パタンから成るアドレスの識別に
有効な特徴として利用される。
For this reason, it has the advantage of being able to extract features that are stable against noise for patterns that have many straight strokes and have clear directions, such as kanji characters. In addition, the feature that combines the feature that accumulates the character line slope information and the surrounding closed state code that is extracted using the feature can be used to extract features that focus on only a part of the character pattern, It is used as an effective feature for identifying addresses that consist of several partial patterns, such as.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は従来のもの則ち侍鹿昭51−85708号に示
す特徴抽出を概念的に説明する説明図、第2図は本発明
の一実施例ブロック図、第3図は第2図額斜コード抽出
部によって得られる黒V点の特徴のビット構成を示す説
明図、第4図は該傾斜コード抽出部によってパタン上の
黒点に形成される額斜コードの例を示す説明図、第5図
は第2図特徴集積部、特徴コード変換部、背景分割線抽
出部、周辺閉じ状態コード抽出部によって得られる白点
の特徴コードのビット構成を示す説明図、第6図は第2
図図示の特徴集積部の一実施例ブロック図、第7図は該
特徴集積部において文字線上の安定傾斜コードを選択し
て集積する処理をパタン領域上で説明する説明図、第8
図は該特徴集積部によって得られる特徴の概念をパタン
領域上で図的に示す説明図、第9図は第5図図示の輪郭
形状判定部の処理をパタン領域について説明する説明図
、第10図は第2図図示の特徴コード変換部の一実施例
ブロック図、第1 1図は該特徴コード変換部の処理を
パタン領域について説明する説明図、第12図は第2図
図示の背景分割線抽出部の一実施例フロック図、第13
図は該背景分割線抽出部の処理をパタン領域について説
明する説明図、第14図は周辺閉じ状態コード抽出部の
一実施例ブロック図、第15図は該周辺閉じ状態コード
抽出部により得られる白点の周辺閉じ状態コードのビッ
ト構成を示す説明図、第16図は該周辺閉じ状態コード
の概念をパタン領域上で図的に示す説明図を示す。 図中、5川ま走査線、60はパタンメモリ部、61はパ
タンメモリ、70は制御部、80は識別部、90は特徴
抽出部、9100は煩斜コード抽出部、920川ま特徴
集積部、930川ま特徴コード変換部、9400は特徴
点抽出部、9500は背景分割線抽出部、9600は周
辺閉じ状態コード抽出部を表わす。 汐3図 ゲワ図 ゲー図 才2図 才4図 グ5図 汐q図 劣5図 沙8図 才′2図 オー5図 矛′○図 オー1図 オー3図 オー4図 オf6図
Figure 1 is an explanatory diagram conceptually explaining the conventional feature extraction shown in Samurai Sho 51-85708, Figure 2 is a block diagram of an embodiment of the present invention, and Figure 3 is the second figure. FIG. 4 is an explanatory diagram showing the bit structure of the feature of the black V point obtained by the diagonal code extraction section; FIG. The figure is an explanatory diagram showing the bit configuration of the feature code of a white point obtained by the feature accumulation section, feature code conversion section, background dividing line extraction section, and peripheral closed state code extraction section in FIG.
FIG. 7 is a block diagram of an embodiment of the feature accumulation unit shown in the figure; FIG.
9 is an explanatory diagram diagrammatically showing the concept of features obtained by the feature accumulating unit on a pattern area, FIG. 9 is an explanatory diagram illustrating the processing of the contour shape determining unit shown in FIG. The figure is a block diagram of an embodiment of the feature code conversion section shown in FIG. 2, FIG. 13th example of a block diagram of a line extraction section
The figure is an explanatory diagram explaining the processing of the background dividing line extraction section for a pattern region, FIG. 14 is a block diagram of an embodiment of the surrounding closed state code extraction section, and FIG. 15 is a diagram showing the processing of the surrounding closed state code extraction section. FIG. 16 is an explanatory diagram showing the bit configuration of a surrounding closed state code of a white point. FIG. 16 is an explanatory diagram that graphically shows the concept of the surrounding closed state code on a pattern area. In the figure, there are 5 scanning lines, 60 a pattern memory section, 61 a pattern memory, 70 a control section, 80 an identification section, 90 a feature extraction section, 9100 a complex code extraction section, and 920 a feature accumulation section. , 930 represents a feature code conversion unit, 9400 represents a feature point extraction unit, 9500 represents a background dividing line extraction unit, and 9600 represents a peripheral closed state code extraction unit. Figure 3

Claims (1)

【特許請求の範囲】 1 黒白の2値パタン1個を含むパタン領域の各点から
予め定めた方向を見た時に見つかる位相幾何的特徴を方
向情報を保存して各点に集積し、この結果得られる高次
特徴を用いて識別する文字読取方式において、上記位相
幾何的特徴として文字線の傾きを検出して当該傾きをパ
タン面上の対応する黒点に傾斜コードとしてコード化し
て形成する傾斜コード抽出手段と、白点に対して予め定
められた方向に存在する文字線上の当該傾斜コードを該
方向に対応づけて集積させて高次特徴を形成させる特徴
集積手段と、をもつ特徴抽出部をそなえ、該特徴抽出部
によつて、白点の周辺の文字線の存在及び傾斜情報を表
わす特徴コードを抽出することを特徴とする特徴抽出処
理方式。 2 ドキユメント上の文字を走査して走査信号を黒白2
値化して出力する走査部と、1個の2値パタンを含むパ
タン領域をその一部として格納する1メツシユあたり複
数ビツトのM×Nメツシユのパタンメモリ部と、当該パ
タンメモリ上に格納されているパタンを処理して位相幾
何的特徴を表わす特徴コードを該2値パタン上の各点に
対応する該パタンメモリ上に形成する特徴抽出部と、該
特徴コードを処理して入力文字がどのカテゴリに属する
かを決定する識別部と、それぞれの回路に必要なタイミ
ング信号を発生する制御部とから構成される文字読取装
置において、上記特徴抽出部は、黒点に傾斜コードを形
成する傾斜コード抽出部、パタン面上を予め定められた
方向に走査し文字線上の前記傾斜コードを選択し該傾斜
コードを該方向に対応づけてパタンメモリ上の白点の集
積させて新たな高次特徴を形成させる特徴集積部とから
構成され、上記識別部は白点上の当該高次特徴をも用い
て総合判断しドキユメント上の文字を識別することを特
徴とする特許請求の範囲第1項記載の特徴抽出処理方式
。 3 文字線の傾きを表現する前記傾斜コードの抽出は、
予め定められた幾つかの方向について文字線上の黒点の
連りの長さを求め、該方向別にその大きさを比較するこ
とにより文字線の傾きを検出する手段と、文字パタンの
文字線部と背景白部との輪郭線を追跡することによりそ
の傾きを検出する手段とを設け、前記2つの手段により
検出される文字線の傾き情報とを組み合せて判定し傾斜
コードを決定することを特徴とする特許請求の範囲第1
項記載の特徴抽出処理方式。 4 白点上の前記高次特徴を該白点から予め定められた
方向にある近傍白点上の高次特徴と比較して、必要なら
ば該高次特徴の一部に集積された傾斜コードを予め定め
られた規則に従つてコード変換を行なう手段を設け、当
該コード変換された新たなる特徴を高次特徴として抽出
することを特徴とする特許請求の範囲第1項記載の特徴
抽出処理方式。 5 文字線上の位相幾何的特徴として端点・分岐点・屈
曲点等の特徴点を抽出し対応する黒点にコード化して形
成する手段と、黒点上の当該特徴点を起点として該特徴
点の種別に応じて予め定められた方向に背景白部を分割
する背景分割線を伸ばし該背景分割線上に位置する白点
に該背景分割線の方向を表わすコードを形成する手段と
、文字線或いは該背景分割線によつて区切られた背景白
部の領域を1つの単位領域とみなして、当該領域から予
め定められた方向に存在する単位領域内白点の高次特徴
の種別、及び当該方向に存在する背景分割線の方向と個
数とから求められ該単位領域周囲の閉じ状態を表わす特
徴を当該領域内の白点にコードとして形成する手段とを
設け、白点上に周囲の文字線の不連続性を反映した位相
幾何的特徴を表わす特徴コードを形成することを特徴と
する特許請求の範囲第1項記載の特徴抽出処理方式。
[Scope of Claims] 1. Topological features found when looking in a predetermined direction from each point of a pattern area including one black-and-white binary pattern are accumulated at each point while storing direction information, and the result is In a character reading method that identifies characters using the obtained higher-order features, a slope code is formed by detecting the slope of a character line as the topological feature and encoding the slope as a slope code on a corresponding black point on a pattern surface. A feature extraction unit comprising an extraction means and a feature accumulation means for forming a higher-order feature by accumulating the slope code on the character line existing in a predetermined direction with respect to the white point in association with the direction. A feature extraction processing method characterized in that the feature extraction unit extracts a feature code representing the presence and slope information of a character line around a white point. 2 Scan the characters on the document and convert the scanning signal into black and white2
a scanning unit that converts and outputs a value; a pattern memory unit that stores a pattern area including one binary pattern as part of an M×N mesh with a plurality of bits per mesh; a feature extraction unit that processes a pattern to form a feature code representing a topological feature on the pattern memory corresponding to each point on the binary pattern; In the character reading device, the character reading device is composed of an identification section that determines whether the black dot belongs to a black dot, and a control section that generates a timing signal necessary for each circuit. , scan the pattern surface in a predetermined direction, select the slope code on the character line, associate the slope code with the direction, accumulate white points on the pattern memory, and form a new higher-order feature. Feature extraction according to claim 1, characterized in that the identification section also uses the higher-order features on the white points to make a comprehensive judgment and identify the characters on the document. Processing method. 3 Extraction of the slope code that expresses the slope of the character line is as follows:
Means for detecting the inclination of a character line by determining the length of a series of black dots on a character line in several predetermined directions and comparing the sizes in each direction; and a means for detecting the inclination of the contour line with respect to the white background part by tracing the contour line, and the inclination code is determined by combining and determining the inclination information of the character line detected by the two means. Claim 1
Feature extraction processing method described in section. 4. Comparing the higher-order features on the white point with higher-order features on neighboring white points in a predetermined direction from the white point, and if necessary, applying a slope code integrated to a part of the higher-order feature. A feature extraction processing method according to claim 1, characterized in that a means for code conversion is provided according to a predetermined rule, and a new feature resulting from the code conversion is extracted as a higher-order feature. . 5. Means for extracting feature points such as end points, branch points, bending points, etc. as topological features on a character line and encoding them into corresponding black points, and a method for determining the type of feature points using the feature points on the black points as a starting point. means for extending a background dividing line that divides a white part of the background in a predetermined direction according to the background dividing line and forming a code representing the direction of the background dividing line at a white point located on the background dividing line; and a character line or the background dividing line. The area of the white part of the background separated by a line is regarded as one unit area, and the type of higher-order feature of the white point within the unit area that exists in a predetermined direction from the area, and the type of higher-order feature that exists in the direction A means for forming a code at a white point in the area, which is determined from the direction and number of the background dividing lines and representing a closed state around the unit area, is provided, and the discontinuity of the surrounding character lines is detected on the white point. 2. The feature extraction processing method according to claim 1, wherein a feature code representing a topological feature reflecting the above is formed.
JP55012208A 1980-02-04 1980-02-04 Feature extraction processing method Expired JPS6037954B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP55012208A JPS6037954B2 (en) 1980-02-04 1980-02-04 Feature extraction processing method

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP55012208A JPS6037954B2 (en) 1980-02-04 1980-02-04 Feature extraction processing method

Publications (2)

Publication Number Publication Date
JPS56110188A JPS56110188A (en) 1981-09-01
JPS6037954B2 true JPS6037954B2 (en) 1985-08-29

Family

ID=11798961

Family Applications (1)

Application Number Title Priority Date Filing Date
JP55012208A Expired JPS6037954B2 (en) 1980-02-04 1980-02-04 Feature extraction processing method

Country Status (1)

Country Link
JP (1) JPS6037954B2 (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP3018891U (en) * 1995-05-31 1995-11-28 有限会社オフィスキャット Portable electrical equipment holder

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH0643979U (en) * 1992-11-17 1994-06-10 エスエムケイ株式会社 Rotary encoder

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP3018891U (en) * 1995-05-31 1995-11-28 有限会社オフィスキャット Portable electrical equipment holder

Also Published As

Publication number Publication date
JPS56110188A (en) 1981-09-01

Similar Documents

Publication Publication Date Title
JP3302147B2 (en) Document image processing method
US4408342A (en) Method for recognizing a machine encoded character
KR19990062829A (en) String Extractor and Pattern Extractor
JP6220770B2 (en) Form definition device, form definition method, and form definition program
CN115439866B (en) A method, apparatus, and storage medium for table structure recognition of three-line tables.
US6947596B2 (en) Character recognition method, program and recording medium
JPH11161736A (en) Character recognition method
CN119445600A (en) Method, device, computer equipment and readable storage medium for identifying tables in images
JP4878057B2 (en) Character recognition method, program, and recording medium
JP4282467B2 (en) Image area separation method
KR100332752B1 (en) Character recognition method
JPS6047636B2 (en) Feature extraction processing method
JPH0535250A (en) How to edit small size character bitmaps using concatenation (run)
JPH0253829B2 (en)
JPH0113583B2 (en)
JP2004334913A (en) Form recognition device and form recognition method
JPS5811662B2 (en) Character/figure recognition method
JP2993533B2 (en) Information processing device and character recognition device
JPH08212292A (en) Frame recognition device
JPH09114925A (en) Optical character reader
JPS596419B2 (en) Character extraction method
JP3037504B2 (en) Image processing method and apparatus
JP2784004B2 (en) Character recognition device
JPH0252313B2 (en)
JPS59128681A (en) Character reader