JPH03670B2 - - Google Patents
Info
- Publication number
- JPH03670B2 JPH03670B2 JP58109187A JP10918783A JPH03670B2 JP H03670 B2 JPH03670 B2 JP H03670B2 JP 58109187 A JP58109187 A JP 58109187A JP 10918783 A JP10918783 A JP 10918783A JP H03670 B2 JPH03670 B2 JP H03670B2
- Authority
- JP
- Japan
- Prior art keywords
- pattern
- character
- sub
- slope
- stroke
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired
Links
- 239000011159 matrix material Substances 0.000 claims description 16
- 230000005484 gravity Effects 0.000 claims description 9
- 238000000034 method Methods 0.000 claims description 9
- 238000000605 extraction Methods 0.000 description 15
- 238000010586 diagram Methods 0.000 description 6
- 239000000284 extract Substances 0.000 description 5
- 230000003287 optical effect Effects 0.000 description 3
- 238000007796 conventional method Methods 0.000 description 2
- 238000001514 detection method Methods 0.000 description 2
- 238000006243 chemical reaction Methods 0.000 description 1
- 230000007547 defect Effects 0.000 description 1
- 230000000694 effects Effects 0.000 description 1
Landscapes
- Character Discrimination (AREA)
Description
【発明の詳細な説明】
(技術分野)
本発明は安定で精度のよい文字認識方式に関す
るものである。DETAILED DESCRIPTION OF THE INVENTION (Technical Field) The present invention relates to a stable and accurate character recognition method.
(背景技術)
従来文字認識装置においては、第1図の例の様
な手書文字の筆者の違いによる文字線の傾斜、又
印字文字の印字の傾斜に起因する抽出した特徴の
ばらつきを吸収するために辞書マスクの複数化の
手段により前記特徴のばらつきを吸収していた。
しかしながらこの手段は識別を行なう際の抽出し
た特徴と辞書との照合の時間が辞書マスクの数に
比例して増大し、装置の処理速度の低下を招いて
いた。この点を除去する為に水平サブパターン及
び垂直サブパターンより傾斜を抽出しその傾斜に
従つて分割領域を定める方法が提案されている。
また、一般に文字パターンを垂直軸又は水平軸に
投影して求めた黒点数の分布を用いて分割点座標
は決定されるので第5図の例では、分割点座標は
Y1,Y2に示すように2本のストロークの中央に
位置する。ここで線aは従来の分割境界を示す。
この様に従来の技術では文字パターンの左辺及び
下辺を基準に前記傾斜に従う分割境界を定めてい
たので第5図のようにストロークが切断され、そ
の特徴が不安定になるという欠点があつた。(Background Art) Conventional character recognition devices absorb variations in extracted features caused by the inclination of character lines due to differences in handwritten characters by different writers, as shown in the example in Figure 1, and the inclination of printed characters. Therefore, variations in the characteristics have been absorbed by creating a plurality of dictionary masks.
However, with this method, the time required to compare the extracted features with the dictionary during identification increases in proportion to the number of dictionary masks, resulting in a reduction in the processing speed of the device. In order to eliminate this point, a method has been proposed in which the slope is extracted from the horizontal sub-pattern and the vertical sub-pattern and divided regions are determined according to the slope.
In addition, the dividing point coordinates are generally determined using the distribution of the number of black dots obtained by projecting the character pattern onto the vertical or horizontal axis, so in the example shown in Figure 5, the dividing point coordinates are
It is located in the center of the two strokes as shown in Y 1 and Y 2 . Here, line a indicates a conventional division boundary.
As described above, in the conventional technique, the dividing boundaries were determined according to the above-mentioned inclination based on the left side and the bottom side of the character pattern, which had the disadvantage that the stroke was cut off as shown in FIG. 5, and its characteristics became unstable.
(発明の課題)
本発明はこのように欠点を除去するために各サ
ブパターンの傾斜を抽出し、さらに特徴マトリク
スを作成する段階において文字パターンの重心を
通る軸上の分割座標を基準に前記傾斜に従う分割
境界を定めることにより、ストロークの切断によ
る特徴の不安定性を除去したものであり、その目
的は安定で精度の良い文字認識方式を提供するこ
とにある。(Problem to be solved by the invention) The present invention extracts the slope of each sub-pattern in order to eliminate defects as described above, and furthermore, in the step of creating a feature matrix, the slope is extracted based on the dividing coordinates on the axis passing through the center of gravity of the character pattern. By defining division boundaries according to the following, instability of features due to stroke cutting is removed, and the purpose is to provide a stable and highly accurate character recognition method.
(発明の構成および作用)
第2図は、本発明の文字認識方式における一実
施例の構成図を示す。図において、文字の光信号
は光信号入力1より光電変換部2において2値の
量子化されたデイジタル電気信号に変換され、パ
ターンレジスタ3に格納される。それと同時に線
幅計算部4において入力パターンの線幅Wが計算
される。サブパターン抽出部5はパターンレジス
タ3について垂直スキヤンを全面に行なつて黒点
(文字線部を黒点とする)の連続の長さと線幅計
算部4において計算された線幅との関係より垂直
サブパターン(VSP)を抽出しサブパターンレ
ジスタに格納する。同様に、水平スキヤンにより
水平サブパターン(HSP)を、右斜め45゜スキヤ
ンにより右斜めにサブパターン(RSP)を、左
斜め45゜スキヤンによる左斜めサブパターン
(LSP)を抽出し各サブパターンレジスタに格納
する。第3図は原パターンと各サブパターンの例
でaは原パターン、bは垂直サブパターン
(VSP)、cは水平サブパターン(HSP)、dは右
斜めサブパターン(RSP)、eは左斜めサブパタ
ーン(LSP)である。(Structure and operation of the invention) FIG. 2 shows a block diagram of an embodiment of the character recognition system of the invention. In the figure, an optical signal of a character is converted into a binary quantized digital electrical signal from an optical signal input 1 by a photoelectric converter 2, and is stored in a pattern register 3. At the same time, the line width W of the input pattern is calculated in the line width calculating section 4. The sub-pattern extraction unit 5 performs a vertical scan over the entire surface of the pattern register 3, and extracts vertical sub-patterns based on the relationship between the continuous length of black dots (character line portions are black dots) and the line width calculated by the line width calculation unit 4. Extract the pattern (VSP) and store it in the subpattern register. Similarly, the horizontal subpattern (HSP) is extracted by horizontal scan, the right diagonal subpattern (RSP) is extracted by right diagonal 45° scan, and the left diagonal subpattern (LSP) is extracted by left diagonal 45° scan. Store in. Figure 3 shows an example of the original pattern and each sub-pattern, where a is the original pattern, b is the vertical sub-pattern (VSP), c is the horizontal sub-pattern (HSP), d is the right diagonal sub-pattern (RSP), and e is the left diagonal. It is a subpattern (LSP).
ストローク抽出部6は各サブパターンレジスタ
における水平又は垂直スキヤンを全面行ない、白
点から黒点、黒点から白点への変化点を検出し、
1列(又は行)前のスキヤンにおける変化点個数
と変化点座標と現列(又は行)の変化点個数と変
化点座標の関係よりストロークを抽出し、抽出し
た各サブパターンレジスタ内のストロークの両端
点のパターンレジスタ3で定義される2次元座標
系における座標(パターンレジスタの左下を原点
とする)を傾斜抽出部7へ送出する。傾斜抽出部
7はストローク抽出部6において抽出した各サブ
パターンレジスタ内のストロークの両端点座標を
参照し、水平サブパターンHSP及び垂直サブパ
ターンVSPの平均傾斜を計算する。即ち水平サ
ブパターンHSPより抽出したストロークの両端
点座標を(HXSn,HYSn)、(HXEn,HYEn)、
但しn=1,…,P,Pはストローク数として(1)
式により傾斜QHを計算する。(但しHXEp>
HXSp)
(1)式中のHLGpは当該ストロークの長さを表わ
し、(2)式の近似式により求める。 The stroke extraction unit 6 performs a horizontal or vertical scan over the entire surface of each sub-pattern register, detects points of change from a white point to a black point, and from a black point to a white point,
Strokes are extracted from the relationship between the number of change points and change point coordinates in the scan of the previous column (or row) and the number of change points and change point coordinates in the current column (or row), and the strokes in each extracted subpattern register are The coordinates of both end points in the two-dimensional coordinate system defined by the pattern register 3 (with the lower left of the pattern register as the origin) are sent to the slope extractor 7. The slope extraction unit 7 refers to the coordinates of both end points of the stroke in each subpattern register extracted by the stroke extraction unit 6, and calculates the average slope of the horizontal subpattern HSP and the vertical subpattern VSP. In other words, the coordinates of both end points of the stroke extracted from the horizontal subpattern HSP are (HXSn, HYSn), (HXEn, HYEn),
However, n=1,..., P, P is the number of strokes (1)
Calculate the slope Q H by the formula. (However, HXEp>
HXSp) HLGp in equation (1) represents the length of the stroke, and is determined by the approximate equation of equation (2).
HLGp=MAX{|HXEp−HXSp|,|HYEp
−HYSp|}+
MIN{|HXEp−HXSp|,|HYEp−HYSp|}/2
(2)
(2)式は2点間の距離を、2点間の水平及び垂直
座標差のうちで小さい方の1/2と他の一方との和
とする近似式である。 HLGp=MAX|HXEp−HXSp|、|HYEp
−HYSp|}+
MIN{|HXEp−HXSp|, |HYEp−HYSp|}/2 (2) Equation (2) calculates the distance between two points as 1/2 of the smaller of the horizontal and vertical coordinate differences between the two points. This is an approximate expression that is the sum of the other one.
同様にθVを(3)式により計算する。但しVZEq>
VZSqとする。 Similarly, θ V is calculated using equation (3). However, VZEq>
Let it be VZSq.
なお、上記式中Qは垂直サブパターンより抽出
したストローク数である。また、ストローク数が
0とのきは傾斜も0とする。またストロークの長
さVLGqは(2)式と同様な計算式により算出する。 Note that Q in the above formula is the number of strokes extracted from the vertical sub-pattern. Furthermore, when the number of strokes is 0, the slope is also 0. Further, the stroke length VLGq is calculated using a formula similar to formula (2).
傾斜抽出部7は上記式(1)〜(3)より計算した各サ
ブパターンの傾斜を特徴マトリクス抽出部10へ
送出する。文字枠検出部8はパターンレジスタ3
内の文字パターンに外接する文字枠を検出し、そ
の結果を文字枠分割決定部9へ送る。 The slope extraction unit 7 sends the slope of each sub-pattern calculated from the above equations (1) to (3) to the feature matrix extraction unit 10. The character frame detection section 8 is the pattern register 3
A character frame circumscribing the character pattern within is detected, and the result is sent to the character frame division determining section 9.
文字枠分割決定部9は検出された文字枠内をM
×Nの領域(M,Nは整定数、本実施例ではM=
N=5)に分割するためのX軸、Y軸上の分割点
座標を決定する。ここでX軸は文字枠の水平方向
を、Y軸は垂直方向をそれぞれ示す。 The character frame division determining unit 9 divides the inside of the detected character frame into M
×N area (M, N are integer constants, in this example, M=
Coordinates of dividing points on the X-axis and Y-axis for dividing into N=5) are determined. Here, the X axis indicates the horizontal direction of the character frame, and the Y axis indicates the vertical direction.
分割点の決定は次のように行う。まず文字パタ
ーンをX軸及びY軸上に投影してそれぞれ黒ビツ
ト数の分布を求めそれら分布の一次モーメント値
を総黒ビツト数で除算することにより該分布の重
心を求める。さらに求めた重心で分布を2分しそ
れぞれの分布の重心を計算するという手順を繰り
返すことにより、複数の重心座標を得る。この重
心座標系列から目的とする分割数に必要な分割座
標を選択し、これと最初に求めた文字パターンの
重心座標XC,YCを特徴マトリクス抽出部10へ
出力する。 The division points are determined as follows. First, the character pattern is projected onto the X-axis and Y-axis to obtain the distribution of the number of black bits, respectively, and the center of gravity of the distribution is determined by dividing the first moment value of the distribution by the total number of black bits. Furthermore, by repeating the procedure of dividing the distribution into two at the obtained center of gravity and calculating the center of gravity of each distribution, a plurality of coordinates of the center of gravity are obtained. The division coordinates necessary for the desired number of divisions are selected from this barycenter coordinate series, and this and the barycenter coordinates XC and YC of the first obtained character pattern are output to the feature matrix extraction section 10.
特徴マトリクス抽出部10は文字枠分割決定部
9により決定された分割点座標X1,X2…,XM-1
及びY1,Y2,…,YN-1及び文字パターンの重心
座標XC,YCと傾斜抽出部7より得られた傾斜量
θH,θVによりVSP,HSP,RSP,LSPの各サブ
パターンレジスタ上の文字枠領域をM×N個の部
分領域に分割する。以下分割領域の決定方法を説
明する。 The feature matrix extraction section 10 extracts the division point coordinates X 1 , X 2 . . . , X M-1 determined by the character frame division determination section 9.
and each sub-pattern of VSP, HSP, RSP, and LSP using Y 1 , Y 2 , ..., Y N-1 , the barycentric coordinates XC, YC of the character pattern, and the tilt amounts θ H , θ V obtained from the tilt extraction unit 7. Divide the character frame area on the register into M×N partial areas. The method for determining divided regions will be explained below.
まず。分割点座標X1,X2…XM-1及びY1…YN-1
及びXC,YCと傾斜量θH,θVとにより文字枠の左
下を原点とする2次元平面上に(4)式で示される
(M+N−2)本の直線を定義する。(但しY1<
Y2…<YN-1,X1<X2…<XM-1)
y=θH(x−XC)+Y1
〓
y=θH(x−XC)+YN-1
x=θV(y−YC)+X1
〓
x=θV(y−YC)+XM-1 (4)
文字枠内の任意の点A(xe,ye)が(5)式の不等
式を満たすときA点は部分領域(m,n)に含ま
れる。 first. Dividing point coordinates X 1 , X 2 ...X M-1 and Y 1 ...Y N-1
Then, (M+N-2) straight lines shown by equation (4) are defined on a two-dimensional plane with the origin at the lower left of the character frame by XC, YC and the inclinations θ H and θ V. (However, Y 1 <
Y 2 ... < Y N - 1 , X 1 < (y-YC) + X 1 〓 x=θ V (y-YC) + Included in partial area (m, n).
{(xe−XC)・θH+Yn- 1}<ye≦{(xe−XC)・
θH+Yn}
かつ
{(ye−YC)・θV+Xm- 1}<xe≦{(ye−YC)・
θV+Xm} (5)
但しY0=X0=0,YN=XM=∞
すなわち(4)式の直線を境界として文字枠内を複
数の部分領域に分割する。第4図はM=N=5の
場合の分割例である。特徴マトリクス抽出部10
は上記の方法で文字枠内を部分領域に分割した各
サブパターンレジスタの各領域の黒点数Bijを計
数し、線幅計算部4で計算した線幅Wを用いて式
(6)により文字線長をあらわす特徴を計算しM×N
×4次元の特徴マトリクスを作成する。 {(xe−XC)・θ H +Yn - 1 }<ye≦{(xe−XC)・
θ H +Yn} and {(ye−YC)・θ V +Xm - 1 }<xe≦{(ye−YC)・
θ V +Xm} (5) However, Y 0 =X 0 =0, Y N =X M =∞ In other words, the inside of the character frame is divided into a plurality of partial areas using the straight line in equation (4) as a boundary. FIG. 4 shows an example of division in the case of M=N=5. Feature matrix extraction unit 10
is calculated by counting the number of black dots Bij in each area of each sub-pattern register in which the character frame is divided into partial areas using the above method, and using the line width W calculated by the line width calculation unit 4 to form an equation.
Calculate the feature representing the character line length using (6) and calculate M×N
×Create a four-dimensional feature matrix.
Lij=Bij/W (6)
さらにVSP特徴マトリクスは文字枠のY軸方
向の長さΔYで、HSP特徴マトリクスはX軸方向
の長さΔXで、RSP及びLSP特徴マトリクスは
(ΔX+ΔY)/2でそれぞれ正規化を行ない最終
的にM×N×4次元の特徴マトリクスを作成す
る。識別部12は特徴マトリクス抽出部11で抽
出された特徴マトリクス(F)とあらかじめ用意
した辞書マスク(f)との間に式(7)で定義される
距離(D)を適用しDが最小の値となる辞書マス
クのカテゴリ名を文字名出力12へ出力するもの
である。 Lij=Bij/W (6) Furthermore, the length of the VSP feature matrix in the Y-axis direction of the character frame is ΔY, the HSP feature matrix is the length in the X-axis direction is ΔX, and the RSP and LSP feature matrices are (ΔX+ΔY)/2. Each is normalized to finally create an M×N×4-dimensional feature matrix. The identification unit 12 applies the distance (D) defined by equation (7) between the feature matrix (F) extracted by the feature matrix extraction unit 11 and the dictionary mask (f) prepared in advance, and calculates the distance (D) defined by equation (7), and determines whether D is the minimum. The category name of the dictionary mask serving as the value is output to the character name output 12.
D=√(−)2 (7)
(発明の効果)
以上説明したように本実施例では水平サブパタ
ーン及び垂直サブパターンのストロークの傾斜を
抽出し、さらに文字の重心を通る軸上に分割座標
を定め、その座標を基準に前記ストロークの傾斜
に従つて分割領域を決定しているので、ストロー
クを切断することなく安定な特徴を得ることがで
きるという利点がある。 D=√(-) 2 (7) (Effects of the Invention) As explained above, in this embodiment, the slope of the stroke of the horizontal sub-pattern and the vertical sub-pattern is extracted, and the dividing coordinates are further set on the axis passing through the center of gravity of the character. Since the dividing area is determined according to the inclination of the stroke using the coordinates as a reference, there is an advantage that stable features can be obtained without cutting the stroke.
換言すれば、文字の重心に着目して傾斜にそつ
て文字枠を分割しているので手書文字の変形に追
従した特徴を抽出できる利点がある。 In other words, since the character frame is divided along the inclination by focusing on the center of gravity of the character, there is an advantage that features that follow the deformation of handwritten characters can be extracted.
本発明は、文字パターン内の各方向のストロー
クの傾斜を抽出し、さらに特徴マトリクスを作成
する段階における文字枠内の分割領域を、文字パ
ターンの重心を通る軸上に定められた分割座標を
通り前記ストロークの傾斜に従う分割境界によつ
て決定しているので、文字線が切断されることが
少なく安定な特徴が得られるという利点があり安
定で認識精度の良い文字認識装置に利用すること
ができる。 The present invention extracts the inclination of strokes in each direction within a character pattern, and furthermore, in the step of creating a feature matrix, dividing regions within a character frame pass through division coordinates determined on an axis passing through the center of gravity of the character pattern. Since the dividing boundary is determined based on the inclination of the stroke, character lines are less likely to be cut and stable features can be obtained, which can be used in a stable character recognition device with high recognition accuracy. .
第1図は手書文字例を示す図、第2図は本発明
の文字認識装置における一実施例を示す構造図、
第3図は原パターンと各サブパターンの例を示す
図、第4図は領域分割の例を示す図、第5図は従
来の技術の欠点を説明する図である。
1:光信号入力、2:光電変換部、3:パター
ンレジスタ、4:線幅計算部、5:サブパターン
抽出部、6:ストローク抽出部、7:傾斜抽出
部、8:文字枠検出部、9:文字枠分割決定部、
10:特徴マトリクス抽出部、11:識別部、1
2:文字名出力。
FIG. 1 is a diagram showing an example of handwritten characters, FIG. 2 is a structural diagram showing an embodiment of the character recognition device of the present invention,
FIG. 3 is a diagram showing an example of the original pattern and each sub-pattern, FIG. 4 is a diagram showing an example of area division, and FIG. 5 is a diagram explaining the drawbacks of the conventional technique. 1: Optical signal input, 2: Photoelectric conversion section, 3: Pattern register, 4: Line width calculation section, 5: Sub pattern extraction section, 6: Stroke extraction section, 7: Slope extraction section, 8: Character frame detection section, 9: Character frame division determination section,
10: Feature matrix extraction section, 11: Identification section, 1
2: Character name output.
Claims (1)
クをあらわすサブパターンを抽出し、各サプパタ
ーンより得られるストロークの傾斜を長さを重み
として加重平均したものを該サブパターンの傾斜
として抽出し、前記各サブパターンを文字枠内に
ついて前記傾斜に従つて分割することによつて得
られる部分領域毎に、該部分領域内の黒ビツト数
を文字線幅とストローク方向に対応した文字枠の
大きさで正規化して得られる量を特徴要素として
特徴マトリクスを作成し、標準文字マスクが当該
特徴マトリクスと同形式で記述されている辞書を
参照して文字パターンの認識を行う文字認識方式
において、 特徴マトリクスを作成する際に、文字パターン
の重心を通る軸上に分割座標を定め、これを通り
前記サブパターンの傾斜に従う分割境界により、
部分領域に分割することを特徴とする文字認識方
式。[Claims] 1. A sub-pattern representing a stroke in a desired direction is extracted from a character/figure pattern, and the slope of the sub-pattern is obtained by weighting the slope of the stroke obtained from each sub-pattern using the length as a weight. For each partial area obtained by extracting and dividing each sub-pattern within the character frame according to the slope, the number of black bits in the partial area is divided into a character frame corresponding to the character line width and stroke direction. In a character recognition method, a feature matrix is created using the quantities obtained by normalizing with the size of the feature element, and character patterns are recognized by referring to a dictionary in which standard character masks are described in the same format as the feature matrix. , When creating a feature matrix, the division coordinates are determined on the axis passing through the center of gravity of the character pattern, and by the division boundary that passes through this and follows the slope of the sub-pattern,
A character recognition method characterized by dividing into partial areas.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP58109187A JPS603070A (en) | 1983-06-20 | 1983-06-20 | Character recognition system |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP58109187A JPS603070A (en) | 1983-06-20 | 1983-06-20 | Character recognition system |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPS603070A JPS603070A (en) | 1985-01-09 |
| JPH03670B2 true JPH03670B2 (en) | 1991-01-08 |
Family
ID=14503839
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP58109187A Granted JPS603070A (en) | 1983-06-20 | 1983-06-20 | Character recognition system |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPS603070A (en) |
-
1983
- 1983-06-20 JP JP58109187A patent/JPS603070A/en active Granted
Also Published As
| Publication number | Publication date |
|---|---|
| JPS603070A (en) | 1985-01-09 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JPH03670B2 (en) | ||
| JPH0545992B2 (en) | ||
| JPS6214277A (en) | Image processing method | |
| JP3095470B2 (en) | Character recognition device | |
| JP3936039B2 (en) | Screened area extraction device | |
| JPH0420227B2 (en) | ||
| JPS5837780A (en) | Character recognizing method | |
| JPH0127474B2 (en) | ||
| JPS60201485A (en) | Handwritten character recognizing device | |
| JPH06223226A (en) | Character recognizing device | |
| JPH0547871B2 (en) | ||
| JPH0656625B2 (en) | Feature extraction method | |
| JPS6363952B2 (en) | ||
| JP2918363B2 (en) | Character classification method and character recognition device | |
| JPS6059488A (en) | Character recognizer | |
| JP3127413B2 (en) | Character recognition device | |
| JPH0632080B2 (en) | Character recognition method | |
| JPS6031683A (en) | Handwritten character recognition device | |
| JPS62154079A (en) | Character recognition system | |
| JPS6038755B2 (en) | Feature extraction method | |
| JPS6262392B2 (en) | ||
| JPH01152586A (en) | Character graphic recognizing method | |
| JPH06195512A (en) | Character feature extraction device | |
| JPS6019287A (en) | Character recognizing method | |
| JPH0545991B2 (en) |