CN106909897A - A kind of text image is inverted method for quick - Google Patents
A kind of text image is inverted method for quick Download PDFInfo
- Publication number
- CN106909897A CN106909897A CN201710090240.9A CN201710090240A CN106909897A CN 106909897 A CN106909897 A CN 106909897A CN 201710090240 A CN201710090240 A CN 201710090240A CN 106909897 A CN106909897 A CN 106909897A
- Authority
- CN
- China
- Prior art keywords
- text
- line
- effective line
- max
- sequence
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/40—Document-oriented image-based pattern recognition
- G06V30/41—Analysis of document content
- G06V30/413—Classification of content, e.g. text, photographs or tables
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/20—Image preprocessing
- G06V10/24—Aligning, centring, orientation detection or correction of the image
- G06V10/242—Aligning, centring, orientation detection or correction of the image by image rotation, e.g. by 90 degrees
Landscapes
- Engineering & Computer Science (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Multimedia (AREA)
- Theoretical Computer Science (AREA)
- Artificial Intelligence (AREA)
- Character Input (AREA)
- Character Discrimination (AREA)
Abstract
本发明涉及一种文本图像倒置快速检测方法,包括对输入的文本图像进行预处理,得到二值化处理结果用B;进行有效文本行检测,得到有效文本行序列;进行文本行分类,方法如下:1)对于有效文本行序列的各个有效文本行s,填充该有效文本行s相邻字符之间的空白;2)计算各个有效文本行s在垂直方向的投影值,用V(c)表示,其中c表示列序号;3)得到有效文本行s的左边界和右边界;4)得到有效文本行序列的左边界和右边界;5)判断“左缩进文本行”和“右缩进文本行”及“非缩进文本行”;文本图像倒置检测。
The present invention relates to a text image inversion rapid detection method, including preprocessing the input text image to obtain the binarization processing result B; performing effective text line detection to obtain an effective text line sequence; and performing text line classification, the method is as follows 1) For each effective text line s of the effective text line sequence, fill the blank space between the adjacent characters of the effective text line s; 2) Calculate the projection value of each effective text line s in the vertical direction, expressed by V(c) , wherein c represents the column number; 3) obtains the left boundary and the right boundary of the effective text line s; 4) obtains the left boundary and the right boundary of the effective text line sequence; 5) judges "left indentation text line" and "right indentation text line" and "non-indented text line"; text image inversion detection.
Description
技术领域technical field
本发明涉及文本图像增强技术,尤其是针对扫描文本图像的方向倒置检测技术。The invention relates to a text image enhancement technology, in particular to a direction inversion detection technology for a scanned text image.
背景技术Background technique
随着计算机技术的不断发展,基于OCR(光学字符识别)的文本图像数字化技术得到了广泛地应用。在完成OCR过程中,文本图像中的文字方向对字符识别性能影响至关重要。当文字存在倾斜时,如果不加以校正,会严重影响文字的识别率。特别是当文字存在倒置情况(即与正常方向偏差180°左右)。因此,在进行OCR之前,必须判断文本图像是否存在倒置情况,针对倒置情况应考虑首先进行旋转处理,以保证后续识别过程正常执行。With the continuous development of computer technology, text image digitization technology based on OCR (Optical Character Recognition) has been widely used. In the process of completing OCR, the text direction in the text image is very important to the performance of character recognition. When the text is tilted, if it is not corrected, it will seriously affect the recognition rate of the text. Especially when the text is inverted (that is, it deviates from the normal direction by about 180°). Therefore, before performing OCR, it is necessary to judge whether there is an inversion of the text image. For the inversion situation, it should be considered to perform rotation processing first to ensure the normal execution of the subsequent recognition process.
针对存在倾斜情况的文本图像,可以借助现有纠偏算法,检测倾斜度并进行相应地校正。但现有文本图像纠偏方法大都假定输入的文本图像倾斜度在一定范围之内,首先获取倾斜角度信息,进而完成倾斜度校正。但当输入文本图像完全倒置时,现有倾斜角度检测方法基本失效。曾凡锋等人提出了一种基于标点符号的文本图像倒置快速检测方法。该方法首先检测文本字符;然后结合中文字符及标点符号结构特征,筛选出文本图像中的标点符号,根据标点符号像素分布特点,判断标点符号类型;最后结合标点符号使用习惯,判断中文文本图像是否倒置。朱敏等人(专利公开号CN102831421A)提出一种基于标点符号的文本上下方向检测方法。该专利所提方法根据标点符号与文本行的相对位置属性来判断文本的方向,其基本思路与曾凡锋所提方法类似。这类基于标点符号的方法完全依靠标点特征,对于标点符号较少的文本图像无效,因此这类方法适用范围有限,不具有普遍性。For text images with skew, existing skew correction algorithms can be used to detect the skew and correct accordingly. However, most of the existing text image correction methods assume that the slope of the input text image is within a certain range, first obtain the slope angle information, and then complete the slope correction. However, when the input text image is completely inverted, the existing tilt angle detection methods basically fail. Zeng Fanfeng et al. proposed a fast detection method for text image inversion based on punctuation marks. This method first detects text characters; then combines the structural characteristics of Chinese characters and punctuation marks to screen out punctuation marks in text images, and judges the type of punctuation marks according to the distribution characteristics of punctuation marks pixels; finally, combines the usage habits of punctuation marks to judge whether Chinese text images inverted. Zhu Min et al. (Patent Publication No. CN102831421A) proposed a text up-down direction detection method based on punctuation marks. The method proposed in this patent judges the direction of the text based on the relative position attributes of punctuation marks and text lines, and its basic idea is similar to the method proposed by Zeng Fanfeng. Such punctuation-based methods rely entirely on punctuation features and are not effective for text images with fewer punctuation marks. Therefore, such methods have limited scope of application and are not universal.
发明内容Contents of the invention
本发明的目的是克服现有技术的上述不足,提供一种面向文本图像的方向倒置快速检测方法。技术方案如下:The purpose of the present invention is to overcome the above-mentioned deficiencies of the prior art, and provide a fast detection method for orientation inversion oriented to text images. The technical scheme is as follows:
一种文本图像倒置快速检测方法,包括下列步骤:A text image inversion fast detection method comprises the following steps:
第一步:对输入的文本图像进行预处理,得到二值化处理结果用B;Step 1: Preprocess the input text image to obtain the binarized processing result B;
第二步:进行有效文本行检测,得到有效文本行序列;Step 2: Perform valid text line detection to obtain a valid text line sequence;
第三步:进行文本行分类,方法如下:Step 3: Carry out text line classification, the method is as follows:
1)对于有效文本行序列的各个有效文本行s,使用矩形结构算子进行膨胀运算,填充该有效文本行s相邻字符之间的空白;1) For each effective text line s of the effective text line sequence, use the rectangular structure operator to perform expansion operation, and fill the blanks between the adjacent characters of the effective text line s;
2)计算各个有效文本行s在垂直方向的投影值,用V(c)表示,其中c表示列序号;2) Calculate the projection value of each valid text line s in the vertical direction, represented by V(c), where c represents the column number;
3)统计满足条件V(c)>0.5×Rhei(s)的c取值,将c的最小值记为cmin,称为该有效文本行s的左边界;将最大值分别记为cmax,称为有效文本行s的右边界,该扫描行的长度为Rleg=cmax-cmin;3) Count the values of c that satisfy the condition V(c)>0.5×R hei (s), and record the minimum value of c as c min , which is called the left boundary of the effective text line s; record the maximum value as c max is called the right boundary of the effective text line s, and the length of the scan line is R leg =c max -c min ;
4)统计同一个有效文本行序列内各有效文本行对应的cmin(m)和cmax(m),将cmin(m)的最小值称为该有效文本行序列的左边界,记为clef;将cmax(m)的最大值称为该有效文本行序列的右边界,记为crgt;4) Count c min (m) and c max (m) corresponding to each effective text line in the same effective text line sequence, and the minimum value of c min (m) is called the left boundary of the effective text line sequence, which is denoted as c lef ; the maximum value of c max (m) is called the right boundary of the effective text line sequence, which is denoted as c rgt ;
5)对于某有效文本行m,如果满足0.6<|cmin(m)-clef|/|cmax(m)-cmin(m)|<0.9,则将该有效文本行m判为“左缩进文本行”;如果满足0.6<|crgt-cmax(m)|/(cmax(m)-cmin(m))<0.9,则将该有效文本行m判为“右缩进文本行”;如果上述两个条件都不满足,则将该文本行判为“非缩进文本行”;5) For a valid text line m, if 0.6<|c min (m)-c lef |/|c max (m)-c min (m)|<0.9 is satisfied, the valid text line m is judged as " Left-indented text line"; if 0.6<|c rgt -c max (m)|/(c max (m)-c min (m))<0.9 is satisfied, the effective text line m will be judged as "right-indented Enter text line”; if the above two conditions are not met, the text line will be judged as “non-indented text line”;
第四步:文本图像倒置检测,方法如下:Step 4: Text image inversion detection, the method is as follows:
统计单幅文本图像中左缩进文本行和右缩进文本行的数目,分别用Nlef和Nrgt表示;使用下式判断文本图像是否存在倒置:Count the number of left-indented text lines and right-indented text lines in a single text image, represented by N lef and N rgt respectively; use the following formula to determine whether there is an inversion of the text image:
优选地,第二步的方法如下:Preferably, the method of the second step is as follows:
1)计算B中各行在水平方向的投影值,用H(r)表示,其中r表示行号序号;1) Calculate the projection value of each row in B in the horizontal direction, denoted by H(r), where r represents the row number;
2)计算H(r)的最大值,用Hmax表示;2) Calculate the maximum value of H(r), represented by H max ;
3)对于第r扫描行,如果满足H(r)>0.5×Hmax,则将该行判为一个有效扫描行;3) For the r-th scanning line, if H(r)>0.5×H max is satisfied, the line is judged as a valid scanning line;
4)统计各有效扫描行的分布情况,如果检测到连续m行被判为有效扫描行,且满足m>M/100,则由这连续m个有效扫描行组成一个有效文本行序列;4) Statistical distribution of each effective scanning line, if detected that continuous m lines are judged as effective scanning lines, and satisfy m>M/100, then form an effective text line sequence by these continuous m effective scanning lines;
确定该有效文本行序列中最上方和最下方有效扫描行的行号,用Rtop(s)和Rbot(s)分别表示该有效文本行序列的上下边界,定义该有效文本行序列的高度为Rhei(s)=|Rtop(s)-Rbot(s)|,符号|·|表示取绝对值符号,式中s是有效文本行的序号。Determine the line numbers of the top and bottom effective scanning lines in this effective text line sequence, represent the upper and lower boundaries of this effective text line sequence respectively with R top (s) and R bot (s), define the height of this effective text line sequence R hei (s)=|R top (s)-R bot (s)|, the symbol |·| represents the absolute value symbol, where s is the serial number of a valid text line.
附图说明Description of drawings
图1是本发明所提方法的流程图Fig. 1 is the flowchart of proposed method of the present invention
图2是本发明所用到的重要定义示意图Fig. 2 is a schematic diagram of important definitions used in the present invention
图3本发明定义的文本行类型示意图Fig. 3 schematic diagram of text line type defined by the present invention
图4本发明实验所用的文本图像左右缩进文本行数目示意图Figure 4 is a schematic diagram of the number of left and right indented text lines in the text image used in the experiment of the present invention
具体实施方式detailed description
首先将输入文本彩色图像进行灰度化、双边滤波、对比度增强、二值化等预处理操作,提高文档图像视觉质量;然后借助水平投影分析,检测文本图像中的有效文本行,并结合文本行的位置和长度特征,对文本行进行分类;最后根据左缩进文本行和右缩进文本行的相对数目,判断文本图像是否存在倒置。图1所示为所提方法的框图。Firstly, preprocessing operations such as grayscale, bilateral filtering, contrast enhancement, and binarization are performed on the input text color image to improve the visual quality of the document image; then, with the help of horizontal projection analysis, effective text lines in the text image are detected and combined with text lines Classify the text lines based on the position and length features of the text; finally, according to the relative number of left-indented text lines and right-indented text lines, it is judged whether there is an inversion of the text image. Figure 1 shows the block diagram of the proposed method.
首先给出若干有用的定义。一幅文本图像由多个段落构成,每一个段落中的字符字体、格式等特征基本一致。本发明将每一段落内各字符所能出现的最左侧和最右侧位置,分别称为“段落左边界”和“段落右边界”。图2给出了一个段落左边界和右边界的示意图。每个段落可能包括一个或多个文本行,对于任意文本行,将其最左侧字符的左侧和最右侧字符的右侧分别称为该文本行的“行左边界”和“行右边界”。图2给出了行左边界和行右边界的示意图。First some useful definitions are given. A text image is composed of multiple paragraphs, and the character fonts and formats in each paragraph are basically the same. In the present invention, the leftmost and rightmost positions where each character can appear in each paragraph are called "paragraph left boundary" and "paragraph right boundary" respectively. Figure 2 shows a diagram of the left and right borders of a paragraph. Each paragraph may include one or more text lines. For any text line, the left side of the leftmost character and the right side of the rightmost character are called the "line left boundary" and "line right boundary" of the text line, respectively. boundary". Figure 2 shows a schematic diagram of the row left boundary and the row right boundary.
对于某个文本行,其左右边界与所属段落的左右边界基本重合,则称该文本行为“完整文本行”。如果对于某个文本行,其左边界距离所属段落的左边界有2~4个字符,同时该文本行右边界与所属段落右边界基本重合,则称该文本行为“左缩进文本行”。对于某个文本行,其右边界距离所属段落右边界有2~4个字符,同时该文本行左边界与所属段落左边界基本重合,则称该文本行为“右缩进文本行”。图3给出了上述三类文本行的示意图。For a text line whose left and right borders basically coincide with the left and right borders of the paragraph to which it belongs, the text line is said to be a "complete text line". If for a certain text line, its left border is 2 to 4 characters away from the left border of the paragraph to which it belongs, and at the same time, the right border of the text line basically coincides with the right border of the paragraph to which it belongs, then the text behavior is called "left indented text line". For a text line, if its right border is 2 to 4 characters away from the right border of the paragraph to which it belongs, and at the same time, the left border of the text line basically coincides with the left border of the paragraph to which it belongs, then the text behavior is called a "right indented text line". Figure 3 shows a schematic diagram of the above three types of text lines.
根据中英文的书写习惯,每一个段落首行字符一般向右缩进2~4个字符,即对于包含两个或两个以上文本行的段落,其必然存在一个左缩进文本行。如果文本图像是正向的,则必然能检测到多个左缩进文本行。反之,如果文本图像是倒置的,则能检测到多个右缩进文本行。本发明正是通过检测和判断文本图像中左缩进文本行和右缩进文本行的相对数目,来判断该文本图像是否存在倒置情况。According to the writing habits of Chinese and English, the characters in the first line of each paragraph are generally indented to the right by 2 to 4 characters, that is, for a paragraph containing two or more text lines, there must be a left indented text line. If the text image is positive, multiple left-indented text lines must be detected. Conversely, multiple right-indented text lines can be detected if the text image is inverted. The present invention judges whether there is an inversion of the text image by detecting and judging the relative numbers of left-indented text lines and right-indented text lines in the text image.
本发明所提方法具体处理过程包括:预处理、文本行检测、文本行分类、文本方向倒置检测等四个主要步骤。The specific processing process of the method proposed in the present invention includes four main steps: preprocessing, text line detection, text line classification, and text direction inversion detection.
1、预处理1. Pretreatment
预处理的目的是提高文档图像的视觉质量,主要包括:灰度化、平滑滤波、对比度增强和二值化等步骤。The purpose of preprocessing is to improve the visual quality of the document image, which mainly includes steps such as grayscale, smoothing filter, contrast enhancement and binarization.
(1)灰度化:(1) Gray scale:
判断输入文本图像是否是灰度图像,如果是灰度图像,则保持不变;如果是彩色图像,用CR、CG和CB分别表示红、绿、蓝三个颜色通道,使用式(1)计算灰度图像,用I表示。Determine whether the input text image is a grayscale image, if it is a grayscale image, it remains unchanged; if it is a color image, use C R , C G and C B to represent the three color channels of red, green and blue respectively, using the formula ( 1) Calculate the grayscale image, denoted by I.
I(x,y)=min{CR(x,y),CG(x,y),CB(x,y)} (1)I(x,y)=min{C R (x,y),C G (x,y),C B (x,y)} (1)
式中,x=0,1,2,...,M-1,y=0,1,2,...,N-1,M和N分别是文本图像的高度和宽度,即扫描总行数和扫描总列数。In the formula, x=0, 1, 2,..., M-1, y=0, 1, 2,..., N-1, M and N are the height and width of the text image respectively, that is, the scanning total line and scan the total number of columns.
(2)平滑滤波(2) smoothing filter
考虑到文本图像在采集及数字化过程中受到噪声污染,采用双边滤波技术对灰度图像I进行滤波处理,降低噪声影响。用G表示经双边滤波处理后的图像。Considering that the text image is polluted by noise in the process of acquisition and digitization, the grayscale image I is filtered by bilateral filtering technology to reduce the influence of noise. G represents the image processed by bilateral filtering.
(3)对比度增强(3) Contrast enhancement
由于光照等原因的影响,文本图像的对比度可能偏低,采用直方图均衡技术对滤波图像G进行增强处理,处理结果用E表示。Due to the influence of lighting and other reasons, the contrast of the text image may be low, and the histogram equalization technology is used to enhance the filtered image G, and the processing result is represented by E.
(4)二值化处理(4) Binary processing
使用经典的Otsu法计算E对应的全局阈值,用Th表示。使用Th对E进行二值化处理,处理结果用B来表示,具体的做法是:The global threshold corresponding to E is calculated using the classic Otsu method, denoted by T h . Use T h to binarize E, and the processing result is represented by B. The specific method is:
其中,B中取值为1的点代表文本点,取值为0的点代表背景点。Among them, the points with a value of 1 in B represent text points, and the points with a value of 0 represent background points.
2、有效文本行检测2. Valid text line detection
使用以下算法完成有效文本行检测:Valid text line detection is done using the following algorithm:
有效文本行检测算法:Valid text line detection algorithm:
1)计算B中各行在水平方向的投影值,用H(r)表示,其中r表示行号序号。1) Calculate the projection value of each row in B in the horizontal direction, denoted by H(r), where r represents the row number.
2)计算H(r)的最大值,用Hmax表示。2) Calculate the maximum value of H(r), denoted by H max .
3)对于第r扫描行,如果满足H(r)>0.5×Hmax,则将该行判为一个有效扫描行。3) For the r-th scanning line, if H(r)>0.5×H max is satisfied, the line is judged as a valid scanning line.
4)统计各有效扫描行的分布情况,如果检测到连续m行被判为有效扫描行,且满足m>M/100,则由这连续m个有效扫描行组成一个有效文本行。4) Count the distribution of each effective scanning line. If it detects that m consecutive lines are judged as effective scanning lines and satisfies m>M/100, then a valid text line is composed of these m continuous effective scanning lines.
5)确定该有效文本行中最上方和最下方有效扫描行的行号,用Rtop(s)和Rbot(s)分别表示该文本行的上下边界,定义该文本行的高度为Rhei(s)=|Rtop(s)-Rbot(s)|,符号|·|表示取绝对值符号,式中s是有效文本行的序号。5) Determine the line numbers of the top and bottom effective scanning lines in the effective text line, represent the upper and lower boundaries of the text line with R top (s) and R bot (s), define the height of the text line as R hei (s)=|R top (s)-R bot (s)|, the symbol |·| represents the absolute value symbol, where s is the sequence number of a valid text line.
3、文本行分类3. Classification of text lines
使用以下算法完成文本行分类:Text line classification is done using the following algorithm:
文本行分类算法:Text line classification algorithm:
1)对于某一个有效文本行s,使用矩形结构算子对该文本行进行膨胀运算,填充该文本行相邻字符之间的空白。矩形结构算子的高度为2个像素,宽度为该文本行高度50%。1) For a valid text line s, use the rectangular structure operator to perform expansion operation on the text line, and fill the blank space between the adjacent characters of the text line. The height of the rectangular structure operator is 2 pixels, and the width is 50% of the height of the text line.
2)计算文本行在垂直方向的投影值,用V(c)表示,其中c表示列序号。2) Calculate the projection value of the text line in the vertical direction, denoted by V(c), where c represents the column number.
3)统计满足条件V(c)>0.5×Rhei(s)的c取值,将c的最小值记为cmin,称为该文本行的左边界;将最大值分别记为cmax,称为该文本行的右边界,该扫描行的长度为Rleg=cmax-cmin。3) Count the values of c that satisfy the condition V(c)>0.5×R hei (s), record the minimum value of c as c min , which is called the left boundary of the text line; record the maximum value as c max , Called the right boundary of the text line, the length of the scan line is R leg =c max - c min .
4)统计同一个段落内各有效文本行对应的cmin(m)和cmax(m),将cmin(m)的最小值称为该段落的左左边界,记为clef;将cmax(m)的最大值称为该段落的右边界,记为crgt。4) Count c min (m) and c max (m) corresponding to each effective text line in the same paragraph, and the minimum value of c min (m) is called the left and left boundary of the paragraph, which is denoted as c lef ; c The maximum value of max (m) is called the right boundary of the paragraph, denoted as c rgt .
5)对于某有效文本行m,如果满足0.6<|cmin(m)-clef|/|cmax(m)-cmin(m)|<0.9,则将该文本行判为“左缩进文本行”;如果满足0.6<|crgt-cmax(m)|/(cmax(m)-cmin(m))<0.9,则将该文本行判为“右缩进文本行”;如果上述两个条件都不满足,则将该文本行判为“非缩进文本行”。5) For a valid text line m, if 0.6<|c min (m)-c lef |/|c max (m)-c min (m)|<0.9 is satisfied, the text line is judged as "left indentation enter the text line"; if 0.6<|c rgt -c max (m)|/(c max (m)-c min (m))<0.9 is satisfied, the text line is judged as "right indented text line"; If the above two conditions are not satisfied, then judge the text line as "non-indented text line".
4、文本图像倒置检测4. Text image inversion detection
统计单幅文本图像中左缩进文本行和右缩进文本行的数目,分别用Nlef和Nrgt表示。使用式(3)判断文本图像是否存在倒置:Count the number of left-indented text lines and right-indented text lines in a single text image, denoted by N lef and N rgt respectively. Use formula (3) to judge whether there is an inversion of the text image:
实施例如下:Examples are as follows:
采用Windows10专业版系统下的Matlab2015a作为实验仿真平台,硬件平台是Intel i5-6200U CPU,8G内存。Matlab2015a under Windows 10 Professional Edition system is used as the experimental simulation platform, and the hardware platform is Intel i5-6200U CPU, 8G memory.
选用专利申请人自行采集的90幅文本图像作为测试集,其中倒置文本图像78幅,正方向文本图像12幅。在全部90幅文本图像中,中文文本图像有56幅,占62%,英文文本图像34幅,占38%。采用本发明提出方法对测试图像进行处理,100%的倒置图像都正常检出。图4给出了90幅文档图像中左缩进文本行和右缩进文本行数目的分布情况。由图可见,对于正向文本图像,其左缩进文本行数明显大于右缩进文本行数;反之,对于倒置方向文本图像,是右缩进文本行数大于左缩进文本行数。很明显分为两类,即倒置文本图像类(在图中用符号“*”标识)和正向文本图像类(在图中用符号“o”标识)。90 text images collected by the patent applicant are selected as the test set, including 78 inverted text images and 12 positive text images. Among all 90 text images, there are 56 Chinese text images, accounting for 62%, and 34 English text images, accounting for 38%. The test image is processed by the method proposed by the invention, and 100% of the inverted images are detected normally. Figure 4 shows the distribution of the number of left-indented text lines and right-indented text lines in 90 document images. It can be seen from the figure that for a forward text image, the number of left-indented text lines is significantly greater than the number of right-indented text lines; conversely, for an inverted text image, the number of right-indented text lines is greater than the number of left-indented text lines. It is obviously divided into two categories, that is, the inverted text image category (identified by the symbol "*" in the figure) and the forward text image category (identified by the symbol "o" in the figure).
测试图像的尺寸是1944×2592分辨率达到了5000万像素,处理一幅图像的平均速度约为2300ms,如果换成执行效率更高的C语言编写算法,处理速度会更快,能够满足实时处理的要求。The size of the test image is 1944×2592 with a resolution of 50 million pixels. The average speed of processing an image is about 2300ms. If the algorithm is written in C language with higher execution efficiency, the processing speed will be faster and can meet real-time processing requirements.
由实验结果可见,采用本发明所述方法,可以快速有效的判断输入的扫描文本图像是否存在倒置情况,并能对包括中英文在内的多种语言类型的文本图像进行处理。It can be seen from the experimental results that the method of the present invention can quickly and effectively judge whether the input scanned text image is inverted, and can process text images in multiple languages including Chinese and English.
本发明的步骤总结如下:The steps of the present invention are summarized as follows:
步骤1:判断输入扫描文本图像类型,如果是灰度图像,则保持不变;如果是彩色图像,则用式(1)转换为灰度图像,用I表示灰度图像。Step 1: Judging the type of the input scanned text image, if it is a grayscale image, it remains unchanged; if it is a color image, then use formula (1) to convert it into a grayscale image, and use I to represent the grayscale image.
步骤2:采用双边滤波技术对灰度图像I进行滤波处理,滤波结果用G表示。Step 2: Use bilateral filtering technology to filter the grayscale image I, and the filtering result is represented by G.
步骤3:采用直方图均衡技术,对滤波结果图像G进行增强处理,处理结果用E表示。Step 3: Use the histogram equalization technique to enhance the filtering result image G, and denote the processing result by E.
步骤4:使用Otsu法计算增强结果图像的全局阈值,并结合式(2)对E进行二值化处理,处理结果用B表示。Step 4: Use the Otsu method to calculate the global threshold of the enhanced result image, and combine Equation (2) to binarize E, and the processing result is denoted by B.
步骤5:采用有效文本行检测算法,检测扫描文本图像中的有效文本行。Step 5: Use a valid text line detection algorithm to detect valid text lines in the scanned text image.
步骤6:使用文本行分类算法,对每一个有效文本行进行分类,确定左缩进文本行和右缩进文本行的数目Nlef和Nrgt。Step 6: Use the text line classification algorithm to classify each valid text line, and determine the numbers N lef and N rgt of left-indented text lines and right-indented text lines.
步骤7:结合式(3),判断扫描文本图像是否存在倒置情况。Step 7: Combining formula (3), determine whether the scanned text image is inverted.
Claims (2)
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201710090240.9A CN106909897B (en) | 2017-02-20 | 2017-02-20 | Text image inversion rapid detection method |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201710090240.9A CN106909897B (en) | 2017-02-20 | 2017-02-20 | Text image inversion rapid detection method |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| CN106909897A true CN106909897A (en) | 2017-06-30 |
| CN106909897B CN106909897B (en) | 2020-03-13 |
Family
ID=59208458
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| CN201710090240.9A Expired - Fee Related CN106909897B (en) | 2017-02-20 | 2017-02-20 | Text image inversion rapid detection method |
Country Status (1)
| Country | Link |
|---|---|
| CN (1) | CN106909897B (en) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107609482A (en) * | 2017-08-15 | 2018-01-19 | 天津大学 | A kind of Chinese text image inversion method of discrimination based on Chinese-character stroke feature |
| CN111414866A (en) * | 2020-03-24 | 2020-07-14 | 上海眼控科技股份有限公司 | Vehicle application form detection method and device, computer equipment and storage medium |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102831421A (en) * | 2012-08-29 | 2012-12-19 | 华东师范大学 | Method for detecting document up-down direction based on punctuation marks |
| CN106097254A (en) * | 2016-06-07 | 2016-11-09 | 天津大学 | A kind of scanning document image method for correcting error |
-
2017
- 2017-02-20 CN CN201710090240.9A patent/CN106909897B/en not_active Expired - Fee Related
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102831421A (en) * | 2012-08-29 | 2012-12-19 | 华东师范大学 | Method for detecting document up-down direction based on punctuation marks |
| CN106097254A (en) * | 2016-06-07 | 2016-11-09 | 天津大学 | A kind of scanning document image method for correcting error |
Non-Patent Citations (2)
| Title |
|---|
| 曾凡锋等: "中文文本图像倒置快速检测算法", 《计算机工程与设计》 * |
| 朱其猛: "基于文字结构特征的文本图像方向的研究与应用", 《中国优秀硕士学位论文全文数据库信息科技辑》 * |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107609482A (en) * | 2017-08-15 | 2018-01-19 | 天津大学 | A kind of Chinese text image inversion method of discrimination based on Chinese-character stroke feature |
| CN111414866A (en) * | 2020-03-24 | 2020-07-14 | 上海眼控科技股份有限公司 | Vehicle application form detection method and device, computer equipment and storage medium |
Also Published As
| Publication number | Publication date |
|---|---|
| CN106909897B (en) | 2020-03-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN111814722B (en) | A form recognition method, device, electronic device and storage medium in an image | |
| CN102208023B (en) | Method for recognizing and designing video captions based on edge information and distribution entropy | |
| CN103258198B (en) | Character extracting method in a kind of form document image | |
| US8548246B2 (en) | Method and system for preprocessing an image for optical character recognition | |
| CN105205488B (en) | Word area detection method based on Harris angle points and stroke width | |
| CN101515325A (en) | Character extracting method in digital video based on character segmentation and color cluster | |
| US20130322757A1 (en) | Document Processing Apparatus, Document Processing Method and Scanner | |
| CN108171104A (en) | A kind of character detecting method and device | |
| CN106446881A (en) | Method for extracting lab test result from medical lab sheet image | |
| CN110598566A (en) | Image processing method, device, terminal and computer readable storage medium | |
| CN111461126B (en) | Method, device, electronic device and storage medium for identifying spaces in text lines | |
| CN106127817B (en) | A kind of image binaryzation method based on channel | |
| CN101122952A (en) | A method of image text detection | |
| CN114495141B (en) | Document paragraph position extraction method, electronic device and storage medium | |
| Al Abodi et al. | An effective approach to offline Arabic handwriting recognition | |
| CN108830269B (en) | A method for determining the width of the central axis of Manchu words | |
| CN108710882A (en) | A kind of screen rendering text recognition method based on convolutional neural networks | |
| CN101727583B (en) | Self-adaption binaryzation method for document images and equipment | |
| CN103606220A (en) | Check printed number recognition system and check printed number recognition method based on white light image and infrared image | |
| CN114821601A (en) | End-to-end English handwritten text detection and recognition technology based on deep learning | |
| US20190266431A1 (en) | Method, apparatus, and computer-readable medium for processing an image with horizontal and vertical text | |
| CN105303190B (en) | A kind of file and picture binary coding method that degrades based on contrast enhancement methods | |
| CN108717544B (en) | An automatic detection method of newspaper sample text based on intelligent image analysis | |
| CN106909897B (en) | Text image inversion rapid detection method | |
| CN107798355B (en) | Automatic analysis and judgment method based on document image format |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PB01 | Publication | ||
| PB01 | Publication | ||
| SE01 | Entry into force of request for substantive examination | ||
| SE01 | Entry into force of request for substantive examination | ||
| GR01 | Patent grant | ||
| CF01 | Termination of patent right due to non-payment of annual fee |
Granted publication date: 20200313 |
|
| CF01 | Termination of patent right due to non-payment of annual fee |