WO2011027113A1 - Procédé et dispositif de segmentation d'images - Google Patents
Procédé et dispositif de segmentation d'images Download PDFInfo
- Publication number
- WO2011027113A1 WO2011027113A1 PCT/GB2010/001661 GB2010001661W WO2011027113A1 WO 2011027113 A1 WO2011027113 A1 WO 2011027113A1 GB 2010001661 W GB2010001661 W GB 2010001661W WO 2011027113 A1 WO2011027113 A1 WO 2011027113A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- segment
- section
- searching
- location
- document
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/40—Extraction of image or video features
- G06V10/42—Global feature extraction by analysis of the whole pattern, e.g. using frequency domain transformations or autocorrelation
- G06V10/422—Global feature extraction by analysis of the whole pattern, e.g. using frequency domain transformations or autocorrelation for representing the structure of the pattern or shape of an object therefor
- G06V10/424—Syntactic representation, e.g. by using alphabets or grammars
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/10—Character recognition
- G06V30/14—Image acquisition
- G06V30/1444—Selective acquisition, locating or processing of specific regions, e.g. highlighted text, fiducial marks or predetermined fields
- G06V30/1448—Selective acquisition, locating or processing of specific regions, e.g. highlighted text, fiducial marks or predetermined fields based on markings or identifiers characterising the document or the area
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/10—Character recognition
Definitions
- This invention relates to a method and an apparatus for obtaining segment information of an image.
- the exam candidates are required to provide their answers in a fixed area within a page of a document so that each answer can be separated in the scanned image and sent to the relevant examiner,
- This method of distributing answers to examiners is not possible with free form answers, where the length of an answer is unpredictable, the questions can be answered in a random order and the answers can be broken into separate sections,
- Marking an examination script from an image stored in a server or online database can be used as part of a quality control system, particularly if the individual answers can be identified from within the individual examination document.
- the present invention provides a method of obtaining segment information of a document image, wherein the document has a first section and a second section, comprising the steps of: searching the first section for a first indicating text and determining the location thereof; in the second section, setting a first position as the start of a segment based on the location of the first indicating text; searching the first section, downstream from the first indicating text, for a second indicating text and determining the location thereof; and in the second section, setting a second position as the end of the segment based on the location of the second indicating text.
- the method described above further comprises the step of searching the second section adjacent to or in the vicinity of the location of the first indicating text for a first blank area having a predetermined height and width and not containing any free text, wherein the first blank area is set as the start of the segment.
- the method described above further comprises, if a first blank area adjacent to or in the vicinity of the first indicating text is not found, searching the second section in an upstream direction for a blank area and setting the first found blank area as the start of the segment.
- the start of a free form text is determined by searching for a blank area across the second section and close to or above the location of the first indicating text (e.g. a question number).
- the first indicating text e.g. a question number
- the start location of the segment will account for sloped writing.
- the methods described above further comprise the step of searching the second section adjacent to or in the vicinity of the location of the second indicating text for a second blank area having a predetermined height and width and not containing any free text, wherein the second blank area is set as the end of the segment.
- the method described above further comprises the step of searching in an upstream direction for further blank areas above the second blank area and setting the end of the segment as the location of a blank area that is adjacent to the free text or the start of the segment.
- the end of a free form answer is determined by searching for a blank area across the second section and close to or above the location of the second indicating text (e.g. another, subsequently answered question number).
- the writer may leave spaces within a segment (e.g. when answering a question), thus a blank area in the second section may not indicate the end of a segment.
- the location of the start of a subsequent segment e.g. an answer to another question
- searching for a blank area close to or above the location of the second indicating text in the first section provides greater certainty that such a blank area is the end of the segment.
- the first section has a predetermined format
- the step of searching the first section for a first indicating text comprises the steps of: searching for any mark in the first section that is different to the predetermined format; applying character recognition to said mark to derive a character; comparing the character obtained from the mark to a list of reference characters; and if the derived character relates to a reference character, setting that character as the first indicating text.
- the characters that are horizontally adjacent to each other form a character string and if the character string relates to reference characters, setting that character string as the first indicating text.
- the first section has a predetermined format
- the step of searching the first section for a second indicating text comprises the steps of: searching for any mark in the first section that is different to the predetermined format; applying character recognition to said mark to derive a character; comparing a character obtained from the mark to a list of reference characters; and if the character relates to a reference character, setting the character as the second indicating text.
- the derived characters that are horizontally adjacent to each other form a character string and if the character string relates to reference characters, setting that character string as the first or second indicating text.
- the predetermined format is known. Thus any marks made by the writer, which are not part of the predetermined format, are easily found. Comparing a recognised character against a list of reference characters (e.g question numbers) allows greater reliably and certainty that the mark and recognised character is a reference character and not an erroneous mark made by the exam candidate.
- reference characters e.g question numbers
- a confidence level that a character is recognised from the mark is higher than a confidence level threshold.
- the predetermined format comprises a plurality of regularly spaced response areas that are distinguished from the surroundings.
- the surroundings are shaded and the response areas are unshaded.
- the response areas indicate to the writer where the reference character (e.g. exam question number) should be written.
- the surroundings may be shaded in a certain colour that is invisible to the scanner.
- the searching means may remove the shading colour from the image, leaving a clear background for character recognition. This clearer image allows easier searching and recognition of the marks.
- Searching the response areas only and not the whole of the first section can increase the speed and reliability of obtaining segmentation information.
- the document has one or more pages, each page having a sequence indicator which indicates the downstream direction.
- a barcode can indicate the sequence of pages in a document.
- the writer is not restricted in the length of text (e.g. an answer to a question) he or she may write.
- the writer may use more than one page for a segment of text.
- the first section is adjacent to and left of the second section.
- the first section is adjacent to and right of the second section.
- the first and second indicating text are handwritten.
- the free text is handwritten.
- the segment information comprises information related to the first indicating text, the location of the start of the segment and the location of the end of the segment.
- the invention also provides a second method comprising the steps of: scanning a document; while scanning the document, obtaining segment information of the document; compressing an image of the document; and storing the segment information and the compressed image related to the document.
- segment information is obtained while scanning the document thereby allowing character recognition to be carried out prior to compressing the image.
- Storing a compressed image and associated segment information data requires less storage space than storing a high resolution image.
- the step of obtaining segment information is done in a buffer memory of a scanner.
- a temporary high resolution image of the scanned document is stored in the buffer memory of the scanner.
- Character recognition can be carried out with greater accuracy on the high resolution image compared with a compressed, low resolution image that is used for storage and transmission purposes.
- segment information can be obtained with greater accuracy from the high resolution image.
- processing for obtaining segment information can be carried out by the scanner, there is no requirement to carry out additional processing of the image by a computer or server connected to the scanner, thus saving time and processing resources.
- the step of obtaining segment information comprises a method according to the method described above.
- the second method described above further comprises a method of displaying a segment, the method of displaying a segment comprising the steps of: requesting a segment; obtaining only parts of the compressed image which correspond to the requested segment; and displaying said parts.
- the second method described above further comprises the step of assigning and submitting, to a database, data related to the displayed parts.
- Assigning submitting and storing data for each segment (e.g. marks for each individual answer), rather than a whole document, allows efficient storage and retrieval such data.
- the marks assigned to an answer can be easily retrieved and compared to marks assigned by other examiners for the same question for quality control purposes.
- said parts are displayed in a sequence according to the segment information and the downstream direction.
- the writer is free to write text associated with an indicating text in different areas of the document (e.g. an answer to a question can be written in different areas of the document).
- viewing text with parts dispersed in different areas is inefficient and difficult (e.g. examining an answer is difficult if it is dispersed in different areas).
- the present invention is able to collate the different parts and display them in sequence or as a single image. This allows, for example, easier and more efficient marking of an answer.
- the step of requesting and displaying can occur at a location remote from a location of the stored segment information and the stored compressed image.
- Storing the compressed images and segment information data in a server allows easy access over a network such as the internet. Thus, for example, multiple examiners can be utilised.
- the document is an examination script.
- the invention also provides an apparatus for obtaining segment information of a document image, wherein the document has a first section and a second section, the apparatus comprising: a means for searching the first section; and a means for setting segment information, wherein: the means for searching the first section searches for a first indicating text and determines the location thereof; the means for setting segment information sets, based on the location of the first indicating text, a first position in the second section as the start of a segment; the means for searching the first section searches, downstream from the first indicating text, for a second indicating text and determines the location thereof; and the means for setting segment information sets, based on the location of the second indicating text, a second position in the second section as the end of the segment.
- the apparatus described above further comprising a means for searching the second section, wherein: the means for searching the second section searches adjacent to or in the vicinity of the location of the first indicating text for a first blank area having a predetermined height and width and not containing any free text; and the means for setting the segment information sets the first blank area as the start of the segment.
- the means for searching the second section in the apparatus described above if the means for searching the second section can not find a first blank area adjacent to or in the vicinity of the first indicating text, the means for searching the second section searches in an upstream direction for a blank area and the means for setting the segment information sets the first found blank area as the start of the segment.
- the means for searching the second section in the apparatus described above the means for searching the second section searches adjacent to or in the vicinity of the location of the second indicating text for a second blank area having a predetermined height and width and not containing any free text; and the means for setting the segment information sets the second blank area as the end of the segment.
- the means for searching the second section searches in an upstream direction for further blank areas above the second blank area and the means for setting the segment information sets the end of the segment as the location of a blank area that is adjacent to the free text or the start of the segment.
- the means for searching the second section in the apparatus described above if the means for searching the second section can not find a second blank area adjacent to or in the vicinity of the second indicating text, the means for searching the second section searches in an upstream direction for a blank area and the means for setting the segment information sets the first found blank area as the end of the segment.
- the means for searching the second section searches, in a downstream direction from the second indicating text for a blank area and the means for setting the segment information sets the first found blank area as the end of the segment.
- the means for searching the second section searches for a blank area that persists to the end of the document and the means for setting segment information sets the start location of the blank area as the end of the segment.
- the means for searching the first section comprises: a means for searching for any mark that is different to the predetermined format; a means for character recognition of said mark; a means for comparing a character obtained from the character recognition means to a list of reference characters; and a means for setting the character as the first or second indicating text if the character relates to a reference character.
- the derived characters that are horizontally adjacent to each other form a character string and if the character string relates to reference characters, the means for setting the character sets that character string as the first or second indicating text.
- the means for character recognition recognises more than one character from the mark.
- the means for character recognition comprises a confidence level threshold, wherein a confidence level that a character is recognised from the mark is higher than the confidence level threshold.
- the predetermined format comprises a plurality of regularly spaced response areas that are distinguished from the surroundings.
- the means for searching the first section only searches the response areas for marks.
- the document in the apparatus described above the document has one or more pages and each page has a sequence indicator which indicates the downstream direction.
- the first section is adjacent to and left of the second section.
- the first section is adjacent to and right of the second section
- the ⁇ first and second indicating text are handwritten.
- the free text is handwritten.
- the segment information comprises information related to the first indicating text, the location of the start of the segment and the location of the end of the segment.
- the invention also provides for a second apparatus comprising: a means for scanning a document; a means for obtaining segment information of the scanned document; a means for compressing an image of the scanned document; and a means for storing the segment information and the compressed image related to the scanned document, wherein the segment information is obtained while scanning the document.
- the segment information is obtained from an image in a buffer memory of a scanner.
- the means for obtaining segment information comprises an apparatus according to any one of the apparatuses described above.
- the second apparatus described above further comprises: a means for requesting a segment; a means for obtaining parts of the compressed image which correspond to the requested segment; and a means for displaying said parts.
- the second apparatus further comprising a means for assigning and submitting, to a database, data related to the displayed parts.
- said parts are displayed in a sequence according to the segment information and the downstream direction.
- means for requesting a segment and the means for displaying the parts are at a location remote from a location of the means for storing the segment information and the compressed image related to the scanned document.
- the document is an examination script.
- the invention also provides for a system comprising: an apparatus, connected to a network, comprising: a means for scanning a document; a means for obtaining segment information of the scanned document while scanning the document according to the apparatus described above; and a means for compressing an image of the scanned document; a network device, connected to the network, comprising: a means for storing the segment information and the compressed image related to the scanned document; and a client computer, connected to the network, comprising: a means for requesting a segment; a means for obtaining parts of the compressed image which correspond to the requested segment from the apparatus; and a means for displaying said parts.
- the client computer further comprises a means for assigning and submitting, to the apparatus, data related to the displayed parts.
- the invention also provides for a computer program comprising instructions that when executed by a computer system, cause the computer system to perform a method according to any one of the above methods.
- Figure 1 depicts an example of a front page of an exam document
- Figure 2 depicts an example of an answer page of an exam document
- Figure 3 depicts a system according to an embodiment of the present invention
- Figure 4 depicts the architecture of a scanner according to an embodiment of the present invention.
- Figures 5A to 5D depict a process of segmenting an image according to an embodiment of the present invention.
- Figure 1 depicts an example of the front page 10 of a type of exam document used in the invention.
- the document has been designed to maximise the ability to read characters using automated data capture techniques.
- the document includes features that permit the document to be analysed when it is passed through an imaging scanner.
- the front page comprises predetermined formatted areas 1 1 (which can, for example, be shaded e.g. coloured), each with a number of response areas 12 (which can, for example, be unshaded) where the exam candidate can write the details of the examination being attempted.
- the details on the front page can include the exam paper number, the candidate number, the examination date, the location number and any other details. These written details can be read by a variety of means, which may be ICR software, and inputted into a central database.
- the front page also comprises an indicator 13.
- This indicator can be a barcode or any other machine-readable image.
- the indicator can be used to identify the orientation, the type of document, the number of pages in the document and any other information.
- the front page also comprises the document version indicator 14, which can be a barcode and identifies the specific document type. These indicators are linked to the written details, such as the candidate number and the exam paper number, stored in the central database.
- Figure 2 depicts an example of an answer page 20 that is part of the exam document.
- Each page can comprise a page number 21 to indicate the page number and order of the pages in the document.
- the pages can comprise a barcode 22, or any other type of machine-readable image, to indicate the page number, the document the page belongs to and any other information.
- the page comprises a first section 23 which is used to enter the characters which indicates the exam question being answered.
- the characters can be numbers, symbols or text in any language (for example, it is possible to use Latin, Greek, Cyrillic, Arabic or Asian characters).
- the first section 23 can comprise a
- predetermined formatted area 24 which can, for example, be shaded
- response areas 25 which can, for example, be unshaded
- the response areas are shown as boxes in Figure 2, however the response areas can be of any shape.
- the exam candidate enters the exam question details in the response areas.
- Figure 2 shows three boxes per line where the exam question details are entered, however one or more response areas per line can be utilised.
- the second section 26 Adjacent to the first section 23 is a second section 26.
- the second section 26 can have a larger area than the first section.
- the second section can comprise horizontal lines 27 that are vertically spaced to correspond to the top and bottom of the response areas in the first section.
- the second section can also comprise horizontal lines in the form of a musical stave.
- the second section can also be a grid design (for example, graph paper).
- the second section can also be blank with no lines.
- Free form text e.g. an answer to an exam question that can be hand written or typed or a diagram
- the exam candidate is instructed to write a question number in the response areas 25 in the first section 23 and begin writing an answer to the question on a line in the second section 26 adjacent to the question number in the first section 23.
- the exam candidate writes the numbers in response areas 25 that are created within the predetermined format area 24 to give structure to the blank document.
- the predetermined format area 24 can be removed from the image when scanned to leave a clear image of the character that needs to be recognised.
- the response areas 25 are regularly repeated down the answer page and the exam candidate writes which answer is being attempted adjacent to the beginning of the answer, the beginning of each answer can be electronically deduced.
- There is some tolerance in the system in case the exam candidate makes a mistake in writing the number and uses nearby space up or down the page.
- the system can be configured to include all of the answer text below detection of a clear line in the second section 26. The candidates are instructed to leave a clear line between each question attempted to assist facilitation this procedure.
- the document can be designed in a variety of formats depending upon the questions being used and the language in which the examination is being completed.
- the first section can be placed on the right hand side of the second section. This is advantageous for languages where text is read and written from right to left.
- the document can be scanned to create an image and associated metadata, which can be used to send the image of that answer to an appropriate qualified examiner (note that it is only that answer, not the entire document).
- the examiner is given tools to assist in marking the candidate's work from the image and to pass those marks and comments back to a central database, where they can be collated and ultimately exported as the candidate's mark.
- Figure 3 depicts a system comprising a scanning device 30, a central database 31 connected to the scanning device 30, a network administering device 32 connected to the central database 31 via a network 33 and a client computer 34 connected to the network administering device 32 via a network 33.
- the scanning device 30, central database 31 and network administering device 32 is known as the technical infrastructure.
- the technical infrastructure can be a single, combined device.
- the central database 31 holds a set of operational instructions (a computer program) to facilitate the purposeful running of the technical infrastructure. More than one scanning device 30 and more than one client computer can be connected to the central database 31 or network administration device 32 at any one time.
- the scanning device 30 scans the front page and the answer pages of the exam document. While scanning the answer pages in the document, the scanning device 30 performs a segmentation routine, where the answer pages are segmented according to the exam questions in the first section 23 and the written answers in the second section 26. Segmentation information data and the image data of the scanned pages are sent to the central database 31.
- the scanning device 30 or the central database 31 can compress the image data to reduce the storage space required to store the image data in the database.
- the segmentation information related to the image data can be correspondingly related to the compressed image data.
- the segment information data can comprise data that indicates the document scanned, the pages scanned, the location of the start and end of each segment, the question number associated with each segment, the exam candidate, the examination paper, the date and time of scanning and the filing reference of the original physical document.
- a client computer 34 can be connected to the technical inf astructure via a local network or the internet.
- An examiner using the client computer 34, can request an answer to a specified question to mark.
- the technical infrastructure identifies, from the segmentation information data, which parts of the images or compressed images correspond to the requested answers.
- the identified parts of the images or compressed images are sent to the client computer 34.
- the client computer 34 displays the answer to the question.
- the examiner can assign a mark, comment or special instruction to the answer and send this to the technical infrastructure.
- the nonadjacent parts will be joined together to display a single image.
- the parts are displayed in an order corresponding to the downstream direction.
- the downstream direction is the direction from the top of a page to the bottom of a page and then continuing to the top of the next page in the sequence of pages in the document.
- the displayed image of a segment can comprise the question number from the first section and the written answer given in the second section. Alternatively, the image of the segment comprises the answer from the second section only.
- ICR requires high resolution images.
- performing segmentation in real time has the benefit of not having to transfer and store large volumes of high resolution image data.
- a high resolution image is captured into buffer memory of the scanner and is segmented to obtained segmentation information data.
- the high resolution image data can then be compressed into low resolution image data and stored in the central database.
- FIG. 4 depicts the architecture of the scanner.
- the scanner comprises an optical scanning means 40 to convert the pages of the document into image data, a buffer memory 41 which temporarily stores the image data, a processor 42 and a network connection terminal 43.
- the image data of the scanned answer pages stored in the buffer memory is segmented by the processor 42.
- the resulting segment information data associated with the answer pages and the image data is sent to the technical infrastructure via the connection terminal 43.
- the pages of the document may be fed through the scanner using a high speed feeder. This allows a large number of pages to be scanned quickly.
- a problem that can arise with high speed scanning is the stretching of areas of a scanned image along the scanning direction.
- This problem can be solved by providing regularly spaced markers of a predetermined distance along the scanning direction of a document page.
- the predetermined distance between the markers is known by the processor.
- the document version indicator 14 on the front page of the document can be used to indicate to the processor 42 the predetermined distance between the markers.
- the processor 42 of the scanner measures the distance between the markers and if the distance between the markers is larger than the predetermined distance, then stretching of the image in that area has occurred.
- Image processing methods can be used to correct the stretching of the image.
- the distance between the markers is related to a number of pixels for a given resolution.
- a greater distance between the markers due to stretching can indicate the number of excess pixels that have arisen along the scanning direction due to stretching.
- the processor can remove the excess pixels along the scanning direction to correct the stretching of the image. For example, where two adjacent rows of pixels are identified, one is deleted.
- the image correction can also be carried out on a processor on computer, for example by a processor on the central database 31. This technique can be used in other methods involving scanning of documents.
- Figures 5A-D depict a process of segmentation of an image. This process can be carried out on an image stored in the buffer memory of a scanner or any other electronic storage means.
- step SI the image is checked for certain indicators (such as rule lines, margins, registration marks, page numbers etc), which indicate that the image is of a known exam page.
- step S2 it is determined whether or not the correct indications are found. If the indicators are not found, then the scanner rejects the page and alerts a user of the scanner, as shown by in step S3.
- step S4 the image of the page in the buffer memory removed and the process is ended. The segmentation process may then start again for another image.
- step S5 the processor begins searching the image of the first answer page of the document.
- the first section is searched for any marks that were not part of the predetermined format, i.e. any marks made by the exam candidate.
- the search begins from the top of the page towards the bottom of the page and then on to the next page in the sequence (the downstream direction).
- the indicators mentioned above can be used to indicate the top and bottom of each page.
- An embodiment can be configured to search the whole of the first area or just in the response areas.
- a segment can span more than one page. From the page indicators the downstream direction can be deduced and the segment continues from the bottom of a page to the top of the next page in the page sequence of the document.
- step S6 a mark is found.
- step S7 the location of the mark is recorded.
- Intelligent Character Recognition is applied to the mark in step S8.
- the processor can remove the imaged artefacts of the predetermined format area (for example, the shading). This improves the accuracy of the character recognition.
- the colour of the predetermined formatted areas (for example, the shading colour) can be removed optically by using a certain wavelength of light while scanning or by using optical filters.
- Character recognition can be carried out on all marks found in the first section, except where these marks may be undesired artefacts of the imagining process which can be electronically removed.
- Step S9 If the confidence level of a character recognised from a mark falls below the acceptable configurable confidence limit as determined in Step S9, then the process moves on to step S10.
- the character is flagged for intervention via a keying process which presents the character to a human operator for a decision on what the character represents.
- step S 10 the search for marks continues in the downstream direction from the location of the last mark found and then continues to step S6.
- step S9 If a character is recognised in step S9, its location, character and confidence rating is recorded. Where the predetermined formatted design includes other response areas in the immediate vicinity of each other, for instance on the same row, any characters that can be recognised from marks in such areas are concatenated and then the process proceeds to step S 1 1.
- step S I 1 the recognised characters are compared with a reference list of characters, which correspond to the nomenclature of the exam questions. If the recognised characters match with characters in the reference list, then the process moves on to step SI 3. In step SI 3, the recognised characters are recorded as the question number.
- step S10 the process proceeds to step S10.
- the character is flagged and checked by a user.
- steps S8 to S12 are carried out on any marks on all three boxes.
- all three marks must be recognised and compared to the list of question numbers.
- the three recognised characters must match with a three-digit question number in the reference list to proceed on to step S13. If the number derived from all three recognised characters is not in the reference list, then the process proceeds to step S 10.
- the characters are flagged and checked by a user. It is also possible that more than one character can be recognised in a single response area.
- step SI 4 the second section is searched for a blank area.
- the search is carried out in a location horizontally adjacent to the location of the question number or alternatively a location where there is a mark that provides suspicion of a question number.
- the parameters of the blank area search can be configured so that the search can be carried out in the vicinity of the location of the question number.
- the search parameters can be configured such that the blank area search is carried out two lines above and two lines below the horizontal location of the question number.
- the blank area can have a predetermined height and width, which can be a blank line across the page, and does not have any marks or written text within it.
- the processor allows for sloped writing when searching for a blank area and does not assume a level/horizontal division between answers.
- Step SI 5 If it is determined in Step SI 5 that a blank area is found, the process proceeds to step SI 8.
- step SI 8 the processor sets the location of the blank space as the start of the segment for the question number or the suspected question number.
- the processor can set the start of the segment directly from the location of the question number. For example, it is possible to configure the processor to set the start of the segment at a location horizontally adjacent to the question number or at a location horizontally adjacent to a predetermined vertical distance from the location of the question number. It is also possible to set the start of the segment at a horizontal line that is adjacent to or in the vicinity of the question number and is drawn across the second section by the exam candidate. Thus, steps S 14 to S 18 would not be required.
- step SI 6 the scanner processor searches a configurable distance in the upstream direction for a blank area. If one is not found, it searches downwards a configurable distance. A blank space that is located below the question number can be set as the start of the question number if there is no other question number adjacent to or in the vicinity of it in the first section. The location of the first blank area that is found is recorded and set as the start of the segment of the question number in step SI 8.
- the start of the segment can be set from the location of the question number in the manner described above in circumstances where a blank space cannot be located in step SI 6.
- step SI 9 the scanner processor then searches the first section, downstream from the location of the question number, for marks which correspond to another question number in the reference list or alternatively another mark.
- Steps S20 to S26 correspond to earlier steps S6 to S12 for finding and recognising a mark.
- the second recognisable character can indicate the location around which the end of the answer of the question is located and the beginning of the answer to another question is located.
- step S27 the second recognisable character is recorded as the subsequent question number attempted by the exam candidate.
- the subsequent question number can be any question number in the reference list and does not have to be the next question in the sequence of question numbers in the exam paper.
- step S28 the second section is searched for a blank area horizontally adjacent to or in the vicinity of the location of the subsequent question number.
- the processor can check for further blank areas above the blank area adjacent to the location of the subsequent question number. If further blank areas are found, the end of the segment will be set as the location of a blank area that is adjacent to the last detected marks in the second section. If further blank areas adjacent to the location of the subsequent question number are not found, the location of the end of segment will be set as the blank area adjacent to the location of the subsequent question number (S32).
- step S30 the scanner processor searches the document image in the upstream direction for a blank area.
- the processor then proceeds to step S32 where the location of the first blank area that is found is set as the end of the segment for the question number. This will also set the start of the segment for the subsequent question number. It is also possible to allow segments to overlap.
- the end of the segment for the question number and the start of the segment for the subsequent question number can be set directly from the location of the subsequent question number.
- the location of the subsequent question number can be used to determine the end of the segment for the question number and the start of the segment for the subsequent question number by plotting the segment coordinates to begin on the upper boundary of response area in the first section and extending through the horizontal plane into the second section.
- step S30 If a blank area is not found in step S30, the location of the subsequent question number can be used to determine the end of a segment in the manner described above. The segmentation process is then repeated for the remaining images of the pages in the document.
- Clip fixing is an exception process where characters that could not be recognised by the standard character recognition approaches and according to the configurable confidence levels in place are escalated to a suitably qualified user in order that a decision on what the mark represents can be made. Characters that have not corresponding segmentation areas, or those that are too thin to practically represent content, or segmentations that overlap others are escalated to a suitably qualified user in order that a decision on what the segment should represent.
- Marking from an image using the approach described herewith can be used as part of a quality control system. Such a system can actively intervene if marking quality is sub-optimal. This reduces the cost in the processing of examination documents because examiners whose work that is not of sufficient standard would not be permitted to continue.
- marking an exam at the level of individual questions brings quality control benefits as it reduces examiner bias that would often be present when the whole document paper is sent to be marked by an individual examiner.
- the present invention can be utilised in areas other than examination documents.
- the invention can be applied to any type of hand written or typed document.
- the present invention is advantageously applied to laboratory books where the images of the laboratory book may be segmented according to a project code, associated data files, the date or any other information.
- Other areas of application can include forms, notepads, project books, diaries, planners, address books, visitors books, time management, stock control, etc.
Landscapes
- Engineering & Computer Science (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Multimedia (AREA)
- Theoretical Computer Science (AREA)
- Computational Linguistics (AREA)
- Character Input (AREA)
- Character Discrimination (AREA)
Abstract
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| AU2010291063A AU2010291063B2 (en) | 2009-09-03 | 2010-09-01 | Method and apparatus for segmenting images |
| ZA2012/01732A ZA201201732B (en) | 2009-09-03 | 2012-03-09 | Method and apparatus for segmenting images |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| GB0915375.0 | 2009-09-03 | ||
| GB0915375A GB2473228A (en) | 2009-09-03 | 2009-09-03 | Segmenting Document Images |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2011027113A1 true WO2011027113A1 (fr) | 2011-03-10 |
Family
ID=41203128
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/GB2010/001661 Ceased WO2011027113A1 (fr) | 2009-09-03 | 2010-09-01 | Procédé et dispositif de segmentation d'images |
Country Status (4)
| Country | Link |
|---|---|
| AU (1) | AU2010291063B2 (fr) |
| GB (1) | GB2473228A (fr) |
| WO (1) | WO2011027113A1 (fr) |
| ZA (1) | ZA201201732B (fr) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112001416A (zh) * | 2020-07-17 | 2020-11-27 | 北京邮电大学 | 一种自适应答题纸序列纠正方法 |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20120200742A1 (en) | 2010-09-21 | 2012-08-09 | King Jim Co., Ltd. | Image Processing System and Imaging Object Used For Same |
| CN103179369A (zh) | 2010-09-21 | 2013-06-26 | 株式会社锦宫事务 | 摄像对象物、图像处理程序及图像处理方法 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP0081767A2 (fr) * | 1981-12-15 | 1983-06-22 | Kabushiki Kaisha Toshiba | Appareil de traitement de caractères et d'images |
| JPH1185895A (ja) * | 1997-09-03 | 1999-03-30 | Olympus Optical Co Ltd | コードパターン読取装置 |
| US6023342A (en) * | 1994-05-30 | 2000-02-08 | Ricoh Company, Ltd. | Image processing device which transfers scanned sheets to a printing region |
| US20030142358A1 (en) * | 2002-01-29 | 2003-07-31 | Bean Heather N. | Method and apparatus for automatic image capture device control |
| US6616038B1 (en) * | 1996-10-28 | 2003-09-09 | Francis Olschafskie | Selective text retrieval system |
| EP1833022A1 (fr) * | 2004-12-28 | 2007-09-12 | Fujitsu Ltd. | Dispositif de traitement d'image detectant la position d'un objet de traitement dans une image |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0757044A (ja) * | 1993-08-20 | 1995-03-03 | Oki Electric Ind Co Ltd | 文字認識装置 |
| JP3391626B2 (ja) * | 1996-03-27 | 2003-03-31 | 株式会社リコー | 画像形成装置 |
| US7324711B2 (en) * | 2004-02-26 | 2008-01-29 | Xerox Corporation | Method for automated image indexing and retrieval |
| US7742199B2 (en) * | 2004-10-06 | 2010-06-22 | Kabushiki Kaisha Toshiba | System and method for compressing and rotating image data |
-
2009
- 2009-09-03 GB GB0915375A patent/GB2473228A/en not_active Withdrawn
-
2010
- 2010-09-01 AU AU2010291063A patent/AU2010291063B2/en not_active Ceased
- 2010-09-01 WO PCT/GB2010/001661 patent/WO2011027113A1/fr not_active Ceased
-
2012
- 2012-03-09 ZA ZA2012/01732A patent/ZA201201732B/en unknown
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP0081767A2 (fr) * | 1981-12-15 | 1983-06-22 | Kabushiki Kaisha Toshiba | Appareil de traitement de caractères et d'images |
| US6023342A (en) * | 1994-05-30 | 2000-02-08 | Ricoh Company, Ltd. | Image processing device which transfers scanned sheets to a printing region |
| US6616038B1 (en) * | 1996-10-28 | 2003-09-09 | Francis Olschafskie | Selective text retrieval system |
| JPH1185895A (ja) * | 1997-09-03 | 1999-03-30 | Olympus Optical Co Ltd | コードパターン読取装置 |
| US20030142358A1 (en) * | 2002-01-29 | 2003-07-31 | Bean Heather N. | Method and apparatus for automatic image capture device control |
| EP1833022A1 (fr) * | 2004-12-28 | 2007-09-12 | Fujitsu Ltd. | Dispositif de traitement d'image detectant la position d'un objet de traitement dans une image |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN112001416A (zh) * | 2020-07-17 | 2020-11-27 | 北京邮电大学 | 一种自适应答题纸序列纠正方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| GB0915375D0 (en) | 2009-10-07 |
| GB2473228A (en) | 2011-03-09 |
| AU2010291063B2 (en) | 2015-01-15 |
| ZA201201732B (en) | 2012-11-28 |
| AU2010291063A1 (en) | 2012-04-05 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US10824801B2 (en) | Interactively predicting fields in a form | |
| US8107727B2 (en) | Document processing apparatus, document processing method, and computer program product | |
| US8155425B1 (en) | Automated check detection and image cropping | |
| US9785627B2 (en) | Automated form fill-in via form retrieval | |
| US8520224B2 (en) | Method of scanning to a field that covers a delimited area of a document repeatedly | |
| US8208737B1 (en) | Methods and systems for identifying captions in media material | |
| US20080218812A1 (en) | Metadata image processing | |
| US9396389B2 (en) | Techniques for detecting user-entered check marks | |
| JP4533273B2 (ja) | 画像処理装置及び画像処理方法、プログラム | |
| US8064703B2 (en) | Property record document data validation systems and methods | |
| US8605297B2 (en) | Method of scanning to a field that covers a delimited area of a document repeatedly | |
| US20190384971A1 (en) | System and method for optical character recognition | |
| CN109726369B (zh) | 一种基于标准文献的智能模板化题录技术实现方法 | |
| CN114596577B (zh) | 图像处理方法、装置、电子设备及存储介质 | |
| WO2011027113A1 (fr) | Procédé et dispositif de segmentation d'images | |
| US8687890B2 (en) | System and method for capturing relevant information from a printed document | |
| US7764923B2 (en) | Material processing apparatus and method for grading material | |
| JP2008022159A (ja) | 文書処理装置及び文書処理方法 | |
| KR20180126352A (ko) | 이미지로부터 텍스트 추출을 위한 딥러닝 기반 인식장치 | |
| US9152885B2 (en) | Image processing apparatus that groups objects within image | |
| JP4347675B2 (ja) | 帳票ocrプログラム、方法及び装置 | |
| JP5134383B2 (ja) | Ocr装置、証跡管理装置及び証跡管理システム | |
| KR100957508B1 (ko) | 광학 문자 인식 시스템 및 방법 | |
| KR20230118321A (ko) | 디지털화 된 참고서 제공 시스템 및 그 방법 | |
| Al-Barhamtoshy et al. | Universal metadata repository for document analysis and recognition |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 10759946 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2010291063 Country of ref document: AU |
|
| ENP | Entry into the national phase |
Ref document number: 2010291063 Country of ref document: AU Date of ref document: 20100901 Kind code of ref document: A |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 10759946 Country of ref document: EP Kind code of ref document: A1 |