WO2020083029A1 - 图形码识别方法、装置、终端及存储介质 - Google Patents
图形码识别方法、装置、终端及存储介质 Download PDFInfo
- Publication number
- WO2020083029A1 WO2020083029A1 PCT/CN2019/110359 CN2019110359W WO2020083029A1 WO 2020083029 A1 WO2020083029 A1 WO 2020083029A1 CN 2019110359 W CN2019110359 W CN 2019110359W WO 2020083029 A1 WO2020083029 A1 WO 2020083029A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- graphic code
- target
- recognition
- code
- graphic
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06K—GRAPHICAL DATA READING; PRESENTATION OF DATA; RECORD CARRIERS; HANDLING RECORD CARRIERS
- G06K7/00—Methods or arrangements for sensing record carriers, e.g. for reading patterns
- G06K7/10—Methods or arrangements for sensing record carriers, e.g. for reading patterns by electromagnetic radiation, e.g. optical sensing; by corpuscular radiation
- G06K7/14—Methods or arrangements for sensing record carriers, e.g. for reading patterns by electromagnetic radiation, e.g. optical sensing; by corpuscular radiation using light without selection of wavelength, e.g. sensing reflected white light
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06K—GRAPHICAL DATA READING; PRESENTATION OF DATA; RECORD CARRIERS; HANDLING RECORD CARRIERS
- G06K7/00—Methods or arrangements for sensing record carriers, e.g. for reading patterns
- G06K7/10—Methods or arrangements for sensing record carriers, e.g. for reading patterns by electromagnetic radiation, e.g. optical sensing; by corpuscular radiation
- G06K7/14—Methods or arrangements for sensing record carriers, e.g. for reading patterns by electromagnetic radiation, e.g. optical sensing; by corpuscular radiation using light without selection of wavelength, e.g. sensing reflected white light
- G06K7/1404—Methods for optical code recognition
- G06K7/1408—Methods for optical code recognition the method being specifically adapted for the type of code
- G06K7/1417—2D bar codes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06K—GRAPHICAL DATA READING; PRESENTATION OF DATA; RECORD CARRIERS; HANDLING RECORD CARRIERS
- G06K7/00—Methods or arrangements for sensing record carriers, e.g. for reading patterns
- G06K7/10—Methods or arrangements for sensing record carriers, e.g. for reading patterns by electromagnetic radiation, e.g. optical sensing; by corpuscular radiation
- G06K7/14—Methods or arrangements for sensing record carriers, e.g. for reading patterns by electromagnetic radiation, e.g. optical sensing; by corpuscular radiation using light without selection of wavelength, e.g. sensing reflected white light
- G06K7/1404—Methods for optical code recognition
- G06K7/1408—Methods for optical code recognition the method being specifically adapted for the type of code
- G06K7/1413—1D bar codes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06K—GRAPHICAL DATA READING; PRESENTATION OF DATA; RECORD CARRIERS; HANDLING RECORD CARRIERS
- G06K7/00—Methods or arrangements for sensing record carriers, e.g. for reading patterns
- G06K7/10—Methods or arrangements for sensing record carriers, e.g. for reading patterns by electromagnetic radiation, e.g. optical sensing; by corpuscular radiation
- G06K7/14—Methods or arrangements for sensing record carriers, e.g. for reading patterns by electromagnetic radiation, e.g. optical sensing; by corpuscular radiation using light without selection of wavelength, e.g. sensing reflected white light
- G06K7/1404—Methods for optical code recognition
- G06K7/1439—Methods for optical code recognition including a method step for retrieval of the optical code
- G06K7/1443—Methods for optical code recognition including a method step for retrieval of the optical code locating of the code in an image
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/20—Image preprocessing
- G06V10/22—Image preprocessing by selection of a specific region containing or referencing a pattern; Locating or processing of specific regions to guide the detection or recognition
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/048—Activation functions
Definitions
- the embodiments of the present application relate to the field of graphic code recognition, and in particular, to a graphic code recognition method, device, terminal, and storage medium.
- the graphic code is a carrier for carrying information or data.
- Common graphic codes include bar codes, two-dimensional codes, and so on.
- Long press recognition is a common way of identifying graphic codes.
- the terminal displays an image containing the two-dimensional code.
- the terminal displays an option to recognize the two-dimensional code.
- the terminal recognizes the QR code in the image and performs corresponding operations according to the recognition result. For example, when the recognition result is a web address, the terminal jumps to a webpage; when the recognition result is a payment link, the terminal displays a payment interface.
- the terminal can only recognize one of the graphic codes. If the recognized graphic code is not the graphic code desired by the user, the user needs to manually extract the desired graphic from the image The code is then recognized by the graphic code, which affects the recognition efficiency of the graphic code.
- a method, device, terminal and storage medium for identifying a graphic code are provided.
- a graphic code recognition method is executed by a terminal.
- the method includes:
- the recognition result of the target graphic code corresponding to the target graphic code is displayed.
- a graphic code recognition device which is installed in a terminal, and includes:
- An image display module for displaying a target image, the target image containing at least two graphic codes
- a position obtaining module configured to obtain the position information of the graphic code of each graphic code in the target image when receiving the graphic code recognition operation on the target image;
- a target determining module configured to determine a target graphic code indicated by the graphic code recognition operation according to the location information of the graphic code, the target graphic code belonging to the at least two graphic codes;
- the result display module is used to display the recognition result of the target graphic code corresponding to the target graphic code.
- a terminal includes a processor and one or more memories.
- the one or more memories store at least one computer-readable instruction, at least one program, code set, or computer-readable instruction set.
- a computer readable instruction, the at least one program, the code set or the computer readable instruction set is executed by the one or more processors to implement the graphic code recognition method as described in the above aspect.
- One or more computer-readable storage media where at least one computer-readable instruction, at least one program, code set, or computer-readable instruction set is stored in the storage medium, the at least one computer-readable instruction, the at least one section
- the program, the code set or the computer readable instruction set is executed by one or more processors to implement the graphic code recognition method as described in the above aspect.
- FIG. 1 shows a schematic diagram of an implementation environment provided by an embodiment of the present application
- Fig. 2 is a flowchart of a process of identifying a graphic code in the related art
- FIG. 3 is a flowchart of a graphic code recognition process in an embodiment of the present application.
- FIG. 4 shows a flowchart of a method for identifying a graphic code provided by an embodiment of the present application
- FIG. 5 shows a flowchart of a method for identifying a graphic code in a long-press identification scenario provided by an embodiment of the present application
- FIG. 6 is an interface schematic diagram of the implementation process of the graphic code recognition method shown in FIG. 5;
- FIG. 7 is a schematic diagram of implementation when determining the distance between the graphic code and the target recognition position
- FIG. 8 shows a flowchart of a method for recognizing a graphic code in a scan code recognition scenario provided by an embodiment of the present application
- FIG. 9 is a schematic view of the interface of the implementation process of the graphic code recognition method shown in FIG. 8;
- FIG. 10 is a schematic structural diagram of a target retrieval model provided by an embodiment of the present application.
- FIG. 11 is a schematic structural diagram of a first residual block provided by an embodiment of the present application.
- FIG. 12 is a schematic structural diagram of a second residual block provided by an embodiment of the present application.
- FIG. 13 shows a flowchart of a method for identifying a graphic code provided by another embodiment of the present application.
- FIG. 14 shows a flowchart of a method for identifying a graphic code provided by another embodiment of the present application.
- FIG. 15 shows a block diagram of a graphic code recognition device provided by an embodiment of the present application.
- FIG. 16 shows a schematic structural diagram of a terminal provided by an embodiment of the present application.
- FIG. 1 illustrates a schematic diagram of an implementation environment provided by an embodiment of the present application.
- the implementation environment includes a terminal 120 and a server 140.
- the terminal 120 is an electronic device with a graphic code recognition function, and the electronic device may be a smart phone, a tablet computer, a personal computer, or the like.
- the terminal 120 is a smartphone as an example.
- the graphic code recognition function may be provided by the operating system of the electronic device, or may be provided by a third-party application installed in the electronic device.
- the third-party application may be a payment application, instant messaging application, or shopping application Programs, video playback applications, browser applications, etc., which are not limited in the embodiments of the present application.
- the graphic code recognized by the terminal 120 may be a picture, such as a picture received in an instant communication application; or an image collected through a camera component, such as an image scanned after the application's graphic code scanning function is enabled.
- the application examples do not limit this.
- the terminal 120 in the embodiment of the present application also has a target detection function.
- the terminal 120 can recognize the position and / or type of each graphic code in the image, so as to subsequently extract the graphic code from the image according to the position of the graphic code, According to the type of the graphic code, the corresponding decoder is used to recognize the graphic code, and the result of the graphic code recognition is obtained.
- the target detection function is implemented by using a target detection model obtained based on deep learning training.
- the target detection model is used to position information and position confidence of each graphic code in the input image according to the input image, And / or, the type of each graphic code and the type confidence.
- the terminal 120 and the server 140 are connected through a wired or wireless network.
- the server 140 is a server cluster or cloud computing center composed of one server and several servers.
- the server 140 is a resource server for providing webpage resources.
- the terminal 120 receives the graphic code recognition operation and completes the target graphic code recognition to obtain the graphic code recognition result
- the web page resource is obtained from the corresponding server 140 according to the graphic code recognition result To display the webpage resource.
- the wireless or wired network described above uses standard communication technologies and / or protocols.
- the network is usually the Internet, but it can also be any network, including but not limited to a local area network (Local Area Network, LAN), a metropolitan area network (Metropolitan Area Network, MAN), a wide area network (Wide Area Network, WAN), mobile, wired or wireless Any combination of networks, private networks or virtual private networks).
- technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged over the network.
- HTML Hyper Text Mark-up Language
- XML Extensible Markup Language
- SSL Secure Socket Layer
- TLS Transport Layer Security
- VPN Virtual Private Network
- IPsec Internet Protocol Security
- the graphic code recognition methods provided by the embodiments of the present application are executed by the terminal 120 in FIG. 1, that is, the terminal 120 realizes the graphic code recognition locally to obtain the graphic code recognition result; of course, in other possible implementation manners, the terminal 120 also
- the image to be recognized can be uploaded to the background server, and the background server performs graphic code recognition on the image and feeds back the graphic code recognition result for the terminal 120 to display, which is not limited in the embodiments of the present application.
- the graphic code recognition method provided in the embodiment of the present application can be used to realize the recognition of a multi-code of one image, that is, to recognize a specified graphic code in an image of multiple graphic codes (referring to an image containing multiple graphic codes). Code to identify the scene and long press to identify the scene.
- the graphic code recognition methods in different application scenarios are described below.
- the terminal may simultaneously collect images containing multiple graphic codes. For example, when scanning for payment, the terminal collects an image that contains both the payment code A (corresponding to the A payment application) and the collection code B (corresponding to the B payment application). In this scenario, the terminal first recognizes the location of each graphic code in the image through the graphic code recognition method, and then determines the target graphic code that meets the user's recognition intention according to the location of the graphic code, thereby identifying the target graphic code and obtaining the corresponding The recognition result of the graphic code.
- the terminal When users use the application, they often see pictures containing graphic codes. For example, when using an instant messaging application, an image containing a graphic code sent by another user is received, and when using a browser application, the displayed webpage image contains a graphic code. At this time, the user can trigger the terminal to recognize the graphic code in the picture by long pressing the picture. In this scenario, the terminal first recognizes the position of the graphic code of each graphic code in the picture through the graphic code recognition method, and then determines the target graphic code that the user desires to recognize based on the graphic code position and the long-press position, so that the target graphic code is identified Recognize, get the corresponding graphic code recognition result.
- the above graphic code recognition method can also be used in other recognition scenarios involving one image and multiple codes, which are not limited in this embodiment of the present application.
- FIG. 2 shows a flowchart of the graphic code recognition process in the related art.
- the terminal After acquiring the image to be recognized (original image), due to the size sensitivity of the graphic code recognition, the terminal needs to downsample the original image by using pyramid sampling to obtain downsampled images of different sizes. As shown in FIG. 2, after pyramid sampling, the original image 21, 1/4 down-sampling image 22 and 1/16 down-sampling image 23 are obtained.
- the terminal For each of the obtained down-sampled images, in order to facilitate subsequent decoding of the graphics code, the terminal performs image binarization processing to obtain a binarized image containing only black and white and two colors.
- the image binarization processing can be mixed Binarization, fast window binarization and adaptive binarization. As shown in FIG. 2, the terminal binarizes the original image 21 and obtains a binarized image 24.
- the terminal tries various decoders to decode the obtained binary map, for example, one-dimensional code decoder is used for one-dimensional code decoding, and two-dimensional code decoder is used for two-dimensional code decoding. If decoding fails with the current decoder, try the next decoder until the decoding is successful. Among them, when decoding by the decoder, it is necessary to traverse each pixel in the binarized image to traverse. As shown in FIG. 2, in the process of decoding the binarized graph by the two-dimensional code decoder, the terminal first recognizes the three vertices 25 of the two-dimensional code, and then decodes the two-dimensional code.
- the terminal can only decode one of the graphic codes (usually the first decoded graphic code in the image). If the user expects to obtain the recognition result of the specified graphic code in the image, the above method will not be achieved.
- the terminal After acquiring the image 31 to be recognized, the terminal first performs target detection on the image to be recognized 31, and obtains the location and type of the graphic code of each graphic code in the image 31 to be recognized, Thus, each graphic code is extracted from the image to be recognized 31 according to the position of the graphic code.
- the terminal samples it based on the principle of small image upsampling and large image downsampling to obtain a sampling map 33, and performs image binarization processing on the sampling map 33 to obtain a binary map 34.
- the terminal decodes according to the graphic code type and uses a graphic code decoder corresponding to the graphic code type to obtain a graphic code recognition result corresponding to the binarized image 34.
- the terminal determines the target graphic code that the user desires to recognize, thereby displaying the recognition result corresponding to the target graphic code, and realizing the recognition of the specified graphic code in the multi-graphic code graphic.
- FIG. 4 shows a flowchart of a method for identifying a graphic code provided by an embodiment of the present application.
- This embodiment is exemplified by the method applied to the terminal 120 in FIG. 1.
- the method may include the following steps:
- step 401 a target image is displayed, and the target image includes at least two graphic codes.
- the target images in different application scenarios are different.
- the target image in the scan code recognition scene, is the image displayed in the terminal's viewfinder frame, that is, the image collected by the terminal through the camera component in real time; in the long press recognition scene, the target image is the picture displayed by the terminal , Such as pictures received in instant messaging applications.
- the target image in the embodiment of the present application contains at least two graphic codes, and the graphic code types of the at least two graphic codes may be the same or different, wherein the graphic code types include one-dimensional codes (also called barcodes) and two At least one of the dimension codes, and the two-dimensional code includes at least one of a dot-shaped two-dimensional code, a ring-shaped two-dimensional code, and a radial two-dimensional code.
- the graphic code types include one-dimensional codes (also called barcodes) and two At least one of the dimension codes
- the two-dimensional code includes at least one of a dot-shaped two-dimensional code, a ring-shaped two-dimensional code, and a radial two-dimensional code.
- Step 402 when receiving a graphic code recognition operation on the target image, acquiring graphic code position information of each graphic code in the target image.
- the graphic code recognition operation is used to instruct the recognition of the graphic code specified in the target image.
- the operation of identifying the image code of the target image is different.
- the graphic code recognition operation in a scan code recognition scene, is a shooting operation; in a long press recognition scene, the graphic code recognition operation is a long press operation on a picture.
- Graphic code position information is used to uniquely represent the position of the graphic code in the target image.
- the position information of the graphic code includes the coordinates of the preset mark point in the graphic code.
- the position information of the graphic code includes the coordinates of the top left corner of the graphic code, or the coordinates of the center point of the graphic code.
- the position information of the graphic code further includes size information of the graphic code. For example, the height and width information of the graphic code.
- the terminal in addition to acquiring the position information of the graphic codes of the respective graphic codes, the terminal also obtains the graphic code types of the respective graphic codes, so as to subsequently use the corresponding decoder to decode according to the graphic code types, and no more need to try Various types of decoders.
- Step 403 According to the position information of the graphic code, determine the target graphic code indicated by the graphic code recognition operation, and the target graphic code belongs to at least two graphic codes.
- this step includes the following steps.
- the target recognition position is expressed in the form of coordinates.
- the terminal calculates the distance between the target recognition position and each graphic code based on the target recognition position and the graphic code position information, and determines the graphic code corresponding to the shortest distance as the target graphic code.
- Step 405 Display the recognition result of the target graphic code corresponding to the target graphic code.
- the terminal recognizes each graphic code to obtain at least two graphic code recognition results, and finally displays the target graphic code recognition result corresponding to the target graphic code;
- the terminal only recognizes the determined target graphic code, and displays the obtained recognition result of the target graphic code.
- the terminal obtains the size of the graphic code and detects whether the size is larger than a preset size (such as 300px ⁇ 300px), and if it is larger, the graphic code is determined It is a large image, and downsample the large image, so that the size of the graphic code after the downsampling is the preset size; if it is smaller, the graphic code is determined to be a small image, and the small image is upsampled (super resolution sampling can be used ) To make the size of the graphic code after upsampling the preset size.
- a preset size such as 300px ⁇ 300px
- the terminal can perform targeted sampling according to the size of the graphic code without using the pyramid sampling method for multiple sampling, thereby reducing the amount of data processing in the recognition process.
- the terminal needs to sample the image to be recognized multiple times, while in FIG. 3, the terminal only needs to upsample the image.
- the terminal After sampling the graphic code, the terminal further performs image binarization processing on the graphic code. Compared with the related art, the entire image needs to be binarized. In this embodiment, the terminal only needs to binarize the extracted graphic code, thereby reducing the amount of data processing in the binarization process.
- the terminal decodes the binarized graphic code. Since the graphic code type of the graphic code can be obtained during target detection, the terminal can use a decoder to decode it without having to try various decoders, thereby reducing the amount of data processing in the decoding process. Schematically, in FIG. 2, the terminal needs to try three types of decoders to decode, while in FIG. 3, the terminal can directly decode the graphic code through the two-dimensional code decoder.
- the terminal when receiving a graphic code recognition operation on a target image containing at least two graphic codes, first obtain the graphic code position information of each graphic code in the target image, and then according to the graphic code position Information, determine the target graphic code indicated by the graphic code recognition operation, so as to display the target graphic code recognition result corresponding to the target graphic code; with the help of the graphic code position recognition mechanism, the terminal can simultaneously recognize multiple graphic codes in the same image, Therefore, according to the position of each graphic code, a target graphic code that meets the user's recognition intention is determined, and then the recognition result of the target graphic code is returned, which improves the recognition efficiency of the graphic code and solves the related art.
- the image contains at least two graphics
- the user needs to manually cut out the graphic code desired to be recognized from the image and then perform graphic code recognition, which leads to the problem of low recognition efficiency of the graphic code.
- the terminal determines the target graphic code from at least two graphic codes in different ways.
- the following two embodiments are used to explain the process of determining the target graphic code in the scenario of long-press recognition and scan code recognition, respectively.
- FIG. 5 shows a flowchart of a method for identifying a graphic code provided by another embodiment of the present application. This embodiment is described by taking the method applied to the long-press recognition scene as an example. The method may include the following steps:
- step 501 a target image is displayed, and the target image includes at least two graphic codes.
- the target image is a picture displayed by the terminal.
- the terminal displays a target image 61 that includes a radial two-dimensional code 62 and a dot two-dimensional code 63.
- Step 502 when receiving a graphic code recognition operation on the target image, acquiring graphic code position information of each graphic code in the target image.
- the graphic code recognition operation is a trigger operation on the target image.
- the terminal when the terminal is a mobile terminal with a touch function, when a long press operation on the target image is received, the terminal displays several operation options, and when a selection operation on a graphic code recognition option is received , The terminal determines to receive the graphic code recognition operation.
- the terminal after receiving the long-press operation on the target image 61, the terminal displays an operation option menu, and when receiving the selection operation of the graphic code recognition option 64, determines that the graphic code recognition operation is received .
- the terminal when the terminal is a terminal such as a PC that includes an external input device (such as a mouse), when receiving a click operation on the target image (executed by the external input device), the terminal displays and displays several operation options, When receiving the selection operation of the graphic code recognition option, the terminal determines that the graphic code recognition operation is received.
- an external input device such as a mouse
- Step 503 Determine the trigger position corresponding to the graphic code recognition operation as the target recognition position.
- the terminal determines the trigger position (such as the long-press position) corresponding to the graphic code recognition operation Identify the location for the target
- the terminal determines the long-press position where the long-press signal is received as the target recognition position 65.
- the terminal acquires the coordinates of the target recognition position on the target image.
- the coordinates of the target recognition position acquired by the terminal are (x pos , y pos ).
- the terminal displays a prompt message instructing the user to long press the graphic code at a different location for recognition, which is not limited in this embodiment.
- Step 504 Determine the distance between the target recognition position and each graphic code according to the position information of the trigger position and the graphic code position information of each graphic code.
- the terminal calculates the code center of the graphic code according to the graphic code position information of each graphic code, and calculates the target recognition position and each graphic code according to the coordinates of the trigger position and the code center the distance between.
- the terminal calculates the code center coordinate of the graphic code according to the vertex coordinates of each vertex in the graphic code position information; or, the terminal according to the vertex coordinates and graphic code size information of at least one vertex in the graphic code position information , Calculate the code center of the graphic code.
- This application does not limit the method of calculating the center coordinates of the code.
- the terminal acquires the coordinates of the first code center 621 of the radial two-dimensional code 62 as (x 1 , y 1 ), and the coordinates of the second code center 631 of the point-like two-dimensional code 63 Is (x 2 , y 2 ), the calculated distance s 1 between the target recognition position 65 and the radial two-dimensional code 62 is
- Step 505 Determine the graphic code corresponding to the shortest distance as the target graphic code.
- the terminal determines the graphic code corresponding to the shortest distance as the target graphic code, that is, the graphic code closest to the target recognition position is determined as the target graphic code.
- the terminal determines the dot-shaped two-dimensional code 63 as the target graphic code.
- Step 506 Display the recognition result of the target graphic code corresponding to the target graphic code.
- the terminal performs graphic code recognition on the dotted two-dimensional code 63, and the obtained target graphic code recognition result is a game download link, so that the game download interface 66 is jump-displayed according to the game download link.
- the terminal determines the trigger position corresponding to the graphic code recognition operation as the target recognition position, and determines the target graphic code indicated by the user by calculating the distance between the target recognition position and each graphic code, so that in the long-press recognition scene To realize the recognition of the graphic code at the long press position in the target image.
- FIG. 8 shows a flowchart of a method for identifying a graphic code provided by another embodiment of the present application.
- the method is applied to scan code identification scenarios as an example for illustration.
- the method may include the following steps:
- step 801 a target image is displayed, and the target image includes at least two graphic codes.
- the target image is the image displayed in the framing frame.
- the target image collected by the camera is displayed in the framing frame 91, and the target image includes a radial two-dimensional code 92 and a dot two-dimensional code 93 .
- Step 802 When receiving a graphic code recognition operation on the target image, obtain graphic code position information of each graphic code in the target image.
- the graphic code recognition operation is a shooting operation on the target image.
- a shooting control is displayed on the terminal interface, and when a click operation on the shooting control is received, the terminal determines that a graphic code recognition operation is received.
- the terminal determines that the graphic code recognition operation is received.
- a duration threshold such as 0.5s
- step 803 the corresponding position of the center of the framing frame in the target image is determined as the target recognition position.
- the user usually moves the terminal so that the target graphic code to be recognized is located at or near the center of the framing frame. Therefore, in a possible implementation manner, the terminal places the center of the framing frame at the corresponding position in the target image Determine the target recognition position.
- the terminal determines the coordinates of the target recognition position as (x pos , y pos ).
- Step 804 Determine the distance between the target recognition position and each graphic code according to the position information of the center of the viewfinder and the graphic code position information of each graphic code.
- the terminal calculates the code center of the graphic code based on the graphic code position information of each graphic code, and calculates the distance between the target recognition position and each graphic code based on the coordinates of the center of the framing frame and the code center.
- the terminal calculates the code center of the graphic code based on the graphic code position information of each graphic code, and calculates the distance between the target recognition position and each graphic code based on the coordinates of the center of the framing frame and the code center.
- Step 805 Determine the graphic code corresponding to the shortest distance as the target graphic code.
- the terminal determines the graphic code corresponding to the shortest distance as the target graphic code, that is, the graphic code closest to the center of the viewfinder frame is determined as the target graphic code.
- the terminal determines the radial two-dimensional code 92 as the target Graphic code.
- Step 806 Display the recognition result of the target graphic code corresponding to the target graphic code.
- the terminal performs graphic code recognition on the radial two-dimensional code 92, and the obtained target graphic code recognition result is applet jump information, so that the applet interface is displayed according to the applet jump information. 94.
- Mini Program is an application program that can be used without downloading and installing. Developers can develop corresponding applets for terminal applications. Applets can be embedded in terminal applications as sub-applications. By running applets in applications, they can provide users with more diversified services.
- the terminal determines the center of the framing frame as the target recognition position, and determines the target graphic code that the user desires to scan by calculating the distance between the target recognition position and each graphic code, thereby realizing the framing under the scan code recognition scene The identification of the specified graphic code among the multiple graphic codes in the frame.
- the terminal stores a pre-trained target detection model, which is obtained through deep learning training, is used to identify the graphic code in the image, and outputs the position information of the graphic code in the image .
- the terminal may include the following steps when acquiring the position information of the graphic code of each graphic code in the target image.
- the target detection model is obtained through deep learning training.
- the predicted position information includes the coordinates of a designated marker point in the graphic code, and the designated marker point may be a vertex or a code center.
- the target detection model includes i serial residual networks and a hole convolution network.
- the target detection model includes three series of residual networks and a hollow convolutional network, namely a first residual network 1010, a second residual network 1020, and a third residual network 1030 And the empty convolutional network 1040.
- each residual network contains a down-sampling block and j first residual blocks.
- the downsampling block is used to obtain the image features after downsampling the input content.
- the first residual block is the basic block in the residual network, and usually includes a residual branch and a short circuit branch.
- the residual branch is used to analyze the residual
- the input of the block is nonlinearly transformed, and the short-circuit branch is used to perform an identity transformation or a linear transformation on the input of the residual block.
- the number of first residual blocks included in each residual network may be the same or different
- the first residual network 1010 includes 3 first residual blocks
- the second residual network 1020 includes 7 first residual blocks
- the third residual network 1030 includes 3 first residual blocks.
- the first residual block may be a conventional residual block or a bottleneck residual block (Bottleneck Residual Block).
- the first residual block adopts a conventional residual block or a bottleneck residual block
- the characteristics of the input first residual block are convolved in the convolutional layers of each layer, and the amount of residual network parameters and calculations are concentrated in the convolution In the layer.
- part of the convolution in the first residual block is replaced by depth depthwise convolution to ensure accurate recognition Under the premise of the rate, reduce the size of the residual network and increase the processing speed of the residual network.
- the bottleneck residual block 1110 contains three convolutional layers, where the first convolutional layer contains several 1 ⁇ 1 convolution kernels and the second convolutional layer contains several 3 ⁇ 3 convolution kernels, the third convolution layer contains several 1 ⁇ 1 convolution kernels.
- Each convolutional layer contains a normalization (Batch Normalization, BN) layer, and the first convolutional layer and the final output both contain an activation layer (Rectified Linear Units, ReLU).
- the first residual block 1120 is obtained.
- Convolution of holes is also called expansion convolution, which is a kind of convolution method in which holes are injected between convolution kernels.
- hollow convolution introduces a hyper-parameter called "dilation rate", which defines the interval between values when the convolution kernel processes data.
- the receptive field can be expanded to achieve more accurate target detection .
- the receptive field is the size of the area where the pixels on the feature map output from the hidden layer in the neural network are mapped on the original image. The larger the receptive field of the pixel on the original image, the larger the range of the mapped original image. It means that it may contain more global and higher-level features.
- k convolutional networks contain k second residual blocks. Schematically, as shown in FIG. 10, the hollow convolutional network contains three second residual blocks.
- dilation convolution is applied in the second residual block to expand the receptive field, and, in order to avoid the lower layer features being directly used as upper layer features, the upper layer features cannot obtain a higher semantic level and visual receptive field,
- the short-circuit branch of the second residual block also contains a convolution transform.
- the bottleneck residual block 1210 contains three convolutional layers, where the first convolutional layer contains several 1 ⁇ 1 convolution kernels, and the second convolutional layer contains several A 3 ⁇ 3 convolution kernel, the third convolution layer contains several 1 ⁇ 1 convolution kernels, each convolution layer contains a BN layer, and both the first convolution layer and the final output contain ReLU.
- the 3 ⁇ 3 convolution kernel in the second convolution layer is transformed into a hollow dilated convolution kernel, and a convolution containing several 1 ⁇ 1 convolution kernels is added to the short circuit branch Layer, the second residual block 1220 is finally obtained.
- the output of the last residual block in the residual network and the output of each second residual block in the hollow convolution network are input to the output network, and the output network performs classification and regression. To improve the accuracy of subsequent classification results.
- each of the second residual network 1020 at the end of the first residual block, the third residual network 1030 at the end of the first residual block, and the hollow convolutional network 1040 The output of the difference block is input to the output network.
- the terminal predicts the position confidence corresponding to the position information corresponding to each graphic code, and determines the predicted position information of the graphic code whose position confidence is greater than the confidence threshold (such as 90%) as the graphic of each graphic code Code location information.
- the confidence threshold such as 90%
- the target detection model in addition to predicting the position of the graphic code, can also predict the graphic code type of the graphic code.
- the terminal uses a graphic code decoder corresponding to the graphic code type Decode to improve decoding efficiency.
- the terminal acquires the prediction type and type confidence of the graphic code output by the target detection model, and determines the graphic code type of each graphic code according to the prediction type and type confidence of the graphic code.
- the graphic code type includes a At least one of a dimension code and a two-dimensional code.
- the terminal determines the prediction type of the graphic code whose type confidence is greater than the confidence threshold (such as 90%) according to the type confidence corresponding to the prediction type of each graphic code as the graphic code type of each graphic code.
- the confidence threshold such as 90%
- step 405 the following steps are also included before step 405:
- Step 404 Recognize the target graphic code by the target decoder corresponding to the target graphic code type to obtain the target graphic code recognition result.
- the target graphic code type is the graphic code type corresponding to the target graphic code; or
- the decoder corresponding to each graphic code type performs graphic code recognition on each graphic code to obtain a graphic code recognition result corresponding to each graphic code; and determines the graphic code recognition result corresponding to the target graphic code as the target graphic code recognition result.
- the terminal extracts the target graphic code from the target image, and according to the target graphic code type corresponding to the target graphic code, uses the target decoder to perform the graphic code Recognition, so as to obtain the recognition result of the target graphic code corresponding to the target graphic code.
- the terminal does not need to perform graphic code recognition, thereby reducing the amount of data processing when the terminal recognizes the graphic code.
- the terminal extracts each graphic code from the target image according to the location information of the graphic code output by the target detection model, and uses a corresponding decoder for each graphic code according to the type of the graphic code corresponding to each graphic code.
- the graphic code performs graphic code recognition, thereby obtaining a graphic code recognition result corresponding to each graphic code. Further, the terminal determines the recognition result of the graphic code corresponding to the target graphic code as the recognition result of the target graphic code for subsequent display.
- the terminal uses corresponding decoders for graphic code recognition according to the type of the recognized graphic code, which improves the recognition efficiency, and Reduced the amount of data processing during recognition.
- the terminal may also determine the target graphic code that the user desires to recognize according to the graphic code recognition result corresponding to each graphic code in the target image.
- step 403 may include the following steps.
- Step 403A Perform graphic code recognition on each graphic code according to the position information of the graphic code, and obtain at least two graphic code recognition results.
- the terminal extracts each graphic code from the target image, and performs graphic code recognition on each graphic code to obtain a graphic code recognition result corresponding to each graphic code.
- Step 403B Determine the target application corresponding to the graphic code recognition operation.
- the target application program is an application program that receives a graphic code recognition operation.
- the instant messaging application A is determined as the target application.
- step 403C if the recognition result of the graphic code belongs to the recognition result supported by the target application program, the graphic code corresponding to the recognition result of the graphic code is determined as the target graphic code.
- the types of recognition results supported by different application programs in the terminal are different, and each application program corresponds to its own recognition result list, and the recognition result list contains the supported recognition result types.
- the target application when the recognition result of the graphic code is a recognition result supported by the target application, the target application can parse the recognition result of the graphic code; otherwise, the target application cannot parse the recognition result of the graphic code.
- the identification result that it supports to display is the payment page of the payment application B.
- the terminal determines that the recognition result of the graphic code corresponding to the first graphic code is the recognition result of the target graphic code, and determines the first graphic code as the target graphic code.
- the recognition result list includes a recognition result keyword. Based on the recognition result list, the terminal detects whether the recognition result of the graphic code includes a recognition result keyword, and if so, determines that the recognition result of the graphic code belongs to a recognition result supported by the current application.
- Subsequent terminals only display the recognition result of the target graphic code, and will not display the recognition result of other graphic codes.
- the terminal determines that the payment application supported by the application is used when scanning the code, and the payment application
- the recognition result corresponding to the payment QR code is displayed to facilitate the user to make quick payment in the current application, and to avoid the problem that payment cannot be caused because the current application cannot display the payment pages of other payment applications.
- FIG. 15 illustrates a block diagram of a graphic code recognition apparatus provided by an embodiment of the present application.
- the device may be the terminal 120 in the implementation environment shown in FIG. 1 or may be provided on the terminal 120.
- the device includes various modules or units, and each module or unit may be implemented in whole or in part by software, hardware, or a combination thereof.
- the device may include:
- the image display module 1501 is configured to display a target image, and the target image includes at least two graphic codes;
- the position obtaining module 1502 is configured to obtain the position information of the graphic code of each graphic code in the target image when receiving the graphic code recognition operation on the target image;
- the target determining module 1503 is configured to determine a target graphic code indicated by the graphic code recognition operation according to the location information of the graphic code, and the target graphic code belongs to the at least two graphic codes;
- the result display module 1504 is used to display the recognition result of the target graphic code corresponding to the target graphic code.
- the target determination module 1503 includes:
- a first determining unit configured to determine a target recognition position indicated by the graphic code recognition operation
- the second determining unit is configured to determine the target graphic code according to the target identification position and the graphic code position information.
- the target image is a picture
- the graphic code recognition operation is a trigger operation on the picture
- the first determining unit is configured to determine the trigger position corresponding to the graphic code recognition operation as the target recognition position
- the second determining unit is configured to determine the distance between the target recognition position and each graphic code according to the position information of the trigger position and the graphic code position information of each graphic code; the graphic corresponding to the shortest distance The code is determined to be the target graphic code.
- the target image is an image displayed in a framing frame
- the graphic code recognition operation is a shooting operation on the target image
- the first determining unit is configured to determine the corresponding position of the center of the framing frame in the target image as the target recognition position;
- the second determining unit is configured to determine the distance between the target recognition position and each graphic code according to the position information of the center of the framing frame and the graphic code position information of each graphic code; the shortest distance corresponds to The graphic code is determined as the target graphic code.
- the location acquisition module 1502 includes:
- the input unit is used to input the target image into the target detection model to obtain the predicted position information and position confidence of the graphical code of each graphical code;
- a third determining unit is configured to determine the position information of the graphic code of each graphic code based on the predicted position information of the graphic code and the position confidence level.
- the target detection model includes i concatenated residual networks and a hole convolution network, wherein each residual network contains a down-sampling block and j first residual blocks , And the first residual block includes deep depthwise convolution; the hollow convolutional network includes k second residual blocks, and the second residual block includes hollow dilated convolutions, i, j, k is an integer greater than or equal to 2.
- the device further includes:
- a type acquisition module used to obtain the prediction type and type confidence of each graphic code output by the target detection model
- a type determining module configured to determine the type of the graphic code of each graphic code according to the prediction type of the graphic code and the type confidence, the graphic code type including at least one of a one-dimensional code and a two-dimensional code;
- the device also includes:
- the first decoding module is configured to perform graphic code recognition on the target graphic code through a target decoder corresponding to the target graphic code type to obtain the target graphic code recognition result, and the target graphic code type corresponds to the target graphic code Type of graphic code;
- the second decoding module is used to identify the graphics code of each graphics code through the decoder corresponding to the graphics code type of each graphics code to obtain the graphics code recognition result corresponding to each graphics code; the graphics corresponding to the target graphics code
- the code recognition result is determined as the target graphic code recognition result.
- the target determination module 1503 further includes:
- a result recognition unit configured to perform graphic code recognition on each graphic code according to the graphic code position information, and obtain at least two graphic code recognition results;
- a fourth determining unit configured to determine a target application corresponding to the graphic code recognition operation
- a fifth determining unit is configured to determine, if the graphic code recognition result belongs to a recognition result supported by the target application program, the graphic code corresponding to the graphic code recognition result as the target graphic code.
- the terminal when receiving a graphic code recognition operation on a target image containing at least two graphic codes, first obtain the graphic code position information of each graphic code in the target image, and then according to the graphic code position Information, determine the target graphic code indicated by the graphic code recognition operation, so as to display the target graphic code recognition result corresponding to the target graphic code; with the help of the graphic code position recognition mechanism, the terminal can simultaneously recognize multiple graphic codes in the same image, Therefore, according to the position of each graphic code, a target graphic code that meets the user's recognition intention is determined, and then the recognition result of the target graphic code is returned, which improves the recognition efficiency of the graphic code and solves the related art.
- the user needs to manually cut out the graphic code desired to be recognized from the image and then perform graphic code recognition, which leads to the problem of low recognition efficiency of the graphic code.
- FIG. 16 shows a schematic structural diagram of a terminal provided by an embodiment of the present application.
- the terminal may be implemented as the terminal 120 in the implementation environment shown in FIG. 1 to implement the account recommendation method provided in the above embodiment. Specifically:
- the terminal includes: a processor 1601 and a memory 1602.
- the processor 1601 may include one or more processing cores, such as a 4-core processor, an 8-core processor, and so on.
- the processor 1601 may use at least one hardware form of DSP (Digital Signal Processing, digital signal processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). achieve.
- the processor 1601 may also include a main processor and a co-processor.
- the main processor is a processor for processing data in a wake-up state, also known as a CPU (Central Processing Unit); the co-processor is A low-power processor for processing data in the standby state.
- the processor 1601 may be integrated with a GPU (Graphics Processing Unit, image processor), and the GPU is used to render and draw content that needs to be displayed on the display screen.
- the processor 1601 may further include an AI (Artificial Intelligence, Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
- AI Artificial Intelligence, Artificial Intelligence
- the memory 1602 may include one or more computer-readable storage media, which may be tangible and non-transitory.
- the memory 1602 may also include high-speed random access memory, and non-volatile memory, such as one or more magnetic disk storage devices, flash memory storage devices.
- non-transitory computer-readable storage medium in the memory 1602 is used to store at least one computer-readable instruction that is executed by the processor 1601 to implement the application provided in the present application The recognition method of the graphic code.
- the terminal may optionally include a peripheral device interface 1603 and at least one peripheral device.
- the peripheral device includes at least one of a radio frequency circuit 1604, a touch display screen 1605, a camera 1606, an audio circuit 1607, a positioning component 1608, and a power supply 1609.
- the peripheral device interface 1603 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1601 and the memory 1602.
- the processor 1601, the memory 1602, and the peripheral device interface 1603 are integrated on the same chip or circuit board; in some other embodiments, any one of the processor 1601, the memory 1602, and the peripheral device interface 1603 or Both can be implemented on a separate chip or circuit board, which is not limited in this embodiment.
- the radio frequency circuit 1604 is used to receive and transmit RF (Radio Frequency) signals, also called electromagnetic signals.
- the radio frequency circuit 1604 communicates with a communication network and other communication devices through electromagnetic signals.
- the radio frequency circuit 1604 converts the electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal.
- the radio frequency circuit 1604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on.
- the radio frequency circuit 1604 can communicate with other terminals through at least one wireless communication protocol.
- the wireless communication protocol includes but is not limited to: World Wide Web, Metropolitan Area Network, Intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks.
- the radio frequency circuit 1604 may further include circuits related to NFC (Near Field Communication), which is not limited in this application.
- the touch screen 1605 is used to display a UI (User Interface, user interface).
- the UI may include graphics, text, icons, video, and any combination thereof.
- the touch display 1605 also has the ability to collect touch signals on or above the surface of the touch display 1605.
- the touch signal can be input to the processor 1601 as a control signal for processing.
- the touch display screen 1605 is used to provide virtual buttons and / or virtual keyboards, also called soft buttons and / or soft keyboards.
- touch display screen 1605 there may be one touch display screen 1605, which is provided with the front panel of the terminal; in other embodiments, there may be at least two touch display screens 1605, which are respectively provided on different surfaces of the terminal or have a folded design; In still other embodiments, the touch display screen 1605 may be a flexible display screen, which is disposed on a curved surface or a folding surface of the terminal. Even, the touch display screen 1605 can also be set as a non-rectangular irregular figure, that is, a shaped screen.
- the touch display 1605 can be made of LCD (Liquid Crystal), Liquid Crystal Display (OLED), Organic Light-Emitting Diode (Organic Light Emitting Diode) and other materials.
- the camera assembly 1606 is used to collect images or videos.
- the camera assembly 1606 includes a front camera and a rear camera.
- the front camera is used for video calls or selfies
- the rear camera is used for photos or videos.
- the camera assembly 1606 may also include a flash.
- the flash can be a single-color flash or a dual-color flash. Dual color temperature flash refers to the combination of warm flash and cold flash, which can be used for light compensation at different color temperatures.
- the audio circuit 1607 is used to provide an audio interface between the user and the terminal.
- the audio circuit 1607 may include a microphone and a speaker.
- the microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals and input them to the processor 1601 for processing, or input them to the radio frequency circuit 1604 to implement voice communication.
- the microphone can also be an array microphone or an omnidirectional acquisition microphone.
- the speaker is used to convert the electrical signal from the processor 1601 or the radio frequency circuit 1604 into sound waves.
- the speaker can be a traditional thin-film speaker or a piezoelectric ceramic speaker.
- the speaker When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible by humans, but also convert electrical signals into sound waves inaudible to humans for ranging and other purposes.
- the audio circuit 1607 may also include a headphone jack.
- the positioning component 1608 is used to locate the current geographic location of the terminal to implement navigation or LBS (Location Based Service, location-based service).
- the positioning component 1608 may be a positioning component based on the GPS (Global Positioning System) of the United States, the Beidou system of China, or the Galileo system of Russia.
- the power supply 1609 is used to supply power to various components in the terminal.
- the power supply 1609 may be alternating current, direct current, disposable batteries, or rechargeable batteries.
- the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery.
- the wired rechargeable battery is a battery charged through a wired line
- the wireless rechargeable battery is a battery charged through a wireless coil.
- the rechargeable battery can also be used to support fast charging technology.
- the terminal further includes one or more sensors 1610.
- the one or more sensors 1610 include, but are not limited to: an acceleration sensor 1611, a gyro sensor 1612, a pressure sensor 1613, a fingerprint sensor 1614, an optical sensor 1615, and a proximity sensor 1616.
- FIG. 16 does not constitute a limitation on the terminal, and may include more or fewer components than shown, or combine certain components, or adopt different component arrangements.
- Embodiments of the present application also provide a computer-readable storage medium, where the storage medium stores at least one computer-readable instruction, at least one program, code set, or computer-readable instruction set, the at least one computer-readable instruction, The at least one program, the code set, or the computer readable instruction set is executed by the processor to implement the graphic code recognition method provided by the foregoing various embodiments.
- the present application also provides a computer program product containing computer-readable instructions, which when run on a computer, causes the computer to execute the graphic code recognition method described in the foregoing embodiments.
- steps in the flowcharts of the above embodiments are displayed sequentially according to the arrows, the steps are not necessarily executed in the order indicated by the arrows. Unless clearly stated in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above embodiments may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but may be executed at different times. These sub-steps or stages The execution order of is not necessarily sequential, but may be executed in turn or alternately with at least a part of other steps or sub-steps or stages of other steps.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Electromagnetism (AREA)
- Toxicology (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Biomedical Technology (AREA)
- Molecular Biology (AREA)
- Software Systems (AREA)
- Biophysics (AREA)
- Computational Linguistics (AREA)
- Data Mining & Analysis (AREA)
- Evolutionary Computation (AREA)
- Life Sciences & Earth Sciences (AREA)
- Computing Systems (AREA)
- General Engineering & Computer Science (AREA)
- Mathematical Physics (AREA)
- Multimedia (AREA)
- User Interface Of Digital Computer (AREA)
- Image Analysis (AREA)
Abstract
一种图形码识别方法,包括:显示目标图像,目标图像中包含至少两个图形码;当接收到对目标图像的图形码识别操作时,获取目标图像中各个图形码的图形码位置信息;根据图形码位置信息,确定图形码识别操作指示的目标图形码;显示目标图形码对应的目标图形码识别结果。
Description
本申请要求于2018年10月22日提交中国专利局,申请号为2018112316520,申请名称为“图形码识别方法、装置、终端及存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请实施例涉及图形码识别领域,特别涉及一种图形码识别方法、装置、终端及存储介质。
图形码是一种用于承载信息或数据的载体,常见的图形码包括条形码、二维码等等。
长按识别是一种常见的图形码识别方式。以二维码为例,终端显示包含二维码的图像,当接收到对图像的长按操作时,终端显示识别二维码选项。用户点击识别二维码选项后,终端即对图像中的二维码进行识别,并根据识别结果进行相应操作。比如,当识别结果为网址时,终端即进行网页跳转;当识别结果是支付链接时,终端即显示支付界面。
然而,当图像中包含至少两个图形码时,终端仅能够识别其中的一个图形码,若识别出的图形码不是用户期望识别的图形码,用户需要从该图像中手动截取出期望识别的图形码再进行图形码识别,影响图形码的识别效率。
发明内容
根据本申请提供的各种实施例,提供了一种图形码识别方法、装置、终端及存储介质。
一种图形码识别方法,由终端执行,所述方法包括:
显示目标图像,所述目标图像中包含至少两个图形码;
当接收到对所述目标图像的图形码识别操作时,获取所述目标图像中各个图形码的图形码位置信息;
根据所述图形码位置信息,确定所述图形码识别操作指示的目标图形码,所述目标图形码属于所述至少两个图形码;
显示所述目标图形码对应的目标图形码识别结果。
一种图形码识别装置,设置于终端,所述装置包括:
图像显示模块,用于显示目标图像,所述目标图像中包含至少两个图形码;
位置获取模块,用于当接收到对所述目标图像的图形码识别操作时,获取所述目标图像中各个图形码的图形码位置信息;
目标确定模块,用于根据所述图形码位置信息,确定所述图形码识别操作指示的目标图形码,所述目标图形码属于所述至少两个图形码;
结果显示模块,用于显示所述目标图形码对应的目标图形码识别结果。
一种终端,所述终端包括处理器和一个或多个存储器,所述一个或多个存储器中存储有至少一条计算机可读指令、至少一段程序、代码集或计算机可读指令集,所述至少一条计算机可读指令、所述至少一段程序、所述代码集或计算机可读指令集由所述一个或多个处理器执行以实现如上述方面所述的图形码识别方法。
一个或多个计算机可读存储介质,所述存储介质中存储有至少一条计算机可读指令、至少一段程序、代码集或计算机可读指令集,所述至少一条计算机可读指令、所述至少一段程序、所述代码集或计算机可读指令集由一个或多个处理器执行以实现如上述方面所述的图形码识别方法。
本申请的一个或多个实施例的细节在下面的附图和描述中提出。基于本申请的说明书、附图以及权利要求书,本申请的其它特征、目的和优点将变得更加明显。
为了更清楚地说明本申请实施例中的技术方案,下面将对实施例描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1示出了本申请一个实施例提供的实施环境的示意图;
图2是相关技术中图形码识别过程的流程图;
图3是本申请实施例中图形码识别过程的流程图;
图4示出了本申请一个实施例提供的图形码识别方法的流程图;
图5示出了本申请一个实施例提供的长按识别场景下图形码识别方法的流程图;
图6是图5所示图形码识别方法实施过程的界面示意图;
图7是确定图形码与目标识别位置之间距离时的实施示意图;
图8示出了本申请一个实施例提供的扫码识别场景下图形码识别方法的流程图;
图9是图8所示图形码识别方法实施过程的界面示意图;
图10是本申请一个实施例提供的目标检索模型的结构示意图;
图11是本申请一个实施例提供的第一残差块的结构示意图;
图12是本申请一个实施例提供的第二残差块的结构示意图;
图13示出了本申请另一个实施例提供的图形码识别方法的流程图;
图14示出了本申请另一个实施例提供的图形码识别方法的流程图;
图15示出了本申请一个实施例提供的图形码识别装置的框图;
图16示出了本申请一个实施例提供的终端的结构示意图。
为使本申请的目的、技术方案和优点更加清楚,下面将结合附图对本申请实施方式作进一步地详细描述。应当理解,此处所描述的具体实施方式仅仅用以解释本申请,并不用于限定本申请。
请参考图1,其示出了本申请一个实施例提供的实施环境的示意图。该实施环境中包括终端120和服务器140。
终端120是具有图形码识别功能的电子设备,该电子设备可以是是智能手机、平板电脑或个人计算机等等。图1中以终端120是智能手机为例进行说明。
其中,该图形码识别功能可以由电子设备的操作系统提供,也可以由电子设备中安装的第三方应用程序提供,该第三方应用程序可以是支付类应用程序、即时通信应用程序、购物类应用程序、视频播放类应用程序、浏览器应用程序等等,本申请实施例对此不做限定。
并且,终端120识别的图形码可以是图片,比如即时通信应用程序中收到的图片;也可以是通过摄像组件采集到的图像,比如启用应用程序的图形码扫描功能后扫描到的图像,本申请实施例对此不做限定。
本申请实施例中的终端120还具有目标检测功能,借助该目标检测功能,终端120能够识别图像中各个图形码的位置和/或类型,以便后续根据图形码 的位置从图像中提取图形码,根据图形码的类型采用相应的解码器进行图形码识别,得到图形码识别结果。
在一种可能的实施方式中,该目标检测功能借助基于深度学习训练得到的目标检测模型实现,该目标检测模型用于根据输入的图像,输入图像中各个图形码的位置信息及位置置信度,和/或,各个图形码的类型以及类型置信度。
终端120与服务器140之间通过有线或无线网络相连。
服务器140是一台服务器、若干台服务器构成的服务器集群或云计算中心。
在一个实施例中,服务器140为资源服务器,用于提供网页资源。在一种可能的应用场景下,当终端120接收到图形码识别操作,并完成目标图形码识别而得到图形码识别结果后,即根据该图形码识别结果,从相应的服务器140处获取网页资源,进而对该网页资源进行显示。
在一个实施例中,上述的无线网络或有线网络使用标准通信技术和/或协议。网络通常为因特网、但也可以是任何网络,包括但不限于局域网(Local Area Network,LAN)、城域网(Metropolitan Area Network,MAN)、广域网(Wide Area Network,WAN)、移动、有线或者无线网络、专用网络或者虚拟专用网络的任何组合)。在一些实施例中,使用包括超文本标记语言(Hyper Text Mark-up Language,HTML)、可扩展标记语言(Extensible Markup Language,XML)等的技术和/或格式来代表通过网络交换的数据。此外还可以使用诸如安全套接字层(Secure Socket Layer,SSL)、传输层安全(Transport Layer Security,TLS)、虚拟专用网络(Virtual Private Network,VPN)、网际协议安全(Internet Protocol Security,IPsec)等常规加密技术来加密所有或者一些链路。在另一些实施例中,还可以使用定制和/或专用数据通信技术取代或者补充上述数据通信技术。
本申请各个实施例提供的图形码识别方法由图1中的终端120执行,即由终端120在本地实现图形码识别,得到图形码识别结果;当然,在其他可能的实施方式中,终端120还可以将待识别的图像上传至后台服务器,由后台服务器对图像进行图形码识别并反馈图形码识别结果,供终端120进行显示,本申请实施例并不对此进行限定。
本申请实施例提供的图形码识别方法可用于实现一图多码的识别,即识别多图形码图像(指包含多个图形码的图像)中的指定图形码,其适用的应用场景可以包括扫码识别场景和长按识别场景。下面对不同应用场景下的图形码识别方法进行说明。
扫码识别场景
日常生活中,经常需要使用终端进行扫码,比如,通过扫描公众号二维码关注公众号,通过扫描点单二维码进行自助点单,通过扫描收款二维码进行支付等等。在某些扫码识别场景下,终端可能会同时采集到包含多个图形码的图像。比如,在扫码付款时,终端采集到同时包含收款码A(对应A支付应用程序)以及收款码B(对应B支付应用程序)的图像。在此场景下,终端通过图形码识别方法首先识别图像中各个图形码的图形码位置,然后根据图形码位置确定出符合用户识别意图的目标图形码,从而对该目标图形码进行识别,得到相应的图形码识别结果。
长按识别场景
用户在使用应用程序的过程中,经常会查看到包含图形码的图片。比如,在使用即时通信应用程序时,接收到其他用户发送的包含图形码的图片,在使用浏览器应用程序时,显示的网页图片中包含图形码。此时,用户可以通过长按图片触发终端对图片中的图形码进行识别。在此场景下,终端通过图形码识别方法首先识别图片中各个图形码的图形码位置,然后根据图形码位置和长按位置,确定出用户期望识别的目标图形码,从而对该目标图形码进行识别,得到相应的图形码识别结果。
当然除了应用于上述场景外,上述图形码识别方法还可以用于其它涉及一图多码的识别场景,本申请实施例并不对此构成限定。
以长按识别图形码为例,如图2所示,其示出了相关技术中图形码识别过程的流程图。
获取到待识别图像(原图)后,由于图形码识别的尺寸敏感性,终端需要采用金字塔采样的方式对原图进行下采样,从而得到不同尺寸的下采样图。如图2,进行金字塔采样后,得到原图21、1/4下采样图22以及1/16下采样图23。
对于得到的各张下采样图,为了方便后续图形码解码,终端对其进行图 像二值化处理,得到仅包含黑白两色的二值化图,其中,进行图像二值化处理时可以采用混合二值化、快速窗口二值化和自适应二值化。如图2,终端对原图21进行图像二值化处理后,得到二值化图24。
进一步的,终端尝试各种解码器对得到的二值化图进行解码,比如,分别利用一维码解码器进行一维码解码,利用二维码解码器进行二维码解码。若利用当前解码器解码失败,则尝试下一种解码器,直至解码成功。其中,利用解码器进行解码时,需要遍历二值化图中的各个像素点进行遍历。如图2所示,终端利用二维码解码器对二值化图进行解码过程中,首先识别出二维码的三个顶点25,然后对二维码进行解码。
然而,当图像中包含至少两个图形码时,采用上述方法识别图像中的图形码时,终端只能对其中的一个图形码进行解码(通常为图像中第一个解码成功的图形码)。若用户期望得到图像中指定图形码的识别结果,采用上述方法将无法实现。
而本申请实施例提供的图形码识别方法中,通过一套全新的图形码识别流程来解决上述问题。在该图形码识别流程中,如图3所示,终端获取到待识别图像31后,首先对待识别图像31进行目标检测,获取待识别图像31中各个图形码的图形码位置以及图形码类型,从而根据图形码位置从待识别图像31中提取出各个图形码。对于提取出的图形码32,终端基于小图上采样,大图下采样的原则对其进行采样,得到采样图33,并对采样图33进行图像二值化处理,得到二值化图34。对于得到的二值化图34,终端根据图形码类型,采用图形码类型对应的图形码解码器进行解码,得到二值化图34对应的图形码识别结果。
进一步的,终端确定用户期望识别的目标图形码,从而对该目标图形码对应的识别结果进行显示,实现多图形码图形中指定图形码的识别。
请参考图4,其示出了本申请一个实施例提供的图形码识别方法的流程图。本实施例以该方法应用于图1中的终端120来举例说明,该方法可以包括以下几个步骤:
步骤401,显示目标图像,目标图像中包含至少两个图形码。
其中,不同应用场景下的目标图像不同。在一个实施例中,在扫码识别场景下,该目标图像为终端取景框内显示的图像,即终端通过摄像组件实时 采集的图像;在长按识别场景下,该目标图像为终端显示的图片,比如即时通信应用程序中接收到的图片。
本申请实施例中的目标图像中包含至少两个图形码,且至少两个图形码的图形码类型可以相同,也可以不同,其中,图形码类型包括一维码(又称为条形码)和二维码中的至少一种,且二维码包括点状二维码、环状二维码和辐射状二维码中的至少一种,本申请并不对图形码的具体类型进行限定。
步骤402,当接收到对目标图像的图形码识别操作时,获取目标图像中各个图形码的图形码位置信息。
在一个实施例中,图形码识别操作用于指示对目标图像中指定的图形码进行识别。
不同应用场景下,对目标图像的图形码识别操作不同。在一个实施例中,在扫码识别场景下,该图形码识别操作是对拍摄操作;在长按识别场景下,该图形码识别操作是对图片的长按操作。
图形码位置信息用于唯一表示图形码在目标图像中的位置。在一种可能的实施方式中,图形码位置信息中包括图形码中预设标志点的坐标,比如,图形码位置信息包括图形码左上角顶点的坐标,或者,图形码中心点的坐标。
在一个实施例中,该图形码位置信息中还包含图形码的尺寸信息。比如,图形码的高度和宽度信息。
在一个实施例中,除了获取到各个图形码的图形码位置信息外,终端还获取到各个图形码的图形码类型,以便后续根据图形码类型采用相应的解码器进行解码,而不再需要尝试各种类型的解码器。
步骤403,根据图形码位置信息,确定图形码识别操作指示的目标图形码,目标图形码属于至少两个图形码。
通常情况下,图形码识别操作所指示的识别位置通常靠近或倾向于用户期望识别的目标图形码,因此,在一种可能的实施方式中,本步骤包括如下步骤。
一、确定图形码识别操作指示的目标识别位置。
在一个实施例中,该目标识别位置采用坐标的形式表示。
二、根据目标识别位置和图形码位置信息,确定目标图形码。
在一个实施例中,终端根据目标识别位置和图形码位置信息,计算目标 识别位置与各个图形码之间的距离,并将最短距离对应的图形码确定为目标图形码。
步骤405,显示目标图形码对应的目标图形码识别结果。
在一种可能的实施方式中,终端对各个图形码进行识别,得到至少两条图形码识别结果,并最终显示目标图形码对应的目标图形码识别结果;
在另一种可能的实施方式中,终端仅对确定出的目标图形码进行识别,并对得到的目标图形码识别结果进行显示。
针对图形码识别过程,如图3所示,对于识别出的各个图形码,终端获取图形码的尺寸,并检测该尺寸是否大于预设尺寸(比如300px×300px),若大于,则确定图形码为大图,并对大图进行下采样,使下采样后图形码的尺寸为预设尺寸;若小于,则确定图形码为小图,并对小图进行上采样(可以采用超分辨率采样),使上采样后图形码的尺寸为预设尺寸。由于借助目标检测能够确定出图像中图形码的尺寸,因此终端能够根据图形码的尺寸针对性的进行采样,无需采用金字塔采样法进行多次采样,从而降低了识别过程的数据处理量。示意性的,图2中,终端需要对待识别图像进行多次采样,而图3中,终端仅需要对图像进行小图上采样。
对图形码进行采样后,终端进一步对图形码进行图像二值化处理。相较于相关技术中,需要对整张图像进行二值化处理,本实施例中,终端仅需要对提取出的图形码进行二值化处理,从而降低了二值化过程的数据处理量。
完成图像二值化处理后,终端对二值化后的图形码进行解码。由于目标检测时可以获取到图形码的图形码类型,因此,终端可以针对性的采用解码器进行解码,无需尝试各种解码器,从而降低了解码过程的数据处理量。示意性的,图2中,终端需要尝试三种解码器进行解码,而图3中,终端直接通过二维码解码器对图形码进行解码即可。
综上所述,本申请实施例中,当接收到对包含至少两个图形码的目标图像的图形码识别操作时,首先获取目标图像中各个图形码的图形码位置信息,然后根据图形码位置信息,确定图形码识别操作所指示的目标图形码,从而对目标图形码对应的目标图形码识别结果进行显示;借助图形码位置识别机制,终端能够同时识别出同一图像中的多个图形码,从而根据各个图形码各自的位置确定出符合用户识别意图的目标图形码,进而返回目标图形码的识 别结果,提高了图形码的识别效率,解决了相关技术中,当图像中包含至少两个图形码时,用户需要从该图像中手动截取出期望识别的图形码再进行图形码识别,导致图形码识别效率较低的问题。
在不同的应用场景下,终端从至少两个图形码中确定出目标图形码的方式不同。下面采用两个实施例,分别对长按识别和扫码识别场景下目标图形码确定过程进行说明。
请参考图5,其示出了本申请另一个实施例提供的图形码识别方法的流程图。本实施例以该方法应用于长按识别场景为例进行说明,该方法可以包括以下几个步骤:
步骤501,显示目标图像,目标图像中包含至少两个图形码。
本实施例中,该目标图像为终端显示的图片。示意性的,如图6所示,终端显示有目标图像61,该目标图像61中包含辐射状二维码62和点状二维码63。
步骤502,当接收到对目标图像的图形码识别操作时,获取目标图像中各个图形码的图形码位置信息。
本实施例中,图形码识别操作是对目标图像的触发操作。
在一种可能的实施方式中,当终端是具有触摸功能的移动终端时,当接收到对目标图像的长按操作时,终端显示若干操作选项,当接收到对图形码识别选项的选择操作时,终端确定接收到图形码识别操作。
示意性的,如图6所示,接收到对目标图像61的长按操作后,终端显示操作选项菜单,并在接收到对图形码识别选项64的选择操作时,确定接收到图形码识别操作。
在其他可能的实施方式中,当终端是PC一类包含外部输入设备(比如鼠标)的终端时,当接收到对目标图像的点击操作(外部输入设备执行)时,终端显示显示若干操作选项,当接收到对图形码识别选项的选择操作时,终端确定接收到图形码识别操作。
步骤503,将图形码识别操作对应的触发位置确定为目标识别位置。
在长按识别场景下,用户通常会在需要识别的图形码上执行长按操作,因此,在一种可能的实施方式中,终端将图形码识别操作对应的触发位置(比如长按位置)确定为目标识别位置
示意性的,如图7所示,终端将接收到长按信号的长按位置确定为目标识别位置65。
在一个实施例中,为了方便后续计算图形码与目标识别位置之间的距离,终端获取目标识别位置在目标图像上的坐标。比如,终端获取到目标识别位置的坐标为(x
pos,y
pos)。
在一个实施例中,首次使用长按识别图形码功能时,终端显示提示信息,指示用户长按不同位置的图形码以进行识别,本实施例对此不做限定。
步骤504,根据触发位置的位置信息以及各个图形码的图形码位置信息,确定目标识别位置与各个图形码之间的距离。
在一个实施例中,对于识别出的各个图形码,终端根据各个图形码的图形码位置信息,计算图形码的码中心,并根据触发位置与码中心的坐标,计算目标识别位置与各个图形码之间的距离。
在一种可能的实施方式中,终端根据图形码位置信息中各个顶点的顶点坐标,计算图形码的码中心坐标;或者,终端根据图形码位置信息中至少一个顶点的顶点坐标和图形码尺寸信息,计算图形码的码中心。本申请并不对计算码中心坐标的方式进行限定。
示意性的,如图7所示,终端获取到辐射状二维码62的第一码中心621的坐标为(x
1,y
1),点状二维码63的第二码中心631的坐标为(x
2,y
2),计算得到目标识别位置65与辐射状二维码62之间的距离s
1为
步骤505,将最短距离对应的图形码确定为目标图形码。
进一步的,终端将最短距离对应的图形码确定为目标图形码,即将距离目标识别位置最近的图形码确定为目标图形码。
示意性的,如图7所示,由于s
1<s
2,因此,终端将点状二维码63确定为目标图形码。
步骤506,显示目标图形码对应的目标图形码识别结果。
示意性的,如图6所示,终端对点状二维码63进行图形码识别,得到的目标图形码识别结果为游戏下载链接,从而根据游戏下载链接跳转显示游戏下载界面66。
本实施例中,终端将图形码识别操作对应的触发位置确定为目标识别位置,并通过计算目标识别位置与各个图形码之间的距离确定用户指示的目标图形码,从而在长按识别场景下,实现对目标图像中长按位置处图形码的识别。
请参考图8,其示出了本申请另一个实施例提供的图形码识别方法的流程图。本实施例以该方法应用于扫码识别场景为例进行说明,该方法可以包括以下几个步骤:
步骤801,显示目标图像,目标图像中包含至少两个图形码。
本实施例中,该目标图像为取景框内显示的图像。示意性的,如图9所示,终端开启扫码识别功能后,在取景框91内显示通过摄像头采集到的目标图像,该目标图像中包含辐射状二维码92和点状二维码93。
步骤802,当接收到对目标图像的图形码识别操作时,获取目标图像中各个图形码的图形码位置信息。
本实施例中,图形码识别操作是对目标图像的拍摄操作。
在一种可能的实施方式中,终端界面中显示有拍摄控件,当接收到对该拍摄控件的点击操作时,终端确定接收到图形码识别操作。
在其他可能的实施方式中,当检测到终端保持稳定,且稳定达到时长阈值(比如0.5s)时,终端确定接收到图形码识别操作。本申请实施例并不对此进行限定。
步骤803,将取景框中心在目标图像中对应的位置确定为目标识别位置。
在扫码识别场景下,用户通常会移动终端,使需要识别的目标图形码位于或靠近取景框中心,因此,在一种可能的实施方式中,终端将取景框中心在目标图像中对应的位置确定为目标识别位置。
比如,终端确定目标识别位置的坐标为(x
pos,y
pos)。
步骤804,根据取景框中心的位置信息以及各个图形码的图形码位置信息,确定目标识别位置与各个图形码之间的距离。
在一个实施例中,终端根据各个图形码的图形码位置信息,计算图形码的码中心,并根据取景框中心与码中心的坐标,计算目标识别位置与各个图形码之间的距离。其中,计算目标识别位置与各个图形码之间距离的过程可以参考上述步骤504,本实施例在此不再赘述。
步骤805,将最短距离对应的图形码确定为目标图形码。
进一步的,终端将最短距离对应的图形码确定为目标图形码,即将距离取景框中心最近的图形码确定为目标图形码。
示意性的,如图9所示,由于辐射状二维码92与取景框中心的距离小于点状二维码93与取景框中心的距离,因此,终端将辐射状二维码92确定为目标图形码。
步骤806,显示目标图形码对应的目标图形码识别结果。
示意性的,如图9所示,终端对辐射状二维码92进行图形码识别,得到的目标图形码识别结果为小程序跳转信息,从而根据小程序跳转信息跳转显示小程序界面94。
其中,小程序(Mini program)是一种不需要下载安装即可使用的应用程序。开发者可以为终端的应用开发相应的小程序,小程序可以作为子应用被嵌入终端的应用中,通过运行应用内的小程序能够为用户提供更多样化的服务。
本实施例中,终端将取景框中心确定为目标识别位置,并通过计算目标识别位置与各个图形码之间的距离确定用户期望扫描的目标图形码,从而在扫码识别场景下,实现对取景框内多个图形码中指定图形码的识别。
在一种可能的实施方式中,终端中存储有预先训练得到的目标检测模型,该目标检测模型通过深度学习训练得到,用于识别图像中的图形码,并输出图形码在图像中的位置信息。相应的,上述各个实施例中,终端获取目标图像中各个图形码的图形码位置信息时可以包括如下步骤。
一、将目标图像输入目标检测模型,得到各个图形码的图形码预测位置信息以及位置置信度,目标检测模型通过深度学习训练得到。
其中,位置置信度越高,预测位置信息所指示位置处为图形码的概率越高,反之,该预测位置信息所指示位置处为图形码的概率越低。在一个实施例中,该预测位置信息包含图形码中指定标志点的坐标,该指定标志点可以为顶点或者码中心。
在一种可能的实施方式中,该目标检测模型中包括i个串联的残差网络和一个空洞卷积网络。示意性的,如图10所示,目标检测模型中包括3个串联的残差网络和空洞卷积网络,分别为第一残差网络1010、第二残差网络 1020、第三残差网络1030以及空洞卷积网络1040。
其中,每个残差网络中包含一个下采样块和j个第一残差块。下采样块用于对输入内容进行下采样后得到图像特征,第一残差块是残差网络中的基础块,通常包括残差支路和短路支路,残差支路用于对残差块的输入进行非线性变换,短路支路用于对残差块的输入进行恒等变换或线性变换。
在一个实施例中,每个残差网络中包含的第一残差块的数量可以相同,也可以不同
示意性的,如图10所示,第一残差网络1010中包含3个第一残差块,第二残差网络1020中包含7个第一残差块,第三残差网络1030中包含3个第一残差块。
在一个实施例中,该第一残差块可以采用常规残差块或者瓶颈残差块(Bottleneck Residual Block)。
当第一残差块采用常规残差块或瓶颈残差块时,输入第一残差块的特征分别在各层卷积层进行卷积处理,残差网络参数量和计算量集中在卷积层中。为了进一步减小整个目标检测模型的大小,使其能够应用于终端中,在一种可能的实施方式中,第一残差块中的部分卷积被替换为深度depthwise卷积,在保证识别准确率的前提下,降低残差网络的大小,提高残差网络的处理速度。
示意性的,如图11所示,瓶颈残差块1110中包含三个卷积层,其中,第一卷积层中包含若干个1×1的卷积核,第二卷积层中包含若干个3×3的卷积核,第三卷积层中包含若干个1×1的卷积核。每个卷积层均包含归一化(Batch Normalization,BN)层,且第一卷积层和最终输出均包含激活层(Rectified Linear Units,ReLU)。对瓶颈残差块1110进行改造时,将第二卷积层中的3×3卷积核变换为深度depthwise卷积核后,即得到第一残差块1120。
空洞卷积也称为扩张卷积,是一种在卷积核之间注入空洞的一种卷积方式。相较于普通卷积,空洞卷积引入了一个称为“扩张率(dilation rate)”的超参数,该参数定义了卷积核处理数据时各值的间距。通过空洞卷积处理,一方面能够保持图像特征的空间尺度不变,从而避免因减少了图像特征的像素的信息而导致的信息损失,另一方面能够扩大感受野,从而实现更加精准 的目标检测。其中,感受野是神经网络中的隐藏层输出的特征图上的像素点在原始图像上映射的区域大小,像素在原始图像上的感受野越大,表示其映射的原始图像范围越大,也意味着其可能蕴含更为全局、语义层次更高的特征。
在一个实施例中,空洞卷积网络中包含k个第二残差块。示意性的,如图10所示,空洞卷积网络中包含3个第二残差块。
在一种可能的实施方式中,第二残差块中应用dilation卷积来扩大感受野,并且,为了避免底层特征被直接作为上层特征,导致上层特征无法获取更高语义层级和视觉感受野,第二残差块的短路支路中还包含卷积变换。
示意性的,如图12所示,瓶颈残差块1210中包含三个卷积层,其中,第一卷积层中包含若干个1×1的卷积核,第二卷积层中包含若干个3×3的卷积核,第三卷积层中包含若干个1×1的卷积核,每个卷积层均包含BN层,且第一卷积层和最终输出均包含ReLU。对瓶颈残差块1210进行改造时,将第二卷积层中的3×3卷积核变换为空洞dilated卷积核,并在短路支路增加包含若干个1×1卷积核的卷积层,最终得到第二残差块1220。
并且,本申请实施例中,残差网络中最末端的第一残差块的输出,以及空洞卷积网络中各个第二残差块的输出,均输入输出网络,由输出网络进行分类和回归,提高后续分类结果的准确性。
示意性的,如图10所示,第二残差网络1020最末端的第一残差块、第三残差网络1030最末端的第一残差块、空洞卷积网络1040中各个第二残差块的输出均输入到输出网络中。
二、根据图形码预测位置信息和位置置信度,确定各个图形码的图形码位置信息。
在一种可能的实施方式中,终端根据各条图形码预测位置信息对应的位置置信度,将位置置信度大于置信度阈值(比如90%)的图形码预测位置信息确定为各个图形码的图形码位置信息。
本申请实施例中,目标检测模型除了能够预测图形码的位置外,还可以预测图形码的图形码类型,相应的,后续进行图形码识别时,终端采用图形码类型对应的图形码解码器进行解码,提高解码效率。
在一种可能的实施方式中,终端获取目标检测模型输出的图形码预测类 型以及类型置信度,并根据图形码预测类型和类型置信度,确定各个图形码的图形码类型,图形码类型包括一维码和二维码中的至少一种。
在一个实施例中,终端根据各个图形码预测类型对应的类型置信度,将类型置信度大于置信度阈值(比如90%)的图形码预测类型确定为各个图形码的图形码类型。
相应的,在图4的基础上,如图13所示,步骤405之前还包括如下步骤:
步骤404,通过目标图形码类型对应的目标解码器对目标图形码进行图形码识别,得到目标图形码识别结果,目标图形码类型为目标图形码对应的图形码类型;或,通过各个图形码的图形码类型各自对应的解码器,对各个图形码进行图形码识别,得到各个图形码对应的图形码识别结果;将目标图形码对应的图形码识别结果确定为目标图形码识别结果。
在一个实施例中,根据目标检测模型输出的图形码位置信息,终端从目标图像中提取目标图形码,并根据目标图形码对应的目标图形码类型,采用目标解码器对目标图形码进行图形码识别,从而得到目标图形码对应的目标图形码识别结果。
在一个实施例中,对于目标图像中的非目标图形码,终端则无需进行图形码识别,从而降低终端识别图形码时的数据处理量。
在另一种可能的实施方式中,终端根据目标检测模型输出的图形码位置信息,从目标图像中提取各个图形码,并根据各个图形码各自对应的图形码类型,采用相应的解码器对各个图形码进行图形码识别,从而得到各个图形码对应的图形码识别结果。进一步的,终端将目标图形码对应的图形码识别结果确定为目标图形码识别结果,以供后续显示。
相较于相关技术中,需要尝试各种解码器进行图形码识别,本实施例中,终端根据识别出的图形码类型,针对性采用相应的解码器进行图形码识别,提高了识别效率,并降低了识别时的数据处理量。
除了根据图形码的位置信息确定目标图形码外,在另一种可能的实施方式中,终端还可以根据目标图像中各个图形码对应的图形码识别结果,确定用户期望识别的目标图形码,在一个实施例中,在图4的基础上,如图14所示,步骤403可以包括如下步骤。
步骤403A,根据图形码位置信息对各个图形码进行图形码识别,得到至 少两条图形码识别结果。
根据各个图形码对应的图形码位置,终端从目标图像中提取各个图形码,并对各个图形码进行图形码识别,得到各个图形码对应的图形码识别结果。
步骤403B,确定图形码识别操作对应的目标应用程序。
在一种可能的实施方式中,目标应用程序为接收到图形码识别操作的应用程序。
比如,当通过即时通信应用程序A的图形码识别功能进行扫码时,将即时通信应用程序A确定为目标应用程序。
步骤403C,若图形码识别结果属于目标应用程序支持的识别结果,则将图形码识别结果对应的图形码确定为目标图形码。
在一种可能的实施方式中,终端中不同应用程序支持的识别结果的类型不同,且各个应用程序对应各自的识别结果列表,该识别结果列表中包含支持的识别结果的类型。
在一个实施例中,当图形码识别结果为目标应用程序支持的识别结果时,目标应用程序能够解析该图形码识别结果,反之,目标应用程序无法解析该图形码识别结果。
比如,对于即时通信应用程序A,其支持显示的识别结果为B支付应用的支付页面。当使用即时通信应用程序A的扫码功能进行扫码时,若扫码图像中,第一图形码对应的图形码识别结果为B支付应用的支付页面,第二图形码对应的图形码识别结果为C支付应用的支付页面,终端则确定第一图形码对应的图形码识别结果为目标图形码识别结果,并将第一图形码确定为目标图形码。
在一个实施例中,该识别结果列表中包含识别结果关键词。终端基于识别结果列表,检测图形码识别结果是否包含识别结果关键词,若包含,则确定图形码识别结果属于当前应用程序支持的识别结果。
后续终端仅显示目标图形码识别结果,而不会对其他图形码识别结果进行显示。
在实际应用过程中,当用户使用终端进行扫码支付,且同时扫描到不同支付应用对应的支付二维码时,终端确定当前扫码时使用应用程序所支持的支付应用,并对该支付应用对应支付二维码的识别结果进行显示,方便用户 在当前应用程序中进行快速支付,并避免因当前应用程序无法显示其他支付应用的支付页面而导致无法支付的问题。
下述为本申请装置实施例,可以用于执行本申请方法实施例。对于本申请装置实施例中未披露的细节,请参照本申请方法实施例。
请参考图15,其示出了本申请一个实施例提供的图形码识别装置的框图。该装置可以是图1所示实施环境中的终端120,也可以设置在终端120上。该装置中包括各个模块或单元,每个模块或单元可全部或部分通过软件、硬件或其组合来实现。该装置可以包括:
图像显示模块1501,用于显示目标图像,所述目标图像中包含至少两个图形码;
位置获取模块1502,用于当接收到对所述目标图像的图形码识别操作时,获取所述目标图像中各个图形码的图形码位置信息;
目标确定模块1503,用于根据所述图形码位置信息,确定所述图形码识别操作指示的目标图形码,所述目标图形码属于所述至少两个图形码;
结果显示模块1504,用于显示所述目标图形码对应的目标图形码识别结果。
在一个实施例中,所述目标确定模块1503,包括:
第一确定单元,用于确定所述图形码识别操作指示的目标识别位置;
第二确定单元,用于根据所述目标识别位置和所述图形码位置信息,确定所述目标图形码。
在一个实施例中,所述目标图像为图片,所述图形码识别操作是对所述图片的触发操作;
所述第一确定单元,用于将所述图形码识别操作对应的触发位置确定为所述目标识别位置;
所述第二确定单元,用于根据所述触发位置的位置信息以及各个图形码的所述图形码位置信息,确定所述目标识别位置与各个图形码之间的距离;将最短距离对应的图形码确定为所述目标图形码。
在一个实施例中,所述目标图像为取景框内显示的图像,所述图形码识别操作是对所述目标图像的拍摄操作;
所述第一确定单元,用于将取景框中心在所述目标图像中对应的位置确 定为所述目标识别位置;
所述第二确定单元,用于根据所述取景框中心的位置信息以及各个图形码的所述图形码位置信息,确定所述目标识别位置与各个图形码之间的距离;将最短距离对应的图形码确定为所述目标图形码。
在一个实施例中,所述位置获取模块1502,包括:
输入单元,用于将所述目标图像输入目标检测模型,得到各个图形码的图形码预测位置信息以及位置置信度;
第三确定单元,用于根据所述图形码预测位置信息和所述位置置信度,确定各个图形码的所述图形码位置信息。
在一个实施例中,所述目标检测模型中包括i个串联的残差网络和一个空洞卷积网络,其中,每个所述残差网络中包含一个下采样块和j个第一残差块,且所述第一残差块中包含深度depthwise卷积;所述空洞卷积网络中包含k个第二残差块,所述第二残差块中包含空洞dilated卷积,i,j,k为大于等于2的整数。
在一个实施例中,所述装置还包括:
类型获取模块,用于获取所述目标检测模型输出的各个图形码的图形码预测类型以及类型置信度;
类型确定模块,用于根据所述图形码预测类型和所述类型置信度,确定各个图形码的图形码类型,所述图形码类型包括一维码和二维码中的至少一种;
所述装置还包括:
第一解码模块,用于通过目标图形码类型对应的目标解码器对所述目标图形码进行图形码识别,得到所述目标图形码识别结果,所述目标图形码类型为所述目标图形码对应的图形码类型;
或,
第二解码模块,用于通过各个图形码的图形码类型各自对应的解码器,对各个图形码进行图形码识别,得到各个图形码对应的图形码识别结果;将所述目标图形码对应的图形码识别结果确定为所述目标图形码识别结果。
在一个实施例中,所述目标确定模块1503还包括:
结果识别单元,用于根据所述图形码位置信息对各个图形码进行图形码 识别,得到至少两条图形码识别结果;
第四确定单元,用于确定所述图形码识别操作对应的目标应用程序;
第五确定单元,用于若所述图形码识别结果属于所述目标应用程序支持的识别结果,则将所述图形码识别结果对应的图形码确定为所述目标图形码。
综上所述,本申请实施例中,当接收到对包含至少两个图形码的目标图像的图形码识别操作时,首先获取目标图像中各个图形码的图形码位置信息,然后根据图形码位置信息,确定图形码识别操作所指示的目标图形码,从而对目标图形码对应的目标图形码识别结果进行显示;借助图形码位置识别机制,终端能够同时识别出同一图像中的多个图形码,从而根据各个图形码各自的位置确定出符合用户识别意图的目标图形码,进而返回目标图形码的识别结果,提高了图形码的识别效率,解决了相关技术中,当图像中包含至少两个图形码时,用户需要从该图像中手动截取出期望识别的图形码再进行图形码识别,导致图形码识别效率较低的问题。
请参考图16,其示出了本申请一个实施例提供的终端的结构示意图。该终端可以实现成为图1所示实施环境中的终端120,以实施上述实施例提供的账号推荐方法。具体来讲:
终端包括有:处理器1601和存储器1602。
处理器1601可以包括一个或多个处理核心,比如4核心处理器、8核心处理器等。处理器1601可以采用DSP(Digital Signal Processing,数字信号处理)、FPGA(Field-Programmable Gate Array,现场可编程门阵列)、PLA(Programmable Logic Array,可编程逻辑阵列)中的至少一种硬件形式来实现。处理器1601也可以包括主处理器和协处理器,主处理器是用于对在唤醒状态下的数据进行处理的处理器,也称CPU(Central Processing Unit,中央处理器);协处理器是用于对在待机状态下的数据进行处理的低功耗处理器。在一些实施例中,处理器1601可以在集成有GPU(Graphics Processing Unit,图像处理器),GPU用于负责显示屏所需要显示的内容的渲染和绘制。一些实施例中,处理器1601还可以包括AI(Artificial Intelligence,人工智能)处理器,该AI处理器用于处理有关机器学习的计算操作。
存储器1602可以包括一个或多个计算机可读存储介质,该计算机可读存储介质可以是有形的和非暂态的。存储器1602还可包括高速随机存取存储 器,以及非易失性存储器,比如一个或多个磁盘存储设备、闪存存储设备。在一些实施例中,存储器1602中的非暂态的计算机可读存储介质用于存储至少一个计算机可读指令,该至少一个计算机可读指令用于被处理器1601所执行以实现本申请中提供的图形码识别方法。
在一些实施例中,终端还可选包括有:外围设备接口1603和至少一个外围设备。具体地,外围设备包括:射频电路1604、触摸显示屏1605、摄像头1606、音频电路1607、定位组件1608和电源1609中的至少一种。
外围设备接口1603可被用于将I/O(Input/Output,输入/输出)相关的至少一个外围设备连接到处理器1601和存储器1602。在一些实施例中,处理器1601、存储器1602和外围设备接口1603被集成在同一芯片或电路板上;在一些其他实施例中,处理器1601、存储器1602和外围设备接口1603中的任意一个或两个可以在单独的芯片或电路板上实现,本实施例对此不加以限定。
射频电路1604用于接收和发射RF(Radio Frequency,射频)信号,也称电磁信号。射频电路1604通过电磁信号与通信网络以及其他通信设备进行通信。射频电路1604将电信号转换为电磁信号进行发送,或者,将接收到的电磁信号转换为电信号。在一个实施例中,射频电路1604包括:天线系统、RF收发器、一个或多个放大器、调谐器、振荡器、数字信号处理器、编解码芯片组、用户身份模块卡等等。射频电路1604可以通过至少一种无线通信协议来与其它终端进行通信。该无线通信协议包括但不限于:万维网、城域网、内联网、各代移动通信网络(2G、3G、4G及5G)、无线局域网和/或WiFi(Wireless Fidelity,无线保真)网络。在一些实施例中,射频电路1604还可以包括NFC(Near Field Communication,近距离无线通信)有关的电路,本申请对此不加以限定。
触摸显示屏1605用于显示UI(User Interface,用户界面)。该UI可以包括图形、文本、图标、视频及其它们的任意组合。触摸显示屏1605还具有采集在触摸显示屏1605的表面或表面上方的触摸信号的能力。该触摸信号可以作为控制信号输入至处理器1601进行处理。触摸显示屏1605用于提供虚拟按钮和/或虚拟键盘,也称软按钮和/或软键盘。在一些实施例中,触摸显示屏1605可以为一个,设置终端的前面板;在另一些实施例中,触摸显示屏 1605可以为至少两个,分别设置在终端的不同表面或呈折叠设计;在再一些实施例中,触摸显示屏1605可以是柔性显示屏,设置在终端的弯曲表面上或折叠面上。甚至,触摸显示屏1605还可以设置成非矩形的不规则图形,也即异形屏。触摸显示屏1605可以采用LCD(Liquid Crystal Display,液晶显示器)、OLED(Organic Light-Emitting Diode,有机发光二极管)等材质制备。
摄像头组件1606用于采集图像或视频。在一个实施例中,摄像头组件1606包括前置摄像头和后置摄像头。通常,前置摄像头用于实现视频通话或自拍,后置摄像头用于实现照片或视频的拍摄。在一些实施例中,后置摄像头为至少两个,分别为主摄像头、景深摄像头、广角摄像头中的任意一种,以实现主摄像头和景深摄像头融合实现背景虚化功能,主摄像头和广角摄像头融合实现全景拍摄以及VR(Virtual Reality,虚拟现实)拍摄功能。在一些实施例中,摄像头组件1606还可以包括闪光灯。闪光灯可以是单色温闪光灯,也可以是双色温闪光灯。双色温闪光灯是指暖光闪光灯和冷光闪光灯的组合,可以用于不同色温下的光线补偿。
音频电路1607用于提供用户和终端之间的音频接口。音频电路1607可以包括麦克风和扬声器。麦克风用于采集用户及环境的声波,并将声波转换为电信号输入至处理器1601进行处理,或者输入至射频电路1604以实现语音通信。出于立体声采集或降噪的目的,麦克风可以为多个,分别设置在终端的不同部位。麦克风还可以是阵列麦克风或全向采集型麦克风。扬声器则用于将来自处理器1601或射频电路1604的电信号转换为声波。扬声器可以是传统的薄膜扬声器,也可以是压电陶瓷扬声器。当扬声器是压电陶瓷扬声器时,不仅可以将电信号转换为人类可听见的声波,也可以将电信号转换为人类听不见的声波以进行测距等用途。在一些实施例中,音频电路1607还可以包括耳机插孔。
定位组件1608用于定位终端的当前地理位置,以实现导航或LBS(Location Based Service,基于位置的服务)。定位组件1608可以是基于美国的GPS(Global Positioning System,全球定位系统)、中国的北斗系统或俄罗斯的伽利略系统的定位组件。
电源1609用于为终端中的各个组件进行供电。电源1609可以是交流电、直流电、一次性电池或可充电电池。当电源1609包括可充电电池时,该可充 电电池可以是有线充电电池或无线充电电池。有线充电电池是通过有线线路充电的电池,无线充电电池是通过无线线圈充电的电池。该可充电电池还可以用于支持快充技术。
在一些实施例中,终端还包括有一个或多个传感器1610。该一个或多个传感器1610包括但不限于:加速度传感器1611、陀螺仪传感器1612、压力传感器1613、指纹传感器1614、光学传感器1615以及接近传感器1616。
本领域技术人员可以理解,图16中示出的结构并不构成对终端的限定,可以包括比图示更多或更少的组件,或者组合某些组件,或者采用不同的组件布置。
本申请实施例还提供一种计算机可读存储介质,所述存储介质中存储有至少一条计算机可读指令、至少一段程序、代码集或计算机可读指令集,所述至少一条计算机可读指令、所述至少一段程序、所述代码集或计算机可读指令集由所述处理器执行以实现上述各个实施例提供的图形码识别方法。
本申请还提供了一种包含计算机可读指令的计算机程序产品,当其在计算机上运行时,使得计算机执行上述各个实施例所述的图形码识别方法。
应该理解的是,虽然上述各实施例的流程图中的各个步骤按照箭头的指示依次显示,但是这些步骤并不是必然按照箭头指示的顺序依次执行。除非本文中有明确的说明,这些步骤的执行并没有严格的顺序限制,这些步骤可以以其它的顺序执行。而且,上述各实施例中的至少一部分步骤可以包括多个子步骤或者多个阶段,这些子步骤或者阶段并不必然是在同一时刻执行完成,而是可以在不同的时刻执行,这些子步骤或者阶段的执行顺序也不必然是依次进行,而是可以与其它步骤或者其它步骤的子步骤或者阶段的至少一部分轮流或者交替地执行。
上述本申请实施例序号仅仅为了描述,不代表实施例的优劣。本领域普通技术人员可以理解实现上述实施例的无线局域网的参数配置方法中全部或部分步骤可以通过硬件来完成,也可以通过程序来计算机可读指令相关的硬件完成,所述的程序可以存储于一种计算机可读存储介质中,上述提到的存储介质可以是只读存储器,磁盘或光盘等。以上所述仅为本申请的较佳实施例,并不用以限制本申请,凡在本申请的精神和原则之内,所作的任何修改、等同替换、改进等,均应包含在本申请的保护范围之内。
Claims (20)
- 一种图形码识别方法,由终端执行,其特征在于,所述方法包括:显示目标图像,所述目标图像中包含至少两个图形码;当接收到对所述目标图像的图形码识别操作时,获取所述目标图像中各个图形码的图形码位置信息;根据所述图形码位置信息,确定所述图形码识别操作指示的目标图形码,所述目标图形码属于所述至少两个图形码;显示所述目标图形码对应的目标图形码识别结果。
- 根据权利要求1所述的方法,其特征在于,所述根据所述图形码位置信息,确定所述图形码识别操作指示的目标图形码,包括:确定所述图形码识别操作指示的目标识别位置;根据所述目标识别位置和所述图形码位置信息,确定所述目标图形码。
- 根据权利要求2所述的方法,其特征在于,所述目标图像为图片,所述图形码识别操作是对所述图片的触发操作;所述确定所述图形码识别操作指示的目标识别位置,包括:将所述图形码识别操作对应的触发位置确定为所述目标识别位置;所述根据所述目标识别位置和所述图形码位置信息,确定所述目标图形码,包括:根据所述触发位置的位置信息以及各个图形码的所述图形码位置信息,确定所述目标识别位置与各个图形码之间的距离;将最短距离对应的图形码确定为所述目标图形码。
- 根据权利要求2所述的方法,其特征在于,所述目标图像为取景框内显示的图像,所述图形码识别操作是对所述目标图像的拍摄操作;所述确定所述图形码识别操作指示的目标识别位置,包括:将取景框中心在所述目标图像中对应的位置确定为所述目标识别位置;所述根据所述目标识别位置和所述图形码位置信息,确定所述目标图形码,包括:根据所述取景框中心的位置信息以及各个图形码的所述图形码位置信息,确定所述目标识别位置与各个图形码之间的距离;将最短距离对应的图形码确定为所述目标图形码。
- 根据权利要求1至4任一所述的方法,其特征在于,所述获取所述目标图像中各个图形码的图形码位置信息,包括:将所述目标图像输入目标检测模型,得到各个图形码的图形码预测位置信息以及位置置信度;根据所述图形码预测位置信息和所述位置置信度,确定各个图形码的所述图形码位置信息。
- 根据权利要求5所述的方法,其特征在于,所述目标检测模型中包括i个串联的残差网络和一个空洞卷积网络,其中,每个所述残差网络中包含一个下采样块和j个第一残差块,且所述第一残差块中包含深度depthwise卷积;所述空洞卷积网络中包含k个第二残差块,所述第二残差块中包含空洞dilated卷积,i,j,k为大于等于2的整数。
- 根据权利要求5所述的方法,其特征在于,所述方法还包括:获取所述目标检测模型输出的各个图形码的图形码预测类型以及类型置信度;根据所述图形码预测类型和所述类型置信度,确定各个图形码的图形码类型;所述显示所述目标图形码对应的目标图形码识别结果之前,所述方法还包括:通过目标图形码类型对应的目标解码器对所述目标图形码进行图形码识别,得到所述目标图形码识别结果,所述目标图形码类型为所述目标图形码对应的图形码类型。
- 根据权利要求5所述的方法,其特征在于,所述方法还包括:所述目标检测模型输出的各个图形码的图形码预测类型以及类型置信度;根据所述图形码预测类型和所述类型置信度,确定各个图形码的图形码类型;所述显示所述目标图形码对应的目标图形码识别结果之前,所述方法还包括:通过各个图形码的图形码类型各自对应的解码器,对各个图形码进行图形码识别,得到各个图形码对应的图形码识别结果;将所述目标图形码对应 的图形码识别结果确定为所述目标图形码识别结果。
- 根据权利要求1所述的方法,其特征在于,所述根据所述图形码位置信息,确定所述图形码识别操作指示的目标图形码,包括:根据所述图形码位置信息对各个图形码进行图形码识别,得到至少两条图形码识别结果;确定所述图形码识别操作对应的目标应用程序;若所述图形码识别结果属于所述目标应用程序支持的识别结果,则将所述图形码识别结果对应的图形码确定为所述目标图形码。
- 一种图形码识别装置,设置于终端,其特征在于,所述装置包括:图像显示模块,用于显示目标图像,所述目标图像中包含至少两个图形码;位置获取模块,用于当接收到对所述目标图像的图形码识别操作时,获取所述目标图像中各个图形码的图形码位置信息;目标确定模块,用于根据所述图形码位置信息,确定所述图形码识别操作指示的目标图形码,所述目标图形码属于所述至少两个图形码;结果显示模块,用于显示所述目标图形码对应的目标图形码识别结果。
- 根据权利要求10所述的装置,其特征在于,所述目标确定模块,包括:第一确定单元,用于确定所述图形码识别操作指示的目标识别位置;第二确定单元,用于根据所述目标识别位置和所述图形码位置信息,确定所述目标图形码。
- 根据权利要求11所述的装置,其特征在于,所述目标图像为图片,所述图形码识别操作是对所述图片的触发操作;所述第一确定单元,用于将所述图形码识别操作对应的触发位置确定为所述目标识别位置;所述第二确定单元,用于根据所述触发位置的位置信息以及各个图形码的所述图形码位置信息,确定所述目标识别位置与各个图形码之间的距离;将最短距离对应的图形码确定为所述目标图形码。
- 根据权利要求11所述的装置,其特征在于,所述目标图像为取景框内显示的图像,所述图形码识别操作是对所述目标图像的拍摄操作;所述第一确定单元,用于将取景框中心在所述目标图像中对应的位置确定为所述目标识别位置;所述第二确定单元,用于根据所述取景框中心的位置信息以及各个图形码的所述图形码位置信息,确定所述目标识别位置与各个图形码之间的距离将最短距离对应的图形码确定为所述目标图形码。
- 根据权利要求10至13任一所述的装置,其特征在于,所述位置获取模块,包括:输入单元,用于将所述目标图像输入目标检测模型,得到各个图形码的图形码预测位置信息以及位置置信度;第三确定单元,用于根据所述图形码预测位置信息和所述位置置信度,确定各个图形码的所述图形码位置信息。
- 根据权利要求14所述的装置,其特征在于,所述目标检测模型中包括i个串联的残差网络和一个空洞卷积网络,其中,每个所述残差网络中包含一个下采样块和j个第一残差块,且所述第一残差块中包含深度depthwise卷积;所述空洞卷积网络中包含k个第二残差块,所述第二残差块中包含空洞dilated卷积,i,j,k为大于等于2的整数。
- 根据权利要求14所述的装置,其特征在于,所述装置还包括:类型获取模块,用于获取所述目标检测模型输出的各个图形码的图形码预测类型以及类型置信度;类型确定模块,用于根据所述图形码预测类型和所述类型置信度,确定各个图形码的图形码类型;所述装置还包括:第一解码模块,用于通过目标图形码类型对应的目标解码器对所述目标图形码进行图形码识别,得到所述目标图形码识别结果,所述目标图形码类型为所述目标图形码对应的图形码类型。
- 根据权利要求14所述的装置,其特征在于,所述装置还包括:类型获取模块,用于所述目标检测模型输出的各个图形码的图形码预测类型以及类型置信度;类型确定模块,用于根据所述图形码预测类型和所述类型置信度,确定各个图形码的图形码类型;所述装置还包括:第二解码模块,用于通过各个图形码的图形码类型各自对应的解码器,对各个图形码进行图形码识别,得到各个图形码对应的图形码识别结果;将所述目标图形码对应的图形码识别结果确定为所述目标图形码识别结果。
- 根据权利要求9所述的装置,其特征在于,所述目标确定模块还包括:结果识别单元,用于根据所述图形码位置信息对各个图形码进行图形码识别,得到至少两条图形码识别结果;第四确定单元,用于确定所述图形码识别操作对应的目标应用程序;第五确定单元,用于若所述图形码识别结果属于所述目标应用程序支持的识别结果,则将所述图形码识别结果对应的图形码确定为所述目标图形码。
- 一种终端,其特征在于,所述终端包括一个或多个处理器和存储器,所述存储器中存储有至少一条计算机可读指令、至少一段程序、代码集或计算机可读指令集,所述至少一条计算机可读指令、所述至少一段程序、所述代码集或计算机可读指令集由所述一个或多个处理器执行以实现如权利要求1至9所述的图形码识别方法。
- 一个或多个计算机可读存储介质,其特征在于,所述存储介质中存储有至少一条计算机可读指令、至少一段程序、代码集或计算机可读指令集,所述至少一条计算机可读指令、所述至少一段程序、所述代码集或计算机可读指令集由所述处理器执行以实现如权利要求1至9任一所述的图形码识别方法。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP19876710.5A EP3767519B1 (en) | 2018-10-22 | 2019-10-10 | Graphic code recognition method and apparatus, and terminal, and storage medium |
| JP2021508059A JP7118244B2 (ja) | 2018-10-22 | 2019-10-10 | グラフィックコード認識方法及び装置、並びに、端末及びプログラム |
| US17/105,119 US11200395B2 (en) | 2018-10-22 | 2020-11-25 | Graphic code recognition method and apparatus, terminal, and storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201811231652.0A CN109409161B (zh) | 2018-10-22 | 2018-10-22 | 图形码识别方法、装置、终端及存储介质 |
| CN201811231652.0 | 2018-10-22 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/105,119 Continuation US11200395B2 (en) | 2018-10-22 | 2020-11-25 | Graphic code recognition method and apparatus, terminal, and storage medium |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020083029A1 true WO2020083029A1 (zh) | 2020-04-30 |
Family
ID=65468230
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/110359 Ceased WO2020083029A1 (zh) | 2018-10-22 | 2019-10-10 | 图形码识别方法、装置、终端及存储介质 |
Country Status (5)
| Country | Link |
|---|---|
| US (1) | US11200395B2 (zh) |
| EP (1) | EP3767519B1 (zh) |
| JP (1) | JP7118244B2 (zh) |
| CN (1) | CN109409161B (zh) |
| WO (1) | WO2020083029A1 (zh) |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113643024A (zh) * | 2021-07-06 | 2021-11-12 | 阿里巴巴新加坡控股有限公司 | 图形码处理方法、装置及电子设备 |
| CN113761961A (zh) * | 2021-09-07 | 2021-12-07 | 杭州海康威视数字技术股份有限公司 | 一种二维码识别方法和装置 |
| US20220262089A1 (en) * | 2020-09-30 | 2022-08-18 | Snap Inc. | Location-guided scanning of visual codes |
| TWI876640B (zh) * | 2023-10-31 | 2025-03-11 | 州巧科技股份有限公司 | 圖形碼品質檢測系統與方法 |
Families Citing this family (21)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109409161B (zh) | 2018-10-22 | 2020-08-11 | 腾讯科技(深圳)有限公司 | 图形码识别方法、装置、终端及存储介质 |
| CN110245537B (zh) * | 2019-05-28 | 2020-10-02 | 北京三快在线科技有限公司 | 图形码解码方法,装置,存储介质及电子设备 |
| CN110222550B (zh) * | 2019-07-11 | 2024-10-29 | 上海肇观电子科技有限公司 | 信息播报方法、电路、播报设备、存储介质、智能眼镜 |
| CN112241640B (zh) * | 2019-07-18 | 2023-06-30 | 杭州海康威视数字技术股份有限公司 | 一种图形码确定方法、装置和工业相机 |
| CN110532826B (zh) * | 2019-08-21 | 2022-09-30 | 厦门壹普智慧科技有限公司 | 一种基于人工智能语义分割的条码识别装置与方法 |
| CN110532825B (zh) * | 2019-08-21 | 2022-09-30 | 厦门壹普智慧科技有限公司 | 一种基于人工智能目标检测的条码识别装置与方法 |
| KR102273198B1 (ko) * | 2019-10-22 | 2021-07-05 | 라인플러스 주식회사 | 시각적으로 코딩된 패턴 인식 방법 및 장치 |
| CN110971820B (zh) * | 2019-11-25 | 2021-03-26 | Oppo广东移动通信有限公司 | 拍照方法、拍照装置、移动终端及计算机可读存储介质 |
| CN111159542B (zh) * | 2019-12-12 | 2023-05-05 | 中国科学院深圳先进技术研究院 | 一种基于自适应微调策略的跨领域序列推荐方法 |
| CN111274842B (zh) * | 2020-02-25 | 2024-05-24 | 维沃移动通信有限公司 | 编码图像的识别方法及电子设备 |
| CN111507122A (zh) * | 2020-04-22 | 2020-08-07 | Oppo广东移动通信有限公司 | 图形码识别方法、装置、存储介质及终端 |
| CN111553673B (zh) * | 2020-05-07 | 2021-09-24 | 支付宝(杭州)信息技术有限公司 | 一种基于图形码识别的信息展示方法及装置 |
| CN112131898A (zh) * | 2020-09-23 | 2020-12-25 | 创新奇智(青岛)科技有限公司 | 一种扫码设备及图形码识别方法 |
| CN112651475B (zh) * | 2021-01-06 | 2022-09-23 | 北京字节跳动网络技术有限公司 | 二维码显示方法、装置、设备及介质 |
| CN113569338B (zh) * | 2021-08-06 | 2022-10-14 | 大连理工大学 | 一种基于时间扩张卷积网络的压气机旋转失速预警方法 |
| CN113627389B (zh) * | 2021-08-30 | 2024-08-23 | 京东方科技集团股份有限公司 | 一种目标检测的优化方法及设备 |
| CN115496084A (zh) * | 2022-09-16 | 2022-12-20 | 阿里巴巴(中国)有限公司 | 图像识别方法、设备、存储介质及程序产品 |
| CN115618905B (zh) * | 2022-10-13 | 2023-12-12 | 东莞市生海科技有限公司 | 一种汽车制造零部件的追溯管理方法及系统 |
| CN117998253A (zh) * | 2022-11-07 | 2024-05-07 | 神基科技股份有限公司 | 语音活动检测装置及方法 |
| JP7323901B1 (ja) | 2022-11-09 | 2023-08-09 | ソノー電機工業株式会社 | 情報処理プログラム及び情報処理端末 |
| CN116306732A (zh) * | 2023-02-10 | 2023-06-23 | 深圳市皮爬爬信息技术有限公司 | 一种适用于近眼显示设备的图形码识别方法及相关装置 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20150016670A1 (en) * | 2013-07-10 | 2015-01-15 | Tencent Technology (Shenzhen) Company Limited | Methods and systems for image recognition |
| CN204155287U (zh) * | 2012-11-15 | 2015-02-11 | 手持产品公司 | 标记读取设备 |
| CN107609437A (zh) * | 2017-08-17 | 2018-01-19 | 阿里巴巴集团控股有限公司 | 一种目标图形码识别方法和装置 |
| CN108537197A (zh) * | 2018-04-18 | 2018-09-14 | 吉林大学 | 一种基于深度学习的车道线检测预警装置及预警方法 |
| CN108563972A (zh) * | 2018-03-09 | 2018-09-21 | 广东欧珀移动通信有限公司 | 图形码识别方法、装置、移动终端及存储介质 |
| CN109409161A (zh) * | 2018-10-22 | 2019-03-01 | 腾讯科技(深圳)有限公司 | 图形码识别方法、装置、终端及存储介质 |
Family Cites Families (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP3675289B2 (ja) | 2000-03-23 | 2005-07-27 | 株式会社デンソー | 情報コード読取装置 |
| US6702183B2 (en) * | 2001-04-17 | 2004-03-09 | Ncr Corporation | Methods and apparatus for scan pattern selection and selective decode inhibition in barcode scanners |
| JP4192847B2 (ja) | 2004-06-16 | 2008-12-10 | カシオ計算機株式会社 | コード読取装置およびプログラム |
| JP2011209805A (ja) * | 2010-03-29 | 2011-10-20 | Konica Minolta Opto Inc | 映像表示装置 |
| US8439260B2 (en) * | 2010-10-18 | 2013-05-14 | Jiazheng Shi | Real-time barcode recognition using general cameras |
| US20130341401A1 (en) * | 2012-06-26 | 2013-12-26 | Symbol Technologies, Inc. | Methods and apparatus for selecting barcode symbols |
| US8985461B2 (en) * | 2013-06-28 | 2015-03-24 | Hand Held Products, Inc. | Mobile device having an improved user interface for reading code symbols |
| JP2016143158A (ja) | 2015-01-30 | 2016-08-08 | パナソニックIpマネジメント株式会社 | 情報提供装置、情報提供方法、及び情報取得プログラム |
-
2018
- 2018-10-22 CN CN201811231652.0A patent/CN109409161B/zh active Active
-
2019
- 2019-10-10 EP EP19876710.5A patent/EP3767519B1/en active Active
- 2019-10-10 JP JP2021508059A patent/JP7118244B2/ja active Active
- 2019-10-10 WO PCT/CN2019/110359 patent/WO2020083029A1/zh not_active Ceased
-
2020
- 2020-11-25 US US17/105,119 patent/US11200395B2/en active Active
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN204155287U (zh) * | 2012-11-15 | 2015-02-11 | 手持产品公司 | 标记读取设备 |
| US20150016670A1 (en) * | 2013-07-10 | 2015-01-15 | Tencent Technology (Shenzhen) Company Limited | Methods and systems for image recognition |
| CN107609437A (zh) * | 2017-08-17 | 2018-01-19 | 阿里巴巴集团控股有限公司 | 一种目标图形码识别方法和装置 |
| CN108563972A (zh) * | 2018-03-09 | 2018-09-21 | 广东欧珀移动通信有限公司 | 图形码识别方法、装置、移动终端及存储介质 |
| CN108537197A (zh) * | 2018-04-18 | 2018-09-14 | 吉林大学 | 一种基于深度学习的车道线检测预警装置及预警方法 |
| CN109409161A (zh) * | 2018-10-22 | 2019-03-01 | 腾讯科技(深圳)有限公司 | 图形码识别方法、装置、终端及存储介质 |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP3767519A4 |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20220262089A1 (en) * | 2020-09-30 | 2022-08-18 | Snap Inc. | Location-guided scanning of visual codes |
| CN113643024A (zh) * | 2021-07-06 | 2021-11-12 | 阿里巴巴新加坡控股有限公司 | 图形码处理方法、装置及电子设备 |
| CN113761961A (zh) * | 2021-09-07 | 2021-12-07 | 杭州海康威视数字技术股份有限公司 | 一种二维码识别方法和装置 |
| CN113761961B (zh) * | 2021-09-07 | 2023-08-04 | 杭州海康威视数字技术股份有限公司 | 一种二维码识别方法和装置 |
| TWI876640B (zh) * | 2023-10-31 | 2025-03-11 | 州巧科技股份有限公司 | 圖形碼品質檢測系統與方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN109409161A (zh) | 2019-03-01 |
| CN109409161B (zh) | 2020-08-11 |
| JP2021520017A (ja) | 2021-08-12 |
| EP3767519A1 (en) | 2021-01-20 |
| US20210150170A1 (en) | 2021-05-20 |
| JP7118244B2 (ja) | 2022-08-15 |
| EP3767519A4 (en) | 2021-08-04 |
| US11200395B2 (en) | 2021-12-14 |
| EP3767519B1 (en) | 2024-12-04 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN109409161B (zh) | 图形码识别方法、装置、终端及存储介质 | |
| KR102444085B1 (ko) | 휴대용 통신 장치 및 휴대용 통신 장치의 영상 표시 방법 | |
| KR102545642B1 (ko) | 효율적인 병렬 광학 흐름 알고리즘 및 gpu 구현 | |
| CN108664829B (zh) | 用于提供与图像中对象有关的信息的设备 | |
| US11095727B2 (en) | Electronic device and server for providing service related to internet of things device | |
| KR102367828B1 (ko) | 이미지 운용 방법 및 이를 지원하는 전자 장치 | |
| KR102173123B1 (ko) | 전자장치에서 이미지 내의 특정 객체를 인식하기 위한 방법 및 장치 | |
| CN113886609B (zh) | 多媒体资源推荐方法、装置、电子设备及存储介质 | |
| CN113076814B (zh) | 文本区域的确定方法、装置、设备及可读存储介质 | |
| KR20220011207A (ko) | 이미지 처리 방법 및 장치, 전자 기기 및 저장 매체 | |
| CN110110787A (zh) | 目标的位置获取方法、装置、计算机设备及存储介质 | |
| US11645758B2 (en) | Object identification in digital images | |
| WO2020048392A1 (zh) | 应用程序的病毒检测方法、装置、计算机设备及存储介质 | |
| CN113888432B (zh) | 一种图像增强方法、装置和用于图像增强的装置 | |
| CN115131789A (zh) | 文字识别方法、设备及存储介质 | |
| KR20230051696A (ko) | 이미지 기반 브라우저 내비게이션 | |
| CN110232417B (zh) | 图像识别方法、装置、计算机设备及计算机可读存储介质 | |
| US11501528B1 (en) | Selector input device to perform operations on captured media content items | |
| KR20230000932A (ko) | 이미지를 분석하는 방법 및 분석 장치 | |
| CN111753813B (zh) | 图像处理方法、装置、设备及存储介质 | |
| CN113835582A (zh) | 一种终端设备、信息显示方法和存储介质 | |
| CN119731639A (zh) | 上下文测试代码生成 | |
| CN112053360A (zh) | 图像分割方法、装置、计算机设备及存储介质 | |
| CN114332118A (zh) | 图像处理方法、装置、设备及存储介质 | |
| US12131221B2 (en) | Fast data accessing system using optical beacons |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19876710 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2019876710 Country of ref document: EP Effective date: 20201014 |
|
| ENP | Entry into the national phase |
Ref document number: 2021508059 Country of ref document: JP Kind code of ref document: A |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
