WO2024179575A1 - 一种数据处理方法、设备以及计算机可读存储介质 - Google Patents
一种数据处理方法、设备以及计算机可读存储介质 Download PDFInfo
- Publication number
- WO2024179575A1 WO2024179575A1 PCT/CN2024/079587 CN2024079587W WO2024179575A1 WO 2024179575 A1 WO2024179575 A1 WO 2024179575A1 CN 2024079587 W CN2024079587 W CN 2024079587W WO 2024179575 A1 WO2024179575 A1 WO 2024179575A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- sample image
- image
- category
- probability
- sample
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/776—Validation; Performance evaluation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/214—Generating training patterns; Bootstrap methods, e.g. bagging or boosting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/047—Probabilistic or stochastic networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/20—Image preprocessing
- G06V10/25—Determination of region of interest [ROI] or a volume of interest [VOI]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/20—Image preprocessing
- G06V10/26—Segmentation of patterns in the image field; Cutting or merging of image elements to establish the pattern region, e.g. clustering-based techniques; Detection of occlusion
- G06V10/267—Segmentation of patterns in the image field; Cutting or merging of image elements to establish the pattern region, e.g. clustering-based techniques; Detection of occlusion by performing operations on regions, e.g. growing, shrinking or watersheds
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/764—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/70—Labelling scene content, e.g. deriving syntactic or semantic representations
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/778—Active pattern-learning, e.g. online learning of image or video features
Definitions
- the present application relates to the field of Internet technology, and in particular to a data processing method, device, and computer-readable storage medium.
- the mainstream image recognition method is mainly based on the machine learning method, that is, an image recognition initial model is trained through a training sample set to obtain an image recognition model.
- the training data can determine the category label recognition accuracy and recognition generalization of the image recognition model.
- the enhancement method for training data is to directly superimpose randomly generated noise on the training sample image.
- the randomly generated noise and the training sample image are independent data from each other, so the noise features in the training sample image superimposed with random noise are independent of the features of the image itself.
- the image recognition model can easily distinguish between the noise features and the features of the image itself. Therefore, the training samples amplified by the prior art can only expand the number of training data, but cannot increase the diversity of training data. Therefore, after the image recognition model is trained again with the training samples amplified by the prior art, the category label recognition accuracy and recognition generalization of the trained image recognition model cannot be improved.
- the embodiments of the present application provide a data processing method, device, and computer-readable storage medium, which can improve the accuracy and generalization of category label recognition.
- An embodiment of the present application provides a data processing method, including:
- a sample image sets and an image recognition model where A is a positive integer greater than 1, the A sample image sets respectively correspond to one of A different category labels, the image recognition model is pre-trained based on the A sample image sets, and the A category labels include error category labels;
- the first probability vector includes A probability elements, each of the A probability elements refers to one of the A category labels, and one of the probability elements is used to represent the probability that the sampling sample image belongs to the referred category label;
- the A probability elements in the first probability vector obtain a first probability element referring to the wrong category label, and adjust the sampled sample image according to the first probability element to obtain an adversarial sample image corresponding to the sampled sample image, and the category prediction result obtained by performing category prediction on the adversarial sample image by the image recognition model is the wrong category label.
- An embodiment of the present application provides a data processing method, including:
- the image recognition optimization model is obtained by continuing to optimize the image recognition model according to the mixed sample training set;
- the mixed sample training set includes A sample image sets, A category labels, correct category labels and adversarial sample images;
- the A sample image sets correspond to one of A different category labels respectively, and the A sample image sets include collected sample images;
- the correct category label belongs to the A category labels;
- A is a positive integer greater than 1;
- the image recognition model is obtained by pre-training the A sample image sets;
- the adversarial sample image is obtained by referring to the wrong category label in the A probability elements in the first probability vector
- the first probability vector is obtained by adjusting the sampled sample image by the first probability element of the signature, the first probability vector is generated by the image recognition model for the sampled sample image, the wrong category label is different from the category label corresponding to the sample image set to which the sampled sample image belongs, and the correct category label is the category label corresponding to the sample image set to which the sampled sample image belongs;
- the first probability vector includes A probability
- the category label corresponding to the maximum probability element in the target probability vector is determined as the category prediction result of the target image.
- An embodiment of the present application provides a data processing device, including:
- a data acquisition module used to acquire A sample image sets and an image recognition model, where A is a positive integer greater than 1, the A sample image sets respectively correspond to one of A different category labels, the image recognition model is pre-trained based on the A sample image sets, and the A category labels include error category labels;
- a first input module is used to obtain a sampled sample image from the A sample image sets, and input the sampled sample image and the erroneous category label into the image recognition model to generate a first probability vector of the sampled sample image for the A category labels, wherein the erroneous category label is different from the category label corresponding to the sample image set to which the sampled sample image belongs;
- the first probability vector includes A probability elements, each of the A probability elements refers to one of the A category labels, and one of the probability elements is used to represent the probability that the sampled sample image belongs to the referred category label;
- the first adjustment module is used to obtain a first probability element referring to the wrong category label from the A probability elements in the first probability vector, and adjust the sampled sample image according to the first probability element to obtain an adversarial sample image corresponding to the sampled sample image, wherein the category prediction result obtained by the image recognition model performing category prediction on the adversarial sample image is the wrong category label.
- the first adjustment module is specifically used to generate a negative probability element corresponding to the first probability element, sum the negative probability element and the maximum probability element in the first probability vector, obtain the label loss value of the sampling sample image for the wrong category label, and adjust the sampling sample image according to the label loss value to obtain the adversarial sample image corresponding to the sampling sample image.
- the first adjustment module is used to adjust the sample image according to the label loss value to obtain the adversarial sample image corresponding to the sample image. It is specifically used to generate an initial gradient value for the sample image according to the label loss value, obtain the numerical sign of the initial gradient value, generate a symbol unit value corresponding to the initial gradient value according to the numerical sign, multiply the initial learning rate and the symbol unit value to obtain a first value to be cropped, crop the first value to be cropped through a first cropping interval generated by the attack intensity to obtain a gradient value for adjusting the sample image, perform difference processing on the sample image and the gradient value to obtain a second value to be cropped, crop the second value to be cropped through a second cropping interval generated by the pixel level, and obtain the adversarial sample image corresponding to the sample image.
- the first adjustment module includes:
- a first adjustment unit configured to adjust the sample image according to the first probability element to obtain an initial adversarial sample image corresponding to the sample image
- a first generating unit configured to input the erroneous category label and the initial adversarial sample image into the image recognition model, and generate a second probability vector of the initial adversarial sample image for A category labels through the image recognition model;
- a first acquisition unit configured to acquire a second probability element for the error category label from the A probability elements included in the second probability vector
- the first adjustment unit is further configured to continue adjusting the initial adversarial sample image according to the second probability element if the second probability element does not satisfy the boundary constraint condition or the category prediction result obtained by performing category prediction on the adversarial sample image is not the wrong category label;
- the first acquisition unit is further used to determine the initial adversarial sample image as the adversarial sample image corresponding to the sampling sample image if the second probability element satisfies the boundary constraint condition and the category prediction result obtained by performing category prediction on the adversarial sample image is the wrong category label; when the category prediction result is the wrong category label, the second probability element is the maximum value of the A probability elements contained in the second probability vector.
- the first acquisition unit is further used to compare the second probability element with a probability constraint value; the probability constraint value is the reciprocal of A;
- the first acquisition unit is further configured to determine that the second probability element does not satisfy a boundary constraint condition if the second probability element is less than the probability constraint value;
- the first acquisition unit is further configured to determine that the second probability element satisfies a boundary constraint condition if the second probability element is equal to or greater than the probability constraint value.
- the data processing device further includes:
- a second input module is used to input the A sample image sets, the A category labels, the correct category label and the adversarial sample image into the image recognition model;
- the correct category label refers to the category label corresponding to the sample image set to which the sample image belongs;
- the second input module is further used to generate, through the image recognition model, first prediction categories corresponding to the A sample image sets and second prediction categories corresponding to the adversarial sample images;
- the second adjustment module is used to generate a first loss value according to the first predicted category and the A category labels, and
- the image recognition optimization model is obtained by adjusting the parameters in the image recognition model according to the first loss value and the second loss value; and the image recognition optimization model performs category prediction on the adversarial sample image and obtains the correct category label.
- the A sample image sets include a sample image set Cb; b is a positive integer and b is less than or equal to A; the first input module includes:
- a first determining unit is used to determine a sampling ratio, and perform sampling processing on sample images in the sample image set Cb according to the sampling ratio to obtain sampled sample images;
- the second acquisition unit is used to acquire a category label different from the category label of the sample image set Cb from among the A category labels, and determine the acquired category label as an erroneous category label.
- the data acquisition module includes:
- a third acquisition unit is used to acquire A sample image sets, and input the A category labels and the A sample image sets into the image recognition initial model respectively;
- a second generating unit is used to generate predicted initial categories corresponding to the A sample image sets respectively through an image recognition initial model
- a second determining unit configured to determine, according to the predicted initial category and the A category labels, category loss values corresponding to the A sample image sets respectively;
- a third determining unit is used to determine a total loss value corresponding to the image recognition initial model according to the category loss values corresponding to the A sample image sets respectively;
- the second adjustment unit is used to adjust the parameters in the initial image recognition model according to the total loss value to obtain an image recognition model;
- the category prediction result obtained by the image recognition model for the sample image set Cb is the category label of the sample image set Cb;
- the A sample image sets include the sample image set Cb;
- b is a positive integer and b is less than or equal to A.
- the third acquisition unit includes:
- the image acquisition subunit is used to acquire A original image sets, and input the A original image sets into the object detection model; the A original image sets correspond to one of A different category information respectively; the A original image sets include an original image set Db; the original image set Db includes an original image Ef; f is a positive integer, and f is less than or equal to the total number of original images in the original image set Db;
- the second determination subunit is used to determine the region coordinates of the key region in the original image Ef through the object detection model, and generate a sample image to be labeled corresponding to the original image Ef according to the region coordinates;
- a third determining subunit is used to determine the sample image to be labeled corresponding to each original image in the original image set Db as the sample image set to be labeled;
- the fourth determining subunit is used to generate a category label Hb according to the category information corresponding to the original image set Db, and determine the set of sample images to be labeled with the category label Hb as the sample image set Cb.
- the second determining subunit includes:
- the third processing subunit is used to generate an initial detection frame including the key area according to the area coordinates, and to expand the initial detection frame to obtain a detection frame including the target object; the key area belongs to the target object;
- a fourth processing subunit is used to obtain, in the detection frame, an image to be scaled including the target object, obtain a first image size, and scale the image to be scaled according to the first image size to obtain an image to be cropped;
- the fifth processing subunit is used to obtain a second image size, and perform cropping processing on the image to be cropped according to the second image size to obtain a sample image to be labeled corresponding to the original image Ef; the second image size is smaller than the first image size.
- An embodiment of the present application provides a data processing device, including:
- An image acquisition module is used to acquire a target image and input the target image into an image recognition optimization model; the image recognition optimization model is obtained by continuing to optimize the image recognition model according to a mixed sample training set; the mixed sample training set includes A sample image sets, A category labels, correct category labels and adversarial sample images; the A sample image sets correspond to one of A different category labels respectively, and the A sample image sets include collected sample images; the correct category label belongs to the A category labels; A is a positive integer greater than 1; the image recognition model is obtained by pre-training the A sample image sets; the adversarial sample image is obtained by referring to the wrong category label in the A probability elements in the first probability vector The first probability element of the wrong category label is adjusted to obtain the sampled sample image, the first probability vector is generated by the image recognition model for the sampled sample image, the wrong category label is different from the category label corresponding to the sample image set to which the sampled sample image belongs, and the correct category label is the category label corresponding to the sample image set to which the sampled sample image belongs; the first probability vector
- a vector generation module is used to generate a target probability vector for a target image with respect to A category labels through an image recognition optimization model
- the category determination module is used to determine the category label corresponding to the maximum probability element in the target probability vector as the category prediction result of the target image.
- the present application provides a computer device, including: a processor, a memory, and a network interface;
- the above-mentioned processor is connected to the above-mentioned memory and the above-mentioned network interface, wherein the above-mentioned network interface is used to provide a data communication function, the above-mentioned memory is used to store a computer program, and the above-mentioned processor is used to call the above-mentioned computer program so that the computer device executes the method in the embodiment of the present application.
- an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored.
- the computer program is suitable for being loaded by a processor and executing the method in the embodiment of the present application.
- an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium; a processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method in the embodiment of the present application.
- a computer device can obtain A sample image sets and an image recognition model; wherein the A sample image sets respectively correspond to one of A different category labels, and the image recognition model is pre-trained based on the A sample image sets, so the image recognition model can accurately determine the category labels of the sample images in the A sample image sets; further, the computer device obtains a sampled sample image from the A sample image sets, and inputs an erroneous category label and the sampled sample image into the image recognition model, wherein the erroneous category label belongs to the A category labels, and the erroneous category label is different from the category label corresponding to the sample image set to which the sampled sample image belongs; a first probability vector of the sampled sample image for the A category labels can be generated through the image recognition model, and the first probability vector includes A probability elements, and the A probability elements respectively refer to one of the A category labels.
- a probability element is used to characterize the probability that the sampled sample image belongs to the referred category label; since the image recognition model can accurately determine the category labels of the sample images in the A sample image sets, in the first probability vector, the probability element referring to the correct category label is the largest, or in other words, the probability element referring to the correct category label is much larger than the first probability element referring to the wrong category label, and the correct category label refers to the category label corresponding to the sample image set to which the sampled sample image belongs; further, according to the first probability element, the computer device can adjust the sampled sample image to obtain an adversarial sample image corresponding to the sampled sample image, wherein the category prediction result obtained by the image recognition model for category prediction of the adversarial sample image is the wrong category label, that is, by adjusting the sampled sample image, the image recognition model can incorrectly determine the category label of the adversarial sample image.
- an embodiment of the present application proposes a method for generating adversarial sample images.
- adversarial sample images that are incorrectly recognized by an image recognition model can be generated.
- the adversarial sample images can improve the sample diversity of a training set used to optimize the training of the image recognition model.
- the optimization training accuracy of the image recognition model can be improved, thereby improving the category label recognition accuracy and recognition generalization of the optimized trained image recognition model.
- FIG1 is a schematic diagram of a system architecture provided by an embodiment of the present application.
- FIG2a is a schematic diagram of a data processing scenario provided by an embodiment of the present application.
- FIG2b is a second schematic diagram of a data processing scenario provided by an embodiment of the present application.
- FIG3 is a flowchart of a data processing method provided in an embodiment of the present application.
- FIG4 is a schematic diagram of a classification boundary surface trained by clean sample data provided in an embodiment of the present application.
- FIG5 is a schematic diagram of a classification boundary surface trained by first mixed sample data provided by an embodiment of the present application.
- FIG6 is a schematic diagram of a classification boundary surface trained by second mixed sample data provided in an embodiment of the present application.
- FIG7 is a second flow chart of a data processing method provided in an embodiment of the present application.
- FIG8 is a third flow chart of a data processing method provided in an embodiment of the present application.
- FIG9 is a first structural diagram of a data processing device provided in an embodiment of the present application.
- FIG10 is a second structural diagram of a data processing device provided in an embodiment of the present application.
- FIG. 11 is a schematic diagram of the structure of a computer device provided in an embodiment of the present application.
- Artificial Intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
- artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence.
- Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
- Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies.
- Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operating/interactive systems, mechatronics and other technologies.
- Artificial intelligence software technologies mainly include computer vision technology, speech processing technology, natural language processing technology, as well as machine learning/deep learning, autonomous driving, smart transportation and other major directions.
- Computer vision technology is a science that studies how to make machines "see”. To put it more specifically, it refers to machine vision such as using cameras and computers to replace human eyes to identify and measure targets, and further perform graphic processing to make computer processing into images that are more suitable for human eye observation or transmission to instruments for detection.
- Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content/behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and mapping, automatic driving, smart transportation and other technologies.
- OCR optical character recognition
- computer vision technology can be used to identify category labels in images (such as people, cats, dogs, or real people, paper people, mold people, etc.).
- Machine Learning is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithmic complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance.
- Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence.
- Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
- the image recognition model and the object detection model are both AI models based on machine learning.
- the image recognition model can be used to identify and process images, and the object detection model can be used to detect key areas of target objects in images.
- Figure 1 is a schematic diagram of a system architecture provided by an embodiment of the present application.
- the system may include a business server 100 and a terminal device cluster, and the terminal device cluster may include: terminal device 200a, terminal device 200b, terminal device 200c, ..., terminal device 200n. It can be understood that the above system may include one or more terminal devices, and the present application does not limit the number of terminal devices.
- any terminal device in the terminal device cluster may have a communication connection with the business server 100, for example, there is a communication connection between the terminal device 200a and the business server 100, wherein the above communication connection does not limit the connection method, and may be directly or indirectly connected by wired communication, directly or indirectly connected by wireless communication, or by other methods, which are not limited in this application.
- each terminal device in the terminal device cluster shown in FIG1 may be installed with an application client, and when the application client runs in each terminal device, it may respectively perform data interaction with the business server 100 shown in FIG1, i.e., the above-mentioned communication connection.
- the application client may be an application client having a target image recognition function, such as a video application, a live broadcast application, a social application, an instant messaging application, a game application, a music application, a shopping application, a novel application, a browser, etc.
- the application client may be an independent client or an embedded sub-client integrated in a client (e.g., a social client, an educational client, a multimedia client, etc.), which is not limited here.
- the business server 100 may be a collection of multiple servers including a background server corresponding to the shopping application, a data processing server, etc. Therefore, each terminal device may perform data transmission with the business server 100 through the application client corresponding to the shopping application. For example, each terminal device may upload a target image to the business server 100 through the application client of the shopping application, and then the business server 100 may determine the category label of the target image.
- the embodiment of the present application can select a terminal device as the target terminal device in the terminal device cluster shown in FIG1, for example, the terminal device 200a is used as the target terminal device.
- the terminal device 200a can send a model optimization request for the image recognition model to the business server 100 through the application client.
- the image recognition model refers to a neural network model used to recognize an image and determine the category information of the image.
- the business server 100 can obtain A sample image sets and an image recognition model, wherein the A sample image sets are multiple sample image sets, and the A sample image sets correspond to different category labels respectively. Therefore, the number of category labels can also be A, that is, the A sample image sets correspond to one of the A different category labels.
- a category label is used to indicate a category information, so the category information corresponding to the A sample image sets is different from each other.
- the embodiment of the present application does not limit the category information and can be set according to the actual application scenario.
- the category information can be object information. In this case, the object information corresponding to the A sample image sets is different from each other.
- the object information of the first sample image set in the A sample image set is a cat (that is, each sample image in the first sample image set is related to a cat), and the object information of the second sample image set in the A sample image set is a dog (that is, each sample image in the second sample image set is related to a dog).
- the category information may be object carrier information.
- the object information corresponding to each of the A sample image sets may be the same or different, but the object carrier information corresponding to each of the A sample image sets is different from each other.
- the object carrier information of the third sample image set among the A sample image sets is a copied screen
- the object carrier information of the fourth sample image set among the A sample image sets is a key feature mold
- the object carrier information of the fifth sample image set among the A sample image sets is a piece of paper.
- the embodiment of the present application does not limit the way in which the business server 100 obtains A sample image sets and image recognition models, and can be set according to the actual application scenario.
- the model optimization request sent by the terminal device 200a carries A sample image sets or image recognition models, or carries A sample image sets and image recognition models at the same time, such as obtaining A sample image sets and image recognition models from a database, such as obtaining A sample image sets and image recognition models from a blockchain network.
- the image recognition model is obtained by pre-training based on A sample image sets. The generation process of the image recognition model is not described here. Please refer to the description of step S101 in the embodiment corresponding to Figure 3 below.
- the embodiment of the present application takes the sample image set C1 in the A sample image set as an example for description. It can be understood that the processing process of the remaining sample image sets in the A sample image set is the same as the processing process of the sample image set C1 described below.
- the business server 100 obtains the sampled sample images in the sample image set C1 (can be obtained randomly or by polling), and then obtains the wrong category label in the A category labels.
- the wrong category label belongs to the category label in the A category labels except the category label of the sample image set C1 (can be called the remaining category label).
- the category label randomly obtained in the remaining category label can be determined as the wrong category label, for example, the wrong category label is the category label of the sample image set C2.
- the wrong category label and the sampled sample image can be input into the image recognition model, wherein the wrong category label and the sampled sample image can be respectively input into the image recognition model as two independent data, and there is a mapping relationship between the wrong category label and the sampled sample image; or, the wrong category label can be added to the file header or file footer of the sampled sample image, and then the sampled sample image carrying the wrong category label can be input into the image recognition model.
- the business server 100 can generate a first probability vector for the sampled sample image for A category labels through the image recognition model; the first probability vector includes A probability elements, each of which refers to one of the A category labels, and one of the probability elements is used to characterize the probability that the sampled sample image belongs to the referred category label.
- the vector dimension of the first probability vector is equal to A, that is, the elements on a vector dimension (also referred to as probability elements in the embodiment of the present application) can characterize the predicted probability value of a category label, that is, the larger the value of a certain probability element, the greater the possibility that the sampled sample image belongs to the category label referred to by the probability element. Therefore, the category label referred to by the probability element with the largest value in the first probability vector can usually be determined as the category prediction result using the sample image. Therefore, in the first probability vector, the business server 100 can obtain the first probability element that refers to the wrong category label, and further, according to the first probability element, the sampled sample image can be adjusted to obtain the adversarial sample image corresponding to the sampled sample image.
- the category prediction result obtained by the image recognition model for the adversarial sample image is the wrong category label, that is, the image recognition model cannot accurately identify the adversarial sample image, and then obtains an incorrect category classification.
- the sample image is an image containing a cat
- its category label is the label corresponding to the cat.
- the corresponding adversarial sample image is obtained.
- the optimizer of the image recognition model can accurately judge that the adversarial sample image is an image containing a cat, but the image recognition model will mistakenly judge that the adversarial sample image is an image containing a dog.
- the business server 100 can generate adversarial sample images. Subsequently, through A sample image sets, correct category labels and adversarial sample images, the image recognition model can be optimized and trained to obtain an image recognition optimization model.
- the correct category label refers to the category label of the sample image set C1. It can be understood that when training the image recognition model, the model loss is calculated by the predicted category label and the correct category label of the adversarial sample image. Therefore, the trained image recognition optimization model can accurately identify the adversarial sample image, and then accurately identify that the predicted category result of the adversarial sample image is the correct category label. Therefore, the image recognition optimization model can have a stronger anti-interference ability.
- the business server 100 may generate a model optimization completion message for the image recognition optimization model, and send the model optimization completion message to the terminal device 200a.
- the terminal device 200a may obtain the image recognition optimization model.
- the embodiment of the present application does not limit the manner in which the terminal device 200a obtains the image optimization recognition model, and the terminal device 200a may obtain the image optimization recognition model according to the actual application.
- the terminal device 200a stores the above image recognition model and A sample image sets locally, and the terminal device 200a has offline computing capabilities, then when the terminal device 200a receives the model optimization instruction for the image recognition model, it can obtain the sample images in the sample image set C1 locally and use the obtained sample images as the sampled sample images.
- the terminal device 200a inputs the error category label and the sampled sample images into the image recognition model, and the subsequent processing process is the same as the process described above, so it is not repeated here.
- the embodiment of the present application assigns an incorrect category label to the sampled sample image, and obtains a first probability element referring to the incorrect category label through the image recognition model, and adjusts the sampled sample image through the first probability element to obtain an adversarial sample image that allows the image recognition model to incorrectly predict the category.
- the embodiment of the present application uses adversarial sample images to amplify and enhance the training samples of A sample image sets (i.e., training data), so it can improve the richness of the sample training set used to optimize the training image recognition model, thereby improving the recognition accuracy of the image recognition optimization model.
- terminal device 200a can all be blockchain nodes in the blockchain network, and the data described in the full text (such as A sample image sets and image recognition models) can be stored.
- the storage method can be that the blockchain node generates blocks according to the data and adds the blocks to the blockchain for storage.
- Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism and encryption algorithm. It is mainly used to organize data in chronological order and encrypt it into a ledger to make it impossible to be tampered with or forged. At the same time, it can verify, store and update data.
- Blockchain is essentially a decentralized database. Each node in the database stores the same blockchain.
- the blockchain network can divide nodes into core nodes, data nodes and light nodes. Core nodes, data nodes and light nodes together constitute blockchain nodes. Among them, the core node is responsible for the consensus of the entire blockchain network, that is, the core node is the consensus node in the blockchain network.
- the process of writing transaction data in the blockchain network into the ledger can be that the data node or light node in the blockchain network obtains the transaction data, and passes the transaction data in the blockchain network (that is, the node passes it in the form of a baton) until the consensus node receives the transaction data, and the consensus node then packs the transaction data into a block, executes consensus on the block, and writes the transaction data into the ledger after the consensus is completed.
- a sample image sets and an image recognition model are used as examples of transaction data.
- the business server 100 (blockchain node) generates a block based on the transaction data and stores the block in the blockchain network.
- the blockchain node can obtain the block containing the transaction data in the blockchain network, and further obtain the transaction data in the block.
- the method provided in the embodiments of the present application can be executed by a computer device, which includes but is not limited to a terminal device or a business server.
- the business server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud databases, cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
- Terminal devices include but are not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc.
- the terminal device and the business server can be directly or indirectly connected by wire or wireless means, and the embodiments of the present application are not limited here.
- Figure 2a is a schematic diagram of a data processing scenario provided by an embodiment of the present application.
- the embodiment of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, audio and video, etc.
- the embodiment of the present application can be applied to business scenarios such as image search scenarios, image generation scenarios, image recommendation scenarios, image distribution scenarios, and specific business scenarios will not be listed one by one here.
- the implementation process of the data processing scenario can be carried out in a business server, or in a terminal device, or interactively in a terminal device and a business server, without limitation here.
- the terminal device can be any terminal device in the terminal device cluster of the embodiment corresponding to Figure 1 above
- the business server can be the business server 100 of the embodiment corresponding to Figure 1 above.
- the embodiment of the present application is described by taking the operation in the business server as an example.
- the business server 100 can obtain A sample image sets 20a and image recognition models 20d, where A is a positive integer greater than 1.
- A is a positive integer greater than 1.
- the embodiment of the present application takes A as 4, that is, the number of sample image sets is 4, such as the sample image set 201a, the sample image set 202a, the sample image set 203a, and the sample image set 204a shown in FIG2a (the number of sample images contained in the sample image set 201a, the sample image set 202a, the sample image set 203a, and the sample image set 204a can be the same or different).
- the category information corresponding to the four sample image sets is different from each other.
- the category information of the sample image set 201a shown in FIG2a is radish
- the category information of the sample image set 202a is pepper
- the category information of the sample image set 203a is pumpkin
- the category information of the sample image set 204a is onion. Therefore, the business server 100 can set different category labels for the above four sample image sets, respectively, setting the category label of radish to 0, that is, the category label of sample image set 201a is 0, setting the category label of pepper to 1, that is, the category label of sample image set 202a is 1, setting the category label of pumpkin to 2, that is, the category label of sample image set 203a is 2, and setting the category label of onion to 3, that is, the category label of sample image set 204a is 3.
- the embodiment of the present application is for the purpose of description and understanding, and uses radish, pepper, pumpkin and onion as examples of category information, and in actual application, the category information should be set according to the actual scenario.
- the embodiment of the present application is for the purpose of description and understanding, and uses 0-3 numerical examples of category labels, and in actual application, the category labels should be set according to the actual scenario.
- sample image 201b is obtained from sample image set 201a
- sample image 202b is obtained from sample image set 202a
- sample image 203b is obtained from sample image set 203a
- sample image 204b is obtained from sample image set 204a.
- the embodiment of the present application does not limit the total number of sample images 201b, which may be one or more.
- the embodiment of the present application does not limit the total number of sample images 202b, sample images 203b, and sample images 204b, which may be one or more.
- the service server 100 inputs the sample image 201b, the sample image 202b, the sample image 203b and the sample image 204b into the image recognition model 20d, wherein the image recognition model 20d is pre-trained based on the above four sample image sets. It can be understood that the service server 100 uses the image recognition model 20d to process each sample image in the same manner, so the following description is made using a sample image (such as the sample image 201c shown in FIG. 2a) as an example, and the processing of the remaining sample images can refer to the processing of the sample image 201c below.
- a sample image such as the sample image 201c shown in FIG. 2a
- the sample image 201c and the wrong category label can be input into the image recognition model 20d.
- the sample image 201c is an image with the category information of radish, that is, the sample image 201c belongs to the sample image set 201a.
- the correct category label of the sample image 201c is the category label of radish, that is, 0.
- FIG2a illustrates the wrong category label of the sample image 201c with the category label of pepper (that is, 1) (the wrong category label can be obtained by randomly selecting from pepper, pumpkin and onion). It should be noted that the wrong category label of the sample image 201c is randomly generated.
- the business server 100 can generate the first probability vector 201e for the four category labels of the sample image 201c through the image recognition model 20d.
- the dimension of the first probability vector 201e is the same as the total number of category labels, and the element value on each dimension can represent the predicted probability value of the corresponding category label. Therefore, in Figure 2a, 0.9 on the first dimension of the first probability vector 201e can represent the predicted probability of category label 0, 0.02 on the second dimension can represent the predicted probability of category label 1, and 0.03 on the third dimension can represent the predicted probability of category label 2; 0.05 on the fourth dimension can represent the predicted probability of category label 3.
- the predicted probabilities on these four dimensions can also be called four probability elements.
- the image recognition model 20d can determine that the numerical value of the predicted probability referring to category label 0 is the largest, so it can be determined that the category prediction result of the sample image 201c is category label 0, which is different from the wrong category label (i.e. 1) input to the image recognition model 20d.
- This result is produced because the image recognition model 20d is pre-trained by the sample image set 201a, the sample image set 202a, the sample image set 203a and the sample image set 204a in Figure 2a, and the sample sample image 201c comes from the sample image set 201a, so the image recognition model 20d recognizes that the category prediction result of the sample sample image 201c is category label 0.
- the service server 100 can obtain the first probability element referring to the wrong category label (category label 1 as shown in FIG2a ) in the first probability vector 201e through the image recognition model 20d, such as 0.02 as shown in FIG2a .
- the service server 100 can adjust the sample image 201c according to the first probability element to obtain the adversarial sample image corresponding to the sample image 201c.
- the sampling sample image 201c is adjusted to obtain the specific process of the adversarial sample image corresponding to the sampling sample image 201c, which may include: the business server 100 may adjust the sampling sample image 201c according to the first probability element to obtain the initial adversarial sample image corresponding to the sampling sample image 201c.
- Figure 2b is a schematic diagram of a data processing scenario provided by an embodiment of the present application.
- the business server 100 generates a negative probability element corresponding to the first probability element (0.02 as shown in Figure 2a), that is, -0.02 in Figure 2b, and sums the negative probability element and the maximum probability element in the first probability vector 201e (0.9 as shown in Figure 2a and Figure 2b) to obtain the label loss value of the sampling sample image 201c for the wrong category label, that is, 0.88 in Figure 2b. Further, the business server 100 adjusts the sampling sample image 201c according to the label loss value to obtain the initial adversarial sample image 202c corresponding to the sampling sample image 201c.
- step S103 The specific process of adjusting the sample image 201c according to the label loss value to obtain the initial adversarial sample image 202c corresponding to the sample image 201c is not described in detail in the embodiment of the present application. Please refer to the description of step S103 in the embodiment corresponding to Figure 3 below.
- the business server 100 inputs the wrong category label (category label 1 as shown in Figures 2a and 2b) and the initial adversarial sample image 202c into the image recognition model 20d.
- the business server 100 can generate a second probability vector 202e for the initial adversarial sample image 202c for the above four category labels through the image recognition model 20d.
- the meaning of the second probability vector 202e is the same as that of the first probability vector 201e, so the meaning of the second probability vector 202e is not explained here. Please refer to the description of the meaning of the first probability vector 201e above.
- the service server 100 obtains a second probability element indicating an error category label (i.e., category label 1 as shown in FIG. 2a and FIG. 2b ) in the second probability vector 202e, such as 0.08 as shown in FIG. 2b .
- the service server 100 can, based on the second probability element, Determine the adversarial sample image corresponding to the sample image 201c. That is, the business server 100 can analyze whether the second probability element is the probability element with the largest value in the second probability vector 202e.
- the initial adversarial sample image 202c can be determined as the adversarial sample image corresponding to the sample image 201c (that is, at this time, the image recognition model 20d can mistakenly identify the category prediction result of the adversarial sample image as the wrong category label 1); if it is analyzed that the second probability element is not the probability element with the largest value in the second probability vector 202e, the initial adversarial sample image 202c will be further adjusted.
- the purpose of the adjustment is to increase the probability element referring to the wrong category label in the third probability vector output by the image recognition model 20d for the adjusted initial adversarial sample image, until the image recognition model 20d can mistakenly identify the adversarial sample image after multiple rounds of adjustment as the wrong category label 1.
- the category prediction result of the image recognition model 20d for the adversarial sample image is the wrong category label 1, that is, the adversarial sample image is still an image containing radish, but the image recognition model 20d will mistakenly identify the category prediction result of the adversarial sample image as the category label 1 of pepper.
- the business server 100 can continue to optimize the training of the image recognition model 20d through the correct category label (that is, the category label 0 of radish) and the adversarial sample image to obtain an optimized image recognition model.
- the image optimization recognition model that is, the optimized trained image recognition model
- the embodiment of the present application proposes a method for generating adversarial sample images, through which adversarial sample images that are incorrectly recognized by an image recognition model can be generated.
- adversarial sample images the sample diversity of the training set used to optimize the training of the image recognition model can be improved.
- the optimization training accuracy of the image recognition model can be improved, and then the category label recognition accuracy and recognition generalization of the optimized trained image recognition model can be improved.
- the adversarial sample image is an image with strong disturbance (i.e., an image that will cause the image recognition model to incorrectly recognize it), and the image optimization recognition model can overcome the strong disturbance, i.e., the image optimization recognition model can accurately identify the category prediction result of the adversarial sample image as the correct category label (i.e., category label 0) through a stronger anti-interference ability.
- Figure 3 is a flow chart of a data processing method provided in an embodiment of the present application.
- the data processing method can be executed by a business server (for example, the business server 100 shown in Figure 1 above), or by a terminal device (for example, the terminal device 200a shown in Figure 1 above), or by a business server and a terminal device interacting with each other.
- a business server for example, the business server 100 shown in Figure 1 above
- a terminal device for example, the terminal device 200a shown in Figure 1 above
- the embodiment of the present application is described by taking the method being executed by a business server as an example.
- the data processing method can at least include the following steps S101-S103.
- Step S101 obtain A sample image sets and an image recognition model, where A is a positive integer greater than 1, and the A sample image sets respectively correspond to one of A different category labels.
- the image recognition model is pre-trained based on the A sample image sets, and the A category labels include error category labels.
- the business server obtains A sample image sets, and the category information corresponding to the A sample image sets is different.
- the category information of a sample image set is radish (that is, each sample image in the sample image set is related to radish), and the category information of another sample image set is pepper (that is, each sample image in the sample image set is related to pepper). Therefore, the business server can set different category labels for each sample image set, such as setting the category label of radish to 0, that is, the category label of the sample image set with the category information of radish is 0; and setting the category label of pepper to 1, that is, the category label of the sample image set with the category information of pepper is 1.
- the business server inputs A sample image sets into the image recognition initial model respectively; generates the predicted initial category corresponding to each sample image set through the image recognition initial model; determines the category loss value corresponding to each sample image set according to the predicted initial category corresponding to each sample image set and the category label corresponding to each sample image set; determines the total loss value corresponding to the image recognition initial model according to the category loss values corresponding to the A sample image sets; adjusts the parameters in the image recognition initial model according to the total loss value to obtain the image recognition model; the image recognition model has the ability to accurately identify the category prediction results of the A sample image sets, such as the image recognition model can accurately identify the category prediction result of the sample image set with category information of radish as the category label of 0 (the category prediction result can also be understood as the category prediction result for each sample image in the sample image set).
- the embodiment of the present application does not limit the number of sample images in the sample image set, which can be one or more. It can be understood that the category information corresponding to the sample images in the same sample image set is the same, so the category labels corresponding to the sample images in the same sample image set are the same. The embodiment of the present application also does not limit the number of sample image sets, which can be two or more.
- the category label in the embodiment of the present application refers to a label used to identify category information.
- the embodiment of the present application does not limit the category information and can be set according to the actual application scenario.
- the category information can be object information.
- each sample image set contains different object information, such as cats, dogs, cars, etc.
- the category information can be object carrier information.
- each sample image contains different object carrier information, such as mask objects, mold objects, photo objects, paper objects, real objects, etc.
- the embodiment of the present application does not limit the model type of the image recognition initial model, and can be composed of any one or more neural networks, such as convolutional neural networks (CNN), residual networks (ResNet), Convolutional neural networks with extended channel number learning mechanism (Wide Residual Networks, Wide-ResNet), High-Resolution Net (HRNet), etc.
- CNN convolutional neural networks
- ResNet residual networks
- HRNet High-Resolution Net
- the specific process of the business server training the image recognition initial model can be: inputting A sample image sets into the image recognition initial model respectively.
- the image recognition initial model processes each sample image in the A sample image sets in the same way, so the following description is taken as an example to describe the first sample image carrying the first category label, and the processing process of the remaining sample images is described below.
- the first sample image is any sample image in the A sample image sets. If the first sample image belongs to the sample image set C1, the first category label is the category label corresponding to the sample image set C1; if the first sample image belongs to the sample image set C2, the first category label is the category label corresponding to the sample image set C2; other situations are as described above and will not be repeated one by one.
- the business server can obtain the first initial probability vector corresponding to the first sample image through the image recognition initial model.
- the dimension of the first initial probability vector is the same as A (that is, one dimension corresponds to one category label).
- the element value on each dimension in the first initial probability vector (also referred to as a probability element) represents the predicted probability of the corresponding category label, and the sum of the element values on each dimension is 1.
- the category label corresponding to the largest element value in the first initial probability vector that is, the predicted initial category corresponding to the first sample image
- the business server can generate a category loss value corresponding to the first sample image based on the predicted initial category corresponding to the first sample image and the first category label.
- the business server can obtain the category loss value corresponding to each sample image in the A sample image set, sum the obtained category loss values, and obtain the total loss value corresponding to the image recognition initial model.
- the business server adjusts the parameters in the image recognition initial model according to the total loss value to obtain a pre-trained image recognition model, which can accurately identify the category prediction results of each sample image in the A sample image set, such as the image recognition model can accurately identify the category prediction result of the first sample image as the first category label.
- the embodiment of the present application does not limit the adjustment method of the parameters in the image recognition initial model, and can be set according to the actual application scenario.
- the embodiment of the present application does not limit the number of iterations (epochs) of the initial image recognition model, which can be set according to the actual application scenario, and does not limit the total loss threshold of the model, which can be set according to the actual application scenario.
- Step S102 obtaining a sampling sample image from the A sample image set, inputting the sampling sample image and the erroneous category label into the image recognition model to generate a first probability vector of the sampling sample image for the A category labels, wherein the erroneous category label is different from the category label corresponding to the sample image set to which the sampling sample image belongs;
- the first probability vector includes A probability elements, each of the A probability elements refers to one of the A category labels, and one of the probability elements is used to characterize the probability that the sampling sample image belongs to the referred category label.
- a sample image set service server Cb in A sample image sets is taken as an example for description, where b is a positive integer and b is less than or equal to A.
- the service server may determine a sampling ratio, and perform sampling processing on the sample images in the sample image set Cb according to the sampling ratio to obtain sampled sample images; among the A category labels, a category label different from the category label of the sample image set Cb is obtained (which may be obtained in a random manner), and the obtained category label is determined as an incorrect category label.
- the business server 100 obtains sample images from the above four sample image sets.
- a feasible sampling method is as follows: determine the sampling ratio, for example, 5%, and the business server 100 randomly samples images in the sample image set 201a to obtain 5% of the sample images in the sample image set 201a, and compose the sample images 201b with the sample images. For example, if there are 10,000 sample images in the sample image set 201a, the number of sample images 201b is 500.
- the generation process of the sample images corresponding to other sample image sets can refer to the generation process of the sample image 201b, which will not be described in detail here.
- the sampling sample image 201b in the example of FIG2a of the embodiment of the present application includes a first sampling sample image and a second sampling sample image. It can be understood that the first sampling sample image and the second sampling sample image are both sample images containing radish, so the correct category labels corresponding to the first sampling sample image and the second sampling sample image are both the category labels of radish, that is, the category label 0 illustrated in FIG2a.
- the sampling sample image 202b in the example of FIG2a of the embodiment of the present application includes a third sampling sample image and a fourth sampling sample image.
- the third sampling sample image and the fourth sampling sample image are both sample images containing peppers, so the correct category labels corresponding to the third sampling sample image and the fourth sampling sample image are both the category labels of peppers, that is, the category label 1 illustrated in FIG2a.
- the business server 100 randomly extracts a category label, such as category label 1, from category label 1, category label 2, and category label 3 illustrated in FIG2a, and sets it as the wrong category label for the first sampling sample image. Similarly, the business server 100 randomly extracts a category label, such as category label 2, from category label 1, category label 2, and category label 3 illustrated in FIG2a, and sets it as the wrong category label for the second sampling sample image. The business server 100 randomly extracts a category label, such as category label 2, from category label 0, category label 2, and category label 3 illustrated in FIG2a, and sets it as the wrong category label for the third sampling sample image.
- a category label such as category label 1, from category label 1, category label 2, and category label 3 illustrated in FIG2a
- the business server 100 randomly extracts a category label, such as category label 2, from category label 0, category label 2, and category label 3 illustrated in FIG2a, and sets it as the wrong category label for the third sampling sample image.
- the class label is 0, which is set as the wrong class label for the fourth sampled image.
- A is exemplified as 3, that is, 3 sample image sets, namely sample image set 3a, sample image set 3b and sample image set 3c.
- the category information of sample image set 3a is real objects, and its category label is 30a; the category information of sample image set 3b is paper objects, and its category label is 30b; the category information of sample image set 3c is mask objects, and its category label is 30c.
- sample image set 3b and sample image set 3c can be regarded as attack sample image sets of sample image set 3a.
- the process of the business server obtaining the sample sample images corresponding to the sample image set 3a, the sample image set 3b and the sample image set 3c is the same as the process described above, so it will not be repeated here.
- the sample sample images of the sample image set 3a include the fifth sample sample image and the sixth sample sample image
- the sample sample images of the sample image set 3b include the seventh sample sample image and the eighth sample sample image
- the sample sample images of the sample image set 3c include the ninth sample sample image and the tenth sample sample image.
- the business server 100 randomly extracts a category label, such as category label 30b, from the category label 30b and the category label 30c, and sets it as the wrong category label of the fifth sample sample image.
- the business server 100 randomly extracts a category label, such as category label 30c, from the category label 30b and the category label 30c, and sets it as the wrong category label of the sixth sample sample image.
- the business server sets the wrong category labels corresponding to the seventh sample sample image, the eighth sample sample image, the ninth sample sample image and the tenth sample sample image, respectively, to category label 30a. Therefore, the embodiment of the present application provides two methods for setting error category labels. In actual application, one of the two setting methods can be selected according to actual needs.
- the service server inputs the sample images corresponding to the A sample image sets into the image recognition model. It is understandable that the image recognition model processes each sample image in the same way, so please refer to the description of the sample image 201c in the embodiment corresponding to FIG. 2a above.
- the image recognition model can obtain the first probability vector corresponding to each sample image, and each first probability vector includes A probability values (also referred to as probability elements). Since the image recognition model is pre-trained by A sample image sets, the image recognition model can output the correct probability distribution for each sample image (the correct probability distribution indicates that the probability value of the correct category label in the output first probability vector is the largest). For ease of understanding, please refer to Figure 2a again.
- the first element (also referred to as a probability element, which is used to refer to the category label 0, and which is used to characterize the probability that the sample image belongs to the referred category label 0) of the first probability vector output by the image recognition model 20d can be infinitely close to 1, and the remaining three elements will be infinitely close to 0.
- the second element of the first probability vector output by the image recognition model 20d can be infinitely close to 1, and the remaining three elements (specifically, the first element, the third element, and the fourth element) are infinitely close to 0.
- the third element of the first probability vector output by the image recognition model 20d can be infinitely close to 1, and the remaining three elements (specifically, the first element, the second element, and the fourth element) are infinitely close to 0.
- the fourth element of the first probability vector output by the image recognition model 20d can be infinitely close to 1, and the remaining three elements (specifically, the first element, the second element, and the fourth element) are infinitely close to 0.
- Step S103 obtaining a first probability element indicating the wrong category label from the A probability elements in the first probability vector, adjusting the sample image according to the first probability element, and obtaining an adversarial sample image corresponding to the sample image, wherein the category prediction result obtained by performing category prediction on the adversarial sample image by the image recognition model is the wrong category label.
- the business server adjusts the sampled sample image according to the first probability element to obtain an initial adversarial sample image corresponding to the sampled sample image; inputs the error category label and the initial adversarial sample image into the image recognition model, wherein the error category label and the initial adversarial sample image can be respectively input into the image recognition model as two independent data, and there is a mapping relationship between the error category label and the initial adversarial sample image; or, the error category label can be added to the file header or file footer of the initial adversarial sample image, and then the initial adversarial sample image carrying the error category label can be input into the image recognition model.
- a second probability vector for the initial adversarial sample image for A category labels is generated, and the second probability vector also contains A probability elements, and the A probability elements respectively refer to one of the A category labels, and one of the probability elements is used to characterize the probability that the initial adversarial sample image belongs to the referred category label; in the second probability vector, obtain the second probability element referring to the error category label, and determine the adversarial sample image corresponding to the sampled sample image according to the second probability element.
- the business server can analyze whether the second probability element is a second probability vector.
- the initial adversarial sample image can be determined as the adversarial sample image corresponding to the sampling sample image (that is, at this time the image recognition model can mistakenly identify the category prediction result of the adversarial sample image as the wrong category label); if it is analyzed that the second probability element is not the probability element with the largest value in the second probability vector, the initial adversarial sample image will be further adjusted.
- the purpose of the adjustment is to hope that the image recognition model outputs a probability element referring to the wrong category label in the third probability vector for the adjusted initial adversarial sample image, which becomes larger, until the image recognition model can mistakenly identify the adversarial sample image after multiple rounds of adjustment as the wrong category label.
- the specific process of adjusting the sample image according to the first probability element to obtain the initial adversarial sample image corresponding to the sample image may include: generating a negative probability element corresponding to the first probability element, summing the negative probability element and the maximum probability element in the first probability vector to obtain the label loss value of the sample image for the wrong category label; adjusting the sample image according to the label loss value to obtain the initial adversarial sample image corresponding to the sample image. If the initial adversarial sample image can be directly used as the adversarial sample image corresponding to the sample image, it can also be understood that the sample image is adjusted according to the label loss value to directly obtain the adversarial sample image corresponding to the sample image.
- the sample image is adjusted to obtain the specific process of the initial adversarial sample image corresponding to the sample image, which may include: generating an initial gradient value for the sample image according to the label loss value, obtaining the numerical sign of the initial gradient value, and generating a symbol unit value corresponding to the initial gradient value according to the numerical sign; performing product processing on the initial learning rate and the symbol unit value to obtain a first value to be cropped, and cropping the first value to be cropped through a first cropping interval generated by the attack intensity to obtain a gradient value for adjusting the sample image; performing difference processing on the sample image and the gradient value (the gradient value may also be referred to as a perturbation value, and the difference processing may be understood as adding the perturbation value to the sample image to obtain an image containing noise) to obtain a second value to be cropped, and cropping the second value to be cropped through a second cropping interval generated by the pixel level to obtain the initial adversarial sample image corresponding to the sample image (if the initial adversarial
- the business server uses the maximum probability element in the first probability vector as the first loss value, uses the first probability element corresponding to the wrong category label in the first probability vector as the second loss value, and uses the difference between the first loss value and the second loss value as the label loss value of the sampled image for the wrong category label. Furthermore, the business server can adjust the sampled image according to the label loss value to obtain the initial adversarial sample image of the sampled image.
- the specific implementation method can adopt the Project Gradient Descent (an iterative attack algorithm, referred to as PGD) attack algorithm. Among them, the specific implementation process of the PGD attack algorithm can be described by the following formula (1).
- t in formula (1) represents the current iteration, and t is a positive integer, which means that the image recognition model will undergo at least one iteration, and that the sample image will undergo at least one adjustment.
- Represents the result data of the tth iteration such as the initial adversarial sample image generated after the first iteration, or the adversarial sample image generated after the last iteration.
- the input data of the first iteration is the sampled sample image.
- the input data of the second iteration is the initial adversarial sample image, that is, the result data of the first iteration; similarly, if there are more iterations, the input data of the next iteration is the result data of the previous iteration.
- ⁇ represents the learning rate.
- sign(.) represents the sign function, which can change each element of the initial gradient value (map) to ⁇ -1, 0, 1 ⁇ .
- eps represents the attack strength, that is, the absolute value of each element value of the gradient map used for adjustment each time cannot be greater than the attack strength.
- the attack strength is an adjustable parameter, generally set to 8/255, 16/255, 32/255, etc.
- the clip(.,.) function represents clipping, that is, the clip(.,.) function can control all values within the specified closed interval. [-eps,eps] represents the first clipping interval, and [0,255] represents the second clipping interval.
- the embodiments of the present application do not limit the method of adjusting the image.
- the above-mentioned PGD attack method may also be replaced by other attack methods, such as Fast Gradient Sign Method (an algorithm based on gradient generation adversarial samples, referred to as FGSM), Iterative Fast Gradient Sign Method (an improved version of FGSM, referred to as I-FGSM), Momentum Iterative Fast Gradient Sign Method (another improved version of FGSM, referred to as MI-FGSM), etc.
- FGSM Fast Gradient Sign Method
- I-FGSM Iterative Fast Gradient Sign Method
- MI-FGSM Momentum Iterative Fast Gradient Sign Method
- the business server inputs the error category label and the initial adversarial sample image into the image recognition model. Generate the second probability vector of the initial adversarial sample image for A category labels.
- ⁇ and the number of iterations t must meet a certain number, so as to ensure that the adjusted sampled sample image can cross the interface of the image recognition model and become a sample image that can be recognized by the image recognition model as a wrong category label (in order to distinguish it from the adversarial sample image, the sample image here is called a strong adversarial sample image, that is, the probability element of the image recognition model output for the strong adversarial sample image that refers to the wrong category label is close to 1).
- Figure 4 is a schematic diagram of a classification boundary surface trained by clean sample data provided in an embodiment of the present application
- Figure 5 is a schematic diagram of a classification boundary surface trained by first mixed sample data provided in an embodiment of the present application.
- the clean sample data refers to the above-mentioned A sample image sets
- the first mixed sample data includes the above-mentioned A sample image sets and the above-mentioned strong adversarial sample images.
- A can be equal to 2, that is, A category labels include the first category label and the second category label illustrated in Figures 4 and 5.
- A is greater than 2, but one category label is used as the true category label, and the remaining one or more category labels are used as attack category labels.
- the category feature distribution area corresponding to the first category label i.e., the area where the symbol “+” is located
- the category feature distribution area corresponding to the second category label i.e., the area where the symbol “-” is located
- the category feature distribution area corresponding to the first category label i.e., the area where the symbol “+” is located and the area where the symbol “ ⁇ ” is located
- the category feature distribution area corresponding to the second category label i.e., the area where the symbol “-” is located and the area where the symbol “ ⁇ ” is located
- the image recognition model will make the interface jagged. Although this improves the robustness of the optimized image recognition model to the strong adversarial sample data, some sample images with low confidence close to the interface are also misclassified.
- the purpose of the embodiment of the present application should be to improve the accuracy and generalization of the image recognition model while improving the robustness. Therefore, the embodiment of the present application adds boundary constraints to screen adversarial sample images. In other words, the adversarial sample images used for data enhancement should not cross the boundary too far.
- another way to determine the adversarial sample image may be: the business server adjusts the sampled sample image according to the first probability element to obtain the initial adversarial sample image corresponding to the sampled sample image; the wrong category label and the initial adversarial sample image are input into the image recognition model, and the image recognition model is used to generate a second probability vector for the A category labels of the initial adversarial sample image; among the A probability elements contained in the second probability vector, the second probability element for the wrong category label is obtained; if the second probability element does not meet the boundary constraint, or the category prediction result obtained by performing category prediction on the adversarial sample image is not the wrong category label, then the initial adversarial sample image is continued to be predicted according to the second probability element.
- the adversarial sample image is adjusted first (the process of continuing to adjust the initial adversarial sample image is the same as the specific implementation process of adjusting the sampled sample image according to the first probability element to obtain the initial adversarial sample image, so it will not be repeated here, please refer to the description above.
- the image adjustment method is the same as the adjustment method of the sampled sample image); if the second probability element satisfies the boundary constraint condition, and the category prediction result obtained by performing category prediction on the adversarial sample image is the wrong category label, then the initial adversarial sample image is determined as the adversarial sample image corresponding to the sampled sample image; when the category prediction result is the wrong category label, the second probability element is the maximum value of the A probability elements contained in the second probability vector.
- the adversarial sample image desired by the present application is one that can both satisfy the boundary constraint condition and ensure that the image recognition model can incorrectly identify the adversarial sample image.
- the business server can also compare the second probability element and the probability constraint value; the probability constraint value is the inverse of A; if the second probability element is less than the probability constraint value, it is determined that the second probability element does not meet the boundary constraint condition; if the second probability element is equal to or greater than the probability constraint value, it is determined that the second probability element meets the boundary constraint condition. That is, the present application controls the learning rate and the number of iterations.
- the adversarial sample image obtained in the embodiment of the present application can be located near the classification boundary surface, rather than in the distribution of the wrong category label, and the classification boundary surface of the image recognition model optimized and trained by the adversarial sample image will be softer.
- Figure 6 is a schematic diagram of a classification boundary surface trained by the second mixed sample data provided by the embodiment of the present application.
- the data includes the above-mentioned A sample image sets and the above-mentioned adversarial sample images.
- the category feature distribution area corresponding to the first category label i.e., the area where the symbol "+” is located and the area where the symbol " ⁇ ” is located
- the category feature distribution area corresponding to the second category label i.e., the area where the symbol "-” is located and the area where the symbol " ⁇ ” is located
- the optimized image recognition model can accurately identify the first category label and the second category label, that is, it can ensure that the optimized image recognition model (referred to as the image recognition optimization model) maintains or improves the original recognition accuracy, and can also supplement the deficiencies of the training data, so the defense capability and security of the image recognition optimization model can be improved.
- the embodiment of the present application proposes a new adversarial data enhancement based on boundary constraints, which utilizes the particularity of the distribution of adversarial data at the boundary to amplify and enhance the training data, thereby improving the defense capability and generalization of the optimized image recognition model.
- a computer device can obtain A sample image sets and an image recognition model; wherein the A sample image sets respectively correspond to one of A different category labels, and the image recognition model is pre-trained based on the A sample image sets, so the image recognition model can accurately determine the category labels of the sample images in the A sample image sets; further, the computer device obtains a sampled sample image from the A sample image sets, and inputs an erroneous category label and the sampled sample image into the image recognition model, wherein the erroneous category label belongs to the A category labels, and the erroneous category label is different from the category label corresponding to the sample image set to which the sampled sample image belongs; a first probability vector of the sampled sample image for the A category labels can be generated through the image recognition model, and the first probability vector includes A probability elements, and the A probability elements respectively refer to one of the A category labels.
- a probability element is used to characterize the probability that the sampled sample image belongs to the referred category label; since the image recognition model can accurately determine the category labels of the sample images in the A sample image sets, in the first probability vector, the probability element referring to the correct category label is the largest, or in other words, the probability element referring to the correct category label is much larger than the first probability element referring to the wrong category label, and the correct category label refers to the category label corresponding to the sample image set to which the sampled sample image belongs; further, according to the first probability element, the computer device can adjust the sampled sample image to obtain an adversarial sample image corresponding to the sampled sample image, wherein the category prediction result obtained by the image recognition model for category prediction of the adversarial sample image is the wrong category label, that is, by adjusting the sampled sample image, the image recognition model can incorrectly determine the category label of the adversarial sample image.
- the embodiment of the present application proposes a method for generating adversarial sample images.
- adversarial sample images that are incorrectly recognized by an image recognition model can be generated.
- the adversarial sample images can improve the sample diversity of the training set used to optimize the training of the image recognition model.
- the training set with sample diversity can improve the optimization training accuracy of the image recognition model, and then improve the category label recognition accuracy and recognition generalization of the optimized trained image recognition model.
- the classification boundary of the image recognition model after optimization training can be made softer, further ensuring the model accuracy.
- Figure 7 is a flow chart of a data processing method provided in an embodiment of the present application.
- the method can be executed by a business server (for example, the business server 100 shown in Figure 1 above), or by a terminal device (for example, the terminal device 200a shown in Figure 1 above), or by a business server and a terminal device interacting with each other.
- a business server for example, the business server 100 shown in Figure 1 above
- a terminal device for example, the terminal device 200a shown in Figure 1 above
- the embodiment of the present application is described by taking the method being executed by a business server as an example.
- the method can at least include the following steps.
- Step S201 obtain A sample image sets and an image recognition model, where A is a positive integer greater than 1, and the A sample image sets respectively correspond to one of A different category labels.
- the image recognition model is pre-trained based on the A sample image sets, and the A category labels include error category labels.
- the business service obtains A original image sets, and inputs the A original image sets into the object detection model; the A original image sets correspond to one of A different category information respectively; the A original image sets include the original image set D b ; the original image set D b includes the original image E f ; f is a positive integer, and f is less than or equal to the total number of original images in the original image set D b ; through the object detection model, the region coordinates of the key region in the original image E f are determined, and according to the region coordinates, the sample image to be labeled corresponding to the original image E f is generated; the sample image to be labeled corresponding to each original image in the original image set D b is determined as the sample image set to be labeled; according to the category information corresponding to the original image set D b, a category label H b is generated, and the sample image set to be labeled with the category label H b is determined as the sample image set C b.
- the specific process of generating a sample image to be labeled corresponding to the original image Ef according to the region coordinates may include: generating an initial detection frame including a key area according to the region coordinates, expanding the initial detection frame to obtain a detection frame including the target object; the key area belongs to the target object; in the detection frame, obtaining an image to be scaled including the target object, obtaining a first image size, scaling the image to be scaled according to the first image size, to obtain an image to be cropped; obtaining a second image size, cropping the image to be cropped according to the second image size, to obtain a sample image to be labeled corresponding to the original image Ef; the second image size is smaller than the first image size.
- the business server does not limit the model type of the object detection model and can set it according to the actual application scenario, such as Region with CNN feature (a target detection model, referred to as Faster R-CNN), Single Shot MultiBox Detector (another target detection model, referred to as SSD), You Only Look Once (a simple and fast target detection algorithm, referred to as YOLO), etc.
- Region with CNN feature a target detection model, referred to as Faster R-CNN
- SSD Single Shot MultiBox Detector
- YOLO a simple and fast target detection algorithm
- the key area belongs to the target object.
- the embodiment of the present application does not limit the target object.
- it can be a cat, a dog, a person, a table, Pepper, etc.
- the key area is not limited, and it should be set according to the target object.
- the expansion multiple of the initial detection frame should be set according to the key area, so the embodiment of the application does not limit the expansion multiple.
- the second image size refers to the input size of the image recognition model.
- Step S202 obtaining a sampling sample image from the A sample image set, inputting the sampling sample image and the erroneous category label into the image recognition model to generate a first probability vector of the sampling sample image for the A category labels, wherein the erroneous category label is different from the category label corresponding to the sample image set to which the sampling sample image belongs;
- the first probability vector includes A probability elements, each of the A probability elements refers to one of the A category labels, and one of the probability elements is used to characterize the probability that the sampling sample image belongs to the referred category label.
- Step S203 obtaining a first probability element indicating the wrong category label from among the A probability elements in the first probability vector, adjusting the sample image according to the first probability element, and obtaining an adversarial sample image corresponding to the sample image, wherein the category prediction result obtained by performing category prediction on the adversarial sample image by the image recognition model is the wrong category label.
- step S202 to step S203 please refer to step S102 and step S103 in the embodiment corresponding to FIG. 3 above, which will not be described in detail here.
- Step S204 inputting the A sample image sets, the A category labels, the correct category label and the adversarial sample image into the image recognition model;
- the correct category label refers to the category label corresponding to the sample image set to which the sampled sample image belongs.
- the A sample image sets, the A category labels, the correct category labels and the adversarial sample images can be respectively input into the image recognition model as four independent data, and there is a mapping relationship between the A sample image sets and the A category labels, and there is a mapping relationship between the correct category labels and the adversarial sample images; or, the corresponding category labels can be added to the file header or file footer of the sample images in the sample image set, and the correct category labels can be added to the file header or file footer of the adversarial sample images, and then the sample images carrying category labels and the adversarial sample images carrying correct category labels can be input into the image recognition model.
- Step S205 Generate, by using the image recognition model, first prediction categories corresponding to the A sample image sets and second prediction categories corresponding to the adversarial sample images.
- Step S206 generating a first loss value according to the first predicted category and the A category labels, generating a second loss value according to the second predicted category and the correct category label, adjusting the parameters in the image recognition model according to the first loss value and the second loss value (specifically, the first loss value and the second loss value can be weighted and summed to obtain a total loss value, and then the parameters in the image recognition model are adjusted based on the total loss value) to obtain the image recognition optimization model; the category prediction result obtained by the image recognition optimization model for category prediction of the adversarial sample image is the correct category label, and the category prediction result obtained by the image recognition optimization model for category prediction of the sample image in the sample image set is still the category label corresponding to the sample image set.
- the optimization training process of the image recognition model by the business server is the same as the training process of the initial image recognition model by the business server. Therefore, for the specific implementation process of steps S204 to S206, please refer to step S101 in the embodiment corresponding to Figure 3 above, which will not be repeated here.
- the image recognition optimization model proposed in the embodiment of the present application can be deployed on a terminal device, for example, to detect an image input for real object recognition (such as real person recognition). If it is a real object, it passes and enters the subsequent recognition process. If it is not a real object, such as a photo object, an error is reported and a retry prompt is given.
- the embodiment of the present application can be applied to all object recognition related applications, including but not limited to: online payment, offline payment, access control unlocking system, mobile phone unlocking recognition, automatic object recognition clearance, etc.
- an embodiment of the present application proposes a method for generating adversarial sample images.
- adversarial sample images that are incorrectly recognized by an image recognition model can be generated.
- the sample diversity of the training set used to optimize the training of the image recognition model can be improved.
- the optimization training accuracy of the image recognition model can be improved, and then the category label recognition accuracy and recognition generalization of the optimized trained image recognition model can be improved.
- Figure 8 is a flow chart of a data processing method provided by an embodiment of the present application.
- the method can be executed by a business server (for example, the business server 100 shown in Figure 1 above), or by a terminal device (for example, the terminal device 200a shown in Figure 1 above), or by a business server and a terminal device interacting with each other.
- a business server for example, the business server 100 shown in Figure 1 above
- a terminal device for example, the terminal device 200a shown in Figure 1 above
- the embodiment of the present application is described by taking the method being executed by a business server as an example.
- the method may at least include the following steps.
- Step S301 obtain a target image, and input the target image into an image recognition optimization model;
- the image recognition optimization model is obtained by continuing to optimize the image recognition model according to the mixed sample training set;
- the mixed sample training set includes A sample image sets, A category labels, correct category labels and adversarial sample images;
- the A sample image sets correspond to one of A different category labels respectively, and the A sample image sets include collected sample images;
- the correct category label belongs to the A category labels;
- A is a positive integer greater than 1;
- the image recognition model is obtained by pre-training the A sample image sets;
- the adversarial sample image is obtained by adjusting the sampled sample image by the first probability element of the A probability elements in the first probability vector that refers to the wrong category label, and the first probability
- the vector is generated by the image recognition model for the sampled sample image, the erroneous category label is different from the category label corresponding to the sample image set to which the sampled sample image belongs, and the correct category label is the category label corresponding to the sample image set
- Step S302 Generate a target probability vector of the target image for A category labels through an image recognition optimization model.
- Step S303 determine the category label corresponding to the maximum probability element in the target probability vector as the category prediction result of the target image.
- the image recognition model of the embodiment of the present application can be used to identify real objects, such as real people.
- the image recognition model can be called a real object detection model.
- the image recognition model can be called a real object detection model.
- the image recognition model can be used to identify real objects, such as real people.
- the image recognition model can be called a real object detection model.
- the image recognition model can be called a real object detection model.
- the real object detection model there are many types of attacks brought by illegal elements, and the attack forms are novel and difficult to defend.
- the real object detection model can be bypassed to achieve the attack. This problem is also faced in high-precision 3D attacks. Once a new material appears, the real object detection model will face a situation where it cannot be defended.
- the embodiment of the present application proposes a new real object detection technology based on boundary constraint adversarial data enhancement, which enhances the real object training data by designing a special adversarial sample image, thereby improving the overall performance of real object detection.
- boundary constraint adversarial data enhancement which enhances the real object training data by designing a special adversarial sample image, thereby improving the overall performance of real object detection.
- an embodiment of the present application proposes a method for generating adversarial sample images.
- adversarial sample images that are incorrectly recognized by an image recognition model can be generated.
- the sample diversity of the training set used to optimize the training of the image recognition model can be improved.
- the optimization training accuracy of the image recognition model can be improved, and then the category label recognition accuracy and recognition generalization of the optimized trained image recognition model can be improved.
- FIG. 9 is a structural schematic diagram of a data processing device provided in an embodiment of the present application.
- the above-mentioned data processing device 1 can be used to execute the corresponding steps in the method provided in an embodiment of the present application.
- the data processing device 1 may include: a data acquisition module 11, a first input module 12, and a first adjustment module 13.
- a data acquisition module 11 is used to acquire A sample image sets and an image recognition model, where A is a positive integer greater than 1, the A sample image sets correspond to one of A different category labels respectively, the image recognition model is pre-trained based on the A sample image sets, and the A category labels include error category labels;
- a first input module 12 is used to obtain a sampled sample image from the A sample image sets, and input the sampled sample image and the erroneous category label into the image recognition model to generate a first probability vector of the sampled sample image for the A category labels, wherein the erroneous category label is different from the category label corresponding to the sample image set to which the sampled sample image belongs;
- the first probability vector includes A probability elements, each of the A probability elements refers to one of the A category labels, and one of the probability elements is used to represent the probability that the sampled sample image belongs to the referred category label;
- the first adjustment module 13 is used to obtain a first probability element referring to the wrong category label from the A probability elements in the first probability vector, and adjust the sampled sample image according to the first probability element to obtain an adversarial sample image corresponding to the sampled sample image, wherein the category prediction result obtained by the image recognition model performing category prediction on the adversarial sample image is the wrong category label.
- the first adjustment module 13 can be specifically used to generate a negative probability element corresponding to the first probability element, sum the negative probability element and the maximum probability element in the first probability vector, obtain the label loss value of the sampling sample image for the wrong category label, and adjust the sampling sample image according to the label loss value to obtain the adversarial sample image corresponding to the sampling sample image.
- the first adjustment module 13 is used to adjust the sample image according to the label loss value to obtain the adversarial sample image corresponding to the sample image. Specifically, it is used to generate an initial gradient value for the sample image according to the label loss value, obtain the numerical sign of the initial gradient value, generate a symbol unit value corresponding to the initial gradient value according to the numerical sign, multiply the initial learning rate and the symbol unit value to obtain a first value to be cropped, crop the first value to be cropped through a first cropping interval generated by the attack intensity to obtain a gradient value for adjusting the sample image, perform difference processing on the sample image and the gradient value to obtain a second value to be cropped, crop the second value to be cropped through a second cropping interval generated by the pixel level, and obtain the adversarial sample image corresponding to the sample image.
- the first adjustment module 13 may include: a first adjustment unit 131 , a first generation unit 132 , and a first acquisition unit 133 .
- a first adjustment unit 131 is used to adjust the sample image according to the first probability element to obtain an initial adversarial sample image corresponding to the sample image;
- a first generating unit 132 is used to input the wrong category label and the initial adversarial sample image into the image recognition model, and generate a second probability vector of the initial adversarial sample image for A category labels through the image recognition model;
- a first acquiring unit 133 is configured to acquire a second probability element for the error category label from the A probability elements included in the second probability vector;
- the first adjustment unit 131 is further configured to continue adjusting the initial adversarial sample image according to the second probability element if the second probability element does not satisfy the boundary constraint condition, or the category prediction result obtained by performing category prediction on the adversarial sample image is not the wrong category label;
- the first acquisition unit 133 is further used to determine the initial adversarial sample image as the adversarial sample image corresponding to the sampling sample image if the second probability element satisfies the boundary constraint condition and the category prediction result obtained by performing category prediction on the adversarial sample image is the wrong category label; when the category prediction result is the wrong category label, the second probability element is the maximum value of the A probability elements contained in the second probability vector.
- the specific functional implementation of the first adjustment unit 131 , the first generation unit 132 and the first acquisition unit 133 may refer to step S103 in the embodiment corresponding to FIG. 3 , which will not be described in detail here.
- the first acquisition unit 133 is further used to compare the second probability element with a probability constraint value; the probability constraint value is the reciprocal of A;
- the first acquisition unit 133 is further configured to determine that the second probability element does not satisfy a boundary constraint condition if the second probability element is less than the probability constraint value;
- the first acquiring unit 133 is further configured to determine that the second probability element satisfies a boundary constraint condition if the second probability element is equal to or greater than the probability constraint value.
- the data processing device 1 may further include: a second input module 14 and a second adjustment module 15 .
- a second input module 14 is used to input the A sample image sets, the A category labels, the correct category label and the adversarial sample image into the image recognition model;
- the correct category label refers to the category label corresponding to the sample image set to which the sample image belongs;
- the second input module 14 is further used to generate the first prediction categories corresponding to the A sample image sets and the second prediction category corresponding to the adversarial sample image through the image recognition model;
- the second adjustment module 15 is used to generate a first loss value according to the first predicted category and the A category labels, generate a second loss value according to the second predicted category and the correct category label, and adjust the parameters in the image recognition model according to the first loss value and the second loss value to obtain the image recognition optimization model; the category prediction result obtained by the image recognition optimization model for category prediction of the adversarial sample image is the correct category label.
- the specific functional implementation of the second input module 14 and the second adjustment module 15 can refer to steps S204 to S206 in the embodiment corresponding to FIG. 7 above, which will not be described in detail here.
- the first input module 12 may include: a first determining unit 121 and a second obtaining unit 122 .
- the A sample image sets include a sample image set Cb; b is a positive integer and b is less than or equal to A;
- the first determining unit 121 is used to determine a sampling ratio, and perform sampling processing on the sample images in the sample image set Cb according to the sampling ratio to obtain a sampled sample image;
- the second acquisition unit 122 is used to acquire a category label different from the category label of the sample image set Cb from among the A category labels, and determine the acquired category label as an erroneous category label.
- the specific functional implementation of the first determining unit 121 and the second acquiring unit 122 can refer to step S102 in the embodiment corresponding to FIG. 3 above, which will not be described in detail here.
- the data acquisition module 11 may include: a third acquisition unit 111 , a second generation unit 112 , a second determination unit 113 , a third determination unit 114 and a second adjustment unit 115 .
- a third acquisition unit 111 is used to acquire A sample image sets, and input A category labels and the A sample image sets into the image recognition initial model respectively;
- a second generating unit 112 is used to generate predicted initial categories corresponding to the A sample image sets respectively through an image recognition initial model
- a second determining unit 113 is used to determine the category loss values corresponding to the A sample image sets respectively according to the predicted initial category and the A category labels;
- the third determining unit 114 is used to determine the image recognition initial model corresponding to the category loss values corresponding to the A sample image sets. Total loss value;
- the second adjustment unit 115 is used to adjust the parameters in the initial image recognition model according to the total loss value to obtain an image recognition model; the category prediction result obtained by the image recognition model for the sample image set Cb is the category label of the sample image set Cb; the A sample image sets include the sample image set Cb; b is a positive integer and b is less than or equal to A.
- step S101 in the embodiment corresponding to Figure 3 above, and will not be repeated here.
- the third acquisition unit 111 may include: an image acquisition subunit 1111 , a second determination subunit 1112 , a third determination subunit 1113 , and a fourth determination subunit 1114 .
- the image acquisition subunit 1111 is used to acquire A original image sets, and input the A original image sets into the object detection model; the category information corresponding to the A original image sets are different from each other; the A original image sets include an original image set Db; the original image set Db includes an original image Ef; f is a positive integer, and f is less than or equal to the total number of original images in the original image set Db;
- the second determining subunit 1112 is used to determine the region coordinates of the key region in the original image Ef through the object detection model, and generate a sample image to be labeled corresponding to the original image Ef according to the region coordinates;
- the third determining subunit 1113 is used to determine the sample image to be labeled corresponding to each original image in the original image set Db as the sample image set to be labeled;
- the fourth determining subunit 1114 is used to generate a category label Hb according to the category information corresponding to the original image set Db, and determine the set of sample images to be labeled with the category label Hb as the sample image set Cb.
- the specific functional implementation of the image acquisition subunit 1111, the second determination subunit 1112, the third determination subunit 1113 and the fourth determination subunit 1114 can refer to step S201 in the embodiment corresponding to Figure 7 above, and will not be repeated here.
- the second determining subunit 1112 may include: a third processing subunit 11121 , a fourth processing subunit 11122 , and a fifth processing subunit 11123 .
- the third processing subunit 11121 is used to generate an initial detection frame including the key region according to the region coordinates, and perform expansion processing on the initial detection frame to obtain a detection frame including the target object; the key region belongs to the target object;
- the fourth processing subunit 11122 is used to obtain the image to be scaled including the target object in the detection frame, obtain the first image size, and scale the image to be scaled according to the first image size to obtain the image to be cropped;
- the fifth processing subunit 11123 is used to obtain a second image size, and perform cropping processing on the image to be cropped according to the second image size to obtain a sample image to be labeled corresponding to the original image Ef; the second image size is smaller than the first image size.
- the specific functional implementation of the third processing sub-unit 11121, the fourth processing sub-unit 11122 and the fifth processing sub-unit 11123 can refer to step S201 in the embodiment corresponding to Figure 7 above, and will not be repeated here.
- the embodiment of the present application proposes a method for generating adversarial sample images, through which adversarial sample images that are incorrectly recognized by an image recognition model can be generated.
- adversarial sample images the sample diversity of the training set used to optimize the training of the image recognition model can be improved.
- the optimization training accuracy of the image recognition model can be improved, and then the category label recognition accuracy and recognition generalization of the optimized trained image recognition model can be improved.
- the second probability element for the wrong category label predicted by the image recognition model for the adversarial sample image satisfies the boundary constraint condition, the classification boundary of the image recognition model after optimization training can be made softer, further ensuring the model accuracy.
- FIG. 10 is a second structural diagram of a data processing device provided in an embodiment of the present application.
- the above-mentioned data processing device 2 can be used to execute the corresponding steps in the method provided in an embodiment of the present application.
- the data processing device 2 may include: an image acquisition module 21, a vector generation module 22, and a category determination module 23.
- the image acquisition module 21 is used to acquire a target image and input the target image into an image recognition optimization model; the image recognition optimization model is obtained by continuing to optimize the image recognition model according to the mixed sample training set; the mixed sample training set includes A sample image sets, A category labels, correct category labels and adversarial sample images; the A sample image sets correspond to one of A different category labels respectively, and the A sample image sets include collected sample images; the correct category label belongs to the A category labels; A is a positive integer greater than 1; the image recognition model is obtained by pre-training the A sample image sets; the adversarial sample image is represented by the A probability elements in the first probability vector The first probability element of the wrong category label is obtained by adjusting the sampled sample image, the first probability vector is generated by the image recognition model for the sampled sample image, the wrong category label is different from the category label corresponding to the sample image set to which the sampled sample image belongs, and the correct category label is the category label corresponding to the sample image set to which the sampled sample image belongs; the first probability vector includes A probability elements, the A
- a vector generation module 22 used to generate a target probability vector of a target image for A category labels through an image recognition optimization model
- the category determination module 23 is used to determine the category label corresponding to the maximum probability element in the target probability vector as the category prediction result of the target image.
- an embodiment of the present application proposes a method for generating adversarial sample images.
- adversarial sample images that are incorrectly recognized by an image recognition model can be generated.
- the sample diversity of the training set used to optimize the training of the image recognition model can be improved.
- the optimization training accuracy of the image recognition model can be improved, and then the category label recognition accuracy and recognition generalization of the optimized trained image recognition model can be improved.
- the computer device 1000 may include: at least one processor 1001, such as a CPU, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002.
- the communication bus 1002 is used to realize the connection and communication between these components.
- the user interface 1003 may include a display screen (Display), a keyboard (Keyboard), and the network interface 1004 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface).
- the memory 1005 may be a high-speed RAM memory, or it may be a non-volatile memory (non-volatile memory), such as at least one disk memory.
- the memory 1005 may also be optionally at least one storage device located away from the aforementioned processor 1001.
- the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a device control application.
- the network interface 1004 can provide a network communication function;
- the user interface 1003 is mainly used to provide an input interface for the user; and
- the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:
- a sample image sets and an image recognition model where A is a positive integer greater than 1, the A sample image sets respectively correspond to one of A different category labels, the image recognition model is pre-trained based on the A sample image sets, and the A category labels include error category labels;
- a sampling sample image is obtained from the A sample image sets, and the sampling sample image and the erroneous category label are input into the image recognition model to generate a first probability vector of the sampling sample image for the A category labels, wherein the erroneous category label is different from the category label corresponding to the sample image set to which the sampling sample image belongs;
- the first probability vector includes A probability elements, each of the A probability elements refers to one of the A category labels, and one of the probability elements is used to represent the probability that the sampling sample image belongs to the referred category label;
- the A probability elements in the first probability vector obtain a first probability element referring to the wrong category label, and adjust the sampled sample image according to the first probability element to obtain an adversarial sample image corresponding to the sampled sample image, and the category prediction result obtained by performing category prediction on the adversarial sample image by the image recognition model is the wrong category label.
- the processor 1001 may be configured to call a device control application stored in the memory 1005 to implement:
- the image recognition optimization model is obtained by continuing to optimize the image recognition model according to the mixed sample training set;
- the mixed sample training set includes A sample image sets, A category labels, correct category labels and adversarial sample images;
- the A sample image sets correspond to one of A different category labels respectively, and the A sample image sets include collected sample images;
- the correct category label belongs to the A category labels;
- A is a positive integer greater than 1;
- the image recognition model is obtained by pre-training the A sample image sets;
- the adversarial sample image is obtained by referring to the wrong category label in the A probability elements in the first probability vector
- the first probability vector is obtained by adjusting the sampled sample image by the first probability element of the signature, the first probability vector is generated by the image recognition model for the sampled sample image, the wrong category label is different from the category label corresponding to the sample image set to which the sampled sample image belongs, and the correct category label is the category label corresponding to the sample image set to which the sampled sample image belongs;
- the first probability vector includes A probability
- the category label corresponding to the maximum probability element in the target probability vector is determined as the category prediction result of the target image.
- the present application also provides a computer-readable storage medium, which stores a computer program.
- a computer program When the computer program is executed by a processor, the description of the data processing method or device in the above embodiments is implemented, which will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated.
- the computer-readable storage medium may be the data processing device provided in any of the aforementioned embodiments or the internal storage unit of the computer device, such as a hard disk or memory of the computer device.
- the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (smart media card, SMC), a secure digital (secure digital, SD) card, a flash card (flash card), etc. equipped on the computer device.
- the computer-readable storage medium may also include both the internal storage unit of the computer device and an external storage device.
- the computer-readable storage medium is used to store the computer program and other programs and data required by the computer device.
- the computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
- the present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium.
- the processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device can execute the description of the data processing method or device in the above embodiments, which will not be repeated here.
- the description of the beneficial effects of the same method will not be repeated.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Evolutionary Computation (AREA)
- Artificial Intelligence (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Software Systems (AREA)
- Computing Systems (AREA)
- Multimedia (AREA)
- Databases & Information Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Medical Informatics (AREA)
- Data Mining & Analysis (AREA)
- Computational Linguistics (AREA)
- General Engineering & Computer Science (AREA)
- Life Sciences & Earth Sciences (AREA)
- Molecular Biology (AREA)
- Biophysics (AREA)
- Biomedical Technology (AREA)
- Mathematical Physics (AREA)
- Probability & Statistics with Applications (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Evolutionary Biology (AREA)
- Image Analysis (AREA)
Abstract
Description
Claims (16)
- 一种图像处理方法,其特征在于,包括:获取A个样本图像集以及图像识别模型,A为大于1的正整数,所述A个样本图像集分别对应A个不同的类别标签中的一个,所述图像识别模型是根据所述A个样本图像集预训练得到的,所述A个类别标签包括错误类别标签;在所述A个样本图像集中获取采样样本图像,将所述采样样本图像与所述错误类别标签输入至所述图像识别模型,以生成所述采样样本图像针对所述A个类别标签的第一概率向量,所述错误类别标签与所述采样样本图像所属的样本图像集对应的类别标签不同;所述第一概率向量包括A个概率元素,所述A个概率元素分别指代所述A个类别标签中的一个,一个所述概率元素用于表征所述采样样本图像属于所指代的类别标签的概率;在所述第一概率向量中的所述A个概率元素中,获取指代所述错误类别标签的第一概率元素,根据所述第一概率元素,对所述采样样本图像进行调整,得到所述采样样本图像对应的对抗样本图像,所述对抗样本图像被所述图像识别模型进行类别预测所得到的类别预测结果为所述错误类别标签。
- 根据权利要求1所述的方法,其特征在于,所述根据所述第一概率元素,对所述采样样本图像进行调整,得到所述采样样本图像对应的对抗样本图像,包括:生成所述第一概率元素对应的负概率元素,对所述负概率元素以及所述第一概率向量中的最大概率元素进行求和处理,得到所述采样样本图像针对所述错误类别标签的标签损失值;根据所述标签损失值,对所述采样样本图像进行调整,得到所述采样样本图像对应的对抗样本图像。
- 根据权利要求1或2所述的方法,其特征在于,所述根据所述标签损失值,对所述采样样本图像进行调整,得到所述采样样本图像对应的对抗样本图像,包括:根据所述标签损失值,生成针对所述采样样本图像的初始梯度值,获取所述初始梯度值的数值符号,根据所述数值符号生成所述初始梯度值对应的符号单位数值;对初始学习率以及符号单位数值进行乘积处理,得到第一待裁剪数值,通过由攻击强度所生成的第一裁剪区间,对所述第一待裁剪数值进行裁剪处理,得到用于调整所述采样样本图像的梯度值;对所述采样样本图像以及所述梯度值进行求差处理,得到第二待裁剪数值,通过由像素等级所生成的第二裁剪区间,对所述第二待裁剪数值进行裁剪处理,得到所述采样样本图像对应的对抗样本图像。
- 根据权利要求1至3任一项所述的方法,其特征在于,所述根据所述第一概率元素,对所述采样样本图像进行调整,得到所述采样样本图像对应的对抗样本图像,包括:根据所述第一概率元素,对所述采样样本图像进行调整,得到所述采样样本图像对应的初始对抗样本图像;将所述错误类别标签和所述初始对抗样本图像输入至所述图像识别模型,通过所述图像识别模型,生成所述初始对抗样本图像针对所述A个类别标签的第二概率向量;在所述第二概率向量所包含的A个概率元素中,获取针对所述错误类别标签的第二概率元素;若所述第二概率元素不满足边界约束条件,或针对所述对抗样本图像进行类别预测所得到的类别预测结果不为所述错误类别标签,则根据所述第二概率元素继续对所述初始对抗样本图像进行调整;若所述第二概率元素满足边界约束条件,且针对所述对抗样本图像进行类别预测所得到的类别预测结果为所述错误类别标签,则将所述初始对抗样本图像确定为所述采样样本图像对应的对抗样本图像;在所述类别预测结果为所述错误类别标签的情况下,所述第二概率元素为所述第二概率向量所包含的A个概率元素中的最大数值。
- 根据权利要求1至4任一项所述的方法,其特征在于,还包括:将所述第二概率元素以及概率约束值进行对比;所述概率约束值为A的倒数;若所述第二概率元素小于所述概率约束值,则确定所述第二概率元素不满足边界约束条件;若所述第二概率元素等于或大于所述概率约束值,则确定所述第二概率元素满足边界约束条件。
- 根据权利要求1至5任一项所述的方法,其特征在于,所述方法还包括:将所述A个样本图像集、所述A个类别标签、正确类别标签以及所述对抗样本图像均输入至所述图像识别模型;所述正确类别标签是指所述采样样本图像所属的样本图像集对应的类别标签;通过所述图像识别模型,生成所述A个样本图像集分别对应的第一预测类别,以及所述对抗样本图像对应的第二预测类别;根据所述第一预测类别和所述A个类别标签生成第一损失值,根据所述第二预测类别以及所述正确类 别标签生成第二损失值,根据所述第一损失值和所述第二损失值对所述图像识别模型中的参数进行调整,得到所述图像识别优化模型;所述图像识别优化模型针对所述对抗样本图像进行类别预测所得到的类别预测结果为所述正确类别标签。
- 根据权利要求1至6任一项所述的方法,其特征在于,所述A个样本图像集包括样本图像集Cb;b为正整数且b小于或等于A;所述在所述A个样本图像集中获取采样样本图像,包括:确定采样比例,根据所述采样比例,对所述样本图像集Cb中的样本图像进行采样处理,得到采样样本图像;在所述A个类别标签中,获取与所述样本图像集Cb的类别标签不同的类别标签,将获取到的类别标签确定为所述错误类别标签。
- 根据权利要求1至7任一项所述的方法,其特征在于,所述获取A个样本图像集以及图像识别模型,包括:获取A个样本图像集,将A个类别标签和所述A个样本图像集分别输入至所述图像识别初始模型;通过所述图像识别初始模型,生成所述A个样本图像集分别对应的预测初始类别;根据所述预测初始类别以及所述A个类别标签,确定所述A个样本图像集分别对应的类别损失值;根据所述A个样本图像集分别对应的类别损失值,确定所述图像识别初始模型对应的总损失值;根据所述总损失值,对所述图像识别初始模型中的参数进行调整,得到图像识别模型;所述图像识别模型针对样本图像集Cb进行类别预测所得到的类别预测结果为所述样本图像集Cb的类别标签;所述A个样本图像集包括样本图像集Cb;b为正整数且b小于或等于A。
- 根据权利要求8所述的方法,其特征在于,所述获取A个样本图像集,包括:获取A个原始图像集,将所述A个原始图像集均输入至对象检测模型;所述A个原始图像集分别对应A个不同的类别信息中的一个;所述A个原始图像集包括原始图像集Db;所述原始图像集Db包括原始图像Ef;f为正整数,且f小于或等于所述原始图像集Db中原始图像的总数量;通过所述对象检测模型,确定所述原始图像Ef中关键区域的区域坐标,根据所述区域坐标,生成所述原始图像Ef对应的待标注样本图像;将所述原始图像集Db中每个原始图像对应的待标注样本图像,确定为待标注样本图像集;根据所述原始图像集Db对应的类别信息,生成类别标签Hb,将标注有所述类别标签Hb的待标注样本图像集,确定为所述样本图像集Cb。
- 根据权利要求9所述的方法,其特征在于,所述根据所述区域坐标,生成所述原始图像Ef对应的待标注样本图像,包括:根据所述区域坐标,生成包括所述关键区域的初始检测框,对所述初始检测框进行扩张处理,得到包括目标对象的检测框;所述关键区域属于所述目标对象;在所述检测框中,获取包括所述目标对象的待缩放图像,获取第一图像尺寸,根据所述第一图像尺寸,对所述待缩放图像进行缩放处理,得到待裁剪图像;获取第二图像尺寸,根据所述第二图像尺寸,对所述待裁剪图像进行裁剪处理,得到所述原始图像Ef对应的待标注样本图像;所述第二图像尺寸小于所述第一图像尺寸。
- 一种数据处理方法,其特征在于,包括:获取目标图像,将所述目标图像输入至图像识别优化模型;所述图像识别优化模型是根据混合样本训练集继续对图像识别模型进行优化训练所得到的;所述混合样本训练集包括A个样本图像集、A个类别标签、正确类别标签以及对抗样本图像;所述A个样本图像集分别对应A个不同的类别标签中的一个,所述A个样本图像集包括采集样本图像;所述正确类别标签属于所述A个类别标签;A为大于1的正整数;所述图像识别模型是根据所述A个样本图像集预训练得到的;所述对抗样本图像是通过第一概率向量中的A个概率元素中指代错误类别标签的第一概率元素对采样样本图像进行调整得到,所述第一概率向量是由所述图像识别模型针对所述采样样本图像所生成的,所述错误类别标签与所述采样样本图像所属的样本图像集对应的类别标签不同,所述正确类别标签为所述采样样本图像所属的样本图像集对应的类别标签;所述第一概率向量包括A个概率元素,所述A个概率元素分别指代所述A个类别标签中的一个,一个所述概率元素用于表征所述采样样本图像属于所指代的类别标签的概率;所述对抗样本图像被所述图像识别模型进行类别预测所得到的类别预测结果为所述错误类别标签;通过所述图像识别优化模型,生成所述目标图像针对所述A个类别标签的目标概率向量;将所述目标概率向量中的最大概率元素对应的类别标签,确定为所述目标图像的类别预测结果。
- 一种数据处理装置,其特征在于,包括:数据获取模块,用于获取A个样本图像集以及图像识别模型,A为大于1的正整数,所述A个样本图像集分别对应A个不同的类别标签中的一个,所述图像识别模型是根据所述A个样本图像集预训练得到的,所述A个类别标签包括错误类别标签;第一输入模块,用于在所述A个样本图像集中获取采样样本图像,将所述采样样本图像与所述错误类别标签输入至所述图像识别模型,以生成所述采样样本图像针对所述A个类别标签的第一概率向量,所述错误类别标签与所述采样样本图像所属的样本图像集对应的类别标签不同;所述第一概率向量包括A个概率元素,所述A个概率元素分别指代所述A个类别标签中的一个,一个所述概率元素用于表征所述采样样本图像属于所指代的类别标签的概率;第一调整模块,用于在所述第一概率向量中的所述A个概率元素中,获取指代所述错误类别标签的第一概率元素,根据所述第一概率元素,对所述采样样本图像进行调整,得到所述采样样本图像对应的对抗样本图像,所述对抗样本图像被所述图像识别模型进行类别预测所得到的类别预测结果为所述错误类别标签。
- 一种数据处理装置,其特征在于,包括:图像获取模块,用于获取目标图像,将所述目标图像输入至图像识别优化模型;所述图像识别优化模型是根据混合样本训练集继续对图像识别模型进行优化训练所得到的;所述混合样本训练集包括A个样本图像集、A个类别标签、正确类别标签以及对抗样本图像;所述A个样本图像集分别对应A个不同的类别标签中的一个,所述A个样本图像集包括采集样本图像;所述正确类别标签属于所述A个类别标签;A为大于1的正整数;所述图像识别模型是根据所述A个样本图像集预训练得到的;所述对抗样本图像是通过第一概率向量中的A个概率元素中指代错误类别标签的第一概率元素对采样样本图像进行调整得到,所述第一概率向量是由所述图像识别模型针对所述采样样本图像所生成的,所述错误类别标签与所述采样样本图像所属的样本图像集对应的类别标签不同,所述正确类别标签为所述采样样本图像所属的样本图像集对应的类别标签;所述第一概率向量包括A个概率元素,所述A个概率元素分别指代所述A个类别标签中的一个,一个所述概率元素用于表征所述采样样本图像属于所指代的类别标签的概率;所述对抗样本图像被所述图像识别模型进行类别预测所得到的类别预测结果为所述错误类别标签;向量生成模块,用于通过所述图像识别优化模型,生成所述目标图像针对所述A个类别标签的目标概率向量;类别确定模块,用于将所述目标概率向量中的最大概率元素对应的类别标签,确定为所述目标图像的类别预测结果。
- 一种计算机设备,其特征在于,包括:处理器、存储器以及网络接口;所述处理器与所述存储器、所述网络接口相连,其中,所述网络接口用于提供数据通信功能,所述存储器用于存储计算机程序,所述处理器用于调用所述计算机程序,以使得所述计算机设备执行权利要求1至11任一项所述的方法。
- 一种计算机可读存储介质,其特征在于,所述计算机可读存储介质中存储有计算机程序,所述计算机程序适于由处理器加载并执行,以使得具有所述处理器的计算机设备执行权利要求1-11任一项所述的方法。
- 一种计算机程序产品,其特征在于,所述计算机程序产品包括计算机程序,所述计算机程序存储在计算机可读存储介质中,所述计算机程序适于由处理器读取并执行,以使得具有所述处理器的计算机设备执行权利要求1-11任一项所述的方法。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24763255.7A EP4614456A4 (en) | 2023-03-02 | 2024-03-01 | DATA PROCESSING METHOD, AND COMPUTER-READABLE STORAGE DEVICE AND MEDIUM |
| US19/072,375 US20250200954A1 (en) | 2023-03-02 | 2025-03-06 | Data processing method, device, and computer-readable storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310237098.1A CN116977692A (zh) | 2023-03-02 | 2023-03-02 | 一种数据处理方法、设备以及计算机可读存储介质 |
| CN202310237098.1 | 2023-03-02 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US19/072,375 Continuation US20250200954A1 (en) | 2023-03-02 | 2025-03-06 | Data processing method, device, and computer-readable storage medium |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024179575A1 true WO2024179575A1 (zh) | 2024-09-06 |
Family
ID=88471972
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/079587 Ceased WO2024179575A1 (zh) | 2023-03-02 | 2024-03-01 | 一种数据处理方法、设备以及计算机可读存储介质 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US20250200954A1 (zh) |
| EP (1) | EP4614456A4 (zh) |
| CN (1) | CN116977692A (zh) |
| WO (1) | WO2024179575A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN119004222A (zh) * | 2024-10-23 | 2024-11-22 | 中国科学技术大学 | 提升人体行为识别系统对抗攻击识别准确率的方法及系统 |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116977692A (zh) * | 2023-03-02 | 2023-10-31 | 腾讯科技(深圳)有限公司 | 一种数据处理方法、设备以及计算机可读存储介质 |
| CN120782732B (zh) * | 2025-06-30 | 2026-02-03 | 广东工业大学 | 基于机器视觉的输电线路绝缘子缺陷检测方法和系统 |
| CN121259493B (zh) * | 2025-12-05 | 2026-03-03 | 中国人民解放军国防科技大学 | 数据变换触发增强数字域对抗样本生成方法、装置及设备 |
Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110210617A (zh) * | 2019-05-15 | 2019-09-06 | 北京邮电大学 | 一种基于特征增强的对抗样本生成方法及生成装置 |
| CN111046380A (zh) * | 2019-12-12 | 2020-04-21 | 支付宝(杭州)信息技术有限公司 | 一种基于对抗样本增强模型抗攻击能力的方法和系统 |
| WO2021189364A1 (zh) * | 2020-03-26 | 2021-09-30 | 深圳先进技术研究院 | 一种对抗图像生成方法、装置、设备以及可读存储介质 |
| CN113569611A (zh) * | 2021-02-08 | 2021-10-29 | 腾讯科技(深圳)有限公司 | 图像处理方法、装置、计算机设备和存储介质 |
| EP4075395A2 (en) * | 2021-08-25 | 2022-10-19 | Beijing Baidu Netcom Science Technology Co., Ltd. | Method and apparatus of training anti-spoofing model, method and apparatus of performing anti-spoofing, and device |
| US20230022943A1 (en) * | 2021-07-22 | 2023-01-26 | Xidian University | Method and system for defending against adversarial sample in image classification, and data processing terminal |
| CN116977692A (zh) * | 2023-03-02 | 2023-10-31 | 腾讯科技(深圳)有限公司 | 一种数据处理方法、设备以及计算机可读存储介质 |
-
2023
- 2023-03-02 CN CN202310237098.1A patent/CN116977692A/zh active Pending
-
2024
- 2024-03-01 EP EP24763255.7A patent/EP4614456A4/en active Pending
- 2024-03-01 WO PCT/CN2024/079587 patent/WO2024179575A1/zh not_active Ceased
-
2025
- 2025-03-06 US US19/072,375 patent/US20250200954A1/en active Pending
Patent Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110210617A (zh) * | 2019-05-15 | 2019-09-06 | 北京邮电大学 | 一种基于特征增强的对抗样本生成方法及生成装置 |
| CN111046380A (zh) * | 2019-12-12 | 2020-04-21 | 支付宝(杭州)信息技术有限公司 | 一种基于对抗样本增强模型抗攻击能力的方法和系统 |
| WO2021189364A1 (zh) * | 2020-03-26 | 2021-09-30 | 深圳先进技术研究院 | 一种对抗图像生成方法、装置、设备以及可读存储介质 |
| CN113569611A (zh) * | 2021-02-08 | 2021-10-29 | 腾讯科技(深圳)有限公司 | 图像处理方法、装置、计算机设备和存储介质 |
| US20230022943A1 (en) * | 2021-07-22 | 2023-01-26 | Xidian University | Method and system for defending against adversarial sample in image classification, and data processing terminal |
| EP4075395A2 (en) * | 2021-08-25 | 2022-10-19 | Beijing Baidu Netcom Science Technology Co., Ltd. | Method and apparatus of training anti-spoofing model, method and apparatus of performing anti-spoofing, and device |
| CN116977692A (zh) * | 2023-03-02 | 2023-10-31 | 腾讯科技(深圳)有限公司 | 一种数据处理方法、设备以及计算机可读存储介质 |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP4614456A4 * |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN119004222A (zh) * | 2024-10-23 | 2024-11-22 | 中国科学技术大学 | 提升人体行为识别系统对抗攻击识别准确率的方法及系统 |
Also Published As
| Publication number | Publication date |
|---|---|
| EP4614456A1 (en) | 2025-09-10 |
| US20250200954A1 (en) | 2025-06-19 |
| CN116977692A (zh) | 2023-10-31 |
| EP4614456A4 (en) | 2025-09-10 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20200342314A1 (en) | Method and System for Detecting Fake News Based on Multi-Task Learning Model | |
| CN112559800B (zh) | 用于处理视频的方法、装置、电子设备、介质和产品 | |
| US20210295114A1 (en) | Method and apparatus for extracting structured data from image, and device | |
| CN109034069B (zh) | 用于生成信息的方法和装置 | |
| EP4614456A1 (en) | Data processing method, and device and computer-readable storage medium | |
| CN112131881B (zh) | 信息抽取方法及装置、电子设备、存储介质 | |
| US20240320807A1 (en) | Image processing method and apparatus, device, and storage medium | |
| CN114219971B (zh) | 一种数据处理方法、设备以及计算机可读存储介质 | |
| CN116975347B (zh) | 图像生成模型训练方法及相关装置 | |
| CN116980541A (zh) | 视频编辑方法、装置、电子设备以及存储介质 | |
| CN114282258A (zh) | 截屏数据脱敏方法、装置、计算机设备及存储介质 | |
| CN113361629A (zh) | 一种训练样本生成的方法、装置、计算机设备及存储介质 | |
| CN116977457A (zh) | 一种数据处理方法、设备以及计算机可读存储介质 | |
| CN116910199A (zh) | 基于人工智能的智能问答处理方法、装置、设备及介质 | |
| CN114820885A (zh) | 图像编辑方法及其模型训练方法、装置、设备和介质 | |
| CN116304146A (zh) | 图像处理方法及相关装置 | |
| CN111491209A (zh) | 视频封面确定方法、装置、电子设备和存储介质 | |
| CN118823184A (zh) | 图像生成、大模型的训练、图像处理方法及装置、设备和介质 | |
| CN111353554B (zh) | 预测缺失的用户业务属性的方法及装置 | |
| CN111553748B (zh) | 一种基于用户场景的Android微服务推荐方法与系统 | |
| CN116150415B (zh) | 用户画像的构建方法、装置和电子设备 | |
| CN115880702A (zh) | 数据处理方法、装置、设备、程序产品及存储介质 | |
| CN115909390A (zh) | 低俗内容识别方法、装置、计算机设备以及存储介质 | |
| CN117540306B (zh) | 一种多媒体数据的标签分类方法、装置、设备及介质 | |
| CN115952295B (zh) | 基于图像的实体关系标注模型处理方法及其相关设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24763255 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2024763255 Country of ref document: EP |
|
| ENP | Entry into the national phase |
Ref document number: 2024763255 Country of ref document: EP Effective date: 20250605 |
|
| WWP | Wipo information: published in national office |
Ref document number: 2024763255 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |