WO2021036715A1 - 一种图文融合方法、装置及电子设备 - Google Patents

一种图文融合方法、装置及电子设备 Download PDF

Info

Publication number
WO2021036715A1
WO2021036715A1 PCT/CN2020/106900 CN2020106900W WO2021036715A1 WO 2021036715 A1 WO2021036715 A1 WO 2021036715A1 CN 2020106900 W CN2020106900 W CN 2020106900W WO 2021036715 A1 WO2021036715 A1 WO 2021036715A1
Authority
WO
WIPO (PCT)
Prior art keywords
text
image
pixel
parameter
feature
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2020/106900
Other languages
English (en)
French (fr)
Inventor
张文杰
钟伟才
胡靓
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Huawei Technologies Co Ltd
Original Assignee
Huawei Technologies Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Huawei Technologies Co Ltd filed Critical Huawei Technologies Co Ltd
Priority to EP20858455.7A priority Critical patent/EP3996046A4/en
Priority to US17/634,002 priority patent/US12254544B2/en
Publication of WO2021036715A1 publication Critical patent/WO2021036715A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T11/00—Two-dimensional [2D] image generation
    • G06T11/60—Creating or editing images; Combining images with text
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T11/00—Two-dimensional [2D] image generation
    • G06T11/10—Texturing; Colouring; Generation of textures or colours
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T5/00—Image enhancement or restoration
    • G06T5/20—Image enhancement or restoration using local operators
    • G06T5/30—Erosion or dilatation, e.g. thinning
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T5/00—Image enhancement or restoration
    • G06T5/40—Image enhancement or restoration using histogram techniques
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T5/00—Image enhancement or restoration
    • G06T5/70—Denoising; Smoothing
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00—Image analysis
    • G06T7/10—Segmentation; Edge detection
    • G06T7/11—Region-based segmentation
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00—Image analysis
    • G06T7/10—Segmentation; Edge detection
    • G06T7/13—Edge detection
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00—Image analysis
    • G06T7/10—Segmentation; Edge detection
    • G06T7/136—Segmentation; Edge detection involving thresholding
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00—Image analysis
    • G06T7/40—Analysis of texture
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00—Arrangements for image or video recognition or understanding
    • G06V10/40—Extraction of image or video features
    • G06V10/44—Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00—Arrangements for image or video recognition or understanding
    • G06V10/40—Extraction of image or video features
    • G06V10/46—Descriptors for shape, contour or point-related descriptors, e.g. scale invariant feature transform [SIFT] or bags of words [BoW]; Salient regional features
    • G06V10/462—Salient features, e.g. scale invariant feature transforms [SIFT]
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00—Arrangements for image or video recognition or understanding
    • G06V10/40—Extraction of image or video features
    • G06V10/54—Extraction of image or video features relating to texture
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00—Arrangements for image or video recognition or understanding
    • G06V10/40—Extraction of image or video features
    • G06V10/56—Extraction of image or video features relating to colour
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00—Arrangements for image or video recognition or understanding
    • G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
    • G06V10/80—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00—Scenes; Scene-specific elements
    • G06V20/60—Type of objects
    • G06V20/62—Text, e.g. of license plates, overlay texts or captions on TV images
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16—Human faces, e.g. facial parts, sketches or expressions
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
    • G06V40/16—Human faces, e.g. facial parts, sketches or expressions
    • G06V40/168—Feature extraction; Face representation
    • Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02D—CLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
    • Y02D10/00—Energy efficient computing, e.g. low power processors, power management or thermal management

Definitions

  • the embodiments of the present application relate to the technical field of digital image processing, and in particular, to a method, device and electronic device for fusion of images and text.
  • the lock screen wallpaper of the mobile phone the advertisement video when the application (APP) starts, the floating advertisement frame of the video window, etc.
  • the embodiment of the present application provides a method for fusion of graphics and text, which can obtain a better typesetting effect when typesetting text into an image.
  • a method for image-text fusion includes: acquiring a first image and a first text to be typeset into the first image; determining the feature value of each pixel in the first image; wherein, one pixel The feature value of a point is used to characterize the probability that the pixel is paid attention to by the user.
  • the feature value of each pixel in the first image determines the multiple first typesetting layouts of the first text in the first image; wherein, when the first text is typeset into the first image according to each first typesetting layout, the first text is not blocked A pixel with a feature value greater than the first threshold; according to the cost parameters of the plurality of first typesetting formats, a second typesetting format is determined from the plurality of first typesetting formats; wherein, a cost parameter of the first typesetting format is used to characterize When the first text is typeset into the first image according to the first typesetting format, the size of the feature value of the pixel that the first text occludes, and the distribution of the feature value of the pixel of each area in the first image typeset with the first text The degree of balance; the first text is typeset to the first image according to the second typesetting format to obtain the second image.
  • the technical solution provided by the above-mentioned first aspect can determine the typesetting positions of multiple candidate text templates and corresponding multiple texts in the image, so that the text typeset into the image does not obscure the visually significant images with higher feature values Subjects, such as human faces, buildings, etc. Then according to different text templates and the corresponding layout position in the image to typeset the image into the image, the size of the feature value of the pixel shaded by the text, and the balance of the feature value distribution of the pixel in each area of the image typeset with the text Determine the final text template of the text and the typesetting position of the text in the image to obtain a better typesetting effect.
  • determining the feature value of each pixel in the first image includes: determining the visually significant parameters of each pixel in the first image, face feature parameters, edge feature parameters, and text feature parameters. At least two parameters of a pixel; among them, the visual saliency parameter of a pixel is used to characterize the possibility that the pixel is the pixel corresponding to the visual saliency feature, and the facial feature parameter of a pixel is used to characterize the one Pixel is the probability of the pixel corresponding to the face.
  • the edge feature parameter of a pixel is used to characterize the probability that the pixel is the pixel corresponding to the contour of the object.
  • the text feature parameter of a pixel is used In order to characterize the probability that the pixel is the pixel corresponding to the text; the visually significant parameters of each pixel in the first image, the facial feature parameters, the edge feature parameters, and the text feature parameters are determined respectively. The two parameters are weighted and summed to determine the characteristic value of each pixel in the first image. At least two of the visually significant parameters, face feature parameters, edge feature parameters, and text feature parameters, as well as the likelihood of the feature corresponding to each parameter being paid attention to by the user, can be comprehensively considered to determine the feature value of each pixel. It is more likely that the typesetting position of the text determined according to the feature value will not obscure the salient features.
  • the method further includes: according to the determined visually significant parameters, facial feature parameters, edge feature parameters, and text feature parameters of each pixel in the first image.
  • At least two parameters are used to generate at least two feature maps; the pixel value of each pixel in each feature map is the corresponding parameter of the corresponding pixel; the visually significant parameters of each pixel in the first image are determined, and the human face
  • the feature parameter, the weighted summation of at least two parameters of the edge feature parameter and the text feature parameter to determine the feature value of each pixel in the first image includes: the pixel value of each pixel in the at least two feature maps described above Perform a weighted summation to determine the feature value of each pixel in the first image.
  • determining multiple first typesetting layouts of the first text in the first image includes: according to the first text The size of the text box of the first text when typesetting using one or more text templates, and the characteristic value of each pixel in the first image, determine a plurality of first typesetting layouts.
  • the method further includes: obtaining one or more text templates, each of which specifies the line spacing, line width, font size, font, text thickness, alignment, position of decorative lines, and At least one of the thickness of the decorative line.
  • determining the second typesetting format from the plurality of first typesetting formats according to the cost parameters of the plurality of first typesetting formats includes: determining that the first text is respectively typeset according to the plurality of first typesetting formats.
  • the text box of the first text blocks the texture feature parameter of the image area of the first image, and the texture feature parameter is used to characterize the number of texture features in the image corresponding to the image area;
  • the first typesetting format a plurality of first typesetting formats corresponding to the image area whose texture feature parameter is less than the second threshold is selected; according to the selected cost parameter of each first typesetting format, from the selected multiple first typesetting formats
  • the second layout is determined in the layout.
  • the method further includes: for each first typesetting format in the plurality of first typesetting formats, performing at least two of step a, step b, and step c, and step d to obtain each first typesetting format.
  • the first parameter is the sum of the feature values of each pixel in the image area where the first text occludes the first image
  • the second parameter is the area of the image area, or the second parameter is the number of pixels in the image area
  • the total number, or, the second parameter is the product of the total number of pixels in the image area and the preset value; step b.
  • step c calculate that when the first text is typeset into the first image according to a first typesetting format, the first The visual balance parameter of the text; the visual balance parameter is used to characterize the degree of influence of the first text on the balance of the feature value distribution of the pixel points in each area of the first image typeset with the first text; step d, calculate according to At least two of the text intrusion parameter, the visual space occupancy parameter, and the visual balance parameter of the first text of the first text are calculated, and the cost parameter of a first typesetting format is calculated.
  • determining the second typesetting format from the multiple first typesetting formats according to the cost parameters of the plurality of first typesetting formats includes: determining the minimum cost parameter of the plurality of first typesetting formats The first typesetting format corresponding to the cost parameter of is the second typesetting format.
  • the method further includes: determining the color parameter of the first text; the color parameter of the first text is that when the first text is typeset into the first image according to the second typesetting format, the first image is The derivative color of the main color of the image area occluded by a text; the derivative color of the main color refers to the same hue as the main color, but the hue, saturation and lightness are different from the HSV of the main color; according to the color parameters of the first text , Color the first text in the second image to obtain the third image.
  • the color of the first text is more coordinated with the background image, and the display is clearer.
  • the main color of the image area where the first image is obscured by the first text is based on the fact that the first image is obscured by the first text when the first text is typeset into the first image according to the second typesetting format.
  • the hue, saturation, and brightness of the three primary colors RGB in the image area are determined in the HSV space; the dominant color is the hue with the highest hue proportion in the image area.
  • the method further includes: if at least one of the following conditions 1 and 2 is satisfied, determining to perform rendering processing on the second image; condition 1: the first text is typeset according to the second typesetting format to After the first image, the texture feature parameter of the image area where the first text occludes the first image is greater than the fourth threshold; the texture feature parameter is used to characterize the number of texture features in the image corresponding to the image area; condition 2: The dominant color ratio is less than the fifth threshold; the mask layer is overlaid on the second image; or, the mask parameters are determined, and the second image is processed according to the determined mask parameters; or, the first text is projected and rendered.
  • condition 1 the first text is typeset according to the second typesetting format to After the first image, the texture feature parameter of the image area where the first text occludes the first image is greater than the fourth threshold; the texture feature parameter is used to characterize the number of texture features in the image corresponding to the image area; condition 2: The dominant color ratio is less than the fifth threshold; the mask layer is overlaid
  • the method further includes: if at least one of the following conditions 1 and 2 is satisfied, determining to perform rendering processing on the third image;
  • Condition 1 the first text is typeset according to the second typesetting format to After the first image, the texture feature parameter of the image area where the first text occludes the first image is greater than the fourth threshold; the texture feature parameter is used to characterize the number of texture features in the image corresponding to the image area;
  • condition 2 The dominant color ratio is less than the fifth threshold; the mask layer is overlaid on the third image; or, the mask parameters are determined, and the third image is processed according to the determined mask parameters; or, the first text is projected and rendered.
  • an image-text fusion device in a second aspect, includes: an information acquisition unit for acquiring a first image and a first text to be typeset into the first image; an analysis unit for determining that the first image is The feature value of each pixel; among them, the feature value of a pixel is used to characterize the probability of the pixel being paid attention to by the user.
  • the first text is typeset according to each first typesetting layout
  • the first text does not block pixels whose feature value is greater than the first threshold
  • the second typesetting is determined from the plurality of first typesetting formats Layout; among them, a cost parameter of the first typesetting format is used to characterize the size of the feature value of the pixel shaded by the first text when the first text is typeset into the first image according to the first typesetting format, and the first text is typeset
  • a cost parameter of the first typesetting format is used to characterize the size of the feature value of the pixel shaded by the first text when the first text is typeset into the first image according to the first typesetting format
  • the first text is typeset The degree of balance of the feature value distribution of the pixel points in each area of the first image; the processing unit typesets the first text to the first image according to the second typesetting format to obtain the second image.
  • the device provided in the second aspect described above can determine the typesetting positions of multiple candidate text templates and corresponding multiple texts in the image, so that the text typeset into the image does not obscure the visually significant subjects with higher feature values in the image , Such as human faces, buildings, etc. Then according to different text templates and the corresponding layout position in the image to typeset the image into the image, the size of the feature value of the pixel shaded by the text, and the balance of the feature value distribution of the pixel in each area of the image typeset with the text Determine the final text template of the text and the typesetting position of the text in the image to obtain a better typesetting effect.
  • the analysis unit determining the feature value of each pixel in the first image includes: the analysis unit determines the visually significant parameter of each pixel in the first image, face feature parameters, edge feature parameters, and At least two parameters in the text feature parameters; among them, the visually significant parameter of a pixel is used to characterize the probability that the pixel is the pixel corresponding to the visually significant feature, and the facial feature parameter of a pixel is used In order to characterize the probability that the pixel is the pixel corresponding to the human face, the edge feature parameter of a pixel is used to characterize the probability that the pixel is the pixel corresponding to the contour of the object.
  • the text feature parameters are used to characterize the probability that the pixel is the pixel corresponding to the text; the analysis unit respectively determines the visually significant parameters of each pixel in the first image, face feature parameters, edge feature parameters and At least two of the text feature parameters are weighted and summed to determine the feature value of each pixel in the first image. At least two of the visually significant parameters, face feature parameters, edge feature parameters, and text feature parameters, as well as the likelihood of the feature corresponding to each parameter being paid attention to by the user, can be comprehensively considered to determine the feature value of each pixel. It is more likely that the typesetting position of the text determined according to the feature value will not obscure the salient features.
  • the analysis unit respectively performs weighted calculation on at least two of the determined visually significant parameters, face feature parameters, edge feature parameters, and text feature parameters of each pixel in the first image. And, before determining the feature value of each pixel in the first image, the analysis unit is also used to: respectively determine the visually significant parameters, face feature parameters, edge feature parameters, and text of each pixel in the first image.
  • At least two of the feature parameters are used to generate at least two feature maps of the first image; the pixel value of each pixel in each feature map is the corresponding parameter of the corresponding pixel; and each of the determined first image
  • the weighted summation of at least two of the visually significant parameters of the pixels, the face feature parameters, the edge feature parameters, and the text feature parameters to determine the feature value of each pixel in the first image includes: The pixel value of each pixel in the feature map is weighted and summed to determine the feature value of each pixel in the first image.
  • the analysis unit determines multiple first typesetting formats of the first text in the first image according to the first text and the feature values of each pixel in the first image, including: the analysis unit according to When the first text is typeset using one or more text templates, the size of the text box of the first text and the feature value of each pixel in the first image determine a plurality of first typesetting layouts.
  • the information obtaining unit is also used to: obtain one or more text templates, each of which specifies the line spacing, line width, font size, font, text thickness, alignment, and decoration of the text. At least one of the line position and the thickness of the decorative line.
  • the analyzing unit determines the second typesetting format from the plurality of first typesetting formats according to the cost parameters of the plurality of first typesetting formats, including: the analyzing unit determines that the first text is in accordance with the above-mentioned multiple When the first typesetting layout is typeset to the first image, the text box of the first text blocks the texture feature parameter of the image area of the first image, and the texture feature parameter is used to represent the number of texture features in the image corresponding to the image area ; The analysis unit selects a plurality of first typesetting formats corresponding to image regions with texture feature parameters less than the second threshold from a plurality of first typesetting formats; according to the selected cost parameter of each first typesetting format, select from The second typesetting format is determined from the multiple first typesetting formats.
  • the analysis unit is further configured to: for each first typesetting format in the plurality of first typesetting formats, perform at least two of step a, step b, and step c, and step d to obtain The cost parameter of each first typesetting format;
  • Step a Calculate the text intrusion parameters of the first text when the first text is typeset into the first image according to a first typesetting format; the text intrusion parameters are the first parameter and the second The ratio of the parameters;
  • the first parameter is the sum of the feature values of each pixel in the image area where the first text occludes the first image;
  • the second parameter is the area of the image area, or the second parameter is the pixel in the image area Or, the second parameter is the product of the total number of pixels in the image area and the preset value; step b.
  • step c calculate that when the first text is typeset to the first image according to a first typesetting format A visual balance parameter of the text; the visual balance parameter is used to represent the degree of influence of the first text on the balance of the feature value distribution of the pixel points in each area of the first image typeset with the first text; step d, according to calculation At least two of the text intrusion parameter, the visual space occupancy parameter, and the visual balance parameter of the first text are calculated, and a cost parameter of the first typesetting format is calculated.
  • the cost parameter T i of the format where E s (L i ) is the text intrusion parameter of the first text when the first text is typeset into the first image according to the first typesetting format, and Eu (L i ) is the first when the text according to a first layout format to the first image
  • the analyzing unit determines the second typesetting format from the plurality of first typesetting formats according to the cost parameters of the plurality of first typesetting formats, including: the analyzing unit determines the cost of the plurality of first typesetting formats Among the parameters, the first typesetting format corresponding to the smallest cost parameter is the second typesetting format.
  • the analysis unit is further used to: determine the color parameter of the first text; the color parameter of the first text is that when the first text is typeset into the first image according to the second typesetting format, the first image is The derivative color of the main color of the image area occluded by the first text; the derivative color of the main color refers to the same hue as the main color, but the hue, saturation and lightness are different from the HSV of the main color; and according to the first text Color parameter, color the first text in the second image to obtain the third image.
  • the color of the first text is more coordinated with the background image, and the display is clearer.
  • the main color of the image area where the first image is obscured by the first text is based on the fact that the first image is obscured by the first text when the first text is typeset into the first image according to the second typesetting format.
  • the hue, saturation, and brightness of the three primary colors RGB in the image area are determined in the HSV space; the dominant color is the hue with the highest hue proportion in the image area.
  • the analysis unit is further configured to: if at least one of the following conditions 1 and 2 is satisfied, determine to perform rendering processing on the second image; condition 1: the first text is typeset according to the second typesetting format After the first image, the texture feature parameter of the image area where the first text occludes the first image is greater than the fourth threshold; the texture feature parameter is used to characterize the number of texture features in the image corresponding to the image area; condition 2: image area The main color ratio of is less than the fifth threshold; the processing unit is also used to overlay the mask layer on the second image; or, determine the mask parameters, and process the second image according to the determined mask parameters; or The first text is projectively rendered. Through mask rendering or projection rendering, the clarity and salience of the first text can be improved.
  • the analysis unit is further configured to: if at least one of the following conditions 1 and 2 is satisfied, determine to perform rendering processing on the third image; condition 1: the first text is typeset according to the second typesetting format After the first image, the texture feature parameter of the image area where the first text occludes the first image is greater than the fourth threshold; the texture feature parameter is used to characterize the number of texture features in the image corresponding to the image area; condition 2: image area The main color ratio of is less than the fifth threshold; the processing unit is also used to overlay the mask layer on the third image; or, determine the mask parameters, and process the third image according to the determined mask parameters; or, The first text is projectively rendered. Through mask rendering or projection rendering, the clarity and salience of the first text can be improved.
  • an electronic device in a third aspect, includes: an information acquisition unit for acquiring a first image and a first text to be typeset into the first image; an analysis unit for determining each of the first images The feature value of a pixel; among them, the feature value of a pixel is used to characterize the probability of the pixel being paid attention to by the user.
  • the first text and the feature value of each pixel in the first image determine multiple first typesetting layouts of the first text in the first image; wherein, the first text is typeset according to each first typesetting layout to In the case of the first image, the first text does not block pixels whose feature values are greater than the first threshold; and, according to the cost parameters of the plurality of first typesetting formats, the second typesetting format is determined from the plurality of first typesetting formats ; Among them, a cost parameter of the first typesetting format is used to characterize the size of the feature value of the pixel shaded by the first text when the first text is typeset into the first image according to the first typesetting format, and the typeset of the first text The degree of balance of the feature value distribution of the pixel points in each area in the first image; the processing unit typesets the first text into the first image according to the second typesetting format to obtain the second image.
  • the electronic device provided in the above third aspect can determine the typesetting positions of multiple candidate text templates and corresponding multiple texts in the image, so that the text typeset into the image will not obscure the visually significant images with higher feature values.
  • Subjects such as human faces, buildings, etc.
  • the size of the feature value of the pixel shaded by the text, and the balance of the feature value distribution of the pixel in each area of the image typeset with the text Determine the final text template of the text and the typesetting position of the text in the image to obtain a better typesetting effect.
  • the analysis unit determining the feature value of each pixel in the first image includes: the analysis unit determines the visually significant parameter of each pixel in the first image, face feature parameters, edge feature parameters, and At least two parameters in the text feature parameters; among them, the visually significant parameter of a pixel is used to characterize the probability that the pixel is the pixel corresponding to the visually significant feature, and the facial feature parameter of a pixel is used In order to characterize the probability that the pixel is the pixel corresponding to the human face, the edge feature parameter of a pixel is used to characterize the probability that the pixel is the pixel corresponding to the contour of the object.
  • the text feature parameters are used to characterize the probability that the pixel is the pixel corresponding to the text; the analysis unit respectively determines the visually significant parameters of each pixel in the first image, face feature parameters, edge feature parameters and At least two of the text feature parameters are weighted and summed to determine the feature value of each pixel in the first image. At least two of the visually significant parameters, face feature parameters, edge feature parameters, and text feature parameters, as well as the likelihood of the feature corresponding to each parameter being paid attention to by the user, can be comprehensively considered to determine the feature value of each pixel. It is more likely that the typesetting position of the text determined according to the feature value will not obscure the salient features.
  • the analysis unit respectively performs weighted calculation on at least two of the determined visually significant parameters, face feature parameters, edge feature parameters, and text feature parameters of each pixel in the first image. And, before determining the feature value of each pixel in the first image, the analysis unit is also used to: respectively determine the visually significant parameters, face feature parameters, edge feature parameters, and text of each pixel in the first image.
  • At least two parameters in the feature parameters generate at least two feature maps; the pixel value of each pixel in each feature map is the corresponding parameter of the corresponding pixel; the visual significance of each pixel in the determined first image is significant
  • the weighted summation of at least two parameters among parameters, face feature parameters, edge feature parameters, and text feature parameters to determine the feature value of each pixel in the first image includes: performing a weighted summation on each pixel in the at least two feature maps. The pixel values of the points are weighted and summed to determine the characteristic value of each pixel in the first image.
  • the analysis unit determines multiple first typesetting formats of the first text in the first image according to the first text and the feature values of each pixel in the first image, including: the analysis unit according to When the first text is typeset using one or more text templates, the size of the text box of the first text and the feature value of each pixel in the first image determine a plurality of first typesetting layouts. By comprehensively analyzing the size of the area occupied by the first text in different text templates, and the feature value of each pixel in the first image, it can be ensured that the first text is typeset at the determined text layout position without blocking the first text. Distinctive features in the image.
  • the information obtaining unit is also used to: obtain one or more text templates, each of which specifies the line spacing, line width, font size, font, text thickness, alignment, and decoration of the text. At least one of the line position and the thickness of the decorative line.
  • the analyzing unit determines the second typesetting format from the plurality of first typesetting formats according to the cost parameters of the plurality of first typesetting formats, including: the analyzing unit determines that the first text is in accordance with the above-mentioned multiple When the first typesetting layout is typeset to the first image, the text box of the first text blocks the texture feature parameter of the image area of the first image, and the texture feature parameter is used to represent the number of texture features in the image corresponding to the image area ; The analysis unit selects a plurality of first typesetting formats corresponding to image regions with texture feature parameters less than the second threshold from a plurality of first typesetting formats; according to the selected cost parameter of each first typesetting format, select from The second typesetting format is determined from the multiple first typesetting formats.
  • the analysis unit is further configured to: for each first typesetting format in the plurality of first typesetting formats, perform at least two of step a, step b, and step c, and step d to obtain The cost parameter of each first typesetting format;
  • Step a Calculate the text intrusion parameters of the first text when the first text is typeset into the first image according to a first typesetting format; the text intrusion parameters are the first parameter and the second The ratio of the parameters;
  • the first parameter is the sum of the feature values of each pixel in the image area where the first text occludes the first image;
  • the second parameter is the area of the image area, or the second parameter is the pixel in the image area Or, the second parameter is the product of the total number of pixels in the image area and the preset value; step b.
  • step c calculate that when the first text is typeset to the first image according to a first typesetting format A visual balance parameter of the text; the visual balance parameter is used to represent the degree of influence of the first text on the balance of the feature value distribution of the pixel points in each area of the first image typeset with the first text; step d, according to calculation At least two of the text intrusion parameter, the visual space occupancy parameter, and the visual balance parameter of the first text are calculated, and a cost parameter of the first typesetting format is calculated.
  • the cost parameter T i of the format where E s (L i ) is the text intrusion parameter of the first text when the first text is typeset into the first image according to the first typesetting format, and Eu (L i ) is the first when the text according to a first layout format to the first image
  • the analyzing unit determines the second typesetting format from the plurality of first typesetting formats according to the cost parameters of the plurality of first typesetting formats, including: the analyzing unit determines the cost of the plurality of first typesetting formats Among the parameters, the first typesetting format corresponding to the smallest cost parameter is the second typesetting format.
  • the analysis unit is further used to: determine the color parameter of the first text; the color parameter of the first text is that when the first text is typeset into the first image according to the second typesetting format, the first image is The derivative color of the main color of the image area occluded by the first text; the derivative color of the main color refers to the same hue as the main color, but the hue, saturation and lightness are different from the HSV of the main color; and according to the first text Color parameter, color the first text in the second image to obtain the third image.
  • the color of the first text is more coordinated with the background image, and the display is clearer.
  • the main color of the image area where the first image is obscured by the first text is based on the fact that the first image is obscured by the first text when the first text is typeset into the first image according to the second typesetting format.
  • the hue, saturation, and lightness of the three primary colors RGB in the image area are determined in the HSV space; the dominant color is the hue color with the highest hue proportion in the image area.
  • the analysis unit is further configured to: if at least one of the following conditions 1 and 2 is satisfied, determine to perform rendering processing on the second image; condition 1: the first text is typeset according to the second typesetting format After the first image, the texture feature parameter of the image area where the first text occludes the first image is greater than the fourth threshold; the texture feature parameter is used to characterize the number of texture features in the image corresponding to the image area; condition 2: image area The main color ratio of is less than the fifth threshold; the processing unit is also used to overlay the mask layer on the second image; or, determine the mask parameters, and process the second image according to the determined mask parameters; or The first text is projectively rendered. Through mask rendering or projection rendering, the clarity and salience of the first text can be improved.
  • the analysis unit is further configured to: if at least one of the following conditions 1 and 2 is satisfied, determine to perform rendering processing on the third image; condition 1: the first text is typeset according to the second typesetting format After the first image, the texture feature parameter of the image area where the first text occludes the first image is greater than the fourth threshold; the texture feature parameter is used to characterize the number of texture features in the image corresponding to the image area; condition 2: image area The main color ratio of is less than the fifth threshold; the processing unit is also used to overlay the mask layer on the third image; or, determine the mask parameters, and process the third image according to the determined mask parameters; or, The first text is projectively rendered. Through mask rendering or projection rendering, the clarity and salience of the first text can be improved.
  • a graphic-text fusion device in a fourth aspect, includes: a memory for storing one or more computer programs; a processor for executing one or more computer programs stored in the memory to make the graphic-text fusion device Realize the image and text fusion method in any possible implementation manner of the first aspect.
  • an electronic device in a fifth aspect, includes: a memory for storing one or more computer programs; a processor for executing one or more computer programs stored in the memory, so that the image-text fusion device realizes Such as the image and text fusion method in any possible implementation of the first aspect.
  • a computer-readable storage medium is provided, and computer-executable instructions are stored on the computer-readable storage medium. Text fusion method.
  • a chip system in a seventh aspect, includes a processor and a memory, and instructions are stored in the memory; when the instructions are executed by the processor, the implementation is as in any one of the possible implementation manners of the first aspect.
  • Image and text fusion method The chip system can be composed of chips, or it can include chips and other discrete devices.
  • a computer program product is provided, and a computer program product is provided, which when running on a computer, enables the image-text fusion method in any possible implementation manner of the first aspect.
  • the computer may be at least one storage node.
  • FIG. 1 is an example diagram of image-text fusion provided by an embodiment of the application
  • FIG. 2 is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the application.
  • FIG. 3 is an example of an application scenario of image and text fusion provided by an embodiment of the application
  • FIG. 4 is an example of another application scenario of image and text fusion provided by an embodiment of the application.
  • FIG. 5 is an example of another application scenario of image and text fusion provided by an embodiment of the application.
  • FIG. 6 is a flowchart of a method for image-text fusion provided by an embodiment of the application.
  • FIG. 7 is an example diagram of the same image corresponding to different texts according to an embodiment of the application.
  • FIG. 8 is an example diagram of a process of generating a feature map of a salient region provided by an embodiment of the application.
  • FIG. 9 is an example diagram of another salient region feature map generation process provided by an embodiment of the application.
  • FIG. 10 is a schematic diagram of a process of generating a visual saliency feature map provided by an embodiment of this application.
  • FIG. 11 is a flowchart of a face detection algorithm provided by an embodiment of the application.
  • FIG. 12 is a flowchart of an edge detection algorithm provided by an embodiment of this application.
  • FIG. 13 is a flowchart of a text detection algorithm provided by an embodiment of the application.
  • FIG. 14 is an example diagram of a typesetting specification in JSON format provided by an embodiment of the application.
  • FIG. 15 is a schematic diagram of a layout position of a text box provided by an embodiment of the application.
  • FIG. 16 is an example diagram of several text templates provided by an embodiment of the application.
  • Figure 17 is an example diagram of several candidate typesetting layouts provided by an embodiment of the application.
  • FIG. 19 is a flowchart of another method for determining a second typesetting format provided by an embodiment of the application.
  • FIG. 20 is a flowchart of another image-text fusion method provided by an embodiment of the application.
  • FIG. 21 is a comparison diagram of several image-text fusion images provided by an embodiment of the application.
  • FIG. 22 is a schematic structural diagram of an electronic device provided by an embodiment of the application.
  • the embodiment of the present application provides a method for merging graphics and text, which can be applied to the process of typesetting text into an image (such as the first image) to achieve the fusion of graphics and text.
  • an image such as the first image
  • the text will generally cover a part of the image.
  • the text typeset into the image will not obscure the salient features in the image.
  • the first image is an image including building features
  • the text to be typeset into the first image is a text with a theme of "Glimpse of the Countryside”.
  • the text with the theme "Glimpse of the Countryside” can be typeset to an appropriate position in the first image including the features of the building, so that the text minimizes the occlusion of the salient features in the image, and further It can also make the text typeset to the first image to obtain a better degree of visual balance.
  • salient features refer to image features that are more likely to be paid attention to by users.
  • salient features may include facial features, human body features, building features, object features (such as animal features, tree features, flower features, etc.), text features, river features, and mountain features.
  • object features such as animal features, tree features, flower features, etc.
  • text features such as river features, and mountain features.
  • the salient feature is the building feature.
  • the image-text fusion method in the embodiment of the present application can be applied to terminal electronic devices that can provide image display. Including desktop devices, laptop devices, handheld devices, wearable devices, etc. For example, it is applied to mobile phones, tablet computers, personal computers, smart cameras, netbooks, personal digital assistants (PDAs), smart watches, AR (augmented reality)/VR (virtual reality) devices, etc.
  • PDAs personal digital assistants
  • AR augmented reality
  • VR virtual reality
  • the image-text fusion method of the embodiment of the present application can also be applied to an image processing device or server-type electronic device (for example, an application server) with or without an image display function.
  • an image processing device or server-type electronic device for example, an application server
  • the embodiment of the present application does not limit the specific type and structure of the electronic device that executes the graphic-text fusion method of the embodiment of the present application.
  • the electronic device 200 may include a processor 210, a memory (including an external memory interface 220 and an internal memory 221), a universal serial bus (USB) interface 230, a charging management module 240, and a power management module 241, battery 242, antenna 1, antenna 2, mobile communication module 250, wireless communication module 260, audio module 270, speaker 270A, receiver 270B, microphone 270C, earphone jack 270D, sensor module 280, buttons 290, motor 291, indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.
  • SIM subscriber identification module
  • the sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, an air pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a gravity sensor 280I, a temperature sensor 280J, and a touch sensor.
  • the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the electronic device 200.
  • the electronic device 200 may include more or fewer components than shown, or combine certain components, or split certain components, or arrange different components.
  • the illustrated components can be implemented in hardware, software, or a combination of software and hardware.
  • the processor 210 may include one or more processing units.
  • the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), and an image signal processor. (image signal processor, ISP), controller, video codec, digital signal processor (digital signal processor, DSP), baseband processor, and/or neural-network processing unit (NPU), etc.
  • AP application processor
  • modem processor modem processor
  • GPU graphics processing unit
  • image signal processor image signal processor
  • ISP image signal processor
  • controller video codec
  • digital signal processor digital signal processor
  • DSP digital signal processor
  • NPU neural-network processing unit
  • the different processing units may be independent devices or integrated in one or more processors.
  • the controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
  • a memory may also be provided in the processor 210 for storing instructions and data.
  • the memory in the processor 210 is a cache memory.
  • the memory can store instructions or data that have just been used or recycled by the processor 210. If the processor 210 needs to use the instruction or data again, it can be directly called from the memory. Repeated access is avoided, the waiting time of the processor 210 is reduced, and the efficiency of the system is improved.
  • the processor 210 may include one or more interfaces.
  • Interfaces can include integrated circuit (I2C) interfaces, integrated circuit built-in audio (inter-integrated circuit sound, I2S) interfaces, pulse code modulation (PCM) interfaces, universal asynchronous transmitters receiver/transmitter, UART) interface, mobile industry processor interface (MIPI), general-purpose input/output (GPIO) interface, subscriber identity module (SIM) interface, and / Or Universal Serial Bus (USB) interface, etc.
  • I2C integrated circuit
  • I2S integrated circuit built-in audio
  • PCM pulse code modulation
  • UART mobile industry processor interface
  • MIPI mobile industry processor interface
  • GPIO general-purpose input/output
  • SIM subscriber identity module
  • USB Universal Serial Bus
  • the I2C interface is a bidirectional synchronous serial bus, which includes a serial data line (SDA) and a serial clock line (SCL).
  • the processor 210 may include multiple sets of I2C buses.
  • the processor 210 may be coupled to the touch sensor 280K, charger, flash, camera 293, etc., respectively through different I2C bus interfaces.
  • the processor 210 may couple the touch sensor 280K through an I2C interface, so that the processor 210 and the touch sensor 280K communicate through the I2C bus interface to implement the touch function of the electronic device 200.
  • the MIPI interface can be used to connect the processor 210 with the display screen 294, the camera 293 and other peripheral devices.
  • the MIPI interface includes a camera serial interface (camera serial interface, CSI), a display serial interface (display serial interface, DSI), and so on.
  • the processor 210 and the camera 293 communicate through a CSI interface to implement the shooting function of the electronic device 200.
  • the processor 210 and the display screen 294 communicate through a DSI interface to realize the display function of the electronic device 200.
  • the GPIO interface can be configured through software.
  • the GPIO interface can be configured as a control signal or as a data signal.
  • the GPIO interface can be used to connect the processor 210 with the camera 293, the display screen 294, the wireless communication module 260, the audio module 270, the sensor module 280, and so on.
  • the GPIO interface can also be configured as an I2C interface, I2S interface, UART interface, MIPI interface, etc.
  • the USB interface 230 is an interface that complies with the USB standard specifications, and specifically may be a Mini USB interface, a Micro USB interface, a USB Type C interface, and so on.
  • the USB interface 230 can be used to connect a charger to charge the electronic device 200, and can also be used to transfer data between the electronic device 200 and peripheral devices. It can also be used to connect earphones and play audio through earphones. This interface can also be used to connect to other electronic devices, such as AR devices.
  • the interface connection relationship between the modules illustrated in the embodiment of the present invention is merely a schematic description, and does not constitute a structural limitation of the electronic device 200.
  • the electronic device 200 may also adopt different interface connection modes in the foregoing embodiments, or a combination of multiple interface connection modes.
  • the charging management module 240 is used to receive charging input from the charger.
  • the charger can be a wireless charger or a wired charger.
  • the charging management module 240 may receive the charging input of the wired charger through the USB interface 230.
  • the charging management module 240 may receive the wireless charging input through the wireless charging coil of the electronic device 200. While the charging management module 240 charges the battery 242, it can also supply power to the electronic device through the power management module 241.
  • the power management module 241 is used to connect the battery 242, the charging management module 240 and the processor 210.
  • the power management module 241 receives input from the battery 242 and/or the charging management module 240, and supplies power to the processor 210, the internal memory 221, the display screen 294, the camera 293, and the wireless communication module 260.
  • the power management module 241 can also be used to monitor parameters such as battery capacity, battery cycle times, and battery health status (leakage, impedance).
  • the power management module 241 may also be provided in the processor 210.
  • the power management module 241 and the charging management module 240 may also be provided in the same device.
  • the wireless communication function of the electronic device 200 can be implemented by the antenna 1, the antenna 2, the mobile communication module 250, the wireless communication module 260, the modem processor, and the baseband processor.
  • the antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals.
  • Each antenna in the electronic device 200 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization.
  • Antenna 1 can be multiplexed as a diversity antenna of a wireless local area network.
  • the antenna can be used in combination with a tuning switch.
  • the mobile communication module 250 may provide a wireless communication solution including 2G/3G/4G/5G and the like applied to the electronic device 200.
  • the mobile communication module 250 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), and the like.
  • the mobile communication module 250 can receive electromagnetic waves by the antenna 1, and perform processing such as filtering, amplifying and transmitting the received electromagnetic waves to the modem processor for demodulation.
  • the mobile communication module 250 can also amplify the signal modulated by the modem processor, and convert it into electromagnetic wave radiation via the antenna 1.
  • at least part of the functional modules of the mobile communication module 250 may be provided in the processor 210.
  • at least part of the functional modules of the mobile communication module 250 and at least part of the modules of the processor 210 may be provided in the same device.
  • the modem processor may include a modulator and a demodulator.
  • the modulator is used to modulate the low frequency baseband signal to be sent into a medium and high frequency signal.
  • the demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Then the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After the low-frequency baseband signal is processed by the baseband processor, it is passed to the application processor.
  • the application processor outputs a sound signal through an audio device (not limited to a speaker 270A, a receiver 270B, etc.), or displays an image or video through the display screen 294.
  • the modem processor may be an independent device.
  • the modem processor may be independent of the processor 210 and be provided in the same device as the mobile communication module 250 or other functional modules.
  • the wireless communication module 260 can provide applications on the electronic device 200 including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), bluetooth (BT), and global navigation satellites.
  • WLAN wireless local area networks
  • BT wireless fidelity
  • GNSS global navigation satellite system
  • FM frequency modulation
  • NFC near field communication technology
  • infrared technology infrared, IR
  • the wireless communication module 260 may be one or more devices integrating at least one communication processing module.
  • the wireless communication module 260 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 210.
  • the wireless communication module 260 may also receive a signal to be sent from the processor 210, perform frequency modulation, amplify, and convert it into electromagnetic waves to radiate through the antenna 2.
  • the antenna 1 of the electronic device 200 is coupled with the mobile communication module 250, and the antenna 2 is coupled with the wireless communication module 260, so that the electronic device 200 can communicate with the network and other devices through wireless communication technology.
  • the wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), broadband Code division multiple access (wideband code division multiple access, WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC , FM, and/or IR technology, etc.
  • the GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), Beidou navigation satellite system (BDS), quasi-zenith satellite system (quasi -zenith satellite system, QZSS) and/or satellite-based augmentation systems (SBAS).
  • GPS global positioning system
  • GLONASS global navigation satellite system
  • BDS Beidou navigation satellite system
  • QZSS quasi-zenith satellite system
  • SBAS satellite-based augmentation systems
  • the electronic device 200 implements a display function through a GPU, a display screen 294, and an application processor.
  • the GPU is an image processing microprocessor, which is connected to the display screen 294 and the application processor.
  • the GPU is used to perform mathematical and geometric calculations for graphics rendering.
  • the processor 210 may include one or more GPUs that execute program instructions to generate or change display information.
  • the electronic device 200 may complete the image and text fusion through the GPU, and may display the image after the image and text fusion through the display screen 294.
  • the display screen 294 is used to display images, videos, and the like.
  • the display screen 294 includes a display panel.
  • the display panel can adopt liquid crystal display (LCD), organic light-emitting diode (OLED), active matrix organic light-emitting diode or active-matrix organic light-emitting diode (active-matrix organic light-emitting diode).
  • LCD liquid crystal display
  • OLED organic light-emitting diode
  • active-matrix organic light-emitting diode active-matrix organic light-emitting diode
  • AMOLED flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-oLed, quantum dot light-emitting diode (QLED), etc.
  • the electronic device 200 may include one or N display screens 294, and N is a positive integer greater than one.
  • the electronic device 200 may implement a shooting function through an ISP, a camera 293, a video codec, a GPU, a display screen 294, and an application processor.
  • the ISP is used to process the data fed back by the camera 293. For example, when taking a picture, the shutter is opened, the light is transmitted to the photosensitive element of the camera through the lens, the light signal is converted into an electrical signal, and the photosensitive element of the camera transmits the electrical signal to the ISP for processing and is converted into an image visible to the naked eye.
  • ISP can also optimize the image noise, brightness, and skin color. ISP can also optimize the exposure, color temperature and other parameters of the shooting scene.
  • the ISP may be provided in the camera 293.
  • the camera 293 is used to capture still images or videos.
  • the object generates an optical image through the lens and is projected to the photosensitive element.
  • the photosensitive element may be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor.
  • CMOS complementary metal-oxide-semiconductor
  • the photosensitive element converts the optical signal into an electrical signal, and then transfers the electrical signal to the ISP to convert it into a digital image signal.
  • ISP outputs digital image signals to DSP for processing.
  • DSP converts digital image signals into standard RGB, YUV and other formats of image signals.
  • the electronic device 200 may include 1 or N cameras 293, and N is a positive integer greater than 1.
  • Digital signal processors are used to process digital signals. In addition to digital image signals, they can also process other digital signals. For example, when the electronic device 200 selects a frequency point, the digital signal processor is used to perform Fourier transform on the energy of the frequency point.
  • Video codecs are used to compress or decompress digital video.
  • the electronic device 200 may support one or more video codecs. In this way, the electronic device 200 can play or record videos in multiple encoding formats, such as: moving picture experts group (MPEG) 1, MPEG2, MPEG3, MPEG4, and so on.
  • MPEG moving picture experts group
  • MPEG2 MPEG2, MPEG3, MPEG4, and so on.
  • NPU is a neural-network (NN) computing processor.
  • NN neural-network
  • applications such as intelligent cognition of the electronic device 200 can be realized, such as image recognition, face recognition, voice recognition, text understanding, and so on.
  • the external memory interface 220 may be used to connect an external memory card, such as a Micro SD card, so as to expand the storage capacity of the electronic device 200.
  • the external memory card communicates with the processor 210 through the external memory interface 220 to realize the data storage function. For example, save music, video and other files in an external memory card.
  • the internal memory 221 may be used to store computer executable program code, where the executable program code includes instructions.
  • the internal memory 221 may include a storage program area and a storage data area.
  • the storage program area can store an operating system, at least one application program (such as a sound playback function, an image playback function, etc.) required by at least one function.
  • the data storage area can store data (such as audio data, phone book, etc.) created during the use of the electronic device 200.
  • the internal memory 221 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like.
  • the processor 210 executes various functional applications and data processing of the electronic device 200 by running instructions stored in the internal memory 221 and/or instructions stored in a memory provided in the processor.
  • the electronic device 200 can implement audio functions through an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone interface 270D, and an application processor. For example, music playback, recording, etc.
  • the audio module 270 is used for converting digital audio information into an analog audio signal for output, and also for converting an analog audio input into a digital audio signal.
  • the audio module 270 can also be used to encode and decode audio signals.
  • the audio module 270 may be provided in the processor 210, or part of the functional modules of the audio module 270 may be provided in the processor 210.
  • the speaker 270A also called “speaker” is used to convert audio electrical signals into sound signals.
  • the electronic device 200 can listen to music through the speaker 270A, or listen to a hands-free call.
  • the receiver 270B also called “earpiece” is used to convert audio electrical signals into sound signals.
  • the electronic device 200 answers a call or voice message, it can receive the voice by bringing the receiver 270B close to the human ear.
  • Microphone 270C also called “microphone”, “microphone”, is used to convert sound signals into electrical signals.
  • the user can make a sound by approaching the microphone 270C through the human mouth, and input the sound signal into the microphone 270C.
  • the electronic device 200 may be provided with at least one microphone 270C.
  • the electronic device 200 may be provided with two microphones 270C, which can implement noise reduction functions in addition to collecting sound signals.
  • the electronic device 200 may also be provided with three, four or more microphones 270C to collect sound signals, reduce noise, identify sound sources, and realize directional recording functions.
  • the earphone interface 270D is used to connect wired earphones.
  • the earphone interface 270D may be a USB interface 230, or a 3.5mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association (cellular telecommunications industry association of the USA, CTIA) standard interface.
  • OMTP open mobile terminal platform
  • CTIA cellular telecommunications industry association
  • the pressure sensor 280A is used to sense the pressure signal and can convert the pressure signal into an electrical signal.
  • the pressure sensor 280A may be provided on the display screen 294.
  • the capacitive pressure sensor may include at least two parallel plates with conductive material. When a force is applied to the pressure sensor 280A, the capacitance between the electrodes changes.
  • the electronic device 200 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen 294, the electronic device 200 detects the intensity of the touch operation according to the pressure sensor 280A.
  • the electronic device 200 may also calculate the touched position based on the detection signal of the pressure sensor 280A.
  • touch operations that act on the same touch position but have different touch operation strengths may correspond to different operation instructions. For example, when a touch operation whose intensity of the touch operation is less than the first pressure threshold is applied to the short message application icon, an instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, an instruction to create a new short message is executed.
  • the button 290 includes a power-on button, a volume button, and so on.
  • the button 290 may be a mechanical button. It can also be a touch button.
  • the electronic device 200 may receive key input, and generate key signal input related to user settings and function control of the electronic device 200.
  • the motor 291 can generate vibration prompts.
  • the motor 291 can be used for incoming call vibration notification, and can also be used for touch vibration feedback.
  • touch operations that act on different applications can correspond to different vibration feedback effects.
  • Acting on touch operations in different areas of the display screen 294, the motor 291 can also correspond to different vibration feedback effects.
  • Different application scenarios for example: time reminding, receiving information, alarm clock, games, etc.
  • the touch vibration feedback effect can also support customization.
  • the indicator 292 can be an indicator light, which can be used to indicate the charging status, power change, and can also be used to indicate messages, missed calls, notifications, and so on.
  • the SIM card interface 295 is used to connect to the SIM card.
  • the SIM card can be inserted into the SIM card interface 295 or pulled out from the SIM card interface 295 to achieve contact and separation with the electronic device 200.
  • the electronic device 200 may support 1 or N SIM card interfaces, and N is a positive integer greater than 1.
  • the SIM card interface 295 may support Nano SIM cards, Micro SIM cards, SIM cards, etc.
  • the same SIM card interface 295 can insert multiple cards at the same time. The types of the multiple cards can be the same or different.
  • the SIM card interface 295 can also be compatible with different types of SIM cards.
  • the SIM card interface 295 may also be compatible with external memory cards.
  • the electronic device 200 interacts with the network through the SIM card to realize functions such as call and data communication.
  • the electronic device 200 adopts an eSIM, that is, an embedded SIM card.
  • the eSIM card can be embedded in the electronic device 200 and cannot be separated from the electronic device 200.
  • the graphics and text fusion methods in the embodiments of the present application can all be implemented in an electronic device with the foregoing hardware structure or an electronic device with a similar structure.
  • the electronic device may be a mobile phone, a tablet computer, a personal computer, or a netbook.
  • the resulting image-text fusion image (for example, the second image or the third image) can be directly displayed on the display screen of the electronic device, or it can be used as For other purposes, this application does not limit it.
  • the graphic and text fusion method is applied to other types of electronic equipment, for example, server electronic equipment, the graphic and text fusion image obtained by the server electronic equipment can be pushed to the display screen of the terminal electronic equipment, or it can be For other purposes.
  • Example 1 Graphic and text fusion images are used as wallpapers of electronic devices. For example, lock screen wallpaper, main interface wallpaper, or chat interface background.
  • FIG. 3 it is an example of a picture-text fusion image as the lock screen wallpaper of a mobile phone.
  • (b) in Figure 3 it is an example of a picture-text fusion image as the wallpaper on the main interface of a mobile phone.
  • (c) in Figure 3 it is an example of a picture-text fusion image as the background of the WeChat chat interface of a mobile phone.
  • Example 2 The graphic and text fusion image is used as the application startup page or guide page. For example, the interface when the application starts.
  • the startup page can also be called a splash screen page.
  • the design of the startup page can effectively use the blank interface in the application initialization process, enhance the user's perception of the application can be quickly launched and put into use immediately, and thereby enhance the user experience when the application is launched.
  • brand display, advertisements, events, etc. can be displayed on the startup page, and the display mode can be static pictures, dynamic pictures, animations, and other methods.
  • the initialization time of the application generally does not exceed 5 seconds
  • the control duration of the startup page usually does not exceed 5 seconds.
  • (a) in Figure 4 it is an example of a fusion image with text and graphics as an APP startup page, and the startup page control duration of the APP is 3 seconds (seconds, s).
  • the guide page is a page used to guide users to learn application usage or understand the role of the application, and its core lies in the word "guidance”.
  • Guide pages generally appear on applications of new concepts or after product iterations.
  • Figure 4(b) it is an example of applying a fusion image with text and graphics to the guide page of an APP.
  • the guide page of the APP is composed of 3 guide pictures (as shown in (c) in Figure 4).
  • the mobile phone can switch the guide pictures in response to the user's left/right sliding operation on the touch screen.
  • Example 3 Graphic and text fusion images are used in the application interface. For example, the floating ad frame of the video playback window.
  • the mobile phone when the mobile phone detects that the video playback pause button is clicked, the mobile phone can display a floating advertisement frame stop interface in the video playback window, as shown in Figure 5.
  • Example 4 Graphic and text fusion images are displayed on traditional image communication media. For example, it can be displayed in newspapers, magazines, TV or outdoor advertising spaces.
  • the following takes the application of the graphic-text fusion method of the embodiment of the present application to a mobile phone having the hardware structure shown in FIG. 2 as an example to specifically describe the graphic-text fusion method provided by the embodiment of the present application.
  • the mobile phone can perform some or all of the steps in the embodiments of the present application, and these steps or operations are only examples, and the embodiments of the present application may also perform other operations or variations of various operations.
  • each step may be executed in a different order presented in the embodiments of the present application, and it may not be necessary to perform all the operations in the embodiments of the present application.
  • the image-text fusion method in the embodiment of the present application may include S601-S605:
  • the mobile phone acquires the first image and the first text to be typeset into the first image.
  • the first image is an image to which text is to be added.
  • the first text is the text to be typeset into the first image.
  • the first text may correspond to the image content of the first image.
  • the first text corresponds to the image content of the first image, which means that the first text can be used to explain and describe the image information in the first image; or the first text and the image in the first image
  • the main body of the information description is the same; or the first text is consistent with the artistic conception conveyed by the image information in the first image.
  • the first image is an image that includes the features of the moon image and the image of the Eiffel Tower.
  • the first text can be a text with the subject "Jiaojiaobai Jade Plate", and the content of the text includes: "Half covered by the Eiffel Tower The full moon is like a shy girl who is still holding a pipa and half-hidden. @ ⁇ ".
  • the text with the theme "Jiaojiao Baiyu Pan" in FIG. 6 is an explanation and description of the image information in the first image in FIG. 6.
  • the first image and the first text may be obtained by the mobile phone from a third party.
  • the mobile phone periodically obtains the first image and the first text from the server of the mobile phone manufacturer.
  • the first image may be a picture taken locally by the mobile phone, and the first text may be a user-defined text received by the mobile phone.
  • the first image is the determined image when the mobile phone receives the user's picture selection operation; the first text is the text input by the user received by the mobile phone.
  • the embodiment of the present application does not limit the specific sources of the first image and the first text number.
  • the first image may only correspond to a set of text data.
  • the image including building features in Figure 1 only corresponds to the text with the theme "Glimpse of the Country".
  • the first text obtained by the mobile phone is the text whose subject is "Glimpse of the Country”.
  • the first image may correspond to multiple sets of text data.
  • the first images corresponding to (a) in FIG. 7 and (b) in FIG. 7 are the same, but the corresponding text data is different.
  • the mobile phone when the mobile phone obtains the first text, it is specifically which set of the multiple sets of text data is obtained. It can be determined according to the ranking of the degree of matching between the first image and each set of text data, it can also be determined randomly, or it can be determined in a certain order. In this regard, the embodiment of the present application does not limit it.
  • the mobile phone when making a guide page as shown in (c) in Figure 4, can first obtain the first image and the text data with the subject "Know you", and then use the image-text fusion method of this application embodiment to merge the subject Type the text data of "Know You" into the first image, and get the first guide page. Then obtain the text data with the themes "fashion” and “trust” in turn, and use the graphic fusion method of the embodiment of this application to obtain the second guide page, the third guide page, the fourth guide page, and the fifth guide in turn. Pages etc.
  • the mobile phone determines the feature value of each pixel in the first image.
  • the feature value of a pixel is used to characterize the likelihood of the pixel being paid attention to by the user.
  • determining the feature value of each pixel in the first image by the mobile phone may include: the mobile phone performs feature detection on the first image to determine the visually significant parameters and facial feature parameters of each pixel in the first image, At least two of the edge feature parameters and the text feature parameters; then, the mobile phone determines at least two of the visually significant parameters of each pixel in the first image, the face feature parameters, the edge feature parameters, and the text feature parameters The parameters are weighted and summed to determine the characteristic value of each pixel in the first image.
  • feature detection is used to identify image features in an image. For example, recognize facial features, human features, building features, object features (such as animal features, tree features, flower features, etc.), text features, river features, and mountain features in the image.
  • object features such as animal features, tree features, flower features, etc.
  • text features such as river features, and mountain features in the image.
  • the mobile phone may obtain at least two feature maps of the first feature map, the second feature map, the third feature map, and the fourth feature map of the first image after performing feature detection on the first image.
  • the pixel value of each pixel in the first feature map is the visually significant parameter of the corresponding pixel
  • the pixel value of each pixel in the second feature map is the face feature parameter of the corresponding pixel
  • each pixel in the third feature map The pixel value of the point is the edge feature parameter of the corresponding pixel
  • the pixel value of each pixel in the fourth feature map is the text feature parameter of the corresponding pixel.
  • the mobile phone can perform a weighted summation on at least two feature maps of the first feature map, the second feature map, the third feature map, and the fourth feature map to obtain the feature map of the first image.
  • the pixel value of each pixel represents the probability of the pixel being paid attention to by the user.
  • the mobile phone performs a weighted summation of the first feature map, the second feature map, the third feature map, and the fourth feature map to obtain the feature map of the first image, which may specifically include: the mobile phone versus the first feature map, The pixel values of each pixel in the second, third, and fourth feature maps are weighted and summed to determine the feature value of each pixel in the first image; then, according to the first The feature value of each pixel in the image obtains the feature map of the first image.
  • the mobile phone may use a visual saliency detection algorithm to perform visual saliency feature detection on the first image, and determine the visual saliency parameters of each pixel in the first image.
  • the principle of the visual saliency detection algorithm is to determine the visual saliency feature by calculating the difference between pixels in the image.
  • the mobile phone can obtain the image features that the human visual system pays attention to through the visual saliency detection algorithm.
  • Figure 8 (a) it is an example of a first image.
  • the mobile phone uses the visual saliency area detection algorithm to perform visual saliency feature detection on the first image to obtain the first feature map (as shown in (b) in Figure 8).
  • FIG. 8 it is an example of a first image.
  • the mobile phone uses the visual saliency area detection algorithm to perform visual saliency feature detection on the first image to obtain the first feature map (as shown in (b) in Figure 8).
  • the mobile phone uses the visual saliency region detection algorithm to perform visual saliency feature detection on the first image to obtain the first feature map (as shown in (b) in FIG. 9).
  • the brightness of each pixel is used to identify the feature value of the image feature of the pixel, and the greater the brightness, the feature value of the image feature of the pixel. Bigger.
  • the visual saliency detection algorithm may specifically be a visual saliency detection algorithm based on multi-feature absorbing Markov chain, a visual saliency detection algorithm based on global color contrast, and the corner convex hull and shell A visual saliency detection algorithm combined with Yeess inference, or a visual saliency detection algorithm based on deep learning (for example, based on a convolutional neural network), etc., which are not limited in the embodiment of the present application.
  • FIG. 10 a flow chart of a visual saliency detection algorithm based on a multi-feature absorbing Markov chain provided by this embodiment of the application.
  • the algorithm is mainly divided into two steps: the first step is to extract the superpixels and features of the image, and the second step is to establish a Markov chain based on the extracted superpixels and features.
  • a simple linear iterative clustering (SLIC) algorithm can be used to divide the first image into a number of super pixels.
  • the color of all pixels in the superpixel is fitted into a CIELab three-dimensional normal distribution, and the directional value of each pixel in 4 directions (0°, 45°, 90° and 135°) is fitted to get Four-dimensional normal distribution characteristics.
  • the visual saliency detection algorithm based on multi-feature absorption Markov chain can use super pixels as the basic processing unit. Establish connections for all super pixels in the image, and then use the final stationary distribution as the first feature map.
  • the relatively bright pixels in the first feature map correspond to the salient features of the first image, and the brighter the pixel, the brighter the pixel is seen by the user. The higher the likelihood of attention.
  • the mobile phone may use a face detection algorithm to perform face feature detection on the first image, and determine the face feature parameters of each pixel in the first image.
  • the face detection algorithm is used to detect the face features in the image.
  • the mobile phone uses a face detection algorithm to perform face detection on the first image shown in (a) in FIG. 9 to obtain a second feature map (as shown in (c) in FIG. 9).
  • the pixel point distribution area with greater brightness corresponds to the face area in the first image.
  • the mobile phone uses a face detection algorithm to perform face detection on the first image shown in Figure 8 (a), and finds that there is no face in the image. Therefore, the second feature map is obtained (As shown in (c) in Figure 8), there are no facial features.
  • the face detection algorithm may specifically be a computer vision face detection algorithm (for example, a face detection algorithm based on Histogram of Oriented Gradients (HOG)), or may be a face detection algorithm based on deep learning (For example, a face detection algorithm based on a convolutional neural network), etc., which are not limited in the embodiment of the present application.
  • a computer vision face detection algorithm for example, a face detection algorithm based on Histogram of Oriented Gradients (HOG)
  • HOG Histogram of Oriented Gradients
  • deep learning for example, a face detection algorithm based on a convolutional neural network
  • FIG. 11 it is a flowchart of a HOG-based face detection algorithm provided by an embodiment of this application.
  • image normalization refers to a process of performing a series of standard processing transformations on an image to transform the image into a fixed standard form, and the standard image is called a normalized image.
  • the normalized image has invariant characteristics to affine transformations such as translation, rotation, and scaling. Then the first-order differential is used to calculate the image gradient of the normalized image.
  • the normalized image is divided into HOG blocks, a gradient histogram is drawn for each cell in the divided HOG structure, and then the gradient histogram of each cell is projected with a prescribed weight. Then, the feature vector of each cell block is normalized, so that the feature vector space of each cell block has invariant characteristics to illumination, shadow and edge changes. Finally, the histogram vectors of all HOG blocks are combined into one HOG feature vector to obtain the HOG feature vector of the first image, which is the second feature map.
  • the mobile phone may use an edge detection algorithm to perform edge feature detection on the first image, and determine the edge feature parameters of each pixel in the first image.
  • the edge detection algorithm can be used to extract the edge gradient features of regions with complex textures (such as grass, woods, mountains, and water waves) in the image.
  • the mobile phone uses an edge detection algorithm to perform edge detection on the first image shown in (a) in FIG. 8 to obtain a third feature map (as shown in (d) in FIG. 8).
  • the third feature map is used to identify the structure texture of the Eiffel Tower in the first image.
  • the mobile phone uses an edge detection algorithm to perform edge detection on the first image shown in Figure 9 (a) to obtain a third feature map (as shown in Figure 9 (d)) .
  • the third feature map is used to identify the contour of the human body in the first image.
  • the edge detection algorithm may specifically be a Laplacian operator, a Canny algorithm, a Prewitt operator, or a Sobel operator, etc., which is not limited in the embodiment of the present application.
  • the Canny algorithm is a method of smoothing first and then derivation.
  • the Canny algorithm can be used to smooth the first image using a Gaussian filter.
  • smoothing processing refers to image noise reduction (for example, suppressing image noise, suppressing high frequency interference, etc.), so that the brightness of the image gradually changes, reduces the sudden gradient, and improves the image quality.
  • the finite difference of the first-order partial derivative can be used to calculate the gradient amplitude and direction of the edge in the noise-free first image.
  • non-maximum suppression is performed on the gradient amplitude of the edge in the noise-free first image.
  • the double threshold algorithm is used to detect and connect the edges to obtain the third feature map.
  • the mobile phone may use a text detection algorithm to perform text feature detection on the first image, and determine the text feature parameters of each pixel in the first image.
  • the text detection algorithm can be used to extract text or character features existing in images (such as calendars, solar terms-related images, advertisements, posters, etc.).
  • the mobile phone uses a text detection algorithm to perform text detection on the first image shown in (a) in FIG. 8 to obtain a fourth feature map (as shown in (e) in FIG. 8).
  • the mobile phone uses a text detection algorithm to perform text detection on the first image shown in (a) in FIG. 9 to obtain a fourth feature map (as shown in (e) in FIG. 9). Since the first image examples shown in (a) in FIG. 8 and (a) in FIG. 9 do not contain text, there are none in (e) in FIG. 8 and (e) in FIG. 9 obtained. Text characteristics.
  • the text detection algorithm may specifically be a text detection algorithm based on computer gradient and expansion and corrosion operations, or may be a text detection algorithm based on deep learning (for example, based on convolutional neural network), etc.
  • the implementation of this application The example does not limit this.
  • FIG. 13 it is a flowchart of a text detection algorithm based on computer gradient and expansion and corrosion operations provided by this embodiment of the present application.
  • the text detection algorithm based on computer gradient and expansion and erosion operations can firstly grayscale the first image to reduce the amount of calculation. Then calculate the gradient feature of the gray image, and binarize the gradient feature, and then use the image expansion and erosion operation to process the binarized gradient feature. If the gradient feature of the image area meets the given threshold, the area can be considered It is the text feature area.
  • the text feature map can be obtained by using the above method.
  • the visual saliency detection algorithm based on the multi-feature absorbing Markov chain as shown in FIG. 10 in the embodiment of the present application, as shown in FIG. 11, is based on the HOG face detection algorithm, as shown in FIG.
  • the edge detection algorithm based on the Canny algorithm and the text detection algorithm based on computer gradient and expansion and erosion operations shown in Figure 13 are just examples of calculations.
  • the embodiments of the present application do not limit specific visual saliency detection algorithms, face detection algorithms, edge detection algorithms, and text detection algorithms.
  • the embodiment of the present application is based on the multi-feature absorption Markov chain visual saliency detection algorithm shown in FIG. 10, such as the HOG-based face detection algorithm shown in FIG. 11, and the Canny-based face detection algorithm shown in FIG.
  • the specific processing details of the edge detection algorithm of the algorithm and the text detection algorithm based on computer gradient and expansion and corrosion operations as shown in Figure 13 are also not limited.
  • the above algorithm and process can also have other variations.
  • For details and processing procedures in the conventional technology reference may be made to the details and processing procedures in the conventional technology, and the examples of this application will not be repeated here.
  • the mobile phone compares each pixel in the first image
  • the weighted summation of the visually significant parameters, face feature parameters, edge feature parameters, and text feature parameters to determine the feature value of each pixel in the first image can be achieved by the following formula:
  • F c (x,y) ⁇ *F sal (x,y)+ ⁇ *F face (x,y)+ ⁇ *F edge (x,y)+ ⁇ *F text (x,y).
  • F c (x, y) is the feature value of the pixel (x, y) in the first image
  • F sal (x, y) is the visually significant parameter of the pixel (x, y) in the first image
  • F face (x, y) is the face feature parameter of the pixel (x, y) in the first image
  • F edge (x, y) is the edge feature parameter of the pixel (x, y) in the first image
  • F text (x, y) is the text feature parameter of the pixel (x, y) in the first image.
  • ⁇ , ⁇ , ⁇ , and ⁇ are visually significant parameters, face feature parameters, edge feature parameters and weight parameters corresponding to text feature parameters, respectively.
  • the specific values of ⁇ , ⁇ , ⁇ , and ⁇ may be determined according to the importance of different parameters.
  • the ranking of importance of visually significant parameters, facial feature parameters, edge feature parameters, and text feature parameters may be: facial feature parameters>text feature parameters>visually significant parameters>edge feature parameters.
  • ⁇ , ⁇ , ⁇ , and ⁇ may be set to 0.2, 0.4, 0.1, and 0.3, respectively.
  • ⁇ , ⁇ , ⁇ , and ⁇ may be set to 0.2, 0.4, 0, and 0.4, respectively.
  • ⁇ is 0, it can also be understood that when determining the feature value of each pixel in the first image, the edge feature parameter is not considered.
  • the F c (x, y) of each pixel corresponds to the feature map of the first image ((f) in FIG. 8 and (f) in FIG. 9).
  • the brightness of the pixel (x, y) represents the F c (x, y) value of the pixel. The greater the brightness, the greater the F c (x, y) value of the pixel (x, y).
  • the mobile phone determines a plurality of first typesetting layouts of the first text in the first image according to the first text and the feature value of each pixel in the first image.
  • the first typesetting format is the candidate typesetting format described above.
  • the first typesetting format is used to at least characterize the text template of the first text and the position of the first text typeset to the first image.
  • the text template can at least specify one or more of the following: the line spacing, line width, font size, font, text thickness, alignment, typesetting form, and the position of the decorative line and the thickness of the decorative line of the text subject and the text body Wait.
  • the typesetting form can include at least vertical and horizontal.
  • the text template can be saved in JavaScript Object Notation (JSON) format, Extensible Markup Language (XML) format, code fragments, or other text formats, which are not limited in this embodiment of the application. .
  • JSON JavaScript Object Notation
  • XML Extensible Markup Language
  • FIG. 14 an example diagram of a typesetting specification in JSON format provided by this embodiment of the application.
  • the multiple text templates in the embodiments of the present application may be pre-designed, stored in the mobile phone, or stored in the server of the mobile phone manufacturer.
  • the text template may also be stored in other corresponding locations, which is not limited in the embodiment of the present application.
  • the position of the first text typeset to the first image is used to indicate that when the first text is typeset to the first image according to the first typesetting format, the first text is located at the relative position of the first image.
  • the position of the first text layout to the first image may include at least any one of: upper left, upper right, centered, lower left, lower right, centered at the top, and centered at the bottom.
  • FIG. 15 a schematic diagram of the position of the first text typeset to the first image provided in this embodiment of the application. As shown in Figure 15, the position of the first text typeset to the first image is "bottom left".
  • the position where the first text is typeset to the first image may be the default typesetting position.
  • the default typesetting position is the preset typesetting position, and the default typesetting position can be any of the top left, top right, center, bottom left, bottom right, top center, and bottom center position.
  • the mobile phone when the mobile phone can typeset the first text with a certain text template, the size of the text box of the first text is combined with the feature value of each pixel in the first image (or the feature map of the first image). ), determining the position where the first text is typeset to the first image, so that the first text does not block pixels in the first image whose feature value is greater than the first threshold.
  • the image feature corresponding to the pixel point in the first image whose feature value is greater than the first threshold is the salient feature mentioned above, that is, the image feature that is more likely to be paid attention to by the user. Therefore, the first text determined by the mobile phone is typeset to the position of the first image, so that the first text does not obscure the salient features in the first image.
  • the mobile phone uses a certain text template to typeset, there may be multiple typesetting positions that meet the above conditions.
  • the size of the text box of the first text is different.
  • FIG. 16 there are shown example diagrams in which several types of first texts provided by the embodiments of the present application are typeset with different text templates. Therefore, the position where the first text is typeset with different text templates to the first image may also be different. Therefore, the mobile phone can determine multiple layout positions of the first text to the first image.
  • the mobile phone has determined multiple second typesetting layouts.
  • candidate layouts ie, multiple first layouts determined by the mobile phone are introduced.
  • the mobile phone can determine 3 candidate layouts based on the first text to be merged: Candidate layout 1 (( in Figure 17) a)), candidate layout 2 (as shown in (b) in Figure 17) and candidate layout 3 (as shown in (c) in Figure 17).
  • the mobile phone needs to further determine the optimal typesetting format of the first text from the candidate typesetting format 1, the candidate typesetting format 2 and the candidate typesetting format 3, that is, the phone executes S604.
  • the mobile phone determines a second typesetting format from the plurality of first typesetting formats according to the cost parameters of the plurality of first typesetting formats.
  • a cost parameter of the first typesetting format is used to characterize the size of the feature value of the pixel shaded by the first text when the first text is typeset in the first image according to the typesetting format, and the first image typeset with the first text The degree of balance of the feature value distribution of the pixel points in each area in the.
  • the mobile phone can have multiple first typesetting formats, so that when the first text is typeset to the first image in a certain first typesetting format, the feature value of the pixels blocked by the first text is the smallest, and the typesetting
  • the first typesetting format in which the feature value distribution of the pixel points in each area in the first image with the first text is the most balanced is the second typesetting format.
  • the mobile phone typesets the first text into the first image according to the second typesetting format to obtain the second image.
  • the mobile phone determines the second typesetting format from the plurality of first typesetting formats according to the cost parameters of the plurality of first typesetting formats (i.e., S604), which may include S1801-S1806:
  • the text box of the first text blocks the image area of the first image.
  • the typesetting layout i is any first typesetting layout determined by the mobile phone.
  • the mobile phone calculates the texture feature parameters of the above-mentioned image area.
  • the texture feature parameter of the image area is used to characterize the number of texture features in the image corresponding to the image area.
  • the mobile phone may first obtain the grayscale image corresponding to the image area. Then calculate the texture feature parameters of the gray image.
  • the gray-scale image refers to the image obtained through gray-scale processing. This processing is used to reduce the amount of calculation.
  • the grayscaled image can present a distribution of white ⁇ gray ⁇ black. Pixels with a gray value of 0 are displayed as white, and pixels with a gray value of 255 are displayed as black.
  • the mobile phone may calculate the texture feature parameter of the image area based on the angular second moment, contrast, entropy, etc. of the gray level co-occurrence matrix, or the variance of the image gray level and the gradient corresponding to the image area.
  • the mobile phone judges whether the texture feature parameter of the image corresponding to the image area occluded by the first text is greater than a second threshold. If the texture feature parameter of the image corresponding to the image area occluded by the first text is less than the second threshold, the mobile phone executes S1804. If the texture feature parameter of the image corresponding to the image area occluded by the first text is greater than the second threshold, the mobile phone discards the layout format i, sets i+1, and re-executes S1801 until S1801 to S1803 are executed for each first layout format.
  • the second threshold is a preset threshold.
  • the second threshold is 125. It is understandable that for the first typesetting format where the texture feature parameter of the image corresponding to the image area occluded by the first text is greater than 125, it can be considered that the texture feature of the image area is too complicated. If the first text is typeset here, it may be At the same time affect the texture characteristics and the display of the first text. In this case, the mobile phone can abandon using the first typesetting format to typeset the first text.
  • the mobile phone calculates that when the first text is typeset into the first image according to the typesetting format i, at least two of the text intrusion parameter, the visual space occupancy parameter and the visual balance parameter of the first text are calculated.
  • the text intrusion parameter of the first text is the ratio of the first parameter to the second parameter.
  • the first parameter is the sum of the feature values of each pixel in the image area where the first text occludes the first image.
  • the second parameter is the area of the image area, or the second parameter is the total number of pixels in the image area, or the second parameter is the product of the total number of pixels in the image area and a preset value.
  • the visual space occupancy parameter of the first text is used to characterize the proportion of pixels in the image area whose feature value is less than the third threshold.
  • the visual space occupancy parameter Eu (L i ) of the first text is calculated.
  • Im (xy) is the pixel value of the pixel in the image area of the first image after the first text is typeset into the first image with the typesetting format i
  • t is the maximum feature value in the first image.
  • the visual balance parameter of the first text is used to characterize the degree of influence of the first text on the balance of the feature value distribution of the pixel points in each area of the first image typeset with the first text.
  • E n (L i) b1 + b2 + b3 is calculated in accordance with a first layout format when text i layout to the first image, the first text visual balance parameter E n (L i).
  • b1 is the distance between the image center of gravity of the feature map of the image region of the first image and the image center of gravity of the feature map of the first image when the first text is typeset into the first image in the typesetting format i.
  • b2 is the distance between the image center of the image area of the first image and the center of the first image when the first text is typeset into the first image in the typesetting format i.
  • b3 is the distance between the image center of the image area where the first text blocks the first image and the image center of gravity of the feature map of the image area where the first text blocks the first image when the first text is typeset into the first image in the typesetting format i .
  • the image center of gravity (X c , Y c ) can be based on the formula And formula Calculated.
  • P1 is the feature map of the image area of the first image that the first text blocks when the first text is typeset into the first image with the typesetting layout i
  • P2 is the image center of gravity of the feature map of the first image.
  • P3 is that when the first text is typeset to the first image in the typesetting format i, the first text blocks the image center of the image area of the first image.
  • P4 is the image center of the first image. It can be obtained that for the candidate typesetting format shown in (c) in FIG.
  • b1 is the distance between P1 and P2
  • b2 is the distance between P1 and P4
  • b3 is the distance between P1 and P3.
  • the same method can be calculated for each candidate layout format corresponding to b1, b2 and b3, each candidate layout to obtain further visual layout balance parameter E n (L i).
  • the mobile phone calculates multiple cost parameters of the first typesetting format according to at least two of the calculated text intrusion parameters, visual space occupancy parameters, and visual balance parameters of the first text.
  • the mobile phone may use any one of the following formula 1, formula 2, or formula 3 to calculate the cost parameter of the typesetting format i.
  • T i ⁇ 1 *E s (L i )+ ⁇ 2 *E u (L i )+ ⁇ 3 *E n (L i ).
  • T i ( ⁇ 1 *E s (L i )+ ⁇ 2 *E u (L i ))*E n (L i ).
  • T i E s (L i )*E u (L i )*E n (L i ).
  • ⁇ 1, ⁇ 2 and ⁇ 3 are the E s (L i), the E u (L i) and the corresponding weight E n (L i) weight parameters.
  • the value range of ⁇ 1 , ⁇ 2 and ⁇ 3 can be any value from 0 to 1.
  • a certain weight parameter is 0, it can also be understood that when determining the cost parameter of the first typesetting format, the parameter corresponding to the weight parameter is not considered. For example, if ⁇ 1 is 0, it can be understood that when determining the cost parameter of the first typesetting format, the text intrusion parameter of the first text is not considered.
  • the mobile phone uses the same cost parameter calculation formula (such as the above formula 1, formula 2 or formula 3) to calculate its cost parameter for each first typesetting format.
  • the mobile phone can use the above formula 1 to calculate the cost parameter of each first typesetting format.
  • the mobile phone determines multiple cost parameters of the first typesetting format, and the first typesetting format corresponding to the smallest cost parameter is the second typesetting format.
  • the mobile phone before S1801, the mobile phone can also perform:
  • the mobile phone judges whether the typesetting layout i is the default typesetting layout (for example, the bottom is centered).
  • the mobile phone can directly execute S1804. If the typesetting layout i is not the default typesetting layout, the mobile phone can continue to execute S1801. As shown in Figure 19.
  • the image-text fusion method of the embodiment of the present application may further include: the mobile phone determines the color parameter of the first text.
  • the mobile phone may determine the color parameter of the first text before S605. In this case, after S605, the mobile phone may color the first text in the second image according to the color parameters of the first text to obtain the third image.
  • S605 may also be: the mobile phone typesets the first text into all the first images according to the second typesetting format and the color parameters of the first text to obtain the second image.
  • the mobile phone may determine the color parameter of the first text after S605. In this case, after the mobile phone determines the color parameters of the first text, the mobile phone can color the first text in the second image according to the color parameters of the first text to obtain the third image.
  • the color parameter of the first text is used to color each text in the first text.
  • the mobile phone can first determine the hue, saturation, and lightness of the three primary colors RGB of the image area where the first image is blocked by the first text when the first text is typeset into the first image according to the second typesetting format. The dominant color of the image area. Then, the HSV derived color of the main color is used as the color parameter of the first text.
  • the RGB color mode is a color standard in the industry, which is obtained by changing the three color channels of red (R), green (G), and blue (B) and superimposing them with each other.
  • s color is a color space created based on the intuitive characteristics of colors, also called Hexcone Model.
  • the color parameters in the hexagonal pyramid model may at least include hue (H), saturation (S) and lightness (V).
  • the main color of the image area is the hue with the highest hue proportion in the image corresponding to the image area.
  • the dominant color of the image area can be determined by the statistical hue histogram.
  • hue histogram statistics techniques reference may be made to conventional statistics techniques, which will not be described in detail in the embodiment of the present application.
  • the derived color of the main color refers to a color whose hue is the same as the main color, but whose hue, saturation, and lightness are different from the HSV of the main color.
  • the first text is blurred because the texture of the image area occluded by the first text is too complicated, or the color tone is too complicated, or Problems that are not prominent enough.
  • the image-text fusion method of the embodiment of the present application may further include:
  • S2001 The mobile phone determines whether it is necessary to perform rendering processing on the second image.
  • the mobile phone determines whether the third image needs to be rendered.
  • the rendering processing may include at least one of mask rendering and projection rendering.
  • the mobile phone can determine whether the second image or the third image needs to be rendered according to whether at least one of the following conditions is met:
  • Condition 1 After the first text is typeset into the first image according to the second typesetting format, the texture feature parameter of the image area where the first text blocks the first image is greater than the fourth threshold.
  • the texture feature parameter is used to characterize the number of texture features in the image corresponding to the image region.
  • texture feature parameters you can refer to the edge feature detection algorithm mentioned above, or other conventional edge feature detection algorithms or texture feature detection algorithms, which will not be repeated here.
  • Condition 2 After the first text is typeset into the first image according to the second typesetting format, the main color ratio of the image area where the first text blocks the first image is less than the fifth threshold.
  • the dominant color ratio of the first image can be determined by counting the hue histogram of the first image.
  • hue histogram statistics techniques reference may be made to conventional statistics techniques, which will not be described in detail in the embodiment of the present application.
  • the mobile phone can render the second image to highlight the first text in the second image.
  • the mobile phone executes S2002.
  • S2002 The mobile phone performs mask rendering or projection rendering on the second image.
  • the mask rendering of the second image by the mobile phone specifically refers to the mask rendering of the text box area of the first text in the second image by the mobile phone.
  • the projection and rendering of the second image by the mobile phone specifically refers to that the mobile phone adds a text shadow to the text of the first text in the second image.
  • mask rendering of the second image by the mobile phone may include: the mobile phone overlays a mask layer on the second image. Specifically, the mobile phone may cover the mask layer on the text frame area of the first text in the second image.
  • the method for generating the mask layer may include: on the basis of the size of the text box of the first text in the second image, expanding the size H1 in at least one of up, down, left and right respectively, and determining The size of the mask layer. Then, the transparency process is performed in the order of transparency threshold 1 ⁇ transparency threshold 2 ⁇ transparency threshold 3 to obtain a mask layer.
  • the transparency can be processed in a gradual direction from top to bottom, from left to right, from bottom to top, or from right to left.
  • the specific gradient direction can be determined by the specific typesetting format. For example, if the second typesetting format is top-centered, the transparency can be processed in a gradual direction from top to bottom. The embodiment of the application does not limit this.
  • performing mask rendering on the second image by the mobile phone may include: the mobile phone determines the mask parameters, and then the mobile phone processes the second image according to the determined mask parameters.
  • the mask parameters may at least include mask size and mask transparency parameters.
  • the mask size can be determined according to the following method: on the basis of the size of the text box of the first text in the second image, at least one of up, down, left, and right is expanded by a size H1, which is the mask size.
  • the mask transparency parameter can be determined according to the following method: the mask transparency parameter is determined in the order of transparency threshold value 1 ⁇ transparency threshold value 2 ⁇ transparency threshold value 3. Among them, the mask transparency parameter can be determined in a gradual direction from top to bottom, from left to right, from bottom to top, or from right to left.
  • the specific gradient direction can be determined by the specific typesetting format. For example, if the second typesetting format is top-centered, the mask transparency parameter can be determined in a gradual direction from top to bottom. The embodiment of the application does not limit this.
  • FIG. 21 there are several image-text fusion image comparison diagrams provided in this embodiment of the application.
  • (b1) in FIG. 21 adopts the typesetting method of the embodiment of the present application.
  • the text box setting position is more scientific and the text is more prominent. Stronger sex.
  • (b2) in FIG. 21 adopts the typesetting method of the embodiment of the present application.
  • the conflict between the image and text colors is smaller. The prominence of the text is stronger.
  • performing projection rendering on the second image by the mobile phone may include: the mobile phone determines the text projection parameters, and then the mobile phone processes the second image according to the determined text projection parameters.
  • the text projection parameters may include at least projection color, projection displacement, and projection blur value.
  • the projection color can be consistent with the text color.
  • the projection displacement can be a preset displacement parameter.
  • the projection blur value may be a preset blur value.
  • the projection blur value may gradually change with different displacement positions.
  • the electronic device includes a hardware structure and/or software module corresponding to each function.
  • the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed by hardware or computer software-driven hardware depends on the specific application and design constraint conditions of the technical solution. Professionals and technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered beyond the scope of this application.
  • the embodiments of the present application may divide the electronic equipment into functional modules.
  • each functional module may be divided corresponding to each function, or two or more functions may be integrated into one processing module.
  • the above-mentioned integrated modules can be implemented in the form of hardware or software function modules. It should be noted that the division of modules in the embodiments of the present application is illustrative, and is only a logical function division, and there may be other division methods in actual implementation.
  • FIG. 22 it is a schematic structural diagram of an electronic device provided in an embodiment of this application.
  • the electronic device may include an information acquisition unit 2210, an analysis unit 2220, and a processing unit 2230.
  • the information obtaining unit 2210 can be used to support the electronic device to perform the above step S601, and to obtain multiple text templates, and/or other processes used in the technology described herein;
  • the analysis unit 2220 can be used to support the electronic device to perform the above steps S602, S603, S604, and S2001; or collecting the first data, and/or other processes used in the technology described herein;
  • the processing unit 2230 is used to support the electronic device to perform the above steps S605 and S2002, and/or used in this article Describe the other processes of the technology.
  • the above electronic device may also include a radio frequency circuit.
  • the electronic device can receive and send wireless signals through a radio frequency circuit.
  • the radio frequency circuit includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, and the like.
  • the radio frequency circuit can also communicate with other devices through wireless communication.
  • the wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communications, General Packet Radio Service, Code Division Multiple Access, Wideband Code Division Multiple Access, Long Term Evolution, Email, Short Message Service, etc.
  • the computer program product includes one or more computer instructions.
  • the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
  • the computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium.
  • the computer instructions may be transmitted from a website, computer, server, or data center.
  • the computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center integrated with one or more available media.
  • the usable medium may be a magnetic medium, (for example, a floppy disk, a hard disk, and a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)).
  • the steps of the method or algorithm described in the embodiments of the present application may be implemented in a hardware manner, or may be implemented in a manner in which a processor executes software instructions.
  • Software instructions can be composed of corresponding software modules, which can be stored in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, mobile hard disk, CD-ROM or any other form of storage known in the art Medium.
  • An exemplary storage medium is coupled to the processor, so that the processor can read information from the storage medium and can write information to the storage medium.
  • the storage medium may also be an integral part of the processor.
  • the processor and the storage medium may be located in the ASIC.
  • the ASIC may be located in the detection device.
  • the processor and the storage medium may also exist as discrete components in the detection device.
  • the disclosed user equipment and method may be implemented in other ways.
  • the device embodiments described above are only illustrative.
  • the division of the modules or units is only a logical function division.
  • there may be other division methods for example, multiple units or components may be It can be combined or integrated into another device, or some features can be omitted or not implemented.
  • the displayed or discussed mutual coupling or direct coupling or communication connection may be indirect coupling or communication connection through some interfaces, devices or units, and may be in electrical, mechanical or other forms.
  • the units described as separate parts may or may not be physically separate.
  • the parts displayed as units may be one physical unit or multiple physical units, that is, they may be located in one place, or they may be distributed to multiple different places. . Some or all of the units may be selected according to actual needs to achieve the objectives of the solutions of the embodiments.
  • the functional units in the various embodiments of the present application may be integrated into one processing unit, or each unit may exist alone physically, or two or more units may be integrated into one unit.
  • the above-mentioned integrated unit can be implemented in the form of hardware or software functional unit.
  • the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium.
  • the technical solutions of the embodiments of the present application are essentially or the part that contributes to the prior art, or all or part of the technical solutions can be embodied in the form of a software product, and the software product is stored in a storage medium. It includes several instructions to make a device (may be a single-chip microcomputer, a chip, etc.) or a processor (processor) execute all or part of the steps of the methods described in the various embodiments of the present application.
  • the aforementioned storage media include: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), magnetic disk or optical disk and other media that can store program code .

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Multimedia (AREA)
  • Health & Medical Sciences (AREA)
  • Oral & Maxillofacial Surgery (AREA)
  • General Health & Medical Sciences (AREA)
  • Human Computer Interaction (AREA)
  • Artificial Intelligence (AREA)
  • Computing Systems (AREA)
  • Databases & Information Systems (AREA)
  • Evolutionary Computation (AREA)
  • Medical Informatics (AREA)
  • Software Systems (AREA)
  • Processing Or Creating Images (AREA)
  • Image Analysis (AREA)

Abstract

本申请公开了一种图文融合方法、装置及电子设备,涉及数字图像处理技术领域,可以在将文本排版至图像时,使得文本最小化地遮挡图像中的显著特征,以及使得文本排版至第一图像后,获得较佳的视觉平衡程度,从而获得较佳的排版效果。采用本申请的方法,首先可以确定出多个候选的文本模板以及对应的多个文本在图像中的排版位置,使得排版至图像中的文本不会遮挡图像中特征值较高的视觉显著主体,例如人脸、建筑物等。然后根据文本以不同文本模板以及在图像中对应的排版位置排版至图像时,文本遮挡的像素点的特征值的大小,以及排版有该文本的图像中各个区域的像素点的特征值分布的平衡程度等,确定文本最终的文本模板以及文本在图像中的排版位置。

Description

一种图文融合方法、装置及电子设备
本申请要求于2019年8月23日提交国家知识产权局、申请号为201910783866.7、发明名称为“一种图文融合方法、装置及电子设备”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请实施例涉及数字图像处理技术领域,尤其涉及一种图文融合方法、装置及电子设备。
背景技术
随着多媒体技术和互联网络技术的快速发展,文字与图像的融合排版使用越来越广泛。例如,手机的锁屏壁纸、应用(Application,APP)启动时的广告视频等、视频窗口的悬浮广告框等。
在进行文本与图像的融合排版时,如何使文本准确地避开图像中的显著主体(如人脸、花、建筑等),以及如何在避开显著主体的文本的候选区域,选取文本排版的较佳方式,是需要解决的问题。
发明内容
本申请实施例提供一种图文融合方法,可以在将文本排版至图像时获得较佳的排版效果。
为达到上述目的,本申请实施例采用如下技术方案:
第一方面,提供一种图文融合方法,该方法包括:获取第一图像和待排版至该第一图像的第一文本;确定该第一图像中各个像素点的特征值;其中,一个像素点的特征值用于表征该一个像素点被用户关注的可能性的高低,像素点的特征值越高,该像素点被用户关注的可能性越高;根据第一文本,以及该第一图像中各个像素点的特征值,确定第一文本在第一图像的多个第一排版版式;其中,第一文本按照每个第一排版版式排版至该第一图像时,该第一文本不遮挡特征值大于第一阈值的像素点;根据多个第一排版版式的代价参数,从该多个第一排版版式中确定出第二排版版式;其中,一个第一排版版式的代价参数用于表征第一文本按照该第一排版版式排版至第一图像时,第一文本遮挡的像素点的特征值的大小,以及排版有第一文本的第一图像中各个区域的像素点的特征值分布的平衡程度;按照该第二排版版式将第一文本排版至第一图像,得到第二图像。
上述第一方面提供的技术方案,可以确定出多个候选的文本模板以及对应的多个文本在图像中的排版位置,使得排版至图像中的文本不会遮挡图像中特征值较高的视觉显著主体,例如人脸、建筑物等。然后根据文本以不同文本模板以及在图像中对应的排版位置排版至图像时,文本遮挡的像素点的特征值的大小,以及排版有该文本的图像中各个区域的像素点的特征值分布的平衡程度等,确定文本最终的文本模板以及文本在图像中的排版位置,获得较佳的排版效果。
在一种可能的实现方式中,确定第一图像中各个像素点的特征值,包括:确定该第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数;其中,一个像素点的视觉显著参数用于表征该一个像素点是视觉显著性特征对应的像素点的可能性的高低,一个像素点的人脸特征参数用于表征该一个像素点是人脸对应的像素点的可能性的高低,一个像素点的边缘特征参数用于表征该一个像素点是物体轮廓对应的像素点的可能性的高低,一个像素点的文本特征参数用于表征该一个像素点是文本对应的像素点的可能性的高低;分别对确定出的第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数进行加权求和,确定第一图像中各个像素点的特征值。可以综合考虑视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数,以及每一个参数对应的特征被用户关注的可能性程度,确定每一个像素点的特征值。使得根据该特征值确定出的文本的排版位置不遮挡显著特征的可能性更高。
在一种可能的实现方式中,在分别对确定出的第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数进行加权求和,确定该第一图像中各个像素点的特征值之前,该方法还包括:分别根据确定出的第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数生成至少两个特征图;每一个特征图中各个像素点的像素值为对应像素点的对应参数;分别对确定出的该第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数进行加权求和,确定该第一图像中各个像素点的特征值,包括:对上述至少两个特征图中各个像素点的像素值进行加权求和,确定该第一图像中各个像素点的特征值。通过对分别用于表征视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数的特征图中的至少两个,结合每一个参数对应的特征被用户关注的可能性程度进行加权求和,确定每一个像素点的特征值。使得根据该特征值确定出的文本的排版位置不遮挡显著特征的可能性更高。
在一种可能的实现方式中,根据第一文本,以及第一图像中各个像素点的特征值,确定第一文本在该第一图像的多个第一排版版式,包括:根据该第一文本以一个或多个文本模板排版时该第一文本的文本框的大小,以及该第一图像中各个像素点的特征值,确定多个第一排版版式。通过综合分析第一文本以不同文本模板排版时,所占区域的大小,以及第一图像中各个像素点的特征值,可以保证第一文本在确定的文本的排版位置排版时,不遮挡第一图像中的显著特征。
在一种可能的实现方式中,该方法还包括:获取一个或多个文本模板,每个文本模板规定了文本的行间距、行宽、字号、字体、文字粗细、对齐方式、装饰线位置和装饰线粗细中的至少一种。本申请支持第一文本以不同的文本模板排版,灵活度高,图文融合效果更好。
在一种可能的实现方式中,根据多个第一排版版式的代价参数,从该多个第一排版版式中确定出第二排版版式,包括:确定第一文本分别按照上述多个第一排版版式排版至第一图像时,该第一文本的文本框遮挡第一图像的图像区域的纹理 特征参数,该纹理特征参数用于表征所述图像区域对应的图像中纹理特征的多少;从多个第一排版版式中,选择出纹理特征参数小于第二阈值的图像区域对应的多个第一排版版式;根据选择出的每个第一排版版式的代价参数,从选择出的多个第一排版版式中确定出第二排版版式。通过舍弃纹理复杂区域用作文本排版位置,可以避免该区域的纹理特征对第一文本显著性的影响,以及避免第一文本排版至该区域时,对该区域纹理特征的遮挡。
在一种可能的实现方式中,该方法还包括:针对多个第一排版版式中每个第一排版版式,执行步骤a、步骤b和步骤c中的至少两个以及步骤d,以得到每个第一排版版式的代价参数;步骤a:计算第一文本按照一个第一排版版式排版至第一图像时,该第一文本的文本入侵参数;该文本入侵参数是第一参数与第二参数的比值;第一参数是所述第一文本遮挡第一图像的图像区域中各个像素点的特征值之和;第二参数是图像区域的面积,或者,第二参数是图像区域中像素点的总数,或者,第二参数是图像区域中像素点的总数与预设数值的乘积;步骤b、计算该第一文本按照一个第一排版版式排版至第一图像时,该第一文本的视觉空间占用参数;该视觉空间占用参数用于表征图像区域中特征值小于第三阈值的像素点的比例;步骤c:计算该第一文本按照一个第一排版版式排版至第一图像时,该第一文本的视觉平衡参数;该视觉平衡参数用于表征第一文本对排版有该第一文本的第一图像中各个区域的像素点的特征值分布的平衡程度的影响程度;步骤d、根据计算出的第一文本的文本入侵参数、视觉空间占用参数和视觉平衡参数中的至少两个,计算一个第一排版版式的代价参数。通过综合考虑第一文本以不同候选排版版式排版时,第一文本遮挡图像区域特征值的具体情况,以及第一文本对第一图像中各个区域的像素点的特征值分布的平衡程度的影响程度,可以从多个候选排版版式中确定出较佳的排版版式。
在一种可能的实现方式中,根据计算出的第一文本的文本入侵参数、视觉空间占用参数和视觉平衡参数中的至少两个,计算一个第一排版版式的代价参数,包括:采用T i=λ 1*E s(L i)+λ 2*E u(L i)+λ 3*E n(L i),或者,T i=(λ 1*E s(L i)+λ 2*E u(L i))*E n(L i),或者,T i=E s(L i)*E u(L i)*E n(L i)计算该一个第一排版版式的代价参数T i;其中,E s(L i)为第一文本按照该一个第一排版版式排版至第一图像时,第一文本的文本入侵参数,E u(L i)为第一文本按照该一个第一排版版式排版至第一图像时,第一文本的视觉空间占用参数,E n(L i)为第一文本按照该一个第一排版版式排版至第一图像时,第一文本的视觉平衡参数;λ 1、λ 2和λ 3分别为E s(L i)、E u(L i)和E n(L i)对应的权重参数。通过采用上述计算方法,综合考虑第一文本以不同候选排版版式排版时,第一文本遮挡图像区域特征值的具体情况,计算得每个候选排版版式对应的代价参数。
在一种可能的实现方式中,根据多个第一排版版式的代价参数,从多个第一排版版式中确定出第二排版版式,包括:确定多个第一排版版式的代价参数中,最小的代价参数对应的第一排版版式为第二排版版式。通过确定代价参数最小值对应的候选排版版式为最终的文本排版版式,可以最大程度的确保第一文本排版至第一图像后的美学效果。
在一种可能的实现方式中,该方法还包括:确定第一文本的颜色参数;该第一文本的颜色参数为第一文本按照第二排版版式排版至第一图像时,第一图像被第一文本遮挡的图像区域的主色的衍生颜色;该主色的衍生颜色是指与主色色相相同,但是色调,饱和度和明度与主色的HSV不同的颜色;根据第一文本的颜色参数,对第二图像中的第一文本着色,获得第三图像。通过以第一文本可能遮挡第一图像的图像区域的主色的衍生颜色对第一文本着色,可以使第一文本排版至第一图像后,其颜色与背景图像更加协调,显示更加清晰。
在一种可能的实现方式中,第一图像被该第一文本遮挡的图像区域的主色是基于第一文本按照第二排版版式排版至第一图像时,第一图像被第一文本遮挡的图像区域的三原色光RGB在HSV空间中的色调、饱和度和明度确定的;该主色为该图像区域中色相占比最高的色相。通过根据第一文本可能遮挡第一图像的图像区域的色调、饱和度和明度确定该图像区域的主色,进而可以根据图像区域的主色的衍生颜色对第一文本着色,使第一文本排版至第一图像后,其颜色与背景图像更加协调,显示更加清晰。
在一种可能的实现方式中,该方法还包括:若满足以下条件1和条件2中的至少一个,确定对该第二图像进行渲染处理;条件1:第一文本按照第二排版版式排版至该第一图像后,第一文本遮挡该第一图像的图像区域的纹理特征参数大于第四阈值;该纹理特征参数用于表征图像区域对应的图像中纹理特征的多少;条件2:图像区域的主色占比小于第五阈值;在所第二图像上覆盖蒙版图层;或者,确定蒙版参数,根据确定的蒙版参数处理所述第二图像;或者,对第一文本进行投影渲染。通过蒙版渲染或者投影渲染,可以提高第一文本的清晰度和显著性。
在一种可能的实现方式中,该方法还包括:若满足以下条件1和条件2中的至少一个,确定对该第三图像进行渲染处理;条件1:第一文本按照第二排版版式排版至该第一图像后,第一文本遮挡该第一图像的图像区域的纹理特征参数大于第四阈值;该纹理特征参数用于表征图像区域对应的图像中纹理特征的多少;条件2:图像区域的主色占比小于第五阈值;在所第三图像上覆盖蒙版图层;或者,确定蒙版参数,根据确定的蒙版参数处理所述第三图像;或者,对第一文本进行投影渲染。通过蒙版渲染或者投影渲染,可以提高第一文本的清晰度和显著性。
第二方面,提供一种图文融合装置,该装置包括:信息获取单元,用于获取第一图像和待排版至该第一图像的第一文本;分析单元,用于确定该第一图像中各个像素点的特征值;其中,一个像素点的特征值用于表征该一个像素点被用户关注的可能性的高低,像素点的特征值越高,该像素点被用户关注的可能性越高;以及,根据第一文本,以及该第一图像中各个像素点的特征值,确定第一文本在第一图像的多个第一排版版式;其中,第一文本按照每个第一排版版式排版至该第一图像时,该第一文本不遮挡特征值大于第一阈值的像素点;以及,根据多个第一排版版式的代价参数,从该多个第一排版版式中确定出第二排版版式;其中,一个第一排版版式的代价参数用于表征第一文本按照该第一排版版式排版至第一图像时,第一文本遮挡的像素点的特征值的大小,以及排版有第一文本的第一图像中各个区域的像素点的特征值分布的平衡程度;处理单元,按照该第二排版版式 将第一文本排版至第一图像,得到第二图像。
上述第二方面提供的装置,可以确定出多个候选的文本模板以及对应的多个文本在图像中的排版位置,使得排版至图像中的文本不会遮挡图像中特征值较高的视觉显著主体,例如人脸、建筑物等。然后根据文本以不同文本模板以及在图像中对应的排版位置排版至图像时,文本遮挡的像素点的特征值的大小,以及排版有该文本的图像中各个区域的像素点的特征值分布的平衡程度等,确定文本最终的文本模板以及文本在图像中的排版位置,获得较佳的排版效果。
在一种可能的实现方式中,分析单元确定第一图像中各个像素点的特征值,包括:分析单元确定该第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数;其中,一个像素点的视觉显著参数用于表征该一个像素点是视觉显著性特征对应的像素点的可能性的高低,一个像素点的人脸特征参数用于表征该一个像素点是人脸对应的像素点的可能性的高低,一个像素点的边缘特征参数用于表征该一个像素点是物体轮廓对应的像素点的可能性的高低,一个像素点的文本特征参数用于表征该一个像素点是文本对应的像素点的可能性的高低;分析单元分别对确定出的第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数进行加权求和,确定第一图像中各个像素点的特征值。可以综合考虑视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数,以及每一个参数对应的特征被用户关注的可能性程度,确定每一个像素点的特征值。使得根据该特征值确定出的文本的排版位置不遮挡显著特征的可能性更高。
在一种可能的实现方式中,在分析单元分别对确定出的第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数进行加权求和,确定该第一图像中各个像素点的特征值之前,该分析单元还用于:分别根据确定出的第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数生成该第一图像的至少两个特征图;每一个特征图中各个像素点的像素值为对应像素点的对应参数;分别对确定出的该第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数进行加权求和,确定该第一图像中各个像素点的特征值,包括:对上述至少两个特征图中各个像素点的像素值进行加权求和,确定该第一图像中各个像素点的特征值。通过对分别用于表征视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数的特征图中的至少两个,结合每一个参数对应的特征被用户关注的可能性程度进行加权求和,确定每一个像素点的特征值。使得根据该特征值确定出的文本的排版位置不遮挡显著特征的可能性更高。
在一种可能的实现方式中,分析单元根据第一文本,以及第一图像中各个像素点的特征值,确定第一文本在该第一图像的多个第一排版版式,包括:分析单元根据该第一文本以一个或多个文本模板排版时该第一文本的文本框的大小,以及该第一图像中各个像素点的特征值,确定多个第一排版版式。通过综合分析第一文本以不同文本模板排版时,所占区域的大小以及第一图像中各个像素点的特征值,可以保证第一文本在确定的文本的排版位置排版时,不遮挡第一图像中的显著特 征。在一种可能的实现方式中,该信息获取单元还用于:获取一个或多个文本模板,每个文本模板规定了文本的行间距、行宽、字号、字体、文字粗细、对齐方式、装饰线位置和装饰线粗细中的至少一种。本申请支持第一文本以不同的文本模板排版,灵活度高,图文融合效果更好。
在一种可能的实现方式中,分析单元根据多个第一排版版式的代价参数,从该多个第一排版版式中确定出第二排版版式,包括:分析单元确定第一文本分别按照上述多个第一排版版式排版至第一图像时,该第一文本的文本框遮挡第一图像的图像区域的纹理特征参数,该纹理特征参数用于表征所述图像区域对应的图像中纹理特征的多少;分析单元从多个第一排版版式中,选择出纹理特征参数小于第二阈值的图像区域对应的多个第一排版版式;根据选择出的每个第一排版版式的代价参数,从选择出的多个第一排版版式中确定出第二排版版式。通过舍弃纹理复杂区域用作文本排版位置,可以避免该区域的纹理特征对第一文本显著性的影响,以及避免第一文本排版至该区域时,对该区域纹理特征的遮挡。
在一种可能的实现方式中,分析单元还用于:针对多个第一排版版式中每个第一排版版式,执行步骤a、步骤b和步骤c中的至少两个以及步骤d,以得到每个第一排版版式的代价参数;步骤a:计算第一文本按照一个第一排版版式排版至第一图像时,该第一文本的文本入侵参数;该文本入侵参数是第一参数与第二参数的比值;第一参数是所述第一文本遮挡第一图像的图像区域中各个像素点的特征值之和;第二参数是图像区域的面积,或者,第二参数是图像区域中像素点的总数,或者,第二参数是图像区域中像素点的总数与预设数值的乘积;步骤b、计算该第一文本按照一个第一排版版式排版至第一图像时,该第一文本的视觉空间占用参数;该视觉空间占用参数用于表征图像区域中特征值小于第三阈值的像素点的比例;步骤c:计算该第一文本按照一个第一排版版式排版至第一图像时,该第一文本的视觉平衡参数;该视觉平衡参数用于表征第一文本对排版有该第一文本的第一图像中各个区域的像素点的特征值分布的平衡程度的影响程度;步骤d、根据计算出的第一文本的文本入侵参数、视觉空间占用参数和视觉平衡参数中的至少两个,计算一个第一排版版式的代价参数。通过综合考虑第一文本以不同候选排版版式排版时,第一文本遮挡图像区域特征值的具体情况,以及第一文本对第一图像中各个区域的像素点的特征值分布的平衡程度的影响程度,可以从多个候选排版版式中确定出较佳的排版版式。
在一种可能的实现方式中,分析单元根据计算出的第一文本的文本入侵参数、视觉空间占用参数和视觉平衡参数中的至少两个,计算一个第一排版版式的代价参数,包括:分析单元采用T i=λ 1*E s(L i)+λ 2*E u(L i)+λ 3*E n(L i),或者,T i=(λ 1*E s(L i)+λ 2*E u(L i))*E n(L i),或者,T i=E s(L i)*E u(L i)*E n(L i)计算该一个第一排版版式的代价参数T i;其中,E s(L i)为第一文本按照该一个第一排版版式排版至第一图像时,第一文本的文本入侵参数,E u(L i)为第一文本按照该一个第一排版版式排版至第一图像时,第一文本的视觉空间占用参数,E n(L i)为第一文本按照该一个第一排版版式排版至第一图像时,第一文本的视觉平衡参数;λ 1、λ 2和λ 3分别为E s(L i)、E u(L i)和E n(L i)对应的权重参数。通过采用上述计算方法,综合考虑第一 文本以不同候选排版版式排版时,第一文本遮挡图像区域特征值的具体情况,计算得每个候选排版版式对应的代价参数。
在一种可能的实现方式中,分析单元根据多个第一排版版式的代价参数,从多个第一排版版式中确定出第二排版版式,包括:分析单元确定多个第一排版版式的代价参数中,最小的代价参数对应的第一排版版式为第二排版版式。通过确定代价参数最小值对应的候选排版版式为最终的文本排版版式,可以最大程度的确保第一文本排版至第一图像后的美学效果。
在一种可能的实现方式中,分析单元还用于:确定第一文本的颜色参数;该第一文本的颜色参数为第一文本按照第二排版版式排版至第一图像时,第一图像被第一文本遮挡的图像区域的主色的衍生颜色;该主色的衍生颜色是指与主色色相相同,但是色调,饱和度和明度与主色的HSV不同的颜色;以及根据第一文本的颜色参数,对第二图像中的第一文本着色,获得第三图像。通过以第一文本可能遮挡第一图像的图像区域的主色的衍生颜色对第一文本着色,可以使第一文本排版至第一图像后,其颜色与背景图像更加协调,显示更加清晰。
在一种可能的实现方式中,第一图像被该第一文本遮挡的图像区域的主色是基于第一文本按照第二排版版式排版至第一图像时,第一图像被第一文本遮挡的图像区域的三原色光RGB在HSV空间中的色调、饱和度和明度确定的;该主色为该图像区域中色相占比最高的色相。通过根据第一文本可能遮挡第一图像的图像区域的色调、饱和度和明度确定该图像区域的主色,进而可以根据图像区域的主色的衍生颜色对第一文本着色,使第一文本排版至第一图像后,其颜色与背景图像更加协调,显示更加清晰。
在一种可能的实现方式中,分析单元还用于:若满足以下条件1和条件2中的至少一个,确定对该第二图像进行渲染处理;条件1:第一文本按照第二排版版式排版至该第一图像后,第一文本遮挡该第一图像的图像区域的纹理特征参数大于第四阈值;该纹理特征参数用于表征图像区域对应的图像中纹理特征的多少;条件2:图像区域的主色占比小于第五阈值;处理单元还用于,在所第二图像上覆盖蒙版图层;或者,确定蒙版参数,根据确定的蒙版参数处理所述第二图像;或者,对第一文本进行投影渲染。通过蒙版渲染或者投影渲染,可以提高第一文本的清晰度和显著性。
在一种可能的实现方式中,分析单元还用于:若满足以下条件1和条件2中的至少一个,确定对该第三图像进行渲染处理;条件1:第一文本按照第二排版版式排版至该第一图像后,第一文本遮挡该第一图像的图像区域的纹理特征参数大于第四阈值;该纹理特征参数用于表征图像区域对应的图像中纹理特征的多少;条件2:图像区域的主色占比小于第五阈值;处理单元还用于,在所第三图像上覆盖蒙版图层;或者,确定蒙版参数,根据确定的蒙版参数处理所述第三图像;或者,对第一文本进行投影渲染。通过蒙版渲染或者投影渲染,可以提高第一文本的清晰度和显著性。
第三方面,提供一种电子设备,该电子设备包括:信息获取单元,用于获取第一图像和待排版至该第一图像的第一文本;分析单元,用于确定该第一图像中各个 像素点的特征值;其中,一个像素点的特征值用于表征该一个像素点被用户关注的可能性的高低,像素点的特征值越高,该像素点被用户关注的可能性越高;以及,根据第一文本,以及该第一图像中各个像素点的特征值,确定第一文本在第一图像的多个第一排版版式;其中,第一文本按照每个第一排版版式排版至该第一图像时,该第一文本不遮挡特征值大于第一阈值的像素点;以及,根据多个第一排版版式的代价参数,从该多个第一排版版式中确定出第二排版版式;其中,一个第一排版版式的代价参数用于表征第一文本按照该第一排版版式排版至第一图像时,第一文本遮挡的像素点的特征值的大小,以及排版有第一文本的第一图像中各个区域的像素点的特征值分布的平衡程度;处理单元,按照该第二排版版式将第一文本排版至第一图像,得到第二图像。
上述第三方面提供的电子设备,可以确定出多个候选的文本模板以及对应的多个文本在图像中的排版位置,使得排版至图像中的文本不会遮挡图像中特征值较高的视觉显著主体,例如人脸、建筑物等。然后根据文本以不同文本模板以及在图像中对应的排版位置排版至图像时,文本遮挡的像素点的特征值的大小,以及排版有该文本的图像中各个区域的像素点的特征值分布的平衡程度等,确定文本最终的文本模板以及文本在图像中的排版位置,获得较佳的排版效果。
在一种可能的实现方式中,分析单元确定第一图像中各个像素点的特征值,包括:分析单元确定该第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数;其中,一个像素点的视觉显著参数用于表征该一个像素点是视觉显著性特征对应的像素点的可能性的高低,一个像素点的人脸特征参数用于表征该一个像素点是人脸对应的像素点的可能性的高低,一个像素点的边缘特征参数用于表征该一个像素点是物体轮廓对应的像素点的可能性的高低,一个像素点的文本特征参数用于表征该一个像素点是文本对应的像素点的可能性的高低;分析单元分别对确定出的第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数进行加权求和,确定第一图像中各个像素点的特征值。可以综合考虑视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数,以及每一个参数对应的特征被用户关注的可能性程度,确定每一个像素点的特征值。使得根据该特征值确定出的文本的排版位置不遮挡显著特征的可能性更高。
在一种可能的实现方式中,在分析单元分别对确定出的第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数进行加权求和,确定该第一图像中各个像素点的特征值之前,该分析单元还用于:分别根据确定出的第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数生成至少两个特征图;每一个特征图中各个像素点的像素值为对应像素点的对应参数;分别对确定出的该第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数进行加权求和,确定该第一图像中各个像素点的特征值,包括:对上述至少两个特征图中各个像素点的像素值进行加权求和,确定该第一图像中各个像素点的特征值。通过对分别用于表征视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数的特征图 中的至少两个,结合每一个参数对应的特征被用户关注的可能性程度进行加权求和,确定每一个像素点的特征值。使得根据该特征值确定出的文本的排版位置不遮挡显著特征的可能性更高。
在一种可能的实现方式中,分析单元根据第一文本,以及第一图像中各个像素点的特征值,确定第一文本在该第一图像的多个第一排版版式,包括:分析单元根据该第一文本以一个或多个文本模板排版时该第一文本的文本框的大小,以及该第一图像中各个像素点的特征值,确定多个第一排版版式。通过综合分析第一文本以不同文本模板排版时,所占区域的大小,以及第一图像中各个像素点的特征值,可以保证第一文本在确定的文本的排版位置排版时,不遮挡第一图像中的显著特征。
在一种可能的实现方式中,该信息获取单元还用于:获取一个或多个文本模板,每个文本模板规定了文本的行间距、行宽、字号、字体、文字粗细、对齐方式、装饰线位置和装饰线粗细中的至少一种。本申请支持第一文本以不同的文本模板排版,灵活度高,图文融合效果更好。
在一种可能的实现方式中,分析单元根据多个第一排版版式的代价参数,从该多个第一排版版式中确定出第二排版版式,包括:分析单元确定第一文本分别按照上述多个第一排版版式排版至第一图像时,该第一文本的文本框遮挡第一图像的图像区域的纹理特征参数,该纹理特征参数用于表征所述图像区域对应的图像中纹理特征的多少;分析单元从多个第一排版版式中,选择出纹理特征参数小于第二阈值的图像区域对应的多个第一排版版式;根据选择出的每个第一排版版式的代价参数,从选择出的多个第一排版版式中确定出第二排版版式。通过舍弃纹理复杂区域用作文本排版位置,可以避免该区域的纹理特征对第一文本显著性的影响,以及避免第一文本排版至该区域时,对该区域纹理特征的遮挡。
在一种可能的实现方式中,分析单元还用于:针对多个第一排版版式中每个第一排版版式,执行步骤a、步骤b和步骤c中的至少两个以及步骤d,以得到每个第一排版版式的代价参数;步骤a:计算第一文本按照一个第一排版版式排版至第一图像时,该第一文本的文本入侵参数;该文本入侵参数是第一参数与第二参数的比值;第一参数是所述第一文本遮挡第一图像的图像区域中各个像素点的特征值之和;第二参数是图像区域的面积,或者,第二参数是图像区域中像素点的总数,或者,第二参数是图像区域中像素点的总数与预设数值的乘积;步骤b、计算该第一文本按照一个第一排版版式排版至第一图像时,该第一文本的视觉空间占用参数;该视觉空间占用参数用于表征图像区域中特征值小于第三阈值的像素点的比例;步骤c:计算该第一文本按照一个第一排版版式排版至第一图像时,该第一文本的视觉平衡参数;该视觉平衡参数用于表征第一文本对排版有该第一文本的第一图像中各个区域的像素点的特征值分布的平衡程度的影响程度;步骤d、根据计算出的第一文本的文本入侵参数、视觉空间占用参数和视觉平衡参数中的至少两个,计算一个第一排版版式的代价参数。通过综合考虑第一文本以不同候选排版版式排版时,第一文本遮挡图像区域特征值的具体情况,以及第一文本对第一图像中各个区域的像素点的特征值分布的平衡程度的影响程度,可以从多个候选排 版版式中确定出较佳的排版版式。
在一种可能的实现方式中,分析单元根据计算出的第一文本的文本入侵参数、视觉空间占用参数和视觉平衡参数中的至少两个,计算一个第一排版版式的代价参数,包括:分析单元采用T i=λ 1*E s(L i)+λ 2*E u(L i)+λ 3*E n(L i),或者,T i=(λ 1*E s(L i)+λ 2*E u(L i))*E n(L i),或者,T i=E s(L i)*E u(L i)*E n(L i)计算该一个第一排版版式的代价参数T i;其中,E s(L i)为第一文本按照该一个第一排版版式排版至第一图像时,第一文本的文本入侵参数,E u(L i)为第一文本按照该一个第一排版版式排版至第一图像时,第一文本的视觉空间占用参数,E n(L i)为第一文本按照该一个第一排版版式排版至第一图像时,第一文本的视觉平衡参数;λ 1、λ 2和λ 3分别为E s(L i)、E u(L i)和E n(L i)对应的权重参数。通过采用上述计算方法,综合考虑第一文本以不同候选排版版式排版时,第一文本遮挡图像区域特征值的具体情况,计算得每个候选排版版式对应的代价参数。
在一种可能的实现方式中,分析单元根据多个第一排版版式的代价参数,从多个第一排版版式中确定出第二排版版式,包括:分析单元确定多个第一排版版式的代价参数中,最小的代价参数对应的第一排版版式为第二排版版式。通过确定代价参数最小值对应的候选排版版式为最终的文本排版版式,可以最大程度的确保第一文本排版至第一图像后的美学效果。
在一种可能的实现方式中,分析单元还用于:确定第一文本的颜色参数;该第一文本的颜色参数为第一文本按照第二排版版式排版至第一图像时,第一图像被第一文本遮挡的图像区域的主色的衍生颜色;该主色的衍生颜色是指与主色色相相同,但是色调,饱和度和明度与主色的HSV不同的颜色;以及根据第一文本的颜色参数,对第二图像中的第一文本着色,获得第三图像。通过以第一文本可能遮挡第一图像的图像区域的主色的衍生颜色对第一文本着色,可以使第一文本排版至第一图像后,其颜色与背景图像更加协调,显示更加清晰。
在一种可能的实现方式中,第一图像被该第一文本遮挡的图像区域的主色是基于第一文本按照第二排版版式排版至第一图像时,第一图像被第一文本遮挡的图像区域的三原色光RGB在HSV空间中的色调、饱和度和明度确定的;该主色为该图像区域中色相占比最高的色相色。通过根据第一文本可能遮挡第一图像的图像区域的色调、饱和度和明度确定该图像区域的主色,进而可以根据图像区域的主色的衍生颜色对第一文本着色,使第一文本排版至第一图像后,其颜色与背景图像更加协调,显示更加清晰。
在一种可能的实现方式中,分析单元还用于:若满足以下条件1和条件2中的至少一个,确定对该第二图像进行渲染处理;条件1:第一文本按照第二排版版式排版至该第一图像后,第一文本遮挡该第一图像的图像区域的纹理特征参数大于第四阈值;该纹理特征参数用于表征图像区域对应的图像中纹理特征的多少;条件2:图像区域的主色占比小于第五阈值;处理单元还用于,在所第二图像上覆盖蒙版图层;或者,确定蒙版参数,根据确定的蒙版参数处理所述第二图像;或者,对第一文本进行投影渲染。通过蒙版渲染或者投影渲染,可以提高第一文本的清晰度和显著性。
在一种可能的实现方式中,分析单元还用于:若满足以下条件1和条件2中的至少一个,确定对该第三图像进行渲染处理;条件1:第一文本按照第二排版版式排版至该第一图像后,第一文本遮挡该第一图像的图像区域的纹理特征参数大于第四阈值;该纹理特征参数用于表征图像区域对应的图像中纹理特征的多少;条件2:图像区域的主色占比小于第五阈值;处理单元还用于,在所第三图像上覆盖蒙版图层;或者,确定蒙版参数,根据确定的蒙版参数处理所述第三图像;或者,对第一文本进行投影渲染。通过蒙版渲染或者投影渲染,可以提高第一文本的清晰度和显著性。
第四方面,提供一种图文融合装置,该装置包括:存储器,用于存储一个或多个计算机程序;处理器,用于执行存储器存储的一个或多个计算机程序,使得该图文融合装置实现如第一方面任一种可能的实现方式中的图文融合方法。
第五方面,提供一种电子设备,该电子设备包括:存储器,用于存储一个或多个计算机程序;处理器,用于执行存储器存储的一个或多个计算机程序,使得该图文融合装置实现如第一方面任一种可能的实现方式中的图文融合方法。
第六方面,提供一种计算机可读存储介质,该计算机可读存储介质上存储有计算机执行指令,该计算机执行指令被处理器执行时实现如第一方面任一种可能的实现方式中的图文融合方法。
第七方面,提供一种芯片系统,该芯片系统包括处理器、存储器,存储器中存储有指令;所述指令被所述处理器执行时,实现如第一方面任一种可能的实现方式中的图文融合方法。该芯片系统可以由芯片构成,也可以包含芯片和其他分立器件。
第八方面,提供一种计算机程序产品,提供一种计算机程序产品,当其在计算机上运行时,使得第一方面任一种可能的实现方式中的图文融合方法。例如,该计算机可以是至少一个存储节点。
附图说明
图1为本申请实施例提供的一种图文融合示例图;
图2为本申请实施例提供的一种电子设备硬件结构示意图;
图3为本申请实施例提供的一种图文融合应用场景示例;
图4为本申请实施例提供的另一种图文融合应用场景示例;
图5为本申请实施例提供的再一种图文融合应用场景示例;
图6为本申请实施例提供的一种图文融合方法流程图;
图7为本申请实施例提供的一种相同图像对应不同文本的示例图;
图8为本申请实施例提供的一种显著区域特征图生成过程示例图;
图9为本申请实施例提供的另一种显著区域特征图生成过程示例图;
图10为本申请实施例提供的一种视觉显著性特征图生成过程示意图;
图11为本申请实施例提供的一种人脸检测算法流程图;
图12为本申请实施例提供的一种边缘检测算法流程图;
图13为本申请实施例提供的一种文本检测算法流程图;
图14为本申请实施例提供的一种JSON格式的排版规格示例图;
图15为本申请实施例提供的一种文本框排版位置示意图;
图16为本申请实施例提供的几种文字模板示例图;
图17为本申请实施例提供的几种候选排版版式示例图;
图18为本申请实施例提供的一种确定第二排版版式的方法流程图;
图19为本申请实施例提供的另一种确定第二排版版式的方法流程图;
图20为本申请实施例提供的另一种图文融合方法流程图;
图21为本申请实施例提供的几种图文融合图像对比图;
图22为本申请实施例提供的一种电子设备结构示意图。
具体实施方式
本申请实施例提供一种图文融合方法,该方法可以应用于将文本排版至图像(如第一图像)中,实现图文融合的过程中。可以理解,将文本排版至图像后,通常情况下该文本会遮挡图像的一部分区域。通过本申请实施例的方法将文本排版至图像后,排版至图像中的文本,不会遮挡图像中的显著特征。示例性的,如图1所示,第一图像为包括建筑物特征的图像,待排版至第一图像的文本为主题为“乡野一瞥”的文本。采用本申请实施例的图文融合方法可以将主题为“乡野一瞥”的文本排版至包括建筑物特征的第一图像中的合适位置,使得文本最小化地遮挡图像中的显著特征,进一步地还可以使得文本排版至第一图像后,获得较佳的视觉平衡程度。
其中,显著特征是指被用户关注的可能性较高的图像特征。例如,显著特征可以包括人脸特征、人体特征、建筑物特征、事物特征(如动物特征、树木特征、花朵特征等)、文字特征、河流特征和山川特征等。如图1所示,显著特征为建筑物特征。
需要说明的是,本申请实施例的图文融合方法可以应用于能够提供图像展示的终端类电子设备。包括桌面型设备、膝上型设备、手持型设备、可穿戴设备等。例如,应用于手机、平板电脑、个人计算机、智能相机、上网本、个人数字助理(Personal Digital Assistant,PDA)、智能手表、AR(增强现实)/VR(虚拟现实)设备等。
或者,本申请实施例的图文融合方法还可以应用于具备或不具备图像展示功能的图像处理装置或服务器类电子设备(例如,应用服务器)等。本申请实施例对执行本申请实施例的图文融合方法的电子设备的具体类型和结构等不作限定。
请参考图2,如图2所示,为本申请实施例提供的一种终端类电子设备200的硬件结构示意图。如图2所示,电子设备200可以包括处理器210,存储器(包括外部存储器接口220和内部存储器221),通用串行总线(universal serial bus,USB)接口230,充电管理模块240,电源管理模块241,电池242,天线1,天线2,移动通信模块250,无线通信模块260,音频模块270,扬声器270A,受话器270B,麦克风270C,耳机接口270D,传感器模块280,按键290,马达291,指示器292,摄像头293,显示屏294,以及用户标识模块(subscriber identification module,SIM)卡接口295等。其中,传感器模块280可以包括压力传感器280A,陀螺仪传感器280B,气压传感器280C,磁传感器280D,加速度传感器280E,距离传感器280F,接近光传感器280G,指纹传感器280H,重力传感器280I,温度传感器280J,触摸传感器280K,环境光传感器280L和骨传导传感器280M等。
可以理解的是,本发明实施例示意的结构并不构成对电子设备200的具体限定。 在本申请另一些实施例中,电子设备200可以包括比图示更多或更少的部件,或者组合某些部件,或者拆分某些部件,或者不同的部件布置。图示的部件可以以硬件,软件或软件和硬件的组合实现。
处理器210可以包括一个或多个处理单元,例如:处理器210可以包括应用处理器(application processor,AP),调制解调处理器,图形处理器(graphics processing unit,GPU),图像信号处理器(image signal processor,ISP),控制器,视频编解码器,数字信号处理器(digital signal processor,DSP),基带处理器,和/或神经网络处理器(neural-network processing unit,NPU)等。其中,不同的处理单元可以是独立的器件,也可以集成在一个或多个处理器中。
控制器可以根据指令操作码和时序信号,产生操作控制信号,完成取指令和执行指令的控制。
处理器210中还可以设置存储器,用于存储指令和数据。在一些实施例中,处理器210中的存储器为高速缓冲存储器。该存储器可以保存处理器210刚用过或循环使用的指令或数据。如果处理器210需要再次使用该指令或数据,可从所述存储器中直接调用。避免了重复存取,减少了处理器210的等待时间,因而提高了系统的效率。
在一些实施例中,处理器210可以包括一个或多个接口。接口可以包括集成电路(inter-integrated circuit,I2C)接口,集成电路内置音频(inter-integrated circuit sound,I2S)接口,脉冲编码调制(pulse code modulation,PCM)接口,通用异步收发传输器(universal asynchronous receiver/transmitter,UART)接口,移动产业处理器接口(mobile industry processor interface,MIPI),通用输入输出(general-purpose input/output,GPIO)接口,用户标识模块(subscriber identity module,SIM)接口,和/或通用串行总线(universal serial bus,USB)接口等。
I2C接口是一种双向同步串行总线,包括一根串行数据线(serial data line,SDA)和一根串行时钟线(derail clock line,SCL)。在一些实施例中,处理器210可以包含多组I2C总线。处理器210可以通过不同的I2C总线接口分别耦合触摸传感器280K,充电器,闪光灯,摄像头293等。例如:处理器210可以通过I2C接口耦合触摸传感器280K,使处理器210与触摸传感器280K通过I2C总线接口通信,实现电子设备200的触摸功能。
MIPI接口可以被用于连接处理器210与显示屏294,摄像头293等外围器件。MIPI接口包括摄像头串行接口(camera serial interface,CSI),显示屏串行接口(display serial interface,DSI)等。在一些实施例中,处理器210和摄像头293通过CSI接口通信,实现电子设备200的拍摄功能。处理器210和显示屏294通过DSI接口通信,实现电子设备200的显示功能。
GPIO接口可以通过软件配置。GPIO接口可以被配置为控制信号,也可被配置为数据信号。在一些实施例中,GPIO接口可以用于连接处理器210与摄像头293,显示屏294,无线通信模块260,音频模块270,传感器模块280等。GPIO接口还可以被配置为I2C接口,I2S接口,UART接口,MIPI接口等。
USB接口230是符合USB标准规范的接口,具体可以是Mini USB接口,Micro USB接口,USB Type C接口等。USB接口230可以用于连接充电器为电子设备200充电, 也可以用于电子设备200与外围设备之间传输数据。也可以用于连接耳机,通过耳机播放音频。该接口还可以用于连接其他电子设备,例如AR设备等。
可以理解的是,本发明实施例示意的各模块间的接口连接关系,只是示意性说明,并不构成对电子设备200的结构限定。在本申请另一些实施例中,电子设备200也可以采用上述实施例中不同的接口连接方式,或多种接口连接方式的组合。
充电管理模块240用于从充电器接收充电输入。其中,充电器可以是无线充电器,也可以是有线充电器。在一些有线充电的实施例中,充电管理模块240可以通过USB接口230接收有线充电器的充电输入。在一些无线充电的实施例中,充电管理模块240可以通过电子设备200的无线充电线圈接收无线充电输入。充电管理模块240为电池242充电的同时,还可以通过电源管理模块241为电子设备供电。
电源管理模块241用于连接电池242,充电管理模块240与处理器210。电源管理模块241接收电池242和/或充电管理模块240的输入,为处理器210,内部存储器221,显示屏294,摄像头293,和无线通信模块260等供电。电源管理模块241还可以用于监测电池容量,电池循环次数,电池健康状态(漏电,阻抗)等参数。在其他一些实施例中,电源管理模块241也可以设置于处理器210中。在另一些实施例中,电源管理模块241和充电管理模块240也可以设置于同一个器件中。
电子设备200的无线通信功能可以通过天线1,天线2,移动通信模块250,无线通信模块260,调制解调处理器以及基带处理器等实现。
天线1和天线2用于发射和接收电磁波信号。电子设备200中的每个天线可用于覆盖单个或多个通信频带。不同的天线还可以复用,以提高天线的利用率。例如:可以将天线1复用为无线局域网的分集天线。在另外一些实施例中,天线可以和调谐开关结合使用。
移动通信模块250可以提供应用在电子设备200上的包括2G/3G/4G/5G等无线通信的解决方案。移动通信模块250可以包括至少一个滤波器,开关,功率放大器,低噪声放大器(low noise amplifier,LNA)等。移动通信模块250可以由天线1接收电磁波,并对接收的电磁波进行滤波,放大等处理,传送至调制解调处理器进行解调。移动通信模块250还可以对经调制解调处理器调制后的信号放大,经天线1转为电磁波辐射出去。在一些实施例中,移动通信模块250的至少部分功能模块可以被设置于处理器210中。在一些实施例中,移动通信模块250的至少部分功能模块可以与处理器210的至少部分模块被设置在同一个器件中。
调制解调处理器可以包括调制器和解调器。其中,调制器用于将待发送的低频基带信号调制成中高频信号。解调器用于将接收的电磁波信号解调为低频基带信号。随后解调器将解调得到的低频基带信号传送至基带处理器处理。低频基带信号经基带处理器处理后,被传递给应用处理器。应用处理器通过音频设备(不限于扬声器270A,受话器270B等)输出声音信号,或通过显示屏294显示图像或视频。在一些实施例中,调制解调处理器可以是独立的器件。在另一些实施例中,调制解调处理器可以独立于处理器210,与移动通信模块250或其他功能模块设置在同一个器件中。
无线通信模块260可以提供应用在电子设备200上的包括无线局域网(wireless local area networks,WLAN)(如无线保真(wireless fidelity,Wi-Fi)网络),蓝牙(bluetooth, BT),全球导航卫星系统(global navigation satellite system,GNSS),调频(frequency modulation,FM),近距离无线通信技术(near field communication,NFC),红外技术(infrared,IR)等无线通信的解决方案。无线通信模块260可以是集成至少一个通信处理模块的一个或多个器件。无线通信模块260经由天线2接收电磁波,将电磁波信号调频以及滤波处理,将处理后的信号发送到处理器210。无线通信模块260还可以从处理器210接收待发送的信号,对其进行调频,放大,经天线2转为电磁波辐射出去。
在一些实施例中,电子设备200的天线1和移动通信模块250耦合,天线2和无线通信模块260耦合,使得电子设备200可以通过无线通信技术与网络以及其他设备通信。所述无线通信技术可以包括全球移动通讯系统(global system for mobile communications,GSM),通用分组无线服务(general packet radio service,GPRS),码分多址接入(code division multiple access,CDMA),宽带码分多址(wideband code division multiple access,WCDMA),时分码分多址(time-division code division multiple access,TD-SCDMA),长期演进(long term evolution,LTE),BT,GNSS,WLAN,NFC,FM,和/或IR技术等。所述GNSS可以包括全球卫星定位系统(global positioning system,GPS),全球导航卫星系统(global navigation satellite system,GLONASS),北斗卫星导航系统(beidou navigation satellite system,BDS),准天顶卫星系统(quasi-zenith satellite system,QZSS)和/或星基增强系统(satellite based augmentation systems,SBAS)。
电子设备200通过GPU,显示屏294,以及应用处理器等实现显示功能。GPU为图像处理的微处理器,连接显示屏294和应用处理器。GPU用于执行数学和几何计算,用于图形渲染。处理器210可包括一个或多个GPU,其执行程序指令以生成或改变显示信息。在本申请实施例中,电子设备200可以通过GPU完成图文融合,可以通过显示屏294显示图文融合后的图像。
显示屏294用于显示图像,视频等。显示屏294包括显示面板。显示面板可以采用液晶显示屏(liquid crystal display,LCD),有机发光二极管(organic light-emitting diode,OLED),有源矩阵有机发光二极体或主动矩阵有机发光二极体(active-matrix organic light emitting diode的,AMOLED),柔性发光二极管(flex light-emitting diode,FLED),Miniled,MicroLed,Micro-oLed,量子点发光二极管(quantum dot light emitting diodes,QLED)等。在一些实施例中,电子设备200可以包括1个或N个显示屏294,N为大于1的正整数。
电子设备200可以通过ISP,摄像头293,视频编解码器,GPU,显示屏294以及应用处理器等实现拍摄功能。
ISP用于处理摄像头293反馈的数据。例如,拍照时,打开快门,光线通过镜头被传递到摄像头感光元件上,光信号转换为电信号,摄像头感光元件将所述电信号传递给ISP处理,转化为肉眼可见的图像。ISP还可以对图像的噪点,亮度,肤色进行算法优化。ISP还可以对拍摄场景的曝光,色温等参数优化。在一些实施例中,ISP可以设置在摄像头293中。
摄像头293用于捕获静态图像或视频。物体通过镜头生成光学图像投射到感光元件。感光元件可以是电荷耦合器件(charge coupled device,CCD)或互补金属氧化物半导体(complementary metal-oxide-semiconductor,CMOS)光电晶体管。感光元件把光信 号转换成电信号,之后将电信号传递给ISP转换成数字图像信号。ISP将数字图像信号输出到DSP加工处理。DSP将数字图像信号转换成标准的RGB,YUV等格式的图像信号。在一些实施例中,电子设备200可以包括1个或N个摄像头293,N为大于1的正整数。
数字信号处理器用于处理数字信号,除了可以处理数字图像信号,还可以处理其他数字信号。例如,当电子设备200在频点选择时,数字信号处理器用于对频点能量进行傅里叶变换等。
视频编解码器用于对数字视频压缩或解压缩。电子设备200可以支持一种或多种视频编解码器。这样,电子设备200可以播放或录制多种编码格式的视频,例如:动态图像专家组(moving picture experts group,MPEG)1,MPEG2,MPEG3,MPEG4等。
NPU为神经网络(neural-network,NN)计算处理器,通过借鉴生物神经网络结构,例如借鉴人脑神经元之间传递模式,对输入信息快速处理,还可以不断的自学习。通过NPU可以实现电子设备200的智能认知等应用,例如:图像识别,人脸识别,语音识别,文本理解等。
外部存储器接口220可以用于连接外部存储卡,例如Micro SD卡,实现扩展电子设备200的存储能力。外部存储卡通过外部存储器接口220与处理器210通信,实现数据存储功能。例如将音乐,视频等文件保存在外部存储卡中。
内部存储器221可以用于存储计算机可执行程序代码,所述可执行程序代码包括指令。内部存储器221可以包括存储程序区和存储数据区。其中,存储程序区可存储操作系统,至少一个功能所需的应用程序(比如声音播放功能,图像播放功能等)等。存储数据区可存储电子设备200使用过程中所创建的数据(比如音频数据,电话本等)等。此外,内部存储器221可以包括高速随机存取存储器,还可以包括非易失性存储器,例如至少一个磁盘存储器件,闪存器件,通用闪存存储器(universal flash storage,UFS)等。处理器210通过运行存储在内部存储器221的指令,和/或存储在设置于处理器中的存储器的指令,执行电子设备200的各种功能应用以及数据处理。
电子设备200可以通过音频模块270,扬声器270A,受话器270B,麦克风270C,耳机接口270D,以及应用处理器等实现音频功能。例如音乐播放,录音等。
音频模块270用于将数字音频信息转换成模拟音频信号输出,也用于将模拟音频输入转换为数字音频信号。音频模块270还可以用于对音频信号编码和解码。在一些实施例中,音频模块270可以设置于处理器210中,或将音频模块270的部分功能模块设置于处理器210中。
扬声器270A,也称“喇叭”,用于将音频电信号转换为声音信号。电子设备200可以通过扬声器270A收听音乐,或收听免提通话。
受话器270B,也称“听筒”,用于将音频电信号转换成声音信号。当电子设备200接听电话或语音信息时,可以通过将受话器270B靠近人耳接听语音。
麦克风270C,也称“话筒”,“传声器”,用于将声音信号转换为电信号。当拨打电话或发送语音信息时,用户可以通过人嘴靠近麦克风270C发声,将声音信号输入到麦克风270C。电子设备200可以设置至少一个麦克风270C。在另一些实施例中,电子设备200可以设置两个麦克风270C,除了采集声音信号,还可以实现降噪功能。在 另一些实施例中,电子设备200还可以设置三个,四个或更多麦克风270C,实现采集声音信号,降噪,还可以识别声音来源,实现定向录音功能等。
耳机接口270D用于连接有线耳机。耳机接口270D可以是USB接口230,也可以是3.5mm的开放移动电子设备平台(open mobile terminal platform,OMTP)标准接口,美国蜂窝电信工业协会(cellular telecommunications industry association of the USA,CTIA)标准接口。
压力传感器280A用于感受压力信号,可以将压力信号转换成电信号。在一些实施例中,压力传感器280A可以设置于显示屏294。压力传感器280A的种类很多,如电阻式压力传感器,电感式压力传感器,电容式压力传感器等。电容式压力传感器可以是包括至少两个具有导电材料的平行板。当有力作用于压力传感器280A,电极之间的电容改变。电子设备200根据电容的变化确定压力的强度。当有触摸操作作用于显示屏294,电子设备200根据压力传感器280A检测所述触摸操作强度。电子设备200也可以根据压力传感器280A的检测信号计算触摸的位置。在一些实施例中,作用于相同触摸位置,但不同触摸操作强度的触摸操作,可以对应不同的操作指令。例如:当有触摸操作强度小于第一压力阈值的触摸操作作用于短消息应用图标时,执行查看短消息的指令。当有触摸操作强度大于或等于第一压力阈值的触摸操作作用于短消息应用图标时,执行新建短消息的指令。
按键290包括开机键,音量键等。按键290可以是机械按键。也可以是触摸式按键。电子设备200可以接收按键输入,产生与电子设备200的用户设置以及功能控制有关的键信号输入。
马达291可以产生振动提示。马达291可以用于来电振动提示,也可以用于触摸振动反馈。例如,作用于不同应用(例如拍照,音频播放等)的触摸操作,可以对应不同的振动反馈效果。作用于显示屏294不同区域的触摸操作,马达291也可对应不同的振动反馈效果。不同的应用场景(例如:时间提醒,接收信息,闹钟,游戏等)也可以对应不同的振动反馈效果。触摸振动反馈效果还可以支持自定义。
指示器292可以是指示灯,可以用于指示充电状态,电量变化,也可以用于指示消息,未接来电,通知等。
SIM卡接口295用于连接SIM卡。SIM卡可以通过插入SIM卡接口295,或从SIM卡接口295拔出,实现和电子设备200的接触和分离。电子设备200可以支持1个或N个SIM卡接口,N为大于1的正整数。SIM卡接口295可以支持Nano SIM卡,Micro SIM卡,SIM卡等。同一个SIM卡接口295可以同时插入多张卡。所述多张卡的类型可以相同,也可以不同。SIM卡接口295也可以兼容不同类型的SIM卡。SIM卡接口295也可以兼容外部存储卡。电子设备200通过SIM卡和网络交互,实现通话以及数据通信等功能。在一些实施例中,电子设备200采用eSIM,即:嵌入式SIM卡。eSIM卡可以嵌在电子设备200中,不能和电子设备200分离。
本申请实施例中的图文融合方法均可以在具有上述硬件结构的电子设备或者具有类似结构的电子设备中实现。例如,该电子设备可以是手机、平板电脑、个人计算机或上网本等。
在本申请实施例中,电子设备将文本排版至第一图像后,得到的图文融合图像(例 如第二图像或第三图像),可以直接在该电子设备的显示屏上展示,也可以作为其他用途,本申请对此不作限定。同样的,若该图文融合方法应用于其他类型的电子设备,例如,服务器类电子设备,该服务器类电子设备得到的图文融合图像可以推送至终端类电子设备的显示屏上展示,也可以作为其他用途。
请参考以下示例,以下几种示例为本申请实施例得到的图文融合图像的几种可能的应用场景示例:
示例1:图文融合图像作为电子设备壁纸。例如,锁屏壁纸、主界面壁纸或聊天界面背景等。
如图3中的(a)所示,为图文融合图像作为手机的锁屏壁纸的示例。如图3中的(b)所示,为图文融合图像作为手机的主界面壁纸的示例。如图3中的(c)所示,为图文融合图像作为手机的微信聊天界面背景的示例。
示例2:图文融合图像作为应用启动页或引导页。例如,应用启动时的界面。
可以理解的是,启动页,也可以称为闪屏页。启动页的设计可以有效利用应用初始化过程中的空白界面,增强用户对应用能够快速启动并立即投入使用的感知度,进而增强应用启动时的用户体验。例如,可以在启动页进行品牌展现、广告、活动等展示,展示方式可以为静态图片、动态图片、动画等多种方式。由于应用的初始化时间一般不会超过5秒,因此启动页的控制时长也通常不超过5秒。如图4中的(a)所示,为图文融合图像作为APP启动页的示例,该APP的启动页控制时长为3秒(seconds,s)。
可以理解的是,引导页是用于引导用户学习应用用法或了解应用作用的页面,其核心在于“引导”二字。引导页一般会出现在全新概念的应用上,或是产品的迭代之后。如图4中的(b)所示,为图文融合图像应用于APP引导页的示例。其中,该APP的引导页由3张引导图片构成(如图4中的(c)所示),手机响应于用户在触摸屏的向左/向右的滑动操作,可以切换引导图片。
示例3:图文融合图像应用于应用界面中。例如,视频播放窗口的悬浮广告框。
例如,当手机检测到视频播放暂停按钮被点击时,手机可以在视频播放窗口显示悬浮广告框停的界面上,图5所示。
示例4:图文融合图像展示在传统图像传播媒介上。例如,展示在报纸、杂志、电视或户外广告位等。
对于该示例,可参考常规的传统传播媒介上的图像展示,本申请实施例这里不予赘述。
需要说明的是,上述示例1~示例4仅作为几种图文融合图像(如第二图像或第三图像)可能的应用场景示例。该图文融合图像还可以应用于其他场景,本申请实施例对此不作限定。
以下以本申请实施例的图文融合方法应用于具有图2所示硬件结构的手机为例,对本申请实施例提供的图文融合方法进行具体阐述。
可以理解的,本申请实施例中,手机可以执行本申请实施例中的部分或全部步骤,这些步骤或操作仅是示例,本申请实施例还可以执行其它操作或者各种操作的变形。此外,各个步骤可以按照本申请实施例呈现的不同的顺序来执行,并且有可能并非要 执行本申请实施例中的全部操作。
如图6所示,本申请实施例的图文融合方法可以包括S601-S605:
S601、手机获取第一图像和待排版至第一图像的第一文本。
其中,第一图像为待添加文本的图像。第一文本为待排版至第一图像的文字。该第一文本可以与第一图像的图像内容相对应。
可以理解的是,第一文本与第一图像的图像内容相对应,是指第一文本可以用来对第一图像中的图像信息进行解释、说明;或者第一文本与第一图像中的图像信息描述的主体一致;或者第一文本与第一图像中的图像信息所传达的意境相通。如图6所示,第一图像为包括月亮图像特征和埃菲尔铁塔图像特征的图像,第一文本可以是主题为“皎皎白玉盘”的文本,该文本的内容包括:“被埃菲尔铁塔遮挡了一半的满月,就像犹抱琵琶半遮面的羞涩女孩儿。@壹刻传媒”。可知,图6中主题为“皎皎白玉盘”的文本是对图6中的第一图像中的图像信息的解释和说明。
在一些实施例中,第一图像与第一文本可以是手机从第三方获取的。例如,手机周期性地从手机厂商的服务器获取第一图像以及第一文本。
或者,第一图像可以是手机本地拍摄的图片,第一文本可以是手机接收用户自定义的文字。例如,第一图像是手机接收用户的图片选择操作,确定的图像;第一文本是手机接收到的用户输入的文字。本申请实施例对第一图像和第一文本数的具体来源不作限定。
在一些实施例中,第一图像可以仅对应一套文本数据。例如,图1中包括建筑物特征的图像仅对应主题为“乡野一瞥”的文本。在这种情况下,手机获取的第一文本即该主题为“乡野一瞥”的文本。
在另一些实施例中,第一图像可以对应多套文本数据。例如,图7中的(a)和图7中的(b)对应的第一图像是相同的,但是对应的文本数据却不同。
当第一图像对应多套文本数据时,在手机获取第一文本时,具体是获取该多套文本数据中的哪一套。可以根据第一图像与每一套文本数据的匹配度的排名确定,也可以是随机确定的,还可以是按照某一次序确定。对此,本申请实施例不作限定。
例如,在制作如图4中的(c)所示的引导页时,手机可以先获取第一图像和主题为“懂你”的文本数据,然后采用本申请实施例的图文融合方法将主题为“懂你”的文本数据排版至第一图像,得到第一张引导页。然后依次获取主题为“时尚”和“信赖”的文本数据,并依次采用本申请实施例的图文融合方法得到第二张引导页、第三张引导页、第四张引导页和第五引导页等。
S602、手机确定第一图像中各个像素点的特征值。
其中,一个像素点的特征值用于表征该像素点被用户关注的可能性的高低。该像素点的特征值越高,则该像素点被用户关注的可能性越高。
在一些实施例中,手机确定第一图像中各个像素点的特征值,可以包括:手机通过对第一图像进行特征检测,确定第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数;然后,手机对确定出的第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数 和文本特征参数中的至少两个参数进行加权求和,确定第一图像中各个像素点的特征值。
其中,特征检测用于识别图像中的图像特征。例如,识别图像中的人脸特征、人体特征、建筑物特征、事物特征(如动物特征、树木特征、花朵特征等)、文字特征、河流特征和山川特征等。
或者,手机可以在对第一图像进行特征检测后,分别获取第一图像的第一特征图、第二特征图、第三特征图和第四特征图中的至少两个特征图。其中,第一特征图中各个像素点的像素值为对应像素点的视觉显著参数,第二特征图中各个像素点的像素值为对应像素点的人脸特征参数,第三特征图中各个像素点的像素值为对应像素点的边缘特征参数,第四特征图中各个像素点的像素值为对应像素点的文本特征参数。
对应的,手机可以对第一特征图、第二特征图、第三特征图和第四特征图中的至少两个特征图进行加权求和,得到第一图像的特征图。其中,第一图像的特征图中,各个像素点的像素值表征该像素点被用户关注的可能性的高低。
可以理解的是,手机对第一特征图、第二特征图、第三特征图和第四特征图进行加权求和,得到第一图像的特征图,具体可以包括:手机对第一特征图、第二特征图、第三特征图和第四特征图中的至少两个特征图中各个像素点的像素值进行加权求和,确定第一图像中各个像素点的特征值;然后,根据第一图像中各个像素点的特征值得到第一图像的特征图。
示例性的,在本申请实施例中,手机可以采用视觉显著性检测算法对第一图像进行视觉显著性特征检测,确定第一图像中各个像素点的视觉显著参数。可以理解的是,视觉显著性检测算法的原理是通过计算图像中像素之间的差异来确定视觉显著性特征。手机通过视觉显著性检测算法可以获取人类视觉系统关注的图像特征。如图8中的(a)所示,为一种第一图像示例。手机采用视觉显著性区域检测算法对第一图像进行视觉显著性特征检测,得到第一特征图(如图8中的(b)所示)。又如,如图9中的(a)所示,为一种第一图形示例。手机采用视觉显著性区域检测算法对第一图像进行视觉显著性特征检测,得到第一特征图(如图9中的(b)所示)。其中,图8中的(b)和图9中的(b)中,每个像素的亮度用于标识该像素点的图像特征的特征值,亮度越大则代表该像素点的图像特征特征值越大。
其中,本申请实施例中,视觉显著性检测算法具体可以为基于多特征的吸收马尔科夫链的视觉显著性检测算法,基于全局颜色对比的视觉显著性检测算法,将角点凸包和贝叶斯推断相结合的视觉显著性检测算法,或者可以为基于深度学习(例如基于卷积神经网络)的视觉显著性检测算法等,本申请实施例对此不作限定。
示例性的,如图10所示,为本申请实施例提供的一种基于多特征的吸收马尔科夫链的视觉显著性检测算法流程图。如图10所示,该算法主要分为两步:第一步是提取图像的超像素及其特征,第二步是基于提取的超像素及特征建立马尔科夫链。其中,在第一步,可以采用简单线性迭代分割(Simple linear iterative clustering, SLIC)算法将第一图像分割为若干超像素。然后,将超像素中的所有像素点的颜色拟合成CIELab三维正态分布,并加入每个像素点在4个方向(0°,45°,90°和135°)的方向值拟合得到四维正态分布特征。在第二步,由于人类在观察图像时是以区域为基本单位的。借鉴这种视觉特性,基于多特征的吸收马尔科夫链的视觉显著性检测算法可以以超像素为基本处理单位。对图像所有超像素建立联系,然后使用最终的平稳分布作为第一特征图。
其中,在图10所示的获取的第一特征图中,可以理解的是,第一特征图中相对亮的像素点对应的是第一图像的显著特征,越亮则代表该像素被用户视觉关注的可能性越高。
示例性的,在本申请实施例中,手机可以采用人脸检测算法对第一图像进行人脸特征检测,确定第一图像中各个像素点的人脸特征参数。可以理解的是,人脸检测算法用于检测图像中的人脸特征。如图9所示,手机采用人脸检测算法对图9中的(a)所示的第一图像进行人脸检测,得到第二特征图(如图9中的(c)所示)。其中,如图9中的(c)中,亮度较大的像素点分布区域对应的是第一图像中的人脸区域。又如,如图8所示,手机采用人脸检测算法对图8中的(a)所示的第一图像进行人脸检测,发现该图像中无人脸,因此,得到的第二特征图(如图8中的(c)所示)中无人脸特征。
其中,本申请实施例中,人脸检测算法具体可以为计算机视觉人脸检测算法(例如,基于方向梯度直方图(Histogram of oriented Gradients,HOG)的人脸检测算法),或者可以为基于深度学习(例如基于卷积神经网络)的人脸检测算法等,本申请实施例对此不作限定。
示例性的,如图11所示,为本申请实施例提供的一种基于HOG的人脸检测算法流程图。如图11所示,在获取第一图像后,可以先对第一图像进行图像归一化。其中,图像归一化是指对图像进行一系列标准的处理变换,使该图像变换为一固定标准形式的过程,该标准图像称作归一化图像。该归一化图像对平移、旋转、缩放等仿射变换具有不变特性。然后采用一阶微分计算归一化图像的图像梯度。接着,对归一化图像进行HOG块划分,对划分后的HOG结构中的每一个单元格绘制梯度直方图,然后对每一个单元格的梯度直方图进行规定权重的投影。然后,对每个单元格块的特征向量进行归一化处理,使得每个单元格块的特征向量空间对光照,阴影和边缘变化具有具有不变特性。最后,将所有HOG块的直方图向量组合成一个HOG特征向量,得到第一图像的HOG特征向量,即为第二特征图。
示例性的,在本申请实施例中,手机可以采用边缘检测算法对第一图像进行边缘特征检测,确定第一图像中各个像素点的边缘特征参数。可以理解的是,边缘检测算法可以用于提取图像中纹理复杂的区域(如,草地、树林、山、和水波等)提的边缘梯度特征。如图8所示,手机采用边缘检测算法对图8中的(a)所示的第一图像进行边缘检测,得到第三特征图(如图8中的(d)所示)。该第三特征图用于标识第一图像中埃菲尔铁塔的结构纹理。又如,如图9所示,手机采用边缘检测算法对图9中的(a)所示的第一图像进行边缘检测,得到的第三特征 图(如图9中的(d)所示)。该第三特征图用于标识第一图像中的人体轮廓。
其中,本申请实施例中,边缘检测算法具体可以为拉普拉斯算子、Canny算法、Prewitt算子或者索贝尔(Sobel)算子等,本申请实施例对此不作限定。
示例性的,如图12所示,为本申请实施例提供的一种基于Canny算法的边缘检测算法流程图。Canny算法是一种先平滑再求导的方法。如图12所示,采用Canny算法可以先使用高斯滤波器对第一图像进行平滑处理。其中,平滑处理是指对图像进行减噪(例如,抑制图像噪声、抑制干扰高频等),使得图像亮度平缓渐变,减小突变梯度,改善图像质量。在获取无噪声第一图像后,可以采用一阶偏导的有限差分计算无噪声第一图像中边缘的梯度幅值和方向。然后,对无噪声第一图像中边缘的梯度幅值进行非极大值抑制。最后,采用双阈值算法检测和连接边缘,得到第三特征图。
示例性的,在本申请实施例中,手机可以采用文本检测算法对第一图像进行文本特征检测,确定第一图像中各个像素点的文本特征参数。可以理解的是,文本检测算法可以用于提取图像(例如日历、节气相关的图像,广告、海报等图像)中存在的文字或者字符特征。如图8所示,手机采用文本检测算法对图8中的(a)所示的第一图像进行文本检测,得到第四特征图(如图8中的(e)所示)。又如,如图9所示,手机采用文本检测算法对图9中的(a)所示的第一图像进行文本检测,得到第四特征图(如图9中的(e)所示)。由于图8中的(a)与图9中的(a)所示的第一图像示例中均不包含文字,因此得到的图8中的(e)和图9中的(e)中均无文本特征。
其中,本申请实施例中,文本检测算法具体可以为基于计算机梯度以及膨胀与腐蚀操作的文本检测算法,或者可以为基于深度学习(例如基于卷积神经网络)的文本检测算法等,本申请实施例对此不作限定。
示例性的,如图13所示,为本申请实施例提供的一种基于计算机梯度以及膨胀与腐蚀操作的文本检测算法流程图。如图13所示,基于计算机梯度以及膨胀与腐蚀操作的文本检测算法可以先对第一图像灰度化处理,用于降低计算量。然后计算灰度图的梯度特征,并将梯度特征二值化,然后使用图像膨胀与腐蚀操作对二值化的梯度特征进行处理,若有图像区域的梯度特征满足给定阈值,可认为该区域为文本特征区域。采用上述方法可以得到文本特征图。
需要说明的是,本申请实施例中如图10所示的基于多特征的吸收马尔科夫链的视觉显著性检测算法,如图11所示的基于HOG的人脸检测算法,如图12所示的基于Canny算法的边缘检测算法以及如图13所示的基于计算机梯度以及膨胀与腐蚀操作的文本检测算法,仅作为几种计算示例。本申请实施例对具体的视觉显著性检测算法、人脸检测算法、边缘检测算法和文本检测算法不作限定。
另外,本申请实施例对于图10所示的基于多特征的吸收马尔科夫链的视觉显著性检测算法,如图11所示的基于HOG的人脸检测算法,如图12所示的基于Canny算法的边缘检测算法以及如图13所示的基于计算机梯度以及膨胀与腐蚀操作的文本检测算法中的具体处理细节也不做限定,上述算法及过程还可以有其他变形,关于上述算法的具体细节及处理过程等,可以参考常规技术中的细节及处 理过程等,本申请实例这里不予赘述。
在本申请实施例中,示例性的,若手机确定出了第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数,手机对第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数进行加权求和,确定第一图像中各个像素点的特征值,可以通过以下公式实现:
F c(x,y)=α*F sal(x,y)+β*F face(x,y)+γ*F edge(x,y)+η*F text(x,y)。
其中,F c(x,y)是第一图像中像素点(x,y)的特征值,F sal(x,y)是第一图像中像素点(x,y)的视觉显著参数,F face(x,y)是第一图像中像素点(x,y)的人脸特征参数,F edge(x,y)是第一图像中像素点(x,y)的边缘特征参数,F text(x,y)是第一图像中像素点(x,y)的文本特征参数。α、β、γ和η分别为视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数对应的权重参数。
在本申请实施例中,α、β、γ和η的具体值可以视不同参数的重要程度而定。示例性,视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数的重要程度排名可以为:人脸特征参数>文本特征参数>视觉显著参数>边缘特征参数。在这种情况下,示例性的,α、β、γ和η可以分别设置为0.2、0.4、0.1和0.3。或者,视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数的重要程度排名可以为:人脸特征参数=文本特征参数>视觉显著参数>边缘特征参数。在这种情况下,示例性的,α、β、γ和η可以分别设置为0.2、0.4、0和0.4。其中,γ为0也可以理解为在确定第一图像中各个像素点的特征值时,不考虑边缘特征参数。
在本申请实施例中,每一个像素点的F c(x,y),对应到第一图像的特征图(如图8中的(f)和图9中的(f))中。其像素点(x,y)的明暗程度,代表该像素点的F c(x,y)值大小。亮度越大则代表该像素点(x,y)的F c(x,y)值越大。
S603、手机根据第一文本,以及第一图像中各个像素点的特征值,确定第一文本在第一图像的多个第一排版版式。
可以理解的是,第一排版版式即上文中所述的候选排版版式。其中,第一排版版式至少用于表征第一文本的文本模板和第一文本排版至第一图像的位置。
其中,文本模板至少可以规定以下中的一种或多种:文本主题和文本正文的行间距、行宽、字号、字体、文字粗细、对齐方式、排版形式,以及装饰线位置和装饰线粗细程度等。其中,排版形式至少可以包括竖版和横版等。
在一些实施例中,文本模板可以以对象简谱(JavaScript Object Notation,JSON)格式、可扩展标记语言(Extensible Markup Language,XML)格式、代码片段或者其他文本格式保存,本申请实施例对此不作限定。
如图14所示,为本申请实施例提供的一种JSON格式的排版规格示例图。
在一些实施例中,本申请实施例中的多个文本模板可以是预先设计好的,保存在手机中,或者存储在手机厂商的服务器中。或者,对于第二图像的其他应用场景实施例,文本模板还可以存储在其他对应位置,本申请实施例对此不作限定。
在本申请实施例中,第一文本排版至第一图像的位置用于表征第一文本按照第一排版版式排版至第一图像时,第一文本位于第一图像的相对位置。该第一文 本排版至第一图像的位置至少可以包括:左上、右上、居中、左下、右下、顶部居中和底部居中等中的任一种。示例性的,如图15所示,为本申请实施例提供的一种第一文本排版至第一图像的位置示意图。如图15所示,第一文本排版至第一图像的位置为“左下”。
在一些实施例中,第一文本排版至第一图像的位置可以为默认排版位置。其中,默认排版位置为预设排版位置,默认排版位置可以为左上、右上、居中、左下、右下、顶部居中和底部居中等排版位置中的任一种。
在本申请实施例中,手机可以根据第一文本以某一文本模板排版时,该第一文本的文本框的大小,结合第一图像中各个像素点的特征值(或者第一图像的特征图),确定第一文本排版至第一图像的位置,使得该第一文本不遮挡第一图像中特征值大于第一阈值的像素点。
可以理解的是,第一图像中特征值大于第一阈值的像素点对应的图像特征为上文中的显著特征,即被用户关注的可能性较大的图像特征。因此,手机确定的第一文本排版至第一图像的位置,可以使得第一文本不遮挡第一图像中的显著特征。
由于手机以某一文本模板排版时,符合上述条件的排版位置可能会有多个。另外第一文本在以不同的文本模板排版时,该第一文本的文本框的大小不同。如图16中的(a)、图16中的(b)和图16中的(c)所示,示出了本申请实施例提供的几种第一文本以不同文本模板排版的示例图。因此第一文本以不同文本模板排版排版至第一图像的位置也可能不同。因此,手机可以确定出第一文本排版至第一图像的多个排版位置。对应的,手机便确定了多个第二排版版式。
示例性的,以图8中的(f)所示的第一图像的特征图为例,介绍手机确定的多个候选排版版式(即多个第一排版版式)。如图17所示,手机根据图8中的(a)所示的第一图像,结合待融合的第一文本,可以确定出3个候选排版版式:候选排版版式1(如图17中的(a)所示)、候选排版版式2(如图17中的(b)所示)和候选排版版式3(如图17中的(c)所示)。
在如图17所示的情况下,手机需要进一步从候选排版版式1、候选排版版式2和候选排版版式3中确定出第一文本的最优排版版式,即手机执行S604。
S604、手机根据多个第一排版版式的代价参数,从该多个第一排版版式中确定出第二排版版式。
其中,一个第一排版版式的代价参数用于表征第一文本按照该排版版式在第一图像排版时,第一文本遮挡的像素点的特征值的大小,以及排版有第一文本的第一图像中各个区域的像素点的特征值分布的平衡程度。
在一些实施例中,手机可以将多个第一排版版式中,能够使得第一文本以某一第一排版版式排版至第一图像时,第一文本遮挡的像素点的特征值最小,以及排版有第一文本的第一图像中各个区域的像素点的特征值分布最均衡的第一排版版式为第二排版版式。
S605、手机按照第二排版版式将第一文本排版至第一图像,得到第二图像。
在一些实施例中,如图18所示,手机根据多个第一排版版式的代价参数,从 该多个第一排版版式中确定出第二排版版式(即S604),可以包括S1801-S1806:
S1801、手机确定第一文本按照一个第一排版版式排版至第一图像时,第一文本的文本框遮挡第一图像的图像区域。
其中,排版版式i为手机确定出的任一个第一排版版式。
S1802、手机计算上述图像区域的纹理特征参数。
其中,上述图像区域的纹理特征参数用于用于表征该图像区域对应的图像中纹理特征的多少。
在一种可能的实现方式中,手机可以先获取该图像区域对应的灰度图像。然后计算该灰度图像的纹理特征参数。
其中,灰度图像是指经过灰度化处理获得的图像。该处理用于降低计算量。灰度化处理后的图像可以呈现出白→灰→黑的分布。灰度值为0的像素点显示为白色,灰度值为255的像素点显示为黑色。
示例性的,手机可以基于灰度共生矩阵的角二阶矩、对比度、熵等,或者图像区域对应的图像灰度与梯度的方差计算图像区域的纹理特征参数。或者,可以参考上文中边缘特征检测算法计算,以及其他常规的边缘特征检测算法或纹理特征检测算法,这里不作赘述。
S1803、手机判断第一文本遮挡的图像区域对应的图像的纹理特征参数是否大于第二阈值。若第一文本遮挡的图像区域对应的图像的纹理特征参数小于第二阈值,手机执行S1804。若第一文本遮挡的图像区域对应的图像的纹理特征参数大于第二阈值,手机舍弃排版版式i,令i+1,重新执行S1801,直至针对每一个第一排版版式执行完成S1801-S1803。
其中,第二阈值为预设阈值。示例性的,第二阈值为125。可以理解的是,对于第一文本遮挡的图像区域对应的图像的纹理特征参数大于125的第一排版版式,可以认为该图像区域的纹理特征过于复杂,若将第一文本排版至此处,可能会同时影响纹理特征和第一文本的展示。在这种情况下,手机可以放弃采用该第一排版版式排版第一文本。
S1804、手机计算第一文本按照排版版式i排版至第一图像时,第一文本的文本入侵参数、视觉空间占用参数和视觉平衡参数中的至少两种。
其中,第一文本的文本入侵参数是第一参数与第二参数的比值。第一参数是第一文本遮挡第一图像的图像区域中各个像素点的特征值之和。第二参数是图像区域的面积,或者,第二参数是图像区域中像素点的总数,或者,第二参数是图像区域中像素点的总数与预设数值的乘积。
示例性的,可以根据公式
Figure PCTCN2020106900-appb-000001
计算第一文本按照排版版式i排版至第一图像时,第一文本的文本入侵参数E s(L i)。其中,R(L)为第一文本以排版版式i排版至第一图像时,第一文本遮挡第一图像的图像区域。x,y分别为第一文本遮挡第一图像的图像区域内,像素点的横坐标和竖坐标。F c(x,y)为像素点(x,y)的特征值。
第一文本的视觉空间占用参数用于表征图像区域中特征值小于第三阈值的像素点的比例。
示例性的,可以根据公式
Figure PCTCN2020106900-appb-000002
计算第一文本按照排版版式i排版至第一图像时,第一文本的视觉空间占用参数E u(L i)。其中,I m(xy)为第一文本以排版版式i排版至第一图像后,第一文本遮挡第一图像的图像区域内像素点的像素值,t为第一图像中的最大特征值。
第一文本的视觉平衡参数用于表征第一文本对排版有该第一文本的第一图像中各个区域的像素点的特征值分布的平衡程度的影响程度。
示例性的,可以根据公式E n(L i)=b1+b2+b3计算第一文本按照排版版式i排版至第一图像时,第一文本的视觉平衡参数E n(L i)。其中b1为第一文本以排版版式i排版至第一图像时,第一文本遮挡第一图像的图像区域的特征图的图像重心与第一图像的特征图的图像重心之间的距离。b2为第一文本以排版版式i排版至第一图像时,第一文本遮挡第一图像的图像区域的图像中心与第一图像中心之间的距离。b3为第一文本以排版版式i排版至第一图像时,第一文本遮挡第一图像的图像区域的图像中心与第一文本遮挡第一图像的图像区域的特征图的图像重心之间的距离。
其中,图像重心(X c,Y c)可以根据公式
Figure PCTCN2020106900-appb-000003
和公式
Figure PCTCN2020106900-appb-000004
计算得到。示例性的,以图17中的(c)所示的候选排版版式为例,P1为第一文本以排版版式i排版至第一图像时,第一文本遮挡第一图像的图像区域的特征图的图像重心。P2为第一图像的特征图的图像重心。P3为第一文本以排版版式i排版至第一图像时,第一文本遮挡第一图像的图像区域的图像中心。P4为第一图像的图像中心。可以得到对于图17中的(c)所示的候选排版版式,b1为P1与P2的距离,b2为P1与P4的距离,b3为P1与P3的距离。同样的方法,可以计算得到每一个候选排版版式对应的b1、b2和b3,进而得到每一个候选排版版式的视觉平衡参数E n(L i)。
S1805、手机根据计算出的第一文本的文本入侵参数、视觉空间占用参数和视觉平衡参数中的至少两个,计算多个第一排版版式的代价参数。
在一些实施例中,手机可以采用以下公式1、公式2或公式3中的任一个计算排版版式i的代价参数。
公式1:T i=λ 1*E s(L i)+λ 2*E u(L i)+λ 3*E n(L i)。
公式2:T i=(λ 1*E s(L i)+λ 2*E u(L i))*E n(L i)。
公式3:T i=E s(L i)*E u(L i)*E n(L i)。
其中,λ 1、λ 2和λ 3分别为所述E s(L i)、所述E u(L i)和所述E n(L i)对应的权重参数。用于标识所述E s(L i)、所述E u(L i)和所述E n(L i)的相对重要程度。因此,λ 1、λ 2和λ 3的取值范围可以为0~1中的任一数值。其中,某一权重参数为0,也可以理解为在确定 第一排版版式的代价参数时,不考虑该权重参数对应的参数。例如,λ 1为0,则可以理解为在确定第一排版版式的代价参数时,不考虑第一文本的文本入侵参数。
需要注意的是,本申请实施例中,手机针对每个第一排版版式,采用相同的代价参数计算公式(如上述公式1,公式2或者公式3)计算其代价参数。例如,手机可以采用上述公式1,计算每个第一排版版式的代价参数。
在本申请实施例中,在手机计算得到每一个第一排版版式的视觉代价T i之后,手机执行S1806。
S1806、手机确定多个第一排版版式的代价参数中,最小的代价参数对应的第一排版版式为第二排版版式。
在一些实施例中,在S1801之前,手机还可以执行:
S1807、手机判断排版版式i是否为默认排版版式(例如:底部居中)。
若排版版式i为默认排版版式,手机可以直接执行S1804。若排版版式i不是默认排版版式,手机可以继续执行S1801。如图19所示。
在一些实施例中,本申请实施例的图文融合方法还可以包括:手机确定第一文本的颜色参数。
在一些实施例中,手机可以在S605之前,确定第一文本的颜色参数。在这种情况下,在S605之后,手机可以根据第一文本的颜色参数,对第二图像中的第一文本着色,获得第三图像。或者,S605还可以为:手机按照第二排版版式,以及第一文本的颜色参数,将第一文本排版至所第一图像,得到第二图像。
在另一些实施例中,手机可以在S605之后,确定第一文本的颜色参数。在这种情况下,在手机确定第一文本的颜色参数之后,手机可以根据第一文本的颜色参数,对第二图像中的第一文本着色,获得第三图像。
其中,第一文本的颜色参数用于对第一文本中每个文字进行着色。
在一些实施例中,手机可以先根据第一文本按照第二排版版式排版至第一图像时,第一图像被第一文本遮挡的图像区域的三原色光RGB的色调、饱和度和明度,确定该图像区域的主色。然后,将该主色的HSV衍生颜色作为第一文本的颜色参数。
其中,RGB色彩模式是工业界的一种颜色标准,是通过对红(R)、绿(G)、蓝(B)三个颜色通道的变化以及它们相互之间的叠加来得到各式各样的颜色。HSV是根据颜色的直观特性创建的一种颜色空间,也称六角锥体模型(Hexcone Model)。该六角锥体模型中颜色的参数至少可以包括:色调(H),饱和度(S)和明度(V)。
在本申请实施例中,图像区域的主色为所述图像区域对应的图像中色相占比最高的色相。该图像区域的主色可以通过统计色相直方图确定。具体的色相直方图统计技术,可以参考常规统计技术,本申请实施例不作赘述。
在本申请实施例中,主色的衍生颜色是指与主色色相相同,但是色调,饱和度和明度与主色的HSV不同的颜色。
在一些实施例中,第一文本按照第二排版版式排版至第二图像后,对于由于第一文本遮挡的图像区域的纹理过于复杂,或者色调过于复杂等原因导致的第一文本显示模糊,或者显示不够突出的问题。如图20所示,本申请实施例的图文融 合方法还可以包括:
S2001、手机判断是否需要对第二图像进行渲染处理。
或者,对于上文中的第三图像,手机判断是否需要对第三图像进行渲染处理。其中,渲染处理至少可以包括蒙版渲染和投影渲染中的至少一种。
在一些实施例中,手机可以根据是否满足以下条件中的至少一种判断是否需要对第二图像或第三图像进行渲染处理:
条件1:第一文本按照第二排版版式排版至第一图像后,第一文本遮挡第一图像的图像区域的纹理特征参数大于第四阈值。
其中,所述纹理特征参数用于表征该图像区域对应的图像中纹理特征的多少。对于纹理特征参数的计算方法,可以参考上文中的边缘特征检测算法,或者其他常规的边缘特征检测算法或纹理特征检测算法,这里不作赘述。
条件2:第一文本按照第二排版版式排版至第一图像后,第一文本遮挡第一图像的图像区域的主色占比小于第五阈值。
其中,第一图像的主色占比可以通过统计第一图像的色相直方图确定。具体的色相直方图统计技术,可以参考常规统计技术,本申请实施例不作赘述。
可以理解的是,若第一文本遮挡第一图像的图像区域的纹理特征参数大于预设值,表示该图像区域的纹理较复杂,可能会影响文字的突出性。以及若第一文本遮挡第一图像的图像区域的主色占比小于预设阈值,表示该图像区域的色调复杂,也可能也会影响文字的突出性。在这种情况下,手机可以对第二图像进行渲染处理,突出第二图像中第一文本。
若手机满足条件1和条件2中的至少一种,手机执行S2002。
S2002、手机对第二图像进行蒙版渲染或者投影渲染。
其中,手机对第二图像进行蒙版渲染具体是指手机对第二图像中第一文本的文字框区域进行蒙版渲染。手机对第二图像进行投影渲染具体是指手机对第二图像中第一文本的文字添加文字阴影。
在一些实施例中,手机对第二图像进行蒙版渲染可以包括:手机在第二图像上覆盖蒙版图层。具体的,手机可以在第二图像中第一文本的文字框区域上覆盖蒙版图层。
在本申请实施例中,蒙版图层的生成方法可以包括:在第二图像中第一文本的文字框的大小基础上,分别向上、下、左、右中的至少一个方向扩展尺寸H1,确定蒙版图层的尺寸。然后,以透明度阈值1→透明度阈值2→透明度阈值3的顺序进行透明度处理,获得蒙版图层。其中,可以以从上到下、从左到右、从下到上或者从右到左等渐变方向进行透明度处理。具体的渐变方向,可以视具体的排版版式而定。例如,第二排版版式为顶部居中,则可以以从上到下的渐变方向进行透明度处理。本申请实施例对此不作限定。
在另一些实施例中,手机对第二图像进行蒙版渲染可以包括:手机确定蒙版参数,然后手机根据确定的蒙版参数处理第二图像。
在本申请实施例中,蒙版参数至少可以包括蒙版尺寸和蒙版透明度参数。其中,蒙版尺寸可以根据以下方法确定:在第二图像中第一文本的文字框的大小基 础上,分别向上、下、左、右中的至少一个方向扩展尺寸H1,即为蒙版尺寸。
在本申请实施例中,蒙版透明度参数可以根据以下方法确定:以透明度阈值1→透明度阈值2→透明度阈值3的顺序确定蒙版透明度参数。其中,可以以从上到下、从左到右、从下到上或者从右到左等渐变方向确定蒙版透明度参数。具体的渐变方向,可以视具体的排版版式而定。例如,第二排版版式为顶部居中,则可以以从上到下的渐变方向确定蒙版透明度参数。本申请实施例对此不作限定。
通过蒙版处理的方法,可以保证图文融合图像中文字的突出性,保证文字清晰可读。如图21所示,为本申请实施例提供的几种图文融合图像对比图。例如,图21中的(b1)采用了本申请实施例的排版方法,其相比于图21中的(a1)所采用现有排版方法排版的图像,文本框设置位置更加科学、文字的突出性更强。又如,图21中的(b2)采用了本申请实施例的排版方法,其相比于图21中的(a2)所采用现有排版方法排版的图像,图文颜色的冲突性更小,文字的突出性更强。
在一些实施例中,手机对第二图像进行投影渲染可以包括:手机确定文字投影参数,然后手机根据确定的文字投影参数处理第二图像。
在本申请实施例中,文字投影参数至少可以包括投影颜色、投影位移和投影模糊数值等。其中,投影颜色可以与文字颜色一致。投影位移可以为预设的位移参数。投影模糊数值可以是预设的模糊数值,示例性的,投影模糊数值可以随着不同位移位置渐变。对于文字投影参数的具体确定方法和方式,可以参考常规的投影渲染技术,本申请实施例不作限定。
可以理解的是,电子设备为了实现上述任一个实施例的功能,其包含了执行各个功能相应的硬件结构和/或软件模块。本领域技术人员应该很容易意识到,结合本文中所公开的实施例描述的各示例的单元及算法步骤,本申请能够以硬件或硬件和计算机软件的结合形式来实现。某个功能究竟以硬件还是计算机软件驱动硬件的方式来执行,取决于技术方案的特定应用和设计约束条件。专业技术人员可以对每个特定的应用来使用不同方法来实现所描述的功能,但是这种实现不应认为超出本申请的范围。
本申请实施例可以对电子设备进行功能模块的划分,例如,可以对应各个功能划分各个功能模块,也可以将两个或两个以上的功能集成在一个处理模块中。上述集成的模块既可以采用硬件的形式实现,也可以采用软件功能模块的形式实现。需要说明的是,本申请实施例中对模块的划分是示意性的,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式。
比如,以采用集成的方式划分各个功能模块的情况下,如图22所示,为本申请实施例提供的一种电子设备的结构示意图。该电子设备可以包括信息获取单元2210、分析单元2220和处理单元2230。
其中,信息获取单元2210可以用于支持电子设备执行上述步骤S601,以及获取多个文本模板,和/或用于本文所描述的技术的其他过程;分析单元2220可以用于支持电子设备执行上述步骤S602、S603、S604和S2001;或采集第一数据,和/或用于本文所描述的技术的其他过程;处理单元2230用于支持电子设备执行上述步骤S605和S2002,和/或用于本文所描述的技术的其他过程。
需要说明的是,上述方法实施例涉及的各步骤的所有相关内容均可以援引到对应 功能模块的功能描述,在此不再赘述。
需要说明的是,上述电子设备还可以包括射频电路。具体的,电子设备可以通过射频电路进行无线信号的接收和发送。通常,射频电路包括但不限于天线、至少一个放大器、收发信机、耦合器、低噪声放大器、双工器等。此外,射频电路还可以通过无线通信和其他设备通信。所述无线通信可以使用任一通信标准或协议,包括但不限于全球移动通讯系统、通用分组无线服务、码分多址、宽带码分多址、长期演进、电子邮件、短消息服务等。
在一种可选的方式中,当使用软件实现数据传输时,可以全部或部分地以计算机程序产品的形式实现。所述计算机程序产品包括一个或多个计算机指令。在计算机上加载和执行所述计算机程序指令时,全部或部分地实现本申请实施例所述的流程或功能。所述计算机可以是通用计算机、专用计算机、计算机网络、或者其他可编程装置。所述计算机指令可以存储在计算机可读存储介质中,或者从一个计算机可读存储介质向另一个计算机可读存储介质传输,例如,所述计算机指令可以从一个网站站点、计算机、服务器或数据中心通过有线(例如同轴电缆、光纤、数字用户线(DSL))或无线(例如红外、无线、微波等)方式向另一个网站站点、计算机、服务器或数据中心进行传输。所述计算机可读存储介质可以是计算机能够存取的任何可用介质或者是包含一个或多个可用介质集成的服务器、数据中心等数据存储设备。所述可用介质可以是磁性介质,(例如软盘、硬盘、磁带)、光介质(例如DVD)、或者半导体介质(例如固态硬盘Solid State Disk(SSD))等。
结合本申请实施例所描述的方法或者算法的步骤可以硬件的方式来实现,也可以是由处理器执行软件指令的方式来实现。软件指令可以由相应的软件模块组成,软件模块可以被存放于RAM存储器、闪存、ROM存储器、EPROM存储器、EEPROM存储器、寄存器、硬盘、移动硬盘、CD-ROM或者本领域熟知的任何其它形式的存储介质中。一种示例性的存储介质耦合至处理器,从而使处理器能够从该存储介质读取信息,且可向该存储介质写入信息。当然,存储介质也可以是处理器的组成部分。处理器和存储介质可以位于ASIC中。另外,该ASIC可以位于探测装置中。当然,处理器和存储介质也可以作为分立组件存在于探测装置中。
通过以上的实施方式的描述,所属领域的技术人员可以清楚地了解到,为描述的方便和简洁,仅以上述各功能模块的划分进行举例说明,实际应用中,可以根据需要而将上述功能分配由不同的功能模块完成,即将装置的内部结构划分成不同的功能模块,以完成以上描述的全部或者部分功能。
在本申请所提供的几个实施例中,应该理解到,所揭露的用户设备和方法,可以通过其它的方式实现。例如,以上所描述的装置实施例仅是示意性的,例如,所述模块或单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如多个单元或组件可以结合或者可以集成到另一个装置,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口,装置或单元的间接耦合或通信连接,可以是电性,机械或其它的形式。
所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是一个物理单元或多个物理单元,即可以位于一个地方,或者也可以分 布到多个不同地方。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本申请各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。上述集成的单元既可以采用硬件的形式实现,也可以采用软件功能单元的形式实现。
所述集成的单元如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个可读取存储介质中。基于这样的理解,本申请实施例的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的全部或部分可以以软件产品的形式体现出来,该软件产品存储在一个存储介质中,包括若干指令用以使得一个设备(可以是单片机,芯片等)或处理器(processor)执行本申请各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(Read-Only Memory,ROM)、随机存取存储器(Random Access Memory,RAM)、磁碟或者光盘等各种可以存储程序代码的介质。
以上所述,仅为本申请的具体实施方式,但本申请的保护范围并不局限于此,任何在本申请揭露的技术范围内的变化或替换,都应涵盖在本申请的保护范围之内。因此,本申请的保护范围应以所述权利要求的保护范围为准。

Claims (30)

  1. 一种图文融合方法,其特征在于,所述方法包括:
    获取第一图像和待排版至所述第一图像的第一文本;
    确定所述第一图像中各个像素点的特征值;其中,一个像素点的特征值用于表征所述一个像素点被用户关注的可能性的高低,所述像素点的特征值越高,所述像素点被用户关注的可能性越高;
    根据所述第一文本,以及所述第一图像中各个像素点的特征值,确定所述第一文本在所述第一图像的多个第一排版版式;其中,所述第一文本按照每个第一排版版式排版至所述第一图像时,所述第一文本不遮挡特征值大于第一阈值的像素点;
    根据所述多个第一排版版式的代价参数,从所述多个第一排版版式中确定出第二排版版式;其中,第一排版版式的代价参数用于表征所述第一文本按照所述第一排版版式排版至所述第一图像时,所述第一文本遮挡的像素点的特征值的大小,以及排版有所述第一文本的第一图像中各个区域的像素点的特征值分布的平衡程度;
    按照所述第二排版版式将所述第一文本排版至所述第一图像,得到第二图像。
  2. 根据权利要求1所述的方法,其特征在于,所述确定所述第一图像中各个像素点的特征值,包括:
    确定所述第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数;其中,一个像素点的视觉显著参数用于表征所述一个像素点是视觉显著性特征对应的像素点的可能性的高低,所述一个像素点的人脸特征参数用于表征所述一个像素点是人脸对应的像素点的可能性的高低,所述一个像素点的边缘特征参数用于表征所述一个像素点是物体轮廓对应的像素点的可能性的高低,所述一个像素点的文本特征参数用于表征所述一个像素点是文本对应的像素点的可能性的高低;
    分别对确定出的所述第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数进行加权求和,确定所述第一图像中各个像素点的特征值。
  3. 根据权利要求2所述的方法,其特征在于,在所述分别对确定出的所述第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数进行加权求和,确定所述第一图像中各个像素点的特征值之前,所述方法还包括:
    分别根据确定出的所述第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数生成至少两个特征图;每一个所述特征图中各个像素点的像素值为对应像素点的对应参数;
    所述分别对确定出的所述第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数进行加权求和,确定所述第一图像中各个像素点的特征值,包括:
    对所述至少两个特征图中各个像素点的像素值进行加权求和,确定所述第一 图像中各个像素点的特征值。
  4. 根据权利要求1-3中任一项所述的方法,其特征在于,所述根据所述第一文本,以及所述第一图像中各个像素点的特征值,确定所述第一文本在所述第一图像的多个第一排版版式,包括:
    根据所述第一文本以一个或多个文本模板排版时所述第一文本的文本框的大小,以及所述第一图像中各个像素点的特征值,确定所述多个第一排版版式。
  5. 根据权利要求4所述的方法,其特征在于,所述方法还包括:
    获取所述一个或多个文本模板,每个所述文本模板规定了文本的行间距、行宽、字号、字体、文字粗细、对齐方式、装饰线位置和装饰线粗细中的至少一种。
  6. 根据权利要求1-5中任一项所述的方法,其特征在于,所述根据所述多个第一排版版式的代价参数,从所述多个第一排版版式中确定出第二排版版式,包括:
    确定所述第一文本分别按照所述多个第一排版版式排版至所述第一图像时,所述第一文本的文本框遮挡所述第一图像的图像区域的纹理特征参数,所述纹理特征参数用于表征所述图像区域对应的图像中纹理特征的多少;
    从所述多个第一排版版式中,选择出纹理特征参数小于第二阈值的图像区域对应的多个第一排版版式;
    根据选择出的每个第一排版版式的代价参数,从所述选择出的多个第一排版版式中确定出所述第二排版版式。
  7. 根据权利要求1-6中任一项所述的方法,其特征在于,所述方法还包括:
    针对所述多个第一排版版式中每个第一排版版式,执行步骤a、步骤b和步骤c中的至少两个以及步骤d,以得到所述每个第一排版版式的代价参数;
    步骤a:计算所述第一文本按照一个第一排版版式排版至所述第一图像时,所述第一文本的文本入侵参数;所述文本入侵参数是第一参数与第二参数的比值;所述第一参数是所述第一文本遮挡所述第一图像的图像区域中各个像素点的特征值之和;所述第二参数是所述图像区域的面积,或者,所述第二参数是所述图像区域中像素点的总数,或者,所述第二参数是所述图像区域中像素点的总数与预设数值的乘积;
    步骤b、计算所述第一文本按照一个第一排版版式排版至所述第一图像时,所述第一文本的视觉空间占用参数;所述视觉空间占用参数用于表征所述图像区域中特征值小于第三阈值的像素点的比例;
    步骤c:计算所述第一文本按照一个第一排版版式排版至所述第一图像时,所述第一文本的视觉平衡参数;所述视觉平衡参数用于表征所述第一文本对排版有所述第一文本的第一图像中各个区域的像素点的特征值分布的平衡程度的影响程度;
    步骤d、根据计算出的所述第一文本的文本入侵参数、视觉空间占用参数和视觉平衡参数中的至少两个,计算所述一个第一排版版式的代价参数。
  8. 根据权利要求7所述的方法,其特征在于,所述根据计算出的所述第一文本的文本入侵参数、视觉空间占用参数和视觉平衡参数中的至少两个,计算所述一个第一排版版式的代价参数,包括:
    采用
    T i=λ 1*E s(L i)+λ 2*E u(L i)+λ 3*E n(L i),或者,
    T i=(λ 1*E s(L i)+λ 2*E u(L i))*E n(L i),或者,
    T i=E s(L i)*E u(L i)*E n(L i)
    计算所述一个第一排版版式的代价参数T i;
    其中,E s(L i)为所述第一文本按照所述一个第一排版版式排版至所述第一图像时,所述第一文本的文本入侵参数,E u(L i)为所述第一文本按照所述一个第一排版版式排版至所述第一图像时,所述第一文本的视觉空间占用参数,E n(L i)为所述第一文本按照所述一个第一排版版式排版至所述第一图像时,所述第一文本的视觉平衡参数;λ 1、λ 2和λ 3分别为所述E s(L i)、所述E u(L i)和所述E n(L i)对应的权重参数。
  9. 根据权利要求8所述的方法,其特征在于,所述根据所述多个第一排版版式的代价参数,从所述多个第一排版版式中确定出第二排版版式,包括:
    确定所述多个第一排版版式的代价参数中,最小的代价参数对应的第一排版版式为第二排版版式。
  10. 根据权利要求1-9中任一项所述的方法,其特征在于,所述方法还包括:
    确定第一文本的颜色参数;所述第一文本的颜色参数为所述第一文本按照所述第二排版版式排版至所述第一图像时,所述第一图像被所述第一文本遮挡的图像区域的主色的衍生颜色;所述主色的衍生颜色是指与主色色相相同,但是色调,饱和度和明度与主色的HSV不同的颜色;
    根据所述第一文本的颜色参数,对所述第二图像中的所述第一文本着色,获得第三图像。
  11. 根据权利要求10所述的方法,其特征在于,
    所述第一图像被所述第一文本遮挡的图像区域的主色是基于所述第一文本按照所述第二排版版式排版至所述第一图像时,所述第一图像被所述第一文本遮挡的图像区域的三原色光RGB在HSV空间中的色调、饱和度和明度确定的;所述主色为所述图像区域中占比最高的色相。
  12. 根据权利要求1-9任一项所述的方法,其特征在于,所述方法还包括:
    若满足以下条件1和条件2中的至少一个,确定对所述第二图像进行渲染处理;
    条件1:所述第一文本按照所述第二排版版式排版至所述第一图像后,所述第一文本遮挡所述第一图像的图像区域的纹理特征参数大于第四阈值;所述纹理特征参数用于表征所述图像区域对应的图像中纹理特征的多少;
    条件2:所述图像区域的主色占比小于第五阈值;
    在所述第二图像上覆盖蒙版图层;或者,确定蒙版参数,根据确定的所述蒙版参数处理所述第二图像;或者,对所述第一文本进行投影渲染。
  13. 根据权利要求10或11所述的方法,其特征在于,所述方法还包括:
    若满足以下条件1和条件2中的至少一个,确定对所述第三图像进行渲染处理;
    条件1:所述第一文本按照所述第二排版版式排版至所述第一图像后,所述第 一文本遮挡所述第一图像的图像区域的纹理特征参数大于第四阈值;所述纹理特征参数用于表征所述图像区域对应的图像中纹理特征的多少;
    条件2:所述图像区域的主色占比小于第五阈值;
    在所述第三图像上覆盖蒙版图层;或者,确定蒙版参数,根据确定的所述蒙版参数处理所述第三图像;或者,对所述第一文本进行投影渲染。
  14. 一种图文融合装置,其特征在于,所述装置包括:
    信息获取单元,用于获取第一图像和待排版至所述第一图像的第一文本;
    分析单元,用于确定所述第一图像中各个像素点的特征值;其中,一个像素点的特征值用于表征所述一个像素点被用户关注的可能性的高低,所述像素点的特征值越高,所述像素点被用户关注的可能性越高;以及,
    根据所述第一文本,以及所述第一图像中各个像素点的特征值,确定所述第一文本在所述第一图像的多个第一排版版式;其中,所述第一文本按照每个第一排版版式排版至所述第一图像时,所述第一文本不遮挡特征值大于第一阈值的像素点;以及,
    根据所述多个第一排版版式的代价参数,从所述多个第一排版版式中确定出第二排版版式;其中,一个第一排版版式的代价参数用于表征所述第一文本按照所述第一排版版式排版至所述第一图像时,所述第一文本遮挡的像素点的特征值的大小,以及排版有所述第一文本的第一图像中各个区域的像素点的特征值分布的平衡程度;
    处理单元,用于按照所述第二排版版式将所述第一文本排版至所述第一图像,得到第二图像。
  15. 根据权利要求14所述的装置,其特征在于,所述分析单元确定所述第一图像中各个像素点的特征值,包括:
    所述分析单元确定所述第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数;其中,一个像素点的视觉显著参数用于表征所述一个像素点是视觉显著性特征对应的像素点的可能性的高低,所述一个像素点的人脸特征参数用于表征所述一个像素点是人脸对应的像素点的可能性的高低,所述一个像素点的边缘特征参数用于表征所述一个像素点是物体轮廓对应的像素点的可能性的高低,所述一个像素点的文本特征参数用于表征所述一个像素点是文本对应的像素点的可能性的高低;
    所述分析单元分别对确定出的所述第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数进行加权求和,确定所述第一图像中各个像素点的特征值。
  16. 根据权利要求15所述的装置,其特征在于,在所述分析单元分别对确定出的所述第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数进行加权求和,确定所述第一图像中各个像素点的特征值之前,所述分析单元还用于:
    分别根据确定出的所述第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数生成至少两个特征图;每一 个所述特征图中各个像素点的像素值为对应像素点的对应参数;
    所述分别对确定出的所述第一图像中各个像素点的视觉显著参数,人脸特征参数,边缘特征参数和文本特征参数中的至少两个参数进行加权求和,确定所述第一图像中各个像素点的特征值,包括:
    对所述至少两个特征图中各个像素点的像素值进行加权求和,确定所述第一图像中各个像素点的特征值。
  17. 根据权利要求14-16任一项所述的装置,其特征在于,所述分析单元根据所述第一文本,以及所述第一图像中各个像素点的特征值,确定所述第一文本在所述第一图像的多个第一排版版式,包括:
    所述分析单元根据所述第一文本以所述一个或多个文本模板排版时所述第一文本的文本框的大小,以及所述第一图像中各个像素点的特征值,确定所述多个第一排版版式。
  18. 根据权利要求17所述的装置,其特征在于,所述信息获取单元还用于:
    获取所述一个或多个文本模板,每个所述文本模板规定了文本的行间距、行宽、字号、字体、文字粗细、对齐方式、装饰线位置和装饰线粗细中的至少一种。
  19. 根据权利要求14-18任一项所述的装置,其特征在于,所述分析单元根据所述多个第一排版版式的代价参数,从所述多个第一排版版式中确定出第二排版版式,包括:
    所述分析单元确定所述第一文本分别按照所述多个第一排版版式排版至所述第一图像时,所述第一文本的文本框遮挡所述第一图像的图像区域的纹理特征参数,所述纹理特征参数用于表征所述图像区域对应的图像中纹理特征的多少;
    所述分析单元从所述多个第一排版版式中,选择出纹理特征参数小于第二阈值的图像区域对应的多个第一排版版式;
    所述分析单元根据选择出的每个第一排版版式的代价参数,从所述选择出的多个第一排版版式中确定出所述第二排版版式。
  20. 根据权利要求14-19任一项所述的装置,其特征在于,所述分析单元还用于:
    针对所述多个第一排版版式中每个第一排版版式,执行步骤a、步骤b和步骤c中的至少两个以及步骤d,以得到所述每个第一排版版式的代价参数;
    步骤a:计算所述第一文本按照一个第一排版版式排版至所述第一图像时,所述第一文本的文本入侵参数;所述文本入侵参数是第一参数与第二参数的比值;所述第一参数是所述第一文本遮挡所述第一图像的图像区域中各个像素点的特征值之和;所述第二参数是所述图像区域的面积,或者,所述第二参数是所述图像区域中像素点的总数,或者,所述第二参数是所述图像区域中像素点的总数与预设数值的乘积;
    步骤b、计算所述第一文本按照一个第一排版版式排版至所述第一图像时,所述第一文本的视觉空间占用参数;所述视觉空间占用参数用于表征所述图像区域中特征值小于第三阈值的像素点的比例;
    步骤c:计算所述第一文本按照一个第一排版版式排版至所述第一图像时,所述第一文本的视觉平衡参数;所述视觉平衡参数用于表征所述第一文本对排版有 所述第一文本的第一图像中各个区域的像素点的特征值分布的平衡程度的影响程度;
    步骤d、根据计算出的所述第一文本的文本入侵参数、视觉空间占用参数和视觉平衡参数中的至少两个,计算所述一个第一排版版式的代价参数。
  21. 根据权利要求20所述的装置,其特征在于,所述分析单元根据计算出的所述第一文本的文本入侵参数、视觉空间占用参数和视觉平衡参数中的至少两个,计算所述一个第一排版版式的代价参数,包括:
    所述分析单元采用
    T i=λ 1*E s(L i)+λ 2*E u(L i)+λ 3*E n(L i),或者,
    T i=(λ 1*E s(L i)+λ 2*E u(L i))*E n(L i),或者,
    T i=E s(L i)*E u(L i)*E n(L i)
    计算所述一个第一排版版式的代价参数T i;
    其中,E s(L i)为所述第一文本按照所述一个第一排版版式排版至所述第一图像时,所述第一文本的文本入侵参数,E u(L i)为所述第一文本按照所述一个第一排版版式排版至所述第一图像时,所述第一文本的视觉空间占用参数,E n(L i)为所述第一文本按照所述一个第一排版版式排版至所述第一图像时,所述第一文本的视觉平衡参数;λ 1、λ 2和λ 3分别为所述E s(L i)、所述E u(L i)和所述E n(L i)对应的权重参数。
  22. 根据权利要求21所述的装置,其特征在于,所述分析单元根据所述多个第一排版版式的代价参数,从所述多个第一排版版式中确定出第二排版版式,包括:
    所述分析单元确定所述多个第一排版版式的代价参数中,最小的代价参数对应的第一排版版式为第二排版版式。
  23. 根据权利要求14-22任一项所述的装置,其特征在于,所述处理单元还用于:
    确定第一文本的颜色参数;所述第一文本的颜色参数为所述第一文本按照所述第二排版版式排版至所述第一图像时,所述第一图像被所述第一文本遮挡的图像区域的主色的衍生颜色;所述主色的衍生颜色是指与主色色相相同,但是色调,饱和度和明度与主色的HSV不同的颜色;以及
    根据所述第一文本的颜色参数,对所述第二图像中的所述第一文本着色,获得第三图像。
  24. 根据权利要求23所述的装置,其特征在于,
    所述第一图像被所述第一文本遮挡的图像区域的主色是基于所述第一文本按照所述第二排版版式排版至所述第一图像时,所述第一图像被所述第一文本遮挡的图像区域三原色光RGB在HSV空间中的色调、饱和度和明度确定的;所述主色为所述图像区域中色相占比最高的色相。
  25. 根据权利要求14-22任一项所述的装置,其特征在于,所述分析单元还用于:
    若满足以下条件1和条件2中的至少一个,确定对所述第二图像进行渲染处理;
    条件1:所述第一文本按照所述第二排版版式排版至所述第一图像后,所述第一文本遮挡所述第一图像的图像区域的纹理特征参数大于第四阈值;所述纹理特征参数用于表征所述图像区域对应的图像中纹理特征的多少;
    条件2:所述图像区域的主色占比小于第五阈值;
    所述处理单元还用于,在所述第二图像上覆盖蒙版图层;或者,确定蒙版参数,根据确定的所述蒙版参数处理所述第二图像;或者,对所述第一文本进行投影渲染。
  26. 根据权利要求23或24所述的装置,其特征在于,所述分析单元还用于:
    若满足以下条件1和条件2中的至少一个,确定对所述第三图像进行渲染处理;
    条件1:所述第一文本按照所述第二排版版式排版至所述第一图像后,所述第一文本遮挡所述第一图像的图像区域的纹理特征参数大于第四阈值;所述纹理特征参数用于表征所述图像区域对应的图像中纹理特征的多少;
    条件2:所述图像区域的主色占比小于第五阈值;
    所述处理单元还用于,在所述第三图像上覆盖蒙版图层;或者,确定蒙版参数,根据确定的所述蒙版参数处理所述第三图像;或者,对所述第一文本进行投影渲染。
  27. 一种电子设备,其特征在于,所述电子设备包括:
    存储器,用于存储一个或多个计算机程序;
    处理器,用于执行所述存储器存储的一个或多个计算机程序,使得所述电子设备实现如权利要求1-13任一项所述的图文融合方法。
  28. 一种计算机可读存储介质,其特征在于,所述计算机可读存储介质上存储有计算机执行指令,所述计算机执行指令被处理电路执行时实现如权利要求1-13任一项所述的图文融合方法。
  29. 一种芯片系统,其特征在于,所述芯片系统包括处理器、存储器,所述存储器中存储有指令;所述指令被所述处理器执行时,实现如权利要求1-13任一项所述的图文融合方法。
  30. 一种计算机程序产品,所述计算机程序产品包括程序指令,所述程序指令被执行时,以实现权利要求1-13任一项所述的图文融合方法。
PCT/CN2020/106900 2019-08-23 2020-08-04 一种图文融合方法、装置及电子设备 Ceased WO2021036715A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
EP20858455.7A EP3996046A4 (en) 2019-08-23 2020-08-04 METHOD AND DEVICE FOR IMAGE-TEXT MERGING AND ELECTRONIC DEVICE
US17/634,002 US12254544B2 (en) 2019-08-23 2020-08-04 Image-text fusion method and apparatus, and electronic device

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201910783866.7A CN110706310B (zh) 2019-08-23 2019-08-23 一种图文融合方法、装置及电子设备
CN201910783866.7 2019-08-23

Publications (1)

Publication Number Publication Date
WO2021036715A1 true WO2021036715A1 (zh) 2021-03-04

Family

ID=69193924

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2020/106900 Ceased WO2021036715A1 (zh) 2019-08-23 2020-08-04 一种图文融合方法、装置及电子设备

Country Status (4)

Country Link
US (1) US12254544B2 (zh)
EP (1) EP3996046A4 (zh)
CN (1) CN110706310B (zh)
WO (1) WO2021036715A1 (zh)

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113158875A (zh) * 2021-04-16 2021-07-23 重庆邮电大学 基于多模态交互融合网络的图文情感分析方法及系统
CN113591972A (zh) * 2021-07-28 2021-11-02 北京百度网讯科技有限公司 图像处理方法、装置、电子设备以及存储介质
CN115619696A (zh) * 2022-11-07 2023-01-17 湖南师范大学 一种基于结构相似性与l2范数优化的图像融合方法
CN119334923A (zh) * 2024-12-17 2025-01-21 陕西金翼服装有限责任公司 一种基于紫外线的纺织物关键点高精度检测方法

Families Citing this family (14)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110706310B (zh) 2019-08-23 2021-10-22 华为技术有限公司 一种图文融合方法、装置及电子设备
CN111311554B (zh) * 2020-01-21 2023-09-01 腾讯科技(深圳)有限公司 图文内容的内容质量确定方法、装置、设备及存储介质
WO2021166120A1 (ja) * 2020-02-19 2021-08-26 三菱電機株式会社 情報処理装置、情報処理方法及び情報処理プログラム
CN113362424A (zh) * 2020-03-04 2021-09-07 阿里巴巴集团控股有限公司 图像合成、商品广告图像合成方法、设备及存储介质
CN111612683B (zh) * 2020-04-08 2025-06-17 西安万像电子科技有限公司 数据处理方法及系统
CN111859893B (zh) * 2020-07-30 2021-04-09 广州云从洪荒智能科技有限公司 图文排版方法、装置、设备及介质
CN113989404B (zh) * 2021-11-05 2024-06-25 北京字节跳动网络技术有限公司 图片处理方法、装置、设备、存储介质和程序产品
CN114429637B (zh) * 2022-01-14 2023-04-07 北京百度网讯科技有限公司 一种文档分类方法、装置、设备及存储介质
CN115019330A (zh) * 2022-06-16 2022-09-06 特赞(上海)信息科技有限公司 一种漫画翻译匹配方法、系统、电子设备及存储介质
CN117761075B (zh) * 2023-11-13 2024-07-23 江苏嘉耐高温材料股份有限公司 一种长寿命功能材料的微孔分布形态检测系统、方法
CN118212326B (zh) * 2024-05-21 2024-09-03 腾讯科技(深圳)有限公司 视觉文本生成方法、装置、设备和存储介质
CN119206374B (zh) * 2024-11-22 2025-03-14 宁波科达精工科技股份有限公司 一种汽车卡钳生产流水线的切削加工质量评估方法及系统
CN119625010B (zh) * 2025-02-13 2025-05-09 南京城铁信息技术有限公司 一种低功耗4g无缝线路钢轨视觉位移检测方法及装置
CN121504871A (zh) * 2025-11-17 2026-02-10 江苏新丝路纺织科技有限公司 一种ai视觉的纺织品染色缺陷智能检测方法及系统

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101123002A (zh) * 2007-09-14 2008-02-13 北大方正集团有限公司 一种图文的自动排版方法
US20110173532A1 (en) * 2010-01-13 2011-07-14 George Forman Generating a layout of text line images in a reflow area
CN107103635A (zh) * 2017-03-20 2017-08-29 中国科学院自动化研究所 图像排版配色方法
CN109493399A (zh) * 2018-09-13 2019-03-19 北京大学 一种图文结合的海报生成方法和系统
CN110009712A (zh) * 2019-03-01 2019-07-12 华为技术有限公司 一种图文排版方法及其相关装置
CN110706310A (zh) * 2019-08-23 2020-01-17 华为技术有限公司 一种图文融合方法、装置及电子设备

Family Cites Families (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPWO2013005366A1 (ja) * 2011-07-05 2015-02-23 パナソニック株式会社 アンチエイリアス画像生成装置およびアンチエイリアス画像生成方法
CN102890826B (zh) * 2011-08-12 2015-09-09 北京多看科技有限公司 一种扫描版文档重排版的方法
US9626768B2 (en) * 2014-09-30 2017-04-18 Microsoft Technology Licensing, Llc Optimizing a visual perspective of media
US9454694B2 (en) * 2014-12-23 2016-09-27 Lenovo (Singapore) Pte. Ltd. Displaying and inserting handwriting words over existing typeset
WO2019227300A1 (zh) * 2018-05-29 2019-12-05 优视科技新加坡有限公司 版面元素的处理方法、装置、存储介质及电子设备/终端/服务器
CN109117713B (zh) * 2018-06-27 2021-11-12 淮阴工学院 一种全卷积神经网络的图纸版面分析与文字识别方法
US11189066B1 (en) * 2018-11-13 2021-11-30 Adobe Inc. Systems and methods of learning visual importance for graphic design and data visualization

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101123002A (zh) * 2007-09-14 2008-02-13 北大方正集团有限公司 一种图文的自动排版方法
US20110173532A1 (en) * 2010-01-13 2011-07-14 George Forman Generating a layout of text line images in a reflow area
CN107103635A (zh) * 2017-03-20 2017-08-29 中国科学院自动化研究所 图像排版配色方法
CN109493399A (zh) * 2018-09-13 2019-03-19 北京大学 一种图文结合的海报生成方法和系统
CN110009712A (zh) * 2019-03-01 2019-07-12 华为技术有限公司 一种图文排版方法及其相关装置
CN110706310A (zh) * 2019-08-23 2020-01-17 华为技术有限公司 一种图文融合方法、装置及电子设备

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See also references of EP3996046A4

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113158875A (zh) * 2021-04-16 2021-07-23 重庆邮电大学 基于多模态交互融合网络的图文情感分析方法及系统
CN113158875B (zh) * 2021-04-16 2022-07-01 重庆邮电大学 基于多模态交互融合网络的图文情感分析方法及系统
CN113591972A (zh) * 2021-07-28 2021-11-02 北京百度网讯科技有限公司 图像处理方法、装置、电子设备以及存储介质
CN115619696A (zh) * 2022-11-07 2023-01-17 湖南师范大学 一种基于结构相似性与l2范数优化的图像融合方法
CN119334923A (zh) * 2024-12-17 2025-01-21 陕西金翼服装有限责任公司 一种基于紫外线的纺织物关键点高精度检测方法

Also Published As

Publication number Publication date
CN110706310B (zh) 2021-10-22
US12254544B2 (en) 2025-03-18
EP3996046A1 (en) 2022-05-11
US20220319077A1 (en) 2022-10-06
CN110706310A (zh) 2020-01-17
EP3996046A4 (en) 2022-10-19

Similar Documents

Publication Publication Date Title
US12254544B2 (en) Image-text fusion method and apparatus, and electronic device
CN115658191B (zh) 一种生成主题壁纸的方法及电子设备
CN112712470B (zh) 一种图像增强方法及装置
US12086957B2 (en) Image bloom processing method and apparatus, and storage medium
WO2020192417A1 (zh) 图像渲染方法及装置、电子设备
CN111179282A (zh) 图像处理方法、图像处理装置、存储介质与电子设备
WO2020077511A1 (zh) 一种拍摄场景下的图像显示方法及电子设备
US12008761B2 (en) Image processing method and apparatus, and device
WO2020134877A1 (zh) 一种皮肤检测方法及电子设备
WO2020078026A1 (zh) 一种图像处理方法、装置与设备
CN113096022B (zh) 图像虚化处理方法、装置、存储介质与电子设备
WO2020113534A1 (zh) 一种拍摄长曝光图像的方法和电子设备
WO2024082976A1 (zh) 文本图像的ocr识别方法、电子设备及介质
CN117132477B (zh) 图像处理方法及电子设备
US20250166121A1 (en) Image rendering method and apparatus
CN114529663B (zh) 一种消除阴影的方法和电子设备
WO2024037211A1 (zh) 着色方法、着色装置和电子设备
CN114911546A (zh) 图像显示方法、电子设备及存储介质
CN120339337B (zh) 运动对象检测方法及相关设备
CN118764727B (zh) 图像处理方法及相关设备
CN115701129B (zh) 一种图像处理方法及电子设备
RU2791810C2 (ru) Способ, аппаратура и устройство для обработки и изображения
WO2025200862A1 (zh) 摄像头颜色校准的方法、装置和电子设备
WO2026082099A1 (zh) 图像处理方法、摄像头及电子设备
CN121704798A (zh) 检测方法及电子设备

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20858455

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2020858455

Country of ref document: EP

Effective date: 20220204

NENP Non-entry into the national phase

Ref country code: DE

WWG Wipo information: grant in national office

Ref document number: 17634002

Country of ref document: US