WO2024221925A1 - 数据处理方法、装置、电子设备、计算机可读存储介质及计算机程序产品 - Google Patents

数据处理方法、装置、电子设备、计算机可读存储介质及计算机程序产品 Download PDF

Info

Publication number
WO2024221925A1
WO2024221925A1 PCT/CN2023/135891 CN2023135891W WO2024221925A1 WO 2024221925 A1 WO2024221925 A1 WO 2024221925A1 CN 2023135891 W CN2023135891 W CN 2023135891W WO 2024221925 A1 WO2024221925 A1 WO 2024221925A1
Authority
WO
WIPO (PCT)
Prior art keywords
features
sparse
type
graphics processor
embedded
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2023/135891
Other languages
English (en)
French (fr)
Inventor
弓静
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Tencent Technology (Shenzhen) Co Ltd
Original Assignee
Tencent Technology (Shenzhen) Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Tencent Technology (Shenzhen) Co Ltd filed Critical Tencent Technology (Shenzhen) Co Ltd
Priority to EP23935092.9A priority Critical patent/EP4614327A4/en
Publication of WO2024221925A1 publication Critical patent/WO2024221925A1/zh
Priority to US19/235,945 priority patent/US20250307980A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T1/00General purpose image data processing
    • G06T1/20Processor architectures; Processor configuration, e.g. pipelining
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00Arrangements for program control, e.g. control units
    • G06F9/06Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/46Multiprogramming arrangements
    • G06F9/54Interprogram communication
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00Arrangements for program control, e.g. control units
    • G06F9/06Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/46Multiprogramming arrangements
    • G06F9/50Allocation of resources, e.g. of the central processing unit [CPU]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00Arrangements for program control, e.g. control units
    • G06F9/06Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/46Multiprogramming arrangements
    • G06F9/50Allocation of resources, e.g. of the central processing unit [CPU]
    • G06F9/5005Allocation of resources, e.g. of the central processing unit [CPU] to service a request
    • G06F9/5027Allocation of resources, e.g. of the central processing unit [CPU] to service a request the resource being a machine, e.g. CPUs, Servers, Terminals
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • G06N20/10Machine learning using kernel methods, e.g. support vector machines [SVM]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/044Recurrent networks, e.g. Hopfield networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • G06N3/0455Auto-encoder networks; Encoder-decoder networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0475Generative networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/06Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons
    • G06N3/063Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons using electronic means
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/084Backpropagation, e.g. using gradient descent
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/088Non-supervised learning, e.g. competitive learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/09Supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/098Distributed learning, e.g. federated learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/0985Hyperparameter optimisation; Meta-learning; Learning-to-learn
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N5/00Computing arrangements using knowledge-based models
    • G06N5/01Dynamic search techniques; Heuristics; Dynamic trees; Branch-and-bound
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N7/00Computing arrangements based on specific mathematical models
    • G06N7/01Probabilistic graphical models, e.g. probabilistic networks

Definitions

  • the present application relates to artificial intelligence technology, and in particular to a data processing method, device, electronic device, computer-readable storage medium, and computer program product.
  • Artificial Intelligence is a comprehensive technology in computer science. By studying the design principles and implementation methods of various intelligent machines, machines are given the functions of perception, reasoning and decision-making. Analysis based on medical images and medical texts is one of the important applications in the field of artificial intelligence.
  • the medical analysis system refers to a system that uses computers to process, analyze and understand medical images and medical texts to identify targets and objects of various different patterns.
  • the embedding feature is learned based on the Parameter Server method.
  • the embedding feature representation (embedding) of the sparse feature is stored on the central processing unit (CPU), so the embedding training is also completed in the CPU.
  • the training speed is slow.
  • the embodiments of the present application provide a data processing method, device, electronic device, computer-readable storage medium and computer program product based on artificial intelligence, which can improve the update processing speed through a graphics processor under the premise of realizing large-scale embedded feature storage.
  • the embodiment of the present application provides an artificial intelligence-based data processing method, the data processing method is applied to a first graphics processor, the first graphics processor stores embedded features of a first type of full sparse features, and the second graphics processor stores embedded features of a second type of full sparse features, the first type is different from the second type, the method includes:
  • the first graphics processor acquires sparse features of a target sample, where the sparse features of the target sample include a first sparse feature of the first type and a second sparse feature of the second type;
  • the first graphics processor obtains an embedded feature corresponding to the first sparse feature from the embedded features of the full amount of sparse features of the first type;
  • the first graphics processor obtains, from the second graphics processor, an embedded feature corresponding to the second sparse feature obtained by querying the embedded features of the full amount of sparse features of the second type;
  • the first graphics processor combines the embedded features corresponding to the first sparse features and the embedded features corresponding to the second sparse features into embedded features corresponding to the sparse features of the target sample;
  • the first graphics processor performs probability mapping processing on the embedded features of the sparse features corresponding to the target sample, and generates an update instruction according to the probability mapping result, wherein the update instruction is used to instruct the second graphics processor to update the embedded features of the full amount of sparse features of the second type.
  • the embodiment of the present application provides an artificial intelligence-based data processing device, wherein the data processing method is applied to a first graphics processor, the first graphics processor stores embedded features of a first type of full sparse features, and the second graphics processor stores embedded features of a second type of full sparse features, the first type is different from the second type, and the device includes:
  • a first receiving module is configured to obtain a sparse feature of a target sample for the first graphics processor, where the sparse feature of the target sample includes a first sparse feature of the first type and a second sparse feature of the second type;
  • a first acquisition module is configured to enable the first graphics processor to acquire, from the second graphics processor, an embedded feature corresponding to the second sparse feature obtained by querying the embedded features of the full amount of sparse features of the second type;
  • the first returning module is configured to obtain from the second graphics processor an embedded feature corresponding to the second sparse feature obtained by querying the embedded features of all sparse features of the second type;
  • a first determining module is configured to use the first graphics processor to combine the embedded features corresponding to the first sparse features and the embedded features corresponding to the second sparse features into embedded features corresponding to the sparse features of the target sample;
  • the first update module is configured so that the first graphics processor performs probability mapping processing on the embedded features of the sparse features corresponding to the target sample, and generates an update instruction based on the probability mapping result, wherein the update instruction is used to instruct the second graphics processor to update the embedded features of the full amount of sparse features of the second type.
  • An embodiment of the present application provides a data processing method, the data processing method is applied to a second graphics processor, the first graphics processor stores embedded features of a first type of full sparse features, the second graphics processor stores embedded features of a second type of full sparse features, the first type is different from the second type;
  • the method comprises:
  • the second graphics processor receives the second sparse feature of the second type sent by the first graphics processor
  • the second graphics processor queries the embedded features of all sparse features of the second type to obtain the embedded features corresponding to the second sparse features;
  • the second graphics processor transmits the embedded features corresponding to the second sparse features to the first graphics processor, so that the first graphics processor performs probability mapping processing on the embedded features corresponding to the first sparse features of the first type and the embedded features corresponding to the second sparse features, wherein the embedded features corresponding to the first sparse features are obtained by the first graphics processor from the embedded features of the full amount of sparse features of the first type;
  • the second graphics processor receives an update instruction generated according to the probability mapping result from the first graphics processor, and updates the embedded features of the second type of full-quantity sparse features based on the update instruction.
  • the embodiment of the present application provides a data processing device, wherein the data processing method is applied to a second graphics processor, wherein the first graphics processor stores embedded features of a first type of full amount of sparse features, and the second A second graphics processor stores embedded features of a full amount of sparse features of a second type, wherein the first type is different from the second type;
  • the device comprises:
  • a second receiving module configured to receive, by the second graphics processor, the second sparse features of the second type sent by the first graphics processor
  • a query module configured to query, by the second graphics processor, the embedded features corresponding to the second sparse features from the embedded features of the full amount of sparse features of the second type;
  • a second returning module is configured for the second graphics processor to transmit the embedded features corresponding to the second sparse features to the first graphics processor, so that the first graphics processor performs a probability mapping process on the embedded features corresponding to the first sparse features of the first type and the embedded features corresponding to the second sparse features, wherein the embedded features corresponding to the first sparse features are obtained by the first graphics processor from the embedded features of the full amount of sparse features of the first type;
  • the second update module is configured so that the second graphics processor receives an update instruction generated according to the probability mapping result from the first graphics processor, and updates the embedded features of the second type of full-quantity sparse features based on the update instruction.
  • An embodiment of the present application provides a graphics processor, which is used to execute the data processing method provided in the embodiment of the present application.
  • An embodiment of the present application provides an electronic device, the electronic device comprising:
  • a memory for storing computer executable instructions
  • the graphics processor is used to implement the data processing method provided in the embodiment of the present application when executing the computer executable instructions stored in the memory.
  • An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions.
  • the computer-executable instructions are executed by a processor, the data processing method provided by the embodiment of the present application is implemented.
  • the processor is a central processing unit or a graphics processing unit.
  • An embodiment of the present application provides a computer program product, including computer executable instructions, which, when executed by a processor, implement the data processing method provided in the embodiment of the present application, wherein the processing area is a central processing unit or a graphics processing unit.
  • the first graphics processor obtains different types of sparse features of the target sample.
  • the first graphics processor obtains the first type of embedded features locally, and obtains the second type of embedded features from the second graphics processor.
  • the classification according to type makes it convenient to obtain embedded features from the corresponding graphics processor, which can improve acquisition efficiency and facilitate feature management.
  • the first graphics processor performs probability mapping processing on the two types of embedded features, and transmits update instructions to the second graphics processor according to the probability mapping results.
  • the update instructions are used to instruct the second graphics processor to update the second type of embedded features, thereby realizing the update of embedded features under the premise of distributed storage, making full use of the high efficiency of graphics processor calculations, and greatly improving the update speed of embedded features while meeting storage requirements.
  • FIG1A is a schematic diagram of the structure of an artificial intelligence-based data processing system provided in an embodiment of the present application.
  • FIG1B is an architecture diagram of a data processing system provided in an embodiment of the present application.
  • FIG2 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
  • 3A-3D are schematic flow charts of a data processing method based on artificial intelligence provided in an embodiment of the present application.
  • FIG4 is a schematic diagram of a framework of a recommendation system provided in an embodiment of the present application.
  • FIG5 is a schematic diagram of the physical architecture of the Parameter Server provided in an embodiment of the present application.
  • FIG6 is a schematic diagram of a single working node provided in an embodiment of the present application.
  • FIG7 is a schematic diagram of data flow of an artificial intelligence-based data processing method provided in an embodiment of the present application.
  • FIG8 is a schematic diagram of data flow of an artificial intelligence-based data processing method provided in an embodiment of the present application.
  • first ⁇ second ⁇ third involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that “first ⁇ second ⁇ third” can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
  • Embedding is a method of using a numerical vector to represent an object.
  • the object here can be a word, an item, or a movie, etc.
  • TensorFlow It is a symbolic mathematical system based on data flow programming. It is widely used in the programming implementation of various machine learning algorithms. It has a multi-level structure and can be deployed on various servers, PC terminals and web pages. It supports GPU and TPU high-performance numerical calculations and is widely used in product development and scientific research in various fields.
  • Parameter Server It is a programming framework that facilitates the writing of distributed parallel programs, with a focus on supporting the distributed storage and coordination of large-scale parameters.
  • the nodes in the cluster of the server can be divided into two types: computing nodes and parameter service nodes.
  • the computing nodes are responsible for computing and learning the training data (blocks) assigned to their local area and updating the corresponding parameters;
  • the parameter service nodes use distributed storage to store part of the global parameters and accept parameter query and update requests from computing nodes as service providers.
  • Sparse features are those that do not appear continuously in the data set, and most of the values are zero. For example, if there are 100 account numbers, then for each account number, 99 digits of the sparse features will be 0, and only 1 digit will be 1. They are called sparse features because they have very few non-zero values in the data set.
  • the full sparse features represent the sparse features of all samples. For example, for account types, all samples have a total of 100 accounts, and the 100 sparse features corresponding to these 100 accounts are the full sparse features of the account type. Dense features are generally relative to sparse features, and the proportion of 0 in dense features is small, or even non-existent.
  • Type The type provided in the embodiments of the present application may be a content type or a data format type.
  • the content type is distinguished based on the semantics represented by sparse features
  • the data format type is distinguished based on the data format of sparse features.
  • the embedded features are learned based on the Parameter Server method.
  • the embedded feature representation (embedding) of the sparse features is stored on the central processing unit (CPU), and the embedding training is also completed in the CPU.
  • the training speed will be limited.
  • the applicant found that if the embedding training is deployed in the graphics processing unit (GPU), the training speed can be effectively improved, but the storage capacity of the GPU cannot meet the storage requirements of large-scale embedding.
  • the embodiments of the present application provide an artificial intelligence-based data processing method, device, electronic device, and computer-readable storage medium, which can improve the update processing speed through a graphics processor while achieving large-scale embedded feature storage.
  • the data processing method provided in the embodiments of the present application can be implemented by the terminal/server alone; or it can be implemented by the terminal and the server in collaboration, for example, the terminal or the server alone undertakes the data processing method described below, or the terminal sends a data sample to the server, and the server executes according to the received data sample. Data processing methods.
  • the electronic device for data processing may be various types of terminals or servers, wherein the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services; the terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto.
  • the terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in the present application.
  • the server can be a server cluster deployed in the cloud, opening artificial intelligence cloud services (AI as a Service, AiaaS) to users.
  • AI artificial intelligence cloud services
  • AiaaS artificial intelligence cloud services
  • the platform will split several common AI services and provide independent or packaged services in the cloud.
  • This service model is similar to an AI theme mall. All users can access one or more artificial intelligence services provided by the AIaaS platform through an application programming interface.
  • one of the artificial intelligence cloud services may be a data processing service, that is, a server in the cloud is encapsulated with a data processing program provided by the embodiment of the present application.
  • the user calls the data processing service in the cloud service through the terminal, so that the server deployed in the cloud calls the encapsulated data processing program.
  • FIG. 1A is a schematic diagram of an application scenario of a data processing system provided in an embodiment of the present application.
  • a terminal 400 is connected to a server 200 via a network 300.
  • the network 300 may be a wide area network or a local area network, or a combination of the two.
  • Four GPUs are deployed in the server 200, and GPU1 and GPU2 are used for illustration below.
  • the data sample may be a data sample of a recommendation model.
  • a single data sample may include relevant object data of an account logged into a news client, where the relevant object data includes an account identifier (account ID), account age, attribute tags representing account interests, etc.
  • the embedded features of the sparse features of the first type (e.g., account identifier type) in the full data sample are stored in GPU1 deployed in server 200, and the embedded features of the sparse features of the second type (e.g., age group type) in the full data sample are stored in GPU2 deployed in server 200.
  • the terminal 400 sends a matching request to the server 200, and the server 200 calls the recommendation model.
  • the recommendation model is deployed in any GPU (for example, GPU1), and the input of the recommendation model is the account of the target object. Therefore, GPU1 needs to query the embedded features of the target object's age group type from GPU2 based on the sparse features of the age group type.
  • the output of the recommendation model is the recommendation probability of the target object clicking on a certain news item. When the recommendation probability is greater than the probability threshold, server 200 returns the news item to terminal 400, and the news client pushes the news item to the target object.
  • the following describes the process of updating each type of embedded features.
  • Server 200 receives target sample A, where target sample A is input to GPU1, GPU1 obtains sparse features of target sample A, the sparse features of target sample A include sparse features of target sample A in account identification type and sparse features of target sample A in age group type, GPU1 obtains embedded features of sparse features of target sample A in account identification type from embedded features of full sparse features of account identification type; GPU1 obtains embedded features of sparse features of target sample A in age group type from GPU2, the embedded features of sparse features of target sample A in age group type are obtained by querying embedded features of full sparse features of age group type by GPU2; GPU1 determines embedded features of sparse features of target sample A based on embedded features of sparse features of target sample A in account identification type and embedded features of sparse features of target sample A in age group type; GPU1 forward propagates based on embedded features of sparse features of target sample A, and transmit
  • the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
  • the terminal 400 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto.
  • the terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present invention.
  • the terminal or server can implement the artificial intelligence-based data processing method provided in the embodiments of the present application by running a computer program.
  • the computer program can be a native program or software module in the operating system; it can be a native application (APP, Applic ation), that is, a program that needs to be installed in the operating system to run, such as a live broadcast APP or an instant messaging APP; it can also be a mini-program, that is, a program that can be run by just downloading it to a browser environment; it can also be a mini-program that can be embedded in any APP.
  • the above-mentioned computer program can be any form of application, module or plug-in.
  • FIG. 2 is a schematic diagram of the structure of the electronic device for data processing provided by the embodiment of the present application, and the electronic device is a server 200 as an example.
  • the server 200 for data processing shown in FIG. 2 includes: at least one processor 210 (the processor 210 may be a graphics processor or a central processing unit), a memory 250, and at least one network interface 220.
  • the various components in the server 200 are coupled together through a bus system 240. It is understandable that the bus system 240 is used to realize the connection and communication between these components.
  • the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are marked as bus systems 240 in FIG.
  • the processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., wherein the general-purpose processor can be a microprocessor or any conventional processor, etc.
  • DSP digital signal processor
  • the memory 250 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories.
  • the non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM).
  • the memory 250 described in the embodiment of the present application is intended to include any suitable type of memory.
  • the memory 250 optionally includes one or more storage devices that are physically far away from the processor 210.
  • memory 250 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.
  • the operating system 251 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic businesses and process hardware-based tasks.
  • hardware-related tasks such as a framework layer, a core library layer, a driver layer, etc.
  • the network communication module 252 is used to communicate with the network interface 220 via one or more (wired or wireless) As for other electronic devices, exemplary network interface 220 includes: Bluetooth, Wireless LAN (WiFi), and Universal Serial Bus (USB).
  • exemplary network interface 220 includes: Bluetooth, Wireless LAN (WiFi), and Universal Serial Bus (USB).
  • the data processing device provided in the embodiment of the present application can be implemented in a software manner, for example, it can be a data processing plug-in in the terminal described above, or it can be a data processing service in the server described above. Of course, it is not limited to this.
  • the data processing device provided in the embodiment of the present application can be provided as various software embodiments, including various forms including application programs, software, software modules, scripts or codes.
  • Figure 2 shows a data processing device 255-1 stored in the memory 250, which can be software in the form of programs and plug-ins, such as an image processing plug-in, and includes a series of modules, including a first receiving module 2551, a first acquisition module 2552, a first return module 2553, a first determination module 2554 and a first update module 2555.
  • Figure 2 also shows a data processing device 255-2 stored in the memory 250, which can be software in the form of programs and plug-ins, such as an image processing plug-in, and includes a series of modules, including a second receiving module 2556, a query module 2557, a second return module 2558, and a second update module 2559.
  • the data processing method provided in the embodiments of the present application can be implemented by various types of electronic devices.
  • an electronic device including multiple graphics processors in which a first graphics processor and a second graphics processor are deployed (the first graphics processor and the second graphics processor can also be deployed on different electronic devices), for example, an electronic device including multiple central processing units, in which a first central processing unit and a second central processing unit are deployed, wherein the computing power of the graphics processor is better than that of the central processing unit.
  • the first graphics processor stores embedded features of a first type of full sparse features
  • the second graphics processor stores embedded features of a second type of full sparse features, where the first type is different from the second type.
  • the field of the semantics represented by the sparse feature can be obtained (the field can be a keyword, for example).
  • the content type corresponding to the preset field is used as the content type of the sparse feature.
  • the field of the semantics represented by the sparse feature can be "account", and the preset field can be the account type. Then the field here matches the preset field, and the type of the sparse feature is the account type.
  • the distributed storage method based on the distinction between content types can help to quickly obtain the embedded features of sparse features of different content types, so that in recommendation scenarios that require feature diversity and richness, the embedded features of the corresponding content types can be queried more quickly, improving the efficiency of acquiring embedded features during training in distributed scenarios.
  • the type here can also be a data format type that is distinguished based on the data format of sparse features.
  • the data format type can be a text type, an image type, a voice type, etc.
  • the distributed storage method based on the distinction of data format types can help to quickly obtain the embedded features of sparse features of different data format types, so that the embedded features of the corresponding data format type can be queried more quickly, thereby improving the efficiency of acquiring embedded features during training in distributed scenarios.
  • the sparse features are used to describe the user's data, so the type can be an object type obtained by dividing based on different users. For example, if there are 10 sparse features that are used to describe user A, then the type to which these 10 sparse features belong is object type A. If there are 10 sparse features that are used to describe user B, then the type to which these 10 sparse features belong is object type B, that is, each user corresponds to an object type, and the embedded features of the sparse features belonging to the same object type are stored in the same graphics processor.
  • the types provided by the embodiments of the present application can also be source types obtained by dividing based on the source of the sparse features, etc.
  • the types provided by the embodiments of the present application can also be time period types obtained by dividing based on the acquisition period of the sparse features, which will not be repeated here.
  • the first graphics processor and the second graphics memory store all sparse features of multiple content types, and a total of 100 sparse features are stored.
  • the content type here can be an account type, an age group type, etc.
  • the first type here can be an account type
  • the second type here can be an age group type.
  • the embedded features of the full amount of sparse features of the identity type stored in the first graphics processor represent that the first graphics processor stores the sparse features of the identity type of all samples.
  • the second graphics processor stores the full amount of embedded features of the sparse features of the age group type, representing that the second graphics processor stores the sparse features of the age group type of all samples.
  • the first graphics processor and the second graphics processor belong to multiple graphics processors, and the multiple graphics processors store embedded features of multiple types of full-quantity sparse features, the first type includes at least one type of the multiple types, and the second type includes at least one type of the multiple types.
  • the embedded features of multiple types of full-quantity sparse features can be divided into different graphics processors for storage according to type, so that when querying, it can have clear directionality without having to query from all graphics processors.
  • the graphics processor system includes multiple graphics processors, for example, 4 graphics processors, wherein the first graphics processor and the second graphics processor are 2 of the 4 graphics processors, and 100 embedded features corresponding to the 100 sparse features will be distributedly stored in the 4 graphics processors.
  • the first type here can be an account type and an interest type
  • the second type here can be an age group type and a gender type.
  • the first graphics processor stores the embedded features of all sparse features of the identity type and the embedded features of all sparse features of the interest type
  • the second graphics processor stores the embedded features of all sparse features of the age group type and the embedded features of all sparse features of the gender type.
  • the storage location of the embedded features of the full amount of sparse features of the first type and the storage location of the embedded features of the full amount of sparse features of the second type are divided when the amount of data of the embedded features of the full amount of sparse features of multiple types is greater than a threshold.
  • distributed storage is only performed when the amount of data is greater than the threshold, thereby improving the utilization rate of the storage resources of the graphics processor and avoiding excessive load on the storage resources of a single graphics processor.
  • the embedded features of the full amount of sparse features of multiple types are stored in each graphics processor of the multiple graphics processors.
  • a single graphics processor can be used to store the embedded features of all sparse features when the amount of data is not greater than the threshold, so that each graphics processor does not need to obtain embedded features from other graphics processors in the subsequent training stage, which can improve the training efficiency.
  • the data volume here can be the total number of embedded features.
  • the embedded features of the first type of full sparse features and the embedded features of the second type of full sparse features are stored in a distributed manner in In two graphics processors, when the total number of embedded features of multiple types of full sparse features is not greater than a threshold, the embedded features of multiple types of full sparse features, that is, all embedded features, will be saved in each graphics processor, that is, each graphics processor saves a full amount of embedded features.
  • the embodiment of the present application uses the dynamic embedding data structure of the tfra component for storage of sparse features.
  • the number of sparse feature parameters is large, the number of corresponding embedded features will be large (for example, greater than the threshold), and the GPU single card memory cannot accommodate it.
  • the embeddings corresponding to the sparse features can be divided into the memory of multiple GPU cards by type. When the total number of sparse features is small, it means that the total number of corresponding embedded features is small (for example, not greater than the threshold), and the GPU single card memory can accommodate it.
  • Using the Replica method a full amount of embedded features are placed in the memory of each GPU card.
  • the embedding layer usually maps high-dimensional sparse features to low-dimensional dense vectors.
  • the low-dimensional dense vector obtained here is the embedded feature.
  • the embedded feature here has the same meaning as the corresponding sparse feature, but occupies a smaller storage space and has a lower data type, and then the model is trained end-to-end.
  • the embodiment of the present application uses the component tfra.dynamic_embedding as the storage method of embedding.
  • the component uses tf.lookup.MutableHashTable to save multiple parameters (fullweights), that is, to save embedded features, and can reuse the native optimizer of tensorflow.
  • the mapping process for each sparse feature is as follows: the type number of the sparse feature is divided by a set value to obtain a corresponding remainder.
  • the remainder obtained based on the type number of the sparse feature is the processor identifier of the sparse feature, so that the sparse feature is saved in the graphics processor indicated by the corresponding processor identifier.
  • the embedded features of sparse features can be stored in the corresponding GPUs according to their types, and the storage resources of multiple GPUs can be reasonably allocated. Moreover, compared with the technical solution that each GPU stores the embedded features of all types of sparse features, in the present application, each GPU is only responsible for storing the embedded features of the specific type of sparse features it is assigned. With the embodiment of the present application, the embedded features can be flexibly stored, and a balance can be achieved between the storage resource utilization rate of each graphics processor and the interactive resource occupancy rate between graphics processors.
  • Figure 3A is a flow chart of the artificial intelligence-based data processing method provided in an embodiment of the present application, which will be explained in conjunction with steps 101 to 105 shown in Figure 3A.
  • a first graphics processor obtains sparse features of a target sample, where the sparse features of the target sample include a first sparse feature of a first type and a second sparse feature of a second type.
  • multiple graphics processors can be deployed in the server, and step 101 here can be implemented by the first graphics processor.
  • the server can be a single machine or a server cluster composed of multiple machines, that is, multiple graphics processors can be deployed in one server or in multiple servers.
  • the target sample is input into the first graphics processor, and the first graphics processor obtains sparse features of the target sample from the target sample, such as sparse features representing the account ID and sparse features representing the age group, that is, the sparse features of the target sample here include the first sparse features and the second sparse features, where the first sparse features are the sparse features of the target sample in the full amount of sparse features of the first type described above, and the second sparse features are the sparse features of the target sample in the full amount of sparse features of the second type described above, for example, the target sample is user A, the first type is the account type, and the second type is age.
  • the first type of full sparse features are sparse features of the account type of all users. If the number of users is 100 and each user has one account, then the first type of full sparse features are sparse features of the account type of 100 users.
  • the second type of full sparse features are sparse features of the age group type of all users. There are 50 age groups in total, so the second type of full sparse features are sparse features of 50 age group types.
  • the first sparse feature is the sparse feature of user A in the account type (sparse feature representing an account) and the sparse feature of user A in the age group type (sparse feature representing an age group).
  • the embedded features are pre-stored in the first graphics processor and the second graphics processor, and the embedded features are obtained by embedding and compressing the sparse features, the process of obtaining the sparse features of multiple data samples before storage, that is, the process of obtaining the full amount of sparse features, is described in detail below.
  • multiple object accounts that log in to the recommendation client and object data of each object account are obtained, the object data of the multiple object accounts are used as multiple data samples, feature analysis processing is performed on the object data of the multiple object accounts to obtain object features of the multiple object accounts, and the object features of the multiple object accounts are used as sparse features of the multiple data samples.
  • the recommendation client can be a news client, a video client, a shopping client, etc., and the news client is used as an example for explanation.
  • the object account is an account that has logged into the news client before, and the account is held by the user (hereinafter, the user is uniformly replaced by the object).
  • object A holds object account A
  • object B holds object account B.
  • the attribute data and operation data of object account A can be obtained as the object data of object account A.
  • the object data of object account A can be used as a data sample, and multiple data samples can be composed of the object data of all objects.
  • the above-mentioned feature analysis processing of the object data of multiple object accounts to obtain the object features of the multiple object accounts can be achieved through the following technical solution: performing the following processing for each object account: obtaining at least one of the following from the object data of the object account: account data corresponding to the object account, biological data corresponding to the object account, location data corresponding to the object account, interest data of the object account; performing one-hot encoding processing on the obtained data, and using the obtained encoding result as the object feature of the object account.
  • the attribute data and operation data of object account A can be obtained as the object data of object account A, and the account identifier (such as account ID), biological identifier (such as gender) and location identifier (such as long-term residence) are obtained from the attribute data.
  • the interest identifier is obtained from the operation data.
  • the operation data can be a browsing operation, a collection operation, etc.
  • the interest identifier can be an interest tag of the information being operated. For example, if the information being operated is sports news, the interest tag can be sports.
  • These identifiers are encoded by one-hot encoding. For example, the gender male is encoded as (0, 1), and the gender female is encoded as (1, 0).
  • the encoding result obtained by one-hot encoding can be used as the object feature.
  • the object feature In the storage process, for some objects, even if the objects are different, they have the same gender or the same age. Therefore, for some types of sparse features, the number of sparse features is fixed.
  • sparse features that accurately describe the object can be obtained, and the sparse features can comprehensively characterize the object, so that the embedded features can effectively characterize the object to improve the object characterization capability of the embedded features.
  • a historical target sample previously obtained by the first graphics processor from multiple data samples is obtained; data samples other than the historical target sample from the multiple data samples are used as other data samples; and at least one data sample is randomly obtained from the other data samples as a target sample.
  • a graphics processor system composed of graphics processors will perform multiple trainings.
  • Each training process can be performed on multiple data samples, or on some data samples in multiple data samples. There is no restriction on this.
  • the embodiment of the present application provides multiple graphics processors, so that each graphics processor can share some data samples.
  • this training process will involve 10 data samples, and these 10 data samples can be multiple data samples, or some data samples in multiple data samples.
  • the first graphics processor A will receive 2 data samples out of the 10 data samples.
  • the data samples (historical target samples) received by the first graphics processor in the last training process are obtained. This acquisition process can be performed by the first graphics processor itself.
  • the data samples received in the last training process are data samples A and data samples B. Then, this training process will randomly obtain any number of data samples other than data samples A and data samples B from multiple data samples. Two data samples can be obtained, or other numbers of data samples can be obtained.
  • the embodiments of the present application can avoid a single graphics processor from repeatedly processing the same data samples during multiple training processes, thereby improving the training efficiency of a single graphics processor.
  • multiple data samples are evenly divided based on the number of graphics processors to obtain multiple sample sets, wherein each sample set includes at least one data sample, and the number of sample sets is the same as the number of graphics processors; a one-to-one correspondence between multiple sample sets and multiple graphics processors is determined; and a data sample in a sample set that has a corresponding relationship with a first graphics processor is used as a target sample corresponding to the first graphics processor.
  • the load balancing of each graphics processor can be ensured by even division.
  • the embodiments of the present application can avoid a single graphics processor from repeatedly processing the same data samples during multiple training processes, thereby optimizing the training effect.
  • each graphics processor can also be ensured by average division.
  • multiple data samples include 10 data samples, and the graphics processor system includes 5 graphics processors. Five sample sets can be obtained by average division processing, and each sample set includes 2 data samples.
  • the correspondence between the sample set and the graphics processor can be pre-saved and determined directly by acquisition.
  • the matching degree between the sample set and the graphics processor can be determined, and the matching degree with each sample set is determined for each graphics processor in any random order.
  • the matching degree with the first graphics processor A and each sample set is determined, and then the sample set corresponding to the highest matching degree is used as the sample set having a corresponding relationship with the graphics processor.
  • the matching degree is negatively correlated with the repetition degree of the historical target sample data of the sample set and the first graphics processor.
  • step 102 the first graphics processor obtains an embedded feature corresponding to the first sparse feature from the embedded features of all sparse features of the first type.
  • the first graphics processor obtains the embedded features of the target sample in the first sparse feature of the first type from the embedded features of the full amount of sparse features of the first type.
  • the full amount represents all meanings
  • each sparse feature corresponds to an embedded feature, which is equivalent to that the embedded features of the sparse features of the first type of all samples (including the target sample) are stored in the first graphics processor. Therefore, it is equivalent to querying from the embedded features of the full amount.
  • the first graphics processor queries the embedded features of the first sparse features of the first type of the target sample from the embedded features corresponding to all the sparse features of the first type, which is equivalent to querying the embedded features of the target sample in the first type from all the embedded features of the first type.
  • the embedded features of the first sparse features of the first type of the target sample can be directly obtained locally, thereby improving the efficiency of acquiring embedded features.
  • the first type is the account identification type, where the full amount represents all meanings
  • the first graphics processor stores the embedded features corresponding to all sparse features belonging to the account identification type.
  • 100 embedded features corresponding to the sparse features of 100 account types are stored.
  • the detector queries the embedded features of the target sample from all embedded features of the account identification type, such as the embedded features of object A in the account identification type.
  • the query here is based on the sparse features of the target sample.
  • the sparse features of the target sample (object A) are (0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0).
  • the corresponding embedded features are queried from the hash table using the sparse features as keys.
  • the embedded features here can be the value 9.
  • the embedded features are equivalent to values.
  • the sparse features and embedded features are stored in the hash table as key-value pairs.
  • the first graphics processor obtains, from the second graphics processor, an embedded feature corresponding to the second sparse feature obtained from the embedded feature query of the full amount of sparse features of the second type.
  • the first graphics processor obtains the embedded features of the second sparse features of the second type of the target sample from the second graphics processor, and the embedded features of the second sparse features of the second type of the target sample are obtained by the second graphics processor from the embedded features of the full amount of sparse features of the second type.
  • the first graphics processor obtains from the second graphics processor the embedded features corresponding to the second sparse features obtained from the embedded feature query of the full amount of sparse features of the second type, which can be implemented by steps 1031 to 1033 shown in FIG. 3B .
  • step 1031 the first graphics processor transmits the second sparse features to the second graphics processor.
  • step 1032 the second graphics processor searches for an embedded feature corresponding to the second sparse feature from the embedded features of the full set of sparse features of the second type.
  • Each sparse feature corresponds to an embedded feature, which is equivalent to that the embedded features of the sparse features of the second type of all samples (including the target sample) are stored in the second graphics processor. Therefore, this is equivalent to querying from the full amount of embedded features.
  • the second graphics processor queries the embedded features of the second sparse features corresponding to the target sample from the embedded features corresponding to all the sparse features of the second type, which is equivalent to querying the embedded features of the target sample in the second type from all the embedded features of the second type.
  • step 1033 the first graphics processor receives an embedded feature corresponding to the second sparse feature.
  • the first type is the account identification type
  • the second type is the age group type.
  • the first GPU transmits the sparse features of the age group type of the target sample to the second GPU, and the second GPU queries the age group type of the target sample from all embedded features belonging to the age group type.
  • the second graphics processor transmits the embedded features of the age group type of the target sample to the first graphics processor, and the first graphics processor receives the embedded features of the age group type of the target sample.
  • the embodiment of the present application can realize the exchange of embedded features.
  • the embedded features that are not stored locally in the first graphics processor can be obtained from the second graphics processor, and the embedded features can be shared under the premise of distributed storage, thereby improving the utilization rate of the embedded features.
  • the first graphics processor combines the embedded features corresponding to the first sparse features and the embedded features corresponding to the second sparse features into embedded features corresponding to the sparse features of the target sample.
  • the first graphics processor combines the embedded features of the target sample in the first type of sparse features and the embedded features of the target sample in the second type of sparse features into the embedded features of the sparse features of the target sample.
  • the first graphics processor combines the embedded features of the target sample in the account ID type of sparse features and the embedded features of the target sample in the age group type of sparse features into the embedded features of the sparse features of the target sample as input for subsequent model inference.
  • the first graphics processor performs probability mapping processing on the embedded features of the sparse features corresponding to the target sample, and generates an update instruction according to the probability mapping result, wherein the update instruction is used to instruct the second graphics processor to update the embedded features of the full amount of sparse features of the second type.
  • probability mapping refers to the mapping process of converting embedded features into probabilities.
  • the probability mapping process can be a linear transformation process or a nonlinear transformation process, that is, the probability mapping process is actually a linear transformation or a nonlinear transformation of the embedded features to transform the embedded features into a probability value.
  • probability mapping There are many ways of probability mapping, which are not specifically limited in this application.
  • the probability mapping processing for linear transformation involves weight parameters (the data form of weight parameters is column vectors) and bias parameters (scalars).
  • the weight parameters are dot-multiplied with embedded features (the data form of embedded features is row vectors), and the dot-multiplication result (the dot-multiplication result here is a scalar) is added to the bias parameter.
  • the added value is used as the probability obtained by mapping (probability mapping result).
  • the added value can also be normalized and the normalized value is used as the probability mapping result.
  • the probability mapping process can also be specifically referred to in formula (1):
  • x is the embedded feature
  • -W T is the weight parameter
  • b is the bias parameter.
  • -W T x+b can be directly used as the probability obtained by probability mapping.
  • -W T x+b can also be normalized according to formula (1) to obtain h(x), and h(x) is used as the probability mapping result.
  • the first graphics processor performs probability mapping processing on the embedded features corresponding to the sparse features of the target sample, which can be achieved by executing step 1051 shown in FIG. 3C , and in step 105 , generating update instructions according to the probability mapping results can be achieved by executing steps 1052 to 1053 shown in FIG. 3C .
  • step 1051 the embedded features corresponding to the sparse features of the target sample are used as input to the recommendation model, and the recommendation model performs probability mapping processing on the embedded features corresponding to the sparse features of the target sample to obtain the recommendation probability corresponding to the target sample as the probability mapping result.
  • the recommendation model can be a deep neural network (DNN).
  • DNN deep neural network
  • the recommendation model includes at least a fully connected layer. Through the fully connected layer, probability mapping processing (which can be a linear transformation or a nonlinear transformation) can be performed on the embedded features corresponding to the target sample to obtain the recommendation probability of each target sample.
  • the first graphics processor determines an error between a recommendation probability corresponding to the target sample and a label value corresponding to the target sample, and determines a gradient of an embedded feature of a second type of full sparse feature of the target sample based on the error.
  • the error here is the reason obtained by substituting the recommendation probability and the label value into the loss function.
  • the gradient of the embedded features of the second type of full sparse features of the target sample is calculated by the error.
  • the calculation methods include but are not limited to the following: numerical calculation method (calculation by function derivation) and analytical calculation method.
  • step 1053 the first graphics processor generates an update instruction carrying the gradient of the embedded feature of the full amount of sparse features of the second type.
  • an update instruction without any information is first generated here to instruct the second graphic
  • the processor updates the embedded features of the full sparse features of the second type. Since the update needs to be performed based on the gradient of the embedded features of the full sparse features of the second type, the gradient of the embedded features of the full sparse features of the second type is added to the above-mentioned update instruction that does not carry any information.
  • the first graphics processor transmits an update instruction to the second graphics processor, and the second graphics processor updates the embedded features of the second type of full sparse features based on the gradients of the embedded features of the second type of full sparse features.
  • the second graphics processor can update the locally stored embedded features, which is equivalent to realizing the embedded feature update under distributed storage, breaking through the information isolation barrier caused by distributed storage, and splitting the training process into different graphics processors to complete in a distributed manner, which can effectively improve the training efficiency.
  • the first graphics processor updates the embedded features of the full amount of sparse features of the first type according to the probability mapping result.
  • the probability mapping result is the recommendation probability
  • the first graphics processor determines the error between the recommendation probability and the label value, and determines the gradient of the embedded features of the full amount of sparse features of the first type based on the error, and then updates the embedded features of the full amount of sparse features of the first type according to the gradient of the embedded features of the full amount of sparse features of the first type.
  • the first graphics processor can update the locally stored embedded features. Since the probability mapping result here is obtained by the first graphics processor, the first graphics processor can directly update the locally stored embedded features based on the probability mapping result, thereby improving training efficiency.
  • TensorFlow provides a declarative programming interface. You don't need to worry about the details of the derivation. You only need to define the model to get the loss equation, and then use the various optimizers (Op, Optimizer) implemented by TensorFlow to perform operations to update the new embedded features.
  • Op Optimizer
  • the automatic derivation part is strung together by various Ops.
  • TensorFlow's approach is that each Op includes its gradient calculation formula when building the graph.
  • the calculation graph of the reverse part is automatically established. The input and output of the forward calculation will be retained and deleted when used up in the backward calculation, and then an Op (such as GradientDescentOptimizer, AdamOptimizer) is added at the end.
  • the data processing method provided in the embodiments of the present application is applied to a second graphics processor, the first graphics processor stores embedded features of a first type of full sparse features, and the second graphics processor stores embedded features of a second type of full sparse features, and the first type is different from the second type.
  • step 201 the second graphics processor receives a second sparse feature of a second type sent by the first graphics processor.
  • the second graphics processor receives the second sparse features of the second type of the target sample sent by the first graphics processor.
  • step 202 the second graphics processor queries the embedded features of all sparse features of the second type to obtain the embedded features corresponding to the second sparse features.
  • step 203 the second graphics processor transmits the embedded features corresponding to the second sparse features to the first graphics processor.
  • the first graphics processor performs a probability mapping process on the embedded features corresponding to the first sparse features and the embedded features corresponding to the second sparse features.
  • the embedded feature corresponding to the first sparse feature is obtained by the first graphics processor from the embedded features of all sparse features of the first type according to the first sparse feature of the first type of the target sample.
  • the embedded features corresponding to the first sparse features and the embedded features corresponding to the second sparse features are composed of embedded features corresponding to the sparse features of the target sample, and probability mapping processing is performed on the embedded features corresponding to the sparse features of the target sample to obtain a probability mapping result.
  • the second graphics processor receives an update instruction generated according to the probability mapping result from the first graphics processor, and updates the embedded features of the full amount of sparse features of the second type based on the update instruction.
  • the first GPU and the second GPU belong to a plurality of GPUs, and the plurality of GPUs store embedded features of a plurality of types of full-quantity sparse features, wherein the first type includes a plurality of The first type includes at least one of the first and second types, and the second type includes at least one of the plurality of types.
  • the storage locations of the embedded features of the first type of all sparse features and the storage locations of the embedded features of the second type of all sparse features are divided when the data amount of the embedded features of multiple types of all sparse features is greater than a threshold.
  • the embedded features of the multiple types of full amount sparse features are stored in each of the multiple graphics processors.
  • the first graphics processor obtains different types of sparse features of the target sample.
  • the first graphics processor obtains the first type of embedded features locally, and obtains the second type of embedded features from the second graphics processor.
  • the classification according to type makes it convenient to obtain embedded features from the corresponding graphics processor, which can improve acquisition efficiency and facilitate feature management.
  • the first graphics processor performs probability mapping processing on the two types of embedded features, and transmits update instructions to the second graphics processor according to the probability mapping results.
  • the update instructions are used to instruct the second graphics processor to update the second type of embedded features, thereby realizing the update of embedded features under the premise of distributed storage, making full use of the high efficiency of GPU computing, and greatly improving the update speed of embedded features while meeting storage requirements.
  • the recommendation model in order to enable the news client to provide users with news recommendations that meet their interests, it is necessary to train and deploy a news recommendation model (hereinafter referred to as the recommendation model).
  • the input of the recommendation model is the embedded features of the target object
  • the output of the recommendation model is the probability that the target object clicks on a certain news.
  • the probability is greater than the probability threshold
  • the news client pushes the news to the target object.
  • the embedded features are obtained by mapping sparse features such as account ID. Therefore, the representation ability of the embedded features for sparse features greatly affects the recommendation accuracy, so we need to update the embedded features so that the embedded features have excellent expression capabilities.
  • the data sample can be a data sample of a recommendation model.
  • a single data sample can include the relevant object data of an account that logs into a news client.
  • the relevant object data here includes the account identifier (account ID), account age, and the data that represents the account. Attribute labels of interest, etc.
  • the embedded features corresponding to the first type of sparse features in the full data sample are stored in GPU1 deployed in the server 200, and the embedded features corresponding to the second type of sparse features in the full data sample are stored in GPU2 deployed in the server 200.
  • Server 200 receives target sample A, where target sample A is input to GPU1, GPU1 obtains sparse features of target sample A, the sparse features of target sample A include sparse features of target sample A in account identification type and sparse features of target sample A in age group type, GPU1 obtains embedded features of sparse features of target sample A in account identification type from embedded features of full sparse features of account identification type; GPU1 obtains embedded features of sparse features of target sample A in age group type from GPU2, the embedded features of sparse features of target sample A in age group type are obtained by GPU2 from embedded features of full sparse features of age group type; GPU1 determines embedded features of sparse features of target sample A based on embedded features of sparse features of target sample A in account identification type and embedded features of sparse features of target sample A in age group type; GPU1 forward propagates based on embedded features of sparse features of target sample A, and transmits update instructions to GPU2 based on probability mapping results, the update instructions are used to instruct GPU2 to update embedded features
  • the recommendation system will help users select satisfactory products, interesting news, most suitable courses, short videos that match users' interests, and music that meets users' needs. See Figure 4.
  • the problem that the recommendation system has to deal with can be defined in a more formal way: for user U, in a specific scene C, for a large amount of item information, construct a function f(U, I, C) to predict the user's preference for a specific candidate item I, and then sort all candidate items according to the preference to generate a recommendation list.
  • the embeddings in recommendation scenarios are stored on the CPU, and embedding updates are also completed in the CPU.
  • the CPU performs complex and large-scale computing operations on the embedding layer, which not only limits the speed of training, but also makes it impossible to use more complex models in actual production, because the use of complex models will cause the CPU to take too long to calculate the given input and cannot respond to requests in time.
  • GPU training has achieved great success in applications such as image recognition and text processing. GPU training has greatly improved the speed of training deep neural networks with its unique efficiency advantages in mathematical operations such as convolution.
  • Parameter Server a distributed and scalable parameter service (Parameter Server) solution was proposed, which almost perfectly solved the distributed training problem of machine learning models.
  • Parameter Server is not only directly applied to the machine learning platform, but also integrated into mainstream deep learning frameworks such as TensorFlow and MXNet as an important solution for distributed training of machine learning.
  • FIG 5 it can be seen that Parameter Server consists of server nodes (server nodes) and worker nodes (worker nodes). The main functions of the server node are to save model parameters, accept local gradients calculated by worker nodes, summarize and calculate global gradients, and update model parameters.
  • the main function of the worker node is to save part of the training data, pull the latest model parameters from the server node, calculate local gradients based on the training data, and upload them to the server node.
  • the single-machine training of TensorFlow is performed on a single worker node.
  • the worker node performs parallel calculations between different GPU and CPU nodes in the form of a task relationship graph.
  • the role of the embedding layer in a deep learning network is to convert sparse input vectors into dense vectors, but the existence of the embedding layer often slows down the convergence of the entire neural network.
  • the number of parameters in the embedding layer is huge. Assuming that the dimension of the input layer is 100,000, the output dimension of the embedding layer is 32, and then 5 layers of 32-dimensional fully connected layers are added, and the final output layer dimension is 10, then the number of parameters from the input layer to the embedding layer is 3,200,000, and the total number of parameters of all other layers is 4,416.
  • the total weight of the embedding layer accounts for 99.86%, which means that the weight of the embedding layer accounts for the vast majority of the weight of the entire network.
  • the training process most of the training time and computational overhead are occupied by the embedding layer; because the input features are too sparse, during the process of stochastic gradient descent, only the weights of the embedding layer connected to the non-zero features (embedded features) will be updated, which further reduces the convergence speed of the embedding layer.
  • the use of large-scale sparse discrete features leads to a sharp increase in the amount of data in the embedding layer of the deep model.
  • the recommendation models with TB size were once popular in major business scenarios in the industry. In most cases, because the recommendation tasks are very sensitive to latency, the model structure of the coarse-grained scenario is usually "short and fat", with a large number of feature inputs and then mapped to embedding vectors, and the depth of the neural network part is relatively low. Therefore, the analysis and embedding of the feature sample content account for a relatively high proportion of the overall time consumption.
  • the embedding of the recommendation scenario is stored on the CPU, and the embedding training is also completed on the CPU side, which not only limits the training speed, but also because the embedding accounts for a relatively high proportion of the overall time consumption, it is impossible to support the calculation of models with higher depth. Therefore, more complex models cannot be used in actual production, because the use of complex models will cause the CPU to calculate the given input for too long and fail to respond to requests in time.
  • the data processing method provided in the embodiment of the present application supports GPU training of large-scale embedding.
  • the embodiment of the present application provides a training solution based on the GPU system.
  • the model can be selected as NVIDIAV100.
  • the hardware topology of the server includes 8 GPUs, and the characteristics of the hardware are fully considered to give full play to the performance advantages.
  • the GPU system mainly includes three core modules: data module, computing module and communication module.
  • the data module relies on Tensorflow's data processing API to complete pre-reading (prefetch) and dump (dumpfile) and other logic; the computing module is used to control each GPU to start a TensorFlow training process to perform training; in the communication module, the Horovod process is used for inter-card communication of distributed training, and a Horovod process is started on each node to perform the corresponding communication tasks.
  • GPU0 and GPU1 respectively execute the processing logic of prefetch and dumpfile through their respective IO interfaces, so that GPU0 obtains input sample 0, GPU1 obtains input sample 1, GPU0 performs feature analysis on sample 0 to obtain sparse feature 0 of the first type and sparse feature 1 of the second type, GPU1 performs feature analysis on sample 1 to obtain sparse feature 0 of the first type and sparse feature 1 of the second type, GPU0 locally stores sparse feature 0 of the first type of all samples, GPU1 locally stores sparse feature 1 of the second type of all samples, therefore GPU0 transmits sparse feature 1 of the second type of sample 0 to GPU1, GPU1 transmits sparse feature 0 of the first type of sample 1 to GPU0.
  • GPU0 searches in hash table 0 based on sparse feature 0 of the first type of sample 0 and sparse feature 0 of the first type of sample 1, and obtains embedded feature 0 of the first type of sample 0 and embedded feature 0 of the first type of sample 1
  • GPU1 searches in hash table 0 based on sparse feature 1 of the second type of sample 0 and sample 1, and obtains embedded feature 0 of the first type of sample 0 and embedded feature 0 of the first type of sample 1.
  • the second type of sparse feature 1 of sample 1 is searched in hash table 1 to obtain the second type of embedded feature 1 of sample 0 and the second type of embedded feature 1 of sample 1, where the sparse features are keys (key), for example, key1, key2 and key3, and the embedded features are values (value), for example, value1, value2, value3).
  • GPU0 transmits the first type of sparse feature 0 of sample 1 to GPU1
  • GPU1 transmits the second type of sparse feature 1 of sample 0 to GPU0, so that GPU0 performs reasoning processing based on the deep model according to the first type of embedded feature 0 of sample 0 and the second type of embedded feature 1 of sample 0, and GPU1 performs reasoning processing based on the deep model according to the first type of embedded feature 0 of sample 1 and the second type of embedded feature 1 of sample 1.
  • the entire training process of the embodiment of the present application involves several key parts such as embedding storage, sparse feature processing and embedding segmentation, and inter-card communication.
  • the embodiment of the present application uses the component tfra.dynamic_embedding as the storage method for embedding.
  • the component uses tf.lookup.MutableHashTable to save multiple parameters (fullweights), and can reuse tensorflow's native optimizer.
  • the embodiment of the present application uses the dynamic embedding data structure storage of the tfra component for sparse features.
  • the number of sparse feature parameters is large, the number of corresponding embedded features will be large (for example, greater than the threshold), and the GPU single card memory cannot accommodate it.
  • the embeddings corresponding to the sparse features can be divided into the memory of multiple GPU cards by type. When the total number of sparse features is small, it means that the total number of corresponding embedded features is small (for example, not greater than the threshold), and the GPU single card memory can accommodate it.
  • Use the Replica method to place a full amount of embedded features in the memory of each GPU card.
  • Sparse features are large in scale, so hash tables are used to store them. Since the input sample data of each GPU is different, the embedding feature vectors corresponding to the input sparse features may be stored on other GPU cards.
  • each card queries the embedding feature vectors corresponding to the sparse features in the internal hash table. Then, through the AllToAll communication between GPUs, the feature vectors obtained from other GPUs in the first AllToAll are returned to the original path.
  • the sparse features of the sample input of each GPU can obtain the corresponding embedded features.
  • the gradients of the embedded features will be returned to other GPUs again through the AllToAll communication between GPUs. After each GPU obtains its own gradient, it performs optimization processing through the optimizer to complete the optimization of large-scale embedded features.
  • GPU0 performs feature analysis on sample 0 to obtain sparse features 0 of the first type and sparse features 1 of the second type.
  • GPU1 performs feature analysis on sample 1 to obtain sparse features 0 of the first type and sparse features 1 of the second type. Sparse features 0 of the first type of all samples are locally stored in GPU0, and sparse features 1 of the second type of all samples are locally stored in GPU1. Therefore, GPU0 transmits sparse features 1 of the second type of sample 0 to GPU1, and GPU1 transmits sparse features 0 of the first type of sample 1 to GPU0.
  • GPU0 searches hash table 0 to obtain embedded features 0 of the first type of sample 0 and embedded features 0 of the first type of sample 1.
  • GPU1 searches hash table 1 to obtain embedded features 1 of the second type of sample 0 and embedded features 1 of the second type of sample 1.
  • GPU0 transmits the first type of sparse feature 0 of sample 1 to GPU1, and GPU1 transmits the second type of sparse feature 1 of sample 0 to GPU0, so that GPU0 performs reasoning processing based on the deep model according to the first type of embedded feature 0 of sample 0 and the second type of embedded feature 1 of sample 0 to obtain a probability mapping result, and determines the gradient 0 corresponding to the first type of embedded feature 0 and the gradient 1 corresponding to the second type of embedded feature 1 based on the error between the probability mapping result and the label value, and GPU1 performs reasoning processing based on the deep model according to the first type of embedded feature 0 of sample 1 and the second type of embedded feature 1 of sample 1 to obtain a probability mapping result, and determines the gradient 0 corresponding to the first type of embedded feature 0 and the gradient 1 corresponding to the second type of embedded feature 1 based on the error between the probability mapping result and the label value.
  • the gradient 0 of the embedded feature 0 and the gradient 1 of the embedded feature 1 of the second type are transmitted by GPU0 to GPU1, and the optimizer of GPU1 updates the embedded feature 1 of the second type based on the gradient 1 corresponding to the embedded feature 1 of the second type.
  • GPU1 transmits the gradient 0 corresponding to the embedded feature 0 of the first type to GPU0, and the optimizer of GPU0 updates the embedded feature 0 of the first type based on the gradient 0 corresponding to the embedded feature 0 of the first type.
  • the embodiment of the present application proposes a GPU training solution that can support large-scale embedding, which has the following advantages: using GPU to store embedding parameters, making full use of the high efficiency of GPU training, and greatly improving the training speed of the model; supporting large-scale embedding parameter storage, when the embedding parameters cannot be stored on a single GPU card, the embedding split storage logic is implemented through efficient GPUNCCL communication; without losing ease of use, it can seamlessly connect to third-party frameworks, such as TensorFlow, Pytorch and other mainstream frameworks.
  • third-party frameworks such as TensorFlow, Pytorch and other mainstream frameworks.
  • the data processing method is applied to a second graphics processor.
  • the first graphics processor stores embedded features of a first type of full sparse features
  • the second graphics processor stores embedded features of a second type of full sparse features.
  • the first type is different from the second type.
  • the software module stored in the artificial intelligence-based data processing device 255-1 of the memory 250 may include: a first receiving module 2551, configured for the first graphics processor to obtain sparse features of a target sample, the sparse features of the target sample include the first sparse features of the first type and the second sparse features of the second type; a first acquisition module 2552, configured for the first graphics processor to obtain from the second graphics processor the embedded features of the full sparse features of the second type obtained from the query; The embedded features corresponding to the second sparse features obtained; the first return module 2553 is configured for the first graphics processor to obtain the embedded features corresponding to the second sparse features obtained from the embedded features of the second type of full sparse features from the second graphics processor; the first determination module 2554 is configured for the first graphics processor to combine the embedded features corresponding to the first sparse features and the embedded features corresponding to the second sparse features to form the embedded features corresponding to the sparse features of the target sample; the first update module 2555 is configured for the first graphics processor
  • the first updating module 2555 is further configured as the first graphics processor to update the embedded features of the first type of full-quantity sparse features according to the probability mapping result.
  • the first graphics processor and the second graphics processor belong to a plurality of graphics processors, and the plurality of graphics processors store embedded features of a plurality of types of full-quantity sparse features, the first type includes at least one type of the plurality of types, and the second type includes at least one type of the plurality of types.
  • the storage locations of the embedded features of the first type of all sparse features and the storage locations of the embedded features of the second type of all sparse features are divided when the data amount of the embedded features of multiple types of all sparse features is greater than a threshold.
  • the embedded features of the multiple types of full amount sparse features are stored in each of the multiple graphics processors.
  • the first return module 2553 is further configured as: the first graphics processor transmits the second sparse feature to the second graphics processor, so that the second graphics processor queries the embedded feature corresponding to the second sparse feature from the embedded features of the full amount of sparse features of the second type; the first graphics processor receives the embedded feature corresponding to the second sparse feature.
  • the first returning module 2553 is further configured to: the first graphics processor queries the embedded features corresponding to the first sparse features from the embedded features of the full amount of sparse features of the first type.
  • the first receiving module 2551 is further configured to obtain a historical target sample previously obtained by the first graphics processor from multiple data samples; use data samples other than the historical target sample from the multiple data samples as other data samples; and randomly obtain at least one data sample from the other data samples as the target sample.
  • the first receiving module 2551 is further configured to perform an average division process on the multiple data samples based on the number of graphics processors to obtain multiple sample sets, wherein each sample set includes at least one data sample, and the number of sample sets is the same as the number of graphics processors; determine a one-to-one correspondence between the multiple sample sets and the multiple graphics processors; and use the data samples in the sample set that have a corresponding relationship with the first graphics processor as the target samples corresponding to the first graphics processor.
  • the first update module 2555 is configured to use the embedded features corresponding to the sparse features of the target sample as input to the recommendation model, perform probability mapping processing on the embedded features corresponding to the sparse features of the target sample through the recommendation model, and obtain the recommendation probability corresponding to the target sample as the probability mapping result; the first graphics processor determines the error between the recommendation probability corresponding to the target sample and the label value corresponding to the target sample, and determines the gradient of the embedded features of the second type of full sparse features of the target sample based on the error; and generates an update instruction carrying the gradient of the embedded features of the second type of full sparse features.
  • the data processing method is applied to a second graphics processor.
  • the first graphics processor stores embedded features of a first type of full sparse features
  • the second graphics processor stores embedded features of a second type of full sparse features.
  • the first type is different from the second type.
  • the software module stored in the artificial intelligence-based data processing device 255-2 of the memory 250 may include: a second receiving module 2556, configured for the second graphics processor to receive the second sparse features of the second type sent by the first graphics processor; a query module 2557, configured for the second graphics processor to query the embedded features of the second type of full sparse features to obtain the embedded features corresponding to the second sparse features; a second return module 2558, configured for the second graphics processor to transmit the embedded features corresponding to the second sparse features to the first graphics processor.
  • the device enables the first graphics processor to perform probability mapping processing on the embedded features corresponding to the first sparse features of the first type and the embedded features corresponding to the second sparse features, wherein the embedded features corresponding to the first sparse features are obtained by the first graphics processor from the embedded features of the full amount of sparse features of the first type;
  • the second update module 2559 is configured so that the second graphics processor receives an update instruction generated according to the probability mapping result from the first graphics processor, and updates the embedded features of the full amount of sparse features of the second type based on the update instruction.
  • An embodiment of the present application provides a graphics processor for executing the artificial intelligence-based data processing method described above in the embodiment of the present application.
  • the embodiment of the present application provides a computer program product, which includes computer executable instructions, which are stored in a computer-readable storage medium.
  • the processor of the electronic device reads the computer executable instructions from the computer-readable storage medium, and the processor executes the computer executable instructions, so that the electronic device executes the data processing method based on artificial intelligence described in the embodiment of the present application, and the processor is a graphics processor or a central processing unit.
  • An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein computer-executable instructions are stored.
  • the processor will execute the artificial intelligence-based data processing method provided by the embodiment of the present application, for example, the artificial intelligence-based data processing method shown in Figures 3A-3D, where the processor is a graphics processor or a central processing unit.
  • the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or it may be various devices including one or any combination of the above memories.
  • computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.
  • computer-executable instructions may, but need not necessarily, correspond to a file in a file system, may be stored as part of a file that holds other programs or data, such as in a hypertext markup file.
  • the program may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).
  • HTML Hyper Text Markup Language
  • computer executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed at multiple sites and interconnected by a communication network.
  • the first graphics processor obtains different types of sparse features of the target sample.
  • the first graphics processor obtains the first type of embedded features locally, and obtains the second type of embedded features from the second graphics processor.
  • the classification according to type makes it convenient to obtain embedded features from the corresponding graphics processor, which can improve acquisition efficiency and facilitate feature management.
  • the first graphics processor performs probability mapping processing on the two types of embedded features, and transmits update instructions to the second graphics processor according to the probability mapping results.
  • the update instructions are used to instruct the second graphics processor to update the second type of embedded features, thereby realizing the update of embedded features under the premise of distributed storage, making full use of the high efficiency of GPU computing, and greatly improving the update speed of embedded features while meeting storage requirements.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Software Systems (AREA)
  • General Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • Artificial Intelligence (AREA)
  • Computing Systems (AREA)
  • Data Mining & Analysis (AREA)
  • Evolutionary Computation (AREA)
  • Mathematical Physics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Biomedical Technology (AREA)
  • Biophysics (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Molecular Biology (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Medical Informatics (AREA)
  • Probability & Statistics with Applications (AREA)
  • Algebra (AREA)
  • Computational Mathematics (AREA)
  • Mathematical Analysis (AREA)
  • Mathematical Optimization (AREA)
  • Pure & Applied Mathematics (AREA)
  • Neurology (AREA)
  • Complex Calculations (AREA)
  • Information Transfer Between Computers (AREA)
  • Processing Or Creating Images (AREA)
  • Image Processing (AREA)

Abstract

本申请提供了一种基于人工智能的数据处理方法、装置、电子设备、计算机可读存储介质及计算机程序产品;数据处理方法应用于第一图形处理器,第一图形处理器存储有第一类型的全量稀疏特征的嵌入特征,第二图形处理器存储有第二类型的全量稀疏特征的嵌入特征,第一类型与第二类型不同;方法包括:第一图形处理器获取目标样本的稀疏特征,目标样本的稀疏特征包含第一类型的第一稀疏特征和第二类型的第二稀疏特征;第一图形处理器从第一类型的全量稀疏特征的嵌入特征中,获取对应于第一稀疏特征的嵌入特征;第一图形处理器从第二图形处理器获取从第二类型的全量稀疏特征的嵌入特征查询得到的对应于第二稀疏特征的嵌入特征;第一图形处理器将对应于第一稀疏特征的嵌入特征和对应于第二稀疏特征的嵌入特征,组成对应于目标样本的稀疏特征的嵌入特征;第一图形处理器对对应于目标样本的稀疏特征的嵌入特征执行概率映射处理,并根据概率映射结果生成更新指令,更新指令用于指示第二图形处理器更新第二类型的全量稀疏特征的嵌入特征。

Description

数据处理方法、装置、电子设备、计算机可读存储介质及计算机程序产品
相关申请的交叉引用
本申请基于申请号为202310492698.2、申请日为2023年04月28日的中国专利申请提出,并要求中国专利申请的优先权,中国专利申请的全部内容在此引入本申请作为参考。
技术领域
本申请涉及人工智能技术,尤其涉及一种数据处理方法、装置、电子设备、计算机可读存储介质及计算机程序产品。
背景技术
人工智能(Artificial Intelligence,AI)是计算机科学的一个综合技术,通过研究各种智能机器的设计原理与实现方法,使机器具有感知、推理与决策的功能。基于医疗图像以及医疗文本的分析是人工智能领域的重要应用之一,医疗分析系统是指利用计算机对医疗图像和医疗文本进行处理、分析和理解,以识别出各种不同模式的目标和对象的系统。
相关技术中基于Parameter Server方式对嵌入特征进行学习,通常稀疏特征的嵌入特征表示(embedding)会存放在中央处理器(CPU,Central Processing Unit)上,因此embedding训练也会在CPU内完成。但是由于embedding运算的规模较大,从而导致训练速度较慢。
发明内容
本申请实施例提供一种基于人工智能的数据处理方法、装置、电子设备、计算机可读存储介质及计算机程序产品,能够在实现大规模嵌入特征存储的前提下通过图形处理器提高更新处理速度。
本申请实施例的技术方案是这样实现的:
本申请实施例提供一种基于人工智能的数据处理方法,所述数据处理方法应用于第一图形处理器,所述第一图形处理器存储有第一类型的全量稀疏特征的嵌入特征,第二图形处理器存储有第二类型的全量稀疏特征的嵌入特征,所述第一类型与所述第二类型不同,所述方法包括:
所述第一图形处理器获取目标样本的稀疏特征,所述目标样本的稀疏特征包含所述第一类型的第一稀疏特征和所述第二类型的第二稀疏特征;
所述第一图形处理器从所述第一类型的全量稀疏特征的嵌入特征中,获取对应于所述第一稀疏特征的嵌入特征;
所述第一图形处理器从所述第二图形处理器获取从所述第二类型的全量稀疏特征的嵌入特征查询得到的对应于所述第二稀疏特征的嵌入特征;
所述第一图形处理器将对应于所述第一稀疏特征的嵌入特征和对应于所述第二稀疏特征的嵌入特征,组成对应于所述目标样本的稀疏特征的嵌入特征;
所述第一图形处理器对所述对应于所述目标样本的稀疏特征的嵌入特征执行概率映射处理,并根据概率映射结果生成更新指令,所述更新指令用于指示所述第二图形处理器更新所述第二类型的全量稀疏特征的嵌入特征。
本申请实施例提供一种基于人工智能的数据处理装置,所述数据处理方法应用于第一图形处理器,第一图形处理器存储有第一类型的全量稀疏特征的嵌入特征,所述第二图形处理器存储有第二类型的全量稀疏特征的嵌入特征,所述第一类型与所述第二类型不同,所述装置包括:
第一接收模块,配置为所述第一图形处理器获取目标样本的稀疏特征,所述目标样本的稀疏特征包含所述第一类型的第一稀疏特征和所述第二类型的第二稀疏特征;
第一获取模块,配置为所述第一图形处理器从所述第二图形处理器获取从所述第二类型的全量稀疏特征的嵌入特征查询得到的对应于所述第二稀疏特征的嵌入特征;
第一返回模块,配置为所述第一图形处理器从所述第二图形处理器获取从 所述第二类型的全量稀疏特征的嵌入特征查询得到的对应于所述第二稀疏特征的嵌入特征;
第一确定模块,配置为所述第一图形处理器将对应于所述第一稀疏特征的嵌入特征和对应于所述第二稀疏特征的嵌入特征,组成对应于所述目标样本的稀疏特征的嵌入特征;
第一更新模块,配置为所述第一图形处理器对所述对应于所述目标样本的稀疏特征的嵌入特征执行概率映射处理,并根据概率映射结果生成更新指令,所述更新指令用于指示所述第二图形处理器更新所述第二类型的全量稀疏特征的嵌入特征。
本申请实施例提供一种数据处理方法,所述数据处理方法应用于第二图形处理器,第一图形处理器存储有第一类型的全量稀疏特征的嵌入特征,所述第二图形处理器存储有第二类型的全量稀疏特征的嵌入特征,所述第一类型与所述第二类型不同;
所述方法包括:
所述第二图形处理器接收所述第一图形处理器发送的,所述第二类型的第二稀疏特征;
所述第二图形处理器从所述第二类型的全量稀疏特征的嵌入特征中查询得到对应于所述第二稀疏特征的嵌入特征;
所述第二图形处理器将对应于所述第二稀疏特征的嵌入特征传输至所述第一图形处理器,以使所述第一图形处理器对对应于所述第一类型的第一稀疏特征的嵌入特征和对应于所述第二稀疏特征的嵌入特征,执行概率映射处理,其中,对应于所述第一稀疏特征的嵌入特征是所述第一图形处理器从所述第一类型的全量稀疏特征的嵌入特征中获取得到的;
所述第二图形处理器从第一图形处理器接收根据概率映射结果生成的更新指令,并基于所述更新指令更新所述第二类型的全量稀疏特征的嵌入特征。
本申请实施例提供一种数据处理装置,所述数据处理方法应用于第二图形处理器,第一图形处理器存储有第一类型的全量稀疏特征的嵌入特征,所述第 二图形处理器存储有第二类型的全量稀疏特征的嵌入特征,所述第一类型与所述第二类型不同;
所述装置包括:
第二接收模块,配置为所述第二图形处理器接收所述第一图形处理器发送的,所述第二类型的第二稀疏特征;
查询模块,配置为所述第二图形处理器从所述第二类型的全量稀疏特征的嵌入特征中查询得到对应于所述第二稀疏特征的嵌入特征;
第二返回模块,配置为所述第二图形处理器将对应于所述第二稀疏特征的嵌入特征传输至所述第一图形处理器,以使所述第一图形处理对对应于所述第一类型的第一稀疏特征的嵌入特征和对应于所述第二稀疏特征的嵌入特征,执行概率映射处理,其中,对应于所述第一稀疏特征的嵌入特征是所述第一图形处理器从所述第一类型的全量稀疏特征的嵌入特征中获取得到的;
第二更新模块,配置为所述第二图形处理器从第一图形处理器接收根据概率映射结果生成的更新指令,并基于所述更新指令更新所述第二类型的全量稀疏特征的嵌入特征。
本申请实施例提供一种图形处理器,所述图形处理器用于执行本申请实施例提供的数据处理方法。
本申请实施例提供一种电子设备,所述电子设备包括:
存储器,用于存储计算机可执行指令;
图形处理器,用于执行所述存储器中存储的计算机可执行指令时,实现本申请实施例提供的数据处理方法。
本申请实施例提供一种计算机可读存储介质,存储有计算机可执行指令,所述计算机可执行指令被处理器执行时实现本申请实施例提供的数据处理方法,所述处理器是中央处理器或者图形处理器。
本申请实施例提供一种计算机程序产品,包括计算机可执行指令,所述计算机可执行指令被处理器执行时实现本申请实施例提供的数据处理方法,所述处理区是中央处理器或者图形处理器。
本申请实施例具有以下有益效果:
按照类型将嵌入特征存储在不同的图形处理器中,可以实现嵌入特征的分布式存储,第一图形处理器获取目标样本的不同类型的稀疏特征,第一图形处理器从本地获取第一类型的嵌入特征,并从第二图形处理器,获取第二类型的嵌入特征,按照类型划分方便从对应图形处理器中获取嵌入特征,可以提高获取效率,并且便于特征管理,第一图形处理器对两个类型的嵌入特征进行概率映射处理,并根据概率映射结果向第二图形处理器传输更新指令,更新指令用于指示第二图形处理器更新第二类型的嵌入特征,从而在分布式存储的前提下实现了嵌入特征的更新,充分利用图形处理器计算的高效性,在满足存储需求的同时极大提升嵌入特征更新速度。
附图说明
图1A是本申请实施例提供的基于人工智能的数据处理系统的结构示意图;
图1B是本申请实施例提供的数据处理系统的架构图;
图2是本申请实施例提供的电子设备的结构示意图;
图3A-图3D是本申请实施例提供的基于人工智能的数据处理方法的流程示意图;
图4是本申请实施例提供的推荐系统的框架示意图;
图5是本申请实施例提供的Parameter Server的物理架构示意图;
图6是本申请实施例提供的单工作节点的示意图;
图7是本申请实施例提供的基于人工智能的数据处理方法的数据流转示意图;
图8是本申请实施例提供的基于人工智能的数据处理方法的数据流转示意图。
具体实施方式
为了使本申请的目的、技术方案和优点更加清楚,下面将结合附图对本申 请作进一步地详细描述,所描述的实施例不应视为对本申请的限制,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其它实施例,都属于本申请保护的范围。
在以下的描述中,涉及到“一些实施例”,其描述了所有可能实施例的子集,但是可以理解,“一些实施例”可以是所有可能实施例的相同子集或不同子集,并且可以在不冲突的情况下相互结合。
在以下的描述中,所涉及的术语“第一\第二\第三”仅仅是是区别类似的对象,不代表针对对象的特定排序,可以理解地,“第一\第二\第三”在允许的情况下可以互换特定的顺序或先后次序,以使这里描述的本申请实施例能够以除了在这里图示或描述的以外的顺序实施。
除非另有定义,本文所使用的所有的技术和科学术语与属于本申请的技术领域的技术人员通常理解的含义相同。本文中所使用的术语只是为了描述本申请实施例的目的,不是旨在限制本申请。
对本申请实施例进行进一步详细说明之前,对本申请实施例中涉及的名词和术语进行说明,本申请实施例中涉及的名词和术语适用于如下的解释。
1)嵌入特征(Embedding):Embedding就是用一个数值向量表示一个对象的方法,这里说的对象可以是一个词、一个物品,也可以是一部电影等等。推荐场景中的类别、ID型特征非常多,大量使用独热编码会导致样本特征向量极度稀疏,而深度学习的结构特点又不利于稀疏特征向量的处理,因此几乎所有深度学习推荐模型都会由嵌入层负责将稀疏高维特征向量转换成稠密低维特征向量,也即嵌入特征Embedding。
2)TensorF1ow:是一个基于数据流编程的符号数学系统,被广泛应用于各类机器学习算法的编程实现,其拥有多层级结构,可部署于各类服务器、PC终端和网页,并支持GPU和TPU高性能数值计算,被广泛应用于产品开发和各领域的科学研究。
3)参数服务器(Parameter Server):是个编程框架,用于方便分布式并行程序的编写,其中重点是对大规模参数的分布式存储和协同的支持,参数服务 器的集群中的节点可以分为计算节点和参数服务节点两种。其中,计算节点负责对分配到自己本地的训练数据(块)计算学习,并更新对应的参数;参数服务节点采用分布式存储的方式,各自存储全局参数的一部分,并作为服务方接受计算节点的参数查询和更新请求。
4)稀疏特征与稠密特征:稀疏特征是那些在数据集中不连续出现的特征,并且大多数值为零,例如账号标识,如果存在100个账号,那么对于每个账号标识而言,稀疏特征中会有99位数值为0,只有1位数值为1,之所以称为稀疏特征,是因为它们在数据集中只有很少的非零值,全量稀疏特征表征所有样本的稀疏特征,例如,对于账号类型而言,所有样本总计100个账号,则这100个账号分别对应的100个稀疏特征即为账号类型的全量稀疏特征。稠密特征一般是相对稀疏特征来说的,稠密特征中数值0的占比较少,甚至没有。
5)类型:本申请实施例提供的类型可以是内容类型或者是数据格式类型,内容类型是基于稀疏特征所表征的语义进行区分的,数据格式类型是基于稀疏特征的数据格式进行区分的。
相关技术中基于Parameter Server方式对嵌入特征进行学习,通常稀疏特征的嵌入特征表示(embedding)会存放在中央处理器(CPU,Central Processing Unit)上,embedding训练也在CPU内完成。但是由于embedding运算的规模较大,如果embedding训练在CPU内完成,会限制训练速度。申请人在实施本申请实施例时发现如果将embedding训练部署在图形处理器(GPU,graphics processing unit)中,可以有效提高训练速度,但是GPU的存储能力无法满足大规模embedding的存储要求。
本申请实施例提供一种基于人工智能的数据处理方法、装置、电子设备和计算机可读存储介质,能够在实现大规模嵌入特征存储的前提下通过图形处理器提高更新处理速度。
本申请实施例所提供的数据处理方法,可以由终端/服务器独自实现;也可以由终端和服务器协同实现,例如终端或者服务器独自承担下文所述的数据处理方法,或者,终端向服务器发送数据样本,服务器根据接收的数据样本执行 数据处理方法。
本申请实施例提供的用于数据处理的电子设备可以是各种类型的终端或服务器,其中,服务器可以是独立的物理服务器,也可以是多个物理服务器构成的服务器集群或者分布式系统,还可以是提供云计算服务的云服务器;终端可以是智能手机、平板电脑、笔记本电脑、台式计算机、智能音箱、智能手表、车载终端等,但并不局限于此。终端以及服务器可以通过有线或无线通信方式进行直接或间接地连接,本申请在此不做限制。
以服务器为例,例如可以是部署在云端的服务器集群,向用户开放人工智能云服务(AI as a Service,AiaaS),平台会把几类常见的AI服务进行拆分,并在云端提供独立或者打包的服务,这种服务模式类似于一个AI主题商城,所有的用户都可以通过应用程序编程接口的方式来接入使用AIaaS平台提供的一种或者多种人工智能服务。
作为示例,其中的一种人工智能云服务可以为数据处理服务,即云端的服务器封装有本申请实施例提供的数据处理的程序。使用者通过终端调用云服务中的数据处理服务,以使部署在云端的服务器调用封装的数据处理的程序。
参见图1A,图1A是本申请实施例提供的数据处理系统的应用场景示意图,终端400通过网络300连接服务器200,网络300可以是广域网或者局域网,又或者是二者的组合,服务器200中部署有4个GPU,下面以GPU1和GPU2进行说明。
数据样本可以是推荐模型的数据样本,例如针对新闻推荐模型,单个数据样本可以包括登录新闻客户端的账号的相关对象数据,这里的相关对象数据包括账号标识(账号ID)、账号年龄、表征账号兴趣的属性标签等等。全量数据样本中对应第一类型(例如账号标识类型)的稀疏特征的嵌入特征保存在服务器200内部署的GPU1内,全量数据样本中对应第二类型(例如年龄段类型)的稀疏特征的嵌入特征保存在服务器200内部署的GPU2内。
终端400将匹配请求发送至服务器200,服务器200调用推荐模型,任意一个GPU(例如GPU1)中部署有推荐模型,推荐模型的输入是目标对象的账 号标识类型的嵌入特征以及年龄段类型的嵌入特征,因此GPU1需要从GPU2中基于年龄段类型的稀疏特征查询目标对象的年龄段类型的嵌入特征,推荐模型的输出是目标对象点击某个新闻的推荐概率,当推荐概率大于概率阈值时,服务器200向终端400返回该新闻,新闻客户端向目标对象推送该新闻。为了能够提高推荐准确度,下面介绍各个类型的嵌入特征更新的过程。
参见图1B,图1B是本申请实施例提供的数据处理系统的架构图。服务器200接收目标样本A,这里的目标样本A被输入至GPU1,GPU1获取目标样本A的稀疏特征,目标样本A的稀疏特征包含目标样本A在账号标识类型的稀疏特征和目标样本A在年龄段类型的稀疏特征,GPU1从账号标识类型的全量稀疏特征的嵌入特征中,获取目标样本A在账号标识类型的稀疏特征的嵌入特征;GPU1从GPU2,获取目标样本A在年龄段类型的稀疏特征的嵌入特征,目标样本A在年龄段类型的稀疏特征的嵌入特征是GPU2从年龄段类型的全量稀疏特征的嵌入特征查询得到的;GPU1根据目标样本A在账号标识类型的稀疏特征的嵌入特征和目标样本A在年龄段类型的稀疏特征的嵌入特征,确定目标样本A的稀疏特征的嵌入特征;GPU1根据目标样本A的稀疏特征的嵌入特征正向传播,并根据概率映射结果向GPU2传输更新指令,更新指令用于指示GPU2更新年龄段类型的全量稀疏特征的嵌入特征。
在一些实施例中,服务器200可以是独立的物理服务器,也可以是多个物理服务器构成的服务器集群或者分布式系统,还可以是提供云服务、云数据库、云计算、云函数、云存储、网络服务、云通信、中间件服务、域名服务、安全服务、CDN、以及大数据和人工智能平台等基础云计算服务的云服务器。终端400可以是智能手机、平板电脑、笔记本电脑、台式计算机、智能音箱、智能手表等,但并不局限于此。终端以及服务器可以通过有线或无线通信方式进行直接或间接地连接,本发明实施例中不做限制。
在一些实施例中,终端或服务器可以通过运行计算机程序来实现本申请实施例提供的基于人工智能的数据处理方法。举例来说,计算机程序可以是操作系统中的原生程序或软件模块;可以是本地(Native)应用程序(APP,Applic ation),即需要在操作系统中安装才能运行的程序,如直播APP或者即时通信APP;也可以是小程序,即只需要下载到浏览器环境中就可以运行的程序;还可以是能够嵌入至任意APP中的小程序。总而言之,上述计算机程序可以是任意形式的应用程序、模块或插件。
下面说明本申请实施例提供的用于数据处理的电子设备的结构,参见图2,图2是本申请实施例提供的用于数据处理的电子设备的结构示意图,以电子设备是服务器200为例说明,图2所示的用于数据处理的服务器200包括:至少一个处理器210(处理器210可以是图形处理器或者中央处理器)、存储器250、至少一个网络接口220。服务器200中的各个组件通过总线系统240耦合在一起。可理解,总线系统240用于实现这些组件之间的连接通信。总线系统240除包括数据总线之外,还包括电源总线、控制总线和状态信号总线。但是为了清楚说明起见,在图2中将各种总线都标为总线系统240。
处理器210可以是一种集成电路芯片,具有信号的处理能力,例如通用处理器、数字信号处理器(DSP,Digital Signal Processor),或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等,其中,通用处理器可以是微处理器或者任何常规的处理器等。
存储器250包括易失性存储器或非易失性存储器,也可包括易失性和非易失性存储器两者。其中,非易失性存储器可以是只读存储器(ROM,Read Only Memory),易失性存储器可以是随机存取存储器(RAM,Random Access Memory)。本申请实施例描述的存储器250旨在包括任意适合类型的存储器。存储器250可选地包括在物理位置上远离处理器210的一个或多个存储设备。
在一些实施例中,存储器250能够存储数据以支持各种操作,这些数据的示例包括程序、模块和数据结构或者其子集或超集,下面示例性说明。
操作系统251,包括用于处理各种基本系统服务和执行硬件相关任务的系统程序,例如框架层、核心库层、驱动层等,用于实现各种基础业务以及处理基于硬件的任务。
网络通信模块252,用于经由一个或多个(有线或无线)网络接口220到 达其他电子设备,示例性的网络接口220包括:蓝牙、无线相容性认证(WiFi)、和通用串行总线(USB,Universal Serial Bus)等。
在一些实施例中,本申请实施例提供的数据处理装置可以采用软件方式实现,例如,可以是上文所述的终端中的数据处理插件,可以是上文所述的服务器中数据处理服务。当然,不局限于此,本申请实施例提供的数据处理装置可以提供为各种软件实施例,包括应用程序、软件、软件模块、脚本或代码在内的各种形式。图2示出了存储在存储器250中的数据处理装置255-1,其可以是程序和插件等形式的软件,例如图像处理插件,并包括一系列的模块,包括第一接收模块2551、第一获取模块2552、第一返回模块2553、第一确定模块2554以及第一更新模块2555。图2还示出了存储在存储器250中的数据处理装置255-2,其可以是程序和插件等形式的软件,例如图像处理插件,并包括一系列的模块,包括第二接收模块2556、查询模块2557、第二返回模块2558、第二更新模块2559。
如前,本申请实施例提供的数据处理方法可以由各种类型的电子设备实施。例如包括多个图形处理器的电子设备,电子设备中部署有第一图形处理器和第二图形处理器(第一图形处理器和第二图形处理器还可以部署在不同的电子设备上),例如包括多个中央处理器的电子设备,电子设备中部署有第一中央处理器和第二中央处理器,其中,图形处理器的计算能力优于中央处理器。
在一些实施例中,第一图形处理器存储有第一类型的全量稀疏特征的嵌入特征,第二图形处理器存储有第二类型的全量稀疏特征的嵌入特征,第一类型与第二类型不同。通过本申请实施例可以将多个类型的全量稀疏特征的嵌入特征划分到不同的图形处理器中存储,从而可以缓解每个图形处理器的存储压力。
作为示例,若这里的类型是基于稀疏特征所表征的语义进行区分的内容类型,可以获取稀疏特征所表征的语义的字段(字段比如可以是关键词),当字段与预设字段匹配时,则将预设字段对应的内容类型作为稀疏特征的内容类型。例如,稀疏特征所表征的语义的字段可以是“账号”,预设字段可以是账号类型,那么这里的字段与预设字段匹配,稀疏特征的类型即为账号类型。按照内容类 型的区分进行分布式存储的方式可以有助于快速获取不同内容类型的稀疏特征的嵌入特征,从而在对特征多样性以及丰富度有要求的推荐场景下可以更快查询到相应内容类型的嵌入特征,提高分布式场景下训练过程中嵌入特征的获取效率。
作为示例,这里的类型还可以是基于稀疏特征的数据格式进行区分的数据格式类型,数据格式类型可以是文本类型、图像类型、语音类型等等,按照数据格式类型的区分进行分布式存储的方式可以有助于快速获取不同数据格式类型的稀疏特征的嵌入特征,从而可以更快查询到相应数据格式类型的嵌入特征,提高分布式场景下训练过程中嵌入特征的获取效率。
本申请实施例所提供的类型不局限于上述示例,在推荐场景下,稀疏特征是用于描述用户的数据,那么类型可以是基于不同用户进行划分得到的对象类型,例如有10个稀疏特征均是用于描述用户A,那么这10个稀疏特征所属的类型即为对象类型A,有10个稀疏特征均是用于描述用户B,那么这10个稀疏特征所属的类型即为对象类型B,即每个用户对应一种对象类型,将属于相同对象类型的稀疏特征的嵌入特征存储在相同图形处理器内,按照这种分布式存储方式,在推荐场景下,且参与训练的目标样本来源于不同用户时,可以更快查询到相应用户的所有相关嵌入特征,提高分布式场景下训练过程中嵌入特征的获取效率。本申请实施例提供的类型还可以是基于稀疏特征的来源进行划分得到的来源类型等等,本申请实施例提供的类型还可以是基于稀疏特征的获取时段进行划分得到的时段类型,在此不再进行赘述。
下文主要以内容类型为例进行说明,第一图形处理器和第二图形存储器存储有多个内容类型的所有稀疏特征,合计存储有100个稀疏特征,例如这里的内容类型可以是账号类型、年龄段类型等等,这里的第一类型可以是账号类型,这里的第二类型可以是年龄段类型,那么第一图形处理器中存储有身份类型的全量稀疏特征的嵌入特征表征第一图形处理器存储有所有样本的属于身份类型的稀疏特征,那么第二图形处理器中存储有年龄段类型的稀疏特征的全量嵌入特征,表征第二图形处理器存储有所有样本的属于年龄段类型的稀疏特征。
在一些实施例中,第一图形处理器和第二图形处理器属于多个图形处理器,多个图形处理器存储有多个类型的全量稀疏特征的嵌入特征,第一类型包含多个类型中的至少一个类型,第二类型包含多个类型中的至少一个类型。通过本申请实施例可以按照类型来将多个类型的全量稀疏特征的嵌入特征划分到不同的图形处理器中存储,从而在查询的时候可以具有明确的指向性,而不必从所有图形处理器中进行查询。
承接上述示例,图形处理器系统包括多个图形处理器,例如包括4个图形处理器,其中,第一图形处理器和第二图形处理器是4个图形处理器中的2个图形处理器,100个稀疏特征对应的100个嵌入特征会分布式存储在4个图形处理器中,这里的第一类型可以是账号类型和兴趣类型,这里的第二类型可以是年龄段类型和性别类型,那么第一图形处理器中存储有身份类型的所有稀疏特征的嵌入特征以及兴趣类型的所有稀疏特征的嵌入特征,那么第二图形处理器中存储有年龄段类型的所有稀疏特征的嵌入特征以及性别类型的所有稀疏特征的嵌入特征。
在一些实施例中,第一类型的全量稀疏特征的嵌入特征的存储位置和第二类型的全量稀疏特征的嵌入特征的存储位置,是在多个类型的全量稀疏特征的嵌入特征的数据量大于阈值的情况下划分的。通过本申请实施例仅在数据量大于阈值的情况下进行分布式存储,从而可以提高图形处理器的存储资源利用率,避免单个图形处理器的存储资源的负载过大。在多个类型的全量稀疏特征的嵌入特征的数据量不大于阈值的情况下,多个类型的全量稀疏特征的嵌入特征保存在多个图形处理器中的每个图形处理器中。通过本申请实施例可以在数据量不大于阈值的情况利用单个图形处理器来存储所有稀疏特征的嵌入特征,从而可以后续训练阶段不需要各个图形处理器不需要再从其他图形处理器获取嵌入特征,可以提高训练效率。
作为示例,这里的数据量可以是嵌入特征的总数目,在多个类型的全量稀疏特征的嵌入特征的总数目大于阈值的情况下,第一类型的全量稀疏特征的嵌入特征以及第二类型的全量稀疏特征的嵌入特征会以分布式的方式分别保存在 两个图形处理器中,在多个类型的全量稀疏特征的嵌入特征的总数目不大于阈值的情况下,多个类型的全量稀疏特征的嵌入特征,即所有的嵌入特征会被保存在每个图形处理器中,即每个图形处理器中均保存有全量的嵌入特征。
作为示例,在实际业务场景中,通常稀疏特征的总量较多,需要的处理较为复杂。本申请实施例针对稀疏特征使用tfra组件的动态embedding数据结构存储。当稀疏特征参数量较大,会导致对应的嵌入特征的数量较大(例如,大于阈值),GPU单卡显存无法容纳时,可以通过按照类型划分的方式将稀疏特征对应的embedding划分到多张GPU卡的显存中存放;当稀疏特征的总数量较少,则意味着对应的嵌入特征的总数量较小(例如不大于阈值),GPU单卡显存可以容纳时,使用Replica的方式,每张GPU卡的显存都放置一份全量的嵌入特征。
作为示例,在推荐系统领域,嵌入特征(embedding)已成为处理稀疏特征的常用手段,作为一种“函数映射”,embedding层通常将高维稀疏特征映射为低维稠密向量,这里所得到的低维稠密向量即为嵌入特征,这里的嵌入特征与对应的稀疏特征所表示的信息含义相同,但是所占据的存储空间更小且数据类型更低,再进行模型端到端训练。本申请实施例利用组件tfra.dynamic_embedding作为embedding的存储方式,该组件用tf.lookup.MutableHashTable保存多个参数(fullweights),也即保存嵌入特征,并且可以复用tensorflow原生的优化器。
作为示例,针对每个稀疏特征的映射处理的过程如下,将稀疏特征的类型编号与某个设定数值进行相除处理,得到对应的余数,基于稀疏特征的类型编号得到的余数即为稀疏特征的处理器标识,从而将稀疏特征保存在对应处理器标识所指示的图形处理器中。
通过本申请实施例可以将稀疏特征的嵌入特征按照类型区分保存在相应的图形处理器,合理分配多个图形处理器的存储资源。并且,相比于每个图形处理器存储所有类型的稀疏特征的嵌入特征这一技术方案,在本申请中,对于每个图形处理器来说,只需要负责所分配的特定类型的稀疏特征的嵌入特征保存 与更新,降低存储和更新开销,降低成本。通过本申请实施例可以灵活存储嵌入特征,在每个图形处理器的存储资源利用率以及图形处理器之间的交互资源占用率之间取得平衡。
参见图3A,图3A是本申请实施例提供的基于人工智能的数据处理方法的流程示意图,将结合图3A示出的步骤101至步骤105进行说明。
在步骤101中,第一图形处理器获取目标样本的稀疏特征,目标样本的稀疏特征包含第一类型的第一稀疏特征和第二类型的第二稀疏特征。
作为示例,服务器内可以部署有多个图形处理器,这里步骤101可以是由第一图形处理器实现,同时服务器可以是单个机器或者是由多个机器组成的服务器集群,即多个图形处理器可以部署在一个服务器内,也可以部署在多个服务器内。
作为示例,在进行训练时,目标样本被输入至第一图形处理器,第一图形处理器从目标样本中获取目标样本的稀疏特征,例如表征账号标识的稀疏特征以及表征年龄段的稀疏特征,即这里的目标样本的稀疏特征是包括第一稀疏特征和第二稀疏特征,这里第一稀疏特征是上文所述的第一类型的全量稀疏特征中对应目标样本的稀疏特征,这里的第二稀疏特征是上文所述的第二类型的全量稀疏特征中对应目标样本的稀疏特征,例如,目标样本是用户A,第一类型是账号类型,第二类型是年龄段类型,第一类型的全量稀疏特征是所有用户的账号类型的稀疏特征,如果用户的数目是100,每个用户有一个账号,那么第一类型的全量稀疏特征是100个用户的账号类型的稀疏特征,第二类型的全量稀疏特征是所有用户的年龄段类型的稀疏特征,总共有50个年龄段,那么第二类型的全量稀疏特征是50个年龄段类型的稀疏特征,这里第一稀疏特征是用户A在账号类型的稀疏特征(表征一个账号的稀疏特征)以及用户A在年龄段类型的稀疏特征(表征一个年龄段的稀疏特征)。
由于嵌入特征是被预先保存在第一图形处理器以及第二图形处理器中,而嵌入特征又是对稀疏特征进行嵌入压缩处理得到的,因此下面具体说明在存储之前获取多个数据样本的稀疏特征的过程,即获取全量稀疏特征的过程。
在一些实施例中,获取登录推荐客户端的多个对象账号以及每个对象账号的对象数据,将多个对象账号的对象数据作为多个数据样本,对多个对象账号的对象数据进行特征解析处理,得到多个对象账号的对象特征,并将多个对象账号的对象特征作为多个数据样本的稀疏特征。
作为示例,推荐客户端可以是新闻客户端、视频客户端、购物客户端等等,以新闻客户端为例进行说明,对象账号是之前登录过新闻客户端的账号,账号由用户持有(下文统一以对象来代替用户)。例如,对象A持有对象账号A,对象B持有对象账号B,可以获取对象账号A的属性数据以及操作数据作为对象账号A的对象数据。针对对象账号A而言,可以将对象账号A的对象数据作为一个数据样本,多个数据样本可以是由所有对象的对象数据构成的。
通过本申请实施例可以获取准确描述对象的稀疏特征,从而使得嵌入特征可以有效表征对象以提高嵌入特征的对象表征能力。
在一些实施例中,上述对多个对象账号的对象数据进行特征解析处理,得到多个对象账号的对象特征,可以通过以下技术方案实现:针对每个对象账号执行以下处理:从对象账号的对象数据中获取以下至少之一:对应对象账号的账号数据、对应对象账号的生物数据、对应对象账号的位置数据、对象账号的兴趣数据;对获取的数据进行独热编码处理,并将得到的编码结果作为对象账号的对象特征。
作为示例,针对对象A而言,可以获取对象账号A的属性数据以及操作数据作为对象账号A的对象数据,从属性数据中获取账号标识(例如账号ID)、生物标识(例如,性别)以及位置标识(例如长期居住地),从操作数据中获取兴趣标识,操作数据可以是浏览操作收藏操作等等,兴趣标识可以是被操作的信息的兴趣标签,例如被操作的信息是体育新闻,则兴趣标签可以是体育。通过独热编码处理对这些标识进行编码处理,例如,将性别男编码为(0,1),将性别女编码为(1,0),通过独热编码得到的编码结果可以作为对象特征。在存储过程中,对于部分对象而言,即便对象不同,他们也具有相同的性别或者相同的年龄,因此对于部分类型的稀疏特征而言,稀疏特征的数目是固定的。
通过本申请实施例可以获取准确描述对象的稀疏特征,并且稀疏特征可以全面表征对象,从而使得嵌入特征可以有效表征对象以提高嵌入特征的对象表征能力。
在一些实施例中,执行步骤101之前,获取第一图形处理器前一次从多个数据样本中获取到的历史目标样本;将多个数据样本中除历史目标样本的数据样本作为其他数据样本;从其他数据样本中随机获取至少一个数据样本作为目标样本。
作为示例,为了不断提高嵌入特征的表达能力,由图形处理器构成的图形处理器系统会进行多次训练,每次训练过程可以均是针对多个数据样本进行的,也可以针对多个数据样本中的部分数据样本进行,对此不进行限制,但是为了提高训练效率,本申请实施例提供多个图形处理器,从而每个图形处理器可以分担部分数据样本,例如,本次训练过程中会涉及到10个数据样本,这10个数据样本可以是多个数据样本,也可以是多个数据样本中的部分数据样本,第一图形处理器A会接收10个数据样本中的2个数据样本。针对每次训练过程,获取第一图形处理器上一次训练过程中接收到的数据样本(历史目标样本),这个获取的过程可以是第一图形处理器本身来执行,例如,上一次训练过程中接收到的数据样本是数据样本A和数据样本B,那么本次训练过程就会从多个数据样本中随机获取除数据样本A和数据样本B之外的任意数目的数据样本,可以获取两个数据样本,或者获取其他数目的数据样本。
通过本申请实施例可以避免多次训练过程中单个图形处理器重复处理相同的数据样本,提高单个图形处理器的训练效率。
在一些实施例中,执行步骤101之前,基于图形处理器的数目对多个数据样本进行平均划分处理,得到多个样本集合,其中,每个样本集合包括至少一个数据样本,样本集合的集合数目是与图形处理器的数目相同;确定多个样本集合与多个图形处理器之间的一一对应关系;将与第一图形处理器之间具有对应关系的样本集合中的数据样本,作为对应第一图形处理器的目标样本。通过本申请实施例可以通过平均划分的方式来保证每个图形处理器的负载均衡。通 过本申请实施例可以避免多次训练过程中单个图形处理器重复处理相同的数据样本,从而优化训练效果。
作为示例,除了按照上述思路来避免多次训练过程中单个图形处理器重复处理相同的数据样本,还可以通过平均划分的方式来保证每个图形处理器的负载均衡,例如多个数据样本包括10个数据样本,图形处理器系统中包括5个图形处理器,通过平均划分处理可以得到5个样本集合,每个样本集合中包括2个数据样本;样本集合和图形处理器之间的对应关系可以是预先保存好的,直接通过获取的方式来确定,除此之外还可以确定出样本集合和图形处理器之间的匹配度,按照任意随机的顺序依次为每个图形处理器确定出与每个样本集合的匹配度,例如针对第一图形处理器A,确定出第一图形处理器A与每个样本集合的匹配度,再将最高匹配度对应的样本集合作为与该图形处理器具有对应关系的样本集合,匹配度与样本集合与第一图形处理器的历史目标样本数据的重复度负相关。
在步骤102中,第一图形处理器从第一类型的全量稀疏特征的嵌入特征中,获取对应于第一稀疏特征的嵌入特征。
作为示例,第一图形处理器从第一类型的全量稀疏特征的嵌入特征中,获取目标样本在第一类型的第一稀疏特征的嵌入特征。
作为示例,全量表征所有的含义,每个稀疏特征对应有一个嵌入特征,相当于所有样本(包括目标样本)在第一类型的稀疏特征的嵌入特征均存储在第一图形处理器,因此这里相当于是从全量的嵌入特征中进行查询,第一图形处理器从第一类型的所有的稀疏特征分别对应的嵌入特征中查询出目标样本在第一类型的第一稀疏特征的嵌入特征,相当于是从第一类型的所有的嵌入特征中查询出目标样本在第一类型的嵌入特征。通过本申请实施例可以直接从本地获取目标样本在第一类型的第一稀疏特征的嵌入特征,提高嵌入特征的获取效率。
例如,第一类型是账号标识类型,这里的全量表征所有的含义,第一图形处理器存储有属于账号标识类型的所有稀疏特征分别对应的嵌入特征,例如,存储有100个账号类型的稀疏特征分别对应的100个嵌入特征,第一图形处理 器从所有属于账号标识类型的嵌入特征中查询出目标样本的嵌入特征,例如对象A在账号标识类型的嵌入特征,这里查询的时候是根据目标样本的稀疏特征进行查询,例如,目标样本(对象A)的稀疏特征是(0,0,0,0,0,0,0,0,1,0,0,0)以稀疏特征为键(key)从哈希表中查询到对应的嵌入特征,例如,这里的嵌入特征可以是数值9,这里的嵌入特征相当于是值(value),稀疏特征和嵌入特征以键值对的方式存储在哈希表中。
在步骤103中,第一图形处理器从第二图形处理器获取从第二类型的全量稀疏特征的嵌入特征查询得到的对应于第二稀疏特征的嵌入特征。
作为示例,第一图形处理器从第二图形处理器,获取目标样本在第二类型的第二稀疏特征的嵌入特征,目标样本在第二类型的第二稀疏特征的嵌入特征是第二图形处理器从第二类型的全量稀疏特征的嵌入特征查询得到的。
在一些实施例中,参见图3B,步骤103中第一图形处理器从第二图形处理器获取从第二类型的全量稀疏特征的嵌入特征查询得到的对应于第二稀疏特征的嵌入特征可以通过图3B示出的步骤1031至步骤1033实现。
在步骤1031中,第一图形处理器向第二图形处理器传输第二稀疏特征。
在步骤1032中,第二图形处理器从第二类型的全量稀疏特征的嵌入特征中查询出对应于第二稀疏特征的嵌入特征。
全量表征所有的含义,每个稀疏特征对应有一个嵌入特征,相当于所有样本(包括目标样本)在第二类型的稀疏特征的嵌入特征均存储在第二图形处理器,因此这里相当于是从全量的嵌入特征中进行查询,第二图形处理器从第二类型的所有的稀疏特征分别对应的嵌入特征中查询出对应于目标样本的第二稀疏特征的嵌入特征,相当于是从第二类型的所有的嵌入特征中查询出目标样本在第二类型的嵌入特征。
在步骤1033中,第一图形处理器接收,对应于第二稀疏特征的嵌入特征。
作为示例,第一类型是账号标识类型,第二类型是年龄段类型,第一图形处理器将目标样本的年龄段类型的稀疏特征传输至第二图形处理器,第二图形处理器从所有属于年龄段类型的所有嵌入特征中查询出目标样本的年龄段类型 的嵌入特征,第二图形处理器将目标样本的年龄段类型的嵌入特征传输至第一图形处理器,第一图形处理器接收目标样本的年龄段类型的嵌入特征。通过本申请实施例可以实现嵌入特征的交换,对于第一图形处理器而言,可以从第二图形处理器获取第一图形处理器本地未存储的嵌入特征,可以在分布式存储的前提下实现嵌入特征的共享,提高嵌入特征的利用率。
在步骤104中,第一图形处理器将对应于第一稀疏特征的嵌入特征和对应于第二稀疏特征的嵌入特征,组成对应于目标样本的稀疏特征的嵌入特征。
作为示例,第一图形处理器将目标样本在第一类型的稀疏特征的嵌入特征以及目标样本在第二类型的稀疏特征的嵌入特征组成目标样本的稀疏特征的嵌入特征,例如,第一图形处理器将目标样本在账号标识类型的稀疏特征的嵌入特征以及目标样本在年龄段类型的稀疏特征的嵌入特征组成目标样本的稀疏特征的嵌入特征,作为后续模型推理的输入。
在步骤105中,第一图形处理器对对应于目标样本的稀疏特征的嵌入特征执行概率映射处理,并根据概率映射结果生成更新指令,更新指令用于指示第二图形处理器更新第二类型的全量稀疏特征的嵌入特征。
作为示例,概率映射是指将嵌入特征转化为概率的映射处理,概率映射处理可以是线性变换处理或者非线性变换处理,即概率映射处理实际上是对嵌入特征进行线性变换或者非线性变换,以将嵌入特征变换为一个概率值,概率映射的方式较多,本申请不具体限定。
作为示例,针对线性变换的概率映射处理,会涉及到权重参数(权重参数的数据形式是列向量)以及偏置参数(标量),将权重参数与嵌入特征(嵌入特征的数据形式是行向量)进行点乘处理,再将点乘结果(这里的点乘结果是标量)与偏置参数相加,最后将相加得到的数值作为映射得到的概率(概率映射结果),也可以对相加得到的数值进行归一化处理,将归一化得到的数值作为概率映射结果。
作为示例,概率映射的过程具体还可以参见公式(1):
其中,x是嵌入特征,-WT是权重参数,b是偏置参数,可以直接将-WTx+b作为概率映射得到的概率,还可以根据公式(1)对-WTx+b执行归一化处理,得到h(x),并将h(x)作为概率映射结果。
在一些实施例中,参见图3C,步骤105中第一图形处理器对对应于目标样本的稀疏特征的嵌入特征执行概率映射处理,可以通过执行图3C示出的步骤1051实现,步骤105中根据概率映射结果生成更新指令可以通过执行图3C示出的步骤1052至步骤1053实现。
在步骤1051中,以对应于目标样本的稀疏特征的嵌入特征作为推荐模型的输入,通过推荐模型对对应于目标样本的稀疏特征的嵌入特征执行概率映射处理,得到对应目标样本的推荐概率作为概率映射结果。
作为示例,推荐模型可以是深度神经网络(DNN,Deep Neural Networks),对于任何一个目标样本,可以对应有多个类型的嵌入特征,将这些类型的嵌入特征组成对应于目标样本的嵌入特征,以对应于目标样本的嵌入特征为推荐模型的输入,推荐模型至少包括全连接层,通过全连接层可以对对应于目标样本的嵌入特征执行概率映射处理(可以是线性变换或者是非线性变换),得到每个目标样本的推荐概率。
在步骤1052中,第一图形处理器确定对应目标样本的推荐概率与对应目标样本的标签值之间的误差,并基于误差确定目标样本的第二类型的全量稀疏特征的嵌入特征的梯度。
作为示例,这里的误差是将推荐概率与标签值代入损失函数得到的理,通过误差计算目标样本的第二类型的全量稀疏特征的嵌入特征的梯度,计算方式包括但不限于以下几种:数值计算法(通过函数求导进行计算)、解析计算法。
在步骤1053中,第一图形处理器生成携带有第二类型的全量稀疏特征的嵌入特征的梯度的更新指令。
作为示例,这里首先生成未携带任何信息的更新指令,用于指示第二图形 处理器对第二类型的全量稀疏特征的嵌入特征进行更新,由于更新需要基于第二类型的全量稀疏特征的嵌入特征的梯度执行,因此将第二类型的全量稀疏特征的嵌入特征的梯度添加至上述未携带任何信息的更新指令。
作为示例,第一图形处理器将更新指令传输至第二图形处理器,第二图形处理器基于第二类型的全量稀疏特征的嵌入特征的梯度对第二类型的全量稀疏特征的嵌入特征进行更新。
通过本申请实施例可以使得第二图形处理器对本地存储的嵌入特征进行更新,相当于实现了分布式存储下的嵌入特征更新,突破了分布式存储造成的信息隔离的壁垒,将训练过程也拆分到不同图形处理器以分布式的方式完成,可以有效提高训练效率。
在一些实施例中,第一图形处理器根据概率映射结果更新第一类型的全量稀疏特征的嵌入特征。具体而言,概率映射结果是推荐概率,第一图形处理器确定出推荐概率与标签值之间的误差,并基于误差确定出第一类型的全量稀疏特征的嵌入特征的梯度,再根据第一类型的全量稀疏特征的嵌入特征的梯度更新第一类型的全量稀疏特征的嵌入特征。通过本申请实施例可以使得第一图形处理器对本地存储的嵌入特征进行更新,由于这里的概率映射结果是由第一图形处理器得到的,因此第一图形处理器可以直接基于概率映射结果更新本地存储的嵌入特征,提高训练效率。
作为示例,TensorFlow提供的是声明式的编程接口,不需要关心求导细节,只需要定义好模型得到损失loss方程,然后使用TensorFlow实现的各种优化器(Op,Optimizer)来进行运算即可更新得到新嵌入特征。TensorFlow的Python代码层面,自动求导的部分是靠各种各样的Op串起来的,TensorFlow的做法是每一个Op在建图的时候就同时包含了它的梯度计算公式,构成前向计算图的时候会自动建立反向部分的计算图,前向计算出来的输入输出会保留下来,留到后向计算的时候用完了才删除,然后在最后加上一个Op(例如GradientDescentOptimizer、AdamOptimizer)。Op都继承自Optimizer这个类,这个类的方法非常多,几个重要方法是minimize、compute_gradients、apply_gradients、sl ot系列。具体梯度如何更新到变量,由_apply_dense、_resource_apply_dense、_apply_sparse、_resource_apply_spars这四个方法实现。
在一些实施例中,本申请实施例提供的数据处理方法,应用于第二图形处理器,第一图形处理器存储有第一类型的全量稀疏特征的嵌入特征,所述第二图形处理器存储有第二类型的全量稀疏特征的嵌入特征,所述第一类型与所述第二类型不同。
参见图3D,图3D是本申请实施例提供的数据处理方法的流程示意图。
在步骤201中,第二图形处理器接收第一图形处理器发送的第二类型的第二稀疏特征
第二图形处理器接收第一图形处理器发送的,目标样本的第二类型的第二稀疏特征。
在步骤202中,第二图形处理器从第二类型的全量稀疏特征的嵌入特征中查询得到对应于第二稀疏特征的嵌入特征。
在步骤203中,第二图形处理器将对应于第二稀疏特征的嵌入特征传输至第一图形处理器。
在步骤204中,所述第一图形处理器对对应于第一稀疏特征的嵌入特征和对应于第二稀疏特征的嵌入特征,执行概率映射处理。
作为示例,对应于第一稀疏特征的嵌入特征是第一图形处理器根据目标样本的第一类型的第一稀疏特征从第一类型的全量稀疏特征的嵌入特征中获取得到的。
作为示例,将对对应于第一稀疏特征的嵌入特征和对应于第二稀疏特征的嵌入特征组成对应于目标样本的稀疏特征的嵌入特征,对对应于目标样本的稀疏特征的嵌入特征执行概率映射处理,得到概率映射结果。
在步骤205中,第二图形处理器从第一图形处理器接收根据概率映射结果生成的更新指令,并基于更新指令更新第二类型的全量稀疏特征的嵌入特征。
在一些实施例中,第一图形处理器和第二图形处理器属于多个图形处理器,多个图形处理器存储有多个类型的全量稀疏特征的嵌入特征,第一类型包含多 个类型中的至少一个类型,第二类型包含多个类型中的至少一个类型。
在一些实施例中,第一类型的全量稀疏特征的嵌入特征的存储位置和第二类型的全量稀疏特征的嵌入特征的存储位置,是在多个类型的全量稀疏特征的嵌入特征的数据量大于阈值的情况下划分的。
在一些实施例中,在多个类型的全量稀疏特征的嵌入特征的数据量不大于阈值的情况下,多个类型的全量稀疏特征的嵌入特征保存在多个图形处理器中的每个图形处理器中。
按照类型将嵌入特征存储在不同的图形处理器中,可以实现嵌入特征的分布式存储,第一图形处理器获取目标样本的不同类型的稀疏特征,第一图形处理器从本地获取第一类型的嵌入特征,并从第二图形处理器,获取第二类型的嵌入特征,按照类型划分方便从对应图形处理器中获取嵌入特征,可以提高获取效率,并且便于特征管理,第一图形处理器对两个类型的嵌入特征进行概率映射处理,并根据概率映射结果向第二图形处理器传输更新指令,更新指令用于指示第二图形处理器更新第二类型的嵌入特征,从而在分布式存储的前提下实现了嵌入特征的更新,充分利用GPU计算的高效性,在满足存储需求的同时极大提升嵌入特征更新速度。
下面,将说明本申请实施例在一个实际的应用场景中的示例性应用。
在一些实施例中,为了使得新闻客户端可以向用户提供符合用户兴趣的新闻推荐,需要训练并部署新闻推荐模型(下文简称推荐模型),推荐模型的输入是目标对象的嵌入特征,推荐模型的输出是目标对象点击某个新闻的概率,当概率大于概率阈值时,新闻客户端向目标对象推送该新闻,在上述过程中嵌入特征是通过对账号ID等稀疏特征进行映射处理得到的,因此嵌入特征对于稀疏特征的表征能力极大影响推荐准确度,所以我们需要对嵌入特征进行更新以使得嵌入特征具有优秀的表达能力。
下面介绍嵌入特征的更新过程,数据样本可以是推荐模型的数据样本,例如针对新闻推荐模型,单个数据样本可以包括登录新闻客户端的账号的相关对象数据,这里的相关对象数据包括账号标识(账号ID)、账号年龄、表征账号 兴趣的属性标签等等。全量数据样本中对应第一类型的稀疏特征的嵌入特征保存在服务器200内部署的GPU1内,全量数据样本中对应第二类型的稀疏特征的嵌入特征保存在服务器200内部署的GPU2内。
服务器200接收目标样本A,这里的目标样本A被输入至GPU1,GPU1获取目标样本A的稀疏特征,目标样本A的稀疏特征包含目标样本A在账号标识类型的稀疏特征和目标样本A在年龄段类型的稀疏特征,GPU1从账号标识类型的全量稀疏特征的嵌入特征中,获取目标样本A在账号标识类型的稀疏特征的嵌入特征;GPU1从GPU2,获取目标样本A在年龄段类型的稀疏特征的嵌入特征,目标样本A在年龄段类型的稀疏特征的嵌入特征是GPU2从年龄段类型的全量稀疏特征的嵌入特征查询得到的;GPU1根据目标样本A在账号标识类型的稀疏特征的嵌入特征和目标样本A在年龄段类型的稀疏特征的嵌入特征,确定目标样本A的稀疏特征的嵌入特征;GPU1根据目标样本A的稀疏特征的嵌入特征正向传播,并根据概率映射结果向GPU2传输更新指令,更新指令用于指示GPU2更新年龄段类型的全量稀疏特征的嵌入特征。
推荐系统会帮用户挑选满意的商品、感兴趣的新闻、最适合的课程、符合用户兴趣的短视频、符合用户需求的音乐,参见图4,在获知“用户信息”、“物品信息”以及“场景信息”的基础上,推荐系统要处理的问题可以较形式化地定义为:对于用户U,在特定场景C下,针对海量物品信息,构建函数f(U,I,C),预测用户对特定候选物品I的喜好程度,再根据喜好程度对所有候选物品进行排序,生成推荐列表。
当采用参数服务(Parameter Server)方式训练时,推荐场景下的Embedding存放在CPU上,embedding更新也在CPU内完成。由CPU进行Embedding层的复杂大型运算操作,这既限制了训练的速度,又导致实际生产中无法使用比较复杂的模型,因为使用复杂模型会导致CPU对给定输入计算时间过长,无法及时响应请求。GPU训练已在图像识别、文字处理等应用上取得巨大成功。GPU训练以其在卷积等数学运算上的独特效率优势,极大地提升了训练深度神经网络的速度。然而在推荐场景中,由于存在大量的稀疏样本(例如用户ID), 每个用户ID在输入推荐模型之前需要被映射为对应的embedding向量,因此完整的推荐模型通常体积巨大以至于单GPU无法进行存储。相关技术中要么为了利用高吞吐电子设备做计算而受限于电子设备的存储容量(比如GPU的显存),导致embedding规模难以形成大规模,要么为了容纳大规模的embedding(比如TB级),不得不使用CPU训练,无法发挥数据并行训练模式的性能优势。
针对推荐系统工程化的特点,分布式可扩展的参数服务(Parameter Server)方案被提出,几乎完美地解决了机器学习模型的分布式训练问题,Parameter Server不仅被直接应用在机器学习平台上,也被集成在TensorF1ow、MXNet等主流的深度学习框架中,作为机器学习分布式训练重要的解决方案。参见图5,可以看出,Parameter Server由服务器节点(server节点)和工作节点(worker节点)组成,server节点的主要功能是保存模型参数、接受worker节点计算出的局部梯度、汇总计算全局梯度,并更新模型参数。worker节点的主要功能是保存部分训练数据,从server节点拉取最新的模型参数,根据训练数据计算局部梯度,上传给server节点。参见图6,TensorFlow的单机训练是在单个worker节点上进行的,worker节点内部按照任务关系图的方式在不同GPU和CPU节点间进行并行计算。
下面介绍大规模embedding,深度学习网络中embedding层的作用是将稀疏输入向量转换成稠密向量,但embedding层的存在往往会拖慢整个神经网络的收敛速度。embedding层的参数数量巨大,假设输入层的维度是100000,embedding层输出维度是32,再加5层32维的全连接层,最后输出层维度是10,那么输入层到embedding层的参数数量是3200000,其余所有层的参数总数是4416,embedding层的权重总数占比是99.86%,也就是说embedding层的权重占了整个网络权重的绝大部分。那么训练过程可想而知,大部分的训练时间和计算开销都被embedding层占据;由于输入特征过于稀疏,在随机梯度下降的过程中,只有与非零特征相连的embedding层权重(嵌入特征)会被更新,这进一步降低了embedding层的收敛速度。
大规模稀疏离散特征的使用,导致深度模型的embedding层的数据量急剧 膨胀,数TB大小的推荐模型一度流行于业界各大头部业务场景。在大多数情况下,由于推荐类的任务对延时非常敏感,粗排场景的模型结构通常比较“矮胖”,有大量的特征输入然后映射为embedding向量,神经网络部分的深度相对较低。因此,特征样本内容的解析和embedding在整体耗时中的占比相对较高。当采用Parameter Server方式训练时,推荐场景的embedding存放在CPU上,embedding训练也在CPU端完成,既限制了训练速度,又因为embedding在整体耗时中的占比相对较高导致,从而无法支持深度较高的模型的计算,因此实际生产中无法使用比较复杂的模型,因为使用复杂模型会导致CPU对给定输入计算时间过长,无法及时响应请求。
本申请实施例提供的数据处理方法支持大规模embedding的GPU训练,本申请实施例提供基于GPU系统的训练方案,可以选用机型为NVIDIAV100,服务器的硬件拓扑包含8张GPU,并且充分考虑硬件的特性以充分发挥性能的优势。GPU系统主要包括三个核心模块:数据模块、计算模块以及通信模块。数据模块依托Tensorflow的数据处理API完成预读取(prefetch)以及转储(dumpfile)等逻辑;计算模块用于控制每张GPU启动一个TensorFlow训练进程执行训练;在通信模块中,使用Horovod进程来做分布式训练的卡间通信,在每个节点上启动一个Horovod进程来执行对应的通信任务。
参见图7,GPU0和GPU1分别通过各自的IO接口执行prefetch和dumpfile的处理逻辑,从而GPU0获取输入的样本0,GPU1获取输入的样本1,GPU0对样本0进行特征解析,得到第一类型的稀疏特征0和第二类型的稀疏特征1,GPU1对样本1进行特征解析,得到第一类型的稀疏特征0和第二类型的稀疏特征1,GPU0中本地存储有所有样本的第一类型的稀疏特征0,GPU1中本地存储有所有样本的第二类型的稀疏特征1,因此GPU0将样本0的第二类型的稀疏特征1传输至GPU1,GPU1将样本1的第一类型的稀疏特征0传输至GPU0。GPU0基于样本0的第一类型的稀疏特征0以及样本1的第一类型的稀疏特征0在哈希表0中进行查找,得到样本0的第一类型的嵌入特征0以及样本1的第一类型的嵌入特征0,GPU1基于样本0的第二类型的稀疏特征1以及样 本1的第二类型的稀疏特征1在哈希表1中进行查找,得到样本0的第二类型的嵌入特征1以及样本1的第二类型的嵌入特征1,上述稀疏特征是键(key),例如,key1、key2和key3,嵌入特征是值(value),例如,value1、value2、value3)。GPU0将样本1的第一类型的稀疏特征0传输至GPU1,GPU1将样本0的第二类型的稀疏特征1传输至GPU0,从而GPU0根据样本0的第一类型的嵌入特征0和样本0的第二类型的嵌入特征1执行基于深度模型的推理处理,GPU1根据样本1的第一类型的嵌入特征0和样本1的第二类型的嵌入特征1执行基于深度模型的推理处理。
本申请实施例的整个训练流程涉及embedding存储、稀疏特征处理与embedding切分、卡间通信等几个关键部分。
首先介绍embedding参数存储,在推荐系统领域,embedding已成为处理身份类稀疏特征的常用手段,作为一种“函数映射”,Embedding通常将高维稀疏特征映射为低维稠密向量,再进行模型端到端训练。在TensorFlow框架中,全部以稠密Tensor为基本数据单元对数据进行计算、存储以及传输。TensorFlow也基于稠密Tensor提供了静态Embedding机制,用于存储Embedding的Tensorshape固定为[vocabulary_size,embedding_dimension],vocabulary_size通常由身份空间决定。本申请实施例利用组件tfra.dynamic_embedding作为embedding的存储方式,该组件用tf.lookup.MutableHashTable保存多个参数(fullweights),并且可以复用tensorflow原生的优化器。
接下来介绍稀疏特征处理与embeddding切分,在实际业务场景中,通常稀疏特征的总量较多,需要的处理较为复杂。本申请实施例针对稀疏特征使用tfra组件的动态embedding数据结构存储。当稀疏特征参数量较大,会导致对应的嵌入特征的数量较大(例如,大于阈值),GPU单卡显存无法容纳时,可以通过按照类型划分的方式将稀疏特征对应的embedding划分到多张GPU卡的显存中存放;当稀疏特征的总数量较少,则意味着对应的嵌入特征的总数量较小(例如不大于阈值),GPU单卡显存可以容纳时,使用Replica的方式,每张GPU卡的显存都放置一份全量的嵌入特征。
最后介绍卡间通信,稀疏特征(例如,身份标识类特征)由于规模较大,因此使用哈希表存储稀疏特征,由于每张GPU的输入样本数据不同,因此输入的稀疏特征对应的embedding特征向量,可能存放在其他GPU卡上。在训练的前向传播过程中,通过GPU间AllToAll通信,每张卡在内部的哈希表中查询稀疏特征对应的embedding特征向量。然后再通过GPU间AllToAll通信,将第一次AllToAll从其他GPU上拿到的特征向量原路返回,通过两次GPU间AllToAll通信,每张GPU的样本输入的稀疏特征都可以拿到对应的嵌入特征。在训练的反向更新过程中,会再次通过GPU间AllToAll通信,将嵌入特征的梯度返回至其他GPU中,每张GPU获取到各自的梯度后再通过优化器执行优化处理,完成大规模嵌入特征的优化。
参见图8,GPU0对样本0进行特征解析,得到第一类型的稀疏特征0和第二类型的稀疏特征1,GPU1对样本1进行特征解析,得到第一类型的稀疏特征0和第二类型的稀疏特征1,GPU0中本地存储有所有样本的第一类型的稀疏特征0,GPU1中本地存储有所有样本的第二类型的稀疏特征1,因此GPU0将样本0的第二类型的稀疏特征1传输至GPU1,GPU1将样本1的第一类型的稀疏特征0传输至GPU0。GPU0基于样本0的第一类型的稀疏特征0以及样本1的第一类型的稀疏特征0在哈希表0中进行查找,得到样本0的第一类型的嵌入特征0以及样本1的第一类型的嵌入特征0,GPU1基于样本0的第二类型的稀疏特征1以及样本1的第二类型的稀疏特征1在哈希表1中进行查找,得到样本0的第二类型的嵌入特征1以及样本1的第二类型的嵌入特征1。GPU0将样本1的第一类型的稀疏特征0传输至GPU1,GPU1将样本0的第二类型的稀疏特征1传输至GPU0,从而GPU0根据样本0的第一类型的嵌入特征0和样本0的第二类型的嵌入特征1执行基于深度模型的推理处理,得到概率映射结果,基于概率映射结果与标签值的误差确定出对应第一类型的嵌入特征0的梯度0以及对应第二类型的嵌入特征1的梯度1,GPU1根据样本1的第一类型的嵌入特征0和样本1的第二类型的嵌入特征1执行基于深度模型的推理处理,得到概率映射结果,基于概率映射结果与标签值的误差确定出对应第一类型的 嵌入特征0的梯度0以及对应第二类型的嵌入特征1的梯度1。GPU0将对应第二类型的嵌入特征1的梯度1传输至GPU1,GPU1的优化器基于对应第二类型的嵌入特征1的梯度1更新第二类型的嵌入特征1,GPU1将对应第一类型的嵌入特征0的梯度0传输至GPU0,GPU0的优化器基于对应第一类型的嵌入特征0的梯度0更新第一类型的嵌入特征0。
现有的推荐系统工程技术方案要么为了利用高吞吐设备做计算,而受限于设备的存储容量(比如GPU的显存),导致embedding规模难以scale。要么为了容纳大规模的embedding(比如TB级),不得不使用CPU训练,无法发挥数据并行训练模式的性能优势。
本申请实施例提出了一种能支持大规模embedding的GPU训练方案,具有如下优点:使用GPU存储embedding参数,充分利用GPU训练的高效,极大提升模型的训练速度;支持大规模embedding参数存储,当embedding参数单卡GPU显存放不下时,通过高效的GPUNCCL通信,实现embedding拆分存储逻辑;不失易用性,可以无缝的对接第三方框架,如TensorFlow、Pytorch等主流框架。
可以理解的是,在本申请实施例中,涉及到用户信息等相关的数据,当本申请实施例运用到具体产品或技术中时,需要获得用户许可或者同意,且相关数据的收集、使用和处理需要遵守相关国家和地区的相关法律法规和标准。
下面继续说明本申请实施例提供的基于人工智能的数据处理装置255-1的实施为软件模块的示例性结构,数据处理方法应用于第二图形处理器,第一图形处理器存储有第一类型的全量稀疏特征的嵌入特征,第二图形处理器存储有第二类型的全量稀疏特征的嵌入特征,第一类型与第二类型不同,在一些实施例中,如图2所示,存储在存储器250的基于人工智能的数据处理装置255-1中的软件模块可以包括:第一接收模块2551,配置为第一图形处理器获取目标样本的稀疏特征,目标样本的稀疏特征包含所述第一类型的第一稀疏特征和所述第二类型的第二稀疏特征;第一获取模块2552,配置为所述第一图形处理器从所述第二图形处理器获取从所述第二类型的全量稀疏特征的嵌入特征查询得 到的对应于所述第二稀疏特征的嵌入特征;第一返回模块2553,配置为所述第一图形处理器从所述第二图形处理器获取从所述第二类型的全量稀疏特征的嵌入特征查询得到的对应于所述第二稀疏特征的嵌入特征;第一确定模块2554,配置为所述第一图形处理器将对应于所述第一稀疏特征的嵌入特征和对应于所述第二稀疏特征的嵌入特征,组成对应于所述目标样本的稀疏特征的嵌入特征;第一更新模块2555,配置为所述第一图形处理器对所述对应于所述目标样本的稀疏特征的嵌入特征执行概率映射处理,并根据概率映射结果生成更新指令,所述更新指令用于指示所述第二图形处理器更新所述第二类型的全量稀疏特征的嵌入特征。
在一些实施例中,第一更新模块2555,还配置为第一图形处理器根据概率映射结果更新第一类型的全量稀疏特征的嵌入特征。
在一些实施例中,第一图形处理器和第二图形处理器属于多个图形处理器,多个图形处理器存储有多个类型的全量稀疏特征的嵌入特征,第一类型包含多个类型中的至少一个类型,第二类型包含多个类型中的至少一个类型。
在一些实施例中,第一类型的全量稀疏特征的嵌入特征的存储位置和第二类型的全量稀疏特征的嵌入特征的存储位置,是在多个类型的全量稀疏特征的嵌入特征的数据量大于阈值的情况下划分的。
在一些实施例中,在多个类型的全量稀疏特征的嵌入特征的数据量不大于阈值的情况下,多个类型的全量稀疏特征的嵌入特征保存在多个图形处理器中的每个图形处理器中。
在一些实施例中,第一返回模块2553,还配置为:所述第一图形处理器向所述第二图形处理器传输所述第二稀疏特征,使得所述第二图形处理器从所述第二类型的全量稀疏特征的嵌入特征中查询出对应于所述第二稀疏特征的嵌入特征;所述第一图形处理器接收对应于所述第二稀疏特征的嵌入特征。
在一些实施例中,第一返回模块2553,还配置为:所述第一图形处理器从所述第一类型的全量稀疏特征的嵌入特征中查询出对应所述第一稀疏特征的嵌入特征。
在一些实施例中,在第一图形处理器获取目标样本的稀疏特征之前,第一接收模块2551,还配置为获取第一图形处理器前一次从多个数据样本中获取到的历史目标样本;将多个数据样本中除历史目标样本的数据样本作为其他数据样本;从其他数据样本中随机获取至少一个数据样本作为目标样本。
在一些实施例中,在第一图形处理器获取目标样本的稀疏特征之前,第一接收模块2551,还配置为基于图形处理器的数目对多个数据样本进行平均划分处理,得到多个样本集合,其中,每个样本集合包括至少一个数据样本,样本集合的集合数目是与图形处理器的数目相同;确定多个样本集合与多个图形处理器之间的一一对应关系;将与第一图形处理器之间具有对应关系的样本集合中的数据样本,作为对应第一图形处理器的目标样本。
在一些实施例中,第一更新模块2555,配置为以所述对应于所述目标样本的稀疏特征的嵌入特征作为推荐模型的输入,通过所述推荐模型对所述对应于所述目标样本的稀疏特征的嵌入特征执行概率映射处理,得到对应所述目标样本的推荐概率作为所述概率映射结果;所述第一图形处理器确定对应所述目标样本的推荐概率与对应所述目标样本的标签值之间的误差,并基于所述误差确定所述目标样本的第二类型的全量稀疏特征的嵌入特征的梯度;生成携带有所述第二类型的全量稀疏特征的嵌入特征的梯度的更新指令。
下面继续说明本申请实施例提供的基于人工智能的数据处理装置255-2的实施为软件模块的示例性结构,数据处理方法应用于第二图形处理器,第一图形处理器存储有第一类型的全量稀疏特征的嵌入特征,第二图形处理器存储有第二类型的全量稀疏特征的嵌入特征,第一类型与第二类型不同,在一些实施例中,如图2所示,存储在存储器250的基于人工智能的数据处理装置255-2中的软件模块可以包括:第二接收模块2556,配置为所述第二图形处理器接收所述第一图形处理器发送的,所述第二类型的第二稀疏特征;查询模块2557,配置为所述第二图形处理器从所述第二类型的全量稀疏特征的嵌入特征中查询得到对应于所述第二稀疏特征的嵌入特征;第二返回模块2558,配置为所述第二图形处理器将对应于所述第二稀疏特征的嵌入特征传输至所述第一图形处理 器,以使所述第一图形处理器对对应于所述第一类型的第一稀疏特征的嵌入特征和对应于所述第二稀疏特征的嵌入特征,执行概率映射处理,其中,对应于所述第一稀疏特征的嵌入特征是所述第一图形处理器从所述第一类型的全量稀疏特征的嵌入特征中获取得到的;第二更新模块2559,配置为所述第二图形处理器从第一图形处理器接收根据概率映射结果生成的更新指令,并基于所述更新指令更新所述第二类型的全量稀疏特征的嵌入特征。
本申请实施例提供一种图形处理器,用于执行本申请实施例上述的基于人工智能的数据处理方法。
本申请实施例提供了一种计算机程序产品,该计算机程序产品包括计算机可执行指令,该计算机可执行指令存储在计算机可读存储介质中。电子设备的处理器从计算机可读存储介质读取该计算机可执行指令,处理器执行该计算机可执行指令,使得该电子设备执行本申请实施例上述的基于人工智能的数据处理方法,处理器是图形处理器或者中央处理器。
本申请实施例提供一种存储有计算机可执行指令的计算机可读存储介质,其中存储有计算机可执行指令,当计算机可执行指令被处理器执行时,将引起处理器执行本申请实施例提供的基于人工智能的数据处理方法,例如,如图3A-图3D示出的基于人工智能的数据处理方法,处理器是图形处理器或者中央处理器。
在一些实施例中,计算机可读存储介质可以是FRAM、ROM、PROM、EPROM、EEPROM、闪存、磁表面存储器、光盘、或CD-ROM等存储器;也可以是包括上述存储器之一或任意组合的各种设备。
在一些实施例中,计算机可执行指令可以采用程序、软件、软件模块、脚本或代码的形式,按任意形式的编程语言(包括编译或解释语言,或者声明性或过程性语言)来编写,并且其可按任意形式部署,包括被部署为独立的程序或者被部署为模块、组件、子例程或者适合在计算环境中使用的其它单元。
作为示例,计算机可执行指令可以但不一定对应于文件系统中的文件,可以可被存储在保存其它程序或数据的文件的一部分,例如,存储在超文本标记 语言(HTML,Hyper Text Markup Language)文档中的一个或多个脚本中,存储在专用于所讨论的程序的单个文件中,或者,存储在多个协同文件(例如,存储一个或多个模块、子程序或代码部分的文件)中。
作为示例,计算机可执行指令可被部署为在一个电子设备上执行,或者在位于一个地点的多个电子设备上执行,又或者,在分布在多个地点且通过通信网络互连的多个电子设备上执行。
综上所述,按照类型将嵌入特征存储在不同的图形处理器中,可以实现嵌入特征的分布式存储,第一图形处理器获取目标样本的不同类型的稀疏特征,第一图形处理器从本地获取第一类型的嵌入特征,并从第二图形处理器,获取第二类型的嵌入特征,按照类型划分方便从对应图形处理器中获取嵌入特征,可以提高获取效率,并且便于特征管理,第一图形处理器对两个类型的嵌入特征进行概率映射处理,并根据概率映射结果向第二图形处理器传输更新指令,更新指令用于指示第二图形处理器更新第二类型的嵌入特征,从而在分布式存储的前提下实现了嵌入特征的更新,充分利用GPU计算的高效性,在满足存储需求的同时极大提升嵌入特征更新速度。
以上所述,仅为本申请的实施例而已,并非用于限定本申请的保护范围。凡在本申请的精神和范围之内所作的任何修改、等同替换和改进等,均包括在本申请的保护范围之内。

Claims (20)

  1. 一种数据处理方法,所述数据处理方法应用于第一图形处理器,所述第一图形处理器存储有第一类型的全量稀疏特征的嵌入特征,第二图形处理器存储有第二类型的全量稀疏特征的嵌入特征,所述第一类型与所述第二类型不同;
    所述方法包括:
    所述第一图形处理器获取目标样本的稀疏特征,所述目标样本的稀疏特征包含所述第一类型的第一稀疏特征和所述第二类型的第二稀疏特征;
    所述第一图形处理器从所述第一类型的全量稀疏特征的嵌入特征中,获取对应于所述第一稀疏特征的嵌入特征;
    所述第一图形处理器从所述第二图形处理器获取从所述第二类型的全量稀疏特征的嵌入特征查询得到的对应于所述第二稀疏特征的嵌入特征;
    所述第一图形处理器将对应于所述第一稀疏特征的嵌入特征和对应于所述第二稀疏特征的嵌入特征,组成对应于所述目标样本的稀疏特征的嵌入特征;
    所述第一图形处理器对所述对应于所述目标样本的稀疏特征的嵌入特征执行概率映射处理,并根据概率映射结果生成更新指令,所述更新指令用于指示所述第二图形处理器更新所述第二类型的全量稀疏特征的嵌入特征。
  2. 根据权利要求1所述的方法,其中,所述方法还包括:
    所述第一图形处理器根据所述概率映射结果更新所述第一类型的全量稀疏特征的嵌入特征。
  3. 根据权利要求1或2所述的方法,其中,所述第一图形处理器和所述第二图形处理器属于多个图形处理器,所述多个图形处理器存储有多个类型的全量稀疏特征的嵌入特征,所述第一类型包含所述多个类型中的至少一个类型,所述第二类型包含所述多个类型中的至少一个类型。
  4. 根据权利要求3所述的方法,其中,所述第一类型的全量稀疏特征的嵌入特征的存储位置和所述第二类型的全量稀疏特征的嵌入特征的存储位置,是在所述多个类型的全量稀疏特征的嵌入特征的数据量大于阈值的情况下划分的。
  5. 根据权利要求3所述的方法,其中,在所述多个类型的全量稀疏特征的 嵌入特征的数据量不大于阈值的情况下,所述多个类型的全量稀疏特征的嵌入特征保存在所述多个图形处理器中的每个图形处理器中。
  6. 根据权利要求1至5中任意一项所述的方法,其中,所述第一图形处理器从所述第二图形处理器获取从所述第二类型的全量稀疏特征的嵌入特征查询得到的对应于所述第二稀疏特征的嵌入特征,包括:
    所述第一图形处理器向所述第二图形处理器传输所述第二稀疏特征,使得所述第二图形处理器从所述第二类型的全量稀疏特征的嵌入特征中查询出对应于所述第二稀疏特征的嵌入特征;
    所述第一图形处理器接收对应于所述第二稀疏特征的嵌入特征。
  7. 根据权利要求1至6中任意一项所述的方法,其中,所述第一图形处理器从所述第一类型的全量稀疏特征的嵌入特征中,获取对应于所述第一稀疏特征的嵌入特征,包括:
    所述第一图形处理器从所述第一类型的全量稀疏特征的嵌入特征中查询出对应所述第一稀疏特征的嵌入特征。
  8. 根据权利要求1至7中任意一项所述的方法,其中,在所述第一图形处理器获取目标样本的稀疏特征之前,所述方法还包括:
    获取所述第一图形处理器前一次从所述多个数据样本中获取到的历史目标样本;
    将所述多个数据样本中除所述历史目标样本的数据样本作为其他数据样本;
    从所述其他数据样本中随机获取至少一个数据样本作为所述目标样本。
  9. 根据权利要求1至8中任意一项所述的方法,其中,在所述第一图形处理器获取目标样本的稀疏特征之前,所述方法还包括:
    基于所述图形处理器的数目对所述多个数据样本进行平均划分处理,得到多个样本集合,其中,每个所述样本集合包括至少一个数据样本,所述样本集合的集合数目与所述图形处理器的数目相同;
    确定所述多个样本集合与所述多个图形处理器之间的一一对应关系;
    将与所述第一图形处理器之间具有所述对应关系的样本集合中的数据样本, 作为对应所述第一图形处理器的目标样本。
  10. 根据权利要求1至9中任意一项所述的方法,其中,
    所述第一图形处理器根据所述对应于所述目标样本的稀疏特征的嵌入特征执行概率映射处理,包括:
    以所述对应于所述目标样本的稀疏特征的嵌入特征作为推荐模型的输入,通过所述推荐模型对所述对应于所述目标样本的稀疏特征的嵌入特征执行概率映射处理,得到对应所述目标样本的推荐概率作为所述概率映射结果;
    所述根据概率映射结果生成更新指令,包括:
    所述第一图形处理器确定对应所述目标样本的推荐概率与对应所述目标样本的标签值之间的误差,并基于所述误差确定所述目标样本的第二类型的全量稀疏特征的嵌入特征的梯度;
    生成携带有所述第二类型的全量稀疏特征的嵌入特征的梯度的更新指令。
  11. 一种数据处理方法,所述数据处理方法应用于第二图形处理器,第一图形处理器存储有第一类型的全量稀疏特征的嵌入特征,所述第二图形处理器存储有第二类型的全量稀疏特征的嵌入特征,所述第一类型与所述第二类型不同;
    所述方法包括:
    所述第二图形处理器接收所述第一图形处理器发送的,所述第二类型的第二稀疏特征;
    所述第二图形处理器从所述第二类型的全量稀疏特征的嵌入特征中查询得到对应于所述第二稀疏特征的嵌入特征;
    所述第二图形处理器将对应于所述第二稀疏特征的嵌入特征传输至所述第一图形处理器,以使所述第一图形处理器对对应于第一稀疏特征的嵌入特征和对应于所述第二稀疏特征的嵌入特征,执行概率映射处理,其中,所述对应于第一稀疏特征的嵌入特征是所述第一图形处理器从所述第一类型的全量稀疏特征的嵌入特征中获取得到的;
    所述第二图形处理器从第一图形处理器接收根据概率映射结果生成的更新 指令,并基于所述更新指令更新所述第二类型的全量稀疏特征的嵌入特征。
  12. 根据权利要求11所述的方法,其中,所述第一图形处理器和所述第二图形处理器属于多个图形处理器,所述多个图形处理器存储有多个类型的全量稀疏特征的嵌入特征,所述第一类型包含所述多个类型中的至少一个类型,所述第二类型包含所述多个类型中的至少一个类型。
  13. 根据权利要求12所述的方法,其中,所述第一类型的全量稀疏特征的嵌入特征的存储位置和所述第二类型的全量稀疏特征的嵌入特征的存储位置,是在所述多个类型的全量稀疏特征的嵌入特征的数据量大于阈值的情况下划分的。
  14. 根据权利要求12所述的方法,其中,在所述多个类型的全量稀疏特征的嵌入特征的数据量不大于阈值的情况下,所述多个类型的全量稀疏特征的嵌入特征保存在所述多个图形处理器中的每个图形处理器中。
  15. 一种数据处理装置,所述数据处理方法应用于第一图形处理器,第一图形处理器存储有第一类型的全量稀疏特征的嵌入特征,所述第二图形处理器存储有第二类型的全量稀疏特征的嵌入特征,所述第一类型与所述第二类型不同,所述装置包括:
    第一接收模块,配置为所述第一图形处理器获取目标样本的稀疏特征,所述目标样本的稀疏特征包含所述第一类型的第一稀疏特征和所述第二类型的第二稀疏特征;
    第一获取模块,配置为所述第一图形处理器从所述第二图形处理器获取从所述第二类型的全量稀疏特征的嵌入特征查询得到的对应于所述第二稀疏特征的嵌入特征;
    第一返回模块,配置为所述第一图形处理器从所述第二图形处理器获取从所述第二类型的全量稀疏特征的嵌入特征查询得到的对应于所述第二稀疏特征的嵌入特征;
    第一确定模块,配置为所述第一图形处理器将对应于所述第一稀疏特征的嵌入特征和对应于所述第二稀疏特征的嵌入特征,组成对应于所述目标样本的 稀疏特征的嵌入特征;
    第一更新模块,配置为所述第一图形处理器对所述对应于所述目标样本的稀疏特征的嵌入特征执行概率映射处理,并根据概率映射结果生成更新指令,所述更新指令用于指示所述第二图形处理器更新所述第二类型的全量稀疏特征的嵌入特征。
  16. 一种数据处理装置,所述数据处理方法应用于第二图形处理器,第一图形处理器存储有第一类型的全量稀疏特征的嵌入特征,所述第二图形处理器存储有第二类型的全量稀疏特征的嵌入特征,所述第一类型与所述第二类型不同;
    所述装置包括:
    第二接收模块,配置为所述第二图形处理器接收所述第一图形处理器发送的,所述第二类型的第二稀疏特征;
    查询模块,配置为所述第二图形处理器从所述第二类型的全量稀疏特征的嵌入特征中查询得到对应于所述第二稀疏特征的嵌入特征;
    第二返回模块,配置为所述第二图形处理器将对应于所述第二稀疏特征的嵌入特征传输至所述第一图形处理器,以使所述第一图形处理器对对应于第一稀疏特征的嵌入特征和对应于所述第二稀疏特征的嵌入特征,执行概率映射处理,其中,所述对应于第一稀疏特征的嵌入特征是所述第一图形处理器从所述第一类型的全量稀疏特征的嵌入特征中获取得到的;
    第二更新模块,配置为所述第二图形处理器从第一图形处理器接收根据概率映射结果生成的更新指令,并基于所述更新指令更新所述第二类型的全量稀疏特征的嵌入特征。
  17. 一种图形处理器,所述图形处理器用于执行权利要求1至10或者权利要求11至14任一项所述的数据处理方法。
  18. 一种电子设备,所述电子设备包括:
    存储器,用于存储计算机可执行指令;
    图形处理器,用于执行所述存储器中存储的计算机可执行指令时,实现权利要求1至10或者权利要求11至14任一项所述的数据处理方法。
  19. 一种计算机可读存储介质,存储有计算机可执行指令,所述计算机可执行指令被处理器执行时实现权利要求1至10或者权利要求11至14任一项所述的数据处理方法,所述处理区是中央处理器或者图形处理器。
  20. 一种计算机程序产品,包括计算机可执行指令,所述计算机可执行指令被处理器执行时实现权利要求1至10或者权利要求11至14任一项所述的数据处理方法,所述处理区是中央处理器或者图形处理器。
PCT/CN2023/135891 2023-04-28 2023-12-01 数据处理方法、装置、电子设备、计算机可读存储介质及计算机程序产品 Ceased WO2024221925A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
EP23935092.9A EP4614327A4 (en) 2023-04-28 2023-12-01 DATA PROCESSING METHOD AND APPARATUS, ELECTRONIC DEVICE, COMPUTER-READABLE STORAGE MEDIA AND COMPUTER PROGRAM PRODUCT
US19/235,945 US20250307980A1 (en) 2023-04-28 2025-06-12 Data processing method and apparatus, electronic device, computer-readable storage medium, and computer program product

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202310492698.2A CN116974781A (zh) 2023-04-28 2023-04-28 数据处理方法、装置、设备、存储介质及程序产品
CN202310492698.2 2023-04-28

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US19/235,945 Continuation US20250307980A1 (en) 2023-04-28 2025-06-12 Data processing method and apparatus, electronic device, computer-readable storage medium, and computer program product

Publications (1)

Publication Number Publication Date
WO2024221925A1 true WO2024221925A1 (zh) 2024-10-31

Family

ID=88478519

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2023/135891 Ceased WO2024221925A1 (zh) 2023-04-28 2023-12-01 数据处理方法、装置、电子设备、计算机可读存储介质及计算机程序产品

Country Status (4)

Country Link
US (1) US20250307980A1 (zh)
EP (1) EP4614327A4 (zh)
CN (1) CN116974781A (zh)
WO (1) WO2024221925A1 (zh)

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN116974781A (zh) * 2023-04-28 2023-10-31 腾讯科技(深圳)有限公司 数据处理方法、装置、设备、存储介质及程序产品
CN118244997B (zh) * 2024-05-28 2024-08-30 山东云海国创云计算装备产业创新中心有限公司 一种固态硬盘数据处理方法、装置、电子设备及存储介质
CN121029778B (zh) * 2025-10-30 2026-02-24 浙江星汉信息技术股份有限公司 基于稀疏注意力的大模型底层数据处理方法及系统

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109919436A (zh) * 2019-01-29 2019-06-21 华融融通(北京)科技有限公司 一种基于稀疏特征嵌入的违约用户概率预测方法
CN112884157A (zh) * 2019-11-29 2021-06-01 北京达佳互联信息技术有限公司 一种模型训练方法、模型训练节点及参数服务器
US20210319078A1 (en) * 2020-04-14 2021-10-14 Microsoft Technology Licensing, Llc Set operations using multi-core processing unit
CN113971428A (zh) * 2020-07-24 2022-01-25 北京达佳互联信息技术有限公司 数据处理方法、系统、设备、程序产品及存储介质
CN116974781A (zh) * 2023-04-28 2023-10-31 腾讯科技(深圳)有限公司 数据处理方法、装置、设备、存储介质及程序产品

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN115423092B (zh) * 2022-08-26 2025-12-19 北京潞晨科技有限公司 一种基于分布式异构计算的大规模推荐系统训练方法

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109919436A (zh) * 2019-01-29 2019-06-21 华融融通(北京)科技有限公司 一种基于稀疏特征嵌入的违约用户概率预测方法
CN112884157A (zh) * 2019-11-29 2021-06-01 北京达佳互联信息技术有限公司 一种模型训练方法、模型训练节点及参数服务器
US20210319078A1 (en) * 2020-04-14 2021-10-14 Microsoft Technology Licensing, Llc Set operations using multi-core processing unit
CN113971428A (zh) * 2020-07-24 2022-01-25 北京达佳互联信息技术有限公司 数据处理方法、系统、设备、程序产品及存储介质
CN116974781A (zh) * 2023-04-28 2023-10-31 腾讯科技(深圳)有限公司 数据处理方法、装置、设备、存储介质及程序产品

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See also references of EP4614327A4 *

Also Published As

Publication number Publication date
CN116974781A (zh) 2023-10-31
US20250307980A1 (en) 2025-10-02
EP4614327A4 (en) 2026-03-25
EP4614327A1 (en) 2025-09-10

Similar Documents

Publication Publication Date Title
US11416268B2 (en) Aggregate features for machine learning
CN111368210B (zh) 基于人工智能的信息推荐方法、装置以及电子设备
US9934260B2 (en) Streamlined analytic model training and scoring system
KR102794506B1 (ko) 자연어 질문에 대해 검색 증강 생성(rag) 활용을 기초로 챗봇 서비스를 제공하는 시스템
CN114154048B (zh) 构建推荐模型的方法、装置、电子设备及存储介质
US20250307980A1 (en) Data processing method and apparatus, electronic device, computer-readable storage medium, and computer program product
WO2020258487A1 (zh) 一种问答关系排序方法、装置、计算机设备及存储介质
CN114329029B (zh) 对象检索方法、装置、设备及计算机存储介质
WO2025026013A1 (zh) 多模态的任务处理和对话任务处理的方法、系统及设备
CN111680799B (zh) 用于处理模型参数的方法和装置
Li et al. A fast distributed stochastic gradient descent algorithm for matrix factorization
CN110334067A (zh) 一种稀疏矩阵压缩方法、装置、设备及存储介质
CN116643814A (zh) 模型库构建方法、基于模型库的模型调用方法和相关设备
CN112114968A (zh) 推荐方法、装置、电子设备及存储介质
US20250238411A1 (en) Database index recommendation generation by large language model
US20250110979A1 (en) Distributed orchestration of natural language tasks using a generate machine learning model
CN116263659A (zh) 数据处理方法、装置、计算机程序产品、设备及存储介质
CN114282002A (zh) 基于人工智能的知识生成方法、装置、设备及存储介质
CN116126856B (zh) 数据库业务处理方法、装置、计算机设备和存储介质
CN118861203A (zh) 一种基于向量数据库的文本搜索方法、系统、设备及介质
CN115599987B (zh) 业务处理方法、装置、设备及存储介质
HK40098123A (zh) 数据处理方法、装置、设备、存储介质及程序产品
US11544240B1 (en) Featurization for columnar databases
US20200372108A1 (en) Natural language skill generation for digital assistants
US20260072947A1 (en) Automated determining of metadata tags

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23935092

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 2023935092

Country of ref document: EP

ENP Entry into the national phase

Ref document number: 2023935092

Country of ref document: EP

Effective date: 20250603

WWP Wipo information: published in national office

Ref document number: 2023935092

Country of ref document: EP

NENP Non-entry into the national phase

Ref country code: DE