WO2020029592A1 - 转换方法、装置、计算机设备及存储介质 - Google Patents

转换方法、装置、计算机设备及存储介质 Download PDF

Info

Publication number
WO2020029592A1
WO2020029592A1 PCT/CN2019/080510 CN2019080510W WO2020029592A1 WO 2020029592 A1 WO2020029592 A1 WO 2020029592A1 CN 2019080510 W CN2019080510 W CN 2019080510W WO 2020029592 A1 WO2020029592 A1 WO 2020029592A1
Authority
WO
WIPO (PCT)
Prior art keywords
model
attribute information
computer device
offline model
initial
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2019/080510
Other languages
English (en)
French (fr)
Inventor
刘少礼
梁军
郭崎
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Cambricon Technologies Corp Ltd
Original Assignee
Cambricon Technologies Corp Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Cambricon Technologies Corp Ltd filed Critical Cambricon Technologies Corp Ltd
Priority to CA3062949A priority Critical patent/CA3062949C/en
Priority to AU2019268193A priority patent/AU2019268193B2/en
Priority to JP2019564530A priority patent/JP6829327B2/ja
Priority to KR1020197034745A priority patent/KR20210033874A/ko
Priority to EP19790111.9A priority patent/EP3640825B8/en
Priority to US16/667,593 priority patent/US11314507B2/en
Publication of WO2020029592A1 publication Critical patent/WO2020029592A1/zh
Anticipated expiration legal-status Critical
Priority to US17/703,757 priority patent/US11853760B2/en
Ceased legal-status Critical Current

Links

Images

Classifications

    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00—Arrangements for program control, e.g. control units
    • G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
    • G06F9/30003—Arrangements for executing specific machine instructions
    • G06F9/3004—Arrangements for executing specific machine instructions to perform operations on memory
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F8/00—Arrangements for software engineering
    • G06F8/30—Creation or generation of source code
    • G06F8/35—Creation or generation of source code model driven
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00—Arrangements for program control, e.g. control units
    • G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/44—Arrangements for executing specific programs
    • G06F9/445—Program loading or initiating
    • G06F9/44505—Configuring for program initiating, e.g. using registry, configuration files
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00—Arrangements for program control, e.g. control units
    • G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/46—Multiprogramming arrangements
    • G06F9/50—Allocation of resources, e.g. of the central processing unit [CPU]
    • G06F9/5005—Allocation of resources, e.g. of the central processing unit [CPU] to service a request
    • G06F9/5011—Allocation of resources, e.g. of the central processing unit [CPU] to service a request the resources being hardware resources other than CPUs, Servers and Terminals
    • G06F9/5016—Allocation of resources, e.g. of the central processing unit [CPU] to service a request the resources being hardware resources other than CPUs, Servers and Terminals the resource being the memory
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00—Arrangements for program control, e.g. control units
    • G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/46—Multiprogramming arrangements
    • G06F9/50—Allocation of resources, e.g. of the central processing unit [CPU]
    • G06F9/5005—Allocation of resources, e.g. of the central processing unit [CPU] to service a request
    • G06F9/5027—Allocation of resources, e.g. of the central processing unit [CPU] to service a request the resource being a machine, e.g. CPUs, Servers, Terminals
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00—Machine learning
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G06N3/08—Learning methods
    • Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02D—CLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
    • Y02D10/00—Energy efficient computing, e.g. low power processors, power management or thermal management

Definitions

  • the present application relates to the field of computer technology, and in particular, to a conversion method, device, computer equipment, and storage medium.
  • the traditional conversion method is: developers can configure multiple different conversion models for the same neural network model, and multiple conversion models can be suitable for different processors.
  • computer equipment needs to receive the neural network model and its For the corresponding multiple conversion models, the user needs to select a model from the above-mentioned multiple models according to the type of the current computer equipment, so that the neural network model can run on the current computer equipment.
  • the data input amount and data processing amount of the above conversion method are both large, and the large data input amount and data processing amount easily exceed the storage capacity and processing limit of the computer equipment, resulting in the processing speed of the computer equipment being low or even unable to work normally.
  • the present application provides conversion methods, apparatuses, computer equipment, and storage media.
  • This application provides a model conversion method, which includes the following steps:
  • the initial offline model is converted to the original offline model according to the hardware attribute information of the computer device and a preset model conversion rule.
  • the target offline model of hardware attribute information matching of computer equipment is described.
  • the step of converting the initial offline model into a target offline model that matches the hardware attribute information of the computer device according to the hardware attribute information of the computer device and a preset model conversion rule include:
  • one of the model conversion rules is selected from a plurality of the preset model conversion rules as The steps of the target model conversion rule include:
  • the step of prioritizing more than one of the available model conversion rules includes the following steps:
  • Process parameters for converting the initial offline model to the target offline model using each of the available model conversion rules are obtained, where the process parameters include conversion speed, power consumption, memory usage, and disk I / O usage One or more of the rates;
  • the model attribute information of the initial offline model, the model attribute information of the target offline model, and the available model conversion rules are stored in a one-to-one correspondence.
  • the step of determining whether the model attribute information of the initial offline model and the hardware attribute information of the computer device match based on the initial offline model and the hardware attribute information of the computer device includes:
  • the computer device can support the running of the initial offline model according to the model attribute information of the initial offline model and the hardware attribute information of the computer device, determine the model attribute information of the initial offline model and the Matching hardware attribute information of computer equipment;
  • the step of obtaining an initial offline model includes:
  • the initial offline model and hardware attribute information of the computer device are obtained through application software on the computer device.
  • the method further includes the following steps:
  • the target offline model is stored in a first memory or a second memory of the computer device.
  • the method further includes the following steps:
  • the target offline model includes network weights, instructions corresponding to each offline node in the original network, and other calculations of each offline node and the original network Interface data between nodes.
  • This application provides a model conversion device, and the device includes:
  • An acquisition module for acquiring initial offline model and hardware attribute information of computer equipment
  • a judging module configured to judge whether the model attribute information of the initial offline model and the hardware attribute information of the computer device match according to the initial offline model and the hardware attribute information of the computer device;
  • a conversion module configured to, when the model attribute information of the initial offline model does not match the hardware attribute information of the computer device, take the initial offline according to the hardware attribute information of the computer device and a preset model conversion rule The model is converted into a target offline model that matches the hardware attribute information of the computer device.
  • the present application provides a computer device including a memory and a processor.
  • the memory stores a computer program
  • the processor implements the steps of the foregoing conversion method when the processor executes the computer program.
  • the processor includes an operation unit and a controller unit, and the operation unit includes a master processing circuit and a plurality of slave processing circuits;
  • the controller unit is configured to obtain data, a machine learning model, and a calculation instruction
  • the controller unit is further configured to parse the calculation instruction to obtain a plurality of operation instructions, and send the plurality of operation instructions and the data to the main processing circuit;
  • the master processing circuit is configured to perform pre-processing on the data and data and operation instructions transmitted between the master processing circuit and the plurality of slave processing circuits;
  • the multiple slave processing circuits are configured to perform intermediate operations in parallel according to data transmitted from the master processing circuit and operation instructions to obtain multiple intermediate results, and transmit the multiple intermediate results to the master processing circuit;
  • the main processing circuit is further configured to perform subsequent processing on the plurality of intermediate results to obtain a calculation result of the calculation instruction.
  • a computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps of the above conversion method are implemented.
  • FIG. 1 is a structural block diagram of a computer device in an embodiment
  • FIG. 2 is a structural block diagram of an embodiment of a processor in FIG. 1;
  • FIG. 3 is a structural block diagram of an embodiment of a processor in FIG. 1;
  • FIG. 4 is a structural block diagram of an embodiment of a processor in FIG. 1;
  • FIG. 5 is a schematic flowchart of a model conversion method in an embodiment
  • FIG. 6 is a schematic flowchart of a model conversion method in an embodiment
  • FIG. 7 is a schematic flowchart of an offline model generation method according to an embodiment
  • FIG. 8 is a schematic flowchart of an offline model generation method according to an embodiment
  • FIG. 9 is a network structure diagram of a network model according to an embodiment
  • FIG. 10 is a schematic diagram of an offline model generation process of the network model in FIG. 9.
  • FIG. 1 is a block diagram of a computer device according to an embodiment.
  • the computer device may be a mobile terminal such as a mobile phone or a tablet computer, or a terminal such as a desktop computer, a board card, or a cloud server.
  • the computer device can be applied to robots, printers, scanners, driving recorders, navigators, cameras, camcorders, projectors, watches, mobile storage, wearable devices, vehicles, home appliances, and / or medical devices.
  • the transportation means may include airplanes, ships and / or vehicles; household appliances may include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, cooker hoods; medical equipment may include nuclear magnetic resonance instruments, Ultrasound and / or electrocardiograph, etc.
  • the computer device may include a processor 100, a first memory 200 and a second memory 300 connected to the processor 100.
  • the processor 100 may be a general-purpose processor, such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or a DSP (Digital Signal Processing). 100 can also be a network model processor such as IPU (Intelligence Processing Unit).
  • the processor 100 may also be an instruction set processor, a related chipset, a special-purpose microprocessor (such as an application-specific integrated circuit (ASIC)), or an on-board memory used for caching, and the like.
  • ASIC application-specific integrated circuit
  • the processor 100 may include a controller unit 110 and an operation unit 120, where the controller unit 110 is connected to the operation unit 120, and the operation unit 120 may include a main processing circuit 121 and multiple processors.
  • the controller unit 110 is configured to obtain data, a machine learning model, and calculation instructions.
  • the machine learning model may specifically include a network model, and the network model may be a neural network model and / or a non-neural network model.
  • the controller unit 110 is further configured to analyze the calculation instruction obtained by the controller unit 110 to obtain an operation instruction, and send a plurality of operation instructions and data to the main processor circuit.
  • the main processing circuit is configured to perform pre-processing on data and data and operation instructions transmitted between the main processing circuit and a plurality of slave processing circuits.
  • Multiple slave processing circuits are used to perform intermediate operations in parallel according to data transmitted from the main processing circuit and operation instructions to obtain multiple intermediate results, and transmit the multiple intermediate results to the main processing circuit; the main processing circuit is also used to As a result, subsequent processing is performed to obtain the calculation result of the calculation instruction.
  • the controller unit 110 may include an instruction storage unit 111, an instruction processing unit 112, and a storage queue unit 114; the instruction storage unit 111 is configured to store calculation instructions associated with the machine learning model; and the instruction processing unit 112 is configured to perform calculations on the calculation instructions. A plurality of operation instructions are obtained through analysis; the storage queue unit 114 is configured to store an instruction queue, and the instruction queue includes a plurality of operation instructions or calculation instructions to be executed in the order of the queue.
  • the controller unit 110 may further include a dependency relationship processing unit 113 for determining whether there is an association relationship between the first operation instruction and the zeroth operation instruction before the first operation instruction when there are multiple operation instructions; for example, The first operation instruction is associated with the zeroth operation instruction, the first operation instruction is buffered in the instruction storage unit, and after the zeroth operation instruction is executed, the first operation instruction is extracted from the instruction storage unit and transmitted to the operation unit. Specifically, if the dependency processing unit 113 extracts a first storage address range of data (for example, a matrix) required in the first operation instruction according to the first operation instruction, extracts a first storage address range of the required matrix in the zeroth operation instruction according to the zeroth operation instruction.
  • a dependency relationship processing unit 113 for determining whether there is an association relationship between the first operation instruction and the zeroth operation instruction before the first operation instruction when there are multiple operation instructions; for example, The first operation instruction is associated with the zeroth operation instruction, the first operation instruction is buffered in the instruction storage unit, and after the zeroth operation instruction is executed
  • Zero memory address interval if the first memory address interval and the zeroth memory address interval have overlapping areas, determine that the first operation instruction and the zeroth operation instruction have an associated relationship; for example, if the first memory address interval and the zeroth memory address interval do not If there is an overlapped area, it is determined that the first operation instruction and the zeroth operation instruction have no correlation.
  • the arithmetic unit 120 may further include a branch processing circuit 123, wherein the main processing circuit 121 is connected to the branch processing circuit 123, and the branch processing circuit 123 is connected to a plurality of slave processing circuits 122;
  • the processing circuit 123 is configured to execute data or instructions transmitted between the master processing circuit 121 and the slave processing circuit 122.
  • the main processing circuit 121 is specifically configured to allocate an input neuron into a plurality of data blocks, and assign at least one data block, a weight, and at least one operation instruction of the plurality of operation instructions in the plurality of data blocks.
  • the branch processing circuit 123 is used to forward the data blocks, weights, and operation instructions between the main processing circuit 121 and the multiple slave processing circuits 122; the multiple slave processing circuits 122 are used to receive the data according to the operation instruction; The obtained data blocks and weights perform an operation to obtain an intermediate result, and transmit the intermediate result to the branch processing circuit 123.
  • the main processing circuit 121 is further configured to perform subsequent processing on the intermediate result sent by the branch processing circuit to obtain the result of the calculation instruction.
  • the result of the calculation instruction is sent to the controller unit.
  • the operation unit 120 may include a master processing circuit 121 and a plurality of slave processing circuits 122.
  • a plurality of slave processing circuits are distributed in an array; each slave processing circuit is connected to an adjacent slave processing circuit, and the master processing circuit connects k slave processing circuits of the plurality of slave processing circuits, and the k slave processing circuits are: The n slave processing circuits in the first row, the n slave processing circuits in the m row, and the m slave processing circuits in the first column.
  • the K slave processing circuits shown in FIG. 1C include only the first slave processing circuit.
  • the n slave processing circuits in the row, the n slave processing circuits in the m row, and the m slave processing circuits in the first column, that is, the k slave processing circuits are slaves directly connected to the main processing circuit among the multiple slave processing circuits. Processing circuit.
  • the K slave processing circuits are used to transfer data and instructions between the master processing circuit and a plurality of slave processing circuits.
  • the above-mentioned main processing circuit 121 may include one or any combination of a conversion processing circuit, an activation processing circuit, and an addition processing circuit; the conversion processing circuit is configured to execute a data block or an intermediate result received by the main processing circuit to perform a first Interchange between the data structure and the second data structure (such as the conversion of continuous data and discrete data); or perform the interchange between the first data type and the second data type on the data block or intermediate result received by the main processing circuit (Such as conversion between fixed-point type and floating-point type); the activation processing circuit is used to perform the activation operation of the data in the main processing circuit; the addition processing circuit is used to perform the addition operation or the accumulation operation.
  • a conversion processing circuit is configured to execute a data block or an intermediate result received by the main processing circuit to perform a first Interchange between the data structure and the second data structure (such as the conversion of continuous data and discrete data); or perform the interchange between the first data type and the second data type on the data block or intermediate result received by the main processing circuit (Such as
  • the first memory 200 or the second memory 300 may store a computer program, which is used to implement the model conversion method provided in the embodiment of the present application.
  • the model conversion method is used to convert the initial offline model to match the hardware attribute information of the computer device when the model attribute information of the initial offline model received by the computer device does not match the hardware attribute information of the computer device.
  • Target offline model so that the computer device can run the initial offline model.
  • the data input amount of the computer equipment is small, and the conversion from the initial offline model to the target offline model can be completed automatically without human intervention.
  • the conversion process is simple and the conversion efficiency is high.
  • the computer program stored on the first memory 200 or the second memory 300 can also be used to implement the offline model generation method provided in the embodiments of the present application.
  • the first memory 200 may be used to store relevant data during the running of the network model, such as network input data, network output data, network weights and instructions, and so on.
  • the first memory 200 may be an internal memory, such as a volatile memory such as a cache.
  • the second memory 300 may be used to store the offline model corresponding to the network model.
  • the second memory 300 may be a non-volatile memory.
  • the computer device may further include a processor and a memory, and the computer device may include a memory connected to the processor and the processor.
  • the processor may use the processor shown in FIG. 2-4, and for a specific structure thereof, refer to the description about the processor 100 above.
  • the memory may include multiple storage units.
  • the memory may include a first storage unit, a second storage unit, and a third storage unit.
  • the first storage unit may be used to store a computer program, and the computer program is used to implement the present invention.
  • the model conversion method provided in the application examples.
  • the second storage unit may be used to store related data during the operation of the original network
  • the third storage unit is used to store the offline model corresponding to the original network, the target offline model corresponding to the offline model, the preset model conversion rule, and so on .
  • the number of storage units included in the memory may be more than three, which is not specifically limited herein.
  • an embodiment of the present application provides a model conversion method, which can be applied to the foregoing computer equipment.
  • the model conversion method is used to convert the initial offline model to a target offline that matches the hardware attribute information of the computer device when the model attribute information of the initial offline model received by the computer device does not match the hardware attribute information of the computer device.
  • Model so that the computer device can run the initial offline model.
  • the above method may include the following steps:
  • the foregoing initial offline model refers to an offline model directly obtained by a computer device, and the initial offline model may be stored in the second memory 300.
  • the offline model may include necessary network structure information such as network weights and instructions of each computing node in an original network.
  • the instructions may be used to indicate which computing function the computing node is to perform, which may specifically include the original network.
  • Information such as the computing attributes of each computing node and the connection relationship between each computing node.
  • the original network may be a network model, such as a neural network model, etc., as shown in FIG. 9.
  • the computer device can realize the computing function of the original network by running the offline model corresponding to the original network, and does not need to repeatedly compile and operate the same original network, thereby reducing the running time of the processor when running the network, thereby improving processing. Processing speed and efficiency.
  • the initial offline model in the embodiment of the present application may be an offline model directly generated based on the original network, or an offline model obtained after one or more conversions of the offline model with other model attribute information. Be specific.
  • S200 Determine whether the model attribute information of the initial offline model matches the hardware attribute information of the computer device according to the initial offline model and the hardware attribute information of the computer device.
  • the model attribute information of the initial offline model may include a data structure, a data type, and the like of the initial offline model.
  • the hardware attribute information of the computer device includes the model of the computer device, the type of data that the computer device can process (such as fixed-point or floating-point, etc.), the data structure, and so on.
  • the computer device can determine whether the computer device can support the running of the initial offline model (that is, whether the computer device can run the initial offline model) according to the model attribute information of the initial offline model and the hardware attribute information of the computer device, so as to determine Whether the model attribute information of the initial offline model matches the hardware attribute information of the computer device.
  • the computer device when the computer device can support the running of the initial offline model, it is determined that the model attribute information of the initial offline model matches the hardware attribute information of the computer device.
  • the computer device does not support the running of the initial offline model, it is determined that the model attribute information of the initial offline model does not match the hardware attribute information of the computer device.
  • the computer device may determine the data type and data structure that the computer device can process according to its hardware attribute information, and determine the data type and data structure of the initial offline model according to the model attribute information of the initial offline model. If the computer device determines, based on its own hardware attribute information and the model attribute information of the initial offline model, that the computer device can support the data type and data structure of the initial offline model, the model attribute information of the initial offline model may be determined. Match the hardware attribute information of the computer device. If the computer device determines that the computer device cannot support the data type and data structure of the initial offline model according to its own hardware attribute information and the model attribute information of the initial offline model, the model of the initial offline model may be determined. The attribute information does not match the hardware attribute information of the computer device.
  • step S300 is performed: according to the hardware attribute information of the computer device and a preset model conversion rule, the initial offline model is converted to the hardware attribute of the computer device Information matching target offline model.
  • the initial offline model needs to be converted to process
  • the computer device can use the initial offline model to implement corresponding operations. That is, when the model attribute information of the initial offline model does not match the hardware attribute information of the computer device, the initial offline model may be converted into the computer device according to the hardware attribute information of the computer device and a preset model conversion rule.
  • the hardware attribute information matches the target offline model.
  • the model attribute information of the target offline model matches the hardware attribute information of the computer device, and the computer device can support the operation of the target offline model.
  • the target offline model also includes the network weights and instructions of the computing nodes in the original network and other necessary network structure information.
  • the above-mentioned preset model conversion rule may be stored in advance in the first memory or the second memory of the computer device. If the model attribute information of the initial offline model does not match the hardware attribute information of the computer device, the computer device may obtain a corresponding model conversion rule from the first memory or the second memory to convert the initial offline model into the target offline model.
  • there may be more than one preset model conversion rule and more than one model conversion rule may be stored in a one-to-one correspondence with the above-mentioned initial offline model and target offline model.
  • the above-mentioned initial offline model, target offline model, and preset model conversion rules may be correspondingly stored in a mapping table manner.
  • the model attribute information of the initial offline model, the model attribute information of the target offline model, and the available model conversion rules are stored in a one-to-one correspondence.
  • step S400 may be performed: the computer device may directly run the initial offline model it receives, that is, the computer device may use the network weights included in the initial offline model. And instructions to perform calculations to achieve the calculation functions of the original network. Specifically, if the model attribute information of the initial offline model matches the hardware attribute information of the computer device, it indicates that the computer device can support the operation of the initial offline model. At this time, there is no need to perform conversion processing on the initial offline model. The computer device Ability to run this initial offline model.
  • directly running the initial offline model refers to using the initial offline model to run a machine learning algorithm (such as a neural network algorithm) corresponding to the original network, and to implement a target application of the algorithm by performing a forward operation.
  • a machine learning algorithm such as a neural network algorithm
  • a target application of the algorithm by performing a forward operation.
  • artificial intelligence applications such as speech recognition
  • a computer device only needs to receive an initial offline model, and according to a preset model conversion rule and hardware attribute information of the computer device, the initial offline model can be converted into a target offline model without the need of Obtaining multiple different model data greatly reduces the data input volume of computer equipment, avoids problems such as exceeding the storage capacity of computer equipment caused by excessive data input volume, and ensures the normal operation of computer equipment.
  • the above model conversion method can reduce the data processing amount of the computer equipment, thereby further improving processing efficiency and reducing power consumption.
  • no manual intervention is required, the degree of automation is high, and the use is convenient.
  • step S300 may include:
  • S310 Determine model attribute information of the target offline model according to hardware attribute information of the computer device.
  • the computer device may determine the data type and data structure that the computer device can support according to its own hardware attribute information, so as to determine the model attribute information such as the data type and data structure of the target offline model. That is, the computer device can determine the model attribute information of the target offline model according to its hardware attribute information.
  • the model attribute information of the target offline model may include information such as the data structure and data type of the target offline model.
  • a model conversion rule is selected as a target model conversion rule from a plurality of preset model conversion rules.
  • the target model can be determined according to the model attribute information of the initial offline model and the model attribute information of the target offline model.
  • the model conversion rule may include a data type conversion method and a data structure conversion method.
  • S330 Convert the initial offline model into the target offline model according to the model attribute information of the initial offline model and the target model conversion rule.
  • the computer device can convert the initial offline model into the target offline model according to the conversion method provided by the target model conversion rule, so that the computer device can run the target offline model for calculation.
  • step S320 further includes the following steps:
  • more than one available model conversion rule is selected from a plurality of preset model conversion rules.
  • the above-mentioned available model conversion rule refers to a conversion rule capable of converting an initial offline model into a target offline model.
  • the above-mentioned priority sorting method may be preset or user-defined.
  • the computer device may obtain process parameters corresponding to each available model conversion rule, and prioritize more than one available model conversion rule according to the process parameters corresponding to each available model conversion rule.
  • the process parameter may be a performance parameter of a computer device involved in converting the initial offline model to the target offline model by using the available model conversion rule.
  • the process parameters may include one or more of conversion speed, power consumption, memory usage, and disk I / O usage.
  • the computer device may perform a combination of one or more process parameters such as the conversion speed, conversion power consumption, memory usage, and disk I / O usage during the conversion of the initial offline model into the target offline model.
  • the model conversion rule can be used for scoring (for example, weighting calculation of each reference factor to obtain a score value), and the available model conversion rule with the highest score is used as the highest available model conversion rule. That is, the available model conversion rule with the highest score can be used as the target conversion rule.
  • step S330 can be performed to convert the initial offline model into the target offline model according to the model attribute information of the initial offline model and the target model conversion rule. In this way, a better model conversion rule can be selected by combining equipment performance factors in the model conversion process to improve the processing speed and efficiency of computer equipment.
  • the above method further includes the following steps:
  • S500 Store the target offline model in the first memory or the second memory of the computer device.
  • the computer device may store the target offline model obtained by the computer device on a local memory (such as the first memory) of the computer device.
  • the computer device may further store the target offline model obtained by it in an external storage (such as a second storage) connected to the computer device, and the external storage may be a cloud storage or other storage, or the like.
  • the process in which the computer device executes the above-mentioned target offline model may include the following steps:
  • the target offline model includes necessary network structure information such as network weights, instructions corresponding to each offline node in the original network, and interface data between each offline node and other computing nodes in the original network.
  • a computer device when a computer device needs to implement operations using the target offline model, it may directly obtain the target offline model from the first memory or the second memory, and perform operations according to network weights and instructions in the target offline model. So as to realize the computing function of the original network. In this way, when the computer device needs to repeatedly use the target offline model, there is no need to repeatedly perform the above-mentioned conversion operation, and the corresponding operation can be realized by simply reading the target offline model directly.
  • directly running the target offline model means that the target offline model is used to run a machine learning algorithm (such as a neural network algorithm) corresponding to the original network, and the target application of the algorithm is realized by performing forward operation (Such as artificial intelligence applications such as speech recognition).
  • a machine learning algorithm such as a neural network algorithm
  • the target application of the algorithm is realized by performing forward operation (Such as artificial intelligence applications such as speech recognition).
  • an application software is installed on the computer device, and the initial offline model can be obtained through the application software on the computer device.
  • the application software can provide a method for reading the initial offline model from the memory 200 or an external memory. Interface so that the initial offline model can be obtained through the application software.
  • the application software may further provide an interface for reading hardware attribute information of the computer device, so that the hardware attribute information of the computer device may be obtained through the application software.
  • the application software may further provide an interface capable of reading a preset model rule, so that the preset model conversion rule may be obtained through the application software.
  • the computer device may also provide an input / output interface (such as an I / O interface), etc.
  • the input / output interface is used to obtain an initial offline model or output a target offline model.
  • the hardware attribute information of the computer device may be preset stored in the computer device.
  • an embodiment of the present application further provides a method for generating an offline model according to the original network model, and is configured to generate and store the offline model of the original network according to the obtained related data of the original network. Therefore, when the processor runs the original network again, the offline model corresponding to the original network can be directly run without compiling and other operations on the same original network again, thereby shortening the running time of the processor when running the network, thereby improving the processor Processing speed and efficiency.
  • the original network may be a network model such as a neural network or a non-neural network.
  • the original network may be a network shown in FIG. 9.
  • the above method includes the following steps:
  • the processor of the computer device can obtain the model data set and model structure parameters of the original network, and the network structure diagram of the original network can be obtained through the model data set and model structure parameters of the original network.
  • the model data set includes data such as network weights corresponding to each computing node in the original network. W1 to W6 in the neural network shown in FIG. 9 are used to represent the network weights of the computing nodes.
  • the model structure parameters include the connection relationship between multiple computing nodes in the original network and the computing attributes of each computing node. Among them, the connection relationship between computing nodes is used to indicate whether there is data transfer between the computing nodes. For example, when multiple computing nodes When there is a data flow between them, it can be explained that there is a connection relationship between multiple computing nodes.
  • connection relationship of the computing nodes may include an input relationship, an output relationship, and the like.
  • FIG. 9 if the output of the computing node F1 is used as the input of the computing nodes F4 and F5, it can be explained that there is a connection relationship between the computing node F1 and the computing node F4, and there is a connection relationship between the computing node F1 and the computing node F4.
  • there is no data transfer between the computing node and the computing node F2 it can be explained that there is no connection relationship between the computing node F1 and the computing node F2.
  • the computing attributes of each computing node can include the computing type and computing parameters of the corresponding computing node, where the computing node's computing type refers to what kind of computing the computing node uses to complete.
  • the computing type of a computing node can include addition operations, subtraction operations, and Convolution operations and so on.
  • the computing node may be a computing node for implementing an addition operation, a computing node for implementing a subtraction operation, or a computing node for implementing a convolution operation.
  • the computing parameters of the computing node may be necessary parameters required to complete the type of computing corresponding to the computing node.
  • the calculation type of a calculation node may be a calculation node used to implement an addition operation.
  • the calculation parameter of the calculation node may be an addition number in an addition operation, and the added number in the addition operation may be obtained as input data through Obtained by the module, or the added number in the addition operation may be output data of a previous computing node of the computing node, and so on.
  • S020 Run the original network according to the model data set and model structure parameters of the original network, and obtain instructions corresponding to each computing node in the original network.
  • the processor of the computer device may run the original network according to the model data set and model structure parameters of the original network, and obtain instructions corresponding to each computing node in the original network. Further, the processor may also obtain input data of the original network, run the original network according to the input data of the original network, the network model data set, and the model structure parameters, and obtain instructions corresponding to each computing node in the original network. Furthermore, the above process of running the original network to obtain the instructions of each computing node is essentially a compilation process, and the compilation process may be implemented by a processor of a computer device or a virtual device. That is, the processor or virtual device of the computer device runs the original network according to the model data set and model structure parameters of the original network. Among them, the virtual device refers to virtualizing a section of processor running space in the memory memory space of the memory.
  • running the original network in this embodiment means that the processor runs some kind of machine learning algorithm (such as a neural network algorithm) using artificial neural network model data, and achieves the target application of the algorithm (such as speech by performing a forward operation).
  • Artificial intelligence applications such as identification).
  • S030 Generate an offline model corresponding to the original network according to the network weights and instructions corresponding to the computing nodes of the original network, and store the offline model corresponding to the original network in a non-volatile memory (database).
  • control module of the processor may generate an offline model corresponding to the original network according to the network weights and instructions corresponding to the computing nodes of the original network.
  • the control module of the processor may convert the computing nodes of the original network.
  • Corresponding network weights and instructions are stored in non-volatile memory to achieve the generation and storage of offline models.
  • For each computing node of the original network a one-to-one correspondence between the network weights and instructions of the computing node is stored. In this way, when the original network is run again, the offline model corresponding to the original network can be directly obtained from the non-volatile memory, and the original network is run according to the corresponding offline model, without performing online calculations on each computing node of the original network. Compile to obtain instructions, which improves the system's operating speed and efficiency.
  • directly running the offline model corresponding to the original network refers to using the offline model to run a machine learning algorithm (such as a neural network algorithm) corresponding to the original network, and achieve the goal of the algorithm by performing a forward operation.
  • Applications such as artificial intelligence applications such as speech recognition).
  • step S200 may include:
  • the processor may obtain the execution order of each computing node in the original network according to the model structure parameters of the original network, and obtain the execution order of each computing node in the original network according to the connection relationship of each computing node in the original network.
  • the input data of the computing node F4 is the output data of the computing node F1 and the output data of the computing node F2
  • the input data of the computing node F6 is the output data of the computing node F4 and the output data of the computing node F5. Therefore, the execution order of each computing node in the neural network shown in FIG. 9 may be F1-F2-F3-F4-F5-F6 or F1-F3-F2-F5-F4-F6 and so on.
  • the computing nodes F1, F2, and F3 can be executed in parallel
  • the computing nodes F4 and F5 can also be executed in parallel.
  • the execution order is not specifically limited.
  • S022 Run the original network according to the execution order of each computing node in the original network, and obtain instructions corresponding to each computing node in the original network.
  • the processor may run the original network according to the execution order of the computing nodes in the original network to obtain instructions corresponding to the computing nodes in the original network, that is, the processor may compile data such as the model data set of the original network to obtain each Instructions corresponding to a computing node.
  • the processor may compile data such as the model data set of the original network to obtain each Instructions corresponding to a computing node.
  • the instructions corresponding to each computing node it is possible to know what computing function the computing node is used to achieve, that is, to obtain computing properties such as the computing type and computing parameters of the computing node.
  • step S300 further includes:
  • S031 Obtain a memory allocation method of the original network according to the model data set and model structure parameters of the original network.
  • the processor may obtain the memory allocation method of the original network according to the model data set and model structure parameters of the original network; obtain the execution order of each computing node in the original network according to the model structure parameters of the original network; and according to the original network, The execution order of each computing node determines the current network memory allocation method. For example, according to the execution order of each computing node, the relevant data of each computing node in the running process is saved to a stack.
  • the memory allocation method refers to determining a storage location of data (including input data, output data, network weight data, intermediate result data, and the like) of each computing node in the original network on a memory space (such as a first memory).
  • a data table may be used to store the mapping relationship between the data of each computing node (input data, output data, network weight data, intermediate result data, etc.) and the memory space.
  • the related data during the operation of the original network is stored in the first storage, where the related data during the operation of the original network includes the network weights and instructions corresponding to the computing nodes of the original network. , Input data, intermediate calculation results, and output data.
  • X1 and X2 represent the input data of the neural network
  • Y represents the output data of the neural network.
  • the processor can convert the output data of the neural network into control commands that control the robot or different digital interfaces.
  • W1 to W6 are used to represent the network weights corresponding to the computing nodes F1, F2, and F3, and the output data of the computing nodes F1 to F5 can be used as intermediate calculation results.
  • the processor can store the relevant data during the operation of the original network to the first memory, such as volatile memory such as internal memory or cache, according to the determined memory allocation mode. For the specific storage mode, see the left half of FIG. 10 storage.
  • S033 Obtain the network weights and instructions corresponding to the computing nodes of the original network from the first memory, and store the network weights and instructions corresponding to the computing nodes of the original network in the second memory to generate an offline model.
  • the second memory may be a non-volatile memory such as an external memory.
  • the generation process of the offline model can be specifically shown in FIG. 10.
  • the corresponding offline model of the original network is stored in the storage space in the right half of FIG. 10.
  • the processor can obtain the model data set, model structure parameters, and input data of the original network, so as to obtain the network structure diagram of the original network according to the model data set and model structure parameters of the original network, as shown in FIG. 9.
  • the processor can obtain the connection relationship of the computing nodes of the original network according to the model structure parameters of the original network, and obtain the execution order of the computing nodes in the original network and the original network during operation according to the connection relationships of the computing nodes.
  • Memory allocation mode so as to obtain the storage location of related data during the operation of the original network. As shown in the left half of the storage space in FIG. 10, related data during the running of the original network can be stored in a stack in accordance with the execution order of each computing node.
  • the processor may store the network weights and instructions corresponding to the computing nodes of the original network in a non-volatile second memory to generate an offline model.
  • the offline model For the storage method of the offline model, see the right half of the storage in Figure 8. Space shown.
  • the offline model only contains data such as network weights and instructions necessary to run the original network, and does not need to store input data, output data, or intermediate calculation results during the operation of the original network, thereby reducing the The consumption of storage space in the second memory.
  • the offline model also includes node interface data, and the node interface data is used to represent the connection relationships of the computing nodes of the original network.
  • the node interface data may include input data sources and output data sources of each computing node.
  • the node interface data may include computing nodes F1, F2, and F3 as starting computing nodes, whose inputs are preset input data, and the output data of computing node F1 is used as computing node F4 and computing node F5. Input data and so on.
  • steps in the flowcharts of FIGS. 5-8 are sequentially displayed in accordance with the directions of the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless explicitly stated in this document, the execution of these steps is not strictly limited, and these steps can be performed in other orders. Moreover, at least a part of the steps in FIG. 5-8 may include multiple sub-steps or stages. These sub-steps or stages are not necessarily performed at the same time, but may be performed at different times. These sub-steps or stages The execution order of is not necessarily performed sequentially, but may be performed in turn or alternately with at least a part of another step or a sub-step or stage of another step.
  • Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory.
  • Volatile memory can include random access memory (RAM) or external cache memory.
  • RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous chain Synchlink DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
  • SRAM static RAM
  • DRAM dynamic RAM
  • SDRAM synchronous DRAM
  • DDRSDRAM dual data rate SDRAM
  • ESDRAM enhanced SDRAM
  • SLDRAM synchronous chain Synchlink DRAM
  • Rambus direct RAM
  • DRAM direct memory bus dynamic RAM
  • RDRAM memory bus dynamic RAM
  • An embodiment of the present application further provides a model conversion device.
  • the above model conversion device includes an acquisition module, a judgment module, and a conversion module. among them,
  • the acquisition module is used to acquire the initial offline model and the hardware attribute information of the computer equipment
  • the judging module is configured to judge whether the model attribute information of the initial offline model and the hardware attribute information of the computer device match according to the initial offline model and the hardware attribute information of the computer device;
  • the conversion module is used to convert the initial offline model to the hardware attributes of the computer device when the model attribute information of the initial offline model does not match the hardware attribute information of the computer device. Information matching target offline model.
  • the model conversion device may be application software (Application) installed on the computer device.
  • the application software can provide an interface for reading the initial offline model from the memory 200 or an external memory, so that the initial offline model can be obtained through the application software.
  • the application software may further provide an interface for reading hardware attribute information of the computer device, so that the hardware attribute information of the computer device may be obtained through the application software.
  • the application software may further provide an interface capable of reading a preset model rule, so that the preset model conversion rule may be obtained through the application software.
  • the application software may, according to the initial offline model and the hardware attribute information of the computer device, and a preset model replacement rule, and when the model attribute information of the initial offline model does not match the hardware attribute information of the computer device, the initial offline model
  • the offline model is converted into a target offline model that matches the hardware attribute information of the computer device, and the target offline model is stored in the first memory or the second memory.
  • Each module in the above model conversion device may be implemented in whole or in part by software, hardware, and a combination thereof.
  • the above-mentioned modules may be embedded in the hardware form or independent of the processor in the computer device, or may be stored in the memory of the computer device in the form of software, so that the processor calls and performs the operations corresponding to the above modules.
  • an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored.
  • the computer program is executed by a processor, the steps of the method of any one of the foregoing claims are implemented.
  • the computer-readable storage medium may include non-volatile and / or volatile memory.
  • Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory.
  • Volatile memory can include random access memory (RAM) or external cache memory.
  • RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous chain Synchlink DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
  • SRAM static RAM
  • DRAM dynamic RAM
  • SDRAM synchronous DRAM
  • DDRSDRAM dual data rate SDRAM
  • ESDRAM enhanced SDRAM
  • SLDRAM synchronous chain Synchlink DRAM
  • Rambus direct RAM
  • DRAM direct memory bus dynamic RAM
  • RDRAM memory bus dynamic RAM
  • the foregoing initial offline model refers to an offline model directly obtained by a computer device, and the initial offline model may be stored in the second memory 300.
  • the offline model may include necessary network structure information such as network weights and instructions of each computing node in an original network.
  • the instructions may be used to indicate which computing function the computing node performs, and it may specifically include the original Information such as the computing attributes of each computing node in the network and the connection relationship between the computing nodes.
  • the initial offline model in the embodiment of the present application may be an offline model directly generated from the original network, or an offline model with other model attribute information is obtained after one or more conversions, and is not performed here. Specific limitations.
  • S200 Determine whether the model attribute information of the initial offline model matches the hardware attribute information of the computer device according to the initial offline model and the hardware attribute information of the computer device.
  • the model attribute information of the initial offline model may include a data structure, a data type, and the like of the initial offline model.
  • the hardware attribute information of the computer device includes the model of the computer device, the type of data that the computer device can process (such as fixed-point or floating-point, etc.), the data structure, and so on.
  • the computer device may determine the data type and data structure that the computer device can process according to its hardware attribute information, and determine the data type and data structure of the initial offline model according to the model attribute information of the initial offline model. If the computer device determines, based on its own hardware attribute information and the model attribute information of the initial offline model, that the computer device can support the data type and data structure of the initial offline model, the model attribute information of the initial offline model may be determined. Match the hardware attribute information of the computer device. If the computer device determines that the computer device cannot support the data type and data structure of the initial offline model according to its own hardware attribute information and the model attribute information of the initial offline model, the model of the initial offline model may be determined. The attribute information does not match the hardware attribute information of the computer device.
  • step S300 is performed: according to the hardware attribute information of the computer device and a preset model conversion rule, the initial offline model is converted to the hardware attribute of the computer device Information matching target offline model.
  • the initial offline model needs to be converted to process
  • the computer device can use the initial offline model to implement corresponding operations. That is, when the model attribute information of the initial offline model does not match the hardware attribute information of the computer device, the initial offline model may be converted into the computer device according to the hardware attribute information of the computer device and a preset model conversion rule.
  • the hardware attribute information matches the target offline model.
  • the model attribute information of the target offline model matches the hardware attribute information of the computer device, and the computer device can support the operation of the target offline model.
  • the target offline model also includes the network weights and instructions of the computing nodes in the original network and other necessary network structure information.
  • step S400 may be performed: the computer device may directly run the initial offline model it receives, that is, the computer device may use the network weights included in the initial offline model. And instructions to perform calculations to achieve the calculation functions of the original network. Specifically, if the model attribute information of the initial offline model matches the hardware attribute information of the computer device, it indicates that the computer device can support the operation of the initial offline model. At this time, there is no need to perform conversion processing on the initial offline model. The computer device Ability to run this initial offline model.
  • Any process or method description in a flowchart or otherwise described herein can be understood as a module, fragment, or portion of code that includes one or more executable instructions for implementing a particular logical function or step of a process
  • the scope of the alternative implementations of this application includes additional implementations, in which the functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved, which It should be understood by those skilled in the art to which the embodiments of the present application belong.
  • each part of the application may be implemented by hardware, software, firmware, or a combination thereof.
  • multiple steps or methods may be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system.
  • a suitable instruction execution system For example, if implemented in hardware, as in another embodiment, it may be implemented using any one or a combination of the following techniques known in the art: Discrete logic circuits, application-specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.
  • a person of ordinary skill in the art can understand that all or part of the steps carried by the methods in the foregoing embodiments can be implemented by a program instructing related hardware.
  • the program can be stored in a computer-readable storage medium.
  • the program is When executed, one or a combination of the steps of the method embodiment is included.
  • each functional unit in each embodiment of the present application may be integrated into one processing module, or each unit may exist separately physically, or two or more units may be integrated into one module.
  • the above integrated modules may be implemented in the form of hardware or software functional modules. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
  • the aforementioned storage medium may be a read-only memory, a magnetic disk, or an optical disk.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Software Systems (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Mathematical Physics (AREA)
  • Evolutionary Computation (AREA)
  • Computing Systems (AREA)
  • Data Mining & Analysis (AREA)
  • Artificial Intelligence (AREA)
  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Medical Informatics (AREA)
  • General Health & Medical Sciences (AREA)
  • Biomedical Technology (AREA)
  • Biophysics (AREA)
  • Computational Linguistics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Molecular Biology (AREA)
  • Stored Programmes (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)
  • Supply And Distribution Of Alternating Current (AREA)
  • Power Sources (AREA)
  • Advance Control (AREA)
  • Devices For Executing Special Programs (AREA)

Abstract

一种模型转换方法、装置、计算机设备及存储介质,该方法能够将初始离线模型转换为目标离线模型。上述的模型转换方法、装置、计算机设备和存储介质,极大的降低了计算机设备的数据输入量,转换过程简单,从而可以降低该计算机设备的数据处理量,进而可以提高处理效率,降低功耗。

Description

转换方法、装置、计算机设备及存储介质
相关申请
本申请要求2018年08月10日申请的,申请号为201810913895.6,名称为“转换方法、装置、计算机设备及存储介质”的中国专利申请的优先权,在此将其全文引入作为参考。
技术领域
本申请涉及计算机技术领域,具体涉及一种转换方法、装置、计算机设备及存储介质。
背景技术
随着人工智能技术的发展,如今深度学习已无处不在且必不可少,并随之产生了许多可扩展的深度学习系统,例如,TensorFlow、MXNet、Caffe和PyTorch等等,上述深度学习系统可以用于提供各种能够在CPU或GPU等处理器上运行的神经网络模型或其他机器学习模型。当神经网络模型等机器学习模型需在不同的处理器上运行时,往往需要对神经网络模型等机器学习模型进行转换。
传统的转换方法是:开发者可以为同一神经网络模型配置多个不同的转换模型,多个转换模型可以适用于不同的处理器,在实际使用中,计算机设备需同时接收该神经网络模型及其对应的多个转换模型,用户需要根据当前计算机设备的类型从上述的多种模型中选择一个模型,从而使得该神经网络模型能够在当前计算机设备上运行。但上述转换方法的数据输入量及数据处理量均较大,上述较大的数据输入量及数据处理量容易超出计算机设备的存储容量及处理极限,导致计算机设备的处理速度低甚至无法正常工作。
发明内容
为至少在一定程度上克服相关技术中存在的问题,本申请提供转换方法、装置、计算机设备和存储介质。
本申请提供了一种模型转换方法,所述方法包括如下步骤:
获取初始离线模型及计算机设备的硬件属性信息;
根据所述初始离线模型和所述计算机设备的硬件属性信息,判断所述初始离线模型的模型属性信息与所述计算机设备的硬件属性信息是否匹配;
若所述初始离线模型的模型属性信息与所述计算机设备的硬件属性信息不匹配,则根据所述计算机设备的硬件属性信息和预设的模型转换规则,将所述初始离线模型转换为与所述计算机设备的硬件属性信息匹配的目标离线模型。
在一个实施例中,所述的根据所述计算机设备的硬件属性信息和预设的模型转换规则,将所述初始离线模型转换为与所述计算机设备的硬件属性信息匹配的目标离线模型的步骤,包括:
根据所述计算机设备的硬件属性信息确定所述目标离线模型的模型属性信息;
根据所述初始离线模型的模型属性信息和所述目标离线模型的模型属性信息,从多个所述预设的模型转换规则中选择一个所述模型转换规则作为目标模型转换规则;
根据所述初始离线模型的模型属性信息和所述目标模型转换规则,将所述初始离线模型转换为所述目标离线模型。
在一个实施例中,所述的根据所述初始离线模型的模型属性信息和所述目标离线模型的模型属性信息,从多个所述预设的模型转换规则中选择一个所述模型转换规则作为目标模型转换规则的步骤,包括:
根据所述初始离线模型的模型属性信息和所述目标离线模型的模型属性信息,从多个所述预设的模型转换规则中选择出一个以上的可用模型转换规则;
对所述一个以上的可用模型转换规则进行优先级排序,并将优先级最高的可用模型转换规则作为所述目标转换规则。
在一个实施例中,所述的对一个以上的所述可用模型转换规则进行优先级排序的步骤,包括如下步骤:
分别获取采用各个所述可用模型转换规则,将所述初始离线模型转换为所述目标离线模型的过程参数,其中,所述过程参数包括转换速度、功耗、内存占用率及磁盘I/O占用率中的一种或多种;
根据各个所述可用模型转换规则的过程参数,对一个以上的可用模型转换规则进行优先级排序。
在一个实施例中,所述初始离线模型的模型属性信息、所述目标离线模型的模型属性信息及所述可用模型转换规则三者之间一一对应存储。
在一个实施例中,所述的根据所述初始离线模型和所述计算机设备的硬件属性信息,判断所述初始离线模型的模型属性信息与所述计算机设备的硬件属性信息是否匹配的步骤包括:
若根据所述初始离线模型的模型属性信息和所述计算机设备的硬件属性信息,确定所述计算机设备能够支持所述初始离线模型的运行时,则判定所述初始离线模型的模型属性信息与所述计算机设备的硬件属性信息匹配;
若根据所述初始离线模型的模型属性信息和所述计算机设备的硬件属性信息,确定所述计算机设备不支持所述初始离线模型的运行时,则判定所述初始离线模型的模型属性信息与所述计算机设备的硬件属性信息不匹配。
在一个实施例中,所述的获取初始离线模型的步骤,包括:
通过所述计算机设备上的应用软件获取所述初始离线模型和所述计算机设备的硬件属性信息。
在一个实施例中,所述方法还包括如下步骤:
将所述目标离线模型存储于所述计算机设备的第一存储器或第二存储器。
在一个实施例中,所述方法还包括如下步骤:
获取所述目标离线模型,并运行所述目标离线模型,其中,所述目标离线模型中包含原始网络中各个离线节点对应的网络权值、指令以及各个离线节点与所述原始网络中的其他计算节点之间的接口数据。
本申请提供了一种模型转换装置,所述装置包括:
获取模块,用于获取初始离线模型及计算机设备的硬件属性信息;
判断模块,用于根据所述初始离线模型和所述计算机设备的硬件属性信息,判断所述初始离线模型的模型属性信息与所述计算机设备的硬件属性信息是否匹配;
转换模块,用于在所述初始离线模型的模型属性信息与所述计算机设备的硬件属性信息不匹配时,根据所述计算机设备的硬件属性信息和预设的模型转换规则,将所述初始离 线模型转换为与所述计算机设备的硬件属性信息匹配的目标离线模型。
本申请提供了一种计算机设备,包括存储器和处理器,所述存储器存储有计算机程序,所述处理器执行所述计算机程序时实现上述转换方法的步骤。
在一个实施例中,所述处理器包括运算单元和控制器单元,所述运算单元包括主处理电路和多个从处理电路;
所述控制器单元用于获取数据、机器学习模型以及计算指令;
所述控制器单元还用于解析所述计算指令得到多个运算指令,并将所述多个运算指令以及所述数据发送给所述主处理电路;
所述主处理电路用于对所述数据以及在所述主处理电路与所述多个从处理电路之间进行传输的数据和运算指令执行前序处理;
所述多个从处理电路用于依据从所述主处理电路传输的数据以及运算指令并行执行中间运算得到多个中间结果,并将多个中间结果传输给所述主处理电路;
所述主处理电路还用于对所述多个中间结果执行后续处理得到所述计算指令的计算结果。
一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现上述转换方法的步骤。
应当理解的是,以上的一般描述和后文的细节描述仅是示例性和解释性的,并不能限制本申请。
附图说明
此处的附图被并入说明书中并构成本说明书的一部分,示出了符合本申请的实施例,并与说明书一起用于解释本申请的原理。
图1为一个实施例中计算机设备的结构框图;
图2为图1中处理器的一个实施例的结构框图;
图3为图1中处理器的一个实施例的结构框图;
图4为图1中处理器的一个实施例的结构框图;
图5为一个实施例中模型转换方法的流程示意图;
图6为一个实施例中模型转换方法的流程示意图;
图7为一个实施例中离线模型生成方法的流程示意图;
图8为一个实施例中离线模型生成方法的流程示意图;
图9为一实施例的网络模型的网络结构图;
图10为图9中网络模型的离线模型生成过程示意图。
具体实施方式
这里将详细地对示例性实施例进行说明,其示例表示在附图中。下面的描述涉及附图时,除非另有表示,不同附图中的相同数字表示相同或相似的要素。以下示例性实施例中所描述的实施方式并不代表与本申请相一致的所有实施方式。相反,它们仅是与如所附权利要求书中所详述的、本申请的一些方面相一致的装置和方法的例子。
图1为一实施例的计算机设备的框图,该计算机设备可以是手机或平板电脑等移动终端,或台式电脑、板卡或云端服务器等终端。该计算机设备可以应用于机器人、打印机、 扫描仪、行车记录仪、导航仪、相机、摄像机、投影仪、手表、移动存储、可穿戴设备、交通工具、家用电器、和/或医疗设备。其中,交通工具可以包括飞机、轮船和/或车辆;家用电器可以包括电视、空调、微波炉、冰箱、电饭煲、加湿器、洗衣机、电灯、燃气灶、油烟机;医疗设备可以包括核磁共振仪、B超仪和/或心电图仪等等。
该计算机设备可以包括处理器100、与该处理器100连接的第一存储器200及第二存储器300。可选地,处理器100可以是通用处理器,如CPU(Central Processing Unit,中央处理器)、GPU(Graphics Processing Unit,图形处理器)或DSP(Digital Signal Processing,数字信号处理),该处理器100还可以为IPU(Intelligence Processing Unit,智能处理器)等网络模型处理器。当然,该处理器100还可以是指令集处理器、相关芯片组、专用微处理器(如,专用集成电路(ASIC))或用于缓存用途的板载存储器等等。
可选地,如图2所示,该处理器100可以包括控制器单元110和运算单元120,其中,控制器单元110与运算单元120连接,该运算单元120可以包括一个主处理电路121和多个从处理电路122。控制器单元110用于获取数据、机器学习模型以及计算指令。该机器学习模型具体可以包括网络模型,该网络模型可以为神经网络模型和/或非神经网络模型。控制器单元110还用于解析其获取的计算指令得到运算指令,并将多个运算指令以及数据发送给主处理器电路。主处理电路用于对数据以及该主处理电路与多个从处理电路之间传输的数据和运算指令执行前序处理。多个从处理电路用于依据从主处理电路传输的数据以及运算指令并行执行中间运算得到多个中间结果,并将多个中间结果传输给主处理电路;主处理电路还用于对多个中间结果执行后续处理得到计算指令的计算结果。
可选地,该控制器单元110可以包括指令存储单元111、指令处理单元112和存储队列单元114;指令存储单元111用于存储机器学习模型关联的计算指令;指令处理单元112用于对计算指令解析得到多个运算指令;存储队列单元114用于存储指令队列,该指令队列包括:按该队列的前后顺序待执行的多个运算指令或计算指令。可选地,该控制器单元110还可以包括依赖关系处理单元113,用于在具有多个运算指令时,确定第一运算指令与第一运算指令之前的第零运算指令是否存在关联关系;如第一运算指令与第零运算指令存在关联关系,则将第一运算指令缓存在指令存储单元内,在第零运算指令执行完毕后,从指令存储单元提取第一运算指令传输至运算单元。具体地,若依赖关系处理单元113依据第一运算指令提取第一运算指令中所需数据(例如矩阵)的第一存储地址区间,依据第零运算指令提取第零运算指令中所需矩阵的第零存储地址区间;如第一存储地址区间与第零存储地址区间具有重叠的区域,则确定第一运算指令与第零运算指令具有关联关系;如第一存储地址区间与第零存储地址区间不具有重叠的区域,则确定第一运算指令与第零运算指令不具有关联关系。
在一个实施例中,如图3所示,运算单元120还可以包括分支处理电路123,其中,主处理电路121与分支处理电路123连接,分支处理电路123与多个从处理电路122连接;分支处理电路123,用于执行转发主处理电路121与从处理电路122之间的数据或指令。在此实施例中,主处理电路121具体用于将一个输入神经元分配成多个数据块,将多个数据块中的至少一个数据块、权值以及多个运算指令中的至少一个运算指令发送给分支处理电路;分支处理电路123用于转发主处理电路121与多个从处理电路122之间的数据块、权值以及运算指令;多个从处理电路122用于依据该运算指令对接收到的数据块以及权值执行运算得到中间结果,并将中间结果传输给分支处理电路123;主处理电路121还 用于将分支处理电路发送的中间结果进行后续处理得到该计算指令的结果,将该计算指令的结果发送给所述控制器单元。
在另一种可选实施例中,如图4所示,运算单元120可以包括一个主处理电路121和多个从处理电路122。其中,多个从处理电路呈阵列分布;每个从处理电路与相邻的其他从处理电路连接,主处理电路连接多个从处理电路中的k个从处理电路,k个从处理电路为:第1行的n个从处理电路、第m行的n个从处理电路以及第1列的m个从处理电路,需要说明的是,如图1C所示的K个从处理电路仅包括第1行的n个从处理电路、第m行的n个从处理电路以及第1列的m个从处理电路,即该k个从处理电路为多个从处理电路中直接与主处理电路连接的从处理电路。K个从处理电路用于在主处理电路以及多个从处理电路之间的数据以及指令的转发。
可选地,上述的主处理电路121可以包括转换处理电路、激活处理电路、加法处理电路中的一种或任意组合;转换处理电路用于将主处理电路接收的数据块或中间结果执行第一数据结构与第二数据结构之间的互换(例如连续数据与离散数据的转换);或将主处理电路接收的数据块或中间结果执行第一数据类型与第二数据类型之间的互换(例如定点类型与浮点类型的转换);激活处理电路用于执行主处理电路内数据的激活运算;加法处理电路用于执行加法运算或累加运算。
该第一存储器200或第二存储器300可以存储有计算机程序,该计算机程序用于实现本申请实施例中提供的模型转换方法。具体地,该模型转换方法用于在计算机设备接收的初始离线模型的模型属性信息与该计算机设备的硬件属性信息不匹配时,将该初始离线模型转换为与该计算机设备的硬件属性信息相匹配的目标离线模型,以便该计算机设备能够运行该初始离线模型。上述的模型转换方法中,计算机设备的数据输入量少,且无需人为干预即可自动完成初始离线模型到目标离线模型的转换,转换过程简单,转换效率高。进一步地,第一存储器200或第二存储器300上存储的计算机程序还能够用于实现本申请实施例中提供的离线模型生成方法。
可选地,第一存储器200可以用于存储网络模型运行过程中的相关数据,如网络输入数据、网络输出数据、网络权值及指令等等。该第一存储器200可以是内存储器,如缓存等易失性存储器。该第二存储器300可以用于存储网络模型对应的离线模型,例如,第二存储器300可以是非易失性存储器。
上述计算机设备的工作原理与下文中的模型转换方法中各个步骤的执行过程一致,该计算机设备的处理器100在执行存储器200中的计算机程序时,实现模型转换方法中的各个步骤,具体可参见下文中的描述。
当然,在其他实施例中,该计算机设备还可以包含处理器和一个存储器,该计算机设备可以包含处理器与该处理器连接的存储器。该处理器可以采用图2-4所示的处理器,其具体结构可参见上文中关于处理器100的描述。该存储器可以包括多个存储单元,例如该存储器可以包括第一存储单元、第二存储单元和第三存储单元,其中,该第一存储单元可以用于存储计算机程序,该计算机程序用于实现本申请实施例中提供的模型转换方法。该第二存储单元可以用于存储原始网络运行过程中相关数据,该第三存储单元用于存储原始网络对应的离线模型及其该离线模型对应的目标离线模型及预设的模型转换规则等等。进一步地,该存储器包含的存储单元的数量还可以大于三个,此处不做具体限定。
如图5所示,本申请实施例提供了一种模型转换方法,可应用于上述的计算机设备中。 该模型转换方法用于在计算机设备接收的初始离线模型的模型属性信息与该计算机设备的硬件属性信息不匹配时,将该初始离线模型转换为与该计算机设备的硬件属性信息相匹配的目标离线模型,以便该计算机设备能够运行该初始离线模型。具体地,上述方法可以包括如下步骤:
S100:获取初始离线模型及计算机设备的硬件属性信息;
具体地,上述的初始离线模型是指计算机设备直接获取的离线模型,该初始离线模型可以存储于第二存储器300中。其中,离线模型可以包含一原始网络中各个计算节点的网络权值以及指令等必要的网络结构信息,指令可以用于表明该计算节点用于执行何种计算功能,其具体可以包括该原始网络中各个计算节点的计算属性以及各个计算节点之间的连接关系等信息。该原始网络可以是网络模型,如神经网络模型等等,参见图9所示。因此,计算机设备可以通过运行该原始网络对应的离线模型,实现该原始网络的运算功能,无需重复对同一原始网络进行编译等操作,从而可以缩短处理器运行该网络时的运行时间,进而提高处理器的处理速度及效率。可选地,本申请实施例中的初始离线模型可以是根据原始网络直接生成的离线模型,也可以是具有其他模型属性信息的离线模型经过一次或多次转换后获得的离线模型,此处不做具体限定。
S200:根据初始离线模型和计算机设备的硬件属性信息,判断初始离线模型的模型属性信息与计算机设备的硬件属性信息是否匹配。
具体地,该初始离线模型的模型属性信息可以包括该初始离线模型的数据结构及数据类型等等。例如,初始离线模型中各个网络权值及指令的摆放方式、各个网络权值的类型及各个指令的类型等等。上述的计算机设备的硬件属性信息包括该计算机设备的型号、该计算机设备能够处理的数据类型(如定点或浮点等等)及数据结构等等。计算机设备可以根据上述初始离线模型的模型属性信息和计算机设备的硬件属性信息,判断该计算机设备能否支持上述初始离线模型的运行(即该计算机设备能否运行上述的初始离线模型),从而确定该初始离线模型的模型属性信息与计算机设备的硬件属性信息是否匹配。
可选地,当该计算机设备能够支持上述初始离线模型的运行时,则判定该初始离线模型的模型属性信息与计算机设备的硬件属性信息匹配。当该计算机设备不支持上述初始离线模型的运行时,则判定该初始离线模型的模型属性信息与该计算机设备的硬件属性信息不匹配。
例如,该计算机设备可以根据其硬件属性信息确定该计算机设备能够处理的数据类型及数据结构,根据该初始离线模型的模型属性信息确定该初始离线模型的数据类型及数据结构。若该计算机设备根据其自身的硬件属性信息以及上述初始离线模型的模型属性信息,确定该计算机设备能够支持上述初始离线模型的数据类型及数据结构时,则可以判定该初始离线模型的模型属性信息与该计算机设备的硬件属性信息匹配。若该计算机设备根据其自身的硬件属性信息以及上述的初始离线模型的模型属性信息,确定该计算机设备不能够支持上述初始离线模型的数据类型及数据结构时,则可以判定该初始离线模型的模型属性信息与该计算机设备的硬件属性信息不匹配。
若初始离线模型的模型属性信息与计算机设备的硬件属性信息不匹配,则执行步骤S300:根据计算机设备的硬件属性信息和预设的模型转换规则,将初始离线模型转换为与计算机设备的硬件属性信息匹配的目标离线模型。
具体地,若该初始离线模型的模型属性信息与该计算机设备的硬件属性信息不匹配, 则说明该计算机设备不能支持该初始离线模型的运行,此时需对该初始离线模型进行转换处理,以便该计算机设备能够使用该初始离线模型实现相应的运算。即当该初始离线模型的模型属性信息与计算机设备的硬件属性信息不匹配时,则可以根据计算机设备的硬件属性信息和预设的模型转换规则,将上述的初始离线模型转换为与该计算机设备的硬件属性信息匹配的目标离线模型。其中,该目标离线模型的模型属性信息与该计算机设备的硬件属性信息匹配,该计算机设备能够支持该目标离线模型的运行。同上,该目标离线模型也包括一原始网络中各个计算节点的网络权值以及指令等必要的网络结构信息。
可选地,上述的预设的模型转换规则可以预先存储于计算机设备的第一存储器或第二存储器中。若初始离线模型的模型属性信息与计算机设备的硬件属性信息不匹配时,计算机设备可以从第一存储器或第二存储器中获得相应的模型转换规则,以将初始离线模型转换为目标离线模型。可选地,上述预设的模型转换规则可以是一个以上,一个以上的模型转换规则可以与上述的初始离线模型及目标离线模型一一对应存储。例如,上述的初始离线模型、目标离线模型以及预设的模型转换规则可以通过映射表的方式进行对应存储。可选地,初始离线模型的模型属性信息、目标离线模型的模型属性信息及可用模型转换规则三者之间一一对应存储。
若初始离线模型的模型属性信息与计算机设备的硬件属性信息匹配,则可以执行步骤S400:计算机设备可以直接运行其接收到的初始离线模型,即计算机设备可以根据该初始离线模型包含的网络权值及指令进行运算,以实现原始网络的运算功能。具体地,若该初始离线模型的模型属性信息与该计算机设备的硬件属性信息匹配,则说明该计算机设备能够支持该初始离线模型的运行,此时无需对该初始离线模型进行转换处理,计算机设备能够运行该初始离线模型。应当清楚的是,本实施例中,直接运行该初始离线模型是指,使用该初始离线模型运行该原始网络对应的机器学习算法(如神经网络算法),通过执行前向运算实现算法的目标应用(如语音识别等人工智能应用)。
本申请实施例中的模型转换方法,计算机设备只需接收一个初始离线模型,即可根据预设的模型转换规则和该计算机设备的硬件属性信息,将该初始离线模型转换为目标离线模型,无需获得多个不同的模型数据,极大的降低了计算机设备的数据输入量,避免了数据输入量过大导致的超出计算机设备的存储容量等问题,保证计算机设备的正常运行。同时,上述模型转换方法可以降低该计算机设备的数据处理量,进而可以提高处理效率,降低功耗。同时,上述的模型转换过程中,无需人工干预,自动化程度高,使用便捷。
可选地,如图6所示,上述的步骤S300可以包括:
S310:根据计算机设备的硬件属性信息确定目标离线模型的模型属性信息。
具体地,该计算机设备可以根据其自身的硬件属性信息,确定该计算机设备能够支持的数据类型及数据结构等等,从而可以确定目标离线模型的数据类型及数据结构等模型属性信息。即该计算机设备能够根据其硬件属性信息确定目标离线模型的模型属性信息,该目标离线模型的模型属性信息可以包括目标离线模型的数据结构及数据类型等信息。
S320:根据初始离线模型的模型属性信息和目标离线模型的模型属性信息,从多个预设的模型转换规则中选择一个模型转换规则作为目标模型转换规则。
具体地,由于初始离线模型、目标离线模型和预设的模型转换规则存在一一对应的映射关系,因此,可以根据该初始离线模型的模型属性信息和目标离线模型的模型属性信息确定出目标模型转换规则。其中,该模型转换规则可以包括数据类型的转换方法及数据结 构的转换方法等等。
S330:根据初始离线模型的模型属性信息和目标模型转换规则,将初始离线模型转换为目标离线模型。
具体地,计算机设备可以根据该目标模型转换规则提供的转换方法,将初始离线模型转换为目标离线模型,从而使得该计算机设备能够运行该目标离线模型进行运算。
可选地,若初始离线模型和目标离线模型之间存在多个可用的模型转换规则时,计算机设备可以通过根据预设算法自动选定目标离线模型。具体地,上述的步骤S320进一步包括如下步骤:
根据初始离线模型的模型属性信息和目标离线模型的模型属性信息,从多个预设的模型转换规则中选择出一个以上的可用模型转换规则。
对一个以上的可用模型转换规则进行优先级排序,并将优先级最高的可用模型转换规则作为目标转换规则。
具体地,上述的可用模型转换规则是指能够将初始离线模型转换为目标离线模型的转换规则。当存在一个以上的可用模型转换规则时,说明将初始离线模型转换为目标离线模型的方式有多种不同的方式。上述的优先级排序方法可以是预设的,也可以是用户自定义的。
可选地,计算机设备可以获取各个可用模型转换规则对应的过程参数,并根据各个可用模型转换规则对应的过程参数,对一个以上的可用模型转换规则进行优先级排序。其中,该过程参数可以是采用该可用模型转换规则将初始离线模型转换为目标离线模型时所涉及的计算机设备的性能参数。可选地,该过程参数可以包括转换速度、功耗、内存占用率及磁盘I/O占用率中的一种或多种。
例如,计算机设备可以根据将初始离线模型转换为目标离线模型过程中的转换速度、转换功耗、内存占用率及磁盘I/O占用率等一种或多种过程参数的组合,对上述的各个可用模型转换规则进行评分(如对各个参考因素进行加权计算以获得评分值),并将评分最高的可用模型转换规则作为优先级最高的可用模型转换规则。即可以将评分最高的可用模型转换规则作为该目标转换规则。之后,可以执行上述步骤S330,根据初始离线模型的模型属性信息和目标模型转换规则,将初始离线模型转换为目标离线模型。这样,可以通过结合模型转换过程中的设备性能因素选择一种较优的模型转换规则,以提高计算机设备的处理速度及效率。
可选地,如图3所示,上述方法还包括如下步骤:
S500:将目标离线模型存储于计算机设备的第一存储器或第二存储器。
具体地,计算机设备可以将其获得的目标离线模型存储于该计算机设备的本地存储器上(如第一存储器)。可选地,计算机设备还可以将其获得目标离线模型存储于与该计算机设备连接的外部存储器(如第二存储器),该外部存储器可以是云端存储器或其他存储器等等。这样,当该计算机设备需要重复使用该目标离线模型时,无需重复执行上述的转换操作,只需直接读取该目标离线模型即可实现相应的运算。
可选地,计算机设备执行上述的目标离线模型的过程可以包括如下步骤:
获取目标离线模型,并运行目标离线模型,以执行原始网络。其中,目标离线模型中包含原始网络中各个离线节点对应的网络权值、指令以及各个离线节点与原始网络中的其他计算节点之间的接口数据等必要的网络结构信息。
具体地,当计算机设备需要利用该目标离线模型实现运算时,其可以直接从第一存储器或第二存储器中获取该目标离线模型,并根据目标离线模型中的网络权值及指令等进行运算,从而实现原始网络的运算功能。这样,当该计算机设备需要重复使用该目标离线模型时,无需重复执行上述的转换操作,只需直接读取该目标离线模型即可实现相应的运算。
应当清楚的是,本实施例中,直接运行该目标离线模型是指,使用该目标离线模型运行该原始网络对应的机器学习算法(如神经网络算法),通过执行前向运算实现算法的目标应用(如语音识别等人工智能应用)。
可选地,该计算机设备上安装有应用软件(Application),可以通过计算机设备上的应用软件获取初始离线模型,该应用软件(Application)能够提供从存储器200或外部存储器中读取初始离线模型的接口,从而可以通过该应用软件获取该初始离线模型。可选地,该应用软件还可以提供读取计算机设备的硬件属性信息的接口,从而可以通过该应用软件获取该计算机设备的硬件属性信息。可选地,该应用软件还可以提供能够读取预设的模型规则的接口,从而可以通过该应用软件获取预设的模型转换规则。
当然,在其他实施例中,该计算机设备还可以提供输入/输出接口(如I/O接口)等,该输入/输出接口用于获取初始离线模型,或输出目标离线模型。该计算机设备的硬件属性信息可以是预设存储于该计算机设备中的。
在一个实施例中,如图7所示,本申请实施例还提供了一种根据原始网络模型生成离线模型的方法,用于根据获取的原始网络的相关数据生成并存储该原始网络的离线模型,从而在处理器再次运行该原始网络时,可以直接运行该原始网络对应的离线模型,无需再次对同一原始网络进行编译等操作,从而缩短处理器运行该网络时的运行时间,进而提高处理器的处理速度及效率。其中,原始网络可以是神经网络或非神经网络等网络模型,例如,原始网络可以是图9所示的网络。具体地,上述方法包括如下步骤:
S010:获取原始网络的模型数据集及模型结构参数。
具体地,计算机设备的处理器可以获取原始网络的模型数据集及模型结构参数,通过该原始网络的模型数据集及模型结构参数可以获得该原始网络的网络结构图。其中,模型数据集包括原始网络中各个计算节点对应的网络权值等数据,图9所示的神经网络中的W1~W6即用于表示计算节点的网络权值。模型结构参数包括原始网络中多个计算节点的连接关系及各个计算节点的计算属性,其中,计算节点之间的连接关系用于表示计算节点之间是否有数据传递,例如,当多个计算节点之间具有数据流的传递时,则可以说明多个计算节点之间具有连接关系。进一步地,计算节点的连接关系可以包括输入关系和输出关系等等。如图9所示,计算节点F1输出作为计算节点F4和F5的输入,则可以说明计算节点F1和计算节点F4之间具有连接关系,计算节点F1和计算节点F4之间具有连接关系。再如,计算节点和计算节点F2之间没有数据传递,则可以说明计算节点F1和计算节点F2之间不存在连接关系。
各个计算节点的计算属性可以包括相应计算节点的计算类型及计算参数,其中计算节点的计算类型是指该计算节点用于完成何种计算,如计算节点的计算类型可以包括加法运算、减法运算及卷积运算等等,相应的,该计算节点可以是用于实现加法运算的计算节点、用于实现减法运算的计算节点或用于实现卷积运算的计算节点等等。计算节点的计算参数可以是完成该计算节点对应的计算类型所需的必要参数。例如,计算节点的计算类型可以是用于实现加法运算的计算节点,相应的,该计算节点的计算参数可以为加法运算中的加 数,该加法运算中的被加数可以作为输入数据通过获取模块获取,或者,该加法运算中的被加数可以是该计算节点的上一计算节点的输出数据等等。
S020:根据原始网络的模型数据集和模型结构参数运行原始网络,获得原始网络中各个计算节点对应的指令。
具体地,计算机设备的处理器可以根据原始网络的模型数据集和模型结构参数运行该原始网络,并获得原始网络中各个计算节点对应的指令。进一步地,处理器还可以获取该原始网络的输入数据,根据原始网络的输入数据、网络模型数据集和模型结构参数运行原始网络,获得该原始网络中各个计算节点对应的指令。更进一步地,上述运行该原始网络获得各个计算节点的指令的过程实质上是编译的过程,该编译过程可以通过计算机设备的处理器或虚拟设备实现。即计算机设备的处理器或虚拟设备根据原始网络的模型数据集和模型结构参数运行原始网络。其中,虚拟设备指的是在存储器的内存空间中虚拟出一段处理器运行空间。
应当清楚的是,本实施例中的运行原始网络是指,处理器使用人工神经网络模型数据运行某种机器学习算法(如神经网络算法),通过执行前向运算实现算法的目标应用(如语音识别等人工智能应用)。
S030:根据原始网络的各个计算节点对应的网络权值及指令,生成原始网络对应的离线模型,并将原始网络对应的离线模型存储至非易失性存储器(数据库)中。
具体地,该处理器的控制模块可以根据原始网络的各个计算节点对应的网络权值和指令,生成该原始网络对应的离线模型,例如,该处理器的控制模块可以将原始网络的各个计算节点对应的网络权值和指令存储至非易失性存储器中,以实现离线模型的生成及存储。其中,针对原始网络的每个计算节点,该计算节点的网络权值及指令二者之间一一对应进行存储。这样,当再次运行该原始网络时,可以直接从非易失性存储器中获取该原始网络对应的离线模型,并根据与其对应的离线模型运行原始网络,无需在线对该原始网络的各个计算节点进行编译获得指令,提高了系统的运行速度及效率。
应当清楚的是,本实施例中,直接运行该原始网络对应的离线模型是指,使用离线模型运行该原始网络对应的机器学习算法(如神经网络算法),通过执行前向运算实现算法的目标应用(如语音识别等人工智能应用)。
可选地,如图8所示,上述步骤S200可以包括:
S021:根据原始网络的模型结构参数,获得原始网络中各个计算节点的执行顺序。
具体地,处理器可以根据原始网络的模型结构参数,获得原始网络中各个计算节点的执行顺序,根据原始网络中各个计算节点的连接关系,获得原始网络中各个计算节点的执行顺序。例如,如图9所示,计算节点F4的输入数据为计算节点F1的输出数据以及计算节点F2的输出数据,计算节点F6的输入数据为计算节点F4的输出数据以及计算节点F5的输出数据。因此,图9所示的神经网络中各个计算节点的执行顺序可以为F1-F2-F3-F4-F5-F6或F1-F3-F2-F5-F4-F6等等。当然,计算节点F1、F2和F3可以并行执行,计算节点F4和F5也可以并行执行,此处仅举例说明,并不具体限定其执行顺序。
S022:按照原始网络中各个计算节点的执行顺序运行原始网络,分别获得原始网络中各个计算节点对应的指令。
具体地,处理器可以根据原始网络中各个计算节点的执行顺序运行该原始网络,以获得原始网络中各个计算节点对应的指令,即处理器可以将原始网络的模型数据集等数据进 行编译获得各个计算节点对应的指令,通过各个计算节点对应的指令可以获知该计算节点用于实现何种计算功能,即可以获得该计算节点的计算类型及计算参数等计算属性。
进一步地,如图8所示,上述步骤S300还包括:
S031:根据原始网络的模型数据集和模型结构参数,获得原始网络的内存分配方式。
具体地,处理器可以根据原始网络的模型数据集和模型结构参数,获得原始网络的内存分配方式;根据原始网络的模型结构参数,获得原始网络中各个计算节点的执行顺序;并根据原始网络中各个计算节点的执行顺序确定当前网络的内存分配方式。例如,按各个计算节点的执行顺序将各个计算节点在运行过程中的相关数据保存至一个栈内。其中,内存分配方式是指确定原始网络中各个计算节点相关的数据(包括输入数据、输出数据、网络权值数据及中间结果数据等等)在内存空间(如第一存储器)上的存储位置。例如,可以采用数据表存储各个计算节点相关的数据(输入数据、输出数据、网络权值数据及中间结果数据等等)和内存空间的映射关系。
S032:根据原始网络的内存分配方式,将原始网络运行过程中的相关数据存储至第一存储器中,其中,原始网络运行过程中的相关数据包括原始网络的各个计算节点对应的网络权值、指令、输入数据、中间计算结果及输出数据等等。例如,如图9所示,X1和X2表示该神经网络的输入数据,Y表示该神经网络的输出数据,处理器可以将该神经网络的输出数据转换为控制机器人或不同数字接口的控制命令。W1~W6用于表示计算节点F1、F2和F3对应的网络权值,计算节点F1~F5的输出数据可以作为中间计算结果。处理器可以根据已确定的内存分配方式,将原始网络运行过程中的相关数据存储至第一存储器,如内存储器或缓存等易失性存储器,其具体的存储方式可参见图10中左半部分存储空间。
S033:从第一存储器中获取原始网络的各个计算节点对应的网络权值及指令,并将原始网络的各个计算节点对应的网络权值及指令存储于第二存储器中,生成离线模型。其中,第二存储器可以为外部存储器等非易失性存储器。该离线模型的生成过程具体可参见图10所示,图10中右半部分的存储空间内存储的即为原始网络的对应的离线模型。
如图9和图10所示,下面结合附图说明上述的离线模型生成过程:
首先,处理器可以获得该原始网络的模型数据集、模型结构参数以及输入数据,从而根据该原始网络的模型数据集和模型结构参数可以获得该原始网络的网络结构图,如图9所示。
其次,处理器可以根据原始网络的模型结构参数,获得原始网络各个计算节点的连接关系,并根据各个计算节点的连接关系获得原始网络中各个计算节点的执行顺序,以及原始网络在运行过程中的内存分配方式,从而可以获得原始网络在运行过程中相关数据的存储位置。如图10中左半部分存储空间所示,原始网络在运行过程中的相关数据可以按照各个计算节点执行顺序存储在一个栈中。
最后,处理器可以将原始网络的各个计算节点对应的网络权值及指令存储于非易失性的第二存储器中,生成离线模型,该离线模型的存储方式可参见图8中右半部分存储空间所示。并且,该离线模型仅仅包含运行该原始网络所必需的网络权值及指令等数据,而不需对原始网络运行过程中的输入数据、输出数据或中间计算结果等进行存储,从而可以减小第二存储器中的存储空间的消耗。
作为进一步地改进,离线模型中还包括节点接口数据,节点接口数据用于表示原始网络的各个计算节点的连接关系。具体地,节点接口数据可以包括各个计算节点的输入数据 来源和输出数据来源。例如,如图9所示,节点接口数据可以包括计算节点F1、F2和F3为起始计算节点,其输入分别为预设的输入数据,计算节点F1的输出数据作为计算节点F4和计算节点F5的输入数据等等。这样,在再次运行该原始网络时,只需获得原始网络的起始计算节点和输入数据,之后,便可以根据该原始网络对应的离线模型执行该原始网络。
应该理解的是,虽然图5-8的流程图中的各个步骤按照箭头的指示依次显示,但是这些步骤并不是必然按照箭头指示的顺序依次执行。除非本文中有明确的说明,这些步骤的执行并没有严格的顺序限制,这些步骤可以以其它的顺序执行。而且,图5-8中的至少一部分步骤可以包括多个子步骤或者多个阶段,这些子步骤或者阶段并不必然是在同一时刻执行完成,而是可以在不同的时刻执行,这些子步骤或者阶段的执行顺序也不必然是依次进行,而是可以与其它步骤或者其它步骤的子步骤或者阶段的至少一部分轮流或者交替地执行。
本领域普通技术人员可以理解实现上述实施例方法中的全部或部分流程,是可以通过计算机程序来指令相关的硬件来完成,所述的计算机程序可存储于一非易失性计算机可读取存储介质中,该计算机程序在执行时,可包括如上述各方法的实施例的流程。其中,本申请所提供的各实施例中所使用的对存储器、存储、数据库或其它介质的任何引用,均可包括非易失性和/或易失性存储器。非易失性存储器可包括只读存储器(ROM)、可编程ROM(PROM)、电可编程ROM(EPROM)、电可擦除可编程ROM(EEPROM)或闪存。易失性存储器可包括随机存取存储器(RAM)或者外部高速缓冲存储器。作为说明而非局限,RAM以多种形式可得,诸如静态RAM(SRAM)、动态RAM(DRAM)、同步DRAM(SDRAM)、双数据率SDRAM(DDRSDRAM)、增强型SDRAM(ESDRAM)、同步链路(Synchlink)DRAM(SLDRAM)、存储器总线(Rambus)直接RAM(RDRAM)、直接存储器总线动态RAM(DRDRAM)、以及存储器总线动态RAM(RDRAM)等。
本申请实施例还提供了一种模型转换装置,上述模型转换装置包括获取模块、判断模块和转换模块。其中,
获取模块用于获取初始离线模型及计算机设备的硬件属性信息;
判断模块用于根据初始离线模型和计算机设备的硬件属性信息,判断初始离线模型的模型属性信息与计算机设备的硬件属性信息是否匹配;
转换模块用于在初始离线模型的模型属性信息与计算机设备的硬件属性信息不匹配时,根据计算机设备的硬件属性信息和预设的模型转换规则,将初始离线模型转换为与计算机设备的硬件属性信息匹配的目标离线模型。
可选地,该模型转换装置可以是安装于该计算机设备上的应用软件(Application)。该应用软件(Application)能够提供从存储器200或外部存储器中读取初始离线模型的接口,从而可以通过该应用软件获取该初始离线模型。可选地,该应用软件还可以提供读取计算机设备的硬件属性信息的接口,从而可以通过该应用软件获取该计算机设备的硬件属性信息。可选地,该应用软件还可以提供能够读取预设的模型规则的接口,从而可以通过该应用软件获取预设的模型转换规则。进一步地,该应用软件可以根据初始离线模型和计算机设备的硬件属性信息以及预设的模型换换规则,在该初始离线模型的模型属性信息与计算机设备的硬件属性信息不匹配时,将该初始离线模型转换为与该计算机设备的硬件属性信息匹配的目标离线模型,并将该目标离线模型存储至第一存储器或第二存储器中。
关于模型转换装置的具体限定可以参见上文中对于模型转换方法的限定,在此不再赘述。上述模型转换装置中的各个模块可全部或部分通过软件、硬件及其组合来实现。上述各模块可以硬件形式内嵌于或独立于计算机设备中的处理器中,也可以以软件形式存储于计算机设备中的存储器中,以便于处理器调用执行以上各个模块对应的操作。
此外,本申请实施例还提供了一种计算机可读存储介质,其上存储有计算机程序,上述计算机程序被处理器执行时实现权利要求上述任一实施例的方法的步骤。可选地,该计算机可读存储介质可以包括非易失性和/或易失性存储器。非易失性存储器可包括只读存储器(ROM)、可编程ROM(PROM)、电可编程ROM(EPROM)、电可擦除可编程ROM(EEPROM)或闪存。易失性存储器可包括随机存取存储器(RAM)或者外部高速缓冲存储器。作为说明而非局限,RAM以多种形式可得,诸如静态RAM(SRAM)、动态RAM(DRAM)、同步DRAM(SDRAM)、双数据率SDRAM(DDRSDRAM)、增强型SDRAM(ESDRAM)、同步链路(Synchlink)DRAM(SLDRAM)、存储器总线(Rambus)直接RAM(RDRAM)、直接存储器总线动态RAM(DRDRAM)、以及存储器总线动态RAM(RDRAM)等。
具体地,上述计算机程序被处理器执行时实现如下步骤:
S100:获取初始离线模型及计算机设备的硬件属性信息;
具体地,上述的初始离线模型是指计算机设备直接获取的离线模型,该初始离线模型可以存储于第二存储器300中。其中,离线模型可以包含一原始网络中各个计算节点的网络权值以及指令等必要的网络结构信息,其中,指令可以用于表明该计算节点用于执行何种计算功能,其具体可以包括该原始网络中各个计算节点的计算属性以及各个计算节点之间的连接关系等信息。可选地,本申请实施例中的初始离线模型可以是根据原始网络直接生成的离线模型,也可以是具有其他模型属性信息的离线模型经过一次或多次转换后获得离线模型,此处不做具体限定。
S200:根据初始离线模型和计算机设备的硬件属性信息,判断初始离线模型的模型属性信息与计算机设备的硬件属性信息是否匹配。
具体地,该初始离线模型的模型属性信息可以包括该初始离线模型的数据结构及数据类型等等。例如,初始离线模型中各个网络权值及指令的摆放方式、各个网络权值的类型及各个指令的类型等等。上述的计算机设备的硬件属性信息包括该计算机设备的型号、该计算机设备能够处理的数据类型(如定点或浮点等等)及数据结构等等。
例如,该计算机设备可以根据其硬件属性信息确定该计算机设备能够处理的数据类型及数据结构,根据该初始离线模型的模型属性信息确定该初始离线模型的数据类型及数据结构。若该计算机设备根据其自身的硬件属性信息以及上述初始离线模型的模型属性信息,确定该计算机设备能够支持上述初始离线模型的数据类型及数据结构时,则可以判定该初始离线模型的模型属性信息与该计算机设备的硬件属性信息匹配。若该计算机设备根据其自身的硬件属性信息以及上述的初始离线模型的模型属性信息,确定该计算机设备不能够支持上述初始离线模型的数据类型及数据结构时,则可以判定该初始离线模型的模型属性信息与该计算机设备的硬件属性信息不匹配。
若初始离线模型的模型属性信息与计算机设备的硬件属性信息不匹配,则执行步骤S300:根据计算机设备的硬件属性信息和预设的模型转换规则,将初始离线模型转换为与计算机设备的硬件属性信息匹配的目标离线模型。
具体地,若该初始离线模型的模型属性信息与该计算机设备的硬件属性信息不匹配,则说明该计算机设备不能支持该初始离线模型的运行,此时需对该初始离线模型进行转换处理,以便该计算机设备能够使用该初始离线模型实现相应的运算。即当该初始离线模型的模型属性信息与计算机设备的硬件属性信息不匹配时,则可以根据计算机设备的硬件属性信息和预设的模型转换规则,将上述的初始离线模型转换为与该计算机设备的硬件属性信息匹配的目标离线模型。其中,该目标离线模型的模型属性信息与该计算机设备的硬件属性信息匹配,该计算机设备能够支持该目标离线模型的运行。同上,该目标离线模型也包括一原始网络中各个计算节点的网络权值以及指令等必要的网络结构信息。
若初始离线模型的模型属性信息与计算机设备的硬件属性信息匹配,则可以执行步骤S400:计算机设备可以直接运行其接收到的初始离线模型,即计算机设备可以根据该初始离线模型包含的网络权值及指令进行运算,以实现原始网络的运算功能。具体地,若该初始离线模型的模型属性信息与该计算机设备的硬件属性信息匹配,则说明该计算机设备能够支持该初始离线模型的运行,此时无需对该初始离线模型进行转换处理,计算机设备能够运行该初始离线模型。
应当清楚的是,处理器执行上述计算机程序的过程,与上述实施例的模型转换方法的执行过程一致,具体可参见上文中的描述,此处不再赘述。
可以理解的是,上述各实施例中相同或相似部分可以相互参考,在一些实施例中未详细说明的内容可以参见其他实施例中相同或相似的内容。
需要说明的是,在本申请的描述中,术语“第一”、“第二”等仅用于描述目的,而不能理解为指示或暗示相对重要性。此外,在本申请的描述中,除非另有说明,“多个”的含义是指至少两个。
流程图中或在此以其他方式描述的任何过程或方法描述可以被理解为,表示包括一个或更多个用于实现特定逻辑功能或过程的步骤的可执行指令的代码的模块、片段或部分,并且本申请的可选实施方式的范围包括另外的实现,其中可以不按所示出或讨论的顺序,包括根据所涉及的功能按基本同时的方式或按相反的顺序,来执行功能,这应被本申请的实施例所属技术领域的技术人员所理解。
应当理解,本申请的各部分可以用硬件、软件、固件或它们的组合来实现。在上述实施方式中,多个步骤或方法可以用存储在存储器中且由合适的指令执行系统执行的软件或固件来实现。例如,如果用硬件来实现,和在另一实施方式中一样,可用本领域公知的下列技术中的任一项或它们的组合来实现:具有用于对数据信号实现逻辑功能的逻辑门电路的离散逻辑电路,具有合适的组合逻辑门电路的专用集成电路,可编程门阵列(PGA),现场可编程门阵列(FPGA)等。
本技术领域的普通技术人员可以理解实现上述实施例方法携带的全部或部分步骤是可以通过程序来指令相关的硬件完成,所述的程序可以存储于一种计算机可读存储介质中,该程序在执行时,包括方法实施例的步骤之一或其组合。
此外,在本申请各个实施例中的各功能单元可以集成在一个处理模块中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个模块中。上述集成的模块既可以采用硬件的形式实现,也可以采用软件功能模块的形式实现。所述集成的模块如果以软件功能模块的形式实现并作为独立的产品销售或使用时,也可以存储在一个计算机可读取存储介质中。
上述提到的存储介质可以是只读存储器,磁盘或光盘等。
在本说明书的描述中,参考术语“一个实施例”、“一些实施例”、“示例”、“具体示例”、或“一些示例”等的描述意指结合该实施例或示例描述的具体特征、结构、材料或者特点包含于本申请的至少一个实施例或示例中。在本说明书中,对上述术语的示意性表述不一定指的是相同的实施例或示例。而且,描述的具体特征、结构、材料或者特点可以在任何的一个或多个实施例或示例中以合适的方式结合。
尽管上面已经示出和描述了本申请的实施例,可以理解的是,上述实施例是示例性的,不能理解为对本申请的限制,本领域的普通技术人员在本申请的范围内可以对上述实施例进行变化、修改、替换和变型。

Claims (13)

  1. 一种模型转换方法,其特征在于,所述方法包括如下步骤:
    获取初始离线模型及计算机设备的硬件属性信息;
    根据所述初始离线模型和所述计算机设备的硬件属性信息,判断所述初始离线模型的模型属性信息与所述计算机设备的硬件属性信息是否匹配;
    若所述初始离线模型的模型属性信息与所述计算机设备的硬件属性信息不匹配,则根据所述计算机设备的硬件属性信息和预设的模型转换规则,将所述初始离线模型转换为与所述计算机设备的硬件属性信息匹配的目标离线模型。
  2. 根据权利要求1所述的模型转换方法,其特征在于,所述的根据所述计算机设备的硬件属性信息和预设的模型转换规则,将所述初始离线模型转换为与所述计算机设备的硬件属性信息匹配的目标离线模型的步骤,包括:
    根据所述计算机设备的硬件属性信息确定所述目标离线模型的模型属性信息;
    根据所述初始离线模型的模型属性信息和所述目标离线模型的模型属性信息,从多个所述预设的模型转换规则中选择一个所述模型转换规则作为目标模型转换规则;
    根据所述初始离线模型的模型属性信息和所述目标模型转换规则,将所述初始离线模型转换为所述目标离线模型。
  3. 根据权利要求2所述的模型转换方法,其特征在于,所述的根据所述初始离线模型的模型属性信息和所述目标离线模型的模型属性信息,从多个所述预设的模型转换规则中选择一个所述模型转换规则作为目标模型转换规则的步骤,包括:
    根据所述初始离线模型的模型属性信息和所述目标离线模型的模型属性信息,从多个所述预设的模型转换规则中选择出一个以上的可用模型转换规则;
    对所述一个以上的可用模型转换规则进行优先级排序,并将优先级最高的可用模型转换规则作为所述目标转换规则。
  4. 根据权利要求3所述的模型转换方法,其特征在于,所述的对一个以上的可用模型转换规则进行优先级排序的步骤,包括如下步骤:
    分别获取采用各个所述可用模型转换规则,将所述初始离线模型转换为所述目标离线模型的过程参数,其中,所述过程参数包括转换速度、功耗、内存占用率及磁盘I/O占用率中的一种或多种;
    根据各个所述可用模型转换规则的过程参数,对一个以上的所述可用模型转换规则进行优先级排序。
  5. 根据权利要求3所述的模型转换方法,其特征在于,所述初始离线模型的模型属性信息、所述目标离线模型的模型属性信息及所述可用模型转换规则三者之间一一对应存储。
  6. 根据权利要求1-5任一项所述的模型转换方法,其特征在于,所述的根据所述初始离线模型和所述计算机设备的硬件属性信息,判断所述初始离线模型的模型属性信息与所述计算机设备的硬件属性信息是否匹配的步骤包括:
    若根据所述初始离线模型的模型属性信息和所述计算机设备的硬件属性信息,确定所述计算机设备能够支持所述初始离线模型的运行时,则判定所述初始离线模型的模型属性信息与所述计算机设备的硬件属性信息匹配;
    若根据所述初始离线模型的模型属性信息和所述计算机设备的硬件属性信息,确定所 述计算机设备不支持所述初始离线模型的运行时,则判定所述初始离线模型的模型属性信息与所述计算机设备的硬件属性信息不匹配。
  7. 根据权利要求1-5任一项所述的模型转换方法,其特征在于,所述的获取初始离线模型的步骤,包括:
    通过所述计算机设备上的应用软件获取所述初始离线模型和所述计算机设备的硬件属性信息。
  8. 根据权利要求1-5任一项所述的模型转换方法,其特征在于,所述方法还包括如下步骤:
    将所述目标离线模型存储于所述计算机设备的第一存储器或第二存储器。
  9. 根据权利要求8所述的模型转换方法,其特征在于,所述方法还包括如下步骤:
    获取所述目标离线模型,并运行所述目标离线模型,其中,所述目标离线模型中包含原始网络中各个离线节点对应的网络权值、指令以及各个离线节点与所述原始网络中的其他计算节点之间的接口数据。
  10. 一种模型转换装置,其特征在于,所述装置包括:
    获取模块,用于获取初始离线模型及计算机设备的硬件属性信息;
    判断模块,用于根据所述初始离线模型和所述计算机设备的硬件属性信息,判断所述初始离线模型的模型属性信息与所述计算机设备的硬件属性信息是否匹配;
    转换模块,用于在所述初始离线模型的模型属性信息与所述计算机设备的硬件属性信息不匹配时,根据所述计算机设备的硬件属性信息和预设的模型转换规则,将所述初始离线模型转换为与所述计算机设备的硬件属性信息匹配的目标离线模型。
  11. 一种计算机设备,包括存储器和处理器,所述存储器存储有计算机程序,其特征在于,所述处理器执行所述计算机程序时实现权利要求1至7中任一项所述方法的步骤。
  12. 根据权利要求11所述的计算机设备,其特征在于,所述处理器包括运算单元和控制器单元,所述运算单元包括主处理电路和多个从处理电路;
    所述控制器单元用于获取数据、机器学习模型以及计算指令;
    所述控制器单元还用于解析所述计算指令得到多个运算指令,并将所述多个运算指令以及所述数据发送给所述主处理电路;
    所述主处理电路用于对所述数据以及在所述主处理电路与所述多个从处理电路之间进行传输的数据和运算指令执行前序处理;
    所述多个从处理电路用于依据从所述主处理电路传输的数据以及运算指令并行执行中间运算得到多个中间结果,并将多个中间结果传输给所述主处理电路;
    所述主处理电路还用于对所述多个中间结果执行后续处理得到所述计算指令的计算结果。
  13. 一种计算机可读存储介质,其上存储有计算机程序,其特征在于,所述计算机程序被处理器执行时实现权利要求1至7中任一项所述的方法的步骤。
PCT/CN2019/080510 2018-08-10 2019-03-29 转换方法、装置、计算机设备及存储介质 Ceased WO2020029592A1 (zh)

Priority Applications (7)

Application Number Priority Date Filing Date Title
CA3062949A CA3062949C (en) 2018-08-10 2019-03-29 Conversion method, apparatus, computer device, and storage medium
AU2019268193A AU2019268193B2 (en) 2018-08-10 2019-03-29 Conversion method, apparatus, computer device, and storage medium
JP2019564530A JP6829327B2 (ja) 2018-08-10 2019-03-29 変換方法、装置、コンピューターデバイス及び記憶媒体
KR1020197034745A KR20210033874A (ko) 2018-08-10 2019-03-29 변환 방법, 장치, 컴퓨터 장치 및 저장 매체
EP19790111.9A EP3640825B8 (en) 2018-08-10 2019-03-29 Conversion method, apparatus, computer device, and storage medium
US16/667,593 US11314507B2 (en) 2018-08-10 2019-10-29 Model conversion method, device, computer equipment, and storage medium
US17/703,757 US11853760B2 (en) 2018-08-10 2022-03-24 Model conversion method, device, computer equipment, and storage medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201810913895.6 2018-08-10
CN201810913895.6A CN109492241B (zh) 2018-08-10 2018-08-10 转换方法、装置、计算机设备和存储介质

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US16/667,593 Continuation US11314507B2 (en) 2018-08-10 2019-10-29 Model conversion method, device, computer equipment, and storage medium

Publications (1)

Publication Number Publication Date
WO2020029592A1 true WO2020029592A1 (zh) 2020-02-13

Family

ID=65690403

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2019/080510 Ceased WO2020029592A1 (zh) 2018-08-10 2019-03-29 转换方法、装置、计算机设备及存储介质

Country Status (8)

Country Link
US (2) US11314507B2 (zh)
EP (1) EP3640825B8 (zh)
JP (1) JP6829327B2 (zh)
KR (1) KR20210033874A (zh)
CN (2) CN109492241B (zh)
AU (1) AU2019268193B2 (zh)
CA (1) CA3062949C (zh)
WO (1) WO2020029592A1 (zh)

Families Citing this family (23)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109492241B (zh) * 2018-08-10 2020-03-10 中科寒武纪科技股份有限公司 转换方法、装置、计算机设备和存储介质
CN109754072B (zh) * 2018-12-29 2020-06-23 中科寒武纪科技股份有限公司 网络离线模型的处理方法、人工智能处理装置及相关产品
CN109976751B (zh) * 2019-03-28 2022-12-27 中科寒武纪科技股份有限公司 模型操作方法、相关装置及计算机可读存储介质
US20200334522A1 (en) 2019-04-18 2020-10-22 Cambricon Technologies Corporation Limited Data processing method and related products
CN111832738B (zh) 2019-04-18 2024-01-09 中科寒武纪科技股份有限公司 一种数据处理方法及相关产品
CN111832356B (zh) * 2019-04-19 2024-12-03 中科寒武纪科技股份有限公司 信息处理装置、方法及相关产品
CN112396186B (zh) * 2019-08-12 2024-05-03 上海寒武纪信息科技有限公司 执行方法、装置及相关产品
CN110532291B (zh) * 2019-07-25 2022-07-12 中国科学院计算技术研究所 基于最小执行代价的深度学习框架间模型转换方法及系统
CN110458285B (zh) * 2019-08-14 2021-05-14 中科寒武纪科技股份有限公司 数据处理方法、装置、计算机设备和存储介质
CN113435591B (zh) * 2019-08-14 2024-04-05 中科寒武纪科技股份有限公司 数据处理方法、装置、计算机设备和存储介质
CN111967568B (zh) * 2020-06-29 2023-09-01 北京百度网讯科技有限公司 深度学习模型的适配方法、装置及电子设备
KR102840299B1 (ko) * 2020-06-29 2025-07-30 베이징 바이두 넷컴 사이언스 앤 테크놀로지 코., 엘티디. 딥 러닝 모델의 적응 방법, 장치 및 전자 기기
CN111783642B (zh) * 2020-06-30 2023-10-13 北京百度网讯科技有限公司 一种图像识别方法、装置、电子设备及存储介质
CN111966361B (zh) * 2020-09-25 2024-04-05 北京百度网讯科技有限公司 用于确定待部署模型的方法、装置、设备及其存储介质
CN112256279B (zh) * 2020-09-29 2024-05-14 深圳市广和通无线股份有限公司 源码转换方法、装置、计算机设备及可读存储介质
CN114529005A (zh) * 2020-11-03 2022-05-24 华为技术有限公司 机器学习模型管理方法、装置和系统
CN112799895B (zh) * 2021-01-27 2024-07-26 北京嘀嘀无限科技发展有限公司 硬件评测方法、装置、电子设备、存储介质和程序产品
KR102760500B1 (ko) * 2021-11-01 2025-01-24 고려대학교 산학협력단 인공 신경망 분석 장치 및 방법
US12517725B2 (en) * 2022-02-28 2026-01-06 Palantir Technologies Inc. Systems and methods of application builders for offline-capable application
CN115511060B (zh) * 2022-10-14 2026-03-20 浙江大华技术股份有限公司 模型的转换方法、装置、存储介质及电子装置
CN115827595A (zh) * 2022-11-08 2023-03-21 深圳市有方科技股份有限公司 数据管理方法、装置及计算机设备
KR102645690B1 (ko) * 2023-06-13 2024-03-11 주식회사 노타 노드에 대응되는 인공지능 기반의 모델을 제공하기 위한 방법 및 디바이스
CN119963279A (zh) * 2023-11-08 2025-05-09 北京京东乾石科技有限公司 一种分拣方法、装置、电子设备及存储介质

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101430647A (zh) * 2008-12-02 2009-05-13 北京中星微电子有限公司 一种硬件设备及其驱动安装方法
CN101937351A (zh) * 2010-09-15 2011-01-05 深圳市任子行网络技术股份有限公司 一种自动安装应用软件的方法和系统
CN106331858A (zh) * 2016-09-08 2017-01-11 北京奇虎科技有限公司 程序安装适配性的检测方法、装置及系统
CN109492241A (zh) * 2018-08-10 2019-03-19 北京中科寒武纪科技有限公司 转换方法、装置、计算机设备和存储介质

Family Cites Families (31)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7222302B2 (en) * 2003-06-05 2007-05-22 International Business Machines Corporation Method and apparatus for generating it level executable solution artifacts from the operational specification of a business
WO2006043012A1 (en) * 2004-10-22 2006-04-27 New Technology/Enterprise Limited Data processing system and method
US8538786B2 (en) * 2006-06-07 2013-09-17 International Business Machines Corporation Method, system and program product for generating an implementation of a business rule including a volatile portion
US8516435B2 (en) * 2008-06-19 2013-08-20 International Business Machines Corporation System and method for generating implementation artifacts for contextually-aware business applications
US8495559B2 (en) * 2008-09-09 2013-07-23 International Business Machines Corporation Extracting platform independent models from composite applications
US8813024B2 (en) * 2008-09-22 2014-08-19 International Business Machines Corporation System and a method for cross-platform porting of business application and making them contextually-aware on target platforms
EP2350815A1 (en) * 2008-10-21 2011-08-03 Accenture Global Services Limited Model transformation unit
KR101770292B1 (ko) * 2014-11-27 2017-08-22 주식회사 엘지씨엔에스 컴퓨터 수행 가능한 모델 역공학 방법 및 장치
CN104679511A (zh) * 2015-02-10 2015-06-03 北京系统工程研究所 基于MDE模型转换的MapReduce代码生成方法
US20160328644A1 (en) * 2015-05-08 2016-11-10 Qualcomm Incorporated Adaptive selection of artificial neural networks
US10366687B2 (en) * 2015-12-10 2019-07-30 Nuance Communications, Inc. System and methods for adapting neural network acoustic models
CN105930354B (zh) * 2016-04-08 2020-02-14 四川师范大学 存储模型转换方法和装置
CN107622427B (zh) * 2016-07-13 2021-04-06 阿里巴巴集团控股有限公司 深度学习的方法、装置及系统
JP2018018451A (ja) * 2016-07-29 2018-02-01 富士通株式会社 機械学習方法、機械学習プログラム及び情報処理装置
CN106650922B (zh) * 2016-09-29 2019-05-03 清华大学 硬件神经网络转换方法、计算装置、软硬件协作系统
US11106969B2 (en) * 2017-01-19 2021-08-31 International Business Machines Corporation Method and apparatus for driver identification leveraging telematics data
EP3786786B1 (en) * 2017-04-19 2023-06-07 Shanghai Cambricon Information Technology Co., Ltd Processing device, processing method, chip, and electronic apparatus
US20200380357A1 (en) * 2017-09-13 2020-12-03 Intel Corporation Incremental network quantization
US10469327B1 (en) * 2017-10-18 2019-11-05 Vulcan Inc. Operations environment model and simulation to reduce operator error
CN108038544B (zh) * 2017-12-04 2020-11-13 华南师范大学 基于大数据和深度学习的神经网络深度学习方法和系统
JP7299846B2 (ja) * 2017-12-29 2023-06-28 カンブリコン テクノロジーズ コーポレイション リミティド ニューラルネットワーク処理方法、コンピュータシステム及び記憶媒体
CN108053031A (zh) * 2018-01-02 2018-05-18 安徽大学 一种深度学习系统及深度学习识别方法
CN108205708A (zh) * 2018-01-02 2018-06-26 安徽大学 一种新型可扩展的深度学习系统及数据识别方法
US20190286972A1 (en) * 2018-03-14 2019-09-19 Microsoft Technology Licensing, Llc Hardware accelerated neural network subgraphs
US20210042621A1 (en) * 2018-04-17 2021-02-11 Shenzhen Corerain Technologies Co., Ltd. Method for operation of network model and related product
CN109716288A (zh) * 2018-04-17 2019-05-03 深圳鲲云信息科技有限公司 网络模型编译器及相关产品
JP7440420B2 (ja) * 2018-05-07 2024-02-28 グーグル エルエルシー 包括的機械学習サービスを提供するアプリケーション開発プラットフォームおよびソフトウェア開発キット
CN112016668A (zh) * 2019-05-31 2020-12-01 苹果公司 机器学习模型在运行时期间的可变参数
US20210117859A1 (en) * 2019-10-20 2021-04-22 Nvidia Corporation Live updating of machine learning models
US20220012575A1 (en) * 2020-07-09 2022-01-13 Femtosense, Inc. Methods and apparatus for localized processing within multicore neural networks
US20230214638A1 (en) * 2021-12-30 2023-07-06 AiM Future Inc. Apparatus for enabling the conversion and utilization of various formats of neural network models and method thereof

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN101430647A (zh) * 2008-12-02 2009-05-13 北京中星微电子有限公司 一种硬件设备及其驱动安装方法
CN101937351A (zh) * 2010-09-15 2011-01-05 深圳市任子行网络技术股份有限公司 一种自动安装应用软件的方法和系统
CN106331858A (zh) * 2016-09-08 2017-01-11 北京奇虎科技有限公司 程序安装适配性的检测方法、装置及系统
CN109492241A (zh) * 2018-08-10 2019-03-19 北京中科寒武纪科技有限公司 转换方法、装置、计算机设备和存储介质

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See also references of EP3640825A4 *

Also Published As

Publication number Publication date
CN109492241A (zh) 2019-03-19
EP3640825A4 (en) 2020-08-26
CN109492241B (zh) 2020-03-10
JP2020532778A (ja) 2020-11-12
US20220214875A1 (en) 2022-07-07
CA3062949A1 (en) 2020-02-10
US20200104129A1 (en) 2020-04-02
EP3640825B1 (en) 2025-03-05
AU2019268193B2 (en) 2022-04-21
CN111309486A (zh) 2020-06-19
JP6829327B2 (ja) 2021-02-10
AU2019268193A1 (en) 2020-02-27
CN111309486B (zh) 2024-01-12
US11314507B2 (en) 2022-04-26
EP3640825A1 (en) 2020-04-22
US11853760B2 (en) 2023-12-26
EP3640825B8 (en) 2025-04-09
KR20210033874A (ko) 2021-03-29
CA3062949C (en) 2023-01-24
EP3640825C0 (en) 2025-03-05

Similar Documents

Publication Publication Date Title
CN109492241B (zh) 转换方法、装置、计算机设备和存储介质
CN114897133B (zh) 一种通用可配置的Transformer硬件加速器及其实现方法
WO2020211205A1 (zh) 一种数据处理方法及相关产品
CN110689138A (zh) 运算方法、装置及相关产品
CN111966361B (zh) 用于确定待部署模型的方法、装置、设备及其存储介质
WO2020042739A1 (zh) 数据预处理方法、装置、计算机设备和存储介质
CN114580606B (zh) 数据处理方法、装置、计算机设备和存储介质
JP7299846B2 (ja) ニューラルネットワーク処理方法、コンピュータシステム及び記憶媒体
CN108694441B (zh) 一种网络处理器和网络运算方法
CN116185377B (zh) 计算图的优化方法、计算装置及相关产品
WO2023143080A1 (zh) 一种数据处理方法以及相关设备
CN115329923B (zh) 用于神经网络模型的编译方法和相关产品
CN112766475B (zh) 处理部件及人工智能处理器
WO2021077284A1 (zh) 神经网络运行系统和方法
CN111260070A (zh) 运算方法、装置及相关产品
CN111258641A (zh) 运算方法、装置及相关产品
CN111260046A (zh) 运算方法、装置及相关产品
WO2026020775A1 (zh) 数据处理方法、装置、系统及相关设备
CN118261206A (zh) 涉及优化神经网络模型的方法及设备
CN115543329A (zh) 对运行于人工智能芯片上的区域候选网络进行优化的编译方法及其相关产品
CN112394987A (zh) 短整形转半精度浮点指令处理装置、方法及相关产品
CN119440536A (zh) 数据处理方法、装置、芯片、设备和介质
CN112394986A (zh) 半精度浮点转浮点指令处理装置、方法及相关产品
CN112394902A (zh) 半精度浮点转浮点指令处理装置、方法及相关产品
CN112394993A (zh) 半精度浮点转短整形指令处理装置、方法及相关产品

Legal Events

Date Code Title Description
ENP Entry into the national phase

Ref document number: 2019790111

Country of ref document: EP

Effective date: 20191030

ENP Entry into the national phase

Ref document number: 2019564530

Country of ref document: JP

Kind code of ref document: A

ENP Entry into the national phase

Ref document number: 20197034745

Country of ref document: KR

Kind code of ref document: A

ENP Entry into the national phase

Ref document number: 3062949

Country of ref document: CA

Kind code of ref document: A

ENP Entry into the national phase

Ref document number: 2019268193

Country of ref document: AU

Date of ref document: 20190329

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

WWG Wipo information: grant in national office

Ref document number: 201917049014

Country of ref document: IN