WO2024198748A1 - 数据处理方法、系统、芯片及终端 - Google Patents

数据处理方法、系统、芯片及终端 Download PDF

Info

Publication number
WO2024198748A1
WO2024198748A1 PCT/CN2024/076589 CN2024076589W WO2024198748A1 WO 2024198748 A1 WO2024198748 A1 WO 2024198748A1 CN 2024076589 W CN2024076589 W CN 2024076589W WO 2024198748 A1 WO2024198748 A1 WO 2024198748A1
Authority
WO
WIPO (PCT)
Prior art keywords
task
synchronization lock
data processing
hardware
memory access
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2024/076589
Other languages
English (en)
French (fr)
Inventor
任子木
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Tencent Technology (Shenzhen) Co Ltd
Original Assignee
Tencent Technology (Shenzhen) Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Tencent Technology (Shenzhen) Co Ltd filed Critical Tencent Technology (Shenzhen) Co Ltd
Priority to EP24777542.2A priority Critical patent/EP4586100A4/en
Publication of WO2024198748A1 publication Critical patent/WO2024198748A1/zh
Priority to US19/190,147 priority patent/US20250258789A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00—Arrangements for program control, e.g. control units
    • G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/46—Multiprogramming arrangements
    • G06F9/52—Program synchronisation; Mutual exclusion, e.g. by means of semaphores
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F13/00—Interconnection of, or transfer of information or other signals between, memories, input/output devices or central processing units
    • G06F13/14—Handling requests for interconnection or transfer
    • G06F13/20—Handling requests for interconnection or transfer for access to input/output bus
    • G06F13/28—Handling requests for interconnection or transfer for access to input/output bus using burst mode transfer, e.g. direct memory access DMA, cycle steal
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00—Arrangements for program control, e.g. control units
    • G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/46—Multiprogramming arrangements
    • G06F9/50—Allocation of resources, e.g. of the central processing unit [CPU]
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00—Arrangements for program control, e.g. control units
    • G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/46—Multiprogramming arrangements
    • G06F9/52—Program synchronisation; Mutual exclusion, e.g. by means of semaphores
    • G06F9/526—Mutual exclusion algorithms
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F2213/00—Indexing scheme relating to interconnection of, or transfer of information or other signals between, memories, input/output devices or central processing units
    • G06F2213/28—DMA
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00—Arrangements for program control, e.g. control units
    • G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
    • G06F9/30003—Arrangements for executing specific machine instructions
    • G06F9/30076—Arrangements for executing specific machine instructions to perform miscellaneous control operations, e.g. NOP
    • G06F9/30087—Synchronisation or serialisation instructions
    • Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02D—CLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
    • Y02D10/00—Energy efficient computing, e.g. low power processors, power management or thermal management

Definitions

  • the embodiments of the present application relate to the field of computer technology, and in particular to a data processing method, system, chip and terminal.
  • the processor's DMA (Direct Memory Access) unit usually moves the data stored in the processor's external memory to the processor's internal storage unit.
  • the processor's VPU (Vector Process Unit) then processes the data stored in the internal storage unit. Therefore, how to control the timing relationship between the direct memory access unit moving data and the vector processing unit processing data is the key to improving the processor's processing efficiency.
  • the SPU Scalar Process Unit
  • the SPU Scalar Process Unit
  • the SPU Scalar Process Unit
  • the SPU Scalar Process Unit
  • the SPU Scalar Process Unit
  • the SPU Scalar Process Unit
  • the scalar processing unit sends tasks to the direct memory access unit and the vector processing unit at the corresponding time point, thereby realizing the control of the timing relationship between the direct memory access unit and the vector processing unit by software.
  • the scalar processing unit needs to wait for the direct memory access unit to complete the task before sending the task to the vector processing unit.
  • the scalar processing unit can only be in an idle waiting state and cannot process other tasks. Therefore, the utilization rate of the scalar processing unit is low, resulting in low efficiency of the processor in processing data.
  • the embodiments of the present application provide a data processing method, system, chip and terminal, which can improve the efficiency of a processor in processing data.
  • the technical solution is as follows:
  • a data processing method is provided, which is performed by a data processing system, wherein the data processing system includes a direct memory access unit, a vector processing unit, a scalar processing unit, and a hardware synchronization lock management unit, and the method includes:
  • the scalar processing unit sends a first task instruction to the direct memory access unit and sends a second task instruction to the vector processing unit;
  • the direct memory access unit determines, based on the first task instruction sent by the scalar processing unit, a first distributed synchronization lock indicated by a first synchronization lock identifier included in the first task instruction in the direct memory access unit, wherein the first distributed synchronization lock is used to determine whether to execute a first data processing task included in the first task instruction;
  • the direct memory access unit executes the first data processing task included in the first task instruction, and sends a task completion instruction to the hardware synchronization lock management unit, when the working state value of the first distributed synchronization lock is a preset value representing an idle state and the task type value of the first distributed synchronization lock meets the execution condition of the first data processing task;
  • the hardware synchronization lock management unit updates the task type value of the first hardware synchronization lock indicated by the first synchronization lock identifier in the hardware synchronization lock sequence based on the task completion instruction;
  • the vector processing unit executes the second data processing task included in the second task instruction sent by the scalar processing unit.
  • a data processing system comprising a direct memory access unit, a vector processing unit, a scalar processing unit, and a hardware synchronization lock management unit;
  • the scalar processing unit is used to send a first task instruction to the direct memory access unit, and send a second task instruction to the vector processing unit;
  • the direct memory access unit is used to determine, based on a first task instruction sent by the scalar processing unit, a first distributed synchronization lock indicated by a first synchronization lock identifier included in the first task instruction in the direct memory access unit. lock, the first distributed synchronization lock is used to determine whether to execute the first data processing task included in the first task instruction;
  • the direct memory access unit is further configured to execute the first data processing task included in the first task instruction and send a task completion instruction to the hardware synchronization lock management unit when the working state value of the first distributed synchronization lock is a preset value representing an idle state and the task type value of the first distributed synchronization lock meets the execution condition of the first data processing task;
  • the hardware synchronization lock management unit is used for updating the task type value of the first hardware synchronization lock indicated by the first synchronization lock identifier in the hardware synchronization lock sequence based on the task completion instruction;
  • the vector processing unit is used for executing the second data processing task included in the second task instruction sent by the scalar processing unit when the task type value of the first hardware synchronization lock is updated.
  • a chip comprising a data processing system.
  • a terminal comprising a processor and a memory, the memory storing a data processing method, and the processor being used to implement the data processing method as described in the above aspects through a data processing system.
  • FIG1 is a schematic diagram of the structure of a data processing system provided in an embodiment of the present application.
  • FIG2 is a flow chart of a data processing method provided in an embodiment of the present application.
  • FIG3 is an interactive flow chart of a data processing method provided in an embodiment of the present application.
  • FIG4 is a schematic diagram of a judgment logic provided in an embodiment of the present application.
  • FIG5 is a flow chart of a data processing method provided by an embodiment of the present application.
  • FIG6 is a schematic diagram of a distributed synchronization lock provided in an embodiment of the present application.
  • FIG7 is a schematic diagram of a first hardware synchronization lock provided in an embodiment of the present application.
  • FIG8 is a flowchart of another data processing provided by an embodiment of the present application.
  • FIG. 9 is a schematic diagram of the structure of a terminal provided in an embodiment of the present application.
  • the term "at least one” means one or more, and the term “plurality” means two or more.
  • the information including but not limited to user device information, user personal information, etc.
  • data including but not limited to data used for analysis, stored data, displayed data, etc.
  • signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
  • the first task instruction and the second task instruction involved in this application are obtained with full authorization.
  • Artificial Intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
  • artificial intelligence is a comprehensive technology of computer science that attempts to understand the nature of intelligence.
  • Artificial intelligence is the study of the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
  • Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies.
  • Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operating/interactive systems, mechatronics and other technologies.
  • Artificial intelligence software technologies mainly include computer vision technology, speech processing technology, natural language processing technology, as well as machine learning/deep learning, autonomous driving, smart transportation and other major directions.
  • FIG1 is a schematic diagram of the structure of a data processing system provided in an embodiment of the present application.
  • the data processing system 10 includes a scalar processing unit 101 , a direct memory access 102 , a vector processing unit 103 , and a hardware synchronization lock management unit 104 .
  • the data processing system 10 can be configured in any type of processor such as an AI (Artificial Intelligence) processor, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and a DSP (Digital Signal Processor).
  • the data processing system 10 can also be configured in an AI (Artificial Intelligence) chip, a SOC (System On Chip), a PCB (Printed Circuit Board) carrying the chip, or a terminal.
  • the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, and other types of devices, which are not limited in the embodiments of the present application.
  • the scalar processing unit 101 is used to send a first task instruction to the direct memory access 102, and send a second task instruction to the vector processing unit 103, wherein the first task instruction and the second task instruction include the same synchronization lock identifier. Accordingly, the direct memory access 102 and the vector processing unit 103 determine the task to be executed and the time to execute the task based on the received task instruction.
  • the direct memory access 102 determines the first distributed synchronization lock in the direct memory access 102 according to the first synchronization lock identifier in the first task instruction sent by the scalar processing unit 101. Then, the direct memory access 102 determines whether to execute the first data processing task in the first task instruction according to the working state value of the first distributed synchronization lock and the task type value of the first distributed synchronization lock.
  • the direct memory access 102 After the direct memory access 102 completes the task, it can send a task completion instruction to the hardware synchronization lock management unit 104. Accordingly, after receiving the task completion instruction, the hardware synchronization lock management unit 104 updates the task type value of the first hardware synchronization lock in the hardware synchronization lock sequence.
  • the vector processing unit 103 can execute the second data processing task in the second task instruction when the task type value of the first hardware synchronization lock is updated.
  • the number of the above terminals may be more or less.
  • the above terminal may be only one, or the above terminals may be dozens or hundreds, or more.
  • the embodiment of the present application does not limit the number of terminals and device types.
  • a data processing method is provided, which is executed by a data processing system, the data processing system including a direct memory access unit, a vector processing unit, a scalar processing unit and a hardware synchronization lock management unit, the method including: the scalar processing unit sends a first task instruction to the direct memory access unit, and sends a second task instruction to the vector processing unit, wherein the first task instruction includes a first data processing task and a first synchronization lock identifier indicating a first distributed synchronization lock, and the second task instruction includes a second data processing task; the direct memory access unit identifies a working state value of the first distributed synchronization lock based on the first synchronization lock identifier from the first task instruction; when the working state value of the first distributed synchronization lock is a preset value representing an idle state and the task type value of the first distributed synchronization lock meets the execution condition of the first data processing task, the direct content access unit executes the first data processing task and sends a task completion instruction to the
  • FIG2 is a flow chart of a data processing method provided in an embodiment of the present application, wherein the method is applied to a data processing system, wherein the data processing system includes a direct memory access unit (Direct Memory Access, direct memory access), a vector processing unit (Vector Process Unit, vector processing unit), and a scalar processing unit (Scalar Process Unit, scalar processing unit). Unit) and a hardware synchronization lock management unit.
  • the method includes:
  • a scalar processing unit sends a first task instruction to a direct memory access unit, and sends a second task instruction to a vector processing unit.
  • the data processing system can be configured in any type of processor such as an AI processor, a central processing unit, a graphics processor, and a digital signal processor.
  • the data processing system can also be configured in an AI chip, a system-level chip, a PCB carrying a chip, or a terminal, which is not limited in the embodiment of the present application.
  • the direct memory access unit in the data processing system can move data between the external memory and the memory inside the data system.
  • the embodiment of the present application is explained by taking the example of the direct memory access unit moving the data to be processed in the external memory to the inside of the data processing system.
  • the vector processing unit can process the data to be processed moved to the inside of the data processing system by the direct memory access unit.
  • the scalar processing unit is used to send task instructions to the direct memory access unit and the vector processing unit respectively. Accordingly, the direct memory access unit and the vector processing unit can determine the task to be executed, the task type of the task, the execution condition of the task and other information through the received task instructions.
  • the direct memory access unit and the vector processing unit need to perform more tasks, and there is a time-dependent relationship between the tasks performed by the direct memory access unit and the tasks performed by the vector processing unit.
  • the vector processing unit performs the task of processing data, and the execution result of the task of moving data by the direct memory access unit is required. Therefore, it is necessary to control the timing relationship between the direct memory access unit and the vector processing unit to perform tasks to ensure the efficiency of the data processing system in processing data.
  • the timing relationship between the direct memory access unit and the vector processing unit in performing tasks is controlled by a hardware synchronization lock.
  • the task instructions sent by the scalar processing unit to the direct memory access unit and the vector processing unit include the same synchronization lock identifier, that is, the first task instruction and the second task instruction include the same synchronization lock identifier.
  • the synchronization lock identifier can indicate any idle hardware synchronization lock in the hardware synchronization lock sequence.
  • the above timing relationship can be controlled by the hardware synchronization lock, that is, before the direct memory access unit executes the task of moving data, the task type value of the corresponding hardware synchronization lock is set to 0.
  • the hardware synchronization lock management unit updates the task type value of the hardware synchronization lock to 1.
  • the vector processing unit can determine whether the direct memory access unit has completed the task through the task type value of the hardware synchronization lock. When the task type value of the hardware synchronization lock is updated to 1, the vector processing unit determines that the direct memory access unit has completed the task, and the vector processing unit can execute the task of processing the target data.
  • the direct memory access unit determines, based on the first task instruction sent by the scalar processing unit, a first distributed synchronization lock indicated by a first synchronization lock identifier included in the first task instruction in the direct memory access unit.
  • the direct memory access unit includes multiple distributed synchronization locks corresponding to the hardware synchronization lock sequence. After the direct memory access unit receives the first task instruction sent by the scalar processing unit, the direct memory access unit determines the first distributed synchronization lock corresponding to the first synchronization lock identifier among the multiple distributed synchronization locks in the direct memory access unit based on the first synchronization lock identifier included in the first task instruction. Since the distributed synchronization lock is consistent with the corresponding hardware synchronization lock in the hardware synchronization lock sequence. Therefore, the direct memory access unit can directly determine whether to execute the first data processing task based on the relevant information of the first distributed synchronization lock.
  • the direct memory access unit executes the first data processing task included in the first task instruction, and sends a task completion instruction to the hardware synchronization lock management unit after the first data processing task is completed.
  • the direct memory access unit determines whether to execute the first data processing task currently according to the working state value of the current first distributed synchronization lock and the task type value of the first distributed synchronization lock.
  • the working state value of the first distributed synchronization lock includes an idle state and a non-idle state.
  • the idle state indicates that the previous task corresponding to the distributed synchronization lock has been completed, and the direct memory access unit can currently execute the first data processing task.
  • the non-idle state indicates that the previous task corresponding to the distributed synchronization lock is being executed, and the direct memory access unit cannot currently execute the first data processing task.
  • the execution condition is used to control the relationship between the task type value of the first distributed synchronization lock and the execution of the first data processing task.
  • the execution condition is used to control the timing of the direct memory access unit executing the task through the distributed synchronization lock. Therefore, when the working state value of the first distributed synchronization lock is a preset value representing the idle state, and the task type value of the first distributed synchronization lock meets the execution condition of the first data processing task, it indicates that the first data processing task can be executed at present.
  • the direct memory access unit moves the data to be processed from the external memory to the inside of the data processing system by executing the first data processing task. After the direct memory access unit completes the data movement, the direct memory access unit sends a task completion instruction to the hardware synchronization lock management unit.
  • the task completion instruction is used to notify the hardware synchronization lock management unit that the current first data processing task has been completed.
  • the hardware synchronization lock management unit updates the task type value of the first hardware synchronization lock indicated by the first synchronization lock identifier in the hardware synchronization lock sequence based on the task completion instruction.
  • the hardware synchronization lock management unit is used to manage multiple hardware synchronization locks in the hardware synchronization lock sequence. After receiving the task completion instruction, the hardware synchronization lock management unit determines the first hardware synchronization lock corresponding to the first synchronization lock identifier in the hardware synchronization lock sequence. The hardware synchronization lock management unit updates the task type value of the first hardware synchronization lock.
  • the vector processing unit executes the second data processing task included in the second task instruction sent by the scalar processing unit.
  • the vector processing unit can determine whether to execute the task of processing data according to the updated task type value of the first hardware synchronization lock.
  • the vector processing unit determines that the first data processing task has been completed, and if the execution condition of the second data processing task is currently met, the vector processing unit executes the second data processing task.
  • the second data processing task is the vector processing unit processing the data to be processed that the direct memory access unit moves to the data processing system.
  • the embodiment of the present application provides a data processing method, which is executed by a data processing system.
  • a scalar processing unit can send task instructions to a direct memory access unit and a vector processing unit respectively.
  • the scalar processing unit does not need to wait for the direct memory access unit to complete the task before sending the task to the vector processing unit, which can improve the processing efficiency of the scalar processing unit.
  • the direct memory access unit can determine the first distributed synchronization lock in the direct memory access unit according to the first synchronization lock identifier in the first task instruction sent by the scalar processing unit. Then the direct memory access unit can determine whether the first data processing task is currently executed according to the working state value of the first distributed synchronization lock and the task type value of the first distributed synchronization lock.
  • the scalar processing unit does not need to determine when the direct memory access unit executes the first data processing task, which realizes decoupling from the scalar processing unit. And after the direct memory access unit completes the task, the direct memory access unit sends a task completion instruction to the hardware synchronization lock management unit. Based on the task completion instruction, the hardware synchronization lock management unit can update the task type value of the first hardware synchronization lock in the hardware synchronization lock sequence. In response to the update of the task type value of the first hardware synchronization lock, the vector processing unit executes the second data processing task in time, which improves the efficiency of the processor in processing data.
  • FIG3 is an interactive flow chart of a data processing method provided in an embodiment of the present application, the method is applied to a data processing system, the data processing system includes a direct memory access unit (Direct Memory Access, direct memory access), a vector processing unit (Vector Process Unit, vector processing unit), a scalar processing unit (Scalar Process Unit, scalar processing unit) and a hardware synchronization lock management unit.
  • the method includes:
  • a hardware synchronization lock management unit determines at least one idle hardware synchronization lock and an identifier of each hardware synchronization lock in a hardware synchronization lock sequence.
  • the data processing system can be configured in any type of processor such as an AI processor, a central processing unit, a graphics processor, and a digital signal processor.
  • the data processing system can also be configured in an AI chip, a system-level chip, a PCB carrying a chip, or a terminal, which is not limited in the embodiment of the present application.
  • the data processing system calls the hardware synchronization lock management unit to determine an idle hardware synchronization lock in the hardware synchronization lock sequence to assist in executing the data processing task.
  • the at least one idle hardware synchronization lock is used to control the timing relationship between the direct memory access unit and the vector processing unit in executing the data processing task.
  • the hardware synchronization lock management unit sends an identifier of any hardware synchronization lock as a first synchronization lock identifier to the scalar processing unit.
  • the data processing system after the data processing system receives the data processing task, it is necessary to call the direct memory access unit and the vector processing unit to perform the data processing task.
  • the data processing task is to process the target data stored in the external memory.
  • the direct memory access unit is required to first move the target data stored in the external memory to the inside of the data processing system, and then the vector processing unit processes the target data inside the data processing system. Therefore, the data management system needs to send corresponding tasks to the direct memory access unit and the vector processing unit, and control the timing relationship between the direct memory access unit and the vector processing unit in executing tasks through an idle hardware synchronization lock.
  • the hardware synchronization lock management unit can send any idle hardware synchronization lock to the scalar processing unit. Accordingly, the hardware synchronization lock management unit first determines at least one idle hardware synchronization lock from the hardware synchronization lock sequence. Then, after determining at least one idle hardware synchronization lock, the hardware synchronization lock management unit sends the identifier of any idle hardware synchronization lock as the first synchronization lock identifier to the scalar processing unit.
  • the first synchronization lock identifier is used to indicate the hardware synchronization lock in the hardware synchronization lock sequence used to control the above-mentioned timing relationship.
  • the scalar processing unit is used to send task instructions to the direct memory access unit and the vector processing unit.
  • the scalar processing unit can send the same first synchronization lock identifier as part of the task instruction to the direct memory access unit and the vector processing unit during the process of sending the task instruction to the direct memory access unit and the vector processing unit. Thereby ensuring the timing relationship between the direct memory access unit and the vector processing unit in executing the data processing task.
  • the hardware sync lock management unit can determine the idle hardware sync lock by the first zero detection method. Accordingly, the hardware sync lock management unit detects the hardware sync lock sequence from the low bit and uses the first detected idle hardware sync lock identifier as the first sync lock identifier.
  • the scalar processing unit In response to the scalar processing unit receiving the first synchronization lock identifier, the scalar processing unit generates a first task instruction and a second task instruction.
  • the scalar processing unit can generate a first task instruction sent to the direct memory access unit and a second task instruction sent to the vector processing unit.
  • the scalar processing unit puts the first synchronization lock identifier and the first data processing task that the direct memory access unit needs to execute into the first task instruction.
  • the scalar processing unit puts the second synchronization lock identifier and the second data processing task that the vector processing unit needs to execute into the second task instruction.
  • the first data processing task is to move the target data stored in the external memory to the inside of the data processing system.
  • the second data processing task is to process the target data inside the data processing system.
  • the direct memory access unit and the vector processing unit can determine the corresponding hardware synchronization lock according to the first synchronization lock identifier in the task instruction.
  • the direct memory access unit and the vector processing unit can determine the timing of executing the corresponding data processing tasks according to the hardware synchronization lock, thereby controlling the timing of the direct memory access unit and the vector processing unit executing tasks through the hardware synchronization lock, thereby improving the efficiency of the data processing system in processing data.
  • the scalar processing unit sends a first task instruction to the direct memory access unit, and sends a second task instruction to the vector processing unit.
  • the scalar processing unit sends a first task instruction to the direct memory access unit, and sends a second task instruction to the vector processing unit.
  • the scalar processing unit may send the first task instruction and the second task instruction to the direct memory access unit and the vector processing unit respectively at the same time.
  • the scalar processing unit may also send the first task instruction to the direct memory access unit first, and then send the second task instruction to the vector processing unit.
  • the scalar processing unit may also send the second task instruction to the vector processing unit first, and then send the first task instruction to the direct memory access unit.
  • the embodiment of the present application does not limit the timing of sending the first task instruction and the second task instruction.
  • the direct memory access unit determines, based on the first task instruction sent by the scalar processing unit, a first distributed synchronization lock indicated by a first synchronization lock identifier included in the first task instruction in the direct memory access unit.
  • the direct memory access unit receives the first task instruction sent by the scalar processing unit.
  • the memory access unit determines a first synchronization lock identifier included in the first task instruction.
  • the direct memory access unit includes a plurality of distributed synchronization locks corresponding to the hardware synchronization lock sequence.
  • the direct memory access unit can determine a first distributed synchronization lock corresponding to the first synchronization lock identifier from among a plurality of distributed synchronization locks in the direct memory access unit based on the first synchronization lock identifier.
  • the distributed synchronization lock and the corresponding hardware synchronization lock in the hardware synchronization lock sequence remain consistent in most clock cycles. Therefore, the direct memory access unit can directly determine whether to execute the first data processing task included in the first task instruction based on the working state value and value of the first distributed synchronization lock.
  • the working state value of the first distributed synchronization lock includes an idle state and a non-idle state.
  • the idle state indicates that the previous task corresponding to the distributed synchronization lock has been executed, and the direct memory access unit can currently execute the first data processing task.
  • the non-idle state indicates that the previous task corresponding to the distributed synchronization lock is being executed, and the direct memory access unit cannot currently execute the first data processing task.
  • the execution condition is used to control the relationship between the task type value of the first distributed synchronization lock and the execution of the first data processing task. By configuring the execution condition for the task, it is possible to control the timing of the direct memory access unit executing the task through the distributed synchronization lock.
  • the direct memory access unit determines the working state value and value of the first distributed synchronization lock through two multiplexers. For the first multiplexer in the direct memory access unit, the first multiplexer determines the first distributed synchronization lock and the working state value of the first distributed synchronization lock corresponding to the first synchronization lock identifier in the first synchronization lock sequence based on the first synchronization lock identifier included in the first task instruction. Wherein, the first synchronization lock sequence is used to indicate the working state values of multiple distributed synchronization locks in the direct memory access unit.
  • the state includes an idle state and a non-idle state.
  • the second multiplexer determines the task type value of the first distributed synchronization lock in the second synchronization lock sequence based on the first synchronization lock identifier.
  • the second synchronization lock sequence is used to indicate the task type values of multiple distributed synchronization locks in the direct memory access unit.
  • the direct memory access unit determines the execution condition of the first data processing task based on the task type of the first data processing task included in the first task instruction. The execution condition is used to indicate the condition to be satisfied by the task type value of the first distributed synchronization lock.
  • the task type includes a read task type and a write task type.
  • the first distributed synchronization lock, the working state value of the first distributed synchronization lock and the task type value of the first distributed synchronization lock can be determined relatively quickly from the two synchronization lock sequences, thereby improving the efficiency of the direct memory access unit in executing the first data processing task.
  • the task type of the first data processing task is a read task type.
  • the direct memory access unit uses the task type value of the first distributed synchronization lock as a first value as an execution condition.
  • the first data processing task of the read task type is that the direct memory access unit reads the target data to be processed stored in the external memory and moves the target data to the memory.
  • the direct memory access unit can determine whether the execution condition of the first data processing task is currently met according to the working state value of the first distributed synchronization lock, the task type value of the first distributed synchronization lock, and the task type of the first data processing task through the schematic diagram of the judgment logic of the execution condition shown in FIG4.
  • the judgment logic of the execution condition includes an AND logic gate 401, an AND logic gate 402, and a multiplexer 403.
  • the negation symbol 404 is used to indicate that the input is inverted before the working state value (busy) of the synchronization lock or the task type value (sync) of the synchronization lock is input to the AND logic gate. It should be noted that the negation symbol 401 in FIG4 is only a schematic example.
  • the direct memory access unit can determine whether to invert the working state value of the synchronization lock or the task type value of the synchronization lock according to actual needs.
  • the multiplexer can determine whether the execution condition of the data processing task is currently met according to the output of the AND logic gate and the task type (sync_set_value), and output the judgment result (issue_en).
  • busy includes two values: 0 and 1.
  • the working state value of the synchronization lock busy is 0, indicating that the working state of the synchronization lock is non-idle, and the working state value of the synchronization lock busy is 1, indicating that the working state of the synchronization lock is idle.
  • sync_set_value also includes two values 0 and 1.
  • sync_set_value is 1, indicating that the task type is a read task
  • sync_set_value is 0, indicating that the task type is a write task.
  • the direct memory access unit can invert the task type value of the synchronization lock input to the logic gate 401, and the sync input to the logic gate 401 is 1.
  • the direct memory access unit does not invert the working state value of the synchronization lock in the logic gate 401, and the busy input to the logic gate 401 is 1. Both inputs of the logic gate 401 are 1, and the output of the logic gate 401 is 1.
  • the direct memory access unit can invert both of them in the logic gate 402, or it can not invert them, or it can only invert busy and not invert sync.
  • the output of the logic gate 402 in the above three cases is 0. Therefore, the multiplexer 403 can determine that the above data processing task can be executed currently based on the output value of the AND logic gate 401 being 1 and the output value of the AND logic gate 402 being 0, and the above judgment logic of the direct memory access unit is consistent with the execution logic of the data processing task.
  • the direct memory access unit can determine the execution condition of the first data processing task according to the task type indicated by sync_set_value, that is, determine the sync when the execution condition is met. For example, when the sync_set_value of the first data processing task is 1, it indicates that the first data processing task is a read task.
  • the direct memory access unit uses sync to be 0 as the execution condition, that is, the first value is 0. Therefore, the execution condition of the first data processing task is met only when busy is 1 and sync is 0. Accordingly, when busy is 1 and sync is 0, the multiplexer 403 outputs the judgment result that the execution condition is currently met. When busy and sync are other values, the multiplexer 403 outputs the judgment result that the execution condition is not currently met.
  • the direct memory access unit executes the first data processing task included in the first task instruction and sends a task completion instruction to the hardware synchronization lock management unit.
  • the direct memory access unit determines that the working state value of the first distributed synchronization lock is a preset value representing an idle state, and the task type value of the first distributed synchronization lock meets the execution condition of the first data processing task, it indicates that the first data processing task can be executed at present.
  • the direct memory access unit executes the first data processing task and moves the target data to be processed from the external memory to the inside of the data processing system. After the direct memory access unit completes the data movement, the direct memory access unit sends a task completion instruction to the hardware synchronization lock management unit.
  • the task completion instruction is used to notify the hardware synchronization lock management unit that the current first data processing task has been executed.
  • the direct memory access unit initiates the req_interface instruction to make an access request to the external memory to move the target data from the external memory.
  • the access request is sent and the access request is processed by the external memory, it indicates that the direct memory access unit has completed the movement of the target data. Then, the direct memory access unit sends a task completion instruction to the hardware synchronization lock management unit.
  • the hardware synchronization lock management unit determines a first hardware synchronization lock in the hardware synchronization lock sequence based on the first synchronization lock identifier in the task completion instruction.
  • the hardware synchronization lock management unit receives a task completion instruction sent by the direct memory access unit.
  • the hardware synchronization lock management unit determines a first synchronization lock identifier included in the task completion instruction.
  • the hardware synchronization lock management unit determines a first hardware synchronization lock corresponding to the first synchronization lock identifier in the hardware synchronization lock sequence.
  • the first hardware synchronization lock and the first distributed synchronization lock in the direct memory access unit module both correspond to the first synchronization lock identifier.
  • the hardware synchronization lock management unit updates the task type value of the first hardware synchronization lock after the second number of clock cycles.
  • the hardware synchronization lock management unit and the direct memory access unit are located in different positions in the data processing system, and the positions are distributed far away. Therefore, it is necessary to delay the update of the task type value of the first hardware synchronization lock through a delayed beater to meet the timing requirements of the hardware synchronization lock management unit and the direct memory access unit.
  • the number of delayed clock cycles is associated with the number of delayed beaters.
  • the hardware synchronization lock management unit can update the task type value of the first hardware synchronization lock after a second number of clock cycles through a second number of delayed beaters.
  • the second number can be a preset value, such as 2, 3 or 5, which is not limited in the embodiment of the present application.
  • the task type value of the first hardware synchronization lock is updated, so that the execution status of the task executed by the direct memory access unit can be represented by the task type value of the first hardware synchronization lock, and the timing of the task executed by the direct memory access unit can be controlled by the task type value of the first hardware synchronization lock.
  • the hardware synchronization lock management unit can determine the updated value of the first hardware synchronization lock according to the type of the first data processing task.
  • the hardware synchronization lock management unit updates the task type value of the first hardware synchronization lock from the first value to the second value after a second number of clock cycles. For example, after three clock cycles, the hardware synchronization lock management unit updates the task type value of the first hardware synchronization lock from 0 to 1.
  • the hardware synchronization lock management unit can determine the updated value of the first data processing task and the first hardware synchronization lock through sync_set_value.
  • sync_set_value is 1, the task type of the first data processing task is a read task, and the updated value of the first hardware synchronization lock is 1.
  • sync_set_value is 0, the task type of the first data processing task is a write task, and the updated value of the first hardware synchronization lock is 0.
  • the hardware sync lock management unit determines at least one idle hardware sync lock and the identifier of each hardware sync lock in the hardware sync lock sequence. Then, the hardware sync lock management unit sends the identifier of any hardware sync lock as the first sync lock identifier to the scalar processing unit. The scalar processing unit receives the first sync lock identifier, and the scalar processing unit sends the first sync lock identifier and the first data processing task as the first task instruction to the direct memory access unit.
  • the direct memory access unit determines the first distributed sync lock corresponding to the first sync lock identifier and the working state value (busy) of the first distributed sync lock in the first sync lock sequence through two multiplexers (mux), and determines the task type value (sync) of the first distributed sync lock in the second sync lock sequence. Then, the direct memory access unit determines whether to execute the first data processing task according to the working state value of the first distributed sync lock, the task type value of the first distributed sync lock, and the task type of the first data processing task.
  • the first data processing task can only be executed when the task type of the first data processing task is a read task, the working state value of the first synchronization lock is a preset value representing the idle state, that is, busy is 1, and the task type value of the first synchronization lock is 0.
  • the direct memory access unit executes the first data processing task.
  • the direct memory access unit accesses the external memory and moves the target data to be processed in the external memory to the inside of the data processing system. After the direct memory access unit completes the data movement, the direct memory access unit sends a task completion instruction to the hardware synchronization lock management unit.
  • the hardware synchronization lock management unit determines the first hardware synchronization lock in the hardware synchronization lock sequence according to the first synchronization lock identifier in the task completion instruction.
  • the hardware synchronization lock management unit updates the task type value of the first hardware synchronization lock from 0 to 1 after three clock cycles through three delayed beaters. Then, through three delayed beaters, after three clock cycles, the task type value of the first distributed synchronization lock corresponding to the first synchronization lock identifier in the second synchronization lock sequence is updated from 0 to 1.
  • the vector processing unit determines, based on the second task instruction sent by the scalar processing unit, a second distributed synchronization lock indicated by the first synchronization lock identifier included in the second task instruction.
  • a vector processing unit receives a second task instruction sent by a scalar processing unit.
  • the vector processing unit determines a first synchronization lock identifier included in the second task instruction.
  • the vector processing unit Similar to the direct memory access unit, the vector processing unit also includes a plurality of distributed synchronization locks corresponding to a hardware synchronization lock sequence.
  • the vector processing unit can determine a second distributed synchronization lock corresponding to the first synchronization lock identifier among a plurality of distributed synchronization locks in the vector processing unit based on the first synchronization lock identifier.
  • the second distributed synchronization lock in the vector processing unit, the first distributed synchronization lock in the direct memory access unit, and the first hardware synchronization lock are consistent in most clock cycles.
  • the vector processing unit can directly determine whether to execute the first data processing task included in the second task instruction according to the working state value of the second distributed synchronization lock and the task type value of the second distributed synchronization lock. Therefore, the vector processing unit does not need to determine the working state value and value of the first hardware synchronization lock in the external hardware synchronization lock sequence, thereby achieving decoupling from the external hardware synchronization lock sequence.
  • the vector processing unit executes the second data processing task and sends a task completion instruction to the hardware synchronization lock management unit after the second data processing task is completed.
  • the vector processing unit determines that the working state value of the second distributed synchronization lock is a preset value representing an idle state, and the task type value of the second distributed synchronization lock meets the execution condition of the second data processing task, it indicates that the second data processing task can be executed at present.
  • the vector processing unit executes the second data processing task and processes the target data inside the data processing system. After the vector processing unit completes processing the target data, the vector processing unit sends a task completion instruction to the hardware synchronization lock management unit.
  • the task completion instruction is used to notify the hardware synchronization lock management unit that the current first data processing task has been executed.
  • the vector processing unit starts the req_interface instruction to make an access request to the internal storage unit of the data processing system to process the target data.
  • the access request is sent and the access request is processed by the internal storage unit, it indicates that the vector processing unit has completed the processing of the target data. Then, the vector processing unit sends a task completion instruction to the hardware synchronization lock management unit.
  • the vector processing unit updates the working state value of the second distributed synchronization lock to a preset value representing the idle state, and the task type value of the second distributed synchronization lock is updated to the task type value of the first hardware synchronization lock a first number of clock cycles after the task type value of the first hardware synchronization lock is updated.
  • the vector processing unit sends a task completion instruction to the hardware synchronization lock management unit after executing and completing the second data processing task. Accordingly, after receiving the task completion instruction, the hardware synchronization lock management unit delays several clock cycles to update the task type value of the first hardware synchronization lock through the delay beater. In the case where the task type value of the first hardware synchronization lock is updated, the task type value of the second distributed synchronization lock is updated to the task type value of the first hardware synchronization lock after the first number of clock cycles through the first number of delay beaters.
  • the first number can be a preset value, such as 2, 3 or 5, which is not limited in the embodiment of the present application.
  • the vector processing unit updates the working state value of the second distributed synchronization lock to represent the idle state, so as to indicate that the current vector processing unit has completed the second data processing task. Furthermore, the vector processing unit can determine whether to execute the next task according to the new state of the second synchronization lock.
  • the vector processing unit includes a third synchronization lock sequence and a fourth synchronization lock sequence.
  • the third synchronization lock sequence is used to indicate the working state values of multiple distributed synchronization locks in the vector processing unit.
  • the fourth synchronization lock sequence is used to indicate the task type values of multiple distributed synchronization locks in the vector processing unit. Therefore, in the case where the task type value of the first hardware synchronization lock is updated, after a first number of clock cycles, the task type value of the second distributed synchronization lock in the fourth synchronization lock sequence is updated to the task type value of the first hardware synchronization lock. In the case where the task type value of the second distributed synchronization lock is updated, the vector processing unit updates the working state value of the second distributed synchronization lock in the third synchronization lock sequence to represent an idle state.
  • the task type value of the first hardware synchronization lock is first delayed to update, and then the task type value of the second distributed synchronization lock is delayed to update. Therefore, there is a delay between the completion of the task by the vector processing unit and the update of the task type value of the second distributed synchronization lock.
  • the task type value of the first hardware synchronization lock is updated from 1 to 0.
  • the time when the task type value of the second distributed synchronization lock is updated lags behind the time when the task type value of the first hardware synchronization lock is updated.
  • the task type value of the second distributed synchronization lock is still 1 and has not been updated to 0. If the vector processing unit determines that the second data processing task can be executed based on the current task type value of the second distributed synchronization lock, it starts to execute the second data processing task, which will cause a task conflict problem.
  • the above-mentioned task conflict problem can be avoided.
  • the working state value of the second distributed synchronization lock in the fourth synchronization lock sequence is a non-idle state, that is, busy is 0.
  • the working state value of the second distributed synchronization lock is updated to represent the idle state, that is, busy is updated from 0 to 1.
  • the update of the task type value of the first hardware synchronization lock is shown in Figure 7.
  • the direct memory access unit determines that the task type value of the first distributed synchronization lock is 0 and the working state value of the first distributed synchronization lock is a preset value representing the idle state
  • the direct memory access unit executes the first data processing task.
  • the direct memory access unit instructs the hardware synchronization lock management unit to update the task type value of the first hardware synchronization lock to 1.
  • the vector processing unit determines that the task type value of the second distributed synchronization lock is 1 and the working state value of the second distributed synchronization lock is a preset value representing the idle state, the vector processing unit executes the second data processing task. After the vector processing unit completes the second data processing task, it instructs the hardware synchronization lock management unit to update the task type value of the first hardware synchronization lock to 0.
  • the hardware synchronization lock management unit determines at least one idle hardware synchronization lock and the identifier of each hardware synchronization lock in the hardware synchronization lock sequence. Then, the hardware synchronization lock management unit sends the identifier of any hardware synchronization lock as the first synchronization lock identifier to the scalar processing unit.
  • the scalar processing unit receives the first synchronization lock identifier, and the scalar processing unit sends the first synchronization lock identifier and the second data processing task as the second task instruction to the vector processing unit. Based on the received second task instruction, the vector processing unit determines the working state value (busy) of the second distributed synchronization lock and the second distributed synchronization lock corresponding to the second synchronization lock identifier in the third synchronization lock sequence through two multiplexers (mux), and determines the task type value (sync) of the second distributed synchronization lock in the fourth synchronization lock sequence.
  • the vector processing unit determines whether to execute the second data processing task according to the working state value of the second distributed synchronization lock, the task type value of the second distributed synchronization lock, and the type of the second data processing task. For example, when the task type of the second data processing task is a write task, the working state value of the second synchronization lock is a preset value representing the idle state, that is, busy is 1, and the task type value of the second synchronization lock is 1, the second data processing task can be executed. When the judgment result indicates that the second data processing task can be executed, the vector processing unit executes the second data processing task. The vector processing unit processes the target data inside the data processing system.
  • the vector processing unit After the vector processing unit completes the processing of the target data, the vector processing unit sends a task completion instruction to the hardware synchronization lock management unit.
  • the hardware synchronization lock management unit determines the first hardware synchronization lock in the hardware synchronization lock sequence according to the first synchronization lock identifier in the task completion instruction.
  • the hardware synchronization lock management unit updates the task type value of the first hardware synchronization lock from 1 to 0 after three clock cycles through three delayed beaters. Then, through three delayed beaters, after three clock cycles, the task type value of the second distributed synchronization lock corresponding to the first synchronization lock identifier in the second synchronization lock sequence is updated from 1 to 0.
  • the hardware synchronization lock management unit can also merge the task completion instructions sent by the direct memory access unit and the vector processing unit, and update the task type value of the corresponding hardware synchronization lock in the hardware synchronization lock sequence according to the merged multiple task completion instructions.
  • the hardware synchronization lock management unit can also receive task completion instructions sent by other units other than the direct memory access unit and the vector processing unit. Accordingly, the hardware synchronization lock management unit merges the task completion instructions of different units and updates the task type value of the corresponding hardware synchronization lock in the hardware synchronization lock sequence, which can improve the efficiency of updating the task type value of the hardware synchronization lock.
  • the above steps 301-311 are described by taking the direct memory access unit performing a read task as an example. That is, the direct memory access unit reads the target data to be processed from the external memory and moves the target data to the inside of the data processing system.
  • the read task may also be performed by the vector processing unit. Accordingly, the vector processing unit reads the target data to be processed inside the data processing system and processes the target data. The direct memory access unit then writes the target data processed by the vector processing unit into the external memory.
  • the process of the vector processing unit executing the read task includes: the scalar processing unit sends a read request to the vector processing unit.
  • the vector processing unit sends a third task instruction to the element and sends a fourth task instruction to the direct memory access unit, the third task instruction includes a third data processing task and a third distributed synchronization lock indicated by the second synchronization lock identifier, and the fourth task instruction includes a fourth data processing task;
  • the vector processing unit executes the third data processing task when the working state value of the third distributed synchronization lock is a preset value representing an idle state and the task type value of the third distributed synchronization lock meets the execution condition of the third data processing task, and sends a task completion instruction to the hardware synchronization lock management unit after the third data processing task is completed;
  • the hardware synchronization lock management unit updates the task type value of the second hardware synchronization lock indicated by the second synchronization lock identifier in the hardware synchronization lock sequence based on the task completion instruction;
  • the direct memory access unit executes
  • the scalar processing unit sends a third task instruction to the vector processing unit, and sends a fourth task instruction to the direct memory access unit.
  • the data processing system After receiving the data processing task, the data processing system needs to call the direct memory access unit and the vector processing unit to execute the data processing task.
  • the data processing system calls the hardware synchronization lock management unit to determine at least one idle hardware synchronization lock and the identifier of each hardware synchronization lock in the hardware synchronization lock sequence.
  • the hardware synchronization lock management unit sends the identifier of any hardware synchronization lock as the second synchronization lock identifier to the scalar processing unit.
  • the scalar processing unit After the scalar processing unit receives the first synchronization lock identifier, the scalar processing unit uses the second synchronization lock identifier and the third data processing task that the vector processing unit needs to execute as the third task instruction.
  • the scalar processing unit uses the second synchronization lock identifier and the fourth data processing task that the direct memory access unit needs to execute as the second task instruction.
  • the scalar processing unit sends the third task instruction to the vector processing unit and sends the fourth task instruction to the direct memory access unit.
  • the vector processing unit determines, based on the third task instruction sent by the scalar processing unit, a third distributed synchronization lock indicated by a second synchronization lock identifier included in the third task instruction in the vector processing unit, wherein the third distributed synchronization lock is used to determine whether to execute a third data processing task included in the third task instruction.
  • the vector processing unit executes the third data processing task included in the third task instruction to process the target data inside the data system. After the vector processing unit completes the execution of the third data processing task, it sends a task completion instruction to the hardware synchronization lock management unit.
  • the task type of the third data processing task is a read task type.
  • the vector processing unit uses the task type value of the third distributed synchronization lock as the first value as the execution condition of the third data processing task.
  • the vector processing unit uses the task type value of the third distributed synchronization lock as 0 as the execution condition of the third data processing task.
  • the hardware synchronization lock management unit updates the task type value of the second hardware synchronization lock indicated by the second synchronization lock identifier in the hardware synchronization lock sequence based on the task completion instruction.
  • the task type of the third data processing task is a read task type.
  • the hardware synchronization lock management unit determines the second hardware synchronization lock in the hardware synchronization lock sequence based on the second synchronization lock identifier in the task completion instruction. After a third number of clock cycles, the hardware synchronization lock management unit updates the task type value of the second hardware synchronization lock from the first value to the second value. For example, after three clock cycles, the hardware synchronization lock management unit updates the task type value of the second hardware synchronization lock from 0 to 1. By updating the second hardware synchronization lock to a different value through the task type of the data processing task, it is possible to determine whether the execution conditions of tasks of different task types are met according to the task type value of the second hardware synchronization lock. This realizes the control of the execution timing of tasks according to the hardware synchronization lock.
  • the direct memory access unit executes the fourth data processing task included in the fourth task instruction sent by the scalar processing unit.
  • the direct memory access unit writes the target data processed by the vector processing unit into the external memory.
  • the embodiment of the present application provides a data processing method, which is executed by a data processing system.
  • a scalar processing unit can send task instructions to a direct memory access unit and a vector processing unit respectively.
  • the scalar processing unit does not need to wait for the direct memory access unit to complete the task before sending the task to the vector processing unit, which can improve the processing efficiency of the scalar processing unit.
  • the direct memory access unit can determine the first distributed synchronization lock in the direct memory access unit according to the first synchronization lock identifier in the first task instruction sent by the scalar processing unit.
  • the direct memory access unit can determine whether the first data processing task is currently being executed according to the working state value of the first distributed synchronization lock and the task type value of the first distributed synchronization lock.
  • the scalar processing unit does not need to determine when the direct memory access unit executes the first data processing task, thereby achieving decoupling from the scalar processing unit. And after the direct memory access unit completes the task, the direct memory access unit sends a task completion instruction to the hardware synchronization lock management unit. Based on the task completion instruction, the hardware synchronization lock management unit can update the task type value of the first hardware synchronization lock in the hardware synchronization lock sequence. In response to the update of the task type value of the first hardware synchronization lock, the vector processing unit executes the second data processing task in a timely manner, thereby improving the efficiency of the processor in processing data.
  • the embodiment of the present application further provides a data processing system.
  • the data processing system includes: a scalar processing unit 801 , a direct memory access unit 802 , a hardware sync lock management unit 803 , and a vector processing unit 804 .
  • the scalar processing unit 801 is used to send a first task instruction to the direct memory access unit 802 and send a second task instruction to the vector processing unit 804;
  • the direct memory access unit 802 is used to determine, based on the first task instruction sent by the scalar processing unit 801, a first distributed synchronization lock indicated by a first synchronization lock identifier included in the first task instruction in the direct memory access unit 802, wherein the first distributed synchronization lock is used to determine whether to execute a first data processing task included in the first task instruction;
  • the direct memory access unit 802 is further configured to execute the first data processing task included in the first task instruction when the working state value of the first distributed synchronization lock is a preset value representing an idle state and the task type value of the first distributed synchronization lock meets the execution condition of the first data processing task, and send a task completion instruction to the hardware synchronization lock management unit 803;
  • the hardware synchronization lock management unit 803 is used to update the task type value of the first hardware synchronization lock indicated by the first synchronization lock identifier in the hardware synchronization lock sequence 8031 based on the task completion instruction;
  • the vector processing unit 804 is configured to execute the second data processing task included in the second task instruction sent by the scalar processing unit 801 when the task type value of the first hardware synchronization lock is updated.
  • the vector processing unit 804 is used to determine, based on the second task instruction sent by the scalar processing unit 801, a second distributed synchronization lock indicated by the first synchronization lock identifier included in the second task instruction in the vector processing unit 804, and the second distributed synchronization lock is used to determine whether to execute the second data processing task included in the second task instruction; when the task type value of the second distributed synchronization lock is updated, the working state value of the second distributed synchronization lock is updated to a preset value representing an idle state, and the task type value of the second distributed synchronization lock is updated to the task type value of the first hardware synchronization lock after the task type value of the first hardware synchronization lock is updated for a first number of clock cycles; when the working state value of the second distributed synchronization lock is the preset value representing the idle state, and the task type value of the second distributed synchronization lock meets the execution condition of the second data processing task, the second data processing task included in the second task instruction is executed.
  • the hardware synchronization lock management unit 803 includes a hardware synchronization lock sequence 8031; the hardware synchronization lock management unit 803 is used to determine at least one idle hardware synchronization lock and the identifier of each hardware synchronization lock in the hardware synchronization lock sequence 8031; the identifier of any hardware synchronization lock is sent as a first synchronization lock identifier to the scalar processing unit 801; the scalar processing unit 801 is used to generate a first task instruction and a second task instruction in response to receiving the first synchronization lock identifier.
  • the direct memory access unit 802 includes a first multiplexer 8021, a second multiplexer 8022, a first synchronization lock sequence 8023, and a second synchronization lock sequence 8024; the first multiplexer 8021 is used to determine the working state values of the first distributed synchronization lock and the first distributed synchronization lock in the first synchronization lock sequence 8023 based on the first synchronization lock identifier included in the first task instruction, and the first synchronization lock sequence 8023 is used to indicate the working state values of multiple distributed synchronization locks in the direct memory access unit 802, and the state includes an idle state and a non-idle state, and the idle state is used to indicate that the direct memory access unit 802 is currently capable of executing The first data processing task, the non-idle state is used to indicate that the direct memory access unit 802 is currently executing the data processing task corresponding to the first distributed synchronization lock; the second multiplexer 8022 is used to determine the task type value of the first distributed synchronization lock in the
  • the task type of the first data processing task is a read task type
  • the direct memory access unit 802 is used to set the task type value of the first distributed synchronization lock to a first value as an execution condition.
  • the hardware sync lock management unit 803 is used to determine the first hardware sync lock in the hardware sync lock sequence 8031 based on the first sync lock identifier in the task completion instruction; and update the task type value of the first hardware sync lock after a second number of clock cycles.
  • the hardware sync lock management unit 803 is configured to update the task type value of the first hardware sync lock from the first value to the second value after a second number of clock cycles.
  • the scalar processing unit 801 is used to send a third task instruction to the vector processing unit 803 and send a fourth task instruction to the direct memory access unit 802;
  • the vector processing unit 804 is used to determine, based on the third task instruction sent by the scalar processing unit 801, a third distributed synchronization lock indicated by a second synchronization lock identifier included in the third task instruction in the vector processing unit 804, and the third distributed synchronization lock is used to determine whether to execute a third data processing task included in the third task instruction;
  • the vector processing unit 804 is also used to execute the third data processing task included in the third task instruction when the working state value of the third distributed synchronization lock is a preset value representing an idle state and the task type value of the third distributed synchronization lock meets the execution condition of the third data processing task, and send a task completion instruction to the hardware synchronization lock management unit 803;
  • the hardware synchronization lock management unit 803 is used to update the task type value of the second hardware synchronization lock indicated by the
  • the task type of the third data processing task is a read task type
  • the vector processing unit 804 is used to set the task type value of the third distributed synchronization lock to the first value as the execution condition of the third data processing task.
  • the task type of the third data processing task is a read task type
  • the hardware synchronization lock management unit 803 is used to determine the second hardware synchronization lock in the hardware synchronization lock sequence 8031 based on the second synchronization lock identifier in the task completion instruction; after a third number of clock cycles, the task type value of the second hardware synchronization lock is updated from the first value to the second value.
  • the data processing system provided in the above embodiment is only illustrated by the division of the above functional modules.
  • the above functions can be assigned to different functional modules as needed, that is, the internal structure of the terminal is divided into different functional modules to complete all or part of the functions described above.
  • the data processing system and data processing method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
  • An embodiment of the present application further provides a terminal, which includes a processor and a memory, wherein a data processing method is stored in the memory, and the processor is used to implement the data processing method provided in the above embodiment through a data processing system.
  • FIG. 9 is a schematic diagram of the structure of a terminal provided in an embodiment of the present application.
  • the terminal 900 includes a processor 901 and a memory 902 .
  • the processor 901 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc.
  • the processor 901 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array).
  • the processor 901 may also include a main processor and a coprocessor.
  • the main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state.
  • the processor 901 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen.
  • the processor 901 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
  • AI Artificial Intelligence
  • the memory 902 may include one or more computer-readable storage media, which may be non-transitory.
  • the memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more Multiple disk storage devices, flash memory storage devices.
  • the non-transitory computer-readable storage medium in the memory 902 is used to store at least one computer program, which is used by the processor 901 to implement the data processing method provided in the method embodiment of the present application.
  • FIG. 9 does not limit the terminal 900 , and may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
  • An embodiment of the present application further provides a computer-readable storage medium, in which at least one computer program is stored.
  • the at least one computer program is loaded and executed by a processor to implement the data processing method provided in the above embodiment.
  • An embodiment of the present application also provides a computer program product, including a computer program, which is loaded and executed by a processor to implement the data processing method provided in the above embodiment.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Software Systems (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Multi Processors (AREA)

Abstract

一种数据处理方法,由数据处理系统执行,数据处理系统包括直接内存访问单元、向量处理单元、标量处理单元以及硬件同步锁管理单元,方法包括:标量处理单元向直接内存访问单元发送第一任务指令,并向向量处理单元发送第二任务指令(201);直接内存访问单元基于标量处理单元发送的第一任务指令,在直接内存访问单元中确定第一任务指令包括的第一同步锁标识所指示的第一分布式同步锁(202);直接内存访问单元在第一分布式同步锁的工作状态值为表征空闲状态的预设值,且第一分布式同步锁的任务类型值满足第一数据处理任务的执行条件的情况下,执行第一任务指令包括的第一数据处理任务,并在第一数据处理任务完成后向硬件同步锁管理单元发送任务完成指令(203);硬件同步锁管理单元基于任务完成指令,更新硬件同步锁序列中第一同步锁标识所指示的第一硬件同步锁的任务类型值(204);及向量处理单元在第一硬件同步锁的任务类型值发生更新的情况下,执行标量处理单元发送的第二任务指令包括的第二数据处理任务(205)。

Description

数据处理方法、系统、芯片及终端
相关申请
本申请要求2023年3月30日申请的,申请号为2023103614485,名称为“数据处理方法、系统、芯片及终端”的中国专利申请的优先权,在此将其全文引入作为参考。
技术领域
本申请实施例涉及计算机技术领域,特别涉及一种数据处理方法、系统、芯片及终端。
背景技术
随着计算机技术的发展,处理器需要处理的数据量越来越大。处理器处理数据的过程中,通常由处理器的DMA(Direct Memory Access,直接内存访问)单元将处理器外部的存储器中存储的数据搬移至处理器的内部存储单元。再由处理器的VPU(Vector Process Unit,向量处理单元)对内部存储单元中存储的数据进行处理。因此,如何控制直接内存访问单元搬移数据和向量处理单元处理数据之间的时序关系,是提升处理器的处理效率的关键。
相关技术中,通常是由处理器中用于标量计算和程序流程控制的SPU(Scalar Process Unit,标量处理单元)根据开发人与编写的软件程序,确定直接内存访问单元搬移数据的时间点和向量处理单元处理数据的时间点。标量处理单元在对应的时间点发送任务给直接内存访问单元和向量处理单元,从而实现了通过软件的方式控制直接内存访问单元和向量处理单元之间的时序关系。
但是,相关技术中标量处理单元需要等待直接内存访问单元执行任务完成,才能发送任务给向量处理单元。在标量处理单元等待期间,标量处理单元只能处于空等状态,无法处理其他的任务。因此,标量处理单元的利用率较低,导致处理器处理数据的效率低下。
发明内容
本申请实施例提供了一种数据处理方法、系统、芯片及终端,能够提高处理器处理数据的效率。所述技术方案如下:
一方面,提供了一种数据处理方法,由数据处理系统执行,所述数据处理系统包括直接内存访问单元、向量处理单元、标量处理单元以及硬件同步锁管理单元,所述方法包括:
所述标量处理单元向所述直接内存访问单元发送第一任务指令,并向所述向量处理单元发送第二任务指令;
所述直接内存访问单元基于所述标量处理单元发送的第一任务指令,在所述直接内存访问单元中确定所述第一任务指令包括的第一同步锁标识所指示的第一分布式同步锁,所述第一分布式同步锁用于确定是否执行所述第一任务指令包括的第一数据处理任务;
所述直接内存访问单元在所述第一分布式同步锁的工作状态值为表征空闲状态的预设值,且所述第一分布式同步锁的任务类型值满足所述第一数据处理任务的执行条件的情况下,执行所述第一任务指令包括的第一数据处理任务,向所述硬件同步锁管理单元发送任务完成指令;
所述硬件同步锁管理单元基于所述任务完成指令,更新硬件同步锁序列中所述第一同步锁标识所指示的第一硬件同步锁的任务类型值;及
所述向量处理单元在所述第一硬件同步锁的任务类型值发生更新的情况下,执行所述标量处理单元发送的第二任务指令包括的第二数据处理任务。
另一方面,提供了一种数据处理系统,所述数据处理系统包括直接内存访问单元、向量处理单元、标量处理单元以及硬件同步锁管理单元;
所述标量处理单元用于向所述直接内存访问单元发送第一任务指令,并向所述向量处理单元发送第二任务指令;
所述直接内存访问单元用于基于所述标量处理单元发送的第一任务指令,在所述直接内存访问单元中确定所述第一任务指令包括的第一同步锁标识所指示的第一分布式同步 锁,所述第一分布式同步锁用于确定是否执行所述第一任务指令包括的第一数据处理任务;
所述直接内存访问单元还用于在所述第一分布式同步锁的工作状态值为表征空闲状态的预设值,且所述第一分布式同步锁的任务类型值满足所述第一数据处理任务的执行条件的情况下,执行所述第一任务指令包括的第一数据处理任务,向所述硬件同步锁管理单元发送任务完成指令;
所述硬件同步锁管理单元用于基于所述任务完成指令,更新硬件同步锁序列中所述第一同步锁标识所指示的第一硬件同步锁的任务类型值;
所述向量处理单元用于在所述第一硬件同步锁的任务类型值发生更新的情况下,执行所述标量处理单元发送的第二任务指令包括的第二数据处理任务。
另一方面,提供了一种芯片,所述芯片包括数据处理系统。
另一方面,提供了一种终端,所述终端包括处理器和存储器,所述存储器中存储有数据处理方法,所述处理器用于通过数据处理系统实现如上述方面所述的数据处理方法。
本申请的一个或多个实施例的细节在下面的附图和描述中提出。本申请的其它特征和优点将从说明书、附图以及权利要求书变得明显。
附图说明
为了更清楚地说明本申请实施例或传统技术中的技术方案,下面将对实施例或传统技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本申请的实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据公开的附图获得其他的附图。
图1是本申请实施例提供的一种数据处理系统的结构示意图;
图2是本申请实施例提供的一种数据处理方法的流程图;
图3是本申请实施例提供的一种数据处理方法的交互流程图;
图4是本申请实施例提供的一种判断逻辑的示意图;
图5是本申请实施例提供的一种数据处理的流程图;
图6是本申请实施例提供的一种分布式同步锁的示意图;
图7是本申请实施例提供的一种第一硬件同步锁的示意图;
图8是本申请实施例提供的另一种数据处理的流程图;
图9是本申请实施例提供的一种终端的结构示意图。
具体实施方式
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。
本申请中术语“第一”“第二”等字样用于对作用和功能基本相同的相同项或相似项进行区分,应理解,“第一”、“第二”、“第n”之间不具有逻辑或时序上的依赖关系,也不对数量和执行顺序进行限定。
本申请中术语“至少一个”是指一个或多个,“多个”的含义是指两个或两个以上。
需要说明的是,本申请所涉及的信息(包括但不限于用户设备信息、用户个人信息等)、数据(包括但不限于用于分析的数据、存储的数据、展示的数据等)以及信号,均为经用户授权或者经过各方充分授权的,且相关数据的收集、使用和处理需要遵守相关国家和地区的相关法律法规和标准。例如,本申请中涉及到的第一任务指令和第二任务指令是在充分授权的情况下获取的。
为了便于理解,以下,对本申请涉及的术语进行解释。
人工智能(Artificial Intelligence,AI)是利用数字计算机或者数字计算机控制的机器模拟、延伸和扩展人的智能,感知环境、获取知识并使用知识获得最佳结果的理论、方法、技术及应用系统。换句话说,人工智能是计算机科学的一个综合技术,它企图了解智能的 实质,并生产出一种新的能以人类智能相似的方式做出反应的智能机器。人工智能也就是研究各种智能机器的设计原理与实现方法,使机器具有感知、推理与决策的功能。
人工智能技术是一门综合学科,涉及领域广泛,既有硬件层面的技术也有软件层面的技术。人工智能基础技术一般包括如传感器、专用人工智能芯片、云计算、分布式存储、大数据处理技术、操作/交互系统、机电一体化等技术。人工智能软件技术主要包括计算机视觉技术、语音处理技术、自然语言处理技术以及机器学习/深度学习、自动驾驶、智慧交通等几大方向。
图1是本申请实施例提供的一种数据处理系统的结构示意图,参见图1,数据处理系统10包括标量处理单元101、直接内存访问102、向量处理单元103以及硬件同步锁管理单元104。
可选地,数据处理系统10可以配置于AI(Artificial Intelligence,人工智能)处理器、CPU(Central Processing Unit,中央处理器)、GPU(Graphics Processing Unit,图形处理器)以及DSP(Digital Signal Processor,数字信号处理器)等任意类型的处理器中。数据处理系统10也可以配置于AI(Artificial Intelligence,人工智能)芯片、SOC(System On Chip,系统级芯片)、承载芯片的PCB(Printed Circuit Board,印制电路板)或者终端中。终端可以为智能手机、平板电脑、笔记本电脑、台式计算机、智能音箱、智能手表、智能语音交互设备、智能家电、车载终端等多种类型的设备,本申请实施例对此不进行限制。
标量处理单元101用于为直接内存访问102发送第一任务指令,并向向量处理单元103发送第二任务指令,第一任务指令和第二任务指令包括相同的同步锁标识。相应地,直接内存访问102和向量处理单元103基于接收到的任务指令,确定需要执行的任务和执行任务的时间。
例如,直接内存访问102根据标量处理单元101发送的第一任务指令中的第一同步锁标识,在直接内存访问102中确定第一分布式同步锁。进而直接内存访问102根据第一分布式同步锁的工作状态值和第一分布式同步锁的任务类型值,确定当前是否执行第一任务指令中第一数据处理任务。
直接内存访问102还能够执行任务完成后,向硬件同步锁管理单元104发送任务完成指令。相应地,硬件同步锁管理单元104收到的任务完成指令之后,更新硬件同步锁序列中第一硬件同步锁的任务类型值。向量处理单元103能够在第一硬件同步锁的任务类型值发生更新的情况下,执行第二任务指令中的第二数据处理任务。
本领域技术人员可以知晓,上述终端的数量可以更多或更少。比如上述终端可以仅为一个,或者上述终端为几十个或几百个,或者更多数量。本申请实施例对终端的数量和设备类型不加以限定。
在一些实施例中,提供了一种数据处理方法,由数据处理系统执行,数据处理系统包括直接内存访问单元、向量处理单元、标量处理单元以及硬件同步锁管理单元,方法包括:标量处理单元向直接内存访问单元发送第一任务指令,并向向量处理单元发送第二任务指令,其中,第一任务指令包括第一数据处理任务和指示第一分布式同步锁的第一同步锁标识,第二任务指令包括第二数据处理任务;直接内存访问单元从第一任务指令中基于第一同步锁标识识别第一分布式同步锁的工作状态值;在第一分布式同步锁的工作状态状值为表征空闲状态的预设值,且第一分布式同步锁的任务类型值满足第一数据处理任务的执行条件的情况下,直接内容访问单元执行第一数据处理任务,并在第一数据处理任务完成后向硬件同步锁管理单元发送任务完成指令;硬件同步锁管理单元基于任务完成指令,更新硬件同步锁序列中第一同步锁标识所指示的第一硬件同步锁的任务类型值;及向量处理单元在第一硬件同步锁的任务类型值发生更新的情况下,执行第二数据处理任务。
图2是本申请实施例提供的一种数据处理方法的流程图,该方法应用于数据处理系统,数据处理系统包括直接内存访问单元(Direct Memory Access,直接内存访问)、向量处理单元(Vector Process Unit,向量处理单元)、标量处理单元(Scalar Process Unit,标量处理 单元)以及硬件同步锁管理单元。参见图2,该方法包括:
201、标量处理单元向直接内存访问单元发送第一任务指令,并向向量处理单元发送第二任务指令。
在本申请实施例中,数据处理系统可以配置于AI处理器、中央处理器、图形处理器以及数字信号处理器等任意类型的处理器中。数据处理系统也可以配置于AI芯片、系统级芯片、承载芯片的PCB或者终端中,本申请实施例对此不进行限制。
数据处理系统中的直接内存访问单元能够在外部存储器与数据系统内部的存储器之间进行数据的搬移。本申请实施例以直接内存访问单元将外部存储器中待处理的数据搬移至数据处理系统内部为例进行说明。向量处理单元能够对直接内存访问单元搬移至数据处理系统内部的待处理数据进行处理。标量处理单元用于向直接内存访问单元和向量处理单元分别发送任务指令。相应地,直接内存访问单元和向量处理单元能够通过接收到的任务指令,确定需要执行的任务、任务的任务类型、任务的执行条件等信息。
由于直接内存访问单元和向量处理单元需要执行的任务较多,且直接内存访问单元执行的任务和向量处理单元执行的任务存在时序上的依赖关系。例如,向量处理单元执行处理数据的任务,需要直接内存访问单元执行搬移数据的任务的执行结果。因此,需要控制直接内存访问单元和向量处理单元执行任务的时序关系,以保证数据处理系统处理数据的效率。在本申请实施例中,通过硬件同步锁来控制直接内存访问单元和向量处理单元执行任务的时序关系。相应地,标量处理单元为直接内存访问单元和向量处理单元发送的任务指令中包括相同的同步锁标识,即第一任务指令和第二任务指令包括相同的同步锁标识。同步锁标识能够指示硬件同步锁序列中任一空闲的硬件同步锁。通过为直接内存访问单元和向量处理单元指定同一个硬件同步锁,就能够通过硬件同步锁控制直接内存访问单元和向量处理单元执行任务的时序关系。
例如,通过硬件同步锁控制上述时序关系可以为,在直接内存访问单元执行搬移数据的任务之前,将对应的硬件同步锁的任务类型值设为0。直接内存访问单元执行任务完成之后,硬件同步锁管理单元将硬件同步锁的任务类型值更新为1。向量处理单元能够通过硬件同步锁的任务类型值,确定直接内存访问单元是否执行完成任务。在硬件同步锁的任务类型值更新为1时,向量处理单元确定直接内存访问单元执行任务完成,向量处理单元就能够执行处理目标数据的任务。
202、直接内存访问单元基于标量处理单元发送的第一任务指令,在直接内存访问单元中确定第一任务指令包括的第一同步锁标识所指示的第一分布式同步锁。
在本申请实施例中,直接内存访问单元中包括与硬件同步锁序列对应的多个分布式同步锁。直接内存访问单元在接收到标量处理单元发送的第一任务指令之后,直接内存访问单元基于第一任务指令中包括的第一同步锁标识,在直接内存访问单元中的多个分布式同步锁中确定与第一同步锁标识对应的第一分布式同步锁。由于,分布式同步锁和硬件同步锁序列中对应的硬件同步锁保持一致。因此,直接内存访问单元能够直接根据第一分布式同步锁的相关信息,确定是否执行第一数据处理任务。
203、直接内存访问单元在第一分布式同步锁的工作状态值为表征空闲状态的预设值,且第一分布式同步锁的任务类型值满足第一数据处理任务的执行条件的情况下,执行第一任务指令包括的第一数据处理任务,并在第一数据处理任务完成后向硬件同步锁管理单元发送任务完成指令。
在本申请实施例中,直接内存访问单元根据当前第一分布式同步锁的工作状态值和第一分布式同步锁的任务类型值,确定当前是否执行第一数据处理任务。第一分布式同步锁的工作状态值包括空闲状态和非空闲状态。空闲状态表示分布式同步锁对应的上一个任务已经执行完成,直接内存访问单元当前能够执行第一数据处理任务。非空闲状态表示分布式同步锁对应的上一个任务正在执行,直接内存访问单元当前不能执行第一数据处理任务。
通过第一分布式同步锁的任务类型值,能够确定当前是否满足第一数据处理任务的执 行条件。执行条件用于控制第一分布式同步锁的任务类型值与执行第一数据处理任务之间的关系。通过为任务配置执行条件,能够实现通过分布式同步锁控制直接内存访问单元执行任务的时序。因此,在第一分布式同步锁的工作状态值为表征空闲状态的预设值,且第一分布式同步锁的任务类型值满足第一数据处理任务的执行条件的情况下,表明当前能够执行第一数据处理任务。直接内存访问单元通过执行第一数据处理任务,实现从外部存储器搬移待处理的数据至数据处理系统内部。直接内存访问单元搬移数据完成之后,直接内存访问单元向硬件同步锁管理单元发送任务完成指令。任务完成指令用于通知硬件同步锁管理单元当前第一数据处理任务已经执行完成。
204、硬件同步锁管理单元基于任务完成指令,更新硬件同步锁序列中第一同步锁标识所指示的第一硬件同步锁的任务类型值。
在本申请实施例中,硬件同步锁管理单元用于管理硬件同步锁序列中的多个硬件同步锁。硬件同步锁管理单元接收到任务完成指令之后,在硬件同步锁序列中确定上述第一同步锁标识对应的第一硬件同步锁。硬件同步锁管理单元更新第一硬件同步锁的任务类型值。
205、向量处理单元在第一硬件同步锁的任务类型值发生更新的情况下,执行标量处理单元发送的第二任务指令包括的第二数据处理任务。
在本申请实施例中,向量处理单元能够根据更新后的第一硬件同步锁的任务类型值,确定是否执行处理数据的任务。在第一硬件同步锁的任务类型值发生更新的情况下,向量处理单元确定第一数据处理任务已执行完毕,若当前满足第二数据处理任务的执行条件,则向量处理单元执行第二数据处理任务。其中,第二数据处理任务即为向量处理单元处理直接内存访问单元搬移至数据处理系统内部的待处理数据。
本申请实施例提供了一种数据处理方法,由数据处理系统执行。标量处理单元能够向直接内存访问单元和向量处理单元分别发送任务指令。而无需标量处理单元等待直接内存访问单元执行任务完成后再向向量处理单元发送任务,能够提高标量处理单元的处理效率。相应地,直接内存访问单元能够根据标量处理单元发送的第一任务指令中的第一同步锁标识,在直接内存访问单元中确定第一分布式同步锁。进而直接内存访问单元能够根据第一分布式同步锁的工作状态值和第一分布式同步锁的任务类型值,确定当前是否执行第一数据处理任务。而无需标量处理单元确定直接内存访问单元何时执行第一数据处理任务,实现了与标量处理单元的解耦。并且在直接内存访问单元完成任务之后直接内存访问单元向硬件同步锁管理单元发送任务完成指令。硬件同步锁管理单元基于任务完成指令,能够更新硬件同步锁序列中第一硬件同步锁的任务类型值。响应于第一硬件同步锁的任务类型值发生更新,向量处理单元及时执行第二数据处理任务,提高了处理器处理数据的效率。
图3是本申请实施例提供的一种数据处理方法的交互流程图,该方法应用于数据处理系统,数据处理系统包括直接内存访问单元(Direct Memory Access,直接内存访问)、向量处理单元(Vector Process Unit,向量处理单元)、标量处理单元(Scalar Process Unit,标量处理单元)以及硬件同步锁管理单元。参见图3,该方法包括:
301、硬件同步锁管理单元在硬件同步锁序列中确定至少一个空闲的硬件同步锁和每个硬件同步锁的标识。
在本申请实施例中,数据处理系统可以配置于AI处理器、中央处理器、图形处理器以及数字信号处理器等任意类型的处理器中。数据处理系统也可以配置于AI芯片、系统级芯片、承载芯片的PCB或者终端中,本申请实施例对此不进行限制。
响应于数据处理系统接收到数据处理任务,数据处理系统调用硬件同步锁管理单元,在硬件同步锁序列中确定一个空闲的硬件同步锁来辅助执行该数据处理任务。上述至少一个空闲的硬件同步锁用于控制直接内存访问单元和向量处理单元执行上述数据处理任务的时序关系。
302、硬件同步锁管理单元将任一硬件同步锁的标识作为第一同步锁标识发送至标量处理单元。
在本申请实施例中,数据处理系统接收到数据处理任务之后,需要调用直接内存访问单元和向量处理单元执行数据处理任务。例如,数据处理任务为处理外部存储器中存储的目标数据。在执行上述数据处理任务的过程中,需要直接内存访问单元先将外部存储器存储的目标数据搬移至数据处理系统内部,再由向量处理单元对数据处理系统内部的目标数据进行处理。因此,数据管理系统需要向直接内存访问单元和向量处理单元发送相应的任务,并通过空闲的硬件同步锁来控制直接内存访问单元和向量处理单元执行任务的时序关系。
在一些实施例中,硬件同步锁管理单元可以将任一空闲的硬件同步锁,发送至标量处理单元。相应的,硬件同步锁管理单元先从硬件同步锁序列中确定至少一个空闲硬件同步锁。然后,硬件同步锁管理单元在确定至少一个空闲硬件同步锁之后,将任一空闲硬件同步锁的标识作为第一同步锁标识发送至标量处理单元。第一同步锁标识用于指示硬件同步锁序列中用于控制上述时序关系的硬件同步锁。标量处理单元用于向直接内存访问单元和向量处理单元发送任务指令。通过向标量处理单元发送第一同步锁标识,能够使得标量处理单元在向直接内存访问单元和向量处理单元发送任务指令的过程中,将相同的第一同步锁标识作为任务指令的一部分发送至直接内存访问单元和向量处理单元。进而保证直接内存访问单元和向量处理单元执行数据处理任务的时序关系。
在一些实施例中,硬件同步锁管理单元能够通过首位零检测方法,来确定空闲的硬件同步锁。相应的,硬件同步锁管理单元从低位开始检测硬件同步锁序列,将第一个检测到的空闲硬件同步锁的标识作为第一同步锁标识。
303、响应于标量处理单元接收到第一同步锁标识,标量处理单元生成第一任务指令和第二任务指令。
在本申请实施例中,标量处理单元接收到第一同步锁标识之后,标量处理单元能够生成为直接内存访问单元发送的第一任务指令和为向量处理单元发送的第二任务指令。在生成第一任务指令的过程中,标量处理单元将第一同步锁标识和直接内存访问单元需要执行的第一数据处理任务放入第一任务指令中。在生成第二任务指令的过程中,标量处理单元将第二同步锁标识和向量处理单元需要执行的第二数据处理任务放入第二任务指令中。其中,第一数据处理任务为将外部存储器存储的目标数据搬移至数据处理系统内部。第二数据处理任务为对数据处理系统内部的目标数据进行处理。通过将第一同步锁标识和第一数据处理任务作为第一任务指令发送至直接内存访问单元,将第一同步锁标识和第二数据处理任务作为第二任务指令发送至向量处理单元,能够使得直接内存访问单元和向量处理单元根据任务指令中的第一同步锁标识确定对应的硬件同步锁。进而直接内存访问单元和向量处理单元能够根据硬件同步锁确定执行相应的数据处理任务的时机,实现了通过硬件同步锁控制直接内存访问单元和向量处理单元执行任务的时序,从而提高了数据处理系统处理数据的效率。
304、标量处理单元向直接内存访问单元发送第一任务指令,并向向量处理单元发送第二任务指令。
在本申请实施例中,标量处理单元向直接内存访问单元发送第一任务指令,向向量处理单元发送第二任务指令。
可选地,标量处理单元可以同时向直接内存访问单元和向量处理单元分别发送第一任务指令和第二任务指令。标量处理单元也可以先向直接内存访问单元发送第一任务指令,再向向量处理单元发送第二任务指令。标量处理单元也可以先向向量处理单元发送第二任务指令,再向直接内存访问单元发送第一任务指令。本申请实施例对第一任务指令和第二任务指令的发送时机不进行限制。
305、直接内存访问单元基于标量处理单元发送的第一任务指令,在直接内存访问单元中确定第一任务指令包括的第一同步锁标识所指示的第一分布式同步锁。
在本申请实施例中,直接内存访问单元接收标量处理单元发送的第一任务指令。直接 内存访问单元确定第一任务指令中包括的第一同步锁标识。直接内存访问单元中包括与硬件同步锁序列对应的多个分布式同步锁。直接内存访问单元能够基于第一同步锁标识,在直接内存访问单元中的多个分布式同步锁中确定与第一同步锁标识对应的第一分布式同步锁。并且,分布式同步锁和硬件同步锁序列中对应的硬件同步锁在大多数时钟周期内保持一致。因此,直接内存访问单元能够直接根据第一分布式同步锁的工作状态值和值,确定是否执行第一任务指令包括的第一数据处理任务。
在一些实施例中,第一分布式同步锁的工作状态值包括空闲状态和非空闲状态。空闲状态表示分布式同步锁对应的上一个任务已经执行完成,直接内存访问单元当前能够执行第一数据处理任务。非空闲状态表示分布式同步锁对应的上一个任务正在执行,直接内存访问单元当前不能执行第一数据处理任务。通过第一分布式同步锁的任务类型值,能够确定当前是否满足第一数据处理任务的执行条件。执行条件用于控制第一分布式同步锁的任务类型值与执行第一数据处理任务之间的关系。通过为任务配置执行条件,能够实现通过分布式同步锁控制直接内存访问单元执行任务的时序。
在一些实施例中,直接内存访问单元通过两个多路选择器确定第一分布式同步锁的工作状态值和值。对于直接内存访问单元中的第一多路选择器,第一多路选择器基于第一任务指令包括的第一同步锁标识,在第一同步锁序列中确定第一同步锁标识对应的第一分布式同步锁和第一分布式同步锁的工作状态值。其中,第一同步锁序列用于指示直接内存访问单元中的多个分布式同步锁的工作状态值。状态包括空闲状态和非空闲状态。对于直接内存访问单元中的第二多路选择器,第二多路选择器基于第一同步锁标识,在第二同步锁序列中确定第一分布式同步锁的任务类型值。其中,第二同步锁序列用于指示直接内存访问单元中的多个分布式同步锁的任务类型值。直接内存访问单元基于第一任务指令包括的第一数据处理任务的任务类型,确定第一数据处理任务的执行条件。执行条件用于指示第一分布式同步锁的任务类型值所要满足的条件。其中,任务类型包括读任务类型和写任务类型。通过直接内存访问单元中的两个多路选择器,能够较为快速的从两个同步锁序列中确定第一分布式同步锁、第一分布式同步锁的工作状态值以及第一分布式同步锁的任务类型值。从而提高了直接内存访问单元执行第一数据处理任务的效率。
在一些实施例中,第一数据处理任务的任务类型为读任务类型。直接内存访问单元将第一分布式同步锁的任务类型值为第一数值作为执行条件。其中,读任务类型的第一数据处理任务为直接内存访问单元读取外部存储器中存储的待处理的目标数据,并将目标数据搬移至存储器内部。通过数据处理任务的任务类型,将第一分布式同步锁不同的值作为执行条件,能够为不同任务类型的任务确定不同的执行条件,从而实现控制任务之间的执行时序。
在一些实施例中,直接内存访问单元能够通过图4所示的执行条件的判断逻辑的示意图,根据第一分布式同步锁的工作状态值、第一分布式同步锁的任务类型值以及第一数据处理任务的任务类型,确定当前是否满足第一数据处理任务的执行条件。如图4所示,执行条件的判断逻辑包括与逻辑门401、与逻辑门402和多路选择器403。取非符号404用于表示在将同步锁的工作状态值(busy)或者同步锁的任务类型值(sync)输入至与逻辑门之前,将输入取反。需要说明的是,图4中的取非符号401仅为一种示意性的举例。直接内存访问单元能够根据实际需要,确定是否将同步锁的工作状态值或者同步锁的任务类型值取反。以实现多路选择器根据与逻辑门的输出结果,判断当前是否能够执行数据处理任务。因此,多路选择器能够根据与逻辑门的输出和任务类型(sync_set_value),判断当前是否满足数据处理任务的执行条件,并输出判断结果(issue_en)。其中,busy包括0和1两个取值。同步锁的工作状态值busy为0,表示同步锁的工作状态为非空闲状态,同步锁的工作状态值busy为1,表示同步锁的工作状态为空闲状态。sync_set_value也包括0和1两个取值。sync_set_value为1表示任务类型为读任务,sync_set_value为0表示任务类型为写任务。
例如,对于读任务类型(sync_set_value为1)的数据处理任务,根据上述执行逻辑可知,在同步锁的工作状态值为表征空闲状态的预设值(busy为1),且同步锁的任务类型值(sync)为0的情况下,当前能够执行该数据处理任务。因此,直接内存访问单元能够在与逻辑门401对输入的同步锁的任务类型值进行取反,输入到与逻辑门401的sync即为1。直接内存访问单元在与逻辑门401不对同步锁的工作状态值进行取反,输入到与逻辑门401的busy即为1。与逻辑门401的两个输入均为1,与逻辑门401的输出即为1。相应地,由于当前busy为1,且sync为0,因此直接内存访问单元在与逻辑门402可以对二者都进行取反,也可以都不进行取反,还可以只对busy取反,对sync不取反。上述三种情况与逻辑门402的输出均为0。因此,多路选择器403根据与逻辑门401输出的值为1,与逻辑门402输出的值为0,能够确定当前能够执行上述数据处理任务,并且直接内存访问单元上述判断的逻辑与该数据处理任务的执行逻辑一致。
在一些实施例中,直接内存访问单元根据sync_set_value所指示的任务类型,能够确定第一数据处理任务的执行条件,也即是确定满足执行条件时的sync。例如,在第一数据处理任务的sync_set_value为1时,表明第一数据处理任务为读任务。直接内存访问单元将sync为0作为执行条件,也即是第一数值为0。因此,在busy为1,且sync为0时,才满足第一数据处理任务的执行条件。相应地,在busy为1,且sync为0时,多路选择器403输出当前满足执行条件的判断结果。在busy和sync为其他数值时,多路选择器403输出当前不满足执行条件的判断结果。
306、直接内存访问单元在第一分布式同步锁的工作状态值为表征空闲状态的预设值,且第一分布式同步锁的任务类型值满足第一数据处理任务的执行条件的情况下,执行第一任务指令包括的第一数据处理任务,向硬件同步锁管理单元发送任务完成指令。
在本申请实施例中,在直接内存访问单元确定第一分布式同步锁的工作状态值为表征空闲状态的预设值,且第一分布式同步锁的任务类型值满足第一数据处理任务的执行条件的情况下,表明当前能够执行第一数据处理任务。直接内存访问单元执行第一数据处理任务,从外部存储器搬移待处理的目标数据至数据处理系统内部。直接内存访问单元搬移数据完成之后,直接内存访问单元向硬件同步锁管理单元发送任务完成指令。任务完成指令用于通知硬件同步锁管理单元当前第一数据处理任务已经执行完成。
在一些实施例中,直接内存访问单元执行第一数据处理任务的过程中,直接内存访问单元启动req_interface指令进行外部存储器的访问请求,以从外部存储器搬移目标数据。在访问请求发送完毕,且访问请求被外部存储器处理完毕的情况下,表明直接内存访问单元完成了目标数据的搬移。然后,直接内存访问单元向硬件同步锁管理单元发送任务完成指令。
307、硬件同步锁管理单元基于任务完成指令中的第一同步锁标识,在硬件同步锁序列中确定第一硬件同步锁。
在本申请实施例中,硬件同步锁管理单元接收直接内存访问单元发送的任务完成指令。硬件同步锁管理单元确定任务完成指令中包括的第一同步锁标识。硬件同步锁管理单元在硬件同步锁序列中确定与第一同步锁标识对应的第一硬件同步锁。第一硬件同步锁与直接内存访问单元模块中的第一分布式同步锁均与第一同步锁标识对应。
308、硬件同步锁管理单元在第二数量个时钟周期之后,更新第一硬件同步锁的任务类型值。
在本申请实施例中,硬件同步锁管理单元和直接内存访问单元位于数据处理系统中不同的位置,且位置分布较远。因此需要通过延迟打拍器延迟更新第一硬件同步锁的任务类型值,以满足硬件同步锁管理单元和直接内存访问单元的时序要求。延迟的时钟周期的数量与延迟打拍器的数量关联。硬件同步锁管理单元通过第二数量个延迟打拍器,能够在第二数量个时钟周期之后,更新第一硬件同步锁的任务类型值。其中,第二数量可以为一个预设的数值,如2,3或者5,本申请实施例对此不进行限制。通过在直接内存访问单元完 成第一数据处理任务之后,更新第一硬件同步锁的任务类型值,能够实现通过第一硬件同步锁的任务类型值表示直接内存访问单元执行任务的执行情况,以及通过第一硬件同步锁的任务类型值控制直接内存访问单元执行任务的时序。
在一些实施例中,硬件同步锁管理单元能够根据第一数据处理任务的类型,确定第一硬件同步锁更新的值。在第一数据处理任务的任务类型为读任务类型的情况下,硬件同步锁管理单元在第二数量个时钟周期之后,将第一硬件同步锁的任务类型值由第一数值更新为第二数值。例如,硬件同步锁管理单元在三个时钟周期之后,将第一硬件同步锁的任务类型值由0更新为1。通过数据处理任务的任务类型,将第一硬件同步锁更新为不同的值,能够根据第一硬件同步锁的任务类型值,确定是否满足不同任务类型的任务的执行条件。而实现了根据硬件同步锁控制任务的执行时序。
例如,硬件同步锁管理单元能够通过sync_set_value,确定第一数据处理任务和第一硬件同步锁更新之后的值。在sync_set_value为1的情况下,第一数据处理任务的任务类型为读任务,第一硬件同步锁更新之后的值为1。在sync_set_value为0的情况下,第一数据处理任务的任务类型为写任务,第一硬件同步锁更新之后的值为0。
为了更清楚的说明上述直接内存访问单元执行第一数据处理任务的过程,下面结合图5所示的数据处理的流程图,对上述过程进行说明。如图5所示,响应于数据处理系统接收到数据处理任务,硬件同步锁管理单元在硬件同步锁序列中确定至少一个空闲的硬件同步锁和每个硬件同步锁的标识。然后,硬件同步锁管理单元将任一硬件同步锁的标识作为第一同步锁标识发送至标量处理单元。标量处理单元接收到第一同步锁标识,标量处理单元将第一同步锁标识和第一数据处理任务作为第一任务指令发送到直接内存访问单元。直接内存访问单元基于接收到的第一任务指令,通过两个多路选择器(mux),在第一同步锁序列中确定第一同步锁标识对应的第一分布式同步锁和第一分布式同步锁的工作状态值(busy),在第二同步锁序列中确定第一分布式同步锁的任务类型值(sync)。然后,直接内存访问单元根据第一分布式同步锁的工作状态值、第一分布式同步锁的任务类型值以及第一数据处理任务的任务类型,判断是否执行第一数据处理任务。例如,在第一数据处理任务的任务类型为读任务,第一同步锁的工作状态值为表征空闲状态的预设值,也即是busy为1,第一同步锁的任务类型值为0的情况下,才能够执行第一数据处理任务。在判断结果指示能够执行第一数据处理任务的情况下,直接内存访问单元执行第一数据处理任务。直接内存访问单元访问外部存储器,搬移外部存储器中待处理的目标数据至数据处理系统内部。直接内存访问单元搬移数据完成之后,直接内存访问单元向硬件同步锁管理单元发送任务完成指令。硬件同步锁管理单元根据任务完成指令中的第一同步锁标识,在硬件同步锁序列中确定第一硬件同步锁。硬件同步锁管理单元通过三个延迟打拍器,在三个时钟周期之后将第一硬件同步锁的任务类型值由0更新为1。再通过三个延迟打拍器,在三个时钟周期之后,将第二同步锁序列中第一同步锁标识对应的第一分布式同步锁的任务类型值由0更新为1。
309、向量处理单元基于标量处理单元发送的第二任务指令,在向量处理单元中确定第二任务指令包括的第一同步锁标识所指示的第二分布式同步锁。
在本申请实施例中,向量处理单元接收标量处理单元发送的第二任务指令。向量处理单元确定第二任务指令中包括的第一同步锁标识。与直接内存访问单元类似,向量处理单元中也包括与硬件同步锁序列对应的多个分布式同步锁。向量处理单元能够基于第一同步锁标识,在向量处理单元中的多个分布式同步锁中确定与第一同步锁标识对应的第二分布式同步锁。并且,向量处理单元中的第二分布式同步锁、直接内存访问单元中的第一分布式同步锁和第一硬件同步锁在大多数时钟周期内保持一致。因此,向量处理单元能够直接根据第二分布式同步锁的工作状态值和第二分布式同步锁的任务类型值,确定是否执行第二任务指令包括的第一数据处理任务。因此,向量处理单元无需确定外部硬件同步锁序列中第一硬件同步锁的工作状态值和值,从而实现了与外部硬件同步锁序列的解耦。
310、向量处理单元在第二分布式同步锁的工作状态值为表征空闲状态的预设值,且第二分布式同步锁的任务类型值满足第二数据处理任务的执行条件的情况下,执行第二数据处理任务,并在第二数据处理任务完成后向硬件同步锁管理单元发送任务完成指令。
在本申请实施例中,在向量处理单元确定第二分布式同步锁的工作状态值为表征空闲状态的预设值,且第二分布式同步锁的任务类型值满足第二数据处理任务的执行条件的情况下,表明当前能够执行第二数据处理任务。向量处理单元执行第二数据处理任务,对数据处理系统内部的目标数据进行处理。向量处理单元处理目标数据完成之后,向量处理单元向硬件同步锁管理单元发送任务完成指令。任务完成指令用于通知硬件同步锁管理单元当前第一数据处理任务已经执行完成。
在一些实施例中,向量处理单元执行第二数据处理任务的过程中,向量处理单元启动req_interface指令进行数据处理系统内部存储单元的访问请求,以处理目标数据。在访问请求发送完毕,且访问请求被内部存储单元处理完毕的情况下,表明向量处理单元完成了对目标数据的处理。然后,向量处理单元向硬件同步锁管理单元发送任务完成指令。
311、向量处理单元在第二分布式同步锁的任务类型值发生更新的情况下,更新第二分布式同步锁的工作状态值为表征空闲状态的预设值,第二分布式同步锁的任务类型值,在第一硬件同步锁的任务类型值发生更新的第一数量个时钟周期之后,更新为第一硬件同步锁的任务类型值。
在本申请实施例中,与直接内存访问单元同理,向量处理单元在执行完成第二数据处理任务之后,向硬件同步锁管理单元发送任务完成指令。相应地,硬件同步锁管理单元接收到任务完成指令后,通过延迟打拍器延迟若干个时钟周期更新第一硬件同步锁的任务类型值。在第一硬件同步锁的任务类型值发生更新的情况下,通过第一数量个延迟打拍器,在第一数量个时钟周期之后,将第二分布式同步锁的任务类型值更新为第一硬件同步锁的任务类型值。其中,第一数量可以为一个预设的数值,如2,3或者5,本申请实施例对此不进行限制。向量处理单元在第二分布式同步锁的任务类型值发生更新的情况下,将第二分布式同步锁的工作状态值更新表征空闲状态,以表明当前向量处理单元执行完成第二数据处理任务。进而,向量处理单元能够根据第二同步锁新的状态,确定是否执行下一个任务。
在一些实施例中,与直接内存访问单元类似,向量处理单元中包括第三同步锁序列和第四同步锁序列。第三同步锁序列用于指示向量处理单元中的多个分布式同步锁的工作状态值。第四同步锁序列用于指示向量处理单元中的多个分布式同步锁的任务类型值。因此,在第一硬件同步锁的任务类型值发生更新的情况下,在第一数量个时钟周期之后,将第四同步锁序列中第二分布式同步锁的任务类型值更新为第一硬件同步锁的任务类型值。向量处理单元在第二分布式同步锁的任务类型值发生更新的情况下,将第三同步锁序列中第二分布式同步锁的工作状态值更新表征空闲状态。
由此可见,在向量处理单元执行任务完成后,先延迟更新第一硬件同步锁的任务类型值,再延迟更新第二分布式同步锁的任务类型值。因此在向量处理单元执行完成任务和第二分布式同步锁的任务类型值发生更新之间存在延时。通过在第二分布式同步锁的任务类型值发生更新的情况下,才更新第二分布式同步锁的工作状态值为表征空闲状态的预设值,能够避免在两条连续的数据处理任务均对应第二分布式同步锁时,产生执行任务冲突问题。例如,在没有第三同步锁序列的情况下,向量处理单元执行完成第一条数据处理任务之后,将第一硬件同步锁的任务类型值由1更新为0。第二分布式同步锁的任务类型值发生更新的时间滞后于第一硬件同步锁的任务类型值发生更新的时间。此时第二分布式同步锁的任务类型值仍然为1,还未更新为0。如果向量处理单元根据当前的第二分布式同步锁的任务类型值,确定能够执行第二条数据处理任务,就开始执行第二数据处理任务,会产生任务冲突问题。通过引入分布式同步锁的工作状态值,并将分布式同步锁的工作状态值也作为判断任务是否能够执行的判断条件之一,能够避免上述任务冲突问题。如图6所示,第一 硬件同步锁的任务类型值由1更新为0,但第四同步锁序列中的第二分布式同步锁的任务类型值还未更新为0的情况下,第三同步锁序列中的第二分布式同步锁的工作状态值为非空闲状态,也即是busy为0。在第二分布式同步锁的任务类型值由1更新为0时,第二分布式同步锁的工作状态值更新表征空闲状态,也即是busy由0更新为1。
其中,在直接内存访问单元执行第一数据处理任务和向量处理单元执行第二数据处理任务过程中,第一硬件同步锁的任务类型值的更新情况如图7所示。直接内存访问单元在确定第一分布式同步锁的任务类型值为0,且第一分布式同步锁的工作状态值为表征空闲状态的预设值的情况下,直接内存访问单元执行第一数据处理任务。在直接内存访问单元执行完成第一数据处理任务之后,指示硬件同步锁管理单元将第一硬件同步锁的任务类型值更新为1。向量处理单元在确定第二分布式同步锁的任务类型值为1,且第二分布式同步锁的工作状态值为表征空闲状态的预设值的情况下,向量处理单元执行第二数据处理任务。在向量处理单元执行完成第二数据处理任务之后,指示硬件同步锁管理单元将第一硬件同步锁的任务类型值更新为0。
为了更清楚的说明数据处理系统中直接内存访问单元和向量处理单元执行数据处理任务的过程,下面结合图8所示的数据处理流程图,对上述过程进行说明。如图8所示,直接内存访问单元执行第一数据处理任务的流程与图5所示的流程图同理,在此不进行赘述。响应于数据处理系统接收到数据处理任务,硬件同步锁管理单元在硬件同步锁序列中确定至少一个空闲的硬件同步锁和每个硬件同步锁的标识。然后,硬件同步锁管理单元将任一硬件同步锁的标识作为第一同步锁标识发送至标量处理单元。标量处理单元接收到第一同步锁标识,标量处理单元将第一同步锁标识和第二数据处理任务作为第二任务指令发送到向量处理单元。向量处理单元基于接收到的第二任务指令,通过两个多路选择器(mux),在第三同步锁序列中确定第二同步锁标识对应的第二分布式同步锁和第二分布式同步锁的工作状态值(busy),在第四同步锁序列中确定第二分布式同步锁的任务类型值(sync)。然后,向量处理单元根据第二分布式同步锁的工作状态值、第二分布式同步锁的任务类型值以及第二数据处理任务的类型,判断是否执行第二数据处理任务。例如,在第二数据处理任务的任务类型为写任务,第二同步锁的工作状态值为表征空闲状态的预设值,也即是busy为1,第二同步锁的任务类型值为1的情况下,才能够执行第二数据处理任务。在判断结果指示能够执行第二数据处理任务的情况下,向量处理单元执行第二数据处理任务。向量处理单元对数据处理系统内部的目标数据进行处理。向量处理单元处理完成目标数据之后,向量处理单元向硬件同步锁管理单元发送任务完成指令。硬件同步锁管理单元根据任务完成指令中的第一同步锁标识,在硬件同步锁序列中确定第一硬件同步锁。硬件同步锁管理单元通过三个延迟打拍器,在三个时钟周期之后将第一硬件同步锁的任务类型值由1更新为0。再通过三个延迟打拍器,在三个时钟周期之后,将第二同步锁序列中第一同步锁标识对应的第二分布式同步锁的任务类型值由1更新为0。其中,硬件同步锁管理单元还能够合并直接内存访问单元和向量处理单元发送的任务完成指令,并根据合并的多条任务完成指令,更新硬件同步锁序列中对应的硬件同步锁的任务类型值。在一些实施例中,硬件同步锁管理单元还能够接收直接内存访问单元和向量处理单元以外其他单元发送的任务完成指令。相应的,硬件同步锁管理单元合并不同单元的任务完成指令,并更新硬件同步锁序列中对应的硬件同步锁的任务类型值,能够提高更新硬件同步锁的任务类型值的效率。
需要说明的是,上述步骤301-311是以直接内存访问单元执行读任务为例进行说明。也即是直接内存访问单元从外部存储器读取待处理的目标数据,将目标数据搬移至数据处理系统内部。在一些实施例中,也可以由向量处理单元执行读任务。相应地,向量处理单元读取数据处理系统内部待处理的目标数据,并对目标数据进行处理。再由直接内存访问单元将向量处理单元处理之后的目标数据写入外部存储器。
在一些实施例中,向量处理单元执行读任务的过程包括:标量处理单元向向量处理单 元发送第三任务指令,并向直接内存访问单元发送第四任务指令,第三任务指令包括第三数据处理任务和第二同步锁标识所指示的第三分布式同步锁,第四任务指令包括第四数据处理任务;向量处理单元在第三分布式同步锁的工作状态值为表征空闲状态的预设值,且第三分布式同步锁的任务类型值满足第三数据处理任务的执行条件的情况下,执行第三数据处理任务,并在第三数据处理任务完成后向硬件同步锁管理单元发送任务完成指令;硬件同步锁管理单元基于任务完成指令,更新硬件同步锁序列中第二同步锁标识所指示的第二硬件同步锁的任务类型值;直接内存访问单元在第二硬件同步锁的任务类型值发生更新的情况下,执行第四数据处理任务。
下面通过步骤(1)-(5)对上述向量处理单元执行读任务的过程进行说明。
(1)标量处理单元向向量处理单元发送第三任务指令,并向直接内存访问单元发送第四任务指令。
数据处理系统接收到数据处理任务之后,需要调用直接内存访问单元和向量处理单元执行数据处理任务。数据处理系统调用硬件同步锁管理单元,在硬件同步锁序列中确定至少一个空闲的硬件同步锁和每个硬件同步锁的标识。硬件同步锁管理单元将任一硬件同步锁的标识作为第二同步锁标识发送至标量处理单元。标量处理单元接收到第一同步锁标识之后,标量处理单元将第二同步锁标识和向量处理单元需要执行的第三数据处理任务作为第三任务指令。标量处理单元将第二同步锁标识和直接内存访问单元需要执行的第四数据处理任务作为第二任务指令。标量处理单元向向量处理单元发送第三任务指令,向直接内存访问单元发送第四任务指令。
(2)向量处理单元基于标量处理单元发送的第三任务指令,在向量处理单元中确定第三任务指令包括的第二同步锁标识所指示的第三分布式同步锁,第三分布式同步锁用于确定是否执行第三任务指令包括的第三数据处理任务。
(3)向量处理单元在第三分布式同步锁的工作状态值为表征空闲状态的预设值,且第三分布式同步锁的任务类型值满足第三数据处理任务的执行条件的情况下,执行第三任务指令包括的第三数据处理任务,对数据系统内部的目标数据进行处理。向量处理单元执行完成第三数据处理任务之后,向硬件同步锁管理单元发送任务完成指令。
其中,第三数据处理任务的任务类型为读任务类型。向量处理单元将第三分布式同步锁的任务类型值为第一数值作为第三数据处理任务的执行条件。例如,向量处理单元将第三分布式同步锁的任务类型值为0作为第三数据处理任务的执行条件。通过数据处理任务的任务类型,将第三分布式同步锁不同的值作为执行条件,能够为不同任务类型的任务确定不同的执行条件,从而实现控制任务之间的执行时序。
(4)硬件同步锁管理单元基于任务完成指令,更新硬件同步锁序列中第二同步锁标识所指示的第二硬件同步锁的任务类型值。
其中,第三数据处理任务的任务类型为读任务类型。硬件同步锁管理单元基于任务完成指令中的第二同步锁标识,在硬件同步锁序列中确定第二硬件同步锁。硬件同步锁管理单元在第三数量个时钟周期之后,将第二硬件同步锁的任务类型值由第一数值更新为第二数值。例如,硬件同步锁管理单元在三个时钟周期之后,将第二硬件同步锁的任务类型值由0更新为1。通过数据处理任务的任务类型,将第二硬件同步锁更新为不同的值,能够根据第二硬件同步锁的任务类型值,确定是否满足不同任务类型的任务的执行条件。而实现了根据硬件同步锁控制任务的执行时序。
(5)直接内存访问单元在第二硬件同步锁的任务类型值发生更新的情况下,执行标量处理单元发送的第四任务指令包括的第四数据处理任务。直接内存访问单元将向量处理单元处理之后的目标数据写入外部存储器。
本申请实施例提供了一种数据处理方法,由数据处理系统执行。标量处理单元能够向直接内存访问单元和向量处理单元分别发送任务指令。而无需标量处理单元等待直接内存访问单元执行任务完成后再向向量处理单元发送任务,能够提高标量处理单元的处理效率。 相应地,直接内存访问单元能够根据标量处理单元发送的第一任务指令中的第一同步锁标识,在直接内存访问单元中确定第一分布式同步锁。进而直接内存访问单元能够根据第一分布式同步锁的工作状态值和第一分布式同步锁的任务类型值,确定当前是否执行第一数据处理任务。而无需标量处理单元确定直接内存访问单元何时执行第一数据处理任务,实现了与标量处理单元的解耦。并且在直接内存访问单元完成任务之后直接内存访问单元向硬件同步锁管理单元发送任务完成指令。硬件同步锁管理单元基于任务完成指令,能够更新硬件同步锁序列中第一硬件同步锁的任务类型值。响应于第一硬件同步锁的任务类型值发生更新,向量处理单元及时执行第二数据处理任务,提高了处理器处理数据的效率。
本申请实施例还提供了一种数据处理系统。参见图8,数据处理系统包括:标量处理单元801、直接内存访问单元802、硬件同步锁管理单元803以及向量处理单元804。
标量处理单元801,用于向直接内存访问单元802发送第一任务指令,并向向量处理单元804发送第二任务指令;
直接内存访问单元802,用于基于标量处理单元801发送的第一任务指令,在直接内存访问单元802中确定第一任务指令包括的第一同步锁标识所指示的第一分布式同步锁,第一分布式同步锁用于确定是否执行第一任务指令包括的第一数据处理任务;
直接内存访问单元802,还用于在第一分布式同步锁的工作状态值为表征空闲状态的预设值,且第一分布式同步锁的任务类型值满足第一数据处理任务的执行条件的情况下,执行第一任务指令包括的第一数据处理任务,向硬件同步锁管理单元803发送任务完成指令;
硬件同步锁管理单元803,用于基于任务完成指令,更新硬件同步锁序列8031中第一同步锁标识所指示的第一硬件同步锁的任务类型值;
向量处理单元804,用于在第一硬件同步锁的任务类型值发生更新的情况下,执行标量处理单元801发送的第二任务指令包括的第二数据处理任务。
在一些实施例中,向量处理单元804,用于基于标量处理单元801发送的第二任务指令,在向量处理单元804中确定第二任务指令包括的第一同步锁标识所指示的第二分布式同步锁,第二分布式同步锁用于确定是否执行第二任务指令包括的第二数据处理任务;在第二分布式同步锁的任务类型值发生更新的情况下,更新第二分布式同步锁的工作状态值为表征空闲状态的预设值,第二分布式同步锁的任务类型值在第一硬件同步锁的任务类型值发生更新的第一数量个时钟周期之后更新为第一硬件同步锁的任务类型值;在第二分布式同步锁的工作状态值为表征空闲状态的预设值,且第二分布式同步锁的任务类型值满足第二数据处理任务的执行条件的情况下,执行第二任务指令包括的第二数据处理任务。
在一些实施例中,硬件同步锁管理单元803包括硬件同步锁序列8031;硬件同步锁管理单元803,用于在硬件同步锁序列8031中确定至少一个空闲的硬件同步锁和每个硬件同步锁的标识;将任一硬件同步锁的标识作为第一同步锁标识发送至标量处理单元801;标量处理单元801,用于响应于接收到第一同步锁标识,生成第一任务指令和第二任务指令。
在一些实施例中,直接内存访问单元802包括第一多路选择8021、第二多路选择器8022、第一同步锁序列8023以及第二同步锁序列8024;第一多路选择器8021,用于基于第一任务指令包括的第一同步锁标识,在第一同步锁序列8023中确定第一分布式同步锁和第一分布式同步锁的工作状态值,第一同步锁序列8023用于指示直接内存访问单元802中的多个分布式同步锁的工作状态值,状态包括空闲状态和非空闲状态,空闲状态用于指示直接内存访问单元802当前能够执行第一数据处理任务,非空闲状态用于指示直接内存访问单元802当前正在执行第一分布式同步锁对应的数据处理任务;第二多路选择器8022,用于基于第一同步锁标识,在第二同步锁序列8024中确定第一分布式同步锁的任务类型值,第二同步锁序列8024用于指示直接内存访问单元802中的多个分布式同步锁的任务类型值;直接内存访问单元802,还用于基于第一任务指令包括的第一数据处理任务的任务类型,确定第一数据处理任务的执行条件,执行条件用于指示第一分布式同步锁的任务 类型值所要满足的条件。
在一些实施例中,第一数据处理任务的任务类型为读任务类型,直接内存访问单元802,用于将第一分布式同步锁的任务类型值为第一数值作为执行条件。
在一些实施例中,硬件同步锁管理单元803,用于基于任务完成指令中的第一同步锁标识,在硬件同步锁序列8031中确定第一硬件同步锁;在第二数量个时钟周期之后,更新第一硬件同步锁的任务类型值。
在一些实施例中,硬件同步锁管理单元803,用于在第二数量个时钟周期之后,将第一硬件同步锁的任务类型值由第一数值更新为第二数值。
在一些实施例中,标量处理单元801,用于向向量处理单元803发送第三任务指令,并向直接内存访问单元802发送第四任务指令;向量处理单元804,用于基于标量处理单元801发送的第三任务指令,在向量处理单元804中确定第三任务指令包括的第二同步锁标识所指示的第三分布式同步锁,第三分布式同步锁用于确定是否执行第三任务指令包括的第三数据处理任务;向量处理单元804,还用于在第三分布式同步锁的工作状态值为表征空闲状态的预设值,且第三分布式同步锁的任务类型值满足第三数据处理任务的执行条件的情况下,执行第三任务指令包括的第三数据处理任务,向硬件同步锁管理单元803发送任务完成指令;硬件同步锁管理单元803,用于基于任务完成指令,更新硬件同步锁序列中第二同步锁标识所指示的第二硬件同步锁的任务类型值;直接内存访问单元802,用于在第二硬件同步锁的任务类型值发生更新的情况下,执行标量处理单元801发送的第四任务指令包括的第四数据处理任务。
在一些实施例中,第三数据处理任务的任务类型为读任务类型,向量处理单元804,用于将第三分布式同步锁的任务类型值为第一数值作为第三数据处理任务的执行条件。
在一些实施例中,第三数据处理任务的任务类型为读任务类型,硬件同步锁管理单元803,用于基于任务完成指令中的第二同步锁标识,在硬件同步锁序列8031中确定第二硬件同步锁;在第三数量个时钟周期之后,将第二硬件同步锁的任务类型值由第一数值更新为第二数值。
需要说明的是:上述实施例提供的数据处理系统,仅以上述各功能模块的划分进行举例说明,实际应用中,可以根据需要而将上述功能分配由不同的功能模块完成,即将终端的内部结构划分成不同的功能模块,以完成以上描述的全部或者部分功能。另外,上述实施例提供的数据处理系统和数据处理方法实施例属于同一构思,其具体实现过程详见方法实施例,这里不再赘述。
本申请实施例还提供了一种终端,该终端包括处理器和存储器,存储器中存储有数据处理方法,处理器用于通过数据处理系统实现上述实施例提供的数据处理方法。
图9是本申请实施例提供的一种终端的结构示意图。
终端900包括有:处理器901和存储器902。
处理器901可以包括一个或多个处理核心,比如4核心处理器、8核心处理器等。处理器901可以采用DSP(Digital Signal Processing,数字信号处理)、FPGA(Field Programmable Gate Array,现场可编程门阵列)、PLA(Programmable Logic Array,可编程逻辑阵列)中的至少一种硬件形式来实现。处理器901也可以包括主处理器和协处理器,主处理器是用于对在唤醒状态下的数据进行处理的处理器,也称CPU(Central Processing Unit,中央处理器);协处理器是用于对在待机状态下的数据进行处理的低功耗处理器。在一些实施例中,处理器901可以集成有GPU(Graphics Processing Unit,图像处理的交互器),GPU用于负责显示屏所需要显示的内容的渲染和绘制。一些实施例中,处理器901还可以包括AI(Artificial Intelligence,人工智能)处理器,该AI处理器用于处理有关机器学习的计算操作。
存储器902可以包括一个或多个计算机可读存储介质,该计算机可读存储介质可以是非暂态的。存储器902还可包括高速随机存取存储器,以及非易失性存储器,比如一个或 多个磁盘存储设备、闪存存储设备。在一些实施例中,存储器902中的非暂态的计算机可读存储介质用于存储至少一条计算机程序,该至少一条计算机程序用于被处理器901所具有以实现本申请中方法实施例提供的数据处理方法。
本领域技术人员可以理解,图9中示出的结构并不构成对终端900的限定,可以包括比图示更多或更少的组件,或者组合某些组件,或者采用不同的组件布置。
本申请实施例还提供了一种计算机可读存储介质,该计算机可读存储介质中存储有至少一条计算机程序,该至少一条计算机程序由处理器加载并执行,以实现上述实施例提供的数据处理方法。
本申请实施例还提供了一种计算机程序产品,包括计算机程序,计算机程序由处理器加载并执行,以实现如上述实施例提供的数据处理方法。
本领域普通技术人员可以理解实现上述实施例的全部或部分步骤可以通过硬件来完成,也可以通过程序来指令相关的硬件完成,上述的程序可以存储于一种计算机可读存储介质中,上述提到的存储介质可以是只读存储器,磁盘或光盘等。
以上实施例的各技术特征可以进行任意的组合,为使描述简洁,未对上述实施例中的各个技术特征所有可能的组合都进行描述,然而,只要这些技术特征的组合不存在矛盾,都应当认为是本说明书记载的范围。
以上所述实施例仅表达了本申请的几种实施方式,其描述较为具体和详细,但并不能因此而理解为对发明专利范围的限制。应当指出的是,对于本领域的普通技术人员来说,在不脱离本申请构思的前提下,还可以做出若干变形和改进,这些都属于本申请的保护范围。因此,本申请专利的保护范围应以所附权利要求为准。

Claims (20)

  1. 一种数据处理方法,由数据处理系统执行,所述数据处理系统包括直接内存访问单元、向量处理单元、标量处理单元以及硬件同步锁管理单元,所述方法包括:
    所述标量处理单元向所述直接内存访问单元发送第一任务指令,并向所述向量处理单元发送第二任务指令,其中,所述第一任务指令包括第一数据处理任务和指示第一分布式同步锁的第一同步锁标识,所述第二任务指令包括第二数据处理任务;
    所述直接内存访问单元从所述第一任务指令中基于所述第一同步锁标识识别所述第一分布式同步锁的工作状态值;
    在所述第一分布式同步锁的所述工作状态状值为表征空闲状态的预设值,且所述第一分布式同步锁的任务类型值满足所述第一数据处理任务的执行条件的情况下,所述直接内容访问单元执行所述第一数据处理任务,并在所述第一数据处理任务完成后向所述硬件同步锁管理单元发送任务完成指令;
    所述硬件同步锁管理单元基于所述任务完成指令,更新硬件同步锁序列中所述第一同步锁标识所指示的第一硬件同步锁的任务类型值;及
    所述向量处理单元在所述第一硬件同步锁的任务类型值发生更新的情况下,执行所述第二数据处理任务。
  2. 根据权利要求1所述的方法,所述向量处理单元在所述第一硬件同步锁的任务类型值发生更新的情况下,执行所述第二数据处理任务,包括:
    所述向量处理单元基于所述第二任务指令,在所述向量处理单元中确定所述第二任务指令包括的第一同步锁标识所指示的第二分布式同步锁;
    所述向量处理单元在所述第二分布式同步锁的任务类型值发生更新的情况下,更新所述第二分布式同步锁的工作状态值为表征空闲状态的预设值,所述第二分布式同步锁的任务类型值,在所述第一硬件同步锁的任务类型值发生更新的第一数量个时钟周期之后,更新为所述第一硬件同步锁的任务类型值;
    所述向量处理单元在所述第二分布式同步锁的工作状态值为表征空闲状态的预设值,且所述第二分布式同步锁的任务类型值满足所述第二数据处理任务的执行条件的情况下,执行所述第二任务指令包括的第二数据处理任务。
  3. 根据权利要求1或2所述的方法,所述方法还包括:
    所述硬件同步锁管理单元在所述硬件同步锁序列中确定至少一个空闲的硬件同步锁和每个硬件同步锁的标识;
    所述硬件同步锁管理单元将任一硬件同步锁的标识作为所述第一同步锁标识发送至所述标量处理单元;
    响应于所述标量处理单元接收到所述第一同步锁标识,所述标量处理单元生成所述第一任务指令和所述第二任务指令。
  4. 根据权利要求1至3任一项所述的方法,所述方法还包括:
    所述直接内存访问单元中的第一多路选择器基于所述第一任务指令包括的第一同步锁标识,在第一同步锁序列中确定所述第一分布式同步锁和所述第一分布式同步锁的工作状态值,所述第一同步锁序列用于指示所述直接内存访问单元中的多个分布式同步锁的工作状态值,所述状态包括空闲状态和非空闲状态,所述空闲状态用于指示所述直接内存访问单元当前能够执行所述第一数据处理任务,所述非空闲状态用于指示所述直接内存访问单元当前正在执行所述第一分布式同步锁对应的数据处理任务;
    所述直接内存访问单元中的第二多路选择器基于所述第一同步锁标识,在第二同步锁序列中确定所述第一分布式同步锁的任务类型值,所述第二同步锁序列用于指示所述直接内存访问单元中的多个分布式同步锁的任务类型值;
    所述直接内存访问单元基于所述第一任务指令包括的第一数据处理任务的任务类型,确定所述第一数据处理任务的执行条件,所述执行条件用于指示所述第一分布式同步锁的 任务类型值所要满足的条件。
  5. 根据权利要求1至4任一项所述的方法,所述第一数据处理任务的任务类型为读任务类型,所述直接内存访问单元基于所述第一任务指令包括的第一数据处理任务的任务类型,确定所述第一数据处理任务的执行条件,包括:
    所述直接内存访问单元将所述第一分布式同步锁的任务类型值为第一数值作为所述执行条件。
  6. 根据权利要求1至5任一项所述的方法,所述硬件同步锁管理单元基于所述任务完成指令,更新硬件同步锁序列中所述第一同步锁标识所指示的第一硬件同步锁的任务类型值,包括:
    所述硬件同步锁管理单元基于所述任务完成指令中的所述第一同步锁标识,在所述硬件同步锁序列中确定所述第一硬件同步锁;
    所述硬件同步锁管理单元在第二数量个时钟周期之后,更新所述第一硬件同步锁的任务类型值。
  7. 根据权利要求1至6任一项所述的方法,所述第一数据处理任务的任务类型为读任务类型,所述硬件同步锁管理单元在第二数量个时钟周期之后,更新所述第一硬件同步锁的任务类型值,包括:
    所述硬件同步锁管理单元在所述第二数量个时钟周期之后,将所述第一硬件同步锁的任务类型值由第一数值更新为第二数值。
  8. 根据权利要求1至7任一项所述的方法,所述方法还包括:
    所述标量处理单元向所述向量处理单元发送第三任务指令,并向所述直接内存访问单元发送第四任务指令,所述第三任务指令包括第三数据处理任务和第二同步锁标识所指示的第三分布式同步锁,所述第四任务指令包括第四数据处理任务;
    所述向量处理单元在所述第三分布式同步锁的工作状态值为表征空闲状态的预设值,且所述第三分布式同步锁的任务类型值满足所述第三数据处理任务的执行条件的情况下,执行所述第三数据处理任务,并在所述第三数据处理任务完成后向所述硬件同步锁管理单元发送任务完成指令;
    所述硬件同步锁管理单元基于所述任务完成指令,更新硬件同步锁序列中所述第二同步锁标识所指示的第二硬件同步锁的任务类型值;
    所述直接内存访问单元在所述第二硬件同步锁的任务类型值发生更新的情况下,执行所述第四数据处理任务。
  9. 根据权利要求1至8任一项所述的方法,所述第三数据处理任务的任务类型为读任务类型,所述方法还包括:
    所述向量处理单元将所述第三分布式同步锁的任务类型值为第一数值作为所述第三数据处理任务的执行条件。
  10. 根据权利要求1至8任一项所述的方法,所述第三数据处理任务的任务类型为读任务类型,所述硬件同步锁管理单元基于所述任务完成指令,更新硬件同步锁序列中所述第二同步锁标识所指示的第二硬件同步锁的任务类型值,包括:
    所述硬件同步锁管理单元基于所述任务完成指令中的所述第二同步锁标识,在所述硬件同步锁序列中确定所述第二硬件同步锁;
    所述硬件同步锁管理单元在第三数量个时钟周期之后,将所述第二硬件同步锁的任务类型值由第一数值更新为第二数值。
  11. 一种数据处理系统,所述数据处理系统包括直接内存访问单元、向量处理单元、标量处理单元以及硬件同步锁管理单元;
    所述标量处理单元用于向所述直接内存访问单元发送第一任务指令,并向所述向量处理单元发送第二任务指令,其中,所述第一任务指令包括第一数据处理任务和指示第一分布式同步锁的第一同步锁标识,所述第二任务指令包括第二数据处理任务;
    所述直接内存访问单元用于从所述第一任务指令中基于所述第一同步锁标识识别所述第一分布式同步锁的工作状态值
    所述直接内存访问单元还用于在所述第一分布式同步锁的所述工作状态状值为表征空闲状态的预设值,且所述第一分布式同步锁的任务类型值满足所述第一数据处理任务的执行条件的情况下,所述直接内容访问单元执行所述第一数据处理任务,并在所述第一数据处理任务完成后向所述硬件同步锁管理单元发送任务完成指令;
    所述硬件同步锁管理单元用于基于所述任务完成指令,更新硬件同步锁序列中所述第一同步锁标识所指示的第一硬件同步锁的任务类型值;及
    所述向量处理单元用于在所述第一硬件同步锁的任务类型值发生更新的情况下,执行所述第二数据处理任务。
  12. 根据权利要求11所述的系统,所述向量处理单元用于:基于所述第二任务指令,在所述向量处理单元中确定所述第二任务指令包括的第一同步锁标识所指示的第二分布式同步锁;在所述第二分布式同步锁的任务类型值发生更新的情况下,更新所述第二分布式同步锁的工作状态值为表征空闲状态的预设值,所述第二分布式同步锁的任务类型值,在所述第一硬件同步锁的任务类型值发生更新的第一数量个时钟周期之后,更新为所述第一硬件同步锁的任务类型值;在所述第二分布式同步锁的工作状态值为表征空闲状态的预设值,且所述第二分布式同步锁的任务类型值满足所述第二数据处理任务的执行条件的情况下,执行所述第二任务指令包括的第二数据处理任务。
  13. 根据权利要求11或12所述的系统,所述硬件同步锁单元包括所述硬件同步锁序列;
    所述硬件同步锁管理单元用于在所述硬件同步锁序列中确定至少一个空闲的硬件同步锁和每个硬件同步锁的标识;
    所述硬件同步锁管理单元还用于将任一硬件同步锁的标识作为所述第一同步锁标识发送至所述标量处理单元;
    所述标量处理单元用于响应于接收到所述第一同步锁标识,生成所述第一任务指令和所述第二任务指令。
  14. 根据权利要求11至13任一项所述的系统,所述直接内存访问单元包括第一多路选择器、第二多路选择器、第一同步锁序列以及第二同步锁序列;
    所述第一多路选择器用于基于所述第一任务指令包括的第一同步锁标识,在第一同步锁序列中确定所述第一分布式同步锁和所述第一分布式同步锁的工作状态值,所述第一同步锁序列用于指示所述直接内存访问单元中的多个分布式同步锁的工作状态值,所述状态包括空闲状态和非空闲状态,所述空闲状态用于指示所述直接内存访问单元当前能够执行所述第一数据处理任务,所述非空闲状态用于指示所述直接内存访问单元当前正在执行所述第一分布式同步锁对应的数据处理任务;
    所述第二多路选择器用于基于所述第一同步锁标识,在第二同步锁序列中确定所述第一分布式同步锁的任务类型值,所述第二同步锁序列用于指示所述直接内存访问单元中的多个分布式同步锁的任务类型值;
    所述直接内存访问单元用于基于所述第一任务指令包括的第一数据处理任务的任务类型,确定所述第一数据处理任务的执行条件,所述执行条件用于指示所述第一分布式同步锁的任务类型值所要满足的条件。
  15. 根据权利要求11至14任一项所述的系统,所述第一数据处理任务的任务类型为读任务类型;
    所述直接内存访问单元用于将所述第一分布式同步锁的任务类型值为第一数值作为所述执行条件。
  16. 根据权利要求11至15任一项所述的系统,所述硬件同步锁单元包括所述硬件同步锁序列;
    所述硬件同步锁管理单元用于:基于所述任务完成指令中的所述第一同步锁标识,在所述硬件同步锁序列中确定所述第一硬件同步锁;在第二数量个时钟周期之后,更新所述第一硬件同步锁的任务类型值。
  17. 根据权利要求11至16任一项所述的系统,所述第一数据处理任务的任务类型为读任务类型;
    所述硬件同步锁管理单元用于在所述第二数量个时钟周期之后,将所述第一硬件同步锁的任务类型值由第一数值更新为第二数值。
  18. 根据权利要求11至17任一项所述的系统,所述向量处理单元包括第三分布式同步锁;
    所述标量处理单元用于向所述向量处理单元发送第三任务指令,并向所述直接内存访问单元发送第四任务指令,所述第三任务指令包括第三数据处理任务和第二同步锁标识所指示的第三分布式同步锁,所述第四任务指令包括第四数据处理任务;
    所述向量处理单元用于在所述第三分布式同步锁的工作状态值为表征空闲状态的预设值,且所述第三分布式同步锁的任务类型值满足所述第三数据处理任务的执行条件的情况下,执行所述第三数据处理任务,并在所述第三数据处理任务完成后向所述硬件同步锁管理单元发送任务完成指令;
    所述硬件同步锁管理单元用于基于所述任务完成指令,更新硬件同步锁序列中所述第二同步锁标识所指示的第二硬件同步锁的任务类型值;
    所述直接内存访问单元用于在所述第二硬件同步锁的任务类型值发生更新的情况下,执行所述第四数据处理任务。
  19. 一种芯片,所述芯片包括如权利要求11至18任一项所述的数据处理系统。
  20. 一种终端,所述终端包括处理器和存储器,所述存储器中存储有如权利要求1至10任一项所述的数据处理方法,所述处理器用于通过数据处理系统实现所述数据处理方法。
PCT/CN2024/076589 2023-03-30 2024-02-07 数据处理方法、系统、芯片及终端 Ceased WO2024198748A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
EP24777542.2A EP4586100A4 (en) 2023-03-30 2024-02-07 Data processing method and system, chip, and terminal
US19/190,147 US20250258789A1 (en) 2023-03-30 2025-04-25 Data processing method and system, chip, and terminal

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202310361448.5 2023-03-30
CN202310361448.5A CN116955244A (zh) 2023-03-30 2023-03-30 数据处理方法、系统、芯片及终端

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US19/190,147 Continuation US20250258789A1 (en) 2023-03-30 2025-04-25 Data processing method and system, chip, and terminal

Publications (1)

Publication Number Publication Date
WO2024198748A1 true WO2024198748A1 (zh) 2024-10-03

Family

ID=88445050

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/076589 Ceased WO2024198748A1 (zh) 2023-03-30 2024-02-07 数据处理方法、系统、芯片及终端

Country Status (4)

Country Link
US (1) US20250258789A1 (zh)
EP (1) EP4586100A4 (zh)
CN (1) CN116955244A (zh)
WO (1) WO2024198748A1 (zh)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN116955244A (zh) * 2023-03-30 2023-10-27 腾讯科技(深圳)有限公司 数据处理方法、系统、芯片及终端

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9785471B1 (en) * 2016-09-27 2017-10-10 International Business Machines Corporation Execution of serialized work data with interlocked full area instructions
CN108536526A (zh) * 2017-03-02 2018-09-14 腾讯科技(深圳)有限公司 一种基于可编程硬件的资源管理方法以及装置
CN112799791A (zh) * 2021-01-22 2021-05-14 平安普惠企业管理有限公司 分布式锁的调用方法、装置、电子设备和存储介质
CN113986534A (zh) * 2021-10-15 2022-01-28 腾讯科技(深圳)有限公司 任务调度方法、装置、计算机设备和计算机可读存储介质
CN114691376A (zh) * 2022-04-11 2022-07-01 深圳Tcl新技术有限公司 一种线程执行方法、装置、电子设备和存储介质
CN116955244A (zh) * 2023-03-30 2023-10-27 腾讯科技(深圳)有限公司 数据处理方法、系统、芯片及终端

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2015055083A1 (en) * 2013-10-14 2015-04-23 International Business Machines Corporation Adaptive process for data sharing with selection of lock elision and locking
CN104699631B (zh) * 2015-03-26 2018-02-02 中国人民解放军国防科学技术大学 Gpdsp中多层次协同与共享的存储装置和访存方法

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US9785471B1 (en) * 2016-09-27 2017-10-10 International Business Machines Corporation Execution of serialized work data with interlocked full area instructions
CN108536526A (zh) * 2017-03-02 2018-09-14 腾讯科技(深圳)有限公司 一种基于可编程硬件的资源管理方法以及装置
CN112799791A (zh) * 2021-01-22 2021-05-14 平安普惠企业管理有限公司 分布式锁的调用方法、装置、电子设备和存储介质
CN113986534A (zh) * 2021-10-15 2022-01-28 腾讯科技(深圳)有限公司 任务调度方法、装置、计算机设备和计算机可读存储介质
CN114691376A (zh) * 2022-04-11 2022-07-01 深圳Tcl新技术有限公司 一种线程执行方法、装置、电子设备和存储介质
CN116955244A (zh) * 2023-03-30 2023-10-27 腾讯科技(深圳)有限公司 数据处理方法、系统、芯片及终端

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See also references of EP4586100A4 *

Also Published As

Publication number Publication date
EP4586100A1 (en) 2025-07-16
US20250258789A1 (en) 2025-08-14
EP4586100A4 (en) 2026-02-25
CN116955244A (zh) 2023-10-27

Similar Documents

Publication Publication Date Title
US9582320B2 (en) Computer systems and methods with resource transfer hint instruction
CN110968415B (zh) 多核处理器的调度方法、装置及终端
CN112363763A (zh) 数据处理方法、装置及计算机可读存储介质
CN111143272A (zh) 异构计算平台的数据处理方法、装置及可读存储介质
US10997102B2 (en) Multidimensional address generation for direct memory access
CN109213607B (zh) 一种多线程渲染的方法和装置
JPS58151655A (ja) 情報処理装置
JPH0535453B2 (zh)
JPH0535454B2 (zh)
CN116360941A (zh) 一种面向多核dsp的并行计算资源自组织调度方法及系统
CN112559403B (zh) 一种处理器及其中的中断控制器
US20250258789A1 (en) Data processing method and system, chip, and terminal
CN100361119C (zh) 计算系统
TW202243714A (zh) 資訊顯示方法、裝置、終端、儲存媒體及電腦程式產品
US8185722B2 (en) Processor instruction set for controlling threads to respond to events
US20120151145A1 (en) Data Driven Micro-Scheduling of the Individual Processing Elements of a Wide Vector SIMD Processing Unit
WO2023086204A1 (en) Reducing latency in highly scalable hpc applications via accelerator-resident runtime management
US20080229310A1 (en) Processor instruction set
CN118796701A (zh) 代码测试方法、装置、设备及介质
CN115391005A (zh) 一种Linux实时处理方法及装置、设备和介质
CN118916123A (zh) 多核处理器的任务调度方法、装置、设备、介质和产品
CN110716750B (zh) 用于部分波前合并的方法和系统
KR100639146B1 (ko) 카테시안 제어기를 갖는 데이터 처리 시스템
CN112988355A (zh) 程序任务的调度方法、装置、终端设备及可读存储介质
JP3014605B2 (ja) ファジィ・コンピュータ

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24777542

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 2024777542

Country of ref document: EP

ENP Entry into the national phase

Ref document number: 2024777542

Country of ref document: EP

Effective date: 20250409

WWP Wipo information: published in national office

Ref document number: 2024777542

Country of ref document: EP

NENP Non-entry into the national phase

Ref country code: DE