WO2023087640A1 - 数据处理方法、装置、缓存器、处理器及电子设备 - Google Patents
数据处理方法、装置、缓存器、处理器及电子设备 Download PDFInfo
- Publication number
- WO2023087640A1 WO2023087640A1 PCT/CN2022/092980 CN2022092980W WO2023087640A1 WO 2023087640 A1 WO2023087640 A1 WO 2023087640A1 CN 2022092980 W CN2022092980 W CN 2022092980W WO 2023087640 A1 WO2023087640 A1 WO 2023087640A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- data
- main memory
- data processing
- memory addresses
- cache
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F12/00—Accessing, addressing or allocating within memory systems or architectures
- G06F12/02—Addressing or allocation; Relocation
- G06F12/08—Addressing or allocation; Relocation in hierarchically structured memory systems, e.g. virtual memory systems
- G06F12/0802—Addressing of a memory level in which the access to the desired data or data block requires associative addressing means, e.g. caches
- G06F12/0893—Caches characterised by their organisation or structure
- G06F12/0897—Caches characterised by their organisation or structure with two or more cache hierarchy levels
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F13/00—Interconnection of, or transfer of information or other signals between, memories, input/output devices or central processing units
- G06F13/14—Handling requests for interconnection or transfer
- G06F13/16—Handling requests for interconnection or transfer for access to memory bus
- G06F13/1668—Details of memory controller
- G06F13/1673—Details of memory controller using buffers
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F12/00—Accessing, addressing or allocating within memory systems or architectures
- G06F12/02—Addressing or allocation; Relocation
- G06F12/08—Addressing or allocation; Relocation in hierarchically structured memory systems, e.g. virtual memory systems
- G06F12/0802—Addressing of a memory level in which the access to the desired data or data block requires associative addressing means, e.g. caches
- G06F12/0806—Multiuser, multiprocessor or multiprocessing cache systems
- G06F12/0811—Multiuser, multiprocessor or multiprocessing cache systems with multilevel cache hierarchies
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F12/00—Accessing, addressing or allocating within memory systems or architectures
- G06F12/02—Addressing or allocation; Relocation
- G06F12/08—Addressing or allocation; Relocation in hierarchically structured memory systems, e.g. virtual memory systems
- G06F12/0802—Addressing of a memory level in which the access to the desired data or data block requires associative addressing means, e.g. caches
- G06F12/0806—Multiuser, multiprocessor or multiprocessing cache systems
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F12/00—Accessing, addressing or allocating within memory systems or architectures
- G06F12/02—Addressing or allocation; Relocation
- G06F12/08—Addressing or allocation; Relocation in hierarchically structured memory systems, e.g. virtual memory systems
- G06F12/0802—Addressing of a memory level in which the access to the desired data or data block requires associative addressing means, e.g. caches
- G06F12/0844—Multiple simultaneous or quasi-simultaneous cache accessing
- G06F12/0846—Cache with multiple tag or data arrays being simultaneously accessible
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F12/00—Accessing, addressing or allocating within memory systems or architectures
- G06F12/02—Addressing or allocation; Relocation
- G06F12/08—Addressing or allocation; Relocation in hierarchically structured memory systems, e.g. virtual memory systems
- G06F12/0802—Addressing of a memory level in which the access to the desired data or data block requires associative addressing means, e.g. caches
- G06F12/0862—Addressing of a memory level in which the access to the desired data or data block requires associative addressing means, e.g. caches with prefetch
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F12/00—Accessing, addressing or allocating within memory systems or architectures
- G06F12/02—Addressing or allocation; Relocation
- G06F12/08—Addressing or allocation; Relocation in hierarchically structured memory systems, e.g. virtual memory systems
- G06F12/0802—Addressing of a memory level in which the access to the desired data or data block requires associative addressing means, e.g. caches
- G06F12/0893—Caches characterised by their organisation or structure
- G06F12/0895—Caches characterised by their organisation or structure of parts of caches, e.g. directory or tag array
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F12/00—Accessing, addressing or allocating within memory systems or architectures
- G06F12/02—Addressing or allocation; Relocation
- G06F12/08—Addressing or allocation; Relocation in hierarchically structured memory systems, e.g. virtual memory systems
- G06F12/12—Replacement control
- G06F12/121—Replacement control using replacement algorithms
- G06F12/123—Replacement control using replacement algorithms with age lists, e.g. queue, most recently used [MRU] list or least recently used [LRU] list
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F12/00—Accessing, addressing or allocating within memory systems or architectures
- G06F12/02—Addressing or allocation; Relocation
- G06F12/04—Addressing variable-length words or parts of words
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F12/00—Accessing, addressing or allocating within memory systems or architectures
- G06F12/02—Addressing or allocation; Relocation
- G06F12/08—Addressing or allocation; Relocation in hierarchically structured memory systems, e.g. virtual memory systems
- G06F12/0802—Addressing of a memory level in which the access to the desired data or data block requires associative addressing means, e.g. caches
- G06F12/0864—Addressing of a memory level in which the access to the desired data or data block requires associative addressing means, e.g. caches using pseudo-associative means, e.g. set-associative or hashing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F2212/00—Indexing scheme relating to accessing, addressing or allocation within memory systems or architectures
- G06F2212/10—Providing a specific technical effect
- G06F2212/1016—Performance improvement
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F2212/00—Indexing scheme relating to accessing, addressing or allocation within memory systems or architectures
- G06F2212/30—Providing cache or TLB in specific location of a processing system
- G06F2212/305—Providing cache or TLB in specific location of a processing system being part of a memory device, e.g. cache DRAM
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F2212/00—Indexing scheme relating to accessing, addressing or allocation within memory systems or architectures
- G06F2212/60—Details of cache memory
- G06F2212/6028—Prefetching based on hints or prefetch instructions
Definitions
- a processor and multiple levels of memory are usually included.
- main memory stores instructions and data.
- the processor fetches instructions and corresponding data from main memory, executes the instructions, and writes the resulting data back to main memory.
- cache memory There is usually one or more levels of cache memory (Cache) between the processor and the main memory. Cache memory is used to reduce the time it takes for the processor to fetch instructions and data.
- the processor needs to read data at an address in main memory, it first checks to see if the data exists in the cache memory. If the cache contains the data, the processor reads the data directly from the cache, which is much faster than reading the data from main memory. Otherwise, the cache reads the data from the main memory, stores it in the cache, and returns it to the main memory.
- the width of the data channel between the caches is multiplied by the number of times the channel can transfer data per unit time, which is the bandwidth, which shows that the cache can be used per unit time.
- At least one embodiment of the present disclosure provides a data processing method, including: receiving a data processing request, the data requested by the data processing request includes data suitable for being stored in at least two cache units, and the data of each of the cache units
- the main memory addresses are continuous; when the main memory address information of each cache unit that satisfies the mapping relationship includes all the main memory addresses, data processing is performed on the data corresponding to each of the main memory addresses at the same time, wherein the satisfying mapping
- Each cache unit in the relationship refers to the cache unit corresponding to the main memory address in the data processing request.
- the cache unit to be replaced is determined by the least recently used principle (RLU).
- RLU least recently used principle
- the data processing request includes a data read request
- the main memory address information of each cache unit satisfying the mapping relationship includes all the main memory addresses
- the step of simultaneously performing data processing on the data corresponding to each of the main memory addresses includes: when the main memory address information of each cache unit that satisfies the mapping relationship includes all the main memory addresses, returning all the main memory addresses at the same time The data corresponding to the address.
- the step of generating a missing data read request for reading data corresponding to each of the missing main memory addresses includes: acquiring the missing data of each of the main memory addresses Address continuous main memory addresses in the address to obtain each continuous main memory address; according to each of the continuous main memory addresses, generate a continuous missing data read request for reading data corresponding to each of the missing continuous main memory addresses;
- the data processing method further includes: when all the data corresponding to the missing main memory addresses are received, returning the data corresponding to all the main memory addresses at the same time.
- the data processing request includes a data write request
- the main memory address information of each cache unit that satisfies the mapping relationship includes all the main memory addresses
- the step of simultaneously performing data processing on the data corresponding to each of the main memory addresses includes: when the main memory address information of each cache unit that satisfies the mapping relationship includes all of the main memory addresses, simultaneously receiving each of the main memory addresses The address corresponds to the data, and each of the data is written into the corresponding cache unit.
- the data processing request includes a data write request, and when the main memory address information of each cache unit satisfying the mapping relationship does not include all the main memory addresses , after the step of writing each of the missing main memory addresses into each buffer unit to be processed, it also includes: simultaneously receiving the data corresponding to each of the main memory addresses, and writing each of the data into the corresponding The cache unit.
- the data processing request further includes a data processing request type identifier, and the data processing request type identifier is used to identify that the data requested by the data processing request includes Data stored in at least two cache units.
- the implementation of the data processing request type identification includes increasing the number of main memory address identification bits or increasing the type of request type identification.
- At least one embodiment of the present disclosure further provides a data processing device, including: a data processing request module adapted to receive a data processing request, the data requested by the data processing request includes data suitable for being stored in at least two cache units, And the main memory addresses of the data of each of the cache units are continuous; the data processing module is adapted to process each of the main memory addresses when all the main memory addresses are included in the main memory address information of each cache unit that satisfies the mapping relationship. Data processing is performed on the corresponding data at the same time, wherein each cache unit satisfying the mapping relationship refers to the cache unit corresponding to the main memory address in the data processing request.
- the data processing module is further adapted to, when the main memory address information of each cache unit that satisfies the mapping relationship does not include all of the main memory addresses, the missing The main memory address is respectively written into each cache unit to be processed, wherein the original main memory address information of each cache unit to be processed is different from the main memory address of the data requested by the data processing request.
- the data processing request module includes: a data read request module adapted to receive a data read request, and the data requested by the data read request includes data suitable for being stored in at least The data of two cache units, and the main memory addresses of the data of each cache unit are continuous; the data processing module includes: a data read operation module, adapted to include in the main memory address information of each cache unit that satisfies the mapping relationship When all the main memory addresses are used, the data corresponding to all the main memory addresses are returned at the same time.
- the data processing request module includes: a data read request module adapted to receive a data read request, and the data requested by the data read request includes data suitable for being stored in at least The data of two cache units, and the main memory addresses of the data of each of the cache units are continuous; the data processing module includes: a data read operation module, adapted to satisfy the mapping relationship when there is no When all the main memory addresses are included, the missing main memory addresses are respectively written into each cache unit to be processed, and a missing data read request for reading data corresponding to the missing main memory addresses is generated and sent .
- the data read operation module is adapted to generate and send a missing data read request for reading missing data corresponding to each of the main memory addresses, including: generating respective missing data read requests for reading data corresponding to the missing main memory addresses, the number of the missing data read requests is the same as the number of the missing main memory addresses;
- the data read operation module is further adapted to return the data corresponding to all the main memory addresses simultaneously when all the data corresponding to the missing main memory addresses are received.
- the data read operation module is adapted to generate and send a missing data read request for reading missing data corresponding to each of the main memory addresses, including: obtaining The main memory addresses with continuous addresses in each of the missing main memory addresses are obtained to obtain each continuous main memory address; according to each of the continuous main memory addresses, the continuous data corresponding to each of the missing continuous main memory addresses is generated to read Missing data read request.
- the data processing request module includes: a data write request module adapted to receive a data write request, and the data requested by the data write request includes data suitable for being stored in at least The data of two cache units, and the main memory addresses of the data of each cache unit are continuous; the data processing module includes: a data write operation module, adapted to include in the main memory address information of each cache unit that satisfies the mapping relationship When all the main memory addresses are used, the data corresponding to all the main memory addresses are written at the same time.
- the data processing request module includes a data write request module
- the data processing device includes: a data write operation module, adapted to satisfy the mapping relationship of each cache unit
- the data processing request further includes a cache unit index of each cache unit and a data processing request type identifier, and the data processing request type identifier is used to identify the
- the data requested by the data processing request includes data suitable for storage in at least two cache units.
- At least one embodiment of the present disclosure further provides a cache, including a first-level cache and a second-level cache, at least two cache units of the first-level cache and at least two cache units of the second-level cache are simultaneously mapped, and the Both the first-level cache and the second-level cache include the data processing device according to the embodiments of the present disclosure.
- At least one embodiment of the present disclosure further provides a processor, the processor executes computer-executable instructions to implement the data processing method according to the embodiments of the present disclosure.
- At least one embodiment of the present disclosure further provides an electronic device, including the processor described in the embodiments of the present disclosure.
- Fig. 1 is a structural schematic diagram of a computer structure
- Fig. 2 is a schematic structural diagram of a cache unit
- FIG. 3 is a schematic structural diagram of a data processing request
- FIG. 4 is a flowchart of a data processing method provided by an embodiment of the present disclosure.
- FIG. 5 is a schematic structural diagram of a data processing request of a data processing method provided by an embodiment of the present disclosure
- Fig. 6 is another flow chart of the data processing method provided by the embodiment of the present disclosure.
- FIG. 8 is another example diagram of a data processing method provided by an embodiment of the present disclosure.
- FIG. 9 is another flow chart of the data processing method provided by the embodiment of the present disclosure.
- Fig. 10 is a block diagram of a data processing device provided by an embodiment of the present disclosure.
- FIG. 1 shows a schematic structural diagram of a computer structure.
- the computer structure mainly includes a processor 100 , a cache and a main memory 130 , and the cache includes a first-level cache 110 and a second-level cache 120 .
- the processor 100 may be a central processing unit, or a graphics processor or a general graphics processor. All the processors described in the embodiments of the present disclosure refer to processors in a functional sense, that is, an operation and control core with logic operations and control functions, rather than a package including a cache in terms of product packaging
- the processor described in the embodiments of the present disclosure may correspond to the physical core of the processor described in some documents.
- High-speed caches are divided into multi-level caches, the first-level cache 110 has the smallest storage capacity, and the second-level cache 120 is larger than the first-level cache.
- Instruction cache and data cache only the data cache can store data, which further reduces the size of the data that can be stored in the L1 cache.
- one processor 100 exclusively shares one level-1 cache 110, and multiple processors share one level-2 cache 120. In some embodiments, one processor may exclusively share one level-1 cache 110 and one level-2 cache.
- the cache 120 multiple processors share a L3 cache, and by analogy, the processors 100 realize mutual data exchange through the shared cache.
- the cache access speed is faster than the main memory, and the storage capacity is larger than the processor.
- Main memory 130 the main memory is a storage unit for storing all instructions and data, which has a large storage capacity but a slow access speed, and usually uses a DRAM (Dynamic Random Access Memory) type memory chip.
- DRAM Dynamic Random Access Memory
- the processor 100 is connected to the exclusive first-level cache 110, the first-level cache 110 is connected to the second-level cache 120, and finally connected to the main memory 130.
- the exclusive first-level cache 110 generally has a small storage space and only stores the current data of the corresponding processor. Processed Linked Data.
- the second-level cache 120 is searched for.
- the L2 cache 120 serves multiple processors, but can only serve one processor at a time, and different processors access each other and the main memory 130 through the L2 cache.
- the L2 cache 120 is still dedicated to one processor, and the L3 cache is shared.
- the processor 100 avoids direct access to the main memory 130 by accessing the cache, thereby avoiding time waste caused by huge speed differences and improving efficiency.
- FIG. 2 shows a schematic structural diagram of a cache unit.
- a cache is composed of several cache units (specifically, cache lines), and each cache unit has the same structure, and a cache line is used as an example for description below. As shown in Figure 2, there is an index in front of each cache line, such as the index shown in Figure 2: 0.
- the index also known as the offset, is the relative position of the data in the cache structure.
- a cache line mainly includes data 202 for storing data, address bits 201 , and valid bits 200 .
- the data 202 is used to store data.
- the data in the embodiments of the present disclosure refers to data information and instruction information in general.
- the data stored in each cache unit corresponds to the data of a main memory block in the main memory.
- the address bit 201 is suitable for indicating the address of the data stored in the cache unit in the main memory.
- the index of the cache unit corresponding to any main memory address can be obtained through a specific location conversion algorithm. Because the same cache unit It is mapped with multiple main memory addresses, so when the cache unit stores data, the data address of the data in the main memory must also be marked, also known as the data address of the cache unit.
- the valid bit 200 is used to indicate whether the cache line is valid. If the value of the valid bit 200 is invalid, no matter whether there is data in the current cache unit, the processor will access the next level step by step and reload the data.
- Each cache line has an index, and there are multiple mapping relationships between the cache line and the main memory. In some embodiments, it can be direct mapping.
- the cache line compares the index of the cache line according to the received address to find the exact Cache line, and then compare the address bit 201, if the address bit 201 is the same as the address information in the received request, it is a hit, and perform corresponding operations on this cache line, such as reading back data operations or writing data operations; if not the same , it is missing. For data read requests, it can only be found from the next-level cache or main memory. For data write operations, it is necessary to determine the cache line that can write data.
- Fig. 3 shows a schematic structural diagram of a data processing request.
- the data processing request mainly includes a request type identification bit 300 , other control bits 310 , and data bits 320 .
- the request type identification type bit 300 is used to indicate whether the request type is a read operation request or a write operation request.
- the read/write control bit is a two-bit bit size, and 00 is used to indicate a read operation, and 11 is used to indicate a write operation .
- Data bit 320 is suitable for identifying the data information or data main memory address information requested by the data processing request. It should be noted that if it is a read operation request, there is no data and only main memory address information. The main memory identified by data bit 320 The address information looks for the cache line in the cache and determines whether it is a hit.
- the other control bits 310 play other control functions.
- the width of the data channel between the caches at all levels and the width of the data channel between the cache and the main memory are consistent with the data bandwidth of the cache line.
- a data processing request only requests one cache unit (cache line ), to judge whether a cache unit is hit or not.
- the data transmission efficiency needs to be further improved, and therefore, the bandwidth needs to be increased, thereby improving the data transmission efficiency.
- the embodiments of the present disclosure provide a data processing method, which can simultaneously process multiple cache units (cache lines) of continuous data addresses, and can realize the improvement of data transmission efficiency by widening the bandwidth.
- Fig. 4 shows a flowchart of a data processing method provided by an embodiment of the present disclosure.
- step S400 a data processing request is received.
- the data processing request is first received. It is easy to understand that the receiving data processing request can be a cache or main memory, and when it is a cache, it can be specifically a first-level cache, a second-level cache or a third-level cache. One of the level caches.
- the data requested by the data processing request includes data suitable for being stored in at least two cache units, and the main memory addresses of the data in each cache unit are continuous.
- the data requested in the data processing request described herein includes data suitable for storage in at least two cache units, including that the requested data is data suitable for storage in two cache units, or that the requested data is Suitable for data stored in more than two cache units.
- the main memory can be divided into several main memory blocks.
- the size of each main memory block is the same as the size of the cache unit.
- the main memory address of the data is the main memory block where the data is stored, and the offset within the block.
- the data requested by the data processing request is stored in at least two cache units, and correspondingly stored in the main memory needs to be stored in at least two main memory blocks, and the main memory addresses of the data in each cache unit are continuous, not in one Two or more words or bytes in a cache unit or a main memory block are continuous, but it means that the data is located in a continuous main memory block, and the cache unit is not necessarily physically continuous, but the main memory block is continuous of.
- data 1 and data 2 For example, two data need to be processed, respectively data 1 and data 2, data 1 is the 4th word stored in the main memory block 12, data 2 is the 7th word stored in the main memory block 13, and the main memory block
- the index of the cache unit corresponding to 12 is 3, and the index of the cache line corresponding to the main memory block 13 is 8. Since the main memory block 12 and the main memory block 13 are continuous, the data processing request for requesting data 1 and data 2 is the disclosure of the data processing request. It should be noted that all the continuations mentioned in the embodiments of the present disclosure refer to the continuation manner described in this paragraph.
- the data requested by the data processing request includes data suitable for storage in at least two cache units, which can be identified by the data processing request type identifier, that is, the data processing request includes the data processing request type identifier, so that when all levels of high-speed After receiving the data processing request, the cache or main memory can determine that data stored in at least two cache units needs to be processed.
- the implementation of the identification of the data processing request type may include increasing the number of identification bits of the main memory address or increasing the type of identification of the request type.
- Fig. 5 shows a schematic structural diagram of a data processing request of a data processing method provided by an embodiment of the present disclosure. As shown in FIG. 5 , the processing request may include: main memory address quantity identification bit 1000 , request type identification type bit 1010 , other control bits 1020 , and data 1030 .
- main memory address number identification bit 1000 can be added to indicate the particularity of the request.
- the value is 1, it means that the requested data is stored in two cache units.
- the address in the data processing request still has only one main memory address, the requested data is the main memory address in the data processing request.
- data in another main memory address adjacent to the main memory address in the data processing request of course, when its value is 0, it means that the requested data is stored in a cache unit.
- the number of main memory address number identification bits 1000 can be increased as required.
- the requested data includes data suitable for storage in two cache units
- 2 bits can be used to indicate the identification of the request type
- 00 indicates an ordinary read data request, that is, the requested data
- 01 indicates a normal write data request, that is, a write request that the requested data includes data suitable for storage in one cache unit
- 10 indicates that the requested data is stored in two
- 11 represents a write request for data stored in two cache units, of course, when the requested data includes data suitable for storage in two cache units, the data processing request in The address still has only one main storage address, but the requested data is the main storage address in the data processing request and the data in another main storage address adjacent to the main storage address in the data processing request.
- the data processing request when the data processing request is a data reading request, the obtained data needs to be returned to the module that sends the data processing request.
- the data processing request may also include the cache unit index of each cache unit, so as to ensure Data processing requests the smooth return of the requested data.
- the embodiments of the present disclosure also provide an identification method for data processing requests. By analyzing the received requests, the receiving module finds that the addresses of the data requested by consecutive requests are For a continuous processing request, it returns to the superior module only when all hits.
- step S410 it is judged whether the main memory address information of each cache unit satisfying the mapping relationship includes all the main memory addresses, if yes, execute step S420, if not, execute step S430.
- the main memory address information of the cache unit includes all the main memory addresses, that is, the requested data is all backed up in the cache line, and it is determined that there is data stored in the corresponding module. At this time, it is hit, and step S420 is executed; otherwise, as long as the requested If the main memory address of a main memory block does not match, it is missing, and step S430 is executed.
- each cache unit that satisfies the mapping relationship refers to the cache unit corresponding to the main memory address in the data processing request, and the corresponding cache units are also different based on different mapping methods.
- associative mapping each cache unit that satisfies the mapping relationship refers to all cache units, and the main memory address can be mapped to all cache units; when the mapping method is group associative mapping, each cache unit that satisfies the mapping relationship refers to the main memory address A group of cache units that the address can correspond to, that is, the main memory address can only be mapped to a part of the cache units; therefore, when searching, the corresponding search ranges are different.
- the main memory block pointed to by the main memory address information stores the data requested by the processing request. If it hits, the main memory address information of the cache unit includes the address of this data, and the addresses of all the requested data include , that is, all data is stored in the cache.
- step S420 data processing is performed simultaneously on the data corresponding to each of the main memory addresses.
- the main memory address information of each cache unit that satisfies the mapping relationship includes all the main memory addresses.
- the secondary cache takes the read operation as an example.
- a data read request received by the secondary cache it is requested to read data with consecutive main memory addresses M1 and M2.
- the main memory addresses of M1 and M2 correspond to the secondary cache respectively.
- cache units S1 and S2 the second-level cache selects the S1 cache unit and the S2 cache unit at the same time after receiving the data read request, and transmits the data content of the two cache units to the first-level cache at the same time, and the first-level cache simultaneously The data of the two buffer units are received.
- the number of simultaneously effective cache units is the same as the number of cache units corresponding to the requested data in the data processing request.
- each cache level can also be divided into multiple caches, and multiple cache units are set in each cache.
- one data processing request received corresponds to at least two cache units, and each cache unit simultaneously judges whether it is a hit, and returns the data of each cache unit at the same time when it is hit.
- the width of the data channel between the caches at all levels, between the cache and the computing unit, and between the cache and the main memory can be increased, the data transmission efficiency can be improved, and the widening of the cache unit can be avoided.
- the change in the required mapping relationship and the resulting increase in workload can improve the bandwidth on the basis of a small workload, thereby improving the efficiency of data transmission; further, the ratio of the number of read and write requests In the case of inconsistency, the data transmission bandwidth ratio of the read and write requests can also be changed between adjacent two-level transmissions, so that the data transmission can better meet the usage requirements.
- step S430 write each missing main memory address into each cache unit to be processed respectively.
- main memory address information of each cache unit that satisfies the mapping relationship only includes part of the main memory address in the data processing request, or does not include it at all, it is necessary to first determine some cache units, and the determined cache units can allow missing The main memory address is written to the cache unit to be processed.
- the cache unit to be processed may be a cache unit to be replaced that is already occupied but can be replaced, or an unoccupied idle cache unit, that is, there is no main memory address information therein.
- main memory address information in the cache unit it is easy to understand that if there is main memory address information in the cache unit to be replaced, then the original main memory address information therein must be different from the respective main memory addresses corresponding to the data requested by the data processing, otherwise there will be an unexpected After all the hit data is read into the cache, the cache still does not contain all the requested data. Of course, there may also be no main memory address information, that is, an empty cache unit.
- the cache unit to be processed may be determined based on the RLU principle (least recently used principle).
- the cache unit to be processed is determined, and the missing main memory address is written to each cache unit to be processed, but the data is not written to each cache unit to be processed, so that when the corresponding data is obtained, it is written to the corresponding main memory address cache unit.
- missing main memory address mentioned herein refers to the base address stored in each cache unit in the missing address information.
- main memory address information of the cache unit to be processed is different from that of each main memory address corresponding to the data of the data processing request, but after each replacement, the Check whether the cache contains all the requested data, if not, repeat this process until it is detected that all the requested data are backed up in the cache.
- FIG. 6 shows another flow of the data processing method provided by the embodiment of the present disclosure picture.
- FIG. 6 is a flow chart in the specific case where the data processing request is a data read request. Since most of the steps in the figure are similar to the steps in FIG. 4 , the description of this part will not be expanded. As shown in Fig. 6, the process may include the following steps.
- step S500 a data read request is received.
- step S500 For the specific content of step S500, please refer to the description of step S400 shown in FIG. 4 .
- the data processing request is a data reading request, and other content will not be repeated here.
- step S510 it is judged whether the main memory address information of each cache unit satisfying the mapping relationship includes all the main memory addresses, if yes, execute step S520, if not, execute step S530.
- step S510 For the specific content of step S510, please refer to the description of step S410 shown in FIG. 4 , which will not be repeated here.
- step S520 the data corresponding to all the main memory addresses are returned at the same time.
- step S520 For the specific content of step S520, please refer to the description of step S420 shown in FIG. 4 .
- the data processing request is a data read request, when the data is processed, the data corresponding to all main memory addresses are returned at the same time.
- step S530 the missing main memory addresses are respectively Write to each pending cache location.
- step S530 For the specific content of step S530, please refer to the description of step S430 shown in FIG. 4 , which will not be repeated here.
- step S540 a missing data read request for reading data corresponding to each of the missing main memory addresses is generated and sent.
- the missing main memory address is multiple addresses, there will be many situations depending on the continuity and discontinuity, for example, the missing data is also continuous data, the missing data is partly continuous and partly discontinuous, or the missing data is all discontinuous.
- a new merged data read request may be issued for data corresponding to several missing continuous main memory addresses.
- a data read request is sent separately for the data of each cache unit. Different processing methods affect subsequent processing, and two different implementation examples are given below.
- Fig. 7 shows an example diagram of the implementation of the data processing method provided by the embodiment of the present disclosure.
- FIG. 7 is similar to the steps in FIG. 6 , and this part of the content will not be described. As shown in Fig. 7, this example may include the following steps.
- step S600 a data read request is received.
- step S610 it is judged whether the main memory address information of each cache unit satisfying the mapping relationship contains all the main memory addresses, if yes, execute step S620, and if not, execute step S630.
- step S620 the data corresponding to all the main memory addresses are returned at the same time.
- step S630 the missing addresses of the main memory are respectively written into the cache units to be processed.
- step S600-step S630 For the specific content of step S600-step S630, please refer to the description of step S400-step S430 shown in FIG. 4, which will not be repeated here.
- step S640 each missing data read request for reading the missing data corresponding to each of the main memory addresses is generated.
- a missing data read request is generated for the data corresponding to the main memory address of each cache unit, and sent to the next-level module one by one.
- Each data request only requests to read the missing data corresponding to the main memory address of a cache unit, the number of missing data read requests is the same as the number of the missing main memory addresses, and the next-level module can be a cache Or main memory, if the sending module is the lowest level cache, then the next level module is main memory. If the sending module is a first-level cache, the next-level module is a second-level cache.
- step S650 it is judged whether all the data corresponding to the missing main memory addresses are received, if so, execute step S660, otherwise, wait and judge again.
- next-level module is a lower-level cache
- the request will be sent to the next-level module. It is easy to understand that the specific The process is consistent with the above process, and will not be repeated here.
- step S660 data corresponding to all main memory addresses are returned at the same time.
- step S660 For the specific content of step S660, please refer to the description of step S520 shown in FIG. 5 , which will not be repeated here.
- Fig. 8 shows another example diagram of the data processing method provided by the embodiment of the present disclosure. It should be noted that the difference between the example in FIG. 8 and that in FIG. 7 mainly lies in the processing of missing data, so other steps will not be described. As shown in Fig. 8, this example may include the following steps.
- step S700 a data read request is received.
- step S710 it is determined whether the main memory address information of each cache unit satisfying the mapping relationship includes all the main memory addresses. If yes, execute step S720, and if no, execute step S730.
- step S720 the data corresponding to all the main memory addresses are returned at the same time.
- step S730 the missing addresses of the main memory are respectively written into the cache units to be processed.
- step S700-step S730 For the specific content of step S700-step S730, please refer to the description of the corresponding part above, and will not repeat them here.
- step S740 acquire the main memory addresses with continuous addresses among the missing main memory addresses, obtain each continuous main memory address, and generate each missing continuous main memory address according to each of the continuous main memory addresses Consecutive missing data read requests for the corresponding data.
- missing main memory addresses may be further obtained, and then new continuous main memory addresses are determined, and continuous missing data read requests are generated according to the continuous main memory addresses.
- the main memory addresses of the missing data are M1, M2 and N1 respectively, where the main memory address M1 and the main memory address M2 are continuous, then a continuous missing data read request is generated, and the main memory address N1 is not continuous with M1 and M2, then , generate missing data read requests, and then send consecutive missing data read requests corresponding to main memory addresses M1 and M2 and missing data read requests corresponding to main memory address N1 to the next-level module.
- step S750 it is judged whether all the data corresponding to the missing main memory addresses have been received, if yes, step S760 is executed, and if not, the judgment is repeated.
- step S760 data corresponding to all main memory addresses are returned at the same time.
- the generated missing data read request may still include reading the missing continuous main memory The data corresponding to the address is requested, thereby further improving the data transmission efficiency during the data reading process.
- FIG. 9 shows Another flowchart of the data processing method provided by the embodiment of the present disclosure is shown.
- the process may include the following processes.
- Step S800 receiving a data writing request.
- step S800 For the specific content of step S800, please refer to the description of step S400 shown in FIG. 4 .
- the data processing request is a data write request, and other content will not be repeated here.
- Step S810 judging whether the main memory address information of each cache unit satisfying the mapping relationship includes all the main memory addresses. If yes, execute step S820, and if no, execute step S830.
- step S810 For the specific content of step S810, please refer to the description of step S410 shown in FIG. 4 , which will not be repeated here.
- Step S820 receiving data corresponding to each of the main memory addresses at the same time, and writing each of the data into the corresponding cache unit.
- step S820 For the specific content of step S820, please refer to the description of step S420 shown in Figure 4.
- the data processing request is a data write request
- the data is simultaneously written into the cache corresponding to the main memory address unit or main memory.
- a received data write request corresponds to at least two cache units, and each cache unit simultaneously judges whether it is a hit, and when it hits, simultaneously writes to each cache unit
- the data can increase the data channel width between caches at all levels, between caches and computing units, and between caches and main memory without changing the size of the cache unit, and improve data transmission efficiency, thereby avoiding the need for widening the cache unit.
- the change of the required mapping relationship and the resulting increase in workload can improve the bandwidth and improve the efficiency of data transmission on the basis of a small workload.
- the writing of data corresponding to each main memory address can be realized, and since the data can be written at the same time, the efficiency of data writing can be improved.
- step S830 may also be included, writing each missing main memory address into each cache unit to be processed.
- step S830 For the specific content of step S830, please refer to the description of step S430 shown in FIG. Can.
- Step S840 receiving data corresponding to each of the main memory addresses at the same time, and writing each of the data into the corresponding cache unit.
- a received data write request corresponds to at least two cache units, and each cache unit simultaneously judges whether it is a hit, and when it misses, writes all missing Write the above main memory address into each cache unit to be processed, and then write the data of each cache unit at the same time, which can increase the size of caches at all levels, between caches and computing units, and between caches and main memory without changing the size of cache units.
- the width of the data channel between the memory can improve the efficiency of data transmission, so as to avoid the change of the mapping relationship required to widen the cache unit, and the increase in workload caused by it, and improve the bandwidth on the basis of a small workload. , thereby improving the efficiency of data transmission.
- Embodiments of the present disclosure also provide a data processing device, which can be regarded as a functional module required to implement the data processing method provided by the embodiments of the present disclosure.
- the content of the device described herein and the content of the method described above refer to each other.
- the data processing request module 910 is adapted to receive a data processing request, the data requested by the data processing request includes data suitable for being stored in at least two cache units, and the main memory addresses of the data in each cache unit are continuous.
- the data processing module 920 is adapted to simultaneously perform data processing on the data corresponding to each of the main memory addresses when the main memory address information of each cache unit satisfying the mapping relationship includes all the main memory addresses, wherein the satisfying Each cache unit in the mapping relationship refers to the cache unit corresponding to the main memory address in the data processing request.
- the data processing module 920 is further adapted to, when the main memory address information of each cache unit that satisfies the mapping relationship does not include all the main memory addresses, write each missing main memory address into each The cache units to be processed, wherein the original main storage address information of each cache unit to be processed is different from the main storage address of the data requested by the data processing request.
- the data processing module 920 may be a data read operation module, adapted to return data corresponding to all the main memory addresses at the same time when the main memory address information of each cache unit satisfying the mapping relationship includes all the main memory addresses.
- the data read operation module is further adapted to write each missing main memory address when the main memory address information of each cache unit satisfying the mapping relationship does not include all the main memory addresses. input each cache unit to be processed, and generate and send a missing data read request for reading the missing data corresponding to each of the main memory addresses.
- the data read operation module is adapted to generate and send a missing data read request for reading data corresponding to each of the missing main memory addresses, including: generating and respectively reading each of the missing main memory addresses For each missing data read request of the data corresponding to the address, the number of the missing data read requests is the same as the number of the missing main memory addresses.
- the data read operation module is further adapted to return the data corresponding to all the main memory addresses simultaneously when all the data corresponding to the missing main memory addresses are received.
- the data read operation module is adapted to generate and send a missing data read request for reading data corresponding to each of the missing main memory addresses, including: obtaining addresses in each of the missing main memory addresses Continuous main memory addresses, each continuous main memory address is obtained; according to each of the continuous main memory addresses, a continuous missing data read request for reading data corresponding to each of the missing continuous main memory addresses is generated; a data read operation module It is also suitable for returning the data corresponding to all the main memory addresses at the same time when all the data corresponding to the missing main memory addresses are received.
- the data processing request module 910 may be a data write request module, adapted to receive a data write request, the data requested by the data write request includes data suitable for being stored in at least two cache units, and each The main storage addresses of the data in the cache unit are continuous.
- the data processing module 920 may be a data writing module, adapted to simultaneously write data corresponding to all the main memory addresses when the main memory address information of each cache unit satisfying the mapping relationship includes all the main memory addresses.
- the data writing operation module is further adapted to write the missing main memory addresses respectively when the main memory address information of each cache unit that satisfies the mapping relationship does not include all the main memory addresses. After entering each to-be-processed cache unit, simultaneously receive the data corresponding to each of the main memory addresses, and write each of the data into the corresponding cache unit.
- one data processing request received by the data processing request module 910 corresponds to at least two cache units, and the data processing module 920 simultaneously judges whether each cache unit is a hit, And when it hits, return the data of each cache unit at the same time, without changing the size of the cache unit, the width of the data channel between the caches at all levels, between the cache and the computing unit, and between the cache and the main memory can be increased.
- At least one embodiment of the present disclosure also provides a cache, which may include a first-level cache and a second-level cache, and at least two cache units of the first-level cache and at least two cache units of the second-level cache are simultaneously Mapping, and both the first-level cache and the second-level cache include the data processing device provided by the embodiments of the present disclosure.
- the data of at least two cache units can be transmitted simultaneously, and without changing the size of the cache unit, Increase the data channel width between caches at all levels to improve data transmission efficiency.
- At least one embodiment of the present disclosure further provides a processor, where the processor executes computer-executable instructions to implement the data processing method provided by the embodiments of the present disclosure.
- At least one embodiment of the present disclosure further provides an electronic device, and the electronic device may include the processor provided above in the embodiments of the present disclosure.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Memory System Of A Hierarchy Structure (AREA)
Abstract
一种数据处理方法、装置、缓存器、处理器及电子设备。该数据处理方法包括接收数据处理请求,所述数据处理请求所请求的数据包括适于存储于至少两个缓存单元的数据,且各个所述缓存单元的数据的主存地址连续;当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,对各个所述主存地址对应的数据同时进行数据处理。该方法可以提高带宽,进而提高数据传输效率。
Description
本申请要求于2021年11月17日递交的中国专利申请第202111363242.3号的优先权,在此全文引用上述中国专利申请公开的内容以作为本申请的一部分。
本公开的实施例涉及一种数据处理方法、装置、缓存器、处理器及电子设备。
在典型计算机结构中,通常包含处理器和多级存储器。多级存储器中,主存储器存储指令和数据。处理器从主存储器中获取指令和相应的数据,执行指令,并将结果数据写回到主存储器中。此外,在处理器和主存储器之间,通常有一级或多级高速缓冲存储器(Cache)。高速缓冲存储器用于降低处理器读取指令和数据的时间。当处理器需要读取主存储器中某个地址数据时,它首先检查该数据是否存在于高速缓存器中。如果高速缓存器中包含该数据,则处理器直接从高速缓存器中读取该数据,从高速缓存器中读取数据要远快于从主存储器中读取数据。否则的话,高速缓存器从主存储器中读取该数据,存放在高速缓存器中,并返回给主存储器。
当处理器运算速度提高,对数据需求量加大时,需要提升数据获取效率,cache间的数据通道宽度乘以单位时间通道可以传递数据的次数,即为带宽,其表明了cache在单位时间可以获取的最大数据量,为了提升数据获取效率就需要提高带宽。
因此,如何提高带宽,进而提高数据传输效率,就成为本领域技术人员需要解决的技术问题。
发明内容
本公开至少一实施例提供一种数据处理方法,包括:接收数据处理请求,所述数据处理请求所请求的数据包括适于存储于至少两个缓存单元的 数据,且各个所述缓存单元的数据的主存地址连续;当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,对各个所述主存地址对应的数据同时进行数据处理,其中,所述满足映射关系的各个缓存单元,是指所述数据处理请求中的主存地址所对应的缓存单元。
例如,本公开一实施例提供的数据处理方法还包括:当满足映射关系的各个缓存单元的主存地址信息中未包括全部所述主存地址时,将缺失的各个所述主存地址分别写入各个待处理缓存单元,其中,各个所述待处理缓存单元的原有的主存地址信息,均异于所述数据处理请求的数据所对应的各个所述主存地址。
例如,在本公开一实施例提供的数据处理方法中,所述待处理缓存单元包括待替换缓存单元和空闲缓存单元中的至少一者。
例如,在本公开一实施例提供的数据处理方法中,所述待替换缓存单元通过最近最少使用原则RLU确定。
例如,在本公开一实施例提供的数据处理方法中,所述数据处理请求包括数据读取请求,所述当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,对各个所述主存地址对应的数据同时进行数据处理的步骤,包括:当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,同时返回全部所述主存地址对应的数据。
例如,在本公开一实施例提供的数据处理方法中,所述数据处理请求包括数据读取请求,所述当满足映射关系的各个缓存单元的主存地址信息中未包括全部所述主存地址时,将缺失的各个所述主存地址分别写入各个待处理缓存单元的步骤之后,还包括:生成并发送读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求。
例如,在本公开一实施例提供的数据处理方法中,所述生成读取缺失的各个所述主存地址所对应的数据的新数据处理请求的步骤包括:生成分别读取缺失的各个所述主存地址所对应的数据的各个缺失数据读取请求,所述缺失数据读取请求的数量与所述缺失的各个所述主存地址的数量相同;所述数据处理方法还包括:当接收到全部与缺失的各个所述主存地址对应的数据时,同时返回全部主存地址对应的数据。
例如,在本公开一实施例提供的数据处理方法中,所述生成读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求的步骤包括:获取 缺失的各个所述主存地址中地址连续的主存地址,得到各个连续主存地址;根据各个所述连续主存地址,生成读取缺失的各个所述连续主存地址所对应的数据的连续缺失数据读取请求;所述数据处理方法还包括:当接收到全部与缺失的各个所述主存地址对应的数据时,同时返回全部主存地址对应的数据。
例如,在本公开一实施例提供的数据处理方法中,所述数据处理请求包括数据写入请求,所述当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,对各个所述主存地址对应的数据同时进行数据处理的步骤,包括:当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,同时接收各个所述主存地址所对应的数据,并将各个所述数据写入相对应的所述缓存单元。
例如,在本公开一实施例提供的数据处理方法中,所述数据处理请求包括数据写入请求,所述当满足映射关系的各个缓存单元的主存地址信息中未包括全部所述主存地址时,将缺失的各个所述主存地址分别写入各个待处理缓存单元的步骤之后,还包括:同时接收各个所述主存地址所对应的数据,并将各个所述数据写入相对应的所述缓存单元。
例如,在本公开一实施例提供的数据处理方法中,所述数据处理请求还包括数据处理请求类型标识,所述数据处理请求类型标识用于标识所述数据处理请求所请求的数据包括适于存储于至少两个缓存单元的数据。
例如,在本公开一实施例提供的数据处理方法中,所述数据处理请求类型标识的实现方式包括增加主存地址数量标识位或增加请求类型标识种类。
本公开至少一实施例还提供一种数据处理装置,包括:数据处理请求模块,适于接收数据处理请求,所述数据处理请求所请求的数据包括适于存储于至少两个缓存单元的数据,且各个所述缓存单元的数据的主存地址连续;数据处理模块,适于当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,对各个所述主存地址对应的数据同时进行数据处理,其中,所述满足映射关系的各个缓存单元,是指所述数据处理请求中的主存地址所对应的缓存单元。
例如,在本公开一实施例提供的数据处理装置中,数据处理模块,还适于当满足映射关系的各个缓存单元的主存地址信息中未包括全部所述主 存地址时,将缺失的各个所述主存地址分别写入各个待处理缓存单元,其中,各个所述待处理缓存单元的原有的主存地址信息,均异于所述数据处理请求所请求数据的主存地址。
例如,在本公开一实施例提供的数据处理装置中,所述数据处理请求模块包括:数据读请求模块,适于接收数据读请求,所述数据读请求所请求的数据包括适于存储于至少两个缓存单元的数据,且各个所述缓存单元的数据的主存地址连续;所述数据处理模块包括:数据读操作模块,适于当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,同时返回全部所述主存地址对应的数据。
例如,在本公开一实施例提供的数据处理装置中,所述数据处理请求模块包括:数据读请求模块,适于接收数据读请求,所述数据读请求所请求的数据包括适于存储于至少两个缓存单元的数据,且各个所述缓存单元的数据的主存地址连续;所述数据处理模块包括:数据读操作模块,适于当满足映射关系的各个缓存单元的主存地址信息中未包括全部所述主存地址时,将缺失的各个所述主存地址分别写入各个待处理缓存单元,生成并发送读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求。
例如,在本公开一实施例提供的数据处理装置中,所述数据读操作模块,适于生成并发送读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求,包括:生成分别读取缺失的各个所述主存地址所对应的数据的各个缺失数据读取请求,所述缺失数据读取请求的数量与所述缺失的各个所述主存地址的数量相同;所述数据读操作模块,还适于当接收到全部与缺失的各个所述主存地址对应的数据时,同时返回全部主存地址对应的数据。
例如,在本公开一实施例提供的数据处理装置中,所述数据读操作模块,适于生成并发送读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求包括:获取缺失的各个所述主存地址中地址连续的主存地址,得到各个连续主存地址;根据各个所述连续主存地址,生成读取缺失的各个所述连续主存地址所对应的数据的连续缺失数据读取请求。
例如,在本公开一实施例提供的数据处理装置中,所述数据处理请求模块包括:数据写请求模块,适于接收数据写请求,所述数据写请求所请求的数据包括适于存储于至少两个缓存单元的数据,且各个所述缓存单元 的数据的主存地址连续;所述数据处理模块包括:数据写操作模块,适于当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,同时写入全部所述主存地址对应的数据。
例如,在本公开一实施例提供的数据处理装置中,所述数据处理请求模块包括数据写请求模块,所述数据处理装置包括:数据写操作模块,适于当满足映射关系的各个缓存单元的主存地址信息中未包括全部所述主存地址时,将缺失的各个所述主存地址分别写入各个待处理缓存单元后,同时接收各个所述主存地址所对应的数据,并将各个所述数据写入相对应的所述缓存单元。
例如,在本公开一实施例提供的数据处理装置中,所述数据处理请求还包括各个所述缓存单元的缓存单元索引和数据处理请求类型标识,所述数据处理请求类型标识用于标识所述数据处理请求所请求的数据包括适于存储于至少两个缓存单元的数据。
例如,在本公开一实施例提供的数据处理装置中,所述数据处理请求类型标识的实现方式包括增加主存地址数量标识位或增加请求类型标识种类。
本公开至少一实施例还提供一种缓存器,包括一级缓存和二级缓存,所述一级缓存的至少两个缓存单元和所述二级缓存的至少两个缓存单元同时映射,且所述一级缓存和所述二级缓存均包括如本公开的实施例所述的数据处理装置。
本公开至少一实施例还提供一种处理器,所述处理器执行计算机可执行指令,以实现如本公开的实施例所述的数据处理方法。
本公开至少一实施例还提供一种电子设备,包括如本公开的实施例所述的处理器。
为了更清楚地说明本公开的实施例,下面将对实施例中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本公开的的实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据提供的附图获得其他的附图。
图1为一种计算机结构的结构示意图;
图2为一种缓存单元的结构示意图;
图3为一种数据处理请求的结构示意图;
图4为本公开的实施例所提供的数据处理方法的一流程图;
图5为本公开的实施例所提供的数据处理方法的数据处理请求的结构示意图;
图6为本公开的实施例所提供的数据处理方法的另一流程图;
图7为本公开的实施例所提供的数据处理方法的实现示例图;
图8为本公开的实施例所提供的数据处理方法的另一示例图;
图9为本公开的实施例所提供的数据处理方法的再一流程图;以及
图10为本公开的实施例所提供的数据处理装置的框图。
下面将结合本公开的实施例中的附图,对本公开的实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本公开的一部分实施例,而不是全部的实施例。基于本公开的中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本公开的保护的范围。
下面对一种缓存技术进行介绍。图1示出了一种计算机结构的结构示意图。如图1所示,该计算机结构主要包括处理器100、高速缓存和主存130,高速缓存包括一级高速缓存110和二级高速缓存120。
其中,处理器100,可以是中央处理器,也可以是图形处理器或通用图形处理器。本公开的实施例所有所述的处理器都是指功能意义上的处理器,即具有逻辑运算和控制功能的运算和控制核心,而不是从产品封装上看的包含高速缓存在内的一个封装盒,本公开的实施例所述的处理器可能对应某些文献里所描述的处理器的物理核心。
高速缓存,分为多级高速缓存,一级高速缓存110存储量最小,二级高速缓存120比一级大,后面以此类推,在极小存储量的一级高速缓存110中有时还会区分成指令高速缓存和数据高速缓存,只有数据高速缓存可以存储数据,这进一步减小了一级高速缓存可以存储的数据大小。一般一个处理器100独享一个一级高速缓存110,多个处理器共享一个二级高速缓存120,在一些实施例中,也可以是一个处理器独享一个一级高速缓存110 和一个二级高速缓存120,多个处理器共享一个三级高速缓存,后面以此类推,处理器100通过共享高速缓存实现了相互间的数据交流。高速缓存存取速度比主存快,存储容量比处理器大。
主存130,主存是存储所有指令和数据的存储单元,存储容量大但存取速度慢,通常采用DRAM(动态随机存取存储器)类型的存储芯片。
处理器100后连接专属的一级高速缓存110,一级高速缓存110连接二级高速缓存120,最后连接主存130,专属的一级高速缓存110一般存储空间较小,只保存对应处理器当下处理的关联数据。当一级高速缓存110中没有命中时才会去二级高速缓存120中寻找。例如在一些实施例中,二级高速缓存120为多个处理器服务,但一次只能为一个处理器提供服务,不同的处理器通过二级高速缓存实现互相访问以及与主存130的访问。在另一些实施例中二级高速缓存120仍为一个处理器所专属,共享的是三级高速缓存。处理器100通过对高速缓存的访问避免了对主存130的直接访问,从而避免了巨大速度差异带来的时间浪费,提高了效率。
图2示出了一种缓存单元的结构示意图。高速缓存由数个缓存单元(具体可以为缓存行)构成,每个缓存单元的结构相同,下面以一个缓存行为例进行说明。如图2所示,每个缓存行的前面有索引,比如图2中所示索引:0。
索引,也称为偏移量,即是数据在缓存结构中的相对位置。
一个缓存行主要包含存储数据的数据202、地址位201、有效位200。
其中,数据202用于存储数据,本公开的实施例所述的数据指的是数据信息、指令信息的统称,每个缓存单元所存储的数据对应主存中的一个主存块的数据。
地址位201,适于表示存储于缓存单元中的数据在主存中的地址,由于计算机系统中,通过特定的位置转换算法可以获取任意主存地址对应的缓存单元的索引,由于同一个缓存单元与多个主存地址映射,因此在缓存单元存储数据时还必须标记数据在主存中的数据地址,也称为缓存单元的数据地址。
有效位200,用于表明该缓存行是否有效,如果有效位200的值所表示的意义为无效,无论当前缓存单元中是否有数据,处理器都会逐级向下一级访问,重新加载数据。
每个缓存行都有一个索引,缓存行与主存间有多种映射关系,在某些实施例中,可以为直接映射,缓存行根据接收到的地址,对比缓存行的索引,找到确切的缓存行,然后对比地址位201,如果地址位201和所接收的请求中的地址信息相同,则为命中,对此缓存行进行相应操作,比如读回数据操作或者写入数据操作;如果不相同,则为缺失,对于数据读取请求,只能从下一级高速缓存或者主存中寻找,对于数据写入操作,则需要确定可以写入数据的缓存行。
图3示出了一种数据处理请求的结构示意图。如图3所示,数据处理请求主要包括请求类型标识种类位300、其它控制位310、数据位320。
其中,请求类型标识种类位300,用于标明请求类型是读操作请求还是写操作请求,在一些实施例中,读写控制位为两位bit大小,用00表示读操作,用11表示写操作。
数据位320,适于标识数据处理请求所请求的数据信息或数据主存地址信息,需要注意的是,如果是读操作请求则没有数据只有主存地址信息,通过数据位320所标识的主存地址信息在缓存中寻找缓存行,并判断是否命中。
其它控制位310,起到其余控制作用。
各级缓存之间的数据通道宽度、缓存与主存之间的数据通道宽度均与缓存行的数据带宽一致,在对数据进行处理的过程中,一个数据处理请求只请求一个缓存单元(缓存行),对一个缓存单元进行是否命中的判断。
然而,某些情况下,还需要进一步提高数据的传输效率,为此,需要增加带宽,进而提高数据传输效率。
为此,本公开的实施例提供一种数据处理方法,可以对连续数据地址的多个缓存单元(缓存行)同时进行处理,可以实现通过加宽带宽,提高数据传输效率。
图4示出了本公开的实施例所提供的数据处理方法的一流程图。
如图4所示,该流程可以包括如下步骤:在步骤S400中,接收数据处理请求。
在数据处理时,首先接收数据处理请求,容易理解的是,接收数据处理请求的可以为高速缓存或者主存,而当为高速缓存时,具体可以为一级高速缓存、二级高速缓存或者三级高速缓存中的一者。
需要说明的是,在接收数据处理请求的过程中,所述数据处理请求所请求的数据包括适于存储于至少两个缓存单元的数据,且各个所述缓存单元的数据的主存地址连续。
本文所述的数据处理请求所请求的数据包括适于存储于至少两个缓存单元的数据,既包括所请求的数据为适于存储于两个缓存单元的数据,也可以包括所请求的数据为适于存储于多于两个缓存单元的数据。
如前所述,主存可以分为数个主存块,每个主存块的大小与缓存单元大小相同,数据的主存地址即为存储数据的主存块,以及块内偏移。数据处理请求所请求的数据存储于至少两个缓存单元,那么对应存储于主存中需存储于至少两个主存块,而各个所述缓存单元的数据的主存地址连续,不是指在一个缓存单元中或一个主存块中的两个或多个字或字节连续,而是指数据位于连续的主存块中,且缓存单元在物理上不一定连续的,但主存块是连续的。
例如,需要处理两个数据,分别为数据1和数据2,数据1是存储于主存块12的第4个字,数据2是存储于主存块13的第7个字,与主存块12对应的缓存单元索引是3,与主存块13对应的缓存行索引是8,由于主存块12和主存块13连续,那么请求数据1和数据2的数据处理请求,即为本公开的所述的数据处理请求。需要注意的是,本公开的实施例提到的所有连续,表达的都是指本段所描述的连续方式。
具体地,数据处理请求所请求的数据包括适于存储于至少两个缓存单元的数据,可以通过数据处理请求类型标识进行标识,即数据处理请求中包括数据处理请求类型标识,从而当各级高速缓存或者主存接收到数据处理请求后,能够确定需要对存储于至少两个缓存单元的数据进行处理。
在一种具体实施方式中,数据处理请求类型标识的实现方式可以包括增加主存地址数量标识位或增加请求类型标识种类。图5示出了本公开的实施例所提供的数据处理方法的数据处理请求的结构示意图。如图5所示,该处理请求可以包括:主存地址数量标识位1000,请求类型标识种类位1010,其它控制位1020,数据1030。
对于增加主存地址数量标识位1000,比如:如果所请求的数据包括适于存储于两个缓存单元的数据,可以通过增加一个主存地址数量标识位1000以表示该请求的特殊性,当其取值为1时,表示所请求的数据是存储 于两个缓存单元的数据,尽管数据处理请求中的地址仍然仅有一个主存地址,但所请求的数据为数据处理请求中的主存地址以及数据处理请求中的主存地址相邻的另一个主存地址中的数据;当然,当其取值为0时,表示所请求的数据是存储于一个缓存单元的数据。
容易理解的是,如果所请求的数据包括适于存储于多于两个缓存单元的数据,那么主存地址数量标识位1000的位数可以根据需要增加。
对于增加请求类型标识种类,比如:如果所请求的数据包括适于存储于两个缓存单元的数据,可以通过2比特来表示请求类型的标识,00表示普通的读数据请求,即所请求的数据包括适于存储于一个缓存单元的数据的读请求,01表示普通的写数据请求,即所请求的数据包括适于存储于一个缓存单元的数据的写请求,10表示所请求的数据存储于两个缓存单元的数据的读请求,11表示所请求的数据存储于两个缓存单元的数据的写请求,当然,当所请求的数据包括适于存储于两个缓存单元的数据,数据处理请求中的地址仍然仅有一个主存地址,但所请求的数据为数据处理请求中的主存地址以及数据处理请求中的主存地址相邻的另一个主存地址中的数据。
进一步地,当数据处理请求为数据读取请求时,所得到的数据需要返回至发送数据处理请求的模块,为此,数据处理请求中还可以包括各个所述缓存单元的缓存单元索引,从而保证数据处理请求所请求数据的顺利返回。
值得注意的是,本公开的实施例还提供一种数据处理请求的标识方式,接收模块通过对接收到的请求进行分析,发现连续的数个请求所请求的数据的地址是连续的,便视为一个连续的处理请求,只有在全部命中的时候才返回给上级模块。
在步骤S410中,判断满足映射关系的各个缓存单元的主存地址信息中是否包括全部所述主存地址,如果是,则执行步骤S420,如果否,则执行步骤S430。
当接收到数据处理请求后,需要基于数据处理请求处理数据,为此需要确定数据处理请求所对应的数据在接收请求的模块(各级高速缓存或主存)中存储,如果满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址,即所请求的数据全部备份在缓存行中,确定对应模块中 存储有数据,此时命中,执行步骤S420,否则,只要请求中的一个主存块的主存地址未命中,则为缺失,执行步骤S430。
容易理解的是,满足映射关系的各个缓存单元,是指所述数据处理请求中的主存地址所对应的缓存单元,基于映射方式的不同,所对应的缓存单元也不同,当映射方式为全关联映射时,满足映射关系的各个缓存单元是指全部的缓存单元,主存地址能够和全部的缓存单元进行映射;当映射方式为组关联映射时,满足映射关系的各个缓存单元是指主存地址所能够对应的一组缓存单元,即主存地址能够仅能和一部分缓存单元映射;从而,在进行查找时,所对应的查找范围不同。
主存地址信息指向的主存块中储存有处理请求所请求的数据,如果命中,那么所述的缓存单元的主存地址信息中包括了此数据的地址,所请求的全部数据的地址都包括,即所有的数据都储存在了缓存中。
在步骤S420中,对各个所述主存地址对应的数据同时进行数据处理。
满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址,尽管数据处理请求所请求的数据涉及多个缓存单元,但仍同时对多个缓存单元的数据同时进行处理。
以读操作为例,例如二级高速缓存接收到的一个数据读取请求中请求读取M1、M2这两个主存地址连续的数据,M1和M2的主存地址分别对应着二级高速缓存的缓存单元S1和S2,则二级高速缓存接收到数据读取请求后同时选中S1缓存单元和S2缓存单元,同时传输该两个缓存单元的数据内容给一级高速缓存,一级高速缓存同时接收到该两个缓存单元的数据。
容易理解的是,为了实现对各个所述主存地址对应的数据同时进行数据处理,在结构上,发送数据处理请求和接收数据处理请求的两级模块之间,至少两个缓存单元同时映射,即至少两个缓存单元的映射同时有效。
当数据处理请求所请求的数据多于两个缓存单元的数据时,映射同时有效的缓存单元的数量与数据处理请求所述请求的数据对应的缓存单元的数量相同。
在一些实施例中,还可以将每级高速缓存分为多个高速缓存器,每个高速缓存器中设置多个缓存单元。
可见,本公开的实施例所提供的数据处理方法,所接收的一个数据处 理请求至少对应两个缓存单元,各个缓存单元同时进行是否命中的判断,并在命中时,同时返回各个缓存单元的数据,可以在无需改变缓存单元大小的情况下,增大各级高速缓存间、高速缓存与计算单元间以及高速缓存与主存间的数据通道宽度,提高数据传输效率,并且可以避免加宽缓存单元所需要的映射关系改变,以及由此带来的工作量的增加,可以实现在较小工作量的基础上,提高带宽,进而提高数据传输效率的提高;进一步地,在读写请求的数量比例不一致的情况下,还可以在相邻两级传输之间改变读写请求的数据传输带宽比例,使得数据传输更满足使用要求。
例如,在另一实施例中,为了保证数据的顺利处理,还可以包括如下步骤:在步骤S430中,将缺失的各个所述主存地址分别写入各个待处理缓存单元。
满足映射关系的各个缓存单元的主存地址信息中仅包括部分数据处理请求中的所述主存地址,或者完全未包括时,需要首先确定一些缓存单元,所确定的缓存单元能够允许将缺失的主存地址写入到待处理缓存单元。
当然,待处理缓存单元既可以为已经被占用但可以被替换的待替换缓存单元,也可以为未被占用的空闲缓存单元,即其中不存在主存地址信息。
容易理解的是,待替换缓存单元中如果存在主存地址信息,那么其中的原有的主存地址信息需与数据处理请求的数据所对应的各个所述主存地址不同,否则会出现将未命中的数据全部读入缓存后,缓存中仍然没有包含全部所请求数据的问题。当然,其中也可以不存在主存地址信息,即为空的缓存单元。
具体地,待处理缓存单元可以基于RLU原则(最近最少使用原则)进行确定。
确定了待处理缓存单元,将缺失的主存地址写入各个待处理缓存单元,但并非将数据写入各个待处理缓存单元,以便当获取到对应的数据时,将其写入对应主存地址的缓存单元。
另外,容易理解的是,本文所述的缺失的主存地址是指缺失的地址信息中的各个缓存单元中存储的基地址。
在一种替代实现中,对待处理缓存单元,没有其主存地址信息异于所述数据处理请求的数据所对应的各个所述主存地址的要求,但在每次替换完后,都要再进行缓存中是否包含所有所请求数据的检查,如果没有,就 重复此过程,直到检测到所有所述请求数据都包含备份到了缓存中。
这样,通过将缺失的各个所述主存地址分别写入各个待处理缓存单元,可以在提高数据传输效率的基础上,保证数据读取或者写入的顺利执行。
为了方便理解,本公开的实施例还提供了一种数据处理方法,以说明在数据读取时的具体处理方法,图6示出了本公开的实施例所提供的数据处理方法的另一流程图。
容易理解的是,图6所示流程是数据处理请求为数据读取请求这一具体情况时的流程图。由于图中大部分步骤都与图4中的步骤相类似,对这部分内容不展开描述。如图6所示,该流程可以包括如下步骤。
在步骤S500中,接收数据读取请求。
步骤S500的具体内容请参考图4所示的步骤S400的描述,当然,数据处理请求为数据读取请求,其他内容,在此不再赘述。
在步骤S510中,判断满足映射关系的各个缓存单元的主存地址信息中是否包括全部所述主存地址,若是,执行步骤S520,若否,执行步骤S530。
步骤S510的具体内容,请参考图4所示的步骤S410的描述,在此不再赘述。
在步骤S520中,同时返回全部所述主存地址所对应的数据。
步骤S520的具体内容,请参考图4所示的步骤S420的描述,当然,由于数据处理请求为数据读取请求,因此,在对数据进行处理时,同时返回全部主存地址对应的数据。
这样,可以实现对于各个主存地址对应数据的读取,并且,由于可以同时返回数据,可以提高数据读取的效率。
在另一种具体实施方式中,当存在缺失的各个所述主存地址时,为了保证数据的顺利读取,还可以包括以下步骤:在步骤S530中,将缺失的各个所述主存地址分别写入各个待处理缓存单元。
步骤S530的具体内容,请参考图4所示的步骤S430的描述,在此不再赘述。
在步骤S540中,生成并发送读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求。
当存在缺失的各个所述主存地址时,为了实现数据的读取还需要生成发送读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求,并向下一级模块(比如:二级缓存或者主存)发送缺失数据读取请求。
可以看出,对于数据读取请求,在满足映射关系的各个缓存单元的主存地址信息中不包括全部所述主存地址时,还生成并发送读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求,从而保证实现数据的读取。
在缺失的主存地址是多个地址时,根据连续不连续会出现多种情况,例如缺失的数据也是连续的数据,缺失的数据部分连续部分不连续或者缺失的数据是全部不连续的。在一些实施例中,可以针对数个缺失的连续的主存地址对应的数据发出新的合并的数据读取请求。在另一些实施例中,不管有没有连续的缺失数据,都对每个缓存单元的数据单独的发送一个数据读取请求。不同的处理方法影响后续的处理,下面举出两个不同的实现示例。
图7示出了本公开的实施例所提供的数据处理方法的实现示例图。
需要说明的是,图7中大部分内容与图6中的步骤类似,对这部分内容不展开描述。如图7所示,该示例可以包括如下步骤。
在步骤S600中,接收数据读取请求。
在步骤S610中,判断满足映射关系的各个缓存单元的主存地址信息中是否包含全部所述主存地址,如果是,则执行步骤S620,如果否,则执行步骤S630。
在步骤S620中,同时返回全部所述主存地址对应的数据。
在步骤S630中,将缺失的各个所述主存地址分别写入各个待处理缓存单元。
步骤S600-步骤S630的具体内容,请参考图4所示的步骤S400-步骤S430的描述,在此不再赘述。
在步骤S640中,生成分别读取缺失的各个所述主存地址所对应的数据的各个缺失数据读取请求。
不论缺失数据是一个还是数个,是数个全部连续还是数个部分连续,都对每一个缓存单元的主存地址对应的数据分别生成一个缺失数据读取请求,挨个发送给下一级模块。每个数据请求只请求读取一个缓存单元的主 存地址对应的缺失数据,缺失数据读取请求的数量与所述缺失的各个所述主存地址的数量相同,下一级模块可以是高速缓存或主存,如果发送模块是最低级高速缓存,则下一级模块是主存。若发送模块是一级缓存,则下一级模块是二级缓存。
在步骤S650中,判断是否接收到全部与缺失的各个所述主存地址对应的数据,若是,执行步骤S660,否则,再次等待和判断。
当向下一级模块发送获取数据的缺失数据读取请求后,还需确定是否接收到了全部的缺失数据,直至得到全部的缺失数据后,同时返回全部主存地址对应的数据。
当然,下一级模块如果是更低一级的缓存,仍然存在下一级模块中缺失所请求的部分数据的情况,此时会再向下一级模块发送请求,容易理解的是,具体的流程与上述流程一致,在此不再赘述。
在步骤S660中,同时返回全部主存地址对应的数据。
步骤S660的具体内容,请参考图5所示的步骤S520的描述,在此不再赘述。
这样,分别生成各个缺失数据读取请求,可以降低数据处理方法的复杂度,并且在接收到数据后,同时返回全部所述主存地址对应的数据,也可以提高数据传输效率。
图8示出了本公开的实施例所提供的数据处理方法的另一示例图。需要说明的是,图8的示例与图7的区别主要在于对于缺失的数据的处理,因此对其他步骤不展开描述。如图8所示,该示例可以包括如下步骤。
在步骤S700中,接收数据读取请求。
在步骤S710中,判断满足映射关系的各缓存单元的主存地址信息中是否包括全部所述主存地址。如果是,则执行步骤S720,如果否,则执行步骤S730。
在步骤S720中,同时返回全部所述主存地址对应的数据。
在步骤S730中,将缺失的各个所述主存地址分别写入各个待处理缓存单元。
步骤S700-步骤S730的具体内容,请参考前述对应部分的描述,在此不再赘述。
在步骤S740中,获取缺失的各个所述主存地址中地址连续的主存地 址,得到各个连续主存地址,根据各个所述连续主存地址,生成读取缺失的各个所述连续主存地址所对应的数据的连续缺失数据读取请求。
在生成缺失数据读取请求时,为了进一步提高数据传输效率,可以进一步获取缺失的主存地址,然后,确定新的连续主存地址,并根据连续主存地址生成连续缺失数据读取请求。
当然,容易理解的是,对于无法确定为连续主存地址的主存地址,则仅需根据各个主存地址分别生成缺失数据读取请求即可。
为方便理解,现举例如下。
比如:缺失的数据的主存地址分别为M1、M2和N1,其中主存地址M1和主存地址M2连续,那么生成连续缺失数据读取请求,主存地址N1与M1和M2不连续,那么,生成缺失数据读取请求,然后向下一级模块发送对应主存地址M1和M2的连续缺失数据读取请求和对应主存地址N1的缺失数据读取请求即可。
在步骤S750中,判断是否接收到全部与缺失的各个所述主存地址对应的数据,如果是,则执行步骤S760,如果否,则重复进行判断。
在步骤S760中,同时返回全部主存地址对应的数据。
可以看出,当满足映射关系的各个缓存单元的主存地址信息中未包括全部所述主存地址时,所生成的缺失数据读取请求中仍然可以包括读取缺失的各个所述连续主存地址所对应的数据的请求,从而,在数据读取过程中,进一步提高数据传输效率。
除了能够提高数据读取时的数据传输效率,当数据写入时,也可以提高数据的传输效率,为此,本公开的实施例还提供一种数据处理方法,请参考图9,图9示出本公开的实施例所提供的数据处理方法的再一流程图。
如图9所示,该流程可以包括如下过程。
步骤S800,接收数据写入请求。
步骤S800的具体内容请参考图4所示的步骤S400的描述,当然,数据处理请求为数据写入请求,其他内容,在此不再赘述。
步骤S810,判断满足映射关系的各个缓存单元的主存地址信息中是否包括全部所述主存地址。如果是,则执行步骤S820,如果否,则执行步骤S830。
步骤S810的具体内容,请参考图4所示的步骤S410的描述,在此不 再赘述。
步骤S820,同时接收各个所述主存地址所对应的数据,并将各个所述数据写入相对应的所述缓存单元。
步骤S820的具体内容,请参考图4所示的步骤S420的描述,当然,由于数据处理请求为数据写入请求,因此,在对数据进行处理时,同时将数据写入对应主存地址的缓存单元或者主存即可。
这样,本公开的实施例所提供的数据处理方法,所接收的一个数据写入请求至少对应两个缓存单元,各个缓存单元同时进行是否命中的判断,并在命中时,同时写入各个缓存单元的数据,可以在无需改变缓存单元大小的情况下,增大各级缓存间、缓存与计算单元间以及缓存与主存间的数据通道宽度,提高数据传输效率,从而可以避免加宽缓存单元所需要的映射关系改变,以及由此带来的工作量的增加,可以实现在较小工作量的基础上,提高带宽,进而提高数据传输效率。
可以实现对于各个主存地址对应数据的写入,并且,由于可以同时写入数据,可以提高数据写入的效率。
在另一实施例中,还可以包括步骤S830中,将缺失的各个所述主存地址分别写入各个待处理缓存单元。
步骤S830的具体内容,请参考图4所示的步骤S430的描述,当然,由于数据处理请求为数据写入请求,因此,将缺失的各个所述主存地址分别写入各个待处理缓存单元即可。
步骤S840,同时接收各个所述主存地址所对应的数据,并将各个所述数据写入相对应的所述缓存单元。
将所有缺失的主存地址都已被写入到高级缓存中后,将接收到的数据分别写入对应的缓存单元里即可。
可见,本公开的实施例所提供的数据处理方法,所接收的一个数据写入请求至少对应两个缓存单元,各个缓存单元同时进行是否命中的判断,并在未命中时,将缺失的各个所述主存地址分别写入各个待处理缓存单元,然后同时写入各个缓存单元的数据,可以在无需改变缓存单元大小的情况下,增大各级缓存间、缓存与计算单元间以及缓存与主存间的数据通道宽度,提高数据传输效率,从而可以避免加宽缓存单元所需要的映射关系改变,以及由此带来的工作量的增加,可以实现在较小工作量的基础上,提 高带宽,进而提高数据传输效率。
上文描述了本公开的多个实施例,各实施例介绍的各可选方式可在不冲突的情况下相互结合、交叉引用,从而延伸出多种可能的实施例,这些均可认为是本公开的实施例披露。
本公开的实施例还提供一种数据处理装置,该装置可以认为是实现本公开的实施例提供的数据处理方法所需设置的功能模块。本文描述的装置内容与上文描述方法内容相互对应参照。
图10示出了本公开的实施例所提供的数据处理装置的框图。该装置可适用于本公开的实施例所提供的数据处理方法。参照图10,该装置可以包括:
数据处理请求模块910,适于接收数据处理请求,所述数据处理请求所请求的数据包括适于存储于至少两个缓存单元的数据,且各个所述缓存单元的数据的主存地址连续。
数据处理模块920,适于当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,对各个所述主存地址对应的数据同时进行数据处理,其中,所述满足映射关系的各个缓存单元,是指所述数据处理请求中的主存地址所对应的缓存单元。
在一些实施例中,数据处理模块920还适于当满足映射关系的各个缓存单元的主存地址信息中未包括全部所述主存地址时,将缺失的各个所述主存地址分别写入各个待处理缓存单元,其中,各个所述待处理缓存单元的原有的主存地址信息,均异于所述数据处理请求所请求数据的主存地址。
在一些实施例中,数据处理请求模块910可以为数据读请求模块,适于接收数据读请求,所述数据读请求所请求的数据包括适于存储于至少两个缓存单元的数据,且各个所述缓存单元的数据的主存地址连续。
数据处理模块920可以为数据读操作模块,适于当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,同时返回全部所述主存地址对应的数据。
在另一些实施例中,该数据读操作模块还适于当满足映射关系的各个缓存单元的主存地址信息中未包括全部所述主存地址时,将缺失的各个所述主存地址分别写入各个待处理缓存单元,生成并发送读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求。
在一些实施例中,数据读操作模块,适于生成并发送读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求,包括:生成分别读取缺失的各个所述主存地址所对应的数据的各个缺失数据读取请求,所述缺失数据读取请求的数量与所述缺失的各个所述主存地址的数量相同。
数据读操作模块,还适于当接收到全部与缺失的各个所述主存地址对应的数据时,同时返回全部主存地址对应的数据。
在一些实施例中,数据读操作模块,适于生成并发送读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求,包括:获取缺失的各个所述主存地址中地址连续的主存地址,得到各个连续主存地址;根据各个所述连续主存地址,生成读取缺失的各个所述连续主存地址所对应的数据的连续缺失数据读取请求;数据读操作模块,还适于当接收到全部与缺失的各个所述主存地址对应的数据时,同时返回全部主存地址对应的数据。
在另一些实施例中,数据处理请求模块910可以为数据写请求模块,适于接收数据写请求,所述数据写请求所请求的数据包括适于存储于至少两个缓存单元的数据,且各个所述缓存单元的数据的主存地址连续。
数据处理模块920可以为数据写操作模块,适于当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,同时写入全部所述主存地址对应的数据。
在另一些实施例中,该数据写操作模块还适于当满足映射关系的各个缓存单元的主存地址信息中未包括全部所述主存地址时,将缺失的各个所述主存地址分别写入各个待处理缓存单元后,同时接收各个所述主存地址所对应的数据,并将各个所述数据写入相对应的所述缓存单元。
可以看出,本公开的实施例所提供的数据处理装置,数据处理请求模块910所接收的一个数据处理请求至少对应两个缓存单元,数据处理模块920对各个缓存单元同时进行是否命中的判断,并在命中时,同时返回各个缓存单元的数据,可以在无需改变缓存单元大小的情况下,增大各级高速缓存间、高速缓存与计算单元间以及高速缓存与主存间的数据通道宽度,提高数据传输效率,并且可以避免加宽缓存单元所需要的映射关系改变,以及由此带来的工作量的增加,可以实现在较小工作量的基础上,提高带宽,进而提高数据传输效率;进一步地,在读写请求的数量比例不一致的情况下,还可以在相邻两级传输之间改变读写请求的数据传输带宽比例, 使得数据传输更满足使用要求。
本公开至少一实施例还提供一种缓存器,该缓存器可以包括一级缓存和二级缓存,所述一级缓存的至少两个缓存单元和所述二级缓存的至少两个缓存单元同时映射,且所述一级缓存和所述二级缓存均包括本公开的实施例所提供的数据处理装置。
这样,通过使一级缓存的至少两个缓存单元和二级缓存的至少两个缓存单元同时映射,可以实现可以同时传输至少两个缓存单元的数据,可以在无需改变缓存单元大小的情况下,增大各级高速缓存间的数据通道宽度,提高数据传输效率。
本公开至少一实施例还提供一种处理器,该处理器执行计算机可执行指令,以实现本公开的实施例提供的数据处理方法。
本公开至少一实施例还提供一种电子设备,该电子设备可以包括本公开的实施例上述提供的处理器。
虽然本公开的实施例披露如上,但本公开并非限定于此。任何本领域技术人员,在不脱离本公开的精神和范围内,均可作各种更动与修改,因此本公开的保护范围应当以权利要求所限定的范围为准。
Claims (25)
- 一种数据处理方法,包括:接收数据处理请求,所述数据处理请求所请求的数据包括适于存储于至少两个缓存单元的数据,且各个所述缓存单元的数据的主存地址连续;当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,对各个所述主存地址对应的数据同时进行数据处理,其中,所述满足映射关系的各个缓存单元,是指所述数据处理请求中的主存地址所对应的缓存单元。
- 如权利要求1所述的数据处理方法,还包括:当满足映射关系的各个缓存单元的主存地址信息中未包括全部所述主存地址时,将缺失的各个所述主存地址分别写入各个待处理缓存单元,其中,各个所述待处理缓存单元的原有的主存地址信息,均异于所述数据处理请求的数据所对应的各个所述主存地址。
- 如权利要求2所述的数据处理方法,其中,所述待处理缓存单元包括待替换缓存单元和空闲缓存单元中的至少一者。
- 如权利要求3所述的数据处理方法,其中,所述待替换缓存单元通过最近最少使用原则RLU确定。
- 如权利要求1-4任一所述的数据处理方法,其中,所述数据处理请求包括数据读取请求,所述当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,对各个所述主存地址对应的数据同时进行数据处理的步骤,包括:当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,同时返回全部所述主存地址对应的数据。
- 如权利要求2-4任一所述的数据处理方法,其中,所述数据处理请求包括数据读取请求,所述当满足映射关系的各个缓存单元的主存地址信息中未包括全部所述主存地址时,将缺失的各个所述主存地址分别写入各个待处理缓存单元的步骤之后,还包括:生成并发送读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求。
- 如权利要求6所述的数据处理方法,其中,所述生成读取缺失的各 个所述主存地址所对应的数据的新数据处理请求的步骤包括:生成分别读取缺失的各个所述主存地址所对应的数据的各个缺失数据读取请求,所述缺失数据读取请求的数量与所述缺失的各个所述主存地址的数量相同;所述数据处理方法还包括:当接收到全部与缺失的各个所述主存地址对应的数据时,同时返回全部主存地址对应的数据。
- 如权利要求6或7所述的数据处理方法,其中,所述生成读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求的步骤包括:获取缺失的各个所述主存地址中地址连续的主存地址,得到各个连续主存地址;根据各个所述连续主存地址,生成读取缺失的各个所述连续主存地址所对应的数据的连续缺失数据读取请求;所述数据处理方法还包括:当接收到全部与缺失的各个所述主存地址对应的数据时,同时返回全部主存地址对应的数据。
- 如权利要求1-8任一所述的数据处理方法,其中,所述数据处理请求包括数据写入请求,所述当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,对各个所述主存地址对应的数据同时进行数据处理的步骤,包括:当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,同时接收各个所述主存地址所对应的数据,并将各个所述数据写入相对应的所述缓存单元。
- 如权利要求2-4任一所述的数据处理方法,其中,所述数据处理请求包括数据写入请求,所述当满足映射关系的各个缓存单元的主存地址信息中未包括全部所述主存地址时,将缺失的各个所述主存地址分别写入各个待处理缓存单元的步骤之后,还包括:同时接收各个所述主存地址所对应的数据,并将各个所述数据写入相对应的所述缓存单元。
- 如权利要求1-10任一项所述的数据处理方法,其中,所述数据处理请求还包括数据处理请求类型标识,所述数据处理请求类型标识用于标 识所述数据处理请求所请求的数据包括适于存储于至少两个缓存单元的数据。
- 如权利要求11所述的数据处理方法,其中,所述数据处理请求类型标识的实现方式包括增加主存地址数量标识位或增加请求类型标识种类。
- 一种数据处理装置,包括:数据处理请求模块,适于接收数据处理请求,所述数据处理请求所请求的数据包括适于存储于至少两个缓存单元的数据,且各个所述缓存单元的数据的主存地址连续;数据处理模块,适于当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,对各个所述主存地址对应的数据同时进行数据处理,其中,所述满足映射关系的各个缓存单元,是指所述数据处理请求中的主存地址所对应的缓存单元。
- 如权利要求13所述的数据处理装置,其中,数据处理模块,还适于当满足映射关系的各个缓存单元的主存地址信息中未包括全部所述主存地址时,将缺失的各个所述主存地址分别写入各个待处理缓存单元,其中,各个所述待处理缓存单元的原有的主存地址信息,均异于所述数据处理请求所请求数据的主存地址。
- 如权利要求13或14所述的数据处理装置,其中,所述数据处理请求模块包括:数据读请求模块,适于接收数据读请求,所述数据读请求所请求的数据包括适于存储于至少两个缓存单元的数据,且各个所述缓存单元的数据的主存地址连续;所述数据处理模块包括:数据读操作模块,适于当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,同时返回全部所述主存地址对应的数据。
- 如权利要求14所述的数据处理装置,其中,所述数据处理请求模块包括:数据读请求模块,适于接收数据读请求,所述数据读请求所请求的数据包括适于存储于至少两个缓存单元的数据,且各个所述缓存单元的数据的主存地址连续;所述数据处理模块包括:数据读操作模块,适于当满足映射关系的各个缓存单元的主存地址信息中未包括全部所述主存地址时,将缺失的各个所述主存地址分别写入各个待处理缓存单元,生成并发送读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求。
- 如权利要求16所述的数据处理装置,其中,所述数据读操作模块,适于生成并发送读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求,包括:生成分别读取缺失的各个所述主存地址所对应的数据的各个缺失数据读取请求,所述缺失数据读取请求的数量与所述缺失的各个所述主存地址的数量相同;所述数据读操作模块,还适于当接收到全部与缺失的各个所述主存地址对应的数据时,同时返回全部主存地址对应的数据。
- 如权利要求16所述的数据处理装置,其中,所述数据读操作模块,适于生成并发送读取缺失的各个所述主存地址所对应的数据的缺失数据读取请求包括:获取缺失的各个所述主存地址中地址连续的主存地址,得到各个连续主存地址;根据各个所述连续主存地址,生成读取缺失的各个所述连续主存地址所对应的数据的连续缺失数据读取请求。
- 如权利要求13-18任一所述的数据处理装置,其中,所述数据处理请求模块包括:数据写请求模块,适于接收数据写请求,所述数据写请求所请求的数据包括适于存储于至少两个缓存单元的数据,且各个所述缓存单元的数据的主存地址连续;所述数据处理模块包括:数据写操作模块,适于当满足映射关系的各个缓存单元的主存地址信息中包括全部所述主存地址时,同时写入全部所述主存地址对应的数据。
- 如权利要求14所述的数据处理装置,其中,所述数据处理请求模块包括数据写请求模块,所述数据处理装置包括:数据写操作模块,适于当满足映射关系的各个缓存单元的主存地址信 息中未包括全部所述主存地址时,将缺失的各个所述主存地址分别写入各个待处理缓存单元后,同时接收各个所述主存地址所对应的数据,并将各个所述数据写入相对应的所述缓存单元。
- 如权利要求13-20任一项所述的数据处理装置,其中,所述数据处理请求还包括各个所述缓存单元的缓存单元索引和数据处理请求类型标识,所述数据处理请求类型标识用于标识所述数据处理请求所请求的数据包括适于存储于至少两个缓存单元的数据。
- 如权利要求21所述的数据处理装置,其中,所述数据处理请求类型标识的实现方式包括增加主存地址数量标识位或增加请求类型标识种类。
- 一种缓存器,包括一级缓存和二级缓存,所述一级缓存的至少两个缓存单元和所述二级缓存的至少两个缓存单元同时映射,且所述一级缓存和所述二级缓存均包括如权利要求13-22任一项所述的数据处理装置。
- 一种处理器,其中,所述处理器执行计算机可执行指令,以实现如权利要求1-12任一项所述的数据处理方法。
- 一种电子设备,包括如权利要求24所述的处理器。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US18/566,583 US20240289275A1 (en) | 2021-11-17 | 2022-05-16 | Data processing method and apparatus, and cache, processor and electronic device |
| EP22894185.2A EP4332781A4 (en) | 2021-11-17 | 2022-05-16 | DATA PROCESSING METHOD AND APPARATUS, AND CACHE MEMORY, PROCESSOR AND ELECTRONIC DEVICE |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202111363242.3 | 2021-11-17 | ||
| CN202111363242.3A CN114036089B (zh) | 2021-11-17 | 2021-11-17 | 数据处理方法、装置、缓存器、处理器及电子设备 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023087640A1 true WO2023087640A1 (zh) | 2023-05-25 |
Family
ID=80138018
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2022/092980 Ceased WO2023087640A1 (zh) | 2021-11-17 | 2022-05-16 | 数据处理方法、装置、缓存器、处理器及电子设备 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US20240289275A1 (zh) |
| EP (1) | EP4332781A4 (zh) |
| CN (1) | CN114036089B (zh) |
| WO (1) | WO2023087640A1 (zh) |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114036089B (zh) * | 2021-11-17 | 2022-10-14 | 海光信息技术股份有限公司 | 数据处理方法、装置、缓存器、处理器及电子设备 |
| CN115809208B (zh) * | 2023-01-19 | 2023-07-21 | 北京象帝先计算技术有限公司 | 缓存数据刷新方法、装置、图形处理系统及电子设备 |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110874328A (zh) * | 2018-08-31 | 2020-03-10 | 爱思开海力士有限公司 | 控制器及其操作方法 |
| CN111694770A (zh) * | 2019-03-15 | 2020-09-22 | 杭州宏杉科技股份有限公司 | 一种处理io请求的方法及装置 |
| WO2020199061A1 (zh) * | 2019-03-30 | 2020-10-08 | 华为技术有限公司 | 一种处理方法、装置及相关设备 |
| CN113157606A (zh) * | 2021-04-21 | 2021-07-23 | 上海燧原科技有限公司 | 一种缓存器实现方法、装置和数据处理设备 |
| CN113222115A (zh) * | 2021-04-30 | 2021-08-06 | 西安邮电大学 | 面向卷积神经网络的共享缓存阵列 |
| CN114036089A (zh) * | 2021-11-17 | 2022-02-11 | 海光信息技术股份有限公司 | 数据处理方法、装置、缓存器、处理器及电子设备 |
Family Cites Families (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2000259497A (ja) * | 1999-03-12 | 2000-09-22 | Fujitsu Ltd | メモリコントローラ |
| US6834327B2 (en) * | 2002-02-08 | 2004-12-21 | Hewlett-Packard Development Company, L.P. | Multilevel cache system having unified cache tag memory |
| US20050125614A1 (en) * | 2003-12-09 | 2005-06-09 | Royer Robert J.Jr. | Adaptive layout cache organization to enable optimal cache hardware performance |
| US8593474B2 (en) * | 2005-12-30 | 2013-11-26 | Intel Corporation | Method and system for symmetric allocation for a shared L2 mapping cache |
| US8954672B2 (en) * | 2012-03-12 | 2015-02-10 | Advanced Micro Devices, Inc. | System and method for cache organization in row-based memories |
| US9361236B2 (en) * | 2013-06-18 | 2016-06-07 | Arm Limited | Handling write requests for a data array |
| US20180189179A1 (en) * | 2016-12-30 | 2018-07-05 | Qualcomm Incorporated | Dynamic memory banks |
| US10423540B2 (en) * | 2017-09-27 | 2019-09-24 | Intel Corporation | Apparatus, system, and method to determine a cache line in a first memory device to be evicted for an incoming cache line from a second memory device |
| US11144466B2 (en) * | 2019-06-06 | 2021-10-12 | Intel Corporation | Memory device with local cache array |
-
2021
- 2021-11-17 CN CN202111363242.3A patent/CN114036089B/zh active Active
-
2022
- 2022-05-16 EP EP22894185.2A patent/EP4332781A4/en active Pending
- 2022-05-16 WO PCT/CN2022/092980 patent/WO2023087640A1/zh not_active Ceased
- 2022-05-16 US US18/566,583 patent/US20240289275A1/en active Pending
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110874328A (zh) * | 2018-08-31 | 2020-03-10 | 爱思开海力士有限公司 | 控制器及其操作方法 |
| CN111694770A (zh) * | 2019-03-15 | 2020-09-22 | 杭州宏杉科技股份有限公司 | 一种处理io请求的方法及装置 |
| WO2020199061A1 (zh) * | 2019-03-30 | 2020-10-08 | 华为技术有限公司 | 一种处理方法、装置及相关设备 |
| CN113157606A (zh) * | 2021-04-21 | 2021-07-23 | 上海燧原科技有限公司 | 一种缓存器实现方法、装置和数据处理设备 |
| CN113222115A (zh) * | 2021-04-30 | 2021-08-06 | 西安邮电大学 | 面向卷积神经网络的共享缓存阵列 |
| CN114036089A (zh) * | 2021-11-17 | 2022-02-11 | 海光信息技术股份有限公司 | 数据处理方法、装置、缓存器、处理器及电子设备 |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP4332781A4 * |
Also Published As
| Publication number | Publication date |
|---|---|
| EP4332781A1 (en) | 2024-03-06 |
| US20240289275A1 (en) | 2024-08-29 |
| EP4332781A4 (en) | 2024-10-02 |
| CN114036089A (zh) | 2022-02-11 |
| CN114036089B (zh) | 2022-10-14 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7340326B2 (ja) | メンテナンス動作の実行 | |
| CN113424160B (zh) | 一种处理方法、装置及相关设备 | |
| CN112527395B (zh) | 数据预取方法和数据处理装置 | |
| US10042576B2 (en) | Method and apparatus for compressing addresses | |
| JP6859361B2 (ja) | 中央処理ユニット(cpu)ベースシステムにおいて複数のラストレベルキャッシュ(llc)ラインを使用してメモリ帯域幅圧縮を行うこと | |
| CN101221538B (zh) | 实现对缓存中数据快速查找的系统和方法 | |
| US10198357B2 (en) | Coherent interconnect for managing snoop operation and data processing apparatus including the same | |
| KR102355374B1 (ko) | 이종 메모리를 이용하여 메모리 주소 변환 테이블을 관리하는 메모리 관리 유닛 및 이의 메모리 주소 관리 방법 | |
| US6745292B1 (en) | Apparatus and method for selectively allocating cache lines in a partitioned cache shared by multiprocessors | |
| CN110046107B (zh) | 存储器地址转换装置和方法 | |
| CN114546896A (zh) | 系统内存管理单元、读写请求处理方法、电子设备和片上系统 | |
| CN109983538B (zh) | 存储地址转换 | |
| EP4332781A1 (en) | Data processing method and apparatus, and cache, processor and electronic device | |
| US7797492B2 (en) | Method and apparatus for dedicating cache entries to certain streams for performance optimization | |
| CN116069719A (zh) | 处理器、内存控制器、片上系统芯片和数据预取方法 | |
| WO2021008552A1 (zh) | 数据读取方法和装置、计算机可读存储介质 | |
| CN116126747B (zh) | 一种缓存方法、缓存架构、异构架构及电子设备 | |
| CN115080464B (zh) | 数据处理方法和数据处理装置 | |
| CN117389914A (zh) | 缓存系统、缓存写回方法、片上系统及电子设备 | |
| CN116611992A (zh) | 一种基于聚合标签阵列的gpu共享l1缓存 | |
| CN114556335B (zh) | 一种片内缓存及集成芯片 | |
| US10942859B2 (en) | Computing system and method using bit counter | |
| JP7311959B2 (ja) | 複数のデータ・タイプのためのデータ・ストレージ | |
| CN121187711B (zh) | 处理器的PCIe控制器和用于处理器的传输控制方法 | |
| CN102147772B (zh) | 快取存储器的替换装置及方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 22894185 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18566583 Country of ref document: US Ref document number: 2022894185 Country of ref document: EP |
|
| ENP | Entry into the national phase |
Ref document number: 2022894185 Country of ref document: EP Effective date: 20231201 |