WO2024060710A1 - 一种页面换入方法以及装置 - Google Patents

一种页面换入方法以及装置 Download PDF

Info

Publication number
WO2024060710A1
WO2024060710A1 PCT/CN2023/100492 CN2023100492W WO2024060710A1 WO 2024060710 A1 WO2024060710 A1 WO 2024060710A1 CN 2023100492 W CN2023100492 W CN 2023100492W WO 2024060710 A1 WO2024060710 A1 WO 2024060710A1
Authority
WO
WIPO (PCT)
Prior art keywords
host
page
target page
swap
memory
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2023/100492
Other languages
English (en)
French (fr)
Inventor
钟刊
崔文林
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Huawei Technologies Co Ltd
Original Assignee
Huawei Technologies Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Huawei Technologies Co Ltd filed Critical Huawei Technologies Co Ltd
Priority to EP23867006.1A priority Critical patent/EP4582943A4/en
Publication of WO2024060710A1 publication Critical patent/WO2024060710A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00—Arrangements for program control, e.g. control units
    • G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/44—Arrangements for executing specific programs
    • G06F9/455—Emulation; Interpretation; Software simulation, e.g. virtualisation or emulation of application or operating system execution engines
    • G06F9/45533—Hypervisors; Virtual machine monitors
    • G06F9/45558—Hypervisor-specific management and integration aspects
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F12/00—Accessing, addressing or allocating within memory systems or architectures
    • G06F12/02—Addressing or allocation; Relocation
    • G06F12/08—Addressing or allocation; Relocation in hierarchically structured memory systems, e.g. virtual memory systems
    • G06F12/10—Address translation
    • G06F12/1009—Address translation using page tables, e.g. page table structures
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F12/00—Accessing, addressing or allocating within memory systems or architectures
    • G06F12/02—Addressing or allocation; Relocation
    • G06F12/08—Addressing or allocation; Relocation in hierarchically structured memory systems, e.g. virtual memory systems
    • G06F12/10—Address translation
    • G06F12/109—Address translation for multiple virtual address spaces, e.g. segmentation
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00—Arrangements for program control, e.g. control units
    • G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/44—Arrangements for executing specific programs
    • G06F9/455—Emulation; Interpretation; Software simulation, e.g. virtualisation or emulation of application or operating system execution engines
    • G06F9/45533—Hypervisors; Virtual machine monitors
    • G06F9/45558—Hypervisor-specific management and integration aspects
    • G06F2009/45583—Memory management, e.g. access or allocation

Definitions

  • the present application relates to the field of communication technology, and in particular, to a page switching method and device.
  • the processor in the host can swap out the inactive pages in the memory of the host to the swap partition, thereby releasing the memory space of the host, thereby realizing the memory expansion of the host.
  • the swap A partition is considered a type of virtual memory, and the swap partition will typically be located on the host's hard drive.
  • the processor in the host needs to call these pages that have been swapped out to the swap partition, the processor will obtain the page from the swap partition and swap the page into the host's memory.
  • swapping pages from the swap partition to the host's memory will occupy processor resources, making the processor unable to provide more computing power to support business-type work in the host.
  • the present application provides a page switching method and device to reduce the occupation of the processor by the page switching operation.
  • the present application provides a page swapping method, in which an acceleration device connected to a host communication system bus can replace the processor in the host to swap pages from a storage device into the memory of the host.
  • the acceleration device in the host receives a swap-in command, and the swap-in command is used to instruct the target page to be swapped into the memory of the host.
  • the acceleration device obtains the target page from the storage device, swaps the target page into the memory, and notifies the host (such as the processor or MMU in the host) that the target page has been swapped into the memory.
  • the page swap operation can be performed by the acceleration device.
  • the processor in the host no longer needs to swap the target page from the storage device into the host's memory, effectively releasing the processor's computing power and reducing the need for page swaps.
  • the occupation of the processor ensures the execution efficiency of business work in the host.
  • the acceleration device may receive the swap instruction from a processor in the host or an MMU in the host. For example, when the MMU finds a page missing in the PT, it can send a swap command to the acceleration device. After the acceleration device swaps the target page into the host's memory, the acceleration device notifies the MMU that the target page has been swapped into the memory. For another example, when the processor determines that there is enough memory space in the memory, it can send a swap command to the acceleration device. After the acceleration device swaps the target page into the memory of the host, the acceleration device notifies the processor that the target page has been swapped into the memory. middle.
  • the processor and MMU in the host only need to send a swap command, and the target page can be swapped into the memory of the host through the acceleration device without the participation of the processor.
  • the page swap method is simpler and faster.
  • the acceleration device when the acceleration device obtains the target page from the storage device, it first obtains the corresponding relationship between the virtual address of the target page and the exchange address of the target page.
  • the exchange address of the target page is the address of the target page in the storage device. .
  • the acceleration device determines the exchange address of the target page according to the corresponding relationship. Then obtain the target page from the storage device according to the exchange address of the target page.
  • the acceleration device can communicate with the storage device and send Send a data request carrying the exchange address of the target page to request to obtain the target page. After receiving the data request, the storage device can feed back the target page to the acceleration device.
  • the acceleration device After the acceleration device obtains the corresponding relationship, it can use the corresponding relationship to obtain the exchange address of the target page, so as to communicate with the storage device and request to obtain the target page.
  • the storage device can provide storage space as the virtual memory of the host.
  • the storage device may be a storage device deployed outside the host and connected to the host through a network.
  • the storage device can also be local to the host.
  • the storage device can also be deployed in the cloud.
  • the storage device can be storage space pre-allocated for the host in a private cloud or public cloud.
  • the corresponding relationship may be recorded when the acceleration device swaps out the target page from the memory to the storage device.
  • the corresponding relationship may also be obtained by the acceleration device from the processor of the host.
  • the corresponding relationship may be recorded when the processor swaps out the target page from the memory to the storage device.
  • the host further includes a processor, and the acceleration device is connected to the processor through a system bus.
  • the acceleration device swaps the target page into the memory, it can receive a notification sent by the processor, and the notification is used to instruct the processor.
  • the physical address configured for the target page in the memory; after the acceleration device obtains the physical address, it can write the target page to the physical address through DMA.
  • the acceleration device After the acceleration device obtains the physical address, it can use DMA to write the target page into the memory, bypassing the processor and further reducing the processor occupancy.
  • the acceleration device may also perform a swap-out operation of the target page.
  • the acceleration device receives a swap-out instruction sent by the processor of the host.
  • the swap-out instruction is used to instruct the target page in the host memory to be swapped out to the storage device.
  • the swap-out instruction includes the virtual address of the target page and the target page; the acceleration device After receiving the swap-out command, the target page is swapped out to the storage device and the corresponding relationship is recorded.
  • the acceleration device undertakes the main operations of page swapping in and page swapping out, which can reduce the occupation of the processor and ensure that the host processor can provide sufficient computing power to support business-type work.
  • the acceleration device also has a data decompression function.
  • the acceleration device obtains the compressed target page from the storage device, the acceleration device decompresses the compressed target page to obtain the target page, and then converts the target page The page is swapped into the host's memory.
  • the acceleration device can decompress the compressed target page by itself without the participation of the processor, further releasing the computing power of the processor.
  • the acceleration device when the acceleration device swaps out the target page to the storage device, in order to reduce the storage space occupied by the target page, the acceleration device compresses the target page and then swaps it out to the storage device.
  • this application also provides a page switching device.
  • the page switching device has the function of implementing the behavior in the method example of the first aspect.
  • Functions can be implemented by hardware, or by hardware executing corresponding software.
  • Hardware or software includes one or more units corresponding to the above functions.
  • the structure of the device includes a receiving module (for receiving swap-in instructions or swap-out instructions), and a swap-out module (for swapping out the target page to the storage device).
  • it also includes Swap-in modules (used to swap the target page from the storage device into the host's memory). These modules can perform the corresponding functions of the DPU in the above-mentioned method examples in the first aspect. For details, please refer to the detailed description in the method examples, which will not be described here. .
  • the present application also provides an acceleration device, which includes a processor with processing functions such as a DPU, GPU, NPU, or TPU.
  • a processor with processing functions such as a DPU, GPU, NPU, or TPU.
  • the types of processors listed above are only examples, and this application does not limit the types of processors in the acceleration device.
  • the specific types of processors included are described below only by taking the processor included in the acceleration device as a DPU as an example.
  • it also includes a power supply circuit and memory. The power supply circuit is used to power the DPU.
  • the DPU has the function of implementing the behavior of the DPU in the method examples in the above-mentioned first aspect and various possible implementation methods of the first aspect.
  • the beneficial effects can be found in the description of the first aspect and will not be repeated here.
  • computer program instructions are stored in a memory, and the DPU is coupled to the memory.
  • the DPU can call computer execution instructions stored in the memory to execute the method executed by the DPU in the above-mentioned first aspect and various possible implementations of the first aspect.
  • embodiments of the present application also provide a computing device.
  • the computing device may be the host mentioned in the first aspect.
  • the beneficial effects can be found in the description of the first aspect and will not be described again here.
  • the computing device includes an accelerator. device and memory.
  • the acceleration device receives the swap-in instruction, which is used to instruct the target page to be swapped into the memory of the computing device. After the acceleration device receives the swap-in instruction, the acceleration device obtains the target page from the storage device; swaps the target page into Memory, notifies the computing device that the target page has been swapped into memory.
  • the acceleration device receives a swap instruction from a processor or MMU of the computing device.
  • the acceleration device when the acceleration device obtains the target page from the storage device, it first obtains the correspondence between the virtual address of the target page and the exchange address of the target page.
  • the exchange address of the target page is the address of the target page in the storage device.
  • the acceleration device determines the exchange address of the target page based on the corresponding relationship.
  • the acceleration device communicates with the storage device and obtains the target page from the storage device according to the exchange address of the target page.
  • the storage device is a storage device deployed outside the computing device and connected to the computing device through a network, or is a local storage device of the computing device, or is allocated to the computing device in a public cloud or a private cloud. of storage space.
  • the corresponding relationship is recorded when the acceleration device swaps out the target page from the memory to the storage device.
  • the corresponding relationship may also be obtained by the acceleration device from the processor of the computing device, and the corresponding relationship may be recorded when the processor swaps out the target page from the memory to the storage device.
  • the computing device further includes a processor, and the acceleration device is connected to the processor through a system bus.
  • the processor may send a notification to the acceleration device, and the notification is used to indicate the physical address configured by the processor for the target page in the memory. After receiving the notification, the acceleration device writes the target page to the physical address through DMA.
  • the processor of the computing device sends a swap-out instruction to the acceleration device.
  • the swap-out instruction is used to instruct the target page in the memory of the computing device to be swapped out to the storage device.
  • the swap-out instruction includes the target page. Virtual address and target page.
  • the acceleration device After receiving the swap-out command, the acceleration device swaps out the target page to the storage device and records the corresponding relationship.
  • the acceleration device obtains the compressed target page from the storage device, decompresses the compressed target page, obtains the target page, and swaps the target page into the memory.
  • the acceleration device when the acceleration device swaps out the target page to the storage device, it compresses the target page and then swaps it out to the storage device.
  • this application provides a data processing system.
  • the data processing system includes a host and a storage device.
  • the storage device communicates with the host through an internal bus or network, and the host is equipped with an acceleration device.
  • the host sends a swap-out command to the acceleration device, and the swap-out command contains data.
  • the acceleration device compresses the data and saves the compressed data in the memory of the acceleration device or in the memory of the host.
  • the acceleration device compresses the data so that the data takes up less space in the memory of the acceleration device or in the host memory. Reduce or improve the utilization of acceleration devices in memory or host memory.
  • the compressed data is stored in the memory of the acceleration device.
  • the acceleration device can ensure that the free memory space in the memory of the acceleration device is less than the first free threshold, or the access frequency of the compressed data is less than When the threshold is reached, the compressed data is migrated to the storage device, which can reduce the memory usage in the acceleration device.
  • the compressed data is stored in the memory of the host, and the acceleration device can be used when the free memory space in the memory of the host is less than the second idle threshold, or the access frequency of the compressed data is lower than the threshold.
  • the compressed data is migrated to the memory of the acceleration device, which can reduce the memory usage in the host.
  • the host sends a swap instruction to the acceleration device, and the swap instruction instructs to obtain data.
  • the acceleration device After receiving the swap command, the acceleration device obtains the compressed data, decompresses the compressed data, and after decompression, stores the decompressed data in the memory of the host.
  • the acceleration device obtains the compressed data from the storage device, and if the compressed data is stored in the memory of the acceleration device or the memory of the host, the acceleration device The device can accelerate the memory in the device or the memory of the host to obtain the compressed data.
  • the present application further provides a computer-readable storage medium, in which instructions are stored, and when the computer-readable storage medium is run on a computer, the computer executes the method in the above-mentioned first aspect and various possible implementations of the first aspect.
  • the present application also provides a computer program product containing instructions that, when run on a computer, cause the computer to execute the method in the above-mentioned first aspect and each possible implementation of the first aspect.
  • this application also provides a computer chip, which is connected to a memory.
  • the chip is used to read and execute the software program stored in the memory, and execute the method in the above-mentioned first aspect and each possible implementation of the first aspect. .
  • Figure 1 is a schematic diagram of the virtualization architecture of a host provided by this application.
  • Figure 2 is a schematic structural diagram of a host provided by this application.
  • FIG. 3 is a schematic diagram of a page swapping method provided by this application.
  • Figure 4 is a schematic diagram of a page switching method provided by this application.
  • Figure 5 is a schematic structural diagram of a page switching device provided by this application.
  • Virtualization technology computing instance, container, virtual machine (VM).
  • Virtualization is a resource management technology that abstracts and transforms the host's physical resources, such as processors, memory, and interfaces, and presents them. Virtualization is a resource configuration method from a logical perspective and is a logical abstraction of physical resources.
  • the host can use virtualization technology to form a software module with an independent operating environment on the host.
  • the software module with an independent operating environment formed on the host is called a computing instance.
  • the computing instance can be a virtual machine or a container.
  • a virtual machine is a "complete computer" with complete hardware system functions simulated through virtual machine technology and running in a completely isolated environment. Everything that can be done on a physical computer can be done on a virtual machine.
  • a virtual machine has components such as a processor (the processor is also called a virtual processor), memory, and a hard disk.
  • the processor, memory, and hard disk of the virtual machine are The software is virtualized by the host's processor, memory, hard disk and other components.
  • An operating system is installed on the virtual machine, and the operating system on the virtual machine is independent of the operating system of the host itself. In order to distinguish between these two different operating systems, the operating system on the host is usually called the host operating system (hostOS), and the operating system on the virtual machine is usually called the guest operating system (guestOS). ).
  • hostOS host operating system
  • guestOS guest operating system
  • a container is an independent operating environment simulated through virtualization technology.
  • the container is similar to a lightweight sandbox, which shields the software and hardware outside the container.
  • the container implements virtualization at the operating system level and directly replicates the environment. Use the host operating system.
  • a compute instance is viewed as a special "process.” This "process” will perform computing tasks and occupy the host's processor, memory, hard disk and other resources.
  • the host operating system configures a memory space dedicated to the computing instance for the computing instance, and the computing instance occupies the memory space to support the computing tasks that the computing instance needs to perform.
  • MMU Memory management unit
  • MMU is also called paged memory management unit (PMMU).
  • PMMU paged memory management unit
  • the MMU is a hardware device located between the core of the processor and the bus connecting cache and memory.
  • the MMU is usually considered to be part of the processor. In this case, the operations performed by the MMU can be considered as part of the processor. The operation performed.
  • the MMU is an independent hardware device and is only explained as an example.
  • the MMU is mainly used to process memory access requests initiated by the processor.
  • the MMU has an address translation function and can convert between the virtual address of the memory and the physical address of the memory.
  • the MMU can parse the access request issued by the processor, convert the virtual address of the memory carried in the access request into the physical address of the memory, and write the data carried in the access request to the physical address of the memory (access A request is used to request to write data at the virtual address of the memory); or to read data from the physical address of the memory and feed the data back to the processor (an access request is used to request to read data at the virtual address of the memory) ).
  • the MMU also has memory protection functions and cache control functions for the processor.
  • MMU can manage memory space at page granularity (that is, page memory management). In the management of memory space at page granularity, the memory space is divided into blocks of fixed size, and one block is a page. The MMU can allocate, manage, and protect memory in units of pages.
  • the MMU In order to convert a virtual address to a physical address, the MMU needs to obtain an address translation table, which is called a page table (PT).
  • the page table includes one or more page table entries (PTE). .
  • PTE page table entries
  • Each page table entry corresponds to a page and records the correspondence between the virtual address of the page and the physical address of the page as well as some control information.
  • the address translation table is maintained by the host operating system and stored in the host's memory. That is to say, the host operating system can perform some updates on the PT, such as adding page table entries, deleting page table entries, modifying page table entries, etc. operation and save the updated page table in the host's memory.
  • the MMU needs to perform address translation, it calls the PT from the host's memory.
  • the virtual addresses of pages used by different processes running on the host may be the same, in order to avoid confusion between the virtual addresses used by different processes.
  • the host operating system maintains the corresponding PT for each process on the host (the process here also includes computing instances).
  • the PT corresponding to the computing instance records the mapping relationship between the virtual address and the physical address of the memory space of the host occupied by the computing instance.
  • Table 1 abstracts PT into a table, and each row of the table indicates a PTE. Each row records the virtual address (VA) of a page and the physical address (PA) of the page.
  • VA virtual address
  • PA physical address
  • the PTE also includes a flag bit. This flag bit is used to identify whether the PTE is accessed. In the embodiment of this application, the flag bit is 1, which indicates that the PTE is accessed.
  • the access of the PTE indicates that the memory space indicated by the physical address in the PTE is accessed; the flag bit indicates that the PTE is accessed. If the flag bit is 0, it means that the PTE has not been accessed. If the PTE is accessed, it means that the memory space indicated by the physical address in the PTE has not been accessed.
  • Each time the MMU reads a PTE it will set the flag bit in the PTE to 1. If the PTE is not read again within a specific period of time after the flag bit is set to 1, the flag bit of the PTE will flip to 0 on its own.
  • the computing instance When the computing instance needs to write data to virtual address V1, the computing instance will send access request 1 to the MMU (from the perspective of hardware interaction, the access request 1 is sent to the MMU by the processor), and the access request 1 Used to request to write data at virtual address V1.
  • This access request 1 also carries the virtual address V1 of the page and the data that needs to be written.
  • the MMU finds the PT corresponding to the computing instance, and searches the PT to see whether there is a physical address corresponding to the virtual address V1. If the PTE recording the virtual address V1 is not found in the PT, the MMU will trigger the page fault process and send an exception signal indicating the page fault interrupt to the host operating system. The host operating system calculates the page fault in the memory.
  • the instance allocates a section of free memory space (that is, a free page), and the physical address P1 of the memory space is used as the physical address corresponding to the virtual address V1.
  • the host operating system updates the PT, that is, adds a PTE to the PT, and the PTE records the correspondence between the physical address P1 and the virtual address V1.
  • the host operating system saves the updated PT in the host's memory and notifies the MMU to continue processing.
  • the MMU obtains the updated PT, searches for the PTE that records the virtual address V1, determines the physical address P1 corresponding to the virtual address V1, and sets the flag position in the PTE to 1.
  • the MMU writes the data to the physical address P1.
  • the computing instance When the computing instance needs to read data from virtual address V1, the computing instance will send access request 2 to the MMU (from the perspective of hardware interaction, the access request 2 is sent to the MMU by the processor), and the access request 2 is used to A request is made to read data at virtual address V1.
  • This access request 2 also carries the virtual address V1 of the page.
  • the MMU finds the PT corresponding to the computing instance, and searches the PT to see whether there is a physical address corresponding to the virtual address V1. If the PTE recording the virtual address V1 is found in the PT, the MMU reads the data from the physical address corresponding to the virtual address, feeds the data back to the computing instance, and sets the flag bit to 1.
  • Scenario 1 There is a switching mechanism in the host and a switching module is deployed.
  • the MMU can trigger the page fault interrupt process and obtain the data on the virtual address V1 through the switching module.
  • the process by which the MMU obtains the data at the virtual address V1 through the switching module can be found in the following description of the switching mechanism, which will not be described again here.
  • Scenario 2 There is no switching mechanism in the host. The MMU will trigger the page fault interrupt process and send an exception signal indicating the page fault interrupt to the host operating system, which will be handed over to the host operating system for continued processing.
  • the swap mechanism refers to memory swapping.
  • the swap mechanism proposes two concepts: physical memory and virtual memory.
  • the physical memory is the actual memory of the host, which may also be referred to as the host's memory in the present embodiment of the application.
  • the virtual memory is the swap partition, which is usually located in the local storage device of the host.
  • the local storage device of the host refers to the storage device connected to the host through the system bus, such as the hard disk of the host.
  • the virtual memory may also be located in a remote storage device.
  • the remote storage device refers to the storage device connected to the host through the network, such as a storage node in a remote storage system.
  • the memory can also be deployed in a private cloud or a public cloud, that is, it can also be located in a cloud data center.
  • the virtual memory is a storage space pre-allocated for the host in the cloud data center, and the storage space is distributed on one or more devices. That is, the virtual memory can be deployed in the cloud, and the device where the virtual memory is located is located in the cloud.
  • the exchange mechanism it is allowed to open up a part of the storage space from the host's storage device (the storage device here includes the host's local storage device, remote storage device, and device in the cloud), and virtualize this part of the storage space into the host's memory space as virtual memory.
  • the host's memory space is insufficient, data is exchanged between the physical memory and the virtual memory.
  • the switching module is essentially a software module running on the host and is a part of the host operating system.
  • the exchange mechanism is implemented as follows:
  • the switch module may swap out some pages in the physical memory to the virtual memory, and save the data in these pages to the virtual memory to free up some memory space in the physical memory. This process is called page swap out, and these swapped pages can be called swapped-out pages.
  • the PTE of the swapped-out page in the PT i.e., the PTE that records the correspondence between the virtual address of the swapped-out page and the physical address of the swapped-out page
  • the switch module needs to record the correspondence between the virtual address of the swapped-out page and the address of the swapped-out page in the virtual memory.
  • the swap module needs to call these pages on the host (for example, the MMU receives access request 2 initiated by a process on the host, and the access request carries the virtual address of the swapped-out page), or there is enough free space in the host's physical memory. (For example, if the size of free memory space in physical memory is greater than the preset threshold), the swap module can swap out pages from virtual memory into physical memory. This process is page swap in. These swapped-in pages may be called swapped-in pages. In this process, since the swap-in page is saved in physical memory, the host operating system needs to allocate a new physical address for the swap-in page in physical memory (this physical address is called the physical address of the swap-in page).
  • the MMU uses The PTE about the swapped-in page will be added to the PT (that is, the PTE that records the previous correspondence between the virtual address of the swapped-in page and the physical address of the swapped-in page).
  • the swap module will also delete the previously recorded correspondence between the virtual address of the swapped-in page and the address of the swapped-in page in the virtual memory.
  • the following combines the workflow of the MMU and the switching module to further explain the processing flow of a common access request initiated for a computing instance to read page A. Since the switching module is usually considered a part of the host operating system, this processing In the description of the process, the switching module and the host operating system are no longer distinguished, and they are collectively referred to as the host operating system.
  • the computing instance initiates an access request, which is used to request to read the data in page A.
  • the access request carries the virtual address V1 of page A.
  • the MMU obtains the access request and obtains the PT corresponding to the computing instance from the memory.
  • the MMU searches for the PTE that records the virtual address V1 from the PT.
  • the MMU determines the physical address P1 corresponding to the virtual address V1, sends the physical address P1 to the memory, and feeds back the data returned by the memory on the physical address P1. Give calculation examples.
  • the MMU will trigger the page fault process and send a page fault exception signal to the host operating system, and the host operating system will suspend the operation of the computing instance.
  • the host operating system looks for the correspondence between the virtual address V1 recorded when page A was swapped out and the address X1 of the data in page A in the virtual memory, swaps the data in page A from the virtual memory to the host's memory according to the address X1, and reallocates a free memory space for page A.
  • the host operating system reallocates the physical address P2 of the free memory space for the page, and uses the physical address P2 as the physical address corresponding to the virtual address V1 to update the PT.
  • the host operating system adds a PTE to the PT, and the PTE records the correspondence between the physical address P2 and the virtual address V1.
  • the host operating system PT saves the updated PT in the host's memory, notifies the MMU to continue processing, and starts the operation of the computing instance.
  • the MMU obtains the updated PT, searches for the PTE that records the virtual address V1, determines the physical address P2 corresponding to the virtual address V1, and sets the flag position in the PTE to 1.
  • the MMU sends the physical address P2 to the host's memory, and feeds back the data fed back from the host's memory to the computing instance.
  • the computing instances can occupy the memory space of the host (that is, occupy some pages in the memory), but the computing instances (especially virtual machines) use the occupied memory space. situation, the host operating system cannot sense it.
  • the host operating system determines which page or pages in memory will be swapped out.
  • a common method is that the host operating system will use the least recently used (LRU) algorithm to determine the least recently used page or pages in the memory, and swap out the determined page or pages to the virtual memory. in memory.
  • LRU least recently used
  • the host operating system cannot obtain the usage of the pages occupied by the computing instance (such as a virtual machine). For example, the host operating system cannot obtain whether each page occupied by the computing instance is the core of the computing instance. The host operating system cannot know the occupied pages or pages occupied by applications in the computing instance. The host operating system cannot know the frequency of access to the occupied pages by the computing instance. In view of this, the host operating system will use some of the more important pages of the computing instance as swap-out pages and swap them out into the virtual memory. When the computing instance needs to call these pages, the MMU will trigger the page fault interrupt process, and the page fault interrupt process will be accompanied by The page swap operation increases the delay, causing the computing instance to freeze and affecting the work efficiency of the computing instance.
  • embodiments of the present application propose a page swapping method.
  • the computing instance can send a page tag to the host operating system.
  • the page tag indicates the importance of each page occupied by the computing instance in the host's memory.
  • the host operating system determines the target page in the host's memory that needs to be swapped out based on the page label, and transfers the determined The target page is swapped out to this storage device.
  • the computing instance can inform the host operating system of the importance of the occupied pages, the host operating system can avoid swapping out the pages of the computing instance that occupy the higher importance to the storage device.
  • the MMU will not trigger the page fault interrupt process, and the computing instance will not be suspended.
  • the host operating system when the host operating system swaps out pages to the storage device, the host operating system sends a swap out command to the host's acceleration device.
  • the swap out command is used to instruct the acceleration device to swap out the page.
  • the swap out command carries the virtual address of the page as well as the page.
  • the acceleration device obtains the page, stores the page in the storage device, and records the correspondence between the virtual address of the page and the address of the page in the storage device. That is to say, the acceleration device can replace the host operating system (ie, the processor) to swap pages from the host's memory to the storage device, which can reduce the occupation of the processor and release the computing power of the processor.
  • the acceleration device of the host can replace the processor in the host (which can also be understood as the host operating system) to swap the page from the storage device to Host memory.
  • the MMU in the host sends a swap command to the acceleration device.
  • the swap command is used to instruct the page to be swapped into the memory.
  • the swap command can carry the virtual address of the page.
  • the acceleration device receives the swap command.
  • the page can be swapped from the storage device into the memory of the host according to the swap instruction, and the MMU is notified that the page has been swapped into the memory of the host.
  • the acceleration device can determine the address of the page in the storage device based on the recorded correspondence, obtain the page from the storage device based on the address of the page in the storage device, and write the page into the memory of the host. . This can further reduce processor usage.
  • swapping out pages to a storage device outside the memory refers to swapping out pages to a swap partition, where the storage device is a storage device with a swap partition deployed. That is to say, the two expressions of swapping out pages to a storage device outside the memory and swapping out pages to a swap partition are essentially the same. In the following, they are unified as swapping out pages to a swap partition. However, it should be understood that swapping out the page to the swap partition means swapping the page into the storage device and storing the page in the storage device.
  • swapping pages from the storage device into the host's memory means swapping pages from the swap partition into the host's memory, swapping pages from the storage device into the host's memory, and swapping pages from the swap partition into the host.
  • the essence expressed by the two expressions of memory is the same. In the following, they are unified as pages being swapped from the swap partition to the host's memory. That is to say, in order to make the expression more clear, in the process of swapping in and swapping out pages, the swap partition will be used to refer to the storage device where the swap partition is located.
  • FIG1 a schematic diagram of a virtualization architecture of a host provided in an embodiment of the present application is shown.
  • Virtualization technology is applied to a host 10 to form a virtualization architecture as shown in FIG1 .
  • the virtualization architecture of the host 10 includes underlying hardware 100, a host operating system 200, a computing instance management unit 300, and at least one computing instance 400.
  • the host 10 is a computing device.
  • the host 10 can be a computing device such as a server, a mobile terminal, or a tablet computer.
  • the underlying hardware 100 refers to some hardware components in the host 10, such as a processor 110, a memory 120, an input output (I/O) interface 130, etc.
  • the underlying hardware 100 can be understood as physical resources on the host 10, which are objects that need to be virtualized in the virtualization technology.
  • the host 10 runs software modules such as a host operating system 200, a computing instance management unit 300, and a computing instance 400.
  • the host operating system 200 runs on the processor 110 and is used to implement the basic functions of the host 10; the basic functions include but are not limited to: management functions for the underlying hardware 100, input/output devices (such as monitors, keyboards) connected to the host 10 , mouse, etc.) and the management of processes in the host 10.
  • the embodiments of this application mainly involve management functions for the underlying hardware 100.
  • the host operating system 200 can monitor the occupancy of the memory 120 and allocate memory space; the host operating system 200 can monitor the occupancy of the memory 120 and allocate memory space. 200 can also detect the access frequency of pages in the memory 120; the host operating system 200 can implement page exchange between the memory 120 of the host 10 and the swap partition 160.
  • the host operating system 200 can obtain the page tag of the computing instance 400 from the computing instance 400.
  • the page tag indicates the importance of the page occupied by the computing instance 400 in the memory 120 of the host 10.
  • the host operating system 200 can also obtain the page access frequency of each page occupied by the computing instance 400.
  • the host operating system 200 can swap out pages from the host 10 based on some or all of the page labels, page access frequencies, and user-configured page swap parameters. Determine the target page within the inner page.
  • the host operating system 200 swaps out the determined target page to the swap partition 160 .
  • the host operating system 200 may also swap the pages swapped out to the swap partition 160 into the memory 120 of the host 10 .
  • the host operating system 200 includes an instance page identification module 210 and a switching module 220 .
  • the instance page identification module 210 can obtain the page tag from the computing instance 400, and can also obtain the page access frequency of each page occupied by the computing instance 400 according to the access status of the MMU to the PT corresponding to the computing instance 400.
  • the instance page identification module 210 can also provide a parameter configuration interface for users, allowing users to configure page exchange parameters.
  • the page exchange parameters indicate the filter conditions of the target page.
  • the page swap parameters include but are not limited to: the identification of the computing instance 400 that is allowed to swap out the page, the page scanning frequency, the cold page determination threshold, and the maximum memory space that the computing instance 400 is allowed to occupy.
  • the instance page identification module 210 When a page needs to be swapped out from the memory 120 of the host 10 to the swap partition 160, the instance page identification module 210 Some or all of the page tags, page access frequencies, and page exchange parameters of the computing instance 400 are used to filter the target page from the pages occupied by the computing instance 400.
  • the target page is a page that is allowed to be swapped out among the pages occupied by the computing instance 400. .
  • the instance page identification module 210 sends the virtual address of the target page to the switching module 220.
  • the swap module 220 is used to implement page swapping between the memory 120 of the host 10 and the swap partition 160 . That is, the swap module 220 can swap out pages in the memory 120 of the host 10 to the swap partition 160 and swap pages in the swap partition 160 into the memory 120 of the host 10 .
  • the swap module 220 When the swap module 220 needs to swap a page in the memory 120 of the host 10 to the swap partition 160 , the swap module 220 swaps out the target page to the swap partition 160 .
  • the exchange module 220 can also update the PT corresponding to the computing instance 400 and delete the PTE related to the target page.
  • the swap module 220 When the swap module 220 swaps out the target page to the swap partition 160, it may send the target page to the swap partition 160, that is, the swap module 220 itself completes swapping out the target page.
  • the swap module 220 may also instruct the acceleration device 150 of the host 10 to perform swapping out of the target page. For example, the swap module 220 sends a swap out instruction to the acceleration device 150 of the host 10, telling the acceleration device 150 to swap out the target page to the swap partition 160.
  • the swap out instruction carries the virtual address of the target page.
  • the swap module 220 can directly obtain the target page from the swap partition 160 and swap the acquired target page into the memory 120 of the host 10 .
  • the swap module 220 may also instruct the acceleration device 150 of the host 10 to perform page swapping. For example, the swap module 220 sends a swap command to the acceleration device 150 of the host 10, telling the acceleration device 150 to swap the target page into the memory 120 of the host 10.
  • the swap command carries the virtual address of the target page.
  • each functional module in the embodiment of the present application can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
  • the above integrated modules can be implemented in the form of hardware or software function modules.
  • the computing instance management unit 300 is used to manage the computing instance 400.
  • the computing instance management unit 300 can virtualize the underlying hardware 100 to provide a running environment for the computing instance 400.
  • the computing instance management unit 300 provides a running environment to the computing instance 400.
  • the computing instance management unit 300 can simulate the running environment for the computing instance 400 through software.
  • the computing instance 400 may pass the underlying hardware 100 (such as the memory 120) directly to the computing instance 400.
  • the computing instance management unit 300 When the computing instance 400 is a virtual machine, the computing instance management unit 300 includes QEMU (quick emulator) and a virtual machine monitor (VMM). When the computing instance 400 is a container, the computing instance management unit 300 can be a container engine.
  • QEMU quick emulator
  • VMM virtual machine monitor
  • the computing instance management unit 300 may transmit the page tag of the computing instance 400 to the host operating system 200 .
  • the computing instance management unit 300 may be built into the host operating system 200.
  • the host operating system 200 has the function of the computing instance management unit 300 . That is to say, the host operating system 200 can directly obtain the page tag of the computing instance 400 from the computing instance 400 .
  • the computing instance 400 has an independent operating environment.
  • the computing instance 400 occupies the physical resources of the host 10 and can perform various tasks and implement related services based on the occupied physical resources.
  • the computing instances 400 are independent of each other and do not affect each other.
  • the computing instance 400 may be a virtual machine, a container, or other modules formed by virtualization of the physical resources of the host 10 .
  • the computing instance 400 can evaluate the memory occupied by the computing instance 400 in the memory 120 of the host 10 Based on the importance of the page, a page label of the calculation instance 400 is generated.
  • the computing instance 400 may also send the page information of the computing instance 400 to the host operating system 200 .
  • the host 10 includes an I/O interface 130, a processor 110, a memory 120, and Acceleration device 150.
  • the I/O interface 130, the processor 110, the memory 120, and the acceleration device 150 may be connected through a system bus.
  • the system bus may be a peripheral component interconnect express (PCIe) bus or a computing bus.
  • PCIe peripheral component interconnect express
  • Fast interconnect compute express link, CXL
  • universal serial bus universal serial bus
  • USB universal serial bus
  • Figure 2 illustrates one of the connection methods.
  • the acceleration device 150 can be directly inserted into the card slot on the motherboard of the host 10, and exchanges data with the processor 110 through the PCIe bus 140.
  • the I/O interface 130 is used to communicate with devices located external to the host 10 . For example, data sent by a device other than the host 10 is received through the I/O interface 130 or data is sent to a device other than the host 10 through the I/O interface 130 .
  • the processor 110 is the computing core and control core of the host 10. It can be a central processing unit (CPU) or other specific integrated circuits.
  • the processor 110 can also be other general-purpose processors, digital signal processing (DSP), application specific integrated circuit (ASIC), field programmable gate array (field programmable gate array, FPGA) or other Programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
  • DSP digital signal processing
  • ASIC application specific integrated circuit
  • FPGA field programmable gate array
  • Programmable logic devices discrete gate or transistor logic devices, discrete hardware components, etc.
  • the memory 120 is generally used to store computer program instructions related to the host operating system 200, the computer program instructions related to the computing instance management unit 300, and the computing instance 400, and during the running process of the host operating system 200, the computing instance management unit 300, and the computing instance 400. the data generated.
  • the memory 120 has the advantage of fast access speed.
  • the memory 120 usually uses dynamic random access memory (DRAM).
  • DRAM dynamic random access memory
  • the memory 120 can also be other random access memories, such as static random access memory (Static random access memory, SRAM), storage class memory (storage class memory, SCM), etc.
  • the memory 120 may also be a read only memory (ROM).
  • read-only memory can be programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), etc.
  • the memory 120 may also be a dual in-line memory module (dual in-line memory module, DIMM), flash memory medium (FLASH), hard disk drive (hard disk drive, HDD) or solid state drive (solid state disk). , SSD), etc.
  • the processor 110 is connected to the memory 120 through a double data rate (DDR) bus or other types of buses.
  • DDR double data rate
  • the memory 120 is understood as the memory 120 (internal memory) of the host 10, and the memory 120 is also called main memory (main memory).
  • the processor 110 can form software modules such as the host operating system 200, the computing instance management unit 300, and the computing instance 400 in the host 10 by calling computer program instructions in the memory 120. As the computing core and control core of the host 10 , the processor 110 can support the steps executed by the host operating system 200 and the computing instance 400 in the embodiment shown in FIG. 3 .
  • the acceleration device 150 is connected to the host 10 , and can be used as an external device of the host 10 ; the acceleration device 150 can also be deployed inside the host 10 , for example, the acceleration device 150 is located on the motherboard or backplane of the host 10 .
  • FIG. 2 is a schematic diagram of the acceleration device 150 deployed inside the host 10 .
  • the acceleration device 150 can be used as a module with a data processing function attached to the host 10 and assume part of the functions of the host 10 . That is to say, part of the functions of the host 10 are offloaded to the acceleration device 150, and the acceleration device 150 replaces the host.
  • the machine 10 (such as the processor 110 in the host 10) processes data and performs some tasks to reduce the pressure on the processor 110 in the host 10 and release the computing power of the processor 110.
  • the acceleration device 150 can realize data access to the swap partition 160.
  • the data access here includes swapping out pages from the memory 120 of the host 10 to the swap partition 160, and swapping in pages from the swap partition 160 to the swap partition 160.
  • Host 10 has memory 120.
  • the acceleration device 150 may receive a swap-out instruction sent by the host operating system 200.
  • the swap-out instruction carries the virtual address of the page and the page (the page may be the target page in the embodiment shown in FIG. 3).
  • the swap out instruction also includes the identification of the process (the process may be the computing instance 400).
  • the acceleration device 150 sends the page to the storage device where the swap partition 160 is located, so as to save the page. It is stored in the swap partition 160 and saves the corresponding relationship between the virtual address of the page (optionally, also including the identity of the process) and the address of the page in the swap partition 160 .
  • the acceleration device 150 may receive a swap-in instruction sent by the MMU, which instructs to swap a page (the page may be a target page in the embodiment shown in FIG. 4 ) into the memory 120 of the host 10.
  • the acceleration device 150 swaps the page from the swap partition 160 into the memory 120 of the host 10 according to the swap-in instruction.
  • the swap-in instruction received by the acceleration device 150 carries the virtual address of the page.
  • the swap-out instruction also includes the identifier of the process (the process may be the computing instance 400).
  • the acceleration device 150 determines the address of the page in the swap partition 160 based on the virtual address of the page, obtains the page from the swap partition 160 based on the address of the page in the swap partition 160, and writes the page into the memory 120.
  • the acceleration device 150 also has data decompression and data compression functions. For example, the acceleration device 150 may first compress the pages that need to be swapped out to the swap partition 160 (that is, the data in the pages), and then swap the compressed pages into the swap partition 160 . For another example, the acceleration device 150 can decompress the compressed pages obtained from the swap partition 160 and swap the pages into the memory 120 of the host 10 .
  • the swap partition 160 is located in the storage device 20 outside the memory 120 , and the storage device 20 may be a remote storage device 30 located outside the host 10 .
  • the storage device 20 may also be the local storage device 170 of the host 10 .
  • the storage device 20 may also be located in a cloud, such as a public cloud or a private cloud.
  • the swap partition 160 is a storage space pre-allocated by the cloud for the host 10 .
  • the so-called “remote storage device 30” refers to a storage device connected to the host 10 through a network and located outside the host 10.
  • the remote storage device 30 can store data.
  • the embodiment of the present application is not limited to the type of the remote storage device 30.
  • the storage medium of the remote storage device 30 can be volatile memory (volatile memory), such as RAM, DRAM, SCM. , SRAM.
  • the storage medium of the remote storage device 30 may also be non-volatile memory (non-volatile memory), such as ROM, flash memory, HDD, SSD, SCM, etc.
  • the so-called “local storage device 170” refers to the storage device used for persistent storage in the host 10.
  • the local storage device 170 is connected to the host 10 through the system bus.
  • the local storage device 170 may be a non-volatile memory such as ROM, flash memory, HDD, SSD, etc.
  • the so-called “storage device located in the cloud” means that the storage device 20 is deployed in the cloud data center.
  • the storage device 20 can be understood as the storage space allocated by the cloud to the host.
  • the storage device 20 can be a certain device in the cloud data center, that is, the device can provide storage space as the swap partition 160.
  • the storage device 20 can also be multiple devices in the cloud data center, that is, the multiple devices jointly provide storage space as the swap partition 160.
  • the access method of the acceleration device 150 to the swap partition 160 is not limited.
  • the acceleration device 150 can access through the remote direct memory access (RDMA) 120 ) to access the remote storage device 30.
  • RDMA remote direct memory access
  • the swap partition 160 is located on the remote storage device 30 and the remote storage
  • the storage medium of the device 30 is a non-volatile memory such as HDD or SSD, or the swap partition 160 is located in the local storage device 170.
  • the acceleration device 150 can be based on Internet small computer system interface (iSCSI) or fiber channel (fibre). Network protocols such as Fiber Channel (FC) or Fiber Channel over Ethernet (FCoE) access the switching partition 160.
  • iSCSI Internet small computer system interface
  • FC Fiber Channel
  • FCoE Fiber Channel over Ethernet
  • the acceleration device 150 may provide a key-value (KV) interface to the processor 110 or the MMU in the host 10, and the KV interface allows the acceleration device 150 to interact with the processor 110, or the acceleration device 150 to interact with the MMU in a KV structure.
  • KV key-value
  • the key (key, K) in the key-value pair may be a virtual address of a page, or a combination of a virtual address of a page and an identifier of a process.
  • the value (value, K) in the key-value pair may be the page itself.
  • the swap command sent by the MMU to the acceleration device 150 can be in the form of a KV.
  • the acceleration device 150 receives the virtual address including the page's virtual address (or a combination of the page's virtual address and the process identifier) through the KV interface.
  • the acceleration device 150 may consider that the page needs to be swapped out to the swap partition 160 .
  • the swap out instruction sent by the processor 110 of the host 10 (which can also be understood as the host operating system 200) to the acceleration device 150 can also be expressed in the form of a KV.
  • the acceleration device 150 receives the virtual address (or page) including the page at the KV interface. When the combination of the virtual address plus the identifier of the process) is the key and the value is blank or invalid data, the acceleration device 150 can therefore consider that the page needs to be swapped from the swap partition 160 to the memory 120 of the host 10 .
  • the acceleration device 150 includes a processor, which may be a data processing unit (DPU) 151, an image processor (graphics processing unit, GPU), a tensor processing unit (TPU), or a neural network processor. (neural network processing unit, NPU) and other processors with data processing functions.
  • the processor included in the acceleration device 150 is the DPU 151 as an example.
  • the acceleration device 150 also includes a memory 152, a power supply circuit, etc.
  • the DPU 151 and the memory 152 are connected through a system bus.
  • the system bus can be a PCIe-based line, or a CXL, USB protocol, or other protocol bus.
  • DPU 151 is the main computing unit of the acceleration device 150 and the core unit of the acceleration device 150. DPU 151 undertakes the main functions of the acceleration device 150. In the embodiment of the present application, DPU 151 can access data of the swap partition 160 and can also perform data compression and decompression.
  • a page swapping method and a page swapping method provided by the embodiment of the present application will be described below in conjunction with Figures 3 and 4. As shown in Figure 3, a page swapping method provided by the embodiment of the present application is shown in Figure 4. The illustrated embodiment of the present application provides a page switching method.
  • the page swapping method shown in Figure 3 and the page swapping method shown in Figure 4 exist independently.
  • the page swapping out method shown in Figure 3 and the page swapping in method shown in Figure 4 can be used in combination. That is to say, when it is necessary to swap out the pages in the memory 120 of the host 10 to the swap partition 160 (that is, to swap out the pages in the memory 120 of the host 10 to the storage device 20), the method shown in Figure 3 can be used.
  • the page swapping method shown below If the pages in the swap partition 160 need to be swapped into the memory 120 of the host 10 later (that is, the pages in the storage device 20 need to be swapped into the memory 120 of the host 10), the page shown in Figure 4 can be used. Swap in method.
  • the page swapping out method shown in Figure 3 and the page swapping in method shown in Figure 4 can also be used independently. That is to say, when the page swapping method shown in Figure 3 is used to swap out the page, other methods can also be used to swap in the page. For other methods of swapping in the page, please refer to the aforementioned description of the swap mechanism. Before using the page swapping method shown in Figure 4 to swap in the page, other methods can also be used to swap out the page. For other methods of swapping out the page, please refer to the aforementioned description of the swapping mechanism.
  • a page swapping method provided in an embodiment of the present application includes:
  • Step 301 The computing instance 400 classifies the pages occupied by the computing instance 400 in the memory 120 of the host 10, and generates page tags for the computing instance 400.
  • the page label of the computing instance 400 indicates the memory 120 of the host 10 Calculate the importance of each page occupied by instance 400.
  • the computing instance 400 can classify the pages occupied by the computing instance 400 according to the importance of the pages.
  • the computing instance 400 can classify the pages occupied by the computing instance 400 from the perspective of the objects occupying the pages. For example, the computing instance 400 can classify the pages occupied by the operating system (such as guestOS) of the computing instance 400 into one category, and consider that the pages of this category have a higher importance; the computing instance 400 can classify the pages occupied by the applications on the computing instance 400 into one category, and consider that the pages of this category have a lower importance.
  • the computing instance 400 can classify the pages occupied by the computing instance 400 from the calling frequency (also called the access frequency) of the pages.
  • the computing instance 400 can classify the pages frequently called by the computing instance 400 (such as pages with a calling frequency greater than a first frequency threshold) into one category, and consider that the pages of this category have a higher importance; the computing instance 400 can classify the pages that are not frequently called by the computing instance 400 (such as pages with a calling frequency less than the first frequency threshold) into one category, and consider that the pages of this category have a lower importance.
  • the embodiment of the present application does not limit the manner in which the computing instance 400 classifies the pages occupied by the computing instance 400 and the evaluation criteria for the importance of the pages.
  • the computing instance 400 After the computing instance 400 classifies the pages occupied by the computing instance 400 in the memory 120 of the host 10, it can set an importance value for each type of page according to the importance of each type of page, and calculate the importance of the page label page of the instance 400. degree value.
  • the page tag of the computing instance 400 also includes the virtual address of the page (or a combination of the virtual address of the page and the identification of the computing instance 400).
  • the page label of the computing instance 400 includes the category to which the page belongs.
  • the page tag of the computing instance 400 also includes the virtual address of the page (or a combination of the virtual address of the page and the identification of the computing instance 400).
  • the categories of pages occupied by the virtual machine may include pages occupied by the operating system (such as guestOS) of the virtual machine, pages on the virtual machine, The page occupied by the application (application, APP) and the virtual machine management page.
  • the category of pages occupied by the operating system (such as guestOS) of the computing instance 400 can be identified as kernel (kernel), and the pages occupied by applications on the virtual machine can be identified as APP.
  • the page occupied by the virtual machine operating system is a page that stores data or computer program instructions related to the virtual machine operating system.
  • the virtual machine management page refers to a page that stores virtual machine management data (such as virtual machine attribute information, virtual machine identification, etc.).
  • the page occupied by an application on a virtual machine stores application-related data or computer program instructions. page.
  • the categories of pages occupied by the container may include pages occupied by applications on the container and container management pages.
  • the pages occupied by the container's operating system are pages that store data related to the container's operating system or computer program instructions.
  • the container management page refers to a page that stores container management data (such as container attribute information, container identification, etc.), and the pages occupied by applications on the container are pages that store data related to the application or computer program instructions.
  • the categories of pages occupied by the computing instance may include cold pages and hot pages.
  • Cold pages are pages with a low degree of access (eg, the access frequency of the page is lower than a certain value)
  • hot pages are pages with a high degree of access (eg, the access frequency of the page is greater than a certain value).
  • Step 302 The computing instance 400 sends the page tag of the computing instance 400 to the host operating system 200.
  • the page tag may be received by the instance page identification module 210 in the host operating system 200 .
  • the computing instance 400 can directly pass the page label to the host operating system 200, or it can manage the page through the computing instance 400.
  • the processing unit 300 passes the page tag to the host operating system 200.
  • the virtual address of the page in the page tag received by the host operating system 200 should be a virtual address that the host operating system 200 can recognize.
  • the page virtual address used may be the guest virtual address (GVA) configured by guestOS for the page.
  • the PT corresponding to the computing instance 400 obtained by the MMU includes two page tables.
  • One page table is the guest page table (GPT), which records the guest virtual address and the guest physical address (GPA).
  • the other page table is the extended page table (EPT), which records the correspondence between the client physical address and the host physical address (HPA).
  • the MMU can query the GPT and EPT based on this GPV. , determine the HPA.
  • the virtual address of the page that the host operating system 200 can recognize is the host virtual address (host virtual address, HVA).
  • the computing instance management unit 300 can convert GPA into HVA, and the virtual address of the page carried in the page tag initially generated by the computing instance 400 can be GPA (GPT is stored internally in the computing instance 400, and when generating the page tag internally, the computing instance 400 can Convert the GVA of the page to GPA by itself).
  • the computing instance management unit 300 can convert the GPA in the page label to HVA, update the page label, and update the page label.
  • the page tag is sent to the host operating system 200. In this way, the virtual address of the page in the page tag received by the host operating system 200 is the HVA that the host operating system 200 can recognize.
  • the computing instance 400 can periodically or in real time evaluate and classify the pages occupied by the computing instance 400 to determine the importance of the pages occupied by the computing instance 400, and then generate the page tags.
  • the computing instance 400 may periodically send the page tag to the host operating system 200 .
  • the computing instance 400 may also send the changed page label or the page label generated after the importance change to the host operating system 200 when the page label changes (or the importance of the page occupied by the computing instance 400 changes).
  • Step 303 The host operating system 200 obtains the page access frequency of each page occupied by the computing instance 400. This step may be performed by the instance page identification module 210.
  • the MMU every time the MMU queries a PTE in the PT, it will set the flag bit in the PTE to 1. If the PTE is not accessed within a period of time, the flag bit will return to 0.
  • the host operating system 200 may periodically scan the flag bits of each PTE in the PT corresponding to the computing instance 400 to determine the page access frequency of each page. For any PTE in the PT corresponding to the computing instance 400, the host operating system 200 checks the flag bit of the PTE when it reaches a monitoring time point. If the host operating system 200 finds that the flag bit of the PTE is 1, it records the page corresponding to the PTE. The number of accesses is incremented by one, and the host operating system 200 sets the flag bit to 0. If the host operating system 200 finds that the flag bit of the PTE is 0, it is considered that the page corresponding to the PTE has not been accessed, and the number of accesses to the page corresponding to the PTE remains unchanged. By resending the above operation within a period of time, the host operating system 200 can obtain the number of visits to the page corresponding to the PTE within the period of time, and thereby determine the page access frequency of the page.
  • Step 304 When the host operating system 200 needs to swap out a page to the swap partition 160, it filters the target page from the pages occupied by the computing instance 400.
  • the target page is a page that is allowed to be swapped out among the pages occupied by the computing instance 400. This step may be performed by the instance page identification module 210.
  • the host operating system 200 needs to swap pages to the swap partition 160.
  • the embodiment of the present application does not limit the specific scenario in which the host operating system 200 needs to swap pages to the swap partition 160. For example, when the host operating system 200 receives an exception signal sent by the MMU to indicate a page fault interrupt and determines that there is no free memory space in the memory 120 of the host 10 or that the free memory space is less than a preset free threshold, It may be determined that pages in the memory 120 of the host 10 need to be swapped out to the swap partition 160 . Another example.
  • Host operating system 200 may also predict the possible presence of host 10 in the future In the case of insufficient memory space, for example, the host operating system 200 detects that a new computing instance 400 needs to be created, or for example, the guest OS in the computing instance 400 needs to be updated. Host operating system 200 may determine that pages in memory 120 of host 10 need to be swapped out to swap partition 160 .
  • the host operating system 200 can also provide a parameter configuration interface for users, allowing users to configure page exchange parameters.
  • the page exchange parameters indicate the filter conditions of the target page.
  • the page swap parameters include but are not limited to: the identification of the computing instance 400 that is allowed to swap out the page, the page scanning frequency, the cold page determination threshold, and the maximum memory space that the computing instance 400 is allowed to occupy.
  • the host operating system 200 may consider all computing instances 400 of the host 10 Both allow swapping out pages.
  • the host operating system 200 can periodically scan the flag bits of each PTE in the PT corresponding to the computing instance 400.
  • the frequency of scanning each PTE in the PT corresponding to the computing instance 400 is the page scanning frequency. Excessive page scanning frequency will cause the host operating system 200 to frequently scan the corresponding PT of the computing instance 400, which will increase the burden of the host operating system 200. If the scanning frequency is too small, the host operating system 200 will not be able to accurately monitor the number of times the PTE flag position is 1, making the access frequency of the ultimately obtained page less accurate.
  • the host operating system 200 allows users to configure the page scanning frequency according to their own experience or needs.
  • the host operating system 200 can scan the flag bits of each PTE in the PT corresponding to the computing instance 400 at the preset page scanning frequency. .
  • Pages whose access frequency is less than a certain threshold are usually called cold pages, and this threshold is the cold page determination threshold.
  • Cold pages are accessed less frequently and can usually be swapped out to the swap partition 160 as pages that need to be swapped out.
  • the host operating system 200 can determine whether the page is a cold page based on a default threshold, or can provide the user with an interface to configure the threshold, and the user can configure the threshold.
  • the host operating system 200 may be preset with a maximum memory space of 120 that the computing instance 400 is allowed to occupy, and the memory space of 120 allocated by the host operating system 200 to the computing instance 400 is not allowed to exceed the maximum memory space of 120 .
  • the user can also configure the maximum memory space allowed to be occupied by the computing instance 400, and provide the user with an interface for configuring the maximum memory space allowed to be occupied by the computing instance 400. After the user configures the maximum memory space 120 that the computing instance 400 is allowed to occupy, Finally, the memory space 120 allocated by the host operating system 200 to the computing instance 400 is not allowed to exceed the maximum memory space 120 configured by the user.
  • the host operating system 200 can display the configuration interface of the page exchange parameters to the user when the computing instance 400 is created or during the operation of the computing instance 400, so that the user can complete the configuration of the page exchange parameters in the configuration interface of the page exchange parameters.
  • the host operating system 200 can obtain three types of information: page tag, page access frequency, and page exchange parameters of the computing instance 400 .
  • the host operating system 200 can perform various operations on the computing instance 400 based on one or more of these three types of information. Sort the occupied pages and select the target page from them.
  • the host operating system 200 can sort the pages occupied by the computing instance 400 based on one or more of the three types of information. There are many ways to select the target page, which are not limited by this application.
  • the host operating system 200 can sort the pages occupied by the computing instance 400 according to the importance of the pages occupied by the computing instance 400 indicated by the page tags, with pages with higher importance being ranked higher.
  • the host operating system 200 uses the N pages at the last sorted position as target pages, where N is a preset positive integer.
  • the target pages include pages occupied by applications on the computing instance 400 and may also include cold pages of the computing instance 400 .
  • the host operating system 200 can sort the pages occupied by the computing instance 400 in ascending order of access frequency according to the access frequency of each page occupied by the computing instance 400 indicated by the page access frequency.
  • the host operating system 200 uses the M pages at the last sorted position as target pages, where M is a preset positive integer.
  • the host operating system 200 can sort the pages occupied by the computing instance 400 according to the importance of the pages occupied by the computing instance 400 indicated by the page tags, with pages with higher importance ranked higher. For pages with the same importance level, the host operating system 200 can sort them in descending order according to the access frequency of each page with the importance level in the page access frequency. The host operating system 200 will sort the P at the end. Pages are used as target pages, and P is a preset positive integer.
  • the host operating system 200 can sort the pages occupied by the computing instance 400 according to page tags and page access frequencies, and then sort the last Q cold pages according to the cold page judgment threshold in the page exchange parameters.
  • Q is a preset positive integer.
  • the host operating system 200 can sort the pages occupied by the computing instance 400 according to the page label and page access frequency, and then sort the pages occupied by the computing instance 400 at the bottom according to the maximum memory space allowed to be occupied by the computing instance 400 in the page exchange parameters.
  • H pages are used as target pages, and H is a positive integer.
  • the size of the memory space of the remaining pages except the H pages is equal to the maximum memory space allowed to be occupied by the computing instance 400, or the computing instance
  • the size of the memory space of the remaining pages except the H pages among the pages occupied by 400 is less than the maximum memory space of 120 allowed to be occupied by the computing instance 400, and the difference between the two is within the preset range.
  • not all pages in the memory 120 of the host 10 are occupied by the computing instance 400.
  • the host operating system 200 can also select pages from the memory 120 of the host 10 other than the pages occupied by the computing instance 400 as pages that need to be swapped out.
  • step 305 only the pages occupied by the computing instance 400 are used.
  • the page is used as the target page for explanation.
  • the host operating system 200 can also select the pages that need to be swapped out.
  • the page is swapped out to the swap partition 160 in a similar manner to step 305.
  • Step 305 The host operating system 200 (such as the swap module 220 in the host operating system 200) swaps out the target page from the memory 120 of the host 10 to the swap partition 160.
  • the host operating system 200 updates the PT corresponding to the computing instance 400, and deletes the PTE that records the virtual address of the target page in the PT corresponding to the computing instance 400.
  • the host operating system 200 may directly send the target page in the memory 120 of the host 10 to the swap partition 160.
  • the host operating system 200 may send a first data request to the storage device 20 where the swap partition 160 is located.
  • the first data request is used to request that the target page be stored in the storage device 20 .
  • the first data request carries the target. page, and the address of the target page in the swap partition 160 set by the host operating system 200 (for convenience of explanation, the address of the target page in the swap partition 160 or the address of the target page in the storage device 20 is called the target page. exchange address,), the host operating system 200 can record the virtual address of the target page The corresponding relationship with the exchange address of the target page.
  • the corresponding relationship recorded by the host operating system 200 may be the corresponding relationship between the combination of the virtual address of the target page plus the identification of the computing instance 400 and the exchange address of the target page.
  • the corresponding relationship between the virtual address of the target page and the swap address of the target page that appears below all has this designated meaning, but for convenience of explanation, the relationship between the virtual address of the target page and the swap address of the target page is still used. Expression of corresponding relationships.
  • the corresponding relationship can be the corresponding relationship between the virtual address of the target page and the exchange address of the target page (for example, in a scenario where the virtual addresses used by different processes or computing instances 400 on the host 10 are completely different) , the corresponding relationship can be a combination of the virtual address of the target page plus the identification of the computing instance 400, and the corresponding relationship between the exchange address of the target page (such as the virtual address used by different processes or computing instances 400 on the host 10 Possibly under the same scenario).
  • the host operating system 200 may maintain a page exchange table for the swapped out pages.
  • the page exchange table includes multiple entries. Each entry corresponds to a page that is swapped out to the swap partition 160. The entry records the virtual address of the page. The corresponding relationship with the swap address of the page (the swap address of the page is the address of the page in the swap partition 160).
  • the host operating system 200 swaps out the target page to the swap partition 160, the host operating system 200 adds an entry corresponding to the target page in the page swap table for recording the virtual address of the target page and the swap address of the target page. correspondence between them.
  • the storage device 20 After receiving the first data request, the storage device 20 allocates a storage location for the target page, saves the target page in the storage location, and records the correspondence between the exchange address of the target page and the storage location. .
  • the target page may also be swapped out from the memory 120 of the host 10 to the swap partition 160 through the acceleration device 150 of the host 10 .
  • the host operating system 200 sends a swap out instruction to the acceleration device 150 of the host 10.
  • the swap out instruction instructs the acceleration device 150 to swap out the target page to the swap partition 160.
  • the swap out instruction carries the virtual address of the target page and
  • the target page optionally, also carries the identifier of the computing instance 400 in the swap out instruction.
  • the acceleration device 150 After receiving the swap out instruction, the acceleration device 150 sends a second data request to the storage device 20 where the swap partition 160 is located.
  • the second data request is used to request that the target page be stored in the storage device 20.
  • the second data request is used to request that the target page be stored in the storage device 20.
  • the data request carries the target page and the swap address of the target page configured by the acceleration device 150.
  • the acceleration device 150 records the correspondence between the virtual address of the target page and the swap address of the target page.
  • the acceleration device 150 can also maintain a page exchange table for the swapped out pages.
  • the page exchange table includes multiple entries, each entry corresponds to a page that is swapped out to the swap partition 160, and the entry records the page. The corresponding relationship between the virtual address and the swap address of the page.
  • the acceleration device 150 swaps out the target page to the swap partition 160, the acceleration device 150 adds an entry corresponding to the target page in the page swap table for recording the virtual address of the target page and the swap address of the target page. correspondence between.
  • the storage device 20 After receiving the second data request, the storage device 20 allocates a storage location for the target page, saves the target page in the storage location, and records the correspondence between the exchange address of the target page and the storage location. .
  • the acceleration device 150 may directly swap out the target page to the swap partition 160.
  • the acceleration device 150 may also first compress the target page and swap out the compressed target page to the swap partition 160.
  • a page switching method is provided in an embodiment of the present application.
  • a scenario in which the computing instance 400 calls a target page is used as an example for explanation. It should be noted that in the scenario where other process calls of the host 10 are swapped out to the swap partition 160, the embodiment shown in FIG. 4 is also applicable.
  • the computing instance 400 is understood to be other pages of the host 10.
  • the target page is understood to be the page that is swapped out to the swap partition 160 by other processes.
  • the method includes:
  • Step 401 The computing instance 400 initiates an access request, and the access request carries the virtual address of the target page. This access request is used to request the data in the target page.
  • Step 402 The MMU obtains the access request and queries the relevant PTE in the PT corresponding to the computing instance 400.
  • the MMU can query the PT corresponding to the computing instance 400. Since the target page has been swapped out to the swap partition 160 before, the PT does not have the PTE corresponding to the target page (that is, the target page is recorded). PTE of the virtual address), the MMU will find that the page table entry is missing in the PT.
  • the MMU will trigger the following two processes.
  • One process is the page remapping process of the host operating system 200 (see steps 403 to 405). In this process, the host operating system 200 needs to reconfigure the memory space for the data in the target page, that is, the data in the target page. The data searches for a blank page (ie, a free page) in the memory 120 of the host 10 .
  • Another process is the page swapping process of the acceleration device 150 (see steps 406 to 409). In this process, the acceleration device 150 obtains the target page from the swap partition 160 and writes the target page into the memory 120 of the host 10 middle. The two processes can proceed simultaneously.
  • Process 1 Page remapping process of the host operating system 200 (see steps 403 to 405).
  • Step 403 The MMU sends an exception signal indicating a page fault interrupt to the host operating system 200.
  • Step 404 The host operating system 200 allocates memory space for the target page and updates the PT corresponding to the computing instance 400.
  • the host operating system 200 allocates memory space for the target page in the memory 120 of the host 10, and the physical address of the memory space is the new physical address of the target page.
  • the host operating system 200 adds a PTE to the PT corresponding to the computing instance 400, and the PTE records the correspondence between the virtual address of the target page and the new physical address of the target page.
  • Step 405 The host operating system 200 notifies the acceleration device 150 that the page allocation is complete, that is, the host operating system 200 notifies the acceleration device 150 that the storage space has been reallocated for the target page in the memory 120 of the host 10, and notifies the acceleration device 150 of the new physical address of the target page.
  • the host operating system 200 sends a notification message through the system bus connected to the DPU 151.
  • the notification message indicates that the storage space for the target page has been reallocated in the memory 120 of the host 10.
  • the notification message also carries the new information of the target page. physical address.
  • a shared memory can be set up between the host operating system 200 and the acceleration device 150 .
  • the shared memory can be located in the memory 120 of the host 10 or in the memory 152 of the acceleration device 150 .
  • the shared memory 120 refers to a storage space that can be used by both the host operating system 200 and the acceleration device 150 .
  • An allocation completion queue (completion queue) jointly maintained by the host operating system 200 and the acceleration device 150 is stored in the shared memory.
  • the host operating system 200 reallocates the memory space for the target page in the memory 120 of the host 10.
  • a completion instruction can be added to the completion queue.
  • the completion instruction is used to indicate The page allocation is completed, and the completion indication also carries the new physical address of the target page.
  • Process 2 Page switching process of the acceleration device 150 (see steps 406 to 409).
  • Step 406 The MMU sends a swap instruction to the acceleration device 150.
  • the swap instruction carries the virtual address of the target page.
  • the swap command also carries the identifier of the computing instance 400 .
  • Step 407 The acceleration device 150 determines the swap address of the target page according to the virtual address of the target page.
  • the acceleration device 150 may obtain the corresponding relationship between the virtual address of the target page and the swap address of the target page, and determine the swap address corresponding to the virtual address of the target page.
  • the corresponding relationship between the virtual address of the target page and the swap address of the target page may be the host operating system 200 recorded, as recorded by the host operating system 200 in step 305 of the embodiment shown in FIG. 3 .
  • the acceleration device 150 may obtain the corresponding relationship between the virtual address of the page and the swap address of the page from the host operating system 200 .
  • the acceleration device 150 may obtain the page exchange table maintained by the host operating system 200 from the host operating system 200 .
  • the corresponding relationship between the virtual address of the target page and the exchange address of the target page may also be recorded by the acceleration device 150, as recorded by the acceleration device 150 in step 305 of the embodiment shown in FIG. 3 .
  • the acceleration device 150 queries the page exchange table for the corresponding relationship between the virtual address of the target page and the exchange address of the target page.
  • Step 408 The acceleration device 150 obtains the target page from the swap partition 160 according to the swap address of the target page.
  • the acceleration device 150 sends a third data request to the storage device 20 where the swap partition 160 is located.
  • the third data request is used to request to obtain the target page from the storage device 20 .
  • the third data request carries the swap address of the target page.
  • the storage device 20 After receiving the third data request, the storage device 20 determines the storage location where the target page is stored based on the exchange address of the target page, reads the target page from the storage location, and feeds the target page back to the acceleration device 150 .
  • Step 409 After receiving the notification from the host operating system 200, the acceleration device 150 writes the target page into the memory 120 of the host 10 and notifies the MMU that the page swap is completed.
  • the acceleration device 150 can receive the notification message from the host operating system 200 and determine that the host operating system 200 has reallocated the memory space for the target page. The acceleration device 150 can pass the new physical address of the target page carried in the notification message. Direct memory access (DMA) writes the target page to the new physical address. After writing the target page to the new physical address, the acceleration device 150 notifies the MMU that the page swap is completed.
  • DMA Direct memory access
  • the acceleration device 150 can also check whether there is a completion instruction from the allocation completion queue of the shared memory. If there is a completion instruction in the allocation completion queue, the acceleration device 150 takes out the completion instruction and obtains the new physical address of the target page. The acceleration device 150 can write the target page to the new physical address through DMA according to the new physical address of the target page. After writing the target page to the new physical address, the acceleration device 150 notifies the MMU Page switching is completed. If there is no completion indication in the allocation completion queue, the acceleration device 150 may suspend writing the target page into the memory 120 of the host 10 until a completion indication is detected in the allocation completion queue.
  • the target page obtained by the acceleration device 150 from the swap partition 160 may be the target page itself, that is, the target page is not compressed. In this case, the acceleration device 150 can directly swap the target page into the memory 120 of the host 10 .
  • the target page obtained by the acceleration device 150 from the swap partition 160 is a compressed target page
  • the compression of the target page may be performed by the acceleration device 150 before the target page is swapped to the swap partition 160, or the target page may be compressed before being swapped to the swap partition 160.
  • the acceleration device 150 can first decompress the compressed target page, obtain the target page, and swap the target page into the memory 120 of the host 10 .
  • the target page is swapped from the swap partition 160 into the memory 120 of the host 10 .
  • the MMU After receiving the notification from the acceleration device 150, the MMU obtains the updated PT corresponding to the computing instance 400, searches for the PTE that records the virtual address of the target page, obtains the target page from the memory 120 of the host 10, and converts the target page Feedback to calculation instance 400.
  • the MMU searches for the PTE that records the virtual address of the target page from the updated PT, determines the new physical address of the target page, reads the target page from the new physical address, and feeds the target page back to the computing instance 400 .
  • the swap-out of the target page and the swap-out of the target page are aimed at the data in the target page.
  • the target page For convenience of expression, in the embodiment of the present application, they are all referred to as the target page.
  • the embodiment of the present application also provides a page switching device.
  • the input device is used to execute the method executed by the DPU 151 or the acceleration device 150 in the above method embodiment shown in FIG. 3 or FIG. 4. Relevant features can be found in the above method embodiment and will not be described again here.
  • the page switching device 500 includes a receiving module 501 and a switching module 502 .
  • a swap out module 503 may also be included.
  • the receiving module 501 is configured to receive a swap-in instruction, which is used to instruct the target page to be swapped into the memory of the host.
  • the swap-in module 502 is used to obtain a target page from a storage device, swap the target page into the memory, and notify the host that the target page has been swapped into the memory.
  • the receiving module 501 can receive a swap instruction sent by the MMU or processor in the host.
  • the swap module 502 knows that the host's MMU or processor target page has been swapped into the memory.
  • the swap module 502 when the swap module 502 obtains the target page from the storage device, it obtains the corresponding relationship between the virtual address of the target page and the exchange address of the target page.
  • the exchange address of the target page is the address of the target page in the storage device. address.
  • the swap-in module 502 determines the swap address of the target page according to the corresponding relationship, and obtains the target page from the storage device according to the swap address of the target page.
  • the corresponding relationship may be recorded when the page swapping device swaps out the target page from the memory to the storage device.
  • the corresponding relationship may also be obtained by the swap-in module 502 from the processor of the host.
  • the swap module 502 when swapping the target page into the memory, can receive a notification sent by the processor of the host, and the notification is used to instruct the processor to configure the physical address of the target page in the memory. After obtaining the physical address, the swap-in module 502 writes the target page to the physical address through DMA.
  • the receiving module 501 receives a swap instruction sent by a processor of the host, the swap instruction is used to instruct to swap out a target page in the host memory to a storage device, and the swap instruction includes a virtual address of the target page and the target page.
  • the swap module 503 swaps out the target page to the storage device and records the corresponding relationship.
  • the swap-in module 502 when the swap-in module 502 obtains the target page from the storage device, if it obtains the compressed target page from the storage device; the swap-in module 502 can decompress the compressed target page to obtain the target page. , and then swap the target page into memory.
  • the swap out module 503 may compress the target page and then swap out the target page to the storage device.
  • each functional module in the embodiment of the present application can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
  • the above integrated modules can be implemented in the form of hardware or software function modules.
  • the above embodiments may be implemented in whole or in part by software, hardware, firmware, or any other combination.
  • the above-described embodiments may be implemented in whole or in part in the form of a computer program product.
  • the computer program product includes one or more computer instructions.
  • the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part.
  • the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
  • the computer instructions may be stored in or transmitted from one computer-readable storage medium to another, e.g., the computer instructions may be transferred from a website, computer, server, or data center Transmission to another website, computer, server or data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means.
  • the computer-readable storage medium may be any available medium that a computer can access, or a data storage device such as a server or a data center that contains one or more sets of available media.
  • the available media may be magnetic media (eg, floppy disk, hard disk, tape), optical media (eg, DVD), or semiconductor media.
  • Semiconductor medium It can be a solid state drive (SSD).
  • embodiments of the present application may be provided as methods, systems, or computer program products. Accordingly, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment that combines software and hardware aspects. Furthermore, the present application may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) having computer-usable program code embodied therein.
  • computer-usable storage media including, but not limited to, disk storage, CD-ROM, optical storage, etc.
  • These computer program instructions may also be stored in a computer-readable memory that causes a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including the instruction means, the instructions
  • the device implements the functions specified in a process or processes of the flowchart and/or a block or blocks of the block diagram.
  • These computer program instructions may also be loaded onto a computer or other programmable data processing device, causing a series of operating steps to be performed on the computer or other programmable device to produce computer-implemented processing, thereby executing on the computer or other programmable device.
  • Instructions provide steps for implementing the functions specified in a process or processes of a flowchart diagram and/or a block or blocks of a block diagram.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Software Systems (AREA)
  • Memory System Of A Hierarchy Structure (AREA)

Abstract

一种页面换入方法以及装置,本申请中,主机中的DPU代替主机中的处理器将页面从存储设备换入至主机的内存中。主机中的DPU接收换入指令,换入指令用于指示将目标页面换入至主机的内存。DPU在接收到该换入指令后,从存储设备获取目标页面,将目标页面换入至内存,通知主机(如主机中的处理器或MMU)目标页面已换入到内存中。页面的换入操作可以由DPU执行,主机中的处理器不再需要将目标页面从存储设备换入到主机的内存中,有效释放处理器的算力,减少对页面换入操作对处理器的占用,进而保证了主机中业务类工作的执行效率。

Description

一种页面换入方法以及装置
相关申请的交叉引用
本申请要求在2022年09月20日提交中国专利局、申请号为202211142315.0、申请名称为“一种页面换入方法以及装置”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及通信技术领域,尤其涉及一种页面换入方法以及装置。
背景技术
为了对该主机的内存进行扩展,主机中的处理器可以将主机的内存中不活跃的页面换出到交换分区中,以此释放主机的内存的空间,从而实现对主机的内存扩展,该交换分区被认为是一种虚拟内存,该交换分区通常会位于该主机的硬盘中。
当主机中的处理器需要调用这些换出至交换分区中的页面时,处理器会从交换分区中获取页面,将该页面换入到主机的内存。
但页面从交换分区换入到主机的内存会占用处理器资源,使得处理器无法提供较多的算力支持主机中业务类的工作。
发明内容
本申请提供一种页面换入方法以及装置,用以减少页面换入操作对处理器的占用。
第一方面,本申请提供了一种页面换入方法,在该方法中与主机通信系统总线连接的加速装置可以代替主机中的处理器将页面从存储设备换入至主机的内存中。该方法中,主机中的加速装置接收换入指令,换入指令用于指示将目标页面换入至主机的内存。加速装置在接收到该换入指令后,从存储设备获取目标页面,将目标页面换入至内存,通知主机(如主机中的处理器或MMU)目标页面已换入到内存中。
通过上述方法,页面的换入操作可以由加速装置执行,主机中的处理器不再需要将目标页面从存储设备换入到主机的内存中,有效释放处理器的算力,减少对页面换入对处理器的占用,进而保证了主机中业务类工作的执行效率。
在一种可能的实施方式中,加速装置可以从主机中的处理器或主机中MMU接收该换入指令。例如当MMU发现PT中缺页时,可以向加速装置发送换入指令,加速装置将目标页面换入到主机的内存后,加速装置通知MMU目标页面已换入到内存中。又例如,当处理器确定内存中有足够的内存空间时,可以向加速装置发送换入指令,加速装置将目标页面换入到主机的内存后,加速装置通知处理器目标页面已换入到内存中。
通过上述方法,主机中的处理器以及MMU仅需发送换入指令,即可通过加速装置将目标页面换入到主机的内存中,无需处理器的参与,页面换入的方式更加简单、快捷。
在一种可能的实施方式中,加速装置在从存储设备获取目标页面时,先获取目标页面的虚拟地址与目标页面的交换地址的对应关系,目标页面的交换地址为目标页面在存储设备的地址。加速装置在获取该对应关系后,根据对应关系确定目标页面的交换地址。再根据目标页面的交换地址从存储设备获取目标页面。加速装置可以与存储设备进行通信,向存储设备 发送携带有目标页面的交换地址的数据请求,以请求获取该目标页面,存储设备在接收到该数据请求后,可以向该加速装置反馈该目标页面。
通过上述方法,加速装置获取对应关系后,能够利用该对应关系获取目标页面的交换地址,以便与存储设备进行通信请求获取该目标页面。
在一种可能的实施方式中,存储设备能够提供存储空间作为该主机的虚拟内存,该存储设备可以是部署在主机之外,与主机通过网络连接的存储设备。存储设备也可以是主机的本地存储设备。存储设备还可以是部署在云端,如该存储设备可以是私有云或公有云中为该主机预先分配的存储空间。
通过上述方法,存储设备的具体形态有多种,有效地扩展了页面换入方法所适用的场景。
在一种可能的实施方式中,对应关系可以是加速装置将目标页面从内存换出至存储设备时、记录的。对应关系也可以是加速装置从主机的处理器获取的,该对应关系可以是处理器将目标页面从内存换出至存储设备时、记录的。
通过上述方法,加速装置获取该对应关系的方式有很多种,适用于不同的场景。
在一种可能的实施方式中,主机还包括处理器,加速装置通过系统总线与处理器连接,加速装置将目标页面换入至内存时,可以接收处理器发送的通知,通知用于指示处理器在内存中为目标页面配置的物理地址;加速装置在获取该物理地址后,可以通过DMA将目标页面写入到物理地址。
通过上述方法,加速装置获取物理地址后可以利用DMA将目标页面写入到内存中,绕过了处理器,进一步减少了对处理器的占用。
在一种可能的实施方式中,加速装置接收换入指令之前,加速装置还可以执行目标页面的换出操作。例如,加速装置接收主机的处理器发送的换出指令,换出指令用于指示将主机内存中的目标页面换出至存储设备,换出指令包括目标页面的虚拟地址以及目标页面;加速装置在接收到该换出指令后,将目标页面换出至存储设备,记录对应关系。
通过上述方法,加速装置承担了页面换入以及页面换出的主要操作,能够减少对处理器的占用,保证主机的处理器能够提供足够的算力以支持业务类的工作。
在一种可能的实施方式中,加速装置还具备数据解压缩功能,加速装置从存储设备获取压缩后的目标页面时,加速装置对压缩后的目标页面解压,获得目标页面,之后,再将目标页面换入至主机的内存中。
通过上述方法,加速装置可自行对压缩后的目标页面进行解压,无需处理器参与,进一步释放了处理器的算力。
在一种可能的实施方式中,加速装置将目标页面换出至存储设备时,为了能够降低该目标页面所占用的存储空间,加速装置将目标页面压缩后,换出至存储设备。
第二方面,本申请还提供了一种页面换入装置,该页面换入装置具有实现上述第一方面的方法实例中行为的功能,有益效果可以参见第一方面的描述此处不再赘述。功能可以通过硬件实现,也可以通过硬件执行相应的软件实现。硬件或软件包括一个或多个与上述功能相对应的单元。在一个可能的设计中,装置的结构中包括接收模块(用于接收换入指令或换出指令)、以及换出模块(用于将目标页面换出至存储设备),可选的,还包括换入模块(用于将目标页面从存储设备换入至主机的内存),这些模块可以执行上述第一方面方法示例中DPU的相应功能,具体参见方法示例中的详细描述,此处不做赘述。
第三方面,本申请还提供了一种加速装置,该加速装置包括DPU、GPU、NPU、或TPU等具有处理功能的处理器。前述列举的处理器的类型仅是举例,本申请并不限定加速装置中 所包括的处理器的具体类型,下面仅是以加速装置所包括的处理器为DPU为例进行说明。可选的,还包括供电电路以及存储器。供电电路用于对DPU进行供电。
在一种可能的实施方式中,该DPU具有实现上述第一方面以及第一方面的各个可能的实现方式中的方法实例中DPU的行为的功能,有益效果可以参见第一方面的描述此处不再赘述。
在另一种可能的实施方式中,存储器中存储计算机程序指令,DPU与存储器耦合,DPU可调用该存储器中存储的计算机执行指令,执行上述第一方面以及第一方面的各个可能的实现方式中DPU所执行的方法。
第四方面,本申请实施例还提供了一种计算设备,该计算设备可以为第一方面中提及的主机,有益效果可以参见第一方面的描述此处不再赘述,该计算设备包括加速装置和内存。
加速装置接收换入指令,换入指令用于指示将目标页面换入至计算设备的内存,加速装置在接收到该换入指令后,加速装置从存储设备获取目标页面;将目标页面换入至内存,通知计算设备目标页面已换入到内存中。
在一种可能的实施方式中,加速装置从计算设备的处理器或MMU接收换入指令。
在一种可能的实施方式中,加速装置从存储设备获取目标页面时,先获取目标页面的虚拟地址与目标页面的交换地址的对应关系,目标页面的交换地址为目标页面在存储设备的地址。加速装置根据对应关系确定目标页面的交换地址。加速装置与存储设备进行通信,根据目标页面的交换地址从存储设备获取目标页面。
在一种可能的实施方式中,存储设备为部署在计算设备之外,与计算设备通过网络连接的存储设备,或为计算设备的本地存储设备,或为公有云或私有云中为计算设备分配的存储空间。
在一种可能的实施方式中,对应关系是加速装置将目标页面从内存换出至存储设备时、记录的。对应关系是也可以是加速装置从计算设备的处理器获取的,该对应关系可以是处理器将目标页面从内存换出至存储设备时、记录的。
在一种可能的实施方式中,计算设备还包括处理器,加速装置通过系统总线与处理器连接。
加速装置将目标页面换入至内存时,处理器可以向加速装置发送通知,该通知用于指示处理器在内存中为目标页面配置的物理地址。加速装置在接收到该通知后,通过DMA将目标页面写入到物理地址。
在一种可能的实施方式中,计算设备的处理器向加速装置发送的换出指令,换出指令用于指示将计算设备内存中的目标页面换出至存储设备,换出指令包括目标页面的虚拟地址以及目标页面。
加速装置接收换出指令后,将目标页面换出至存储设备,记录对应关系。
在一种可能的实施方式中,加速装置从存储设备获取压缩后的目标页面,对压缩后的目标页面解压,获得目标页面,将目标页面换入至内存中。
在一种可能的实施方式中,加速装置将目标页面换出至存储设备时,将目标页面压缩后,换出至存储设备。
第五方面,本申请提了一种数据处理系统,该数据处理系统包括主机和存储设备,存储设备通过内部总线或网络与主机通信,主机设置有加速装置。
主机向加速装置发送换出指令,换出指令中包含数据。加速装置在从换入指令获取该数据后,对数据进行压缩,将压缩后的数据保存在加速装置的内存中,或者保存在主机的内存中。加速装置通过对数据进行压缩,使得数据在加速装置在内存中或主机内存中占用的空间 减少,提升加速装置在内存中或主机内存的利用率。
在一种可能的实施方式中,压缩后的数据保存在加速装置的内存中,加速装置可以在加速装置的内存中空闲的内存空间小于第一空闲阈值、或压缩后的数据的访问频率低于阈值的情况下,将压缩后的数据迁移至存储设备,这样可以减少对加速装置中内存的占用。
在一种可能的实施方式中,压缩后的数据保存在主机的内存中,加速装置可以在主机的内存中空闲的内存空间小于第二空闲阈值、或压缩后的数据的访问频率低于阈值的情况下,将压缩后的数据迁移至加速装置的内存中,这样可以减少对主机中内存的占用。
在一种可能的实施方式中,主机向加速装置发送换入指令,换入指令指示获取数据。加速装置在接收到该换入指令后,获取压缩后的数据,对压缩后的数据进行解压,在解压后,将解压后的数据保存在主机的内存中。
在一种可能的实施方式中,若压缩后的数据已迁移至存储设备,加速装置从存储设备获取压缩后的数据,若压缩后的数据保存在加速装置中的内存或主机的内存中,加速装置可以加速装置中的内存或主机的内存获取该压缩后的数据。
第六方面,本申请还提供一种计算机可读存储介质,计算机可读存储介质中存储有指令,当其在计算机上运行时,使得计算机执行上述第一方面以及第一方面的各个可能的实现方式中的方法。
第七方面,本申请还提供一种包含指令的计算机程序产品,当其在计算机上运行时,使得计算机执行上述第一方面以及第一方面的各个可能的实现方式中的方法。
第八方面,本申请还提供一种计算机芯片,芯片与存储器相连,芯片用于读取并执行存储器中存储的软件程序,执行上述第一方面以及第一方面的各个可能的实现方式中的方法。
附图说明
图1为本申请提供的一种主机的虚拟化架构示意图;
图2为本申请提供的一种主机的结构示意图;
图3为本申请提供的一种页面换出方法示意图;
图4为本申请提供的一种页面换入方法示意图;
图5为本申请提供的一种页面换入装置的结构示意图。
具体实施方式
在对本申请实施例提供的一种页面换入方法以及装置介绍之前,对本申请涉及的一些概念进行说明:
(1)、虚拟化技术、计算实例、容器(container)、虚拟机(virtual machine,VM)。
虚拟化(virtualization)是一种资源管理技术,通过虚拟化技术将主机的各种物理资源,如处理器、内存及接口等,予以抽象、转换后呈现出来。虚拟化是一种从逻辑角度出发的资源配置方法,是物理资源的逻辑抽象。
主机借助虚拟化技术能够在主机上形成具备独立的运行环境的软件模块,在本申请实施例中将主机上形成的具备独立的运行环境的软件模块称为计算实例。该计算实例可以为虚拟机,也可以为容器。
虚拟机是通过虚拟机化技术模拟的具有完整硬件系统功能的、运行在一个完全隔离环境中的“完整计算机”。实体计算机中能够完成的工作在虚拟机中都能够实现。虚拟机具备处理器(该处理器也被称为虚拟处理器)、内存、硬盘等组件,虚拟机的处理器、内存、硬盘等组 件均是由主机的处理器、内存、硬盘等组件虚拟而成的。虚拟机上安装有操作系统,虚拟机上的操作系统与主机自身的操作系统相互独立。为了区分这两种不同的操作系统,主机的上的操作系统通常被称为主机操作系统(host operating system,hostOS),虚拟机上的操作系统通常被称为客户操作系统(guest operating system,guestOS)。
容器是通过虚拟化技术模拟的一种独立的运行环境,容器类似于一个轻量级的沙盒,实现对容器外部的软件以及硬件进行屏蔽,容器是在操作系统层面上实现虚拟化,直接复用主机的操作系统。
对主机来说,计算实例会被看做是一种特殊的“进程”。该“进程”会执行计算任务、还会占用主机的处理器、内存、硬盘等资源。
举例来说,主机操作系统会为计算实例配置专属于该计算实例的内存空间,计算实例会占用该内存空间,以用于支持计算实例所需执行的计算任务。
(2)、内存管理单元(memory management unit,MMU)。
MMU也被称为分页内存管理单元(paged memory management unit,PMMU)。主机中,MMU是一个位于处理器的内核和连接高速缓存以及内存的总线之间的硬件装置,MMU通常也被认为是处理器的一部分,这种情况下MMU所执行的操作可以认为是处理器所执行的操作。在本申请实施例中,仅是以MMU是独立的硬件装置为例进行说明。MMU主要用于处理处理器发起的针对内存的访问请求,MMU具备地址翻译功能,能够实现内存的虚拟地址与内存的物理地址之间的转换。例如,MMU能够解析处理器发出的访问请求,将该访问请求中携带的内存的虚拟地址转换为内存的物理地址,以将该访问请求中携带的数据写入到该内存的物理地址上(访问请求用于请求在该内存的虚拟地址上写入数据);或者从该内存的物理地址上读取数据,将数据反馈给处理器(访问请求用于请求读取该内存的虚拟地址上的数据)。除此之外,MMU还具备内存保护功能、以及针对处理器的缓存的控制功能。
MMU可以以页为粒度对内存的空间进行管理(也即页式内存管理)。在以页为粒度对内存的空间的管理中,将内存空间划分为大小固定的块,一块即为一个页面。MMU可以以页面为单元分配、管理、以及保护内存。
为了实现虚拟地址到物理地址的转换,MMU需要获取地址翻译表,该地址翻译表称为页表(page table,PT),页表中包括一个或多个页表项(page table entry,PTE)。每个页表项对应一个页面,记录了该页面的虚拟地址与该页面的物理地址的对应关系以及一些控制信息。该地址翻译表是由主机操作系统维护并保存在主机的内存中的,也就是说,主机操作系统能够对该PT执行如增加页表项、删除页表项、修改页表项等的一些更新操作,并将更新后的页表保存在主机的内存中。MMU在需要进行地址翻译时,从主机的内存中调取该PT。
主机上运行的不同进程所使用的页面的虚拟地址可能相同,为了避免不同进程所使用的虚拟地址之间混淆。主机中,主机操作系统为主机上的每个进程(这里的进程也包括计算实例)维护对应的PT。
以计算实例为例,该计算实例对应的PT中记录了该计算实例所占用主机的内存空间的虚拟地址与物理地址之间的映射关系。如表1所示,为本申请实施例提供的一种PT的抽象表格,表1中将PT抽象为一个表格,该表格的每一行指示了一个PTE。每一行记录了一个页面的虚拟地址(virtual address,VA)、和该页面的物理地址(physical address,PA),该PTE中还包括一个标志位。该标志位用于标识该PTE是否被访问。在本申请实施例中以标识位为1,表征该PTE被访问,PTE被访问说明该PTE中的物理地址所指示的内存空间被访问;标 识位为0,表征为该PTE未被访问,PTE被访问说明该PTE中的物理地址所指示的内存空间未被访问。MMU每读取一次PTE,会将该PTE中的标志位置1,若标志位置1之后的特定时间内PTE未再次被读取,该PTE的标志位会自行翻转变为0。
表1
当计算实例需要将数据写入到虚拟地址V1,该计算实例会向MMU发送访问请求1(从硬件交互的角度来看,该访问请求1是由处理器发送给MMU的),该访问请求1用于请求在虚拟地址V1写入数据,该访问请求1还携带了页面的虚拟地址V1、以及需要写入的数据。MMU在接收到该访问请求1后,找到与该计算实例对应的PT,在该PT中查找是否存在该虚拟地址V1对应的物理地址。若在PT中未查找到记录了虚拟地址V1的PTE,MMU会触发缺页中断(page fault)流程,向主机操作系统发出用于指示缺页中断的异常信号,主机操作系统在内存中为计算实例分配一段空闲的内存空间(也即空闲页面),该内存空间的物理地址P1作为该虚拟地址V1对应的物理地址。主机操作系统更新PT,也就是说在该PT中增加PTE,该PTE记录该物理地址P1与该虚拟地址V1的对应关系。主机操作系统将更新后的PT保存在主机的内存中,通知MMU继续处理。MMU获取更新后的PT,查找记录了虚拟地址V1的PTE,确定该虚拟地址V1对应的物理地址P1,将该PTE中的标志位置1。MMU将该数据写入到该物理地址上P1。
当计算实例需要从虚拟地址V1读取数据,该计算实例会向MMU发送访问请求2(从硬件交互的角度看,该访问请求2是由处理器发送给MMU的),该访问请求2用于请求读取虚拟地址V1的数据,该访问请求2还携带了页面的虚拟地址V1。MMU在接收到该访问请求2后,找到与该计算实例对应的PT,在该PT中查找是否存在与该虚拟地址V1对应的物理地址。若在PT中查找到了记录了虚拟地址V1的PTE,MMU从与该虚拟地址对应的物理地址上读取数据,将数据反馈给该计算实例,并将标志位置1。若在PT中未查找到记录了虚拟地址V1的PTE,说明该数据未存储在主机的内存中。这种情况存在两种可能的情况:情况一、在该主机中存在交换机制,部署有交换模块的场景下,该MMU可以触发缺页中断流程,通过交换模块获取该虚拟地址V1上的数据,MMU通过交换模块获取该虚拟地址V1上的数据的过程可以参见下述关于交换机制的相关描述,此处不再赘述。情况二、在该主机中不存在交换机制,MMU会触发缺页中断流程,向主机操作系统发出用于指示缺页中断的异常信号,交由主机操作系统继续处理。
(3)、交换(swap)机制、交换分区(swap area)、交换模块。
交换机制指的是内存交换,交换机制中提出了物理内存以及虚拟内存这两个概念,物理内存为该主机的实际内存,在本申请实施例中也可以称为主机的内存。该虚拟内存即为交换分区,通常位于主机的本地存储设备中,主机的本地存储设备指与主机通过系统总线连接的存储设备,如主机的硬盘等。在本申请实施例中该虚拟内存也可以位于远端存储设备中,远端存储设备是指与主机通过网络连接的存储设备,如远端存储系统中的存储节点等,该虚拟 内存也可以部署在私有云或公有云中,也即也可以位于云数据中心中,例如,该虚拟内存是云数据中心中为该主机预先分配的存储空间,该存储空间为分布在一个或多个设备上。也即该虚拟内存可以部署在云端,虚拟内存所在的设备位于云端。交换机制中,允许从主机的存储设备(这里的存储设备包括主机的本地存储设备、远端存储设备以及云中的设备)中开辟出一部分存储空间,将该部分存储空间虚拟为主机的内存空间,作为虚拟内存。当主机的内存空间不足时,在物理内存和虚拟内存之间进行数据交换。
为了方便说明交换机制的实施方式,这里将实施该交换机制的执行方称为交换模块,该交换模块实质上是主机上运行一种软件模块,是主机操作系统中的一部分。
交换机制的实施方式如下:
交换模块在主机的物理内存中空闲空间较少(如物理内存中空闲的内存空间大小低于预设的阈值),或者交换模块预期后续可能存在物理内存中空闲内存空间不足的情况,交换模块可以将物理内存中一些页面换出到虚拟内存中,将这些页面中的数据保存到虚拟内存中,以释放物理内存的一些内存空间,这个过程即为页面的换出(swap out),这些换出的页面可以称为换出页面。在这种过程中,由于换出页面不再保存到物理内存中,PT中关于换出页面的PTE(也即记录了换出页面的虚拟地址与换出页面的物理地址之前的对应关系的PTE)也会被删除。而交换模块需要记录该换出页面的虚拟地址与该换出页面在虚拟内存中的地址的对应关系。
交换模块在主机需要调用这些页面(如MMU接收到主机上的进程发起的访问请求2,该访问请求携带了虚拟地址为换出页面的虚拟地址),或者主机的物理内存存在足够多的空闲空间(如物理内存中空闲的内存空间大小大于预设的阈值)的情况下,交换模块可以将换出页面从虚拟内存换入至物理内存中,这个过程即为页面的换入(swap in),这些换入的页面可以称为换入页面。在这种过程中,由于换入页面保存到了物理内存中,主机操作系统需要在物理内存中为换入页面分配新的物理地址(该物理地址称为换入页面的物理地址),MMU使用的PT中会增加关于换入页面的PTE(也即记录了换入页面的虚拟地址与换入页面的物理地址之前的对应关系的PTE)。而交换模块也会删除之前记录该换入页面的虚拟地址与该换入页面在虚拟内存中的地址的对应关系。
下面结合MMU以及交换模块的工作流程,对常见的针对计算实例发起的用于读取页面A的访问请求的处理流程进行进一步说明,由于交换模块通常被认为是主机操作系统的一部分,这该处理流程的说明中,不再区分交换模块以及主机操作系统,统一称为主机操作系统。
计算实例发起访问请求,该访问请求用于请求读取页面A中的数据,该访问请求携带了页面A的虚拟地址V1。MMU获取该访问请求,从内存中获取该计算实例对应的PT。MMU从PT查找记录了该虚拟地址V1的PTE。
若MMU在该PT找到了记录了该虚拟地址V1的PTE,MMU确定该虚拟地址V1对应的物理地址P1,将该物理地址P1发送给内存,将内存返回的、该物理地址P1上的数据反馈给计算实例。
若MMU在该PT未找到了记录了该虚拟地址V1的PTE,MMU会触发page fault流程,向主机操作系统发出缺页中断的异常信号,主机操作系统会暂停该计算实例的运行。主机操作系统查找在将页面A换出时、记录的该虚拟地址V1与页面A中的数据在虚拟内存中的地址X1的对应关系,根据该地址X1将该页面A中的数据从虚拟内存换入到主机的内存中、并为该页面A重新分配一段空闲的内存空间。主机操作系统为该页面重新分配的空闲的内存空间的物理地址P2,该物理地址P2作为该虚拟地址V1对应的物理地址,更新PT。也就是 说,主机操作系统在该PT中增加PTE,该PTE记录该物理地址P2与该虚拟地址V1的对应关系。主机操作系统PT将更新后的PT保存在主机的内存中,通知MMU继续处理,并启动该计算实例的运行。MMU获取更新后的PT,查找记录了虚拟地址V1的PTE,确定该虚拟地址V1对应的物理地址P2,将该PTE中的标志位置1。MMU将该物理地址P2发送至主机的内存,将从主机的内存反馈的数据反馈至计算实例。
在主机应用虚拟化技术、部署计算实例的场景下,计算实例能够占用主机的内存空间(也即占用内存中的一些页面),但计算实例(尤其是虚拟机)对所占用的内存空间的使用情况,主机操作系统是无法感知的。
在页面换出时,主机操作系统会决定内存中的哪一个或哪几个页面会被换出。常见的方式是,主机操作系统会利用最近最少使用(least recently used,LRU)算法,确定内存中页面最近最少使用的一个或多个页面,将所确定的该一个或多个页面换出到虚拟内存中。
但这种方式中,由于主机操作系统无法获取计算实例(如虚拟机)对该计算实例所占用的页面的使用情况,如主机操作系统无法获取计算实例所占用的各个页面是否为计算实例的内核占用的页面、或计算实例中应用所占用的页面,主机操作系统也无法获知计算实例对所占用的页面的访问频率。鉴于此,导致主机操作系统会将计算实例一些较为重要的页面作为换出页面,换出到虚拟内存中,当计算实例需要调用这些页面时,MMU会触发缺页中断流程,缺页中断流程伴随着页面的换入操作,增加时延,导致计算实例的卡顿,影响计算实例的工作效率。
为此本申请实施例提出了一种页面换出方法,该方法中,计算实例可以向主机操作系统发送页面标签,该页面标签指示了主机的内存中计算实例所占用的各个页面的重要程度。主机操作系统在需要将该内存中的页面换出至内存之外、部署有交换分区的存储设备时,主机操作系统根据该页面标签确定主机的内存中需要换出的目标页面,将所确定的目标页面换出至该存储设备。在本申请实施例中,由于计算实例能够将所占用的页面的重要程度告知主机操作系统,使得主机操作系统能够避免将计算实例的占用的重要程度较高的页面换出到存储设备。计算实例在需要调用重要程度较高的页面时,MMU不会触发缺页中断流程,计算实例也不会被暂停运行。
另外,为了减少主机操作系统的压力,主机操作系统在将页面换出至该存储设备时,主机操作系统向主机的加速装置发送换出指令,该换出指令用于指示加速装置将页面换出至该存储设备,该换出指令携带了页面的虚拟地址以及该页面。加速装置获取该页面,将该页面存储至该存储设备,并记录该页面的虚拟地址与该页面在该存储设备中的地址的对应关系。也即是说,加速装置可以代替主机操作系统(也即处理器)将页面从主机的内存中交换至该存储设备中,能够减少对处理器的占用,释放处理器的算力。
在将该内存中的页面换出至该存储设备后,可能会存在需要将该存储设备中的页面换入至主机的内存的情况。本申请实施例中提出了一种页面换入方法,在本申请实施例中,主机的加速装置可以代替主机中的处理器(也可以理解为主机操作系统)将页面从该存储设备换入至主机的内存。主机中的MMU向该加速装置发送换入指令,该换入指令用于指示将页面换入至内存中,该换入指令中可以携带有页面的虚拟地址,加速装置在接收到该换入指令后,可以根据该换入指令将页面从存储设备换入至主机的内存中,并通知MMU页面已换入至主机的内存中。例如,加速装置可以根据已记录的对应关系确定该页面在该存储设备中的地址,根据该页面在该存储设备中的地址从该存储设备获取该页面,将该页面写入到主机的内存中。 这样可以进一步减少对处理器的占用。
在本申请实施例中,页面换出至内存之外的存储设备是指将页面换出至交换分区,其中,该存储设备是部署有交换分区的存储设备。也就是说,将页面换出至内存之外的存储设备以及将页面换出至交换分区的两种表述方式所表达的实质是相同的,在下文中,统一为将页面换出至交换分区。但应需理解,将页面换出至交换分区是指将页面交换到存储设备中,将页面存储在该存储设备中。类似的,将页面从存储设备换入至主机的内存是指将交换分区中的页面换入至主机的内存中,页面从存储设备换入至主机的内存以及页面从交换分区中换入至主机的内存的两种表述方式所表达的实质是相同的,在下文中,统一为页面从交换分区换入至主机的内存。也就是说,为了表述上更加明确,在下文中涉及到页面的换入以及换出的过程中用交换分区指代该交换分区所在的存储设备。
如图1所示,为本申请实施例提供的一种主机的虚拟化架构示意图。将虚拟化技术应用于主机10,形成了如图1所示的虚拟化架构,主机10的虚拟化架构包括底层硬件100、主机操作系统200、计算实例管理单元300、以及至少一个计算实例400。主机10为一种计算设备。该主机10可以为服务器、移动终端、平板电脑等计算设备。
底层硬件100是指该主机10中的一些硬件组件,如处理器110、内存120、输入输出(input output,I/O)接口130等。底层硬件100可以理解为主机10上的物理资源,是虚拟化技术中需要虚拟的对象。
以底层硬件100为基础,该主机10运行了主机操作系统200、计算实例管理单元300、计算实例400等软件模块。
主机操作系统200运行在处理器110上,用于实现主机10的基础功能;基础功能包括但不限于:针对底层硬件100的管理功能、针对与主机10连接的输入/输出设备(如显示器、键盘、鼠标等)的控制功能、对主机10中进程的管理。本申请实施例中主要涉及针对底层硬件100的管理功能,以主机操作系统200对主机10的内存120的管理为例,主机操作系统200能够监控内存120的占用情况,分配内存空间;主机操作系统200还能够检测内存120中页面的访问频率;主机操作系统200能够实现主机10的内存120与交换分区160之间的页面交换。
在本申请实施例,主机操作系统200能够从计算实例400获取计算实例400的页面标签,该页面标签指示了主机10的内存120中计算实例400所占用的页面的重要程度。主机操作系统200还能够获取计算实例400所占用的各个页面的页面访问频率。主机操作系统200在需要将主机10的内存120的页面换出到交换分区160时,主机操作系统200可以根据页面标签、页面访问频率以及用户配置的页面交换参数中的部分或全部从主机10的内页中确定目标页面。主机操作系统200将所确定的目标页面换出至交换分区160。主机操作系统200也可以将被换出至交换分区160的页面换入至主机10的内存120中。
从逻辑功能上,该主机操作系统200包括实例页面识别模块210、交换模块220。
实例页面识别模块210能够获取来自计算实例400的页面标签,还能够根据MMU对该计算实例400对应的PT的访问情况获得计算实例400占用的各个页面的页面访问频率。
实例页面识别模块210还能够面向用户提供参数配置接口,允许用户配置页面交换参数。页面交换参数指示了目标页面的筛选条件。该页面交换参数包括但不限于:允许换出页面的计算实例400的标识、页面扫描频率、冷页面的判断阈值、以及计算实例400允许占用的最大内存空间。
在需要将页面从主机10的内存120换出至交换分区160时,实例页面识别模块210根据 计算实例400的页面标签、页面访问频率以及页面交换参数中的部分或全部从计算实例400所占用的页面中筛选目标页面,该目标页面为该计算实例400所占用的页面中允许换出的页面。实例页面识别模块210向交换模块220发送目标页面的虚拟地址。
交换模块220用于实现主机10的内存120与交换分区160之间的页面交换。也即交换模块220能够将主机10的内存120中的页面换出至交换分区160、以及将交换分区160中的页面换入至主机10的内存120。
交换模块220在需要将主机10的内存120中的页面交换至交换分区160时,交换模块220将该目标页面换出到交换分区160。交换模块220还可以更新该计算实例400对应的PT,从中删除与该目标页面相关的PTE。
交换模块220将该目标页面换出到交换分区160时,可以向将该目标页面发送到交换分区160,也即交换模块220自身完成目标页面的换出。交换模块220也可以指示该主机10的加速装置150执行目标页面的换出。例如,交换模块220向该主机10的加速装置150发送换出指令,告知加速装置150将目标页面换出至交换分区160,该换出指令携带了目标页面的虚拟地址。
交换模块220在需要交换分区160中目标页面换入至主机10的内存120时,交换模块220可以直接从交换分区160获取目标页面,将获取的目标页面换入至主机10的内存120中。交换模块220也可以指示该主机10的加速装置150执行页面的换入。例如,交换模块220向该主机10的加速装置150发送换入指令,告知加速装置150将目标页面换入至主机10的内存120,该换入指令携带了目标页面的虚拟地址。
本申请实施例中对模块的划分是示意性的,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式。在本申请的实施例中的各功能模块可以集成在一个处理模块中,也可以是各个模块单独物理存在,也可以两个或两个以上模块集成在一个模块中。上述集成的模块既可以采用硬件的形式实现,也可以采用软件功能模块的形式实现。
计算实例管理单元300用于管理计算实例400,该计算实例管理单元300能够将底层硬件100进行虚拟化,以向计算实例400提供运行环境。计算实例管理单元300向计算实例400提供运行环境的方式有很多种,例如,计算实例管理单元300能够通过软件的方式为该计算实例400模拟运行环境。又例如,计算实例400可以将底层硬件100(如内存120)直通给计算实例400。
当计算实例400为虚拟机时,该计算实例管理单元300包括QEMU(quick emulator)以及虚拟机管理器(virtual machine monitor,VMM)。当计算实例400为容器时,该计算实例管理单元300可以为容器引擎(container engine)。
在本申请实施例中,计算实例管理单元300可以将计算实例400的页面标签传输给主机操作系统200。
在一些可能的场景中,计算实例管理单元300可以内置在主机操作系统200中。这种场景下,主机操作系统200具备了计算实例管理单元300的功能。也就是说,主机操作系统200可以直接从计算实例400获取该计算实例400的页面标签。
计算实例400具备独立的运行环境,计算实例400占用了主机10的物理资源,能够基于所占用的物理资源执行各种任务,实现相关业务。计算实例400之间相互独立,彼此不会影响。在本申请实施例中该计算实例400可以为虚拟机,也可以为容器,还可以是其他借助主机10的物理资源的虚拟化形成的模块。
在本申请实施例中,计算实例400能够评价主机10的内存120中该计算实例400所占用 的页面的重要程度,生成该计算实例400的页面标签。计算实例400还可以将该计算实例400的页面信息发送给主机操作系统200。
下面对主机10包括的硬件组件进行说明,如图2所示,为本申请实施例提供的一种主机10的结构示意图,主机10包括I/O接口130、处理器110、存储器120、以及加速装置150。I/O接口130、处理器110、存储器120、以及加速装置150之间可通过系统总线连接,该系统总线可以为快捷外围部件互连标准(peripheral component interconnect express,PCIe)总线,也可以为计算快速互联(compute express link,CXL)、通用串行总线(universal serial bus,USB)协议或其他协议的总线。
图2示例性的展示了其中一种连接方式,图2中,加速装置150可以直接插在主机10的主板上的卡槽中,通过PCIe总线140与处理器110交换数据。
I/O接口130用于与位于主机10外部的设备通信。例如,通过I/O接口130接收主机10之外的设备发送的数据或通过I/O接口130向主机10之外的设备发送数据。
处理器110是主机10的运算核心和控制核心,它可以是中央处理器(central processing unit,CPU),也可以是其他特定的集成电路。处理器110还可以是其他通用处理器、数字信号处理器(digital signal processing,DSP)、专用集成电路(application specific integrated circuit,ASIC)、现场可编程门阵列(field programmable gate array,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。
存储器120通常用来存放主机操作系统200相关的计算机程序指令、计算实例管理单元300以及计算实例400相关的计算机程序指令、以及主机操作系统200、计算实例管理单元300、计算实例400在运行过程中所产生的数据。存储器120具备访问速度快的优点。存储器120通常采用动态随机存取存储器(dynamic random access memory,DRAM)。除了DRAM之外,存储器120还可以是其他随机存取存储器,例如静态随机存取存储器(Static random access memory,SRAM)、存储级存储器(storage class memory,SCM)等。另外,存储器120也可以是只读存储器(read only memory,ROM)。而对于只读存储器,举例来说,可以是可编程只读存储器(programmable read only memory,PROM)、可抹除可编程只读存储器(erasable programmable read only memory,EPROM)等。存储器120还可以为双列直插式存储器模块或双线存储器模块(dual in-line memory module,DIMM)、闪存介质(FLASH)、硬盘驱动器(hard disk drive,HDD)或固态驱动器(solid state disk,SSD)等。
处理器110通过双倍速率(double data rate,DDR)总线或者其他类型的总线和存储器120相连。存储器120理解为主机10的内存120(internal memory),内存120又称为主存(main memory)。
处理器110可以通过调用该内存120中的计算机程序指令在主机10中形成主机操作系统200、计算实例管理单元300以及计算实例400等软件模块。处理器110作为主机10的运算核心以及控制核心能够支持如图3所示的实施例中主机操作系统200以及计算实例400所执行的步骤。
加速装置150与主机10连接,该加速装置150可以作为该主机10的外接设备;加速装置150也可以部署在该主机10内部,如加速装置150位于在主机10的主板或背板上。图2为加速装置150部署在主机10内部的一种示意图。
该加速装置150可以作为主机10附带的具有数据处理功能的模块,承担该主机10的部分功能。也就是说,主机10的部分功能卸载到该加速装置150上,由该加速装置150代替主 机10(如主机10中的处理器110)处理数据,执行部分的任务,以减轻主机10中处理器110的压力,释放该处理器110的算力。在本申请实施例中该加速装置150能够实现对交换分区160的数据访问,这里的数据访问包括将页面从主机10的内存120换出至交换分区160、以及将页面从交换分区160换入至主机10的内存120。
例如,加速装置150可以接收主机操作系统200发送的换出指令,该换出指令中携带了页面的虚拟地址以及该页面(该页面可以为图3所示实施例中的目标页面),可选的,该换出指令还包括进程的标识(该进程可以为计算实例400),加速装置150在接收到该换出指令后,将该页面发送至交换分区160所在的存储设备,以将该页面存储在该交换分区160,并保存页面的虚拟地址(可选的,还包括进程的标识)、与页面在交换分区160的地址之间的对应关系。
又例如,加速装置150可以接收MMU发送的换入指令,该换入指令指示将页面(该页面可以为图4所示实施例中的目标页面)换入至主机10的内存120中,加速装置150根据该换入指令,将页面从交换分区160中换入至主机10的内存120中。
加速装置150接收的换入指令中携带了页面的虚拟地址,可选的,该换出指令还包括进程的标识(该进程可以为计算实例400),加速装置150在接收到该换入出指令后,根据该页面的虚拟地址确定与页面在交换分区160的地址,根据页面在交换分区160的地址从交换分区160获取页面,将该页面写入到内存120中。
此外,该加速装置150还具备数据解压以及数据压缩功能。例如加速装置150可以对先对需要换出至交换分区160的页面(也即页面中的数据)进行压缩,将压缩后的页面换入至交换分区160。又例如,加速装置150可以对从交换分区160中获取的压缩后的页面进行解压,将页面换入至主机10的内存120中。
在本申请实施例中,该交换分区160位于内存120之外的存储设备20,该存储设备20可以为位于该主机10之外的远端存储设备30。该存储设备20也可以为该主机10的本地存储设备170。该存储设备20也可以位于云端,如位于公有云或私有云。该交换分区160为云端预先为该主机10分配的存储空间。
所谓“远端存储设备30”是指与主机10通过网络连接、位于主机10之外的存储设备。远端存储设备30能够存储数据,本申请实施例并不限远端存储设备30的类型,该远端存储设备30的存储介质可以为易失性存储器(volatile memory),例如RAM、DRAM、SCM、SRAM。该远端存储设备30的存储介质也可以为非易失性存储器(non-volatile memory),例如ROM、快闪存储器、HDD、SSD、SCM等。
所谓“本地存储设备170”是指主机10中、用于持久化存储的存储设备,本地存储设备170通过系统总线连接到主机10上。该本地存储设备170可以为ROM、快闪存储器、HDD、SSD等非易失性存储器。
所谓“位于云端的存储设备”,是指该存储设备20部署在云数据中心中。该存储设备20可以理解为云端为该主机分配的存储空间。该存储设备20可以为该云数据中心中的某一个设备,也即该设备能够提供存储空间以作为交换分区160。该存储设备20也可以为该云数据中心中的多个设备,也即该多个设备共同提供存储空间以作为交换分区160。
在本申请实施例中并不限定加速装置150对交换分区160的访问方式。例如,当该交换分区160位于远端存储设备30,该远端存储设备30的存储介质为DRAM、SCM等易失性存储器,加速装置150可以通过远程直接内存120访问(remote direct memory access,RDMA)的方式访问远端存储设备30。又例如,当交换分区160位于远端存储设备30且该远端存储 设备30的存储介质为HDD、SSD等非易失性存储器、或者交换分区160位于本地存储设备170,加速装置150可以基于互联网小型计算机系统接口(internet small computer system interface,iSCSI)、光纤通道(fibre channel,FC)、或以太网光纤通道(fibre channel over ethernet,FCoE)等网络协议访问交换分区160。
加速装置150可以面向主机10中的处理器110或MMU提供键值对(key value,KV)接口,该KV接口允许加速装置150与处理器110之间,或加速装置150与MMU之间以KV结构的数据结构进行数据交互。其中,键值对中的键(key,K)可以为页面的虚拟地址、或者页面的虚拟地址加进程的标识的组合。键值对中的值(value,K)可以为页面本身。
举例来说,MMU向加速装置150发送的换入指令,即可表现为KV的形式,加速装置150在该KV接口接收到包括页面的虚拟地址(或者页面的虚拟地址加进程的标识的组合)为键,页面为值的数据时,加速装置150由此可认为需要将该页面换出至交换分区160。
主机10的处理器110(也可以理解为主机操作系统200)向加速装置150发送的换出指令同样可以表示为KV的形式,加速装置150在该KV接口接收到包括页面的虚拟地址(或者页面的虚拟地址加进程的标识的组合)为键,值为空白或无效数据时,加速装置150由此可认为需要将该页面从交换分区160换入至主机10的内存120。
加速装置150包括处理器,该处理器可以为数据处理单元(data process unit,DPU)151图像处理器(graphics processing unit,GPU)、张量处理器(tensor processing unit,TPU)或神经网络处理器(neural network processing unit,NPU)等具有数据处理功能的处理器,在图2中仅是以加速装置150包括的处理器为DPU151为例进行说明。可选的,该加速装置150还包括存储器152、供电电路等,DPU151与存储器152通过系统总线连接,该系统总线可以为基于PCIe的线路,也可以为CXL、USB协议或其他协议的总线。
DPU151为加速装置150的主要运算单元,是加速装置150的核心单元,DPU151承担了加速装置150的主要功能。在本申请实施例中,DPU151能够对交换分区160进行数据访问,还能够进行数据压缩以及解压。
下面结合附图3以及图4对本申请实施例提供的一种页面换出方法以及页面换入方法进行说明,如图3所示的本申请实施例提供的一种页面换出方法,如图4所示的本申请实施例提供的一种页面换入方法。
需要说明的是,图3所示的页面换出方法以及图4的页面换入方法是独立存在的。图3所示的页面换出方法以及图4的页面换入方法可以结合使用。也就是说,在需要将主机10的内存120中的页面换出至交换分区160(也即是说需要将主机10的内存120中的页面换出至存储设备20)时,可以采用图3所示的页面换出方法。若后续需要将交换分区160中的页面换入至主机10的内存120(也即是说需要将存储设备20中的页面换入至主机10的内存120中),可以采用图4所示的页面换入方法。图3所示的页面换出方法以及图4的页面换入方法也可以独立使用。也就是说,当采用图3所示的页面换出方法进行页面换出时,也可以采用其他方式进行页面换入,其他页面换入的方式可以参见前述关于交换机制的描述。在采用图4所示的页面换入方法进行页面换入之前,也可以采用其他方式,进行页面换出,其他页面换出的方式可以参见前述关于交换机制的描述。
如图3所示,为本申请实施例提供的一种页面换出方法,该方法包括:
步骤301:计算实例400对主机10的内存120中该计算实例400所占用的页面进行分类,生成该计算实例400的页面标签。该计算实例400的页面标签指示了主机10的内存120中计 算实例400所占用的各个页面的重要程度。
计算实例400可以按照页面的重要程度对该计算实例400所占用的页面进行分类。计算实例400可以从占用页面的对象为角度,对该计算实例400所占用的页面进行分类。例如,计算实例400可以将被该计算实例400的操作系统(如guestOS)所占用的页面归为一类,认为该类页面具备较高的重要程度;计算实例400可以将被该计算实例400上的应用所占用的页面归为一类,认为该类页面具备较低的重要程度。计算实例400可以从页面的调用频率(也可以称为访问频率),对该计算实例400所占用的页面进行分类,该计算实例400可以将该计算实例400经常调用的页面(如调用频率大于第一频率阈值的页面)归为一类,认为该类页面具备较高的重要程度;该计算实例400可以将该计算实例400不经常调用的页面(如调用频率小于第一频率阈值的页面)归为一类,认为该类页面具备较低的重要程度。
本申请实施例并不限定该计算实例400对该计算实例400所占用的页面进行分类的方式以及页面的重要程度的评价标准。
计算实例400对主机10的内存120中该计算实例400所占用的页面进行分类后,可以按照各类页面的重要程度为每一类页面设置一个重要程度数值,计算实例400的页面标签页面的重要程度数值。可选的,该计算实例400的页面标签还包括页面的虚拟地址(或页面的虚拟地址与计算实例400的标识的组合)。
计算实例400对主机10的内存120中该计算实例400所占用的页面进行分类后,也可以直接将页面所属的类别作为该页面的重要程度,计算实例400的页面标签包括页面的所属的类别。可选的,该计算实例400的页面标签还包括页面的虚拟地址(或页面的虚拟地址与计算实例400的标识的组合)。
例如,从占用页面的对象的角度,当计算实例400为虚拟机时,该虚拟机所占用的页面的类别可以包括被该虚拟机的操作系统(如guestOS)所占用的页面、该虚拟机上的应用(application,APP)所占用的页面、虚拟机管理页面。被该计算实例400的操作系统(如guestOS)所占用的页面的类别为可以标识为内核(kernel),该虚拟机上的应用所占用的页面可以标识为APP。
其中,虚拟机的操作系统(如guestOS)所占用的页面是在存储有与虚拟机操作系统相关数据或计算机程序指令的页面。虚拟机管理页面是指存储有虚拟机管理数据(如虚拟机的属性信息、虚拟机的标识等)的页面,虚拟机上的应用所占用的页面是存储有与应用相关数据或计算机程序指令的页面。
类似的,当计算实例400为容器时,该容器所占用的页面的类别可以包括该容器上的应用所占用的页面、容器管理页面。
其中,容器的操作系统(如guestOS)所占用的页面是在存储有与容器操作系统相关数据或计算机程序指令的页面。容器管理页面是指存储有容器管理数据(如容器的属性信息、容器标识等)的页面,容器上的应用所占用的页面是存储有与应用相关数据或计算机程序指令的页面。
又例如,从页面的访问情况,该计算实例所占用的页面的类别可以包括冷页面、热页面。冷页面是页面的访问程度较低的页面(如页面的访问频率低于某一值),热页面是页面访问程度较高的页面(如页面的访问频率大于某一值)。
步骤302:计算实例400向主机操作系统200发送该计算实例400的页面标签。主机操作系统200中可以由实例页面识别模块210接收到页面标签。
计算实例400可以直接将该页面标签传递给主机操作系统200,也可以通过计算实例管 理单元300将该页面标签传递给主机操作系统200。
需要说明的是,该主机操作系统200所接收到的页面标签中页面的虚拟地址应为该主机操作系统200能够识别的虚拟地址。
举例来说,当计算实例400为虚拟机时,该计算实例400在访问页面时,所采用的页面虚拟地址可以是guestOS为该页面所配置的客户机虚拟地址(guest virtual address,GVA)。MMU获取的该计算实例400对应的PT包括两种页表,一个页表为客户端页表(guest page table,GPT),其中记录了客户端虚拟地址与客户端物理地址(guest physical address,GPA)之间的对应关系。另一个页表为扩展页表(extended page table,EPT),其中记录了客户端物理地址与主机10物理地址(host physical address,HPA)之间的对应关系,MMU可以根据该GPV查询GPT以及EPT,确定HPA。而对于主机操作系统200所能识别的页面的虚拟地址为主机虚拟地址(host virtual address,HVA)。计算实例管理单元300能够将GPA转换为HVA,计算实例400最初生成的页面标签中携带的页面的虚拟地址可以为GPA(计算实例400内部保存有GPT,计算实例400内部在生成页面标签时,可以自行将页面的GVA转换为GPA),计算实例400在经过计算实例管理单元300传输该页面标签时,计算实例管理单元300可以将该页面标签中的GPA转换为HVA,更新页面标签,将更新后的页面标签发送给主机操作系统200。这样主机操作系统200所接收的页面标签中页面的虚拟地址即为该主机操作系统200能够识别的HVA。
本申请实施例中计算实例400可以周期性或实时评估对该计算实例400所占用的页面进行分类,以确定该计算实例400所占用的页面的重要程度,进而生成该页面标签。计算实例400可以将该页面标签周期性的发送给主机操作系统200。计算实例400也可以在页面标签发生变化(或计算实例400所占用的页面的重要程度发生变化)的情况下,向主机操作系统200发送变化后的页面标签或重要程度变化后生成的页面标签。
步骤303:主机操作系统200获取该计算实例400所占的各个页面的页面访问频率。该步骤可以由实例页面识别模块210执行。
在前面关于MMU的描述可知,MMU每查询到PT中的一个PTE,会将该PTE中的标志位置1,若该PTE在之后一段时间内未被访问,标志位将恢复为0。
主机操作系统200可以周期性的扫描该计算实例400对应的PT中各个PTE的标志位,以确定该各个页面的页面访问频率。针对该计算实例400对应的PT中任一PTE,主机操作系统200到达一个监控时间点时查看PTE的标志位,若主机操作系统200发现PTE的标志位为1时,记录该PTE对应的页面的访问次数加一,主机操作系统200并将该标志位置0。若主机操作系统200发现PTE的标志位为0时,认为该PTE对应的页面未被访问,该PTE对应的页面的访问次数维持不变。在一段时间内重发执行上述操作,主机操作系统200能够获得该PTE对应的页面在该段时间内的访问次数,进而确定出该页面的页面访问频率。
步骤304:主机操作系统200在需要将页面换出至交换分区160时,从计算实例400所占用的页面中筛选目标页面。该目标页面为该计算实例400所占用的页面中允许换出的页面。该步骤可以由实例页面识别模块210执行。
主机操作系统200需要将页面换出至交换分区160的情况有很多种,本申请实施例并不限定主机操作系统200需要将页面交换至交换分区160的具体场景。例如,主机操作系统200在接收到MMU发出的、用于指示缺页中断的异常信号时,若确定主机10的内存120中无空闲的内存空间或空闲的内存空间小于预设的空闲阈值时,可以确定需要将主机10的内存120中的页面换出至交换分区160。又例如。主机操作系统200也可以预测未来可能存在主机10 的内存空间不足的情况,如主机操作系统200检查到需要创建新的计算实例400,又如,计算实例400中的guest OS需要进行版本更新。主机操作系统200可以确定需要将主机10的内存120中的页面换出至交换分区160。
主机操作系统200还能够面向用户提供参数配置接口,允许用户配置页面交换参数。页面交换参数指示了目标页面的筛选条件。该页面交换参数包括但不限于:允许换出页面的计算实例400的标识、页面扫描频率、冷页面的判断阈值、以及计算实例400允许占用的最大内存空间。
由于计算实例400所占的页面的换入换出总是在一些情况下影响计算实例400的工作效率,并不是所有计算实例400均允许所占用的页面被换出,总是存在一些不允许换出页面的计算实例400。例如,在主机10中一些承担重要业务的计算实例400,或者该主机10中由重要用户所租用的计算实例400。用户在主机操作系统200为用户提供的参数配置接口中配置允许换出页面的计算实例400的标识,或者不允许换出页面的计算实例400的标识。若用户未配置允许换出页面的计算实例400的标识或者主机操作系统200未向用户提供配置该允许换出页面的计算实例400的标识的接口,主机操作系统200可以认为主机10所有计算实例400均允许换出页面。
在步骤303中提及主机操作系统200能够周期性的扫描该计算实例400对应的PT中各个PTE的标志位,扫描该计算实例400对应的PT中各个PTE的频率即为页面扫描频率。过大的页面扫描频率会导致主机操作系统200需要频繁的扫描计算实例400的对应的PT,会增大主机操作系统200的负担。而过小的扫描频率会导致主机操作系统200不能准确的监控到PTE的标志位置1的次数,使得最终获得的页面的访问频率的准确性较差。主机操作系统200允许用户根据自身经验或者需求配置该页面扫描频率。若用户未配置页面扫描频率或者主机操作系统200未向用户提供配置该页面扫描频率的接口,主机操作系统200可以以预设的页面扫描频率扫描该计算实例400对应的PT中各个PTE的标志位。
页面的访问频率小于某一阈值的页面通常被称为冷页面,该阈值即为冷页面的判断阈值。冷页面的访问频率较低,通常能够作为需要换出的页面换出到交换分区160。主机操作系统200可以根据默认的阈值判断页面是否属于冷页面,也可以向用户提供配置该阈值的接口,由用户配置该阈值。
由于计算实例400会占用主机10内存120中的一些内存120空间,若计算实例400占用的内存120空间较大会影响其他计算实例400或主机操作系统200自身的运行。若计算实例400占用的内存120空间较小,该计算实例400的工作效率也会受到影响。主机操作系统200可以预设有计算实例400允许占用的最大内存120空间,主机操作系统200为该计算实例400所分配的内存120空间不允许超过该最大内存120空间。当然,也可以由用户配置该计算实例400允许占用的最大内存空间,向用户提供配置该计算实例400允许占用的最大内存空间的接口,在用户配置了该计算实例400允许占用的最大内存120空间后,主机操作系统200为该计算实例400所分配的内存120空间不允许超过用户配置的最大内存120空间。
应需理解的是,用户配置页面交换参数的时机应早于步骤304,如主机操作系统200可以在计算实例400创建时或计算实例400运行过程中,向用户显示该页面交换参数的配置界面,以使得用户能够在该页面交换参数的配置界面中,完成页面交换参数的配置。
主机操作系统200能够获取该计算实例400的页面标签、页面访问频率以及页面交换参数这三种信息。主机操作系统200能够根据这三种信息中的一种或多种,对计算实例400所 占用的页面进行排序,从中选出目标页面。
主机操作系统200能够根据这三种信息中的一种或多种,对计算实例400所占用的页面进行排序,从中选出目标页面的方式有很多种,本申请并不限定。
例如,主机操作系统200能够根据页面标签指示的该计算实例400所占用的各个页面的重要程度对计算实例400所占用的各个页面进行排序,重要程度较高的页面排在靠前的位置。主机操作系统200将排序位置在最后的N个页面作为目标页面,N为预设的正整数。以步骤301所列举的几种类别为例,目标页面包括计算实例400上应用所占用的页面,也可以包括计算实例400的冷页面。
又例如,主机操作系统200能够根据页面访问频率指示的该计算实例400所占用的各个页面的访问频率、按照访问频率由大到小顺序对计算实例400所占用的各个页面进行排序。主机操作系统200将排序位置在最后的M个页面作为目标页面,M为预设的正整数。
又例如,主机操作系统200能够根据页面标签指示的该计算实例400所占用的各个页面的重要程度对计算实例400所占用的各个页面进行排序,重要程度较高的页面排在靠前的位置,对于重要程度相同的页面,主机操作系统200能够根据页面访问频率中该重要程度下的各个页面的访问频率、按照访问频率由大到小顺序进行排序,主机操作系统200将排序位置在最后的P个页面作为目标页面,P为预设的正整数。
又例如,主机操作系统200能够根据页面标签以及页面访问频率对计算实例400所占用的各个页面进行排序后,根据页面交换参数中的冷页面的判断阈值,将排序位置在最后的Q个冷页面作为目标页面,Q为预设的正整数。
又例如,主机操作系统200能够根据页面标签以及页面访问频率对计算实例400所占用的各个页面进行排序后,根据页面交换参数中的计算实例400允许占用的最大内存空间,将排序位置在最后的H个页面作为目标页面,H为正整数,其中,计算实例400所占用的页面中除该H个页面外的剩余页面的内存空间的大小等于计算实例400允许占用的最大内存空间,或计算实例400所占用的页面中除该H个页面外的剩余页面的内存空间的大小小于计算实例400允许占用的最大内存120空间,这两者的差值处于预设范围。
在实际应用中,主机10的内存120中的页面并非均是由计算实例400所占用的,主机10的内存120中的页面还存在被主机操作系统200所占用的页面,以及该主机10上的应用所占用的页面。主机操作系统200也可以从主机10的内存120中除计算实例400所占用的页面之外的页面选出页面,作为需要被换出的页面,在下面步骤305中仅是以计算实例400所占用的页面作为目标页面为例进行说明,但对于主机操作系统200从主机10的内存120中除计算实例400所占用的页面之外的页面选出的需要换出的页面,主机操作系统200也可以采用与步骤305类似的方式将页面换出到交换分区160。
步骤305:主机操作系统200(如主机操作系统200中的交换模块220)将目标页面从主机10的内存120中换出至交换分区160。主机操作系统200更新该计算实例400对应的PT,删除该计算实例400对应的PT中记录该目标页面的虚拟地址的PTE。
主机操作系统200在执行该步骤305时,主机操作系统200可以直接将主机10的内存120中的目标页面发送到交换分区160。主机操作系统200可以向交换分区160所在的存储设备20发送第一数据请求,该第一数据请求用于请求将该目标页面存储在该存储设备20中,该第一数据请求中携带了该目标页面,以及主机操作系统200设置的该目标页面在交换分区160中的地址(为方便说明,该目标页面在交换分区160中的地址或该目标页面在存储设备20中的地址称为目标页面的交换地址,),主机操作系统200可以记录该目标页面的虚拟地址 与该目标页面的交换地址之间的对应关系。
应需理解的是,由于主机10上不同进程所使用的虚拟地址可能相同,为了区分不同进程。主机操作系统200记录的对应关系可以为该目标页面的虚拟地址加计算实例400的标识的组合、与该目标页面的交换地址之间的对应关系。下文出现的该目标页面的虚拟地址与该目标页面的交换地址之间的对应关系均具备该指定意义,但为了方便说明,仍采用该目标页面的虚拟地址与该目标页面的交换地址之间的对应关系的表述。也就是说,该对应关系可以为该目标页面的虚拟地址与该目标页面的交换地址之间的对应关系(如在主机10上不同进程或计算实例400所使用的虚拟地址完全不同的场景下),该对应关系可以为该目标页面的虚拟地址加计算实例400的标识的组合、与该目标页面的交换地址之间的对应关系(如在主机10上不同进程或计算实例400所使用的虚拟地址可能相同的场景下)。
主机操作系统200可以为换出的页面维护页面交换表,页面交换表包括多个表项,每个表项对应一个被换出至交换分区160的页面,该表项记录了该页面的虚拟地址与页面的交换地址(页面的交换地址即为该页面在交换分区160的地址)的对应关系。主机操作系统200在该目标页面换出至交换分区160时,主机操作系统200在页面交换表中增加与该目标页面对应的表项,用于记录目标页面的虚拟地址与该目标页面的交换地址之间的对应关系。
该存储设备20在接收到该第一数据请求后,为该目标页面分配存储位置,将该目标页面保存在该存储位置中,并记录该目标页面的交换地址与该存储位置之间的对应关系。
主机操作系统200在执行该步骤305时,也可以通过主机10的加速装置150将目标页面从主机10的内存120中换出至交换分区160。例如,主机操作系统200向该主机10的加速装置150发送换出指令,该换出指令中指示加速装置150将目标页面换出至交换分区160,该换出指令携带了目标页面的虚拟地址以及该目标页面,可选的,还换出指令中还携带了该计算实例400的标识。
加速装置150在接收到该换出指令后,向交换分区160所在的存储设备20发送第二数据请求,该第二数据请求用于请求将该目标页面存储在该存储设备20中,该第二数据请求中携带了该目标页面,以及加速装置150所配置的该目标页面的交换地址,加速装置150记录该目标页面的虚拟地址与该目标页面的交换地址之间的对应关系。
类似的,加速装置150也可以为换出的页面维护页面交换表,页面交换表包括多个表项,每个表项对应一个被换出至交换分区160的页面,该表项记录了该页面的虚拟地址与页面的交换地址的对应关系。加速装置150在将该目标页面换出至交换分区160时,加速装置150在页面交换表中增加与该目标页面对应的表项,用于记录目标页面的虚拟地址与该目标页面的交换地址之间的对应关系。
该存储设备20在接收到该第二数据请求后,为该目标页面分配存储位置,将该目标页面保存在该存储位置中,并记录该目标页面的交换地址与该存储位置之间的对应关系。
在由加速装置150将页面从主机10的内存120换出至交换分区160时,加速装置150可以直接将该目标页面换出至该交换分区160。加速装置150也可以先对该目标页面进行压缩,将压缩后的目标页面换出至该交换分区160。
至此,页面的换出流程结束,当计算实例400需要调用该目标页面时,会触发页面的换入流程,具体可以参见图4。
如图4所示,为本申请实施例提供的一种页面换入方法,在图4中以计算实例400调用目标页面的场景为例进行说明。需要说明的是,在该主机10的其他进程调用被换出至交换分区160的页面的场景,如图4所示的实施例也同样适用,将计算实例400理解为主机10的其 他进程,目标页面理解为该其他进程被换出至交换分区160的页面。该方法包括:
步骤401:计算实例400发起访问请求,该访问请求中携带了目标页面的虚拟地址。该访问请求用于请求获取该目标页面中的数据。
步骤402:MMU获取该访问请求,在该计算实例400对应的PT中查询相关的PTE。
MMU在获取该访问请求后,可以查询该计算实例400对应的PT,由于该目标页面之前已被换出至交换分区160,PT并不存在该目标页面对应的PTE(也即记录了该目标页面的虚拟地址的PTE),MMU会发现该PT中缺少页表项。
若MMU在计算实例400对应的PT中未查询到相关的PET,MMU会触发如下两个流程。一个流程为主机操作系统200的页面重映射流程(参见步骤403~步骤405),在该流程中主机操作系统200需要为该目标页面中的数据重新配置内存空间,也即为该目标页面中的数据在主机10的内存120中找寻空白页面(也即空闲页面)。另一个流程为加速装置150的页面换入流程(参见步骤406~步骤409),在该流程中,加速装置150从交换分区160中获取目标页面,并将目标页面写入到主机10的内存120中。这两个流程可以同步进行。
流程一、主机操作系统200的页面重映射流程(参见步骤403~步骤405)。
步骤403:MMU向主机操作系统200发送用于指示缺页中断的异常信号。
步骤404:主机操作系统200为该目标页面分配内存空间,更新该计算实例400对应的PT。
主机操作系统200在主机10的内存120为该目标页面分配内存空间,内存空间的物理地址即为该目标页面新的物理地址。
主机操作系统200在该计算实例400对应的PT中增加PTE,该PTE记录该目标页面的虚拟地址与该目标页面新的物理地址之间的对应关系。
步骤405:主机操作系统200通知加速装置150页面分配完成,也即主机操作系统200告知加速装置150已在该主机10的内存120中为该目标页面重新分配的存储空间,并告知该加速装置150该目标页面新的物理地址。
主机操作系统200通知加速装置150页面分配完成的方式有很多种。例如,主机操作系统200通过与DPU151连接的系统总线发送通知消息,该通知消息指示已在该主机10的内存120中为该目标页面重新分配存储空间,该通知消息还携带了该目标页面新的物理地址。
又例如,主机操作系统200与加速装置150之间可以设置共享内存,该共享内存可以位于主机10的内存120中,也可以位于该加速装置150的存储器152中。共享内存120是指主机操作系统200与加速装置150均能够使用的存储空间。在该共享内存中保存了由主机操作系统200与加速装置150共同维护的分配完成队列(completion queue)。主机操作系统200在该主机10的内存120中为该目标页面重新分配的内存空间,更新了该计算实例400对应的PT后,可在该完成队列中加入一个完成指示,该完成指示用于指示页面分配完成,该完成指示还携带有该目标页面新的物理地址。
流程二、加速装置150的页面换入流程(参见步骤406~步骤409)。
步骤406:MMU向加速装置150发送换入指令,该换入指令携带了目标页面的虚拟地址。可选的,该换入指令还携带了计算实例400的标识。
步骤407:加速装置150根据该目标页面的虚拟地址确定该目标页面的交换地址。
加速装置150可以获取目标页面的虚拟地址以及目标页面的交换地址的对应关系,确定与该目标页面的虚拟地址对应的交换地址。
其中,目标页面的虚拟地址以及目标页面的交换地址的对应关系可以是主机操作系统200 记录的,如图3所示的实施例的步骤305中主机操作系统200记录的。加速装置150可以从主机操作系统200获取该页面的虚拟地址以及页面的交换地址的对应关系。例如,加速装置150可以从主机操作系统200获取主机操作系统200维护的页面交换表。
目标页面的虚拟地址以及目标页面的交换地址的对应关系也可以是加速装置150记录的,如图3所示的实施例的步骤305中加速装置150记录的。例如,加速装置150从页面交换表查询目标页面的虚拟地址以及目标页面的交换地址的对应关系。
步骤408:加速装置150根据该目标页面的交换地址从该交换分区160获取该目标页面。
加速装置150向交换分区160所在的存储设备20发送第三数据请求,该第三数据请求用于请求从存储设备20获取该目标页面,该第三数据请求中携带了该目标页面的交换地址。
该存储设备20在接收到该第三数据请求后,根据该目标页面的交换地址确定存储该目标页面的存储位置,从存储位置读取该目标页面,将该目标页面反馈给加速装置150。
步骤409:加速装置150在接收到主机操作系统200的通知后,将该目标页面写入到该主机10的内存120中,通知MMU页面换入完成。
加速装置150可以接收该主机操作系统200的通知消息,确定该主机操作系统200已经为该目标页面重新分配的内存空间,加速装置150可以根据该通知消息中携带的目标页面新的物理地址、通过直接内存访问(direct memory access,DMA),将该目标页面写入到该新的物理地址处,加速装置150在将该目标页面写入到该新的物理地址后,通知MMU页面换入完成。
加速装置150也可以从共享内存的分配完成队列中查看是否有完成指示,若该分配完成队列中有完成指示,加速装置150将该完成指示取出,获取该目标页面新的物理地址。加速装置150可以根据该目标页面新的物理地址、通过DMA,将该目标页面写入到该新的物理地址处,加速装置150在将该目标页面写入到该新的物理地址后,通知MMU页面换入完成。若该分配完成队列中不存在完成指示,加速装置150可以暂缓将该目标页面写入到该主机10的内存120中,直至检测到该分配完成队列中出现完成指示。
在由加速装置150将页面从交换分区160换出至主机10的内存120时,当加速装置150从交换分区160获取的目标页面可以是该目标页面本身,也即该目标页面未被压缩,这种情况下加速装置150可以直接将该目标页面换入至主机10的内存120中。当加速装置150从交换分区160获取的目标页面是压缩后的目标页面,该目标页面的压缩可是在该目标页面被交换到交换分区160之前由加速装置150执行的,也可以是目标页面在被换出至交换分区160后,由交换分区160所在的存储设备20执行的。这种情况下加速装置150可以先对该压缩后的目标页面进行解压,获得目标页面,将目标页面换入至主机10的内存120中。
至此,目标页面从交换分区160换入到主机10的内存120中。
MMU在接收到加速装置150的通知后,获取更新后的、该计算实例400对应的PT,查找记录该目标页面的虚拟地址的PTE,从主机10的内存120中获取目标页面,将该目标页面反馈给计算实例400。
MMU从更新后的PT中查找记录该目标页面的虚拟地址的PTE,确定该目标页面新的物理地址,从该新的物理地址上读取该目标页面,将该目标页面反馈给计算实例400。
需要说明的是,在本申请实施例中目标页面的换出以及目标页面的换出,针对的是该目标页面中的数据,为了方便表述,在本申请实施例中均是以目标页面指代该目标页面中的数据。
基于与方法实施例同一发明构思,本申请实施例还提供了一种页面换入装置,该页面换 入装置用于执行上述如图3或图4所示的方法实施例中DPU151或加速装置150执行的方法,相关特征可参见上述方法实施例,此处不再赘述。如图5所示,页面换入装置500包括接收模块501以及换入模块502。可选的,还可以包括换出模块503。
接收模块501,用于接收换入指令,换入指令用于指示将目标页面换入至主机的内存。
换入模块502,用于从存储设备获取目标页面;将目标页面换入至内存,通知主机目标页面已换入到内存中。
在一种可能的实施方式中,接收模块501可以接收主机中的MMU或处理器发送的换入指令。相应的,换入模块502知主机的MMU或处理器目标页面已换入到内存中。
在一种可能的实施方式中,换入模块502在从存储设备获取目标页面时,获取目标页面的虚拟地址与目标页面的交换地址的对应关系,目标页面的交换地址为目标页面在存储设备的地址。换入模块502根据对应关系确定目标页面的交换地址,根据目标页面的交换地址从存储设备获取目标页面。
在一种可能的实施方式中,对应关系可以是页面换入装置将目标页面从内存换出至存储设备时、记录的。对应关系也可以是换入模块502从主机的处理器获取的。
在一种可能的实施方式中,换入模块502在将目标页面换入至内存时,可以接收主机的处理器发送的通知,通知用于指示处理器在内存中为目标页面配置的物理地址。换入模块502在获取该物理地址后,通过DMA将目标页面写入到物理地址。
一种可能的实施方式中,接收模块501接收主机的处理器发送的换出指令,换出指令用于指示将主机内存中的目标页面换出至存储设备,换出指令包括目标页面的虚拟地址以及目标页面。
换出模块503将目标页面换出至存储设备,记录对应关系。
在一种可能的实施方式中,换入模块502在从存储设备获取目标页面时,若从存储设备获取了压缩后的目标页面;换入模块502可以对压缩后的目标页面解压,获得目标页面,再将该目标页面换入到内存中。
在一种可能的实施方式中,换出模块503在将目标页面换出至存储设备时,可以将目标页面压缩后,换出至存储设备。
需要说明的是,本申请实施例中对模块的划分是示意性的,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式。在本申请的实施例中的各功能模块可以集成在一个处理模块中,也可以是各个模块单独物理存在,也可以两个或两个以上模块集成在一个模块中。上述集成的模块既可以采用硬件的形式实现,也可以采用软件功能模块的形式实现。
上述实施例,可以全部或部分地通过软件、硬件、固件或其他任意组合来实现。当使用软件实现时,上述实施例可以全部或部分地以计算机程序产品的形式实现。所述计算机程序产品包括一个或多个计算机指令。在计算机上加载或执行所述计算机程序指令时,全部或部分地产生按照本发明实施例所述的流程或功能。所述计算机可以为通用计算机、专用计算机、计算机网络、或者其他可编程装置。所述计算机指令可以存储在计算机可读存储介质中,或者从一个计算机可读存储介质向另一个计算机可读存储介质传输,例如,所述计算机指令可以从一个网站站点、计算机、服务器或数据中心通过有线(例如同轴电缆、光纤、数字用户线(DSL))或无线(例如红外、无线、微波等)方式向另一个网站站点、计算机、服务器或数据中心进行传输。所述计算机可读存储介质可以是计算机能够存取的任何可用介质或者是包含一个或多个可用介质集合的服务器、数据中心等数据存储设备。所述可用介质可以是磁性介质(例如,软盘、硬盘、磁带)、光介质(例如,DVD)、或者半导体介质。半导体介质 可以是固态硬盘(solid state drive,SSD)。
本领域内的技术人员应明白,本申请的实施例可提供为方法、系统、或计算机程序产品。因此,本申请可采用完全硬件实施例、完全软件实施例、或结合软件和硬件方面的实施例的形式。而且,本申请可采用在一个或多个其中包含有计算机可用程序代码的计算机可用存储介质(包括但不限于磁盘存储器、CD-ROM、光学存储器等)上实施的计算机程序产品的形式。
本申请是参照根据本申请的方法、设备(系统)、和计算机程序产品的流程图和/或方框图来描述的。应理解可由计算机程序指令实现流程图和/或方框图中的每一流程和/或方框、以及流程图和/或方框图中的流程和/或方框的结合。可提供这些计算机程序指令到通用计算机、专用计算机、嵌入式处理机或其他可编程数据处理设备的处理器以产生一个机器,使得通过计算机或其他可编程数据处理设备的处理器执行的指令产生用于实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能的装置。
这些计算机程序指令也可存储在能引导计算机或其他可编程数据处理设备以特定方式工作的计算机可读存储器中,使得存储在该计算机可读存储器中的指令产生包括指令装置的制造品,该指令装置实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能。
这些计算机程序指令也可装载到计算机或其他可编程数据处理设备上,使得在计算机或其他可编程设备上执行一系列操作步骤以产生计算机实现的处理,从而在计算机或其他可编程设备上执行的指令提供用于实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能的步骤。
显然,本领域的技术人员可以对本申请进行各种改动和变型而不脱离本申请范围。这样,倘若本申请的这些修改和变型属于本申请权利要求及其等同技术的范围之内,则本申请也意图包含这些改动和变形在内。

Claims (22)

  1. 一种页面换入出方法,其特征在于,所述方法包括:
    加速装置接收主机发送的换入指令,所述换入指令用于指示将目标页面换入至所述主机的内存,所述加速装置通过系统总线与所述主机连接;
    所述加速装置从存储设备获取所述目标页面;
    所述加速装置将所述目标页面换入至所述内存,通知所述主机所述目标页面已换入到所述内存中。
  2. 如权利要求1所述的方法,其特征在于,所述主机中的加速装置接收换入指令,包括:
    所述加速装置接收所述主机中的内存管理单元MMU发送的换入指令。
  3. 如权利要求1或2所述的方法,其特征在于,所述加速装置从存储设备获取所述目标页面,包括:
    所述加速装置获取所述目标页面的虚拟地址与所述目标页面的交换地址的对应关系,所述目标页面的交换地址为所述目标页面在所述存储设备的地址;
    所述加速装置根据所述对应关系确定所述目标页面的交换地址;
    所述加速装置根据所述目标页面的交换地址从所述存储设备获取所述目标页面。
  4. 如权利要求3所述的方法,其特征在于,所述对应关系是所述加速装置将所述目标页面从所述内存换出至所述存储设备时、记录的。
  5. 如权利要求1-4任一项所述的方法,其特征在于,所述主机还包括处理器,所述加速装置通过系统总线与所述处理器连接,所述加速装置将所述目标页面换入至所述内存,包括:
    所述加速装置接收所述处理器发送的通知,所述通知用于指示所述处理器在所述内存中为所述目标页面配置的物理地址;
    所述加速装置通过直接内存访问DMA将所述目标页面写入到所述物理地址。
  6. 如权利要求5所述的方法,其特征在于,所述加速装置接收换入指令之前,还包括:
    所述加速装置接收所述主机的处理器发送的换出指令,所述换出指令用于指示将所述主机内存中的所述目标页面换出至所述存储设备,所述换出指令包括所述目标页面的虚拟地址以及所述目标页面;
    所述加速装置将所述目标页面换出至所述存储设备,记录所述对应关系。
  7. 如权利要求1-6任一项所述的方法,其特征在于,所述加速装置从存储设备获取所述目标页面,包括:
    所述加速装置从所述存储设备获取压缩后的所述目标页面;
    所述加速装置对压缩后的所述目标页面解压,获得所述目标页面。
  8. 如权利要求6所述的方法,其特征在于,所述加速装置将所述目标页面换出至所述存储设备,包括:
    所述加速装置将所述目标页面压缩后,换出至所述存储设备。
  9. 如权利要求3所述的方法,其特征在于,所述对应关系是所述加速装置从所述主机的处理器获取的,所述加速装置通过系统总线与所述处理器连接。
  10. 如权利要求1-9任一项所述的方法,其特征在于,所述存储设备为部署在所述主机之外,与所述主机通过网络连接的存储设备,或为所述主机的本地存储设备,或为公有云或私有云中为所述主机分配的存储空间。
  11. 一种页面换入装置,其特征在于,所述装置位于主机中,所述装置包括:
    接收模块,用于接收换入指令,所述换入指令用于指示将目标页面换入至所述主机的内存;
    换入模块,用于从存储设备获取所述目标页面;将所述目标页面换入至所述内存,通知所述主机所述目标页面已换入到所述内存中。
  12. 如权利要求11所述的装置,其特征在于,所述接收模块,用于:
    接收所述主机中的内存管理模块MMU发送的换入指令。
  13. 如权利要求11或12所述的装置,其特征在于,所述换入模块在从所述存储设备获取所述目标页面,用于:
    获取所述目标页面的虚拟地址与所述目标页面的交换地址的对应关系,所述目标页面的交换地址为所述目标页面在所述存储设备的地址;
    根据所述对应关系确定所述目标页面的交换地址;
    根据所述目标页面的交换地址从所述存储设备获取所述目标页面。
  14. 如权利要求13所述的装置,其特征在于,所述对应关系是所述换入模块将所述目标页面从所述内存换出至所述存储设备时、记录的。
  15. 如权利要求11-14任一项所述的装置,其特征在于,所述换入模块在将所述目标页面换入至所述内存,用于:
    接收所述主机的处理器发送的通知,所述通知用于指示所述处理器在所述内存中为所述目标页面配置的物理地址;
    通过直接内存访问DMA将所述目标页面写入到所述物理地址。
  16. 如权利要求15所述的装置,其特征在于,所述装置还包括换出模块,
    所述接收模块,还用于:接收所述主机的处理器发送的换出指令,所述换出指令用于指示将所述主机内存中的所述目标页面换出至所述存储设备,所述换出指令包括所述目标页面的虚拟地址以及所述目标页面;
    所述换出模块,用于将所述目标页面换出至所述存储设备,记录所述对应关系。
  17. 如权利要求11-16任一项所述的装置,其特征在于,所述换入模块在从所述存储设备获取所述目标页面,用于:
    从所述存储设备获取压缩后的所述目标页面;
    对压缩后的所述目标页面解压,获得所述目标页面。
  18. 如权利要求16所述的装置,其特征在于,所述换出模块在将所述目标页面换出至所述存储设备,用于:
    将所述目标页面压缩后,换出至所述存储设备。
  19. 如权利要求13所述的装置,其特征在于,所述对应关系是所述换入模块从所述主机的处理器获取的。
  20. 如权利要求11-19任一项所述的装置,其特征在于,所述存储设备为部署在所述主机之外,与所述主机通过网络连接的存储设备,或为所述主机的本地存储设备,或为公有云或私有云中为所述主机分配的存储空间。
  21. 一种加速装置,其特征在于,所述加速装置包括数据处理单元DPU和供电电路,所述供电电路用于对所述DPU供电,所述DPU用于执行如权利要求1-10任一项所述的方法。
  22. 一种计算机存储介质,其特征在于,所述计算机可读存储介质存储有计算机可执行指令,所述计算机可执行指令用于使计算机执行权利要求1-10任一项所述的方法。
PCT/CN2023/100492 2022-09-20 2023-06-15 一种页面换入方法以及装置 Ceased WO2024060710A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
EP23867006.1A EP4582943A4 (en) 2022-09-20 2023-06-15 METHOD AND APPARATUS FOR PAGE SWITCHING

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202211142315.0A CN117785371A (zh) 2022-09-20 2022-09-20 一种页面换入方法以及装置
CN202211142315.0 2022-09-20

Publications (1)

Publication Number Publication Date
WO2024060710A1 true WO2024060710A1 (zh) 2024-03-28

Family

ID=90378568

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2023/100492 Ceased WO2024060710A1 (zh) 2022-09-20 2023-06-15 一种页面换入方法以及装置

Country Status (3)

Country Link
EP (1) EP4582943A4 (zh)
CN (1) CN117785371A (zh)
WO (1) WO2024060710A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20240152358A1 (en) * 2022-11-09 2024-05-09 Lemon Inc. Offloading data processing and knowledge synthesis

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120670338B (zh) * 2024-08-30 2026-04-14 华为技术有限公司 一种内存管理方法与电子设备

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109670345A (zh) * 2018-12-21 2019-04-23 成都海光集成电路设计有限公司 内存页面换入换出的保护方法、加速器模块和soc芯片
CN111967065A (zh) * 2020-08-17 2020-11-20 海光信息技术有限公司 一种数据保护方法、处理器及电子设备
CN113590509A (zh) * 2020-04-30 2021-11-02 华为技术有限公司 一种页交换的方法、存储系统和电子设备
WO2022121866A1 (zh) * 2020-12-09 2022-06-16 第四范式(北京)技术有限公司 一种基于加速卡的服务运行方法、装置、电子设备及计算机可读存储介质

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109670345A (zh) * 2018-12-21 2019-04-23 成都海光集成电路设计有限公司 内存页面换入换出的保护方法、加速器模块和soc芯片
CN113590509A (zh) * 2020-04-30 2021-11-02 华为技术有限公司 一种页交换的方法、存储系统和电子设备
CN111967065A (zh) * 2020-08-17 2020-11-20 海光信息技术有限公司 一种数据保护方法、处理器及电子设备
WO2022121866A1 (zh) * 2020-12-09 2022-06-16 第四范式(北京)技术有限公司 一种基于加速卡的服务运行方法、装置、电子设备及计算机可读存储介质

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See also references of EP4582943A4 *

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20240152358A1 (en) * 2022-11-09 2024-05-09 Lemon Inc. Offloading data processing and knowledge synthesis
US12164920B2 (en) * 2022-11-09 2024-12-10 Lemon Inc. Offloading data processing and knowledge synthesis

Also Published As

Publication number Publication date
EP4582943A1 (en) 2025-07-09
CN117785371A (zh) 2024-03-29
EP4582943A4 (en) 2025-11-05

Similar Documents

Publication Publication Date Title
US11397690B2 (en) Virtualized cache implementation method and physical machine
CN107209681B (zh) 一种存储设备访问方法、装置和系统
US9612966B2 (en) Systems, methods and apparatus for a virtual machine cache
US9959074B1 (en) Asynchronous in-memory data backup system
WO2024060711A1 (zh) 一种页面换出方法、装置、设备及数据处理系统
CN114207596B (zh) 将中断从输入-输出存储器管理单元提供到访客操作系统
EP4439312A1 (en) Data storage method and system, storage access configuration method and related device
US20250094203A1 (en) Method and apparatus for creating container, and storage medium
US20210216232A1 (en) Memory data migration method and apparatus
US11243877B2 (en) Method, apparatus for data management, and non-transitory computer-readable storage medium for storing program
WO2015180598A1 (zh) 对存储设备的访问信息处理方法和装置、系统
US12260120B2 (en) Guest operating system buffer and log accesses by an input-output memory management unit
US20250217037A1 (en) Data Migration Method and Apparatus, Chip, and Computer-Readable Storage Medium
EP4582943A1 (en) Page swap-in method and apparatus
US10402333B2 (en) Computer system including plurality of types of memory devices and method
CN120104043A (zh) 一种数据处理方法、装置和计算设备
CN110018879A (zh) 应用于分布式系统的延迟加载方法及装置
US20250377931A1 (en) Container live migration method, processor, host, chip, and interface card
WO2024193272A1 (zh) 一种数据共享方法、装置及设备
CN110209354B (zh) 用于处理数据的方法、装置、设备和介质
CN115794296A (zh) 基于硬件卸载的链接克隆方法、系统、设备及存储介质
CN107832097A (zh) 数据加载方法及装置
CN118860622A (zh) 数据处理系统和内存动态分配方法
WO2024082702A1 (zh) 数据处理方法、装置、芯片以及计算机可读存储介质
WO2017113329A1 (zh) 一种主机集群中缓存管理方法及主机

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23867006

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 2023867006

Country of ref document: EP

ENP Entry into the national phase

Ref document number: 2023867006

Country of ref document: EP

Effective date: 20250402

NENP Non-entry into the national phase

Ref country code: DE

WWP Wipo information: published in national office

Ref document number: 2023867006

Country of ref document: EP