JP7748409B2 - 光ネットワークを用いた再構成可能な計算ポッド - Google Patents
光ネットワークを用いた再構成可能な計算ポッドInfo
- Publication number
- JP7748409B2 JP7748409B2 JP2023035587A JP2023035587A JP7748409B2 JP 7748409 B2 JP7748409 B2 JP 7748409B2 JP 2023035587 A JP2023035587 A JP 2023035587A JP 2023035587 A JP2023035587 A JP 2023035587A JP 7748409 B2 JP7748409 B2 JP 7748409B2
- Authority
- JP
- Japan
- Prior art keywords
- workload
- dimension
- building blocks
- building block
- compute nodes
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Active
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F15/00—Digital computers in general; Data processing equipment in general
- G06F15/16—Combinations of two or more digital computers each having at least an arithmetic unit, a program unit and a register, e.g. for a simultaneous processing of several programs
- G06F15/163—Interprocessor communication
- G06F15/173—Interprocessor communication using an interconnection network, e.g. matrix, shuffle, pyramid, star, snowflake
- G06F15/17356—Indirect interconnection networks
- G06F15/17368—Indirect interconnection networks non hierarchical topologies
- G06F15/17381—Two dimensional, e.g. mesh, torus
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F15/00—Digital computers in general; Data processing equipment in general
- G06F15/16—Combinations of two or more digital computers each having at least an arithmetic unit, a program unit and a register, e.g. for a simultaneous processing of several programs
- G06F15/163—Interprocessor communication
- G06F15/173—Interprocessor communication using an interconnection network, e.g. matrix, shuffle, pyramid, star, snowflake
- G06F15/17356—Indirect interconnection networks
- G06F15/17368—Indirect interconnection networks non hierarchical topologies
- G06F15/17387—Three dimensional, e.g. hypercubes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/46—Multiprogramming arrangements
- G06F9/50—Allocation of resources, e.g. of the central processing unit [CPU]
- G06F9/5005—Allocation of resources, e.g. of the central processing unit [CPU] to service a request
- G06F9/5027—Allocation of resources, e.g. of the central processing unit [CPU] to service a request the resource being a machine, e.g. CPUs, Servers, Terminals
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/46—Multiprogramming arrangements
- G06F9/50—Allocation of resources, e.g. of the central processing unit [CPU]
- G06F9/5061—Partitioning or combining of resources
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/46—Multiprogramming arrangements
- G06F9/50—Allocation of resources, e.g. of the central processing unit [CPU]
- G06F9/5061—Partitioning or combining of resources
- G06F9/5072—Grid computing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/46—Multiprogramming arrangements
- G06F9/50—Allocation of resources, e.g. of the central processing unit [CPU]
- G06F9/5083—Techniques for rebalancing the load in a distributed system
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L45/00—Routing or path finding of packets in data switching networks
- H04L45/02—Topology update or discovery
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L45/00—Routing or path finding of packets in data switching networks
- H04L45/46—Cluster building
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L47/00—Traffic control in data switching networks
- H04L47/10—Flow control; Congestion control
- H04L47/12—Avoiding congestion; Recovering from congestion
- H04L47/125—Avoiding congestion; Recovering from congestion by balancing the load, e.g. traffic engineering
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L47/00—Traffic control in data switching networks
- H04L47/70—Admission control; Resource allocation
- H04L47/72—Admission control; Resource allocation using reservation actions during connection setup
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L47/00—Traffic control in data switching networks
- H04L47/70—Admission control; Resource allocation
- H04L47/78—Architectures of resource allocation
- H04L47/781—Centralised allocation of resources
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L47/00—Traffic control in data switching networks
- H04L47/70—Admission control; Resource allocation
- H04L47/80—Actions related to the user profile or the type of traffic
- H04L47/803—Application aware
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L47/00—Traffic control in data switching networks
- H04L47/70—Admission control; Resource allocation
- H04L47/82—Miscellaneous aspects
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L47/00—Traffic control in data switching networks
- H04L47/70—Admission control; Resource allocation
- H04L47/82—Miscellaneous aspects
- H04L47/821—Prioritising resource allocation or reservation requests
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L49/00—Packet switching elements
- H04L49/65—Re-configuration of fast packet switches
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L67/00—Network arrangements or protocols for supporting network services or applications
- H04L67/01—Protocols
- H04L67/10—Protocols in which an application is distributed across nodes in the network
- H04L67/1001—Protocols in which an application is distributed across nodes in the network for accessing one among a plurality of replicated servers
- H04L67/1004—Server selection for load balancing
- H04L67/1008—Server selection for load balancing based on parameters of servers, e.g. available memory or workload
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04L—TRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
- H04L67/00—Network arrangements or protocols for supporting network services or applications
- H04L67/50—Network services
- H04L67/56—Provisioning of proxy services
- H04L67/563—Data redirection of data network streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04Q—SELECTING
- H04Q11/00—Selecting arrangements for multiplex systems
- H04Q11/0001—Selecting arrangements for multiplex systems using optical switching
- H04Q11/0005—Switch and router aspects
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04Q—SELECTING
- H04Q11/00—Selecting arrangements for multiplex systems
- H04Q11/0001—Selecting arrangements for multiplex systems using optical switching
- H04Q11/0062—Network aspects
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F2209/00—Indexing scheme relating to G06F9/00
- G06F2209/50—Indexing scheme relating to G06F9/50
- G06F2209/505—Clust
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04Q—SELECTING
- H04Q11/00—Selecting arrangements for multiplex systems
- H04Q11/0001—Selecting arrangements for multiplex systems using optical switching
- H04Q11/0005—Switch and router aspects
- H04Q2011/0052—Interconnection of switches
- H04Q2011/0058—Crossbar; Matrix
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04Q—SELECTING
- H04Q11/00—Selecting arrangements for multiplex systems
- H04Q11/0001—Selecting arrangements for multiplex systems using optical switching
- H04Q11/0062—Network aspects
- H04Q2011/0064—Arbitration, scheduling or medium access control aspects
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04Q—SELECTING
- H04Q11/00—Selecting arrangements for multiplex systems
- H04Q11/0001—Selecting arrangements for multiplex systems using optical switching
- H04Q11/0062—Network aspects
- H04Q2011/0079—Operation or maintenance aspects
- H04Q2011/0081—Fault tolerance; Redundancy; Recovery; Reconfigurability
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04Q—SELECTING
- H04Q11/00—Selecting arrangements for multiplex systems
- H04Q11/0001—Selecting arrangements for multiplex systems using optical switching
- H04Q11/0062—Network aspects
- H04Q2011/0086—Network resource allocation, dimensioning or optimisation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04Q—SELECTING
- H04Q11/00—Selecting arrangements for multiplex systems
- H04Q11/0001—Selecting arrangements for multiplex systems using optical switching
- H04Q11/0062—Network aspects
- H04Q2011/009—Topology aspects
- H04Q2011/0098—Mesh
Landscapes
- Engineering & Computer Science (AREA)
- Computer Networks & Wireless Communication (AREA)
- Signal Processing (AREA)
- Theoretical Computer Science (AREA)
- Software Systems (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Computer Hardware Design (AREA)
- Mathematical Physics (AREA)
- Data Exchanges In Wide-Area Networks (AREA)
- Multi Processors (AREA)
- Hardware Redundancy (AREA)
Description
いくつかの計算ワークロード、例えば機械学習トレーニングは、ワークロードを効率的に処理するために多くの処理ノードを必要とする。処理ノードは、相互接続ネットワークを介して互いに通信することができる。例えば、機械学習トレーニングの場合、処理ノードは、互いに通信することによって、最適な深層学習モデルに収束することができる。相互接続ネットワークは、処理ユニットが収束を達成する速度および効率にとって重要である。
本明細書は、光ネットワークを用いて、ワークロードクラスタを生成する計算ノードのスーパポートを再構成できる技術を説明する。
含む。いくつかの態様において、光ネットワークは、n次元の各次元について、当該次元に沿った計算ノードの間にデータをルーティングする当該光ネットワークの1つ以上の光回路スイッチを含む。各ビルディングブロックは、当該ビルディングブロックの各次元に沿った複数のセグメントの計算ノードを含むことができる。光ネットワークは、各次元の各セグメントについて、ワークロードクラスタ内の各ビルディングブロックに対応する計算ノードセグメントの間にデータをルーティングする当該光ネットワークの光回路スイッチを含むことができる。
行するための計算ノードの適切な数をより効率で割り当てることができ、各ワークロードを実行するための計算ノードの構成を最適化(または改善)することができる。複数の種類の計算ノードを含むスーパポッドを使用して、例えば、データセンタまたは他の場所において互いに物理的に近接する(例えば、同一のラックにおいて互いに接続されるおよび/または隣接する)計算ノードに限定されず、計算ノードの適切な数および構成だけでなく、各ワークロードを実行するための計算ノードの適切な種類を含むワークロードクラスタを生成することができる。代わりに、光ネットワークは、様々な形状のワークロードクラスタを可能にする。これらのワークロードクラスタにおいて、計算ノードは、互いに任意の物理的位置に配置されても、互いに隣接するように動作する。
様々な図面において、同様の参照番号および名称は、同様の要素を示す。
ィングブロック160を含む。しかしながら、スーパポッド152~158は、他の数のビルディングブロック160、例えば、20、50、100、または別の適切な数のビルディングブロック160を含むことができる。また、スーパポッド152~158は、異なる数のビルディングブロック160を含むことができる。例えば、スーパポッド152は、64個のビルディングブロックを含むことができ、スーパポッド152は、100個のビルディングブロックを含む。
成は、同様のまたは異なる計算ノードを含むことができる。例えば、TPUを含むビルディングブロックは、GPUを含むビルディングブロックとは異なる構成を有してもよい。
ドの優先度に基づいて、スーパポッドのビルディングブロックから、ワークロードクラスタを生成するように接続すべき1組のビルディングブロックを選択することができる。例えば、以下で説明するように、ワークロードスケジューラがスーパポッド内の利用可能且つ健全なビルディングブロックの数よりも多くのビルディングブロックを含むワークロードクラスタを要求する要求を受信する場合、ワークロードスケジューラは、より低い優先度ワークロードを実行するためのビルディングブロックを、要求されたワークロードクラスタに割り当て直すことができる。ワークロードスケジューラは、選択されたビルディングブロックを特定するデータを、OCSマネージャに提供することができる。OCSマネージャは、ビルディングブロックを互いに接続するように1つ以上のOCSスイッチを構成することによって、ワークロードクラスタを生成することができる。その後、ワークロードスケジューラは、ワークロードクラスタ内の計算ノード上でワークロードを実行することができる。
計算ノードと、z次元に沿った4つの計算ノードとを有する。各ビルディングブロックが各次元に沿って4つの計算ノードを有するため、ワークロードクラスタ220は、x次元に沿った2つのビルディングブロックと、y次元に沿った2つのビルディングブロックと、z次元に沿った1つのビルディングブロックとを含む。
る。
する追加の外部リンクがある。また、1つ以上のOCSスイッチは、3次元の全てに沿って、各セグメントの一端と各セグメントの他端との間にラップアラウンドリンク341~343を形成するように構成されてもよい。
クタ、例えばOSFPコネクタにルーティングされてもよい。この例において、ポート410は、電気リンク412を介して光モジュール420に接続される。光モジュール420は、必要に応じて、大きなデータセンタに配置された計算ノード間のデータ通信を提供するために、電気リンクを、外部リンクの長さで延在する、例えば1キロメートル(km)を超えて延長する光リンクに変換することができる。光モジュールの種類は、ビルディングブロックとOCSスイッチとの間の必要な長さおよびリンクの所望の速度および帯域幅に基づいて、変更されてもよい。
光ファイバケーブルに取り付けられた光ファイバモジュールを収容することができる。
ってもよい。例示的なスーパポッド900は、64個の4×4×4ビルディングブロック960を含み、これらのビルディングブロック960を使用して、計算ワークロード、例えば機械学習ワークロードを実行するためのワークロードクラスタを形成することができる。上述したように、各4×4×4ビルディングブロック960は、3次元の各次元に沿って配置された4つの計算ノードからなる32個の計算ノードを含む。例えば、ビルディングブロック960は、上述したビルディングブロック310、ワークロードクラスタ320、またはビルディングブロック700と同様であってもよく、類似であってもよい。
適切な範囲で数値的に表すことができる。例えば、ワークロードスケジューラ910は、ユーザ装置またはセルスケジューラ、例えば図1のユーザ装置110またはセルスケジューラ140から、要求データを受信することができる。上述したように、要求データは、計算ノードのn次元目標構成、例えば計算ノードを含むビルディングブロックの目標構成を指定することができる。
ロードのためのビルディングブロックを解放することができる。
、第1のビルディングブロックを第2のビルディングブロックに接続しようとすると仮定する。OCSマネージャ920は、第1のビルディングブロックのx次元セグメントと第2のビルディングブロックのx次元セグメントとの間にデータをルーティングするために、x次元のOCSスイッチ950のルーティングテーブルを更新する。ビルディングブロックの各x次元セグメントを接続する必要があるため、OCSマネージャ920は、各OCSスイッチ950のルーティングテーブルを更新することができる。
CSスイッチは、ワークロードクラスタのビルディングブロックの間にデータをルーティングすることができる。構成されたOCSスイッチは、計算ノードが目標構成において物理的に接続されていなくても物理的に接続されていたように、ビルディングブロックの計算ノードの間にデータをルーティングすることができる。
ットの論理スポットに論理的に配置することができる。上述したように、OCSスイッチのルーティングテーブルは、あるビルディングブロックのセグメントに接続されたOCSスイッチの物理ポートを、対応する別のビルディングブロックのセグメントに接続されたOCSスイッチの物理ポートにマッピングすることができる。この場合、システムは、故障したビルディングブロックではなく、特定された利用可能なビルディングブロックの対応するセグメントとのマッピングを更新することによって、置換を行うことができる。
らに、コンピュータは、ユーザによって使用されている装置に文書を送信し、当該装置から文書を受信することによって、例えば、ウェブブラウザから受信された要求に応答して、ユーザのクライアント装置上のウェブブラウザにウェブページを送信することによって、ユーザと対話することができる。
は、望ましい結果を達成するために、必ずしも図示された特定の順序でまたは順次に実行される必要がない。いくつかの実装形態において、マルチタスクおよび並列処理が有利であり得る。
Claims (10)
- 1つ以上のデータ処理装置によって実行される方法であって、前記方法は、
計算ワークロードを実行するための計算ノードの目標構成を特定することと、
前記計算ノードの目標構成に一致する構成を有する計算ノードのワークロードクラスタを生成することとを含み、
前記生成することは、
n次元の各次元について、当該次元に配置された複数のビルディングブロックの計算ノードが当該次元の1つ以上の光回路スイッチを介して互いに通信するように、当該次元の光ネットワークの前記1つ以上の光回路スイッチのそれぞれを構成することを含み、前記複数のビルディングブロックは、前記光ネットワークに通信可能に接続され、各ビルディングブロックは、n次元構成の計算ノードを含み、複数の計算ノードは、前記n次元(nは、2以上である)の各次元に配置され、
前記ワークロードクラスタの前記計算ノードに、前記計算ワークロードを実行させることを含む、方法。 - 各次元の前記1つ以上の光回路スイッチのそれぞれを構成することは、各次元の前記1つ以上の光スイッチのそれぞれのルーティングデータを構成することを含み、
前記それぞれのルーティングデータは、前記次元に沿った計算ノードの間に前記計算ワークロードのデータをどのようにルーティングするかを指定する、請求項1に記載の方法。 - 前記計算ワークロードのビルディングブロックのセットから、ビルディングブロックのサブセットを前記複数のビルディングブロックとして選択することをさらに含む、請求項1または2に記載の方法。
- 各ビルディングブロックは、前記n次元の各次元に沿った複数のセグメントの計算ノードを含み、
前記光ネットワークは、前記複数のビルディングブロックの各次元の各セグメントに対応する光回路スイッチを含む、請求項1から3のいずれか1項に記載の方法。 - 各ビルディングブロックは、3次元トーラス状計算ノードまたはメッシュ状計算ノードのうちの1つを含む、請求項1から4のいずれか1項に記載の方法。
- 前記ワークロードクラスタの特定のビルディングブロックが故障したことを示すデータを受信することと、
前記複数のビルディングブロックのうちの利用可能なビルディングブロックを用いて前記特定のビルディングブロックを置換することとをさらに含み、
前記置換することは、前記利用可能なビルディングブロックの前記計算ノードが1つ以上の次元に対応する前記ワークロードクラスタの他の計算ノードと通信するように、前記1つ以上の次元に対応する前記1つ以上の光回路スイッチのそれぞれを再構成することを含む、請求項1から5のいずれか1項に記載の方法。 - 前記計算ワークロードを実行するための前記計算ノードの目標構成を特定することは、前記計算ノードの目標構成および前記計算ノードの目標構成に含まれる異なる種類の計算ノードを指定する要求データを受信することを含む、請求項1から6のいずれか1項に記載の方法。
- 前記要求データによって指定された各種類の計算ノードについて、前記指定された種類の1つ以上の計算ノードを含むビルディングブロックを決定することをさらに含む、請求項7に記載の方法。
- システムであって、
データ処理装置と、
コンピュータプログラムを格納したコンピュータ記憶媒体とを備え、
前記コンピュータプログラムは、前記データ処理装置によって実行されると、前記データ処理装置に請求項1から8のいずれか1項に記載の方法を実施させる、システム。 - 請求項1から8のいずれか1項に記載の方法をコンピュータに実行させるプログラム。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2025075406A JP2025114657A (ja) | 2019-03-06 | 2025-04-30 | 光ネットワークを用いた再構成可能な計算ポッド |
Applications Claiming Priority (6)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US201962814757P | 2019-03-06 | 2019-03-06 | |
| US62/814,757 | 2019-03-06 | ||
| US16/381,951 US11042416B2 (en) | 2019-03-06 | 2019-04-11 | Reconfigurable computing pods using optical networks |
| US16/381,951 | 2019-04-11 | ||
| JP2021522036A JP7242847B2 (ja) | 2019-03-06 | 2019-12-18 | 光ネットワークを用いた再構成可能な計算ポッド |
| PCT/US2019/067100 WO2020180387A1 (en) | 2019-03-06 | 2019-12-18 | Reconfigurable computing pods using optical networks |
Related Parent Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP2021522036A Division JP7242847B2 (ja) | 2019-03-06 | 2019-12-18 | 光ネットワークを用いた再構成可能な計算ポッド |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP2025075406A Division JP2025114657A (ja) | 2019-03-06 | 2025-04-30 | 光ネットワークを用いた再構成可能な計算ポッド |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JP2023078228A JP2023078228A (ja) | 2023-06-06 |
| JP7748409B2 true JP7748409B2 (ja) | 2025-10-02 |
Family
ID=72336372
Family Applications (3)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP2021522036A Active JP7242847B2 (ja) | 2019-03-06 | 2019-12-18 | 光ネットワークを用いた再構成可能な計算ポッド |
| JP2023035587A Active JP7748409B2 (ja) | 2019-03-06 | 2023-03-08 | 光ネットワークを用いた再構成可能な計算ポッド |
| JP2025075406A Pending JP2025114657A (ja) | 2019-03-06 | 2025-04-30 | 光ネットワークを用いた再構成可能な計算ポッド |
Family Applications Before (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP2021522036A Active JP7242847B2 (ja) | 2019-03-06 | 2019-12-18 | 光ネットワークを用いた再構成可能な計算ポッド |
Family Applications After (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP2025075406A Pending JP2025114657A (ja) | 2019-03-06 | 2025-04-30 | 光ネットワークを用いた再構成可能な計算ポッド |
Country Status (9)
| Country | Link |
|---|---|
| US (3) | US11042416B2 (ja) |
| EP (1) | EP3853732B1 (ja) |
| JP (3) | JP7242847B2 (ja) |
| KR (2) | KR102583771B1 (ja) |
| CN (2) | CN117873727A (ja) |
| BR (1) | BR112021007538A2 (ja) |
| DK (1) | DK3853732T3 (ja) |
| FI (1) | FI3853732T3 (ja) |
| WO (1) | WO2020180387A1 (ja) |
Families Citing this family (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9847918B2 (en) * | 2014-08-12 | 2017-12-19 | Microsoft Technology Licensing, Llc | Distributed workload reassignment following communication failure |
| US11042416B2 (en) * | 2019-03-06 | 2021-06-22 | Google Llc | Reconfigurable computing pods using optical networks |
| US11847012B2 (en) * | 2019-06-28 | 2023-12-19 | Intel Corporation | Method and apparatus to provide an improved fail-safe system for critical and non-critical workloads of a computer-assisted or autonomous driving vehicle |
| US12335099B2 (en) | 2020-10-19 | 2025-06-17 | Google Llc | Enhanced reconfigurable interconnect network |
| US11516087B2 (en) * | 2020-11-30 | 2022-11-29 | Google Llc | Connecting processors using twisted torus configurations |
| CN121864567A (zh) * | 2024-10-12 | 2026-04-14 | 华为技术有限公司 | 计算集群、数据传输方法以及计算模块 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2006146864A (ja) | 2004-11-17 | 2006-06-08 | Raytheon Co | 高性能計算(hpc)システムにおけるスケジューリング |
| JP2016504668A (ja) | 2012-11-21 | 2016-02-12 | コーヒレント・ロジックス・インコーポレーテッド | 分散型プロセッサを有する処理システム |
| JP2016091069A (ja) | 2014-10-30 | 2016-05-23 | 富士通株式会社 | ジョブ管理プログラム、ジョブ管理方法、およびジョブ管理装置 |
| JP2017527031A (ja) | 2014-08-18 | 2017-09-14 | アドバンスト・マイクロ・ディバイシズ・インコーポレイテッドAdvanced Micro Devices Incorporated | セルオートマトンを用いたクラスタサーバの構成 |
Family Cites Families (48)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US4598400A (en) * | 1983-05-31 | 1986-07-01 | Thinking Machines Corporation | Method and apparatus for routing message packets |
| US6853635B1 (en) * | 2000-07-24 | 2005-02-08 | Nortel Networks Limited | Multi-dimensional lattice network |
| US20020176131A1 (en) | 2001-02-28 | 2002-11-28 | Walters David H. | Protection switching for an optical network, and methods and apparatus therefor |
| CN1742512A (zh) | 2002-12-04 | 2006-03-01 | 康宁股份有限公司 | 用快速竞争解决设计的快速切换可升级光互连 |
| US8401385B2 (en) | 2003-10-02 | 2013-03-19 | Trex Enterprises Corp. | Optically switched communication network |
| US9178784B2 (en) * | 2004-04-15 | 2015-11-03 | Raytheon Company | System and method for cluster management based on HPC architecture |
| US7518120B2 (en) | 2005-01-04 | 2009-04-14 | The Regents Of The University Of Michigan | Long-distance quantum communication and scalable quantum computation |
| JP2006215816A (ja) * | 2005-02-03 | 2006-08-17 | Fujitsu Ltd | 情報処理システムおよび情報処理システムの制御方法 |
| US8194638B2 (en) | 2006-07-27 | 2012-06-05 | International Business Machines Corporation | Dual network types solution for computer interconnects |
| CN101330413B (zh) * | 2007-06-22 | 2012-08-08 | 上海红神信息技术有限公司 | 基于环绕网络与超立方网络架构的混合多阶张量扩展方法 |
| CN101354694B (zh) * | 2007-07-26 | 2010-10-13 | 上海红神信息技术有限公司 | 基于mpu架构的超高扩展超级计算系统 |
| US8687975B2 (en) | 2007-10-23 | 2014-04-01 | Hewlett-Packard Development Company, L.P. | Integrated circuit with optical interconnect |
| DK2083532T3 (da) * | 2008-01-23 | 2014-02-10 | Comptel Corp | Konvergerende formidlingssystem med forbedret dataoverføring |
| US7856544B2 (en) * | 2008-08-18 | 2010-12-21 | International Business Machines Corporation | Stream processing in super node clusters of processors assigned with stream computation graph kernels and coupled by stream traffic optical links |
| US8296419B1 (en) | 2009-03-31 | 2012-10-23 | Amazon Technologies, Inc. | Dynamically modifying a cluster of computing nodes used for distributed execution of a program |
| US8270830B2 (en) | 2009-04-01 | 2012-09-18 | Fusion-Io, Inc. | Optical network for cluster computing |
| US8619605B2 (en) * | 2009-05-13 | 2013-12-31 | Avaya Inc. | Method and apparatus for maintaining port state tables in a forwarding plane of a network element |
| US8260840B1 (en) | 2010-06-28 | 2012-09-04 | Amazon Technologies, Inc. | Dynamic scaling of a cluster of computing nodes used for distributed execution of a program |
| US8719415B1 (en) | 2010-06-28 | 2014-05-06 | Amazon Technologies, Inc. | Use of temporarily available computing nodes for dynamic scaling of a cluster |
| US8824491B2 (en) | 2010-10-25 | 2014-09-02 | Polytechnic Institute Of New York University | Distributed scheduling for variable-size packet switching system |
| KR101254706B1 (ko) * | 2011-09-27 | 2013-04-15 | 성균관대학교산학협력단 | 3차원 네트워크 온 칩 |
| US20130156425A1 (en) * | 2011-12-17 | 2013-06-20 | Peter E. Kirkpatrick | Optical Network for Cluster Computing |
| US8867915B1 (en) | 2012-01-03 | 2014-10-21 | Google Inc. | Dynamic data center network with optical circuit switch |
| EP2810451A1 (en) | 2012-02-03 | 2014-12-10 | Lightfleet Corporation | Scalable optical broadcast interconnect |
| US9465632B2 (en) | 2012-02-04 | 2016-10-11 | Global Supercomputing Corporation | Parallel hardware hypervisor for virtualizing application-specific supercomputers |
| US9229163B2 (en) | 2012-05-18 | 2016-01-05 | Oracle International Corporation | Butterfly optical network with crossing-free switches |
| US9479219B1 (en) * | 2012-09-24 | 2016-10-25 | Google Inc. | Validating a connection to an optical circuit switch |
| US9332323B2 (en) * | 2012-10-26 | 2016-05-03 | Guohua Liu | Method and apparatus for implementing a multi-dimensional optical circuit switching fabric |
| US10394611B2 (en) | 2012-11-26 | 2019-08-27 | Amazon Technologies, Inc. | Scaling computing clusters in a distributed computing system |
| KR101465420B1 (ko) * | 2013-10-08 | 2014-11-27 | 성균관대학교산학협력단 | 네트워크 온 칩 및 네트워크 온 칩의 신호를 라우팅하는 방법 |
| CN103580771B (zh) * | 2013-11-11 | 2016-01-20 | 清华大学 | 基于时间同步的全光时片交换方法 |
| CN104731796B (zh) * | 2013-12-19 | 2017-12-19 | 秒针信息技术有限公司 | 数据存储计算方法和系统 |
| US11290524B2 (en) * | 2014-08-13 | 2022-03-29 | Microsoft Technology Licensing, Llc | Scalable fault resilient communications within distributed clusters |
| US10200292B2 (en) * | 2014-08-25 | 2019-02-05 | Intel Corporation | Technologies for aligning network flows to processing resources |
| US9521089B2 (en) | 2014-08-30 | 2016-12-13 | International Business Machines Corporation | Multi-layer QoS management in a distributed computing environment |
| US20160241474A1 (en) * | 2015-02-12 | 2016-08-18 | Ren Wang | Technologies for modular forwarding table scalability |
| CN107710702B (zh) * | 2015-03-23 | 2020-09-01 | 艾易珀尼斯公司 | 一种用于在数据中心网络中路由数据的系统 |
| US10390114B2 (en) * | 2016-07-22 | 2019-08-20 | Intel Corporation | Memory sharing for physical accelerator resources in a data center |
| US10389800B2 (en) * | 2016-10-11 | 2019-08-20 | International Business Machines Corporation | Minimizing execution time of a compute workload based on adaptive complexity estimation |
| US10834484B2 (en) * | 2016-10-31 | 2020-11-10 | Ciena Corporation | Flat, highly connected optical network for data center switch connectivity |
| US10243687B2 (en) * | 2016-11-17 | 2019-03-26 | Google Llc | Optical network unit wavelength tuning |
| CN106851442B (zh) | 2017-01-19 | 2019-05-21 | 西安电子科技大学 | 一种超级计算机中的光互连网络系统及通信方法 |
| CN107094270A (zh) * | 2017-05-11 | 2017-08-25 | 中国科学院计算技术研究所 | 可重构的互连系统及其拓扑构建方法 |
| JP6885193B2 (ja) * | 2017-05-12 | 2021-06-09 | 富士通株式会社 | 並列処理装置、ジョブ管理方法、およびジョブ管理プログラム |
| CN107241660B (zh) * | 2017-06-26 | 2021-07-06 | 国网信息通信产业集团有限公司 | 面向智能电网业务的全光灵活粒度的交换网络架构及方法 |
| US10552227B2 (en) * | 2017-10-31 | 2020-02-04 | Calient Technologies, Inc. | Reconfigurable computing cluster with assets closely coupled at the physical layer by means of an optical circuit switch |
| US11042416B2 (en) * | 2019-03-06 | 2021-06-22 | Google Llc | Reconfigurable computing pods using optical networks |
| US11122347B2 (en) * | 2019-07-01 | 2021-09-14 | Google Llc | Reconfigurable computing pods using optical networks with one-to-many optical switches |
-
2019
- 2019-04-11 US US16/381,951 patent/US11042416B2/en active Active
- 2019-12-18 CN CN202410134763.9A patent/CN117873727A/zh active Pending
- 2019-12-18 EP EP19839305.0A patent/EP3853732B1/en active Active
- 2019-12-18 DK DK19839305.0T patent/DK3853732T3/da active
- 2019-12-18 KR KR1020217011905A patent/KR102583771B1/ko active Active
- 2019-12-18 KR KR1020237032481A patent/KR102625118B1/ko active Active
- 2019-12-18 BR BR112021007538-0A patent/BR112021007538A2/pt unknown
- 2019-12-18 FI FIEP19839305.0T patent/FI3853732T3/fi active
- 2019-12-18 WO PCT/US2019/067100 patent/WO2020180387A1/en not_active Ceased
- 2019-12-18 CN CN201980069191.8A patent/CN112889032B/zh active Active
- 2019-12-18 JP JP2021522036A patent/JP7242847B2/ja active Active
-
2021
- 2021-05-27 US US17/332,769 patent/US11537443B2/en active Active
-
2022
- 2022-12-05 US US18/075,332 patent/US12182628B2/en active Active
-
2023
- 2023-03-08 JP JP2023035587A patent/JP7748409B2/ja active Active
-
2025
- 2025-04-30 JP JP2025075406A patent/JP2025114657A/ja active Pending
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2006146864A (ja) | 2004-11-17 | 2006-06-08 | Raytheon Co | 高性能計算(hpc)システムにおけるスケジューリング |
| JP2016504668A (ja) | 2012-11-21 | 2016-02-12 | コーヒレント・ロジックス・インコーポレーテッド | 分散型プロセッサを有する処理システム |
| JP2017527031A (ja) | 2014-08-18 | 2017-09-14 | アドバンスト・マイクロ・ディバイシズ・インコーポレイテッドAdvanced Micro Devices Incorporated | セルオートマトンを用いたクラスタサーバの構成 |
| JP2016091069A (ja) | 2014-10-30 | 2016-05-23 | 富士通株式会社 | ジョブ管理プログラム、ジョブ管理方法、およびジョブ管理装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN117873727A (zh) | 2024-04-12 |
| BR112021007538A2 (pt) | 2021-07-27 |
| KR102583771B1 (ko) | 2023-09-27 |
| US20230161638A1 (en) | 2023-05-25 |
| KR20230141921A (ko) | 2023-10-10 |
| JP2025114657A (ja) | 2025-08-05 |
| JP2022522320A (ja) | 2022-04-18 |
| KR102625118B1 (ko) | 2024-01-12 |
| EP3853732A1 (en) | 2021-07-28 |
| WO2020180387A1 (en) | 2020-09-10 |
| US12182628B2 (en) | 2024-12-31 |
| KR20210063382A (ko) | 2021-06-01 |
| US11042416B2 (en) | 2021-06-22 |
| JP2023078228A (ja) | 2023-06-06 |
| CN112889032B (zh) | 2024-02-06 |
| FI3853732T3 (fi) | 2026-01-23 |
| US11537443B2 (en) | 2022-12-27 |
| US20200285524A1 (en) | 2020-09-10 |
| JP7242847B2 (ja) | 2023-03-20 |
| US20210286656A1 (en) | 2021-09-16 |
| EP3853732B1 (en) | 2025-10-29 |
| CN112889032A (zh) | 2021-06-01 |
| DK3853732T3 (da) | 2025-12-15 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11451889B2 (en) | Reconfigurable computing pods using optical networks with one-to-many optical switches | |
| US12182628B2 (en) | Reconfigurable computing pods using optical networks | |
| EP3993359B1 (en) | Enhanced reconfigurable interconnect network | |
| US10606651B2 (en) | Free form expression accelerator with thread length-based thread assignment to clustered soft processor cores that share a functional circuit | |
| WO2017003837A1 (en) | Deep neural network partitioning on servers | |
| WO2017003887A1 (en) | Convolutional neural networks on hardware accelerators | |
| KR20070085089A (ko) | 고성능 연산(hpc) 시스템에서의 온디맨드 인스턴스화 | |
| HK40108484A (zh) | 使用光网络的可重新配置的计算平台 | |
| HK40081215A (en) | Reconfigurable computing pods using optical networks with one-to-many optical switches | |
| HK40053729A (en) | Reconfigurable computing pods using optical networks | |
| HK40053729B (zh) | 使用光网络的可重新配置的计算平台 | |
| HK40044238A (en) | Reconfigurable computing pods using optical networks with one-to-many optical switches | |
| HK40044238B (en) | Reconfigurable computing pods using optical networks with one-to-many optical switches | |
| HK40081215B (zh) | 使用具有一对多光交换机的光网络的可重新配置的计算平台 | |
| Vijaya Kumar et al. | Reconfigurable Torus Fabrics for Multi-tenant ML |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| A521 | Request for written amendment filed |
Free format text: JAPANESE INTERMEDIATE CODE: A523 Effective date: 20230404 |
|
| A621 | Written request for application examination |
Free format text: JAPANESE INTERMEDIATE CODE: A621 Effective date: 20230404 |
|
| A131 | Notification of reasons for refusal |
Free format text: JAPANESE INTERMEDIATE CODE: A131 Effective date: 20240702 |
|
| A02 | Decision of refusal |
Free format text: JAPANESE INTERMEDIATE CODE: A02 Effective date: 20250107 |
|
| A521 | Request for written amendment filed |
Free format text: JAPANESE INTERMEDIATE CODE: A523 Effective date: 20250430 |
|
| TRDD | Decision of grant or rejection written | ||
| A01 | Written decision to grant a patent or to grant a registration (utility model) |
Free format text: JAPANESE INTERMEDIATE CODE: A01 Effective date: 20250826 |
|
| A61 | First payment of annual fees (during grant procedure) |
Free format text: JAPANESE INTERMEDIATE CODE: A61 Effective date: 20250919 |
|
| R150 | Certificate of patent or registration of utility model |
Ref document number: 7748409 Country of ref document: JP Free format text: JAPANESE INTERMEDIATE CODE: R150 |