GPU Cluster Architect
Quick Summary
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a GPU Cluster Architect based in the United States. This is a remote,
This is a remote, high-impact architecture role focused on designing next-generation AI infrastructure at massive scale.
You’ll make end-to-end architectural decisions spanning GPU compute, high-performance networking, storage, reliability, and control planes.
Your work will shape how tens of thousands of GPUs are interconnected, powered, cooled, monitored, and optimized across multiple data center sites.
You’ll model demanding AI and machine-learning workloads, including large language model training and inference, to guide critical performance and infrastructure tradeoffs.
The role combines deep systems expertise with hands-on collaboration across networking, storage, site reliability, and data center engineering teams.
You’ll operate in a fast-moving, engineering-led environment where scalability, performance, reliability, and innovation are central to the work.
This is an opportunity to influence the architecture of large-scale AI infrastructure while solving complex distributed systems challenges.
-
Architect scalable GPU cluster topologies encompassing compute nodes, high-performance interconnects, storage systems, and control planes.
-
Define and evaluate infrastructure architectures capable of supporting large-scale AI and machine-learning workloads across multiple data center sites.
-
Model workload requirements for applications such as large language model training and inference, using latency, bandwidth, GPU density, and other performance factors to guide architectural decisions.
-
Design and validate high-throughput, low-latency networking architectures at both POD and data-center scale, including InfiniBand and Ethernet-based environments.
-
Work with network architecture teams to evaluate and validate technologies such as InfiniBand HDR/NDR and RoCEv2.
-
Partner with storage engineering teams to optimize infrastructure for training datasets, checkpointing, and other demanding AI workloads.
-
Analyze monitoring and telemetry signals to identify design issues, reliability risks, and opportunities for architectural improvement.
-
Collaborate closely with site reliability, networking, storage, and data center engineering teams to operationalize, deploy, and scale infrastructure architectures.
-
Contribute to automation and telemetry initiatives that improve the visibility, performance, and reliability of large-scale GPU environments.
-
Make end-to-end architectural decisions that balance scalability, performance, reliability, operational complexity, and infrastructure efficiency.
Requirements
~1 min read-
5+ years of experience designing and architecting large-scale computing or GPU clusters.
-
Deep understanding of modern GPU architectures, including NVIDIA, AMD, or comparable platforms.
-
Strong experience with high-performance computing interconnects, particularly InfiniBand and RoCE.
-
Solid background in systems architecture, networking, hardware infrastructure, and hardware reliability.
-
Understanding of GPU cluster design principles, including compute topology, network architecture, storage integration, and control-plane considerations.
-
Experience evaluating infrastructure performance and making architecture decisions based on workload characteristics such as latency, bandwidth, and compute density.
-
Experience with scripting or software development for automation, telemetry, monitoring, or infrastructure tooling using languages such as Python or Go.
-
Ability to analyze technical signals and operational data to identify infrastructure issues and inform design improvements.
-
Strong cross-functional collaboration skills, with the ability to work effectively with networking, storage, site reliability, and data center engineering teams.
-
Strong analytical and problem-solving abilities, with a practical approach to complex infrastructure challenges.
-
Ability to work independently in a fast-moving environment while taking ownership of significant architectural decisions.
-
Excellent communication skills and the ability to explain complex technical architectures and tradeoffs to technical stakeholders.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- First seen
- September 30, 2026
- Last seen
- September 30, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 68%
- Scored at
- September 30, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.