GPU Cluster Architect
Nebius · India; Singapore · 2026-07-29
About this role
About Nebius:
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
The role
We are seeking a GPU Cluster Architect to drive the design of our next-generation AI infrastructure. In this high-impact, hands-on role, you will make end-to-end architectural decisions across compute, networking, and storage — ensuring our platforms can meet the massive scale, performance, and reliability requirements of modern AI workloads.
This is a high-impact, hands-on architecture role where you’ll define how tens of thousands of GPUs are interconnected, cooled down, powered, and optimized across multiple data center sites.
Your responsibilities will include:
• Cluster Design: Architect scalable GPU cluster topologies including compute nodes, interconnect (InfiniBand, Ethernet), storage, and control planes.
• Performance Modeling: Analyze AI/ML workloads (e.g. LLM training, inference) to inform design tradeoffs across latency, bandwidth, and GPU density.
• Network Architecture: Align with network architect relevant design and validate low-latency, high-throughput interconnects (e.g., InfiniBand HDR/NDR, RoCEv2) at POD and DC scale.
• Storage Integration: Work with storage teams to optimize performance for training datasets, checkpointing, and others.
• Reliability & Monitoring: Understand and analyze signal from monitoring systems to the detect flows in design
• Collaboration: Partner with site reliability, networking, storage, and DC engineering teams to operationalize and scale your architecture.
We expect you to have:
• 5+ years of experience designing clusters.
• Deep understanding of modern GPU architecture (NVIDIA, AMD, etc.).
• Experience with HPC interconnects (InfiniBand & RoCE).
• Solid background in systems architecture, networking, and hardware reliability.
• Experience in scripting for automation and telemetry pipelines (Python, Go, etc.)
Benefits & Perks:
• Competitive compensation
• Career growth and learning opportunities
• Flexibility and ownership
• Collaborative and innovative culture
• Opportunity to work on impactful AI projects
• International environment and talented teams
What's it like to work at Nebius:
Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI
Equal Opportunity Statement:
Nebius is an equal opportunity employer. We are…
Skills asked for
- r
- llm
- python
- go
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.