Lead Software Systems Engineer - GPU Performance
Nebius · Remote - United States · 2026-07-27
About this role
<div class="content-intro"><p><strong>About Nebius:</strong></p> <p>Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.</p> <p>Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.</p> <p>Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&amp;D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&amp;D.</p></div><p>We are looking for a Lead Software Systems Engineer - GPU Performance to play a key role in building our hyperscaler platform, working across its core components while analyzing and optimizing the performance of large-scale GPU clusters at the intersection of hardware and software.</p> <p>You will operate across the full stack—from hardware and system software to networking (InfiniBand/RoCE), virtualization (KVM/QEMU), and distributed communication layers (e.g., MPI, NCCL).</p> <p><strong>In this role you will </strong></p> <ul> <li>Focus on understanding system behavior across multiple layers, identifying performance bottlenecks, and driving improvements that shape how our clusters are built, operated, tuned, and validated.</li> <li>Investigate and troubleshoot performance issues of GPU cluster under real workloads (training and inference)</li> <li>Evaluate and integrate new hardware, system configurations and tuning approaches through software stack</li> <li>Support complex performance-related escalations from internal teams and customers</li> <li>Work closely with infrastructure, software engineering and hardware vendor teams (e.g. NVIDIA, Mellanox, Intel)</li> <li>Contribute to hardware and cluster qualification (acceptance), ensuring systems meet performance expectations</li> </ul> <p><br><strong>We expect you to have:</strong></p> <ul> <li>5+ years of professional experience in&nbsp;<strong>system-level software development</strong>&nbsp;(focused on performance optimization, low-level programming). </li> <li>3+ years of hands-on experience with&nbsp;<strong>Linux systems</strong>&nbsp;(administration, troubleshooting, and performance tuning). </li> <li><strong>In-depth understanding</strong> of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems. </li> <li>Strong proficiency in one or more <strong>performance-oriented programming languages</strong>&nbsp;(C/C++, Go, Python).</li> </ul> <p><em>We conduct coding interviews as part of the process.<br><br></em></p> <p><strong>Key employee benefits:</strong></p> <ul> <li><strong>Health insurance:&nbsp;</strong>100% company-paid medical, dental and vision coverage for employees and families.</li> <li><strong>401(k) plan:</strong>&nbsp;Up to 4% company match with immediate vesting.</li> <li><strong>Parental leave</strong>: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.</li> <li><strong>Remote work reimbursement:</strong>&nbsp;Up to $85/month for mobile and internet.</li> <li><strong>Disability &amp; life insurance:</strong>&nbsp;Company-paid short-term, long-term and life insurance coverage.</li> </ul> <p><span style="color: rgb(236, 240, 241);">#LI-LH2</span></p> <p>&nbsp;</p><div class="content-pay-transparency"><div class="pay-input"><div class="description"><p><strong><span style="font-size: 18px;">Pay Transparency</span></strong></p> <p>We offer competitive compensation and benefits packages. Actual compensation will be determined based on job-related factors, including experience, skills, qualifications, the level at which the candidate is hired, and geographic location, consistent with applicable law.</p></div><div class="title">Base Compensation Range</div><div class="pay-range"><span>$170,000</span><span class="divider">&mdash;</span><span>$300,000 USD</span></div></div></div><div class="content-conclusion"><p><strong>Benefits &amp; Perks:</strong></p> <ul> <li>Competitive compensation</li> <li>Career growth and learning opportunities</li> <li>Flexibility and ownership</li> <li>Collaborative and innovative culture</li> <li>Opportunity to work on impactful AI projects</li> <li>International environment and talented teams</li> </ul> <p><strong>What's it like to work at Nebius:</strong></p> <p>Fast moving&nbsp;- Bold thinking&nbsp;- Constant growth&nbsp;- Meaningful impact&nbsp;- Trust and real ownership&nbsp;- Opportunity to shape the future of AI&nbsp;</p> <p><strong>Equal Opportunity Statement:</strong></p> <p>Nebius is an equal opportunity employer. We are…
Skills asked for
- r
- linux
- c++
- go
- python
Similar jobs
- Staff Software Engineer Data - DC Tech LeadAfresh · Remote - United States
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.