Site Reliability Engineer
Nebius · Remote - United States · 2026-07-27
About this role
About Nebius:
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
The role
Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. You’re welcome to work in our office in Amsterdam.
Hardware Infrastructure team designs, develops and supports systems involved in the data-centers lifecycle:
• Serving functional and load testing system.
• Monitoring of engineering equipment located in our data centers (power supply, air and water cooling, etc.)
• Monitoring of IT equipment: racks, servers, JBODs, JBOGs, power shelves, network devices, etc.
• Asset tracking.
• Hardware repairs tasks tracking.
• Server production.
In this position, your responsibility will be to:
• Ensure fault-tolerance, scale and uninterrupted operations for our services.
• Use cutting-edge technology to solve a variety of infrastructure problems.
• Implement and improve CI/CD processes.
We expect you to have:
• Proficiency in Linux systems, with expertise in Python and Bash scripting for automation.
• Demonstrated ability to troubleshoot complex system issues, including hardware, software and networking problems.
• Strong analytical and problem-solving skills, with a focus on optimizing system performance.
• Working proficiency in English.
It would be an added bonus if you had:
• Desire to be involved in backend development.
• Experience designing, developing and running high-load distributed systems.
Working conditions:
• Primarily remote
• Occasional travel to data centers required, especially if not located near one
• Collaboration with globally distributed engineering and operations teams
Key employee benefits:
• Health insurance: 100% company-paid medical, dental, and vision coverage for employees and families
• 401(k) plan: up to 4% company match with immediate vesting
• Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers
• Remote work reimbursement: up to $85/month for mobile and internet
• Disability & life insurance: company-paid short-term, long-term, and life insurance coverage
Compensation
•
We offer competitive salaries, ranging from $130k- $180k base + quarterly performance bonuses.
Join Nebius and help operate the systems that power next-generation AI
infrastructure.
Benefits & Perks:
• Competitive compensation
• Career growth and learning opportunities
• Flexibility and ownership
• Collaborative and innovative culture
• Opportunity to work on impactful AI projects
• International environment and talented teams
What's it like to work at Nebius:
Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI
Equal Opportunity Statement:
Nebius is an equal opportunity employer.…
Skills asked for
- r
- ci/cd
- linux
- python
- bash
Similar jobs
- Senior Site Reliability Engineer, InfrastructureVultr · Remote - United States
- Staff Site Reliability EngineerDatavant2 · Remote - United States
- Site Reliability EngineerDatavant2 · Remote - United States
- Senior Site Reliability EngineerCribl · Remote - United States
- Staff Site Reliability EngineerAlphasense · Remote - United States
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.