AI Infrastructure Engineer
Aifoundry · Seattle, WA · 2026-05-20
About this role
<p><strong>Compensation:</strong> $200K - $325K+</p> <p><strong>Travel Requirement:</strong> International travel required to India 8+ times per year</p> <p><strong><em>Multiple Openings:</em></strong><em> We are hiring multiple AI Infrastructure Engineers across data center infrastructure teams. Exact level, scope, and team placement may vary based on experience and business needs.</em></p> <p>&nbsp;</p> <p>AI Foundry is an AI building and consulting business working with some of the largest companies in the world. AI Foundry is partnered with a major global company operating across finance, consumer, and retail verticals, leading their AI efforts from the ground up.</p> <p>The team is seeking builders who want to be part of something massive: designing and building multi-agent AI systems that will reshape how hundreds of millions of people interact with financial services, retail, insurance, and consumer products.</p> <p>AI Foundry is looking for AI Infrastructure Engineers to help build and operate large-scale GPU infrastructure for training and inference workloads. This is hands-on systems work across Linux, networking, storage, orchestration, observability, and automation in a data center environment.</p> <p><em>If you want to build the infrastructure foundation for serious AI workloads rather than operate a small lab cluster, we want to hear from you.</em></p> <h3>What You'll Do</h3> <ul> <li>Build, configure, and operate GPU cluster infrastructure for training and inference workloads.</li> <li>Work across Linux systems, high-performance networking, storage, orchestration, monitoring, and automation.</li> <li>Support provisioning, configuration management, job scheduling, workload management, and lifecycle operations.</li> <li>Implement observability, alerting, incident response, and operational runbooks for infrastructure health.</li> <li>Partner with hardware, facilities, vendors, and engineering teams to resolve performance, reliability, and capacity issues.</li> <li>Automate repetitive operational tasks and improve the reliability of cluster operations.</li> <li>Help evaluate hardware, networking, storage, and platform components for AI workloads.</li> <li>Document systems clearly so global teams can operate and troubleshoot consistently.</li> <li>Travel to India 8+ times per year to work directly with infrastructure and client teams.</li> </ul> <h3>Who You Are</h3> <ul> <li>You have strong Linux systems engineering experience and enjoy hands-on infrastructure work.</li> <li>You understand GPU infrastructure or are deeply motivated to build expertise in NVIDIA systems, CUDA, NCCL, and distributed AI workloads.</li> <li>You have experience with networking, storage, Kubernetes, Slurm, observability, automation, or related infrastructure tooling.</li> <li>You can troubleshoot complex systems methodically and communicate what you are seeing.</li> <li>You care about reliability, operational clarity, and repeatable systems.</li> <li>You are comfortable working with incomplete information and learning quickly.</li> <li>You can collaborate with infrastructure, facilities, vendor, and engineering teams across time zones.</li> <li>You write useful documentation and runbooks.</li> <li>You are energized by building greenfield infrastructure at serious scale.</li> <li>You use AI and modern tools to improve how you build, debug, document, and operate systems.</li> </ul> <h3>How We Work</h3> <p>We are looking for people who use AI to raise the speed, quality, and ambition of the work while staying grounded in customers, teammates, and measurable outcomes.</p> <ul> <li>We are hands-on with AI and technical tools, and we make ideas tangible through prototypes, clear artifacts, and working systems.</li> <li>We move quickly, share work early, and keep refining until the product experience feels useful, polished, and memorable.</li> <li>We focus on the hardest problems and the outcomes that matter, with the judgment to change course when the evidence points somewhere better.</li> <li>We communicate clearly enough that teammates, leaders, and AI-assisted workflows can act on our thinking across time zones.</li> <li>We collaborate with openness and respect, build on other people's ideas, and treat shared ownership as part of doing excellent work.</li> <li>We speak up directly and constructively when something looks wrong, then commit fully once a decision is made.</li> </ul> <h3>Benefits &amp; More</h3> <ul> <li>Medical and Dental benefits</li> <li>401K</li> <li>Opportunity to shape AI strategy for 500M+ users</li> </ul>
Skills asked for
- linux
- kubernetes
Similar jobs
- Senior Software Engineer, InfrastructureDocker · Seattle
- Software Development Engineer II - InfrastructureAmperity · Seattle
- Software Engineer - InfrastructureMotherduck · Seattle
- Software Engineer, Infrastructure - Core ExperimentationOpenAI · Seattle
- Senior Staff Backend Engineer - Cloud InfrastructureCoupang · Seattle
- VP, Infrastructure EngineeringDocker · Seattle
- Senior Test Infrastructure Software EngineerEndurance Energy · Seattle
- Staff Infrastructure EngineerAurelian · Seattle
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.