JobBobsReal-time global job discoveryLive

Tech Lead Manager, GPU Cluster Infrastructure

Far Ai · Remote (International) · 2026-09-11

FullTimeexecutiveRemote
Apply on the employer's site

About this role

ABOUT US

FAR.AI http://FAR.AI is a non-profit AI research institute working to ensure advanced AI is safe and beneficial for everyone. Our mission is to facilitate breakthrough AI safety research, advance global understanding of AI risks and solutions, and foster a coordinated global response.

We’re structured to support that work from early research through real-world adoption:

- Independent by design. We can pursue what's most impactful based on our theory of change and share what we find publicly.

- A portfolio approach. Rather than focus on one single direction, we run diverse bets across the safety stack. We take promising ideas from initial experiments to deployment, informed by red-team partnerships with frontier labs and governments.

- Serious infrastructure for ambitious research. A dedicated engineering team runs our compute cluster and experiment-scaling stack, so researchers spend their time on research instead of on infra.

- Setting the standard. Our events convene key decision makers; our red-team works with frontier developers and governments; and our communications inform the public. Together, this drives adoption and sets the new standard in safety.

Since our founding in July 2022, we've grown to 50+ staff https://www.far.ai/about/team, published 40+ academic papers https://scholar.google.com/citations?user=FVJ24k8AAAAJ, and convened leading AI safety events https://far.ai/events/. Our work is recognized globally, with publications at premier venues such as NeurIPS, ICML including a Best Paper Honorable Mention in 2026 https://icml.cc/virtual/2026/oral/71065, and ICLR, and features in the Financial Times https://www.ft.com/content/175e5314-a7f7-4741-a786-273219f433a1, Nature News https://www.nature.com/articles/d41586-024-02218-7, Wired Magazine https://www.wired.com/story/jailbreaking-ai-models-google-anthropic-openai-spacexai/ and MIT Technology Review https://www.technologyreview.com/2020/02/28/905615/reinforcement-learning-adversarial-attack-gaming-ai-deepmind-alphazero-selfdriving-cars/. We conduct pre-deployment testing on behalf of frontier developers such as OpenAI and independent evaluations for governments including the EU AI Office https://www.far.ai/news/far-ai-selected-to-lead-eu-ai-act-cbrn-risk-consortium and publish the AI Security Leaderboard https://leaderboard.far.ai/ based on our red-teaming expertise. We help steer and grow the AI safety field through developing https://arxiv.org/abs/2405.06624 research https://arxiv.org/abs/2506.20702 roadmaps https://www.researchgate.net/publication/396910034_Open_Technical_Problems_in_Open-Weight_AI_Model_Risk_Management with renowned researchers such as Yoshua Bengio; running FAR.Labs https://www.far.ai/programs/far-labs, an AI safety-focused co-working space in Berkeley housing 40+ members; and supporting the community through targeted grants https://www.far.ai/programs/grantmaking to technical researchers.

ABOUT THE TEAM

Foundations is FAR.AI http://FAR.AI's infrastructure and engineering team. Our remit is broad: we run the compute platform, build the tools and frameworks researchers work in, automate research workflows, and help teams scale experiments well past what they'd manage alone. Our job is to accelerate the research. We do so by working directly with researchers through embedded engagements and day-to-day consulting, and building systems that can scale with the organization as it grows.

Foundations is growing quickly, and our infrastructure portfolio is growing fastest. We run FAR.AI http://FAR.AI's research on a mix of bare-metal and managed Kubernetes GPU clusters from multiple providers. We rent the hardware and operate the platform ourselves. The fleet has grown from dozens to hundreds of GPUs this year and it's continuing to grow quickly: we're adding providers, taking on users beyond our own researchers, and moving experiments onto frontier open-weight models. A large amount of research is now being done by AI agents working directly on the cluster, which is driving updates to our platform infrastructure and security.

Running it well now takes dedicated specialists, so we're standing up an infrastructure sub-team that owns the cluster fleet: adding capacity, designing and managing the networking and storage under it, infrastructure as code, and the security posture, plus some of the platform layer above it. It works directly with research teams as their needs change.

ABOUT THE ROLE

You'd be the infrastructure sub-team's lead and one of its engineers, with at least half your time on technical work. You'd own the platform's technical direction and roadmap, deciding what we build and how we run it in collaboration with the research teams and the rest of Foundations, and you'd hire and grow the team.

In frontier AI research, working out the infrastructure is often part of the science. You'd work directly with researchers and other engineers to keep our large-scale experiments performant and fault-tolerant.

We're also hiring a Software Engineer, GPU Cluster Infrastructure https://jobs.ashbyhq.com/far.ai/a94080f7-288b-4a98-ad4a-aaaa4bbf9968for this team. If you want the hands-on work without the management responsibilities, take a look there instead.

WHAT YOU'LL DO

- Set the platform's technical direction and own its roadmap. You decide which systems we run, how we schedule and store across providers, and what we measure, from utilization and queue wait to failure rates.

- Stay in the work. You own architecture and the scheduling and storage design, and you debug the failures that cross layers, such as node health, GPU and fabric faults, and multi-node job hangs.

- Hire and grow a small team of senior engineers. You set priorities and ownership, scope projects, and give regular feedback and coaching.

- Set security direction for a shared cluster where AI agents run experiments, covering identity and access, workload isolation, and sandboxing.

- Decide…

Skills asked for

Apply on the employer's site

Your next role is already in here.

Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.