AI Platform Support Engineer (EMEA)
Lightningai · London, England, United Kingdom · 2026-07-29
About this role
Who We Are
Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.
Through our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.
We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.
The Way We Work
The people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:
• Move with Urgency: We move quickly, make thoughtful decisions, and keep momentum. We value action over perfection and learn by shipping.
• Take Ownership: We own outcomes, not just our individual work. We make decisions that move the company forward and follow through.
• Communicate Openly: We communicate directly, seek to understand, and create clarity for others. Honest conversations help us move faster together.
• Build Great Teams: We lead by example, empower others, and create healthy teams where people can do their best work.
• Raise the Bar: We're always improving ourselves. We learn from feedback, consistently challenge ourselves to grow, and focus on the work that matters most.
• Think Long-Term: We design for what's next. We create scalable systems, simplify complexity, and use AI and automation to amplify our impact.
What We’re Looking For
Lightning AI is looking to hire AI Platform Support Engineers to join our EMEA Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production environments.
This role is not a ticket router or traditional support engineer. You are a technical partner to ML teams - helping diagnose failures, improve reliability, and guide customers through complex distributed systems problems.The problems range from Kubernetes scheduling and GPU orchestration to distributed PyTorch failures, inference latency, networking bottlenecks, storage performance, and platform reliability. You’ll gain exposure to a wide variety of real world AI workloads across industries and help shape the infrastructure powering the next generation of ML applications.
We are currently hiring for two EMEA shifts (9AM–7PM CET/CEST):
• Saturday–Tuesday
• Thursday–Sunday
This role is hybrid out of our London office, with an in-office requirement of at least 2 days per week and occasional team and company offsites. We are not able to provide visa sponsorship for this role at this time.
What You'll Do
Work Directly With ML Engineers
• Partner directly with customer engineering teams running training and inference workloads in production
• Help customers diagnose and resolve complex distributed systems and ML infrastructure issues
• Act as a technical advisor during high impact incidents and platform degradation events
• Translate infrastructure level issues into actionable guidance for ML…
Skills asked for
- pytorch
- kubernetes
- linux
- prometheus
- grafana
- machine learning
- python
Similar jobs
- Digital Platform Support AnalystEvelyn Partners · London
- Wealth Platform Support AnalystEvelyn Partners · London
- Platform Support AnalystEntain · London
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.