AI Platform Support Engineer (US)
Lightningai · New York, New York, United States; San Francisco, California, United States; Seattle, Washington, United States · 2026-07-29
About this role
<div class="content-intro"><h2>Who <strong>We Are</strong></h2> <p>Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.</p> <p>Through our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.</p> <p>We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.</p> <h2 class="PDq2pG_selectionAnchorContainer" data-section-id="o4nw1n" data-start="0" data-end="18">The Way We Work</h2> <p data-start="20" data-end="222">The people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:</p> <ul> <li data-section-id="vkpt18" data-start="319" data-end="461"><strong data-start="321" data-end="343">Move with Urgency:</strong> We move quickly, make thoughtful decisions, and keep momentum. We value action over perfection and learn by shipping.</li> <li data-section-id="1c7e9p5" data-start="462" data-end="598"><strong data-start="464" data-end="483">Take Ownership:</strong> We own outcomes, not just our individual work. We make decisions that move the company forward and follow through.</li> <li data-section-id="1wm6yoj" data-start="599" data-end="751"><strong data-start="601" data-end="624">Communicate Openly:</strong> We communicate directly, seek to understand, and create clarity for others. Honest conversations help us move faster together.</li> <li data-section-id="1w7531e" data-start="752" data-end="874"><strong data-start="754" data-end="776">Build Great Teams:</strong> We lead by example, empower others, and create healthy teams where people can do their best work.</li> <li data-section-id="jb6f5l" data-start="875" data-end="1036"><strong data-start="877" data-end="895">Raise the Bar:</strong> We're always improving ourselves. We learn from feedback, consistently challenge ourselves to grow, and focus on the work that matters most.</li> <li data-section-id="12lc0uc" data-start="1037" data-end="1184"><strong data-start="1039" data-end="1059">Think Long-Term:</strong> We design for what's next. We create scalable systems, simplify complexity, and use AI and automation to amplify our impact.</li> </ul> <div class="c-message_actions__container c-message__actions">&nbsp;</div></div><h2><strong>What We’re Looking For</strong></h2> <p>Lightning AI is looking to hire an <strong>AI</strong> <strong>Platform Support Engineer&nbsp;</strong>to join our US Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production environments.</p> <p>This role sits at the intersection of ML systems, cloud infrastructure, Kubernetes, and customers. You’ll support engineers training models, deploying inference systems, and scaling GPU workloads in production.You are not a ticket router or traditional support engineer. You are a technical partner to ML teams - helping diagnose failures, improve reliability, and guide customers through complex distributed systems problems.</p> <p>The problems range from Kubernetes scheduling and GPU orchestration to distributed PyTorch failures, inference latency, networking bottlenecks, storage performance, and platform reliability. You’ll gain exposure to a wide variety of real world AI workloads across industries and help shape the infrastructure powering the next generation of ML applications.</p> <p><em>This role is hybrid out of our Seattle, San Francisco, or New York office hubs, with an in-office requirement of at least 2 days per week and occasional team and company offsites. The role follows a Monday–Friday schedule, with working hours from 8:00 AM to 5:00 PM PST. We are not able to provide visa sponsorship for this role at this time.&nbsp;</em></p> <h2><br><strong>What You'll Do</strong></h2> <p><strong>Work Directly With ML Engineers</strong></p> <ul> <li>Partner directly with customer engineering teams running training and inference workloads in production</li> <li>Help customers diagnose and resolve complex distributed systems and ML infrastructure issues</li> <li>Act as a technical advisor during high impact incidents and platform degradation events</li> <li>Translate infrastructure level issues into actionable guidance for ML engineers</li> <li>Build credibility with customers through strong technical reasoning and clear communication</li> </ul> <p><strong>Debug ML Infrastructure &amp; Distributed…
Skills asked for
- pytorch
- kubernetes
- linux
- prometheus
- grafana
- machine learning
- python
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.