AI Platform Support Engineer (APAC)
Lightningai · Philippines · 2026-07-31
About this role
<div class="content-intro"><h2>Who <strong>We Are</strong></h2> <p>Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.</p> <p>Through our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.</p> <p>We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.</p> <h2 class="PDq2pG_selectionAnchorContainer" data-section-id="o4nw1n" data-start="0" data-end="18">The Way We Work</h2> <p data-start="20" data-end="222">The people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:</p> <ul> <li data-section-id="vkpt18" data-start="319" data-end="461"><strong data-start="321" data-end="343">Move with Urgency:</strong> We move quickly, make thoughtful decisions, and keep momentum. We value action over perfection and learn by shipping.</li> <li data-section-id="1c7e9p5" data-start="462" data-end="598"><strong data-start="464" data-end="483">Take Ownership:</strong> We own outcomes, not just our individual work. We make decisions that move the company forward and follow through.</li> <li data-section-id="1wm6yoj" data-start="599" data-end="751"><strong data-start="601" data-end="624">Communicate Openly:</strong> We communicate directly, seek to understand, and create clarity for others. Honest conversations help us move faster together.</li> <li data-section-id="1w7531e" data-start="752" data-end="874"><strong data-start="754" data-end="776">Build Great Teams:</strong> We lead by example, empower others, and create healthy teams where people can do their best work.</li> <li data-section-id="jb6f5l" data-start="875" data-end="1036"><strong data-start="877" data-end="895">Raise the Bar:</strong> We're always improving ourselves. We learn from feedback, consistently challenge ourselves to grow, and focus on the work that matters most.</li> <li data-section-id="12lc0uc" data-start="1037" data-end="1184"><strong data-start="1039" data-end="1059">Think Long-Term:</strong> We design for what's next. We create scalable systems, simplify complexity, and use AI and automation to amplify our impact.</li> </ul> <div class="c-message_actions__container c-message__actions">&nbsp;</div></div><h2><strong>What We’re Looking For</strong></h2> <p>Lightning AI is looking to hire an <strong>AI Platform Support Engineer </strong>to join our APAC Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production environments.</p> <p>This role is not a ticket router or traditional support engineer. You are a technical partner to ML teams - helping diagnose failures, improve reliability, and guide customers through complex distributed systems problems.The problems range from Kubernetes scheduling and GPU orchestration to distributed PyTorch failures, inference latency, networking bottlenecks, storage performance, and platform reliability. You’ll gain exposure to a wide variety of real world AI workloads across industries and help shape the infrastructure powering the next generation of ML applications.</p> <p><em>This role is remote and open to candidates based in the Philippines. We are hiring for a Sunday-Wednesday shift schedule, with working hours from 7:00 AM to 5:00 PM local time (UTC+8).</em></p> <h2><br><strong>What You'll Do</strong></h2> <p><strong>Work Directly With ML Engineers</strong></p> <ul> <li>Partner directly with customer engineering teams running training and inference workloads in production</li> <li>Help customers diagnose and resolve complex distributed systems and ML infrastructure issues</li> <li>Act as a technical advisor during high impact incidents and platform degradation events</li> <li>Translate infrastructure level issues into actionable guidance for ML engineers</li> <li>Build credibility with customers through strong technical reasoning and clear communication</li> </ul> <p><strong>Debug ML Infrastructure &amp; Distributed Workloads</strong></p> <ul> <li>Investigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems</li> <li>Troubleshoot PyTorch, CUDA, NCCL, and inference serving related issues</li> <li>Analyze logs, metrics, traces, and system behavior to isolate root causes</li> <li>Debug containerized workloads running across Kubernetes and bare metal GPU…
Skills asked for
- pytorch
- kubernetes
- linux
- prometheus
- grafana
- machine learning
- python
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.