Sr. Staff Observability Engineer (GPU Cloud & Telemetry Platform)
Coupanginternal · Seoul, South Korea · 2026-08-01
About this role
<p>본 공고는 재직 임직원을 대상으로 하는 사내 공모 전용입니다. (임직원 추천은 <span style="text-decoration: underline;"><a href="https://coupang.eightfold.ai/refer?pid=38569611&amp;domain=coupang.com&amp;triggerGoButton=false">Link</a></span>로 진행)</p> <p>This posting is exclusively for internal employees. (Employee referrals are submitted via the&nbsp;<span style="text-decoration: underline;"><a href="https://coupang.eightfold.ai/refer?query=%2A&amp;pid=38569611&amp;domain=coupang.com">Link</a></span>)</p> <p><br>지원 시에는 반드시 첨부된 ‘<a href="https://coupang.service-now.com/sp?id=kb_article&amp;sysparm_article=KB0010204"><span style="text-decoration: underline;">사내 공모 지원서 양식</span></a>’을 작성한 후, <span style="text-decoration: underline;"><strong>쿠팡 이메일 계정</strong></span>으로 접수해 주시기 바랍니다.</p> <p>To apply, please complete the attached ‘<a href="https://coupang.service-now.com/sp?id=kb_article&amp;sysparm_article=KB0010204"><span style="text-decoration: underline;">Internal Transfer Request Form</span></a>’ and submit it via <strong><span style="text-decoration: underline;">your Coupang email address</span></strong>.</p> <p>• 사내 공모 정책: [<span style="text-decoration: underline;"><a href="https://coupang.service-now.com/sp?id=kb_article_view&amp;sys_id=c1f29839475cc7504b66c9aa216d4316&amp;table=kb_knowledge&amp;searchTerm=%EC%82%AC%EB%82%B4%EA%B3%B5%EB%AA%A8%EC%A0%95%EC%B1%85">Link</a></span>]</p> <p>• Internal Transfer Policy: [<a href="https://coupang.service-now.com/sp?id=kb_article_view&amp;sys_id=c1f29839475cc7504b66c9aa216d4316&amp;table=kb_knowledge&amp;searchTerm=%EC%82%AC%EB%82%B4%EA%B3%B5%EB%AA%A8%EC%A0%95%EC%B1%85"><span style="text-decoration: underline;">Link</span></a>]</p> <p><span data-contrast="auto">&nbsp;</span></p> <hr> <p><strong>About Coupang</strong></p> <p>We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, we’re collectively disrupting the multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce.</p> <p>We are proud to have the best of both worlds — a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurial surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day.</p> <p>Our mission to build the future of commerce is real. We push the boundaries of what’s possible to solve problems and break traditional tradeoffs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world.</p> <p>&nbsp;</p> <p><strong><span data-ccp-props="{&quot;134233279&quot;:true,&quot;134245417&quot;:false,&quot;335559738&quot;:240,&quot;335559739&quot;:240}">Role Overview</span></strong></p> <div data-olk-copy-source="MessageBody">We are seeking a&nbsp;<strong>Sr. Staff Observability Engineer</strong>&nbsp;to lead the design and evolution of our observability platform for a&nbsp;<strong>GPU-as-a-Service (GPUaaS)</strong>&nbsp;infrastructure. This role will own the end-to-end telemetry strategy—from high-throughput metric ingestion to log pipelines and real-time visualization—powering deep insights into GPU clusters, datacenter systems, and distributed workloads.</div> <div>You will architect and operate&nbsp;<strong>planet-scale telemetry pipelines leveraging Grafana Alloy, Mimir, Loki, and Vector</strong>, ensuring high-fidelity observability across GPU workloads, Kubernetes clusters, and datacenter infrastructure.</div> <div>&nbsp;</div> <div><hr></div> <div><strong data-olk-copy-source="MessageBody">Key Responsibilities</strong></div> <div>&nbsp;</div> <div><strong>&lt;Architectural Leadership &amp; Strategy&gt;</strong></div> <div> <ul> <li><strong>End-to-End Observability Platform Ownership</strong>: Design and scale telemetry pipelines using: <ul> <li><strong>Grafana Alloy</strong> for metrics collection (Prometheus-compatible pipelines)</li> <li><strong>Datadog Vector</strong> for high-throughput log ingestion and transformation</li> <li><strong>Grafana Mimir</strong> for scalable time-series storage</li> <li><strong>Grafana Loki</strong> for log aggregation and querying</li> </ul> </li> <li><strong>Strategic Roadmap</strong>: Define the multi-year vision for GPU infrastructure observability, transitioning from reactive monitoring to&nbsp;<strong>SLO-driven, predictive, and automated observability</strong>.</li> <li><strong>High-Cardinality Telemetry…
Skills asked for
- grafana
- kubernetes
- prometheus
- datadog
- sre
- ci/cd
- terraform
- go
Similar jobs
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.