Sr. Director, Back-End Engineering
Coupang · Seoul, South Korea · 2026-08-23
About this role
Sr. Director, Site Reliability Engineering
Coupang operates one of the largest and most complex technology platforms in the world. We are seeking a Senior Director, Site Reliability Engineering (Head of SRE) to define and lead company-wide reliability, resilience, scalability, and operational excellence. This leader will transform reliability from a collection of team-specific practices into platform mechanisms that services inherit by tier, while advancing incident response toward an intelligent, AI-assisted, and increasingly autonomous operating model. We are looking for a visionary, industry-recognized technology leader who has previously conceived, built, and scaled a comparable SRE, production engineering, resilience, or autonomous-operations organization at a leading global technology company. The successful candidate must combine deep technical credibility with the organizational leadership required to align executives, influence architecture across the company, and build a world-class leadership bench.
Key Responsibilities
• Set a bold, multi-year vision for company-wide reliability, resilience, and autonomous operations, and translate that vision into an executable roadmap with measurable business outcomes.
• Define and own the SRE strategy, operating model, engineering standards, and reliability governance across Coupang.
• Build platform mechanisms that allow services to inherit reliability requirements based on service tier rather than recreate them independently.
• Lead initiatives that materially improve availability, resilience, scalability, performance, and operational readiness.
• Partner with engineering, product, infrastructure, security, finance, and business leaders to align reliability investments with customer and business priorities.
• Own executive reliability metrics, including availability, detection and recovery performance, change risk, incident recurrence, capacity readiness, and operational toil.
• Build and scale a world-class SRE organization capable of influencing engineering practices across the company.
Reliability Strategy, SLOs & Engineering Governance
• Establish and evolve service-tier definitions, SLOs, SLAs, error budgets, reliability scorecards, and objective certification mechanisms such as RBD/RBO.
• Create clear reliability requirements for Tier 0, Tier 1, and Tier 2 services, including redundancy, load testing, disaster recovery, observability, and incident response.
• Ensure reliability governance is embedded in architecture, development, release, and production operations rather than applied as a final review.
• Drive systematic reduction of recurring incidents, reliability risks, operational debt, and unsafe change patterns.
• Influence company-wide architecture for graceful degradation, fault isolation, load shedding, circuit breaking, and failure containment.
Incident Management & Autonomous Operations
• Transform incident management into a fast, disciplined, data-driven, and increasingly autonomous operating model.
• Enable AI-assisted detection, event correlation, triage, escalation, root-cause drafting, remediation recommendations, and selected guardrailed auto-remediation.
• Improve incident command, on-call quality, escalation mechanisms, communication, post-incident learning, and corrective-action completion.
• Reduce noisy alerts, manual on-call work, repeated diagnosis, and time spent coordinating across fragmented systems.
• Use incident and telemetry data to continuously improve platform standards, testing, capacity models, and engineering roadmaps.
Disaster Recovery, Resilience & Capacity
• Own the strategy and execution model for disaster recovery, regional resilience, availability-zone loss, capacity-constrained recovery, and critical business continuity.
• Build reusable DR and failover mechanisms that services inherit from the platform rather than implement as bespoke projects.
• Establish objective RPO/RTO targets, automated readiness gates, regular game days, fault injection, and evidence-based recovery certification.
• Drive proactive and intelligent capacity management using forecasting, reservations, workload prioritization, and automated response to demand and failure scenarios.
• Partner with compute, traffic, networking, storage, and application leaders to enable safe zone evacuation, regional failover, and surge readiness.
Observability, Testing & Reliability Intelligence
• Partner with Observability and TestOps leaders to integrate logs, metrics, traces, continuous profiling, testing, and incident intelligence into one reliability feedback loop.
• Ensure every critical service has actionable telemetry, meaningful SLOs, release-quality signals, and production-readiness evidence.
• Use production incidents and operational patterns to drive targeted integration, load, resilience, and regression testing.
• Establish executive reliability dashboards that provide trusted views of service health, risk, capacity, and operational effectiveness.
Talent Leadership & Organization
• Lead multiple layers of SRE leaders, including senior managers, directors, principal engineers, and senior individual contributors.
• Own organizational design, global hiring strategy, leadership development, succession planning, and the creation of a strong leadership bench.
• Attract exceptional SRE, distributed systems, resilience, incident-management, and capacity-engineering talent from best-in-class technology organizations.
• Build an empowered organization with clear accountability, strong technical judgment, high execution velocity, and a company-wide perspective.
• Act as a force multiplier by mentoring technical and organizational leaders and raising reliability capabilities across engineering.
Technical Leadership & Architecture
• Own reliability architecture decisions across large-scale distributed systems and cloud-native infrastructure.
• Define resilient patterns for redundancy, failover,…
Skills asked for
- sre
- kubernetes
- ci/cd
Similar jobs
- Sr. Director, Back-End EngineeringCoupanginternal · Seoul
- Director, Back-end Engineering (Rocket Pay)Coupang · Seoul
- Director, Back-End Engineering (E-commerce Engineering)Coupang · Seoul
- Director, Back-End EngineeringCoupanginternal · Seoul
- [Coupang Pay] Director, Back-End EngineeringCoupanginternal · Seoul
- [Coupang Pay] Director, Back-End EngineeringCoupang · Seoul
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.