JobBobsBúsqueda de empleo en tiempo realEn vivo

Senior SRE Engineer

Codeway · Barcelona · 2026-07-15

FullTimeleadTeletrabajo
Inscríbete en la web de la empresa

Sobre el puesto

ABOUT CODEWAY

Codeway is a global consumer tech company with more than 400M users worldwide.

Since 2020, we’ve built and scaled 60+ mobile apps across creativity, productivity, wellness, language learning, and entertainment.

Our flagship apps — Retake AI, Cleanup, Learna, and DramaPops — and many of them lead their categories globally. In 2024, we became the most downloaded app publisher on iOS, driven by cutting-edge AI research, sharp data-driven execution, and a relentless focus on product and marketing.

We’re a team of 300+ people across İstanbul and Barcelona who bring curiosity, passion, trust, and ownership to everything we build. Recognized as a #1 LinkedIn Top Startup and a Great Place to Work in Europe, Codeway is where ambitious people do their life’s best work.

We’re building the next generation of consumer tech and reimagining what mobile apps can be.

This is Codeway. This is our way. Join us.

POSITION

We’re looking for a Senior Site Reliability Engineer to own and mature reliability, performance, and security across our growing platform. This role sits at the intersection of Engineering, Infrastructure, and Security, helping design, operate, and continuously improve the systems that keep dozens of consumer apps running for users around the world.

You’ll work closely with product engineering teams to make reliability measurable rather than assumed. That means defining and enforcing SLIs, SLOs, and error budgets; operating and hardening a multi-cluster Kubernetes environment; building the observability that catches problems before users feel them; and leading incident response when things break. It’s a hands-on role with real ownership over how reliability and security evolve as we scale.

Several parts of our reliability practice are still early in their maturity. We’re looking for someone who enjoys building the standards, processes, tooling, and automation that will form the foundation of our SRE function — not someone waiting for a playbook to already exist.

We welcome applicants from all backgrounds and experiences. If you’re excited about running systems at consumer scale and believe you could be a strong fit, we encourage you to apply, even if your experience doesn’t align perfectly with every qualification listed below.

WHAT YOU’LL BE DOING

Reliability, SLOs, Observability

- Define, instrument, and report on SLIs, SLOs, and error budgets across critical services, so reliability decisions are driven by data rather than opinion.

- Own observability end-to-end — metrics, logs, traces, dashboards, and alerting — and drive measurable reductions in detection and resolution times.

- Reduce alert noise and false positives so on-call engineers can trust what wakes them up.

- Run reliability reviews and an error-budget policy that shapes how teams prioritize between shipping and stability.

Kubernetes & Platform Operations

- Operate, scale, and upgrade our multi-cluster Kubernetes (GKE) environment: cluster lifecycle, autoscaling, networking, ingress, and resource management.

- Act as the deep-expertise escalation point for cluster and platform issues across dozens of services.

- Own capacity planning, performance, and cloud cost efficiency, balancing spend against reliability targets.

- Build self-service platform tooling that lets product teams move quickly without needing to become infrastructure experts.

Security & Resilience

- Embed security into the platform through RBAC and least-privilege, secrets management, image and dependency scanning, network policies, and a disciplined patching cadence.

- Partner with the security function on vulnerability remediation, audit readiness, and secure-by-default infrastructure.

- Own disaster recovery: define and regularly validate RTO/RPO targets through DR drills and failure testing.

- Contribute to architecture and production-readiness reviews so reliability and security are designed in, not bolted on.

Incident Response & Automation

- Lead the on-call rotation and act as incident commander during production incidents.

- Run blameless postmortems, quantify impact, and track corrective actions through to closure so the same failure doesn’t recur.

- Build and maintain Infrastructure as Code (Terraform) and CI/CD pipelines, enforcing GitOps and progressive delivery with automated rollbacks.

- Systematically identify, measure, and eliminate operational toil through automation, protecting engineering time for high-leverage work.

WHAT YOU’LL BRING?

- Experience operating high-traffic, always-on production systems at meaningful scale, typically gained over 5–8 years in SRE, Platform, or DevOps roles.

- Hands-on production Kubernetes experience — you’ve run clusters day to day, through upgrades, autoscaling, and real troubleshooting under load, not just deployed to them.

- A strong cloud engineering background, along with solid Linux and networking fundamentals.

- A track record of defining and operating with SLOs and error budgets, and comfort being measured on reliability outcomes.

- Experience with Infrastructure as Code and CI/CD pipeline design — you treat infrastructure and delivery as code.

- Depth in observability tooling: instrumentation, dashboarding, and alert design.

- A genuine security-first mindset, where least-privilege, secrets hygiene, and vulnerability management are habits rather than afterthoughts.

- Scripting and automation fluency in at least one language, used to build tooling and remove toil.

- Incident-command experience: owning on-call, running blameless postmortems, and driving resolution times down over time.

- Ability to communicate clearly with both engineers and leadership, especially under pressure.

NICE TO HAVE

- Experience with high-scale consumer or mobile app backends, or with AI/ML inference workloads and their scaling characteristics.

- Experience with GitOps and progressive-delivery patterns such as canary and…

Competencias solicitadas

Inscríbete en la web de la empresa

Tu próximo puesto ya está aquí.

Busca ofertas en directo de miles de empresas, guarda las que merecen una segunda mirada y deja que JobBob vigile el resto.