Sr. Site Reliability Engineer
Backblaze · Remote - US · 2026-04-01
About this role
<p><span style="font-family: helvetica, arial, sans-serif;"><strong>About Backblaze</strong></span></p> <p>Backblaze is the object storage leader in the open cloud movement, fueling customer success with cloud storage built purposefully to unlock budgets, unburden administrators, and unleash innovators. Together with our partners, we’re helping customers break free from the restrictive, overpriced legacy solutions that hold them back, and blaze forward with the full power of the open cloud in their hands.</p> <p>Founded in 2007, we scaled the business with less than $3 million in outside funding until 2021, when we did a traditional IPO on the Nasdaq stock exchange. Today, Backblaze generates over $100m in revenue and is the leading specialized storage cloud - managing over three billion gigabytes of data storage for 500K+ customers in 175+ countries, including businesses, developers, IT professionals, and individuals.<br><br>But while there is a lot to celebrate in our past, there is almost as much opportunity ahead of us. We’re seeking a<strong> Sr. Site Reliability Engineer </strong>to join our team!</p> <p><span style="font-family: helvetica, arial, sans-serif;"><strong>About the Role:</strong></span></p> <p><span style="font-family: helvetica, arial, sans-serif;">We are seeking a Senior Site Reliability Engineer (SRE) to help ensure the stability, scalability, and reliability of our services and infrastructure. This role focuses on building automation, maintaining observability, and supporting incident response to keep customer-facing systems performing at their best.</span><br><br><span style="font-family: helvetica, arial, sans-serif;">The SRE will collaborate with engineering, product, and operations teams to embed reliability practices into day-to-day development and operations while contributing to tools and processes that improve efficiency and reduce manual effort.</span></p> <p><strong><span style="font-family: helvetica, arial, sans-serif;">What You'll Do:</span></strong></p> <ul> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Service Reliability &amp; Operations</span></li> <ul> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Own and drive the availability, durability, and performance of critical services across all production environments.</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Lead and champion complex projects from problem discovery through complete, cross-functional resolution, demonstrating high-level technical ownership.</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Define, establish, and enforce service health standards, including working with engineering leadership to implement SLIs, SLOs, and error budget policies for multiple services.</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Lead critical incident response and post-incident reviews, translating findings into strategic, long-term service improvements and architectural changes.</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Mentor others and act as a subject matter expert in following and evolving established ITIL/OSS processes (incident, change, problem, and capacity management).</span></li> </ul> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Automation &amp; Tooling</span></li> <ul> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Design and architect scalable automation solutions to eliminate toil and improve the efficiency of operational tasks across the entire platform.</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Drive the strategic direction of monitoring, logging, and alerting frameworks (e.g., Prometheus, Grafana, Catchpoint, ELK), and integrate them for comprehensive observability.</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Build, maintain, and secure advanced CI/CD pipelines, configuration management, and complex infrastructure as code solutions (Terraform, Ansible, Jenkins).</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Write production-grade code (Bash, Python, Go, etc.) to develop new reliability tools and enhance existing systems.</span></li> </ul> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Collaboration</span></li> <ul> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Act as a principal partner to engineering, product,…
Skills asked for
- sre
- prometheus
- grafana
- ci/cd
- terraform
- ansible
- jenkins
- bash
Similar jobs
- Site Reliability EngineerMyfitnesspal · Remote - US
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.