Senior Site Reliability Engineer
Blinkhealth · India · 2026-06-30
About this role
<div class="content-intro"><p><strong>Company Overview:<br><br></strong><a href="https://www.blinkhealth.com/">Blink Health </a>is the fastest growing healthcare technology company that builds products to make prescriptions accessible and affordable to everybody.&nbsp; Our two primary products – BlinkRx and Quick Save – remove traditional roadblocks within the current prescription supply chain, resulting in better access to critical medications and improved health outcomes for patients.&nbsp;<br><br><a href="https://www.blinkhealth.com/blinkrx">BlinkRx </a>is the world’s first pharma-to-patient cloud that offers a digital concierge service for patients who are prescribed branded medications. Patients benefit from transparent low prices, free home delivery, and world-class support on this first-of-its-kind centralized platform. With BlinkRx, never again will a patient show up at the pharmacy only to discover that they can’t afford their medication, their doctor needs to fill out a form for them, or the pharmacy doesn’t have the medication in stock.&nbsp;<br><br>We are a highly collaborative team of builders and operators who invent new ways of working in an industry that historically has resisted innovation. Join us!</p></div><h3><strong>Responsibilities</strong></h3> <ul> <li><strong>Establish and evolve SRE best practices</strong> across the organization, including reliability principles, error budgets, incident response, postmortems, and operational readiness standards.<br><br></li> <li><strong>Define and drive observability strategy</strong> for system health, performance, and reliability, including SLIs/SLOs, alerting quality, dashboards, and service health indicators.<br><br></li> <li>Design and implement <strong>software-driven solutions within the infrastructure domain</strong>, automating manual processes and eliminating operational complexity and toil.<br><br></li> <li>Act as a <strong>technical leader and force multiplier</strong>, helping set priorities and influencing decision-making across core cloud infrastructure, reliability tooling, and platform architecture.<br><br></li> <li>Take ownership of <strong>large, ambiguous initiatives</strong>, driving them from concept to delivery while aligning stakeholders across engineering, security, and product.<br><br></li> <li>Combine deep knowledge of <strong>software development, infrastructure, and security</strong> to improve platform resilience, scalability, performance, and compliance.<br><br></li> <li>Proactively identify systemic risks and reliability gaps, <strong>recommending and leading platform upgrades</strong> and architectural improvements before they become incidents.<br><br></li> <li>Partner with engineering teams to improve <strong>developer workflows, tooling, and operational maturity</strong>, increasing productivity while reducing cognitive load.<br><br></li> <li>Provide <strong>technical mentorship</strong>, architecture guidance, and high-quality design and code reviews for engineers across infrastructure and product teams.<br><br></li> <li>Lead by example in <strong>documentation and knowledge sharing</strong>, ensuring systems and processes are well-understood and not dependent on individual ownership.<br><br></li> <li>Participate in and help mature <strong>incident response</strong>, escalation practices, and post-incident learning across the organization.<br><br></li> </ul> <h3><strong>Desired Experience</strong></h3> <ul> <li>Bachelor’s or Master’s degree in Computer Science or equivalent practical experience.<br><br></li> <li>10+ years of experience in site reliability engineering, infrastructure engineering, or platform engineering roles, with demonstrated impact at scale.<br><br></li> </ul> <h4><strong>Reliability &amp; Troubleshooting</strong></h4> <ul> <li>Expert-level, methodical troubleshooting across the <strong>entire stack</strong>, from application to kernel to network.<br><br></li> <li>Strong command-line proficiency and deep expertise in <strong>Linux systems and operating system fundamentals</strong>.<br><br></li> <li>Advanced understanding of networking concepts including <strong>load balancing, proxies, DNS, TCP/IP, NAT, and service-to-service communication</strong>.<br><br></li> </ul> <h4><strong>Software &amp; Automation</strong></h4> <ul> <li>Experience working across multiple languages (e.g., <strong>Python, Go, Bash</strong>, and familiarity troubleshooting application stacks such as React or similar).<br><br></li> <li>Strong track record of <strong>automating repetitive and complex operational work</strong> to reduce toil and increase reliability.<br><br></li> <li>Ability to design and build internal tools (Python or Go) that <strong>standardize and scale engineering practices</strong>.<br><br></li> <li>Comfortable operating in an <strong>agile environment</strong>, with disciplined testing and quality practices.<br><br></li> </ul> <h4><strong>Cloud &amp; Platform…
Skills asked for
- sre
- linux
- python
- go
- bash
- react
- agile
- aws
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.