Senior Manager, Site Reliability Engineering
Claritas Rx · United States · 2026-07-09
About this role
Who We Are Claritas Rx uses AI and predictive modeling to help rare disease and specialty brands remove the barriers that keep patients from accessing and staying on the treatments they need. By uniting the most complete view of the patient journey with purpose-built technologies, we predict and resolve access challenges before they disrupt care, combining advanced analytics, real-world data, AI, and CRM capabilities to increase start and refill rates, reduce abandonment, and improve brand performance. Our mission is to ensure patients with chronic, life-threatening diseases receive the support that enables the greatest benefit from their therapy. Simply put, our promise is progress for every patient journey. This is the opportunity to help shape a first-in-industry digital health solution alongside a team of mission-driven professionals. We were named one of Inc.'s Best Workplaces in 2025 and recognized on the 2025 Inc. 5000 list, and our team genuinely respects and supports each other. We thrive on being fast-paced, innovative, and results-driven, and our employees enjoy a flexible, collaborative work environment, unlimited PTO, stock options, and a growing set of tools and technology to drive innovation for our customers.
The Position We are seeking an experienced and hands-on Sr. Manager of Site Reliability Engineering to lead our SRE and DevOps function. Reporting to the Sr. Director of Software Engineering, you will be responsible for the reliability, availability, scalability, and operational excellence of our AWS-hosted SaaS platform — a system that handles sensitive patient and commercial data for some of the world's leading biopharmaceutical companies. You will lead a small, high-impact team of full-time SRE/DevOps engineers and offshore contractors, setting technical direction, establishing operational standards, and rolling up your sleeves to solve hard infrastructure and reliability problems alongside your team. This is a player-coach role: you are expected to lead by example, contribute directly to infrastructure design and tooling, and grow a team that the broader engineering organization depends on. You will partner closely with Software Engineering, Product Management, and Security to ensure our platform operates with the reliability and compliance posture our customers and regulators require.
Key Accountabilities Platform Reliability & Operations Own the reliability, availability, performance, and scalability of Claritas Rx's AWS-hosted SaaS platform, with accountability for SLA/SLO commitments made to customers. Define and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across all production services; use error budgets to drive engineering prioritization conversations. Lead incident response and on-call operations: triage, coordinate resolution, communicate to stakeholders, and conduct thorough post-incident reviews with actionable corrective actions. Drive a proactive reliability culture — identifying risks before they become incidents through load testing, chaos engineering, and systematic failure mode analysis. Infrastructure & Cloud Engineering Architect, build, and maintain AWS cloud-native infrastructure using infrastructure-as-code (AWS CDK, Terraform, or equivalent), ensuring environments are reproducible, auditable, and secure. Oversee and continuously improve CI/CD pipelines (GitHub Actions) to enable rapid, safe, and consistent delivery of application and infrastructure changes across environments. Manage and optimize core AWS services including ECS, EC2, Aurora RDS, DynamoDB, Lambda, S3, SQS, EventBridge, Cognito, Secrets Manager, and CloudFront. Ensure robust observability across the stack — centralizing logs, metrics, traces, and alerts using CloudWatch, Sentry, and related tooling — so the team can detect and respond to issues quickly. Manage platform capacity planning, cost optimization, and cloud spend governance. Security, Compliance & Data Protection Ensure all infrastructure design and operational practices meet HIPAA, SOC 2, and HITRUST requirements, given the PHI our platform processes. Partner with the Security function on vulnerability management, infrastructure hardening, secrets management, and access control. Maintain and regularly test disaster recovery (DR) and business continuity plans, including defined RTO/RPO targets for all production systems. Support audit readiness and evidence collection for compliance certifications. Team Leadership & Offshore Coordination Lead, mentor, and grow a blended team of full-time SREs/DevOps engineers and offshore contractors, fostering a culture of ownership, continuous improvement, and operational excellence. Manage distributed team dynamics effectively — establishing clear communication rhythms, documentation standards, and handoff protocols to ensure offshore resources are productive and well-integrated. Conduct regular 1:1s, set clear goals and development plans for direct reports, and advocate for your team's growth and recognition. Build and maintain a healthy on-call rotation with appropriate tooling, runbooks, and escalation paths to protect team sustainability. Cross-Functional Partnership Collaborate with Software Engineering teams to embed reliability practices into the SDLC — including production readiness reviews, deployment standards, and shared observability tooling. Work with Product Management and Engineering leadership to balance feature delivery velocity against operational risk and technical debt. Contribute to architecture decisions across the platform, providing an operational and reliability perspective on new services and major technical changes. Champion automation-first thinking: eliminate toil through tooling, scripting, and process improvement wherever possible.
Our Stack Cloud: AWS (ECS, EC2, Aurora RDS, DynamoDB, Lambda, S3, SQS, EventBridge, Cognito, Secrets Manager, CloudFront, WAF) Infrastructure as Code: AWS CDK (primary); familiarity…
Skills asked for
- sre
- devops
- aws
- terraform
- ci/cd
- github actions
- dynamodb
- nestjs
Similar jobs
- Technology Risk Senior ManagerFuku · Singapore
- Senior Sales & Marketing Manager Property DeveloperFuku · Kuala Lumpur
- Senior Property ManagerFuku · Singapore
- Senior Project Manager Railway / Rail InfrastructureFuku · Kuala Lumpur
- Senior Project Manager - MEP & Mission CriticalFuku · Johor Bahru
- Senior Product Manager, SaaSFuku · Bangkok
- Senior Product Manager Growth & RetentionFuku · Kuala Lumpur
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.