JobBobsReal-time global job discoveryLive

Staff Site Reliability Engineer

Arcadiacareers · Chennai, Tamil Nadu, India · 2026-06-25

executive
Apply on the employer's site

About this role

Staff Site Reliability Engineer

Who we are:

Arcadia is the AI-powered energy intelligence platform for businesses. We replace fragmented tools and manual workflows with one platform to pay utility bills, buy energy, and advance sustainability — across every location, at enterprise scale.

Trusted by Fortune 2000 companies, Arcadia combines unified data, AI-powered analytics, and expert advisory to help enterprise teams save money, mitigate risk, and cut carbon.

We deliver this through three comprehensive solutions:

• Utility Bill Management: Automating the entire utility bill lifecycle — from data capture and validation to payment processing and auditing.

• Energy Procurement Advisory: Bringing together comprehensive data, AI-powered analytics, market expertise, and a strong partner network to make sophisticated procurement options accessible to all. .

• Sustainability Reporting — Verified emissions data with seamless integration into leading sustainability platforms.

Tackling the world's most complex energy challenges requires diverse thinking. We're building teams of people from different backgrounds, industries, and disciplines — united by a belief that energy management should be simple, intelligent, and a genuine driver of business value.

What we’re looking for:

We are seeking a Staff Site Reliability Engineer (L4) to join our SRE/Platform Engineering team in India. This is a senior technical leadership role — not people management, but engineering leadership through execution, mentorship, and architectural ownership.

Our India SRE team is growing, and this role is central to that growth. As we scale, we need a technical anchor in the India timezone who can independently own multi-week SRE projects from problem statement to production, make sound architectural decisions under ambiguity, and elevate the team around them. You will be the person engineers lean on for design reviews, debugging escalations, and “how should we approach this?” conversations. You’ll bring the depth and experience to drive execution autonomously in the India timezone while collaborating closely with US-based SRE leadership on roadmap priorities, incident response, and platform strategy.

This is a role for someone who doesn’t wait for direction — you identify reliability gaps, propose solutions, build consensus, and ship.

Our infrastructure is primarily AWS-based, managed by Terraform and CloudFormation, and deployed using CI/CD best practices. In your application, please include a link to GitHub or another place where your code is published, though we understand that not everyone has public code online.

What you’ll do:

• Own and deliver SRE projects end-to-end — from scoping and design through implementation, testing, rollout, and documentation

• Serve as a technical anchor for the India SRE team — conduct design reviews, pair on complex debugging, and mentor engineers to develop the judgment to work through ambiguous problems independently

• Design and implement infrastructure solutions across AWS (EKS, VPC, RDS, IAM, CloudWatch, CloudTrail, GuardDuty, S3, CloudFront, Lambda, SQS) using Terraform and CloudFormation, with an emphasis on making the right tradeoffs between speed, reliability, and cost

• Lead Kubernetes operations including cluster upgrades, capacity planning, CNI troubleshooting, workload scaling, Helm chart packaging, and GitOps deployments — and build the runbooks and automation so these become repeatable rather than one-off heroics

• Evolve CI/CD pipelines across Jenkins (Groovy scripting), GitHub Actions, AWS CodePipeline, ArgoCD, and FluxCD — with an emphasis on reducing manual deployment steps and improving rollback safety

• Drive observability stack enhancements — deliver the infrastructure and architectural direction necessary for engineering teams to leverage Prometheus, Grafana, and CloudWatch effectively

• Identify and execute FinOps initiatives — find zombie resources, right-size instances, enforce tagging standards, and present cost-reduction recommendations with data to back them up

• Manage database reliability across MySQL and PostgreSQL including backup validation, performance tuning, replication health, failover testing, and operational runbooks

• Strengthen security posture through IAM least-privilege enforcement, CSPM reviews, GuardDuty/CloudTrail monitoring, secrets management (Vault, AWS Secrets Manager, Parameter Store), and audit readiness

• Troubleshoot complex cross-cutting production issues spanning networking, Kubernetes, compute, databases, and CI/CD — and then turn the fix into a runbook or automation so the same issue doesn’t require the same person next time

• Write the documentation the team actually needs — architecture decision records, operational runbooks, troubleshooting guides, and post-incident action items that get closed, not just…

Skills asked for

Apply on the employer's site

Your next role is already in here.

Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.