Senior Site Reliability Engineer (AUS)
Climavision · Australia · 2026-10-02
About this role
Senior Site Reliability Engineer
Remote | Australia
About Climavision
At Climavision, we’re rebuilding climate technology from the ground up and changing the way we see weather. We merge the power of our proprietary, high-resolution weather radar and satellite network with advanced weather prediction modelling and decades of industry expertise to reduce existing coverage gaps and drastically improve forecasting ability. Our revolutionary new approach to climate technology weather solutions is poised to help reduce the economic risks of climate change on companies, governments, and societies alike. We are backed by The Rise Fund, the world’s largest global impact platform committed to achieving measurable, positive social and environmental outcomes alongside competitive financial returns. Climavision is headquartered in Louisville, KY, with research and development operations in Raleigh, NC.
The Work
Are you an experienced Site Reliability Engineer who thrives at the intersection of software engineering and production operations? Do you take pride in keeping mission-critical customer systems reliable under real-world operational pressure? Are you looking for an opportunity to own production reliability for a modern hybrid infrastructure platform spanning cloud, colocation, and edge environments?
If so, we have an exceptional opportunity for you.
Climavision is seeking a Senior Site Reliability Engineer to contribute towards reliability, operational excellence, and production resilience across the company's platform and data services. This role sits on a shared SRE team that supports the full business rather than a single product line, covering both the radar network and the weather intelligence sides of the company as priorities shift. A central focus of this role is building the observability layer that puts the health of the full fleet in one place, and then automating recovery so that systems heal themselves. Multi-cluster and multi-replica high availability across our distributed edge fleet remains a core part of the work.
This is a hands-on engineering role for someone who is equally comfortable troubleshooting Kubernetes clusters, leading incident response, and improving operational maturity across the organization. The successful candidate will combine deep production operations expertise with a disciplined approach to reliability engineering and strong automation skills.
Climavision operates a hybrid infrastructure footprint spanning Microsoft Azure, colocation data centers, and edge Kubernetes clusters, deployed alongside weather radar systems. This role will drive production reliability across Azure, colocation, and edge environments. Right-sizing cluster resources and migrating workloads off Azure to reduce spend are active priorities for the team.
35% Kubernetes Platform Reliability and Operations
30% Production Reliability Engineering and Incident Response
20% Observability, Monitoring, and Alerting
15% Automation, Recovery, and Cost Optimization
Primary Responsibilities:
• Own production reliability for Climavision's customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments.
• Work as part of a shared SRE function supporting the whole company rather than a single product line, taking on work across both the radar network and the weather intelligence sides of the business as priorities shift.
• Contribute to the definition and improvement of SLIs, SLOs, alerting standards, and operational metrics used to measure platform reliability.
• Build and own the observability layer for the fleet. Today the underlying metrics exist but are only reachable from the command line inside each cluster. This role is responsible for surfacing that data in shared dashboards and building the alerting that tells the team something is going wrong before a customer does.
• Design and build automated recovery and self-healing for production systems, so that common failure modes are detected and remediated without human intervention.
• Optimize cluster resourcing and cost, including right-sizing workloads and nodes and supporting the migration of workloads off Azure to reduce spend.
• Support and coordinate production incident response efforts, including troubleshooting, mitigation, communication, and postmortem analysis.
• Diagnose and resolve complex production issues across application services, Kubernetes infrastructure, storage, and distributed systems.
• Drive multi-replica and multi-cluster high availability across Climavision's services, including workload placement, scheduling, and deployment patterns that allow services to run safely as multiple replicas across multiple clusters.
• Contribute to the multi-cluster high-availability strategy across Climavision's hybrid fleet, including active-active and active-passive failover behavior, traffic routing, data replication considerations, and graceful degradation when a cluster becomes unavailable.
• Operate and improve Climavision's self-managed Kubernetes platform spanning cloud-hosted, colocation, and edge clusters, with a focus on availability, resiliency, recovery, and operational performance.
• Ensure Kubernetes platform lifecycle activities including upgrades, patching, cluster health, node management, and production change management are executed in a manner that preserves service availability and minimizes customer-facing risk.
• Improve reliability and operational maturity of production platform services, including observability, autoscaling, ingress, and distributed storage. Partner with the teams responsible for the underlying networking and security primitives rather than owning those areas directly.
• Design and validate Kubernetes workloads for resiliency, scalability, and operational efficiency, including autoscaling behavior, workload placement, resource management, and graceful degradation strategies.
• Partner with software engineering teams across the…
Skills asked for
- sre
- kubernetes
- azure
- helm
- devops
- terraform
- ansible
- ci/cd
Similar jobs
- Senior Site Reliability EngineerPriority
- Senior Data Center Technician (On-site)Trace3 · Las Vegas
- Senior Site Reliability EngineerPlatformsh · Remote • Australia
- Senior Site Reliability EngineerSentinellabs · Czech Republic
- Senior Site Reliability EngineerSentinellabs · Brno
- Senior Site Reliability Engineer (Agentic Search)Nebius · Israel
- Senior Project Services Project Manager (On-site / Pampa,TX)Mullins · Abernathy
- Senior Project Services Project Manager (On-site / Pampa, TX)Mullins · Dallas-Fort Worth
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.