Senior Site Reliability Engineer, AI Research
Algolia · Remote - Australia · 2026-06-01
About this role
<div class="content-intro"><p>At Algolia, we’re proud to be a pioneer and market leader in AI Search, empowering 17,000+ businesses to deliver blazing-fast, predictive search and browse experiences at internet scale. Every week, we power over 30 billion search requests — four times more than Microsoft Bing, Yahoo, Baidu, Yandex, and DuckDuckGo combined.</p> <p>In 2021, we raised $150 million in Series D funding, quadrupling our valuation to $2.25 billion. This strong foundation enables us to keep investing in our market-leading platform and serving incredible customers like Under Armour, PetSmart, Stripe, Gymshark, and Walgreens.</p></div><h4><strong>About the AI Research Team</strong></h4> <p>The AI Research team at Algolia combines <strong>fundamental research</strong> with <strong>product engineering</strong> to deliver customer-facing AI-powered features.</p> <p>The team is highly cross-functional, made up of PhD researchers, full-stack engineers, and infrastructure specialists working together to explore new ideas, validate impact, and bring successful research outcomes into production. While the work is research-driven, the output is real, customer-facing systems.</p> <h4><strong>The Opportunity</strong></h4> <p>We are looking for an <strong>embedded Senior Site Reliability Engineer</strong> to join the AI Research team as a full member of the group. In this role, you will support both the research and product-engineering aspects of the team by ensuring the stability, scalability, and operability of the infrastructure that enables this work.</p> <p>This is a <strong>classic SRE role</strong> focused on <strong>cloud-first, service-oriented architectures</strong> running on Google Cloud Platform. While the team builds AI-powered systems, <strong>AI or ML experience is not required</strong> for this role. Our priority is strong SRE fundamentals, experience operating production services, and comfort working in an environment with ambiguity and high ownership.</p> <p>You will play an important role in <strong>day-to-day execution</strong> as well as in <strong>longer-term (12-month) planning</strong>, helping shape how the team builds and operates its platforms over time.</p> <h4><strong>What You’ll Work On</strong></h4> <h4><strong>Platform Reliability &amp; Enablement</strong></h4> <ul> <li>Support and evolve the reliability of platforms used by the AI Research team. Examples of our infrastructure work to date include:</li> <ul> <li>A production inference service (embedding model serving API)</li> <li>AI data feature store</li> <li>Internal tools used for novel research and experimentation</li> <li>Infrastructure that combines the above to enable offline testing of customer deployments to agentically discover configuration improvements.<br><br></li> </ul> <li>Ensure production services meet expectations for availability, latency, and operational readiness, particularly for systems that sit on customer-critical paths<br><br></li> <li>Design infrastructure and operational patterns that prioritize <strong>iteration speed</strong> while maintaining appropriate safeguards for production systems<br><br></li> </ul> <h4><strong>Embedded Collaboration</strong></h4> <ul> <li>Work closely with researchers and engineers in a cross-functional setting, acting as an advisor on infrastructure, reliability, and operational concerns<br><br></li> <li>Participate directly in team planning and execution, from early exploration through production rollout<br><br></li> <li>Help researchers self-serve infrastructure safely and effectively, without becoming a bottleneck<br><br></li> </ul> <h4><strong>Cloud Infrastructure &amp; Operations</strong></h4> <ul> <li>Build and maintain Kubernetes-based services on GCP using infrastructure-as-code and GitOps (Terraform, ArgoCD)<br><br></li> <li>Own and improve CI/CD pipelines for services written primarily in Go, with some Python-based services<br><br></li> <li>Design and operate observability systems using tools such as Datadog<br><br></li> <li>Participate in an on-call rotation (relatively light), responding to incidents and helping improve systems over time<br><br></li> </ul> <h4><strong>What We’re Looking For</strong></h4> <h4><strong>Required Experience</strong></h4> <ul> <li>Strong experience operating <strong>cloud-first infrastructure</strong><br><br></li> <li>Hands-on experience running production services on Kubernetes<br><br></li> <li>Proficiency with infrastructure-as-code (Terraform) and CI/CD systems<br><br></li> <li>Experience supporting production services written in <strong>Go</strong> (Python experience is a plus)<br><br></li> <li>Solid grounding in service reliability, incident response, and operational best practices<br><br></li> <li>Comfort working in environments with ambiguity, where problems are not always well-defined upfront<br><br></li> </ul> <h4><strong>Nice to Have</strong></h4> <ul> <li>Experience supporting <strong>mission-critical internal…
Skills asked for
- sre
- google cloud
- kubernetes
- gcp
- terraform
- argocd
- ci/cd
- go
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.