Staff ML Engineer
Buildkite · ANZ Region · 2026-07-01
About this role
<h2 class="text-text-100 mt-3 -mb-1 text-[1.125rem] font-bold">About Buildkite</h2> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">Buildkite's CI platform is trusted by the world's leading engineering teams, shipping software to over 1,000,000,000 daily users.</p> <h2 class="text-text-100 mt-3 -mb-1 text-[1.125rem] font-bold">Job Overview</h2> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">We're hiring a Staff Engineer (ML) to join our Test Engine team. In this role, you'll define and lead the technical strategy for machine learning within Test Engine — specifically, building the models and infrastructure behind predictive test selection: using code changes to determine which tests actually need to run.</p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">Staff Engineers at Buildkite are hands-on technical leaders. You'll influence how we design, build, and scale systems while supporting other engineers to deliver their best work. You'll be the most senior ML practitioner in the company, setting the technical direction for how we approach test selection and establishing the patterns and infrastructure that the broader ML effort builds on.</p> <h2 class="text-text-100 mt-3 -mb-1 text-[1.125rem] font-bold">🔧 About the Team</h2> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">The Test Engine team helps engineering teams ship faster by giving them visibility and control over their test suites. Today, that means real-time flaky test detection and management, intelligent test splitting across parallel jobs, and performance analytics and tracing — all working across any CI/CD platform, not just Buildkite Pipelines.</p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">Test Engine already ingests billions of test runs. We have deep visibility into test suites, codebases, and the relationships between them. The next step is using that data to answer a fundamental question: for a given code change, which tests are most likely to fail?</p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">We believe the industry is moving away from running full test suites on every change. The teams that can shift their outer testing loop into a fast, precise inner loop — running only the tests that matter — will ship value to their customers dramatically faster. For many of our customers, that speed is existential. Switching costs are low, competition is fierce, and the teams with faster feedback loops win.</p> <p class="font-claude-response-body break-words whitespace-normal leading-[1.7]">This is where ML comes in. If we can model the relationship between code changes and test failures, we can give engineering teams a fundamentally faster development cycle. We're not trying to optimise individual tests — we're trying to build a generalised solution to test selection that works across codebases, frameworks, and languages.</p> <h2 class="text-text-100 mt-3 -mb-1 text-[1.125rem] font-bold">🚀 What You'll Do</h2> <h3 class="text-text-100 mt-2 -mb-1 text-base font-bold">Own Technical Direction for ML in Test Engine</h3> <ul class="[li_&amp;]:mb-0 [li_&amp;]:mt-1 [li_&amp;]:gap-1 [&amp;:not(:last-child)_ul]:pb-1 [&amp;:not(:last-child)_ol]:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3"> <li class="whitespace-normal break-words pl-2">Lead and define the ML strategy for predictive test selection — from early experimentation through to models running reliably in production at scale</li> <li class="whitespace-normal break-words pl-2">Lead the technical investigation into how we build a generalised test selection model, and shape the approach based on what the data tells you</li> <li class="whitespace-normal break-words pl-2">Lead the design of the ML architecture end-to-end: feature engineering from code changes and test history, model training and evaluation, serving infrastructure, and feedback loops for continuous improvement</li> <li class="whitespace-normal break-words pl-2">Drive key decisions around model operationalisation — latency constraints (test selection has to be fast enough to sit in the critical path), prediction accuracy trade-offs, and graceful degradation when confidence is low</li> <li class="whitespace-normal break-words pl-2">Shape how ML capabilities integrate with Test Engine's existing data infrastructure — billions of ingested test runs, test-to-code mapping, and the intelligent splitting engine</li> </ul> <h3 class="text-text-100 mt-2 -mb-1 text-base font-bold">Build and Scale the ML Platform</h3> <ul class="[li_&amp;]:mb-0 [li_&amp;]:mt-1 [li_&amp;]:gap-1 [&amp;:not(:last-child)_ul]:pb-1 [&amp;:not(:last-child)_ol]:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3"> <li class="whitespace-normal break-words pl-2">Build the ML platform layer so that getting a model into production is fast and repeatable</li> <li class="whitespace-normal break-words pl-2">Design, build, and maintain the data pipelines that feed ML workloads — connecting code change signals with test execution history at scale</li> <li class="whitespace-normal break-words pl-2">Train, evaluate, and deploy models, taking ownership through to monitoring and retraining in production</li> <li class="whitespace-normal break-words pl-2">Instrument production models…
Skills asked for
- machine learning
- ci/cd
- python
- spark
- aws
- docker
- kubernetes
- ruby
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.