JobBobsReal-time global job discoveryLive

Lead Staff Systems Reliability Engineer (Linux & Distributed Systems)

Thetradedesk · London · 2026-07-20

lead
Apply on the employer's site

About this role

The Trade Desk is a global technology company and the world’s leading independent platform for digital advertising, with nearly 4,000 employees across more than 30 offices. Our technology helps advertisers reach the right audiences across the open internet — from streaming TV and podcasts to mobile apps, news, and more.

Advertising powers the content people love. By making it more transparent, effective, and responsible, we help support trusted journalism, quality entertainment, and creators worldwide. The world’s brands and agencies rely on us to reach their customers and grow their businesses responsibly.

The scale of our platform brings unique technical challenges — from processing massive datasets in real time to building systems that operate reliably on a global scale. When you work here, your impact is worldwide. We welcome diverse perspectives, encourage curiosity, and build teams that learn from one another. If you’re driven to solve meaningful challenges, we’d love to meet you.

What we do

We are looking to hire a Lead Systems Reliability Engineer to join our engineering team to continue building and maintaining our data-driven platform. We leverage technologies like Aerospike, MongoDB, and Kafka to perform many real time activities, translating to with a p99 latency under 1 millisecond on the back end!

Do you enjoy tuning, performance testing, troubleshooting, automation, and operating at scale? Does testing next-gen hardware, evaluating data access patterns, and designing automation around distributed systems excite you?

What makes this role different:

• First in the Industry: The Trade Desk is the first company to run over 5MM QPS to NVMe in Aerospike on a single node, forcing core software redesigns to achieve this scale.

• Work on Cutting-Edge Hardware: Design clusters with nodes featuring 300TB of NVMe, 3TB RAM, and 512 cores, delivering a global 2,500GB/s throughput directly from flash.

• Shape the Future of Infrastructure: Spec your own systems and collaborate directly with AMD and NoSQL vendors to run PoCs and optimize bleeding-edge technology for internet-scale workloads.

• Deep Performance Engineering: Dive into kernel, hardware, and system interactions, leveraging tools like flamegraphs, NUMA counters, BIOS tuning, and synthetic testing to achieve world-class performance.

• Push Hardware Endurance Limits: Build clusters engineered to withstand over 1 zettabyte of endurance.

What you’ll do:

• Lead a team to influence, manage, and plan work streams, systems, and data structures at scale within a global ecosystem, spanning multiple infrastructure providers (cloud and traditional datacenters).

• Encourage, improve, and build infrastructure automation in a way that works with stateful systems at scale.

• Own operations for Linux-based systems running Aerospike, Kafka, and Mongo.

• Serve as a point of contact to review new use cases, answer questions, and participate in on-call rotation.

• Learn to be a NoSQL SME. You do not need experience to apply – we will train you.

• Benchmark and analyze next generation hardware offerings.

Who you are:

Skills and…

Skills asked for

Similar jobs

Apply on the employer's site

Your next role is already in here.

Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.