Senior Systems Engineer, OS Automation
Coreweave · Livingston, NJ / New York City, NY/ Sunnyvale, CA/ Bellevue, WA · 2026-07-27
About this role
<div class="content-intro"><div> <div> <div class="gmail_quote"> <div> <div><span id="m_1770241969069985273m_-2746164444908759431gmail-docs-internal-guid-131e4fb0-7fff-b4e9-ff50-e8cf32449b1b">CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at&nbsp;<a href="http://www.coreweave.com/" target="_blank" data-saferedirecturl="https://www.google.com/url?q=http://www.coreweave.com&amp;source=gmail&amp;ust=1762613132717000&amp;usg=AOvVaw3D-UOhNaqEvF5BEWxjYyAU">www.coreweave.com</a>.</span></div> </div> </div> </div> </div></div><p><strong>What You'll Do:</strong></p> <p>&nbsp;</p> <p>HAVOCK owns the host software stack that turns a freshly provisioned bare-metal machine into a healthy CoreWeave Kubernetes worker — everything from power-on to a node joining the cluster: the OS image, boot-time configuration, and the services that decide which combination of software is valid for a given piece of hardware. As that problem space keeps growing, the only way to stay agile is to write software that manages the complexity, checks our work, and validates our assumptions — that's what this team builds.</p> <p>&nbsp;</p> <p><strong>About the role:</strong></p> <p>&nbsp;</p> <p>As a Senior Software Engineer on the Automation team, you will design, build, and operate the services, APIs, and libraries that sit behind our OS image, payload, and boot-configuration systems — the software platform other HAVOCK engineers and partner teams rely on to release, test, and ship node software quickly and safely. You'll work on a constraint-solver–based service that resolves compatibility between images, kernels, drivers, payloads, and hardware into a single validated configuration; an end-to-end test framework that validates OS images on real hardware; a library suite for declaratively configuring node storage; and natural-language tooling that lets stakeholders query and interact with our systems. This is a software- and platform-engineering role first, with a clear forward trajectory toward AI-assisted automation — log triage, regression detection, natural-language interfaces to infrastructure — but the core of the job is designing and shipping reliable services and APIs.</p> <p>&nbsp;</p> <p><strong>Some of what you'll work on:</strong></p> <p>&nbsp;</p> <ul> <li>Own and evolve a boot-configuration service that models complex compatibility and dependency relationships between OS images, kernels, drivers, payloads, and instance types as a constraint-solved graph, exposed through clean, well-specified interfaces.</li> <li>Extend our Kubernetes-native, end-to-end test framework that validates OS images and configuration on real hardware, plus the broader testing and validation story for the team.</li> <li>Build and maintain a library suite for declaratively configuring node storage — filesystems, mount options, block-device selection — with configurable strictness.</li> <li>Ship changes to our versioned, boot-time payload system (networking, storage, Kubernetes join) that's published as artifacts and consumed during node bring-up.</li> <li>Grow our natural-language / chat interface that lets stakeholders query and interact with the team's systems.</li> <li>Design and evolve versioned service contracts (gRPC / Connect-RPC, Protobuf) with strong correctness guarantees and robust validation.</li> <li>Build tooling that meaningfully shortens the build-and-release loop</li> <li>Lower the barrier to entry for everyone who touches this software, and lay the groundwork — clean interfaces, structured build/test metadata — for future AI-assisted automation across build triage, regression detection, and natural-language infrastructure tooling.</li> <li>Operate the services you build: participate in an on-call rotation for the team's services and own their reliability.</li> </ul> <p>&nbsp;</p> <p><strong>Who You Are:</strong></p> <p>&nbsp;</p> <ul> <li>3+ years of professional software engineering experience building and operating backend services, platforms, or developer/infrastructure tooling.</li> <li>Strong proficiency in <strong>Go</strong> and/or <strong>Python</strong>, with the ability to work fluently across both.</li> <li>Experience designing and maintaining APIs and service contracts (REST, gRPC, or similar), with an eye for clean, well-specified, versioned interfaces.</li> <li>A demonstrated instinct for data modeling — representing relationships, constraints, and dependencies in code (graphs, constraint solving, relational models, or similar).</li> <li>Solid testing discipline: you write services that are testable, and you build the automation that proves they work.</li> <li>Comfort operating in a Kubernetes-based environment and reasoning about how software is built, packaged, deployed, and released.</li> <li>A working understanding of how Linux systems boot and are configured (the OS image / cloud-init / provisioning…
Skills asked for
- kubernetes
- agile
- grpc
- go
- python
- rest
- linux
- rust
Similar jobs
- Senior Systems Engineer, VirtualizationCoreweave · Livingston
- Senior Systems Engineer, Test Frameworks & Validation PlatformCoreweave · Livingston
- Senior Systems Engineer, CKS PerformanceCoreweave · Livingston
- Senior Salesforce Administrator (GTM Systems)Coreweave · Livingston
- Senior Business Systems AdministratorCoreweave · Livingston
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.