JobBobsReal-time global job discoveryLive

Senior Software Engineer, Server Fleet Infrastructure

Coreweave · Livingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA · 2026-07-27

executive
Apply on the employer's site

About this role

<div class="content-intro"><div> <div> <div class="gmail_quote"> <div> <div><span id="m_1770241969069985273m_-2746164444908759431gmail-docs-internal-guid-131e4fb0-7fff-b4e9-ff50-e8cf32449b1b">CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at <a href="http://www.coreweave.com/" target="_blank" data-saferedirecturl="https://www.google.com/url?q=http://www.coreweave.com&source=gmail&ust=1762613132717000&usg=AOvVaw3D-UOhNaqEvF5BEWxjYyAU">www.coreweave.com</a>.</span></div> </div> </div> </div> </div></div><p><strong>What You’ll Do: </strong></p> <p>The Fleet Provisioning Automation (FPA) team is responsible for the automated provisioning and lifecycle management of CoreWeave’s rapidly growing fleet of hardware nodes and node types. The team streamlines and coordinates node bring-up, hardware RMA, data center operations, and platform services into a cohesive, high-reliability engine of fleet management. This group sits at the intersection of hardware, data center operations, and platform engineering, building the software that keeps our global fleet healthy and ready for customer workloads.</p> <p><strong>About the Role: </strong></p> <p>As a Senior Software Engineer on the Fleet Provisioning Automation team, you will design and build backend services and APIs that automate provisioning, configuration, and lifecycle operations for CoreWeave’s globally distributed bare metal fleet. You will primarily work in Go to implement gRPC APIs that integrate with Kubernetes, vendor and internal services, and data center tooling. Your work will focus on turning complex, multi-step operational procedures into simple, safe, and auditable automation for internal users operating at hyperscale.</p> <p>In this role, you will:</p> <ul> <li>Design and implement backend services and APIs (primarily gRPC in Go) that orchestrate provisioning, configuration, and lifecycle operations across CoreWeave’s global server fleet.</li> <li>Develop Kubernetes custom resource definitions (CRDs) to automate provisioning and lifecycle management of CoreWeave’s entire server fleet.</li> <li>Model and evolve RPC schemas and data contracts used by other engineering and operations teams to integrate with fleet provisioning workflows.</li> <li>Build integrations with vendor and internal APIs to make hardware and data center processes robust, transparent, and easy to operate at scale.</li> <li>Collaborate closely with hardware engineering, data center operations, and platform teams to design solutions to problems of scale for multi-site deployment and management of CoreWeave’s global hardware fleet.</li> <li>Create test plans, deployment automation, and supporting tooling that enable safe rollouts and continuous improvement of fleet provisioning services.</li> <li>Participate in the Fleet Provisioning Automation on-call rotation and help drive incident response and post-incident improvements.</li> </ul> <p>Who You Are:</p> <ul> <li>5+ years of experience in software or infrastructure engineering building and operating production backend or infrastructure services.</li> <li>Proficiency in Go for building networked services and APIs (gRPC and REST) in production environments.</li> <li>Experience designing and implementing distributed systems that operate reliably at scale, including concurrency, failure handling, and resiliency patterns.</li> <li>Strong understanding of Linux systems and how to debug issues across processes, networking, and storage.</li> <li>Experience with Kubernetes or similar container orchestration platforms and their APIs (for example, interacting with custom resources, controllers, or operators).</li> <li>Familiarity with CI/CD tooling (such as Argo, Flux, or GitHub Actions) to ship and operate services safely and frequently.</li> <li>Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.</li> </ul> <p>Preferred:</p> <ul> <li>Experience designing and operating services that automate lifecycle management for large fleets of physical servers or other hardware.</li> <li>Experience with infrastructure automation and configuration management tools (for example, Ansible, Puppet, Chef, or Salt).</li> <li>Experience integrating with vendor APIs and internal systems to coordinate multi-step operational workflows.</li> <li>Experience with multi-datacenter or regionally distributed systems, including thinking about failure domains, capacity, and rollout strategies.</li> </ul> <p>Wondering if you’re a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match. Here are a few qualities we’ve found compatible with our team. If some of this describes you, we’d love to talk.</p> <ul> <li>You enjoy building backend APIs and distributed services that automate real-world infrastructure operations at…

Skills asked for

Similar jobs

Apply on the employer's site

Your next role is already in here.

Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.