Lead Cloud Infrastructure Engineer / Site Reliability Engineer (SRE)
Corelight · North America · 2026-06-30
About this role
<div class="content-intro"><p><strong>Be part of the team that defends the networks the world depends on</strong></p> <p>Corelight defends the world’s most sensitive networks—from global commerce to national defense—quietly, relentlessly, and with resolve. As cyber threats grow faster and smarter, we serve as the trusted force behind network resilience, putting elite defense within reach.</p> <p>By transforming digital footprints from physical, virtual, and cloud networks into actionable insights, we empower defenders to illuminate blind spots and stay ahead of an evolving threat landscape. Built on open-source innovations and fueled by industry leading agentic AI technology, Corelight helps teams to detect advanced threats and close cases with unprecedented clarity and precision.</p></div><p><span style="font-family: helvetica, arial, sans-serif;">As a<strong>&nbsp;</strong>Lead Cloud Infrastructure Engineer / Site Reliability Engineer (SRE), you will ensure the stability, performance, and security of our Federal region’s cloud platform. You’ll manage infrastructure and operations with a focus on availability, latency, performance optimization, monitoring, incident response, and capacity planning. This role requires maintaining a FedRAMP-compliant environment and working closely with teams to meet the highest standards of security and compliance.</span></p> <p><span style="font-family: helvetica, arial, sans-serif;">We adopt an "everything as code" approach, leveraging automation and best practices to create an efficient, reliable, and scalable infrastructure. You will be instrumental in maintaining core infrastructure services that are robust, secure, and capable of processing high volumes of data seamlessly.</span></p> <p><span style="font-family: helvetica, arial, sans-serif;"><strong>The successful candidate must be a U.S. citizen and may need to perform work that the U.S. government has specified can only be carried out by a U.S. citizen on U.S. soil.</strong></span></p> <p><span style="font-family: helvetica, arial, sans-serif;"><strong>Must be located in the PST time zone and able to work PST times including some off hours.</strong></span></p> <h3><span style="font-family: helvetica, arial, sans-serif;"><strong>Responsibilities</strong></span></h3> <ul> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Collaborate with software engineering teams to ensure the reliability, performance, and security of the Federal region’s infrastructure.</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Design, deploy, and scale AI/ML/LLM infrastructure across cloud platforms (AWS, Azure, or GCP) ensuring high reliability and performance.</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Manage and optimize Kubernetes environments (EKS, AKS, GKE) for AI services, data pipelines, and model operations.</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Build and automate end-to-end data and model pipelines for fine-tuning, inference, and RAG workloads using Terraform, Python, and CI/CD tooling.</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Utilize automation tools such as GitOps, CI/CD pipelines, and containerization technologies (Docker, Kubernetes) to streamline ML/LLM tasks across the Large Language Model lifecycle.</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Implement monitoring, observability, and reliability best practices using Prometheus, Grafana, ELK/EFK, Langfuse, and SLI/SLO/SLA frameworks.</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Participate in 24x7 on-call rotations, leading incident response, performance tuning, and cost optimization across SaaS Platform and production workloads</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Own infrastructure end to end, leading scaling initiatives, deployments, and automation, and providing technical leadership across the team</span></li> </ul> <h3><strong><span style="font-family: helvetica, arial, sans-serif;">Qualifications/Requirements:</span></strong></h3> <ul> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">Bachelor’s or Master’s degree in Computer Science, Engineering, or related field, or equivalent experience.</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">8+ years in SRE, DevOps, Platform Engineering, MLOps, or Cloud Infrastructure roles.</span></li> <li style="font-family: helvetica, arial, sans-serif;"><span style="font-family: helvetica, arial, sans-serif;">4+ years of production…
Skills asked for
- sre
- llm
- aws
- azure
- gcp
- kubernetes
- terraform
- python
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.