Senior Technical Lead, AI Benchmarking and Evaluation Research (Remote, ROU)
CrowdStrike · Remote · 2026-10-07
About this role
As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn't changed - we're here to stop breaches, and we've redefined modern security with the world's most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We're always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you.
About the Role
The CrowdStrike OCAIO team is looking for a senior technical lead to own how we measure AI models and agentic systems that perform cybersecurity tasks. This is a hands-on leadership role. You will lead a small team of researchers and you will personally help design, build and defend the benchmarks they ship.
Our mission is to establish rigorous, reproducible standards for how well AI and agentic systems perform real security work: malware analysis, reverse engineering, threat intelligence reasoning, alert triage, investigation and response. We build our ground truth with security subject matter experts, not with other models, and we score agents on what they actually did, not only on what they said.
As the Senior Lead, you define what "good" looks like for an AI system operating in a security workflow, you build the datasets, rubrics and judges that encode that definition, and you run the standing scorecards that engineering, product and research teams use to decide what ships. You set the technical bar and you meet it yourself.
What You'll Do
Hands-on benchmark work (about half of your time):
- Design and build benchmark datasets for cybersecurity agent tasks, from task definition through expert-labeled ground truth, scoring rubrics and release.
- Build and calibrate judges (LLM-based and programmatic) against expert labels, measure agreement and publish where they disagree.
- Build reproducible evaluation harnesses that capture agent actions and traces, not only transcripts, so results are defensible and comparable across model versions.
- Run and publish standing scorecards of CrowdStrike agents and frontier models on the same tasks, with protection, correctness and usefulness reported together.
- Investigate failure modes personally: where agents collapse, why, and what evidence proves it.
- Write the research: internal reports and external papers or talks that establish the benchmark's credibility.
Leading the team and the function:
- Lead, mentor and grow a team of researchers and engineers; review their designs, labels and code.
- Set the strategy, roadmap and success metrics for evaluating AI and agentic systems across security use cases.
- Define and enforce the evaluation methodology and the shared task and trace formats used across benchmarking and red teaming.
- Secure and coordinate subject matter expert labeling time from SOC, threat research and malware analysis teams, and set the quality bar for every label.
- Collaborate with engineering, product and threat research to turn evaluation findings into agreed pass/fail criteria and shipped improvements.
- Communicate results, trade-offs and limitations clearly to technical and executive audiences. Report what was not tested as plainly as what was.
What You'll Need
- Hands-on cybersecurity expertise in at least one of: malware analysis and reverse engineering, incident response and threat hunting, detection engineering, SOC operations. You can label ground truth yourself and judge whether an expert's label is right.
- Demonstrated experience building evaluations, benchmarks or datasets for AI or LLM systems, with shareable artifacts (papers, repositories, internal benchmarks with measured adoption).
- Strong Python and data tooling skills. You write production-quality evaluation code, harnesses and analysis, and you review others' code.
- Practical experience with LLMs and agentic systems: tool use, multi-step planning, sandboxing, and how to instrument and trace agent behavior.
- At least 2 years leading technical people (as a team lead, tech lead or manager), with a record of growing researchers and engineers while remaining hands-on.
- Ability to define, design and standardize evaluation methodologies and reproducible testing pipelines, and to defend them under scrutiny.
- Broad knowledge of the cybersecurity landscape: attack techniques, defensive controls, and the analyst workflows that AI is meant to augment.
- Exceptional written and spoken communication. You can present measured findings and their limitations to executives and to researchers.
- Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.
Bonus Points
- Published research or talks on AI evaluation, LLM judging, or AI security.
- Experience with adversarial testing or red teaming of AI systems, and with scoring agents from observed side effects rather than text.
- Experience building LLM judges and measuring judge agreement with human experts.
- Experience with SOC platforms, SIEM/SOAR tooling and detection engineering.
- Understanding of MITRE…
Skills asked for
- cybersecurity
- llm
- python
Similar jobs
- Senior Technical Architect / Technical Architect Director - Data Foundation DeliverySalesforce · Flexible / Remote
- Senior Technical ConsultantSalesforce · Flexible / Remote
- Senior Professional Services Technical Architect - SecurityGitLab · Flexible / Remote
- Senior Technical ArchitectSalesforce · Flexible / Remote
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.