Senior Staff / Principal Machine Learning Scientist, AI Inference & Optimization
Netskope · Santa Clara, California, United States · 2026-07-28
About this role
Join the Future of Security at Netskope
Netskope (NASDAQ: NTSK) is a leader in modern security and networking for the cloud and AI era. We secure and accelerate cloud, data, and AI in real time, everywhere. Thousands of customers, including more than 30 of the Fortune 100, trust the Netskope One platform, its Zero Trust Engine, and the powerful NewEdge network to gain full visibility and control without performance trade-offs.
At Netskope, our technology is driven by our greatest strength: our people. We believe that belonging powers innovation, and success is both personal and organizational. We embrace differences in gender, ethnicity, beliefs, ability, and identity, creating an environment where every voice is heard and respected. We empower our employees to bring their authentic selves to work, grow their careers through continuous education and mentorship, and lead with transparency and curiosity. Join a team where you belong, where you are encouraged to be an entrepreneur, and where together, we continue to redefine the landscape of security.
Visit Careers at Netskope to learn more. Follow us on LinkedIn and Instagram.
Positions are available at Senior Staff and above. Candidates are assessed individually and leveled according to their specific skills and background.
About the role
As a Senior Staff Machine Learning Scientist, you own the inference and optimization layer that makes AI in agentic workflows fast, efficient, and production-grade. You fine-tune and evaluate models, push latency and throughput on real hardware, and build the runtime that executes bounded AI tasks, validated against usage from Netskope’s large customer base so you optimize where the data points, not where you guess.
What’s in it for you
• High-impact ownership. You own the model layer of a net-new product that changes the performance and economics of agentic AI.
• Cutting-edge, unusual stack. The hard, interesting inference problems live here: quantization, KV-cache and memory management, sparsity, fine-tuning, and hardware acceleration under real-world resource constraints.
• Real scale to build against. Netskope’s customer footprint gives you production signals most teams never see, so you deploy, validate, and iterate fast.
What you will be doing
• Build and optimize the model inference path: quantization, KV-cache optimization, batching, and latency/memory/throughput tuning on constrained, commodity hardware.
• Fine-tune and evaluate models for bounded tasks; build eval harnesses that gate a capability to release on real accuracy, latency, and security relevance.
• Design and grow the task execution runtime (bounded sub-agents), pushing toward dynamic task generation and context compaction.
• Drive hardware acceleration / sparsity and support for larger models as the platform matures.
• Partner with the systems and backend engineers to ship capabilities end-to-end and iterate on real production signals.
Required skills and experience
• 10+ years of overall industry experience, with 4+ years hands-on in ML/AI (model development, fine-tuning, and inference optimization).
• Hands-on with fine-tuning (e.g. LoRA/QLoRA), quantization (GGUF/AWQ/GPTQ), and inference runtimes (vLLM/SGLang, TensorRT-LLM, ONNX Runtime, llama.cpp, or MLX/CoreML). On-device or edge inference experience is a strong plus.
• Strong Python; comfort reaching into C++ for low-level interop is a plus.
• Solid grasp of transformer internals and the levers that move real inference performance and cost: KV cache, attention, batching, memory footprint.
• Fluency with agentic coding systems and genuine curiosity about agent harnesses like Claude Code, Pi, and Codex, so you should already be building with them, or itching to.
• Clear communication: able to distill a model or infra bottleneck into an actionable concept for cross-functional teammates.
Education
• MS in Computer Science, Machine Learning, Electrical Engineering, or equivalent technical degree required, with a focus in AI/ML research; PhD in a related field strongly preferred.
Compensation:
At Netskope, salary is one component of our competitive total rewards package. The salary range for this position is as listed below. This is a national range. For purposes of complying with applicable laws, the range applies to…
Skills asked for
- machine learning
- llm
- python
- c++
Similar jobs
- Senior/Staff Software Engineer - Remote AssistanceGatikaiinc · Santa Clara
- Senior Staff LLM Inference EngineerD Matrix · Santa Clara
- Senior Staff Product Manager, AI PlatformServiceNow · Santa Clara
- Senior Staff Technical Program ManagerServiceNow · Santa Clara
- Senior Staff Software Engineer, Developer and Qualification ToolsD Matrix · Santa Clara
- Senior Staff Database EngineerServiceNow · Santa Clara
- Senior Staff Operations Program ManagerD Matrix · Santa Clara
- Senior Staff Power Performance Architect Accelerator DesignD Matrix · Santa Clara
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.