JobBobsReal-time global job discoveryLive

Senior Machine Learning Engineer, LLM Inference Optimization

Nebius · Palo Alto, California, United States · 2026-07-30

executive
Apply on the employer's site

About this role

About Nebius:

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

The role

Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering.

A Senior MLE owns substantial model and endpoint optimization projects end to end. They are deeply hands-on, can debug difficult serving problems independently, and can deliver measurable improvements without needing heavy supervision.

Your responsibilities:

•
Own optimization work for specific model families, customer endpoints, or serving backends.

•
Run engine comparisons and recommend practical serving configurations for specific workloads.

•
Debug model quality or performance regressions during production rollouts.

•
Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.

•
Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, or similar systems.

•
Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.

•
Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.

•
Build reproducible benchmark harnesses for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token.

•
Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers.

•
Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations.

Must-haves:

Skills asked for

Similar jobs

Apply on the employer's site

Your next role is already in here.

Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.