Senior Software Developer: Models Team (Token Factory)
Nebius · Amsterdam, Netherlands · 2026-07-27
Over deze functie
<div class="content-intro"><p><strong>About Nebius:</strong></p> <p>Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.</p> <p>Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.</p> <p>Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&amp;D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&amp;D.</p></div><p><strong><span data-contrast="none"><span data-ccp-parastyle="heading 3">About the Product</span></span></strong><span data-ccp-props="{&quot;134233117&quot;:false,&quot;134233118&quot;:false,&quot;134245418&quot;:true,&quot;134245529&quot;:true,&quot;335559738&quot;:281,&quot;335559739&quot;:281}">&nbsp;</span></p> <p><span data-contrast="auto">Token Factory is focused on building a next-generation platform that enables companies to seamlessly integrate AI into their products and workflows. Our vision is to create a powerful, open, and scalable alternative for deploying and managing AI systems—making advanced AI infrastructure more accessible to both fast-growing startups and large enterprises.</span><span data-ccp-props="{&quot;134233117&quot;:false,&quot;134233118&quot;:false,&quot;335559738&quot;:240,&quot;335559739&quot;:240}">&nbsp;</span></p> <p><span data-contrast="auto">We work with a wide range of customers, from AI-first companies to established technology organisations, helping them run AI workloads reliably at scale. Our goal is to become a leading platform for high-performance AI inference, delivering predictable latency, strong reliability, and the ability to scale to meet demanding production needs.</span><span data-ccp-props="{&quot;134233117&quot;:false,&quot;134233118&quot;:false,&quot;335559738&quot;:240,&quot;335559739&quot;:240}">&nbsp;</span></p> <p><span data-contrast="auto">Customer feedback plays&nbsp;a central role&nbsp;in how we build—our development process is highly iterative and closely aligned with real-world use cases.</span><span data-ccp-props="{&quot;134233117&quot;:false,&quot;134233118&quot;:false,&quot;335559738&quot;:240,&quot;335559739&quot;:240}">&nbsp;</span></p> <h3><span style="font-size: 12pt;">About the Team</span></h3> <p><span data-contrast="auto"><span data-ccp-charstyle="Strong">The Models Team is responsible for onboarding state-of-the-art (SOTA) open-source models into Nebius TokenFactory, including models such as DeepSeek V4 Pro, GLM 5.1, Kimi K2.6, and Minimax M2.7. </span></span><span data-contrast="auto"><span data-ccp-charstyle="Strong">A major focus of the team is serving large-scale AI models efficiently and reliably in production. </span></span></p> <p><strong><span data-contrast="auto"><span data-ccp-charstyle="Strong">To achieve this, we work on advanced inference and systems optimization techniques, including:<br></span></span></strong></p> <ul> <li>Cache-aware routing</li> <li>NUMA-aware deployments</li> <li>KV-cache offloading</li> <li>Disaggregated serving architectures</li> <li>Autoscaling with high-speed model loading over InfiniBand / RoCE</li> </ul> <p><span data-contrast="auto"><span data-ccp-charstyle="Strong">The team maintains and extends forks of leading inference frameworks such as <strong>vLLM</strong> and <strong>TRT-LLM</strong>. We have deep expertise in production-scale model serving and regularly support the Solutions Architects team on the most demanding customer PoCs.</span></span></p> <p><span data-contrast="auto"><span data-ccp-charstyle="Strong">To operate efficiently at scale, we invest heavily in tooling and automation.&nbsp;</span></span></p> <p><strong><span data-contrast="auto"><span data-ccp-charstyle="Strong">Examples include:<br></span></span></strong></p> <ul> <li>Performance, quality, and smoke-testing frameworks</li> <li>Hyperparameter optimization for inference framework configurations</li> <li>Gibberish detection systems</li> <li>Automated rollout pipelines for inference framework upgrades</li> <li>Diagnostics and observability tooling</li> <li>Traffic replay systems</li> <li>Automated search for optimal serverless deployment configurations</li> </ul> <p><span data-contrast="auto"><span data-ccp-charstyle="Strong">We collaborate closely with model builders, open-source communities, Nebius Cloud teams, and hardware vendors to continuously improve our serving infrastructure.<br>The team is highly goal-oriented and outcome-driven, with a…
Gevraagde vaardigheden
- r
- llm
- go
- python
- kubernetes
Vergelijkbare vacatures
- Senior Software Engineer Document IntelligenceDatasnipper · Amsterdam
- Senior Software Engineer (Financial Statement Suite)Datasnipper · Amsterdam
- Staff / Senior Software Engineer (Agentic Search) - RuntimeNebius · Amsterdam
- Staff / Senior Software Engineer (Agentic Search) - IndexNebius · Amsterdam
- Staff / Senior Software Engineer (Agentic Search) - CrawlerNebius · Amsterdam
- Senior Software Engineer (Serverless)Nebius · Amsterdam
- Senior Software Engineer (Managed PostgreSQL)Nebius · Amsterdam
- Senior Software Engineer (Token Factory)Nebius · Amsterdam
Uw volgende baan staat er al tussen.
Doorzoek actuele vacatures van duizenden werkgevers, bewaar de interessante en laat JobBob de rest in de gaten houden.