Senior Software Developer: Models Team (Token Factory)
Nebius · Amsterdam, Netherlands · 2026-07-27
Over deze functie
About Nebius:
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
About the Product
Token Factory is focused on building a next-generation platform that enables companies to seamlessly integrate AI into their products and workflows. Our vision is to create a powerful, open, and scalable alternative for deploying and managing AI systems—making advanced AI infrastructure more accessible to both fast-growing startups and large enterprises.
We work with a wide range of customers, from AI-first companies to established technology organisations, helping them run AI workloads reliably at scale. Our goal is to become a leading platform for high-performance AI inference, delivering predictable latency, strong reliability, and the ability to scale to meet demanding production needs.
Customer feedback plays a central role in how we build—our development process is highly iterative and closely aligned with real-world use cases.
About the Team
The Models Team is responsible for onboarding state-of-the-art (SOTA) open-source models into Nebius TokenFactory, including models such as DeepSeek V4 Pro, GLM 5.1, Kimi K2.6, and Minimax M2.7. A major focus of the team is serving large-scale AI models efficiently and reliably in production.
To achieve this, we work on advanced inference and systems optimization techniques, including:
• Cache-aware routing
• NUMA-aware deployments
• KV-cache offloading
• Disaggregated serving architectures
• Autoscaling with high-speed model loading over InfiniBand / RoCE
The team maintains and extends forks of leading inference frameworks such as vLLM and TRT-LLM. We have deep expertise in production-scale model serving and regularly support the Solutions Architects team on the most demanding customer PoCs.
To operate efficiently at scale, we invest heavily in tooling and automation.
Examples include:
• Performance, quality, and smoke-testing frameworks
• Hyperparameter optimization for inference framework configurations
• Gibberish detection systems
• Automated rollout pipelines for inference framework upgrades
• Diagnostics and observability tooling
• Traffic replay systems
• Automated search for optimal serverless deployment configurations
We collaborate closely with model builders, open-source communities, Nebius Cloud teams, and hardware vendors to continuously improve our serving infrastructure.
The team is highly goal-oriented and outcome-driven, with a…
Gevraagde vaardigheden
- r
- llm
- go
- python
- kubernetes
Vergelijkbare vacatures
- Senior Software Engineer: Platform SREFlexport · Amsterdam
- Senior Software Engineer - AI Compute, Together CloudTogetherai · Amsterdam
- Senior Software Engineer (Java) - Unified PlatformAdyen · Amsterdam
- Senior Software Engineer, Developer Experience (Platform Team)Wetravel · Amsterdam
- Senior Software Engineer - LLM Ops & EvalsDatasnipper · Amsterdam
- Senior Software Engineer, C++ , Core PlatformFlowtraders · Amsterdam
- Senior Scala software engineerXebia · Amsterdam
- Senior Software Engineer - ObservabilityAdyen · Amsterdam
Uw volgende baan staat er al tussen.
Doorzoek actuele vacatures van duizenden werkgevers, bewaar de interessante en laat JobBob de rest in de gaten houden.