Senior AI Engineer
Dvtrading · Chicago; Hong Kong; London; New York; Singapore · 2026-07-20
About this role
<p data-renderer-start-pos="1" data-local-id="5d6d5c6caa36"><strong data-renderer-mark="true">About Us</strong>:</p> <p data-renderer-start-pos="1" data-local-id="5d6d5c6caa36">Founded 20 years ago and headquartered in Chicago, the&nbsp;<span data-highlighted="true" data-vc="highlighted-text"><span class="_kqswh2mm"><span class="_5pioz8co _189e1dm9 _1il9buyh _19lc184f _d0altlke" data-testid="definition-highlighter">DV</span></span></span> Group of financial services firms has grown to more than 600 people operating throughout North America, Europe and Asia. Since spinning out of a large brokerage firm in 2016, <span data-highlighted="true" data-vc="highlighted-text">DV</span> Trading has rapidly scaled as an independent proprietary trading firm utilizing its own capital, trading strategies, and risk management methodologies to provide liquidity to worldwide financial markets and hedging opportunities to commodity producers and users. Now, <span data-highlighted="true" data-vc="highlighted-text">DV</span> group affiliates include two broker dealers, a cryptocurrency market making firm, and a bourgeoning investment adviser.</p> <p data-renderer-start-pos="641" data-local-id="3db7b8c101de"><strong data-renderer-mark="true">Overview: </strong></p> <p data-renderer-start-pos="641" data-local-id="3db7b8c101de">DV Trading is building a centralized AI function and is now hiring for the model layer. The long-term goal is&nbsp;for DV to own its model capability — not to be permanently dependent on what frontier providers choose to&nbsp;offer, at what price, for how long. This role is how that happens: fine-tuning and distilling open-weight&nbsp;models for DV-specific tasks, operating the inference infrastructure to run them on-prem, and building the<br>model gateway that routes intelligently across open and closed providers. The near-term result is lower&nbsp;cost and better latency. The long-term result is a firm that controls its own AI stack.</p> <p data-renderer-start-pos="706" data-local-id="6d17a07f558d"><strong data-renderer-mark="true">Job Responsibilities:</strong></p> <ul> <li>Build and operate a model gateway routing inference across open and closed models with cost, latency, and quality tracking</li> <li>Design and run distillation pipelines: use frontier model outputs to generate training data for task-specific open models</li> <li>Fine-tune and evaluate open-weight models (Llama, Qwen, Mistral, or similar) for DV-specific tasks</li> <li>Deploy and maintain on-prem inference infrastructure (vLLM, TGI, or equivalent) on KubernetesBuild model evaluation frameworks for quality, cost, latency, and regression</li> <li>Define criteria and tooling for model selection: when open models are production-ready vs. when to use closed APIs</li> <li>Partner with the agent engineering team to ensure the model layer meets agent workload</li> </ul> <p data-renderer-start-pos="731" data-local-id="db73e6d3f6c4"><strong data-renderer-mark="true">Requirements:</strong></p> <ul> <li data-renderer-start-pos="731" data-local-id="db73e6d3f6c4">5+ years software engineering; strong Python</li> <li data-renderer-start-pos="731" data-local-id="db73e6d3f6c4">Production fine-tuning or distillation of open-weight models (not just inference API wrappers)</li> <li data-renderer-start-pos="731" data-local-id="db73e6d3f6c4">Experience serving LLMs on-prem (vLLM, TGI, Triton, or equivalent)</li> <li data-renderer-start-pos="731" data-local-id="db73e6d3f6c4">Experience managing GPU infrastructure (provisioning, scheduling, utilization monitoring) in a&nbsp;production environment</li> <li data-renderer-start-pos="731" data-local-id="db73e6d3f6c4">Model evaluation and regression testing in production</li> <li data-renderer-start-pos="731" data-local-id="db73e6d3f6c4">Kubernetes and GPU workload management</li> <li data-renderer-start-pos="731" data-local-id="db73e6d3f6c4">Strong grasp of the tradeoffs between open and closed models across cost, quality, latency, and data&nbsp;sensitivity</li> </ul> <p data-renderer-start-pos="731" data-local-id="db73e6d3f6c4"><strong>Preferred:</strong></p> <ul> <li data-renderer-start-pos="731" data-local-id="db73e6d3f6c4">Quantization, PEFT/LoRA, or other efficient training techniques</li> <li data-renderer-start-pos="731" data-local-id="db73e6d3f6c4">Model gateway or inference proxy design (routing, fallback, rate limiting)</li> <li data-renderer-start-pos="731" data-local-id="db73e6d3f6c4">Financial services or other regulated/sensitive-data environments</li> <li data-renderer-start-pos="731" data-local-id="db73e6d3f6c4">Familiarity with the open model ecosystem (Hugging Face, model cards, licensing</li> </ul> <p data-renderer-start-pos="748" data-local-id="1ac93e0b889c"><strong data-renderer-mark="true">Benefits: </strong></p> <ul> <li>Discretionary bonus eligibility</li> <li>Medical, dental, and vision insurance</li> <li>HSA, FSA, and Dependent…
Skills asked for
- python
- kubernetes
Your next role is already in here.
Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.