JobBobsReal-time global job discoveryLive

Sr Machine Learning Engineering Manager - AI Quality and Governance

Workiva · United States · 2026-08-03

managerRemote
Apply on the employer's site

About this role

Join Workiva as aSr Machine Learning Engineering Manager - AI Quality and Governance and help establish how we build, evaluate, release, and operate trustworthy AI products at scale. You will lead a multidisciplinary team of software, machine learning, and quality engineers responsible for two connected missions: advancing end-to-end quality across Workiva's AI platform and products, and building shared evaluation and governance capabilities that make our AI systems measurable, observable, reliable, and ready for enterprise use.
Your team's scope spans generative AI and agentic products, including AI platform services, agent frameworks and runtimes, conversational experiences, and RAG/knowledge systems. You will partner across Product, Engineering, Data Science, Security, Risk, and Legal to establish practical quality standards and embed evaluation and governance throughout the AI development lifecycle.
What You'll Do
Leadership & Team Development

Lead, mentor, and develop a multidisciplinary team of software, ML, and quality engineers

Build a culture of technical excellence, quality ownership, experimentation, and continuous improvement

Establish clear team priorities while balancing platform investments, product needs, and enterprise risk

Recruit engineers with complementary expertise across software quality, ML evaluation, platform engineering, and governance automation

AI Product Quality

Define and drive a comprehensive quality strategy for Workiva's AI platform and products, spanning unit, integration, end-to-end, performance, resilience, security, and production testing

Establish measurable quality bars, release-readiness criteria, and automated quality gates for AI and agentic capabilities

Advance testing approaches for nondeterministic systems, including RAG pipelines, agents, prompts, models, tools, and multi-step workflows

Detect regressions, model or data drift, unsafe behavior, and degraded customer experiences before and after release

AI Evaluation Platform

Lead architecture and delivery of a scalable, self-service evaluation platform for generative AI, RAG, and agentic systems

Enable teams to create, manage, version, and reuse evaluation datasets, golden test sets, task-specific metrics, graders, and benchmarks

Support deterministic checks, statistical metrics, model-based graders, human evaluation, adversarial testing, and domain-expert review

Build capabilities for offline evaluation, pre-release regression testing, online experimentation, production sampling, and continuous evaluation

Ensure evaluation results are reproducible, explainable, actionable, and integrated into developer workflows, CI/CD pipelines, and operational dashboards

AI Governance & Assurance

Translate Workiva's Responsible AI principles into practical engineering controls and platform capabilities

Build governance into the AI lifecycle through traceability, lineage, versioning, documentation, risk classification, approval workflows, and auditable evidence

Partner with Security, Legal, Privacy, Compliance, and Risk teams to define controls that support enterprise and regulated use cases

Enable inventories and traceability across models, prompts, datasets, evaluations, tools, knowledge sources, and deployed AI features

Cross-Functional Leadership

Collaborate with Product, Program Management, UX, UXR, Data Science, Security, Legal, Risk, and engineering leaders to define quality expectations and roadmaps

Influence engineering teams across Workiva to adopt shared evaluation standards, testing practices, observability, and release controls

Communicate complex technical tradeoffs, quality signals, and risk findings clearly to technical and non-technical audiences

Operational Excellence

Ensure the evaluation and governance platform is secure, scalable, reliable, observable, and cost-effective

Define service-level objectives and meaningful operational and quality metrics

Champion production readiness, incident response, root-cause analysis, and continuous operational improvement

What You'll Need
Minimum Qualifications

Bachelor's degree in Computer Science, Engineering, Data Science, or related field (or equivalent experience)

10+ years in software engineering, ML engineering, quality engineering, or related roles, including 4+ years leading an engineering team

Strong software engineering and systems-design fundamentals, with experience delivering and operating production SaaS or platform capabilities

Demonstrated experience establishing automated quality practices for distributed, cloud-based products

Practical understanding of the generative AI development lifecycle and challenges of evaluating nondeterministic systems

Experience with generative AI concepts: LLMs, RAG, embeddings, vector/hybrid search, agents, tool use, and prompt orchestration

Experience defining measurable quality criteria using data, experimentation, telemetry, and production signals

Experience with cloud-native architectures on AWS, Azure, or GCP.

Proven ability to lead senior individual contributors, navigate tehhnical disagreements, and build high-performance cultures

Strong communication and cross-functional leadership skills

Preferred Qualifications

Master's degree in Computer Science, Engineering, ML, Data Science, or related field.

Experience building or operating AI/ML evaluation, experimentation, observability, model-governance, or ML platform capabilities

Experience evaluating RAG and agentic systems, including retrieval quality, groundedness, task completion, tool use, and safety

Familiarity with evaluation techniques: golden datasets, statistical metrics, model-based graders, human evaluation, red teaming, A/B testing, and drift/regression detection

Working knowledge of ML/AI lifecycle practices: dataset management, model/prompt versioning, experiment tracking, deployment, monitoring, and feedback loops

Experience translating Responsible AI, model-risk, privacy, security, or…

Skills asked for

Similar jobs

Apply on the employer's site

Your next role is already in here.

Search live openings from thousands of employers, save the ones worth a second look, and let JobBob keep watch for the rest.