Senior Engineering Manager, Agentic & Generative AI Benchmarking and Evaluations

ServiceNow

Santa Clara (CA)

On-site

USD 201,300 - 352,300

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ServiceNow is seeking a Senior Engineering Manager to establish and lead AI evaluation practices for Now Assist agentic workflows. You will build infrastructure, rigorous validation frameworks, and benchmarks that quantify performance across multi-step orchestration, tool-calling, and enterprise-grounded reasoning.

You will bridge frontier AI research and production metrics, directly impacting trust and adoption of autonomous workflows for millions of users.

Qualifications

  • 8+ years of professional software engineering or machine learning experience.
  • 3+ years managing or technically leading AI/ML teams.
  • Experience integrating AI into enterprise work processes.
  • Strong knowledge of frontier AI SDKs and agentic architectures.
  • Experience implementing AI metrics at enterprise scale.
  • Bachelor’s or higher degree in Computer Science, Data Science, ML, or related quantitative field.

Responsibilities

  • Build and scale AI evaluation infrastructure, unit/integration tests, and production drift monitors.
  • Define enterprise AI benchmarks for complex business workflows and multi-agent orchestration.
  • Validate RAG pipelines, hybrid search, and semantic re-ranking systems.
  • Benchmark frontier LLMs for latency, context-window efficiency, and costs.
  • Lead a high-performing AI-native engineering team and promote best practices.
  • Collaborate with Core Product, ML Platforms, and Engineering leads to drive product improvements.

Skills

Senior AI/ML leadership
Python proficiency
AI evaluation metrics
Multi-agent orchestration
Data analytics (SQL, Pandas, NumPy)
MLOps tracking platforms

Education

Bachelor's or higher in CS/DS/ML

Tools

Python
SQL
Pandas
NumPy
MLOps platforms

Job description

Job Description

It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.

Join us to put AI to work for people.

The Advanced Technology Group (ATG) at ServiceNow is a customer-focused innovation group building intelligent software and smart user experiences using existing and latest advanced technologies to enable end-to-end, industry-leading work experiences for customers. We are a group of researchers, applied scientists, engineers, and product managers with a dual mission. We build and evolve the AI platform, and partner with teams to build products and end-to-end AI-powered work experiences. In equal measure, we lay the foundations, research, experiment, and de-risk AI technologies that unlock new work experiences in the future.

Job Description

We are seeking an exceptional, data-driven Senior Engineering Manager, Agentic & GenAI Benchmarking and Evaluations to establish and lead AI evaluation practices for both ServiceNow and our customers. As ServiceNow shifts enterprise workflows from simple generation to complex, autonomous agents, ensuring system reliability, safety, and accuracy is paramount.

In this role, you will lead a specialized team of AI evaluation engineers and data scientists. Your team will build the infrastructure, rigorous validation frameworks, and benchmarks that quantify the performance of Now Assist agentic workflows across multi-step orchestration, tool-calling, and enterprise-grounded reasoning. You will bridge the gap between frontier AI research and hard production metrics, directly impacting the trust and adoption of autonomous workflows for millions of enterprise users.

What You Get To Do In This Role
  • Build the Evaluation Infrastructure: Design, own, and scale automated testing and evaluation harnesses (unit evals, integration evals, and production drift monitors) to measure agent quality and eliminate regressions.
  • Define Enterprise AI Benchmarks: Create standard, repeatable evaluation frameworks tailored to complex business workflows—assessing multi-agent orchestration, intent routing, multi-step planning loops, and long-term memory accuracy.
  • Validate Grounding & RAG Pipelines: Partner with search and data fabric teams to systematically evaluate Retrieval-Augmented Generation (RAG) pipelines, hybrid search, and semantic re-ranking systems.
  • Model Selection Optimization: Rigorously benchmark frontier LLMs (e.g., OpenAI, Anthropic, Google, and proprietary ServiceNow models) to evaluate trade-offs across execution capabilities, latency, context-window efficiency, and inference costs.
  • Lead a High-Performing Team: Recruit, mentor, and foster an AI-native engineering team, driving engineering best practices, prompt-infrastructure stability, and production-grade rigor.
  • Cross-Functional Leadership: Collaborate with Core Product, Machine Learning Platforms, and Engineering leads to translate baseline performance statistics into actionable product improvements and model fine-tuning targets.
To be successful in this role you have:
  • 8+ years of professional software engineering or machine learning experience, including 3+ years managing or technically leading high-performing AI/ML teams.
  • Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
  • Strong foundational knowledge of frontier AI SDKs and deep experience deploying or testing agentic/probabilistic software architectures (multi-agent orchestration, tool execution, and probabilistic feature deployment).
  • Demonstrated experience implementing rigorous AI metrics (e.g., ROUGE, BLEU, G-Eval, LLM-as-a-judge patterns, and custom deterministic evaluation code) at an enterprise scale.
  • Proficiency in Python and familiarity with data analytics infrastructures (SQL, Pandas, NumPy) alongside standard MLOps tracking platforms.
  • Experience with complex knowledge infrastructure, SaaS platform architectures, or relational datasets (e.g., Knowledge Graphs, CMDBs).
  • Ability to translate deeply technical evaluation data into executive-level risk assessments, ROI summaries, and strategic roadmap recommendations.
  • Bachelor’s or higher degree in Computer Science, Data Science, Machine Learning, or a highly quantitative field (Master's or Ph.D. is a plus).

For positions in this location, we offer a base pay of $201,300 - $352,300, plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

Equal Opportunity Employer

ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.

Accommodations

We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact globaltalentss@servicenow.com for assistance.

Export Control Regulations

For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.

From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Staff, Product Manager- AI Evaluation & Quality
Senior Staff, Product Manager- AI Evaluation & Quality

ServiceNow • California (MO)

On-site
USD 190,000 - 335,000
Health plans
401(k) Plan with company match
ESPP
+3
Senior Software Engineer - Agent Development
Senior Software Engineer - Agent Development

ServiceNow • California (MO)

On-site
USD 143,000 - 244,000
Equity
401(k) Plan with company match
Employee Stock Purchase Plan (ESPP)
+3
Staff AI Engineer - Conversational & Agentic AI
Staff AI Engineer - Conversational & Agentic AI

ServiceNow • Santa Clara (CA)

Hybrid
USD 176,000 - 309,000
Health plans
401(k) Plan with company match
Flexible time away plan
+1
Principal Machine Learning Engineer
Principal Machine Learning Engineer

ServiceNow • San Diego (CA)

On-site
USD 216,000 - 379,000
Health plans
401(k) Plan with company match
ESP P
+3
Software Engineering Manager - Build Agent
Software Engineering Manager - Build Agent

ServiceNow • California (MO)

On-site
USD 166,500 - 291,400
Staff AI Engineer - Conversational & Agentic AI
Staff AI Engineer - Conversational & Agentic AI

ServiceNow • Santa Clara (CA)

On-site
USD 176,000 - 309,000
Health plans
401(k) Plan with company match
Flexible time away plan
+1
Senior Staff Machine Learning Engineer - Agentic AI
Senior Staff Machine Learning Engineer - Agentic AI

Servicenow • Santa Clara (CA)

On-site
USD 201,000 - 352,000
Equity (when applicable)
Health plans
401(k) Plan with company match
+1
Software Engineering Manager - Build Agent
Software Engineering Manager - Build Agent

ServiceNow • Santa Clara (CA)

On-site
USD 166,500 - 291,400
Health plans
401(k) Plan with company match
ESPP
+3
Solution Architect - AI & Data
Solution Architect - AI & Data

Servicenow • Denver (CO)

Hybrid
USD 173,000 - 271,000
401(k) Plan with company match
Flexible spending accounts
Health plans
+1
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Latitude • Santa Clara (CA), Northern (KY)

Hybrid
USD 176,000 - 308,000
Health plans
401(k) with company match
ESPP
+2