Senior Research Scientist, Agent Evaluation

ServiceNow

Telangana

Hybrid

INR 2,500,000 - 4,500,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

ServiceNow is seeking a skilled ML/AI engineer to own core parts of the evaluation platform for AI agents, ensuring scalable, model-based judging and reliable scoring for customer-facing evaluation.

You will design benchmarks, curate datasets, and ship end-to-end features, leveraging Python and LLM APIs, with a focus on accuracy, latency, and cost tradeoffs in a distributed work environment.

Qualifications

  • 5+ years in ML, applied AI, prompt engineering, agentic AI, including shipping something real to users.
  • Strong Python skills.
  • Practical depth in agentic AI and context engineering: planning, reasoning, memory, tool use, retrieval, long-context.
  • Experience designing evaluation methodologies, not just running evaluations.
  • Hands-on production work with LLM APIs — prompt engineering, structured output, cost and latency tradeoffs.
  • Clear communication with technical and non-technical audiences.
  • Nice to have experience with AI-assisted development tools (Claude code, Windsurf, or similar).

Responsibilities

  • Design evaluation methodologies and benchmarks for agent reasoning, planning, tool use, reliability, and safety — across LLM-as-a-Judge, trajectory-based, and human evaluation.
  • Take problems from research question to prototype to shipped feature, owning them end to end.
  • Build and harden the pipelines and scoring logic behind customer-facing evaluation.
  • Curate synthetic and real-world datasets; measure the evaluator itself for consistency and agreement with human labels.

Skills

Python
LLM APIs
Agentic AI
Prompt engineering
Communication
AI tooling

Tools

Claude code
Windsurf

Job description

Company Description

It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.

Join us to put AI to work for people.

Job Description

We build the evaluation layer that validates AI agents before they reach customers — an automated system that scores across large volumes of agent traces.
You'll own core parts of that platform: the pipeline that runs traces through model-based judges at scale, and the scoring logic that turns raw output into results teams can act on.

What you'll do

  • Design evaluation methodologies and benchmarks for agent reasoning, planning, tool use, reliability, and safety — across LLM-as-a-Judge, trajectory-based, and human evaluation
  • Take problems from research question to prototype to shipped feature, owning them end to end
  • Build and harden the pipelines and scoring logic behind customer-facing evaluation
  • Curate synthetic and real-world datasets; measure the evaluator itself for consistency and agreement with human labels
Qualifications

What we're looking for

  • 5+ years in ML, applied AI, prompt engineering, agentic AI, including shipping something real to users
  • Strong Python skills
  • Practical depth in agentic AI and context engineering: planning, reasoning, memory, tool use, retrieval, long-context
  • Experience designing evaluation methodologies, not just running evaluations
  • Hands-on production work with LLM APIs — prompt engineering, structured output, cost and latency tradeoffs
  • Clear communication with technical and non-technical audiences
  • Good to have Experience with AI-assisted development tools (Claude code, Windsurf, or similar)
Additional Information

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

Equal Opportunity Employer

ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.

Accommodations

We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact globaltalentss@servicenow.com for assistance.

Export Control Regulations

For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.

From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Research Scientist, Agent Evaluation
Senior Research Scientist, Agent Evaluation

SmartRecruiters, Inc. • Hyderabad

On-site
INR 4,000,000 - 8,000,000
Senior Research Scientist, Agentic AI
Senior Research Scientist, Agentic AI

ServiceNow • Hyderabad

On-site
INR 4,800,000 - 6,500,000
Senior Research Scientist, Agent Evaluation
Senior Research Scientist, Agent Evaluation

ServiceNow, Inc. • Hyderabad

On-site
INR 4,000,000 - 7,500,000
Senior Research Engineer/Scientist
Senior Research Engineer/Scientist

ServiceNow • Hyderabad

Hybrid
INR 2,600,000 - 4,200,000
Staff Software Engineer
Staff Software Engineer

Servicenow • Telangana

On-site
INR 2,500,000 - 5,000,000
Senior Software AIML Engineer
Senior Software AIML Engineer

Servicenow • Hyderabad

On-site
INR 4,000,000 - 6,000,000
Staff Technical Product Manager – AI/LLM expertise + AI Evaluation Science
Staff Technical Product Manager – AI/LLM expertise + AI Evaluation Science

Servicenow • Hyderabad

On-site
INR 3,500,000 - 6,500,000
Staff Data Engineer
Staff Data Engineer

Servicenow • Telangana

On-site
INR 4,200,000 - 7,000,000
Staff Software Engineer
Staff Software Engineer

ServiceNow • Hyderabad

On-site
INR 4,200,000 - 6,500,000
Senior Manager, Software Engineering
Senior Manager, Software Engineering

ServiceNow • Hyderabad

On-site
INR 1,600,000 - 2,400,000