Staff Machine Learning Engineer

Worky

Santa Clara (CA)

On-site

USD 176,000 - 308,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health plans
401(k) Plan with company match
ESPP
Matching donations
Flexible time away plan
Family leave programs

Job summary

ServiceNow is seeking a Senior Staff engineer to shape the next generation of agentic systems. You will own cross-cutting architecture, tool boundaries, and streaming protocols while collaborating with teams across Java/Spring Boot and React/TypeScript frontend.

Strong capability in multi-agent coordination, context management, and real-time execution is essential. The role expects hands-on fluency across the stack, customer-facing deployments, and a drive to push the platform forward with

Qualifications

  • 8+ years building production software, with several years specifically shipping LLM-powered / agentic systems.
  • Experience integrating AI into workflows and decision-making.
  • Deep, hands-on expertise with LangGraph and/or LangChain.
  • Strong understanding of MCP and building MCP servers.

Responsibilities

  • Lead multi-agent orchestration across LangGraph/LangChain for transaction editing and multi-product quote planning.
  • Own the Harness agent execution runtime and shape next steps.
  • Shape MCP-first vs Harness surface for external interoperability.
  • Manage A2A protocol and streaming to external systems.
  • Collaborate with Forward Deployed Engineering with customer deployments.
  • Address RAG/context engineering and embedding pipelines.
  • Read and modify Spring Boot services and React components end-to-end.

Skills

Multi-agent orchestration
Full-stack fluency
Architectural judgement
Reading Spring/Bom services
Ambiguity tolerance
Forward Deployed Engineering
A2A protocol design
Real-time streaming

Tools

LangGraph
LangChain
FastAPI
WebSockets
Spring Boot
React
TypeScript
Python asyncio

Job description

Company Description

It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.

Join us to put AI to work for people.

Job Description
Team Overview

We build the AI layer of our CPQ platform — a set of Python services that let users configure, quote, and manage transactions through natural language instead of forms. This isn't a thin LLM wrapper. We're running multiple production agent architectures concurrently (ReAct-style tool-calling agents, hand-rolled LangGraph state machines, and the Harness — our from-scratch, industry-leading agent execution runtime). Our systems are backed by a first-party MCP surface into admin/product/rules/transaction systems and interoperate with other AI agents over the A2A protocol. Below that sits a conventional Java/Spring Boot microservices fleet and a React/TypeScript + Lit frontend that the agents ultimately drive.

Role Overview

We're looking for someone who already operates at a Senior-Staff bar in the agentic/LLM domain but is building out breadth across the rest of the stack. You'll be one of the most senior technical voices on how agentic systems get designed here — state management, tool boundaries, streaming protocols, prompt/context architecture, and multi-agent coordination — while staying credible end-to-end: able to read a Spring Boot service, unblock a frontend integration, or reason about a classical ML model pipeline when the problem calls for it.

What you get in this role:
  • Multi-agent orchestration — LangGraph/LangChain agents over frontier LLMs for transaction editing, conversational configuration, and multi-product quote planning with plan/approve/refine loops and parallel task execution
  • The Harness — we're crystallizing our own agent execution runtime into an industry-leading, state-of-the-art harness. Full-duplex sessions where a user can interrupt, redirect, or answer a clarifying question mid-execution while other work keeps streaming, built on a from-scratch async runtime rather than a bolted-on wrapper around someone else's agent loop. This is as much a performance and UX problem as a backend one — low-latency streaming, backpressure, live progress, partial results, graceful cancellation — and it's the part of the stack we're most invested in owning outright. You'd be a primary owner of where this goes next.
  • MCP as a secondary interface — we maintain a first-party MCP server and clients into our admin/product/rules/transaction systems, but as the Harness matures it becomes the primary way our own agents interact with the platform, with MCP kept as the secondary, standards-based surface for external interop. You'd help decide what stays MCP-first and what moves onto the Harness.
  • A2A protocol — agent-to-agent task delegation and streaming, surfaced through an external gateway so other systems (including core ServiceNow) can drive our agents directly
  • Forward Deployed Engineering — expect real time embedded with customer- and product-facing teams against live deployments. Adapting the Harness and our agents to actual customer workflows under real constraints, not just building platform capability in the abstract
  • RAG / context engineering — tenant-uploaded document ingestion, categorization, and aggregation into agent context. Prefix-cacheable prompt design for cost/latency
  • Classical ML, when the problem isn't a good fit for an LLM — we have a separate PyTorch/scikit-learn training and serving pipeline (field-value prediction) that a whole-stack ML engineer should be able to read, extend, or evaluate against LLM-based alternatives
  • Full-stack fluency — enough comfort in Spring Boot/Java services and the React/TypeScript + Lit frontend to unblock an integration end-to-end without waiting on a handoff
Qualifications
To be successful in this role you have:
  • 8+ years building production software, with several years specifically shipping LLM-powered / agentic systems (not just API wrapper calls — real tool-use loops, state management, multi-turn orchestration)
  • Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
  • Deep, hands‑on expertise with LangGraph and/or LangChain (or the judgment to know when to skip them and hand‑roll something better)
  • Strong understanding of MCP — ideally having built an MCP server, not just consumed one
  • Production async Python (FastAPI, asyncio) — comfortable with WebSockets, streaming, and the failure modes of long-lived stateful connections
  • Track record of making real architecture decisions on agent systems — tool boundaries, context/state design, cost/latency tradeoffs, when a bounded agent beats a fully autonomous one
  • Enough range outside Python to read/modify a Spring Boot service and a React component without hand‑holding — this is explicitly a whole‑stack role, not "Python specialist with a frontend allergy"
  • Comfort operating with ambiguity and setting technical direction, not just executing a spec — this is a Staff‑level bar on judgment
  • A demonstrated habit of pulling the latest from the industry — new agent frameworks, protocol standards, model capabilities — into production quickly
  • Willingness and ability to build real fluency in the Core ServiceNow platform. Our systems increasingly need to interoperate with it directly, and this role is expected to help drive that, not just react to it
  • Willingness to work directly with customers/deployments as part of Forward Deployed Engineering efforts — this isn't a purely internal‑platform role
Desired Qualifications
  • Experience with classical ML (PyTorch/scikit‑learn) in addition to LLM-based systems
  • Experience with A2A or other agent-to-agent interop protocols
  • Experience with RAG / knowledge‑graph systems (embeddings, vector or graph‑based retrieval)
  • Multi‑tenant SaaS experience, especially around per‑tenant isolation of stateful connections/resources
  • CPQ, quoting, or transaction/pricing domain experience
  • Prior Forward Deployed Engineering experience, or time spent embedded with customers shipping bespoke solutions exposure

For positions in this location, we offer a base pay of $176,100 - $308,200, plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.

We also offer

  • health plans, including flexible spending accounts
  • a 401(k) Plan with company match
  • ESPP
  • matching donations
  • a flexible time away plan
  • family leave programs
Additional Information
Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third‑party service.

Equal Opportunity Employer

ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.

Accommodations

We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact for assistance.

Export Control Regulations

For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.

From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.

Job Details
  • Employee Type: Regular
  • Region: AMS - North America and Canada
  • Work Persona: Flexible
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Machine Learning Engineer
Staff Machine Learning Engineer

ServiceNow • Santa Clara (CA), Northern (KY)

On-site
USD 176,000 - 308,000
Staff Machine Learning Engineer
Staff Machine Learning Engineer

ServiceNow, Inc. • Santa Clara (CA)

On-site
USD 176,000 - 308,000
Staff AI Engineer - Conversational & Agentic AI
Staff AI Engineer - Conversational & Agentic AI

ServiceNow • Santa Clara (CA)

Hybrid
USD 176,000 - 309,000
Health plans
401(k) Plan with company match
Flexible time away plan
+1
Forward Deployed Solution Engineer – AppliedAI FDE
Forward Deployed Solution Engineer – AppliedAI FDE

Worky • Santa Clara (CA)

On-site
USD 201,000 - 352,000
Assoc Applications Dev Engineer
Assoc Applications Dev Engineer

Worky • Santa Clara (CA)

On-site
USD 109,000 - 140,000
Health plans
401(k) plan
ESPP
+1
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Latitude • Santa Clara (CA), Northern (KY)

Hybrid
USD 176,000 - 308,000
Health plans
401(k) with company match
ESPP
+2
Staff Full Stack Software Engineer
Staff Full Stack Software Engineer

Worky • Santa Clara (CA)

Hybrid
USD 167,000 - 291,000
Health plans
401(k) with company match
ESPP
+3
Senior Software Engineer - Agent Development
Senior Software Engineer - Agent Development

ServiceNow • California (MO)

On-site
USD 143,000 - 244,000
Equity
401(k) Plan with company match
Employee Stock Purchase Plan (ESPP)
+3
Senior Staff Inbound Product Manager, ServiceNow SDK -Platform Agentic Build Lifecycle Team
Senior Staff Inbound Product Manager, ServiceNow SDK -Platform Agentic Build Lifecycle Team

ServiceNow • San Diego (CA)

Hybrid
USD 171,000 - 301,000
Health plans
401(k) Plan
ESPP
+2
Senior Staff Machine Learning Engineer - Agentic AI
Senior Staff Machine Learning Engineer - Agentic AI

Servicenow • Santa Clara (CA)

On-site
USD 201,000 - 352,000
Equity (when applicable)
Health plans
401(k) Plan with company match
+1