Senior AI Devops Engineer

Demandbase

Hyderabad

On-site

INR 4,500,000 - 7,000,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

LLM gateway experience
LLM observability tooling
Eval frameworks
Guardrails/PII filtering

Job summary

Demandbase seeks an experienced AI Runtime Platform Engineer to own the gateway, reliability, and cost-controls for AI features in production. You will drive LLM-centric tooling, observability, and guardrails across the runtime and internal tooling.

The role requires strong Kubernetes, AWS, Terraform, and Python skills; on-call and incident-response experience are essential. You will work in Hyderabad, contributing to a fast-growing, security-minded AI platform.

Qualifications

  • Production infra background with systems you were paged for.
  • Experience owning LLM-driven or AI-enabled services.
  • Ability to design reliable, observable runtimes and cost controls.
  • Familiar with incident response and postmortems.

Responsibilities

  • Own the LLM gateway and model-provider routing and quotas.
  • Maintain reliability with SLOs, on-call, and incident reviews.
  • Implement cost attribution and budgets per team/model.
  • Improve observability with traces, metrics and billing visibility.
  • Build eval infra for prompt/model changes and experiments.
  • Develop caching and guardrails for safety and efficiency.

Skills

LLM systems
Python
Observability
On-call
Cost control
Incident response
Strong infra background

Tools

Kubernetes
AWS
Terraform
GitOps

Job description

Demandbase is the pipeline AI platform that empowers go-to-market teams to automate growth at scale. By bringing together data, insights, actions, and outcomes in one platform, Demandbase helps B2B enterprises align and execute account-based GTM strategies with confidence.

Thousands of businesses trust Demandbase to maximize revenue, reduce waste, and consolidate their data and technology stacks. We are equally committed to building world-class technology and growing meaningful careers. Demandbase has been recognized as one of the Best Places to Work in the San Francisco Bay Area by Fortune and one of the 60 Best Companies to Sell For by Selling Power. Our offices are located in San Francisco, New York, Austin, Seattle, India, and the United Kingdom.

About the Role

AI Runtime Services owns the shared runtime that AI at Demandbase runs on — two ways. First, it's the layer that enables and supports the AI features in the Demandbase platform: product teams ship LLM features on top of it without reinventing model access, spend control, safety, and observability. Second, the team builds and supports internal AI solutions — the tooling and services the company uses to work with LLMs day to day. The pillars underneath both are: the LLM gateway, cost controls, observability, evals infra, caching, and guardrails.

The surface is broad and the team is early. You won't be handed a narrow slice: expect to move across the gateway one week and the ingest pipeline the next — and across product-facing runtime and internal tooling — and to own what you build in production.

What you'll own
  • The LLM gateway. The single front door to every model provider we use — keys, routing, rate limits, failover, provider quotas. When it's down, every AI feature at Demandbase is down.
  • Reliability, for real. SLOs, on‑call, incident response, and the postmortems for the runtime. This is a genuine SRE ownership role, not "build it and let someone else run it."
  • Cost control. Per‑team and per‑model attribution, budgets, quotas. LLM spend is the kind of line item that quietly triples; your job is to make it legible and then make it smaller.
  • LLM observability. Traces, spans, prompt/response capture, ingest and billing visibility for every AI feature in the company.
  • Evals and experimentation infra. The shared harnesses, datasets, and scoring plumbing product teams use to know whether a prompt or model change actually made things better.
  • Caching and guardrails. Response and semantic caching to cut latency and redundant spend; input/output safety, PII handling, and policy enforcement at the gateway.
What we need
  • Strong production infrastructure background: Kubernetes, AWS, Terraform, GitOps. You've operated systems that people were paged for, and you've been the one paged.
  • Python, at a level where you're comfortable owning services and tooling in it — not just scripting.
  • Real on‑call and incident‑response experience. You can talk about an incident you ran, what you got wrong, and what you changed afterward.
  • Observability fluency beyond "we have dashboards": you've instrumented systems, chased cardinality and cost in a metrics/logging backend, and built alerts that fire when they should.
  • Hands‑on familiarity with how LLM systems actually work. You don't need production ML experience. But you need to have built something with these tools — an agent, a RAG pipeline, an internal tool — and to have used evals or experiments to decide whether it was any good. Tokens, context windows, prompt/response tracing, and why an eval suite is fundamental should all be familiar ground. Candidates who have only read about this are not a fit.
Nice to have
  • An LLM gateway or proxy (LiteLLM or similar) in production.
  • LLM observability tooling: Datadog LLM Observability, LangSmith, Braintrust, Arize, or similar.
  • Eval frameworks and LLM‑as‑judge scoring in a real workflow, not a demo.
  • Guardrails, PII detection, or content‑filtering systems.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps Engineer - AI Runtime Services
DevOps Engineer - AI Runtime Services

Demandbase • Hyderabad

On-site
INR 3,800,000 - 6,200,000
Group Medical
Personal Accident Insurance
Term Life Insurance
+4
DevOps Engineer, AI Runtime Services
DevOps Engineer, AI Runtime Services

demandbase • Hyderabad

On-site
INR 2,400,000 - 4,200,000
Group Medical insurance
Personal Accident insurance
Term Life Insurance
+5
DevOps Engineer - AI Runtime Services
DevOps Engineer - AI Runtime Services

Demandbase, Inc. • Hyderabad

On-site
INR 2,500,000 - 4,000,000
Group Medical
Personal Accident
Term Life Insurance
+4
Machine Learning Engineer II
Machine Learning Engineer II

Latitude • India

On-site
INR 2,400,000 - 4,200,000
Group medical insurance
Mental health support
Fitness benefit
+1
Machine Learning Engineer II
Machine Learning Engineer II

Demandbase • Hyderabad

On-site
INR 1,500,000 - 2,100,000
Group Medical Insurance
Personal Accident Insurance
Term Life Insurance
+4
Sr. AI Engineer
Sr. AI Engineer

Solytics Partners Careers • Pune District

On-site
INR 1,200,000 - 1,800,000
Lead AI Engineer
Lead AI Engineer

The Ksquare Group • Hyderabad

On-site
INR 3,500,000 - 6,000,000
Senior AI Engineer
Senior AI Engineer

Spice Money • Dadri

On-site
INR 1,800,000 - 3,200,000
Staff AI & Automation Data Analyst
Staff AI & Automation Data Analyst

Pure Storage • Bengaluru

On-site
INR 2,800,000 - 4,600,000
Senior AI/ML Engineer
Senior AI/ML Engineer

ApplyMint • Bengaluru

On-site
INR 1,800,000 - 2,800,000