Manager, AI Platform Engineer, DTS - Global Capability Center

Alvarez & Marsal

Gurugram District

On-site

INR 1,500,000 - 2,600,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Alvarez & Marsal is seeking an experienced software/AI engineer to lead the development of AI evaluation frameworks and production-grade agentic applications. You will work with experts to define outcomes, calibrate automated judging, and ensure robust safety controls.

The role emphasizes observability, cost-aware inference, and cross-provider model deployments in a dynamic consulting environment. You will collaborate with clients and internal teams to translate requirements into trusted

Qualifications

  • 8+ years of experience in software, data or AI/ML engineering.
  • Strong Python capability and hands-on experience building production-grade agentic or LLM applications.
  • Experience designing AI evaluations and measuring retrieval quality.
  • Practical experience implementing model-safety controls and adversarial testing.
  • Experience with multi-provider model access, cost, latency and capability trade-offs.
  • Consulting or client-facing delivery experience.
  • Hands-on OpenTelemetry instrumentation for nested multi-step agents.
  • Experience with LLM observability tools and platform integrations.

Responsibilities

  • Build the evaluation framework to determine whether an AI application is fit to go live, with defined pass thresholds.
  • Collaborate with business subject-matter experts to define outcomes and convert them into benchmark datasets.
  • Design and calibrate automated judging, validating it against human review for meaningful scores.
  • Own repeatable adversarial testing for prompts, jailbreaks, tool misuse, data leakage and updates.
  • Implement runtime safety controls including content filtering, redaction and guardrails.
  • Instrument applications to monitor quality, latency, tokens and model cost per request.
  • Own model-cost optimisation through routing, caching and prompt efficiency.
  • Monitor production quality and drift, maintain a failure taxonomy and feed failures into benchmarks.
  • Validate model upgrades before cutover to avoid regressions.

Skills

Python
LangGraph
LLM apps
Tool calling
Retrieval quality
OpenTelemetry
Observability

Tools

Langfuse
LangSmith
Arize Phoenix
Weights & Biases Weave
Grafana
Datadog
Azure Monitor
CloudWatch

Job description

Description
About Alvarez & Marsal

Alvarez & Marsal (A&M) is a global consulting firm with entrepreneurial, action and results-oriented professionals. We take a hands-on approach to solving our clients' problems and assisting them in reaching their potential. Our culture celebrates independent thinkers and doers who positively impact our clients and shape our industry. The collaborative environment and engaging work - guided by A&M's core values of Integrity, Quality, Objectivity, Fun, Personal Reward, and Inclusive Diversity - are why our people love working at A&M.

The Team

Our DTS team provides following services to clients:

  • Data & Applied Intelligence - Helping clients in harnessing the power of data and cutting-edge analytics to drive intelligent decision-making and transform businesses.
  • Product and Innovation - Empowering clients to innovate, develop, and launch products that drives growth and competitive advantage.
  • Technology M&A and Strategy - Assist clients to manage the technology aspects and business enablement of complex M&A, integrations and carve-outs.
  • Technology Transformation - A&M helps clients create a scalable, cost-effective IT function that delivers the company's strategic vision and priorities.
How you will contribute
  • Build the evaluation framework used to determine whether an AI application is fit to go live, covering component, trajectory and end-to-end outcome measurement, with defined pass thresholds integrated into the delivery pipeline.
  • Work directly with business subject-matter experts to define correct outcomes and convert them into maintained benchmark datasets.
  • Design and calibrate automated judging, validating it against human review so that reported scores are meaningful and defensible.
  • Own repeatable adversarial testing for prompt injection, jailbreaks, tool misuse, guardrail bypass and data leakage, and rerun the suite after every material model or prompt change.
  • Implement runtime safety controls, including content filtering, sensitive-data redaction, tool permissions, spend caps and circuit breakers.
  • Instrument applications so quality, latency, tokens and model cost are observable per request, application and tenant.
  • Own model-cost optimisation through routing, caching and prompt efficiency, using evaluation evidence to demonstrate when a lower-cost model is sufficient.
  • Monitor production quality and drift, maintain a failure taxonomy, and feed observed failures back into benchmark datasets.
  • Validate model upgrades before cutover so vendor model changes do not degrade client-facing systems.
Qualifications
  • 8+ years of experience in software, data or AI/ML engineering, with strong Python capability and hands-on experience building production-grade agentic or LLM applications for real users using LangGraph or equivalent frameworks, tool calling and retrieval.
  • Demonstrable experience designing AI evaluations and measuring retrieval quality, including defining measures and scoring methods, assessing groundedness and relevance, diagnosing failures, distinguishing genuine regressions from noise, and running online evaluations over sampled production traces.
  • Practical experience implementing model-safety controls and adversarial testing, including prompt injection, jailbreaks, tool misuse, data leakage, filtering, redaction, guardrails, spend limits and loop controls.
  • Experience accessing and deploying models across multiple providers, understanding cost, latency and capability trade-offs, and working with self-hosted or open-weight models, fine-tuning, distillation or prompt optimisation.
  • Consulting or client-facing delivery experience, including the ability to work with non-technical subject-matter experts to define correct outcomes and translate their judgement into maintained benchmark datasets.
  • Hands-on OpenTelemetry instrumentation for nested, multi-step agents, including tool-call spans, session and conversation tracing, trace-context propagation across asynchronous workers and queues, and correlation IDs through background processing.
  • Experience with LLM observability tools such as Langfuse, LangSmith, Arize Phoenix or Weights & Biases Weave; integration with platforms such as Azure Monitor, CloudWatch, Grafana or Datadog; and structured telemetry for token usage, latency and cost by request, application and tenant.
Your journey at A&M

We recognize that our people are the driving force behind our success, which is why we prioritize an employee experience that fosters each person’s unique professional and personal development. Our robust performance development process promotes continuous learning, rewards your contributions, and fosters a culture of meritocracy. With top-notch training and on-the-job learning opportunities, you can acquire new skills and advance your career. We prioritize your well-being, providing benefits and resources to support you on your personal journey. Our people consistently highlight the growth opportunities, our unique, entrepreneurial culture, and the fun we have together as their favorite aspects of working at A&M. The possibilities are endless for high-performing and passionate professionals.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Associate, Agentic/GenAI, DTS - Global Capability Center
Senior Associate, Agentic/GenAI, DTS - Global Capability Center

Alvarez & Marsal • Gurugram District

On-site
INR 1,800,000 - 3,000,000
Associate Director, Deployment Architect, DTS - Global Capability Center
Associate Director, Deployment Architect, DTS - Global Capability Center

Alvarez & Marsal • Gurugram District

On-site
INR 3,000,000 - 4,500,000
Associate Director, Technology Transformation Advisory, DTS - Global Capability Center
Associate Director, Technology Transformation Advisory, DTS - Global Capability Center

Alvarez & Marsal • Gurugram District

On-site
INR 2,400,000 - 3,600,000
Senior Director - Data & AI - Platform Engineering Lead, DTS - Global Capability Center
Senior Director - Data & AI - Platform Engineering Lead, DTS - Global Capability Center

Alvarez & Marsal • Gurugram District

On-site
INR 3,500,000 - 5,600,000
Azure DevOps Manager - DTS - Global Capability Center
Azure DevOps Manager - DTS - Global Capability Center

Alvarez & Marsal • Gurugram District

On-site
INR 3,500,000 - 5,500,000
Manager, Technology Transformation Advisory, DTS - Global Capability Center
Manager, Technology Transformation Advisory, DTS - Global Capability Center

Alvarez & Marsal • Gurugram District

On-site
INR 400,000 - 700,000
Manager, Cloud Platform Engineer, DTS - Global Capability Center
Manager, Cloud Platform Engineer, DTS - Global Capability Center

Alvarez & Marsal • Gurugram District

On-site
INR 3,200,000 - 5,200,000
Associate Director, GenAI & Data Solution Architect - Global Capability Center
Associate Director, GenAI & Data Solution Architect - Global Capability Center

Alvarez & Marsal • Gurugram District

On-site
INR 4,000,000 - 7,000,000
Manager, Resposible AI, GESS - Global Capability Center
Manager, Resposible AI, GESS - Global Capability Center

Alvarez & Marsal • Gurugram District

On-site
INR 400,000 - 900,000
Senior Associate, Security, Governance, Risk & Compliance (AI)
Senior Associate, Security, Governance, Risk & Compliance (AI)

Alvarez & Marsal • Gurugram District

On-site
INR 1,800,000 - 2,800,000