Senior AI/ML Engineer: LLM & Agent Stack

TrueFoundry

Bengaluru

On-site

INR 4,500,000 - 6,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TrueFoundry is hiring a Senior AI/ML Engineer focused on LLM & Agent Stack to build enterprise-grade orchestration for production agentic AI. You’ll enable multi-agent workflows, integrate model runtimes, and enforce governance across tools and data.

You will drive reliable, observable LLM deployments with cost controls, security, and data residency considerations while collaborating with customers and cross-functional teams.

Qualifications

  • 4–10 years of software engineering with distributed systems experience.
  • Deep practical experience deploying LLMs in production (RAG, embeddings pipelines).
  • Hands-on experience with agent orchestration frameworks (LangGraph / LangChain) and stateful workflow design.
  • Proven track record building observability, cost controls, and policy enforcement for production services.

Responsibilities

  • Architect scalable agent orchestration patterns for production workloads.
  • Own critical integrations: model adapters, LLM gateway hooks, vector DBs, tools & external APIs, and the platform’s LLMops flows.
  • Build and improve tracing, benchmarking and observability for LLMs and agents, token/cost accounting, latency p95, throughput.
  • Drive design for safety/guardrails: moderation hooks, human-in-the-loop checkpoints, replayable audit trails and policy enforcement.
  • Mentor junior engineers, run design reviews, and improve engineering practices (testing, CI/CD, chaos testing for agents).
  • Work directly with strategic customers to prototype complex agentic solutions and translate them into product features.

Skills

Distributed systems
LLMs in production
Agent orchestration
Observability & cost controls

Tools

LangGraph
LangChain
Vector stores

Job description

About TrueFoundry:

Every production AI system, whether it's powering customer support, writing code, analyzing financial data, or diagnosing medical conditions, needs the same foundational infrastructure.

A way to route between models. A way to manage tools and integrate them securely. A way to orchestrate agents and enforce governance. A unified compute layer to run it all.

That infrastructure layer is being built right now.

We're TrueFoundry, and we're building it. We're looking for a Senior AI/ML Engineer: LLM & Agent Stack to join the team.

The Problem We're Solving

Companies are moving beyond simple chatbots to production agentic systems. These systems route between OpenAI, Anthropic, Google, and self-hosted models. They integrate dozens of tools via protocols like MCP. They orchestrate multi-agent workflows where agents coordinate with other agents.

The infrastructure to support this doesn't exist yet. You can't just duct-tape together a few API calls and call it production-ready.

You need a control plane that handles:

  • Intelligent routing with observability, cost policies, and fallback logic
  • Centralized tool and MCP server management with security and lifecycle controls
  • Agent orchestration with governance and guardrails
  • A unified compute layer to run self-hosted models, custom tools, and agents

We've built two products to solve this:

AI Gateway is the control plane, a five-composable components (Prompts, LLM Gateway, MCP Gateway, Guardrails, Agent Gateway) that handle routing, orchestration, and governance.

AI Deploy is the compute layer, a Kubernetes-based platform that abstracts ML workloads as standard software primitives, so everything runs on unified infrastructure.

We're Series A, backed by Intel Capital and Sequoia. Companies like CVS, Mastercard, Siemens, Paytm, Synopsys, and Zscaler run production AI workloads on our platform.

Role summary

You’ll design and own core components that enable enterprise customers to run production agentic AI safely and efficiently on TrueFoundry. This includes building robust orchestration for multi-step agents (graph/stateful workflows), model/routing logic, observability and policy enforcement (cost, data residency, rate limiting), and integrating upstream tooling like LangGraph, LangChain, vector stores, and specialized LLM runtimes.

What you’ll do
  • Architect and implement scalable agent orchestration patterns (graph-based executors, state management, multi-agent coordination) for production workloads.

  • Own critical integrations: model adapters, LLM gateway hooks, vector DBs, tools & external APIs, and the platform’s LLMops flows.

  • Build and improve tracing, benchmarking and observability for LLMs and agents, token/cost accounting, latency p95, throughput, and correctness checks.

  • Drive design for safety/guardrails: moderation hooks, human-in-the-loop checkpoints, replayable audit trails and policy enforcement.

  • Mentor junior engineers, run design reviews, and improve engineering practices (testing, CI/CD, chaos testing for agents).

  • Work directly with strategic customers to prototype complex agentic solutions and translate them into product features.

Must-have
  • 4–10 years of software engineering with substantial experience building distributed systems, infra, or ML platforms.

  • Deep practical experience integrating and deploying LLMs in production (RAG, retrieval, embeddings pipelines).

  • Hands-on experience with agent orchestration frameworks (LangGraph / LangChain or custom agent runtimes) and stateful workflow design.

  • Proven track record building observability, cost controls, and policy enforcement for production services.

Preferred / differentiators
  • Experience building or contributing to open-source LLM orchestration tools (LangGraph, LangChain, or similar).

  • Familiarity with enterprise constraints: on-prem/cloud hybrid deployments, data residency, compliance requirements.

  • Background in security, privacy, or model governance for LLMs.

  • Demonstrated leadership in cross-functional projects and direct customer engagement.

Qualifications & signals we like
  • Open-source contributions, architecture blogs, or public talks on agentic LLMs or LLMops.

  • Examples of productizing research or shipping complex infra features.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI/ML Engineer (Customer Facing)
Senior AI/ML Engineer (Customer Facing)

TrueFoundry • Bengaluru

Hybrid
INR 4,000,000 - 6,000,000
Health Insurance
Flexible hybrid work
Lunch and snacks provided
+2
Senior Software Engineer (Core Team)
Senior Software Engineer (Core Team)

TrueFoundry • Bengaluru

Hybrid
INR 3,500,000 - 7,000,000
Health insurance
Flexible hybrid work
Lunch and snacks
Principal Staff Engineer
Principal Staff Engineer

TrueFoundry • Bengaluru

Hybrid
INR 4,500,000 - 7,500,000
Health insurance
Hybrid work
Lunch snacks
+1
Founding Engineer, AI Hyderabad · onsite · Full Time Apply →
Founding Engineer, AI Hyderabad · onsite · Full Time Apply →

CENNA Systems Inc. • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Senior Engineer
Senior Engineer

Western Digital • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Software Engineering Manager — AI & Agentic Systems
Software Engineering Manager — AI & Agentic Systems

Readyly • India

On-site
INR 4,000,000 - 7,000,000
AI/ML Engineer
AI/ML Engineer

Recrew AI • Bengaluru Urban

On-site
INR 1,200,000 - 1,800,000
Ownership of agentic AI systems
Access to latest LLM models and AI工具
Global delivery network collaboration
+1
Forward Deployed Software Engineer
Forward Deployed Software Engineer

TrueFoundry • Bengaluru

Hybrid
INR 1,500,000 - 2,800,000
Health insurance for you and family
Flexible hybrid work: 3 days in office
Lunch and snacks
+1
AI Engineer
AI Engineer

Andpayments • India

On-site
INR 1,800,000 - 3,000,000
AI/ML Engineer – Agentic AI & LLM Systems
AI/ML Engineer – Agentic AI & LLM Systems

Recrew AI • Bengaluru Urban

On-site
INR 1,500,000 - 2,400,000