Senior AI/ML Engineer: LLM & Agent Stack

Socotra, Inc.

Hinoba-an

Hybrid

PHP 7,528,000 - 11,292,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health insurance
Hybrid work
Lunch provided
Office swap program

Job summary

TrueFoundry is seeking a Senior AI/ML Engineer: LLM & Agent Stack to design and own core components for production agentic AI, including multi-step agent orchestration, model routing, observability, and policy enforcement.

You’ll integrate LangGraph, LangChain, vector stores, and LLM runtimes, build scalable distributed systems, and mentor engineers while engaging with strategic customers to translate complex needs into product features.

Qualifications

  • 4–10 years in software engineering with distributed systems, infra, or ML platforms.
  • Deep practical experience deploying LLMs in production (RAG, embeddings pipelines).
  • Hands-on with agent orchestration frameworks (LangGraph / LangChain) or custom runtimes.
  • Proven observability, cost controls, and policy enforcement for production services.

Responsibilities

  • Architect and implement scalable agent orchestration patterns (graph-based executors, state management, multi-agent coordination) for production workloads.
  • Own critical integrations: model adapters, LLM gateway hooks, vector DBs, tools & external APIs, and the platform’s LLMops flows.
  • Build and improve tracing, benchmarking and observability for LLMs and agents, token/cost accounting, latency p95, throughput, and correctness checks.
  • Drive design for safety/guardrails: moderation hooks, human-in-the-loop checkpoints, replayable audit trails and policy enforcement.
  • Mentor junior engineers, run design reviews, and improve engineering practices (testing, CI/CD, chaos testing for agents).
  • Work directly with strategic customers to prototype complex agentic solutions and translate them into product features.

Skills

Distributed systems
LLM deployment
Agent orchestration
Observability
Security governance
Mentoring

Tools

LangGraph
LangChain
Vector stores
LLM runtimes
Model adapters
MCP gateway

Job description

About TrueFoundry:

Every production AI system, whether it's powering customer support, writing code, analyzing financial data, or diagnosing medical conditions, needs the same foundational infrastructure.

A way to route between models. A way to manage tools and integrate them securely. A way to orchestrate agents and enforce governance. A unified compute layer to run it all.

That infrastructure layer is being built right now.

We are looking for a Senior AI/ML Engineer: LLM & Agent Stack to join the team.

The Problem We're Solving

Companies are moving beyond simple chatbots to production agentic systems. These systems route between OpenAI, Anthropic, Google, and self-hosted models. They integrate dozens of tools via protocols like MCP. They orchestrate multi-agent workflows where agents coordinate with other agents.

The infrastructure to support this doesn't exist yet. You can't just duct-tape together a few API calls and call it production-ready.

You need a control plane that handles:

  • Intelligent routing with observability, cost policies, and fallback logic
  • Centralized tool and MCP server management with security and lifecycle controls
  • Agent orchestration with governance and guardrails
  • A unified compute layer to run self-hosted models, custom tools, and agents

AI Gateway is the control plane, a five-composable components (Prompts, LLM Gateway, MCP Gateway, Guardrails, Agent Gateway) that handle routing, orchestration, and governance.

We're Series A, backed by Intel Capital and Sequoia. Companies like CVS, Mastercard, Siemens, Paytm, Synopsys, and Zscaler run production AI workloads on our platform.

Role summary

You’ll design and own core components that enable enterprise customers to run production agentic AI safely and efficiently on TrueFoundry. This includes building robust orchestration for multi-step agents (graph/stateful workflows), model/routing logic, observability and policy enforcement (cost, data residency, rate limiting), and integrating upstream tooling like LangGraph, LangChain, vector stores, and specialized LLM runtimes.

What you’ll do
  • Architect and implement scalable agent orchestration patterns (graph-based executors, state management, multi-agent coordination) for production workloads.

  • Own critical integrations: model adapters, LLM gateway hooks, vector DBs, tools & external APIs, and the platform’s LLMops flows.

  • Build and improve tracing, benchmarking and observability for LLMs and agents, token/cost accounting, latency p95, throughput, and correctness checks.

  • Drive design for safety/guardrails: moderation hooks, human-in-the-loop checkpoints, replayable audit trails and policy enforcement.

  • Mentor junior engineers, run design reviews, and improve engineering practices (testing, CI/CD, chaos testing for agents).

  • Work directly with strategic customers to prototype complex agentic solutions and translate them into product features.

Must-have
  • 4–10 years of software engineering with substantial experience building distributed systems, infra, or ML platforms.

  • Deep practical experience integrating and deploying LLMs in production (RAG, retrieval, embeddings pipelines).

  • Hands-on experience with agent orchestration frameworks (LangGraph / LangChain or custom agent runtimes) and stateful workflow design.

  • Proven track record building observability, cost controls, and policy enforcement for production services.

Preferred / differentiators
  • Experience building or contributing to open-source LLM orchestration tools (LangGraph, LangChain, or similar).

  • Familiarity with enterprise constraints: on-prem/cloud hybrid deployments, data residency, compliance requirements.

  • Background in security, privacy, or model governance for LLMs.

  • Demonstrated leadership in cross-functional projects and direct customer engagement.

Qualifications & signals we like
  • Open-source contributions, architecture blogs, or public talks on agentic LLMs or LLMops.

  • Examples of productizing research or shipping complex infra features.

Traits we are looking for:Ownership, ability to execute, hustle and think out of the box, data-driven decision making, be comfortable with more unknowns than knowns.

Perks of Working at TrueFoundry
  • Join a fast-growing Series A, Bay Area-based startup building cutting-edge AI infrastructure.
  • Comprehensive health insurance for you and your family.
  • Flexible hybrid work, 3 days a week in the office (Tuesday, Wednesday and Thursday), with flexibility around the schedule.
  • Lunch and snacks are on us whenever you come into the office.
  • Work from another TrueFoundry office for up to 3 months each year, whether that's our London or Bay Area office, if you'd like a change of scenery.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Account Executive (Focus on Enterprise Sales)
Account Executive (Focus on Enterprise Sales)

TrueFoundry • San Mateo

On-site
PHP 2,466,000 - 4,316,000
Senior AI/ML Engineer - LLM & Agent Orchestration
Senior AI/ML Engineer - LLM & Agent Orchestration

Socotra, Inc. • Hinoba-an

Hybrid
PHP 7,528,000 - 11,292,000
Health insurance
Hybrid work
Lunch provided
+1
Staff Backend AI Platform Engineer
Staff Backend AI Platform Engineer

Freshworks • San Mateo

Hybrid
PHP 13,846,000 - 17,108,000
Health insurance
Dental & Vision
Equity + ESPP
+4
AI Engineer
AI Engineer

Dry Ground • Philippines

On-site
PHP 2,995,000 - 4,794,000
Competitive salary and performance-based incentives
Flexible work environment
Collaborative innovation-driven culture
Production AI Systems Engineer
Production AI Systems Engineer

LangChain • Boston

On-site
PHP 9,231,000 - 15,385,000
Principal AI Observability & Evals Platform Engineer
Principal AI Observability & Evals Platform Engineer

LangChain • Boston

On-site
PHP 14,154,000 - 16,615,000
Principal Software Engineer, AI Observability & Evals Platform
Principal Software Engineer, AI Observability & Evals Platform

LangChain • Boston

On-site
PHP 14,154,000 - 16,615,000
Medical coverage
Dental coverage
Vision coverage
+2
Staff Software Engineer
Staff Software Engineer

Socket.dev • Hinoba-an

On-site
PHP 1,800,000 - 2,800,000
GenAI Engineer - Hybrid Role with Growth & Benefits
GenAI Engineer - Hybrid Role with Growth & Benefits

Thinking Machines Data Science • Philippines

Hybrid
PHP 800,000 - 1,200,000
Hybrid Set-Up
Health benefits
Professional development budget
+1
PH - GenAI Engineer
PH - GenAI Engineer

Thinking Machines Data Science • Taguig

Hybrid
Competitive salary
Hybrid work model
Health benefits
+2