AgenticOps Platform Engineer Lead

Bridge AI

India

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

BridgeAI is seeking a senior, hands-on AgentOps Platform Engineer to design, build, and operate the cloud-native infrastructure powering our AI agents at scale. You will own Terraform-based IaC, lead by example, and sit at the crossroads of DevOps, MLOps, and AgentOps across GCP and multi-cloud.

In this lead role you’ll optimize resources, build CI/CD for infra and agent services, and drive reliability, security, and observability with dashboards and traces.

Qualifications

  • 5+ years in DevOps, Platform Engineering, SRE, or MLOps.
  • Hands-on experience with GCP and multi-cloud design.
  • Terraform is mandatory.
  • Experience with CI/CD and automation.
  • Security, observability, and reliability focus.

Responsibilities

  • Design, build, and operate production-grade infrastructure for AI agents and LLM services.
  • Own Terraform-based IaC for all environments.
  • Lead infrastructure decisions by hands-on implementation.
  • Build scalable foundations for Agent orchestration, Inference services, RAG pipelines.
  • Design agent runtime environments with isolation, failover, and controlled rollout.
  • Develop CI/CD pipelines for infrastructure and agent services.
  • Automate workflows for vector DB updates and model/version rollouts.
  • Create dashboards and telemetry using Prometheus, Grafana, OpenTelemetry.
  • Ensure cloud security, IAM best practices, and operational safety.
  • Collaborate with AI engineers, product, and CTO office; set SLOs and reliability targets.

Skills

DevOps
SRE
MLOps
Cloud architecture
Automation

Tools

GCP
Terraform
Kubernetes
GitHub Actions
Prometheus Grafana

Job description

We are looking for a senior, hands‑on AgentOps Platform Engineer to design, build, and operate the cloud‑native infrastructure that powers our AI agents at scale.

This is a lead‑by‑example role:

  • You write the Terraform
  • You build the pipelines
  • You own the platform in production

GCP is your primary environment, but you will design with multi‑cloud in mind (AWS, Azure), ensuring portability, resilience, and long‑term flexibility. This role sits at the intersection of DevOps, MLOps, and AgentOps, with deep responsibility for reliability, security, observability, and cost.

KEY RESPONSIBILITIES
Platform & Infrastructure Ownership
  • Design, build, and operate production‑grade infrastructure for AI agents and LLM services
  • Own Terraform‑based Infrastructure as Code for all environments (dev, uat, prod)
  • Lead infrastructure decisions through hands‑on implementation, not diagrams
  • Build scalable foundations for: Agent orchestration, Inference services, RAG pipelines, Vector stores
  • Optimise cloud resources for performance and cost efficiency
AgentOps & AI Platform Enablement
  • Enable safe, continuous operation of autonomous agents
  • Design agent runtime environments with: Isolation & sandboxing, Failover and recovery strategies, Controlled rollout mechanisms
  • Support prompt/version management, agent configuration, and tool/plugin lifecycle
  • Work closely with Agentic RAG engineers to operationalise research into production
CI/CD & Automation
  • Build and maintain CI/CD pipelines for: Infrastructure, Agent services, Prompt and config changes, Model/version rollouts
  • Automate workflows for: Vector DB updates, RAG index refreshes, Agent memory stores, Tool registration and validation
  • Reduce manual ops toil aggressively through automation
Observability & Production Readiness
  • Design and implement deep observability for agent systems: Platform health, Agent execution metrics, Latency, cost, and throughput, Failure modes and retries
  • Build dashboards, alerts, and telemetry using: Prometheus, Grafana, OpenTelemetry (or equivalent)
  • Enable visibility into agent decision traces and runtime behavior
Security, Safety & Reliability
  • Implement secure cloud architecture and IAM best practices
  • Own production reliability, incident response, and recovery
  • Enforce operational guardrails and safety controls for agent APIs
  • Support responsible AI practices from an infrastructure and runtime perspective
Collaboration & Technical Leadership
  • Work closely with: Agentic RAG engineers, AI engineers, Product & CTO Office
  • Define SLOs, reliability targets, and operational metrics
  • Set the technical bar for AgentOps at BridgeAI
  • Mentor engineers by example and code, not process overhead
REQUIRED SKILLS & EXPERIENCE
Core Platform & DevOps
  • 5+ years in DevOps, Platform Engineering, SRE, or MLOps
  • Strong, hands‑on experience with GCP: GKE / Compute Engine, Cloud Run / Functions, Cloud Storage, Pub/Sub, Vertex AI (or equivalent)
  • Deep experience with Terraform (mandatory)
Containers, CI/CD & Automation
  • CI/CD tooling (GitHub Actions, Jenkins, ArgoCD)
  • Python and Bash for automation and platform glue code
Agentic & AI Systems
  • Experience supporting LLM‑based systems in production
  • Understanding of: Prompt/version management, Context handling & caching, Model rollout strategies
  • Hands‑on experience with vector databases (Weaviate, FAISS, Pinecone)
  • Familiarity with RAG pipelines and agent execution patterns
Observability & Security
  • Monitoring and telemetry using Prometheus, Grafana, OpenTelemetry
  • Strong understanding of cloud security, IAM, and operational safety
NICE TO HAVE
  • Multi‑cloud experience (AWS, Azure)
  • Exposure to agent frameworks (LangChain, LangGraph, AutoGen, CrewAI)
  • Experience with responsible AI operations or safety monitoring
WHAT SUCCESS LOOKS LIKE
  • Infrastructure is reproducible, observable, and boring (in a good way)
  • Agent failures are visible, debuggable, and recoverable
  • Cloud costs are understood and controlled
  • Engineers trust the platform and move faster because of it
  • You are the go‑to authority for AgentOps at BridgeAI
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform & AgentOps Engineer
Senior AI Platform & AgentOps Engineer

Tredence • Bengaluru

On-site
INR 400,000 - 700,000
Senior AI Platform & AgentOps Engineer
Senior AI Platform & AgentOps Engineer

Tredence Inc. • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Devops Engineer
Devops Engineer

Tredence • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Agentic AI Product developer
Agentic AI Product developer

Tredence • Bengaluru

On-site
INR 3,000,000 - 5,500,000
AI Engineer
AI Engineer

Tredence • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Software AI Engineer
Software AI Engineer

Tredence • Bengaluru

On-site
INR 2,500,000 - 4,500,000
Software Engineer
Software Engineer

Tredence • Bengaluru

On-site
INR 2,500,000 - 4,200,000
AI Solutions and Platforms Operations Engineer
AI Solutions and Platforms Operations Engineer

PepsiCo • Hyderabad

On-site
INR 1,000,000 - 2,000,000
Lead Platform Engineer
Lead Platform Engineer

EPAM Systems • India

On-site
INR 3,000,000 - 5,500,000
AI Platform Lead
AI Platform Lead

Virtualyyst • Mumbai

On-site
INR 4,000,000 - 8,000,000