Senior AI Platform & AgentOps Engineer

Tredence Inc.

Bengaluru

Hybrid

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tredence Inc. is seeking a Senior AI Platform & AgentOps Engineer in Bangalore to own deployment architectures and the operations of AI agents, workflow engines, and AI runtime services.

You will lead Kubernetes-based platforms, cloud infrastructure automation, and platform reliability across cloud and enterprise environments. The role requires 5–8 years of hands-on experience in cloud/platform engineering, Kubernetes, Terraform, and GitOps, with a focus on secure, scalable, and observable

Qualifications

  • 5–8 years of hands-on cloud/platform engineering experience.
  • Proven Kubernetes production ops experience.
  • Experience with Terraform and GitOps tooling.

Responsibilities

  • Own deployment and operational architecture for AI Agents, Agent Servers, Workflow Engines, RAG Platforms, Model Gateways, and AI Runtime Services.
  • Design, build, and operate Kubernetes-based platforms for large-scale AI workloads.
  • Lead cloud infrastructure engineering, environment provisioning, deployment automation, release management, and platform reliability initiatives.
  • Establish CI/CD, GitOps, Infrastructure-as-Code, observability, security, disaster recovery, and operational standards across the platform.
  • Drive production readiness reviews, scalability planning, performance optimization, capacity management, and incident response processes.
  • Collaborate with AI and Platform Engineering teams to ensure cloud-native, secure, observable, enterprise-ready applications.
  • Support customer deployments across cloud, hybrid-cloud, and enterprise-managed environments.

Skills

Cloud & Platform Engineering
Kubernetes Platform Ops
Observability & Reliability
Platform Security & IAM
CI/CD & GitOps
Distributed Systems
Platform Architecture

Tools

GKE
Cloud Run
BigQuery
Cloud Storage
Pub/Sub
IAM
VPC Networking
Cloud Monitoring
Secret Manager
Cloud Functions
API Gateway
Terraform
GitHub Actions
GitLab CI/CD
ArgoCD
FluxCD
Kubernetes
Docker
Helm

Job description

Role Description

Senior AI Platform & AgentOps Engineer

Location:Bangalore, India

Experience:5–8 Years

Type:Full-Time

4 days to Office

Role Overview

We are building a next-generation Enterprise AI Platform powering AI Agents, Multi-Agent Systems, AI Workflows, RAG Applications, Knowledge Systems, and Enterprise Integrations. We are looking for a highly capable Platform & AgentOps Engineer to lead the deployment, operations, scalability, security, and reliability of our AI platform across cloud and customer environments.

This is a hands-on engineering role for builders who enjoy owning production systems end-to-end. You will be responsible for designing deployment architectures, operating Kubernetes platforms, automating cloud infrastructure, establishing operational excellence, and ensuring that AI applications can be deployed, monitored, governed, and scaled reliably in enterprise environments.

Key Responsibilities
  • Own the deployment and operational architecture for AI Agents, Agent Servers, Workflow Engines, RAG Platforms, Model Gateways, and AI Runtime Services.
  • Design, build, and operate Kubernetes-based platforms supporting large-scale AI workloads.
  • Lead cloud infrastructure engineering, environment provisioning, deployment automation, release management, and platform reliability initiatives.
  • Establish CI/CD, GitOps, Infrastructure-as-Code, observability, security, disaster recovery, and operational standards across the platform.
  • Drive production readiness reviews, scalability planning, performance optimization, capacity management, and incident response processes.
  • Work closely with AI and Platform Engineering teams to ensure applications are cloud-native, secure, observable, and enterprise-ready.
  • Support customer deployments across cloud, hybrid-cloud, and enterprise-managed environments.
Mandatory Skills & Experience

Cloud & Platform Engineering

  • Strong hands-on Google Cloud Platform (GCP) experience is mandatory.
  • Deep expertise with services such as GKE, Cloud Run, BigQuery, Cloud Storage, Pub/Sub, IAM, VPC Networking, Cloud Monitoring, Secret Manager, Cloud Functions, and API Gateway.
  • Strong experience with Azure and/or AWS is highly desirable.
  • Experience designing and operating secure, scalable, multi-environment cloud architectures.
Kubernetes & Container Platforms
  • Deep production experience with Kubernetes (GKE preferred).
  • Strong expertise in Docker, Helm, Ingress Controllers, Autoscaling, Stateful Workloads, Service Discovery, Networking, Storage, and Secrets Management.
  • Experience operating multi-cluster and multi-environment Kubernetes deployments.
  • Strong troubleshooting capabilities across Kubernetes, networking, infrastructure, and application layers.
Infrastructure as Code & Deployment Automation
  • Strong Terraform expertise is mandatory.
  • Experience building reusable infrastructure modules and enterprise-scale Infrastructure-as-Code frameworks.
  • Hands-on experience with GitHub Actions, GitLab CI/CD, Azure DevOps, ArgoCD, FluxCD, and GitOps deployment models.
  • Experience implementing automated provisioning, release automation, environment management, and deployment governance.
AgentOps & AI Infrastructure
  • Experience deploying and operating AI applications in production.
  • Understanding of Agent Servers, AI Workflow Engines, Model Gateways, RAG Services, Vector Databases, AI Runtime Infrastructure, and LLM-based applications.
  • Experience supporting OpenAI, Azure OpenAI, Claude, Gemini, Ollama, vLLM, TGI, or similar AI serving environments.
  • Understanding of AI workload scaling, model routing, latency optimization, throughput planning, token consumption monitoring, and operational governance.
Observability, Reliability & Operations
  • Strong experience with OpenTelemetry, Prometheus, Grafana, ELK, Datadog, Cloud Monitoring, or Application Insights.
  • Deep understanding of monitoring, distributed tracing, logging, ing, SRE practices, reliability engineering, and operational excellence.
  • Experience leading incident management, root-cause analysis, disaster recovery planning, backup strategies, and production support.
Distributed Systems & Messaging
  • Strong experience with Redis, Redis Pub/Sub, Redis Streams, Kafka, RabbitMQ, Google Pub/Sub, Azure Service Bus, or AWS SQS/SNS.
  • Understanding of event-driven architectures, distributed systems, fault tolerance, high availability, and scalability patterns.
Security & Enterprise Deployments
  • Experience implementing RBAC, IAM, SSO, SAML, OIDC, Secret Management, Network Security, Compliance Controls, and Secure Software Delivery practices.
  • Experience deploying applications into enterprise, regulated, and customer-managed environments.
Highly Desirable
  • MCP (Model Context Protocol) infrastructure and integrations.
  • Self-hosted LLM platforms using vLLM, Ray Serve, Ollama, TGI, or Kubernetes-based inference infrastructure.
  • Vector databases such as Pinecone, Qdrant, Weaviate, Chroma, pgvector, and Vertex AI Search.
  • Service Mesh technologies including Istio and Linkerd.
  • Airflow, Temporal, Dagster, Prefect, or workflow orchestration platforms.
  • Platform Engineering, Internal Developer Platforms (IDP), Developer Experience tooling, and FinOps.
Who We Are Looking For
  • Engineers who can independently own mission-critical production platforms.
  • Individuals with strong architectural thinking and exceptional troubleshooting abilities.
  • Builders who thrive in fast-moving environments and enjoy solving infrastructure, reliability, scalability, and deployment challenges.
  • Professionals capable of evolving into Platform Leads, Cloud Architects, or Infrastructure Architects.

This is not a traditional DevOps role. We are looking for a Platform Engineer who can own the operational backbone of an enterprise AI platform, establish deployment excellence, and build the foundation on which large-scale Agentic AI systems run reliably in production.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform & AgentOps Engineer
Senior AI Platform & AgentOps Engineer

Tredence • Bengaluru

On-site
INR 400,000 - 700,000
Devops Engineer
Devops Engineer

Tredence • Bengaluru

On-site
INR 4,000,000 - 7,000,000
AI Engineer
AI Engineer

Tredence • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Software AI Engineer
Software AI Engineer

Tredence • Bengaluru

On-site
INR 2,500,000 - 4,500,000
Software Engineer
Software Engineer

Tredence • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Sr Software AI Engineer
Sr Software AI Engineer

Tredence Inc. • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Agentic AI Product developer
Agentic AI Product developer

Tredence • Bengaluru

On-site
INR 3,000,000 - 5,500,000
AI Platform Architect
AI Platform Architect

CLOUDSUFI • Dadri

On-site
INR 4,200,000 - 7,000,000
AgenticOps Platform Engineer Lead
AgenticOps Platform Engineer Lead

Bridge AI • India

On-site
INR 4,000,000 - 7,000,000
Senior Agentic AI Platform Engineer – MCP & Agent Runtime
Senior Agentic AI Platform Engineer – MCP & Agent Runtime

ITOrizon • Bengaluru

On-site
INR 1,200,000 - 1,800,000