Devops Engineer

Tredence

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tredence is seeking a Senior AI Platform & AgentOps Engineer to own production platforms for AI agents, workflows, and runtimes in a cloud-native, enterprise context.

You will design deployment architectures, operate GKE-based platforms, implement CI/CD pipelines, and scale AI workloads across multi-cloud environments. Collaboration with AI and Platform teams is essential to deliver secure, observable, and reliable services.

Qualifications

  • Strong hands-on Google Cloud Platform experience is mandatory.
  • Deep expertise with GKE, Cloud Run, BigQuery, Cloud Storage, Pub/Sub, IAM, VPC Networking, Cloud Monitoring, Secret Manager, Cloud Functions, and API Gateway.
  • Strong experience with Azure and/or AWS is highly desirable.
  • Experience designing and operating secure, scalable multi-environment cloud architectures.

Responsibilities

  • Own the deployment and operational architecture for AI Agents, Agent Servers, Workflow Engines, RAG Platforms, Model Gateways, and AI Runtime Services.
  • Design, build, and operate Kubernetes-based platforms supporting large-scale AI workloads.
  • Lead cloud infrastructure engineering, environment provisioning, deployment automation, release management, and platform reliability initiatives.
  • Establish CI/CD, GitOps, Infrastructure-as-Code, observability, security, disaster recovery, and operational standards across the platform.
  • Drive production readiness reviews, scalability planning, performance optimization, capacity management, and incident response processes.
  • Work closely with AI and Platform Engineering teams to ensure applications are cloud-native, secure, observable, and enterprise-ready.

Skills

GCP experience
Kubernetes
CI/CD
Terraform
GitOps
Observability
Security
Incident response

Tools

GKE
Docker
Helm
Ingress Controllers
Autoscaling
Stateful Workloads
Service Discovery
Networking
Storage
Secrets Management
ArgoCD
FluxCD

Job description

Senior AI Platform & AgentOps Engineer

Location: Bangalore, India

Experience: 58 Years
Type: Full-Time

4 days to Office

Role Overview

We are building a next-generation Enterprise AI Platform powering AI Agents, Multi-Agent Systems, AI Workflows, RAG Applications, Knowledge Systems, and Enterprise Integrations. We are looking for a highly capable Platform & AgentOps Engineer to lead the deployment, operations, scalability, security, and reliability of our AI platform across cloud and customer environments.

This is a hands-on engineering role for builders who enjoy owning production systems end-to-end. You will be responsible for designing deployment architectures, operating Kubernetes platforms, automating cloud infrastructure, establishing operational excellence, and ensuring that AI applications can be deployed, monitored, governed, and scaled reliably in enterprise environments.

Key Responsibilities
  • Own the deployment and operational architecture for AI Agents, Agent Servers, Workflow Engines, RAG Platforms, Model Gateways, and AI Runtime Services.
  • Design, build, and operate Kubernetes-based platforms supporting large-scale AI workloads.
  • Lead cloud infrastructure engineering, environment provisioning, deployment automation, release management, and platform reliability initiatives.
  • Establish CI/CD, GitOps, Infrastructure-as-Code, observability, security, disaster recovery, and operational standards across the platform.
  • Drive production readiness reviews, scalability planning, performance optimization, capacity management, and incident response processes.
  • Work closely with AI and Platform Engineering teams to ensure applications are cloud-native, secure, observable, and enterprise-ready.
  • Support customer deployments across cloud, hybrid-cloud, and enterprise-managed environments.
Mandatory Skills & Experience
Cloud & Platform Engineering
  • Strong hands-on Google Cloud Platform (GCP) experience is mandatory.
  • Deep expertise with services such as GKE, Cloud Run, BigQuery, Cloud Storage, Pub/Sub, IAM, VPC Networking, Cloud Monitoring, Secret Manager, Cloud Functions, and API Gateway.
  • Strong experience with Azure and/or AWS is highly desirable.
  • Experience designing and operating secure, scalable, multi-environment cloud architectures.
Kubernetes & Container Platforms
  • Deep production experience with Kubernetes (GKE preferred).
  • Strong expertise in Docker, Helm, Ingress Controllers, Autoscaling, Stateful Workloads, Service Discovery, Networking, Storage, and Secrets Management.
  • Experience operating multi-cluster and multi-environment Kubernetes deployments.
  • Strong troubleshooting capabilities across Kubernetes, networking, infrastructure, and application layers.
Infrastructure as Code & Deployment Automation
  • Strong Terraform expertise is mandatory.
  • Experience building reusable infrastructure modules and enterprise-scale Infrastructure-as-Code frameworks.
  • Hands-on experience with GitHub Actions, GitLab CI/CD, Azure DevOps, ArgoCD, FluxCD, and GitOps deployment models.
  • Experience implementing automated provisioning, release automation, environment management, and deployment governance.
AgentOps & AI Infrastructure
  • Experience deploying and operating AI applications in production.
  • Understanding of Agent Servers, AI Workflow Engines, Model Gateways, RAG Services, Vector Databases, AI Runtime Infrastructure, and LLM-based applications.
  • Experience supporting OpenAI, Azure OpenAI, Claude, Gemini, Ollama, vLLM, TGI, or similar AI serving environments.
  • Understanding of AI workload scaling, model routing, latency optimization, throughput planning, token consumption monitoring, and operational governance.
Observability, Reliability & Operations
  • Strong experience with OpenTelemetry, Prometheus, Grafana, ELK, Datadog, Cloud Monitoring, or Application Insights.
  • Deep understanding of monitoring, distributed tracing, logging, ing, SRE practices, reliability engineering, and operational excellence.
  • Experience leading incident management, root-cause analysis, disaster recovery planning, backup strategies, and production support.
Distributed Systems & Messaging
  • Strong experience with Redis, Redis Pub/Sub, Redis Streams, Kafka, RabbitMQ, Google Pub/Sub, Azure Service Bus, or AWS SQS/SNS.
  • Understanding of event-driven architectures, distributed systems, fault tolerance, high-availability, and scalability patterns.
Security & Enterprise Deployments
  • Experience implementing RBAC, IAM, SSO, SAML, OIDC, Secret Management, Network Security, Compliance Controls, and Secure Software Delivery practices.
  • Experience deploying applications into enterprise, regulated, and customer-managed environments.
Highly Desirable
  • MCP (Model Context Protocol) infrastructure and integrations.
  • Self-hosted LLM platforms using vLLM, Ray Serve, Ollama, TGI, or Kubernetes-based inference infrastructure.
  • Vector databases such as Pinecone, Qdrant, Weaviate, Chroma, pgvector, and Vertex AI Search.
  • Service Mesh technologies including Istio and Linkerd.
  • Airflow, Temporal, Dagster, Prefect, or workflow orchestration platforms.
  • Platform Engineering, Internal Developer Platforms (IDP), Developer Experience tooling, and FinOps.
Who We Are Looking For
  • Engineers who can independently own mission-critical production platforms.
  • Individuals with strong architectural thinking and exceptional troubleshooting abilities.
  • Builders who thrive in fast-moving environments and enjoy solving infrastructure, reliability, scalability, and deployment challenges.
  • Professionals capable of evolving into Platform Leads, Cloud Architects, or Infrastructure Architects.

This is not a traditional DevOps role. We are looking for a Platform Engineer who can own the operational backbone of an enterprise AI platform, establish deployment excellence, and build the foundation on which large-scale Agentic AI systems run reliably in production.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform & AgentOps Engineer
Senior AI Platform & AgentOps Engineer

Tredence Inc. • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Senior AI Platform & AgentOps Engineer
Senior AI Platform & AgentOps Engineer

Tredence • Bengaluru

On-site
INR 400,000 - 700,000
Software AI Engineer
Software AI Engineer

Tredence • Bengaluru

On-site
INR 2,500,000 - 4,500,000
Software Engineer
Software Engineer

Tredence • Bengaluru

On-site
INR 2,500,000 - 4,200,000
AI Engineer
AI Engineer

Tredence • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Agentic AI Product developer
Agentic AI Product developer

Tredence • Bengaluru

On-site
INR 3,000,000 - 5,500,000
Lead DevOps Engineer (5-8 Years)
Lead DevOps Engineer (5-8 Years)

airtel • Gurugram District

On-site
INR 1,500,000 - 2,000,000
Sr Software AI Engineer
Sr Software AI Engineer

Tredence Inc. • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
AgenticOps Platform Engineer Lead
AgenticOps Platform Engineer Lead

Bridge AI • India

On-site
INR 4,000,000 - 7,000,000
Senior Agentic AI Platform Engineer – MCP & Agent Runtime
Senior Agentic AI Platform Engineer – MCP & Agent Runtime

ITOrizon • Bengaluru

On-site
INR 1,200,000 - 1,800,000