Agentic AI developer

Tredence

Bengaluru

On-site

INR 1,800,000 - 3,200,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Tredence is seeking a Platform & AgentOps Engineer to own deployment, operations, and reliability of our AI platform across cloud and on-prem environments. You will design deployment architectures, run Kubernetes platforms, and automate cloud infrastructure for enterprise-grade AI workloads.

You will lead CI/CD, GitOps, IaC, and observability efforts, while ensuring secure, scalable, and observable deployments that meet enterprise requirements.

Qualifications

  • Hands-on GCP experience and strong multi-environment cloud architectures.
  • Expertise with GKE, Cloud Run, BigQuery, and IAM.
  • Proficiency in Terraform and GitOps tooling.
  • Experience deploying secure, scalable platforms for enterprise workloads.
  • Ability to design CI/CD and observability pipelines for AI workloads.
  • Knowledge of security, reliability, and incident response in production.

Responsibilities

  • Own deployment and operation architecture for AI Agents, Agents Servers, and runtimes.
  • Build and operate Kubernetes-based platforms for large-scale AI workloads.
  • Lead cloud infrastructure engineering and release management.
  • Establish CI/CD, GitOps, and IaC standards across the platform.
  • Drive production readiness, scalability planning, and incident response.
  • Collaborate with AI and Platform teams to ensure cloud-native, secure deployments.

Skills

GCP
GKE
Cloud Run
BigQuery
Cloud Storage
IAM
VPC Networking
Cloud Monitoring
Secret Manager
Cloud Functions
API Gateway
Azure/AWS
Kubernetes
Docker
Helm
Ingress Controllers
Autoscaling
Stateful Workloads
Service Discovery
Networking
Secrets Management
Terraform
GitHub Actions
GitLab CI/CD
ArgoCD
FluxCD
GitOps
OpenTelemetry
Prometheus
Grafana
ELK
Datadog
Incident Management
Disaster Recovery
RBAC
AI Infrastructure

Tools

Terraform
GitHub Actions
GitLab CI/CD
ArgoCD
FluxCD
Kubernetes
Docker
Helm
Ingress Controllers
Istio
Linkerd
Airflow
Temporal
Dagster
Prefect
OpenTelemetry
Prometheus
Grafana
Terraform Modules

Job description

Role & responsibilities
Role & responsibilities
Role Overview

We are building a next-generation Enterprise AI Platform powering AI Agents, Multi-Agent Systems, AI Workflows, RAG Applications, Knowledge Systems, and Enterprise Integrations. We are looking for a highly capable Platform & AgentOps Engineer to lead the deployment, operations, scalability, security, and reliability of our AI platform across cloud and customer environments.

This is a hands-on engineering role for builders who enjoy owning production systems end-to-end. You will be responsible for designing deployment architectures, operating Kubernetes platforms, automating cloud infrastructure, establishing operational excellence, and ensuring that AI applications can be deployed, monitored, governed, and scaled reliably in enterprise environments.

Key Responsibilities
  • Own the deployment and operational architecture for AI Agents, Agent Servers, Workflow Engines, RAG Platforms, Model Gateways, and AI Runtime Services.
  • Design, build, and operate Kubernetes-based platforms supporting large-scale AI workloads.
  • Lead cloud infrastructure engineering, environment provisioning, deployment automation, release management, and platform reliability initiatives.
  • Establish CI/CD, GitOps, Infrastructure-as-Code, observability, security, disaster recovery, and operational standards across the platform.
  • Drive production readiness reviews, scalability planning, performance optimization, capacity management, and incident response processes.
  • Work closely with AI and Platform Engineering teams to ensure applications are cloud-native, secure, observable, and enterprise-ready.
  • Support customer deployments across cloud, hybrid-cloud, and enterprise-managed environments.
Mandatory Skills & Experience
Cloud & Platform Engineering
  • Strong hands-on Google Cloud Platform (GCP) experience is mandatory.
  • Deep expertise with services such as GKE, Cloud Run, BigQuery, Cloud Storage, Pub/Sub, IAM, VPC Networking, Cloud Monitoring, Secret Manager, Cloud Functions, and API Gateway.
  • Strong experience with Azure and/or AWS is highly desirable.
  • Experience designing and operating secure, scalable, multi-environment cloud architectures.
Kubernetes & Container Platforms
  • Deep production experience with Kubernetes (GKE preferred).
  • Strong expertise in Docker, Helm, Ingress Controllers, Autoscaling, Stateful Workloads, Service Discovery, Networking, Storage, and Secrets Management.
  • Experience operating multi-cluster and multi-environment Kubernetes deployments.
  • Strong troubleshooting capabilities across Kubernetes, networking, infrastructure, and application layers.
Infrastructure as Code & Deployment Automation
  • Strong Terraform expertise is mandatory.
  • Experience building reusable infrastructure modules and enterprise-scale Infrastructure-as-Code frameworks.
  • Hands-on experience with GitHub Actions, GitLab CI/CD, Azure DevOps, ArgoCD, FluxCD, and GitOps deployment models.
  • Experience implementing automated provisioning, release automation, environment management, and deployment governance.
AgentOps & AI Infrastructure
  • Experience deploying and operating AI applications in production.
  • Understanding of Agent Servers, AI Workflow Engines, Model Gateways, RAG Services, Vector Databases, AI Runtime Infrastructure, and LLM-based applications.
  • Experience supporting OpenAI, Azure OpenAI, Claude, Gemini, Ollama, vLLM, TGI, or similar AI serving environments.
  • Understanding of AI workload scaling, model routing, latency optimization, throughput planning, token consumption monitoring, and operational governance.
Observability, Reliability & Operations
  • Strong experience with OpenTelemetry, Prometheus, Grafana, ELK, Datadog, Cloud Monitoring, or Application Insights.
  • Deep understanding of monitoring, distributed tracing, logging, ing, SRE practices, reliability engineering, and operational excellence.
  • Experience leading incident management, root-cause analysis, disaster recovery planning, backup strategies, and production support.
Distributed Systems & Messaging
  • Strong experience with Redis, Redis Pub/Sub, Redis Streams, Kafka, RabbitMQ, Google Pub/Sub, Azure Service Bus, or AWS SQS/SNS.
  • Understanding of event-driven architectures, distributed systems, fault tolerance, high availability, and scalability patterns.
Security & Enterprise Deployments
  • Experience implementing RBAC, IAM, SSO, SAML, OIDC, Secret Management, Network Security, Compliance Controls, and Secure Software Delivery practices.
  • Experience deploying applications into enterprise, regulated, and customer-managed environments.
Highly Desirable
  • MCP (Model Context Protocol) infrastructure and integrations.
  • Self-hosted LLM platforms using vLLM, Ray Serve, Ollama, TGI, or Kubernetes-based inference infrastructure.
  • Vector databases such as Pinecone, Qdrant, Weaviate, Chroma, pgvector, and Vertex AI Search.
  • Service Mesh technologies including Istio and Linkerd.
  • Airflow, Temporal, Dagster, Prefect, or workflow orchestration platforms.
  • Platform Engineering, Internal Developer Platforms (IDP), Developer Experience tooling, and FinOps.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Agentic AI Product developer
Agentic AI Product developer

Tredence • Bengaluru

On-site
INR 3,000,000 - 5,500,000
Software Engineer
Software Engineer

Tredence • Bengaluru

On-site
INR 2,500,000 - 4,200,000
AI Orchestration / Platform Engineer
AI Orchestration / Platform Engineer

NTT DATA BUSINESS SOLUTIONS • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Agentic AI Engineer
Agentic AI Engineer

EY • Hyderabad, Bengaluru, Delhi

On-site
INR 6,000,000 - 9,000,000
AgenticOps Platform Engineer Lead
AgenticOps Platform Engineer Lead

Bridge AI • India

On-site
INR 4,000,000 - 7,000,000
Lead Software Engineer Agentic AI Platform & AI Research
Lead Software Engineer Agentic AI Platform & AI Research

WNS Holdings • Bengaluru

On-site
INR 2,600,000 - 3,800,000
Agentic AI Engineer
Agentic AI Engineer

Deloitte Shared Services India • Gurugram District, Delhi, Mumbai

On-site
INR 1,800,000 - 3,200,000
Principal AI Solutions Architect
Principal AI Solutions Architect

Metatron Hr Solutions Coimbatore • Chennai District

On-site
INR 3,600,000 - 6,000,000
Senior AI Engineer - Agentic AI & Knowledge Systems
Senior AI Engineer - Agentic AI & Knowledge Systems

Ontio AI • Pune District

On-site
INR 3,000,000 - 5,000,000
Lead AI Platform Engineer
Lead AI Platform Engineer

Channel Fusion • India

On-site
INR 2,400,000 - 4,200,000