Software Engineer

Tredence

Bengaluru

On-site

INR 2,500,000 - 4,200,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tredence in Bangalore, India is seeking an AI Engineering Lead for AI Platform & Agent Systems. You will drive end-to-end design, architecture, and productionization of AI infra and applications at scale across cloud and on-prem environments.

You will lead AI Platform & AgentOps teams, defining best practices, orchestrating LLM-based workloads, and delivering production-grade AI solutions with reliability, security, and cost-awareness.

Qualifications

  • Strong experience building and deploying LLM-based applications and AI platforms.
  • Deep understanding of RAG architectures, vector databases, and AI inference pipelines.
  • Experience with AI model APIs and/or self-hosted inference.
  • Knowledge of agent-based systems, workflow engines, and model orchestration patterns.
  • Hands-on experience with Google Cloud Platform (GCP) services.

Responsibilities

  • Architect AI platforms for agents, workflows, and RAG systems.
  • Provide technical leadership and mentorship to the AI infra team.
  • Collaborate with AI Research, Product, and Platform teams to deliver production-grade solutions.
  • Define engineering standards and patterns for the team.
  • Own CI/CD and GitOps deployment models for AI systems.
  • Ensure security, reliability, and enterprise readiness.

Skills

LLM-based apps
RAG architectures
Vector databases
AI inference pipelines
AgentOps
GKE
GCP
CI/CD
GitOps
Terraform
Kubernetes
Monitoring
SRE principles
Security & IAM
Cost optimization

Tools

GKE
Cloud Run
BigQuery
Pub/Sub
Terraform
ArgoCD
Istio
Temporal
Airflow
Dagster

Job description

AI Engineering Lead AI Platform & Agent Systems

Location: Bangalore, India

Experience: 2 to 8 Years

Type: Full-Time

Role Overview

We are building a next-generation Enterprise AI Platform powering AI Agents, Multi-Agent Systems, RAG Applications, AI Workflows, and Knowledge Platforms.

We are looking for an AI Engineering Lead who will drive the design, architecture, and productionization of AI systems at scale. This is a hands‑on leadership role responsible for owning the end‑to‑end lifecycle of AI infrastructure and applications—from design and development to deployment, operations, and continuous optimization.

You will lead the AgentOps and AI Platform engineering efforts, ensuring that AI systems are scalable, reliable, secure, and enterprise‑ready across cloud and hybrid environments.

Key Responsibilities
AI Platform & Architecture Leadership
  • Lead the architecture and development of AI platforms supporting agents, workflows, RAG systems, and LLM-based applications.
  • Define best practices for AI system design, model orchestration, inference pipelines, and runtime infrastructure.
  • Drive the evolution of AgentOps frameworks for managing AI agents at scale.
Engineering Leadership
  • Provide technical leadership and mentorship to a team of engineers working on AI infrastructure and platform systems.
  • Collaborate cross‑functionally with AI Research, Product, and Platform teams to deliver production‑grade AI solutions.
  • Establish engineering standards, design patterns, and development practices.
AI Systems & AgentOps
  • Design and manage AI agent architectures, workflow orchestration, and multi‑agent systems.
  • Build and operate Model Gateways and LLM routing layers across providers (OpenAI, Azure OpenAI, Claude, Gemini, etc.).
  • Lead development of RAG systems with vector databases and retrieval pipelines.
  • Optimize latency, throughput, and cost of AI workloads.
Platform & Infrastructure Engineering
  • Own Kubernetes‑based platforms for scalable AI workloads (GKE preferred).
  • Design and implement cloud‑native architectures across GCP (primary), AWS/Azure (secondary).
  • Lead Infrastructure‑as‑Code (Terraform) and platform automation initiatives.
  • Establish CI/CD and GitOps‑based deployment models for AI systems.
Reliability, Observability & Operations
  • Define and implement SRE practices including monitoring, ing, and incident management.
  • Architect observability using OpenTelemetry, Prometheus, Grafana, ELK, or Cloud Monitoring.
  • Drive production readiness, scalability planning, and disaster recovery strategies.
Enterprise Deployments & Security
  • Ensure AI platform compliance with enterprise‑grade security and governance standards.
  • Implement IAM, RBAC, SSO, secrets management, and network security controls.
  • Support customer deployments across cloud, hybrid, and on‑prem environments.
Mandatory Skills & Experience
AI Engineering & Systems
  • Strong experience building and deploying LLM-based applications and AI platforms.
  • Deep understanding of RAG architectures, vector databases, and AI inference pipelines.
  • Experience with AI model APIs and/or self‑hosted inference (vLLM, TGI, Ollama, etc.).
  • Knowledge of agent‑based systems, workflow engines, and model orchestration patterns.
Cloud & Platform Engineering
  • Strong hands‑on experience with Google Cloud Platform (GCP).
  • Expertise in services such as GKE, Cloud Run, BigQuery, Pub/Sub, IAM, VPC, and Monitoring.
  • Experience designing scalable, secure, multi‑environment cloud architectures.
Kubernetes & Distributed Systems
  • Deep production experience with Kubernetes (GKE preferred).
  • Strong understanding of containerization, autoscaling, networking, storage, and high availability.
  • Experience with distributed systems and event‑driven architectures.
Infrastructure & DevOps
  • Strong expertise in Terraform and Infrastructure‑as‑Code frameworks.
  • Experience with CI/CD, GitOps (ArgoCD, Flux), and release automation.
  • Proven ability to build automated, scalable platform infrastructure.
Observability & Reliability
  • Experience with monitoring, logging, and tracing systems (Prometheus, Grafana, OTEL, etc.).
  • Strong understanding of SRE principles, incident management, and reliability engineering.
Leadership & Ownership
  • Proven experience leading complex engineering initiatives or mentoring teams.
  • Ability to own end‑to‑end system delivery and operational excellence.
Highly Desirable
  • Experience with self-hosted LLM infrastructure (vLLM, Ray Serve, TGI).
  • Hands‑on with vector databases (Pinecone, Qdrant, Weaviate, pgvector).
  • Knowledge of Service Mesh (Istio, Linkerd).
  • Experience with workflow orchestration tools (Temporal, Airflow, Dagster).
  • Exposure to Platform Engineering / Internal Developer Platforms (IDP).
  • Understanding of FinOps and cost optimization for AI workloads.
Required Skills
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software AI Engineer
Software AI Engineer

Tredence • Bengaluru

On-site
INR 2,500,000 - 4,500,000
AI Engineer
AI Engineer

Tredence • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior AI Platform & AgentOps Engineer
Senior AI Platform & AgentOps Engineer

Tredence • Bengaluru

On-site
INR 400,000 - 700,000
Devops Engineer
Devops Engineer

Tredence • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior AI Platform & AgentOps Engineer
Senior AI Platform & AgentOps Engineer

Tredence Inc. • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Lead Engineer AI Platform
Lead Engineer AI Platform

Tutor Cloud Pvt Ltd • Bengaluru

On-site
INR 3,000,000 - 5,200,000
Agentic AI Product developer
Agentic AI Product developer

Tredence • Bengaluru

On-site
INR 3,000,000 - 5,500,000
Senior AI platform Engineer
Senior AI platform Engineer

Qloron Pvt Ltd • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Lead AI Engineer
Lead AI Engineer

Impetus Career Consultants • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Lead AI Engineer
Lead AI Engineer

Srijan Technologies PVT LTD • Gurugram District

On-site
INR 1,500,000 - 2,500,000