Senior DevOps Engineer (Kubernetes & AI Infra)

Navikenz

Bengaluru

Hybrid

INR 1,500,000 - 2,500,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

A technology solutions company is searching for an experienced Senior DevOps Engineer in Bengaluru, India. The successful candidate will be responsible for designing and maintaining scalable cloud infrastructure while automating AI/ML workflows using Kubernetes and related tools. Candidates should have over 8 years of experience in DevOps, strong knowledge of Kubernetes fundamentals, and hands-on experience with cloud platforms like AWS or Azure. Strong problem-solving skills and scripting proficiency are essential for this role.

Qualifications

  • 8+ years of experience in DevOps, SRE, or Platform Engineering roles.
  • Strong knowledge of Kubernetes fundamentals, especially for AI workloads.
  • Hands-on experience with AWS, Azure, or GCP in Kubernetes-based environments.

Responsibilities

  • Design and maintain scalable cloud infrastructure using Kubernetes.
  • Automate AI/ML workflows with tools like Kubeflow and MLflow.
  • Monitor Kubernetes clusters and AI workloads for performance tuning.

Skills

Kubernetes fundamentals
Scripting with Python
AWS or Azure or GCP
CI/CD & MLOps tools
Problem-solving skills
Azure
GCP
MLOps

Tools

Kubeflow
Terraform
Prometheus
Grafana
Prometheus

Job description

We’re looking for an experienced Senior DevOps Engineer who loves working with Kubernetes and AI-driven applications. In this role, you’ll be responsible for designing, implementing, and maintaining scalable cloud infrastructure while supporting MLOps pipelines for AI workloads.

What You’ll Be Doing:
  • Building Scalable Infrastructure: You’ll design, implement, and maintain cloud infrastructure using Kubernetes to handle AI and non-AI workloads efficiently.
  • Developing CI/CD & MLOps Pipelines: Help us automate AI/ML workflows using tools like Kubeflow, MLflow, or Argo Workflows, ensuring seamless deployment and monitoring of AI models.
  • Optimizing AI Model Deployments: Work with ML engineers to fine-tune LLM models, AI-driven applications, and containerized environments for smooth operation.
  • Monitoring & Performance Tuning: Keep an eye on Kubernetes clusters and AI workloads, using tools like Prometheus, Grafana, and Loki to ensure high availability and performance.
  • Automating Everything: Whether it’s infrastructure provisioning (Terraform, Helm) or Kubernetes security best practices, you’ll help enforce efficiency and compliance.
Staying Ahead of the Curve:

You’ll have the opportunity to explore and implement emerging AI infrastructure trends, including KServe, Ray, and Triton Inference Server.

What We’re Looking For:
  • 8+ years of experience in DevOps, SRE, or Platform Engineering role, with expertise in Kubernetes and cloud-native DevOps
  • Strong knowledge of Kubernetes fundamentals (deployments, services, ingress, storage, GPU scheduling, multi-cluster management).
  • Proficiency in scripting & automation with Python, Bash, or Go, particularly for AI-related workflows.
  • Hands‑on experience with AWS, Azure, or GCP, especially in Kubernetes-based AI/ML infrastructure (e.g., Amazon SageMaker, GKE with AI, Azure ML).
  • Hands‑on experience with model deployment frameworks (NVIDIA Triton, vLLM. TGI etc.)
  • Experience with Distributed computing, multi‑GPU training on kubernetes and on‑prem GPU clusters.
  • Experience managing resource allocation and autoscaling for large training/inference workloads. (KEDA, HPA etc.)
  • Experience with CI/CD & MLOps tools such as Jenkins, Argo CD, Kubeflow, MLflow, or Tekton.
  • Familiarity with GenAI model deployment, including fine‑tuning, inference optimization, and A/B testing.
  • Hands‑on experience with managed ML services (AWS bedrock, Vertext AI models etc.)
  • Strong problem‑solving skills and a mindset of automating repetitive tasks.
  • Excellent communication skills to collaborate with ML engineers, data scientists, and software teams.
Bonus Points If You Have:
  • Experience with LLMOps (Large Language Model Operations) and deploying LLM‑based applications at scale.
  • Knowledge of Vector Databases (FAISS, Weaviate, Qdrant) for AI‑driven applications.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead DevOps AI Engineer
Lead DevOps AI Engineer

Trine Infotech • India

Remote
INR 2,500,000 - 6,000,000
AI_DevOps Engineer
AI_DevOps Engineer

RIA Advisory • Pune District

On-site
INR 1,000,000 - 1,500,000
AI / ML Ops Engineer
AI / ML Ops Engineer

Zoho • India

Remote
INR 1,800,000 - 2,800,000
Devops Engineer
Devops Engineer

Airtel • Gurugram District

On-site
INR 800,000 - 1,400,000
Senior Backend & MLOps Engineer
Senior Backend & MLOps Engineer

Patch Infotech Pvt Ltd • Bengaluru

On-site
INR 4,000,000 - 7,000,000
AI / ML Ops Engineer
AI / ML Ops Engineer

Zoho • Bengaluru South

On-site
INR 1,800,000 - 3,200,000
MLOps+DevOps Engineer
MLOps+DevOps Engineer

Teambees Corp • Pune District

On-site
INR 4,000,000 - 7,000,000
AI DevOps Engineer
AI DevOps Engineer

RIA Advisory LLC. • Pune District

On-site
INR 1,000,000 - 1,500,000
Senior AI DevOps Engineer - Azure/AWS & Python
Senior AI DevOps Engineer - Azure/AWS & Python

NewVision Software • Pune District

Hybrid
INR 2,600,000 - 4,200,000
AI-ML Engineer
AI-ML Engineer

KanthamAi • Mumbai

On-site
INR 2,000,000 - 3,000,000