Senior DevOps Engineer (Kubernetes & AI Infra)

Navikenz

Bengaluru

Hybrid

INR 1,500,000 - 2,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology solutions company is searching for an experienced Senior DevOps Engineer in Bengaluru, India. The successful candidate will be responsible for designing and maintaining scalable cloud infrastructure while automating AI/ML workflows using Kubernetes and related tools. Candidates should have over 8 years of experience in DevOps, strong knowledge of Kubernetes fundamentals, and hands-on experience with cloud platforms like AWS or Azure. Strong problem-solving skills and scripting proficiency are essential for this role.

Qualifications

  • 8+ years of experience in DevOps, SRE, or Platform Engineering roles.
  • Strong knowledge of Kubernetes fundamentals, especially for AI workloads.
  • Hands-on experience with AWS, Azure, or GCP in Kubernetes-based environments.

Responsibilities

  • Design and maintain scalable cloud infrastructure using Kubernetes.
  • Automate AI/ML workflows with tools like Kubeflow and MLflow.
  • Monitor Kubernetes clusters and AI workloads for performance tuning.

Skills

Kubernetes fundamentals
Scripting with Python
AWS or Azure or GCP
CI/CD & MLOps tools
Problem-solving skills
Azure
GCP
MLOps

Tools

Kubeflow
Terraform
Prometheus
Grafana
Prometheus

Job description

We’re looking for an experienced Senior DevOps Engineer who loves working with Kubernetes and AI-driven applications. In this role, you’ll be responsible for designing, implementing, and maintaining scalable cloud infrastructure while supporting MLOps pipelines for AI workloads.

What You’ll Be Doing:
  • Building Scalable Infrastructure: You’ll design, implement, and maintain cloud infrastructure using Kubernetes to handle AI and non-AI workloads efficiently.
  • Developing CI/CD & MLOps Pipelines: Help us automate AI/ML workflows using tools like Kubeflow, MLflow, or Argo Workflows, ensuring seamless deployment and monitoring of AI models.
  • Optimizing AI Model Deployments: Work with ML engineers to fine-tune LLM models, AI-driven applications, and containerized environments for smooth operation.
  • Monitoring & Performance Tuning: Keep an eye on Kubernetes clusters and AI workloads, using tools like Prometheus, Grafana, and Loki to ensure high availability and performance.
  • Automating Everything: Whether it’s infrastructure provisioning (Terraform, Helm) or Kubernetes security best practices, you’ll help enforce efficiency and compliance.
Staying Ahead of the Curve:

You’ll have the opportunity to explore and implement emerging AI infrastructure trends, including KServe, Ray, and Triton Inference Server.

What We’re Looking For:
  • 8+ years of experience in DevOps, SRE, or Platform Engineering role, with expertise in Kubernetes and cloud-native DevOps
  • Strong knowledge of Kubernetes fundamentals (deployments, services, ingress, storage, GPU scheduling, multi-cluster management).
  • Proficiency in scripting & automation with Python, Bash, or Go, particularly for AI-related workflows.
  • Hands‑on experience with AWS, Azure, or GCP, especially in Kubernetes-based AI/ML infrastructure (e.g., Amazon SageMaker, GKE with AI, Azure ML).
  • Hands‑on experience with model deployment frameworks (NVIDIA Triton, vLLM. TGI etc.)
  • Experience with Distributed computing, multi‑GPU training on kubernetes and on‑prem GPU clusters.
  • Experience managing resource allocation and autoscaling for large training/inference workloads. (KEDA, HPA etc.)
  • Experience with CI/CD & MLOps tools such as Jenkins, Argo CD, Kubeflow, MLflow, or Tekton.
  • Familiarity with GenAI model deployment, including fine‑tuning, inference optimization, and A/B testing.
  • Hands‑on experience with managed ML services (AWS bedrock, Vertext AI models etc.)
  • Strong problem‑solving skills and a mindset of automating repetitive tasks.
  • Excellent communication skills to collaborate with ML engineers, data scientists, and software teams.
Bonus Points If You Have:
  • Experience with LLMOps (Large Language Model Operations) and deploying LLM‑based applications at scale.
  • Knowledge of Vector Databases (FAISS, Weaviate, Qdrant) for AI‑driven applications.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI_DevOps Engineer
AI_DevOps Engineer

RIA Advisory • Pune District

Hybrid
INR 1,000,000 - 1,500,000
MLOps Engineer
MLOps Engineer

Insight Global • Hyderabad

On-site
INR 1,800,000 - 2,400,000
AI DevOps Engineer — Mid/Senior Level
AI DevOps Engineer — Mid/Senior Level

Stackular • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Devops Engineer
Devops Engineer

Airtel • Gurugram District

On-site
INR 800,000 - 1,400,000
MLOps Engineer
MLOps Engineer

Codvo Private Limited • Pune District

On-site
INR 600,000 - 1,000,000
MLOps Engineer
MLOps Engineer

Codvo.ai • Pune District

On-site
INR 2,000,000 - 3,500,000
Senior Architect - DevOps and ML Op's
Senior Architect - DevOps and ML Op's

Metaplore Solutions Pvt Ltd • Bengaluru

On-site
INR 3,500,000 - 6,000,000
AI DevOps Engineer
AI DevOps Engineer

RIA Advisory LLC. • Pune District

Hybrid
INR 1,000,000 - 1,500,000
Senior DevOps Engineer
Senior DevOps Engineer

Anaptyss • Dadri

On-site
INR 2,400,000 - 4,200,000
AI-ML Engineer
AI-ML Engineer

KanthamAi • Mumbai

On-site
INR 2,000,000 - 3,000,000