Lead MLOps Engineer

Cloudkeeper

Garhi

On-site

INR 3,800,000 - 6,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cloudkeeper seeks a senior ML/AI infrastructure leader to drive R&D and engineering for GPU-driven optimization in its FinOps for AI platform. You will design engines for right-sizing, idle shutdown, and batching, and extend to LLM workloads with caching, routing, and dynamic batching.

You will partner with Lens AI to translate signals into customer-ready optimization recommendations, collaborate with product and customer success, set engineering standards, and mentor a growing ML/MLOps team.

Qualifications

  • 7+ years hands-on engineering experience (B.E./B.Tech./M.Tech./MCA)
  • Production GPU workloads experience with measurable optimization
  • Strong Python, Linux and systems fundamentals
  • ML model lifecycle: training, serving, inference
  • MLOps fluency: deployment, monitoring, observability, GPU ops
  • Hands-on with cloud GPU instances and Kubernetes GPU orchestration
  • Familiarity with modern LLM inference frameworks
  • Lead experience managing 3+ engineers

Responsibilities

  • Drive R&D and engineering for AI infrastructure optimization on GPU/ML workloads
  • Design optimization engines for GPU right-sizing, idle shutdown, batching
  • Extend optimization stack to LLM workloads: caching, routing, dynamic batching
  • Translate GPU/ML signals into actionable optimization recommendations
  • Collaborate with product, platform and customer success to ship features
  • Set engineering standards and mentor ML/MLOps engineers
  • Hire and grow a team of ML infrastructure engineers

Skills

Python
Linux
ML lifecycle
MLOps
Kubernetes
Leadership
Communication

Education

B.E./B.Tech./M.Tech./MCA with 7+ years experience

Tools

Kubernetes-based GPU orchestration
NVIDIA GPU Operator
Run:ai
vLLM

Job description

Responsibilities:
  • Drive R&D and engineering for AI Infrastructure optimization within CloudKeeper's FinOps for AI platform building the Tuner AI / Commit AI capability on GPU and ML workloads
  • Design and build optimization engines for GPU right-sizing, idle shutdown, spot migration with checkpoint/resume automation, inference batching, quantization, and model placement
  • Extend the optimization stack to LLM-era workloads — caching, model routing, dynamic batching, prompt optimization, RAG-aware architectures
  • Partner with the Lens AI team to translate GPU and ML workload signals into actionable, dollar-quantified optimization recommendations for customers
  • Work cross-functionally with product, platform, and customer success teams to ship optimization features end-to-end (data ingestion optimization engine customer-facing recommendation)
  • Lead technical direction for AI workload optimization, set engineering standards, and mentor the ML / MLOps engineering bench as the AI Infrastructure pillar scales
  • (Lead level) Hire, ramp, and grow a team of ML infrastructure engineers as headcount expands
Must Have:
  • B.E / B.Tech / M.Tech / MCA with 7+ years of hands‑on engineering experience
  • Production experience with GPU workloads — has measurably optimized GPU utilization, throughput, or cost in a real production environment, not just academic / lab work
  • Strong performance engineering background — must come ready with a concrete optimization story including before/after metrics (latency, throughput, or cost reduction)
  • Strong Python + Linux + systems fundamentals
  • Solid understanding of the ML model lifecycle — training, serving, inference — able to reason about what is running on the GPU and why
  • MLOps fluency — model deployment, monitoring, observability, GPU cluster operations
  • Hands‑on with cloud GPU instances (AWS P5 / G6, Azure ND series, GCP A3, or equivalent) and Kubernetes-based GPU orchestration (EKS / AKS / GKE GPU node pools, Karpenter, Run:ai, NVIDIA GPU Operator, or similar)
  • Familiarity with at least one modern LLM inference framework — vLLM, TGI, Triton, SGLang, Ray Serve, or BentoML
  • Strong communication skills — able to translate deep technical optimization into customer / business outcomes
  • (Lead level) Experience managing or technically leading a team of 3+ engineers
Good to Have:
  • Deep LLM-era optimization expertise — KV caching, semantic caching, model routing, dynamic batching, quantization (FP16 INT8 INT4), model distillation, structured outputs
  • Familiarity with LLM workload patterns — RAG, agents, embeddings, vector databases (Pinecone, Weaviate, Qdrant)
  • CUDA, NCCL, mixed‑precision training and inference
  • Experience with managed ML training platforms — SageMaker, Azure ML, Vertex AI, Databricks Mosaic
  • Exposure to GPU‑native clouds — CoreWeave, Lambda Labs, RunPod, Crusoe
  • Open source contributions to ML infrastructure projects — vLLM, llama.cpp, TGI, Ray, Triton, KubeRay
  • Adjacent experience in cloud cost optimization / FinOps — Spot.io, ScaleOps, Granulate, CAST AI
  • Comfort with Agile methodology and modern engineering practices (CI/CD, code review, observability)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior MLOps Engineer
Senior MLOps Engineer

Unico Connect LLP. • Mumbai

On-site
INR 2,500,000 - 3,500,000
CloudKeeper - Associate Director - Java
CloudKeeper - Associate Director - Java

CloudKeeper • Varanasi

On-site
INR 2,000,000 - 3,500,000
MLOps Platform Engineer (Chennai / Pune)
MLOps Platform Engineer (Chennai / Pune)

Money Forward India • Chennai District

On-site
INR 3,500,000 - 6,000,000
Machine Learning Engineer
Machine Learning Engineer

Tranzeal • Bengaluru

On-site
INR 3,500,000 - 7,500,000
AI/ML Engineer
AI/ML Engineer

Ambifo • Bengaluru

On-site
INR 2,000,000 - 4,200,000
MLOps Manager
MLOps Manager

Anblicks • Hyderabad

On-site
INR 2,000,000 - 3,000,000
AI Engineer
AI Engineer

Aziro • Chennai District, Bengaluru, Pune District

Hybrid
INR 1,200,000 - 1,800,000
MLOps Engineer
MLOps Engineer

GyanSys Inc. • Bengaluru

Hybrid
INR 1,400,000 - 2,100,000
Distributed Training & Inference Optimization Engineer
Distributed Training & Inference Optimization Engineer

Winzons • India

On-site
INR 3,000,000 - 5,000,000
Senior MLOps Engineer
Senior MLOps Engineer

Vitric Business Solutions • Pune District, Mumbai, Indore District

On-site
INR 180,000 - 300,000