Lead MLOps Engineer

Cloudkeeper

Dadri

On-site

INR 4,500,000 - 7,000,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Cloudkeeper is seeking an experienced ML infrastructure engineer to lead AI workload optimization for GPU-powered FinOps. You will drive R&D, design optimization engines, and extend capabilities to LLM workloads while partnering with product and customer success to deliver measurable improvements.

You will mentor engineers, define standards, and help scale the AI infrastructure pillar as the team grows. Strong Python, Linux, and cloud GPU experience are essential.

Qualifications

  • 7+ years of hands-on engineering experience.
  • Production experience with GPU workloads and optimization in real environments.
  • Concrete optimization story with before/after metrics.
  • Strong Python, Linux and systems fundamentals.
  • Understanding of ML model lifecycle: training, serving, inference.
  • MLOps fluency: deployment, monitoring, observability, GPU ops.
  • Hands-on with cloud GPU instances and Kubernetes GPU orchestration.
  • Familiarity with at least one modern LLM inference framework.
  • Strong communication skills to translate tech Optimization into business outcomes.
  • Lead experience: managing 3+ engineers.

Responsibilities

  • Drive R&D and engineering for AI optimization in CloudKeeper FinOps.
  • Build optimization engines for GPU right-sizing, idle shutdown and autoscaling.
  • Extend optimization to LLM workloads: caching, routing, dynamic batching.
  • Translate GPU/ML signals into actionable customer recommendations.
  • Collaborate cross-functionally to ship end-to-end optimization features.
  • Set engineering standards and mentor ML/MLOps engineers.
  • Hire, ramp, and grow ML infrastructure team as headcount expands.

Skills

GPU workloads optimization
Performance optimization
Python
Linux
Systems fundamentals
ML lifecycle understanding
MLOps
Cloud GPUs
Kubernetes
LLM inference frameworks
Communication
People leadership
LLM optimization
LLM workload patterns
CUDA
NCCL
Mixed-precision training
ML platforms
GPU-native clouds
Open source contributions
FinOps
Agile

Education

B.E / B.Tech / M.Tech / MCA

Tools

Kubernetes

Job description

Responsibilities
  • Drive R&D and engineering for AI Infrastructure optimization within CloudKeeper's FinOps for AI platform building the Tuner AI / Commit AI capability on GPU and ML workloads
  • Design and build optimization engines for GPU right-sizing, idle shutdown, spot migration with checkpoint/resume automation, inference batching, quantization, and model placement
  • Extend the optimization stack to LLM-era workloads - caching, model routing, dynamic batching, prompt optimization, RAG-aware architectures
  • Partner with the Lens AI team to translate GPU and ML workload signals into actionable, dollar-quantified optimization recommendations for customers
  • Work cross-functionally with product, platform, and customer success teams to ship optimization features end-to-end (data ingestion optimization engine customer-facing recommendation)
  • Lead technical direction for AI workload optimization, set engineering standards, and mentor the ML / MLOps engineering bench as the AI Infrastructure pillar scales
  • (Lead level) Hire, ramp, and grow a team of ML infrastructure engineers as headcount expands
Must Have
  • B.E / B.Tech / M.Tech / MCA with 7+ years of hands-on engineering experience
  • Production experience with GPU workloads — has measurably optimized GPU utilization, throughput, or cost in a real production environment, not just academic / lab work
  • Strong performance engineering background — must come ready with a concrete optimization story including before/after metrics (latency, throughput, or cost reduction)
  • Strong Python + Linux + systems fundamentals
  • Solid understanding of the ML model lifecycle — training, serving, inference — able to reason about what is running on the GPU and why
  • MLOps fluency — model deployment, monitoring, observability, GPU cluster operations
  • Hands-on with cloud GPU instances (AWS P5 / G6, Azure ND series, GCP A3, or equivalent) and Kubernetes-based GPU orchestration (EKS / AKS / GKE GPU node pools, Karpenter, Run:ai, NVIDIA GPU Operator, or similar)
  • Familiarity with at least one modern LLM inference framework — vLLM, TGI, Triton, SGLang, Ray Serve, or BentoML
  • Strong communication skills — able to translate deep technical optimization into customer / business outcomes
  • (Lead level) Experience managing or technically leading a team of 3+ engineers
Good to Have
  • Deep LLM-era optimization expertise — KV caching, semantic caching, model routing, dynamic batching, quantization (FP16 INT8 INT4), model distillation, structured outputs
  • Familiarity with LLM workload patterns — RAG, agents, embeddings, vector databases (Pinecone, Weaviate, Qdrant)
  • CUDA, NCCL, mixed-precision training and inference
  • Experience with managed ML training platforms — SageMaker, Azure ML, Vertex AI, Databricks Mosaic
  • Exposure to GPU-native clouds — CoreWeave, Lambda Labs, RunPod, Crusoe
  • Open source contributions to ML infrastructure projects — vLLM, llama.cpp, TGI, Ray, Triton, KubeRay
  • Adjacent experience in cloud cost optimization / FinOps — Spot.io, ScaleOps, Granulate, CAST AI
  • Comfort with Agile methodology and modern engineering practices (CI/CD, code review, observability)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI / ML Ops Engineer
AI / ML Ops Engineer

Zoho • India

Remote
INR 1,800,000 - 2,800,000
MLOps Engineer
MLOps Engineer

Unico Connect LLP. • Mumbai

On-site
INR 1,000,000 - 1,500,000
Technical Lead-Machine learning
Technical Lead-Machine learning

Myntra • Bengaluru

Hybrid
INR 400,000 - 700,000
AI / ML Ops Engineer
AI / ML Ops Engineer

Zoho • Bengaluru South

On-site
INR 1,800,000 - 3,200,000
MLOps Lead / Engineer
MLOps Lead / Engineer

National e-Governance Division (NeGD), Digital India Corporation • Delhi

On-site
INR 3,000,000 - 6,000,000
Senior MLOps Engineer
Senior MLOps Engineer

Unico Connect LLP. • Mumbai

On-site
INR 2,500,000 - 3,500,000
Lead MLOps DevOps Engineer
Lead MLOps DevOps Engineer

Sonata Software • Hyderabad, Chennai District, Bengaluru

On-site
INR 4,000,000 - 7,000,000
Member of Technical Staff
Member of Technical Staff

eBay • Bengaluru

On-site
INR 4,000,000 - 7,500,000
ML Ops Lead
ML Ops Lead

Corover • Delhi

Hybrid
INR 1,500,000 - 2,500,000
Lead / Principal MLOps Engineer
Lead / Principal MLOps Engineer

Nextloop Technologies LLP • Chennai District

On-site
INR 4,000,000 - 7,000,000