Senior MLOps / AI Platform Engineer

Keka Technologies Private Limited

Coimbatore District

On-site

INR 3,000,000 - 6,000,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Aivar Innovations is seeking a Senior MLOps / AI Platform Engineer to design and build an enterprise-grade MLOps and AIOps platform on Kubernetes. You will deploy, manage, scale, and observe ML/DL and generative AI workloads across cloud and on‑prem environments with GPU support.

You will work with Kubernetes, GPUs, model-serving frameworks, and cloud-native technologies, collaborating with Product, Engineering, AI/ML, DevOps, and Customer Delivery teams to translate complex AI needs into

Qualifications

  • 5–8 years in software/platform engineering, DevOps, SRE, or MLOps.
  • Hands-on Kubernetes with Operators and CRDs.
  • Experience deploying ML/DL/AI models in production.
  • GPU-accelerated workloads on Kubernetes; GPU scheduling and memory management.
  • Familiar with model-serving frameworks (KServe, Triton, vLLM, etc.).
  • Strong Go or Python production API/backend experience.
  • Containers, Helm, CI/CD, IaC, and observability toolchains.
  • Solid Linux, networking, storage, security, and distributed systems knowledge.
  • Excellent debugging, communication, and cross-functional collaboration.

Responsibilities

  • Design and build an enterprise-grade MLOps platform on Kubernetes.
  • Deploy, manage, scale, and observe ML/DL and generative AI workloads.
  • Collaborate with Product, Engineering, AI/ML, DevOps, and Customer Delivery teams.
  • Develop infrastructure capabilities for reliable production move of AI workloads.
  • Implement secure, cloud-native infrastructure across cloud and on‑prem environments.

Skills

Kubernetes
Go or Python
Distributed systems
GPU workloads
CI/CD
Observability
Linux fundamentals
Networking basics
Problem solving

Tools

Kubebuilder
Operator SDK
AWS (EKS, S3, EC2, ECR, IAM)
KServe / Triton / TorchServe
NVIDIA GPU Operator
Prometheus / Grafana / OpenTelemetry
Terraform / IaC
Docker / Containers
CUDA C/C++ (plus)

Job description

Overview

Aivar Innovations is looking for a Senior MLOps / AI Platform Engineer to help design and build an enterprise-grade MLOps and AIOps software platform running on Kubernetes. You will be responsible for developing the infrastructure and platform capabilities required to deploy, manage, scale, and observe machine learning, deep learning, and generative AI workloads across cloud and on-premises environments. You will work extensively with Kubernetes, GPUs, model-serving frameworks, distributed systems, and cloud-native technologies. You will collaborate closely with Product, Engineering, AI/ML, DevOps, and Customer Delivery teams to transform complex AI infrastructure requirements into reliable, secure, and easy-to-use platform capabilities. This role requires strong hands-on engineering experience and a practical understanding of how machine learning models move from experimentation into reliable production environments.

Requirements

5–8 years of experience in software engineering, platform engineering, DevOps, SRE, MLOps, or related infrastructure roles. Strong hands-on experience with Kubernetes, including writing Kubernetes Operators and Custom Resource Definitions using frameworks such as Kubebuilder, Operator SDK, or equivalent. Experience designing and operating cloud-native infrastructure on AWS, particularly Amazon EKS, EC2, S3, ECR, IAM, VPC, and CloudWatch. Experience deploying and operating machine learning, deep learning, or generative AI models in production. Experience running and troubleshooting GPU-accelerated workloads on Kubernetes, with an understanding of GPU scheduling, utilization, memory constraints, and performance. Familiarity with model-serving frameworks such as KServe, NVIDIA Triton Inference Server, vLLM, Ray Serve, TorchServe, or equivalent technologies. Strong programming experience in Go or Python, with experience building production-grade APIs, controllers, or distributed backend services. Experience with containers, Helm, CI/CD, infrastructure as code, and observability tools such as Prometheus, OpenTelemetry, and Grafana. Strong understanding of Linux, networking, storage, security, and distributedsystem fundamentals. Strong debugging, problem-solving, communication, and cross-functional collaboration skills.

Preferred Qualifications

Experience building an MLOps platform, AI infrastructure platform, internal developer platform, or Kubernetes-based enterprise product. Experience writing GPU kernels or performance-critical code using CUDA C/C++ or Triton is a plus. Experience with large language model serving, distributed inference, batching, quantization, or inference-performance optimization. Experience with NVIDIA GPU Operator, MIG, GPU time-slicing, Dynamic Resource Allocation, or similar GPU-management technologies. Familiarity with AWS Inferentia, Trainium, SageMaker, or Amazon Bedrock. Experience operating AI platforms across hybrid-cloud, on-premises, airgapped, or multi-tenant environments. Why You’ll Love Working at Aivar Build a Core AI Platform: Help create a Kubernetes-native platform that enables enterprises to deploy and operate AI workloads at scale. Solve Challenging Infrastructure Problems: Work on Kubernetes, AWS, GPUs, distributed systems, model serving, and enterprise AI operations. Influence

Product Direction:

Work closely with Product, Engineering, and Leadership to shape the platform architecture and roadmap. Modern Technology Stack: Work with cloud-native infrastructure, GPU technologies, generative AI systems, observability platforms, and modern engineering practices. High Ownership: Lead major technical initiatives and take platform capabilities from architecture through production deployment. Accelerated Growth: Build your career in a fast-growing AI startup where your technical decisions will have visible and lasting impact.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

DevOps & Site Reliability Engineer
DevOps & Site Reliability Engineer

Aivar Innovations • Coimbatore District

On-site
INR 1,500,000 - 2,300,000
Mentorship
Ownership of greenfield projects
Modern cloud-native tech exposure
+2
Associate AI Architect
Associate AI Architect

Keka Technologies Private Limited • India

On-site
INR 4,000,000 - 6,500,000
Senior DevOps Engineer (Kubernetes & AI Infra)
Senior DevOps Engineer (Kubernetes & AI Infra)

Navikenz • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Senior AI/ML Engineer
Senior AI/ML Engineer

Aivar Innovations • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Learn from Experts
Direct Ownership of projects
Modern Generative AI stacks
+2
Senior DevOps Engineer
Senior DevOps Engineer

Capitolis • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Senior Software Engineer / Software Engineering Lead
Senior Software Engineer / Software Engineering Lead

Keka Inc. • Bengaluru

On-site
INR 4,000,000 - 6,400,000
Senior Engineer - AI Platform
Senior Engineer - AI Platform

NetConnectGlobal • Bengaluru

On-site
INR 4,200,000 - 6,000,000
Infrastructure Engineer
Infrastructure Engineer

AION • Bengaluru

On-site
INR 3,800,000 - 6,000,000
Competitive compensation
Flexible work options
Wellness benefits
Senior MLOps Engineer
Senior MLOps Engineer

Unico Connect LLP. • Mumbai

On-site
INR 2,500,000 - 3,500,000
Senior Platform Engineer
Senior Platform Engineer

EPAM Systems • Hyderabad

On-site
INR 2,600,000 - 5,200,000