Senior Member of Technical Staff: ML Systems and Infrastructure

DevRev

Bengaluru

On-site

INR 1,500,000 - 2,300,000

Full time

7 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

DevRev is building an AI-first infrastructure platform in Bengaluru. The role focuses on designing and owning end-to-end ML infrastructure, from training to low-latency inference, coordinating with AI research teams, and delivering robust CI/CD for model lifecycle.

You will work with Kubernetes, modern serving stacks, and observability tooling to ensure scalable, reliable deployments across complex ML workloads.

Qualifications

  • 5+ years in infrastructure or software engineering for ML platforms.
  • 2+ years in MLOps or ML infra for large-scale systems.
  • Bachelor’s or Master’s degree in CS/Engineering or related field.
  • Hands-on with Kubernetes in production and cloud-native tools.

Responsibilities

  • Architect end-to-end AI infrastructure for ML model lifecycle.
  • Design and scale inference stacks for low latency and high availability.
  • Partner with AI research and data teams to streamline experiments and deployments.
  • Build CI/CD/CT pipelines (Argo Workflows, ArgoCD, GitHub Actions) for model validation and rollout.

Skills

5+ years in infrastructure
2+ years in MLOps or ML infra
Strong coding in Python or Go
Observability mindset

Education

Bachelor’s or Master’s in CS/Engineering

Tools

Kubernetes
Helm
ArgoCD
Argo Workflows
GitHub Actions
Prometheus
Grafana
OpenTelemetry
vLLM
SGLang
Triton Inference Server
Ray Serve
PyTorch
Jax
TensorFlow

Job description

About DevRev

At DevRev, we're building the future of work with

About DevRev

At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform, giving employees real-time insights, proactive suggestions, and powerful agentic actions. It extends your existing software with AI-native apps and agents that work alongside your teams and customers – updating workflows, coordinating across teams, and eliminating repetitive work. We call this Team Intelligence: human-AI collaboration that breaks down silos, brings people back together, and frees you to solve bigger problems. Backed by Khosla Ventures and Mayfield with $150M+ raised, DevRev is trusted by global companies across industries.

What You’ll Do
  • Architect the Future of AI Infrastructure: You will design, build, and own the end-to-end platform that supports the entire lifecycle of our ML models—from massive-scale distributed training to ultra-low-latency, highly-available inference.
  • Optimize and Serve Cutting-Edge Models: You'll implement and scale sophisticated inference stacks for LLMs using frameworks like vLLM, TensorRT-LLM, or SGLang. You’ll solve complex challenges in throughput, latency, token streaming, and automated scaling to deliver a seamless user experience.
  • Empower AI Innovation: You will act as a strategic partner to our AI Research and Data Science teams. You’ll create a seamless developer experience that accelerates their ability to experiment, fine-tune, and deploy groundbreaking models with velocity and confidence.
  • Automate Everything: You'll develop robust CI/CD/CT (Continuous Training) pipelines using tools like Argo Workflows, ArgoCD, and GitHub Actions to automate model validation, deployment, and lifecycle management, ensuring our systems are both agile and rock-solid.
What Are We Looking For
  • Experience: 5+ years in infrastructure or software engineering, with at least 2+ years laser-focused on MLOps or ML infrastructure for large-scale distributed systems.
  • Education: A Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • Kubernetes & Cloud Native Expertise: Deep, hands-on expertise with Kubernetes in production. You are fluent in the cloud-native ecosystem, including Helm, ArgoCD, and Argo Workflows.
  • GPU & Cloud Mastery: Optimize the platform’s performance and scalability, considering factors such as GPU resource utilization, data ingestion, model training, and deployment.
  • Modern LLM Serving Experience: Hands-on experience with modern LLM inference serving frameworks (e.g., vLLM, SGLang, Triton Inference Server, Ray Serve). You understand the unique challenges of serving generative models.
  • Strong Coder: Strong programming proficiency in Python or Go, with experience using ML frameworks like PyTorch, Jax, TensorFlow.
  • Observability Mindset: A passion for building observable and resilient systems using modern monitoring tools (e.g., Prometheus, Grafana, OpenTelemetry).
We Would Love To See
  • Deep performance optimization skills, including writing custom inference kernels in CUDA or Triton to accelerate model performance beyond what off-the-shelf frameworks provide.
  • Experience with model optimization techniques like quantization, distillation, and speculative decoding.
  • Exposure to training and serving multi-modal models (e.g., text-to-image, vision-language).
  • Knowledge of AI safety and evaluation frameworks for monitoring model performance for things like bias, toxicity, and hallucinations.

As part of our hiring process, shortlisted candidates will undergo a Background Verification (BGV). By applying, you consent to sharing personal information required for this process. Any offer made will be subject to successful completion of the BGV.

DevRev is an equal opportunity employer and does not discriminate on the basis of race, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition, or any other basis protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Ops Engineer
AI/ML Ops Engineer

revolte.ai • Chennai District

On-site
INR 900,000 - 1,800,000
Member of Applied AI - Strategist
Member of Applied AI - Strategist

DevRev • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Forward Deployed Engineer
Forward Deployed Engineer

DevRev • India

On-site
INR 2,000,000 - 3,200,000
Forward Deployed Engineer
Forward Deployed Engineer

DevRev • Tamil Nadu

On-site
INR 1,200,000 - 1,800,000
Senior DevOps Engineer (Kubernetes & AI Infra)
Senior DevOps Engineer (Kubernetes & AI Infra)

Navikenz • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Senior Technical Specialist, AI Solutions : Delhi
Senior Technical Specialist, AI Solutions : Delhi

DevRev • India

On-site
INR 3,000,000 - 5,400,000
AI/ML Engineer
AI/ML Engineer

Jash Data Sciences Pvt. Ltd. • Pune District

On-site
INR 800,000 - 1,200,000
Competitive salary
Learning opportunities
Exposure to latest AI technologies
AI/ML Engineer
AI/ML Engineer

agilisium • Chennai District

On-site
INR 3,000,000 - 6,000,000
AI/ML Engineer
AI/ML Engineer

Creuto Cloud • Khordha

On-site
INR 900,000 - 1,500,000
Solution Engineer : Mumbai
Solution Engineer : Mumbai

DevRev • Mumbai

On-site
INR 1,800,000 - 3,000,000