Machine Learning Engineer

Valiance Solutions

Bengaluru Urban

On-site

INR 3,500,000 - 5,500,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Valiance seeks a Senior LLMOps Engineer to own the end-to-end LLM inference infrastructure, running on GPUs, with a focus on reducing cost and latency while ensuring reliability for enterprise and government clients. You will lead production-grade pipelines and collaborate with ML engineers on optimizations.

This high-ownership role offers access to large-scale GPU infrastructure and the opportunity to shape Valiance's AI platform as we scale, with competitive compensation and performance-linked

Qualifications

  • 3+ years of hands-on production LLM inference experience.
  • Deep expertise with vLLM in production and batching.
  • Proficiency with Docker and Kubernetes for GPU inference.
  • Experience building REST/gRPC APIs for model serving at scale.
  • Experience evaluating open-source LLMs for cost vs. quality tradeoffs.

Responsibilities

  • Design and operate production-grade LLM inference pipelines on GPU clusters.
  • Evaluate and deploy small-to-medium open-source LLMs as cost-efficient alternatives.
  • Tune and manage vLLM deployments — batching, attention, tensor parallelism, quantization.
  • Build and maintain model-serving APIs with observability dashboards.
  • Architect Kubernetes-based autoscaling strategies for inference workloads.
  • Run structured A/B experiments comparing model variants and batching strategies.
  • Collaborate with applied ML engineers to identify latency and cost bottlenecks.
  • Establish and enforce SLOs for inference reliability and runbooks for incidents.

Skills

LLMOps
Docker
Kubernetes
REST APIs
Model serving

Tools

vLLM

Job description

Valiance is a deeptech AI company building sovereign and mission-critical AI solutions for enterprises, public sector, and government institutions. From predictive maintenance and demand planning to sovereign AI for citizen services, we design systems that thrive in high-stakes environments. Recognized with the NASSCOM AI Game Changers Award and the Aegis Graham Bell Award, and a certified Google Cloud Partner, our 200+ engineers and data scientists are shaping the future of industries and societies through responsible AI.

The Role

We are looking for a senior LLMOps Engineer who has taken LLM inference optimization from idea to production — not just proof of concept. You will own the end-to-end efficiency of our LLM inference infrastructure running on GPUs, driving down cost and latency while maintaining the reliability our enterprise and government clients demand. This is a high-ownership, high-impact role on a team building some of India's most consequential AI systems.

What You Will Do
  • Design and operate production-grade LLM inference pipelines on GPU clusters, optimizing for throughput, latency, and cost per token.
  • Evaluate and deploy small-to-medium open-source LLMs (e.g., Mistral, Llama, Phi, Gemma) as cost-efficient alternatives to large models without sacrificing output quality.
  • Tune and manage vLLM deployments — including continuous batching, paged attention, tensor parallelism, and quantization (GPTQ, AWQ, FP8) — in production environments.
  • Build and maintain model-serving APIs with robust observability: latency percentiles, GPU utilization, queue depths, and cost-per-request dashboards.
  • Architect Kubernetes-based autoscaling strategies for inference workloads, balancing cold-start penalties against cost at scale.
  • Run structured A/B experiments comparing model variants, quantization levels, and batching strategies using production traffic — not synthetic benchmarks.
  • Collaborate with applied ML engineers and solution architects to identify latency and cost bottlenecks across the model serving stack.
  • Establish and enforce SLOs for inference reliability, and build alerting and runbooks for production incidents.
What We Are Looking For
Non-Negotiables
  • 3+ years of hands-on experience operating LLM inference in production — demonstrable cost and latency improvements, not POC results.
  • Deep expertise with vLLM in production: batching strategies, memory management, quantization tradeoffs.
  • Proficiency with Docker and Kubernetes for deploying and scaling GPU inference workloads.
  • Experience building and maintaining REST/gRPC APIs for model serving at scale.
  • Hands-on experience with open-source LLMs and the ability to evaluate model-quality vs. cost tradeoffs for real use cases.
Strong Advantages
  • Experience with GPU memory profiling and optimization (CUDA-level awareness a plus).
  • Familiarity with model distillation, speculative decoding, or flash attention implementations.
  • Exposure to multi-GPU and multi-node inference setups.
  • Experience with inference frameworks beyond vLLM: TGI, TensorRT-LLM, Triton Inference Server.
  • Familiarity with sovereign AI or air-gapped deployment constraints.
Why Valiance
  • You will work on AI systems that are actually deployed at scale — used by government institutions and large enterprises, not just demoed.
  • Direct access to GPU infrastructure with meaningful compute budgets — no GPU rationing.
  • A culture that rewards engineering depth and production ownership over slide decks.
  • Competitive compensation with performance-linked incentives.
  • Opportunity to define how Valiance builds its AI platform as we scale.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff
Member of Technical Staff

eBay • Bengaluru

On-site
INR 4,000,000 - 7,500,000
LLM Ops Engineer
LLM Ops Engineer

gnani.ai • Bengaluru

On-site
INR 2,800,000 - 4,800,000
LLM Engineer (Large Language Models)
LLM Engineer (Large Language Models)

Fospe UK Ltd • Bengaluru

On-site
INR 2,500,000 - 5,200,000
Competitive compensation with bonuses
Hybrid work at Bangalore Innovation Cn
Health, dental, wellness insurance
+3
Senior Forward Deployed Engineer I Ai Inference Digitalocean Inc Bengaluru
Senior Forward Deployed Engineer I Ai Inference Digitalocean Inc Bengaluru

Vibehackers • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Travel up to 30%
Open-source contributions
MTS 2, AI Platform Professional
MTS 2, AI Platform Professional

The Networker • Bengaluru

On-site
INR 3,000,000 - 5,200,000
Interesting Job Opportunity: Carelon - Artificial Intelligence Engineer - Python/LLM
Interesting Job Opportunity: Carelon - Artificial Intelligence Engineer - Python/LLM

Carelon Global Solutions India • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior AI/ML Data Scientist or Engineer - LLM Fine-Tuning, Evaluation & Inference
Senior AI/ML Data Scientist or Engineer - LLM Fine-Tuning, Evaluation & Inference

Vamstar • India

On-site
INR 600,000 - 1,200,000
AI/ML Engineer
AI/ML Engineer

Zohorecruit • Bengaluru

On-site
INR 1,500,000 - 2,100,000
AI/ML Engineer
AI/ML Engineer

Zohorecruit • Bengaluru Urban

On-site
INR 2,500,000 - 4,500,000
AI Developer
AI Developer

Salvo Software • Bengaluru

On-site
INR 1,800,000 - 3,000,000