Senior Data Scientist ( Kernel Optimisation & Inference Engineer)

Eka.Care

Bengaluru

On-site

INR 1,200,000 - 2,400,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical Insurance & Accidental Ins

Job summary

EkaCare, a leading Indian healthcare platform, is hiring a Senior Data Scientist specializing in kernel optimisation and inference engineering. You will push the boundaries of performance from training-time kernels to fast, low-latency on-device inference for clinical contexts.

You will work with the performance lead on 30B MoE training, 2B/4B on-device tiers, and guided roofline-based optimisations. A strong CUDA background and hands-on serving stack experience are essential.

Qualifications

  • 2–5 years of GPU performance work with profiling and shipped optimisations.
  • Fluency in CUDA and memory-hierarchy reasoning (coalescing, occupancy, SRAM tiling).
  • Experience with a modern serving stack beyond running it (e.g., vLLM, TensorRT-LLM).
  • Measurement discipline: profile before optimising and verify correctness after.

Responsibilities

  • Profile training and inference workloads and hunt utilisation gaps across kernels, memory and comms.
  • Write and tune CUDA kernels where existing ops leave real performance on the table.
  • Optimise MoE-specific paths: grouped GEMMs, all-to-all communication, expert load imbalance.
  • Build the fast inference path: vLLM-class serving, continuous batching, prompt/prefix caching for clinical-context workloads, speculative decoding.
  • Own quantisation for the 2B/4B variants (AWQ/GPTQ-class, fp8) — with eval-parity verification, not just perplexity.
  • Make on-device inference real for the hardware Indian clinics actually have.

Skills

CUDA
Memory hierarchy
GPU performance
Profiling workloads
Quantisation/FP8
On-device inference

Tools

vLLM
TensorRT-LLM
SGLang

Job description

Senior Data Scientist ( Kernel Optimisation & Inference Engineer)

2 - 5

Full-Time

Kernel Optimisation & Inference Engineer

Somewhere between the model and the silicon, 10–20% of a training budget goes missing. Your job is to go get it back, and then make the same model fast enough to run in a clinic, and small enough to run on a phone.

About EkaCare and the mission

EkaCare is India's connected healthcare platform: an EMR that doctors run their practices on, a personal health record used by millions of Indians, and one of the deepest integrations with India's ABDM digital-health rails. Our Parrotlet family of medical models already serves Indian doctors in production, and we open-source our work where it counts.

The role:

You’ll work with our performance lead on making everything fast: training-side fused kernels and MFU on the 30B MoE, inference-side latency and throughput, and the quantised 2B/4B on-device tier. Hardware-up: profiler first, roofline reasoning always, custom kernels when the math says so.

What you’ll do
  • Profile training and inference workloads and hunt utilisation gaps across kernels, memory and comms.
  • Write and tune CUDA kernels where existing ops leave real performance on the table, and know when they don't.
  • Optimise MoE-specific paths: grouped GEMMs, all-to-all communication, expert load imbalance.
  • Build the fast inference path: vLLM-class serving, continuous batching, prompt/prefix caching for clinical-context workloads, speculative decoding.
  • Own quantisation for the 2B/4B variants (AWQ/GPTQ-class, fp8) — with eval-parity verification, not just perplexity.
  • Make on-device inference real for the hardware Indian clinics actually have.
What we look for
  • 2–5 years in GPU performance work; you've profiled real workloads and shipped optimisations with before/after numbers you can defend.
  • Working fluency in CUDA, and memory-hierarchy reasoning (coalescing, occupancy, SRAM tiling; you can explain why FlashAttention is fast).
  • Hands-on with a modern serving stack (vLLM, TensorRT-LLM, SGLang or similar) beyond just running it.
  • Measurement discipline: you profile before optimising and verify correctness after.
Bonus
  • fp8 experience on H100/H200-class hardware; torch.compile/inductor internals.
  • Quantisation research or on-device/mobile inference experience.
  • Open-source kernels or serving contributions.
Why this is a rare gig
  • Open source, with your name on it : weights and technical reports ship publicly.
  • India-scale mission : models for a billion people in their own languages.
  • Compute that’s rare to fine : dedicated multi-node H200 training under a national grant.
  • Small senior team: you work with the people who own the recipe.
  • A live deployment path : Government institutes, EkaCare's doctors and patients use what you ship.
Full-Time Employee Benefits
  • Medical Insurance & Accidental Insurance
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer
Senior ML Engineer

Eka.Care • Bengaluru

On-site
INR 3,000,000 - 5,000,000
Medical Insurance & Accidental
ML Engineer
ML Engineer

Eka.Care • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Medical Insurance
Accidental Insurance
Performance Engineer, Kernels
Performance Engineer, Kernels

Sarvam • Bengaluru

On-site
INR 4,000,000 - 7,000,000
ML Engineer (Training Infra), Foundational Models
ML Engineer (Training Infra), Foundational Models

Sarvam • Bengaluru

On-site
INR 1,200,000 - 2,000,000
High ownership and impact
AI-first approach
Foundational Model Engineer — Multimodal & Agentic Medical AI Systems
Foundational Model Engineer — Multimodal & Agentic Medical AI Systems

SAIGroup • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Competitive compensation
Opportunity to shape technical architecture
Collaboration with researchers and clinicians
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Bengaluru

Hybrid
INR 5,500,000 - 9,000,000
Senior ML Engineer
Senior ML Engineer

IDfy • Mumbai

On-site
INR 2,500,000 - 6,000,000
Performance Engineer, On-Device Inference
Performance Engineer, On-Device Inference

Sarvam • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Performance Engineer, Inference
Performance Engineer, Inference

Sarvam • Bengaluru

On-site
INR 3,500,000 - 7,500,000
ML Engineer, TorchBridge
ML Engineer, TorchBridge

Cloudly Inc • India

On-site
INR 100,000 - 150,000
Competitive salary
Performance-based commission
Two annual bonuses
+3