Backend / ML-Ops Engineer — Speech Model Deployment & Inference Optimization

OutcomesAI

Bengaluru

On-site

INR 1,500,000 - 2,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OutcomesAI, based in Bengaluru, India, is seeking a candidate to own the infrastructure for integrating AI speech models into production. The role includes containerizing models, developing CI/CD pipelines, and optimizing inference performance on GPU. Candidates should have 4–6 years of ML-ops experience, a degree in Computer Science, and proficiency in tools like Triton and Docker. Opportunities for collaboration with AI teams and involvement in cutting-edge healthcare technology are offered.

Qualifications

  • 4–6 years of experience in backend or ML-ops, with at least 1–2 years in GPU inference pipelines.
  • Proven experience deploying models to production environments with measurable latency gains.
  • Interest in inference optimization, mixed precision, and quantization.

Responsibilities

  • Containerize and deploy speech models using Triton Inference Server with TensorRT/FP16 optimizations.
  • Develop and manage CI/CD pipelines for model promotion.
  • Configure autoscaling on Kubernetes based on active calls or streaming sessions.
  • Build health and observability dashboards for monitoring.
  • Integrate LM bias APIs and implement inference paths for low-latency scenarios.

Skills

Triton
TensorRT
Docker
Kubernetes
GPU scheduling
Python
Bash
AWS
GCP
Azure

Education

B.Tech / M.Tech in Computer Science or related field

Tools

Prometheus
Grafana
ELK

Job description

OutcomesAI is a healthcare technology company building an AI-enabled nursing platform designed to augment clinical teams, automate routine workflows, and safely scale nursing capacity. Our solution combines AI voice agents and licensed nurses to handle patient communication, symptom triage, remote monitoring, and post-acute care — reducing administrative burden and enabling clinicians to focus on direct patient care.

Our core product suite includes:

  • Glia Voice Agents – multimodal conversational agents capable of answering patient calls, triaging symptoms using evidence-based protocols (e.g., Schmitt-Thompson), scheduling visits, and delivering education and follow-ups
  • Glia Productivity Agents – AI copilots for nurses that automate charting, scribing, and clinical decision support by integrating directly into EHR systems such as Epic and Athena
  • AI-Enabled Nursing Services – a hybrid care delivery model where AI and licensed nurses work together to deliver virtual triage, remote patient monitoring, and specialty patient support programs (e.g., oncology, dementia, dialysis)

Our AI infrastructure leverages multimodal foundation models — incorporating speech recognition (ASR), natural language understanding, and text-to-speech (TTS) — fine‑tuned for healthcare environments to ensure safety, empathy, and clinical accuracy. All models operate within a HIPAA‑compliant and SOC 2–certified framework. OutcomesAI partners with leading health systems and virtual care organizations to deploy and validate these capabilities at scale. Our goal is to create the world’s first AI + nurse hybrid workforce, improving access, safety, and efficiency across the continuum of care.

Own the infrastructure and pipelines for integrating trained ASR/TTS/Speech‑LLM models into production. Focus on scalable serving, GPU optimization, monitoring, and continuous improvement of inference latency and reliability.

What You’ll Do
  • Containerize and deploy speech models using Triton Inference Server with TensorRT/FP16 optimizations
  • Develop and manage CI/CD pipelines for model promotion (staging → production)
  • Configure autoscaling on Kubernetes (GPU pools) based on active calls or streaming sessions
  • Build health and observability dashboards: latency, token delay, WER drift, SNR/packet loss monitors
  • Integrate LM bias APIs, failover logic, and model switchers for fallback to larger/cloud models
  • Implement on‑device or edge inference paths for low‑latency scenarios
  • Collaborate with AI team to expose APIs for context biasing, rescoring, and diagnostics
  • Optimize GPU/CPU utilization, cost optimization, and memory footprint for concurrent ASR/TTS/Speech LLM workloads
  • Maintain data and model versioning pipelines with MLflow, DVC, or internal registries
Desired Skills
  • Experience with Triton, TensorRT, Docker, Kubernetes, and GPU scheduling
  • Familiarity with speech inference (streaming ASR, TTS pipelines)
  • Proficient in Python, Bash, and cloud services (AWS/GCP/Azure)
  • Understanding of observability stacks (Prometheus, Grafana, ELK)
  • Knowledge of DevSecOps, access policies, and PHI‑safe environments
  • Interest in inference optimization, mixed precision, and quantization
Qualifications
  • B.Tech / M.Tech in Computer Science or related field
  • 4–6 years of experience in backend or ML‑ops; at least 1–2 years with GPU inference pipelines
  • Proven experience deploying models to production environments with measurable latency gains

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech Lead — ASR / TTS / Speech LLM (IC + Mentor)
Tech Lead — ASR / TTS / Speech LLM (IC + Mentor)

OutcomesAI • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Applied AI Scientist - TTS
Applied AI Scientist - TTS

FutureLeap Search • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Senior DevOps Engineer (Kubernetes & AI Infra)
Senior DevOps Engineer (Kubernetes & AI Infra)

Navikenz • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Senior AI Engineer
Senior AI Engineer

xponentiate • Bengaluru

On-site
INR 2,500,000 - 3,500,000
Staff Engineer
Staff Engineer

Stryker • Bengaluru

Hybrid
INR 3,500,000 - 5,200,000
Senior AI Engineer
Senior AI Engineer

2070Health • Bengaluru

On-site
INR 1,500,000 - 2,500,000
AI Engineer (Voice Applications)
AI Engineer (Voice Applications)

Durus Consulting • Chennai District

On-site
INR 400,000 - 750,000
Competitive compensation
Cutting-edge Voice AI projects
Career growth opportunities
Staff AI Engineer
Staff AI Engineer

Get Well • Bengaluru

On-site
INR 2,000,000 - 3,500,000
Senior Engineer
Senior Engineer

Stryker Group • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Senior backend Engineer
Senior backend Engineer

Infer • Karnataka

On-site
INR 1,000,000 - 2,000,000