Applied AI Engineer

Adam

Bengaluru

On-site

INR 2,500,000 - 4,200,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Infer seeks an Applied AI Engineer to own evaluation harnesses for voice agents and to drive end-to-end quality improvements. You will work on the system that catches transcriptions, quotes, and disclosures before customers see them, while building benchmarks for rapid model upgrades.

The role centers on ML engineering, LLM prompts, and audio processing, with an emphasis on self-improving loops and production feedback.

Qualifications

  • Production ML engineering experience shipping systems.
  • Strong Python and working ML stack (PyTorch, HuggingFace, pandas, scikit-learn).
  • Hands-on experience designing LLM-based agents and multi-turn state.
  • Experience building eval pipelines for ML/LLM/voice systems.
  • Practical ASR/STT experience; TTS models familiarity.
  • Comfort with audio data, sampling rates, codecs, and alignment.
  • Ability to curate datasets and run ML benchmarks.

Responsibilities

  • Build and maintain the eval framework scoring voice agent quality across transcription, reasoning, tool use, TTS, and outcomes.
  • Design voice agent behavior: prompts, tools, flow, error recovery, guardrails.
  • Drive STT and TTS accuracy via provider comparisons and AB tests.
  • Improve TTS quality, voice selection, latency, and prosody.
  • Curate evaluation datasets from production traffic and hard-case mining.
  • Build benchmarks to test new models within days and run red-team probes.
  • Integrate eval signals into CI with backend teams for regression prevention.
  • Develop self-improvement loops to feed production failures back into prompts.

Skills

ML engineering
Python
LLM-based agents
eval frameworks
ASR/STT
TTS systems
audio data handling

Tools

PyTorch
HuggingFace
Pandas
scikit-learn
Whisper

Job description

About us

Infer is building the operating system for insurance agencies. We make AI agents(including voice agents) that handle the work agencies have always done by hand: qualifying inbound leads, helping producers during live calls, auditing calls after, running renewals, and bringing churned customers back.

Our long bet is that AI eventually sells insurance directly. Agencies are the wedge because that is where the work, the data, and the customer relationships actually live. Get good there, and the rest follows.

We are a YC company and have raised from Stellaris Venture partners and others. Founders are: Vaibhav, Urvin and Suneel. Vaibhav was an architect and AI researcher(at Purdue) now a licensed insurance agent. Urvin worked at BCG, is a surfer with six pack abs. Suneel is an IITian and a philomath.

A few reasons to join us:

  • We like pushing each other on team to test the limits because that's when you rediscover yourself.
  • We’re paranoid about making customers succeed (we challenge whats already good)
  • We love people who question-challenge-build.
  • We're highly transparent founders to work with & love getting challenged.
  • Finally, we love people who’re interdisciplinary.

About the role

We're hiring an Applied AI Engineer to own the system that tells us whether our voice agents are getting better, and to keep them getting better on their own.

Voice quality is the product. If an agent stutters, hallucinates a quote, or misses a disclosure, we lose trust, deals, and sometimes compliance footing. The system that catches all of that before customers do is the most important infrastructure we will build this year.

Today we run thousands of conversations a day with real prospects. We need a harness that scores every change end to end, a benchmark suite that runs against any new model the day it drops, a red-team pipeline that probes our agents for failure modes, and self-improvement loops that feed production failures back into the eval set.

This is an evals and infrastructure role with deep LLM work. You will touch audio, but the center of gravity is the harness and the loops around it. Think of the harness as CI for voice conversations: it runs synthetic and real calls through our stack and scores agent behavior at every layer (STT, LLM, tools, TTS, full call outcomes), so we catch regressions before customers do. New models are coming out every few weeks, so the question is not just whether ours is good today, but whether we can tell within a week if a new open source release should replace it.

What you'll do

  • Building and maintaining the eval framework that scores voice agent quality across transcription, LLM reasoning, tool use, TTS, and full-conversation outcomes
  • Design voice agent behavior: system prompts, tool use, conversation flow, error recovery, and guardrails for real-time interactions
  • Drive STT and TTS accuracy improvements by comparing providers, tuning configurations, and running rigorous A/B experiments the team can act on.
  • Drive TTS quality improvements voice selection, latency vs. fidelity tradeoffs, prosody, edge cases
  • Curate and grow our evaluation datasets, including hard-case mining from production traffic
  • You’ll build benchmarks we can run against any new model in days, run a red-team pipeline that probes for jailbreaks, hallucinated quotes, and compliance failures,
  • Partner with backend engineers to wire eval signals into CI so regressions get caught before they ship
  • Wire eval signals into CI so regressions block merges, and build self-improvement loops where hard cases from production auto-feed the eval set and our prompts optimize themselves over time.
What success looks like
Day 30
  • You understand how our agents work across prompts, tools, evals, telephony, and customer systems.
  • You have shipped a v1 of evals with at least one end-to-end metric the team trusts.
  • You are sitting in on customer call reviews and tagging failure modes by hand to learn where the real problems live.
  • You have one new model (open or closed) benchmarked against our production stack with numbers we can defend.
Day 60
  • The eval system runs on updates and blocks merges that regress on a known set of cases.
  • We have a first red-team suite covering at least three classes of failure modes (jailbreaks, hallucinated quotes, compliance), running on a schedule.
  • Hard-case mining from production calls is automated, so the eval set grows without anyone triaging every example by hand.
  • At least one open source model (Qwen, DeepSeek, or similar) is benchmarked against our production stack with a defensible recommendation on whether to switch.
Day 90
  • We can swap in any new LLM and have a numbers-backed answer on whether to ship it within a week.
  • DSPy or GEPA-style prompt optimization is running over at least one production voice flow, and you have shown measurable lift.
  • Self-improvement v1 is live for at least one failure pattern. The same problem does not get solved twice because the system feeds the fix back into the platform.
  • You are spotting failure patterns across customer accounts and turning them into product fixes the rest of the team builds on.

Must-haves

  • ML engineering experience shipping production systems
  • Strong Python and a working ML stack (PyTorch, Huggingface, pandas, scikit-learn)
  • Hands-on experience designing LLM-based agents: prompting, tool/function calling, multi-turn state, structured outputs
  • Hands-on experience building evals or eval frameworks for ML, LLM, or voice systems. Built LLM-as-judge eval pipelines and know their failure modes
  • Practical experience with ASR/STT comparing providers, fine-tuning, or running open models like Whisper
  • Practical experience with TTS systems (ElevenLabs or open models)
  • Comfortable working with audio data: sample rates, codecs, noise, alignment

Nice-to-haves

  • Designed voice agents specifically handled barge-in, interruption recovery, disfluencies, and natural turn-taking at the prompt/behavior layer
  • Experience with diarization, VAD, or endpointing models
  • Audio dataset curation, labeling, or annotation pipelines
  • Trained or fine-tuned ASR or TTS models from scratch or on domain audio
  • Experience with active learning or data-flywheel patterns over production traffic
  • Open-source contributions to AI/ML frameworks
  • Familiarity with cost/latency tradeoffs across model providers for real-time voice
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Applied AI Engineer
Applied AI Engineer

Synth (YC S21) • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Senior ML Research Scientist, Speech
Senior ML Research Scientist, Speech

Blue Machines AI • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Senior Software Engineer
Senior Software Engineer

Zingly • Bengaluru

On-site
INR 2,600,000 - 3,800,000
ML Research Engineer, Speech
ML Research Engineer, Speech

Blue Machines AI • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Technical Project Manager
Technical Project Manager

Bigship • Dehradun

On-site
INR 1,800,000 - 3,000,000
GPU compute resources
AI research budget
Flexible working arrangements
Software Engineer - 2 (Voicebot)
Software Engineer - 2 (Voicebot)

Exotel • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Principal AI Engineer
Principal AI Engineer

VoiceCare AI • Bengaluru

On-site
INR 3,000,000 - 5,000,000
Senior Voice AI Engineer
Senior Voice AI Engineer

Keka Technologies Private Limited • Bengaluru

On-site
INR 4,000,000 - 7,000,000
ML Research Engineer Speech
ML Research Engineer Speech

Blue Machines AI • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Lead AI Engineer
Lead AI Engineer

United States Digital Space LLC • Bengaluru

On-site
INR 3,000,000 - 5,500,000