Lead Artificial Intelligence Engineer

Vahan.ai

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

15 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Vahan.ai is building India's first AI-powered recruitment marketplace for blue-collar workers. This role owns outcomes across the Samvaadini system—voice, WhatsApp agents, the eval suite, and the conversion funnel.

You will ship production ML-enabled software, debug non-determinism, and defend latency and cost decisions with data. Expect to work across code, models, and data in a fast-moving, high-impact environment.

Qualifications

  • Shipped an LLM-backed system in production with real users.
  • Experience with async Python and building services, not notebooks.
  • Ability to reason about latency and cost with data-driven decisions.

Responsibilities

  • Own the outcome across the Samvaadini system: voice, WhatsApp, eval suite, funnel, data traces.
  • Identify and fix bottlenecks across code, agent behavior, models and data.
  • Build and defend the eval suite; ensure regression gates are in place.
  • Debug production non-determinism; read traces and reproduce rare failures.
  • Define and implement latency/cost improvements and justify decisions.
  • Work across voice pipeline, ASR, and tool surfaces for robust interactions.

Skills

LLM-backed system
Async Python
Production software
Debugging
Latency optimization
Cost optimization
Coding agents
Data-driven decision making

Tools

Claude Code
Python

Job description

Vahan is building India's first AI-powered recruitment marketplace for the country's 300 million-strong blue-collar workforce — already the largest platform of its kind, backed by Khosla Ventures, Y Combinator, LemmaTree.

Our mission: impact a billion lives globally by giving blue-collar professionals not just jobs, but a real path to economic prosperity. If that vision excites you, here's where you'd make your mark.

Scope: You own outcomes across the system, not one component: the voice and WhatsApp agents, the eval suite, the conversion funnel they feed, and root-cause analysis across code, agent behaviour, models and data.

Experience: 5+ years building production software, with meaningful time spent on an LLM-backed system that ran with real users, long enough to break.

What Samvaadini Is

Samvaadini is Vahan's hybrid voice and WhatsApp agent. It places 100,000+ outbound calls a day to people looking for delivery and warehouse work across 900+ Indian cities, qualifies them, answers what they actually want to know about pay and shift timings, and hands the serious ones to a recruiter. It talks in Hindi, Hinglish and a widening set of regional languages — mostly to people on a low-end Android phone in a noisy street, many of whom have never used a chatbot before.

Vahan is a marketplace for blue-collar work. Zomato, Zepto, Swiggy and others hire through us; hundreds of recruitment agencies source through us. When Samvaadini gets better, more people get placed in a job this week instead of next month. That is the point of the team.

What You Will Actually Do

Own the outcome, not the ticket. At this level you will more often be handed a number that moved than a task to complete. Qualified handoffs are down 12% — go. Localise it across code, agent behaviour, a silent model or provider change, upstream data, or a genuine market shift, and then argue for the fix with the most leverage. Sometimes that argument is for doing nothing, and that is a real answer here.

Know the funnel cold — outreach, connect, qualify, handoff, interview, placement. Write your own queries. An engineer who has to ask someone else for every number moves at that person's speed, and the gap between a metric moving and us knowing why is where this system loses money.

Own agent behaviour end to end — not "write the prompt," but own the harness: the loop, the tool surface, what goes into context and what gets compacted out, stop conditions, retries, and what happens when a tool times out mid-conversation.

Build and defend the eval suite. Every behaviour change ships behind a golden set and a regression gate. You will spend as much time deciding what "better" means as making it better.

Fight for latency and cost. Time-to-first-audio is the difference between a conversation and a hang-up. Inference cost at 100K calls a day is a real line item. Both are yours to defend.

Work the voice pipeline — streaming ASR on code-switched Hinglish over a bad line, turn detection that does not cut a hesitant speaker off, barge-in that actually cancels the in-flight generation, and graceful failure when the network drops mid-turn.

Debug production non-determinism. Read traces, reproduce a failure that only happens on 0.3% of calls, write the postmortem, and close the loop with an eval so it cannot recur silently.

Use coding agents hard and well. Claude Code from Day 1. We expect you to move fast with it and to know exactly where you stopped trusting it.

What We Are Looking For
  • You have shipped an LLM-backed system that ran in production, with real users, long enough to break.
  • You can take a business metric that moved and find the cause, across whichever layer it turns out to live in — and you are as comfortable in a warehouse query as in a stack trace.
  • When you look at a system, you can say what would improve it most — and defend that against the three more obvious things it is not.
  • You are comfortable in async Python and you write services, not notebooks.
  • You bring up evaluation before we do.
  • You have opinions about model choice grounded in latency and cost numbers — and you change them when shown better numbers.
  • You are fluent with coding agents, and can name the last time one confidently produced something wrong and how you caught it.
  • You can explain a technical trade-off to a recruiter or an ops lead without dumbing it down.
  • You learn fast in public. This field moves faster than any of us can individually keep up with; we hire for the rate of learning, not the current stock of it.
Nice to Have — Genuinely Not Required

Real-time voice stacks (LiveKit, Pipecat, Deepgram, ElevenLabs, Cartesia, telephony). Hindi/Hinglish or other code-switched NLP. Fine-tuning or distilling small models to beat big ones on a narrow task. Durable workflow orchestration. ClickHouse or comparable analytical SQL. Open-source work or public writing. MCP server authoring.

Why This Is Worth Your Time

Small team and a very short path from your commit to a rider getting a call. Real scale from day one — the systems you touch run six figures of conversations daily, so your latency and cost decisions show up in a dashboard the same week. And the users are people for whom a job this week rather than next month genuinely changes the year.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Engineer
Lead AI Engineer

Vahan • Karnataka

On-site
INR 4,000,000 - 7,000,000
Lead AI Engineer
Lead AI Engineer

United States Digital Space LLC • Bengaluru

On-site
INR 3,000,000 - 5,500,000
Senior AI Engineer - Voice AI & Agentic Systems
Senior AI Engineer - Voice AI & Agentic Systems

Danish Mullaji • Gurugram District

On-site
INR 4,500,000 - 7,500,000
Agentic AI Engineer | AI Labs | Founder's Office - GO2026
Agentic AI Engineer | AI Labs | Founder's Office - GO2026

GrabOn is Registered Trademark of Inspirelabs Solutions Pvt. Ltd. • Hyderabad

On-site
INR 1,000,000 - 1,500,000
Member of Technical Staff - Applied AI Research
Member of Technical Staff - Applied AI Research

Vaya • Delhi

On-site
INR 1,500,000 - 2,500,000
Founding Engineer
Founding Engineer

Blaugarnet Inc. • Pune District

Hybrid
INR 5,000,000 - 6,000,000
Founding equity
AI Client Success (YOE: 2 yrs+)
AI Client Success (YOE: 2 yrs+)

Bolna AI • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Forward Deployed Engineer (Contract) 3–5 years Remote, Bengaluru
Forward Deployed Engineer (Contract) 3–5 years Remote, Bengaluru

Realfast • Bengaluru

Remote
INR 1,200,000 - 2,400,000
Software Engineer (Junior/Mid) — Agentic AI (Quant Research Platform) | Chennai (on-site)
Software Engineer (Junior/Mid) — Agentic AI (Quant Research Platform) | Chennai (on-site)

Jnaara • Chennai District

On-site
INR 1,500,000 - 3,000,000
Relocation support available
Equity-based grant
Clear path up: comp/scope re-benchmark
Agent Engineer | AI
Agent Engineer | AI

Stealth Video Generation Startup • India

On-site
INR 2,000,000 - 3,000,000
Competitive Pay
Work-Life Balance
Cutting-Edge AI Work
+1