Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Vahan.ai is building India's first AI-powered recruitment marketplace for blue-collar workers. This role owns outcomes across the Samvaadini system—voice, WhatsApp agents, the eval suite, and the conversion funnel.
You will ship production ML-enabled software, debug non-determinism, and defend latency and cost decisions with data. Expect to work across code, models, and data in a fast-moving, high-impact environment.
Vahan is building India's first AI-powered recruitment marketplace for the country's 300 million-strong blue-collar workforce — already the largest platform of its kind, backed by Khosla Ventures, Y Combinator, LemmaTree.
Our mission: impact a billion lives globally by giving blue-collar professionals not just jobs, but a real path to economic prosperity. If that vision excites you, here's where you'd make your mark.
Scope: You own outcomes across the system, not one component: the voice and WhatsApp agents, the eval suite, the conversion funnel they feed, and root-cause analysis across code, agent behaviour, models and data.
Experience: 5+ years building production software, with meaningful time spent on an LLM-backed system that ran with real users, long enough to break.
Samvaadini is Vahan's hybrid voice and WhatsApp agent. It places 100,000+ outbound calls a day to people looking for delivery and warehouse work across 900+ Indian cities, qualifies them, answers what they actually want to know about pay and shift timings, and hands the serious ones to a recruiter. It talks in Hindi, Hinglish and a widening set of regional languages — mostly to people on a low-end Android phone in a noisy street, many of whom have never used a chatbot before.
Vahan is a marketplace for blue-collar work. Zomato, Zepto, Swiggy and others hire through us; hundreds of recruitment agencies source through us. When Samvaadini gets better, more people get placed in a job this week instead of next month. That is the point of the team.
Own the outcome, not the ticket. At this level you will more often be handed a number that moved than a task to complete. Qualified handoffs are down 12% — go. Localise it across code, agent behaviour, a silent model or provider change, upstream data, or a genuine market shift, and then argue for the fix with the most leverage. Sometimes that argument is for doing nothing, and that is a real answer here.
Know the funnel cold — outreach, connect, qualify, handoff, interview, placement. Write your own queries. An engineer who has to ask someone else for every number moves at that person's speed, and the gap between a metric moving and us knowing why is where this system loses money.
Own agent behaviour end to end — not "write the prompt," but own the harness: the loop, the tool surface, what goes into context and what gets compacted out, stop conditions, retries, and what happens when a tool times out mid-conversation.
Build and defend the eval suite. Every behaviour change ships behind a golden set and a regression gate. You will spend as much time deciding what "better" means as making it better.
Fight for latency and cost. Time-to-first-audio is the difference between a conversation and a hang-up. Inference cost at 100K calls a day is a real line item. Both are yours to defend.
Work the voice pipeline — streaming ASR on code-switched Hinglish over a bad line, turn detection that does not cut a hesitant speaker off, barge-in that actually cancels the in-flight generation, and graceful failure when the network drops mid-turn.
Debug production non-determinism. Read traces, reproduce a failure that only happens on 0.3% of calls, write the postmortem, and close the loop with an eval so it cannot recur silently.
Use coding agents hard and well. Claude Code from Day 1. We expect you to move fast with it and to know exactly where you stopped trusting it.
Real-time voice stacks (LiveKit, Pipecat, Deepgram, ElevenLabs, Cartesia, telephony). Hindi/Hinglish or other code-switched NLP. Fine-tuning or distilling small models to beat big ones on a narrow task. Durable workflow orchestration. ClickHouse or comparable analytical SQL. Open-source work or public writing. MCP server authoring.
Small team and a very short path from your commit to a rider getting a call. Real scale from day one — the systems you touch run six figures of conversations daily, so your latency and cost decisions show up in a dashboard the same week. And the users are people for whom a job this week rather than next month genuinely changes the year.