Our Healthcare Engineering team builds the AI systems that run directly inside hospital operations not pilots, not demos. Every Voice AI agent, LLM pipeline, and clinical workflow we ship carries real patient safety stakes. This role sits at the center of that responsibility: you decide what's actually ready to go live, and what isn't.
About the Role
You will be the last line of defense before our AI agents touch real patients and real hospital operations. This isn't generic QA it's building the evaluation infrastructure, benchmarks, and simulation frameworks that decide whether a Voice AI agent, an LLM pipeline, or a clinical workflow is safe and reliable enough to ship. You will own the "does this actually work?" question for every AI system we put into production, across voice, language, and healthcare integrations. If you care deeply about rigor, edge cases, and building things that don't break when it matters most, this role is built for you.
What You'll Do
- Design and run automated test suites for Voice AI agents, LLM pipelines, and prompt chains, covering functional correctness, edge cases, and failure modes.
- Build real conversation simulation frameworks that stress test voice agents against realistic and adversarial hospital scenarios before they ever reach a live patient or provider.
- Develop benchmarking pipelines that score model and agent performance across accuracy, latency, hallucination rate, tone, and task completion.
- Own prompt evaluation and regression testing, catching quality drops the moment prompts, models, or pipelines change.
- Build and maintain evaluation datasets and golden answer sets specific to clinical and hospital workflows.
- Evaluate integrations FHIR/HL7 data flows, EHR connections, telephony for correctness, data integrity, and failure handling.
- Continuously monitor production agents for quality drift, and build dashboards and alerts that surface issues before they become incidents.
- Partner with FDEs, the Voice & Agentic AI team, and clinical advisors to define what "good" looks like for each workflow, and translate that into measurable test criteria.
- Own healthcare specific compliance checks within the QA process, ensuring evaluations account for patient safety critical flows, not just generic product bugs.
- Drive root cause analysis on failures and feed learnings back into prompt design, model selection, and agent architecture.
What We're Looking For
Technical
- Strong programming fundamentals in Python, with the ability to build test harnesses, scripts, and automation pipelines from scratch.
- Working knowledge of LLM evaluation techniques hallucination detection, factuality scoring, prompt regression testing, LLM-as-judge patterns.
- Familiarity with Voice AI evaluation transcription accuracy (WER), latency benchmarking, conversational quality scoring and enough understanding of real-time voice pipelines (e.g. LiveKit, Pipecat, or other WebRTC-based frameworks) to know where a voice agent is actually likely to fail.
- Understanding of testing frameworks and CI-style automated pipelines (pytest or equivalent) applied to AI systems rather than traditional software.
- Basic statistics and data analysis skills you're comfortable turning raw logs into a clear quality score or trend.
- Familiarity with dashboarding/monitoring tools (Grafana, simple internal dashboards, or spreadsheet-based reporting at minimum).
Good to Have
- Healthcare Interoperability & Coding Standards: Exposure to healthcare data exchange standards such as FHIR and HL7v2, or medical terminologies including ICD-10 and SNOMED CT.
- Adversarial AI & Red Teaming: Hands-on experience with red teaming, adversarial prompt-injection testing, jailbreak testing, edge-case analysis, and simulation-based evaluation of non-deterministic AI systems.
- LLM Council & Multi-Agent Testing: Familiarity with LLM Council-style evaluation, where multiple LLMs or specialized evaluators assess, critique, and compare model responses to improve reliability, consistency, and evaluation coverage.
- Real-World AI Simulations & Scenario Testing: Experience designing and running realistic simulations that reproduce production-like user interactions, multi-turn conversations, edge cases, failure scenarios, and complex agent workflows for evaluating AI systems before and after deployment.
- Modern LLM Evaluation Methods: Awareness of emerging approaches to LLM and agent evaluation, including model-as-a-judge, pairwise evaluation, rubric-based grading, synthetic test generation, automated evaluation, human-in-the-loop assessment, trajectory evaluation, and regression testing.
- End-to-End Voice Agent Evaluation: Direct experience evaluating real-time WebRTC-based or telephony-integrated voice agents using technologies such as SIP, Twilio, or Exotel, including conversation quality, latency, interruption handling, tool execution, and end-to-end task completion.
- AIOps & Observability Pipelines: Familiarity with tracing and analyzing multi-turn agentic trajectories using observability and evaluation platforms such as LangSmith, Phoenix, DeepEval, Braintrust, or similar tools.
- Continuous AI Evaluation: Understanding of building reusable evaluation datasets, automated test suites, benchmark pipelines, and regression checks to continuously monitor model quality as prompts, models, tools, and agent workflows evolve.
- Emerging AI Testing Practices: Willingness and ability to explore new AI testing tools, evaluation frameworks, simulation techniques, benchmarks, and methodologies as the LLM/agent ecosystem evolves.
Mindset
- Obsessive attention to detail : you notice the 1% edge case everyone else scrolls past.
- Extreme ownership : quality is your problem to solve, not something you wait to be assigned.
- Skeptical by default : you don't trust an AI system works until you've tried to break it.
- Clear, structured communicator : able to turn a messy failure into a precise, actionable bug report.
- Comfortable with ambiguity : evaluation frameworks for agentic healthcare AI barely exist yet; you'll be building the playbook, not following one.
- Calm under pressure : you understand that in healthcare, "good enough" isn't good enough.
- Fast learner with a continuous learning habit : new eval technique, new benchmark, new failure mode in the wild: you’re testing it before you’re asked to.
- Experimenter by default : you don't wait for a ticket to try a better way of catching a bug; if you spot a gap in coverage, you build the test for it.
Qualifications
- Final year student with no backlogs and CGPA above 7.5 (or equivalent academic standing).
- A student, recent graduate, or self-taught builder with demonstrable projects (GitHub, hackathons, personal agents/bots, bug bounty work) what you've shipped matters more than the degree.
- Someone who has tested, broken, or evaluated AI/ML systems coursework, personal projects, hackathons, or bug bounties all count.
- Genuine satisfaction from finding the flaw nobody else caught.
- Motivated, not intimidated, by the fact that a missed AI error in healthcare has real human consequences.
- Ready to build evaluation infrastructure from scratch in a fast-moving startup, not follow an existing QA manual.
Why This Role
Evaluating agentic, voice-driven AI in a regulated, patient-safety-critical domain is a genuinely unsolved engineering problem most teams doing this well are a handful of frontier labs and startups. You'll own that problem directly, with real autonomy, and your evaluation calls will decide what actually ships into hospitals.
What You'll Get
- High-Impact Stipend : ₹25,000 - ₹30,000 / month,
- FastTrack Career Growth: Top performers convert to a full-time role with a package of ₹6-10 LPA, based on impact and ownership shown during the internship.
- Own the Quality Bar : Your evaluations decide what AI ships into real hospitals no sandbox testing, no toy datasets.
- Field-Tested Engineering : Work shoulder to shoulder with FDEs and the Voice & Agentic AI team, stress-testing agents against real-world hospital scenarios.
- Master a Rare, In-Demand Skill : Fast-track your mastery of LLM evaluation, red teaming, and benchmarking one of the most sought-after specializations in applied AI right now.
- Ground-Floor Access : Get an insider's seat to building the trust and safety layer of a category-defining healthcare AI company from scratch.
Our Culture
This internship is more than a temporary role. We are building the future of healthcare AI and expect everyone who joins to approach the work with seriousness and commitment.From day one, you help shape the culture, protect our standards, and raise the talent bar. You become part of the bridge to what this company will become.
Our Non-Negotiables
- Uncompromising Honesty : especially when it is difficult, especially when it costs you something.
- Hard Work & Humility : existing in the exact same person, at the exact same time. Leave your ego at the door.
- Impact for Humanity : what we build must tangibly leave this world better than we found it.
- Relentless Excellence : excellence is our only acceptable baseline. Raise the bar, then raise it again. The last 1% of detail is where true trust is won.
- Extreme Ownership : take absolute responsibility. There is no "that is not my job" here.
- Bias to Action : most decisions are reversible. Make them fast, execute fiercely, and learn even faster.
- Solve Aggressively, Ask Fearlessly : bring a proposed fix, and raise your hand early when stuck.
- Resourcefulness & Frugality : constraints breed genius. Embrace them.
- Disagree Openly, Commit Fully : debate passionately to find the truth, then move forward as one unbreakable front.
Note :
Equal Opportunity Employer: AI.Prof is proud to be an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees regardless of race, color, religion, sex, sexual orientation, gender identity, national origin, veteran, or disability status.