Stand out for this role — generate a tailored resume and cover letter in about a minute.
Aircall, an AI-powered customer communications platform, seeks an ML Engineer to lead evaluation of its AI agents across voice, chat, and messaging. You design evaluation frameworks, train voice models, and drive trustworthy AI metrics.
You will build automated pipelines, calibrate LLM-based judges, monitor production quality, and collaborate with cross-functional teams. A BS in CS/ML, 3+ years in ML engineering, 8+ years total, and hands-on experience with TTS/ASR are required.
Aircall is a unicorn, AI-powered customer communications platform used by 22,000+ companies worldwide to drive revenue, resolve issues faster, and scale customer-facing teams. We’re redefining customer communications by bringing voice, SMS, WhatsApp, and AI together into one seamless workspace. Our momentum comes from a simple idea: help teams work smarter, not harder. Aircall’s AI Voice Agent automates routine calls, AI Assist streamlines post-call work, and AI Assist Pro delivers real-time guidance so people can do their best work. The result is higher revenue, faster resolutions, and teams that scale with confidence. Aircall is headquartered in Paris, our European HQ, with a strong North American presence anchored in Seattle, our North American HQ, and teams across Madrid, London, Berlin, San Francisco, New York City, Sydney, and Mexico City. We’ve built a product customers love and a business that’s scaling quickly, backed by world‑class investors and driven by rapid AI innovation across multiple product lines. At Aircall, you’ll join a company in motion. We’re ambitious, product-driven, and execution-focused, with visible impact, fast decisions, and real growth.
We’re customer-obsessed, data-driven, and focused on delivering meaningful outcomes. We value ownership, continuous learning, and thoughtful speed. If you thrive in a collaborative, fast-moving environment where trust and impact matter, you’ll feel at home here. Aircall's AI suite includes an AI Voice Agent and AI Messaging Agent that autonomously handle calls, WhatsApp, and SMS, plus AI Assist, which delivers real-time coaching, call summaries, and CRM automation for sales and support teams. We are looking for someone that can build out the evaluation foundation across all of these products and other agentic products. You'll work on voice models, agent capability evals, benchmark design, LLM-as-judge systems, failure analysis, and the infrastructure that ties it together by establishing shared metrics, test sets, and tooling to measure accuracy, resolution quality, and safety consistently across products. You will set up repeatable pipelines for regression testing and benchmarking as models and features evolve so teams can ship confidently without re-inventing evaluation methodology for each product.
Base salary range: $181,000 - $250,000 USD
DE&CI Statement: At Aircall, we believe diversity, equity and inclusion – irrespective of origins, identity, background and orientations – are core to our journey. We pride ourselves on promoting active inclusion within our business to foster a strong sense of belonging for all. We’re working to create a place filled with diverse people who can enrich and learn from one another. We’re committed to ensuring that everyone not only has a seat at the table but is valued and respected at it by providing equal opportunities to develop and thrive. We will constantly challenge ourselves to make sure that we live up to our ambitions around diversity, equity and inclusion, and keep this conversation open. Above all else, we understand and acknowl