A complete application in a minute — tailored resume and cover letter, ready to send.
Tekion Corp in Bengaluru is seeking a Staff SDET – AI Evaluation to join Tekion’s AI Platform team. You will define and build Tekion’s AI evaluation capabilities as a shared platform service, collaborating with ML Engineers, Data Scientists, and Product Management to design eval datasets and quality metrics.
You will own offline benchmarks, online evaluation, CI/CD gates, and dashboards that make AI quality visible for ML and product teams, while scaling evaluation with automated judges and
Positively disrupting an industry that has not seen any innovation in over 50 years, Tekion has challenged the paradigm with the first and fastest cloud‑native automotive platform that includes the revolutionary Automotive Retail Cloud (ARC) for retailers, Automotive Enterprise Cloud (AEC) for manufacturers and other large automotive enterprises and Automotive Partner Cloud (APC) for technology and industry partners. Tekion connects the entire spectrum of the automotive retail ecosystem through one seamless platform. The transformative platform uses cutting‑edge technology, big data, machine learning, and AI to seamlessly bring together OEMs, retailers/dealers and consumers. With its highly configurable integration and greater customer engagement capabilities, Tekion is enabling the best automotive retail experiences ever. Tekion employs close to 3,000 people across North America, Asia and Europe.
We are looking for a highly motivated Staff SDET – AI Evaluation to join Tekion’s AI Platform team. Evaluation is the backbone of trustworthy AI: as Tekion scales from a handful of AI agents
to 100+ across Service, Sales, F&I, and Analytics, this role builds the evaluation platform and frameworks that let every ML team measure, trust, and improve the quality of AI outputs.In this role, you will be responsible for defining and building Tekion’s AI evaluation capabilities as a shared platform service. You will work closely with ML Engineers, Data Scientists, the AI Platform team, and Product Management to design evaluation datasets, automated scoring pipelines, and quality metrics that quantify the accuracy, consistency, and safety of AIgenerated outputs across the organization. You will own the systems that answer “is this model or agent good enough to ship, and is it
staying good in production?” — from offline benchmarks and LLM-as-judge pipelines to online evaluation and continuous quality monitoring. You will also use AI and LLMs to scale evaluation itself, building automated judges and synthetic datasets that expand coverage faster than manual review ever could.