Staff Software Development Test Engineer - AI Evaluation

Tekion Corporation

Bengaluru

On-site

INR 2,500,000 - 4,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tekion Corporation in Bengaluru is seeking a Senior SDET – AI Evaluation to join our AI Platform team. You will build the evaluation capabilities that measure accuracy, consistency, and safety of AI outputs across the organization.

Responsibilities include designing evaluation datasets, automated scoring pipelines, and quality metrics for offline and online evaluation, including LLM‑as‑judge. Strong Python skills and ML evaluation fluency are required.

Qualifications

  • 5–8 years in SDET/quality engineering, ML engineering, or data science with evaluation systems.
  • Strong Python and testing pipelines for robust evaluations.
  • Deep understanding of ML/LLM evaluation, benchmarks, and non‑deterministic systems.
  • Hands‑on with LLM‑as‑judge, rubric scoring, or human‑in‑the‑loop eval.
  • Excellent communication to translate quality signals to ML/product teams.

Responsibilities

  • Define and build Tekion’s AI evaluation platform and shared capabilities.
  • Design evaluation datasets, automated scoring pipelines, and metrics.
  • Develop offline and online evaluation, including CI/CD gates.
  • Create dashboards and reporting to visualize AI quality for teams.
  • Identify hallucinations, bias, safety issues, and edge cases.
  • Collaborate with ML engineers, data scientists, and product managers.

Skills

Python
SDET
ML/LLM evaluation
Data analysis
Communication

Education

Bachelor's degree in Computer Science or related field

Tools

Ragas
LangSmith
TruLens
Promptfoo
HELM

Job description

Staff Software Development Test Engineer - AI Evaluation

Positively disrupting an industry that has not seen any innovation in over 50 years, Tekion has challenged the paradigm with the first and fastest cloud-native automotive platform that includes the revolutionary Automotive Retail Cloud (ARC) for retailers, Automotive Enterprise Cloud (AEC) for manufacturers and other large automotive enterprises and Automotive Partner Cloud (APC) for technology and industry partners. Tekion connects the entire spectrum of the automotive retail ecosystem through one seamless platform. The transformative platform uses cutting-edge technology, big data, machine learning, and AI to seamlessly bring together OEMs, retailers/dealers and consumers. With its highly configurable integration and greater customer engagement capabilities, Tekion is enabling the best automotive retail experiences ever. Tekion employs close to 3,000 people across North America, Asia and Europe.

About the Role

We are looking for a highly motivated Senior SDET – AI Evaluation to join Tekion’s AI Platform team. Evaluation is the backbone of trustworthy AI: as Tekion scales from a handful of AI agents to 100+ across Service, Sales, F&I, and Analytics, this role builds the evaluation platform and frameworks that let every ML team measure, trust, and improve the quality of AI outputs. In this role, you will be responsible for defining and building Tekion’s AI evaluation capabilities as a shared platform service. You will work closely with ML Engineers, Data Scientists, the AI Platform team, and Product Management to design evaluation datasets, automated scoring pipelines, and quality metrics that quantify the accuracy, consistency, and safety of AIgenerated outputs across the organization. You will own the systems that answer “is this model or agent good enough to ship, and is it staying good in production?” — from offline benchmarks and LLM-as-judge pipelines to online evaluation and continuous quality monitoring. You will also use AI and LLMs to scale evaluation itself, building automated judges and synthetic datasets that expand coverage faster than manual review ever could.

Develop a deep understanding of Tekion’s AI agents, ML models, and the quality dimensions that matter for each business domain.

Design, enhance, and own Tekion’s AI evaluation infrastructure as a shared capability used across ML teams.

Create, curate, and maintain evaluation datasets (evals) and golden/ground-truth sets across use cases and domains.

Define quality metrics for AI outputs — accuracy, relevance, faithfulness/groundedness, consistency, safety, and task success.

Build automated scoring pipelines, including LLM-as-judge, rubric-based, and referencebased evaluation methods.

Validate user intents and measure response accuracy and consistency for AI-powered capabilities such as the Analytics Agent.

Identify hallucinations, unsafe or biased outputs, and edge cases; design targeted eval suites to catch them.

Build both offline evaluation (pre-release benchmarking) and online evaluation (production quality monitoring, A/B, drift detection).

Establish evaluation gates in CI/CD so model, prompt, or data changes are quality-checked before release.

Develop dashboards and reporting that make AI quality visible and actionable for ML and product teams.

Use AI/LLMs to scale evaluation — automated judges, synthetic data generation, and eval tooling.

Champion evaluation and responsible‑AI quality best practices across the organization.

What We’re Looking For

5–8 years in SDET, quality engineering, ML engineering, or data science, with hands‑on experience building evaluation or measurement systems — or a strong SDET background with deep LLM/ML fluency.

Strong programming skills in Python, with the ability to build robust, reusable evaluation pipelines and tooling.

Deep understanding of ML/LLM evaluation, benchmark design, and the pitfalls of evaluating non‑deterministic systems.

Hands‑on experience with LLM-as-judge, rubric-based scoring, or human‑in‑the‑loop evaluation.

Solid grasp of LLM/agent concepts — prompting, RAG, embeddings, tool use — and generative failure modes (hallucination, drift, prompt sensitivity, bias).

Experience designing and curating datasets, including labeling/annotation strategy and data quality.

Strong statistical intuition for interpreting eval results and significance.

Excellent communication skills to translate quality signals into decisions for ML and product teams.

Nice to Have
  • Experience with eval frameworks/tools such as Ragas, DeepEval, LangSmith, TruLens, Promptfoo, HELM, or provider eval suites.
  • Experience building online evaluation, guardrails, or production model monitoring.
  • Familiarity with responsible AI / safety evaluation and red‑teaming.
  • Experience with experiment tracking (MLflow, Weights & Biases) and A/B testing.
  • Prior work standing up evaluation as a platform capability for multiple teams

Effective 4 Aug 2026, Current Tekion Employees should apply via the Internal Job Board in Ashby

Tekion is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, gender (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, victim of violence or having a family member who is a victim of violence, the intersectionality of two or more protected categories, or other applicable legally protected characteristics.

For more information on our privacy practices, please refer to our Applicant Privacy Notice here .

Wherever you are in the process, we’re ready to help drive your business forward.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Development Test Engineer - AI Evaluation
Staff Software Development Test Engineer - AI Evaluation

Tekion Corp • Bengaluru

On-site
INR 900,000 - 1,300,000
Senior Software Development Test Engineer - AI
Senior Software Development Test Engineer - AI

Tekion Corp • Bengaluru

On-site
INR 3,200,000 - 5,200,000
Senior Software Development Test Engineer - AI
Senior Software Development Test Engineer - AI

Tekion Corporation • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Senior Software Development Test Enigneer
Senior Software Development Test Enigneer

Tekion Corp • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Opportunity to work with AI experts
Global cloud-native platform
Innovative culture
Senior Software Development Test Enigneer
Senior Software Development Test Enigneer

Tekion Corporation • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Staff Software Development Test Engineer
Staff Software Development Test Engineer

Tekion Corporation • Bengaluru

On-site
INR 2,500,000 - 4,200,000
Senior Software Engineer - Artificial Intelligence Platform
Senior Software Engineer - Artificial Intelligence Platform

Tekion India • Bengaluru

On-site
INR 4,500,000 - 6,500,000
Staff Software Engineer - AI Engineer
Staff Software Engineer - AI Engineer

Tekion Corp • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Staff Software Engineer
Staff Software Engineer

Tekion Corporation • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Tekion Corp • Bengaluru

On-site
INR 4,000,000 - 7,000,000