AI Evaluation Engineer

Wipro Technologies

San Diego (CA)

On-site

USD 60,000 - 149,000

Full time

35 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical and dental benefits
Disability insurance
Paid time off

Job summary

Wipro Limited is seeking an engineer to help build and scale tools that measure AI-powered software development performance. The role involves creating evaluation harnesses, automating benchmarks, and ensuring reproducible results across environments.

You will work with engineering and data teams to document findings for technical and leadership audiences. The position emphasizes hands-on automation, versioned workflows, and calibration against human judgments, with exposure to Docker, Git, and

Qualifications

  • Strong software engineering background with automation and tooling experience.
  • Experience evaluating AI coding agents and calibrating against human judgments.
  • Experience building reproducible evaluation workflows (test execution, env setup, validation).
  • Familiar with Git, CI/CD, and containerized workloads.

Responsibilities

  • Build and integrate evaluation harnesses and automation for software dev use cases.
  • Create versioned, repeatable evaluation processes with pinned deps and containerized runs.
  • Validate evaluation approaches against human judgment for accuracy.
  • Support benchmarking across quality, productivity, and cost metrics.
  • Analyze results for variance, failures, and cost per outcome; improve workflows.
  • Collaborate with engineering and data teams to document evaluations for technical and leadership audiences.

Skills

Software engineering
Automation
Dev tooling
Test & validation
AI evaluation concepts
LLM as a judge
Python/Java/JS
Git & CI/CD
Troubleshooting
AI coding tools

Tools

Docker
Git
Claude Code
Devin
Cursor

Job description

Wipro Limited (NYSE: WIT, BSE: 507685, NSE: WIPRO) is a leading technology services and consulting company focused on building innovative solutions that address clients’ most complex digital transformation needs. Leveraging our holistic portfolio of capabilities in consulting, design, engineering, and operations, we help clients realize their boldest ambitions and build future-ready, sustainable businesses. With over 230,000 employees and business partners across 65 countries, we deliver on the promise of helping our customers, colleagues, and communities thrive in an ever-changing world. For additional information, visit us at www.wipro.com.

Job Description
Role Overview

Help build and scale the tooling we use to measure how well AI-powered software development tools actually perform. You'll develop evaluation harnesses, automate benchmark runs, and help make sure the results we produce are reproducible and hold up to scrutiny. This is an engineering role, but a lot of the work is about getting the measurement right, not just automating it.

Key Responsibilities
  • Build and integrate evaluation harnesses and automation for software development use cases, including turning real engineering artifacts like merged pull requests into repeatable benchmark tasks.
  • Build versioned, repeatable processes to evaluate AI tools, models, and harnesses, with reproducible run environments (pinned dependencies, containerized runs, isolated worktrees) so results stay comparable over time.
  • Validate and calibrate evaluation approaches against human judgment, so scores are consistent and correct rather than just repeatable.
  • Support execution-based benchmarking across quality, productivity, and e iciency measures, including cost and latency.
  • Analyze results across repeated runs, looking at variance, failure patterns, and cost per outcome, and find ways to make the workflows more reliable and more automated.
  • Work with engineering and data teams to improve the tooling, and document how the evaluations work and what they found for both technical and leadership audiences.
Required Skills & Experience
  • Strong software engineering background, with real experience building automation, developer tooling, or test and validation systems.
  • Coding-agent evaluation: Evaluating AI coding agents that modify code repositories, including validating generated code changes against expected outcomes.
  • Evaluation harnesses: Building automated, reproducible evaluation workflows, including test execution, environment setup, and result validation.
  • Benchmarking and reliability: Establishing baselines, measuring run-to-run variance, analyzing failures, and ensuring consistent evaluation results.
  • Git and CI/CD integration: Working with repository history, branches, pull requests, and automated testing within CI/CD pipelines.
  • LLM-as-a-Judge: Using LLMs to evaluate coding-agent outputs, including calibration against human assessments.
  • Proficient in at least one general-purpose language such as Python, Java, or JavaScript — the specific language background is flexible.
  • Solid working knowledge of Git, including how branches, history, and working trees behave, and of containerization with Docker.
  • Experience with APIs, development environments, CI/CD pipelines, and typical engineering workflows.
  • Understanding of how AI, LLM, or agent evaluation works and where it goes wrong, such as why a judge can be consistent but still wrong, why a single run can mislead, and how benchmark contamination happens.
  • Able to troubleshoot technical problems, think clearly about whether a measurement is valid, and analyze results carefully.
  • Hands-on experience using AI coding tools and agentic harnesses such as Claude Code, Devin, or Cursor, and command of the best practices for working with them effectively.

Mandatory Skills: Cloud Product & Platform Testing.Experience: 5-8 Years.The expected compensation for this role ranges from $60,000 to $148,500 .Final compensation will depend on various factors, including your geographical location, minimum wage obligations, skills, and relevant experience. Based on the position, the role is also eligible for Wipro's standard benefits including a full range of medical and dental benefits options, disability insurance, paid time off (inclusive of sick leave), other paid and unpaid leave options.Applicants are advised that employment in some roles may be conditioned on successful completion of a post-offer drug screening, subject to applicable state law.

Wipro provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws. Applications from veterans and people with disabilities are explicitly welcome.

If you encounter any suspicious mail, advertisements, or persons who offer jobs at Wipro, please email us at helpdesk.recruitment@wipro.com . Do not email your resume to this ID as it is not monitored for resumes and career applications.

Any complaints or concerns regarding unethical/unfair hiring practices should be directed to our Ombuds Group at ombuds.person@wipro.com .

We are an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, caste, creed, religion, gender, marital status, age, ethnic and national origin, gender identity, gender expression, sexual orientation, political orientation, disability status, protected veteran status, or any other characteristic protected by law.

Wipro is committed to creating an accessible, supportive, and inclusive workplace. Reasonable accommodation will be provided to all applicants including persons with disabilities, throughout the recruitment and selection process. Accommodations must be communicated in advance of the application, where possible, and will be reviewed on an individual basis. Wipro provides equal opportunities to all and values diversity.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer — Benchmarking Automation Lead
AI Evaluation Engineer — Benchmarking Automation Lead

Wipro Technologies • San Diego (CA)

On-site
USD 60,000 - 149,000
Medical and dental benefits
Disability insurance
Paid time off
AI LEAD (CONTRACT)
AI LEAD (CONTRACT)

Wipro Technologies • Tampa (FL)

On-site
USD 80,000 - 158,000
AI LEAD L1
AI LEAD L1

Wipro Technologies • Bolingbrook (IL)

On-site
USD 60,000 - 135,000
AI/ML App Testing
AI/ML App Testing

Wipro Technologies • Marlborough (MA)

On-site
USD 80,000 - 158,000
Medical benefits
Dental benefits
Paid time off
AI Researcher Lead 1
AI Researcher Lead 1

Wipro Technologies • San Francisco (CA)

On-site
USD 240,000 - 375,000
Medical and dental benefits
Disability insurance
Paid time off
Domain AI Researcher
Domain AI Researcher

Wipro Technologies • San Francisco (CA)

On-site
USD 240,000 - 413,000
Medical benefits
Dental benefits
Paid time off
AI Engineer
AI Engineer

Wipro Technologies • Austin (TX)

On-site
USD 60,000 - 135,000
Medical & dental
Disability insurance
Paid time off
+1
Software Engineer - AI Eval & Automation
Software Engineer - AI Eval & Automation

ServiceNow • San Diego (CA)

On-site
USD 150,000 - 190,000
Magnit Global benefits
AI LEAD L1
AI LEAD L1

Wipro • Tampa (FL)

On-site
USD 60,000 - 135,000
Medical and dental benefits
Disability insurance
Paid time off
Industry Cloud & Digital (ICD) - AI Leader
Industry Cloud & Digital (ICD) - AI Leader

Wipro Technologies • Atlanta (GA)

On-site
USD 200,000 - 280,000
Medical benefits
Dental benefits
Paid time off
+1