AI Evaluation Engineer — Benchmarking Automation Lead

Wipro Technologies

San Diego (CA)

On-site

USD 60,000 - 149,000

Full time

33 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical and dental benefits
Disability insurance
Paid time off

Job summary

Wipro Limited is seeking an engineer to help build and scale tools that measure AI-powered software development performance. The role involves creating evaluation harnesses, automating benchmarks, and ensuring reproducible results across environments.

You will work with engineering and data teams to document findings for technical and leadership audiences. The position emphasizes hands-on automation, versioned workflows, and calibration against human judgments, with exposure to Docker, Git, and

Qualifications

  • Strong software engineering background with automation and tooling experience.
  • Experience evaluating AI coding agents and calibrating against human judgments.
  • Experience building reproducible evaluation workflows (test execution, env setup, validation).
  • Familiar with Git, CI/CD, and containerized workloads.

Responsibilities

  • Build and integrate evaluation harnesses and automation for software dev use cases.
  • Create versioned, repeatable evaluation processes with pinned deps and containerized runs.
  • Validate evaluation approaches against human judgment for accuracy.
  • Support benchmarking across quality, productivity, and cost metrics.
  • Analyze results for variance, failures, and cost per outcome; improve workflows.
  • Collaborate with engineering and data teams to document evaluations for technical and leadership audiences.

Skills

Software engineering
Automation
Dev tooling
Test & validation
AI evaluation concepts
LLM as a judge
Python/Java/JS
Git & CI/CD
Troubleshooting
AI coding tools

Tools

Docker
Git
Claude Code
Devin
Cursor

Job description

Wipro Limited is seeking an engineer to help build and scale tools that measure AI-powered software development performance. The role involves creating evaluation harnesses, automating benchmarks, and ensuring reproducible results across environments.

You will work with engineering and data teams to document findings for technical and leadership audiences. The position emphasizes hands-on automation, versioned workflows, and calibration against human judgments, with exposure to Docker, Git, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer
AI Evaluation Engineer

Wipro Technologies • San Diego (CA)

On-site
USD 60,000 - 149,000
Medical and dental benefits
Disability insurance
Paid time off
AI Benchmarking & Automation Engineer
AI Benchmarking & Automation Engineer

ServiceNow • San Diego (CA)

On-site
USD 150,000 - 190,000
Magnit Global benefits
AI Tool Evaluation Engineer: Benchmarking & Automation
AI Tool Evaluation Engineer: Benchmarking & Automation

Akraya, Inc. • San Diego (CA)

On-site
USD 90,000 - 101,000
AI Benchmarking & Performance Architect
AI Benchmarking & Performance Architect

CoreWeave • Sunnyvale (CA)

On-site
USD 206,000 - 333,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+2
Senior AI/ML Evaluation Engineer — Benchmarks (Remote)
Senior AI/ML Evaluation Engineer — Benchmarks (Remote)

OpenTeams • Washington, Denver (CO), Colorado Springs (CO)

Hybrid
USD 145,000 - 250,000
401(k) Match – Up to 5% with full vest
Unlimited PTO – 15 days minimum
Fully Remote Setup – up to $3,000 for
+3
AI Evaluation Engineer: Benchmarking & Automation
AI Evaluation Engineer: Benchmarking & Automation

Akraya, Inc. • San Diego (CA)

On-site
USD 90,000 - 101,000
AI Software Engineer - Evaluation & Benchmarks
AI Software Engineer - Evaluation & Benchmarks

Hire Feed • San Francisco (CA)

On-site
USD 120,000 - 190,000
AI Eval & Benchmarking Manager - 8-Engineer Team Lead
AI Eval & Benchmarking Manager - 8-Engineer Team Lead

ServiceNow • California (MO)

On-site
USD 166,500 - 291,400
Senior AI Benchmarking & Performance Engineer
Senior AI Benchmarking & Performance Engineer

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with generous employer match
Tuition Reimbursement
+2
AI Benchmarking & Strategy Lead
AI Benchmarking & Strategy Lead

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Competitive compensation