Lead AI Evaluation & Safety Architect

BrightClaim

Philippines

On-site

PHP 1,800,000 - 2,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Genpact in the Philippines is seeking a Lead AI Engineer to drive evaluation frameworks for AI models, including performance, safety, and user experience. You will lead large-scale human evaluations, curate test datasets, and translate results into strategic guidance for research, engineering, and product teams.

The role requires expertise in AI/ML evaluation, strong Python and data-analysis skills, and the ability to communicate insights to executives and stakeholders.

Qualifications

  • Experience in AI/ML evaluation, model testing, or related technical roles.
  • Strong understanding of machine learning concepts, model behavior, and evaluation metrics.
  • Experience designing and running human‑in‑the‑loop evaluations or annotation workflows.
  • Proficiency with Python, data analysis tools, and experiment‑tracking frameworks.
  • Demonstrated ability to synthesize complex findings into clear, actionable insights.

Responsibilities

  • Develop comprehensive evaluation frameworks, benchmarks, and success metrics for AI models (text, agents).
  • Define evaluation methodologies for performance, safety, robustness, fairness, and user experience.
  • Build scalable processes for continuous model assessment throughout the development lifecycle.
  • Lead large‑scale human evaluations, including designing tasks, guidelines, and quality controls.
  • Create and maintain high‑quality test datasets, including adversarial, edge‑case, and domain‑specific scenarios.
  • Analyze evaluation results to identify failure patterns, risks, and opportunities for improvement.
  • Translate findings into actionable recommendations for research, engineering, and product teams.
  • Partner with model researchers, data scientists, product managers, and safety teams to align evaluation goals with product requirements.
  • Communicate insights clearly to technical and non‑technical stakeholders, including executives.
  • Influence model development roadmaps based on evaluation outcomes and risk assessments.
  • Ensure evaluations meet internal safety standards and external regulatory expectations.
  • Contribute to the development of responsible AI policies, red‑team strategies, and risk‑mitigation plans.
  • Monitor emerging risks and industry best practices to keep evaluation approaches current.

Skills

AI/ML evaluation
Python
Data analysis
Experiment-tracking
Insight synthesis

Education

Bachelor’s or Master’s degree in Computer Science / ML / Data Science

Job description

Genpact in the Philippines is seeking a Lead AI Engineer to drive evaluation frameworks for AI models, including performance, safety, and user experience. You will lead large-scale human evaluations, curate test datasets, and translate results into strategic guidance for research, engineering, and product teams.

The role requires expertise in AI/ML evaluation, strong Python and data-analysis skills, and the ability to communicate insights to executives and stakeholders.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Engineer: Evaluation & Safety Architect
Lead AI Engineer: Evaluation & Safety Architect

Latitude • Philippines

On-site
PHP 2,000,000 - 3,400,000
Lead AI Engineer 4C
Lead AI Engineer 4C

Latitude • Philippines

On-site
PHP 2,000,000 - 3,400,000
Senior Data Scientist: Lead AI Solutions & Impact
Senior Data Scientist: Lead AI Solutions & Impact

BrightClaim • Philippines

On-site
PHP 1,800,000 - 2,400,000
Lead AI Engineer 4C
Lead AI Engineer 4C

BrightClaim • Philippines

On-site
PHP 1,800,000 - 2,400,000
AI-Driven Python Fullstack Engineer
AI-Driven Python Fullstack Engineer

Genpact Services LLC • Philippines

On-site
PHP 700,000 - 1,100,000
AI Evaluation & Validation Engineer — Test Automation
AI Evaluation & Validation Engineer — Test Automation

NNIT • Philippines

Hybrid
PHP 1,200,000 - 1,800,000
Competitive Compensation
13th Month Pay
Hybrid set up
+2
AI Quality & Evaluation Architect for LLM Systems
AI Quality & Evaluation Architect for LLM Systems

DutchTech • España

Hybrid
PHP 4,291,000 - 6,132,000
Stock options
Work & Swim program in Cyprus
Flexible work model across Europe
Lead QA Automation Engineer - Java/Selenium & AI Testing
Lead QA Automation Engineer - Java/Selenium & AI Testing

Genpact • Makati

On-site
PHP 1,200,000 - 1,800,000
GenAI Engineer: Build & Scale Enterprise AI Solutions
GenAI Engineer: Build & Scale Enterprise AI Solutions

EPAM Systems • Mexico

On-site
International projects with top brands
Paid time off and sick leave
Upskilling and certification courses
+1
AI Engineer: Build Scalable, Production-Ready AI Systems
AI Engineer: Build Scalable, Production-Ready AI Systems

SYSGEN RPO • Philippines

On-site
PHP 1,200,000 - 1,900,000