Technical AI Evaluation Specialist

Innodata India

United States

Remote

USD 90,000 - 130,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Innodata India is seeking Technical AI Evaluation Specialists to benchmark frontier LLMs for code generation, multi-turn reasoning, and software architecture. You will assess model-generated code, UI fidelity, and SxS comparisons against detailed rubrics to identify failures and edge cases.

Applicants should have 2+ years of software development, strong English, and experience with front-end technologies (HTML5, CSS, React/TypeScript).

Qualifications

  • 2+ years of professional software development experience.
  • Experience translating designs into responsive front-end code (HTML5, CSS, React, TypeScript).
  • Experience with AI model evaluation/RLHF pipelines is a plus.
  • Proficient written and spoken English (C1).

Responsibilities

  • Evaluate AI-generated code for correctness, complexity, and security.
  • Assess image-to-code and UI verification from wireframes and designs.
  • Perform pairwise (SxS) analyses using multi-dimensional rubrics.
  • Identify failure modes, edge cases, and susceptibilities to vulnerabilities.

Skills

Software development
Front-End/UI
English proficiency (C1)
RLHF
Code evaluation

Education

Computer Science / Software Engineering degree

Tools

Git
Docker
CI/CD pipelines
AST parsers
Static code analysis tools

Job description

ROLE OVERVIEW & OBJECTIVE:

We are seeking rigorous Technical AI Evaluation Specialists to benchmark, evaluate, and align frontier Large Language Models (LLMs) specialized in code generation, multi-turn technical reasoning, and software architecture. In this role, you will analyze model-generated

code against strict correctness, complexity, security, and UI fidelity standards across text-to-code, image-to-code, and side-by-side (SxS)

comparison tasks. You will be responsible for uncovering subtle failure modes, edge-case vulnerabilities, and producing evidence-based,

defensible rationales to guide model fine-tuning and Reinforcement Learning from Human Feedback (RLHF).

KEY RESPONSIBILITIES & CORE WORKFLOWS
  • Model-Generated Code Evaluation: Evaluate AI-generated code for syntactical validity, execution accuracy, algorithmic complexity,

and architectural best practices across diverse languages.

  • Image-to-Code & UI Verification: Assess model capability in rendering pixel-perfect, responsive front-end components from

wireframes, mockups, and UI design screenshots.

  • Text-to-Code & Pairwise Analysis: Perform rigorous side-by-side (SxS) evaluations to determine model preference, scoring

completions against granular multi-dimensional rubrics.

  • Failure Mode & Edge Case Reasoning: Stress-test model completions against corner cases, boundary conditions, race conditions,

memory leaks, and input sanitization vulnerabilities.

  • Evidence-Based Rationales: Write authoritative, structured C1-level technical rationales explaining exact point deductions, execution

trace errors, and counterfactual fixes.

  • Guideline Calibration & Feedback: Collaborate with research engineers and prompt authors to refine evaluation rubrics, establish

baseline test harnesses, and identify emerging model degradation patterns.

Mandatory Requirements:
  • Experience: 2+ years of professional software development

experience OR a strong Computer Science / Software Engineering

degree with demonstrable coding proficiency.

  • Front-End / UI Exposure: Hands-on experience translating

design mockups to responsive front-end code (HTML5, modern

CSS, React, or TypeScript frameworks).

  • Code Review & RLHF Exposure: Proven background in

structured code reviews, automated testing, or prior experience in

AI model evaluation/RLHF pipelines.

  • Language & Communication: C1-equivalent professional English

proficiency with exceptional technical articulation and defensive

writing discipline.

Preferred Qualifications:
  • Multi-Language Fluency: Strong capability in Python,

JavaScript/TypeScript, Java, C++, or Go, with an aptitude for

rapidly reading and debugging unfamiliar frameworks.

  • DevOps & Tooling: Familiarity with Git, Docker, CI/CD pipelines,

AST parsers, or static code analysis tools.

  • Advanced CS Foundations: Deep understanding of data

structures, algorithms, concurrency, and secure coding practices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer - FL
AI Engineer - FL

LawPro.ai • United States

On-site
USD 140,000 - 230,000
Senior LLM Evaluation Engineer
Senior LLM Evaluation Engineer

Aspire, Jordan • Egypt (PA)

On-site
USD 140,000 - 200,000
AI Engineer - OH
AI Engineer - OH

LawPro.ai • Kentucky

On-site
USD 150,000 - 190,000
AI Engineer - TX
AI Engineer - TX

LawPro.ai • Town of Texas (WI)

On-site
USD 140,000 - 210,000
AI Engineer - NC
AI Engineer - NC

LawPro.ai • North Carolina

On-site
USD 140,000 - 190,000
AI Engineer - VA
AI Engineer - VA

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000
AI Engineer - FL
AI Engineer - FL

LawPro.ai • Town of Florida (NY)

On-site
USD 140,000 - 210,000
AI Engineer - GA
AI Engineer - GA

LawPro.ai • Georgia

On-site
USD 140,000 - 210,000
Principal AI Engineer
Principal AI Engineer

Stellantis Financial Services • Auburn Hills (MI)

On-site
USD 180,000 - 240,000
Principal AI Engineer
Principal AI Engineer

Stellantis • Auburn Hills (MI)

On-site
USD 180,000 - 280,000