C++ Systems Engineer – AI Model Evaluation & Code Review | Remote

Crossing Hurdles

United States

On-site

USD 82,656 - 137,760

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology evaluation firm is seeking software engineering experts with a strong background in C++ and data science. This role involves evaluating LLM-generated content, conducting accuracy assessments, and ensuring alignment with technical standards. Candidates should hold at least a bachelor's degree in computer science and have real-world experience in software engineering. The position offers flexibility between full-time or part-time commitments and can be performed remotely.

Qualifications

  • Real-world experience in software engineering or related technical roles.
  • Ability to solve HackerRank or LeetCode Medium and Hard–level problems independently.
  • Experience contributing to open-source projects.
  • Prior experience with RLHF, model evaluation, or data annotation work preferred.

Responsibilities

  • Evaluate LLM-generated responses for accuracy and completeness.
  • Conduct fact-checking using trusted public sources.
  • Conduct accuracy testing by executing code.
  • Annotate model responses for strengths and improvements.
  • Assess code quality and alignment with guidelines.

Skills

C++ programming language
Problem-solving (HackerRank, LeetCode)
Attention to detail
Complex technical reasoning evaluation
Experience with LLMs

Education

BS, MS, or PhD in Computer Science

Job description

Position: Software Engineering, Data Science, and Systems Design Experts, C++ (5+ YOE)
Type: Hourly contract
Compensation: $60-$100 per hour
Location: Remote
Commitment: Full-time or Part-time Contract Work
Role Responsibilities
  • Evaluate LLM-generated responses to coding and software engineering queries for accuracy, reasoning, clarity, and completeness
  • Conduct fact-checking using trusted public sources and authoritative references
  • Conduct accuracy testing by executing code and validating outputs using appropriate tools
  • Annotate model responses by identifying strengths, areas of improvement, and factual or conceptual inaccuracies
  • Assess code quality, readability, algorithmic soundness, and explanation quality
  • Ensure model responses align with expected conversational behavior and system guidelines
  • Apply consistent evaluation standards by following taxonomies, benchmarks, and detailed evaluation guidelines
Requirements
  • BS, MS, or PhD in Computer Science or a closely related field
  • Real-world experience in software engineering or related technical roles
  • Expertise in C++ programming language
  • Ability to solve HackerRank or LeetCode Medium and Hard–level problems independently
  • Experience contributing to open-source projects, including merged pull requests
  • Experience using LLMs while coding and understanding their strengths and failure modes
  • Strong attention to detail and ability to evaluate complex technical reasoning and identify bugs or logical flaws
  • Prior experience with RLHF, model evaluation, or data annotation work preferred
  • Track record in competitive programming preferred
  • Experience reviewing code in production environments preferred
  • Familiarity with multiple programming paradigms or ecosystems preferred
  • Experience explaining complex technical concepts to non-expert audiences preferred
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Embedded / Systems Engineer (C) – AI Code Analysis | Remote
Embedded / Systems Engineer (C) – AI Code Analysis | Remote

Crossing Hurdles • United States

Remote
Backend / NET Engineer (C#) – AI Systems & Code Quality Evaluation | Remote
Backend / NET Engineer (C#) – AI Systems & Code Quality Evaluation | Remote

Crossing Hurdles • United States

Remote
iOS / Swift Developer (AI Model Evaluation & Code Quality) | Remote
iOS / Swift Developer (AI Model Evaluation & Code Quality) | Remote

Crossing Hurdles • United States

Remote
C++ Systems Engineer for AI Model Evaluation & Code Review
C++ Systems Engineer for AI Model Evaluation & Code Review

Crossing Hurdles • United States

On-site
USD 60,000 - 80,000
Language Model Evaluator - Fully Remote | Upto $20/hr Part-time
Language Model Evaluator - Fully Remote | Upto $20/hr Part-time

United States Digital Space LLC • United States

Remote
USD 21,000 - 28,000
Software Engineer (Python/C++) – Code Evaluation | Remote
Software Engineer (Python/C++) – Code Evaluation | Remote

Crossing Hurdles • United States

Remote
GBP 25,000 - 41,000
Mobile Software Engineer – Swift & AI Evaluation | Remote
Mobile Software Engineer – Swift & AI Evaluation | Remote

Crossing Hurdles • United States

Remote
LLM Evaluator (Prompt Engineer) | $65/hr Remote
LLM Evaluator (Prompt Engineer) | $65/hr Remote

Crossing Hurdles • United States

On-site
USD 70,000 - 110,000
Remote, Hourly C++ Systems Engineer for AI Evaluation
Remote, Hourly C++ Systems Engineer for AI Evaluation

SME Careers • Idaho

On-site
Senior Software Engineer - 35501
Senior Software Engineer - 35501

Turing • New York (NY)

Remote