Remote AI Model Evaluation Engineer

Intelliswift - An LTTS Company

United States

On-site

USD 95,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Intelliswift - An LTTS Company is seeking a Software Engineer/Research Engineer to design, build, and evaluate pipelines for AI model evaluation and systems in a remote U.S. setting (PST preferred).

The role combines software engineering with AI research, contributing to benchmarks, test suites, and validation workflows for next-generation AI systems. The candidate will work with researchers to diagnose model performance, implement scalable tooling in Python, and ensure high-quality,

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Machine Learning, AI, or a related field.
  • 1-3 years of experience in AI/ML engineering, research engineering, or software engineering.
  • Strong programming skills in Python.
  • Hands-on experience training and evaluating machine learning models.
  • Experience with PyTorch and modern deep learning frameworks.
  • Experience developing benchmarks, validation frameworks, or model evaluation workflows.
  • Strong debugging, analytical, and problem-solving abilities.
  • Demonstrated experience managing technical projects from implementation through validation and delivery.
  • Familiarity with AI-assisted development tools such as Claude Code, Cursor, Codex, or GitHub Copilot.

Responsibilities

  • Design, implement, and maintain evaluation frameworks for AI and machine learning systems.
  • Develop benchmarks, test suites, and validation workflows used to measure model performance.
  • Investigate model behavior, performance discrepancies, and system-level issues through data-driven experimentation.
  • Build scalable Python-based tooling and infrastructure to support AI research and evaluation.
  • Collaborate with researchers and cross-functional teams to improve model quality and performance.
  • Document findings, methodologies, and technical decisions to ensure reproducibility.
  • Contribute to deployment, monitoring, and operational excellence of AI systems.
  • Drive engineering best practices around testing, reliability, and code quality.

Skills

Python
AI/ML engineering
Model evaluation
Debugging
Project delivery

Education

Bachelor's degree in Computer Science or related field

Tools

PyTorch
Docker
Kubernetes
GitHub Copilot
Claude Code
Cursor
Codex

Job description

Intelliswift - An LTTS Company is seeking a Software Engineer/Research Engineer to design, build, and evaluate pipelines for AI model evaluation and systems in a remote U.S. setting (PST preferred).

The role combines software engineering with AI research, contributing to benchmarks, test suites, and validation workflows for next-generation AI systems. The candidate will work with researchers to diagnose model performance, implement scalable tooling in Python, and ensure high-quality,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Architect - Data Science Expert (Remote)
AI Evaluation Architect - Data Science Expert (Remote)

Weekday AI (YC W21) • United States

On-site
Fully remote
Weekly payments
Senior AI Evaluation Engineer - Remote Data Pipelines
Senior AI Evaluation Engineer - Remote Data Pipelines

Socket.dev • United States

On-site
USD 196,000 - 227,000
Remote AI Engineer (Contract) - RL Environments
Remote AI Engineer (Contract) - RL Environments

Appsierra Group • United States

On-site
USD 47,000 - 94,000
Remote AI Evaluation Software Engineer - Freelance
Remote AI Evaluation Software Engineer - Freelance

Feedinkoo • United States

Remote
USD 80,000 - 120,000
Remote AI Engineer: RLHF & Evaluation Pipelines
Remote AI Engineer: RLHF & Evaluation Pipelines

Rex.zone • United States

On-site
USD 100,000 - 140,000
Remote Senior AI Evaluation Engineer
Remote Senior AI Evaluation Engineer

YO AI Labs • Atlanta (GA)

Remote
USD 165,000 - 220,000
Remote Software Engineer, AI Benchmarking & Evaluation
Remote Software Engineer, AI Benchmarking & Evaluation

Epoch AI • United States

Remote
USD 125,000 - 200,000
Comprehensive health insurance
Flexible work environment
Generous paid time off
+1
AI Software Engineer - Code & Model Evaluation (Contractor)
AI Software Engineer - Code & Model Evaluation (Contractor)

Turing • San Francisco (CA)

On-site
USD 83,000 - 138,000
Remote Senior AI Evaluation Engineer – Software Tasks
Remote Senior AI Evaluation Engineer – Software Tasks

YO AI Labs • Miami (FL)

Remote
USD 96,000 - 152,000
Senior Software Engineer - AI Evaluation & System Design
Senior Software Engineer - AI Evaluation & System Design

CodeGeniusRecruit • California Hot Springs (CA)

On-site
USD 120,000 - 200,000