Senior Software Engineer – Open Source & SWE-Bench Evaluation

anyone-ai

United States

Remote

USD 83,000 - 124,000

Part time

13 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Anyone AI in the United States is seeking experienced Machine Learning Engineers to review and evaluate ML challenges used in model training. You will analyze experiments, datasets, metrics, and pipelines to determine technical soundness and reproducibility.

The role focuses on ensuring challenges reward good ML reasoning, not just brute-force tuning. Remote, part-time, project-based consulting, with a strong emphasis on written feedback and clear recommendations.

Qualifications

  • 3+ years of hands-on applied ML experience.
  • Strong ability to evaluate ML experiments and datasets.
  • Ability to explain findings with written feedback.

Responsibilities

  • Review ML challenges for design quality and solvability.
  • Evaluate datasets for signals and learnability.
  • Identify shortcuts and artifacts in synthetic data.
  • Assess task calibration and difficulty.
  • Review metrics and thresholds for improvements.
  • Detect data leakage and metric gaming.
  • Verify reproducibility across data → model → eval.
  • Provide actionable recommendations to improve tasks.

Skills

Machine learning
Experiment design
Statistical testing
Data analysis
Written feedback

Tools

Python
TensorFlow
PyTorch
Jupyter

Job description

Anyone AI is recruiting experienced Machine Learning Engineers for a specialized project focused on reviewing and evaluating machine learning challenges used in AI model training and evaluation.

The work involves analyzing ML experiments, datasets, metrics, and pipelines to determine whether challenges are technically sound, reproducible, appropriately difficult, and genuinely require strong machine learning reasoning.

What You’ll Work On

You’ll review ML challenges involving:

  • Experiment design and model selection
  • Small and synthetic datasets
  • Data quality and preprocessing
  • Distribution shift and data contamination
  • Label noise and feature leakage
  • Model evaluation and metric selection
  • Hyperparameter tuning
  • Train / validation / test methodology
  • Reproducibility and deterministic pipelines
  • Statistical significance of model improvements

A key part of the role is determining whether a challenge actually rewards good ML reasoning, rather than simply being solvable through brute-force model selection or large hyperparameter searches.

What We’re Looking For
  • 3+ years of hands-on applied machine learning experience
  • Strong experience with:
  • Strong understanding of train, validation, and test splits
  • Ability to identify:
  • Experience evaluating whether performance improvements are statistically meaningful rather than random fluctuations
  • Strong understanding of ML evaluation metrics and when different metrics are appropriate
  • Experience debugging ML workloads across CPU and GPU environments
  • Ability to analyze technical problems and provide clear written feedback
Nice to Have
  • Experience creating or participating in Kaggle, DrivenData, or similar ML competitions
  • Experience designing benchmark datasets or ML challenges
  • Background in data-centric AI or dataset quality
  • Experience with synthetic data generation and validation
  • Familiarity with statistical testing, confidence intervals, and effect sizes
  • Experience with ML evaluation pipelines, RLHF, or AI model evaluation
  • Experience developing ML curricula or technical assessments
  • Understanding of common ML failure modes such as:
What You’ll Be Responsible For
  • Reviewing ML challenges and determining whether they are well designed and technically solvable
  • Evaluating whether datasets contain meaningful and learnable signals
  • Identifying unintended shortcuts or artifacts in synthetic datasets
  • Determining whether tasks require genuine diagnosis of the underlying ML problem
  • Reviewing evaluation metrics and improvement thresholds
  • Detecting metric gaming, data leakage, and evaluation flaws
  • Verifying reproducibility across the complete data → model → evaluation pipeline
  • Assessing whether challenge difficulty is appropriately calibrated
  • Providing clear recommendations for improving, recalibrating, or excluding problematic tasks
Engagement

Work Type: Remote
Engagement: Part‑time, project‑based consulting
Focus: Applied machine learning, experiment design, data quality, and model evaluation

This role is a strong fit for ML engineers who enjoy debugging experiments, understanding why models succeed or fail, identifying problems in datasets and evaluation pipelines, and designing rigorous machine learning experiments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Expert - Fully Remote | Upto $90/hr
Machine Learning Expert - Fully Remote | Upto $90/hr

Obsidian • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
ML Research Engineer - PhD - AI Trainer
ML Research Engineer - PhD - AI Trainer

Obsidian • Seattle (WA)

Hybrid
USD 100,000 - 150,000
Remote ML Evaluation Architect (Part-Time)
Remote ML Evaluation Architect (Part-Time)

anyone-ai • United States

On-site
USD 83,000 - 124,000
ML Engineer Specialist - Freelance AI Trainer Project
ML Engineer Specialist - Freelance AI Trainer Project

Meridial • United States

On-site
Confidential
Flexibility to set your schedule
Remote work environment
Impact on cutting-edge AI
Remote | ML Engineer — $100–$150/hour
Remote | ML Engineer — $100–$150/hour

24-Mag Llc • New York (NY), Northern (KY)

Hybrid
USD 138,000 - 207,000
AI/ ML Engineer
AI/ ML Engineer

Crate and Barrel • Northbrook (IL)

Hybrid
USD 80,000 - 120,000
Senior Software Engineer - 35501
Senior Software Engineer - 35501

Turing • New York (NY)

On-site
USD 68,880 - 206,640
Applied AI Engineer
Applied AI Engineer

SherlockTalent • Miami (FL)

On-site
USD 120,000 - 140,000
Solid Benefits
Referral bonus of $2,500
Machine Learning Engineer | Remote
Machine Learning Engineer | Remote

Crossing Hurdles • United States

Remote
GBP 102,000 - 153,000
Machine Learning Engineering Evaluator
Machine Learning Engineering Evaluator

OpenTrain AI, Inc. • United States

Remote
USD 138,000 - 207,000
Remote work
Flexible schedule
Portfolio building