Machine Learning Engineer – ML Evaluation & Experiment Design

Anyone AI Inc.

Argentina

A distancia

ARS 1.200.000 - 2.000.000

A tiempo parcial

hace 47 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo — un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Remote work
Part-time project-based consulting

Descripción de la vacante

Anyone AI Inc. is seeking experienced Machine Learning Engineers to review and evaluate ML challenges used in AI model training and evaluation. You will analyze experiments, datasets, metrics, and pipelines to determine technical soundness and reproducibility.

The role focuses on ensuring challenges reward strong ML reasoning rather than brute-force model tuning, with a remote, project-based, part-time engagement.

Formación

  • 3+ years hands-on applied machine learning experience.
  • Strong ability to analyze ML challenges, datasets, and evaluation pipelines.
  • Experience identifying data leakage, distribution shift, and spurious correlations.
  • Ability to provide clear, written feedback on technical problems.

Responsabilidades

  • Review ML challenges to ensure they are well designed and solvable.
  • Evaluate whether datasets contain meaningful signals and learnable patterns.
  • Identify shortcuts or artifacts in synthetic datasets.
  • Verify reproducibility across the data–model–evaluation pipeline.
  • Provide actionable recommendations to improve or recalibrate tasks.

Conocimientos

Applied ML experience
Data preprocessing
ML evaluation metrics
Experiment design
Debugging experiments

Descripción del empleo

Anyone AI is recruiting experienced Machine Learning Engineers for a specialized project focused on reviewing and evaluating machine learning challenges used in AI model training and evaluation.

The work involves analyzing ML experiments, datasets, metrics, and pipelines to determine whether challenges are technically sound, reproducible, appropriately difficult, and genuinely require strong machine learning reasoning.

What You’ll Work On

You’ll review ML challenges involving:

  • Small and synthetic datasets
  • Data quality and preprocessing
  • Distribution shift and data contamination
  • Label noise and feature leakage
  • Model evaluation and metric selection
  • Train / validation / test methodology
  • Reproducibility and deterministic pipelines
  • Statistical significance of model improvements

A key part of the role is determining whether a challenge actually rewards good ML reasoning, rather than simply being solvable through brute-force model selection or large hyperparameter searches.

What We’re Looking For

3+ years of hands-on applied machine learning experience

Strong experience with:

  • Data preprocessing and validation
  • Strong understanding of train, validation, and test splits
  • Ability to identify:
  • Data leakage
  • Distribution shift
  • Spurious correlations
  • Feature leakage
  • Experience evaluating whether performance improvements are statistically meaningful rather than random fluctuations
  • Strong understanding of ML evaluation metrics and when different metrics are appropriate
  • Experience debugging ML workloads across CPU and GPU environments
  • Ability to analyze technical problems and provide clear written feedback
Nice to Have
  • Experience creating or participating in Kaggle, DrivenData, or similar ML competitions
  • Experience designing benchmark datasets or ML challenges
  • Background in data-centric AI or dataset quality
  • Experience with synthetic data generation and validation
  • Familiarity with statistical testing, confidence intervals, and effect sizes
  • Experience with ML evaluation pipelines, RLHF, or AI model evaluation
  • Experience developing ML curricula or technical assessments
  • Understanding of common ML failure modes such as:
  • Shortcut learning
  • Spurious correlations
  • Metric gaming
What You’ll Be Responsible For
  • Reviewing ML challenges and determining whether they are well designed and technically solvable
  • Evaluating whether datasets contain meaningful and learnable signals
  • Identifying unintended shortcuts or artifacts in synthetic datasets
  • Determining whether tasks require genuine diagnosis of the underlying ML problem
  • Reviewing evaluation metrics and improvement thresholds
  • Detecting metric gaming, data leakage, and evaluation flaws
  • Verifying reproducibility across the complete data - model - evaluation pipeline
  • Assessing whether challenge difficulty is appropriately calibrated
  • Providing clear recommendations for improving, recalibrating, or excluding problematic tasks
Engagement

Work Type: Remote

Engagement: Part-time, project-based consulting

Focus: Applied machine learning, experiment design, data quality, and model evaluation

This role is a strong fit for ML engineers who enjoy debugging experiments, understanding why models succeed or fail, identifying problems in datasets and evaluation pipelines, and designing rigorous machine learning experiments.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Remote ML Engineer — Challenge Evaluation & Data Quality
Remote ML Engineer — Challenge Evaluation & Data Quality

Anyone AI Inc. • Argentina

A distancia
ARS 1.200.000 - 2.000.000
Remote work
Part-time project-based consulting
Applied AI/ML Engineer
Applied AI/ML Engineer

FutureProofing • Buenos Aires

A distancia
ARS 150.912.000 - 226.367.000
Senior Software Engineer, AI Training - Argentina
Senior Software Engineer, AI Training - Argentina

G2i Inc. • Argentina

A distancia
ARS 209.054.000 - 418.107.000
Software Engineering AI Trainer
Software Engineering AI Trainer

Chromestarchemicals • Comisión de Fomento de Perú

Presencial
ARS 125.533.000 - 251.066.000
Sr. Machine Learning Engineer
Sr. Machine Learning Engineer

Promtior • Buenos Aires

Presencial
ARS 135.892.000 - 196.289.000
20 paid days off per year
Hybrid work model
Flex Days: monthly team activities
+2
AI Quality Engineer
AI Quality Engineer

Movéo Technologies Corporation • Rosario

Presencial
ARS 1.800.000 - 3.000.000
Python Engineer (Remote)
Python Engineer (Remote)

Hired • Argentina

Presencial
ARS 135.892.000 - 226.487.000
Senior AI Engineer
Senior AI Engineer

MAS Global Consulting • Argentina

Presencial
ARS 83.351.856 - 111.135.808
Principal AI Engineer
Principal AI Engineer

Creative Chaos • Argentina

Presencial
ARS 1.800.000 - 3.000.000
Lead AI/ML Engineer
Lead AI/ML Engineer

N-iX • Argentina

Presencial
ARS 83.351.856 - 111.135.808
Flexible working format
Competitive salary
Education reimbursement
+1