Machine Learning Research Engineer, Model Evaluation

WindBorne Systems

Atlanta (GA)

Hybrid

USD 140,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401(k)
Dental, health, and vision insurance
Unlimited PTO
Stock Option Plan
Office food and beverages

Job summary

WindBorne Systems is seeking a Machine Learning Research Engineer focused on model evaluation for our global weather forecasting platform. You will design evaluation strategies, run large-scale experiments, and build reusable tooling to quantify forecast quality across regions and lead times.

You will communicate results with clear scorecards and visualizations, collaborating with meteorology and research teams to ensure rigorous, credible assessments of AI and physics-based models.

Qualifications

  • Excellent scientific judgment and healthy skepticism.
  • Strong experimental taste to identify meaningful evaluation.
  • Systems thinking to build reusable infrastructure.
  • Experience evaluating ML systems with large, scientific, geospatial, multidimensional, or time-series data.
  • Able to investigate ambiguous results independently, synthesize evidence, and communicate conclusions clearly.

Responsibilities

  • Evaluation strategy — develop a rigorous, meteorologically valid strategy for comparing WeatherMesh with leading models.
  • Fast feedback and durable systems — build quick evaluations and reusable infrastructure with reproducibility in mind.
  • Agentic tooling for evals — improve evaluation infrastructure, including AI-based tools for forecasts analysis.
  • Technical communication — produce clear scorecards, visualizations, and explanations for researchers and leadership.

Skills

Scientific judgment
Experimental taste
Systems thinking
ML system evaluation
Independent analysis & communication
Weather/Climate context helpful

Tools

PyTorch
NumPy
pandas
xarray

Job description

Machine Learning Research Engineer, Model Evaluation

WindBorne Systems is supercharging weather forecasts with a proprietary data source: a global constellation of next-generation smart weather balloons targeting critical atmospheric data. We design, manufacture, and operate our own balloons, using their observations to generate otherwise unattainable weather intelligence.

Our mission is to eliminate weather uncertainty and help humanity adapt to climate change—whether by predicting hurricanes or speeding the adoption of renewables. The founding team of Stanford engineers was named Forbes 2019 30 Under 30 and is backed by top-tier investors, including Khosla Ventures and Footwork VC.

WindBorne builds AI weather models that run 24/7, producing global forecasts every 20 minutes. Evaluating these models is much harder than producing a headline accuracy number: performance varies across regions, lead times, weather regimes, and customer use cases, while standard metrics often fail to capture what makes a forecast meteorologically sound or useful.

We need someone with excellent scientific taste to determine where our models excel, where they fail, and which results we should trust. You will work at the intersection of machine learning and meteorology, combining fast analyses with robust systems that make future research faster and more reliable.

Responsibilities
  • Evaluation strategy — Work with our Meteorology team to develop a rigorous, meteorologically valid strategy for comparing WeatherMesh with leading AI and physics-based models. Choose the metrics, datasets, baselines, and case studies that provide an honest picture of forecast quality.
  • Fast feedback and durable systems — Build quick evaluations that give researchers useful signals, then turn recurring analyses into reliable, reusable infrastructure. Think systematically about reproducibility, provenance, and how evaluation tools fit into the broader research workflow.
  • Agentic tooling for evals — Improve our existing evaluation infrastructure, including agentic AI-based tools for investigating forecasts and synthesizing results.
  • Technical communication — Produce clear scorecards, visualizations, and explanations for researchers, leadership, customers, and external partners. Communicate model performance precisely, including uncertainty and important caveats.
Requirements
  • Excellent scientific judgment and healthy skepticism. You ask whether a comparison is fair, what else could explain a result, and what evidence would change your mind.
  • Strong experimental taste: you can identify the evaluation that answers the question that matters and distinguish robust improvement from noise.
  • Systems thinking: you can solve today’s problem while recognizing what should become reusable infrastructure for future work.
  • Experience evaluating ML systems using large, scientific, geospatial, multidimensional, or time-series datasets.
  • Strong Python skills and experience with scientific and ML tools such as PyTorch, NumPy, pandas, or xarray.
  • Able to investigate ambiguous results independently, synthesize evidence, and communicate conclusions clearly.
  • Experience with weather, climate, forecasting, physical science, or AI-assisted research tools is helpful, but not required.
Benefits
  • 401(k)
  • Dental, health, and vision insurance
  • Unlimited PTO
  • Stock Option Plan
  • Office food and beverages
Salary
  • $140k–$240k. We consider a range of backgrounds and experience levels and adjust offers to be competitive with market rates.
Location

1600 Bridge Pkwy, Redwood City, CA. Hybrid or in-person.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Research Engineer, Model Evaluation
Machine Learning Research Engineer, Model Evaluation

WindBorne Systems • Redwood City (CA)

Hybrid
USD 140,000 - 240,000
401(k)
Dental, health, and vision insurance
Unlimited PTO
+2
Machine Learning Research Engineer, Model Evaluation
Machine Learning Research Engineer, Model Evaluation

WindBorne Systems • Palo Alto (CA)

On-site
USD 140,000 - 240,000
401(k)
Dental, health, and vision insurance
Unlimited PTO
+2
Machine Learning Research Engineer, Model Evaluation
Machine Learning Research Engineer, Model Evaluation

Convectivecapital • Redwood City (CA)

Hybrid
USD 140,000 - 240,000
401(k)
Dental, health, and vision insurance
Unlimited PTO
+2
Machine Learning Research Engineer
Machine Learning Research Engineer

WindBorne Systems • Redwood City (CA)

Hybrid
USD 140,000 - 240,000
401(k)
Dental, health, and vision insurance
Unlimited PTO
+2
Machine Learning Research Engineer
Machine Learning Research Engineer

Convectivecapital • Redwood City (CA)

Hybrid
USD 140,000 - 240,000
401(k)
Health benefits
Unlimited PTO
+2
Machine Learning Research Engineer, Applied Research
Machine Learning Research Engineer, Applied Research

WindBorne Systems • Atlanta (GA)

Hybrid
USD 140,000 - 240,000
401(k)
Dental insurance
Health insurance
+4
Machine Learning Infrastructure Engineer
Machine Learning Infrastructure Engineer

WindBorne Systems • Palo Alto (CA)

Hybrid
USD 140,000 - 240,000
401(k)
Dental, health, and vision insurance
Unlimited PTO
+2
Machine Learning Infrastructure Engineer
Machine Learning Infrastructure Engineer

WindBorne Systems • Atlanta (GA)

Hybrid
USD 140,000 - 240,000
401(k)
Dental, health, and vision insurance
Unlimited PTO
+2
Machine Learning Research Engineer, Applied Research
Machine Learning Research Engineer, Applied Research

Convectivecapital • Redwood City (CA)

Hybrid
USD 140,000 - 240,000
401(k)
Dental insurance
Health insurance
+4
Machine Learning Research Engineer, Applied Research
Machine Learning Research Engineer, Applied Research

WindBorne Systems • Redwood City (CA)

Hybrid
USD 140,000 - 240,000
401(k)
Dental insurance
Health insurance
+4