Machine Learning Engineering Manager

GoFundMe

San Francisco (CA)

Hybrid

USD 210,000 - 270,000

Full time

48 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Fitness and wellness stipend
Health insurance options
Weekly massages
Standing desks
Fully-stocked kitchens and lunches
Team off-sites and social events
Dog-friendly offices
Parental leave enhancements
Paid holidays and vacation time
Commuter benefits

Job summary

GoFundMe is seeking a Manager, Machine Learning Engineering to lead the ML/AI operations team responsible for infrastructure, pipelines, and observability in production. You will ensure reliability, scalability, and safety of ML systems and own training pipelines, feature stores, and model serving.

You will hire and mentor engineers, define technical direction, collaborate with data science, product, design, and legal teams, and manage vendor relationships.

Qualifications

  • 5+ years leading engineers in ML/AI ops or infrastructure.
  • Familiarity with generative AI/LLM infrastructure and safety considerations.
  • Advanced degree preferred in CS/Statistics/Data Science.
  • Proven experience implementing ML monitoring and production deployment.
  • Experience designing real-time model serving with scalable infra.

Responsibilities

  • Lead and grow a team of ML/AI operations engineers.
  • Own reliability, scalability, and observability of ML/AI production systems.
  • Partner with data science to streamline model development to production.
  • Drive CI/CD, packaging, versioning, and safe rollback strategies.
  • Establish on-call processes, SLOs/SLAs, and incident response.

Skills

Team leadership
ML monitoring
Production ML
Python programming
PyTorch
TensorFlow
Scikit-learn
Data engineering
SQL
Spark/Databricks
Snowflake
CI/CD
API design
Problem solving

Education

Advanced degree in CS/Stats/Data Science

Tools

PyTorch
TensorFlow
Scikit-learn
Spark
Databricks
Snowflake

Job description

  • Join GoFundMe as our next Manager, Machine Learning Engineering (ML and AI Operations)
  • In this role, you will lead the team responsible for the infrastructure, pipelines, and operational rigor that keep GoFundMe’s machine learning and AI systems reliable, scalable, and safe in production
  • This role requires strong technical judgment across the ML lifecycle (data → training → online inference → monitoring), a strong understanding of how to enable AI applications to operate safely at scale, and a proven ability to build and lead a high performance team that operates production ML/AI systems with the same rigor as core infrastructure
  • Own the reliability, scalability, and operational health of ML/AI production systems across GoFundMe, including training pipelines, feature stores, model serving, and monitoring/observability infrastructure
  • Lead, hire, and grow a team of ML/AI operations engineers, setting technical direction through design reviews, architecture decisions, and shared best practices for production ML and AI systems
  • Partner with data science and ML engineering teams to streamline the path from model development to production deployment, including CI/CD for ML, model packaging, versioning, and rollback strategies
  • Establish ML operational excellence org-wide by driving standards for model observability (latency, errors, drift, calibration, business KPI deltas), automated retraining triggers, and incident response playbooks
  • Build and mature on-call processes, SLOs/SLAs, and postmortem practices for ML/AI systems, treating model incidents with the same discipline as production infrastructure incidents
  • Drive operational strategy for GoFundMe’s generative AI systems alongside traditional ML, balancing innovation velocity with safety, compliance, cost, and reliability
  • Collaborate cross-functionally with Product, Engineering, Design, and Legal/Privacy stakeholders to translate business goals into team priorities and measurable operational outcomes
  • Manage vendor and platform relationships (e.g., cloud ML platforms, LLM providers) and make build-vs-buy calls that balance cost, control, and speed
  • Report on team health, system reliability metrics, and operational risk to senior engineering leadership
  • Employ a diverse set of tools and platforms, including Python, AWS, Databricks, Docker, Kubernetes, Terraform, Snowflake, and GitHub, to guide your team in developing, deploying, and maintaining scalable and robust machine learning systems
Benefits
  • $600 annual fitness and wellness reimbursement
  • Wide range of health insurance options, including medical, dental, and vision (GoFundMe covers 100% of employee premiums, and 80% of spouse and dependents)
  • Weekly massages
  • Standing desks
  • Fully-stocked kitchens & daily lunches
  • Team off-sites & monthly social events
  • Many of our offices are dog friendly
  • Enhanced parental leaves
  • 10 paid holidays, 17 days of accrued vacation per year, unlimited sick time & three volunteer days
  • Caltrain GoPasses for our Bay Area commuters
  • $50/month for employees commuting to and from work (public transit and/or parking)
  • Quarterly volunteer events in each office to give back to our local communities
  • “Gives Back” program, where employees nominate fundraisers weekly for donations from GoFundMe
  • 401(k) retirement plan with company matching
  • Access to learning tools and resources, including a subscription to Udemy, guest speakers, and internal brown bag sessions
  • 1-3+ years of experience directly managing engineers, ideally in an MLOps, ML platform, or infrastructure context, with a track record of hiring and developing strong teams
  • Familiarity with generative AI/LLM infrastructure and operational considerations (latency, cost, safety guardrails) is a strong plus
  • Advanced degree (Master’s or Ph.D.) in Computer Science, Statistics, Data Science, or a related technical field is preferred
  • Proven experience implementing ML monitoring for both technical and business metrics (drift, calibration, segment performance, latency, error budgets) and running models reliably in production
  • Sense of humor is optional but appreciated
  • Experience designing and operating real-time model serving at scale, including containerization, scalable inference, feature retrieval, and safe rollout strategies (canaries, shadowing, backward-compatible schema evolution)
  • Strong proficiency in Python and ML libraries/frameworks such as PyTorch, TensorFlow, Scikit-learn, plus strong software engineering fundamentals (testing, code review, CI/CD, API design, performance, and reliability) — enough depth to stay hands-on and credible with your team
  • Strong data engineering fluency: building reliable datasets and features using SQL, Spark/Databricks, and warehouse technologies (e.g., Snowflake), with an understanding of event semantics, identity resolution, and data quality controls
  • Strong leadership and mentoring skills and a proven ability to raise the bar on architecture, engineering quality, and operational rigor for production ML/AI systems
  • 7+ years of hands-on experience building and shipping production machine learning systems, with demonstrated ownership of backend services and ML pipelines in a high-availability environment
  • Ability to break down ambiguous, high-impact problems, define crisp interfaces and success metrics, and deliver iteratively while managing stakeholder expectations across engineering leadership, product, and data science
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manager, Machine Learning Engineering New San Francisco, CA
Manager, Machine Learning Engineering New San Francisco, CA

GoFundMe • San Francisco (CA), Northern (KY)

Hybrid
USD 219,000 - 329,000
Equity
Healthcare
Dental
+3
Manager, Machine Learning Engineering
Manager, Machine Learning Engineering

GoFundMe • San Francisco (CA)

On-site
USD 219,000 - 329,000
Equity
Healthcare benefits
Dental benefits
+2
Senior Manager, Machine Learning Engineering
Senior Manager, Machine Learning Engineering

GoFundMe • San Francisco (CA)

On-site
USD 225,000 - 337,000
Competitive pay
Comprehensive healthcare benefits
Financial assistance for hybrid work
+3
Manager of Data Science (Marketing)
Manager of Data Science (Marketing)

GoFundMe • San Francisco (CA)

On-site
USD 180,000 - 240,000
Fitness & wellness reimbursement
Health insurance including premiums
Weekly massages
+10
ML Ops Engineer — Agentic AI Lab (Founding Team)
ML Ops Engineer — Agentic AI Lab (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Meaningful equity
Senior Staff MLOps Engineer
Senior Staff MLOps Engineer

Grindr • United States

On-site
USD 180,000 - 240,000
Premium health benefits
401k Matching
Flexible PTO
+7
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Harnham • Tampa (FL)

On-site
USD 150,000 - 200,000
Staff MLOps Engineer – ML Platform
Staff MLOps Engineer – ML Platform

BrightAI Corporation • Palo Alto (CA)

On-site
USD 150,000 - 190,000
Machine Learning Engineer
Machine Learning Engineer

Errgo • Town of Boston (NY)

Hybrid
USD 120,000 - 160,000
Medical, dental, and vision insurance
401(k)
Equity
+2
Senior Software Engineer, Machine Learning Infrastructure (Tinder LLC, West Hollywood, California)
Senior Software Engineer, Machine Learning Infrastructure (Tinder LLC, West Hollywood, California)

Match Group • West Hollywood (CA)

On-site
USD 190,000 - 246,000