Senior Machine Learning Engineer, ML Infrastructure- Online

LE130 Unity Technologies SF

United States

À distance

USD 210 000 - 273 000

Plein temps

Il y a 3 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Une candidature complète en une minute — un CV et une lettre de motivation personnalisés, prêts à envoyer.

Passez les filtres ATS

Avantages offerts par ce poste

Health insurance
Stock options
Retirement plans
Commuting subsidy
Wellbeing programs

Résumé du poste

Unity Technologies seeks a Senior ML Engineer to design and evolve Unity Vector's online model inference platform. You will build reliable infrastructure for serving ML models in production, optimize inference performance, and enable safe experimentation across high-traffic systems.

You will partner with ML engineers and platform teams to deploy, monitor, and iterate models, shaping packaging, serving, validation, and automated rollback across teams.

Qualifications

  • Strong experience building production ML inference systems.
  • Experience with Kubernetes and cloud-native infrastructure.
  • Proficiency in Python for production-grade ML platforms.

Responsabilités

  • Design and operate large-scale online inference infrastructure with low latency.
  • Develop infrastructure for distributed training workflows.
  • Integrate ML pipelines with workflow orchestration systems.

Connaissances

Python
Distributed systems
Kubernetes
GKE
PyTorch
Ray

Outils

NVIDIA Triton
TorchServe
Ray Serve
TensorFlow Serving

Description du poste

The Role

We are seeking a Senior ML engineer to design and evolve Unity Vector's online model inference platform. This role focuses on building reliable infrastructure for serving machine learning models in production, optimizing inference performance, and enabling safe, efficient experimentation across high-traffic online systems. You will work closely with ML engineers, platform teams, and product stakeholders to ensure models can be deployed, scaled, monitored, and iterated on efficiently. You will play a key role in shaping how models are packaged, served, validated, monitored, and optimized in production environments. This role requires strong systems thinking, deep experience with production ML infrastructure, and the ability to drive architectural improvements across teams.

What you'll be doing
  • Design and operate large-scale online inference infrastructure that serves production ML models with low latency and high reliability, such as PyTorch, Triton Inference Server, Kubernetes, GKE, Ray, or similar distributed serving frameworks.
  • Develop infrastructure that supports distributed training workflows using technologies such as Pytorch, Ray Data, and Ray Train, etc.
  • Integrate ML pipelines with workflow orchestration systems (e.g., Flyte, Airflow, or similar) to enable reliable multi-stage training workflows.
  • Optimize model performance through model compilation, GPU/CPU utilization improvements, request scheduling, kernel fusion, and runtime-level tuning.
  • Improve observability of ML systems through latency, throughput, error-rate, cost, saturation, and model-health monitoring.
  • Partner closely with ML engineers to support faster model iteration while maintaining production safety, scalability, and cost efficiency.
  • Improve the reliability and reproducibility of model serving workflows, including model packaging, artifact validation, compatibility testing, and deployment automation.
  • Lead architectural improvements that make the online ML platform more robust, user-friendly, scalable, and cost-efficient.
What we’re looking for
  • Experience building and operating production-grade online ML inference systems, such as NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or similar systems.
  • Experience with model serving frameworks such as NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or similar systems.
  • Experience optimizing inference workloads using techniques such as dynamic batching, model compilation, quantization, GPU acceleration, GPU kernel optimization, caching, or runtime tuning.
  • Strong experience with distributed systems, Kubernetes, autoscaling, service reliability, and production observability.
  • Strong programming skills in Python, with practical experience working on production ML systems and high-scale services.
  • Experience with PyTorch and modern model deployment workflows, including model packaging, validation, and serving lifecycle management.
  • Experience designing infrastructure for safe model rollout, canary testing, A/B experimentation, and automated rollback.
  • Strong systems thinking, with the ability to reason about latency, throughput, reliability, scalability, and cost tradeoffs in online systems.
  • Proven ability to lead technical direction and influence architectural decisions across teams without formal authority.

Additional information

Zone A: $210,300 - $273,400

Zone B: $187,200 - $243,300

Zone C: $165,600 - $215,200

This range reflects the anticipated base salary for this position. Beyond base salary, this role may be eligible for equity awards and participation in our company incentive plans (such as annual discretionary bonuses or sales commissions). The final offer amount will depend on several factors, including geographic location and the candidate's relevant experience, professional background, and skill set.

Benefits
  • Comprehensive health, life, and disability insurance
  • Commute subsidy
  • Employee stock ownership
  • Competitive retirement/pension plans
  • Generous vacation and personal days
  • Support for new parents through leave and family-care programs
  • Office food snacks
  • Mental Health and Wellbeing programs and support
  • Employee Resource Groups
  • Global Employee Assistance Program
  • Training and development programs
  • Volunteering and donation matching program
Life at Unity

Unity [NYSE: U] is the world's leading game engine, powering play for more than 3 billion consumers each month. The top mobile games in the world, the most played PC indie titles, the most innovative console games, and virtually all of the top XR and Web Games are developed, deployed, and grown in Unity. Unity also enables teams across industries like automotive, manufacturing, and healthcare to design, simulate, and collaborate in 3D - closing the gap between ideas and reality. For more information, please visit www.unity.com.

Unity is a proud equal opportunity employer. We are committed to fostering an inclusive, innovative environment and celebrate our employees across age, race, color, ancestry, national origin, religion, disability, sex, gender identity or expression, sexual orientation, or any other protected status in accordance with applicable law. Our differences are strengths that enable us to support the growing and evolving needs of our customers, partners, and collaborators.

If you have a disability that means there are preparations or accommodations we can make to help ensure you have a comfortable and positive interview experience, please fill out this form to let us know.

This position requires the incumbent to have a sufficient knowledge of English to have professional verbal and written exchanges in this language since the performance of the duties related to this position requires frequent and regular communication with colleagues and partners located worldwide and whose common language is English.

Culture & Values

At Unity, we're committed to creating a workplace that fosters collaboration and teamwork that allows you to harness your unique skills. Join us in creating industry tools to help creators of all levels bring their projects to reality. Learn more about our culture and values, and get a head start by reading about how our hiring process works.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Senior Machine Learning Engineer, ML Infrastructure- Online
Senior Machine Learning Engineer, ML Infrastructure- Online

Unity • Washington

Hybride
USD 187 000 - 243 000
Health insurance
Stock ownership
Retirement plan
+3
Senior Machine Learning Engineer, ML Infrastructure- Online
Senior Machine Learning Engineer, ML Infrastructure- Online

Unity • États-Unis

À distance
USD 187 000 - 260 000
Health insurance
Stock options
Retirement plans
+1
Senior Machine Learning Engineer, ML Infrastructure- Online
Senior Machine Learning Engineer, ML Infrastructure- Online

Unity Enterprise • Northern (KY)

Hybride
USD 210 000 - 273 000
Comprehensive health insurance
Employee stock ownership
Generous vacation and personal days
+1
Staff Machine Learning Engineer
Staff Machine Learning Engineer

LE130 Unity Technologies SF • États-Unis

À distance
USD 218 000 - 284 000
Health and life insurance
Commuter subsidy
Employee stock ownership
+8
Senior Machine Learning Engineer, Data Infrastructure
Senior Machine Learning Engineer, Data Infrastructure

LE130 Unity Technologies SF • Mountain View (CA)

Sur place
USD 210 000 - 273 000
Health insurance
Commuter subsidy
Stock options
+8
Staff Backend Engineer, Vector AI
Staff Backend Engineer, Vector AI

LE130 Unity Technologies SF • Mountain View (CA)

Sur place
USD 193 000 - 318 000
Health insurance
Stock ownership
Retirement plans
+8
Senior Machine Learning Engineer, Data Infrastructure
Senior Machine Learning Engineer, Data Infrastructure

Unity • Mountain View (CA)

Sur place
USD 166 000 - 273 000
Health insurance
Stock options
Retirement plans
+3
Senior Machine Learning Engineer, Data Infrastructure
Senior Machine Learning Engineer, Data Infrastructure

Unity • Bellevue (WA)

Sur place
USD 165 000 - 273 000
Staff Machine Learning Engineer, ML Infrastructure
Staff Machine Learning Engineer, ML Infrastructure

Unity Technologies • Bellevue (WA)

Sur place
USD 210 000 - 284 000
Health insurance
Stock options
Retirement plans
+1
Staff Machine Learning Engineer
Staff Machine Learning Engineer

Unity • California

Hybride
USD 218 000 - 284 000
Health insurance
Life insurance
Stock ownership
+6