ML Ops Engineer

Lilly

South San Francisco (CA)

Hybrid

USD 147,000 - 268,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Company bonus
401(k) plan
Health benefits

Job summary

Lilly 在 Silicon Valley 的AI合作实验室正在招聘ML Ops Engineer。您将构建并运营端到端机器学习生命周期的平台,确保模型的可部署、可扩展与可重复生产。与工程与科研团队紧密协作,以实现高效的AI能力落地。

该职位要求扎实的Python与ML框架经验,熟练容器化与CI/CD流程,具备云与本地GPU基础设施经验。工作模式为三天在现场、两天远程,注重成果与协作。

Qualifications

  • Python技能与ML框架(如PyTorch、JAX、TensorFlow)
  • 将生产级ML平台部署、运行与扩展
  • 具备ML Ops工具(MLflow、Weights & Biases、KServe)经验
  • 熟练容器化与编排(Docker、Kubernetes、Slurm、Ray)
  • 熟悉部署、托管与优化大规模AI平台的技术
  • 具备基础自动化与CI/CD工作流(Terraform、Ansible、GitHub Actions)
  • 具备云平台(AWS、Azure、GCP)及本地GPU基础设施经验
  • 具备观测性与监控(指标、日志、追踪、性能调优)
  • 具备通过自动化解决系统、基础设施与模型性能问题的能力
  • 能与研究科学家、AI工程师、基础设施团队协同在快节奏环境中工作

Responsibilities

  • 领导ML模型的部署、监控与可靠性生命周期
  • 运行与优化支撑科研与AI工作负载的大规模推理平台
  • 确保模型可在生产环境中部署、扩展、监控与维护
  • 测试、改进模型精度
  • 与数据科学家和业务分析师合作将ML模型融入更广泛策略
  • 通过基础设施即代码与CI/CD实现平台自动化并撰写清晰文档

Skills

Strong Python skills
Production ML deployment
Collaboration with researchers/enginee
Observability & monitoring
Automation mindset

Education

Bachelor’s in Computer Science/Engineering/Statistics/Mathematics

Tools

MLflow
Weights & Biases
KServe
Docker
Kubernetes
Slurm
Ray
Terraform
Ansible
GitHub Actions
AWS
Azure
GCP
TensorRT-LLM
Triton
vLLM

Job description

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life‑changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.

If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.

Where AI Meets Medicine: Build the Future of Drug Discovery in the Heart of Silicon Valley!

Making medicine that’s never been made means doing what’s never been done. If you’re an engineer, scientist, or builder who thrives on problems no one has solved before, this is your invitation; we want you on the team. We are ready to challenge the status quo and push medicine forward, all in the name of health. Are you up for the challenge? If so, join us!

About the Lilly and NVIDIA Partnership

Lilly and NVIDIA are launching a new AI co-innovation lab in the heart of Silicon Valley — an up-to-$1 billion, multi-year commitment to solve drug discovery’s toughest challenges. The lab brings Lilly scientists, technologists, chemists and biologists together with NVIDIA engineers under one roof. Together, we are building purpose-built foundation and frontier AI models trained on Lilly data at scale, tightening the feedback loop between automated wet labs and computational dry labs, designing the next generation of medicines for millions of patients across the globe.

What You’ll Be Doing

As an ML Ops Engineer, you build and operate the platforms that run the end-to-end machine learning lifecycle. You enable reliable model deployment, operation, monitoring, retraining, and reproducibility at scale. You optimize infrastructure and GPU resources to support research and discovery workloads. You will work closely with engineering and scientific teams to deliver production-ready AI capabilities.

How You’ll Succeed
  • Lead the operational lifecycle of ML models, including deployment, monitoring, and ongoing reliability.

  • Operate and optimize large-scale inference platforms that support scientific discovery and AI workloads.

  • Ensure models can be deployed, scaled, monitored, and maintained in production environments.

  • Test, refine, and improve model accuracy.

  • Work with data scientists, business analysts and partners to integrate ML models into broader strategies.

  • Automate the platform with infrastructure-as-code and CI/CD, and document it well enough that someone else can operate it.

What You Should Bring
  • Strong Python skills and experience working with machine learning frameworks such as PyTorch, JAX, or TensorFlow.

  • Experience deploying, operating, and scaling production machine learning platforms, including model serving, monitoring, and large-scale inference workloads.

  • Experience with MLOps platforms and tools like MLflow, Weights & Biases, KServe, or similar technologies.

  • Proficiency with containerization, orchestration, and distributed compute environments (Docker, Kubernetes, Slurm, Ray).

  • Experience operating large-scale AI platforms that deploy, host, and optimize machine learning models for production use, using technologies such as Triton, vLLM, or TensorRT-LLM.

  • Experience with infrastructure automation and CI/CD practices using tools such as Terraform, Ansible, GitHub Actions, or related.

  • Experience supporting cloud platforms (AWS, Azure, or GCP) and on-premises GPU infrastructure.

  • Knowledge of observability and operational monitoring, including metrics, logging, tracing, and performance tuning.

  • Ability to identify and address system, infrastructure, and model performance issues through automation and continuous improvement.

  • Ability to collaborate effectively with research scientists, AI engineers, and infrastructure teams in a fast-paced environment.

Your Basic Qualifications
  • Bachelor’s in Computer Science, Engineering, Statistics, Mathematics, or a related technical field

  • 4+ years of experience in machine learning engineering, ML Ops, or platform engineering.

Location & Work Flexibility

This role is based at our Silicon Valley Hub. We offer a flexible hybrid work model, with three days onsite and two days working remotely each week, supporting both collaboration and work‑life balance.

Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form (https://careers.lilly.com/us/en/workplace-accommodation) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.

Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.

Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (BE@Lilly), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Women’s Initiative for Leading at Lilly (WILL).

Actual compensation will depend on a candidate’s education, experience, skills, and geographic location. The anticipated wage for this position is

$147,000 - $268,400

Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well‑being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities). Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly’s compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.

#WeAreLilly

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Ops Engineer
ML Ops Engineer

BioSpace • San Francisco (CA)

Hybrid
USD 147,000 - 268,000
Bonus program
401(k)
Medical/Dental/Vision
ML Ops Engineer
ML Ops Engineer

Eli Lilly and Company • Indianapolis (IN)

Hybrid
USD 147,000 - 268,000
Company bonus
401(k) plan
Pension
+2
ML Ops Engineer
ML Ops Engineer

Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 147,000 - 268,000
401(k)
Pension
Vacation benefits
+3
AI Engineer
AI Engineer

BioSpace • South San Francisco (CA)

Hybrid
USD 141,000 - 253,000
HPC Systems Administrator
HPC Systems Administrator

BioSpace • San Francisco (CA)

Hybrid
USD 150,000 - 230,000
Bonus potential
401(k) plan
Medical benefits
+3
HPC Systems Administrator
HPC Systems Administrator

Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 141,000 - 231,000
Hybrid work schedule
Comprehensive benefits
Sr. AI Science Lead
Sr. AI Science Lead

Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 267,000 - 392,000
401(k) matching
Pension plan
Health insurance
+2
HPC Systems Administrator
HPC Systems Administrator

Initial Therapeutics, Inc. • San Francisco (CA)

Hybrid
USD 141,000 - 231,000
401(k)
Pension
Vacation benefits
+5
Sr. AI Science Lead
Sr. AI Science Lead

BioSpace • San Francisco (CA)

Hybrid
USD 260,000 - 381,000
401(k)
Comprehensive health benefits
Vacation & leave benefits
HPC Systems Administrator
HPC Systems Administrator

100 Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 141,000 - 231,000
401(k)
Pension
Vacation benefits
+1