ML Ops Engineer

100 Eli Lilly and Company

South San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Eli Lilly and Company in California seeks an ML Ops Engineer to build and operate platforms for the end-to-end machine learning lifecycle. You will enable reliable model deployment, monitoring, retraining, and reproducibility at scale, collaborating with engineering and scientific teams to deliver production-ready AI capabilities.

You will optimize infrastructure and GPU resources, implement infrastructure-as-code and CI/CD, and ensure production readiness for injectable AI workloads.

Qualifications

  • Bachelor’s degree in a technical field or equivalent experience.
  • Strong Python and ML framework experience required.
  • Experience deploying ML models at scale and operating AI platforms.

Responsibilities

  • Lead the operational lifecycle of ML models, including deployment and monitoring.
  • Operate large-scale inference platforms supporting AI workloads.
  • Ensure production deployment, scaling, and reliability of models.
  • Collaborate with researchers and engineers to integrate ML models into workflows.
  • Automate infrastructure with code and CI/CD practices; document processes.
  • Improve model accuracy and performance through ongoing refinements.

Skills

Python programming
ML frameworks experience
Model deployment
CI/CD automation
Cloud platforms
Collaboration

Education

Bachelor's in Computer Science, Engineering, Statistics

Tools

Docker
Kubernetes
Slurm
Ray
Terraform
Ansible
GitHub Actions

Job description

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work but it's work worth doing. If you're driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. Where AI Meets Medicine: Build the Future of Drug Discovery in the Heart of Silicon Valley! Making medicine that is never been made means doing what's never been done. If you're an engineer, scientist, or builder who thrives on problems no one has solved before, this is your invitation; we want you on the team. We are ready to challenge the status quo and push medicine forward, all in the name of health. Are you up for the challenge? If so, join us!

About the Lilly and NVIDIA Partnership

Lilly and NVIDIA are launching a new AI co-innovation lab in the heart of Silicon Valley - an up-to-$1 billion, multi-year commitment to solve drug discovery's toughest challenges. The lab brings Lilly scientists, technologists, chemists and biologists together with NVIDIA engineers under one roof. Together, we are building purpose-built foundation and frontier AI models trained on Lilly data at scale, tightening the feedback loop between automated wet labs and computational dry labs, designing the next generation of medicines for millions of patients across the globe.

What You'll Be Doing

As an ML Ops Engineer, you build and operate the platforms that run the end-to-end machine learning lifecycle. You enable reliable model deployment, operation, monitoring, retraining, and reproducibility at scale. You optimize infrastructure and GPU resources to support research and discovery workloads. You will work closely with engineering and scientific teams to deliver production-ready AI capabilities.

How You'll Succeed
  • Lead the operational lifecycle of ML models, including deployment, monitoring, and ongoing reliability.
  • Operate and optimize large-scale inference platforms that support scientific discovery and AI workloads.
  • Ensure models can be deployed, scaled, monitored, and maintained in production environments.
  • Test, refine, and improve model accuracy.
  • Work with data scientists, business analysts and partners to integrate ML models into broader strategies.
  • Automate the platform with infrastructure-as-code and CI/CD, and document it well enough that someone else can operate it.
What You Should Bring
  • Strong Python skills and experience working with machine learning frameworks such as PyTorch, JAX, or TensorFlow.
  • Experience deploying, operating, and scaling production machine learning platforms, including model serving, monitoring, and large-scale inference workloads.
  • Experience with MLOps platforms and tools like MLflow, Weights & Biases, KServe, or similar technologies.
  • Proficiency with containerization, orchestration, and distributed compute environments (Docker, Kubernetes, Slurm, Ray).
  • Experience operating large-scale AI platforms that deploy, host, and optimize machine learning models for production use, using technologies such as Triton, vLLM, or TensorRT-LLM.
  • Experience with infrastructure automation and CI/CD practices using tools such as Terraform, Ansible, GitHub Actions, or related.
  • Experience supporting cloud platforms (AWS, Azure, or GCP) and on-premises GPU infrastructure.
  • Knowledge of observability and operational monitoring, including metrics, logging, tracing, and performance tuning.
  • Ability to identify and address system, infrastructure, and model performance issues through automation and continuous improvement.
  • Ability to collaborate effectively with research scientists, AI engineers, and infrastructure teams in a fast-paced environment.
Your Basic Qualifications
  • Bachelor's in Computer Science, Engineering, Statistics,
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Ops Engineer
ML Ops Engineer

BioSpace • San Francisco (CA)

Hybrid
USD 147,000 - 268,000
Bonus program
401(k)
Medical/Dental/Vision
ML Ops Engineer
ML Ops Engineer

Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 147,000 - 268,000
AI Engineer
AI Engineer

100 Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 141,000 - 253,000
401(k)
Pension
Vacation benefits
+1
AI Engineer
AI Engineer

Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 141,000 - 253,000
401(k) plan
Pension
Medical, dental, vision benefits
+1
HPC Systems Administrator
HPC Systems Administrator

Initial Therapeutics, Inc. • San Francisco (CA)

Hybrid
USD 141,000 - 231,000
401(k)
Pension
Vacation benefits
+5
HPC Systems Administrator
HPC Systems Administrator

100 Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 141,000 - 231,000
401(k)
Pension
Vacation benefits
+1
HPC Systems Administrator
HPC Systems Administrator

BioSpace • San Francisco (CA)

Hybrid
USD 150,000 - 230,000
Bonus potential
401(k) plan
Medical benefits
+3
HPC Systems Administrator
HPC Systems Administrator

Eli Lilly and Company • South San Francisco (CA)

Hybrid
USD 141,000 - 231,000
Hybrid work schedule
Comprehensive benefits
Data Engineer
Data Engineer

Eli Lilly and Company • Indianapolis (IN)

Hybrid
CAD 220,000 - 322,000
Data Engineer
Data Engineer

BioSpace • San Francisco (CA)

Hybrid
USD 158,000 - 231,000
Company bonus
401(k)