Senior Applied ML Production Engineer

ByteDance

San Jose (CA)

On-site

USD 162,000 - 388,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ByteDance is seeking an experienced engineer to own production stability and reliability for AML training, inference, and storage systems in our San Jose location. You will drive SLOs, observability, and on-call processes while improving CI/CD, canary releases, and auto-rollback capabilities.

The role emphasizes cross-functional collaboration, performance analysis, and automation platform development to optimize resource usage and overall R&D efficiency.

Qualifications

  • Bachelor’s degree or above in Computer Science, Software Engineering, Artificial Intelligence, or related fields.
  • Familiar with Linux and proficient in at least one language: Shell, Python, Go, or C++.
  • Understanding of machine learning training/inference architectures, Kubernetes, GPU clusters, or distributed storage systems.
  • Proven experience in online troubleshooting, performance analysis, and building automation platforms.
  • Strong sense of responsibility, clear logical thinking, and ability to drive resolution of complex issues across cross-functional teams.

Responsibilities

  • System Stability & Production Management for AML training, inference, and storage systems.
  • Reliability Engineering: SLO/SLA, observability, alerting, on-call processes, fault diagnosis, auto-healing, disaster recovery, post-mortems.
  • Engineering Excellence: CI/CD, canary releases, auto-rollback, automated inspections, pre-flight checks, capacity forecasting, elastic auto-scaling.
  • Resource & Cost Management: governance across GPU/CPU/storage/network, quota management, cost attribution, performance tuning.

Skills

Strong sense of responsibility
Clear logical thinking
Cross-functional collaboration
Troubleshooting

Education

Bachelor’s degree or above in Computer Science, Software Engineering, AI, or related fields

Tools

Linux
Kubernetes
GPU clusters
Distributed storage systems

Job description

ByteDance is seeking an experienced engineer to own production stability and reliability for AML training, inference, and storage systems in our San Jose location. You will drive SLOs, observability, and on-call processes while improving CI/CD, canary releases, and auto-rollback capabilities.

The role emphasizes cross-functional collaboration, performance analysis, and automation platform development to optimize resource usage and overall R&D efficiency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Platform Engineer (Graduate)
ML Platform Engineer (Graduate)

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Graduate ML Systems Engineer - AML Orchestration
Graduate ML Systems Engineer - AML Orchestration

Bytedance • San Jose (CA), Northern (KY)

Hybrid
USD 162,000 - 317,000
Health insurance
401(k) with company match
Parental leave
+6
ML Engineer for AI-native Risk Intelligence
ML Engineer for AI-native Risk Intelligence

ByteDance • San Jose (CA)

On-site
USD 162,000 - 388,000
Medical insurance
401(k) with company match
Parental leave
+2
Lead Engineer, AI Compute Infrastructure
Lead Engineer, AI Compute Infrastructure

ByteDance • San Jose (CA)

On-site
USD 244,800 - 450,000
Day‑one health benefits
401(k) with company match
Parental leave
+1
Graduate ML Platform Engineer – Research & Recommendations
Graduate ML Platform Engineer – Research & Recommendations

Pangle • San Jose (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
AI Risk & Multi-Agent ML Engineer (Real-Time Insights)
AI Risk & Multi-Agent ML Engineer (Real-Time Insights)

Bytedance • San Jose (CA)

Hybrid
USD 150,000 - 230,000
ML Production Engineering Intern - Build Scalable Systems
ML Production Engineering Intern - Build Scalable Systems

ByteDance • San Jose (CA)

On-site
USD 52,000 - 72,000
Health insurance from day one
Wellbeing benefits
Housing allowance for non-remote
Graduate ML Engineer: Scalable AI Infrastructure
Graduate ML Engineer: Scalable AI Infrastructure

ByteDance • San Jose (CA)

On-site
USD 162,000 - 317,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+6
Senior ML Systems Scientist — High-Performance Inference
Senior ML Systems Scientist — High-Performance Inference

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+5
Senior AI Infra Engineer - Developer Tooling
Senior AI Infra Engineer - Developer Tooling

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+8