Tech Lead, Machine Learning Infrastructure Engineer

TikTok USDS Joint Venture

Seattle (WA)

On-site

USD 198,000 - 416,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TikTok USDS Joint Venture LLC is seeking a seasoned software engineer for leading large-scale ML training and inference systems powering TikTok recommendations. You will drive architecture, optimize distributed training across GPUs, and collaborate with research and data teams to deliver scalable, secure platforms.

The role emphasizes real-time ML platforms, multi-node parallelism, and performance profiling across complex GPU clusters, with a strong focus on reliability and compliance in a

Qualifications

  • Bachelor’s or Master's degree in Computer Science, Computer Engineering, or a related technical discipline.
  • 5+ years of professional software engineering experience with deep expertise in Python, C++/Java and a proven track record of designing large-scale distributed systems.
  • 3+ years of direct experience building and maintaining machine learning infrastructure at enterprise scale (managing large GPU clusters, Kubernetes, or native Slurm environments).
  • Deep technical familiarity with the internals of core ML frameworks (PyTorch / Tensorflow) and a strong understanding of low-level GPU memory management, CUDA interactions, and networking topologies (InfiniBand/RoCE).
  • Solid understanding of production-grade LLM training and inference tools, with hands-on profiling skills to eliminate I/O, compute, or network bottlenecks.
  • Strong system-level troubleshooting and debugging skills, with experience profiling and eliminating I/O, compute, or network bottlenecks.

Responsibilities

  • Technical Leadership: Drive the technical roadmap for large-scale distributed real-time ML training and inferencing platforms with business impact.
  • Large-Scale Parallelism Architecture: Architect and scale multi-node distributed training systems and 3D parallelism strategies.
  • Production Inference & Serving: Build low-latency, high-throughput model serving infrastructure for massive live traffic.
  • Algorithm-Infra Co-Design: Partner with Applied ML Research to co-design generative recommendation systems.
  • Resiliency & Fault Tolerance: Design automated fault-detection and checkpointing across thousands of GPUs.
  • Cross-Functional Collaboration: Work with researchers and data teams on high-throughput data processing and storage engines.
  • Security & Compliance Hardening: Enforce encryption, ACLs, and tenant isolation across compute stack.

Skills

Python
C++/Java
Distributed systems
GPU clusters
Kubernetes
Slurm
CUDA
PyTorch/TensorFlow

Education

Bachelor’s or Master’s degree in Computer Science/Engineering

Tools

Megatron
DeepSpeed
CUDA
Triton

Job description

Responsibilities

We are a group of applied machine learning engineers that focus on TikTok recommendations and search engineers powering multiple product areas such as For-You-Page (FYP), Live Streaming, Global E-commerce, Local Services and more. We are developing innovative algorithms and techniques to improve user engagement and satisfaction, converting creative ideas into business-impacting solutions. We are interested in and excited about pushing the envelope of State-of-the-Art (SOTA) large scale machine learning to solve various real-world problems.

What You'll Do
  • Technical Leadership: Drive the technical roadmap for our large-scale (in billions parameters) distributed real-time ML training and inferencing platforms that power the TikTok recommendation and search engines, Short Form Video (SFV) ecosystem. Have a direct business impact on Live, Global E-Commerce, Local Services and many businesses domains.
  • Large-Scale Parallelism Architecture: Architect and scale multi-node distributed training systems, implementing advanced 3D parallelism strategies (Data, Tensor, Pipeline) to maximize compute efficiency and model scalability. Lead the architecture, scale-testing, and maintenance of massive distributed computing foundations, GPU cluster configurations, and orchestration pipelines, establish robust SLI/SLO frameworks while maximizing hardware utilization and cluster efficiency to expedite innovation.
  • Production Inference & Serving: Build and scale low-latency, high-throughput model serving infrastructure, optimizing inference pipelines and leveraging low-level execution paths to handle massive live traffic under strict boundary isolation.
  • Algorithm-Infra Co-Design: Partner closely with Applied ML Research teams to co-design and pioneer next-generation Generative Recommendation systems, abstracting general-purpose components to support advanced generative paradigms in production.
  • Resiliency & Fault Tolerance: Design robust, automated fault-detection systems and asynchronous checkpointing mechanisms to gracefully handle hardware drops, silent data corruption (SDC), or network-switch failures in multi-thousand GPU clusters.
  • Cross-Functional Collaboration: Partner with Applied ML researchers and data platform teams to engineer high-throughput, secure multi-modal data processing and storage engines that prevent compliance friction.
  • Security & Compliance Hardening: Implement and enforce strict encryption-at-rest/in-transit controls, access control lists (ACLs), and secure tenant isolation protocols across the entire compute stack to meet compliance objectives.
Qualifications
Minimum Qualifications
  • Bachelor’s or Master's degree in Computer Science, Computer Engineering, or a related technical discipline.
  • 5+ years of professional software engineering experience with deep expertise in Python, C++/Java and a proven track record of designing large-scale distributed systems.
  • 3+ years of direct experience building and maintaining machine learning infrastructure at enterprise scale (managing large GPU clusters, Kubernetes, or native Slurm environments).
  • Deep technical familiarity with the internals of core ML frameworks (PyTorch / Tensorflow) and a strong understanding of low-level GPU memory management, CUDA interactions, and networking topologies (InfiniBand/RoCE).
  • Solid understanding of production-grade LLM training and inference tools, with hands-on profiling skills to eliminate I/O, compute, or network bottlenecks.
  • Strong system-level troubleshooting and debugging skills, with experience profiling and eliminating I/O, compute, or network bottlenecks.
Preferred Qualifications
  • Experience optimizing high-performance training loops to maximize Model Flops Utilization (MFU) through advanced communication-computation overlap and zero-bubble pipeline scheduling across large-scale distributed clusters using industry-standard frameworks (e.g., Megatron, DeepSpeed).
  • Proven track record of scaling LLM training or inference workloads across hundreds of GPUs, with deep familiarity in advanced serving techniques such as KV Cache management, Prefill-Decoding (PD) separation, and model quantization.
  • Experience working in highly regulated industries, sovereign cloud environments, or dealing with federal data security compliance frameworks.
  • Strong flavor in low-level kernel development and graph compilation, with experience in CUDA, Triton, Cutlass, TensorRT, or Triton Inference Server being a huge plus.
  • Active background or interest in keeping up with the latest industry breakthroughs in MLOps, MoE (Mixture of Experts) routing infrastructure, and specialized hardware optimization.
  • Excellent technical leadership skills, with the ability to mentor junior engineers, draft clear architecture designs, and communicate complex infrastructure needs to non-technical stakeholders.
About USDS

TikTok USDS Joint Venture LLC is dedicated to the safety and security of millions of Americans who create, discover, and connect with what they love on the apps we operate. The Joint Venture has been established in compliance with the Executive Order signed by President Trump on September 25, 2025. Our foundation is a comprehensive data privacy and cybersecurity program we operate under defined safeguards to protect national security and secure U.S. user data, apps and the algorithm. We safeguard the U.S. content ecosystem, holding decision-making authority for trust and safety policies and moderation. USDS Joint Venture helps ensure Americans can continue to express their creativity, discover new hobbies and interests, and build thriving communities and businesses on a global scale.

On-site presence across teams allows the company to operate with greater speed, alignment, and agility — especially in areas like real-time decision-making, team development, and integrated execution. As such, the company is shifting from a hybrid work model to a fully in-person schedule up to 5 days a week.

Why Join Us

Inspiring creativity is at the core of TikTok's mission. Our innovative product is built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and bring joy - a mission we work towards every day.

We strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. Every challenge is an opportunity to learn and innovate as one team. We're resilient and embrace challenges as they come. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our company, and our users. When we create and grow together, the possibilities are limitless. Join us.

Diversity & Inclusion

TikTok is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At TikTok, our mission is to inspire creativity and bring joy. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

USDS Reasonable Accommodation

USDS is committed to providing reasonable accommodations in our recruitment processes for candidates with disabilities, pregnancy, sincerely held religious beliefs or other reasons protected by applicable laws. If you need assistance or a reasonable accommodation, please reach out to us at https://tinyurl.com/USDS-RA

Job Information

【For Pay Transparency】Compensation Description (Annually)

The base salary range for this position in the selected city is $198360 - $416100 annually.

Compensation may vary outside of this range depending on a number of factors, including a candidate’s qualifications, skills, competencies and experience, and location. Base pay is one part of the Total Package that is provided to compensate and recognize employees for their work, and this role may be eligible for additional discretionary bonuses/incentives, and restricted stock units.

Benefits may vary depending on the nature of employment and the country work location. Employees have day one access to medical, dental, and vision insurance, a 401(k) savings plan with company match, paid parental leave, short-term and long-term disability coverage, life insurance, wellbeing benefits, among others. Employees also receive 10 paid holidays per year, 10 paid sick days per year and 17 days of Paid Personal Time (prorated upon hire with increasing accruals by tenure).

The Company reserves the right to modify or change these benefits programs at any time, with or without notice.

For Los Angeles County (unincorporated) Candidates:

Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state, and local laws including the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Our company believes that criminal history may have a direct, adverse and negative relationship on the following job duties, potentially resulting in the withdrawal of the conditional offer of employment:

  1. 1. Interacting and occasionally having unsupervised contact with internal/external clients and/or colleagues;
  2. 2. Appropriately handling and managing confidential information including proprietary and trade secret information and access to information technology systems; and
  3. 3. Exercising sound judgment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech Lead, Machine Learning Infrastructure Engineer - USDS
Tech Lead, Machine Learning Infrastructure Engineer - USDS

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 187,000 - 438,000
Site Reliability Engineer, AI Infrastructure
Site Reliability Engineer, AI Infrastructure

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 123,000 - 259,000
Senior Site Reliability Engineer, Platform Responsibility - USDS
Senior Site Reliability Engineer, Platform Responsibility - USDS

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 178,000 - 342,000
Software Engineer, AI Data Application – USDS
Software Engineer, AI Data Application – USDS

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 137,000 - 360,000
Engineering Manager, Data Platform
Engineering Manager, Data Platform

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 209,000 - 438,000
Senior Site Reliability Engineer, Reliability Team - USDS
Senior Site Reliability Engineer, Reliability Team - USDS

TikTok USDS Joint Venture • San Jose (CA)

On-site
USD 187,000 - 360,000
Health insurance
401(k) with company match
Parental leave
+3
Senior Site Reliability Engineer, AI Infrastructure
Senior Site Reliability Engineer, AI Infrastructure

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 178,000 - 342,000
Senior Site Reliability Engineer, Compute - USDS
Senior Site Reliability Engineer, Compute - USDS

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 178,000 - 342,000
Medical, dental, vision insurance
401(k) with company match
Paid parental leave
+1
Senior Data Scientist, Product & Operation Analytics - USDS
Senior Data Scientist, Product & Operation Analytics - USDS

TikTok USDS Joint Venture • Los Angeles (CA)

On-site
USD 136,800 - 312,867
Machine Learning Engineer, Recommendations - USDS E-commerce Alliance
Machine Learning Engineer, Recommendations - USDS E-commerce Alliance

TikTok USDS Joint Venture • Seattle (WA)

On-site
USD 130,000 - 246,000