ML Tech Lead Manager: LLM/TPU Optimization & Leadership

Socket.dev

Sunnyvale (CA)

On-site

USD 215,000 - 292,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Benefits
Bonus target

Job summary

Socket.dev is seeking a Software Engineering Manager focused on ML/LLM systems to lead multiple teams, shaping strategy and delivery on Google Cloud. You will drive bring-up and performance tuning for LLM/Non-LLM workloads on TPUs, manage up to six engineers as a TLM, and guide model onboarding and optimization efforts across the org.

Responsibilities include implementing advanced sharding, TorchTPU support, disaggregated serving, RL inference, and close collaboration with ML research and

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of software development experience.
  • 5 years of experience with one or more ML areas (speech/audio, reinforcement learning, ML infrastructure, or related field).
  • 5 years of experience leading ML design and optimizing ML infrastructure (model deployment, evaluation, data processing, debugging, fine tuning).
  • 3 years of experience in a technical leadership role.
  • 2 years of experience in people management or team leadership.

Responsibilities

  • Work with LLM/Non-LLM models bringup to performance tuning/optimization on Google Cloud TPUs. Manage up to 6 engineers to drive and deliver results as a TLM.
  • Manage onboarding emerging open-weights models in industry.
  • Drive performance optimization for massive-scale models by engineering advanced sharding strategies (2D/3D parallelism), developing custom Pallas kernels, and maximizing utilization across single-host and multi-host TPU.
  • Design and implement key features such as TorchTPU support, disaggregated serving, RL inference, etc.
  • Collaborate with the ML research, ML performance, model optimization tooling, and other optimization teams. Participate in ML Perf Inference submissions.

Skills

Software development
ML design leadership
Technical leadership
People management
Cross-functional collaboration

Education

Bachelor’s degree or equivalent practical experience
Master’s degree or PhD in Engineering/CS or related field

Job description

Socket.dev is seeking a Software Engineering Manager focused on ML/LLM systems to lead multiple teams, shaping strategy and delivery on Google Cloud. You will drive bring-up and performance tuning for LLM/Non-LLM workloads on TPUs, manage up to six engineers as a TLM, and guide model onboarding and optimization efforts across the org.

Responsibilities include implementing advanced sharding, TorchTPU support, disaggregated serving, RL inference, and close collaboration with ML research and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech Lead Manager, TPU Inference at Scale
Tech Lead Manager, TPU Inference at Scale

Socket.dev • Sunnyvale (CA)

On-site
USD 215,000 - 292,000
Equity
Benefits
Bonus target
Lead ML Performance Engineering Manager
Lead ML Performance Engineering Manager

Google • Kirkland (WA)

On-site
USD 207,000 - 300,000
Health insurance
401(k) with company match
Paid Time Off: 20 days per year
+4
Tech Lead Manager, TPU Inference at Scale
Tech Lead Manager, TPU Inference at Scale

Google Inc. • Sunnyvale (CA)

Hybrid
USD 207,000 - 300,000
Senior TPU Performance Co-Design Engineer (LLM Serving)
Senior TPU Performance Co-Design Engineer (LLM Serving)

Google • Sunnyvale (CA)

On-site
USD 174,000 - 252,000
Technical Lead Manager, Scalable ML Runtime — Edge & Cloud
Technical Lead Manager, Scalable ML Runtime — Edge & Cloud

Waymo • Mountain View (CA)

On-site
USD 251,000 - 310,000
Senior TPU Co-Design & Performance Engineer - LLM Serving
Senior TPU Co-Design & Performance Engineer - LLM Serving

Google • United States

On-site
USD 174,000 - 252,000
Health insurance
Engineering Manager, ML Performance
Engineering Manager, ML Performance

Google • Kirkland (WA)

On-site
USD 207,000 - 300,000
Health insurance
401(k) with company match
Paid Time Off: 20 days per year
+4
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Software Development Manager: LLM Inference Enablement
Software Development Manager: LLM Inference Enablement

Amazon • Cupertino (CA)

On-site
USD 213,000 - 288,000
RSUs (Restricted Stock Units)
Health insurance
401(k) plan
Tech Lead Manager for Scalable LLM Training Platform
Tech Lead Manager for Scalable LLM Training Platform

Scale AI • San Francisco (CA)

On-site
USD 120,000 - 160,000