Post-Training ML Research Engineer - Scale Transformers
Baseten
San Francisco (CA)
On-site
USD 200,000 - 275,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Competitive compensation
100% medical coverage
Generous PTO policy
Paid parental leave
401(k)
Exposure to ML startups
Job summary
A leading AI company in San Francisco seeks a research engineer to enhance post-training ML models. You will develop in-house tools, work with diverse model architectures, and interact with various technologies including Kubernetes and GPU computing. The role offers competitive compensation, comprehensive benefits including medical coverage and PTO, and a collaborative environment to drive innovation in AI.
Qualifications
Deep understanding of ML model training techniques.
Experience with GPU computations and performance profiling.
Willingness to tackle complex problems collaboratively.
Responsibilities
Build in-house tooling for model training efficiency.
Collaborate across technical stacks for systems-level concepts.
Support customers’ post-trained ML models.
Skills
Understanding of modern ML techniques
Advanced experience in PyTorch
Understanding of transformer training parallelism
Profiling distributed GPU programs
Ability to perform roofline analysis
Familiarity with HPC and distributed computing
Solid fundamentals in operating systems
Creativity and problem-solving skills
Tools
PyTorch
TensorFlow
Jax
Kubernetes
Slurm
Dask
Job description
A leading AI company in San Francisco seeks a research engineer to enhance post-training ML models. You will develop in-house tools, work with diverse model architectures, and interact with various technologies including Kubernetes and GPU computing. The role offers competitive compensation, comprehensive benefits including medical coverage and PTO, and a collaborative environment to drive innovation in AI.