Member of ML Technical Staff

Pragmatike

San Francisco (CA)

On-site

USD 200,000 - 350,000

Full time

38 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Pragmatike in San Francisco is seeking a highly technical Member of Technical Staff to advance LLM research and large-scale training infrastructures.

You will design and optimize pre-training, post-training pipelines, and distributed GPU training across clusters, collaborating across research and engineering to push state-of-the-art models.

Qualifications

  • Strong technical depth in ML/AI research and development.
  • Experience with LLMs beyond API usage.
  • Background in large-scale training or research infrastructure.

Responsibilities

  • Research and implement techniques for training and improving large language models.
  • Build and optimize pre-training and post-training pipelines.
  • Work on distributed training across large GPU clusters.
  • Design and optimize model-parallel and data-parallel training strategies.

Skills

Python
PyTorch
CUDA
Triton

Education

Bachelor's degree in CS/Engineering/Math

Tools

JAX

Job description

Member of Technical Staff — LLM Research & Training
About the Role

We are looking for an exceptional Member of Technical Staff specializing in Machine Learning and Large Language Models to join an early-stage AI company building and training state-of-the-art foundation models.

This role sits at the intersection of LLM research, large-scale training infrastructure, post-training, and GPU/kernel optimization.

We are particularly interested in highly motivated researchers and engineers who want to contribute directly to training powerful models — whether their strengths are in theoretical model research, training systems, distributed infrastructure, or low-level performance optimization.

You will work in a small, highly technical team where researchers and engineers collaborate closely and are expected to take ownership across the stack.

Responsibilities
  • Research, design, and implement new techniques for training and improving large language models.
  • Build and optimize large-scale pre-training and post-training pipelines.
  • Improve model training efficiency, throughput, stability, and scalability.
  • Work on distributed training across large GPU clusters.
  • Design and optimize model-parallel training strategies, including tensor, pipeline, sequence, and data parallelism.
  • Optimize GPU workloads using technologies such as CUDA and Triton.
  • Improve inference and training kernels when necessary.
  • Explore new model architectures, training methodologies, and post-training techniques.
  • Run experiments, analyze results, and rapidly iterate on research ideas.
  • Collaborate on software/hardware co-design to maximize training throughput.
  • Contribute to internal research infrastructure and potentially open-source initiatives.
What We're Looking For
LLM / ML Research Experience
  • At least 1+ years of experience in theoretical LLM research or as an ML researcher/engineer at a highly technical AI or technology organization.
  • Hands-on experience working with large language models beyond simply consuming existing APIs.
  • Experience with one or more of:
    • LLM architecture research
    • Pre-training
    • Post-training
    • Reinforcement learning / preference optimization
    • Training framework development
    • Kernel or inference optimization
    • Large-scale distributed training

Experience working on language models at organizations or research environments comparable to OpenAI, Google DeepMind, Mistral AI, Qwen, DeepSeek, Z.ai, Allen Institute for AI, or leading academic labs is highly relevant.

Large-Scale Training

Strong understanding of large-scale AI infrastructure and at least some of the following:

  • Distributed GPU training
  • Model parallelism
  • Tensor parallelism
  • Pipeline parallelism
  • Sequence parallelism
  • Data parallelism
  • Communication optimization
  • Memory optimization
  • Training throughput optimization
  • Software/hardware co-design

Experience contributing to initiatives such as NanoGPT Speedrun, Marin, or similar open-source model-training projects is a strong plus.

Technical Skills

Strong proficiency with:

  • Python
  • PyTorch
  • CUDA
  • Triton

Experience with JAX is highly valued.

Additional experience with distributed training frameworks, custom kernels, GPU profiling, compiler optimization, or high-performance computing is a plus.

Research Background

We value candidates who have demonstrated strong technical depth through one or more of:

  • ML/AI research during undergraduate, master's, or PhD studies
  • Publications or meaningful research contributions
  • Open-source ML contributions
  • Competitive programming
  • Building large-scale ML systems from first principles

A strong undergraduate degree is expected, ideally from a highly selective technical university. Advanced degrees are welcome but not required.

What Makes Someone Successful Here

You are likely to thrive in this role if you:

  • Have extremely strong technical fundamentals.
  • Are genuinely interested in understanding how modern language models work internally.
  • Prefer building and improving models rather than simply applying existing LLMs to business use cases.
  • Are comfortable moving between research and engineering.
  • Have high energy, intellectual curiosity, and low ego.
  • Enjoy working in small, fast-moving teams.
  • Are comfortable tackling problems that do not yet have established solutions.
  • Can independently turn research ideas into working systems and experiments.
Nice to Have
  • Experience at an early-stage AI startup.
  • Contributions to open-source ML frameworks or research projects.
  • Experience optimizing GPU kernels or inference engines.
  • Experience building training infrastructure from scratch.
  • Experience training models across large GPU clusters.
  • Strong systems engineering or HPC background.
Not a Fit If

This role is probably not the right fit if your experience is primarily:

  • Integrating existing LLM APIs into applications.
  • Building RAG or chatbot applications without working on the underlying models.
  • Prompt engineering without model training experience.
  • Working exclusively in large, highly structured engineering organizations with narrowly defined responsibilities.
Location

San Francisco, CA

This is an on-site position, 5 days per week, based in San Francisco's Financial District.

Visa Sponsorship

Visa transfers may be supported, including candidates currently on statuses such as OPT or H-1B, depending on individual circumstances.

Compensation

Base Salary: $200,000 – $350,000

Plus competitive equity.

Compensation will depend on experience, technical depth, research background, and expected impact.

Compensation Range: $200K - $350K

Hiring Plan

We are looking to hire multiple exceptional engineers and researchers for this team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Researcher
Machine Learning Researcher

Multicoin • San Francisco (CA)

On-site
USD 250,000 - 350,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Machine Learning Researcher
Machine Learning Researcher

SOLANA FOUNDATION • San Francisco (CA)

Hybrid
USD 250,000 - 350,000
Equity in a high-growth startup
Comprehensive benefits
Machine Learning Engineer, LLM Post-Training
Machine Learning Engineer, LLM Post-Training

GoTo Meeting • Mountain View (CA)

On-site
USD 150,000 - 230,000
Health, dental, and vision care for you and your family
Top-tier 401(K) plan with company matching
Paid time off and paid holidays
+2
Senior Machine Learning Engineer (LLMs)
Senior Machine Learning Engineer (LLMs)

Albiware Inc. • Chicago (IL)

On-site
USD 140,000 - 210,000
Competitive salary
Generous PTO
Medical, dental, and vision coverage
+2
LLM Training & Model Development Engineer
LLM Training & Model Development Engineer

InOpTra Digital • United States

Remote
USD 90,000 - 120,000
Competitive salary
Opportunity for remote work
Health benefits
Senior Machine Learning Engineer (LLMs)
Senior Machine Learning Engineer (LLMs)

Albi • Chicago (IL)

On-site
USD 150,000 - 210,000
Competitive salary
Generous PTO
Medical, dental, and vision coverage
+6
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Take2 Consulting, LLC • San Jose (CA)

Hybrid
USD 175,000 - 185,000
Hybrid work environment
Continuous learning and professional development opportunities
Competitive compensation package
Senior/Principal Local LLM & Generative AI Platform Engineer
Senior/Principal Local LLM & Generative AI Platform Engineer

Parallel Wireless • Northern (KY)

Hybrid
USD 150,000 - 210,000
Machine Learning Engineer
Machine Learning Engineer

HUG • New York (NY)

On-site
USD 175,000 - 300,000
Competitive compensation
Excellent benefits
Career growth
Senior/Principal Local LLM & Generative AI Platform Engineer
Senior/Principal Local LLM & Generative AI Platform Engineer

Parallel Wireless • United States

On-site
USD 180,000 - 280,000