MTS, Research Engineer

Fireworks AI

San Mateo (CA)

On-site

USD 250,000 - 400,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Fireworks AI in San Mateo, CA seeks a Research Engineer at the intersection of model research and training infrastructure. You will design architectures, reproduce state-of-the-art results, and build distributed training systems to enable scalable AI research.

Responsibilities include implementing new models, optimizing training loops, data pipelines, and coordinating with researchers to translate ideas into robust, efficient code for large GPU clusters.

Qualifications

  • Strong programming skills and writing clean, maintainable code.
  • Deep practical knowledge of ML frameworks (PyTorch, JAX, or TensorFlow).
  • Experience with large distributed systems and parallel computing (CUDA, NCCL, MPI).
  • Strong foundation in linear algebra, calculus, probability, and statistics.
  • Proven track record of implementing complex deep learning algorithms from scratch.

Responsibilities

  • Conduct Open-Ended Research: Explore new model architectures, training objectives, and optimization techniques.
  • Reproduce and Extend State-of-the-Art: Implement and reproduce results from recent ML papers and scale methods.
  • Build and Scale Training Infrastructure: Design, implement, and maintain high-performance distributed ML systems.
  • Bridge Science and Engineering: Translate mathematical concepts into robust, efficient code.
  • Collaborate Cross-Functionally: Work with Research Scientists to unblock experiments with tooling and hardware-aware co-design.

Skills

Python
C++
Rust
PyTorch
JAX
TensorFlow
Distributed systems
CUDA
NCCL
MPI
Linear algebra
Calculus
Probability & statistics
Deep learning algorithms

Education

Master's degree or PhD in Computer Science, Machine Learning, Physics, Mathematics, or related field

Job description

About Us

Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed, NVIDIA, Sequoia Capital, and TCV, Fireworks powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. Today, Fireworks is a Series D company valued at $17.5 billion, bringing together an ambitious, collaborative team that’s building the future of enterprise AI.

Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed, NVIDIA, Sequoia Capital, and TCV, Fireworks powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. Today, Fireworks is a Series D company valued at $17.5 billion, bringing together an ambitious, collaborative team that’s building the future of enterprise AI.

About The Role

We are looking for a Research Engineer to join our team, operating at the critical intersection of model research and training infrastructure.

In this role, your time will be split between tackling open-ended research problems—such as designing novel architectures and improving algorithmic efficiency—in and building the distributed training systems required to make those research breakthroughs a reality. You won’t just be handed a paper to implement; you will be expected to reproduce state-of-the-art results from the literature, identify their limitations, and build the infrastructure needed to push beyond them.

The most significant advances in deep learning require massive scale. We need engineers who are as comfortable reasoning about gradient descent and loss landscapes as they are about distributed systems, GPU cluster utilization, and data pipelines.

What You'll Do
  • Conduct Open-Ended Research: Explore new model architectures, training objectives, and optimization techniques. Formulate hypotheses, design experiments, and iterate quickly based on empirical results.
  • Reproduce and Extend State-of-the-Art: Implement and reproduce results from recent machine learning papers. Identify bottlenecks, propose improvements, and scale these methods to larger datasets and models.
  • Build and Scale Training Infrastructure: Design, implement, and maintain high-performance, distributed machine learning systems. Optimize training loops, data loaders, and communication overhead across large GPU clusters.
  • Bridge Science and Engineering: Translate abstract mathematical concepts and research ideas into robust, bug-free, and efficient code.
  • Collaborate Cross-Functionally: Work closely with Research Scientists to unblock their experiments by providing tooling, optimizing code, and co-designing experiments that are hardware-aware.
We Expect You To Have
  • Strong programming skills (Python, C++, or Rust) and a commitment to writing clean, maintainable code.
  • Deep practical knowledge of machine learning frameworks (PyTorch, JAX, or TensorFlow).
  • Experience working with large distributed systems and parallel computing (e.g., CUDA, NCCL, MPI).
  • A strong foundation in linear algebra, calculus, probability, and statistics.
  • A proven track record of implementing complex deep learning algorithms from scratch.
Nice To Have
  • A Master's or PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related field (or equivalent industry experience).
  • Experience with low-level GPU programming (CUDA/Triton) or hardware co-design.
  • Familiarity with the challenges of training Large Language Models (LLMs)
  • Familiarity with the challenges of inference, and OSS inference engines such as SGLang and vLLM
Why Fireworks?
  • Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving.
  • Build What’s Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally.
  • Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI—no bureaucracy, just results.
  • Learn from the Best: Collaborate with world‑class engineers and AI researchers who thrive on curiosity and innovation.

Fireworks AI is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all innovators.

Compensation Range: $250K - $400K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

MTS, Research Engineer
MTS, Research Engineer

Fireworks • San Mateo (CA)

On-site
USD 140,000 - 200,000
Member of Technical Staff, Research
Member of Technical Staff, Research

Fireworks AI • San Mateo (CA)

On-site
USD 175,000 - 240,000
Meaningful equity in a fast-growing startup
Competitive salary
Comprehensive benefits package
+1
Member of Technical Staff, Cloud Infrastructure
Member of Technical Staff, Cloud Infrastructure

Fireworks AI • San Mateo (CA)

On-site
USD 175,000 - 220,000
Applied Machine Learning Engineer
Applied Machine Learning Engineer

Fireworks AI • New York (NY)

On-site
USD 170,000 - 240,000
Member of Technical Staff, Cloud Infrastructure
Member of Technical Staff, Cloud Infrastructure

Fireworks AI • New York (NY)

On-site
USD 175,000 - 220,000
Software Engineer, LLM Infrastructure
Software Engineer, LLM Infrastructure

Fireworks AI • San Mateo (CA)

On-site
USD 175,000 - 220,000
Applied Machine Learning Engineer
Applied Machine Learning Engineer

Fireworks AI • San Mateo (CA)

On-site
USD 170,000 - 240,000
MTS, Security
MTS, Security

Fireworks AI • San Mateo (CA)

On-site
USD 180,000 - 220,000
Member of Technical Staff, Performance Optimization
Member of Technical Staff, Performance Optimization

Fireworks AI • San Mateo (CA)

On-site
USD 175,000 - 220,000
Competitive compensation
Inclusive environment
Ownership & impact
Member of Technical Staff, Cloud Infrastructure
Member of Technical Staff, Cloud Infrastructure

Fireworks • New York (NY), San Mateo (CA)

On-site
USD 140,000 - 210,000