ML Systems & Performance Engineer

Engram

San Francisco (CA)

On-site

USD 180,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive cash compensation
Startup equity

Job summary

Engram is seeking an ML Systems Engineer to design, optimize, and scale training and inference workloads across GPUs and distributed systems. You’ll collaborate with researchers and performance engineers to bridge cutting‑edge AI research with production systems.

The role requires a Bachelor’s degree and 5+ years of experience in ML or training/inference, contributing to personalization features, memory retrieval, and low‑latency serving in a fast‑paced startup environment based in San Francisco.

Qualifications

  • 5+ years of experience with training or inference systems.
  • Strong engineering fundamentals and ability to ship high‑quality code in a fast‑paced environment.
  • Deep understanding of ML frameworks, GPUs, distributed systems, and infrastructure.
  • Experience with open‑source ML or systems projects is a plus.

Responsibilities

  • Design and optimize training and inference workloads for personalization and continual learning APIs.
  • Collaborate with researchers to turn prototypes into scalable production systems.
  • Improve serving paths for per‑user state and low latency requirements.
  • Work on distributed training across GPUs and nodes.

Skills

Training/inference systems
ML frameworks PyTorch/JAX
GPU computing & distributed systems
Open-source ML experience

Education

Bachelor’s degree or equivalent experience in CS/Engineering

Tools

PyTorch
JAX

Job description

About Engram

EngramToday’s AI is a brilliant stranger: it can solve the world’s hardest math problems, but it knows next to nothing about you and your work. It rereads your files to answer even basic questions, burns an enormous amount of tokens when sifting through large corpuses, and between sessions, it retains scraps at best. We train models to study your world and anticipate your questions in advance, forming engrams: compact memories that capture your knowledge and history. Our approach opens a new axis of scaling. The more we study your context at training time, the better we become at inference time. We’re already working with leaders in AI like Microsoft, Notion, and Harvey, and just raised $98M from General Catalyst, Kleiner Perkins, Sequoia, Factory, Modern, Amplify, Neo and others. Our investors and advisors include Assaf Rappaport, Andrej Karpathy, and Pieter Abbeel. AI has spent years learning everything about the world. Now it should learn something about yours.

About this role

You’ll be among the first ML Systems Engineer, joining a team of machine learning researchers and performance engineers in building our personalization and continual learning API, which powers models and agents that learn from user context. This role is focused on designing, optimizing, and scaling training and inference workloads —bridging the gap between cutting‑edge AI research and production. This includes:

Responsibilities
  • Design and execute new frameworks, techniques, and systems to improve performance, reliability, latency, and efficiency.
  • Partner closely with researchers, turning prototypes into systems that run at scale and feeding systems constraints back into research decisions.
  • Optimize serving paths for personalization and memory retrieval, where per‑user state and low latency both matter.
  • Work on distributed training – data and model parallelism, communication scheduling, and scaling efficiency across multiple GPUs and nodes.
Qualifications
  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • 5+ years of experience with training or inference systems, optimized workloads with measurable results.
  • Strong engineering foundation, with demonstrated excellence navigating complex technical environments and shipping high‑quality code in a fast‑paced environment.
  • Deep understanding of ML framework (e.g., PyTorch, JAX), GPUs, distributed systems, and infrastructure.
  • Operate well in ambiguous environments — you will have real ownership and be responsible for steering the ship in a novel sector of the industry.
  • Bias toward action and knack for turning research concepts into concrete, executable plans.
  • Bonus points if you have early‑stage experience.
  • Experience in open‑source ML or systems infrastructure projects.
What we offer
  • Competitive cash compensation and startup equity.

Engram is based in San Francisco. This role is in‑person in our SF office.

Equal Opportunity Statement

Engram is an equal‑opportunity employer. We’re building a team that reflects a range of backgrounds and perspectives, and we welcome applicants regardless of race, color, religion, national origin, gender, gender identity, sexual orientation, age, disability, or veteran status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer
Platform Engineer

Engram Lab • San Francisco (CA)

On-site
USD 180,000 - 260,000
Startup equity
Competitive cash compensation
Research Scientist
Research Scientist

Engram • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Systems Engineer: Scale Training & Inference
ML Systems Engineer: Scale Training & Inference

Doist • San Francisco (CA)

On-site
USD 180,000 - 230,000
Competitive cash compensation
Startup equity
Platform Engineer - ML Data Infrastructure
Platform Engineer - ML Data Infrastructure

Engram Lab • San Francisco (CA)

On-site
USD 180,000 - 260,000
Startup equity
Competitive cash compensation
Systems ML Engineer (Member of the Technical Staff)
Systems ML Engineer (Member of the Technical Staff)

Transfyr Bio • Cambridge (MA)

On-site
USD 160,000 - 230,000
Systems ML Engineer (Member of the Technical Staff)
Systems ML Engineer (Member of the Technical Staff)

S27a • Cambridge (MA)

On-site
USD 170,000 - 240,000
Competitive compensation
Full benefits
Research Engineer Infrastructure Training Systems
Research Engineer Infrastructure Training Systems

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health benefits
Unlimited PTO
Paid parental leave
+1
Systems ML Engineer (Member of the Technical Staff)
Systems ML Engineer (Member of the Technical Staff)

Breakout Ventures • Cambridge (MA)

On-site
USD 170,000 - 260,000
System Software Engineer - AI
System Software Engineer - AI

Delos Data • Palo Alto (CA)

Hybrid
USD 140,000 - 200,000
Equity
401k
Benefits
Software Engineer, Systems Generalist
Software Engineer, Systems Generalist

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1