AI Systems Engineer

Transluce

San Francisco (CA)

On-site

USD 350,000 - 600,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Transluce, a fast-moving research lab in San Francisco, is seeking an exceptional AI systems engineer to lead the design and development of our core ML stack, building scalable systems that can leverage thousands of GPUs and handle trillion-token databases.

As an early member of a highly collaborative team, you will move fast, innovate from the ground up, and deliver high-impact tooling with cross-organizational reach, including open-source components that support AI evaluation and public policy

Qualifications

  • Expert Python programmer with GPU/parallel programming experience.
  • Experience engineering distributed systems at scale.
  • Leadership in code health and team impact.
  • Familiarity with LLM pipelines and AI tooling.

Responsibilities

  • Set overall code culture and tooling for a fast-growing org.
  • Address core infra challenges across verticals (Docent, Interpretability, RL).
  • Build internal tools to speed up the team and improve reliability.
  • Help shape best practices for infra decisions and path-setting.

Skills

Python
Distributed systems
GPU programming
Code quality
Leadership
LLM pipelines

Job description

Salary range:

$350,000 - $600,000/year + benefits


Description:

Transluce is a fast-moving research lab building the public tech stack for understanding and debugging AI systems. We build world-class, AI-backed analysis tools and use these to set industry standards for evaluation. We are a non-profit with a mission to steer the development of AI for the public good.


About the role:

We are looking for an exceptional AI systems engineer to lead the design and development of our core ML stack, building systems that can scale to thousands of GPUs and performantly query trillion-token databases.


As an early member of a highly collaborative team, you will be free to innovate and move fast, building high-impact systems from the ground up. As part of a mission-focused non-profit, your work will have high direct impact (e.g. used by governments to inform AI policy) and cross-organisational reach (open-source tools the entire community can build on).


Core responsibilities:


  • Set overall code culture and tooling for a fast-growing org

  • Help to solve our core technical challenges across verticals. Examples include:

    • Docent:

      • High-concurrency container-based evals with quick ability to iterate on interventions to agentic trajectories

      • Deterministic sandbox execution of code that can efficiently restore state from checkpoints



    • Interpretability:

      • Inference stacks that are as performant as vLLM but flexible enough to allow complex model introspection and intervention, steering, configurable sampling, etc., and that can scale to 400B+ parameter models



    • Behavior elicitation:

      • Distributed RL training and roll-outs allowing thousands of concurrent rollouts across machines



    • Build great internal tools to speed up the team



  • Help tone-set in the organization around best practices for building and path-set on what infra we should build

  • Help other team members think through infra challenges


Qualities of a strong candidate:


  • Exceptional programmer fluent in Python

  • Bare metal optimization: know GPUs, other accelerators in and out (low-level performance + optimization + parallel programming)

  • Experience engineering at scale (distributed systems, reliability, architecture design)

  • Leader on global code quality and health (designing good primitives, managing complexity and scale)

  • Bonus: can set up LLM pipelines, e.g. multiple specialized LLMs interacting with each other in a performant and reliable way

  • Bonus: experience with open-source community management


We are located in San Francisco and enthusiastic to work together in-person. We are open to sponsoring international visas.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist/Research Engineer
Research Scientist/Research Engineer

Transluce • San Francisco (CA)

On-site
USD 250,000 - 500,000
Full-Stack Product Engineer
Full-Stack Product Engineer

Transluce • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 500,000
AI Engineering Tech Lead / Architect
AI Engineering Tech Lead / Architect

Faros AI, Inc. • San Mateo (CA)

On-site
USD 210,000 - 250,000
Distributed Systems Engineer, Data & Inference Platform
Distributed Systems Engineer, Data & Inference Platform

OpenTalent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Flexible work
Adaption Passport
Lunch Stipend
+1
Lead AI Systems Engineer — Scale Core ML & GPU Clusters
Lead AI Systems Engineer — Scale Core ML & GPU Clusters

Transluce • San Francisco (CA)

On-site
USD 350,000 - 600,000
Member of Technical Staff, MLSys
Member of Technical Staff, MLSys

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Machine Learning Systems Engineer
Machine Learning Systems Engineer

Recruiting From Scratch • Palo Alto (CA)

On-site
USD 200,000 - 300,000
Competitive equity
Cutting-edge diffusion models
Direct collaboration with researchers
Software Engineer, Applied AI
Software Engineer, Applied AI

Sobek AI • Seattle (WA)

Hybrid
USD 170,000 - 230,000
Company-paid health coverage
Equity opportunities
Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Physical Superintelligence • Boston (MA)

Hybrid
USD 140,000 - 210,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Netpreme • Northern (KY)

Hybrid
USD 150,000 - 210,000
Performance bonus
Equity grant
Health, dental, vision fully paid
+5