Member of Technical Staff, Performance and Scale

Inferact

San Francisco (CA)

Hybrid

USD 200,000 - 400,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Generous health, dental, and vision benefits
401(k) company match
Equity options

Job summary

A leading AI technology company located in San Francisco is seeking an infrastructure engineer to build distributed systems for their AI inference engine. The role involves designing systems that ensure minimal latency and maximum reliability. Candidates should have a strong background in systems programming (Rust, Go, C++) and experience with distributed systems. The compensation ranges from $200,000 to $400,000 USD annually, accompanied by benefits and possible visa sponsorship.

Qualifications

  • Bachelor's degree or equivalent in computer science or engineering.
  • Strong systems programming skills in Rust, Go, or C++.
  • Experience with high-performance distributed systems.
  • Understanding of network protocols and the ability to debug issues.

Responsibilities

  • Design and implement systems that power inference at global scale.
  • Enable vLLM to serve models across thousands of accelerators.
  • Absorb complexity into the infrastructure being built.

Skills

Systems programming skills in Rust, Go, or C++
Designing and building high-performance distributed systems
Understanding of network protocols and high-performance I/O
Ability to debug complex distributed systems issues

Education

Bachelor's degree or equivalent experience in computer science, engineering, or similar

Tools

GPU programming models
NVLink
InfiniBand
RoCE

Job description

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About The Role

We're looking for an infrastructure engineer to build the distributed systems that power inference at global scale. You'll design and implement the foundational layers that enable vLLM to serve models across thousands of accelerators with minimal latency and maximum reliability. Tomorrow, deploying a frontier model at scale should be as straightforward as spinning up a serverless database. The complexity doesn't disappear as it gets absorbed into the infrastructure you're building.

Skills And Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Strong systems programming skills in Rust, Go, or C++.
  • Experience designing and building high-performance distributed systems at scale.
  • Understanding of network protocols and high-performance I/O.
  • Ability to debug complex distributed systems issues.

Preferred qualifications:

  • Experience with ML serving infrastructure and disaggregated inference architecture.
  • Familiarity with GPU programming models and memory hierarchies.
  • Knowledge of GPU interconnects (NVLink, InfiniBand, RoCE) and their performance characteristics.
  • Track record of improving system reliability and performance at scale.

Bonus points if you have:

  • Prior experience in supporting large‑scale model training or inference environments.

Logistics

  • Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
  • Visa sponsorship: We sponsor visas on a case-by-case basis.
  • Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.

Compensation Range: $200K - $400K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of Engineering
Head of Engineering

Inferact • San Francisco (CA)

On-site
USD 260,000 - 380,000
Health, dental, vision benefits
401(k) company match
Member of Technical Staff, Cluster Administration
Member of Technical Staff, Cluster Administration

Inferact Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health benefits
Dental benefits
Vision benefits
+1
Member of Technical Staff, Exceptional Generalist (Remote)
Member of Technical Staff, Exceptional Generalist (Remote)

Inferact • United States

Remote
USD 180,000 - 240,000
Competitive salary and equity
Visa sponsorship
Health coverage where applicable
Member of Technical Staff, Inference
Member of Technical Staff, Inference

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Health, dental, and vision benefits
401(k) company match
Visa sponsorship on case-by-case basis
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
AI Inference Engineer
AI Inference Engineer

Premier Global Links LLC • Palo Alto (CA), Northern (KY)

Hybrid
USD 230,000 - 350,000
Equity 0.5%
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
AI Inference Engineer
AI Inference Engineer

Socket.dev • Palo Alto (CA)

On-site
USD 230,000 - 350,000
Equity opportunity (0.5%)
Professional growth
High-impact work
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1