Software Engineer, Systems Generalist

Thinking Machines Lab Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 350,000 - 475,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Visa sponsorship
Relocation support
Health benefits

Job summary

Thinking Machines Lab Inc. in San Francisco, California seeks generalist infrastructure and systems engineers to build the foundations powering our AI models and to support research and product teams.

You will join a small, high-impact group responsible for architecting and scaling core infrastructure, solving distributed systems challenges, and delivering robust, scalable platforms. This role offers visa sponsorship, competitive compensation—$350,000 to $475,000 annually—and relocation support

Qualifications

  • Bachelor’s degree or equivalent in CS/engineering or similar.
  • Proficiency in Python or Rust as a backend language.
  • Experience operating large-scale clusters and container orchestration (Kubernetes/Slurm).
  • Comfort operating across the stack and owning projects end-to-end.
  • Thrive in a collaborative, cross-functional environment.
  • Bias for action and initiative to drive shipping across stacks.

Responsibilities

  • Architect and scale core infrastructure powering AI models.
  • Collaborate with researchers to accelerate experiments and improve efficiency.
  • Build robust, scalable platforms across the full technology stack.
  • Support data and research teams with reliable infrastructure and tooling.

Skills

Python
Rust
Kubernetes
End-to-end ownership
Collaboration

Education

Bachelor’s degree in CS/Engineering

Tools

Spark
Slurm
CI/CD
GPU workloads

Job description

The mission of Thinking Machines is to build AI that extends human will and judgment.


About the Role

We’re looking for generalist infrastructure and systems engineers to help build the systems that power our foundation models and the internal teams on research and product development to be able to create the models and ship the products powered by our models.


You'll join a small, high-impact team responsible for architecting and scaling the core infrastructure behind everything we do. You’ll work across the full technical stack, solving complex distributed systems problems and building robust, scalable platforms.


Infrastructure is critical to us: it's the bedrock that enables every breakthrough. You'll work directly with researchers to accelerate experiments, improve infrastructure efficiency, and enable key insights across our models, products, and data assets.


What You’ll Do

We interview generally, but during project selection we’ll take into account your interests and experience alongside organizational needs. This flexible approach allows us to match talented engineers with the infrastructure teams where they'll have the greatest impact and growth potential.


Here are example areas you may contribute to depending on your area of expertise and interest:



  • Core Infrastructure: We support teams that train, research, and ultimately serve AI models and build the underlying infrastructure for the clusters to reliably and safely train frontier models. Examples might include building systems and running large Kubernetes clusters with GPU workloads, or building infrastructure to support Tinker.


  • Data Infrastructure: We build and maintain the data systems for our research and products. You'll design and optimize data pipelines using tools like Spark and other modern data infrastructure technologies. You’ll build scalable, reliable, data infrastructure while embedding governance best practices.


  • Developer Productivity: We care deeply about research and engineering productivity and our ability to continue shipping quickly. We build tooling, systems, frameworks, and systems to make sure everyone gets well configured, optimized developer environments.



Skills and Qualifications

Minimum qualifications:



  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.


  • Proficiency in at least one backend language (we use Python or Rust).


  • Experience operating large‑scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).


  • Comfort operating across the stack and owning projects end-to-end.


  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.


  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.



Preferred qualifications — we encourage you to apply if you meet some but not all of these:



  • Strong debugging across application, OS, and network layers.


  • Proficiency in Python or Rust (or similar), containers, and modern CI.


  • Experience with Kubernetes, controllers/operators, or performance profiling.


  • Familiarity with GPU/ML workflows or large‑scale data/eval pipelines.



Logistics


  • Location: This role is based in San Francisco, California.


  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.


  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.


  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.


Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Systems Generalist
Software Engineer, Systems Generalist

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Engineer, Infrastructure, Training Systems
Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Software Engineer, Research Acceleration
Software Engineer, Research Acceleration

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Software Engineer, Data Infrastructure
Software Engineer, Data Infrastructure

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Software Engineer, Research Acceleration
Software Engineer, Research Acceleration

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Engineer Infrastructure Training Systems
Research Engineer Infrastructure Training Systems

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health benefits
Unlimited PTO
Paid parental leave
+1
Engineering Manager
Engineering Manager

Thinkingmachines • San Francisco (CA)

On-site
USD 400,000 - 500,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Software Engineer, Data Infrastructure
Software Engineer, Data Infrastructure

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision insurance
Unlimited paid time off (PTO)
Paid parental leave
+1
Software Engineer, Developer Productivity, AI Tools
Software Engineer, Developer Productivity, AI Tools

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Software Engineer, Full Stack
Software Engineer, Full Stack

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1