Senior Backend/Platform Engineer

TypeSafe AI

San Francisco (CA)

On-site

USD 180,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Base salary of $180k-280k plus equity
100% covered health insurance
Daily lunch and dinner
Visa sponsorships
401K plans

Job summary

TypeSafe AI in San Francisco is seeking a skilled ML Infrastructure and Reliability Engineer to build and maintain its API platform for inference. This role involves creating robust infrastructure across multiple cloud providers, ensuring reliability and low error rates, and maintaining monitoring systems.

The ideal candidate will have 5+ years in software engineering, experience with Kubernetes, and a passion for collaboration. TypeSafe AI offers competitive compensation, health insurance, and a supportive team environment.

Qualifications

  • Experience designing resilient systems and improving on-call experience.
  • Collaborate well with others on technical and product design.
  • Experience with challenges of ML orchestration.

Responsibilities

  • Create robust infrastructure for serving inference across multiple cloud providers.
  • Ensure inference infrastructure is reliable with low error rates.
  • Create and maintain infrastructure for monitoring and alerting.

Skills

Backend services operation
Resilient system design
Debugging skills
Team collaboration

Education

5+ years of professional software engineering experience
3+ years in infrastructure/backend

Tools

Kubernetes
AWS

Job description

About the Role

As an ML infrastructure and reliability engineer, you will join the team responsible for building and maintaining TypeSafe’s API platform for inference. These APIs will be user facing, latency sensitive, and (once we ship) have uptime, reliability and backwards compatibility requirements.

The role is wide ranging and you will wear many hats. Responsibilities include:

  • Create robust infrastructure for serving inference across multiple cloud providers.
  • Work to ensure inference infrastructure and services are reliable and have low error rates.
  • Create and maintain infrastructure for monitoring and alerting on requests, improving our debugging and operations stance.
  • Ramp up oncall engineers on error handling in production via playbooks, mentorship, “fire drill” exercises, etc.

Our tech stack is primarily Python. We also use TypeScript, Next.js, and Tailwind CSS for frontend, with Kubernetes for orchestration. We empower developers to use any tooling they find helpful for getting their job done, including Claude Code and Cursor.

Qualifications
  • Are responsible, ownership-inclined, and a team player – you believe there is no such thing as "other people's code", and will complain only the normal amount about pager duty.
  • Experience building and operating backend services at scale with continuous delivery.
  • Experience designing resilient systems and improving on-call experience on production systems.
  • Are excellent under pressure – especially debugging/resolving outages.
  • Collaborate well with others on technical and product design, advocating for what you need, and adjusting specs to changing requirements and feedback.
  • Are mission aligned and excited to go all-in.
  • Love being part of a team.
Preferred Qualifications
  • Have previously built big things.
  • Are detail-oriented to the point of helpful paranoia.
  • Have 5+ years of professional software engineering experience (3+ years of infra/backend) with a team with rigorous engineering standards.
  • Have experience with Kubernetes, cloud providers, and AWS in particular.
  • Have experience with the challenges of ML orchestration.
  • Have experienced LLMs’ capabilities and limitations from implementing them in the past.
Benefits
  • Base salary of $180k-280k plus equity, based on leveling.
  • 100% covered health insurance.
  • Daily lunch and dinner.
  • Visa sponsorships.
  • 401K plans.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Backend Engineer
Backend Engineer

Space Executive • Berkeley (CA)

Remote
USD 125,000 - 225,000
Medical, dental, vision
401(k)
Unlimited PTO
+2
Distributed Systems Engineer, Data & Inference Platform
Distributed Systems Engineer, Data & Inference Platform

adaption • Redwood City (CA)

Hybrid
USD 120,000 - 160,000
Flexible work
Annual travel stipend
Weekly meal allowance
+1
Distributed Systems Engineer, Data & Inference Platform
Distributed Systems Engineer, Data & Inference Platform

Adaption Labs • San Francisco (CA)

Hybrid
USD 120,000 - 180,000
Flexible work
Annual travel stipend
Weekly meal allowance
+1
Distributed Systems Engineer, Data & Inference Platform
Distributed Systems Engineer, Data & Inference Platform

adaption • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Flexible work
Annual travel stipend
Weekly meal allowance
+1
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Cash compensation range of $150-300k
Flexible work arrangement (remote or San Francisco office)
Full visa sponsorship and relocation support
+3
AI Infrastructure Engineer, Model Serving Platform
AI Infrastructure Engineer, Model Serving Platform

Scale AI, Inc. • New York (NY)

On-site
USD 180,000 - 225,000
Comprehensive health coverage
Equity compensation
Learning and development stipend
+2
Infrastructure Engineer, Kubernetes Specialist
Infrastructure Engineer, Kubernetes Specialist

TypeSafe AI • San Francisco (CA)

On-site
USD 180,000 - 280,000
Base salary of $180k-280k plus equity
100% covered health insurance
Daily lunch and dinner
+2