Get more replies from employers
Send a job-specific resume in minutes.
TypeSafe AI in San Francisco is seeking a skilled ML Infrastructure and Reliability Engineer to build and maintain its API platform for inference. This role involves creating robust infrastructure across multiple cloud providers, ensuring reliability and low error rates, and maintaining monitoring systems.
The ideal candidate will have 5+ years in software engineering, experience with Kubernetes, and a passion for collaboration. TypeSafe AI offers competitive compensation, health insurance, and a supportive team environment.
As an ML infrastructure and reliability engineer, you will join the team responsible for building and maintaining TypeSafe’s API platform for inference. These APIs will be user facing, latency sensitive, and (once we ship) have uptime, reliability and backwards compatibility requirements.
The role is wide ranging and you will wear many hats. Responsibilities include:
Our tech stack is primarily Python. We also use TypeScript, Next.js, and Tailwind CSS for frontend, with Kubernetes for orchestration. We empower developers to use any tooling they find helpful for getting their job done, including Claude Code and Cursor.