GPU Infra Engineer: Scale Massive Clusters & Observability
Exa
San Francisco (CA)
On-site
USD 120,000 - 160,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A cutting-edge tech company in San Francisco seeks infrastructure engineers to enhance the tooling and systems that power its AI applications. Responsibilities include building GPU orchestration, scaling cloud batchjob systems, and designing efficient scheduling software. Candidates should have experience with large-scale infrastructure and a strong focus on reliability and observability. This position is in-person, and international sponsorship is available.
Qualifications
Experience with large-scale infrastructure, such as GPU clusters or cloud batchjob systems.
Obsessive focus on reliability, observability, and optimization.
Responsibilities
Build GPU cluster orchestration.
Scale AWS batchjob systems to handle extensive jobs.
Design GPU scheduling software for maximum utilization.
Develop observability tooling for production systems.
Skills
Designing and operating large-scale infrastructure
Reliability and observability mindset
Optimization across the entire stack
Tools
Kubernetes
AWS batchjob systems
GPU scheduling software
Job description
A cutting-edge tech company in San Francisco seeks infrastructure engineers to enhance the tooling and systems that power its AI applications. Responsibilities include building GPU orchestration, scaling cloud batchjob systems, and designing efficient scheduling software. Candidates should have experience with large-scale infrastructure and a strong focus on reliability and observability. This position is in-person, and international sponsorship is available.