Senior Inference Systems Engineer — Large-Scale GPUs
RadixArk
Palo Alto (CA)
On-site
USD 190,000 - 260,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Competitive compensation
Meaningful equity
Comprehensive benefits
Flexible work arrangements
Job summary
A leading AI infrastructure company in California is seeking a Member of Technical Staff — Inference to design and optimize large-scale AI inference systems. The role demands 5+ years in systems engineering and expertise in large-scale inference systems. Successful candidates will enhance GPU utilization and work closely with various teams to debug and drive the reliability of infrastructure. Competitive compensation and flexible work arrangements are offered, alongside a commitment to equity and diversity.
Qualifications
5+ years of experience in systems engineering, ML infrastructure, or performance-critical backend systems.
Strong expertise in large-scale inference systems for LLMs or generative models.
Deep understanding of GPU architecture and performance characteristics.
Experience optimizing latency- and throughput-critical production systems.
Strong knowledge of distributed systems and networking fundamentals.
Proficiency in C++, Rust, Go, or Python for production systems.
Experience profiling and optimizing compute-intensive workloads.
Responsibilities
Design and build large-scale inference systems for frontier AI models.
Optimize latency, throughput, and GPU utilization in production inference.
Develop and improve model serving architectures and runtimes.
Work on batching, scheduling, and memory management strategies.
Collaborate with kernel, compiler, and systems teams on performance optimization.
Debug performance bottlenecks across the stack.
Drive reliability and scalability of inference infrastructure.
Build tooling for observability, profiling, and performance analysis.
Contribute to long-term inference architecture and strategy.
Job description
A leading AI infrastructure company in California is seeking a Member of Technical Staff — Inference to design and optimize large-scale AI inference systems. The role demands 5+ years in systems engineering and expertise in large-scale inference systems. Successful candidates will enhance GPU utilization and work closely with various teams to debug and drive the reliability of infrastructure. Competitive compensation and flexible work arrangements are offered, alongside a commitment to equity and diversity.