Founding ML Inference Engineer — Ultra-Low Latency AI
Reactor
San Francisco (CA)
On-site
USD 180,000 - 280,000
Full time
14 days+
Application generator
Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Get past ATS filters
Benefits offered by this job
Competitive salary
Early equity
Health, dental, and vision coverage
Relocation support
Job summary
A media technology company in San Francisco is seeking a Founding Engineer specializing in ML Inference. This highly technical role requires expertise in the ML infrastructure stack and aims to optimize generative media performance. The ideal candidate will drive innovations in real-time model performance, design in-house inference runtimes, and optimize models through advanced techniques. Competitive salary and relocation support are offered, along with generous health coverage.
Qualifications
Strong foundation in systems programming with a track record of identifying and resolving bottlenecks.
Deep expertise in ML infrastructure stack, including PyTorch, TensorRT, and more.
Knowledge of model compilation and quantization techniques.
Responsibilities
Drive performance for real-time model performance for diffusion models.
Design a high-performance in-house inference runtime.
Optimize models for inference through quantization and pruning.
Benchmark model performance to identify bottlenecks.
Skills
Systems programming
PyTorch
TensorRT
Model compilation
Quantization
NVIDIA GPU hardware knowledge
Transformer architectures
Job description
A media technology company in San Francisco is seeking a Founding Engineer specializing in ML Inference. This highly technical role requires expertise in the ML infrastructure stack and aims to optimize generative media performance. The ideal candidate will drive innovations in real-time model performance, design in-house inference runtimes, and optimize models through advanced techniques. Competitive salary and relocation support are offered, along with generous health coverage.