Inference Systems Engineer for Transformers & Low-Latency HPC
Etched
San Jose (CA)
On-site
USD 180,000 - 240,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for those moving to San Jose
Daily lunch + dinner in the office
Job summary
An innovative AI hardware company in San Jose is looking for talented engineers to support the porting of state-of-the-art AI models to their architecture. Candidates should be proficient in C++ or Rust and have a strong understanding of performance-sensitive distributed software systems. This full-time position offers competitive benefits including a housing subsidy and wellness programs designed to support team members both professionally and personally. Join the team committed to redefining AI infrastructure.
Qualifications
Proficiency in C++ or Rust.
Understanding of complex distributed software systems like Linux internals and accelerator architectures.
Familiarity with deep learning frameworks such as PyTorch or JAX.
Responsibilities
Support porting state‑of‑the‑art models to our architecture.
Build, enhance, and scale Sohu’s runtime for multi‑node inference.
Utilize performance profiling tools to identify bottlenecks.
Skills
Proficiency in C++ or Rust
Understanding of performance-sensitive distributed software systems
Familiarity with PyTorch or JAX
Job description
An innovative AI hardware company in San Jose is looking for talented engineers to support the porting of state-of-the-art AI models to their architecture. Candidates should be proficient in C++ or Rust and have a strong understanding of performance-sensitive distributed software systems. This full-time position offers competitive benefits including a housing subsidy and wellness programs designed to support team members both professionally and personally. Join the team committed to redefining AI infrastructure.