Inference Systems Engineer — High-Performance ML Runtime

The Consensus

San Jose (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical/dental/vision benefits
Housing subsidy
Relocation support
Wellness programs
Free lunch/dinner

Job summary

Etched is building hardware for frontier intelligence and seeks engineers to port state-of-the-art models to our architecture and to build scalable runtimes for multi-node inference. You will optimize communication, state management, and error handling while working with Linux internals and accelerator architectures.

Strong candidates will have experience with C++ or Rust, and familiarity with PyTorch or JAX.

Qualifications

  • Proficiency in C++ or Rust.
  • Understanding of performance-sensitive or complex distributed software systems.
  • Familiarity with PyTorch or JAX.
  • Ported applications to non-standard accelerator hardware or hardware platforms.

Responsibilities

  • Port models to our architecture and help build testing capabilities for rapid iteration.
  • Develop and scale runtime for multi-node inference and state management.
  • Optimize routing and communication layers across accelerators.
  • Use profiling tools to identify bottlenecks and correctness issues.

Skills

C++
Rust
PyTorch
JAX

Tools

Linux internals
Accelerator architectures
High-speed interconnects
Compilers

Job description

Etched is building hardware for frontier intelligence and seeks engineers to port state-of-the-art models to our architecture and to build scalable runtimes for multi-node inference. You will optimize communication, state management, and error handling while working with Linux internals and accelerator architectures.

Strong candidates will have experience with C++ or Rust, and familiarity with PyTorch or JAX.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Systems Engineer — Inference Runtime Lead
Staff Systems Engineer — Inference Runtime Lead

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 405,000 - 485,000
Inference Software Engineer
Inference Software Engineer

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical/dental/vision benefits
Housing subsidy
Relocation support
+2
Inference engineer
Inference engineer

Garuda Ventures • Hermosa Beach (CA)

On-site
USD 100,000 - 140,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
High-Performance Inference Engineer (ML Systems)
High-Performance Inference Engineer (ML Systems)

Garuda Ventures • Hermosa Beach (CA)

On-site
USD 100,000 - 140,000
Staff Inference Runtime Architect (Rust/Python)
Staff Inference Runtime Architect (Rust/Python)

Anthropic • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
Inference Systems Engineer for Transformers & Low-Latency HPC
Inference Systems Engineer for Transformers & Low-Latency HPC

Etched • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for those moving to San Jose
+1
Staff Inference Runtime Architect (Rust/Python)
Staff Inference Runtime Architect (Rust/Python)

Anthropic • New York (NY)

Hybrid
USD 405,000 - 485,000
Generous vacation
Parental leave
Flexible working hours
+1
Embedded ML Inference & Optimization Engineer
Embedded ML Inference & Optimization Engineer

Applied Intuition Inc. • Sunnyvale (CA)

Hybrid
USD 159,000 - 200,000
Staff Inference Systems Engineer — High-Throughput AI
Staff Inference Systems Engineer — High-Throughput AI

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000