Inference Systems Engineer for Transformers & Low-Latency HPC

Etched

San Jose (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for those moving to San Jose
Daily lunch + dinner in the office

Job summary

An innovative AI hardware company in San Jose is looking for talented engineers to support the porting of state-of-the-art AI models to their architecture. Candidates should be proficient in C++ or Rust and have a strong understanding of performance-sensitive distributed software systems. This full-time position offers competitive benefits including a housing subsidy and wellness programs designed to support team members both professionally and personally. Join the team committed to redefining AI infrastructure.

Qualifications

  • Proficiency in C++ or Rust.
  • Understanding of complex distributed software systems like Linux internals and accelerator architectures.
  • Familiarity with deep learning frameworks such as PyTorch or JAX.

Responsibilities

  • Support porting state‑of‑the‑art models to our architecture.
  • Build, enhance, and scale Sohu’s runtime for multi‑node inference.
  • Utilize performance profiling tools to identify bottlenecks.

Skills

Proficiency in C++ or Rust
Understanding of performance-sensitive distributed software systems
Familiarity with PyTorch or JAX

Job description

An innovative AI hardware company in San Jose is looking for talented engineers to support the porting of state-of-the-art AI models to their architecture. Candidates should be proficient in C++ or Rust and have a strong understanding of performance-sensitive distributed software systems. This full-time position offers competitive benefits including a housing subsidy and wellness programs designed to support team members both professionally and personally. Join the team committed to redefining AI infrastructure.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Software Engineer: High-Performance Transformers
Inference Software Engineer: High-Performance Transformers

Etched.ai, Inc. • San Jose (CA)

On-site
USD 120,000 - 180,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support
+2
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Inference Systems Engineer, High-Throughput AI Serving
Inference Systems Engineer, High-Throughput AI Serving

Future Ventures • Palo Alto (CA)

On-site
USD 135,000 - 160,000
Comprehensive medical, vision, dental coverage
401(k) retirement plan
Paid parental leave
+1
System Software Engineer — AI Compute & HPC
System Software Engineer — AI Compute & HPC

Etched.ai, Inc. • San Jose (CA)

On-site
USD 120,000 - 160,000
Medical insurance with generous premium coverage
Housing subsidy of $2k per month
Relocation support for moving to San Jose
+2
Staff Engineer, Scalable AI Inference Infrastructure
Staff Engineer, Scalable AI Inference Infrastructure

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Senior ML Engineer - Real-Time Inference & Systems
Senior ML Engineer - Real-Time Inference & Systems

Inworld AI • Germany (OH)

On-site
USD 120,000 - 180,000
Inference Systems Engineer — High-Performance ML Runtime
Inference Systems Engineer — High-Performance ML Runtime

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical/dental/vision benefits
Housing subsidy
Relocation support
+2
Senior AI Inference Engineer - GPU, Rust & CUDA
Senior AI Inference Engineer - GPU, Rust & CUDA

Perplexity • San Francisco (CA)

On-site
USD 220,000 - 485,000
Senior AI Model Serving Engineer — Low-Latency Inference
Senior AI Model Serving Engineer — Low-Latency Inference

Menlo Ventures • San Francisco (CA)

On-site
USD 166,000 - 225,000
Annual performance bonus
Equity options
Comprehensive benefits package
AI Systems Intern - Supercomputing & Hardware Co-Design
AI Systems Intern - Supercomputing & Hardware Co-Design

Etched • San Jose (CA)

On-site
Generous housing support for those relocating
Daily lunch and dinner in the office
Direct mentorship from industry leaders