Inference Systems Engineer for Transformers & Low-Latency HPC

Etched

San Jose (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support for those moving to San Jose
Daily lunch + dinner in the office

Job summary

An innovative AI hardware company in San Jose is looking for talented engineers to support the porting of state-of-the-art AI models to their architecture. Candidates should be proficient in C++ or Rust and have a strong understanding of performance-sensitive distributed software systems. This full-time position offers competitive benefits including a housing subsidy and wellness programs designed to support team members both professionally and personally. Join the team committed to redefining AI infrastructure.

Qualifications

  • Proficiency in C++ or Rust.
  • Understanding of complex distributed software systems like Linux internals and accelerator architectures.
  • Familiarity with deep learning frameworks such as PyTorch or JAX.

Responsibilities

  • Support porting state‑of‑the‑art models to our architecture.
  • Build, enhance, and scale Sohu’s runtime for multi‑node inference.
  • Utilize performance profiling tools to identify bottlenecks.

Skills

Proficiency in C++ or Rust
Understanding of performance-sensitive distributed software systems
Familiarity with PyTorch or JAX

Job description

An innovative AI hardware company in San Jose is looking for talented engineers to support the porting of state-of-the-art AI models to their architecture. Candidates should be proficient in C++ or Rust and have a strong understanding of performance-sensitive distributed software systems. This full-time position offers competitive benefits including a housing subsidy and wellness programs designed to support team members both professionally and personally. Join the team committed to redefining AI infrastructure.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Software Engineer: High-Performance Transformers
Inference Software Engineer: High-Performance Transformers

Etched.ai, Inc. • San Jose (CA)

On-site
USD 120,000 - 180,000
Medical, dental, and vision packages
Housing subsidy of $2k per month
Relocation support
+2
Principal Transformer Inference Engineer - High Performance
Principal Transformer Inference Engineer - High Performance

Oho Group • San Francisco (CA)

On-site
USD 180,000 - 240,000
System Software Engineer — AI Compute & HPC
System Software Engineer — AI Compute & HPC

Etched.ai, Inc. • San Jose (CA)

On-site
USD 120,000 - 160,000
Medical insurance with generous premium coverage
Housing subsidy of $2k per month
Relocation support for moving to San Jose
+2
Staff Engineer, Scalable AI Inference Infrastructure
Staff Engineer, Scalable AI Inference Infrastructure

Inferact • San Francisco (CA)

Hybrid
USD 200,000 - 400,000
Generous health, dental, and vision benefits
401(k) company match
Equity options
Inference Systems Engineer — High-Performance ML Runtime
Inference Systems Engineer — High-Performance ML Runtime

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical/dental/vision benefits
Housing subsidy
Relocation support
+2
Performance Engineer, Inference Engine - High-Performance AI
Performance Engineer, Inference Engine - High-Performance AI

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior AI Model Serving Engineer — Low-Latency Inference
Senior AI Model Serving Engineer — Low-Latency Inference

Neon • San Francisco (CA)

On-site
USD 166,000 - 225,000
Annual performance bonus
Equity options
Comprehensive benefits package
Infra Software Engineer: AI ASIC & Scalable HPC
Infra Software Engineer: AI ASIC & Scalable HPC

Etched • San Jose (CA)

On-site
USD 120,000 - 160,000
Full medical, dental, and vision packages
Housing subsidy of $2,000/month
Daily lunch and dinner provided
+1
On-Device AI Inference Engineer — Ultra-Low Latency
On-Device AI Inference Engineer — Ultra-Low Latency

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 450,000
AI Systems Intern - Supercomputing & Hardware Co-Design
AI Systems Intern - Supercomputing & Hardware Co-Design

Etched • San Jose (CA)

On-site
USD 34,440 - 55,104
Generous housing support for those relocating
Daily lunch and dinner in the office
Direct mentorship from industry leaders