ML Inference Engineer - High-Performance AI Systems

Together Computer Inc

San Francisco (CA)

On-site

USD 200,000 - 300,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Startup equity
Health insurance
Competitive compensation

Job summary

Together AI is seeking a Machine Learning Engineer to strengthen our Inference Engine team. You will design production systems powering AI inference at scale, optimize runtime services for large-scale AI applications, and collaborate with researchers, engineers, and product teams to bring cutting-edge features to market.

The role requires 3+ years of production-quality coding experience, proficiency in Python and PyTorch, and a strong grasp of low-level OS concepts, multithreading, memory

Qualifications

  • 3+ years of experience writing production-quality code.
  • Proficiency with Python and PyTorch.
  • Experience building high-performance libraries and tooling.
  • Strong understanding of low-level OS concepts including multi-threading, memory management, networking, storage, performance, and scale.

Responsibilities

  • Design and build the production systems that power the Together AI inference engine, enabling reliability and performance at scale.
  • Develop and optimize runtime inference services for large-scale AI applications.
  • Collaborate with researchers, engineers, product managers, and designers to bring new features and research capabilities to the world.
  • Conduct design and code reviews to ensure high standards of quality.
  • Create services, tools, and developer documentation to support the inference engine.
  • Implement robust and fault-tolerant systems for data ingestion and processing.

Skills

Python
PyTorch
High-performance libraries
Multi-threading
CUDA/Triton

Tools

TGI
vLLM
TensorRT-LLM
Optimum
Rust
Cython
Compilers

Job description

Together AI is seeking a Machine Learning Engineer to strengthen our Inference Engine team. You will design production systems powering AI inference at scale, optimize runtime services for large-scale AI applications, and collaborate with researchers, engineers, and product teams to bring cutting-edge features to market.

The role requires 3+ years of production-quality coding experience, proficiency in Python and PyTorch, and a strong grasp of low-level OS concepts, multithreading, memory

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
ML Inference & Performance Engineer
ML Inference & Performance Engineer

Bonfirevc • Palo Alto (CA)

On-site
USD 180,000 - 250,000
Machine Learning Engineer - Inference
Machine Learning Engineer - Inference

Fintal Partners • New York (NY)

On-site
USD 140,000 - 220,000
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
LLM AI Inference Performance Engineer
LLM AI Inference Performance Engineer

Intel • California (MO)

Hybrid
USD 171,000 - 315,000
Software Engineer, Inference Runtime
Software Engineer, Inference Runtime

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000