AI Performance Engineer – HPC, ARM & Distributed Inference

EngineersOfAI

Austin (TX)

On-site

USD 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

EngineersOfAI is seeking a candidate to optimize AI workloads across various architectures. The role involves collaboration with hardware and software teams to ensure efficient performance of AI/ML models, as well as benchmarking and troubleshooting at scale.

The ideal applicant has a BS/MS in Computer Science or a related field, with strong programming skills in C++ and Python, and experience with distributed systems. This position offers opportunities to work on cutting-edge technology in Austin, Texas.

Qualifications

  • Strong programming skills in C++ and Python.
  • Experience with distributed systems and communication libraries.
  • Experience profiling and optimizing HPC or AI/ML workloads.

Responsibilities

  • Analyze ML models’ compute and memory requirements.
  • Collaborate across hardware and software teams.
  • Benchmark and troubleshoot system performance.

Skills

C++
Python
Distributed systems
Communication libraries (MPI, NCCL, UCX)
Profiling and optimizing HPC or AI/ML workloads

Education

BS/MS in Computer Science, Electrical Engineering, or related field

Tools

ML frameworks (PyTorch, TensorFlow)
HPC networking technologies (InfiniBand, RoCE)

Job description

EngineersOfAI is seeking a candidate to optimize AI workloads across various architectures. The role involves collaboration with hardware and software teams to ensure efficient performance of AI/ML models, as well as benchmarking and troubleshooting at scale.

The ideal applicant has a BS/MS in Computer Science or a related field, with strong programming skills in C++ and Python, and experience with distributed systems. This position offers opportunities to work on cutting-edge technology in Austin, Texas.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff AI Performance Engineer: Distributed HPC on ARM
Staff AI Performance Engineer: Distributed HPC on ARM

Graphcore • Austin (TX)

On-site
USD 100,000 - 150,000
Medical, dental, and vision coverage
401(k) retirement plan
Flexible Spending Accounts (FSAs)
+2
Senior AI Field Applications Engineer – GPUs & HPC
Senior AI Field Applications Engineer – GPUs & HPC

AMD • Austin (TX)

On-site
USD 120,000 - 160,000
Remote work options
Travel opportunities
Staff AI Performance Engineer
Staff AI Performance Engineer

EngineersOfAI • Austin (TX)

On-site
USD 90,000 - 120,000
Embedded Systems Performance Engineer (AI Devices)
Embedded Systems Performance Engineer (AI Devices)

OpenAI • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Fellow AI Performance & Reliability Engineer (Hybrid)
Fellow AI Performance & Reliability Engineer (Hybrid)

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 180,000 - 240,000
AI Inference Engineer - Performance & API
AI Inference Engineer - Performance & API

Topazlabs • Dallas (TX)

On-site
USD 120,000 - 180,000
Full medical/dental/vision coverage
15 days PTO
5 personal days + holidays
+1
AI/HPC Cluster Architect - Scalable Data Center Design
AI/HPC Cluster Architect - Scalable Data Center Design

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 140,000 - 190,000
AMD benefits
Equal opportunity employer
Visa sponsorship not available
AI/HPC Cluster Architect
AI/HPC Cluster Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
AI/HPC Cluster Architect: Power, Network & Systems
AI/HPC Cluster Architect: Power, Network & Systems

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits