Applications Engineer, AI Server Software Performance

Akash Systems

Emeryville (CA)

On-site

USD 140,000 - 210,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
401(k) with company match
Annual performance bonus

Job summary

Akash Systems, Inc. seeks an Applications Engineer, AI Server Software Performance to bridge engineering, customers, and partners.

You will lead training and inference benchmarking, support POC engagements, and contribute to software development initiatives. You will drive performance evaluation across training/inference workloads, own end-to-end deployment on CUDA/ROCm, and help translate hardware capabilities into optimized software solutions for AI workloads.

Qualifications

  • Bachelor's degree in Computer Engineering or Computer Science; Master’s degree preferred.
  • Minimum of 5 years of relevant industry experience.
  • Familiarity with CUDA and RocM software platforms.
  • Strong written and verbal communication skills.
  • Experience analyzing, interpreting, and presenting technical data.

Responsibilities

  • Evaluate AI server performance across training and inference benchmarks.
  • Build and maintain a catalog of performance and benchmark data.
  • Lead POC demonstrations and customer testing.
  • Own end-to-end LLM deployment across CUDA and ROCm platforms.
  • Drive internal/external benchmarking initiatives (e.g., MLPerf).
  • Bridge OEM partners and chip manufacturers to optimize software solutions.
  • Identify and resolve performance bottlenecks across server stack.

Skills

CUDA/RocM familiarity
Communication skills
Project management
Analytical thinking

Education

Bachelor's degree in Computer Engineering or Computer Science
Master’s degree preferred

Tools

CUDA toolkit
ROCm
MLPerf benchmarking

Job description

Akash Systems, Inc. is a venture-backed, late-stage, Bay Area company that makes and sells Diamond Cooled GPU-based servers to AI companies, Cloud Service Providers, and data centers worldwide.The deep tech company pioneered Diamond Cooling® technology wherein the world’s most thermally conductive material, lab-grown diamond, is brought close to the GPU chip – the densest heat source in a modern high compute server. The resulting Akash server exhibits more compute in FLOPS (>50%) and FLOPS per Watt (by 2x) than any other server in the market today. Akash’s Diamond Cooled servers also reduce the energy cost of cooling the entire data center and maintain performance in high ambient temperatures. The company’s lead investors include Khosla Ventures and Founders Fund. Akash Systems has deployed its Diamond Cooling® technology in space, with many Akash satellite radios in orbit.

Role Description

We are seeking an Applications Engineer, AI Server Software Performance to serve as the technical bridge between our engineering teams, customers, and strategic partners. This individual will lead training and inference benchmarking activities, support customer proof‑of‑concept (POC) engagements, and contribute to software development initiatives.

Responsibilities
  • Evaluate performance of AI servers across training and inference testing benchmarks
  • Build and maintain catalog of performance and benchmark data across benchmark models
  • Lead implementation and support customer testing / proof of concept demonstrations of inference and training models on servers
  • Own E2E LLM deployment across NVIDIA and AMD platforms, from kernel optimization on CUDA and ROCm to production ready inference stacks.
  • Drive internal and external testing and benchmarking initiatives, including 3rd party benchmarks like MLCommons’ MLPerf Inference and MLPerf Training workloads.
  • Serve as a technical bridge between OEM partners and chip manufacturers (e.g., NVIDIA, AMD, Supermicro, Dell), translating hardware capabilities into optimized software solutions for training and inference workloads.
  • Identify and resolve performance bottlenecks across the full server stack, from driver and firmware compatibility to model quantization (FP8/FP4) and throughput tuning.
Required Qualifications
  • Bachelor's degree in Computer Engineering or Computer Science; Master’s Degree preferred.
  • Minimum of 5 years of relevant industry experience following completion of degree.
  • Familiarity with CUDA and RocM software platforms.
  • Strong written and verbal communication skills.
  • Experience analyzing, interpreting, and presenting technical data to both internal and external reporting.
  • Excellent project management skills.

We offer a competitive compensation package commensurate with experience, including:

  • Base salary: competitive and based on experience, skills, and qualifications
  • Annual performance bonus
  • Comprehensive health, dental, and vision insurance
  • 401(k) with company match
Work Authorization

Akash Systems does not sponsor or take over sponsorship for employment visas for this role (e.g., H‑1B, O‑1, TN, F‑1/OPT, etc.). Eligible candidates are U.S. citizens or U.S. lawful permanent residents (Green Card Holders).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Server Performance Applications Engineer – Benchmarking
AI Server Performance Applications Engineer – Benchmarking

Akash Systems • Emeryville (CA)

On-site
USD 140,000 - 210,000
Health insurance
Dental insurance
Vision insurance
+2
Frontier AI Workloads - Performance and Scalability Engineer
Frontier AI Workloads - Performance and Scalability Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 200,000
Sr Hardware Development Engineer, High Performance AI & ML Servers
Sr Hardware Development Engineer, High Performance AI & ML Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 216,000
Test Engineer
Test Engineer

Pegatron Technologies LLC • Georgetown (TX)

On-site
USD 65,000 - 85,000
Senior/Staff AI Engineer
Senior/Staff AI Engineer

Data Direct Networks • California (MO)

On-site
USD 150,000 - 230,000
Vacation plans
Paid holidays
Bonus programs
+5
Performance/ Benchmark Engineer - NVIDIA GPU Systems
Performance/ Benchmark Engineer - NVIDIA GPU Systems

Yoh, A Day & Zimmermann Company • California (MO)

On-site
USD 250,000 - 300,000
Medical+Vision
HSA
Life & Disability
+5
Staff AI Performance Engineer
Staff AI Performance Engineer

EngineersOfAI • Austin (TX)

On-site
USD 90,000 - 120,000
Senior GPU Inference Performance Engineer
Senior GPU Inference Performance Engineer

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Staff AI Performance Engineer
Staff AI Performance Engineer

Graphcore • Austin (TX)

On-site
USD 100,000 - 150,000
Medical, dental, and vision coverage
401(k) retirement plan
Flexible Spending Accounts (FSAs)
+2
Machine Learning Performance Engineer - Offboard Training & Inference
Machine Learning Performance Engineer - Offboard Training & Inference

Decisive Point • Sunnyvale (CA)

On-site
USD 180,000 - 240,000