ML Inference Engineer - Custom Accelerator Data Plane

Amazon

Cupertino (CA)

On-site

USD 165,000 - 224,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Amazon Cupertino is seeking a Software Development Engineer for the MLIL DataPlane team to own design and implement inference data plane for large models on custom hardware. You will work on model execution, memory management, data movement, and serving integration.

You will develop and optimize compute kernels for a custom ML accelerator, validate architectures end-to-end, and build profiling infrastructure while driving performance across the stack from simulation to production on GPUs/ASICs.

Qualifications

  • Bachelor's degree or equivalent.
  • 4+ years of full software development life cycle experience.
  • Knowledge of computer architecture, operating systems, and parallel computing.
  • Knowledge of Linux fundamentals.
  • Strong proficiency in C/C++.
  • Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
  • Proven track record of owning and delivering complex software features end-to-end.

Responsibilities

  • Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
  • Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
  • Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
  • Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
  • Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bring-up.
  • Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
  • Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.

Skills

C/C++
Linux fundamentals
Parallel computing
Computer architecture
Software development lifecycle

Education

Bachelor's degree or equivalent

Job description

Amazon Cupertino is seeking a Software Development Engineer for the MLIL DataPlane team to own design and implement inference data plane for large models on custom hardware. You will work on model execution, memory management, data movement, and serving integration.

You will develop and optimize compute kernels for a custom ML accelerator, validate architectures end-to-end, and build profiling infrastructure while driving performance across the stack from simulation to production on GPUs/ASICs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Inference Systems Engineer - Custom Accelerator
Senior ML Inference Systems Engineer - Custom Accelerator

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs
401(k) matching
Senior ML Inference Engineer - Custom Hardware, LLMs
Senior ML Inference Engineer - Custom Hardware, LLMs

Socket.dev • Cupertino (CA)

On-site
USD 193,000 - 262,000
LLM Inference Systems Engineer — Custom Accelerator
LLM Inference Systems Engineer — Custom Accelerator

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
401(k) matching
+2
Senior ML Software Engineer, Data Plane
Senior ML Software Engineer, Data Plane

Socket.dev • Cupertino (CA)

On-site
USD 193,000 - 262,000
ML Software Engineer, Data Plane
ML Software Engineer, Data Plane

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
401(k) matching
+2
Senior ML Software Engineer, Data Plane (AWS)
Senior ML Software Engineer, Data Plane (AWS)

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs
401(k) matching
Senior ML Software Engineer, Data Plane
Senior ML Software Engineer, Data Plane

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+1
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
ML Accelerator Engineer - Software/Hardware Co-Design
ML Accelerator Engineer - Software/Hardware Co-Design

Google Inc. • Mountain View (CA), San Francisco (CA)

On-site
USD 174,000 - 252,000
Equity
Benefits
Bonus target
ML Inference Performance Visibility Engineer
ML Inference Performance Visibility Engineer

Etched • San Jose (CA)

On-site
USD 150,000 - 210,000
Housing subsidy
Relocation support
Medical benefits
+1