Senior AI Inference Systems Engineer - Global, Low Latency

Amazon Web Services (AWS)

New York (NY)

On-site

USD 168,000 - 227,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

RSUs
401(k) matching
Paid time off
Parental leave

Job summary

Amazon.com Services LLC is seeking a Senior Software Development Engineer for the Mantle team to drive the vision of a global, high-performance ML inference platform that serves Bedrock customers. You will shape architecture, mentoring, and security-focused design decisions that impact latency, reliability, and cost at scale.

The role emphasizes cross-team collaboration, open AI model onboarding, and operating at millisecond latency with strong security guarantees across AWS regions.

Qualifications

  • 5+ years of non-internship professional software development experience.
  • 5+ years of programming experience in at least one modern language (Java, C++, Python, Go, or Rust).
  • Experience leading design/architecture of large-scale systems.
  • Experience mentoring, tech lead roles, or leading an engineering team.

Responsibilities

  • Architect reliable, scalable, and secure solutions for Mantle's distributed inference engine.
  • Evolve Mantle's architecture including capacity management and onboarding pipelines.
  • Drive cross-organizational initiatives to deliver OpenAI-compatible API experiences with ZOA security guarantees.
  • Lead technical strategy for scaling inference and maintaining global availability and performance SLAs.

Skills

Software development
Distributed systems
Security
Mentorship/tech lead
API design

Education

Bachelor's degree in Computer Science, Engineering or related field
Master's degree or PhD in CS/ML/Engineering or related field

Tools

vLLM
TensorRT
Triton Inference Server

Job description

Amazon.com Services LLC is seeking a Senior Software Development Engineer for the Mantle team to drive the vision of a global, high-performance ML inference platform that serves Bedrock customers. You will shape architecture, mentoring, and security-focused design decisions that impact latency, reliability, and cost at scale.

The role emphasizes cross-team collaboration, open AI model onboarding, and operating at millisecond latency with strong security guarantees across AWS regions.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SDE — Global ML Inference Engine (Mantle)
Senior SDE — Global ML Inference Engine (Mantle)

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+2
Senior SDE - AI Inference Infra for Global Mantle Platform
Senior SDE - AI Inference Infra for Global Mantle Platform

Amazon • New York (NY)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+1
Senior Software Development Engineer, AWS Mantle
Senior Software Development Engineer, AWS Mantle

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+2
Senior Software Development Engineer, AWS Mantle
Senior Software Development Engineer, AWS Mantle

Amazon Web Services (AWS) • New York (NY)

On-site
USD 168,000 - 227,000
RSUs
401(k) matching
Paid time off
+1
Senior Software Development Engineer, AWS Mantle
Senior Software Development Engineer, AWS Mantle

Socket.dev • Seattle (WA)

On-site
USD 168,100 - 227,400
Comprehensive health benefits
401(k) matching
Paid time off
+1
Senior Software Development Engineer, AWS Mantle
Senior Software Development Engineer, AWS Mantle

Amazon • New York (NY)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+1
Enterprise AI Solutions Architect – Strategic & Production‑Ready
Enterprise AI Solutions Architect – Strategic & Production‑Ready

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 183,000 - 247,000
Senior Systems Engineer, AI Inference & HPC Networking
Senior Systems Engineer, AI Inference & HPC Networking

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
RSUs
Health insurance
401(k) matching
+1
Senior ML Inference Engineer — Open-Source LLM Serving
Senior ML Inference Engineer — Open-Source LLM Serving

Amazon • Seattle (WA), Northern (KY)

Hybrid
USD 210,000 - 320,000
Senior AI/ML Systems Engineer - Flexible Hours & Networking
Senior AI/ML Systems Engineer - Flexible Hours & Networking

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000