Sr. SDE, Edge AI ML Platform, Edge AI and Science

Amazon

Vancouver

On-site

CAD 140,000 - 200,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health Insurance
RRSP
DPSP
Paid Time Off

Job summary

Amazon Devices (Lab126) is hiring a Senior Software Development Engineer to lead the architecture and delivery of core Edge AI ML Platform capabilities. You will work with scientists, ML engineers, and hardware teams to enable training, optimization, evaluation, and deployment of models on devices and in the cloud.

You will design and implement distributed training, compression pipelines, and artifact workflows, while driving CI/CD, observability, and reliability for GPU‑intensive workloads

Qualifications

  • 5+ years of non-internship professional software development experience.
  • 5+ years of programming with at least one software programming language.
  • 5+ years of leading design or architecture of new and existing systems.
  • Experience as a mentor, tech lead or leading an engineering team.
  • Experience designing or building distributed or high‑performance computing systems.

Responsibilities

  • Lead design and delivery of distributed ML platform services and libraries across model onboarding, optimization, training, evaluation, packaging and deployment.
  • Define stable APIs and architecture boundaries to decouple research from training and deployment.
  • Design distributed training across data, tensor, pipeline and model parallelism.
  • Scale workflows on multi-node GPU clusters, improve throughput, memory, and recovery.
  • Develop infrastructure connecting distributed training with distillation, quantization, pruning and other techniques.
  • Build evaluation and artifact workflows and carry validated models to target hardware.
  • Implement CI/CD, observability, and release mechanisms for GPU workloads.
  • Profile and optimize end-to-end performance with scientists and kernel engineers.
  • Establish on-call runbooks and incident practices for production services.
  • Collaborate with model, compiler, runtime, hardware and product teams on multi‑team programs.
  • Write designs, evaluate trade-offs, and build consensus when tech strategy is unclear.
  • Mentor engineers, improve code review practices, and help recruit in Vancouver.

Skills

Software development
Programming experience
Architecture leadership
Mentor/Tech lead
Distributed systems

Tools

PyTorch
TensorFlow
JAX
NeMo
Megatron

Job description

Sr. SDE, Edge AI ML Platform, Edge AI and Science

Amazon Devices (Lab126) builds products and services that delight millions of customers globally. The Edge AI ML Platform and Infrastructure team is building the platform that enables Amazon teams to train, optimize, evaluate, and deploy generative AI models on devices and in the cloud.

Today, optimizing a large model for a new hardware target requires experts to connect model onboarding, distributed training, compression, evaluation, compilation, and deployment systems by hand. We are turning that work into a repeatable, self-service workflow. Our platform supports large language, vision, audio, multimodal, and mixture-of-experts models. It gives scientists and engineers the tools to move new optimization techniques from research code into reliable production workflows.

We are looking for a Senior Software Development Engineer to lead the architecture and delivery of core ML platform capabilities. You will solve problems across distributed training on multi-node GPU clusters, model onboarding, compression pipelines, evaluation, GPU performance, artifact management, CI/CD, observability, and operational reliability. You will work with applied scientists, ML engineers, GPU kernel engineers, compiler and runtime teams, hardware teams, and product teams to deliver systems for models with hundreds of billions of parameters.

This role combines hands‑on software development with technical leadership. You will write and review code, define architecture, resolve ambiguous requirements, lead projects that span multiple engineers and teams, and raise the engineering bar for an evolving ML platform.

Key job responsibilities
  • Lead the design and delivery of distributed ML platform services and libraries across model ingestion, optimization, training, evaluation, packaging, and deployment.
  • Define stable APIs and architecture boundaries that allow scientists to add algorithms without coupling research code to training, infrastructure, or deployment implementations.
  • Design distributed training capabilities across data, tensor, pipeline, and model parallelism for large language and multimodal models.
  • Scale workflows on multi-node GPU clusters while improving training throughput, GPU utilization, memory efficiency, communication performance, failure recovery, and developer iteration time.
  • Develop infrastructure that connects distributed training with distillation, quantization, pruning, and other model optimization techniques.
  • Build evaluation and artifact workflows that measure model quality and system performance, then carry validated models through deployment on target hardware.
  • Build automated validation, CI/CD, regression testing, observability, and release mechanisms for GPU-intensive ML workloads.
  • Profile and optimize end‑to‑end system performance with applied scientists and GPU kernel engineers. Translate bottlenecks into durable platform improvements.
  • Establish operational mechanisms, including metrics, alarms, runbooks, on‑call practices, and root‑cause correction for production platform services.
  • Partner with model, compiler, runtime, hardware, security, and infrastructure teams to clarify requirements, manage technical dependencies, and deliver multi‑team programs.
  • Write technical designs, evaluate trade‑offs, and build consensus when the customer need is clear but the technology strategy is not.
  • Mentor engineers, improve code and design review practices, and help recruit and develop a strong engineering team in Vancouver.
A day in the life

You will move between architecture and implementation. Your work will include reviewing designs for model onboarding interfaces, investigating failures in distributed training runs, profiling GPU workloads with scientists, leading cross‑team reviews of end‑to‑end deployment paths, simplifying platform abstractions, and improving the release and regression mechanisms used by multiple model teams.

You will use performance, reliability, and developer productivity data to prioritize platform investments. You will make incremental deliveries while protecting long‑term architecture, and you will ensure that the team resolves recurring problems at their root.

About the team

The Edge AI ML Platform and Infrastructure team brings together software engineers, ML infrastructure engineers, and GPU performance specialists. We build reusable model training, optimization, and deployment capabilities for Amazon product teams, working closely with applied scientists across Edge AI. Our customers need to adapt rapidly changing model architectures to constrained hardware and production workloads without rebuilding the toolchain for every model.

The team owns the platform foundations that connect model development to deployment. Our end‑to‑end scope lets us improve training, compression, evaluation, and deployment as one system. We value clear interfaces, measurable performance, automated quality gates, and direct collaboration between science and engineering.

Basic Qualifications
  • 5+ years of non-internship professional software development experience
  • 5+ years of programming with at least one software programming language experience
  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • Experience as a mentor, tech lead or leading an engineering team
  • Experience designing or building distributed systems or high‑performance computing systems.
Preferred Qualifications
  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Experience building distributed ML training, inference, evaluation, or data platforms using frameworks such as PyTorch, TensorFlow, JAX, NeMo, or Megatron.
  • Experience with containers, Kubernetes, AWS infrastructure, CI/CD, observability, and production operations.
  • Experience with model compression, quantization, knowledge distillation, model compilation, or edge deployment.
  • Experience designing extensible platform APIs and delivering systems with science, hardware, compiler, or product teams.

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. As a total compensation company, Amazon's package may include other elements such as sign‑on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon offers comprehensive benefits including health insurance (medical, dental, vision, prescription, basic life & AD&D insurance), Registered Retirement Savings Plan (RRSP), Deferred Profit Sharing Plan (DPSP), paid time off, and other resources to improve health and well‑being. We thank all applicants for their interest, however only those interviewed will be advised as to hiring status.

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Applied Scientist, Silicon and Systems Group Edge AI, Edge AI Platform
Sr. Applied Scientist, Silicon and Systems Group Edge AI, Edge AI Platform

Amazon Science • Vancouver

On-site
CAD 195,900 - 327,200
Health insurance
RRSP
DPSP
+2
Sr. Applied Scientist, Workforce Solutions
Sr. Applied Scientist, Workforce Solutions

Amazon Science • Vancouver

On-site
CAD 196,000 - 327,000
Machine Learning Engineer , Amazon Customer Service
Machine Learning Engineer , Amazon Customer Service

Amazon • Vancouver

On-site
CAD 115,000 - 192,000
Health insurance
RRSP
DPSP
+1
Software Engineering Manager, ML Kernel Performance, AWS Neuron, Annapurna Labs
Software Engineering Manager, ML Kernel Performance, AWS Neuron, Annapurna Labs

Amazon • Toronto

On-site
CAD 171,000 - 286,000
Sr. Applied Scientist, Silicon and Systems Group Edge AI, Edge AI Platform
Sr. Applied Scientist, Silicon and Systems Group Edge AI, Edge AI Platform

Amazon • Vancouver

On-site
CAD 196,000 - 327,000
Health insurance
RRSP
Paid time off
Software Dev Engineer II, Prime MG Tech - Turing and Foundational Tech / Prime Pixel
Software Dev Engineer II, Prime MG Tech - Turing and Foundational Tech / Prime Pixel

Amazon • Vancouver

On-site
CAD 115,000 - 192,000
Sr. Applied Scientist, Workforce Solutions
Sr. Applied Scientist, Workforce Solutions

Socket.dev • Vancouver

On-site
CAD 196,000 - 327,000
Health insurance (medical, dental, and
RRSP
DPSP
+1
Senior ML Kernel Performance Engineer
Senior ML Kernel Performance Engineer

Amazon • Toronto

On-site
CAD 151,000 - 252,000
Software Engineering Manager, ML Kernel Performance, AWS Neuron, Annapurna Labs
Software Engineering Manager, ML Kernel Performance, AWS Neuron, Annapurna Labs

Socket.dev • Toronto

On-site
CAD 171,000 - 286,000
Software Engineering Manager, ML Kernel Performance, AWS Neuron, Annapurna Labs
Software Engineering Manager, ML Kernel Performance, AWS Neuron, Annapurna Labs

Amazon Web Services (AWS) • Toronto

On-site
CAD 171,000 - 286,000