SDE II: ML Diagnostics & Fleet Analytics

Amazon

Austin (TX)

On-site

USD 144,000 - 194,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance
RSUs
Sign-on bonus
401(k) matching
Parental leave

Job summary

Amazon MLA Vetting seeks a Software Development Engineer II to build data pipelines ingesting diagnostic results from thousands of servers, dashboards for test behavior, and metrics that indicate when tests flag failures before capacity loss. You will own alarms, production code, and anomaly detection, collaborating with hardware, firmware, and data-center teams.

You will use Python and SQL to produce scalable solutions, leveraging Redshift, Spark, and BI tools to deliver actionable insights and

Qualifications

  • 3+ years of non-internship professional software development experience.
  • 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience.
  • Bachelor's degree or foreign equivalent in Computer Science, Engineering, Mathematics, or a related field.
  • Experience programming with at least one software programming language.

Responsibilities

  • Design, build, and operate production data pipelines that ingest diagnostic, telemetry, and repair-ticket data from the Trainium and Inferentia fleet into a warehouse other teams query with confidence.
  • Build and own the metrics, alarms, and anomaly detection that surface test regressions, failure-rate shifts, and new failure signatures across hardware generations without a human going looking for them.
  • Build dashboards and visualizations that make fleet and test health legible to engineers, hardware partners, and leadership, covering failure rates, failure-signature breakdowns, first pass yield, and repair latency.
  • Analyze large-scale fleet data to find root cause behind failure trends, and separate genuine hardware faults from software defects and test noise — a distinction that decides whether a failure reaches a technician or an engineer.
  • Define the evidence standard that gates operational decisions, including whether a diagnostic has soaked long enough and cleanly enough in the fleet to move from observation mode into blocking production.
  • Improve data quality and pipeline reliability so downstream consumers trust the numbers without re-deriving them.
  • Write clear analyses and design documents for technical and non-technical readers, including leadership.

Skills

Python
SQL
Data pipelines

Education

Bachelor's degree or foreign equivalent in Computer Science, Engineering, Mathematics, or a related field

Tools

Amazon Redshift
Apache Spark
Grafana
QuickSight

Job description

Amazon MLA Vetting seeks a Software Development Engineer II to build data pipelines ingesting diagnostic results from thousands of servers, dashboards for test behavior, and metrics that indicate when tests flag failures before capacity loss. You will own alarms, production code, and anomaly detection, collaborating with hardware, firmware, and data-center teams.

You will use Python and SQL to produce scalable solutions, leveraging Redshift, Spark, and BI tools to deliver actionable insights and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SDE II: ML Acceleration Data Pipelines & Analytics
SDE II: ML Acceleration Data Pipelines & Analytics

Amazon Inc. • Austin (TX)

On-site
USD 144,000 - 194,000
Software Engineer, ML Diagnostics & Fleet Analytics
Software Engineer, ML Diagnostics & Fleet Analytics

Annapurna Labs (U.S.) Inc. • Austin (TX)

On-site
USD 120,000 - 180,000
SDE II: Infra for ML Workloads & Cloud Services
SDE II: Infra for ML Workloads & Cloud Services

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+1
SDE II: ML Infra – Capacity, Scheduling & Fleet Orchestration
SDE II: ML Infra – Capacity, Scheduling & Fleet Orchestration

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 150,000 - 210,000
Sign-on payments
Restricted stock units (RSUs)
Health insurance
+14
SDE II — Data Platform & AI Catalog at Scale
SDE II — Data Platform & AI Catalog at Scale

Amazon Inc. • Seattle (WA)

On-site
USD 144,000 - 194,000
SDE II: Secure Data Storage & AI Analytics
SDE II: Secure Data Storage & AI Analytics

Amazon • Santa Clara (CA)

On-site
USD 165,000 - 224,000
RSUs
Health insurance
401(k) matching
SDE II: Data Center Fleet Automation & Telemetry
SDE II: Data Center Fleet Automation & Telemetry

Amazon Inc. • Seattle (WA)

On-site
USD 144,000 - 194,000
SDE II: AI-Ready Marketing Data Pipelines
SDE II: AI-Ready Marketing Data Pipelines

Amazon Inc. • Boulder (CO)

On-site
USD 144,000 - 194,000
Health insurance
RSUs & 401(k)
Software Development Engineer, ML Acceleration, Trainium AI Systems, Annapurna Labs
Software Development Engineer, ML Acceleration, Trainium AI Systems, Annapurna Labs

Annapurna Labs (U.S.) Inc. • Austin (TX)

On-site
USD 120,000 - 180,000
SDE II — Ads AI Core Infra & Scalable Data Lake
SDE II — Ads AI Core Infra & Scalable Data Lake

Amazon • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1