Software Engineer, ML Diagnostics & Fleet Analytics

Annapurna Labs (U.S.) Inc.

Austin (TX)

In loco

USD 120.000 - 180.000

Tempo pieno

3 giorni fa
Candidati tra i primi
Generatore di candidature

Distinguiti per questa posizione — genera un curriculum e una lettera di presentazione personalizzati in circa un minuto.

Supera i filtri ATS

Descrizione del lavoro

Annapurna Labs (U.S.) Inc. is hiring a Software Development Engineer II to build data pipelines and dashboards that aggregate diagnostic results across tens of thousands of servers.

You will own the metrics, alarms, and anomaly detection that surface test health and failure trends, with a focus on actionable insights for production gates. You will write production code, collaborate with hardware, firmware, provisioning, and data center operations, and ensure data quality.

Competenze

  • 3+ years of non-internship professional software development experience.
  • 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • Bachelor's degree or foreign equivalent in Computer Science, Engineering, Mathematics, or a related field
  • Experience programming with at least one software programming language

Mansioni

  • Design, build, and operate production data pipelines that ingest diagnostic, telemetry, and repair-ticket data from the Trainium and Inferentia fleet into a warehouse other teams query with confidence.
  • Build and own the metrics, alarms, and anomaly detection that surface test regressions, failure-rate shifts, and new failure signatures across hardware generations without a human going looking for them.
  • Build dashboards and visualizations that make fleet and test health legible to engineers, hardware partners, and leadership, covering failure rates, failure-signature breakdowns, first pass yield, and repair latency.
  • Analyze large-scale fleet data to find root cause behind failure trends, and separate genuine hardware faults from software defects and test noise - a distinction that decides whether a failure reaches a technician or an engineer.
  • Define the evidence standard that gates operational decisions, including whether a diagnostic has soaked long enough and cleanly enough in the fleet to move from observation mode into blocking production.
  • Improve data quality and pipeline reliability so downstream consumers trust the numbers without re-deriving them.
  • Write clear analyses and design documents for technical and non-technical readers, including leadership.

Conoscenze

Python
SQL
Data pipelines
Distributed systems

Formazione

Bachelor's degree in CS/Engineering/Math

Strumenti

Amazon Redshift
Apache Spark
Grafana
QuickSight

Descrizione del lavoro

Annapurna Labs (U.S.) Inc. is hiring a Software Development Engineer II to build data pipelines and dashboards that aggregate diagnostic results across tens of thousands of servers.

You will own the metrics, alarms, and anomaly detection that surface test health and failure trends, with a focus on actionable insights for production gates. You will write production code, collaborate with hardware, firmware, provisioning, and data center operations, and ensure data quality.

Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

ML Hardware Systems Engineer – Fleet Automation & Debugging
ML Hardware Systems Engineer – Fleet Automation & Debugging

Amazon • Austin (TX)

In loco
USD 136.000 - 184.000
Health insurance
401(k) matching
Paid time off
+1
Systems Development Engineer II — Automation & Fleet Ops
Systems Development Engineer II — Automation & Fleet Ops

Amazon Inc. • Austin (TX), Northern (KY)

Ibrido
USD 129.000 - 175.000
SDE II: ML Diagnostics & Fleet Analytics
SDE II: ML Diagnostics & Fleet Analytics

Amazon • Austin (TX)

In loco
USD 144.000 - 194.000
Health insurance
RSUs
Sign-on bonus
+2
Automation & Fleet Systems Engineer II
Automation & Fleet Systems Engineer II

Amazon • Austin (TX)

In loco
USD 129.000 - 175.000
Machine Learning Hardware Platform Engineer
Machine Learning Hardware Platform Engineer

Amazon Web Services (AWS) • Austin (TX)

In loco
USD 136.000 - 184.000
Health insurance
RSUs / restricted stock units
401(k) match
+2
Software Engineer I – Cloud & ML Systems
Software Engineer I – Cloud & ML Systems

Amazon • Seattle (WA)

In loco
USD 129.000 - 148.000
EAP
Mental health support
Medical advice line
+1
SDE II: ML Acceleration Data Pipelines & Analytics
SDE II: ML Acceleration Data Pipelines & Analytics

Amazon Inc. • Austin (TX)

In loco
USD 144.000 - 194.000
AI Hardware Systems Engineer, Annapurna Labs, Trainium Machine Learning Fleet Operations
AI Hardware Systems Engineer, Annapurna Labs, Trainium Machine Learning Fleet Operations

Amazon Web Services (AWS) • Austin (TX)

In loco
USD 136.000 - 184.000
Health insurance
RSUs / restricted stock units
401(k) match
+2
ML Systems Software Engineer II — Cloud & Hardware
ML Systems Software Engineer II — Cloud & Hardware

Amazon Web Services (AWS) • Cupertino (CA)

In loco
USD 165.000 - 224.000
Health insurance
RSUs
401(k) matching
+2
Software Development Engineer, ML Acceleration, Trainium AI Systems, Annapurna Labs
Software Development Engineer, ML Acceleration, Trainium AI Systems, Annapurna Labs

Annapurna Labs (U.S.) Inc. • Austin (TX)

In loco
USD 120.000 - 180.000