Systems Development Engineer , Operations Infrastructure Services

Amazon

Tennessee

On-site

USD 123,000 - 166,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, Dental, Vision
401(k) Match
Paid Time Off

Job summary

Amazon in Nashville, TN seeks a Systems Development Engineer to design, build, and operate scalable monitoring infrastructure for high-volume device telemetry and network data. You will own end-to-end monitoring systems—from collection to deployment and ongoing maintenance—working with infrastructure teams to reduce downtime and improve operator experience.

The role emphasizes collaboration, design reviews, and hands-on implementation, with opportunities to apply AI techniques for incident

Qualifications

  • 3+ years designing or architecting scalable, reliable systems.
  • Experience automating, deploying, and supporting large-scale infrastructure.
  • Proficiency in at least one modern language (Python, Ruby, Golang, Java, C++, C#, Rust).
  • Experience with Linux/Unix and CI/CD pipelines.

Responsibilities

  • Design, build, and operate scalable monitoring infrastructure on AWS for device telemetry and network data.
  • Own end-to-end monitoring system components from build to deployment and maintenance.
  • Collaborate with stakeholders to identify monitoring gaps and implement automated solutions.
  • Apply AI/ML and generative AI techniques to improve detection and reduce alarm noise.
  • Improve system reliability and automation through thorough testing and architectural improvements.

Skills

System design
Reliability engineering
Scaling
Automation
Deployment
CI/CD
Python
Java
C++
Linux/Unix

Job description

Systems Development Engineer , Operations Infrastructure Services

Job ID: 10490709 | Amazon.com Services LLC

Join us in building infrastructure monitoring applications that power Amazon's global operations. You'll design and deliver services that process high-volume device telemetry, evaluate network linkages in real time, and provide operators with clear, actionable insights across thousands of sites worldwide.

We're an agile development team within Operations Infrastructure Services (OIS), part of Amazon Robotics, assisting fulfillment centers, delivery stations, and sortation centers globally. Our monitoring product gives operators a single, reliable view of device and network health—from live infrastructure telemetry to reliance-aware alarming that points to root causes instead of overwhelming teams with duplicate alerts. You'll have the opportunity to apply modern AI technologies, from AI-assisted incident detection to generative AI tooling that accelerates how we build and operate our network.

You might start your day reviewing a design for integrating metrics from a new device type being onboarded across OIS sites worldwide, then inspect runtime metrics to tune collection thresholds and data quality before shipping a monitoring improvement operators experience immediately. You'll partner with teammates on design reviews, participate in quick feedback loops, and own features end-to-end — working alongside infrastructure teams to deliver the monitoring experience needed to reduce building downtime and assistance. Throughout the day, you'll balance designing scalable monitoring solutions with hands‑on implementation, working across back-end services and data pipelines while assisting each other's growth and bringing your authentic perspective to the team.

Key job responsibilities
  • Design, build, and operate scalable monitoring infrastructure on AWS that ingests and processes high-volume device telemetry and network topology data
  • Own systems end-to-end across monitoring tooling and data pipelines, from build through deployment, operation, and ongoing maintenance
  • Partner with infrastructure and operations stakeholders to grasp monitoring gaps and translate them into reliable, automated solutions that reduce manual intervention
  • Apply AI/ML and generative AI techniques to improve detection quality, reduce alarm noise, and streamline operational workflows
  • Raise the bar on system reliability, operational excellence, and automation through thoughtful design, thorough testing, and continuous improvement
A day in the life

You might start by reviewing a design for integrating metrics from a new device type being onboarded across all Robotics buildings worldwide, then inspect runtime metrics to tune metric collection and quality before shipping a service improvement our operators experience immediately. Our team values partnership, quick feedback loops, and clear ownership.

Amazon offers a full range of benefits that assist you and eligible family members, including domestic partners. Benefits can vary by location, the number of regularly scheduled hours you work, length of employment, and job status such as seasonal or temporary employment. The benefits that generally apply to regular, full-time employees include:

  • 1. Medical, Dental, and Vision Coverage
  • 2. Maternity and Parental Leave Options
  • 3. Paid Time Off (PTO)
  • 4. 401(k) Plan

If you are not sure that every qualification on the list above describes you exactly, we'd still love to hear from you! At Amazon, we value people with unique backgrounds, experiences, and skillsets.

About the team

We're a close-knit, agile team that owns infrastructure monitoring for OIS within Amazon Robotics and ships to a worldwide operational fleet. We care deeply about the craft of software and foster each other's growth. Our vision centers on delivering reliable, intelligent monitoring that helps operators grasp their infrastructure at scale. You'll work with product and operations stakeholders to turn complex monitoring challenges into simple, elegant solutions. We're excited to welcome someone who align with our standards for well-architected software and collective problem-solving.

Basic Qualifications
  • 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
  • Experience in automating, deploying, and supporting large‑scale infrastructure
  • Experience programming with at least one modern language such as Python, Ruby, Golang, Java, C++, C#, Rust
  • Experience with Linux/Unix
  • Experience with CI/CD pipelines build processes
Preferred Qualifications
  • 3+ years of non‑internship professional software development experience
  • Experience with distributed systems at scale

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign‑on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

USA, TN, Nashville - 122,800.00 - 166,100.00 USD annually

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems Development Engineer , Operations Infrastructure Services
Systems Development Engineer , Operations Infrastructure Services

Socket.dev • Nashville (TN)

On-site
USD 123,000 - 166,000
Health benefits
401(k) matching
Paid time off
Systems Development Engineer , Operations Infrastructure Services
Systems Development Engineer , Operations Infrastructure Services

Amazon • Nashville (TN)

On-site
USD 123,000 - 166,000
Health insurance
401(k) matching
Paid time off
+1
System Development Engineer II, OIS Command Center
System Development Engineer II, OIS Command Center

Amazon • Nashville (TN)

On-site
USD 123,000 - 166,000
Health insurance
401(k) matching
Paid Time Off
+1
System Development Engineer II, OIS Command Center
System Development Engineer II, OIS Command Center

Amazon • Arlington (VA)

On-site
USD 129,000 - 175,000
Health insurance
401(k) matching
Paid time off
Software Development Engineer, Data Center Host Monitoring
Software Development Engineer, Data Center Host Monitoring

Amazon • Seattle (WA)

On-site
USD 143,700 - 194,400
Health insurance
RSUs
401(k) matching
+2
System Development Engineer II, Customer Experience Infrastructure
System Development Engineer II, Customer Experience Infrastructure

Socket.dev • Austin (TX)

On-site
USD 129,000 - 175,000
Health insurance
401(k) matching
Paid time off
+1
Systems Development Engineer I, OTS - Global Portfolio Strategy
Systems Development Engineer I, OTS - Global Portfolio Strategy

Amazon • Austin (TX)

On-site
USD 99,000 - 160,000
Medical, Dental, and Vision Coverage
401(k) Plan
Paid Time Off
Systems Development Engineer II, AWS Managed Operations (MO) AWSOM Team
Systems Development Engineer II, AWS Managed Operations (MO) AWSOM Team

Amazon Web Services (AWS) • Herndon (VA)

On-site
USD 129,000 - 175,000
Site Reliability Engineer - Software Ops and Scaling , One Material Handling System - Software,[...]
Site Reliability Engineer - Software Ops and Scaling , One Material Handling System - Software,[...]

Amazon • Bellevue (WA)

On-site
USD 129,000 - 175,000
System Development Engineer II, OIS Command Center
System Development Engineer II, OIS Command Center

Socket.dev • Nashville (TN)

On-site
USD 123,000 - 166,000
Medical, Dental, and Vision Coverage
Paid Time Off (PTO)
401(k) Plan
+1