Software Dev Engineer II, ECS - Experience

Amazon Development Center U.S., Inc.

Jersey City (NJ)

On-site

USD 140,000 - 210,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Amazon Development Center U.S., Inc. is seeking a Software Development Engineer II to own and build the distributed systems behind the ECS Scheduler at massive scale in AWS.

You will design scheduling algorithms, drive operational excellence, and ship features that millions of customers rely on daily. You will design, develop, and operate highly available distributed systems for container scheduling, mentor L4 engineers, and contribute to performance improvements across services used by AWS

Qualifications

  • 3+ years of professional software development experience.
  • 2+ years of design or architecture experience for large systems.
  • 1+ years of software development engineer experience.
  • 1+ years designing and developing large-scale, multi-tiered systems.

Responsibilities

  • Design, develop, and operate highly available, scalable distributed systems for container scheduling and task placement at massive scale (115M+ daily launches).
  • Own end-to-end design and delivery of ECS Scheduler features from design docs to production.
  • Build high-throughput, low-latency scheduling algorithms in Java for real-time placement across large compute fleets.
  • Design and implement resource management to optimize utilization and handle capacity constraints.
  • Lead operational excellence, on-call rotations, and production health.
  • Drive architectural discussions and review for new scheduler capabilities.
  • Mentor more junior engineers through code reviews and design guidance.
  • Identify and resolve performance bottlenecks in AWS-scale systems.
  • Collaborate across ECS teams on cross-cutting projects.
  • Build tooling to improve diagnostics and observability.

Skills

Java
Design patterns
Distributed systems
Performance optimization
Mentor engineers
On-call

Job description

What happens when a customer tells AWS ECS to keep 500 copies of their application running, spread across three Availability Zones, and then deploys a new version with zero downtime? The ECS Scheduler makes it happen. We are the team behind the service scheduling engine in Amazon Elastic Container Service, managing over 10 million customer services and processing 115 million task launches daily across every AWS region. We are looking for a Software Development Engineer II to own and build the distributed systems that make this work at scale. You will design scheduling algorithms, drive operational excellence, and ship features that millions of customers depend on every day.

Key job responsibilities
  • Design, develop, and operate highly available, scalable distributed systems for container scheduling and task placement at massive scale (115M+ daily launches)
  • Own the end-to-end design and delivery of features in the ECS Scheduler service, from design doc through production deployment
  • Build high-throughput, low-latency scheduling algorithms in Java that make real-time placement decisions across large compute fleets
  • Design and implement resource management systems that optimize utilization, handle capacity constraints, and support diverse workload types (services, tasks, daemon sets)
  • Lead operational excellence for services you own, including on-call rotations, COE investigations, deployment safety, and production health
  • Drive technical design documents and lead design reviews for new scheduler capabilities
  • Mentor L4 engineers through code reviews, design guidance, and technical leadership
  • Identify and resolve performance bottlenecks and scalability challenges in a system that operates at AWS-wide scale
  • Collaborate across ECS teams (Control Plane, Data Plane, Agent, Capacity) on cross-cutting projects
  • Build operational tooling and automation that improve diagnostics, accelerate root-cause analysis, and enhance system observability
A day in the life

Your morning might start with a code review for a teammate's placement algorithm change, followed by a quick check on the deployment pipeline. Mid-morning, you dive into a design document for a new scheduling capability. You sketch out the system interactions, model the expected throughput, and post your proposal for team review. After lunch, you pair with a teammate to debug a subtle production issue where a specific task placement pattern is causing higher-than-expected latency in one region. You trace the request flow through the scheduler, reproduce it locally, and draft a fix. Later in the afternoon, you work on operational tooling you have been building. Before wrapping up, you join a quick sync with the Capacity team to align on a new instance type integration, then review the on-call dashboard. Some days you are deep in scheduling algorithm internals optimizing bin-packing efficiency; other days you are designing a new API for service deployment controls or building automation that makes the next on-call shift smoother. No two days are the same, but the thread that ties them together is making container orchestration reliable and fast - so customers can focus on their applications, not their infrastructure.

About the team

Amazon Elastic Container Service (ECS) lets customers deploy containerized applications at scale. The ECS Scheduler team owns two core primitives:

ECS Services - When a customer creates an ECS service, the service scheduler runs and maintains the specified number of tasks simultaneously. If a task fails or stops, the scheduler launches a replacement. It spreads tasks across Availability Zones and manages rolling deployments. We serve over 10 million ECS services and process 115 million task launches daily across all AWS regions today.

ECS Managed Daemons - A new primitive that deploys exactly one daemon task on every managed instance in a capacity provider and manages the daemon lifecycle. When a managed instance registers with a cluster, ECS automatically starts daemon tasks before scheduling other tasks.

Every placement decision, every deployment rollout, every recovery from a failed task flows through the systems we build and operate. This is infrastructure that AWS customers trust with their production workloads.

You will work alongside engineers who care deeply about getting distributed systems right at scale. You will have the autonomy to own features end-to-end, the support of a team that takes operational excellence seriously, and the opportunity to see your work running across every AWS region in the world.

Basic Qualifications:
  • 3+ years of non-internship professional software development experience
  • 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • 1+ years of software development engineer or related occupational experience
  • 1+ years of designing and developing large-scale, multi-tiered, multi-threaded, embedded or dist
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Dev Engineer II, ECS - Experience
Software Dev Engineer II, ECS - Experience

Amazon • Jersey City (NJ)

On-site
USD 158,000 - 214,000
Senior Software Dev Engineer II - ECS Scheduler at Scale
Senior Software Dev Engineer II - ECS Scheduler at Scale

Amazon • Jersey City (NJ)

On-site
USD 158,000 - 214,000
Software Engineer II - ECS Scheduler, Scheduling at Scale
Software Engineer II - ECS Scheduler, Scheduling at Scale

Amazon Development Center U.S., Inc. • Jersey City (NJ)

On-site
USD 140,000 - 210,000
Software Development Engineer II, ECS Developer Platforms
Software Development Engineer II, ECS Developer Platforms

Amazon.com Services LLC • Jersey City (NJ)

On-site
USD 140,000 - 180,000
Software Development Engineer II, ECS Developer Platforms
Software Development Engineer II, ECS Developer Platforms

Socket.dev • New Jersey

On-site
USD 158,000 - 214,000
Health insurance
401(k) matching
Paid time off
+1
Software Development Engineer II, ECS Developer Platforms
Software Development Engineer II, ECS Developer Platforms

Amazon • Jersey City (NJ)

On-site
USD 158,000 - 214,000
Health insurance
401(k) matching
Paid time off
+1
Software Development Engineer, Amazon ECS
Software Development Engineer, Amazon ECS

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Software Development Manager, Elastic Container Service (ECS) (AWS)
Software Development Manager, Elastic Container Service (ECS) (AWS)

Amazon • Jersey City (NJ)

On-site
USD 203,000 - 275,000
Health insurance
401(k) matching
Paid time off
+1
Sr Manager, Software Dev, Elastic Container Service
Sr Manager, Software Dev, Elastic Container Service

Socket.dev • Seattle (WA)

On-site
USD 220,000 - 298,000
Health insurance
401(k) matching
Paid time off
+1
Software Development Manager, Elastic Container Service (ECS)
Software Development Manager, Elastic Container Service (ECS)

Amazon Web Services (AWS) • Jersey City (NJ)

On-site
USD 203,000 - 275,000
RSUs and sign-on payments
Comprehensive health insurance