Senior Software Engineer, Platform (DevOps)

ECP

Chicago (IL)

On-site

USD 120,000 - 160,000

Full time

47 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

ECP is seeking a Senior Infrastructure Engineer to own production AWS services, automate deployment, and improve observability. You will collaborate with delivery teams to provision environments and ensure security and audit readiness. The role emphasizes on-call responsibilities, self-service tooling, and AI-assisted incident workflows.

You will work primarily with Node.js/TypeScript? Not applicable here; focus on AWS infra and platform engineering practices to keep production reliable and

Qualifications

  • Built production AWS infrastructure as code using CloudFormation, Terraform, CDK, or Pulumi.
  • Designed networking and IAM: VPCs, cross-account access, data isolation.
  • Proficient with containers in production on ECS or Kubernetes, including capacity decisions.
  • Implemented security controls: least privilege, credential rotation, patch cadence.
  • Automated evidence for SOC2/HIPAA/ISO/PCI audits.
  • On-call experience and ability to walk through incidents end-to-end.
  • Applied AI in operational workflows for triage, alerts, or runbooks.
  • Built observability that helps engineers debug systems while protecting sensitive logs.
  • Wrote self-service tooling and runbooks to reduce tickets.

Responsibilities

  • Own defined production systems and their infrastructure code.
  • Handle patching, upgrades, and capacity planning; automate where possible.
  • Ensure security controls and audit readiness for production environments.
  • Improve observability with metrics, traces, and logs accessible to delivery teams.
  • Collaborate with delivery teams to provision queues, services, or environments ahead of needs.

Skills

AWS Infra as Code
Networking & IAM design
Containers in production
Security hardening
Audit evidence automation
On-call incident handling
AI in operations
Observability for debugging
Self-service & runbooks
Production code in Python/Bash/Go

Education

Bachelor’s degree in related field

Tools

CloudFormation
Terraform
CDK
Pulumi
ECS
Kubernetes

Job description

ECP is a market-leading SaaS solution that enables senior living communities to better care for their residents. ECP is used in over 8,000 communities.We'relooking to further expand by increasing the number of customers that use our software and increasing the scope of how we serve our customers by developing and releasing new products.

Senior living is deeply under-penetrated withsoftwareand ECP is one of the largest and fastest-growing software companies in the industry. We recently raised a growth round of equity to reinvest in our product, technology, and go-to-market. Our mission is to build world-class software that improves the quality of life for seniors and improves clinical, business, compliance, and operational performance for our customers.

About ECP

ECP is a market-leading SaaS solution that enables senior living communities to better care for their residents. ECP is used in over 8,000 communities.We'relooking to further expand by increasing the number of customers that use our software and increasing the scope of how we serve our customers by developing and releasing new products.

Senior living is deeply under-penetrated withsoftwareand ECP is one of the largest and fastest-growing software companies in the industry. We recently raised a growth round of equity to reinvest in our product, technology, and go-to-market. Our mission is to build world-class software that improves the quality of life for seniors and improves clinical, business, compliance, and operational performance for our customers.

The Role

Foundation Engineering is a new capability in ECP's Platform organization. Its charter is the path from code to production: the environment an engineer develops in, the pipeline that builds and ships their work, the infrastructure it runs on, and the tooling that makes all of it fast. Core Services is the four-person team inside it that runs ECP's AWS infrastructure, and you'd report to the cloud architect who leads it.

We're hiring because the team is behind. Some of the four get pulled into application work, and the delivery teams end up waiting on us for environments, services, and access they should be able to get on their own. You'd take on a share of the production estate, run it well, and work with those teams so that fewer of their requests need to come through us at all.

Everything runs on AWS. Most workloads are containers on ECS with Fargate, with EC2 for legacy services and our SQL Server clusters, and we use a wide range of AWS services beyond that. Infrastructure is CloudFormation, generated through an internal framework that handles templating and enforces our naming and tagging conventions. Application code is Node and TypeScript on the new platform and ColdFusion on the one we're migrating away from, and both will be in production for a while. You won't write ColdFusion, but you'll run the infrastructure under it. Databases are SQL Server and PostgreSQL. ECP is a multi-tenant HIPAA product, so what an engineer is allowed to see in production is a constraint on almost everything you'd design.

What You’ll Build

You'd own a defined set of production systems: designing them, writing the infrastructure code, operating them, and carrying the pager for them. That includes the routine work. Patching, runtime upgrades, certificate rotation, and adding capacity before a customer hits a limit are all part of the job, and we want as much of it automated as you can manage.

Security on those systems is yours as well. IAM and least privilege, rotating credentials and keys, keeping images patched, and building controls so they produce their own audit evidence. We are SOC2 compliant, infrastructure controls are a good part of what gets evidenced, and controls that document themselves save everyone a bad week before the audit.

The biggest thing we want from you in the first year is observability the delivery teams can use on their own. Right now they can't see enough of what they run in production. They need traces, metrics, and logs they can read without asking us, and because resident data can't be in what they see, this is more of a design problem than a purchasing one. On-call goes with this. It runs in two tiers today, delivery teams first and Core Services behind them for infrastructure. Our incident reviews are blameless and work reasonably well. The alerting in front of them is where the improvement is, and we want AI in that loop where it helps: triage, correlating signals nobody has time to read, a first draft of the postmortem.

The rest of the job is working with the delivery teams directly. When a team needs a queue, a new service, a database, or an environment that behaves like production, we want you in that conversation early, before it turns into a ticket. Common requests should end up in the provisioning framework so they stop being requests. And write things down. Too much of how this estate runs lives in one or two people's heads, and runbooks and documentation count as real work here.

Requirements
What We're Looking For

Required

  • Production infrastructure on AWS that you built as code and kept running. CloudFormation, Terraform, CDK, or Pulumi, we don't mind which.
  • Networking and IAM you designed yourself: VPCs, cross-account access, and keeping one workload's data isolated from another's.
  • Containers in production on ECS or Kubernetes, including the deployment strategy and the capacity decisions.
  • Hands-on security work on infrastructure you ran: least privilege, credential rotation, patching on a cadence.
  • Producing evidence for an audit, whether SOC 2, HIPAA, ISO, or PCI, ideally by automating it rather than assembling it by hand at the end.
  • You've been on call for something other teams depended on, and you can walk us through an incident you handled from start to finish.
  • You've used AI in an operational workflow, for triage, correlating alerts, running runbooks, or drafting incident writeups, and you can say what it changed.
  • You've built observability that other engineers used to debug their own systems, in an environment where log contents were sensitive.
  • You've cut the requests coming to your team by making things self-service, and you write runbooks that people other than you use.
  • You write real code, in Python, Bash, Go, or whatever your infrastructure is in. Much of this job is software.
  • You've moved other engineers along in how they use AI. It's already how you work yourself, and you know where it holds up and where it needs a human check. We expect our engineers to advocate for this.

Preferred

  • Running a production database, with restores and failover you've exercised yourself
  • A migration between AWS accounts or architectures done without downtime
  • Cost work: reserved capacity, savings plans, enterprise agreements
  • Healthcare or another regulated industry
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Platform (Developer Experience)
Senior Software Engineer, Platform (Developer Experience)

Ecp123 • Chicago (IL), Northern (KY)

On-site
USD 140,000 - 210,000
Senior Software Engineer, Clinical
Senior Software Engineer, Clinical

ECP • Chicago (IL)

On-site
USD 120,000 - 160,000
AWS DevOps Engineer
AWS DevOps Engineer

Decca Consulting LLC • Norwalk (CT)

On-site
USD 120,000 - 160,000
Senior DevOps Engineer
Senior DevOps Engineer

Decca Consulting LLC • United States

Remote
USD 140,000 - 210,000
Sr DevOps Engineer
Sr DevOps Engineer

Decca Consulting LLC • United States

Remote
USD 140,000 - 190,000
DevOps Engineer
DevOps Engineer

Decca Consulting • Norwalk (CT)

On-site
USD 120,000 - 180,000
Health insurance
Paid time off
Competitive salary
Cloud / Deployment Engineer (Nigeria)-Contract
Cloud / Deployment Engineer (Nigeria)-Contract

Talent Hackers, LLC • United States

Remote
USD 120,000 - 200,000
Platform Engineer II
Platform Engineer II

S27a • Northern (KY)

On-site
USD 120,000 - 180,000
Platform Engineer (Cloud Services)
Platform Engineer (Cloud Services)

Nexcess • United States

Remote
USD 145,000 - 195,000
Visa support
Relocation assistance
Daily lunches in the office
+4
Principal Software Engineer
Principal Software Engineer

American Bureau of Shipping • Spring (TX)

On-site
USD 140,000 - 210,000