SRE / DevOps Engineer

HeadHR

Town of Poland (NY)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

HeadHR in Town of Poland, NY is seeking a Senior SRE/DevOps to take ownership of CI/CD processes and AWS infrastructure management. The role involves extensive auditing and operational responsibilities, aiming for enhanced observability and cost control.

Applicants should have over 5 years of experience in production environments with deep AWS knowledge, particularly in ECS or EKS. This position offers an opportunity to lead crucial technical projects, enhance security procedures, and contribute to onboarding new team members.

Qualifications

  • 5+ years SRE / DevOps / Platform Engineering experience in production.
  • Deep knowledge of AWS services including ECS or EKS.
  • Experience with Infrastructure-as-Code using Terraform or Pulumi.

Responsibilities

  • Own CI/CD processes across services.
  • Manage AWS infrastructure with Terraform/Pulumi.
  • Control costs and maintain security baselines.

Skills

SRE/DevOps/Platform Engineering
AWS (ECS, EKS, IAM, RDS, ElastiCache)
Infrastructure-as-code (Terraform, Pulumi)
GitHub Actions
Container fundamentals (Docker)
Linux operations
Datadog in production
Incident response
Observability for Node.js and Python
Secrets management
Working English

Job description

Engagement context

Takeover of production AI mobile coaching platform. Runtimes: Node.js/NestJS, Python/FastAPI. Datastores: MongoDB, Postgres, Redis. Infra: AWS (ECS/EKS, RDS, ElastiCache, S3, VPC, IAM). CI: GitHub Actions. Observability: Datadog. Push: OneSignal. Errors: Crashlytics. Deep links: Branch. Vendors: Auth0, ElevenLabs, OpenAI, Amplitude, Terra, Strava.

Role summary

Senior SRE/DevOps. Owner: CI→production, IaC, deploy automation, observability, on-call, cost control, secrets, security baselines. Phase 1: measure and document. Phase 2: operate and transfer ownership.

First 90 days
  • Audit CI/CD (GitHub Actions): duration, flakiness, failure modes, secrets handling
  • Audit AWS: ECS/EKS topology, IAM posture, VPC layout, RDS, ElastiCache, S3
  • Audit Datadog: dashboards, tracked metrics, SLO/SLI gaps
  • Audit incidents (12m): count, severity, MTTR, RCA patterns
  • Vendor inventory: Auth0, OneSignal, ElevenLabs, OpenAI, Branch, Amplitude, Terra, Strava, Crashlytics — owners, billing, MFA, recovery plans
Ongoing
  • Own CI/CD across services
  • Own AWS infra (Terraform/Pulumi where suitable)
  • Cost control (OpenAI token spend, AWS rightsizing)
  • Security baselines: least-privilege IAM, secrets rotation, dependency scanning
  • Build onboarding for second SRE/DevOps hire
Kogo poszukujemy?
Must-have skills
  • 5+ years SRE / DevOps / Platform Engineering in production
  • AWS at depth — ECS or EKS, IAM (assume-role patterns, scoped policies), VPC, RDS, ElastiCache, S3, CloudWatch
  • Infrastructure-as-code — Terraform (preferred) or Pulumi
  • GitHub Actions — building reusable workflows, secret handling, reproducible builds
  • Container fundamentals — Dockerfile authoring, multi-stage builds, image hardening
  • Linux operations
  • Datadog in production — logs, APM, metrics, dashboards, monitors, SLO/SLI definition
  • Incident response — leading or co-leading real production incidents, writing post-mortems
  • Observability for both Node.js and Python services
  • Secrets management — AWS Secrets Manager, SOPS, or comparable
  • Working English
Nice-to-haves
  • Cost-optimisation discipline (FinOps, AWS Cost Explorer, Reserved Instance planning)
  • LLM cost monitoring (per-route OpenAI token spend dashboards)
  • Kubernetes specifically (we may or may not be on EKS)
  • Security baseline experience — CIS benchmarks, dependency scanning (Snyk, Dependabot), SAST tools
  • GDPR / data-residency considerations for cross-border data flows (US PL)
  • Mobile CI considerations — Fastlane, app-signing automation, TestFlight / Google Play internal tracks
  • On-call playbook authoring
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, DevOps
Member of Technical Staff, DevOps

Reactor • San Francisco (CA)

On-site
USD 100,000 - 160,000
Competitive salary and early equity
Visa sponsorship
Generous health, dental, and vision coverage
Senior SRE, Platform Software Engineer
Senior SRE, Platform Software Engineer

Jobtailor • California (MO)

On-site
USD 150,000 - 200,000
DevOps Engineer
DevOps Engineer

Oxbridge Health, Inc. • Norwalk (CT)

On-site
USD 120,000 - 160,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Platform Engineer
Platform Engineer

Relativity Space • Long Beach (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision coverage
401(k)
Generous parental leave
+1
Weekend Site Reliability Engineer
Weekend Site Reliability Engineer

Sporty Group • United States

Remote
USD 120,000 - 180,000
Remote first
Bonuses (quarterly)
28 days leave
+4
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

HTEC Group • United States

Hybrid
USD 90,000 - 130,000
SRE DevOps Engineer
SRE DevOps Engineer

Highbrow LLC • Atlanta (GA), San Francisco (CA), Overland Park (KS)

On-site
USD 110,000 - 130,000
Infrastructure & SRE Engineer — Secure AI Platform, Equity
Infrastructure & SRE Engineer — Secure AI Platform, Equity

Crosscheck Staffing • San Francisco (CA)

On-site
Head of SRE
Head of SRE

Wand AI • Palo Alto (CA)

On-site
USD 130,000 - 180,000