Senior Site Reliability Engineer (SRE

Govserviceshub

New York (NY)

On-site

USD 130,000 - 160,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

A financial technology startup is seeking a Senior Site Reliability Engineer (SRE) to lead the establishment of SRE practices and ensure reliability and performance of critical systems. This remote position is ideal for individuals who bring expertise in AWS, Infrastructure as Code, and CI/CD automation. The role offers the opportunity to build and own the entire SRE practice within a fast-paced, innovative environment focused on consumer finance.

Qualifications

  • Proven experience leading incident response and postmortem processes for high-availability production systems.
  • Deep expertise in designing highly available architectures such as EC2 and Fargate.
  • Strong experience with AWS cloud infrastructure and IaC tools.

Responsibilities

  • Lead incident response and develop sustainable on-call practices.
  • Build and maintain self-service observability tools.
  • Create and maintain Infrastructure as Code (IaC) using Terraform or CloudFormation.

Skills

Incident response
High availability architectures
AWS cloud infrastructure
CI/CD automation
Observability and monitoring stacks
Scripting/programming in Python

Tools

Terraform
CloudFormation
GitHub Actions
Datadog
Prometheus
ELK

Job description

New York, United States | Posted on 11/13/2025

Title: Senior Site Reliability Engineer (SRE) Location: Remote

AboutJanuary

AtJanuary, we’re transforming the lives of borrowers by bringing humanity to consumer finance. Our data-driven products empower financial institutions to streamline collections and help borrowers regain financial stability and control over their lives. We’re not just expanding access to credit — we’re restoring dignity and paving the way for millions to achieve financial freedom.

Aboutthe Role

As a Senior Site Reliability Engineer (SRE), you will establish SRE practices from the ground up — ensuring reliability, scalability, and performance as January scales from thousands to millions of borrowers. You’ll architect resilient infrastructure, design modern observability solutions, and build sustainable on-call processes that evolve with our rapid growth.

Your work will directly address scaling challenges including database optimization, async workflow infrastructure, and data pipeline reliability — enabling the engineering team to ship confidently and efficiently.

KeyResponsibilities
  • Lead incident response and develop sustainable on-call practices, including runbooks, blameless postmortems, and continuous improvement to reduce MTTR.
  • Build and maintain self-service observability tools (Datadog, Prometheus, ELK) for proactive monitoring and troubleshooting.
  • Create and maintain Infrastructure as Code (IaC) using Terraform or CloudFormation for consistent, secure AWS environments.
  • Partner with development teams to architect resilient, scalable infrastructure for critical components like databases, networking, async workflows, and data pipelines.
  • Design and implement robust CI/CD pipelines (GitHub Actions) with advanced deployment strategies (blue/green, canary).
  • Drive best practices in reliability and performance early in the design phase to future-proof January’s systems.
RequiredSkills & Experience
  • Proven experience leading incident response and postmortem processes for high-availability production systems.
  • Deep expertise in designing highly available architectures (EC2, Fargate, auto-scaling, health checks, graceful degradation).
  • Strong experience with AWS cloud infrastructure and IaC tools (Terraform, CloudFormation).
  • Hands-on experience with CI/CD automation using GitHub Actions or equivalent tools.
  • Proficiency in observability and monitoring stacks (Datadog, Prometheus, ELK).
  • Solid scripting/programming skills in Python (for automation, tooling, and debugging).
  • Excellent communication and documentation skills, with the ability to collaborate across engineering and platform teams.
Requirements
  • Cloud: AWS
  • IaC: Terraform, CloudFormation
  • CI/CD: GitHub Actions
  • Languages: Python
  • Infrastructure: EC2, Fargate
AdditionalDetails
  • Remote role (NYC-based preferred for hybrid collaboration).
  • Opportunity to build and own the entire SRE practice for a growing FinTech startup.
  • Fast-paced, innovative environment working on AI-forward consumer finance products.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE, Software Engineering (AWS / Scaling Infrastructure)
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)

PulseRise Technologies • New York (NY)

Hybrid
USD 130,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Mission Staffing • New York (NY)

Hybrid
USD 140,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Discretionary incentive plan
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

OutSolve • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work environment
Competitive compensation
Professional development opportunities
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

National Black MBA Association • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Annual discretionary plan
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Health, dental, vision insurance
Generous PTO
Learning & development
+2
Senior Software Engineer, Site Reliability
Senior Software Engineer, Site Reliability

Upstart • United States

Hybrid
USD 166,000 - 231,000
Competitive compensation
401(k) matching
Employee Stock Purchase Plan
+5
Site Reliability Engineer
Site Reliability Engineer

Longbridge Singapore • New York (NY)

On-site
USD 140,000 - 190,000
Competitive compensation
Growth opportunities