Senior Software Engineer - SRE

Socure

Town of Concord (NY)

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Socure seeks exceptional Site Reliability Engineers to own and operate mission-critical, production-grade systems in the cloud. You will focus on preventing incidents and raising the reliability bar through automation and observability.

You will manage end-to-end AWS infrastructure, Kubernetes (EKS) platforms, and CI/CD pipelines. This role emphasizes design, automation, and proactive incident remediation in a fast-paced environment.

Qualifications

  • Experience building production-grade infrastructure on AWS.
  • Proficient in Kubernetes internals, scheduling, networking, and storage.
  • Strong coding skills in Go or Python for automation.
  • Hands-on with GitOps workflows and CI/CD pipelines.
  • Ability to define and use SLIs/SLOs for reliability decisions.

Responsibilities

  • Own end-to-end AWS infrastructure to keep services highly available and scalable.
  • Design and operate Kubernetes platforms (EKS) and automate platform tasks.
  • Enhance reliability via observability, automation, and incident response.

Skills

Go
Python
Automation
SRE fundamentals
Incident response
Observability
SLIs/SLOs
Cloud security

Tools

Terraform
Kubernetes
Amazon EKS
GitHub Actions
ArgoCD
Datadog

Job description

Why Socure?

Socure is building the identity trust infrastructure for the digital economy — verifying 100% of good identities in real time and stopping fraud before it starts. The mission is big, the problems are complex, and the impact is felt by businesses, governments, and millions of people every day.

We hire people who want that level of responsibility. People who move fast, think critically, act like owners, and care deeply about solving customer problems with precision. If you want predictability or narrow scope, this won’t be your place. If you want to help build the future of identity with a team that holds a high bar for itself — keep reading.

We are hiring exceptional Site Reliability Engineers who take pride in building and operating mission‑critical, production‑grade systems. This role is for engineers who own what they build, thrive in high‑pressure environments, and continuously raise the reliability and operational bar.

You will work at the intersection of cloud infrastructure, Kubernetes, automation, and observability, with a strong focus on preventing incidents rather than reacting to them.

What You’ll Own
  • End-to-end ownership of highly available, scalable AWS infrastructure

  • Design, operation, and continuous improvement of Kubernetes (EKS) platforms

  • Reliability of production systems through strong observability, automation, and SLOs

  • CI/CD systems that enable safe, fast, and repeatable deployments

  • Infrastructure defined and enforced through Terraform and GitOps

  • Incident response, root cause analysis, and long-term remediation

  • Raising operational standards through automation, documentation, and best practices

Technical Requirements:

We’re looking for engineers who have actually built, run, and scaled real production systems in the following areas:

Cloud & Infrastructure
  • Deep AWS expertise - networking, compute, IAM, scaling, security

  • Strong experience managing infrastructure using Terraform at scale

Kubernetes & Platform Engineering
  • Very strong Kubernetes fundamentals (internals, scheduling, networking, storage)

  • Hands‑on experience operating Amazon EKS in production environments

  • Experience troubleshooting complex, multi‑layer Kubernetes issues

Coding & Automation
  • Ability to write clean, maintainable, production-quality code in: Go/ Python

  • Strong automation mindset — eliminating toil through code

CI/CD & GitOps
  • Proven experience building and operating CI/CD pipelines

  • Hands‑on experience with:

    • GitHub (Actions or integrations)

    • ArgoCD and GitOps-based deployment workflows

Observability & Reliability
  • Strong understanding of observability principles: metrics, logs, traces, and alerting

  • Hands‑on experience with Datadog or similar tool for:

    • Infrastructure and Kubernetes monitoring

    • Application performance monitoring (APM)

    • Alerting, dashboards, and incident detection

  • Experience defining and using SLIs/SLOs to drive reliability decisions

Ability to turn observability data into actionable operational improvements

Socure is an equal opportunity employer that values diversity in all its forms within our company. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. If you need an accommodation during any stage of the application or hiring process—including interview or onboarding support—please reach out to your Socure recruiting partner directly.

Follow Us!

YouTube | LinkedIn | X (Twitter) | Facebook

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer - SRE
Senior Software Engineer - SRE

Socure • City of Albany (NY)

On-site
USD 160,000 - 180,000
Senior SRE: Cloud, Kubernetes & Automation
Senior SRE: Cloud, Kubernetes & Automation

Socure • Carson City (NV)

On-site
USD 150,000 - 190,000
Senior Backend Engineer
Senior Backend Engineer

Socure • New York (NY)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer - Cloud, Kubernetes & Automation
Senior Site Reliability Engineer - Cloud, Kubernetes & Automation

Socure • Town of Concord (NY)

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer – Cloud, Kubernetes & Automation
Senior Site Reliability Engineer – Cloud, Kubernetes & Automation

Socure • Seattle (WA)

On-site
USD 160,000 - 180,000
Software Development Engineer in Test
Software Development Engineer in Test

Socure • San Francisco (CA), Seattle (WA), New York (NY)

Hybrid
USD 170,000 - 230,000
Software Development Engineer in Test
Software Development Engineer in Test

Socure • Town of Concord (NY)

On-site
USD 120,000 - 180,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Engg • Tempe (AZ)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer — Cloud & Kubernetes
Senior Site Reliability Engineer — Cloud & Kubernetes

Socure • New York (NY)

On-site
USD 160,000 - 180,000
Senior Business Systems Engineer – Product & Engineering Platforms
Senior Business Systems Engineer – Product & Engineering Platforms

Apply • Northern (KY)

Hybrid
USD 120,000 - 180,000