Senior SRE Engineer

flowcode

New York (NY)

Hybrid

USD 140,000 - 190,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Flowcode is seeking a Senior Site Reliability Engineer (SRE) to strengthen reliability and infrastructure across our platforms in a hybrid New York setting. You will help grow our infrastructure strategy, improve observability, and partner with engineering teams to ensure scalable, resilient systems.

You will develop and operate cloud infrastructure, drive deployment best practices, and contribute to incident response, postmortems, and durable fixes in a fast-growing environment.

Qualifications

  • 4+ years in SRE, DevOps, or Platform Engineering.
  • Kubernetes expertise including troubleshooting and controllers/CRDs.
  • Terraform/OpenTofu proficiency and production state management.
  • GitOps workflows via ArgoCD and Helm deployments.
  • Production code in Go or Python with shell scripting.

Responsibilities

  • Improve availability, scalability, and resilience across Flowcode platforms.
  • Own key pieces of EKS-based infrastructure end-to-end.
  • Contribute to incident response and postmortems with durable fixes.
  • Support engineering teams with infra questions and day-to-day blockers.
  • Design and scale deployment pipelines with GitHub Actions.
  • Develop monitoring, logging, metrics, and SLOs for services.

Skills

Kubernetes
Terraform/OpenTofu
GitOps (ArgoCD)
Go or Python
Shell scripting
AWS (EKS, VPC, RDS, IAM)
GitHub Actions
Infrastructure leadership
Distributed systems

Tools

Datadog
Prometheus
Crossplane
Karpenter

Job description

Senior SRE Reliability Engineer

Location: New York, NY (Hybrid) / Remote

Department: Engineering

The Role

Flowcode is seeking a Senior Site Reliability Engineer (SRE) to work on reliability and infrastructure efforts across our platforms. This role will help grow and drive our infrastructure strategy, operational rigor and observability while building and supporting the systems and tooling required to support Flowcode's continued growth.

As an individual contributor within our engineering organization, you will develop and operate scalable cloud infrastructure, establish best practices around deployment and reliability, and partner closely with engineering teams to ensure systems are scalable, resilient and observable.

What You'll Do
Reliability & Infrastructure
  • Improve system availability, scalability, and resilience across Flowcode's platforms
  • Own key pieces of our EKS-based infrastructure end-to-end
  • Contribute to incident response and postmortems, turning findings into durable fixes
  • Support engineering teams with infrastructure questions, escalations, and day-to-day unblocking
Cloud & Platform Engineering
  • Manage and scale our core AWS footprint (EKS, VPC, RDS) through Infrastructure as Code (Terraform)
  • Enhance disaster recovery and failover mechanisms to protect mission-critical workloads
  • Collaborate with product engineering to streamline and optimize internal developer experience
CI/CD & Deployment Automation
  • Design and scale deployment pipelines using GitHub Actions
  • Expand GitOps practices and tooling through ArgoCD
  • Facilitate secure delivery with automated validation and progressive rollout strategies
Observability & Monitoring
  • Oversee and optimize the organization's monitoring, logging, and alerting infrastructure
  • Develop high-signal metrics, tracing, and visualization dashboards while minimizing operational noise
  • Establish and monitor Service Level Objectives for managed platform components
Qualifications
Required
  • 4+ years of professional experience across SRE, DevOps, or Platform Engineering domains
  • Technical proficiency in Kubernetes, including cluster troubleshooting and managing controllers or CRDs
  • Advanced Terraform or OpenTofu expertise, encompassing module architecture and production state management
  • Hands-on operational experience with GitOps workflows via ArgoCD and Helm-based deployments
  • Ability to author production-grade code in Go or Python alongside robust shell scripting
  • Mastery of core AWS services, specifically EKS, Networking/VPC, RDS, and IAM
  • Experience maintaining and scaling CI/CD automation using GitHub Actions within collaborative environments
  • Proven track record of leading infrastructure initiatives from initial design through to long-term operation
  • Background in supporting large-scale distributed systems within high-availability production environments
  • Adept at navigating interrupt-driven workflows, balancing strategic project delivery with day-to-day operational support
Preferred
  • Exposure to Crossplane or alternative Kubernetes-native solutions for infrastructure provisioning
  • Deep observability experience utilizing Datadog or Prometheus to engineer SLOs, high-signal dashboards, and intelligent alerting
  • Practical knowledge of modern secrets management frameworks and implementation
  • Experience optimizing cluster efficiency through autoscaling technologies such as Karpenter or Cluster Autoscaler

Flowcode is not for everyone. We hire with a pinhole lens - only those with the rare combination of intellectual horsepower, execution velocity, and uncompromising drive will thrive here. If you are seeking to operate at the highest levels of performance and impact, we want to meet you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE Engineer New New York, Hybrid
Senior SRE Engineer New New York, Hybrid

Flowcode • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior SRE Engineer: Cloud Reliability & Platform
Senior SRE Engineer: Cloud Reliability & Platform

flowcode • New York (NY)

Hybrid
USD 140,000 - 190,000
Senior SRE: Scale Cloud Infra, Observability & Pipelines
Senior SRE: Scale Cloud Infra, Observability & Pipelines

Flowcode • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior SRE
Senior SRE

Selby Jennings • Austin (TX)

On-site
USD 140,000 - 190,000
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)

PulseRise Technologies • New York (NY)

Hybrid
USD 130,000 - 160,000
Software Engineer, Infrastructure – San Francisco
Software Engineer, Infrastructure – San Francisco

Flow Engineering • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive salary and equity
Health, dental, and vision coverage
Flexible time off
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Health, dental, vision insurance
Generous PTO
Learning & development
+2
Site Reliability Engineer
Site Reliability Engineer

Harrison Clarke • New York (NY)

On-site
USD 120,000 - 160,000