Site Reliability Engineer

Stott and May

New York (NY)

On-site

USD 120,000 - 140,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Stott and May seeks a Site Reliability Engineer to be the first technical hire in the US, based in New York with 3 days in the office. The role focuses on reliability, scaling, and incident response for a fast-growing FX hedging platform.

You will work closely with engineering, infrastructure and business stakeholders, bridging US and UK teams, and contributing to runbooks, automation, and continuous improvement of support processes.

Qualifications

  • Strong background in application support, site reliability engineering or production support.
  • Experience supporting production environments and resolving live incidents.
  • Understanding of SRE principles, including SLIs, SLOs, error budgets and blameless post-mortems.
  • Experience with observability and monitoring tools such as Datadog.
  • Experience with cloud platforms such as AWS or Azure.
  • Proficiency in at least one programming or scripting language (TypeScript, .NET, Bash, PowerShell).
  • Knowledge of SQL and databases.
  • Understanding of networking fundamentals, Linux and Windows systems.
  • Excellent communication skills for technical and non-technical stakeholders.

Responsibilities

  • Investigate, own and resolve incidents across the platform and infrastructure.
  • Work with cross-functional teams for rapid production issue response.
  • Build and maintain runbooks, post-mortems and internal documentation.
  • Improve observability through monitoring, alerting and dashboards.
  • Reduce operational overhead through automation and reliability-focused engineering.
  • Contribute to the improvement of support processes, tooling and applications.
  • Train, coach and support UK on-call support rota.

Skills

SRE/Production Support
Incident response
Observability (Datadog)
AWS/Azure
TypeScript
.NET
Bash
PowerShell
SQL
Linux
Windows
Communication skills

Tools

Datadog
Terraform
CloudFormation
CDK
Docker
TeamCity

Job description

Site Reliability Engineer – AWS, Azure, IaC, Typescript, .NET – NY (office 3 days a week) - $120,000 - $140,000

Do you want to be the first technical hire in the US for a fast-growing leading FX hedging platform?

Do you want to be the cornerstone for the team as it grows?

This is a fantastic opportunity to a well-established UK business that has recently received some significant investment and is now breaking into the US market. The UK has been comfortably managing the UK coverage; however, they are growing quickly and they need a full-time US presence.

They are now looking for an Site Reliability Engineer to become the first technical hire in the US. This is a high-impact role for someone who enjoys solving complex production issues, improving platform reliability and working closely with engineering, infrastructure and business stakeholders. As Site Reliability Engineer, you will play a key role in ensuring the reliability, scalability and smooth operation of the company’s platform and supporting infrastructure. You will contribute through incident response, proactive monitoring, automation, documentation and continuous improvement of support processes. You will also act as the on-the-ground engineering presence in the New York office, extending technical coverage into US working hours and helping bridge collaboration between US-based stakeholders and the UK engineering team.

Duties and responsibilities
  • Investigate, own and resolve incidents across the platform and infrastructure.
  • Work closely with cross-functional teams to ensure a rapid and effective response to production issues.
  • Build and maintain runbooks, post-mortems and internal documentation to support knowledge sharing.
  • Improve observability through monitoring, alerting and dashboards.
  • Reduce operational overhead through automation and reliability-focused engineering.
  • Contribute to the ongoing improvement of support processes, tooling and applications.
  • Train, coach and support members of the UK on-call support rota.
Key skills
  • A strong background in application support, site reliability engineering or production support.
  • Experience supporting production environments and resolving live incidents.
  • Understanding of SRE principles, including SLIs, SLOs, error budgets and blameless post-mortems.
  • Experience with observability and monitoring tools such as Datadog.
  • Experience with cloud platforms such as AWS or Azure.
  • Proficiency in at least one programming or scripting language, such as TypeScript, .NET, Bash or PowerShell.
  • Knowledge of SQL and experience working with databases.
  • Understanding of networking fundamentals, Linux and Windows systems.
  • Strong problem-solving skills and the ability to remain calm under pressure.
  • Excellent communication skills, with the ability to explain technical concepts to technical and non-technical stakeholders.
Nice to have
  • Experience with infrastructure-as-code tools such as Terraform, CloudFormation or CDK.
  • Knowledge of DevOps deployment practices and tooling such as TeamCity, Octopus Deploy or Docker.
  • Experience in the financial services or fintech sector.
  • Exposure to AI/ML platforms, prompt engineering or integrations with LLM APIs.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hybrid SRE - AWS/Azure | First US Tech Hire (NY)
Hybrid SRE - AWS/Azure | First US Tech Hire (NY)

Stott and May • New York (NY)

Hybrid
USD 120,000 - 140,000
Site Reliability Engineer
Site Reliability Engineer

Longbridge Singapore • New York (NY)

On-site
USD 140,000 - 190,000
Competitive compensation
Growth opportunities
Site Reliability Engineer
Site Reliability Engineer

Longbridge • Dallas (TX)

On-site
USD 120,000 - 170,000
Experienced SRE | Diversified Strategies Hedge Fund
Experienced SRE | Diversified Strategies Hedge Fund

Techfellow Limited • New York (NY)

Hybrid
USD 117,000 - 250,000
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)

PulseRise Technologies • New York (NY)

Hybrid
USD 130,000 - 160,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Health, dental, vision insurance
Generous PTO
Learning & development
+2
Specialist Site Reliability Engineer
Specialist Site Reliability Engineer

Harvey Nash • New York (NY)

Hybrid
USD 120,000 - 160,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

Remote
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Site Reliability Engineer
Site Reliability Engineer

Longbridge Securities • Town of Texas (WI)

On-site
USD 100,000 - 130,000
Competitive compensation package
Growth opportunities
Site Reliability Engineer
Site Reliability Engineer

Piper Sandler & Co. in • New York (NY)

On-site
USD 120,000 - 150,000