Senior Software Engineer - SRE

Embedded Shishya

San Francisco, New York, Portland (CA, NY, OR)

Hybrid

USD 201,000 - 251,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

Mercury is seeking a Site Reliability Engineer to join a newly formed SRE team that embeds with product groups. You will drive reliability, observability, and incident preparedness across services written in Haskell and TypeScript, helping product teams mature their operational practices.

You will run game days, refine runbooks, and steer SLOs toward customer outcomes, while championing reliability through reviews and cross-team collaboration.

Qualifications

  • Past Site Reliability Engineering or DevOps experience.
  • Measurable examples of influencing an organization towards greater reliability.
  • Significant experience with PostgreSQL.
  • Authored and operated Temporal workflows.
  • Experience with observability platforms like Grafana or Honeycomb; familiarity with OpenTelemetry.

Responsibilities

  • Embed with product teams to improve operational maturity around on-call, monitoring, alerting, and run books.
  • Run regular game day exercises with product teams to improve incident response.
  • Work with Haskell & TypeScript code to implement reliability techniques such as retries, better error handling, better logging, circuit breaking.
  • Steer SLOs towards meaningful customer outcomes that product teams are accountable for.
  • Champion reliability practices through design document and code reviews.
  • Identify observability gaps that hinder debugging, incident response, and business intelligence.
  • Advocate for longer-term improvements that non-product engineering teams can drive.
  • Participate in the product team's on-call rotation while embedding or as part of a broader engineering rotation.

Skills

Site Reliability Engineering
DevOps
PostgreSQL
Temporal workflows
OpenTelemetry
Observability
Grafana
Honeycomb

Tools

Temporal workflows
OpenTelemetry
Grafana
Honeycomb

Job description

When the Tarr Steps, a footbridge assembled of heavy stones in Exmoor National Park in England, washed away in a flood in 1942, the Royal Engineers rebuilt it. Then it washed away again in 1952. So the Royal Engineers heaved in more stones: on and on, like a Sisyphean lesson in absurdity.

Of course, they do this because the Tarr Steps are a historical monument, believed to be built during the Bronze Age. Modern bridge building looks a lot different and generally requires less upkeep. While we appreciate the charm of ancient things, Mercury is engineering the future of banking*. "Quaint" and "archaic" are not values we seek in our systems. Rebooting a tumbling server and hoping a flood of requests doesn't wash it away can be an emergency tactic, but we actively seek out durable solutions over monotonous ops work. You'll bring this ethos and teach others how to live it too.

Up to this point, the Stability team at Mercury has primarily built platform-level constructs that product teams adopt. However, those teams are asking for us to work more closely with them to mature their implementations and practices. We are creating an SRE team that will rotate through product teams. You will understand the team's domain, identify opportunities for improvements in reliability/observability/performance/preparedness and create self-reinforcing, virtuous cycles.

As part of this role, you will:
  • Embed with product teams, helping them improve their operational maturity by setting up and refining practices around on-call, monitoring, alerting, and run books
  • Run regular game day exercises with product teams, helping them feel more prepared to investigate and quickly remediate incidents
  • Jump into application code written in Haskell & TypeScript and implement reliability techniques such as retries, better error handling, better logging, circuit breaking, etc
  • Steer SLOs towards meaningful customer outcomes that product teams are accountable for. Those SLOs become a strong signal for whether the product is working as intended
  • Champion reliability practices through design document and code reviews
  • Identify observability gaps that hinder debugging, incident response, and business intelligence and help close those
  • Advocate for longer-term improvements that non-product engineering teams can drive
  • Participate in the product team's on-call rotation while embedding or as part of a more general engineering rotation, helping us improve processes and how we learn from incidents
The ideal candidate for the role:
  • Has past Site Reliability Engineering or DevOps experience
  • Has measurable examples of influencing an organization towards greater reliability
  • Has significant experience with PostgreSQL
  • Has authored and operated Temporal workflows
  • Has experience with observability platforms like Grafana or Honeycomb
  • Has familiarity with OpenTelemetry

If this role interests you, we invite you to explore our public demo at personal-demo.mercury.com.

*Mercury is a fintech company, not an FDIC-insured bank. Banking services provided through Choice Financial Group and Column N.A., Members FDIC.

Mercury values diversity & belonging and is proud to be an Equal Employment Opportunity employer. All individuals seeking employment at Mercury are considered without regard to race, color, religion, national origin, age, sex, marital status, ancestry, physical or mental disability, veteran status, gender identity, sexual orientation, or any other legally protected characteristic. We are committed to providing reasonable accommodations throughout the recruitment process for applicants with disabilities or special needs. If you need assistance, or an accommodation, please let your recruiter know once you are contacted about a role.

#LI-GC1

Total Rewards
The total rewards package at Mercury includes base salary, equity (stock options/RSUs), and benefits.

Our salary and equity ranges are highly competitive within the SaaS and fintech industry and are updated regularly using the most reliable compensation survey data for our industry. New hire offers are made based on a candidate’s experience, expertise, geographic location, and internal pay equity relative to peers.

Our target new hire base salary ranges for this role are the following:

US employees (any location):

$200,700 — $250,900 USD

Canadian employees (any location):

$189,700 — $237,100 CAD

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer - SRE
Senior Software Engineer - SRE

Mercury • Portland (OR)

On-site
USD 201,000 - 251,000
Senior Software Engineer - SRE
Senior Software Engineer - SRE

Mercury • New York (NY)

On-site
USD 201,000 - 251,000
Senior Software Engineer - SRE
Senior Software Engineer - SRE

Mercury • San Francisco (CA)

On-site
USD 201,000 - 251,000
Software Engineer - Infrastructure
Software Engineer - Infrastructure

Mercury • Portland (OR)

On-site
USD 122,000 - 158,000
Software Engineer - Infrastructure
Software Engineer - Infrastructure

Mercury • New York (NY)

On-site
USD 122,000 - 158,000
Software Engineer - Infrastructure
Software Engineer - Infrastructure

Mercury • San Francisco (CA)

On-site
USD 122,000 - 158,000
Software Engineer - Infrastructure
Software Engineer - Infrastructure

Mercury Technologies, Inc. • San Francisco (CA), Northern (KY)

On-site
USD 122,000 - 158,000
Stock options/RSUs
Benefits
Senior Software Engineer - Investments
Senior Software Engineer - Investments

Mercury • San Francisco (CA), New York (NY), Portland (OR)

On-site
USD 201,000 - 251,000
Software Engineer - Product
Software Engineer - Product

Mercury • New York (NY)

On-site
USD 122,000 - 158,000
Engineering Manager - Bank Accounts
Engineering Manager - Bank Accounts

Engg • United States

Remote
USD 201,000 - 251,000