Senior Site Reliability Engineer (Remote)

Fathom.ai

United States

On-site

USD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Supportive environment for personal growth
Dynamic and collaborative team

Job summary

Fathom.ai is seeking a passionate Site Reliability Engineer (SRE) to leverage data and automation in scaling dynamic infrastructure. The role offers a unique blend of internal tooling development and infrastructure management to enhance customer experience.

Ideal candidates will have a strong background in observability, GitOps tooling, and familiarity with technologies such as Google Compute, Kubernetes, and Prometheus. Join a collaborative team and enjoy competitive compensation and growth opportunities.

Qualifications

  • Proficiency with Infrastructure as Code / GitOps tooling.
  • Foundation in Observability best practices and implementation.
  • Experience in a SaaS or PaaS environment.
  • Familiarity with tech stack: Google Compute, Kubernetes, Message Queues, Prometheus, ClickHouse, ArgoCD, Github Actions, Golang (Ruby / Rails is a bonus).

Responsibilities

  • Scale existing tools and improve automation for infrastructure.
  • Support platform across regions and enhance multi-regional capabilities.
  • Work with engineering to improve platform observability.

Skills

Infrastructure as Code / GitOps
Observability best practices
SaaS or PaaS environment experience
Google Compute
Kubernetes
Message Queues
Prometheus
ClickHouse
ArgoCD
Github Actions
Golang
Ruby / Rails

Job description

Role Overview

We are looking for an SRE who is passionate about leveraging data and automation to drive a highly dynamic infrastructure. The role is a unique blend of infrastructure and internal tooling to reduce friction at every step of delivering an amazing customer experience.

As part of our team, you'll play a pivotal role in scaling our infrastructure, reducing toil through automation, and contributing to our culture of innovation and continuous improvement.

What You’ll Do

By 30 Days:

  • Use your observability background to help scale our existing tools to new heights as we continue to grow the platform.
  • Enhance our existing automation for scaling our infrastructure and improve the development experience.

By 90 Days:

  • Play a key role in continuing to diversify and scale our platform across additional regions.
  • Evaluate options to replace our existing real-time data pipeline for enhanced multi-regional capabilities.
  • Provide platform support to all of engineering, using data-driven decision-making.

By 1 Year:

  • Work with engineering to re-evaluate what observability means for the Fathom platform, and drive improvements to remove friction
  • Help us design and implement improvements to our elastic multi-regional storage platform.
  • Drive platform improvements to enhance reliability and efficiency
Requirements
Hard Skills
  • Proficiency with Infrastructure as Code / GitOps tooling.
  • Foundation in Observability best practices and implementation.
  • Experience in a SaaS or PaaS environment.
  • Familiarity with our tech stack: Google Compute, Kubernetes, Message Queues, Prometheus, ClickHouse, ArgoCD, Github Actions, Golang. (Ruby / Rails is a bonus)
Soft Skills
  • Curiosity-driven with a focus on delivering results.
  • A generalist mindset with the ability to dive deep into a wide range of challenges.
  • Resilience and an ability to grind through complex problems.
  • Openness to disagreement and commitment to decisions once made.
  • Strong collaborative skills, with the ability to explain complex insights in an accessible manner.
  • Independence in managing one’s workload and priorities.
What You’ll Get
  • The opportunity to shape the dynamic platform of a growing company.
  • A role that balances scaling infrastructure, enabling development teams, and internal tooling development.
  • A chance to work with a dynamic and collaborative team.
  • Competitive compensation and benefits.
  • A supportive environment that encourages innovation and personal growth.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Engineer
Infrastructure Engineer

Pantera Capital • San Francisco (CA)

Remote
USD 180,000 - 240,000
Competitive compensation
Supportive work environment
Opportunities for innovation
Senior SRE: Scale Infra, Automate, Elevate Reliability
Senior SRE: Scale Infra, Automate, Elevate Reliability

Fathom.ai • United States

On-site
USD 100,000 - 130,000
Competitive compensation
Supportive environment for personal growth
Dynamic and collaborative team
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)
Senior SRE, Software Engineering (AWS / Scaling Infrastructure)

PulseRise Technologies • New York (NY)

Hybrid
USD 130,000 - 160,000
Platform SRE Engineer: Observability & Automation
Platform SRE Engineer: Observability & Automation

Fathom • United States

Remote
USD 120,000 - 180,000
Senior Site Reliability Engineer II
Senior Site Reliability Engineer II

Juniper Square • United States

On-site
USD 165,000 - 195,000
Health, dental, and vision care
Life insurance
Mental wellness coverage
+3
Site Reliability Engineer
Site Reliability Engineer

your Jared • Northern (KY)

Hybrid
USD 150,000 - 210,000
Remote work
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Health, dental, vision insurance
Generous PTO
Learning & development
+2
Senior SRE Engineer
Senior SRE Engineer

Compunnel, Inc. • Alpharetta (GA)

On-site
USD 140,000 - 190,000
Senior SRE Engineer
Senior SRE Engineer

flowcode • New York (NY)

Hybrid
USD 140,000 - 190,000