Senior Infra Engineer: Observability

Railway

United States

On-site

USD 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Full health benefits including dependents
Strong equity grants
Equipment stipend
Autonomy with few meetings
Opportunity for creative solutions
Focus on growth and talent development

Job summary

Railway is seeking a Software Engineer to build scalable services and manage complex distributed systems. This high-impact role offers an environment rich in autonomy and ownership, where engineers can thrive while addressing novel problems. The position emphasizes collaboration, communication skills, and resilience in a startup atmosphere. Railway provides competitive salaries, health benefits for dependents, strong equity grants, and a unique culture that fosters growth and creativity.

Qualifications

  • Experience in building resilient and scalable distributed systems.
  • Familiarity with observability stacks including VictoriaMetrics and ClickHouse.
  • Ability to document engineering requirements and communicate effectively.

Responsibilities

  • Build ingestion pipelines for high-volume logs and metrics.
  • Create scalable alerting engines for real-time notifications.
  • Define infrastructure for immutable deployments.
  • Interface APIs for internal and external consumption.

Skills

Distributed systems
Fault tolerance
Scalable services
Communication skills
Prioritization in ambiguity
Problem-solving

Tools

Terraform
Ansible
Golang
Rust
TypeScript
GraphQL

Job description

Job description

Our core mission at Railway is to make software engineers higher leverage. We believe that people should be given powerful tools so that they can spend less time setting up to do, and more time doing.

Many infrastructure platforms simply focus on how you deploy your singular application, and now how these applications function in concert. Questions like “How do you build systems for zero downtime deployment”, “How do you do service-to-service communications”, etc are usually left up to the engineers to define.

At Railway, our goal is to be an all encompassing solution to all these problems. As such, we take special care as we define our networking infrastructure.

Note: Networking falls under the platform engineering umbrella. If you’re specialized, we’d love to chat! That said, we’d also like it noted you’re probably going to do a lot of non-networking + platform things

“But the world would be a better place if more engineers, like me, hated technology. The stuff I design, if I'm successful, nobody will ever notice. Things will just work, and will be self-managing”

- Radia Perlman

About The Role
  1. Build ingestion pipelines to consume 1M+ RPS streams of logs, metrics, and other telemetry
  2. Build scalable, fault tolerant alerting engines for notifying users, in real-time, of threshold breaches
  3. Craft rich backend observability APIs, working with product to build amazing experiences for instantly grokking their application
  4. Provide APIs to access realtime log/metrics streams to be consumed by the Dashboard and Product Teams
  5. Build Golang/Rust GRPC services from scratch capable of supporting tens of thousands of users, and the million+ to come.
  6. Define infrastructure that can be torn down, failed over, and reconstituted from scratch using principle of immutable infrastructure using Terraform and Ansible.
  7. Write Engineering Requirement Documents to take something from idea, to defined tasks, to implementation, to monitoring it’s success.
  8. Interface with our TypeScript and GraphQL edge to expose your microservice APIs for both internal and potentially external consumption

This is a high impact, high agency role with direct effect on company culture, trajectory, and outcome.

About You
  • A strong understanding of distributed systems. You enjoy building fault tolerant, resilient, and scalable services
  • Interests in VictoriaMetrics, ClickHouse, and other systems for building observability stacks from the ground up
  • A solid intuition about how long your solutions will last. All systems age. In startups, we can hope for 2-3 orders of magnitude, or 12-18mo.
  • The tact to implement your solution, creator monitors for it’s error boundaries, and document any requirements for when you’re not around
  • A great sense of direction and prioritization when it comes to dealing with the ambiguity of an early stage startup
  • A sense of grit to dive into a problem, implement a solution, scale that solution, and replace it when needed
  • A great set of communication skills for getting your point across, solution implemented, and beyond

We value and love to work with diverse persons from all backgrounds

Benefits and perks

At Railway, we provide best in class benefits. Great salary, full health benefits including dependents, strong equity grants, equipment stipend, and much more. For more details, check back on the main careers page.

Beyond compensation, there are a few things that we believe that make working at Railway truly unique:

  • Autonomy: We have very few meetings. Just a Monday and a Friday to go over the Company Board. We think your time is sacred, whether it's at work, or outside of work.
  • Ownership: We're a company with a high ownership, high autonomy culture. We hope that you'll come in, help us, and over the course of many years do the best work of your life. When we bring you onboard, we expect you to change the company.
  • Novel problems/solutions: We're a startup that's well funded, with cool problems, which lets us implement novel solutions! We abhor “busywork” and think, whether it's community, engineering, operations, etc there's always opportunity for creative and high leverage solutions.
  • Growth: We want you to grow with us, but we know that talent is loaned, so when you figure out what area you want to grow in next, whether it's at Railway or outside, we'll make sure you land there.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Infra Engineer: Baremetal Orchestration
Senior Infra Engineer: Baremetal Orchestration

Railway • United States

Remote
USD 90,000 - 130,000
Great salary
Full health benefits including dependents
Strong equity grants
+3
Infra Engineer - Datacenters
Infra Engineer - Datacenters

Railway • San Francisco (CA)

Remote
USD 140,000 - 190,000
Health benefits
Equity grants
Equipment stipend
Infra Engineer - Datacenters
Infra Engineer - Datacenters

Railway • United States

Remote
USD 120,000 - 180,000
Great salary
Full health benefits including dependents
Strong equity grants
+1
Remote Datacenter Infra Engineer - Ownership & Impact
Remote Datacenter Infra Engineer - Ownership & Impact

Railway • San Francisco (CA)

On-site
USD 140,000 - 190,000
Senior Observability Engineer: Scale Telemetry & Resilience
Senior Observability Engineer: Scale Telemetry & Resilience

Railway • United States

Remote
USD 90,000 - 130,000
Full health benefits including dependents
Strong equity grants
Equipment stipend
+3
Software Engineer, Observability
Software Engineer, Observability

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 175,000
Equity
Healthcare
Mentorship & events
+2
Senior Site Reliability Engineer, Observability New York, NY, United States
Senior Site Reliability Engineer, Observability New York, NY, United States

Ripple • New York (NY)

On-site
USD 160,000 - 200,000
Senior Systems Engineer
Senior Systems Engineer

Cox Automotive Inc. • Atlanta (GA)

On-site
USD 92,000 - 154,000
Senior Solution Engineer (Observability & Linux, North America, Remote)
Senior Solution Engineer (Observability & Linux, North America, Remote)

VictoriaMetrics • Seattle (WA)

Remote
USD 140,000 - 180,000
Remote-first culture
Flexible PTO
Startup atmosphere
Engineering Manager, Site Reliability
Engineering Manager, Site Reliability

Radar Labs • New York (NY)

On-site
USD 180,000 - 240,000
Stock options
401(k) with match
HQ in Flatiron, NYC
+6