Senior Infra Engineer: Baremetal Orchestration

Railway

United States

Remote

USD 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Great salary
Full health benefits including dependents
Strong equity grants
Equipment stipend
High ownership culture
Support for personal growth

Job summary

Railway is seeking a highly skilled engineer to build and maintain provisioning tools that empower software engineers. The role is focused on enhancing our internal infrastructure, with responsibilities that include optimizing performance, developing internal tools and observability systems, and creating resilient services in Golang and Rust. Candidates should have a strong understanding of distributed systems, hands-on experience with bare metal provisioning, and excellent communication skills. The position offers a high degree of autonomy and ownership, with best-in-class benefits.

Qualifications

  • Strong understanding of distributed systems and fault-tolerant architectures.
  • Experience with bare metal provisioning and configuration management.
  • Ability to build and operate internal tools focusing on developer experience.
  • Intuition about system longevity in start-up environments.
  • Skills in implementing solutions with proper monitoring.
  • Ability to prioritize and navigate ambiguity in startups.
  • Grit to tackle problems and iterate on solutions.
  • Strong communication skills for effective collaboration.

Responsibilities

  • Build and maintain host provisioning stack using PXE boot and Ansible.
  • Evolve orchestration engine for managing clusters and containers.
  • Optimize bin packing algorithm to enhance performance and reduce costs.
  • Own tooling for Railway engineers to facilitate their work.
  • Develop internal observability and alerting systems.
  • Design CI pipelines for infrastructure code.
  • Implement infrastructure using Terraform and Ansible.
  • Develop GRPC services in Golang/Rust.

Skills

Distributed systems
Bare metal provisioning
Configuration management
Internal tools development
Communication skills

Tools

Golang
Rust
Terraform
Ansible

Job description

Job description

Our core mission at Railway is to make software engineers higher leverage. We believe that people should be given powerful tools so that they can spend less time setting up to do, and more time doing.

Many infrastructure platforms simply focus on how you deploy your singular application, and now how these applications function in concert. Questions like “How do you build systems for zero downtime deployment”, “How do you do service‑to‑service communications”, etc are usually left up to the engineers to define.

At Railway, our goal is to be an all encompassing solution to all these problems. As such, we take special care as we define our networking infrastructure.

“But the world would be a better place if more engineers, like me, hated technology. The stuff I design, if I'm successful, nobody will ever notice. Things will just work, and will be self‑managing”.

- Radia Perlman

About the role
  • Build and maintain our host provisioning stack: PXE boot, Ansible, and burn-in agents that bring new bare metal online quickly and confidently
  • Continue to evolve our homegrown orchestration engine to manage clusters, containers, and VMs through a single lens
  • Optimize the efficiency of our bin packing algorithm to maximize utilization/performance and minimize costs
  • Own the internal tooling that Railway engineers use to interact with our fleet every day
  • Build out internal observability and alerting so we catch fleet problems before customers feel them
  • Design and maintain the CI pipelines that ship our infrastructure code safely
  • Define infrastructure that can be torn down, failed over, and reconstituted from scratch using principle of immutable infrastructure using Terraform and Ansible
  • Build Golang/Rust GRPC services from scratch capable of supporting millions of users
  • Write Engineering Requirement Documents to take something from idea, to defined tasks, to implementation, to monitoring its success

The arc of this role is more internal‑facing than user‑facing. You're building the platform that Railway engineers run on. This is a high impact, high agency role with direct effect on company culture, trajectory, and outcome.

About you
  • A strong understanding of distributed systems and what it takes to operate them. You enjoy building fault‑tolerant, resilient, and scalable services, and you care about what happens when they break at 3am
  • Hands‑on experience with bare metal provisioning, configuration management, and the unglamorous‑but‑critical work of getting hardware production‑ready
  • Comfort building and operating internal tools. You understand that developer experience inside the company matters as much as the product outside it
  • A solid intuition about how long your solutions will last. All systems age. In startups, we can hope for 2‑3 orders of magnitude, or 12‑18mo
  • The tact to implement your solution, create monitors for its error boundaries, and document any requirements for when you're not around
  • A great sense of direction and prioritization when it comes to dealing with the ambiguity of an early stage startup
  • A sense of grit to dive into a problem, implement a solution, scale that solution, and replace it when needed
  • A great set of communication skills for getting your point across, solution implemented, and beyond
Benefits and perks

At Railway, we provide best in class benefits. Great salary, full health benefits including dependents, strong equity grants, equipment stipend, and much more.

Beyond compensation, there are a few things that we believe that make working at Railway truly unique:

  • Autonomy: We have very few meetings. Just a Monday and a Friday to go over the Company Board. We think your time is sacred, whether it's at work, or outside of work.
  • Ownership: We're a company with a high ownership, high autonomy culture. We hope that you'll come in, help us, and over the course of many years do the best work of your life. When we bring you onboard, we expect you to change the company.
  • Novel problems/solutions: We're a startup that's well funded, with cool problems, which lets us implement novel solutions! We abhor “busywork” and think, whether it's community, engineering, operations, etc there's always opportunity for creative and high leverage solutions.
  • Growth: We want you to grow with us, but we know that talent is loaned, so when you figure out what area you want to grow in next, whether it's at Railway or outside, we'll make sure you land there.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Infra Engineer: Observability
Senior Infra Engineer: Observability

Railway • United States

Remote
USD 90,000 - 130,000
Full health benefits including dependents
Strong equity grants
Equipment stipend
+3
Developer Relations
Developer Relations

Railway • San Francisco (CA)

Remote
USD 90,000 - 150,000
Remote Developer Relations Architect
Remote Developer Relations Architect

Railway • San Francisco (CA)

On-site
USD 90,000 - 150,000
Infra Engineer - Datacenters
Infra Engineer - Datacenters

Railway • San Francisco (CA)

Remote
USD 140,000 - 190,000
Health benefits
Equity grants
Equipment stipend
Developer Relations
Developer Relations

Railway • United States

Remote
USD 80,000 - 120,000
Full health benefits, including dependents
Strong equity grants
Equipment stipend
+2
Remote Datacenter Infra Engineer - Ownership & Impact
Remote Datacenter Infra Engineer - Ownership & Impact

Railway • San Francisco (CA)

On-site
USD 140,000 - 190,000
Infra Engineer - Datacenters
Infra Engineer - Datacenters

Railway • United States

Remote
USD 120,000 - 180,000
Great salary
Full health benefits including dependents
Strong equity grants
+1
Infrastructure Engineer
Infrastructure Engineer

Monograph • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health Insurance
401k Plan
HSA/FSA
+7
Platform Engineer
Platform Engineer

AirOps • San Francisco (CA), New York (NY)

On-site
USD 130,000 - 160,000
Equity in a fast-growing startup
Competitive benefits package
Flexible time off policy
+2
Site Reliability / Infrastructure Engineer
Site Reliability / Infrastructure Engineer

General Intuition & Medal • New York (NY)

On-site
USD 120,000 - 160,000
Competitive salary and meaningful equity
Comprehensive medical, dental, and vision coverage
401(k)
+5