Lead Infrastructure Engineer

RapidFort, Inc.

United States

Hybrid

USD 180,000 - 225,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision insurance
401(k) retirement plan
Paid time off
Company-provided equipment
Professional development

Job summary

RapidFort, Inc. is seeking a Lead Infrastructure Engineer in the Bay Area with hybrid work.

The role emphasizes hands-on Linux systems engineering (about 70%) and data center hardware (about 30%), including patching, automation, and a deputy leadership stance when the manager is unavailable. You will own the golden image pipeline, write Bash/Python automation, manage Ansible content, and lead incident response.

Qualifications

  • 7+ years engineering and administering Linux in production environments.
  • Strong Bash and Python — you write production automation, not just modify it.
  • Ansible at fleet scale, with opinions on idempotency and inventory structure.
  • Demonstrated Linux performance troubleshooting depth.
  • Hands-on data center experience: server lifecycle, firmware, out-of-band management.
  • Enterprise storage: SAN/NAS, multipathing, capacity management.
  • Hypervisor cluster experience (VMware/Hyper-V/Proxmox).
  • Solid networking fundamentals and ability to read a packet capture.
  • Production on-call experience under incident and change processes.
  • Sound escalation judgment: know which calls to make.

Responsibilities

  • Design, build, and harden Linux systems and golden images.
  • Write production automation in Bash and Python for provisioning and remediation.
  • Own configuration management content in Ansible: roles, playbooks, inventory.
  • Own OS patching cadence, testing, rollout, and end-of-life upgrades.
  • Troubleshoot performance issues at kernel, I/O, memory, and filesystem levels.
  • Reduce manual work by building tooling and automation.

Skills

Bash
Python
Linux troubleshooting
On-call support
Escalation judgment

Tools

Ansible
VMware
Hyper-V
Proxmox
Out-of-band mgmt

Job description

Title Lead Infrastructure Engineer

Department Information Technology — Infrastructure Engineering

Reports To Infrastructure Engineering Manager

Location SFO / Bay area- Hybrid

Employment Type Full-time

Travel periodic data center and colocation site visits

On-Call Participates in a shared infrastructure on-call rotation

On-site requirement: this role has hands-on data center responsibility and cannot be performed fully remote.

About the Role

RapidFort Inc. runs its core systems on infrastructure we own and operate. We are looking for a Lead Infrastructure Engineer to be the deepest technical resource on that estate — and the person our Infrastructure Engineering Manager can hand the keys to.

This is a hands-on senior engineering role with a deliberate weighting: roughly 70% Linux systems engineering and automation, 30% data center hardware, storage, and network. You will spend most of your week making the Linux environment more consistent and less manual, and a meaningful part of it with your hands on physical infrastructure.

It is also a deputy role. When the Infrastructure Engineering Manager is unavailable, you carry technical decision authority, change approval, and incident escalation. You will not manage people or budgets — this is technical leadership, not a management track — but you will be the reason the team does not stall when the manager is out.

We are being specific about the split because the combination is uncommon. If you are strong on Linux automation but have never racked a server or traced a bad uplink, this is probably not the right fit. If you are a strong data center engineer who reaches for a script before a runbook, it likely is.

What You’ll Do
Linux systems and automation (the majority of your time)
  • Design, build, and harden Linux systems (RHEL/Rocky and Ubuntu) against defined baselines, and own the golden image pipeline that produces them.
  • Write production automation in Bash and Python — provisioning, remediation, reporting, and evidence collection.
  • Own our configuration management content in Ansible: roles, playbooks, inventory structure, and the standards they encode.
  • Own patching and OS lifecycle: cadence, testing, staged rollout, exception tracking, and end-of-life upgrade planning.
  • Root-cause performance problems — kernel, I/O, memory, filesystem, systemd — without waiting on a vendor.
  • Replace repeatable manual work with tooling, and be able to show the hours you removed.
Data center, hardware, and network
  • Execute the server hardware lifecycle: staging, racking, cabling, firmware and BIOS, out-of-band management, and secure decommissioning.
  • Diagnose hardware faults and drive vendor RMA and support cases to resolution.
  • Configure and maintain enterprise storage: SAN/NAS presentation, multipathing, replication, and capacity management.
  • Maintain the hypervisor estate — host builds, cluster membership, patch currency, datastore and resource pool configuration.
  • Configure and troubleshoot data center networking: VLANs, trunking, link aggregation, routing, firewalls, load balancers, and VPN.
  • Troubleshoot at the packet level and keep IPAM, topology, and asset records accurate.
Operations and technical leadership
  • Participate in the infrastructure on-call rotation and serve as the escalation tier for complex incidents.
  • Act as second reviewer on infrastructure designs before they are built.
  • Mentor other engineers and raise the team’s automation and documentation standards.
  • Deputize for the Infrastructure Engineering Manager: change approval, intake triage, maintenance window approval, and incident escalation during absence.
What You’ll Bring
Required
  • 7+ years engineering and administering Linux in production environments.
  • Strong Bash and Python — you write production automation, not just modify it.
  • Ansible (or equivalent) at fleet scale, with real opinions about idempotency and inventory structure.
  • Demonstrated Linux performance troubleshooting depth.
  • Hands‑on data center experience: server lifecycle, firmware, out-of-band management, racking and structured cabling.
  • Enterprise storage: SAN/NAS presentation, multipathing, capacity management.
  • Hypervisor platform experience at cluster scale (VMware, Hyper‑V, Proxmox, or equivalent).
  • Solid networking fundamentals — VLANs, routing, link aggregation, DNS/DHCP, firewall rule design — and the ability to read a packet capture.
  • Production on-call experience under a formal incident and change process.
  • Sound escalation judgment: you know which calls to make and which to hold.
Nice to Have
  • RHCE or equivalent Linux certification; CCNA or equivalent networking certification; a virtualization platform certification.
  • Prior experience acting as technical lead or deputy for a manager.
  • Experience in an environment operating under ISO 27001 or SOC 2 Type 2.
  • Colocation management, data center migration, or consolidation experience.
  • Enough cloud literacy to hand work off cleanly to our Cloud Operations team — depth is not required.
What Success Looks Like

By 90 days — you know the estate well enough to run an incident, and you have identified the three most expensive manual processes on the team.

By six months — at least one of those processes is automated and measurably gone, you are a routine design reviewer, and you have covered a manager absence without escalation.

By twelve months — the manager can be out for two consecutive weeks without delivery or escalation impact. That is the real test of this role.

Why This Role
  • Genuine ownership. We run our own infrastructure. Decisions about how it is built are made here, not by a provider.
  • Real technical authority. The deputy delegation is written down, not implied — you will know exactly which decisions are yours.
  • Automation is the mandate, not a side project. Reducing manual work is how this role is measured.
Compensation & Benefits
  • Base Salary: $180,000–$225,000 USD annually, depending on experience, qualifications, and location
  • Annual performance-based bonus
  • Equity: Stock options in RapidFort
  • Medical, dental, and vision insurance
  • 401(k) retirement plan
  • Paid time off and company holidays
  • Paid sick leave
  • Company-provided equipment
  • Professional development and growth opportunities

RapidFort is an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, national origin, age, disability, veteran status, or any other characteristic protected by applicable law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Infrastructure Engineer
Infrastructure Engineer

RAPIDFORT • United States

On-site
USD 100,000 - 125,000
Stock options
Medical, dental, and vision insurance
401(k) retirement plan
+4
Staff Software Engineer – Build Platform
Staff Software Engineer – Build Platform

RapidFort, Inc. • United States

Hybrid
USD 150,000 - 200,000
Annual performance bonus
Equity participation
Medical, dental, vision coverage
+5
Senior IT Systems and Infrastructure Engineer
Senior IT Systems and Infrastructure Engineer

NightDragon Acquisition Corp. • Sterling (VA)

On-site
USD 150,000 - 175,000
Health benefits
PTO 20 days
Parental leave 7 weeks
+2
Senior IT Systems and Infrastructure Engineer
Senior IT Systems and Infrastructure Engineer

Forterra • Arlington (VA)

On-site
USD 150,000 - 175,000
Healthcare benefits
PTO 20 days
Parental leave 7 weeks
+3
Senior/ Staff IT Systems Engineer
Senior/ Staff IT Systems Engineer

Forterra • Arlington (VA)

On-site
USD 150,000 - 175,000
Premium Healthcare Benefits
PTO 20 days (annual)
7 weeks parental leave
+2
Staff IT Infrastructure Engineer
Staff IT Infrastructure Engineer

Forterra • Clarksburg (MA)

On-site
USD 150,000 - 175,000
Premium Healthcare Benefits
Generous PTO
Parental leave
Senior/ Staff IT Systems Engineer
Senior/ Staff IT Systems Engineer

Koitecc Solutions • Arlington (VA), Northern (KY)

On-site
USD 150,000 - 175,000
Premium Healthcare Benefits
20 days PTO per year
7 weeks parental leave
+2
Senior/Staff Infrastructure & Platform Engineer (Bay Area)
Senior/Staff Infrastructure & Platform Engineer (Bay Area)

Cerebras • Santa Clara (CA)

On-site
USD 155,000 - 230,000
Senior/Staff Infrastructure & Platform Engineer (Bay Area)
Senior/Staff Infrastructure & Platform Engineer (Bay Area)

Fortanix • Santa Clara (CA), Northern (KY)

On-site
USD 155,000 - 230,000
Unlimited PTO
Volunteer Time Off (40 hours/year)
Internet stipend
+1
Lead Infra Engineer: Linux Automation & Data Center
Lead Infra Engineer: Linux Automation & Data Center

RapidFort, Inc. • United States

Hybrid
USD 180,000 - 225,000
Medical, dental, vision insurance
401(k) retirement plan
Paid time off
+2