Founding Site Reliability Engineer for AI Platform

exiger

Jersey City (NJ)

Hybrid

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Discretionary time off
16 weeks parental leave
Hybrid work model

Job summary

Exiger is seeking a Site Reliability Engineer in the United States (Hybrid) to stand up the SRE function for its 1Exiger platform. You will set reliability standards, tooling, and playbooks to keep the platform reliable for 550+ customers, including Fortune 500 firms and government agencies.

You will own reliability across the full service lifecycle—from design and capacity planning to deployment, monitoring, and incident response—plus build automation to scale without expanding headcount.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, a related field, or equivalent practical experience.
  • 6 years of experience in software or systems engineering, including at least 4 years in a dedicated Site Reliability Engineering, production engineering, or platform reliability role.
  • 4 years of experience designing, analyzing, and troubleshooting large-scale distributed systems.
  • Strong grounding in Unix/Linux internals and networking fundamentals.
  • Hands-on experience establishing core SRE practices from the ground up: SLIs, SLOs, monitoring, capacity planning, and automation.
  • Experience with chaos engineering or fault-injection testing to validate system resilience.
  • Proven incident management experience: on-call ownership, leading response under pressure, and driving blameless postmortems.
  • Experience in troubleshooting and supporting applications like web services, data storage, databases, and data pipelines, with Linux/Unix or other operating systems.
  • Familiarity with cloud platforms (AWS) and secure system integration.
  • Comfort integrating AI coding assistants into daily engineering workflow.
  • Ability to translate ambiguous mission problems into structured technical solutions.
  • Ability to operate independently in dynamic, high-stakes environments.
  • Willingness to travel as needed to support customer engagements.

Responsibilities

  • Establish the SRE function: define SLIs, SLOs, and error budgets, and set reliability standards adopted by other teams.
  • Build and own observability: instrument services for availability, latency, and health; convert signals into actionable insight.
  • Drive decisions with data: form hypotheses, measure impact of changes, and rely on metrics to set reliability priorities.
  • Own the reliability of production services from design through steady-state operation.
  • Eliminate repetitive manual operations through automation and infrastructure as code.
  • Plan for scale: capacity planning, performance analysis, and changes that improve reliability and delivery velocity.
  • Improve resilience through chaos engineering and fault-injection testing to prove graceful degradation.
  • Lead blameless incident response and on-call rotation; lead postmortems to root cause.
  • Leverage AI-assisted development tooling to accelerate automation and investigation work.

Skills

Site Reliability Engineering
Distributed systems
Unix/Linux internals
Networking fundamentals
Monitoring & observability
Automation
Incident management
Go or C programming
AWS / cloud platforms
AI coding assistants integration

Education

Bachelor’s or Master’s degree in Computer Science

Tools

Chaos Monkey
Gremlin
LitmusChaos
Claude
Codex

Job description

Exiger is seeking a Site Reliability Engineer in the United States (Hybrid) to stand up the SRE function for its 1Exiger platform. You will set reliability standards, tooling, and playbooks to keep the platform reliable for 550+ customers, including Fortune 500 firms and government agencies.

You will own reliability across the full service lifecycle—from design and capacity planning to deployment, monitoring, and incident response—plus build automation to scale without expanding headcount.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

exiger • Jersey City (NJ)

Hybrid
USD 150,000 - 210,000
Discretionary time off
16 weeks parental leave
Hybrid work model
Site Reliability Engineer
Site Reliability Engineer

Exiger • McLean (VA)

Hybrid
USD 150,000 - 210,000
Discretionary Time Off
Health, vision, dental benefits
16 weeks parental leave
+1
Founding SRE Engineer — Hybrid, Scale an AI Platform
Founding SRE Engineer — Hybrid, Scale an AI Platform

Exiger • McLean (VA)

Hybrid
USD 150,000 - 210,000
Discretionary Time Off
Health, vision, dental benefits
16 weeks parental leave
+1
Staff SRE: AI-Driven Reliability & Platform Leader
Staff SRE: AI-Driven Reliability & Platform Leader

WEX, Inc. • San Francisco (CA)

On-site
USD 121,000 - 151,000
Staff SRE: AI‑Driven Reliability Leader
Staff SRE: AI‑Driven Reliability Leader

WEX Inc. • United States

On-site
USD 121,000 - 151,000
Health insurance
Retirement savings plan
Paid time off
+1
Remote Senior SRE: Cloud Reliability, CI/CD & AI-Driven
Remote Senior SRE: Cloud Reliability, CI/CD & AI-Driven

IDEXX • Boston (MA)

On-site
USD 100,000 - 125,000
Health benefits
401k matching
Pet Insurance
+1
Senior SRE - Cloud Reliability & AI-Enhanced Ops
Senior SRE - Cloud Reliability & AI-Enhanced Ops

IDEXX • Concord (NH)

Hybrid
USD 100,000 - 125,000
Health/Dental/Vision benefits
5% 401(k) matching
Annual cash bonus
Founding SRE — Build Reliable, Scalable Systems
Founding SRE — Build Reliable, Scalable Systems

Incident IQ • Atlanta (GA)

On-site
USD 120,000 - 190,000
Medical benefits
Dental benefits
Vision benefits
+3
Senior SRE: Platform Reliability & AI-Driven Ops
Senior SRE: Platform Reliability & AI-Driven Ops

Block • New York (NY)

On-site
USD 170,000 - 284,000
Healthcare coverage
Retirement plans
Employee Stock Purchase Program
+1
Staff SRE: Reliability Architect for AI-Driven Platform
Staff SRE: Reliability Architect for AI-Driven Platform

Slope • Costa Mesa (CA)

On-site
USD 191,000 - 253,000
Benefits package