Site Reliability Engineer

exiger

Jersey City (NJ)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Discretionary time off
16 weeks parental leave
Hybrid work model

Job summary

Exiger is seeking a Site Reliability Engineer in the United States (Hybrid) to stand up the SRE function for its 1Exiger platform. You will set reliability standards, tooling, and playbooks to keep the platform reliable for 550+ customers, including Fortune 500 firms and government agencies.

You will own reliability across the full service lifecycle—from design and capacity planning to deployment, monitoring, and incident response—plus build automation to scale without expanding headcount.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, a related field, or equivalent practical experience.
  • 6 years of experience in software or systems engineering, including at least 4 years in a dedicated Site Reliability Engineering, production engineering, or platform reliability role.
  • 4 years of experience designing, analyzing, and troubleshooting large-scale distributed systems.
  • Strong grounding in Unix/Linux internals and networking fundamentals.
  • Hands-on experience establishing core SRE practices from the ground up: SLIs, SLOs, monitoring, capacity planning, and automation.
  • Experience with chaos engineering or fault-injection testing to validate system resilience.
  • Proven incident management experience: on-call ownership, leading response under pressure, and driving blameless postmortems.
  • Experience in troubleshooting and supporting applications like web services, data storage, databases, and data pipelines, with Linux/Unix or other operating systems.
  • Familiarity with cloud platforms (AWS) and secure system integration.
  • Comfort integrating AI coding assistants into daily engineering workflow.
  • Ability to translate ambiguous mission problems into structured technical solutions.
  • Ability to operate independently in dynamic, high-stakes environments.
  • Willingness to travel as needed to support customer engagements.

Responsibilities

  • Establish the SRE function: define SLIs, SLOs, and error budgets, and set reliability standards adopted by other teams.
  • Build and own observability: instrument services for availability, latency, and health; convert signals into actionable insight.
  • Drive decisions with data: form hypotheses, measure impact of changes, and rely on metrics to set reliability priorities.
  • Own the reliability of production services from design through steady-state operation.
  • Eliminate repetitive manual operations through automation and infrastructure as code.
  • Plan for scale: capacity planning, performance analysis, and changes that improve reliability and delivery velocity.
  • Improve resilience through chaos engineering and fault-injection testing to prove graceful degradation.
  • Lead blameless incident response and on-call rotation; lead postmortems to root cause.
  • Leverage AI-assisted development tooling to accelerate automation and investigation work.

Skills

Site Reliability Engineering
Distributed systems
Unix/Linux internals
Networking fundamentals
Monitoring & observability
Automation
Incident management
Go or C programming
AWS / cloud platforms
AI coding assistants integration

Education

Bachelor’s or Master’s degree in Computer Science

Tools

Chaos Monkey
Gremlin
LitmusChaos
Claude
Codex

Job description

Who We Are

Exiger transforms supply chains into a strategic advantage, advancing our mission to make the world a safer and more transparent place to succeed. Our AI platform, 1Exiger, delivers instant visibility into complex supplier ecosystems, leveraging proprietary data and advanced AI to surface risk, automate compliance, and unlock efficiencies and cost savings to strengthen long‑term resilience. Trusted by 550+ global customers, including Fortune 500 companies and U.S. government agencies, Exiger is a recognized, award‑winning leader in supply chain AI and a FedRAMP authorized provider to the federal government.

Site Reliability Engineer

Location: U.S. (Hybrid)

This role requires U.S. citizenship and eligibility for a U.S. security clearance.

Role Summary

Exiger is transforming how governments and global enterprises manage supply chain, defense, and geopolitical risk. Our AI‑powered platform equips the world’s most important institutions with the intelligence they need to protect critical infrastructure, secure national interests, and make data‑driven operational decisions.

From identifying counterfeit parts in defense supply chains to anticipating geopolitical risk exposure, Exiger enables mission owners to act with clarity and confidence in complex, high‑stakes environments.

This is our first dedicated Site Reliability Engineering hire and a founding role. You will help stand up the SRE function at Exiger: setting the standards, tooling, and practices that keep 1Exiger reliable for our 550+ customers, including Fortune 500 companies and U.S. government agencies. You will own reliability across the full service lifecycle, from design and capacity planning through deployment, monitoring, and incident response, and build the automation that lets the platform scale without scaling headcount. Because you are first, we need someone who has practiced SRE before and can bring the playbook, not learn it on the job.

You will use your expertise in coding, algorithms, complexity analysis, and large‑scale distributed system design to solve the reliability challenges that are unique to operating a mission‑critical AI platform in regulated and government environments.

SRE’s culture of intellectual curiosity, problem solving and openness is key to its success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame‑free environment. We promote self‑direction to work on meaningful projects, while we also strive to create an environment that provides the support and mentorship needed to learn and grow.

What You’ll Do
  • Establish the SRE function: define SLIs, SLOs, and error budgets, and set reliability standards that other engineering teams adopt.
  • Build and own observability: instrument services for availability, latency, and system health, and turn that signal into actionable insight.
  • Drive decisions with data: form hypotheses, measure the impact of every change, and let metrics rather than intuition set reliability priorities.
  • Own the reliability of production services from design consulting and launch reviews through steady‑state operation.
  • Eliminate repetitive manual operations through automation and infrastructure as code, replacing them with reliable, self‑service tooling.
  • Plan for scale: capacity planning, performance analysis, and driving changes that improve both reliability and delivery velocity.
  • Improve resilience through chaos engineering and fault‑injection testing, running game days that prove the platform degrades gracefully and recovers from failure.
  • Lead sustainable, blameless incident response and postmortems, and stand up and participate in an on‑call rotation.
  • Leverage AI‑assisted development tooling (such as Codex and Claude) to accelerate automation, tooling, and investigation work, and help the team adopt these tools effectively.
What You Need
  • Bachelor’s or Master’s degree in Computer Science, a related field, or equivalent practical experience.
  • 6 years of experience in software or systems engineering, including at least 4 years in a dedicated Site Reliability Engineering, production engineering, or platform reliability role. As our first SRE hire, you must have practiced SRE before and be ready to establish the function.
  • 4 years of experience designing, analyzing, and troubleshooting large‑scale distributed systems.
  • Strong grounding in Unix/Linux internals (filesystems, processes, system calls) and networking fundamentals (TCP/IP, DNS, routing, load balancing).
  • Hands‑on experience establishing core SRE practices from the ground up: SLIs, SLOs, and error budgets, monitoring and observability, capacity planning, and automation that removes repetitive manual work.
  • A rigorous, empirical mindset: you form hypotheses, measure outcomes, and make metrics‑driven decisions rather than relying on intuition or anecdote.
  • Experience with chaos engineering or fault‑injection testing (for example game days, Chaos Monkey, Gremlin, or LitmusChaos) to validate system resilience.
  • Proven incident management experience: on‑call ownership, leading response under pressure, and driving blameless postmortems to root cause.
  • Experience in troubleshooting and supporting applications like web services, data storage, databases, and data pipelines, with Linux/Unix or other operating systems.
  • Familiarity with cloud platforms (AWS) and secure system integration.
  • Comfort integrating AI coding assistants (such as Claude and Codex) into your daily engineering workflow.
  • Ability to translate ambiguous mission problems into structured technical solutions.
  • Ability to operate independently in dynamic, high‑stakes environments.
  • Willingness to travel as needed to support customer engagements.
Nice to Have
  • 4 years of experience programming in Go or C (Java also welcome), with the ability to debug, optimize, and automate rather than just script.
  • Experience supporting ML or data platforms in production.
  • Familiarity with data warehouses such as Snowflake, Redshift, and/or Apache Iceberg.
  • Experience operating in FedRAMP or other regulated or government environments.
Why You’ll Love Working at Exiger
  • High‑performance culture rooted in accountability, collaboration, and a shared commitment to excellence.
  • Discretionary Time Off for all employees, with no maximum limits on time off.
  • Industry‑leading health, vision, and dental benefits.
  • Competitive compensation package.
  • 16 weeks of fully paid parental leave.
  • Flexible, hybrid approach to working from home and in the office where applicable.
  • Focus on wellness and employee health through stipends and dedicated wellness programming.
  • Purposeful career development programs with reimbursement provided for educational certifications.

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability or protected veteran status, or any other legally protected basis, in accordance with applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Exiger • McLean (VA)

Hybrid
USD 150,000 - 210,000
Discretionary Time Off
Health, vision, dental benefits
16 weeks parental leave
+1
Founding SRE Engineer — Hybrid, Scale an AI Platform
Founding SRE Engineer — Hybrid, Scale an AI Platform

Exiger • McLean (VA)

Hybrid
USD 150,000 - 210,000
Discretionary Time Off
Health, vision, dental benefits
16 weeks parental leave
+1
Tech Lead/Principal Engineer, Public Sector
Tech Lead/Principal Engineer, Public Sector

Exiger • McLean (VA)

Hybrid
USD 120,000 - 180,000
Discretionary Time Off
Comprehensive benefits
Senior Forward Deployed Engineer
Senior Forward Deployed Engineer

Exiger • McLean (VA)

On-site
USD 180,000 - 230,000
Hybrid work model
Discretionary Time Off
Health, vision, and dental benefits
+1
Platform Engineer
Platform Engineer

Exiger • Jersey City (NJ)

Hybrid
USD 85,000 - 110,000
Discretionary Time Off
Industry-leading health, vision, and dental benefits
Competitive compensation package
+3
Forward Deployed Engineer South Ogden, Utah, United States
Forward Deployed Engineer South Ogden, Utah, United States

Exiger • Ogden (UT)

Hybrid
USD 90,000 - 130,000
Discretionary Time Off
Industry-leading health benefits
16 weeks fully paid parental leave
+1
Government Pre-Sales Engineer
Government Pre-Sales Engineer

Exiger • McLean (VA)

On-site
USD 100,000 - 140,000
Discretionary Time Off
Industry-leading health benefits
16 weeks of fully paid parental leave
+2
Platform Engineer
Platform Engineer

Exiger • Richmond (VA)

Hybrid
USD 100,000 - 110,000
Discretionary Time Off
Industry leading health benefits
Competitive compensation package
+3
Platform Engineer
Platform Engineer

Exiger • McLean (VA)

Hybrid
USD 100,000 - 110,000
Discretionary time off
Health, vision, and dental benefits
Parental leave (16 weeks)
Forward Deployed Engineer
Forward Deployed Engineer

Exiger • South Ogden (UT)

Hybrid
USD 100,000 - 130,000
Discretionary Time Off
Health, vision, and dental benefits
16 weeks of fully paid parental leave
+1