Site Reliability Engineer

PVH (Tommy Hilfiger/Calvin Klein)

Alpharetta (GA)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Incident IQ in Atlanta is seeking a Site Reliability Engineer to build from scratch and define reliability for our production systems. You will work with cutting-edge observability tooling and shape how the engineering org ships with confidence.

We value data-driven, collaborative problem solving, and you taking total ownership while staying pragmatic. Expect startup-speed execution and deep dives to resolve complex reliability challenges.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent.

Responsibilities

  • SLI/SLO definition and Grafana dashboards to drive reliability.
  • Own incident management practices including on-call readiness.
  • Own the observability stack across metrics, logs, traces, and RUM.
  • Coach teams on SRE practices and improve service reliability.
  • Automate toil away with infrastructure as code.
  • Design and run load tests and chaos engineering days.

Education

Bachelor's degree in Computer Science, Computer Engineering, or equivalent

Job description

Company Overview:

About Us: Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12 districts. Trusted by over 2,000 districts, Incident IQ powers mission-critical services for more than 12 million students and educators nationwide. By connecting technology and operational workflows, Incident IQ enables schools to streamline processes, reduce administrative burdens, and focus on what matters most: supporting students.

Purpose: Incident IQ is committed to creating a future where every K-12 district operates with seamless efficiency. When operations are unified on a single platform, districts gain the clarity and control needed to build a stronger foundation for student success. We're focused on delivering the tools, support, and partnerships that help make that vision a reality.

Mission: Incident IQ is on a mission to eliminate the friction of disconnected systems and clunky workflows that slow schools down. We're reimagining the critical work that happens behind the scenes, bringing visibility, efficiency, and impact to the processes that keep classrooms running. By streamlining the complex, automating the routine, and surfacing the insights that matter most, we can create the conditions for educators to teach, students to thrive, and districts to shape the future of education.

Site Reliability Engineer (SRE) Overview:

We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-from-zero role at startup speed. You're our first dedicated Site Reliability Engineer, and you'll be defining what "reliable" means for our production systems, not maintaining someone else's playbook. You'll work with leading-edge observability and reliability tooling, and the calls you make will directly shape how confidently the whole engineering org ships.

Expect real engineering deep dives, not top-down mandates. We love digging into a hard problem together, and we want you to bring a strong point of view, back it up with data and sound reasoning, and enjoy the back-and-forth as we work toward the best answer. Good persuasion skills matter here as much as technical depth, since good ideas still have to win the room. We move at startup speed: we'd rather figure something out in a few hours than plan it for weeks. We're a collaborative, respectful team: we debate ideas hard, never people.

We care much more about a proven track record running big, ambiguous projects efficiently than about years of tenure or a wall of certifications. You should be genuinely comfortable working independently: we won't hand-hold you or chase you for status updates. We expect you to take total ownership of outcomes and drive them without being asked twice, and without running your own separate agenda. This work is relentless, juggling several things at once under real time pressure is normal here, and the right candidate is passionate about SRE and thrives on that intensity, not just tolerates it.

Site Reliability Engineer (SRE) Responsibilities:
  • SLI/SLO Definition & Grafana Implementation: Drive the definition of Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for our core services, translating them into insightful Grafana dashboards and actionable, burn-rate-based alerting, so pages are precise and noise stays low.
  • Incident Management: Stand up our incident management practice (tooling such as PagerDuty, on-call training, incident command), then own and continuously improve it, stepping in personally only for the most severe incidents.
  • Observability Stack Ownership: Own the observability stack end to end: metrics, logs, traces, Real User Monitoring (RUM), and synthetic checks across the user journey, alerting whenever a signal deviates from baseline.
  • Team Enablement: Partner with engineering teams to refine SLIs, SLOs, and error budgets as services evolve, and coach teams on SRE and observability best practices.
  • Toil Reduction: Identify and automate away manual, repetitive operational work through infrastructure as code and tooling.
  • Chaos & Performance Engineering: Design and run load/performance tests and chaos engineering game days to proactively surface weaknesses before they cause incidents.

For example: in your first few days, you might stand up an SLO and a burn-rate alert in Grafana for our highest-traffic service. Within a couple of weeks, PagerDuty on-call is configured and the rotation is trained on incident command. That's the pace we operate at here: hours and days, not weeks.

Site Reliability Engineer (SRE) Requirements:
  • Education & Systems Foundations: Bachelor's degree in Computer Science, Computer Engineering, or equivalent f
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Incident IQ • Atlanta (GA)

On-site
USD 120,000 - 190,000
Medical benefits
Dental benefits
Vision benefits
+3
Site Reliability Engineer
Site Reliability Engineer

TrulyHired • Atlanta (GA)

On-site
USD 140,000 - 190,000
Medical insurance
401k match
PTO
Site Reliability Engineer
Site Reliability Engineer

Incident IQ • Alpharetta (GA)

On-site
USD 120,000 - 180,000
Medical, Dental, Vision
401k Match
PTO
+1
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Founding SRE — Build Reliable, Scalable Systems
Founding SRE — Build Reliable, Scalable Systems

Incident IQ • Atlanta (GA)

On-site
USD 120,000 - 190,000
Medical benefits
Dental benefits
Vision benefits
+3
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Methodic • San Francisco (CA)

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

R&D • United States

On-site
USD 90,000 - 140,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Virtual Tech Gurus • Puerto Rico

On-site
USD 140,000 - 210,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

OutSolve • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work environment
Competitive compensation
Professional development opportunities
+1