Senior Site Reliability Engineer (DevOps)

Dover

Northern (KY)

Hybrid

USD 120,000 - 160,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Kintsugi is seeking a Senior Site Reliability Engineer (DevOps) to scale and harden production infrastructure across managed Kubernetes and AWS. You’ll own reliability, reduce toil, and build tooling that lets engineers move fast without sacrificing stability.

You will collaborate with Platform Engineering, Product, and QA to design resilient architectures, improve deployment pipelines, and extend internal tools that support engineering velocity while preserving security and observability.

Qualifications

  • 5-8 years in Site Reliability Engineering, DevOps, or Infrastructure Engineering roles with ownership of a production system at meaningful scale
  • Fluent agent-first daily work — directing coding agents to do real engineering work
  • Strong foundation in AWS-hosted data & networking (RDS/Postgres, ElastiCache/Redis, VPC) and managed Kubernetes
  • Track record of building tools that remove manual work (scripts, services, internal platforms)
  • Hands-on experience with CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation)
  • Observability expertise (metrics, tracing, logging) and modern monitoring practices
  • Familiarity with cloud security/compliance (SOC 2, GDPR a plus)
  • Collaborative mindset to empower developers to move fast safely
  • Experience with developer enablement — internal tooling and platforms for productivity
  • Nice to have: multi-cloud operations across providers

Responsibilities

  • Own the reliability of production infrastructure on Kubernetes and AWS
  • Work agent-first day to day: build, debug, and automate with agentic coding workflows
  • Build internal tools and automation to eliminate recurring toil
  • Develop and operate monitoring, alerting, and observability across the stack
  • Partner with engineering teams to design for reliability and performance from the start
  • Automate infrastructure management via infrastructure-as-code and improve CI/CD
  • Lead and evolve incident response, including postmortems and blameless learning
  • Optimize infrastructure for cost efficiency while maintaining high availability and security
  • Contribute to security, compliance, and disaster recovery efforts
  • Support developer enablement with in-house tooling and internal platforms

Skills

Agent-first mindset
AWS & Kubernetes
CI/CD pipelines
Observability
Infrastructure as code
Security/compliance awareness

Tools

Terraform
CloudFormation
Kubernetes

Job description

About Kintsugi

Kintsugi is revolutionizing sales tax automation with our AI-powered platform designed specifically for e-commerce and SaaS businesses. Our solution reduces tax preparation time by 75 percent and cuts compliance costs by 50 percent, allowing finance teams to focus on strategic initiatives rather than routine calculations. As we continue to grow and disrupt the tax automation space, we're building a world-class team to help us scale with purpose.

The Role

We're looking for a Senior Site Reliability Engineer (DevOps) to help scale and harden the infrastructure that powers Kintsugi. This role sits at the intersection of software engineering and operations: you'll keep production reliable under real load, and you'll build the tooling and automation that reduce the manual work of doing that, rather than absorbing more of it yourself as we grow.

You'll work closely with Platform Engineering, Product, and QA to design resilient architectures, improve deployment pipelines, and build the internal tools and guardrails that let the team move quickly without sacrificing stability. You'll operate across our managed Kubernetes and AWS-hosted data layer (Postgres, Redis, networking), and you'll be a key force in shaping the reliability and developer-experience foundation of our engineering org.

We're an agentic-coding-first team — coding agents already do real engineering and operations work here, not just autocomplete on the side. We expect this role to build and extend that practice, not just adopt it.

What You'll Do
  • Own the reliability of production infrastructure running on managed Kubernetes and AWS, keeping a high-traffic system up and catching issues before customers do

  • Work agent-first day to day: build, debug, and automate using agentic coding workflows as your default mode, not a fallback tool

  • Build internal tools and automation that eliminate recurring manual work (toil) for the team, rather than just documenting runbooks around it

  • Develop and operate monitoring, alerting, and observability systems (metrics, tracing, logging) across the stack

  • Partner with engineering teams to design for reliability and performance from the start, not bolt it on after incidents

  • Automate infrastructure management through infrastructure-as-code, and improve CI/CD pipelines and local developer workflows

  • Lead and evolve incident response practices, including postmortems and blameless learning

  • Optimize infrastructure for cost efficiency while maintaining high availability and security standards

  • Contribute to security, compliance, and disaster recovery efforts as the platform scales

  • Support developer enablement: build and improve the in-house tooling, local dev workflows, and internal platforms other engineers rely on

What We’re Looking For
  • 5-8 years in Site Reliability Engineering, DevOps, or Infrastructure Engineering roles, with real ownership of a production system at meaningful scale — someone who drives reliability and tooling initiatives rather than waiting to be assigned them

  • Fluent working agent-first day to day — directing coding agents to do real engineering work, not just occasional autocomplete

  • Strong foundation in AWS-hosted data and networking services (RDS/Postgres, ElastiCache/Redis, VPC/networking) and experience running workloads on managed Kubernetes

  • A track record of building tools, not just running playbooks: scripts, services, or internal platforms that removed manual work for a team

  • Hands-on experience with CI/CD pipelines and infrastructure-as-code (e.g., Terraform, CloudFormation)

  • Expertise in observability stacks (metrics, tracing, logging) and modern monitoring practices

  • Familiarity with security and compliance in cloud environments (SOC 2, GDPR, etc. a plus)

  • A collaborative mindset with a passion for empowering developers to move fast safely

  • Experience with (or strong interest in) developer enablement — building the internal tooling and platforms that make other engineers more productive

  • Nice to have: experience operating across multiple cloud providers, as our infrastructure footprint expands

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE & DevOps - Agent-First Reliability Engineer
Senior SRE & DevOps - Agent-First Reliability Engineer

Dover • Northern (KY)

Hybrid
USD 120,000 - 160,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

United States Digital Space LLC • Michigan

On-site
USD 150,000 - 190,000
Senior Site Reliability Engineer, Forward Deployed - Remote USA ONLY
Senior Site Reliability Engineer, Forward Deployed - Remote USA ONLY

Ardan Labs • United States

On-site
USD 140,000 - 210,000
Senior Engineer, Platform & Site Reliability
Senior Engineer, Platform & Site Reliability

ICE Clear Europe Limited • Jacksonville (FL)

On-site
USD 120,000 - 170,000
Senior Engineer, Platform & Site Reliability
Senior Engineer, Platform & Site Reliability

ICE • Jacksonville (FL)

On-site
USD 140,000 - 190,000
Senior Engineer, Platform & Site Reliability
Senior Engineer, Platform & Site Reliability

Intercontinental Exchange Holdings, Inc. • Jacksonville (FL)

On-site
USD 140,000 - 180,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Specter • San Francisco (CA)

On-site
USD 180,000 - 230,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • United States

On-site
USD 140,000 - 190,000