Lead Site Reliability Engineer

CardWorks Servicing LLC

Pittsburgh (Allegheny County)

On-site

USD 146,032 - 162,257

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive base pay
Medical, Dental and Vision coverage
401(k) Plan with Company Match
Paid Vacation and Sick Days

Job summary

CardWorks Servicing LLC seeks a skilled Site Reliability Engineer to establish and oversee an SRE operating model. This role involves implementing AI solutions, improving service reliability, and collaborating across teams to foster a strong culture of reliability.

The ideal candidate will have extensive experience in SRE and infrastructure as code, along with a Master's degree in a relevant field. Compensation for this position ranges from $146,032 to $162,257, with hybrid or remote options available.

Qualifications

  • 7+ years of experience in Site Reliability Engineering.
  • Experience applying AI/ML to operations with measurable outcomes.
  • Strong observability and telemetry expertise.

Responsibilities

  • Establish the SRE operating model and ensure adoption across teams.
  • Own the centralized SRE service engagement model.
  • Design automation and observability solutions to reduce operational toil.

Skills

Site Reliability Engineering
AI/ML Operations
Observability and Telemetry
Infrastructure as Code
CI/CD Pipeline Design
Containerization

Education

Master’s degree in Computer Science or Engineering

Tools

Terraform
Ansible
Docker
Kubernetes
Azure DevOps
GitHub Actions

Job description

Become an everyday champion — and build a career where your impact fuels financial progress. CardWorks Financial Group is a diversified financial services platform building ethical solutions across credit, lending, and the full customer lifecycle. Through our family of companies, CardWorks Financial Group tackles the complex challenges that larger financial institutions leave behind. We’re embedded throughout the credit card ecosystem as a lender, servicer, and merchant acquirer.

We operate as a bank (Merrick Bank), a servicing provider (CardWorks Servicing), and a debt recovery provider (Carson Smithfield). With nearly 40 years of operating history, we are a top‑three non‑prime focused general‑purpose card issuer and a top‑fifteen U.S. merchant acquirer.

Job Responsibilities
  • Establish the SRE operating model, including service onboarding, engagement model, governance, reliability reviews, production readiness standards, and quarterly planning, and ensure adoption across teams.
  • Identify, pilot, and operationalize AI‑enabled reliability use cases (e.g., alert noise reduction, incident summarization, root‑cause hypothesis generation, runbook assistance, auto‑remediation with human approval) with appropriate guardrails.
  • Define, implement, and operationalize reliability metrics by establishing and managing SLIs, SLOs, and error budgets to quantify and continuously improve service reliability.
  • Own the centralized SRE service engagement model by defining service tiers, onboarding criteria, reliability standards, and a transparent intake/prioritization process aligned to business criticality.
  • Define and enforce error‑budget policies, including escalation paths and release risk decisions, in partnership with Product and Engineering to guide trade‑offs between feature velocity and reliability.
  • Establish and maintain centralized “paved road” reliability standards and assets (instrumentation conventions, golden signals, alerting standards, runbook templates, SLO dashboards) for easy adoption by product teams.
  • Design the on‑call and escalation model for a centralized SRE team, ensuring clear ownership boundaries and minimizing single‑team dependency.
  • Design and engineer automation and observability solutions, developing tooling, dashboards, and systems to reduce operational toil, enhance system visibility, and accelerate delivery.
  • Participate in incident and problem management as incident coordinator for high‑severity events, driving cross‑functional responses, conducting blameless root‑cause analyses, and running post‑incident reviews with clear owners and due dates.
  • Oversee operational readiness and performance by managing capacity planning, validating disaster recovery, conducting production readiness reviews, and ensuring systems meet availability, scalability, and recovery expectations.
  • Partner with security, risk, and compliance teams to align reliability goals with governance and compliance requirements, ensuring secure, auditable, and well‑documented practices.
  • Collaborate across the organization with end users, product management, development, architecture, and IT operational teams to embed reliability principles throughout the software development lifecycle.
  • Champion reliability as a core product feature by promoting it throughout all phases of development, advocating for continuous improvement, and communicating key metrics and potential customer impact to stakeholders.
  • Train, mentor, and upskill engineering teams, coaching engineers in SRE practices and fostering a culture of shared ownership and accountability for reliability.
  • Stay current on the latest SRE trends and best practices, including observability, AI‑enabled operations, and SLO management, and implement these methodologies to support business outcomes.
  • Evaluate AI tools for reliability with security, privacy, and compliance guardrails, and measure impact.
  • Participate in on‑call rotations and operational support for SRE‑supported systems and products.
Required Qualifications
  • Experience in Site Reliability Engineering with a track record of delivering measurable improvements in uptime, scalability, release stability, and overall reliability in complex enterprise environments.
  • Demonstrated experience establishing or significantly maturing an SRE practice (operating model, service engagement, production readiness, incident/post‑mortem program, and reliability roadmap).
  • Hands‑on experience applying AI/ML to operations (AIOps) or GenAI in production support workflows, with measurable outcomes (MTTD/MTTR, alert fatigue reduction, change failure rate) and responsible use controls.
  • Proven ability to establish Service Level Indicators (SLIs) and SLOs in production environments and to implement them.
  • Demonstrated background in production incident response, leading resolution efforts, conducting blameless post‑incident reviews, and implementing actionable remediation strategies.
  • Strong observability and telemetry expertise: designing instrumentation, building actionable dashboards and alerts, and delivering proactive reliability insights using metrics, logs, and traces.
  • Infrastructure engineering experience with strong Infrastructure as Code skills using Terraform and Ansible.
  • Thorough understanding of CI/CD pipeline design, optimization, and troubleshooting with modern tooling (Azure DevOps, GitHub Actions, Jenkins, or GitLab CI), emphasizing speed, reliability, and security.
  • Practical knowledge of containerization and platform modernization: architecting and operating containerized workloads with Docker, VMware, and Kubernetes (or comparable orchestration platforms).
  • Knowledge of emerging reliability practices, such as SLO automation platforms, AIOps, or predictive operations to advance proactive reliability management.
  • Preferred certifications: AWS Professional, Terraform, Ansible, Azure DevOps, Octopus Deploy, or other automation‑focused credentials.
Education and Experience
  • Master’s degree in Computer Science, Engineering, or equivalent practical experience designing and operating production systems at scale.
  • 7+ years of experience in Site Reliability Engineering.
  • Ideal location: Woodbury, NY; Pittsburgh, PA; Orlando, FL; South Jordan, UT. Hybrid or fully remote is possible based on hiring manager decisions.
Compensation

Salary range for this position (NY Metro/NY State): $146,032 – $162,257. The range will vary for other geographic areas.

Benefits
  • Competitive base pay with a Bonus Target or Variable Pay Incentive Program.
  • Medical, Dental and Vision coverage (plus additional benefits).
  • 401(k) Plan with Company Match.
  • Short‑and Long‑Term Disability, Wellness Programs, Group Life and AD&D Insurance.
  • Paid Vacation, Sick Days and bank Holidays.
  • Employee Engagement Activities, Employee Appreciation Day, DEI Employee Resource Groups, Corporate Social Responsibility initiatives, Service Recognition.
Equal Opportunity Employer

We are proud to be an equal‑opportunity employer. All qualified applicants will receive consideration without regard to age, race, color, sex, gender identity, sexual orientation, religion or creed, ancestry, citizenship, national origin, disability, military or veteran status, marital status, genetic information, or any other characteristic protected by applicable law. We do not tolerate discrimination, harassment, or retaliation. Employment decisions are based solely on qualifications, merit, and business needs. Everyone is welcome here, and we hire based on your ability to do the job, not any protected characteristics. If you need help or reasonable accommodation during the application or hiring process, please let your TA Partner know.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

CardWorks Servicing • Pittsburgh

On-site
USD 146,000 - 163,000
Medical, Dental, and Vision benefits
401(k) Plan with Company Match
Paid Vacation and Sick Days
+1
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks, Inc. • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay including Bonus
401(k) Plan with Company Match
Paid Vacation and Sick Days
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Empower Retirement • Greenwood Village (CO)

Hybrid
USD 114,000 - 166,000
Medical insurance
401(k) with company match
Tuition reimbursement
+3
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Hispanic Alliance for Career Enhancement • United States

Hybrid
USD 114,000 - 166,000
Medical, dental, vision and life
401(k) with company match
Tuition reimbursement
+3
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Plano (TX)

On-site
USD 150,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

hardrockdigital • United States

Hybrid
USD 120,000 - 160,000
Competitive pay and benefits
Flexible vacation allowance
Startup culture with global brand support
+1
Senior Software Engineer: Site Reliability Engineering
Senior Software Engineer: Site Reliability Engineering

Jack Henry & Associates, Inc. • Northern (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineering Lead
Site Reliability Engineering Lead

Fayette Chamber of Commerce • Atlanta (GA)

On-site
USD 140,000 - 190,000
Senior Software Engineer: Site Reliability Engineering
Senior Software Engineer: Site Reliability Engineering

Jack Henry & Associates, Inc. • United States

On-site
USD 140,000 - 180,000
Comprehensive benefits