Lead Site Reliability Engineer

CardWorks Servicing

Pittsburgh (Allegheny County)

On-site

USD 146,032 - 162,257

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, Dental, and Vision benefits
401(k) Plan with Company Match
Paid Vacation and Sick Days
Employee Engagement Activities

Job summary

CardWorks Servicing is seeking an experienced Site Reliability Engineer to join our team in Pittsburgh, PA. This role involves defining and operationalizing reliability metrics, enhancing service reliability, and adopting AI-enabled solutions. Ideal candidates should have over 7 years of experience, a Master’s degree in a relevant field, and skills in infrastructure as code, CI/CD, and observability practices. Competitive salary range is $146,032 to $162,257, with additional benefits including medical insurance and a 401(k) plan.

Qualifications

  • 7+ years of experience in Site Reliability Engineering.
  • Demonstrated ability to improve uptime and reliability.
  • Strong understanding of automation tools and platforms.

Responsibilities

  • Define and implement reliability metrics and error budgets.
  • Establish and maintain SRE operating model.
  • Participate in incident and problem management.

Skills

Site Reliability Engineering
AI/ML in operations
Service Level Indicators (SLIs)
Infrastructure as Code
CI/CD pipeline design
Observability and telemetry
Containerization

Education

Master’s degree in Computer Science or Engineering

Tools

Terraform
Ansible
Azure DevOps
Jenkins
Docker
Kubernetes

Job description

Join our team - and take the next step in achieving a fulfilling career!What We DoAt CardWorks, we aim to help people connect with possibility and opportunity using our financial servicing expertise. Building meaningful, long-term relationships with consumers, our employees, and our clients is what matters most.Who We AreCardWorks, Inc. is a diversified consumer finance service provider and parent company of CardWorks Servicing, LLC, Merrick Bank and Carson Smithfield, LLC.CardWorks Servicing, LLC provides end-to end operational servicing functions for credit cards, secured cards, and installment loans. We service consumer and small business loans across the credit spectrum and offers backup servicing and due diligence services to capital providers and trustees.Merrick Bank is an FDIC-insured Utah Industrial Loan Bank. Merrick operates three main business lines: credit cards, recreational lending, and merchant services.Carson Smithfield, LLC provides a variety of post-charge-off debt recovery services, including digital self-service, IVR, live agent, and external agency management.Essential Functions:Establish the SRE operating model (service onboarding, engagement model, governance, reliability reviews, production readiness standards, and quarterly planning) and ensure it is adopted across teams.Identify, pilot, and operationalize AI-enabled reliability use cases (e.g., alert noise reduction, incident summarization, correlation/root-cause hypothesis generation, runbook assistance, and auto-remediation with human approval) with appropriate guardrails.Define, implement, and operationalize reliability metrics by establishing and managing SLIs, SLOs, and error budgets to quantify and continuously improve service reliability, supporting engineering and business decisions.Own the centralized SRE service engagement model by defining service tiers, onboarding criteria, reliability standards, and a transparent intake/prioritization process aligned to business criticality.Define and enforce error budget policies (including escalation paths and release risk decisions) in partnership with Product/Engineering, using SLO attainment to guide trade-offs between feature velocity and reliabilityEstablish and maintain centralized “paved road” reliability standards and assets (instrumentation conventions, golden signals, alerting standards, runbook templates, SLO dashboards) that product teams can adopt with minimal friction.Design the on-call and escalation model for a centralized SRE team (e.g., SRE overlay for major incidents, defined handoffs with service owners, and clear ownership boundaries) to improve response quality without creating single-team dependency.Design and engineer automation and observability solutions by developing tooling, dashboards, and systems to reduce operational toil (measure, report, and drive toil down over time), enhance system visibility, and accelerate delivery.Participates in incident and problem management by serving as incident coordinator for high-severity events, driving cross-functional responses, conducting blameless root cause analysis, running post-incident reviews (postmortems) with clear owners and due dates, ensuring remedial actions drive reliability improvements.Oversee operational readiness and performance by managing capacity planning, validating disaster recovery, conducting production readiness reviews, and ensuring systems meet availability, scalability, and recovery expectations.Partner with security, risk, and compliance teams to align reliability goals with governance and compliance requirements, ensuring secure, auditable, and well-documented practices.Collaborate across the organization by working closely with end users, product management, development, architecture, and IT Operational teams to embed reliability principles throughout the software development lifecycle, including service onboarding, reliability reviews, and shared SLO ownership.Champion reliability as a core product feature by promoting reliability throughout all phases of development, advocating for continuous improvement, and communicating key metrics and potential customer impact to stakeholders.Train, mentor, and upskill engineering teams by coaching engineers in SRE practices, supporting junior team members, and fostering a culture of shared ownership and accountability for reliability, including influencing teams without direct authority through standards, data, and executive-aligned priorities. Remain current on the latest SRE trends and best practices, including observability, AI-enabled operations (AIOps), and SLO management, and implement these methodologies to effectively support desired business outcomes. Evaluate AI tools for reliability with security/privacy/compliance guardrails (e.g., data handling, prompt/content controls, auditability) and measure impact.Participate in on-call rotations and operational support for SRE-supported systems and products.Summary of Qualifications:Experience in Site Reliability Engineering with a track record of delivering measurable improvements in uptime, scalability, release stability, and overall reliability in complex enterprise environments.Demonstrated experience standing up or significantly maturing an SRE practice (operating model, SRE/service engagement, production readiness, incident/postmortem program, and reliability roadmap).Hands-on experience applying AI/ML to operations (AIOps) or GenAI in production support workflows, with a focus on measurable outcomes (MTTD/MTTR, alert fatigue reduction, change failure rate) and responsible use controls.Proven ability to establish Service Level Indicators (SLIs) and SLOs in production environments, including hands-on definition and implementation.Demonstrated background in production incident response, leading resolution efforts, conducting blameless post-incident reviews, and implementing actionable remediation strategies.Strong observability and telemetry expertise in designing instrumentation, building actionable dashboards and alerts, and delivering proactive reliability insights using metrics, logs, and traces.Infrastructure engineering experience with strong Infrastructure as Code skills using tools such as Terraform and Ansible.Thorough understanding and practical experience in CI/CD pipeline design, optimization, and troubleshooting using modern tooling and platforms such as Azure DevOps, GitHub Actions, Jenkins, or GitLab CI, with an emphasis on speed, reliability, and security.Practical knowledge of containerization and platform modernization, including architecting and operating containerized workloads with Docker, VMware, and Kubernetes (or comparable orchestration platforms) to modernize legacy applications and improve fault tolerance.Knowledge of emerging reliability practices, including SLO automation platforms, AIOps, or predictive operations to advance proactive reliability management.Preferred certifications include AWS Professional, Terraform, Ansible, Azure DevOps, Octopus Deploy or other automation-focused credentials that demonstrate continuous technical development.Education and Experience:Master’s degree in computer science, Engineering, or equivalent practical experience designing and operating production systems at scale.7+ years of experience in Site Reliability Engineering.Ideally, the qualified candidate will work at the following location(s): Woodbury, NY; Pittsburgh, PA, Orlando, Fl, South Jordan, UT. A hybrid work model or fully remote model can be considered based on hiring manager decision and priorities of the role.The salary range for this position, if located in NY Metro/NY State is $146,032 to $162,257. However, please note that the salary range will vary for other geographic areas.#INDHPOur Employee Value PropositionCompetitive Pay, including a Bonus Target or Variable Pay Incentive ProgramBenefits Package -Medical, Dental, and Vision (plus much more)401(k) Plan with Company MatchShort- & Long-Term DisabilityWellness ProgramsGroup Life and AD&D InsurancePaid Vacation, Sick Days and bank HolidaysEmployee Engagement Activities including Employee Appreciation Day, DEI Employee Resource Groups, Corporate Social Responsibility, Service RecognitionWe offer a total rewards package comprised of a competitive base rate of pay, variable pay incentive programs based on the role, and a comprehensive benefit suite. Offered rates of pay are determined based on job-related knowledge, relevant experience, skills, certifications, and geographic location.We are an equal opportunity employer, and we evaluate qualified applicants without regard to race, color, religion, sex, national origin, disability, veteran status or any other legally protected characteristic. We will conduct a thorough background check for all hires in compliance with applicable laws.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

CardWorks Servicing LLC • Pittsburgh

On-site
USD 146,000 - 163,000
Competitive base pay
Medical, Dental and Vision coverage
401(k) Plan with Company Match
+1
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks, Inc. • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay including Bonus
401(k) Plan with Company Match
Paid Vacation and Sick Days
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Manager, Security Compliance
Manager, Security Compliance

CardWorks, Inc. • Orlando (FL)

On-site
USD 128,000 - 143,000
Competitive Pay
401(k) Plan with Company Match
Paid Vacation, Sick Days and Bank Holidays
+1
Staff Site Reliability Engineer
Staff Site Reliability Engineer

EarnIn • Mountain View (WY)

Hybrid
USD 252,000 - 308,000
Equity
Hybrid work model
Site Reliability Engineer/L3 Support
Site Reliability Engineer/L3 Support

SS&C Technologies, Inc. • New York (NY)

Remote
USD 110,000 - 120,000
401k Matching
Professional Development Reimbursement
Paid Holidays
+2
SRE Architect, AI-Powered Reliability
SRE Architect, AI-Powered Reliability

WEX Inc. • Portland (ME)

On-site
USD 200,000 - 251,000
Health insurance
Dental and vision insurances
Retirement savings plan
+1
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Empower Retirement • Greenwood Village (CO)

Hybrid
USD 114,000 - 166,000
Medical insurance
401(k) with company match
Tuition reimbursement
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Koitecc Solutions • Jersey City (NJ)

On-site
USD 152,600 - 191,500
Discretionary incentive eligible
Benefits package
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Cognizant • Bentonville (AR)

On-site
USD 90,000 - 101,000
Medical/Dental/Vision/Life Insurance
Paid holidays & PTO
401(k) plan
+3