Lead Site Reliability Engineer

Luxoft

United States

On-site

USD 140,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Luxoft seeks a Senior Site Reliability Engineer to ensure reliability, scalability, and operational excellence of banking platforms. You will design, implement, and improve SRE practices across the SDLC, partnering with development, infrastructure, and business teams to boost resiliency through automation and observability.

You will lead incident responses, perform RCAs, and drive automation and tooling to improve performance, capacity, and fault tolerance in a cloud-native, Azure-based

Qualifications

  • First 2 weeks onsite in Buffalo or Wilmington, then remote with occasional travel.
  • Strong experience in observability and monitoring with hands-on expertise.
  • Proficient in OpenTelemetry, Dynatrace, and centralized logging.
  • Experience designing automated regression testing frameworks.
  • Proven IaC expertise with Terraform.
  • Experience with CI/CD pipelines and deployment automation.
  • Expert knowledge of production systems monitoring and incident management.
  • Understanding of cloud-native architecture and distributed systems.
  • Experience with Azure including resource groups, scaling, deployment, and lifecycle management.
  • Familiarity with Azure-native tooling such as Application Insights, Azure Monitor, and Log Analytics.

Responsibilities

  • Design, implement, and support highly available, scalable, and resilient apps and cloud infra following SRE best practices.
  • Lead initiatives to improve reliability, availability, and operational maturity through automation.
  • Define and monitor SLOs, SLIs, and error budgets for critical services.
  • Develop observability strategies with Dynatrace, OTel, metrics, logs, dashboards, and alerts.
  • Design end-to-end monitoring solutions for health of apps, infra, and customer experience.
  • Analyze production telemetry to identify bottlenecks and capacity constraints.
  • Lead incident response for high-severity events and coordinate cross-functional teams.
  • Perform Root Cause Analysis and implement corrective actions.
  • Drive operational excellence via automation of tasks, workflows, deployments, and recovery procedures.
  • Collaborate with development to build observable services across the SDLC.
  • Develop automated regression testing strategies to validate stability after deployments.
  • Review architectural designs to improve resiliency and cloud optimization.
  • Lead capacity planning, performance tuning, and workload optimization.
  • Maintain runbooks, incident playbooks, and standard operating procedures.
  • Partner with engineering, infrastructure, cybersecurity, architecture, and support teams for continuous improvement.
  • Communicate system health and reliability trends to stakeholders.
  • Present reliability initiatives at architecture reviews and leadership meetings.
  • Mentor engineers on observability, cloud engineering, automation, and SRE principles.

Skills

Observability
Monitoring
Dynatrace
OpenTelemetry
Distributed tracing
Logging
Dashboards
Alerting
Regression testing
Terraform
CI/CD
Azure
Incident management
SRE
Automation
Cloud-native
Distributed systems
Performance tuning
Capacity planning

Tools

Terraform
Azure Monitor
Application Insights
Log Analytics

Job description

Responsible at the expert level for ensuring the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. Serves as a senior individual contributor responsible for designing, implementing, and improving Site Reliability Engineering (SRE) practices across the software development lifecycle. Works closely with application development, infrastructure, platform engineering, and business teams to enhance system resiliency through automation, observability, testing, and proactive operational management while coaching and influencing others.

Responsibilities
  • :Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise technology standards and SRE best practices
  • .Lead initiatives to improve system reliability, availability, performance, and operational maturity through automation and engineering excellence
  • .Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical business services
  • .Develop comprehensive observability strategies leveraging Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting solutions
  • .Design and maintain end-to-end monitoring solutions that provide actionable insights into application, infrastructure, and customer experience health
  • .Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints
  • .Lead incident response activities for high-severity production events, coordinating cross-functional teams to restore services and minimize customer impact
  • .Perform and facilitate Root Cause Analysis (RCA) activities, ensuring corrective and preventive actions are identified, prioritized, and implemented
  • .Drive operational excellence through automation of repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls
  • .Partner with development teams to build reliable and observable services throughout the Software Development Lifecycle (SDLC)
  • .Design, develop, and execute automated regression testing strategies to validate application stability, reliability, and performance following deployments and infrastructure changes
  • .Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation
  • .Create, maintain, and improve Infrastructure as Code (IaC) solutions using Terraform for cloud infrastructure provisioning, configuration management, and environment standardization
  • .Support and optimize Microsoft Azure environments, including Azure App Services, resource management, scaling strategies, deployment automation, and application lifecycle management
  • .Utilize Azure-native tools such as Azure Monitor, Application Insights, Log Analytics, and related services to improve platform visibility and reliability
  • .Drive implementation of performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness within assigned domains
  • .Establish operational readiness standards and ensure applications meet reliability, scalability, observability, and supportability requirements before production deployment
  • .Review architectural designs and provide recommendations to improve platform resiliency, operational efficiency, and cloud optimization
  • .Lead capacity planning, performance tuning, and workload optimization efforts across production environments
  • .Develop and maintain operational runbooks, incident playbooks, knowledge articles, and standard operating procedures
  • .Serve as a key partner with engineering, infrastructure, cybersecurity, architecture, and support teams to identify and implement continuous process improvements spanning organizational boundaries
  • .Communicate system health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders
  • .Present reliability initiatives, operational metrics, and engineering recommendations at architecture reviews, technical forums, and leadership meetings
  • .Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational best practices
  • .Understand and adhere to the Company's risk and regulatory standards, policies, and controls in accordance with the Company's Risk Appetite
  • .Identify reliability, operational, and technology risks requiring escalation to management
  • .Promote an environment that supports a culture of belonging and reflects the Client brand
  • .Maintain Client internal control standards, including timely implementation of internal and external audit findings and regulatory requirements as applicable
  • .Complete other related duties as assigned
Mandatory Skills Descriptio

First 2 weeks onsite in Buffalo or Wilmington ,after that remote with occasional travelin

  • g Strong experience in observability and monitoring, including hands-on expertise wit
  • h:Dynatra
  • ceOpenTelemetry (OTe
  • l)Distributed traci
  • ngMetrics collection and analys
  • isCentralized logging and log aggregati
  • onAlerting and dashboard developme
  • ntProven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployment
  • s.Strong proficiency in Infrastructure as Code (IaC) using Terrafor
  • m.Experience with CI/CD pipelines, deployment automation, and operational toolin
  • g.Expert knowledge of production systems monitoring, incident management, and operational troubleshootin
  • g.Strong understanding of application performance management, distributed systems, and modern cloud-native architecture
  • seStrong experience with Microsoft Azure, includin
  • esResource Grou
  • tsScaling and performance optimizati
  • onDeployment and release manageme
  • ntExperience leveraging Azure-native operational tooling such a
  • orApplication Insigh
  • csAzure dashboards and alerti
  • ngExperience supporting cloud-native and hybrid infrastructure environment

s.Reliability & Engineering Practic

  • esDemonstrated experience implementing and operating SRE practices, includin
  • ntRoot Cause Analysis (RC
  • A)Reliability automati
  • onAbility to improve system reliability throug
  • h:Performance tuni
  • ngCapacity planni
  • onReliability engineering initiativ
  • esExperience developing automated recovery mechanisms and self-healing solution
  • s.Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architecture
Nice-to-Have Skills Descripti
  • on:Experience supporting large-scale enterprise applications in regulated environmen
  • ts.Strong analytical and troubleshooting skills related to production systems and distributed architectur
  • es.Experience working in Agile and DevOps operating mode
  • ls.Ability to work autonomously and lead complex reliability initiativ
  • es.Strong organizational and time management skil
  • ls.Advanced verbal and written communication skil
  • ls.Experience driving project milestones and delivery commitmen
  • ts.Proven experience leading major incident response and post-incident improvement effor
  • ts.Experience partnering with architecture, infrastructure, cybersecurity, and application development tea
  • ms.Experience with scripting and automation using PowerShell, Python, Bash, or similar technologi
  • es.Industry certifications in Azure, Terraform, Cloud Engineering, or Site Reliability Engineering preferr
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Buffalo (NY)

On-site
USD 140,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • Wilmington (DE)

On-site
USD 140,000 - 190,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Chandler (MN)

On-site
USD 100,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Chandler (AZ)

On-site
USD 140,000 - 190,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Software Engineering Manager – Site Reliability Center
Software Engineering Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

System One • Pittsburgh

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000