Senior Site Reliability Engineer - Azure & Observability

Luxoft

United States

On-site

USD 140,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Luxoft seeks a Senior Site Reliability Engineer to ensure reliability, scalability, and operational excellence of banking platforms. You will design, implement, and improve SRE practices across the SDLC, partnering with development, infrastructure, and business teams to boost resiliency through automation and observability.

You will lead incident responses, perform RCAs, and drive automation and tooling to improve performance, capacity, and fault tolerance in a cloud-native, Azure-based

Qualifications

  • First 2 weeks onsite in Buffalo or Wilmington, then remote with occasional travel.
  • Strong experience in observability and monitoring with hands-on expertise.
  • Proficient in OpenTelemetry, Dynatrace, and centralized logging.
  • Experience designing automated regression testing frameworks.
  • Proven IaC expertise with Terraform.
  • Experience with CI/CD pipelines and deployment automation.
  • Expert knowledge of production systems monitoring and incident management.
  • Understanding of cloud-native architecture and distributed systems.
  • Experience with Azure including resource groups, scaling, deployment, and lifecycle management.
  • Familiarity with Azure-native tooling such as Application Insights, Azure Monitor, and Log Analytics.

Responsibilities

  • Design, implement, and support highly available, scalable, and resilient apps and cloud infra following SRE best practices.
  • Lead initiatives to improve reliability, availability, and operational maturity through automation.
  • Define and monitor SLOs, SLIs, and error budgets for critical services.
  • Develop observability strategies with Dynatrace, OTel, metrics, logs, dashboards, and alerts.
  • Design end-to-end monitoring solutions for health of apps, infra, and customer experience.
  • Analyze production telemetry to identify bottlenecks and capacity constraints.
  • Lead incident response for high-severity events and coordinate cross-functional teams.
  • Perform Root Cause Analysis and implement corrective actions.
  • Drive operational excellence via automation of tasks, workflows, deployments, and recovery procedures.
  • Collaborate with development to build observable services across the SDLC.
  • Develop automated regression testing strategies to validate stability after deployments.
  • Review architectural designs to improve resiliency and cloud optimization.
  • Lead capacity planning, performance tuning, and workload optimization.
  • Maintain runbooks, incident playbooks, and standard operating procedures.
  • Partner with engineering, infrastructure, cybersecurity, architecture, and support teams for continuous improvement.
  • Communicate system health and reliability trends to stakeholders.
  • Present reliability initiatives at architecture reviews and leadership meetings.
  • Mentor engineers on observability, cloud engineering, automation, and SRE principles.

Skills

Observability
Monitoring
Dynatrace
OpenTelemetry
Distributed tracing
Logging
Dashboards
Alerting
Regression testing
Terraform
CI/CD
Azure
Incident management
SRE
Automation
Cloud-native
Distributed systems
Performance tuning
Capacity planning

Tools

Terraform
Azure Monitor
Application Insights
Log Analytics

Job description

Luxoft seeks a Senior Site Reliability Engineer to ensure reliability, scalability, and operational excellence of banking platforms. You will design, implement, and improve SRE practices across the SDLC, partnering with development, infrastructure, and business teams to boost resiliency through automation and observability.

You will lead incident responses, perform RCAs, and drive automation and tooling to improve performance, capacity, and fault tolerance in a cloud-native, Azure-based

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Azure SRE: Automation, Observability & Reliability
Senior Azure SRE: Automation, Observability & Reliability

Koitecc Solutions • Plano (TX), Northern (KY)

Hybrid
USD 153,000 - 192,000
Discretionary incentive
Benefits package
Senior Azure SRE: Reliability, Automation & Observability
Senior Azure SRE: Reliability, Automation & Observability

Bank of America • Chandler (AZ)

On-site
USD 140,000 - 190,000
Senior Azure SRE - Platform Reliability & Automation Lead
Senior Azure SRE - Platform Reliability & Automation Lead

Hobbsnews • Plano (TX)

On-site
USD 152,000 - 192,000
Discretionary incentive eligible
Benefits eligible
Senior Site Reliability Engineer: Cloud & Automation
Senior Site Reliability Engineer: Cloud & Automation

LSEG (London Stock Exchange Group) • Creve Coeur (MO)

On-site
USD 120,000 - 180,000
Senior Observability Engineer: Build Reliability & Dashboards
Senior Observability Engineer: Build Reliability & Dashboards

Ripple • New York (NY)

On-site
USD 160,000 - 200,000
Senior Observability & SRE Engineer
Senior Observability & SRE Engineer

Hidden Road • Chicago (IL)

Hybrid
USD 160,000 - 200,000
Equity
Bonuses
Healthcare
+4
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 190,000
Senior Site Reliability Engineer - Cloud & Automation
Senior Site Reliability Engineer - Cloud & Automation

LSEG (London Stock Exchange Group) • St. Louis (MO)

On-site
USD 120,000 - 180,000
Azure SRE: Reliability, Observability & Incident Leadership
Azure SRE: Reliability, Observability & Incident Leadership

Veriipro • Deerfield (IL)

On-site
Azure SRE Lead: Reliability Strategy & Incident Leadership
Azure SRE Lead: Reliability Strategy & Incident Leadership

Tata Consultancy Services • Bellevue (WA)

On-site
USD 38,000 - 121,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
Maternal & Parental Leaves
+4