Site Reliability Engineer (SRE) Lead

The Matlen Silver Group, Inc.

Jersey City (NJ)

On-site

USD 140,000 - 190,000

Full time

6 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

The Matlen Silver Group, Inc. in Jersey City, NJ is seeking a Lead Site Reliability Engineer (SRE) to partner with development and production teams, driving reliable cloud infrastructure and automated tooling for enterprise-scale systems.

You will own monitoring designs, implement observability, mentor engineers, and contribute to incident response and problem management to reduce toil and enable scalable software delivery in a financial services environment.

Qualifications

  • Hands-on experience with Microsoft Azure infrastructure and platform services.
  • Experience with Terraform, Terraform Enterprise, or comparable infrastructure-as-code tooling.
  • Working knowledge of Azure networking, including VNets, subnets, route tables, firewalls, DNS, private endpoints, and load balancing.
  • Experience with observability platforms (Azure Monitor, Dynatrace, Splunk, Prometheus, Grafana).
  • Experience with CI/CD tooling (Jenkins, Azure DevOps, GitHub Actions).
  • Scripting or programming experience (Python, PowerShell, Bash).

Responsibilities

  • Collaborate with Development and Infrastructure teams to implement monitoring capabilities.
  • Mentor SRE resources on reliability practices.
  • Develop and maintain reliability scripts, tools, libraries.
  • Implement code changes to leverage reliability libraries and tools.
  • Participate in on-call rotations and incident triage with Problem Management.
  • Identify vulnerabilities and opportunities for reliability improvement.

Skills

Azure infrastructure
Terraform
CI/CD tooling
Python/PowerShell
Observability
Cloud nephrology?

Education

Microsoft Azure certification

Tools

Azure Monitor
Azure Log Analytics
Dynatrace
Splunk
Prometheus
Grafana
Jenkins
Azure DevOps
GitHub Actions

Job description

Site Reliability Engineer (SRE) Lead (BH-110924)

Location Jersey City, United States Sector Financial Services

Position Summary: The individual in this role is responsible for directly partnering with Application Development and Production Support teams to implement the measures prescribed through the collaboration of the Site Reliability Engineer (SRE) Lead or Senior SRE and their partners. This individual will ensure the appropriate instrumentation, tooling, ticketing, alerting and on-call routines are in place for key services. This role will be engaged in production triage efforts and work with Problem Management in the identification of root cause for issues as required, using the knowledge gained in those efforts to partner closely with the Senior SRE to address any gaps in the reliability measurements and dashboards. This role will also focus heavily on software development activities, with a focus toward delivering automated solutions to eliminate toil and suggest code enhancements to the Application Development teams.

Key responsibilities
  • Collaborate with Development and Infrastructure teams to understand technical solutions and to implement the monitoring capabilities outlined in the application and system monitoring designs put forward by the SRE Lead.
  • Mentor SRE resources on reliability practices and established tools/capabilities.
  • Develop and maintain a catalog of extensible reliability scripts, tools and libraries that can be leveraged for common instrumentation, automation, and operational needs.
  • Partner to implement code changes to make use of common reliability libraries and tools and help Application Production Services (APS) and Application Development teammates understand how to use them.
  • Partner with infrastructure engineers and application teams to implement the necessary code changes to make use of common reliability libraries and tools and help the APS and Application Development teammates understand how to use them.
  • Engage as a subject matter expert (SME) in major incident triage efforts, failure scenario modelling and work with Problem Manager to diagnose root causes for major incident / problem management investigations.
  • Identify vulnerabilities and opportunities for reliability improvement, such as investigating low level error rates and 'noise' in monitoring, and to help define solutions to reduce manual support effort and/or improve system reliability. Participate regularly in an on-call rotation with Production Support teammates to learn more about reliability issues affecting their portfolio
Primary Skill: Site Reliability Engineering
  • Hands-on experience with Microsoft Azure infrastructure and platform services.
  • Experience with Terraform, Terraform Enterprise, or comparable infrastructure-as-code tooling.
  • Working knowledge of Azure networking, including VNets, subnets, route tables, firewalls, DNS, private endpoints, and load balancing.
  • Experience with Azure Monitor, Azure Log Analytics, Dynatrace, Splunk, Prometheus, Grafana, or similar observability platforms.
  • Experience with CI/CD tooling such as Jenkins, Azure DevOps, GitHub Actions, or comparable pipeline frameworks.
  • Scripting or programming experience with Python, PowerShell, Bash, or similar languages.
  • Understanding of cloud resiliency, high availability, logging, monitoring, and operational support practices.
  • Strong analytical, troubleshooting, organizational, and communication skills.
  • Ability to collaborate effectively with globally distributed engineering and operations teams.
Desired Qualifications
  • Microsoft Azure certification.
  • Experience in financial services, regulated technology, or other enterprise-scale environments.
  • Experience with AKS, ACR, Kubernetes, container registries, or container-based deployment models.
  • Experience with Azure AI Foundry, OpenAI/GenAI platform operations, model-serving observability, or AI platform readiness.
  • Familiarity with SRE practices such as SLIs, SLOs, incident response, problem management, post-incident reviews, and toil reduction.
  • Exposure to IAM, cloud security, policy-as-code, vulnerability remediation, and governance dashboards.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer -Jersey City, NJ & Dallas, TX
Site Reliability Engineer -Jersey City, NJ & Dallas, TX

StradIT • Jersey City (NJ)

Hybrid
USD 120,000 - 160,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Jobtailor • New Jersey

On-site
USD 120,000 - 180,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Discretionary incentive plan
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

New York Technology Partners • Chicago (IL)

On-site
USD 120,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Cosm Inc. • El Segundo (CA), Northern (KY)

Hybrid
USD 110,000 - 145,000
Senior Engineer - Site Reliability Engineering
Senior Engineer - Site Reliability Engineering

LSEG (London Stock Exchange Group) • Allen (TX)

On-site
USD 140,000 - 190,000