Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance

Irving (TX)

Hybrid

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mindlance is seeking a Lead Site Reliability Engineer (SRE) to work in a hybrid role in Irving, TX and Charlotte, NC. The ideal candidate will possess at least 8 years of SRE experience and expertise in Linux, container platforms like Kubernetes, and multiple cloud environments such as AWS, GCP, or Azure.

Responsibilities include implementing automated tooling, enhancing system availability, and supporting critical applications. This position offers an opportunity to innovate with AIOps and facilitate Agile-based remediation efforts.

Qualifications

  • 8+ years minimum experience as a Site Reliability Engineer (SRE).
  • Expertise in Linux and container platforms.
  • Experience with multiple cloud platforms.

Responsibilities

  • Design and implement automated tooling to optimize operations.
  • Enhance system availability in a multi-cloud environment.
  • Support critical applications and lead remediation efforts.

Skills

SRE experience
Database knowledge
Observability tools

Tools

Kubernetes
Oracle
AWS
GCP
Azure

Job description

Overview

Title: Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Location: Irving, TX & Charlotte, NC - Hybrid Role

Duration: 18+ Months (s) Contract to hire, or possibility to extension

We are seeking a Senior Site Reliability Engineer (SRE) with a strong background in software engineering and a passion for solving complex problems at scale. This role blends software engineering with operational expertise to deliver stable, scalable, and resilient services, while reducing toil and shifting operations left.

Runs support for Shared Services Operations Technology. Split amongst Payment Evaluations, Regulatory Operations, Financial Crimes, and Business and Real Estate Evaluation. Supports systems that do KYC and AML supporting financial crimes. Have about 85 apps they support, about 75 of those have no SLOs and SLI s, so they\'d like those defined. Also getting into automation with RPA and chatbots. Hoping to find someone who could apply to any one of the domains. High volume of tickets in the org, but this person would be expected to be working more proactively on projects. Right now, that person may be "firefighting" 60% of the time and doing prevention the other 40%, but would like to improve to 80% prevention.

OCP is highly preferable for cloud experience since it\'s being implemented across the organization.

Backfilling an FTE with someone they\'d like to try out. May be some weekends that require system support, overtime could be an occasional possibility. May work weekends once a month or two months on a rotation, depending on if they\'re assigned to that rotation as an SRE.

Key Responsibilities
  • Design and implement automated tooling to eliminate manual toil and optimize operations.
  • Build and enhance monitoring, alerting and overall observability.
  • Champion the SRE practice within COO Technology by modeling best practices, mentoring peers, and collaborating with embedded platform SRE teams.
  • Enhance system availability in a multi-cloud environment by evolving resiliency patterns.
  • Introduce and scale AIOps, including self-healing and autonomic systems using AI/ML, RPA, and unified communications.
  • Automate key SRE metrics and IT service operations processes, including customer impact analysis, availability tracking, SLO/SLI adherence, error budgeting, and incident response.
  • Support critical applications and customer journeys, lead Agile-based remediation efforts, conduct blameless postmortems, and drive root cause analysis to eliminate recurring issues.
  • Implement and guide through Non-Functional Requirements (NFRs) during modernization and uplift initiatives.
  • Help define, govern and enforce Permit to Operate.
Top Skills
  • 8+ years minimum SRE experience
  • Database knowledge
  • Observability tools
Nice to Have
  • Autosys
  • A good SRE will likely be interested in AI
Infrastructure & Cloud
  • Expertise in Linux and container platforms (Kubernetes)
  • Experience with cloud platforms: PCF, AWS, GCP, or Azure
CI/CD & Automation
Observability & AIOps
Operations & Data
  • Data platforms: Oracle, DB2, SQL, MongoDB, Hadoop, Cloudera, Spark, Teradata
EEO

Mindlance is an Equal Opportunity Employer and does not discriminate in employment on the basis of – Minority/Gender/Disability/Religion/LGBTQI/Age/Veterans.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

OutSolve • Mission (KS)

Remote
USD 90,000 - 130,000
100% remote work environment
Competitive compensation
Professional development opportunities
+1
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer -Jersey City, NJ & Dallas, TX
Site Reliability Engineer -Jersey City, NJ & Dallas, TX

StradIT • Jersey City (NJ)

Hybrid
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Optomi • Seattle (WA)

On-site
USD 140,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Optomi • New York (NY)

On-site
USD 140,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The Glove • Cleveland (OH)

On-site
USD 120,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Discretionary incentive plan
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

JPS Tech Solutions • Colorado

On-site
USD 160,000 - 230,000