Lead Associate – Service Reliability Engineering (SRE) – $140-170K Plus 20% Bonus

ACCsurance, LLC

Town of Texas (WI)

On-site

USD 140,000 - 170,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

A leading IT consultancy is seeking a Lead Associate in Service Reliability Engineering. This role involves designing and implementing reliable software solutions while collaborating with cross-functional teams in a dynamic environment. The ideal candidate will have over 4 years of experience and proficiency in full-stack development, cloud technologies like AWS, and reliability metrics. Key responsibilities include improving system performance, coaching junior staff, and driving the adoption of reliability best practices. Competitive salary and benefits offered.

Qualifications

  • 4+ years of relevant professional experience.
  • Experience advising teams in Performance and Chaos Engineering strategies.
  • Skilled in establishing and maintaining technology products' health.

Responsibilities

  • Determine customer needs and resolve conflicting requirements.
  • Design and develop software solutions to meet customer needs.
  • Implement new software technology across teams.

Skills

Excellent verbal and written communication skills
Full-stack development (Java, Python)
Cross-functional collaboration
Negotiation and persuasion skills
Managing project plans in Agile/Scrum
Experience with performance engineering
Cloud technologies (AWS)
Reliability Metrics measurement
DevOps/CICD tools (Jenkins, Terraform)
Containerization (Docker, Kubernetes)

Education

Bachelor’s Degree or Equivalent

Tools

JMeter
Load Runner
Splunk
Dynatrace
Datadog
Jenkins
Terraform

Job description

Lead Associate – Service Reliability Engineering (SRE) – $140-170K Plus 20% Bonus

Lead Associate – Service Reliability Engineering (SRE) – $140-170K Plus 20% Bonus — 3 OPENINGS

As a valued contributor to our Fortune 200 Client’s team, you will design, produce, test, or implement software, technology, or processes across multiple projects, programs, or products, as well as create and maintain IT architecture, large scale data stores, and cloud-based systems.

You will apply expertise in software and systems engineering to ensure that both their internally critical and externally visible systems meet the appropriate performance needs of their users. In this role, you will be expected to: drive technical capabilities for increasing SRE value proposition within the portfolio; strategize portfolio / program reliability by working with cross-functional IT organizations and build roadmaps to drive reliability into the product; enable the portfolio to standardize and adopt application reliability metrics and improve application health; serve as a change agent in educating portfolio team members on reliability, evangelize SRE service capabilities, influence customers in adopting SRE services and best practices, measure and showcase the added value through metrics; educate and coach junior resources on application and infrastructure reliability best practices; and become the go-to person for all technical needs of the portfolio SREs.

THE IMPACT YOU WILL MAKE

The Service Reliability Engineering (SRE) Lead role will offer you the flexibility to make each day your own, while working alongside people who care, so that you can deliver on the following responsibilities:

  • Independently determine the needs of the customer while identifying and resolving conflicting or complementary needs across customer groups.
  • Applying advanced skill, knowledge and experience, design and develop software solutions to meet customer needs.
  • Use a process-driven approach to leading design solutions.
  • Implement new software technology and coordinate simultaneous implementation tasks across teams.
  • May maintain or oversee the maintenance of existing software.
THE EXPERIENCE YOU BRING TO THE TEAM
Required Experience:
  • 4+ years of relevant professional experience;
  • Excellent verbal and written communication skills with experience presenting information and/or ideas to an audience in a way that is engaging and easy to understand;
  • Experience as a full-stack developer with hands on knowledge of languages like Java, Python etc. and exposure with application / infrastructure architecture;
  • Experience collaborating cross-functionally on availability / performance issues in order to identify root-cause, determine areas for improvement, and drive those actions to closure through effective solutions;
  • Extensive knowledge of principles, advanced techniques, and theories to suggest and implement solutions on a specific project, program, or product;
  • Experience identifying and selecting strategic options, and identifying resources to meet the defined objectives;
  • Influencing skills to include negotiation, persuasion of others, meeting facilitation, and conflict resolution;
  • Skilled in deriving business insight for the purposes of advising stakeholders and project team members, designing business models, interpreting customer and market insights, forecasting, benchmarking, etc.;
  • Adept at managing project plans, resources, and people to ensure successful project completion in an Agile / Scrum environment in order to facilitate the design / development of performance engineering and resiliency methodologies through collaboration with engineering and product teams to implement shift left techniques on test design & automation;
  • Experience advising teams in the writing of Performance and Chaos Engineering strategies and scripts with a strong emphasis on automated deployment, infrastructure automation solutions, and continuous integration & delivery processes;
  • Ability to identify gaps in the code from a non-functional viewpoint and experience assisting developers to fix the code and promote relevant reliability pattern implementations;
  • Skilled in establishing and maintaining the overall health, availability, performance, resiliency, and capacity of technology products with specific experience in performance engineering and validations using JMeter, Load Runner, etc.;
  • Skilled in cloud technologies and cloud computing to include Amazon Web Services (AWS) offerings, development, and networking platforms;
  • Experience defining, measuring, and improving Reliability Metrics (SLO/SLI), Observability (Monitoring, Logging-Tracing solutions), Operations Processes (Incident, Problem Management), and Operations Toil Reduction through Automation;
  • Experience designing, building and implementing necessary dashboards from application and infrastructure health perspectives using tools such as Splunk, Dynatrace, Datadog, etc. to provide a single pane view of all critical business and operational information to relevant stakeholders;
  • Experience architecting solutions for the design and implementation of applications in the cloud;
  • Experience in activities like architecture reviews, code reviews, creating platforms and frameworks, capacity planning, etc.;
  • Experience designing & developing highly available systems that utilize load balancing, horizontal scalability, and high availability;
  • Knowledge on Cloud technologies and containerization using Docker & Kubernetes;
  • Excellent understanding and demonstrated experience in the use of DevOps / CICD tools like Jenkins, Terraform, Jules and automated deployment tools;
  • Familiarity with Blue Prism, Selenium, or Ansible playbooks and programming languages like Java, Perl, Python or PowerShell scripting and Ansible playbook;
  • Experience implementing resiliency design pattern frameworks and validation.
Desired Experience :
  • Bachelor’s Degree or Equivalent;
  • Relevant certifications such as AWS Certified Solutions Architect, AWS Certified SysOps Administrator, Splunk Certified Developer, Dynatrace, Sun Certified Java Programmer, etc.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

On-site
USD 120,000 - 160,000
Sr. Site Reliability Engineer(Local to Atlanta GA Only)
Sr. Site Reliability Engineer(Local to Atlanta GA Only)

Trigint Solutions LLC • Atlanta (GA)

Hybrid
USD 124,000 - 220,000
Lead SRE Associate — Reliability & Cloud (20% Bonus)
Lead SRE Associate — Reliability & Cloud (20% Bonus)

ACCsurance, LLC • Town of Texas (WI)

On-site
USD 140,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas
Engineering - SRE Platforms - SRE Engineer - Associate - Dallas

The Goldman Sachs Group • Dallas (TX)

On-site
USD 110,000 - 140,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

On-site
USD 146,032 - 162,257
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Sr. Director, Site Reliability and Platform Engineering
Sr. Director, Site Reliability and Platform Engineering

Optomi • Tacoma (WA)

On-site
USD 150,000 - 200,000