Site reliability engineering(SRE) Operations at Re Focus LLC O Fallon, MO

Re Focus LLC

O’Fallon (MO)

On-site

USD 75,000 - 105,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Re Focus LLC is seeking a skilled IT Operations Specialist to review and resolve incidents from monitoring and deployment activities. You will coordinate with development and regional support teams to fix issues, manage change tickets, and perform root cause analysis for high-severity incidents to prevent recurrence.

The role requires L2 support experience, Unix shell scripting and SQL proficiency, and familiarity with Splunk and Dynatrace for troubleshooting.

Qualifications

  • Unix Shell Scripting and SQL are essential.
  • Ability to troubleshoot using logs and monitoring tools such as Splunk or Dynatrace.
  • ITSM Incident, Change and Problem Management experience is required.
  • L2 support experience is a must.
  • Familiarity with Snowflake is a plus.

Responsibilities

  • Review and resolve incidents arising from monitoring and deployment processes.
  • Coordinate with development and support teams to fix deployment issues.
  • Manage change tickets and obtain CAB approvals as needed.
  • Assist in root cause analysis for high-severity incidents and preventive actions.
  • Support UAT and customer onboarding activities.
  • Create automation scripts to reduce incidents and improve processes.
  • Participate in war rooms affecting availability or customer impact.

Skills

Unix Shell Scripting
SQL
Troubleshooting logs
Splunk
Dynatrace
ITSM Incident Change Problem
L2 Support
Snowflake

Tools

Remedy
Rally
Splunk
Dynatrace
WinSCP
CyberArk/Putty
Toad Querying Tool

Job description

Job Description
Roles/responsibilities:
  • Incident Resolution - Review and resolve the Incidents arising from
  • Operation Command Center Alerts
  • Alerts from Enterprise Monitoring Operations (EM Operations).
  • OMNIBUS and Splunk Alerts
  • Change Implementation - Deploying the application related artifacts to the production environments in the slotted approved release window
  • Reporting the issues with the deployments and coordinating with the Development Teams to fix deployment issues
  • Work Orders - Resolve Work orders in form of Business/functional queries, adhoc testing, verification and validation etc, from Regional product team and customer support teams.
  • Traffic Routing perform traffic routing in support of infrastructure maintenance
  • Perform Root Cause Analysis in detail for High severity Incidents and take action on fixing the underlying cause of the high severity issues. Take necessary preventive actions also.
  • Supporting the UAT testing by the Product team and Regional customer support team.
  • Configuring application/artifacts and supporting the new customer onboarding to the platform
  • Raise new change tickets and arrange for approvals, including CAB approvals
  • Review and approve change tickets.
  • Work with customers on ad-hoc queries
  • Work with Development / Testing team for defect analysis (with Production simulated data)
  • Build automation scripts that reduce the number of Incidents and/or improves processes followed
  • Support customer to fill in the Post Incident Report (PIR) when any high impacting Incidents affecting customers occurred.
  • Participate / Initiate in War Room calls that impacts application availability or has a customer impact
  • Willing to work on shifts (Morning & Afternoon shifts) & Weekend support
Must have skills:
  • Unix Shell Scripting, SQL
  • Troubleshooting using logs, Splunk / Dynatrace
  • ITSM Incident, Change and Problem Management
  • L2 Support experience is a must
  • Snowflake
Good to have skills:
  • PCF Cloud knowledge
  • CI/CD, Jenkins, Git & Maven
Tools Used:
  • Remedy Ticketing Tool
  • Rally (For Story and Bug Tracking)
  • Splunk and Dynatrace for Monitoring
  • WinSCP (file movement/ validation)
  • CyberArk/Putty
  • Toad Querying Tool for DB
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
Solutions Architect(SRE)
Solutions Architect(SRE)

Business Needs Inc. • Fort Mill (SC)

On-site
USD 100,000 - 130,000
SRE Production Support
SRE Production Support

SelectMinds LLC • Livonia (MI)

On-site
USD 100,000 - 140,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Knack Solutions • Reston (VA)

On-site
USD 120,000 - 160,000
Lead Site Reliability Engineer(SRE) -St Louis, MO(Hybrid)
Lead Site Reliability Engineer(SRE) -St Louis, MO(Hybrid)

The Dignify Solutions, LLC • St. Louis (MO)

On-site
USD 90,000 - 120,000
Software Engineering Manager – Site Reliability Center
Software Engineering Manager – Site Reliability Center

Jobtailor • Alabama

On-site
USD 120,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

System One • Pittsburgh

On-site
USD 140,000 - 190,000
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

Hybrid
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
T Site Reliability Engineer TEKsystems Phoenix, Arizona, US
T Site Reliability Engineer TEKsystems Phoenix, Arizona, US

Artha Nexgen • Phoenix (AZ), Northern (KY)

Hybrid
USD 120,000 - 180,000
Vice President, SRE Lead (Incident Management), Application Production Services & Engineering
Vice President, SRE Lead (Incident Management), Application Production Services & Engineering

Bank of America • United States

On-site
USD 110,000 - 140,000