Site Reliability Engineer

System One

Dallas (TX)

On-site

USD 130,000 - 170,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

System One in Dallas, TX is seeking a Senior Site Reliability Engineer for a permanent role. You will monitor distribution systems, assist with on‑call troubleshooting, and manage outages with a focus on rapid restoration and root cause analysis.

You will lead incident reviews, identify automation opportunities, collaborate with multi geolocation teams, and apply AI Ops concepts to improve reliability across customer systems.

Qualifications

  • 5+ years IT experience.
  • 2–3 years analytics and technical lead experience.
  • Production support with client facing experience.
  • Knowledge of system management and analytics tools.
  • Self‑initiative and detail orientation.
  • Ability to propose efficient, cost‑effective solutions.
  • Ability to work with multi-geographical teams.
  • Understanding of AI Ops and NLP.
  • Thought leadership for process improvements and knowledge sharing.
  • Problem solving and solution delivery.
  • Excellent written and oral communication with leadership and SMEs.
  • DB knowledge: PL/SQL, Oracle, DB2.
  • SOAP/REST in production environments.
  • Tomcat/Apache/WAS/IIS knowledge.
  • Linux and Windows server experience.
  • Agile methodologies.
  • Monitoring tools such as Dynatrace, GlassBox, ITCAMS, TrueSight, OEM.

Responsibilities

  • Monitor distribution systems and notify potential issues.
  • Assist with on‑call troubleshooting.
  • Manage and track outages.
  • Lead incident analysis meetings and postmortems.
  • Identify automation opportunities for automation specialists.
  • Oversee in‑scope systems to ensure customer reliability.
  • Ensure process follow‑ups and event completion.
  • Analyze log files for production issues.
  • Participate in critical issue calls and provide expertise.
  • Communicate with client leadership and SMEs effectively.
  • Research and replicate user experience to validate solutions.
  • Provide technical solutions and lead problem resolution.
  • Propose innovative ideas for complex problems.
  • Support extended hours during critical timelines.

Skills

IT experience
Technical lead
Production support
System management
Analytical thinking
AI Ops
NLP
Communication
SQL databases
SOAP/REST
Tomcat/Apache/IIS
Linux
Windows/Linux
Agile
Dynatrace/Monitoring tools

Tools

SOAP/REST
Dynatrace
TrueSight

Job description

Senior Site Reliability Engineer (Permanent Role)

Cleveland, OH, Pittsburgh, PA, or Dallas, TX

Your future duties and responsibilities
  • Monitoring distribution systems and notifying them of any potential issues.
  • Assisting with troubleshooting on call.
  • Managing and tracking incidents such as outages.
  • Facilitating analysis meetings to discuss incidents. Lead
  • Identify the automation opportunities for automation specialists to review and execute.
  • Watching over in scope and related systems to ensure reliability for customers.
  • Ensuring process follow ups and coordination until the event completion.
  • Extensive applications support experience with troubleshooting issues including analyzing log files.
  • Should be able to participate in critical issue calls and contribute expertise to resolve the issues.
  • Very strong in technical concepts, Analytical and Problem‑solving skills, with ability to confidently interact with other client‑side teams.
  • Research and replicate the user experience to validate the solutions.
  • Resolve issues that other junior team members are unable to solve. Lead
  • Provide technical solutions to problems encountered in the team. Lead
  • Provide innovative ideas and solutions to complex as well as rather simple looking problems. Lead
  • Should be flexible to extended and weekend hours during critical timelines to provide periodical oversights.
  • Should be able to work in business hour overlapping shifts for critical meetings and activities.
Required qualifications to be successful in this role
  1. Overall 5+ years of IT experience
  2. 2 to 3 years of business analytics and technical lead experience
  3. Solid understanding of production support activities with client facing experience
  4. Knowledge of system management and support analytic tools
  5. Self initiative, details oriented, capable of bringing new ideas to the table.
  6. Demonstrated ability to propose the most efficient and cost effective solutions.
  7. Ability to work with multi geological teams and collect the inputs for analysis.
  8. Good understanding of AI Ops and NLP.
  9. Ability to provide thought leadership for process improvements and documenting the knowledge.
  10. An inherent ability to solve problems and provide solutions.
  11. Excellent oral and written communication skills to interact and communicate with the client leaderships and technical SMEs.
  12. Excellent knowledge of database (PL/SQL, Oracle, DB2)
  13. Experience developing and consuming SOAP and REST web services for an operational/production environment.
  14. Excellent Knowledge of Tomcat, Apache, WAS, and IIS
  15. Extensive experience of Linux systems.
  16. Good knowledge of Windows and Linux server configurations.
  17. Researches and evaluates alternative technical solutions.
  18. Demonstrated ability to propose the most efficient and cost effective solutions.
  19. Understanding on Agile Methodology
  20. Experience on monitoring tools like Dynatrace, DT Managed, GlassBox, ITCAMS, TrueSight, OEM

Ref: #404-IT Pittsburgh

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

System One • Pittsburgh

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

Inherent Technologies • Schaumburg (IL)

Hybrid
USD 120,000 - 150,000
Site Reliability Lead
Site Reliability Lead

System One • Pittsburgh

Hybrid
USD 120,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Senior Vice President, Site Reliability Automation Engineer
Senior Vice President, Site Reliability Automation Engineer

BNY Mellon • Town of Florida (NY)

Hybrid
USD 100,000 - 130,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Mike Albert Fleet Solutions • Cincinnati (OH)

Hybrid
USD 100,000 - 135,000
Site Reliability Engineer
Site Reliability Engineer

SSi People • Pittsburgh

On-site
USD 90,000 - 120,000
Site Reliability Engineer
Site Reliability Engineer

TechDigital Group • Houston (TX), Juno Beach (FL)

On-site
USD 120,000 - 180,000
Senior AIOps and Incident Management / Site Reliability Engineering C2C jobs
Senior AIOps and Incident Management / Site Reliability Engineering C2C jobs

Tech Mirrors • Fort Mill (SC)

Hybrid
USD 140,000 - 190,000