Senior Site Reliability Engineer

Ll Oefentherapie

Nashville (TN)

On-site

USD 110,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Ll Oefentherapie in Nashville seeks an experienced Site Reliability Engineer to design and operate reliable, secure infrastructure. You will implement scalable solutions across Windows and Linux, automate repetitive tasks, and support incident investigations to minimize downtime.

Responsibilities include patching, vulnerability remediation, and maintaining runbooks while collaborating with software teams to improve reliability. This role offers strong growth in a production environment.

Qualifications

  • Experience administering Windows Server and Linux systems.
  • Ability to troubleshoot OS, service, and app-level issues.
  • Familiarity with STIGs, security hardening, and compliance.

Responsibilities

  • Design and architect reliable, secure infrastructure and services.
  • Install, configure, deploy, validate applications on Windows/Linux.
  • Troubleshoot OS, service, and application issues; perform incident response.
  • Lead or support incident investigations and root cause analyses.
  • Maintain runbooks, deployment procedures, and documentation.

Skills

Windows/Linux Admin
SRE/Incident Response
Scripting & Automation
Cloud & OCI Familiarity
Network Troubleshooting
Cybersecurity & Compliance
Documentation & Runbooks

Tools

PowerShell
Bash
Python
Ansible
Chef

Job description

Takes proactive steps to design and architect infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Performs data collection to maintain and optimize operations and reliability. Leverages knowledge to perform incident response and/or maintenance tasks. Provides health and performance reporting. Identifies opportunities for automation. Communicates about services and identifies and explains the potential impact of changes. Provides support for technology and documents incidents. Experiments with new tools and assesses potential impact and develops knowledge of site reliability trends.Key ResponsibilitiesDesign and architect reliable, secure, and maintainable infrastructure and services. Take proactive steps to ensure solutions meet defined reliability and functionality requirements.Identify operational risks, dependencies, performance issues, and potential failure points before they affect service.Translate business and application requirements into practical technical solutions.Install, configure, deploy, and validate applications across Windows Server and Linux environments.Perform structured builds and deployments using runbooks, scripts, readiness checks, and post-build validation.Troubleshoot operating system, service, application, installation, patching, permissions, certificate, and connectivity issues.Monitor system performance and implement improvements to availability, reliability, and operational efficiency.Develop and maintain scripts and automation that reduce manual effort and improve consistency.Plan and execute operating system, middleware, and application patching using established change-control and rollback procedures.Lead or support incident investigations, root cause analyses, and corrective actions.Support vulnerability remediation, system hardening, STIG compliance, and other security-driven changes.Maintain accurate runbooks, deployment procedures, troubleshooting guides, and operational records.Provide technical guidance and mentorship to junior engineers.Communicate status, risks, blockers, and escalation details clearly to stakeholders.Core Skills and QualificationsWindows and Linux System AdministrationHands-on experience administering Windows Server and/or Linux systems.Ability to access deployed hosts and perform post-deployment configuration and validation.Experience installing, configuring, and validating applications in Windows Server and Linux environments.Ability to troubleshoot operating-system-level, service-level, and application-level issues.Working knowledge of system services, permissions, configuration files, logs, and resource utilization.Manual Build and Deployment ExperienceExperience performing structured build and deployment tasks using runbooks, deployment guides, and technical procedures.Ability to execute scripts, validate outputs, and correct common build or configuration issues.Familiarity with build handoffs, environment-readiness checks, deployment validation, and post-build verification.Ability to follow detailed implementation steps while identifying and documenting exceptions or deviations.Troubleshooting and Operational SupportAbility to investigate failed services, installation errors, patching failures, application startup problems, permissions issues, and connectivity failures.Experience reviewing logs, event viewers, service status, configuration files, ports, certificates, and permissions.Strong analytical and problem-solving skills, with the ability to isolate root causes and recommend practical solutions.Ability to escalate issues clearly by documenting symptoms, troubleshooting steps, findings, impact, and recommended actions.Experience supporting production or other business-critical environments.Scripting and AutomationHands-on experience with one or more of the following:PowerShellBashPythonAnsibleChefCandidates should be able to run, modify, validate, and troubleshoot existing scripts and understand basic automation concepts.Patching and Software MaintenanceExperience applying operating system, middleware, and application patches.Ability to follow patching procedures, validate successful completion, and troubleshoot failures.Understanding of maintenance windows, change control, rollback planning, and post-change validation.Cloud and OCI FamiliarityFamiliarity with Oracle Cloud Infrastructure or another major cloud platform.Understanding of cloud compute, storage, networking, identity, and environment-provisioning concepts.Experience supporting applications in cloud-hosted or hybrid environments.Network TroubleshootingWorking knowledge of DNS, firewalls, routing, load balancers, ports, and certificates.Ability to identify basic connectivity issues between hosts, applications, and services.Familiarity with standard network diagnostic tools.Cybersecurity and ComplianceFamiliarity with STIGs, vulnerability remediation, system hardening, and compliance-driven configuration.Ability to support cybersecurity remediation activities without disrupting application functionality.Experience in federal, government-hosted, or regulated environments is highly valued.Documentation and CommunicationAbility to follow detailed technical instructions, runbooks, and change procedures.Strong attention to detail when documenting completed work, issues, deviations, and validation results.Experience working in ticketing, incident-management, or change-management systems.Clear written and verbal communication skills for status updates, handoffs, and escalations.Ability to collaborate effectively with engineering, operations, cybersecurity, networking, and client-facing teams.Preferred QualificationsExperience supporting federal clients or government-hosted environments.Knowledge of cybersecurity workflows, STIG implementation, or federal compliance requirements.Experience with Citrix technologies.Legacy infrastructure or application-support experience.Millennium or Cerner application knowledge.Experience with infrastructure-as-code or configuration-management tools.Production support, incident response, or SRE operational experience.Knowledge of monitoring, alerting, centralized logging, observability, and reliability practices.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer
Principal Site Reliability Engineer

Ll Oefentherapie • Nashville (TN)

On-site
USD 120,000 - 160,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobtailor • Arlington (VA)

On-site
USD 140,000 - 200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

ARA • Albuquerque (NM)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

SCIGON • Naperville (IL)

Hybrid
USD 110,000 - 170,000
IT CONSULTANT SR
IT CONSULTANT SR

First Horizon Corp. • Memphis (TN), Northern (KY)

Hybrid
USD 120,000 - 160,000
Mgr IT Site Reliability Eng
Mgr IT Site Reliability Eng

Kforce, Inc • Town of Florida (NY)

On-site
USD 110,000 - 140,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Brooksource • San Antonio (TX)

On-site
USD 80,000 - 120,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • United States

On-site
USD 140,000 - 180,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 190,000