Infrastructure Services Engineer (Hybrid Eligible)

UT-Battelle

Oak Ridge (TN)

Hybrid

USD 90,000 - 130,000

Full time

29 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

UT-Battelle’s Oak Ridge National Laboratory seeks an Infrastructure Services Engineer focused on monitoring and observability across on‑prem, cloud, and containerized environments. You will design dashboards, alerts, and telemetry pipelines while collaborating with cross‑functional teams to enhance system health.

The role emphasizes automation, AI-assisted anomaly detection, and on‑call support within a hybrid work model at ORNL. A BS degree with 2 years of experience is required.

Qualifications

  • BS degree in information technology or related field and 2 years of relevant experience.
  • Experience operating enterprise monitoring platforms for infrastructure, applications, networks, or cloud environments.
  • Experience supporting Windows and Linux server environments, including performance analysis and troubleshooting.
  • Experience developing automated solutions using PowerShell, Python, or similar scripting tools.
  • Working knowledge of cloud infrastructure, container platforms, orchestration technologies, virtualization, and VM lifecycle operations.
  • Understanding of networking fundamentals, telemetry, and diagnostic methodologies.

Responsibilities

  • Design, implement, administer, and maintain enterprise monitoring and observability across on‑prem, cloud, and containerized environments.
  • Develop alerts, dashboards, reports, and telemetry pipelines to identify degradation early and accelerate incident triage and root-cause analysis.
  • Automate monitoring deployment, configuration, data collection, and remediation using PowerShell, Python, or other scripting and automation tools.
  • Collaborate with infrastructure, network, application, security, and other teams to improve system health and observability.
  • Provide escalated support and participate in on-call rotations as required.
  • Establish and maintain monitoring standards, topology diagrams, runbooks, and team procedures.

Skills

Monitoring platforms
PowerShell
Python
Cloud platforms
Linux administration
Windows administration
Observability concepts
Automation

Education

BS degree in information technology or related field

Tools

Prometheus
Grafana
Elastic
SolarWinds
Dynatrace
Terraform

Job description

Overview

We are seeking an Infrastructure Services Engineer who will focus on specializing in monitoring and observability. This position resides in the Infrastructure Operations Center (IOC) in the Digital Services Infrastructure & Operations division of the Information Technology Services Directorate, at Oak Ridge National Laboratory (ORNL).


Major Duties/Responsibilities


  • Design, implement, administer, and maintain enterprise monitoring and observability solutions across on-premises, cloud, and containerized environments.

  • Develop and optimize alerts, dashboards, reports, synthetic monitors, and telemetry pipelines to identify degradation early and accelerate incident triage and root-cause analysis.

  • Evaluate monitoring coverage, gaps, overlaps, and underused capabilities, and implement tools, integrations, and data sources that improve operational visibility.

  • Automate monitoring deployment, configuration, data collection, and remediation using PowerShell, Python, or other scripting and automation tools.

  • Evaluate and apply AI-assisted capabilities for anomaly detection, predictive analytics, and operational efficiency.

  • Collaborate with infrastructure, network, application, security, and other technical teams to improve system health and observability.

  • Support incident and problem management by providing relevant metrics, logs, performance trends, and historical analysis.

  • Work with vendors and internal subject matter experts to troubleshoot monitoring agents, collectors, integrations, and platform components.

  • Establish and maintain monitoring standards, topology diagrams, technical documentation, runbooks, and team procedures.

  • Support patching, backup, upgrade, and lifecycle activities for monitoring platforms and related infrastructure components.

  • Continuously improve alert thresholds, dashboards, data quality, automated remediations, and monitoring workflows to reduce noise and repetitive operational work.

  • Provide escalated support for monitoring-related issues and participate in an on-call or planned maintenance rotation as required.

  • Deliver ORNL’s mission by aligning behaviors, priorities, and interactions with our core values of Impact, Integrity, Teamwork, Safety, and Service. Promote equal opportunity by fostering a respectful workplace – in how we treat one another, work together, and measure success.


Basic Qualifications


  • BS degree in information technology or a related technical field and 2 years of relevant experience.

  • Experience operating, administering, or engineering enterprise monitoring platforms for infrastructure, applications, networks, or cloud environments.

  • Experience supporting enterprise Windows and Linux server environments, including performance analysis and troubleshooting.

  • Experience developing automated solutions using PowerShell, Python, or similar scripting tools.

  • Working knowledge of cloud infrastructure, container platforms, orchestration technologies, virtualization, and virtual-machine lifecycle operations.

  • Understanding of networking fundamentals, system performance indicators, telemetry, and diagnostic methodologies.


Preferred Qualifications


  • Strong analytical and problem-solving skills, including the ability to use operational data to identify issues and recommend improvements.

  • Strong written and verbal communication, customer service, collaboration, and technical documentation skills.

  • Ability to prioritize responsibilities and balance project work, operational support, and incident response in a fast-paced environment.

  • Demonstrated experience automating repetitive work or improving technical and operational processes.

  • Experience engineering and administering one or more enterprise-scale monitoring platforms, such as Prometheus, Grafana, Elastic, SolarWinds, or Dynatrace.

  • Experience with observability concepts and technologies, including metrics, logs, traces, baselining, synthetic monitoring, and service-level objectives.

  • Experience with anomaly detection, predictive analytics, AIOps, or automated remediation.

  • Knowledge of automation and infrastructure-as-code frameworks, such as Ansible, Terraform, or Azure Automation.

  • Experience using version-control systems to maintain scripts, configurations, dashboards, or infrastructure code.

  • Experience with virtualized or clustered compute environments, including performance tuning and lifecycle automation.

  • Familiarity with enterprise storage technologies, including direct-attached, SAN, and object storage, and their monitoring requirements.

  • Knowledge of enterprise server, storage, network hardware, and platform-level instrumentation.

  • Experience with enterprise backup, patching, configuration, or lifecycle management practices.

  • Understanding of change management and controlled operational workflows.

  • Experience working in regulated, scientific, government, or similarly complex technical environments.

  • Motivated self-starter with the ability to work independently and participate creatively in collaborative teams across the laboratory.


Special Requirements


  • Visa sponsorship: Visa sponsorship is not available for this position.

  • Security, Credentialing, and Eligibility Requirements: Q Clearance: This position requires the ability to obtain and maintain a clearance from the Department of Energy. As such, this position is a Workplace Substance Abuse (WSAP) testing designated position. WSAP positions require passing a pre-placement drug test and participation in an ongoing random drug testing program.

  • For Hybrid eligible positions: In addition, we offer a flexible work environment that supports both the organization and the employee. A hybrid/onsite working arrangement may be available with this position.


About ORNL

As a U.S. Department of Energy (DOE) Office of Science national laboratory, ORNL has an impressive 80-year legacy of addressing the nation’s most pressing challenges. Our team is made up of over 7,000 dedicated and innovative individuals! Our goal is to create an environment where a variety of perspectives and backgrounds are valued, ensuring ORNL is known as a top choice for employment. These principles are essential for supporting our broader mission to drive scientific breakthroughs and translate them into solutions for energy, environmental, and security challenges facing the nation.


ORNL offers competitive pay and benefits programs to attract and retain individuals who demonstrate exceptional work behaviors. The laboratory provides a range of employee benefits, including medical and retirement plans and flexible work hours, to support the well-being of you and your family. Employee amenities such as on-site fitness, banking, and cafeteria facilities are also available for added convenience.


Other benefits include the following: Prescription Drug Plan, Dental Plan, Vision Plan, 401(k) Retirement Plan, Contributory Pension Plan, Life Insurance, Disability Benefits, Generous Vacation and Holidays, Parental Leave, Legal Insurance with Identity Theft Protection, Employee Assistance Plan, Flexible Spending Accounts, Health Savings Accounts, Wellness Programs, Educational Assistance, Relocation Assistance, and Employee Discounts.


This position will remain open for a minimum of 5 days after which it will close when a qualified candidate is identified and/or hired.


ORNL is an equal opportunity employer.


All qualified applicants, including individuals with disabilities and protected veterans, are encouraged to apply.


UT-Battelle is an E-Verify employer.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Services Engineer (Hybrid Eligible)
Infrastructure Services Engineer (Hybrid Eligible)

Oak Ridge National Laboratory • Oak Ridge (TN)

Hybrid
USD 85,000 - 120,000
Medical and retirement plans
Flexible work hours
On-site amenities
Business Relationship Manager
Business Relationship Manager

UT-Battelle • Oak Ridge (TN)

On-site
USD 110,000 - 140,000
Medical plans
Retirement plans
Flexible work hours
+2
DevOps Build & Release Engineer (Hybrid Eligible)
DevOps Build & Release Engineer (Hybrid Eligible)

Knoxville Technology Council • Oak Ridge (TN)

On-site
USD 90,000 - 130,000
Classified HPC Systems Engineer (Hybrid Eligible)
Classified HPC Systems Engineer (Hybrid Eligible)

Oak Ridge National Laboratory • Oak Ridge (TN)

Hybrid
USD 110,000 - 150,000
401(k) Retirement Plan
Medical Plan
Relocation Assistance
+2
Research Operations Support Professional
Research Operations Support Professional

Oak Ridge National Laboratory • Oak Ridge (TN)

On-site
USD 65,000 - 95,000
Data Center Engineer
Data Center Engineer

Oak Ridge National Laboratory • Oak Ridge (TN)

On-site
USD 90,000 - 130,000
Communications Specialist (BESSD)
Communications Specialist (BESSD)

Oak Ridge National Laboratory • Oak Ridge (TN)

On-site
USD 75,000 - 95,000
Medical benefits
Retirement plans
Flexible work hours
DevOps Engineering Support Team Lead
DevOps Engineering Support Team Lead

Oak Ridge National Laboratory • Oak Ridge (TN)

On-site
USD 140,000 - 190,000
Business Relationship Manager
Business Relationship Manager

Oak Ridge National Laboratory • Oak Ridge (TN)

On-site
USD 90,000 - 150,000
Knowledge Management Team Lead
Knowledge Management Team Lead

UT-Battelle • Oak Ridge (TN)

Hybrid
USD 120,000 - 165,000
Hybrid work eligibility