Reliability & Observability Analyst I

IREN

Fort Worth (TX)

On-site

USD 65,000 - 90,000

Full time

43 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance
PTO & holidays
401(k) plan
Disability insurance
Employee assistance

Job summary

IREN is seeking an IOC Reliability & Observability Analyst I to join our Operations team in a 100% onsite role in Dallas/Fort Worth. You will analyze incident data, validate operational signals, and support AIOps enabled automation across GPU clusters and facilities.

This entry‑level position emphasizes learning, data quality, and improving detection quality under established processes. You will work with Linux environments, observe telemetry from monitoring tools, and collaborate with IOC,

Qualifications

  • Bachelor's degree in Computer Science, Data Science, Statistics, or equivalent hands‑on experience.
  • Exposure to 24/7 production environments supporting infraestrutura or data center operations.
  • Foundational awareness of SRE concepts such as service health, MTTR/MTTD, and the incident lifecycle.

Responsibilities

  • Analyze incident data, system behaviors, and operational signals across GPU clusters, networks, and facilities.
  • Identify detection gaps, alert delays, false positives, and under‑monitored systems.
  • Validate ticketing and incident data for accuracy, completeness, and reporting integrity.
  • Support continuous improvement of observability by evaluating metrics, logs, alerts, and dashboards.
  • Assist in refining operational views focused on service health, reliability, and signal quality.
  • Generate post‑incident insights highlighting trends, risks, and improvement opportunities.
  • Support AIOps‑enabled capabilities by reviewing outputs from anomaly detection, alert correlation, and event clustering.
  • Validate automated insights and elevate tuning or accuracy issues to IOC and engineering teams.
  • Assist with testing automation related to alert routing, enrichment, and suppression.
  • Produce and maintain SLA/KPI dashboards and reliability reports using established templates and data sources.

Skills

Linux basics
Monitoring & observability
Python/Bash scripting
Incident analysis
Communication skills

Education

Bachelor's degree in CS or related

Tools

Splunk
Datadog
Prometheus

Job description

Job Type: Full-Time l Location: Dallas / Fort Worth, TX l Department: Operations l Reporting to: Data Center Manager | Work Location Type: #onsite

IREN is a vertically integrated AI Cloud provider, delivering large‑scale data centers and GPU clusters for AI training and inference. IREN’s platform is underpinned by its expansive portfolio of grid‑connected land and power in renewable‑rich regions across North America, Europe and APAC.

Job Description

Job Type: Full-Time l Location: Dallas / Fort Worth, TX l Department: Operations l Reporting to: Data Center Manager | Work Location Type: #onsite

IREN is a vertically integrated AI Cloud provider, delivering large‑scale data centers and GPU clusters for AI training and inference. IREN’s platform is underpinned by its expansive portfolio of grid‑connected land and power in renewable‑rich regions across North America, Europe and APAC.

With 100% renewable energy, we build, own and operate our data centers and take pride in being at the forefront of sustainable solutions for the ever‑evolving applications of high‑performance compute. We believe that human progress is invaluable, but it should be done in the right way – responsibly, sustainably and having a positive impact on the communities we operate in.

We are seeking an IOC Reliability & Observability Analyst I with a strong reliability, observability, and automation mindset to support our 24/7 HPC Data Center Operations. The role focuses on analyzing operational signals, improving incident quality, and supporting AIOps enabled automation and tooling and is designed for candidates early in their careers who want to grow into Site Reliability, Infrastructure Operations, or Platform Engineering paths. This is an entry‑level (Level 1) IOC role focused on operational analysis, data quality, and reliability signal validation rather than system design or engineering ownership. You will support IOC, engineering, and operations teams by analyzing incidents, validating operational signals, and identifying opportunities to improve detection quality and operational reliability under established processes and guidance.

Job Requirements
  • 1-3 years of experience in IOC, NOC, SRE‑adjacent operations, systems analysis, or technical support roles
  • Bachelor's degree in Computer Science, Data Science, Statistics, or equivalent hands‑on experience
  • Exposure to 24/7 production environments supporting infrastructure, cloud, or data center operations
  • Foundational awareness of SRE concepts such as service health, MTTR/MTTD, and the incident lifecycle, with the ability to apply these concepts in operational analysis.
  • Working knowledge of Linux‑based systems, basic networking concepts, and infrastructure dependencies
  • Experience working with metrics, logs, and alerting systems across infrastructure or application environments
  • Familiarity with observability platforms (e.g., Splunk, Datadog, Prometheus‑style metrics)
  • Ability to assess alert quality, identify noise, and recognize monitoring gaps
  • Awareness of AIOps concepts such as anomaly detection, event correlation, and alert noise reduction, primarily for the purpose of reviewing and validating automated insights
  • Experience validating automated insights and supporting alerting or observability automation
  • Ability to read automation artifacts (Python, Bash, or configuration‑based workflows) and assist with minor updates under documented procedures and guidance
  • Ability to analyze incident trends and system behaviors with strong attention to data accuracy, signal integrity, and identify recurring issues or improvement opportunities
  • Clear communication skills and comfort working cross‑functionally with operations and engineering teams
Other important requirements
  • Pre‑employment screening, including background check and substance testing may be required according to company policies
Job Responsibilities
  • Analyze incident data, system behaviors, and operational signals across GPU clusters, networks, and facilities to identify risks and trends
  • Identify detection gaps, alert delays, false positives, and under‑monitored systems, and document findings for review by IOC leadership or engineering teams
  • Validate ticketing and incident data for accuracy, completeness, and reporting integrity
  • Support continuous improvement of observability by evaluating metrics, logs, alerts, and dashboards
  • Assist in refining operational views focused on service health, reliability, and signal quality
  • Generate post‑incident insights highlighting trends, risks, and improvement opportunities
  • Support AIOps‑enabled capabilities by reviewing outputs from anomaly detection, alert correlation, and event clustering, and flagging accuracy or data‑quality issues
  • Validate automated insights and elevate tuning or accuracy issues to IOC and engineering teams
  • Assist with testing automation related to alert routing, enrichment, and suppression, and submit recommended changes through established change and review processes
  • Produce and maintain SLA/KPI dashboards and reliability reports using established templates, definitions, and data sources
  • Provide data‑driven insights and recommendations to inform preventive measures, workflow improvements, and monitoring enhancements
  • Contribute to runbook updates, operational documentation, and reliability initiatives in partnership with IOC and engineering teams
  • Develop foundational SRE skills in preparation for expanded operational responsibility
  • This role operates under defined IOC processes and supervision, with increasing responsibility as skills and experience develop
Job Benefits

At IREN, we offer a comprehensive, market‑competitive total rewards package designed to support employees’ well‑being, career advancement, and financial wealth. Our offerings reflect our commitment to Proceed with Purpose while rewarding high performance and long‑term growth.

Compensation
  • Actual compensation will be determined based on factors such as experience, qualifications.
  • Overtime compensation for non‑exempt workers for hours worked over 40 per week
Health & Wellness
  • 100% company paid health insurance premiums (medical, dental, and vision) for employees, 75% company paid coverage for dependents
  • Company‑paid short‑term and long‑term disability insurance
  • Voluntary life, critical illness, and accident coverage available
  • Health Savings Accounts (HSA) – when combined with the High‑Deductible Health Plan
  • Employee Assistance Program and wellness resources
Retirement & Financial Wealth
  • 401(k) retirement plan with company match
  • Paid professional development and access to financial planning and legal services
Time Off & Leave Programs
  • Paid Time Off (PTO) and paid holidays
Growth & Development
  • Professional development to support certifications, continuing education, or role related training
Community & Culture
  • Company events and team‑building activities

We value diverse perspectives and believe that skills can be developed. If you’re passionate about this role, we want to hear from you — whether you meet every criteria or not. Your unique experiences might be exactly what we need!

IE US Operations Inc., the employingentityand proud member of the IREN group is an equal opportunity employer that is committed to creating an inclusive workplace. We are committed to evaluating qualified applicants and do not discriminate against protected characteristics under applicable legislation.

We participate in E-Verify and will provide the federal government with your Form I-9 information to confirm that you are authorized to work in the U.S. E-Verify Participation Notice .

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Center Operations Lead
Data Center Operations Lead

Iris Energy • Childress (TX)

On-site
USD 90,000 - 120,000
Health insurance
401(k) match
Paid time off
+2
Data Center Operations Lead
Data Center Operations Lead

IREN • Childress (TX)

On-site
USD 75,000 - 95,000
100% company paid health insurance premiums
401(k) retirement plan with company match
Paid Time Off (PTO) and paid holidays
Technical Program Manager
Technical Program Manager

IREN • United States

On-site
USD 135,000 - 165,000
Health insurance
401(k) with company match
Paid time off
+1
Senior Operations Manager
Senior Operations Manager

IREN • Childress (TX)

On-site
USD 180,000 - 240,000
Health insurance (100% employee) with
401(k) with company match
PTO & holidays
+1
Data Centre Security Specialist
Data Centre Security Specialist

IREN • Childress (TX)

On-site
USD 85,000 - 125,000
Medical, dental, vision insurance
401(k) with company match
Paid time off
Construction Data Analyst
Construction Data Analyst

Iris Energy • Childress (TX)

On-site
USD 90,000 - 120,000
Construction Supervisor LOTO - Mechanical
Construction Supervisor LOTO - Mechanical

Facade Today • Sweetwater (TX)

On-site
USD 62,000 - 84,000
Company-paid health insurance
PTO and holidays
Relocation assistance
Senior Operations Manager
Senior Operations Manager

IREN • Sweetwater (TX)

On-site
USD 174,000 - 210,000
Health insurance
401(k) match
Paid time off
Electrical Engineer
Electrical Engineer

Iris Energy • Fort Worth (TX)

Hybrid
USD 110,000 - 170,000
Medical, dental, and vision insurance
Company-paid life and disability保险
401(k) retirement plan with company-m\
HPC - Senior Data Center Technician (Supervisory)
HPC - Senior Data Center Technician (Supervisory)

IREN • Childress (TX)

On-site
USD 57,000 - 81,000
Health insurance
401(k) retirement plan with company-m

Paid time off
+2