Site Reliability Engineer

Spectraforce Technologies

Austin (TX)

On-site

USD 120,000 - 155,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Spectraforce Technologies seeks a Site Reliability Engineer (Contractor) in the Austin area for a 12‑month onsite engagement, 4 days per week. Strong automation, cloud infra, and production operations experience are required.

You will develop Python automation, automate across Linux, Windows, Kubernetes, and GCP, and enhance CI/CD reliability while boosting observability with Splunk, Grafana, and Prometheus. Apply advanced RCA and AIOps concepts in a fast-paced environment.

Qualifications

  • Bachelor's degree in CS, engineering, or related field (or equivalent).
  • 3–5 years of SRE/DevOps or platform engineering experience.
  • Strong Python for automation and tooling development.
  • Experience with Kubernetes and cloud platforms (GCP, AWS, or Azure).
  • Familiarity with monitoring/observability tools (Splunk, Grafana, Prometheus, Datadog).
  • Linux, networking, distributed apps, and problem-solving skills.

Responsibilities

  • Develop Python automation to reduce manual operational effort.
  • Automate infrastructure management across Linux, Windows, Kubernetes, GCP, and cloud-native environments.
  • Integrate tools via APIs and client libraries; assist infrastructure automation (Terraform/Ansible).
  • Support CI/CD automation and deployment reliability initiatives.
  • Monitor production systems and participate in incident response and RCA.
  • Build and maintain observability using Splunk, Grafana, Prometheus, and related tools.
  • Explore AI/ML-driven improvements and emerging AIOps capabilities.

Skills

Python programming
Kubernetes experience
Linux fundamentals
Analytical thinking
Troubleshooting

Education

Bachelor's degree in Computer Science/Engineering or related field

Tools

Splunk
Grafana
Prometheus
Datadog
Terraform
Ansible
CI/CD tooling

Job description

Role: Site Reliability Engineer

Location: Southlake / Austin, TX - Onsite 4 days weekly

Duration: 12 Months

Job Summary

We are seeking a motivated Site Reliability Engineer (Contractor) with 3 to 5 years of experience in automation, cloud infrastructure, and production operations. The ideal candidate has a strong automation-first mindset and enjoys solving complex operational challenges. This role focuses on improving reliability, observability, operational efficiency, and automation across on-premises and cloud platforms.

Key Responsibilities
Automation & Platform Engineering
  • Develop Python-based automation solutions to reduce manual operational effort.
  • Automate infrastructure management across Linux, Windows, Kubernetes, GCP, and cloud-native environments.
  • Integrate tools and platforms through APIs and client libraries.
  • Assist in implementing infrastructure automation using Ansible, Terraform, or similar technologies.
  • Support CI/CD automation and deployment reliability initiatives.
Reliability & Operations
  • Monitor and maintain production systems to meet reliability and availability objectives.
  • Participate in incident response, troubleshooting, and root cause analysis activities.
  • Develop automation and operational improvements to prevent recurring issues.
  • Support disaster recovery, failover testing, and operational readiness activities.
  • Perform performance analysis and system health reviews.
Observability & Monitoring
  • Build and maintain dashboards, alerts, and monitoring solutions using Splunk, Grafana, Prometheus, GCP Operations Suite, or similar tools.
  • Improve visibility into application and infrastructure health through metrics, logs, and traces.
  • Investigate alerts and identify opportunities to reduce noise and improve detection.
AIOps & Intelligence
  • Explore AI/ML-driven operational improvements such as anomaly detection, intelligent alerting, and log analytics.
  • Assist in developing automation solutions that leverage AI to improve operational efficiency.
  • Participate in evaluating emerging AIOps capabilities and observability technologies.
Required Qualifications
  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent experience.
  • 3 to 5 years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or Platform Engineering.
  • Strong programming skills in Python for automation and tooling development.
  • Experience supporting Kubernetes and cloud platforms (GCP, AWS, or Azure).
  • Familiarity with infrastructure automation and configuration management tools.
  • Experience with monitoring and observability platforms such as Splunk, Grafana, Prometheus, Datadog, or similar.
  • Understanding of Linux systems, networking, and distributed applications.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Ability to work effectively in fast-paced, mission-critical environments.
Preferred Qualifications
  • Experience with Terraform, Ansible, or Infrastructure as Code solutions.
  • Exposure to OpenTelemetry and modern observability practices.
  • Experience with CI/CD pipelines and deployment automation.
  • Knowledge of AI/ML, AIOps, or intelligent operational tooling.
  • Experience supporting highly available production systems in regulated or enterprise environments.
Must Have:
  • 3-5 years of hands‑on Site Reliability Engineering or Production Engineering experience
  • Must have supported production systems at scale.
  • Operations ownership, incident response, and reliability engineering experience required.
  • Pure DevOps, build/release, or CI/CD-only backgrounds are not a fit.
  • Strong Python and operation automation Development Skills
  • Demonstrated experience building automation tools, scripts, frameworks, or operational solutions.
  • Candidate should be able to provide examples of automation they personally developed.
  • Python must be a primary skill, not just basic scripting.
  • Production Operations & Incident Management Experience
  • Experience troubleshooting critical production incidents.
  • Root Cause Analysis (RCA) participation and problem remediation.
  • Experience reducing operational toil through automation.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

Hybrid
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

TechDigital Group • Houston (TX), Juno Beach (FL)

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

Hybrid
USD 150,000 - 155,000
Senior AIOps and Incident Management / Site Reliability Engineering C2C jobs
Senior AIOps and Incident Management / Site Reliability Engineering C2C jobs

Tech Mirrors • Fort Mill (SC)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer – AI & Automation.
Senior Site Reliability Engineer – AI & Automation.

Veriipro • Miami (FL)

On-site
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

System One • Dallas (TX)

On-site
USD 130,000 - 170,000
Systems Analyst 3 529601671
Systems Analyst 3 529601671

LMG Technology Services LLC • Austin (TX)

Hybrid
USD 120,000 - 160,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Jobtailor • Arizona

On-site
USD 180,000 - 240,000