Senior Site Reliability Engineer

Spectraforce Technologies

Austin (TX)

Hybrid

USD 130,000 - 170,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Spectraforce Technologies in Austin, TX is seeking a Senior Site Reliability Engineer to design and implement scalable, automated operations for complex enterprise systems. You will advance AI/ML‑driven observability and reduce toil across cloud and on‑prem environments.

This role emphasizes scripting, CI/CD automation, incident response, capacity planning, and collaboration with engineering teams. Ideal candidates have 6–8 years of enterprise IT experience, strong Linux/Windows administration,

Qualifications

  • 6-8 years of enterprise-level administration and support

Responsibilities

  • Evangelize SRE mindset and solve problems through systematization
  • Identify opportunities to build innovative tools and solve operations problems on large enterprise and mission-critical applications
  • Create scripts to automate operational tasks and incorporate solutions into infrastructure; architect and own production automation solutions
  • Design and implement AI/ML-driven automation pipelines, observability enhancements, and proactive operational response systems
  • Lead expansion of automation coverage across deployment, monitoring, alerting, and self-healing workflows for Cloud and Login Platforms
  • Collaborate with Engineering, Scrum, and Ops resources to provide technical expertise and support on key initiatives for system availability and reliability
  • Triage alerts and diagnose/resolve critical issues; manage implementation of changes with minimal risk
  • Develop tools, frameworks, and instrumentation to validate and increase rollout success for applications; leverage AI/ML capabilities to enhance operational visibility
  • Champion AIOps platform adoption and ML-assisted observability practices across the team
  • Coordinate capacity planning using data-driven trend analysis and ML-informed forecasting
  • Develop CI/CD orchestration systems to reduce friction for software delivery to production; drive adoption of GitOps and AI-assisted pipeline optimization
  • Real-time troubleshooting of mission-critical application workflows and incorporate feedback into product development
  • Participate in on-call support

Skills

Automation scripting
Monitoring dashboards
SDLC
Linux administration
Windows administration
Cloud configuration
Networking
Distributed systems
Programming: .NET/PowerShell/Java
Observability tools
AIOps/AI/ML

Education

Bachelor's degree in Computer Science

Tools

Splunk
AppDynamics
Kubernetes
GitHub Actions
Jenkins
Kafka
RabbitMQ
IBM MQ
Solace
GCP

Job description

Title: Senior Site Reliability Engineer
Duration: 06 Months
Location: Austin, TX - Hybrid 4 days weekly onsite
Our Opportunity

We are looking for a skilled engineer with disciplines that incorporate aspects of software systems engineering and operations. We are combining these skills to come up with better ways of managing and operating applications - including AI/ML-driven approaches to observability and reliability.

What you'll do
  • Evangelize SRE mindset and solve problems through systematization.
  • Identify opportunities to build innovative tools and solve unique operations problems on large enterprise and mission-critical applications.
  • Create scripts to automate operational tasks and incorporate solutions into infrastructure; architect and own production automation solutions that measurably reduce manual toil and improve operational throughput.
  • Design and implement AI/ML-driven automation pipelines, observability enhancements, and proactive operational response systems - including anomaly detection and predictive alerting to improve platform reliability.
  • Lead expansion of automation coverage across deployment, monitoring, alerting, and self-healing workflows for Cloud and Login Platforms.
  • Collaborate with Engineering, Scrum, and Ops resources to provide technical expertise and support on key initiatives for system availability and reliability.
  • Triage alerts and diagnose/resolve critical issues; manage implementation of changes with clear communication and minimal risk.
  • Develop tools, frameworks, and instrumentation to validate and increase rollout success for applications; leverage AI/ML capabilities to enhance operational visibility and rollout validation at scale.
  • Champion AIOps platform adoption and ML-assisted observability practices across the team.
  • Coordinate capacity planning using data-driven trend analysis and ML-informed forecasting.
  • Develop CI/CD orchestration systems to reduce friction for software delivery to production; drive adoption of GitOps concepts and AI-assisted pipeline optimization.
  • Real-time troubleshooting of mission-critical application workflows and incorporate feedback into product development.
  • Participate in on-call support.
Required Skills
  • 6-8 years of experience with enterprise-level administration and support.
  • 6-8 years of experience writing automation scripts, building application dashboards for proactive monitoring, and setting up alerts for early issue determination.
  • 6-8 years practicing SDLC, process improvements.
  • Hands-on enterprise systems administration, monitoring, and deployment activities.
  • Experience with Windows 2019/2022 and Linux hosted via Virtual Machine.
  • Experience in Cloud application configuration, deployment, support, and migration - GCP/PCF is a plus.
  • Knowledge of IP networking including DNS, DHCP, firewalls, IP routing, etc.
  • Familiarity with large-scale distributed systems and high-availability architecture.
  • Linux and Windows system administration, troubleshooting, and tuning.
  • Development experience in one or more programming languages: .NET, PowerShell, Java, Python, Bash.
  • Knowledge of one or more of SQL, Oracle, MongoDB databases.
  • Working knowledge of Actimize.
  • Knowledge of one or more Message Brokers: Solace, RabbitMQ, IBM MQ, Kafka.
  • Knowledge of Splunk, AppDynamics, or similar observability tools.
  • Demonstrated experience applying AI/ML or AIOps approaches (e.g., anomaly detection, predictive alerting, ML-assisted observability) in production environments.
  • Bachelor's degree in computer science or related discipline.
Helpful Skills
  • Financial services industry experience.
  • Agile methodologies.
  • Hands-on experience with AIOps platforms or ML-driven observability tooling.
  • Experience integrating AI/ML capabilities into CI/CD or operational automation workflows.
  • Familiarity with CI/CD tools (Harness, Jenkins, GitHub Actions) or GitOps concepts.
  • Exposure to container orchestration (Kubernetes, OpenShift) or cloud platforms (AWS, Azure, GCP).
Personal Skills
  • Strong customer orientation with an affinity to proactively own, communicate, and follow through on projects and issues.
  • Extreme sense of ownership to resolve problems in a distributed environment.
  • Gritty resolve to dig deeper into technical issues in a complex login ecosystem.
  • A self-starter with the ability and confidence to independently resolve issues and bring results back to the team.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

On-site
USD 120,000 - 155,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE) / Principal Site Reliability Engineer (SRE)

Mindlance • Irving (TX)

Hybrid
USD 120,000 - 160,000
SRE - Site Reliability Engineer - Senior
SRE - Site Reliability Engineer - Senior

ManpowerGroup Global, Inc. • Austin (TX)

On-site
USD 66,000 - 90,000
Senior Site Reliability Engineer – AI & Automation.
Senior Site Reliability Engineer – AI & Automation.

Veriipro • Miami (FL)

On-site
USD 130,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior AIOps and Incident Management / Site Reliability Engineering C2C jobs
Senior AIOps and Incident Management / Site Reliability Engineering C2C jobs

Tech Mirrors • Fort Mill (SC)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • Austin (TX)

On-site
USD 140,000 - 190,000
Associate Engineer, Site Reliability
Associate Engineer, Site Reliability

R&D • United States

On-site
USD 90,000 - 140,000