Site Reliability Engineer

Aquent

Southlake (TX)

On-site

USD 120,000 - 180,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Subsidized health, vision, and dental
Paid sick leave
Retirement plan with company match

Job summary

Aquent Talent is seeking a Site Reliability Engineer to join our client’s team, focusing on reliability of mission-critical systems across on-prem and cloud environments. You will drive automation, incident response, and scalable operations in a fast-paced setting.

You will design and implement Python-based tools, Kubernetes automation, and cloud-based strategies, contributing to continuous delivery and improved system resilience for millions of users.

Qualifications

  • Bachelor’s degree or equivalent practical experience.
  • 3–5 years in SRE/DevOps with production-scale systems.
  • Strong Python automation skills required.
  • Experience with Kubernetes and cloud platforms (GCP/AWS/Azure).

Responsibilities

  • Develop Python automation to reduce manual operational effort.
  • Automate infrastructure across Linux, Windows, Kubernetes, and clouds.
  • Integrate tools via APIs and client libraries to streamline workflows.
  • Assist in implementing infrastructure automation using standard technologies.
  • Support CI/CD automation and deployment reliability initiatives.
  • Monitor and maintain production systems to meet reliability goals.
  • Engage in incident response, RCA, and troubleshooting to prevent recurrence.
  • Develop proactive automation and operational improvements.
  • Support disaster recovery, failover testing, and DR readiness.
  • Perform performance analysis and health reviews to optimize systems.
  • Build dashboards and monitoring with Splunk, Grafana, and Prometheus.
  • Improve visibility with metrics, logs, and traces; reduce noise in alerts.
  • Explore AI/ML-driven observations and anomaly detection for ops.
  • Develop automation that leverages AI for efficiency gains.
  • Evaluate emerging AIOps and observability technologies.

Skills

Python automation
Kubernetes
Cloud platforms
CI/CD
Incident response
Root Cause Analysis
Monitoring/observability
Linux
IaC (Terraform/Ansible)
OpenTelemetry

Education

Bachelor's degree in Computer Science or related

Tools

Splunk
Grafana
Prometheus
Datadog
Terraform
Ansible
OpenTelemetry

Job description

The hiring company is a leader in the financial services industry, dedicated to empowering individuals and institutions to achieve their financial goals. They are at the forefront of innovation, constantly evolving their platforms and services to provide unparalleled value and security to their clients. This is an opportunity to join a dynamic team that is passionate about leveraging technology to drive business success and enhance the client experience.

We are seeking a highly motivated and experienced Site Reliability Engineer to join a pivotal team, partnering with our client to elevate the reliability and operational excellence of critical production systems. In this dynamic contract role, you will be instrumental in shaping the future of our infrastructure, driving automation, and ensuring seamless operations across both on-premises and cloud platforms. Your expertise will directly impact the stability, performance, and scalability of systems that serve millions, making a tangible difference in the daily lives of our clients. If you thrive on solving complex challenges, have an automation-first mindset, and are eager to contribute to a high-impact environment, this is your chance to shine.

Key Responsibilities
  • Develop Python-based automation solutions to reduce manual operational effort and enhance efficiency.
  • Automate infrastructure management across Linux, Windows, Kubernetes, cloud platforms, and cloud-native environments.
  • Integrate various tools and platforms through APIs and client libraries to streamline workflows.
  • Assist in implementing robust infrastructure automation using industry-standard technologies.
  • Support CI/CD automation and deployment reliability initiatives to ensure smooth and consistent releases.
  • Monitor and maintain production systems to consistently meet and exceed reliability and availability objectives.
  • Actively participate in incident response, thorough troubleshooting, and root cause analysis activities to prevent recurrence.
  • Develop proactive automation and operational improvements to eliminate recurring issues and improve system resilience.
  • Support disaster recovery, failover testing, and other operational readiness activities to ensure business continuity.
  • Perform in-depth performance analysis and regular system health reviews to optimize system behavior.
  • Build and maintain comprehensive dashboards, alerts, and monitoring solutions using Splunk, Grafana, Prometheus, or similar tools.
  • Improve visibility into application and infrastructure health through enhanced metrics, logs, and traces.
  • Investigate alerts thoroughly and identify opportunities to reduce noise and improve detection accuracy.
  • Explore AI/ML-driven operational improvements such as anomaly detection, intelligent alerting, and log analytics.
  • Assist in developing automation solutions that leverage AI to significantly improve operational efficiency.
  • Participate in evaluating emerging AIOps capabilities and cutting-edge observability technologies.
Required Qualifications
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 3 to 5 years of hands-on experience in Site Reliability Engineering, Production Engineering, DevOps, Systems Engineering, or Platform Engineering, with a strong focus on production operations.
  • Must have supported production systems at scale, demonstrating a deep understanding of operational challenges in large environments.
  • Operations ownership, incident response, and reliability engineering experience are essential.
  • Strong programming skills in Python for automation and tooling development; Python must be a primary skill, not just basic scripting.
  • Demonstrated experience building automation tools, scripts, frameworks, or operational solutions, with the ability to provide examples of personally developed automation.
  • Experience supporting Kubernetes and cloud platforms (GCP, AWS, or Azure).
  • Familiarity with infrastructure automation and configuration management tools.
  • Experience with monitoring and observability platforms such as Splunk, Grafana, Prometheus, Datadog, or similar.
  • Understanding of Linux systems, networking, and distributed applications.
  • Strong analytical, troubleshooting, and problem-solving skills, with experience in critical production incident management and Root Cause Analysis (RCA).
  • Ability to work effectively in fast-paced, mission-critical environments, reducing operational toil through automation.
Preferred Qualifications
  • Experience with Terraform, Ansible, or other Infrastructure as Code solutions.
  • Exposure to OpenTelemetry and modern observability practices.
  • Experience with CI/CD pipelines and deployment automation.
  • Knowledge of AI/ML, AIOps, or intelligent operational tooling.
  • Experience supporting highly available production systems in regulated or enterprise environments.
About Aquent Talent
  • Our eligible talent get access to amazing benefits like subsidized health, vision, and dental plans, paid sick leave, and retirement plans with a match.

Aquent Talent connects the best talent in marketing, creative, and design with the world’s biggest brands.

Aquent is an equal-opportunity employer.

We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, and other legally protected characteristics.

We’re about creating an inclusive environment-one where different backgrounds, experiences, and perspectives are valued, and everyone can contribute, grow their careers, and thrive.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Skill • Southlake (TX)

On-site
USD 66,000 - 73,000
Health insurance
Vision insurance
Dental insurance
+2
Site Reliability Engineer
Site Reliability Engineer

Aquent • Austin (TX)

On-site
USD 140,000 - 175,000
Subsidized health plan
Vision plan
Dental plan
+2
Site Reliability Engineer
Site Reliability Engineer

Skill • Austin (TX)

On-site
USD 140,000 - 190,000
Subsidized health plan
Retirement plan with match
Paid sick leave
Site Reliability Engineer
Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

On-site
USD 120,000 - 155,000
Site Reliability Engineer - Automation & Cloud Reliability
Site Reliability Engineer - Automation & Cloud Reliability

Aquent • Southlake (TX)

On-site
USD 120,000 - 180,000
Subsidized health, vision, and dental
Paid sick leave
Retirement plan with company match
Contract Site Reliability Engineer - Automation & Observability
Contract Site Reliability Engineer - Automation & Observability

Skill • Southlake (TX)

On-site
USD 66,000 - 73,000
Health insurance
Vision insurance
Dental insurance
+2
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
System Engineer (212375)
System Engineer (212375)

Aquent • Redmond (WA)

On-site
USD 140,000 - 190,000
Health benefits
Vision plan
Dental plan
+2
Site Reliability Engineer
Site Reliability Engineer

Inclusion Services S.A • Chicago (IL)

On-site
USD 90,000 - 130,000
100% company-covered health insurance
401k plan with 4% match
15 days paid time off
+3
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

MeridianLink • United States

Remote
USD 140,000 - 190,000