Site Reliability Engineer [AQ-18637]

Aquent

Southlake (TX)

On-site

USD 120,000 - 170,000

Full time

10 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

subsidized health
vision
dental plans
paid sick leave
retirement plans with a match
free online training through AquentGym

Job summary

Aquent is assisting a leading financial services innovator in seeking a Senior Reliability Engineer focused on automation and platform excellence. You will enhance production systems, drive automation, and ensure reliable operations across on-prem and cloud environments.

The role emphasizes building scalable automation, improving incident response, and advancing observability using Splunk, Grafana, Prometheus, and related tools. Candidates should have strong Python and SRE experience.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field or equivalent experience.

Responsibilities

  • Develop Python-based automation solutions to reduce manual operational effort.
  • Automate infrastructure management across Linux, Windows, Kubernetes, cloud platforms, and cloud-native environments.
  • Integrate tools via APIs and client libraries to streamline workflows.
  • Support CI/CD automation and deployment reliability initiatives.
  • Monitor and maintain production systems to meet reliability goals.
  • Participate in incident response, troubleshooting, and RCA activities.
  • Develop proactive automation and operational improvements to reduce recurring issues.
  • Support disaster recovery and failover testing to ensure business continuity.
  • Perform in-depth performance analysis and health reviews to optimize systems.
  • Build dashboards and monitoring with Splunk, Grafana, Prometheus, or similar.
  • Improve visibility with metrics, logs, and traces for proactive detection.
  • Investigate alerts to reduce noise and improve detection accuracy.
  • Explore AI/ML-driven operational improvements such as anomaly detection and intelligent alerting.
  • Assist in AI-enabled automation solutions for efficiency.
  • Evaluate emerging observability technologies to keep systems at the forefront.

Skills

Python
Kubernetes
GCP
AWS
Azure
Splunk
Grafana
Prometheus
Datadog
Linux
Incident management
RCA
CI/CD
OpenTelemetry
Terraform
Ansible
AI/ML (AIOps)

Education

Bachelor's degree or equivalent

Tools

Terraform
Ansible
OpenTelemetry
CI/CD
Datadog
Splunk
Grafana
Prometheus

Job description

Join a leading innovator in the financial services industry, a company dedicated to empowering individuals and institutions to achieve their financial goals. This organization stands at the forefront of technological advancement, continuously evolving its platforms and services to deliver unparalleled value and security to millions of clients. Aquent is proud to partner with this client, offering you a unique opportunity to contribute to their mission-critical operations and drive significant impact. Are you a highly motivated and experienced reliability professional passionate about automation and operational excellence? This is your chance to join a pivotal team where you will be instrumental in shaping the future of our client’s infrastructure. In this dynamic role, you will elevate the reliability and operational excellence of critical production systems, driving automation and ensuring seamless operations across both on-premises and cutting-edge cloud platforms. Your expertise will directly impact the stability, performance, and scalability of systems that serve millions, making a tangible difference in the daily lives of clients. If you thrive on solving complex challenges, possess an automation-first mindset, and are eager to contribute to a high-impact environment, we invite you to shine with us!

What You’ll Do
  • Develop Python-based automation solutions to reduce manual operational effort and enhance efficiency across diverse environments.
  • Automate infrastructure management across Linux, Windows, Kubernetes, cloud platforms, and cloud-native environments.
  • Integrate various tools and platforms through APIs and client libraries to streamline workflows and improve system interoperability.
  • Assist in implementing robust infrastructure automation using industry-standard technologies to build scalable and resilient systems.
  • Support CI/CD automation and deployment reliability initiatives to ensure smooth, consistent, and secure software releases.
  • Monitor and maintain production systems to consistently meet and exceed reliability and availability objectives, ensuring seamless service delivery.
  • Actively participate in incident response, thorough troubleshooting, and root cause analysis activities to prevent recurrence and enhance system stability.
  • Develop proactive automation and operational improvements to eliminate recurring issues and continuously improve system resilience.
  • Support disaster recovery, failover testing, and other operational readiness activities to ensure business continuity and minimize downtime.
  • Perform in-depth performance analysis and regular system health reviews to optimize system behavior and identify areas for improvement.
  • Build and maintain comprehensive dashboards, alerts, and monitoring solutions using Splunk, Grafana, Prometheus, or similar industry-leading tools.
  • Improve visibility into application and infrastructure health through enhanced metrics, logs, and traces for proactive issue detection.
  • Investigate alerts thoroughly and identify opportunities to reduce noise and improve detection accuracy, enhancing operational responsiveness.
  • Explore AI/ML-driven operational improvements such as anomaly detection, intelligent alerting, and log analytics to advance operational capabilities.
  • Assist in developing automation solutions that leverage AI to significantly improve operational efficiency and predictive insights.
  • Participate in evaluating emerging AIOps capabilities and cutting-edge observability technologies to keep our systems at the forefront of innovation.
What You’ll Bring
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 3 to 5 years of hands-on experience in Site Reliability Engineering, Production Engineering, DevOps, Systems Engineering, or Platform Engineering, with a strong focus on production operations.
  • Must have supported production systems at scale, demonstrating a deep understanding of operational challenges in large-scale environments.
  • Proven experience in operations ownership, incident response, and reliability engineering.
  • Strong programming skills in Python for automation and tooling development; Python must be a primary skill, not just basic scripting.
  • Demonstrated experience building automation tools, scripts, frameworks, or operational solutions, with the ability to provide examples of personally developed automation.
  • Experience supporting Kubernetes and cloud platforms (GCP, AWS, or Azure).
  • Familiarity with infrastructure automation and configuration management tools.
  • Experience with monitoring and observability platforms such as Splunk, Grafana, Prometheus, Datadog, or similar.
  • Solid understanding of Linux systems, networking, and distributed applications.
  • Strong analytical, troubleshooting, and problem-solving skills, with experience in critical production incident management and Root Cause Analysis (RCA).
  • Ability to work effectively in fast-paced, mission-critical environments, consistently reducing operational toil through automation.
What Will Set You Apart
  • Experience with Terraform, Ansible, or other Infrastructure as Code solutions.
  • Exposure to OpenTelemetry and modern observability practices.
  • Experience with CI/CD pipelines and deployment automation.
  • Knowledge of AI/ML, AIOps, or intelligent operational tooling.
  • Experience supporting highly available production systems in regulated or enterprise environments.
About Aquent Talent

Aquent Talent connects the best talent in marketing, creative, and design with the world’s biggest brands.

  • subsidized health
  • vision
  • dental plans
  • paid sick leave
  • retirement plans with a match
  • free online training through Aquent Gymnasium

Aquent is an equal-opportunity employer. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, and other legally protected characteristics. We’re about creating an inclusive environment-one where different backgrounds, experiences, and perspectives are valued, and everyone can contribute, grow their careers, and thrive.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Aquent • Southlake (TX)

On-site
USD 120,000 - 180,000
Subsidized health, vision, and dental
Paid sick leave
Retirement plan with company match
Site Reliability Engineer
Site Reliability Engineer

Skill • Southlake (TX)

On-site
USD 66,000 - 73,000
Health insurance
Vision insurance
Dental insurance
+2
Site Reliability Engineer
Site Reliability Engineer

Aquent • Austin (TX)

On-site
USD 140,000 - 175,000
Subsidized health plan
Vision plan
Dental plan
+2
Site Reliability Engineer [AQ-18653]
Site Reliability Engineer [AQ-18653]

Aquent • Austin (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Skill • Austin (TX)

On-site
USD 140,000 - 190,000
Subsidized health plan
Retirement plan with match
Paid sick leave
DevOps Engineer [AQ-18200]
DevOps Engineer [AQ-18200]

Aquent • Phoenix (AZ)

On-site
USD 110,000 - 140,000
Subsidized health, vision, and dental
Retirement with matching
Free online training
+1
Site Reliability Engineer
Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

On-site
USD 120,000 - 155,000
Site Reliability Engineer - Automation & Cloud Reliability
Site Reliability Engineer - Automation & Cloud Reliability

Aquent • Southlake (TX)

On-site
USD 120,000 - 180,000
Subsidized health, vision, and dental
Paid sick leave
Retirement plan with company match
.net Developer
.net Developer

Skill • Austin (TX)

On-site
USD 84,000 - 93,000
Subsidized health plan
Vision plan
Dental plan
+3
.net Developer
.net Developer

Skill • Southlake (TX)

On-site
USD 127,000 - 141,000
Subsidized health plans
Vision and dental plans
Paid sick leave
+2