Site Reliability Engineer

Aquent

Austin (TX)

On-site

USD 140,000 - 175,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Subsidized health plan
Vision plan
Dental plan
Paid sick leave
Retirement plan with match

Job summary

Aquent partners with a leading financial services institution to advance reliability and performance of critical applications in a high-availability environment. You will design and deploy AI/ML–driven automation, improve observability, and drive proactive operational responses for millions of users.

You will champion the SRE mindset, build tools, and extend automation across deployment, monitoring, and self-healing workflows while collaborating with cross-functional teams on complex login

Qualifications

  • Bachelor's degree in computer science or related field.
  • Hands-on enterprise systems administration experience.
  • Experience with AI/ML or AIOps approaches in production.
  • Strong automation, tooling, and observability skills.

Responsibilities

  • Champion the SRE mindset and drive systematic problem solving.
  • Architect and implement production automation to reduce manual toil.
  • Design AI/ML-driven observability and proactive response systems.
  • Expand automation across deployment, monitoring, and self-healing workflows.
  • Collaborate with Engineering, Scrum, and Operations on high availability.

Skills

Automation scripting
AI/ML in production
On-call incident response
Problem solving
Customer orientation

Education

Bachelor's degree in computer science

Tools

PowerShell
Python
Java
SQL
Splunk
AppDynamics
Kafka
Kubernetes
GitHub Actions

Job description

Aquent is partnering with a leading financial services institution that is revolutionizing how individuals manage their financial futures. This organization is at the forefront of integrating cutting-edge technology, including AI/ML, to enhance the reliability and performance of its critical applications. By joining this team, you will play a pivotal role in shaping the future of financial technology, driving innovation, and ensuring seamless, high-availability experiences for millions of users. Your contributions will directly impact the stability, efficiency, and intelligence of the platforms that power financial success.

Unleash Your Expertise as a Site Reliability Engineer

Are you a skilled engineer passionate about combining software systems engineering with robust operations, especially through AI/ML-driven approaches? We are seeking an innovative Site Reliability Engineer to join a dynamic team dedicated to managing and optimizing large enterprise and mission-critical applications. This is an exciting opportunity to evangelize the SRE mindset, build groundbreaking tools, and implement intelligent automation that significantly reduces manual toil and elevates operational throughput. You will be at the heart of designing and deploying advanced AI/ML-driven automation pipelines, enhancing observability, and creating proactive operational response systems that set new industry standards for reliability.

What You’ll Do
  • Champion the SRE mindset and drive problem-solving through systematic approaches and innovative solutions.
  • Identify and seize opportunities to develop unique tools and resolve complex operational challenges within critical enterprise applications.
  • Architect and implement production automation solutions that measurably reduce manual effort and boost operational efficiency.
  • Design and deploy AI/ML-driven automation pipelines, observability enhancements, and proactive operational response systems, including anomaly detection and predictive alerting, to elevate platform reliability.
  • Lead the expansion of automation coverage across deployment, monitoring, alerting, and self-healing workflows for Cloud and Login Platforms.
  • Collaborate extensively with Engineering, Scrum, and Operations teams, providing crucial technical expertise and support for key initiatives focused on system availability and reliability.
  • Efficiently triage alerts, diagnose, and resolve critical issues, managing change implementations with clear communication and minimal risk.
  • Develop essential tools, frameworks, and instrumentation to validate and enhance the success of application rollouts, leveraging AI/ML capabilities for operational visibility and validation at scale.
  • Advocate for AIOps platform adoption and ML-assisted observability practices across the team.
  • Coordinate robust capacity planning through data-driven trend analysis and ML-informed forecasting.
  • Develop CI/CD orchestration systems to streamline software delivery to production, championing GitOps concepts and AI-assisted pipeline optimization.
  • Perform real-time troubleshooting of mission-critical application workflows, integrating feedback directly into product development cycles.
  • Participate actively in on-call support rotations, ensuring continuous operational excellence.
What You’ll Bring
  • 6-8 years of hands‑on experience in enterprise-level administration and support.
  • 6-8 years of experience crafting automation scripts, developing application dashboards for proactive monitoring, and configuring alerts for early issue detection.
  • 6-8 years of practical experience with SDLC and process improvement methodologies.
  • Extensive hands‑on experience with enterprise systems administration, monitoring, and deployment activities.
  • Proficiency with Windows 2019/2022 and Linux operating systems hosted via Virtual Machine.
  • Experience in Cloud application configuration, deployment, support, and migration.
  • Solid understanding of IP networking fundamentals, including DNS, DHCP, firewalls, and IP routing.
  • Familiarity with large‑scale distributed systems and high‑availability architectures.
  • Expertise in Linux and Windows system administration, troubleshooting, and performance tuning.
  • Development experience in one or more programming languages such as .NET, PowerShell, Java, Python, or Bash.
  • Knowledge of one or more database systems, including SQL, Oracle, or MongoDB.
  • Working knowledge of Actimize.
  • Familiarity with one or more Message Brokers, such as Solace, RabbitMQ, IBM MQ, or Kafka.
  • Experience with Splunk, AppDynamics, or similar observability tools.
  • Demonstrated experience applying AI/ML or AIOps approaches (e.g., anomaly detection, predictive alerting, ML‑assisted observability) in production environments.
  • A Bachelor’s degree in computer science or a related discipline.
  • Strong customer orientation with an affinity for proactively owning, communicating, and following through on projects and issues.
  • An extreme sense of ownership to meticulously resolve problems within a distributed environment.
  • A gritty resolve to delve deep into technical issues within a complex login ecosystem.
  • A self‑starter mentality with the ability and confidence to independently resolve issues and deliver results to the team.
Bonus Points
  • Experience within the financial services industry.
  • Proficiency with Agile methodologies.
  • Hands‑on experience with AIOps platforms or ML-driven observability tooling.
  • Experience integrating AI/ML capabilities into CI/CD or operational automation workflows.
  • Familiarity with CI/CD tools (e.g., Harness, Jenkins, GitHub Actions) or GitOps concepts.
  • Exposure to container orchestration platforms (e.g., Kubernetes, OpenShift) or cloud platforms (e.g., AWS, Azure, GCP).
About Aquent Talent

Aquent Talent connects the best talent in marketing, creative, and design with the world’s biggest brands.

Our eligible talent get access to amazing benefits like subsidized health, vision, and dental plans, paid sick leave, and retirement plans with a match.

Aquent is an equal‑opportunity employer. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, and other legally protected characteristics. We’re about creating an inclusive environment‑one where different backgrounds, experiences, and perspectives are valued, and everyone can contribute, grow their careers, and thrive.

#LI-LP1

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Skill • Austin (TX)

On-site
USD 140,000 - 190,000
Subsidized health plan
Retirement plan with match
Paid sick leave
Java Developer
Java Developer

Skill • Southlake (TX)

On-site
USD 83,000 - 91,000
Health benefits
Vision benefits
Dental benefits
+2
.net Developer
.net Developer

Aquent • Austin (TX)

On-site
USD 83,000 - 92,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

Hybrid
USD 130,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hard Rock Digital • United States

Hybrid
USD 150,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Systems Engineer
Systems Engineer

Skill • Austin (TX)

On-site
USD 120,000 - 180,000
Subsidized health, vision, and dental
Paid sick leave
Retirement plan with match
AI Software Test Engineer (SDET) [AQ-13260]
AI Software Test Engineer (SDET) [AQ-13260]

Aquent Talent • Ann Arbor (MI)

On-site
USD 120,000 - 150,000
Health insurance
Vision plan
Dental plan
+3
AI Software Test Engineer (SDET)
AI Software Test Engineer (SDET)

Aquent • Ann Arbor (MI)

On-site
USD 110,000 - 150,000
Subsidized health plan
Vision plan
Dental plan
+1