Site Reliability Engineer

Skill

Austin (TX)

On-site

USD 140,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Subsidized health plan
Retirement plan with match
Paid sick leave

Job summary

Aquent Talent partners with a leading financial services institution to redefine reliability for millions of users. The Site Reliability Engineer blends software systems engineering with robust operations, applying AI/ML-driven automation, observability, and proactive response to elevate platform reliability.

You will design and deploy AI-powered pipelines, triage alerts, implement self-healing workflows, and drive GitOps-based CI/CD in a cloud environment.

Qualifications

  • Senior SRE with enterprise-scale experience.
  • Proven automation scripting and proactive monitoring.
  • Strong knowledge of cloud deployment and networking fundamentals.
  • Experience applying AI/ML-driven observability in production.

Responsibilities

  • Champion the SRE mindset and drive systematic problem solving.
  • Architect production automation to reduce manual toil.
  • Design AI/ML-driven observability and proactive alerting.
  • Lead automation across deployment, monitoring, and self-healing workloads.
  • Collaborate with Engineering, Scrum, and Operations teams to improve reliability.

Skills

SRE Leadership
AI/ML automation
Observability
Anomaly detection
Cloud platforms
CI/CD automation
GitOps
On-call support

Education

Bachelor's degree in Computer Science or related field

Tools

Splunk
AppDynamics
Kafka
Kubernetes
Jenkins
GitHub Actions

Job description

Aquent is partnering with a leading financial services institution that is revolutionizing how individuals manage their financial futures. This organization is at the forefront of integrating cutting-edge technology, including AI/ML, to enhance the reliability and performance of its critical applications. By joining this team, you will play a pivotal role in shaping the future of financial technology, driving innovation, and ensuring seamless, high-availability experiences for millions of users. Your contributions will directly impact the stability, efficiency, and intelligence of the platforms that power financial success.

Unleash Your Expertise as a Site Reliability Engineer

Are you a skilled engineer passionate about combining software systems engineering with robust operations, especially through AI/ML-driven approaches? We are seeking an innovative Site Reliability Engineer to join a dynamic team dedicated to managing and optimizing large enterprise and mission-critical applications. This is an exciting opportunity to evangelize the SRE mindset, build groundbreaking tools, and implement intelligent automation that significantly reduces manual toil and elevates operational throughput. You will be at the heart of designing and deploying advanced AI/ML-driven automation pipelines, enhancing observability, and creating proactive operational response systems that set new industry standards for reliability.

What You'll Do

  • Champion the SRE mindset and drive problem-solving through systematic approaches and innovative solutions.
  • Identify and seize opportunities to develop unique tools and resolve complex operational challenges within critical enterprise applications.
  • Architect and implement production automation solutions that measurably reduce manual effort and boost operational efficiency.
  • Design and deploy AI/ML-driven automation pipelines, observability enhancements, and proactive operational response systems, including anomaly detection and predictive alerting, to elevate platform reliability.
  • Lead the expansion of automation coverage across deployment, monitoring, alerting, and self-healing workflows for Cloud and Login Platforms.
  • Collaborate extensively with Engineering, Scrum, and Operations teams, providing crucial technical expertise and support for key initiatives focused on system availability and reliability.
  • Efficiently triage alerts, diagnose, and resolve critical issues, managing change implementations with clear communication and minimal risk.
  • Develop essential tools, frameworks, and instrumentation to validate and enhance the success of application rollouts, leveraging AI/ML capabilities for operational visibility and validation at scale.
  • Advocate for AIOps platform adoption and ML-assisted observability practices across the team.
  • Coordinate robust capacity planning through data-driven trend analysis and ML-informed forecasting.
  • Develop CI/CD orchestration systems to streamline software delivery to production, championing GitOps concepts and AI-assisted pipeline optimization.
  • Perform real-time troubleshooting of mission-critical application workflows, integrating feedback directly into product development cycles.
  • Participate actively in on-call support rotations, ensuring continuous operational excellence.

What You'll Bring

  • 6-8 years of hands-on experience in enterprise-level administration and support.
  • 6-8 years of experience crafting automation scripts, developing application dashboards for proactive monitoring, and configuring alerts for early issue detection.
  • 6-8 years of practical experience with SDLC and process improvement methodologies.
  • Extensive hands-on experience with enterprise systems administration, monitoring, and deployment activities.
  • Proficiency with Windows 2019/2022 and Linux operating systems hosted via Virtual Machine.
  • Experience in Cloud application configuration, deployment, support, and migration.
  • Solid understanding of IP networking fundamentals, including DNS, DHCP, firewalls, and IP routing.
  • Familiarity with large-scale distributed systems and high-availability architectures.
  • Expertise in Linux and Windows system administration, troubleshooting, and performance tuning.
  • Development experience in one or more programming languages such as .NET, PowerShell, Java, Python, or Bash.
  • Knowledge of one or more database systems, including SQL, Oracle, or MongoDB.
  • Working knowledge of Actimize.
  • Familiarity with one or more Message Brokers, such as Solace, RabbitMQ, IBM MQ, or Kafka.
  • Experience with Splunk, AppDynamics, or similar observability tools.
  • Demonstrated experience applying AI/ML or AIOps approaches (e.g., anomaly detection, predictive alerting, ML-assisted observability) in production environments.
  • A Bachelor's degree in computer science or a related discipline.
  • Strong customer orientation with an affinity for proactively owning, communicating, and following through on projects and issues.
  • An extreme sense of ownership to meticulously resolve problems within a distributed environment.
  • A gritty resolve to delve deep into technical issues within a complex login ecosystem.
  • A self-starter mentality with the ability and confidence to independently resolve issues and deliver results to the team.

Bonus Points

  • Experience within the financial services industry.
  • Proficiency with Agile methodologies.
  • Hands-on experience with AIOps platforms or ML-driven observability tooling.
  • Experience integrating AI/ML capabilities into CI/CD or operational automation workflows.
  • Familiarity with CI/CD tools (e.g., Harness, Jenkins, GitHub Actions) or GitOps concepts.
  • Exposure to container orchestration platforms (e.g., Kubernetes, OpenShift) or cloud platforms (e.g., AWS, Azure, GCP).

About Aquent Talent:

Aquent Talent connects the best talent in marketing, creative, and design with the world's biggest brands.

Our eligible talent get access to amazing benefits like subsidized health, vision, and dental plans, paid sick leave, and retirement plans with a match.

Aquent is an equal-opportunity employer. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, and other legally protected characteristics. We're about creating an inclusive environment-one where different backgrounds, experiences, and perspectives are valued, and everyone can contribute, grow their careers, and thrive.

Helping all candidates find great careers is our goal. The information you provide here is secure and confidential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Aquent • Austin (TX)

On-site
USD 140,000 - 175,000
Subsidized health plan
Vision plan
Dental plan
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Spectraforce Technologies • Austin (TX)

Hybrid
USD 130,000 - 170,000
Java Developer
Java Developer

Skill • Southlake (TX)

On-site
USD 83,000 - 91,000
Health benefits
Vision benefits
Dental benefits
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hard Rock Digital • United States

Hybrid
USD 150,000 - 210,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
AI-Driven SRE Engineer for Cloud & Automation
AI-Driven SRE Engineer for Cloud & Automation

Skill • Austin (TX)

On-site
USD 140,000 - 190,000
Subsidized health plan
Retirement plan with match
Paid sick leave
.net Developer
.net Developer

Aquent • Austin (TX)

On-site
USD 83,000 - 92,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

Inclusion Services S.A • Chicago (IL)

On-site
USD 90,000 - 130,000
100% company-covered health insurance
401k plan with 4% match
15 days paid time off
+3