Step into a pivotal role where your expertise will directly shape the future of financial technology, driving innovation and ensuring seamless, high-availability experiences for millions of users. Your contributions will make a tangible difference in the stability, efficiency, and intelligence of platforms that power financial success, revolutionizing how individuals manage their financial futures.
Aquent is proud to partner with a pioneering financial services institution dedicated to empowering individuals to take control of their financial destinies. This organization is at the forefront of integrating cutting-edge technology, including AI/ML, to enhance the reliability and performance of its critical applications. They are deeply committed to leveraging advanced technology to build robust, high-performing platforms that deliver exceptional user experiences, setting new benchmarks for reliability and efficiency in the industry.
Are you a skilled engineer passionate about combining software systems engineering with robust operations, especially through AI/ML-driven approaches? We are seeking an innovative professional to join a dynamic team dedicated to managing and optimizing large enterprise and mission-critical applications. This is an exciting opportunity to evangelize the Site Reliability Engineering (SRE) mindset, build groundbreaking tools, and implement intelligent automation that significantly reduces manual toil and elevates operational throughput. You will be at the heart of designing and deploying advanced AI/ML-driven automation pipelines, enhancing observability, and creating proactive operational response systems that set new industry standards for reliability.
What You’ll Do
- Champion the SRE mindset and drive problem-solving through systematic approaches and innovative solutions.
- Identify and seize opportunities to develop unique tools and resolve complex operational challenges within critical enterprise applications.
- Architect and implement production automation solutions that measurably reduce manual effort and boost operational efficiency.
- Design and deploy AI/ML-driven automation pipelines, observability enhancements, and proactive operational response systems, including anomaly detection and predictive alerting, to elevate platform reliability.
- Lead the expansion of automation coverage across deployment, monitoring, alerting, and self-healing workflows for Cloud and Login Platforms.
- Collaborate extensively with Engineering, Scrum, and Operations teams, providing crucial technical expertise and support for key initiatives focused on system availability and reliability.
- Efficiently triage alerts, diagnose, and resolve critical issues, managing change implementations with clear communication and minimal risk.
- Develop essential tools, frameworks, and instrumentation to validate and enhance the success of application rollouts, leveraging AI/ML capabilities for operational visibility and validation at scale.
- Advocate for AIOps platform adoption and ML-assisted observability practices across the team.
- Coordinate robust capacity planning through data-driven trend analysis and ML-informed forecasting.
- Develop CI/CD orchestration systems to streamline software delivery to production, championing GitOps concepts and AI-assisted pipeline optimization.
- Perform real-time troubleshooting of mission-critical application workflows, integrating feedback directly into product development cycles.
- Participate actively in on-call support rotations, ensuring continuous operational excellence.
Must-Have Qualifications
- Extensive hands-on experience (6-8 years) in enterprise-level administration, support, and deployment activities, including crafting automation scripts, developing proactive monitoring dashboards, configuring alerts for early issue detection, and applying SDLC and process improvement methodologies.
- Proficiency with Windows 2019/2022 and Linux operating systems hosted via Virtual Machine.
- Experience in Cloud application configuration, deployment, support, and migration.
- Solid understanding of IP networking fundamentals, including DNS, DHCP, firewalls, and IP routing.
- Familiarity with large-scale distributed systems and high-availability architectures.
- Expertise in Linux and Windows system administration, troubleshooting, and performance tuning.
- Development experience in one or more programming languages such as .NET, PowerShell, Java, Python, or Bash.
- Knowledge of one or more database systems, including SQL, Oracle, or MongoDB.
- Working knowledge of Actimise.
- Familiarity with one or more Message Brokers, such as Solace, RabbitMQ, IBM MQ, or Kafka.
- Experience with Splunk, AppDynamics, or similar observability tools.
- Demonstrated experience applying AI/ML or AIOps approaches (e.g., anomaly detection, predictive alerting, ML-assisted observability) in production environments.
- A Bachelor’s degree in computer science or a related discipline.
- Strong customer orientation with an affinity for proactively owning, communicating, and following through on projects and issues.
- An extreme sense of ownership to meticulously resolve problems within a distributed environment.
- A gritty resolve to delve deep into technical issues within a complex ecosystem.
- A self-starter mentality with the ability and confidence to independently resolve issues and deliver results to the team.
Nice-to-Have Qualifications
- Experience within the financial services industry.
- Proficiency with Agile methodologies.
- Hands-on experience with AIOps platforms or ML-driven observability tooling.
- Experience integrating AI/ML capabilities into CI/CD or operational automation workflows.
- Familiarity with CI/CD tools (e.g., Harness, Jenkins, GitHub Actions) or GitOps concepts.
- Exposure to container orchestration platforms (e.g., Kubernetes, OpenShift) or cloud platforms (e.g., AWS, Azure, GCP).
Aquent is an equal-opportunity employer. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, and other legally protected characteristics. We’re about creating an inclusive environment-one where different backgrounds, experiences, and perspectives are valued, and everyone can contribute, grow their careers, and thrive.