Senior Site Reliability Engineer

Avaloq

Town of Florida (NY)

Hybrid

USD 140,000 - 200,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Hybrid work model

Job summary

Avaloq is hiring a Senior Site Reliability Engineer for its Fort Lauderdale team. You will help define SLIs/SLOs, shape the observability strategy, and lead incident response in a distributed cloud native setup. You will collaborate with Zurich and product teams to improve deployment and resilience in a global SaaS environment.

The role emphasizes automation, scalable operations, and strong cross-team communication, with regular travel to Switzerland and a hybrid work model.

Qualifications

  • 5+ years in Site Reliability Engineering, DevOps, or production ops.
  • Hands-on AWS experience and serverless architecture knowledge.
  • Strong incident response background and post-mortem practice.
  • Clear view on observability for distributed systems and metrics.
  • Automation-focused with scripting in Rust/TypeScript/Python.

Responsibilities

  • Define and evolve SLIs/SLOs and use them to guide reliability decisions.
  • Design observability for a distributed serverless system (metrics/logs/traces).
  • Lead incident response: on-call model, escalation, post-mortems.
  • Build reliability automation to detect and remediate issues before impact.
  • Improve CI/CD pipelines and deployment with GitHub Actions.

Skills

SRE
AWS
Observability
Incident response
CI/CD
Communication

Tools

GitHub Actions
Terraform
OpenTofu
DynamoDB
Lambda
Rust
TypeScript
Python

Job description

Founded and headquartered in Switzerland, Avaloq is continuously expanding its global footprint with around 2,500 colleagues in 10 countries, and more than 170 clients in 35 countries. We are an industry-leading provider of wealth management technology and services for financial institutions around the world, including private banks and wealth managers, investment managers, as well as retail and neo banks. Our research led approach and continual innovation is powered by the passion and creativity of our colleagues.

We are always looking for talented people to join us on our mission to orchestrate the financial ecosystem and democratize access to wealth management. Avaloq offers the opportunity to work closely with some of the world’s leading financial institutions as we jointly develop and shape careers. Championing a collaborative, supportive and flexible work environment empowers our colleagues to reach their full potential.

Avaloq’s R&D Lab is building a SaaS, API-first, composable banking platform. As a Senior Site Reliability Engineer you will help build the reliability practice behind it.

This is amongst the first technical roles we are hiring for our Fort Lauderdale site. You will partner with our Senior SRE in Zurich and work alongside our platform engineers and product teams.

The first three to six months:

Some of what you can expect early on:

  • Learning our AWS environment, our serverless platform, and how our product teams build and release
  • Contributing to the direction on observability tooling, together with the platform engineers, the SRE team, and our developers
  • Working with the Zurich team on the incident and on-call operating model
  • Taking on increasing responsibility as you build context, always in alignment with the Zurich team
What you will do:
  • Define and evolve SLIs, SLOs, and error budgets with product teams, and use them to drive reliability decisions
  • Collaborate on the design of our observability approach for a distributed serverless system, covering metrics, logs, and traces
  • Build the incident response practice with us: on-call model, escalation, blameless post-mortems, and the loop that turns findings into hardening work
  • Build reliability automation that detects and remediates issues before they reach clients
  • Improve CI/CD pipelines and deployment automation on GitHub Actions to reduce operational toil and release risk
  • Work with product teams on resilient design, capacity planning, and progressive delivery approaches such as canary and blue-green
  • Contribute to disaster recovery design and testing for client-facing environments
  • Partner with Security and Compliance so that operational practice holds up to regulatory and audit expectations
  • Help us define and implement operational readiness for client go-live
  • Mentor colleagues and support a culture of shared operational ownership
Our stack:

Serverless-first on AWS: Lambda, DynamoDB, SQS, S3, and Bedrock. Terraform and OpenTofu for infrastructure, GitHub Actions for CI/CD. Product teams work in Rust, TypeScript, Python, Vue, and Angular. We do not run containers.

  • 5+ years in Site Reliability Engineering, DevOps, or production operations for distributed cloud systems, with substantial hands-on AWS experience
  • Practical experience defining and working with SLIs, SLOs, and error budgets
  • Strong incident response background, including on-call, triage under pressure, and post-mortem practice
  • A clear point of view on observability for distributed systems, and the reasoning behind it
  • Solid automation and scripting ability. We are language-agnostic; Rust, TypeScript, or Python all work here
  • Excellent collaboration and communication skills, with a pragmatic approach to balancing speed and stability
It would be a real bonus if you have:
  • Experience with AWS serverless and event-driven architecture, including idempotency, retries, dead-letter queues, and failure handling. Strong AWS generalists who want to go deep here are welcome
  • Experience taking a platform from pre-launch to production, including defining operational readiness criteria
  • Experience building AI agents to monitor platform health
  • Strong CI/CD background, ideally with GitHub Actions, including secrets management, least-privilege permissions, and deployment controls
  • Infrastructure as Code experience with Terraform or OpenTofu
  • Experience with progressive deployment approaches such as canary, blue-green, or rollback automation
  • Disaster recovery design and testing for production SaaS
  • Experience operating across more than one cloud provider, or designing for portability
  • Experience operating a B2B SaaS platform in a regulated environment, including audit-sensitive release processes
  • Familiarity with SOC 2, PCI DSS, or GDPR as they apply to operational practice
  • Experience with AI-assisted engineering tools
  • AWS certification, particularly DevOps Engineer - Professional, Solutions Architect - Professional, or Security - Specialty
This role will require regular travel to Switzerland

We realize that managing work life balance is a challenge we all face in our daily lives and in order to support with this we are pleased to offer hybrid and flexible working for most of our Avaloqers to maintain work life balance and still continue our fantastic Avaloq culture in our global offices.

In Avaloq we are proud to embrace diversity and understand the success of our business is built on the power of different opinions, we are whole heartedly committed to fostering an equal opportunity environment and inclusive culture where you can be your true authentic self.

We hire, compensate and promote regardless of origin, age, gender identity, sexual orientation or any other fantastic traits that make us all unique, we have done our best to write this advert in an inclusive and neutral way.

Please be aware that we will not accept speculative CV submissions for any of our roles from recruitment agencies, and any unsolicited candidate submissions will be exempt from any payment expectations.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Avaloq • Fort Lauderdale (FL)

Hybrid
USD 130,000 - 180,000
Hybrid work model
Travel to Switzerland
A Senior Site Reliability Engineer Avaloq Fort Lauderdale, Florida, US
A Senior Site Reliability Engineer Avaloq Fort Lauderdale, Florida, US

Artha Nexgen • Fort Lauderdale (FL), Northern (KY)

On-site
USD 140,000 - 190,000
Hybrid work model
Global team collaboration
Travel to Switzerland
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Avaloq Group • Fort Lauderdale (FL)

Hybrid
USD 140,000 - 190,000
Hybrid & flexible work
Travel to Switzerland
Senior Platform Engineer
Senior Platform Engineer

Avaloq Group • Fort Lauderdale (FL)

On-site
USD 140,000 - 190,000
Hybrid work model
Flexible working
Senior Platform Engineer
Senior Platform Engineer

Avaloq • Fort Lauderdale (FL)

Hybrid
USD 140,000 - 180,000
Hybrid work model
Flexible working hours
Senior Platform Engineer
Senior Platform Engineer

Avaloq • Town of Florida (NY)

Hybrid
USD 150,000 - 210,000
Hybrid work model
Global collaboration
Flexible working
Senior Solutions Architect
Senior Solutions Architect

Avaloq • Fort Lauderdale (FL)

Hybrid
USD 150,000 - 210,000
Hybrid work model
Travel opportunities
Senior Solutions Architect
Senior Solutions Architect

Avaloq • Town of Florida (NY)

Hybrid
USD 150,000 - 210,000
Hybrid work model
Travel to Switzerland
Senior Solutions Architect
Senior Solutions Architect

Avaloq Group • Fort Lauderdale (FL)

Hybrid
USD 140,000 - 190,000
Hybrid work model
Flexible working
Senior Test Engineer
Senior Test Engineer

Avaloq • Town of Florida (NY)

Hybrid
USD 110,000 - 170,000
Hybrid work model
Flexible working hours
Global offices