Senior Platform Reliability Engineer – AI-Driven FinServ

interface.ai

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

100% paid health, dental & vision
401(k) & financial wellness
Daily meals on us
Commuter benefit
Discretionary PTO + parental leave
Claude Enterprise + frontier AI tools

Job summary

interface.ai is seeking a hands-on, senior IC to own the reliability and security backbone of our AI-driven platform for financial services. You will define SLOs, lead incident response, and implement GitOps-based delivery across a multi-service AWS/Kubernetes stack.

You will drive observability and AI-native operations, ensure disaster recovery readiness, and maintain security posture for a trusted production environment used by banks and credit unions.

Qualifications

  • BS/BA in Computer Science required; MS or PhD a strong plus. San Francisco-based and committed to working onsite; on-call participation. H1B transfers welcome.
  • Deep production Kubernetes on AWS, including service mesh, with GitOps-based delivery across many services.
  • Infrastructure-as-code at multi-account scale, including taking over and reshaping a large existing estate.
  • You've built an SLO and error-budget practice that actually changed release decisions, with alerting tuned to burn rate rather than noise.
  • You've delivered multi-region or DR capability with defined RTO/RPO and proven it with real failover tests.
  • Security as daily practice, not a checklist — least-privilege IAM, secrets management, admission and network policy, software supply-chain controls.

Responsibilities

  • Reliability & SLOs — define customer-facing SLIs and SLOs across product surfaces, run an error-budget program that governs release decisions.
  • Resilience & disaster recovery — own the DR strategy: regional failover, written RTO/RPO per tier, resilience against third-party dependency failure.
  • Deploy & delivery — a GitOps deploy path with progressive delivery, automated analysis, and one-click rollback for every service.
  • Cloud & infrastructure-as-code — AWS foundation, Kubernetes platform and service mesh; capacity, cost, and scale.
  • Incident management — end to end: paging and severity policy, incident-commander rotation, status-page automation.
  • Observability — metrics, logging, and distributed tracing across services; dashboards and alerts as code.
  • AI-native operations — build automation and guardrails for safe AI-enabled deployment and self-healing templates.

Skills

Production reliability
Kubernetes on AWS
GitOps
SLOs & error budgets
Incident management
Observability & monitoring
Security best practices
TypeScript/Python/Bash

Education

BS/BA in Computer Science

Tools

Kubernetes
AWS
GitOps tooling

Job description

interface.ai is seeking a hands-on, senior IC to own the reliability and security backbone of our AI-driven platform for financial services. You will define SLOs, lead incident response, and implement GitOps-based delivery across a multi-service AWS/Kubernetes stack.

You will drive observability and AI-native operations, ensure disaster recovery readiness, and maintain security posture for a trusted production environment used by banks and credit unions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Senior Platform Engineer: AI-Driven Infra & Security
Senior Platform Engineer: AI-Driven Infra & Security

7AI • Boston (MA)

On-site
USD 120,000 - 160,000
Senior Backend Reliability Engineer — AI‑Driven Platform (Remote)
Senior Backend Reliability Engineer — AI‑Driven Platform (Remote)

Affirm • Riverside (OH)

Remote
USD 173,000 - 233,000
Health coverage
FSA Wallets
Time off
+1
AI-Driven SRE Engineer for Cloud & Automation
AI-Driven SRE Engineer for Cloud & Automation

Skill • Austin (TX)

On-site
USD 140,000 - 190,000
Subsidized health plan
Retirement plan with match
Paid sick leave
Senior Backend Reliability Engineer - AI-Driven Platform
Senior Backend Reliability Engineer - AI-Driven Platform

Affirm • Phoenix (AZ)

On-site
USD 173,000 - 233,000
Health care coverage
Flexible Spending Wallets
Time off
+1
Chief Platform Reliability Architect for AI Infrastructure
Chief Platform Reliability Architect for AI Infrastructure

The Consensus • San Jose (CA)

On-site
USD 210,000 - 320,000
Medical, dental, and vision packages
Housing subsidy
Relocation support
+3
Senior Backend Engineer - Reliability Platform with AI
Senior Backend Engineer - Reliability Platform with AI

Affirm • Charlotte (NC)

On-site
USD 173,000 - 233,000
Health care coverage
Flexible Spending Wallets
Time off
+1
Senior Infrastructure Engineer — AI-Driven Fintech (Equity)
Senior Infrastructure Engineer — AI-Driven Fintech (Equity)

Cascading AI Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Impact & Ownership
Collaborative culture
Competitive compensation
+2
Senior Platform Engineer
Senior Platform Engineer

7AI, Inc. • Boston (MA)

On-site
USD 120,000 - 150,000
Platform Reliability Leader for AI Infrastructure
Platform Reliability Leader for AI Infrastructure

Etched.ai, Inc. • San Jose (CA)

On-site
USD 210,000 - 320,000
Medical, dental and vision coverage
Housing subsidy
Relocation support
+3