Описание
Payward builds financial infrastructure platforms to advance an open, global financial system.
Задачи
- Scale and tune high-throughput MariaDB clusters for global expansion and new product initiatives
- Strengthen high-availability, disaster-recovery, and backup approaches through sound design and regularly validated procedures
- Reduce manual operational work by building automation, improving process consistency, and enabling safe, low-friction database workflows
- Improve observability and alert quality by championing meaningful metrics, reducing noise, and ensuring operational clarity across environments
- Enhance database security through robust access controls, disciplined patching and upgrade practices, and secure operational patterns
- Contribute to platform initiatives, including containerized environments, infrastructure-as-code workflows, and reproducible deployments
- Partner with service teams to improve performance patterns, operational readiness, and database hygiene across the engineering organization
- Participate in on-call rotations, with a long-term focus on making on-call predictable through high-quality signals and preventative engineering
Требования
- 5+ Years operating MariaDB/MySQL in high-volume production environments, including tuning, replication, and troubleshooting
- Strong understanding of high-availability patterns, distributed database behavior, and practical failure scenarios
- Hands-on experience with database proxies or load-balancing layers such as HAProxy, ProxySQL, or MaxScale
- Practical experience with CI/CD, GitOps, and Infrastructure-as-Code workflows; Terraform experience is ideal
- Solid cloud, Linux, and networking fundamentals
- Experience building container images and managing Kubernetes workloads at scale
- Strong security instincts around access control, upgrade processes, and safe operational workflows
- Expertise in monitoring, alerting hygiene, and incident response readiness
- Strong communication and collaboration skills, including partnering with stakeholders, negotiating long-term plans, writing formal documentation, and tying success to specific metrics and KPIs
Будет плюсом
- SRE methodologies such as error budgets, operational reviews, or reliability programs
- Python or Rust scripting or programming for automation and internal tools
- GitOps tooling such as ArgoCD, GitHub Actions, or GitLab CI
- multi-region or disaster-recovery designs
- interest in cryptocurrency or decentralized systems
Условия
Applications are accepted on an ongoing basis unless a specific application deadline is stated Candidates may complete job-related skills or work-style assessments as part of the hiring process Qualified applicants with criminal histories are considered consistently with the requirements of the San Francisco Fair Chance Ordinance