Optomi, in partnership with a leading global financial services organization, is seeking an experienced Senior Site Reliability Engineers (SRE) to join a newly formed Application Reliability team. These are contract to hire positions with a strong preference for candidates in the Dallas/Fort Worth or Plano, TX area (hybrid schedule).
What the Right Professional Will Enjoy:
- Playing a foundational role in building a greenfield SRE practice focused on mission-critical financial applications
- Working in a highly collaborative environment alongside development, product, and infrastructure teams to drive measurable reliability improvements
- Opportunity to design and implement cutting-edge observability, automation, and self-healing capabilities at scale
- Mentoring engineers and influencing modern reliability practices across the enterprise
Apply today if your background includes:
- 5+ years in Site Reliability Engineering, DevOps, Platform Engineering, or equivalent production engineering roles
- Hands‑on experience designing, building, and maintaining large‑scale, highly available systems in AWS (EKS, ECS, or similar)
- Proven expertise with Infrastructure as Code (Terraform) and modern CI/CD platforms (GitHub Actions, Harness, ArgoCD, or equivalent)
- Strong programming and automation skills with Python (or similar) for operational tooling and runbooks
- Deep experience with container orchestration (Kubernetes/EKS highly preferred) and cloud‑native ecosystems
- Track record of implementing observability solutions (Dynatrace, CloudWatch, Prometheus, Grafana, OpenTelemetry, etc.) that drive actionable insights
- Solid understanding of SLOs, error budgets, SLAs, and blameless post‑incident processes
- Demonstrated ability to lead major incident response, root cause analysis, and preventive remediation efforts
- Experience embedding reliability practices into application architecture and development lifecycles
Preferred Qualifications:
- Familiarity with GitOps workflows, secrets management (HashiCorp Vault, AWS Secrets Manager), and advanced monitoring/alerting best practices
- History of building self‑healing and automated remediation systems