Tiger Analytics is seeking a Site Reliability Engineer (SRE) to maintain and ensure the reliability of our production AI platforms. The role requires expertise in Kubernetes, MLOps, and automation tools like Terraform. You will define and monitor service level objectives, architect scalable solutions, and play a crucial part in performance engineering for AI services. This hybrid position provides significant career development opportunities in a fast-growing environment.
Qualifications
Experience in reliability and performance engineering, particularly with AI/ML services.
Familiarity with automation tools for infrastructure management.
Strong background in networking and cloud architectures.
Responsibilities
Define, monitor, and maintain SLOs for AI/ML services.
Architect auto-scaling strategies for Kubernetes.
Ensure high availability of Vertex AI endpoints.
Skills
Kubernetes (GKE)
Python
Terraform
MLOps
CI/CD
Grafana
Prometheus
Tools
Vertex AI
Kubeflow
Google Cloud Operations Suite (Stackdriver)
Job description
Tiger Analytics is seeking a Site Reliability Engineer (SRE) to maintain and ensure the reliability of our production AI platforms. The role requires expertise in Kubernetes, MLOps, and automation tools like Terraform. You will define and monitor service level objectives, architect scalable solutions, and play a crucial part in performance engineering for AI services. This hybrid position provides significant career development opportunities in a fast-growing environment.