Tiger Analytics, LLC is seeking a Site Reliability Engineer (SRE) in Washington, D.C. to ensure the reliability and performance of complex AI platforms. This hybrid role combines software engineering and systems architecture with a strong focus on MLOps. Key responsibilities include managing SLAs, optimizing AI infrastructure, automating tasks, and incident response. The position offers significant career growth opportunities in a fast-paced environment.
Qualifications
Expert-level knowledge of Kubernetes and Docker.
Strong proficiency in Python for automation and scripting.
Familiarity with MLOps tools like Kubeflow and Vertex AI.
Responsibilities
Define and maintain Service Level Objectives and Indicators.
Manage auto-scaling strategies for Kubernetes infrastructure.
Ensure high availability of Vertex AI endpoints and services.
Skills
Kubernetes (GKE)
Terraform
Python
MLOps
CI/CD
Tools
Prometheus
Grafana
Vertex AI
Kubeflow
Google Cloud Operations Suite
Job description
Tiger Analytics, LLC is seeking a Site Reliability Engineer (SRE) in Washington, D.C. to ensure the reliability and performance of complex AI platforms. This hybrid role combines software engineering and systems architecture with a strong focus on MLOps. Key responsibilities include managing SLAs, optimizing AI infrastructure, automating tasks, and incident response. The position offers significant career growth opportunities in a fast-paced environment.