Tiger Analytics is seeking a Site Reliability Engineer (SRE) in Washington, DC to ensure the resilience and performance of data-driven AI platforms. The role involves a blend of software engineering and systems architecture, focusing on MLOps to maintain high availability for AI models. Responsibilities include managing SLAs, Kubernetes scaling, and automating operations. Ideal candidates will have expertise in Kubernetes, Docker, and Python. Benefits include significant career development opportunities in a fast-growing environment.
Qualifications
Expert-level knowledge of Kubernetes and Docker required.
Strong proficiency in Python and Bash for automation.
Familiarity with MLOps tools like Kubeflow or Vertex AI is needed.
Responsibilities
Ensure production ecosystems are resilient, scalable, and performant.
Define and maintain SLOs and SLIs for AI/ML services.
Optimize resources for Large Language Models.
Skills
Kubernetes
Docker
Python
Bash
MLOps tools (Kubeflow, Vertex AI)
Data Systems management
Networking (VPCs, Load Balancers)
Tools
Terraform
Prometheus
Grafana
Google Cloud Operations Suite
Job description
Tiger Analytics is seeking a Site Reliability Engineer (SRE) in Washington, DC to ensure the resilience and performance of data-driven AI platforms. The role involves a blend of software engineering and systems architecture, focusing on MLOps to maintain high availability for AI models. Responsibilities include managing SLAs, Kubernetes scaling, and automating operations. Ideal candidates will have expertise in Kubernetes, Docker, and Python. Benefits include significant career development opportunities in a fast-growing environment.