Site Reliability Engineer - Research Compute for AI
Astera Institute
United States
On-site
USD 120,000 - 190,000
Full time
14 days+
Application generator
Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Get past ATS filters
Job summary
A leading research foundation in the United States seeks a Site Reliability Engineer to manage its digital infrastructure, ensuring reliable access to compute resources for research. The role involves automating processes, enhancing resource visibility, and managing access efficiently. Candidates should have strong experience with technologies like Kubernetes and Docker, and a keen understanding of systems operations. This position is on-site in Emeryville, with potential visa sponsorship for qualified applicants.
Qualifications
Comfortable being accountable for cluster health and capacity.
Understanding of schedulers, containers, and networking interactions.
Values observability, reproducibility, and clear operational boundaries.
Can support experimental workloads and adapt to project needs.
Responsibilities
Ensure easy access to compute resources for researchers.
Provide visibility into resource utilization and cluster health.
Enable automatic scaling of compute resources.
Manage access to resources effectively.
Drive towards reproducible research environments.
Automate operational processes to increase efficiency.
Skills
Ansible
Kubernetes
Docker
Python
Grafana
Prometheus
Tailscale
Talos Linux
Job description
A leading research foundation in the United States seeks a Site Reliability Engineer to manage its digital infrastructure, ensuring reliable access to compute resources for research. The role involves automating processes, enhancing resource visibility, and managing access efficiently. Candidates should have strong experience with technologies like Kubernetes and Docker, and a keen understanding of systems operations. This position is on-site in Emeryville, with potential visa sponsorship for qualified applicants.