A global cloud technology firm is looking for a Senior Site Reliability Engineer (SRE) to join their team in Kraków, Poland. The role involves owning reliability workstreams for a serverless inference platform, building automation, and contributing to architecture decisions. Candidates should have expertise in managing large-scale distributed systems, proficiency in Kubernetes, and skills in automation using Python or Go. Flexible working arrangements are available.
Qualifications
Expertise in managing large-scale distributed systems.
Proficient in CI/CD pipelines and deployment safety.
Ability to independently resolve issues and ensure accountability.
Responsibilities
Build and maintain observability for AI workloads.
Write automation to improve deployment safety.
Integrate AI workloads into incident management processes.
Skills
SRE expertise
Kubernetes
Python
Go
AI/ML infrastructure
Observability tools (Prometheus, Grafana)
Infrastructure-as-code (Terraform)
Job description
A global cloud technology firm is looking for a Senior Site Reliability Engineer (SRE) to join their team in Kraków, Poland. The role involves owning reliability workstreams for a serverless inference platform, building automation, and contributing to architecture decisions. Candidates should have expertise in managing large-scale distributed systems, proficiency in Kubernetes, and skills in automation using Python or Go. Flexible working arrangements are available.