Production experience in SRE / Infrastructure / ops for large-scale systems
Strong programming/scripting skills (Python, Go, Java, or equivalent)
Deep experience with containerization (Docker), orchestration (Kubernetes)
Familiarity with GPU / AI compute clusters
Experience with monitoring / observability tools
Networking & systems engineering knowledge
Experience in capacity planning and incident response
Excellent communication and collaboration skills
Proven track record of reducing operational toil via automation