Junior Site Reliability Engineer

EPAM Systems

México

Remote

MXN 420,000 - 660,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

EPAM Systems, Inc. is seeking a Junior Site Reliability Engineer to help keep cloud services reliable, observable, and automated across Kubernetes and Microsoft Azure environments. You will support incident response, improve monitoring, and reduce toil through scripting while collaborating with development teams.

This entry‑level role offers exposure to international projects and a collaborative, learning‑focused culture within a global engineering organization.

Qualifications

  • 1+ years of site reliability or cloud operations experience.
  • Hands-on experience with Microsoft Azure services and core cloud concepts.
  • Practical experience with Kubernetes and container platforms.
  • Strong scripting skills in Python, Bash, or PowerShell.

Responsibilities

  • Operate containerized workloads on Kubernetes in Microsoft Azure.
  • Troubleshoot production issues end-to-end across network, OS, and application layers.
  • Automate repetitive operational tasks with scripting to reduce toil.
  • Define and track SLIs and SLOs to improve service reliability.
  • Build and maintain monitoring, alerting, and dashboards using the Elastic Stack.
  • Collaborate with development teams to improve resilience and operability.
  • Document runbooks and incident learnings to strengthen operational readiness.

Skills

Azure
Kubernetes
Python
Bash
PowerShell
Linux
Networking
Incident response

Tools

Elastic Stack
Istio
ArgoCD
Windows Administration

Job description

We are looking for a Junior Site Reliability Engineer to help keep cloud services reliable, observable, and automated across Kubernetes and Microsoft Azure environments. You will support incident response, improve monitoring, and reduce toil through scripting while collaborating with development teams.ResponsibilitiesOperate containerized workloads on Kubernetes in Microsoft AzureTroubleshoot production issues end-to-end across network, OS, and application layersAutomate repetitive operational tasks with scripting languages to reduce toilDefine and track SLIs and SLOs to improve service reliabilityBuild and maintain monitoring, alerting, and dashboards using the Elastic StackCollaborate with development teams to improve resilience and operabilityDocument runbooks and incident learnings to strengthen operational readinessRequirements1+ years experience with site reliability or cloud operationsHands-on experience with Microsoft Azure services and core cloud conceptsPractical experience with Kubernetes and container platformsStrong scripting skills in Python, Bash, or PowerShellSolid infrastructure fundamentals in networking and operating systemsStrong Linux administration skillsStrong debugging skills and calm incident response mindsetUpper-Intermediate English proficiency (B2)Clear communication skills for cross-team collaborationNice to haveArgo CDCursorElastic StackIstioWindows AdministrationWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedInEPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Junior SRE: Cloud Reliability & Automation (Azure/K8s)
Junior SRE: Cloud Reliability & Automation (Azure/K8s)

EPAM Systems • Mexico

Remote
MXN 420,000 - 660,000
Lead Data DevOps
Lead Data DevOps

EPAM Systems • Mexico

Remote
MXN 900,000 - 1,500,000
Healthcare benefits
Paid time off and sick leave
Upskilling & certification courses
+2
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Rosarito

On-site
MXN 870,019 - 1,305,028
Professional growth
Competitive compensation
Exciting projects
+1
Site Reliability Engineer ID60188
Site Reliability Engineer ID60188

AgileEngine • Ciudad de México

On-site
MXN 1,049,685 - 1,399,580
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Lead Infrastructure Engineer (Automation & Orchestration)
Lead Infrastructure Engineer (Automation & Orchestration)

EPAM Systems • Mexico

Remote
MXN 900,000 - 1,300,000
Healthcare benefits
Professional development & training
Paid time off
Functional Tester
Functional Tester

EPAM Systems, Inc. • Mexico

Remote
MXN 360,000 - 480,000
Senior Full-Stack Developer (.NET)
Senior Full-Stack Developer (.NET)

EPAM Systems • Mexico

Remote
MXN 700,000 - 1,100,000
Junior Support Specialist
Junior Support Specialist

EPAM Systems • Guadalajara

Hybrid
MXN 180,000 - 240,000
Healthcare benefits
Paid time off
LinkedIn Learning access
+1
DevOps Engineer ID89052
DevOps Engineer ID89052

AgileEngine, LLC. • Región Centro

Hybrid
MXN 420,000 - 700,000
Professional growth
Competitive USD-based compensation
Challenging projects with Fortune 500/
DevOps Engineer ID89052
DevOps Engineer ID89052

AgileEngine, LLC. • Monterrey

Hybrid
MXN 1,262,000 - 1,983,000
Professional growth
Competitive USD pay
Exciting projects
+1