Working Hours: 5 PM to 1 AM EST (Mon – Fri)
About the Role & Team
We are expanding our Site Reliability Engineering (SRE) organization. This role is part of newly established offshore SRE teams that will work in close partnership with our US-based engineering teams to ensure the reliability, availability, and performance of critical production systems.
This is a high-impact, front-line operations role focused on real-time incident response, proactive prevention, and continuous automation. Every minute matters—our SREs act decisively to prevent service degradation and protect the customer experience.
What You’ll Do
- Act as the first responder to alerts and production incidents, rapidly assessing severity and initiating mitigation actions
- Serve as Incident Commander during major incidents, leading bridge calls with clarity and urgency
- Drive root cause isolation within 30 minutes for critical incidents whenever possible
- Communicate effectively across engineering, product, and leadership during high-pressure situations
- Maintain a strong presence on incident bridges—this role requires confidence, ownership, and clear decision-making
Proactive Reliability Engineering
- Identify patterns, trends, and signals to prevent incidents before they occur
- Continuously improve alert quality, reduce noise, and increase signal fidelity
- Partner with engineering teams to enhance system resilience and reliability
Automation & Toil Reduction
- Eliminate manual work by automating operational tasks, ticket handling, and repetitive workflows
- Build and improve tooling across incident response, observability, and operations
- Leverage AI-assisted development tools (e.g., Cursor, Claude) where they provide clear value
Platform & Systems Support
Troubleshoot across a hybrid ecosystem including:
- Cloud platforms (AWS, GCP, Azure)
Diagnose and resolve issues across:
- Networking (connectivity, latency, DB access interruptions)
- CDN and traffic management layers (Akamai, waiting rooms – plus)
Required Technical Skills & Experience
Core Engineering & Operations
- Strong experience in incident management and triage in production environments
- Proven ability to troubleshoot complex distributed systems under pressure
- Solid understanding of Linux systems administration (including performance, networking, NTP, etc.)
- Hands-on experience with AWS core services (S3, Lambda, Load Balancers, ECS, EC2)
- Familiarity with GCP and/or Azure environments
- Experience operating in multi-cloud and hybrid environments
Containers & Orchestration
- Understanding of containerized application architectures
DevOps & CI/CD
Strong knowledge of DevOps practices and CI/CD pipelines
Hands-on experience with:
Working knowledge of:
- Java, Node.js, React-based applications
Understanding of database connectivity and dependencies across:
- Oracle, MariaDB, MSSQL (no DBA ownership, but strong troubleshooting awareness required)
Networking
Strong foundational knowledge of:
- Load balancing and network troubleshooting
- Diagnosing connectivity issues between services and databases
Preferred Qualifications
- Experience in large-scale enterprise (Fortune 500) environments supporting mission-critical applications
- Prior experience as an Incident Commander or similar leadership role during outages
- Familiarity with Akamai CDN and traffic management tools
- Experience in high-volume, high-availability production environments