Engineer Gurugram Full Time Mid Level
ABOUT SQUAREOPS
SquareOps is a managed DevOps and SRE company. We run production infrastructure for cloud-native product teams startups, scale-ups, and enterprises across AWS, GCP, and Azure. Our clients build on modern stacks: Kubernetes, ECS, serverless, and AI-powered applications. Several of our clients operate in fintech and healthcare, where uptime, security, and information discipline are non-negotiable.
This is not a legacy IT shop. You will operate infrastructure that is genuinely cutting edge the kind of stack most engineers only read about.
We are scaling our 24x7 shared SRE operations team and hiring L1 SRE Engineers.
WHAT YOU WILL DO
- Monitor cloud infrastructure (AWS primary, GCP and Azure secondary) across multiple client environments simultaneously, using dashboards, alerting tools, and logs to spot and diagnose issues
- Provide hands-on production support investigate degraded services, failed jobs, resource exhaustion, and connectivity issues, and take corrective action
- Work directly in AWS, Azure, and container environments (ECS/Kubernetes) to check service health, restart/scale workloads, inspect logs, and validate fixes
- Use observability tooling (CloudWatch, Grafana, Prometheus, Loki, or client-specific equivalents) to build situational awareness and catch issues early
- Write and run operational scripts (Bash/Python) to automate checks, gather diagnostics, and remediate known issues
- Execute infrastructure changes during change windows deployments, config updates, scaling actions safely and accurately
- Handle client credentials, access, and operational data with strict infosec discipline
- Coordinate with L2 engineers and client stakeholders during escalations, communicating technical findings clearly
WHAT WE ARE LOOKING FOR
Must Have (Technical)
- Linux, hands-on process management, resource monitoring (CPU/memory/disk), log navigation, systemd/services, filesystem troubleshooting
- Networking fundamentals, applied can actually diagnose DNS resolution issues, TCP/IP connectivity problems, HTTP/S errors, and load balancer misrouting not just definitions
- AWS, hands-on comfortable working inside EC2, ECS, RDS, CloudWatch, IAM, and S3; can navigate the console and CLI to investigate and fix issues, not just describe services
- Monitoring & observability real experience reading dashboards, setting or tuning alerts, and correlating metrics/logs to root-cause an issue (CloudWatch, Grafana, Prometheus, Loki, or equivalent)
- Scripting Bash and/or Python at an operational level: can write a script to check service status, parse logs, or automate a repetitive diagnostic task
- Troubleshooting mindset can walk through how they'd debug a production issue (e.g., "service is returning 502s" or "pod is crash-looping") from first principles
- Strong written and verbal English technical communication with clients is part of the role
- Comfortable with rotational shifts including nights