We are looking for a Senior Data Site Reliability Engineer (SRE) to join our Data Platform team. This role focuses on building highly reliable, scalable, and automated data systems across cloud environments.
You will work at the intersection of Data engineering, SRE, and Platform Automation, ensuring our data pipelines, ingestion systems, and infrastructure run efficiently with minimal manual intervention.
Key Responsibilities
- Design and build automation-first solutions for data platform operations
- Own reliability, availability, and performance of data pipelines and ingestion systems
- Develop and maintain self-healing systems, monitoring, and alerting frameworks
- Build tools to automate:
- Job orchestration and recovery
- Deployment workflows (CI/CD for data pipelines)
- Environment provisioning and configuration
- Work with GCP services (GKE, Cloud Run, Pub/Sub, BigQuery, etc.)
- Improve system observability using logs, metrics, and tracing
- Drive incident management, root cause analysis (RCA), and postmortems
- Optimize infrastructure costs and performance
- Collaborate with data engineers and platform teams to enforce SRE best practices
- Build internal tools (CLI/UI/AI-based agents) for operational efficiency
Required Skills & Experience
- 7+ years in SRE / DevOps / Data Platform Engineering/ Data Engineering
- Strong experience in Python (automation, tooling, scripting)
- Hands-on experience with:
- Kubernetes (GKE preferred)
- Cloud platforms (GCP strongly preferred)
- CI/CD pipelines
- Experience with data systems:
- Batch & streaming pipelines
- Pub/Sub, Kafka, or similar messaging systems
- Strong understanding of:
- Monitoring (Datadog, Cloud Monitoring, etc.)
- Logging and observability
- Experience in infrastructure as code (Terraform or equivalent)
- Solid debugging and production troubleshooting skills
Good to Have
- Experience with AlloyDB / PostgreSQL / Database reliability
- Experience building internal platforms or developer tools
- Knowledge of Helm, container lifecycle, and image management
- Experience with large-scale data ingestion systems
What Were Looking For
- Strong ownership mindset - you build it, you run it
- Automation-first thinking - eliminate manual processes wherever possible
- Ability to work in ambiguous environments and define solutions
- Focus on reliability, scalability, and long-term maintainability
- Good communication and collaboration skills
Location Chennai : Resource has to work based out of Chennai Ingram office(Guindy) daily.
Locations
Chennai
Skill Set
Python,Continuous integration (CI) and continuous delivery (CD),Kubernetes (GKE preferred),Kafka,CKAD,Terraform,GCP Cloud