6 years
Full-Time
Location: Offshore — India (Chennai / Bangalore)
Experience: 6+ years
Candidate Skills :
Big Data Platform Engineering Hands-on experience with Big Data ecosystems, including Hadoop distributions, Apache Spark, and distributed ETL pipelines. Supporting, maintaining, and extending large-scale Big Data platform infrastructure under production workloads. Designing, building, testing, and troubleshooting Big Data environments to sustain high reliability and uptime. Conducting capacity planning and scaling infrastructure as data ingestion volumes and compute workloads expand. Documenting platform configurations, infrastructure standards, operational runbooks, and architectures.
Platform Automation & AI Tooling
- Familiarity with applying AI technologies and automation tools within platform operations and data environments.
- Automating recurring operational tasks, monitoring routines, and administrative runbooks using AI-enabled tooling.
- Strong operational scripting and automation expertise using Python, Bash, or similar languages.
- Evaluating and benchmarking open-source technologies to improve efficiency, performance, and operational cost.
Systems, Reliability & Observability
- Solid working knowledge of Linux systems administration and OS-level performance tuning.
- Hands-on experience configuring monitoring, log aggregation, and alerting platforms for data infrastructure.
- Monitoring platform health metrics, job execution performance, and hardware/cluster resource utilization.
- Resolving, debugging, and escalating distributed system errors, job failures, and resource contention issues.
Delivery & Engineering Standards
- Working within Agile delivery frameworks to iteratively deliver platform capabilities.
- Researching and benchmarking emerging tooling against existing platform capabilities.
- Prior experience working within banking, financial services, or regulated enterprise environments is preferred.
- Self-motivated approach to identifying gaps and driving platform enhancements with minimal supervision.
Capabilities
- Support, scale, and extend enterprise Hadoop and Spark Big Data platforms.
- Introduce AI-driven operational tooling to automate repetitive platform tasks and runbooks.
- Build custom automation tools and operational workflows utilizing Python and Bash scripting.
- Evaluate, integrate, and maintain open-source platform technologies to lower operational overhead.
- Administer Linux-based cluster infrastructure, monitor distributed jobs, and troubleshoot cluster bottlenecks.
- Translate evolving business and data engineering needs into reliable platform architecture capabilities.
Responsibilities
- Maintain and extend the Hadoop-based Big Data platform, Spark clusters, and underlying ETL infrastructure.
- Implement and operationalize AI-enabled tooling to automate day-to-day platform maintenance.
- Monitor system health, cluster resource allocation, and Spark/ETL job performance to maintain high availability.
- Test and debug platform components continuously, driving prompt resolution of infrastructure incidents.
- Support capacity planning, cluster sizing, and resource management as enterprise data volumes scale.
- Evaluate, benchmark, and deploy open-source technologies to continuously modernize the platform.
- Maintain operational runbooks, infrastructure documentation, and configuration standards.
- Collaborate with data engineers, solution architects, and business stakeholders to design and deploy platform capabilities.