Sr. Platform Engineer, Kubernetes

Comcast

West Chester (Chester County)

On-site

USD 140,000 - 180,000

Full time

48 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Comcast in the United States is seeking a Sr. Platform Engineer to build, manage, and optimize the infrastructure that runs large-scale data workloads.

You will design systems for metrics collection with Prometheus and visualization with Grafana, ensuring a robust data processing environment for engineers and scientists. Responsibilities include managing Spark platforms on Kubernetes and AWS, packaging workloads with Docker, deploying with Terraform/Ansible, and writing Python/Scala.

Qualifications

  • Bachelor's degree or equivalent experience in computer science or related field; typically 7 years in a DevOps or Systems Engineering role.
  • Expertise in Apache Spark: deep understanding of Spark architecture, RDDs, DataFrames, execution hierarchy, lazy evaluation, shuffling, fault tolerance.
  • Proficiency in Python, PySpark, and Scala/Java for Spark development and automation.
  • Proficiency in Linux scripting (Bash).
  • Proficient in SQL writing.
  • Experience with CI/CD tools and GitHub.
  • Experience with observability tools like Prometheus, Grafana.
  • Strong knowledge of networking protocols (TCP/IP, DNS) and hardware basics.
  • Automation via Terraform/Ansible; cloud/on-prem experience (AWS/Azure/GCP); Docker/Kubernetes.

Responsibilities

  • Architect and manage Spark platforms on Kubernetes and cloud services (AWS EKS).
  • Package Spark workloads and integrate with orchestration systems like Flyte.
  • Deploy infrastructure using Terraform/Ansible.
  • Troubleshoot and optimize job failures and resource issues for cost efficiency.
  • Develop tools in Python/Java/Scala and shell scripting to boost productivity.
  • Write medium to complex SQL queries as needed.
  • Implement and maintain monitoring, logging, and alerting with Prometheus and Grafana.
  • Develop data catalog capabilities (Iceberg, Unity Catalog) for authorization and lineage.
  • Collaborate with Data Stewards, Analysts, and Scientists to address data needs.

Skills

Apache Spark
Python
Scala/Java
Linux Bash
SQL
CI/CD
Prometheus Grafana
Terraform/Ansible
AWS/Azure/GCP
Docker Kubernetes
Data Governance
Delta Lake
Apache Iceberg
Apache Kafka
Databricks
Unity Catalog
Alluxio
IAM/VPC
Networking TCP/IP
Hadoop

Education

Bachelor's degree in Computer Science or related field

Tools

Docker
Kubernetes
Databricks
Snowflake
Unity Catalog

Job description

Make your mark at Comcast -- a Fortune 30 global media and technology company. From the connectivity and platforms we provide, to the content and experiences we create, we reach hundreds of millions of customers, viewers, and guests worldwide. Become part of our award-winning technology team that turns big ideas into cutting-edge products, platforms, and solutions that our customers love. We create space to innovate, and we recognize, reward, and invest in your ideas, while ensuring you can proudly bring your authentic self to the workplace. Join us. You’ll do the best work of your career right here at Comcast. (In most cases, Comcast prefers to have employees on-site collaborating unless the team has been designated as virtual due to the nature of their work. If a position is listed with both office locations and virtual offerings, Comcast may be willing to consider candidates who live greater than 100 miles from the office for the remote option.)

Job Summary

As a Sr. Platform Engineer, you will be responsible for building, managing, and optimizing the underlying infrastructure and tools that enable efficient, scalable, and reliable execution of large-scale data processing workloads. Designing systems for collecting metrics (Prometheus) and visualizing data (Grafana) to provide deep insights into application/infrastructure performance. This role is a specialized subset of data platform engineering, ensuring the environment where data engineers and data scientists run their Spark jobs is robust and cost-efficient.

Responsibilities and Duties
  • Architecting and managing the platforms where Spark runs, such as Kubernetes clusters, or cloud services like AWS (EKS).
  • Packaging Spark workloads (often via Docker/Kubernetes) and integrating them with orchestration systems like Apache Flyte.
  • Deploying Infrastructure via Terraform/Ansible
  • Troubleshooting and resolving job failures, memory/resource issues, and execution anomalies. This includes optimizing Spark configurations to reduce cloud compute and storage costs.
  • Building automation and tools in languages like Python, Java, or Scala, Linux Scripting (Bash) to increase the productivity of development teams.
  • Write medium to complex SQL Queries as needed.
  • Implementing and maintaining systems for monitoring, logging, and alerting (e.g., Prometheus, Grafana) to ensure platform stability and reliability.
  • Develop and optimize the data catalog platform (e.g., Apache Iceberg, Unity Catalog) for authorization, search, and lineage.
  • Automate workflows, monitoring, and incident resolution.
  • Collaborate with Data Stewards, Analysts, and Scientists to address data needs and issues.
  • Promote best practices and assess emerging technologies.
  • Working closely with data engineers, data scientists, and other engineering teams to define requirements, advise on best practices, and ensure successful delivery of data objectives.
  • Engaging with open-source communities (like Apache Spark, Delta Lake, or Apache Iceberg) to discuss technical challenges and contribute improvements.
  • Create and maintain comprehensive documentation for Kubernetes infrastructure, processes, and procedures. Provide training and support to team members as needed.
Qualifications
  • Bachelor's degree in computer science or a related field, or equivalent experience, typically 7 years in a DevOps or Systems Engineering role.
  • Expertise in Apache Spark: Deep understanding of Spark architecture, including RDDs, DataFrames, execution hierarchy, lazy evaluation, shuffling, and fault tolerance.
  • Proficiency in languages used for Spark development and automation, such as Python, Pyspark and Scala/Java.
  • Proficient in Linux Scripting (Bash).
  • Proficient in writing SQL.
  • Experience in CI/CD tools, Github.
  • Experience in setting up and using observability tools like Prometheus, Grafana etc.,
  • Strong knowledge on Networking Protocols (TCP/IP, DNS, Load Balancer etc.,) and hardware components,
  • Automation via Terraform/Ansible
  • Hands-on experience with on-prem and major cloud providers (AWS, Azure, GCP) and container orchestration tools like Docker and Kubernetes.
  • Handson experience setting up IAM, VPC, EC2 etc.,
  • Familiarity with related technologies and formats like Delta Lake, Apache Iceberg, Apache Kafka, Hadoop, and various data storage systems (S3, HDFS, etc.).
  • Hands-on experience with Databricks, Snowflake, Apache Iceberg, Unity Catalog, or similar tools.
  • Solid understanding of data lakes and governance.
  • Experience setting up, maintaining caching layers like Alluxio.
  • Strong analytical skills for debugging complex distributed systems issues.
  • Strong communication and collaboration abilities.

Disclaimer: This information has been designed to indicate the general nature and level of work performed by employees in this role. It is not designed to contain or be interpreted as a comprehensive inventory of all duties, responsibilities and qualifications.

Comcast is an equal opportunity workplace. We will consider all qualified applicants for employment without regard to race, color, religion, age, sex, sexual orientation, gender identity, national origin, disability, veteran status, genetic information, or any other basis protected by applicable law.

Skills

Kubernetes; Terraform (Software); Amazon S3; Databricks Platform; Linux; Adobe Spark; Go Programming Language

Base pay is one part of the Total Rewards that Comcast provides to compensate and recognize employees for their work. Most sales positions are eligible for a Commission under the terms of an applicable plan, while most non-sales positions are eligible for a Bonus. Additionally, Comcast provides best-in-class Benefits to eligible employees. We believe that benefits should connect you to the support you need when it matters most, and should help you care for those who matter most. That’s why we provide an array of options, expert guidance and always-on tools, that are personalized to meet the needs of your reality - to help support you physically, financially and emotionally through the big milestones and in your everyday life. Please visit the compensation and benefits summary on our careers site for more details.

Education

Bachelor's Degree

Relevant Work Experience

7-10 Years

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Platform Engineer, Kubernetes
Sr. Platform Engineer, Kubernetes

Blueface Ltd • West Whiteland Township

On-site
USD 140,000 - 190,000
Sr. Platform Engineer, Kubernetes
Sr. Platform Engineer, Kubernetes

Blueface Ltd • Chester

On-site
USD 140,000 - 190,000
Platform Data Engineer - (DataBricks, PySpark, AWS)
Platform Data Engineer - (DataBricks, PySpark, AWS)

Blueface Ltd • West Whiteland Township

On-site
USD 120,000 - 160,000
Platform Data Engineer - (DataBricks, PySpark, AWS)
Platform Data Engineer - (DataBricks, PySpark, AWS)

Blueface Ltd • Chester

On-site
USD 120,000 - 180,000
Data Engineer
Data Engineer

Comcast • Chester

On-site
USD 140,000 - 190,000
Data Engineer
Data Engineer

Blueface Ltd • Chester

On-site
USD 120,000 - 170,000
Data Engineer
Data Engineer

Blueface Ltd • West Whiteland Township

On-site
USD 120,000 - 180,000
Senior Platform Engineer: Kubernetes & Spark
Senior Platform Engineer: Kubernetes & Spark

Blueface Ltd • West Whiteland Township

On-site
USD 140,000 - 190,000
Lead Software Engineer - Kubernetes Platform Management - Freewheel
Lead Software Engineer - Kubernetes Platform Management - Freewheel

Blueface Ltd • Reston (VA)

On-site
USD 167,000 - 251,000
Backend Platform Engineer - AI Operations Platform
Backend Platform Engineer - AI Operations Platform

Comcast • Philadelphia

On-site
USD 120,000 - 170,000
On-site collaboration