As a Spark Technical Solutions Engineer, you will provide a deep dive technical and consulting related solutions for the challenging Spark/ML/AI/Delta/Streaming/Lakehouse reported issues by our customers and resolve any challenges involving the Databricks unified analytics platform with your highly comprehensive technical and customer communication skills. You will assist our customers in their Databricks journey and provide them with the guidance, knowledge, and expertise that they need to realize value and achieve their strategic objectives using our products.
Responsibilities
- Perform initial level analysis and troubleshoot issues in Spark using Spark UI metrics, DAG, and event logs for various customer reported job slowness issues.
- Troubleshoot, resolve and suggest deep code‑level analysis of Spark to address customer issues related to Spark core internals, Spark SQL, Structured Streaming, Delta, Lakehouse, and other Databricks runtime features.
- Assist customers in setting up reproducible Spark problems with solutions in the areas of Spark SQL, Delta, Memory Management, Performance tuning, Streaming, Data Science, and Data Integration areas in Spark.
- Participate in the Designated Solutions Engineer program and drive one or two of strategic customer day‑to‑day Spark and Cloud issues.
- Plan and coordinate with Account Executives, Customer Success Engineers and Resident Solution Architects for coordinating customer issues and best practices guidelines.
- Participate in screen sharing meetings, answer Slack channel conversations with internal stakeholders and customers, helping drive the major Spark issues at an individual contributor level.
- Build an internal wiki and knowledge base with technical documentation and manuals for the support team and customers. Participate in creation and maintenance of company documentation and knowledge base articles.
- Coordinate with Engineering and Backline Support teams to provide assistance in identifying, reporting product defects.
- Participate in weekend and weekday on‑call rotation and run escalations during Databricks runtime outages, incident situations, with ability to multitask and plan day‑to‑day activities and provide escalated level of support for critical customer operational issues.
- Provide best practices guidance around Spark runtime performance and usage of Spark core libraries and APIs for custom‑built solutions developed by Databricks customers.
- Be a true proponent of customer advocacy.
- Contribute in the development of tools/automation initiatives.
- Provide front‑line support on third‑party integrations with Databricks.
- Review Engineering JIRA tickets and proactively inform the support leadership team for following up on action items.
- Manage assigned Spark cases on a daily basis and adhere to committed SLAs.
- Achieve above and beyond expectations of the support organization KPIs.
- Strengthen your AWS/Azure and Databricks platform expertise through continuous learning and internal training programs.
Qualifications
- Minimum 6 years of experience designing, building, testing, and maintaining Python, Java, and Scala based applications in typical project delivery and consulting environments.
- 3+ years of hands‑on experience developing at least two of the Big Data, Hadoop, Spark, Machine Learning, Artificial Intelligence, Streaming, Kafka, Data Science, or Elasticsearch related industry use cases at production scale. Spark experience is mandatory.
- Hands‑on experience in performance tuning and troubleshooting Hive and Spark based applications at production scale.
- Proven real‑time experience in JVM and memory management techniques such as garbage collection and Heap/Thread Dump analysis.
- Working and hands‑on experience with any SQL‑based database, data warehousing/ETL technologies such as Informatica, DataStage, Oracle, Teradata, SQL Server, MySQL and SCD type use cases.
- Hands‑on experience with AWS, Azure, or GCP is preferred.
- Excellent written and oral communication skills.
- Linux/Unix administration skills is a plus.
- Working knowledge of data lakes and preferably SCD type use cases at production scale.
- Demonstrated analytical and problem‑solving skills, particularly those that apply to a "Distributed Big Data Computing" environment.