Senior Data Engineer for Foundation Model Pipelines
Avacend Inc
San Jose (CA)
On-site
USD 150,000 - 210,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
A technology company in San Jose is looking for a skilled data engineer to design and scale data pipelines essential for model training and production systems. The ideal candidate should have over 5 years of software engineering experience, strong proficiency in Python, and hands-on expertise with Apache Spark. Responsibilities include building distributed data pipelines, optimizing performance, and collaborating with ML engineers. A strong emphasis on debugging and data quality is essential for success in this role.
Qualifications
5+ years of software engineering experience.
Strong proficiency in Python.
Experience working with time series, logs, or high-volume event data.
Responsibilities
Build and scale distributed data pipelines for large-scale time series and log data.
Design reliable, high-performance Spark/Python workflows for model training datasets.
Analyze and resolve performance bottlenecks.
Skills
Python
Apache Spark
Debugging skills
Performance optimization
Tools
Spark
Kafka
Kubernetes
Job description
A technology company in San Jose is looking for a skilled data engineer to design and scale data pipelines essential for model training and production systems. The ideal candidate should have over 5 years of software engineering experience, strong proficiency in Python, and hands-on expertise with Apache Spark. Responsibilities include building distributed data pipelines, optimizing performance, and collaborating with ML engineers. A strong emphasis on debugging and data quality is essential for success in this role.