Data Infrastructure Engineer - Scalable ML Pipelines
Alljoined
San Francisco (CA)
On-site
USD 140,000 - 180,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Competitive equity compensation
Options for housing support
Visa sponsorship
3% 401k matching
Health insurance
Job summary
A technology startup in San Francisco is seeking a Data Infrastructure Engineer to build backend systems and manage the data lifecycle. You will be responsible for creating pipelines that process large datasets and ensuring the infrastructure supports world-class research efforts. Candidates should have over 3 years of production experience, expertise in languages like Python and C++, and a background in managing compute clusters. Competitive equity compensation and additional benefits are offered, including housing support and visa sponsorship.
Qualifications
3+ years of production software engineering experience with deep expertise in systems-level architecture.
Experience building and maintaining ETL pipelines for terabytes of unstructured data.
Comfortable managing local compute clusters and high-speed networking for ML workloads.
Responsibilities
Own the entire data lifecycle from building pipelines to managing compute clusters.
Work alongside researchers to ensure high-throughput and low-latency pipelines.
Define storage topologies and sync datasets between local servers and the cloud.
Skills
Production software engineering experience
Expertise in Python
Expertise in Rust
Expertise in C++
Expertise in Go
High-performance ETL pipelines
Architecting compute clusters
Managing databases
Handling concurrent data streams
Optimizing I/O bound operations
Tools
TimescaleDB
ClickHouse
AWS
GCP
Azure
FFmpeg
GStreamer
OpenCV
Job description
A technology startup in San Francisco is seeking a Data Infrastructure Engineer to build backend systems and manage the data lifecycle. You will be responsible for creating pipelines that process large datasets and ensuring the infrastructure supports world-class research efforts. Candidates should have over 3 years of production experience, expertise in languages like Python and C++, and a background in managing compute clusters. Competitive equity compensation and additional benefits are offered, including housing support and visa sponsorship.