A complete application in a minute — tailored resume and cover letter, ready to send.
Get past ATS filters
Benefits offered by this job
Medical, dental, vision insurance
401(k)
Equity potential
Onsite Pittsburgh office
Job summary
Qualifications
3+ years of experience in Cloud infrastructure, MLOps, and data engineering at scale.
Hands-on experience designing systems for automated model checkpointing, model registries, and metadata management (e.g., MLflow, Kubeflow etc).
Strong experience with Cloud workflow orchestration tools (Argo Workflows, Airflow, Prefect, or similar) and cloud storage architectures.
Proven track record building self-service ML platforms, pipeline abstraction layers, or automated developer workflows.
Experience with working with embedded engineers.
Responsibilities
Design, build, and maintain scalable ML infrastructure that abstracts away underlying storage, pipeline management, and repetitive environment setup tasks.
Implement robust systems for automated model checkpointing, persistent metadata management, and experiment tracking across distributed training runs.
Create self‑service ML workflows and tooling that empower ML engineers and data scientists to focus on core logic, model architecture, and validation.
Build and maintain automated ETL and data ingestion pipelines that stream and transform raw robot telemetry into clean datasets for training and analytics.
Contribute to infrastructure‑as‑code (Terraform) and CI/CD automation for model deployment, data processing, and cloud services.
Partner with cloud and embedded engineers to streamline model deployment to fleets of robots in the field and participate in on‑call rotations for platform reliability.
Monitoring & Cost: Track model drift, system throughput, and optimize cloud compute costs.