Infrastructure Engineer - Platform

Jaide Health

Toronto

On-site

CAD 150,000 - 200,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity package
Salary CAD 150k–200k
Access to sporting events
Ownership of high-impact projects
Flexible PTO
Health, dental, vision

Job summary

Peripheral is building the cloud infrastructure to power large-scale ML training and reliable inference for live events. You will work with the CTO and research/engineering teams to design scalable production systems, spanning data pipelines, training clusters, and model serving with a strong focus on reliability and cost.

The role requires deep expertise in AWS/GCP, infrastructure-as-code, containers, and orchestration, plus mentoring and documentation to support rapid growth.

Qualifications

  • Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.
  • 4+ years of experience building and operating production cloud infrastructure, ideally for ML workloads.
  • Deep expertise in AWS or GCP, with the ability to work effectively across both.
  • Experience building distributed training infrastructure, including GPU/TPU clusters, job scheduling, orchestration, and large-scale training workloads.
  • Experience deploying and operating production ML inference systems with strong requirements around reliability, latency, scalability, and cost.
  • Track record of operating production systems at meaningful user scale and maintaining reliability under growing load.
  • Strong experience with infrastructure-as-code, containers, and orchestration tools such as Terraform, Kubernetes, and Docker.
  • Strong understanding of cloud security, networking, observability, and operational best practices.
  • Comfortable working cross-functionally across research, product, and hardware teams that generate and consume data.
  • Must have the legal right to work in Canada and be willing to relocate to Toronto for an in-office role. We are unable to provide immigration sponsorship at this time.

Responsibilities

  • Design, build, and operate cloud infrastructure for large-scale ML training, including GPU/TPU compute, job orchestration, and experiment tracking.
  • Partner with research to scale foundation model training workflows, including distributed PyTorch training and efficient multi-view video data loading.
  • Build and operate production infrastructure for model deployment and serving, with a focus on reliability, latency, scalability, and cost.
  • Design reliable data pipelines for moving multi-view video, LiDAR, and other sensor data from capture through customer-facing outputs.
  • Work across research and engineering to translate infrastructure needs into scalable, self-serve systems and developer tooling.
  • Establish standards for infrastructure-as-code, CI/CD, monitoring, observability, security, and cloud operations.
  • Improve the reliability and resilience of customer-facing systems during live, time-sensitive workloads.
  • Mentor engineers and interns as the infrastructure function grows.

Skills

Systems thinking
Cross-functional collaboration
Documentation
Mentorship

Education

Bachelor’s/Master’s/PhD in CS/CE or related

Tools

Terraform
Kubernetes
Docker
AWS
GCP

Job description

WHO WE ARE:

Peripheral is developing spatial intelligence, starting in live sports and entertainment. Our models generate spatial data, used for advanced sports analytics and immersive media experiences. We’re solving key research challenges in 3D computer vision, creating the foundations for the next generation of robotic perception and embodied intelligence.

We’re backed by top investors, including Khosla Ventures, Daybreak, and Entrepreneurs First, and working with some of the biggest names in sports. Our team includes engineers and researchers from leading technology companies and research institutions, and we’re building technology at the intersection of AI, graphics, and the future of live entertainment. We’re ambitious and looking to win.

THE OPPORTUNITY:

We’re seeking an experienced Infrastructure Engineer to architect and build the cloud infrastructure that powers Peripheral, from large-scale foundation model training to reliable inference during live events.

You’ll work directly with our CTO and in collaboration with research and engineering teams to turn models and modules into scalable production systems. This is a highly architectural role that requires strong systems thinking and deep experience in DevOps and MLOps. You should understand how backend systems fail and evolve as usage grows and anticipate and design the architecture, processes and tooling to scale reliably.

You’ll help define how systems across Peripheral work together to deliver spatial experiences to customers. At the same time, you’ll build the infrastructure that keeps our research velocity high, including setting up and maintaining ML training clusters, distributed training infrastructure with PyTorch, and optimized multi-view video data loading.

The role is highly cross-functional and hands-on. You’ll build working knowledge across capture systems, robotics, foundation model training, and production inference to support teams across Peripheral while keeping customer-facing systems reliable. We’re also looking for someone who can mentor others as the team grows, maintain clear documentation in a fast-moving environment, and ship high-quality infrastructure with strong attention to detail.

WHO YOU ARE:

You’ve built and scaled production cloud infrastructure and are comfortable owning systems end to end, from large-scale training workloads to reliable model serving.

You think in terms of systems, tradeoffs, and failure modes. Cost, latency, reliability, security, and developer productivity are all first-class concerns, and you know how to balance long-term architecture with pragmatic execution.

You collaborate well across research and engineering, communicate clearly, value strong documentation, and can provide technical leadership and mentorship as the infrastructure function grows.

You’re excited to help shape Peripheral’s infrastructure culture from the ground up, establishing the tooling, standards, and practices the broader engineering organization will build on as the company scales.

WHAT YOU’LL BE DOING:
  • Design, build, and operate cloud infrastructure for large-scale ML training, including GPU/TPU compute, job orchestration, and experiment tracking.
  • Partner with research to scale foundation model training workflows, including distributed PyTorch training and efficient multi-view video data loading.
  • Build and operate production infrastructure for model deployment and serving, with a focus on reliability, latency, scalability, and cost.
  • Design reliable data pipelines for moving multi-view video, LiDAR, and other sensor data from capture through customer-facing outputs.
  • Work across research and engineering to translate infrastructure needs into scalable, self-serve systems and developer tooling.
  • Establish standards for infrastructure-as-code, CI/CD, monitoring, observability, security, and cloud operations.
  • Improve the reliability and resilience of customer-facing systems during live, time-sensitive workloads.
  • Mentor engineers and interns as the infrastructure function grows.
REQUIREMENTS
  • Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.
  • 4+ years of experience building and operating production cloud infrastructure, ideally for ML workloads.
  • Deep expertise in AWS or GCP, with the ability to work effectively across both.
  • Experience building distributed training infrastructure, including GPU/TPU clusters, job scheduling, orchestration, and large-scale training workloads.
  • Experience deploying and operating production ML inference systems with strong requirements around reliability, latency, scalability, and cost.
  • Track record of operating production systems at meaningful user scale and maintaining reliability under growing load.
  • Strong experience with infrastructure-as-code, containers, and orchestration tools such as Terraform, Kubernetes, and Docker.
  • Strong understanding of cloud security, networking, observability, and operational best practices.
  • Comfortable working cross-functionally across research, product, and hardware teams that generate and consume data.
  • Must have the legal right to work in Canada and be willing to relocate to Toronto for an in-office role. We are unable to provide immigration sponsorship at this time.
NICE TO HAVE
  • Experience with AWS and GCP services such as S3, EC2, ECR, Batch, ParallelCluster, SageMaker, GCS, Cluster Toolkit, or equivalent infrastructure.
  • Experience with ML infrastructure and tooling such as Weights & Biases, MLflow, CVAT, and Hugging Face.
  • Experience with high-throughput ingestion or streaming systems such as Kafka, Pub/Sub, or Kinesis, particularly for video, point clouds, or other large sensor data.
  • Experience with real-time video, low-latency streaming, CDNs, or large-scale data delivery systems.
  • Experience optimizing ML systems, including data loading, distributed training, GPU utilization, or custom kernels.
  • Prior ML research or ML systems experience, including publications, patents or open‑source contributions.
  • Experience as an early or founding infrastructure or platform engineer at a startup.
  • Experience mentoring junior engineers or intern
WHY YOU’LL LOVE WORKING HERE
  • Competitive equity package as an early team member.
  • Annual salary of $150K–$200K CAD plus performance bonuses, commensurate with experience.
  • Unparalleled access to premier global sporting events and iconic venues.
  • Full ownership of high-impact projects shaping the future of spatial intelligence and 3D media.
  • Flexible Paid Time Off (PTO).
  • Comprehensive health, dental, vision, and wellness benefits.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer - Reconstruction
Senior Machine Learning Engineer - Reconstruction

Jaide Health • Toronto

On-site
CAD 250,000 - 300,000
Competitive equity package
CAD salary 250k–300k + bonuses
Access to premier sporting events
+3
Software Engineer – Viewer
Software Engineer – Viewer

Peripheral • Toronto

On-site
CAD 100,000 - 150,000
Competitive equity package
Access to premier global sportingand
Software Engineer - Viewer
Software Engineer - Viewer

Jaide Health • Toronto

On-site
CAD 100,000 - 150,000
Equity
Health, dental, vision benefits
Flexible PTO
Software Engineer - Viewer
Software Engineer - Viewer

S27a • Toronto

Hybrid
CAD 100,000 - 150,000
Equity package
Salary CAD 100K–150K plus bonuses
Access to premier sporting events
+3
Software Engineer - Viewer
Software Engineer - Viewer

Socket.dev • Toronto

On-site
CAD 100,000 - 150,000
Competitive equity package
Flexible PTO
Health benefits
Embedded Systems Engineer - Capture Systems
Embedded Systems Engineer - Capture Systems

Jaide Health • Toronto

On-site
CAD 100,000 - 150,000
Equity package
Salary CAD 100k–150k (CAD)
Access to premier events
+3
Senior Robotics Engineer - Capture Systems
Senior Robotics Engineer - Capture Systems

Jaide Health • Toronto

On-site
CAD 200,000 - 250,000
Equity
Performance bonuses
Flexible PTO
+1
Machine Learning Engineering Intern - Motion Capture (Fall 2026)
Machine Learning Engineering Intern - Motion Capture (Fall 2026)

Jaide Health • Toronto

On-site
CAD 25,000 - 30,000
Flexible PTO
Machine Learning Engineering Intern - Motion Capture (Fall 2026)
Machine Learning Engineering Intern - Motion Capture (Fall 2026)

S27a • Toronto

Hybrid
CAD 25,000 - 35,000
Flexible Paid Time Off
Machine Learning Engineering Intern - Motion Capture (Fall 2026)
Machine Learning Engineering Intern - Motion Capture (Fall 2026)

Socket.dev • Toronto

On-site
CAD 20,000 - 27,000