Software Engineer, ML Infrastructure

Prudence Holdings

New York (NY)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity package
Medical, dental, and vision coverage
Agentic tooling budget
Luncheons and dinners provided

Job summary

Rebar, a startup reshaping AI-driven construction workflows, is seeking a Senior ML Infrastructure Engineer to own and expand the ML infrastructure platform used by our ML engineers. You will build scalable services, manage model registries, and drive production-grade pipelines for training, evaluation, deployment, and monitoring.

You should be obsessed with developer experience, have strong Python/PyTorch skills, and bring a track record of optimizing ML workloads at scale.

Qualifications

  • 6+ years of experience building production backend systems, with significant time on internal developer platforms or ML platforms.
  • Strong Python and PyTorch, including profiling and debugging below the model code.
  • 3+ years of experience with cloud infrastructure (AWS preferred).
  • Hands-on GPU performance and inference work with production optimizations (batching, mixed precision, quantization).
  • Experience managing a model zoo / model registry - versioning, promotion, governance from experiment to production.

Responsibilities

  • Platform and Developer Experience: build services as the single front door to our ML platform.
  • Model and Data Governance: instrument active learning pipelines, retrainings, model registry.
  • Infrastructure Consolidation and Integration: smooth GPU inference and cloud stack.
  • Observability and Operations: extend DataDog and ensure clear GPU util metrics.
  • Collaboration and Roadmap: work with ML engineers to turn scripts into self-serve platform features and guide architecture.

Skills

Python
PyTorch
AWS
GPU performance
Model registry

Tools

ONNX Runtime
TensorRT
Torch.compile

Job description

Senior ML Infrastructure Engineer
About Rebar

Rebar is building the AI operating system for commercial HVAC, Electrical, and Plumbing.

Over the past year our quoting platform has processed tens of thousands of projects across North America and we’re continuing that growth. Our customers include many of the top firms in the industry. Some of these companies are running billion dollar construction projects on workflows that still look like it's 1985.

Construction is 10% of GDP and still massively underserved by software. We are changing that.

We recently raised a $14M Series A from leading construction tech investors and are entering our next phase of growth. We are building a set of AI native products that will define how this industry operates.

About the Engineering Org

The role of the engineer is rapidly changing. We're aware and we're being very intentional of ensuring we adapt with it. We are fostering an engineering culture of growth and development. We strongly emphasize care of craft and winning together. To echo our values, everyone operates like an owner, we find a way, and we win together.

About the Role

We're hiring a Software Engineer to own and expand the ML infrastructure platform our ML engineers depend on to rapidly iterate, experiment, and ship models. This work will span feature pipelines, training infra, evaluation, deployment, and monitoring. You should be well versed with GPU architecture and performance focused - understanding what is blocking the engineers as well as our MFU. You'll be joining a small group of talented engineers focused on delivering practical, production-ready ML systems in a fast-moving startup context.

The role is ideal for someone who is obsessive over the developer experience of the engineers they support, is interested in training, and meticulous over performance. Our team is still (somewhat) lean and engineers wear many hats. Your purview will span the whole ML lifecycle.

Responsibilities
  • Platform and Developer Experience

    • You'll help build services that act as the single front door to our ML platform

  • Model and Data Governance

    • Explore and help instrument our active learning pipelines, retrainings, model registry, and more

  • Infrastructure Consolidation and Integration

    • Smooth the seams for GPU inference and our cloud stack

  • Observability and Operations

    • We utilize DataDog for our observability. You should be comfortable extending DD and Modal support to have clear GPU util metrics as well as clean insights into our entire inference flow

  • Collaboration and Roadmap

    • You will work closely with ML engineers to understand their workflows, turn one-off scripts into self-serve platform features, and participate in architecture and roadmap decisions.

What We're Looking For

You should feel confident designing developer-facing APIs and SDKs, integrating disparate cloud and SaaS services into coherent systems, and obsessing over the experience of the engineers who use what you build.

We're seeking someone with strong platform-engineering instincts who enjoys turning fragmented workflows into products teams actually want to use. This role is a great fit if you have taste in abstractions, opinions about developer experience, and a track record of making ML or data teams meaningfully more productive.

Qualifications
  • 6+ years of experience building production backend systems, with significant time on internal developer platforms, ML platforms, or integration-heavy infrastructure work.

  • Strong Python + PyTorch, including profiling and debugging below the model code

  • 3+ years of experience with cloud infrastructure (AWS preferred)

  • Hands on GPU performance and inference work. Ideally you have optimizations in production: batching, mixed precision, quantization, ONNX Runtime/TensorRT or torch.compile

  • A proven track record operating inference at large scale across a range of model types - detection, segmentation, recognition, and LLM/VLM workloads.

  • Experience managing a model zoo / model registry - versioning, promotion, and governance of models from experiment to production.

Nice to Have
  • Experience integrating common ML tooling - experiment trackers (W&B, MLflow), feature stores, model serving frameworks - into broader platforms.

  • Experience with DAG / workflow orchestration frameworks such as Temporal, Prefect, or Apache Airflow.

  • Built a Backstage-style internal developer portal or comparable internal platform.

  • Familiarity with GPU compute providers (AWS, Lambda Labs, CoreWeave, RunPod).

  • Some ML practitioner background — you've trained or deployed models yourself and understand the workflow from the user's side.

  • Experience with deployment and monitoring pipelines for ML systems.

Compensation and Benefits
  • Salary: Competitive base salary

  • Equity: Meaningful equity package, commensurate with experience

  • Benefits: Comprehensive medical, dental, and vision coverage

  • Perks:

    • agentic tooling budget

    • lunches provided, dinners provided (after a set time)

    • great culture and office banter

This is a salaried, onsite role located in New York City's Flatiron district. We are still a startup! We love working onsite together and believe strongly that this gives us for creative problem-solving, and building strong connections. You'll be at the heart of our fast-paced operations, actively contributing to a culture that values engagement, growth, and teamwork.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infrastructure Engineer
Senior ML Infrastructure Engineer

Rebar • New York (NY)

On-site
USD 120,000 - 160,000
Comprehensive medical, dental, and vision coverage
Free lunches and dinners
Software Engineer, Data Platform
Software Engineer, Data Platform

Prudence Holdings • New York (NY)

On-site
USD 150,000 - 230,000
agentic tooling budget
lunches provided
dinners provided (after a set time)
+1
Software Engineer, Data Platform
Software Engineer, Data Platform

Rebar • New York (NY)

On-site
USD 140,000 - 180,000
agentic tooling budget
lunches provided, dinners provided (as
great culture and office banter
Software Engineer, Data Platform
Software Engineer, Data Platform

Unchain Data • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive base salary
Meaningful equity package
Comprehensive medical, dental, and vis
+3
Machine Learning Infra Engineer
Machine Learning Infra Engineer

Reducto • San Francisco (CA)

On-site
USD 120,000 - 160,000
Unlimited PTO
Free lunch
Reimbursed transportation
+3
Senior Machine Learning Engineer, AI Infra
Senior Machine Learning Engineer, AI Infra

United States Digital Space LLC • Bellevue (CA)

On-site
USD 209,000 - 245,000
Health insurance
Equity ownership
401k matching
+2
Machine Learning Infrastructure Tech Lead
Machine Learning Infrastructure Tech Lead

Reducto • Santa Fe (NM)

On-site
USD 190,000 - 230,000
Unlimited PTO
Daily Lunch
Commuter Reimbursement
+3
Senior ML Platform Engineer
Senior ML Platform Engineer

Rebar • New York (NY)

On-site
USD 120,000 - 160,000
Comprehensive medical, dental, and vision coverage
Free lunches and dinners
Machine Learning Infrastructure Tech Lead
Machine Learning Infrastructure Tech Lead

Reducto • San Francisco (CA)

On-site
USD 180,000 - 260,000
Unlimited PTO
Daily Lunch
Commuter Reimbursement
+3
GTM/Growth Engineer
GTM/Growth Engineer

Socket.dev • New York (NY)

On-site
USD 120,000 - 180,000
Salary: competitive
Equity
Medical, dental, vision
+1