Software Engineer, ML Infrastructure

rebar ukraine

New York (NY)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Agentic tooling budget
Lunches provided
Dinners provided (after a set time)
Great culture and office banter

Job summary

Rebar in New York City is seeking a Senior ML Infrastructure Engineer to own and expand the ML infrastructure platform used by ML engineers to rapidly iterate, experiment, and ship models. You will focus on feature pipelines, training infra, deployment, monitoring, and performance tuning on GPU architectures, with a strong emphasis on developer experience and production readiness.

Join a lean, fast-moving startup that values ownership, collaboration, and practical solutions for real-world

Qualifications

  • 6+ years of experience building production backend systems.
  • 2-3+ years of experience with ML platforms or internal developer platforms is a plus.
  • Strong Python and PyTorch experience including profiling and debugging.
  • 3+ years of cloud infrastructure experience (AWS preferred).
  • Hands-on GPU performance work and inference optimization in production.
  • Proven track record operating inference at large scale across multiple model types including LLM/VLM.

Responsibilities

  • Platform and Developer Experience: build services that serve as the single front door to the ML platform.
  • Model and Data Governance: instrument active learning pipelines, retrainings, model registry and governance.
  • Infrastructure Consolidation and Integration: smooth GPU inference and cloud stack seams.
  • Observability and Operations: extend observability with metrics for GPUUtil and inference flow.
  • Collaboration and Roadmap: work with ML engineers to turn scripts into self-serve platform features and shape architecture decisions.

Skills

Python
PyTorch
Backend systems
Platform engineering
GPU performance
Inference optimization
Cloud infrastructure

Tools

AWS

Job description

Senior ML Infrastructure Engineer
About Rebar

Rebar is building the AI operating system for commercial HVAC, Electrical, and Plumbing.

Over the past year our quoting platform has processed tens of thousands of projects across North America and we’re continuing that growth. Our customers include many of the top firms in the industry. Some of these companies are running billion dollar construction projects on workflows that still look like it's 1985.

Construction is 10% of GDP and still massively underserved by software. We are changing that.

We recently raised a $14M Series A from leading construction tech investors and are entering our next phase of growth. We are building a set of AI native products that will define how this industry operates.

About the Engineering Org

The role of the engineer is rapidly changing. We're aware and we're being very intentional of ensuring we adapt with it. We are fostering an engineering culture of growth and development. We strongly emphasize care of craft and winning together. To echo our values, everyone operates like an owner, we find a way, and we win together.

About the Role

We're hiring a Software Engineer to own and expand the ML infrastructure platform our ML engineers depend on to rapidly iterate, experiment, and ship models. This work will span feature pipelines, training infra, evaluation, deployment, and monitoring. You should be well versed with GPU architecture and performance focused - understanding what is blocking the engineers as well as our MFU. You'll be joining a small group of talented engineers focused on delivering practical, production-ready ML systems in a fast-moving startup context.

The role is ideal for someone who is obsessive over the developer experience of the engineers they support, is interested in training, and meticulous over performance. Our team is still (somewhat) lean and engineers wear many hats. Your purview will span the whole ML lifecycle.

Responsibilities
  • Platform and Developer Experience

    • You'll help build services that act as the single front door to our ML platform

  • Model and Data Governance

    • Explore and help instrument our active learning pipelines, retrainings, model registry, and more

  • Infrastructure Consolidation and Integration

    • Smooth the seams for GPU inference and our cloud stack

  • Observability and Operations

    • We utilize DataDog for our observability. You should be comfortable extending DD and Modal support to have clear GPU util metrics as well as clean insights into our entire inference flow

  • Collaboration and Roadmap

    • You will work closely with ML engineers to understand their workflows, turn one-off scripts into self-serve platform features, and participate in architecture and roadmap decisions.

What We're Looking For

You should feel confident designing developer-facing APIs and SDKs, integrating disparate cloud and SaaS services into coherent systems, and obsessing over the experience of the engineers who use what you build.

We're seeking someone with strong platform-engineering instincts who enjoys turning fragmented workflows into products teams actually want to use. This role is a great fit if you have taste in abstractions, opinions about developer experience, and a track record of making ML or data teams meaningfully more productive.

Qualifications
  • 6+ years of experience building production backend systems, with significant time on internal developer platforms, ML platforms, or integration-heavy infrastructure work.

  • Strong Python + PyTorch, including profiling and debugging below the model code

  • 3+ years of experience with cloud infrastructure (AWS preferred)

  • Hands on GPU performance and inference work. Ideally you have optimizations in production: batching, mixed precision, quantization, ONNX Runtime/TensorRT or torch.compile

  • A proven track record operating inference at large scale across a range of model types - detection, segmentation, recognition, and LLM/VLM workloads.

  • Experience managing a model zoo / model registry - versioning, promotion, and governance of models from experiment to production.

Nice to Have
  • Experience integrating common ML tooling - experiment trackers (W&B, MLflow), feature stores, model serving frameworks - into broader platforms.

  • Experience with DAG / workflow orchestration frameworks such as Temporal, Prefect, or Apache Airflow.

  • Built a Backstage-style internal developer portal or comparable internal platform.

  • Familiarity with GPU compute providers (AWS, Lambda Labs, CoreWeave, RunPod).

  • Some ML practitioner background — you've trained or deployed models yourself and understand the workflow from the user's side.

  • Experience with deployment and monitoring pipelines for ML systems.

Compensation and Benefits
  • Salary: Competitive base salary

  • Equity: Meaningful equity package, commensurate with experience

  • Benefits: Comprehensive medical, dental, and vision coverage

  • Perks:

    • agentic tooling budget

    • lunches provided, dinners provided (after a set time)

    • great culture and office banter

This is a salaried, onsite role located in New York City's Flatiron district. We are still a startup! We love working onsite together and believe strongly that this gives us for creative problem-solving, and building strong connections. You'll be at the heart of our fast-paced operations, actively contributing to a culture that values engagement, growth, and teamwork.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, Applied AI
Software Engineer, Applied AI

Rebar • New York (NY)

On-site
USD 140,000 - 210,000
Agentic tooling budget
Lunches provided
Dinners provided
+1
Software Engineer, Data Platform
Software Engineer, Data Platform

rebar ukraine • New York (NY)

On-site
USD 140,000 - 210,000
Agentic tooling budget
Lunches provided
Dinners provided
+1
Software Engineer, Applied AI
Software Engineer, Applied AI

rebar ukraine • New York (NY)

On-site
USD 150,000 - 190,000
Agentic tooling budget
Lunches provided
Dinners provided
+1
Software Engineer, Data Platform
Software Engineer, Data Platform

Prudence Holdings • New York (NY)

On-site
USD 150,000 - 230,000
agentic tooling budget
lunches provided
dinners provided (after a set time)
+1
Software Engineer, Data Platform
Software Engineer, Data Platform

Rebar • New York (NY)

On-site
USD 140,000 - 180,000
agentic tooling budget
lunches provided, dinners provided (as
great culture and office banter
Software Engineer, Data Platform
Software Engineer, Data Platform

Unchain Data • New York (NY), Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive base salary
Meaningful equity package
Comprehensive medical, dental, and vis
+3
Software Engineer, Product
Software Engineer, Product

rebar ukraine • New York (NY)

On-site
USD 140,000 - 190,000
agentic tooling budget
lunches provided
dinners provided (after a set time)
+1
Machine Learning Infrastructure Tech Lead
Machine Learning Infrastructure Tech Lead

Reducto • Santa Fe (NM)

On-site
USD 190,000 - 230,000
Unlimited PTO
Daily Lunch
Commuter Reimbursement
+3
Machine Learning Infrastructure Tech Lead
Machine Learning Infrastructure Tech Lead

Reducto • San Francisco (CA)

On-site
USD 180,000 - 260,000
Unlimited PTO
Daily Lunch
Commuter Reimbursement
+3
Founding Engineer - ML Infrastructure
Founding Engineer - ML Infrastructure

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k)
Paid time off
+2