Senior ML Infrastructure Engineer

Rebar

New York (NY)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Comprehensive medical, dental, and vision coverage
Free lunches and dinners

Job summary

Rebar is seeking a Senior ML Infrastructure Engineer to build and optimize the platform for ML engineers. The role focuses on designing APIs, integrating cloud systems, and ensuring operational excellence in a startup environment.

Ideal candidates have a solid background in backend systems, ML workflows, and cloud infrastructure, with a bachelor's degree in a relevant field. This onsite position offers a competitive salary and substantial equity.

Qualifications

  • 3+ years of experience building production backend systems.
  • Experience with cloud-based infrastructure management.
  • Proven track record operating ML inference at large scale.

Responsibilities

  • Design and build CLI, SDK, and services for the ML platform.
  • Integrate cloud and SaaS stack into a unified system.
  • Monitor and ensure reliable production model serving.

Skills

Expert-level Python
Cloud infrastructure (AWS)
Infrastructure-as-code (IaC) tooling
Managed ML inference and serving platforms
APIs and SDKs development
ML workflows understanding

Education

Bachelor’s degree in Computer Science, Electrical Engineering, or related field

Tools

Terraform
AWS SageMaker
GCP Vertex AI

Job description

Senior ML Infrastructure Engineer
Background

Rebar is building the next-generation operating system for commercial HVAC, electrical, and plumbing suppliers and subcontractors. Over the past year, our V1 quoting product has scaled to thousands of quotes completed weekly, doubled revenue in 2026, and gained adoption across many of the top suppliers in North America. Fresh off a $14M Series A backed by leading construction tech investors, we’re entering our next phase of growth — with AI at the center of everything we build next.

We’re looking for a Senior ML Infrastructure Engineer to build the platform our ML engineers depend on to rapidly iterate, experiment, and ship models — spanning feature pipelines, training infrastructure, evaluation, deployment, and monitoring. You’ll be joining a small, highly capable team focused on delivering practical, production-ready ML systems in a fast-moving startup context.

This role is ideal for someone who enjoys designing clean abstractions, integrating disparate systems into coherent platforms, and obsessing over the developer experience of the engineers they support. Our work spans the full ML lifecycle, and we’re building the platform that makes it all hang together.

Responsibilities

Platform & Developer Experience: Design and build the CLI, SDK, and services that serve as the single front door to our ML platform. Make launching a training job, tracking an experiment, or shipping a model feel like one coherent product.

Infrastructure Integration: Wire together our cloud and SaaS stack — compute providers, storage, experiment tracking, model serving — into a unified system, codified as version‑controlled infrastructure‑as‑code. Own the abstractions for compute orchestration, feature store, model registry, and model deployment.

Observability & Operations: Build cost attribution, usage dashboards, and monitoring across the platform. Surface what’s running where, catch problems early, and keep production model serving — across detection, segmentation, recognition, and LLM/VLM workloads — reliable and cost‑efficient at scale.

Collaboration and Roadmap: Work closely with ML engineers to understand their workflows, turn one‑off scripts into self‑serve platform features, and participate in architecture and roadmap decisions.

What We’re Looking For

You should feel confident designing developer‑facing APIs and SDKs, integrating disparate cloud and SaaS services into coherent systems, and obsessing over the experience of the engineers who use what you build.

We’re seeking someone with strong platform‑engineering instincts who enjoys turning fragmented workflows into products teams actually want to use. This role is a great fit if you have taste in abstractions, opinions about developer experience, and a track record of making ML or data teams meaningfully more productive.

Required Qualifications
  • Bachelor’s degree or higher in Computer Science, Electrical Engineering, or other relevant field — or equivalent industry experience.

  • 3+ years of experience building production backend systems, with significant time on internal developer platforms, ML platforms, or integration‑heavy infrastructure work.

  • Expert‑level Python; comfortable picking up other languages as the tooling demands.

  • 2+ years of experience with cloud infrastructure (AWS preferred), including IAM, networking, and cost management.

  • Proficiency with infrastructure‑as‑code (IaC) tooling such as Terraform, AWS CDK, or Pulumi for managing reproducible, version‑controlled cloud environments.

  • Hands‑on experience with managed ML inference and serving platforms such as AWS SageMaker and GCP Vertex AI.

  • A proven track record operating inference at large scale across a range of model types — detection, segmentation, recognition, and LLM/VLM workloads.

  • Experience managing a model zoo / model registry — versioning, promotion, and governance of models from experiment to production.

  • Proven ability to design clean, composable APIs and SDKs that internal users adopt willingly.

  • Deep understanding of ML related workflows and requirements.

Nice to Have
  • Experience integrating common ML tooling — experiment trackers (W&B, MLflow), feature stores, model serving frameworks — into broader platforms.

  • Experience with DAG / workflow orchestration frameworks such as Temporal, Prefect, or Apache Airflow.

  • Built a Backstage‑style internal developer portal or comparable internal platform.

  • Familiarity with GPU compute providers (AWS, Lambda Labs, CoreWeave, RunPod).

  • Some ML practitioner background — you’ve trained or deployed models yourself and understand the workflow from the user’s side.

  • Experience with deployment and monitoring pipelines for ML systems.

Compensation and Benefits
  • Salary: Competitive

  • Equity: Meaningful equity package, commensurate with experience

  • Benefits: Comprehensive medical, dental, and vision coverage

  • Perks: Free lunches and dinners provided

This is a salaried, onsite role located in New York City’s beautiful Flatiron district, just minutes away from Madison Square Park and Union Square. Working onsite offers invaluable opportunities for real‑time collaboration, creative problem‑solving, and building strong connections within our talented and dynamic team. You’ll be at the heart of our fast‑paced operations, actively contributing to a culture that values engagement, growth, and teamwork.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Data Platform
Software Engineer, Data Platform

Prudence Holdings • New York (NY)

On-site
USD 150,000 - 230,000
agentic tooling budget
lunches provided
dinners provided (after a set time)
+1
Backend Software Engineer (ML Infra)
Backend Software Engineer (ML Infra)

Rockstar • San Francisco (CA)

On-site
USD 100,000 - 130,000
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Senior ML Ops Engineer
Senior ML Ops Engineer

United States Digital Space LLC • New York (NY)

On-site
USD 180,000 - 240,000
Equity
Fully paid health coverage
Dental and vision
+7
Software Engineer, Data Platform
Software Engineer, Data Platform

Rebar • New York (NY)

On-site
USD 140,000 - 180,000
agentic tooling budget
lunches provided, dinners provided (as
great culture and office banter
Staff Software Engineer, ML/AI Platform
Staff Software Engineer, ML/AI Platform

United States Digital Space LLC • United States

On-site
USD 200,000 - 300,000
Equity
Health & wellness benefits
Backend Engineer
Backend Engineer

Space Executive • Berkeley (CA)

Remote
USD 125,000 - 225,000
Medical, dental, vision
401(k)
Unlimited PTO
+2
Machine Learning Engineer
Machine Learning Engineer

Good Inside • New York (NY)

Hybrid
USD 205,000 - 235,000
Competitive Compensation
Company Equity
Comprehensive benefits package
+2
Machine Learning Infra Engineer
Machine Learning Infra Engineer

Reducto • San Francisco (CA)

On-site
USD 120,000 - 160,000
Unlimited PTO
Free lunch
Reimbursed transportation
+3