Machine Learning Training Infrastructure Engineer

HushOne, Inc.

Kirkland (WA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Stock options
401(k) with company match
Health/dental/vision
Gym membership
Remote-friendly
AI tokens

Job summary

HushOne, Inc. seeks a senior engineer to build robust ML training systems, enabling distributed execution and reproducible experiments on local resources. You will ensure memory accounting, numerical stability, and a clear boundary between local and remote workloads.

You will plan recoverable training runs, coordinate local inference with training, and implement checkpointing with attention to data permissions and safe preemption. Join a remote-friendly team in a hands-on environment.

Qualifications

  • Practical ML training systems experience and distributed computing fundamentals.
  • Memory accounting and numerical stability in practice.
  • Understand fine-tuning vs large-scale training and design for the first.
  • Make runs reproducible, including parts people usually omit.
  • Nice to have: parameter-efficient fine-tuning methods.
  • Nice to have: checkpointing and preemption handling.
  • Nice to have: federated or on-device training.

Responsibilities

  • Plan and implement training pipelines, checkpoints, optimizer-state handling, and distributed execution.
  • Coordinate local inference with background training and enforce dataset permissions.
  • Ensure a recoverable local fine-tuning workflow and clear boundary with remote workloads.
  • Plan runs that survive power interruptions while preserving baseline comparability.

Skills

ML training systems
distributed computing
memory accounting
numerical stability
reproducible runs
fine-tuning vs large-scale training

Tools

Checkpointing
Federated learning
On-device training

Job description

You let people adapt models and run appropriately sized training jobs on resources they control, and you make those jobs reproducible, interruptible and accountable. Interruptible matters more than it sounds: this is somebody's own machine, and they get to want it back.

The work: Build training pipelines, checkpoints, optimizer-state handling, resource scheduling and distributed execution where the network supports it. Coordinate local inference with background training. Enforce dataset permissions and explicit limits before a job moves to remote compute.

What good looks like: In your first 90 days, deliver a recoverable local fine-tuning workflow and a documented boundary between supported local jobs and remote workloads.

Evidence we look for: Bring practical ML training systems experience and distributed-computing fundamentals. Understand memory accounting, numerical stability and the difference between fine-tuning and large-scale pretraining.

Required:

  • Practical ML training systems experience plus distributed-computing fundamentals
  • Memory accounting and numerical stability in practice
  • You understand the difference between fine-tuning and large-scale training, and design for the first honestly
  • You make runs reproducible, including the parts people usually leave out

Nice to have:

  • Parameter-efficient fine-tuning methods
  • Checkpointing and preemption handling
  • Federated or on-device training

The exercise: Plan a training run that must survive a power interruption while preserving the ability to compare results with the original baseline.

How we work: We work together in the office, five days a week, and you can be based at any of our garages: Kirkland Garage (Kirkland, WA); UAE Garage (Dubai, Dubai). We hire across the United States, India and the UAE. We are remote-friendly around family: if you need to work from home some days to look after the people you love, we arrange that with you one person at a time, and we encourage people to use it rather than tough it out.

Pay: We publish what we pay. Indicative ranges by market and level are on our compensation page, and they are realistic going market rates rather than headline numbers. Wherever we hire we pay at least the local market rate, and for full-time roles our floor is a living wage, never the statutory minimum. Compensation is reviewed every year and on promotion.

Equity: Every full-time teammate gets stock options. Four-year vesting with a one-year cliff, sized to role, level and impact, and confirmed in writing at offer. High performers earn refresh grants.

Bonus and commission: An annual performance bonus tied to clear company and personal goals, indicatively 10 to 20 percent of base for non-sales roles. Customer-facing roles carry on-target earnings, typically a 50/50 split of base and variable, with uncapped commission and accelerators above quota.

Health and peace of mind: Medical, dental and vision for you and your family, plus life and disability cover. We have chosen the highest plan tier available to us rather than the cheapest one that clears the bar, because the point of this is that you never have to think about it. The specific plan numbers are confirmed in your offer letter.

401(k) with company matching: A 401(k) with a company match, so the years you spend here compound into something that is yours whatever happens next. The match formula is confirmed in your offer letter.

AI tokens: A budget of AI tokens of your own, because a company that says you should own your AI cannot be the company that rations it. Use them on the work and on whatever you are curious about.

Gym, and the everyday things: A gym membership, and a corporate benefits programme with its own app, where you redeem real discounts with a long list of retailers on the ordinary purchases of a life.

Family: We are remote-friendly around your family, arranged one person at a time, and we would rather you took it than toughed it out.

Referrals: Refer someone we hire full-time who stays a year and you get $1,000 plus $10,000 in referral stock-based equity, on top of your own package.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Privacy Preserving Machine Learning Scientist
Privacy Preserving Machine Learning Scientist

HushOne, Inc. • Kirkland (WA)

Hybrid
USD 120,000 - 180,000
Stock options
401(k) with company match
Gym membership
+2
Efficient Models Research Scientist
Efficient Models Research Scientist

HushOne, Inc. • Kirkland (WA)

Hybrid
USD 120,000 - 180,000
Stock options
401(k) with company match
Health, dental and vision
+2
Distributed Systems Engineer
Distributed Systems Engineer

HushOne, Inc. • Kirkland (WA)

Hybrid
USD 140,000 - 210,000
Stock options
Annual bonus
Health insurance
+2
Chief Scientist
Chief Scientist

HushOne, Inc. • Kirkland (WA)

On-site
USD 180,000 - 280,000
Stock options
401(k) match
Health, dental, vision
+2
ML Data and Experiment Infrastructure Engineer
ML Data and Experiment Infrastructure Engineer

HushOne, Inc. • Kirkland (WA)

On-site
USD 150,000 - 210,000
Stock options
401(k) with company matching
Health, dental and vision coverage
+4
Formal Methods and Verified Software Engineer
Formal Methods and Verified Software Engineer

HushOne, Inc. • Kirkland (WA)

On-site
USD 140,000 - 210,000
Stock options
Four-year vesting
Annual performance bonus
+7
Performance and Energy Benchmark Engineer
Performance and Energy Benchmark Engineer

HushOne, Inc. • Kirkland (WA)

Hybrid
USD 120,000 - 180,000
Stock options
401(k) with company match
Health insurance
+2
Android and Edge Devices Engineer
Android and Edge Devices Engineer

HushOne, Inc. • Kirkland (WA)

Hybrid
USD 110,000 - 140,000
Stock options
401(k) with company match
Health, dental, vision
Open Source and Developer Relations Engineer
Open Source and Developer Relations Engineer

HushOne, Inc. • Kirkland (WA)

Hybrid
USD 120,000 - 180,000
Stock options
401(k) with match
Medical, dental, and vision
+1
CTO and Chief Systems Architect
CTO and Chief Systems Architect

HushOne, Inc. • Kirkland (WA)

On-site
USD 180,000 - 260,000
Stock options
Vesting schedule