Infrastructure Tech Lead

Omnifold

San Francisco (CA)

On-site

USD 150,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Omnifold is seeking an Infrastructure Tech Lead / Principal Engineer based in San Francisco. You'll own the systems critical for AI model training and deployment while ensuring security and efficient resource management.

The ideal candidate has significant experience in cloud computing, specifically with GPU workloads, and a strong computer science background. This role requires a proven ability to lead technology initiatives in a fast-paced environment to optimize processes and monitor system performance.

Qualifications

  • Experience with cloud computing, especially GPU workloads.
  • Familiarity with ML workflows preferred.
  • At least 10 years of experience in a relevant field.
  • 3+ years in a tech lead role.

Responsibilities

  • Own reliable processes for getting models and services into production.
  • Manage data isolation and product security.
  • Oversee cloud resource management and cost optimization.
  • Build visibility into system performance through monitoring.
  • Implement ETL pipelines and lifecycle management for models.
  • Create automated testing infrastructure.

Skills

Cloud computing
CI/CD infrastructure-as-code
Security fundamentals
Python
Machine Learning workflows

Education

Strong Computer Science background

Job description

Infrastructure Tech Lead / Principal Engineer

Omnifold trains custom AI models that help planners forecast the future. We are hiring our first infrastructure tech lead, who will own the systems that make everything else possible.

What makes this job interesting:
  • We train a unique model for each customer, which means model training and inference work differently here than at any other company. You’ll never get more reps building model training infrastructure!
  • Our team has very fast iteration speed but needs robust monitoring to pick up signal on user patterns. This is especially important as our application interface for AI-driven forecasting is unique on the market.
What you’ll own
  • Deployment: Reliable processes for getting models and services into production
  • Security: Data isolation between customers, product security, infrastructure hardening (SOC2 compliance and beyond)
  • Cloud resource management: GPU allocation, instance sizing, cost optimization
  • Monitoring and logging: Visibility into what's running, what's failing, and why
  • Data and ML ops: ETL pipelines from varied customer data sources, model versioning and lifecycle management
  • Automated testing: Building the test infrastructure that lets us ship with confidence
What we’re looking for
  • Experience with cloud computing (especially GPU workloads), CI/CD infrastructure-as-code. We run on AWS
  • Familiarity with or interest in ML workflows
  • Security fundamentals: encryption, access controls, compliance basics
  • Python proficiency
  • Ideally ~10 years of experience, including startup experience, with at least 3 years in a tech lead role. 5+ years in infrastructure, DevOps, or platform engineering roles
  • Must have a strong Computer Science background

Location: San Francisco (in-person, 5 days per week)

Omnifold’s Mission

Every bad forecast has a physical consequence. Unnecessary goods are manufactured, shipped, and stored. Emergency air freight is needed for misallocated products. Poor production planning means workers show up with nothing to do, or work frantic overtime. Inefficiency is everywhere.

Our mission is to eliminate waste and accelerate growth for every company with physical products.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Tech Lead — AI/ML Ops & Cloud Security
Infrastructure Tech Lead — AI/ML Ops & Cloud Security

Omnifold • San Francisco (CA)

On-site
USD 150,000 - 200,000
MTS - ML Research Engineer
MTS - ML Research Engineer

Omnifold • San Francisco (CA)

On-site
USD 120,000 - 160,000
Director of Applied AI
Director of Applied AI

Socket.dev • San Francisco (CA)

On-site
USD 120,000 - 190,000
Lead Data and ML Infrastructure Engineer
Lead Data and ML Infrastructure Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 260,000
Competitive compensation
Meaningful equity
Autonomy and ownership
Infrastructure Engineer – San Francisco
Infrastructure Engineer – San Francisco

Syndesus • San Francisco (CA)

On-site
USD 150,000 - 300,000
Equity
On-site (SF/NYC)
Software Engineer, Cloud Infrastructure
Software Engineer, Cloud Infrastructure

DatologyAI • Redwood City (CA)

On-site
USD 180,000 - 250,000
100% covered health benefits
401(k) plan with 4% match
Unlimited PTO
+3
Member of Technical Staff: Infrastructure at Observable Intuition
Member of Technical Staff: Infrastructure at Observable Intuition

Jack & Jill • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Founding engineer opportunity
Director of Applied AI
Director of Applied AI

Omnifold • San Francisco (CA)

On-site
USD 180,000 - 240,000
Infrastructure Engineer
Infrastructure Engineer

Overland AI • Seattle (WA)

On-site
USD 130,000 - 225,000
Competitive salary: $130K – $225K annually
Equity compensation
Best-in-class healthcare, dental, and vision plans
+3
Software Engineer, Infrastructure
Software Engineer, Infrastructure

DatologyAI • Redwood City (CA)

On-site
USD 180,000 - 250,000
100% covered health benefits
401(k) plan with 4% company match
Unlimited PTO
+3