Member of Technical Staff, Data & Training Infrastructure

Arena Physica

New York (NY)

On-site

USD 150,000 - 230,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Premium medical, vision, dental
401(k) plan
Unlimited PTO
Relocation support
Lunch from local restaurants

Job summary

Arena Physica is building a team to advance electromagnetic superintelligence. You will join the Platform team as a Machine Learning Engineer focused on Data & Training Infrastructure, shaping how simulation data, measurements, and expert workflows feed scalable training data for the first electromagnetic foundation model.

Work closely with applied researchers, electrical engineers, and product engineers to orchestrate solver farms, data corpora, and fast, reproducible model development in a

Qualifications

  • 5+ years of software engineering or ML infrastructure experience at a venture-backed startup or high-performing research org.

Responsibilities

  • Scale electromagnetic training data generation infrastructure.
  • Design dataset APIs, schemas, lineage systems, and quality gates.
  • Own high-throughput training pipelines for large models.
  • Collaborate with researchers to translate model needs into data programs.
  • Make ML experiments reproducible with dataset snapshots and traceability.
  • Integrate EM solvers and measurements into production-grade pipelines.

Skills

Python
ML infrastructure
Distributed systems
Data processing
Experimentation & reproducibility
Communication
Ownership mindset
Ambiguity handling
Collaboration
Hardware/EM knowledge (preferred)

Tools

PyTorch
JAX
Slurm
Ray
Kubernetes
AWS Batch
Elastic Fabric Adapter

Job description

Arena Physica is on a mission to accelerate hardware innovation that powers human progress. Our name is inspired by Theodore Roosevelt's 'Citizenship in a Republic' speech. To us, entering the Arena means committing fully and accepting the risk of failure in pursuit of an audacious, worthy cause. We believe the future belongs to those brave enough to build it.

Our team of 50 combines AI engineering and applied physics expertise with deep experience in enterprise deployments. We're headquartered in NYC with presences in San Francisco and Los Angeles, backed by ~$90M from Initialized, Founders Fund, Goldcrest Capital, Fifth Down Capital, and Shield Capital.

If you're ready to do the most important work of your career, join us in the Arena.

Who we are

Arena Physica is on a mission to accelerate hardware innovation that powers human progress. Our name is inspired by Theodore Roosevelt's 'Citizenship in a Republic' speech. To us, entering the Arena means committing fully and accepting the risk of failure in pursuit of an audacious, worthy cause. We believe the future belongs to those brave enough to build it.

Our team of 50 combines AI engineering and applied physics expertise with deep experience in enterprise deployments. We're headquartered in NYC with presences in San Francisco and Los Angeles, backed by ~$90M from Initialized, Founders Fund, Goldcrest Capital, Fifth Down Capital, and Shield Capital.

If you're ready to do the most important work of your career, join us in the Arena.

What we do

At Arena Physica, we're building electromagnetic superintelligence. Our AI platform Atlas operationalizes physics-grounded intelligence to verify, debug, and optimize hardware across its lifecycle. Atlas is already trusted globally by the world's most advanced hardware companies, including AMD, Anduril, and Bausch & Lomb, for applications across R&D, integration testing, production assembly, and field repair.

About the role

As a Machine Learning Engineer focused on Data & Training Infrastructure, you will join Arena's Platform team to build the data generation and training substrate behind the first electromagnetic foundation model. You will design the systems that turn simulation, measurement, expert workflows, and customer-grounded hardware problems into high-quality training data at scale.

This role sits at the intersection of large-scale ML infrastructure, scientific computing, and applied physics. You will work closely with applied researchers, electrical engineers, and product engineers to orchestrate solver farms, standardize multimodal data corpora, improve training throughput, and make foundation model development faster, more reproducible, and more reliable.

How you will contribute
  • Scale electromagnetic training data generation - Build infrastructure for generating, validating, versioning, and replaying synthetic and physically grounded EM datasets across solver farms, hardware-in-the-loop campaigns, and customer-inspired design spaces.
  • Build the data foundation for Heaviside - Design dataset APIs, schemas, lineage systems, and quality gates for simulations, measurements, design files, S-parameters, meshes, fields, telemetry, documents, and expert annotations.
  • Own high-throughput training pipelines - Develop distributed data loading, preprocessing, sharding, caching, and observability systems that keep large model training jobs performant across GPU and HPC environments.
  • Partner with applied research - Work with researchers building electromagnetic foundation models to translate coverage targets, evaluation failures, and model needs into new data generation programs and infrastructure capabilities.
  • Make ML experimentation reproducible - Build tooling for dataset snapshots, experiment traceability, regression detection, and training run analysis so Arena can move quickly without losing scientific rigor.
  • Operationalize physics-grounded intelligence - Connect EM solvers, lab measurements, and platform services into production-grade pipelines that let Atlas learn from the real workflows hardware engineers use every day.
  • Travel domestically and internationally (10-20% of your time).
  • Work in person at Arena's NYC HQ when not traveling.
You have
  • 5+ years of software engineering or ML infrastructure experience at a venture-backed startup, top technology company, frontier AI lab, or high-performing research organization
  • Strong experience building distributed systems, data platforms, or training infrastructure for large-scale machine learning workloads
  • Proficiency with Python and at least one modern ML framework such as PyTorch or JAX. Experience with large-scale data processing, storage, orchestration, and observability systems across cloud or HPC environments
  • Deep practical judgment around reliability, throughput, reproducibility, and developer ergonomics for research and production systems
  • Comfort working with ambiguous research requirements and turning them into robust, scalable platform primitives
  • Strong communication skills and the ability to collaborate across ML researchers, electrical engineers, software engineers, and customer-facing teams
  • Self-directed ownership mindset and excitement for building foundational infrastructure in a fast-moving environment
  • [Preferred] Experience with scientific ML, neural operators, physics simulation, EDA tooling, EM solvers, or hardware design workflows
  • [Preferred] Experience with Slurm, Ray, Kubernetes, AWS Batch, Elastic Fabric Adapter, distributed filesystems, or high-performance data loading for GPU clusters.
  • [Preferred] Familiarity with synthetic data generation, active learning, data quality evaluation, or model-driven dataset curation
  • [Preferred] An interest in electromagnetic systems, RF, signal integrity, power integrity, or the application of AI to real-world engineering work
Benefits & Perks Include:
  • 100% of the monthly premiums covered with Aetna medical vision, and dental insurance for you and your dependents
  • 401(k) Retirement Plan
  • Unlimited PTO
  • Lunch every day from local restaurants via Sharebite
  • Relocation support provided
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Product
Member of Technical Staff, Product

Arena Physica • New York (NY)

Hybrid
USD 150,000 - 210,000
Health insurance (Aetna)
401(k) Retirement Plan
Unlimited PTO
+2
Electrical Engineer, RF/EM
Electrical Engineer, RF/EM

Arena Physica • New York (NY)

On-site
USD 120,000 - 170,000
Medical insurance
Vision & dental
401(k) retirement plan
+3
Deployment Strategist, Integrated Systems
Deployment Strategist, Integrated Systems

Arena Physica • New York (NY)

On-site
USD 120,000 - 160,000
100% of the monthly premium for Aetna
401(k) Retirement Plan
Unlimited PTO
+2
Strategy & Business Operations
Strategy & Business Operations

Arena Physica • New York (NY)

On-site
USD 120,000 - 180,000
Healthcare benefits
401(k)
Unlimited PTO
+2
ML Infrastructure Engineer — Data & Training
ML Infrastructure Engineer — Data & Training

Arena Physica • New York (NY)

On-site
USD 150,000 - 230,000
Premium medical, vision, dental
401(k) plan
Unlimited PTO
+2
Infrastructure Engineer, TL
Infrastructure Engineer, TL

Arena Intelligence, Inc. • United States

On-site
USD 140,000 - 210,000
Competitive compensation and equity
Health benefits
Cutting-edge AI work
+1
Infrastructure Engineer – TL
Infrastructure Engineer – TL

Arena • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive compensation
Health and wellness benefits
Cutting-edge AI projects
+1
Staff Software Engineer, Product
Staff Software Engineer, Product

arena • San Francisco (CA)

On-site
USD 130,000 - 160,000
Competitive compensation
Health and wellness benefits
Opportunity to work on cutting-edge AI
Site Reliability Engineer
Site Reliability Engineer

Arena • San Francisco (CA)

On-site
USD 180,000 - 280,000
Equity
Health benefits
Cutting-edge AI work
+1
Site Reliability Engineer
Site Reliability Engineer

Arena Intelligence, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive compensation
Equity
Health benefits
+1