Senior ML Quantization Engineer for Optical AI Compute

Neurophos, Inc.

Sunnyvale, Northern (TX, KY)

Hybrid

USD 150,000 - 240,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health coverage
HSA contributions
Unlimited PTO
401(k) matching
Stock options
Dental, Vision, Life coverage

Job summary

Neurophos, Inc. seeks an experienced ML engineer to advance post-training quantization for language and diffusion models on our optical inference engines. You will bridge ML research and hardware, enabling customers to deploy AI workloads on Neurophos hardware.

Ideal candidates have deep expertise in model quantization, transformer models, and numerical optimization, with hands-on PyTorch experience and a track record of rigorous experiments. This role is onsite in Austin or Sunnyvale.

Qualifications

  • PhD or equivalent research experience in ML, optimization, numerical analysis, or CS.
  • 5+ years in ML engineering with at least 3 years focused on model optimization and deployment.
  • Experience in neural network quantization, model compression, or efficient inference.
  • Strong knowledge of numerical linear algebra, including matrix factorizations and iterative methods.
  • Experience with non-convex optimization, discrete optimization, or second-order methods.
  • Strong proficiency in PyTorch; familiarity with JAX, Triton, and TensorFlow.
  • Hands-on experience with transformer architectures, LLMs, and diffusion models.
  • Experience designing controlled numerical experiments and distinguishing algorithmic improvements from artifacts.
  • Strong written communication and collaboration skills.

Responsibilities

  • Develop hardware-aware post-training methods for full model quantization.
  • Investigate preconditioning and formulate quantization as non-convex, discrete, constrained, or second-order optimization.
  • Contribute to refining Neurophos's quantization strategy.
  • Design controlled numerical experiments to understand potential improvements and secondary effects.
  • Build research-quality implementations and reproducible experiment harnesses for testing candidate methods.
  • Adapt models from open-source repositories and customer private models.
  • Work with models in PyTorch, JAX, Triton, and TensorFlow.
  • Design and execute re-quantization, retraining, and other model adaptation techniques.
  • Optimize GEMM operations for high-throughput execution.
  • Collaborate with hardware, software, and architecture teams to co-optimize model architectures for optical compute characteristics.
  • Publish research papers on novel optimization techniques and methodologies, with IP protection.

Skills

Machine learning engineering
Model quantization
Numerical optimization
Transformer models
PyTorch
JAX
Triton
LLMs
Diffusion models
Experimental design
Communication

Education

PhD in ML / related field

Tools

CUDA
NumPy
SciPy
Profiling tools

Job description

Neurophos, Inc. seeks an experienced ML engineer to advance post-training quantization for language and diffusion models on our optical inference engines. You will bridge ML research and hardware, enabling customers to deploy AI workloads on Neurophos hardware.

Ideal candidates have deep expertise in model quantization, transformer models, and numerical optimization, with hands-on PyTorch experience and a track record of rigorous experiments. This role is onsite in Austin or Sunnyvale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer — Optical AI Inference & Quantization
Senior ML Engineer — Optical AI Inference & Quantization

Socket.dev • Austin (TX)

On-site
USD 150,000 - 230,000
Health premium coverage
Unlimited PTO
401(k) matching
+2
Senior Applied Scientist — Optical AI Quantization
Senior Applied Scientist — Optical AI Quantization

Neurophos • Austin (TX)

On-site
USD 180,000 - 260,000
Health plan premiums coverage
Unlimited PTO
401(k) matching
+2
Senior ML Engineer, Optimization
Senior ML Engineer, Optimization

Neurophos, Inc. • Sunnyvale (TX), Northern (KY)

Hybrid
USD 150,000 - 240,000
Health coverage
HSA contributions
Unlimited PTO
+3
Senior ML Engineer: Quantized Inference & Pipelines
Senior ML Engineer: Quantized Inference & Pipelines

NVIDIA AI • Redmond (WA)

On-site
USD 140,000 - 200,000
Equity
Benefits
Senior/Staff Applied Scientist, Numerical Optimization & Quantization
Senior/Staff Applied Scientist, Numerical Optimization & Quantization

Neurophos • Austin (TX)

On-site
USD 180,000 - 260,000
Health plan premiums coverage
Unlimited PTO
401(k) matching
+2
Senior ML Engineer, Optimization
Senior ML Engineer, Optimization

Socket.dev • Austin (TX)

On-site
USD 150,000 - 230,000
Health premium coverage
Unlimited PTO
401(k) matching
+2
Senior Software Engineer, Quantized Inference
Senior Software Engineer, Quantized Inference

NVIDIA AI • Redmond (WA)

On-site
USD 140,000 - 200,000
Equity
Benefits
Senior ML Infrastructure Engineer - Autonomy & Optimization
Senior ML Infrastructure Engineer - Autonomy & Optimization

Nuro • California (MO)

On-site
USD 194,000 - 352,000
Embedded ML Runtime Optimization Engineer
Embedded ML Runtime Optimization Engineer

Applied Intuition • Sunnyvale (CA)

Hybrid
USD 159,000 - 199,000
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage