AI Engineer: Model Training, Inference & GPU Infra

Agentrys

San Jose, Northern (CA, KY)

Hybrid

USD 170,000 - 210,000

Full time

28 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Agentrys in San Jose seeks an exceptional AI Engineer to own model training, inference, and infrastructure powering its agentic design workforce. You will drive the full model lifecycle—from data pipelines and pretraining to reinforcement learning and high-performance serving.

You will operate scalable GPU infrastructure, optimize performance, and build evaluation loops that improve models from real execution feedback in production environments.

Qualifications

  • > Strong programming in Python and proficiency in a systems language such as C++ or Rust.
  • < Deep experience with ML frameworks such as PyTorch or JAX.
  • < Hands-on experience with large-scale distributed model training and production inference.

Responsibilities

  • > Train, post-train, and fine-tune large language models for agentic engineering workflows, including supervised fine-tuning and RLHF.
  • < Build scalable data pipelines for pretraining, post-training, and evaluation across private engineering data.
  • < Design and operate distributed training on multi-node GPU clusters with parallelism strategies.
  • < Build high-throughput, low-latency inference systems with advanced batching and KV-cache optimization.
  • < Write and optimize custom GPU kernels (CUDA, Triton) and profile end-to-end performance.
  • < Create core model infrastructure: orchestration, scheduling, checkpointing, observability, cost tracking.
  • < Develop automated evaluation systems and reward models to measure agent capability.
  • < Integrate models with agent runtimes and production serving stacks.
  • < Improve reliability, throughput, and cost efficiency of training and inference.

Skills

Python
C++ or Rust
PyTorch or JAX
Distributed training
GPU optimization
Software engineering

Education

PhD or Master in CS/EE

Tools

Megatron-LM
DeepSpeed
FSDP
Ray
CUDA
Triton

Job description

Agentrys in San Jose seeks an exceptional AI Engineer to own model training, inference, and infrastructure powering its agentic design workforce. You will drive the full model lifecycle—from data pipelines and pretraining to reinforcement learning and high-performance serving.

You will operate scalable GPU infrastructure, optimize performance, and build evaluation loops that improve models from real execution feedback in production environments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer, Model Training, Inference & Infra
AI Engineer, Model Training, Inference & Infra

Agentrys • San Jose (CA), Northern (KY)

Hybrid
USD 170,000 - 210,000
Senior Training Infra Engineer - 800+ GPU Scale
Senior Training Infra Engineer - 800+ GPU Scale

Figureai • San Jose (CA)

On-site
USD 150,000 - 350,000
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
AI Training Performance Engineer — Large-Scale GPU Training
AI Training Performance Engineer — Large-Scale GPU Training

Figure • San Jose (CA)

On-site
USD 200,000 - 400,000
Senior AI Infrastructure Engineer — Scale GPU Clusters
Senior AI Infrastructure Engineer — Scale GPU Clusters

AI Breaking Wire • San Francisco (CA)

On-site
USD 280,000 - 400,000
Equity
Medical, dental, and vision benefits
Unlimited PTO
+2
Staff Engineer, GPU AI Inference & RL Infrastructure
Staff Engineer, GPU AI Inference & RL Infrastructure

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
+2
AI Inference Performance & Scale Engineer
AI Inference Performance & Scale Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 210,000
Benefits at a glance
Staff Software Engineer (AI Infrastructure)
Staff Software Engineer (AI Infrastructure)

DeepRec.ai • Palo Alto (CA)

On-site
USD 180,000 - 320,000
AI Infrastructure Performance Modeler
AI Infrastructure Performance Modeler

OpenAI • California (MO)

Hybrid
USD 120,000 - 180,000
Relocation assistance
Hybrid work model
Office in San Francisco
AI Inference Infrastructure Engineer
AI Inference Infrastructure Engineer

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1