Model API Engineer - High-Performance Inference

Baseten

New York (NY)

On-site

USD 180,000 - 360,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation including equity
100% medical, dental, and vision insurance
Generous PTO including Winter Break
Paid parental leave
Company-facilitated 401(k)

Job summary

A leading AI platform provider in New York seeks a Model API Engineer focused on infrastructure for hosted AI models. Responsibilities include optimizing performance, building benchmarking frameworks, and collaborating on robust model serving solutions. Ideal candidates have 3+ years in distributed systems and a strong background in backend services. With a competitive compensation range of $180K - $360K and a comprehensive benefits package, this role offers a unique opportunity to shape the future of AI within a diverse and inclusive team.

Qualifications

  • 3+ years experience in building and operating distributed systems or large-scale APIs.
  • Proven track record with low-latency, reliable backend services.
  • Comfortable debugging complex systems, with strong written communication skills.

Responsibilities

  • Design and operate the Model APIs surface with advanced inference capabilities.
  • Profile and optimize TensorRT-LLM kernels for performance.
  • Build benchmarking frameworks to measure real-world performance.

Skills

Distributed systems
Low-latency backend services
Performance profiling
Debugging complex systems
Strong written communication

Tools

Kubernetes
API gateways

Job description

A leading AI platform provider in New York seeks a Model API Engineer focused on infrastructure for hosted AI models. Responsibilities include optimizing performance, building benchmarking frameworks, and collaborating on robust model serving solutions. Ideal candidates have 3+ years in distributed systems and a strong background in backend services. With a competitive compensation range of $180K - $360K and a comprehensive benefits package, this role offers a unique opportunity to shape the future of AI within a diverse and inclusive team.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, High-Performance AI Inference APIs
Software Engineer, High-Performance AI Inference APIs

Dormont Manufacturing Co • San Francisco (CA)

On-site
USD 150,000 - 230,000
Competitive compensation package
Inclusive work culture
Exposure to various ML startups
Model API Engineer — High-Performance Inference
Model API Engineer — High-Performance Inference

The Consensus • New York (NY)

On-site
USD 120,000 - 160,000
Competitive compensation including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
Software Engineer, Model Performance & HPC Tools
Software Engineer, Model Performance & HPC Tools

Baseten • New York (NY)

On-site
USD 160,000 - 200,000
ML Model Performance Engineer - Inference and Acceleration
ML Model Performance Engineer - Inference and Acceleration

Baseten • New York (NY)

On-site
USD 200,000 - 275,000
Senior Model Inference Engineer for Production-Scale AI
Senior Model Inference Engineer for Production-Scale AI

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 490,000
Senior Model API Engineer – Low-Latency AI Serving
Senior Model API Engineer – Low-Latency AI Serving

Baseten • United States

Remote
USD 150,000 - 230,000
Senior AI Model Serving Engineer — Low-Latency Inference
Senior AI Model Serving Engineer — Low-Latency Inference

Menlo Ventures • San Francisco (CA)

On-site
USD 166,000 - 225,000
Annual performance bonus
Equity options
Comprehensive benefits package
AI Infrastructure: Performance Modeling Engineer
AI Infrastructure: Performance Modeling Engineer

OpenAI • Seattle (WA)

Hybrid
USD 266,000 - 445,000
Relocation assistance
Principal AI Infrastructure Engineer
Principal AI Infrastructure Engineer

Postman • Cupertino (CA)

On-site
USD 256,000 - 276,000
Full medical coverage
Flexible PTO
Wellness reimbursement
+1
Remote AI Inference Engineer — Edge Model Deployment & Optimization
Remote AI Inference Engineer — Edge Model Deployment & Optimization

Quadric Inc. • Burlingame (CA)

On-site
Health Care Plan (Medical, Dental & Vision)
Retirement Plan (401k, IRA)
Life Insurance (Basic, Voluntary & AD&D)
+7