Model API Engineer — High-Performance Inference

The Consensus

New York (NY)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
Paid parental leave
Fertility and family-building stipend
Company-facilitated 401(k)

Job summary

The Consensus is seeking a skilled engineer to join our Model Performance team, focusing on building and optimizing APIs used in AI applications. You will contribute to enhancing the performance and reliability of our model serving infrastructure, ensuring it meets the needs of developers.

This role requires experience with distributed systems and a strong foundation in backend services, allowing you to work effectively within our small, high-impact team. Enjoy competitive compensation, health insurance, flexible PTO, and an inclusive workplace at The Consensus.

Qualifications

  • 3+ years of experience building and operating distributed systems or large-scale APIs.
  • Proven track record of low-latency, reliable backend services.
  • Comfortable debugging complex systems from runtime internals to GPU execution traces.

Responsibilities

  • Design, build, and operate Model APIs with advanced inference capabilities.
  • Profile and optimize TensorRT-LLM kernels and analyze CUDA kernel performance.
  • Productionize performance improvements across runtimes with deep understanding.
  • Build comprehensive benchmarking frameworks for real-world performance measurement.
  • Collaborate closely with teams to deliver robust model serving experiences.

Skills

Distributed systems
Large-scale APIs
Profiling and tracing
Debugging complex systems
Strong written communication

Tools

TensorRT
CUDA
Kubernetes

Job description

The Consensus is seeking a skilled engineer to join our Model Performance team, focusing on building and optimizing APIs used in AI applications. You will contribute to enhancing the performance and reliability of our model serving infrastructure, ensuring it meets the needs of developers.

This role requires experience with distributed systems and a strong foundation in backend services, allowing you to work effectively within our small, high-impact team. Enjoy competitive compensation, health insurance, flexible PTO, and an inclusive workplace at The Consensus.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Model API Engineer - High-Performance Inference
Model API Engineer - High-Performance Inference

Baseten • New York (NY)

On-site
USD 180,000 - 360,000
Software Engineer, High-Performance AI Inference APIs
Software Engineer, High-Performance AI Inference APIs

Dormont Manufacturing Co • San Francisco (CA)

On-site
USD 150,000 - 230,000
Competitive compensation package
Inclusive work culture
Exposure to various ML startups
Senior Model API Engineer – Low-Latency AI Serving
Senior Model API Engineer – Low-Latency AI Serving

Baseten • United States

Remote
USD 150,000 - 230,000
Engineering Manager, Forward Deployed AI
Engineering Manager, Forward Deployed AI

The Consensus • New York (NY)

On-site
USD 130,000 - 160,000
Competitive compensation with equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
Software Engineer - AI Inference Platform
Software Engineer - AI Inference Platform

The Consensus • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation with equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
AI Inference Platform Engineer
AI Inference Platform Engineer

BaseTen • New York (NY), San Francisco (CA)

On-site
USD 140,000 - 210,000
Equity
Medical, dental and vision insurance (
Flexible PTO including Winter Break
+4
Staff Engineer – High-Performance Model Inference
Staff Engineer – High-Performance Model Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
Software Engineer - Model APIs
Software Engineer - Model APIs

Baseten • New York (NY)

On-site
USD 180,000 - 360,000
Competitive compensation including equity
100% medical, dental, and vision insurance
Generous PTO including Winter Break
+2
Software Engineer, Inference - Performance Optimization
Software Engineer, Inference - Performance Optimization

OpenAI • Los Angeles (CA)

On-site
USD 295,000 - 555,000
Equity