Get more replies from employers
Send a job-specific resume in minutes.
IC Resources seeks an engineer to accelerate production AI systems, focusing on speed and efficiency of large language model inference. You will optimize GPU-heavy pipelines and scale distributed GPU environments in a fast-moving, innovative AI company.
You will work with research and infrastructure teams to translate cutting-edge models into reliable production systems, improving latency and throughput while leveraging modern hardware.
Compensation: $200K–$290K + Equity
Location: San Jose, California
An innovative AI company is expanding its systems engineering team and is searching for an engineer who enjoys making large language models faster, more scalable, and more efficient in production.
If your interests include GPU optimization, distributed computing, and extracting every ounce of performance from modern hardware, this role offers the opportunity to work on some of today's most demanding inference challenges.
Responsibilities
Your primary focus will be improving the speed and efficiency of production AI systems. You'll evaluate performance bottlenecks, build benchmarking tools, optimize inference pipelines, and help scale distributed GPU environments.
Working closely with research and infrastructure teams, you'll transform new modeling techniques into reliable production systems while continuously improving latency, throughput, and hardware utilization.
We're Looking For Someone Who Has
Additional Experience That Stands Out
What's Offered