- As an AI Field Engineer for Strategic Partnerships, you will be one of the technical owners of Fireworks’ most strategic partnership. You’ll work closely with Strategic Partner field teams, Partner-aligned ISVs, and the SIs that run enterprise AI transformation programs to make Fireworks the default inference and fine-tuning layer in every Partner AI architecture
- The role sits at the intersection of engineering, partner development, and customer delivery. You build reference architectures, run benchmarks, debug production integrations, and co-develop POCs — all while holding your own in executive-level conversations about strategy, roadmap, and business outcomes
- You spend most of your time building and enabling. You ship code, run joint POCs with Partner field teams, and architect deployments that span Strategic Partners and Fireworks. But you also lead discovery conversations, align partner stakeholders, and translate field signals into product improvements that compress the feedback loop from partner to roadmap
- As a Field Engineer aligned with our Partnerships team you own the technical relationship between Fireworks and the Partner ecosystems, Partner field teams, ISVs building on Strategic Partners, and the SIs that deliver AI transformation programs on Strategic Partners. As an example, the Microsoft partnership is a core go-to-market bet: clients like UIPath, Stack Blitz, Motif run via Fireworks on Foundry.. Your job is to scale that pattern across the partner ecosystem
- These engagements involve large, multi-stakeholder organizations, so you will need to navigate both the enterprise buyer (IT, security, compliance) and the builder (ML engineers, platform teams, app developers), while building the trusted-advisor relationships inside Microsoft’s field that multiply your reach
- Technical Delivery and Deployment
- Be the technical lead on co-sell motions with Strategic Partners — joint reference architectures, partner integration patterns, and shared POCs for strategic accounts
- Build end-to-end POCs and MVPs alongside partner engineering teams, working inside their codebases, infrastructure, and constraints
- Run load tests and establish latency, throughput, and cost baselines against realistic customer traffic profiles, and tune deployments to hit those targets
- Deploy and validate new model families on inference frameworks (vLLM, SGLang), determining optimal shapes, quantization configs, and serving patterns across workloads
- Model Strategy and Fine-Tuning
- Guide customers on model selection, fine-tuning strategy (SFT, DPO, RFT), and evaluation methodology
- Build and run fine-tuning pipelines directly with customers, navigating trade-offs between model families, compute cost, and quality targets
- Design and implement evaluation frameworks that measure production-quality metrics, not just benchmark scores
- Product Feedback and Platform Improvement
- Own the feedback loop — surface partner-driven product gaps to Fireworks engineering, and translate the roadmap back into partner messaging
- Ship external technical content: reference architectures, integration guides, and benchmark posts that make it easy for partners to win deals with us
- Track pipeline health; flag risks and opportunities to Field leadership weekly
Real experience with fine-tuning — LoRA at minimum, RFT a strong plus. You understand when SFT is enough and when it isn’tHands-on fluency with LLM inference: latency/throughput tradeoffs, batching strategies, quantization, structured outputs, function calling. You can explain why 50ms p99 matters to an enterprise CTODeep familiarity with the Azure AI stack: Azure Foundry, Azure OpenAI Service, Azure ML, AKS, Entra/RBAC for AI workloads. You know where Fireworks fits and where it doesn’tStrong Python skills. Comfortable reading, writing, and debugging production code. Familiarity with Kubernetes and infrastructure engineering3+ years in a pre-sales, partner engineering, forward-deployed, or technical consulting roleExceptional communication: able to run a sharp discovery call, present to a VP, and debug a latency issue with an ML engineer in the same afternoonDemonstrated ability to build production software with customers, not just advise on it. You have shipped code running in someone else’s production environmentExperience with inference serving frameworks (vLLM, SGLang, TensorRT-LLM) and tuning deployments for real workloads5+ years in technical field or engineering roles where you’ve owned a technical relationship with a hyperscaler or major SI, not just supported onePrior role at a hyperscaler, AI-native cloud, or inference providerExperience with agentic frameworks (LangChain, LlamaIndex, or custom tool-use pipelines) — you understand how inference latency and reliability shapes agent behavior at scaleDeep familiarity with other strategic partner stacks: Platforms for AI workloads, network, and identity integration patterns. You know where Fireworks fits and where it doesn’tBackground in model evaluation — you understand why benchmark gaming is rampant and what rigorous evals actually look likeYou’ve written a technical blog post or reference architecture that people actually readTrack record taking GenAI POCs from prototype to production-scale deployments