5+ years in technical field or engineering roles where you've owned a technical relationship with a hyperscaler or major SI, not just supported one
Experience with inference serving frameworks (vLLM, SGLang, TensorRT-LLM) and tuning deployments for real workloads
Prior role at a hyperscaler, AI-native cloud, or inference provider
Deep familiarity with other strategic partner stacks: Platforms for AI workloads, network, and identity integration patterns. You know where Fireworks fits and where it doesn't
Experience with agentic frameworks (LangChain, LlamaIndex, or custom tool-use pipelines) — you understand how inference latency and reliability shapes agent behavior at scale
Background in model evaluation — you understand why benchmark gaming is rampant and what rigorous evals actually look like
You've written a technical blog post or reference architecture that people actually read
Track record taking GenAI POCs from prototype to production-scale deployments
On-Target Expectations (Plus Equity)
Total compensation also includes meaningful equity in a fast-growing startup, along with a competitive salary and comprehensive benefits package. Base salary is determined by a range of factors including individual qualifications, experience, skills, interview performance, market data, and work location
Fireworks AI is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all innovators
WHY FIREWORKS?
Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving
Build What’s Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally
Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI—no bureaucracy, just results
Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation
3+ years in a pre-sales, partner engineering, forward-deployed, or technical consulting role
Demonstrated ability to build production software with customers, not just advise on it. You have shipped code running in someone else's production environment
Strong Python skills. Comfortable reading, writing, and debugging production code. Familiarity with Kubernetes and infrastructure engineering
Hands-on fluency with LLM inference: latency/throughput tradeoffs, batching strategies, quantization, structured outputs, function calling. You can explain why 50ms p99 matters to an enterprise CTO
Real experience with fine-tuning — LoRA at minimum, RFT a strong plus. You understand when SFT is enough and when it isn't
Deep familiarity with the Azure AI stack: Azure Foundry, Azure OpenAI Service, Azure ML, AKS, Entra/RBAC for AI workloads. You know where Fireworks fits and where it doesn't
Exceptional communication: able to run a sharp discovery call, present to a VP, and debug a latency issue with an ML engineer in the same afternoon