Build end-to-end POCs and MVPs alongside customer engineering teams, working inside their codebases, infrastructure, and constraints
For customers whose core product is built on GenAI, architect the inference foundations that capability depends on, and size deployments so they can scale in their market without infrastructure becoming the bottleneck
Run load tests and establish latency, throughput, and cost baselines against realistic customer traffic profiles, and tune deployments to hit those targets
Deploy and validate new model families on inference frameworks (vLLM, SGLang), determining optimal shapes, quantization configs, and serving patterns across workloads
Identify recurring customer pain points and translate them into concrete product proposals, working directly with engineering and product to ship fixes and features
Codify repeatable deployment patterns and contribute them back to internal tooling, documentation, and the platform itself
Feed customer signals (deployment patterns, failure modes, feature gaps) back into the product roadmap with specificity and urgency
Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving
Build What’s Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally
Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI—no bureaucracy, just results
Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation
Fireworks AI is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all innovators