Maintain and evolve Command's prompt architecture: the system prompt, skill loading system, session context, and the policy and compliance layers underneath
Tune model behavior: reasoning effort, prompt caching strategy, fallback chains, and the streaming patterns that make the product feel fast
Stay current with how models are evolving and bring that knowledge back to how Command is built
Has 7 or more years of software engineering experience, with deep technical expertise building and scaling LLM-powered applications in production
Has gone beyond shipping a first version: you have scaled an LLM-powered product, dealt with the reliability and performance problems that come with real usage, and made it better over time
Has experience designing agentic systems and has opinions about how to architect multi-step workflows that are reliable, explainable, and safe to run on behalf of real users
Has built eval infrastructure and can write cases that actually measure whether the product works, not just whether the model outputs something plausible
Understands the real tradeoffs in LLM deployments: latency, cost, compliance, and what breaks in production that doesn't show up in demos
Has opinions about what makes an AI product trustworthy, not just impressive, and can build toward that bar
Is comfortable with TypeScript and willing to learn Haskell for backend tool work, or already comfortable with both
Can work across the full stack of an AI product, from the system prompt to the streaming frontend
Has a track record of mentoring engineers and raising the technical bar of their team
$189,700—$237,100 CAD
Our target new hire base salary ranges for this role are the following
$200,700—$250,900 USD