Design, build, and optimize real-time speech-to-text pipelines (streaming ASR, VAD, audio processing)
Improve transcription accuracy through context injection (user names, teams, custom vocabulary, language detection)
Develop and maintain LLM-powered post-processing (grammar correction, filler removal, mention resolution, formatting)
Build voice-to-action systems that parse natural language into structured workspace commands
Evaluate, benchmark, and integrate ASR models (Whisper, AssemblyAI, Fireworks, etc.) for cost, latency, and accuracy
Collaborate with product and platform teams to ship voice features across MAX Desktop, Mobile, Web, and Browser Extension
Explore multimodal AI capabilities (screen + voice + text) for next-gen assistant experiences
Privacy Notice
If you are a Philippine Job Applicant, please also see our Philippine Data Privacy Notice for further details
Visa Sponsorship
Please note we are unable to sponsor or take over sponsorship of an employment visa for roles outside of engineering and product at this time. Sponsorship for engineering and product roles is not guaranteed, but is instead based on the business needs for that specific role at that time. Please reach out to the recruiter with any questions
Fraud Alert
ClickUp Talent Acquisition will only initiate contact via an .com email or through our official careers portal on clickup.com We will never request fees, payments, or sensitive personal information. Please disregard any offers received outside these channels and report them to support .com
AI Processing Notice