Чем предстоит заниматься
You'll build a new team from the ground up, owning multimodality at Poolside: teaching our frontier coding model to see. Multimodality is becoming increasingly important in frontier models, and this is a chance to shape this capability in Poolside’s models from an early stage. To deliver it, you'll draw on our powerful model factory, thousands of GPUs, and our strong research team
The near-term focus is image input — the capability most useful to SWE and general agents — reading a design and speccing it out, writing and verifying the code that makes that design real, understanding the plots and diagrams in a paper. You'll set the technical direction, starting with pragmatic adapter-based approaches and evolving toward native multimodality over time. You'll partner closely with across teams inside Applied Research, and work in a newly-formed team alongside one of our talented founding engineers to bootstrap our multimodality efforts
YOUR MISSION
To bring multimodality to Poolside's models, starting with image input and building toward native multimodal understanding
Own Poolside's multimodality direction and capability adoption
Ship our first image-input capabilities
Chart and drive the path from adapter-based to native multimodality
Collaborate on custom evaluations and datasets for multimodal SWE capabilities
Run experiments end to end: hypothesis, implementation, training at scale, analysis
Partner with evals, architecture, data, and post-training teams to land multimodality in Poolside models