synthesia - Staff Research Engineer - Multimodal Generative Modelling
Requirements
• Having shipped a generative model into a live product used at meaningful scale, not just published or prototyped it. • Working on conversational or interactive systems where latency, responsiveness, and user experience were first-class constraints, not afterthoughts. • Working on LLMs with large scale trainings leadings to models with decent reasoning capabilities • Owning a research problem end to end: from architecture proposal through pretraining, post-training, and production deployment. • Collaborating across modalities or teams (e.g. audio and video, or research and product) to ship a unified system. • Experience with real-time or streaming architectures. • Familiarity with state-of-the-art architectures in audio and speech generation, such as diffusion models, neural codecs, flow-matching models, or autoregressive decoders. • Excellence in one or more of the following modalities: voice, text, video. • Evidence of original research contributions, such as publications or open-source work at top-tier venues (e.g. NeurIPS, CVPR, ICML, ICLR, Interspeech).
Responsibilities
• Shape our roadmap to create new model capabilities and unlock new functionality for our customer base, on both short and long time horizons. • Propose novel multi-modal system architectures (especially text and voice). • Develop and evaluate streaming and conversational systems for low-latency, interactive voice-video synthesis. • Design solutions that reinforce emotional expressiveness and natural interaction. • Implement and bring designs to life, from pretraining through post-training. • Integrate and test novel architectures (neural codecs, diffusion, flow-matching) to enhance realism and responsiveness. • Define new evaluation metrics for conversational systems, including latency-aware and interaction-based measurements. • Track the latest research in audio-visual diffusion, autoregressive models, neural codecs, and multimodal LLMs. • Curate new datasets to complement existing data. • Lead post-training initiatives like DPO, fine-tuning, and distillation to bring models to shipping quality. • Ship models to production with optimised runtime to serve customers, and address their feedback thereafter. • The ability to bring novel ideas and designs that advance the field of interactive multimodal systems. • Strong understanding of generative modelling, ideally applied to sequential or multimodal data. • Hands-on experience with large language models or similar transformer-based architectures. • High proficiency in PyTorch, including distributed training and model optimization. • A solid grasp of time-series modeling and tokenization, preferably in the context of audio, speech, or video. • A demonstrated ability to prototype quickly, test hypotheses, and iterate efficiently. • Proven experience training deep learning models end-to-end, from data preparation through evaluation. • Strong general software engineering skills, enabling contributions to a large, shared research infrastructure.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT