hcompany - Research Engineer / Scientist, Post-training & Reinforcement Learning
Requirements
• You have strong programming skills in Python, Rust, or similar; and strong software engineering fundamentals building performant and reliable systems. • Proficient in deep learning frameworks (Pytorch, JAX, TensorFlow). • You can work on different layers of the stack from low-level training backends, data ingestion to ML/RL algorithmic design and implementation. • You know when and where to be rigorous and slower versus when to break and iterate quickly. • You have trained LLMs/VLMs with techniques such as SFT, DPO, RLHF/RLVR, reward modelling, offline RL, distillation, etc. • You have experience with offline and online reinforcement learning in or outside of the context of language models. • Publications in top-tier AI conferences (e.g., NeurIPS, ICML, CVPR, ACL, ICCV, AAMAS, ...) • Advanced degree (PhD or MSc) in a relevant field (e.g., ML, DL, NLP, CV) • Experience with large-scale distributed training and inference (multi-node, large models, MoE, parallelism strategies, etc) • Experience training models for computer use or other multi-turn and/or multimodal agentic settings. • Extensive experience with reinforcement learning with sparse rewards. • Experience with multi-domain training, data mixture design, curriculum learning, model merging, distillation. • You are a good communicator, collaborative and low-ego. • You are able to handle a controlled-chaotic environment with a high-degree of between-teams dependencies and collaboration. • You have a go-do attitude and can balance personal conviction/interests with wider team needs. • You don't shy away from hard research or engineering problems. • Paris or London. • This role is hybrid, and you are expected to be in the office 3 days a week on average. • Please expect some travel between offices on a reasonable cadence (e.g., every 4-6 weeks).
Responsibilities
• Develop and train advanced LLMs and VLMs, including multimodal architectures • Research and implement training methods for enhanced capabilities like instruction following and tool use • Design and optimize data pipelines and training systems for large-scale distributed training • Collaborate with cross-functional teams to integrate models into agentic AI systems • Evaluate model performance and communicate findings to stakeholders • Stay current with advancements in LLMs, VLMs, and related fields
Benefits
• Join the exciting journey of shaping the future of AI, and be part of the early days of one of the hottest AI startups • Collaborate with a fun, dynamic and multicultural team, working alongside world-class AI talent in a highly collaborative environment • Unlock opportunities for professional growth, continuous learning, and career development • If you want to change the status quo in AI, join us.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT