relationrx - Research Data Engineer
Requirements
• Deep experience with cloud-native data infrastructure (AWS S3 / GCS, plus the surrounding ecosystem) and Infrastructure-as-Code (Terraform or equivalent). • A track record of designing data pipelines and storage layouts for large, heterogeneous datasets. • Experience building scalable analytical data processing workflows using modern engines and frameworks (e.g. Spark, Polars, Dask, DuckDB, or equivalent), with an understanding of their performance and architectural trade-offs. • Hands-on experience with workflow orchestration (e.g. Airflow, Dagster, Prefect, or equivalent) and containerised environments (e.g. docker, k8s). • Working knowledge of modern columnar / scientific data formats (Parquet, Zarr, TileDB, HDF5) and lakehouse technologies. • Experience partnering closely with scientific or research users and comfortable with the messiness of real-world experimental data. • Bonus experience: biomedical or genomics data (BAM, FASTQ, AnnData, OME-Zarr); regulated or pharma-partnered environments; data governance, FAIR principles, or research data management; feature store implementations. • PERSONALLY, YOU • Are comfortable working in a matrixed environment, balancing multiple stakeholders and contributing effectively across teams. • Take ownership of your work, proactively seek opportunities to contribute, and enable others to do their best work. • Communicate openly and directly, give and receive feedback constructively, and handle challenging conversations with respect. • Actively seek out diverse perspectives, build strong working relationships, and contribute to shared goals across teams. • Embrace challenges with openness and resilience, set high standards for yourself, and strive to deliver meaningful outcomes. • WORKING STYLE & CULTURE AT RELATION • At Relation, we operate in a matrixed, interdisciplinary environment, where impact is driven through collaboration across scientific, technical, and operational domains. We collaborate, and you will partner with colleagues across multiple teams and projects, contributing your expertise while aligning to shared company priorities. We work together and win together! The patient is waiting! • RECRUITMENT AGENCIES • Please note that Relation does not accept unsolicited resumes from agencies. Resumes should not be forwarded to our job aliases or employees. Relation will not be liable for any fees associated with unsolicited CVs.
Responsibilities
• Design, build, and maintain scalable data pipelines that ingest multi-modal scientific data. • Optimise data movement, storage layouts, and access patterns for analytical and ML workloads. • Stand up and evolve cloud-native data lake / lakehouse infrastructure. • Implement data versioning, lineage, and quality monitoring. • Partner with data scientists day-to-day to ensure fast and seamless data workflows. • Collaborate closely with ML scientists and research engineers to design data representations, storage layouts, and access patterns that enable efficient model training and experimentation. • Build and operate workflow orchestration for both production pipelines and large-scale batch jobs. • Ensure infrastructure meets security, audit, and governance requirements. • Champion engineering best practices across the data platform. • Contribute to architecture decisions across the broader ML platform, including how compute, data, and training systems integrate. • A degree in Computer Science, Engineering, or a related quantitative discipline; significant industry experience in data engineering, MLOps, or data platform roles.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT