preply - Data Engineer
Requirements
• Hands-on experience building components of large, high-scale applications (e.g., data pipelines, well-structured APIs, efficient algorithms). • Solid experience working in platform or data engineering teams (or equivalent) with the ability to deliver within a multi-stakeholder environment. • Familiarity with cloud platforms (AWS/GCP or equivalent) and modern DevOps practices. • Hands-on experience designing and implementing real-time and batch data processing pipelines using modern frameworks like Spark, Flink, Spark Streaming, Kafka, Debezium, etc. • Experience with orchestration tools such as Airflow, dbt, or similar. • Exceptional problem-solving skills paired with a proactive, innovative mindset focused on continuous improvement. • Strong communication and cross-functional collaboration skills (English level B2+)
Responsibilities
• Contribute to trusted ingestion & enrichment foundations (Data Lake and Data as a Product): • Build and maintain components of Preply's data lake. Ensure every dataset has clear ownership, purpose, schemas, and quality expectations from first ingestion through downstream consumption by analytics, product, and ML teams. Treat trust, correctness, and predictability as first-class features of the platform. • Develop end-to-end ingestion pipelines (batch & streaming): • Data quality, contracts & early validation: • Implement data contracts between producers and consumers, covering schema, freshness, volume, and quality guarantees. Embed validation, anomaly detection, and quality checks early in the ingestion lifecycle to catch issues before they propagate. Apply standardized quality metrics. • Enrichment, modeling & lifecycle management: • Build enrichment logic that joins, standardizes, and contextualizes data across domains using shared definitions and reusable patterns. Support historical tracking, point-in-time correctness, and dataset versioning so downstream users can confidently analyze changes and impacts over time. • Observability, reliability & operational excellence: • Instrument ingestion pipelines with strong observability: freshness, latency, data quality, and cost metrics. Contribute to SLOs, alerting, and incident response playbooks so data failures are visible, diagnosable, and recoverable. Help move the platform from reactive firefighting to proactive reliability management. • Governance & compliance by design: • Enable self-service & standardization: • Contribute to standardized ingestion templates, shared libraries, and platform tooling that enable teams to onboard new data sources independently. Improve discoverability, documentation, and metadata so datasets you own are easy to find and trust without relying on tribal knowledge. • Cross-team collaboration & ownership: • Work closely with Product, Backend, Analytics, and ML partners to align on ingestion requirements and trade-offs. Build strong working relationships across teams. Mentor junior team members and actively contribute to a culture of shared data quality standards and data contracts.
Benefits
• An open, collaborative, dynamic, and diverse culture; • A generous monthly allowance for lessons on http://preply.comPreply.com http://Preply.com, Learning & Development budget, and time off for your self-development. • A competitive financial package with equity, leave allowance, and health insurance; • Access to free mental health support platforms;
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT