Neurons Lab - Data Engineer (UA/RU Language speaking)
Requirements
• Strong Python and solid SQL • Python • Unstructured-data pipelines: transcripts, mail, chat, documents — parsing, normalisation, deduplication • Unstructured-data pipelines • Embedding / retrieval infrastructure: chunking strategies, vector stores (pgvector, OpenSearch, Pinecone-class), plus loading a graph store • Embedding / retrieval infrastructure • API and connector integration at scale: Google Workspace / M365, Slack, CRM; rate limits, pagination, incremental cursors, webhooks • API and connector integration • Entity resolution / record linkage (deterministic + fuzzy) without a clean shared key • Entity resolution / record linkage • Orchestration: Airflow, Step Functions or equivalent; idempotent, restartable jobs • Orchestration • AWS and/or GCP data stack; comfortable in a private / VPC deployment • AWS and/or GCP • PII detection, redaction, encryption and retention in practice • PII detection, redaction, encryption and retention • Clear written English; documents for handover and works well async in a small distributed pod • Knowledge • GDPR applied to employee-generated data (mail, chat, meeting recordings) and EU data residency across multiple jurisdictions • Data lineage, provenance and audit patterns — and why an AI system needs them more, not less • lineage, provenance and audit • How retrieval quality depends on ingestion quality — enough understanding of RAG to make the right upstream choices • retrieval quality depends on ingestion quality • Well-Architected security and cost practice; awareness of financial-services expectations — a plus • Well-Architected • 4+ years in data engineering, with real unstructured / semi-structured work (not only warehouse modelling) • 4+ years • unstructured / semi-structured • Demonstrated experience integrating many third-party APIs into one coherent store, including historical backfill • integrating many third-party APIs • Experience building pipelines feeding an LLM / retrieval system — strong plus • LLM / retrieval system • Experience handling sensitive personal data in a regulated or security-sensitive environment • sensitive personal data • Comfortable being the only data engineer on a small (2.5-FTE) pod, at part-time allocation, without hand-holding • only data engineer
Responsibilities
• Stand up capture by default: notetaker on every call with speaker attribution, plus ingestion from mail, Slack and messengers — designed as opt-out, not opt-in, and reversible if the client changes their mind. • capture by default • Backfill the archive: years of historical email, Slack, board protocols, decks and portfolio updates — parsed, deduplicated and dated correctly. • Backfill the archive • Build document parsing for the awkward long tail: PDFs, scanned board packs, spreadsheets, slide decks, forwarded attachments. • document parsing • Implement identity / entity resolution: the same person across Slack handle, mail alias and calendar invite; the same portfolio company across a deck, a mail thread and a CRM record. • identity / entity resolution • Build chunking and embedding pipelines and load the vector + graph stores behind the ontology the architect defines. • chunking and embedding pipelines • Implement incremental sync through the connector layer (MCP / Composio-class) — no full re-crawls, no silent drift, clear handling of edits and deletions. • incremental sync • Attach access scope and provenance to every record at ingestion, so permission-aware retrieval and audit are possible downstream rather than bolted on. • access scope and provenance to every record at ingestion • Run PII detection, redaction and retention logic; evidence to the client's security function what is stored, where, and for how long. • PII detection, redaction and retention • Orchestrate with Airflow / Step Functions; build repeatable, monitored pipelines rather than scripts, with alerting when a source stops flowing. • Airflow / Step Functions • repeatable, monitored pipelines rather than scripts • Keep cost and latency under control at volume — batching, incremental embedding, storage tiering — and report the unit economics. • cost and latency under control • Write runbooks so the client's own team can operate this after handover. • runbooks
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT