Infinity - Data Engineer, Red Tape Index
Requirements
• Our stack is deliberately modern (Python 3.14, uv, ruff, ty, polars, Prefect 3, marimo). We don't filter on those exact tools; we hire for Python and data depth and expect a short ramp. • Strong Python and SQL; you have designed Postgres schemas and owned migrations (SQLAlchemy and Alembic, or equivalents) in production • Data pipeline experience with a lakehouse/medallion mindset: idempotent ingestion, content hashing, and lineage are habits, not aspirations • Web scraping beyond requests: anti-bot evasion, browser automation, and resilience against messy or hostile sources • Statistics literacy for index methodology: winsorization, normalization, weighting, and sensitivity testing, and you can reason about whether an index's math supports its claims • Comfort with modern Python tooling and CI discipline: typing, linting, coverage gates, and conventional commits • Product discovery instincts: you talk with partners in plain language, assess data feasibility before committing, and flag what is proven versus assumed • End-to-end ownership: you are a pragmatic generalist who moves across data, backend, infrastructure, and basic product decisions in an uncertain environment • Prefect experience, or Airflow/Dagster with willingness to switch • AWS (ECS, S3) and Terraform • polars, pyarrow, and marimo or a Jupyter background • LLM-in-pipeline experience (pydantic-ai, AWS Bedrock, evals) • Actuarial, quantitative research, or data science background in ranking or index construction • Experience with government open data (permits, energy, environmental, or economic datasets) • Comfort working alongside AI tooling; our repos are agent-forward (Claude agent teams, spec-driven docs)
Responsibilities
• Ship scrapers and ingestion flows against messy, sometimes adversarial sources, using HTTP/2 clients, TLS-fingerprint evasion, and browser automation fallbacks, and keep them resilient as sources change • Own Postgres schema design and migrations end to end across per-country and per-domain schemas • Build and maintain medallion (bronze → silver → gold) transforms that are idempotent, content-hashed, and lineage-tracked • Implement and defend index methodology: normalization, weighting, and composite construction where the math verifiably says what it claims (our scoring core is held to 100% test coverage) • Assess data feasibility early, clarify requirements with partners, and convert ambiguous index ideas into executable plans • Take an index end to end: sourcing, validation, methodology, publication, and refresh planning • Operate pipelines on our orchestration stack (Prefect dispatching per-flow ECS Fargate tasks) with observability everywhere
Benefits
• High-impact work at the intersection of AI and critical infrastructure regulation • End-to-end ownership of indices, from raw source to published methodology • Small team with outsized influence; your feasibility calls shape what we build • Modern AI-native development environment (Claude Code, Cursor, multi-model orchestration) • Values We Hire For • Values We Hire For • Character: integrity and trustworthiness above all • Character: • Competency: evoking trust and reliably delivering • Competency: • Togetherness: family-level support and alignment • Togetherness: • Impact: meaningful outcomes over activity • Impact: • Commitment: ownership and follow-through • Commitment: • Equal Opportunity Statement
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT