Vytalize Health - Data Reliability Engineer
Requirements
• Education • Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent professional experience. • 5+ years of experience working with data platforms, data pipelines, or distributed data systems in production environments. • Demonstrated experience improving reliability, observability, or operational quality of data systems with measurable SLI/SLO/SLA improvements. • Hands-on experience supporting both data ingestion pipelines and downstream data consumption or delivery patterns. • 1+ years of hands-on experience with machine learning-based monitoring, anomaly detection, or AI-assisted observability tools. • Demonstrated experience with data quality testing, validation frameworks, and quality metrics definition. • Strong understanding of modern data architectures, including data lakehouse patterns and multi-layer (bronze/silver/gold) data models. • Experience with cloud-based data platforms (AWS, Databricks, or similar). • Proficiency in Python and SQL, with experience building or supporting production-grade data pipelines. • Experience implementing data quality frameworks, monitoring tools, and alerting systems. • Demonstrated expertise with workflow orchestration tools (e.g., Databricks Workflows, Airflow) and version-controlled deployment practices. • Familiarity with SRE and reliability engineering concepts including SLIs, SLOs, error budgets, and blameless postmortem culture. • Strong troubleshooting and root cause analysis skills across complex, distributed systems. • Experience designing and operating observability systems for data pipelines (metrics, logs, traces, alerts). • Ability to communicate clearly with both technical and non-technical stakeholders during incidents, postmortems, and requirements discussions. • Understanding of healthcare data, EMR integrations, or regulated data environments is strongly preferred. • Experience defining and measuring data quality metrics; ability to establish and track reliability KPIs. • Hands-on experience with ML-based anomaly detection frameworks or tools (e.g., Datadog Anomaly Detection, cloud-native monitoring ML, custom model development). • Experience leveraging LLMs or AI-assisted tools (e.g., Claude Code, ChatGPT, GitHub Copilot) to accelerate development of monitoring code, incident response workflows, and documentation. • Familiarity with healthcare data standards: FHIR, HL7, CCD, claims data formats, and value-based care metrics. • Experience operating observability and incident management platforms (e.g., DataDog, New Relic, Sumo Logic, PagerDuty). • On-call experience and demonstrated comfort with incident response, runbook creation, and blameless postmortem analysis. • Experience with policy-as-code and data governance frameworks. • Background in a startup or high-growth environment with exposure to scaling data systems. • Familiarity with Tuva or similar clinical data normalization and quality frameworks. • This job description is not designed to cover or contain a comprehensive listing of activities, duties, or responsibilities that are required of the employee. Other duties, responsibilities, and activities may change or be assigned at any time with or without notice.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT