Vytalize Health logo

Data Reliability Engineer (Remote)

Vytalize Health
Remote
Remote· about 3 hours ago

Straight from Vytalize Health’s careers page. Apply on the company site — no recruiter, no middleman.

Data Reliability Engineer

Department: Technology

Location: Remote

Employment Type: FullTime

Description of the Role

The Data Reliability Engineer (DRE) at Vytalize Health is responsible for ensuring the end-to-end reliability, quality, and operational health of data across the full data lifecycle — from ingestion through downstream delivery and consumption. This role sits at the intersection of Data Engineering and Data Services, with a primary focus on building confidence that data is accurate, timely, observable, and dependable for both internal and external consumers.

The DRE role emphasizes proactive monitoring, automation, failure prevention, and rapid recovery. This individual partners closely with Data Engineering, Data Services, DevOps, Product, and Analytics teams to define and enforce reliability standards and operational practices for mission-critical healthcare data pipelines and data products.

Given the sensitive and regulated nature of healthcare data, this role plays a key part in ensuring data reliability while maintaining strict compliance with security, privacy, and regulatory requirements. You will be metrics-driven — establishing clear reliability targets (SLIs, SLOs, SLAs) and measuring success through data quality, freshness, and delivery timeliness.

 

Essential Functions of the Role

Data Pipeline Reliability & Operations

  • Own and continuously improve the reliability of data pipelines across ingestion, transformation, and delivery layers, ensuring data is accurate, complete, and delivered on schedule.

  • Design, implement, and maintain comprehensive monitoring and observability frameworks for data pipelines, datasets, and data services with clear visibility into freshness, volume, schema changes, and data quality.

  • Establish data quality metrics and KPIs; measure and track data accuracy, completeness, timeliness, and consistency across pipelines.

  • Partner with engineering teams and technical stakeholders to establish data quality standards and reduce user-identified bugs.

  • Advocate for a culture of data ownership, operational accountability, and continuous improvement across data teams through documentation, knowledge sharing, and mentorship.

AI-Powered Observability & Anomaly Detection

  • Leverage machine learning and AI-assisted tools to detect data anomalies, quality issues, and reliability risks before they impact downstream consumers—including drift detection, schema validation, and volume/freshness alerting.

  • Establish patterns and best practices for integrating AI-driven observability into data systems while maintaining explainability and human oversight of critical alerts and decisions.

Compliance & Operational Excellence

  • Ensure data reliability practices align with healthcare security, privacy, and compliance requirements, including auditability, traceability, and regulatory reporting.

  • Support capacity planning and scaling efforts by analyzing pipeline performance, usage patterns, and failure modes to identify infrastructure and architectural improvements.

  • Maintain comprehensive documentation of reliability standards, SLAs, incident runbooks, and observability architecture for both technical and non-technical stakeholders.

Qualifications

Education

Bachelors degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent professional experience.

Experience

  • 5+ years of experience working with data platforms, data pipelines, or distributed data systems in production environments.

  • Demonstrated experience improving reliability, observability, or operational quality of data systems with measurable SLI/SLO/SLA improvements.

  • Hands-on experience supporting both data ingestion pipelines and downstream data consumption or delivery patterns.

  • 1+ years of hands-on experience with machine learning-based monitoring, anomaly detection, or AI-assisted observability tools.

Knowledge, Skills, and Abilities

  • Strong understanding of modern data architectures, including data lakehouse patterns and multi-layer (bronze/silver/gold) data models.

  • Experience with cloud-based data platforms (AWS, Databricks, or similar).

  • Proficiency in Python and SQL, with experience building or supporting production-grade data pipelines.

  • Experience implementing data quality frameworks, monitoring tools, and alerting systems.

  • Demonstrated expertise with workflow orchestration tools (e.g., Databricks Workflows, Airflow) and version-controlled deployment practices.

  • Understanding of healthcare data, EMR integrations, or insurance/underwriting data.

  • Strong troubleshooting and root cause analysis skills across complex, distributed systems.

  • Experience designing and operating observability systems for data pipelines (metrics, logs, traces, alerts).

  • Ability to communicate clearly with both technical and non-technical stakeholders during incidents, postmortems, and requirements discussions.

  • Experience defining and measuring data quality metrics; ability to establish and track reliability KPIs.

Preferred Qualifications

  • Hands-on experience with ML-based anomaly detection frameworks or tools (e.g., Datadog Anomaly Detection, cloud-native monitoring ML, custom model development).

  • Experience leveraging LLMs or AI-assisted tools (e.g., Claude Code, ChatGPT, GitHub Copilot) to accelerate development of monitoring code, incident response workflows, and documentation.

  • Familiarity with healthcare data standards: FHIR, HL7, CCD, claims data formats, and value-based care metrics.

  • Experience operating observability and incident management platforms (e.g., DataDog, New Relic, Sumo Logic, PagerDuty).

  • On-call experience and demonstrated comfort with incident response, runbook creation, and blameless postmortem analysis.

  • Experience with policy-as-code and data governance frameworks.

  • Background in a startup or high-growth environment with exposure to scaling data systems.

  • Familiarity with Tuva or similar clinical data normalization and quality frameworks.

Similar remote jobs

More like this →
Appriss Retail logo

Appriss Retail

Data Science Manager (Remote)

Remote
$160k–$170k
✓ From careers page· about 1 hour ago
SDG Group logo

SDG Group

Data Platform & Cloud Engineer

Remote
Santiago de Compostela
✓ From careers page· about 1 hour ago
Qureos Inc logo

Qureos Inc

Data Analyst (Remote)

Remote
Texas City, TX$52k–$121k
✓ From careers page· about 1 hour ago
Qureos Inc logo

Qureos Inc

DevOps Engineer (Remote)

Remote
Palo Alto, CA
✓ From careers page· about 1 hour ago

Discover More than 100,000 Hidden Remote Jobs Before Everyone Else

Unlock All Remote Jobs Today

Simple pricing. Big savings on Quarterly and Yearly.

Monthly Access

$19/month
  • Instant access to fresh remote jobs from 500+ companies
  • New opportunities added hourly, often 3-7 days before anywhere else
  • Advanced filtering by role type, stack, pay, and location
  • Priority customer support
Start 7-day trial — $2.95
Most Popular

Yearly Access

$59/year
  • Everything in Monthly
  • Save $169 (~74%) vs paying monthly
  • Average job search takes ~6 months - get covered for the whole journey
  • Less than the cost of one lunch per month for competitive advantage
  • Equivalent to just ~$4.92/month
Start 7-day trial — $2.95