Site Reliability Engineer
Straight from Vida Health’s careers page. Apply on the company site — no recruiter, no middleman.
Site Reliability Engineer III
Team: All Teams
Location: United States
Commitment: Full Time- Exempt
Workplace Type: remote
Vida has been operating and growing for years, and our infrastructure reflects that. We run on GCP with a production GKE cluster hosting around 50 workloads, from Django applications to scheduled Airflow jobs. Our data layer includes Cloud SQL (MySQL and PostgreSQL), Redis, and Firestore. Our infrastructure is defined in two Terraform repositories, one for core GCP infrastructure and one for our data platform, and both have grown across many contributors over time. Until now, our infrastructure has been managed by backend engineers with deep infrastructure experience, and this role adds our first dedicated SRE to that group.
Youll be Vidas first dedicated Site Reliability Engineer. Youll join the Enablement Team, which owns the platform and tooling our Engineering Teams build on. Youll report to the Engineering Manager and work closely with the teams Lead Engineer, who sets technical direction and will mentor you. This is a fully remote role with no time zone restrictions.
Youll modernize, consolidate, and scale our infrastructure as Vida takes on a wave of new enterprise contracts starting January 1. Youll also help shape what SRE looks like at Vida going forward.
Repsonsibilities:
- Consolidate our Terraform, which has grown into inconsistent patterns across our infrastructure and data repositories, into a clean, well-documented structure the whole team can work in. Establish conventions for state management, module structure, code review, and CI checks.
- Normalize environments, improve build and deploy automation in GitHub Actions, and add drift detection and alerting.
- Apply overdue patches and upgrades across our Cloud SQL databases and application runtimes.
- Right-size compute and database workloads for growth, including connection pooling and scaling improvements for high-traffic services.
- Evaluate our Kubernetes architecture as we grow, including whether and when to move to a multi-cluster setup.
- Improve monitoring and observability in Datadog and Cloud Monitoring so we catch issues before they become incidents.
- Design observability access for contractors and external partners that gives them the visibility they need while keeping protected health information out of view.
- Retire legacy infrastructure and tooling that has been replaced but not yet decommissioned.
- Build repeatable operational processes, including runbooks, an on-call rotation, and escalation documentation.
- Support infrastructure readiness for Vidas January 1 enterprise launches.
- Additional responsibilities as needed.
Qualifications:
- Bachelors degree at a minimum.
- 5+ years of experience in SRE, DevOps, or infrastructure engineering, with real ownership of production systems.
- Deep hands-on Terraform experience, including structuring modules and managing state across environments.
- Strong working knowledge of GCP, including GKE, Cloud SQL (MySQL and PostgreSQL), IAM, networking and load balancing, and cost management.
- Production Kubernetes experience, including autoscaling, resource management, and judgment about what belongs in the cluster versus outside it.
- Hands-on experience building monitoring, alerting, and dashboards with tools like Datadog or Cloud Monitoring.
- Proficiency in Python for tooling and automation.
- Comfortable working across multiple teams and disciplines, and explaining infrastructure decisions to non-specialists.
Preferred:
- Experience as an early or first SRE hire.
- Experience refactoring or consolidating a large, organically grown Terraform codebase.
- Experience improving observability from a less mature baseline.
- Experience in a HIPAA-regulated or other compliance-driven environment.
- CI/CD experience with GitHub Actions.
- Experience running Django applications or Airflow in production on Kubernetes.
- Experience designing or migrating to multi-cluster Kubernetes architectures.
Similar remote jobs
All Site Reliability Engineer jobs →


Thomson Reuters
Senior Software Engineer

Job Listings at Kinaxis Inc.
Senior AI Software Engineer, Agentic Systems & Data Fabric (Remote)
Discover More than 100,000 Hidden Remote Jobs Before Everyone Else
Unlock All Remote Jobs Today
Simple pricing. Big savings on Quarterly and Yearly.
Monthly Access
- Instant access to fresh remote jobs from 500+ companies
- New opportunities added hourly, often 3-7 days before anywhere else
- Advanced filtering by role type, stack, pay, and location
- Priority customer support
Yearly Access
- Everything in Monthly
- Save $169 (~74%) vs paying monthly
- Average job search takes ~6 months - get covered for the whole journey
- Less than the cost of one lunch per month for competitive advantage
- Equivalent to just ~$4.92/month