Jobgether logo

Director, Site Reliability Engineering

Jobgether
Remote
CanadaRemote$244k–$244k· about 23 hours ago

Straight from Jobgether’s careers page. Apply on the company site — no recruiter, no middleman.

Director, Site Reliability Engineering

Team: IT

Location: Canada

Commitment: Full-time

Workplace Type: remote

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Director, Site Reliability Engineering based in Canada.

As Director, Site Reliability Engineering, you will lead the teams and technical strategy responsible for keeping critical infrastructure reliable, scalable, and secure for millions of users. You will work across software, systems, automation, cloud infrastructure, and operational processes to solve complex reliability challenges at scale. The role combines strategic leadership with hands-on technical depth, including troubleshooting production systems and partnering closely with software engineering teams. You will help shape the future architecture and deployment practices of a large-scale, privacy-focused technology environment. You will also drive improvements in automation, observability, incident response, and engineering efficiency. This is a remote-first leadership opportunity with significant ownership, autonomy, and impact.

Accountabilities:

    • Lead and develop Site Reliability Engineering teams responsible for the reliability, scalability, performance, and operational health of large-scale systems.

    • Define and execute the technical direction for infrastructure, deployment, reliability engineering, automation, and operational practices.

    • Lead high-impact and complex initiatives from initial proposal and planning through implementation, measurement, and postmortem.

    • Investigate and resolve sources of instability across high-traffic, distributed systems, identifying root causes and implementing sustainable remediation.

    • Establish and improve tools, services, monitoring, alerts, incident-response processes, and operational practices that identify and mitigate reliability risks.

    • Partner closely with software engineers to troubleshoot production issues, evaluate performance considerations, and implement appropriate code-level or infrastructure-level solutions.

    • Drive automation for infrastructure provisioning and configuration management to improve efficiency, scalability, consistency, and reliability.

    • Leverage cloud-native architectures and services to strengthen system resilience and support continued growth.

    • Help ensure products and infrastructure meet established reliability standards while minimizing user impact during failures and incidents.

    • Identify emerging technical needs and opportunities to guide the long-term evolution of deployment and infrastructure architecture.

    • Support a culture of ownership, continuous improvement, measurable outcomes, and effective post-incident learning.

    • Requirements:

      • 10+ years of relevant professional experience in Site Reliability Engineering, platform engineering, infrastructure engineering, software engineering, or related fields.

      • 4+ years of experience leading SRE or comparable engineering teams.

      • Experience participating in or managing 24/7 on-call operations for large-scale production environments.

      • Advanced programming experience and the ability to read, write, troubleshoot, and deploy software across production systems.

      • Strong experience with Linux administration and troubleshooting, web technologies, distributed systems, and high-traffic production environments.

      • Demonstrated ability to lead complex technical projects from ambiguous initial requirements through execution and postmortem.

      • Experience developing effective reliability tooling, services, monitoring, alerting, and incident-response capabilities.

      • Strong investigative and root-cause analysis skills, particularly within distributed and high-scale systems.

      • Experience designing and implementing infrastructure automation, provisioning, and configuration-management solutions.

      • Hands-on experience with cloud-native services and architectures, including application packaging and deployment using Docker and Docker Compose.

      • Experience with high-level programming languages such as Go, Perl, TypeScript, Python, or comparable technologies.

      • Experience with AI-driven software development, including the design and implementation of agentic workflows.

      • Strong ability to turn ambiguous or complex problems into practical, innovative solutions with measurable outcomes.

      • Strategic thinking and technical foresight, with the ability to anticipate future infrastructure and reliability requirements.

      • Excellent communication and collaboration skills, with the ability to work effectively across engineering teams and technical disciplines.

      • Strong sense of ownership, autonomy, and accountability in a remote-first working environment.

      • Benefits:

        • Annual compensation of $243,800 USD, plus stock options.

        • Transparent compensation structure, with team members at the same professional level and within the same global region receiving the same compensation regardless of functional team, location, gender, educational background, or years of experience.

        • Fully remote, flexible working arrangement with no core working hours.

        • Average full-time commitment of approximately 40 hours per week.

        • Company-sponsored health benefits for eligible team members based in the United States; these benefits do not extend to team members based in Canada or other countries.

        • Paid parental leave.

        • Support for home-office setup.

        • Co-working allowances.

        • Opportunities to participate in company-wide and team gatherings, with travel expected at least twice per year for an all-hands meeting and a team retreat.

        • Remote-first environment centered on trust, inclusivity, ownership, and empowered project management.

        • Equal employment opportunities and a commitment to an accessible, inclusive hiring process.

        • Reasonable accommodations are available for candidates who require support during the application process.

        • Successful candidates must complete a background check as a condition of employment.

        • The role requires participation in video meetings with cameras enabled.

How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the roles core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
 
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
 
 
## LI-CL1

Similar remote jobs

More like this →
FactSet logo

FactSet

Principal Site Reliability Engineer (Remote)

Remote
Norwalk, CT$190k–$220k
✓ From careers page· about 1 hour ago
FactSet logo

FactSet

Lead Site Reliability Engineer (Remote)

Remote
São Paulo, BR
✓ From careers page· about 1 hour ago
Atria logo

Atria

Senior Software Engineer, DevOps

Remote
✓ From careers page· about 3 hours ago
Mirantis logo

Mirantis

Senior Site Reliability Engineer

Remote
Hyderabad, in
✓ From careers page· about 6 hours ago

Discover More than 100,000 Hidden Remote Jobs Before Everyone Else

Unlock All Remote Jobs Today

Simple pricing. Big savings on Quarterly and Yearly.

Monthly Access

$19/month
  • Instant access to fresh remote jobs from 500+ companies
  • New opportunities added hourly, often 3-7 days before anywhere else
  • Advanced filtering by role type, stack, pay, and location
  • Priority customer support
Start 7-day trial — $2.95
Most Popular

Yearly Access

$59/year
  • Everything in Monthly
  • Save $169 (~74%) vs paying monthly
  • Average job search takes ~6 months - get covered for the whole journey
  • Less than the cost of one lunch per month for competitive advantage
  • Equivalent to just ~$4.92/month
Start 7-day trial — $2.95