Gradle logo

Senior Site Reliability Engineer

Gradle
Remote
Remote$150k–$190k· about 2 months ago

Straight from Gradle’s careers page. Apply on the company site — no recruiter, no middleman.

Explore more remote Site Reliability Engineer jobs — salaries, top companies, and the latest openings.See all →

Senior Site Reliability Engineer

Location: North America (PST)

Department: Develocity Engineering

Who We Are

AI is changing how software gets built. Code production is becoming a commodity. The focus is shifting from writing code to specifying, orchestrating, verifying, and governing change – and the toolchain is the new constraint.

We are at the center of this shift. We build Develocity, a toolchain observability and intelligence platform used by some of the worlds leading software organizations – Netflix, Airbnb, Spotify, SAP, major global banks, and hundreds more. Develocity helps software teams achieve delivery excellence through deep observability, build and test acceleration, and AI-powered intelligence across the entire toolchain – with support for Gradle Build Tool, Apache Maven™, sbt, npm, Python, and Bazel.

We are an AI-native company. AI is not a feature were bolting on – its central to how we work, how we think about our product, and where were heading. Develocity encodes a decade of build and test expertise into a context engineering layer: the ground truth AI agents need to change code safely, and the governance to attribute and gate that change at machine speed. Trust, evidence, and explainability are at the core of everything we build.

We have partnered with the Apache Software Foundation, the Commonhaus Foundation, the Micronaut Foundation, and other OSS projects such as Spring, Quarkus, Kotlin, JUnit, AndroidX, and many more to bring the values of Develocity also to the OSS Community.

Our Values

Seek to Understand: Everything starts with listening and understanding, and we strive to understand different viewpoints, problems, and motivations. Before we take action, we ensure we truly grasp the challenges, perspectives, and goals. 

Know the Why: We approach our work with a clear sense of purpose, ensuring every step is deliberate and focused. We take meaningful action with urgency, but never at the expense of thoughtful consideration. 

Innovate & Iterate: We embrace challenges and are not afraid to try new things, even if they might fail. With deep understanding and a clear purpose, we can develop creative and bold solutions to tackle challenges.

Own the Outcome: We are empowered to take initiative and we maintain transparency in our work and its outcomes. When we execute, we take responsibility for our decisions, measure the success of our innovations, and learn from the results.

Who You Are

Were building a new SRE team and looking for founding members to help shape how we operate. Youll be responsible for the reliability, performance, and availability of Develocity instances serving paying customers, open-source projects, and public-facing services, plus supporting infrastructure like artifact registries.

Youll work on our internally-built Cloud Application Platform, Kubernetes on AWS, and develop deep expertise in it. When incidents happen, youll troubleshoot issues across the stack, from application to infrastructure. Youll collaborate with the Cloud Platform team to improve the tooling you depend on, and with engineering teams to build reliability into how we ship software. If you like automating things and hate doing the same task twice, youll fit in well.

Youll be part of a distributed, remote-first team that values asynchronous communication and written documentation. Strong self-direction and clear communication across time zones are essential.

Responsibilities

  • Operate and maintain all Develocity instances and supporting services.
  • Participate in an on-call rotation, owning incident response and troubleshooting issues across the stack.
  • Drive automation across application deployment, upgrades, monitoring, self-healing, and recovery.
  • Build and maintain observability for all managed services (logging, metrics, tracing, and alerting).
  • Work with engineering teams to build reliability into features from the start.
  • Run incident response and retrospectives, and make sure we learn from them.
  • Own disaster recovery, backups, and business continuity.
  • Communicate with customers during incidents and maintenance windows.
  • Optimize performance, resource usage, and costs.
  • Help evolve our SaaS operations as we grow.

Minimum qualifications

  • 5+ years in SRE, DevOps, or equivalent role operating production services at scale.
  • Strong Kubernetes experience in production environments.
  • Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2).
  • Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform).
  • Track record of incident management and response.
  • Knowledge of SRE best practices (SLAs, SLOs).
  • Scripting proficiency (Python, Bash) for automation.
  • Experience with 24/7 on-call rotations.
  • Strong written and verbal English communication.

Preferred qualifications

  • Experience operating SaaS platforms at scale.
  • Familiarity with Develocity.
  • JVM language experience (Java, Kotlin).
  • Disaster recovery planning and execution experience.
  • Customer-facing incident communication skills.
  • Experience establishing SRE practices in new or growing teams.

What We Offer

  • A ground-floor role in a new SRE team—youll shape how we do things, not inherit someone elses decisions.
  • Real ownership of production systems used by engineers at companies youve heard of.
  • Direct interaction with customers when things go wrong (and when they go right).
  • A culture that values automation over heroics.
  • In-person meetings, such as our annual company offsite and team meetings.
  • Work from home in a remote-first environment.
  • Competitive salaries and equity grants.

Compensation

The US salary range for this position is $180,000-205,000 which reflects the target ranges for all US locations. Within this range, individual pay is determined by geographic location and additional factors including but not limited to experience, relevant skills, qualifications, seniority, performance, and travel requirements. Our recruiting team can share more information about the specific salary range for your location during the hiring process.

Location

  • Remote from anywhere in PST timezone.
  • While our team works remotely and is spread across the globe, we deeply value daily interactions and collaboration.
General Dynamics Mission Systems International logo

General Dynamics Mission Systems International

DevOps Software Engineering Developer

Remote
Calgary, AB$85k–$110k
✓ From careers page· about 2 hours ago
Truelogic logo

Truelogic

Backend Engineer, Security and Safety Solutions

Remote
Mexico City, Mexico
✓ From careers page· about 3 hours ago
Truelogic logo

Truelogic

Backend Engineer, Python (Remote)

Remote
✓ From careers page· about 3 hours ago
Truelogic logo

Truelogic

Backend Engineer, Security and Safety Solutions

Remote
Santo Domingo
✓ From careers page· about 3 hours ago

Discover More than 100,000 Hidden Remote Jobs Before Everyone Else

Unlock All Remote Jobs Today

Simple pricing. Big savings on Quarterly and Yearly.

Monthly Access

$19/month
  • Instant access to fresh remote jobs from 500+ companies
  • New opportunities added hourly, often 3-7 days before anywhere else
  • Advanced filtering by role type, stack, pay, and location
  • Priority customer support
Start 7-day trial — $2.95
Most Popular

Yearly Access

$59/year
  • Everything in Monthly
  • Save $169 (~74%) vs paying monthly
  • Average job search takes ~6 months - get covered for the whole journey
  • Less than the cost of one lunch per month for competitive advantage
  • Equivalent to just ~$4.92/month
Start 7-day trial — $2.95