Site Reliability Engineer (Remote)
Straight from JumpCloud’s careers page. Apply on the company site — no recruiter, no middleman.
Site Reliability Engineer - India
Team: Software Engineering
Location: Bangalore, India - Remote
Commitment: Full Time
Workplace Type: remote
About JumpCloud®
JumpCloud is Intelligent, Secure IT.
About the role:
This is not a traditional operations role.
What you’ll be doing:
-
Design, deploy, and maintain the reliability, availability, and performance of critical JumpCloud systems and APIs across AWS and GCP.
-
Operationalize SLIs, SLOs, and error budgets in direct partnership with core application teams.
-
Build and refine end-to-end observability across microservices and cloud infrastructure using tools like Datadog.
-
Implement actionable monitoring across Golden Signals (Latency, Traffic, Errors, Saturation) to optimize detection (MTTD) and minimize alert fatigue.
-
Participate in on-call rotations, incident response, and blameless post-incident reviews to drive continuous systemic improvements.
-
Manage and operationalize production Kubernetes (EKS) clusters utilizing GitOps delivery workflows (Argo CD, Kargo).
-
Provision and secure multi-cloud infrastructure using modular Terraform (Infrastructure-as-Code).
-
Develop and maintain Disaster Recovery (DR) dashboards, runbooks, multi-region failover automation, and validation tests to ensure alignment with defined RTO and RPO targets.
-
Eliminate operational toil by writing production-grade Python or Go scripts and automation tools.
-
Leverage AI-assisted development tools (Cursor, Claude Code, GitHub Copilot) to accelerate scripting, runbook generation, and incident triage.
We’re looking for:
-
5+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical systems.
-
Python/Go Proficiency: Hands-on capabilities writing code for SRE tools, custom automation, and cloud integrations.
-
Kubernetes Ecosystem: Production experience with Kubernetes cluster operations, container orchestration, and GitOps pipelines (Argo CD).
-
Infrastructure as Code: Solid experience writing, maintaining, and modularizing Terraform configurations.
-
Cloud Architecture: Direct experience operating cloud workloads on AWS (EKS, IAM, VPC networking, Route53, ALB/NLB) or GCP.
-
FinOps & Cost Visibility: Practical experience setting up cost-allocation tagging, resource right-sizing, and building FinOps dashboards to visualize cloud spend.
-
Disaster Recovery & Monitoring: Experience building DR dashboards, running failover drills, and configuring monitoring tools to track system health and recovery metrics
-
Observability & Incident Management: Practical experience with Datadog (or similar), PagerDuty, alerting hygiene, and working within SLI/SLO frameworks.
-
Solid operational experience configuring and troubleshooting production service meshes (Istio or similar) and managing high-availability proxy solutions (HAProxy, NGINX, or similar).
-
Problem Solving & Mindset: Strong troubleshooting skills, effective collaboration, and a track record of driving operational efficiency through code.
-
A strong team player who helps us live by our core values: building connections, thinking big, and getting 1% better every day.
Preferred Qualifications:
-
Experience with CI/CD tools such as GitHub Actions or GitLab Pipelines.
-
Basic understanding of chaos engineering principles or testing resilience in staging/production.
-
Familiarity with secrets management tools (HashiCorp Vault, AWS Secrets Manager, External Secrets Operator).
-
Basic knowledge of DevSecOps tools and scanning/fixing infrastructure-as-code vulnerabilities.
## LI-MS1
Similar remote jobs
All Site Reliability Engineer jobs →



Discover More than 100,000 Hidden Remote Jobs Before Everyone Else
Unlock All Remote Jobs Today
Simple pricing. Big savings on Quarterly and Yearly.
Monthly Access
- Instant access to fresh remote jobs from 500+ companies
- New opportunities added hourly, often 3-7 days before anywhere else
- Advanced filtering by role type, stack, pay, and location
- Priority customer support
Yearly Access
- Everything in Monthly
- Save $169 (~74%) vs paying monthly
- Average job search takes ~6 months - get covered for the whole journey
- Less than the cost of one lunch per month for competitive advantage
- Equivalent to just ~$4.92/month