cover
Full Time

Site Reliability Engineer/ 8 hours ago

Rentsync
Attractive
Application ends: 2026-10-29

Quick Summary

This remote (Canada) Site Reliability Engineer (SRE) role focuses on leading production incident response, root cause analysis, and implementing fixes within extensive AWS and Kubernetes environments, while also preventing recurrence through infrastructure hardening and monitoring enhancements. Candidates require 3+ years in cloud engineering, DevOps, or SRE with hands-on incident response, strong AWS experience (EKS, EC2, RDS, VPC networking, IAM, CloudWatch), and deep production Kubernetes expertise. Essential skills include Infrastructure as Code with Terraform, CI/CD comfort, Linux/networking/container fundamentals, scripting (Bash, Python), and experience with monitoring tools like Prometheus, Grafana, and PagerDuty. The annual salary range is 80,000 - 110,000 CAD.

About Rentsync

Rentsync is an award-winning, high-growth organization providing high-quality websites, marketing services, and software solutions to the rental and property management industry throughout Canada and the United States.

About the Role

We are seeking a hands-on Site Reliability Engineer (SRE) to lead our response to production incidents. You will investigate root causes and implement fixes within our AWS and Kubernetes environments.

Between incidents, you will focus on preventing recurrence by hardening infrastructure, enhancing monitoring, and collaborating with engineering teams on performance and reliability improvements.

Our environment is extensive and diverse, featuring over 10 products and 100 services across multiple Kubernetes clusters. Primarily on AWS, with some Azure and GCP, our systems are built using PHP, Ruby on Rails, JavaScript/TypeScript, .NET, Python, and Rust.

Technologies you will work with: AWS (EKS, EC2, RDS, S3, ALB/NLB, CloudWatch), Azure, GCP, Kubernetes, Terraform, Ansible, GitHub Actions/GitLab CI, PagerDuty, Prometheus/Mimir, Loki, Tempo, Grafana, OpenTelemetry, Cloudflare, Ubuntu & Amazon Linux, MySQL & PostgreSQL, Redis & Memcached, NGINX & Traefik, Bash, Python, and applications built in PHP, Ruby on Rails, JavaScript/TypeScript, .NET, and Rust.

This is a remote position. While all qualified candidates are encouraged to apply, preferential consideration may be given to individuals within a reasonable commuting distance of one of our offices.

Duties & Responsibilities

Incident Response & First-Contact Remediation (Primary Focus)

  • Be the first responder for production alerts and incidents across our services, managing them from triage through to resolution.
  • Diagnose and fix issues directly in AWS (EKS, EC2, RDS, networking, IAM) and Kubernetes, addressing problems such as failing pods, resource exhaustion, bad deploys, networking/DNS, and database/cache issues.
  • Roll back, scale, reconfigure, or patch infrastructure to restore service quickly; escalate to development teams only when a code change is truly needed, providing a clear diagnosis.
  • Own our PagerDuty setup and incident response during business hours, aiming to reduce MTTD (Mean Time To Detect) and MTTR (Mean Time To Resolve).
  • Run blameless post-mortems and personally drive the technical follow-up work, beyond just managing action items.
  • Automate runbooks and repetitive operational tasks, including using AI tools to accelerate triage, investigation, and remediation.

Reliability Engineering (Preventing the Next Incident)

  • Build and maintain monitoring for Kubernetes workloads and services (Prometheus/Mimir, Loki, Tempo, Grafana, OpenTelemetry), ensuring low-noise, high-signal alerts.
  • Monitor new releases in production to catch regressions in latency, errors, or resource use before they escalate into incidents.
  • Create and maintain production test suites: synthetic checks, smoke tests, health checks, and load/performance tests.
  • Partner with engineering teams to identify and resolve performance and reliability issues, and define SLOs, SLIs, and error budgets.
  • Harden our platform after incidents: implement Terraform changes, Kubernetes resource tuning, autoscaling, CI/CD checks, and improve secrets and IAM.
  • Keep service documentation and architecture decisions current to enable any engineer to operate our systems.

Required Knowledge, Skills & Abilities

  • Infrastructure as code with Terraform, and comfort working in CI/CD pipelines.
  • Solid Linux, networking, and container fundamentals.
  • Scripting/automation in Bash, Python, or similar.
  • Calm, clear communication during incidents and across teams.

Essential Qualifications

  • 3+ years in a cloud engineering, DevOps, or SRE role supporting production web applications.
  • Hands-on incident response experience where you diagnosed and fixed production issues yourself, not only coordinated them.
  • Strong, hands-on AWS experience in production (EKS, EC2, RDS, VPC networking, IAM, CloudWatch).
  • Deep production Kubernetes experience, including troubleshooting, debugging, and monitoring Kubernetes workloads.
  • A track record of working with engineering teams to identify and resolve performance and reliability issues.
  • Experience with monitoring and observability tools (e.g., Prometheus, Grafana, Loki, Datadog, CloudWatch) and on-call/alerting tools such as PagerDuty.
  • Experience building automated tests or checks for production reliability (synthetics, smoke, health, or load testing).
  • Willingness to take part in an after-hours on-call rotation as we introduce one in the future.

Additional Preferred Qualifications

  • Using AI tools to accelerate SRE work, such as incident triage, log and metric analysis, runbook automation, or infrastructure code.
  • Azure experience (GCP is a plus too).
  • Supporting many tech stacks across multiple teams (PHP, Ruby on Rails, .NET, Python, Rust, JavaScript).
  • AWS certification (e.g., Solutions Architect, DevOps Engineer, or SysOps).
  • Running LGTM (Loki, Grafana, Tempo, Mimir) or OpenTelemetry at scale.
  • Load testing tools such as k6, Locust, or JMeter.
  • MySQL/PostgreSQL operations and Redis/Memcached tuning.
  • Cloudflare (WAF, Workers, Zero Trust).
  • Cloud cost optimization and capacity planning.

Rentsync is an equal opportunity employer. If you are selected to participate in the interview process and require unique accommodations, please don’t hesitate to let us know.

Successful candidates may be required to complete a criminal background check in the final phase of the interview process.

This is an open-ended job posting and may not represent a specific vacancy within the organization.

Rentsync reserves the right to use Artificial Intelligence to screen and/or assess candidates.

Pay Range: 80,000 - 110,000 CAD per year (Canada)

Share

Rentsync

Rentsync

  • Address
    Ontario
View Profile
Your experience on this site will be improved by allowing cookies Cookie Policy