cover
Full Time

SRE Team lead/ 1 week ago

Kommo
Attractive
Application ends: 2026-08-12

Quick Summary

Kommo, an international IT company, is hiring an SRE Team Lead for a remote role in LATAM (preferably BRT, UTC-3) to establish core SRE practices, build robust platform architecture, and manage international infrastructure. This role requires 3+ years of DevOps/SRE experience in production environments, team leadership experience (hiring/managing), a data-driven mindset, deep familiarity with SRE practices (SLI/SLO/SLA, incident management), and code fluency in Go, Python, or PHP. The tech stack includes Kubernetes, Prometheus, Grafana, and MySQL.

Overview

Kommo, an international IT company, is seeking an SRE Team Lead to join our development department. Kommo is a leading global messenger-based CRM and a top traffic driver for WhatsApp and Instagram in LATAM. We boast a multinational team of over 100 tech specialists and a rapidly expanding infrastructure.

We are currently restructuring our teams to deeply integrate Site Reliability Engineering (SRE) practices into product development. Our primary goal is to maximize stability and predictability under high loads, handling tens of thousands of requests per second (RPS). We require a leader to establish core SRE practices, build robust platform architecture, and ultimately manage our entire international infrastructure, with a focus on LATAM.

Responsibilities

  • Define, build, and implement reliability practices and tools; analyze and optimize current infrastructure approaches.
  • Develop change management strategies to minimize any negative impact on service availability.
  • Collaborate with product teams to advocate for, test, and implement your ideas.
  • Establish and manage SLA/SLO frameworks and Incident Management processes.
  • Evolve our internal technical platform to boost infrastructure operations efficiency.
  • Hire, mentor, and grow your own engineering team while driving cross-functional alignment.

Our current tech stack includes: Bare-metal + LXD, Kubernetes, Prometheus, Grafana, MySQL, Pulsar / Gearman / Beanstalkd, and internal tools. We are always open to new solutions and approaches!

Requirements

  • 3+ years of experience in DevOps/SRE, with a proven track record in production environment operations.
  • Team leadership experience: Ready to take on hiring and managing a small, agile engineering team.
  • Data-driven mindset: Proficient in using infrastructure metrics and understanding their impact on business KPIs.
  • Deep familiarity with SRE practices, including: SLI/SLO/SLA, reliability roadmaps, production readiness, Deployment Reliability, and Resilience Engineering.
  • AI-forward approach: Open to utilizing AI tools for rapid prototyping, utility building, and brainstorming.
  • Code fluency: Prepared to dive into application code (both client-facing and internal tools). Our core company stack is Go / Python / PHP.

What We Offer

  • Remote work from LATAM (Brasília Time (BRT), UTC-3 preferably).
  • Competitive salary.
  • Opportunity to build teams and processes from scratch, or in your own way.
  • Freedom to pitch your own ideas and see them come to life.
  • A global team of 600+ people across multiple countries.

Job Type: Full-time

Work Location: Remote

Share

Kommo

Kommo

  • Address
    Remoto
View Profile
Your experience on this site will be improved by allowing cookies Cookie Policy