cover
Full Time

SRE Team Lead/ 1 week ago

Kommo
Attractive
Application ends: 2026-08-12

Quick Summary

Kommo, an international IT company specializing in messenger-based CRM, is seeking an SRE Team Lead to establish and manage Site Reliability Engineering practices within its development department. This role involves defining reliability practices, developing change management strategies, establishing SLA/SLO frameworks, evolving the technical platform, and hiring/mentoring an engineering team to maximize system stability under high loads (tens of thousands RPS). Candidates need 3+ years of DevOps/SRE experience in production, proven team leadership, a data-driven mindset, deep familiarity with SRE practices (SLI/SLO/SLA), an AI-forward approach, and code fluency in Go, Python, or PHP. This is a full-time remote position, preferably for candidates in LATAM (Brasília Time).

Overview

Kommo, an international IT company, is seeking an SRE Team Lead to join our development department. As a leading global messenger-based CRM and a top traffic driver for WhatsApp and Instagram in LATAM, Kommo boasts a multinational team of over 100 tech specialists and a massive, fast-growing infrastructure.

We are currently restructuring our teams to deeply integrate Site Reliability Engineering (SRE) practices into product development. Our primary goal is to maximize system stability and predictability under high loads, handling tens of thousands of requests per second (RPS). We require a leader to establish core SRE practices, build robust platform architecture, and ultimately manage our entire international infrastructure, with a focus on LATAM operations.

Responsibilities

  • Define, build, and implement reliability practices and tools; analyze and optimize current infrastructure approaches.
  • Develop change management strategies to minimize any negative impact on service availability.
  • Collaborate with product teams to advocate for, test, and implement innovative ideas.
  • Establish and manage SLA/SLO frameworks and Incident Management processes.
  • Evolve our internal technical platform to boost infrastructure operations efficiency.
  • Hire, mentor, and grow your own engineering team while driving cross-functional alignment.

Our current tech stack includes: Bare-metal + LXD, Kubernetes, Prometheus, Grafana, MySQL, Pulsar / Gearman / Beanstalkd, and internal tools. We are always open to new solutions and approaches!

Requirements

  • 3+ years of experience in DevOps/SRE, with a proven track record in production environment operations.
  • Team leadership experience: Ready to take on hiring and managing a small, agile engineering team.
  • Data-driven mindset: Proficient in using infrastructure metrics and understanding their impact on business KPIs.
  • Deep familiarity with SRE practices, including: SLI/SLO/SLA, reliability roadmaps, production readiness, Deployment Reliability, and Resilience Engineering.
  • AI-forward approach: Open to using AI tools for rapid prototyping, utility building, and brainstorming.
  • Code fluency: Prepared to dive into application code (both client-facing and internal tools). Our core company stack is Go / Python / PHP.

What we offer

  • Remote work from LATAM (Brasília Time (BRT), UTC-3 preferably).
  • Competitive salary.
  • Opportunity to build teams and processes from scratch, your own way.
  • Freedom to pitch your own ideas and see them come to life.
  • A global team of 600+ people across multiple countries.
  • Job Type: Full-time.
  • Work Location: Remote.

Share

Kommo

Kommo

  • Address
    Remoto
View Profile
Your experience on this site will be improved by allowing cookies Cookie Policy