cover
Part Time

Senior DevOps Engineer, Reliability & Platform Operations/ 3 hours ago

MedTrainer
Attractive
Application ends: 2026-11-01

Quick Summary

MedTrainer is hiring a Senior DevOps Engineer, Reliability & Platform Operations, to work 100% remotely from Mexico. This role requires 8+ years of experience in DevOps, SRE, or Platform Engineering, focusing on enhancing the reliability, stability, and delivery effectiveness of cloud infrastructure, Kubernetes (AKS), CI/CD pipelines (GitHub Actions), databases (MySQL, ProxySQL, RabbitMQ), and production systems. Key responsibilities include improving reliability engineering practices (SLIs, SLOs), operating AKS clusters, implementing Infrastructure as Code (Terraform/Pulumi, Ansible), enhancing observability, supporting incident response, and strengthening security practices, with a strong emphasis on Azure and Python scripting, and supporting PHP Symfony applications.

Company Description

MedTrainer offers the sole all-in-one compliance platform, specifically designed for healthcare organizations by healthcare professionals. For over 13 years, we have empowered healthcare teams to maintain audit readiness and confidence in regulatory compliance through continuous innovation and profound industry expertise.

Our cloud-based platform integrates learning management, credentialing, and compliance into one intelligent system, simplifying complex healthcare operations. Leveraging automation and AI-driven workflows, MedTrainer facilitates faster onboarding, streamlined processes, and efficient scaling, all while upholding the highest standards of compliance and workforce readiness.

Job Description

We are seeking a Senior DevOps Engineer, Reliability & Platform Operations, to enhance the reliability, operational stability, risk posture, and delivery effectiveness of our cloud infrastructure, Kubernetes platforms, CI/CD workflows, databases, middleware, and production systems.

In this pivotal role, you will collaborate closely with DevOps Engineers, Software Delivery Engineers, Software Engineering, Security, Database providers, Support, Product, and leadership teams. Your efforts will focus on reducing production risk, enhancing observability, strengthening operational standards, automating repetitive tasks, and establishing safer paths to production.

You will offer design guidance, advanced technical support, automation solutions, best practices, and reliability standards to DevOps Engineers and Software Delivery Engineers, who will maintain ownership of application delivery readiness and implementation.

This is a senior individual contributor role for a professional with strong production ownership, deep infrastructure and platform expertise, solid Site Reliability Engineering (SRE) practices, and the ability to proactively identify recurring issues and implement lasting improvements.

Responsibilities

  • Improve reliability engineering practices across production systems, including SLIs, SLOs, error budgets, postmortems, runbooks, service health indicators, and corrective-action tracking.
  • Lead reliability enhancements across cloud platforms, AKS, CI/CD pipelines, application environments, databases, and supporting middleware.
  • Operate and enhance AKS clusters, covering upgrades, autoscaling, node pools, networking, storage, identity, security, reliability, and operational standards.
  • Provide guidance, best practices, and standards for Kubernetes application deployment patterns.
  • Improve delivery reliability via advanced GitHub Actions workflows, reusable automation, secure deployment patterns, rollback support, and pipeline optimization.
  • Develop automation, self-service capabilities, reusable infrastructure patterns, and operational procedures to minimize toil, ticket-driven tasks, manual handoffs, and infrastructure drift.
  • Define, implement, and maintain Infrastructure as Code (IaC) and configuration management practices using approved tools.
  • Own and maintain Ansible-based configuration management for OS, middleware, application-supporting services, and operational automation.
  • Support and improve Azure cloud environments as the primary platform; AWS and GCP experience is a plus.
  • Implement and enhance observability across metrics, logs, traces, dashboards, alerting, APM, capacity planning, performance tuning, and incident dashboards.
  • Support incident response for production, deployment, infrastructure, database, middleware, and application reliability issues, including severity recommendation, triage coordination, stakeholder communication, mitigation support, postmortems, and corrective actions.
  • Operate with strong production awareness and ownership, driving reliability improvements and production risk reduction.
  • Recommend blocking, delaying, or rolling back releases when reliability, security, operational, or business-continuity risks are identified.
  • Understand PHP Symfony monolith behavior to support production reliability and collaborate effectively with software engineering teams.
  • Operate, tune, monitor, back up, troubleshoot, and support MySQL, ProxySQL, and RabbitMQ directly, coordinating with the database provider as needed.
  • Strengthen security and compliance practices across infrastructure, CI/CD, and runtime environments, including secrets management, access control, dependency and image scanning, signing/provenance, CI/CD security controls, cloud security configuration, and audit support.
  • Improve backup, restore, disaster recovery, cloud cost visibility, cost control, and operational resilience across critical systems.
  • Mentor DevOps Engineers and Software Delivery Engineers through technical guidance, design review, standards definition, operational knowledge sharing, and reliability best practices.
  • Produce and maintain runbooks, troubleshooting guides, deployment procedures, reliability standards, architecture notes, and knowledge-base articles.
  • Propose technical initiatives focused on reliability, platform operations, observability, automation, incident reduction, cloud maturity, and operational risk reduction.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent professional experience.
  • 8+ years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, cloud operations, infrastructure automation, or production operations.
  • Strong written and verbal English communication skills.
  • Extensive experience operating production systems with core responsibilities in reliability, observability, incident response, automation, and operational risk reduction.
  • Strong hands-on experience with Azure and production Kubernetes / AKS environments.
  • Experience designing or improving CI/CD, Infrastructure as Code (IaC), configuration management, observability, and deployment automation at scale.
  • Strong production troubleshooting skills across cloud infrastructure, Linux, containers, Kubernetes, databases, middleware, CI/CD pipelines, and application-supporting services.
  • Ability to understand application behavior and support production reliability of PHP Symfony applications.
  • Experience operating or supporting MySQL, ProxySQL, and RabbitMQ in production environments.
  • Strong understanding of secure delivery and infrastructure security practices, including secrets management, least privilege, access reviews, CI/CD security, and audit support.
  • Experience with backup, restore, disaster recovery, business continuity, capacity planning, performance tuning, and cost optimization.
  • Ability to mentor engineers, define technical standards, constructively challenge weak practices, and lead technical initiatives without formal people-management authority.
  • Fully remote work capability, including disciplined written communication, self-management, asynchronous collaboration, and reliable participation in distributed-team workflows.

Essential Technologies and Skills:

  • Cloud platforms: Azure (required); AWS and GCP (preferred).
  • Kubernetes: AKS, cluster operations, troubleshooting, upgrades, autoscaling, networking, storage, identity, security, and platform standards.
  • CI/CD: GitHub Actions, reusable workflows, composite actions, environments, approvals, OIDC, self-hosted runners, rollback workflows, artifacts, caching, and pipeline security controls.
  • Infrastructure as Code: Terraform or Pulumi (Pulumi with Python preferred).
  • Configuration management: Ansible.
  • Scripting and automation: Python (required); Bash (preferred).
  • Containers: Docker / OCI and standalone Docker hosts.
  • Operating systems: Linux administration and troubleshooting.
  • Databases and middleware: MySQL, ProxySQL, RabbitMQ.
  • Observability: New Relic, Logz.io, Azure Metrics, ClickStack (preferred), metrics, logs, traces, dashboards, APM, synthetic checks, and SLO-based alerting.
  • Reliability engineering: SLIs, SLOs, error budgets, postmortems, runbooks, toil reduction, incident response, and corrective actions.
  • Security: HashiCorp Vault, 1Password, secrets management, least privilege, access reviews, dependency scanning, container image scanning, signing/provenance, and CI/CD security controls.
  • Application reliability: Production support context for PHP Symfony applications.
  • Operational resilience: Backup, restore, disaster recovery, business continuity, capacity planning, and performance tuning.
  • Documentation: Runbooks, troubleshooting guides, operational procedures, standards, and knowledge-base articles.
  • Working style: Structured problem-solving, strong ownership, production awareness, mentoring, automation mindset, and proactive risk reduction.

Additional Information

  • 100% remote work from anywhere in Mexico.
  • Competitive monthly salary (after taxes).
  • Major Medical Insurance and healthcare coverage.
  • Home office and ergonomics support, including internet and electricity.
  • Professional development opportunities, including English classes.
  • Wellness benefits such as TotalPass gym discounts.
  • Savings plan.
  • Paid time off, including personal days.
  • Collaborative, international, and growth-oriented environment.
  • Opportunity to work with modern cloud, automation, CI/CD, Kubernetes, reliability, observability, and platform technologies.
  • Collaborative engineering culture focused on automation, reliability, operational excellence, and continuous improvement.

All your information will be kept confidential according to EEO guidelines.

At MedTrainer, every role contributes to building technology that supports safer, more efficient healthcare organizations. We value collaboration, ownership, and continuous improvement across all teams.

If you're excited to grow, make an impact, and be part of a mission-driven company, we encourage you to apply.

Share

MedTrainer

MedTrainer

  • Address
    Querétaro, Querétaro de Arteaga
View Profile
Your experience on this site will be improved by allowing cookies Cookie Policy