cover

Pipeline Engineer (Worldwide Remote)/ 8 hours ago

CloudLinux
Attractive
Application ends: 2026-09-23

Quick Summary

CloudLinux seeks an experienced, worldwide remote Pipeline Engineer with 5+ years in backend/platform/infrastructure engineering to design, build, and operate automated pipelines that convert threat intelligence into website protection. The role requires demonstrable experience with multi-stage data/automation pipelines (CI/CD, ETL, job orchestration), deep proficiency in Python, Go, or Rust, strong systems design judgment, and practical experience with workflow orchestration (e.g., Airflow, Temporal). Key responsibilities include transforming batch jobs into reliable, observable systems, defining SLOs, building guardrails, and ensuring low-maintenance operations, all while working with petabyte-scale data and fleet-scale telemetry. Cybersecurity background is not required; the focus is on orchestration, reliability, and observability.

CloudLinux is a global, remote-first company committed to principles of integrity, employee well-being, and delivering high-volume, low-cost Linux infrastructure and security products. Our solutions enhance operational efficiency for businesses. We foster a supportive team environment where mutual success is prioritized.

Learn more about us: https://cloudlinux.com/

Imunify360 Security Suite, a product of CloudLinux Inc., is renowned as the #1 OS for security and stability among hosting providers. Imunify offers an innovative, automated security solution specifically designed for shared, VPS, and dedicated servers. Its six-layer approach ensures comprehensive and complete attack prevention.

We are seeking an experienced engineer to manage and enhance the automated pipelines responsible for transforming threat intelligence into deployed protection for millions of websites.

When a new vulnerability emerges, a sophisticated chain of automated systems must detect it, acquire the vulnerable source, generate Web Application Firewall (WAF) rules, create tests, validate rule effectiveness without disrupting legitimate traffic, and then deploy it across tens of millions of websites. This system also monitors production to automatically roll back any misbehaving deployments. This critical chain operates 24/7, and any unreliable link could compromise customer protection. We need a strong Engineer to rapidly expand this system while maintaining its reliability and transparency. This involves tackling complex engineering challenges and managing data at scale in an unpredictable environment.

This role is not an analyst or research position, though relevant experience is a plus. We require someone capable of building robust, continuously operating systems: complex internally, yet boringly reliable externally.

This is a fully remote position with flexible hours, offering the freedom to plan your day and work from anywhere globally.

What you would own

Systems currently in production that require significant growth:

  • Automated Protection Pipeline: The end-to-end chain from threat intelligence to a validated, deployed rule. This multi-stage, largely autonomous pipeline must complete within a fixed daily time window.
  • Progressive Release Automation: Controlled, staged rollout of protection across the fleet, featuring automated guardrails that can pause or roll back a stage without human intervention.
  • Quality Gates: Systems that determine, from live production signals, if a deployed component is causing harm, and act on this information before it reaches the next stage. Errors are costly in both directions: missing a problem breaks customer sites, while over-correction silently removes protection.
  • CI at Scale: Validation systems that provision real, disposable environments across a vast matrix of software versions and configurations, delivering trustworthy verdicts quickly enough to meet release windows.
  • LLM Orchestration and Cost Control: Subsystems driven by AI agents within purpose-built harnesses, with evaluation, budgets, and spend accounting as critical components.
  • Observability and Alerting: Comprehensive monitoring ensuring pipelines report their own condition, prove health, and escalate issues autonomously.

These systems operate on a petabyte-scale threat intelligence store and live telemetry from over 60 million websites, adhering to strict end-to-end latency budgets measured in hours. Our primary focus is eliminating silent budget misses.

Specific details will be discussed during the interview process.

Key responsibilities

  • Design, build, and operate the automated pipelines end-to-end.
  • Transform fragile multi-stage batch jobs into resumable, idempotent, observable systems with explicit state machines and recovery paths.
  • Define and enforce latency budgets and Service Level Objectives (SLOs) per stage, making violations visible and actionable.
  • Build the observability layer, including metrics, dashboards, alerting, and health gates, enabling pipelines to report their own condition.
  • Design and implement guardrails: automatic hold and rollback, blast-radius limits, kill switches, and safe-by-default behavior for upstream dependency unavailability.
  • Ensure low-maintenance systems by eliminating manual steps, reducing human oversight, and minimizing operational surface area.
  • Write and maintain unit and integration tests for complex logic involving concurrency, partial failure, external API flakiness, and multi-stage state.
  • Investigate and resolve complex issues across ClickHouse, GitLab CI, S3/object storage, Prometheus/Grafana, and third-party APIs.
  • Collaborate with security analysts and the Server team on architecture, providing critical feedback on designs that may not withstand production realities.

Requirements

  • 5+ years of professional backend, platform, or infrastructure engineering experience.
  • Demonstrable experience building and operating multi-stage data or automation pipelines (e.g., CI/CD systems, ETL/ELT, build and release automation, job orchestration, ML/data platforms). This is the most critical requirement; expect a detailed discussion.
  • Deep proficiency in at least one of Python, Go, or Rust. We use all three, and language specialization is not required. Depth in one language serves as a proxy for real experience; expect specific questions about systems you've designed, shipped, their architecture, failure behavior, and potential improvements.
  • Strong systems design judgment, prioritizing architectural decisions over raw coding throughput. The core challenge is determining what to build, anticipating failure points, and ensuring self-correcting systems.
  • Practical experience with workflow orchestration and job scheduling (e.g., Airflow, Temporal, Prefect, Dagster, Argo, custom schedulers) in production environments.
  • A working instinct for reliability engineering: idempotency, retries with backoff, exactly-once vs. at-least-once semantics, checkpointing, resumability, graceful degradation, backpressure, and safe handling of partial failure.
  • Hands-on observability experience (e.g., Prometheus/Grafana, LGTM stack), including designing metrics, not just consuming existing dashboards.
  • Deep CI/CD experience, ideally with GitLab CI, including dynamic/child pipelines and self-hosted runners; comfort with Docker and container-based test environments.
  • Experience with object storage (S3/Ceph or equivalent) and large-scale analytical stores (ClickHouse or other columnar databases).
  • Comfort designing state machines and long-running processes that persist across restarts, and reasoning about concurrency across multiple in-flight rollouts.
  • Excellent debugging skills across system, network, and data layers.
  • Strong communication skills and comfort working in a distributed team.
  • Proficiency in spoken and written English.

Nice to have

  • Experience with progressive delivery: canary and percentage-based rollouts, feature flags, automated rollback, and blast-radius control.
  • Experience running AI/LLM systems in production, particularly with cost control, token accounting, evaluation harnesses, and managing non-deterministic components within deterministic pipelines.
  • Experience with fleet-scale telemetry and building quality gates on top of noisy production signals.
  • Familiarity with WordPress, PHP, or WAF/ModSecurity concepts.
  • Experience with configuration management (Ansible, Puppet, Salt) and Linux service operations.

A cybersecurity background is not required. The primary challenges in this role involve orchestration, reliability, correctness under concurrency, and observability. Domain knowledge is learnable from our specialists; pipeline engineering judgment is what we cannot substitute.

We value engineers who are

  • Curious and fearless problem solvers: Eager to investigate existing systems, identify root causes, and propose improvements.
  • Skeptical by default: Questioning conclusions and trusting measurements over plausible reasoning.
  • Pragmatic and detail-oriented: Focused on building reliable, maintainable systems, avoiding solutions that rely on human memory.
  • Owners: Comfortable being accountable for pipeline correctness.
  • Effective communicators: Able to articulate ideas clearly, provide constructive feedback, and foster team collaboration.
  • Engaging and proactive: Contributing energy, initiative, and a positive presence to strengthen team culture.

Benefits

What's in it for you?

  • A strong focus on professional development.
  • Engaging and challenging projects.
  • Fully remote work with flexible hours, allowing you to schedule your day and work from any location worldwide.
  • Paid 24 days of vacation per year, 10 national holidays, and unlimited sick leave.
  • Compensation for private medical insurance.
  • Co-working and gym/sports reimbursement.
  • Budget for education.
  • Opportunity to receive a reward for innovative, patentable ideas.

By applying, you consent to the processing of your personal data as described in our Privacy Policy (https://cloudlinux.com/candidate-privacy-notice), which details how we maintain and handle your data.

Share

CloudLinux

CloudLinux

  • Address
    Warszawa, mazowieckie
View Profile
Your experience on this site will be improved by allowing cookies Cookie Policy