Build your online resume. Claim your username
CloudLinux logo

Engineer to own the automated pipelines at CloudLinux

Worldwide 🌍 Work from Anywhere Full time Senior Posted  Apply before Oct 19, 2026

Job Description

About CloudLinux and Imunify360

CloudLinux is a global, remote-first company dedicated to delivering high-volume, low-cost Linux infrastructure and security products. Our core principles are doing the right thing, prioritizing employees, embracing remote work, and ensuring success through mutual support. Discover more at cloudlinux.com.

Imunify360 Security Suite, a product of CloudLinux Inc., is a leading security solution for shared and VPS/Dedicated servers. It provides comprehensive, six-layer attack prevention through an automated, easy to use system. We are seeking an experienced engineer to take ownership of the automated pipelines that transform threat intelligence into robust protection for millions of websites. This critical role involves ensuring our automated systems, which operate 24/7, reliably detect new vulnerabilities, generate WAF rules, create tests, validate rules against legitimate traffic, and deploy across a vast network, with the capability to automatically roll back if issues arise. We need a strong engineer to rapidly expand these systems while maintaining their reliability and transparency. This position involves solving complex engineering challenges and working with data at scale in a dynamic environment. It is not an analyst or research role, though related experience is a plus. We are looking for someone who builds resilient, reliable systems with sophisticated underlying logic.

What You Would Own

You will be responsible for scaling and operating critical production systems:

  • Automated Protection Pipeline: The end to end chain from threat intelligence to a validated, deployed rule. This multi-stage, largely autonomous system must complete within a fixed daily time window.
  • Progressive Release Automation: Controlled, staged rollout of protection across the fleet, incorporating automated guardrails for holding or rolling back stages without human intervention.
  • Quality Gates: Systems for assessing live production signals to determine if deployed features are causing harm and taking immediate corrective action. This requires balancing the risks of missing problems against over-correcting and silently removing protection.
  • CI at Scale: Validation processes that establish real, disposable environments across a broad matrix of software versions and configurations, delivering trustworthy verdicts quickly to meet release windows.
  • LLM Orchestration and Cost Control: Managing several AI agent driven subsystems, including evaluation, budget management, and spend accounting as core components.
  • Observability and Alerting: Building systems for the pipelines to report their own condition, prove health, and escalate issues autonomously. This operates on a petabyte scale threat intelligence store and live telemetry from over 60 million websites, with strict end to end latency budgets measured in hours. A primary goal is to eliminate silent budget misses.

Specific details will be discussed during the interview process.

Key Responsibilities

  • Design, build, and operate the automated pipelines from end to end.
  • Transform fragile multi-stage batch jobs into resumable, idempotent, observable systems with clear state machines and recovery paths.
  • Define and enforce latency budgets and Service Level Objectives (SLOs) per stage, ensuring violations are visible and actionable.
  • Build the observability layer, including metrics, dashboards, alerting, and health gates, to enable the pipeline to self report its condition.
  • Design and implement guardrails: automatic hold and rollback mechanisms, blast radius limits, kill switches, and safe by default behavior when dependencies are unavailable.
  • Develop low maintenance systems, eliminating manual steps, reducing the need for human monitoring, and simplifying operational surface area.
  • Write and maintain unit and integration tests for complex logic involving concurrency, partial failure, external API flakiness, and multi-stage state.
  • Investigate and resolve intricate issues across technologies such as ClickHouse, GitLab CI, S3/object storage, Prometheus/Grafana, and third party APIs.
  • Collaborate with security analysts and the Server team on architecture, providing critical feedback on designs to ensure production readiness.

Requirements

  • 5+ years of professional backend, platform, or infrastructure engineering experience.
  • Demonstrable experience building and operating multi-stage data or automation pipelines, such as CI/CD systems, ETL/ELT, build and release automation, job orchestration, or ML/data platforms. This is the most crucial requirement, and you will be asked to detail your experience.
  • Deep proficiency in at least one of Python, Go (Golang), or Rust. We utilize all three, and the specific language is less important than the depth of your experience, which will be verified through detailed questions about systems you have designed and shipped.
  • Strong systems design judgement, emphasizing architectural decisions and anticipating failure points over raw coding throughput.
  • Practical experience with workflow orchestration and job scheduling tools like Airflow, Temporal, Prefect, Dagster, Argo, or custom schedulers, demonstrated by actual production deployments.
  • A robust understanding of reliability engineering principles: idempotency, retries with backoff, exactly-once vs at-least-once processing, checkpointing, resumability, graceful degradation, backpressure, and safe handling of partial failure.
  • Hands on observability experience with tools like Prometheus/Grafana or equivalent, including designing metrics, not just consuming dashboards.
  • Extensive CI/CD experience, particularly with GitLab CI (including dynamic/child pipelines and self hosted runners) and comfort with Docker and container based test environments.
  • Experience with object storage (S3/Ceph or similar) and large-scale analytical stores like ClickHouse or other columnar databases.
  • Comfort designing state machines and long running processes that withstand restarts, and effectively reasoning about concurrency across multiple in flight rollouts.
  • Excellent debugging skills across system, network, and data layers.
  • Strong communication skills and comfort working effectively in a distributed team.
  • Proficiency in spoken and written English.

Nice to Have

  • Experience with progressive delivery: canary and percentage based rollouts, feature flags, automated rollback, and blast radius control.
  • Experience running AI/LLM systems in production, including cost control, token accounting, evaluation harnesses, and managing non deterministic components within deterministic pipelines.
  • Experience with fleet-scale telemetry and building quality gates on top of noisy production signals.
  • Familiarity with WordPress, PHP, or WAF/ModSecurity concepts.
  • Experience with configuration management (Ansible, Puppet, Salt) and Linux service operations.

A cybersecurity background is not required. The primary challenges in this role involve orchestration, reliability, correctness under concurrency, and observability. Domain knowledge is acquirable, and specialists are available for learning; pipeline engineering judgement is paramount.

We Value Engineers Who Are

  • Curious and Fearless Problem Solvers: Eager to investigate existing systems, identify root causes, and propose improvements.
  • Sceptical by Default: Prioritizing evidence and measurements over plausible reasoning.
  • Pragmatic and Detail Oriented: Focused on building reliable, maintainable systems and averse to solutions requiring manual intervention.
  • Owners: Accepting accountability for pipeline performance and correctness.
  • Effective Communicators: Articulating ideas clearly, providing constructive feedback, and fostering team collaboration.
  • Engaging and Proactive: Contributing positively to team culture with energy and initiative.

Benefits

Joining CloudLinux offers a range of advantages:

  • A strong focus on professional development and growth.
  • Engaging and challenging projects.
  • Fully remote work with flexible hours, allowing you to plan your day and work from any location worldwide.
  • 24 days of paid vacation per year, 10 national holidays, and unlimited sick leaves.
  • Compensation for private medical insurance.
  • Reimbursement for co-working spaces and gym/sports memberships.
  • Dedicated budget for education and learning.
  • Opportunity to receive a reward for innovative, patentable ideas.

By applying for this position, you consent to the processing of your personal data as described in our Privacy Policy.

Ready to Apply?

Take the next step in your career journey.

Apply Now

You will be redirected to the company's application page

Link verified 1 day ago

💜 Please mention that you found the job on True Work From Home, this helps us grow. Thanks!

More Software Development Engineer (SDE) Jobs

Discover similar opportunities that match your skills