Lead Site Reliability Engineer, Imunify Reliability Platform at CloudLinux
Job Description
The Challenge
You will own the reliability of Imunify360, a comprehensive multi layer Linux server security suite. This suite - comprising WAF, IDS/IPS, malware scanning and cleanup, proactive defense, patch management, and reputation systems - operates as an agent on hundreds of thousands of customer servers. It's supported by a cloud infrastructure of scanning, correlation, and signature-delivery services running on our own bare metal. Currently, approximately 70 components lack defined service level indicators (SLIs).
While some are internal services we can monitor, many are agent side subsystems on customer owned machines, reporting via a heartbeat not originally designed for reliability tracking. Existing monitoring and dashboards fail to provide a clear answer to the critical question: "Is this component performing its job correctly right now, and how would we detect if it stopped?" We recently experienced the severe consequences of this gap when a security control was silently disabled across a large portion of the fleet for 61 days. Despite this, all dashboards remained green. Telemetry reported a ruleset version but not whether the consuming control was active, making a configuration change indistinguishable from a broken updater. Three independent safety mechanisms failed because they were gated by the same condition that caused the primary failure.
Your mission is to ensure that such failures are detectable within hours, not months, across the entire product line, and to build a robust system that maintains this detectability as the product evolves. This role offers a greenfield charter within an existing brownfield environment. You will not inherit an SRE team, an SLO framework, or an existing paging culture; instead, you will define these, in collaboration with engineering leads, and ensure their successful implementation.
Key Responsibilities
1. Define "Working" for ~70 Components
- Lead SLI definition sessions with squad leads and senior engineers, facilitating the process and upholding standards. The owning squad will ultimately sign off on the SLI.
- Develop a taxonomy tailored to the product's actual needs, extending beyond basic availability and latency metrics.
- Enforce a strict design rule: an SLI must be measurable externally from the component it monitors. If a control being off also disables its monitoring signal, the SLI is invalid. This principle is a direct outcome of past incidents and the core reason for this role.
- Assign a Service Level Objective (SLO), an error budget, and an owning squad to each component. Tiering is expected, meaning not every component will require a 99.9% target or a dedicated pager.
2. Build the Collection System
- Design and construct the pipeline for collecting these indicators from the fleet into a queryable store. This pipeline should be push based, sampled, privacy constrained, and adhere to a cardinality budget that you will define and defend.
- Collaborate with owning squads to extend agent side and service side instrumentation in Python, Go, and Rust where necessary to capture missing signals.
- Consolidate the current disparate dashboards, ad hoc queries, and reporting paths into a cohesive and maintainable set of instruments, retiring any elements that do not prove their value.
3. Build Alerting and Alert Management
- Implement symptom based, SLO anchored alerting with multi-window burn rate semantics, moving away from a proliferation of basic thresholds.
- Establish a three tier alerting taxonomy (page / ticket / dashboard) with clear rules dictating what events warrant paging a human at 3:00 AM.
- Ensure every alert comes with an owner, a runbook, and a documented failure mode; otherwise, it will not be deployed.
- Foster alert hygiene as a continuous practice: conduct quarterly reviews, celebrate alert deletions as wins, and track actionable alert rates. A persistent failure to perform security relevant refreshes, which currently only logs a warning, should trigger a page.
4. Build Escalation
- Develop and maintain a machine readable, current map of component to owning squad ownership. This will be wired into routing to ensure alerts reach the appropriate individuals, rather than a generic shared channel.
- Design a severity matrix, acknowledgement Service Level Agreements (SLAs), follow the sun rota across UTC−5 to UTC+8, and a clean handoff protocol.
- Implement incident command practices and blameless postmortems within 24 hours. While postmortems are already conducted honestly, including public retractions of incorrect findings, you will elevate the mechanical aspects - timelines, ownership, and action item follow through.
- Architect the escalation system so that squads manage their own pagers. Your role is to build and operate the platform and coach on best practices, not to serve as a buffer for other squads' alerts.
Required Qualifications
- Substantial production engineering or Site Reliability Engineering (SRE) experience, including at least one environment where you initiated and defined the SLO framework, rather than inheriting an existing one. Be prepared to discuss SLIs you personally crafted and how you negotiated them with resistant teams.
- Strong proficiency in Python.
- Comfortable reading and modifying Go or Rust, as our agents are developed in these languages, and instrumentation will involve them.
- Deep practical understanding of time series and event telemetry at scale, including Prometheus/OpenMetrics, Grafana, an Alertmanager class routing layer, and a columnar store for high cardinality fleet data (e.g., ClickHouse).
- Experience with distributed systems debugging on bare metal and long lived hosts. Much of this environment does not use Kubernetes, so skills assuming an orchestrator will not directly transfer.
- Proficiency in configuration management and Continuous Integration (CI) at production scale, using tools like Ansible, GitLab CI, Jenkins, or similar.
- The analytical judgment to design measurement systems for machines you do not own and cannot directly scrape, considering push telemetry, sampling, clock skew, partial reporting, and privacy constraints inherent in running on customer servers.
- Excellent written communication skills for effective asynchronous collaboration. This role is approximately 40% telemetry engineering, 40% achieving consensus among sixty engineers on "healthy" definitions, and 20% ensuring that the agreed answer is not merely an unread dashboard.
Valuable Qualifications (Nice to Haves)
- Background in security products such as WAF, EDR, AV, or vulnerability management, and an intuitive understanding that a security control's SLI is focused on enforcement, not just uptime.
- Experience with monitoring under audit frameworks like SOC 2 CC7.x, ISO 27001 A.8.16, or NIST SP 800-137 for continuous monitoring. Experience writing for this audience is beneficial.
- Familiarity with OpenTelemetry, eBPF, and Sentry.
- Proficiency in cost and cardinality aware telemetry design.
- Fluency with agentic development tooling, as we operate a Cursor/Claude first SDLC with internal and third party MCP servers, and engineers are evaluated on their ability to work with these tools.
- Knowledge of Kubernetes, relevant for the single workload that currently utilizes it.
This Role Is Not
- A DevOps ticket queue.
- Ownership of the build system.
- Cloud cost management.
- The on call rota for other squads' services.
First Year Outcomes
30 Days:
- Completion of component inventory with named owners.
- SLI taxonomy and tiering are agreed upon.
- Three pilot components fully instrumented end to end as the reference implementation.
90 Days:
- Collection pipeline is in production.
- Tier 1 components (those whose failure poses a customer security exposure) have an SLO, alert, runbook, and owner.
- Escalation routing for Tier 1 components is live.
180 Days:
- All approximately 70 components have a defined SLI and an owner.
- Alert taxonomy is enforced; page volume and actionable rate are measured and published.
- Squad on call operations are active.
365 Days:
- Mean time to detect a silent control degradation is under 24 hours, measured against a 61 day baseline.
- Error budget policy influences release decisions.
- The function is documented sufficiently for future hires (#2 and #3) to be additive, rather than performing archaeological investigations.
Our Work Culture
We are a remote first and asynchronous organization operating across nine time zones. Our cadence includes weekly Product Owner (PO) and architecture syncs; monthly demos and OKR reviews; and quarterly architecture summits. Decisions are documented as Architecture Decision Records (ADRs). Every output has a clear owner and a due date. Postmortems are blameless and published, and we transparently correct our own findings when proven wrong.
Benefits
What we offer:
- A strong emphasis on professional development, providing ample opportunities for learning and career advancement.
- Fully remote work with flexible hours, allowing you to structure your day and work from any global location.
- Generous paid time off: 24 days of vacation per year, 10 national holidays, and unlimited sick leaves to promote a healthy work life balance.
- Compensation for private medical insurance.
- Reimbursement for co-working spaces and gym/sports memberships.
- The chance to earn rewards for innovative ideas that the company can patent, fostering a culture of creativity.
By applying for this position, you consent to the processing of your personal data as described in our Privacy Policy (https://cloudlinux.com/candidate-privacy-notice), which details how we manage and protect your data.
Ready to Apply?
Take the next step in your career journey.
Apply NowYou will be redirected to the company's application page
Link verified 2 days ago
💜 Please mention that you found the job on True Work From Home, this helps us grow. Thanks!
More DevOps Jobs
Discover similar opportunities that match your skills
Senior Android Developer
FinOps Cloud Financial Engineer at Supabase
Site Reliability Engineering Manager
Experienced Software Engineer
Engineering Director, Web Platform
Lead Infrastructure Engineer
Software Engineer (Trading)
Platform Engineer, Compute Capacity
About CloudLinux
CloudLinux is a software company that helps hosting providers and data centers make their servers more secure, stable, and efficient.
View Company Profile