Senior Site Reliability Engineer at CloudLinux
Job Description
About CloudLinux
CloudLinux is a remote-first, global company driven by principles of doing the right thing, prioritizing employees, and delivering high volume, low cost Linux infrastructure and security products. Our mission is to enhance operational efficiency for businesses. We foster a supportive team environment where every individual contributes to collective success. Learn more at cloudlinux.com.
The Challenge: Imunify360 Security Observability
Imunify360 is a comprehensive multi-layer Linux server security suite, incorporating WAF, IDS/IPS, malware scanning, proactive defense, and patch management. It operates as an agent on hundreds of thousands of customer servers, supported by a cloud infrastructure for scanning, correlation, and signature delivery. With roughly 70 components lacking defined service level indicators (SLIs), there's a critical gap in knowing if each component is functioning correctly. A recent incident highlighted this when a security control was silently disabled across a large portion of the fleet for 61 days, despite all dashboards appearing green. The telemetry reported ruleset versions but not the active status of the control consuming them, making configuration changes indistinguishable from failures. This role exists to prevent such failures by making them detectable within hours, not months, across the entire product line.
Your Mission
You will lead a greenfield charter within an existing infrastructure, defining the SRE team, SLO framework, and paging culture. Your key responsibilities include:
- Define "Working" for ~70 Components: Collaborate with squad leads and senior engineers to define SLIs. You will facilitate the process and maintain standards, ensuring owning squads sign off on their SLIs. Develop a taxonomy beyond just availability and latency, enforcing the crucial rule that an SLI must be measurable externally. Attach Service Level Objectives (SLOs), error budgets, and owning squads to each component, with appropriate tiering based on criticality.
- Build the Collection System: Design and implement a robust pipeline to collect indicators from the fleet into a queryable store. This system will be push based, sampled, privacy constrained, and adhere to a defined cardinality budget. Extend agent side and service side instrumentation using Python, Go, and Rust in collaboration with owning squads. Consolidate disparate dashboards and reporting paths into a coherent set of instruments.
- Build Alerting and Alert Management: Implement symptom based, SLO anchored alerting with multi window burn rate semantics, moving away from simple threshold based alerts. Establish a three tier alert taxonomy (page / ticket / dashboard) with clear rules for when a human is paged at 3 AM. Ensure every alert includes an owner, a runbook, and a documented failure mode. Maintain alert hygiene through quarterly reviews, celebrating alert deletions, and tracking actionable rates.
- Build Escalation: Create a machine readable, current component to owning squad ownership map, integrated with routing to ensure alerts reach the correct team. Design a severity matrix, acknowledgement SLAs, and follow the sun rota across global time zones (UTC-5 to UTC+8) with a clear handoff protocol. Implement incident command practices and blameless postmortems within 24 hours, enhancing the mechanical aspects of incident response. Design the system so squads manage their own pagers, with you building the platform and coaching on best practices.
What You'll Bring
Required Skills & Experience:
- Substantial SRE Experience: Proven production engineering or SRE experience, specifically having defined an SLO framework from scratch rather than inheriting one. Be prepared to discuss SLIs you personally authored and how you negotiated them with resistant teams.
- Programming Proficiency: Strong command of Python, with comfort in reading and modifying Go or Rust for agent side instrumentation.
- Telemetry Expertise: Deep practical understanding of time series and event telemetry at scale, including Prometheus/OpenMetrics, Grafana, Alertmanager class routing, and columnar stores (e.g., ClickHouse).
- Distributed Systems Debugging: Experience with distributed systems debugging on bare metal and long lived hosts, recognizing that orchestrator based reflexes may not apply.
- Configuration Management & CI: Proficiency with production scale configuration management and CI tools like Ansible, GitLab CI, or Jenkins.
- Measurement Design: The ability to design measurement for systems you do not own or cannot scrape, considering push telemetry, sampling, clock skew, partial reporting, and privacy constraints.
- Communication: Excellent written communication skills for effective asynchronous collaboration. This role involves significant technical work and crucial alignment across sixty engineers on system health.
Valuable (Nice to have):
- Security Product Background: Experience with WAF, EDR, AV, vulnerability management, and an instinct that security control SLIs are about enforcement, not just uptime.
- Monitoring Under Audit: Familiarity with audit standards like SOC 2 CC7.x, ISO 27001 A.8.16, NIST SP 800-137.
- Advanced Telemetry Tools: Experience with OpenTelemetry, eBPF, Sentry.
- Cost & Cardinality Aware Telemetry: Design telemetry with cost and cardinality in mind.
- Agentic Development: Fluency with tools like Cursor/Claude first SDLC for enhanced engineering efficiency.
- Kubernetes Experience: Familiarity with Kubernetes for specific workloads.
This Role is NOT:
- A DevOps ticket queue.
- Build system ownership.
- Cloud cost management.
- On call rota for other squads' services.
First Year Outcomes
- 30 Days: Component inventory with named owners, SLI taxonomy and tiering agreed upon. Three pilot components fully instrumented as a reference.
- 90 Days: Collection pipeline in production. Tier-1 components (critical for customer security) have SLOs, alerts, runbooks, and owners. Escalation routing live for Tier-1.
- 180 Days: All ~70 components have defined SLIs and owners. Alert taxonomy enforced; page volume and actionable rates measured and published. Squad on call operating.
- 365 Days: Mean time to detect silent control degradation is under 24 hours (measured against a 61 day baseline). Error budget policy influences release decisions. The function is documented for future hires.
Our Work Culture
We operate as a remote first and asynchronous team across nine time zones. Our cadence includes weekly PO and architecture syncs, monthly demos and OKR reviews, and quarterly architecture summits. Decisions are documented as ADRs. Every output has a clear owner and due date. Postmortems are blameless, published, and we publicly correct any erroneous findings.
What's in it for You?
- Strong focus on professional development with continuous learning and growth opportunities.
- Fully remote work with flexible hours, allowing you to work from any location worldwide.
- Generous paid time off: 24 vacation days, 10 national holidays, and unlimited sick leaves for a healthy work life balance.
- Compensation for private medical insurance.
- Reimbursement for co working spaces and gym/sports memberships.
- Opportunity for rewards for innovative, patentable ideas.
Ready to Apply?
Take the next step in your career journey.
Apply NowYou will be redirected to the company's application page
💜 Please mention that you found the job on True Work From Home, this helps us grow. Thanks!
More DevOps Jobs
Discover similar opportunities that match your skills
Staff Software Engineer, SecureChain
Security Engineer
Senior Software Engineer, Backend - Core/API & Process Automation
Staff ML and AI Agent Systems Engineer
Offensive Security Engineer
Staff Backend Engineer (Social)
Senior Software Engineer
Senior Software Engineer, Engineering Operations
About CloudLinux
CloudLinux is a software company that helps hosting providers and data centers make their servers more secure, stable, and efficient.
View Company Profile