Build your online resume. Claim your username
CloudLinux logo

Senior Platform Engineer - Remote at CloudLinux

Worldwide Remote 🌍 Work from Anywhere Full time Senior Posted  Apply before Nov 08, 2026

Job Description

CloudLinux is seeking a Senior Platform Engineer to join their Automation & Management Services cell within the Infrastructure Department. In this role, you will be instrumental in solving complex problems across various teams and services, including cloud-cost data, infrastructure inventory, network policies, and capacity workflows. CloudLinux is renowned for building robust Linux infrastructure and security products.

About the Team

The Platform cell is a small, dynamic team responsible for running the company's observability platform, GitLab instances, CI runners, and essential automation for provisioning and configuration. As a Platform Engineer, you will take ownership of key platform services, ensuring their reliability, adherence to service levels, and making changes and recovery processes repeatable. A core focus will be on reducing operational work through automation and self-service solutions.

This position offers significant autonomy; you will choose your implementation approach, defend it during reviews, and be accountable for the service's performance. The role primarily involves Site Reliability Engineering (SRE) principles, ensuring services run seamlessly day-to-day and product team requests are handled efficiently. Expect a balance of building new systems, such as the observability platform and runner cluster built from scratch recently, and operational support, incident response, and assisting other engineering teams.

What You Will Do

  • Operate the Observability Platform: Maintain its health, onboard new teams, monitor costs and capacity, and manage the underlying alerting system.
  • Manage GitLab & CI Runner Fleet: Handle upgrades, capacity planning, access management, and conduct backup and restore drills.
  • Ensure Service Health: Keep all services healthy with robust monitoring and comprehensive runbooks.
  • Deploy New Services: Research options, design solutions, and implement new services from scratch following best practices: as code, monitored, backed up, and well documented.
  • Address Developer Requests: Support developers with access, onboarding, pipeline issues, and new exporters/dashboards. Automate recurring requests into self-service solutions.
  • Lead Incident Response: Diagnose and mitigate impact, safely restore services, perform root-cause analysis and post-mortems, and implement preventive measures.
  • Practice Infrastructure as Code: Ship all changes as code, reviewed via merge requests, with thorough planning and checks before deployment.
  • Create Technical Documentation: Write clear runbooks, onboarding guides, maintenance notices, and status updates for external engineering teams.
  • Leverage AI Agents: Delegate data collection and drafting tasks to AI, review their output, and document learnings for team knowledge sharing.

Requirements

Must Have

  • Senior-level SRE Experience: Extensive experience in infrastructure, platform, or site reliability engineering, including direct responsibility for a production service. Be prepared to discuss challenges, detection methods, and improvements.
  • Linux Systems Administration: Proficient in Linux system administration and debugging on bare metal and virtual machines, as much of the infrastructure is not Kubernetes-based.
  • Production Kubernetes: Hands-on experience with production Kubernetes delivered via GitOps, including performing cluster upgrades.
  • Infrastructure as Code: Expertise with Ansible and Terraform or OpenTofu for infrastructure provisioning, with changes reviewed in merge requests.
  • GitLab Administration: Deep experience with GitLab administration and GitLab CI in production (self-hosted or SaaS). Equivalent depth with another CI system is acceptable.
  • Prometheus & Grafana Ecosystem: Working knowledge of running Prometheus and Grafana for a team, writing alert rules and dashboards, and understanding PromQL.
  • Technical Communication: Proven ability to write clear technical explanations for engineers (runbooks, notices, request responses).
  • Strong Communication & Interpersonal Skills: Excellent communication for understanding product team needs, agreeing on scope, and keeping stakeholders informed. A collaborative and approachable attitude is essential.
  • Advanced AI Engineering Assistant Use: Proficient in using AI tools like Claude and Codex for context provision, task breakdown, agent loop design, and delegating automated execution within defined scopes and permissions. Ability to explain, debug, test, and verify AI-generated output before production use.
  • English Proficiency: Upper-intermediate or higher English for effective team communication.

Nice to Have

  • Alerting design experience (SLOs, burn-rate alerts, data-driven thresholds).
  • MicroVM isolation for CI (Kata Containers, Firecracker, or gVisor).
  • S3-compatible object storage operations (e.g., Ceph RGW).
  • Experience with AWS cost management.
  • Managing self-hosted Sentry or other Kafka, ClickHouse, and Redis-backed applications under load.
  • Proficiency in Python or Go for developing exporters and small internal services.

Solid fundamentals and the ability to quickly understand, make observable, and document unfamiliar services are prioritized over matching every item on this list.

What This Role Is Not Focused On

  • This is not a ticket-queue operator role; recurring requests are automated into self-service solutions.
  • This is not solely a cloud or Kubernetes role; bare metal and virtual machines are a significant part of the infrastructure.
  • This is not a DBA, network engineer, or security engineer role; those teams manage their own systems, with this role providing the monitoring platform.

Benefits & Perks

  • Professional Growth: Strong emphasis on professional development with a dedicated learning budget.
  • Challenging Projects: Engage in interesting and stimulating projects.
  • Fully Remote & Flexible: Enjoy fully remote work with flexible working hours, allowing you to work from any location worldwide.
  • Generous Time Off: Receive 24 days of paid vacation, 10 national holidays, and unlimited sick leaves annually.
  • Health & Wellness: Compensation for private medical insurance, gym/sports reimbursement, and coworking budget.
  • Innovation Rewards: Opportunity to earn rewards for innovative, patentable ideas.

By applying, you consent to the processing of your personal data as outlined in CloudLinux's Privacy Policy, which details data handling practices.

Ready to Apply?

Take the next step in your career journey.

Apply Now

You will be redirected to the company's application page

Link verified about 1 hour ago

💜 Please mention that you found the job on True Work From Home, this helps us grow. Thanks!