Platform Engineer, PaaS at CloudLinux
Job Description
Description
CloudLinux is seeking a Senior Platform Engineer to join their Automation & Management Services cell within the Infrastructure Department. This role involves working with the PaaS team to maintain and improve platform services, focusing on reliability, automation, and self-service. The successful candidate will take ownership of agreed-upon platform services, ensuring their stability, making changes repeatable, and reducing operational overhead through automation. The position offers significant autonomy in implementing solutions.
This role encompasses aspects of site reliability engineering, ensuring services run smoothly day after day, and handling requests from product teams effectively. It also involves building new infrastructure, such as the observability platform and CI runner cluster, which were recently developed from scratch. The work will be split between platform engineering, incident response, and assisting other engineering teams in utilizing the provided services.
What you'll do
- Operate and maintain the observability platform, ensuring its health, onboarding new teams, monitoring costs and capacity, and managing the associated alerting system.
- Manage GitLab and the CI runner fleet, including upgrades, capacity planning, access control, and conducting backup and restore drills.
- Ensure the health of other services by establishing robust monitoring and runbooks required for production systems.
- Deploy new services as requested, which involves researching options, designing solutions, and setting up services from scratch following best practices (as code, monitored, backed up, documented).
- Address developers' requests related to access, onboarding, pipeline issues, and new exporters/dashboards. Transform recurring requests into self-service solutions.
- Lead incident response, diagnosing and mitigating impact, safely restoring services, and conducting root-cause analysis and post-mortems to implement prevention or detection improvements.
- Implement all changes as code, ensuring thorough review via merge requests, and planning/checking before every modification.
- Create technical documentation for external engineers, including runbooks, onboarding guides, maintenance notices, and actionable status updates.
- Collaborate with AI agents for tasks like data collection and drafting, reviewing their output as a colleague's merge request, and documenting learned insights for the team.
Requirements
Must have
- Senior-level experience in infrastructure, platform, or site reliability engineering, including responsibility for at least one production service. Candidates should be prepared to discuss their experience in detail, including incident response and improvements made.
- Proficiency in Linux systems administration and debugging on both bare metal and virtual machines, as a significant portion of the infrastructure is not Kubernetes-based.
- Experience with Kubernetes in production, delivered through GitOps, including performing cluster upgrades.
- Expertise in Infrastructure as Code (IaC) using Ansible and Terraform or OpenTofu, with changes reviewed via merge requests.
- Deep experience with GitLab administration and GitLab CI in production (self-hosted or SaaS). Comparable depth with another CI system is acceptable.
- Working knowledge of the Prometheus and Grafana ecosystem, including running it for a team, writing alert rules and dashboards, and reading PromQL.
- Strong ability to write clear technical explanations for engineers outside the team, such as runbooks, notices, and responses to requests.
- Excellent communication and interpersonal skills, as the role involves significant interaction with product teams to understand needs, agree on scope/priority/timing, provide polite pushback, and keep stakeholders informed.
- Advanced use of AI engineering assistants (e.g., Claude, Codex) for providing context, task breakdown, designing agent loops, and delegating plans for unattended execution within defined scope. This includes explaining, debugging, testing automation, and verifying generated commands/scripts before production deployment.
- Upper-intermediate or higher English proficiency for clear team communication.
Nice to have
- Experience in alerting design, including SLOs, burn-rate alerts, and data-driven threshold sizing.
- Knowledge of MicroVM isolation for CI (Kata Containers, Firecracker, or gVisas).
- Experience with S3-compatible object storage operations (e.g., Ceph RGW).
- AWS experience with real cost optimization work.
- Experience maintaining self-hosted Sentry or other Kafka, ClickHouse, and Redis-backed applications under load.
- Proficiency in Python or Go for developing exporters and small internal services.
Solid fundamentals and the ability to learn and manage unfamiliar services, making them observable and documenting runbooks, are prioritized over matching every specific system on the list.
What this role is not focused on
- This is not a ticket-queue operator role; recurring requests are automated into self-service.
- This is not exclusively a cloud or Kubernetes role; bare metal and virtual machines are a significant part of the infrastructure.
- This is not a DBA, network engineer, or security engineer position; those teams manage their own systems, with the platform provided by this team for monitoring.
- A strong focus on professional development and continuous learning.
- Engaging and challenging projects.
- Fully remote work with flexible working hours, allowing you to work from any location worldwide.
- 24 days of paid vacation per year, 10 national holidays, and unlimited sick leaves.
- Compensation for private medical insurance.
- Reimbursement for co-working spaces and gym/sports activities.
- A dedicated budget for education.
- Opportunity to receive a reward for innovative ideas that the company can patent.
Benefits
What's in it for you?
By applying, you consent to data processing as per CloudLinux's Privacy Policy.
Ready to Apply?
Take the next step in your career journey.
Apply NowYou will be redirected to the company's application page
💜 Please mention that you found the job on True Work From Home, this helps us grow. Thanks!
More DevOps Jobs
Discover similar opportunities that match your skills
Senior Security Engineer - Blue Team
Senior Software Engineer, Quality Engineering
Senior Software Engineer, CI/CD Branching Platform
Engineer to own the automated pipelines
Senior Security Software Engineer
Staff Data Scientist at DuckDuckGo
Staff Backend Engineer (Social)
AWS Practice Lead, Open Source Databases
About CloudLinux
CloudLinux is a software company that helps hosting providers and data centers make their servers more secure, stable, and efficient.
View Company Profile