Senior Database Reliability Engineer at CloudLinux
Job Description
About CloudLinux / TuxCare
CloudLinux / TuxCare is a leading remote-first infrastructure and security company. With a team of over 300 engineers, they develop and operate cutting-edge products utilized by hosting providers, enterprises, and internal service teams globally. The Infrastructure Department is crucial, running platforms for CloudLinux OS, Imunify, KernelCare, TuxCare ELS, and various engineering systems.
The Opportunity: Senior Database Reliability Engineer
We are actively seeking a highly skilled and experienced Senior Database Reliability Engineer to join our Infrastructure DBA cell. This is a hands-on, production-focused ownership role, moving beyond typical ticket processing. Your primary mission will be to ensure the reliability of critical database services, streamline operations through automation, provide essential support to engineering teams, and reduce single-point dependencies across our PostgreSQL, ClickHouse, MongoDB, and Redis environments.
While deep expertise in PostgreSQL is a core requirement, prior experience with ClickHouse is a significant advantage, though not mandatory from day one. We require a senior engineer possessing extensive knowledge in database management, Linux systems, automation tools, and incident response, capable of quickly mastering and safely operating our ClickHouse environment.
Key Responsibilities of a Senior DRE:
- PostgreSQL Reliability: Take full ownership of production PostgreSQL reliability, encompassing HA design, Patroni, PgBouncer, replication strategies, failover procedures, database upgrades, vacuum/bloat control, query optimization, lock management, index tuning, capacity planning, robust backup solutions, Point-In-Time Recovery (PITR), and thorough restore validation processes.
- Disaster Recovery & Operational Excellence: Enhance disaster recovery capabilities and operational transparency by implementing tested restore procedures, documenting clear recovery paths, establishing measurable Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets, creating comprehensive runbooks, and devising safe maintenance plans.
- Database Ecosystem Support: Provide comprehensive support for the broader database estate, including ClickHouse, MongoDB, and Redis. This involves troubleshooting incidents, reviewing access controls and data safety changes, improving monitoring frameworks, and understanding existing production ClickHouse patterns.
- DBA Workflow Automation: Automate critical DBA workflows using tools like Ansible, Terraform/OpenTofu, GitLab CI/CD, custom scripts, and reproducible runbooks for provisioning, grant management, backups, restores, health checks, and ownership metadata.
- Self-Service DBaaS Development: Contribute to building Database-as-a-Service (DBaaS)-style self-service functionalities, enabling engineering teams to request databases, access credentials, and operational checks with minimal manual DBA intervention.
- Observability & Incident Response: Significantly improve observability and incident response through the effective utilization of Grafana, comprehensive metrics, centralized logs, Service Level Objectives (SLOs), precise alert rules, Opsgenie routing, and clear, concise communication during production incidents.
What Success Looks Like in this Role:
- PostgreSQL clusters demonstrate proven backup and restore capabilities, feature insightful dashboards, have clear ownership, and follow documented failover procedures.
- Repetitive DBA tasks are successfully transformed into automated or self-service workflows, enhancing efficiency.
- Operational knowledge of ClickHouse is effectively distributed, eliminating single-person dependencies.
- Database incidents are managed with clear owners, defined runbooks, supporting evidence, and measurable recovery paths.
- Product and engineering teams receive prompt database assistance without compromising on safety, auditability, or reliability standards.
Why Join CloudLinux?
- You will directly contribute to real production infrastructure powering CloudLinux and TuxCare products globally.
- Your work will have a tangible impact on system reliability, incident response, developer experience, and overall operational resilience.
- Thrive in an AI-assisted engineering culture where automation, comprehensive documentation, and AI tools like Claude and Codex are integrated into daily operations, complemented by rigorous human verification.
What We Expect From You:
- PostgreSQL Expertise: Possess deep, hands-on experience with PostgreSQL in business-critical production environments, typically demonstrated by 5+ years of relevant depth.
- PostgreSQL Internals: Strong understanding of PostgreSQL internals and operational nuances, including MVCC, WAL, transactions, locks, indexes, query planning, replication, autovacuum, bloat management, major upgrades, robust backup strategies, PITR, and meticulous restore testing.
- High Availability Databases: Proven experience with highly available database systems and the ability to critically analyze quorum, mitigate split-brain risks, manage failover, execute rollbacks, and ensure effective recovery.
- Linux & Infrastructure Fundamentals: Solid grasp of Linux and core infrastructure principles: systemd, networking protocols, storage solutions, filesystems, identifying CPU/memory/disk bottlenecks, TLS, DNS, firewall configurations, and systematic root-cause troubleshooting.
- Automation Skills: Proficiency in automation with Ansible and scripting. Experience with Terraform/OpenTofu, GitLab CI/CD, and merge-request based delivery methodologies are significant advantages.
- Multi-Database Support: Capability to effectively support multiple database engines. While not requiring immediate ClickHouse expertise, you must be eager to learn it rapidly and assume responsibility for its operation.
- AI Engineering Assistants: Practical experience utilizing AI engineering assistants such as Claude and Codex to enhance speed and quality, coupled with a commitment to personally verify all generated SQL, commands, scripts, and operational conclusions.
- English Fluency: Upper-intermediate or higher English proficiency is essential for clear and effective team communication.
Nice-to-Have Skills:
- ClickHouse Operations: Familiarity with ClickHouse replication, Keeper/ZooKeeper, MergeTree engines, distributed DDL, grant management, row policies, backup procedures, query troubleshooting, and cluster recovery.
- MongoDB & Redis: Experience with MongoDB replica sets and Percona Backup for MongoDB, as well as Redis/Sentinel and understanding broker/cache failure modes.
- Observability & DBaaS: Knowledge of database observability principles, SLOs, golden signals, alert tuning, executable incident runbooks, and experience in building internal platforms, self-service portals, or DBaaS workflows for engineering teams.
Benefits & Perks:
- Professional Growth: A strong emphasis on continuous professional development and career advancement.
- Challenging Work: Engage in interesting and stimulating projects that make a real impact.
- Flexible Remote Work: Enjoy a fully remote work environment with flexible working hours, offering the freedom to schedule your day and work from any location worldwide.
- Generous Time Off: Benefit from 24 paid vacation days per year, 10 national holidays, and unlimited sick leaves.
- Health & Wellness: Compensation for private medical insurance, and reimbursement for co-working spaces and gym/sports activities.
- Learning & Innovation: A dedicated budget for education and the exciting opportunity to earn a reward for innovative ideas that the company can patent.
Ready to Apply?
Take the next step in your career journey.
Apply NowYou will be redirected to the company's application page
Link verified about 3 hours ago
💜 Please mention that you found the job on True Work From Home, this helps us grow. Thanks!
More Backend Jobs
Discover similar opportunities that match your skills
Senior Backend Engineer
Senior Full-Stack Engineer, Internal Tools
Software Engineer - Infrastructure
IT Engineer - Remote
Full Stack Engineer (Frontend Oriented), Identity & Security
Lead Site Reliability Engineer, Imunify Reliability Platform
Senior Software Engineer, Backend - Core/API & Process Automation
Senior Software Engineer, Quality Engineering
About CloudLinux
CloudLinux is a software company that helps hosting providers and data centers make their servers more secure, stable, and efficient.
View Company Profile