Senior Database Reliability Engineer
Bangun dan perbaiki kehandalan layanan database kritis di CloudLinux
Sebagai Senior Database Reliability Engineer di CloudLinux, kamu akan menjaga layanan database PostgreSQL, ClickHouse, MongoDB, dan Redis tetap handal. Kamu akan mengotomatiskan pekerjaan berulang, mendukung tim engineering, dan mengurangi ketergantungan pada satu orang. Kamu akan bekerja dengan teknologi seperti Patroni, PgBouncer, dan Ansible.
Kenapa Menarik?
Bergabung dengan tim yang berfokus pada kehandalan dan keamanan infrastruktur, dengan kesempatan untuk bekerja dengan teknologi terbaru.
Tanggung Jawab Utama
- Mengelola kehandalan PostgreSQL di produksi: desain HA, Patroni, PgBouncer, replikasi, failover, dan upgrade
- Meningkatkan pemulihan bencana dan bukti operasional: restorasi yang diuji, jalur pemulihan yang terdokumentasi, dan rencana pemeliharaan ya
- Mendukung estate database yang lebih luas: ClickHouse, MongoDB, dan Redis
- Mengotomatiskan alur kerja DBA dengan Ansible, Terraform, GitLab CI/CD, dan runbook yang dapat direproduksi
- Membangun kemampuan self-service DBaaS agar tim engineering dapat meminta database, akses, dan kredensial dengan sedikit intervensi DBA manu
- Meningkatkan observabilitas dan tanggapan insiden melalui Grafana, metrik, log, SLO, dan aturan alert
Persyaratan
- Pengalaman senior dalam database, Linux, otomatisasi, dan tanggapan insiden
- Kemampuan untuk belajar dan mengoperasikan lingkungan ClickHouse dengan cepat
- Pengalaman dengan PostgreSQL, Patroni, PgBouncer, dan teknologi terkait
Skills Wajib
Keywords
Lihat Deskripsi Asli dari RemoteOK
Deskripsi asli dari RemoteOK
CloudLinux / TuxCare is a remote-first infrastructure and security company. More than 300 engineers build and operate products used by hosting providers, enterprises, and internal service teams worldwide. Our Infrastructure Department runs the platforms behind CloudLinux OS, Imunify, KernelCare, TuxCare ELS, and our engineering systems. We are hiring a Senior Database Reliability Engineer to join the Infrastructure DBA cell. This is a hands-on production ownership role, not a narrow ticket-processing DBA position. You will keep critical database services reliable, automate repeated work, support engineering teams, and reduce single-person dependency in our PostgreSQL, ClickHouse, MongoDB, and Redis operations. PostgreSQL is the main requirement. ClickHouse experience is a strong plus, but it is not a day-one blocker. We need a senior engineer with enough database, Linux, automation, and incident-response depth to learn our ClickHouse environment quickly and operate it safely. Your Responsibilities: Own production PostgreSQL reliability: HA design, Patroni, PgBouncer, replication, failover, upgrades, vacuum/bloat control, query tuning, locks, indexes, capacity, backups, PITR, and restore validation. Improve disaster recovery and operational evidence: tested restores, documented recovery paths, measurable RTO/RPO targets, runbooks, and safe maintenance plans. Support the wider database estate: ClickHouse, MongoDB, and Redis. You will troubleshoot incidents, review access and data-safety changes, improve monitoring, and learn the production ClickHouse patterns already in use. Automate DBA workflows with Ansible, Terraform/OpenTofu, GitLab CI/CD, scripts, and reproducible runbooks for provisioning, grants, backups, restores, health checks, and ownership metadata. Help build DBaaS-style self-service capabilities so engineering teams can request databases, access, credentials, and operational checks with less manual DBA intervention. Improve observability and incident response through Grafana, metrics, logs, SLOs, alert rules, Opsgenie routing, and clear communication during production issues. What Success Looks Like: PostgreSQL clusters have tested backup and restore paths, useful dashboards, clear ownership, and documented failover procedures. Repeated DBA tickets become automation or self-service workflows. ClickHouse operational knowledge is no longer a single-person dependency. Database incidents have owners, runbooks, evidence, and measurable recovery paths. Product and engineering teams get database help faster without sacrificing safety, auditability, or reliability. Why CloudLinux? You will work on real production infrastructure used across CloudLinux and TuxCare products. You will have a direct impact on reliability, incident response, developer experience, and operational resilience. You will also work in an AI-assisted engineering culture where automation, documentation, Claude, Codex, and careful human verification are part of the daily operating model. What We Expect From You: Deep hands-on PostgreSQL experience in business-critical production environments, typically 5+ years or equivalent depth. Strong understanding of PostgreSQL internals and operations: MVCC, WAL, transactions, locks, indexes, query planning, replication, autovacuum, bloat, major upgrades, backups, PITR, and restore testing. Proven experience with highly available databases and the ability to reason about quorum, split-brain risk, failover, rollback, and recovery. Strong Linux and infrastructure fundamentals: systemd, networking, storage, filesystems, CPU/memory/disk bottlenecks, TLS, DNS, firewalls, and root-cause troubleshooting. Automation skills with Ansible and scripting. Terraform/OpenTofu, GitLab CI/CD, and merge-request based delivery are strong advantages. Ability to support more than one database engine. You do not need to be a ClickHouse expert on day one, but you must be ready to learn it quickly and take responsibility for it. Practical
Situs sumber mungkin diblokir ISP Indonesia
Beberapa ISP Indonesia (Telkomsel, Indihome) memblokir RemoteOK. Kalau tombol Apply tidak terbuka, coba pakai data seluler atau VPN.
Tips: ganti jaringan atau aktifkan VPN, lalu klik Apply lagi.
Pemberi kerja tidak menyatakan keterbukaan lokasi — cek langsung lowongannya
Jelajahi lebih lanjut
Data & laporan pasar
Riset gaji & permintaan skill dari data lowongan kami sendiri.
- Lowongan IT Indonesia vs Remote Global (2026)Analisis data primer 2.049 lowongan: metodologi, klasifikasi, dataset bisa diunduh.
- Permintaan Skill AI: Indonesia vs Global (2026)10.000+ lowongan, classifier taxonomy-first, Wilson CI, pra-registrasi sebelum analisis.
- Laporan Hiring Indonesia: Tech vs Non-TechPermintaan lowongan per bidang dari hitungan agregat — bukan listing per-listing.
- Benchmark Gaji IndonesiaKisaran gaji agregat lintas peran, dengan metodologi dan dataset terbuka.
- Indeks Gaji & Permintaan Kerja Remote untuk IndonesiaBerapa banyak lowongan remote global yang terbuka untuk Indonesia, dan gajinya (USD) per bidang.
- Laporan Kuartalan Pasar Kerja IndonesiaPHK, pendanaan, gaji & skill per kuartal — agregat terbuka.
- Laporan Pasar Remote per PeranLaporan otomatis per kelompok peran — skill, senioritas, perusahaan, gaji.
- Benchmark Gaji Remote GlobalGaji tahunan per bidang & mata uang, plus porsi lowongan terbuka untuk seluruh dunia.