Langsung ke konten utama
Kembali ke Lowongan

Senior Site Reliability Engineer - India

Tingkatkan kestabilan platform produksi Syfe sebagai Senior Site Reliability

Anda akan bertanggung jawab atas kestabilan platform produksi Syfe secara end-to-end. Anda akan mendefinisikan target kestabilan, membangun sistem dan otomatisasi, serta memimpin program on-call dan tanggapan insiden. Kerja ini melibatkan pengembangan SLOs, observabilitas, dan otomatisasi di platform Kubernetes multi-region.

Kenapa Menarik?

Dapat mengelola ketersediaan dan kinerja platform digital wealth management di tiga wilayah (Singapore, Hong Kong, Sydney) menggunakan Kubernetes dan AWS.

Tanggung Jawab Utama

  • Mendefinisikan dan mendorong SLIs/SLOs serta anggaran kesalahan di layanan kritis
  • Memimpin program on-call dan tanggapan insiden, termasuk postmortem tanpa penyalahan
  • Mengelola kestabilan platform EKS, GitOps, Helm, dan IaC di AWS
  • Membangun dan mematangkan stack observabilitas dengan Datadog, Grafana, dan ClickHouse
  • Melakukan perencanaan kapasitas, analisis mode kegagalan, dan latihan game-day

Persyaratan

  • Pengalaman dalam mengelola platform Kubernetes multi-region
  • Kemampuan dalam mengembangkan SLOs dan observabilitas
  • Pengalaman dalam otomatisasi dan pengembangan IaC
  • Kemampuan dalam mengelola program on-call dan tanggapan insiden
  • Pengalaman dalam perencanaan kapasitas dan analisis mode kegagalan

Skills Wajib

kubernetesawsterraformdatadogsloobservability

Konteks Indonesia

Overlap Jam Kerja:
Fleksibel — atur jam kerjamu sendiri
Lihat selisih gaji remote (USD) vs lokal →

Keywords

site-reliability-engineerkubernetesmulti-regionsloobservabilityawsterraformdatadogfull-timeremote
Lihat Deskripsi Asli dari Manatal Career Pages

Deskripsi asli dari Manatal Career Pages

Syfe is APAC's largest and fastest-growing digital wealth platform , trusted with over US$10 billion in assets. We are fundamentally changing how hundreds of thousands of people across Asia-Pacific build wealth through a holistic approach to managing money rather than just pushing investment products. Backed by world-class investors and recognised as a leader in wealthtech, we are a team of passionate builders creating the future of wealth management. About the Role : We are looking for a Senior Site Reliability Engineer to own the reliability of Syfe's production platform end-to-end. Syfe runs a Kubernetes-native, multi-region platform (Singapore, Hong Kong, Sydney) serving a regulated digital wealth- management product, where availability, latency, and trust are first-order product features. This is a senior individual-contributor role, not a managerial one. You will define what "reliable" means in measurable terms, build the systems and automation that keep us there, and own and lead our on-call and incident-response program. You'll spend your time engineering reliability into the platform — through SLOs, observability, and automation — rather than firefighting, and you'll raise the bar for how the whole engineering org operates production What You'll Own Reliability targets. Define and drive SLIs/SLOs and error budgets across critical services; partner with product and engineering teams to make error-budget-based decisions that balance velocity and stability. On-call & incident program. Own the on-call rotation, escalation policies, and paging strategy. Establish incident command, run blameless postmortems, and turn RCAs into tracked, completed reliability work. Drive down MTTD and MTTR. Platform & Kubernetes reliability. Own the reliability of our EKS-based deployment platform — GitOps delivery (ArgoCD), Helm-based release configuration, and Infrastructure as Code (Terraform/OpenTofu) on AWS. Make deployments safe, progressive, and reversible. Observability. Build and mature the observability stack (Datadog for production APM/RUM; Grafana, VictoriaMetrics, and ClickHouse for metrics and logs). Make systems debuggable: meaningful dashboards, actionable alerts, and low alert noise. Resilience. Lead capacity planning, scalability, failure-mode analysis, disaster-recovery and business-continuity planning, and game-day / chaos exercises across regions. Toil reduction. Identify operational toil and eliminate it with automation and self-service tooling, so reliability scales with the platform rather than with headcount. Production safety. Strengthen rollout/rollback paths, deployment guardrails, and secrets handling (HashiCorp Vault), and partner with engineering teams to harden services before they reach production. Engineering influence. Lead by example with hands-on engineering — design reviews, production-readiness reviews, runbooks, documentation, and mentoring — embedding SRE practices across the org Qualifications (Must-have): Bachelor's or Master's degree in Computer Science, Engineering, or a related field, or equivalent practical experience. 4–8 years in SRE, platform engineering, or DevOps, with a strong senior IC track record of owning production systems. Production-grade expertise with Kubernetes and containers, and a cloud platform (AWS preferred) in a distributed- systems environment. Hands-on experience defining and operating against SLIs/SLOs and error budgets, and leading incident response and blameless postmortems. Strong with observability tooling — metrics, logging, tracing, dashboards, and alerting (e.g., Datadog, Prometheus/Grafana, or equivalents). Proficient writing automation and infrastructure code (e.g., Python, Go, Shell, Terraform) and comfortable with GitOps / CI/CD delivery. Solid grasp of Linux/Unix internals, networking, and cloud-native security fundamentals. Strong operational rigor, ownership, and clear written/verbal communication, including during high-pressure incidents. Nice to Have: Experience operating in regulated or high-trust environments (fintech, payments, etc.). Experience running multi-region / multi-cluster Kubernetes at scale. Familiarity with ArgoCD, Helm/Helmfile, OpenTofu, Vault, Cloudflare, or comparable tooling. Experience building self-service developer platforms or internal reliability tooling. Contributions to open-source projects or public technical content (GitHub, blogs, talks). Relevant cloud or Kubernetes certifications (e.g., AWS, CKA/CKS) Come As You Are We believe in the power of diversity and are dedicated to creating a welcoming and innovative environment for all our employees. So we embrace and encourage applications from candidates of all backgrounds and provide equal employment opportunities for all. Due to the volume of applications, we regret that only shortlisted candidates will be notified.

Konteks Gaji

Posisi Engineering serupa di LokerDollar dibayar sekitar $195k/yr (kisaran $36k–300k/yr, dari 258 listing aktif).

Perekrutan di Syfe

Syfe punya 21 lowongan aktif lain di LokerDollar dan telah merekrut di sini sejak 10 Agu 2026 — di kategori Engineering, Marketing, Product.

Lihat semua lowongan Syfe →
ℹ️ Kami belum bisa memverifikasi tautan lamaran ini. Kamu tetap bisa coba melamar — cek lagi nanti kalau gagal terbuka.
Terbuka untuk Indonesia
Perusahaan
Syfe
Sumber
Manatal Career Pages
Tipe Pekerjaan
full time
Lokasi
Remote
Kategori
Level
senior
DipostingNew
10 Agu 2026

Bagikan lowongan ini

Bantu temanmu nemu kerja remote berikutnya.

Pertanyaan yang sering diajukan

Apakah Senior Site Reliability Engineer - India di Syfe bisa dikerjakan remote?
Posisi ini berlokasi di Remote. Detail remote/onsite ada di deskripsi lowongan.
Jenis pekerjaan apa Senior Site Reliability Engineer - India di Syfe?
Posisi ini adalah pekerjaan full time.
Bagaimana cara melamar?
Klik tombol "Lamar" pada halaman ini untuk menuju halaman aplikasi resmi Syfe.

Jelajahi lebih lanjut

Data & laporan pasar

Riset gaji & permintaan skill dari data lowongan kami sendiri.

Dari blog kami