Platform Reliability Engineer
Bantu memastikan platform Apify tetap handal dan stabil
Platform Reliability Engineer di Apify bertugas memantau sistem, mengelola insiden, dan mengatur peringatan agar tim pengembangan dapat mengirimkan produk dengan percaya diri. Anda akan bekerja dengan stack monitoring seperti Prometheus, Grafana, dan OpenTelemetry untuk memastikan sistem berjalan lancar dan efisien.
Kenapa Menarik?
Apify memiliki misi untuk membantu masyarakat dengan teknologi web scraping, seperti mencari anak hilang dan melindungi konsumen dari penipuan.
Tanggung Jawab Utama
- Mengoperasikan dan meningkatkan stack monitoring menggunakan Prometheus, Grafana, dan OpenTelemetry
- Mendefinisikan dan mengatur peringatan agar tim mendapatkan sinyal yang dapat diambil tindakan tanpa kebisingan
- Membantu mendefinisikan proses penanganan insiden dengan komunikasi yang jelas dan dokumentasi yang berguna
- Bekerja sama dengan tim platform dan produk untuk menerapkan standar kehandalan yang praktis
Persyaratan
- Pengalaman praktis dalam memilih metrik yang penting untuk dipantau di produksi
- Pengalaman dengan Prometheus, Grafana, OpenTelemetry, atau alat serupa
- Mampu membaca dan menulis kode untuk mengikuti layanan dan pipa di seluruh stack
- Memahami budaya pasca-insiden yang baik dalam praktik
Skills Wajib
Konteks Indonesia
- Overlap Jam Kerja:
- Fleksibel — atur jam kerjamu sendiri
Keywords
Lihat Deskripsi Asli dari Ashby Job Boards
Deskripsi asli dari Ashby Job Boards
Apify is the largest marketplace of tools for AI. 40,000+ Actors helping people and agents get real-time web data, track competitors, generate leads, or integrate their apps. Actors are built by a global creator community that now earns more than $1.2 million every month. Join us to help people put the web to work. Apify can find missing children https://blog.apify.com/fighting-child-traffickers-with-technology/, protect consumers from fake discounts across the EU https://blog.apify.com/how-web-scraping-ai-and-the-eu-have-come-together-to-sweep-away-fake-discounts-in-europe/, and feed data to AI chatbots https://blog.apify.com/intercom-customer-support-ai-chatbot-web-scraping/. To support our mission, we're looking for a Platform Reliability Engineer with a developer's mindset. You've shipped code and you care what happens when it runs in production (speed, failures, recovery). You'll help us strengthen how Apify monitors systems, handle incidents, and route alerts so engineering teams can ship with confidence. You won't be on-call. This role is focused on sustainable improvement, not after-hours emergency response. WHAT YOU'LL BE WORKING ON: - Monitoring & signals: Operate and improve our monitoring stack (Prometheus, Grafana, OpenTelemetry) - instrument services to expose the right metrics, define what we watch in production, and shape alerting so teams get actionable signals without the noise. - When things go wrong: Help define how we run incidents - clear communication, structured learning afterward, and supporting artifacts (status page, runbooks). - With the team: Work with platform and product engineers to make reliability standards practical - help teams adopt better tooling or practices when things change, and write documentation people actually use. WHO WE'RE LOOKING FOR: - You have hands-on experience choosing what to measure in production - not just reading dashboards, but picking signals that reflect the customer experience. - You're comfortable with incidents and alerts, from early detection through resolution and follow-up so similar issues are less likely to recur. - You have hands-on experience with Prometheus, Grafana, OpenTelemetry, or similar, and with alert-routing tools such as PagerDuty. - You read and write code: you can follow services and pipelines across the stack and collaborate on technical details with the teams building them. - You know what good post-incident culture looks like in practice - blame-free, learning-focused, and actually used to make things better - even if your past title never mentioned reliability. - You can write clear, concise guidance that teams adopt, and you work constructively toward sound decisions. - You're driven to automate repetitive tasks and improve developer workflows. Nice to have: - Meaningful hands-on experience as an application or backend developer - you've built things that run in production and approach observability as someone who needs it as a "user," not just the person who sets it up. - Experience building and maintaining infrastructure on AWS (EC2, EKS, S3, CloudFormation, or similar), and hands-on experience with container technologies. - Some familiarity with CI/CD pipelines or release practices - enough to have an informed opinion on what makes deployments reliable and safe. Don't worry if you don't meet all of the above criteria. We value diverse skills and experience and would love to hear from you. Our tech stack - Infra: AWS Compute (Kubernetes (EKS), EC2, Lambda), Helm, ArgoCD, MongoDB, Redis, DynamoDB, S3, GitHub Actions - Monitoring: Grafana, Prometheus, OpenTelemetry, Mezmo, PagerDuty - Frontend: React.js, styled-components, Storybook, Chromatic, Cypress, Playwright - Backend: TypeScript/Node.js, Nest.js, Next.js, Express.js, Docusaurus, Vitest - Tools: GitHub, Notion, Google Workspace - Editor and AI assistant of your choice (GH Copilot, Cursor, Claude, Gemini, or JetBrains AI) - Process: two-week sprints, code reviews, tests, automating whatever we can, and deploying multiple times per day. BY THE END OF THE FIRST 3 MONTHS, WE EXPECT YOU TO: - Have completed the general onboarding process. - Have built working relationships with platform engineers, engineering leads, and others involved in production response, and aligned on how you'll collaborate. - Understand, in principle, how the Apify platform works, and be able to handle smaller problems, incidents, or bugs on the infrastructure you work with most. - Have mapped how we handle monitoring, incidents, and alerts today - where the friction is and where a focused improvement would help. - Have published initial monitoring, observability, and alerting guidelines - covering signals, naming, key dashboards, and alerting principles (severity, routing, and noise reduction) - aligned with existing tooling. - Be participating in incident reviews and translating patterns into improved playbooks. - Be contributing actively in team ceremonies (planning, grooming) and technical discussions, and in touch with other teams to support their infrastructure needs. BY THE END OF THE FIRST 6 MONTHS, WE EXPECT YOU TO: - Be working on bigger tasks mostly independently (while staying fearless about asking for help). - Have built a network across engineering, stay in touch with other teams on infrastructure initiatives, and gather feedback to find ways to help them in their daily work. - Have teams referencing your guidance when planning higher-risk changes, with measurably less alert noise and duplicate paging. - Have incident documentation (communication, roles, lessons learned) that's easy to find and actually used during real incidents. - Own the monitoring and alerting improvement roadmap end-to-end. - Have agreed with leadership on priorities for monitoring and alerting - tooling, training, and the metrics that actually matter. WHY SHOULD YOU WORK AT APIFY? - Space, support, and autonomy for personal growth, with a direct impact on Apify's success - Full-time position in Prague (Lucerna Palace) or Brno (Titanium) 🏰 - Option to work remotely 🛋️ - Flexible working hours (perfect for both night owls 🦉 and early birds 🐥) - Nobody counts holidays as long as the work gets done 💪 - Unlimited Claude for every Apifier. We don't count tokens. Just use them well 🤖 - Stock options and profit sharing 💰 - We welcome pets, kids, and bikes at the office 🐕👨👧 - Epic team buildings and offsites 🚢 with biking, canoeing, and other adventures 🪂 - Solid education and training budget, conference tickets, internal "Eat & Learn" sessions, and the possibility to work across teams 👩🏼💻👨🏽💻 - Generous hardware budget 💻 - Free lunches every day when you're in the office 🌮🍱🍜🍕🥡 - Unlimited supply of ☕ & 🍺 and snacks - Free entry to the wonderful Prague Zoo 🐘 - Free Multisport card 🏋 - Ping-pong, chess, PS5, lightsabers, foosball league after lunch. For more details about Apify and what it's like to work with us, see our Careers page https://apify.com/jobs.
Akun gratis · tanpa kartu kredit · Masuk
Pro Rp39rb/bln · lamar tanpa batas + resume AI
Pertanyaan yang sering diajukan
- Apakah Platform Reliability Engineer di Apify bisa dikerjakan remote?
- Posisi ini berlokasi di Remote. Detail remote/onsite ada di deskripsi lowongan.
- Jenis pekerjaan apa Platform Reliability Engineer di Apify?
- Posisi ini adalah pekerjaan full time.
- Bagaimana cara melamar?
- Klik tombol "Lamar" pada halaman ini untuk menuju halaman aplikasi resmi Apify.
Jelajahi lebih lanjut
Data & laporan pasar
Riset gaji & permintaan skill dari data lowongan kami sendiri.
- Lowongan IT Indonesia vs Remote Global (2026)Analisis data primer 2.049 lowongan: metodologi, klasifikasi, dataset bisa diunduh.
- Permintaan Skill AI: Indonesia vs Global (2026)10.000+ lowongan, classifier taxonomy-first, Wilson CI, pra-registrasi sebelum analisis.
- Laporan Hiring Indonesia: Tech vs Non-TechPermintaan lowongan per bidang dari hitungan agregat — bukan listing per-listing.
- Benchmark Gaji IndonesiaKisaran gaji agregat lintas peran, dengan metodologi dan dataset terbuka.
- Laporan Kuartalan Pasar Kerja IndonesiaPHK, pendanaan, gaji & skill per kuartal — agregat terbuka.
- Laporan Pasar Remote per PeranLaporan otomatis per kelompok peran — skill, senioritas, perusahaan, gaji.
- Benchmark Gaji Remote GlobalGaji tahunan per bidang & mata uang, plus porsi lowongan terbuka untuk seluruh dunia.
Dari blog kami
- Tren Lowongan Juli 2026: Security Engineer 🚀Analisis tren lowongan remote Juli 2026: permintaan Security Engineer meningkat. Gaji & skill yang dicari? Cek di sini!
- Gaji Senior Tech Role Remote 2026Analisis mendalam range gaji 7 senior tech role remote dari Vercel, Airbnb, Stripe hingga Notion. Bandingkan dengan pasar lokal dan strategi negosiasi.
- Playlist PHK Masuk QueueAnalisis gelombang PHK teknologi Juni 2026: sektor mana yang masih hiring, strategi bertahan, dan 7 lowongan remote yang tetap terbuka.
Akun gratis · tanpa kartu kredit · Masuk
Pro Rp39rb/bln · lamar tanpa batas + resume AI
