Skip to main content
Back to Jobs

G13 - Operations Support Engineer

Design service observability usage model ensuring metrics, logs, traces flow into Elastic Cloud

Design and own the service observability usage model by ensuring all service metrics, logs, and traces flow into Elastic Cloud as the authoritative source. Maintain dashboards and SLOs, evaluate supplemental tooling like CloudWatch and AWS Managed Prometheus/Grafana, and build proactive, noise-reduced alerting and incident response playbooks. Drive post-incident RCA and remediation tracking with closure SLAs, optimize service performance throu...

Why This Role?

Mentor engineers on production readiness and observability patterns while evolving on-call playbooks

Key Responsibilities

  • Design and maintain service observability in Elastic Cloud ensuring metrics, logs, and traces are properly ingested
  • Build noise-reduced alerting and incident response playbooks, drive RCA and remediation tracking
  • Optimize service performance via profiling, caching, autoscaling heuristics, and concurrency tuning
  • Implement secure supply chain and runtime controls including image scanning, SBOM, secrets management, and TLS/mTLS
  • Curate operational runbooks, golden dashboards, and reliability/production readiness checklists
  • Integrate model/guardrail service telemetry into unified Elastic Cloud views and support compliance evidence collection

Requirements

  • 4+ years of experience in SRE, Production Ops, Platform, or Reliability for SaaS or high-throughput services
  • Working knowledge of AWS and Kubernetes for deployment, troubleshooting, and networking concepts
  • Familiarity with Infrastructure as Code and GitOps tools like Terraform and Argo for consuming modules and reviewing changes
  • Observability implementation and usage with Elastic Cloud, including metrics, logs, traces, and Prometheus/OpenTelemetry concepts
  • Proven on-call and incident management experience with triage, MTTR reduction, and RCA authorship
  • Scripting and automation skills in Python, Bash, or Go for ops tooling

Required Skills

awskuberneteselastic-cloudobservabilitypythonbashgoincident responseperformance optimizationsecurity compliancescripting
View Excerpt — Full Description at Manatal Career PagesShow more

Read the full description at Manatal Career Pages

Responsibilities: Design & own service observability usage model: ensure all service metrics, logs, traces flow into Elastic Cloud (authoritative); maintain dashboards & SLOs; evaluate pragmatic use of CloudWatch, AWS Managed Prometheus / Grafana for supplemental or fallback views.…

Openness not stated by employer — check the listing

Salary
Job Type
full time
Location
On-site
Category
Seniority
senior
PostedRecheck the source
Aug 10, 2026

Share this job

Help a friend find their next remote role.

Explore related

Market data & reports

Salary & skill-demand research built from our own listings data.