Site Reliability Engineer - Sapphire Software Solutions Inc
Site Reliability Engineer at the organization.
Key technologies: Kubernetes, Prometheus, Grafana.
Key Responsibilities
Define and track SLOs, SLIs and error budgets
Design and implement observability stacks (metrics, logging, tracing)
Automate toil and improve system reliability through engineering
Conduct post-mortems and drive blameless incident retrospectives
Requirements
3+ years of relevant experience in site reliability engineer
Proficiency with monitoring tools (Prometheus, Grafana, Datadog)
Strong programming skills for automation and tooling
#J-18808-Ljbffr