Site Reliability Engineer - Sapphire Software Solutions Inc
ARCHIVED
We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.
Site Reliability Engineer at the organization.
Key technologies: Kubernetes, Prometheus, Grafana.
Key Responsibilities
Define and track SLOs, SLIs and error budgets
Design and implement observability stacks (metrics, logging, tracing)
Automate toil and improve system reliability through engineering
Conduct post-mortems and drive blameless incident retrospectives
Requirements
3+ years of relevant experience in site reliability engineer
Proficiency with monitoring tools (Prometheus, Grafana, Datadog)
Strong programming skills for automation and tooling
#J-18808-Ljbffr