JOBSEARCHER

Sr. Observability Engineer

MxProvo, UTL6 LeadSeptember 15th, 2026
Overview As Senior Observability Engineer, you will build and operate MX's observability control plane, establishing baselines, signals, and automation to scale across services. You will partner with product engineering to ensure meaningful visibility and reduce noise, acting as Incident Commander when needed. Your work will drive measurable improvements in reliability and incident response, aligning with MX’s mission to empower financial experiences. This role offers hands-on impact in a fast-moving, high-trust environment and a chance to shape how we observe and operate at scale. Compensation / Benefitsonsite mealsgymmother’s loungemeditation roomteam collaboration and office perks ResponsibilitiesBuild and operate an observability control plane using Datadog API and Terraform to automate monitors, dashboards, and tagging.Create detection and dashboard gap packs after incidents, grounded in Datadog patterns and MX investigations.Define and audit 'good' standards for services across Ruby, Go, and Java; assess readiness and enforce contracts.Validate ownership and completeness of alerts/dashboards with service owners; escalate when coverage fails.Own monthly observability and service-catalog health reporting including gaps and coverage trends.Run maturity assessments (baseline through SLO, launch-ready, self-serve) and track progress.Tune alerting to minimize false Sev1/2, improve Sev3/4 actionability, and guide cost/cardinality practices.Develop self-serve onboarding to establish baseline observability for new services on day one.Share on-call pager responsibilities; act as Incident Commander or support during incidents as needed.Close detection loop after incidents with gap packs and new monitors/dashboards to reduce pager load.Lead high-value launch and production-readiness reviews as checkpoints, not ongoing staffing. Key requirements5+ years running production observability, SRE, or DevOps5+ years automation-first engineering in Python, Bash, Go, and/or Terraform; Kubernetes proficiencyDistributed-systems debugging across microservices with Kubernetes, NATS, RabbitMQ, Postgres, RedisExperience with incident management and on-call/or Incident Commander capabilitiesAI-aware/workflow-literate for scalable reviews, audits, and docscross-functional influence without authorityproblem-solving under pressurecommunication and governance reportingDatadogTerraformKubernetes