JOBSEARCHER

Staff Observability Platform Engineer

πŸš€ Staff Observability Platform Engineer | AI/GPU Infrastructure | Remote/HybridπŸ“ Hybrid β€” Seattle, Houston, or New YorkHave you personally operated an observability backend at real production scale β€” not just consumed dashboards built by another team? We're looking for someone who can quantify that scale and explain the engineering decisions behind it.What you'll do: πŸ”Ή Design, build, and operate large-scale metrics, logging, and tracing platforms πŸ”Ή Own observability backend architecture in distributed Kubernetes and AI/GPU environments πŸ”Ή Operate and scale Mimir, Thanos, VictoriaMetrics, Cortex, Loki, or Elasticsearch at production scale πŸ”Ή Design Prometheus-based architectures (remote write, high availability, retention, global querying) πŸ”Ή Build OpenTelemetry Collector pipelines (receivers, processors, exporters, sampling) πŸ”Ή Identify and remediate high-cardinality metrics πŸ”Ή Establish observability standards across engineering teamsWhat we're looking for: βœ… Personal ownership (not just consumption) of a metrics/logs backend at meaningful scale βœ… Ability to quantify that scale: active series, ingestion rate, TB/day, retention, cluster size βœ… Real experience making cardinality, retention, ingestion, and storage-cost trade-offs βœ… Strong hands-on production Kubernetes experience (multi-cluster, networking, autoscaling) βœ… Strong production code-reading/review ability (Go and/or Python) βœ… Understanding of retries, timeouts, backpressure, circuit breaking, and failure handlingNice to have: 🌟 Large-scale GPU fleets, NVIDIA DCGM, InfiniBand/RoCE/NVLink, Slurm/HPC 🌟 Custom Prometheus exporters, custom OpenTelemetry components 🌟 AI/ML infrastructure or distributed training platformsπŸ“ Hybrid β€” Seattle, Houston, or New York