Staff Observability Platform Engineer
π Staff Observability Platform Engineer | AI/GPU Infrastructure | Remote/Hybridπ Hybrid β Seattle, Houston, or New YorkHave you personally operated an observability backend at real production scale β not just consumed dashboards built by another team? We're looking for someone who can quantify that scale and explain the engineering decisions behind it.What you'll do: πΉ Design, build, and operate large-scale metrics, logging, and tracing platforms πΉ Own observability backend architecture in distributed Kubernetes and AI/GPU environments πΉ Operate and scale Mimir, Thanos, VictoriaMetrics, Cortex, Loki, or Elasticsearch at production scale πΉ Design Prometheus-based architectures (remote write, high availability, retention, global querying) πΉ Build OpenTelemetry Collector pipelines (receivers, processors, exporters, sampling) πΉ Identify and remediate high-cardinality metrics πΉ Establish observability standards across engineering teamsWhat we're looking for: β
Personal ownership (not just consumption) of a metrics/logs backend at meaningful scale β
Ability to quantify that scale: active series, ingestion rate, TB/day, retention, cluster size β
Real experience making cardinality, retention, ingestion, and storage-cost trade-offs β
Strong hands-on production Kubernetes experience (multi-cluster, networking, autoscaling) β
Strong production code-reading/review ability (Go and/or Python) β
Understanding of retries, timeouts, backpressure, circuit breaking, and failure handlingNice to have: π Large-scale GPU fleets, NVIDIA DCGM, InfiniBand/RoCE/NVLink, Slurm/HPC π Custom Prometheus exporters, custom OpenTelemetry components π AI/ML infrastructure or distributed training platformsπ Hybrid β Seattle, Houston, or New York