{"schemaVersion":"jobsearcher.job.v1","id":"548427dd7e537d1e7e9108a7","url":"https://jobsearcher.com/jobs/548427dd7e537d1e7e9108a7","canonicalUrl":"https://jobsearcher.com/jobs/548427dd7e537d1e7e9108a7","title":"Staff Observability Platform Engineer","description":"🚀 Staff Observability Platform Engineer | AI/GPU Infrastructure | Remote/Hybrid📍 Hybrid — Seattle, Houston, or New YorkHave you personally operated an observability backend at real production scale — not just consumed dashboards built by another team? We're looking for someone who can quantify that scale and explain the engineering decisions behind it.What you'll do: 🔹 Design, build, and operate large-scale metrics, logging, and tracing platforms 🔹 Own observability backend architecture in distributed Kubernetes and AI/GPU environments 🔹 Operate and scale Mimir, Thanos, VictoriaMetrics, Cortex, Loki, or Elasticsearch at production scale 🔹 Design Prometheus-based architectures (remote write, high availability, retention, global querying) 🔹 Build OpenTelemetry Collector pipelines (receivers, processors, exporters, sampling) 🔹 Identify and remediate high-cardinality metrics 🔹 Establish observability standards across engineering teamsWhat we're looking for: ✅ Personal ownership (not just consumption) of a metrics/logs backend at meaningful scale ✅ Ability to quantify that scale: active series, ingestion rate, TB/day, retention, cluster size ✅ Real experience making cardinality, retention, ingestion, and storage-cost trade-offs ✅ Strong hands-on production Kubernetes experience (multi-cluster, networking, autoscaling) ✅ Strong production code-reading/review ability (Go and/or Python) ✅ Understanding of retries, timeouts, backpressure, circuit breaking, and failure handlingNice to have: 🌟 Large-scale GPU fleets, NVIDIA DCGM, InfiniBand/RoCE/NVLink, Slurm/HPC 🌟 Custom Prometheus exporters, custom OpenTelemetry components 🌟 AI/ML infrastructure or distributed training platforms📍 Hybrid — Seattle, Houston, or New York","company":"Programmingcom","rawCompany":"programmingcom","city":"Seattle","state":"WA","isRemote":false,"isActive":false,"createdAt":"2026-08-27T09:55:52.323Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Staff Observability Platform Engineer","description":"🚀 Staff Observability Platform Engineer | AI/GPU Infrastructure | Remote/Hybrid📍 Hybrid — Seattle, Houston, or New YorkHave you personally operated an observability backend at real production scale — not just consumed dashboards built by another team? We're looking for someone who can quantify that scale and explain the engineering decisions behind it.What you'll do: 🔹 Design, build, and operate large-scale metrics, logging, and tracing platforms 🔹 Own observability backend architecture in distributed Kubernetes and AI/GPU environments 🔹 Operate and scale Mimir, Thanos, VictoriaMetrics, Cortex, Loki, or Elasticsearch at production scale 🔹 Design Prometheus-based architectures (remote write, high availability, retention, global querying) 🔹 Build OpenTelemetry Collector pipelines (receivers, processors, exporters, sampling) 🔹 Identify and remediate high-cardinality metrics 🔹 Establish observability standards across engineering teamsWhat we're looking for: ✅ Personal ownership (not just consumption) of a metrics/logs backend at meaningful scale ✅ Ability to quantify that scale: active series, ingestion rate, TB/day, retention, cluster size ✅ Real experience making cardinality, retention, ingestion, and storage-cost trade-offs ✅ Strong hands-on production Kubernetes experience (multi-cluster, networking, autoscaling) ✅ Strong production code-reading/review ability (Go and/or Python) ✅ Understanding of retries, timeouts, backpressure, circuit breaking, and failure handlingNice to have: 🌟 Large-scale GPU fleets, NVIDIA DCGM, InfiniBand/RoCE/NVLink, Slurm/HPC 🌟 Custom Prometheus exporters, custom OpenTelemetry components 🌟 AI/ML infrastructure or distributed training platforms📍 Hybrid — Seattle, Houston, or New York","datePosted":"2026-08-27T09:55:52.323Z","dateModified":"2026-08-27T09:55:52.323Z","hiringOrganization":{"@type":"Organization","name":"Programmingcom","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Seattle","addressRegion":"WA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"548427dd7e537d1e7e9108a7"},"url":"https://jobsearcher.com/jobs/548427dd7e537d1e7e9108a7"}}