JOBSEARCHER

Observability Engineer

AsobbiDenver, COL6 LeadSeptember 10th, 2026
Senior Observability Engineer, Cloud PlatformRemote (US)To $185k–$210k base + bonus + RSUsOur client is a NASDAQ-listed GPU cloud provider running dense accelerated-compute clusters on bare metal and Kubernetes, across Ethernet and InfiniBand fabrics. They're building a clean-slate, Kubernetes-native GPU cloud and scaling it across new data centre sites.They're hiring a Senior Observability Engineer to build the observability layer for that platform from scratch. This role makes "we detect it before the customer does" the normal case. You'll own the metrics, logs, alerting, dashboards and detection that turn raw signals from multi-vendor GPUs (AMD/NVIDIA), hosts, fabrics and switches into problems caught early, both operator-facing and tenant-facing, and set the standards the rest of the platform instruments by. You'll work alongside SRE, platform, network and security, who own the control plane, fabric and day-to-day ops.The main focus of the role: OpenTelemetry as the platform-wide instrumentation foundation; the Prometheus-based metrics pipeline across GPU/CPU/network/thermal/power including the DCGM profiling suite; proactive detection (GPU failure-rate tracking, Xid, anomaly signals feeding an automated detect-drain-remediate loop); tenant-scoped observability with per-tenant isolation; and SLO/reliability instrumentation (99.9%+ node SLA, MTTA/MTTR).You'll need:Significant experience building production observability at scale, with real platform-build experience (not just running dashboards someone else built).Deep Prometheus/Grafana (PromQL, exporters, alerting rules, Alertmanager) and hands-on OpenTelemetry as a foundational layer (Collector, receivers/exporters, semantic conventions).Time-series at scale, including high-cardinality/long-term storage (Thanos, Mimir, Cortex or VictoriaMetrics).Low-noise, SLO/error-budget-based alerting, strong Python or Go, and solid Linux systems fundamentals.Strongly preferred:Multi-vendor GPU/HPC observability (DCGM & dcgm-exporter, NVML, Xid; AMD Device Metrics Exporter, ROCm, amd-smi; GPU node health checks), network/fabric telemetry (gNMI, RoCE, InfiniBand/UFM), Kubernetes observability, Checkmk, and log/trace pipelines (Loki, Elastic, Vector).If this role is of interest, then submit your resume and get in touch!