JOBSEARCHER

Engineering Manager, Observability

Overview Join CoreWeave as Manager, Observability Engineering to lead a team building and operating observability across metrics, logs, traces, and telemetry pipelines. You will shape strategy and roadmap, improve platform reliability and performance, and guide architecture for scalable observability. Partner with infrastructure, platform, security, and application teams to improve instrumentation and visibility in production AI infrastructure. This role combines technical leadership with people management to scale observability as the business grows. You will help deliver reliable, high-velocity engineering experiences in a fast-growing cloud environment. Compensation / BenefitsMedical, dental, and vision insurance401(k) with generous employer matchFlexible PTOTuition ReimbursementEmployee Stock Purchase Program (ESPP)Mental Wellness Benefits ResponsibilitiesLead the team responsible for observability across metrics, logs, traces, and telemetry pipelinesDefine strategy and roadmap for observability platformsDrive reliability and performance improvements of observability systemsGuide architectural decisions for observability infrastructureCollaborate with infrastructure, platform, security, and application engineering teams to improve instrumentation and production visibilityEnsure observability platforms scale with business and customer needsOversee operational ownership of observability tooling and pipelinesFoster engineering excellence and cross-functional adoption of observability practices Key requirements5+ years of software engineering experience with production systems at scale2+ years of engineering management experienceExperience building and operating observability platforms across logs, metrics, traces, or alerting in distributed systemsKnowledge of reliability engineering concepts including SLOs, SLIs, incident management, error budgets, and fault-tolerant designExperience scaling telemetry systems including storage backends, and query layersExperience with distributed systems, performance engineering, and trade-offs involving scale, resilience, and costExperience partnering with infrastructure, security, and application engineering teams to drive platform adoptionExperience hiring and managing engineering teamsProblem solving across complex systemsCollaborative cross-functional workCuriosity about improving observability and debugging workflowsOpenTelemetryGrafanaPrometheus-compatible systems