JOBSEARCHER

Principal Distributed Systems Engineer - Observability

WorkdayOakland, CAL7 ManagerSeptember 15th, 2026
Overview In this role, you lead the technical vision for Workday’s distributed tracing and Observability Platform, built on ClickHouse and Grafana Tempo, with a multi-petabyte data backbone. You design end-to-end tracing infrastructure, guide cross-team architecture, and drive AI-enabled incident management. You’ll mentor engineers, advocate best practices, and shape the future of Observability AI while operating with high autonomy in a fast-paced environment. Compensation / Benefitsbase salary plus bonus planstock grantsflexible work arrangementsremote/hybrid optionscareer growth and skills developmentequal opportunity employer ResponsibilitiesArchitect and build the distributed tracing platform for multi-petabyte scale with low-latency queriesOwn the big-data pipeline feeding tracing data (Kafka, Spark/Flink, Iceberg-on-S3) including schema, partitioning, and lifecycleOptimize storage and query performance (Parquet/Iceberg, compression, partitioning, indexing)Lead HA/DR design across regions and zones with defined RTO/RPO for a tier-1 platformDesign security architecture for multi-tenant data access across ingestion and query layersEnsure operational excellence (monitoring, logging, alerting, capacity planning, on-call)Evaluate and adopt new cloud-native/open-source technologies to improve scalability and costShape Observability AI by enabling automated anomaly detection and AI-assisted incident managementPublish best practices, mentor teams, and act as a technical thought leader for observability/data stackOperate with high autonomy, aligning technical direction with platform strategy Key requirements14+ years in software development engineering6+ years designing, building, and operating complex distributed systems with high availability (e.g., 99.9% uptime)8+ years in at least two of: Java, Python, Go with production distributed systems experienceBachelor’s degree in Computer Science, Engineering, or related field; Master’s strongly preferred or equivalent experienceleadership and mentorshipcross-team collaborationtechnical writing and documentationDistributed systems architectureAPI development and advanced API patternsLarge-scale data processing (Kafka, Spark/Flink)