JOBSEARCHER

Software Development Engineer(Distributed Systems)

WorkdayOakland, CAL6 LeadSeptember 15th, 2026
Overview Join Workday’s Data Platform and Observability Engineering team to build a next-generation, multi-petabyte scale observability platform. You will develop and scale core features for Workday’s distributed tracing stack, working with ClickHouse, Tempo, Kafka, Spark/Flink, Iceberg, and AWS. You’ll deliver high-performance backend services and contribute to Observability AI for anomaly detection and root-cause analysis. Expect a hands-on role in a collaborative, growth-focused environment that values integrity and bold ideas. Compensation / Benefitsflexible work arrangementsbase salary with range providedbonus plan or commission eligibilitystock grantsopportunity to work with cutting-edge techcomprehensive benefits (Workday offers) ResponsibilitiesDevelop and scale features for the distributed tracing platform on ClickHouse/Tempo with robust ingestion and fast query executionBuild and support high-throughput data pipelines using Kafka, Spark/Flink, and Iceberg-on-S3Tune, optimize, profile, and resolve latency/throughput issues in productionEnsure reliability with robust error handling, retries, and failover for tracing servicesApply security controls for multi-tenant data access within the platform architectureMaintain platform health via monitoring, logging, alerts, and participate in on-call rotationsSupport Observability AI by creating reliable data ingestion paths for AI-driven anomaly detection and root-cause analysisCollaborate on system design reviews, document technical details, and work with senior/ principal engineers to deepen distributed-systems expertise Key requirements5+ years in software development engineering3+ years designing, building, operating complex distributed systems with high availability5+ years programming experience in Java, Python, or GoBachelor’s degree in Computer Science, Engineering, or related field; Master’s strongly preferred or equivalent experienceStrong algorithmic skills for high-throughput data ingestion and sub-second queriesExperience with API development (gRPC, REST, OTLP) and scalable observability APIsExperience with distributed systems testing, CI/CD automation, and telemetry pipelinesHands-on experience with Kafka, Spark, Flink, or ClickHouse; multi-AZ high availability and automated failovercollaborative teamworkanalytical/problem-solving mindsetclear technical writingKafkaSparkFlink