JOBSEARCHER

Data Platform Lead

Location: Santa Clara, CAFunction: Data Platforms / Advanced AnalyticsRole Type: Technical Lead / Solution LeadPrimary Platform: Databricks Lakehouse on CloudScope: EPIC Data Platform (Dedicated Tenant and Multi-Tenant Capabilities)Role PurposeThe EPIC Data Platform Lead owns the technical direction and implementation leadership for a secure, governed, and scalable cloud data platform built on Databricks. This role combines hands-on data engineering and analytical problem-solving with enterprise architecture leadership across tenant isolation, data governance, identity and access, encryption, observability, production readiness, and platform operations.The lead will translate business, security, and engineering requirements into implementable platform capabilities for internal, customer-dedicated, and controlled multi-tenant use cases.Key Expected OutcomesTrusted Data Products: Curated, traceable, and analytics-ready data with clear ownership, quality controls, and validation. Secure Tenant Boundaries: Validated isolation across workspace, catalog, storage, identity, network, compute, jobs, APIs, and data exports. Production-Grade Operations: Observable, supportable, and cost-aware services supported by automated deployment, runbooks, and evidence-based security controls. Key ResponsibilitiesPlatform Architecture & Technical Leadership: Define target architecture, engineering standards, roadmaps, Architecture Decision Records (ADRs), reusable patterns, and non-functional requirements for EPIC Databricks environments. Conduct design reviews and manage trade-offs across performance, security, operability, scalability, and cost. Databricks Implementation: Lead the setup and optimization of workspaces, Unity Catalog, Delta Lake, data pipelines, Databricks Workflows, SQL Warehouses, compute policies, external locations, storage credentials, and deployment topologies. Establish maintainable Medallion Architecture (Bronze/Silver/Gold) and production practices. Data Engineering & Analysis: Design batch, streaming, and event-driven ingestion, source-to-target transformations, reconciliation, data profiling, exploratory analysis, and root-cause analysis. Validate data latency, completeness, entity linking, accuracy, and business logic outcomes. Security by Design: Partner with cybersecurity, IAM, network, and cloud teams to enforce least privilege, SSO/federation, service principal authentication, secrets management, private connectivity (PrivateLink), controlled egress, hardening, and vulnerability remediation. Governance & Data Protection: Implement data classification, taxonomies, metadata management, automated lineage, retention rules, fine-grained access controls, row filters, column masks, controlled sharing, DLP-aligned controls, and evidence-driven compliance. Encryption & Key Management: Design and implement encryption in transit and at rest, Customer-Managed Keys (CMK) / Bring Your Own Key (BYOK) patterns, cloud KMS/HSM integration, key separation, rotation, revocation, monitoring, recovery, and control validation. Dedicated & Multi-Tenant Delivery: Define tenant onboarding, registry, provisioning, configuration, isolation, metadata-driven routing, metering/showback, offboarding, and migration. Prevent unauthorized cross-tenant access and validate isolation via automated negative testing. Observability & Operations: Implement end-to-end logging, auditability, data-quality monitoring, health dashboards, alerting, SIEM integration, incident response runbooks, service-level measures, capacity planning, and cost management. Delivery Leadership: Own backlog refinement, milestone delivery, risk mitigation, release readiness, production cutover, operational handoff, and cross-functional alignment. Mentor engineers and drive execution across data, cloud, security, QA, and business teams. Required QualificationsEducation & Experience: Bachelor’s degree in Computer Science, Engineering, Information Systems, Data Science, or a related field (or equivalent practical experience) with substantial experience leading enterprise cloud data platform implementations. Databricks Depth: Extensive hands-on technical proficiency with Databricks, Apache Spark, Delta Lake, Databricks Workflows, Unity Catalog, SQL, and Python (Scala is a plus). Data Engineering & Analytics: Proven expertise in complex data analysis, profiling, reconciliation, debugging, performance tuning, and root-cause analysis on large-scale datasets. Pipeline Design: Direct experience building production-grade batch, streaming, micro-batch, event, and file-based ingestion pipelines incorporating schema evolution, replayability, backfills, and idempotent processing. Cloud Infrastructure: Strong hands-on knowledge of cloud-native data services, object storage, IAM, private networking, key management, logging/monitoring, IaC, and CI/CD pipelines (AWS preferred; Azure or GCP relevant). Security & Governance: Deep understanding of least privilege, identity federation, service principals, data classification, lineage, retention, masking, audit logging, DLP, controlled sharing, and handling sensitive or regulated data. Multi-Tenancy: Demonstrated experience designing or operating dedicated-tenant, multi-tenant, or customer-isolated platforms (tenant lifecycle, logical/physical isolation, resource governance, and cross-tenant testing). Cryptographic Controls: Hands-on experience implementing encryption at rest/transit, cloud KMS/HSM integration, CMK/BYOK, key rotation, separation of duties, and audit evidence generation. Leadership & Communication: Strong architecture, technical writing, stakeholder management, and team leadership skills to convert ambiguous requirements into executable engineering blueprints. Preferred QualificationsAWS & Databricks Integration: In-depth experience with Databricks on AWS (S3, KMS, PrivateLink, VPC Endpoints, IAM roles, CloudTrail, CloudWatch). Streaming & Event Edge: Experience with Apache Kafka, streaming integration, or edge telemetry platforms. DevSecOps Automation: Experience with Databricks Asset Bundles (DABs), Terraform, Git-based release pipelines, policy-as-code, automated testing, and environment promotion. Advanced Databricks Workloads: Experience with Delta Sharing, MLflow, custom APIs, BI integration, AI/ML workloads, and governed data-product consumption patterns. Security Frameworks: Familiarity with Zero Trust principles, NIST CSF, CIS Controls, threat modeling, penetration testing, and audit evidence practices. Domain Experience: Experience in semiconductor manufacturing, R&D, lab/metrology data, equipment telemetry, or OT-integrated environments is a strong advantage. Certifications: Relevant Databricks, Cloud Architecture, Data Engineering, or Security/Governance certifications.Skills: leadership,data,integration,cloud,management,data engineering,databricks & lakehouse,data engineering & analysis,architecture,tenant architecture,security