JOBSEARCHER

Data Engineer

Data Infrastructure EngineerYou'll build and scale the data infrastructure that powers Surf — the unified data layer that replaces 5+ data providers for our users. This means ingesting, processing, and serving on-chain data from 40+ blockchains, social data from 100K+ crypto KOLs, and market data with 200+ technical indicators — all in real-time.ResponsibilitiesDefine the crypto/web3 data platform roadmap; set and meet SLO/SLI for freshness, latency, availability, accuracy, and costBuild resilient multi-chain pipelines with unified schemas across heterogeneous on-chain, social, and market data sourcesArchitect lake/warehouse/lakehouse stacks; select the right OLTP/OLAP/vector stores (PostgreSQL/pgvector, ClickHouse, etc.) for streaming and batch workloadsBuild ETL/ELT and event-driven pipelines to ingest and normalize on-chain data across wallets, NFTs, DeFi, and social use cases; implement incremental sync, late-arrival handling, upserts, and materialized views for low-latency readsBuild embedding pipelines, retrieval acceleration, and feedback/evaluation loops to improve AI answer qualityDrive canonical modeling, entity resolution, schema versioning, lineage, data contracts, and anomaly detection; enforce consistent naming and version controlEnforce access control, encryption, secrets management, auditing, and secure data sharing across environmentsBuild and operate data APIs (REST, streaming, SQL/GraphQL, webhooks, bulk downloads) with multi-tenant isolation, metering, rate limits, versioned endpoints, SDKs, and SLAsBuild and optimize multi-chain node infrastructure: load balancing, geo-routing, failover, mempool subscriptions, transaction tracking, and gas/fee estimationDeliver wallet risk scoring, NFT/DeFi analytics dashboards, on-chain monitoring/alerting, compliance/audit trails, and curated datasets for AI trainingQualifications5+ years of experience with 3+ years in web3/crypto data engineering or platform rolesProficiency in Python, Go, and SQL; hands-on with Airflow/Dagster/dbt for orchestration and modelingPractical experience with Kafka/Pulsar/Pub-Sub and stream processing (Flink/Spark) for large, heterogeneous datasetsComfortable with Docker/Kubernetes, IaC (Terraform/Helm), CI/CD; cost/perf governance on AWS/GCPApply at careers@cybertinolab.com