JOBSEARCHER

Data Engineer

About The RoleYou will engineer the critical path of CYBERIA's Data Platform — ingestion, transformation, feature pipelines, and vector retrieval — with SLOs for latency, freshness, and quality.ResponsibilitiesBuild and operate batch and streaming pipelines with offline + online parityImplement feature stores with sub-50 ms p99 reads and point-in-time correctnessDevelop pipeline designers that turn documents (PDF, DOCX) into structured, AI-ready dataInstrument everything: lineage, data quality checks, and cost telemetry by defaultHarden prototypes from applied ML teams into production-grade platform featuresRequirements4+ years engineering production data systemsStrong Python plus SQL; experience with Spark, dbt, Airflow/Dagster, or similarFamiliarity with vector databases, embeddings, or retrieval systems a strong plusComfort with distributed systems, observability, and on-callBenefitsFully remote within the USCompetitive base + equityLearning and conference budgetPremium health benefits and 401(k) match