Data Analyst
Responsibilities:
Build, optimize, and operate production Apache Spark pipelines that generate attributed conversion and optimization feeds from large-scale impression and conversion data.
Design and evolve data processing and models across modern data lake and warehouse technologies, orchestrated with a production workflow scheduler on AWS.
Build and operate partner optimization feeds and integrations with external partners, including schemas, identity fields and unique IDs, file delivery, reconciliation, and SLAs.
Build, change, and operate production REST APIs, including PHP APIs, delivering secure, backward-compatible changes end to end (implementation, validation, authorization, data access, documentation, automated testing, and production diagnostics).
Trace and modify behavior across controllers, database queries, templates, JavaScript, and automated tests.
Diagnose production failures, reconcile partner-facing data, execute backfills, and manage staged releases and rollbacks.
Write and maintain unit and integration tests across pipelines, APIs, and feeds.
Qualifications and Education Requirements:
7+ years of professional software or data engineering experience.
4+ years building, optimizing, and maintaining production Apache Spark pipelines, with strong Java and Python skills across JVM Spark and PySpark.
Advanced SQL, relational data modeling, and large-scale data processing experience, including production work with Snowflake, MySQL, and Iceberg.
3+ years operating AWS data workloads using S3 and Airflow (or equivalent production workflow orchestration); container orchestration and data-catalog experience a plus.
3+ years building, changing, and operating production REST APIs, including PHP APIs built with Symfony or a closely equivalent PHP MVC framework.
Demonstrated proficiency using AI development tools to deliver production-quality work efficiently — including AI-assisted coding, code review, documentation, and investigation, working with agentic code harnesses, and spec-based (spec-driven) AI development — with the judgment to know when to rely on AI output and when not to.
Ability to independently deliver secure, backward-compatible API changes, including implementation, validation, authorization, data access, documentation, automated testing, and production diagnostics.
Experience implementing external provider or client integrations involving schemas, APIs, file delivery, identity fields, reconciliation, privacy-sensitive data, and SLAs.
Production full-stack experience with PHP, Symfony (or a comparable framework), Doctrine, Twig (or another server-rendered template system), and JavaScript.
Ability to trace and modify behavior across controllers, database queries, templates, JavaScript, and automated tests.
Experience writing unit and integration tests for data pipelines, APIs, and web applications.
Ability to diagnose production failures, reconcile data, execute backfills, manage staged releases and rollbacks, and support delivery commitments.
Hands-on experience building and running big data systems, with a focus on performance, reliability, and data quality.
Experience in ad tech or advertising measurement, such as attribution, conversions, audience and impression data, identity matching, or partner optimization feeds.
Ability to ramp quickly and deliver with minimal onboarding in an existing, complex codebase.
Preferred Skills:
Direct experience building partner optimization or advertising feeds with external platforms (DSPs, publishers, or measurement partners).
Experience with deterministic and probabilistic identity matching — unique IDs, device IDs, IP-based matching, and identity graphs.
Familiarity with edge/log delivery infrastructure and pixel/impression tracking.
Experience with data lake table formats and query engines at large scale.
Experience with privacy and compliance-driven data workflows (deletion, opt-out, suppression, data retention).
Comfort operating in uncharted territory — turning ambiguity into production systems without a detailed guide.