{"schemaVersion":"jobsearcher.job.v1","id":"dfc524bf1da32c4e33df6b47","url":"https://jobsearcher.com/jobs/dfc524bf1da32c4e33df6b47","canonicalUrl":"https://jobsearcher.com/jobs/dfc524bf1da32c4e33df6b47","title":"Data Acquisition Engineer","description":"We are looking for a Data Acquisition Engineer to lead the ingestion, parsing, cleaning, and structuring of large external datasets used across our financial data platform. You will work with a wide range of inputs and turn them into clean, well-organized, analysis-ready assets.\nWe are looking for a Data Acquisition Engineer to lead the ingestion, parsing, cleaning, and structuring of large external datasets used across our financial data platform. You will work with a wide range of inputs—including public datasets, financial reference data, and alternative data—and turn them into clean, well-organized, analysis-ready assets.\nYou will take datasets that have been identified for integration and handle the full technical onboarding lifecycle: ingesting raw data, designing parsing logic, developing cleaning and normalization workflows, implementing validation checks, and producing high-quality structured outputs. If you enjoy bringing order to messy, heterogeneous datasets and building reliable ingestion pipelines that support financial research and analytics, this role is for you.\nResponsibilities Own the technical onboarding of external datasets, including parsing raw files, transforming fields, and producing clean structured outputs.\nWrite parsing and transformation logic in Python and SQL to handle diverse file formats (CSV, JSON, XML, HTML, XBRL, PDF, etc.).\nDevelop reproducible ETL/ELT workflows that clean, normalize, validate, and structure incoming datasets.\nManage data storage and processing workflows using S3-compatible object storage systems.\nProduce efficient, analytics-ready Parquet datasets, using appropriate partitioning and metadata conventions.\nImplement data-quality checks to detect anomalies, schema drift, missing fields, or unexpected changes in incoming data.\nTroubleshoot and resolve inconsistencies through systematic, transparent cleaning and transformation rules.\nCollaborate with internal data and research teams to understand dataset characteristics, quirks, semantics, and intended uses.\nProvide light technical input during dataset evaluation, offering insight into ingest feasibility and transformation complexity.\nWrite clear documentation describing dataset structure, parsing assumptions, transformation logic, and known limitations.\nSkills & Qualifications Direct experience working with financial datasets at a hedge fund, financial institution, or major financial data provider.\nProven track record onboarding or structuring large, complex financial or finance‑adjacent datasets, such as market data, fundamentals, regulatory data, reference data, or alternative data.\nStrong proficiency in Python for parsing, cleaning, and transformation workflows.\nStrong SQL skills for exploration, validation, and modeling.\nHands‑on experience working with S3-compatible object storage for large dataset management.\nProficiency with Parquet and other columnar storage formats, including partitioning strategies for performance and scale.\nExperience designing ETL/ELT workflows that are reproducible, maintainable, and resilient to upstream dataset changes.\nAbility to interpret messy or loosely documented datasets and design stable parsing logic.\nClear written communication skills for documenting processes, assumptions, and dataset behavior.\nExperience with DataFusion, DuckDB, or other modern analytical engines is a plus.\nExposure to datasets such as regulatory filings, financial reference data, or alternative data is a plus.\nExperience extracting structured information from PDFs or other irregular data sources is a plus.\nFamiliarity with schema validation, metadata management, or data-quality frameworks is a plus.\nAbout Massive At Massive we are on a mission to help developers build the future of fintech. We are committed to democratizing access to the world's financial market data and enabling developers to build the future of fintech. Join us and be part of a team that is revolutionizing the way we interact with money and value.\nMassive is an equal opportunity employer and complies with all applicable federal, state, and local fair employment practices laws. We strictly prohibit and do not tolerate harassment or discrimination against employees, applicants, or any other covered persons because of race, color, religion, creed, national origin or ancestry, ethnicity, sex, gender, gender identity, age, physical or mental disability, citizenship, sexual orientation, past, current or prospective service in the uniformed services. To request a reasonable accommodation, please email careers@massive.com .\nBenefits for full time offers from Massive include, but are not limited to, comprehensive medical plans, 401(k), and unlimited time off. When determining a candidate’s compensation, we consider a number of factors including skillset, experience, job scope, and current market data.\nModernizing Wall St. Reimagining financial market data for the 21st century.\n\n#J-18808-Ljbffr","company":"Polygonio","rawCompany":"polygonio","city":"Brooklyn","state":"NY","isRemote":false,"isActive":true,"createdAt":"2026-08-25T04:02:18.717Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"522320","title":"Financial Transactions Processing, Reserve, and Clearinghouse Activities","slug":"financial-transactions-processing-reserve-and-clearinghouse-activities"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Data Acquisition Engineer","description":"We are looking for a Data Acquisition Engineer to lead the ingestion, parsing, cleaning, and structuring of large external datasets used across our financial data platform. You will work with a wide range of inputs and turn them into clean, well-organized, analysis-ready assets.\nWe are looking for a Data Acquisition Engineer to lead the ingestion, parsing, cleaning, and structuring of large external datasets used across our financial data platform. You will work with a wide range of inputs—including public datasets, financial reference data, and alternative data—and turn them into clean, well-organized, analysis-ready assets.\nYou will take datasets that have been identified for integration and handle the full technical onboarding lifecycle: ingesting raw data, designing parsing logic, developing cleaning and normalization workflows, implementing validation checks, and producing high-quality structured outputs. If you enjoy bringing order to messy, heterogeneous datasets and building reliable ingestion pipelines that support financial research and analytics, this role is for you.\nResponsibilities Own the technical onboarding of external datasets, including parsing raw files, transforming fields, and producing clean structured outputs.\nWrite parsing and transformation logic in Python and SQL to handle diverse file formats (CSV, JSON, XML, HTML, XBRL, PDF, etc.).\nDevelop reproducible ETL/ELT workflows that clean, normalize, validate, and structure incoming datasets.\nManage data storage and processing workflows using S3-compatible object storage systems.\nProduce efficient, analytics-ready Parquet datasets, using appropriate partitioning and metadata conventions.\nImplement data-quality checks to detect anomalies, schema drift, missing fields, or unexpected changes in incoming data.\nTroubleshoot and resolve inconsistencies through systematic, transparent cleaning and transformation rules.\nCollaborate with internal data and research teams to understand dataset characteristics, quirks, semantics, and intended uses.\nProvide light technical input during dataset evaluation, offering insight into ingest feasibility and transformation complexity.\nWrite clear documentation describing dataset structure, parsing assumptions, transformation logic, and known limitations.\nSkills & Qualifications Direct experience working with financial datasets at a hedge fund, financial institution, or major financial data provider.\nProven track record onboarding or structuring large, complex financial or finance‑adjacent datasets, such as market data, fundamentals, regulatory data, reference data, or alternative data.\nStrong proficiency in Python for parsing, cleaning, and transformation workflows.\nStrong SQL skills for exploration, validation, and modeling.\nHands‑on experience working with S3-compatible object storage for large dataset management.\nProficiency with Parquet and other columnar storage formats, including partitioning strategies for performance and scale.\nExperience designing ETL/ELT workflows that are reproducible, maintainable, and resilient to upstream dataset changes.\nAbility to interpret messy or loosely documented datasets and design stable parsing logic.\nClear written communication skills for documenting processes, assumptions, and dataset behavior.\nExperience with DataFusion, DuckDB, or other modern analytical engines is a plus.\nExposure to datasets such as regulatory filings, financial reference data, or alternative data is a plus.\nExperience extracting structured information from PDFs or other irregular data sources is a plus.\nFamiliarity with schema validation, metadata management, or data-quality frameworks is a plus.\nAbout Massive At Massive we are on a mission to help developers build the future of fintech. We are committed to democratizing access to the world's financial market data and enabling developers to build the future of fintech. Join us and be part of a team that is revolutionizing the way we interact with money and value.\nMassive is an equal opportunity employer and complies with all applicable federal, state, and local fair employment practices laws. We strictly prohibit and do not tolerate harassment or discrimination against employees, applicants, or any other covered persons because of race, color, religion, creed, national origin or ancestry, ethnicity, sex, gender, gender identity, age, physical or mental disability, citizenship, sexual orientation, past, current or prospective service in the uniformed services. To request a reasonable accommodation, please email careers@massive.com .\nBenefits for full time offers from Massive include, but are not limited to, comprehensive medical plans, 401(k), and unlimited time off. When determining a candidate’s compensation, we consider a number of factors including skillset, experience, job scope, and current market data.\nModernizing Wall St. Reimagining financial market data for the 21st century.\n\n#J-18808-Ljbffr","datePosted":"2026-08-25T04:02:18.717Z","dateModified":"2026-08-25T04:02:18.717Z","hiringOrganization":{"@type":"Organization","name":"Polygonio","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Brooklyn","addressRegion":"NY","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"dfc524bf1da32c4e33df6b47"},"url":"https://jobsearcher.com/jobs/dfc524bf1da32c4e33df6b47"}}