{"schemaVersion":"jobsearcher.job.v1","id":"0810072f8a4fba1e6ccdc309","url":"https://jobsearcher.com/jobs/0810072f8a4fba1e6ccdc309","canonicalUrl":"https://jobsearcher.com/jobs/0810072f8a4fba1e6ccdc309","title":"Data Engineer","description":"About Troveo\nTroveo builds the data platform that AI labs and model builders need to train the next generation of models. We have created the world's largest licensed platform of scarce, proprietary data for AI, spanning video, audio, text, and business workflows.\nTroveo indexes, enriches, and packages this high-quality data into formats ready for training, fine-tuning, evaluation, and agentic use cases. Backed by top investors, we’re a small, high-impact team solving one of the biggest bottlenecks in AI development.\nRole Overview\nWe are seeking a versatile, hands-on Data Engineer to build and maintain a scalable analytics data warehouse while contributing to data modeling, performing data analysis, and ensuring the reliable delivery of data to downstream teams and systems. This hybrid role combines core data engineering responsibilities with data modeling, analytics, and operational support. You will own the full analytics data lifecycle, from ingestion and transformation to modeling, quality assurance, and timely delivery, while partnering closely with software engineers and business stakeholders.\nKey Responsibilities\nData Pipeline Engineering\nDesign, build, and maintain ELT/ETL data pipelines (batch and streaming), optimizing for performance, reliability, scalability, and cost.\nDesign and implement conceptual, logical, and physical data models (including dimensional modeling, star/snowflake schemas).\nBuild and maintain transformation layers using modern tools (e.g., dbt) to create clean, well-documented, analytics-ready datasets.\nApply data modeling best practices, versioning, testing, and documentation to ensure consistency and reusability.\nData Analysis & Reporting Support\nWrite optimal SQL queries for data exploration, ad-hoc analysis, and troubleshooting.\nSupport the creation of reports, dashboards, and self-service analytics assets in collaboration with data analysts and business teams.\nTranslate business questions into data requirements and deliver actionable insights or datasets.\nOperational Support & Data Deliveries\nMonitor data pipelines and data delivery processes to ensure SLAs for timeliness, freshness, and accuracy are consistently met.\nProactively identify, troubleshoot, and resolve data issues impacting downstream consumers or business operations.\nManage incidents related to data availability and quality; participate in on-call rotations as needed.\nImplement data quality checks, observability, and alerting to maintain high reliability of data deliveries.\nAutomate operational tasks and continuously improve data delivery processes.\nCollaboration & Best Practices\nWork cross-functionally with analysts, data scientists, engineers, and business stakeholders to understand data needs and deliver solutions.\nDocument data pipelines, models, lineage, and processes.\nContribute to data governance, security, and best practices across the data platform.\nRequirements\n7+ years of professional experience in data engineering or a closely related role (analytics engineering experience is highly relevant).\nStrong proficiency in SQL and Python.\nHands-on experience building and maintaining data pipelines and working with cloud data platforms/warehouses. (Snowflake, BigQuery, Redshift, Databricks, etc.).\nExperience with data orchestration tools (Apache Airflow, Dagster, Prefect, or similar).\nSolid understanding of data modeling techniques and dimensional modeling.\nExperience performing data analysis and working with BI/visualization tools (Looker, Tableau, Power BI, or similar).\nProven ability to troubleshoot data issues and support operational reliability/SLAs.\nStrong communication skills and ability to collaborate with both technical and non-technical stakeholders.\nBonus Points\nExperience with DBT for data transformation and modeling.\nKnowledge of data observability/monitoring tools.\nExperience with real-time/streaming data technologies (Kafka, Flink, etc.).\nFamiliarity with CI/CD practices for data pipelines.\nExperience in data quality frameworks and governance.\nBachelor’s degree in Computer Science, Engineering, or a related quantitative field (or equivalent practical experience).\nCompensation\nBase Salary: $100,000 – $140,000 (depending on experience and location)\nEquity: Competitive equity package in a well-funded AI startup with significant upside\nCompensation is location-adjusted for cost of living. We are open to candidates in California, New York, and select other states.\nWhat We Offer\nComprehensive Health Benefits: Medical, dental, and vision coverage (100% employer-paid for employees)\nFlexible PTO & Paid Holidays: Unlimited PTO with encouragement to actually use it\nRemote First Policy: Work from anywhere in the US (with occasional team offsites)\nLearning & Growth: Annual learning stipend, access to top conferences, and direct mentorship from experienced founders\nEquity Ownership: Competitive equity package with clear growth potential as we scale\nModern Tech Stack & Tools: Budget for the best equipment and software\nStrong Culture: High-trust, low-ego environment focused on impact, transparency, and work-life balance. We believe great work happens when people are supported, challenged, and given ownership.\nEqual Opportunity Employer\nTroveo is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.","company":"Troveo Ai","rawCompany":"troveo ai","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-04T19:50:13.505Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Data Engineer","description":"About Troveo\nTroveo builds the data platform that AI labs and model builders need to train the next generation of models. We have created the world's largest licensed platform of scarce, proprietary data for AI, spanning video, audio, text, and business workflows.\nTroveo indexes, enriches, and packages this high-quality data into formats ready for training, fine-tuning, evaluation, and agentic use cases. Backed by top investors, we’re a small, high-impact team solving one of the biggest bottlenecks in AI development.\nRole Overview\nWe are seeking a versatile, hands-on Data Engineer to build and maintain a scalable analytics data warehouse while contributing to data modeling, performing data analysis, and ensuring the reliable delivery of data to downstream teams and systems. This hybrid role combines core data engineering responsibilities with data modeling, analytics, and operational support. You will own the full analytics data lifecycle, from ingestion and transformation to modeling, quality assurance, and timely delivery, while partnering closely with software engineers and business stakeholders.\nKey Responsibilities\nData Pipeline Engineering\nDesign, build, and maintain ELT/ETL data pipelines (batch and streaming), optimizing for performance, reliability, scalability, and cost.\nDesign and implement conceptual, logical, and physical data models (including dimensional modeling, star/snowflake schemas).\nBuild and maintain transformation layers using modern tools (e.g., dbt) to create clean, well-documented, analytics-ready datasets.\nApply data modeling best practices, versioning, testing, and documentation to ensure consistency and reusability.\nData Analysis & Reporting Support\nWrite optimal SQL queries for data exploration, ad-hoc analysis, and troubleshooting.\nSupport the creation of reports, dashboards, and self-service analytics assets in collaboration with data analysts and business teams.\nTranslate business questions into data requirements and deliver actionable insights or datasets.\nOperational Support & Data Deliveries\nMonitor data pipelines and data delivery processes to ensure SLAs for timeliness, freshness, and accuracy are consistently met.\nProactively identify, troubleshoot, and resolve data issues impacting downstream consumers or business operations.\nManage incidents related to data availability and quality; participate in on-call rotations as needed.\nImplement data quality checks, observability, and alerting to maintain high reliability of data deliveries.\nAutomate operational tasks and continuously improve data delivery processes.\nCollaboration & Best Practices\nWork cross-functionally with analysts, data scientists, engineers, and business stakeholders to understand data needs and deliver solutions.\nDocument data pipelines, models, lineage, and processes.\nContribute to data governance, security, and best practices across the data platform.\nRequirements\n7+ years of professional experience in data engineering or a closely related role (analytics engineering experience is highly relevant).\nStrong proficiency in SQL and Python.\nHands-on experience building and maintaining data pipelines and working with cloud data platforms/warehouses. (Snowflake, BigQuery, Redshift, Databricks, etc.).\nExperience with data orchestration tools (Apache Airflow, Dagster, Prefect, or similar).\nSolid understanding of data modeling techniques and dimensional modeling.\nExperience performing data analysis and working with BI/visualization tools (Looker, Tableau, Power BI, or similar).\nProven ability to troubleshoot data issues and support operational reliability/SLAs.\nStrong communication skills and ability to collaborate with both technical and non-technical stakeholders.\nBonus Points\nExperience with DBT for data transformation and modeling.\nKnowledge of data observability/monitoring tools.\nExperience with real-time/streaming data technologies (Kafka, Flink, etc.).\nFamiliarity with CI/CD practices for data pipelines.\nExperience in data quality frameworks and governance.\nBachelor’s degree in Computer Science, Engineering, or a related quantitative field (or equivalent practical experience).\nCompensation\nBase Salary: $100,000 – $140,000 (depending on experience and location)\nEquity: Competitive equity package in a well-funded AI startup with significant upside\nCompensation is location-adjusted for cost of living. We are open to candidates in California, New York, and select other states.\nWhat We Offer\nComprehensive Health Benefits: Medical, dental, and vision coverage (100% employer-paid for employees)\nFlexible PTO & Paid Holidays: Unlimited PTO with encouragement to actually use it\nRemote First Policy: Work from anywhere in the US (with occasional team offsites)\nLearning & Growth: Annual learning stipend, access to top conferences, and direct mentorship from experienced founders\nEquity Ownership: Competitive equity package with clear growth potential as we scale\nModern Tech Stack & Tools: Budget for the best equipment and software\nStrong Culture: High-trust, low-ego environment focused on impact, transparency, and work-life balance. We believe great work happens when people are supported, challenged, and given ownership.\nEqual Opportunity Employer\nTroveo is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.","datePosted":"2026-08-04T19:50:13.505Z","dateModified":"2026-08-04T19:50:13.505Z","hiringOrganization":{"@type":"Organization","name":"Troveo Ai","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"0810072f8a4fba1e6ccdc309"},"url":"https://jobsearcher.com/jobs/0810072f8a4fba1e6ccdc309"}}