{"schemaVersion":"jobsearcher.job.v1","id":"a6325c30271dd8a887c6a743","url":"https://jobsearcher.com/jobs/a6325c30271dd8a887c6a743","canonicalUrl":"https://jobsearcher.com/jobs/a6325c30271dd8a887c6a743","title":"Data Engineer","description":"What sets INVID apart is our collaborative and flexible work environment. We encourage our team to raise the bar in everything they do while maintaining a healthy work-life balance. With our hybrid work model, team members thrive both in the office and remotely. We foster a culture of mutual respect, autonomy, and accountability, where your voice matters and your growth is supported. From structured career paths and paid professional development to access to industry events, we’re committed to your success.\n\nJoin us at INVID, where innovation meets support, and together we deliver excellence.\n\nJob Description\nWe are hiring a Data Engineer to build the data infrastructure powering our predictive analytics initiative. You will create the pipelines that turn raw vessel tracking data into training datasets for ML models. The core challenge: we have rich behavioral data (vessel positions, AIS gaps, ship-to-ship transfers, spoofing events) but limited labeled outcomes (confirmed violations, detentions, seizures). You will build pipelines that create usable training data through proxy labels, data joins, and outcome correlation.\n\nResponsibilities\n\nBuild labeling pipelines that join behavioral events to outcome data (sanctions designations, flag changes,\ndetentions)\n\nImplement proxy labeling strategies that create training signal from observable outcomes\nBuild weak supervision infrastructure to combine multiple noisy labeling rules\nCreate and maintain ML training datasets at scale\nBuild data validation and quality monitoring systems\nImplement versioning for reproducible model training\nIntegrate LRIT position data for prediction validation\nBuild pipelines that compare predicted locations against actual LRIT reports\nCreate feedback loops that improve model accuracy over time\nScale data infrastructure as models and data sources grow\n\nRequired Skill\n\n4+ years data engineering experience\nStrong SQL skills, including complex joins across large datasets\nExperience with Spark, Airflow, or equivalent distributed processing frameworks\nPython for data processing and pipeline orchestration\nAWS experience\nUnderstanding of ML training data requirements\n\nEducation/Certifications\n\nBachelor's Degree in Computer Science, Engineering, or related field\nDesired Skills (Not Required)\n\nExperience with geospatial data (PostGIS, H3, spatial joins)\nMaritime, defense, or intelligence domain experience\nExperience with data labeling infrastructure or weak supervision\nFamiliarity with real-time streaming data systems\n\nImportant:\nMust be a U.S. citizen and a U.S. resident\nThis job works on a hybrid work modality (San Juan, Puerto Rico)\nMust have a valid driver's license\n\nEEO","company":"Invid","rawCompany":"invid","city":"El Paso","state":"TX","isRemote":false,"isActive":false,"createdAt":"2026-07-15T13:49:32.431Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Data Engineer","description":"What sets INVID apart is our collaborative and flexible work environment. We encourage our team to raise the bar in everything they do while maintaining a healthy work-life balance. With our hybrid work model, team members thrive both in the office and remotely. We foster a culture of mutual respect, autonomy, and accountability, where your voice matters and your growth is supported. From structured career paths and paid professional development to access to industry events, we’re committed to your success.\n\nJoin us at INVID, where innovation meets support, and together we deliver excellence.\n\nJob Description\nWe are hiring a Data Engineer to build the data infrastructure powering our predictive analytics initiative. You will create the pipelines that turn raw vessel tracking data into training datasets for ML models. The core challenge: we have rich behavioral data (vessel positions, AIS gaps, ship-to-ship transfers, spoofing events) but limited labeled outcomes (confirmed violations, detentions, seizures). You will build pipelines that create usable training data through proxy labels, data joins, and outcome correlation.\n\nResponsibilities\n\nBuild labeling pipelines that join behavioral events to outcome data (sanctions designations, flag changes,\ndetentions)\n\nImplement proxy labeling strategies that create training signal from observable outcomes\nBuild weak supervision infrastructure to combine multiple noisy labeling rules\nCreate and maintain ML training datasets at scale\nBuild data validation and quality monitoring systems\nImplement versioning for reproducible model training\nIntegrate LRIT position data for prediction validation\nBuild pipelines that compare predicted locations against actual LRIT reports\nCreate feedback loops that improve model accuracy over time\nScale data infrastructure as models and data sources grow\n\nRequired Skill\n\n4+ years data engineering experience\nStrong SQL skills, including complex joins across large datasets\nExperience with Spark, Airflow, or equivalent distributed processing frameworks\nPython for data processing and pipeline orchestration\nAWS experience\nUnderstanding of ML training data requirements\n\nEducation/Certifications\n\nBachelor's Degree in Computer Science, Engineering, or related field\nDesired Skills (Not Required)\n\nExperience with geospatial data (PostGIS, H3, spatial joins)\nMaritime, defense, or intelligence domain experience\nExperience with data labeling infrastructure or weak supervision\nFamiliarity with real-time streaming data systems\n\nImportant:\nMust be a U.S. citizen and a U.S. resident\nThis job works on a hybrid work modality (San Juan, Puerto Rico)\nMust have a valid driver's license\n\nEEO","datePosted":"2026-07-15T13:49:32.431Z","dateModified":"2026-07-15T13:49:32.431Z","hiringOrganization":{"@type":"Organization","name":"Invid","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"El Paso","addressRegion":"TX","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"a6325c30271dd8a887c6a743"},"url":"https://jobsearcher.com/jobs/a6325c30271dd8a887c6a743"}}