{"schemaVersion":"jobsearcher.job.v1","id":"aedd8bf2b8a94525112e8341","url":"https://jobsearcher.com/jobs/aedd8bf2b8a94525112e8341","canonicalUrl":"https://jobsearcher.com/jobs/aedd8bf2b8a94525112e8341","title":"Founding Data Engineer","description":"Position SummaryAirys is an AI-powered platform helping nonprofits and community organizations unlock climate resilience data and funding opportunities that are often buried in fragmented municipal systems. Our mission is to make resilience planning and investment transparent, accessible, and actionable, so nonprofits can better advocate for vulnerable communities, guide equitable infrastructure decisions, and secure critical resources.At the same time, Airys provides the same structured data to the private sector—commercial real estate, insurers, and infrastructure investors—who need clear visibility into how cities are preparing for climate risks. By connecting nonprofit and community priorities with private-sector diligence needs, we create a shared data backbone that drives investment toward projects that strengthen resilience for all.Building the first comprehensive database and AI-powered platform that aggregates climate resilience insights from various planning documents across all US states. The platform will enable investment professionals to quickly research climate adaptation efforts for due diligence.Trial: 1-3 weeks, 20 hrs/week starting immediately$40-62.5/hr based on experienceRemote or hybrid SF Bay AreaPotential founding team role based on performance/mutual fit What You’ll BuildTransform our scrappy prototype into a scalable system that processes tens of thousands of municipal documents nationwide, extracts structured project data using AI, and provides fast location-based search for climate adaptation projects.Key ResponsibilitiesScale the ArchitectureMigrate prototype to production cloud infrastructure (AWS/GCP)Build distributed systems for parallel document processing and web scrapingDesign scalable databases (relational + vector) with cost/performance optimizationProduction Data PipelineCreate robust ETL with error handling, monitoring, and automated retriesImplement accurate geocoding across inconsistent municipal address formatsStandardize data validation across diverse state/municipal document typesBuild RESTful APIs and efficient search functionalityAI-Powered ExtractionScale LLM-based PDF processing while managing API costsImplement semantic search across infrastructure project databasesEnhance AI extraction accuracy from unstructured municipal documentsTechnical RequirementsLanguages: Python, TypeScript, SQLCloud Platforms: AWS, GCP, or Azure with distributed systems experienceDatabases: PostgreSQL, vector databases (Pinecone, Supabase)AI/ML: LangChain, vector embeddings, RAG, conversational agentsData Pipeline Tools: Apache Airflow (or similar tools)Web Scraping: Scrapy, Selenium, Google Custom Search API Geocoding: Google Maps API, OpenStreetMap, PostGISPDF Processing: Text extraction and document parsing librariesNice to haveTechnical MindsetThinks in systems and can architect for scale from day oneComfortable making technical decisions with limited guidanceExperience debugging production issues and optimizing performanceWants to shape engineering culture and hiring as the team growsEnjoy working in start-up environmentTakes ownership and drives projects to completionAdapts quickly to changing requirements and prioritiesDirect communication style with both technical and non-technical stakeholdersPassionate About Urban Climate Adaptation And ResilienceBackground in municipal/government document analysisFamiliarity with infrastructure planning or environmental dataInterest in climate risk assessment or sustainability techPrevious work with public sector or policy-related datasets","company":"Progressive Data","rawCompany":"progressive data","city":"California","state":"MO","isRemote":false,"isActive":false,"createdAt":"2026-08-14T12:04:39.782Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Founding Data Engineer","description":"Position SummaryAirys is an AI-powered platform helping nonprofits and community organizations unlock climate resilience data and funding opportunities that are often buried in fragmented municipal systems. Our mission is to make resilience planning and investment transparent, accessible, and actionable, so nonprofits can better advocate for vulnerable communities, guide equitable infrastructure decisions, and secure critical resources.At the same time, Airys provides the same structured data to the private sector—commercial real estate, insurers, and infrastructure investors—who need clear visibility into how cities are preparing for climate risks. By connecting nonprofit and community priorities with private-sector diligence needs, we create a shared data backbone that drives investment toward projects that strengthen resilience for all.Building the first comprehensive database and AI-powered platform that aggregates climate resilience insights from various planning documents across all US states. The platform will enable investment professionals to quickly research climate adaptation efforts for due diligence.Trial: 1-3 weeks, 20 hrs/week starting immediately$40-62.5/hr based on experienceRemote or hybrid SF Bay AreaPotential founding team role based on performance/mutual fit What You’ll BuildTransform our scrappy prototype into a scalable system that processes tens of thousands of municipal documents nationwide, extracts structured project data using AI, and provides fast location-based search for climate adaptation projects.Key ResponsibilitiesScale the ArchitectureMigrate prototype to production cloud infrastructure (AWS/GCP)Build distributed systems for parallel document processing and web scrapingDesign scalable databases (relational + vector) with cost/performance optimizationProduction Data PipelineCreate robust ETL with error handling, monitoring, and automated retriesImplement accurate geocoding across inconsistent municipal address formatsStandardize data validation across diverse state/municipal document typesBuild RESTful APIs and efficient search functionalityAI-Powered ExtractionScale LLM-based PDF processing while managing API costsImplement semantic search across infrastructure project databasesEnhance AI extraction accuracy from unstructured municipal documentsTechnical RequirementsLanguages: Python, TypeScript, SQLCloud Platforms: AWS, GCP, or Azure with distributed systems experienceDatabases: PostgreSQL, vector databases (Pinecone, Supabase)AI/ML: LangChain, vector embeddings, RAG, conversational agentsData Pipeline Tools: Apache Airflow (or similar tools)Web Scraping: Scrapy, Selenium, Google Custom Search API Geocoding: Google Maps API, OpenStreetMap, PostGISPDF Processing: Text extraction and document parsing librariesNice to haveTechnical MindsetThinks in systems and can architect for scale from day oneComfortable making technical decisions with limited guidanceExperience debugging production issues and optimizing performanceWants to shape engineering culture and hiring as the team growsEnjoy working in start-up environmentTakes ownership and drives projects to completionAdapts quickly to changing requirements and prioritiesDirect communication style with both technical and non-technical stakeholdersPassionate About Urban Climate Adaptation And ResilienceBackground in municipal/government document analysisFamiliarity with infrastructure planning or environmental dataInterest in climate risk assessment or sustainability techPrevious work with public sector or policy-related datasets","datePosted":"2026-08-14T12:04:39.782Z","dateModified":"2026-08-14T12:04:39.782Z","hiringOrganization":{"@type":"Organization","name":"Progressive Data","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"California","addressRegion":"MO","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"aedd8bf2b8a94525112e8341"},"url":"https://jobsearcher.com/jobs/aedd8bf2b8a94525112e8341"}}