JOBSEARCHER

Founding Data Engineer

OverviewHolly is the HR platform built for city and county government. We help local governments modernize how they hire, classify, and manage their workforce — work that directly shapes public services for millions of people. Our platform is live across 12 states with 60+ jurisdictions representing over 10% of the US population, including major counties like Santa Clara and Contra Costa in the Bay Area, Orange County in LA, and Snohomish County in Washington State. Demand is outpacing our team, and we're building to meet it.Holly is a seed-stage team of 20+ operators with deep roots in public service, civic tech, and AI. We’ve raised $10M from leading investors in government technology and the future of work and have grown over 100X in the last year. We center public servants and impact in every decision, and we bring that same care to how we work here at Holly.About The RoleWe're hiring our Founding Data Engineer to build the data backbone of Holly. Local governments publish enormous amounts of public information - salary schedules, job classifications, MOUs (Labor Union Agreements), budgets - but it's scattered across thousands of websites and buried in messy formats: scanned PDFs, inconsistent HTML, spreadsheets, and everything in between. Turning that chaos into clean, trustworthy, structured data is our single biggest data challenge and one of our deepest moats.You'll design and build the data platform: the pipelines and systems that ingest public government data at scale, normalize and validate it, and serve it as clean, canonical datasets the rest of Holly's product builds on as its source of truth. Over time, you'll make this platform increasingly automated and intelligent - less manual wrangling, more self-healing, monitored, high-quality pipelines.This is a hands-on, high-ownership building role. As the first data hire on a small engineering team, you'll set the direction, make the architecture calls, and establish the standards for how Holly collects, models, and trusts its data for years to come. You'll work directly with the founders and partner closely with our product engineers - your job is to make sure they always have the clean, reliable data they need to build features.If you've worked with large, high-volume data and love the challenge of taming messy real-world inputs into something people can rely on, we'd love to talk.What You'll DoYou'll be our founding data engineer and a core partner to the founding team - owning the data platform end to end, making high-leverage technical decisions, and building the foundation the rest of the product depends on.Own and Build the Data PlatformOwn the data platform end-to-end - from raw public sources to clean, canonical datasets the product consumesDesign the architecture, schemas, and standards for how Holly ingests, models, and trusts its dataPartner with the founders to scope work, make tradeoffs, and drive delivery on our highest-leverage data initiativesSet the long-term direction for our data foundation as the first data hireBuild Ingestion & Normalization PipelinesBuild systems that collect large volumes of public government data from thousands of local-government sources across the webTurn messy, heterogeneous inputs - scanned PDFs, inconsistent HTML, spreadsheets - into structured, normalized data (parsing, extraction, OCR, dedupe, entity resolution, schema mapping)Where it adds leverage, incorporate LLM-assisted extraction and embeddings into the pipelineBuild for freshness, reliability, and scale so data stays current and trustworthyModel & Serve Data for the ProductDesign canonical data models and domain schemas that product engineers build onExpose clean, versioned, well-documented datasets the main app can reliably consumeOwn data quality, validation, lineage, and observability so downstream teams can trust what they're building onMake It Automated & IntelligentEvolve pipelines from manual/one-off toward automated, self-healing, monitored systemsEstablish data-quality checks, alerting, and standards that keep the platform reliable as it growsRaise the bar on how we collect, validate, and serve data across the companyHow We WorkSix principles drive how we build:Work on What Matters, Default to No - every yes has a cost, so we save them for what moves the business and spend time on what matters.Question Everything, Be Opinionated - titles don't settle arguments, the better case wins. Feel empowered to push back to everyone from the Head of Engineering to one of the founders. Obsess Over Craft - quality first, and we don't trade it for a date. If you wouldn't put your name on it, it shouldn’t end up in the codebase.Own It End to End - if you build it, you own it: to production, in tests, and when it breaks. With great power comes great responsibility, with autonomy comes responsibility to make sure you own your work.Ship Small, Ship Often - the smallest thing that stands on its own, kept reversible. Small ships compound, are easier to review and easier to fix if there are issues.Automate the Hurt, Not the Itch - automate the recurring pain, the Toil aka things you do repeatedly that waste time, not the one-off annoyances or what seems “fun” to automate. What You'll HaveWe'd love to talk if you're a strong data engineer with high ownership who has worked with large, high-volume data and knows how to turn messy real-world inputs into clean, trustworthy datasets.The EssentialsHave 5+ years building and shipping production software (or equivalent experience)Are a senior data engineer with a strong track record building and operating production data systems (several years of relevant experience or equivalent)Have worked with large-scale, high-volume data - ideally where lots of sources, users, or records make volume and reliability matterAre strong at data modeling and SQL, with experience designing schemas that others build on (Postgres a plus)Have built and owned ETL/ELT pipelines that handle messy, heterogeneous, real-world inputs (scraped data, PDFs, HTML, spreadsheets)Bring a strong data-quality mindset - validation, testing, monitoring, lineage, and reliability are core to how you workTake ownership and move fast - you work independently, ship often, and thrive in early-stage ambiguityHave a growth mindset - you learn quickly, and raise the bar through collaboration and clear standardsAre pragmatic about tooling and comfortable working in (or ramping quickly into) a modern TypeScript/Postgres codebaseBonus PointsOpen source contributions to or maintainer of a widely used tool.Experience with large-scale web scraping / crawling, document extraction (OCR), or LLM-assisted parsingExperience with embeddings / vector search or supporting ML/AI data workflowsExperience with analytical/columnar or warehouse stacks (ClickHouse, BigQuery, Snowflake) and/or streaming pipelinesExperience in government, public sector, or civic techPrior early-stage startup experienceDon't meet every bullet? Apply anyway. If you're strong on most of this and excited about the work, we want to hear from you - we'll help you ramp on the rest.What You'll GetFoundational ownership. Architect and build the data platform that will define our product and data model for years.Technical influence. Make the high-leverage calls on data architecture, standards, and how we scale.Commitment to Open Source. We are big believers in supporting open source, and provide a monthly day of Open Source where you can work on your favorite tool. In addition to internal hackathons and other projects!Founder-level access. Work directly with the founders, with autonomy to drive major initiatives end-to-end.High-impact scope. Build the data foundation the entire product depends on, and see its power features customers rely on quickly.Public-service impact. Your work improves how local governments operate, helping millions of Americans access public-service careers.Competitive package. $170k -$216k base, 0.15-0.40% equity (L3), comprehensive health benefits (platinum plan with vision and dental), 401(k), paid parental leave, and a professional development stipend.Ready to Join Us? A few important notes:Location: This is an onsite role based out of our New York City HQ, four days a week, with some flexibility depending on the role and the candidate. Candidates must reside in New York or be able to commute to our NYC office. Applicants must be authorized to work in the U.S. without requiring sponsorship.Work Philosophy: We’re an early-stage startup serving government clients with hard deadlines. There may be occasional off-hours work around launches or critical issues (rare and typically planned). We value flexibility and trust you to manage your schedule while maintaining a high bar for responsiveness and customer outcomes.We're excited to build with you!Team Holly 🌆www.hollygov.comHolly is committed to building a diverse company and working with the broadest talent pool possible. We encourage applications from all races, religions, national origins, genders, sexual orientations, gender identities, gender expressions, and ages, as well as veterans and individuals with disabilities. If you need a reasonable accommodation during the application or interview process, let us know.The Pay Range For This Role Is170,000 - 216,000 USD per year (hq)