{"schemaVersion":"jobsearcher.job.v1","id":"6a52dcb6737cac00c158468c","url":"https://jobsearcher.com/jobs/6a52dcb6737cac00c158468c","canonicalUrl":"https://jobsearcher.com/jobs/6a52dcb6737cac00c158468c","title":"VP, Data Engineering","description":"The Mission\nMost data engineering roles are about moving data from A to B. This one is about making 20 years of complex, relational ERP data legible to AI agents — so they can reason over financial transactions, inventory movements, and supply chain events without hallucinating.\nECI is rebuilding how enterprise software is built and operated using an AI-native model. The data layer is the foundation everything else runs on. Without a world-class context engine, the agents are guessing. You are the person who makes sure they never have to.\nThis is a greenfield mandate. You will hire the team, choose the stack, define the architecture, and own the outcome. The CTO is your only direct stakeholder.\nWhat You’ll Own\nYou are not supporting the AI initiative. You are building the infrastructure without which it cannot exist.\nContext Architecture & Retrieval\nDesign and own the retrieval systems that allow AI agents to reason over ERP data with zero hallucinations\nBuild and scale the vector infrastructure — pgvector, Qdrant, or equivalent — with production-grade embedding and reranking pipelines\nOwn the hybrid search strategy: semantic retrieval layered on top of SQL-scoped financial data\nDrive context window optimization — packing the most relevant financial 'truth' into each LLM call efficiently\nKnowledge Graph & MDM\nLead the Master Data Management strategy — golden record survivorship, identity resolution, entity deduplication across ERP entities\nBuild the knowledge graph that maps relationships between Vendors, Purchase Orders, Invoices, GL Entries, and Inventory so agents understand meaning, not just rows\nOwn the semantic layer: translate a 500-table legacy schema into a structured, LLM-readable ontology\nDefine data quality standards and automated validation pipelines that enforce them continuously\nData Platform & Infrastructure\nBuild the core data platform from scratch: ingestion, transformation, storage, and serving layers\nOwn the modern data stack — dbt, Airflow or equivalent, Postgres/SQL Server — with an AI-augmented workflow throughout\nImplement data-centric evals: 'Judge Agents' that verify AI output against ground truth SQL\nBuild synthetic data generation pipelines that produce high-fidelity, relationally consistent ERP data for agent training and testing\nBuilder Data Track\nOwn the Data Builder squad: hire, develop, and hold the team to Builder-level output standards\nPartner with the Dev and QA Builder leads to ensure data systems are the right interface for agentic tool-calling\nRun the Data track of the Builder Bootcamp — define the curriculum, set the graduation bar, make the calls\nPartner with product and engineering on AI feature data requirements — you are the upstream dependency for almost everything\nGovernance & Compliance\nDefine data governance policies for AI-consumed data: lineage, access control, PII handling, audit trails\nOwn compliance requirements relevant to financial data in an ERP context — SOC 2, data residency, retention policies\nBuild the observability layer: OpenTelemetry, Weights & Biases, or equivalent for embedding quality and retrieval performance\nWho you are\nRequirements:\nYou have built and led a data engineering team before — you know how to hire, structure, and technically lead a team that ships production data systems\nKnowledge graph or MDM at scale: you have designed entity resolution, survivorship rules, and ontologies for complex relational domains — not just prototyped them\nAI/ML platform or LLMOps experience: you have operated embedding pipelines, vector stores, and LLM-integrated data systems in production — you understand latency, cost, and quality trade-offs\nYou think in systems: schema design, retrieval architecture, and data contracts are your native language\nYou are comfortable in ambiguity — greenfield means no existing patterns to follow and no team to hand things off to on day one\nHighly Desirable:\nProduction RAG pipelines over structured or financial data — you have gone beyond demos and operated retrieval systems with real precision/recall requirements\nERP, financial, or supply chain data domain — you understand what makes a General Ledger different from a web analytics event stream\nModern data stack depth: dbt, Airflow, Postgres, SQL Server — you have opinions about transformation layer design and know when to break the rules\nExperience working across time zones with an offshore engineering team (India context is a plus)\nThe Stack:\nLanguages\nPython, SQL (Postgres / SQL Server), TypeScript\nAI / Retrieval\nOpenAI / Anthropic APIs, pgvector, Qdrant, LangChain / LangGraph\nData Platform\ndbt (AI-augmented), Apache Airflow, Docker\nGraph / MDM\nNeo4j (primary), with open evaluation of alternatives\nObservability\nWeights & Biases (embedding evals), OpenTelemetry, custom Judge Agents\nInfra\nAWS / GCP, Kubernetes, GitHub Actions\nThe Archetypes we’re looking for:\nThe Data Alchemist — you believe data is only valuable when an AI can reason over it, and you spend time experimenting with embedding models and retrieval techniques to make that true\nThe Manual Mapping Hater — if you have to map two schemas twice, you've already built an agent to do it for you\nRigor over Hype — you know the difference between a vector search demo and a production-grade financial data engine; you care about Precision and Recall\nThe Founding Mindset — you're energized by building from scratch, not managing existing systems, and you make decisions confidently without a playbook\nWhy this role:\nERP data is the hardest data problem in enterprise software — 20 years of relational financial history, undocumented schemas, and zero tolerance for hallucination. If you can solve RAG for an ERP, you have solved the hardest version of the problem.\nGreenfield with real stakes: you are not inheriting someone else's technical debt or org structure. You build what you believe will win.\nDirect line to the CTO — no data governance committee, no analytics manager layer, no 6-month roadmap approval process\nUnlimited context budget: access to frontier models and the compute to run serious embedding and indexing experiments\nThe work matters: every AI feature in the product runs on the infrastructure you build","company":"Eci Software Solutions","rawCompany":"eci software solutions","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-07-27T13:03:42.869Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"},{"code":"11-3021.00","title":"Computer and Information Systems Managers","slug":"computer-and-information-systems-managers"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"VP, Data Engineering","description":"The Mission\nMost data engineering roles are about moving data from A to B. This one is about making 20 years of complex, relational ERP data legible to AI agents — so they can reason over financial transactions, inventory movements, and supply chain events without hallucinating.\nECI is rebuilding how enterprise software is built and operated using an AI-native model. The data layer is the foundation everything else runs on. Without a world-class context engine, the agents are guessing. You are the person who makes sure they never have to.\nThis is a greenfield mandate. You will hire the team, choose the stack, define the architecture, and own the outcome. The CTO is your only direct stakeholder.\nWhat You’ll Own\nYou are not supporting the AI initiative. You are building the infrastructure without which it cannot exist.\nContext Architecture & Retrieval\nDesign and own the retrieval systems that allow AI agents to reason over ERP data with zero hallucinations\nBuild and scale the vector infrastructure — pgvector, Qdrant, or equivalent — with production-grade embedding and reranking pipelines\nOwn the hybrid search strategy: semantic retrieval layered on top of SQL-scoped financial data\nDrive context window optimization — packing the most relevant financial 'truth' into each LLM call efficiently\nKnowledge Graph & MDM\nLead the Master Data Management strategy — golden record survivorship, identity resolution, entity deduplication across ERP entities\nBuild the knowledge graph that maps relationships between Vendors, Purchase Orders, Invoices, GL Entries, and Inventory so agents understand meaning, not just rows\nOwn the semantic layer: translate a 500-table legacy schema into a structured, LLM-readable ontology\nDefine data quality standards and automated validation pipelines that enforce them continuously\nData Platform & Infrastructure\nBuild the core data platform from scratch: ingestion, transformation, storage, and serving layers\nOwn the modern data stack — dbt, Airflow or equivalent, Postgres/SQL Server — with an AI-augmented workflow throughout\nImplement data-centric evals: 'Judge Agents' that verify AI output against ground truth SQL\nBuild synthetic data generation pipelines that produce high-fidelity, relationally consistent ERP data for agent training and testing\nBuilder Data Track\nOwn the Data Builder squad: hire, develop, and hold the team to Builder-level output standards\nPartner with the Dev and QA Builder leads to ensure data systems are the right interface for agentic tool-calling\nRun the Data track of the Builder Bootcamp — define the curriculum, set the graduation bar, make the calls\nPartner with product and engineering on AI feature data requirements — you are the upstream dependency for almost everything\nGovernance & Compliance\nDefine data governance policies for AI-consumed data: lineage, access control, PII handling, audit trails\nOwn compliance requirements relevant to financial data in an ERP context — SOC 2, data residency, retention policies\nBuild the observability layer: OpenTelemetry, Weights & Biases, or equivalent for embedding quality and retrieval performance\nWho you are\nRequirements:\nYou have built and led a data engineering team before — you know how to hire, structure, and technically lead a team that ships production data systems\nKnowledge graph or MDM at scale: you have designed entity resolution, survivorship rules, and ontologies for complex relational domains — not just prototyped them\nAI/ML platform or LLMOps experience: you have operated embedding pipelines, vector stores, and LLM-integrated data systems in production — you understand latency, cost, and quality trade-offs\nYou think in systems: schema design, retrieval architecture, and data contracts are your native language\nYou are comfortable in ambiguity — greenfield means no existing patterns to follow and no team to hand things off to on day one\nHighly Desirable:\nProduction RAG pipelines over structured or financial data — you have gone beyond demos and operated retrieval systems with real precision/recall requirements\nERP, financial, or supply chain data domain — you understand what makes a General Ledger different from a web analytics event stream\nModern data stack depth: dbt, Airflow, Postgres, SQL Server — you have opinions about transformation layer design and know when to break the rules\nExperience working across time zones with an offshore engineering team (India context is a plus)\nThe Stack:\nLanguages\nPython, SQL (Postgres / SQL Server), TypeScript\nAI / Retrieval\nOpenAI / Anthropic APIs, pgvector, Qdrant, LangChain / LangGraph\nData Platform\ndbt (AI-augmented), Apache Airflow, Docker\nGraph / MDM\nNeo4j (primary), with open evaluation of alternatives\nObservability\nWeights & Biases (embedding evals), OpenTelemetry, custom Judge Agents\nInfra\nAWS / GCP, Kubernetes, GitHub Actions\nThe Archetypes we’re looking for:\nThe Data Alchemist — you believe data is only valuable when an AI can reason over it, and you spend time experimenting with embedding models and retrieval techniques to make that true\nThe Manual Mapping Hater — if you have to map two schemas twice, you've already built an agent to do it for you\nRigor over Hype — you know the difference between a vector search demo and a production-grade financial data engine; you care about Precision and Recall\nThe Founding Mindset — you're energized by building from scratch, not managing existing systems, and you make decisions confidently without a playbook\nWhy this role:\nERP data is the hardest data problem in enterprise software — 20 years of relational financial history, undocumented schemas, and zero tolerance for hallucination. If you can solve RAG for an ERP, you have solved the hardest version of the problem.\nGreenfield with real stakes: you are not inheriting someone else's technical debt or org structure. You build what you believe will win.\nDirect line to the CTO — no data governance committee, no analytics manager layer, no 6-month roadmap approval process\nUnlimited context budget: access to frontier models and the compute to run serious embedding and indexing experiments\nThe work matters: every AI feature in the product runs on the infrastructure you build","datePosted":"2026-07-27T13:03:42.869Z","dateModified":"2026-07-27T13:03:42.869Z","hiringOrganization":{"@type":"Organization","name":"Eci Software Solutions","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"6a52dcb6737cac00c158468c"},"url":"https://jobsearcher.com/jobs/6a52dcb6737cac00c158468c"}}