{"schemaVersion":"jobsearcher.job.v1","id":"f1a322ec101f7a00a686fb95","url":"https://jobsearcher.com/jobs/f1a322ec101f7a00a686fb95","canonicalUrl":"https://jobsearcher.com/jobs/f1a322ec101f7a00a686fb95","title":"Senior Data Architect (Hands on)","description":"Stratus, deriving from the Latin term meaning 'layer', offers an advanced set of MEP specific solutions that seamlessly layer across a contractor's entire workflow from design to fabrication to installation. Our team of seasoned industry experts, skilled technology leaders, innovators, and entrepreneurs understands that fabrication does not occur in isolation, and increasingly, it may not happen within your own fabrication shop. Through close relationships with our customers—who include some of the most innovative and largest MEP contractors—we have developed a suite of Stratus tools to digitize, automate, and optimize piping, plumbing, sheet metal, and electrical contracting. Stratus provides the software layer an MEP Contractor needs to optimize profits with true \"Data Driven Contracting.\"\n\nGENERAL DESCRIPTION\nThe Senior Data Architect owns our canonical data architecture — the schema, contracts, tenancy, and governance that every product and every AI/ML workload builds on. You are the single owner of the canonical data model: one normalized definition of the core business objects shared across our products, and the standard the rest of engineering builds against. This is a foundational, hands-on role — you design, prototype, and ship reference implementations and in-repo guardrails, not just diagrams.\nOur approach to AI is to build durable, domain-specific data assets rather than commodity model infrastructure: we don't pretrain foundation models and we don't ship thin wrappers around someone else's. The differentiated value lives in how our data is modeled, governed, and made trustworthy for AI — and that is the layer you own.\nKEY RESPONSIBILITIES\nAI/ML readiness\nArchitect the data layer so AI/ML workloads — vector search, embeddings pipelines, RAG-grounded retrieval, model training — run on a clean, governed substrate.\nMake production data AI-ready: well-modeled, contract-enforced, lineage-tracked, and drift-detectable.\nDesign the data-side integration patterns these workloads depend on, such as feature-store and vector-store patterns across document, relational, and embedding data.\nData architecture\nOwn the canonical data model — the normalized definition of the core business objects shared across our products — and decide what is canonical versus tenant-specific.\nEstablish data architecture standards, data contracts, and schema discipline the rest of engineering builds against, enforced in-repo.\nExercise strong polyglot-persistence judgment: what belongs in document vs. relational vs. vector stores, and how to migrate between them without big-bang rewrites.\nDefine the multi-tenant data architecture: tenancy isolation, data residency posture, and per-tenant cost attribution across storage and compute.\nModernization\nLead staged modernization toward the right mix of stores and patterns for transactional, analytical, and AI/ML use cases — improving scalability, governance, and usability while minimizing disruption.\nOwn the architectural direction of the data pipeline and lake / lakehouse layer: ingestion, transformation, orchestration, and storage tiers.\nLead the move from homegrown pipelines to proven, industry-standard platforms, balancing build-vs-buy and total cost of ownership.\nModernize legacy data-access patterns via incremental, strangler-fig migrations that keep production stable.\nTechnical leadership\nDrive hands-on prototypes, reference implementations, and in-repo guardrails.\nDefine the data, storage, and retrieval patterns the rest of engineering builds against.\nEstablish data quality, testing, lineage, and observability standards for pipelines and AI/ML serving.\nMentor engineers on schema discipline, modern data practices, and AI/ML-readiness patterns.\nMake canonical decisions that are time-boxed, written, and defensible; hold disagree-and-commit rather than letting schema debate become a standing committee.\nUse AI-assisted development tools (Claude Code, Copilot, Cursor) as a force multiplier for schema design, query tuning, and migration scripting.\nCross-team partnership\nPartner with database engineering on production data health while owning long-term architectural direction.\nPartner with ML and application engineering on their data needs — structuring and governing data so it is retrieval-ready and safe to build on.\nPartner with platform / infrastructure on reliability, disaster recovery, residency, and the multi-tenant operational posture.\nQUALIFICATIONS\n8+ years in data architecture, data engineering, database administration, or analytics engineering, with 3+ years in senior / lead roles.\nDemonstrated ownership of a canonical or enterprise data model / cross-product schema — the model and contracts other teams built against.\nHands-on MongoDB at production scale (Atlas M40+ ideal): document modeling, aggregation framework, indexing, change streams, sharding, replica sets — and the judgment to recognize the Mongo-as-RDBMS anti-pattern.\nStrong polyglot-persistence judgment: deciding what belongs in documents vs. relational vs. a vector store, and migrating between them incrementally.\nHands-on relational depth: schema design, indexing strategy, and query tuning, plus familiarity with vector search (Atlas Vector Search, pgvector, or equivalent).\nProduction experience making data AI/ML-ready: data architecture supporting RAG, semantic search, embeddings / vector pipelines, or agentic workloads.\nMulti-tenant architecture experience: data residency and per-tenant cost attribution.\nPipeline / ELT / lake / lakehouse design at scale, with incremental migration strategies that minimize disruption.\nCloud-native data services (Azure, AWS, or GCP).\nStrong grasp of data quality, testing, lineage, and monitoring — including observability for pipelines and AI/ML serving.\nComfortable modeling a complex, specialized domain. MEP / AEC / construction experience is a plus; appetite to learn the domain is required.\nNICE TO HAVE\nKnowledge-graph, ontology, or semantic-layer experience.\nCDC and cross-engine sync (MongoDB Change Streams, Debezium, or equivalent).\nLakehouse platforms (Databricks, Snowflake, or open table formats — Iceberg, Delta, Hudi) and feature stores (Feast or equivalent).\nData governance for AI/agent access to production data: query-cost controls, read-path safety, lineage, and audit for higher-risk use cases.\nSOC 2 and data-classification experience.\nAzure data ecosystem (Data Factory, Synapse, Functions, Event Grid).\nMongoDB certification (Associate DBA / Developer or higher) or substantive MongoDB University coursework.\nWHAT SUCCESS LOOKS LIKE — FIRST YEAR\nThe canonical data model is owned and enforced: teams build against stable, documented contracts instead of bespoke forks.\nWorkloads sit in the right stores, legacy anti-patterns are receding, and reliability targets are holding.\nTenancy is formalized and per-tenant cost attribution is instrumented, so cost and capacity are observable as we scale.\nThe data substrate is AI-ready — model, contracts, and lineage in place — so AI/ML work builds on a solid foundation rather than waiting on data.\nYou've done it in partnership: the data tier is healthier, and engineers build against your contracts.\nBENEFITS\nComprehensive and competitive health benefits plan\nMatching 401k contributions\n20 days annual PTO\nPrimarily remote work with occasional annual team onsites\n\nThis is a fully remote position open to candidates based in the United States.","company":"Stratus","rawCompany":"stratus","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-03T13:32:52.660Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior Data Architect (Hands on)","description":"Stratus, deriving from the Latin term meaning 'layer', offers an advanced set of MEP specific solutions that seamlessly layer across a contractor's entire workflow from design to fabrication to installation. Our team of seasoned industry experts, skilled technology leaders, innovators, and entrepreneurs understands that fabrication does not occur in isolation, and increasingly, it may not happen within your own fabrication shop. Through close relationships with our customers—who include some of the most innovative and largest MEP contractors—we have developed a suite of Stratus tools to digitize, automate, and optimize piping, plumbing, sheet metal, and electrical contracting. Stratus provides the software layer an MEP Contractor needs to optimize profits with true \"Data Driven Contracting.\"\n\nGENERAL DESCRIPTION\nThe Senior Data Architect owns our canonical data architecture — the schema, contracts, tenancy, and governance that every product and every AI/ML workload builds on. You are the single owner of the canonical data model: one normalized definition of the core business objects shared across our products, and the standard the rest of engineering builds against. This is a foundational, hands-on role — you design, prototype, and ship reference implementations and in-repo guardrails, not just diagrams.\nOur approach to AI is to build durable, domain-specific data assets rather than commodity model infrastructure: we don't pretrain foundation models and we don't ship thin wrappers around someone else's. The differentiated value lives in how our data is modeled, governed, and made trustworthy for AI — and that is the layer you own.\nKEY RESPONSIBILITIES\nAI/ML readiness\nArchitect the data layer so AI/ML workloads — vector search, embeddings pipelines, RAG-grounded retrieval, model training — run on a clean, governed substrate.\nMake production data AI-ready: well-modeled, contract-enforced, lineage-tracked, and drift-detectable.\nDesign the data-side integration patterns these workloads depend on, such as feature-store and vector-store patterns across document, relational, and embedding data.\nData architecture\nOwn the canonical data model — the normalized definition of the core business objects shared across our products — and decide what is canonical versus tenant-specific.\nEstablish data architecture standards, data contracts, and schema discipline the rest of engineering builds against, enforced in-repo.\nExercise strong polyglot-persistence judgment: what belongs in document vs. relational vs. vector stores, and how to migrate between them without big-bang rewrites.\nDefine the multi-tenant data architecture: tenancy isolation, data residency posture, and per-tenant cost attribution across storage and compute.\nModernization\nLead staged modernization toward the right mix of stores and patterns for transactional, analytical, and AI/ML use cases — improving scalability, governance, and usability while minimizing disruption.\nOwn the architectural direction of the data pipeline and lake / lakehouse layer: ingestion, transformation, orchestration, and storage tiers.\nLead the move from homegrown pipelines to proven, industry-standard platforms, balancing build-vs-buy and total cost of ownership.\nModernize legacy data-access patterns via incremental, strangler-fig migrations that keep production stable.\nTechnical leadership\nDrive hands-on prototypes, reference implementations, and in-repo guardrails.\nDefine the data, storage, and retrieval patterns the rest of engineering builds against.\nEstablish data quality, testing, lineage, and observability standards for pipelines and AI/ML serving.\nMentor engineers on schema discipline, modern data practices, and AI/ML-readiness patterns.\nMake canonical decisions that are time-boxed, written, and defensible; hold disagree-and-commit rather than letting schema debate become a standing committee.\nUse AI-assisted development tools (Claude Code, Copilot, Cursor) as a force multiplier for schema design, query tuning, and migration scripting.\nCross-team partnership\nPartner with database engineering on production data health while owning long-term architectural direction.\nPartner with ML and application engineering on their data needs — structuring and governing data so it is retrieval-ready and safe to build on.\nPartner with platform / infrastructure on reliability, disaster recovery, residency, and the multi-tenant operational posture.\nQUALIFICATIONS\n8+ years in data architecture, data engineering, database administration, or analytics engineering, with 3+ years in senior / lead roles.\nDemonstrated ownership of a canonical or enterprise data model / cross-product schema — the model and contracts other teams built against.\nHands-on MongoDB at production scale (Atlas M40+ ideal): document modeling, aggregation framework, indexing, change streams, sharding, replica sets — and the judgment to recognize the Mongo-as-RDBMS anti-pattern.\nStrong polyglot-persistence judgment: deciding what belongs in documents vs. relational vs. a vector store, and migrating between them incrementally.\nHands-on relational depth: schema design, indexing strategy, and query tuning, plus familiarity with vector search (Atlas Vector Search, pgvector, or equivalent).\nProduction experience making data AI/ML-ready: data architecture supporting RAG, semantic search, embeddings / vector pipelines, or agentic workloads.\nMulti-tenant architecture experience: data residency and per-tenant cost attribution.\nPipeline / ELT / lake / lakehouse design at scale, with incremental migration strategies that minimize disruption.\nCloud-native data services (Azure, AWS, or GCP).\nStrong grasp of data quality, testing, lineage, and monitoring — including observability for pipelines and AI/ML serving.\nComfortable modeling a complex, specialized domain. MEP / AEC / construction experience is a plus; appetite to learn the domain is required.\nNICE TO HAVE\nKnowledge-graph, ontology, or semantic-layer experience.\nCDC and cross-engine sync (MongoDB Change Streams, Debezium, or equivalent).\nLakehouse platforms (Databricks, Snowflake, or open table formats — Iceberg, Delta, Hudi) and feature stores (Feast or equivalent).\nData governance for AI/agent access to production data: query-cost controls, read-path safety, lineage, and audit for higher-risk use cases.\nSOC 2 and data-classification experience.\nAzure data ecosystem (Data Factory, Synapse, Functions, Event Grid).\nMongoDB certification (Associate DBA / Developer or higher) or substantive MongoDB University coursework.\nWHAT SUCCESS LOOKS LIKE — FIRST YEAR\nThe canonical data model is owned and enforced: teams build against stable, documented contracts instead of bespoke forks.\nWorkloads sit in the right stores, legacy anti-patterns are receding, and reliability targets are holding.\nTenancy is formalized and per-tenant cost attribution is instrumented, so cost and capacity are observable as we scale.\nThe data substrate is AI-ready — model, contracts, and lineage in place — so AI/ML work builds on a solid foundation rather than waiting on data.\nYou've done it in partnership: the data tier is healthier, and engineers build against your contracts.\nBENEFITS\nComprehensive and competitive health benefits plan\nMatching 401k contributions\n20 days annual PTO\nPrimarily remote work with occasional annual team onsites\n\nThis is a fully remote position open to candidates based in the United States.","datePosted":"2026-08-03T13:32:52.660Z","dateModified":"2026-08-03T13:32:52.660Z","hiringOrganization":{"@type":"Organization","name":"Stratus","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"f1a322ec101f7a00a686fb95"},"url":"https://jobsearcher.com/jobs/f1a322ec101f7a00a686fb95"}}