{"schemaVersion":"jobsearcher.job.v1","id":"e0aa3752dcff886dce8fbd7d","url":"https://jobsearcher.com/jobs/e0aa3752dcff886dce8fbd7d","canonicalUrl":"https://jobsearcher.com/jobs/e0aa3752dcff886dce8fbd7d","title":"Senior Data Engineer","description":"Marble is a technology company founded to revolutionize the food processing industry. Marble is seeking a full-time Senior Data Engineer who is ready for a challenge and eager to design, implement, and support automation solutions that are transforming the industry. As a part of the Marble team, you will leverage cutting-edge technologies to develop the next generation of automated solutions for food processing, enhancing resilience in the food supply chain.\r\nJob Summary\r\nAs a Senior Data Engineer at Marble, you will own the design and performance of data pipelines that power everything from real-time classification dashboards to ML training datasets to operational analytics for production facilities. You will work closely with Software, Infrastructure, and Machine Learning teams to ensure data flows efficiently through our pipelines securely, reliably, and at scale.\r\nYou will design for both high-throughput real-time ingestion and large-scale batch processing across on-prem edge nodes and AWS.\r\nResponsibilities\r\nArchitect and build scalable ETL/ELT pipelines for both batch and streaming workloads\r\nDesign real-time ingestion and transformation workflows integrating NATS JetStream and distributed microservices\r\nDevelop robust data models and ETL layers for ClickHouse, enabling high-performance analytics and ML feature extraction\r\nManage and optimize data storage across AWS S3, ClickHouse, and operational datasets generated on-prem\r\nBuild automation workflows for labeling data, CV pipeline pre-annotation, dataset generation, and versioning\r\nEnsure data quality, validation, integrity, and lineage, including automated tests and monitoring across pipelines\r\nCollaborate with ML and backend teams to deliver pipelines for training datasets and annotation tools.\r\nImplement scalable compute workloads for large dataset transformations\r\nDefine and enforce data governance best practices, including schema evolution, retention policies, and compliance requirements\r\nMonitor and improve data pipeline performance across multi-region environments\r\nMinimum Qualifications\r\nB.S. or M.S. in Computer Science, Data Engineering, or related field\r\n4+ years of experience building production-grade data pipelines or distributed systems\r\nStrong proficiency in Python and SQL\r\nProduction experience with at least one major distributed compute framework, Apache Spark, Ray, or Apache Airflow (2+ years preferred)\r\nExperience building streaming pipelines or real-time systems (Kafka, NATS, Redis Streams, or similar)\r\nDeep familiarity with AWS cloud services (S3, Lambda, IAM, EC2, Glue etc.)\r\nExperience with PostgreSQL, MongoDB, ClickHouse or other columnar/NoSQL systems\r\nStrong understanding of data modeling, partitioning, schema evolution, and performance tuning\r\nUnderstanding of data quality, lineage, orchestration, and governance\r\nAbility to design systems in hybrid environments (on-prem + cloud)\r\nExcellent communication, documentation, and teamwork skills\r\nPreferred Qualifications\r\nExperience with NATS JetStream, Kafka, or high-throughput messaging systems\r\nFamiliarity with GPU-based CV pipelines, ML datasets, or annotation workflows\r\nExperience with ClickHouse Materialized Views, Replicated Tables, or S3-backed storage\r\nExperience working in a regulated, safety-critical, or high-uptime environment\r\nExperience with Nomad, Consul, Vault, or HashiCorp ecosystem\r\nJob Type: Full-time\r\nLocation: Lincoln, NE - US, Omaha, NE - US, or Cambridge, MA - US\r\nTeam members can expect occasional travel for in-person meetings and site visits.\r\nMarble is an equal-opportunity employer. We understand the power of a diverse team, celebrate differences, and promote inclusion.\r\nJ-18808-Ljbffr","company":"Marblecom","rawCompany":"marblecom","city":"Lincoln","state":"NE","isRemote":false,"isActive":false,"createdAt":"2026-07-16T02:02:15.935Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Senior Data Engineer","description":"Marble is a technology company founded to revolutionize the food processing industry. Marble is seeking a full-time Senior Data Engineer who is ready for a challenge and eager to design, implement, and support automation solutions that are transforming the industry. As a part of the Marble team, you will leverage cutting-edge technologies to develop the next generation of automated solutions for food processing, enhancing resilience in the food supply chain.\r\nJob Summary\r\nAs a Senior Data Engineer at Marble, you will own the design and performance of data pipelines that power everything from real-time classification dashboards to ML training datasets to operational analytics for production facilities. You will work closely with Software, Infrastructure, and Machine Learning teams to ensure data flows efficiently through our pipelines securely, reliably, and at scale.\r\nYou will design for both high-throughput real-time ingestion and large-scale batch processing across on-prem edge nodes and AWS.\r\nResponsibilities\r\nArchitect and build scalable ETL/ELT pipelines for both batch and streaming workloads\r\nDesign real-time ingestion and transformation workflows integrating NATS JetStream and distributed microservices\r\nDevelop robust data models and ETL layers for ClickHouse, enabling high-performance analytics and ML feature extraction\r\nManage and optimize data storage across AWS S3, ClickHouse, and operational datasets generated on-prem\r\nBuild automation workflows for labeling data, CV pipeline pre-annotation, dataset generation, and versioning\r\nEnsure data quality, validation, integrity, and lineage, including automated tests and monitoring across pipelines\r\nCollaborate with ML and backend teams to deliver pipelines for training datasets and annotation tools.\r\nImplement scalable compute workloads for large dataset transformations\r\nDefine and enforce data governance best practices, including schema evolution, retention policies, and compliance requirements\r\nMonitor and improve data pipeline performance across multi-region environments\r\nMinimum Qualifications\r\nB.S. or M.S. in Computer Science, Data Engineering, or related field\r\n4+ years of experience building production-grade data pipelines or distributed systems\r\nStrong proficiency in Python and SQL\r\nProduction experience with at least one major distributed compute framework, Apache Spark, Ray, or Apache Airflow (2+ years preferred)\r\nExperience building streaming pipelines or real-time systems (Kafka, NATS, Redis Streams, or similar)\r\nDeep familiarity with AWS cloud services (S3, Lambda, IAM, EC2, Glue etc.)\r\nExperience with PostgreSQL, MongoDB, ClickHouse or other columnar/NoSQL systems\r\nStrong understanding of data modeling, partitioning, schema evolution, and performance tuning\r\nUnderstanding of data quality, lineage, orchestration, and governance\r\nAbility to design systems in hybrid environments (on-prem + cloud)\r\nExcellent communication, documentation, and teamwork skills\r\nPreferred Qualifications\r\nExperience with NATS JetStream, Kafka, or high-throughput messaging systems\r\nFamiliarity with GPU-based CV pipelines, ML datasets, or annotation workflows\r\nExperience with ClickHouse Materialized Views, Replicated Tables, or S3-backed storage\r\nExperience working in a regulated, safety-critical, or high-uptime environment\r\nExperience with Nomad, Consul, Vault, or HashiCorp ecosystem\r\nJob Type: Full-time\r\nLocation: Lincoln, NE - US, Omaha, NE - US, or Cambridge, MA - US\r\nTeam members can expect occasional travel for in-person meetings and site visits.\r\nMarble is an equal-opportunity employer. We understand the power of a diverse team, celebrate differences, and promote inclusion.\r\nJ-18808-Ljbffr","datePosted":"2026-07-16T02:02:15.935Z","dateModified":"2026-07-16T02:02:15.935Z","hiringOrganization":{"@type":"Organization","name":"Marblecom","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Lincoln","addressRegion":"NE","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"e0aa3752dcff886dce8fbd7d"},"url":"https://jobsearcher.com/jobs/e0aa3752dcff886dce8fbd7d"}}