{"schemaVersion":"jobsearcher.job.v1","id":"42b601d09de7fd03f1638a61","url":"https://jobsearcher.com/jobs/42b601d09de7fd03f1638a61","canonicalUrl":"https://jobsearcher.com/jobs/42b601d09de7fd03f1638a61","title":"DataOps Engineer","description":"Overview\nWe are looking for a mid‑level engineer to build and operate a data platform that uses Apache Iceberg as the lake‑house table format and Docker‑based micro‑services (Spark, Flink, Presto, etc.). you will own the end‑to‑end delivery pipeline, monitoring, security, and incident response, ensuring the platform runs reliably at scale.\nKey Responsibilities\nIceberg operations: support tables, manage schema changes, partitions, snapshot retention, and keep the catalog (Hive Metastore, AWS Glue, Nessie, …) synchronized.\nDocker image creation & testing: write multi‑stage Dockerfiles for Spark/Flink/Presto, run local test environments with Docker‑Compose, and conduct vulnerability scans (Trivy, Snyk, …).\nData pipeline development: build ETL/ELT jobs that ingest raw data and write to Iceberg tables; add simple streaming components using Kafka, Pulsar, or Kinesis when needed.\nCI/CD automation: configure pipelines (GitHub Actions, GitLab CI, Azure DevOps, …) to lint Dockerfiles, scan images, version Iceberg metadata, and deploy pipelines without downtime.\nAutomation with Ansible/Python: script cluster provisioning, catalog configuration, vacuum/compaction, and other routine housekeeping tasks.\nObservability: instrument services with OpenTelemetry, Prometheus, Grafana, and Loki; create dashboards showing pipeline latency, resource usage, table health, and error rates; set up basic alerts.\nSLA monitoring: measure data freshness, job success rates, and query response times against agreed‑upon targets and report deviations.\nIncident response: join the on‑call rotation, perform first‑line diagnosis and resolution of pipeline failures, Iceberg metadata issues, or container crashes; write concise root‑cause analyses and suggest improvements.\nSecurity & compliance support: help enforce image signing, mTLS, IAM roles, and bucket policies; collaborate with the security team to meet GDPR, HIPAA, or ISO 27001 requirements.\nKnowledge sharing: keep internal documentation up to date and run short tech demos or brown‑bag sessions on Iceberg, Docker best practices, and automation techniques.\nMinimum Requirements\nBachelor’s degree in Computer Science, IT, Data Engineering, or a related field (Master’s a plus).\n~5 years of hands‑on experience building and operating large‑scale data platforms (lake‑house, data‑warehouse, or big‑data ecosystems).\nProven production experience with Apache Iceberg (table creation, partition management, schema evolution, catalog integration).\nStrong Docker skills: multi‑stage builds, Docker‑Compose testing, routine image security scanning.\nExperience with at least one major data‑processing engine (Spark, Flink, or Presto/Trino) and its connection to Iceberg tables.\nProficiency in Python and/or Ansible for automating infrastructure and platform tasks.\nExperience building CI/CD pipelines that include Docker linting, vulnerability scanning, and automated deployment of data‑pipeline code.\nFamiliarity with observability tooling (Prometheus, Grafana, OpenTelemetry, Loki) and ability to create useful alerts and dashboards.\nAbility to respond to incidents, write clear root‑cause analysis reports, and contribute to post‑mortem actions.\nWillingness to participate in an on‑call rotation as a first‑line responder.\nAvailability to work on‑site in New Jersey for the initial assignment and relocate to Dallas by October 2026.\nPreferred Qualifications\nExperience with cloud‑native data services on AWS, Azure, or GCP (EMR, Dataproc, Synapse, etc.).\nFamiliarity with other lake‑house formats such as Delta Lake or Apache Hudi and ability to evaluate trade‑offs against Iceberg.\nKnowledge of streaming platforms (Kafka, Pulsar, Kinesis) and real‑time processing patterns.\nRelevant certifications (Databricks Lakehouse Associate, Google Professional Data Engineer, AWS Certified Data Analytics – Specialty, etc.).\nBackground supporting data platforms in regulated industries (pharma, finance, healthcare) and understanding of associated compliance frameworks.\nPay: $107,867.82 - $129,905.34 per year\nWork Location: In person","company":"Ariescomputersystemsinc","rawCompany":"ariescomputersystemsinc","city":"Plano","state":"TX","isRemote":false,"isActive":false,"createdAt":"2026-08-06T15:43:20.989Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"DataOps Engineer","description":"Overview\nWe are looking for a mid‑level engineer to build and operate a data platform that uses Apache Iceberg as the lake‑house table format and Docker‑based micro‑services (Spark, Flink, Presto, etc.). you will own the end‑to‑end delivery pipeline, monitoring, security, and incident response, ensuring the platform runs reliably at scale.\nKey Responsibilities\nIceberg operations: support tables, manage schema changes, partitions, snapshot retention, and keep the catalog (Hive Metastore, AWS Glue, Nessie, …) synchronized.\nDocker image creation & testing: write multi‑stage Dockerfiles for Spark/Flink/Presto, run local test environments with Docker‑Compose, and conduct vulnerability scans (Trivy, Snyk, …).\nData pipeline development: build ETL/ELT jobs that ingest raw data and write to Iceberg tables; add simple streaming components using Kafka, Pulsar, or Kinesis when needed.\nCI/CD automation: configure pipelines (GitHub Actions, GitLab CI, Azure DevOps, …) to lint Dockerfiles, scan images, version Iceberg metadata, and deploy pipelines without downtime.\nAutomation with Ansible/Python: script cluster provisioning, catalog configuration, vacuum/compaction, and other routine housekeeping tasks.\nObservability: instrument services with OpenTelemetry, Prometheus, Grafana, and Loki; create dashboards showing pipeline latency, resource usage, table health, and error rates; set up basic alerts.\nSLA monitoring: measure data freshness, job success rates, and query response times against agreed‑upon targets and report deviations.\nIncident response: join the on‑call rotation, perform first‑line diagnosis and resolution of pipeline failures, Iceberg metadata issues, or container crashes; write concise root‑cause analyses and suggest improvements.\nSecurity & compliance support: help enforce image signing, mTLS, IAM roles, and bucket policies; collaborate with the security team to meet GDPR, HIPAA, or ISO 27001 requirements.\nKnowledge sharing: keep internal documentation up to date and run short tech demos or brown‑bag sessions on Iceberg, Docker best practices, and automation techniques.\nMinimum Requirements\nBachelor’s degree in Computer Science, IT, Data Engineering, or a related field (Master’s a plus).\n~5 years of hands‑on experience building and operating large‑scale data platforms (lake‑house, data‑warehouse, or big‑data ecosystems).\nProven production experience with Apache Iceberg (table creation, partition management, schema evolution, catalog integration).\nStrong Docker skills: multi‑stage builds, Docker‑Compose testing, routine image security scanning.\nExperience with at least one major data‑processing engine (Spark, Flink, or Presto/Trino) and its connection to Iceberg tables.\nProficiency in Python and/or Ansible for automating infrastructure and platform tasks.\nExperience building CI/CD pipelines that include Docker linting, vulnerability scanning, and automated deployment of data‑pipeline code.\nFamiliarity with observability tooling (Prometheus, Grafana, OpenTelemetry, Loki) and ability to create useful alerts and dashboards.\nAbility to respond to incidents, write clear root‑cause analysis reports, and contribute to post‑mortem actions.\nWillingness to participate in an on‑call rotation as a first‑line responder.\nAvailability to work on‑site in New Jersey for the initial assignment and relocate to Dallas by October 2026.\nPreferred Qualifications\nExperience with cloud‑native data services on AWS, Azure, or GCP (EMR, Dataproc, Synapse, etc.).\nFamiliarity with other lake‑house formats such as Delta Lake or Apache Hudi and ability to evaluate trade‑offs against Iceberg.\nKnowledge of streaming platforms (Kafka, Pulsar, Kinesis) and real‑time processing patterns.\nRelevant certifications (Databricks Lakehouse Associate, Google Professional Data Engineer, AWS Certified Data Analytics – Specialty, etc.).\nBackground supporting data platforms in regulated industries (pharma, finance, healthcare) and understanding of associated compliance frameworks.\nPay: $107,867.82 - $129,905.34 per year\nWork Location: In person","datePosted":"2026-08-06T15:43:20.989Z","dateModified":"2026-08-06T15:43:20.989Z","hiringOrganization":{"@type":"Organization","name":"Ariescomputersystemsinc","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Plano","addressRegion":"TX","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"42b601d09de7fd03f1638a61"},"url":"https://jobsearcher.com/jobs/42b601d09de7fd03f1638a61"}}