{"schemaVersion":"jobsearcher.job.v1","id":"0289f7772314291ca2e717a2","url":"https://jobsearcher.com/jobs/0289f7772314291ca2e717a2","canonicalUrl":"https://jobsearcher.com/jobs/0289f7772314291ca2e717a2","title":"Google Cloud Data Architect IAM Data Modernization","description":"RoleGoogle Cloud Data Architect – IAM Data Modernization\r\nLocationDallas, TX / Charlotte, NC / Iselin, NJ / Chandler, AZ / Ohio, Delaware (Hybrid)\r\nEligibilityMust be a US Citizen/GC only\r\nAbout PositionIdentity & Access Management (IAM) Data Modernization – migration of an on‑premises SQL data warehouse to a target‑state Data Lake on Google Cloud (GCP), enabling metrics & reporting, advanced analytics, and GenAI use cases (natural language querying, accelerated summarization, cross‑domain trend analysis) leveraging PySpark‑based processing, cloud‑native DevOps CI/CD pipelines, and containerized deployments on OpenShift (OCP) to deliver scalable, secure, and high‑performance data solutions.\r\nWhat You'll DoExperience implementing CI/CD pipelines for data and analytics workloads.\r\nFamiliarity with Git‑based source control, build automation, and deployment strategies.\r\nExperience with OpenShift Container Platform (OCP) for deploying data workloads and services.\r\nUnderstanding of containerized architecture, scaling, and environment management.\r\nProven ability to build CI/CD pipelines for data and infrastructure workloads.\r\nExperience managing secrets securely using GCP Secret Manager.\r\nOwnership of observability, SLOs, dashboards, alerts, and runbooks.\r\nProficiency in logging, monitoring, and alerting for data pipelines and platform reliability.\r\nHands‑on experience with PySpark for ETL/ELT, data transformation, and performance optimization.\r\nSolid understanding of distributed data processing concepts.\r\nStrong experience designing data platforms on Google Cloud Platform (GCP).\r\nExperience with Data Lakes, data warehousing, and large‑scale migration programs.\r\nProven experience designing and implementing data lake architectures (e.g., Bronze/Silver/Gold or layered models).\r\nStrong knowledge of Cloud Storage (GCS) design, including bucket layout, naming conventions, lifecycle policies, and access controls.\r\nExperience with Hadoop/HDFS architecture, distributed file systems, and data locality principles.\r\nHands‑on experience with columnar data formats (Parquet, Avro, ORC) and compression techniques.\r\nExpertise in partitioning strategies, backfills, and large‑scale data organization.\r\nAbility to design data models optimized for analytics and BI consumption.\r\nExperience building batch and streaming ingestion pipelines using GCP-native services.\r\nKnowledge of Pub/Sub‑based streaming architectures, event schema design, and versioning.\r\nStrong understanding of incremental ingestion and CDC patterns, including idempotency and deduplication.\r\nHands‑on experience with workflow orchestration tools (Cloud Composer / Airflow).\r\nAbility to design robust error handling, replay, and backfill mechanisms.\r\nExperience developing scalable batch and streaming pipelines using Dataflow (Apache Beam) and/or Spark (Dataproc).\r\nStrong proficiency in BigQuery SQL, including query optimization, partitioning, clustering, and cost control.\r\nHands‑on experience with Hadoop MapReduce and ecosystem tools (Hive, Pig, Sqoop).\r\nAdvanced Python programming skills for data engineering, including testing and maintainable code design.\r\nExperience managing schema evolution while minimizing downstream impact.\r\nExpertise in BigQuery performance optimization and data serving patterns.\r\nExperience building semantic layers and governed metrics for consistent analytics.\r\nFamiliarity with BI integration, access controls, and dashboard standards.\r\nUnderstanding of data exposure patterns via views, APIs, or curated datasets.\r\nExperience implementing data catalogs, metadata management, and ownership models.\r\nUnderstanding of data lineage for auditability and troubleshooting.\r\nStrong focus on data quality frameworks, including validation, freshness checks, and alerting.\r\nExperience defining and enforcing data contracts, schemas, and SLAs.\r\nGood to haveSecurity, Privacy & ComplianceHands‑on experience implementing fine‑grained access controls for BigQuery and GCS.\r\nExperience with Sprint planning and helping team technically.\r\nStrong stakeholder communication and solution‑architecture skills.\r\nExpertise You'll BringExperience: 10–14+ years in DevOps and Data Architecture, 5+ years designing on Pyspark/GCP/OCP at scale; prior on‑prem → cloud migration a must.\r\nEducation: Bachelor’s/Master’s in Computer Science, Information Systems, or equivalent experience.\r\nCertifications: Google Cloud Professional Cloud Architect/DevOps/OCP (required or within 3 months). Plus: Professional Data Engineer, Security Engineer.\r\nFlexible work from home options available.#J-18808-Ljbffr","company":"Vytwo","rawCompany":"vytwo","city":"Prosper","state":"TX","isRemote":false,"isActive":false,"createdAt":"2026-08-21T01:28:07.348Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Google Cloud Data Architect IAM Data Modernization","description":"RoleGoogle Cloud Data Architect – IAM Data Modernization\r\nLocationDallas, TX / Charlotte, NC / Iselin, NJ / Chandler, AZ / Ohio, Delaware (Hybrid)\r\nEligibilityMust be a US Citizen/GC only\r\nAbout PositionIdentity & Access Management (IAM) Data Modernization – migration of an on‑premises SQL data warehouse to a target‑state Data Lake on Google Cloud (GCP), enabling metrics & reporting, advanced analytics, and GenAI use cases (natural language querying, accelerated summarization, cross‑domain trend analysis) leveraging PySpark‑based processing, cloud‑native DevOps CI/CD pipelines, and containerized deployments on OpenShift (OCP) to deliver scalable, secure, and high‑performance data solutions.\r\nWhat You'll DoExperience implementing CI/CD pipelines for data and analytics workloads.\r\nFamiliarity with Git‑based source control, build automation, and deployment strategies.\r\nExperience with OpenShift Container Platform (OCP) for deploying data workloads and services.\r\nUnderstanding of containerized architecture, scaling, and environment management.\r\nProven ability to build CI/CD pipelines for data and infrastructure workloads.\r\nExperience managing secrets securely using GCP Secret Manager.\r\nOwnership of observability, SLOs, dashboards, alerts, and runbooks.\r\nProficiency in logging, monitoring, and alerting for data pipelines and platform reliability.\r\nHands‑on experience with PySpark for ETL/ELT, data transformation, and performance optimization.\r\nSolid understanding of distributed data processing concepts.\r\nStrong experience designing data platforms on Google Cloud Platform (GCP).\r\nExperience with Data Lakes, data warehousing, and large‑scale migration programs.\r\nProven experience designing and implementing data lake architectures (e.g., Bronze/Silver/Gold or layered models).\r\nStrong knowledge of Cloud Storage (GCS) design, including bucket layout, naming conventions, lifecycle policies, and access controls.\r\nExperience with Hadoop/HDFS architecture, distributed file systems, and data locality principles.\r\nHands‑on experience with columnar data formats (Parquet, Avro, ORC) and compression techniques.\r\nExpertise in partitioning strategies, backfills, and large‑scale data organization.\r\nAbility to design data models optimized for analytics and BI consumption.\r\nExperience building batch and streaming ingestion pipelines using GCP-native services.\r\nKnowledge of Pub/Sub‑based streaming architectures, event schema design, and versioning.\r\nStrong understanding of incremental ingestion and CDC patterns, including idempotency and deduplication.\r\nHands‑on experience with workflow orchestration tools (Cloud Composer / Airflow).\r\nAbility to design robust error handling, replay, and backfill mechanisms.\r\nExperience developing scalable batch and streaming pipelines using Dataflow (Apache Beam) and/or Spark (Dataproc).\r\nStrong proficiency in BigQuery SQL, including query optimization, partitioning, clustering, and cost control.\r\nHands‑on experience with Hadoop MapReduce and ecosystem tools (Hive, Pig, Sqoop).\r\nAdvanced Python programming skills for data engineering, including testing and maintainable code design.\r\nExperience managing schema evolution while minimizing downstream impact.\r\nExpertise in BigQuery performance optimization and data serving patterns.\r\nExperience building semantic layers and governed metrics for consistent analytics.\r\nFamiliarity with BI integration, access controls, and dashboard standards.\r\nUnderstanding of data exposure patterns via views, APIs, or curated datasets.\r\nExperience implementing data catalogs, metadata management, and ownership models.\r\nUnderstanding of data lineage for auditability and troubleshooting.\r\nStrong focus on data quality frameworks, including validation, freshness checks, and alerting.\r\nExperience defining and enforcing data contracts, schemas, and SLAs.\r\nGood to haveSecurity, Privacy & ComplianceHands‑on experience implementing fine‑grained access controls for BigQuery and GCS.\r\nExperience with Sprint planning and helping team technically.\r\nStrong stakeholder communication and solution‑architecture skills.\r\nExpertise You'll BringExperience: 10–14+ years in DevOps and Data Architecture, 5+ years designing on Pyspark/GCP/OCP at scale; prior on‑prem → cloud migration a must.\r\nEducation: Bachelor’s/Master’s in Computer Science, Information Systems, or equivalent experience.\r\nCertifications: Google Cloud Professional Cloud Architect/DevOps/OCP (required or within 3 months). Plus: Professional Data Engineer, Security Engineer.\r\nFlexible work from home options available.#J-18808-Ljbffr","datePosted":"2026-08-21T01:28:07.348Z","dateModified":"2026-08-21T01:28:07.348Z","hiringOrganization":{"@type":"Organization","name":"Vytwo","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Prosper","addressRegion":"TX","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"0289f7772314291ca2e717a2"},"url":"https://jobsearcher.com/jobs/0289f7772314291ca2e717a2"}}