{"schemaVersion":"jobsearcher.job.v1","id":"0d01c7fdfb4b1fcb401464ef","url":"https://jobsearcher.com/jobs/0d01c7fdfb4b1fcb401464ef","canonicalUrl":"https://jobsearcher.com/jobs/0d01c7fdfb4b1fcb401464ef","title":"T-Hub - Data Engineer with Java","description":"Responsibilities\r\nDesign and implement batch and streaming pipelines; select Beam/Spark/Flink based on requirements.\r\nData transformation ops for analytics and downstream services.\r\nOptimize jobs (performance, cost) and implement robust testing and observability.\r\nBuild and manage Airflow orchestration.\r\nTroubleshoot production issues, lead incident resolution, and drive root-cause fixes.\r\nContribute to team standards, templates, and reusable libraries.\r\nUsing Java and Apache Beam as the main programming language and library for the streaming pipelines.\r\nDesigns and ships a new pipeline or major refactor with measurable latency/cost improvements\r\nHands-on, close cooperation with the Data Engineering Lead, resulting in quick onboarding and successful code base understanding\r\nIntroduces or improves testing/observability patterns adopted by the team.\r\nWHAT SKILLS WILL BE APPRECIATED?\r\nStrong experience with Google Cloud Platform, including Dataflow, BigQuery performance optimization (partitioning and clustering), Pub/Sub, and Dataproc, as well as Cloud Monitoring and Logging.\r\nGood understanding of IAM, VPC fundamentals, and service accounts.\r\nDeep expertise in Apache Beam, including advanced transformations, windowing and triggering, side inputs and outputs, as well as state and timers for streaming pipelines.\r\nExperience in optimizing resource sizing and performance on Dataflow, including high-throughput and low-latency pipelines, custom IOs, and debugging at runner level.\r\nStrong experience with Apache Spark, including optimization of joins, shuffles, and partitioning, handling schema evolution, and debugging jobs using metrics.\r\nFamiliarity with Apache Flink and ability to implement streaming pipelines.\r\nKnowledge of Docker and Kubernetes for job packaging, especially in Spark and Flink environments.\r\nAdvanced SQL skills, including complex queries, performance tuning, and incremental data processing patterns.\r\nAbility to define and enforce data quality checks and acceptance criteria.\r\nHands-on experience with Airflow, including building and maintaining complex DAGs.\r\nProven ownership of production-grade data pipelines with focus on reliability and scalability.\r\nExtensive proficiency in Java at an expert level, including concurrency, immutability, and performance profiling.\r\nExperience in building reusable libraries and enforcing code quality, testing standards, and review practices.\r\nStrong skills in system design and API design, including performance-critical code.\r\nExperience in designing and architecting batch and streaming data systems, including selecting technologies and defining standards and guardrails.\r\nAbility to provide technical leadership, drive engineering strategy, and ensure scalability, security, and cost efficiency.\r\nProven track record of building platform components such as libraries, templates, governance frameworks, and data quality solutions.\r\nExperience with CI/CD for data (e.g. GitHub Actions, Cloud Build) and infrastructure-as-code tools like Terraform.\r\nMinimum of 6 years of experience in data engineering with demonstrated architectural leadership.\r\nTrack record of delivering large-scale data systems and influencing cross-team outcomes.\r\nAbility to set architectural direction, ensure compliance and security, and drive measurable platform improvements.\r\nWillingness to travel at least four times per year.\r\nOUR OFFER FOR YOU\r\nWorking at T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, we do not only provide a platform to hone your technical skills but also empower you to be a catalyst for innovation.\r\nYou'll have the opportunity to work at the forefront of modern technologies, from 5G to IoT and AI, shaping the future of connectivity.\r\nNo dress code - you can just be yourself here\r\nMedical, sport and life insurance packages at preferential terms\r\nAccess to our products and services at preferential terms\r\nEmployment contract-based cooperation\r\nKnow Talent - receive training or financial bonus for recommending new employees\r\nJ-18808-Ljbffr","company":"Magentateam","rawCompany":"magentateam","city":"Poland","state":"NY","isRemote":false,"isActive":false,"createdAt":"2026-07-04T02:34:29.149Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"T-Hub - Data Engineer with Java","description":"Responsibilities\r\nDesign and implement batch and streaming pipelines; select Beam/Spark/Flink based on requirements.\r\nData transformation ops for analytics and downstream services.\r\nOptimize jobs (performance, cost) and implement robust testing and observability.\r\nBuild and manage Airflow orchestration.\r\nTroubleshoot production issues, lead incident resolution, and drive root-cause fixes.\r\nContribute to team standards, templates, and reusable libraries.\r\nUsing Java and Apache Beam as the main programming language and library for the streaming pipelines.\r\nDesigns and ships a new pipeline or major refactor with measurable latency/cost improvements\r\nHands-on, close cooperation with the Data Engineering Lead, resulting in quick onboarding and successful code base understanding\r\nIntroduces or improves testing/observability patterns adopted by the team.\r\nWHAT SKILLS WILL BE APPRECIATED?\r\nStrong experience with Google Cloud Platform, including Dataflow, BigQuery performance optimization (partitioning and clustering), Pub/Sub, and Dataproc, as well as Cloud Monitoring and Logging.\r\nGood understanding of IAM, VPC fundamentals, and service accounts.\r\nDeep expertise in Apache Beam, including advanced transformations, windowing and triggering, side inputs and outputs, as well as state and timers for streaming pipelines.\r\nExperience in optimizing resource sizing and performance on Dataflow, including high-throughput and low-latency pipelines, custom IOs, and debugging at runner level.\r\nStrong experience with Apache Spark, including optimization of joins, shuffles, and partitioning, handling schema evolution, and debugging jobs using metrics.\r\nFamiliarity with Apache Flink and ability to implement streaming pipelines.\r\nKnowledge of Docker and Kubernetes for job packaging, especially in Spark and Flink environments.\r\nAdvanced SQL skills, including complex queries, performance tuning, and incremental data processing patterns.\r\nAbility to define and enforce data quality checks and acceptance criteria.\r\nHands-on experience with Airflow, including building and maintaining complex DAGs.\r\nProven ownership of production-grade data pipelines with focus on reliability and scalability.\r\nExtensive proficiency in Java at an expert level, including concurrency, immutability, and performance profiling.\r\nExperience in building reusable libraries and enforcing code quality, testing standards, and review practices.\r\nStrong skills in system design and API design, including performance-critical code.\r\nExperience in designing and architecting batch and streaming data systems, including selecting technologies and defining standards and guardrails.\r\nAbility to provide technical leadership, drive engineering strategy, and ensure scalability, security, and cost efficiency.\r\nProven track record of building platform components such as libraries, templates, governance frameworks, and data quality solutions.\r\nExperience with CI/CD for data (e.g. GitHub Actions, Cloud Build) and infrastructure-as-code tools like Terraform.\r\nMinimum of 6 years of experience in data engineering with demonstrated architectural leadership.\r\nTrack record of delivering large-scale data systems and influencing cross-team outcomes.\r\nAbility to set architectural direction, ensure compliance and security, and drive measurable platform improvements.\r\nWillingness to travel at least four times per year.\r\nOUR OFFER FOR YOU\r\nWorking at T Hub will offer you an unique and highly rewarding experience on IT market. As a leader in the telecommunications industry, we do not only provide a platform to hone your technical skills but also empower you to be a catalyst for innovation.\r\nYou'll have the opportunity to work at the forefront of modern technologies, from 5G to IoT and AI, shaping the future of connectivity.\r\nNo dress code - you can just be yourself here\r\nMedical, sport and life insurance packages at preferential terms\r\nAccess to our products and services at preferential terms\r\nEmployment contract-based cooperation\r\nKnow Talent - receive training or financial bonus for recommending new employees\r\nJ-18808-Ljbffr","datePosted":"2026-07-04T02:34:29.149Z","dateModified":"2026-07-04T02:34:29.149Z","hiringOrganization":{"@type":"Organization","name":"Magentateam","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Poland","addressRegion":"NY","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"0d01c7fdfb4b1fcb401464ef"},"url":"https://jobsearcher.com/jobs/0d01c7fdfb4b1fcb401464ef"}}