{"schemaVersion":"jobsearcher.job.v1","id":"3b50e915ca9595fc771576e8","url":"https://jobsearcher.com/jobs/3b50e915ca9595fc771576e8","canonicalUrl":"https://jobsearcher.com/jobs/3b50e915ca9595fc771576e8","title":"Lead Data Engineer – Physical AI Platform, Data Engineering","description":"Career Area:\nTechnology, Digital and Data\nJob Description:\nYour Work Shapes the World at Caterpillar Inc.\nWhen you join Caterpillar, you're joining a global team who cares not just about the work we do – but also about each other. We are the makers, problem solvers, and future world builders who are creating stronger, more sustainable communities. We don't just talk about progress and innovation here – we make it happen, with our customers, where we work and live. Together, we are building a better world, so we can all enjoy living in it.\nHelp Build the Future of Caterpillar –\nAt Caterpillar, technology always has a purpose, which is to solve our customers’ toughest challenges. Through Cat Technology, we are solving problems by building the intelligence layer that connects machines, data, and people to make jobsites safer, more productive, and more sustainable. By combining deep domain expertise in physical systems with software, connectivity, autonomy, and AI, we deliver solutions that work in the real world—on real jobsites, at global scale.\nYou’ll build and deploy against one of the most unique data foundations—over 1.6 million connected assets generating real-world data daily. These data and platform capabilities are enabling the development of AI models, edge computing architectures, and software systems that scale across fleets, products, and industries. The result will be a new generation of machines that continuously learn, improve, and deliver performance at scale.\nBe Part of What’s Next in Autonomous Construction Sites\nConstruction autonomy is one of the most complex challenges in applied AI, and at Caterpillar, advancements in physical AI, simulation, sensing, and edge computing are turning things that once felt impossible—intelligent machines operating in dynamic jobsites—into reality.\nOur connected ecosystem brings together massive volumes of high-quality data to create a foundation where engineers like you can build and deploy against.\nIf this work motivates you, we invite you to join our team. In these roles, you’ll work at the intersection of the physical and digital worlds. You’ll help design and deliver intelligent systems that enable machines to perceive their environment, make informed decisions, and support safer, more productive operations.\nApply today to build the new era of construction autonomy at Caterpillar.\nJob Summary\nAs a Lead Data Engineer, you will design, build, and maintain scalable data pipelines, microservices, and cloud-based data platforms that deliver reliable, high-quality data for business and engineering teams. Working in an agile environment, you will help drive data architecture, performance, reliability, and continuous improvement across modern data solutions.\nWhat You Will Do:\nActively collaborate with Principal Software Engineers and Data Architects to define solution architecture\nLead the solution design and optimization of scalable data pipelines and microservices in Python, enabling both real-time and batch data processing across enterprise platforms\nDrive the development of cloud-native data ingestion and streaming solutions leveraging AWS services including Kinesis, S3, DynamoDB, EventBridge, and related technologies\nOwn the design, implementation, and operational excellence of data integration frameworks and source data pipelines supporting CI Autonomy initiatives\nPartner with business, product, and engineering stakeholders to translate complex requirements into scalable data architectures, workflows, mappings, and system designs\nEstablish and enforce automated testing, data quality controls, and validation frameworks to ensure integrity, reliability, and compliance across distributed data ecosystems\nLead operational monitoring, performance tuning, and root-cause analysis of production data platforms using observability tools such as CloudWatch to maintain high availability and service reliability\nWhat You Will Have:\nDecision Making and Critical Thinking: Ability to lead the analysis and resolution of complex issues within distributed data platforms, designing scalable, and resilient solutions\nEffective Communications: Ability to communicate across teams by sharing feedback constructively, listening to others, and creating documentation that makes data systems and processes easy to understand and support\nSoftware Development: Experience in leading the design and development of backend systems and data pipelines using Python, Java, and modern frameworks, providing technical directions and ensuring the delivery of reliable, scalable solutions\nSoftware Development Life Cycle: Experience leading the delivery of data engineering solutions in an Agile environment by guiding work through the full development lifecycle, translating requirements into technical solutions, and ensuring projects are delivered with quality, reliability, and business value\nSoftware Integration Engineering: Capability to lead the design and integration of APIs, data pipelines, streaming platforms, and databases to enable reliable data exchange across enterprise systems and partner platforms\nSoftware Product Design/Architecture: Expertise leading the design of scalable, event-driven data systems and architectures, guiding technical decisions and ensuring solutions are reliable, maintainable, and aligned with business needs.\nSoftware Product Technical Knowledge: Ability to apply strong knowledge of AWS services and data engineering tools to define requirements, support testing and deployment activities, troubleshoot issues, and ensure data solutions are configured, implemented, and operated effectively across environments\nSoftware Product Testing: Ability to define and implement testing strategies, including functional, performance, and data quality testing, to ensure reliable, scalable, and high-performing data solutions across the development lifecycle.\nTop Candidates Will Have:\nBachelor’s degree in Computer Science, Computer Engineering, or related field\n8+ years of experience in data engineering or related disciplines with increasing responsibility\nExtensive experience on modern, large scale, complex Caterpillar data platforms such as Helios Data Platform\nStrong foundation developing and deploying Python solutions to a production environment\nExperience leading teams to build high-throughput, scalable data pipelines\nStrong hands-on experience with AWS data services (Kinesis, S3, DynamoDB, EventBridge, etc.) at scale\nStrong in SQL, including data quality and validation practices\nExperience in deploying software using CI/CD tools such as Azure DevOps, Jira, Jenkins, etc.\nExperience developing microservices that support real-time data ingestion\nExperience developing software applications using relational and noSQL databases\nAbility to ensure data integrity across distributed and streaming systems\nExperience with monitoring, testing, and automation in large-scale data environments\nSummary Pay Range:\n$128,470.00 - $208,770.00\nCompensation and benefits offered may vary depending on multiple individualized factors, job level, market location, job-related knowledge, skills, individual performance and experience. Please note that salary is only one component of total compensation at Caterpillar.\nBenefits:\nSubject to plan eligibility, terms, and guidelines. This is a summary list of benefits.\nMedical, dental, and vision benefits*\nPaid time off plan (Vacation, Holidays, Volunteer, etc.)*\n401(k) savings plans*\nHealth Savings Account (HSA)*\nFlexible Spending Accounts (FSAs)*\nHealth Lifestyle Programs*\nEmployee Assistance Program*\nVoluntary Benefits and Employee Discounts*\nCareer Development*\nIncentive bonus*\nDisability benefits\nLife Insurance\nParental leave\nAdoption benefits\nTuition Reimbursement\nThese benefits also apply to part-time employees\nThis position requires working onsite five days a week.\nRelocation is available for this position.\nVisa sponsorship is available for eligible applicants.\nPosting Dates:\nAny offer of employment is conditioned upon the successful completion of a drug screen.\nCaterpillar is an Equal Opportunity Employer, Including Veterans and Individuals with Disabilities. Qualified applicants of any age are encouraged to apply.\nNot ready to apply? Join our Talent Community.","company":"Caterpillar","rawCompany":"caterpillar","city":"Arlington","state":"TX","isRemote":false,"isActive":false,"createdAt":"2026-08-03T11:17:08.473Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"}],"industries":[{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Lead Data Engineer – Physical AI Platform, Data Engineering","description":"Career Area:\nTechnology, Digital and Data\nJob Description:\nYour Work Shapes the World at Caterpillar Inc.\nWhen you join Caterpillar, you're joining a global team who cares not just about the work we do – but also about each other. We are the makers, problem solvers, and future world builders who are creating stronger, more sustainable communities. We don't just talk about progress and innovation here – we make it happen, with our customers, where we work and live. Together, we are building a better world, so we can all enjoy living in it.\nHelp Build the Future of Caterpillar –\nAt Caterpillar, technology always has a purpose, which is to solve our customers’ toughest challenges. Through Cat Technology, we are solving problems by building the intelligence layer that connects machines, data, and people to make jobsites safer, more productive, and more sustainable. By combining deep domain expertise in physical systems with software, connectivity, autonomy, and AI, we deliver solutions that work in the real world—on real jobsites, at global scale.\nYou’ll build and deploy against one of the most unique data foundations—over 1.6 million connected assets generating real-world data daily. These data and platform capabilities are enabling the development of AI models, edge computing architectures, and software systems that scale across fleets, products, and industries. The result will be a new generation of machines that continuously learn, improve, and deliver performance at scale.\nBe Part of What’s Next in Autonomous Construction Sites\nConstruction autonomy is one of the most complex challenges in applied AI, and at Caterpillar, advancements in physical AI, simulation, sensing, and edge computing are turning things that once felt impossible—intelligent machines operating in dynamic jobsites—into reality.\nOur connected ecosystem brings together massive volumes of high-quality data to create a foundation where engineers like you can build and deploy against.\nIf this work motivates you, we invite you to join our team. In these roles, you’ll work at the intersection of the physical and digital worlds. You’ll help design and deliver intelligent systems that enable machines to perceive their environment, make informed decisions, and support safer, more productive operations.\nApply today to build the new era of construction autonomy at Caterpillar.\nJob Summary\nAs a Lead Data Engineer, you will design, build, and maintain scalable data pipelines, microservices, and cloud-based data platforms that deliver reliable, high-quality data for business and engineering teams. Working in an agile environment, you will help drive data architecture, performance, reliability, and continuous improvement across modern data solutions.\nWhat You Will Do:\nActively collaborate with Principal Software Engineers and Data Architects to define solution architecture\nLead the solution design and optimization of scalable data pipelines and microservices in Python, enabling both real-time and batch data processing across enterprise platforms\nDrive the development of cloud-native data ingestion and streaming solutions leveraging AWS services including Kinesis, S3, DynamoDB, EventBridge, and related technologies\nOwn the design, implementation, and operational excellence of data integration frameworks and source data pipelines supporting CI Autonomy initiatives\nPartner with business, product, and engineering stakeholders to translate complex requirements into scalable data architectures, workflows, mappings, and system designs\nEstablish and enforce automated testing, data quality controls, and validation frameworks to ensure integrity, reliability, and compliance across distributed data ecosystems\nLead operational monitoring, performance tuning, and root-cause analysis of production data platforms using observability tools such as CloudWatch to maintain high availability and service reliability\nWhat You Will Have:\nDecision Making and Critical Thinking: Ability to lead the analysis and resolution of complex issues within distributed data platforms, designing scalable, and resilient solutions\nEffective Communications: Ability to communicate across teams by sharing feedback constructively, listening to others, and creating documentation that makes data systems and processes easy to understand and support\nSoftware Development: Experience in leading the design and development of backend systems and data pipelines using Python, Java, and modern frameworks, providing technical directions and ensuring the delivery of reliable, scalable solutions\nSoftware Development Life Cycle: Experience leading the delivery of data engineering solutions in an Agile environment by guiding work through the full development lifecycle, translating requirements into technical solutions, and ensuring projects are delivered with quality, reliability, and business value\nSoftware Integration Engineering: Capability to lead the design and integration of APIs, data pipelines, streaming platforms, and databases to enable reliable data exchange across enterprise systems and partner platforms\nSoftware Product Design/Architecture: Expertise leading the design of scalable, event-driven data systems and architectures, guiding technical decisions and ensuring solutions are reliable, maintainable, and aligned with business needs.\nSoftware Product Technical Knowledge: Ability to apply strong knowledge of AWS services and data engineering tools to define requirements, support testing and deployment activities, troubleshoot issues, and ensure data solutions are configured, implemented, and operated effectively across environments\nSoftware Product Testing: Ability to define and implement testing strategies, including functional, performance, and data quality testing, to ensure reliable, scalable, and high-performing data solutions across the development lifecycle.\nTop Candidates Will Have:\nBachelor’s degree in Computer Science, Computer Engineering, or related field\n8+ years of experience in data engineering or related disciplines with increasing responsibility\nExtensive experience on modern, large scale, complex Caterpillar data platforms such as Helios Data Platform\nStrong foundation developing and deploying Python solutions to a production environment\nExperience leading teams to build high-throughput, scalable data pipelines\nStrong hands-on experience with AWS data services (Kinesis, S3, DynamoDB, EventBridge, etc.) at scale\nStrong in SQL, including data quality and validation practices\nExperience in deploying software using CI/CD tools such as Azure DevOps, Jira, Jenkins, etc.\nExperience developing microservices that support real-time data ingestion\nExperience developing software applications using relational and noSQL databases\nAbility to ensure data integrity across distributed and streaming systems\nExperience with monitoring, testing, and automation in large-scale data environments\nSummary Pay Range:\n$128,470.00 - $208,770.00\nCompensation and benefits offered may vary depending on multiple individualized factors, job level, market location, job-related knowledge, skills, individual performance and experience. Please note that salary is only one component of total compensation at Caterpillar.\nBenefits:\nSubject to plan eligibility, terms, and guidelines. This is a summary list of benefits.\nMedical, dental, and vision benefits*\nPaid time off plan (Vacation, Holidays, Volunteer, etc.)*\n401(k) savings plans*\nHealth Savings Account (HSA)*\nFlexible Spending Accounts (FSAs)*\nHealth Lifestyle Programs*\nEmployee Assistance Program*\nVoluntary Benefits and Employee Discounts*\nCareer Development*\nIncentive bonus*\nDisability benefits\nLife Insurance\nParental leave\nAdoption benefits\nTuition Reimbursement\nThese benefits also apply to part-time employees\nThis position requires working onsite five days a week.\nRelocation is available for this position.\nVisa sponsorship is available for eligible applicants.\nPosting Dates:\nAny offer of employment is conditioned upon the successful completion of a drug screen.\nCaterpillar is an Equal Opportunity Employer, Including Veterans and Individuals with Disabilities. Qualified applicants of any age are encouraged to apply.\nNot ready to apply? Join our Talent Community.","datePosted":"2026-08-03T11:17:08.473Z","dateModified":"2026-08-03T11:17:08.473Z","hiringOrganization":{"@type":"Organization","name":"Caterpillar","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Arlington","addressRegion":"TX","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"3b50e915ca9595fc771576e8"},"url":"https://jobsearcher.com/jobs/3b50e915ca9595fc771576e8"}}