{"schemaVersion":"jobsearcher.job.v1","id":"b500624c9342851842b604d5","url":"https://jobsearcher.com/jobs/b500624c9342851842b604d5","canonicalUrl":"https://jobsearcher.com/jobs/b500624c9342851842b604d5","title":"Data Engineer, Forward Deployed","description":"Applied Computing was founded in 2024 to build Orbital, a physics-informed foundation model for energy operations. We’re live across oil and gas, refineries, and petrochemicals, working towards our mission: sustainable\nabundance for a growing planet.\n\nThe hydrocarbon industry keeps the world running. But its complexity has left operators tied to legacy systems, making critical decisions on less than 10% of available data. We built Orbital to change that. It’s a foundation model built specifically for energy that lets companies use AI at scale, harnessing all of their operational\ndata and optimising in real time for any metric. Decisions get faster, operations get safer, and carbon intensity falls.\n\nWe’ve raised over $32 million, including one of the largest seed rounds for an\nAI company in the UK. We’re just getting started\n\nThe Role\n\nAs our Data Engineer, you’ll architect and maintain pipelines that make high-frequency time-series, lab, and historian data into a scalable Lakehouse architecture, usable for both deep learning models and real-time LLMs. You’ll be working across AWS (EKS, S3, EBS, KMS, CloudWatch) and Databricks/PySpark, ensuring data is contextualised, synchronised, and optimised for both deep learning models and real-time LLM workloads.\n\nThis isn’t a traditional ETL role, you’ll be solving problems at the intersection of control systems, industrial data engineering, and AI enablement.\n\nTechnical Requirements\n\nDeep expertise in PostgreSQL (partitioning, indexing, query optimisation, storage design).\n\nStrong proficiency in Python for data processing, scripting, and pipeline orchestration.\n\nHands-on experience with AWS (EKS, S3, EBS, IAM, KMS, CloudWatch, etc.)for secure and scalable data pipelines.\n\nProven ability to work with Databricks and PySpark for large-scale distributed data processing.\n\nFamiliarity with time-series industrial data (control systems, DCS/SCADA logs, process historians).\n\nExperience in unstructured data sync and management within hybrid cloud/on-prem environments.\n\nBonus: Experience working as a data engineer in oil and gas or energy environments\n\nBonus: Knowledge of streaming frameworks (Kafka, Flink, Spark Streaming) or MLOps stacks for data versioning and lineage.\n\nCore Responsibilities\n\n1. Ingest & Contextualise Data\n\nIngest from OPC UA servers, process historians, IoT sensors, LIMS systems, alarms/events, and P&IDs.\n\nMap signals to their physical processes (tags, units, hierarchies) for interpretability in AI pipelines.\n\n2. Data Movement & Accessibility\n\nBuild pipelines that handle real-time streaming and batch ingestion into the Lakehouse.\n\nManage synchronisation between historian archives, unstructured files, and AWS storage (S3/EBS).\n\nOrchestrate Databricks Lakeflow/Connectors for integrating data into Lakebase/Lakehouse.\n\nHandle secure, high-throughput transfers between historian archives and sandbox/live environments.\n\n3. Change Tracking & Integrity\n\nDetect and manage schema changes, signal drift, and inconsistencies acrosstime.\n\nImplement lineage and audit trails across Spark/Databricks and AWS pipelines.\n\n4. Data Preparation for AI\n\nBuild and maintaindual pipelines:\n\nTraining large-scale historical data prep for time-series + LLM training.\n\nInference low-latency, real-time pipelines for anomaly detection, optimisation, and LLM search.\n\nSupport heterogeneous AI workloads (time-series forecasting and retrieval-augmented LLMs).\n\n5. Database Performance & Optimisation\n\nTune PostgreSQLand sparkfor high-throughput time-series workloads (partitioning, indexing, query optimisation).\n\nOptimise pipelines for both fast analytical queries and high-efficiency model training.\n\nDeploy and manage data pipelines in AWS EKS (Kubernetes) with persisten tEBS-backed storage.\n\nWhat Success Looks Like\n\nLive data streams are contextualised,queryable, and AI-ready.\n\nSchema changes and signal drift are detected and handled without breaking downstream workflows.\n\nTraining and inference pipelines run smoothly in parallel, optimised for scale and latency.","company":"Applied Computing","rawCompany":"applied computing","city":"Houston","state":"TX","isRemote":false,"isActive":false,"createdAt":"2026-09-02T08:54:34.216Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1243.00","title":"Database Architects","slug":"database-architects"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Data Engineer, Forward Deployed","description":"Applied Computing was founded in 2024 to build Orbital, a physics-informed foundation model for energy operations. We’re live across oil and gas, refineries, and petrochemicals, working towards our mission: sustainable\nabundance for a growing planet.\n\nThe hydrocarbon industry keeps the world running. But its complexity has left operators tied to legacy systems, making critical decisions on less than 10% of available data. We built Orbital to change that. It’s a foundation model built specifically for energy that lets companies use AI at scale, harnessing all of their operational\ndata and optimising in real time for any metric. Decisions get faster, operations get safer, and carbon intensity falls.\n\nWe’ve raised over $32 million, including one of the largest seed rounds for an\nAI company in the UK. We’re just getting started\n\nThe Role\n\nAs our Data Engineer, you’ll architect and maintain pipelines that make high-frequency time-series, lab, and historian data into a scalable Lakehouse architecture, usable for both deep learning models and real-time LLMs. You’ll be working across AWS (EKS, S3, EBS, KMS, CloudWatch) and Databricks/PySpark, ensuring data is contextualised, synchronised, and optimised for both deep learning models and real-time LLM workloads.\n\nThis isn’t a traditional ETL role, you’ll be solving problems at the intersection of control systems, industrial data engineering, and AI enablement.\n\nTechnical Requirements\n\nDeep expertise in PostgreSQL (partitioning, indexing, query optimisation, storage design).\n\nStrong proficiency in Python for data processing, scripting, and pipeline orchestration.\n\nHands-on experience with AWS (EKS, S3, EBS, IAM, KMS, CloudWatch, etc.)for secure and scalable data pipelines.\n\nProven ability to work with Databricks and PySpark for large-scale distributed data processing.\n\nFamiliarity with time-series industrial data (control systems, DCS/SCADA logs, process historians).\n\nExperience in unstructured data sync and management within hybrid cloud/on-prem environments.\n\nBonus: Experience working as a data engineer in oil and gas or energy environments\n\nBonus: Knowledge of streaming frameworks (Kafka, Flink, Spark Streaming) or MLOps stacks for data versioning and lineage.\n\nCore Responsibilities\n\n1. Ingest & Contextualise Data\n\nIngest from OPC UA servers, process historians, IoT sensors, LIMS systems, alarms/events, and P&IDs.\n\nMap signals to their physical processes (tags, units, hierarchies) for interpretability in AI pipelines.\n\n2. Data Movement & Accessibility\n\nBuild pipelines that handle real-time streaming and batch ingestion into the Lakehouse.\n\nManage synchronisation between historian archives, unstructured files, and AWS storage (S3/EBS).\n\nOrchestrate Databricks Lakeflow/Connectors for integrating data into Lakebase/Lakehouse.\n\nHandle secure, high-throughput transfers between historian archives and sandbox/live environments.\n\n3. Change Tracking & Integrity\n\nDetect and manage schema changes, signal drift, and inconsistencies acrosstime.\n\nImplement lineage and audit trails across Spark/Databricks and AWS pipelines.\n\n4. Data Preparation for AI\n\nBuild and maintaindual pipelines:\n\nTraining large-scale historical data prep for time-series + LLM training.\n\nInference low-latency, real-time pipelines for anomaly detection, optimisation, and LLM search.\n\nSupport heterogeneous AI workloads (time-series forecasting and retrieval-augmented LLMs).\n\n5. Database Performance & Optimisation\n\nTune PostgreSQLand sparkfor high-throughput time-series workloads (partitioning, indexing, query optimisation).\n\nOptimise pipelines for both fast analytical queries and high-efficiency model training.\n\nDeploy and manage data pipelines in AWS EKS (Kubernetes) with persisten tEBS-backed storage.\n\nWhat Success Looks Like\n\nLive data streams are contextualised,queryable, and AI-ready.\n\nSchema changes and signal drift are detected and handled without breaking downstream workflows.\n\nTraining and inference pipelines run smoothly in parallel, optimised for scale and latency.","datePosted":"2026-09-02T08:54:34.216Z","dateModified":"2026-09-02T08:54:34.216Z","hiringOrganization":{"@type":"Organization","name":"Applied Computing","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Houston","addressRegion":"TX","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"b500624c9342851842b604d5"},"url":"https://jobsearcher.com/jobs/b500624c9342851842b604d5"}}