{"schemaVersion":"jobsearcher.job.v1","id":"eb28bc1100fff71a982d25ee","url":"https://jobsearcher.com/jobs/eb28bc1100fff71a982d25ee","canonicalUrl":"https://jobsearcher.com/jobs/eb28bc1100fff71a982d25ee","title":"Data Engineer","description":"DESCRIPTION At Apple, great ideas have a way of becoming phenomenal products, services, and customer experiences very quickly. Our team is building a massive, real-time platform that transforms continuous streams of multimodal data (including structured, image, and log data) into an intelligent, searchable foundation. By enriching this data with language and embedding models, we power critical experiences for billions of Apple customers across multiple downstream applications.\r\nMINIMUM QUALIFICATIONS Masters Degree\r\n10+ years of experience in data engineering, including building and maintaining large-scale ETL/ELT data pipelines\r\nProficiency in data modeling, especially dimensional modeling, and designing schemas optimized for analytics and reporting\r\nExperience with leveraging databases including SQL/NoSQL Databases (including Postgres / Cassandra / Redis)\r\nStrong experience with distributed data processing frameworks including Apache Spark\r\nStrong experience with Parallel processing frameworks: BigTable/Hadoop\r\nStrong software engineering fundamentals and proven experience with Scala, Java\r\nHands-on experience with Apache Kafka, Iceberg, and Flink.\r\nExperience with workflow orchestration tools including Apache Airflow and Beam\r\nExperience with AWS: e.g., S3, EMR, Lambda, Glue, Redshift, BigQuery, Kinesis, or similar services\r\nExperience with Analytics frameworks including Trino (Presto, BigQuery, Snowflake)\r\nHands-on experience with big data lake architectures\r\nExperience with containerization and orchestration (Docker, Kubernetes/EKS) and CI/CD tooling including Jenkins\r\nExperience in Python and PySpark\r\nFamiliarity with graph databases such as TigerGraph\r\nExperience building pipelines that process multimodal data (structured and image) and integrate ML model inference - including LLMs and embedding models - for data enrichment and transformation\r\nHands-on experience deploying, serving, and optimizing LLMs or ML models directly in the production, inference runtimes/compilers (ONNX Runtime, TensorRT/TensorRT-LLM), and serving frameworks (Triton, vLLM, TorchServe or similar).\r\nExperience tuning batching, KV-cache, and GPU utilization for low-latency, high-throughput real-time inference in a data pipeline\r\nKnowledge of data governance principles, data security best practices, and data privacy regulations\r\nPREFERRED QUALIFICATIONS Experience with data versioning tools and frameworks (e.g., DVC, Delta Lake)\r\nExcellent communication skills and a collaborative mindset\r\nExperience storing/serving embeddings (e.g., pgvector, Milvus, FAISS)\r\nJ-18808-Ljbffr","company":"Socket","rawCompany":"socket","city":"Cupertino","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-08-08T01:49:48.125Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Data Engineer","description":"DESCRIPTION At Apple, great ideas have a way of becoming phenomenal products, services, and customer experiences very quickly. Our team is building a massive, real-time platform that transforms continuous streams of multimodal data (including structured, image, and log data) into an intelligent, searchable foundation. By enriching this data with language and embedding models, we power critical experiences for billions of Apple customers across multiple downstream applications.\r\nMINIMUM QUALIFICATIONS Masters Degree\r\n10+ years of experience in data engineering, including building and maintaining large-scale ETL/ELT data pipelines\r\nProficiency in data modeling, especially dimensional modeling, and designing schemas optimized for analytics and reporting\r\nExperience with leveraging databases including SQL/NoSQL Databases (including Postgres / Cassandra / Redis)\r\nStrong experience with distributed data processing frameworks including Apache Spark\r\nStrong experience with Parallel processing frameworks: BigTable/Hadoop\r\nStrong software engineering fundamentals and proven experience with Scala, Java\r\nHands-on experience with Apache Kafka, Iceberg, and Flink.\r\nExperience with workflow orchestration tools including Apache Airflow and Beam\r\nExperience with AWS: e.g., S3, EMR, Lambda, Glue, Redshift, BigQuery, Kinesis, or similar services\r\nExperience with Analytics frameworks including Trino (Presto, BigQuery, Snowflake)\r\nHands-on experience with big data lake architectures\r\nExperience with containerization and orchestration (Docker, Kubernetes/EKS) and CI/CD tooling including Jenkins\r\nExperience in Python and PySpark\r\nFamiliarity with graph databases such as TigerGraph\r\nExperience building pipelines that process multimodal data (structured and image) and integrate ML model inference - including LLMs and embedding models - for data enrichment and transformation\r\nHands-on experience deploying, serving, and optimizing LLMs or ML models directly in the production, inference runtimes/compilers (ONNX Runtime, TensorRT/TensorRT-LLM), and serving frameworks (Triton, vLLM, TorchServe or similar).\r\nExperience tuning batching, KV-cache, and GPU utilization for low-latency, high-throughput real-time inference in a data pipeline\r\nKnowledge of data governance principles, data security best practices, and data privacy regulations\r\nPREFERRED QUALIFICATIONS Experience with data versioning tools and frameworks (e.g., DVC, Delta Lake)\r\nExcellent communication skills and a collaborative mindset\r\nExperience storing/serving embeddings (e.g., pgvector, Milvus, FAISS)\r\nJ-18808-Ljbffr","datePosted":"2026-08-08T01:49:48.125Z","dateModified":"2026-08-08T01:49:48.125Z","hiringOrganization":{"@type":"Organization","name":"Socket","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Cupertino","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"eb28bc1100fff71a982d25ee"},"url":"https://jobsearcher.com/jobs/eb28bc1100fff71a982d25ee"}}