{"schemaVersion":"jobsearcher.job.v1","id":"f262b949246724b3ab942a81","url":"https://jobsearcher.com/jobs/f262b949246724b3ab942a81","canonicalUrl":"https://jobsearcher.com/jobs/f262b949246724b3ab942a81","title":"Software Engineer, Data","description":"Overview\nIn this role you will build data foundations that enable real-time AI features and scalable multimedia workflows. You’ll work closely with ML engineers to design robust pipelines and data lakehouse infrastructure, powering cutting-edge video generation. The position blends backend data engineering with productizing data for front-end and AI agents, supporting high-velocity product development. You will help scale data systems to deliver low-latency, engaging user experiences and new features like PPT-to-video automation and interactive avatars.\n\nCompensation / Benefitscompetitive salaryequity401khealth benefitsgenerous PTOparental leave\nResponsibilitiesDesign, develop, and maintain batch and real-time data pipelines (Python, Go, Spark, Kafka) for large multi-modal data (text, audio, video) to train and run AI modelsCollaborate with ML engineers to implement data structures and APIs for new features requiring low-latency data accessArchitect and manage data lakehouse solutions (Snowflake, Databricks, Apache Iceberg) for efficient storage and querying of unstructured media dataImplement data quality checks, contracts, and monitoring to ensure data reliability and minimize downtime in production video generationTransform raw data into structured, actionable data products for front-end apps, API endpoints, and AI agents\nKey requirementsBachelor’s/Master’s degree in Computer Science, Engineering, or a related field3-5+ years of experience as a Backend Software Engineer with heavy data processing responsibilitiesStrong proficiency in Python (ETL/scripting) and SQL (data modeling)Experience with cloud platforms (AWS/GCP) and data technologies like Kafka, Spark, Snowflake/DatabricksExperience or interest in Computer Vision/Generative AI data processingProactive, owner mindset; ability to operate in a fast-paced, startup environmentOwnership mindsetAdaptability in a fast-paced startup environmentCollaborative mindsetPythonSQLAWS/GCP","company":"Heygen","rawCompany":"heygen","city":"Long Beach","state":"CA","isRemote":false,"isActive":false,"createdAt":"2026-09-15T03:26:52.006Z","occupations":[{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Software Engineer, Data","description":"Overview\nIn this role you will build data foundations that enable real-time AI features and scalable multimedia workflows. You’ll work closely with ML engineers to design robust pipelines and data lakehouse infrastructure, powering cutting-edge video generation. The position blends backend data engineering with productizing data for front-end and AI agents, supporting high-velocity product development. You will help scale data systems to deliver low-latency, engaging user experiences and new features like PPT-to-video automation and interactive avatars.\n\nCompensation / Benefitscompetitive salaryequity401khealth benefitsgenerous PTOparental leave\nResponsibilitiesDesign, develop, and maintain batch and real-time data pipelines (Python, Go, Spark, Kafka) for large multi-modal data (text, audio, video) to train and run AI modelsCollaborate with ML engineers to implement data structures and APIs for new features requiring low-latency data accessArchitect and manage data lakehouse solutions (Snowflake, Databricks, Apache Iceberg) for efficient storage and querying of unstructured media dataImplement data quality checks, contracts, and monitoring to ensure data reliability and minimize downtime in production video generationTransform raw data into structured, actionable data products for front-end apps, API endpoints, and AI agents\nKey requirementsBachelor’s/Master’s degree in Computer Science, Engineering, or a related field3-5+ years of experience as a Backend Software Engineer with heavy data processing responsibilitiesStrong proficiency in Python (ETL/scripting) and SQL (data modeling)Experience with cloud platforms (AWS/GCP) and data technologies like Kafka, Spark, Snowflake/DatabricksExperience or interest in Computer Vision/Generative AI data processingProactive, owner mindset; ability to operate in a fast-paced, startup environmentOwnership mindsetAdaptability in a fast-paced startup environmentCollaborative mindsetPythonSQLAWS/GCP","datePosted":"2026-09-15T03:26:52.006Z","dateModified":"2026-09-15T03:26:52.006Z","hiringOrganization":{"@type":"Organization","name":"Heygen","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Long Beach","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"f262b949246724b3ab942a81"},"url":"https://jobsearcher.com/jobs/f262b949246724b3ab942a81"}}