Big Data Engineer
About Our ClientThe organization operates in the AI-driven data integration industry, addressing the challenge of connecting, transforming, and reasoning over complex datasets through agentic workflows. Its platform operates across cloud and on-premises environments and supports multi-tenant, production-scale use cases. The technology enables organizations to manage and leverage complex data efficiently while supporting scalable and reliable data operations across diverse deployment environments.About the OpportunityThe Big Data Engineer builds and maintains the core data layer of an AI-driven platform supporting agentic workflows. This role focuses on designing data models, implementing data services, and developing scalable, reliable data pipelines. The position also creates practical use cases and data interfaces that make complex information easier to access, understand, and use. The role contributes directly to how data is represented, accessed, and integrated into AI-driven applications.Responsibilities• Design and implement logical and physical data models for complex datasets.• Define schemas and access patterns that support multi-tenant environments and application workflows.• Balance normalization, performance, and flexibility across multiple data storage systems.• Collaborate with product and engineering teams to translate requirements into effective data designs.• Develop platform use cases to validate and extend data capabilities.• Design and build data interfaces, abstractions, and tools that improve user understanding of data.• Contribute to data glossaries, semantic layers, metadata systems, and schema discovery tools.• Define intuitive methods for users to explore and model data.• Translate complex data structures into accessible representations.• Build backend services and APIs for accessing and managing data models.• Implement reliable, maintainable, and performant data access layers.• Contribute to core architecture connecting data and application services.• Write clean, testable, production-grade code.• Design and implement data pipelines for ingestion, transformation, and validation.• Support batch and near-real-time data processing workflows.• Work with structured, semi-structured, and unstructured data systems.• Enable data flows supporting AI-driven and agent-based workflows.• Work with embeddings, context retrieval, and AI-related data representations.• Design systems that make data accessible to autonomous agents.• Implement validation, monitoring, and testing for data systems.• Ensure correctness, consistency, and observability across data pipelines and services.• Diagnose and resolve data issues in production environments.Requirements• Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.• 6+ years of professional experience in software or data engineering roles.• U.S. citizenship and ability to obtain Secret clearance.• Strong programming skills in Python or a similar backend language.• Experience designing and implementing production data models, including dimensional modeling.• Proficiency in SQL and relational databases such as PostgreSQL.• Experience building backend services or APIs that interact with data systems.• Experience designing and operating ETL/ELT data pipelines.• Familiarity with NoSQL databases and diverse data storage paradigms.• Experience working with large datasets and optimizing performance.• Experience with Docker and containerized development workflows.• Familiarity with Kubernetes environments.• Strong understanding of software engineering fundamentals, including testing and version control.Preferred Qualifications• Experience building multi-tenant data systems.• Familiarity with semantic layers, data catalogs, or data discovery systems.• Experience designing data-facing user interfaces or developer tools.• Experience with streaming systems such as Kafka.• Experience with orchestration tools such as Airflow, Dagster, or Prefect.• Experience with AI/ML data pipelines or agent-based systems.• Experience supporting on-premises or hybrid deployments.• Exposure to data governance, access control, and metadata management.• Experience with cloud platforms including AWS, Azure, or GCP.• Familiarity with vector databases and embedding-based retrieval methods.Pay Range and Compensation Package• The pay range and compensation package for this role will be determined based on the candidate’s experience, skills, qualifications, location, and other relevant factors.Equal Opportunity Statement: Our client is an equal opportunity employer. They celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, or national origin.Note:RemoteHunter is not the Employer of Record (EOR) for this role. Our purpose in this opportunity is to connect exceptional candidates with leading employers. We help job seekers worldwide discover roles that match their goals and guide them to complete their full application directly through the hiring company’s career page or ATS.