Big Data Developer
Role: Software Engineer – Java, Big Data & Cloud InfrastructureLocation: Sunnyvale, CA & Austin, TX - OnsiteJob Type: Fulltime / W2About the RoleWe are looking for an experienced Software Engineer to design, build, and maintain scalable backend services and data-driven systems.The ideal candidate has strong expertise in Java, exposure to Big Data technologies (including Apache Iceberg), Kubernetes, and CI/CD practices, and thrives in a fast-paced, collaborative engineering environment.ResponsibilitiesDesign, develop, and maintain robust, scalable Java-based backend applications and services.Build and support data pipelines that read/write large-scale datasets, leveraging Apache Iceberg table format for efficient, reliable storage and querying.Work with Big Data ecosystems (e.g., Spark, Kafka, Hive) as needed to support platform and application features.Containerize applications and manage deployments using Kubernetes, ensuring high availability, scalability, and fault tolerance.Design and maintain CI/CD pipelines (e.g., Jenkins, GitLab CI, GitHub Actions) to automate build, test, and deployment workflows.Collaborate with cross-functional teams including data engineers, DevOps, QA, and product managers to deliver end-to-end solutions.Write clean, well-tested, maintainable code following best software engineering practices.Monitor, troubleshoot, and optimize production systems for performance and reliability.Participate in code reviews, architecture discussions, and technical design decisions.Ensure data quality, security, and compliance across services and pipelines.Required QualificationsBachelor's degree in Computer Science, Engineering, or related field (or equivalent practical experience).3+ years of professional software engineering experience.Strong proficiency in Java (Java 8+), including multithreading, collections, and design patterns.Working knowledge of Big Data frameworks/table formats such as Apache Iceberg, Spark, or Hive.Experience with ETL/data workflow concepts (building, consuming, or integrating with pipelines).Experience deploying and managing containerized applications with Kubernetes (Helm, Docker).Proven experience building and maintaining CI/CD pipelines (Jenkins, GitLab CI/CD, GitHub Actions, or similar).Solid understanding of relational and NoSQL databases (e.g., MySQL, PostgreSQL, Cassandra, MongoDB).Experience with version control systems (Git) and Agile/Scrum development practices.Strong problem-solving, debugging, and analytical skills.Preferred QualificationsExperience with cloud platforms (AWS, Azure, or GCP).Familiarity with data lakehouse architectures and table formats (Iceberg, Delta Lake, Hudi).Knowledge of infrastructure-as-code tools (Terraform, Ansible).Experience with monitoring/observability tools (Prometheus, Grafana, ELK/EFK stack).Exposure to microservices architecture and RESTful API design.Understanding of data security, governance, and compliance best practices.