Data Engineer- Backend Developer
Back-End Developer – Data Engineering, Snowflake & GenBIDirect End Client: California State Teachers' Retirement System (CalSTRS)Duration: 12 Months + Possible 36-Month ExtensionLocation: West Sacramento, CAWork Model: Hybrid – On-site 2–3 business days per week at CalSTRS HeadquartersHours Per Week: 40Scope of Project: CalSTRS is seeking experienced technical resources to support the Investment Data Warehouse (IDW) platform. The project will focus on the design, development, implementation, maintenance, and support of technical solutions across the IDW platform. The team will enhance the IDW architecture and accelerate investment-related use cases involving:Generative Business Intelligence (GenBI)Artificial Intelligence (AI)Machine Learning (ML)Agentic AIAWSSnowflakeEnterprise data and analyticsThe selected resources will collectively provide full-stack development coverage across data engineering, analytics, AI/ML, application development, APIs, semantic modeling, and cloud technologies. A mandatory project objective is knowledge transfer to CalSTRS employees, including providing documentation, technical materials, and other requested content necessary for CalSTRS staff to maintain and operate the solutions.ResponsibilitiesData Engineering & Data WarehouseBuild, maintain, and optimize ETL/ELT pipelines.Design and develop cloud data warehouse solutions.Develop and optimize Snowflake data solutions.Work with Amazon Redshift and Azure Synapse or similar cloud data platforms.Integrate source systems and data feeds into the Investment Data Warehouse (IDW).Develop data warehouse objects, transformations, and integrations.Ensure data quality, reliability, and performance.SQL & Data ModelingDevelop advanced SQL queries and stored procedures.Perform advanced joins, window functions, and query performance tuning.Design and implement star and snowflake schemas.Develop enterprise data models and semantic layers.Support semantic models for business intelligence and AI-driven analytics.ETL/ELT & OrchestrationDevelop and maintain ETL/ELT pipelines.Work with orchestration tools such as Apache Airflow and dbt.Automate data processing and transformation workflows.Monitor pipeline execution, failures, and performance.Implement appropriate logging and error-handling mechanisms.GenBI, AI/ML & GenAISupport development of GenBI and AI-powered analytics solutions.Develop solutions using AI/ML and Generative AI technologies.Work with OpenAI APIs or comparable LLM platforms.Develop natural-language-to-SQL/query solutions.Design prompt strategies for business insights and analytics.Apply prompt engineering techniques to control LLM behavior.Support AI-assisted business intelligence and analytics use cases.Develop automated tests to help prevent Text-to-SQL engine hallucinations.API & Application DevelopmentDevelop APIs supporting GenBI and AI-powered solutions.Support backend services for analytics applications.Work with application teams to integrate data, AI, and analytics services.Support interactive analytics applications and dashboards.Collaborate with front-end developers working with React, Chainlit, and Streamlit.Cloud & InfrastructureWork within AWS and Snowflake ecosystems.Support AWS services including Bedrock and SageMaker.Implement infrastructure automation using Terraform and Ansible.Support containerized applications using Docker and Kubernetes.Support CI/CD pipelines and automated deployments.Monitor cloud infrastructure and application performance.Monitoring, Operations & SupportMonitor data pipelines and platform performance.Implement monitoring and logging for ML, BI, and data systems.Troubleshoot data, application, and infrastructure issues.Support ongoing maintenance and operations of the IDW platform.Optimize query performance, pipeline reliability, and platform availability.Collaboration & Knowledge TransferCollaborate with developers, data engineers, architects, analysts, and business stakeholders.Participate in requirements analysis and technical solution design.Follow applicable SDLC and Agile Development practices.Adhere to CalSTRS Minimum Information Security Requirements (MISR) and AI governance processes.Prepare technical documentation and project materials.Provide knowledge transfer to CalSTRS employees.Required/Preferred SkillsRequired Qualifications7+ years of experience in the Information Technology field with extensive experience in report writing, data analysis, or database querying.5+ years of experience in data engineering, data warehousing, or BI analytics.2+ years of experience with ETL/ELT pipelines.2+ years of experience with cloud data warehouses such as Snowflake, Amazon Redshift, Azure Synapse, and orchestration tools such as Airflow and dbt.2+ years of experience with SQL, including:Advanced joinsWindow functionsPerformance tuning2+ years of experience with data modeling, including:Star schemasSnowflake schemasSemantic layersHands-on experience with at least one major BI tool:Power BITableauLooker2+ years of hands-on experience with:DockerKubernetesCI/CD pipelinesMonitoring and loggingTerraformAnsibleAI/ML & GenAI SkillsExperience with AI/ML or Generative AI technologies.Experience with OpenAI APIs or similar LLM platforms preferred.Experience building natural-language-to-SQL/query systems preferred.Experience with prompt engineering and prompt strategy development.Experience with GenBI or AI-driven analytics preferred.Experience with semantic models supporting natural-language querying.Experience with statistical analysis, predictive analytics, or optimization preferred.AWS AI services experience, including Bedrock and SageMaker, preferred.Front-End / AI Application Skills – PreferredReact development.Chainlit development.Streamlit development.UX/UI development for analytics dashboards.Highly interactive reporting applications.AI/ML or GenAI application development.Complex application state and data handling.Real-time updates.Loading, retry, error, and partial-response handling.Latency handling and asynchronous UX patterns.Explainable AI interfaces, including confidence indicators, sources, and disclaimers.