Infrastructure Engineer
About Our Client:The organization operates in the artificial intelligence and computer vision industry, developing tools, communities, and resources that make the world programmable through AI. Its platform simplifies the development and deployment of computer vision models and is used by more than one million developers across industries including healthcare research, construction, environmental preservation, and autonomous navigation. The organization focuses on practical AI applications, innovation, and enabling developers to build and deploy machine learning solutions more efficiently.About the Opportunity:The Infrastructure Engineer is responsible for maintaining, securing, and scaling the core infrastructure supporting the organization’s AI and machine learning products. This role balances rapid development with reliability, security, and operational excellence across cloud infrastructure, databases, Kubernetes, microservices, and ML pipelines. The position contributes directly to platform stability, scalability, cost efficiency, and the organization’s ability to securely serve a growing user base.Responsibilities:• Secure, scale, and maintain cloud architecture, databases, file storage, search clusters, microservices, and machine learning pipelines.• Collaborate with Product and other teams to resolve operational and customer-facing challenges.• Manage and maintain containerized applications using Kubernetes.• Automate infrastructure deployment and configuration using Infrastructure-as-Code tools such as Terraform and Helm.• Monitor, troubleshoot, and scale applications across cloud environments including AWS and GCP.• Develop and maintain infrastructure and application code using Python and Node.js.• Implement and manage CI/CD pipelines using tools such as GitHub Actions.• Strengthen security practices and support compliance with standards including SOC 2, HIPAA, and GDPR.• Participate in incident response, troubleshooting, and on-call rotations.• Optimize cloud infrastructure costs and improve system observability, monitoring, and alerting.• Identify opportunities to improve infrastructure reliability, automation, and operational efficiency.Requirements:• Hands-on experience managing Kubernetes in production environments.• Proficiency with Infrastructure-as-Code tools such as Terraform and Helm, along with scripting capabilities.• Experience with AWS and/or GCP cloud environments, including infrastructure scaling and site reliability practices.• Strong development skills in Node.js and Python.• Hands-on experience supporting machine learning infrastructure and frameworks such as PyTorch or TensorFlow.• Familiarity with CI/CD automation and deployment tools.• Understanding of cloud security best practices, particularly within fast-paced startup environments.• Ability to leverage AI-assisted tools to improve development, infrastructure, and operational workflows.• Strong troubleshooting, problem-solving, and collaboration skills.Pay Range and Compensation Package:• The pay range and compensation package for this role will be determined based on the candidate’s experience, skills, qualifications, location, and other relevant factors.Equal Opportunity Statement: Our client is an equal opportunity employer. They celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, or national origin.Note:RemoteHunter is not the Employer of Record (EOR) for this role. Our purpose in this opportunity is to connect exceptional candidates with leading employers. We help job seekers worldwide discover roles that match their goals and guide them to complete their full application directly through the hiring company’s career page or ATS.