{"schemaVersion":"jobsearcher.job.v1","id":"8fe8e72c5611265f3cabdd9e","url":"https://jobsearcher.com/jobs/8fe8e72c5611265f3cabdd9e","canonicalUrl":"https://jobsearcher.com/jobs/8fe8e72c5611265f3cabdd9e","title":"Site Reliability Engineer","description":"Site Reliability Engineer\n\nAbout Arango:\nArango delivers a unified, natively multimodel contextual data platform that powers AI agents, assistants, and applications with the unified, current, and trusted business context needed to reason, decide, and act at scale.\n\nThe Arango Contextual Data Platform connects fragmented enterprise data with LLMs, copilots, and AI agents through a simplified architecture delivered out of the box. By combining graph, vector, document, key-value, and search capabilities in a single platform, Arango eliminates the complex stacks many organizations build to operationalize enterprise AI.\n\nTrusted by organizations including NVIDIA, HPE, the London Stock Exchange, PSI CRO, the U.S. Air Force, NIH, Siemens, Transient.AI, Matpriskollen, and Articul8, Arango helps enterprises move from AI pilots to reliable production systems faster while lowering infrastructure complexity and total cost of ownership. Arango is a proud member of the NVIDIA Inception Program and the AWS ISV Accelerate Program. Learn more at arango.ai, LinkedIn, and G2.\n\nValues at ArangoDB:\nInnovation: We continually push boundaries in database technology to meet the evolving needs of developers and organizations.\nCustomer-Centric Focus: We prioritize the success of our users by delivering reliable, scalable, and performant database solutions.\nCollaboration and Growth: We value a supportive work environment where employees can grow, learn, and share knowledge.\nJoining ArangoDB means becoming part of a forward-thinking company that is transforming how data is handled across diverse applications, while working in an inclusive, growth-oriented team.\nJob Overview:\nAt ArangoDB, we are building a robust, cloud-native infrastructure to support our distributed database systems, which power mission-critical applications for a wide range of industries. We are searching for a Site Reliability Engineer (SRE) to ensure the reliability, scalability, and performance of our infrastructure and applications, with a focus on automation, monitoring, and optimizing cloud environments.\nAs a Site Reliability Engineer (SRE), you will be responsible for maintaining and improving the reliability of our distributed database systems running on Kubernetes and cloud environments (AWS, Google Cloud). You will design, implement, and maintain scalable infrastructure solutions, improve and expand observability into these solutions, and troubleshoot complex system issues. It is expected that you will come to work to write clean and efficient code in Golang, working closely with development teams\nYour goal is to ensure high availability and performance of our cloud-based systems, automating repetitive tasks, and enhancing our CI/CD pipelines. If you're passionate about building resilient systems, managing cloud infrastructure, and using Golang to create scalable solutions (or willingness to learn Golang), we want to hear from you!\nKey Responsibilities:\nDesign, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms.\nEnsure the scalability, performance, and reliability of our Kubernetes-based distributed database systems.\nCollaborate with developers to write efficient, production-grade code in Golang to automate infrastructure management and improve system operations.\nOptimize and automate CI/CD pipelines, deployment processes, and monitoring systems to support our production environment.\nDevelop strategies for disaster recovery, high availability, and fault tolerance.\nProactively identify system bottlenecks, troubleshoot, and resolve issues across the stack (network, OS, cloud infrastructure).\nImplement monitoring, logging, and alerting systems to ensure visibility into system health and performance.\nParticipate in on-call rotations to support critical production systems and respond to incidents.\nCollaborate with cross-functional teams to improve overall system reliability and scalability.\nCollaborate with the Customer Success team to resolve customer issues.\nRequired Skills and Qualifications:\nProven experience as an SRE or DevOps Engineer in a cloud-native environment.\nProficiency with Kubernetes in managing large-scale, distributed systems.\nExperience with cloud providers such as AWS and Google Cloud (GCP).\nSolid understanding of networking, security practices, and troubleshooting methods\nUnderstanding of Linux internals (processes, environment variables etc.)\nFamiliarity with containerization technologies (e.g., Docker).\nKnowledge of CI/CD practices and tools (Jenkins, CircleCI, etc.).\nFamiliarity with alerting, monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack).\nStrong troubleshooting and problem-solving skills, with the ability to address complex infrastructure issues.\nExcellent communication and collaboration skills with a focus on continuous improvement and operational excellence.\nStrong ability to self-organize and to work independently as part of a remote team Knowledge of version control systems, particularly Git\nFamiliarity with programming languages such as Golang or Python\n\nNice-to-Have:\nExperience managing distributed databases or large-scale data storage systems. Knowledge of security best practices in cloud environments.\nExperience with scripting languages like Python or Bash.\nExperience with Infrastructure-as-Code (IaC) tools like Terraform is a plus. Experience working with GitOps\nStrong programming skills in Golang, with experience in developing automation tools, scripts, or services.\nLocation: EU Timezone, preferably within the EU itself (Remote)\nWhat Makes Arango Special?\nAt Arango, we believe that AI is only as powerful as the data foundation. Our mission is to help organizations build AI systems that can reason, decide and act based on unified, current, and trusted business context at scale. We are helping define a new category of infrastructure: the contextual data layer for AI.\nWorking at Arango means:\nContributing to cutting-edge AI and data infrastructure\nCollaborating with experienced engineers, marketers, and product leaders\nHelping shape how enterprises build AI-powered applications\nIf you're excited about the intersection of AI, data, and social media, we’d love to hear from you.\n\n6YTiSNbuMj","company":"Arango","rawCompany":"arango","city":"Remote","state":"OR","isRemote":false,"isActive":false,"createdAt":"2026-08-08T13:29:10.079Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"518210","title":"Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services","slug":"computing-infrastructure-providers-data-processing-web-hosting-and-related-services"},{"code":"541519","title":"Other Computer Related Services","slug":"other-computer-related-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Site Reliability Engineer","description":"Site Reliability Engineer\n\nAbout Arango:\nArango delivers a unified, natively multimodel contextual data platform that powers AI agents, assistants, and applications with the unified, current, and trusted business context needed to reason, decide, and act at scale.\n\nThe Arango Contextual Data Platform connects fragmented enterprise data with LLMs, copilots, and AI agents through a simplified architecture delivered out of the box. By combining graph, vector, document, key-value, and search capabilities in a single platform, Arango eliminates the complex stacks many organizations build to operationalize enterprise AI.\n\nTrusted by organizations including NVIDIA, HPE, the London Stock Exchange, PSI CRO, the U.S. Air Force, NIH, Siemens, Transient.AI, Matpriskollen, and Articul8, Arango helps enterprises move from AI pilots to reliable production systems faster while lowering infrastructure complexity and total cost of ownership. Arango is a proud member of the NVIDIA Inception Program and the AWS ISV Accelerate Program. Learn more at arango.ai, LinkedIn, and G2.\n\nValues at ArangoDB:\nInnovation: We continually push boundaries in database technology to meet the evolving needs of developers and organizations.\nCustomer-Centric Focus: We prioritize the success of our users by delivering reliable, scalable, and performant database solutions.\nCollaboration and Growth: We value a supportive work environment where employees can grow, learn, and share knowledge.\nJoining ArangoDB means becoming part of a forward-thinking company that is transforming how data is handled across diverse applications, while working in an inclusive, growth-oriented team.\nJob Overview:\nAt ArangoDB, we are building a robust, cloud-native infrastructure to support our distributed database systems, which power mission-critical applications for a wide range of industries. We are searching for a Site Reliability Engineer (SRE) to ensure the reliability, scalability, and performance of our infrastructure and applications, with a focus on automation, monitoring, and optimizing cloud environments.\nAs a Site Reliability Engineer (SRE), you will be responsible for maintaining and improving the reliability of our distributed database systems running on Kubernetes and cloud environments (AWS, Google Cloud). You will design, implement, and maintain scalable infrastructure solutions, improve and expand observability into these solutions, and troubleshoot complex system issues. It is expected that you will come to work to write clean and efficient code in Golang, working closely with development teams\nYour goal is to ensure high availability and performance of our cloud-based systems, automating repetitive tasks, and enhancing our CI/CD pipelines. If you're passionate about building resilient systems, managing cloud infrastructure, and using Golang to create scalable solutions (or willingness to learn Golang), we want to hear from you!\nKey Responsibilities:\nDesign, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms.\nEnsure the scalability, performance, and reliability of our Kubernetes-based distributed database systems.\nCollaborate with developers to write efficient, production-grade code in Golang to automate infrastructure management and improve system operations.\nOptimize and automate CI/CD pipelines, deployment processes, and monitoring systems to support our production environment.\nDevelop strategies for disaster recovery, high availability, and fault tolerance.\nProactively identify system bottlenecks, troubleshoot, and resolve issues across the stack (network, OS, cloud infrastructure).\nImplement monitoring, logging, and alerting systems to ensure visibility into system health and performance.\nParticipate in on-call rotations to support critical production systems and respond to incidents.\nCollaborate with cross-functional teams to improve overall system reliability and scalability.\nCollaborate with the Customer Success team to resolve customer issues.\nRequired Skills and Qualifications:\nProven experience as an SRE or DevOps Engineer in a cloud-native environment.\nProficiency with Kubernetes in managing large-scale, distributed systems.\nExperience with cloud providers such as AWS and Google Cloud (GCP).\nSolid understanding of networking, security practices, and troubleshooting methods\nUnderstanding of Linux internals (processes, environment variables etc.)\nFamiliarity with containerization technologies (e.g., Docker).\nKnowledge of CI/CD practices and tools (Jenkins, CircleCI, etc.).\nFamiliarity with alerting, monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack).\nStrong troubleshooting and problem-solving skills, with the ability to address complex infrastructure issues.\nExcellent communication and collaboration skills with a focus on continuous improvement and operational excellence.\nStrong ability to self-organize and to work independently as part of a remote team Knowledge of version control systems, particularly Git\nFamiliarity with programming languages such as Golang or Python\n\nNice-to-Have:\nExperience managing distributed databases or large-scale data storage systems. Knowledge of security best practices in cloud environments.\nExperience with scripting languages like Python or Bash.\nExperience with Infrastructure-as-Code (IaC) tools like Terraform is a plus. Experience working with GitOps\nStrong programming skills in Golang, with experience in developing automation tools, scripts, or services.\nLocation: EU Timezone, preferably within the EU itself (Remote)\nWhat Makes Arango Special?\nAt Arango, we believe that AI is only as powerful as the data foundation. Our mission is to help organizations build AI systems that can reason, decide and act based on unified, current, and trusted business context at scale. We are helping define a new category of infrastructure: the contextual data layer for AI.\nWorking at Arango means:\nContributing to cutting-edge AI and data infrastructure\nCollaborating with experienced engineers, marketers, and product leaders\nHelping shape how enterprises build AI-powered applications\nIf you're excited about the intersection of AI, data, and social media, we’d love to hear from you.\n\n6YTiSNbuMj","datePosted":"2026-08-08T13:29:10.079Z","dateModified":"2026-08-08T13:29:10.079Z","hiringOrganization":{"@type":"Organization","name":"Arango","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Remote","addressRegion":"OR","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"8fe8e72c5611265f3cabdd9e"},"url":"https://jobsearcher.com/jobs/8fe8e72c5611265f3cabdd9e"}}