{"schemaVersion":"jobsearcher.job.v1","id":"19c181bc311af19c20a83cec","url":"https://jobsearcher.com/jobs/19c181bc311af19c20a83cec","canonicalUrl":"https://jobsearcher.com/jobs/19c181bc311af19c20a83cec","title":"Devops Engineer","description":"Overview:\n\nProdapt is the largest and fastest-growing specialized player in the Connectedness industry, recognized by Gartner as a Large, Telecom-Native, Regional IT Service Provider across North America, Europe and Latin America. With its singular focus on the domain, Prodapt has built deep expertise in the most transformative technologies that connect our world. Prodapt is a trusted partner for enterprises across all layers of the Connectedness vertical. Prodapt designs, configures, and operates solutions across their digital landscape, network infrastructure, and business operations – and craft experiences that delight their customers. Today, Prodapt’s clients connect 1.1 billion people and 5.4 billion devices, and are among the largest telecom, media, and internet firms in the world. Prodapt works with Google, Amazon, Verizon, Vodafone, Liberty Global, Liberty Latin America, Claro, Lumen, Windstream, Rogers, Telus, KPN, Virgin Media, British Telecom, Deutsche Telekom, Adtran, Samsung, and many more. A “Great Place To Work® Certified™” company, Prodapt employs over 6,000 technology and domain experts in 30+ countries across North America, Latin America, Europe, Africa, and Asia. Prodapt is part of the 130-year-old business conglomerate The Jhaver Group, which employs over 30,000 people across 80+ locations globally.\n\nUnlike a traditional production support role, this position requires strong engineering aptitude and a proactive operational mindset. The ideal candidate combines customer/agent facing application ecosystem knowledge, application monitoring expertise, incident management experience, automation skills, and a passion for leveraging AI and GenAI technologies to improve operational efficiency.\n\nYou will partner closely with engineering teams, release management, security, compliance, and business stakeholders to identify risks, reduce incidents, enhance monitoring capabilities, and drive platform reliability across a complex ecosystem supporting hundreds of enterprise applications.\n\nResponsibilities:\n\nSite Reliability & Operations\n\nServe as a key member of the Engineering Operations organization supporting self-assist Web & Mobile App and related business applications.\nProvide Tier 1/2 operational support for production systems in a 24x7 environment.\nMonitor application health, performance, availability, and customer experience across the platform.\nDrive proactive issue detection and prevention rather than relying solely on customer-reported incidents.\nParticipate in incident response, triage, war rooms, major incident management, and post-incident reviews.\nPerform root cause analysis (RCA) and identify opportunities to improve platform stability and resiliency.\nPartner with Tier 1, Tier 2, Tier 3, infrastructure, security, and application teams to rapidly resolve issues.\nCreate and maintain operational runbooks, knowledge articles, and support documentation.\n\nObservability & Monitoring\n\nBuild, maintain, and optimize monitoring dashboards, alerts, and health checks.\nAnalyze application logs, API activity, transactions, and performance metrics.\nUtilize observability and monitoring platforms including:\nDynatrace\nELK Stack (Elasticsearch, Logstash, Kibana)\nCatchpoint or other synthetic monitoring solutions\nQuantum Metrics or other user session replay solutions\nReduce alert fatigue through automation, threshold tuning, and intelligent event correlation.\nDevelop and enhance monitoring strategies to provide end-to-end visibility across Digital ecosystem and integrated systems.\n\nIncident Management & Problem Management\n\nAct as a technical responder during production incidents and service disruptions.\nCoordinate issue resolution efforts across multiple technical teams and stakeholders.\nManage incident lifecycle activities including:\nDetection\nTriage\nResolution\nCommunication\nRoot cause analysis\nIdentify recurring issues and lead problem management initiatives to eliminate operational inefficiencies.\n\nAutomation & AI Enablement\n\nDevelop innovative approaches to reduce manual operational effort through automation.\nLeverage AI, GenAI, agentic workflows, and intelligent operational tooling where appropriate.\nCreate automation solutions to improve:\nShift handoffs\nIncident reporting\nAlert management\nKnowledge management\nOperational reporting\nContribute to internal AI initiatives that improve engineering productivity and service reliability.\nEvaluate and implement automation opportunities across monitoring, ticketing, collaboration, and support workflows.\nRequirements:\n\nRequired Qualifications\n\nOverall experience: 5+ years experience performing Production Support for Mission Critical, high-performance applications (Customer Care, Retail and eCommerce customer/agent facing application experience preferred)\nBachelor’s degree in Computer Science, Information Technology, Engineering, or a related field\nExperience using Docker, Kubernetes and Microsoft Azure Cloud, Unix, Networking and troubleshooting knowledge\nExperience with enterprise monitoring and observability tools such as:\nApplication & Infrastructure Performance Monitoring tools like Dynatrace\nApplication Log Analytics tools like Elastic\nVisualization tools like Kibana and Grafana. EFK stack experience preferred\nCreation of Dashboards on Dynatrace, ELK and Grafana\nDebugging java log, debugging microservices log\nExperience in Relational & NoSQL databases like Oracle & Cassandra\nStrong experience with:\nIncident Management\nProblem Management\nRoot Cause Analysis\nService restoration processes\nExperience working across multiple technical teams in large-scale enterprise environments.\nStrong written and verbal communication skills.\n\nPreferred Qualifications\n\nSite Reliability Engineering (SRE) experience.\nExperience building operational dashboards and observability platforms.\nHands-on experience with AI, GenAI, Copilot, & operational automation.\nExperience with scripting or automation using Python, PowerShell, JavaScript, or similar technologies.\nExposure to release management, change management, and production deployments.\nUnderstanding of security, SOX, compliance, and enterprise operational controls.\nExperience supporting large-scale environments with hundreds of integrated applications.\nGenerative AI and Workflow Automation skills\nGPT-4o LLM – Advanced\nLangGraph and LangChain\nGoogle DialogFlow(Api.ai)\nGoogle Vertex\nDatabricks, Spark, Snowflake\nPython, Java, SQL\nSelenium, Playwright Robotic Process Automation (RPA)\nPower Automate, Automation Anywhere\nDemonstrated experience leveraging AI-driven tools for automating end-to-end operational workflows\nDemonstrated experience using Text Generative & Code Generative AI Models\nAutomation, Gen AI & Agentic Workflow Technical Skills to include","company":"Prodapt","rawCompany":"prodapt","city":"Irving","state":"TX","isRemote":false,"isActive":false,"createdAt":"2026-09-04T10:33:31.845Z","occupations":[{"code":"15-1299.08","title":"Computer Systems Engineers/Architects","slug":"computer-systems-engineers-architects"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"},{"code":"15-1244.00","title":"Network and Computer Systems Administrators","slug":"network-and-computer-systems-administrators"}],"industries":[{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Devops Engineer","description":"Overview:\n\nProdapt is the largest and fastest-growing specialized player in the Connectedness industry, recognized by Gartner as a Large, Telecom-Native, Regional IT Service Provider across North America, Europe and Latin America. With its singular focus on the domain, Prodapt has built deep expertise in the most transformative technologies that connect our world. Prodapt is a trusted partner for enterprises across all layers of the Connectedness vertical. Prodapt designs, configures, and operates solutions across their digital landscape, network infrastructure, and business operations – and craft experiences that delight their customers. Today, Prodapt’s clients connect 1.1 billion people and 5.4 billion devices, and are among the largest telecom, media, and internet firms in the world. Prodapt works with Google, Amazon, Verizon, Vodafone, Liberty Global, Liberty Latin America, Claro, Lumen, Windstream, Rogers, Telus, KPN, Virgin Media, British Telecom, Deutsche Telekom, Adtran, Samsung, and many more. A “Great Place To Work® Certified™” company, Prodapt employs over 6,000 technology and domain experts in 30+ countries across North America, Latin America, Europe, Africa, and Asia. Prodapt is part of the 130-year-old business conglomerate The Jhaver Group, which employs over 30,000 people across 80+ locations globally.\n\nUnlike a traditional production support role, this position requires strong engineering aptitude and a proactive operational mindset. The ideal candidate combines customer/agent facing application ecosystem knowledge, application monitoring expertise, incident management experience, automation skills, and a passion for leveraging AI and GenAI technologies to improve operational efficiency.\n\nYou will partner closely with engineering teams, release management, security, compliance, and business stakeholders to identify risks, reduce incidents, enhance monitoring capabilities, and drive platform reliability across a complex ecosystem supporting hundreds of enterprise applications.\n\nResponsibilities:\n\nSite Reliability & Operations\n\nServe as a key member of the Engineering Operations organization supporting self-assist Web & Mobile App and related business applications.\nProvide Tier 1/2 operational support for production systems in a 24x7 environment.\nMonitor application health, performance, availability, and customer experience across the platform.\nDrive proactive issue detection and prevention rather than relying solely on customer-reported incidents.\nParticipate in incident response, triage, war rooms, major incident management, and post-incident reviews.\nPerform root cause analysis (RCA) and identify opportunities to improve platform stability and resiliency.\nPartner with Tier 1, Tier 2, Tier 3, infrastructure, security, and application teams to rapidly resolve issues.\nCreate and maintain operational runbooks, knowledge articles, and support documentation.\n\nObservability & Monitoring\n\nBuild, maintain, and optimize monitoring dashboards, alerts, and health checks.\nAnalyze application logs, API activity, transactions, and performance metrics.\nUtilize observability and monitoring platforms including:\nDynatrace\nELK Stack (Elasticsearch, Logstash, Kibana)\nCatchpoint or other synthetic monitoring solutions\nQuantum Metrics or other user session replay solutions\nReduce alert fatigue through automation, threshold tuning, and intelligent event correlation.\nDevelop and enhance monitoring strategies to provide end-to-end visibility across Digital ecosystem and integrated systems.\n\nIncident Management & Problem Management\n\nAct as a technical responder during production incidents and service disruptions.\nCoordinate issue resolution efforts across multiple technical teams and stakeholders.\nManage incident lifecycle activities including:\nDetection\nTriage\nResolution\nCommunication\nRoot cause analysis\nIdentify recurring issues and lead problem management initiatives to eliminate operational inefficiencies.\n\nAutomation & AI Enablement\n\nDevelop innovative approaches to reduce manual operational effort through automation.\nLeverage AI, GenAI, agentic workflows, and intelligent operational tooling where appropriate.\nCreate automation solutions to improve:\nShift handoffs\nIncident reporting\nAlert management\nKnowledge management\nOperational reporting\nContribute to internal AI initiatives that improve engineering productivity and service reliability.\nEvaluate and implement automation opportunities across monitoring, ticketing, collaboration, and support workflows.\nRequirements:\n\nRequired Qualifications\n\nOverall experience: 5+ years experience performing Production Support for Mission Critical, high-performance applications (Customer Care, Retail and eCommerce customer/agent facing application experience preferred)\nBachelor’s degree in Computer Science, Information Technology, Engineering, or a related field\nExperience using Docker, Kubernetes and Microsoft Azure Cloud, Unix, Networking and troubleshooting knowledge\nExperience with enterprise monitoring and observability tools such as:\nApplication & Infrastructure Performance Monitoring tools like Dynatrace\nApplication Log Analytics tools like Elastic\nVisualization tools like Kibana and Grafana. EFK stack experience preferred\nCreation of Dashboards on Dynatrace, ELK and Grafana\nDebugging java log, debugging microservices log\nExperience in Relational & NoSQL databases like Oracle & Cassandra\nStrong experience with:\nIncident Management\nProblem Management\nRoot Cause Analysis\nService restoration processes\nExperience working across multiple technical teams in large-scale enterprise environments.\nStrong written and verbal communication skills.\n\nPreferred Qualifications\n\nSite Reliability Engineering (SRE) experience.\nExperience building operational dashboards and observability platforms.\nHands-on experience with AI, GenAI, Copilot, & operational automation.\nExperience with scripting or automation using Python, PowerShell, JavaScript, or similar technologies.\nExposure to release management, change management, and production deployments.\nUnderstanding of security, SOX, compliance, and enterprise operational controls.\nExperience supporting large-scale environments with hundreds of integrated applications.\nGenerative AI and Workflow Automation skills\nGPT-4o LLM – Advanced\nLangGraph and LangChain\nGoogle DialogFlow(Api.ai)\nGoogle Vertex\nDatabricks, Spark, Snowflake\nPython, Java, SQL\nSelenium, Playwright Robotic Process Automation (RPA)\nPower Automate, Automation Anywhere\nDemonstrated experience leveraging AI-driven tools for automating end-to-end operational workflows\nDemonstrated experience using Text Generative & Code Generative AI Models\nAutomation, Gen AI & Agentic Workflow Technical Skills to include","datePosted":"2026-09-04T10:33:31.845Z","dateModified":"2026-09-04T10:33:31.845Z","hiringOrganization":{"@type":"Organization","name":"Prodapt","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Irving","addressRegion":"TX","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"19c181bc311af19c20a83cec"},"url":"https://jobsearcher.com/jobs/19c181bc311af19c20a83cec"}}