AIOps Lead, Software Engineering
Overview
In this role you lead Zelis’ operational transformation, steering AWS migration and AI-native capabilities to modernize operations. You will define the AIOps strategy, blending cloud operations, observability, automation, SRE, and AI-enabled workflows to boost reliability and efficiency. You’ll collaborate across Engineering, Cloud, Infrastructure, Security, and Operations to implement proactive, agent-assisted operations. This is a high-impact leadership role at a company focused on innovative healthcare financial solutions.
Compensation / Benefits401k with employer matchflexible paid time offhealth benefits (medical, dental, vision)parential leaveslife and disability insurancediscretionary bonus or incentives
ResponsibilitiesDefine and drive the AIOps strategy and architecture for the Price Business Unit during AWS migration and AI-native accelerationArchitect scalable observability across cloud and application environments (metrics, logs, traces, dashboards, alerting)Establish patterns for monitoring, resilience, automation, and governance in AWS ecosystemsDesign and implement AI-driven, agentic solutions to support incident response, anomaly detection, and workflow automationDevelop agentic and multi-agent operational tools coordinating monitoring, diagnostics, remediation, and operator assistanceBuild ChatOps capabilities to enhance collaboration, visibility, and rapid responseEmbed reliability engineering practices with SRE, Security, and Infrastructure teamsSet standards for telemetry, SLOs/SLIs, alert quality, escalation, and post-incident learningPromote governance, security, and auditability in AIOps implementationsCreate scalable playbooks, patterns, and operating models for broad AIOps adoptionMentor engineers in observability, automation, SRE, ChatOps, and AI-assisted operations
Key requirementsBachelor’s or higher with 12+ years (or Master’s with 10+ years) in cloud/platform operations, SRE, or observability initiativesDeep AWS expertise across compute, networking, storage, security, and automationStrong experience with New Relic, OpenSearch, and modern observability toolingProven track record applying AI to operations (event correlation, anomaly detection, remediation), including agentic workflowsExperience designing/implementing AI agents and multi-agent systems for operational useSolid SRE foundation: reliability, SLOs/SLIs, incident management, automation, resilienceExperience building or scaling ChatOps for cross-functional collaborationAutomation and engineering strength: scripting, APIs, event-driven patternsExcellent communication and stakeholder influence across technical and non-technical teamsGovernance, security, and trust in operational AI with oversight controlsLeadership and influenceProblem-solving orientationCross-functional collaborationAWS operational expertiseNew Relic and OpenSearch proficiencyAIOps implementation experience