Senior Software Engineer – AI Platform Reliability
ARCHIVED
We can't find an active application page for this role right now. It may reopen or be listed elsewhere. Use Next Steps to search for an active apply link and similar live jobs.
12-month contract with strong potential for extension or full-time conversionStandard business hours (Monday–Friday, 8:00 AM – 5:00 PM)Business casual work environmentCharlotte, NC – Hybrid schedule (3 days onsite, 2 days remote)W-2 only (not eligible for C2C or 1099); must be authorized to work in the U.S. without sponsorshipBrooksource is seeking a Senior Software Engineer with strong full stack and infrastructure experience and a passion for AI to join our enterprise financial services client in Charlotte, NC.This isn't a traditional senior dev role. You'll be embedded directly into an SRE team working on the latest in AI orchestrated platforms — building the full stack tooling, infrastructure, and reliability patterns that make intelligent systems production-ready. That means frontend interfaces, backend services, Terraform-first environments, and the instrumentation layer that keeps AI workflows honest at scale. If you care deeply about system behavior under load, enjoy tracing a failure across a distributed call graph, and want your code to directly reduce production incidents, this is for you.TECHNICAL SKILLS:Strong backend development experience with Node.js, Python, or Go (must have at least two)An AI-first mindset – this engineer will be expected to use AI tooling aggressively and thoughtfully, from code generation to incident investigation to building that platforms that run AI workflows in production. You should have a clear point of view on which models are appropriate for which tasks.Experience building full stack applications (React or similar frontend frameworks)Strong experience with Terraform and infrastructure as code (IaC)Hands-on experience with AWS (understanding infrastructure layer, not just application layer)Experience with observability and monitoring tools (Dynatrace, Datadog, Splunk, AppDynamics)Experience building or supporting CI/CD pipelines, containers, and automated deploymentsStrong debugging experience within distributed systems environmentsExperience in developing code through the use of AI, ideally through either Claude or GeminiIdeally:Experience supporting or instrumenting AI/LLM-based platforms (LangChain, LangGraph, etc.)Exposure to high-volume transactional or financial systemsExperience with Salesforce, Genesys, or telephony platformsJOB RESPONSIBILITIES:Build and enhance internal engineering platforms, including full stack features across frontend (React) and backend servicesDevelop automation solutions to reduce manual operational processes, including incident detection and self-healing systemsDesign and implement observability frameworks, including alerting pipelines and system health monitoringPartner with DevOps and SRE teams to build, deploy, and maintain applications in AWS using Terraform-first environmentsInvestigate and resolve production issues, performing root cause analysis across distributed systemsCreate dashboards and tooling that provide real-time visibility into system performance and business transactionsContribute to reliability engineering efforts by improving system performance, scalability, and incident response processesEight Eleven Group provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, national origin, age, sex, citizenship, disability, genetic information, gender, sexual orientation, gender identity, marital status, amnesty or status as a covered veteran in accordance with applicable federal, state, and local laws.