Senior Software Engineer, DevOps
Utilidata is a fast-growing NVIDIA-backed AI company enabling AI data centers to dynamically orchestrate power and unlock more compute capacity from existing energy infrastructure. For over a decade, we have applied AI to the electric grid — bringing real-time visibility and power-flow control to complex energy infrastructure. Our Karman platform, built on a custom NVIDIA module, brings that same capability to AI data centers, giving operators a way to better use the power already available to them.
We are seeking a Senior Software Engineer, DevOps to help design, build, and operate Utilidata's off-device platform that ingests, processes, and serves data flowing from edge AI devices. The role will build and maintain infrastructure across on-premises and cloud environments – bridging edge deployments with cloud-based data processing to support analytics, operations, and ML workloads at scale. This is a hands-on development role with deep technical ownership and cross-team visibility. This engineer will build and maintain the systems that keep our platform running, help shape infrastructure and deployment best practices, and support less experienced engineers. This engineer will partner closely with on-device and ML teams to ensure our off-device platform is resilient, well-instrumented, and ready to scale.
Responsibilities
Support the deployment and management of containerized applications using Kubernetes, ensuring optimal performance and availability
Contribute to strategic planning on how infrastructure solutions evolve to match Data Center partner requirements
Design, implement, and maintain scalable and reliable systems on AWS and/or on-premise
Utilize Terraform for infrastructure as code to automate the provisioning and management of cloud resources
Monitor system performance and uptime, ensuring systems meet established service level objectives (SLOs)
Support SOC2 security compliance requirements for data handling
Guide team members in DevOps practices, promoting a culture of reliability and excellence
Advocate for automation of operational tasks to enhance efficiency and reduce manual intervention
Collaborate with cross-functional teams to build and maintain CI/CD pipelines
Troubleshoot and resolve complex production issues, conducting root cause analysis and implementing corrective actions
Participate in on-call rotations and incident response teams
Assist in capacity planning, performance tuning, and technical decision-making
Drive continuous improvement initiatives for processes and infrastructure
Minimum Qualifications
8+ years of development experience including experience in platform engineering, SRE, or distributed systems, with demonstrated senior-level impact
Experience designing and operating infrastructure across on-premises and cloud environments
Strong proficiency in container orchestration, particularly Kubernetes
Strong proficiency with AWS services and architecture
Hands-on experience with Terraform for infrastructure automation
Familiarity with monitoring tools (Prometheus, Grafana, or similar) and observability best practices
Strong problem-solving skills and attention to detail
Strong communication and collaboration skills, with experience contributing to technical outcomes
Willingness to travel up to 20% of time
Enhanced Qualifications (Nice to Have)
Bachelor's degree in Computer Science, Engineering, or a related field
Experience supporting or enabling MLOps platforms, model deployment pipelines, or ML-adjacent infrastructure
AI workload scheduling using Kubernetes
Knowledge of Apache Spark for large-scale data processing
Knowledge of database technologies (SQL, NoSQL)
Understanding of networking concepts and security best practices
Salary Range: $160,000 to $190,000 base compensation depending on experience and stock options. Salary will be commensurate with an individual's skills, training, years of experience, and in line with internal compensation bands.
Location: This position is based at our company headquarters in Ann Arbor, Michigan, with flexibility for occasional remote work.
Our Commitments:
Utilidata values the diversity of our team. We provide equal employment opportunities without regard to race, color, religion, creed, sex, gender, sexual orientation, gender identity or expression, national origin, age, physical disability, mental disability, medical condition, pregnancy or childbirth, sexual orientation, genetics, genetic information, marital status, or status as a covered veteran or any other basis protected by applicable federal, state and local laws.
We are committed to:
Creating a diverse and inclusive workplace that is welcoming, supportive, affirming and respectful
Empowering employees to solve problems and work together to make a difference
Providing mentorship and growth opportunities as part of a collaborative team
A flexible work environment with flexible paid time off
Competitive compensation and benefits, including health, dental, vision, and employer-match 401k
FtX4A40HW7