JOBSEARCHER

Lead Software Platform Engineer | AI Infrastructure & MLOps

A high-growth technology company is seeking a Lead Software Platform Engineer to own the architecture, scalability, and operational excellence of a next-generation AI and machine learning platform. This is a rare opportunity to shape the foundation that enables complex AI workloads to run securely, reliably, and at scale in highly regulated environments.This role sits at the intersection of distributed systems, cloud infrastructure, AI/ML operations, and software platform engineering. You will serve as a technical leader responsible for defining how AI models, large language models, and intelligent agents are deployed, governed, monitored, and operated in production.The platform you help build will be a customer-facing product rather than an internal tool. Your work will directly influence how sophisticated organizations apply AI to mission-critical workflows, requiring a balance of innovation, reliability, security, and cost efficiency.The OpportunityAs a senior technical leader, you will own the strategy and architecture for AI infrastructure that supports production-grade machine learning and generative AI systems. You will collaborate closely with software engineers, data engineers, AI practitioners, and platform teams to design systems that are scalable, observable, secure, and maintainable.This position is ideal for someone who enjoys solving large-scale technical challenges and has experience building platforms that support external users, stringent compliance requirements, and demanding service-level expectations.You will have significant influence over technical direction, platform standards, architectural decisions, and engineering best practices while helping mentor and elevate other engineers across the organization.What You'll Be Responsible ForOwning the architecture and technical roadmap for a cloud-native AI and machine learning platform.Designing and scaling infrastructure that supports model deployment, model lifecycle management, prompt management, and multi-model serving.Building reliable systems for both real-time and batch inference workloads.Developing frameworks for running and operating LLM-based applications, retrieval systems, and intelligent agent architectures in production.Establishing platform-wide standards for security, governance, tenant isolation, and data protection.Creating robust evaluation frameworks that measure model quality, detect regressions, and support safe releases.Defining observability practices including monitoring, alerting, logging, tracing, and performance management.Driving reproducibility, lineage, auditability, and compliance requirements across AI workflows.Contributing to infrastructure automation and deployment processes using modern cloud engineering practices.Leading technical design reviews and acting as a trusted advisor on architecture, scalability, reliability, and platform strategy.Evaluating emerging AI technologies, frameworks, and tooling while making informed build-versus-buy decision.Supporting operational excellence through incident response, production readiness reviews, and continuous improvement initiatives.What We're Looking ForYou are an experienced platform engineer or technical architect who has spent your career designing and operating large-scale cloud systems. You understand what it takes to move AI and machine learning workloads from experimentation into production environments where reliability, cost control, security, and performance matter.We're particularly interested in engineers who have built customer-facing AI infrastructure rather than internal-only tools.Key Requirements10+ years of software engineering and cloud infrastructure experience.Proven success designing and scaling distributed, cloud-native systems in production.Experience serving as a technical lead, architect, or principal-level engineer responsible for major architectural decisions.Deep hands-on experience building and operating AI/ML infrastructure at scale.Strong knowledge of modern LLM architectures and production deployment patterns, including retrieval-augmented generation (RAG), prompt lifecycle management, embeddings, and tool integration.Advanced coding skills in Python and TypeScript, with a strong focus on backend services and API development.Experience with model lifecycle management, deployment workflows, and model-serving infrastructure.Strong understanding of software quality, automated testing, release management, and CI/CD practices.Familiarity with cloud-native technologies including AWS, containers, and infrastructure-as-code tooling.Experience designing secure multi-tenant systems and handling sensitive or regulated data.Expertise in monitoring, observability, reliability engineering, and service-level management.Excellent communication skills with the ability to influence technical direction across multiple teams.Highly Desirable ExperienceProduction experience with agentic AI systems and orchestration frameworks.Knowledge of advanced model serving, latency optimization, and cost management techniques.Experience supporting multimodal AI workloads involving images or other complex data types.Exposure to model optimization techniques such as fine-tuning, distillation, or performance tuning.Background operating AI systems within regulated industries where auditability and compliance are critical.Experience working in data-intensive, scientific, healthcare, pharmaceutical, or similarly complex environments.Why This Role Stands OutThis is a high-impact leadership position where you'll help define the future architecture of a rapidly evolving AI platform. You'll work on meaningful technical challenges involving scalability, reliability, security, governance, and production AI operations while partnering with a team that views engineering excellence as a strategic advantage.If you're excited by the challenge of building the platforms that power next-generation AI applications and enjoy solving difficult infrastructure problems at scale, we'd love to hear from you.Applicants must be authorized to work in the United States. Sponsorship is not available for this position.