JOBSEARCHER

Principal Engineer

Technical Architect — Ai Systems & Platform InternalsAccellor is looking for a Technical Architect — AI Systems, Inference & Platform Internals to help design, scale, and optimize the systems that power AI systems and internal research workloads.This role is focused on the internal AI systems stack, including inference runtime, model serving, GPU infrastructure, distributed systems, context engineering, cost optimization, evaluation gates, observability, release safety, and production reliability.The ideal candidate is a senior hands-on architect who can reason across the full AI platform — from GPU-level performance and distributed inference to product-scale reliability, model deployment, safety, and cost-efficient operations.Key responsibilities include:1. AI Systems ArchitectureDesign and evolve large-scale AI systems that support AI systems and internal research workloads.Define architecture across inference runtime, model serving, request routing, batching, KV-cache handling, GPU scheduling, distributed execution, observability, release gates, and production rollout.Own technical trade-offs across latency, throughput, reliability, correctness, safety, scalability, cost, and infrastructure efficiency.2. Inference Runtime & Model ServingArchitect high-throughput, low-latency inference systems across large-scale GPU clusters.Work across inference engines, serving layers, scheduling systems, caching, streaming, deployment pipelines, and runtime optimization.Partner with engineering teams to improve model-serving efficiency, tail latency, GPU utilization, memory efficiency, correctness under load, and cost per request.Guide architecture decisions involving PyTorch, JAX, Triton, vLLM-style serving, CUDA/Triton kernels, distributed inference, tensor parallelism, pipeline parallelism, model sharding, and long-context serving.3. GPU, Kernel & Distributed PerformanceAnalyze and improve performance across GPU kernels, memory movement, collective communication, orchestration, and runtime scheduling.Guide engineering decisions involving CUDA, Triton, NCCL/RCCL, GPU profiling, memory pressure, compute utilization, tensor layouts, interconnect behavior, and distributed execution.Identify system-level bottlenecks across compute, memory, networking, scheduling, model execution, and data movement.4. Context EngineeringDesign and guide context engineering frameworks that determine what information should be passed to the model, how it should be structured, how much context should be used, and how context quality should be measured.Own architecture patterns for prompt structure, dynamic context assembly, retrieval-augmented generation, long-context management, conversation memory, tool context, agent state, multimodal context, source grounding, permission-aware retrieval, context compression, and context auditability.Ensure AI systems use the right context, from the right source, with the right permissions, at the right cost, and with measurable quality.5. Cost Optimization FrameworksDesign and build cost optimization frameworks for large-scale LLM and GenAI workloads.Create architecture patterns that reduce unnecessary token usage, redundant retrieval, repeated model calls, inefficient inference paths, and avoidable infrastructure spend.Drive model routing, token budgeting, prompt compression, context pruning, semantic caching, response caching, batch inference, async execution, fallback strategies, and cost telemetry across AI workflows.Ensure cost optimization does not compromise quality, safety, grounding, reliability, or user experience.6. Training & Research InfrastructureCollaborate with research and training infrastructure teams to support large-scale model training and post-training workflows.Contribute to architecture around distributed training, checkpointing, orchestration, fault tolerance, observability, data movement, evaluation infrastructure, and experiment velocity.Support frontier model workflows across pre-training, post-training, reinforcement learning, agent training, evaluation harnesses, and large-scale experiment execution.7. Release Safety, Validation & Evaluation GatesArchitect validation and release systems that ensure model updates, inference engine changes, runtime images, prompt changes, context changes, and platform releases are correct, safe, performant, and regression-free.Define release gates across correctness, numerical stability, latency, throughput, token usage, cost regression, context quality, retrieval quality, safety behavior, reliability, and model output quality.Ensure platform optimizations do not reduce safety, grounding, quality, or user trust.8. Reliability, Observability & Production OperationsDesign systems that make AI infrastructure observable, debuggable, reliable, and operationally safe.Define telemetry, tracing, dashboards, alerts, logs, profiling views, runbooks, SLOs, and post-incident learning loops.Provide visibility into prompts, context payloads, retrieved sources, token consumption, model selection, cache behavior, inference latency, GPU utilization, evaluation scores, safety events, cost, and failures.Turn production issues into stronger platform abstractions, safer rollout mechanisms, better automation, and more reliable infrastructure.9. Agentic & Multimodal Platform InternalsSupport architecture for AI agents, tool use, memory, function calling, multimodal interaction, long-running workflows, and internal or external agent deployment.Work across agent harnesses, evaluation pipelines, workflow orchestration, safety controls, state management, tool execution, memory systems, and product-facing runtime constraints.Ensure agentic and multimodal systems are reliable, observable, secure, cost-aware, and safe under real workloads.10. Technical LeadershipWork closely with Research, Inference, Runtime, Infrastructure, Product, Safety, Security, Technical Success, and Deployment teams. Act as a senior technical authority who can cut across layers, resolve ambiguity, identify systemic risks, and drive architecture decisions.Mentor engineers and technical leads on distributed systems, performance engineering, context engineering, cost optimization, production readiness, AI platform design, and architecture trade-offs.Represent architecture decisions through design docs, RFCs, diagrams, technical reviews, operational plans, and leadership-level summaries.