Python Insfrastructure Engineer - Model Evaluation
Python Infrastructure Engineer — Model Evaluation (AI Training)About The RoleWhat if your Python expertise could directly shape how the world's most advanced AI models are built, evaluated, and improved? We're looking for a Senior Python Infrastructure Engineer to design and build the data pipelines, evaluation harnesses, and annotation tooling that power next-generation AI systems at leading research labs.This is a fully remote contract role with serious technical depth — the kind of work that ships to production and influences model quality at scale.Organization: AlignerrType: Hourly ContractLocation: RemoteCommitment: 20–40 hours/weekWhat You'll DoDesign, build, and optimize high-performance Python systems supporting AI data pipelines and model evaluation workflowsDevelop full-stack tooling and backend services for large-scale data annotation, validation, and quality controlBuild and maintain evaluation harnesses that integrate with inference frameworks and benchmarking pipelinesImprove reliability, performance, and safety across existing Python codebasesInstrument systems with observability tooling and metrics collection to monitor model performance and system healthIdentify bottlenecks and edge cases in data and system behavior, and implement scalable, maintainable fixesCollaborate with data, research, and engineering teams through synchronous design reviews and async communicationWho You AreNative or fluent English speaker with clear written and verbal communication skills3–5+ years of professional experience writing production-grade PythonFull-stack developer with a strong systems programming backgroundExperienced building evaluation harnesses for ML models and integrating with inference frameworksStrong grasp of observability, metrics collection, and system reliability practicesAble to commit 20–40 hours per week with consistent availabilityNice to HavePrior experience with data annotation pipelines, data quality systems, or model evaluation infrastructureFamiliarity with AI/ML workflows, model training, or benchmarking frameworksExperience with distributed systems or internal developer toolingBackground working directly with AI labs or ML research teamsWhy Join UsWork on real production systems at the frontier of AI development alongside leading research labsFully remote and flexible — work from wherever you do your best workFreelance autonomy with the structure of high-impact, technically challenging projectsMake a direct, measurable contribution to how next-generation AI models are evaluated and improvedPotential for ongoing work and contract extension as new projects launch