JOBSEARCHER

Fall 2026 Internship- AI/ML Developer

About UsWe’re building a transformative healthcare accreditation platform that is revolutionizing how hospitals manage compliance, quality improvement, and regulatory processes. Our platform combines cutting-edge technology with deep healthcare domain expertise to solve real problems for healthcare organizations nationwide.The OpportunityThe goal is to have interns turn into full-time employees; therefore, you will be given full-time responsibilities from day one. You will be working in a high-velocity growth startup and will be required to move fast. You’ll work directly with our engineering team on production healthcare ML systems, gaining hands-on experience with enterprise-grade pipelines while making real contributions that impact our product and customers.Compensation StructureBase position is unpaid; however, qualified candidates may receive upfront equity compensation based on their experience level and demonstrated capabilities. We evaluate each applicant individually and offer equity packages commensurate with their potential contribution.About the RoleWe’re hiring an ML Pipeline & Data Science Developer. This role is focused on building and operating production machine learning pipelines—clustering, classification, recommendation, and outcome analysis—on top of healthcare event and quality data. You’ll design the data workflows that power risk identification, pattern detection, and corrective action effectiveness analysis across the platform.Requirements• Deep AWS experience: Lambda, Step Functions, EventBridge, S3, SageMaker, and Bedrock for production ML infrastructure• Strong Python ML fundamentals: scikit-learn, pandas, numpy, scipy for building and evaluating models• Pipeline architecture: Designing end-to-end data and ML pipelines—ingestion, transformation, feature engineering, model training, inference, and monitoring• Clustering and classification: Hands-on experience with unsupervised and supervised learning techniques (topic modeling, text clustering, multi-class classification)• Statistical rigor: Hypothesis testing, A/B testing, experimental design, and outcome measurement• Version control: Git for code and model versioning• Collaborative mindset: Ability to work within a specialized team structure and move fast in a startup environmentNice to Have• Experience with BERTopic, sentence-transformers, or similar NLP clustering libraries• Familiarity with recommendation systems or best-practice identification from historical outcome data• Healthcare or regulated industry experience• Time-series forecasting and anomaly detection• Knowledge of data privacy and compliance frameworks• Experience with MongoDB aggregation pipelines or similar document-database analyticsWhat You’ll BuildML Pipelines & Predictive Risk Analytics• Event clustering pipelines: Automated workflows that ingest healthcare event data and cluster by similarity to allow AI inference to surface patterns and systemic risks• Classification systems: Multi-class models that categorize incoming events by type, severity, and process area using both structured metadata and unstructured text• Corrective action recommendation: Models that analyze historical outcomes to identify the most effective actions for a given type of event or similar event, surfacing best practices from past resolutions• Risk scoring and prioritization: Scoring frameworks that rank identified risks by frequency, severity, and trend direction to focus quality improvement efforts• Outcome and effectiveness analysis: Statistical models that measure whether implemented corrective actions actually reduce recurrence, with confidence intervals and significance testingData Pipeline Infrastructure• Ingestion and transformation: Lambda-based pipelines that process incoming event data, normalize fields, and route records through feature engineering stages• Feature engineering: Domain-specific feature stores built on healthcare event attributes—process area, finding type, facility context, temporal patterns, and text-derived features• Batch and real-time inference: Scheduled batch pipelines (EventBridge + Lambda) for nightly AI/ML model runs and on-demand inference endpoints for interactive use• Model monitoring and versioning: MLOps workflows for tracking experiments, model performance drift, and A/B comparisons across model versions• Visualization and reporting: Interactive dashboards and data visualizations that surface pipeline outputs to end users and internal stakeholdersKey Responsibilities• Design and implement end-to-end ML pipelines on AWS—from data ingestion through model inference and output delivery• Build clustering and topic modeling workflows that group healthcare events by similarity and identify systemic patterns• Develop classification models for automated event categorization across process areas and finding types• Analyze historical action data to build recommendation models that surface the most effective interventions for a given scenario• Conduct statistical analysis on action outcomes to measure effectiveness and identify best practices• Engineer features from healthcare event data—combining structured fields, temporal patterns, and text-derived signals• Build and maintain Lambda-based data transformation and inference pipelines with EventBridge scheduling• Implement model versioning, experiment tracking, and performance monitoring using SageMaker• Create visualizations and reports that communicate risk patterns and model outputs to non-technical stakeholders• Collaborate with product and engineering teams to integrate ML outputs into the broader platformRequired QualificationsCandidates must meet all Core Qualifications plus demonstrate depth in the ML Pipeline & Data Science focus area.Core Qualifications• Advanced Python programming skills (2+ years)• Git for version control and collaborative development workflows• Understanding of deployment and production systems• Experience building on AWS (Lambda, S3, EventBridge, Step Functions, SageMaker)ML Pipeline & Data Science• 2+ years with Python ML libraries (scikit-learn, pandas, numpy, scipy)• Experience designing and operating ML pipelines—data ingestion, feature engineering, training, inference• Hands-on experience with clustering (k-means, DBSCAN, hierarchical, or topic modeling such as BERTopic/LDA)• Strong statistical analysis skills: hypothesis testing, A/B testing, experimental design, outcome measurement• Experience with text embeddings and NLP feature extraction (sentence-transformers, spaCy, or similar)• 1+ years with AWS SageMaker or equivalent managed ML platform• Data visualization expertise (matplotlib, seaborn, plotly, Tableau)• SQL proficiency for analytics and data querying• MLflow or equivalent for experiment trackingNice to Have• Experience with recommendation systems or outcome-based ranking models• Healthcare or regulated industry experience• MongoDB aggregation pipelines for document-database analytics• Hugging Face transformers and fine-tuning workflows• Time-series forecasting and anomaly detection• Advanced statistical modeling (survival analysis, propensity scoring)• Knowledge of data privacy and compliance frameworksTechnical StackCategory TechnologiesAWS Infrastructure Lambda, Step Functions, EventBridge, S3, EC2, SageMaker, BedrockML & Data Science Python (scikit-learn, pandas, numpy, scipy, statsmodels), MLflow, BERTopic, sentence-transformersNLP & Text Processing spaCy, NLTK, sentence-transformers, text embeddings, topic modelingData & Analytics SQL, MongoDB aggregation pipelines, Plotly, Tableau, matplotlib, seabornDevOps & Tooling Git, Docker, CI/CD, CloudWatch, infrastructure-as-codeOur Hiring ProcessWe believe in a transparent and thorough selection process that respects your time while ensuring mutual fit:1. Initial Screening Call We’ll discuss your background, experience, and career goals, while providing an overview of the role and our team culture.2. Technical Challenge (issued on a case-by-case basis) You’ll receive a real-world technical challenge to complete within a specified timeframe. We encourage you to leverage all available resources—including AI tools, documentation, and libraries—just as you would in a production environment. This reflects how we actually work and allows you to showcase your problem-solving approach.3. Technical Interview We’ll have an in-depth discussion about your solution and explore related technical concepts. You should be prepared to walk through every aspect of your submission—explaining architectural decisions, code logic, trade-offs, and potential improvements. Whether you wrote specific code sections manually or generated them with AI assistance, you must demonstrate complete ownership and understanding of the entire codebase. This is a production-level assessment: we expect you to discuss, debug, and defend your work as if it were going live tomorrow.We’re looking for engineers who can think critically, adapt their approach, and truly understand the systems they build—not just those who can generate code.Ready to apply? We look forward to hearing from you!MedLaunch is an equal opportunity employer committed to diversity and inclusion.