{"schemaVersion":"jobsearcher.job.v1","id":"0b5dad9e55a4bd72a8975f59","url":"https://jobsearcher.com/jobs/0b5dad9e55a4bd72a8975f59","canonicalUrl":"https://jobsearcher.com/jobs/0b5dad9e55a4bd72a8975f59","title":"Data Engineering Intern (Python)","description":"CompanyCanaria is a technology startup providing large-scale labor market data to the B2B market, powering recruitment, HR analytics, workforce planning, healthcare, research, and finance. Our data mining, computing optimizations, and NLP process over 1 billion unique job postings from 200,000+ data sources, with 100+ fields per record.About the PositionNote: This is an unpaid internship position.A Python-first role on the same production pipelines and crawling system our full-time engineers run, with mentorship from a senior engineer. You ship real code on real data, not a sandbox project.What You Will DoBuild and improve Python data pipelines (ingestion, parsing, cleaning, semantic deduplication, enrichment) over billions of job postings.Add new source integrations to our in-house distributed crawler (custom scheduling, queueing, proxy/session management, and anti-blocking).Work across PostgreSQL, MongoDB, Redis and Aerospike at a scale where query plans and storage layout matter.Test, containerize, and deploy what you build with Docker on AWS/GCP.Who You Are (Required)Currently pursuing a Bachelor's or Master's in Computer Science, Computer Engineering, Software Engineering, or a closely related computing field.Python is your strongest language; comfortable with Unix/Linux and SQL.You use AI coding tools (Cursor, Claude Code, GitHub Copilot) in your actual workflow, not just for chat.Authorized to work in the US (we cannot sponsor for interns) and able to overlap with US business hours.Nice to HaveA scraping or data-pipeline project you built yourself (Scrapy, Playwright, asyncio, pandas).Docker, a cloud platform, or a NoSQL database (MongoDB, Redis).What You GetMentorship and production experience at billion-record scale.Founders and colleagues with FAANG and top-institution backgrounds, across US and EU teams.Access to powerful on-prem and cloud compute.Hiring ProcessFirst call about the role & expectations (15 mins)Basic Python coding interview (1 hour)Technical interview (1 hour)Only shortlisted candidates will be contacted.","company":"Canaria","rawCompany":"canaria","city":"Denver","state":"CO","isRemote":false,"isActive":false,"createdAt":"2026-09-29T10:07:27.927Z","occupations":[{"code":"15-1243.01","title":"Data Warehousing Specialists","slug":"data-warehousing-specialists"},{"code":"15-2051.00","title":"Data Scientists","slug":"data-scientists"},{"code":"15-1252.00","title":"Software Developers","slug":"software-developers"}],"industries":[{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"},{"code":"541512","title":"Computer Systems Design Services","slug":"computer-systems-design-services"},{"code":"513210","title":"Software Publishers","slug":"software-publishers"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Data Engineering Intern (Python)","description":"CompanyCanaria is a technology startup providing large-scale labor market data to the B2B market, powering recruitment, HR analytics, workforce planning, healthcare, research, and finance. Our data mining, computing optimizations, and NLP process over 1 billion unique job postings from 200,000+ data sources, with 100+ fields per record.About the PositionNote: This is an unpaid internship position.A Python-first role on the same production pipelines and crawling system our full-time engineers run, with mentorship from a senior engineer. You ship real code on real data, not a sandbox project.What You Will DoBuild and improve Python data pipelines (ingestion, parsing, cleaning, semantic deduplication, enrichment) over billions of job postings.Add new source integrations to our in-house distributed crawler (custom scheduling, queueing, proxy/session management, and anti-blocking).Work across PostgreSQL, MongoDB, Redis and Aerospike at a scale where query plans and storage layout matter.Test, containerize, and deploy what you build with Docker on AWS/GCP.Who You Are (Required)Currently pursuing a Bachelor's or Master's in Computer Science, Computer Engineering, Software Engineering, or a closely related computing field.Python is your strongest language; comfortable with Unix/Linux and SQL.You use AI coding tools (Cursor, Claude Code, GitHub Copilot) in your actual workflow, not just for chat.Authorized to work in the US (we cannot sponsor for interns) and able to overlap with US business hours.Nice to HaveA scraping or data-pipeline project you built yourself (Scrapy, Playwright, asyncio, pandas).Docker, a cloud platform, or a NoSQL database (MongoDB, Redis).What You GetMentorship and production experience at billion-record scale.Founders and colleagues with FAANG and top-institution backgrounds, across US and EU teams.Access to powerful on-prem and cloud compute.Hiring ProcessFirst call about the role & expectations (15 mins)Basic Python coding interview (1 hour)Technical interview (1 hour)Only shortlisted candidates will be contacted.","datePosted":"2026-09-29T10:07:27.927Z","dateModified":"2026-09-29T10:07:27.927Z","hiringOrganization":{"@type":"Organization","name":"Canaria","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Denver","addressRegion":"CO","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"0b5dad9e55a4bd72a8975f59"},"url":"https://jobsearcher.com/jobs/0b5dad9e55a4bd72a8975f59"}}