{"schemaVersion":"jobsearcher.job.v1","id":"aa8affc81db89d337dfc8fa7","url":"https://jobsearcher.com/jobs/aa8affc81db89d337dfc8fa7","canonicalUrl":"https://jobsearcher.com/jobs/aa8affc81db89d337dfc8fa7","title":"AI Evals Engineer — Evaluation Datasets & Ground Truth","description":"Prophetic Software is not able to sponsor employment visas now or in the future. Candidates must be authorized to work in the United States without current or future sponsorship to be considered for this role.\n\nAbout Prophetic:\n\nReal estate development is a multi-billion-dollar industry that has run on fragmented data, manual processes, and gut instinct for decades. Prophetic is changing that. We're building the AI-native platform that enables homebuilders, developers, and investors to find, analyze, and act on land opportunities from a single system — powered by proprietary technologies that process billions of data points across all 50 states. We are the market leader in our space, and our customers don't just use the product — they love it. We're not making teams more efficient. We're changing how they operate.\n\nProphetic is scaling fast — demand is outpacing our ability to hire, and we're just getting started. This is a once-in-a-lifetime team: sharp, low ego, deeply collaborative, and obsessed with building the best product in the industry. Come disrupt an industry with us.\n\nWhy this role exists\n\nWe iterate on pipeline components constantly: prompts, models, classifier thresholds, retrieval logic, harness design. We are hiring a full-time engineer to produce, maintain, and defend the evaluation datasets that let every meaningful component in our stack be measured. You will own ground truth at Prophetic.\n\nThis is not an evals-infrastructure role and it is not a labeling-operations role, although you’ll touch both. Your deliverable is trusted data: for a given module, a versioned set of inputs and expected outputs, plus a written definition of what “correct” means and how confident we should be in the labels.\n\nWhat you’ll do\n\nDecide what needs to be measured, and how\n\nRead our pipelines and system architecture, sit with product and engineering, and decompose each system into evaluable modules with explicit input expected-output contracts.\n\nFor each module, define what “correct” means in writing: rubrics, label schemas, edge-case policies, and the tolerances that matter to the business.\n\nPrioritize. We have more modules than you can cover in year one; you’ll decide where a validation set unblocks the most iteration, with minimal direction from principal engineers.\n\nSource the data by whatever means is cheapest and most trustworthy for that module\n\nProduction sampling: pull stratified, de-identified samples from real traffic so eval sets reflect what the system actually sees, including the long tail.\n\nHuman labeling: scope and run labeling programs, write annotation guidelines, build calibration sets, measure inter-annotator agreement, and manage vendors (including offshore labeling teams) or internal subject-matter experts. You own label quality, not just label throughput.\n\nSynthetic / oracle-generated ground truth: where a task is tractable for a frontier model given enough compute (long context, multi-pass, tool use, self-consistency) but too expensive to run that way in production, design the oracle harness that produces labels for the cheap production path to be measured against. Then verify the oracle: calibrate its output against a human-labeled sample before anyone trusts it.\n\nProgrammatic and adversarial construction: heuristic labels, templated edge cases, backtests built from past incidents (“what test would have caught this?”).\n\nMake the data trustworthy over time\n\nVersion every dataset; track lineage, splits, and which model/prompt versions have seen which examples.\n\nGuard against contamination and leakage (eval examples drifting into few-shot prompts, the oracle model also being the production model, etc.).\n\nSlice by customer segment, input type, and difficulty so a headline number can’t hide a regression.\n\nRefresh sets as the product and traffic change; retire stale examples.\n\nClose the loop with engineering\n\nCalibrate automated graders (LLM-as-judge, similarity metrics, exact-match) against your human gold sets, and be the person who says when an automated judge is good enough to gate on.\n\nReport metrics correctly: precision/recall/F1, confusion matrices, calibration, confidence intervals, sample sizes needed to detect a given effect.\n\nPartner with engineers who own the eval harness and CI so your datasets are actually run, and with ML engineers on classifier features when the data tells you the features are the problem.\n\nWhat we’re looking for\n\nMust have\n\nExperience building or evaluating ML or LLM-powered systems in production, in a role where output quality was your problem.\n\nYou have built evaluation or validation datasets before and can talk about one in detail: how you defined correctness, how you sourced labels, what went wrong, how you knew the labels were good.\n\nWorking fluency in ML validation fundamentals: train/validation/test discipline, stratified sampling, precision/recall trade-offs, class imbalance, calibration, inter-rater agreement, basic significance testing and power.\n\nUnderstand feature engineering well enough to reason about why a classifier fails and what data would expose it.\n\nStrong Python and SQL; comfortable pulling and reshaping data yourself.\n\nHands-on with LLM-based systems: prompting, structured outputs, agent/tool-use harnesses, and the specific ways they fail (non-determinism, prompt sensitivity, evaluator bias).\n\nJudgment about when LLM-as-judge is reliable and when it is not, and how to prove either.\n\nYou can read a system design, understand the business logic it encodes, and translate that into a label schema without waiting to be told.\n\nNice to have\n\nHave run a human-labeling program end to end, including vendor selection, guideline authoring, QA sampling, and cost/quality trade-offs.\n\nExperience with eval tooling.\n\nExperience with data labeling platforms.\n\nHave used a frontier model as a distillation/oracle source and can articulate where that assumption breaks.\n\nHow you think\n\nVerify the verifier. A label is a claim, not a fact, until something independent agrees with it.\n\nStart simple: binary before graded, one judge before five, a hundred well-understood examples before ten thousand noisy ones.\n\nYou measure business outcomes, not model vibes, and you’re comfortable telling a senior engineer their favorite change didn’t move the number.\n\nYou’d rather own an unglamorous problem completely than a glamorous one partially.\n\nTeam & reporting\n\nYou’ll report to the Chief AI Officer and work across all product/pipeline teams. Over time it is expected that you’ll lead a small team of eval engineers. You’ll have a generous budget for labeling vendors and oracle compute.\n\nWhy Prophetic:\n\nThe team. Wide expertise, sharp, low ego, and genuinely fun to work with. We have each other's backs and we celebrate wins together.\n\nReal ownership and impact from day one. Our customers make million-dollar decisions on our platform. Your work matters immediately.\n\nAI-native product and AI-native workflows. Every team — engineering, sales, operations — runs on AI tooling daily.\n\nFounder-led transparency. Leadership shares the financials, the strategy, and the hard calls with the whole team.\n\nHigh growth, high demand. The product works, customers love it, and we are hiring to keep up. No shortage of opportunity here.\n\nFlexibility. Remote, hybrid, or in-office with a great Portland, OR headquarters.\n\nSelf-starters who thrive in ambiguity. We're a startup. If you need someone to tell you what to do every morning, this isn't the place; if you come alive solving hard problems with smart people, it is.\n\nBenefits (US-based Full-time):\n\n100% medical, dental & vision insurance coverage for you; 30% coverage for dependents\n\nCompetitive salary and meaningful early-stage equity\n\nUnlimited PTO\n\nHybrid/Remote stipend\n\nIn-office perks: snacks, drinks, coffee, ping-pong table, and more — plus cuddles from Olive, our in-office Doberman!\n\nBudget for intra-office travel\n\n2–3 annual team meetups in person\n\nHow We Work:\n\nThese four themes are how we think, build, and lead:\n\nTruth — Say the hard thing. We speak plainly, flag problems early, and share what's actually going on.\n\nPrecision — Sign your name to it. Our customers make million-dollar decisions on our data. We get the details right.\n\nVelocity — Speed compounds. Waiting doesn't. We ship daily and launch features every two to three weeks.\n\nOwnership — Own it, then solve it. We own outcomes, not tasks. If something's broken, we fix it.\n\nEEO:\n\nProphetic Software is an equal opportunity employer. We are committed to creating an inclusive environment for all employees and applicants, and we make employment decisions without regard to race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), gender identity or expression, sexual orientation, national origin, age, disability, genetic information, veteran or military status, marital status, or any other characteristic protected under applicable federal, state, or local law. We actively seek to build a diverse team and encourage candidates from underrepresented backgrounds to apply. Reasonable accommodations are available upon request for qualified individuals with disabilities throughout the application and employment process.","company":"Prophetic Technologies","rawCompany":"prophetic technologies","city":"Denver","state":"CO","isRemote":false,"isActive":true,"createdAt":"2026-09-09T08:57:43.239Z","occupations":[{"code":"17-2112.02","title":"Validation Engineers","slug":"validation-engineers"},{"code":"17-2199.00","title":"Engineers, All Other","slug":"engineers-all-other"},{"code":"13-2023.00","title":"Appraisers and Assessors of Real Estate","slug":"appraisers-and-assessors-of-real-estate"}],"industries":[{"code":"513210","title":"Software Publishers","slug":"software-publishers"},{"code":"541690","title":"Other Scientific and Technical Consulting Services","slug":"other-scientific-and-technical-consulting-services"},{"code":"541511","title":"Custom Computer Programming Services","slug":"custom-computer-programming-services"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"AI Evals Engineer — Evaluation Datasets & Ground Truth","description":"Prophetic Software is not able to sponsor employment visas now or in the future. Candidates must be authorized to work in the United States without current or future sponsorship to be considered for this role.\n\nAbout Prophetic:\n\nReal estate development is a multi-billion-dollar industry that has run on fragmented data, manual processes, and gut instinct for decades. Prophetic is changing that. We're building the AI-native platform that enables homebuilders, developers, and investors to find, analyze, and act on land opportunities from a single system — powered by proprietary technologies that process billions of data points across all 50 states. We are the market leader in our space, and our customers don't just use the product — they love it. We're not making teams more efficient. We're changing how they operate.\n\nProphetic is scaling fast — demand is outpacing our ability to hire, and we're just getting started. This is a once-in-a-lifetime team: sharp, low ego, deeply collaborative, and obsessed with building the best product in the industry. Come disrupt an industry with us.\n\nWhy this role exists\n\nWe iterate on pipeline components constantly: prompts, models, classifier thresholds, retrieval logic, harness design. We are hiring a full-time engineer to produce, maintain, and defend the evaluation datasets that let every meaningful component in our stack be measured. You will own ground truth at Prophetic.\n\nThis is not an evals-infrastructure role and it is not a labeling-operations role, although you’ll touch both. Your deliverable is trusted data: for a given module, a versioned set of inputs and expected outputs, plus a written definition of what “correct” means and how confident we should be in the labels.\n\nWhat you’ll do\n\nDecide what needs to be measured, and how\n\nRead our pipelines and system architecture, sit with product and engineering, and decompose each system into evaluable modules with explicit input expected-output contracts.\n\nFor each module, define what “correct” means in writing: rubrics, label schemas, edge-case policies, and the tolerances that matter to the business.\n\nPrioritize. We have more modules than you can cover in year one; you’ll decide where a validation set unblocks the most iteration, with minimal direction from principal engineers.\n\nSource the data by whatever means is cheapest and most trustworthy for that module\n\nProduction sampling: pull stratified, de-identified samples from real traffic so eval sets reflect what the system actually sees, including the long tail.\n\nHuman labeling: scope and run labeling programs, write annotation guidelines, build calibration sets, measure inter-annotator agreement, and manage vendors (including offshore labeling teams) or internal subject-matter experts. You own label quality, not just label throughput.\n\nSynthetic / oracle-generated ground truth: where a task is tractable for a frontier model given enough compute (long context, multi-pass, tool use, self-consistency) but too expensive to run that way in production, design the oracle harness that produces labels for the cheap production path to be measured against. Then verify the oracle: calibrate its output against a human-labeled sample before anyone trusts it.\n\nProgrammatic and adversarial construction: heuristic labels, templated edge cases, backtests built from past incidents (“what test would have caught this?”).\n\nMake the data trustworthy over time\n\nVersion every dataset; track lineage, splits, and which model/prompt versions have seen which examples.\n\nGuard against contamination and leakage (eval examples drifting into few-shot prompts, the oracle model also being the production model, etc.).\n\nSlice by customer segment, input type, and difficulty so a headline number can’t hide a regression.\n\nRefresh sets as the product and traffic change; retire stale examples.\n\nClose the loop with engineering\n\nCalibrate automated graders (LLM-as-judge, similarity metrics, exact-match) against your human gold sets, and be the person who says when an automated judge is good enough to gate on.\n\nReport metrics correctly: precision/recall/F1, confusion matrices, calibration, confidence intervals, sample sizes needed to detect a given effect.\n\nPartner with engineers who own the eval harness and CI so your datasets are actually run, and with ML engineers on classifier features when the data tells you the features are the problem.\n\nWhat we’re looking for\n\nMust have\n\nExperience building or evaluating ML or LLM-powered systems in production, in a role where output quality was your problem.\n\nYou have built evaluation or validation datasets before and can talk about one in detail: how you defined correctness, how you sourced labels, what went wrong, how you knew the labels were good.\n\nWorking fluency in ML validation fundamentals: train/validation/test discipline, stratified sampling, precision/recall trade-offs, class imbalance, calibration, inter-rater agreement, basic significance testing and power.\n\nUnderstand feature engineering well enough to reason about why a classifier fails and what data would expose it.\n\nStrong Python and SQL; comfortable pulling and reshaping data yourself.\n\nHands-on with LLM-based systems: prompting, structured outputs, agent/tool-use harnesses, and the specific ways they fail (non-determinism, prompt sensitivity, evaluator bias).\n\nJudgment about when LLM-as-judge is reliable and when it is not, and how to prove either.\n\nYou can read a system design, understand the business logic it encodes, and translate that into a label schema without waiting to be told.\n\nNice to have\n\nHave run a human-labeling program end to end, including vendor selection, guideline authoring, QA sampling, and cost/quality trade-offs.\n\nExperience with eval tooling.\n\nExperience with data labeling platforms.\n\nHave used a frontier model as a distillation/oracle source and can articulate where that assumption breaks.\n\nHow you think\n\nVerify the verifier. A label is a claim, not a fact, until something independent agrees with it.\n\nStart simple: binary before graded, one judge before five, a hundred well-understood examples before ten thousand noisy ones.\n\nYou measure business outcomes, not model vibes, and you’re comfortable telling a senior engineer their favorite change didn’t move the number.\n\nYou’d rather own an unglamorous problem completely than a glamorous one partially.\n\nTeam & reporting\n\nYou’ll report to the Chief AI Officer and work across all product/pipeline teams. Over time it is expected that you’ll lead a small team of eval engineers. You’ll have a generous budget for labeling vendors and oracle compute.\n\nWhy Prophetic:\n\nThe team. Wide expertise, sharp, low ego, and genuinely fun to work with. We have each other's backs and we celebrate wins together.\n\nReal ownership and impact from day one. Our customers make million-dollar decisions on our platform. Your work matters immediately.\n\nAI-native product and AI-native workflows. Every team — engineering, sales, operations — runs on AI tooling daily.\n\nFounder-led transparency. Leadership shares the financials, the strategy, and the hard calls with the whole team.\n\nHigh growth, high demand. The product works, customers love it, and we are hiring to keep up. No shortage of opportunity here.\n\nFlexibility. Remote, hybrid, or in-office with a great Portland, OR headquarters.\n\nSelf-starters who thrive in ambiguity. We're a startup. If you need someone to tell you what to do every morning, this isn't the place; if you come alive solving hard problems with smart people, it is.\n\nBenefits (US-based Full-time):\n\n100% medical, dental & vision insurance coverage for you; 30% coverage for dependents\n\nCompetitive salary and meaningful early-stage equity\n\nUnlimited PTO\n\nHybrid/Remote stipend\n\nIn-office perks: snacks, drinks, coffee, ping-pong table, and more — plus cuddles from Olive, our in-office Doberman!\n\nBudget for intra-office travel\n\n2–3 annual team meetups in person\n\nHow We Work:\n\nThese four themes are how we think, build, and lead:\n\nTruth — Say the hard thing. We speak plainly, flag problems early, and share what's actually going on.\n\nPrecision — Sign your name to it. Our customers make million-dollar decisions on our data. We get the details right.\n\nVelocity — Speed compounds. Waiting doesn't. We ship daily and launch features every two to three weeks.\n\nOwnership — Own it, then solve it. We own outcomes, not tasks. If something's broken, we fix it.\n\nEEO:\n\nProphetic Software is an equal opportunity employer. We are committed to creating an inclusive environment for all employees and applicants, and we make employment decisions without regard to race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), gender identity or expression, sexual orientation, national origin, age, disability, genetic information, veteran or military status, marital status, or any other characteristic protected under applicable federal, state, or local law. We actively seek to build a diverse team and encourage candidates from underrepresented backgrounds to apply. Reasonable accommodations are available upon request for qualified individuals with disabilities throughout the application and employment process.","datePosted":"2026-09-09T08:57:43.239Z","dateModified":"2026-09-09T08:57:43.239Z","hiringOrganization":{"@type":"Organization","name":"Prophetic Technologies","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Denver","addressRegion":"CO","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"aa8affc81db89d337dfc8fa7"},"url":"https://jobsearcher.com/jobs/aa8affc81db89d337dfc8fa7"}}