{"schemaVersion":"jobsearcher.job.v1","id":"d16d2d07f144147e28d76da1","url":"https://jobsearcher.com/jobs/d16d2d07f144147e28d76da1","canonicalUrl":"https://jobsearcher.com/jobs/d16d2d07f144147e28d76da1","title":"Research Scientist","description":"If an AI remembers everything two people have said, does it truly understand their relationship?\nIt may recall the facts but miss what mattered. It may sound insightful while making unsupported assumptions. It may tell the user what they want to hear—and mistake agreement for understanding.\nAiloha is building a new kind of intelligence for human relationships: one that captures details before they disappear, understands the context between people, and surfaces what matters at the right moment.\nWe are looking for a researcher who can turn that ambition into something measurable:\nWhat does it mean for an agent to understand a person—and the situation between people—and how can we test it reliably?\nWhat You’ll Own\nIn the first phase, your focus will be agent evaluation and harness design , not model training.\nYou will:\nDefine evaluation dimensions for relationship-signal extraction, memory fidelity, insight depth, shared-context reconstruction, and sycophancy resistance\nBuild benchmarks, rubrics, scenario suites, judges, and regression tests\nDesign an agent harness that makes trajectories, tool calls, memory updates, intermediate reasoning, and failures observable and reproducible\nCreate simulated and real-world relationship scenarios for testing long-horizon agent behavior\nDiagnose failures such as missed signals, memory drift, shallow interpretation, over-inference, and premature advice\nTurn product feedback into structured evaluation cases and measurable agent behaviors\nEstablish the evaluation foundation that can later support reward design and post-training\nThis is not a role for running a predefined eval suite. You will help decide what should be measured, how the agent should be tested, and what evidence counts as improvement.\nQuestions You May Work On\nWhat separates relationship understanding from factual recall?\nHow do we evaluate an agent across an entire trajectory rather than a single response?\nHow should memory, tools, environment state, and user feedback be represented in the harness?\nWhat distinguishes a valuable insight from a persuasive hallucination?\nWhen does empathy become sycophancy?\nCan an offline scenario predict agent behavior in real, long-term interactions?\nWhich signals could eventually become useful rewards without encouraging shortcuts?\nWhat We’re Looking For\nExperience with agents, LLM evaluation, agent harnesses, model behavior analysis, or benchmark design\nStrong Python skills and experience building reliable research infrastructure\nFamiliarity with modern LLM inference frameworks, tool use, memory systems, and multi-step agent workflows\nA solid understanding of reinforcement learning, including environments, trajectories, reward design, credit assignment, and policy behavior\nAbility to formulate hypotheses, design experiments, analyze failures, and communicate findings clearly\nGenuine curiosity about how people understand other people\nYou do not need to have trained frontier models yourself. You do need to understand RL well enough to ensure that today’s evaluations can become tomorrow’s learning signals.\nWe welcome candidates from computer science, psychology, cognitive science, linguistics, sociology, philosophy, and related fields—as long as you can turn ambiguous human questions into rigorous, working systems.\nMainstream AI teams evaluate whether an agent completed a task.\nAt Ailoha, we are asking a harder question:\nAgent Evals & Harnesses\nHuman Understanding & Relationship Intelligence\nIf an AI remembers everything two people have said, does it truly understand their relationship?\nNot necessarily.\nIt may recall the facts but miss what mattered. It may sound insightful while making unsupported assumptions. It may tell the user what they want to hear—and mistake agreement for understanding.\nAiloha is building a new kind of intelligence for human relationships: one that captures details before they disappear, understands the context between people, and surfaces what matters at the right moment.\nWe are looking for a researcher who can turn that ambition into something measurable:\nWhat does it mean for an agent to understand a person—and the situation between people—and how can we test it reliably?\nWhat You’ll Own\nIn the first phase, your focus will be agent evaluation and harness design , not model training.\nYou will:\nDefine evaluation dimensions for relationship-signal extraction, memory fidelity, insight depth, shared-context reconstruction, and sycophancy resistance\nBuild benchmarks, rubrics, scenario suites, judges, and regression tests\nDesign an agent harness that makes trajectories, tool calls, memory updates, intermediate reasoning, and failures observable and reproducible\nCreate simulated and real-world relationship scenarios for testing long-horizon agent behavior\nDiagnose failures such as missed signals, memory drift, shallow interpretation, over-inference, and premature advice\nTurn product feedback into structured evaluation cases and measurable agent behaviors\nEstablish the evaluation foundation that can later support reward design and post-training\nThis is not a role for running a predefined eval suite. You will help decide what should be measured, how the agent should be tested, and what evidence counts as improvement.\nQuestions You May Work On\nWhat separates relationship understanding from factual recall?\nHow do we evaluate an agent across an entire trajectory rather than a single response?\nHow should memory, tools, environment state, and user feedback be represented in the harness?\nWhat distinguishes a valuable insight from a persuasive hallucination?\nWhen does empathy become sycophancy?\nCan an offline scenario predict agent behavior in real, long-term interactions?\nWhich signals could eventually become useful rewards without encouraging shortcuts?\nWhat We’re Looking For\nExperience with agents, LLM evaluation, agent harnesses, model behavior analysis, or benchmark design\nStrong Python skills and experience building reliable research infrastructure\nFamiliarity with modern LLM inference frameworks, tool use, memory systems, and multi-step agent workflows\nA solid understanding of reinforcement learning, including environments, trajectories, reward design, credit assignment, and policy behavior\nAbility to formulate hypotheses, design experiments, analyze failures, and communicate findings clearly\nGenuine curiosity about how people understand other people\nYou do not need to have trained frontier models yourself. You do need to understand RL well enough to ensure that today’s evaluations can become tomorrow’s learning signals.\nWe welcome candidates from computer science, psychology, cognitive science, linguistics, sociology, philosophy, and related fields—as long as you can turn ambiguous human questions into rigorous, working systems.\nAbout Ailoha！\nMainstream AI teams evaluate whether an agent completed a task.\nAt Ailoha, we are asking a harder question:\nDid the agent understand the human situation well enough to act appropriately? You will define the harness, evaluation language, and failure taxonomy for that question—and work directly with the founding, product, and engineering teams to bring it into a real product.\nIf this problem makes you want to design an environment, instrument a trajectory, or argue about what the reward is actually measuring, we would like to hear from you.\n\n#J-18808-Ljbffr","company":"Ailoha","rawCompany":"ailoha","city":"Millbrae","state":"CA","isRemote":false,"isActive":true,"createdAt":"2026-09-16T03:46:36.238Z","occupations":[{"code":"15-1221.00","title":"Computer and Information Research Scientists","slug":"computer-and-information-research-scientists"},{"code":"19-4061.00","title":"Social Science Research Assistants","slug":"social-science-research-assistants"},{"code":"19-3099.00","title":"Social Scientists and Related Workers, All Other","slug":"social-scientists-and-related-workers-all-other"}],"industries":[{"code":"541720","title":"Research and Development in the Social Sciences and Humanities","slug":"research-and-development-in-the-social-sciences-and-humanities"},{"code":"541910","title":"Marketing Research and Public Opinion Polling","slug":"marketing-research-and-public-opinion-polling"},{"code":"541715","title":"Research and Development in the Physical, Engineering, and Life Sciences (except Nanotechnology and Biotechnology)","slug":"research-and-development-in-the-physical-engineering-and-life-sciences-except-nanotechnology-and-biotechnology"}],"jobPosting":{"@context":"https://schema.org","@type":"JobPosting","title":"Research Scientist","description":"If an AI remembers everything two people have said, does it truly understand their relationship?\nIt may recall the facts but miss what mattered. It may sound insightful while making unsupported assumptions. It may tell the user what they want to hear—and mistake agreement for understanding.\nAiloha is building a new kind of intelligence for human relationships: one that captures details before they disappear, understands the context between people, and surfaces what matters at the right moment.\nWe are looking for a researcher who can turn that ambition into something measurable:\nWhat does it mean for an agent to understand a person—and the situation between people—and how can we test it reliably?\nWhat You’ll Own\nIn the first phase, your focus will be agent evaluation and harness design , not model training.\nYou will:\nDefine evaluation dimensions for relationship-signal extraction, memory fidelity, insight depth, shared-context reconstruction, and sycophancy resistance\nBuild benchmarks, rubrics, scenario suites, judges, and regression tests\nDesign an agent harness that makes trajectories, tool calls, memory updates, intermediate reasoning, and failures observable and reproducible\nCreate simulated and real-world relationship scenarios for testing long-horizon agent behavior\nDiagnose failures such as missed signals, memory drift, shallow interpretation, over-inference, and premature advice\nTurn product feedback into structured evaluation cases and measurable agent behaviors\nEstablish the evaluation foundation that can later support reward design and post-training\nThis is not a role for running a predefined eval suite. You will help decide what should be measured, how the agent should be tested, and what evidence counts as improvement.\nQuestions You May Work On\nWhat separates relationship understanding from factual recall?\nHow do we evaluate an agent across an entire trajectory rather than a single response?\nHow should memory, tools, environment state, and user feedback be represented in the harness?\nWhat distinguishes a valuable insight from a persuasive hallucination?\nWhen does empathy become sycophancy?\nCan an offline scenario predict agent behavior in real, long-term interactions?\nWhich signals could eventually become useful rewards without encouraging shortcuts?\nWhat We’re Looking For\nExperience with agents, LLM evaluation, agent harnesses, model behavior analysis, or benchmark design\nStrong Python skills and experience building reliable research infrastructure\nFamiliarity with modern LLM inference frameworks, tool use, memory systems, and multi-step agent workflows\nA solid understanding of reinforcement learning, including environments, trajectories, reward design, credit assignment, and policy behavior\nAbility to formulate hypotheses, design experiments, analyze failures, and communicate findings clearly\nGenuine curiosity about how people understand other people\nYou do not need to have trained frontier models yourself. You do need to understand RL well enough to ensure that today’s evaluations can become tomorrow’s learning signals.\nWe welcome candidates from computer science, psychology, cognitive science, linguistics, sociology, philosophy, and related fields—as long as you can turn ambiguous human questions into rigorous, working systems.\nMainstream AI teams evaluate whether an agent completed a task.\nAt Ailoha, we are asking a harder question:\nAgent Evals & Harnesses\nHuman Understanding & Relationship Intelligence\nIf an AI remembers everything two people have said, does it truly understand their relationship?\nNot necessarily.\nIt may recall the facts but miss what mattered. It may sound insightful while making unsupported assumptions. It may tell the user what they want to hear—and mistake agreement for understanding.\nAiloha is building a new kind of intelligence for human relationships: one that captures details before they disappear, understands the context between people, and surfaces what matters at the right moment.\nWe are looking for a researcher who can turn that ambition into something measurable:\nWhat does it mean for an agent to understand a person—and the situation between people—and how can we test it reliably?\nWhat You’ll Own\nIn the first phase, your focus will be agent evaluation and harness design , not model training.\nYou will:\nDefine evaluation dimensions for relationship-signal extraction, memory fidelity, insight depth, shared-context reconstruction, and sycophancy resistance\nBuild benchmarks, rubrics, scenario suites, judges, and regression tests\nDesign an agent harness that makes trajectories, tool calls, memory updates, intermediate reasoning, and failures observable and reproducible\nCreate simulated and real-world relationship scenarios for testing long-horizon agent behavior\nDiagnose failures such as missed signals, memory drift, shallow interpretation, over-inference, and premature advice\nTurn product feedback into structured evaluation cases and measurable agent behaviors\nEstablish the evaluation foundation that can later support reward design and post-training\nThis is not a role for running a predefined eval suite. You will help decide what should be measured, how the agent should be tested, and what evidence counts as improvement.\nQuestions You May Work On\nWhat separates relationship understanding from factual recall?\nHow do we evaluate an agent across an entire trajectory rather than a single response?\nHow should memory, tools, environment state, and user feedback be represented in the harness?\nWhat distinguishes a valuable insight from a persuasive hallucination?\nWhen does empathy become sycophancy?\nCan an offline scenario predict agent behavior in real, long-term interactions?\nWhich signals could eventually become useful rewards without encouraging shortcuts?\nWhat We’re Looking For\nExperience with agents, LLM evaluation, agent harnesses, model behavior analysis, or benchmark design\nStrong Python skills and experience building reliable research infrastructure\nFamiliarity with modern LLM inference frameworks, tool use, memory systems, and multi-step agent workflows\nA solid understanding of reinforcement learning, including environments, trajectories, reward design, credit assignment, and policy behavior\nAbility to formulate hypotheses, design experiments, analyze failures, and communicate findings clearly\nGenuine curiosity about how people understand other people\nYou do not need to have trained frontier models yourself. You do need to understand RL well enough to ensure that today’s evaluations can become tomorrow’s learning signals.\nWe welcome candidates from computer science, psychology, cognitive science, linguistics, sociology, philosophy, and related fields—as long as you can turn ambiguous human questions into rigorous, working systems.\nAbout Ailoha！\nMainstream AI teams evaluate whether an agent completed a task.\nAt Ailoha, we are asking a harder question:\nDid the agent understand the human situation well enough to act appropriately? You will define the harness, evaluation language, and failure taxonomy for that question—and work directly with the founding, product, and engineering teams to bring it into a real product.\nIf this problem makes you want to design an environment, instrument a trajectory, or argue about what the reward is actually measuring, we would like to hear from you.\n\n#J-18808-Ljbffr","datePosted":"2026-09-16T03:46:36.238Z","dateModified":"2026-09-16T03:46:36.238Z","hiringOrganization":{"@type":"Organization","name":"Ailoha","sameAs":"https://jobsearcher.com"},"jobLocation":{"@type":"Place","address":{"@type":"PostalAddress","addressLocality":"Millbrae","addressRegion":"CA","addressCountry":"US"}},"identifier":{"@type":"PropertyValue","name":"JobSearcher","value":"d16d2d07f144147e28d76da1"},"url":"https://jobsearcher.com/jobs/d16d2d07f144147e28d76da1"}}