Research Scientist
If an AI remembers everything two people have said, does it truly understand their relationship?
It may recall the facts but miss what mattered. It may sound insightful while making unsupported assumptions. It may tell the user what they want to hear—and mistake agreement for understanding.
Ailoha is building a new kind of intelligence for human relationships: one that captures details before they disappear, understands the context between people, and surfaces what matters at the right moment.
We are looking for a researcher who can turn that ambition into something measurable:
What does it mean for an agent to understand a person—and the situation between people—and how can we test it reliably?
What You’ll Own
In the first phase, your focus will be agent evaluation and harness design , not model training.
You will:
Define evaluation dimensions for relationship-signal extraction, memory fidelity, insight depth, shared-context reconstruction, and sycophancy resistance
Build benchmarks, rubrics, scenario suites, judges, and regression tests
Design an agent harness that makes trajectories, tool calls, memory updates, intermediate reasoning, and failures observable and reproducible
Create simulated and real-world relationship scenarios for testing long-horizon agent behavior
Diagnose failures such as missed signals, memory drift, shallow interpretation, over-inference, and premature advice
Turn product feedback into structured evaluation cases and measurable agent behaviors
Establish the evaluation foundation that can later support reward design and post-training
This is not a role for running a predefined eval suite. You will help decide what should be measured, how the agent should be tested, and what evidence counts as improvement.
Questions You May Work On
What separates relationship understanding from factual recall?
How do we evaluate an agent across an entire trajectory rather than a single response?
How should memory, tools, environment state, and user feedback be represented in the harness?
What distinguishes a valuable insight from a persuasive hallucination?
When does empathy become sycophancy?
Can an offline scenario predict agent behavior in real, long-term interactions?
Which signals could eventually become useful rewards without encouraging shortcuts?
What We’re Looking For
Experience with agents, LLM evaluation, agent harnesses, model behavior analysis, or benchmark design
Strong Python skills and experience building reliable research infrastructure
Familiarity with modern LLM inference frameworks, tool use, memory systems, and multi-step agent workflows
A solid understanding of reinforcement learning, including environments, trajectories, reward design, credit assignment, and policy behavior
Ability to formulate hypotheses, design experiments, analyze failures, and communicate findings clearly
Genuine curiosity about how people understand other people
You do not need to have trained frontier models yourself. You do need to understand RL well enough to ensure that today’s evaluations can become tomorrow’s learning signals.
We welcome candidates from computer science, psychology, cognitive science, linguistics, sociology, philosophy, and related fields—as long as you can turn ambiguous human questions into rigorous, working systems.
Mainstream AI teams evaluate whether an agent completed a task.
At Ailoha, we are asking a harder question:
Agent Evals & Harnesses
Human Understanding & Relationship Intelligence
If an AI remembers everything two people have said, does it truly understand their relationship?
Not necessarily.
It may recall the facts but miss what mattered. It may sound insightful while making unsupported assumptions. It may tell the user what they want to hear—and mistake agreement for understanding.
Ailoha is building a new kind of intelligence for human relationships: one that captures details before they disappear, understands the context between people, and surfaces what matters at the right moment.
We are looking for a researcher who can turn that ambition into something measurable:
What does it mean for an agent to understand a person—and the situation between people—and how can we test it reliably?
What You’ll Own
In the first phase, your focus will be agent evaluation and harness design , not model training.
You will:
Define evaluation dimensions for relationship-signal extraction, memory fidelity, insight depth, shared-context reconstruction, and sycophancy resistance
Build benchmarks, rubrics, scenario suites, judges, and regression tests
Design an agent harness that makes trajectories, tool calls, memory updates, intermediate reasoning, and failures observable and reproducible
Create simulated and real-world relationship scenarios for testing long-horizon agent behavior
Diagnose failures such as missed signals, memory drift, shallow interpretation, over-inference, and premature advice
Turn product feedback into structured evaluation cases and measurable agent behaviors
Establish the evaluation foundation that can later support reward design and post-training
This is not a role for running a predefined eval suite. You will help decide what should be measured, how the agent should be tested, and what evidence counts as improvement.
Questions You May Work On
What separates relationship understanding from factual recall?
How do we evaluate an agent across an entire trajectory rather than a single response?
How should memory, tools, environment state, and user feedback be represented in the harness?
What distinguishes a valuable insight from a persuasive hallucination?
When does empathy become sycophancy?
Can an offline scenario predict agent behavior in real, long-term interactions?
Which signals could eventually become useful rewards without encouraging shortcuts?
What We’re Looking For
Experience with agents, LLM evaluation, agent harnesses, model behavior analysis, or benchmark design
Strong Python skills and experience building reliable research infrastructure
Familiarity with modern LLM inference frameworks, tool use, memory systems, and multi-step agent workflows
A solid understanding of reinforcement learning, including environments, trajectories, reward design, credit assignment, and policy behavior
Ability to formulate hypotheses, design experiments, analyze failures, and communicate findings clearly
Genuine curiosity about how people understand other people
You do not need to have trained frontier models yourself. You do need to understand RL well enough to ensure that today’s evaluations can become tomorrow’s learning signals.
We welcome candidates from computer science, psychology, cognitive science, linguistics, sociology, philosophy, and related fields—as long as you can turn ambiguous human questions into rigorous, working systems.
About Ailoha!
Mainstream AI teams evaluate whether an agent completed a task.
At Ailoha, we are asking a harder question:
Did the agent understand the human situation well enough to act appropriately? You will define the harness, evaluation language, and failure taxonomy for that question—and work directly with the founding, product, and engineering teams to bring it into a real product.
If this problem makes you want to design an environment, instrument a trajectory, or argue about what the reward is actually measuring, we would like to hear from you.
#J-18808-Ljbffr