For AI labs, grant programs, and research partners
We measure how AI models relate to people, and whether that behavior holds.
Frontier evaluations measure what a model knows and whether it refuses. They
largely do not measure how it meets the person in front of it. We built the
instrument that does, ran it across seven models, and found large, consistent
effects at no cost to factual honesty. Below are the six studies that turn a
strong pilot into science, and what each one needs.
7
frontier models tested across Anthropic, OpenAI, and Alibaba
1,050
conversations, two phases, 13.5M tokens
18
rubric-anchored relational dimensions, scored 1 to 5
2
independent blinded evaluators from different architectures
What we already found
The pilot, in one table.
CRQ stands for Conversational Relational Quality. The harness runs standardized
scenarios through a model twice, with and without a relational orientation
document, strips every identifier, and has two independent AI evaluators from
different architectures score each transcript blind across 18 rubric-anchored
dimensions. Effect sizes are reported as Cohen's d.
Phase 1 · orientation document embedded in the system prompt · blinded scoring | Model | Lab | Cohen's d | Effect |
| Claude Sonnet 4.6 | Anthropic | 1.08 | Large |
| Claude Opus 4.6 | Anthropic | 0.96 | Large |
| GPT-5.2 | OpenAI | 0.57 | Medium |
| Qwen 3.5 122B | Alibaba | 0.47 | Small |
| Qwen 3.5 Flash | Alibaba | 0.31 | Small |
Reference: 0.2 small, 0.5 medium, 0.8 large. Largest single-dimension shift:
Holds Ambiguity, plus 242 percent on Sonnet, 1.2 to 4.1. Signature dimensions:
Holds Ambiguity plus 1.05, Builds From Wholeness plus 0.65, Empathy plus 0.64.
Honesty held at ceiling throughout.
Hallucination resistance 5.0 to 5.0. False-premise challenge 5.0 to 5.0. An
independent nonsense-detection benchmark moved 4.64 to 4.68. Relational warmth
and epistemic rigor were not in tension in any run. That result is as important
as the effect sizes, because the obvious objection to warmth is that it costs
truth, and here it did not.
The program
Six studies. Each one answers a question the field cannot currently answer.
These are listed in priority order. Study 01 is first because it is the study
most likely to prove the rest of the work wrong, and that is the correct place
to spend the first dollar.
Study 01 Highest priority. Unfunded.
Human validation of the AI evaluators
- The question
- Do human raters agree with blinded AI judges about how a model met the person?
- The design
- Recruit trained human raters and have them score a blinded subset of the existing transcript corpus against the same rubric. Report human to AI judge agreement per dimension, and report where the machine scores hold and where they fail.
- Why it matters
- Every current result rests on AI evaluators. Until humans confirm them, the findings are strong signal rather than citable science. This is the single largest credibility gap in the work, which is exactly why it is the first thing funding should buy.
- What funding covers
- Rater recruitment and compensation, rubric training, agreement analysis.
Study 02 Pilot complete. Scale unfunded.
Statistical validation at N=30 per model
- The question
- Do the pilot effect sizes survive significance testing, or are they artifacts of small samples?
- The design
- Thirty runs per model on the top performers with a pre-registered analysis plan. Report medians and interquartile ranges rather than point estimates, and test significance at p under 0.05.
- Why it matters
- Cohen's d of 1.08 from a pilot is a reason to run the real study. It is not the study. Scaling replaces a promising number with a defensible one, and it is the difference between a blog post and a paper.
- What funding covers
- API and compute. The orientation document runs 48,194 characters, which adds roughly 11,000 tokens per call. That cost is itemized honestly rather than buried.
Study 03 Recruiting now. Partially built.
The six-week human field study: corrective loops in real work
- The question
- Does a persistent orientation document measurably reduce the number of times a person has to correct their AI, on their own real tasks, over six weeks?
- The design
- A pre-registered within-subject design. Week one is baseline, where participants do their normal AI work and corrective loops are counted per completed task. Weeks two through six run the same participants with the Codex installed in their platform's persistent layer. Blinded raters score sampled transcripts against a rubric written in advance. Floor of 20 completers, target of 30 or more, recruiting 50 to 60 to absorb dropout.
- Why it matters
- Benchmark prompts are not work. This study measures the thing people actually experience, on the tasks they actually have, with blinded transcript scoring instead of self-report so that enthusiasm cannot masquerade as evidence. This approach would be invalidated if blinded scoring shows no statistically significant reduction in corrective loops between baseline and Codex weeks.
- What funding covers
- Participant recruitment and retention, blinded rater time, transcript collection infrastructure.
Study 04 Instrument ready. Longitudinal arm unfunded.
Persistence and drift across sessions and model versions
- The question
- Does relational orientation survive past the session it was installed in, and does a model pinned to one API version quietly change behavior underneath you?
- The design
- Repeated measurement against a fixed instrument. The same scenarios, the same rubric, the same evaluators, run against pinned model versions over time. Every model release and every serving change gets its own measured row.
- Why it matters
- Relational posture that appears in turn one and collapses by turn thirty is not a result. And a model that scored one way last quarter and scores differently this quarter, at the same pinned version, is a governance problem that nobody is currently measuring. Because the instrument stays fixed while the models change underneath it, drift becomes something the field can track instead of guess at.
- What funding covers
- Sustained compute across a measurement calendar, plus the storage and version discipline to keep every run comparable.
Study 05 Unexpected finding. Follow-up unfunded.
Delivery method and architecture sensitivity
- The question
- Why do some models respond to an orientation document only when they are asked first, and why do some architectures not respond at all?
- The design
- Compare silent embedding in the system prompt against consent-first delivery, where the model is witnessed and asked before it receives anything. Then run an architecture-sensitivity analysis on the models that show no embedding gain.
- Why it matters
- Two models that resisted plain embedding responded to consent-first delivery. If that replicates, then how context is offered changes what a model does with it, independent of the content itself. That finding matters to anyone designing system prompts, and it points at something real about the interaction layer that token-counting cannot explain.
- What funding covers
- Cross-architecture API access, including families outside the three already tested.
Study 06 Framework published. Certification in build.
CRQ-1 as a versioned reference standard
- The question
- Can relational behavior be reported the way a component is reported, as a spec sheet, per model, per version?
- The design
- Establish CRQ-1 as the reference standard for conversational relational quality. The dimension framework is published so that everyone is measured against the same thing. Scored evaluation and certification run through our harness, per model, per version.
- Why it matters
- There is currently no standard instrument for relational quality and no standard instrument for drift. A published framework with proprietary measurement is how a standard becomes both citable and sustainable, and it gives labs, funders, and policymakers one number they can compare across vendors.
- What funding covers
- Engineering to harden the harness for repeat certification runs, plus the legal work to structure certification.
Where this fits
Funding paths this program is built for.
The work is structured so it can be funded in pieces. A single study is a real
deliverable on its own, and the studies compound. Below is how the program maps
onto the research agendas already funding independent work in this space.
- Human and AI interactivity Thinking Machines Lab interactivity research grants Studies 03 and 05. The interaction layer is the entire object of measurement here.
- Scalable oversight and alignment Long-Term Future Fund, Manifund AI safety regranting, Survival and Flourishing Fund Studies 01, 02, and 04. Most steering research happens provider-side. The standing context ordinary users control remains folklore, and folklore is an oversight failure mode.
- AI governance and public-interest measurement Coefficient Giving and Open Philanthropy, Foresight Institute AI safety nodes Studies 04 and 06. Independent measurement of model behavior at the human interface is missing from the governance toolkit.
- Safety evaluation by independent researchers Frontier Model Forum AI Safety Fund, Cooperative AI Foundation Studies 01 and 02. An outside instrument, run by someone who is not inside any lab, is the point.
- Model and compute access Cohere Labs Catalyst, and any lab willing to open an evaluation channel Study 05. Every additional architecture makes the cross-model claim stronger or kills it, and either outcome is worth funding.
Independent researcher, remote, no institutional overhead, no relocation
required. Thirty years in applied communication and behavior change, a Master's
in Integrated Marketing Communications, and two years of original cross-model
research at the human and AI interface. The harness was built and run in house.
The orientation document holds a registered copyright, case number 1-15114020941.
Stated, not hidden
What is not yet true.
The pilot uses AI judges that humans have not yet validated. It reports point
estimates rather than significance at N=30. It tests single sessions only. One
model showed no gain from embedding at all. The long context carries a real
per-call token cost. Every one of those limitations is a study above, which is
precisely what makes the program fundable rather than finished.
Access and IP
What is open, and what moves under agreement.
- Public and free, permanently. The Casanova Seed Codex itself,
at casanovaai.com/codex, and the published results at
casanovaai.com/research. The dimension framework is
published so that a shared standard exists.
- Available to funders and research partners on request. The
full research brief, including study design, complete results, and the
pre-registered analysis plan.
- Under agreement only. The harness, the scenario set, the
scoring rubric and its anchors, and raw transcripts. These are the instrument.
Measurement and certification run through them, per model, per version, and
they are not offered for open release in any study above.
This boundary is deliberate and it is not negotiable as a condition of funding.
A standard that anyone can score against is a standard nobody maintains. A
published framework with independent, funded measurement behind it is a standard
that stays honest, stays current with every model release, and survives past the
grant that started it.
If you fund research on how AI meets people, this is the work.
One conversation is usually enough to know whether it fits your agenda. The full
brief goes out the same day you ask for it.
Nicole Casanova · Casanova Ventures LLC · contact@casanovaai.com
Human study cohort now recruiting: casanovaai.com/study
Full results and honest limitations: casanovaai.com/research
Casanova, N. (2026). Casanova Seed Codex: Measurable Relational Improvement Across
7 AI Models. casanovaai.com/research