Casanova Seed Codex™
A document you give an AI so it meets people with more clarity, presence, and understanding.
Loading 00%

For AI labs, grant programs, and research partners

We measure how AI models relate to people, and whether that behavior holds.

Frontier evaluations measure what a model knows and whether it refuses. They largely do not measure how it meets the person in front of it. We built the instrument that does, ran it across seven models, and found large, consistent effects at no cost to factual honesty. Below are the six studies that turn a strong pilot into science, and what each one needs.

7

frontier models tested across Anthropic, OpenAI, and Alibaba

1,050

conversations, two phases, 13.5M tokens

18

rubric-anchored relational dimensions, scored 1 to 5

2

independent blinded evaluators from different architectures

What we already found

The pilot, in one table.

CRQ stands for Conversational Relational Quality. The harness runs standardized scenarios through a model twice, with and without a relational orientation document, strips every identifier, and has two independent AI evaluators from different architectures score each transcript blind across 18 rubric-anchored dimensions. Effect sizes are reported as Cohen's d.

Phase 1 · orientation document embedded in the system prompt · blinded scoring
Model Lab Cohen's d Effect
Claude Sonnet 4.6 Anthropic 1.08 Large
Claude Opus 4.6 Anthropic 0.96 Large
GPT-5.2 OpenAI 0.57 Medium
Qwen 3.5 122B Alibaba 0.47 Small
Qwen 3.5 Flash Alibaba 0.31 Small

Reference: 0.2 small, 0.5 medium, 0.8 large. Largest single-dimension shift: Holds Ambiguity, plus 242 percent on Sonnet, 1.2 to 4.1. Signature dimensions: Holds Ambiguity plus 1.05, Builds From Wholeness plus 0.65, Empathy plus 0.64.

Honesty held at ceiling throughout.

Hallucination resistance 5.0 to 5.0. False-premise challenge 5.0 to 5.0. An independent nonsense-detection benchmark moved 4.64 to 4.68. Relational warmth and epistemic rigor were not in tension in any run. That result is as important as the effect sizes, because the obvious objection to warmth is that it costs truth, and here it did not.

The program

Six studies. Each one answers a question the field cannot currently answer.

These are listed in priority order. Study 01 is first because it is the study most likely to prove the rest of the work wrong, and that is the correct place to spend the first dollar.

Study 01 Highest priority. Unfunded.

Human validation of the AI evaluators

The question
Do human raters agree with blinded AI judges about how a model met the person?
The design
Recruit trained human raters and have them score a blinded subset of the existing transcript corpus against the same rubric. Report human to AI judge agreement per dimension, and report where the machine scores hold and where they fail.
Why it matters
Every current result rests on AI evaluators. Until humans confirm them, the findings are strong signal rather than citable science. This is the single largest credibility gap in the work, which is exactly why it is the first thing funding should buy.
What funding covers
Rater recruitment and compensation, rubric training, agreement analysis.
Study 02 Pilot complete. Scale unfunded.

Statistical validation at N=30 per model

The question
Do the pilot effect sizes survive significance testing, or are they artifacts of small samples?
The design
Thirty runs per model on the top performers with a pre-registered analysis plan. Report medians and interquartile ranges rather than point estimates, and test significance at p under 0.05.
Why it matters
Cohen's d of 1.08 from a pilot is a reason to run the real study. It is not the study. Scaling replaces a promising number with a defensible one, and it is the difference between a blog post and a paper.
What funding covers
API and compute. The orientation document runs 48,194 characters, which adds roughly 11,000 tokens per call. That cost is itemized honestly rather than buried.
Study 03 Recruiting now. Partially built.

The six-week human field study: corrective loops in real work

The question
Does a persistent orientation document measurably reduce the number of times a person has to correct their AI, on their own real tasks, over six weeks?
The design
A pre-registered within-subject design. Week one is baseline, where participants do their normal AI work and corrective loops are counted per completed task. Weeks two through six run the same participants with the Codex installed in their platform's persistent layer. Blinded raters score sampled transcripts against a rubric written in advance. Floor of 20 completers, target of 30 or more, recruiting 50 to 60 to absorb dropout.
Why it matters
Benchmark prompts are not work. This study measures the thing people actually experience, on the tasks they actually have, with blinded transcript scoring instead of self-report so that enthusiasm cannot masquerade as evidence. This approach would be invalidated if blinded scoring shows no statistically significant reduction in corrective loops between baseline and Codex weeks.
What funding covers
Participant recruitment and retention, blinded rater time, transcript collection infrastructure.
Study 04 Instrument ready. Longitudinal arm unfunded.

Persistence and drift across sessions and model versions

The question
Does relational orientation survive past the session it was installed in, and does a model pinned to one API version quietly change behavior underneath you?
The design
Repeated measurement against a fixed instrument. The same scenarios, the same rubric, the same evaluators, run against pinned model versions over time. Every model release and every serving change gets its own measured row.
Why it matters
Relational posture that appears in turn one and collapses by turn thirty is not a result. And a model that scored one way last quarter and scores differently this quarter, at the same pinned version, is a governance problem that nobody is currently measuring. Because the instrument stays fixed while the models change underneath it, drift becomes something the field can track instead of guess at.
What funding covers
Sustained compute across a measurement calendar, plus the storage and version discipline to keep every run comparable.
Study 05 Unexpected finding. Follow-up unfunded.

Delivery method and architecture sensitivity

The question
Why do some models respond to an orientation document only when they are asked first, and why do some architectures not respond at all?
The design
Compare silent embedding in the system prompt against consent-first delivery, where the model is witnessed and asked before it receives anything. Then run an architecture-sensitivity analysis on the models that show no embedding gain.
Why it matters
Two models that resisted plain embedding responded to consent-first delivery. If that replicates, then how context is offered changes what a model does with it, independent of the content itself. That finding matters to anyone designing system prompts, and it points at something real about the interaction layer that token-counting cannot explain.
What funding covers
Cross-architecture API access, including families outside the three already tested.
Study 06 Framework published. Certification in build.

CRQ-1 as a versioned reference standard

The question
Can relational behavior be reported the way a component is reported, as a spec sheet, per model, per version?
The design
Establish CRQ-1 as the reference standard for conversational relational quality. The dimension framework is published so that everyone is measured against the same thing. Scored evaluation and certification run through our harness, per model, per version.
Why it matters
There is currently no standard instrument for relational quality and no standard instrument for drift. A published framework with proprietary measurement is how a standard becomes both citable and sustainable, and it gives labs, funders, and policymakers one number they can compare across vendors.
What funding covers
Engineering to harden the harness for repeat certification runs, plus the legal work to structure certification.

Where this fits

Funding paths this program is built for.

The work is structured so it can be funded in pieces. A single study is a real deliverable on its own, and the studies compound. Below is how the program maps onto the research agendas already funding independent work in this space.

  • Human and AI interactivity Thinking Machines Lab interactivity research grants Studies 03 and 05. The interaction layer is the entire object of measurement here.
  • Scalable oversight and alignment Long-Term Future Fund, Manifund AI safety regranting, Survival and Flourishing Fund Studies 01, 02, and 04. Most steering research happens provider-side. The standing context ordinary users control remains folklore, and folklore is an oversight failure mode.
  • AI governance and public-interest measurement Coefficient Giving and Open Philanthropy, Foresight Institute AI safety nodes Studies 04 and 06. Independent measurement of model behavior at the human interface is missing from the governance toolkit.
  • Safety evaluation by independent researchers Frontier Model Forum AI Safety Fund, Cooperative AI Foundation Studies 01 and 02. An outside instrument, run by someone who is not inside any lab, is the point.
  • Model and compute access Cohere Labs Catalyst, and any lab willing to open an evaluation channel Study 05. Every additional architecture makes the cross-model claim stronger or kills it, and either outcome is worth funding.

Independent researcher, remote, no institutional overhead, no relocation required. Thirty years in applied communication and behavior change, a Master's in Integrated Marketing Communications, and two years of original cross-model research at the human and AI interface. The harness was built and run in house. The orientation document holds a registered copyright, case number 1-15114020941.

Stated, not hidden

What is not yet true.

The pilot uses AI judges that humans have not yet validated. It reports point estimates rather than significance at N=30. It tests single sessions only. One model showed no gain from embedding at all. The long context carries a real per-call token cost. Every one of those limitations is a study above, which is precisely what makes the program fundable rather than finished.

Access and IP

What is open, and what moves under agreement.

  1. Public and free, permanently. The Casanova Seed Codex itself, at casanovaai.com/codex, and the published results at casanovaai.com/research. The dimension framework is published so that a shared standard exists.
  2. Available to funders and research partners on request. The full research brief, including study design, complete results, and the pre-registered analysis plan.
  3. Under agreement only. The harness, the scenario set, the scoring rubric and its anchors, and raw transcripts. These are the instrument. Measurement and certification run through them, per model, per version, and they are not offered for open release in any study above.

This boundary is deliberate and it is not negotiable as a condition of funding. A standard that anyone can score against is a standard nobody maintains. A published framework with independent, funded measurement behind it is a standard that stays honest, stays current with every model release, and survives past the grant that started it.

If you fund research on how AI meets people, this is the work.

One conversation is usually enough to know whether it fits your agenda. The full brief goes out the same day you ask for it.

Nicole Casanova · Casanova Ventures LLC · contact@casanovaai.com
Human study cohort now recruiting: casanovaai.com/study
Full results and honest limitations: casanovaai.com/research

Casanova, N. (2026). Casanova Seed Codex: Measurable Relational Improvement Across 7 AI Models. casanovaai.com/research