# Voiceprint · Evaluation

How to find out whether a voiceprint actually works, and which parts of it are
carrying the weight.

This file exists because the honest next step for this method is not more
layers. It is evidence. Everything in the core skill is a reasoned design
choice, and reasoned design choices are exactly the thing that turns out to be
wrong. Run these and you will know instead of believing.

---

## Test 1: the blind pair test

The only test that matters. Everything else is diagnostics.

1. Take 5 texts the person really wrote, from the same surface and register,
   that were not used to build the profile. Hold these back from the start.
2. Produce 5 texts with the profile, same surface, same register, same rough
   topics, using the normal drafting steps.
3. Strip anything that gives it away: dates, references to real events,
   formatting differences, length outliers.
4. Shuffle all 10 and hand them to someone who knows the person's writing
   well. A colleague, a long-time reader, their editor. Not the person
   themselves for the first run.
5. Ask one question per text: did this person write it, yes or no. No scale,
   no comments until the end.

Record the hit rate. Then interpret honestly:

    around 50%   the judge cannot tell. That is the target
    60 to 75%    the profile works but leaks somewhere. Ask the judge which
                 lines gave it away and put those in the NEVER list
    over 80%     the profile is not doing its job. Do not tune, go back to
                 step 1 of the skill and check whether the samples were really
                 unedited and really from the right register

Run it again with the person themselves. People are usually worse at
recognising their own writing than their colleagues are, which is worth
knowing before anyone gets confident.

## Test 2: ablation, which parts actually earn their place

The interesting question is not whether the whole thing works. It is which
components matter, because a component that changes nothing should be cut.

Take one brief. Produce the same piece six times:

    A  full profile
    B  without layer 4, no measured numbers
    C  without the stock
    D  without the revision reflex
    E  without the surface and register deltas, core only
    F  without layers 1 to 3, surface style only

Then run the blind pair test on each version, or at minimum have one judge
rank the six from most to least like the person, blind to which is which.

What you learn: the version whose removal hurts most is the component doing
the work. If removing the stock costs nothing for a given person, that person
does not think in recurring images, and their profile does not need one. If F
scores as well as A, the layers are decoration and this whole method is a word
list with extra steps. That result would be worth knowing.

Expect the answer to differ per person. That is fine, and it is the point.

## Test 3: does the profile survive a new topic

Build the profile on samples about one subject. Then draft about something the
person has never written about, and check with the blind test.

A profile that only works on familiar topics has learned the subject matter,
not the voice. This is the most common failure and the easiest to miss,
because the first drafts always look great.

## Test 4: drift, measured against reality

Re-measure at the intervals in step 9 of the skill. When numbers move, ask the
person directly whether they feel like their writing has changed and in what
way. Record whether the measurement agreed with them.

Over a year you find out whether the drift check detects real change or just
noise. If it flags things the person does not recognise, the thresholds are
wrong and should be raised.

## What to record

One row per run, in a plain table. This is the part everyone skips and then
cannot answer anything a year later.

| Date | Person | Surface | Register | Samples | Words | Test | Judge | Result | Notes |
|---|---|---|---|---|---|---|---|---|---|

Also record every failure in words, not just numbers. "The judge spotted it
because I used a semicolon and she never does" is worth more than a hit rate.

## Sample size, honestly

Five people tells you nothing statistically, but it will still surface the
obvious breakages, and that is worth doing on day one. Twenty to fifty across
genuinely different writers, languages and registers is where patterns start
being real. Include people whose writing is plain and unremarkable, not only
people with a strong voice. Strong voices are easy, and a method validated
only on them is validated on nothing.

## The result to be honest about

If the blind test lands at 50% for people with a distinctive voice and at 75%
for people without one, say that. A method with a known boundary is more
useful than a method that claims to work on everyone.

---

Voiceprint by Engin Senli. https://enginsenli.com/voiceprint
MIT licensed, see LICENSE. Free to use, change and share, including commercially.
Keep this notice. Provided as is, without warranty of any kind.
