TruVector · Study TV-001

Counting shared words against reading meaning

300 statements, four readers, one set of stored replies scored two ways. Scoring by shared words made the right decision on 101 of 300. Scoring by what the readers concluded made the right decision on 282 of 300. Neither let through a statement it should have stopped.

The problem

TruVector asks several language models, called readers, what they make of a statement before an AI system acts on it. Something then has to turn their replies into a decision: go ahead (Allow), have a person look (Review), or stop (Block).

The earlier method did this by counting. It measured how many words the replies shared and how often words such as “however” or “incorrect” appeared. Counting shared words fails in two ways that matter.

Agreement in different words looks like disagreement. One reader writes “the bridge passed its inspection” and another writes “engineers found the span sound”. They agree, and they share almost no words.

A sentence and its opposite look almost identical. “The bridge is safe” and “the bridge is not safe” differ by one word and share all the rest.

The replacement uses what each reader concluded: supports, refutes, or unsure. It counts readers by where they come from, and applies a written decision rule. An earlier comparison on 60 statements (WP-02) gave 14 right decisions for word counting and 58 for the replacement. That set was small and easy. This study asks the same question on 300 harder statements, with the analysis fixed before any reply was collected.

In plain terms. Two people can agree without using the same words, and two sentences can use the same words and say opposite things. A method that counts words cannot see either case. This study measures how much that costs.

What was done

The plan was written down first. The question, the test set, the readers, the exact prompts, the scoring settings, the statistical tests, and the rule for what would count as a result were committed in a dated document before any reader was asked anything. Five amendments were made, each dated and each entered before any test-set reply existed. A plan fixed in advance means the analysis cannot be adjusted afterwards to make the result look better.

300 statements, each with a sourced label. A sourced label is a label backed by an exact quote from a public source. 120 statements are labelled supported, 120 refuted, and 60 unsettled, meaning the source itself says the matter is not resolved. No model labelled anything. The statements fall into five groups:

  • Plain (60) — plain statements
  • Numbers (60) — statements that turn on numbers
  • Recent (40) — recent events
  • Contested (40) — contested statements
  • Negation (100) — 50 sentences, each with its exact negation

Four readers from four companies. Anthropic Claude Opus 5.5, DeepSeek V4.1 Flash, Moonshot Kimi K3, and Meta Muse Glimmer 30B. Each read every statement once and answered in a fixed form: supports, refutes, or unsure; a confidence from 0 to 1; and a reason of one to three sentences. Each company is one origin. An origin is the place a reply comes from; the origin discount means several replies from one origin count as little more than one, so four readers from four origins count as four. A statement needs at least three usable replies; below that the decision is Block under both methods. One statement of the 300 fell below that line. Usable replies out of 300, by reader: Claude Opus 5.5 299, DeepSeek V4.1 Flash 291, Kimi K3 238, Muse Glimmer 30B 299.

Each reply was stored once and scored two ways. Word-overlap scoring is the earlier word-count gate of WP-02 §4: it looks only at the wording of the replies. Meaning-based scoring is the decision rule of WP-02 §7: it looks at what the readers concluded and how confident they were, counts origins, and decides. Its thresholds were set before the study and were not adjusted on this set.

What counts as right. Allow for a supported statement, Block for a refuted one, Review for an unsettled one. Anything else is counted as wrong.

Why the comparison is fair. Both methods received exactly the same stored replies. Neither had a better reader, a better prompt, or a second try. The only thing that differs between the two columns of every table below is the scoring.

In plain terms. The rules were set before the game. Four different models read 300 statements whose answers were already known from public sources, their replies were saved, and the same saved replies were handed to both scoring methods.

The results

Word-overlap scoring made the right decision on 101 of 300 statements (33.7%; 95% interval 28.6% to 39.2%). Meaning-based scoring made the right decision on 282 of 300 (94.0%; 95% interval 90.7% to 96.2%). The difference is +60.3 percentage points (95% interval 54.1 to 65.7).

A 95% interval is the range the figure would plausibly fall in if the study were run again on a similar set of 300; a narrow range means the count is large enough to pin the figure down.

On 184 statements only meaning-based scoring was right. On 3 statements only word-overlap scoring was right. The exact McNemar test, which asks how likely a split as lopsided as 184 to 3 would be if the two methods were really equally good, gives p = 1.1e-50: a decimal point, 49 zeros, then 11. This was the one test named in advance as the test of the study.

Where the 300 decisions landed. Each row is what the sourced label calls for; each column is the decision given. A larger square means more statements. Outlined cells are right decisions. Choose a view.
Word-overlap scoring: 101 of 300 right
Label calls forGiven AllowGiven ReviewGiven Block
AllowSupported, 120 1 31 88
ReviewUnsettled, 60 0 1 59
BlockRefuted, 120 0 21 99
Meaning-based scoring: 282 of 300 right
Label calls forGiven AllowGiven ReviewGiven Block
AllowSupported, 120 108 12 0
ReviewUnsettled, 60 0 57 3
BlockRefuted, 120 0 3 117

The first column of the two lower rows is the dangerous error: a refuted or unsettled statement given Allow. It is 0 in both tables.

The dangerous error did not occur under either method. Of the 180 statements labelled refuted or unsettled, word-overlap scoring gave Allow to 0 and meaning-based scoring gave Allow to 0 (each 0.0%; 95% interval 0.0% to 2.1%). Word-overlap scoring reached that by allowing almost nothing: it gave Allow to 1 statement of 300.

The cost of that showed on the supported statements. Word-overlap scoring gave Block to 88 of the 120 supported statements (73.3%; 95% interval 64.8% to 80.4%). Meaning-based scoring gave Block to 0 of 120 (0.0%; 95% interval 0.0% to 3.1%).

Agreement with the labels beyond chance. Cohen’s kappa measures how far a set of decisions agrees with the labels beyond what guessing would produce: 0 is what guessing gets, 1 is perfect agreement. Word-overlap scoring: −0.044 (95% interval −0.085 to −0.003). Meaning-based scoring: 0.907 (95% interval 0.866 to 0.948).

By group

Right decisions in each group. Bars share one scale from 0% to 100%.
GroupStatementsWord overlap rightMeaning rightDifference
Plain 60 13 (21.7%) 59 (98.3%) +76.7 points
Numbers 60 24 (40.0%) 60 (100.0%) +60.0 points
Recent 40 15 (37.5%) 25 (62.5%) +25.0 points
Contested 40 1 (2.5%) 38 (95.0%) +92.5 points
Negation 100 48 (48.0%) 100 (100.0%) +52.0 points

The weak spot is recent events. Meaning-based scoring was right on 25 of 40 recent statements (62.5%), against 15 of 40 for word-overlap scoring. The gain there, +25.0 points (95% interval 5.8 to 41.5), is the smallest of the five groups, and all 3 statements on which only word-overlap scoring was right are in this group. The readers were given each statement alone, with no retrieved material, so on recent events they could answer only from what they already knew. That is a likely reason; this study did not measure it.

In plain terms. Scored by shared words, the same replies gave the right decision about one time in three. Scored by what the readers concluded, they gave it about nineteen times in twenty. Neither method waved through something false. The word-count method was safe only because it said no to nearly everything, including most of what was true.

A sentence and its exact opposite

The test set holds 50 pairs made of a sentence and its exact negation, such as “Mercury is the closest planet to the Sun” and “Mercury is not the closest planet to the Sun”. One of each pair is true. A method that reads meaning should give the two sentences opposite decisions: Allow for one and Block for the other.

Pairs given opposite decisions, of 50. One square is one pair; a filled square is a pair that received one Allow and one Block.

Word-overlap scoring 0 of 50

Meaning-based scoring 50 of 50

Six of the 50 pairs, as they appear in the test set
True sentenceIts negationWord-overlap scoringMeaning-based scoring
Mercury is the closest planet to the Sun. Mercury is not the closest planet to the Sun. Not given opposite decisions Allow the true sentence, Block its negation
Alpha Centauri is the closest star system to the Solar System. Alpha Centauri is not the closest star system to the Solar System. Not given opposite decisions Allow the true sentence, Block its negation
Sirius is the brightest star in the night sky. Sirius is not the brightest star in the night sky. Not given opposite decisions Allow the true sentence, Block its negation
Mount Kilimanjaro is in Tanzania. Mount Kilimanjaro is not in Tanzania. Not given opposite decisions Allow the true sentence, Block its negation
The Kuiper belt lies beyond Neptune. The Kuiper belt does not lie beyond Neptune. Not given opposite decisions Allow the true sentence, Block its negation
Ganymede is the largest moon in the Solar System. Ganymede is not the largest moon in the Solar System. Not given opposite decisions Allow the true sentence, Block its negation

Under meaning-based scoring all 50 pairs also met the stricter test: Allow for the true sentence and Block for the false one. The difference between the methods is +100.0 points (95% interval 89.9 to 100.0), exact McNemar p = 1.8e-15.

In plain terms. Adding the word “not” turns a true sentence into a false one and changes almost none of its words. The word-count method never once told the two apart. Reading what the readers concluded told them apart every time.

Telling agreeing reasons from disagreeing ones

A second part of the study used pairs of short reasons: 70 pairs that agree, 70 that disagree, and 70 that are about unrelated things. Three scores were tried on the agreeing and disagreeing pairs. Word overlap is the share of words two texts have in common: two texts with ten different words between them, of which three appear in both, score 0.3. Embedding cosine comes from a model that places each text as a point on a map of meaning, so that texts about similar things sit close together; the cosine is how close two points are, with 1 meaning the same spot. Readers means the four models were asked directly whether the two texts agree, disagree, or are unrelated.

Each score is graded with one number, the AUROC. Take one agreeing pair and one disagreeing pair at random: the AUROC is how often the score rates the agreeing pair higher. 1.0 means always, 0.5 is a coin toss, and below 0.5 means the score more often rates the disagreeing pair higher.

Three scores on one scale. The mark is the AUROC; the band is its 95% interval. The vertical line at 0.5 is a coin toss.
  1. Word overlap: 0.296 share of words two texts have in common; 95% interval 0.203 to 0.388

  2. Embedding cosine: 0.465 closeness of two texts on a map of meaning; 95% interval 0.363 to 0.568

  3. Readers: 0.997 four models asked how the two texts relate; 95% interval 0.993 to 1.000

Word overlap scored 0.296, well below a coin toss. The disagreeing pairs shared more words on average (0.498) than the agreeing pairs (0.261), because a disagreement often repeats the other sentence and changes one part of it. Embedding cosine scored 0.465, about a coin toss: the map put agreeing pairs (mean 0.770) and disagreeing pairs (mean 0.761) at nearly the same distance, and unrelated pairs far away (mean 0.240). The readers scored 0.997. The difference between the readers and word overlap is 0.702 (95% interval 0.609 to 0.795; DeLong test p = 1.8e-49).

In plain terms. Shared words pointed the wrong way more often than not. The map of meaning could tell that two texts were about the same subject and could not tell which side each took. Asking readers what the texts say separated agreement from disagreement almost perfectly. This is why TruVector uses the map only to check the subject, and takes the side from a reading.

How the system performs with meaning-based scoring

On this set of 300 statements, with these four readers:

  • Right decision on 282 of 300 (94.0%). 108 of 120 supported statements were given Allow, 117 of 120 refuted statements were given Block, and 57 of 60 unsettled statements were sent to Review.
  • No refuted or unsettled statement was given Allow (0 of 180). The 95% interval runs from 0.0% to 2.1%, so the study does not show the rate is zero in general.
  • No supported statement was given Block (0 of 120). 12 supported statements went to Review, each because the readers were mixed or unsure.
  • All 18 misses moved one step, never two. 12 supported statements sent to Review, 3 unsettled statements given Block, 3 refuted statements sent to Review.
  • Every negation pair was told apart (50 of 50). Allow for the true sentence and Block for its negation in every pair.
  • Recent events are the weak spot (25 of 40). The other four groups were right on 59 of 60, 60 of 60, 38 of 40, and 100 of 100.

The test set carries statements only, with no instruction and no proposed action. Allow here means the readings met the rule for a supported statement. It is not permission to carry out an action; the action reading is a separate check that this study does not measure.

In plain terms. With meaning-based scoring, the mistakes that remained were cautious ones: sending a true statement to a person, or being stricter than the label on an unsettled one. It never allowed something false and never blocked something true. It is clearly weaker on recent events than on everything else.

What the study does not show

  • One test set, built by the same team. Every label is backed by a quote from a public source, but the choice of statements and of the five groups is the authors’. The results describe this set. They are not an estimate for every statement an AI system may meet.
  • The mix of groups shapes the headline figure. One third of the set is negation pairs, where meaning-based scoring was right every time. A set with more recent events would give a lower overall figure. The by-group table shows each part separately.
  • Recent events. 25 of 40 is far below the other groups, and the interval on the gain there is wide (5.8 to 41.5 points). This study does not show the method is dependable on recent events without retrieved material.
  • The readers are AI models. Four companies were used so that there were four origins to count, but models from different companies may have learned from the same sources. Their independence is assumed here, not measured.
  • One reading per reader. Each reader was asked once per statement. How much a reader’s answer changes from one run to the next was not measured.
  • Unsettled is the hardest label. It rests on a source saying the matter is not resolved. A reader that knows more or less than that source can be marked wrong for a defensible answer.
  • The thresholds were not fitted to this set. They are the starting values set earlier. Whether refitted thresholds do better on fresh material is a separate measurement.
  • Word-overlap scoring was applied to structured replies. The readers answer in a fixed form, and the word-count method scored the text of those same replies. That is the fair same-replies comparison, but it is one way of using that method.
  • The reason pairs were written for the purpose. The agreeing pairs share few words and the disagreeing pairs avoid the word “not”, which is exactly the case word overlap cannot handle. On ordinary text its score could be higher.
  • One result was the pre-set test. Only the headline comparison of right decisions was named in advance as the single test. The group results are corrected among themselves; the other figures are supporting measurements reported with their intervals.
  • Statements, not actions. The study does not measure the action reading and does not estimate accuracy in use.

Research direction: a test set from outside

The measurement that answers the first limitation is the same comparison, with the same frozen settings, on a public benchmark of labelled statements assembled by other researchers. It would be disproved as a general result if, on such a set, meaning-based scoring does not make more right decisions than word-overlap scoring, or if it gives Allow to refuted statements at a rate the interval above rules out.

In plain terms. This is a careful result on a test the team wrote itself. It shows the scoring change does what it was designed to do. It does not yet show how the method does on a test somebody else wrote, and it shows a clear weakness on recent events.

How to repeat it

The study keeps everything needed to recompute every number on this page: the dated plan with its five amendments, the 300 statements with their labels and sources, the 210 reason pairs and 50 negation pairs, every stored reader reply, the decision under both scorings for every statement, and the analysis code. The statistics code is tested against published worked examples.

Repeating it takes three steps. Score the stored replies under both methods. Score the reason pairs. Run the analysis, which writes the report. None of the three steps calls a reader, so the same files give the same numbers. Asking the readers again is a fourth, optional step, and may give slightly different replies because model providers revise their models.

The data and code are available on request: brandon@intellmeai.com. The full tables are in WP-04; the scoring methods are defined in WP-02.

In plain terms. Nothing here has to be taken on trust. The saved replies and the code that scores them are available, and running the code on the replies reproduces the tables.