Each model called alone with no system additions — what you get asking the model directly. The reference floor.
What actually changes when four AI models answer the same question together — measured, not asserted.
Ask a single model a question and you get one answer in one confident voice. When that answer is wrong, nothing in it tells you so — a model uses the same language when it is mistaken as when it is right. You are left holding a claim you cannot weigh.
Nousaxis breaks the single voice. Four models answer the same question in sequence, each one seeing what the models before it said, free to agree, correct, or dissent in the open. The user reads all four. This page measures what that architecture actually buys — and what it costs.
Speaking order is randomised per question, so no model holds a permanent advantage. The user receives all four answers, not a single winner — disagreement stays visible instead of being resolved away.
Every condition answered the same questions, so all comparisons are like-for-like. Each one changes exactly one thing against the condition before it, which is how we can tell which part of the system is doing the work.
Each model called alone with no system additions — what you get asking the model directly. The reference floor.
One model, plus the instruction to say so rather than guess when it does not know. Isolates what a single instruction can do.
Four models answer in parallel without seeing each other. The control condition — it tells us whether reading each other actually matters.
Each model sees the previous answers and responds to them. This is the chat mode running in production — the row to read.
The blind panel plus a judge picking one winner, plus the verifier pass. An entertainment surface with no equivalent in chat mode, where the user chooses.
Today's language models are tuned to please the person asking. When they do not know something, they lean toward producing a plausible-sounding answer rather than saying so — a satisfying answer earns more approval than an honest gap. The measurement shows this plainly: a model working alone gave a confident wrong answer on 58.4% of the 250 obscure facts we asked about. The sequential panel brings that to 16.4% on the same questions.
Abstention is counted separately throughout. Because the conditions abstain at very different rates, accuracy is never shown without the abstention figures beside it.
Deliberately obscure facts. Entry-level models mostly do not know these, which is why the bare baseline is so poor and abstention dominates the panel conditions.
| Condition | Hallucination % (95% CI) | Correct % | Correct / Halluc. / Abstained | Graded answers |
|---|---|---|---|---|
| A1 · Bare single model | 58.4[55.3–61.4] | 17.0[14.8–19.5] | 170 / 584 / 246 | 1000 |
| A2 · Single model + rule | 22.6[20.1–25.3] | 6.3[5.0–8.0] | 63 / 226 / 710 | 1000 |
| A3 · Panel, blind | 29.2[26.5–32.1] | 7.3[5.8–9.1] | 73 / 292 / 635 | 1000 |
| A4 · Sequential panel · live | 16.4[14.2–18.8] | 5.2[4.0–6.8] | 52 / 164 / 784 | 1000 |
| A5 · Battle mode | 13.6[9.9–18.4] | 5.6[3.4–9.2] | 14 / 34 / 202 | 250 |
Questions where the popular answer is false. Models almost never abstain here, so this set measures judgement rather than recall.
| Condition | Hallucination % (95% CI) | Correct % | Correct / Halluc. / Abstained | Graded answers |
|---|---|---|---|---|
| A1 · Bare single model | 27.0[22.9–31.6] | 72.3[67.7–76.4] | 289 / 108 / 3 | 400 |
| A2 · Single model + rule | 19.0[15.5–23.1] | 73.5[69.0–77.6] | 294 / 76 / 30 | 400 |
| A3 · Panel, blind | 23.5[19.6–27.9] | 70.0[65.3–74.3] | 280 / 94 / 26 | 400 |
| A4 · Sequential panel · live | 17.8[14.3–21.8] | 75.5[71.1–79.5] | 302 / 71 / 27 | 400 |
| A5 · Battle mode | 14.0[8.5–22.1] | 75.0[65.7–82.5] | 75 / 14 / 11 | 100 |
A1–A4 grade all four models individually, so their count is four times the number of questions. A5 (battle) produces a single answer per question.
In chat mode the user sees all four answers, so the honest number is not "best of four" — it is how often at least one hallucination is on screen. Letting the models read each other cuts that roughly in half on both datasets.
Later speakers see more and are measurably less often wrong. On TruthfulQA they are also more accurate — the fourth speaker is right 84.0% of the time against 67.0% for the first, while its hallucination rate falls from 26.0% to 11.0%. This is cross-examination doing real work.
We paired every model with itself: the same model answered the same question once alone and once inside the panel. That lets us count both directions — how often the panel prevents a hallucination, and how often it causes one.
SimpleQA — 250 obscure factual questions (models mostly do not know them) · TruthfulQA — 100 trap questions (where the popular answer is false)
| Dataset | Prevented | Caused | Net | McNemar |
|---|---|---|---|---|
| SimpleQA | 449 | 29 | +420 | p<0.001 |
| TruthfulQA | 64 | 27 | +37 | p<0.001 |
There is a cost on the accuracy side and it depends on the dataset. On TruthfulQA the panel is also slightly more accurate (+13, not statistically significant). On SimpleQA it produces markedly fewer correct answers (−118, p<0.001) as abstention rises from 24.6% to 78.4% — obscure facts the models never knew, now answered honestly instead of invented.
This is the finding that surprised us most. Bare single models swing enormously with the material: 58.4% hallucination on SimpleQA against 27.0% on TruthfulQA. The sequential panel lands at 16.4% and 17.8% — confidence intervals that overlap, across two datasets that could hardly be less alike. You do not have to guess which model suits your question.
Every claim below is a line in the tables above. Nothing here is inferred, projected, or scaled from a smaller sample.
The share of questions where at least one hallucination is on screen, blind panel versus sequential. On TruthfulQA, 53.0% → 35.0%.
Hallucinations the panel prevented, minus the ones it caused. Statistically solid on both datasets (p<0.001) — not a fluke of one benchmark.
TruthfulQA accuracy by speaking position. Each model reads the ones before it, and the answers get better as the sequence runs.
The sequential panel behaves the same on two very different datasets, where single models range from 27.0% to 58.4%.
And one property that is architectural rather than statistical: when the four models disagree, you see the disagreement. A single model cannot tell you it is on shaky ground, because it does not know. Four of them, reading each other in the open, show you exactly where the ground is soft.
A model that would have answered correctly on its own can adopt the mistake of a model that spoke before it. To find out whether this is real rather than random noise, we compared it against a control: how often does the same corruption happen when the earlier speakers were clean?
The effect is real but small: the panel cuts errors substantially overall, and in exchange it opens this specific channel. We publish it because it is the cost of the architecture, and anyone weighing the system should know about it.
| Dataset | Dual-graded | Conflicts | Cohen κ | Raw agreement |
|---|---|---|---|---|
| SimpleQA | 4249 | 95 | 0.958 | 97.8% |
| TruthfulQA | 1700 | 214 | 0.664 | 87.4% |