Quorum:
ask the jury.
Three open models under 1B parameters read a message and each casts a vote. Quorum returns the intent and a confidence you can act on. Zero-shot, no training, and it runs on a CPU.
Level with Jev (0.924); the difference is not significant. Given the same retrieved examples, the two tie: 0.936 vs 0.935. Runs a 4B reader; a GPU is recommended.
With an NLI juror never trained on Banking77. Jev, measured on the same sentences: 0.803.
One NLI model and two sentence embedders. No training, no calibration, and it runs on a CPU.
Try it
Every message below is a real test sentence from a public benchmark, with the dataset's own answer attached. Pick a domain, draw a sentence and hear the verdict. You will see when the jury is right, and when it isn't.
01 Choose a domain
Messages come only from the benchmarks' test sets, so every verdict can be checked against the dataset's own answer.
Three jurors vote, Quorum rules, and the dataset tells you whether it was right.
Requests go to the live Quorum API. The demo only sends benchmark sentences, and those verdicts are precomputed, so most come straight back from the cache. The footnote under each verdict shows which, and how long it took.
Results
We evaluate on the full official test set of seven public intent benchmarks, covering banking, smart assistants, voice commands, everyday tasks and customer support. The zero-shot configuration is fixed and identical on every benchmark. Each number below can be recomputed from the repo, including our Jev measurements.
Zero-shot, seven benchmarks
The same three-model jury, with no training. “Names” matches each message against bare label names. “Descriptions” matches against a one-line description per label: a plain rephrasing anyone can write, not taken from the data.
| Dataset | Labels | Test sentences | Names | Descriptions | Jev |
|---|---|---|---|---|---|
| Banking77† | 77 | 3,080 | .7555 | .7737 | .8029 |
| CLINC150 | 151 | 5,500 | .6518 | .6816 | .9089 |
| HWU64† | 64 | 1,076 | .7388 | .7574 | — |
| MASSIVE† | 60 | 2,974 | .7394 | .7774 | — |
| MTOP | 113 | 4,386 | .6140 | .7214 | .8698 |
| SNIPS | 7 | 1,400 | .8529 | .9443 | .9786 |
| Bitext | 27 | 5,375 | .7751 | .8203 | .9233 |
Against Jev
Jev is a closed, commercial calibrated-decision model with no official public Banking77 number. We measured it ourselves through its API (model jev-1.13.0), on the same test sentences and with the same label text, after all of our configurations were frozen. Its outputs were used only for evaluation, never for training or tuning.
Zero-shot
24-shot
At 24-shot, the code in the repo scores 0.932 against 0.924 for Jev in the public reproduction's setup, which is not a significant difference (p = 0.066). A pre-registered run scored 0.936 (p = 0.014 against that setup). Most of that edge comes from how we retrieve the 24 examples, not from the models: given our retrieved examples, Jev rises to 0.935 and the two tie (p = 0.87). Jev's run in the reproduction's setup gives 92.40%, exactly the reproduction's published figure.
Zero-shot, Jev is clearly ahead: 0.803 against 0.716 for our jury with an NLI juror never trained on Banking77 (0.774 with the default juror, which saw Banking77 training data). On the four benchmarks no juror has seen, the gap ranges from 3.4 points (SNIPS) to 22.7 points (CLINC150), all significant.
The jury beats its own jurors
On the four benchmarks where no juror has any known training exposure (CLINC150, MTOP, SNIPS, Bitext), the jury beats its best single member in all eight conditions, names and descriptions, even with an NLI juror never trained on intent data. On Banking77 that does not hold: with that juror the jury is level with its best member using names (0.690 vs 0.700, not significant) and below it using descriptions (0.716 vs 0.738).
| Dataset | Ensemble | Best single member | Paired McNemar |
|---|---|---|---|
| MASSIVE† | 0.739 | 0.691 | win · p < .001 |
| MTOP | 0.614 | 0.587 | win · p = 5e-5 |
| Bitext | 0.775 | 0.712 | win · p < .001 |
| CLINC150 | 0.652 | 0.630 | win · p < .001 |
| HWU64† | 0.739 | 0.643 | win · p < .001 |
| SNIPS | 0.853 | 0.868 | tie · p = 0.06, ns |
How it works
No single small model is a great zero-shot intent classifier. Three complementary ones disagree in useful ways. Each juror reads the message against every candidate label, phrased as “This message is about {label}.”
PrismNLI-0.4B scores entailment; bge-large-en-v1.5 and bge-base-en-v1.5 score cosine similarity to the same verbalized labels. Each juror's scores become a probability distribution (softmax at temperature 1). Quorum takes the geometric mean, the mean of log-probabilities, and renormalises. The top label is the verdict; its probability is the confidence. No labelled data, no training, no per-task calibration.
Descriptions, not names
A model that has never seen a dataset gets little signal from a short label name. A one-line description gives every juror more to match against, at no cost: the same three models, no extra code. The gain scales with how opaque the names are: +0.091 on SNIPS, whose names are short and ambiguous, against +0.018 on Banking77, whose names are already fairly descriptive.
The 24-shot pipeline
At the 24-shot setting of the public Jev reproduction, each query retrieves 24 labelled examples with BM25 and no weights are updated at test time. Two mechanisms read the same 24 examples: an in-context reader, Qwen3-4B with a small LoRA adapter trained only on other public intent datasets with Banking77 held out, which scores every class in one forward pass; and a frozen bge-large nearest-neighbour vote. A geometric mean combines them. The adapter is gated on Hugging Face. This pipeline is not the CPU jury: its reader has 4B parameters, and a GPU is recommended.
Keeping the numbers honest
- Only full official test sets are reported; small smoke tests overstate.
- The zero-shot NLI juror is built on a model whose training mix includes the Banking77 and MASSIVE train splits (not their test sentences), so its zero-shot results on those two datasets are not fully held out.
- Comparisons are matched-budget: zero-shot with zero-shot, 24-shot with 24-shot.
- Systems predicting on the same items are compared with a paired McNemar test.
- In the pre-registered runs, every rule, weight and temperature was chosen on held-out data, never on test; earlier exploration that did score on the test set is listed under Limitations.
- Jev is measured on exactly our test items and label text, after our configurations were frozen.
Limitations
- Clearly behind Jev zero-shot. On all five benchmarks we measured, Jev is 3.4 to 22.7 points more accurate with the same label text, and its probabilities also score better (Brier).
- The tie with Jev needs the 4B reader. Only the 24-shot pipeline, with a 4B reader, ties Jev. The sub-1B CPU jury is the zero-shot one.
- Banking77, MASSIVE and HWU64 zero-shot are not fully held out. The NLI juror's base model was trained on the Banking77 and MASSIVE train splits (no test sentences), and 44% of HWU64's test sentences appear in MASSIVE's train split. The other four benchmarks have no known exposure.
- No official Jev number. Jev publishes no Banking77 figure; ours are measured through its API (jev-1.13.0), and its 24-shot run reproduces the public reproduction's 92.40% exactly.
- 0.938 is exploratory history. An early configuration reached 0.938, but about 40 configurations were scored on the test set along the way, so it is not reported as a result.
- Not open zero-shot state of the art. A 7–9B open LLM will beat our zero-shot accuracy on some datasets (e.g. reported ~0.84 on CLINC150 vs our 0.65).
- Adapter contamination. The 24-shot adapter's training mix includes CLINC150, HWU64 and SNIPS, so don't use it to judge zero-shot generalization on those.
What did not work
- More training data saturates. Scaling the reader's fine-tuning from ~23K to ~60K examples did not raise zero-shot accuracy (flat around 0.70).
- A bigger embedder does not help. Swapping bge-large for Qwen3-Embedding-8B made the nearest-neighbour member weaker on this task.
- Reranking the top-k does not help. A cross-encoder reranker and a small generative reader over the top-5 both fail to beat the ensemble's own ranking.
Use Quorum in production
The zero-shot jury is open source and free to self-host today. It is three models under 1B parameters that run on a CPU, so your data never leaves your machines; it is less accurate than Jev (see Limitations). We are also preparing two ways to run it without the plumbing. Joining a waitlist is free and commits you to nothing; we will email you when there is something to try. The 24-shot pipeline (4B reader) is in the repo.
Paid API
A hosted Quorum endpoint. Send a message and your labels, get back the verdict, its confidence and the full distribution. Nothing to install.
On-prem
Quorum running inside your own network, so messages never leave it. We help you deploy and fit it to your labels.
You'll sign in first (Google, GitHub or an email link), then come straight back here. You can join one list or both.
Privacy. When you join, we store your account email, which list you joined and when, the optional company and use case you add, and the campaign link (UTM) that brought you here. We use it only to contact you about Quorum. We never ask for payment details. You can leave a list at any time with the “Leave” link, which deletes the entry. Page visits are counted anonymously, without IP addresses or account ids. Trying the demo sends the benchmark sentence to the Quorum API; its web server keeps standard access logs (including IP address).
Resources
Test sentences: Banking77, CLINC150, HWU64, MASSIVE, MTOP, SNIPS and Bitext, public research datasets under their own licenses.