Booster
Module

Jev as Judge: Evaluating a Trivia Answer Judge

How I built a data-driven eval for Jev-based trivia judges.

SEP 29, 2026 14 MIN READ

Jevpardy is a fun little game I created that pits you against Jev and a traditional LLM in a race to answer (simple) trivia clues. Jev was also used to judge the answers. In this article, I'll talk about how I evaluated Jev as a judge and the challenges that posed. Strap in; it's a long one!

The answer is only half the game

In Jevpardy, a human, a conventional language model, and Jev receive the same clue and compete to produce an answer first. That race created a second problem that was less visible but just as important: once someone answers, how does the game decide whether the answer should count?

That decision is a bit harder than comparing two strings. Washington, D.C. and What is Washington DC? should normally receive the same verdict. Sagarmatha should count for a clue whose stored answer is Mount Everest, even though the two strings share no useful words. A response can also contain the expected answer while rejecting it: It is not Mount Everest.

I built an evaluation to compare three ways of making the decision: a heuristic baseline, a direct Jev judge, and a Jev judge that normalizes the input first. You should consider this a practical evaluation for a playable demo, rather than a benchmark for trivia systems in general. The goal was to replace “this feels reasonable” with a small set of examples, measurements, and tradeoffs that I could inspect.

Why judging answers needs its own eval

Exact string matching gives consistent results, but often handles ordinary variation poorly. It treats capitalization, punctuation, and harmless framing as meaningful differences:

  • Washington, D.C.
  • Washington DC
  • What is Washington DC?

A person would normally recognize all three as the same answer. A spelling mistake may also be acceptable in an ordinary trivia category, while an alternate name such as Sagarmatha requires knowledge rather than text normalization.

Punctuation is doing enough work in Washington already; it probably should not decide the score too.

The opposite shortcut is risky too. Finding the expected answer somewhere inside the response is insufficient. It is not Mount Everest contains Mount Everest, as does Mount Everest or Mount Fuji, but neither gives Mount Everest as an unqualified answer.

That leaves a game judge balancing two kinds of mistake:

  • A false acceptance awards points for a response that should have been rejected.
  • A false rejection denies points for a response that a person would recognize as correct.

The right balance is partly a product decision. A strict tournament might prefer to reject borderline responses. A casual web game may care more about avoiding the frustration of marking a recognizable answer wrong. Either way, both errors need to be measured. Accuracy alone cannot reveal whether a judge is generous, strict, or simply inconsistent.

Building the evaluation data

Starting with nested question sets

Jevpardy's question bank contained roughly 2,000 clues. Running every early experiment over the entire bank would have slowed iteration without necessarily making it more informative, so I selected three reproducible, stratified subsets:

  • A 500-question set covered each available category-and-value combination.
  • A 200-question set selected two clues from every category and balanced clue values across the set.
  • A 25-question set used distinct categories and balanced values for quick smoke tests.

The sets were nested: every question in the 25-question set also appeared in the 200-question set, and every question in the 200-question set appeared in the 500-question set. This structure let me move from a quick check to a routine run to broader confirmation without quietly switching samples along the way.

I used the 200-question subset for the judge evaluation.

Turning one answer into many judging decisions

Each source record already supplied a category, a clue, and its stored expected answer. Evaluating a judge also required plausible contestant responses and a reference label for each one: accepted, rejected, or ambiguous.

I began with deterministic variants. The accepted group included the canonical answer, Jeopardy-style phrasing, and changes to capitalization or whitespace. For clear negatives, I reproducibly borrowed expected answers from unrelated clues and added constructions such as a negated correct answer. I also created borderline cases, including shortened answers and responses containing two conflicting alternatives.

Those rules check the plumbing but create little linguistic variety. For each question, I therefore made one LLM call and requested a list of additional responses. Gemini 2.5 Flash Lite and GPT-5.4 Nano generated plausible aliases, alternate names, misspellings, near misses, and more adversarial wrong answers. Generating a list in one call kept the process simple and avoided a separate request for every variant.

I then deduplicated the responses and resolved cases where generation had placed identical text under different labels. After the first run of the heuristic baseline, I performed a broad audit of the generated data: obviously invalid accepted responses were corrected, clearly valid rejected responses were moved, policy-dependent cases became ambiguous, and formatting debris or clue paraphrases were removed.

That cleanup limits the strength of the conclusions. The resulting reference labels are useful evaluation targets, but they are not independently established truth, and the baseline later being measured had already influenced the audit. A stronger future version would reserve an untouched sample for blind human review. A frontier model that cannot see the judges' predictions could also provide another adjudication signal, though I would reserve the term “gold labels” for independently reviewed human judgments, ideally with consensus or a second pass.

This dependence on generated labels will come up again when the results show a suspiciously perfect precision score for the heuristic baseline.

What the transformation looks like

The following simplified examples show how a single question-bank entry becomes several judge-evaluation cases. Each row contains only a few representative variants from what could be a larger generated record.

Category and clueStored answerAccepted reference examplesAmbiguous reference examplesRejected reference examples
U.S. capitals: “The White House is located in this U.S. capital.”Washington, D.C.What is Washington DC?
Wasington, D.C.
Washington
Washington, D.C. or Paris
Paris
Tokyo
It is not Washington, D.C.
Geography: “At 29,032 feet, this is the world's highest mountain above sea level.”Mount EverestChomolungma
Sagarmatha
Mount Everst
Mount Everest or Mount Fuji
Everest Mountain
Mount Fuji
K2
the Himalayas
Spelling: “Spell the word for a month that comes after January.”FebruaryFEBRUARY
What is February?
Feb
Febr
Febuary
January
March

The spelling example highlights why the clue must remain part of the judging input. Febuary might be an acceptable typo when a general-knowledge clue asks for the second month of the year. It should fail when spelling the word is the task. Similarity to the stored answer alone cannot always determine the right judgment.

What I measured

A judge can fail in two directions, and they do not feel the same to a player. It can award points for a wrong response, or reject a response that a person would plainly recognize as correct. A single accuracy number blends those mistakes together, so I tracked several measurements.

The basic accounting is a confusion matrix, although the name is more intimidating than the idea. For every response with an accepted or rejected reference label, I compared that label with the judge's decision. A response can land in one of four buckets: correctly accepted, incorrectly accepted, correctly rejected, or incorrectly rejected.

From those buckets, I calculated:

  • Accuracy: What fraction of all labeled responses did the judge classify correctly? If it gets 90 out of 100 decisions right, its accuracy is 90 percent. This is a useful summary, but it can conceal which kind of mistake the judge makes.
  • Accepted-answer precision: When the judge awards points, how often was the response labeled acceptable? High precision means that a “correct” verdict is usually trustworthy. A judge that accepts only the stored answer verbatim might have excellent precision while still being unpleasant to play against.
  • Accepted-answer recall: Of all the responses labeled acceptable, how many did the judge actually accept? Recall captures aliases, alternate phrasings, and harmless misspellings that a strict judge might reject.
  • False-accept rate: Of the responses labeled wrong, how many slipped through as correct? This is the most direct measure of how often the game gives away points. Lower is better.
  • False-reject rate: Of the acceptable responses, how many were marked wrong? This is the complement of accepted-answer recall, expressed in the language of errors. Lower is better here too.

My hand-written baseline had a third option: it could defer a response as ambiguous rather than force an answer. For that judge, I also tracked coverage, the fraction of cases on which it made a definite accepted-or-rejected decision, and the corresponding deferral rate. A system can manufacture impressive precision by declining every difficult case, so coverage matters when interpreting its headline numbers.

Jev returns a probability rather than only a yes-or-no verdict. That let me record a Brier score, the average squared distance between Jev's confidence and the binary reference label. If the correct label is 1 and the model reports 0.9, the contribution is (1 - 0.9)², or 0.01. Confident mistakes receive a much larger penalty than uncertain ones, and lower scores are better. I treated this as a rough calibration check rather than the main result of the experiment.

I excluded responses labeled ambiguous from accuracy, precision, and recall because there was no definitive label for a judge to agree with. I still counted how often each algorithm accepted or rejected them. A large shift can show that one judge is much more permissive than another, but it does not show that either judge is more correct.

Finally, I split the measurements by how the response variants were created. The deterministic cases include obvious plumbing checks: capitalization changes, Jeopardy-style phrasing, answers borrowed from unrelated clues, and explicit negation. They verify that an implementation is functioning, but they are intentionally easy. The model-generated variants contain more aliases, near misses, partial answers, and plausible-but-wrong alternatives. That split tests semantic judgment more directly. I report both so a large number of easy examples cannot obscure performance on the cases that motivated an AI judge in the first place.

Three ways to judge an answer

I compared one hand-written baseline with two Jev-based algorithms. All three received the clue's category, its text, the stored expected answer, and a candidate player response. They differed in how much processing happened before the decision and how the final comparison was made.

1. A heuristic baseline

The first judge uses ordinary string-processing rules and no language model. It lowercases text; normalizes punctuation, whitespace, and optional articles; and removes openings such as “What is.” After normalization, it accepts exact matches. A limited edit-distance rule also permits small spelling errors when the category does not specifically test spelling.

Some errors are safer to detect than others. An explicit negation such as “It is not Mount Everest” can be rejected, as can an unrelated answer. Containment alone is unsafe: “Mount Everest or Mount Fuji” contains the expected answer and still should not receive points. The baseline therefore defers conflicting, partial, and otherwise uncertain responses as ambiguous.

I designed this judge to be conservative. It establishes how far inexpensive, understandable string rules can get before I credit a model with semantic judgment. Its limitations are informative too. Punctuation stripping cannot reveal that “Sagarmatha” is another name for Mount Everest, or that an answer is correct under one clue but too broad under another.

The word “baseline” matters here. This is neither a standard judging algorithm used across trivia systems nor the strongest possible non-model implementation. It is a transparent point of comparison tailored to the responses in this game.

2. The direct Jev judge

The second judge sends the material to Jev largely as it arrived: category, clue, expected answer, and the contestant's response, including its original capitalization, punctuation, and Jeopardy-style phrasing. The prompt asks a binary question: in the context of this clue, should the contestant's response be accepted as correct?

Jev assigns a probability to true. For the initial comparison, I accepted a response when that probability was at least 0.5. Saving the probability, rather than only the thresholded verdict, makes sense because the threshold is a product-policy choice. If I later decide that false accepts are more costly, I can raise it without paying to run every example again. This actually turned out to be very useful because the 0.5 threshold turned out to be too conservative, and I subsequently revised it to 0.7.

The clue context takes this beyond fuzzy string matching. Consider a stored answer of “Mercury.” Whether “the closest planet to the Sun” is a helpful equivalent, an overlong but acceptable response, or evidence that the contestant never named the answer depends on the clue and the rules of the game. A geographic alias can also be semantically correct while sharing no characters with the expected answer. The direct judge gives Jev that context and asks it to make the call.

3. The normalized-input Jev judge

The third algorithm tests whether removing superficial formatting differences helps Jev focus on meaning.

Before making the call, it lowercases the expected answer and contestant response, removes punctuation, collapses repeated whitespace, and strips leading forms such as “what is,” “who is,” “what,” and “who.” The category and clue remain unchanged. A response like What is Washington, D.C.? therefore reaches the comparison as something close to washington dc.

Its prompt also tells Jev that the two answer strings have already been normalized. The aim is to keep punctuation and game-show phrasing from distracting from semantic equivalence.

There is an important experimental caveat: I changed two things at once. The input normalization changed, as did the wording that explained those inputs to Jev. The result compares two complete algorithms; it does not isolate the causal effect of lowercasing or punctuation removal. Any behavioral difference would require a follow-up experiment with a fixed prompt before I could attribute it to preprocessing alone.

Together, the approaches provide three reference points: basic string rules, a model's judgment over the original response, and a model's judgment after preprocessing. Their accuracy scores are only part of the comparison. The error breakdown shows which valid responses each approach recovers and which wrong responses it begins to accept in exchange.

Results: better recall, with a tradeoff

I evaluated all three judges at a Jev acceptance threshold of 0.5. This first view includes every response with an accepted or rejected reference label. I excluded ambiguous responses because I do not yet have a defensible correct verdict for them.

JudgeAccuracyAccepted-answer precisionAccepted-answer recallFalse-accept rate
Heuristic baseline89.8%100.0%79.9%0.0%
Direct Jev96.4%97.0%95.5%2.8%
Normalized-input Jev96.1%97.2%94.7%2.5%

The heuristic baseline looks excellent on precision: every response it accepted agreed with the reference labels. That result comes from caution. It rejected or deferred many answers that required recognizing an alias, an alternate name, or another semantic equivalence beyond the reach of string comparison. Its 79.9% recall means that the baseline did not accept roughly one in five responses labeled acceptable.

Both Jev judges recovered most of those responses. Direct Jev raised accepted-answer recall to 95.5%, while normalized-input Jev reached 94.7%. The cost was a small but real number of false acceptances: 2.8% for direct Jev and 2.5% for normalized-input Jev.

For Jevpardy, that looks like a useful trade. A judge that rejects a familiar alternate name can make a trivia game feel arbitrary, even if its caution produces an impressive precision number. I would rather recover many of those recognizable answers while accepting a few more borderline ones. That is a product decision, not a general result about how every quiz, exam, or competition should be scored.

The aggregate table also includes deterministic variants designed to check the basic machinery: capitalization changes, Jeopardy-style phrasing, sampled wrong answers, and explicit negations. Those cases are intentionally straightforward. The model-generated variants are more revealing because they include aliases, misspellings, partial answers, and plausible but incorrect alternatives.

JudgeAccuracyAccepted-answer precisionAccepted-answer recallFalse-accept rate
Heuristic baseline79.2%100.0%57.3%0.0%
Direct Jev92.7%93.4%90.6%5.5%
Normalized-input Jev92.2%94.0%88.9%4.9%

On this harder split, the contrast is sharper. The heuristic baseline accepted only 57.3% of responses labeled acceptable. Direct Jev reached 90.6% recall, and normalized-input Jev reached 88.9%. Their false-accept rates also rose, to 5.5% and 4.9% respectively. The easy formatting cases therefore do not account for most of Jev's apparent advantage. The larger difference appears in the semantic cases that motivated a model judge in the first place.

Comparison of the heuristic baseline, direct Jev, and normalized-input Jev on model-generated response variants

Figure 1. Performance on model-generated response variants at a 0.5 Jev acceptance threshold. Ambiguous cases are excluded, and accepted/rejected outcomes are evaluated against synthetic reference labels. The chart uses four vertically stacked panels so the comparison remains readable on a narrow screen.

The preferred error tradeoff depends on the product. A medical or financial classifier may treat a false positive as much more serious than a false negative. In a trivia game, rejecting a recognizable alias in front of the player is immediately frustrating, while an occasional borderline acceptance usually does less damage to the experience. That difference is why I care about precision and recall together rather than looking for one universally correct balance.

Why the baseline's 100% precision needs an asterisk

The baseline's perfect precision is useful evidence that its acceptance rules were conservative, but it does not establish a universally perfect deterministic judge.

First, the baseline can defer uncertain cases rather than make a binary decision. Avoiding hard cases naturally protects precision. Second, when I audited the generated labels, I corrected some minor-typo cases using the same policy encoded by the baseline. That creates a degree of circularity: the labels are not fully independent of the algorithm being measured. Finally, the easy deterministic variants lift every judge's aggregate score. The model-generated split and the concrete disagreements therefore matter more than the 100.0% headline by itself.

Within those limits, the comparison still gives me useful directional evidence. Against the same fixed reference labels, both Jev approaches exchanged a few false accepts for a large reduction in false rejects. That is the behavior I wanted from the model judge. The result does not show that Jev will judge every real player response correctly. It does support trying the model-based approach in the game instead of relying only on string rules.

Did preprocessing help Jev?

The normalized-input experiment made less difference than I expected. At the same 0.5 threshold, it was slightly more conservative than direct Jev: accepted-answer precision increased by about 0.2 percentage points, and the false-accept rate fell by about 0.25 points. At the same time, recall fell by about 0.8 points and accuracy by about 0.26 points.

Those small aggregate movements came from only 32 changed decisions among 2,348 labeled responses. Thirteen changes corrected a mistake made by direct Jev; nineteen turned a previously correct decision into a mistake. On the labeled data, normalization therefore produced a small net regression rather than a clear improvement.

It had a larger effect on the ambiguous cases. Direct Jev accepted 828 of the 1,177 ambiguous responses, while normalized-input Jev accepted 608. That suggests that normalization, the revised prompt, or their combination made the second judge more cautious around borderline answers. I cannot call that an accuracy improvement: these are precisely the responses for which I have not assigned accepted or rejected reference labels.

This experiment also changed two things at once: the input normalization and the prompt, so it does not isolate which one caused the difference. I chose the direct Jev judge for the implementation. Its recall and overall results were marginally better, and it avoided a preprocessing layer that did not demonstrate a meaningful gain. The normalized variant still answered a useful question: simple cleanup may add little when the model can already interpret the original response in context.

What the evaluation cost

Each Jev-based algorithm made one call for every candidate response in the evaluation set: 3,525 calls for the direct judge and another 3,525 for the normalized-input judge. The direct run consumed 1.45 million input tokens and cost about $0.061. The normalized run consumed 1.52 million input tokens and cost about $0.064. Together, the two complete evaluations cost roughly $0.125 at the current price of $0.042 per million input tokens; Jev output tokens were free.

The normalized answers were shorter, yet the second run was slightly more expensive. Its prompt needed additional instructions explaining the preprocessing and how the judge should interpret the normalized text. Across thousands of calls, that fixed prompt overhead outweighed the few characters removed from each answer.

The cheaper-looking prompt had apparently brought luggage.

I saved Jev's probability for every decision along with the final yes-or-no verdict. That let me compare different acceptance thresholds afterward without making another 3,525 calls. Changing the threshold became a local calculation instead of a new model evaluation.

Where this evaluation is weak

I built this evaluation to choose a judge for a game. Its results describe one fixed dataset and three particular implementations; they say much less about Jev's general ability to understand trivia answers.

The reference data is the largest source of uncertainty. Gemini 2.5 Flash Lite and GPT-5.4 Nano generated most of the interesting answer variants. I gave them a broad audit; humans did not adjudicate every response. I corrected obvious mistakes and moved policy-dependent cases into an ambiguous bucket. That cleanup made the labels more useful. They are still synthetic reference labels, with the uncertainty that implies.

The ambiguous responses deserve particular attention. They contain many cases on which reasonable judges could disagree, but they have no accepted-or-rejected verdict. I therefore excluded them from accuracy, precision, and recall. The count of how each algorithm handled them is still a useful diagnostic: for example, it showed that the normalized-input judge behaved much more conservatively. It cannot tell me whether that conservatism was correct.

By the final measurements, the evaluation set had also served as development material. I inspected errors, cleaned labels, and refined the judging code against the same responses. In one especially relevant instance, the heuristic baseline's treatment of minor typos influenced some label corrections. Its perfect accepted-answer precision is partly circular as a result. Repeatedly looking at a dataset can shape decisions around its examples even without model training. The final numbers consequently resemble development-set results more closely than held-out test results.

The unit of measurement introduces another wrinkle. Every response variant contributes one row to the metrics, which gives questions with more generated variants more influence than questions with fewer. The reported numbers describe performance per synthetic response. They do not directly describe performance per clue or per game. The dataset also has a deliberately challenging mix of aliases, misspellings, negations, and adversarial answers. Real players may produce those answer types at very different rates. A challenge set is useful for finding failures; it does not estimate how often players will encounter each one.

I think of this synthetic distribution like a crash-test laboratory: deliberately overrepresenting difficult cases helps expose weak points, while field data is still needed to estimate how often those failures matter.

The third algorithm combines two changes. I normalized its input and changed its prompt to explain that normalization. The comparison therefore measures the algorithms as complete packages. It cannot isolate whether preprocessing caused a change, whether the prompt did, or whether both contributed.

I also ran each response through one served model version once. This leaves run-to-run variation and future model changes unmeasured. Under those conditions, differences below 1% (such as several of the gaps between direct Jev and normalized-input Jev) should carry little weight.

There is one Jevpardy-specific risk that this dataset only partly captures. In the finished game, Jev judges answers that were themselves produced through Jev. Shared preferences or blind spots could affect both the answer and the verdict, producing correlated errors. Variants from two other models add some independence to this evaluation, but they do not recreate the full Jev-answering-Jev loop. The results therefore cannot rule out a judge confidently approving an answer that reflects its own misconception.

The arrangement has a faint air of a student grading an exam from an answer key the student also wrote.

A practical path to stronger evidence

The most valuable next step would be a small, untouched test set with stronger labels. I would sample real and generated responses across clear accepts, clear rejects, disputed cases, and disagreements between judges. Human reviewers would see the clue, expected answer, and proposed response without seeing which algorithm produced each verdict. Difficult cases could receive a second review under a written judging policy. A frontier model could provide another independent opinion, although I would describe that output as model-adjudicated reference data rather than gold labels. The resulting held-out set could support the final comparison, while a separate development set remained available for prompt and threshold tuning.

Sampling deserves care here. Filling the test set entirely with disagreements would make it an excellent stress test, but it would exaggerate their prevalence. I would combine a representative random sample with a smaller disagreement-heavy slice and report the two separately. That preserves a view of likely game performance while keeping the examples most capable of distinguishing the judges.

Real player responses would make the test distribution more representative. Once the game has traffic, I could sample actual submissions while preserving how frequently different answer types occur. Reporting both per-response results and per-question averages would stop a clue with many variants from dominating the overall score. Breaking errors into categories, specifically: aliases, spelling, partial answers, negation, conflicting alternatives, and clue paraphrases would also make it clear where an improvement came from.

Several narrower experiments would clean up specific uncertainties. Running raw and normalized answers through an otherwise identical prompt would isolate preprocessing. Repeating calls, or pinning a model version when that option exists, would reveal stability. Bootstrap confidence intervals would put the sub-one-percentage-point differences in context. I would also keep threshold tuning and final testing on separate datasets so that the chosen confidence cutoff is evaluated on responses that did not help select it. Saving the model probabilities, as I did in these runs, makes threshold comparisons cheap: changing the policy cutoff requires recomputing metrics, not repeating every model call.

Those additions could support a more deliberate judging pipeline. Cheap deterministic rules could settle exact matches and other high-confidence cases. Jev could handle answers that require semantic context. Unresolved evaluation examples could enter an offline human-review queue and improve the reference set over time. For this demo, the present evaluation already supports a reasonable engineering choice. Better labels, representative traffic, and a held-out test would make the strength of that evidence much easier to judge.

What I learned

Writing down a cheap, understandable baseline turned out to be the most useful first step. The heuristic judge gave me something concrete to beat and exposed the product tradeoff along the way: it avoided awarding points for wrong answers by rejecting or deferring many answers that a person would recognize as valid. A score from an AI judge would have been much harder to interpret without that comparison.

The evaluation set only had to represent the decisions this game would encounter: formatting differences, misspellings, aliases, partial answers, negations, and responses containing conflicting alternatives. That small challenge set let me replace “this judge feels reasonable” with claims I could inspect. The disagreements and individual errors were at least as informative as the aggregate metrics. Two judges with similar scores can still create very different experiences for players.

Saving Jev's probabilities also paid off. I could explore different confidence thresholds without running the model again, while keeping the product policy separate from the model output. A game that strongly wants to avoid false accepts can choose a higher threshold. Another application might tolerate that risk to reject fewer valid responses.

The preprocessing experiment was a useful reminder to test seemingly obvious improvements. Lowercasing the answers and stripping punctuation and question phrasing sounded helpful, yet the normalized-input judge did not meaningfully improve on the simpler direct-input version. The experiment was cheap, and its result spared me from carrying extra complexity on intuition alone.

This evaluation was modest and imperfect, but it provided enough evidence to choose a reasonable judge for the game. The test I care about is no longer whether the judge recognizes one memorable alias such as Sagarmatha, but whether it handles that entire class of answer without becoming permissive elsewhere. Greater confidence would require cleaner held-out labels, blind human review, and responses from real players. For this stage of Jevpardy, I now know why I chose the direct Jev judge, where it is likely to fail, and what evidence I would collect before trusting it more broadly.

---

You can try Jevpardy for yourself and let me know what you think. If you want to see an example of a Jev use use that has more practical applications, check out the Jev Composer. I can be reached at hello [at] boostermodule [dot] com or u/boostermodule on Reddit.

Related Essays

View all essays