Booster
Module

How Jevpardy Works

What happens when a tiny evaluation model races an LLM and a human through a trivia board?

SEP 28, 2026 9 MIN READ

Jevpardy is a trivia game with an intentionally uneven cast: a human player, a conventional large language model, and Jev, a small model designed for structured evaluation rather than free-form writing.

The premise is simple. Give all three contestants the same clue, start them at roughly the same time, and see who can produce a correct response first. The interesting part is not whether one model has memorized more trivia than another. It is what happens when three very different ways of arriving at an answer have to compete under one visible set of rules.

The simplified game

This is inspired by Jeopardy!, not an attempt to reproduce the television game. A game contains three randomly selected categories with five clues each, valued from 100 to 500 points. There are no Daily Doubles, wagers, or rebounds. As in the original game, an incorrect answer gives you negative points.

When you select a square, its clue is revealed. Once the reveal finishes, the race begins:

  • The LLM receives the category and clue and generates one complete Jeopardy-style response.
  • Jev starts assembling its response one character at a time.
  • You can buzz in whenever you think you know the answer, then type and submit it.

The first AI contestant to finish takes the buzzer automatically. A human takes it by clicking first. Everyone else is locked out for that clue. The winning response goes to a judge; a correct answer adds the clue's value to that contestant's score, while an incorrect answer gets points deducted. Either way, the clue is complete and play returns to the board.

That simplification makes the comparison legible. Each round has one clue, one race, one answer, and one verdict.

Why the race is interesting

The LLM contestant is doing the familiar thing: it receives a prompt and generates text autoregressively. It currently uses OpenAI's GPT-5.4 Nano. One request produces one answer.

Jev is not being used as a text-generation model. Its natural interface is a constrained question with structured possible outcomes. Ask it to choose among several labels or estimate a binary proposition, and it returns probabilities over those outcomes. That is useful for classification and evaluation, but it does not naturally provide a sentence to display on a game show.

So Jevpardy turns text generation into a sequence of small decisions. The result is deliberately inefficient, visible, and a little suspenseful. The LLM tends to think in one opaque burst. Jev's answer crawls onto the screen letter by letter. A human may recognize the answer before either system finishes, but still has to decide when there is enough confidence to buzz.

This is not a benchmark. Network latency, model latency, interface timing, and the chosen decoding rules all affect the winner. It is closer to a playable systems demo: the constraints are part of the experience, and the experience makes those constraints easier to notice.

To me, the biggest surprise was that it's a close race. Giving Jev a choice of 10 possible answers is easy and fast, and it would always finish first. But that's not interesting. Having it build the answer one character at a time tends to finish around the same time as the LLM's response, but we had to choose a smaller model. The upshot is that it doesn't always get the answer right -- which makes the game even more interesting. The final element was to make the questions on the slightly easier side. While this doesn't impact the speed of the AI responses, it makes the human just competitive enough to make the game a very close race.

Turning evaluation into generation

Jev's answer algorithm has two stages. First it chooses a conventional Jeopardy opening. Then it repeatedly chooses the next character until it decides the response is complete.

Choosing the opening

The first request asks Jev to select one of four phrases:

what is
what are
who is
who are

The category and clue are included as state, and each phrase has a short description explaining when it is appropriate. This gives the response a grammatical starting point without spending many character-level calls rediscovering the format.

Proposing the next character

After the opening, each step builds a finite candidate set:

STOP
a b c ... z
SPACE

Each candidate is presented as the current answer with that candidate appended. STOP is represented by the unchanged answer. Jev sees the category, the clue, the exact answer written so far, and an instruction to choose the continuation that looks most like the beginning of a concise, correct, grammatical response.

For example, if the response so far is what is par, the options effectively ask Jev to compare what is para, what is parb, and so on, plus a space and the unchanged string. The highest-scoring continuation becomes the next character.

The full candidate strings are clipped to their most recent 40 characters before being sent. That keeps the options compact while preserving the part of the answer where the next character matters most.

Reducing option-order noise

Classification models can be sensitive to the order in which labels appear. To dampen that effect, every character step creates four copies of the same question with independently shuffled option orders. They are submitted together, then their returned probabilities are mapped back to the canonical candidates and averaged.

In rough pseudocode, one step looks like this:

variants = shuffle_the_options_four_times(answer_so_far)
results = ask_jev(variants)
probabilities = average_results_by_character(results)
next = highest_probability(probabilities)

This small ensemble costs more than a single classification, but it makes the output less dependent on an arbitrary presentation detail.

Stopping without stopping too early

STOP competes with every character at every step, but a few decoding rules shape the raw probabilities:

  • Stopping is disabled for the first three generated characters.
  • A second consecutive space is not allowed.
  • The probability of STOP is multiplied by 0.5, making Jev require stronger evidence that the answer is finished.
  • Responses are capped at 30 generated characters after the opening.

The decoder is greedy: after applying those rules, it chooses the highest-probability candidate. If that candidate is STOP, the answer is complete. Otherwise the chosen character is appended, streamed to the browser, and the process repeats.

The stream matters to the game. The server sends each accepted character immediately as a newline-delimited event, so the UI can show the answer forming in real time. If Jev finishes before the LLM response returns or the human buzzes, Jev wins the race and the other in-flight work is cancelled.

Using Jev as the judge

Jev also decides whether the winning response should count. The judge receives the category, clue, stored expected answer, and contestant response. It evaluates a binary proposition: should this response be accepted as correct?

That is a much more natural Jev task than writing an answer. The output is a probability for true versus false; the game accepts the response when the probability of true is at least 0.5. The judging instructions explicitly ignore capitalization, punctuation, optional Jeopardy phrasing, and minor spelling differences while rejecting answers that are materially incomplete or refer to something else.

Using an AI judge is reasonable here because exact string comparison would be brittle. Paris, What is Paris?, and a harmless typo should not produce three different outcomes. The decision is constrained, the expected answer is available, and every contestant is evaluated by the same function.

It is not perfectly neutral. Jev is judging answers produced by Jev, and any model can be overconfident or inconsistent around aliases and ambiguous clues. For a lightweight game that tradeoff is acceptable; for a serious evaluation, the judge would need independent calibration and human-reviewed test cases.

What I would improve next

There are several obvious ways to make both the game and the experiment stronger:

  • Calibrate the judge. Build a held-out set of correct, partially correct, misspelled, and wrong responses, then tune the threshold against human decisions.
  • Use an independent judge. A separate model—or an ensemble that includes deterministic alias rules—would reduce the appearance and risk of Jev grading its own homework.
  • Improve the decoder. The current alphabet cannot emit punctuation or digits, the 30-character cap can truncate long answers, and greedy decoding cannot recover from an early mistake. Beam search or word-level candidates could produce better answers with fewer calls.
  • Measure a fairer race. Today the competition includes provider and network latency. Recording time-to-first-confident-answer separately from end-to-end latency would distinguish reasoning speed from infrastructure speed.
  • Add rebounds and penalties. Letting another contestant try after a wrong answer would create more strategy and make early buzzing meaningfully risky.
  • Audit clue quality. Difficulty labels, alternate accepted answers, and ambiguous wording deserve systematic review as the question bank grows.

The deliberately awkward character loop is the point of the project. Jevpardy takes a model built to evaluate bounded choices and stretches that capability until it can participate in a generative task. Putting it beside an LLM and a human makes the differences concrete: not just what each contestant knows, but how its interface, latency, confidence, and failure modes shape what it can do.

Related Essays

View all essays