Insights
Sep 29, 2026
10
min read

How to turn an LLM into a low-latency classifier on the fly

Sahil Garg
One token, one classification: a prompt goes through an open LLM in one forward pass, producing the probability of every class

Many applications rely on multi-class text classification, which assigns each input to exactly one class from a fixed set. Dedicated services such as Jev, a closed-source API built specifically for classification, are fast because they return a class and a confidence score rather than generated text. It is well known that a general-purpose language model can do the same. What is used less outside machine learning is two by-products of the same forward pass: the probability of every class, and a confidence score read from those probabilities.

The technique requires no fine-tuning, no classification head and no output parser: the prompt describes the classes and asks for a one-letter answer, and a single forward pass returns the probability of every class. Below we walk through the recipe in two steps and measure it on one task, routing cybersecurity prompts. On that task, two Gemma models match Jev's accuracy and answer with lower latency at light load. Their confidence score is 1 − p₂, one minus the runner-up class's probability. Escalate the least confident classifications by this score and the Gemma models' error rate on the rest keeps falling, past the point where Jev's own confidence score stops lowering Jev's.

The recipe

Step 1: Describe the classes and read their probabilities from one token

One classification. The prompt, which holds the class descriptions and the text to classify, enters the open model; a single forward pass yields the next-token probabilities of the 22 letters, one per class (the bars). The most probable letter, here B at 0.77, is the classification.
One classification. The prompt, which holds the class descriptions and the text to classify, enters the open model; a single forward pass yields the next-token probabilities of the 22 letters, one per class (the bars). The most probable letter, here B at 0.77, is the classification.

Give each class a single capital letter, A to V for our 22 classes, and describe each class in the prompt, including what does not belong in it and how to tell it apart from classes it is easily confused with.1 The text to classify follows, with an instruction to answer with one letter. Then have the model generate exactly one token and read the probability it assigns to each letter.2 In the figure, B has the highest probability, 0.77, and is the predicted class. Each classification costs one forward pass, as with a dedicated classifier.

Step 2: Read a confidence score and set an escalation threshold

Left: the probabilities from Step 1 and two confidence scores read from them, the top probability (p₁ = 0.77) and one minus the runner-up probability (1 − p₂ = 1 − 0.15 = 0.85). Right: classifications ranked by confidence score; those below the escalation threshold are escalated, and the model's own classification is kept for the rest.
Left: the probabilities from Step 1 and two confidence scores read from them, the top probability (p₁ = 0.77) and one minus the runner-up probability (1 − p₂ = 1 − 0.15 = 0.85). Right: classifications ranked by confidence score; those below the escalation threshold are escalated, and the model's own classification is kept for the rest.

The same probabilities also measure how decisive the model was. Several confidence scores can be read from them, for example the top probability p₁ or one minus the probability of the runner-up class, 1 − p₂; the results section compares the two. To decide what to escalate, rank the classifications by the chosen score and send those below an escalation threshold to a larger model or a human reviewer, as the right half of the figure shows.3 The error rates we report under escalation in the results cover only the model's own answers on the classifications it keeps: escalated classifications are left out of the count, and no answer from a larger model or a reviewer is included.

Results: routing cybersecurity prompts

Our router assigns each incoming cybersecurity prompt to one of 22 task groups and dispatches it to the model best suited to that group. We compared three classifiers on the same labeled held-out set, each receiving identical input text:

  • Gemma 4 26B A4B, Google's open mixture-of-experts model, used as released.4
  • Gemma 4 E2B, a small language model (SLM) we had already trained with reinforcement learning as a language model that writes out its reasoning, not as a classifier; here it is read exactly like the 26B: one output token, no generated reasoning.5
  • Jev, called through its API.

In deployment we read the E2B the same way and use it as a fast task-drift check within agentic sessions, while the primary router can be a different model. Each Gemma model was served on a single NVIDIA GH200 GPU.6

Accuracy and latency

Median latency in ms (y axis, log scale) against parallel requests in flight, 1 to 128 (x axis), for Gemma 4 26B A4B and Gemma 4 E2B, each on its own GH200, and for Jev over the internet. Bands: slowest to fastest of three runs. Inset: accuracy on every prompt of the labeled held-out set, with nothing escalated.
Median latency in ms (y axis, log scale) against parallel requests in flight, 1 to 128 (x axis), for Gemma 4 26B A4B and Gemma 4 E2B, each on its own GH200, and for Jev over the internet. Bands: slowest to fastest of three runs. Inset: accuracy on every prompt of the labeled held-out set, with nothing escalated.

The three classifiers are on par in accuracy, measured on every held-out prompt with nothing escalated: 97% for Gemma 4 26B A4B, for Jev and for Gemma 4 E2B. With one request in flight, the median latency is 91 ms for the 26B and 47 ms for the E2B, against 325 ms for Jev. As parallel requests grow, each model's GPU saturates and requests queue; even so, the 26B's median latency stays below Jev's up to about a dozen parallel requests, and the E2B's up to a point between 32 and 64.

Throughput

Requests completed per second (y axis, log scale) against parallel requests in flight, 1 to 128 (x axis), for Gemma 4 26B A4B and Gemma 4 E2B, each on its own GH200, and for Jev over the internet. Bands: lowest to highest of three runs.
Requests completed per second (y axis, log scale) against parallel requests in flight, 1 to 128 (x axis), for Gemma 4 26B A4B and Gemma 4 E2B, each on its own GH200, and for Jev over the internet. Bands: lowest to highest of three runs.

On one GH200, throughput levels off at about 50 requests per second for the 26B and about 170 for the E2B; Jev's was still rising, at 352 requests per second at 128 parallel requests, the highest load shown. The Gemma ceiling is the hardware we allocated, not the method. More GPUs or replicas would raise it; we did not measure that.

Probability on the letters

Share of held-out prompts (y axis) by the next-token probability mass that falls outside the 22 letters (x axis, one pair of bars per order of magnitude), for Gemma 4 26B A4B and Gemma 4 E2B. The dashed line marks 1%; the annotation counts the prompts beyond it, 2 for the 26B and 47 for the E2B.
Share of held-out prompts (y axis) by the next-token probability mass that falls outside the 22 letters (x axis, one pair of bars per order of magnitude), for Gemma 4 26B A4B and Gemma 4 E2B. The dashed line marks 1%; the annotation counts the prompts beyond it, 2 for the 26B and 47 for the E2B.

A confidence score read from the letters is meaningful only if the next-token distribution puts nearly all of its probability mass on them, and the figure shows that both Gemma models do. For the median held-out prompt, about 2 in 10 million of the 26B's probability mass falls outside the letters, and about 3 in 100,000 of the E2B's. Fewer than 0.2% of prompts put more than 1% outside the letters for either model, with no additional tuning. Thus, the letter probabilities can be read as returned.7

Confidence and escalation

Error rate of each system's own classifications on the prompts it keeps (y axis, log scale) against the share it escalates, least confident first (x axis, 0% to 30%), on the same labeled held-out set. Escalated prompts are left out of the error rate, not answered by another model. Gemma 4 26B A4B and Gemma 4 E2B rank by 1 − p₂; Jev by its own confidence score. Bands: 95% bootstrap intervals. The hatched region between Jev's curve and the 26B's curve, from 10% escalated onward, is the difference in error rate between the two.
Error rate of each system's own classifications on the prompts it keeps (y axis, log scale) against the share it escalates, least confident first (x axis, 0% to 30%), on the same labeled held-out set. Escalated prompts are left out of the error rate, not answered by another model. Gemma 4 26B A4B and Gemma 4 E2B rank by 1 − p₂; Jev by its own confidence score. Bands: 95% bootstrap intervals. The hatched region between Jev's curve and the 26B's curve, from 10% escalated onward, is the difference in error rate between the two.

Jev returns its own confidence score, so for each system we escalate its least confident X% of classifications by its own score (the figure's x axis) and measure the error rate of its own classifications on the prompts it keeps. The escalated prompts are removed from the measurement rather than passed to a larger model, so the curves show how well each score separates a model's errors from its correct answers, not the accuracy of a system that includes a larger model.8 With nothing escalated, that error rate is 2.6% for the 26B and 3.1% for both Jev and the E2B; escalating the least confident 10% lowers it to 0.8%, 1.1% and 0.8%. Jev's error rate stays at 0.7% from 20% to 30% escalated: in that range its score is escalating classifications that are no more likely to be wrong than the ones it keeps. Under 1 − p₂, the error rate of the classifications the 26B and the E2B keep continues to fall, to 0.3% at 30% escalated.

We rank the Gemma models' classifications by 1 − p₂ rather than p₁ mainly because most errors are confusions between two classes: when the 26B was wrong, the correct class was ranked second 93% of the time, so the runner-up's probability carries the signal. On this task, however, the two scores rank classifications almost identically.9

Choosing a model

Smaller off-the-shelf models were clearly less accurate with our class descriptions, which we wrote while working with the 26B. A small model already trained for the domain, such as our E2B, closes that accuracy gap and serves more than three times the 26B's requests per second on one GPU; without one, a mid-size open model such as the 26B matches Jev's accuracy with no training at all.

Takeaways

  • Any open LLM or SLM can act as a classifier on the fly: describe the classes, have the model generate one token and read the class probabilities, with no fine-tuning and no output parsing. Serving the model still needs a GPU or an API.
  • An off-the-shelf open model can match a purpose-built classification API on accuracy and, at light load, answer with lower latency.
  • The same forward pass also yields a confidence score. On our task, escalating by 1 − p₂ kept lowering the Gemma models' error rate on the classifications they keep, past the point where the closed API's own score stopped lowering Jev's.
  • Throughput under heavy load is determined by the hardware allocated, not by the method.
  • The choice of model matters more than the recipe, so measure candidate models on your own classes before relying on one.

  1. A capital letter is one token in common tokenizers, so each class maps to exactly one token and the whole answer fits in one generation step. An off-the-shelf model knows the classes only through their descriptions, so their clarity largely determines accuracy, and descriptions that overlap split the probability between letters, which lowers either confidence score. ↩
  2. OpenAI-compatible APIs return these probabilities as log-probabilities of the most likely candidate tokens, alongside the generated token. ↩
  3. Choose the escalated share by weighing the cost of an escalation against the cost of a wrong classification; set the threshold on one labeled sample and check the error rate on another. A confidence score serves this purpose as long as it ranks wrong classifications below correct ones. ↩
  4. We chose the 26B for efficiency and price. As a mixture-of-experts model it activates only about 4 billion of its 26 billion parameters per token, and its input tokens cost the same as Jev's, 4.2 cents per million. The output price is immaterial, since every classification generates a single token. ↩
  5. The E2B was trained with reinforcement learning as a language model, not as a classifier. For each training prompt it wrote the task group followed by a few sentences of reasoning; a larger model, Gemma 4 31B, saw the same prompt and the correct task group and accepted an answer only if both the task group and the reasoning were right, and a rejected answer received one revision against the judge's critique. The accepted answers became the fine-tuning data for the next round. Nothing in this training targets the class probabilities or the confidence score, and the model has no classification head, so reading it as a one-token classifier is the same on-the-fly step we apply to the 26B. None of the held-out prompts were used in its training. ↩
  6. Each Gemma model was called from the machine that hosts its GPU, so its latencies exclude the network round trip that Jev's latencies include. We ran each load level, from 1 to 128 parallel requests, three times: latency is the median over all requests from the three runs and throughput is the mean of the three runs' rates. ↩
  7. The probability mass reported outside the letters is an upper bound, because APIs return only the top-k next tokens and any letter missing from that list is counted as outside. ↩
  8. In a deployment the escalated classifications are answered by a larger model or a reviewer, and the accuracy of the whole system then also depends on how often that answer is right. We did not measure that step, so none of the error rates or accuracies in this post include answers from a larger model or a reviewer. ↩
  9. With 30% escalated, the error rate of the classifications kept is 0.27% for the 26B under 1 − p₂ and 0.31% under p₁; for the E2B it is 0.28% under 1 − p₂ and 0.27% under p₁. A second, practical reason favors the runner-up: at the confident end p₁ takes values such as 0.9999997, and for about one prompt in ten with the 26B the log-probability returned for the top token rounds to exactly zero, so p₁ read directly cannot order those prompts. p₂, by contrast, is returned as its own log-probability, retains precision across orders of magnitude and requires only the top two candidate tokens. ↩

Ready to Reduce Cloud Security Noise and Act Faster?

Discover the power of Averlon’s AI-driven insights. Identify and prioritize real threats faster and drive a swift, targeted response to regain control of your cloud. Shrink the time to resolution for critical risk by up to 90%.

CTA image