· 14 min read

CLM decides fast when the options stay the same

CLM reads the situation and each option separately and picks the closest one. It is fast when the options repeat, and weak when the sentence needs nuance.

Keys hanging on wire loops in two rows on a display rack with handwritten labels.
YoNeKeN

I sent this ticket, with two options. “My invoice was charged twice and nobody answers the phone!” CLM returned billing at 0.989 and technical at 0.011. It did not write a reply. It read the ticket, read each option, and kept the closer one.

The README opens on another test, the Chrome dinosaur. CLM and Jev both got through five 60-second tries without dying. CLM answered in a median of 16.5 ms. Jev took 149.8 ms.

The results file for the same test shows the rest. CLM picked the same move as the physics planner in 65.8% of decisions. Jev did in 98.7%. Across CLM’s five tries, the test’s shield stepped in 4,883 times. Across Jev’s, 28. That shield replaces an answer the planner marked unsafe with the most likely safe option.

The two numbers measure different things. Survival includes the shield. Agreement with the planner is the model on its own.

CLM reads the situation and each option separately, keeps what it has already read, and picks the closest option. It is fast when the same options come back. When the sentence needs nuance, the published numbers and the only independent test I found leave CLM behind Jev.

CLM picks the option closest to the situation

Researchers from Stanford and Nvidia, according to VentureBeat, released CLM on September 23, 2026, with code and weights under Apache 2.0. The name is Contrastive Language Models. The first model is CLM-8B.

The input is the situation, the state field, and a closed list of options. Each option has a name and a description. In the ticket above, the situation is the customer’s sentence. The options are billing and technical.

A frozen Qwen3-8B reads each text and returns a list of numbers, the last token’s embedding. Frozen means those 8 billion parameters are not part of training. The README calls the rest two encoders, one for the situation and one for the option. At serving time, both use that same Qwen. What changes is a small projection module on each side. Each module compresses the list to 512 numbers.

An option’s score is the cosine between the two lists, how much they point the same way, multiplied by a scale training learns. A softmax turns those scores into probabilities that add up to 1. That is why 0.989 and 0.011 add up to 1. The highest one wins.

The published file with both modules is 75 MB. Qwen is not in it and runs separately, on a vLLM server that returns one vector per text instead of generating the next word.

“Reset password” produces the same list on every ticket. If an agent always picks among the same 50 actions, Qwen reads each one once and the server keeps the result. On the next decision, only the new situation goes through the large model.

Training uses InfoNCE in both directions. In a batch of pairs, the situation has to land closer to its own action than to the other pairs’ actions. The action also has to land closer to its own situation than to the other situations.

The README describes three stages.

  • Pre-training on about 60 million Nemotron question and answer pairs.
  • Mid-training with about 30 million hard negatives, answers that look like the right one and are still wrong, generated by Gemini 2.5 Flash-Lite.
  • Post-training on about 1 million agent trajectories. Each step becomes a pair of the context and the action taken.

The order shows up in the team’s own numbers. On about 100 thousand questions, with one right answer and ten similar wrong ones, pre-training alone gets 52.1% top-1. A short mid-training run lifts it to 69.2%. Training with those similar answers from the start peaks at 62.4% and then drops. The README calls that drop overfitting.

The API follows TypeSafe’s format. A request written for Jev, with a state and choice, score, or noul questions, runs on clm-serve if you change the URL and the key. choice picks one option. score returns a number. noul returns a yes or no probability. Underneath, each question becomes the situation plus a list of texts. The rank endpoint does the same with free-form candidates, such as N solutions proposed by an agent.

The difference from Jev and Laya is what you can cache

Jev, Laya, and CLM answer those three questions. The path to the score differs.

Laya puts the situation, the question, and the options into the same input text, with one [MASK] marker per option, as I described in the post about it. Each option is read together with the situation. TypeSafe has not published Jev’s architecture.

In CLM, situation and option only meet at the cosine. Jacky Kwok, the paper’s first author, told VentureBeat that Jev and Laya mainly cache the situation’s representation, while CLM computes and caches situation and option separately. The claim about Jev is his. The CLM side is in the code.

The README measures that cache on an RTX 4090 with a fixed action set. With a new situation on every call, the server median was 28.6 ms without the vector cache and 28.0 ms with it. With situations it had already seen, it dropped from 1.7 ms to 0.6 ms. The cache helps when the agent comes back to a situation Qwen has already read. A new situation still waits on the whole encoder.

Jev (TypeSafe)Laya (ConvAI)CLM-8B
WeightsClosed, hosted APIApache 2.0Apache 2.0
ModelArchitecture not publishedModernBERT-large, 421MFrozen Qwen3-8B + two projection modules
Situation and optionsNot publishedIn the same input textRead separately
What you downloadNothingAbout 800 MB (English checkpoint)75 MB of modules + 15.26 GiB of Qwen3-8B
Marginal costUS$0.042 per million input tokensYour own computeYour own compute, with a GPU for an 8B model
Customer fine-tuningNo public offeringOfficial notebookTrains only the projection modules, with Qwen frozen

The README benchmarks compare CLM and Jev

I found no published numbers for CLM against Laya. The README compares against Jev in two blocks.

The first had no extra training on the task. It uses the published projection modules.

TaskCLM latencyJev latencyCLM successJev success
T-Rex16.5 ms149.8 ms5/55/5
Tool calling (BFCL v4)76.8 ms125.5 ms95.2%99.2%
WikiRacing79.8 ms225 ms26/3030/30
Super Mario33.5 ms132.6 ms5/55/5

The “up to 9x faster” figure is T-Rex. 149.8 divided by 16.5. On the other three tasks, the ratio sits between 1.6x and 4x. On tool calling and WikiRacing, Jev got more right.

The latency mixes two measurements. CLM ran on a local RTX 4090. Jev answered through TypeSafe’s API. Both latencies were measured on the client, and the T-Rex example writes that down. One includes the network. The other has almost no network at all.

T-Rex and Super Mario are five tries each. WikiRacing is 26 of 30 for CLM and 30 of 30 for Jev. In T-Rex, the planner already writes “Safe” or “Unsafe” into each option’s text. The model picks among labels the test computed, and the shield can still replace the answer.

The second block uses CLM as a verifier. A larger model generates several solutions per task, Opus 5 on DeepSWE and Fable 5 on Terminal-Bench 2.1. The verifier picks which one to submit. The “random pick” row is pass@1, the rate if the choice among the candidates were random.

DeepSWETerminal-Bench 2.1
Tasks kept out of training3830
Candidates per task45
Random pick73.7%84.0%
CLM with tuned modules81.6%87.6%
Jev71.1%83.1%
CLM / Jev latency, on an H10079 / 449 ms32 / 131 ms

On DeepSWE, the gap over a random pick is three tasks. 31 of 38 against 28 of 38. The checkpoint card publishes the ceiling. A perfect verifier would get 34 of 38. Jev landed at 27 of 38, one task below a random pick.

On Terminal-Bench, CLM scores 87.6% and Jev 83.1%, against 84.0% for a random pick. 87.6% of 30 tasks is not a whole number. The README does not say how many runs the average covers.

The split I would weigh most is the training. The CLM in those rows had its modules tuned on 59 DeepSWE tasks, kept separate from the 38 test tasks. Jev went in without that tuning, because it has no public fine-tuning. It is the same split that already showed up with Laya. A specialist trained on the benchmark against a generalist that never saw it. The result shows that tuning the modules works. It does not show that CLM judges better than Jev.

The independent test I found counts against nuance

An author on Zenn compared CLM-8B, Jev, Kev-4B, and Qwen3.5-4B on an RTX 3090. The author built 13 cases where the sentence looks like one thing and the answer is another. The boy who cries “wolf!” is lying, and he is not disguised. Someone wearing a mask for hay fever is not hiding their identity. CLM got 4 of 13 right on one prompt and 5 of 13 on the other. Jev got 12 and 11. The author blames CLM’s errors on words like “secret”, “lie”, and “mask”. They pull the list of numbers toward the wrong option even when the sentence says the opposite.

Another test had three options, almost certain, uncertain, and impossible. There were two scenarios with ambiguous evidence, an internal investigation and a cancellation forecast. The middle option got 1.7% and 1.1% from CLM. Jev gave it 75% and 67%. If the workflow sends the uncertain case to a person, CLM barely does that on these two scenarios.

CLM did well at rejecting made-up entities and scientific results, with 94.8% to 96.0% probability on “no”.

That is 13 cases and two scenarios, written by one author. It does not support a ranking. The score is similarity between two lists, and sentences with the same words tend to land close together.

Running on Modal takes a 24 GB GPU

There is no hosted CLM API. clm-serve is the API, and it runs wherever you put it. Hugging Face hosts the weights and offers no inference for this checkpoint. The shortest path I found was Modal. It bills the GPU per second and shuts the container down when requests stop.

Qwen3-8B in bf16 took 14.11 GiB on Modal’s L4, according to the vLLM log. That left 5.47 GiB for context, with a 2,048-token limit. A 24 GB L4 is enough. The projection modules run on CPU.

The file starts two processes in the same container. vLLM serves Qwen on the GPU and answers embeddings on localhost. clm-serve computes the scores on CPU and is the only exposed port. The version below leaves out the code that stops the processes and part of the cache configuration.

import subprocess
import time
import urllib.request

import modal

app = modal.App("clm")

image = (
    modal.Image.debian_slim(python_version="3.12")
    .uv_pip_install("vllm==0.21.0", "contrastive-lm==0.1.0")
    .env({
        "HF_HOME": "/cache/huggingface",
        "VLLM_CACHE_ROOT": "/cache/vllm",
        "CLM_CKPT_DIR": "/cache/clm",
    })
)
cache = modal.Volume.from_name("clm-cache", create_if_missing=True)


def wait_for(url: str, timeout_s: int) -> None:
    deadline = time.time() + timeout_s
    while time.time() < deadline:
        try:
            with urllib.request.urlopen(url, timeout=5) as r:
                if r.status == 200:
                    return
        except OSError:
            pass
        time.sleep(5)
    raise TimeoutError(url)


@app.server(
    image=image,
    gpu="L4",
    cpu=4,
    memory=32768,
    volumes={"/cache": cache},
    port=8700,
    startup_timeout=20 * 60,
    scaledown_window=5 * 60,
    max_containers=1,
)
class Server:
    @modal.enter()
    def start(self):
        self.vllm = subprocess.Popen([
            "vllm", "serve", "Qwen/Qwen3-8B",
            "--served-model-name", "qwen3-8b",
            "--runner", "pooling",
            "--pooler-config", '{"task":"embed","pooling_type":"LAST"}',
            "--dtype", "bfloat16",
            "--max-model-len", "2048",
            "--gpu-memory-utilization", "0.90",
            "--enforce-eager",
            "--enable-prefix-caching",
            "--host", "127.0.0.1", "--port", "8090",
        ])
        wait_for("http://127.0.0.1:8090/health", 18 * 60)
        self.clm = subprocess.Popen([
            "clm-serve", "--port", "8700",
            "--emb-url", "http://127.0.0.1:8090/v1/embeddings",
            "--device", "cpu",
            "--action-cache", "0",
            "--no-ui",
        ])
        wait_for("http://127.0.0.1:8700/health", 3 * 60)
        cache.commit()

Three choices in that file change cost and latency. --enforce-eager turns off CUDA graphs and torch.compile. The container boots faster, and each pass through the model gets slower. --device cpu and --action-cache 0 leave the whole GPU to vLLM, so CLM’s vector cache is off. clm-serve still keeps the embeddings it has seen in memory, and forgets the oldest ones when that memory fills up. max_containers=1 caps the bill at one GPU.

To deploy and create the access credential:

modal token new
modal deploy app.py
modal workspace proxy-tokens create

The endpoint sits behind Modal’s proxy auth. Requests carry the Modal-Key and Modal-Secret headers with the token from the last command:

curl -s "https://<workspace>--clm-server.us-east.modal.direct/v1/systemone" \
  -H "Modal-Key: $MODAL_KEY" \
  -H "Modal-Secret: $MODAL_SECRET" \
  -H "Content-Type: application/json" \
  -d '{
    "state": "Customer: my invoice was charged twice and nobody answers the phone!",
    "questions": {
      "department": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
          "billing": "Charges, invoices, refunds",
          "technical": "Bugs and outages"
        }
      }
    }
  }'

For this example, with the two options from the top, the answer came back with billing at 0.989 and technical at 0.011.

From Brazil, the network weighs more than the model

The numbers below come from requests sent from Brazil to a Modal L4 in us-east, using that file.

The first boot took about 4 and a half minutes until the first answer. 150 seconds of that was downloading the Qwen3-8B weights. With the weights already in the volume, the second boot took about 76 seconds between the first request, which came back 503, and the server being ready. vLLM loaded the weights in 12.7 seconds.

With the container warm, I sent two questions per request, a three-option choice and a noul.

CaseRequestsServer p50Client p50Client p95
Situation and options already seen507.3 ms518.5 ms676.9 ms
Never-seen situation30146.4 ms674.8 ms922.4 ms

When the text has already been through Qwen, the server answers in milliseconds. The network and Modal’s proxy add about half a second on top, and this data does not separate one from the other. In the comparison post, hosted Jev had a p50 of 0.37 s, with requests also leaving from Brazil. The days, loads, and requests are different. At this distance, CLM on Modal was slower end to end than Jev.

With a new situation, the server took 146 ms, well above the 28 ms the README measures on an RTX 4090. The L4 is a smaller GPU, and --enforce-eager costs latency. I did not measure how much comes from each.

I also wrote ten short tickets with obvious answers. Four billing, three technical, and three sales. CLM got 8 of 10. Both errors were technical tickets sent to sales. One about a 500 error on the dashboard, the other about the API rejecting a token. The noul question “Is this urgent?” landed between 0.60 and 0.86 on all ten, including 0.70 for someone asking about a nonprofit discount. On those ten, it did not separate urgent from routine. Ten sentences I wrote myself are a sanity check, not a benchmark.

The double-charge ticket had a confidence of 0.979 with two options. I added sales as a third, and the number fell to 0.326. Billing still won. The probabilities are relative to the options you sent, as the model card warns. A cutoff tuned on two options does not hold for three.

The L4 costs US$0.000222 per second on Modal, about US$0.80 per hour. CPU and memory are billed separately. With 4 cores and 32 GiB, that adds about US$0.45 per hour. The container shuts down after 5 minutes without requests, and each boot waits through startup again.

Where I would put CLM

I would use CLM when an agent picks many times among actions that do not change, the situation repeats or arrives in batches, and the GPU sits close to the caller. Tool routing over a fixed catalog and ranking N candidate solutions are the cases the team itself shows.

Tuning the modules is worth it when there is a set of trajectories with known outcomes. Qwen stays frozen. The README’s verifier results come from that tuning.

On its own, I would not leave CLM on a question with an “uncertain” option that routes the case to a person, on a judgment that depends on negation or on who revealed what, or on a noul that has not yet shown it separates the cases in the workflow. On those points, Jev got more right in the numbers I found.

As with Jev and Laya, the answer is still one of the options you sent. If every one of them is wrong, it picks a wrong one.

Are the actions your agent picks the same on every call, or do they change along with the situation?

Sources