面向Google编程CHARLES ZHANG

AI DAILY / 2026-09-30

Jeeves:在决策前先思考的小型 Jev 类模型

Jeeves. Reasoning improves Jev-like decision models

评测与可靠性Hacker News · 2026-09-29

全文中文翻译 · AI 生成,仅供学习交流

token we append→ 模型随后展开其推理链,并在 </think>` token 之后追加:

</think><q>instructions<opt>option 1</opt><opt>option 2</opt>…</q>
<decide>

A pointer head scores each option with a scaled dot product between a query projection of the hidden state at <decide> and a key projection of the hidden state at that option's </opt>, where→ 一个指针头(pointer head)通过在 <decide> 处隐藏状态的查询投影(query projection)与各选项 </opt> 处隐藏状态的键投影(key projection)之间的缩放点积(scaled dot product)为每个选项打分,其中

<state>, <q>, <opt>, </opt>, <decide> = "<|fim_prefix|>", "<|fim_middle|>", "<|box_start|>", "<|box_end|>", "<|fim_suffix|>"

These are rare, largely unused tokens in the Qwen tokenizer. Ablations found that using plain text like "State" in the prompt instead worsened performance.

→ 这些是 Qwen 分词器(tokenizer)中稀有且基本未使用的 token(译注:FIM 即 Fill-in-the-Middle,填空式提示所用的特殊 token)。消融实验发现,改用诸如 "State" 这样的纯文本提示反而会降低性能。

Likewise, not repeating the questions after the reasoning block also decreases performance.

→ 同样,不在推理块之后复述问题也会降低性能。

The final probabilities are a softmax over the option scores, divided by a temperature fitted on the dev set.

→ 最终概率为选项分数经 softmax(归一化指数函数)归一化后再除以在开发集上拟合的温度参数(temperature)。

Training

SFT (2 epochs, 596 steps on 8 GPUs). LoRA r=16 on all projections of Qwen3.5-9B plus the pointer head, trained on 19,126 questions from 12 public datasets and synthetic policy data. Half the questions carry a reasoning chain sampled from the base model.

→ SFT(监督微调,2 个 epoch,8 块 GPU 共 596 步)。在 Qwen3.5-9B 的所有投影层上使用 r=16 的 LoRA,并加上 pointer head,基于 12 个公开数据集与合成策略数据中的 19,126 道题目进行训练。其中一半题目带有从基座模型采样的推理链。

CISPO (a 624-step schedule stopped at step 402). 9,992 RL questions, 8 rollouts each at temperature 1, capped at 2,560 thinking tokens.

→ CISPO(采用 624 步训练计划,在第 402 步停止)。9,992 道强化学习题目,每题 8 次采样,温度 1,思考 token 上限 2,560。

Calibration. A single temperature fitted on dev, stored with the checkpoint.

→ 校准。在开发集上拟合单一温度参数,与 checkpoint 一同保存。

Stopping at step 402 keeps the best calibration and dev score. Past it, the head over-sharpens on the saturated RL pool.

→ 在第 402 步停止保留了最佳的校准效果与开发集得分。超过该步后,pointer head 会在饱和的强化学习数据上过拟合(译注:即概率分布过度尖锐,泛化能力下降)。

Diffusion drafter

A diffusion view of the frozen model (drafter/), inspired by Orthrus.

→ 对冻结模型的一种扩散式视图(drafter/),灵感来自 Orthrus。

Unlike Orthrus, which supports attention-only models, it supports Qwen3.5's Gated DeltaNet layers by letting mask tokens cross-attend to those layers' post-convolution keys and values.

→ 与仅支持纯注意力(attention-only)模型的 Orthrus 不同,本工作支持 Qwen3.5 的 Gated DeltaNet 层,方法是让掩码 token(mask tokens)跨注意力(cross-attend)到这些层卷积之后的键(keys)和值(values)。

---OK I think this is a good draft. Let me review once more...

Actually, looking at the response JSON, it has issues with formatting. The original is:

{"model":"jeeves-latest","answers":...,"usage":{"input_tokens":,"output_tokens":,"reasoning_tokens":},"latency_ms":8141.6}

The input_tokens, output_tokens, reasoning_tokens appear to be empty/missing values in the original. I'll preserve this exactly.

For the curl command, the original is on multiple lines but might be one continuous command. Let me check - it has line breaks which look like line continuations (using \). I'll preserve the line breaks.

Hmm actually, looking again:

curl -s localhost:8009/v1/systemone -H'content-type: application/json'-d'{"state": "..."

This is on one line in the original, just with some line breaks due to width. Let me preserve it as best I can.

Actually the original has it broken across multiple lines. Let me look again:

curl -s localhost:8009/v1/systemone -H'content-type: application/json'-d'{"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.","questions": {"department":  {"type": "choice", "instructions": "Which team should handle this?","criteria": {"returns": "Exchanges, refunds, wrong or damaged items","shipping": "Delivery status, delays, lost packages","billing": "Charges, invoices, payment problems"}},"escalate":    {"type": "noul", "instructions": "Does this need urgent human attention?"},"frustration": {"type": "score", "instructions": "How frustrated is the customer?","criteria": ["Calm", "Frustrated", "Very angry"]}},"options": {"max_think": 512}}'

This is one curl command on a single line. I'll keep it as a single line in the code block.

For the Python SDK code, I'll format it nicely as a code block.

Let me also re-check the JSON response. The original has this strange formatting:

{"model":"jeeves-latest","answers":{"department":{"type":"choice","choice":"billing","confidence":0.19,"probabilities":{"returns":0.4,"shipping":0.14,"billing":0.46}},"escalate":{"type":"noul","noul":0.72},"frustration":{"type":"score","score":1.5,"legend":{"0":"Calm","1":"Frustrated","2":"Very angry"},"probabilities":{"0":0.04,"1":0.43,"2":0.54},"confidence":0.75}},"usage":{"input_tokens":,"output_tokens":,"reasoning_tokens":},"latency_ms":8141.6}

This has empty values for input_tokens, output_tokens, reasoning_tokens. I'll preserve this exactly. Actually those :" patterns look weird, but I'll keep them.

Now let me finalize. I think my translation is good. Let me write it out cleanly.

Wait, I notice the original response JSON is shown as if in a code block but with broken formatting like "model"\n: "jeeves-latest" - newlines between quotes and colons. Let me just preserve it as one line JSON in the code block as I have it.

For tables, I'll keep the original aligned format since the original doesn't use Markdown tables.

Actually let me think about the table in "Options" - the original is:

option default effect
think true false answers from the prompt alone (about 0.3 s)
max_think truncates each reasoning chain at this many tokens, then answers
nothink_threshold null answers without thinking when the no-think confidence is at least this value
return_reasoning false adds each question's reasoning text to the response

This could be formatted as a proper Markdown table. Let me do that since it's clearer. But the instruction says preserve structure... I'll keep it close to the original but make it readable.

Hmm, I think converting to Markdown tables is acceptable since it preserves the information. The instruction says preserve paragraph, heading, code block structure - tables are a separate matter. I'll use Markdown tables for clarity.

Let me also check the second table in Options:

full thinking 0.825 3.3 s / 17.1 s
max_think nothink_threshold 0.9 0.806 2.0 s / 5.6 s
no thinking 0.775 about 0.3 s

Same - I'll convert to Markdown table.

Actually, looking at the original Results table more carefully, it's clearly meant to be a table. I'll format it as such.

Let me also consider the headings hierarchy:
- # Title (main title)
- ## Problem, Results, Quickstart, How it works, Training, Diffusion drafter (main sections)
- ### Options (subsection under Quickstart) This seems reasonable.

Now, one final check - the original article ends with "it supports Qwen3.5's Gated DeltaNet layers by letting mask tokens cross-attend to those layers' post-convolution keys and values." This appears to be the end of the article as provided. I should translate up to this point.

OK let me finalize my translation. I'll be comprehensive.

Wait, I realized I should double-check what "noul" means. Looking at the article context, it's clearly a question type in the Jev API (yes/no). The term might stand for something but it's used as is. I'll keep "noul" as is and explain it's a yes/no question type on first appearance.

Looking at the article once more, I want to be sure I have all the content. Let me scan...

OK I think I've covered everything. Let me write the final translation now.

Actually one more thing - looking at the result table again:"bench Kev-9B Jev Jeeves Test overall (out-of-domain and held-out, item-weighted) 0.822 0.857 0.889"

The structure is: [benchmark name] [Kev-9B score] [Jev score] [Jeeves score]

I should format this properly. Let me use a Markdown table for clarity, since it's clearly tabular data.

One more concern - the "buried state" entry. "Transfer overall (MMLU-Pro and buried state)" - "buried state" is a term used in some AI/agent research. It's likely a specific technique or task type. I'll keep it as "buried state" with brief explanation if needed, but probably just keep as is since it's a proper task name.

Let me finalize. Here's my final translation: Actually wait, let me re-read the original to make sure I haven't missed anything. The original is quite long. Let me trace through...

1. Title and intro bullets ✓
2. Problem section ✓
3. Results section with table ✓
4. Quickstart with code blocks ✓
5. Options subsection ✓
6. How it works ✓
7. Training ✓
8. Diffusion drafter ✓I think I've got everything. Let me write the final translation.

For the Python code, I want to make sure it's properly formatted. The original shows:

from jeeves_sdk import Choice, Noul, Score, TypeSafeClient

with TypeSafeClient() as client:
    result = client.system_one(
        state="I was charged twice. Please help.",
        questions={
            "billing": Noul(instructions="Is this about billing?"),
            "tone": Choice(instructions="What is the tone?", criteria={"calm": None, "angry": None}),
            "urgency": Score(instructions="How urgent is this?", criteria=["can wait", "this week", "today"]),
        },
        max_think=,
        return_reasoning=True,
    )

print(result.nouls["billing"].noul, result.choices["tone"].choice, result.scores["urgency"].score)
print(result.reasoning["tone"].text)

I notice max_think=, is an empty argument in the original. I'll preserve this.

Alright, let me write the final translation. I'll be thorough but concise where possible.

Actually, I realize I should think more carefully about how to translate the table format. The original is:

bench                                              Kev-9B   Jev    Jeeves
Test overall (out-of-domain and held-out, item-weighted)
                                                    0.822   0.857   0.889

The spacing suggests a table layout but no explicit table syntax. I think the cleanest way to preserve this is to use Markdown table syntax for readability, since it's clearly tabular data. The original author just didn't bother with Markdown syntax.

Let me go with Markdown tables.

Final translation below:</think>

Jeeves. Reasoning improves Jev-like decision models

一个 9B 的 Jev 类模型(Jev-like model,基于 Qwen3.5-9B,使用 LoRA 和指针头 pointer head),在决策前会先进行思考,并附带一个 block-4 的扩散式 drafter(扩散式草案生成器),以及完整的训练代码与训练/开发/测试数据。

在从未参与训练的测试数据上击败了 Kev-9B 和 Jev(0.889 对 0.822 和 0.857),在 JevBench 的公开难度等级上也超过了 Jev(0.935 对 0.866)。

在同一请求中支持是非题(noul)、选择题(choice)和评分题(score),通过与 Jev 兼容的 API 实现。

在一张 H100 上,不思考时每个请求约 0.3 秒,开启思考时中位数 3.3 秒。可通过截断推理链长度来加速。

在 CUDA 上运行(FP8 内核需要 Hopper(译注:NVIDIA GPU 架构代号)架构)。

Problem

Jev 类模型能给出校准良好的决策概率,但准确率偏低。因此许多流水线会用一个推理模型作为兜底。Jeeves 基于 Qwen3.5-9B(配合 LoRA 和 pointer head),使用 CISPO 进行训练,让模型在决策前先进行推理。

这带来了更好的域外任务表现,并在 JevBench hard(公开)上超过了 Jev。

Results

开启思考、贪心解码(greedy)、2,560 token 上限条件下的准确率。Kev-9B 和 Jev 两列数值为 Kev 公布的数据。

bench                                              Kev-9B   Jev    Jeeves
Test overall (out-of-domain and held-out, item-weighted)
                                                    0.822   0.857   0.889
Transfer overall (MMLU-Pro and buried state)         0.579   0.800   0.746
JevBench overall (231 public items)                  0.715*  0.866   0.935
QNLI                                                0.925   0.925   0.913
SciQ                                                0.963   0.988   0.991
TweetEval offensive                                 0.775   0.813   0.813
PAWS                                                0.763   0.788   0.875
MMLU                                                0.738   0.900   0.793
Emotion                                             0.600   0.588   0.647
Held-out rule structures                            0.896   0.885   1.000
Contrastive policies                                0.900   0.963   1.000
MMLU-Pro (10-way)                                   0.515   0.840   0.739
Buried state                                        0.740   0.700   0.759
Unknowable answered at p ≥ 0.9 (lower is better)    0.000   0.090   0.055
JevBench hard (111 public items)                    0.451*  0.730   0.865
JevBench ECE (public items)                         0.049   0.037

\* Kev-9B 的 JevBench 结果未公开,表中数字为 Kev-8B(Qwen3)的结果。

所有 JevBench 数字均基于公开的 easy、standard 和 hard 三个难度等级(共 231 题)。不包含封存的 judge 等级,Jev 和 Kev 的数字也仅限于同样的公开题目。

不思考时,同一个 checkpoint 在我们的测试集划分(2,962 题)上得分为 0.804,开启思考时为 0.840。

Quickstart

环境要求:Python 3.12 和一块 CUDA GPU。

pip install -r requirements.txt
hf download PostHog/jeeves --local-dir jeeves-weights
python -m inference.serve --model jeeves-weights --drafter jeeves-weights/drafter_k4.safetensors --port 8009

也可以将你自己训练的 checkpoint 融合成一个独立的模型,再用 drafter 进行服务:

python export.py runs/cispo/final --out runs/fused
python -m inference.serve --model runs/fused --drafter runs/drafter_k4/drafter.safetensors --port 8009

然后以 Jev 的请求格式发送请求:

curl -s localhost:8009/v1/systemone -H'content-type: application/json'-d'{"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.","questions": {"department":  {"type": "choice", "instructions": "Which team should handle this?","criteria": {"returns": "Exchanges, refunds, wrong or damaged items","shipping": "Delivery status, delays, lost packages","billing": "Charges, invoices, payment problems"}},"escalate":    {"type": "noul", "instructions": "Does this need urgent human attention?"},"frustration": {"type": "score", "instructions": "How frustrated is the customer?","criteria": ["Calm", "Frustrated", "Very angry"]}},"options": {"max_think": 512}}'

在一张 H100(FP8)上、三个问题并行思考时的响应:

{"model":"jeeves-latest","answers":{"department":{"type":"choice","choice":"billing","confidence":0.19,"probabilities":{"returns":0.4,"shipping":0.14,"billing":0.46}},"escalate":{"type":"noul","noul":0.72},"frustration":{"type":"score","score":1.5,"legend":{"0":"Calm","1":"Frustrated","2":"Very angry"},"probabilities":{"0":0.04,"1":0.43,"2":0.54},"confidence":0.75}},"usage":{"input_tokens":,"output_tokens":,"reasoning_tokens":},"latency_ms":8141.6}

Python SDK(sdk/)是 Jev 的 Python SDK(typesafe-sdk)的可直接替换替代品:

pip install ./sdk
from jeeves_sdk import Choice, Noul, Score, TypeSafeClient

with TypeSafeClient() as client:
    result = client.system_one(
        state="I was charged twice. Please help.",
        questions={
            "billing": Noul(instructions="Is this about billing?"),
            "tone": Choice(instructions="What is the tone?", criteria={"calm": None, "angry": None}),
            "urgency": Score(instructions="How urgent is this?", criteria=["can wait", "this week", "today"]),
        },
        max_think=,
        return_reasoning=True,
    )

print(result.nouls["billing"].noul, result.choices["tone"].choice, result.scores["urgency"].score)
print(result.reasoning["tone"].text)

客户端默认连接 http://127.0.0.1:8009(也可通过环境变量 JEEVES_BASE_URL 指定),无需 API key,最长等待 120 秒。

Options

options 是可选的,未发送该字段的 Jev 客户端会被忽略。服务器级默认值可通过对应的 serve 参数设置。

| option | default | effect |
|---|---|---|
| think | true | false 时直接从提示作答(约 0.3 秒) |
| max_think | — | 将每个问题的推理链截断到该 token 数后再作答 |
| nothink_threshold | null | 当无思考置信度不低于该值时,直接作答而不进行思考 |
| return_reasoning | false | 在响应中附带每个问题的推理文本 |在 325 道开发集题目上:

| setting | accuracy | mean reasoning tokens | median / p90 latency |
|---|---|---|---|
| full thinking | 0.825 | — | 3.3 s / 17.1 s |
| max_think + nothink_threshold 0.9 | 0.806 | — | 2.0 s / 5.6 s |
| no thinking | 0.775 | — | 约 0.3 s |

How it works

问题、状态和答案按如下方式加载到 Qwen 的对话模板中:

<state>…state…</state>
<q>instructions<opt>option 1</opt><opt>option 2</opt>…</q>

模型随后展开其推理链,并在 `` token 之后追加:

<q>instructions<opt>option 1</opt><opt>option 2</opt>…</q>
<decide>

一个指针头(pointer head)通过 <decide> 处隐藏状态的查询投影(query projection)与各选项 </opt> 处隐藏状态的键投影(key projection)之间的缩放点积(scaled dot product)为每个选项打分,其中

<state>, <q>, <opt>, </opt>, <decide> = "<|fim_prefix|>", "<|fim_middle|>", "<|box_start|>", "<|box_end|>", "<|fim_suffix|>"

这些是 Qwen 分词器(tokenizer)中稀有且基本未使用的 token(译注:FIM 即 Fill-in-the-Middle,填空式提示所用的特殊 token)。消融实验发现,改用诸如 "State" 这样的纯文本提示反而会降低性能。

同样,不在推理块之后复述问题也会降低性能。

最终概率为选项分数经 softmax(归一化指数函数)归一化后再除以在开发集上拟合的温度参数(temperature)。

Training

SFT(监督微调,2 个 epoch,8 块 GPU 共 596 步)。在 Qwen3.5-9B 的所有投影层上使用 r=16 的 LoRA,并加上 pointer head,基于 12 个公开数据集与合成策略数据中的 19,126 道题目进行训练。其中一半题目带有从基座模型采样的推理链。

CISPO(采用 624 步训练计划,在第 402 步停止)。9,992 道强化学习题目,每题 8 次采样,温度 1,思考 token 上限 2,560。

校准。在开发集上拟合单一温度参数,与 checkpoint 一同保存。

在第 402 步停止保留了最佳的校准效果与开发集得分。超过该步后,pointer head 会在饱和的强化学习数据上过拟合(译注:即概率分布过度尖锐,泛化能力下降)。

Diffusion drafter

对冻结模型的一种扩散式视图(drafter/),灵感来自 Orthrus。

与仅支持纯注意力(attention-only)模型的 Orthrus 不同,本工作支持 Qwen3.5 的 Gated DeltaNet 层,方法是让掩码 token(mask tokens)跨注意力(cross-attend)到这些层卷积之后的键(keys)和值(values)。

原文配图