Jeff – 兼容 Jev 的 0.8B 决策模型,家庭训练,约 30 毫秒
Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

原始链接: https://github.com/firelex/jeff

“Jeff” 是一系列经过微调的超轻量级语言模型(Qwen3.5 0.8B/2B 和 Gemma 2 2B),专为高速零样本分类而设计。通过在单次前向传播中输出用户定义选项的校准概率,Jeff 无需进行文本生成或解析。 **主要特性:** * **性能:** 速度极快,在高端硬件上的决策时间约为 22–28 毫秒。在分类任务中,其表现可与 Jev 等较大模型媲美,但不具备深度推理能力。 * **易用性:** 模型作为本地代码的“决策引擎”运行,可处理用户意图、内容审核及游戏指令。它们沿用了 Jev 的请求格式。 * **定制化:** 虽然零样本表现强劲,但用户可以通过简短的领域特定微调实现高准确率(例如从 31.7% 提升至 95.8%)。 * **本地优先:** 完全基于本地硬件构建,采用开源合成数据管道,不依赖闭源模型。 * **部署:** 支持 NVIDIA GPU(PyTorch)和 Apple Silicon(MLX)。 Jeff 是一款专注于分类和落地的专用工具,并非规划或推理引擎。其权重和代码均已开源(Apache 2.0/MIT),为基于 API 的分类服务提供了一种高性能的私有化替代方案。

开发者 "firelex" 发布了 **Jeff**,这是一组针对高速零样本(zero-shot)分类进行优化的开源小型语言模型(0.8B 和 2B 参数)。与标准模型不同,Jeff 的设计旨在跳过文本生成过程,通过单次前向传递直接输出特定选项的校准概率。 在本地硬件上运行,0.8B 模型每次决策的推理速度约为 30 毫秒。虽然这些模型在复杂推理任务(如 BBH 基准测试)中表现不如大型模型,但它们非常适合作为实时应用(如语音导航和游戏代理)的“系统 1”分类器。 **主要内容:** * **效率:** 模型旨在直接嵌入代码中,以实现快速决策。 * **微调:** 虽然零样本性能已相当可观,但进行短时间的微调(30 分钟内)可以大幅提高特定任务的准确性。 * **细微差别:** 在这些特定角色中,小型模型往往表现优于大型模型,因为它们不易出现大型预训练模型中常见的“犹豫”或风险规避倾向。 * **提示词:** 开发者强调精确的措辞至关重要,并指出对选项描述的细微更改可能会显著影响模型性能。
相关文章

原文

Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification: small, fast decision models you slot into your code, with the same request format as Jev. You describe a situation and list the options in plain words; Jeff returns a calibrated probability for each option from a single forward pass. No generated text, no parsing: about 22 ms per decision on an RTX PRO 6000 and 28 ms on an Apple M4 Max (MLX).

Zero-shot means the options can be anything: support queues, user intents, moderation labels, voice commands, game moves. Your categories don't need to appear in the training data; you describe them, and Jeff picks.

What it is, and what it isn't. These are very small models. They make extremely fast, well-calibrated judgement calls between options, and they slot easily into your local code. On benchmarks they approach, and sometimes beat, Jev; but at this size their reasoning won't match Jev's, which runs on a much larger model. If zero-shot accuracy isn't good enough for your purposes, a short fine-tune on your own examples takes you much further: our voice-navigation fine-tune moved held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU.

Built entirely on local hardware. Training on one RTX PRO 6000 workstation GPU (the 0.8B trains in about 2 hours, the 2B in about 3.5), all synthetic training data written by an open model (Qwen3.8-Flash-Next) on two DGX Sparks, testing on a MacBook. No cloud GPUs, and no closed-model output in the training data; a closed model was used only to spot-check the quality of a sample of the synthetic data.

Independent project. Jeff uses the same request format as Jev, but it is not affiliated with or endorsed by TypeSafe, the makers of Jev. Our training code starts from the open-source AutoJev recipe.

Models on Hugging Face: Jeff-Qwen3.5-0.8B · Jeff-Qwen3.5-2B · Jeff-Gemma4-E2B

uv sync
uv run hf download mstrasser/Jeff-Qwen3.5-0.8B --local-dir checkpoints/jeff-0.8b

# NVIDIA GPU or CPU (PyTorch)
JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve
# Apple silicon (MLX, much faster on a Mac; Qwen models only)
uv sync --extra mac
JEFF_BACKEND=mlx JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
  "model": "jeff-latest",
  "state": "Refund request: the customer says the parcel arrived crushed and wants their money back.",
  "questions": {
    "route": {"type": "choice", "instructions": "Which team should handle this?",
              "criteria": {"1": "Refunds and payments", "2": "Damaged or lost parcels", "3": "Account and login problems"}},
    "angry": {"type": "noul", "instructions": "Is the customer angry?"}
  }
}'

Each answer has a probability per option, the chosen option and a confidence. Three question types: choice (pick one of up to 255 options), noul (yes/no, returned as a probability) and score (a point on a scale you describe). Several independent questions in one request are answered together.

4,599 questions from five public benchmarks, plus JevBench's public hard tier (105 items, scored separately):

Accuracy of Jeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B and Jeff-Gemma4-E2B against Jev's published figures, per benchmark

Benchmark Qwen3.5-0.8B untrained Jeff-Qwen3.5-0.8B Qwen3.5-2B untrained Jeff-Qwen3.5-2B Gemma 4 E2B untrained Jeff-Gemma4-E2B Jev (published) AutoJev-27B (published)
Overall (5 benchmarks) 45.3 79.1 46.5 83.1 62.5 81.6 83.0 84.9
BBH 39.5 64.0 46.0 68.0 51.3 66.4 94.3 82.8
Financial PhraseBank 36.0 96.4 53.4 96.3 86.0 96.1 77.0 84.2
JudgeBench 56.6 62.6 57.4 64.6 46.9 60.6 78.6 78.9
RAGTruth 49.1 86.1 35.9 88.9 63.8 87.4 77.3 88.9
WinoGrande 49.2 68.6 52.2 79.0 51.0 77.4 90.7 83.3
JevBench hard (separate) 36.2 47.6 45.7 53.3 41.0 48.6 73.3 70.3

Bold: the winner of Jeff against Jev in each row. Bold italic: AutoJev-27B where it is the best of all models in the row (on RAGTruth, tied with Jeff-Qwen3.5-2B); it is shown for reference, since the head-to-head comparison is with Jev. The published Jev and AutoJev figures were measured on a different sample of the same benchmarks. Jeff's overall score comes from classification and grounding, where it matches or beats the large models; on the reasoning-heavy benchmarks (BBH, JudgeBench, JevBench) it stays well below them, as you would expect at this size.

To test zero-shot performance on tasks unlike anything in the benchmarks, we had Jeff play three games. Games aren't the ideal zero-shot test, since a game's state isn't typical unstructured data; but they are a common, and fun, way to test a System 1 model. Each turn, the code describes the situation and the legal moves in words, and the model picks one. The options state what each move leads to (Frogger: "you would be hit by a car and lose a life"; Doom: "the nearest monster is a little to your left"), but never which move is right. Each result is 20 episodes, seed 1234; ▶ opens a video of the run's first episode.

Jeff-Qwen3.5-0.8B playing, zero-shot (the bold row in the table below; click a clip for the full video):

Model Doom, kills (monster's direction in words) Frogger, crossings (consequences) Pac-Man, pellets of 98 (consequences)
Random moves −0.05 0 11.2
Hand-coded rule bot 6.55 ▶ 10.25 ▶ 94.1 ▶
Qwen3.5-0.8B, untrained 5.0 ▶ 1.0 ▶ 25.8 ▶
Jeff-Qwen3.5-0.8B 6.55 ▶ 10.3 ▶ 57.0 ▶
Qwen3.5-2B, untrained 0.55 ▶ 0.05 ▶ 72.1 ▶
Jeff-Qwen3.5-2B −0.9 ▶ 6.0 ▶ 41.2 ▶
Gemma 4 E2B, untrained −0.55 ▶ 0 ▶ 3.2 ▶
Jeff-Gemma4-E2B 0.55 ▶ 0.15 ▶ 53.2 ▶
Jev (published, Doom) 6.55, told the aiming rule; −0.60 without it — —

Jeff-0.8B decides in 29–49 ms per move on an M4 Max; Jev's published Doom run took 212 ms per call over its API. The two times were not measured on the same hardware. To play them yourself:

uv sync --extra games
uv run python -m jeff.games --game doom --player jeff --criteria situation --url http://127.0.0.1:8765 --video --out runs/games/doom.json
uv run python -m jeff.games --game frogger --player jeff --criteria outcomes --url http://127.0.0.1:8765 --out runs/games/frogger.json
uv run python -m jeff.games --game pacman --player rule --out runs/games/pacman-rule.json

Median time per decision over the same 200 benchmark questions (about 200 input tokens each), one question at a time, from raw text to probabilities:

Model Parameters Weights (16-bit) NVIDIA RTX PRO 6000 Apple M4 Max (MLX) CPU (32 threads)
Jeff-Qwen3.5-0.8B 0.8B 1.7 GB 22 ms 28 ms 463 ms
Jeff-Qwen3.5-2B 2B 4.2 GB 24 ms 60 ms 708 ms
Jeff-Gemma4-E2B 2B effective (4.6B stored) 9.3 GB 29 ms — (MLX runs Qwen only) 1.0 s
AutoJev-27B 27B ~54 GB not published — —
Jev not disclosed API only 114–212 ms per call in published Doom runs, including the network
  • Reason in code, decide with Jeff. It's a classifier, not a planner. State what each option leads to ("this move gets you hit by a car"); asked to forecast ("a car arrives in 2 turns"), it does no better than random.
  • Wording matters enormously. Describe options consistently: giving Frogger's goal option the same words as every other forward option took one episode from 15 crossings to 23.
  • Use short option keys and descriptive text: {"1": "Engagement letter"}, not long IDs, which cost time and add nothing.
  • Ask independent questions together in one request.
  • Fine-tune it if zero-shot isn't enough. A voice-navigation fine-tune on ~11k app-specific examples took about half an hour on one GPU and moved held-out accuracy from 31.7% to 95.8%, at about 40 ms per decision on an M4 Max: autojev-train --initial-checkpoint <jeff> --epochs 1 ....
  • Pick the size for the job. For fast option picking the 0.8B is the sweet spot: the 2B is more cautious and plays the games worse, despite scoring higher on the benchmarks.
uv run autojev-mix ...          # build the training set (public data, synthetic data, leak filter)
scripts/train.sh RUN data/mix/public.jsonl data/mix 5e-6 40 Qwen/Qwen3.5-0.8B <revision> --epochs 1
uv run autojev-evaluate --data data/panel.jsonl --local --checkpoint checkpoints/RUN/selected --output runs/eval/RUN.json

The full pipeline (synthetic data from a local teacher, leak filter, learning-rate sweeps, dashboard) is described in scripts/train_all.sh, and every training source with its licence in docs/data-sources.md. Training recipe: full-weight fine-tuning, one epoch, batches of 256, cross-entropy over the option letters, then one fitted temperature for calibration; checkpoints are chosen on a development set, never on the benchmark panel. At least half of each training family follows the panel's layout conventions (formats only; no panel item is ever trained on).

  • Small models don't reason. Expect fast, calibrated choices between the options you describe, not multi-step reasoning. At 0.8B–2B parameters this holds for every model, not just Jeff.
  • Jeff-2B is a weaker game player than Jeff-0.8B. The untrained 2B already appears more risk-averse than the untrained 0.8B, and our training seems to have made that worse. This needs more investigation.
  • Benchmark scores don't predict game play. The untrained Gemma 4 E2B beats the untrained Qwen models on the benchmarks yet plays the games worst: right most of the time, but not reliably. Training fixed its Pac-Man (3.2 → 53.2 pellets) but not its Doom or Frogger.
  • Prompts matter. Jev's own Doom prompt (a raw bearing number plus an aiming rule) does not work for any of our models; options that state consequences in words do.
  • English and text only.

Jeff began as a fork of AutoJev by Denis Yarats (MIT licence), an open recipe that fine-tunes Qwen3.8-27B to return Jev-style decisions. We kept its core design (one forward pass per decision, a trained answer readout, a fitted temperature for calibration) and built on it: small students (0.8B and 2B Qwen, Gemma 4 E2B), a local synthetic-data pipeline with a leak filter, prompt layouts for domain fine-tunes, MLX serving on Apple silicon, game tests and a training dashboard. The original copyright notice is kept in LICENSE.

Code: MIT (including AutoJev's). Model weights: Apache 2.0. Doom harness adapted from jev-plays-doom (MIT). Training data: see the dataset card; each source keeps its licence and is listed in docs/data-sources.md. We release the weights and code, not the training data; some sources are share-alike (CC BY-SA).

联系我们 contact @ memedata.com