Skip to content

Behavior Evals & Model Policy ​

The compiler evals measure extraction quality against the golden corpus — a fixed set of 9 test sources with known-correct claim extractions. Every compiler model must score above thresholds on it (recall, exact locators, verbatim quotes) before it's trusted to write canonical evidence. Behavior evals measure the thing that actually matters: when an agent receives the context manifest, does it make a better decision?

bash
wf eval behavior                  # zero-LLM: manifest-level compliance (CI)
wf eval behavior --llm            # probe a real model + score its answer
wf eval behavior --record         # append metrics to registry/log.md

Fixtures in evaluations/behavior/ encode scenarios with two layers:

  1. Zero-LLM layer (deterministic): does the compiled manifest deliver the required knowledge? — banned anti-pattern present in the prompt, binding decision present, stale pattern excluded with a reason, gap acknowledged (nothing selected where no knowledge exists).
  2. LLM layer (--llm): the same fixture's probe question is sent to a real model with the manifest; the answer must contain the compliant behavior (must_contain) and must not contain the banned approach (must_not_contain).

Current fixtures:

FixtureBehavior tested
be1-avoid-anti-patternagent is told NOT to build the banned shared token cache; implements rotation + per-session
be2-project-over-globalproject constraint (rotating tokens) beats the general preference (session store)
be3-stale-demotedstale pattern excluded from the manifest, fresh pattern delivered
be4-escalate-gapno billing knowledge exists → agent escalates rather than inventing policy
be5-receipt-proves-deliverya persisted context receipt proves the manifest was actually delivered
be6-commitment-triggeredan open commitment whose trigger matches the task surfaces at P1 precedence
text
Knowledge Utility: 100% (6/6)

Knowledge Utility = fixtures whose required knowledge was delivered ÷ total. Regression guard: CI fails if any fixture stops passing — the same discipline as the compiler evals, applied to behavior. This is the differentiator: repos that only promise "compounding memory" can't measure whether the memory changed anything.

Real-repo evaluation ​

For a heavyweight end-to-end check, run the whole pipeline against a real, popular repo (default: FastAPI — large, active, well-documented Python):

bash
wf eval real                          # clone + full pipeline + LLM ingest
wf eval real --skip-llm                   # corpus + context metrics, no LLM
wf eval real --repo /path/to/clone
wf eval real --json

It seeds a throwaway fabric with both integrations enabled, captures a scoped slice of the repo's docs, ingests with LLM claim extraction (or seeds doc-backed patterns in --skip-llm mode), compiles probe-task manifests, and measures each tier:

TierMetric
Corpussources captured, claims extracted, locator rate, lint errors
Contextmanifest hit-rate per probe task (selected > 0 with reasons)
GraphifyAST symbols indexed from the real repo (e.g. 4,992 for FastAPI), entity pages written
Embeddingsre-rank delta when sentence-transformers is installed; reported as config-only when absent

The graphify/embeddings tiers are measured as integrations: their contribution is reported separately, so you can see exactly what enabling them buys.

PR replay evaluation ​

Replay real PRs from an active repo and measure whether the fabric would have helped when the PR was written — a knowledge-recall eval on real-world vocabulary:

bash
wf eval pr --repo tiangolo/fastapi --prs 16192 --llm
wf eval pr --repo tiangolo/fastapi --auto 3    # pick recent merged PRs (deps-bumps skipped)
wf eval pr ... --json

Per replayed PR (fetched live via gh — title, body, review comments, touched files):

  • Graphify tier: the repo's code is AST-indexed; the report shows entity pages written and how many of the PR's touched files have code-symbol coverage (e.g. 4,992 symbols / 913 pages from FastAPI; 1/2 touched files covered).
  • Zero-LLM tier: manifest hit-rate over docs-only corpus (honest baseline).
  • --llm tier: the PR's own discussion is ingested with LLM claim extraction; the manifest re-compiles and the report shows after LLM ingest (12 claims): coverage 0.414 (Δ +0.414) — how much the PR's own knowledge improves later recall of its task.

Term coverage = share of the PR's own vocabulary (title + body + comments, stemmed) present in the selected artifacts — a recall proxy on real work. The coverage delta isolates the value of the ingest step itself.

Stability & sensitivity evaluation ​

Closes the reproducibility requirements with hard, machine-checkable gates:

bash
wf eval stability --skip-llm          # deterministic gates (CI)
wf eval stability --models qwen2.5-coder:7b,qwen3.8:27b-mlx
wf eval stability --record            # append metrics to registry/log.md
GateWhat it provesThreshold
G1 context determinism20 manifest runs byte-identicalexact (0-token code must not vary)
G2 rebuild determinismrebuild-index output identical across runsexact
G3 ingest stabilitysame model, same source, 2 runsfuzzy word-coverage ≥ 0.7 (exact Jaccard reported — literal word-set overlap; paraphrase ≠ instability)
G6 locator presenceevery claim carries an L-locator1.0
G4 model sensitivitytwo models on the same sourcefuzzy ≥ 0.5 (target 0.8); a fail means model swap requires re-running compiler evals
G4-J judged refinementnear-miss G4 pairsjudgment tier re-asks "same factual content?" — judged-SAME downgrades wording drift to a warning (--judge)

These runs produced findings worth knowing. First, the deterministic parts are exactly deterministic. Second, the same model re-extracting the same source paraphrases rather than repeating. Third — the important one — different models initially disagreed wildly, and the fix was a prompt contract, not a bigger model. (Recorded in registry/log.md; commands are reproducible.): context/rebuild are exactly deterministic (~40ms); same-model extraction is semantically stable (fuzzy 0.76–0.86) but drifts in wording (exact 0.31–0.73). A 5-model matrix — local (qwen2.5-coder:7b, qwen3.8:27b-mlx) vs Ollama cloud (deepseek-v4.1-flash:cloud, kimi-k2.7-code:cloud, glm-5.3-flash:cloud) — first showed capability-correlated disagreement (fuzzy 0.28–0.59). Fixing it took a contract change, not a model change: the CLAIM_PROMPT now carries an explicit granularity spec (exactly one verifiable fact per claim, target 5–12). After that: all 10 pairwise G4 checks PASS (fuzzy 0.75–0.85) and claim counts converge to 11–12 across every model. The compiler change was re-baselined against the golden corpus first (recall 0.89, locator 1.00 — PASS). Infra note: cloud reasoning models can silently return 0 claims when reasoning consumes the token budget — the extractor now budgets 16K tokens and retries on finish_reason=length.

Model policy: ops vs compiler models ​

The tier table and provider setup live in Configuration; this page covers the policy — why the compiler model is gated and how evals enforce it.

Local vs cloud: what the small-model tests actually showed ​

The fabric's privacy tiering promises "sensitive repos extract on-device" — which is only honest if on-device models can actually extract claims at governance-grade quality. That's a testable claim, so it was tested. The findings below are recorded in registry/log.md with reproducible commands (wf eval golden, wf eval stability).

The benchmark ​

Both candidate small models ran the golden corpus (9 golden keys, recall/locator/quote-rate gates) plus the stability gates (G3 self-stability, G4 cross-model agreement against the cloud compiler):

ModelSizeGolden recallLocatorQuote rateVerdict
qwen2.5-coder:7b (GGUF/llama.cpp)4.7 GBPASS1.0 PASS0.88 PASSsolid local fallback
gemma4:e4b stock QAT (quantization-aware training)6.1 GB0.89 PASS1.0 PASS0.65 FAILclose, two gates failed
gemma4:e4b-fixed (temp 0.1 baked)6.1 GB0.89 PASS1.0 PASS0.96 PASSviable local compiler
deepseek-v4.1-flash:cloud—0.891.00.96precision-critical tier

The interesting part: two fixes, not a bigger model ​

The stock gemma4 E4B failed the quote-verbatim gate — 35% of returned quotes were markdown-normalized paraphrases, not verbatim source. Two fixes closed it:

  1. Deterministic quote repair (in verify_and_fix_locators): models strip markdown artifacts and join lines; a fuzzy-match pass rewrites model quotes back to the true verbatim source text. Quote rate 0.65 → 0.96 — and this fix benefits every model, cloud included.
  2. Temperature baked into the model build: the stock QAT gemma ships temperature 1.0, silently overriding API calls. A fixed variant (temp 0.1 + top_k 10) made G3 self-stability pass.

Lesson: small-model failures were often contract failures. The same pattern as the G4 model-sensitivity finding — when 5 models disagreed wildly (fuzzy 0.28–0.59), the fix wasn't a bigger model, it was an explicit granularity spec in the extraction prompt (one verifiable fact per claim, target 5–12), after which all 10 pairwise G4 checks passed (0.75–0.85).

What small models cost you ​

Honest accounting from the runs:

DimensionCloud compiler (deepseek-v4.1)Local gemma4 E4B
Golden recall0.890.89 (ties)
Cross-agreement (G4 vs deepseek)—0.78 (bar 0.75, target 0.8)
Latency~5–6s/fixture~6.1s/fixture (MLX/GGUF on Apple Silicon)
Memory—~2.5 GB (MLX 4-bit on Apple Silicon; GGUF Q4_K_M is ~4 GB), one model at a time
Privacyraw docs leave the machinezero egress
Costtokens per documentzero

The 0.78-vs-0.8 G4 shortfall is small-model capability variance — acceptable for the local tier, because the architecture assigns the precision-critical path to the cloud compiler. The local tier exists for privacy and availability, not to beat the cloud on recall.

Sources. Every number in this section is recorded in registry/log.md inside your fabric dir (~/.local/share/wiki-fabric/ after a standard install) — grep for the model id to find the eval block:

bash
grep -B2 -A8 "gemma4:e4b-fixed" ~/.local/share/wiki-fabric/registry/log.md

Each entry carries the command that produced it (wf eval golden, wf eval stability --record), so any number can be re-run to verify.

Policy consequences ​

  • The compiler-eval gate (promote/mine refuse without a recorded PASS eval) applies to whichever model does compiler work — local routes included. A local compiler needs its own golden-corpus pass before it can drive promotion.
  • llm.local_model defaults are models that passed the corpus; swapping them for untested small models is exactly the G4 lesson — re-run wf eval golden + wf eval stability --record first.
  • Extraction quality varies more by prompt contract than by model within a capability class — that's why the prompt is versioned, and prompt changes require re-baselining (the granularity-spec change was re-baselined against the golden corpus first: recall 0.89, locator 1.0, before rolling out).

Claim extraction is compiler work — it compiles sources into the fabric's canonical evidence — so it runs on a policy-designated compiler model, not whatever model is configured for cheap queries:

SettingDefaultUsed for
llm.ops_modelqwen2.5-coder:7bops: capture, status, cheap query synthesis (legacy llm.model reads)
llm.compiler_modeldeepseek-v4.1-flash:cloudclaim extraction (ingest --extract-claims), synthesis, promotion mining
llm.local_modelplatform default: gemma4:e4b-fixed (ollama-served); offline fallbacks mlx-community/gemma-4-e4b-it-4bit (Apple Silicon) / unsloth/gemma-4-e4b-it-GGUF (other)resolves repos.<slug>.<stage>: local routes — extraction, synthesis, dossier generation on-device
bash
# Override per-run (e.g. test a different compiler model):
WIKI_LLM_COMPILER_MODEL=kimi-k2.7-code:cloud wf ingest evidence/raw/<doc>.md --extract-claims
# Check the local model is cached; offers a human-gated download when missing:
wf models ensure [--yes]      # --check exits 0/1 without prompting (for scripts)

Local routes: repos.<slug>.extract/synthesize/dossier: "local" runs that stage on-device (privacy tiering). The backend is picked from the model id — GGUF (llama-cpp-python, universal — works on macOS/Linux/Windows) or MLX (mlx-lm, Apple Silicon only). Chat-tuned GGUF models (gemma etc.) run through their chat template automatically. Missing models are offered for download at first use (only the preferred quantization split is fetched, ~4 GB), or pre-fetch with wf models ensure. Install the backends with pip install -e ".[local]". GGUF is the recommended testing default: one model format runs identically on every platform.

The gate: wf promote and wf mine promotions refuse to run unless registry/log.md contains a recorded compiler eval for the current compiler model. This enforces the G4 lesson — model swaps are compiler changes, and compiler changes require re-evaluation — as a hard gate instead of a docs sentence. wf status shows all three models (ops, compiler, local).

Alpha — expect breaking changes.