Skip to content

Troubleshooting ​

Symptom-first index of the failure modes we know about. Anything not here — file an issue.

Installation & environment ​

wf command not found after install ​

The one-liner symlinks wf into ~/.local/bin/. If that directory isn't on your PATH:

bash
export PATH="$HOME/.local/bin:$PATH"   # add to .zshrc / .bashrc

uv can't be installed → fallback mode ​

wf falls back to plain python3 when uv is unavailable. Core commands (status, query, context, lint) work; LLM-dependent commands need deps:

bash
pip install pyyaml openai anthropic   # core deps (pyproject.toml is the source)

Tests fail with ModuleNotFoundError: huggingface_hub ​

The fast venv (core deps from pyproject.toml) doesn't include model-backend deps by design. The ensure/download tests skip automatically. To run them:

bash
pip install -e ".[local]"         # huggingface_hub + backend for your platform
python3 -m pytest tests/test_local_model.py -m live

LLM providers ​

OpenAI-compatible extraction failed: <error> ​

Check, in order:

  1. Endpoint reachable: curl $WIKI_LLM_BASE_URL/models (or the base_url from fabric.yaml)
  2. Model name spelled as the server expects — Ollama needs the exact tag (qwen2.5-coder:7b); OpenRouter/OpenAI take their own ids
  3. WIKI_LLM_TIMEOUT (default 600s) — reasoning models on long extractions may need more
  4. API key: Ollama ignores it; anything else needs a real key

Extraction chains through fallbacks automatically (OpenAI-compatible → Anthropic → opencode CLI). The stderr messages tell you which layer failed.

Reasoning models return 0 claims ​

DeepSeek-style reasoning models can spend the whole token budget thinking and never emit the JSON array. The extractor already retries on finish_reason=length — if you still get zero claims, lower WIKI_MLX_REASONING=low (on-device) or pick a non-reasoning compiler model.

On-device models (GGUF / MLX) ​

mlx backend unavailable / mlx-lm not installed ​

mlx-lm only exists on Apple Silicon macOS. On other platforms use the GGUF backend: set llm.local_model: unsloth/gemma-4-e4b-it-GGUF (or any *.gguf id) in fabric.yaml — llama-cpp-python runs everywhere. Install:

bash
pip install -e ".[local]"

local model '<id>' not found locally; download with: wf models ensure ​

The model isn't in your HuggingFace cache. Interactive runs offer a y/N prompt automatically; non-interactive environments (CI, scripts) print the hint and continue. Pre-fetch explicitly:

bash
wf models ensure --yes            # download llm.local_model without prompting
wf models ensure --model <hf-id>  # specific model
wf models ensure --check          # exit 0/1 for scripts, no prompt

no .gguf file found for '<id>' ​

GGUF repos ship many quantization splits; the resolver prefers Q4_K_M at the repo root. If your repo keeps quants in subdirectories, download the file manually and point llm.local_model at a local path instead of an HF id.

download failed: 401/403 (private/gated models) ​

Gated HF repos need a token: export HF_TOKEN=... (huggingface_hub reads it), or download manually and point llm.local_model at the local directory.

GGUF output is empty / synthesis falls back ​

Some chat-tuned models return nothing from raw text completion. The on-device backend routes chat-tuned models through their chat template automatically — if you hit this with an explicit local dir (not an HF id), make sure the dir contains the model's *.gguf plus tokenizer files. The failure path prints local synthesis failed ... — falling back; extraction then uses your cloud compiler model (check stderr to confirm which backend actually ran).

Git hooks ​

Hook doesn't fire / capture never runs ​

bash
wf hook status    # per-repo hook state
  • The hook skips commits touching only fabric-owned paths (evidence/, registry/, …) — that's the anti-loop rule, not a bug.
  • WIKI_SKIP_HOOK=1 git commit ... skips once per command.
  • Hook output goes to ~/.cache/wiki-fabric-hook.log (detached; commit returns instantly).

Hook triggered ingest on every commit ​

You installed with --extract-claims; each drift commit runs an LLM call. Re-install without the flag for capture-only (0 tokens), ingest manually with wf ingest --changed <slug> --extract-claims when you want extraction.

Second machine has no project config after cloning ​

The project's .wiki-overlay.md was never committed (bootstraps older than H4/#180 didn't track it). wf doctor names the repos; the fix:

bash
wf overlay-track        # stages + commits every connected project's overlay

After the commit, push from each repo; clones carry the config from then on.

Exit codes / project addressing ​

wf capture <slug> (or capture-git, or freshness) exits 3 mentioning a connected list ​

The slug isn't a connected project — the old behavior (exit 0, clean-looking) was the bug (#167). The error prints the connected-project inventory and a closest-match hint:

bash
wf projects        # the connected inventory (slug/path/route/captures)
wf capture <the-exact-slug-from-the-list>

Underscore↔kebab differences still collide correctly (comfyui_mcp and comfyui-mcp address one identity through the canonical seam). If both spellings REALLY appear separately, you're mid-fork — wf lint's IDENTITY gate names the collision.

freshness/hook "drift indistinguishable from success" ​

Not anymore: capture-family exit 2 = drift, 3 = unknown slug; freshness 0/1/2; gate 1 = actionable. The table lives in cli.md ("Exit-code contracts"); scripts (hooks, corpus CI) parse off it.

Config & lint ​

LLM-CONFIG llm.local_model: '...' does not look like an on-device model id ​

Ollama-style tags (qwen2.5-coder:7b) and provider-namespaced ids (openai/gpt-4o) can't run on-device. Use an HF id (mlx-community/*, unsloth/*), a *.gguf file, or an existing local path.

Lint passes locally, okflint fails in CI ​

The external validator reads okf-base.yaml. If you added a new generated/non-fabric directory (like a docs build), add it to exclude_patterns in okf-base.yaml so the two linters agree on what's a concept page.

SOURCE-DRIFT errors after re-capture ​

Hash drift means the raw file changed since its claims were extracted — exactly the signal the fabric is designed to catch. Re-ingest the affected sources (wf ingest --changed <slug>) to produce a change-set, then review and merge. On re-ingest the old revision's claims are mechanically stamped stale_after + flip contested (0 tokens, before the new record is written); wf review --auto-reverify re-grounds the ones whose quote still holds in the new revision.

Claims suddenly contested with stale_after — why? ​

That's the sha256-drift trigger doing its job: re-captured source ⇒ old- revision evidence expired ⇒ derived claims demoted until re-extraction re-grounds them. Recovery: wf ingest --changed <slug> --extract-claims, then wf review --auto-reverify. A claim whose quote genuinely no longer exists stays contested for human review.

IDENTITY repos.<key>: canonical slug '<slug>' collides ... ​

Two config keys fold to the same repo name (an underscore dir name and a kebab fabric.yaml key are ONE identity — the seam enforces it). Rename one of the config keys.

IDENTITY namespace '<raw>' is a non-canonical spelling ... (warning) ​

Your overlay's namespace: is underscore-form. Lookups fold either way, but new projects should use the canonical (kebab) form — rename at leisure, or add an ontology alias if the old spelling must stay matchable.

VOCABULARY <page>: domain '<x>' resolves to no ontology domain ​

The page declares a domain: the ontology doesn't know. Fix: approve the domain (promote-domains --apply), add an alias under the ontology's ## Aliases (if the spelling is a legitimate variant), or drop the field (the page stays unbound in concepts/).

LAYOUT-GUARD <file>:<line>: corpus path re-spelled outside layout.py ​

Code (probably a new script or AI edit) joined CORPUS_ROOT / "evidence"... directly. Compose via scripts/lib/layout.py accessors — the guard exists because re-spelled paths are how the corpus forked into dual trees. (The full AST guard runs in tests; this lint tier catches re-spells at commit time.)

Dual trees for one project (wiki/projects/foo_mcp/ + foo-mcp/) ​

Legacy symptom of the old raw-spelling bug. The project_slug seam + hook fold (S3) prevent new forks; regenerate the wiki (wf export wiki --deep-dives) after confirming the canonical tree — the migration deleted the duplicate. Canonical slugs are kebab.

A source record says status: expired ​

The upstream source vanished (raw file deleted) — detected by wf review --verify-sources. The record is a provenance tombstone: claims derived from it were stale-stamped; nothing re-captures it automatically. If the source is genuinely gone, leave the tombstone; if it moved, re-capture creates a fresh record.

SECRETS <page>: ... pattern detected ​

The #188 lint rule found a credential-shaped string in a corpus page — token prefixes match standalone text only (slugs like …risk-monitoring-setup-md don't trip it), so a match is a real key-shaped string in authored content. Recovery: rotate the credential first, remove/redact the page, re-capture or re-ingest clean. Real keys belong in <fabric>/secrets.env (gitignored); authored pages reference the env-var NAME, never the value. Matches are masked in lint output (the report itself gets committed). evidence/raw/ is exempt — the capture plane holds upstream docs verbatim and never lints for content.

Auto-promote is OFF (tuning.promotion.auto_apply: false) ​

The #190 tier is off until a fabric turns it on explicitly — a policy change must be explicit. Add the tuning.promotion block (configuration.md), then re-run wf promote-patterns --auto. If you ran it from the corpus CI expecting automation: the tier is deliberately a HUMAN command; the scheduled mining pipeline stays promotion-free.

Where to look when something else breaks ​

  • wf status — health, inventory, all three models, lint summary
  • ~/.cache/wiki-fabric-hook.log — hook activity
  • registry/log.md — the fabric's own operation timeline
  • stderr: every extraction/backend fallback prints why it fell back

Source says "ingested" but has no claims ​

A pre-fix bug (#93) flipped pending→ingested even with 0 claims extracted (LLM outage, malformed JSON). Recover:

bash
wf ingest --reclaim <project>     # flips zero-claim records back to pending
wf ingest --pending <project> --extract-claims

Lint's SOURCE-EMPTY warning names any stranded source.

A claim's locator doesn't match the source anymore ​

Sources get re-captured and edited; locators drift.

bash
wf review --verify-locators       # rewrite/repair/restore/contest — 0 tokens

Lint is slow on a big corpus ​

The shared frontmatter cache makes lint parse each file once (#108). If lint is still slow, check for enormous evidence/raw/ subtrees (excluded by default) or run wf lint --orphans scoped checks.

The packaged wf shows STALE ​

It shouldn't: packaged installs carry their own orchestrator (install ≡ behavior). If you see CLI: STALE, you're running a dev-mode checkout copy — wf update pulls the harness; uv tool upgrade wiki-fabric upgrades the packaged tool.

Alpha — expect breaking changes.