Troubleshooting
Symptom-first index of the failure modes we know about. Anything not here — file an issue.
Installation & environment
wf command not found after install
The one-liner symlinks wf into ~/.local/bin/. If that directory isn't on your PATH:
export PATH="$HOME/.local/bin:$PATH" # add to .zshrc / .bashrcuv can't be installed → fallback mode
wf falls back to plain python3 when uv is unavailable. Core commands (status, query, context, lint) work; LLM-dependent commands need deps:
pip install pyyaml openai anthropic # core deps (pyproject.toml is the source)Tests fail with ModuleNotFoundError: huggingface_hub
The fast venv (core deps from pyproject.toml) doesn't include model-backend deps by design. The ensure/download tests skip automatically. To run them:
pip install -e ".[local]" # huggingface_hub + backend for your platform
python3 -m pytest tests/test_local_model.py -m liveLLM providers
OpenAI-compatible extraction failed: <error>
Check, in order:
- Endpoint reachable:
curl $WIKI_LLM_BASE_URL/models(or thebase_urlfrom fabric.yaml) - Model name spelled as the server expects — Ollama needs the exact tag (
qwen2.5-coder:7b); OpenRouter/OpenAI take their own ids WIKI_LLM_TIMEOUT(default 600s) — reasoning models on long extractions may need more- API key: Ollama ignores it; anything else needs a real key
Extraction chains through fallbacks automatically (OpenAI-compatible → Anthropic → opencode CLI). The stderr messages tell you which layer failed.
Reasoning models return 0 claims
DeepSeek-style reasoning models can spend the whole token budget thinking and never emit the JSON array. The extractor already retries on finish_reason=length — if you still get zero claims, lower WIKI_MLX_REASONING=low (on-device) or pick a non-reasoning compiler model.
On-device models (GGUF / MLX)
mlx backend unavailable / mlx-lm not installed
mlx-lm only exists on Apple Silicon macOS. On other platforms use the GGUF backend: set llm.local_model: unsloth/gemma-4-e4b-it-GGUF (or any *.gguf id) in fabric.yaml — llama-cpp-python runs everywhere. Install:
pip install -e ".[local]"local model '<id>' not found locally; download with: wf models ensure
The model isn't in your HuggingFace cache. Interactive runs offer a y/N prompt automatically; non-interactive environments (CI, scripts) print the hint and continue. Pre-fetch explicitly:
wf models ensure --yes # download llm.local_model without prompting
wf models ensure --model <hf-id> # specific model
wf models ensure --check # exit 0/1 for scripts, no promptno .gguf file found for '<id>'
GGUF repos ship many quantization splits; the resolver prefers Q4_K_M at the repo root. If your repo keeps quants in subdirectories, download the file manually and point llm.local_model at a local path instead of an HF id.
download failed: 401/403 (private/gated models)
Gated HF repos need a token: export HF_TOKEN=... (huggingface_hub reads it), or download manually and point llm.local_model at the local directory.
GGUF output is empty / synthesis falls back
Some chat-tuned models return nothing from raw text completion. The on-device backend routes chat-tuned models through their chat template automatically — if you hit this with an explicit local dir (not an HF id), make sure the dir contains the model's *.gguf plus tokenizer files. The failure path prints local synthesis failed ... — falling back; extraction then uses your cloud compiler model (check stderr to confirm which backend actually ran).
Git hooks
Hook doesn't fire / capture never runs
wf hook status # per-repo hook state- The hook skips commits touching only fabric-owned paths (
evidence/,registry/, …) — that's the anti-loop rule, not a bug. WIKI_SKIP_HOOK=1 git commit ...skips once per command.- Hook output goes to
~/.cache/wiki-fabric-hook.log(detached; commit returns instantly).
Hook triggered ingest on every commit
You installed with --extract-claims; each drift commit runs an LLM call. Re-install without the flag for capture-only (0 tokens), ingest manually with wf ingest --changed <slug> --extract-claims when you want extraction.
Second machine has no project config after cloning
The project's .wiki-overlay.md was never committed (bootstraps older than H4/#180 didn't track it). wf doctor names the repos; the fix:
wf overlay-track # stages + commits every connected project's overlayAfter the commit, push from each repo; clones carry the config from then on.
Exit codes / project addressing
wf capture <slug> (or capture-git, or freshness) exits 3 mentioning a connected list
The slug isn't a connected project — the old behavior (exit 0, clean-looking) was the bug (#167). The error prints the connected-project inventory and a closest-match hint:
wf projects # the connected inventory (slug/path/route/captures)
wf capture <the-exact-slug-from-the-list>Underscore↔kebab differences still collide correctly (comfyui_mcp and comfyui-mcp address one identity through the canonical seam). If both spellings REALLY appear separately, you're mid-fork — wf lint's IDENTITY gate names the collision.
freshness/hook "drift indistinguishable from success"
Not anymore: capture-family exit 2 = drift, 3 = unknown slug; freshness 0/1/2; gate 1 = actionable. The table lives in cli.md ("Exit-code contracts"); scripts (hooks, corpus CI) parse off it.
Config & lint
LLM-CONFIG llm.local_model: '...' does not look like an on-device model id
Ollama-style tags (qwen2.5-coder:7b) and provider-namespaced ids (openai/gpt-4o) can't run on-device. Use an HF id (mlx-community/*, unsloth/*), a *.gguf file, or an existing local path.
Lint passes locally, okflint fails in CI
The external validator reads okf-base.yaml. If you added a new generated/non-fabric directory (like a docs build), add it to exclude_patterns in okf-base.yaml so the two linters agree on what's a concept page.
SOURCE-DRIFT errors after re-capture
Hash drift means the raw file changed since its claims were extracted — exactly the signal the fabric is designed to catch. Re-ingest the affected sources (wf ingest --changed <slug>) to produce a change-set, then review and merge. On re-ingest the old revision's claims are mechanically stamped stale_after + flip contested (0 tokens, before the new record is written); wf review --auto-reverify re-grounds the ones whose quote still holds in the new revision.
Claims suddenly contested with stale_after — why?
That's the sha256-drift trigger doing its job: re-captured source ⇒ old- revision evidence expired ⇒ derived claims demoted until re-extraction re-grounds them. Recovery: wf ingest --changed <slug> --extract-claims, then wf review --auto-reverify. A claim whose quote genuinely no longer exists stays contested for human review.
IDENTITY repos.<key>: canonical slug '<slug>' collides ...
Two config keys fold to the same repo name (an underscore dir name and a kebab fabric.yaml key are ONE identity — the seam enforces it). Rename one of the config keys.
IDENTITY namespace '<raw>' is a non-canonical spelling ... (warning)
Your overlay's namespace: is underscore-form. Lookups fold either way, but new projects should use the canonical (kebab) form — rename at leisure, or add an ontology alias if the old spelling must stay matchable.
VOCABULARY <page>: domain '<x>' resolves to no ontology domain
The page declares a domain: the ontology doesn't know. Fix: approve the domain (promote-domains --apply), add an alias under the ontology's ## Aliases (if the spelling is a legitimate variant), or drop the field (the page stays unbound in concepts/).
LAYOUT-GUARD <file>:<line>: corpus path re-spelled outside layout.py
Code (probably a new script or AI edit) joined CORPUS_ROOT / "evidence"... directly. Compose via scripts/lib/layout.py accessors — the guard exists because re-spelled paths are how the corpus forked into dual trees. (The full AST guard runs in tests; this lint tier catches re-spells at commit time.)
Dual trees for one project (wiki/projects/foo_mcp/ + foo-mcp/)
Legacy symptom of the old raw-spelling bug. The project_slug seam + hook fold (S3) prevent new forks; regenerate the wiki (wf export wiki --deep-dives) after confirming the canonical tree — the migration deleted the duplicate. Canonical slugs are kebab.
A source record says status: expired
The upstream source vanished (raw file deleted) — detected by wf review --verify-sources. The record is a provenance tombstone: claims derived from it were stale-stamped; nothing re-captures it automatically. If the source is genuinely gone, leave the tombstone; if it moved, re-capture creates a fresh record.
SECRETS <page>: ... pattern detected
The #188 lint rule found a credential-shaped string in a corpus page — token prefixes match standalone text only (slugs like …risk-monitoring-setup-md don't trip it), so a match is a real key-shaped string in authored content. Recovery: rotate the credential first, remove/redact the page, re-capture or re-ingest clean. Real keys belong in <fabric>/secrets.env (gitignored); authored pages reference the env-var NAME, never the value. Matches are masked in lint output (the report itself gets committed). evidence/raw/ is exempt — the capture plane holds upstream docs verbatim and never lints for content.
Auto-promote is OFF (tuning.promotion.auto_apply: false)
The #190 tier is off until a fabric turns it on explicitly — a policy change must be explicit. Add the tuning.promotion block (configuration.md), then re-run wf promote-patterns --auto. If you ran it from the corpus CI expecting automation: the tier is deliberately a HUMAN command; the scheduled mining pipeline stays promotion-free.
Where to look when something else breaks
wf status— health, inventory, all three models, lint summary~/.cache/wiki-fabric-hook.log— hook activityregistry/log.md— the fabric's own operation timeline- stderr: every extraction/backend fallback prints why it fell back
Source says "ingested" but has no claims
A pre-fix bug (#93) flipped pending→ingested even with 0 claims extracted (LLM outage, malformed JSON). Recover:
wf ingest --reclaim <project> # flips zero-claim records back to pending
wf ingest --pending <project> --extract-claimsLint's SOURCE-EMPTY warning names any stranded source.
A claim's locator doesn't match the source anymore
Sources get re-captured and edited; locators drift.
wf review --verify-locators # rewrite/repair/restore/contest — 0 tokensLint is slow on a big corpus
The shared frontmatter cache makes lint parse each file once (#108). If lint is still slow, check for enormous evidence/raw/ subtrees (excluded by default) or run wf lint --orphans scoped checks.
The packaged wf shows STALE
It shouldn't: packaged installs carry their own orchestrator (install ≡ behavior). If you see CLI: STALE, you're running a dev-mode checkout copy — wf update pulls the harness; uv tool upgrade wiki-fabric upgrades the packaged tool.