Configuration
Install first: Getting Started. Everything in wiki-fabric is configured through one gitignored file, fabric.yaml (template: fabric.yaml.example), plus environment variables that override it per-run.
Where it lives: ~/.local/share/wiki-fabric/fabric.yaml after a standard install (the fabric dir — content and config live there, never in the harness clone). WIKI_FABRIC_DIR relocates the whole fabric. Dev mode: a harness clone holding its own fabric.yaml doubles as the fabric.
The complete fabric.yaml (annotated)
A full-featured example showing every key in context. Copy the sections you need — every field is optional except owner; defaults come from fabric_config.py (the source of truth for defaults).
# ─── Identity ────────────────────────────────────────────────────────────────
owner: your-name # your handle; used in actor conventions
# (agent/<owner>/<model>) across all pages
# ─── LLM: three model roles ─────────────────────────────────────────────────
llm:
# Endpoint: any OpenAI-compatible /v1 API (Ollama, OpenAI, OpenRouter,
# Together, LM Studio, vLLM, llama.cpp server). See provider table above.
base_url: http://localhost:11434/v1
api_key: ollama # Ollama ignores it; cloud providers need a real key
ops_model: qwen2.5-coder:7b # OPS model — cheap queries, capture, status
# (legacy key `model:` still reads — renamed #190-adjacent)
# COMPILER model — claim extraction, synthesis, promotion mining.
# The most consequential knob: the fabric's canonical evidence is compiled
# by this model, and swapping it requires a recorded compiler eval.
compiler_model: deepseek-v4.1-flash:cloud
# LOCAL model — resolves repos.<slug>.<stage>: "local" routes (on-device,
# zero egress). Defaults per platform if unset:
# Apple Silicon: mlx-community/gemma-4-e4b-it-4bit (MLX, ~2.5 GB)
# elsewhere: unsloth/gemma-4-e4b-it-GGUF (GGUF, ~4 GB Q4_K_M)
# Missing models are offered for download at first use (human-gated y/N),
# or pre-fetch with: wf models ensure [--yes]
local_model: gemma4:e4b-fixed # ollama-served (local tier); MLX/HF ids work offline too
# ─── Connected repos (namespaces): each project gets its own folder ────────
# inside the fabric, so its claims stay separate from global knowledge ─────
repos:
# Every connected project gets an entry. path is relative to the fabric
# root (or absolute). The fabric itself can be its own namespace ("path: .")
wiki-fabric:
path: .
graph_dir: graphify-out # optional: per-repo graphify graph dir
my-oss-project:
path: ../my-oss-project # all stages cloud (default) — fastest
my-private-repo: # privacy tiering per stage
path: ../my-private-repo
graph_dir: graphify-out
extract: local # raw docs never leave the machine (highest sensitivity)
synthesize: local # extracted claims stay local too
dossier: cloud # experience events are safe for cloud
# Values per stage: "cloud" (default) | "local" (uses llm.local_model)
# | an explicit model id (e.g. "qwen2.5-coder:7b" or an HF id)
# Stages: extract (sees raw docs) · synthesize (sanitized claims)
# · dossier (experience events)
# OpenAI directly instead of Ollama? Override per-fabric or per-repo:
# llm:
# base_url: https://api.openai.com/v1
# api_key: sk-your-key-here
# compiler_model: gpt-4o
# ─── Ignores (exclude-side filter) ──────────────────────────────────────────
# Applied by capture, entity index, context, lint, rebuild-index, okf export.
# One list, auto-classified: globs by default, regex when metacharacters
# appear, `regex:` prefix as escape hatch.
ignore:
patterns:
- "vendor/**" # glob: fnmatch-style, ** crosses dirs
- "docs/generated/**"
- "**/*.min.js"
- "_archive\\d+/" # regex: \d triggers auto-classification
- "regex:^node_modules/" # regex: prefix forces regex matching
projects: # per-repo patterns, unioned with global
my-godot-game:
patterns:
- "addons/generated/**"
# ─── Optional integrations (all off by default) ─────────────────────────────
integrations:
graphify:
enabled: false # true → call-graph staleness, claim
graph_dir: graphify-out # enrichment, query graph expansion
embeddings:
enabled: false # semantic re-rank boost (off by default; shipped)
model: all-MiniLM-L6-v2
# ─── Domain taxonomy ────────────────────────────────────────────────────────
# Drives classification + context scoping. Grow with: wf propose-domains
domains:
agent-systems:
signals: [agent, mcp, fastmcp, opencode, claude]
web-systems:
signals: [fastapi, flask, react, nextjs, supabase, postgresql]Cheat sheet — who reads what:
| Key | Read by |
|---|---|
llm.ops_model | query synthesis, capture, status (legacy key llm.model still reads) |
llm.compiler_model | ingest extraction, synthesize, mine-promotions, compiler-eval gate |
llm.local_model | any local route (extract/synthesize/dossier) |
repos.<slug>.path | capture, entity index, graphify bridge, hooks |
repos.<slug>.<stage> | ingest, synthesize, mine-promotions routing |
ignore.patterns / .globs / .regexes | capture, context, lint, rebuild-index, okf export |
integrations.* | graphify bridge, skills — fabric-global (no per-repo override; judgment included) |
domains.* | classification, context scoping, propose-domains (the ontology is the vocabulary; config signals are overrides) |
tuning.git_history | capture-git window/budget (since, budget:int|"all") |
repos.<slug>.git_history | per-repo window/budget override |
tuning.ingest.budget | ingest --changed/--pending budget (0 = uncapped) |
tuning.promotion.auto_apply / .auto_threshold | #190 judged auto-apply tier (inbox candidates): opt-in switch (default off) + confident floor (0.90); near-band escalates |
LLM providers (any OpenAI-compatible endpoint)
The fabric talks to any OpenAI-compatible /v1 endpoint. Set it once in fabric.yaml:
llm:
base_url: http://localhost:11434/v1 # any OpenAI-compatible endpoint
api_key: ollama # Ollama ignores this; cloud providers need a real key
ops_model: qwen2.5-coder:7b # ops model (see "Model tiers" below; legacy `model:` reads)Tested endpoints:
| Provider | base_url | Notes |
|---|---|---|
| Ollama (default) | http://localhost:11434/v1 | api_key: ollama (ignored) |
| OpenAI | https://api.openai.com/v1 | api_key: sk-... |
| OpenRouter | https://openrouter.ai/api/v1 | 400+ models, one key |
| Together AI | https://api.together.xyz/v1 | open models, cheap |
| LM Studio | http://localhost:1234/v1 | local server |
| vLLM | http://localhost:8000/v1 | self-hosted |
| llama.cpp server | http://localhost:8080/v1 | self-hosted |
# Example — OpenAI directly:
llm:
base_url: https://api.openai.com/v1
api_key: sk-your-key-here
compiler_model: gpt-4oEnvironment variables (override fabric.yaml per-run)
| Variable | Default | Purpose |
|---|---|---|
WIKI_LLM_BASE_URL | http://localhost:11434/v1 | OpenAI-compatible endpoint |
WIKI_LLM_API_KEY | ollama | API key |
WIKI_LLM_OPS_MODEL (legacy WIKI_LLM_MODEL) | qwen2.5-coder:7b | Ops model name |
WIKI_LLM_COMPILER_MODEL | deepseek-v4.1-flash:cloud | Compiler model name |
WIKI_LLM_LOCAL_MODEL | platform-split (below) | On-device model id |
WIKI_LLM_TIMEOUT | 600 | Request timeout (seconds) |
WIKI_LLM_BACKEND | — | Set to mlx to force the on-device backend |
WIKI_MLX_MODEL | — | Explicit on-device model (wins over llm.local_model) |
WIKI_MLX_REASONING | low | Reasoning effort for templates that accept it |
WIKI_LOCAL_MAX_TOKENS | 4096 | Synthesis token budget on-device |
WIKI_LOCAL_N_CTX | 16384 | GGUF context window |
WIKI_INGEST_WORKERS | 1 | Concurrent extraction threads (cloud tolerates 6–12) |
HF_HOME / HF_HUB_CACHE | ~/.cache/huggingface | Where local models are cached |
WIKI_FABRIC_DIR | ~/.local/share/wiki-fabric | Fabric location (content + config) |
Model tiers: ops vs compiler vs local
Three model roles, each independently configurable:
| Setting | Default | Used for |
|---|---|---|
llm.ops_model | qwen2.5-coder:7b | ops: capture, status, cheap query synthesis (legacy llm.model reads) |
llm.compiler_model | deepseek-v4.1-flash:cloud | claim extraction, synthesis, promotion mining (compiler work) |
llm.local_model | gemma4:e4b-fixed (ollama-served, egress-free) — offline/HF alternatives: mlx-community/gemma-4-e4b-it-4bit (Apple Silicon) / unsloth/gemma-4-e4b-it-GGUF (other) | resolves repos.<slug>.<stage>: local routes; ollama tags need ollama pull <tag> |
Why a separate compiler model? Claim extraction compiles sources into canonical evidence — the fabric's most sensitive operation. Cross-model extraction disagreement is capability-correlated, so compiler work runs on the most capable model, and swapping it requires a recorded compiler eval (Model Policy & Evals).
The design reasoning — which model class fits which stage, and why compiler swaps are eval-gated — is in Why: Model splits.
On-device (local) routes: privacy tiering
Per-repo, per-stage routing sends sensitive work to on-device models:
repos:
my-private-repo:
path: ../my-private-repo
extract: local # raw docs never leave the machine
synthesize: local # extracted claims stay local too
dossier: cloud # experience events are safe for cloudStages: extract (sees raw docs — highest sensitivity), synthesize (sanitized statements), dossier (experience events). Values: "cloud" (default) · "local" (on-device via llm.local_model) · any explicit model id.
Integrations are fabric-global. There is deliberately no per-repo integrations: override (an earlier version deep-merged repos.<slug>.integrations.<name> over the global block so a privacy-sensitive repo could run judgment locally — the seam is gone): the judgment tier's G-J calibration gate is keyed to ONE judge identity per fabric, and per-repo route pinning would multiply identities and silently fork calibrations. Route a specific repo's extraction instead — repos.<slug>.extract: local keeps that repo's raw docs on-device while the judge stays one shared setting:
repos:
sensitive-repo:
extract: local # raw docs never leave the machine (stage routing)
integrations:
judgment: {enabled: true, route: local} # one judge, whole fabricgraphify stays repo-adjacent only through its per-repo graph_dir (a path, not an enable/route decision); embeddings and obsidian are correctly global (one index, one vault).
Backends, chosen by model-id shape:
| Backend | Runs on | Model ids |
|---|---|---|
| GGUF (llama-cpp-python) | macOS / Linux / Windows — the universal testing default | *.gguf ids, e.g. unsloth/gemma-4-e4b-it-GGUF |
| MLX (mlx-lm) | Apple Silicon only — fastest on M-series | mlx-community/* ids, e.g. mlx-community/gemma-4-e4b-it-4bit |
Chat-tuned models (gemma etc.) run through their chat template automatically. Missing models are offered for download at first use (human-gated y/N, only the preferred quantization split is fetched), or pre-fetch:
wf models ensure [--yes] # check + offer download; --check exits 0/1 for scriptsInstall the backends: pip install -e ".[local]".
Quality expectations: the default local models pass the golden-corpus eval (recall 0.89, quote-verbatim 0.96) and agree with the cloud compiler at 0.68–0.78 on the cross-model agreement gate (G4 — see Model Policy & Evals) — the full local-vs-cloud test story with numbers lives in Model Policy & Evals.
Connecting repos: config lives in the project
Every project carries its own config in .wiki-overlay.md (written by wf bootstrap, versioned with the project repo) — identity, domains, capture globs, and the LLM routing: block:
# /path/to/my-project/.wiki-overlay.md (frontmatter)
project: my-project
namespace: my-project
routing:
extract: local # raw docs never leave the machine
synthesize: local
dossier: cloudThe fabric discovers these automatically: it scans its sibling directories for .wiki-overlay.md files (respects namespace: for the slug, folded to its canonical form — comfyui_mcp and comfyui-mcp are one identity: lookups, hooks, and wf status address either spelling; lint's IDENTITY gate errors on canonical collisions and warns on non-canonical namespaces), so a bootstrap'd project appears in wf status with no fabric.yaml edit. Explicit repos: entries merge over the overlay — explicit keys win (matched by canonical fold: a kebab key configures an underscore-named project):
repos:
# auto_discover: siblings is the default (set false to disable)
# Only exceptions need entries here:
my-private-repo:
path: ../elsewhere/my-private-repo # non-sibling location
extract: local # explicit routing overrides overlay
wiki-fabric:
path: . # the fabric itselfStage and value semantics are the same as On-device routes above: stages extract / synthesize / dossier; values "cloud" (default), "local", or an explicit model id.
Integrations are fabric-global (no per-repo integrations: override — the judgment seam was removed; see On-device routes above for the rationale and the stage-routing alternative). A stale repos.<slug>.integrations: block in an old fabric.yaml is inert. graphify remains repo-adjacent via its per-repo graph_dir; embeddings and obsidian are correctly global (one index, one vault).
Migrating an existing fabric: wf repos migrate --dry-run shows which per-repo keys would move into overlays; --apply --prune writes them and removes the migrated keys from fabric.yaml (backup at fabric.yaml.bak). Discoverable siblings become fully overlay-driven — their repos: entries disappear entirely. After pruning, a multi-project fabric.yaml is ~15 lines: owner, llm, integrations, domains, and repos: entries only for the fabric itself and non-sibling repos.
The vault is standalone output: generated wiki content lands there via wf export wiki, never copies/symlinks of corpus content. A machine can have multiple vaults by pointing vault.path at different targets (see CLI Reference).
Ignoring files
Exclusion filter applied by capture, entity index, context, lint, rebuild-index, and okf export.
One list, auto-classified (preferred — ignore.patterns):
ignore:
patterns:
- "vendor/**" # glob (default): fnmatch-style, ** crosses directories
- "docs/generated/**"
- "**/*.min.js"
- "_archive\\d+/" # regex — auto-detected by metacharacters (\d)
- "regex:^temp-\\d+\\.tmp$" # `regex:` prefix forces regex (escape hatch)Classification rules for patterns entries:
regex:/re:prefix → regex (prefix stripped) — use this for ambiguous strings- contains a regex-only metacharacter (
\(){}+|^) → regex (re.searchon the posix path) - anything else → glob (
fnmatchon the full path or basename;**crosses directories)
The explicit keys still work and win for ambiguous strings:
ignore:
globs: # always fnmatch-style
- "vendor/**"
regexes: # always re.search
- "_archive\\d+/"
projects: # per-repo patterns, merged (union) with global
my-godot-game:
patterns: # unified key works here too
- "addons/generated/**"
globs:
- "addons/third-party/**"Invalid regexes don't crash anything (they're skipped at match time) but lint flags them as IGNORE-CONFIG errors — so typos surface at the 0-error gate instead of silently never matching.
Integrations
All off by default; see Optional Integrations:
integrations:
graphify:
enabled: true # call-graph staleness, claim enrichment, graph expansion
graph_dir: graphify-out
embeddings:
enabled: false # semantic re-rank boost (shipped; off by default)
model: all-MiniLM-L6-v2
judgment:
enabled: false # decision-model judging for eval gates (Jev / Laya)
route: cloud # "cloud" | "local"
cloud_model: jev-latest
obsidian:
enabled: false # two-way vault (Local REST API plugin required)
api_url: https://127.0.0.1:27124
api_key_env: OBSIDIAN_REST_KEYSync distribution
sync:
mode: solo # "solo" (default, sticky: direct push) | "team" (one PR per push)
evidence_prs: auto # team mode: evidence-only PRs auto-merge on green CI | "review"Judgment tier
integrations:
judgment:
enabled: true
route: cloud # "cloud" (TypeSafe Jev) | "local" (local backends)
cloud_model: jev-latest
api_key: <jev key> # or TYPESAFE_API_KEY env; gitignored file OK
# base_url: http://localhost:8000 # self-hosted judge (stuntd/laya-serve)
local_backend: ollama # ollama (tev1/nimble — System One on localhost; local-first)
# laya (auto) | laya-mlx | laya-torch | generic
local_model: tev1:latest # the ollama decision-model tag (judge tier)The judgment tier (integrations.judgment) enables low-variance decision-model judging for evaluation surfaces — wf eval behavior --judge. It is never used by wf query or lint (0-token core); wf context gains it only behind the explicit --judge-borderline flag (a test enforces the gate). route: cloud uses the TypeSafe Jev API — POST /v1/systemone, key via TYPESAFE_API_KEY (env or the machine-local secrets.env), endpoint overridable (base_url / TYPESAFE_BASE_URL, e.g. a self-hosted stuntd/laya-serve). route: local picks per machine — the ollama decision-model tier (local-first default suggestion: tev1:latest/nimble:latest via ollama's own /v1/systemone, egress-free, no key, no weights download), Laya-MLX typed heads on Apple Silicon, upstream laya (torch/ONNX) on Linux/Windows, or the generic GGUF route via llm.local_model. Auto-picked by platform unless you pin local_backend. Every judgment records backend, model, and probability — low-variance judgment, not determinism; near-threshold values escalate to human review. See Judgment — the decision-model tier.
Domains
Domain taxonomy drives classification and context scoping. The ontology (domains/ontology.md) is the vocabulary single-truth — canonical domain names (## Domains), alternate spellings (## Aliases: `godot-systems` -> `godot` — pages carry the spelling their synthesis wrote, the alias folds them without re-tagging), and the shared tag lexicon (## Shared tag set, the mining expansion vocabulary). Every consumer reads it through the shared parser (scripts/lib/ontology.py): context's domain tier binds domain: frontmatter through the alias map, synthesize/log-experience/mine-promotions/hubs bind from it, and lint's VOCABULARY gate warns when a page's domain: resolves to no ontology domain. Domain-bound pages live in physical homes (domains/<domain>/{concepts,questions,syntheses} — relocate-concepts.py migrates legacy flat pages).
wf propose-domains discovers new domains from evidence signals (ontology names fold into the signal lexicon; config signals below are overrides) and writes them as pending-review proposal dossiers (registry/domain-proposals/); merge an approved one into the ontology with wf promote-domains --apply <dossier> (aliases: a dossier may carry aliases: to record alternate spellings). wf gate surfaces pending proposals at session start:
domains:
agent-systems:
signals: [agent, mcp, fastmcp, opencode, claude]
web-ui:
signals: [fastapi, flask, react, nextjs, supabase, postgresql]
# web-systems (legacy spelling): alias-mapped to web-ui in the ontology —
# config keeps working while the vocabulary canon movesSecurity & privacy
Where secrets live. API keys go in fabric.yaml only — the file is gitignored, and the repo's .gitignore ships with that exclusion. Never put keys in fabric.yaml.example (committed), page frontmatter, or claims. The fabric's own pages never store credentials; lint has no secret scanner yet (roadmap), so the discipline is: keys only in the config file or env vars.
What data flows where. The fabric sends content to exactly two places:
| Operation | What leaves the machine | To where |
|---|---|---|
ingest --extract-claims on a cloud-routed repo | raw document text | your base_url endpoint |
ingest --extract-claims on a extract: local repo | nothing (on-device) | — |
wf synthesize-class stages (cloud route) | sanitized claim statements | your endpoint |
| local-route stages | nothing | — |
| dossier generation (cloud route) | experience events | your endpoint |
wf query | nothing — deterministic retrieval, 0 tokens | — |
wf context | nothing — deterministic compilation | — |
wf models ensure | model download request | huggingface.co (metadata + weights only) |
The threat model for sensitive repos: raw documents are the highest sensitivity tier (they may contain internal URLs, credentials-in-comments, customer names). That's why extract is the stage you route local — a private repo's raw docs never reach a cloud endpoint. Synthesized claims are deliberately sanitized statements, so synthesize: local is a second lock; dossier input (experience events) is usually safe for cloud.
Practical setup for mixed fleets:
# cloud compiler for everything...
llm:
compiler_model: deepseek-v4.1-flash:cloud
repos:
public-oss-project: {} # all stages cloud — fast
internal-project:
extract: local # raw docs never leave
synthesize: local # claims stay local too
dossier: cloud # experience events are fineInbound content is screened too. wf okf import runs a deterministic injection screen on every imported page (zero-width chars, control chars, "ignore previous instructions" patterns); hits are quarantined to evidence/_inbox/<scope>-okf/ for human review instead of entering the fabric. Trust tiers are recorded, never inherited — see OKF v0.2.
Team sync pushes knowledge content to a git remote you own. Use a private repo for anything sensitive; the harness (scripts/schemas) never syncs.