qwen38-iq3s reasons extensively before emitting the skill body. A max_tokens cap truncated the hidden reasoning (finish_reason: length) leaving content empty, so the body write failed with 'skill body is empty'. Omit max_tokens so the model runs to completion (~7 min); reasoning is filtered automatically since only content is read. Client timeout default raised to 1200s.
14 KiB
AGENTS.md
Guidance for working on the semif agent. Read this before touching code.
What this is
A local desktop CLI agent whose entire control flow is a single decision model
(SemIf). Inputs are gated, scored for urgency, queued, and dispatched through a
skill tree. Every SemIf decision is logged as a labeled training row; the
dream pass computes the prediction-vs-observation cost (cross-entropy / NLL +
ECE) that later drives fine-tuning.
The design spec is IDEA.md. The key point: semantic ifs, not text
generation, do the routing. SemIf returns probabilities conditional on the
supplied options; an LLM is used only for generation and self-assessment.
Architecture map
cli.py argparse: run (REPL / --script), dream, skills, status, relabel,
dashboard
scheduler.py gate -> choice(tau) -> score -> queue; preempt + requeue
queue.py urgency max-heap (desc weight, FIFO seq), age pulls toward 1.0
skills.py tree + registry (email.compose, response.reject, tracking.check),
navigation = SemIf choices per level (logged), create_category
and create_skill author + register stubs via the decision model
in generation mode; SkillBodyStore + materialize_skill persist
and hot-load runnable skill bodies from data/skills/
skill.py loop: observe -> predict -> act -> observe -> assess (LLM)
engine.py SemIfEngine -> semif_phase1.llamacpp_backend (lazy import)
codegen.py CodegenClient (OpenAI-compatible) writes runnable skill bodies
against SKILL.md; parse/validate (compile + predict/act)
llm.py OpenAI-compatible client for self-assessment (stdlib urllib)
log.py decisions.jsonl rows {state, question, options, predicted_probs,
selected, observed_outcome, label_source}
trace.py runs.jsonl lifecycle events keyed by run_id (submit/queued/
preempted/assessed/...); decisions reference run_id in extra
dream.py NLL of observed outcome per row; weighted CE, accuracy, ECE
decisions.py contract dataclasses (Option, DecisionRequest, DecisionResult,
Request)
dashboard.py stdlib http.server + JSON API (tree/trace/dream/status +
POST submit/relabel); static/ frontend served at /
Run / verify
Dev machine is a thin client (no GPU, ~1.4G disk): only pure stdlib unit tests
run here (python3 -m pytest tests/ -q --ignore=tests/integration).
Git / sync
- Canonical repo lives on Gitea:
git.manyworlds.fit, SSH on port 222 (ssh://git@git.manyworlds.fit:222/gabby/semif-agent.git). Key~/.ssh/id_ed25519is registered there. The boxguppykeeps a working copy at~/semif-agent; the dev machine at~/Repos/semif-agent. Push/pull from Gitea — never rsync/tar the code. config.jsonis gitignored and per-machine (dev and the box use different engine/LLM paths). Copyconfig.example.jsontoconfig.jsonand edit.data/decisions.jsonl,data/runs.jsonl, anddata/drafts/are runtime artifacts and gitignored too.- The Gitea instance was rebuilt fresh (Sep 2026) after its git pack transfer
broke (
sh: bad option '--oneshot'— a stray systemuploadpack.packObjectsHookkilled pack generation; the old AGENTS.md note blamingauthorized_keyswas a misdiagnosis). Both the dev key (shitass@nunya) and the box key (guppy@semif-agent) are re-registered; clone/fetch/push all work now. If a fresh machine can't pull, the durable fix is on the Gitea host:git config --system --unset-all uploadpack.packObjectsHook. - Decision rows logged before the
run_idthreading landed show up under run_id"?"in the dashboard — that's expected, not a bug.
The AMD box guppy (abby@192.168.8.181) is the real run target. Key facts:
- ssh key
~/.ssh/id_ed25519is passphrase-protected. Load it into an agent at a fixed socket before connecting (the default flatpakSSH_AUTH_SOCKrefuses):The agent dies if this machine restarts; redo it each session.SOCK=/tmp/opencode/ssh-agent.sock; rm -f "$SOCK"; eval $(ssh-agent -a "$SOCK") printf '#!/bin/sh\necho "<PASSPHRASE>"\n' > /tmp/opencode/askpass.sh; chmod 700 /tmp/opencode/askpass.sh SSH_ASKPASS=/tmp/opencode/askpass.sh SSH_ASKPASS_REQUIRE=force setsid -w ssh-add ~/.ssh/id_ed25519 semif_agentis editable-installed into the box venv (pip install -e ~/semif-agent --no-deps), sopython -m semif_agent.cli ...works from any directory on the box, not just the repo root.- Run the agent on the box:
cd ~/semif-agent && export HF_HOME=/home/abby/hf ~/semif-venv/bin/python -m semif_agent.cli run # REPL ~/semif-venv/bin/python -m semif_agent.cli run --script demo.jsonl ~/semif-venv/bin/python -m semif_agent.cli dream # cost report ~/semif-venv/bin/python -m semif_agent.cli relabel <id> <outcome> ~/semif-venv/bin/python -m semif_agent.cli dashboard --port 8765 - Dashboard: runs as a systemd user service on the box
(
semif-dashboard.service, linger enabled, binds0.0.0.0:8765), so it's up after reboots with no manual launch — browser UI at http://192.168.8.181:8765/. Manage it withsystemctl --user status/restart semif-dashboard.service. Note the submit/relabel POST endpoints are therefore open to the whole LAN. It works in live mode on the box (submit runs the real engine + LLM) and in replay mode anywhere (--replay; reads decisions.jsonl + runs.jsonl; submit degrades to a JSON error without the engine). Relabeling in the UI writes a human override (3x weight in dream) viaPOST /api/relabel. On the dev machine, replay mode:cd ~/Repos/semif-agent && python3 -m semif_agent.cli dashboard --replay(binds 127.0.0.1:8765). - Integration tests (real engine + real LLM) only run on the box:
~/semif-venv/bin/python -m pytest tests/integration -q -sThey take ~100s (model load ~34s). Run them in the background and poll — long-lived ssh sessions get SIGHUP'd and kill the run.
Roadmap
v1 (done)
Core loop, urgency queue, skill tree, skill loop with real SemIf + real LLM
self-assessment, decision logging, dream cost pass, REPL + JSONL CLI,
unit tests (24) + box integration tests (2).
v2
- Real fine-tuning from
decisions.jsonlat a regular interval ("dreaming"): accumulate labeled rows, compute cost, fine-tune the decision model, CI/CD validate (accuracy/ECE on a held-out slice, prompt-hash regression), swap the pinned model revision. GPU offload: train on a beefier GPU; the running agent keeps a frozen inference revision until a swap validates. create_skillbranch: live, mirroringcreate_category. The decision model, driven in normal generation mode viaSemIfEngine.generate, proposes a specific skill title + description for the chosen category; the stub is persisted todata/categories.json(under that category'sskillslist) and merged into the running tree as a leaf. Since Sep 2026 the leaf also gets a real runnable body: a larger OpenAI-compatible model (codegen, defaultqwen38-iq3s) writespredict/actcode againstSKILL.md, persisted todata/skills/and hot-loaded, then the request re-dispatches to the new skill. Authoring is still a single pass — validating/reusing written bodies across runs is future work.- Queue persistence (durable across restarts).
- Event/timer intake sources beyond typed input.
- Concurrency: SemIf shared-state mode (
score_shared/SerialPrefixScorer) for parallel decisions; single execution slot remains for processes. - Dashboard: run-requeue cross-linking (child run references its parent), scheduler sim controls (busy/idle/tau) as a first-class panel.
Later / open questions
- Safety/authority: which inputs may interrupt high-stakes processes; is interrupt a per-skill permission?
- Calibration: SemIf ships per-workload temperature scaling; adopt it before treating probabilities as confidence.
- Enumerate the intake source taxonomy and per-source gating.
Constraints & gotchas (learned the hard way)
Environment
- Dev box: Python 3.13, GTX 780M (Kepler, useless), ~1.4G disk free. Never pip-install heavy deps here.
- guppy box: Python 3.14, AMD RX 6950 XT (gfx1030), 32 cores, 30G RAM,
passwordless sudo. SemIf runs via llama.cpp CPU backend (its llamacpp
backend forces
n_gpu_layers=0), so the GPU is NOT used by the decision engine — it IS used by ollama. - The SemIf tokenizer is fetched from HF (
Qwen/Qwen3.5-4Bat the pinned revision). SetHF_HOME=/home/abby/hfor the tokenizer re-downloads.
SemIf install (box)
- SemIf hard-pins
torch==2.10.0,numpy==2.2.6, etc. The llamacpp path does not need torch (torch is imported lazily insidedirect.score). Install with--no-depsand bring only what's needed:pip install -e ~/semif --no-deps, then numpy 2.3.5, transformers 5.17.0, tokenizers 0.23.2, huggingface-hub, llama-cpp-python 0.3.35. numpy==2.2.6has no cp314 wheel → pip tries a source build that fails withoutpkg-config+python3-dev. Use numpy 2.3.5 (has cp314 wheels).llama-cpp-python==0.3.35builds from source. With all 32 cores it OOM-kills gcc (internal compiler error: Segmentation fault). Limit parallelism:CMAKE_BUILD_PARALLEL_LEVEL=6 MAKEFLAGS=-j6 pip install llama-cpp-python==0.3.35. Do NOT bump the llama-cpp-python version — SemIf calls specific llama.cpp C APIs that change between versions.- The pinned GGUF:
Qwen3.5-4B-Q4_K_M.gguffrom bartowski (2.8G) at~/models/. Load ~34s; score ~0.9s/decision on CPU at 8 threads.
ollama (box)
- Installed at
/home/abby/ollama/bin/ollama(not on PATH), systemd serviceollama.service, ROCm backend withHSA_OVERRIDE_GFX_VERSION=10.3.0and KV cache q4_0 + flash attention. This is expected, not a bug. semif-hermes/semif-hermes-v3are for ANOTHER project (hermes agent) — ignore them; they spew "token repeat limit" errors.- Use
qwen3.5:4bfor self-assessment (works, ~3s).qwen38-iq3s(12G 27B) also works but is huge/slow. - If generation hangs with no log output, restart the service
(
sudo systemctl restart ollama) — the ROCm runner can wedge.
codegen (skill bodies, box)
- Skill bodies are written by a separate OpenAI-compatible model, configured
under
codegenin config.json (default modelqwen38-iq3s, the 12G 27B IQ3_S GGUF — huge/slow). Title + description for new skills still come from the small decision model (engine.generate); only the runnable code body uses codegen. - Do NOT cap
max_tokenson the codegen call. qwen38-iq3s reasons first and a cap truncates the hidden reasoning, leavingcontentempty (finish_reason: length) and the body write fails with "skill body is empty". Unbounded, it runs to completion in ~7 min (~40k chars of reasoning then the code); the client reads onlycontent, so reasoning is filtered automatically. The client default timeout is 1200s — raisecodegen.timeoutin config if a harder prompt needs more. - Bodies are persisted to
data/skills/<category>/<name>.py(gitignored) and loaded back at startup viaimportlib, so skills stay runnable across restarts.SKILL.mdat the repo root is the contract the codegen model is prompted with — change it only with intent, it shapes every generated body. - Trust boundary: generated skill code is executed locally (it is imported
as a module and its
predict/actrun in-process). The box is the intended target; treat the endpoint as trusted. - Flow in
scheduler._create_skill: small model authors title+description → traceskill_writing(dashboard shows title/description + a "writing skill body…" badge) → sync codegen write →materialize_skill→ hot-merge into the tree → bounded re-dispatch so the request runs the new skill. Codegen failure leaves a navigable stub and returns a gracefulcreate_skillresult.
Code principles
- No mocking. The decision engine is always real SemIf; the LLM is always a real endpoint. Pure unit tests touch data-structure math only (queue ordering, dream cost, contract serialization). Engine-dependent behavior is verified by integration tests on the box.
- Engine import is lazy (
engine.py) so the rest of the package stays pure stdlib and testable without SemIf installed. Keep it that way. DecisionLog.appendlabels a row with the selected option by default (self-consistent, near-zero cost). Real labels come fromrelabel(human, weight 3x indream) — failures alone don't produce correct labels.- Navigation decisions ARE logged (
navigate:category,navigate:leafin skills.py) and therefore count toward dream cost. This is intended per the design; don't silently drop them. - Queue ordering: urgency desc, then FIFO (
seq). Recency is stored but is NOT in the sort key (it's anti-correlated with FIFO). Ageing pulls weights toward the max (1.0) so low items catch up; uniform additive boosts do nothing. - CLI subcommands must not crash when the engine is unavailable —
submitcatchesEngineUnavailableand returns("error", ...). - Keep deps stdlib-only in the core; heavy deps live on the box venv.
Testing
python3 -m pytest tests/ -q --ignore=tests/integration— anywhere, fast. Includes the dashboard API tests (tests/test_dashboard_api.py), which spin up the stdlib HTTP server on an ephemeral port with the engine never loaded, andtests/test_codegen.pyfor prompt/parse/validate + body store round-trips.tests/integration/— box only; requires real SemIf + real ollama.- After touching scheduler/skills/codegen/engine, re-run both; the integration tests are the only end-to-end verification.