Split the extension into partials, persist chats on the data volume, park/warm the chat model around Generate, and drop dual raw/persona dump paths. Co-authored-by: Cursor <cursoragent@cursor.com>
148 lines
8.1 KiB
Markdown
148 lines
8.1 KiB
Markdown
# Swarm Assistent
|
||
|
||
SwarmUI extension for **collaborative Krea 2** prompting via **Ollama**: chat + board (Generate | Refs tabs), LoRA chips, **persona presets** (`Config/personas/`), **vector memory**, model cards with Civitai fetch, img2img/inpaint, slash commands, auto Generate.
|
||
|
||
**Version 0.8.2** — Hybrid memory (FTS5 + cosine, kind quotas, min_score) plus model hops `memory_get` / `memory_search` / `lookup_tags`. Danbooru csv is a shared FTS catalog (no embeddings); Krea prompts stay prose.
|
||
|
||
## Layout
|
||
|
||
- **Left — Board tabs:** **Generate** (full-height live view) | **Refs** (reference grid + badge `N · vision M`); **Посмотри результат** attaches the finished frame and asks for a verdict
|
||
- **Splitter:** drag to resize panes
|
||
- **Right:** Chat | Cards; persona / pack / Ollama chat model; **Ollama health** badge; settings gear (memory model + skills + **Память**)
|
||
- **Chips / slash:** loaded from `Config/_base/ui.json` (persona can override)
|
||
|
||
## Config (bundled + overlay)
|
||
|
||
```
|
||
Config/
|
||
_base/ # defaults (assistant, ui, models/krea2, exact.json, core, packs, skills, memory-seed, identity)
|
||
personas/<id>/ # sparse preset: persona/voice/likes/dislikes/rules + optional exact.json / memory-seed / overrides
|
||
```
|
||
|
||
Disk overlay (wins over bundled): `/mnt/swarm_data/Assistent/` — same folder layout as `Config/`, i.e. drop `_base/…` and `personas/<id>/…` files to override any bundled preset. Plus this extension's own state:
|
||
|
||
```
|
||
Assistent/
|
||
_base/ personas/<id>/ # overlay presets — same names as Config/, sparse
|
||
settings.json # embed_model, base_url, per-persona skills
|
||
ui-state.json # pack / persona / auto_* / pane_width / models — seeds a fresh browser
|
||
taste.json # learned taste profile (wins over localStorage)
|
||
chats/<id>.json # chat history, newest 60 kept
|
||
ollama-roles.json # chat vs memory model tags
|
||
memory/assistent.sqlite # vector store (shared + per-persona)
|
||
```
|
||
|
||
Copy `personas/cinema/` → `noir/`, edit only differing JSON files. Persona prompt overrides belong in `personas/<id>/` — the old flat `personas.json` is legacy and only read when no overlay folder exists for that id.
|
||
|
||
## Exact memory (KV)
|
||
|
||
- `Config/_base/exact.json` — canonical generation defaults, profiles (turbo/raw), aspect table, short facts
|
||
- Persona / disk overlays merge via DeepMerge (matching keys overwrite)
|
||
- Always injected into the system prompt; UI fills **empty** SwarmUI fields from Exact (no LLM call)
|
||
- Chat-session overrides (`session_exact`) last until persona change or clear chat — not written to disk
|
||
- Priority: core contract → current user → session_exact → exact (+ persona) → live fields → vector `memory_hits`
|
||
|
||
## Vector memory
|
||
|
||
Two layers in `memory/assistent.sqlite` (`persona` column; empty = shared):
|
||
|
||
- **Shared** — `Config/_base/memory-seed/`, model cards, `scope: "shared"` upserts. Visible to every persona.
|
||
- **Personal** — `Config/personas/<id>/memory-seed/` and chat upserts (default). Never copied into shared. Other personas do not retrieve it.
|
||
- Retrieve = shared ∪ this persona (and `extends` parents). Hybrid **FTS5 + cosine**, kind quotas (e.g. 3 cards / 3 pitfalls / 4 notes), `min_score`. Same `kind`+`key`: personal overwrites parent overwrites shared.
|
||
- Tools: `memory_get`, `memory_search`, `lookup_tags` (Danbooru csv in `Data/Autocompletions`, FTS, **no embeddings**).
|
||
- SQLite + Ollama `/api/embed` (default `nomic-embed-text`, pick in ⚙)
|
||
- Soft notes only — Exact and the user beat RAG for params
|
||
- ⚙ → **Память** lists every row (scope · source · date) with a per-row forget; bundled rows are read-only because reseed brings them back
|
||
|
||
## Chats on disk
|
||
|
||
- Every chat is written to `Assistent/chats/<id>.json` (messages + a Generate params snapshot), so History survives a cleared browser and follows the data volume across VMs
|
||
- localStorage stays as a fast cache; on first run with an empty `chats/` the old `swarm_assistent_chats_v1` store is migrated up once
|
||
- `ui-state.json` seeds a **fresh** browser only — anything already in localStorage wins, and `auto_download` is never restored as on
|
||
|
||
## VRAM handover
|
||
|
||
- Before every Generate the chat model is unloaded (`keep_alive: 0`) so Krea 2 gets the whole GPU
|
||
- Back in the Chat tab it is warmed again with a 1-token request (`keep_alive 15m`, `num_ctx` from `assistant.json`)
|
||
- Embed / memory models are never parked — reloading them would stall every retrieve
|
||
|
||
## UX
|
||
|
||
- **Send to Assistent** under Generate/History → Ref + Assistent tab
|
||
- Enter sends; Shift+Enter newline; Interrupt cancels chat epoch
|
||
- Manual **Apply + Generate** / `/gen` always generate; Auto-generate checkbox only for LLM auto-path
|
||
- **Посмотри результат** / auto-critique wait for a real Generate frame — model previews and unfinished batches are skipped
|
||
- Civitai Confirm required (unless auto-download); queued-but-missing models show a `⏳ wanted` badge in Cards
|
||
|
||
### Slash commands (client-side, no LLM)
|
||
|
||
| Command | Effect |
|
||
| --- | --- |
|
||
| `/help` | List commands |
|
||
| `/debug` | Short UI/Exact dump (no LLM) |
|
||
| `/debug ask` / `/why` | Dump + short model explanation |
|
||
| `/gen` | Generate now |
|
||
| `/look generate\|refN` | Attach that board window + ask the LLM to look |
|
||
| `/init` `/mask` `/clear` | Same as board buttons |
|
||
| `/interrupt` | Stop generation / cancel chat |
|
||
| `/aspect 16:9` | Set size from the official 1K table |
|
||
| `/seed lock\|random` | Lock or randomize seed |
|
||
| `/vary` | New seed, same prompt (+ generate if auto) |
|
||
| `/pack write\|critique\|…` | Switch pack |
|
||
| `/civitai <query>` | Ask LLM to search Civitai |
|
||
| `/inventory` | Rescan models + refresh LoRA list |
|
||
|
||
## Requirements
|
||
|
||
- SwarmUI with a **Krea 2** checkpoint selected
|
||
- Ollama on `http://127.0.0.1:11434` **on the GPU VM** (gpu-rent `LLM_RUNTIME=ollama`)
|
||
- Chat model + memory embed (`use: memory` in `ollama-models.yaml`; gpu-rent creates CPU variant)
|
||
- Optional: Civitai API key in SwarmUI User Settings
|
||
|
||
## Install
|
||
|
||
```yaml
|
||
swarmui:
|
||
- url: https://gitea.hsrv.site/mrleo1nid/swarm-assistent.git
|
||
ref: main
|
||
dir: swarm-assistent
|
||
requires: ollama
|
||
```
|
||
|
||
Restart / rebuild SwarmUI after clone. gpu-rent: `seed-extensions` + restart.
|
||
|
||
## Packs & skills
|
||
|
||
**Packs** (one active): `write_prompt`, `critique_image`, `compose_scene`, `fix_params`, `inpaint_edit`, `describe_ref`, `catalog_card`.
|
||
|
||
**Skills** (checkboxes): `prompting`, `creativity_sliders`, `memory` — procedures; encyclopedia numbers live in Exact, soft notes in memory-seed / RAG.
|
||
|
||
**Personas:** `neutral`, `lewd`, `aggressive`, `cinema`, `terse` under `Config/personas/`.
|
||
|
||
## API routes
|
||
|
||
| Route | Role |
|
||
| --- | --- |
|
||
| `AssistentListModels` | Ollama tags → `models` (chat) + `memory_models` |
|
||
| `AssistentGetConfig` | Merged preset for persona (ui, packs, skills, identity) |
|
||
| `AssistentGetSettings` / `AssistentSaveSettings` | Overlay settings (skills, embed_model) |
|
||
| `AssistentListInventory` | LoRA / checkpoint / wildcard inventory |
|
||
| `AssistentListPersonas` | Persona catalog |
|
||
| `AssistentGetPacks` | Prompt pack texts |
|
||
| `AssistentGetCard` / `AssistentSaveCard` | `.assistent.json` cards (+ memory ingest) |
|
||
| `AssistentGetCardMeta` | Local sidecar + optional Civitai by-hash |
|
||
| `AssistentEnqueueWanted` / `AssistentListWanted` | Wanted YAML queue (write / read + count) |
|
||
| `AssistentGetTaste` / `AssistentSaveTaste` | Persistent taste profile |
|
||
| `AssistentSearchCivitai` | Civitai LoRA search |
|
||
| `AssistentChat` / `AssistentChatWS` | Chat (+ hybrid memory + Civitai/tag hops) |
|
||
| `AssistentListMemory` / `AssistentUpsertMemory` / `AssistentForgetMemory` | Vector store (optional `scope` / `persona`) |
|
||
| `AssistentSearchMemory` / `AssistentGetMemory` | Hybrid search / exact kind+key |
|
||
| `AssistentLookupTags` | Danbooru csv FTS (no embeddings) |
|
||
| `AssistentListChats` / `AssistentGetChat` / `AssistentSaveChat` / `AssistentDeleteChat` | `Assistent/chats/<id>.json` |
|
||
| `AssistentGetUiState` / `AssistentSaveUiState` | `Assistent/ui-state.json` |
|
||
| `AssistentParkLlm` / `AssistentWarmLlm` | Unload / reload the chat model in VRAM |
|
||
|
||
## License
|
||
|
||
MIT
|