Files

509 lines
40 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CLAUDE.md
Guidance for Claude Code (and other agents) working in this repository.
## Memory: EchoVault (read this first)
This project uses the **EchoVault** MCP server for persistent, cross-session memory.
Prior sessions store architectural decisions, fixed bugs, and gotchas there. Follow
this protocol every session:
1. **At session start — load context.** Call `memory_context` (project is
auto-detected from cwd) before doing any work. Use `memory_search` for specific
topics (e.g. "detector model", "classic-cv", "false positives").
2. **During work — search before re-deciding.** When the task touches an area that
may have prior context, `memory_search` it first instead of re-deriving decisions.
3. **Before ending a session — save what matters.** Call `memory_save` when you made
a design decision, fixed a bug (include root cause + fix), found a non-obvious
gotcha, or the user corrected/clarified a requirement. Pick the right `category`
(`decision` / `bug` / `pattern` / `learning` / `context`). Do **not** save trivia,
things obvious from the code, or duplicates.
EchoVault is the source of truth for *why* things are the way they are; this file is
the stable, high-level map. When they disagree, trust on-disk code first, then
EchoVault, then this file — and update whichever is stale.
## What this project is
**HVideoTool** is a Windows-first desktop GUI utility that **detects already-applied
censorship** (mosaic, pixelation, blur, black bars) in **images**, and draws outlines
over the detected censored regions. You open a **project** (see below); it runs each
image through a detector, draws the regions, and shows a detailed per-image list of
what it found.
> **Projects (this session).** Work is organized into **projects** — a project is a
> folder holding `project.json` (metadata + per-project settings) · `frames/` (the
> images) · `detections.json` (the detection cache) · `collections/Избранное` (the
> single default "favorites" collection). See `core/project.py`. The per-project settings (detector, model,
> threshold, restore engine) live in `project.json`; the global `settings.json` only
> seeds the **defaults for new projects**. The old "open a bare folder" flow is now
> "Импортировать папку как проект…" (copies images into a new project's `frames/`).
> **Scope is still narrow.** It used to extract frames from video, detect, and play
> back with overlays (a thread-based "project model"). That video *playback* pipeline
> was **removed**; the new "projects" are just an on-disk layout, NOT worker threads or
> playback. The tool remains a synchronous, single-threaded **image inspector**.
Keep this scope sharp:
- Primary job is **detection + overlay/inspection**. A **restoration** ("расцензурить")
step was added later (user-requested): on-demand (single frame **or** whole-project
batch into `restored/`), behind a `Restorer` interface. **Detection is YOLO-only.**
Restoration has **two engine families**: **DeepMosaics** (default — *reconstructs*
mosaic, locates it itself) and **diffusion-inpaint** (SwarmUI — *regenerates* the masked
region; opt-in, see below). The noisy classic-CV detector (+ the `combined` composite)
and the cv2 `inpaint` baseline (filled but didn't reconstruct) were **removed** as
"works poorly". Detection is **multi-model** (ADetailer-style): drop YOLO
weights under `models/yolo/<category>/`, tick which ones are active in the toolbar
"Модели" menu, and a detect runs **all** ticked models and merges results (each tagged
with its category → its own overlay colour). DeepMosaics has two engines: **image** (per-frame) and
**video** (BVDNet, temporal — uses neighbour frames). Its GPL-3.0 network code is
**vendored** under `core/restore/_deepmosaics/` and run in-process (user supplies only
the weights). Because of that vendoring the **whole project is GPL-3.0**. LADA
(BasicVSR++) is a possible future engine, not wired.
- **Diffusion-inpaint is now allowed** (the earlier ban was lifted by the user). It is a
*second* restoration engine (`restorer="diffusion"`), NOT a replacement for DeepMosaics
and NOT the default. It **regenerates** the censored region with an external diffusion
server (SwarmUI) over HTTP — it does *not* reconstruct the original, it draws plausible
new content from a prompt + the YOLO mask. So it's best where DeepMosaics is helpless
(black bars / solid fill), per-frame only (a video sequence flickers), and **needs
detections** (the mask). The diffusion model runs in SwarmUI's process, so this path
adds **no torch dependency** to the app. The backend is abstract (`DiffusionBackend`);
SwarmUI is the first impl — ComfyUI/A1111 could be added later as another backend.
(`xinsir/controlnet-union-sdxl-1.0` is still not used — it's a *conditioning* model, a
poor fit; "diffusion-inpaint" here means a standard SD/SDXL inpaint via SwarmUI.)
- It detects **already-censored** regions, not "content that should be censored"
(i.e. not an NSFW classifier).
- Video is only a one-shot frame-extraction convenience (see below): "Создать из
ролика…" makes a new project and decodes the clip into its `frames/`. Detection and
restoration operate on the project's images.
## Target environment
- **OS:** Windows 11 x64 (primary). Use PowerShell syntax in commands.
- **Python:** 3.11+.
- **GPU:** NVIDIA + CUDA via PyTorch, for the YOLO detector and the DeepMosaics
restorer. CPU fallback works but is slow (esp. DeepMosaics / the temporal BVDNet).
## Tech stack (decided)
| Concern | Choice |
|----------------|---------------------------------------------|
| GUI | PySide6 (Qt 6) — LGPL |
| Image IO | OpenCV (`opencv-python`) + NumPy, unicode-safe via `core/imageio.py` |
| Detector | Ultralytics YOLO, multi-model ensemble (models/yolo/<category>/*.pt) |
Torch/CUDA + Ultralytics enter with the YOLO detector. Keep that dependency optional
(the `yolo` extra in `pyproject.toml` pulls only Ultralytics; torch is installed
separately per the README). Both detection (YOLO) and restoration (DeepMosaics) now
require torch — there's no longer a torch-free detector.
## Architecture (as implemented)
> This reflects the actual code on disk. The GUI is mostly synchronous, but the
> **heavy compute (detection + restoration) runs on a background thread** so the UI
> stays responsive — see `ui/workers.py` and the "Background jobs" bullet (this reverses
> the earlier "no worker threads" rule; `processEvents` can't unfreeze a single multi-
> second `detector.detect()`/DeepMosaics call). Work is organized into **projects**
> (`core/project.py`): a project folder holds `project.json` (metadata + per-project
> settings), `frames/` (the images), `detections.json` (the detection cache, at the
> project root — no longer a sidecar next to the images), and `collections/Избранное`
> (the favorites collection). The only video touch is a one-shot "Создать из ролика…"
> that creates a new project and decodes a clip into its `frames/` via the **ffmpeg
> CLI** (cv2.VideoCapture fallback; NOT PyAV). When code and this file disagree, trust
> the code.
```
hvideotool/
├── __main__.py # entry point + CLI (optional project path, --detector, --model)
├── app.py # QApplication bootstrap; run(config, target=None) — opens/auto-reopens a project
├── config.py # AppConfig + DetectionConfig/OverlayConfig (thresholds live here)
├── settings_store.py # new-project DEFAULTS + last/recent projects to ~/HVideoTool/settings.json
├── ui/
│ ├── main_window.py # the whole UI: toolbar + [file list | image view | detail table]
│ ├── workers.py # Job (QRunnable): runs detect/restore off-thread, results via Qt signals
│ └── image_view.py # renders an image + draws polygon/bbox overlays (QPainter); can highlight one
└── core/
├── imageio.py # unicode-safe imread/imwrite (np.fromfile + imdecode)
├── torch_info.py # probe torch/CUDA (gather/reason/install_hint) for the device badge; no Qt
├── project.py # Project: layout (project.json/frames/detections.json/collections) + per-project settings
├── video/
│ ├── extract.py # extract_frames(): ffmpeg CLI (cv2 fallback) -> JPGs; keyframe/step modes + downscale
│ └── frame.py # Frame dataclass (image BGR, index, pts) — the detector input type
├── restore/ # "un-censor": DeepMosaics (reconstruct) OR diffusion-inpaint (regenerate)
│ ├── base.py # Restorer ABC: restore(image, dets, should_cancel) + restore_sequence (batch/temporal) + .temporal/.needs_detections flags; Cancelled exc
│ ├── factory.py # build_restorer(name, config) -> deepmosaics | deepmosaics_video | diffusion (lada = TODO); restorer_needs_detections(name)
│ ├── deepmosaics.py # DeepMosaicsRestorer (image, per-frame) + DeepMosaicsVideoRestorer (BVDNet, temporal); in-process, load once; uses _deepmosaics/
│ ├── _deepmosaics/ # VENDORED DeepMosaics models/+util/ (GPL-3.0) — added to sys.path at import
│ ├── mask.py # detections_to_mask() — rasterise dets → uint8 mask (dilate/blur) for diffusion inpaint
│ ├── diffusion.py # DiffusionRestorer (needs_detections=True, per-frame) + DiffusionBackend ABC + InpaintParams
│ └── swarmui.py # SwarmUIBackend — HTTP to a SwarmUI server (stdlib urllib, no torch dep); GetNewSession + GenerateText2Image
└── detection/ # YOLO only, multi-model
├── base.py # Detector ABC: detect(frame) -> list[Detection]
├── factory.py # build_detector(config) -> MultiYoloDetector over config.detector_models
├── registry.py # discover_models()/category_of() — scans models/yolo/<category>/*.pt
├── multi.py # MultiYoloDetector — runs several YoloDetectors, concatenates results
├── types.py # Detection (type + score + bbox + polygon + label/category; .display); CensorType enum
├── cache.py # save/load the detection cache (detections.json); key = set of model basenames + conf/imgsz
└── yolo.py # YoloDetector — Ultralytics YOLO-seg; lazy torch/ultralytics; tags dets with a category label
```
### How it works
- **Toolbar layout (`_build_toolbar`).** To avoid a long flat row spilling into Qt's
"⋯" overflow, related actions are **grouped into `QToolButton` dropdowns** (two helpers:
`_dropdown_button` = InstantPopup menu-only; `_split_button` = MenuButtonPopup, click runs
the primary action, the arrow opens related ones). Top-level groups: **"Проект ▾"**
(создать/открыть/из ролика/импортировать) · **"Детекторы: Модели (N) ▾"** (the model
picker, unchanged) · split **"Рассчитать кадр ▾"** (menu: детектировать все дозапуск/
заново) · split **"Расцензурить кадр ▾"** (menu: расцензурить все/найденное/заново,
движок…, открыть папку результатов, «Сохранить результат…» = `save_restored_action`) ·
**checkable "Показать расцензуренное"** (`toggle_restored_action`, key `R`, kept visible so
the original⇄restored state shows as a pressed button) · **"■ Стоп"** (kept visible,
reachable instantly) · a stretch spacer pushes **"Порог:"** to the right edge. The full
action list also lives in the **menu bar "Файл"** (`_build_menu`). When adding an action,
put it in the matching dropdown — don't add another flat top-level button.
- `MainWindow` holds the config, builds the detector lazily via `build_detector`
(cached by the **selected-model set** + conf/imgsz in `_make_detector`), and keeps
`_results: dict[path -> list[Detection]]` as the detection cache.
- **Multi-model selection.** `config.detector_models` is the list of ticked YOLO weights
(paths under `models/yolo/<category>/`). The toolbar "Модели" `QToolButton`/`QMenu`
(`_rebuild_models_menu`, items grouped by category via `addSection`) toggles them
(`_on_model_toggled` → persist + `_invalidate_results`); "Добавить модель…" copies a
`.pt` into `models/yolo/<category>/`. On project open `_ensure_models` prunes vanished
paths and, if nothing is selected, default-ticks every discovered model. `build_detector`
builds one `YoloDetector` per selected model (tagged `label=category`) wrapped in a
`MultiYoloDetector`. Each `Detection` carries `label` (category) **and `model`** (the
producing `.pt` stem, tagged in `YoloDetector` — shown as its own "Модель" column in the
detail table since a category folder may hold several models); overlay colour + table
group by `Detection.display` (label, else the CensorType) via `OverlayConfig.colors` + a
stable `palette` fallback.
- **Cross-model NMS (optional).** By default `MultiYoloDetector` just **concatenates** all
models' detections (different categories are meant to coexist). The "Модели" menu has a
checkable **"Объединять пересечения (NMS)"** (`config.cross_model_nms` + `nms_iou`,
`_on_nms_toggled`): when on, `MultiYoloDetector(nms_iou=…)` runs a greedy category-agnostic
IoU NMS (`multi._nms`/`_iou`) that drops the lower-score box of any overlapping pair — kills
the duplicate rects you get when overlapping models fire (e.g. penis + cockAndBall). It
changes the detection result, so it's part of the in-memory detector identity
(`_make_detector` key) **and** the on-disk cache key — but `cache.make_key` adds `nms_iou`
**only when NMS is on**, so the default (off) key is unchanged and an existing cache stays
valid; turning NMS on yields a distinct key (recompute) without clobbering the non-NMS
cache. Toggling also `_invalidate_results` (drops the in-memory cache). Off = concatenate.
- **Per-model display thresholds (optional).** The toolbar "Порог" spin is the global overlay
threshold; the "Модели" menu **"Пороги по моделям…"** (`_edit_model_thresholds`) stores
per-model overrides in `config.model_thresholds` (keyed by `.pt` **stem**). `ImageView`
(`set_model_thresholds`/`_eff_threshold`) draws a detection only if its score clears its
model's override, else the global threshold. **Display-only** — does not change detection or
the "с цензурой" counts (a frame is a hit if it has *any* detection, threshold-independent).
- **Background jobs (`ui/workers.py`).** Detection and restoration are CPU-heavy and
would freeze the GUI, so they run on a `QThreadPool` thread via `Job` (a `QRunnable`
wrapping `fn(job)`); results return to the GUI through queued Qt signals
(`tick`/`progress`/`done`/`failed`). `MainWindow._start_job(fn, total, on_tick, on_done)`
starts one (only one at a time — `_busy` guards entry points), `_finish_job`/
`_on_job_failed` end it. `_make_detector`/`_make_restorer`, image reads, and
`engine.detect/restore` all run **inside the worker** (`_compute` is the pure
read+detect helper); the `fn` must touch NO Qt widgets — it emits plain data that the
GUI-thread slots (`_apply_detection`, restore `tick`) apply. `_begin_busy` disables
`model_action` for the duration (it'd race the running detector). This is the
deliberate exception to the old single-threaded rule (CPU YOLO/DeepMosaics per-call
latency can't be hidden with `processEvents`). Progress messages get an **ETA** suffix
("· осталось ~Xм Yс") from `_eta_suffix(done, total)` using `_job_start` (set in
`_start_job`, `time.monotonic`) — average-rate estimate, blank at 0 %/100 %.
- **Device badge + CUDA diagnostics.** A clickable status-bar chip (`device_badge`) shows
"⚡ CUDA" (green) or "🖥 CPU" (orange). `_probe_device` runs `core/torch_info.gather()` in
a background `Job` at startup (it imports torch AND shells out to `nvidia-smi`, so it's off
the GUI thread) → `_set_device_badge`. `gather()` collects torch facts (version, built_cuda,
cuda_available, device_name) **and** NVIDIA facts (gpus, driver_version, max cuda_driver).
Clicking (`_show_device_info`) opens a diagnostic **QDialog** (not QMessageBox — its text
wasn't copyable): a read-only monospace `QPlainTextEdit` with `torch_info.analyze(info)`
`{summary, details, steps, command}` — a verdict on *why* it's on CPU (CPU-only `+cpu`
build / no GPU / driver-too-old-for-built-CUDA) and the exact pip fix. Buttons: "Скопировать
команду установки" (`_copy_install_command` → the recommended cu121/cu118 command, picked by
`recommend_channel` from the driver's CUDA) and "Проверить заново" (re-runs `_probe_device`).
`core/torch_info.py` is pure (no Qt); subprocess uses `CREATE_NO_WINDOW` on Windows.
- **Projects (`core/project.py`).** `MainWindow._project` is the open `Project`; its
frames come from `project.frames_dir` (navigation/cache/tags work off `_files`). Entry
points: "Создать проект…"
(`_create_project`), "Открыть проект…" (`_open_project_dialog`), "Импортировать папку
как проект…" (`_import_folder_as_project` — copies images into a new project's
`frames/`, carries over an old `.hvideotool_detections.json` sidecar if present), and
a "Недавние проекты" submenu. `_open_project(project)` is the core open: it
`apply_to_config`s the project's settings, syncs the toolbar widgets without signal
loops (`_sync_settings_ui`), titles the window, records last/recent, and lists
`frames/`. On startup `app.run` opens the CLI `target` or auto-reopens
`settings_store.last_project()` (`_auto_open_last`).
- **Per-project settings.** Detector/model/threshold/restore engine live in
`project.json` (`Project.settings`, `_SETTING_KEYS`). `_persist_settings()` writes both
the global defaults (for new projects) **and** the open project. Global `settings.json`
is now only defaults + last/recent projects.
- Listing image files (`_IMAGE_EXTS`) from `frames/` uses a progress bar (bulk insert
with `setUpdatesEnabled(False)` + periodic `processEvents`), so a big project doesn't
freeze silently.
- **Viewing and detecting are decoupled on purpose** (so browsing stays instant even
with a slow CPU detector): selecting a file only *shows* it with its cached result
(header reads "не рассчитано" if none). Detection runs on **double-click**, the
"Рассчитать кадр" action (Space → `_recompute_current`, force-recomputes current), or
"Детектировать все" (whole folder, progress bar). Do NOT re-add auto-detect-on-select.
Results cache in `_results`; the file-list row gets a count suffix when computed.
Switching the model clears the cache (`_choose_model``_invalidate_results`).
- **Detection cache (persisted).** `_results` is mirrored to the project's
`detections.json` (`core/detection/cache.py`, at `project.cache_path`; keyed by
**basename** so it survives moving the project). Both `detections.json` and
`project.json` are written **atomically** (sibling `.tmp` + `os.replace`, see
`cache._atomic_write_text` and `Project.save`) — a crash mid-write can't corrupt/truncate
a large cache (29k entries) and lose all detection work. `cache.save_results`/`load_results`
take the cache file and the image `base_dir` (= `project.frames_dir`) separately, since
the cache lives at the project root, not next to the images. It's tagged with
the detector identity (`_results_key` = detector + model + conf/imgsz); on open,
`_load_cached_results` reloads it **only on a key match** (else ignored, not shown as
current). Saved (`_save_results`, skipped when `_results` is empty so it never clobbers
a good cache with nothing) after detect-all (incl. cancel → partial), single
detect/recompute, move-to-collection, and on `closeEvent`.
**"Детектировать все" is incremental** (skips already-cached frames → resumes/top-ups);
**"Все заново"** (`_detect_all(force=True)`) clears the cache first (full regen);
**"Рассчитать кадр"/Space** always recomputes the one current frame. Empty list in the
cache = "checked, clean" (tinted green, no mark); `in _results` distinguishes it from
"not computed".
- **Favorites (curation).** Curation was simplified (user request) to a **single
default collection** — no create/select/browse UI. The "★ В избранное" button (left
pane, under the file list) / "В избранное" menu item / Ctrl+M → `_move_to_favorites`
**moves** (`shutil.move`, not copy) the `ExtendedSelection`-selected frames into
`project.favorites_dir` (`collections/Избранное`, `FAVORITES_DIR` in `core/project.py`,
created lazily on first move), removing them from list/`_files`/cache. `_unique_dest`
avoids clobbering (`foo.jpg``foo (1).jpg`). Use case: flag good frames into a
training/example set while inspecting detections.
- **Restoration ("Расцензурить кадр" / "Расцензурить все").** The **default DeepMosaics**
engine **locates the mosaic itself** — so for it restoration is **decoupled from
detection**: no detector runs, detections are passed as `[]`. (The **diffusion** engine
is the exception — it masks the detections, so the UI feeds it `_results`; see the
diffusion bullet above and `restorer.needs_detections`.) Single-frame: a toolbar
action reads the current frame + runs `self._restorer` (built via `build_restorer`) **on
a background job** (`_restore_current`; `done` stores `_restored[path]`, **auto-saves it
to the project's `restored/`** (so a single restore persists like the batch, not just in
memory), and shows it). **"Показать расцензуренное/оригинал" (`_toggle_restored`, key `R`)
is a GLOBAL view mode** (`_showing_restored`): when on, `_show` displays each frame's
restored version if one exists — loaded lazily from memory `_restored` **or disk
`restored/`** via `_restored_image_for` (so the whole batch result is browsable, not just
the last frame) — else falls back to the original; overlays are hidden on restored. The
mode persists across navigation (reset to off on project open); a batch/single restore
auto-switches it on. `_has_restored`/`_restored_disk_path` are the cheap (no-decode)
existence checks driving the toggle/save enabled-state. "Сохранить результат" additionally
exports `<stem>_restored.jpg` beside the frame (an explicit one-off export, via
`_restored_image_for`). "Открыть папку результатов" opens `restored/` in Explorer.
**Batch ("Расцензурить все" / "Все заново" /
"Расцензурить найденное", `_restore_all(force, only_detected)`)** mirrors `_detect_all`:
a single background job restores frames and writes results to the project's
**`restored/`** dir (`Project.restored_dir`, basename-mirrored, kept OUT of `frames/` so
outputs aren't re-listed/re-restored); the per-frame engine **skips frames already in
`restored/`** unless `force` (resume). **`only_detected`** ("Расцензурить найденное")
uses the YOLO detection cache to skip frames known clean: per-frame restores only the
frames with detections; the temporal engine restricts the run to the contiguous span
`[first hit … last hit]` (recurrence needs continuity). It's a separate, *faster* action
— NOT the default — because LADA misses some mosaic (esp. anime), so it can miss
censorship YOLO didn't flag; "Расцензурить все" stays the thorough option. Returns early
(status hint) if detection isn't computed or no frame has a detection. The
engines are **DeepMosaics** (`restore/deepmosaics.py`), run **in-process** from the
vendored `_deepmosaics/` code, loading the BiSeNet locator + generator **once** (lazy,
cached on the instance):
- `deepmosaics` (image, per-frame): reproduces `cleanmosaic_img_server` (locate mosaic
→ run generator on the crop → feather back), ~0.3 s/frame cached on CPU. Image model
`clean_youknow_resnet_9blocks.pth`.
- `deepmosaics_video` (**temporal, BVDNet**): `DeepMosaicsVideoRestorer`, `.temporal=True`.
Reproduces `cleanmosaic_video_fusion` — per target frame it feeds the net a window of
`T=5` neighbour frames sampled at step `S=3` around it (`N=2` each side, clamped at the
sequence edges) **plus its own previous output (recurrent)**, for temporal coherence.
Because of that recurrence it must run a **contiguous, ordered range** — it implements
`restore_sequence(count, get_frame, get_dets, emit, should_cancel)` (the batch run uses
it; single-frame `restore` degrades to a window of the same frame). Needs the **video**
weights `clean_youknow_video.pth` (+ `mosaic_position.pth` beside). `INPUT_SIZE=256`.
**`feed_restored`** (config `dm_feed_restored`, default on, checkbox in `RestoreDialog`):
when set, the **past** neighbours in the window (`j<i`) are taken from the engine's own
already-restored outputs (a small rolling cache, `reach=N*S` deep) instead of the original
mosaic frames — stronger temporal coherence. The **centre** (the frame being cleaned) and
**future** neighbours (`j>i`, not yet restored) stay original. Slightly out-of-distribution
for the BVDNet (trained on mosaic windows), so it's a toggle; off = faithful DeepMosaics.
`restore_sequence` is on the `Restorer` ABC (default = independent per-frame loop);
`_restore_all` dispatches on `restorer.temporal` (temporal → `restore_sequence` over the
whole range; per-frame → resumable loop with skip-existing). `should_cancel`
(= `lambda: job.cancelled`) is polled so "■ Стоп" stops it; engines raise `Cancelled`,
which `Job.run` reports as a clean cancel. Engine + weights are set in `RestoreDialog`
(Файл → Движок восстановления…) — the model dropdown shows image vs video weights per
selected engine — persisted, and built lazily/cached in `_make_restorer`. NOTE:
DeepMosaics locates mosaics itself (its `mosaic_position.pth`) — our detections aren't
passed to it. To add another engine (e.g. LADA), implement `core/restore/base.Restorer`
(set `.temporal` + override `restore_sequence` if it needs neighbours) and register it in
`restore/factory.build_restorer`.
**Diffusion-inpaint engine (`restorer="diffusion"`, `core/restore/diffusion.py`).** A
second engine *family* that **regenerates** the censored region instead of reconstructing
it. `DiffusionRestorer` (`needs_detections=True`, per-frame): builds an inpaint mask from
the frame's YOLO detections (`mask.detections_to_mask`, with `diff_mask_dilate`/
`diff_mask_blur`) and hands `(image, mask, InpaintParams)` to a pluggable
`DiffusionBackend`. First backend is `SwarmUIBackend` (`swarmui.py`): stdlib-`urllib`
HTTP to a running SwarmUI server (`GetNewSession``GenerateText2Image` with base64
init+mask images, `diff_prompt`/`diff_negative`/`diff_steps`/`diff_cfg`/`diff_denoise`/
`diff_seed`/`diff_model`) → decode the returned image. The diffusion model runs in
SwarmUI's process, so **no torch dep is added here**. Because it needs a mask, the UI
feeds it real detections (`_restore_current` captures `_results[path]`; `_restore_all`
builds `dets_by_index` for the hit frames) and gates it like "Расцензурить найденное"
(requires detection computed + at least one hit) via `restorer_needs_detections`. Frames
with no detection come back unchanged. Per-frame only → flickers on video; best for
black bars / solid fill where DeepMosaics can't help. Config fields `diff_*` persist in
`project.json` + `settings.json`; engine chosen in `RestoreDialog` (its diffusion field
group shows when the engine is selected). The dialog has a **"Проверить соединение"**
button (`RestoreDialog._test_connection``SwarmUIBackend.ping()`, a fresh
`GetNewSession` with a short 15s timeout) that reports ✓/✗ inline — lets the user verify
SwarmUI is reachable without running a restore.
- **Navigation bar** under the image (`_build_nav_bar`): prev/next frame (◀ ▶, keys
`,`/`.`), a scrubber `frame_slider` across the whole sequence, a clickable `pos_label`
(a flat `QPushButton` "row / n" → `_jump_to_frame`, a "go to frame N" `QInputDialog`
needed on 29k-frame projects where the scrubber is ~50 frames/px), and jump-to-detection
(◀ детекция / детекция ▶, keys `[`/`]`, `_step_hit` scans `_results` for the next
non-empty frame). The slider and file list are kept in sync via `_update_nav` guarded by
`_nav_sync` (avoids signal loops); all navigation ultimately drives
`file_list.setCurrentRow`. `_step` (◀ ▶) **skips rows hidden by the filter**.
- **File-list filter (`filter_combo`, left pane).** A combobox above the list shows only a
subset of the (possibly huge) frame list: Все · С цензурой · Чистые · Не рассчитано ·
Расцензуренные · Без расцензуривания. `_filter_mode` + `_row_matches_filter(path)` (over
`_results`/`_row_restored`); `_on_filter_changed` re-labels (each `_relabel_row` calls
`item.setHidden(...)`) and jumps off a now-hidden current row. Pure show/hide — doesn't
touch `_files`/cache. The scrubber is a custom
`MarkerSlider` (`ui/marker_slider.py`) that paints **two mark layers**: cyan ticks
(upper half) at frames with detections (`_refresh_marks` projects `_results`) and
**green ticks (lower half) at restored frames** (`_refresh_restored_marks` scans
`restored/` + in-memory `_restored`; called on load and after each restore op, not
per-frame); per-pixel deduped so big folders stay cheap. Under the scrubber a
**progress summary** `stats_label` reads "Кадров: N · детектировано: D/N (с цензурой:
H) · расцензурено: R/N" (`_update_counts_label`, cheap counts; `_restored_count`
cached by `_refresh_restored_marks`). File-list rows are labelled too via a single
`_relabel_row` (used by `_tag_file`/`_relabel_all`): tint red = censorship found, green =
checked & clean; a trailing **✓** marks frames with a restored version (`_row_restored`,
populated by `_refresh_restored_marks`). The ✓ is independent of detection — it survives
`_clear_results`. Detection tints reset on `_invalidate_results`.
- **Cancellation (cooperative).** A single "■ Стоп" toolbar action (Esc) cancels the
running op. `_begin_busy(total)` / `_end_busy()` toggle `self._busy` + the Stop button +
the progress bar (`total=None` → indeterminate). For **background jobs** (detection,
restore) `_request_cancel` calls `self._job.cancel()`; the worker loop checks
`job.cancelled` between frames and restore polls it via `should_cancel`. The still-
synchronous loops (`_load_folder` listing, import copy, video extraction — its
`progress` cb returns `not self._cancel`) check `self._cancel` between `processEvents`
ticks. Entry points guard with `if self._busy: return` (notably `_move_to_favorites`,
which mutates `_files` that a detect-all job reads — so a snapshot/pending list is used).
`closeEvent` cancels a running job and `waitForDone(3000)` before tearing down.
- `image_view.ImageView` draws the image scaled-to-fit plus overlays. Overlay
visibility/threshold are applied at paint time. Selecting a row in the detail table
calls `set_highlight(i)` — that detection is drawn boldly (even below threshold) and
the rest dim. The detail table lists ALL detections (sorted by score), so sub-threshold
hits are still visible for debugging; the threshold only affects what's drawn.
### Separation of concerns
- `ui/` must not import `torch` / `ultralytics` directly. It builds detectors only via
`core/detection/factory.build_detector` and talks to `core/` through the `Detector`
interface and the `Detection`/`CensorType` types.
- Detection is YOLO-only (multi-model); restoration is DeepMosaics (default) **or**
diffusion-inpaint (SwarmUI). New detection kinds plug in by adding more `.pt` under
`models/yolo/<category>/` — no code change. The toolbar shows a "Модели" menu of checkable
models (no detector dropdown); restoration engines are chosen in `RestoreDialog`. A new
restoration *engine kind* = implement `core/restore/base.Restorer` and register it in
`restore/factory.build_restorer`; a new *diffusion backend* = implement
`core/restore/diffusion.DiffusionBackend` (keep it out-of-process — no torch dep in the
app). If you re-add a different detector kind, implement `core/detection/base.Detector`
and register it in its `factory`. (No legacy-settings migration is kept while in active development —
old `settings.json`/`project.json` keys are simply ignored, not coerced.)
## Commands
```powershell
python -m venv .venv; .\.venv\Scripts\Activate.ps1
pip install -e ".[yolo]" # YOLO needs ultralytics; install torch separately (README)
python -m hvideotool # reopen the last project (or create/open one in-app)
python -m hvideotool "C:\path\to\MyProject" --model models\lada_mosaic_detection_model_v4_accurate.pt
```
No formal test suite, but there's a **headless smoke test**: `scripts/smoke_test.py`
(run with `QT_QPA_PLATFORM=offscreen` + `PYTHONIOENCODING=utf-8`) covers the pure core
(atomic cache round-trip, cross-model NMS, list-filter predicate, ETA formatting,
extract-dialog options, per-project settings round-trip) and an offscreen `MainWindow`
build on a throwaway project — no torch/weights (detections are injected into `_results`).
Exits non-zero on failure; run it after touching core/UI plumbing. Ad-hoc check: build a
`MainWindow`, `Project.create(tmp)` + copy a few images into `frames/`,
`_open_project(project)`, drive `file_list.setCurrentRow(...)`, and read
`detail_table` / `detail_header`. Or run `build_detector(config).detect(...)` on a
frame directly.
## Conventions
- Match the style of surrounding code; keep `core/` free of Qt where reasonable.
- Type hints on public functions and the `Detector` interface.
- Model weights (`.pt`) and large media are **not** committed — keep them in `models/`
and `.gitignore`d.
- User-facing strings / README are in Russian; code identifiers and this file in English.
## Gotchas
- **Models live under `models/yolo/<category>/`.** Discovery (`detection/registry.py`)
scans that tree; the **category folder is the detection label + overlay colour** (e.g.
`models/yolo/mosaic/lada.pt` → "mosaic", `models/yolo/face/…` → "face"). On open
`_ensure_models` default-ticks all discovered models if the project has no selection.
Put a **censorship** model in `mosaic/` (LADA `lada_mosaic_detection_model_v4_accurate.pt`);
a generic COCO model (e.g. `yolo11n-seg.pt`) would detect people/objects → noise.
Selecting many models multiplies per-frame time (each runs in turn).
- **classic-CV / inpaint were removed (worked poorly).** The classic-CV detector was
noisy/approximate on real footage (false positives on skin/hair/fabric/JPEG; missed
real mosaic after downscale) and the `combined` mode + cv2 `inpaint` baseline went with
it. Detection is YOLO-only; restoration is DeepMosaics (default) or diffusion-inpaint.
(The removed cv2 `inpaint` was a *classic* fill; the new diffusion-inpaint is a different
thing — a real generative SD/SDXL inpaint via SwarmUI.) No backward-compat shims
while in active development — stale keys in old `settings.json`/`project.json` are just
ignored (a project with no valid model selection default-ticks all discovered models).
- **Domain matters.** LADA is trained on REAL video (JAV). It detects some anime mosaic
but not all. The real anime fix is *retraining* a YOLO11-seg (see `scripts/training/`).
- **YOLO detector = LADA weights** ([HF `ladaapp/lada`](https://huggingface.co/ladaapp/lada)).
YOLO **segmentation** model, classes `{0: mosaic_nsfw, 1: mosaic_sfw_head}` → both map
to `CensorType.MOSAIC` (`_name_to_type` matches "mosaic" in the class name). Detects
mosaic only. Weights + Ultralytics are AGPL-3.0 (accepted). `yolo.py` lazy-imports
`torch`/`ultralytics`.
- **No model weights in the repo.** Code must fail with a clear, actionable message
when the model path is missing — not a raw stack trace (`factory._require_model`,
`YoloDetector.__init__`).
- **CUDA/torch install is environment-specific.** Don't add torch to core deps; it
stays out (the `yolo` extra pulls only Ultralytics) and is installed separately.
- **CPU-only torch must not request CUDA.** A `+cpu` torch build raises "Torch not
compiled with CUDA enabled" the moment something calls `.cuda()`. Both engines guard
for this: `YoloDetector` picks `cuda` only when `torch.cuda.is_available()` (even an
explicit `yolo_device="cuda"` is downgraded to cpu); both `DeepMosaicsRestorer._ensure_loaded`
and `DeepMosaicsVideoRestorer._ensure_loaded` force `gpu_id="-1"` when CUDA is absent (the
vendored `model_util.todevice` / `data.im2tensor`/`to_tensor` call `.cuda()` for any
`gpu_id != "-1"`, e.g. the `dm_gpu="0"` default). So a wrong/CPU-only torch falls back to
CPU instead of crashing (the temporal BVDNet engine is heavy on CPU, though).
- **QImage from a numpy buffer must be `.copy()`d** (see `ImageView.set_image`),
otherwise it aliases a buffer that gets freed → garbage/crash.
- **Always use `core/imageio.py`** (`imread_unicode`/`imwrite_unicode`) for images —
`cv2.imread`/`imwrite` silently fail on non-ASCII Windows paths.
- **Diffusion-inpaint runs out-of-process (SwarmUI), so `ui/` and the diffusion path add
no torch/diffusers dependency** — `swarmui.py` uses only stdlib `urllib`. Keep it that
way: the diffusion model lives in the SwarmUI server, we just POST image+mask+prompt.
Don't add `diffusers`/in-process SD to the app. The diffusion engine **needs detections**
(it masks them) — the UI feeds them via `dets_by_index` / captured `_results` and gates
it like "Расцензурить найденное" (`restorer_needs_detections` + `Restorer.needs_detections`);
DeepMosaics still gets `[]` (it self-locates). A new diffusion backend = another
`DiffusionBackend` impl, not new app deps.
- Don't reintroduce the removed video *playback pipeline* (PyAV, producer/consumer worker
threads, player, project session). (Generative/diffusion inpaint via an external server
IS now allowed — see the diffusion engine; the old blanket "no generative" ban is lifted.)
(The new `core/project.py` is an on-disk layout, not that thread-based "project
model".) NOTE: a **single** background `Job` thread for detect/restore (`ui/workers.py`)
IS in scope now (keeps the GUI responsive) — that's different from the rejected
multi-thread video pipeline. The one allowed video touch is
`core/video/extract.py` (one-shot decode → a new project's `frames/`, behind "Создать
из ролика…"): ffmpeg CLI — `_find_ffmpeg()` prefers PATH, else the binary bundled by
the `imageio-ffmpeg` dep, else cv2 fallback. Keyframe-only `-skip_frame nokey` is
~10× faster than every-frame; `-hwaccel` does NOT help (GPU transfer overhead). Use
ffmpeg/cv2, not PyAV, and keep it synchronous. Decoding every frame is the inherent
cost — the speed lever is decoding *fewer* frames (keyframes). `ExtractDialog.options()`
returns `(keyframes_only, step, max_dim, jpg_quality)`; **jpg_quality** (1100, default 92,
via `-q:v` `_quality_to_qscale` / cv2 `IMWRITE_JPEG_QUALITY`) trades quality for a bit of
encode speed + smaller files. Default sampling is **every frame** (step=1, not keyframes).