Files
HVideoTool/CLAUDE.md
T

341 lines
24 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CLAUDE.md
Guidance for Claude Code (and other agents) working in this repository.
## Memory: EchoVault (read this first)
This project uses the **EchoVault** MCP server for persistent, cross-session memory.
Prior sessions store architectural decisions, fixed bugs, and gotchas there. Follow
this protocol every session:
1. **At session start — load context.** Call `memory_context` (project is
auto-detected from cwd) before doing any work. Use `memory_search` for specific
topics (e.g. "detector model", "classic-cv", "false positives").
2. **During work — search before re-deciding.** When the task touches an area that
may have prior context, `memory_search` it first instead of re-deriving decisions.
3. **Before ending a session — save what matters.** Call `memory_save` when you made
a design decision, fixed a bug (include root cause + fix), found a non-obvious
gotcha, or the user corrected/clarified a requirement. Pick the right `category`
(`decision` / `bug` / `pattern` / `learning` / `context`). Do **not** save trivia,
things obvious from the code, or duplicates.
EchoVault is the source of truth for *why* things are the way they are; this file is
the stable, high-level map. When they disagree, trust on-disk code first, then
EchoVault, then this file — and update whichever is stale.
## What this project is
**HVideoTool** is a Windows-first desktop GUI utility that **detects already-applied
censorship** (mosaic, pixelation, blur, black bars) in **images**, and draws outlines
over the detected censored regions. You open a **project** (see below); it runs each
image through a detector, draws the regions, and shows a detailed per-image list of
what it found.
> **Projects (this session).** Work is organized into **projects** — a project is a
> folder holding `project.json` (metadata + per-project settings) · `frames/` (the
> images) · `detections.json` (the detection cache) · `collections/Избранное` (the
> single default "favorites" collection). See `core/project.py`. The per-project settings (detector, model,
> threshold, restore engine) live in `project.json`; the global `settings.json` only
> seeds the **defaults for new projects**. The old "open a bare folder" flow is now
> "Импортировать папку как проект…" (copies images into a new project's `frames/`).
> **Scope is still narrow.** It used to extract frames from video, detect, and play
> back with overlays (a thread-based "project model"). That video *playback* pipeline
> was **removed**; the new "projects" are just an on-disk layout, NOT worker threads or
> playback. The tool remains a synchronous, single-threaded **image inspector**.
Keep this scope sharp:
- Primary job is **detection + overlay/inspection**. A **restoration** ("расцензурить")
step was added later (user-requested): per-frame, on-demand, behind a `Restorer`
interface. Two engines: a cv2 **inpaint baseline** (fills, does NOT reconstruct) and
**DeepMosaics** — real generative mosaic removal, its GPL-3.0 network code **vendored**
under `core/restore/_deepmosaics/` and run in-process (user supplies only the weights).
Because of that vendoring the **whole project is GPL-3.0**. LADA (BasicVSR++) is a
possible future engine, not wired.
- Still **no diffusion / ControlNet / SDXL**. (`xinsir/controlnet-union-sdxl-1.0` was
rejected early — a generative *conditioning* model, not a censorship restorer. Don't
reintroduce it.) Restoration, if upgraded, uses a mosaic-removal model (DeepMosaics/
LADA), not a general text-to-image diffusion pipeline.
- It detects **already-censored** regions, not "content that should be censored"
(i.e. not an NSFW classifier).
- Video is only a one-shot frame-extraction convenience (see below): "Создать из
ролика…" makes a new project and decodes the clip into its `frames/`. Detection and
restoration operate on the project's images.
## Target environment
- **OS:** Windows 11 x64 (primary). Use PowerShell syntax in commands.
- **Python:** 3.11+.
- **GPU:** NVIDIA + CUDA via PyTorch, only for the YOLO detector. CPU fallback works
but is slow. The `classic` detector needs no torch and no GPU.
## Tech stack (decided)
| Concern | Choice |
|----------------|---------------------------------------------|
| GUI | PySide6 (Qt 6) — LGPL |
| Image IO | OpenCV (`opencv-python`) + NumPy, unicode-safe via `core/imageio.py` |
| Detector | classic-CV heuristic; Ultralytics YOLO (LADA) behind a pluggable interface |
Torch/CUDA + Ultralytics enter only with the YOLO detector. Keep that dependency
optional (the `yolo` extra in `pyproject.toml` pulls only Ultralytics; torch is
installed separately per the README). The classic detector must keep running with no
torch present.
## Architecture (as implemented)
> This reflects the actual code on disk. The GUI is mostly synchronous, but the
> **heavy compute (detection + restoration) runs on a background thread** so the UI
> stays responsive — see `ui/workers.py` and the "Background jobs" bullet (this reverses
> the earlier "no worker threads" rule; `processEvents` can't unfreeze a single multi-
> second `detector.detect()`/DeepMosaics call). Work is organized into **projects**
> (`core/project.py`): a project folder holds `project.json` (metadata + per-project
> settings), `frames/` (the images), `detections.json` (the detection cache, at the
> project root — no longer a sidecar next to the images), and `collections/Избранное`
> (the favorites collection). The only video touch is a one-shot "Создать из ролика…"
> that creates a new project and decodes a clip into its `frames/` via the **ffmpeg
> CLI** (cv2.VideoCapture fallback; NOT PyAV). When code and this file disagree, trust
> the code.
```
hvideotool/
├── __main__.py # entry point + CLI (optional project path, --detector, --model)
├── app.py # QApplication bootstrap; run(config, target=None) — opens/auto-reopens a project
├── config.py # AppConfig + DetectionConfig/OverlayConfig (thresholds live here)
├── settings_store.py # new-project DEFAULTS + last/recent projects to ~/HVideoTool/settings.json
├── ui/
│ ├── main_window.py # the whole UI: toolbar + [file list | image view | detail table]
│ ├── workers.py # Job (QRunnable): runs detect/restore off-thread, results via Qt signals
│ └── image_view.py # renders an image + draws polygon/bbox overlays (QPainter); can highlight one
└── core/
├── imageio.py # unicode-safe imread/imwrite (np.fromfile + imdecode)
├── torch_info.py # probe torch/CUDA (gather/reason/install_hint) for the device badge; no Qt
├── project.py # Project: layout (project.json/frames/detections.json/collections) + per-project settings
├── video/
│ ├── extract.py # extract_frames(): ffmpeg CLI (cv2 fallback) -> JPGs; keyframe/step modes + downscale
│ └── frame.py # Frame dataclass (image BGR, index, pts) — the detector input type
├── restore/ # "un-censor" detected regions (per-frame)
│ ├── base.py # Restorer ABC: restore(image, detections, should_cancel=None) -> image; Cancelled exc
│ ├── factory.py # build_restorer(name, config) -> inpaint | deepmosaics (lada = TODO)
│ ├── inpaint.py # InpaintRestorer (cv2) — baseline, fills not reconstructs
│ ├── deepmosaics.py # DeepMosaicsRestorer — in-process, loads models once; uses _deepmosaics/
│ ├── _deepmosaics/ # VENDORED DeepMosaics models/+util/ (GPL-3.0) — added to sys.path at import
│ └── mask.py # detections_to_mask(shape, dets, dilate)
└── detection/
├── base.py # Detector ABC: detect(frame) -> list[Detection]
├── factory.py # build_detector(config) -> classic | yolo | combined
├── types.py # Detection (+ to_dict/from_dict), CensorType enum
├── cache.py # save/load the project detection cache (detections.json): cache_file + base_dir args
├── classic_cv.py # ClassicCVDetector — heuristic; accepts a `types` filter
├── yolo.py # YoloDetector — Ultralytics YOLO-seg; lazy-imports torch/ultralytics
└── composite.py # CompositeDetector — merge detectors + IoU dedup
```
### How it works
- `MainWindow` holds the config, builds the detector lazily via `build_detector`
(cached by detector+model+conf in `_make_detector`), and keeps `_results: dict[path
-> list[Detection]]` as the detection cache.
- **Background jobs (`ui/workers.py`).** Detection and restoration are CPU-heavy and
would freeze the GUI, so they run on a `QThreadPool` thread via `Job` (a `QRunnable`
wrapping `fn(job)`); results return to the GUI through queued Qt signals
(`tick`/`progress`/`done`/`failed`). `MainWindow._start_job(fn, total, on_tick, on_done)`
starts one (only one at a time — `_busy` guards entry points), `_finish_job`/
`_on_job_failed` end it. `_make_detector`/`_make_restorer`, image reads, and
`engine.detect/restore` all run **inside the worker** (`_compute` is the pure
read+detect helper); the `fn` must touch NO Qt widgets — it emits plain data that the
GUI-thread slots (`_apply_detection`, restore `tick`) apply. `_begin_busy` disables
`detector_combo`/`model_action` for the duration (they'd race the running detector).
This is the deliberate exception to the old single-threaded rule (CPU YOLO/DeepMosaics
per-call latency can't be hidden with `processEvents`).
- **Device badge.** A clickable status-bar chip (`device_badge`) shows "⚡ CUDA" (green)
or "🖥 CPU" (orange). `_probe_device` runs `core/torch_info.gather()` in a background
`Job` at startup (importing torch is slow, so it's off the GUI thread) → `_set_device_badge`.
Clicking (`_show_device_info`) opens a diagnostic dialog: `torch_info.reason()` explains
why CPU (no torch / CPU-only `+cpu` build / built-with-CUDA-but-no-GPU) plus
`install_hint()` (the cu121 pip command). `core/torch_info.py` is pure (no Qt).
- **Projects (`core/project.py`).** `MainWindow._project` is the open `Project`;
`_folder` is kept as a synonym for `project.frames_dir` so the rest of the code
(navigation, cache, tags) didn't need rewiring. Entry points: "Создать проект…"
(`_create_project`), "Открыть проект…" (`_open_project_dialog`), "Импортировать папку
как проект…" (`_import_folder_as_project` — copies images into a new project's
`frames/`, carries over an old `.hvideotool_detections.json` sidecar if present), and
a "Недавние проекты" submenu. `_open_project(project)` is the core open: it
`apply_to_config`s the project's settings, syncs the toolbar widgets without signal
loops (`_sync_settings_ui`), titles the window, records last/recent, and lists
`frames/`. On startup `app.run` opens the CLI `target` or auto-reopens
`settings_store.last_project()` (`_auto_open_last`).
- **Per-project settings.** Detector/model/threshold/restore engine live in
`project.json` (`Project.settings`, `_SETTING_KEYS`). `_persist_settings()` writes both
the global defaults (for new projects) **and** the open project. Global `settings.json`
is now only defaults + last/recent projects.
- Listing image files (`_IMAGE_EXTS`) from `frames/` uses a progress bar (bulk insert
with `setUpdatesEnabled(False)` + periodic `processEvents`), so a big project doesn't
freeze silently.
- **Viewing and detecting are decoupled on purpose** (so browsing stays instant even
with a slow CPU detector): selecting a file only *shows* it with its cached result
(header reads "не рассчитано" if none). Detection runs on **double-click**, the
"Рассчитать кадр" action (Space → `_recompute_current`, force-recomputes current), or
"Детектировать все" (whole folder, progress bar). Do NOT re-add auto-detect-on-select.
Results cache in `_results`; the file-list row gets a count suffix when computed.
Switching detector/model clears the cache (`_invalidate_results`).
- **Detection cache (persisted).** `_results` is mirrored to the project's
`detections.json` (`core/detection/cache.py`, at `project.cache_path`; keyed by
**basename** so it survives moving the project). `cache.save_results`/`load_results`
take the cache file and the image `base_dir` (= `project.frames_dir`) separately, since
the cache lives at the project root, not next to the images. It's tagged with
the detector identity (`_results_key` = detector + model + conf/imgsz); on open,
`_load_cached_results` reloads it **only on a key match** (else ignored, not shown as
current). Saved (`_save_results`, skipped when `_results` is empty so it never clobbers
a good cache with nothing) after detect-all (incl. cancel → partial), single
detect/recompute, move-to-collection, and on `closeEvent`.
**"Детектировать все" is incremental** (skips already-cached frames → resumes/top-ups);
**"Все заново"** (`_detect_all(force=True)`) clears the cache first (full regen);
**"Рассчитать кадр"/Space** always recomputes the one current frame. Empty list in the
cache = "checked, clean" (tinted green, no mark); `in _results` distinguishes it from
"not computed".
- **Favorites (curation).** Curation was simplified (user request) to a **single
default collection** — no create/select/browse UI. The "★ В избранное" button (left
pane, under the file list) / "В избранное" menu item / Ctrl+M → `_move_to_favorites`
**moves** (`shutil.move`, not copy) the `ExtendedSelection`-selected frames into
`project.favorites_dir` (`collections/Избранное`, `FAVORITES_DIR` in `core/project.py`,
created lazily on first move), removing them from list/`_files`/cache. `_unique_dest`
avoids clobbering (`foo.jpg``foo (1).jpg`). Use case: flag good frames into a
training/example set while inspecting detections.
- **Restoration ("Расцензурить кадр").** Toolbar action runs `self._restorer` (built via
`build_restorer`) on the current frame's detections (computing them first if needed),
**on a background job** (`_restore_current` builds an `fn` that detects-if-needed +
restores in the worker; a `tick` caches freshly-computed detections, `done` stores
`_restored[path]` + shows it). "Показать оригинал/
результат" toggles (`_showing_restored`); "Сохранить результат" writes
`<stem>_restored.jpg` beside the frame. The baseline
is cv2 inpaint; the real engine is **DeepMosaics** (`restore/deepmosaics.py`), run
**in-process** from the vendored `_deepmosaics/` code: it loads the BiSeNet locator +
clean generator **once** (lazy, cached on the instance) and per frame reproduces their
`cleanmosaic_img_server` (locate mosaic → run generator on the crop → feather back),
~0.3 s/frame cached on CPU vs ~7 s when it spawned a subprocess. Use the **image**
model `clean_youknow_resnet_9blocks.pth` — the video model (BVDNet) is rejected per
frame (needs a neighbour). `should_cancel` (= `lambda: job.cancelled`) is polled so
"■ Стоп" stops it; the engine raises `Cancelled`, which `Job.run` reports as a clean cancel.
The engine + weights are set in `RestoreDialog` (Файл → Движок восстановления…),
persisted, and built lazily/cached in `_make_restorer` (like `_make_detector`). NOTE:
DeepMosaics locates mosaics itself (its `mosaic_position.pth`, expected beside the
clean weights) — our detections aren't passed to it. `_show` resets `_showing_restored`
+ `_update_restore_actions`. To add another engine (e.g. LADA), implement
`core/restore/base.Restorer` and register it in `restore/factory.build_restorer`.
- **Navigation bar** under the image (`_build_nav_bar`): prev/next frame (◀ ▶, keys
`,`/`.`), a scrubber `frame_slider` across the whole sequence, a `pos_label`
("row / n"), and jump-to-detection (◀ детекция / детекция ▶, keys `[`/`]`,
`_step_hit` scans `_results` for the next non-empty frame). The slider and file list
are kept in sync via `_update_nav` guarded by `_nav_sync` (avoids signal loops); all
navigation ultimately drives `file_list.setCurrentRow`. The scrubber is a custom
`MarkerSlider` (`ui/marker_slider.py`) that paints cyan ticks at frames with
detections (`_refresh_marks` projects `_results` onto row indices; per-pixel deduped
so big folders stay cheap). File-list rows are tinted too (`_tag_file`): red =
censorship found, green = checked & clean. Both reset on `_invalidate_results`.
- **Cancellation (cooperative).** A single "■ Стоп" toolbar action (Esc) cancels the
running op. `_begin_busy(total)` / `_end_busy()` toggle `self._busy` + the Stop button +
the progress bar (`total=None` → indeterminate). For **background jobs** (detection,
restore) `_request_cancel` calls `self._job.cancel()`; the worker loop checks
`job.cancelled` between frames and restore polls it via `should_cancel`. The still-
synchronous loops (`_load_folder` listing, import copy, video extraction — its
`progress` cb returns `not self._cancel`) check `self._cancel` between `processEvents`
ticks. Entry points guard with `if self._busy: return` (notably `_move_to_favorites`,
which mutates `_files` that a detect-all job reads — so a snapshot/pending list is used).
`closeEvent` cancels a running job and `waitForDone(3000)` before tearing down.
- `image_view.ImageView` draws the image scaled-to-fit plus overlays. Overlay
visibility/threshold are applied at paint time. Selecting a row in the detail table
calls `set_highlight(i)` — that detection is drawn boldly (even below threshold) and
the rest dim. The detail table lists ALL detections (sorted by score), so sub-threshold
hits are still visible for debugging; the threshold only affects what's drawn.
### Separation of concerns
- `ui/` must not import `torch` / `ultralytics` directly. It builds detectors only via
`core/detection/factory.build_detector` and talks to `core/` through the `Detector`
interface and the `Detection`/`CensorType` types.
- New detector kinds: implement `core/detection/base.Detector`, register the string in
`core/detection/factory.build_detector`, and add it to `_DETECTORS` in
`ui/main_window.py`.
## Commands
```powershell
python -m venv .venv; .\.venv\Scripts\Activate.ps1
pip install -e . # classic detector needs no torch/CUDA
python -m hvideotool # reopen the last project (or create/open one in-app)
python -m hvideotool "C:\path\to\MyProject" --detector yolo --model models\lada_mosaic_detection_model_v4_accurate.pt
pip install -e ".[yolo]" # + install torch separately, see README
```
No formal test suite. Headless sanity check: set `QT_QPA_PLATFORM=offscreen`, build a
`MainWindow`, `Project.create(tmp)` + copy a few images into `frames/`,
`_open_project(project)`, drive `file_list.setCurrentRow(...)`, and read
`detail_table` / `detail_header`. Or run `build_detector(config).detect(...)` on a
frame directly.
## Conventions
- Match the style of surrounding code; keep `core/` free of Qt where reasonable.
- Type hints on public functions and the `Detector` interface.
- Model weights (`.pt`) and large media are **not** committed — keep them in `models/`
and `.gitignore`d.
- User-facing strings / README are in Russian; code identifiers and this file in English.
## Gotchas
- **Wrong YOLO model = "noise".** The YOLO detector needs a **censorship** model
(LADA `models\lada_mosaic_detection_model_v4_accurate.pt`). If `model_path` points at
a generic COCO model (e.g. the `yolo11n-seg.pt` in the repo root, which Ultralytics
auto-downloads / is the training base), it detects people/objects and maps them to
`CensorType.UNKNOWN` → purple boxes that look like noise. This was a real user trap.
**Switching to yolo/combined without a model auto-picks one** via
`MainWindow._auto_find_model()`: it scans `./models/**.pt` and matches only filenames
containing `lada`/`mosaic` (so it skips the COCO `yolo11n-seg.pt` trap), no prompt; it
falls back to the "Модель…" file dialog only when nothing suitable is found.
- **classic-CV is approximate and noisy on real video.** Its mosaic heuristic (low
block-reconstruction residual + 2D gradient + contrast) fires on textured real
footage (skin/hair/fabric/JPEG) → many false positives, while simultaneously missing
real mosaic after the `proc_max_dim=720` downscale softens block edges (measured:
contrast/grad fall below `mosaic_contrast_min`/`mosaic_grad_min`). For real-video
mosaic use `yolo`/`combined` + LADA. For anime there is no good public model.
- **Domain matters.** LADA is trained on REAL video (JAV). It detects some anime mosaic
but not all. The real anime fix is *retraining* a YOLO11-seg (see `scripts/training/`),
not tuning more classic thresholds.
- **YOLO detector = LADA weights** ([HF `ladaapp/lada`](https://huggingface.co/ladaapp/lada)).
YOLO **segmentation** model, classes `{0: mosaic_nsfw, 1: mosaic_sfw_head}` → both map
to `CensorType.MOSAIC` (`_name_to_type` matches "mosaic" in the class name). Detects
mosaic only; black bars / blur stay with classic. Weights + Ultralytics are AGPL-3.0
(accepted). `yolo.py` lazy-imports `torch`/`ultralytics`.
- **No model weights in the repo.** Code must fail with a clear, actionable message
when the model path is missing — not a raw stack trace (`factory._require_model`,
`YoloDetector.__init__`).
- **CUDA/torch install is environment-specific.** Don't add torch to core deps; it
stays out (the `yolo` extra pulls only Ultralytics) and is installed separately.
- **CPU-only torch must not request CUDA.** A `+cpu` torch build raises "Torch not
compiled with CUDA enabled" the moment something calls `.cuda()`. Both engines guard
for this: `YoloDetector` picks `cuda` only when `torch.cuda.is_available()` (even an
explicit `yolo_device="cuda"` is downgraded to cpu); `DeepMosaicsRestorer._ensure_loaded`
forces `gpu_id="-1"` when CUDA is absent (its vendored `model_util.todevice` /
`data.im2tensor` call `.cuda()` for any `gpu_id != "-1"`, e.g. the `dm_gpu="0"` default).
So a wrong/CPU-only torch falls back to CPU instead of crashing.
- **QImage from a numpy buffer must be `.copy()`d** (see `ImageView.set_image`),
otherwise it aliases a buffer that gets freed → garbage/crash.
- **Always use `core/imageio.py`** (`imread_unicode`/`imwrite_unicode`) for images —
`cv2.imread`/`imwrite` silently fail on non-ASCII Windows paths.
- Don't reintroduce any generative / ControlNet dependency, nor the removed video
*playback pipeline* (PyAV, producer/consumer worker threads, player, project session).
(The new `core/project.py` is an on-disk layout, not that thread-based "project
model".) NOTE: a **single** background `Job` thread for detect/restore (`ui/workers.py`)
IS in scope now (keeps the GUI responsive) — that's different from the rejected
multi-thread video pipeline. The one allowed video touch is
`core/video/extract.py` (one-shot decode → a new project's `frames/`, behind "Создать
из ролика…"): ffmpeg CLI — `_find_ffmpeg()` prefers PATH, else the binary bundled by
the `imageio-ffmpeg` dep, else cv2 fallback. Keyframe-only `-skip_frame nokey` is
~10× faster than every-frame; `-hwaccel` does NOT help (GPU transfer overhead). Use
ffmpeg/cv2, not PyAV, and keep it synchronous. Decoding every frame is the inherent
cost — the speed lever is decoding *fewer* frames (keyframes).