# CLAUDE.md Guidance for Claude Code (and other agents) working in this repository. ## Memory: EchoVault (read this first) This project uses the **EchoVault** MCP server for persistent, cross-session memory. Prior sessions store architectural decisions, fixed bugs, and gotchas there. Follow this protocol every session: 1. **At session start — load context.** Call `memory_context` (project is auto-detected from cwd) before doing any work. Use `memory_search` for specific topics (e.g. "detector model", "classic-cv", "false positives"). 2. **During work — search before re-deciding.** When the task touches an area that may have prior context, `memory_search` it first instead of re-deriving decisions. 3. **Before ending a session — save what matters.** Call `memory_save` when you made a design decision, fixed a bug (include root cause + fix), found a non-obvious gotcha, or the user corrected/clarified a requirement. Pick the right `category` (`decision` / `bug` / `pattern` / `learning` / `context`). Do **not** save trivia, things obvious from the code, or duplicates. EchoVault is the source of truth for *why* things are the way they are; this file is the stable, high-level map. When they disagree, trust on-disk code first, then EchoVault, then this file — and update whichever is stale. ## What this project is **HVideoTool** is a Windows-first desktop GUI utility that **detects already-applied censorship** (mosaic, pixelation, blur, black bars) in **images**, and draws outlines over the detected censored regions. You open a folder of images; it runs each through a detector, draws the regions, and shows a detailed per-image list of what it found. > **Scope was deliberately narrowed (this session).** It used to extract frames from > video, detect, and play back with overlays (a "project model" with worker threads). > That whole video pipeline was **removed** — the tool is now a simple **image-folder > inspector** for viewing/debugging detector output on test images. Pre-extract video > to frames externally if you need that. Keep this scope sharp: - It is a **detection + overlay/inspection** tool. It does **not** remove, restore, or reconstruct censored content. - It does **not** generate images. There is **no** ControlNet / SDXL / diffusion pipeline. (`xinsir/controlnet-union-sdxl-1.0` was considered early but rejected — a generative model, not a detector. Do not reintroduce it.) - It detects **already-censored** regions, not "content that should be censored" (i.e. not an NSFW classifier). - It does **not** decode video. No PyAV. Input is image files only. ## Target environment - **OS:** Windows 11 x64 (primary). Use PowerShell syntax in commands. - **Python:** 3.11+. - **GPU:** NVIDIA + CUDA via PyTorch, only for the YOLO detector. CPU fallback works but is slow. The `classic` detector needs no torch and no GPU. ## Tech stack (decided) | Concern | Choice | |----------------|---------------------------------------------| | GUI | PySide6 (Qt 6) — LGPL | | Image IO | OpenCV (`opencv-python`) + NumPy, unicode-safe via `core/imageio.py` | | Detector | classic-CV heuristic; Ultralytics YOLO (LADA) behind a pluggable interface | Torch/CUDA + Ultralytics enter only with the YOLO detector. Keep that dependency optional (the `yolo` extra in `pyproject.toml` pulls only Ultralytics; torch is installed separately per the README). The classic detector must keep running with no torch present. ## Architecture (as implemented) > This reflects the actual code on disk. It is a synchronous, single-threaded GUI app > — no worker threads, no project/cache files. The only video touch is a one-shot > "Создать из ролика…" that decodes a clip to a folder of JPGs via the **ffmpeg CLI** > (cv2.VideoCapture fallback; NOT PyAV); detection still works on image folders only. > Detection runs on the GUI > thread (lazily per image, or via "Детектировать все"). When code and this file > disagree, trust the code. ``` hvideotool/ ├── __main__.py # entry point + CLI (optional folder arg, --detector, --model) ├── app.py # QApplication bootstrap; run(config, folder=None) ├── config.py # AppConfig + DetectionConfig/OverlayConfig (thresholds live here) ├── settings_store.py # persist detector/model/threshold/last_dir to ~/HVideoTool/settings.json ├── ui/ │ ├── main_window.py # the whole UI: toolbar + [file list | image view | detail table] │ └── image_view.py # renders an image + draws polygon/bbox overlays (QPainter); can highlight one └── core/ ├── imageio.py # unicode-safe imread/imwrite (np.fromfile + imdecode) ├── video/ │ ├── extract.py # extract_frames(): ffmpeg CLI (cv2 fallback) -> JPGs; keyframe/step modes + downscale │ └── frame.py # Frame dataclass (image BGR, index, pts) — the detector input type └── detection/ ├── base.py # Detector ABC: detect(frame) -> list[Detection] ├── factory.py # build_detector(config) -> classic | yolo | combined ├── types.py # Detection (+ to_dict/from_dict), CensorType enum ├── classic_cv.py # ClassicCVDetector — heuristic; accepts a `types` filter ├── yolo.py # YoloDetector — Ultralytics YOLO-seg; lazy-imports torch/ultralytics └── composite.py # CompositeDetector — merge detectors + IoU dedup ``` ### How it works - `MainWindow` holds the config, builds the detector lazily via `build_detector` (cached by detector+model+conf in `_make_detector`), and keeps `_results: dict[path -> list[Detection]]` as the detection cache. - "Открыть папку" lists image files (`_IMAGE_EXTS`) with a progress bar (bulk insert with `setUpdatesEnabled(False)` + periodic `processEvents`), so a big folder doesn't freeze silently. - **Viewing and detecting are decoupled on purpose** (so browsing stays instant even with a slow CPU detector): selecting a file only *shows* it with its cached result (header reads "не рассчитано" if none). Detection runs on **double-click**, the "Рассчитать кадр" action (Space → `_recompute_current`, force-recomputes current), or "Детектировать все" (whole folder, progress bar). Do NOT re-add auto-detect-on-select. Results cache in `_results`; the file-list row gets a count suffix when computed. Switching detector/model clears the cache (`_invalidate_results`). - **Collections (curation).** "Создать коллекцию…" makes a destination folder (`_collections_base()` = the opened folder's parent, else `~/HVideoTool/collections`) and marks it active. The file list is `ExtendedSelection`; "В коллекцию" / Ctrl+M **moves** (`shutil.move`, not copy) the selected frames there, removing them from the list/`_files`/cache. `_unique_dest` avoids clobbering (`foo.jpg` → `foo (1).jpg`). Use case: sort frames into a training/example set while inspecting detections. - `image_view.ImageView` draws the image scaled-to-fit plus overlays. Overlay visibility/threshold are applied at paint time. Selecting a row in the detail table calls `set_highlight(i)` — that detection is drawn boldly (even below threshold) and the rest dim. The detail table lists ALL detections (sorted by score), so sub-threshold hits are still visible for debugging; the threshold only affects what's drawn. ### Separation of concerns - `ui/` must not import `torch` / `ultralytics` directly. It builds detectors only via `core/detection/factory.build_detector` and talks to `core/` through the `Detector` interface and the `Detection`/`CensorType` types. - New detector kinds: implement `core/detection/base.Detector`, register the string in `core/detection/factory.build_detector`, and add it to `_DETECTORS` in `ui/main_window.py`. ## Commands ```powershell python -m venv .venv; .\.venv\Scripts\Activate.ps1 pip install -e . # classic detector needs no torch/CUDA python -m hvideotool # open a folder in-app python -m hvideotool "C:\path\to\images" --detector yolo --model models\lada_mosaic_detection_model_v4_accurate.pt pip install -e ".[yolo]" # + install torch separately, see README ``` No formal test suite. Headless sanity check: set `QT_QPA_PLATFORM=offscreen`, build a `MainWindow`, `open_path(folder)`, drive `file_list.setCurrentRow(...)`, and read `detail_table` / `detail_header`. Or run `build_detector(config).detect(...)` on a frame directly. ## Conventions - Match the style of surrounding code; keep `core/` free of Qt where reasonable. - Type hints on public functions and the `Detector` interface. - Model weights (`.pt`) and large media are **not** committed — keep them in `models/` and `.gitignore`d. - User-facing strings / README are in Russian; code identifiers and this file in English. ## Gotchas - **Wrong YOLO model = "noise".** The YOLO detector needs a **censorship** model (LADA `models\lada_mosaic_detection_model_v4_accurate.pt`). If `model_path` points at a generic COCO model (e.g. the `yolo11n-seg.pt` in the repo root, which Ultralytics auto-downloads / is the training base), it detects people/objects and maps them to `CensorType.UNKNOWN` → purple boxes that look like noise. This was a real user trap. - **classic-CV is approximate and noisy on real video.** Its mosaic heuristic (low block-reconstruction residual + 2D gradient + contrast) fires on textured real footage (skin/hair/fabric/JPEG) → many false positives, while simultaneously missing real mosaic after the `proc_max_dim=720` downscale softens block edges (measured: contrast/grad fall below `mosaic_contrast_min`/`mosaic_grad_min`). For real-video mosaic use `yolo`/`combined` + LADA. For anime there is no good public model. - **Domain matters.** LADA is trained on REAL video (JAV). It detects some anime mosaic but not all. The real anime fix is *retraining* a YOLO11-seg (see `scripts/training/`), not tuning more classic thresholds. - **YOLO detector = LADA weights** ([HF `ladaapp/lada`](https://huggingface.co/ladaapp/lada)). YOLO **segmentation** model, classes `{0: mosaic_nsfw, 1: mosaic_sfw_head}` → both map to `CensorType.MOSAIC` (`_name_to_type` matches "mosaic" in the class name). Detects mosaic only; black bars / blur stay with classic. Weights + Ultralytics are AGPL-3.0 (accepted). `yolo.py` lazy-imports `torch`/`ultralytics`. - **No model weights in the repo.** Code must fail with a clear, actionable message when the model path is missing — not a raw stack trace (`factory._require_model`, `YoloDetector.__init__`). - **CUDA/torch install is environment-specific.** Don't add torch to core deps; it stays out (the `yolo` extra pulls only Ultralytics) and is installed separately. - **QImage from a numpy buffer must be `.copy()`d** (see `ImageView.set_image`), otherwise it aliases a buffer that gets freed → garbage/crash. - **Always use `core/imageio.py`** (`imread_unicode`/`imwrite_unicode`) for images — `cv2.imread`/`imwrite` silently fail on non-ASCII Windows paths. - Don't reintroduce any generative / ControlNet dependency, nor the removed video *pipeline* (PyAV, project/cache, worker threads, playback). The one allowed video touch is `core/video/extract.py` (one-shot decode → JPG folder, behind "Создать из ролика…"): ffmpeg CLI — `_find_ffmpeg()` prefers PATH, else the binary bundled by the `imageio-ffmpeg` dep, else cv2 fallback. Keyframe-only `-skip_frame nokey` is ~10× faster than every-frame; `-hwaccel` does NOT help (GPU transfer overhead). Use ffmpeg/cv2, not PyAV, and keep it synchronous. Decoding every frame is the inherent cost — the speed lever is decoding *fewer* frames (keyframes).