12 KiB
CLAUDE.md
Guidance for Claude Code (and other agents) working in this repository.
Memory: EchoVault (read this first)
This project uses the EchoVault MCP server for persistent, cross-session memory. Prior sessions store architectural decisions, fixed bugs, and gotchas there. Follow this protocol every session:
- At session start — load context. Call
memory_context(project is auto-detected from cwd) before doing any work. Usememory_searchfor specific topics (e.g. "detector model", "classic-cv", "false positives"). - During work — search before re-deciding. When the task touches an area that
may have prior context,
memory_searchit first instead of re-deriving decisions. - Before ending a session — save what matters. Call
memory_savewhen you made a design decision, fixed a bug (include root cause + fix), found a non-obvious gotcha, or the user corrected/clarified a requirement. Pick the rightcategory(decision/bug/pattern/learning/context). Do not save trivia, things obvious from the code, or duplicates.
EchoVault is the source of truth for why things are the way they are; this file is the stable, high-level map. When they disagree, trust on-disk code first, then EchoVault, then this file — and update whichever is stale.
What this project is
HVideoTool is a Windows-first desktop GUI utility that detects already-applied censorship (mosaic, pixelation, blur, black bars) in images, and draws outlines over the detected censored regions. You open a folder of images; it runs each through a detector, draws the regions, and shows a detailed per-image list of what it found.
Scope was deliberately narrowed (this session). It used to extract frames from video, detect, and play back with overlays (a "project model" with worker threads). That whole video pipeline was removed — the tool is now a simple image-folder inspector for viewing/debugging detector output on test images. Pre-extract video to frames externally if you need that.
Keep this scope sharp:
- It is a detection + overlay/inspection tool. It does not remove, restore, or reconstruct censored content.
- It does not generate images. There is no ControlNet / SDXL / diffusion
pipeline. (
xinsir/controlnet-union-sdxl-1.0was considered early but rejected — a generative model, not a detector. Do not reintroduce it.) - It detects already-censored regions, not "content that should be censored" (i.e. not an NSFW classifier).
- It does not decode video. No PyAV. Input is image files only.
Target environment
- OS: Windows 11 x64 (primary). Use PowerShell syntax in commands.
- Python: 3.11+.
- GPU: NVIDIA + CUDA via PyTorch, only for the YOLO detector. CPU fallback works
but is slow. The
classicdetector needs no torch and no GPU.
Tech stack (decided)
| Concern | Choice |
|---|---|
| GUI | PySide6 (Qt 6) — LGPL |
| Image IO | OpenCV (opencv-python) + NumPy, unicode-safe via core/imageio.py |
| Detector | classic-CV heuristic; Ultralytics YOLO (LADA) behind a pluggable interface |
Torch/CUDA + Ultralytics enter only with the YOLO detector. Keep that dependency
optional (the yolo extra in pyproject.toml pulls only Ultralytics; torch is
installed separately per the README). The classic detector must keep running with no
torch present.
Architecture (as implemented)
This reflects the actual code on disk. It is a synchronous, single-threaded GUI app — no worker threads, no project/cache files. The only video touch is a one-shot "Создать из ролика…" that decodes a clip to a folder of JPGs via the ffmpeg CLI (cv2.VideoCapture fallback; NOT PyAV); detection still works on image folders only. Detection runs on the GUI thread (lazily per image, or via "Детектировать все"). When code and this file disagree, trust the code.
hvideotool/
├── __main__.py # entry point + CLI (optional folder arg, --detector, --model)
├── app.py # QApplication bootstrap; run(config, folder=None)
├── config.py # AppConfig + DetectionConfig/OverlayConfig (thresholds live here)
├── settings_store.py # persist detector/model/threshold/last_dir to ~/HVideoTool/settings.json
├── ui/
│ ├── main_window.py # the whole UI: toolbar + [file list | image view | detail table]
│ └── image_view.py # renders an image + draws polygon/bbox overlays (QPainter); can highlight one
└── core/
├── imageio.py # unicode-safe imread/imwrite (np.fromfile + imdecode)
├── video/
│ ├── extract.py # extract_frames(): ffmpeg CLI (cv2 fallback) -> JPGs; keyframe/step modes + downscale
│ └── frame.py # Frame dataclass (image BGR, index, pts) — the detector input type
└── detection/
├── base.py # Detector ABC: detect(frame) -> list[Detection]
├── factory.py # build_detector(config) -> classic | yolo | combined
├── types.py # Detection (+ to_dict/from_dict), CensorType enum
├── classic_cv.py # ClassicCVDetector — heuristic; accepts a `types` filter
├── yolo.py # YoloDetector — Ultralytics YOLO-seg; lazy-imports torch/ultralytics
└── composite.py # CompositeDetector — merge detectors + IoU dedup
How it works
MainWindowholds the config, builds the detector lazily viabuild_detector(cached by detector+model+conf in_make_detector), and keeps_results: dict[path -> list[Detection]]as the detection cache.- "Открыть папку" lists image files (
_IMAGE_EXTS) with a progress bar (bulk insert withsetUpdatesEnabled(False)+ periodicprocessEvents), so a big folder doesn't freeze silently. - Viewing and detecting are decoupled on purpose (so browsing stays instant even
with a slow CPU detector): selecting a file only shows it with its cached result
(header reads "не рассчитано" if none). Detection runs on double-click, the
"Рассчитать кадр" action (Space →
_recompute_current, force-recomputes current), or "Детектировать все" (whole folder, progress bar). Do NOT re-add auto-detect-on-select. Results cache in_results; the file-list row gets a count suffix when computed. Switching detector/model clears the cache (_invalidate_results). - Collections (curation). "Создать коллекцию…" makes a destination folder
(
_collections_base()= the opened folder's parent, else~/HVideoTool/collections) and marks it active. The file list isExtendedSelection; "В коллекцию" / Ctrl+M moves (shutil.move, not copy) the selected frames there, removing them from the list/_files/cache._unique_destavoids clobbering (foo.jpg→foo (1).jpg). Use case: sort frames into a training/example set while inspecting detections. image_view.ImageViewdraws the image scaled-to-fit plus overlays. Overlay visibility/threshold are applied at paint time. Selecting a row in the detail table callsset_highlight(i)— that detection is drawn boldly (even below threshold) and the rest dim. The detail table lists ALL detections (sorted by score), so sub-threshold hits are still visible for debugging; the threshold only affects what's drawn.
Separation of concerns
ui/must not importtorch/ultralyticsdirectly. It builds detectors only viacore/detection/factory.build_detectorand talks tocore/through theDetectorinterface and theDetection/CensorTypetypes.- New detector kinds: implement
core/detection/base.Detector, register the string incore/detection/factory.build_detector, and add it to_DETECTORSinui/main_window.py.
Commands
python -m venv .venv; .\.venv\Scripts\Activate.ps1
pip install -e . # classic detector needs no torch/CUDA
python -m hvideotool # open a folder in-app
python -m hvideotool "C:\path\to\images" --detector yolo --model models\lada_mosaic_detection_model_v4_accurate.pt
pip install -e ".[yolo]" # + install torch separately, see README
No formal test suite. Headless sanity check: set QT_QPA_PLATFORM=offscreen, build a
MainWindow, open_path(folder), drive file_list.setCurrentRow(...), and read
detail_table / detail_header. Or run build_detector(config).detect(...) on a
frame directly.
Conventions
- Match the style of surrounding code; keep
core/free of Qt where reasonable. - Type hints on public functions and the
Detectorinterface. - Model weights (
.pt) and large media are not committed — keep them inmodels/and.gitignored. - User-facing strings / README are in Russian; code identifiers and this file in English.
Gotchas
- Wrong YOLO model = "noise". The YOLO detector needs a censorship model
(LADA
models\lada_mosaic_detection_model_v4_accurate.pt). Ifmodel_pathpoints at a generic COCO model (e.g. theyolo11n-seg.ptin the repo root, which Ultralytics auto-downloads / is the training base), it detects people/objects and maps them toCensorType.UNKNOWN→ purple boxes that look like noise. This was a real user trap. - classic-CV is approximate and noisy on real video. Its mosaic heuristic (low
block-reconstruction residual + 2D gradient + contrast) fires on textured real
footage (skin/hair/fabric/JPEG) → many false positives, while simultaneously missing
real mosaic after the
proc_max_dim=720downscale softens block edges (measured: contrast/grad fall belowmosaic_contrast_min/mosaic_grad_min). For real-video mosaic useyolo/combined+ LADA. For anime there is no good public model. - Domain matters. LADA is trained on REAL video (JAV). It detects some anime mosaic
but not all. The real anime fix is retraining a YOLO11-seg (see
scripts/training/), not tuning more classic thresholds. - YOLO detector = LADA weights (HF
ladaapp/lada). YOLO segmentation model, classes{0: mosaic_nsfw, 1: mosaic_sfw_head}→ both map toCensorType.MOSAIC(_name_to_typematches "mosaic" in the class name). Detects mosaic only; black bars / blur stay with classic. Weights + Ultralytics are AGPL-3.0 (accepted).yolo.pylazy-importstorch/ultralytics. - No model weights in the repo. Code must fail with a clear, actionable message
when the model path is missing — not a raw stack trace (
factory._require_model,YoloDetector.__init__). - CUDA/torch install is environment-specific. Don't add torch to core deps; it
stays out (the
yoloextra pulls only Ultralytics) and is installed separately. - QImage from a numpy buffer must be
.copy()d (seeImageView.set_image), otherwise it aliases a buffer that gets freed → garbage/crash. - Always use
core/imageio.py(imread_unicode/imwrite_unicode) for images —cv2.imread/imwritesilently fail on non-ASCII Windows paths. - Don't reintroduce any generative / ControlNet dependency, nor the removed video
pipeline (PyAV, project/cache, worker threads, playback). The one allowed video
touch is
core/video/extract.py(one-shot decode → JPG folder, behind "Создать из ролика…"): ffmpeg CLI —_find_ffmpeg()prefers PATH, else the binary bundled by theimageio-ffmpegdep, else cv2 fallback. Keyframe-only-skip_frame nokeyis ~10× faster than every-frame;-hwacceldoes NOT help (GPU transfer overhead). Use ffmpeg/cv2, not PyAV, and keep it synchronous. Decoding every frame is the inherent cost — the speed lever is decoding fewer frames (keyframes).