Parallel rendering pipeline, ring vertex buffer, phase timings

Five optimizations measured on the 100k-entity stress scene (Release,
vsync off): 103 FPS baseline -> 297 FPS.

- Sprite submission and vertex building run on all cores above
  Renderer2DOptions.ParallelThreshold (default 8192). Work is sliced
  into 4096-entity segments: a Friflo chunk holds a whole archetype,
  so per-chunk parallelism degenerates to one thread. Segments merge
  in deterministic order, preserving radix sort stability.
- Vertex buffer is ring-written with SetDataOptions.NoOverwrite
  (GPU buffer 2x frame size); Discard only on wrap-around.
- Texture2DRegion precomputes UVs - four float divisions per sprite
  per frame removed.
- Renderer2D exposes per-phase timings (submit/sort/build/upload/draw),
  shown in the sample HUD - all further optimization is data-driven.
- Sample BounceSystem parallelized the same segmented way.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Leonid Pershin
2026-06-11 05:06:55 +03:00
co-authored by Claude Fable 5
parent af319d1276
commit e06f24a319
10 changed files with 589 additions and 132 deletions
-1
View File
@@ -23,4 +23,3 @@
- DevTools-модуль на ImGui.NET: инспектор сущностей, дебаг-панели
- Бенчмарки BenchmarkDotNet для систем (сейчас производительность контролируется стресс-сценой)
- Spatial hash для culling на очень больших мирах (если профилирование покажет необходимость)
- Параллельная запись вершин (Parallel.For по чанкам), если упрёмся в CPU на ещё больших сценах