Parallel rendering pipeline, ring vertex buffer, phase timings

Five optimizations measured on the 100k-entity stress scene (Release,
vsync off): 103 FPS baseline -> 297 FPS.

- Sprite submission and vertex building run on all cores above
  Renderer2DOptions.ParallelThreshold (default 8192). Work is sliced
  into 4096-entity segments: a Friflo chunk holds a whole archetype,
  so per-chunk parallelism degenerates to one thread. Segments merge
  in deterministic order, preserving radix sort stability.
- Vertex buffer is ring-written with SetDataOptions.NoOverwrite
  (GPU buffer 2x frame size); Discard only on wrap-around.
- Texture2DRegion precomputes UVs - four float divisions per sprite
  per frame removed.
- Renderer2D exposes per-phase timings (submit/sort/build/upload/draw),
  shown in the sample HUD - all further optimization is data-driven.
- Sample BounceSystem parallelized the same segmented way.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Leonid Pershin
2026-06-11 05:06:55 +03:00
co-authored by Claude Fable 5
parent af319d1276
commit e06f24a319
10 changed files with 589 additions and 132 deletions
+4 -1
View File
@@ -37,7 +37,10 @@ dotnet run --project samples/MrGameEng.Sample
- ECS-first: components are plain data (`struct` implementing `IComponent`),
behavior goes into Friflo systems (`QuerySystem`), wired through `SystemRoot`.
No `Update()` methods on game objects, no inheritance-based entities.
- Hot paths (per-frame systems) must be allocation-free.
- Hot paths (per-frame systems) must be allocation-free below the renderer's parallel
threshold; above it Parallel.For scheduler overhead is the accepted trade.
A Friflo chunk holds a whole archetype — parallelize by slicing chunks into segments,
never by chunk alone. Measure in Release only, using the Renderer2D phase timings.
- Rendering: custom batcher in `Graphics` (vertex buffers, layer→depth→texture sort,
atlas support); `SpriteBatch` is not used in engine code. Draw systems write vertices
directly from Friflo chunk iteration. Orthographic camera (one active per scene),