Video summary
More Speed & Simplicity: Practical Data-Oriented Design in C++ - Vittorio Romeo - CppCon 2025
Main summary
Key takeaways
Main ideas, concepts, and lessons
1) Data-oriented design (DoD) vs OOP (OP) mindset
- Speaker’s framing: A common OOP approach models a world of autonomous objects (entities) that own their data and perform their own behavior via virtual methods and polymorphism.
- DoD alternative mindset: Treat the program as a pipeline of data transformations:
- Data is the centerpiece.
- Behavior (update/draw) is centralized in a “world/system” that processes batches of data.
- Data layout and access patterns are primary drivers of performance and simplicity.
2) Performance comes from memory layout, not just algorithms
- Key claim: If the data layout causes cache misses, the CPU may spend most time waiting for memory.
- Cache line concept:
- The smallest unit moved between memory and CPU cache is a cache line (typically ~64 bytes).
- Even if you need a few bytes, the CPU still loads an entire cache line.
- Practical implication:
- Contiguous “flat” storage improves spatial locality.
- Predictable access patterns enable hardware prefetching.
- Sometimes a “worse” algorithm can outperform a “better” one due to memory behavior.
3) The rocket demo as a controlled experiment
- What the demo simulates: Many rockets with attached smoke/fire emitters and particles.
- Requirements:
- Physics-like integration (position/velocity/acceleration).
- Particles and emitters support time-varying properties (opacity, scale, rotation).
- Each entity/particle is identifiable via identity/handles.
- System must be extensible (new actors/effects).
- What changes in the experiment: not the operations, but the data layout.
- Observed outcome: an OOP-style layout tanks frame rate; a data-oriented layout makes it playable.
Methodology / implementation path shown (detailed)
Step A: Implement using OOP-style hierarchy (baseline)
- Create polymorphic base type:
Entitywith:virtual update(delta_time)virtual draw(render_target)virtual destructor- state fields like position/velocity/acceleration
- reference to
Worldto query other entities / spawn entities - an
aliveboolean to signal removal
- Derive concrete types:
Particle:- fields like scale/opacity/rotation and rates of change
alivebecomes true only while opacity > 0 (fade-out cleanup)- overrides
drawto render smoke/fire textures
Emitter:- stores spawn timer/rate
virtual spawn_particle()overridden bySmokeEmitter/FireEmitter
Rocket:- has attached emitters (smoke + fire) and moves with physics
Worldstores:std::vector<std::unique_ptr<Entity>>(heterogeneous heap allocations)
- Update/draw loop:
- Iterate all entities in the vector and call
update()anddraw().
- Iterate all entities in the vector and call
- Cleanup:
- Use erase/remove-style logic with predicates (mentions C++20
erase_if).
- Use erase/remove-style logic with predicates (mentions C++20
Why it performs poorly (as explained):
- Scattered heap allocations → many cache misses.
- Virtual dispatch overhead (vtable lookups).
- Frequent dynamic allocation churn when spawning/destroying.
Step B: Move toward DoD by flattening and decoupling (reduce indirection)
- Goals of the refactor:
- Remove per-entity heap allocations.
- Remove inheritance polymorphic trees.
- Decouple data from logic:
- Entities become plain structs (data only).
- The world/system becomes the behavior executor.
- Replace single heterogeneous container with multiple homogeneous containers:
std::vector<Particle>std::vector<Rocket>std::vector<optional<Emitter>>(slot array)
- Handle emitter/rocket relationships using indices, not pointers:
- Store emitter references in rockets as indices.
- Use
optional“slots” so indices remain stable:- if an emitter is destroyed, clear the slot but don’t reorder.
- Emitter spawning:
- World loops over emitters by type (initial version uses an enum-like branching per type).
- When emitting, push particles into the particle vector.
- Updates become batch operations:
- World loops through all particles and applies physics/integration in bulk.
- World updates emitters and spawns particles from timers.
- World updates rockets and positions attached emitters using indices.
Benchmarked improvement:
- Same operations, different data layout → large speedup (about 70% less update time on average in their benchmark set).
- Rendering also improves (mentioned: bulk sending to GPU).
Step C: Reduce branching and shrink hot data (still AOS initially)
- Optimization idea:
- Avoid branching in hot loops by grouping data by type before processing.
- Implementation:
- Split particles into separate containers:
- one for smoke particles
- one for fire particles
- Split emitters similarly.
- Use lambdas/local callbacks to reduce repeated code while staying explicit.
- Split particles into separate containers:
- Reduce common type sizes / padding:
- Example: change some fields from
size_t-like touint16_twhen safe (to fit more in cache lines / reduce footprint).
- Example: change some fields from
- Result:
- Smaller improvement overall (~5.7%) in update time (since particles are main bottleneck and branching wasn’t dominant there).
- Rendering improves further (mentioned: ~12%).
Step D: Change data layout to SOA (structure of arrays)
- Transform from AoS (array of structs) to SoA (struct of arrays):
- Instead of
Particle { pos, vel, acc, opacity, scale, ... }per element, - Use
ParticleSOAcontaining separate vectors per field:positions[],velocities[],accelerations[], etc.
- Instead of
- Implicit invariant:
- All field vectors are the same length and grow/shrink in sync.
- Execution:
- Update loop iterates by index and updates all relevant vectors.
- Cleanup in SOA:
- Standard algorithms don’t fit well, so implement a custom erase:
- Use a predicate to determine which indices should be removed (opacity fade-out).
- Perform a partition-like movement:
- move kept elements from left/right,
- swap/move across all vectors simultaneously
- then resize all vectors at once (to keep them aligned)
- Standard algorithms don’t fit well, so implement a custom erase:
- Benchmarked result:
- Large update-time reduction (~32% decrease).
- Rendering becomes ~2x faster by mapping grouped fields to GPU buffers more directly.
Why SOA helped even in a “worst case” scenario:
- Explained overhead from AOS vs SOA vectorization differences:
- AOS may require gather/scatter when applying SIMD to strided data.
- SOA tends to enable more continuous SIMD-friendly access patterns.
Step E: Maintain simplicity despite SOA awkwardness (proposed techniques)
- Problem:
- SOA makes inserting/updating particles verbose (e.g., multiple
push_backcalls per field).
- SOA makes inserting/updating particles verbose (e.g., multiple
- Suggested direction (not fully implemented in the talk):
- Create a “magic” wrapper/template that:
- accepts a “particle-like” aggregate
- automatically distributes fields into the SOA vectors
- Use compile-time reflection-like techniques:
- In principle: C++20/17 techniques (e.g.,
boost::pfr) for aggregate field access - Anticipated benefit from C++26 reflection for more ergonomic APIs using parameter names/types.
- In principle: C++20/17 techniques (e.g.,
- Create a “magic” wrapper/template that:
Additional lessons: correctness, extensibility, and trade-offs
“SoA is not always the answer”
- DoD is a mindset:
- Data layout and access patterns drive the choice.
- Consider:
- fields accessed together frequently → store together
- cold fields accessed rarely → store separately
- target platform constraints (cache size, SIM length, etc.)
- Hybrid experimentation:
- Switching layouts at compile time can allow benchmarking without rewriting logic.
Don’t reject OOP entirely
- OOP can be valuable at the wrong level vs right level distinction:
- If OOP is used to wrap data-oriented engines (e.g., “particle manager” API), it can be beneficial.
- Virtual overhead can be negligible when calls aren’t in hot loops.
- Team collaboration:
- Clear interfaces and SRP (single responsibility) can improve understanding and parallel work.
Side benefits of DoD
- Serialization becomes easier:
- If state is “just bytes in structures/arrays,” saving/loading and networking snapshots are straightforward.
- Testability/debugging/tooling:
- Save and replay state for deterministic repro.
- Build tools/UI since state is explicit data.
- Multi-threading:
- Centralized batch loops can be more easily parallelized than polymorphic object callbacks.
Q&A highlights (key points only)
- Allocators as intermediate solution:
- The speaker believes allocators could significantly recover some performance for OOP designs.
- Loss of explicit relationships (pointer-based ties):
- Indices make relationships explicit during update; relationships become visible in the central data-processing loops.
- Testability comparison:
- DoD helps by making it easy to store/load test cases as data.
- OOP can still be fine when used for higher-level abstractions; performance-critical storage should remain data-oriented.
- How to find where to optimize:
- Use profilers (Intel VTune, Valgrind tools, perf) to detect cache/memory bottlenecks and CPU idle time.
- Views/ranges overhead and practical use:
- Ranges/functional abstractions may add overhead without optimization; efficiency depends on inlining/compile flags.
- Random/non-batch access:
- Depends on application; can still improve cache friendliness by aligning data layout with access patterns (even for graph-like workflows).
Speakers / sources featured (as identified in the subtitles)
Speaker
- Vittorio Romeo (main speaker; keynote by Vittorio Romeo)
Referenced people / sources (mentioned as examples or in recommendations)
- Jason (addressed by name; appears to have introduced/asked a prompt at the start)
- Mike Acton (gave a keynote in 2014; cited as inspiration)
- John Leos (co-author)
- Russell Lapenov (co-author)
- Alistair Meredith (co-author)
- Scott Meyers (recommended talk: “CPU Caches and Why You Care” from 2014)
- Jonathan Mueller (recommended talk at CppCon: “Cache Friendly C++”)
- Barry (referenced as giving a relevant talk about SOA/reflection; “Barry’s excellent talk on Monday”)
Libraries / standards / tools mentioned
- SFML (render target; also modernizing to C++17 mentioned)
- SDL
- C++ standard / features (e.g.,
std::erase_if, C++17/20, “reflection in C++26”,std::optional,std::vector) - Intel VTune (profiler)
- Valgrind (suite/profiling tools)
- perf
- OpenMP (suggested for parallelization)
- Open source / tech context: SFML, SDL, Steam, Factorio (for test/state replay example)
boost::pfr(mentioned as a way to do aggregate reflection-like behavior in C++17)
Rate this summary
Your feedback will help improve summaries.
Improve this summary
Reprocess with a stronger model when the summary feels incomplete or inaccurate.
Translate summary in another language
Ask questions to this video
Chat for follow-up questions, clarifications, and source-backed answers.