Generate schema-specialized binary readers - #14
Merged
Conversation
jallum
force-pushed
the
perf/generated-reader
branch
from
August 31, 2026 18:35
8e48b13 to
cdfbd60
Compare
jallum
added a commit
that referenced
this pull request
Aug 31, 2026
Stacked on #14 (and #13). ## Summary - generate schema-specialized reverse-builder dispatch during `use Flatbuffer` expansion - keep schema-independent mechanics shared in the internal `Flatbuffer.Writer`: final framing, string pooling, vector materialization, scalar validation, aligned pushes, and vtable construction - compile table fields and IDs, scalar defaults, enum/union dispatch, struct padding, vector element sizes, and alignment into the caller - precompute alignment-friendly table emission order while independently assembling vtable locations in field-ID order - preserve `to_iolist/1` and `to_binary/1`, compact/shared vtables, shared strings, atom/binary input keys, safe schemas, recursive tables, and existing error tuples - extend flatc interoperability so the C++ verifier reads buffers from both interpreted and generated writers The shared-runtime refactor removed 291 lines (about 30%) from the writer generator; generated code now contains only the schema-specialized control flow and measurements. ## TDD and verification The first test traced `Flatbuffer.to_iolist/2` and `Flatbuffer.to_binary/2`; it failed on the original macro delegation and passes only when both calls use generated code. - Elixir 1.18 / OTP 27: 138 tests, 0 failures, 2 excluded - Elixir 1.19 / OTP 28: 136 tests, 0 failures, 2 excluded - Elixir 1.20 / OTP 29 coverage: 136 tests, 0 failures, 99.5% - flatc interoperability: 2 tests, 0 failures; C++ verifies interpreted and generated writer buffers - `mix deps.unlock --check-unused` - `mix compile --warnings-as-errors` - `mix format --check-formatted` - `git diff --check` ## Benchee M4 Pro, Elixir 1.19.4, OTP 28.3, 100 integers and 20 child tables, after shared-runtime extraction: | writer | average | median | memory | reductions | |---|---:|---:|---:|---:| | interpreted | 11.45 µs | 11.21 µs | 55.45 KB | 4.62 K | | generated | 10.30 µs | 10.08 µs | 45.80 KB | 3.56 K | The generated writer remains about 10% faster, allocates 17.4% less memory, and uses 22.9% fewer reductions. Both produce a 1,384-byte representative buffer. ## flatc review and alignment ordering flatc generates fixed `add_*` calls and groups scalar emission by wire size/alignment. The Elixir generator follows the same useful property, but computes true schema alignment for structs as well. Reference objects are prepared first; inline fields are emitted by descending alignment with a deterministic field-ID tiebreaker. Each emitted location has its own generated local, so the vtable is assembled in ID order without a runtime sort. This is why alignment-friendly ordering is practical in either implementation but only wins here after the ordering work itself is moved to compile time. ## Adversarial review - alternate valid field order changes non-canonical bytes, so tests compare decoded semantics and flatc verification rather than assuming one byte layout - sparse/absent fields still produce compact vtables because generated location collection is ID-ordered and stops at the highest present field - union vectors retain the existing explicit error behavior - safe-schema code generation never interns schema field names; hygienic locals avoid field-number-derived atoms - shared strings/vtables, default omission, nested structs, recursive tables, malformed scalar/enum/table/union values, binary keys, and iolist output have dedicated parity coverage - the module boundary was benchmarked after extraction; only schema-dependent dispatch remains quoted
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #13.
Summary
use Flatbufferexpansion@external_resourceTDD and verification
The first test traced
Flatbuffer.read/2and failed until binary reads stopped delegating to the interpreted implementation. The trace uses its own OTP trace session so it cannot interfere with ExCoveralls.mix compile --warnings-as-errorsmix format --check-formattedgit diff --checkBenchee
M4 Pro, Elixir 1.19.4, OTP 28.3, 100 integers and 20 child tables:
That is 2.62x faster with 46% less allocation.
Adversarial review
The fast path is deliberately restricted to binaries; flattening arbitrary iodata would trade hidden copying for benchmark speed. Iodata keeps the existing implementation. Generated clauses preserve compact-vtable schema evolution, unknown enum/union behavior, safe binary keys, identifier errors, and recursive table schemas.
OTP 29 exposed exponential type analysis when each generated field directly refined the row map shape. Generated code now routes puts through a small helper, keeping runtime map construction while presenting a stable type boundary to the compiler. Bitstring sizes are explicitly pinned for OTP 29 correctness warnings.