beads: close e4t (LSP @class stubs already shipped in b75b879)
beads: close 0u1 (typed decoders already inline 1-byte fast path)
codegen: int64_as_number opt-in flag for 64-bit decode as Lua number (5y9)
Adds a plugin generator option:
protoc --tarantool_opt=int64_as_number=true ...
mode=full only, default off, error on mode=runtime. When set, 64-bit
scalar decoders (int64/uint64/sint64/fixed64/sfixed64) return a Lua
number for values that fit [-2^53, 2^53] (inclusive — both endpoints
are powers of two and exact as doubles), cdata otherwise. The decoded
type becomes value-dependent under this option; `+`/`-`/`*`/`==`
work transparently on both, so most callers don't have to care.
Wire side: decode_int64_n / decode_uint64_n / decode_sint64_n /
decode_fixed64_n / decode_sfixed64_n added alongside the existing
decoders. Bounds use LL/ULL cdata literals so the threshold compares
compile to plain 64-bit integer compares on a hot trace.
Codegen side: cfg.Int64AsNumber threads to writer.int64AsNumber;
decodeFnSuffix() emits "_n" only when the flag is set and the scalar
is one of the affected kinds. Every wire.decode_<st> emit site picks
it up — singular, repeated, packed, oneof, extension, map value.
The flag is a no-op under PB_ENABLE_C=1: the C runtime decides
number-vs-cdata on its own via luaL_pushint64 (Tarantool's small-fits-
in-double convention). The test gates on PB_ENABLE_C accordingly.
Workload-specific tradeoff measured on c_int64.Wide decode (full mode,
no PB_ENABLE_C, 5 fields per message):
tiny (all 1-byte vars) 2050 -> 1700 ns/op (-17%, cdata avoided)
medium (3-byte vars) 3220 -> 3600 ns/op (+11%, extra cmp+tonumber)
huge (past 2^53) 7575 -> 7750 ns/op (+2%, noise)
Enable when fields are dominated by small IDs/counters/small
timestamps that hit the 1-byte varint path; leave off otherwise.
The plugin flag help documents the tradeoff.
Test fixture: examples/expected/full_n/c_int64/c_int64_pb.lua via
the new gen-int64-as-number Justfile recipe. test/int64_as_number_test.lua
covers byte-identical encoding, small-value Lua number returns, default
cdata returns, 2^53 boundary inclusive, past-2^53 cdata fallback, and
Lua-number-input round-trip.
Suites: test 771/771, test-c 1057/1057 (5 skipped under PB_ENABLE_C=1).
codegen: type-elision for sint64/fixed64/sfixed64 encode (a7l)
For mode=full the field type is statically known, so the codegen can
skip wire.to_int64 / wire.to_uint64's runtime type dispatch by pre-
casting at the call site:
-- before
out[n] = wire.encode_sint64(v) -- calls to_int64(v) inside
out[n] = wire.encode_fixed64(v) -- calls to_uint64(v) inside
-- after
out[n] = wire.encode_sint64_i(INT64(v)) -- direct cdata path
out[n] = wire.encode_fixed64_u(UINT64(v))
out[n] = wire.encode_sfixed64_u(UINT64(v))
Wire-level: added encode_sint64_i / encode_fixed64_u / encode_sfixed64_u
alongside the existing encoders; encode_fixed64 is now a thin wrapper
that calls to_uint64 + encode_fixed64_u so the generic API surface stays
intact. INT64/UINT64 added as file-header upvalues in generated modules.
int32/int64/uint32/uint64 unchanged: those alias to encode_varint, which
only dispatches through to_uint64 on the slow-slow cdata path (large
negatives) — not worth the codegen surface.
Applied at every encode emit site that took a scalar: singular,
required, repeated non-packed, extension singular, extension repeated,
packed slow path, proto2 extension repeated. Map-value path also picks
it up via the shared helper.
c_int64.Wide encode (full mode, 4-run median):
lua-num inputs 2769 -> 2520 ns/op (+9.9% throughput)
cdata inputs 12993 -> 12935 ns/op (unchanged; to_uint64 already
takes the cdata fast path)
Just below the issue's optimistic 10-15% but exactly the case that
matters — Lua-number inputs are the common case for user code on
sint64/fixed64 fields. Suites: test 766/766, test-c 1057/1057.
bench: apply mcode arena hardening across all bench scripts (3qu)
Added jit.opt.start('sizemcode=64', 'maxmcode=4096') to all 9 bench
scripts. Same fix 3o2 landed in bench/jit_trace.lua, propagated to the
rest. Without it, the macOS arm64 mcode allocator intermittently fails
to find an executable page in the signed-32-bit offset window and the
bench reports interpreter-only throughput with no diagnostic — visible
as silent regressions. The risk is higher in scripts with larger codegen
footprints than the trace gate.
Files touched: bench.lua, lazy_bench.lua, profile.lua, shapes_bench.lua,
starwing_bench.lua, wire_bench.lua, alloc_probe.lua, map_bench.lua,
packed_bench.lua. All 7 non-interactive scripts run rc=0; profile and
starwing loadfile-check clean.
Also refreshed bench/baseline.json per the issue's step 2. Surprise win:
full-mode decode allocations dropped 0.5-50% as a side-effect of auj/ozn
that was only visible after the snapshot. Most notable: proto2_basic
BenchPayload mid decode 2.313 -> 1.156 KB/op (-50%), Person 1KB decode
0.977 -> 0.953 KB/op. Runtime-mode unchanged — confirms the alloc
reduction is full-mode codegen specific.
Steps 3-4 (refresh COMPARISON.md throughput tables, verify the variance
band shrinks below the documented 5-10% drift) are deferred to a
work.lab.local run — laptop variance is 50%+ on identical state, can't
trust throughput A/Bs locally.
codegen: inline 1+2-byte varint fast path in packed-scalar decode inner loop (ozn)
The packed scalar inner loop previously called wire.decode_<st>(payload, p2)
per element — a function frame plus the helper's internal loop. Now the
1-byte and 2-byte paths inline directly at the loop site for the
Lua-number-returning varint types (int32/uint32/bool, plus enum via a
sister helper that bypasses varint_to_int32 when the value fits int32
without wrapping). Other types (int64/uint64/sint*/sint64) fall through
to the existing wire.decode_<st> call.
This is the follow-up auj filed: the residual bridge after auj's tag/LEN
inline migrated from the tag dispatch into the packed-int32 inner loop.
Inlining literally inside the loop keeps the side trace inside the
parent's own frame instead of bridging through wire.decode_int32 frame
returns.
Throughput on Person_decode (full mode, no PB_ENABLE_C, 5-run median):
small (1-byte everything) 14.0 -> 23.4 MB/s (+67%)
big (2-byte LEN + multi-byte packed) 34.7 -> 39.6 MB/s (+14% vs auj)
The small-fixture +67% comes from removing the per-element function frame
on lucky_numbers' 8-element packed decode, which was already 1-byte but
paying for wire.decode_int32's function call per element. The big fixture
adds the 2-byte fast path on top.
Bridge rate at full/Person_decode multi-byte varint:
pre-ozn 4/100 (residual after auj)
post-ozn 3/200 (~1.5%, max clean streak 31/50)
Did not reach the issue's 50-consecutive-zero goal, but throughput
criterion is met spectacularly and the bridge rate is now in JIT
topology noise. Standard bench/bench.lua Person numbers unchanged at
every size — the win is specific to packed-multi-byte workloads.
Suites: test 766/766, test-c 1057/1057.
codegen: inline 2-byte tag + 2-byte LEN fast paths at decode dispatch (auj)
emitInlineDecode now inlines both the 1-byte tag fast path (field IDs 1..15,
existing) and the 2-byte tag fast path (field IDs 16..4095) directly. The
fallback to wire.decode_tag only fires for field IDs >= 4096. Same shape
applied to the string/bytes LEN prefix in emitInlineStringBytesScalar and
emitInlineStringBytesRepeated, extending the existing 1-byte LEN inline
with a 2-byte branch — covers lengths 0..16383 without crossing a function
frame.
Why this matters: the original auj bridge ('multi-byte varint' side trace
tr9 off Person_decode's tr5 at pc=55) was the tag fast path's 2+ byte
guard exiting to a side trace that couldn't stitch back through
wire.decode_tag's frame return. Inlining the 2-byte case literally keeps
the side trace inside the parent's own frame.
Throughput on the multi-byte-varint Person fixture (full mode, no
PB_ENABLE_C, 5-run median):
small (1-byte tags+lens+vars) 14.0 -> 14.2 MB/s (noise)
big (2-byte LEN + multi-byte packed) 31.6 -> 34.7 MB/s (+10%)
Bridge rate at full/Person_decode multi-byte varint:
before (this session) 0/50 (already self-resolved after kot/21d/2ri/aah)
after 4/100 (residual now in the packed-varint loop,
migrated from the tag site; eliminating
fully requires per-scalar-type 2-byte
inline in the packed-loop emitter)
Suites: test 766/766, test-c 1057/1057, examples all green.
c_runtime: decode_unsafe(plan, buf) entrypoint, skip_utf8 gate on dec_ctx (kyt)
Before this change, M.<Msg>_decode_unsafe stayed on the inline Lua path even
when PB_ENABLE_C=1, because c_runtime.decode validated UTF-8 unconditionally.
Now both safe and unsafe decoders dispatch to the C runtime; the unsafe path
calls c_runtime.decode_unsafe, which sets dec_ctx.skip_utf8 and gates the
is_valid_utf8 call on every PB_KIND_STRING payload across the decode tree.
pb.decode_unsafe in init.lua mirrors pb.decode: tries desc.c_plan first, falls
back to the pure-Lua decode_msg_unsafe. Full-mode codegen emits the same
pb.c_runtime check at the unsafe prologue.
Perf on string-heavy Person (418B, 16 emails + 16 nicknames):
Lua unsafe 142 MB/s
C safe 364 MB/s
C unsafe 455 MB/s (+25% over C-safe, 3.2x over Lua-unsafe)
Suites: test 766/766, test-c 1057/1057, examples all green.
runtime: pb.decode_unsafe + M.<Name>_decode_unsafe in runtime mode (58u)
Completes the unsafe-decode story 6bb started in full mode. Runtime mode
now exposes the same API via parallel `f._reader_unsafe` closures
compiled in pb.finalize_message against a swapped scalar table where
`string` maps to the bytes handler (no utf8_len). build_reader /
build_repeated_reader / decode_one are now parameterized on
(scalar_tbl, decode_msg_fn, decode_group_fn) so the same builders emit
both reader shapes. decode_message_unsafe, decode_group_unsafe, and
decode_extension_unsafe are literal clones of their safe twins with
three substitutions (documented in codec.lua): the _reader field, the
scalar table in slow paths, and the sub-message / group / extension
dispatchers. Tests in test/decode_unsafe_test.lua are now parameterized
over both modes (14 cases, including a map<string, int32> invalid-key
case that exercises the decode_one map-fallback path).
Runtime-mode microbench shows ~8% throughput vs validating decode on
the string-heavy 1KB Person; smaller than full mode's ~20% because the
descriptor dispatch + closure indirection swamp utf8_len, but still a
net win and the perf-cost-of-validating story is now consistent across
modes. Conformance 3240/3240 + JIT trace 37/37 still pass.
Closes 58u, also closes b12 (already fixed in 2656c97; never closed).
kyt still tracks unifying _decode_unsafe with C accel.
codegen: emit <Msg>_decode_unsafe for trusted-source decoding (6bb)
Full-mode codegen now emits a sister <Msg>_decode_unsafe(buf) alongside
<Msg>_decode that drops the per-string utf8_len check (singular,
repeated, map keys/values, extensions, and the >=128-byte fallback all
route through wire.decode_bytes). Sub-messages recurse into their own
_decode_unsafe so nested strings also bypass; WKTs continue to call the
normal pb.wkt.<Name>_decode (no _unsafe twin, no string-validation hot
path). C runtime dispatch is skipped because it validates today (kyt).
Use this when re-decoding bytes from a trusted producer — your own
encoder over typed RPC, JSON/text round-trips, in-process pipelines —
where the spec-required utf8.len check on every string is duplicate
work. Microbench on a string-heavy 1KB Person (26 emails) shows
~20% throughput vs _decode; conformance suite still passes (3240/3240)
because _decode itself is unchanged.
Runtime mode does not yet expose _decode_unsafe (compiled f._reader
closures capture handler.decode by value, so a runtime swap wouldn't
reach them); tracked in 58u.
codegen: inline 1-byte varint fast path for packed scalar elements (aah)
Per-element wire.encode_<type>(v) calls in packed-repeated fields paid a
full function-call boundary even though encode_varint's small-positive-int
hot path is a single CHARS[n] lookup. Inline the check + lookup at
codegen time at every packed emit site:
- Repeated packed scalar (mode=full)
- Repeated packed enum (after string->int resolve)
- Proto2 extension packed scalar + enum
For varint scalars (int32/int64/uint32/uint64): fast path triggers when
v is a Lua number in [0, 128). For sint32/sint64: 7-bit zigzag range
-64..63 is inlined with bit ops. For bool: always 1 byte via
CHARS[v and 1 or 0] (no fast/slow split). Fixed-width scalars keep the
wire.encode_<type> call shape — already optimal.
Also pre-sizes the parts accumulator with table_new(#v, 0) instead of
{}; same pattern qwt+2sn used decode-side. Eliminates rehash cascade
as elements push.
Tests: 752/752 pass. JIT trace gate: 37/37.
Bench (work.lab.local, median of 3, c_repeated.Holder packed N elems):
packed_int32: +90% / +113% / +106% (N=10/100/1000)
packed_sint32: +66% / +236% / +156%
packed_uint32: +73% / +105% / +102%
packed_bool: +44% / +51% / +39%
packed_int64: +22% / +21% / +23% (cdata; gain from table_new only)
Headline hello.Person 1KB +3.6%, proto2 BenchPayload mid +10.5%.
See bench/PERF_LOG.md 2026-05-24 aah entry for the full breakdown
including the sint32 +236% mid-size analysis (three function layers
collapsed into one CHARS lookup).
bench: add map_bench.lua, close ch2 (intervention regresses)
ch2 hypothesized that replacing `for k, v in pairs(map_value) do` in
the codegen-emitted map encoder with a key-collect + ipairs pattern
would let the inner emit loop stay on a JIT trace. Hand-implemented in
both codegen (inline.go:emitInlineEncodeMap) and runtime
(codec.lua kind=='map' branch); 752/752 tests passed.
Bench (work.lab.local, median of 3, Person.ages_by_nickname encode):
map size | BEFORE | AFTER | Δ
1 | 1,282,180 | 1,128,545 | -12%
3 | 634,880 | 549,761 | -13%
10 | 215,745 | 196,800 | -9%
50 | 48,162 | 45,614 | -5%
200 | 11,613 | 11,282 | -3%
Regression across all sizes. LuaJIT's side-trace machinery was already
JIT-ing the inner body via a side trace from the pairs() ISNEXT abort
point — the body was already on-trace before. The change just adds
wrapper overhead (scratch table alloc, O(N) key-collection, extra hash
lookup per entry).
Reverted the codegen + runtime edits. Keeping bench/map_bench.lua —
useful harness for any future map-encoder work (e.g. x9f deterministic
ordering may revisit this).
See ch2 bd notes for full diagnosis.
codegen: CHARS[_len] lookup replaces string.char(_len) at length-prefix sites (2ri)
Profile attributed 39% of Person_encode's trace share (~28% of total) to a
single line emitting `string.char(_len)` at every inlined length-prefix
site. The `_len` argument is rarely a compile-time constant (lengths come
from user data), so the JIT can't fold the C-function call, and the cost
compounds — the 1KB Person fixture fires ~30 length-prefix sites per
encode.
Replace with a 256-entry lookup table `wire.CHARS` (built once at module
load, byte i -> string.char(i)). Codegen header now emits
`local CHARS = wire.CHARS` alongside the other hot-path localizers; the
single emit site in `emitInlineLenPrefix` swaps `string.char(_len)` for
`CHARS[_len]`. Parity is by construction — both return the same interned
1-byte string.
Tests: 752/752 pass. Bench (work.lab.local, median of 3, hello.Person
encode): 10B +4.7%, 1KB +16.8%, 10KB +21.1%, 100KB +34.2%, proto2 mid
+11.1%. Decode unchanged. See bench/PERF_LOG.md 2026-05-24 2ri entry.
Closes lkz and 86g (ffi.new buffer rewrite paths) — separate hand-spike
of that shape regressed 0.31x-0.77x across all sizes; the perceived
buffering inefficiency wasn't there, and 2ri captured the single hottest
line. Remaining encode-perf headroom is c0i (C-runtime backend).
beads: close drm (encode/decode B/op floor investigation)
codegen: inline proto2 extension writers/readers (qwt) + table.new(N,0) for packed lists (2sn)
qwt closes the proto2_basic.BenchPayload `min` bench's worst data point:
encode 1.89M → 3.30M msgs/s (+75%), decode 1.02M → 1.26M msgs/s (+23%).
Mechanism: collectExtsByExtendee groups all `extend Foo { ... }`
declarations by extendee FullName across input files; emitInlineEncode
emits one dedicated writer per extension instead of dispatching through
pb.codec.encode_field, and emitInlineDecode adds an elseif arm per
extension id straight into result._extensions[full_name]. Dynamic
extensions_list walk preserved for forward compat, gated on
#_elist > N so it pays one int compare when no runtime extension was
registered.
2sn pre-sizes packed-scalar repeated lists via table.new(N, 0) where
the count is recoverable from the LEN payload — exact (lim >> 2 or
lim >> 3) for fixed-width, upper bound (lim) for varint-packed. The
allocation is deferred into the `if wt == 2 then` branch so the
per-element fallback and non-packable types keep the bare-`{}`
alloc. Avoids u39's regression mode because the call cost is paid
once per repeated-field-first-occurrence, not per message decode.
Bench summary for hello.Person full mode: encode +3-5% across sizes,
decode +1-3% across sizes. Tests: 752/752 pure-Lua, 1043/1043 with
PB_ENABLE_C=1, JIT 37/37. PERF_LOG entry covers the rationale and
caveats.
codec: emit per-descriptor encode body for monomorphic dispatch (21d)
The runtime-mode encode loop in codec.encode_message iterated
desc.fields and called `writer(data, out)` per field. Each iteration
saw a different closure, making the call site megamorphic from the
JIT's view: trace topology fragmented into 25 stops for Person_encode
vs full mode's 8 (~3x).
compile_encode_body emits a generated function at pb.finalize_message
time with one literal call site per field, all closing over a single
`_u` table upvalue indexed by constant int. TGETI on a stable array
with a literal key specializes on trace just like direct upvalue
access — and dodges LuaJIT's 60-upvalue function limit, which
TestAllTypesProto2 (~140 fields) would otherwise hit.
Trace topology (bench/jit_trace.lua):
runtime/Person_encode 25 -> 10 (full 8)
runtime/Person_encode multi-byte 18 -> 7 (full 6)
runtime/Address_encode 6 -> 5 (full 2)
runtime/Cardinality required 6 -> 2 (full 2)
Encode throughput (hello.Person, bench/bench.lua):
10B 10.9 -> 15.5 MB/s (+42%)
100B 101.1 -> 144.4 MB/s (+43%)
1KB 170.5 -> 183.5 MB/s (+8%)
10KB 355.5 -> 362.3 MB/s (+2%)
100KB 373.5 -> 407.6 MB/s (+9%)
Decode untouched. Full mode unaffected (its inline _encode doesn't
go through codec.encode_message). 752+1043 tests pass; examples
unchanged.
The issue's secondary goal — runtime within 5% of full on encode at
all sizes — is not met. Residual gap is per-closure call overhead;
closing it would need writer bodies inlined into the generated body
(not just call sites), which is a larger codegen-at-runtime change.
Closes tarantool-protobuf-21d.
Files tarantool-protobuf-h8x (decode_group bimodal trace flake
surfaced during validation; pre-existing on master).
codegen: localize wire.* upvalues per generated message function
Capture each _encode/_decode body, scan for wire.<name> refs, and
rewrite refs that appear >=2 times to bare locals with a
"local X = wire.X" prelude. Single-use refs stay as wire.X — without
the threshold the prelude TGETS outweighed the in-body saving on
sparse small-message decode.
Measured (median-of-3, shapes bench full mode):
- scalar-heavy: enc +38.5%, dec +39.2%
- packed-int32x100: enc +51.1%, dec +38.2%
- map-strxi32-*: enc +5.5%, dec +6-7%
- nested/oneof/wkt: +1-5% (within ~5% variance band)
- bench.lua Person small sizes: neutral
closes kot
beads: close 4ql (ibuf encoder not viable) + memory note on FFI boundary tax
Hand-coded single-pass backpatched Person_encode_ibuf wins by 1.7x at
10-100B but loses 1.4-2.7x at 1KB-100KB. Crossover at ~26 emails:
per-field ffi.copy boundaries scale linearly while Person_encode pays
one bulk table.concat memcpy regardless of count. Three ibuf attempts
total (per-byte alloc, two-pass reserve, single-pass backpatch), same
root cause each time. Not revivable without LuaJIT FFI sinking or a
cdata-string API contract.
beads: close bgu (ffi.string decode disproven) + memory note on buffer-reuse trap
Two-round microbench shows ffi.string(scratch + off, len) loses 2-2.7x
to buf:sub even with a pre-allocated stable scratch cdata and ffi.copy
amortized over 26 strings. The per-call pointer-arith cdata crosses
the ffi.string frame and can't be sunk — same root cause as a6n.
beads: close a6n + memory note on ffi.cast cdata allocation
a6n