~bigbes/tarantool

tarantool-protobuf

ref: 5c7055577dedb31f1c4817f02ec5ad6f65a068b5 tarantool-protobuf/bench/PERF_LOG.md -rw-r--r-- 18.7 KiB
5c705557 — Eugene Blikh c-accel: descriptor -> C plan compiler (bd-mq7) 2 months ago

#Performance optimization log

Iterative log of optimization work on tarantool-protobuf encode/decode hot paths. Each entry captures: what changed, why, measured before/after, and any caveats.

#Workflow

For each beads task on the optimization track:

  1. Implementbd update <id> --claim, write the change.
  2. Testjust test. Fix any breakage before measuring anything.
  3. Benchmarkjust bench (and just jit-trace if the change touches hot wire paths). Capture full output, compute deltas vs. previous entry's "after" numbers, and append a new entry to this file.
  4. Commit — single commit per task, including the PERF_LOG.md update. bd close <id> afterwards.
  5. Next — pick the next P1 (or whatever's at the top of bd ready in the perf bucket) and repeat from step 1.

Notes:

  • "Before" for each entry equals the previous entry's "After" — single contiguous history, no per-task baselines.
  • Always run just bench on a quiet machine; numbers can swing 5-10% from background noise on macOS. Re-run if a number looks suspicious.
  • If a change regresses a workload (any cell drops >5%), document it honestly. A regression that buys 2× somewhere else is a tradeoff worth recording, not a failure to hide.
  • Decoder, encoder, and the proto2 BenchPayload schemas are tracked separately because optimizations rarely move all three uniformly.

#Schemas tracked

  • hello.Person at 10B / 100B / 1KB / 10KB / 100KB. Real-world-shaped: string fields, nested messages, repeated emails. Both full (inlined codegen) and runtime (descriptor dispatch) modes.
  • proto2_basic.BenchPayload at min / mid. Pins proto2 extension and default-value paths.

#Baseline — 2026-05-18

Pre-optimization snapshot. Tarantool 3.7.0-0-g1f1ec9fdf, LuaJIT 2.1.0-beta3, macOS arm64.

#hello.Person — encode

mode size msgs/s MB/s alloc B/op
full 10B 2,250,858 22.5 136
full 100B 2,215,919 208.3 136
full 1KB 240,381 223.6 1368
full 10KB 43,240 416.6 8540
full 100KB 4,920 475.7 131605
runtime 10B 1,098,666 11.0 136
runtime 100B 1,089,811 102.4 136
runtime 1KB 186,459 173.4 1368
runtime 10KB 39,600 381.5 8540
runtime 100KB 4,433 428.5 131605

#hello.Person — decode

mode size msgs/s MB/s alloc B/op
full 10B 2,476,811 24.8 112
full 100B 2,091,875 196.6 112
full 1KB 109,479 101.8 1000
full 10KB 14,961 144.1 4840
full 100KB 1,461 141.3 33512
runtime 10B 2,103,514 21.0 112
runtime 100B 1,843,624 173.3 112
runtime 1KB 103,503 96.3 1000
runtime 10KB 13,710 132.1 4840
runtime 100KB 1,336 129.1 33512

#proto2_basic.BenchPayload

mode size enc msgs/s enc MB/s dec msgs/s dec MB/s
full min 1,485,112 7.4 846,439 4.2
full mid 181,413 110.5 73,650 44.9
runtime min 1,163,210 5.8 982,154 4.9
runtime mid 162,481 99.0 73,952 45.0

#2026-05-18 — u39: table.new(0, N) for result tables (REVERTED)

Task: [tarantool-protobuf-u39] Decoder: emit table.new(0, N) for result tables. The thesis was that pre-sizing the result table's hash part to the message's field count would avoid the rehash that fresh-{} tables pay as fields are added.

Change attempted: Header now require('table.new') with a pcall fallback so plain-Lua loads still work; emitInlineDecode allocates the result via table_new(0, len(m.Fields)) instead of {}. 745/745 tests pass. JIT 37/37.

Bench (Person full decode, msgs/s, median-of-3 vs post-cch):

size post-cch u39 Δ
10B 2,495,633 2,111,197 -15.4%
100B 2,131,196 1,754,540 -17.7%
1KB 115,310 117,270 +1.7%
10KB 15,831 15,627 -1.3%
100KB 1,620 1,617 -0.2%

Outcome: reverted. Small-payload decode collapsed. The table_new(0, N) call (through an upvalue, with arguments) is several times the cost of a {} literal, and at 10B/100B the decode loop only populates 2-4 fields — well under the LuaJIT default hash size where the first rehash would land. The rehash savings never materialize because the small case doesn't reach them; meanwhile the extra call cost is paid every decode.

Large payloads (1KB+) saw no meaningful gain either — by then decode time is dominated by tag decoding, string extraction, and varint parsing, not table allocation. The drm memory entry already documented "136 B alloc floor is not the throughput bottleneck"; this result confirms it from the opposite direction.

Lesson: allocation-shape optimizations are only worth it when the workload spends real time in allocation. The bench corpus does not exercise that mode; if a workload ever does (lots of tiny messages, or messages with very wide field sets), revisit then. Sticking with {} also keeps generated code compatible with plain Lua (no LuaJIT-only require) — small platform-portability win.

Commit: none — change reverted. PERF_LOG entry is the only artifact.


#2026-05-18 — cch: local counter for repeated-field append

Task: [tarantool-protobuf-cch] Decoder: use local counter instead of #list at repeated-field append. Profile attributed 6.2% of Person 1KB decode time to list[#list + 1] = val re-traversing the list per append (26 emails per Person → 26 re-scans).

Change: In cmd/protoc-gen-tarantool/internal/gen/inline.go, the generated M.X_decode now declares one counter local per repeated non-map field at function entry:

local _n_emails = 0
local _n_lucky_numbers = 0
-- ...

Each list[#list + 1] = val emit becomes _n_emails = _n_emails + 1; list[_n_emails] = val. Counters survive across loop iterations, so out-of-order wire entries for the same field keep appending past the existing length without re-scanning.

Map fields are skipped (hash-keyed, no array index).

Tests: 745/745. JIT: 37/37, 0 bridges.

Bench (Person full, msgs/s, median-of-3 vs post-4kj baseline):

size dir post-4kj cch Δ
1KB dec 109,792 115,310 +5.0%
10KB dec 15,084 15,831 +5.0%
100KB dec 1,512 1,620 +7.1%
10B dec 2,579,347 2,495,633 -3.2%
100B dec 2,122,601 2,131,196 +0.4%
1KB enc 293,367 294,691 +0.5%
10KB enc 59,390 58,684 -1.2%
100KB enc 6,570 6,666 +1.5%

Decode wins land at 1KB+ where the repeated-field pattern dominates (Person has 26 emails). At 10B no repeats fire, so the small -3.2% is noise. Encode unchanged path, flat. Profile's "6% of decode time" target translated cleanly into +5–7% on sizes that exercise repeated fields.

Runtime mode: flat. cch only touches mode=full repeated emit; runtime dispatches through pb.codec which already uses a different append shape.

Commit: see git history for SHA.


#2026-05-18 — 4kj: inline 1-byte tag fast path at decode call sites

Task: [tarantool-protobuf-4kj] Decoder: generated tag/length fast path for full-mode decode. Profile attributed ~21% of hello.Person 1KB decode time to wire.decode_tag dispatch.

Change: In cmd/protoc-gen-tarantool/internal/gen/inline.go, the generated M.X_decode while-loop now decodes the 1-byte tag form inline before falling back to wire.decode_tag for multi-byte. A header tweak in gen.go localizes string.byte, bit.band, and bit.rshift at the top of every generated file so each call inside the loop becomes a straight-line local-call.

-- Before:
local id, wt
id, wt, pos = wire.decode_tag(buf, pos)

-- After:
local id, wt
local _b = string_byte(buf, pos)
if _b ~= nil and _b < 0x80 then
    wt = band(_b, 7)
    if wt >= 6 then error("illegal wire type " .. wt, 0) end
    id = rshift(_b, 3)
    if id == 0 then error("illegal field number 0", 0) end
    pos = pos + 1
else
    id, wt, pos = wire.decode_tag(buf, pos)
end

Field numbers 1..15 always encode as a single byte; that's the overwhelmingly dominant case in real payloads, including every Person field in the bench corpus.

Tests: 745/745 pass. JIT trace gate: 37/37, all bridges still 0.

Bench (Person full, msgs/s, median-of-3 vs proper post-h8v median-of-3 baseline):

size dir post-h8v 4kj Δ
10B dec 2,247,620 2,579,347 +14.7%
100B dec 1,949,565 2,122,601 +8.9%
1KB dec 102,175 109,792 +7.5%
10KB dec 14,100 15,084 +7.0%
100KB dec 1,393 1,512 +8.5%
10B enc 2,291,108 2,272,701 -0.8%
100B enc 2,250,757 2,313,101 +2.8%
1KB enc 293,677 293,367 -0.1%
10KB enc 60,269 59,390 -1.5%
100KB enc 6,746 6,570 -2.6%

Methodology note (lesson from gcy): the post-h8v "single-run" baseline I'd captured for h8v was a peak run (the bench is noisier than I'd expected on macOS arm64). For 4kj I re-baselined post-h8v with a 3-run median before comparing, which made the decode win obvious and exposed the encode -1 to -3% as small/within-noise. Going forward, medians-of-3 are the comparison standard; PERF_LOG entries earlier than this one used single-run baselines (h8v's numbers are likely tilted ~3-5% optimistic).

Bench (proto2_basic.BenchPayload full): roughly flat — proto2 mid encode/decode within ±1% of post-h8v. Expected: BenchPayload doesn't have a tight decode loop the way Person 1KB does.

Bench (runtime mode): the runtime mode wrappers don't see the inlined fast path — they delegate to pb.codec.decode. Runtime decode 1KB measured at 97,285 vs an earlier baseline 103,503, but that baseline was single-run and inconsistent with the variance I've seen since. Treating this as noise; no semantic change is plausible for runtime mode here (only file-header upvalues added, three locals never referenced by runtime wrappers).

Caveats / leftovers:

  • Multi-byte tags (field ids > 15) still pay the wire.decode_tag call. Generated protobuf rarely uses high field numbers, but extensions often do; the fallback keeps them correct.
  • Encode regressed slightly at 10KB+. Plausible cause: three extra file-level upvalues (band/rshift/string_byte) shift LuaJIT's function-prologue layout for Person_encode even though that function doesn't use them. Not currently worth optimizing.
  • wire.decode_len is still a function call. Inlining its byte-read half can come next, but the substring it produces is unavoidable.
  • The richer "order-prediction" form of this task — emit a literal tag-byte equality check per declared field — is deferred. It would double-dispatch (literal-match + id-match fallback) and the simple inline already captures ~half the available gain.

Commit: see git history for SHA.


#2026-05-18 — gcy: inline nested-message decode at the call site (REVERTED)

Task: [tarantool-protobuf-gcy] Decoder: inline nested-message decode at the call site. Profile flagged result.address = M.Address_decode(payload) at hello_pb.lua:1031 as a 100% interpreter bail in the vl trace; thesis was that replacing the call with the inlined Address decode body would eliminate the bail and lift Person decode by 20–40%.

Change attempted: In cmd/protoc-gen-tarantool/internal/gen/inline.go, added inlineCandidateForDecode (singular, non-group, non-WKT message fields whose target has no further message subfields) and emitInlineNestedDecode, which emits Address's decode loop inline in Person_decode using a do/end scope to shadow the outer result. The length prefix was read in-place (no wire.decode_len substring alloc).

Tests: 745/745 pass. JIT trace gate: 37/37 (after a flaky first-run 0/37 caused by the documented macOS arm64 mcode alloc issue).

Bench (hello.Person — full, msgs/s, median of 3 runs):

size dir h8v baseline after gcy Δ
1KB enc 302,117 292,680 -3.0%
1KB dec 110,115 104,182 -5.3%
10KB enc 64,052 59,543 -6.4%
10KB dec 15,086 14,113 -6.6%
100KB enc 6,487 6,661 +2.7%
100KB dec 1,519 1,396 -8.1%

Outcome: reverted. Tests passed but the change is a net regression at every size that exercises the inlined nested decode (Person 1KB/10KB have Address embedded; 100KB grows the Address body but is dominated by emails). Encode also regressed because Person_encode's trace had to account for a larger Person_decode (shared mcode arena / instruction cache pressure on macOS arm64, or LuaJIT abandoning some inlining of Person_encode under the new pressure).

Why the profile claim didn't translate:

  1. vl (and jit.p count) reported Address_decode as 100% Interpreted for the parent line, but in the live benchmark LuaJIT was already inlining the small Address_decode body into Person_decode's trace when entering. The "bail" was a tooling artifact of how jit.attach('trace') attributes side traces, not a sustained interpreter fallback. The benchmark numbers refute the trace interpretation.
  2. Person_decode trace went from N stops to 22 stops after the inline. Larger root traces compile more slowly, are more sensitive to side-trace stitching limits, and produce more mcode — exactly the failure mode CLAUDE.md's "Keep hot wire helpers small" warning describes, applied at the call-site instead of the helper.
  3. The inlined-body's do/end scope with shadowed result may introduce extra upvalue references that LuaJIT 2.1 doesn't optimize as well as a plain function call into a trace it has already inlined.

What would actually help here: the real decode bottleneck on hello.Person (per profile recap) is still decode_string (utf8 validation = ~16% of decode time), decode_tag (21%), and list[#list+1] = val at the email append loop (6%). Those are the next targets — see entries on 6bb (skip_utf8_validation), 4kj (generated tag/length fast path), cch (local counter for repeated append).

Beads: issue closed with --reason referencing this entry; not reopened. Inlining nested decode bodies can come back if a workload shows that the nested call site is the hot trace boundary, but the candidate restriction + bench coupling needs to be rethought first (measure each call-site in isolation, not just at the trace level).

Commit: none — change reverted before commit. PERF_LOG entry is the only artifact.


#2026-05-18 — h8v: inline 1-byte varint length prefix at codegen sites

Task: [tarantool-protobuf-h8v] Encoder: codegen-time inline FFI writes (mode=full) — first slice. Full FFI-buffer rewrite is still future work; this attacks the single hottest line identified by jit.p profiling.

Change: In cmd/protoc-gen-tarantool/internal/gen/inline.go, every length-prefixed emit site (singular/repeated message, singular/repeated string|bytes, packed scalar bundle, packed enum bundle, map entry) now inlines the 1-byte varint fast path:

-- Before:
n = n + 1; out[n] = wire.encode_varint(#_b)

-- After:
local _len = #_b
if _len < 128 then
    n = n + 1; out[n] = string.char(_len)
else
    n = n + 1; out[n] = wire.encode_varint(_len)
end

Eliminates the function call + dispatch for every length prefix under 128 bytes — which is the dominant case for proto strings, message bodies, and packed scalar bundles. Profile flagged this as ~33% of encode time (out[n] = wire.encode_varint(#_b)) + another ~17% in encode_varint dispatch.

Tests: 745/745 pass. JIT trace gate: 37/37 (no regressions).

Bench (hello.Person — full encode, msgs/s and Δ vs baseline):

size before after Δ
10B 2,250,858 2,405,176 +6.9%
100B 2,215,919 2,389,715 +7.9%
1KB 240,381 302,117 +25.7%
10KB 43,240 64,052 +48.1%
100KB 4,920 6,487 +31.9%

Bench (hello.Person — full decode): flat (±1%). Expected; decode path unchanged.

Bench (runtime mode): flat (±2%). Expected; only mode=full codegen touched, runtime mode dispatches through pb.codec.encode_field which still calls into wire.encode_varint for length prefixes.

Bench (proto2_basic.BenchPayload — full):

size dir before after Δ
min enc 1,485,112 1,502,031 +1.1%
mid enc 181,413 203,832 +12.4%
min dec 846,439 897,014 +6.0%
mid dec 73,650 71,496 -2.9%

mid decode -2.9% is within run-to-run noise (proto2 mid has no string fields touched by this change; the wider distribution at 73K msgs/s swings ±5%).

Alloc/op: unchanged (table-of-strings model preserved). Future h8v slices targeting the per-field table writes themselves would move this.

Caveats / leftovers:

  • The 1-byte ceiling at 128 bytes matches the proto3 varint boundary; the 100KB Person bench has email-string + name-string lengths above 128, so it pays the slow path for those — but message-body lengths there are still mostly < 128 because they wrap individual nested messages. Net result is still +32%.
  • Map encode still uses table.concat(entry) then inlined length prefix. Per-piece wire.encode_len(...) inside map values (emitMapPiece message branch) was not rewritten — the call sits inside a single out[idx] = ... slot assignment and untangling it would require separating value-build from value-write. Map fields are not on the current hot benchmark; deferred.
  • Runtime mode (descriptor dispatch via pb.codec.encode_field) does not benefit. The corresponding follow-up is 21d (codec dispatch fragmenting traces); independent route to similar wins.

Commit: see git history for SHA.