codec: proto2 groups (SGROUP/EGROUP wire format)
Plugin: detect protoreflect.GroupKind, emit kind='group' in the
field descriptor and a SGROUP opening tag in inline codegen. The
group body bypasses the LEN-prefix path entirely — encoder emits
start-tag, body bytes, end-tag in three slots; decoder calls
pb.codec.decode_group(desc, buf, pos, field_id) which reads inner
tags until the matching EGROUP id.
Wire layer already understood SGROUP/EGROUP for skip_field; the new
decode_group reuses the per-tag dispatch from decode_message but
stops on EGROUP instead of end-of-buffer. Repeated groups bracket
each element with its own SGROUP/EGROUP pair.
Runtime parser now desugars `optional|required|repeated group Name
= id { body }` into (a) a nested message named Name and (b) a
synthetic field of kind=group whose lowercased name is `name` — so
source-parsed proto2 schemas behave the same as build-time codegen.
Text format renders groups under the submessage's capitalized name
(`SingleGroup { ... }`) rather than the lowercase field name,
matching mainline protoc's convention.
Adds a vendored, MessageSet-stripped test_messages_proto2.proto in
test/conformance/proto/ (header comment documents the patch).
10 new luatest cases (group_*) cover singular + repeated + text +
descriptor kind, across both codegen modes. 739 tests pass.
codegen: proto2 baseline — required, optional, custom defaults
Lifts the proto3-only syntax gate in the plugin and threads three
new field-descriptor attributes through codegen and the codec:
* required=true — fields declared with the proto2 `required` keyword.
Inline codegen and the runtime codec both error when
a required field is missing on encode (vs the silent
elide that proto3 implicit-presence fields get).
* optional=true — already wired for proto3 explicit `optional`; in
proto2 every singular field carries it via the
existing HasOptionalKeyword() check, giving presence
semantics without a separate emission path.
* default_value=… — proto2 [default = X] from the field descriptor,
rendered as a Lua literal (cdata for 64-bit ints,
symbolic name for enums) so consumers can surface
it; the codec itself does not auto-materialize
defaults on decode, matching how proto3 absent
fields stay nil.
Packed-by-default already flips correctly because we ask
protoreflect's `IsPacked()`, which is syntax-aware.
Adds test/proto/proto2_basic.proto with 33 luatest cases covering
required validation, optional presence, custom defaults, the proto2
unpacked-by-default repeated rule, nested-required messages, and
full-vs-runtime mode parity. `just gen-proto2-tests` regenerates the
fixture into examples/expected/{full,runtime}/.
Out of scope: extensions, extend, group; conformance harness still
skips TestAllTypesProto2.
codegen: preserve descriptor options as `options = { ... }`
Every populated *Options message — FileOptions, MessageOptions,
FieldOptions, OneofOptions, EnumOptions, EnumValueOptions,
ServiceOptions, MethodOptions — surfaces on the generated descriptor as
a plain Lua sub-table named `options` (or `oneof_options` /
`value_options` for the per-member shapes). Standard fields use their
proto name as a bare Lua key; extensions use their fully-qualified
extension name as a bracket-quoted string key.
The walker is generic — no extension-specific code paths. Consumers
pull whatever they care about: `(google.api.http)` for REST routing,
`(versionpb.etcd_version_*)` for compatibility gates, `[deprecated =
true]` for migration tooling, and any in-house extension without pb
knowing about them. Standard fields sort alphabetically before
extensions (also alphabetical by full name) so codegen output stays
byte-identical across runs.
The `options` key is only emitted when at least one field is populated,
so option-free protos produce zero-diff output to before. Resolver
re-links *Options* messages so in-file extensions surface via the
protoreflect walker — protogen builds f.Desc before in-file extensions
are registered, and only f.Proto gets the post-pass fix-up, so we
rebuild the resolver manually.
UninterpretedOption is treated as a codegen-time error: a populated
entry means protoc couldn't resolve the extension, and emitting
opaque parser state would hide the problem.
codegen: bracket-quote Lua-keyword field names
The plugin emitted bare-identifier field names in three positions:
the `M.<Type>_fields` table key, the inline encoder's `v = t.<name>`
load, and the inline decoder's `result.<name> = ...` store. When a
proto field name collided with a Lua reserved word the generated
`*_pb.lua` failed to load with `'(' expected near '<keyword>'`.
The real-world hit is pprof's profile.proto, which declares
`repeated Function function = 5` — that breaks all three emit sites.
Route every user-named identifier through new luaTableKey /
luaFieldAccess helpers that bracket-quote reserved words. Covers
field_names + oneof descriptors, oneof presence pre-pass, inline
encode/decode field bodies (including repeated and map paths), and
optional has/clear accessors.
codegen: resolve (tarantool.lua_package) via global type registry
The custom file option was being dropped silently: protoc encoded it
correctly into the FileDescriptorProto, but protobuf-go parked the
unknown extension in the message's unknown-fields tail because
E_LuaPackage was never registered with protoregistry.GlobalTypes. As a
result proto.GetExtension returned "" and every override the README
documented was a no-op — every caller fell through to the default
"<pkg>.<file>_pb" path resolution.
Register E_LuaPackage in init() and add a luatest regression that
asserts both the output path and the cross-file require strings honor
the option, parameterized over both codegen modes.
codegen: emit strict <Type>_fields / _oneofs constants for lazy view
Lazy-view callers passing a typo'd field name to :get / :has / :set /
:clear / :which got `nil` back, indistinguishable from a legitimately-
absent optional field. Failures surfaced as missing data downstream.
Each generated message now exports a M.<Type>_fields table mapping
each field name to itself (and M.<Type>_oneofs for oneof groups),
wrapped by a new pb.field_names() helper that errors on unknown-key
reads and on any write. Routing field-name arguments through these
tables turns a typo into a load-time error at the read site.
Eager _encode / _decode keep round-tripping plain Lua tables — the
constants table is a lazy-view contract, documented in
docs/api-modes.md. README and lazy_test.lua converted to the new
pattern; three new tests cover typo / read-only / oneof-typo errors.
codegen: propagate proto comments into generated _pb.lua
Leading `//` comments on messages, enums, enum values, fields, services,
and RPC methods now surface in the generated Lua:
- Message / enum descriptions become `---` blocks above `---@class` /
`---@alias`, so LuaLS shows them on hover.
- Per-field comments collapse onto a trailing `@ description` on the
matching `---@field` line.
- Enum values, service banners, and per-method entries get plain `--`
lines inside the table literals so readers without an LSP still see
the proto-side context.
Covers both `mode=full` and `mode=runtime` (emission is mode-independent
but the new luatest group runs against both).
json: strict validation pass closes the proto3 conformance suite
Six classes of relaxation that the proto3 JSON conformance corpus
flagged are now enforced. All as Recommended.* tests; combined with
the -0.0 codec fix this empties known_failures.txt and brings the
proto3 binary+JSON suite to 1493 ✓ / 1313 skipped / 0 expected
failures / 0 unexpected failures.
1. Duplicate JSON keys. Tarantool's json.decode is hash-backed and
silently collapses `{"foo":1,"foo":2}` to one entry. A small
byte-walker `find_duplicate_json_keys` runs before json.decode,
tracks per-object brace frames and key sets, errors on the
second occurrence. Closes Recommended.FieldNameDuplicate.
2. camelCase / snake_case aliases of the same proto field appearing
side-by-side. Detected inside decode_message via a `field_seen`
set keyed by proto-name; second hit errors. Closes
FieldNameDuplicateDifferentCasing{1,2}.
3. JSON null inside repeated arrays and map values. Previously
silently dropped; now errors before decode_field_value. Closes
RepeatedField{Message,Primitive}ElementIsNull and
MapFieldValueIsNull.
4. Unknown enum *names* (not integers). decode_enum used to return
nil so callers silently dropped them; now raises by default and
returns nil only when M.decode's `ignore_unknown_fields=true`
opt is set. Conformance dispatch in cmd/conformance/core.lua
forwards this flag when req.test_category ==
JSON_IGNORE_UNKNOWN_PARSING_TEST. Closes
RejectUnknownEnumStringValueIn{Optional,Repeated,Map} and the
paired IgnoreUnknownEnumStringValueIn* tests.
5. google.protobuf.NullValue JSON canonical form. The single enum
value renders as the literal JSON `null` (not the string
"NULL_VALUE"); decode accepts either, encode emits null. The
decode_field_value null-handling path also treats a JSON null
on a NullValue-typed field as "set" rather than "absent" so a
oneof gets marked active. Closes
NullValueInOtherOneof{New,Old}Format.Validator.
6. FieldMask strict round-trip. Path validity is checked on both
sides: the snake_case wire form rejects uppercase letters,
consecutive underscores, trailing underscore, and underscore
followed by anything other than a lowercase letter — these
break the snake↔camel round-trip. The JSON form rejects any
underscore in the input (must be lowerCamelCase). Closes
FieldMask{TooManyUnderscore,PathsDontRoundTrip,
NumbersDontRoundTrip}.JsonOutput and JsonInput.FieldMaskInvalidCharacter.
The pre-existing "drop unknown enum strings" unit regressions in
test/conformance_test.lua were inverted to assert the new error
shape. New strict-validation regressions in test/json_test.lua pin
all six categories so they don't regress; the `json.strict` group
runs across both codegen modes via the shared descriptor table.
codec: preserve -0.0 for proto3 float/double scalars
Proto3 default-elision dropped any singular float/double whose value
compared equal to 0 — `-0.0 == 0.0` in IEEE so the `if v ~= 0` guard
silently elided negative zero. The wire bytes for -0 differ from +0
and the TextFormatInput conformance corpus pins that -0 must survive
a round-trip; the failure surfaced as 10 unexpected text-suite
regressions covering FloatFieldNegativeZero (3 spellings × 2 outputs)
and Neg{Float,Double}FieldLargeNegativeExponentParsesAsNegZero
(2 types × 2 outputs).
Sign-bit guard via `1/v == math.huge` (positive zero yields +inf,
negative zero yields -inf). Applied in three layers:
* runtime/pb/codec.lua: is_default_scalar + the specialized
monomorphic writer for non-optional numeric scalars
* runtime/pb/text.lua: is_proto3_default (text encoder elision)
* inline codegen: scalarNotDefaultExpr emits the guard for
float/double fields in full-mode `_encode` functions
Proto3 text-format conformance now sits at 416 ✓ / 18 skipped / 0
expected failures; binary+JSON holds at 1478 ✓. Unit-test regressions
cover both directions (preserved on encode + decode round-trip, +0
still elided) and use a runtime-computed -0.0 sentinel because LuaJIT
can constant-fold the literal `-0.0` to a sign-less zero in some
load paths.
text: add pb.text.decode and wire it into conformance dispatch
Hand-written recursive-descent parser for the textproto grammar
covering every bucket the proto3 conformance suite exercises:
decimal/hex/octal integer literals with full 32/64-bit range checks,
float specials (inf/infinity/nan any case, oversize exponents
saturating to ±inf, underflows to ±0), C-style + \u/\U string escapes
with adjacent-literal concat and surrogate rejection, aggregate {} /
<> bodies, repeated short-form `[a, b, c]`, `key: K value: V` map
entries, the `[type.googleapis.com/...]` inline Any form alongside
the direct `type_url:`/`value:` form, enum-by-name-or-number,
reserved-name silent drop, numeric-field-ID tolerance, and
duplicate-singular-field rejection.
Plugin gains a small reserved_names emitter so the parser can match
mainline TextFormat::Parser's "silently drop reserved" rule. The Any
WKT descriptor advertises its real fields (type_url + string,
value + bytes) so the generic body walker can populate it directly
when the input doesn't use the inline-URL form.
cmd/conformance/core.lua stops short-circuiting text_payload to
`skipped` and runs it through pb.text.decode. The proto3 TextFormat
input suite climbs from 8 ✓ / 426 skipped to 406 ✓ / 18 skipped / 10
expected failures. The 10 surviving failures all share one cause
(proto3 -0.0 elision in the codec, not a parser bug — documented in
test/conformance/known_failures_text.txt). Binary+JSON conformance
holds at 1478 ✓. 591 unit tests pass across both codegen modes.
Closes the text-conformance-output branch.
text: render captured unknown fields, tolerate SGROUP in skip
Two changes close the proto3 text-format conformance suite:
1. `wire.skip_field` learns SGROUP/EGROUP. Wire 3 recurses through
inner tags until a matching EGROUP, with field_id checked against
the SGROUP's id. Callers (codec.lua, lazy.lua, wkt.lua, generated
full-mode `_pb.lua`) now pass the tag's field_id so groups inside
unknown-field skips don't error.
2. `pb.text.encode` walks the captured `_unknown_fields` buffer when
`opts.print_unknown_fields=true` and emits each entry in
TextFormat numeric-field form:
VARINT -> "<id>: <uint64>"
I64/I32 -> "<id>: 0x<hex>"
LEN -> speculative "<id> { <recurse> }"; rolls back to
byte-string form if the inner bytes don't parse as a
sub-message
SGROUP -> "<id> { <recurse> }" through matching EGROUP
`cmd/conformance/core.lua` threads `req.print_unknown_fields` into
`pb.text.encode` so `_Drop` tests drop unknowns and `_Print` tests
render them.
Conformance: text-format suite goes from 2 ✓ / 6 expected fails to
8 ✓ / 0 expected fails. All eight regression tests in
`conformance_test.lua` (one per upstream test, plus the fixed field-1011
tag bytes that were miscomputed earlier) now assert the target output.
text: wire pb.text into conformance runner TEXT_FORMAT output
The runner short-circuited every TEXT_FORMAT request to `skipped`, even
though pb.text.encode has been encode-capable since M7. Plug it into
cmd/conformance/core.lua so protobuf/JSON input → text output exercises
the existing encoder end-to-end. Text-format input remains deferred
(pb.text is encode-only).
Text-format suite: 0 ✓ / 430 skipped / 4 expected fails →
2 ✓ / 426 skipped / 6 expected fails (Scalar/Message
*_Print added — unknown-field rendering still missing).
codegen+wire: split length-prefix emission + 2-byte varint fast path
Two complementary encode optimizations: codegen + runtime writers no
longer pay a per-field string concat for length-prefixed fields, and
encode_varint_slow now handles 2/3/4-byte values without dropping into
the uint64 cdata path.
Composite bench/bench.lua, full mode:
1KB encode: 229 K -> 332 K msgs/s (1.45x)
10KB encode: 38 K -> 56.5 K (1.49x)
100KB encode: 4.1 K -> 6.7 K (1.64x)
Runtime mode now matches full mode at 10KB+:
1KB encode: 203 K -> 276 K (1.36x)
10KB encode: 35 K -> 55 K (1.56x)
bench/shapes_bench.lua biggest swings (both modes):
packed-int32x1000 enc: 1.7 K -> 13 K (7.0-7.6x)
packed-int32x100 enc: 21 K -> 113 K (5.3x)
scalar-heavy enc: 582 K -> 1.0 M (1.7x)
Per-helper, bench/wire_bench.lua:
encode_varint(150) [2-byte]: 529 ns -> 68 ns (7.8x)
encode_varint(20K) [3-byte]: 841 ns -> 88 ns (9.6x)
encode_tag(16, LEN) [2-byte]: 531 ns -> 78 ns (6.8x)
encode_string(200B) [2-byte L]: 560 ns -> 109 ns (5.1x)
encode_varint(127) [1-byte]: 29 ns -> 28 ns (no regression)
1. Split length-prefix emission. Length-delimited fields (string/bytes
scalars, nested messages, packed scalars/enums) used to emit
`tag → encode_len(body)` where `encode_len(body)` returns
`encode_varint(#body) .. body`. The string concat allocated a copy
of the body per field. Now the codegen and runtime writers emit
three separate `out` slots — `tag`, `varint(#body)`, `body` — and
let `table.concat(out)` join them in one pass at the end of encode.
Applied to inline.go (full-mode codegen) for: singular and repeated
nested messages, singular and repeated string/bytes scalars, packed
scalars, packed enums. Applied to codec.lua build_writer / build_
repeated_writer for the same set. Maps still go through encode_len
pending a separate pass.
2. encode_varint_slow Lua-number fast paths. The outer encode_varint
stays at the tiny `1-byte check + tail call` shape that LuaJIT
inlines into hot traces. The slow function (not inlined into hot
traces, so its body size is unconstrained) now handles non-negative
Lua numbers up to 2^28 directly via bit.rshift / string.char with no
cdata allocation. Values in [2^28, 2^53) emit one byte and recurse
on the smaller residue. Only cdata inputs, negative Lua numbers
(sign-extended to 10-byte varint), and the rare > 2^53 case still
take the uint64 cdata loop. Net effect: every multi-byte varint
encode that fits in a Lua number — including the length prefix for
any string >= 128 bytes, every tag for field IDs >= 16, and every
negative-zigzag sint — drops from ~500 ns to ~70 ns.
497/497 luatest pass. 23/23 jit-trace gates pass (the two new
multi-byte varint gates added in the bench infra commit confirm
encode_varint_slow JIT-compiles cleanly). Conformance suite (binary
+ text) shows 1478 expected passes, 0 unexpected failures.
json: enable conformance JSON output + close 459 tests
Three coupled changes that turn the JSON output path on for the
conformance harness:
1. string_to_int64 accepts cdata: Tarantool's json.decode parses raw
JSON integer literals outside double range as int64_t/uint64_t cdata,
not Lua numbers. The decoder errored "expected JSON string or number
for int64" on any unquoted 64-bit value (Int64FieldMaxValueNotQuoted
et al). cdata is now cast through directly, preserving precision.
2. encode_message marks output as a map: an empty proto3 message
serialized as `[]` because Tarantool's json defaults empty tables to
array shape. jsoncpp's strict comparator threw Json::LogicError and
aborted the whole suite. Setting __serialize='map' on the output
gives `{}` and unblocks all JsonOutput tests.
3. PB_CONFORMANCE_SKIP_JSON gate is opt-in by default. core.lua now
matches exactly "1" (so docker -e VAR= disables it), and the
Dockerfile no longer hard-codes "=1" — JSON output runs end-to-end
for everyone unless they re-enable the gate.
Conformance moves from 930 / 1869 / 11 to 1389 / 1313 / 79
(successes / skipped / expected fails). The 75 new expected fails
are canonical-form edge cases (Duration formatting sign handling,
Timestamp out-of-range rejection, double precision digits, NaN
canonicalization, JSON-input strict rejection) — left for a follow-up.
Drops 7 entries from test/conformance/known_failures.txt that this
change closes; adds 75 newly-visible ones.
json: treat null fields as absent (and Value's null as a real value)
Per the proto3 JSON spec, a null on any field means "use the field's
default" — encoded as missing — with the lone exception of
google.protobuf.Value, where JSON null is itself a Value carrying
NullValue.NULL_VALUE.
Three coupled bugs surfaced together:
1. decode_field_value used to fall through with v = box.NULL, leaving a
useless box.NULL sitting in the result table for scalars. Now it
returns nil for non-Value fields, PB_NULL for Value fields.
2. decode_message's repeated and map branches called `#jv` and
`pairs(jv)` unconditionally; a JSON-null on either type crashed with
"attempt to get length of 'void *'". Now both branches short-circuit
when jv is box.NULL.
3. The nil-skip checks in the decode loop (`if dv ~= nil`) and in the
codec / inline message encoders (`if v == nil then return end`)
evaluated TRUE on box.NULL because Tarantool's cdata __eq aliases
it to nil. Decode now uses rawequal(dv, nil); encode special-cases
message kind by also accepting cdata, so the Value field's
box.NULL sentinel survives all the way through to value_encode.
Drops 3 entries from test/conformance/known_failures.txt
(AllFieldAcceptNull, WrapperTypesWithNullValue, ValueAcceptNull).
Adds 5 regression tests covering scalar / repeated / map / wrapper
null treatment and the Value-NULL_VALUE exception.
codec: recursively merge repeated singular-message wire entries
Per proto3 spec, when the same singular message field (including a oneof
branch) appears twice on the wire, the two values must merge: scalar
fields last-wins, repeated fields concatenate, sub-messages merge
recursively, maps last-wins per key. The previous reader replaced the
prev value wholesale for oneof branches and overwrote repeated/nested
fields with `prev[k] = v` even outside oneofs, losing data unique to
the first occurrence.
Adds wire.codec.merge_message(desc, prev, decoded), a descriptor-driven
recursive merge, and routes both codec.lua's reader and the inline-mode
codegen through it. WKT message fields (custom decode) keep the replace
behavior because their decoded value is not a generic Lua table.
Exposes pb.codec to the generated inline code so the helper is reachable
without a per-call require.
Drops 3 entries from test/conformance/known_failures.txt:
ValidDataOneof.MESSAGE.Merge, ValidDataOneofBinary.MESSAGE.Merge,
RepeatedScalarMessageMerge. Adds 5 regression tests covering scalar
last-wins, oneof merge, recursive sub-message merge, repeated-in-
submessage concat, and oneof sibling clearing after merge.
wkt: auto-register WKT descriptors so Any JSON decode resolves @type
pb.wkt exported Timestamp/Duration/Empty/FieldMask/Any/Struct/Value/
ListValue and the nine Wrapper descriptors, but the REGISTRY they live in
was empty until callers manually invoked pb.register. json_to_any looked
up @type in that empty registry, fell through to the opaque base64
fallback, and errored on every Any-of-WKT JSON payload.
Iterates M for every *_descriptor entry at module load and self-registers
it. Also calls pb.register on TestAllTypesProto3 in the conformance runner
so Any tests that embed the user message type also resolve.
Drops 10 entries from test/conformance/known_failures.txt.
wire: truncate varints to 32 bits in int32/uint32/sint32/enum decode
decode_int32/uint32/sint32 returned full uint64 values when the wire input
carried bits above bit 31, corrupting re-encode on 51 conformance tests that
exercise overlong or over-range varints. Per the proto3 spec the decoder
must keep only the low 32 bits (and sign-extend for signed types).
Adds wire.varint_to_int32 / varint_to_uint32 and routes every enum-varint
decode site through them: wire.lua typed decoders, codec.lua (5 enum
sites), lazy.lua (3 sites), and the generated code via inline.go (3 sites).
Drops 51 entries from test/conformance/known_failures.txt.
conformance: local Docker pipeline + cdata int64 map dedup
Wires up the Google protobuf conformance harness as a local target.
docker/conformance.Dockerfile builds conformance_test_runner from
upstream protobuf v34.1 source (matching the host's libprotoc 34.1)
and bundles Tarantool 3 from the official installer. `just conformance`
regenerates Lua, then runs the harness against cmd/conformance-runner.lua
with the repo mounted as a volume.
Six bugs surfaced and got fixed on the way to green:
1. conformance_test_runner uses execv (not execvp): bare `tarantool`
hits ENOENT. Pass /usr/bin/tarantool in CMD and Justfile.
2. The harness strips LUA_PATH from the child: the runner now
self-bootstraps package.path from debug.getinfo(1, 'S').source.
3. C-stdio buffering on pipe stdin made io.stdin:read(n) wait for a
full BUFSIZ before returning, deadlocking against the parent.
setvbuf('no') on stdin/stdout.
4. v34.1 fetches libjsoncpp via CMake FetchContent under
_deps/jsoncpp-build/...; the runtime image now COPYs the matching
.so* and runs ldconfig.
5. The harness's strict jsoncpp comparator crashes on our currently-
imperfect JSON output (enum numerics, map<K,V> shape, oneof
object form). Gate JSON output behind PB_CONFORMANCE_SKIP_JSON=1,
set in the container ENV; host-side `make test` still exercises
the full JSON path.
6. Codec bug — LuaJIT hashes cdata int64 by pointer, so duplicate-
key map entries (per proto3's "last value wins" semantics) split
across hash buckets even though __eq matches. Codec walks the
map once on insert to find a canonical key, gated by a
precomputed `f.key_dedup` flag so the dedup only fires for
int64/uint64/sint64/fixed64/sfixed64 keys. inline.go emits the
same `for _k in pairs(map) do` walk only when the static key
kind is 64-bit, so string/int32-keyed map decode stays
JIT-traceable.
Watchlists at test/conformance/known_failures.txt (binary + JSON) and
test/conformance/known_failures_text.txt (text-format) hold the
deferred failures. Current baseline:
- Binary + JSON suite: 803 ✓ / 1864 skipped / 139 expected fails
- Text-format suite: 0 ✓ / 430 skipped / 4 expected fails
403/403 luatest green, 19/19 jit-trace gate green.
codegen: protoc-gen-tarantool-doc — Markdown reference plugin
Sibling Go plugin under cmd/protoc-gen-tarantool-doc that emits one
Markdown file per input .proto. Sections (omitted when empty):
- Header (path, package, imports)
- Messages (per-message description + field table:
# | Field | Type | Label | Description)
- Enums (value table)
- Services (method table with unary/client/server/bidi label)
Field type cells render scalar names, full type names for message/enum
references, and `map<K, V>` for maps. Synthetic map-entry messages are
skipped. Leading comments preserved via SourceCodeInfo (squashed to a
single line inside table cells).
Built via `make build-doc`; sample output committed at
examples/docs/hello.md via `make gen-docs`. Smoke tests in
test/doc_test.lua build the plugin if needed and assert the expected
sections and labels appear. 403/403 luatest green.