Leia

Leia Performance And Benchmarks

Leia keeps performance evidence in the repository instead of treating it as an external benchmark notebook. The benchmark harnesses are intended for local optimization, public validation, and regression diagnosis.

Benchmark Layout

Benchmarks live under benchmarks/<domain>/<name>.leia.

Domain Focus
numeric Arithmetic kernels, matrices, nbody, spectral norm.
recursion Recursive calls and recursive data structures.
table Arrays, maps, traversal, sort, and metamethod tables.
calls Closures, varargs, coroutine calls, and method dispatch.
string String, pattern, regexp, UTF-8, and tokenization workloads.
concurrency Goroutine-like tasks, channels, select, cancellation, and sync.
data Data-oriented and SoA kernels.
app Larger application-shaped workloads.
control Control flow and protected execution.

Benchmark selectors use the domain ID, for example numeric/matmul or table/json_table_walk. Historical selector families are not accepted.

LuaJIT reference programs, where meaningful, live under benchmarks/lua_ref/<domain>/.

Commands

The CLI runs the Go benchmark harnesses:

leia bench --quick
leia bench numeric/matmul --runs 5 --warmup 1
leia bench --full --runs 5 --warmup 1 --sort luajit-gap
leia bench strict --runs 3 --warmup 1 \
  --json /tmp/leia_strict.json \
  --markdown /tmp/leia_strict.md
leia bench diagnose --bench table/table_array_access --out-dir /tmp/leia-diag

--sort luajit-gap ranks compare reports by the largest current/LuaJIT median ratio.

The shell gate wraps the same harnesses with repository defaults:

scripts/run.sh perf --syntax-smoke --no-luajit
scripts/run.sh perf --smoke
scripts/run.sh perf --feature-smoke
scripts/run.sh perf --full

Use --no-luajit when LuaJIT is not installed or the workload has no useful Lua reference. Without --no-luajit, script-timed current/LuaJIT rows are validated against --luajit-threshold (default 0.80). The short --smoke profile uses 0.85 to avoid treating 1-2 sample timing jitter as release evidence.

For fast local checks, use leia bench --quick. For release-quality evidence, write JSON and Markdown reports from the full and strict benchmark commands. --jobs N runs different benchmark selectors concurrently while preserving deterministic report row order. Use it for throughput-oriented local sweeps; use --jobs 1 when investigating a single noisy timing row.

Timing Modes

leia bench compare is the main optimization harness. It compares:

It reports median time, coefficient of variation, repeat count, current-vs-HEAD ratio, LuaJIT ratio, Tier 2 attempted/entered/failed counters, and total exits. The same counters are preserved in leia bench triage JSON and Markdown rows.

leia bench diagnose records single-benchmark evidence under an artifact directory. Its JSON rows include runtime_summary and tier2_call_summary objects so automation can distinguish a clean baseline from Tier 2 failures, runtime exits, skipped timing, and unavailable benchmark inputs. Optional evidence requests are reported with separate *_requested and *_effective fields. --warm-dump collects production Tier 2 warm-dump artifacts under the benchmark artifact directory; pprof requests remain explicit because JIT CPU profile collection is blocked by the CLI -cpuprofile/JIT guard until a JIT-safe profile path is wired. Requested pprof controls such as --pprof-min-samples-ms and --pprof-max-runs are carried into the pprof evidence summary.

Low-resolution script timers are not treated as wins. The harness can increase repeat counts and fall back to repeated command wall time when the script time is below the timer resolution.

Use --scale=group/name:KEY=VALUE[,KEY=VALUE...] to run the same benchmark shape with larger or smaller top-level numeric constants. The harness writes temporary scaled Leia sources, applies the same override to the clean --head-ref snapshot, and records the applied scale in the JSON report. --scale-profile NAME reads recommended scales from benchmarks/manifest.json when a workload declares them.

Strict Guard

Strict benchmark mode is the regression truth pass. It runs selected benchmarks in these modes:

Mode Meaning
vm Bytecode VM without JIT.
default Normal execution with configured accelerators.
no_filter JIT-enabled path with optimization filters disabled where supported.
luajit External LuaJIT reference when available.

The strict pass checks output stability, timing quality, suspicious benchmark-only wins, and LuaJIT comparisons where references exist. Suspicious strict thresholds are recorded in JSON and emitted as report warnings; release failure policy remains in the outer gate.

Execution Performance Contract

The interpreter is the semantic baseline. The bytecode VM and ARM64 JIT are execution accelerators, not a promise that every language feature or function runs as native code: supported hot paths may run natively, but unsupported operations must fall back to the VM/runtime without changing visible results, errors, capability checks, resource-budget behavior, or deoptimization behavior.

The LuaJIT comparison is a validation baseline, not a marketing claim. For script-timed rows with a Lua reference, scripts/run.sh perf --full runs the performance submit guard and fails when current / LuaJIT exceeds the configured --luajit-threshold (default 0.80). Use --no-luajit only when LuaJIT is unavailable or the selected workload has no meaningful Lua reference, and record that limitation in release evidence.

Production and release plans keep this bottom line active through scripts/run.sh production --full --release-profile, go run ./cmd/leia ci release --list, and scripts/run.sh perf --full.

Artifacts

Benchmark commands can write JSON and Markdown reports:

leia bench --full \
  --json /tmp/leia_timing.json \
  --markdown /tmp/leia_timing.md

leia bench strict \
  --json /tmp/leia_strict.json \
  --markdown /tmp/leia_strict.md

scripts/run.sh perf --validate-only FILE --json validates an existing timing report and exposes failure_kind_count, failure_kinds, failure_details, validate_target, and captured output_lines for release dashboards. validate_target records the checked timing report path, whether it exists, whether it is a file, and its byte size.

Release and diagnostic scripts collect these artifacts into evidence bundles. Do not commit temporary timing output. Only maintainers should update intentional baseline or history files under benchmarks/data/.

Adding Benchmarks

Add a benchmark when it covers a language feature, runtime mechanism, stdlib path, or user-facing workload shape that is not already represented.

Rules:

Use leia bench coverage to audit semantic-family performance coverage.