Runtime code generation, compose(), size and load time: the contract
Status: BINDING. This document supersedes every conflicting statement in docs/guide/modes.md, docs/guide/macro-mode.md, CHANGELOG.md and commit messages. Where another document disagrees with it, this document is right and the other is a defect to correct.
Recorded 2026-09-27 at the owner's direction, "so we have a record of how the LLM fucked up and how we are correcting it". Owner rulings are quoted verbatim. An agent may not narrow, reinterpret or close any rule here. If a rule looks wrong, escalate to the owner.
Tracking: jesscss/jess#322 (size), jesscss/jess#325 (runtime compose()).
The rules
1. No runtime code generation outside compile()
"there should be no fuckin new Function shit outside of compile() code"
new Function,eval, indirectevaland any other string-to-code step may run at runtime only insidecompile()'s live specialisation. Today that issrc/table/assemble.ts, which catchesEvalErrorand falls back to the closure assembler under a Content Security Policy.- The build-time macro plugin is exempt, because it runs in the bundler, not in the shipped program.
- Every other path — the interpreter, runtime
compose(),composeLeaf(), loading table artifacts, the spec and railroad tooling — must never turn a string into code.
2. Runtime compose() is the pure interpreter path, with zero build cost
"compose() should take zero time to "build" in and of itself, if done properly, it should just e the pure interpreter path"
- At runtime,
compose()only links. It merges the rule maps (a later piece's name wins), binds the composing grammar's trivia, resolvesg.Xreferences, and returns a parser that runs on the interpreter. - It never lowers, fuses, builds a
TableProgram, evaluates source or specialises. - It must not mutate shared rules. Composing trivia is bound per composed grammar, not stamped onto a shared rule's
_meta. - Its cost is about that of merging the maps: well under a millisecond for a real-scale grammar.
- A caller who wants speed opts in with
compile(composed).
3. The three modes, and the browser
Parseman has three modes (modes.md): the interpreter, the build-time macro, and runtime compile().
compile()IS "compile in the browser if available". It specialises once when permitted and falls back to closures under CSP.- A browser consumer ships the small interpreter grammar and runs it through
compile()when it wants the fast form. - A grammar package ships two builds: the macro-lowered one, and a plain
.ts→.jsone. Owner: "each parser is making a lowered macro build and a non-lowered .ts -> .js file so that either/both are composed for downstream grammars so that the less-preview parser build is not downloading a gargantuan thing". - A downstream grammar composes over the matching upstream build: compiled over compiled, interpreter over interpreter. Neither graph pulls in the other.
4. Size: 5× generated bytes per byte of grammar source, on real grammars
"I asked the LLM to MAYBE find instances where it could balance some things with inline generation vs table, but not a whole duplciate build-time materialization of everything!!!!" · "25-30x far far far exceeds the promises made on the parseman docs" · "no, inside a 5x budget"
- Every artifact variant of a real consumer grammar (jess's css, less, scss and jess grammars) stays within 5× raw generated bytes per byte of its grammar source.
- Inline or assembly code is generated only for specific regions where a steady-state measurement shows a real win per byte. Each region pays for its bytes, and nothing is materialised for the whole grammar.
- The size gate measures those real grammars, from vendored snapshots, and fails the build over budget. A gate that only flags growth past its own baseline does not satisfy this rule.
5. Load time is measured at real scale, against an absolute budget
The owner called the missing load benchmark a flaw.
A gated benchmark over the same real grammars measures:
- import and first-parse cost of the macro artifact;
- runtime
compose()cost of the interpreter grammar, including composing over a base; compile()cost, and the CSP-fallback path.
Each has an absolute per-grammar budget, and CI fails on regression.
The numeric budgets don't exist yet: this rule is unmet (see "Correction status"). The change that adds the gate records each budget in this section, together with the CI configuration that enforces it.
6. No performance regression
"please please just fix all of this without ALSO destroying parseman performance"
- Nothing here may make parsing slower. That means steady-state marginal instructions per parse after warmup,
(I(N2) − I(N1)) / (N2 − N1), for every real grammar in AST and CST mode, and cold start as well. - Each release is faster than the last.
- A trade of speed for size needs the owner's sign-off on the specific number.
7. Every documented claim has a test
Each failure below was a claim written into a doc, a changelog or a commit message with no test behind it. A behavioural claim in docs/guide/ (CSP safety, size ceilings, "build-time only", "follows the same rule") must name the test that proves it. A claim with no test is removed until it has one.
What went wrong
| # | Claim that shipped | What was actually true | Why nothing caught it |
|---|---|---|---|
| A | 0.48.0 CHANGELOG (2026-08-14): precompile the terminal composeLeaf AST/no-lines assembly once it reaches 1,024 instruction words, "and the deterministic size ceiling stays green" | It generated the whole grammar a second time as JavaScript. Jess's css grammar ast.js reached 1.8–2.24 MB, 13–14× its source, 80% of it that assembly. It does buy ~1.5× steady-state css parsing: 8.6M vs 13.0M instructions per parse. | The size gate covered only parseman's own probes (largest 11 KB), and none reaches the 1,024-word threshold. |
| B | macro-mode.md: "The ceiling is 10× raw bytes. It's enforced on every PR … and it can't be waived by rebaselining." | Since 0.45 the gate only blocks a fixture that grows past its own baseline (bench/size-guard.ts:103, 525–540). The 10× ceiling is reported, not enforced. | The claim had no test. |
| C | Commit cfa50d7 (2026-07-06), "carry compact IR instead of lowered rule source": "Build-time only (perf-free)" | Runtime compose() rebuilds every carried piece from IR source text with eval / new Function. That is evalRuleMapIR, src/compiler/ir-serialize.ts:158–270, called from src/compiler/linker.ts:318–324 and :511, plus a buildSrc eval at linker.ts:301. It even prefers the text when the live map exists ("IR FIRST: re-evaluating it yields FRESH combinators, so seeding composing trivia onto their _meta cannot leak back"): a shared-state mutation was "fixed" by evaluating source. Measured: 2.4 s to compose the Less grammar in Chromium, 2,825 ms total script load against 50 ms compiled. | No test ran runtime compose() at real scale or timed it. |
| D | modes.md CSP warning: "Runtime compose() follows the same rule" (catch EvalError, fall back) | Runtime compose() throws EvalError under CSP with no fallback, so the interpreter build fails behind a strict CSP. | test/unit/no-function-constructor.test.ts covers the macro path and compile() only, never runtime compose(). |
| E | An agent's report (2026-09-27): parseman "can't compile in the browser"; runtime artifacts "permanently disable" code generation | False. A live compile() parser specialises with new Function (src/table/compile.ts:158–169 → tableRules(prog) → src/table/assemble.ts:3801) and falls back under CSP. The comment it misread (src/table/program.ts:176–197) is about printed artifacts only. | A subagent's interpretation was repeated without checking the source. |
Superseded statements
These statements are wrong until the fixes land, and they must be corrected in the same change that fixes the behaviour:
docs/guide/modes.md:- the mode table's "a large terminal
composeLeafalso embeds one strict assembly"; - "A big enough terminal
composeLeaf()gets one extra thing"; - the CSP warning's "Runtime
compose()follows the same rule"; - "Terminal large
composeLeaf()artifacts carry one ordinary function literal".
- the mode table's "a large terminal
docs/guide/macro-mode.md, "The budget": the 10× ceiling is replaced by rule 4, and "enforced on every PR" is false (row B).docs/guide/extending.md, "How this behaves in each execution mode": runtimecompose()"fuses … using the same code generationcompile()uses" and "falls back to the closure assembler". Both break rules 1 and 2, and the fallback claim is false (row D).docs/guide/performance.md: "a large enough terminal (acomposeLeaf) it may also materialize one strict assembly" (row A).docs/reference/api.md: "Once its canonical table reaches 1,024 instruction words, the default AST/no-lines leaf …" (row A).CHANGELOG.md0.48.0: "the deterministic size ceiling stays green" (row A).- Commit
cfa50d7: "Build-time only (perf-free)" (row C).
Correction status
| Rule | Work | Tracking | Status |
|---|---|---|---|
| 1, 2 | Runtime compose() links live pieces. Remove IR eval and the buildSrc eval from the runtime path. Bind composing trivia without mutating shared rules. Remove railroad.ts's new Function. Add a test that runs compose() under --disallow-code-generation-from-strings. | jess#325 | in progress |
| 4 | Remove the whole-grammar assembly. Add selective inline regions within 5× that keep most of the css win. Vendor the real grammars into a size gate that fails. | jess#322 | in progress |
| 5 | Real-scale load benchmark with absolute budgets, run in CI | jess#322 / jess#325 | in progress |
| 6 | Before/after steady-state, cold-start and browser-load tables for each grammar and mode in the release PR | release PR | in progress |
| 7 | Every modes.md / macro-mode.md claim names its test, or is removed | release PR | in progress |
When a row lands, update its status here with the PR link. Don't delete the "What went wrong" table; it is the record.
