8 Commits

Author SHA1 Message Date
Oleksandr Kozachuk b4a28f342d chore: main is 0.3.0-to-be
CI / check (push) Has been cancelled
2026-08-12 11:02:13 +02:00
Oleksandr Kozachuk 0a0f1e9e95 feat(core): >ORDER, -ORDER, VOCABULARY search-order extensions 2026-08-12 11:01:49 +02:00
Oleksandr Kozachuk 94a0566ce3 release: 0.2.9
CI / check (push) Has been cancelled
2026-08-11 17:24:04 +02:00
Oleksandr Kozachuk 35da69cf7b bench: CrossCalls lane, best-of sampling, 10ms sizes 2026-08-11 17:23:52 +02:00
Oleksandr Kozachuk 392f2d0136 fix(core): inline loop-free callees first so the loop guard can fire
The guard ran before inlining and only ever saw calls.
2026-08-11 17:23:38 +02:00
Oleksandr Kozachuk b1cc93edc6 fix(core): typed entry only for a self-recursive word
Non-recursive words paid an extra wrapper hop. Adds the WAFER_DUMP_WASM dump hook.
2026-08-11 17:23:25 +02:00
Oleksandr Kozachuk 4f96f8860a release: 0.2.8
CI / check (push) Has been cancelled
Ships the self-guard expansion, and corrects what the benchmark tables claim.
Measured with wafer, gforth and SwiftForth all native on x86-64 -- the macOS
sf64 build runs under Rosetta 2 and flatters us -- Fibonacci is 1.16x rather
than 0.83x, so sf64 still wins it and wafer takes the other four. README and
OPTIMIZATIONS now carry both tables.
2026-08-10 12:48:33 +02:00
Oleksandr Kozachuk e963e636d3 perf(core): test a recursive word's base case at the call site
CI / check (push) Has been cancelled
A recursive Forth word almost always opens with a guard that returns early,
so every leaf of the recursion costs a call whose whole body is that test.
`Call(self)` now compiles as `<guard> IF <what the guard returns> ELSE
Call(self) THEN`, which is what the callee would have done on entry anyway.
Half of fib's nodes are leaves: Fibonacci(25) 356 -> 237 us, 1.24x sf64 ->
0.83x, so all five benchmarks now beat it.

The guard runs twice along the recursive path, hence the bounds: at most six
effect-free operations, at most four call sites, never a tail call. WS-018.
2026-08-09 18:27:25 +02:00
14 changed files with 1061 additions and 129 deletions
+120
View File
@@ -5,6 +5,124 @@ All notable changes to WAFER are documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [Unreleased]
### Added
- **Search-order conveniences from common practice** (none are Forth 2012;
all three exist in gforth and friends, and the semantics were checked
against gforth 0.7.3):
- `>ORDER ( wid -- )` pushes a wordlist on top of the search order --
the word the standard forgot when ANS replaced named vocabularies
with anonymous wid handles and left `ALSO` with nothing to name.
- `-ORDER ( wid -- )` removes a wordlist from the search order wherever
it sits (VFX/MPE extension, the inverse of `>ORDER`).
- `VOCABULARY <name>` creates a named wordlist; executing the name
replaces the top of the search order, the same semantics the standard
gives `FORTH`. `ORDER` and `WORDS ALL` now print vocabulary names
instead of `wid#N`, and `MARKER` rollback forgets them along with
the words.
## [0.2.9] - 2026-08-10
### Fixed
- **A word that never recurses no longer gets a typed entry it cannot use.**
Every word with a statically known stack effect was given the typed
wrapper + fast-entry pair. In the JIT path the function-table slot holds
the wrapper and the only caller that can reach the fast entry is
`RECURSE`, so for any other word a cross-word call went
`call_indirect` -> wrapper -> fast entry: one hop more for exactly the
same memory traffic. On a 300k-iteration loop over a callee too big to
inline that cost **1569 µs against 1067 with the convention off** -- an
optimisation making things worse. It is now emitted only when the body
calls itself, which is where it is worth 4x (Fibonacci 242 µs typed
against 636 untyped). `CONSOLIDATE` and the AOT export are unaffected;
they solve their effects separately. Present in 0.2.7 and 0.2.8.
- **The inliner's loop guard has never actually fired.** 0.2.7 added a rule
that a loop-bearing callee must not be inlined into a caller that can
never be promoted, since the loop then loses its registers -- a 7x
pessimisation applied by an optimisation pass. The check ran _before_
inlining, where the caller is nothing but calls: `DROP` is
`Call(WordId(2))`, `CR` is `Call(WordId(38))`. Since the check looks
through calls by design, it called nearly every caller promotable and
the guard did nothing. Inlining now happens in two passes -- loop-free
callees first, then the question, then the rest.
### Added
- **A sixth benchmark, `CrossCalls(300K)`, that measures what `CONSOLIDATE`
does.** The other five have no cross-word call left in their hot loop:
four have their callee inlined away and Fibonacci is self-recursive. So
the `CONSOL` column measured nothing, which is how both bugs above stayed
hidden. With a real call in the loop, consolidation is worth 2.8-4x.
### Changed
- **The benchmark harness stops reporting noise.** It took the median of three
timed repetitions inside one process, and a `samples` field that was never
read. Each measurement is now the mean of the three fastest of seven
repetitions, and that whole process runs three times with the fastest kept.
Benchmark noise is one-sided -- a scheduling hiccup or a busy SMT sibling can
only make a run slower -- so the fastest runs are the honest ones, and only a
fresh process resamples core placement and code layout. On a shared 16-vCPU
box the run-to-run spread went from 20-79% to 1-6%, and Fibonacci after
`CONSOLIDATE` stopped being bimodal (413-419 µs on three reports and 712-770
on two, with nothing in between; now 412-426 across four).
- **Every benchmark is now sized to run about 10 ms**, from the 0.2-2 ms most
of them took. Not for the usual reason -- the timing wrapper already excludes
start-up and compilation, and in the measurements shorter benchmarks were if
anything the _steadier_ ones -- but it buys a comfortable margin over timer
resolution and first-iteration effects for nothing: the report still finishes
in under a minute, and gforth, 3-20x slower than WAFER, is what sets that
clock. Fibonacci went from 25 to 33 rather than into a loop, so it stays pure
recursion; Collatz repeats its 2000-value round 50 times instead of counting
higher, because past ~100000 the sequence peaks near 1.5 billion and `3 * 1+`
overflows WAFER's 32-bit cells while sf64's 64-bit cells carry on -- the two
engines would stop doing the same work. All three engines agree on the results
at the new sizes.
### Explained
- **Why `CONSOLIDATE` makes some promoted loops slower** (NestedLoops
1.7x on x86-64, 1.1x on arm64): not worse code -- the WASM is
byte-identical and the machine code instruction-identical modulo
registers -- but worse placement. A tight loop pays for straddling an
instruction-fetch window (16 bytes on the M1 at ~9%; 32 bytes on
Skylake at up to ~65%, where a fused `cmp+jcc` crossing the boundary
drops the loop out of the uop cache every iteration -- the JCC
erratum). Cranelift never aligns loop headers, and the per-word JIT
module's dead dsp-prologue bytes happen to shift its loops onto
luckier offsets. Verified by a padding sweep that reproduces the full
penalty range on both hosts, including placements where consolidated
code beats the JIT. Details in docs/OPTIMIZATIONS.md; native x86-64
reference numbers in the README re-taken at the new workload sizes.
## [0.2.8] - 2026-08-10
### Added
- **A recursive word tests its base case at the call site.** A recursive Forth
word almost always opens with a guard that returns early --
`: FIB DUP 2 < IF EXIT THEN ... RECURSE ... ;` -- so every leaf of the
recursion costs a call whose entire body is that test. `Call(self)` now
compiles as `<guard> IF <what the guard returns> ELSE Call(self) THEN`,
which computes the same thing: the callee would have run the guard, taken
the branch and returned. In fib's tree the leaves are half of all nodes.
Fibonacci(25) 356 -> 237 µs on the arm64 development machine, where that
reads 1.24x -> 0.83x of SwiftForth `sf64`. Measured again with **both
engines native on x86-64** -- the macOS `sf64` build runs under Rosetta 2,
which flatters WAFER -- Fibonacci is 1.16x, so it remains the one benchmark
of the five that `sf64` wins. See the two tables in the README.
The guard runs twice along the recursive path, so it has to be small (at
most six operations) and free of effects -- no calls, no memory, no
branches. Words with more than four self-call sites are left alone to bound
the code growth, and a `TailCall` is never expanded.
## [0.2.7] - 2026-08-09
### Added
@@ -340,6 +458,8 @@ compliance suite, `CONSOLIDATE` whole-program recompilation, `wafer build`
AOT export (WASM / native / JS loader), browser REPL, SHA-1/256/512 words,
and cross-engine benchmark lanes against gforth and SwiftForth.
[0.2.9]: https://github.com/ok2/wafer/compare/v0.2.8...v0.2.9
[0.2.8]: https://github.com/ok2/wafer/compare/v0.2.7...v0.2.8
[0.2.7]: https://github.com/ok2/wafer/compare/v0.2.6...v0.2.7
[0.2.1]: https://github.com/ok2/wafer/compare/v0.2.0...v0.2.1
[0.2.0]: https://github.com/ok2/wafer/compare/v0.1.0...v0.2.0
+2 -2
View File
@@ -2,7 +2,7 @@
## What is WAFER?
WAFER (WebAssembly Forth Engine in Rust) is an optimizing Forth 2012 compiler targeting WebAssembly. Currently a working Forth system with 200+ words, JIT compilation, 12 word sets at 100% compliance, and a full optimization pipeline (peephole, constant folding, inlining, strength reduction, DCE, tail calls, per-region stack-to-local promotion with DO/BEGIN loop and IF support, self-recursive direct calls, a typed calling convention for words with a known stack effect, consolidation). Beats gforth on all benchmarks in release mode, and SwiftForth `sf64` on four of five. Includes a browser-based REPL via wasm-pack.
WAFER (WebAssembly Forth Engine in Rust) is an optimizing Forth 2012 compiler targeting WebAssembly. Currently a working Forth system with 200+ words, JIT compilation, 12 word sets at 100% compliance, and a full optimization pipeline (peephole, constant folding, inlining, strength reduction, DCE, tail calls, per-region stack-to-local promotion with DO/BEGIN loop and IF support, self-recursive direct calls, a typed calling convention for words with a known stack effect, self-guard expansion for recursive words, consolidation). Beats gforth on every benchmark, and SwiftForth `sf64` on five of six (measured native-vs-native on x86-64; the macOS sf64 build is x86-64 under Rosetta and flatters WAFER). Includes a browser-based REPL via wasm-pack.
## Architecture
@@ -79,7 +79,7 @@ Handle in `interpret_token_immediate()` or `compile_token()` as a special case.
## Testing
- Run `cargo test --workspace` before committing (currently 601 unit + 1 benchmark + 12 compliance + 9 comparison + 5 crypto)
- Run `cargo test --workspace` before committing (currently 611 unit + 1 benchmark + 12 compliance + 9 comparison + 5 crypto)
- Forth 2012 compliance: `cargo test -p wafer-core --test compliance`
- Cross-engine comparison (vs gforth): `cargo test -p wafer-core --test comparison`
- Performance benchmarks (release mode): `cargo test -p wafer-core --test comparison -- --nocapture --ignored`
Generated
+3 -3
View File
@@ -1589,7 +1589,7 @@ checksum = "0b928f33d975fc6ad9f86c8f283853ad26bdd5b10b7f1542aa2fa15e2289105a"
[[package]]
name = "wafer"
version = "0.2.7"
version = "0.3.0"
dependencies = [
"anyhow",
"clap",
@@ -1600,7 +1600,7 @@ dependencies = [
[[package]]
name = "wafer-core"
version = "0.2.7"
version = "0.3.0"
dependencies = [
"anyhow",
"insta",
@@ -1615,7 +1615,7 @@ dependencies = [
[[package]]
name = "wafer-web"
version = "0.2.7"
version = "0.3.0"
dependencies = [
"anyhow",
"js-sys",
+1 -1
View File
@@ -3,7 +3,7 @@ members = ["crates/*"]
resolver = "2"
[workspace.package]
version = "0.2.7"
version = "0.3.0"
edition = "2024"
license = "MIT OR Apache-2.0"
repository = "https://github.com/ok2/wafer"
+85 -30
View File
@@ -8,7 +8,7 @@ An optimizing Forth 2012 compiler targeting WebAssembly. WAFER JIT-compiles each
- **200+ words** across 12 Forth 2012 word sets, all at **100% compliance**
- **Optimizing compiler** with 6 IR passes + stack-to-local promotion (per region, so a hot loop keeps its registers even inside a word that does I/O; `DO` and `BEGIN` loops alike) + consolidation
- **Faster than gforth** on all benchmarks in release mode (2-10x faster)
- **Faster than gforth** on every benchmark, and past SwiftForth `sf64` -- a native-code compiler -- on five of six
- **JIT compilation** — each `:` definition compiles to its own WASM module
- **Self-recursive direct calls** — RECURSE compiles to native `call` instead of `call_indirect`
- **Typed calling convention** — a word with a statically known stack effect passes its stack items as WASM values, so a call keeps them in registers instead of round-tripping through memory
@@ -80,51 +80,106 @@ git submodule update --init
## Performance
WAFER beats gforth (the GNU Forth reference implementation) on all benchmarks in release mode, and is within
reach of SwiftForth `sf64`, which compiles to native code:
WAFER beats gforth (the GNU Forth reference implementation) on every benchmark by 3-20x, and
SwiftForth `sf64` -- which compiles to native code -- on five of the six. Fibonacci is the one it
loses: one call per node, no loop to promote, and `sf64` keeps its stack in registers across a call
the way only a native code generator can.
Measured on the development machine (M1 Ultra, arm64), median of three reports:
```
Benchmark WAFER CONSOL gforth sf64 WAFER/gf WAFER/sf
Fibonacci(25) 356 361 3389 287 0.11x 1.24x
Factorial(12)x100K 479 495 6249 1650 0.08x 0.29x
GCD-bench(20K) 540 559 1801 801 0.30x 0.67x
NestedLoops(50)x1K 509 501 7023 1887 0.07x 0.27x
Collatz(2K) 185 213 3873 610 0.05x 0.30x
Fibonacci(33) 11307 11407 157001 13053 0.07x 0.87x
Factorial(12)x2M 9639 9599 123950 32091 0.08x 0.30x
GCD-bench(400K) 11662 11580 38580 17001 0.30x 0.68x
NestedLoops(50)x20K 8920 9852 140518 36828 0.06x 0.24x
CrossCalls(3M) 10883 3769 87691 8240 0.04x 0.46x
Collatz(2K)x50 8838 8715 189903 28657 0.05x 0.30x
```
Times in microseconds. WAFER/gf < 1.0 means WAFER is faster. CONSOL = after `CONSOLIDATE`.
Times in microseconds; the ratios use the better of `WAFER` and `CONSOL`. Below 1.0 means WAFER is
faster.
Two caveats on the `sf64` column. The SwiftForth build here is x86-64 running under Rosetta 2
while WAFER and gforth are native arm64, so it is a native-vs-emulated comparison; and sf64
uses 64-bit cells to WAFER's 32-bit. WAFER is ahead on the four loop-heavy benchmarks and
behind on Fibonacci, which is one call per node with no loop to promote.
**The `sf64` column here flatters WAFER, and by enough to change an answer.** The only SwiftForth
build for macOS is x86-64 running under Rosetta 2, while WAFER and gforth are native arm64 -- so
that column compares native code against emulated code, and the penalty falls hardest on the
call-heavy benchmark. Measured with all three engines native on x86-64 (Xeon Platinum 8124M,
Ubuntu 22.04; two reports agreed within 1%), Fibonacci reads **1.21x** where the table above says
0.87x; the other five keep their wins. That native comparison is what the "five of six" above
rests on:
```
Benchmark WAFER CONSOL gforth sf64 WAFER/gf WAFER/sf
Fibonacci(33) 19512 19511 129784 16076 0.15x 1.21x
Factorial(12)x2M 22532 16601 137168 57986 0.12x 0.29x
GCD-bench(400K) 34216 34089 66595 51680 0.51x 0.66x
NestedLoops(50)x20K 10729 17827 126687 40469 0.08x 0.27x
CrossCalls(3M) 20457 7412 81303 29264 0.09x 0.25x
Collatz(2K)x50 18686 17328 188592 80857 0.09x 0.21x
```
A second caveat holds on any host: `sf64` uses 64-bit cells to WAFER's 32-bit, so WAFER does less
work per operation.
`CrossCalls` is the only benchmark with a cross-word call left in its hot loop -- the other five
have their callee inlined away or are self-recursive -- so it is the only one that measures what
`CONSOLIDATE` does, and there it is worth 2.9x. `NestedLoops` goes the other way: `CONSOLIDATE`
makes it 1.1x _slower_ on the M1 and 1.7x on x86-64 -- not worse code but worse luck. Both paths
emit identical WASM for the hot word; the delta is where the machine code lands. A tight loop
pays for straddling an instruction-fetch window (16 bytes on the M1, 32 on Skylake, where a fused
branch crossing the boundary drops the loop out of the uop cache -- the JCC erratum), Cranelift
does not align loop headers, and dead prologue bytes in the per-word JIT module happen to shift
its loops into luckier spots. Details in
[docs/OPTIMIZATIONS.md](docs/OPTIMIZATIONS.md#8-consolidation).
Every benchmark is sized to run about 10 ms. Not for the usual reason -- the timing wrapper already
excludes start-up and compilation -- but to keep a comfortable margin over timer resolution and
first-iteration effects without pushing the report past a minute. gforth is 3-20x slower than
WAFER, so it sets the wall clock.
A word whose stack effect is statically known gets a **typed entry point**: its stack items travel in and out
as WASM values instead of through the memory data stack, so cranelift keeps them in registers across a call
the way a native Forth keeps TOS in one. The word also keeps a `( -- )` wrapper, which is what the function
table, `EXECUTE` and the outer interpreter reach, so nothing about the memory ABI changes from the outside.
Call-heavy code is what this pays for -- Fibonacci went from 4.3x slower than `sf64` to 1.2x. Set
`WAFER_TYPED_CALLS=0` to fall back to the memory-stack convention.
Only a caller inside the same module can use the fast entry -- `RECURSE` in the JIT path, every resolvable
call after `CONSOLIDATE` -- so that is exactly when it is emitted. Set `WAFER_TYPED_CALLS=0` to fall back.
Recursive words then get one more thing: their base-case guard is tested at the **call site**, so a
leaf of the recursion costs a comparison instead of a call. `: FIB DUP 2 < IF EXIT THEN ... RECURSE`
compiles its `RECURSE` as `DUP 2 < IF ELSE RECURSE THEN`, which is what the callee would have done
on entry anyway. Half of fib's nodes are leaves, and that is worth 1.4x.
## Testing
Everything below has a `just` target; the raw command is given where it is worth
knowing what the target does.
```bash
# All tests (~628 currently passing)
cargo test --workspace
# Forth 2012 compliance suite
cargo test -p wafer-core --test compliance
# Cross-engine comparison (WAFER vs gforth, requires gforth)
cargo test -p wafer-core --test comparison -- --nocapture --ignored
# Optimization benchmark report (WAFER-internal)
cargo test -p wafer-core --test benchmark_report -- --nocapture --ignored
# Lints
cargo clippy --workspace
just test # all tests (~638 currently passing)
just compliance # Forth 2012 compliance suite
just clippy # lints
just fmt # formatting check (Rust + Markdown)
just ci # everything CI runs
```
Benchmarks are separate, because they are `#[ignore]`d -- they take minutes, and
a debug build would measure nothing useful:
```bash
just bench-compare # WAFER vs gforth vs SwiftForth, the table in Performance
just bench-opts # WAFER against its own optimization settings
just bench # criterion micro-benchmarks
just compare-correctness # same three engines, compared on output instead of time
```
`bench-compare` needs `gforth` and `sf64` on `PATH` -- a missing engine drops its
column rather than failing. Each number in it is the best of three processes, and
each process reports the mean of its three fastest of seven timed repetitions:
benchmark noise is one-sided, so the fastest runs are the honest ones, and only a
fresh process resamples core placement and code layout. Run it on an idle
machine; a busy one produced 20-79% run-to-run spread where an idle one gives
1-6%.
## Architecture
```
@@ -142,7 +197,7 @@ Forth Source -> Outer Interpreter -> IR -> [Optimize] -> WASM Codegen (wasm-enco
- `WebRuntime` — browser WebAssembly API via js-sys, for the browser REPL
- **Subroutine threading** via WASM function tables (`call_indirect` for cross-word, direct `call` for self-recursion)
- **JIT mode**: each new word compiles to a separate WASM module linked to shared memory/globals/table
- **IR-based pipeline** with 6 optimization passes (peephole, constant folding, strength reduction, DCE, tail call detection, inlining) plus per-region stack-to-local promotion (DO and BEGIN loops, IF/ELSE), DO/LOOP index locals, typed entry points for words with a known stack effect, and consolidation
- **IR-based pipeline** with 6 optimization passes (peephole, constant folding, strength reduction, DCE, tail call detection, inlining) plus per-region stack-to-local promotion (DO and BEGIN loops, IF/ELSE), DO/LOOP index locals, typed entry points for words with a known stack effect, self-guard expansion, and consolidation
- **Dictionary**: linked-list word headers in simulated linear memory
## Project Structure
+1 -1
View File
@@ -9,7 +9,7 @@ license.workspace = true
workspace = true
[dependencies]
wafer-core = { path = "../core", version = "0.2.7" }
wafer-core = { path = "../core", version = "0.3.0" }
wasmtime = { workspace = true }
anyhow = { workspace = true }
clap = { version = "4", features = ["derive"] }
+141 -5
View File
@@ -1584,6 +1584,48 @@ fn is_promotable_body(ops: &[IrOp], mode: PromoteMode) -> bool {
true
}
/// Does `ops` call `target` at any nesting depth?
///
/// In the JIT path this decides whether a typed entry is worth emitting at
/// all: the table slot holds the `( -- )` wrapper, so the only caller that can
/// reach the fast entry is the word itself.
fn body_calls(ops: &[IrOp], target: WordId) -> bool {
ops.iter().any(|op| match op {
IrOp::Call(id) | IrOp::TailCall(id) => *id == target,
IrOp::If {
then_body,
else_body,
} => {
body_calls(then_body, target)
|| else_body
.as_deref()
.is_some_and(|eb| body_calls(eb, target))
}
IrOp::DoLoop { body, .. } | IrOp::BeginUntil { body } | IrOp::BeginAgain { body } => {
body_calls(body, target)
}
IrOp::BeginWhileRepeat { test, body } => {
body_calls(test, target) || body_calls(body, target)
}
IrOp::BeginDoubleWhileRepeat {
outer_test,
inner_test,
body,
after_repeat,
else_body,
} => {
body_calls(outer_test, target)
|| body_calls(inner_test, target)
|| body_calls(body, target)
|| body_calls(after_repeat, target)
|| else_body
.as_deref()
.is_some_and(|eb| body_calls(eb, target))
}
_ => false,
})
}
/// Does `ops` contain an `EXIT` at any nesting depth?
fn body_has_exit(ops: &[IrOp]) -> bool {
ops.iter().any(|op| match op {
@@ -3274,12 +3316,16 @@ pub fn compile_word(
let mut module = Module::new();
// A word whose stack effect is statically known gets a second, typed
// entry point; the self-recursive case is the one that pays, since the
// recursion then runs entirely in WASM values. Cross-word typed calls
// need every callee in the same module, which only CONSOLIDATE gives.
// entry point -- but only if it calls itself. Cross-word typed calls need
// every callee in the same module, which only CONSOLIDATE gives, so here
// the table slot holds the `( -- )` wrapper and no other word can reach
// the fast entry. Emitting the pair anyway just puts a wrapper hop in
// front of every call through the table, for the same memory traffic:
// measured at +47% on a 300k-iteration loop over a callee too big to
// inline. `RECURSE` is the one caller that does reach it, and there the
// convention is worth 4x.
let self_id = WordId(config.base_fn_index);
let typed = config
.typed_calls
let typed = (config.typed_calls && body_calls(body, self_id))
.then(|| typed_effect(body, Some(self_id), &HashMap::new()))
.flatten();
@@ -3483,6 +3529,23 @@ pub fn compile_word(
/// The name section carries the Forth word name into wasmtime trap
/// backtraces (best-effort symbolication, WS-008); a typed word names both
/// of its entries so the innermost frame is the one that reports.
/// Write a compiled module to `$WAFER_DUMP_WASM/<name>.wasm` when that
/// variable is set.
///
/// Reading the emitted code is the only way some questions get answered --
/// three separate investigations have needed this and re-added it by hand each
/// time, so it lives here now. `wasm-tools print` turns the output into wat.
fn maybe_dump(name: &str, bytes: &[u8]) {
let Ok(dir) = std::env::var("WAFER_DUMP_WASM") else {
return;
};
let safe: String = name
.chars()
.map(|c| if c.is_alphanumeric() { c } else { '_' })
.collect();
let _ = std::fs::write(format!("{dir}/{safe}.wasm"), bytes);
}
fn finish_word_module(
mut module: Module,
name: &str,
@@ -3504,6 +3567,7 @@ fn finish_word_module(
module.section(&names);
let bytes = module.finish();
maybe_dump(name, &bytes);
// Validate
wasmparser::validate(&bytes).map_err(|e| {
@@ -4113,6 +4177,7 @@ fn compile_multi_word_module(
}
let bytes = module.finish();
maybe_dump("CONSOLIDATED", &bytes);
// Validate
wasmparser::validate(&bytes)
@@ -5299,6 +5364,77 @@ mod tests {
]
}
#[test]
fn body_calls_finds_the_word_at_any_depth() {
let me = WordId(5);
assert!(body_calls(&fib_ir(me), me));
assert!(!body_calls(&fib_ir(me), WordId(6)));
assert!(!body_calls(&[IrOp::Dup, IrOp::Mul], me));
// Nested in every body-bearing op, since the JIT gate reads this to
// decide whether a typed entry can ever be reached.
assert!(body_calls(
&[IrOp::DoLoop {
body: vec![IrOp::Call(me)],
is_plus_loop: false,
}],
me
));
assert!(body_calls(
&[IrOp::If {
then_body: vec![IrOp::Dup],
else_body: Some(vec![IrOp::TailCall(me)]),
}],
me
));
assert!(body_calls(
&[IrOp::BeginWhileRepeat {
test: vec![IrOp::Dup],
body: vec![IrOp::Call(me)],
}],
me
));
assert!(body_calls(
&[IrOp::BeginUntil {
body: vec![IrOp::Call(me)],
}],
me
));
}
#[test]
fn jit_gives_a_typed_entry_only_to_a_self_recursive_word() {
// The table slot holds the `( -- )` wrapper, so nothing but the word
// itself can reach the fast entry. Emitting it for a word that never
// recurses just adds a wrapper hop to every call through the table --
// measured at +47% on a hot cross-word call.
let id = 5;
let cfg = |typed| CodegenConfig {
base_fn_index: id,
table_size: 256,
stack_to_local_promotion: true,
stack_guards: None,
typed_calls: typed,
};
// `DUP *` has a perfectly good effect ( n -- n ) and no self-call.
let square = [IrOp::Dup, IrOp::Mul];
let plain = compile_word("SQ", &square, &cfg(true)).expect("compiles");
let untyped = compile_word("SQ", &square, &cfg(false)).expect("compiles");
assert_eq!(
plain.bytes, untyped.bytes,
"a word that never calls itself must compile the same with typed calls on or off"
);
// FIB does recurse, so it keeps the pair and must differ.
let fib = fib_ir(WordId(id));
let typed = compile_word("FIB", &fib, &cfg(true)).expect("compiles");
let memory = compile_word("FIB", &fib, &cfg(false)).expect("compiles");
assert_ne!(
typed.bytes, memory.bytes,
"self-recursion should still get a fast entry"
);
}
#[test]
fn typed_effect_solves_self_recursion() {
// The recursion makes the equation circular (2d = d), so the
+2
View File
@@ -39,6 +39,7 @@ impl WaferConfig {
strength_reduce: true,
dce: true,
inline: true,
self_guard: true,
},
codegen: CodegenOpts {
stack_to_local_promotion: true,
@@ -58,6 +59,7 @@ impl WaferConfig {
strength_reduce: false,
dce: false,
inline: false,
self_guard: false,
},
codegen: CodegenOpts {
stack_to_local_promotion: false,
+323 -3
View File
@@ -27,6 +27,9 @@ pub struct OptConfig {
pub dce: bool,
/// Enable inlining of small word bodies.
pub inline: bool,
/// Expand a recursive word's base-case guard into its own call sites, so
/// the leaves of the recursion cost a test instead of a call.
pub self_guard: bool,
}
/// Run all enabled optimization passes.
@@ -34,6 +37,7 @@ pub fn optimize(
ops: Vec<IrOp>,
config: &OptConfig,
bodies: &HashMap<WordId, Vec<IrOp>>,
self_id: Option<WordId>,
) -> Vec<IrOp> {
let mut ir = ops;
@@ -57,9 +61,22 @@ pub fn optimize(
// inlined loop down with it, so leave those callees where they are:
// as their own word the loop keeps its registers, and one call is far
// cheaper than a loop's worth of memory traffic.
//
// This takes two passes, because before the primitives are substituted
// the caller is nothing but `Call`s -- `DROP` and `CR` included -- and
// the promotability check deliberately looks through calls. Asked too
// early it says "promotable" about almost anything, which is how this
// guard managed to be a no-op. Inline the loop-free callees first, then
// ask, then let the loop-bearing ones in if the answer was yes.
ir = inline(ir, bodies, 8, true);
let keep_loops_out = !crate::codegen::promotable_modulo_calls(&ir);
ir = inline(ir, bodies, 8, keep_loops_out);
}
if config.self_guard
&& let Some(id) = self_id
{
ir = expand_self_guard(ir, id);
}
if config.peephole {
ir = peephole(ir);
}
@@ -585,6 +602,142 @@ fn detailcall(op: IrOp) -> IrOp {
}
/// Check if an IR body contains a direct call to the given word (recursion guard).
/// Largest guard the expander is willing to run twice, in IR operations.
const MAX_GUARD_OPS: usize = 6;
/// Most self-call sites worth expanding, to bound the code growth.
const MAX_GUARD_SITES: usize = 4;
/// Expand a recursive word's base-case guard into its own call sites.
///
/// A recursive Forth word almost always opens with a guard that returns early
/// -- `: FIB DUP 2 < IF EXIT THEN ... RECURSE ... ;` -- so every leaf of the
/// recursion costs a call whose whole body is that test. Testing at the call
/// site instead removes the call for the leaves, which in fib's tree is half
/// of all nodes.
///
/// `Call(self)` becomes `<guard> IF <what the guard returns> ELSE Call(self)
/// THEN`, which computes the same thing: the callee would have run the guard,
/// taken the branch and returned. The price is that the guard runs twice along
/// the recursive path, which is why it has to be small and free of effects.
fn expand_self_guard(ops: Vec<IrOp>, self_id: WordId) -> Vec<IrOp> {
let Some((cond, base)) = split_guard(&ops) else {
return ops;
};
if count_self_calls(&ops, self_id) > MAX_GUARD_SITES {
return ops;
}
let (cond, base) = (cond.to_vec(), base.to_vec());
replace_self_calls(ops, self_id, &cond, &base)
}
/// Split a body into the condition of its leading base-case guard and what
/// that guard leaves behind, or `None` if it does not open with one.
fn split_guard(ops: &[IrOp]) -> Option<(&[IrOp], &[IrOp])> {
let at = ops.iter().position(|op| matches!(op, IrOp::If { .. }))?;
let cond = &ops[..at];
if at > MAX_GUARD_OPS || !cond.iter().all(is_duplicable) {
return None;
}
let IrOp::If {
then_body,
else_body: None,
} = &ops[at]
else {
return None;
};
// The guard is only a guard if it returns; what precedes the `EXIT` is
// the value it returns, and has to be as harmless as the condition.
let (IrOp::Exit, base) = then_body.split_last()? else {
return None;
};
if base.len() > MAX_GUARD_OPS || !base.iter().all(is_duplicable) {
return None;
}
Some((cond, base))
}
/// Can this operation be duplicated at every call site -- cheap, effect-free,
/// and not itself a call or a branch?
fn is_duplicable(op: &IrOp) -> bool {
matches!(
op,
IrOp::PushI32(_)
| IrOp::Drop
| IrOp::Dup
| IrOp::Swap
| IrOp::Over
| IrOp::Rot
| IrOp::Nip
| IrOp::Tuck
| IrOp::TwoDup
| IrOp::TwoDrop
| IrOp::Add
| IrOp::Sub
| IrOp::Mul
| IrOp::Negate
| IrOp::Abs
| IrOp::Eq
| IrOp::NotEq
| IrOp::Lt
| IrOp::Gt
| IrOp::LtUnsigned
| IrOp::ZeroEq
| IrOp::ZeroLt
| IrOp::And
| IrOp::Or
| IrOp::Xor
| IrOp::Invert
| IrOp::Lshift
| IrOp::Rshift
| IrOp::ArithRshift
)
}
fn count_self_calls(ops: &[IrOp], self_id: WordId) -> usize {
ops.iter()
.map(|op| match op {
IrOp::Call(id) if *id == self_id => 1,
IrOp::If {
then_body,
else_body,
} => {
count_self_calls(then_body, self_id)
+ else_body
.as_deref()
.map_or(0, |eb| count_self_calls(eb, self_id))
}
_ => 0,
})
.sum()
}
/// Wrap every `Call(self_id)` in the guard. Only plain calls: a `TailCall` is
/// followed by a return, and leaving those alone keeps tail-call detection and
/// this pass from having to agree about what tail position means.
fn replace_self_calls(ops: Vec<IrOp>, self_id: WordId, cond: &[IrOp], base: &[IrOp]) -> Vec<IrOp> {
let mut out = Vec::with_capacity(ops.len());
for op in ops {
match op {
IrOp::Call(id) if id == self_id => {
out.extend_from_slice(cond);
out.push(IrOp::If {
then_body: base.to_vec(),
else_body: Some(vec![IrOp::Call(id)]),
});
}
IrOp::If {
then_body,
else_body,
} => out.push(IrOp::If {
then_body: replace_self_calls(then_body, self_id, cond, base),
else_body: else_body.map(|eb| replace_self_calls(eb, self_id, cond, base)),
}),
other => out.push(other),
}
}
out
}
fn contains_call_to(ops: &[IrOp], target: WordId) -> bool {
for op in ops {
match op {
@@ -746,8 +899,173 @@ mod tests {
strength_reduce: true,
dce: true,
inline: false,
self_guard: false,
};
optimize(ops, &config, &HashMap::new())
optimize(ops, &config, &HashMap::new(), None)
}
/// A body shaped like a recursive Forth word: a base-case guard, then the
/// recursive step. `SELF` is the word being compiled.
const SELF: WordId = WordId(9);
fn guarded_body(step: Vec<IrOp>) -> Vec<IrOp> {
let mut ops = vec![
IrOp::Dup,
IrOp::PushI32(2),
IrOp::Lt,
IrOp::If {
then_body: vec![IrOp::Exit],
else_body: None,
},
];
ops.extend(step);
ops
}
#[test]
fn self_guard_moves_the_base_case_to_the_call_site() {
let out = expand_self_guard(guarded_body(vec![IrOp::Call(SELF)]), SELF);
assert_eq!(
out,
guarded_body(vec![
IrOp::Dup,
IrOp::PushI32(2),
IrOp::Lt,
IrOp::If {
then_body: vec![],
else_body: Some(vec![IrOp::Call(SELF)]),
},
])
);
}
#[test]
fn self_guard_carries_the_value_the_guard_returns() {
// `: F DUP 2 < IF DROP 0 EXIT THEN RECURSE ;` -- the base case is not
// "leave the argument", it is "replace it with 0".
let body = vec![
IrOp::Dup,
IrOp::PushI32(2),
IrOp::Lt,
IrOp::If {
then_body: vec![IrOp::Drop, IrOp::PushI32(0), IrOp::Exit],
else_body: None,
},
IrOp::Call(SELF),
];
let out = expand_self_guard(body, SELF);
let IrOp::If { then_body, .. } = &out[7] else {
panic!("expected the expanded guard at index 7, got {out:?}");
};
assert_eq!(then_body, &vec![IrOp::Drop, IrOp::PushI32(0)]);
}
#[test]
fn self_guard_leaves_a_body_without_a_guard_alone() {
// An `IF` with an `ELSE` is a branch, not an early return.
let body = vec![
IrOp::Dup,
IrOp::If {
then_body: vec![IrOp::Drop],
else_body: Some(vec![IrOp::Call(SELF)]),
},
];
assert_eq!(expand_self_guard(body.clone(), SELF), body);
// No `EXIT` in the then-branch: also not a guard.
let body = guarded_body(vec![IrOp::Call(SELF)])
.into_iter()
.map(|op| match op {
IrOp::If { .. } => IrOp::If {
then_body: vec![IrOp::Drop],
else_body: None,
},
other => other,
})
.collect::<Vec<_>>();
assert_eq!(expand_self_guard(body.clone(), SELF), body);
}
#[test]
fn self_guard_refuses_a_condition_it_cannot_run_twice() {
// A guard reached through a call or a memory write would be evaluated
// once at the call site and again inside the callee.
let body = vec![
IrOp::Call(WordId(3)),
IrOp::If {
then_body: vec![IrOp::Exit],
else_body: None,
},
IrOp::Call(SELF),
];
assert_eq!(expand_self_guard(body.clone(), SELF), body);
let body = vec![
IrOp::Dup,
IrOp::Fetch,
IrOp::If {
then_body: vec![IrOp::Exit],
else_body: None,
},
IrOp::Call(SELF),
];
assert_eq!(expand_self_guard(body.clone(), SELF), body);
}
#[test]
fn self_guard_stops_at_the_call_site_budget() {
let step = std::iter::repeat_n(IrOp::Call(SELF), MAX_GUARD_SITES + 1).collect();
let body = guarded_body(step);
assert_eq!(expand_self_guard(body.clone(), SELF), body);
}
#[test]
fn self_guard_leaves_tail_calls_alone() {
let body = guarded_body(vec![IrOp::TailCall(SELF)]);
assert_eq!(expand_self_guard(body.clone(), SELF), body);
}
#[test]
fn a_loop_stays_out_of_a_caller_that_is_only_unpromotable_through_a_call() {
// The shape the benchmark harness uses, and the one that made this
// guard a no-op for its whole life: at the moment the guard runs, the
// caller's `CR` is still `Call(cr_word)`, not `IrOp::Cr`. A test built
// from `IrOp::Cr` directly passes even with the bug.
let cross = WordId(7);
let cr = WordId(9);
let mut bodies = HashMap::new();
bodies.insert(cr, vec![IrOp::Cr]);
bodies.insert(
cross,
vec![
IrOp::PushI32(0),
IrOp::Swap,
IrOp::PushI32(0),
IrOp::DoLoop {
body: vec![IrOp::RFetch, IrOp::Call(WordId(8)), IrOp::Xor],
is_plus_loop: false,
},
],
);
let out = opt_with_inline(
vec![
IrOp::PushI32(300000),
IrOp::Call(cross),
IrOp::Drop,
IrOp::Call(cr),
],
&bodies,
);
assert!(
out.iter()
.any(|op| matches!(op, IrOp::Call(id) if *id == cross)),
"a loop-bearing callee must not be inlined into a caller that cannot \
be promoted -- it would lose its registers: {out:?}"
);
assert!(
out.iter().any(|op| matches!(op, IrOp::Cr)),
"the loop-free callee should still have been inlined: {out:?}"
);
}
fn opt_with_inline(ops: Vec<IrOp>, bodies: &HashMap<WordId, Vec<IrOp>>) -> Vec<IrOp> {
@@ -758,8 +1076,9 @@ mod tests {
strength_reduce: true,
dce: true,
inline: true,
self_guard: false,
};
optimize(ops, &config, bodies)
optimize(ops, &config, bodies, None)
}
// Peephole tests
@@ -1019,8 +1338,9 @@ mod tests {
strength_reduce: false,
dce: false,
inline: true,
self_guard: false,
};
let result = optimize(vec![IrOp::Call(WordId(5))], &config, &bodies);
let result = optimize(vec![IrOp::Call(WordId(5))], &config, &bodies, None);
assert_eq!(result, vec![IrOp::Call(WordId(5))]);
}
+194 -24
View File
@@ -270,6 +270,7 @@ pub(crate) const INTERPRETER_TOKENS: &[&str] = &[
"GILD",
"EMPTY",
"SYNONYM",
"VOCABULARY",
"CONSOLIDATE",
// Parsing words
"'",
@@ -348,6 +349,7 @@ struct MarkerState {
// REPLACES table, ABORT" texts
search_order: Vec<u32>,
next_wid: u32,
wid_names: HashMap<u32, String>,
current_wid: u32,
substitutions: HashMap<String, Vec<u8>>,
abort_messages_len: usize,
@@ -484,6 +486,8 @@ pub struct ForthVM<R: Runtime> {
search_order: Arc<Mutex<Vec<u32>>>,
/// Next wordlist ID to allocate (shared).
next_wid: Arc<Mutex<u32>>,
/// Names of wordlists created by VOCABULARY, for ORDER/WORDS display.
wid_names: HashMap<u32, String>,
/// xorshift64 PRNG state for RANDOM / RND-SEED.
rng_state: Arc<Mutex<u64>>,
/// Stacked compile state for nested definitions (quotations `[: ;]`).
@@ -712,6 +716,7 @@ impl<R: Runtime> ForthVM<R> {
substitutions: Arc::new(Mutex::new(HashMap::new())),
search_order: Arc::new(Mutex::new(vec![1])),
next_wid: Arc::new(Mutex::new(2)),
wid_names: HashMap::new(),
// SystemTime::now() PANICS on wasm32-unknown-unknown (no time
// source), which turned VM construction into an `unreachable`
// trap in the browser. Seed from the wall clock only where one
@@ -1236,21 +1241,14 @@ impl<R: Runtime> ForthVM<R> {
"FVALUE" => return self.define_fvalue(),
"CONSOLIDATE" => return self.consolidate(),
"SYNONYM" => return self.define_synonym(),
"VOCABULARY" => return self.define_vocabulary(),
"ORDER" => {
// wid 1 is FORTH-WORDLIST; other wids are anonymous.
let wid_name = |wid: u32| {
if wid == 1 {
"FORTH".to_string()
} else {
format!("wid#{wid}")
}
};
let so = self.search_order.lock().unwrap();
let names: Vec<String> = so.iter().map(|&w| wid_name(w)).collect();
let order = self.search_order.lock().unwrap().clone();
let names: Vec<String> = order.iter().map(|&w| self.wid_name(w)).collect();
let output = format!(
"Search order: {} Compilation: {}\n",
names.join(" "),
wid_name(self.dictionary.current_wid())
self.wid_name(self.dictionary.current_wid())
);
self.output.lock().unwrap().push_str(&output);
return Ok(());
@@ -2463,8 +2461,13 @@ impl<R: Runtime> ForthVM<R> {
}
/// Run all enabled optimization passes on an IR sequence.
fn optimize_ir(&self, ir: Vec<IrOp>, bodies: &HashMap<WordId, Vec<IrOp>>) -> Vec<IrOp> {
optimize(ir, &self.config.opt, bodies)
fn optimize_ir(
&self,
ir: Vec<IrOp>,
bodies: &HashMap<WordId, Vec<IrOp>>,
self_id: Option<WordId>,
) -> Vec<IrOp> {
optimize(ir, &self.config.opt, bodies, self_id)
}
/// Parse a `{: args | locals -- comment :}` block and compile local
@@ -2576,7 +2579,7 @@ impl<R: Runtime> ForthVM<R> {
let ir = std::mem::take(&mut self.compiling_ir);
let bodies = self.ir_bodies.clone();
let ir = self.optimize_ir(ir, &bodies);
let ir = self.optimize_ir(ir, &bodies, Some(word_id));
self.ir_bodies.insert(word_id, ir.clone());
// Compile to WASM
@@ -2992,7 +2995,7 @@ impl<R: Runtime> ForthVM<R> {
ir_body: Vec<IrOp>,
) -> anyhow::Result<WordId> {
let bodies = self.ir_bodies.clone();
let ir_body = self.optimize_ir(ir_body, &bodies);
let ir_body = self.optimize_ir(ir_body, &bodies, None);
let word_id = self
.dictionary
.create(name, immediate)
@@ -3737,6 +3740,47 @@ impl<R: Runtime> ForthVM<R> {
Ok(())
}
/// VOCABULARY <name> -- create a named wordlist (fig-Forth heritage;
/// gforth/SwiftForth extension, not Forth 2012). Executing the created
/// word replaces the top of the search order with its wordlist, the
/// same semantics the standard gives the word FORTH. The name is
/// remembered so ORDER and WORDS ALL display it instead of wid#N.
fn define_vocabulary(&mut self) -> anyhow::Result<()> {
let name = self
.next_token()
.ok_or_else(|| anyhow::anyhow!("VOCABULARY: expected name"))?;
let set_context_id = self
.dictionary
.find("_SET_CONTEXT_")
.map(|(_, id, _)| id)
.ok_or_else(|| anyhow::anyhow!("_SET_CONTEXT_ not found"))?;
let wid = {
let mut nw = self.next_wid.lock().unwrap();
let wid = *nw;
*nw += 1;
wid
};
self.wid_names.insert(wid, name.to_uppercase());
let word_id = self
.dictionary
.create(&name, false)
.map_err(|e| anyhow::anyhow!("{e}"))?;
let ir_body = vec![IrOp::PushI32(wid as i32), IrOp::Call(set_context_id)];
self.ir_bodies.insert(word_id, ir_body.clone());
self.word_sources
.insert(word_id, format!("VOCABULARY {name}"));
let config = self.codegen_config(word_id.0);
let compiled = compile_word(&name, &ir_body, &config)
.map_err(|e| anyhow::anyhow!("codegen error for VOCABULARY: {e}"))?;
self.instantiate_and_install(&compiled, word_id)?;
self.dictionary.reveal();
self.next_table_index = self.next_table_index.max(word_id.0 + 1);
Ok(())
}
/// IMMEDIATE -- toggle the immediate flag on the most recently defined word.
/// Called via `pending_define` when IMMEDIATE is executed from compiled code.
fn set_immediate(&mut self) -> anyhow::Result<()> {
@@ -3806,6 +3850,7 @@ impl<R: Runtime> ForthVM<R> {
fvalue_words: self.fvalue_words.clone(),
search_order: self.search_order.lock().unwrap().clone(),
next_wid: *self.next_wid.lock().unwrap(),
wid_names: self.wid_names.clone(),
current_wid: self.dictionary.current_wid(),
substitutions: self.substitutions.lock().unwrap().clone(),
abort_messages_len: self.abort_messages.lock().unwrap().len(),
@@ -3829,6 +3874,7 @@ impl<R: Runtime> ForthVM<R> {
self.fvalue_words = state.fvalue_words;
*self.search_order.lock().unwrap() = state.search_order;
*self.next_wid.lock().unwrap() = state.next_wid;
self.wid_names = state.wid_names;
*self.substitutions.lock().unwrap() = state.substitutions;
self.abort_messages
.lock()
@@ -6517,6 +6563,17 @@ impl<R: Runtime> ForthVM<R> {
/// `WORDS ALL` -- grouped full view: one section per wordlist (search
/// order first, then any other populated wids), then internal words.
/// Display name for a wordlist: FORTH, a VOCABULARY name, or wid#N.
fn wid_name(&self, wid: u32) -> String {
if wid == 1 {
return "FORTH".to_string();
}
self.wid_names
.get(&wid)
.cloned()
.unwrap_or_else(|| format!("wid#{wid}"))
}
fn do_words_all(&mut self) {
let entries = self.dictionary.visible_entries();
let mut wids: Vec<u32> = self.search_order.lock().unwrap().clone();
@@ -6525,15 +6582,9 @@ impl<R: Runtime> ForthVM<R> {
wids.push(*wid);
}
}
let wid_name = |wid: u32| {
if wid == 1 {
"FORTH".to_string()
} else {
format!("wid#{wid}")
}
};
let wid_names: Vec<String> = wids.iter().map(|&w| self.wid_name(w)).collect();
let mut out = self.output.lock().unwrap();
for wid in wids {
for (wid, wid_name) in wids.iter().copied().zip(wid_names) {
let names: Vec<&str> = entries
.iter()
.filter(|(_, w, internal)| *w == wid && !internal)
@@ -6542,7 +6593,7 @@ impl<R: Runtime> ForthVM<R> {
if names.is_empty() {
continue;
}
out.push_str(&format!("-- {} ({} words)\n", wid_name(wid), names.len()));
out.push_str(&format!("-- {} ({} words)\n", wid_name, names.len()));
push_wrapped(&mut out, &names);
}
let internals: Vec<&str> = entries
@@ -6848,6 +6899,56 @@ impl<R: Runtime> ForthVM<R> {
self.register_host_primitive("ALSO", false, func)?;
}
// >ORDER ( wid -- ) — push wid on top of the search order.
// Not in Forth 2012; a widely used gforth extension.
{
let so = Arc::clone(&self.search_order);
let func: HostFn = Box::new(move |ctx: &mut dyn HostAccess| {
let sp = host_need(ctx, 1)?;
let wid = ctx.mem_read_i32(sp) as u32;
so.lock().unwrap().insert(0, wid);
ctx.set_dsp(sp + CELL_SIZE);
Ok(())
});
self.register_host_primitive(">ORDER", false, func)?;
}
// -ORDER ( wid -- ) — remove wid from the search order wherever it
// sits (no-op if absent). VFX/MPE extension, the inverse of >ORDER.
{
let so = Arc::clone(&self.search_order);
let func: HostFn = Box::new(move |ctx: &mut dyn HostAccess| {
let sp = host_need(ctx, 1)?;
let wid = ctx.mem_read_i32(sp) as u32;
so.lock().unwrap().retain(|&w| w != wid);
ctx.set_dsp(sp + CELL_SIZE);
Ok(())
});
self.register_host_primitive("-ORDER", false, func)?;
}
// _SET_CONTEXT_ ( wid -- ) — replace the top of the search order.
// Internal carrier for words created by VOCABULARY (same pattern as
// _MARKER_RESTORE_): a vocabulary word compiles to
// `PushI32(wid) Call(_SET_CONTEXT_)`.
{
let so = Arc::clone(&self.search_order);
let func: HostFn = Box::new(move |ctx: &mut dyn HostAccess| {
let sp = host_need(ctx, 1)?;
let wid = ctx.mem_read_i32(sp) as u32;
let mut order = so.lock().unwrap();
if order.is_empty() {
order.push(wid);
} else {
order[0] = wid;
}
drop(order);
ctx.set_dsp(sp + CELL_SIZE);
Ok(())
});
self.register_host_primitive("_SET_CONTEXT_", false, func)?;
}
// PREVIOUS ( -- ) remove top of search order
{
let so = Arc::clone(&self.search_order);
@@ -8290,6 +8391,36 @@ mod tests {
assert_eq!(eval_output(": D 1 2 3 2 0 DO ROT LOOP . . . ; D"), "2 1 3 ");
}
#[test]
fn test_self_guard_expansion_keeps_the_answers() {
// The base-case guard is tested at the call site, so a leaf never
// costs a call. All four verified against gforth.
assert_eq!(
eval_output(
": FIB DUP 2 < IF EXIT THEN DUP 1- RECURSE SWAP 2 - RECURSE + ; \
25 FIB . 0 FIB . 1 FIB . 2 FIB . 10 FIB ."
),
"75025 0 1 1 55 "
);
// A guard that replaces its argument rather than leaving it.
assert_eq!(
eval_output(": G DUP 0= IF DROP 0 EXIT THEN DUP 1- RECURSE + ; 5 G . 0 G . 100 G ."),
"15 0 5050 "
);
assert_eq!(
eval_output(": H DUP 3 < IF DROP 7 EXIT THEN 1- RECURSE 2 * ; 5 H . 2 H . 8 H ."),
"56 7 448 "
);
// Two guards and three call sites, one of them behind the second guard.
assert_eq!(
eval_output(
": ACK OVER 0= IF SWAP DROP 1+ EXIT THEN DUP 0= IF DROP 1- 1 RECURSE EXIT THEN \
OVER SWAP 1- RECURSE SWAP 1- SWAP RECURSE ; 2 3 ACK . 1 2 ACK ."
),
"9 4 "
);
}
#[test]
fn test_promoted_begin_loops() {
// BEGIN loops promote too, so these run entirely in locals.
@@ -10692,6 +10823,45 @@ mod tests {
assert!(!output.contains('['));
}
#[test]
fn test_to_order_pushes_on_top() {
let output = eval_output("WORDLIST >ORDER ORDER");
assert!(output.contains("Search order: wid#2 FORTH"), "{output}");
}
#[test]
fn test_minus_order_removes_wid() {
let output = eval_output("WORDLIST DUP >ORDER -ORDER ORDER");
assert!(output.contains("Search order: FORTH "), "{output}");
assert!(!output.contains("wid#"), "{output}");
}
#[test]
fn test_vocabulary_replaces_top_and_names_order() {
// Vocabulary execution has FORTH-word semantics: replace the top
// of the search order. ALSO first, so FORTH stays underneath.
let output = eval_output("VOCABULARY EDITOR ALSO EDITOR ORDER");
assert!(output.contains("Search order: EDITOR FORTH"), "{output}");
}
#[test]
fn test_vocabulary_definitions_land_in_it_and_words_groups_by_name() {
let output = eval_output("VOCABULARY EDITOR ALSO EDITOR DEFINITIONS : E1 1 ; WORDS ALL");
assert!(output.contains("-- EDITOR (1 words)"), "{output}");
// and the word is findable through the search order
let output = eval_output("VOCABULARY EDITOR ALSO EDITOR DEFINITIONS : E1 42 ; E1 .");
assert!(output.contains("42"), "{output}");
}
#[test]
fn test_marker_rolls_back_vocabulary_name() {
// After rollback the vocabulary is gone; a freshly allocated wid
// with the same number must not inherit its stale name.
let output = eval_output("MARKER M VOCABULARY V0 M WORDLIST >ORDER ORDER");
assert!(!output.contains("V0"), "{output}");
assert!(output.contains("wid#2"), "{output}");
}
// ===================================================================
// Double DOES>: Forth 2012 WEIRD: W1 test
// ===================================================================
+15
View File
@@ -1004,6 +1004,21 @@ pub const WORD_DOCS: &[(&str, &str, &str)] = &[
"New definitions go to the top wordlist.",
),
("ALSO", "( -- )", "Duplicate the top of the search order."),
(
">ORDER",
"( wid -- )",
"Push wid on top of the search order (gforth extension).",
),
(
"-ORDER",
"( wid -- )",
"Remove wid from the search order (VFX extension).",
),
(
"VOCABULARY",
"( \"name\" -- )",
"Create a named wordlist; executing name replaces the top of the search order.",
),
("ONLY", "( -- )", "Reset the search order to the minimum."),
("PREVIOUS", "( -- )", "Drop the top of the search order."),
(
+92 -39
View File
@@ -731,7 +731,6 @@ struct PerfBenchmark {
run_code: &'static str,
verify: &'static str,
expected: i32,
samples: u32, // Number of runs for WAFER median
/// Maximum acceptable WAFER/gforth ratio (< 1.0 = WAFER faster).
/// Test fails if ratio exceeds this. Set ~40-50% above measured baseline.
max_ratio: f64,
@@ -740,54 +739,67 @@ struct PerfBenchmark {
fn perf_benchmarks() -> Vec<PerfBenchmark> {
vec![
PerfBenchmark {
name: "Fibonacci(25)",
name: "Fibonacci(33)",
define: ": FIB DUP 2 < IF EXIT THEN DUP 1- RECURSE SWAP 2 - RECURSE + ;",
run_code: "25 FIB DROP",
verify: "25 FIB",
expected: 75025,
samples: 5,
max_ratio: 0.17,
run_code: "33 FIB DROP",
verify: "33 FIB",
expected: 3524578,
max_ratio: 0.10,
},
PerfBenchmark {
name: "Factorial(12)x100K",
name: "Factorial(12)x2M",
define: ": FACT 1 SWAP 1+ 1 ?DO I * LOOP ; \
: FACT-BENCH 100000 0 DO 12 FACT DROP LOOP ;",
: FACT-BENCH 2000000 0 DO 12 FACT DROP LOOP ;",
run_code: "FACT-BENCH",
verify: "12 FACT",
expected: 479001600,
samples: 5,
max_ratio: 0.12,
},
PerfBenchmark {
name: "GCD-bench(20K)",
name: "GCD-bench(400K)",
define: ": GCD BEGIN DUP WHILE TUCK MOD REPEAT DROP ; \
: GCD-BENCH 0 DO 10000 I 1+ GCD DROP LOOP ;",
run_code: "20000 GCD-BENCH",
run_code: "400000 GCD-BENCH",
verify: "48 36 GCD",
expected: 12,
samples: 5,
max_ratio: 0.45,
},
PerfBenchmark {
name: "NestedLoops(50)x1K",
name: "NestedLoops(50)x20K",
define: ": NESTED 0 SWAP 0 DO I 0 ?DO I J + DROP LOOP LOOP ; \
: NESTED-BENCH 1000 0 DO 50 NESTED DROP LOOP ;",
: NESTED-BENCH 20000 0 DO 50 NESTED DROP LOOP ;",
run_code: "NESTED-BENCH",
verify: "5 NESTED",
expected: 0,
samples: 5,
max_ratio: 0.11,
},
PerfBenchmark {
name: "Collatz(2K)",
// The only benchmark with a cross-word call left in its hot loop:
// WORK is over the inliner's eight-operation budget, so it stays a
// real call. That is what CONSOLIDATE exists to turn into a direct
// one, and without this the CONSOL column measures nothing -- every
// other benchmark has its callee inlined away or self-recursive.
name: "CrossCalls(3M)",
define: ": WORK DUP 3 * OVER XOR SWAP 2 / XOR DUP 7 AND XOR DUP 1 AND XOR ; \
: CROSS-BENCH 0 SWAP 0 DO I WORK XOR LOOP ;",
run_code: "3000000 CROSS-BENCH DROP",
verify: "1000 CROSS-BENCH",
expected: 3176,
// Guards CONSOLIDATE as much as the engine: the ratio uses the
// better of the two columns, so a consolidation regression here
// pushes it from 0.04 to 0.12 and trips the limit.
max_ratio: 0.08,
},
PerfBenchmark {
name: "Collatz(2K)x50",
define: ": COLLATZ 0 SWAP BEGIN DUP 1 > WHILE \
DUP 1 AND IF 3 * 1+ ELSE 2 / THEN \
SWAP 1+ SWAP REPEAT DROP ; \
: COLLATZ-BENCH 0 DO I 1+ COLLATZ DROP LOOP ;",
run_code: "2000 COLLATZ-BENCH",
: COLLATZ-BENCH 0 DO I 1+ COLLATZ DROP LOOP ; \
: COLLATZ-REPEAT 50 0 DO 2000 COLLATZ-BENCH LOOP ;",
run_code: "COLLATZ-REPEAT",
verify: "27 COLLATZ",
expected: 111,
samples: 3,
max_ratio: 0.08,
},
]
@@ -833,15 +845,16 @@ fn measure_wafer_release(wafer: &str, bench: &PerfBenchmark) -> Option<u64> {
let code = format!(
"{define} {run} \
: TIMED-BENCH UTIME {run} UTIME 2SWAP D- DROP . CR ; \
TIMED-BENCH TIMED-BENCH TIMED-BENCH",
{reps}",
define = bench.define,
run = bench.run_code,
reps = repeat_timed(" "),
);
let output = run_via_stdin(wafer, &code)?;
if !output.status.success() {
return None;
}
median_printed_time(&output.stdout)
best_of_printed_times(&output.stdout)
}
/// Measure WAFER execution time after CONSOLIDATE (direct calls between all words).
@@ -849,31 +862,70 @@ fn measure_wafer_consolidated(wafer: &str, bench: &PerfBenchmark) -> Option<u64>
let code = format!(
"{define} CONSOLIDATE {run} \
: TIMED-BENCH UTIME {run} UTIME 2SWAP D- DROP . CR ; \
TIMED-BENCH TIMED-BENCH TIMED-BENCH",
{reps}",
define = bench.define,
run = bench.run_code,
reps = repeat_timed(" "),
);
let output = run_via_stdin(wafer, &code)?;
if !output.status.success() {
return None;
}
median_printed_time(&output.stdout)
best_of_printed_times(&output.stdout)
}
/// Parse the microsecond values printed by TIMED-BENCH (one per line) and
/// return the median.
fn median_printed_time(stdout: &[u8]) -> Option<u64> {
/// How many separate process invocations each measurement takes the best of.
///
/// `REPS`/`BEST_OF` deal with noise *inside* one process. They do not touch
/// the rest: whether a process lands on a core whose SMT sibling is busy, and
/// where its code ends up in memory, are fixed for its lifetime, and they make
/// some benchmarks frankly bimodal -- Fibonacci after CONSOLIDATE measured
/// 413-419 us in three runs of the report and 712-770 in the other two, with
/// nothing in between. Only a fresh process resamples that.
const PROCESS_RUNS: usize = 3;
/// How many timed repetitions each engine runs per benchmark.
const REPS: usize = 7;
/// How many of the fastest repetitions the reported time averages over.
const BEST_OF: usize = 3;
/// Run `measure` in `PROCESS_RUNS` fresh processes and keep the fastest.
///
/// The minimum, not a mean: process-level noise is one-sided too, so the
/// fastest process is the one that ran closest to undisturbed.
fn best_of_processes(mut measure: impl FnMut() -> Option<u64>) -> Option<u64> {
(0..PROCESS_RUNS).filter_map(|_| measure()).min()
}
/// `TIMED-BENCH` repeated `REPS` times, separated by `sep`.
///
/// sf64 needs one statement per line (it truncates input at ~256 characters);
/// the others do not care.
fn repeat_timed(sep: &str) -> String {
["TIMED-BENCH"; REPS].join(sep)
}
/// Parse the microsecond values printed by `TIMED-BENCH` and reduce them to
/// one number: the mean of the fastest `BEST_OF`.
///
/// Not the median, and not the mean of all of them. Benchmark noise on a
/// shared machine is one-sided -- a scheduling hiccup, an SMT sibling or a
/// migration can only ever make a run slower, never faster -- so the fastest
/// repetitions are the ones closest to the cost we are trying to measure.
/// Averaging a few of them rather than taking the single minimum keeps one
/// lucky run from setting the result on its own.
fn best_of_printed_times(stdout: &[u8]) -> Option<u64> {
let stdout = String::from_utf8_lossy(stdout);
let mut times: Vec<u64> = stdout
.trim()
.lines()
.filter_map(|l| l.trim().parse::<u64>().ok())
.collect();
times.sort();
if times.is_empty() {
return None;
}
Some(times[times.len() / 2])
times.sort_unstable();
let n = times.len().min(BEST_OF);
Some(times[..n].iter().sum::<u64>() / n as u64)
}
/// Measure gforth execution time using Forth-level `utime` (excludes startup).
@@ -881,19 +933,19 @@ fn median_printed_time(stdout: &[u8]) -> Option<u64> {
/// Returns microseconds, or None if gforth is unavailable.
fn measure_gforth(gforth: &str, bench: &PerfBenchmark) -> Option<u64> {
// The timing wrapper must be inside a word (DO/LOOP is compile-only in gforth).
// We take the median of 3 runs.
let code = format!(
"{define} {run} \
: TIMED-BENCH utime {run} utime 2swap d- drop . CR ; \
TIMED-BENCH TIMED-BENCH TIMED-BENCH bye",
{reps} bye",
define = bench.define,
run = bench.run_code,
reps = repeat_timed(" "),
);
let output = Command::new(gforth).arg("-e").arg(&code).output().ok()?;
if !output.status.success() {
return None;
}
median_printed_time(&output.stdout)
best_of_printed_times(&output.stdout)
}
/// Measure `SwiftForth` (`sf64`) execution time using Forth-level `ucounter`
@@ -906,15 +958,16 @@ fn measure_sf64(sf64: &str, bench: &PerfBenchmark) -> Option<u64> {
let code = format!(
"{define}\n{run}\n\
: TIMED-BENCH ucounter {run} ucounter 2swap d- drop . cr ;\n\
TIMED-BENCH\nTIMED-BENCH\nTIMED-BENCH\nbye\n",
{reps}\nbye\n",
define = bench.define,
run = bench.run_code,
reps = repeat_timed("\n"),
);
let output = run_via_stdin(sf64, &code)?;
if !output.status.success() {
return None;
}
median_printed_time(&output.stdout)
best_of_printed_times(&output.stdout)
}
#[test]
@@ -987,14 +1040,14 @@ fn performance_report() {
for bench in &benchmarks {
let wafer = wafer_release
.and_then(|w| measure_wafer_release(w, bench))
.and_then(|w| best_of_processes(|| measure_wafer_release(w, bench)))
.unwrap_or(0);
let consol = wafer_release
.and_then(|w| measure_wafer_consolidated(w, bench))
.and_then(|w| best_of_processes(|| measure_wafer_consolidated(w, bench)))
.unwrap_or(0);
let gf = gforth.and_then(|g| measure_gforth(g, bench));
let gf_fast = gforth_fast.and_then(|g| measure_gforth(g, bench));
let sf = sf64.and_then(|s| measure_sf64(s, bench));
let gf = gforth.and_then(|g| best_of_processes(|| measure_gforth(g, bench)));
let gf_fast = gforth_fast.and_then(|g| best_of_processes(|| measure_gforth(g, bench)));
let sf = sf64.and_then(|s| best_of_processes(|| measure_sf64(s, bench)));
let gf_str = gf.map_or_else(|| "-".to_string(), |v| format!("{v}"));
let gf_fast_str = gf_fast.map_or_else(|| "-".to_string(), |v| format!("{v}"));
+1 -1
View File
@@ -12,7 +12,7 @@ workspace = true
crate-type = ["cdylib", "rlib"]
[dependencies]
wafer-core = { path = "../core", version = "0.2.7", default-features = false, features = ["crypto"] }
wafer-core = { path = "../core", version = "0.3.0", default-features = false, features = ["crypto"] }
wasm-bindgen = "0.2"
js-sys = "0.3"
send_wrapper = { workspace = true }
+76 -15
View File
@@ -30,6 +30,7 @@ This document describes every optimization that makes sense for WAFER, why it ma
| 14 | Self-Recursive Direct Call | Codegen | Done | High |
| 15 | Float / Double-Cell | Codegen | Not started | Future |
| 16 | Typed Calling Convention | Codegen | Done | Highest |
| 17 | Self-Guard Expansion | IR pass | Done | Medium |
## 1. Stack-to-Local Promotion
@@ -307,6 +308,28 @@ After interactive development, `CONSOLIDATE` recompiles all defined words into a
| JIT (current) | Interactive development | Per-word modules, `call_indirect`, fast redefine |
| Consolidated | After `CONSOLIDATE` | Single module, direct `call`, no redefine |
### Why the CONSOL column can lose to the JIT column
`NestedLoops` runs 1.1x slower after `CONSOLIDATE` on the M1 and 1.7x slower on a Skylake Xeon,
with **byte-identical WASM** for the hot word in both modes (verified via `WAFER_DUMP_WASM` +
`wasm-tools print`) and instruction-identical machine code modulo register names (verified via
`Engine::precompile_module` + objdump). The whole delta is code placement:
- A tight loop pays for straddling an instruction-fetch window: ~9% for a 16-byte window on the
M1, up to ~65% on Skylake when the fused `cmp+jcc` crosses a 32-byte boundary and the loop
falls out of the uop cache every iteration (the JCC erratum, post-microcode).
- Cranelift never aligns loop headers (`align_basic_block` is an identity default, no ISA
overrides it), so where a loop lands is whatever the code before it leaves behind.
- The per-word JIT module keeps a dead dsp load in its prologue (the store-back is DCE'd, the
load survives), which happens to shift its loops onto luckier offsets than the consolidated
module's cleaner function bodies. A padding experiment that moves the same loop across offsets
reproduces the full penalty range on both hosts, including placements where the consolidated
code **beats** the JIT code.
So the column difference on loop-only benchmarks is an alignment lottery, not an emitter defect;
divider-bound benchmarks (`GCD`) mask it entirely. Fixing it for real means loop-header alignment
upstream in Cranelift.
## 9. Compound IR Operations
**Status: Done.** `TwoDup` and `TwoDrop` IrOp variants with optimized codegen. Peephole converts `Over, Over -> TwoDup` and `Drop, Drop -> TwoDrop`.
@@ -468,7 +491,7 @@ The float stack lives in its own memory region (0x2540--0x2D40). Float operation
## 16. Typed Calling Convention
**Status: Done.** A word whose stack effect is statically known compiles to two entry points: a fast one with signature `(i32 x p) -> (i32 x q)`, carrying its stack items as WASM values, and the usual `( -- )` wrapper that moves those items on and off the memory data stack. The wrapper keeps the function-table slot, so `EXECUTE`, the outer interpreter, host words and `CATCH` see exactly the ABI they saw before; only direct calls inside a module take the fast entry. `WAFER_TYPED_CALLS=0` falls back.
**Status: Done.** A word that calls itself and whose stack effect is statically known compiles to two entry points: a fast one with signature `(i32 x p) -> (i32 x q)`, carrying its stack items as WASM values, and the usual `( -- )` wrapper that moves those items on and off the memory data stack. The wrapper keeps the function-table slot, so `EXECUTE`, the outer interpreter, host words and `CATCH` see exactly the ABI they saw before; only direct calls inside a module take the fast entry. `WAFER_TYPED_CALLS=0` falls back. The self-recursion condition matters: the table slot holds the wrapper, so in the JIT path nothing but `RECURSE` can reach the fast entry, and emitting it for any other word just puts a wrapper hop in front of every call through the table -- measured at +47% before that was fixed in 0.2.9.
### The Problem
@@ -484,32 +507,70 @@ Fibonacci(25) went from 1035 to 366 microseconds, 4.3x slower than `sf64` to 1.2
Untyped by design: anything using `SP@`, `DEPTH`, `EXECUTE`, `>R`/`R>`, floats or locals; anything calling a word that is itself untyped, which in the JIT path means every call except `RECURSE`; mutually recursive words; and words whose effect is not static -- branches that disagree on depth, `EXIT` at the wrong depth, a non-neutral loop body, or a recursion that grows the stack per level.
## 17. Self-Guard Expansion
**Status: Done.** A recursive word's base-case guard is duplicated into its own call sites, so the leaves of the recursion cost a test instead of a call. Implemented in `optimizer.rs::expand_self_guard`, gated on `OptConfig::self_guard`, and applied after inlining so the later passes still run over the result.
### The Shape
A recursive Forth word almost always opens with a guard that returns early:
```forth
: FIB DUP 2 < IF EXIT THEN DUP 1- RECURSE SWAP 2 - RECURSE + ;
```
Every leaf of the recursion is then a call whose entire body is `DUP 2 <`. The pass rewrites each `Call(self)` as
```forth
DUP 2 < IF ( leave it ) ELSE RECURSE THEN
```
which computes the same thing -- the callee would have run the guard, taken the branch and returned. When the guard returns a value rather than its argument (`IF DROP 0 EXIT THEN`), that value moves into the then-branch with it.
### Why It Is Bounded
The guard runs twice along the recursive path: once at the call site, once inside the callee. So it must be small and free of effects -- `MAX_GUARD_OPS` is six, and the operations are restricted to stack shuffles, arithmetic and comparisons; a call, a memory access or a branch disqualifies it. `MAX_GUARD_SITES` caps the expansion at four call sites, since each one replicates the guard. A `TailCall` is never expanded, which keeps this pass and tail-call detection from having to agree about what tail position means.
### Impact
Fibonacci(25): 356 to 237 microseconds on the arm64 development machine. In fib's tree half of all nodes are leaves, which is where the factor comes from. It does not take Fibonacci past `sf64`, though the arm64 table below says otherwise: with both engines native on x86-64, Fibonacci reads 1.16x and stays the one benchmark `sf64` wins.
## Current Performance vs Gforth
All optimizations enabled, release mode, measured with UTIME:
Development machine (M1 Ultra, arm64), median of three reports, every
benchmark sized to about 10 ms:
```
Benchmark WAFER CONSOL gforth sf64 WAFER/gf WAFER/sf
Fibonacci(25) 356 361 3389 287 0.11x 1.24x
Factorial(12)x100K 479 495 6249 1650 0.08x 0.29x
GCD-bench(20K) 540 559 1801 801 0.30x 0.67x
NestedLoops(50)x1K 509 501 7023 1887 0.07x 0.27x
Collatz(2K) 185 213 3873 610 0.05x 0.30x
Fibonacci(33) 11307 11407 157001 13053 0.07x 0.87x
Factorial(12)x2M 9639 9599 123950 32091 0.08x 0.30x
GCD-bench(400K) 11662 11580 38580 17001 0.30x 0.68x
NestedLoops(50)x20K 8920 9852 140518 36828 0.06x 0.24x
CrossCalls(3M) 10883 3769 87691 8240 0.04x 0.46x
Collatz(2K)x50 8838 8715 189903 28657 0.05x 0.30x
```
Times in microseconds. WAFER/gf < 1.0 means WAFER is faster. `sf64` is SwiftForth,
which compiles to native code; two caveats on that column. The install here is an
x86-64 binary under Rosetta 2 while WAFER and gforth are native arm64, so it is a
native-vs-emulated comparison and a native SwiftForth would be faster than these
numbers; and sf64 uses 64-bit cells to WAFER's 32-bit. WAFER is ahead on the four
loop-heavy benchmarks and behind on Fibonacci, which is one call per node with no
loop to promote.
Times in microseconds; ratios take the better of WAFER and CONSOL. The `sf64`
column flatters WAFER: the only SwiftForth build for macOS is x86-64 under
Rosetta 2 while WAFER and gforth are native arm64. Measured with all three
native on x86-64, Fibonacci reads 1.23x rather than 0.87x -- the emulation
penalty lands hardest on the call-heavy benchmark -- while the other five keep
their ratios. One caveat holds on both: sf64 uses 64-bit cells to WAFER's
32-bit. (The native table is being re-taken at these workload sizes.)
`CrossCalls` is the only benchmark with a cross-word call left in its hot loop,
so it is the only one that measures section 8 at all -- the other five have
their callee inlined away or are self-recursive. Note that `CONSOLIDATE` makes
NestedLoops and Collatz _slower_; see the open item below.
## Remaining Opportunities
| Optimization | Status | Potential Impact |
| -------------------------------- | ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Bounded self-inlining | Not started | Measured 1.33x on Fibonacci, the last benchmark behind sf64. Blocked on `EXIT`: the inliner refuses any body containing one, and a recursive Forth word is `... IF EXIT THEN ... RECURSE`. Needs either a scoped exit (compile an inlined `EXIT` as a branch to the end of a block) or guard-only expansion |
| --------------------------------- | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Explain CONSOLIDATE on pure loops | Open defect | Isolated on an idle box: NestedLoops 540 -> 905 us (1.68x) and Collatz 310 -> 360 (1.16x) on x86-64, against 1.07x and 1.05x for the same probes on arm64 -- so the magnitude is strongly architecture-dependent, which points at code size or branch density rather than a gross codegen error. Not the promotion logic (same code path), not inlining (no call left), not the harness (CONSOLIDATE is outside the timed window). Next step is to diff the emitted wat for NESTED-BENCH between the two paths |
| Scoped exit for the inliner | Not started | The inliner still refuses any body containing an `EXIT`, because an inlined one would return from the caller. Compiling it as a branch to the end of a block would unlock inlining for every word with an early return, not just the guard shape section 17 handles |
| BeginDoubleWhileRepeat promotion | Not started | Rare pattern, low priority. Its promoted emitter exists but has no loop fixup and is unverified |
| LEAVE as IR primitive | Not started | Would enable fast-path for loops with LEAVE |
| Float stack-to-local | Not started | Eliminate float stack memory traffic |