Compare commits
3 Commits
3bb613ece0
...
v0.2.8
| Author | SHA1 | Date | |
|---|---|---|---|
| 4f96f8860a | |||
| e963e636d3 | |||
| 645c2dadd7 |
@@ -5,6 +5,29 @@ All notable changes to WAFER are documented in this file.
|
||||
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
||||
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||||
|
||||
## [0.2.8] - 2026-08-10
|
||||
|
||||
### Added
|
||||
|
||||
- **A recursive word tests its base case at the call site.** A recursive Forth
|
||||
word almost always opens with a guard that returns early --
|
||||
`: FIB DUP 2 < IF EXIT THEN ... RECURSE ... ;` -- so every leaf of the
|
||||
recursion costs a call whose entire body is that test. `Call(self)` now
|
||||
compiles as `<guard> IF <what the guard returns> ELSE Call(self) THEN`,
|
||||
which computes the same thing: the callee would have run the guard, taken
|
||||
the branch and returned. In fib's tree the leaves are half of all nodes.
|
||||
|
||||
Fibonacci(25) 356 -> 237 µs on the arm64 development machine, where that
|
||||
reads 1.24x -> 0.83x of SwiftForth `sf64`. Measured again with **both
|
||||
engines native on x86-64** -- the macOS `sf64` build runs under Rosetta 2,
|
||||
which flatters WAFER -- Fibonacci is 1.16x, so it remains the one benchmark
|
||||
of the five that `sf64` wins. See the two tables in the README.
|
||||
|
||||
The guard runs twice along the recursive path, so it has to be small (at
|
||||
most six operations) and free of effects -- no calls, no memory, no
|
||||
branches. Words with more than four self-call sites are left alone to bound
|
||||
the code growth, and a `TailCall` is never expanded.
|
||||
|
||||
## [0.2.7] - 2026-08-09
|
||||
|
||||
### Added
|
||||
@@ -340,6 +363,7 @@ compliance suite, `CONSOLIDATE` whole-program recompilation, `wafer build`
|
||||
AOT export (WASM / native / JS loader), browser REPL, SHA-1/256/512 words,
|
||||
and cross-engine benchmark lanes against gforth and SwiftForth.
|
||||
|
||||
[0.2.8]: https://github.com/ok2/wafer/compare/v0.2.7...v0.2.8
|
||||
[0.2.7]: https://github.com/ok2/wafer/compare/v0.2.6...v0.2.7
|
||||
[0.2.1]: https://github.com/ok2/wafer/compare/v0.2.0...v0.2.1
|
||||
[0.2.0]: https://github.com/ok2/wafer/compare/v0.1.0...v0.2.0
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
## What is WAFER?
|
||||
|
||||
WAFER (WebAssembly Forth Engine in Rust) is an optimizing Forth 2012 compiler targeting WebAssembly. Currently a working Forth system with 200+ words, JIT compilation, 12 word sets at 100% compliance, and a full optimization pipeline (peephole, constant folding, inlining, strength reduction, DCE, tail calls, per-region stack-to-local promotion with DO/BEGIN loop and IF support, self-recursive direct calls, a typed calling convention for words with a known stack effect, consolidation). Beats gforth on all benchmarks in release mode, and SwiftForth `sf64` on four of five. Includes a browser-based REPL via wasm-pack.
|
||||
WAFER (WebAssembly Forth Engine in Rust) is an optimizing Forth 2012 compiler targeting WebAssembly. Currently a working Forth system with 200+ words, JIT compilation, 12 word sets at 100% compliance, and a full optimization pipeline (peephole, constant folding, inlining, strength reduction, DCE, tail calls, per-region stack-to-local promotion with DO/BEGIN loop and IF support, self-recursive direct calls, a typed calling convention for words with a known stack effect, self-guard expansion for recursive words, consolidation). Beats gforth on every benchmark, and SwiftForth `sf64` on four of five (measured native-vs-native on x86-64; the macOS sf64 build is x86-64 under Rosetta and flatters WAFER). Includes a browser-based REPL via wasm-pack.
|
||||
|
||||
## Architecture
|
||||
|
||||
@@ -79,7 +79,7 @@ Handle in `interpret_token_immediate()` or `compile_token()` as a special case.
|
||||
|
||||
## Testing
|
||||
|
||||
- Run `cargo test --workspace` before committing (currently 601 unit + 1 benchmark + 12 compliance + 9 comparison + 5 crypto)
|
||||
- Run `cargo test --workspace` before committing (currently 608 unit + 1 benchmark + 12 compliance + 9 comparison + 5 crypto)
|
||||
- Forth 2012 compliance: `cargo test -p wafer-core --test compliance`
|
||||
- Cross-engine comparison (vs gforth): `cargo test -p wafer-core --test comparison`
|
||||
- Performance benchmarks (release mode): `cargo test -p wafer-core --test comparison -- --nocapture --ignored`
|
||||
|
||||
Generated
+3
-3
@@ -1589,7 +1589,7 @@ checksum = "0b928f33d975fc6ad9f86c8f283853ad26bdd5b10b7f1542aa2fa15e2289105a"
|
||||
|
||||
[[package]]
|
||||
name = "wafer"
|
||||
version = "0.2.7"
|
||||
version = "0.2.8"
|
||||
dependencies = [
|
||||
"anyhow",
|
||||
"clap",
|
||||
@@ -1600,7 +1600,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "wafer-core"
|
||||
version = "0.2.7"
|
||||
version = "0.2.8"
|
||||
dependencies = [
|
||||
"anyhow",
|
||||
"insta",
|
||||
@@ -1615,7 +1615,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "wafer-web"
|
||||
version = "0.2.7"
|
||||
version = "0.2.8"
|
||||
dependencies = [
|
||||
"anyhow",
|
||||
"js-sys",
|
||||
|
||||
+1
-1
@@ -3,7 +3,7 @@ members = ["crates/*"]
|
||||
resolver = "2"
|
||||
|
||||
[workspace.package]
|
||||
version = "0.2.7"
|
||||
version = "0.2.8"
|
||||
edition = "2024"
|
||||
license = "MIT OR Apache-2.0"
|
||||
repository = "https://github.com/ok2/wafer"
|
||||
|
||||
@@ -8,7 +8,7 @@ An optimizing Forth 2012 compiler targeting WebAssembly. WAFER JIT-compiles each
|
||||
|
||||
- **200+ words** across 12 Forth 2012 word sets, all at **100% compliance**
|
||||
- **Optimizing compiler** with 6 IR passes + stack-to-local promotion (per region, so a hot loop keeps its registers even inside a word that does I/O; `DO` and `BEGIN` loops alike) + consolidation
|
||||
- **Faster than gforth** on all benchmarks in release mode (2-10x faster)
|
||||
- **Faster than gforth** on every benchmark, and past SwiftForth `sf64` -- a native-code compiler -- on four of five
|
||||
- **JIT compilation** — each `:` definition compiles to its own WASM module
|
||||
- **Self-recursive direct calls** — RECURSE compiles to native `call` instead of `call_indirect`
|
||||
- **Typed calling convention** — a word with a statically known stack effect passes its stack items as WASM values, so a call keeps them in registers instead of round-tripping through memory
|
||||
@@ -80,24 +80,47 @@ git submodule update --init
|
||||
|
||||
## Performance
|
||||
|
||||
WAFER beats gforth (the GNU Forth reference implementation) on all benchmarks in release mode, and is within
|
||||
reach of SwiftForth `sf64`, which compiles to native code:
|
||||
WAFER beats gforth (the GNU Forth reference implementation) on every benchmark, and SwiftForth
|
||||
`sf64` -- which compiles to native code -- on four of the five.
|
||||
|
||||
Measured with all three engines running **native x86-64**, on an idle 16-vCPU Xeon Platinum 8124M
|
||||
@ 3.0 GHz (median of three runs):
|
||||
|
||||
```
|
||||
Benchmark WAFER gforth sf64 WAFER/gf WAFER/sf
|
||||
Fibonacci(25) 411 3221 355 0.13x 1.16x
|
||||
Factorial(12)x100K 994 7141 3058 0.14x 0.33x
|
||||
GCD-bench(20K) 1591 3211 2423 0.50x 0.66x
|
||||
NestedLoops(50)x1K 889 6824 2342 0.13x 0.38x
|
||||
Collatz(2K) 391 3981 1659 0.10x 0.24x
|
||||
```
|
||||
|
||||
Times in microseconds; WAFER is the better of the JIT and `CONSOLIDATE` runs. Below 1.0 means WAFER
|
||||
is faster. Fibonacci is the one WAFER loses: it is one call per node with no loop to promote, and
|
||||
`sf64` keeps its stack in registers across a call the way only a native code generator can.
|
||||
Fibonacci, GCD and Collatz held to within 2% across the three runs; Factorial and NestedLoops are
|
||||
softer, since `sf64` varied by half there, but they are wide wins either way.
|
||||
|
||||
`just bench-compare` on the development machine (M1 Ultra, arm64) reports different numbers, and
|
||||
they flatter WAFER:
|
||||
|
||||
```
|
||||
Benchmark WAFER CONSOL gforth sf64 WAFER/gf WAFER/sf
|
||||
Fibonacci(25) 356 361 3389 287 0.11x 1.24x
|
||||
Factorial(12)x100K 479 495 6249 1650 0.08x 0.29x
|
||||
GCD-bench(20K) 540 559 1801 801 0.30x 0.67x
|
||||
NestedLoops(50)x1K 509 501 7023 1887 0.07x 0.27x
|
||||
Collatz(2K) 185 213 3873 610 0.05x 0.30x
|
||||
Fibonacci(25) 237 242 3340 287 0.07x 0.83x
|
||||
Factorial(12)x100K 480 479 6109 1594 0.08x 0.30x
|
||||
GCD-bench(20K) 549 541 1830 797 0.30x 0.68x
|
||||
NestedLoops(50)x1K 501 509 7092 1898 0.07x 0.26x
|
||||
Collatz(2K) 196 190 3955 633 0.05x 0.30x
|
||||
```
|
||||
|
||||
Times in microseconds. WAFER/gf < 1.0 means WAFER is faster. CONSOL = after `CONSOLIDATE`.
|
||||
The only SwiftForth build for macOS is x86-64 under Rosetta 2, while WAFER and gforth are native
|
||||
arm64 -- so that `sf64` column is native against emulated. The gap is not small, and it lands
|
||||
exactly where it matters: Fibonacci reads 0.83x there and 1.16x when neither engine is emulated.
|
||||
Treat the arm64 table as what the regression limits in `comparison.rs` are calibrated against, and
|
||||
the x86-64 table as what to believe about the engines.
|
||||
|
||||
Two caveats on the `sf64` column. The SwiftForth build here is x86-64 running under Rosetta 2
|
||||
while WAFER and gforth are native arm64, so it is a native-vs-emulated comparison; and sf64
|
||||
uses 64-bit cells to WAFER's 32-bit. WAFER is ahead on the four loop-heavy benchmarks and
|
||||
behind on Fibonacci, which is one call per node with no loop to promote.
|
||||
A caveat applies to both: `sf64` uses 64-bit cells to WAFER's 32-bit, so WAFER does less work per
|
||||
operation.
|
||||
|
||||
A word whose stack effect is statically known gets a **typed entry point**: its stack items travel in and out
|
||||
as WASM values instead of through the memory data stack, so cranelift keeps them in registers across a call
|
||||
@@ -106,10 +129,15 @@ table, `EXECUTE` and the outer interpreter reach, so nothing about the memory AB
|
||||
Call-heavy code is what this pays for -- Fibonacci went from 4.3x slower than `sf64` to 1.2x. Set
|
||||
`WAFER_TYPED_CALLS=0` to fall back to the memory-stack convention.
|
||||
|
||||
Recursive words then get one more thing: their base-case guard is tested at the **call site**, so a
|
||||
leaf of the recursion costs a comparison instead of a call. `: FIB DUP 2 < IF EXIT THEN ... RECURSE`
|
||||
compiles its `RECURSE` as `DUP 2 < IF ELSE RECURSE THEN`, which is what the callee would have done
|
||||
on entry anyway. Half of fib's nodes are leaves, and that is worth 1.4x.
|
||||
|
||||
## Testing
|
||||
|
||||
```bash
|
||||
# All tests (~628 currently passing)
|
||||
# All tests (~635 currently passing)
|
||||
cargo test --workspace
|
||||
|
||||
# Forth 2012 compliance suite
|
||||
@@ -142,7 +170,7 @@ Forth Source -> Outer Interpreter -> IR -> [Optimize] -> WASM Codegen (wasm-enco
|
||||
- `WebRuntime` — browser WebAssembly API via js-sys, for the browser REPL
|
||||
- **Subroutine threading** via WASM function tables (`call_indirect` for cross-word, direct `call` for self-recursion)
|
||||
- **JIT mode**: each new word compiles to a separate WASM module linked to shared memory/globals/table
|
||||
- **IR-based pipeline** with 6 optimization passes (peephole, constant folding, strength reduction, DCE, tail call detection, inlining) plus per-region stack-to-local promotion (DO and BEGIN loops, IF/ELSE), DO/LOOP index locals, typed entry points for words with a known stack effect, and consolidation
|
||||
- **IR-based pipeline** with 6 optimization passes (peephole, constant folding, strength reduction, DCE, tail call detection, inlining) plus per-region stack-to-local promotion (DO and BEGIN loops, IF/ELSE), DO/LOOP index locals, typed entry points for words with a known stack effect, self-guard expansion, and consolidation
|
||||
- **Dictionary**: linked-list word headers in simulated linear memory
|
||||
|
||||
## Project Structure
|
||||
|
||||
@@ -9,7 +9,7 @@ license.workspace = true
|
||||
workspace = true
|
||||
|
||||
[dependencies]
|
||||
wafer-core = { path = "../core", version = "0.2.7" }
|
||||
wafer-core = { path = "../core", version = "0.2.8" }
|
||||
wasmtime = { workspace = true }
|
||||
anyhow = { workspace = true }
|
||||
clap = { version = "4", features = ["derive"] }
|
||||
|
||||
@@ -39,6 +39,7 @@ impl WaferConfig {
|
||||
strength_reduce: true,
|
||||
dce: true,
|
||||
inline: true,
|
||||
self_guard: true,
|
||||
},
|
||||
codegen: CodegenOpts {
|
||||
stack_to_local_promotion: true,
|
||||
@@ -58,6 +59,7 @@ impl WaferConfig {
|
||||
strength_reduce: false,
|
||||
dce: false,
|
||||
inline: false,
|
||||
self_guard: false,
|
||||
},
|
||||
codegen: CodegenOpts {
|
||||
stack_to_local_promotion: false,
|
||||
|
||||
@@ -27,6 +27,9 @@ pub struct OptConfig {
|
||||
pub dce: bool,
|
||||
/// Enable inlining of small word bodies.
|
||||
pub inline: bool,
|
||||
/// Expand a recursive word's base-case guard into its own call sites, so
|
||||
/// the leaves of the recursion cost a test instead of a call.
|
||||
pub self_guard: bool,
|
||||
}
|
||||
|
||||
/// Run all enabled optimization passes.
|
||||
@@ -34,6 +37,7 @@ pub fn optimize(
|
||||
ops: Vec<IrOp>,
|
||||
config: &OptConfig,
|
||||
bodies: &HashMap<WordId, Vec<IrOp>>,
|
||||
self_id: Option<WordId>,
|
||||
) -> Vec<IrOp> {
|
||||
let mut ir = ops;
|
||||
|
||||
@@ -60,6 +64,11 @@ pub fn optimize(
|
||||
let keep_loops_out = !crate::codegen::promotable_modulo_calls(&ir);
|
||||
ir = inline(ir, bodies, 8, keep_loops_out);
|
||||
}
|
||||
if config.self_guard
|
||||
&& let Some(id) = self_id
|
||||
{
|
||||
ir = expand_self_guard(ir, id);
|
||||
}
|
||||
if config.peephole {
|
||||
ir = peephole(ir);
|
||||
}
|
||||
@@ -585,6 +594,142 @@ fn detailcall(op: IrOp) -> IrOp {
|
||||
}
|
||||
|
||||
/// Check if an IR body contains a direct call to the given word (recursion guard).
|
||||
/// Largest guard the expander is willing to run twice, in IR operations.
|
||||
const MAX_GUARD_OPS: usize = 6;
|
||||
/// Most self-call sites worth expanding, to bound the code growth.
|
||||
const MAX_GUARD_SITES: usize = 4;
|
||||
|
||||
/// Expand a recursive word's base-case guard into its own call sites.
|
||||
///
|
||||
/// A recursive Forth word almost always opens with a guard that returns early
|
||||
/// -- `: FIB DUP 2 < IF EXIT THEN ... RECURSE ... ;` -- so every leaf of the
|
||||
/// recursion costs a call whose whole body is that test. Testing at the call
|
||||
/// site instead removes the call for the leaves, which in fib's tree is half
|
||||
/// of all nodes.
|
||||
///
|
||||
/// `Call(self)` becomes `<guard> IF <what the guard returns> ELSE Call(self)
|
||||
/// THEN`, which computes the same thing: the callee would have run the guard,
|
||||
/// taken the branch and returned. The price is that the guard runs twice along
|
||||
/// the recursive path, which is why it has to be small and free of effects.
|
||||
fn expand_self_guard(ops: Vec<IrOp>, self_id: WordId) -> Vec<IrOp> {
|
||||
let Some((cond, base)) = split_guard(&ops) else {
|
||||
return ops;
|
||||
};
|
||||
if count_self_calls(&ops, self_id) > MAX_GUARD_SITES {
|
||||
return ops;
|
||||
}
|
||||
let (cond, base) = (cond.to_vec(), base.to_vec());
|
||||
replace_self_calls(ops, self_id, &cond, &base)
|
||||
}
|
||||
|
||||
/// Split a body into the condition of its leading base-case guard and what
|
||||
/// that guard leaves behind, or `None` if it does not open with one.
|
||||
fn split_guard(ops: &[IrOp]) -> Option<(&[IrOp], &[IrOp])> {
|
||||
let at = ops.iter().position(|op| matches!(op, IrOp::If { .. }))?;
|
||||
let cond = &ops[..at];
|
||||
if at > MAX_GUARD_OPS || !cond.iter().all(is_duplicable) {
|
||||
return None;
|
||||
}
|
||||
let IrOp::If {
|
||||
then_body,
|
||||
else_body: None,
|
||||
} = &ops[at]
|
||||
else {
|
||||
return None;
|
||||
};
|
||||
// The guard is only a guard if it returns; what precedes the `EXIT` is
|
||||
// the value it returns, and has to be as harmless as the condition.
|
||||
let (IrOp::Exit, base) = then_body.split_last()? else {
|
||||
return None;
|
||||
};
|
||||
if base.len() > MAX_GUARD_OPS || !base.iter().all(is_duplicable) {
|
||||
return None;
|
||||
}
|
||||
Some((cond, base))
|
||||
}
|
||||
|
||||
/// Can this operation be duplicated at every call site -- cheap, effect-free,
|
||||
/// and not itself a call or a branch?
|
||||
fn is_duplicable(op: &IrOp) -> bool {
|
||||
matches!(
|
||||
op,
|
||||
IrOp::PushI32(_)
|
||||
| IrOp::Drop
|
||||
| IrOp::Dup
|
||||
| IrOp::Swap
|
||||
| IrOp::Over
|
||||
| IrOp::Rot
|
||||
| IrOp::Nip
|
||||
| IrOp::Tuck
|
||||
| IrOp::TwoDup
|
||||
| IrOp::TwoDrop
|
||||
| IrOp::Add
|
||||
| IrOp::Sub
|
||||
| IrOp::Mul
|
||||
| IrOp::Negate
|
||||
| IrOp::Abs
|
||||
| IrOp::Eq
|
||||
| IrOp::NotEq
|
||||
| IrOp::Lt
|
||||
| IrOp::Gt
|
||||
| IrOp::LtUnsigned
|
||||
| IrOp::ZeroEq
|
||||
| IrOp::ZeroLt
|
||||
| IrOp::And
|
||||
| IrOp::Or
|
||||
| IrOp::Xor
|
||||
| IrOp::Invert
|
||||
| IrOp::Lshift
|
||||
| IrOp::Rshift
|
||||
| IrOp::ArithRshift
|
||||
)
|
||||
}
|
||||
|
||||
fn count_self_calls(ops: &[IrOp], self_id: WordId) -> usize {
|
||||
ops.iter()
|
||||
.map(|op| match op {
|
||||
IrOp::Call(id) if *id == self_id => 1,
|
||||
IrOp::If {
|
||||
then_body,
|
||||
else_body,
|
||||
} => {
|
||||
count_self_calls(then_body, self_id)
|
||||
+ else_body
|
||||
.as_deref()
|
||||
.map_or(0, |eb| count_self_calls(eb, self_id))
|
||||
}
|
||||
_ => 0,
|
||||
})
|
||||
.sum()
|
||||
}
|
||||
|
||||
/// Wrap every `Call(self_id)` in the guard. Only plain calls: a `TailCall` is
|
||||
/// followed by a return, and leaving those alone keeps tail-call detection and
|
||||
/// this pass from having to agree about what tail position means.
|
||||
fn replace_self_calls(ops: Vec<IrOp>, self_id: WordId, cond: &[IrOp], base: &[IrOp]) -> Vec<IrOp> {
|
||||
let mut out = Vec::with_capacity(ops.len());
|
||||
for op in ops {
|
||||
match op {
|
||||
IrOp::Call(id) if id == self_id => {
|
||||
out.extend_from_slice(cond);
|
||||
out.push(IrOp::If {
|
||||
then_body: base.to_vec(),
|
||||
else_body: Some(vec![IrOp::Call(id)]),
|
||||
});
|
||||
}
|
||||
IrOp::If {
|
||||
then_body,
|
||||
else_body,
|
||||
} => out.push(IrOp::If {
|
||||
then_body: replace_self_calls(then_body, self_id, cond, base),
|
||||
else_body: else_body.map(|eb| replace_self_calls(eb, self_id, cond, base)),
|
||||
}),
|
||||
other => out.push(other),
|
||||
}
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
fn contains_call_to(ops: &[IrOp], target: WordId) -> bool {
|
||||
for op in ops {
|
||||
match op {
|
||||
@@ -746,8 +891,130 @@ mod tests {
|
||||
strength_reduce: true,
|
||||
dce: true,
|
||||
inline: false,
|
||||
self_guard: false,
|
||||
};
|
||||
optimize(ops, &config, &HashMap::new())
|
||||
optimize(ops, &config, &HashMap::new(), None)
|
||||
}
|
||||
|
||||
/// A body shaped like a recursive Forth word: a base-case guard, then the
|
||||
/// recursive step. `SELF` is the word being compiled.
|
||||
const SELF: WordId = WordId(9);
|
||||
|
||||
fn guarded_body(step: Vec<IrOp>) -> Vec<IrOp> {
|
||||
let mut ops = vec![
|
||||
IrOp::Dup,
|
||||
IrOp::PushI32(2),
|
||||
IrOp::Lt,
|
||||
IrOp::If {
|
||||
then_body: vec![IrOp::Exit],
|
||||
else_body: None,
|
||||
},
|
||||
];
|
||||
ops.extend(step);
|
||||
ops
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn self_guard_moves_the_base_case_to_the_call_site() {
|
||||
let out = expand_self_guard(guarded_body(vec![IrOp::Call(SELF)]), SELF);
|
||||
assert_eq!(
|
||||
out,
|
||||
guarded_body(vec![
|
||||
IrOp::Dup,
|
||||
IrOp::PushI32(2),
|
||||
IrOp::Lt,
|
||||
IrOp::If {
|
||||
then_body: vec![],
|
||||
else_body: Some(vec![IrOp::Call(SELF)]),
|
||||
},
|
||||
])
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn self_guard_carries_the_value_the_guard_returns() {
|
||||
// `: F DUP 2 < IF DROP 0 EXIT THEN RECURSE ;` -- the base case is not
|
||||
// "leave the argument", it is "replace it with 0".
|
||||
let body = vec![
|
||||
IrOp::Dup,
|
||||
IrOp::PushI32(2),
|
||||
IrOp::Lt,
|
||||
IrOp::If {
|
||||
then_body: vec![IrOp::Drop, IrOp::PushI32(0), IrOp::Exit],
|
||||
else_body: None,
|
||||
},
|
||||
IrOp::Call(SELF),
|
||||
];
|
||||
let out = expand_self_guard(body, SELF);
|
||||
let IrOp::If { then_body, .. } = &out[7] else {
|
||||
panic!("expected the expanded guard at index 7, got {:?}", out);
|
||||
};
|
||||
assert_eq!(then_body, &vec![IrOp::Drop, IrOp::PushI32(0)]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn self_guard_leaves_a_body_without_a_guard_alone() {
|
||||
// An `IF` with an `ELSE` is a branch, not an early return.
|
||||
let body = vec![
|
||||
IrOp::Dup,
|
||||
IrOp::If {
|
||||
then_body: vec![IrOp::Drop],
|
||||
else_body: Some(vec![IrOp::Call(SELF)]),
|
||||
},
|
||||
];
|
||||
assert_eq!(expand_self_guard(body.clone(), SELF), body);
|
||||
|
||||
// No `EXIT` in the then-branch: also not a guard.
|
||||
let body = guarded_body(vec![IrOp::Call(SELF)])
|
||||
.into_iter()
|
||||
.map(|op| match op {
|
||||
IrOp::If { .. } => IrOp::If {
|
||||
then_body: vec![IrOp::Drop],
|
||||
else_body: None,
|
||||
},
|
||||
other => other,
|
||||
})
|
||||
.collect::<Vec<_>>();
|
||||
assert_eq!(expand_self_guard(body.clone(), SELF), body);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn self_guard_refuses_a_condition_it_cannot_run_twice() {
|
||||
// A guard reached through a call or a memory write would be evaluated
|
||||
// once at the call site and again inside the callee.
|
||||
let body = vec![
|
||||
IrOp::Call(WordId(3)),
|
||||
IrOp::If {
|
||||
then_body: vec![IrOp::Exit],
|
||||
else_body: None,
|
||||
},
|
||||
IrOp::Call(SELF),
|
||||
];
|
||||
assert_eq!(expand_self_guard(body.clone(), SELF), body);
|
||||
|
||||
let body = vec![
|
||||
IrOp::Dup,
|
||||
IrOp::Fetch,
|
||||
IrOp::If {
|
||||
then_body: vec![IrOp::Exit],
|
||||
else_body: None,
|
||||
},
|
||||
IrOp::Call(SELF),
|
||||
];
|
||||
assert_eq!(expand_self_guard(body.clone(), SELF), body);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn self_guard_stops_at_the_call_site_budget() {
|
||||
let step = std::iter::repeat_n(IrOp::Call(SELF), MAX_GUARD_SITES + 1).collect();
|
||||
let body = guarded_body(step);
|
||||
assert_eq!(expand_self_guard(body.clone(), SELF), body);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn self_guard_leaves_tail_calls_alone() {
|
||||
let body = guarded_body(vec![IrOp::TailCall(SELF)]);
|
||||
assert_eq!(expand_self_guard(body.clone(), SELF), body);
|
||||
}
|
||||
|
||||
fn opt_with_inline(ops: Vec<IrOp>, bodies: &HashMap<WordId, Vec<IrOp>>) -> Vec<IrOp> {
|
||||
@@ -758,8 +1025,9 @@ mod tests {
|
||||
strength_reduce: true,
|
||||
dce: true,
|
||||
inline: true,
|
||||
self_guard: false,
|
||||
};
|
||||
optimize(ops, &config, bodies)
|
||||
optimize(ops, &config, bodies, None)
|
||||
}
|
||||
|
||||
// Peephole tests
|
||||
@@ -1019,8 +1287,9 @@ mod tests {
|
||||
strength_reduce: false,
|
||||
dce: false,
|
||||
inline: true,
|
||||
self_guard: false,
|
||||
};
|
||||
let result = optimize(vec![IrOp::Call(WordId(5))], &config, &bodies);
|
||||
let result = optimize(vec![IrOp::Call(WordId(5))], &config, &bodies, None);
|
||||
assert_eq!(result, vec![IrOp::Call(WordId(5))]);
|
||||
}
|
||||
|
||||
|
||||
@@ -2463,8 +2463,13 @@ impl<R: Runtime> ForthVM<R> {
|
||||
}
|
||||
|
||||
/// Run all enabled optimization passes on an IR sequence.
|
||||
fn optimize_ir(&self, ir: Vec<IrOp>, bodies: &HashMap<WordId, Vec<IrOp>>) -> Vec<IrOp> {
|
||||
optimize(ir, &self.config.opt, bodies)
|
||||
fn optimize_ir(
|
||||
&self,
|
||||
ir: Vec<IrOp>,
|
||||
bodies: &HashMap<WordId, Vec<IrOp>>,
|
||||
self_id: Option<WordId>,
|
||||
) -> Vec<IrOp> {
|
||||
optimize(ir, &self.config.opt, bodies, self_id)
|
||||
}
|
||||
|
||||
/// Parse a `{: args | locals -- comment :}` block and compile local
|
||||
@@ -2576,7 +2581,7 @@ impl<R: Runtime> ForthVM<R> {
|
||||
|
||||
let ir = std::mem::take(&mut self.compiling_ir);
|
||||
let bodies = self.ir_bodies.clone();
|
||||
let ir = self.optimize_ir(ir, &bodies);
|
||||
let ir = self.optimize_ir(ir, &bodies, Some(word_id));
|
||||
self.ir_bodies.insert(word_id, ir.clone());
|
||||
|
||||
// Compile to WASM
|
||||
@@ -2992,7 +2997,7 @@ impl<R: Runtime> ForthVM<R> {
|
||||
ir_body: Vec<IrOp>,
|
||||
) -> anyhow::Result<WordId> {
|
||||
let bodies = self.ir_bodies.clone();
|
||||
let ir_body = self.optimize_ir(ir_body, &bodies);
|
||||
let ir_body = self.optimize_ir(ir_body, &bodies, None);
|
||||
let word_id = self
|
||||
.dictionary
|
||||
.create(name, immediate)
|
||||
@@ -8290,6 +8295,36 @@ mod tests {
|
||||
assert_eq!(eval_output(": D 1 2 3 2 0 DO ROT LOOP . . . ; D"), "2 1 3 ");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_self_guard_expansion_keeps_the_answers() {
|
||||
// The base-case guard is tested at the call site, so a leaf never
|
||||
// costs a call. All four verified against gforth.
|
||||
assert_eq!(
|
||||
eval_output(
|
||||
": FIB DUP 2 < IF EXIT THEN DUP 1- RECURSE SWAP 2 - RECURSE + ; \
|
||||
25 FIB . 0 FIB . 1 FIB . 2 FIB . 10 FIB ."
|
||||
),
|
||||
"75025 0 1 1 55 "
|
||||
);
|
||||
// A guard that replaces its argument rather than leaving it.
|
||||
assert_eq!(
|
||||
eval_output(": G DUP 0= IF DROP 0 EXIT THEN DUP 1- RECURSE + ; 5 G . 0 G . 100 G ."),
|
||||
"15 0 5050 "
|
||||
);
|
||||
assert_eq!(
|
||||
eval_output(": H DUP 3 < IF DROP 7 EXIT THEN 1- RECURSE 2 * ; 5 H . 2 H . 8 H ."),
|
||||
"56 7 448 "
|
||||
);
|
||||
// Two guards and three call sites, one of them behind the second guard.
|
||||
assert_eq!(
|
||||
eval_output(
|
||||
": ACK OVER 0= IF SWAP DROP 1+ EXIT THEN DUP 0= IF DROP 1- 1 RECURSE EXIT THEN \
|
||||
OVER SWAP 1- RECURSE SWAP 1- SWAP RECURSE ; 2 3 ACK . 1 2 ACK ."
|
||||
),
|
||||
"9 4 "
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_promoted_begin_loops() {
|
||||
// BEGIN loops promote too, so these run entirely in locals.
|
||||
|
||||
@@ -746,7 +746,7 @@ fn perf_benchmarks() -> Vec<PerfBenchmark> {
|
||||
verify: "25 FIB",
|
||||
expected: 75025,
|
||||
samples: 5,
|
||||
max_ratio: 0.17,
|
||||
max_ratio: 0.10,
|
||||
},
|
||||
PerfBenchmark {
|
||||
name: "Factorial(12)x100K",
|
||||
|
||||
@@ -12,7 +12,7 @@ workspace = true
|
||||
crate-type = ["cdylib", "rlib"]
|
||||
|
||||
[dependencies]
|
||||
wafer-core = { path = "../core", version = "0.2.7", default-features = false, features = ["crypto"] }
|
||||
wafer-core = { path = "../core", version = "0.2.8", default-features = false, features = ["crypto"] }
|
||||
wasm-bindgen = "0.2"
|
||||
js-sys = "0.3"
|
||||
send_wrapper = { workspace = true }
|
||||
|
||||
+63
-20
@@ -30,6 +30,7 @@ This document describes every optimization that makes sense for WAFER, why it ma
|
||||
| 14 | Self-Recursive Direct Call | Codegen | Done | High |
|
||||
| 15 | Float / Double-Cell | Codegen | Not started | Future |
|
||||
| 16 | Typed Calling Convention | Codegen | Done | Highest |
|
||||
| 17 | Self-Guard Expansion | IR pass | Done | Medium |
|
||||
|
||||
## 1. Stack-to-Local Promotion
|
||||
|
||||
@@ -484,33 +485,75 @@ Fibonacci(25) went from 1035 to 366 microseconds, 4.3x slower than `sf64` to 1.2
|
||||
|
||||
Untyped by design: anything using `SP@`, `DEPTH`, `EXECUTE`, `>R`/`R>`, floats or locals; anything calling a word that is itself untyped, which in the JIT path means every call except `RECURSE`; mutually recursive words; and words whose effect is not static -- branches that disagree on depth, `EXIT` at the wrong depth, a non-neutral loop body, or a recursion that grows the stack per level.
|
||||
|
||||
## 17. Self-Guard Expansion
|
||||
|
||||
**Status: Done.** A recursive word's base-case guard is duplicated into its own call sites, so the leaves of the recursion cost a test instead of a call. Implemented in `optimizer.rs::expand_self_guard`, gated on `OptConfig::self_guard`, and applied after inlining so the later passes still run over the result.
|
||||
|
||||
### The Shape
|
||||
|
||||
A recursive Forth word almost always opens with a guard that returns early:
|
||||
|
||||
```forth
|
||||
: FIB DUP 2 < IF EXIT THEN DUP 1- RECURSE SWAP 2 - RECURSE + ;
|
||||
```
|
||||
|
||||
Every leaf of the recursion is then a call whose entire body is `DUP 2 <`. The pass rewrites each `Call(self)` as
|
||||
|
||||
```forth
|
||||
DUP 2 < IF ( leave it ) ELSE RECURSE THEN
|
||||
```
|
||||
|
||||
which computes the same thing -- the callee would have run the guard, taken the branch and returned. When the guard returns a value rather than its argument (`IF DROP 0 EXIT THEN`), that value moves into the then-branch with it.
|
||||
|
||||
### Why It Is Bounded
|
||||
|
||||
The guard runs twice along the recursive path: once at the call site, once inside the callee. So it must be small and free of effects -- `MAX_GUARD_OPS` is six, and the operations are restricted to stack shuffles, arithmetic and comparisons; a call, a memory access or a branch disqualifies it. `MAX_GUARD_SITES` caps the expansion at four call sites, since each one replicates the guard. A `TailCall` is never expanded, which keeps this pass and tail-call detection from having to agree about what tail position means.
|
||||
|
||||
### Impact
|
||||
|
||||
Fibonacci(25): 356 to 237 microseconds on the arm64 development machine. In fib's tree half of all nodes are leaves, which is where the factor comes from. It does not take Fibonacci past `sf64`, though the arm64 table below says otherwise: with both engines native on x86-64, Fibonacci reads 1.16x and stays the one benchmark `sf64` wins.
|
||||
|
||||
## Current Performance vs Gforth
|
||||
|
||||
All optimizations enabled, release mode, measured with UTIME:
|
||||
|
||||
All three engines **native x86-64**, idle 16-vCPU Xeon Platinum 8124M @ 3.0 GHz,
|
||||
median of three runs:
|
||||
|
||||
```
|
||||
Benchmark WAFER CONSOL gforth sf64 WAFER/gf WAFER/sf
|
||||
Fibonacci(25) 356 361 3389 287 0.11x 1.24x
|
||||
Factorial(12)x100K 479 495 6249 1650 0.08x 0.29x
|
||||
GCD-bench(20K) 540 559 1801 801 0.30x 0.67x
|
||||
NestedLoops(50)x1K 509 501 7023 1887 0.07x 0.27x
|
||||
Collatz(2K) 185 213 3873 610 0.05x 0.30x
|
||||
Benchmark WAFER gforth sf64 WAFER/gf WAFER/sf
|
||||
Fibonacci(25) 411 3221 355 0.13x 1.16x
|
||||
Factorial(12)x100K 994 7141 3058 0.14x 0.33x
|
||||
GCD-bench(20K) 1591 3211 2423 0.50x 0.66x
|
||||
NestedLoops(50)x1K 889 6824 2342 0.13x 0.38x
|
||||
Collatz(2K) 391 3981 1659 0.10x 0.24x
|
||||
```
|
||||
|
||||
Times in microseconds. WAFER/gf < 1.0 means WAFER is faster. `sf64` is SwiftForth,
|
||||
which compiles to native code; two caveats on that column. The install here is an
|
||||
x86-64 binary under Rosetta 2 while WAFER and gforth are native arm64, so it is a
|
||||
native-vs-emulated comparison and a native SwiftForth would be faster than these
|
||||
numbers; and sf64 uses 64-bit cells to WAFER's 32-bit. WAFER is ahead on the four
|
||||
loop-heavy benchmarks and behind on Fibonacci, which is one call per node with no
|
||||
loop to promote.
|
||||
The same suite on the arm64 development machine (M1 Ultra), which is what the
|
||||
regression limits in `comparison.rs` are calibrated against:
|
||||
|
||||
```
|
||||
Benchmark WAFER CONSOL gforth sf64 WAFER/gf WAFER/sf
|
||||
Fibonacci(25) 237 242 3340 287 0.07x 0.83x
|
||||
Factorial(12)x100K 480 479 6109 1594 0.08x 0.30x
|
||||
GCD-bench(20K) 549 541 1830 797 0.30x 0.68x
|
||||
NestedLoops(50)x1K 501 509 7092 1898 0.07x 0.26x
|
||||
Collatz(2K) 196 190 3955 633 0.05x 0.30x
|
||||
```
|
||||
|
||||
Times in microseconds. WAFER/gf < 1.0 means WAFER is faster. The two tables
|
||||
disagree because the only SwiftForth build for macOS is x86-64 under Rosetta 2
|
||||
while WAFER and gforth are native arm64 -- and the emulation penalty lands
|
||||
hardest on the call-heavy benchmark, so Fibonacci reads 0.83x on arm64 and 1.16x
|
||||
when neither engine is emulated. Believe the x86-64 table about the engines. One
|
||||
caveat holds for both: sf64 uses 64-bit cells to WAFER's 32-bit.
|
||||
|
||||
## Remaining Opportunities
|
||||
|
||||
| Optimization | Status | Potential Impact |
|
||||
| -------------------------------- | ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Bounded self-inlining | Not started | Measured 1.33x on Fibonacci, the last benchmark behind sf64. Blocked on `EXIT`: the inliner refuses any body containing one, and a recursive Forth word is `... IF EXIT THEN ... RECURSE`. Needs either a scoped exit (compile an inlined `EXIT` as a branch to the end of a block) or guard-only expansion |
|
||||
| BeginDoubleWhileRepeat promotion | Not started | Rare pattern, low priority. Its promoted emitter exists but has no loop fixup and is unverified |
|
||||
| LEAVE as IR primitive | Not started | Would enable fast-path for loops with LEAVE |
|
||||
| Float stack-to-local | Not started | Eliminate float stack memory traffic |
|
||||
| WASM tail calls proposal | Waiting on wasmtime | Would eliminate stack growth for tail-recursive words |
|
||||
| Optimization | Status | Potential Impact |
|
||||
| -------------------------------- | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Scoped exit for the inliner | Not started | The inliner still refuses any body containing an `EXIT`, because an inlined one would return from the caller. Compiling it as a branch to the end of a block would unlock inlining for every word with an early return, not just the guard shape section 17 handles |
|
||||
| BeginDoubleWhileRepeat promotion | Not started | Rare pattern, low priority. Its promoted emitter exists but has no loop fixup and is unverified |
|
||||
| LEAVE as IR primitive | Not started | Would enable fast-path for loops with LEAVE |
|
||||
| Float stack-to-local | Not started | Eliminate float stack memory traffic |
|
||||
| WASM tail calls proposal | Waiting on wasmtime | Would eliminate stack growth for tail-recursive words |
|
||||
|
||||
Reference in New Issue
Block a user