A recursive Forth word almost always opens with a guard that returns early, so every leaf of the recursion costs a call whose whole body is that test. `Call(self)` now compiles as `<guard> IF <what the guard returns> ELSE Call(self) THEN`, which is what the callee would have done on entry anyway. Half of fib's nodes are leaves: Fibonacci(25) 356 -> 237 us, 1.24x sf64 -> 0.83x, so all five benchmarks now beat it. The guard runs twice along the recursive path, hence the bounds: at most six effect-free operations, at most four call sites, never a tail call. WS-018.
This commit is contained in:
@@ -5,6 +5,26 @@ All notable changes to WAFER are documented in this file.
|
||||
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
||||
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
||||
|
||||
## [Unreleased]
|
||||
|
||||
### Added
|
||||
|
||||
- **A recursive word tests its base case at the call site.** A recursive Forth
|
||||
word almost always opens with a guard that returns early --
|
||||
`: FIB DUP 2 < IF EXIT THEN ... RECURSE ... ;` -- so every leaf of the
|
||||
recursion costs a call whose entire body is that test. `Call(self)` now
|
||||
compiles as `<guard> IF <what the guard returns> ELSE Call(self) THEN`,
|
||||
which computes the same thing: the callee would have run the guard, taken
|
||||
the branch and returned. In fib's tree the leaves are half of all nodes.
|
||||
|
||||
Fibonacci(25) 356 -> 237 µs, which takes the last benchmark that was behind
|
||||
SwiftForth `sf64` past it: 1.24x -> 0.83x, five of five.
|
||||
|
||||
The guard runs twice along the recursive path, so it has to be small (at
|
||||
most six operations) and free of effects -- no calls, no memory, no
|
||||
branches. Words with more than four self-call sites are left alone to bound
|
||||
the code growth, and a `TailCall` is never expanded.
|
||||
|
||||
## [0.2.7] - 2026-08-09
|
||||
|
||||
### Added
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
## What is WAFER?
|
||||
|
||||
WAFER (WebAssembly Forth Engine in Rust) is an optimizing Forth 2012 compiler targeting WebAssembly. Currently a working Forth system with 200+ words, JIT compilation, 12 word sets at 100% compliance, and a full optimization pipeline (peephole, constant folding, inlining, strength reduction, DCE, tail calls, per-region stack-to-local promotion with DO/BEGIN loop and IF support, self-recursive direct calls, a typed calling convention for words with a known stack effect, consolidation). Beats gforth on all benchmarks in release mode, and SwiftForth `sf64` on four of five. Includes a browser-based REPL via wasm-pack.
|
||||
WAFER (WebAssembly Forth Engine in Rust) is an optimizing Forth 2012 compiler targeting WebAssembly. Currently a working Forth system with 200+ words, JIT compilation, 12 word sets at 100% compliance, and a full optimization pipeline (peephole, constant folding, inlining, strength reduction, DCE, tail calls, per-region stack-to-local promotion with DO/BEGIN loop and IF support, self-recursive direct calls, a typed calling convention for words with a known stack effect, self-guard expansion for recursive words, consolidation). Beats gforth on all benchmarks in release mode (3-20x), and SwiftForth `sf64` on all five. Includes a browser-based REPL via wasm-pack.
|
||||
|
||||
## Architecture
|
||||
|
||||
@@ -79,7 +79,7 @@ Handle in `interpret_token_immediate()` or `compile_token()` as a special case.
|
||||
|
||||
## Testing
|
||||
|
||||
- Run `cargo test --workspace` before committing (currently 601 unit + 1 benchmark + 12 compliance + 9 comparison + 5 crypto)
|
||||
- Run `cargo test --workspace` before committing (currently 608 unit + 1 benchmark + 12 compliance + 9 comparison + 5 crypto)
|
||||
- Forth 2012 compliance: `cargo test -p wafer-core --test compliance`
|
||||
- Cross-engine comparison (vs gforth): `cargo test -p wafer-core --test comparison`
|
||||
- Performance benchmarks (release mode): `cargo test -p wafer-core --test comparison -- --nocapture --ignored`
|
||||
|
||||
@@ -8,7 +8,7 @@ An optimizing Forth 2012 compiler targeting WebAssembly. WAFER JIT-compiles each
|
||||
|
||||
- **200+ words** across 12 Forth 2012 word sets, all at **100% compliance**
|
||||
- **Optimizing compiler** with 6 IR passes + stack-to-local promotion (per region, so a hot loop keeps its registers even inside a word that does I/O; `DO` and `BEGIN` loops alike) + consolidation
|
||||
- **Faster than gforth** on all benchmarks in release mode (2-10x faster)
|
||||
- **Faster than gforth** on all benchmarks in release mode (3-20x), and past SwiftForth `sf64` on all five
|
||||
- **JIT compilation** — each `:` definition compiles to its own WASM module
|
||||
- **Self-recursive direct calls** — RECURSE compiles to native `call` instead of `call_indirect`
|
||||
- **Typed calling convention** — a word with a statically known stack effect passes its stack items as WASM values, so a call keeps them in registers instead of round-tripping through memory
|
||||
@@ -80,24 +80,23 @@ git submodule update --init
|
||||
|
||||
## Performance
|
||||
|
||||
WAFER beats gforth (the GNU Forth reference implementation) on all benchmarks in release mode, and is within
|
||||
reach of SwiftForth `sf64`, which compiles to native code:
|
||||
WAFER beats gforth (the GNU Forth reference implementation) on all benchmarks in release mode, and
|
||||
SwiftForth `sf64`, which compiles to native code, on all five:
|
||||
|
||||
```
|
||||
Benchmark WAFER CONSOL gforth sf64 WAFER/gf WAFER/sf
|
||||
Fibonacci(25) 356 361 3389 287 0.11x 1.24x
|
||||
Factorial(12)x100K 479 495 6249 1650 0.08x 0.29x
|
||||
GCD-bench(20K) 540 559 1801 801 0.30x 0.67x
|
||||
NestedLoops(50)x1K 509 501 7023 1887 0.07x 0.27x
|
||||
Collatz(2K) 185 213 3873 610 0.05x 0.30x
|
||||
Fibonacci(25) 237 242 3340 287 0.07x 0.83x
|
||||
Factorial(12)x100K 480 479 6109 1594 0.08x 0.30x
|
||||
GCD-bench(20K) 549 541 1830 797 0.30x 0.68x
|
||||
NestedLoops(50)x1K 501 509 7092 1898 0.07x 0.26x
|
||||
Collatz(2K) 196 190 3955 633 0.05x 0.30x
|
||||
```
|
||||
|
||||
Times in microseconds. WAFER/gf < 1.0 means WAFER is faster. CONSOL = after `CONSOLIDATE`.
|
||||
|
||||
Two caveats on the `sf64` column. The SwiftForth build here is x86-64 running under Rosetta 2
|
||||
while WAFER and gforth are native arm64, so it is a native-vs-emulated comparison; and sf64
|
||||
uses 64-bit cells to WAFER's 32-bit. WAFER is ahead on the four loop-heavy benchmarks and
|
||||
behind on Fibonacci, which is one call per node with no loop to promote.
|
||||
uses 64-bit cells to WAFER's 32-bit.
|
||||
|
||||
A word whose stack effect is statically known gets a **typed entry point**: its stack items travel in and out
|
||||
as WASM values instead of through the memory data stack, so cranelift keeps them in registers across a call
|
||||
@@ -106,10 +105,15 @@ table, `EXECUTE` and the outer interpreter reach, so nothing about the memory AB
|
||||
Call-heavy code is what this pays for -- Fibonacci went from 4.3x slower than `sf64` to 1.2x. Set
|
||||
`WAFER_TYPED_CALLS=0` to fall back to the memory-stack convention.
|
||||
|
||||
Recursive words then get one more thing: their base-case guard is tested at the **call site**, so a
|
||||
leaf of the recursion costs a comparison instead of a call. `: FIB DUP 2 < IF EXIT THEN ... RECURSE`
|
||||
compiles its `RECURSE` as `DUP 2 < IF ELSE RECURSE THEN`, which is what the callee would have done
|
||||
on entry anyway. Half of fib's nodes are leaves, and that is the last 1.4x.
|
||||
|
||||
## Testing
|
||||
|
||||
```bash
|
||||
# All tests (~628 currently passing)
|
||||
# All tests (~635 currently passing)
|
||||
cargo test --workspace
|
||||
|
||||
# Forth 2012 compliance suite
|
||||
@@ -142,7 +146,7 @@ Forth Source -> Outer Interpreter -> IR -> [Optimize] -> WASM Codegen (wasm-enco
|
||||
- `WebRuntime` — browser WebAssembly API via js-sys, for the browser REPL
|
||||
- **Subroutine threading** via WASM function tables (`call_indirect` for cross-word, direct `call` for self-recursion)
|
||||
- **JIT mode**: each new word compiles to a separate WASM module linked to shared memory/globals/table
|
||||
- **IR-based pipeline** with 6 optimization passes (peephole, constant folding, strength reduction, DCE, tail call detection, inlining) plus per-region stack-to-local promotion (DO and BEGIN loops, IF/ELSE), DO/LOOP index locals, typed entry points for words with a known stack effect, and consolidation
|
||||
- **IR-based pipeline** with 6 optimization passes (peephole, constant folding, strength reduction, DCE, tail call detection, inlining) plus per-region stack-to-local promotion (DO and BEGIN loops, IF/ELSE), DO/LOOP index locals, typed entry points for words with a known stack effect, self-guard expansion, and consolidation
|
||||
- **Dictionary**: linked-list word headers in simulated linear memory
|
||||
|
||||
## Project Structure
|
||||
|
||||
@@ -39,6 +39,7 @@ impl WaferConfig {
|
||||
strength_reduce: true,
|
||||
dce: true,
|
||||
inline: true,
|
||||
self_guard: true,
|
||||
},
|
||||
codegen: CodegenOpts {
|
||||
stack_to_local_promotion: true,
|
||||
@@ -58,6 +59,7 @@ impl WaferConfig {
|
||||
strength_reduce: false,
|
||||
dce: false,
|
||||
inline: false,
|
||||
self_guard: false,
|
||||
},
|
||||
codegen: CodegenOpts {
|
||||
stack_to_local_promotion: false,
|
||||
|
||||
@@ -27,6 +27,9 @@ pub struct OptConfig {
|
||||
pub dce: bool,
|
||||
/// Enable inlining of small word bodies.
|
||||
pub inline: bool,
|
||||
/// Expand a recursive word's base-case guard into its own call sites, so
|
||||
/// the leaves of the recursion cost a test instead of a call.
|
||||
pub self_guard: bool,
|
||||
}
|
||||
|
||||
/// Run all enabled optimization passes.
|
||||
@@ -34,6 +37,7 @@ pub fn optimize(
|
||||
ops: Vec<IrOp>,
|
||||
config: &OptConfig,
|
||||
bodies: &HashMap<WordId, Vec<IrOp>>,
|
||||
self_id: Option<WordId>,
|
||||
) -> Vec<IrOp> {
|
||||
let mut ir = ops;
|
||||
|
||||
@@ -60,6 +64,11 @@ pub fn optimize(
|
||||
let keep_loops_out = !crate::codegen::promotable_modulo_calls(&ir);
|
||||
ir = inline(ir, bodies, 8, keep_loops_out);
|
||||
}
|
||||
if config.self_guard
|
||||
&& let Some(id) = self_id
|
||||
{
|
||||
ir = expand_self_guard(ir, id);
|
||||
}
|
||||
if config.peephole {
|
||||
ir = peephole(ir);
|
||||
}
|
||||
@@ -585,6 +594,142 @@ fn detailcall(op: IrOp) -> IrOp {
|
||||
}
|
||||
|
||||
/// Check if an IR body contains a direct call to the given word (recursion guard).
|
||||
/// Largest guard the expander is willing to run twice, in IR operations.
|
||||
const MAX_GUARD_OPS: usize = 6;
|
||||
/// Most self-call sites worth expanding, to bound the code growth.
|
||||
const MAX_GUARD_SITES: usize = 4;
|
||||
|
||||
/// Expand a recursive word's base-case guard into its own call sites.
|
||||
///
|
||||
/// A recursive Forth word almost always opens with a guard that returns early
|
||||
/// -- `: FIB DUP 2 < IF EXIT THEN ... RECURSE ... ;` -- so every leaf of the
|
||||
/// recursion costs a call whose whole body is that test. Testing at the call
|
||||
/// site instead removes the call for the leaves, which in fib's tree is half
|
||||
/// of all nodes.
|
||||
///
|
||||
/// `Call(self)` becomes `<guard> IF <what the guard returns> ELSE Call(self)
|
||||
/// THEN`, which computes the same thing: the callee would have run the guard,
|
||||
/// taken the branch and returned. The price is that the guard runs twice along
|
||||
/// the recursive path, which is why it has to be small and free of effects.
|
||||
fn expand_self_guard(ops: Vec<IrOp>, self_id: WordId) -> Vec<IrOp> {
|
||||
let Some((cond, base)) = split_guard(&ops) else {
|
||||
return ops;
|
||||
};
|
||||
if count_self_calls(&ops, self_id) > MAX_GUARD_SITES {
|
||||
return ops;
|
||||
}
|
||||
let (cond, base) = (cond.to_vec(), base.to_vec());
|
||||
replace_self_calls(ops, self_id, &cond, &base)
|
||||
}
|
||||
|
||||
/// Split a body into the condition of its leading base-case guard and what
|
||||
/// that guard leaves behind, or `None` if it does not open with one.
|
||||
fn split_guard(ops: &[IrOp]) -> Option<(&[IrOp], &[IrOp])> {
|
||||
let at = ops.iter().position(|op| matches!(op, IrOp::If { .. }))?;
|
||||
let cond = &ops[..at];
|
||||
if at > MAX_GUARD_OPS || !cond.iter().all(is_duplicable) {
|
||||
return None;
|
||||
}
|
||||
let IrOp::If {
|
||||
then_body,
|
||||
else_body: None,
|
||||
} = &ops[at]
|
||||
else {
|
||||
return None;
|
||||
};
|
||||
// The guard is only a guard if it returns; what precedes the `EXIT` is
|
||||
// the value it returns, and has to be as harmless as the condition.
|
||||
let (IrOp::Exit, base) = then_body.split_last()? else {
|
||||
return None;
|
||||
};
|
||||
if base.len() > MAX_GUARD_OPS || !base.iter().all(is_duplicable) {
|
||||
return None;
|
||||
}
|
||||
Some((cond, base))
|
||||
}
|
||||
|
||||
/// Can this operation be duplicated at every call site -- cheap, effect-free,
|
||||
/// and not itself a call or a branch?
|
||||
fn is_duplicable(op: &IrOp) -> bool {
|
||||
matches!(
|
||||
op,
|
||||
IrOp::PushI32(_)
|
||||
| IrOp::Drop
|
||||
| IrOp::Dup
|
||||
| IrOp::Swap
|
||||
| IrOp::Over
|
||||
| IrOp::Rot
|
||||
| IrOp::Nip
|
||||
| IrOp::Tuck
|
||||
| IrOp::TwoDup
|
||||
| IrOp::TwoDrop
|
||||
| IrOp::Add
|
||||
| IrOp::Sub
|
||||
| IrOp::Mul
|
||||
| IrOp::Negate
|
||||
| IrOp::Abs
|
||||
| IrOp::Eq
|
||||
| IrOp::NotEq
|
||||
| IrOp::Lt
|
||||
| IrOp::Gt
|
||||
| IrOp::LtUnsigned
|
||||
| IrOp::ZeroEq
|
||||
| IrOp::ZeroLt
|
||||
| IrOp::And
|
||||
| IrOp::Or
|
||||
| IrOp::Xor
|
||||
| IrOp::Invert
|
||||
| IrOp::Lshift
|
||||
| IrOp::Rshift
|
||||
| IrOp::ArithRshift
|
||||
)
|
||||
}
|
||||
|
||||
fn count_self_calls(ops: &[IrOp], self_id: WordId) -> usize {
|
||||
ops.iter()
|
||||
.map(|op| match op {
|
||||
IrOp::Call(id) if *id == self_id => 1,
|
||||
IrOp::If {
|
||||
then_body,
|
||||
else_body,
|
||||
} => {
|
||||
count_self_calls(then_body, self_id)
|
||||
+ else_body
|
||||
.as_deref()
|
||||
.map_or(0, |eb| count_self_calls(eb, self_id))
|
||||
}
|
||||
_ => 0,
|
||||
})
|
||||
.sum()
|
||||
}
|
||||
|
||||
/// Wrap every `Call(self_id)` in the guard. Only plain calls: a `TailCall` is
|
||||
/// followed by a return, and leaving those alone keeps tail-call detection and
|
||||
/// this pass from having to agree about what tail position means.
|
||||
fn replace_self_calls(ops: Vec<IrOp>, self_id: WordId, cond: &[IrOp], base: &[IrOp]) -> Vec<IrOp> {
|
||||
let mut out = Vec::with_capacity(ops.len());
|
||||
for op in ops {
|
||||
match op {
|
||||
IrOp::Call(id) if id == self_id => {
|
||||
out.extend_from_slice(cond);
|
||||
out.push(IrOp::If {
|
||||
then_body: base.to_vec(),
|
||||
else_body: Some(vec![IrOp::Call(id)]),
|
||||
});
|
||||
}
|
||||
IrOp::If {
|
||||
then_body,
|
||||
else_body,
|
||||
} => out.push(IrOp::If {
|
||||
then_body: replace_self_calls(then_body, self_id, cond, base),
|
||||
else_body: else_body.map(|eb| replace_self_calls(eb, self_id, cond, base)),
|
||||
}),
|
||||
other => out.push(other),
|
||||
}
|
||||
}
|
||||
out
|
||||
}
|
||||
|
||||
fn contains_call_to(ops: &[IrOp], target: WordId) -> bool {
|
||||
for op in ops {
|
||||
match op {
|
||||
@@ -746,8 +891,130 @@ mod tests {
|
||||
strength_reduce: true,
|
||||
dce: true,
|
||||
inline: false,
|
||||
self_guard: false,
|
||||
};
|
||||
optimize(ops, &config, &HashMap::new())
|
||||
optimize(ops, &config, &HashMap::new(), None)
|
||||
}
|
||||
|
||||
/// A body shaped like a recursive Forth word: a base-case guard, then the
|
||||
/// recursive step. `SELF` is the word being compiled.
|
||||
const SELF: WordId = WordId(9);
|
||||
|
||||
fn guarded_body(step: Vec<IrOp>) -> Vec<IrOp> {
|
||||
let mut ops = vec![
|
||||
IrOp::Dup,
|
||||
IrOp::PushI32(2),
|
||||
IrOp::Lt,
|
||||
IrOp::If {
|
||||
then_body: vec![IrOp::Exit],
|
||||
else_body: None,
|
||||
},
|
||||
];
|
||||
ops.extend(step);
|
||||
ops
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn self_guard_moves_the_base_case_to_the_call_site() {
|
||||
let out = expand_self_guard(guarded_body(vec![IrOp::Call(SELF)]), SELF);
|
||||
assert_eq!(
|
||||
out,
|
||||
guarded_body(vec![
|
||||
IrOp::Dup,
|
||||
IrOp::PushI32(2),
|
||||
IrOp::Lt,
|
||||
IrOp::If {
|
||||
then_body: vec![],
|
||||
else_body: Some(vec![IrOp::Call(SELF)]),
|
||||
},
|
||||
])
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn self_guard_carries_the_value_the_guard_returns() {
|
||||
// `: F DUP 2 < IF DROP 0 EXIT THEN RECURSE ;` -- the base case is not
|
||||
// "leave the argument", it is "replace it with 0".
|
||||
let body = vec![
|
||||
IrOp::Dup,
|
||||
IrOp::PushI32(2),
|
||||
IrOp::Lt,
|
||||
IrOp::If {
|
||||
then_body: vec![IrOp::Drop, IrOp::PushI32(0), IrOp::Exit],
|
||||
else_body: None,
|
||||
},
|
||||
IrOp::Call(SELF),
|
||||
];
|
||||
let out = expand_self_guard(body, SELF);
|
||||
let IrOp::If { then_body, .. } = &out[7] else {
|
||||
panic!("expected the expanded guard at index 7, got {:?}", out);
|
||||
};
|
||||
assert_eq!(then_body, &vec![IrOp::Drop, IrOp::PushI32(0)]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn self_guard_leaves_a_body_without_a_guard_alone() {
|
||||
// An `IF` with an `ELSE` is a branch, not an early return.
|
||||
let body = vec![
|
||||
IrOp::Dup,
|
||||
IrOp::If {
|
||||
then_body: vec![IrOp::Drop],
|
||||
else_body: Some(vec![IrOp::Call(SELF)]),
|
||||
},
|
||||
];
|
||||
assert_eq!(expand_self_guard(body.clone(), SELF), body);
|
||||
|
||||
// No `EXIT` in the then-branch: also not a guard.
|
||||
let body = guarded_body(vec![IrOp::Call(SELF)])
|
||||
.into_iter()
|
||||
.map(|op| match op {
|
||||
IrOp::If { .. } => IrOp::If {
|
||||
then_body: vec![IrOp::Drop],
|
||||
else_body: None,
|
||||
},
|
||||
other => other,
|
||||
})
|
||||
.collect::<Vec<_>>();
|
||||
assert_eq!(expand_self_guard(body.clone(), SELF), body);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn self_guard_refuses_a_condition_it_cannot_run_twice() {
|
||||
// A guard reached through a call or a memory write would be evaluated
|
||||
// once at the call site and again inside the callee.
|
||||
let body = vec![
|
||||
IrOp::Call(WordId(3)),
|
||||
IrOp::If {
|
||||
then_body: vec![IrOp::Exit],
|
||||
else_body: None,
|
||||
},
|
||||
IrOp::Call(SELF),
|
||||
];
|
||||
assert_eq!(expand_self_guard(body.clone(), SELF), body);
|
||||
|
||||
let body = vec![
|
||||
IrOp::Dup,
|
||||
IrOp::Fetch,
|
||||
IrOp::If {
|
||||
then_body: vec![IrOp::Exit],
|
||||
else_body: None,
|
||||
},
|
||||
IrOp::Call(SELF),
|
||||
];
|
||||
assert_eq!(expand_self_guard(body.clone(), SELF), body);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn self_guard_stops_at_the_call_site_budget() {
|
||||
let step = std::iter::repeat_n(IrOp::Call(SELF), MAX_GUARD_SITES + 1).collect();
|
||||
let body = guarded_body(step);
|
||||
assert_eq!(expand_self_guard(body.clone(), SELF), body);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn self_guard_leaves_tail_calls_alone() {
|
||||
let body = guarded_body(vec![IrOp::TailCall(SELF)]);
|
||||
assert_eq!(expand_self_guard(body.clone(), SELF), body);
|
||||
}
|
||||
|
||||
fn opt_with_inline(ops: Vec<IrOp>, bodies: &HashMap<WordId, Vec<IrOp>>) -> Vec<IrOp> {
|
||||
@@ -758,8 +1025,9 @@ mod tests {
|
||||
strength_reduce: true,
|
||||
dce: true,
|
||||
inline: true,
|
||||
self_guard: false,
|
||||
};
|
||||
optimize(ops, &config, bodies)
|
||||
optimize(ops, &config, bodies, None)
|
||||
}
|
||||
|
||||
// Peephole tests
|
||||
@@ -1019,8 +1287,9 @@ mod tests {
|
||||
strength_reduce: false,
|
||||
dce: false,
|
||||
inline: true,
|
||||
self_guard: false,
|
||||
};
|
||||
let result = optimize(vec![IrOp::Call(WordId(5))], &config, &bodies);
|
||||
let result = optimize(vec![IrOp::Call(WordId(5))], &config, &bodies, None);
|
||||
assert_eq!(result, vec![IrOp::Call(WordId(5))]);
|
||||
}
|
||||
|
||||
|
||||
@@ -2463,8 +2463,13 @@ impl<R: Runtime> ForthVM<R> {
|
||||
}
|
||||
|
||||
/// Run all enabled optimization passes on an IR sequence.
|
||||
fn optimize_ir(&self, ir: Vec<IrOp>, bodies: &HashMap<WordId, Vec<IrOp>>) -> Vec<IrOp> {
|
||||
optimize(ir, &self.config.opt, bodies)
|
||||
fn optimize_ir(
|
||||
&self,
|
||||
ir: Vec<IrOp>,
|
||||
bodies: &HashMap<WordId, Vec<IrOp>>,
|
||||
self_id: Option<WordId>,
|
||||
) -> Vec<IrOp> {
|
||||
optimize(ir, &self.config.opt, bodies, self_id)
|
||||
}
|
||||
|
||||
/// Parse a `{: args | locals -- comment :}` block and compile local
|
||||
@@ -2576,7 +2581,7 @@ impl<R: Runtime> ForthVM<R> {
|
||||
|
||||
let ir = std::mem::take(&mut self.compiling_ir);
|
||||
let bodies = self.ir_bodies.clone();
|
||||
let ir = self.optimize_ir(ir, &bodies);
|
||||
let ir = self.optimize_ir(ir, &bodies, Some(word_id));
|
||||
self.ir_bodies.insert(word_id, ir.clone());
|
||||
|
||||
// Compile to WASM
|
||||
@@ -2992,7 +2997,7 @@ impl<R: Runtime> ForthVM<R> {
|
||||
ir_body: Vec<IrOp>,
|
||||
) -> anyhow::Result<WordId> {
|
||||
let bodies = self.ir_bodies.clone();
|
||||
let ir_body = self.optimize_ir(ir_body, &bodies);
|
||||
let ir_body = self.optimize_ir(ir_body, &bodies, None);
|
||||
let word_id = self
|
||||
.dictionary
|
||||
.create(name, immediate)
|
||||
@@ -8290,6 +8295,36 @@ mod tests {
|
||||
assert_eq!(eval_output(": D 1 2 3 2 0 DO ROT LOOP . . . ; D"), "2 1 3 ");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_self_guard_expansion_keeps_the_answers() {
|
||||
// The base-case guard is tested at the call site, so a leaf never
|
||||
// costs a call. All four verified against gforth.
|
||||
assert_eq!(
|
||||
eval_output(
|
||||
": FIB DUP 2 < IF EXIT THEN DUP 1- RECURSE SWAP 2 - RECURSE + ; \
|
||||
25 FIB . 0 FIB . 1 FIB . 2 FIB . 10 FIB ."
|
||||
),
|
||||
"75025 0 1 1 55 "
|
||||
);
|
||||
// A guard that replaces its argument rather than leaving it.
|
||||
assert_eq!(
|
||||
eval_output(": G DUP 0= IF DROP 0 EXIT THEN DUP 1- RECURSE + ; 5 G . 0 G . 100 G ."),
|
||||
"15 0 5050 "
|
||||
);
|
||||
assert_eq!(
|
||||
eval_output(": H DUP 3 < IF DROP 7 EXIT THEN 1- RECURSE 2 * ; 5 H . 2 H . 8 H ."),
|
||||
"56 7 448 "
|
||||
);
|
||||
// Two guards and three call sites, one of them behind the second guard.
|
||||
assert_eq!(
|
||||
eval_output(
|
||||
": ACK OVER 0= IF SWAP DROP 1+ EXIT THEN DUP 0= IF DROP 1- 1 RECURSE EXIT THEN \
|
||||
OVER SWAP 1- RECURSE SWAP 1- SWAP RECURSE ; 2 3 ACK . 1 2 ACK ."
|
||||
),
|
||||
"9 4 "
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_promoted_begin_loops() {
|
||||
// BEGIN loops promote too, so these run entirely in locals.
|
||||
|
||||
@@ -746,7 +746,7 @@ fn perf_benchmarks() -> Vec<PerfBenchmark> {
|
||||
verify: "25 FIB",
|
||||
expected: 75025,
|
||||
samples: 5,
|
||||
max_ratio: 0.17,
|
||||
max_ratio: 0.10,
|
||||
},
|
||||
PerfBenchmark {
|
||||
name: "Factorial(12)x100K",
|
||||
|
||||
+37
-10
@@ -30,6 +30,7 @@ This document describes every optimization that makes sense for WAFER, why it ma
|
||||
| 14 | Self-Recursive Direct Call | Codegen | Done | High |
|
||||
| 15 | Float / Double-Cell | Codegen | Not started | Future |
|
||||
| 16 | Typed Calling Convention | Codegen | Done | Highest |
|
||||
| 17 | Self-Guard Expansion | IR pass | Done | Medium |
|
||||
|
||||
## 1. Stack-to-Local Promotion
|
||||
|
||||
@@ -484,32 +485,58 @@ Fibonacci(25) went from 1035 to 366 microseconds, 4.3x slower than `sf64` to 1.2
|
||||
|
||||
Untyped by design: anything using `SP@`, `DEPTH`, `EXECUTE`, `>R`/`R>`, floats or locals; anything calling a word that is itself untyped, which in the JIT path means every call except `RECURSE`; mutually recursive words; and words whose effect is not static -- branches that disagree on depth, `EXIT` at the wrong depth, a non-neutral loop body, or a recursion that grows the stack per level.
|
||||
|
||||
## 17. Self-Guard Expansion
|
||||
|
||||
**Status: Done.** A recursive word's base-case guard is duplicated into its own call sites, so the leaves of the recursion cost a test instead of a call. Implemented in `optimizer.rs::expand_self_guard`, gated on `OptConfig::self_guard`, and applied after inlining so the later passes still run over the result.
|
||||
|
||||
### The Shape
|
||||
|
||||
A recursive Forth word almost always opens with a guard that returns early:
|
||||
|
||||
```forth
|
||||
: FIB DUP 2 < IF EXIT THEN DUP 1- RECURSE SWAP 2 - RECURSE + ;
|
||||
```
|
||||
|
||||
Every leaf of the recursion is then a call whose entire body is `DUP 2 <`. The pass rewrites each `Call(self)` as
|
||||
|
||||
```forth
|
||||
DUP 2 < IF ( leave it ) ELSE RECURSE THEN
|
||||
```
|
||||
|
||||
which computes the same thing -- the callee would have run the guard, taken the branch and returned. When the guard returns a value rather than its argument (`IF DROP 0 EXIT THEN`), that value moves into the then-branch with it.
|
||||
|
||||
### Why It Is Bounded
|
||||
|
||||
The guard runs twice along the recursive path: once at the call site, once inside the callee. So it must be small and free of effects -- `MAX_GUARD_OPS` is six, and the operations are restricted to stack shuffles, arithmetic and comparisons; a call, a memory access or a branch disqualifies it. `MAX_GUARD_SITES` caps the expansion at four call sites, since each one replicates the guard. A `TailCall` is never expanded, which keeps this pass and tail-call detection from having to agree about what tail position means.
|
||||
|
||||
### Impact
|
||||
|
||||
Fibonacci(25): 356 to 237 microseconds. In fib's tree half of all nodes are leaves, which is where the factor comes from. This is what took the last benchmark past `sf64`.
|
||||
|
||||
## Current Performance vs Gforth
|
||||
|
||||
All optimizations enabled, release mode, measured with UTIME:
|
||||
|
||||
```
|
||||
Benchmark WAFER CONSOL gforth sf64 WAFER/gf WAFER/sf
|
||||
Fibonacci(25) 356 361 3389 287 0.11x 1.24x
|
||||
Factorial(12)x100K 479 495 6249 1650 0.08x 0.29x
|
||||
GCD-bench(20K) 540 559 1801 801 0.30x 0.67x
|
||||
NestedLoops(50)x1K 509 501 7023 1887 0.07x 0.27x
|
||||
Collatz(2K) 185 213 3873 610 0.05x 0.30x
|
||||
Fibonacci(25) 237 242 3340 287 0.07x 0.83x
|
||||
Factorial(12)x100K 480 479 6109 1594 0.08x 0.30x
|
||||
GCD-bench(20K) 549 541 1830 797 0.30x 0.68x
|
||||
NestedLoops(50)x1K 501 509 7092 1898 0.07x 0.26x
|
||||
Collatz(2K) 196 190 3955 633 0.05x 0.30x
|
||||
```
|
||||
|
||||
Times in microseconds. WAFER/gf < 1.0 means WAFER is faster. `sf64` is SwiftForth,
|
||||
which compiles to native code; two caveats on that column. The install here is an
|
||||
x86-64 binary under Rosetta 2 while WAFER and gforth are native arm64, so it is a
|
||||
native-vs-emulated comparison and a native SwiftForth would be faster than these
|
||||
numbers; and sf64 uses 64-bit cells to WAFER's 32-bit. WAFER is ahead on the four
|
||||
loop-heavy benchmarks and behind on Fibonacci, which is one call per node with no
|
||||
loop to promote.
|
||||
numbers; and sf64 uses 64-bit cells to WAFER's 32-bit.
|
||||
|
||||
## Remaining Opportunities
|
||||
|
||||
| Optimization | Status | Potential Impact |
|
||||
| -------------------------------- | ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Bounded self-inlining | Not started | Measured 1.33x on Fibonacci, the last benchmark behind sf64. Blocked on `EXIT`: the inliner refuses any body containing one, and a recursive Forth word is `... IF EXIT THEN ... RECURSE`. Needs either a scoped exit (compile an inlined `EXIT` as a branch to the end of a block) or guard-only expansion |
|
||||
| -------------------------------- | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Scoped exit for the inliner | Not started | The inliner still refuses any body containing an `EXIT`, because an inlined one would return from the caller. Compiling it as a branch to the end of a block would unlock inlining for every word with an early return, not just the guard shape section 17 handles |
|
||||
| BeginDoubleWhileRepeat promotion | Not started | Rare pattern, low priority. Its promoted emitter exists but has no loop fixup and is unverified |
|
||||
| LEAVE as IR primitive | Not started | Would enable fast-path for loops with LEAVE |
|
||||
| Float stack-to-local | Not started | Eliminate float stack memory traffic |
|
||||
|
||||
Reference in New Issue
Block a user