diff --git a/CHANGELOG.md b/CHANGELOG.md index 1926b5e..d56425c 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -27,8 +27,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 convention cost the most. It now handles both. Fibonacci(25) goes from 1035 to 366 µs, 4.3x slower than `sf64` to 1.2x. - Loop-heavy benchmarks are unchanged (they were already promoted, and - already beat `sf64`). Words that keep the memory convention: anything + Loop-heavy benchmarks are unchanged by this entry — see the region + promotion below for those. Words that keep the memory convention: anything using `SP@`, `DEPTH`, `EXECUTE`, `>R`/`R>`, floats or locals; anything calling a word that is itself untyped, which in the JIT path means every call except `RECURSE`; mutually recursive words; and words whose effect is @@ -46,12 +46,93 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 `WAFER_TYPED_CALLS=0` falls back to the memory-stack convention. +- **Promotion is now per region, not per word.** Stack-to-local promotion + used to be all-or-nothing: a single `.`, `CR`, `>R` or host call + anywhere in a definition put the _entire_ body on the memory data + stack, hot loops included. The stack simulator now runs over each + stretch of a word that can live in WASM locals, loading what the + region reads and writing back what it leaves, with the rest of the + word unchanged around it. + + The cliff this removes was steep. The same loop, same build: + + | `: L1 0 5000000 0 DO 1+ LOOP DROP ;` reached as | µs | ns/iter | + | ----------------------------------------------- | ----- | ------- | + | its own word | 1571 | 0.31 | + | inlined into a caller with a `.` in it (before) | 11100 | 2.22 | + | the same, after this change | 1572 | 0.31 | + + 7x, for one `i32.add`: on the memory path the accumulator is stored to + linear memory and reloaded next iteration, so the loop-carried + dependency runs through store-to-load forwarding instead of a + register. + + A region may only use `I` / `J` when the DO loops naming them are + inside the region, since the simulator resolves them against its own + loop stack. Straight-line regions have to be at least three operations + to be worth the load and store either side; a loop always is. + +- **The inliner no longer drags a loop onto the memory stack.** It + inlined any callee of eight IR operations or fewer, so a small + loop-bearing word inlined into a caller that can never be promoted + lost its registers -- an optimisation pass applying the 7x + pessimisation above. Loop-bearing callees now stay put in that case: + one call is far cheaper than a loop's worth of memory traffic. + Straight-line words still inline everywhere. + +- **`BEGIN` loops promote as well.** `BEGIN..UNTIL`, `BEGIN..AGAIN` and + `BEGIN..WHILE..REPEAT` were rejected outright by the eligibility check, + so any word built on the idiomatic Forth loop kept the memory data + stack no matter how hot it was. They are promoted now when the + construct is stack-neutral: `UNTIL` consumes exactly the flag its body + leaves, `AGAIN`'s body is neutral, and for `WHILE..REPEAT` the test and + the body balance separately -- `WHILE` leaves the loop between the two, + so a net that only added up over the pair would give the two exits + different stack shapes. Bodies containing an `EXIT` stay out, the same + rule `DO`/`LOOP` follows. `BEGIN..WHILE..WHILE..REPEAT` is still + excluded. + + GCD 994 -> 540 µs, Collatz 428 -> 185. + + Together these four entries put four of the five cross-engine + benchmarks past SwiftForth `sf64`: Factorial 0.29x, Collatz 0.30x, + NestedLoops 0.27x, GCD 0.67x. Fibonacci stays at 1.24x, being pure + call overhead with no loop to promote. + ### Fixed +- **A promoted loop or `IF` whose branch permutes the stack lost a value.** + At the bottom of a promoted loop the body's results are copied back into + the loop-top locals, and the join after a promoted `IF` copies one + branch's locals into the other's. Both did it one slot at a time in index + order, which is wrong as soon as a destination is also a later source: + `: C 3 4 2 0 DO SWAP LOOP . . ;` printed `4 4` where gforth and + SwiftForth print `4 3`, and `2 0 DO ROT LOOP` over three cells printed + `3 2 3` instead of `2 1 3`. The copies are now ordered so every source is + read before it is overwritten, with one scratch local to break a cycle. + Present since stack-to-local promotion was introduced; reachable from + any `DO` loop or `IF` whose body reorders cells it did not create. + - The Forth 2012 Core suite now also runs against consolidated code (`compliance_core_after_consolidate`). `CONSOLIDATE` had no correctness test at all before -- only benchmarks. +### Changed + +- **Three cross-engine benchmarks were too small to be measured.** GCD ran + in 14 µs, Factorial in 49 and NestedLoops in 51, where per-run scatter is + a good fraction of the total and fixed per-invocation costs in the other + engines dominate. Scaled to Factorial x100K, GCD-bench(20K) and + NestedLoops(50)x1K, all now around 0.5-1 ms. + + This changed a result rather than just steadying it: GCD looked like a + win at 0.42x of `sf64` and was in fact a loss at 1.17x. That is what + pointed at `BEGIN` loops as the remaining gap -- GCD is the one benchmark + whose loop is a `BEGIN ... WHILE ... REPEAT` -- and with those promoted it + now reads 0.67x. The regression limits, which had drifted to 3-6x looser + than the measurements they guard, were retightened to ~45% above the + current ratios. + ## [0.2.6] - 2026-08-07 ### Fixed diff --git a/README.md b/README.md index 4ba3c59..0eb62be 100644 --- a/README.md +++ b/README.md @@ -7,7 +7,7 @@ An optimizing Forth 2012 compiler targeting WebAssembly. WAFER JIT-compiles each ## Highlights - **200+ words** across 12 Forth 2012 word sets, all at **100% compliance** -- **Optimizing compiler** with 6 IR passes + stack-to-local promotion (loops + IF) + consolidation +- **Optimizing compiler** with 6 IR passes + stack-to-local promotion (per region, so a hot loop keeps its registers even inside a word that does I/O; `DO` and `BEGIN` loops alike) + consolidation - **Faster than gforth** on all benchmarks in release mode (2-10x faster) - **JIT compilation** — each `:` definition compiles to its own WASM module - **Self-recursive direct calls** — RECURSE compiles to native `call` instead of `call_indirect` @@ -85,15 +85,20 @@ reach of SwiftForth `sf64`, which compiles to native code: ``` Benchmark WAFER CONSOL gforth sf64 WAFER/gf WAFER/sf -Fibonacci(25) 378 359 3238 296 0.11x 1.21x -Factorial(12)x10K 335 320 633 183 0.51x 1.75x -GCD-bench(500) 14 15 29 31 0.48x 0.45x -NestedLoops(50) 72 69 698 207 0.10x 0.33x -Collatz(2K) 994 997 3940 668 0.25x 1.49x +Fibonacci(25) 356 361 3389 287 0.11x 1.24x +Factorial(12)x100K 479 495 6249 1650 0.08x 0.29x +GCD-bench(20K) 540 559 1801 801 0.30x 0.67x +NestedLoops(50)x1K 509 501 7023 1887 0.07x 0.27x +Collatz(2K) 185 213 3873 610 0.05x 0.30x ``` Times in microseconds. WAFER/gf < 1.0 means WAFER is faster. CONSOL = after `CONSOLIDATE`. +Two caveats on the `sf64` column. The SwiftForth build here is x86-64 running under Rosetta 2 +while WAFER and gforth are native arm64, so it is a native-vs-emulated comparison; and sf64 +uses 64-bit cells to WAFER's 32-bit. WAFER is ahead on the four loop-heavy benchmarks and +behind on Fibonacci, which is one call per node with no loop to promote. + A word whose stack effect is statically known gets a **typed entry point**: its stack items travel in and out as WASM values instead of through the memory data stack, so cranelift keeps them in registers across a call the way a native Forth keeps TOS in one. The word also keeps a `( -- )` wrapper, which is what the function diff --git a/crates/core/src/codegen.rs b/crates/core/src/codegen.rs index e2ef3fa..062f65c 100644 --- a/crates/core/src/codegen.rs +++ b/crates/core/src/codegen.rs @@ -354,6 +354,9 @@ struct EmitCtx { /// Stack of open block labels for flat forward branches (CS-ROLL'd IF/THEN). /// Used by `BranchIfFalse` to compute `br_if` depth. open_blocks: Vec, + /// First WASM local a promoted region may allocate from. Regions run one + /// after another, so they all share this pool. + region_local_base: u32, } /// Decrement the FSP global by 8 (allocate space for one f64). @@ -465,8 +468,148 @@ fn emit_float_cmp(f: &mut Function, ctx: &EmitCtx, wasm_cmp: &Instruction<'_>) { /// Emit all IR operations in `ops` into the WASM function body `f`. fn emit_body(f: &mut Function, ops: &[IrOp], ctx: &mut EmitCtx) { + let mut i = 0; + while i < ops.len() { + let run = promotable_run(&ops[i..]); + if run > 0 && region_is_worth_promoting(&ops[i..i + run]) { + emit_promoted_region(f, &ops[i..i + run], ctx); + } else { + let run = run.max(1); + for op in &ops[i..i + run] { + emit_op(f, op, ctx); + } + } + i += run.max(1); + } +} + +/// Run the stack simulator over one stretch of a word that is otherwise on +/// the memory path: load what the region reads into WASM locals, work there, +/// write the results back. +/// +/// This is what keeps a hot loop in registers inside a word that can never be +/// promoted as a whole -- one `.` or one host call used to put the entire +/// body, loops included, back on the memory data stack. +fn emit_promoted_region(f: &mut Function, ops: &[IrOp], ctx: &mut EmitCtx) { + let (preload, _) = compute_stack_needs(ops); + let mut sim = StackSim::new(ctx.region_local_base).with_move_scratch(); + emit_promoted_prologue(f, preload, &mut sim); for op in ops { - emit_op(f, op, ctx); + emit_promoted_op(f, op, &mut sim); + } + emit_promoted_epilogue(f, &mut sim); +} + +/// Length of the longest prefix of `ops` that can run as a promoted region. +fn promotable_run(ops: &[IrOp]) -> usize { + ops.iter() + .take_while(|op| { + let one = std::slice::from_ref(*op); + is_promotable_body(one, PromoteMode::Memory) && region_loop_refs_resolved(one, 0) + }) + .count() +} + +/// Is a region worth the load/store either side of it? +/// +/// A loop always is -- that is the whole point. Otherwise the prologue and +/// epilogue have to be amortised over enough operations to beat leaving them +/// on the memory stack, which costs roughly two or three accesses each. +fn region_is_worth_promoting(ops: &[IrOp]) -> bool { + ops.len() >= MIN_PROMOTED_REGION || ops.iter().any(is_loop_op) +} + +/// Smallest straight-line region worth promoting. +const MIN_PROMOTED_REGION: usize = 3; + +fn is_loop_op(op: &IrOp) -> bool { + matches!( + op, + IrOp::DoLoop { .. } + | IrOp::BeginUntil { .. } + | IrOp::BeginAgain { .. } + | IrOp::BeginWhileRepeat { .. } + | IrOp::BeginDoubleWhileRepeat { .. } + ) +} + +/// Does every `I` / `J` in `ops` refer to a DO loop that is inside `ops`? +/// +/// The promoted emitter resolves them against its own loop stack, so a region +/// that borrows the index of a loop emitted around it would read the wrong +/// local -- or, for `J` below two levels, silently emit nothing. +fn region_loop_refs_resolved(ops: &[IrOp], depth: u32) -> bool { + ops.iter().all(|op| match op { + IrOp::RFetch => depth >= 1, + IrOp::LoopJ => depth >= 2, + IrOp::DoLoop { body, .. } => region_loop_refs_resolved(body, depth + 1), + IrOp::If { + then_body, + else_body, + } => { + region_loop_refs_resolved(then_body, depth) + && else_body + .as_deref() + .is_none_or(|eb| region_loop_refs_resolved(eb, depth)) + } + IrOp::BeginUntil { body } | IrOp::BeginAgain { body } => { + region_loop_refs_resolved(body, depth) + } + IrOp::BeginWhileRepeat { test, body } => { + region_loop_refs_resolved(test, depth) && region_loop_refs_resolved(body, depth) + } + _ => true, + }) +} + +/// Locals needed by the largest single promoted region in `ops`, walking the +/// body exactly the way [`emit_body`] partitions it. +fn region_local_budget(ops: &[IrOp]) -> u32 { + let mut max = 0; + let mut i = 0; + while i < ops.len() { + let run = promotable_run(&ops[i..]); + if run > 0 && region_is_worth_promoting(&ops[i..i + run]) { + let region = &ops[i..i + run]; + let (preload, _) = compute_stack_needs(region); + max = max.max(count_promoted_locals(region, preload)); + } else { + for op in &ops[i..i + run.max(1)] { + max = max.max(region_local_budget_of_children(op)); + } + } + i += run.max(1); + } + max +} + +/// Largest region budget among an operation's nested bodies. +fn region_local_budget_of_children(op: &IrOp) -> u32 { + match op { + IrOp::If { + then_body, + else_body, + } => { + region_local_budget(then_body).max(else_body.as_deref().map_or(0, region_local_budget)) + } + IrOp::DoLoop { body, .. } | IrOp::BeginUntil { body } | IrOp::BeginAgain { body } => { + region_local_budget(body) + } + IrOp::BeginWhileRepeat { test, body } => { + region_local_budget(test).max(region_local_budget(body)) + } + IrOp::BeginDoubleWhileRepeat { + outer_test, + inner_test, + body, + after_repeat, + else_body, + } => region_local_budget(outer_test) + .max(region_local_budget(inner_test)) + .max(region_local_budget(body)) + .max(region_local_budget(after_repeat)) + .max(else_body.as_deref().map_or(0, region_local_budget)), + _ => 0, } } @@ -1281,6 +1424,30 @@ enum PromoteMode { Typed, } +/// Can this body be promoted once its calls are accounted for? +/// +/// True when the only things standing between it and the register path are +/// calls and `EXIT` -- i.e. the word either gets a typed entry or, failing +/// that, promotes region by region. False means it is stuck on the memory +/// data stack whatever happens, which is what the inliner needs to know. +pub(crate) fn promotable_modulo_calls(ops: &[IrOp]) -> bool { + is_promotable_body(ops, PromoteMode::Typed) +} + +/// Does this body contain a loop at any nesting depth? +pub(crate) fn contains_loop(ops: &[IrOp]) -> bool { + ops.iter().any(|op| { + is_loop_op(op) + || match op { + IrOp::If { + then_body, + else_body, + } => contains_loop(then_body) || else_body.as_deref().is_some_and(contains_loop), + _ => false, + } + }) +} + /// Recursive check for promotable ops. fn is_promotable_body(ops: &[IrOp], mode: PromoteMode) -> bool { let typed = mode == PromoteMode::Typed; @@ -1367,11 +1534,46 @@ fn is_promotable_body(ops: &[IrOp], mode: PromoteMode) -> bool { } } } - // BEGIN loops, BeginDoubleWhileRepeat, flat forward blocks: not promoted - IrOp::BeginUntil { .. } - | IrOp::BeginAgain { .. } - | IrOp::BeginWhileRepeat { .. } - | IrOp::BeginDoubleWhileRepeat { .. } + // BEGIN loops: the construct as a whole is stack-neutral, which is + // what lets the next iteration reuse the loop-top locals. A body + // that is not neutral has no single promoted stack shape, and an + // EXIT out of one would have to unwind the join -- the same rule + // DO/LOOP already follows. + IrOp::BeginUntil { body } | IrOp::BeginAgain { body } => { + if !is_promotable_body(body, mode) || body_has_exit(body) { + return false; + } + if !typed { + // UNTIL consumes a flag the body leaves, AGAIN consumes nothing. + let expected = i32::from(matches!(op, IrOp::BeginUntil { .. })); + let (_, body_net) = compute_stack_needs(body); + if body_net != expected { + return false; + } + } + } + IrOp::BeginWhileRepeat { test, body } => { + if !is_promotable_body(test, mode) + || !is_promotable_body(body, mode) + || body_has_exit(test) + || body_has_exit(body) + { + return false; + } + if !typed { + // WHILE leaves the loop between test and body, so the two + // have to be neutral separately: a net that only balances + // over the pair would give the two exits different shapes. + let (_, test_net) = compute_stack_needs(test); + let (_, body_net) = compute_stack_needs(body); + if test_net != 1 || body_net != 0 { + return false; + } + } + } + // BeginDoubleWhileRepeat has a promoted emitter, but one without a + // loop fixup and never exercised; flat forward blocks have none. + IrOp::BeginDoubleWhileRepeat { .. } | IrOp::Block(_) | IrOp::BranchIfFalse(_) | IrOp::EndBlock(_) @@ -1759,21 +1961,34 @@ fn compute_stack_needs_rec(ops: &[IrOp], st: &mut Needs<'_>) { IrOp::BeginUntil { body } => { let saved = st.depth; compute_stack_needs_rec(body, st); - // Body produces flag, consumed by UNTIL: net 0 for the whole construct + // Body produces the flag UNTIL consumes: net 0 for the whole + // construct, and anything else has no promoted stack shape. + if st.depth != saved + 1 { + st.consistent = false; + } st.depth = saved; } IrOp::BeginAgain { body } => { let saved = st.depth; compute_stack_needs_rec(body, st); + if st.depth != saved { + st.consistent = false; + } st.depth = saved; } IrOp::BeginWhileRepeat { test, body } => { let saved = st.depth; compute_stack_needs_rec(test, st); - // WHILE consumes flag + // WHILE consumes the flag, and leaves the loop right here, so + // test and body have to balance separately rather than as a pair. + if st.depth != saved + 1 { + st.consistent = false; + } st.depth -= 1; compute_stack_needs_rec(body, st); - // Whole construct is stack-neutral + if st.depth != saved { + st.consistent = false; + } st.depth = saved; } IrOp::BeginDoubleWhileRepeat { @@ -1821,7 +2036,8 @@ fn compute_stack_needs_rec(ops: &[IrOp], st: &mut Needs<'_>) { /// DSP and scratch locals). This is an upper bound -- we allocate a fresh /// local for each value-producing operation. fn count_promoted_locals(ops: &[IrOp], preload: u32) -> u32 { - let mut count = preload; + // +1 for the simulator's cycle-breaking scratch (`with_move_scratch`). + let mut count = preload + 1; count_promoted_locals_body(ops, &mut count); count } @@ -1913,6 +2129,10 @@ struct StackSim { /// True once the code emitted so far cannot fall through (an `EXIT` ran). /// The join after an `IF` uses it to take the surviving branch's state. diverged: bool, + /// Spare local reserved for breaking a cycle in `emit_parallel_move`. + /// Reserved before any value local so that the `IF` join, which rewinds + /// `next_local` for the else arm, can never hand it out twice. + move_scratch: Option, } /// What the typed emitter needs beyond the simulator itself. @@ -1932,13 +2152,22 @@ impl StackSim { loop_index_stack: Vec::new(), typed: None, diverged: false, + move_scratch: None, } } + /// Reserve the cycle-breaking local. Every simulator that emits promoted + /// operations needs this; the typed wrapper, which only shuffles the + /// memory stack, does not. `count_promoted_locals` budgets for it. + fn with_move_scratch(mut self) -> Self { + self.move_scratch = Some(self.alloc()); + self + } + /// Simulator for a typed fast entry: params occupy locals `0..params`, /// so fresh locals start above them. fn new_typed(params: u32, results: u32, callees: &Rc>) -> Self { - let mut sim = Self::new(params); + let mut sim = Self::new(params).with_move_scratch(); sim.stack = (0..params).collect(); sim.typed = Some(TypedCtx { results, @@ -2313,14 +2542,10 @@ fn emit_promoted_op(f: &mut Function, op: &IrOp, sim: &mut StackSim) { // join state is the else state, already in sim.stack } else { if !else_diverged { - let else_stack = &sim.stack; - let min_len = then_stack.len().min(else_stack.len()); - for i in 0..min_len { - if then_stack[i] != else_stack[i] { - f.instruction(&Instruction::LocalGet(else_stack[i])); - f.instruction(&Instruction::LocalSet(then_stack[i])); - } - } + let min_len = then_stack.len().min(sim.stack.len()); + let dsts = then_stack[..min_len].to_vec(); + let srcs = sim.stack[..min_len].to_vec(); + emit_parallel_move(f, sim, &dsts, &srcs); } sim.stack = then_stack; } @@ -2438,6 +2663,12 @@ fn emit_promoted_op(f: &mut Function, op: &IrOp, sim: &mut StackSim) { let cond = sim.pop(); f.instruction(&Instruction::LocalGet(cond)); f.instruction(&Instruction::I32Eqz); + // WHILE leaves the loop here, so the loop-top locals have to hold + // the right values on the way out too -- a test that permutes + // (`BEGIN SWAP DUP WHILE`) would otherwise leave them crossed. + // The flag is already on the operand stack, so moving locals + // between it and the `br_if` is safe. + emit_promoted_loop_fixup(f, sim, &loop_top_stack); f.instruction(&Instruction::BrIf(1)); // break to outer block emit_promoted_body(f, body, sim); @@ -2623,16 +2854,69 @@ fn emit_promoted_loop_fixup(f: &mut Function, sim: &mut StackSim, loop_top_stack sim.stack.len(), loop_top_stack.len() ); - for (i, &top_local) in loop_top_stack.iter().enumerate() { - if sim.stack[i] != top_local { - f.instruction(&Instruction::LocalGet(sim.stack[i])); - f.instruction(&Instruction::LocalSet(top_local)); - } - } + let srcs = sim.stack.clone(); + emit_parallel_move(f, sim, loop_top_stack, &srcs); // Reset sim to loop-top state sim.stack = loop_top_stack.to_vec(); } +/// Emit `dsts[i] := srcs[i]` for every `i`, all at once. +/// +/// Copying them in index order is wrong as soon as a destination is also a +/// later source: `BEGIN ... SWAP ... UNTIL` would write the top into the +/// second slot and then read that slot back, so both end up holding the same +/// value. The moves are ordered so that every source is read before it is +/// overwritten, and a cycle -- which has no such order -- is broken by +/// stashing one source in `sim.move_scratch`. +/// +/// One scratch local is enough for any number of cycles: the loop only breaks +/// a new cycle once nothing else can be emitted, and by then the previous +/// cycle has drained and released it. +fn emit_parallel_move(f: &mut Function, sim: &mut StackSim, dsts: &[u32], srcs: &[u32]) { + let mut pending: Vec<(u32, u32)> = dsts + .iter() + .zip(srcs) + .filter(|(d, s)| d != s) + .map(|(d, s)| (*d, *s)) + .collect(); + + while !pending.is_empty() { + let before = pending.len(); + let mut i = 0; + while i < pending.len() { + let (dst, src) = pending[i]; + // Safe to write `dst` now only if nothing still has to read it. + if pending + .iter() + .enumerate() + .all(|(j, (_, s))| j == i || *s != dst) + { + f.instruction(&Instruction::LocalGet(src)); + f.instruction(&Instruction::LocalSet(dst)); + pending.remove(i); + } else { + i += 1; + } + } + if pending.len() == before { + // Everything left is a cycle. Lift one source out of it, which + // frees its local and turns the cycle into a chain. + let (dst, src) = pending.remove(0); + let tmp = sim + .move_scratch + .expect("promoted simulator without a move scratch local"); + f.instruction(&Instruction::LocalGet(src)); + f.instruction(&Instruction::LocalSet(tmp)); + for p in &mut pending { + if p.1 == src { + p.1 = tmp; + } + } + pending.push((dst, tmp)); + } + } +} + /// Emit a promoted binary operation (commutative). fn emit_promoted_binary(f: &mut Function, sim: &mut StackSim, op: &Instruction<'_>) { let b = sim.pop(); @@ -3114,13 +3398,20 @@ pub fn compile_word( let forth_local_count = count_forth_locals(body); let loop_depth = count_loop_depth(body); let loop_local_count = loop_depth * 2; // 2 locals per nesting level (index, limit) + // Words on the memory path still promote what they can, region by + // region, so they need a pool of locals for that on top of everything else. + let region_locals = if promoted { + 0 + } else { + region_local_budget(body) + }; let num_locals = if promoted { let (preload, _) = compute_stack_needs(body); let promoted_count = count_promoted_locals(body, preload); // 1 (cached DSP) + promoted locals (scratch locals not needed in promoted path) 1 + promoted_count + forth_local_count + loop_local_count } else { - 1 + scratch_count + forth_local_count + loop_local_count + 1 + scratch_count + forth_local_count + loop_local_count + region_locals }; let forth_f_local_count = count_forth_f_locals(body); // F: locals need f64 storage, which also implies the f64 scratch pair. @@ -3143,6 +3434,7 @@ pub fn compile_word( 1 + scratch_count }; let loop_local_base = forth_local_base + forth_local_count; + let region_local_base = loop_local_base + loop_local_count; // f64 scratch pair first (indices num_locals, num_locals+1), then F: locals. let forth_f_local_base = num_locals + 2; let mut ctx = EmitCtx { @@ -3155,6 +3447,7 @@ pub fn compile_word( fast_loop_depth: 0, self_word_id: Some(WordId(config.base_fn_index)), open_blocks: Vec::new(), + region_local_base, }; // Prologue: cache $dsp global into local 0 @@ -3164,7 +3457,7 @@ pub fn compile_word( if promoted { let (preload, _) = compute_stack_needs(body); let first_promoted = SCRATCH_BASE; // promoted locals start right after cached_dsp - let mut sim = StackSim::new(first_promoted); + let mut sim = StackSim::new(first_promoted).with_move_scratch(); emit_promoted_prologue(&mut func, preload, &mut sim); for op in body { emit_promoted_op(&mut func, op, &mut sim); @@ -3715,12 +4008,17 @@ fn compile_multi_word_module( let forth_local_count = count_forth_locals(body); let loop_depth = count_loop_depth(body); let loop_local_count = loop_depth * 2; + let region_locals = if promoted { + 0 + } else { + region_local_budget(body) + }; let num_locals = if promoted { let (preload, _) = compute_stack_needs(body); let promoted_count = count_promoted_locals(body, preload); 1 + promoted_count + forth_local_count + loop_local_count } else { - 1 + scratch_count + forth_local_count + loop_local_count + 1 + scratch_count + forth_local_count + loop_local_count + region_locals }; let forth_f_local_count = count_forth_f_locals(body); let has_floats = needs_f64_locals(body) || forth_f_local_count > 0; @@ -3742,6 +4040,7 @@ fn compile_multi_word_module( 1 + scratch_count }; let loop_local_base = forth_local_base + forth_local_count; + let region_local_base = loop_local_base + loop_local_count; let forth_f_local_base = num_locals + 2; let mut ctx = EmitCtx { f64_local_0: num_locals, @@ -3753,6 +4052,7 @@ fn compile_multi_word_module( fast_loop_depth: 0, self_word_id: None, // consolidated module uses direct calls via local_fn_map open_blocks: Vec::new(), + region_local_base, }; // Prologue: cache $dsp global into local 0 @@ -3763,7 +4063,7 @@ fn compile_multi_word_module( // Use stack-to-local promotion (same as compile_word path) let (preload, _) = compute_stack_needs(body); let first_promoted = SCRATCH_BASE; - let mut sim = StackSim::new(first_promoted); + let mut sim = StackSim::new(first_promoted).with_move_scratch(); emit_promoted_prologue(&mut func, preload, &mut sim); for op in body { emit_promoted_op(&mut func, op, &mut sim); diff --git a/crates/core/src/optimizer.rs b/crates/core/src/optimizer.rs index 291d420..e4c1da9 100644 --- a/crates/core/src/optimizer.rs +++ b/crates/core/src/optimizer.rs @@ -53,7 +53,12 @@ pub fn optimize( // Phase 2: inline then simplify again if config.inline { - ir = inline(ir, bodies, 8); + // A caller that can never leave the memory data stack would drag an + // inlined loop down with it, so leave those callees where they are: + // as their own word the loop keeps its registers, and one call is far + // cheaper than a loop's worth of memory traffic. + let keep_loops_out = !crate::codegen::promotable_modulo_calls(&ir); + ir = inline(ir, bodies, 8, keep_loops_out); } if config.peephole { ir = peephole(ir); @@ -496,7 +501,12 @@ fn dce(ops: Vec) -> Vec { /// Inline small word bodies: replaces `Call(id)` with the word's IR body /// if the body is small enough and not recursive. -fn inline(ops: Vec, bodies: &HashMap>, max_size: usize) -> Vec { +fn inline( + ops: Vec, + bodies: &HashMap>, + max_size: usize, + keep_loops_out: bool, +) -> Vec { let mut out = Vec::new(); for op in ops { match &op { @@ -505,6 +515,7 @@ fn inline(ops: Vec, bodies: &HashMap>, max_size: usize) && body.len() <= max_size && !contains_call_to(body, *id) && !contains_exit(body) + && !(keep_loops_out && crate::codegen::contains_loop(body)) { // Inline the body, recursively converting TailCall back to Call // (tail position in the callee is not tail position in the caller). @@ -517,7 +528,7 @@ fn inline(ops: Vec, bodies: &HashMap>, max_size: usize) } _ => { out.push(apply_to_bodies(op, &|inner| { - inline(inner, bodies, max_size) + inline(inner, bodies, max_size, keep_loops_out) })); } } @@ -1012,4 +1023,54 @@ mod tests { let result = optimize(vec![IrOp::Call(WordId(5))], &config, &bodies); assert_eq!(result, vec![IrOp::Call(WordId(5))]); } + + #[test] + fn keeps_a_loop_out_of_a_caller_stuck_on_the_memory_stack() { + // The caller has a `.`, so it can never leave the memory data stack. + // Inlining the loop would drag it down too; as its own word the loop + // keeps its registers and the caller just pays one call. + let mut bodies = HashMap::new(); + bodies.insert( + WordId(5), + vec![IrOp::DoLoop { + body: vec![IrOp::PushI32(1), IrOp::Add], + is_plus_loop: false, + }], + ); + let result = opt_with_inline(vec![IrOp::Call(WordId(5)), IrOp::Dot], &bodies); + assert!( + matches!(result.first(), Some(IrOp::Call(WordId(5)))), + "loop should not have been inlined, got {result:?}" + ); + } + + #[test] + fn still_inlines_a_loop_into_a_caller_that_can_be_promoted() { + let mut bodies = HashMap::new(); + bodies.insert( + WordId(5), + vec![IrOp::DoLoop { + body: vec![IrOp::PushI32(1), IrOp::Add], + is_plus_loop: false, + }], + ); + let result = opt_with_inline(vec![IrOp::Call(WordId(5)), IrOp::Dup], &bodies); + assert!( + !result.iter().any(|op| matches!(op, IrOp::Call(_))), + "loop should have been inlined, got {result:?}" + ); + } + + #[test] + fn still_inlines_straight_line_words_anywhere() { + // Only loops are held back; a small straight-line word is still + // better off inlined even into an unpromotable caller. + let mut bodies = HashMap::new(); + bodies.insert(WordId(5), vec![IrOp::Dup, IrOp::Mul]); + let result = opt_with_inline(vec![IrOp::Call(WordId(5)), IrOp::Dot], &bodies); + assert!( + !result.iter().any(|op| matches!(op, IrOp::Call(_))), + "straight-line word should still inline, got {result:?}" + ); + } } diff --git a/crates/core/src/outer.rs b/crates/core/src/outer.rs index b15eac9..8cf77e7 100644 --- a/crates/core/src/outer.rs +++ b/crates/core/src/outer.rs @@ -8255,6 +8255,88 @@ mod tests { assert_eq!(stack, vec![1, 0]); } + // -- Region promotion (a hot loop inside an unpromotable word) ----- + + #[test] + fn test_loop_in_an_unpromotable_word_still_computes() { + // `.` keeps MIXED off the register path as a whole, but the loop + // inside it is promoted as its own region. Values checked against + // gforth 0.7.3. + assert_eq!( + eval_output(": MIXED 0 1000 0 DO 1+ LOOP . ; MIXED"), + "1000 " + ); + } + + #[test] + fn test_j_in_a_promoted_region_reads_the_right_loop() { + // A region may only use `I` / `J` when the DO loops they name are + // inside the region itself -- otherwise the simulator resolves them + // against its own empty loop stack. gforth prints 9. + assert_eq!( + eval_output(": JT 0 3 0 DO 3 0 DO J + LOOP LOOP . ; JT"), + "9 " + ); + assert_eq!(eval_output(": IT 0 5 0 DO I + LOOP . ; IT"), "10 "); + } + + #[test] + fn test_promoted_loop_body_that_permutes_the_stack() { + // The values a loop body leaves have to reach the loop-top locals all + // at once. Copying them in index order writes the top into the second + // slot and then reads that slot back, so both come out equal -- this + // printed "4 4" and "3 2 3" before. gforth: "4 3" and "2 1 3". + assert_eq!(eval_output(": C 3 4 2 0 DO SWAP LOOP . . ; C"), "4 3 "); + assert_eq!(eval_output(": D 1 2 3 2 0 DO ROT LOOP . . . ; D"), "2 1 3 "); + } + + #[test] + fn test_promoted_begin_loops() { + // BEGIN loops promote too, so these run entirely in locals. + assert_eq!( + eval_output(": GCD BEGIN DUP WHILE TUCK MOD REPEAT DROP . ; 1071 462 GCD"), + "21 " + ); + assert_eq!( + eval_output(": CD BEGIN 1 - DUP 0= UNTIL DROP 42 . ; 5 CD"), + "42 " + ); + // A WHILE test that permutes: the loop is left between test and body, + // so that exit needs the loop-top locals straightened out as well. + assert_eq!( + eval_output(": W BEGIN SWAP DUP WHILE 1 - SWAP REPEAT . . ; 9 3 W"), + "0 3 " + ); + } + + #[test] + fn test_begin_loop_with_an_unbalanced_body_is_not_promoted() { + // `BEGIN DUP 1+ SWAP DUP 5 > UNTIL` leaves one extra cell per pass, so + // there is no fixed promoted stack shape. It has to keep working. + assert_eq!( + eval_output(": U 0 BEGIN 1 + DUP DUP 3 > UNTIL DROP . . . . ; U"), + "4 3 2 1 " + ); + } + + #[test] + fn test_several_regions_in_one_word() { + // Two loops separated by a `.`: each is its own region, and the + // stack has to survive the hand-off through memory between them. + assert_eq!( + eval_output(": M2 0 10 0 DO I + LOOP DUP . 5 0 DO 1+ LOOP . ; M2"), + "45 50 " + ); + } + + #[test] + fn test_region_hands_results_back_to_the_memory_stack() { + // The region computes in locals; what it leaves has to be visible to + // the interpreter afterwards. + let (stack, _) = eval(": R 7 4 0 DO 1+ LOOP ; 100 R"); + assert_eq!(stack, vec![11, 100]); + } + fn eval(input: &str) -> (Vec, String) { let mut vm = ForthVM::::new().unwrap(); vm.evaluate(input).unwrap(); diff --git a/crates/core/tests/comparison.rs b/crates/core/tests/comparison.rs index baa1470..b98d061 100644 --- a/crates/core/tests/comparison.rs +++ b/crates/core/tests/comparison.rs @@ -746,37 +746,37 @@ fn perf_benchmarks() -> Vec { verify: "25 FIB", expected: 75025, samples: 5, - max_ratio: 0.65, + max_ratio: 0.17, }, PerfBenchmark { - name: "Factorial(12)x10K", + name: "Factorial(12)x100K", define: ": FACT 1 SWAP 1+ 1 ?DO I * LOOP ; \ - : FACT-BENCH 10000 0 DO 12 FACT DROP LOOP ;", + : FACT-BENCH 100000 0 DO 12 FACT DROP LOOP ;", run_code: "FACT-BENCH", verify: "12 FACT", expected: 479001600, samples: 5, - max_ratio: 0.75, + max_ratio: 0.12, }, PerfBenchmark { - name: "GCD-bench(500)", + name: "GCD-bench(20K)", define: ": GCD BEGIN DUP WHILE TUCK MOD REPEAT DROP ; \ : GCD-BENCH 0 DO 10000 I 1+ GCD DROP LOOP ;", - run_code: "500 GCD-BENCH", + run_code: "20000 GCD-BENCH", verify: "48 36 GCD", expected: 12, samples: 5, - max_ratio: 0.70, + max_ratio: 0.45, }, PerfBenchmark { - name: "NestedLoops(50)", + name: "NestedLoops(50)x1K", define: ": NESTED 0 SWAP 0 DO I 0 ?DO I J + DROP LOOP LOOP ; \ - : NESTED-BENCH 100 0 DO 50 NESTED DROP LOOP ;", + : NESTED-BENCH 1000 0 DO 50 NESTED DROP LOOP ;", run_code: "NESTED-BENCH", verify: "5 NESTED", expected: 0, - samples: 3, - max_ratio: 0.20, + samples: 5, + max_ratio: 0.11, }, PerfBenchmark { name: "Collatz(2K)", @@ -788,7 +788,7 @@ fn perf_benchmarks() -> Vec { verify: "27 COLLATZ", expected: 111, samples: 3, - max_ratio: 0.45, + max_ratio: 0.08, }, ] }