1 Commits

Author SHA1 Message Date
ok fa7dadbdeb codegen: LOOP exits on index==limit (Forth 2012 boundary crossing) 2026-07-13 13:04:44 +02:00
34 changed files with 1398 additions and 7307 deletions
-3
View File
@@ -3,6 +3,3 @@
*.swp *.swp
.DS_Store .DS_Store
*.bk *.bk
# Local planning notes — never tracked
/plans/
-370
View File
@@ -1,370 +0,0 @@
# Changelog
All notable changes to WAFER are documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [0.2.8] - 2026-08-10
### Added
- **A recursive word tests its base case at the call site.** A recursive Forth
word almost always opens with a guard that returns early --
`: FIB DUP 2 < IF EXIT THEN ... RECURSE ... ;` -- so every leaf of the
recursion costs a call whose entire body is that test. `Call(self)` now
compiles as `<guard> IF <what the guard returns> ELSE Call(self) THEN`,
which computes the same thing: the callee would have run the guard, taken
the branch and returned. In fib's tree the leaves are half of all nodes.
Fibonacci(25) 356 -> 237 µs on the arm64 development machine, where that
reads 1.24x -> 0.83x of SwiftForth `sf64`. Measured again with **both
engines native on x86-64** -- the macOS `sf64` build runs under Rosetta 2,
which flatters WAFER -- Fibonacci is 1.16x, so it remains the one benchmark
of the five that `sf64` wins. See the two tables in the README.
The guard runs twice along the recursive path, so it has to be small (at
most six operations) and free of effects -- no calls, no memory, no
branches. Words with more than four self-call sites are left alone to bound
the code growth, and a `TailCall` is never expanded.
## [0.2.7] - 2026-08-09
### Added
- **A typed calling convention for words with a known stack effect.** Such a
word now compiles to two entry points: a fast one whose signature is
`(i32 x p) -> (i32 x q)`, carrying its stack items as WASM values, and the
usual `( -- )` wrapper that moves those items on and off the memory data
stack. The wrapper keeps the function-table slot, so `EXECUTE`, the outer
interpreter, host words and `CATCH` see exactly the ABI they saw before;
only direct calls inside a module take the fast entry.
This is what the SwiftForth gap was made of. sf64 keeps TOS in `RBX` and
the stack pointer in `RBP`, and both survive a `CALL` untouched, so its
`FIB` is 16 instructions and ~7 memory touches per node. WAFER kept the
whole stack in linear memory and flushed its cached `$dsp` to an imported
global before every call: ~36 memory touches per node. The stack simulator
that already promoted loop and `IF` bodies into WASM locals refused any
body containing a call or an `EXIT` -- exactly the words where the
convention cost the most. It now handles both.
Fibonacci(25) goes from 1035 to 366 µs, 4.3x slower than `sf64` to 1.2x.
Loop-heavy benchmarks are unchanged by this entry — see the region
promotion below for those. Words that keep the memory convention: anything
using `SP@`, `DEPTH`, `EXECUTE`, `>R`/`R>`, floats or locals; anything
calling a word that is itself untyped, which in the JIT path means every
call except `RECURSE`; mutually recursive words; and words whose effect is
not static -- branches that disagree on depth, `EXIT` at the wrong depth,
a non-neutral loop body, or a recursion that grows the stack per level.
`CONSOLIDATE` extends this across words, since it puts them all in one
module: the effects are solved to a fixpoint from the leaves outward, and
105 of 187 words in a booted dictionary end up typed.
Stack guards get cheap as a side effect -- they hang off the memory-stack
push/pop choke points, and a typed word barely has any. The default
guards-on configuration that the REPL and the web build use went from 1631
to 365 µs on the same benchmark.
`WAFER_TYPED_CALLS=0` falls back to the memory-stack convention.
- **Promotion is now per region, not per word.** Stack-to-local promotion
used to be all-or-nothing: a single `.`, `CR`, `>R` or host call
anywhere in a definition put the _entire_ body on the memory data
stack, hot loops included. The stack simulator now runs over each
stretch of a word that can live in WASM locals, loading what the
region reads and writing back what it leaves, with the rest of the
word unchanged around it.
The cliff this removes was steep. The same loop, same build:
| `: L1 0 5000000 0 DO 1+ LOOP DROP ;` reached as | µs | ns/iter |
| ----------------------------------------------- | ----- | ------- |
| its own word | 1571 | 0.31 |
| inlined into a caller with a `.` in it (before) | 11100 | 2.22 |
| the same, after this change | 1572 | 0.31 |
7x, for one `i32.add`: on the memory path the accumulator is stored to
linear memory and reloaded next iteration, so the loop-carried
dependency runs through store-to-load forwarding instead of a
register.
A region may only use `I` / `J` when the DO loops naming them are
inside the region, since the simulator resolves them against its own
loop stack. Straight-line regions have to be at least three operations
to be worth the load and store either side; a loop always is.
- **The inliner no longer drags a loop onto the memory stack.** It
inlined any callee of eight IR operations or fewer, so a small
loop-bearing word inlined into a caller that can never be promoted
lost its registers -- an optimisation pass applying the 7x
pessimisation above. Loop-bearing callees now stay put in that case:
one call is far cheaper than a loop's worth of memory traffic.
Straight-line words still inline everywhere.
- **`BEGIN` loops promote as well.** `BEGIN..UNTIL`, `BEGIN..AGAIN` and
`BEGIN..WHILE..REPEAT` were rejected outright by the eligibility check,
so any word built on the idiomatic Forth loop kept the memory data
stack no matter how hot it was. They are promoted now when the
construct is stack-neutral: `UNTIL` consumes exactly the flag its body
leaves, `AGAIN`'s body is neutral, and for `WHILE..REPEAT` the test and
the body balance separately -- `WHILE` leaves the loop between the two,
so a net that only added up over the pair would give the two exits
different stack shapes. Bodies containing an `EXIT` stay out, the same
rule `DO`/`LOOP` follows. `BEGIN..WHILE..WHILE..REPEAT` is still
excluded.
GCD 994 -> 540 µs, Collatz 428 -> 185.
Together these four entries put four of the five cross-engine
benchmarks past SwiftForth `sf64`: Factorial 0.29x, Collatz 0.30x,
NestedLoops 0.27x, GCD 0.67x. Fibonacci stays at 1.24x, being pure
call overhead with no loop to promote.
### Fixed
- **A promoted loop or `IF` whose branch permutes the stack lost a value.**
At the bottom of a promoted loop the body's results are copied back into
the loop-top locals, and the join after a promoted `IF` copies one
branch's locals into the other's. Both did it one slot at a time in index
order, which is wrong as soon as a destination is also a later source:
`: C 3 4 2 0 DO SWAP LOOP . . ;` printed `4 4` where gforth and
SwiftForth print `4 3`, and `2 0 DO ROT LOOP` over three cells printed
`3 2 3` instead of `2 1 3`. The copies are now ordered so every source is
read before it is overwritten, with one scratch local to break a cycle.
Present since stack-to-local promotion was introduced; reachable from
any `DO` loop or `IF` whose body reorders cells it did not create.
- The Forth 2012 Core suite now also runs against consolidated code
(`compliance_core_after_consolidate`). `CONSOLIDATE` had no correctness
test at all before -- only benchmarks.
### Changed
- **Three cross-engine benchmarks were too small to be measured.** GCD ran
in 14 µs, Factorial in 49 and NestedLoops in 51, where per-run scatter is
a good fraction of the total and fixed per-invocation costs in the other
engines dominate. Scaled to Factorial x100K, GCD-bench(20K) and
NestedLoops(50)x1K, all now around 0.5-1 ms.
This changed a result rather than just steadying it: GCD looked like a
win at 0.42x of `sf64` and was in fact a loss at 1.17x. That is what
pointed at `BEGIN` loops as the remaining gap -- GCD is the one benchmark
whose loop is a `BEGIN ... WHILE ... REPEAT` -- and with those promoted it
now reads 0.67x. The regression limits, which had drifted to 3-6x looser
than the measurements they guard, were retightened to ~45% above the
current ratios.
## [0.2.6] - 2026-08-07
### Fixed
- **An uncaught `ABORT` no longer prints anything.** It used to report
`ABORT (throw -1)`, but the standard defines `ABORT` as "empty the data
stack and perform the function of `QUIT`", and `QUIT` displays no
message. gforth and SwiftForth are both silent here. `CATCH` still
reports -1 as before, and `ABORT"` still prints its text — that is a
different word with a different code (-2).
- **Compile-only words used in interpretation state name the condition.**
`ABORT"`, `IF`, `THEN`, `LOOP`, `LITERAL`, `RECURSE` and the rest of
the compile-time constructs claimed to be an `unknown word`, which is
actively misleading for a word the system obviously knows. They now
report `interpreting a compile-only word: <name> (throw -14)`, the
standard condition both reference engines give. A genuine typo still
reports `unknown word`.
## [0.2.5] - 2026-08-06
### Added
- **`QUIT`** ( -- ) ( R: i\*x -- ), the CORE word that was missing: empty
the return stack, enter interpretation state, hand the input source
back to the user input device and return to the interpreter without a
message. The data stack is deliberately left alone — that is the whole
difference to `ABORT`, which the standard defines as "empty the data
stack, then `QUIT`". It unwinds through nested `EVALUATE` and
`INCLUDE`, abandoning them, and `SOURCE-ID` is restored to 0.
`CATCH` does **not** report it: `QUIT` rides throw code -56, which the
interpreter treats as a return to the prompt rather than an exception.
Both behaviours were checked against gforth 0.7.3 and SwiftForth
`sf64`, which agree — `1 2 ' QUIT CATCH .` prints nothing and leaves
`1 2` on the stack in all three engines.
The gap had gone unnoticed because the Forth 2012 test suite skips it
by its own admission ("I HAVEN'T FIGURED OUT HOW TO TEST KEY, QUIT,
ABORT, OR ABORT\""), and because `HELP`'s coverage lint compares the
dictionary against the docs — a word absent from both looks complete.
`docs/wafer-anki.txt` had been documenting `QUIT` as if it existed.
Note that `ABORT` was already correct: executing it while a definition
is open does clear both stacks and return to interpretation state.
Typing `ABORT` (or `QUIT`) into an unfinished definition compiles it
rather than running it, exactly as in every other Forth; `[` is the
word that gets you out.
## [0.2.4] - 2026-08-06
### Fixed
- **Errors from host words in the browser build read like Forth errors
again.** A host word signals failure by throwing across the JS
boundary, and the browser runtime reported the exception with its
`Debug` form, so an empty-stack `RESIZE` came back as
`call_func(134) failed: JsValue(Error: Stack underflow ...)` trailed by
an engine stack trace. The thrown message is the Forth message, so it
is now surfaced verbatim — `Stack underflow`, exactly what the native
CLI prints. Exceptions that carry no message keep the call context,
since those are genuine runtime faults rather than Forth throws.
`CATCH` was never affected: it reads the throw code from its own
channel, not from the message.
## [0.2.3] - 2026-08-06
### Fixed
- **Release builds of `wafer-web` no longer fail on proc-macro loading.**
Cargo strips debuginfo from release artifacts by default, and on macOS
that also strips the metadata proc-macro dylibs need to be loadable, so
`wasm-pack build --release` died with `can't find crate` for
`rustversion`, `thiserror_impl` and every other proc-macro. Build
scripts and proc-macros gain nothing from stripping, so
`[profile.release.build-override]` now exempts them; release binaries
stay stripped. Debug builds were never affected, which is why the test
suite stayed green while the browser REPL could not be built for
production.
- `wafer-web` and `wafer-cli` requested `wafer-core` version `0.2.1`
while the workspace had moved to `0.2.2`. The caret requirement still
resolved, so nothing broke, but the pin is now kept in step.
## [0.2.2] - 2026-08-06
### Added
- **SwiftForth-style input number conversion.** Punctuation (`,` `.` `+`
`/` `:` and an embedded `-`) anywhere after the leftmost digit now forces
double-cell conversion, so `12.34`, `1,234`, `12:30:45` and `2026-08-06`
all convert as doubles without a custom parser. Previously only a
trailing `.` worked and `1.5` was an "unknown word" error. The
punctuation is a double-cell marker, not a fractional point: every
spelling of `1234` (`1234.`, `123.4`, `.1234`) yields the same value.
- **`DPL`** ( -- addr ): digits to the right of the rightmost punctuation
character in the last converted number, negative when the token carried
none. Seeded at -1024 and bumped once per digit, matching `sf64`.
Together with `<# #>` this is how fixed-point input is scaled.
- **`NH`** ( -- addr ): the high-order cell dropped by a single-cell
conversion, so a token that overflows a cell can be recovered as a
double (`4000000000 NH @ D.`).
Verified token-for-token against SwiftForth `sf64`: DPL values, double
promotion and sign handling agree on every probed form. One deliberate
divergence — WAFER also accepts a sign before a base prefix (`-$FF`), which
`sf64` rejects; the Forth 2012 spelling `$-FF` works in both. A leading `+`
is punctuation rather than a sign in both engines, so `+7` is the double 7
with `DPL` = 1.
## [0.2.1] - 2026-08-06
### Fixed
- **The search order is now authoritative** (Forth 2012 §16.3.3): a word
whose wordlist is not in the search order is no longer findable.
Previously lookup fell back to the newest entry across all wordlists,
making word hiding impossible. Verified against gforth and SwiftForth,
and guarded by a cross-engine corpus program.
- **Host words validate their stack arguments.** Around 40 host-implemented
words (`RND-SEED`, `ACCEPT`, `RESIZE`, `ALLOCATE`, `FREE`, `SEARCH`,
`SUBSTITUTE`, `ROLL`, `M*`, `UM/MOD`, `SF@ SF! DF@ DF!`, `F. FE. FS. F~`,
`2R@`, and friends) performed raw stack-pointer arithmetic with no
underflow check — calling them on an empty stack silently corrupted the
stack pointer (the compiled-code guards from 0.2.0 do not cover host
words). All argument-taking host words now fail with a clean, CATCHable
underflow error, enforced by a class-wide regression test.
## [0.2.0] - 2026-08-06
The usability release: introspection, source files, honest errors, and a
safety net under every compiled word.
### Added
- **Stack guards in compiled code**: under/overflow checks at the
stack-pointer choke points of generated WASM. Faults THROW standard codes
(`-3`..`-6`, `-44`, `-45`), are CATCHable, and print standard messages
instead of silently corrupting memory. Default on; `wafer build` output
stays unguarded; `WAFER_STACK_GUARDS=0|1` overrides.
- **`SEE`**: source-level decompiler. Colon words (including everything in
`boot.fth`) show their captured verbatim source; data words show
synthesized definitions with current values (`9 VALUE X`,
`DEFER D ( IS DUP )`); primitives fall back to a readable IR dump —
`SEE` never dead-ends on a defined word.
- **`SEE-IR`**: post-optimization IR view with resolved callee names and
indented control flow — shows what the optimizer actually did.
- **`HELP`**: stack effect + one-line description for **every** word in a
fresh VM (dictionary words and outer-interpreter tokens alike); coverage
is enforced by a unit test, so an undocumented new word fails the build.
User words echo their leading `( n -- n )` comment.
- **`INCLUDE` / `INCLUDED`**: nestable source-file loading with cycle
detection, depth bound, paths relative to the including file, and
per-level `SOURCE-ID`. The loader is injected (CLI: filesystem; web:
defined error), so the core stays IO-free. `wafer prog.fth` now runs
through the same machinery.
- **`MARKER` extensions**: `REMEMBER` (re-runnable marker), `EMPTY` and
`GILD` (boot-state rollback and re-baselining). Marker rollback now also
restores search order, wordlists, `REPLACES` substitutions, `ABORT"`
texts, and captured word sources — enabling the `REMEMBER` + `INCLUDE`
edit-reload loop.
- **`WORDS`**: optional substring filter (`WORDS FLOAT`), word count, and
`WORDS ALL` — a grouped full view by wordlist plus internal words.
- **Return-stack introspection**: `.RS`, `RDEPTH`, `RP@`.
- **Tools**: `.S` honors `BASE`, `F.S`, `?`, bounds-checked `DUMP`, real
`BYE`, named `ORDER` output.
- **CLI REPL**: persistent history (XDG state dir, `0600`), dictionary-backed
tab completion, prefix history search on Up/Down, Ctrl-C clears the line.
- **Web REPL**: history persisted to localStorage, User Words palette,
`BASE` indicator in the stack bar.
- **Error reporting**: uncaught `THROW` codes map to standard messages;
`ABORT"` text prints only when uncaught; errors inside included files
carry `file.fth:line:` context; uncaught throws are typed
(`WaferError::UncaughtThrow`) for embedding consumers; compiled words
carry WASM name sections, so genuine traps name the faulting word
(`in CRASHER: wasm trap: out of bounds memory access`).
- **SwiftForth correctness lane**: the cross-engine program corpus can run
against sf64 as an oracle (`just compare-correctness`), alongside the
existing gforth lane and the sf64 performance lane.
### Fixed
- Multi-line command output in the CLI REPL starts on its own line
(inline `ok` echo only for single-line output).
- `.S` printed in decimal regardless of `BASE`.
- A bare interpreted `R>` underflowed silently (exposed by the new stack
guards; compliance baseline updated).
- `SPACES` with a negative count now outputs nothing, per Forth 2012
6.1.2230.
### Changed
- `wafer prog.fth` reports errors with `file:line` context and resolves
nested `INCLUDE`s relative to the file.
- Internal words (`_`-prefixed) are flagged in the dictionary and hidden
from `WORDS` and completion (`WORDS ALL` shows them).
- Dependencies upgraded across the board: wasmtime 43 → 47,
wasm-encoder/wasmparser 0.246 → 0.255, plus all semver-compatible
updates.
## [0.1.0] - 2026-08-04
Initial development line (untagged): Forth 2012 core with IR optimizer and
WASM codegen via wasm-encoder/wasmtime, ~300 words across Core, Double,
Float, String, Search-Order, Exception, and Tools word sets, Forth 2012
compliance suite, `CONSOLIDATE` whole-program recompilation, `wafer build`
AOT export (WASM / native / JS loader), browser REPL, SHA-1/256/512 words,
and cross-engine benchmark lanes against gforth and SwiftForth.
[0.2.8]: https://github.com/ok2/wafer/compare/v0.2.7...v0.2.8
[0.2.7]: https://github.com/ok2/wafer/compare/v0.2.6...v0.2.7
[0.2.1]: https://github.com/ok2/wafer/compare/v0.2.0...v0.2.1
[0.2.0]: https://github.com/ok2/wafer/compare/v0.1.0...v0.2.0
[0.1.0]: https://github.com/ok2/wafer/releases/tag/v0.1.0
+2 -2
View File
@@ -2,7 +2,7 @@
## What is WAFER? ## What is WAFER?
WAFER (WebAssembly Forth Engine in Rust) is an optimizing Forth 2012 compiler targeting WebAssembly. Currently a working Forth system with 200+ words, JIT compilation, 12 word sets at 100% compliance, and a full optimization pipeline (peephole, constant folding, inlining, strength reduction, DCE, tail calls, per-region stack-to-local promotion with DO/BEGIN loop and IF support, self-recursive direct calls, a typed calling convention for words with a known stack effect, self-guard expansion for recursive words, consolidation). Beats gforth on every benchmark, and SwiftForth `sf64` on four of five (measured native-vs-native on x86-64; the macOS sf64 build is x86-64 under Rosetta and flatters WAFER). Includes a browser-based REPL via wasm-pack. WAFER (WebAssembly Forth Engine in Rust) is an optimizing Forth 2012 compiler targeting WebAssembly. Currently a working Forth system with 200+ words, JIT compilation, 12 word sets at 100% compliance, and a full optimization pipeline (peephole, constant folding, inlining, strength reduction, DCE, tail calls, stack-to-local promotion with loop/IF support, self-recursive direct calls, consolidation). Beats gforth on all benchmarks in release mode. Includes a browser-based REPL via wasm-pack.
## Architecture ## Architecture
@@ -79,7 +79,7 @@ Handle in `interpret_token_immediate()` or `compile_token()` as a special case.
## Testing ## Testing
- Run `cargo test --workspace` before committing (currently 608 unit + 1 benchmark + 12 compliance + 9 comparison + 5 crypto) - Run `cargo test --workspace` before committing (currently 431 unit + 1 benchmark + 11 compliance + 9 comparison)
- Forth 2012 compliance: `cargo test -p wafer-core --test compliance` - Forth 2012 compliance: `cargo test -p wafer-core --test compliance`
- Cross-engine comparison (vs gforth): `cargo test -p wafer-core --test comparison` - Cross-engine comparison (vs gforth): `cargo test -p wafer-core --test comparison`
- Performance benchmarks (release mode): `cargo test -p wafer-core --test comparison -- --nocapture --ignored` - Performance benchmarks (release mode): `cargo test -p wafer-core --test comparison -- --nocapture --ignored`
Generated
+634 -311
View File
File diff suppressed because it is too large Load Diff
+6 -14
View File
@@ -3,7 +3,7 @@ members = ["crates/*"]
resolver = "2" resolver = "2"
[workspace.package] [workspace.package]
version = "0.2.8" version = "0.1.0"
edition = "2024" edition = "2024"
license = "MIT OR Apache-2.0" license = "MIT OR Apache-2.0"
repository = "https://github.com/ok2/wafer" repository = "https://github.com/ok2/wafer"
@@ -41,21 +41,13 @@ needless_collect = "warn"
or_fun_call = "warn" or_fun_call = "warn"
[workspace.dependencies] [workspace.dependencies]
wasm-encoder = "0.255" wasm-encoder = "0.246"
wasmparser = "0.255" wasmparser = "0.246"
wasmtime = "47" wasmtime = "43"
anyhow = "1" anyhow = "1"
thiserror = "2" thiserror = "2"
proptest = "1" proptest = "1"
insta = "1" insta = "1"
sha1 = "0.10" sha1 = "0.11"
sha2 = "0.10" sha2 = "0.11"
send_wrapper = "0.6" send_wrapper = "0.6"
# Cargo strips debuginfo from release artifacts by default, and on macOS that
# also strips the metadata proc-macro dylibs need to be loadable — release
# builds then fail with "can't find crate" for every proc-macro (rustversion,
# thiserror_impl, ...). Build scripts and proc-macros gain nothing from
# stripping, so exempt them; the release binaries stay stripped.
[profile.release.build-override]
strip = false
-15
View File
@@ -43,14 +43,6 @@ bench:
bench-opts: bench-opts:
cargo test -p wafer-core --test benchmark_report -- --nocapture --ignored cargo test -p wafer-core --test benchmark_report -- --nocapture --ignored
# Cross-engine performance report: WAFER vs gforth vs SwiftForth (sf64)
bench-compare:
CARGO_PROFILE_RELEASE_STRIP=none cargo test -p wafer-core --release --test comparison -- --nocapture --ignored performance_report
# Cross-engine correctness lanes: program corpus vs gforth + sf64 oracles
compare-correctness:
cargo test -p wafer-core --test comparison -- --nocapture --ignored compare_all_programs
# Check dependency licenses and advisories # Check dependency licenses and advisories
deny: deny:
cargo deny check cargo deny check
@@ -66,13 +58,6 @@ ci: fmt clippy deny test
check: check:
cargo check --workspace cargo check --workspace
# Install the wafer CLI (release build) and bat syntax highlighting.
# STRIP=none: Cargo's release default (strip = "debuginfo") emits dylibs that
# macOS 27's dyld rejects ("mis-aligned LINKEDIT string pool"), so proc macros
# fail to load during the build itself.
install: install-syntax
CARGO_PROFILE_RELEASE_STRIP=none cargo install --path crates/cli --locked
# Install bat syntax highlighting for WAFER / Forth # Install bat syntax highlighting for WAFER / Forth
install-syntax: install-syntax:
mkdir -p ~/.config/bat/syntaxes mkdir -p ~/.config/bat/syntaxes
+12 -55
View File
@@ -7,11 +7,10 @@ An optimizing Forth 2012 compiler targeting WebAssembly. WAFER JIT-compiles each
## Highlights ## Highlights
- **200+ words** across 12 Forth 2012 word sets, all at **100% compliance** - **200+ words** across 12 Forth 2012 word sets, all at **100% compliance**
- **Optimizing compiler** with 6 IR passes + stack-to-local promotion (per region, so a hot loop keeps its registers even inside a word that does I/O; `DO` and `BEGIN` loops alike) + consolidation - **Optimizing compiler** with 6 IR passes + stack-to-local promotion (loops + IF) + consolidation
- **Faster than gforth** on every benchmark, and past SwiftForth `sf64` -- a native-code compiler -- on four of five - **Faster than gforth** on all benchmarks in release mode (2-10x faster)
- **JIT compilation** — each `:` definition compiles to its own WASM module - **JIT compilation** — each `:` definition compiles to its own WASM module
- **Self-recursive direct calls** — RECURSE compiles to native `call` instead of `call_indirect` - **Self-recursive direct calls** — RECURSE compiles to native `call` instead of `call_indirect`
- **Typed calling convention** — a word with a statically known stack effect passes its stack items as WASM values, so a call keeps them in registers instead of round-tripping through memory
- **Consolidation mode** — recompile all words into a single optimized WASM module - **Consolidation mode** — recompile all words into a single optimized WASM module
- **Interactive REPL** with line editing (rustyline) - **Interactive REPL** with line editing (rustyline)
- **Browser REPL** — runs entirely in the browser via wasm-pack + js-sys - **Browser REPL** — runs entirely in the browser via wasm-pack + js-sys
@@ -80,64 +79,23 @@ git submodule update --init
## Performance ## Performance
WAFER beats gforth (the GNU Forth reference implementation) on every benchmark, and SwiftForth WAFER beats gforth (the GNU Forth reference implementation) on all benchmarks in release mode:
`sf64` -- which compiles to native code -- on four of the five.
Measured with all three engines running **native x86-64**, on an idle 16-vCPU Xeon Platinum 8124M
@ 3.0 GHz (median of three runs):
``` ```
Benchmark WAFER gforth sf64 WAFER/gf WAFER/sf Benchmark WAFER CONSOL gforth WAFER/gf
Fibonacci(25) 411 3221 355 0.13x 1.16x Fibonacci(25) 1629 1535 3422 0.45x
Factorial(12)x100K 994 7141 3058 0.14x 0.33x Factorial(12)x10K 340 339 638 0.53x
GCD-bench(20K) 1591 3211 2423 0.50x 0.66x GCD-bench(500) 18 15 30 0.50x
NestedLoops(50)x1K 889 6824 2342 0.13x 0.38x NestedLoops(50) 84 73 720 0.10x
Collatz(2K) 391 3981 1659 0.10x 0.24x Collatz(2K) 1212 1202 3914 0.31x
``` ```
Times in microseconds; WAFER is the better of the JIT and `CONSOLIDATE` runs. Below 1.0 means WAFER Times in microseconds. WAFER/gf < 1.0 means WAFER is faster. CONSOL = after `CONSOLIDATE`.
is faster. Fibonacci is the one WAFER loses: it is one call per node with no loop to promote, and
`sf64` keeps its stack in registers across a call the way only a native code generator can.
Fibonacci, GCD and Collatz held to within 2% across the three runs; Factorial and NestedLoops are
softer, since `sf64` varied by half there, but they are wide wins either way.
`just bench-compare` on the development machine (M1 Ultra, arm64) reports different numbers, and
they flatter WAFER:
```
Benchmark WAFER CONSOL gforth sf64 WAFER/gf WAFER/sf
Fibonacci(25) 237 242 3340 287 0.07x 0.83x
Factorial(12)x100K 480 479 6109 1594 0.08x 0.30x
GCD-bench(20K) 549 541 1830 797 0.30x 0.68x
NestedLoops(50)x1K 501 509 7092 1898 0.07x 0.26x
Collatz(2K) 196 190 3955 633 0.05x 0.30x
```
The only SwiftForth build for macOS is x86-64 under Rosetta 2, while WAFER and gforth are native
arm64 -- so that `sf64` column is native against emulated. The gap is not small, and it lands
exactly where it matters: Fibonacci reads 0.83x there and 1.16x when neither engine is emulated.
Treat the arm64 table as what the regression limits in `comparison.rs` are calibrated against, and
the x86-64 table as what to believe about the engines.
A caveat applies to both: `sf64` uses 64-bit cells to WAFER's 32-bit, so WAFER does less work per
operation.
A word whose stack effect is statically known gets a **typed entry point**: its stack items travel in and out
as WASM values instead of through the memory data stack, so cranelift keeps them in registers across a call
the way a native Forth keeps TOS in one. The word also keeps a `( -- )` wrapper, which is what the function
table, `EXECUTE` and the outer interpreter reach, so nothing about the memory ABI changes from the outside.
Call-heavy code is what this pays for -- Fibonacci went from 4.3x slower than `sf64` to 1.2x. Set
`WAFER_TYPED_CALLS=0` to fall back to the memory-stack convention.
Recursive words then get one more thing: their base-case guard is tested at the **call site**, so a
leaf of the recursion costs a comparison instead of a call. `: FIB DUP 2 < IF EXIT THEN ... RECURSE`
compiles its `RECURSE` as `DUP 2 < IF ELSE RECURSE THEN`, which is what the callee would have done
on entry anyway. Half of fib's nodes are leaves, and that is worth 1.4x.
## Testing ## Testing
```bash ```bash
# All tests (~635 currently passing) # All tests (~450 currently passing)
cargo test --workspace cargo test --workspace
# Forth 2012 compliance suite # Forth 2012 compliance suite
@@ -170,7 +128,7 @@ Forth Source -> Outer Interpreter -> IR -> [Optimize] -> WASM Codegen (wasm-enco
- `WebRuntime` — browser WebAssembly API via js-sys, for the browser REPL - `WebRuntime` — browser WebAssembly API via js-sys, for the browser REPL
- **Subroutine threading** via WASM function tables (`call_indirect` for cross-word, direct `call` for self-recursion) - **Subroutine threading** via WASM function tables (`call_indirect` for cross-word, direct `call` for self-recursion)
- **JIT mode**: each new word compiles to a separate WASM module linked to shared memory/globals/table - **JIT mode**: each new word compiles to a separate WASM module linked to shared memory/globals/table
- **IR-based pipeline** with 6 optimization passes (peephole, constant folding, strength reduction, DCE, tail call detection, inlining) plus per-region stack-to-local promotion (DO and BEGIN loops, IF/ELSE), DO/LOOP index locals, typed entry points for words with a known stack effect, self-guard expansion, and consolidation - **IR-based pipeline** with 6 optimization passes (peephole, constant folding, strength reduction, DCE, tail call detection, inlining) plus stack-to-local promotion (with loop and IF/ELSE support), DO/LOOP index locals, and consolidation
- **Dictionary**: linked-list word headers in simulated linear memory - **Dictionary**: linked-list word headers in simulated linear memory
## Project Structure ## Project Structure
@@ -227,7 +185,6 @@ Over 200 words are implemented across the following categories:
| Strings | `COMPARE SEARCH SLITERAL REPLACES SUBSTITUTE UNESCAPE` | | Strings | `COMPARE SEARCH SLITERAL REPLACES SUBSTITUTE UNESCAPE` |
| Floating-Pt | `F+ F- F* F/ FABS FNEGATE FSQRT FSIN FCOS FTAN FEXP FLOG FMIN FMAX` and 55+ more | | Floating-Pt | `F+ F- F* F/ FABS FNEGATE FSQRT FSIN FCOS FTAN FEXP FLOG FMIN FMAX` and 55+ more |
| Case | `CASE OF ENDOF ENDCASE` | | Case | `CASE OF ENDOF ENDCASE` |
| Tools | `WORDS SEE SEE-IR HELP INCLUDE INCLUDED .S F.S ? DUMP MARKER REMEMBER EMPTY GILD BYE` |
## Web REPL ## Web REPL
+1 -1
View File
@@ -9,7 +9,7 @@ license.workspace = true
workspace = true workspace = true
[dependencies] [dependencies]
wafer-core = { path = "../core", version = "0.2.8" } wafer-core = { path = "../core", version = "0.1.0" }
wasmtime = { workspace = true } wasmtime = { workspace = true }
anyhow = { workspace = true } anyhow = { workspace = true }
clap = { version = "4", features = ["derive"] } clap = { version = "4", features = ["derive"] }
+58 -192
View File
@@ -137,9 +137,7 @@ fn cmd_build(
) -> anyhow::Result<()> { ) -> anyhow::Result<()> {
let source = std::fs::read_to_string(file)?; let source = std::fs::read_to_string(file)?;
// Exported modules are production artifacts: no stack guards by default let mut vm = ForthVM::<NativeRuntime>::new()?;
let mut vm = ForthVM::<NativeRuntime>::new_with_config(vm_config(false))?;
vm.set_source_loader(fs_loader());
vm.set_recording(true); vm.set_recording(true);
vm.evaluate(&source)?; vm.evaluate(&source)?;
@@ -262,40 +260,18 @@ fn cmd_run(file: &str) -> anyhow::Result<()> {
Ok(()) Ok(())
} }
/// `WaferConfig` for CLI-created VMs. `WAFER_STACK_GUARDS=0|1` overrides
/// the per-command default (REPL/file execution on, build off);
/// `WAFER_TYPED_CALLS=0` falls back to the memory-stack calling convention.
fn vm_config(default_guards: bool) -> wafer_core::config::WaferConfig {
let mut cfg = wafer_core::config::WaferConfig::all();
cfg.codegen.stack_guards = match std::env::var("WAFER_STACK_GUARDS").ok().as_deref() {
Some("0") => false,
Some(_) => true,
None => default_guards,
};
cfg.codegen.typed_calls = std::env::var("WAFER_TYPED_CALLS").ok().as_deref() != Some("0");
cfg
}
/// Filesystem source loader for INCLUDE/INCLUDED.
fn fs_loader() -> Box<dyn Fn(&str) -> anyhow::Result<String> + Send + Sync> {
Box::new(|path| Ok(std::fs::read_to_string(path)?))
}
/// `wafer` (REPL) or `wafer program.fth` (evaluate and exit) /// `wafer` (REPL) or `wafer program.fth` (evaluate and exit)
fn cmd_eval_or_repl(file: Option<&str>) -> anyhow::Result<()> { fn cmd_eval_or_repl(file: Option<&str>) -> anyhow::Result<()> {
let mut vm = ForthVM::<NativeRuntime>::new_with_config(vm_config(true))?; let mut vm = ForthVM::<NativeRuntime>::new()?;
vm.set_source_loader(fs_loader());
match file { match file {
Some(file) => { Some(file) => {
// Through the include machinery: file:line error context and a let source = std::fs::read_to_string(file)?;
// base directory for nested INCLUDEs. vm.evaluate(&source)?;
let result = vm.include(file);
let output = vm.take_output(); let output = vm.take_output();
if !output.is_empty() { if !output.is_empty() {
print!("{output}"); print!("{output}");
} }
result?;
} }
None => { None => {
if !stdin_is_tty() { if !stdin_is_tty() {
@@ -309,17 +285,66 @@ fn cmd_eval_or_repl(file: Option<&str>) -> anyhow::Result<()> {
if !output.is_empty() { if !output.is_empty() {
print!("{output}"); print!("{output}");
} }
if vm.bye_requested() {
break;
}
} }
Err(e) => { Err(e) => {
eprintln!("Error: {e:#}"); eprintln!("Error: {e}");
} }
} }
} }
} else { } else {
run_repl(&mut vm)?; // Interactive REPL
println!(
"WAFER v{} - WebAssembly Forth Engine in Rust",
env!("CARGO_PKG_VERSION")
);
println!("Type BYE to exit.");
let mut rl = rustyline::DefaultEditor::new()?;
loop {
let prompt = if vm.is_compiling() { " ] " } else { "> " };
match rl.readline(prompt) {
Ok(line) => {
let trimmed = line.trim();
if trimmed.eq_ignore_ascii_case("BYE") {
break;
}
let _ = rl.add_history_entry(&line);
match vm.evaluate(&line) {
Ok(()) => {
let output = vm.take_output();
// PAGE (form feed) clears the terminal
if output.contains('\x0C') {
print!("\x1b[2J\x1b[H");
}
let output = output.replace('\x0C', "");
if !vm.is_compiling() {
// Move cursor back up to end of input line so
// output appears inline, like traditional Forth:
// > 2 2 + . 4 ok
let col = prompt.len() + line.len() + 1;
print!("\x1b[A\x1b[{col}G {output} ok");
println!();
} else if !output.is_empty() {
print!("{output}");
}
}
Err(e) => {
eprintln!("Error: {e}");
}
}
}
Err(
rustyline::error::ReadlineError::Interrupted
| rustyline::error::ReadlineError::Eof,
) => {
break;
}
Err(e) => {
eprintln!("Readline error: {e}");
break;
}
}
}
} }
} }
} }
@@ -332,162 +357,3 @@ fn stdin_is_tty() -> bool {
use std::io::IsTerminal; use std::io::IsTerminal;
std::io::stdin().is_terminal() std::io::stdin().is_terminal()
} }
/// Completes the token under the cursor against the live dictionary.
struct WaferHelper {
words: Vec<String>,
}
impl rustyline::completion::Completer for WaferHelper {
type Candidate = String;
fn complete(
&self,
line: &str,
pos: usize,
_ctx: &rustyline::Context<'_>,
) -> rustyline::Result<(usize, Vec<String>)> {
let start = line[..pos]
.rfind(|c: char| c.is_whitespace())
.map_or(0, |i| i + 1);
let prefix = line[start..pos].to_ascii_uppercase();
let mut matches: Vec<String> = self
.words
.iter()
.filter(|w| w.to_ascii_uppercase().starts_with(&prefix))
.cloned()
.collect();
matches.sort();
matches.dedup();
Ok((start, matches))
}
}
impl rustyline::hint::Hinter for WaferHelper {
type Hint = String;
}
impl rustyline::highlight::Highlighter for WaferHelper {}
impl rustyline::validate::Validator for WaferHelper {}
impl rustyline::Helper for WaferHelper {}
/// History file: `$WAFER_HISTORY`, else `$XDG_STATE_HOME/wafer/history`,
/// else `~/.local/state/wafer/history`.
fn history_path() -> Option<std::path::PathBuf> {
if let Some(p) = std::env::var_os("WAFER_HISTORY") {
return Some(p.into());
}
let base = std::env::var_os("XDG_STATE_HOME")
.map(std::path::PathBuf::from)
.or_else(|| {
std::env::var_os("HOME").map(|h| std::path::PathBuf::from(h).join(".local/state"))
})?;
Some(base.join("wafer/history"))
}
/// Interactive REPL: line editing, persistent history with prefix search
/// on Up/Down, and Tab completion over the live dictionary.
fn run_repl(vm: &mut ForthVM<NativeRuntime>) -> anyhow::Result<()> {
use rustyline::{Cmd, Editor, EventHandler, KeyCode, KeyEvent, Modifiers};
println!(
"WAFER v{} - WebAssembly Forth Engine in Rust",
env!("CARGO_PKG_VERSION")
);
println!("Type BYE to exit.");
let config = rustyline::Config::builder()
.completion_type(rustyline::CompletionType::List)
.history_ignore_dups(true)?
.build();
let mut rl: Editor<WaferHelper, rustyline::history::DefaultHistory> =
Editor::with_config(config)?;
rl.set_helper(Some(WaferHelper {
words: vm.word_names(),
}));
// Up/Down recall only entries starting with the typed prefix
rl.bind_sequence(
KeyEvent(KeyCode::Up, Modifiers::NONE),
EventHandler::Simple(Cmd::HistorySearchBackward),
);
rl.bind_sequence(
KeyEvent(KeyCode::Down, Modifiers::NONE),
EventHandler::Simple(Cmd::HistorySearchForward),
);
let history = history_path();
if let Some(path) = &history {
if let Some(dir) = path.parent() {
let _ = std::fs::create_dir_all(dir);
}
let _ = rl.load_history(path);
}
let save_history = |rl: &mut Editor<WaferHelper, rustyline::history::DefaultHistory>| {
if let Some(path) = &history {
let _ = rl.save_history(path);
#[cfg(unix)]
{
use std::os::unix::fs::PermissionsExt;
let _ = std::fs::set_permissions(path, std::fs::Permissions::from_mode(0o600));
}
}
};
loop {
let prompt = if vm.is_compiling() { " ] " } else { "> " };
match rl.readline(prompt) {
Ok(line) => {
let _ = rl.add_history_entry(&line);
match vm.evaluate(&line) {
Ok(()) => {
let output = vm.take_output();
if vm.bye_requested() {
break;
}
// PAGE (form feed) clears the terminal
if output.contains('\x0C') {
print!("\x1b[2J\x1b[H");
}
let output = output.replace('\x0C', "");
if !vm.is_compiling() {
if output.contains('\n') {
// Multi-line output (DUMP, WORDS, ...):
// print as a block, then ok on its own line
print!("{output}");
if !output.ends_with('\n') {
println!();
}
println!(" ok");
} else {
// Move cursor back up to end of input line so
// output appears inline, like traditional Forth:
// > 2 2 + . 4 ok
let col = prompt.len() + line.len() + 1;
print!("\x1b[A\x1b[{col}G {output} ok");
println!();
}
} else if !output.is_empty() {
print!("{output}");
}
}
Err(e) => {
eprintln!("Error: {e:#}");
}
}
// New definitions may have appeared: refresh completion
if let Some(h) = rl.helper_mut() {
h.words = vm.word_names();
}
save_history(&mut rl);
}
// Ctrl-C abandons the current line, Ctrl-D exits
Err(rustyline::error::ReadlineError::Interrupted) => {}
Err(rustyline::error::ReadlineError::Eof) => break,
Err(e) => {
eprintln!("Readline error: {e}");
break;
}
}
}
save_history(&mut rl);
Ok(())
}
+2 -33
View File
@@ -72,19 +72,6 @@
1- 1-
REPEAT ; REPEAT ;
\ ---------------------------------------------------------------
\ Common extensions (not in Forth 2012, gforth-compatible)
\ ---------------------------------------------------------------
\ -ROT ( x1 x2 x3 -- x3 x1 x2 ) rotate top item to third place
: -ROT ROT ROT ;
\ <= ( n1 n2 -- flag ) true if n1 <= n2 (signed)
: <= > 0= ;
\ >= ( n1 n2 -- flag ) true if n1 >= n2 (signed)
: >= < 0= ;
\ --------------------------------------------------------------- \ ---------------------------------------------------------------
\ Phase 2: Double-cell arithmetic \ Phase 2: Double-cell arithmetic
\ --------------------------------------------------------------- \ ---------------------------------------------------------------
@@ -197,8 +184,8 @@
\ TYPE ( c-addr u -- ) output u characters \ TYPE ( c-addr u -- ) output u characters
: TYPE 0 ?DO DUP C@ EMIT 1+ LOOP DROP ; : TYPE 0 ?DO DUP C@ EMIT 1+ LOOP DROP ;
\ SPACES ( n -- ) output n spaces (nothing for n <= 0, per 6.1.2230) \ SPACES ( n -- ) output n spaces
: SPACES 0 MAX 0 ?DO SPACE LOOP ; : SPACES 0 ?DO SPACE LOOP ;
\ Pictured numeric output constants \ Pictured numeric output constants
\ PICT_BUF_TOP = 0x05C0 = 1472, SYSVAR_HLD = 28 \ PICT_BUF_TOP = 0x05C0 = 1472, SYSVAR_HLD = 28
@@ -243,9 +230,6 @@
\ U. ( u -- ) print unsigned number and space \ U. ( u -- ) print unsigned number and space
: U. 0 <# #S #> TYPE SPACE ; : U. 0 <# #S #> TYPE SPACE ;
\ ? ( a-addr -- ) fetch and print
: ? @ . ;
\ .R ( n width -- ) print right-justified signed number \ .R ( n width -- ) print right-justified signed number
: .R >R DUP ABS 0 <# #S ROT SIGN #> R> OVER - SPACES TYPE ; : .R >R DUP ABS 0 <# #S ROT SIGN #> R> OVER - SPACES TYPE ;
@@ -258,21 +242,6 @@
\ D.R ( d width -- ) print right-justified signed double \ D.R ( d width -- ) print right-justified signed double
: D.R >R SWAP OVER DABS <# #S ROT SIGN #> R> OVER - SPACES TYPE ; : D.R >R SWAP OVER DABS <# #S ROT SIGN #> R> OVER - SPACES TYPE ;
\ ---------------------------------------------------------------
\ Return-stack introspection (debug aids)
\ ---------------------------------------------------------------
\ RDEPTH ( -- n ) number of cells on the return stack
\ RETURN_STACK_TOP = 9728 (0x2600). Only >R temps and loop params
\ live there; return addresses are on the WASM call stack.
: RDEPTH 9728 RP@ - 2 RSHIFT ;
\ .RS ( -- ) print the return stack bottom-to-top, like .S
\ Walks with BEGIN/WHILE (not DO) so the walk itself never pushes
\ onto the return stack it is printing.
: .RS ." R:<" RDEPTH 0 .R ." > "
9728 BEGIN DUP RP@ > WHILE 4 - DUP @ . REPEAT DROP ;
\ --------------------------------------------------------------- \ ---------------------------------------------------------------
\ Phase 6: DEFER support \ Phase 6: DEFER support
\ --------------------------------------------------------------- \ ---------------------------------------------------------------
+129 -1314
View File
File diff suppressed because it is too large Load Diff
-16
View File
@@ -7,16 +7,6 @@ use crate::optimizer::OptConfig;
pub struct CodegenOpts { pub struct CodegenOpts {
/// Enable stack-to-local promotion for straight-line words. /// Enable stack-to-local promotion for straight-line words.
pub stack_to_local_promotion: bool, pub stack_to_local_promotion: bool,
/// Emit stack under/overflow guards in compiled words. Faults throw
/// standard codes (-3/-4/-5/-6/-44/-45) instead of silently
/// corrupting stack pointers. On by default; benchmarks and
/// exported production modules turn it off.
pub stack_guards: bool,
/// Compile words with a statically known stack effect to a typed entry
/// point that carries stack items in WASM values, so a call keeps them
/// in registers instead of round-tripping through the memory stack.
/// On by default; `WAFER_TYPED_CALLS=0` turns it off.
pub typed_calls: bool,
} }
/// Master configuration for all WAFER optimizations. /// Master configuration for all WAFER optimizations.
@@ -39,12 +29,9 @@ impl WaferConfig {
strength_reduce: true, strength_reduce: true,
dce: true, dce: true,
inline: true, inline: true,
self_guard: true,
}, },
codegen: CodegenOpts { codegen: CodegenOpts {
stack_to_local_promotion: true, stack_to_local_promotion: true,
stack_guards: true,
typed_calls: true,
}, },
} }
} }
@@ -59,12 +46,9 @@ impl WaferConfig {
strength_reduce: false, strength_reduce: false,
dce: false, dce: false,
inline: false, inline: false,
self_guard: false,
}, },
codegen: CodegenOpts { codegen: CodegenOpts {
stack_to_local_promotion: false, stack_to_local_promotion: false,
stack_guards: false,
typed_calls: false,
}, },
} }
} }
+9 -9
View File
@@ -21,7 +21,7 @@ mod tests {
// Empty word list should produce nothing (but we guard against this at call site) // Empty word list should produce nothing (but we guard against this at call site)
let words = vec![]; let words = vec![];
let map = HashMap::new(); let map = HashMap::new();
let result = compile_consolidated_module(&words, &map, 16, None, true); let result = compile_consolidated_module(&words, &map, 16);
// Empty is valid -- should produce a valid module with no functions // Empty is valid -- should produce a valid module with no functions
assert!(result.is_ok()); assert!(result.is_ok());
} }
@@ -31,7 +31,7 @@ mod tests {
let words = vec![(WordId(1), vec![IrOp::PushI32(42)])]; let words = vec![(WordId(1), vec![IrOp::PushI32(42)])];
let mut map = HashMap::new(); let mut map = HashMap::new();
map.insert(WordId(1), 1u32); // function index 1 (after emit import) map.insert(WordId(1), 1u32); // function index 1 (after emit import)
let result = compile_consolidated_module(&words, &map, 16, None, true); let result = compile_consolidated_module(&words, &map, 16);
assert!(result.is_ok()); assert!(result.is_ok());
} }
@@ -49,7 +49,7 @@ mod tests {
map.insert(WordId(1), 1u32); map.insert(WordId(1), 1u32);
map.insert(WordId(2), 2u32); map.insert(WordId(2), 2u32);
map.insert(WordId(3), 3u32); map.insert(WordId(3), 3u32);
let result = compile_consolidated_module(&words, &map, 16, None, true); let result = compile_consolidated_module(&words, &map, 16);
assert!(result.is_ok()); assert!(result.is_ok());
} }
@@ -59,7 +59,7 @@ mod tests {
let words = vec![(WordId(3), vec![IrOp::Call(WordId(99))])]; let words = vec![(WordId(3), vec![IrOp::Call(WordId(99))])];
let mut map = HashMap::new(); let mut map = HashMap::new();
map.insert(WordId(3), 1u32); map.insert(WordId(3), 1u32);
let result = compile_consolidated_module(&words, &map, 256, None, true); let result = compile_consolidated_module(&words, &map, 256);
assert!(result.is_ok()); assert!(result.is_ok());
} }
@@ -72,7 +72,7 @@ mod tests {
let mut map = HashMap::new(); let mut map = HashMap::new();
map.insert(WordId(1), 1u32); map.insert(WordId(1), 1u32);
map.insert(WordId(2), 2u32); map.insert(WordId(2), 2u32);
let result = compile_consolidated_module(&words, &map, 16, None, true); let result = compile_consolidated_module(&words, &map, 16);
assert!(result.is_ok()); assert!(result.is_ok());
} }
@@ -95,7 +95,7 @@ mod tests {
let mut map = HashMap::new(); let mut map = HashMap::new();
map.insert(WordId(1), 1u32); map.insert(WordId(1), 1u32);
map.insert(WordId(2), 2u32); map.insert(WordId(2), 2u32);
let result = compile_consolidated_module(&words, &map, 16, None, true); let result = compile_consolidated_module(&words, &map, 16);
assert!(result.is_ok()); assert!(result.is_ok());
} }
@@ -120,7 +120,7 @@ mod tests {
let mut map = HashMap::new(); let mut map = HashMap::new();
map.insert(WordId(1), 1u32); map.insert(WordId(1), 1u32);
map.insert(WordId(2), 2u32); map.insert(WordId(2), 2u32);
let result = compile_consolidated_module(&words, &map, 16, None, true); let result = compile_consolidated_module(&words, &map, 16);
assert!(result.is_ok()); assert!(result.is_ok());
} }
@@ -141,7 +141,7 @@ mod tests {
let mut map = HashMap::new(); let mut map = HashMap::new();
map.insert(WordId(1), 1u32); map.insert(WordId(1), 1u32);
map.insert(WordId(2), 2u32); map.insert(WordId(2), 2u32);
let result = compile_consolidated_module(&words, &map, 16, None, true); let result = compile_consolidated_module(&words, &map, 16);
assert!(result.is_ok()); assert!(result.is_ok());
} }
@@ -163,7 +163,7 @@ mod tests {
let mut map = HashMap::new(); let mut map = HashMap::new();
map.insert(WordId(1), 1u32); map.insert(WordId(1), 1u32);
map.insert(WordId(2), 2u32); map.insert(WordId(2), 2u32);
let result = compile_consolidated_module(&words, &map, 16, None, true); let result = compile_consolidated_module(&words, &map, 16);
assert!(result.is_ok()); assert!(result.is_ok());
} }
} }
+9 -34
View File
@@ -18,8 +18,6 @@ pub mod flags {
pub const IMMEDIATE: u8 = 0x80; pub const IMMEDIATE: u8 = 0x80;
/// Word is hidden (being compiled, not yet findable). /// Word is hidden (being compiled, not yet findable).
pub const HIDDEN: u8 = 0x40; pub const HIDDEN: u8 = 0x40;
/// Word is an implementation detail: findable, but skipped by WORDS.
pub const INTERNAL: u8 = 0x20;
/// Mask for the name length (lower 5 bits). /// Mask for the name length (lower 5 bits).
pub const LENGTH_MASK: u8 = 0x1F; pub const LENGTH_MASK: u8 = 0x1F;
/// Maximum word name length. /// Maximum word name length.
@@ -97,17 +95,11 @@ impl Dictionary {
// Write link field (points to previous LATEST) // Write link field (points to previous LATEST)
self.write_u32_unchecked(entry_start, self.latest); self.write_u32_unchecked(entry_start, self.latest);
// Write flags byte: HIDDEN | length, optionally IMMEDIATE. // Write flags byte: HIDDEN | length, optionally IMMEDIATE
// Underscore-prefixed names are implementation details by repo
// convention (see tools/editor-support): flag them INTERNAL so
// WORDS and completion skip them while FIND still works.
let mut flag_byte = flags::HIDDEN | (name_len as u8 & flags::LENGTH_MASK); let mut flag_byte = flags::HIDDEN | (name_len as u8 & flags::LENGTH_MASK);
if immediate { if immediate {
flag_byte |= flags::IMMEDIATE; flag_byte |= flags::IMMEDIATE;
} }
if name_bytes.first() == Some(&b'_') {
flag_byte |= flags::INTERNAL;
}
self.memory[(entry_start + 4) as usize] = flag_byte; self.memory[(entry_start + 4) as usize] = flag_byte;
// Write name bytes // Write name bytes
@@ -192,9 +184,10 @@ impl Dictionary {
} }
} }
} }
// In no wordlist of the search order: not findable // Fallback: return newest entry across all wordlists
// (Forth 2012 §16.3.3 — the order is authoritative). if let Some(&(_wid, word_addr, fn_index, is_immediate)) = entries.last() {
return None; return Some((word_addr, WordId(fn_index), is_immediate));
}
} }
// Fallback: linked-list walk (for words not yet in the index) // Fallback: linked-list walk (for words not yet in the index)
@@ -417,21 +410,8 @@ impl Dictionary {
} }
/// Return names of all visible (non-hidden) words, newest first. /// Return names of all visible (non-hidden) words, newest first.
/// With `include_internal` false, words flagged INTERNAL are skipped. pub fn visible_words(&self) -> Vec<String> {
pub fn visible_words(&self, include_internal: bool) -> Vec<String> { let mut names = Vec::new();
self.visible_entries()
.into_iter()
.filter(|(_, _, internal)| include_internal || !internal)
.map(|(name, _, _)| name)
.collect()
}
/// All visible (non-hidden) entries, newest first:
/// (name, wordlist id, INTERNAL flag). The wid comes from the hash
/// index (entries themselves store no wid); words missing from the
/// index default to wid 1 (FORTH).
pub fn visible_entries(&self) -> Vec<(String, u32, bool)> {
let mut entries = Vec::new();
let mut addr = self.latest; let mut addr = self.latest;
while addr != 0 { while addr != 0 {
let flags_byte = self.memory[(addr + 4) as usize]; let flags_byte = self.memory[(addr + 4) as usize];
@@ -440,12 +420,7 @@ impl Dictionary {
let name_start = (addr + 5) as usize; let name_start = (addr + 5) as usize;
let name = String::from_utf8_lossy(&self.memory[name_start..name_start + name_len]) let name = String::from_utf8_lossy(&self.memory[name_start..name_start + name_len])
.to_string(); .to_string();
let wid = self names.push(name);
.index
.get(&name)
.and_then(|es| es.iter().find(|e| e.1 == addr))
.map_or(1, |e| e.0);
entries.push((name, wid, flags_byte & flags::INTERNAL != 0));
} }
let link = self.read_u32_unchecked(addr); let link = self.read_u32_unchecked(addr);
if link == addr { if link == addr {
@@ -453,7 +428,7 @@ impl Dictionary {
} }
addr = link; addr = link;
} }
entries names
} }
/// Get a reference to the raw memory buffer. /// Get a reference to the raw memory buffer.
-7
View File
@@ -61,13 +61,6 @@ pub enum WaferError {
#[error("{0}")] #[error("{0}")]
Abort(String), Abort(String),
/// An uncaught Forth THROW as reported to the user. `message` is the
/// full display text (standard message or ABORT" payload); `code`
/// carries the THROW code for typed consumers (CLI exit paths, web
/// REPL styling) via `Error::downcast_ref`.
#[error("{message}")]
UncaughtThrow { code: i32, message: String },
} }
/// Result type alias for WAFER operations. /// Result type alias for WAFER operations.
+2 -9
View File
@@ -120,15 +120,8 @@ pub fn export_module(
metadata_json: metadata_json.as_bytes(), metadata_json: metadata_json.as_bytes(),
}; };
let wasm_bytes = compile_exportable_module( let wasm_bytes = compile_exportable_module(&words, &local_fn_map, table_size, &export_sections)
&words, .map_err(|e| anyhow::anyhow!("export codegen error: {e}"))?;
&local_fn_map,
table_size,
&export_sections,
vm.stack_guard_param(),
vm.typed_calls(),
)
.map_err(|e| anyhow::anyhow!("export codegen error: {e}"))?;
Ok((wasm_bytes, metadata)) Ok((wasm_bytes, metadata))
} }
-2
View File
@@ -159,8 +159,6 @@ pub enum IrOp {
Execute, Execute,
/// Push the current data-stack pointer: ( -- addr ) /// Push the current data-stack pointer: ( -- addr )
SpFetch, SpFetch,
/// Push the current return-stack pointer: ( -- addr )
RpFetch,
// -- Float stack manipulation -- // -- Float stack manipulation --
/// Float duplicate: ( F: r -- r r ) /// Float duplicate: ( F: r -- r r )
-2
View File
@@ -24,8 +24,6 @@ pub mod ir;
pub mod memory; pub mod memory;
pub mod optimizer; pub mod optimizer;
pub mod runtime; pub mod runtime;
pub mod see;
pub mod wordhelp;
// Outer interpreter: runtime-agnostic, works with any Runtime impl // Outer interpreter: runtime-agnostic, works with any Runtime impl
#[allow(trivial_numeric_casts, clippy::unnecessary_cast)] #[allow(trivial_numeric_casts, clippy::unnecessary_cast)]
-22
View File
@@ -106,25 +106,6 @@ pub const SYSVAR_NUM_TIB: u32 = SYSVAR_BASE + 24;
pub const SYSVAR_HLD: u32 = SYSVAR_BASE + 28; pub const SYSVAR_HLD: u32 = SYSVAR_BASE + 28;
/// LEAVE flag: nonzero when LEAVE has been called inside a DO loop. /// LEAVE flag: nonzero when LEAVE has been called inside a DO loop.
pub const SYSVAR_LEAVE_FLAG: u32 = SYSVAR_BASE + 32; pub const SYSVAR_LEAVE_FLAG: u32 = SYSVAR_BASE + 32;
/// Throw code left by a compiled stack-guard fault for `_STACK_FAULT_`.
pub const SYSVAR_FAULT_CODE: u32 = SYSVAR_BASE + 36;
/// DPL: digits right of the rightmost punctuation in the last converted
/// number; negative when the token carried no punctuation.
pub const SYSVAR_DPL: u32 = SYSVAR_BASE + 40;
/// NH: high-order cell of the last single-cell conversion, so an
/// out-of-range token can be recovered as a double.
pub const SYSVAR_NH: u32 = SYSVAR_BASE + 44;
/// Seed for [`SYSVAR_DPL`] before conversion starts.
///
/// `SwiftForth` seeds DPL with a negative value and bumps it once per digit,
/// so an unpunctuated token still ends up negative. Punctuation resets the
/// counter to zero, which makes the final value the digit count right of the
/// rightmost punctuation character.
///
/// The exact seed is observable: `sf64` reports DPL as -1020 after `1234`
/// and -1023 after `-1`, both of which pin it to -1024.
pub const DPL_INIT: i32 = -1024;
#[cfg(test)] #[cfg(test)]
mod tests { mod tests {
@@ -166,9 +147,6 @@ mod tests {
SYSVAR_NUM_TIB, SYSVAR_NUM_TIB,
SYSVAR_HLD, SYSVAR_HLD,
SYSVAR_LEAVE_FLAG, SYSVAR_LEAVE_FLAG,
SYSVAR_FAULT_CODE,
SYSVAR_DPL,
SYSVAR_NH,
]; ];
for offset in all_offsets { for offset in all_offsets {
assert!(offset + CELL_SIZE <= SYSVAR_BASE + SYSVAR_SIZE); assert!(offset + CELL_SIZE <= SYSVAR_BASE + SYSVAR_SIZE);
+6 -336
View File
@@ -27,9 +27,6 @@ pub struct OptConfig {
pub dce: bool, pub dce: bool,
/// Enable inlining of small word bodies. /// Enable inlining of small word bodies.
pub inline: bool, pub inline: bool,
/// Expand a recursive word's base-case guard into its own call sites, so
/// the leaves of the recursion cost a test instead of a call.
pub self_guard: bool,
} }
/// Run all enabled optimization passes. /// Run all enabled optimization passes.
@@ -37,7 +34,6 @@ pub fn optimize(
ops: Vec<IrOp>, ops: Vec<IrOp>,
config: &OptConfig, config: &OptConfig,
bodies: &HashMap<WordId, Vec<IrOp>>, bodies: &HashMap<WordId, Vec<IrOp>>,
self_id: Option<WordId>,
) -> Vec<IrOp> { ) -> Vec<IrOp> {
let mut ir = ops; let mut ir = ops;
@@ -57,17 +53,7 @@ pub fn optimize(
// Phase 2: inline then simplify again // Phase 2: inline then simplify again
if config.inline { if config.inline {
// A caller that can never leave the memory data stack would drag an ir = inline(ir, bodies, 8);
// inlined loop down with it, so leave those callees where they are:
// as their own word the loop keeps its registers, and one call is far
// cheaper than a loop's worth of memory traffic.
let keep_loops_out = !crate::codegen::promotable_modulo_calls(&ir);
ir = inline(ir, bodies, 8, keep_loops_out);
}
if config.self_guard
&& let Some(id) = self_id
{
ir = expand_self_guard(ir, id);
} }
if config.peephole { if config.peephole {
ir = peephole(ir); ir = peephole(ir);
@@ -510,12 +496,7 @@ fn dce(ops: Vec<IrOp>) -> Vec<IrOp> {
/// Inline small word bodies: replaces `Call(id)` with the word's IR body /// Inline small word bodies: replaces `Call(id)` with the word's IR body
/// if the body is small enough and not recursive. /// if the body is small enough and not recursive.
fn inline( fn inline(ops: Vec<IrOp>, bodies: &HashMap<WordId, Vec<IrOp>>, max_size: usize) -> Vec<IrOp> {
ops: Vec<IrOp>,
bodies: &HashMap<WordId, Vec<IrOp>>,
max_size: usize,
keep_loops_out: bool,
) -> Vec<IrOp> {
let mut out = Vec::new(); let mut out = Vec::new();
for op in ops { for op in ops {
match &op { match &op {
@@ -524,7 +505,6 @@ fn inline(
&& body.len() <= max_size && body.len() <= max_size
&& !contains_call_to(body, *id) && !contains_call_to(body, *id)
&& !contains_exit(body) && !contains_exit(body)
&& !(keep_loops_out && crate::codegen::contains_loop(body))
{ {
// Inline the body, recursively converting TailCall back to Call // Inline the body, recursively converting TailCall back to Call
// (tail position in the callee is not tail position in the caller). // (tail position in the callee is not tail position in the caller).
@@ -537,7 +517,7 @@ fn inline(
} }
_ => { _ => {
out.push(apply_to_bodies(op, &|inner| { out.push(apply_to_bodies(op, &|inner| {
inline(inner, bodies, max_size, keep_loops_out) inline(inner, bodies, max_size)
})); }));
} }
} }
@@ -594,142 +574,6 @@ fn detailcall(op: IrOp) -> IrOp {
} }
/// Check if an IR body contains a direct call to the given word (recursion guard). /// Check if an IR body contains a direct call to the given word (recursion guard).
/// Largest guard the expander is willing to run twice, in IR operations.
const MAX_GUARD_OPS: usize = 6;
/// Most self-call sites worth expanding, to bound the code growth.
const MAX_GUARD_SITES: usize = 4;
/// Expand a recursive word's base-case guard into its own call sites.
///
/// A recursive Forth word almost always opens with a guard that returns early
/// -- `: FIB DUP 2 < IF EXIT THEN ... RECURSE ... ;` -- so every leaf of the
/// recursion costs a call whose whole body is that test. Testing at the call
/// site instead removes the call for the leaves, which in fib's tree is half
/// of all nodes.
///
/// `Call(self)` becomes `<guard> IF <what the guard returns> ELSE Call(self)
/// THEN`, which computes the same thing: the callee would have run the guard,
/// taken the branch and returned. The price is that the guard runs twice along
/// the recursive path, which is why it has to be small and free of effects.
fn expand_self_guard(ops: Vec<IrOp>, self_id: WordId) -> Vec<IrOp> {
let Some((cond, base)) = split_guard(&ops) else {
return ops;
};
if count_self_calls(&ops, self_id) > MAX_GUARD_SITES {
return ops;
}
let (cond, base) = (cond.to_vec(), base.to_vec());
replace_self_calls(ops, self_id, &cond, &base)
}
/// Split a body into the condition of its leading base-case guard and what
/// that guard leaves behind, or `None` if it does not open with one.
fn split_guard(ops: &[IrOp]) -> Option<(&[IrOp], &[IrOp])> {
let at = ops.iter().position(|op| matches!(op, IrOp::If { .. }))?;
let cond = &ops[..at];
if at > MAX_GUARD_OPS || !cond.iter().all(is_duplicable) {
return None;
}
let IrOp::If {
then_body,
else_body: None,
} = &ops[at]
else {
return None;
};
// The guard is only a guard if it returns; what precedes the `EXIT` is
// the value it returns, and has to be as harmless as the condition.
let (IrOp::Exit, base) = then_body.split_last()? else {
return None;
};
if base.len() > MAX_GUARD_OPS || !base.iter().all(is_duplicable) {
return None;
}
Some((cond, base))
}
/// Can this operation be duplicated at every call site -- cheap, effect-free,
/// and not itself a call or a branch?
fn is_duplicable(op: &IrOp) -> bool {
matches!(
op,
IrOp::PushI32(_)
| IrOp::Drop
| IrOp::Dup
| IrOp::Swap
| IrOp::Over
| IrOp::Rot
| IrOp::Nip
| IrOp::Tuck
| IrOp::TwoDup
| IrOp::TwoDrop
| IrOp::Add
| IrOp::Sub
| IrOp::Mul
| IrOp::Negate
| IrOp::Abs
| IrOp::Eq
| IrOp::NotEq
| IrOp::Lt
| IrOp::Gt
| IrOp::LtUnsigned
| IrOp::ZeroEq
| IrOp::ZeroLt
| IrOp::And
| IrOp::Or
| IrOp::Xor
| IrOp::Invert
| IrOp::Lshift
| IrOp::Rshift
| IrOp::ArithRshift
)
}
fn count_self_calls(ops: &[IrOp], self_id: WordId) -> usize {
ops.iter()
.map(|op| match op {
IrOp::Call(id) if *id == self_id => 1,
IrOp::If {
then_body,
else_body,
} => {
count_self_calls(then_body, self_id)
+ else_body
.as_deref()
.map_or(0, |eb| count_self_calls(eb, self_id))
}
_ => 0,
})
.sum()
}
/// Wrap every `Call(self_id)` in the guard. Only plain calls: a `TailCall` is
/// followed by a return, and leaving those alone keeps tail-call detection and
/// this pass from having to agree about what tail position means.
fn replace_self_calls(ops: Vec<IrOp>, self_id: WordId, cond: &[IrOp], base: &[IrOp]) -> Vec<IrOp> {
let mut out = Vec::with_capacity(ops.len());
for op in ops {
match op {
IrOp::Call(id) if id == self_id => {
out.extend_from_slice(cond);
out.push(IrOp::If {
then_body: base.to_vec(),
else_body: Some(vec![IrOp::Call(id)]),
});
}
IrOp::If {
then_body,
else_body,
} => out.push(IrOp::If {
then_body: replace_self_calls(then_body, self_id, cond, base),
else_body: else_body.map(|eb| replace_self_calls(eb, self_id, cond, base)),
}),
other => out.push(other),
}
}
out
}
fn contains_call_to(ops: &[IrOp], target: WordId) -> bool { fn contains_call_to(ops: &[IrOp], target: WordId) -> bool {
for op in ops { for op in ops {
match op { match op {
@@ -891,130 +735,8 @@ mod tests {
strength_reduce: true, strength_reduce: true,
dce: true, dce: true,
inline: false, inline: false,
self_guard: false,
}; };
optimize(ops, &config, &HashMap::new(), None) optimize(ops, &config, &HashMap::new())
}
/// A body shaped like a recursive Forth word: a base-case guard, then the
/// recursive step. `SELF` is the word being compiled.
const SELF: WordId = WordId(9);
fn guarded_body(step: Vec<IrOp>) -> Vec<IrOp> {
let mut ops = vec![
IrOp::Dup,
IrOp::PushI32(2),
IrOp::Lt,
IrOp::If {
then_body: vec![IrOp::Exit],
else_body: None,
},
];
ops.extend(step);
ops
}
#[test]
fn self_guard_moves_the_base_case_to_the_call_site() {
let out = expand_self_guard(guarded_body(vec![IrOp::Call(SELF)]), SELF);
assert_eq!(
out,
guarded_body(vec![
IrOp::Dup,
IrOp::PushI32(2),
IrOp::Lt,
IrOp::If {
then_body: vec![],
else_body: Some(vec![IrOp::Call(SELF)]),
},
])
);
}
#[test]
fn self_guard_carries_the_value_the_guard_returns() {
// `: F DUP 2 < IF DROP 0 EXIT THEN RECURSE ;` -- the base case is not
// "leave the argument", it is "replace it with 0".
let body = vec![
IrOp::Dup,
IrOp::PushI32(2),
IrOp::Lt,
IrOp::If {
then_body: vec![IrOp::Drop, IrOp::PushI32(0), IrOp::Exit],
else_body: None,
},
IrOp::Call(SELF),
];
let out = expand_self_guard(body, SELF);
let IrOp::If { then_body, .. } = &out[7] else {
panic!("expected the expanded guard at index 7, got {:?}", out);
};
assert_eq!(then_body, &vec![IrOp::Drop, IrOp::PushI32(0)]);
}
#[test]
fn self_guard_leaves_a_body_without_a_guard_alone() {
// An `IF` with an `ELSE` is a branch, not an early return.
let body = vec![
IrOp::Dup,
IrOp::If {
then_body: vec![IrOp::Drop],
else_body: Some(vec![IrOp::Call(SELF)]),
},
];
assert_eq!(expand_self_guard(body.clone(), SELF), body);
// No `EXIT` in the then-branch: also not a guard.
let body = guarded_body(vec![IrOp::Call(SELF)])
.into_iter()
.map(|op| match op {
IrOp::If { .. } => IrOp::If {
then_body: vec![IrOp::Drop],
else_body: None,
},
other => other,
})
.collect::<Vec<_>>();
assert_eq!(expand_self_guard(body.clone(), SELF), body);
}
#[test]
fn self_guard_refuses_a_condition_it_cannot_run_twice() {
// A guard reached through a call or a memory write would be evaluated
// once at the call site and again inside the callee.
let body = vec![
IrOp::Call(WordId(3)),
IrOp::If {
then_body: vec![IrOp::Exit],
else_body: None,
},
IrOp::Call(SELF),
];
assert_eq!(expand_self_guard(body.clone(), SELF), body);
let body = vec![
IrOp::Dup,
IrOp::Fetch,
IrOp::If {
then_body: vec![IrOp::Exit],
else_body: None,
},
IrOp::Call(SELF),
];
assert_eq!(expand_self_guard(body.clone(), SELF), body);
}
#[test]
fn self_guard_stops_at_the_call_site_budget() {
let step = std::iter::repeat_n(IrOp::Call(SELF), MAX_GUARD_SITES + 1).collect();
let body = guarded_body(step);
assert_eq!(expand_self_guard(body.clone(), SELF), body);
}
#[test]
fn self_guard_leaves_tail_calls_alone() {
let body = guarded_body(vec![IrOp::TailCall(SELF)]);
assert_eq!(expand_self_guard(body.clone(), SELF), body);
} }
fn opt_with_inline(ops: Vec<IrOp>, bodies: &HashMap<WordId, Vec<IrOp>>) -> Vec<IrOp> { fn opt_with_inline(ops: Vec<IrOp>, bodies: &HashMap<WordId, Vec<IrOp>>) -> Vec<IrOp> {
@@ -1025,9 +747,8 @@ mod tests {
strength_reduce: true, strength_reduce: true,
dce: true, dce: true,
inline: true, inline: true,
self_guard: false,
}; };
optimize(ops, &config, bodies, None) optimize(ops, &config, bodies)
} }
// Peephole tests // Peephole tests
@@ -1287,59 +1008,8 @@ mod tests {
strength_reduce: false, strength_reduce: false,
dce: false, dce: false,
inline: true, inline: true,
self_guard: false,
}; };
let result = optimize(vec![IrOp::Call(WordId(5))], &config, &bodies, None); let result = optimize(vec![IrOp::Call(WordId(5))], &config, &bodies);
assert_eq!(result, vec![IrOp::Call(WordId(5))]); assert_eq!(result, vec![IrOp::Call(WordId(5))]);
} }
#[test]
fn keeps_a_loop_out_of_a_caller_stuck_on_the_memory_stack() {
// The caller has a `.`, so it can never leave the memory data stack.
// Inlining the loop would drag it down too; as its own word the loop
// keeps its registers and the caller just pays one call.
let mut bodies = HashMap::new();
bodies.insert(
WordId(5),
vec![IrOp::DoLoop {
body: vec![IrOp::PushI32(1), IrOp::Add],
is_plus_loop: false,
}],
);
let result = opt_with_inline(vec![IrOp::Call(WordId(5)), IrOp::Dot], &bodies);
assert!(
matches!(result.first(), Some(IrOp::Call(WordId(5)))),
"loop should not have been inlined, got {result:?}"
);
}
#[test]
fn still_inlines_a_loop_into_a_caller_that_can_be_promoted() {
let mut bodies = HashMap::new();
bodies.insert(
WordId(5),
vec![IrOp::DoLoop {
body: vec![IrOp::PushI32(1), IrOp::Add],
is_plus_loop: false,
}],
);
let result = opt_with_inline(vec![IrOp::Call(WordId(5)), IrOp::Dup], &bodies);
assert!(
!result.iter().any(|op| matches!(op, IrOp::Call(_))),
"loop should have been inlined, got {result:?}"
);
}
#[test]
fn still_inlines_straight_line_words_anywhere() {
// Only loops are held back; a small straight-line word is still
// better off inlined even into an unpromotable caller.
let mut bodies = HashMap::new();
bodies.insert(WordId(5), vec![IrOp::Dup, IrOp::Mul]);
let result = opt_with_inline(vec![IrOp::Call(WordId(5)), IrOp::Dot], &bodies);
assert!(
!result.iter().any(|op| matches!(op, IrOp::Call(_))),
"straight-line word should still inline, got {result:?}"
);
}
} }
+397 -2558
View File
File diff suppressed because it is too large Load Diff
+2 -21
View File
@@ -98,29 +98,11 @@ impl HostAccess for CallerHostAccess<'_, '_> {
let func = *func_ref let func = *func_ref
.unwrap_func() .unwrap_func()
.ok_or_else(|| anyhow::anyhow!("call_func: null funcref {fn_index}"))?; .ok_or_else(|| anyhow::anyhow!("call_func: null funcref {fn_index}"))?;
func.call(&mut *self.caller, &[], &mut []) func.call(&mut *self.caller, &[], &mut [])?;
.map_err(name_trap_frame)?;
Ok(()) Ok(())
} }
} }
/// Prefix a wasmtime trap error with the innermost named WASM frame.
/// Compiled words carry their Forth name in the module name section, so a
/// genuine trap reads "in <WORD>: wasm trap: ...". THROW-driven unwinds
/// also pass through here, but CATCH and `describe_uncaught` key on the
/// shared `throw_code` cell, never on the message, so the wrap is inert
/// for them.
fn name_trap_frame(e: wasmtime::Error) -> wasmtime::Error {
let name = e
.downcast_ref::<wasmtime::WasmBacktrace>()
.and_then(|bt| bt.frames().iter().find_map(|f| f.func_name()))
.map(str::to_string);
match name {
Some(n) => e.context(format!("in {n}")),
None => e,
}
}
/// Wasmtime-based native runtime. /// Wasmtime-based native runtime.
pub struct NativeRuntime { pub struct NativeRuntime {
engine: Engine, engine: Engine,
@@ -311,8 +293,7 @@ impl Runtime for NativeRuntime {
let func = *r let func = *r
.unwrap_func() .unwrap_func()
.ok_or_else(|| anyhow::anyhow!("word {fn_index} is null funcref"))?; .ok_or_else(|| anyhow::anyhow!("word {fn_index} is null funcref"))?;
func.call(&mut self.store, &[], &mut []) func.call(&mut self.store, &[], &mut [])?;
.map_err(name_trap_frame)?;
Ok(()) Ok(())
} }
-376
View File
@@ -1,376 +0,0 @@
//! IR pretty-printer for `SEE-IR` and the `SEE` fallback path.
//!
//! Renders a post-optimization IR body as indented, one-op-per-line text.
//! Simple ops print as short lowercase mnemonics (Forth glyphs where they
//! are universally recognizable: `@`, `!`, `0=`, `>r`, ...); structured ops
//! print as Forth control words with 2-space indented bodies. Calls resolve
//! `WordId`s to names through an optional resolver so the formatter itself
//! stays independent of the VM.
use crate::dictionary::WordId;
use crate::ir::IrOp;
/// Format an IR body as indented, one-op-per-line text.
pub fn format_ir(ops: &[IrOp]) -> String {
format_ir_with(ops, &|_| None)
}
/// Like [`format_ir`], resolving `Call`/`TailCall`/`Execute` targets to word
/// names via `resolve`; unresolved ids print as `#N`.
pub fn format_ir_with(ops: &[IrOp], resolve: &dyn Fn(WordId) -> Option<String>) -> String {
let mut out = String::new();
write_ops(&mut out, ops, 0, resolve);
out
}
fn line(out: &mut String, depth: usize, text: &str) {
for _ in 0..depth {
out.push_str(" ");
}
out.push_str(text);
out.push('\n');
}
fn callee(id: WordId, resolve: &dyn Fn(WordId) -> Option<String>) -> String {
resolve(id).unwrap_or_else(|| format!("#{}", id.0))
}
fn write_ops(
out: &mut String,
ops: &[IrOp],
depth: usize,
resolve: &dyn Fn(WordId) -> Option<String>,
) {
for op in ops {
write_op(out, op, depth, resolve);
}
}
fn write_op(out: &mut String, op: &IrOp, depth: usize, resolve: &dyn Fn(WordId) -> Option<String>) {
// Exhaustive on purpose: a new IrOp variant must show up here at
// compile time, not silently render wrong.
let simple: String = match op {
// -- Literals --
IrOp::PushI32(v) => format!("push {v}"),
IrOp::PushI64(v) => format!("push64 {v}"),
IrOp::PushF64(v) => format!("fpush {v}"),
// -- Stack manipulation --
IrOp::Drop => "drop".into(),
IrOp::Dup => "dup".into(),
IrOp::Swap => "swap".into(),
IrOp::Over => "over".into(),
IrOp::Rot => "rot".into(),
IrOp::Nip => "nip".into(),
IrOp::Tuck => "tuck".into(),
IrOp::TwoDup => "2dup".into(),
IrOp::TwoDrop => "2drop".into(),
// -- Arithmetic --
IrOp::Add => "add".into(),
IrOp::Sub => "sub".into(),
IrOp::Mul => "mul".into(),
IrOp::DivMod => "divmod".into(),
IrOp::Negate => "negate".into(),
IrOp::Abs => "abs".into(),
// -- Comparison --
IrOp::Eq => "eq".into(),
IrOp::NotEq => "ne".into(),
IrOp::Lt => "lt".into(),
IrOp::Gt => "gt".into(),
IrOp::LtUnsigned => "u<".into(),
IrOp::ZeroEq => "0=".into(),
IrOp::ZeroLt => "0<".into(),
// -- Logic --
IrOp::And => "and".into(),
IrOp::Or => "or".into(),
IrOp::Xor => "xor".into(),
IrOp::Invert => "invert".into(),
IrOp::Lshift => "lshift".into(),
IrOp::Rshift => "rshift".into(),
IrOp::ArithRshift => "arshift".into(),
// -- Memory --
IrOp::Fetch => "@".into(),
IrOp::Store => "!".into(),
IrOp::CFetch => "c@".into(),
IrOp::CStore => "c!".into(),
IrOp::PlusStore => "+!".into(),
// -- Calls --
IrOp::Call(id) => format!("call {}", callee(*id, resolve)),
IrOp::TailCall(id) => format!("tail-call {}", callee(*id, resolve)),
// -- Structured control flow (multi-line) --
IrOp::If {
then_body,
else_body,
} => {
line(out, depth, "if");
write_ops(out, then_body, depth + 1, resolve);
if let Some(eb) = else_body {
line(out, depth, "else");
write_ops(out, eb, depth + 1, resolve);
}
line(out, depth, "then");
return;
}
IrOp::DoLoop { body, is_plus_loop } => {
line(out, depth, "do");
write_ops(out, body, depth + 1, resolve);
line(out, depth, if *is_plus_loop { "+loop" } else { "loop" });
return;
}
IrOp::BeginUntil { body } => {
line(out, depth, "begin");
write_ops(out, body, depth + 1, resolve);
line(out, depth, "until");
return;
}
IrOp::BeginAgain { body } => {
line(out, depth, "begin");
write_ops(out, body, depth + 1, resolve);
line(out, depth, "again");
return;
}
IrOp::BeginWhileRepeat { test, body } => {
line(out, depth, "begin");
write_ops(out, test, depth + 1, resolve);
line(out, depth, "while");
write_ops(out, body, depth + 1, resolve);
line(out, depth, "repeat");
return;
}
IrOp::BeginDoubleWhileRepeat {
outer_test,
inner_test,
body,
after_repeat,
else_body,
} => {
line(out, depth, "begin");
write_ops(out, outer_test, depth + 1, resolve);
line(out, depth, "while");
write_ops(out, inner_test, depth + 1, resolve);
line(out, depth, "while");
write_ops(out, body, depth + 1, resolve);
line(out, depth, "repeat");
write_ops(out, after_repeat, depth + 1, resolve);
if let Some(eb) = else_body {
line(out, depth, "else");
write_ops(out, eb, depth + 1, resolve);
}
line(out, depth, "then");
return;
}
IrOp::Exit => "exit".into(),
IrOp::LoopRestartIfFalse => "loop-restart-if-false".into(),
// -- Flat forward branches --
IrOp::Block(l) => format!("block L{l}"),
IrOp::BranchIfFalse(l) => format!("branch-if-false L{l}"),
IrOp::EndBlock(l) => format!("end-block L{l}"),
// -- Return stack --
IrOp::ToR => ">r".into(),
IrOp::FromR => "r>".into(),
IrOp::RFetch => "r@".into(),
IrOp::LoopJ => "j".into(),
// -- Forth locals --
IrOp::ForthLocalGet(n) => format!("local@ {n}"),
IrOp::ForthLocalSet(n) => format!("local! {n}"),
IrOp::ForthFLocalGet(n) => format!("flocal@ {n}"),
IrOp::ForthFLocalSet(n) => format!("flocal! {n}"),
// -- I/O --
IrOp::Emit => "emit".into(),
IrOp::Dot => ".".into(),
IrOp::Cr => "cr".into(),
IrOp::Type => "type".into(),
// -- System --
IrOp::Execute => "execute".into(),
IrOp::SpFetch => "sp@".into(),
IrOp::RpFetch => "rp@".into(),
// -- Float stack --
IrOp::FDup => "fdup".into(),
IrOp::FDrop => "fdrop".into(),
IrOp::FSwap => "fswap".into(),
IrOp::FOver => "fover".into(),
// -- Float arithmetic --
IrOp::FAdd => "fadd".into(),
IrOp::FSub => "fsub".into(),
IrOp::FMul => "fmul".into(),
IrOp::FDiv => "fdiv".into(),
IrOp::FNegate => "fnegate".into(),
IrOp::FAbs => "fabs".into(),
IrOp::FSqrt => "fsqrt".into(),
IrOp::FMin => "fmin".into(),
IrOp::FMax => "fmax".into(),
IrOp::FFloor => "ffloor".into(),
IrOp::FRound => "fround".into(),
// -- Float comparisons --
IrOp::FZeroEq => "f0=".into(),
IrOp::FZeroLt => "f0<".into(),
IrOp::FEq => "f=".into(),
IrOp::FLt => "f<".into(),
// -- Float memory --
IrOp::FetchFloat => "f@".into(),
IrOp::StoreFloat => "f!".into(),
// -- Conversions --
IrOp::StoF => "s>f".into(),
IrOp::FtoS => "f>s".into(),
};
line(out, depth, &simple);
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn simple_ops_one_per_line() {
let out = format_ir(&[IrOp::Dup, IrOp::Mul, IrOp::PushI32(7)]);
assert_eq!(out, "dup\nmul\npush 7\n");
}
#[test]
fn call_resolves_via_resolver() {
let ops = [IrOp::Call(WordId(12)), IrOp::TailCall(WordId(13))];
assert_eq!(format_ir(&ops), "call #12\ntail-call #13\n");
let named = format_ir_with(&ops, &|id| (id.0 == 12).then(|| "SQ".to_string()));
assert_eq!(named, "call SQ\ntail-call #13\n");
}
#[test]
fn nested_if_inside_do_loop_indents() {
let ops = [IrOp::DoLoop {
body: vec![
IrOp::Dup,
IrOp::If {
then_body: vec![IrOp::Dup, IrOp::Mul],
else_body: Some(vec![IrOp::Drop]),
},
],
is_plus_loop: false,
}];
let expected = "do\n dup\n if\n dup\n mul\n else\n drop\n then\nloop\n";
assert_eq!(format_ir(&ops), expected);
}
#[test]
fn while_loops_and_flat_branches() {
let ops = [
IrOp::BeginWhileRepeat {
test: vec![IrOp::Dup],
body: vec![IrOp::PushI32(1), IrOp::Sub],
},
IrOp::Block(3),
IrOp::BranchIfFalse(3),
IrOp::EndBlock(3),
];
let expected = "begin\n dup\nwhile\n push 1\n sub\nrepeat\nblock L3\nbranch-if-false L3\nend-block L3\n";
assert_eq!(format_ir(&ops), expected);
}
#[test]
fn every_simple_variant_renders() {
// One of each non-structured op; count of output lines must match.
let ops = vec![
IrOp::PushI32(1),
IrOp::PushI64(2),
IrOp::PushF64(1.5),
IrOp::Drop,
IrOp::Dup,
IrOp::Swap,
IrOp::Over,
IrOp::Rot,
IrOp::Nip,
IrOp::Tuck,
IrOp::TwoDup,
IrOp::TwoDrop,
IrOp::Add,
IrOp::Sub,
IrOp::Mul,
IrOp::DivMod,
IrOp::Negate,
IrOp::Abs,
IrOp::Eq,
IrOp::NotEq,
IrOp::Lt,
IrOp::Gt,
IrOp::LtUnsigned,
IrOp::ZeroEq,
IrOp::ZeroLt,
IrOp::And,
IrOp::Or,
IrOp::Xor,
IrOp::Invert,
IrOp::Lshift,
IrOp::Rshift,
IrOp::ArithRshift,
IrOp::Fetch,
IrOp::Store,
IrOp::CFetch,
IrOp::CStore,
IrOp::PlusStore,
IrOp::Call(WordId(1)),
IrOp::TailCall(WordId(2)),
IrOp::Exit,
IrOp::LoopRestartIfFalse,
IrOp::Block(1),
IrOp::BranchIfFalse(1),
IrOp::EndBlock(1),
IrOp::ToR,
IrOp::FromR,
IrOp::RFetch,
IrOp::LoopJ,
IrOp::ForthLocalGet(0),
IrOp::ForthLocalSet(0),
IrOp::ForthFLocalGet(0),
IrOp::ForthFLocalSet(0),
IrOp::Emit,
IrOp::Dot,
IrOp::Cr,
IrOp::Type,
IrOp::Execute,
IrOp::SpFetch,
IrOp::RpFetch,
IrOp::FDup,
IrOp::FDrop,
IrOp::FSwap,
IrOp::FOver,
IrOp::FAdd,
IrOp::FSub,
IrOp::FMul,
IrOp::FDiv,
IrOp::FNegate,
IrOp::FAbs,
IrOp::FSqrt,
IrOp::FMin,
IrOp::FMax,
IrOp::FFloor,
IrOp::FRound,
IrOp::FZeroEq,
IrOp::FZeroLt,
IrOp::FEq,
IrOp::FLt,
IrOp::FetchFloat,
IrOp::StoreFloat,
IrOp::StoF,
IrOp::FtoS,
];
let out = format_ir(&ops);
assert_eq!(out.lines().count(), ops.len());
// Every line non-empty, no accidental blank rendering.
assert!(out.lines().all(|l| !l.trim().is_empty()));
}
}
File diff suppressed because it is too large Load Diff
+67 -214
View File
@@ -1,10 +1,8 @@
#![allow(dead_code)] #![allow(dead_code)]
//! Cross-engine comparison tests: WAFER vs gforth (and `SwiftForth` for perf). //! Cross-engine comparison tests: WAFER vs gforth.
//! //!
//! Validates that WAFER produces identical output to gforth for standard //! Validates that WAFER produces identical output to gforth for standard
//! Forth programs, and benchmarks performance of the engines. `SwiftForth` //! Forth programs, and benchmarks performance of both engines.
//! (`sf64`, native-code commercial compiler) joins the performance report
//! as an upper-bound reference when installed.
//! //!
//! WAFER-only correctness: `cargo test -p wafer-core --test comparison` //! WAFER-only correctness: `cargo test -p wafer-core --test comparison`
//! Full comparison + perf: `cargo test -p wafer-core --test comparison -- --nocapture --ignored` //! Full comparison + perf: `cargo test -p wafer-core --test comparison -- --nocapture --ignored`
@@ -65,48 +63,6 @@ fn find_gforth_fast() -> Option<&'static str> {
.as_deref() .as_deref()
} }
// -----------------------------------------------------------------------
// SwiftForth (sf64) discovery (cached)
// -----------------------------------------------------------------------
static SF64_PATH: OnceLock<Option<String>> = OnceLock::new();
/// Probe sf64 by piping `bye` via stdin — sf64 has no `-e` flag; it takes
/// Forth source from stdin or as bare command-line arguments.
fn probe_sf64(candidate: &str) -> bool {
run_via_stdin(candidate, "bye\n").is_some_and(|o| o.status.success())
}
fn find_sf64() -> Option<&'static str> {
SF64_PATH
.get_or_init(|| {
for candidate in &["/Applications/ForthInc/SwiftForth/bin/macos/sf64", "sf64"] {
if probe_sf64(candidate) {
return Some(candidate.to_string());
}
}
None
})
.as_deref()
}
/// Spawn `binary`, write `input` to its stdin, and collect the output.
fn run_via_stdin(binary: &str, input: &str) -> Option<std::process::Output> {
Command::new(binary)
// Perf lanes measure unguarded code (only the wafer binary reads this)
.env("WAFER_STACK_GUARDS", "0")
.stdin(std::process::Stdio::piped())
.stdout(std::process::Stdio::piped())
.stderr(std::process::Stdio::piped())
.spawn()
.and_then(|mut child| {
use std::io::Write;
child.stdin.take().unwrap().write_all(input.as_bytes())?;
child.wait_with_output()
})
.ok()
}
// ----------------------------------------------------------------------- // -----------------------------------------------------------------------
// Engine runners // Engine runners
// ----------------------------------------------------------------------- // -----------------------------------------------------------------------
@@ -453,26 +409,6 @@ fn programs() -> Vec<Program> {
expected: "99 \n", expected: "99 \n",
category: Category::Definitions, category: Category::Definitions,
}, },
Program {
name: "search-order-hides",
code: "WORDLIST CONSTANT MY-WL\n\
MY-WL SET-CURRENT\n\
: SECRET 42 ;\n\
FORTH-WORDLIST SET-CURRENT\n\
[UNDEFINED] SECRET . CR\n\
GET-ORDER MY-WL SWAP 1+ SET-ORDER\n\
[DEFINED] SECRET . CR\n\
SECRET . CR\n\
-1 SET-ORDER\n\
[UNDEFINED] SECRET . CR",
expected: "-1 \n-1 \n42 \n-1 \n",
category: Category::Definitions,
},
// QUIT is deliberately absent from this corpus: what it abandons is
// "the input source", and each engine here is fed differently (wafer
// line by line, gforth from a file, sf64 from a prompting stdin), so
// a comparison would measure the harness. Its semantics are pinned by
// the QUIT tests in outer.rs, checked by hand against both engines.
// -- Strings -- // -- Strings --
Program { Program {
name: "s-quote-type", name: "s-quote-type",
@@ -644,81 +580,6 @@ fn compare_all_programs() {
); );
} }
// -----------------------------------------------------------------------
// Cross-engine behavioral comparison (requires SwiftForth sf64) -- WS-003
// -----------------------------------------------------------------------
/// Run Forth code through `SwiftForth`. Piped sf64 is quiet (no banner, no
/// `ok` echo), truncates input lines at ~256 chars, and exits 243 after an
/// error, so statements are fed one per line with a final `bye`.
fn run_sf64_code(sf64: &str, code: &str) -> Option<EngineResult> {
let mut input = String::new();
for line in code.lines() {
let t = line.trim();
if !t.is_empty() {
input.push_str(t);
input.push('\n');
}
}
input.push_str("bye\n");
let out = run_via_stdin(sf64, &input)?;
Some(EngineResult {
output: String::from_utf8_lossy(&out.stdout).to_string(),
success: out.status.success(),
})
}
/// Correctness lane against `SwiftForth`: the same program corpus as the
/// gforth comparison, sf64 as the oracle. Skips gracefully when sf64 is
/// not installed (CI/linux). Programs listed in `SF64_SKIP` use words or
/// output conventions `SwiftForth` does not share.
#[test]
#[ignore = "requires SwiftForth sf64 (run with -- --ignored)"]
fn compare_all_programs_sf64() {
// dot-quote: `."` outside a definition is a no-op in SwiftForth
// (compile-only); WAFER supports the interpret-mode extension.
const SF64_SKIP: &[&str] = &["dot-quote"];
let Some(sf64) = find_sf64() else {
eprintln!("SKIP: sf64 not found");
return;
};
let progs = programs();
let mut passed = 0;
let mut skipped = 0;
for prog in &progs {
if SF64_SKIP.contains(&prog.name) {
skipped += 1;
continue;
}
let wafer = run_wafer(prog.code);
assert!(wafer.success, "{}: WAFER execution failed", prog.name);
let Some(sf) = run_sf64_code(sf64, prog.code) else {
skipped += 1;
continue;
};
if !sf.success {
eprintln!(" WARN {}: sf64 execution failed, skipping", prog.name);
skipped += 1;
continue;
}
// SwiftForth prints numbers space-prefixed and echoes piped input
// lines, so byte-exact comparison is meaningless; compare the
// whitespace-token stream (the printed values and strings).
let wafer_tokens: Vec<&str> = wafer.output.split_whitespace().collect();
let sf_tokens: Vec<&str> = sf.output.split_whitespace().collect();
assert_eq!(
wafer_tokens, sf_tokens,
"{}: output differs\n WAFER: {:?}\n sf64: {:?}",
prog.name, wafer.output, sf.output
);
passed += 1;
}
eprintln!(
"\nsf64 behavioral comparison: {passed} passed, {skipped} skipped (of {})",
progs.len()
);
}
// ----------------------------------------------------------------------- // -----------------------------------------------------------------------
// Performance comparison (requires gforth) // Performance comparison (requires gforth)
// ----------------------------------------------------------------------- // -----------------------------------------------------------------------
@@ -746,37 +607,37 @@ fn perf_benchmarks() -> Vec<PerfBenchmark> {
verify: "25 FIB", verify: "25 FIB",
expected: 75025, expected: 75025,
samples: 5, samples: 5,
max_ratio: 0.10, max_ratio: 0.65,
}, },
PerfBenchmark { PerfBenchmark {
name: "Factorial(12)x100K", name: "Factorial(12)x10K",
define: ": FACT 1 SWAP 1+ 1 ?DO I * LOOP ; \ define: ": FACT 1 SWAP 1+ 1 ?DO I * LOOP ; \
: FACT-BENCH 100000 0 DO 12 FACT DROP LOOP ;", : FACT-BENCH 10000 0 DO 12 FACT DROP LOOP ;",
run_code: "FACT-BENCH", run_code: "FACT-BENCH",
verify: "12 FACT", verify: "12 FACT",
expected: 479001600, expected: 479001600,
samples: 5, samples: 5,
max_ratio: 0.12, max_ratio: 0.75,
}, },
PerfBenchmark { PerfBenchmark {
name: "GCD-bench(20K)", name: "GCD-bench(500)",
define: ": GCD BEGIN DUP WHILE TUCK MOD REPEAT DROP ; \ define: ": GCD BEGIN DUP WHILE TUCK MOD REPEAT DROP ; \
: GCD-BENCH 0 DO 10000 I 1+ GCD DROP LOOP ;", : GCD-BENCH 0 DO 10000 I 1+ GCD DROP LOOP ;",
run_code: "20000 GCD-BENCH", run_code: "500 GCD-BENCH",
verify: "48 36 GCD", verify: "48 36 GCD",
expected: 12, expected: 12,
samples: 5, samples: 5,
max_ratio: 0.45, max_ratio: 0.70,
}, },
PerfBenchmark { PerfBenchmark {
name: "NestedLoops(50)x1K", name: "NestedLoops(50)",
define: ": NESTED 0 SWAP 0 DO I 0 ?DO I J + DROP LOOP LOOP ; \ define: ": NESTED 0 SWAP 0 DO I 0 ?DO I J + DROP LOOP LOOP ; \
: NESTED-BENCH 1000 0 DO 50 NESTED DROP LOOP ;", : NESTED-BENCH 100 0 DO 50 NESTED DROP LOOP ;",
run_code: "NESTED-BENCH", run_code: "NESTED-BENCH",
verify: "5 NESTED", verify: "5 NESTED",
expected: 0, expected: 0,
samples: 5, samples: 3,
max_ratio: 0.11, max_ratio: 0.20,
}, },
PerfBenchmark { PerfBenchmark {
name: "Collatz(2K)", name: "Collatz(2K)",
@@ -788,7 +649,7 @@ fn perf_benchmarks() -> Vec<PerfBenchmark> {
verify: "27 COLLATZ", verify: "27 COLLATZ",
expected: 111, expected: 111,
samples: 3, samples: 3,
max_ratio: 0.08, max_ratio: 0.45,
}, },
] ]
} }
@@ -837,11 +698,31 @@ fn measure_wafer_release(wafer: &str, bench: &PerfBenchmark) -> Option<u64> {
define = bench.define, define = bench.define,
run = bench.run_code, run = bench.run_code,
); );
let output = run_via_stdin(wafer, &code)?; let output = Command::new(wafer)
.stdin(std::process::Stdio::piped())
.stdout(std::process::Stdio::piped())
.stderr(std::process::Stdio::piped())
.spawn()
.and_then(|mut child| {
use std::io::Write;
child.stdin.take().unwrap().write_all(code.as_bytes())?;
child.wait_with_output()
})
.ok()?;
if !output.status.success() { if !output.status.success() {
return None; return None;
} }
median_printed_time(&output.stdout) let stdout = String::from_utf8_lossy(&output.stdout);
let mut times: Vec<u64> = stdout
.trim()
.lines()
.filter_map(|l| l.trim().parse::<u64>().ok())
.collect();
times.sort();
if times.is_empty() {
return None;
}
Some(times[times.len() / 2])
} }
/// Measure WAFER execution time after CONSOLIDATE (direct calls between all words). /// Measure WAFER execution time after CONSOLIDATE (direct calls between all words).
@@ -853,17 +734,21 @@ fn measure_wafer_consolidated(wafer: &str, bench: &PerfBenchmark) -> Option<u64>
define = bench.define, define = bench.define,
run = bench.run_code, run = bench.run_code,
); );
let output = run_via_stdin(wafer, &code)?; let output = Command::new(wafer)
.stdin(std::process::Stdio::piped())
.stdout(std::process::Stdio::piped())
.stderr(std::process::Stdio::piped())
.spawn()
.and_then(|mut child| {
use std::io::Write;
child.stdin.take().unwrap().write_all(code.as_bytes())?;
child.wait_with_output()
})
.ok()?;
if !output.status.success() { if !output.status.success() {
return None; return None;
} }
median_printed_time(&output.stdout) let stdout = String::from_utf8_lossy(&output.stdout);
}
/// Parse the microsecond values printed by TIMED-BENCH (one per line) and
/// return the median.
fn median_printed_time(stdout: &[u8]) -> Option<u64> {
let stdout = String::from_utf8_lossy(stdout);
let mut times: Vec<u64> = stdout let mut times: Vec<u64> = stdout
.trim() .trim()
.lines() .lines()
@@ -893,28 +778,18 @@ fn measure_gforth(gforth: &str, bench: &PerfBenchmark) -> Option<u64> {
if !output.status.success() { if !output.status.success() {
return None; return None;
} }
median_printed_time(&output.stdout) let stdout = String::from_utf8_lossy(&output.stdout);
} // Parse the 3 timing values and take the median
let mut times: Vec<u64> = stdout
/// Measure `SwiftForth` (`sf64`) execution time using Forth-level `ucounter` .trim()
/// (double-cell microsecond counter; `2swap d- drop` yields elapsed us — .lines()
/// the same wrapper shape as gforth's `utime`). Timing excludes startup. .filter_map(|l| l.trim().parse::<u64>().ok())
/// sf64 has no `-e` flag, so the program is piped via stdin — one statement .collect();
/// per line, because sf64 truncates input lines at ~256 chars. times.sort();
/// Returns microseconds, or None if sf64 is unavailable or fails. if times.is_empty() {
fn measure_sf64(sf64: &str, bench: &PerfBenchmark) -> Option<u64> {
let code = format!(
"{define}\n{run}\n\
: TIMED-BENCH ucounter {run} ucounter 2swap d- drop . cr ;\n\
TIMED-BENCH\nTIMED-BENCH\nTIMED-BENCH\nbye\n",
define = bench.define,
run = bench.run_code,
);
let output = run_via_stdin(sf64, &code)?;
if !output.status.success() {
return None; return None;
} }
median_printed_time(&output.stdout) Some(times[times.len() / 2])
} }
#[test] #[test]
@@ -955,31 +830,18 @@ fn performance_report() {
); );
} }
let sf64 = find_sf64(); let sep = "=".repeat(80);
if sf64.is_none() { let thin = "-".repeat(80);
eprintln!("NOTE: sf64 (SwiftForth) not found — column skipped");
}
let sep = "=".repeat(100);
let thin = "-".repeat(100);
println!("\n{sep}"); println!("\n{sep}");
println!(" WAFER vs Gforth vs SwiftForth Performance Comparison (release mode)"); println!(" WAFER vs Gforth Performance Comparison (release mode)");
println!("{sep}\n"); println!("{sep}\n");
println!( println!(
"{:<22} {:>9} {:>9} {:>9} {:>9} {:>9} {:>9} {:>9} {:>9}", "{:<22} {:>10} {:>10} {:>10} {:>10} {:>10} {:>10}",
"Benchmark", "Benchmark", "WAFER", "CONSOL", "gforth", "gf-fast", "WAFER/gf", "limit"
"WAFER",
"CONSOL",
"gforth",
"gf-fast",
"sf64",
"WAFER/gf",
"WAFER/sf",
"limit"
); );
println!( println!(
"{:<22} {:>9} {:>9} {:>9} {:>9} {:>9} {:>9} {:>9} {:>9}", "{:<22} {:>10} {:>10} {:>10} {:>10} {:>10} {:>10}",
"", "(us)", "(us)", "(us)", "(us)", "(us)", "", "", "" "", "(us)", "(us)", "(us)", "(us)", "", ""
); );
println!("{thin}"); println!("{thin}");
@@ -994,11 +856,9 @@ fn performance_report() {
.unwrap_or(0); .unwrap_or(0);
let gf = gforth.and_then(|g| measure_gforth(g, bench)); let gf = gforth.and_then(|g| measure_gforth(g, bench));
let gf_fast = gforth_fast.and_then(|g| measure_gforth(g, bench)); let gf_fast = gforth_fast.and_then(|g| measure_gforth(g, bench));
let sf = sf64.and_then(|s| measure_sf64(s, bench));
let gf_str = gf.map_or_else(|| "-".to_string(), |v| format!("{v}")); let gf_str = gf.map_or_else(|| "-".to_string(), |v| format!("{v}"));
let gf_fast_str = gf_fast.map_or_else(|| "-".to_string(), |v| format!("{v}")); let gf_fast_str = gf_fast.map_or_else(|| "-".to_string(), |v| format!("{v}"));
let sf_str = sf.map_or_else(|| "-".to_string(), |v| format!("{v}"));
let best_wafer = if consol > 0 && consol < wafer { let best_wafer = if consol > 0 && consol < wafer {
consol consol
} else { } else {
@@ -1012,15 +872,11 @@ fn performance_report() {
} }
}); });
let ratio = ratio_val.map_or_else(|| "-".to_string(), |r| format!("{r:.2}x")); let ratio = ratio_val.map_or_else(|| "-".to_string(), |r| format!("{r:.2}x"));
let sf_ratio = sf.filter(|&s| s > 0).map_or_else(
|| "-".to_string(),
|s| format!("{:.2}x", best_wafer as f64 / s as f64),
);
let limit_str = format!("{:.2}x", bench.max_ratio); let limit_str = format!("{:.2}x", bench.max_ratio);
println!( println!(
"{:<22} {:>9} {:>9} {:>9} {:>9} {:>9} {:>9} {:>9} {:>9}", "{:<22} {:>10} {:>10} {:>10} {:>10} {:>10} {:>10}",
bench.name, wafer, consol, gf_str, gf_fast_str, sf_str, ratio, sf_ratio, limit_str bench.name, wafer, consol, gf_str, gf_fast_str, ratio, limit_str
); );
// Check regression limits // Check regression limits
@@ -1043,9 +899,6 @@ fn performance_report() {
println!("{thin}"); println!("{thin}");
println!(" WAFER = all optimizations, CONSOL = after CONSOLIDATE"); println!(" WAFER = all optimizations, CONSOL = after CONSOLIDATE");
println!(" WAFER/gf = best(WAFER,CONSOL) vs gforth, < 1.0 means WAFER faster"); println!(" WAFER/gf = best(WAFER,CONSOL) vs gforth, < 1.0 means WAFER faster");
println!(
" WAFER/sf = best(WAFER,CONSOL) vs SwiftForth sf64 (native code; informational, no limit)"
);
println!("{sep}\n"); println!("{sep}\n");
if !regressions.is_empty() { if !regressions.is_empty() {
+1 -35
View File
@@ -105,13 +105,8 @@ fn expected_load_failures(path: &str) -> u32 {
// TRAVERSE-WORDLIST / NAME>COMPILE / NAME>INTERPRET blocks leak as // TRAVERSE-WORDLIST / NAME>COMPILE / NAME>INTERPRET blocks leak as
// unknown-word errors. Fix the SOURCE/`>IN` interaction with // unknown-word errors. Fix the SOURCE/`>IN` interaction with
// line-mode input and drop this to 0. // line-mode input and drop this to 0.
//
// The 38th: line 368 `R> DROP TRUE` runs interpreted (its enclosing
// definition aborted on the missing NAME?), and the bare `R>` used
// to underflow the return stack silently; stack guards now report
// it as "Return stack underflow (throw -6)".
if path.ends_with("/toolstest.fth") { if path.ends_with("/toolstest.fth") {
return 38; return 37;
} }
0 0
} }
@@ -344,32 +339,3 @@ fn compliance_tools() {
let errors = run_suite(&mut vm, "toolstest.fth"); let errors = run_suite(&mut vm, "toolstest.fth");
assert_eq!(errors, 0, "Programming-Tools: {errors} test failures"); assert_eq!(errors, 0, "Programming-Tools: {errors} test failures");
} }
/// The Forth 2012 Core suite against consolidated code.
///
/// `CONSOLIDATE` recompiles the whole dictionary into one WASM module, which
/// is where cross-word typed calls live: a word with a known stack effect
/// gets a fast entry taking and returning its stack items as WASM values,
/// and its `() -> ()` wrapper keeps the table slot. Nothing else covers that
/// path for correctness, so run the suite on top of it.
#[test]
fn compliance_core_after_consolidate() {
let mut vm = ForthVM::<NativeRuntime>::new().expect("Failed to create ForthVM");
let tester_path = format!("{SUITE_DIR}/tester.fr");
let f1 = load_file(&mut vm, &tester_path);
assert_load_fails_within_baseline(&tester_path, f1);
vm.evaluate("CONSOLIDATE").expect("CONSOLIDATE failed");
vm.take_output();
let core_path = format!("{SUITE_DIR}/core.fr");
let f2 = load_file(&mut vm, &core_path);
assert_load_fails_within_baseline(&core_path, f2);
let _ = vm.evaluate("DECIMAL #ERRORS @");
let errors = vm.data_stack().first().copied().unwrap_or(-1);
assert_eq!(
errors, 0,
"Core word set after CONSOLIDATE: {errors} failures"
);
}
+1 -1
View File
@@ -12,7 +12,7 @@ workspace = true
crate-type = ["cdylib", "rlib"] crate-type = ["cdylib", "rlib"]
[dependencies] [dependencies]
wafer-core = { path = "../core", version = "0.2.8", default-features = false, features = ["crypto"] } wafer-core = { path = "../core", version = "0.1.0", default-features = false, features = ["crypto"] }
wasm-bindgen = "0.2" wasm-bindgen = "0.2"
js-sys = "0.3" js-sys = "0.3"
send_wrapper = { workspace = true } send_wrapper = { workspace = true }
+4 -8
View File
@@ -6,9 +6,8 @@ use send_wrapper::SendWrapper;
use wasm_bindgen::prelude::*; use wasm_bindgen::prelude::*;
use wafer_core::config::WaferConfig; use wafer_core::config::WaferConfig;
use wafer_core::memory::{CELL_SIZE, PAD_BASE, PAD_SIZE, SYSVAR_BASE_VAR}; use wafer_core::memory::{CELL_SIZE, PAD_BASE, PAD_SIZE};
use wafer_core::outer::ForthVM; use wafer_core::outer::ForthVM;
use wafer_core::runtime::Runtime;
use wafer_core::runtime::{HostAccess, HostFn}; use wafer_core::runtime::{HostAccess, HostFn};
use crate::runtime_web::WebRuntime; use crate::runtime_web::WebRuntime;
@@ -54,12 +53,9 @@ impl WaferRepl {
/// Get the current number base (10 = decimal, 16 = hex). /// Get the current number base (10 = decimal, 16 = hex).
pub fn base(&mut self) -> u32 { pub fn base(&mut self) -> u32 {
self.vm.runtime_mut().mem_read_i32(SYSVAR_BASE_VAR) as u32 // BASE is stored at SYSVAR_BASE_VAR in WASM memory
} self.vm.take_output(); // no-op side effect; just return base
10 // TODO: read from memory once we have a getter
/// Names of all user-facing words (visible, non-internal), newest first.
pub fn words(&self) -> Vec<String> {
self.vm.word_names()
} }
/// Reset the VM to initial state. /// Reset the VM to initial state.
+2 -19
View File
@@ -38,23 +38,6 @@ impl WebHostAccess {
} }
} }
/// An exception on its way back out of compiled code. Host words rethrow the
/// Forth message (`Stack underflow`, an `ABORT"` text, a `THROW` description),
/// so surface exactly that and nothing else — the JS `Error` carries the whole
/// engine stack in its message, which is noise to a Forth programmer. Anything
/// without a message is a genuine runtime fault and keeps the call context.
fn call_error(fn_index: u32, e: &JsValue) -> anyhow::Error {
match Reflect::get(e, &"message".into())
.ok()
.and_then(|m| m.as_string())
.and_then(|m| m.lines().next().map(str::trim).map(str::to_string))
.filter(|m| !m.is_empty())
{
Some(msg) => anyhow::anyhow!("{msg}"),
None => anyhow::anyhow!("call_func({fn_index}) failed: {e:?}"),
}
}
impl HostAccess for WebHostAccess { impl HostAccess for WebHostAccess {
fn mem_read_i32(&mut self, addr: u32) -> i32 { fn mem_read_i32(&mut self, addr: u32) -> i32 {
let view = js_sys::Int32Array::new(&self.buffer()); let view = js_sys::Int32Array::new(&self.buffer());
@@ -151,7 +134,7 @@ impl HostAccess for WebHostAccess {
.dyn_into() .dyn_into()
.map_err(|_| anyhow::anyhow!("table entry {fn_index} is not a function"))?; .map_err(|_| anyhow::anyhow!("table entry {fn_index} is not a function"))?;
func.call0(&JsValue::NULL) func.call0(&JsValue::NULL)
.map_err(|e| call_error(fn_index, &e))?; .map_err(|e| anyhow::anyhow!("call_func({fn_index}) failed: {e:?}"))?;
Ok(()) Ok(())
} }
} }
@@ -423,7 +406,7 @@ impl Runtime for WebRuntime {
.dyn_into() .dyn_into()
.map_err(|_| anyhow::anyhow!("table entry {fn_index} is not callable"))?; .map_err(|_| anyhow::anyhow!("table entry {fn_index} is not callable"))?;
func.call0(&JsValue::NULL) func.call0(&JsValue::NULL)
.map_err(|e| call_error(fn_index, &e))?; .map_err(|e| anyhow::anyhow!("call_func({fn_index}) failed: {e:?}"))?;
Ok(()) Ok(())
} }
+23 -42
View File
@@ -1,11 +1,8 @@
import init, { WaferRepl } from './pkg/wafer_web.js'; import init, { WaferRepl } from './pkg/wafer_web.js';
let repl = null; let repl = null;
const HISTORY_KEY = 'wafer-history'; const history = [];
const HISTORY_MAX = 200; let historyIdx = -1;
const history = JSON.parse(localStorage.getItem(HISTORY_KEY) || '[]');
let historyIdx = history.length;
let builtinWords = null;
const WORD_CATEGORIES = { const WORD_CATEGORIES = {
'Stack': 'DUP DROP SWAP OVER ROT NIP TUCK 2DUP 2DROP 2SWAP 2OVER PICK ROLL DEPTH .S'.split(' '), 'Stack': 'DUP DROP SWAP OVER ROT NIP TUCK 2DUP 2DROP 2SWAP 2OVER PICK ROLL DEPTH .S'.split(' '),
@@ -42,12 +39,10 @@ function updateStack() {
if (!repl) return; if (!repl) return;
try { try {
const stack = repl.data_stack(); const stack = repl.data_stack();
const base = repl.base();
const suffix = base !== 10 ? ` [base ${base}]` : '';
if (stack.length === 0) { if (stack.length === 0) {
stackBar.textContent = `Stack: (empty)${suffix}`; stackBar.textContent = 'Stack: (empty)';
} else { } else {
stackBar.textContent = `Stack <${stack.length}> ${stack.join(' ')}${suffix}`; stackBar.textContent = `Stack <${stack.length}> ${stack.join(' ')}`;
} }
} catch { } catch {
stackBar.textContent = 'Stack: (error)'; stackBar.textContent = 'Stack: (error)';
@@ -55,25 +50,19 @@ function updateStack() {
} }
function updateUserWords() { function updateUserWords() {
const list = document.getElementById('user-word-list'); const cat = document.getElementById('cat-user');
if (!list || !repl || !builtinWords) return; if (!cat) return;
list.innerHTML = ''; // We'll track user words by checking what the REPL evaluates
for (const w of repl.words()) { // For now, just show the category
if (!builtinWords.has(w)) list.appendChild(wordChip(w));
}
} }
function evaluate(line, record = true) { function evaluate(line) {
if (!repl) return; if (!repl) return;
const trimmed = line.trim(); const trimmed = line.trim();
if (!trimmed) return; if (!trimmed) return;
// Add to history (user-typed lines only; skip consecutive duplicates) // Add to history
if (record && history[history.length - 1] !== trimmed) { history.push(trimmed);
history.push(trimmed);
if (history.length > HISTORY_MAX) history.splice(0, history.length - HISTORY_MAX);
localStorage.setItem(HISTORY_KEY, JSON.stringify(history));
}
historyIdx = history.length; historyIdx = history.length;
try { try {
@@ -94,7 +83,6 @@ function evaluate(line, record = true) {
updatePrompt(); updatePrompt();
updateStack(); updateStack();
updateUserWords();
} }
// Input handling // Input handling
@@ -128,18 +116,6 @@ document.getElementById('btn-toggle-words').addEventListener('click', () => {
document.getElementById('word-panel').classList.toggle('collapsed'); document.getElementById('word-panel').classList.toggle('collapsed');
}); });
function wordChip(w) {
const chip = document.createElement('span');
chip.className = 'word-chip';
chip.textContent = w;
chip.title = w;
chip.addEventListener('click', () => {
input.value += (input.value.length > 0 ? ' ' : '') + w;
input.focus();
});
return chip;
}
function buildWordPanel() { function buildWordPanel() {
const container = document.getElementById('word-categories'); const container = document.getElementById('word-categories');
container.innerHTML = ''; container.innerHTML = '';
@@ -153,7 +129,15 @@ function buildWordPanel() {
const list = document.createElement('div'); const list = document.createElement('div');
list.className = 'word-list'; list.className = 'word-list';
for (const w of words) { for (const w of words) {
list.appendChild(wordChip(w)); const chip = document.createElement('span');
chip.className = 'word-chip';
chip.textContent = w;
chip.title = w;
chip.addEventListener('click', () => {
input.value += (input.value.length > 0 ? ' ' : '') + w;
input.focus();
});
list.appendChild(chip);
} }
cat.appendChild(list); cat.appendChild(list);
container.appendChild(cat); container.appendChild(cat);
@@ -195,7 +179,7 @@ document.getElementById('btn-run-init').addEventListener('click', () => {
if (code.trim()) { if (code.trim()) {
// Run each line separately // Run each line separately
for (const line of code.split('\n')) { for (const line of code.split('\n')) {
if (line.trim()) evaluate(line, false); if (line.trim()) evaluate(line);
} }
} }
localStorage.setItem('wafer-init-code', code); localStorage.setItem('wafer-init-code', code);
@@ -230,7 +214,6 @@ document.getElementById('btn-reset').addEventListener('click', () => {
appendLine('WAFER reset.', 'line-ok'); appendLine('WAFER reset.', 'line-ok');
updatePrompt(); updatePrompt();
updateStack(); updateStack();
updateUserWords();
} catch (e) { } catch (e) {
appendLine(`Reset error: ${e.message}`, 'line-error'); appendLine(`Reset error: ${e.message}`, 'line-error');
} }
@@ -242,8 +225,6 @@ async function boot() {
try { try {
await init(); await init();
repl = new WaferRepl(); repl = new WaferRepl();
// Everything defined at boot is "builtin"; later definitions are user words
builtinWords = new Set(repl.words());
output.innerHTML = ''; output.innerHTML = '';
appendLine('WAFER — WebAssembly Forth Engine in Rust', 'line-output'); appendLine('WAFER — WebAssembly Forth Engine in Rust', 'line-output');
appendLine(`Type Forth at the > prompt. Press ? for help.`, 'line-output'); appendLine(`Type Forth at the > prompt. Press ? for help.`, 'line-output');
@@ -260,7 +241,7 @@ async function boot() {
const initCode = document.getElementById('init-code').value; const initCode = document.getElementById('init-code').value;
if (initCode.trim()) { if (initCode.trim()) {
for (const line of initCode.split('\n')) { for (const line of initCode.split('\n')) {
if (line.trim()) evaluate(line, false); if (line.trim()) evaluate(line);
} }
localStorage.setItem('wafer-init-code', initCode); localStorage.setItem('wafer-init-code', initCode);
} }
@@ -271,7 +252,7 @@ async function boot() {
const code = atob(location.hash.slice(1)); const code = atob(location.hash.slice(1));
document.getElementById('init-code').value = code; document.getElementById('init-code').value = code;
for (const line of code.split('\n')) { for (const line of code.split('\n')) {
if (line.trim()) evaluate(line, false); if (line.trim()) evaluate(line);
} }
} catch { /* ignore bad hash */ } } catch { /* ignore bad hash */ }
} }
+2 -2
View File
@@ -18,11 +18,11 @@ confidence-threshold = 0.8
[bans] [bans]
multiple-versions = "deny" multiple-versions = "deny"
wildcards = "deny" wildcards = "deny"
# Transitive duplicates from wasmtime v47 dependencies # Transitive duplicates from wasmtime v31 -- will resolve when upgrading
skip = [ skip = [
"getrandom", "getrandom",
"syn",
"hashbrown", "hashbrown",
"r-efi",
"thiserror", "thiserror",
"thiserror-impl", "thiserror-impl",
"wasm-encoder", "wasm-encoder",
+17 -92
View File
@@ -14,7 +14,7 @@ This document describes every optimization that makes sense for WAFER, why it ma
| # | Optimization | Level | Status | Impact | | # | Optimization | Level | Status | Impact |
| -- | -------------------------- | ------------ | ----------- | ------- | | -- | -------------------------- | ------------ | ----------- | ------- |
| 1 | Stack-to-Local Promotion | Codegen | Phase 4 | Highest | | 1 | Stack-to-Local Promotion | Codegen | Phase 2 | Highest |
| 2 | Peephole Optimization | IR pass | Done | High | | 2 | Peephole Optimization | IR pass | Done | High |
| 3 | Constant Folding | IR pass | Done | High | | 3 | Constant Folding | IR pass | Done | High |
| 4 | Inlining | IR pass | Done | High | | 4 | Inlining | IR pass | Done | High |
@@ -29,19 +29,12 @@ This document describes every optimization that makes sense for WAFER, why it ma
| 13 | Startup Batching | Architecture | Done | Low | | 13 | Startup Batching | Architecture | Done | Low |
| 14 | Self-Recursive Direct Call | Codegen | Done | High | | 14 | Self-Recursive Direct Call | Codegen | Done | High |
| 15 | Float / Double-Cell | Codegen | Not started | Future | | 15 | Float / Double-Cell | Codegen | Not started | Future |
| 16 | Typed Calling Convention | Codegen | Done | Highest |
| 17 | Self-Guard Expansion | IR pass | Done | Medium |
## 1. Stack-to-Local Promotion ## 1. Stack-to-Local Promotion
**Status: Phase 4 done.** Straight-line code, DO/LOOP, IF/ELSE and the BEGIN loop family use WASM locals instead of the memory stack, per region rather than per word. Stack manipulation ops (Swap, Rot, Nip, Tuck, Dup, Drop) emit zero WASM instructions. Loop index/limit stay in WASM locals (zero return stack traffic). Switchable via `WaferConfig::codegen.stack_to_local_promotion`. **Status: Phase 2 done.** Words with straight-line code, DO/LOOP, and IF/ELSE use WASM locals instead of memory stack. Stack manipulation ops (Swap, Rot, Nip, Tuck, Dup, Drop) emit zero WASM instructions. Loop index/limit kept in WASM locals (zero return stack traffic). Switchable via `WaferConfig::codegen.stack_to_local_promotion`.
- **Phase 1** straight-line code. Phase 1 covered straight-line code only. Phase 2 extends to DO/LOOP (with stack-neutrality check) and IF/ELSE/THEN (with equal-branch-effect check). BEGIN loops and BeginDoubleWhileRepeat are not yet promoted.
- **Phase 2** — DO/LOOP (stack-neutrality check) and IF/ELSE/THEN (equal-branch-effect check).
- **Phase 3**_per region instead of per word_. Promotion used to be all-or-nothing: one `.`, `CR`, `>R` or host call anywhere in a definition put the entire body on the memory stack, hot loops included, which costs 2.2 ns per loop-carried add instead of 0.31 — the accumulator round-trips through store-to-load forwarding rather than staying in a register. `emit_body` now partitions a body into maximal promotable stretches and runs the simulator over each, loading what a region reads and writing back what it leaves. A region may only use `I` / `J` when the DO loops naming them are inside it, and a straight-line region needs at least three operations to pay for its prologue and epilogue; a loop always does.
- **Phase 4**`BEGIN..UNTIL`, `BEGIN..AGAIN` and `BEGIN..WHILE..REPEAT`, when the construct is provably stack-neutral: UNTIL's body nets +1 (the flag it consumes), AGAIN's nets 0, and for WHILE..REPEAT the test and the body must balance _separately_, because WHILE leaves the loop between them and a net that only added up over the pair would give the two exits different stack shapes. Bodies containing an `EXIT` stay out, the same rule DO/LOOP follows.
Still not promoted: `BeginDoubleWhileRepeat`, `>R`/`R>`, floats, `{: :}` locals, `SP@`/`DEPTH`/`EXECUTE`, and the flat forward-block IR ops.
### The Problem ### The Problem
@@ -459,101 +452,33 @@ Fibonacci(25) with ~243K recursive calls:
The optimization is implemented in `emit_op` for `IrOp::Call`: when `ctx.self_word_id == Some(word_id)`, emit `call WORD_FUNC` (function index 1 in the word's own module). The `self_word_id` is derived from `CodegenConfig::base_fn_index`. The optimization is implemented in `emit_op` for `IrOp::Call`: when `ctx.self_word_id == Some(word_id)`, emit `call WORD_FUNC` (function index 1 in the word's own module). The `self_word_id` is derived from `CodegenConfig::base_fn_index`.
The numbers above are the state before section 16: they measure the call instruction, and what dominated turned out to be the calling _convention_ around it. A self-recursive word that is also typed now calls its own fast entry instead, and Fibonacci(25) is 356 microseconds rather than 1.6 ms.
## 15. Float and Double-Cell Stack ## 15. Float and Double-Cell Stack
**Status: Not started.** `PushI64` and `PushF64` exist as IR ops but are stubs in codegen. Float stack operations are currently all host functions. **Status: Not started.** `PushI64` and `PushF64` exist as IR ops but are stubs in codegen. Float stack operations are currently all host functions.
The float stack lives in its own memory region (0x2540--0x2D40). Float operations will have the same memory-based overhead as integer operations, but worse: `f64` values are 8 bytes, doubling the memory traffic per push/pop. Stack-to-local promotion (section 1) is even more impactful for floats because WASM has native `f64` locals and operand stack support. The float stack lives in its own memory region (0x2540--0x2D40). Float operations will have the same memory-based overhead as integer operations, but worse: `f64` values are 8 bytes, doubling the memory traffic per push/pop. Stack-to-local promotion (section 1) is even more impactful for floats because WASM has native `f64` locals and operand stack support.
## 16. Typed Calling Convention
**Status: Done.** A word whose stack effect is statically known compiles to two entry points: a fast one with signature `(i32 x p) -> (i32 x q)`, carrying its stack items as WASM values, and the usual `( -- )` wrapper that moves those items on and off the memory data stack. The wrapper keeps the function-table slot, so `EXECUTE`, the outer interpreter, host words and `CATCH` see exactly the ABI they saw before; only direct calls inside a module take the fast entry. `WAFER_TYPED_CALLS=0` falls back.
### The Problem
This is what the SwiftForth gap was made of. sf64 keeps TOS in `RBX` and the stack pointer in `RBP`, and both survive a `CALL` untouched, so its `FIB` is 16 instructions and about 7 memory touches per node. WAFER kept the whole stack in linear memory and flushed its cached `$dsp` to an imported global before every call: about 36 touches. Section 1's simulator, which already promoted loop and `IF` bodies into locals, refused any body containing a call or an `EXIT` -- exactly the words where the convention cost the most.
### The Effect Fixpoint
Self-recursion makes the stack-effect equation circular (`d = k + m*d`), so the effect is solved by iterating a guess until it reproduces itself: `FIB` settles on `(1,1)` in two rounds, while `: F 1 RECURSE ;` never settles and stays untyped. `CONSOLIDATE` extends this across words, since it puts them all in one module: the effects are solved from the leaves outward, and 105 of 187 words in a booted dictionary end up typed.
### Impact
Fibonacci(25) went from 1035 to 366 microseconds, 4.3x slower than `sf64` to 1.2x. Stack guards became nearly free as a side effect -- they hang off the memory-stack push/pop choke points, and a typed word barely has any -- so the default guards-on configuration that the REPL and the web build use went from 1631 to 365 microseconds on the same benchmark.
Untyped by design: anything using `SP@`, `DEPTH`, `EXECUTE`, `>R`/`R>`, floats or locals; anything calling a word that is itself untyped, which in the JIT path means every call except `RECURSE`; mutually recursive words; and words whose effect is not static -- branches that disagree on depth, `EXIT` at the wrong depth, a non-neutral loop body, or a recursion that grows the stack per level.
## 17. Self-Guard Expansion
**Status: Done.** A recursive word's base-case guard is duplicated into its own call sites, so the leaves of the recursion cost a test instead of a call. Implemented in `optimizer.rs::expand_self_guard`, gated on `OptConfig::self_guard`, and applied after inlining so the later passes still run over the result.
### The Shape
A recursive Forth word almost always opens with a guard that returns early:
```forth
: FIB DUP 2 < IF EXIT THEN DUP 1- RECURSE SWAP 2 - RECURSE + ;
```
Every leaf of the recursion is then a call whose entire body is `DUP 2 <`. The pass rewrites each `Call(self)` as
```forth
DUP 2 < IF ( leave it ) ELSE RECURSE THEN
```
which computes the same thing -- the callee would have run the guard, taken the branch and returned. When the guard returns a value rather than its argument (`IF DROP 0 EXIT THEN`), that value moves into the then-branch with it.
### Why It Is Bounded
The guard runs twice along the recursive path: once at the call site, once inside the callee. So it must be small and free of effects -- `MAX_GUARD_OPS` is six, and the operations are restricted to stack shuffles, arithmetic and comparisons; a call, a memory access or a branch disqualifies it. `MAX_GUARD_SITES` caps the expansion at four call sites, since each one replicates the guard. A `TailCall` is never expanded, which keeps this pass and tail-call detection from having to agree about what tail position means.
### Impact
Fibonacci(25): 356 to 237 microseconds on the arm64 development machine. In fib's tree half of all nodes are leaves, which is where the factor comes from. It does not take Fibonacci past `sf64`, though the arm64 table below says otherwise: with both engines native on x86-64, Fibonacci reads 1.16x and stays the one benchmark `sf64` wins.
## Current Performance vs Gforth ## Current Performance vs Gforth
All optimizations enabled, release mode, measured with UTIME: All optimizations enabled, release mode, measured with UTIME:
All three engines **native x86-64**, idle 16-vCPU Xeon Platinum 8124M @ 3.0 GHz,
median of three runs:
``` ```
Benchmark WAFER gforth sf64 WAFER/gf WAFER/sf Benchmark WAFER CONSOL gforth WAFER/gf
Fibonacci(25) 411 3221 355 0.13x 1.16x Fibonacci(25) 1629 1535 3422 0.45x
Factorial(12)x100K 994 7141 3058 0.14x 0.33x Factorial(12)x10K 340 339 638 0.53x
GCD-bench(20K) 1591 3211 2423 0.50x 0.66x GCD-bench(500) 18 15 30 0.50x
NestedLoops(50)x1K 889 6824 2342 0.13x 0.38x NestedLoops(50) 84 73 720 0.10x
Collatz(2K) 391 3981 1659 0.10x 0.24x Collatz(2K) 1212 1202 3914 0.31x
``` ```
The same suite on the arm64 development machine (M1 Ultra), which is what the Times in microseconds. WAFER/gf < 1.0 means WAFER is faster.
regression limits in `comparison.rs` are calibrated against:
```
Benchmark WAFER CONSOL gforth sf64 WAFER/gf WAFER/sf
Fibonacci(25) 237 242 3340 287 0.07x 0.83x
Factorial(12)x100K 480 479 6109 1594 0.08x 0.30x
GCD-bench(20K) 549 541 1830 797 0.30x 0.68x
NestedLoops(50)x1K 501 509 7092 1898 0.07x 0.26x
Collatz(2K) 196 190 3955 633 0.05x 0.30x
```
Times in microseconds. WAFER/gf < 1.0 means WAFER is faster. The two tables
disagree because the only SwiftForth build for macOS is x86-64 under Rosetta 2
while WAFER and gforth are native arm64 -- and the emulation penalty lands
hardest on the call-heavy benchmark, so Fibonacci reads 0.83x on arm64 and 1.16x
when neither engine is emulated. Believe the x86-64 table about the engines. One
caveat holds for both: sf64 uses 64-bit cells to WAFER's 32-bit.
## Remaining Opportunities ## Remaining Opportunities
| Optimization | Status | Potential Impact | | Optimization | Status | Potential Impact |
| -------------------------------- | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | -------------------------------- | ------------------- | ----------------------------------------------------- |
| Scoped exit for the inliner | Not started | The inliner still refuses any body containing an `EXIT`, because an inlined one would return from the caller. Compiling it as a branch to the end of a block would unlock inlining for every word with an early return, not just the guard shape section 17 handles | | BEGIN loop promotion | Not started | Would speed up GCD-style tight loops further |
| BeginDoubleWhileRepeat promotion | Not started | Rare pattern, low priority. Its promoted emitter exists but has no loop fixup and is unverified | | BeginDoubleWhileRepeat promotion | Not started | Rare pattern, low priority |
| LEAVE as IR primitive | Not started | Would enable fast-path for loops with LEAVE | | LEAVE as IR primitive | Not started | Would enable fast-path for loops with LEAVE |
| Float stack-to-local | Not started | Eliminate float stack memory traffic | | Float stack-to-local | Not started | Eliminate float stack memory traffic |
| WASM tail calls proposal | Waiting on wasmtime | Would eliminate stack growth for tail-recursive words | | WASM tail calls proposal | Waiting on wasmtime | Would eliminate stack growth for tail-recursive words |
+2 -4
View File
@@ -282,13 +282,11 @@ When the compiler encounters a word reference during compilation, it emits:
(call_indirect (type $void) (table 0)) ;; indirect call through the table (call_indirect (type $void) (table 0)) ;; indirect call through the table
``` ```
**Self-recursive optimization**: When a word calls itself (RECURSE), the codegen detects this and emits a direct `call` instead of `call_indirect`, eliminating the table lookup and signature check (~3x faster for recursive words like Fibonacci). When the word is also typed, that direct call goes to its fast entry -- see below. **Self-recursive optimization**: When a word calls itself (RECURSE), the codegen detects this and emits a direct `call` instead of `call_indirect`, eliminating the table lookup and signature check (~3x faster for recursive words like Fibonacci).
**After CONSOLIDATE**: All `call_indirect` between words in the consolidated module are replaced with direct `call` instructions, giving similar benefits for cross-word calls. **After CONSOLIDATE**: All `call_indirect` between words in the consolidated module are replaced with direct `call` instructions, giving similar benefits for cross-word calls.
At runtime, wasmtime resolves the table entry and calls the target function. Because all functions share the same memory, globals, and table, state passes between words through the data stack in linear memory. At runtime, wasmtime resolves the table entry and calls the target function. Because all functions share the same memory, globals, and table, state passes between words through the data stack in linear memory. There are no function parameters or return values at the WASM level -- everything goes through the stack.
**Typed entry points**: that last sentence is the default, not the whole story. A word whose stack effect is statically known also gets a _fast_ entry with signature `(i32 x p) -> (i32 x q)`, which takes its arguments as WASM values and returns its results the same way, so they stay in registers across the call instead of round-tripping through linear memory. The `( -- )` function above is then a wrapper around it, and it is the wrapper that keeps the table slot -- so `EXECUTE`, the outer interpreter, host words and `CATCH` see the memory ABI unchanged. Only a direct call inside the same module takes the fast entry: `RECURSE` in the JIT path, and every resolvable call after `CONSOLIDATE`. See [OPTIMIZATIONS.md](OPTIMIZATIONS.md) section 16.
This is subroutine threading: each word is a subroutine, and calling a word is an indirect function call. This is subroutine threading: each word is a subroutine, and calling a word is an indirect function call.
+10 -32
View File
@@ -33,10 +33,7 @@ contexts:
- include: compare - include: compare
- include: memory - include: memory
- include: io - include: io
- include: pictured
- include: string_ops
- include: float - include: float
- include: tools
- include: dictionary - include: dictionary
- include: exception - include: exception
- include: parsing - include: parsing
@@ -98,31 +95,27 @@ contexts:
# Quotations (Core-Ext 6.2.0455): [: ... ;] compiles an anonymous word. # Quotations (Core-Ext 6.2.0455): [: ... ;] compiles an anonymous word.
- match: '(?i)(?:^|(?<=\s))(\[:|;\]){{ident_break}}' - match: '(?i)(?:^|(?<=\s))(\[:|;\]){{ident_break}}'
scope: keyword.other.definition.forth scope: keyword.other.definition.forth
- match: '(?i)(?:^|(?<=\s))(VARIABLE|2VARIABLE|CONSTANT|2CONSTANT|VALUE|CREATE|DEFER|MARKER|REMEMBER|BUFFER:|FCONSTANT|FVARIABLE)(\s+)(\S+)?' - match: '(?i)(?:^|(?<=\s))(VARIABLE|2VARIABLE|CONSTANT|2CONSTANT|VALUE|CREATE|DEFER|MARKER|BUFFER:|FCONSTANT|FVARIABLE)(\s+)(\S+)?'
captures: captures:
1: keyword.other.defining.forth 1: keyword.other.defining.forth
3: entity.name.constant.forth 3: entity.name.constant.forth
- match: '(?i)(?:^|(?<=\s))(DOES>|IMMEDIATE|RECURSE|POSTPONE|COMPILE,|LITERAL|2LITERAL|FLITERAL|SLITERAL|DEFER!|DEFER@){{ident_break}}' - match: '(?i)(?:^|(?<=\s))(DOES>|IMMEDIATE|RECURSE|POSTPONE|COMPILE,|LITERAL|2LITERAL|FLITERAL|SLITERAL){{ident_break}}'
scope: keyword.other.defining.forth scope: keyword.other.defining.forth
control: control:
- match: '(?i)(?:^|(?<=\s))(IF|THEN|ELSE|BEGIN|UNTIL|WHILE|REPEAT|AGAIN|DO|\?DO|LOOP|\+LOOP|LEAVE|UNLOOP|EXIT|CASE|OF|ENDOF|ENDCASE|QUIT){{ident_break}}' - match: '(?i)(?:^|(?<=\s))(IF|THEN|ELSE|BEGIN|UNTIL|WHILE|REPEAT|AGAIN|DO|\?DO|LOOP|\+LOOP|LEAVE|UNLOOP|EXIT|CASE|OF|ENDOF|ENDCASE|QUIT){{ident_break}}'
scope: keyword.control.forth scope: keyword.control.forth
# Conditional compilation (Tools-ext 15.6.2).
- match: '(?i)(?:^|(?<=\s))(\[IF\]|\[ELSE\]|\[THEN\]|\[DEFINED\]|\[UNDEFINED\]){{ident_break}}'
scope: keyword.control.conditional-compilation.forth
stack_ops: stack_ops:
- match: '(?i)(?:^|(?<=\s))(DUP|\?DUP|DROP|SWAP|OVER|ROT|-ROT|NIP|TUCK|PICK|ROLL|2DUP|2DROP|2SWAP|2OVER|2ROT|DEPTH|SP@){{ident_break}}' - match: '(?i)(?:^|(?<=\s))(DUP|\?DUP|DROP|SWAP|OVER|ROT|-ROT|NIP|TUCK|PICK|ROLL|2DUP|2DROP|2SWAP|2OVER|2ROT|DEPTH|SP@){{ident_break}}'
scope: support.function.stack.forth scope: support.function.stack.forth
return_stack: return_stack:
# RP@ / RDEPTH are WAFER extensions (gforth-style return-stack access). - match: '(?i)(?:^|(?<=\s))(>R|R>|R@|2>R|2R>|2R@|N>R|NR>|I|J|CS-PICK|CS-ROLL){{ident_break}}'
- match: '(?i)(?:^|(?<=\s))(>R|R>|R@|2>R|2R>|2R@|N>R|NR>|I|J|CS-PICK|CS-ROLL|RP@|RDEPTH){{ident_break}}'
scope: support.function.return-stack.forth scope: support.function.return-stack.forth
arithmetic: arithmetic:
- match: '(?i)(?:^|(?<=\s))(\+|-|\*|/|MOD|/MOD|\*/|\*/MOD|NEGATE|ABS|MIN|MAX|1\+|1-|2\*|2/|M\*|M\+|M\*/|UM\*|UM/MOD|FM/MOD|SM/REM|S>D|D>S|D\+|D-|DNEGATE|DABS|DMAX|DMIN|D2\*|D2/){{ident_break}}' - match: '(?i)(?:^|(?<=\s))(\+|-|\*|/|MOD|/MOD|\*/|\*/MOD|NEGATE|ABS|MIN|MAX|1\+|1-|2\*|2/|M\*|M\+|M\*/|UM\*|UM/MOD|FM/MOD|SM/REM|S>D|D>S){{ident_break}}'
scope: keyword.operator.arithmetic.forth scope: keyword.operator.arithmetic.forth
logic: logic:
@@ -130,36 +123,21 @@ contexts:
scope: keyword.operator.logical.forth scope: keyword.operator.logical.forth
compare: compare:
- match: '(?i)(?:^|(?<=\s))(=|<>|<|>|<=|>=|U<|U>|0=|0<>|0<|0>|D<|D=|D0<|D0=|DU<|WITHIN){{ident_break}}' - match: '(?i)(?:^|(?<=\s))(=|<>|<|>|<=|>=|U<|U>|0=|0<>|0<|0>){{ident_break}}'
scope: keyword.operator.comparison.forth scope: keyword.operator.comparison.forth
memory: memory:
- match: '(?i)(?:^|(?<=\s))(@|!|C@|C!|\+!|2@|2!|C,|ALLOT|HERE|ALIGN|ALIGNED|CELL\+|CELLS|CHAR\+|CHARS|UNUSED|MOVE|CMOVE|CMOVE>|FILL|ERASE|BLANK|ALLOCATE|FREE|RESIZE|PAD){{ident_break}}' - match: '(?i)(?:^|(?<=\s))(@|!|C@|C!|\+!|2@|2!|ALLOT|HERE|ALIGN|ALIGNED|CELL\+|CELLS|CHAR\+|CHARS|UNUSED|MOVE|CMOVE|CMOVE>|FILL|ERASE|BLANK|ALLOCATE|FREE|RESIZE|PAD){{ident_break}}'
scope: support.function.memory.forth scope: support.function.memory.forth
io: io:
- match: '(?i)(?:^|(?<=\s))(EMIT|CR|SPACE|SPACES|TYPE|\.|U\.|\.R|U\.R|D\.|D\.R|\?|KEY|KEY\?|PAGE|AT-XY|ACCEPT|EXPECT|\.S|F\.S|\.RS){{ident_break}}' - match: '(?i)(?:^|(?<=\s))(EMIT|CR|SPACE|SPACES|TYPE|\.|U\.|\.R|U\.R|D\.|D\.R|\?|KEY|KEY\?|PAGE|AT-XY|ACCEPT|EXPECT|\.S){{ident_break}}'
scope: support.function.io.forth scope: support.function.io.forth
# Pictured numeric output (6.1: <# # #S #> HOLD SIGN; HOLDS is Core-Ext).
pictured:
- match: '(?i)(?:^|(?<=\s))(<#|#>|#S|#|HOLD|HOLDS|SIGN){{ident_break}}'
scope: support.function.pictured.forth
# String word set (17.6).
string_ops:
- match: '(?i)(?:^|(?<=\s))(COUNT|COMPARE|-TRAILING|/STRING){{ident_break}}'
scope: support.function.string.forth
float: float:
- match: '(?i)(?:^|(?<=\s))(F\+|F-|F\*\*|F\*|F/|FNEGATE|FABS|FMAX|FMIN|FSQRT|FFLOOR|FROUND|FLOOR|FSINCOS|FSINH|FSIN|FCOSH|FCOS|FTANH|FTAN|FASINH|FASIN|FACOSH|FACOS|FATANH|FATAN2|FATAN|FEXPM1|FEXP|FLNP1|FLN|FLOG|FALOG|F=|F<|F0=|F0<|F~|FDUP|FDROP|FSWAP|FOVER|FROT|FNIP|FTUCK|FDEPTH|F@|F!|FE\.|FS\.|F\.|F>D|D>F|F>S|S>F|>FLOAT|REPRESENT|PRECISION|SET-PRECISION|FALIGN|FALIGNED|DFALIGN|DFALIGNED|SFALIGN|SFALIGNED|FLOAT\+|FLOATS|DFLOAT\+|DFLOATS|SFLOAT\+|SFLOATS|DF@|DF!|SF@|SF!){{ident_break}}' - match: '(?i)(?:^|(?<=\s))(F\+|F-|F\*|F/|FNEGATE|FABS|FMAX|FMIN|FSQRT|FFLOOR|FROUND|FSINCOS|F=|F<|F0=|F0<|F~|FDUP|FDROP|FSWAP|FOVER|FROT|FNIP|FTUCK|FDEPTH|F@|F!|FE\.|FS\.|F\.|F>D|D>F|F>S|S>F|>FLOAT|REPRESENT|PRECISION|SET-PRECISION|FALIGNED|DFALIGNED|SFALIGNED|DF@|DF!|SF@|SF!){{ident_break}}'
scope: support.function.float.forth scope: support.function.float.forth
# Interactive/debug tools (Tools word set + WAFER REPL additions).
tools:
- match: '(?i)(?:^|(?<=\s))(SEE-IR|SEE|DUMP|BYE|HELP){{ident_break}}'
scope: support.function.tools.forth
dictionary: dictionary:
- match: "(?i)(?:^|(?<=\\s))('|\\[']|,|>BODY|FIND|WORDS|ONLY|ALSO|PREVIOUS|DEFINITIONS|FORTH|GET-ORDER|SET-ORDER|GET-CURRENT|SET-CURRENT|WORDLIST|SEARCH-WORDLIST|FORTH-WORDLIST|ENVIRONMENT\\?|EXECUTE){{ident_break}}" - match: "(?i)(?:^|(?<=\\s))('|\\[']|,|>BODY|FIND|WORDS|ONLY|ALSO|PREVIOUS|DEFINITIONS|FORTH|GET-ORDER|SET-ORDER|GET-CURRENT|SET-CURRENT|WORDLIST|SEARCH-WORDLIST|FORTH-WORDLIST|ENVIRONMENT\\?|EXECUTE){{ident_break}}"
scope: support.function.dictionary.forth scope: support.function.dictionary.forth
@@ -169,7 +147,7 @@ contexts:
scope: keyword.control.exception.forth scope: keyword.control.exception.forth
parsing: parsing:
- match: '(?i)(?:^|(?<=\s))(PARSE|PARSE-NAME|WORD|REFILL|EVALUATE|INCLUDE|INCLUDED|SOURCE|SOURCE-ID|>IN|BASE|DECIMAL|HEX|STATE|>NUMBER|SEARCH|SUBSTITUTE|UNESCAPE|REPLACES|S){{ident_break}}' - match: '(?i)(?:^|(?<=\s))(PARSE|PARSE-NAME|WORD|REFILL|EVALUATE|SOURCE|SOURCE-ID|>IN|BASE|STATE|>NUMBER|SEARCH|SUBSTITUTE|UNESCAPE|REPLACES|S){{ident_break}}'
scope: support.function.parsing.forth scope: support.function.parsing.forth
literals: literals:
@@ -207,5 +185,5 @@ contexts:
wafer_extras: wafer_extras:
# WAFER-specific extensions beyond the Forth 2012 standard. # WAFER-specific extensions beyond the Forth 2012 standard.
# When the language grows new user-facing non-standard words, add them here. # When the language grows new user-facing non-standard words, add them here.
- match: '(?i)(?:^|(?<=\s))(CONSOLIDATE|RANDOM|RND-SEED|UTIME|READ-PASSWORD|EMPTY|GILD){{ident_break}}' - match: '(?i)(?:^|(?<=\s))(CONSOLIDATE|RANDOM|RND-SEED|UTIME|READ-PASSWORD){{ident_break}}'
scope: support.function.wafer-extra.forth scope: support.function.wafer-extra.forth