# Strict CalcKernel WASM parity suite

This is a separate, versioned comparison suite for CalcKernel WebAssembly,
Clang WebAssembly, and Rust WebAssembly. Its reports use the stable suite ID
`calckernel-wasm-parity-strict-v1`; the compiler release, binary hash, declared
build commit, observed source checkout commit, and dirty source state are
recorded independently. The pre-change and post-change labels produce separate
report filenames, and the runner refuses to overwrite an existing report.

The six implementation routes are fixed in each of the 11 workload records:
CK baseline, Clang baseline, Rust baseline, CK SIMD128, Clang SIMD128, and Rust
SIMD128. A compiler/profile with an incompatible feature set remains in the
report with its exact target feature declarations and is excluded from the
matched-profile reference denominator. The suite always keeps all 11 workloads
and every compiled profile result.

## Requirements and smoke run

Use `ckc`, Clang with a `wasm32-unknown-unknown` target and `wasm-ld`, Rust with
the `wasm32v1-none` core library, Node.js, and Python 3. The Rust target is
available from rustup:

```sh
rustup target add wasm32v1-none
```

From the Website repository root, first build all six modules and validate all
66 workload/profile combinations against the shared fixture/reference set:

```sh
python3 benchmarks/wasm-parity/run.py \
  --wasm-parity --smoke \
  --ckc /path/to/ckc \
  --wasm-clang /path/to/clang \
  --rustc /path/to/rustc \
  --node /path/to/node \
  --work-dir build/wasm-parity-smoke
```

The smoke run writes no report. It checks complete outputs, input immutability,
output guards, exported ABI functions, absence of WASM imports, and pairwise
disjoint fixed memory regions during reusable worker setup. The same regions
and argument views are reused for warmup and later calls.

## Full measurements

Run the released/pre-change compiler first, then the rebuilt/post-change
compiler in a separate invocation. Each run needs an idle host; the default
configuration performs seven or more interleaved rounds, 500 ms warmup per
route, and batches calibrated near 80 ms.

```sh
python3 benchmarks/wasm-parity/run.py \
  --wasm-parity --build-label pre-change \
  --ckc /path/to/pre-change/ckc \
  --compiler-root /path/to/pre-change/CalcKernel/compiler \
  --compiler-commit 0123456789abcdef0123456789abcdef01234567 \
  --wasm-clang /path/to/clang --rustc /path/to/rustc --node /path/to/node

python3 benchmarks/wasm-parity/run.py \
  --wasm-parity --build-label post-change \
  --ckc /path/to/post-change/ckc \
  --compiler-root /path/to/post-change/CalcKernel/compiler \
  --compiler-commit 89abcdef0123456789abcdef0123456789abcdef \
  --wasm-clang /path/to/clang --rustc /path/to/rustc --node /path/to/node
```

The default result path contains the date, build label, and exact `ckc` binary
hash. Pass `--output` to select another new path. The runner refuses to replace
an existing report or the historical six-implementation JSON.

## Comparison contract

All routes use the same `.ck`, C++, and Rust kernel operations, fixed fixture
bytes, dimensions, row/loop order, and Python reference checks. The `f64`
reductions preserve ascending-index order. Clang uses `-O3`, freestanding
linking, `-fno-fast-math`, `-fno-associative-math`, `-fno-reciprocal-math`, and
`-ffp-contract=off`. Rust uses `-C opt-level=3` with a freestanding `wasm32v1-none`
core target and only the selected `simd128` feature. Its generated wrapper adds
`#![no_std]` and a panic-to-trap handler; the kernel bodies remain byte-for-byte
from the checked-in Rust source. No route enables Relaxed SIMD or changes
floating-point reassociation/FMA semantics. Clang additionally disables the
`mutable-globals` declaration and LLVM's `call-indirect-overlong` encoding
feature where supported. Rust requests `-mutable-globals`, but its final
freestanding `wasm32v1-none` artifact still declares that feature; the raw
declaration is retained and assessed from the final module evidence.

Correctness checks compare finite `f64` values bit-for-bit, including signed
zero and subnormals, and compare `u32` outputs exactly. Any two NaNs compare as
the same semantic class regardless of payload. The fixed 11-workload smoke
fixtures produced identical output/reference hashes for all 66 WASM profile
checks; adversarial signed-zero, NaN-payload, subnormal, FMA, and reassociation
cases are unit tests and are never added to timed workloads.

The selected CK capability profiles are `MVP + MULTI_VALUE + BULK_MEMORY` and
that set plus `SIMD128`. CK artifacts carry `ck.wasm.target` profile metadata and
the profile digest. Clang/Rust artifacts are parsed for their final
`target_features` custom section; the report retains the exact raw
`requiredFeatures` and `disallowedFeatures`. `mutable-globals` is the separate
Import/Export of Mutable Globals proposal, not MVP. The runner records all
module imports, global imports, exports, and the mutability Node observes for
each exported global. A declaration with no global imports or mutable global
exports is marked `declaration_only_no_use_observed`; that is a conservative
feature difference, never an exact declaration match. Any observed global
import or mutable global export makes that artifact ineligible for a matched
capability comparison. `bulk-memory-opt` is a bulk-memory subfeature, while
`call-indirect-overlong` is a reference-types encoding subfeature; both raw
declarations remain visible and are labeled conservative differences. The
runner does not infer instructions from a toolchain's declaration metadata.

The worker allocates each fixture/output/workspace range separately in linear
memory and checks all fixed address intervals for overlap once during
initialization. It reuses those addresses and views for subsequent invocations.
This satisfies CK `noalias`, C++ `__restrict`, and Rust's unsafe disjoint-region
precondition. Output guards and unchanged-input hashes are checked during
smoke, correctness, warmup, and every timed sample. The first Wasm call hashes
inputs before invocation, checks input immutability and guards afterward, and
validates its output before any later invocation can mask a first-call defect.

The report contains all per-call samples, calibration repeats, warmup records,
shuffled order, fixture/reference/output hashes, final artifact hashes, build
commands, compiler/toolchain identity, target feature evidence, and timing
boundary. It also records each final Wasm binary's byte length, Node module
compilation time once per profile worker, per-workload instance construction
time, and that instance's first kernel call. These cold phase costs remain
separate from all hot samples and are not folded into a cross-workload score.
The hot timer includes the repeat loop, output/workspace initialization,
and Node-to-WASM calls; it excludes compilation, process startup, fixture
copies, instantiation, memory growth, output hashing, and report generation.
Ratios compare CK only with the faster feature-compatible Clang/Rust route in
the same profile and workload. The raw report keeps these per-workload
comparisons; the website separately computes a geometric mean within each
profile for its summary. Baseline and SIMD128 are never combined.
