# CalcKernel v0.15.1 six-implementation benchmark suite

This is a separate versioned suite for 11 fixed workloads. It measures CK
Native, C++, Rust, Java 21, Node.js JavaScript, and CK WebAssembly on Node.js.
JavaScript is the per-case baseline at exactly `1.0×`. Array-input workloads
share generated fixture bytes; Monte Carlo uses a common fixed seed. The
historical two-workload report and its NumPy measurements remain separate.

## Requirements and invocation

Use the released CalcKernel `ckc 0.15.1` executable, Clang/Clang++, Rust,
Node.js, and a matching OpenJDK 21 `java`/`javac` pair. Python 3 uses only its
standard library. From the Website repository root:

```sh
python3 benchmarks/suite-v0151/run.py \
  --ckc /path/to/released/ckc \
  --output benchmarks/results/m5-max-v0151-six-impl.json
```

For a quick ABI/correctness check with reduced dimensions and no timing report,
run the same command with `--smoke` and a temporary `--work-dir`.

Optional `--cxx`, `--clang`, `--rustc`, `--node`, `--java`, and `--javac`
arguments select toolchains. `--rounds` may increase the seven default rounds,
but cannot reduce them. `--work-dir` selects temporary fixture/build storage;
otherwise the runner uses a system temporary directory and removes it when the
run ends. A report is written atomically only after all 66 results pass full
output validation. A failed rerun leaves any earlier complete report intact and
removes only its temporary report file.

Do not run the final measurement while another CPU-intensive task is active.
Use the checked-in source and the exact released compiler binary. The report
contains the compiler binary SHA-256, source hashes, toolchain versions,
sanitized build commands, fixture hashes, output/reference hashes, every
per-call sample, batch repeat count, warmup duration, and shuffled round order.

## Workloads

All arrays are little-endian `f64` or `u32`. The runner generates fixtures in
this recorded order using a single xorshift32 stream seeded with `0x6d2b79f5`:
fused `a`, fused `b`, sum input, dot `a`, dot `b`, stencil input, matrix `a`,
matrix `b`, normalization input, polynomial input, piecewise input, and memory
transform input. Each f64 fixture value is `(state >> 8) / 16777216 - 0.5`;
sum values use each raw u32 state. Dijkstra has a separate graph seed
`0x1badb002`: a directed edge exists when `(state & 3)==0`, its weight is
`1+((state >> 2)%1024)`, and a positive-weight ring is forced so every vertex
is reachable. Monte Carlo has no input fixture: each timed invocation resets
its LCG seed to `0x6d2b79f5`.

| Workload | Fixed operation |
| --- | --- |
| Element-wise fused arithmetic | `N=262144`, `out[i]=(a[i]*1.25+b[i]*0.5)*(a[i]-b[i])+0.125`, f64. “Fused” describes one loop pass; it does not claim hardware FMA. |
| Reduction / Sum | `N=1048576`, ascending-index u32 addition modulo `2^32`. |
| Dot Product | `N=262144`, ascending-index f64 `sum += a[i]*b[i]`; no reassociation. |
| 3×3 Image Stencil | 1024×1024 f64; Gaussian weights `[1,2,1;2,4,2;1,2,1]/16`; the full output is written and the border is zero. |
| Matrix Multiply | 256×256 row-major f64; row→inner→column loop order; no BLAS or NumPy. |
| Normalization | `N=262144` f64, two-pass min–max normalization; constant input maps to zero. |
| Polynomial Evaluation | `N=262144` f64, degree-seven polynomial in high-to-low Horner order with coefficients `[1/16,-1/8,1/4,-1/2,1/2,-1/4,1/8,-1/16]`. |
| Monte Carlo Simulation | 1048576 points; u32 LCG `state=(1664525*state+1013904223) mod 2^32` twice per point; `x=(state/256)/16777216` and likewise for y; count points where `x*x+y*y<=1`. |
| Branch-heavy Piecewise Kernel | `N=262144` f64 in `[-0.5,0.5)`; `<-0.25`: `x*x+0.5`; `<0`: `x*0.75-0.125`; `<0.25`: `x*x*x+0.25`; else: `(x-0.25)*1.5`. |
| Memory-bound Transform | `N=8388608` f64; `out[i]=input[i]*0.5+0.25`, 128 MiB for input and output together. This is described as a streaming transform; the size alone does not prove DRAM saturation. |
| Dijkstra Shortest Path | 1024 vertices, dense row-major u32 graph, positive weights and zero for no edge; source 0; O(V²) linear-minimum implementation with lowest-index ties. |

All six implementations use the same loop order and calculations. CK Native and
CK Wasm are compiled from the same `.ck` source. Wasm uses the O3 `simd128`
profile; that profile permits SIMD, but an individual kernel can still remain
scalar. Wasm module compilation, instantiation, memory growth, and fixture copy
are outside timing. The timed Wasm call includes the Node.js to Wasm invocation.

## Validation and timing

The Python standard-library reference computes every output element. Its
Dijkstra reference uses adjacency lists and a binary heap, independent of the
measured O(V²) linear-minimum implementation. Results are compared byte-for-byte
first; if a strict f64 byte difference occurs, every element must still be
within `1e-12` relative/absolute tolerance. Inputs are hashed before and after
each invocation, output guard bytes are checked, and every complete output is
validated before timing and after every recorded sample. A mismatch aborts the
run and produces no report.

Each implementation warms for at least 500 ms. Batches are calibrated near
80 ms, then seven shuffled, interleaved rounds are measured. The batch timer
includes the kernel, required output/workspace initialization, repeat loop,
language call overhead, and the Node-to-Wasm call. It excludes compilation,
process startup, fixture reads/copies, output hashing, and report generation.
All runtime and math-library thread limits are set to one; no BLAS or Apple
Accelerate implementation is used.

For each workload, the displayed multiplier is
`JavaScript median duration / implementation median duration`. JavaScript is
exactly `1.0×`. Ratios are per workload; they are not combined into an overall
language score.
