CalcKernel v0.15.2 is a WebAssembly performance and correctness update. It improves code generation for eligible numerical kernels while keeping strict floating-point semantics and the existing public WebAssembly ABI.
WebAssembly performance and correctness#
The optimization updates target hot numerical workloads, including matrix multiplication, stencil processing, and other measured kernels. Vector candidates continue to require compiler-proven legality conditions; when those conditions do not hold, the compiler retains a scalar implementation. The update does not enable Relaxed SIMD, FMA contraction, floating-point reassociation, or threading as an implicit change to the strict comparison semantics.
The WebAssembly caller-owned-memory ABI and the existing CK syntax and CLI contracts are unchanged. Baseline and SIMD128 remain separate target profiles. A module must be run by a host that supports the features declared for its selected profile.
Hot WebAssembly measurements#
The current complete strict comparison is development-candidate evidence, not a benchmark of the official v0.15.2 release binary. It uses the fixed eleven-workload, six-route suite (11×6×7) on Node.js/V8, with baseline and SIMD128 reported separately. Each run used the same development ckc 0.15.1 binary (SHA-256 5db7c4db4146dda00020941f1047a1ade794741374e9446af4b2669de8979e3b). The reports declare a clean source checkout at 08292f18b6374f9636bae5ab07fbfcb8e9e30091; the binary-to-source relation is not independently verifiable, so the exact binary hash is authoritative.
- First complete run: baseline geometric mean 1.117×, lowest kernel 0.920×; SIMD128 geometric mean 1.016×, lowest kernel 0.923×.
- Independent complete repeat: baseline geometric mean 1.118×, lowest kernel 0.914×; SIMD128 geometric mean 1.008×, lowest kernel 0.920×.
These are hot repeated-call throughput ratios against the faster compatible Clang or Rust WASM route within the same profile. The two reports are independent and are not pooled. See the first raw report, repeat report, and detailed methodology and source bundles.
Official release artifact verification#
The published v0.15.2 Darwin ARM64 archive (SHA-256 91951b5af7291e4a2112c05264d6a972befccc6f7d62b4ee598044beb448485c) contains a native-enabled ckc 0.15.2 binary (SHA-256 f7bc1c70b272678814313c1e55440a9a4f318b6751c45f2bf7f6bb9bddc6b611) from tag commit 8c2596e0403c4b08af8797a8dd2be4f3ac1b2985. It generated the frozen suite's baseline module with SHA-256 7e0baa440af31d251445228c6ae179f7e577bf4d9b2c7769daef839eac8e417b and SIMD128 module with SHA-256 e4574b4b532531016c84944f7614053be0f4f673ccc363838ee4022e3e1d5e9e. Those module bytes and selected feature declarations match both development reports. A fresh eleven-case, six-route smoke passed 66/66 correctness checks and 66/66 exact output-hash checks.
This verifies the official compiler's generated modules and correctness for the frozen suite. The hot throughput ratios above remain measurements of the development ckc 0.15.1 binary; the official release binary was not timed in those runs.
Known cold-call and module-size gaps#
Hot-call results exclude module compilation, instantiation, and each workload's first call. In the same development report, CK's baseline combined suite module was 28,325 bytes (Clang 4,683; Rust 592,509); its SIMD128 module was 34,941 bytes (Clang 4,821; Rust 595,342). These figures describe the benchmark artifacts, not a general module-size guarantee.
The baseline matrix-multiplication first call was 43.954 ms for CK, 13.376 ms for Clang, and 30.574 ms for Rust. Under SIMD128, it was 20.578 ms for CK, 8.850 ms for Clang, and 9.485 ms for Rust. Cold-start performance and module size remain follow-up work; neither is included in the hot geometric means.
Download#
Download CalcKernel and its per-archive SHA-256 sidecars from the v0.15.2 GitHub release, or review the public compiler repository. All six published archive checksums were independently verified. For the earlier release, see the preserved v0.15.1 release page.