CalcKernel 0.15.1 includes the v0.15 WebAssembly improvements for numerical kernels. It adds Bulk Memory to the baseline feature set and makes SIMD128 vectorization available as an explicit O3 target option.
WebAssembly optimization#
Bulk Memory in the baseline#
The v0.15 baseline permits WebAssembly Bulk Memory operations, including memory copy and fill. This helps CK compile bulk array operations to the corresponding Wasm instructions. A host runtime must support the features declared by the module's target.
Optional SIMD128#
Build with O3 and the simd128 target option to enable vectorization. The optimizer can vectorize eligible contiguous numeric maps and reductions; loops that do not meet the safety and target conditions keep scalar code. SIMD remains opt in, so teams can choose a target profile that matches their deployment runtimes.
Target metadata and host memory#
CK Wasm target metadata is now schema 2. It records the selected baseline or simd128 profile and its canonical SHA-256 profile digest; it does not enumerate individual Wasm features. Hosts should validate the selected profile and confirm runtime support for that profile's v0.15 capability set. Update tools that read this metadata before adopting the new target descriptor. The caller-owned-memory ABI is unchanged: the host still owns the data memory and passes it to exported CK functions.
Build a SIMD module#
From the CalcKernel compiler repository root, build the included i32 map example with O3 and SIMD128:
mkdir -p build && ckc emit-wasm examples/wasm/i32_map.ck \
--out build/i32-map-simd.wasm --opt-level 3 --wasm-features simd128 \
--overflow unchecked --bounds unchecked
WebAssembly emission supports unchecked overflow and bounds modes; both are written explicitly here. Choose --wasm-features baseline for a module without SIMD128.
A scoped Node.js measurement#
An earlier, scoped local hot-kernel comparison of CK's v0.15 O3 simd128 implementation on an Apple M5 Max with Node.js 24.14.0 measured about 5.95× Node.js JavaScript for a 16,384-element i32 map and 2.79× for a 16,384-element f64 map. This historical measurement is not bound to a verified official v0.15.1 binary identity and should not be read as a result for the published release artifact. The runs used preallocated Int32Array / Float64Array inputs and outputs and preallocated WebAssembly memory; the timed region covered repeated kernel calls, not compilation, module instantiation, or memory setup. Multipliers are JavaScript elapsed time divided by CK Wasm elapsed time.
These are local measurements for two vectorizable kernels, not a general guarantee across algorithms, hardware, or runtimes. A floating-point AXPY loop with a loop-carried checksum was roughly tied with JavaScript in a separate baseline comparison. Measure representative workloads with the same inputs and deployment runtime.
For a broader measured comparison, see the eleven-algorithm performance page. It includes CK Native, CK Wasm on Node.js, C++, Rust, Java, and Node.js JavaScript, with JavaScript set to 1.0× for each workload.
Upgrade notes#
- Check the schema 2
featuresprofile andprofile_sha256digest, then ensure the deployment runtime supports that profile's v0.15 capability set. - SIMD128 modules require a runtime with SIMD support; keep a non-SIMD target when deployments include runtimes without that capability.
- The exported caller-owned-memory contract has not changed.
Download CalcKernel from the v0.15.1 GitHub release, or review the public compiler repository.