Home/Documentation/Performance

Performance

How CK turns focused calculations into efficient code.

CK targets numerical kernels: compact, compute-heavy routines embedded in larger applications. The new v0.15.1 comparison runs eleven fixed algorithms in CK Native, C++, Rust, Java, JavaScript on Node.js, and CK WebAssembly called from Node.js. The earlier two-workload comparison remains below as a separate historical measurement. These results compare specific implementations and build choices, not languages as a whole.

v0.15.1: eleven algorithms, six implementations#

The home-page selector lists these workloads by CK Native speedup over JavaScript. Within each chart, displayed speedups run fastest first; CK is placed first when rounded values tie.

The home-page selector defaults to the top v0.15.1 workload, sum. Each of its eleven workload options shows five rows from the original six-result v0.15.1 report—CK Native, C++, Rust, Java, and Node.js JavaScript—and a strict WASM chart below, with baseline and SIMD128 profile tabs. The raw v0.15.1 report still includes CK WASM called from Node.js; strict WASM separately measures CK, Clang, and Rust. A shared note below both charts labels their references: JavaScript = 1.0× for v0.15.1 and the faster compatible same-profile Clang or Rust route = 1.0× for strict WASM. Their times and ratios are not directly comparable. Strict timings were not measured with the official v0.15.2 binary. The two earlier historical workloads remain separate options.

The CK Wasm row in the v0.15.1 suite uses O3 and the simd128 feature profile with strict floating-point behavior and unchecked overflow and bounds. The profile permits SIMD instructions; individual kernels may remain scalar. Its timed samples are warmed, repeated Node.js-to-Wasm calls, including host dispatch and required output or workspace initialization. Module instantiation, memory setup, fixture reads and copies, compilation, and output hashing are outside that timer.

In the v0.15.1 charts, the home-page selector shows each algorithm with JavaScript on Node.js fixed at 1.0×. Every multiplier there is the JavaScript median time divided by that implementation's median time for the same algorithm and inputs. A value below 1.0× means slower than JavaScript. The raw report includes every timed sample, build and runtime identities, input and output hashes, and the measurement order. No overall multiplier is calculated across different algorithms.

The workloads are element-wise fused arithmetic, u32 reduction/sum, f64 dot product, 3×3 Gaussian image stencil, row–inner–column matrix multiplication, min–max normalization, degree-seven polynomial evaluation, deterministic Monte Carlo π sampling, a four-branch piecewise kernel, a 128 MiB streaming transform, and dense-graph O(V²) Dijkstra shortest path. “Fused” describes combining arithmetic in one traversal; it does not promise a hardware fused multiply-add. The streaming transform size does not, by itself, prove that memory bandwidth is saturated. Matrix multiplication uses direct kernels, with no BLAS or NumPy.

For workloads with array inputs, each implementation receives the same prebuilt little-endian fixture bytes. Monte Carlo instead starts from the same deterministic seed and generates its points inside the timed kernel. Timed work is one full kernel calculation, including any required output/workspace reset. Compilation, process start, fixture loading, Wasm instantiation, memory growth and input copy, and output hashing are outside the timer. The CK Wasm result includes the JavaScript-to-Wasm call inside Node.js. The compiler uses one .ck source for CK Native and CK Wasm; Wasm is built with the O3 SIMD128 profile. That profile permits SIMD on eligible loops; strict floating-point reductions and other loops may still use scalar instructions. JavaScript uses typed arrays, Java uses a warmed OpenJDK 21 process, and every result is checked against an independent complete-output reference.

See the complete v0.15.1 measurements and benchmark sources and exact reproduction commands. Results are specific to the recorded Apple M5 Max system, toolchain versions, algorithms, and input sizes. Measure a representative workload on its target machine before applying them elsewhere.

Strict same-target WebAssembly comparison#

The primary strict WASM report measures all eleven fixed kernels with CK, Clang, and Rust under separate baseline and simd128 profiles. In each home-page workload panel, the added strict chart has tabs for those two profiles and reports median time and throughput relative to the faster compatible Clang or Rust route. Its report identity and reference scale remain separate from the five-visible-row v0.15.1 chart above it. Run 1 measured baseline geometric mean 1.117× (minimum 0.920×; 0/11 below 0.90×) and SIMD128 1.016× (minimum 0.923×; 0/11 below 0.90×). A second complete run of the same ckc binary measured baseline 1.118× (minimum 0.914×; 0/11 below 0.90×) and SIMD128 1.008× (minimum 0.920×; 0/11 below 0.90×). Both runs independently meet the 0.95 geometric-mean and 0.90 per-kernel targets in both profiles; run results are not pooled. Download the primary source bundle, the second raw report, and its source bundle. The source bundle README documents the published source and reproduction commands.

The measured compiler is a development candidate whose version string is ckc 0.15.1; it is neither the official v0.15.1 release artifact nor the official v0.15.2 binary. Its exact binary SHA-256 is 5db7c4db4146dda00020941f1047a1ade794741374e9446af4b2669de8979e3b, built from and measured against clean source checkout 08292f18b6374f9636bae5ab07fbfcb8e9e30091. The raw report says the binary-to-source relationship is not independently verifiable, so the exact binary hash is authoritative. Run 1 report SHA-256: ed56abd762d3d02404ae92b09ce213f8e8ef95f378020dea0f4b7355f302351e; run 2 report SHA-256: 4384dc2262a4a967a40d1b6f04096235def2b51408c12ae9c363c7946243378b. Both were measured on Apple M5 Max, macOS 26.6.2 (Darwin 25), arm64, with Clang 22.1.8, Rust 1.90.0, and Node.js 24.14.0. The reports share the same ckc binary and source hashes, but each preserves its own artifact hashes: the Rust baseline and SIMD modules differ by eight bytes because the generated build path is embedded; their feature/opcode probes and full-output checks agree. This development measurement does not change the identity of the v0.15.1 release. Earlier report runs remain accessible: dc75e053d02e run 1, dc75e053d02e repeat, dc61ebdb37ee run 1, and dc61ebdb37ee repeat, each with a report-specific source bundle. The old dc61 source ZIP mistakenly embedded the dc75 report bytes; its replacement now embeds the dc61 JSON matching its manifest. The dc75 primary source ZIP remains byte-for-byte unchanged. No historical report JSON or hash-verified benchmark source was modified; the repeat reports were not previously included in website build output.

Both profiles use O3 and the same deterministic inputs, dimensions, algorithms, and iteration order. baseline selects MVP, MULTI_VALUE, and BULK_MEMORY without SIMD; simd128 adds SIMD128, with Relaxed SIMD disabled. All routes use strict floating-point semantics: no fast math, reassociation, or FMA contraction. Finite f64 outputs are compared bit-for-bit, including signed zero and subnormals; NaNs compare by class and u32 outputs compare exactly. When the reusable workspace is initialized, the host checks that input, output, and workspace regions are pairwise disjoint; the fixed ranges are then reused for calls. This satisfies CK noalias, C++ __restrict, and Rust's unsafe disjoint-region precondition. Memory bounds and integer overflow are explicitly unchecked in this WASM ABI.

The report compares implementations within the selected capability profiles, while preserving observed module requirements and declaration differences. Restricted validation of the six hashed modules found CK and Clang baseline valid with MVP, Rust baseline valid with MVP plus bulk memory, CK and Clang SIMD valid with MVP plus SIMD128, and Rust SIMD valid with MVP plus bulk memory and SIMD128; both Rust modules contain real memory.fill instructions. The external declarations still differ: Clang declares bulk-memory-opt; Rust declares bulk-memory-opt and mutable-globals. No standalone mutable-global use, multi-value instruction, or Relaxed SIMD use was observed. These results do not make the declarations identical.

Hot timings include required output/workspace initialization, repeat-loop and language-call overhead, including Node-to-Wasm calls. They exclude compilation, process startup, fixture I/O/copy, Wasm instantiation and memory growth, output hashing, and report generation. The cold-cost table records artifact size, one module-compilation observation per profile worker, and medians across eleven single per-workload instantiation and first-call observations; those measurements exclude process or browser startup. Individual cold first calls show a material gap: baseline matmul was 43.954 ms for CK vs 13.376 ms for Clang, and Dijkstra was 26.896 ms for CK vs 4.458 ms for Rust; SIMD128 matmul was 20.578 ms for CK vs 8.850 ms for Clang, and Dijkstra was 27.117 ms for CK vs 5.595 ms for Clang. CK modules were 28,325 bytes (baseline) and 34,941 bytes (SIMD128), compared with Clang at 4,683 and 4,821 bytes. These are cold observations and artifact sizes, not part of the hot geometric means or per-kernel acceptance scores. The raw reports retain every sample, build command, toolchain identity, and cold observation.

Historical same-machine cross-runtime view (two workloads)#

This chart preserves the earlier crosslang-benchmark-m5-max.json report, measured on 2026-09-26 on an Apple M5 Max. It predates the v0.15.1 suite and contains no CK Wasm measurement. The raw report remains downloadable from the chart for its original samples and identity.

For each workload, the home page selects the tuned variant for every displayed implementation: five for matrix multiplication and six for image blur. Each displayed multiplier is the JavaScript tuned median divided by the selected tuned median, rounded for display; JavaScript is therefore 1.0×, and the bars are sorted fastest first. Bar lengths use a logarithmic visual scale to keep large differences readable, so do not read their lengths as linear ratios. This historical cross-runtime view combines specific Native implementations and host runtimes with different tuning methods; it does not compare all implementations at one optimization level or measure Wasm parity. The paired ordinary and tuned results for the displayed implementations are below.

Each cell is the median time for one complete calculation in milliseconds; lower is faster. “Ordinary” and “tuned” name the specific checked-in variants in this suite. They do not describe universal levels of language optimization.

256 × 256 matrix multiplication#

Implementation Ordinary (ms) Tuned (ms)
CK 9.315268 1.869669
C++ 9.329641 1.866093
Rust 9.685109 1.893518
JavaScript (Node.js) 11.875417 10.087641
Java (OpenJDK 21) 9.734412 3.116292

1024 × 1024 Gaussian convolution (3 × 3)#

Implementation Ordinary (ms) Tuned (ms)
CK 1.032351 0.597181
C++ 0.533230 0.532288
Rust 0.552982 0.547729
JavaScript (Node.js) 7.873651 2.203127
Java (OpenJDK 21) 1.595859 1.041701
NumPy 3.306759 3.443799

These kernels expose different performance characteristics. For matrix multiplication, tuned CK, C++, and Rust take about 1.87–1.89 ms, Java 3.116 ms, and JavaScript 10.088 ms. For image convolution, tuned C++ and Rust take about 0.53–0.55 ms, tuned CK 0.597 ms, Java 1.042 ms, and JavaScript 2.203 ms. NumPy's tuned vector-expression version takes 3.444 ms, slightly longer than its ordinary version at 3.307 ms. These results show how algorithm shape and implementation change the outcome; they do not rank languages in general.

What was measured#

The host was an Apple M5 Max running Darwin 25.6.0 on arm64. Toolchains were the CK compiler identified in the raw measurement report, Apple Clang 21.0.0, Rust 1.90.0, Node.js 24.14.0, OpenJDK 21.0.8, Python 3.12.14, and NumPy 2.3.5. The workloads were a 256 × 256 f64 matrix product and a 1024 × 1024 f64 image processed with a 3 × 3 Gaussian filter. The raw report also retains a NumPy matrix measurement, but it is excluded from the displayed matrix comparison because it calls Apple's optimized Accelerate library. For image convolution, the ordinary NumPy case uses vector expressions that create intermediate arrays; the tuned case reuses preallocated scratch arrays. C++, Rust, and CK used O3 builds; their ordinary variants target a portable CPU baseline, while tuned variants enable native CPU features and use a row-contiguous matrix loop. JavaScript and Java run in persistent, warmed-up workers; Java uses OpenJDK 21.

Each variant ran in seven interleaved rounds; the reported value is the median. Compilation, process startup, input setup, initialization, and output hashing were outside the timed interval. Persistent JavaScript, Java, and NumPy workers include their language-level repeat loop and per-call function dispatch inside timed batches. Before timing, the complete output buffer from every implementation was SHA-256 checked against an independent reference. All twelve raw variants matched the reference for each workload; the displayed comparison includes ten matrix and twelve convolution variants. Matrix multiplication reference: a70e525b589edae101181e5ad5f5acc36c8ce2e77fd54a7222a15faf9b4a62ac. Convolution reference: 7744d264c90a3103680087a715d70ab36a8effc8e9a93e58ec0f0f85dbb4c71a.

See the raw M5 Max measurements, the reproducible benchmark method, or view the chart on the home page.

The results depend on the exact algorithms, compiler and runtime versions, machine, and input sizes. The displayed matrix variants are direct kernel implementations; the NumPy/Accelerate result remains only in the raw report for audit. Treat these numbers as a reproducible comparison of this suite, not a prediction for another application. Run your own representative workload on its target machine before drawing a performance conclusion.

For a separate historical measurement on AMD EPYC, see the archived CI report from September 2026. It used different workloads and a different machine, so its numbers should not be combined with the chart above.

Why CK can run efficiently#

Native machine code#

ckc run and ckc build compile a .ck program into native machine code for the target architecture. The standard build uses a portable CPU baseline; CPU-specific variants require an explicit option. The program runs without a CK interpreter in its calculation path. Optimized O3 builds are the default for these commands; O3 is a compiler setting, not a speed multiplier.

Optimizations with defined conditions#

On suitable loops, CK may process independent values together using SIMD (single instruction, multiple data), or reduce repeated work. It applies a transformation only when its correctness conditions and the target processor permit it. Loops with dependencies or unsupported shapes keep their ordinary execution path.

Optional tuning for a real workload#

For a completed program with a representative workload, PGO can use observed execution patterns to guide code generation. It adds build steps and is useful only when the measured workload resembles real use. Start with the standard build, and use PGO or CPU variants only when measurements show they help your application.

Keep learning#

Repository reference links follow the main branch and may describe features newer than the latest downloadable release.

↵ open · esc close