Comptime vs runtime lane

The same query in both lanes, measured

Queries in Mojo describes two expression lanes under one plan IR: col("a", int64) resolves the operand’s dtype at compile time and fuses a subtree into one SIMD loop, while col("a") resolves it at run time and materialises every intermediate column. This page measures what that difference is worth, query by query.

The comptime lane is faster on 8 of 8 queries; speedups range from 1.05× to 9.22×.

Mean over each case’s rounds, 1,000,000 rows. Measured on Apple M4 Max (Darwin 24.6.0, arm64) with Mojo 1.2.0.dev2026092105 (e9569894), at commit 35db696a on 2026-09-27.
Query Comptime Runtime Speedup
aggregate_sum 1.07 ms 1.88 ms 1.75×
aggregate_sum_grouped_1 8.76 ms 9.93 ms 1.13×
aggregate_sum_grouped_100 7.90 ms 8.27 ms 1.05×
aggregate_sum_grouped_4 7.23 ms 9.44 ms 1.31×
filter_computed 3.80 ms 15.57 ms 4.09×
filter_range 4.05 ms 12.24 ms 3.03×
project_float_arithmetic 0.72 ms 6.62 ms 9.22×
project_int_arithmetic 0.91 ms 2.21 ms 2.44×

What is compared

Every row is a pair of benchmarks in marrow/expr/tests/bench_comptime.mojo: bench_comptime_<query> and bench_runtime_<query> build the same plan over the same batch and return the same answer — two contenders, compared the way --competition compares libraries. The only thing that differs is how the leaves are spelled:

table(batch).filter((col("a", int64) > lit(200, int64)) & (col("a", int64) < lit(800, int64)))
table(batch).filter((col("a") > lit(Int64Scalar(200))) & (col("a") < lit(Int64Scalar(800))))

The time covers the whole plan — building it, lowering it to operators and draining it — so it is what a caller of .execute() pays.

Every query is one the comptime lane can fuse: it runs the expression as one loop over the input, where the runtime lane evaluates a * b into a column, then the comparison into another, and so on. An aggregate neither lane can fuse, such as count(DISTINCT a), is left out on purpose: both lanes call the same kernel and the times tie.

The grouped rows win by less, and that is the grouping rather than the fusion. Fusion saves about the same absolute time there as in the whole-table sum, but every row’s key is hashed and probed whatever the number of groups, and both lanes pay that alike. A single group is no cheaper: every row then updates the same slot, one after another.

Reproducing it

pixi run -e dev bench-comptime

That runs the pairs with --competition-winner comptime, which fails the run unless the comptime lane is faster on every query — a tie, or a query only one lane ran, counts as a loss — and --competition-json docs/data/comptime.json, which is what this page renders. The Benchmarks workflow runs the same check on every pull request.

Timings on a shared machine are noisy, and the numbers above are only as good as the machine they were recorded on.

Back to top