Comptime vs runtime lane
The same query in both lanes, measured
Queries in Mojo describes two expression lanes under one plan IR: col("a", int64) resolves the operand’s dtype at compile time and fuses a subtree into one SIMD loop, while col("a") resolves it at run time and materialises every intermediate column. This page measures what that difference is worth, query by query.
The comptime lane is faster on 8 of 8 queries; speedups range from 1.05× to 9.22×.
| Query | Comptime | Runtime | Speedup |
|---|---|---|---|
aggregate_sum |
1.07 ms | 1.88 ms | 1.75× |
aggregate_sum_grouped_1 |
8.76 ms | 9.93 ms | 1.13× |
aggregate_sum_grouped_100 |
7.90 ms | 8.27 ms | 1.05× |
aggregate_sum_grouped_4 |
7.23 ms | 9.44 ms | 1.31× |
filter_computed |
3.80 ms | 15.57 ms | 4.09× |
filter_range |
4.05 ms | 12.24 ms | 3.03× |
project_float_arithmetic |
0.72 ms | 6.62 ms | 9.22× |
project_int_arithmetic |
0.91 ms | 2.21 ms | 2.44× |
What is compared
Every row is a pair of benchmarks in marrow/expr/tests/bench_comptime.mojo: bench_comptime_<query> and bench_runtime_<query> build the same plan over the same batch and return the same answer — two contenders, compared the way --competition compares libraries. The only thing that differs is how the leaves are spelled:
table(batch).filter((col("a", int64) > lit(200, int64)) & (col("a", int64) < lit(800, int64)))
table(batch).filter((col("a") > lit(Int64Scalar(200))) & (col("a") < lit(Int64Scalar(800))))The time covers the whole plan — building it, lowering it to operators and draining it — so it is what a caller of .execute() pays.
Every query is one the comptime lane can fuse: it runs the expression as one loop over the input, where the runtime lane evaluates a * b into a column, then the comparison into another, and so on. An aggregate neither lane can fuse, such as count(DISTINCT a), is left out on purpose: both lanes call the same kernel and the times tie.
The grouped rows win by less, and that is the grouping rather than the fusion. Fusion saves about the same absolute time there as in the whole-table sum, but every row’s key is hashed and probed whatever the number of groups, and both lanes pay that alike. A single group is no cheaper: every row then updates the same slot, one after another.
Reproducing it
pixi run -e dev bench-comptimeThat runs the pairs with --competition-winner comptime, which fails the run unless the comptime lane is faster on every query — a tie, or a query only one lane ran, counts as a loss — and --competition-json docs/data/comptime.json, which is what this page renders. The Benchmarks workflow runs the same check on every pull request.
Timings on a shared machine are noisy, and the numbers above are only as good as the machine they were recorded on.