Marrow provides SIMD-vectorized compute kernels for arithmetic, comparisons, selection and casting. All kernels are null-aware; the exact null handling depends on the operation.
String, temporal, boolean and conditional kernels are implemented in Mojo but not yet bound to Python — they are reachable from a query plan instead.
The compute functions live in marrow.compute, and their names and signatures follow pyarrow.compute for the functions it implements — mc.add, mc.subtract, mc.cast, mc.filter. It is not a drop-in replacement: some options (skip_nulls=False, nan_is_null=True, multi-key sort_keys) raise NotImplementedError, and memory_pool is ignored.
Notemarrow.compute compares like pyarrow; marrow.expr compares like SQL
The two answer differently on NaN, deliberately. marrow.compute mirrors pyarrow.compute, so its comparisons are IEEE — a NaN equals nothing, itself included. marrow.expr is a SQL engine, and DuckDB, DataFusion and Polars all agree that SQL’s comparisons are total:
expression
pc.equal / pc.greater
col("a") == … / > …
nan = nan
False
True
nan <> nan
True
False
nan > 1.0
False
True
So a NaN is its own GROUP BY group, its own join key and its own rank peer, and it sorts after inf. NULL is unaffected and stays three-valued in both.
Float division by zero answers a value — 10.0 / 0.0 is inf, 0.0 / 0.0 is nan — while integer // and % by zero answer NULL, as SQL does. The two genuinely want different answers: division by zero has a value in the reals’ completion and integer division by zero does not.
Null propagation
If either operand at a position is null, the result at that position is null. This mirrors SQL’s three-valued logic.
a = ma.array([1, None, 3, None])b = ma.array([10, 20, 30, 40])result = mc.add(a, b)print(result) # index 1 and 3 are nullprint("null count:", result.null_count)
# Both inputs can contribute nullsa = ma.array([None, 2, 3, None])b = ma.array([10, None, 30, None])print(mc.add(a, b)) # null at 0, 1, and 3
PrimitiveArray[int64]([NULL, NULL, 33, NULL])
Aggregates
Most aggregates are not eager kernel calls (boolean any and all, below, are the exception). sum, mean, min, max, product, count, count_distinct and the rest reduce a column inside a query plan, so that a whole-table reduction and a GROUP BY are the same code path. Wrap a batch with ma.memtable(), describe the reduction, and .collect() runs it:
sum and product do not preserve the input type — integers accumulate in int64 and floats in float64, so a long column cannot silently overflow its own width. mean is always float64:
count counts non-null values; ma.count_star() counts rows, which differs on a nullable column. count_distinct is exact, while approx_count_distinct uses a fixed-size HyperLogLog sketch — far cheaper in memory on a high-cardinality column: