import marrow as ma
from marrow import compute as mcQuickstart
From zero to arrays, compute and tables
This tour takes you through the essentials of Marrow from Python: building arrays, running compute kernels, and assembling columns into a table. Every snippet on this page is executed when the docs are built, so the output you see is real.
Install
Marrow is a Mojo project with a Python binding compiled to a native extension. Clone the repository and build the extension with pixi:
git clone https://github.com/kszucs/marrow
cd marrow
pixi run build_python # compiles the native extension (python/marrow/libmarrow.so)Then bring it in under the customary ma alias:
The API follows PyArrow’s names where it can: pa.array, pa.record_batch and pa.table have direct counterparts, though marrow implements far less of the API.
Create an array
ma.array() turns a Python list into a columnar array. The type is inferred, and None marks a null position:
a = ma.array([1, 2, 3, None, 5])
print(a)
print("length: ", len(a))
print("null count:", a.null_count)
print("type: ", a.type)PrimitiveArray[int64]([1, 2, 3, NULL, 5])
length: 5
null count: 1
type: int64
Pass type= to skip inference and coerce the data to a specific type:
f = ma.array([1, 2, 3], type=ma.float32())
print(f)PrimitiveArray[float32]([1.0, 2.0, 3.0])
Strings, nested lists and structs all work the same way:
print(ma.array(["hello", None, "world"]))
print(ma.array([[1, 2], [3, 4, 5], None]))
print(ma.array([{"x": 1, "y": 1.5}, {"x": 2, "y": 2.5}]))StringArray([hello, NULL, world])
ListArray([PrimitiveArray[int64]([1, 2]), PrimitiveArray[int64]([3, 4, 5]), NULL])
StructArray({'x': PrimitiveArray[int64]([1, 2]), 'y': PrimitiveArray[float64]([1.5, 2.5])})
Compute on arrays
Compute kernels live in marrow.compute (imported here as mc), vectorized and null-aware, with names that mirror pyarrow.compute. A null in either operand propagates to the result:
x = ma.array([1, 2, 3, None, 5])
y = ma.array([10, 20, 30, 40, 50])
print("add: ", mc.add(x, y)) # index 3 is null
print("multiply:", mc.multiply(y, y))add: PrimitiveArray[int64]([11, 22, 33, NULL, 55])
multiply: PrimitiveArray[int64]([100, 400, 900, 1600, 2500])
Most aggregates are not on the eager surface (boolean mc.any and mc.all are). sum, mean, min, max, count_distinct and the rest reduce a column inside a query plan, built with ma.memtable(...) / ma.read_parquet(...), so that a group-by and a whole-table aggregate are the same code path.
Comparisons return a boolean array, which you can feed straight into filter:
big = mc.greater(y, ma.array([15, 15, 15, 15, 15]))
print("mask: ", big)
print("kept: ", mc.filter(y, big))
print("no nulls:", mc.drop_null(x))mask: BoolArray([False, True, True, True, True])
kept: PrimitiveArray[int64]([20, 30, 40, 50])
no nulls: PrimitiveArray[int64]([1, 2, 3, 5])
Sorting handles nulls explicitly and returns either the sorted values or the sort indices:
s = ma.array([3, 1, None, 2, 5])
print("sorted: ", mc.sort(s))
print("desc: ", mc.sort(s, [("", "descending")], null_placement="at_end"))
print("argsort: ", s.argsort())sorted: PrimitiveArray[int64]([1, 2, 3, 5, NULL])
desc: PrimitiveArray[int64]([5, 3, 2, 1, NULL])
argsort: PrimitiveArray[int32]([1, 3, 0, 4, 2])
Assemble a table
A RecordBatch is a schema plus a set of equal-length columns — the workhorse tabular structure:
people = ma.record_batch({
"name": ma.array(["Alice", "Bob", "Carol", "Dave"]),
"age": ma.array([34, 28, 45, 52]),
"salary": ma.array([95_000, 72_000, 120_000, 88_000]),
})
print(people)
print("shape:", people.shape)RecordBatch(num_rows=4, schema=Schema(fields=[name: string, age: int64, salary: int64]))
shape: (4, 3)
Inspect it as native Python with to_pydict() (column-oriented) or to_pylist() (row-oriented):
print(people.to_pydict())
print(people.to_pylist()){'name': ['Alice', 'Bob', 'Carol', 'Dave'], 'age': [34, 28, 45, 52], 'salary': [95000, 72000, 120000, 88000]}
[{'name': 'Alice', 'age': 34, 'salary': 95000}, {'name': 'Bob', 'age': 28, 'salary': 72000}, {'name': 'Carol', 'age': 45, 'salary': 120000}, {'name': 'Dave', 'age': 52, 'salary': 88000}]
Select columns, slice rows, and sort by one or more keys:
print(people.select(["name", "salary"]).to_pydict())
print(people.slice(1, 2).to_pylist())
print(people.sort_by([("age", "descending")]).to_pylist()){'name': ['Alice', 'Bob', 'Carol', 'Dave'], 'salary': [95000, 72000, 120000, 88000]}
[{'name': 'Bob', 'age': 28, 'salary': 72000}, {'name': 'Carol', 'age': 45, 'salary': 120000}]
[{'name': 'Dave', 'age': 52, 'salary': 88000}, {'name': 'Carol', 'age': 45, 'salary': 120000}, {'name': 'Alice', 'age': 34, 'salary': 95000}, {'name': 'Bob', 'age': 28, 'salary': 72000}]
Where to go next
- Build a data pipeline — a worked example that filters, sorts and joins two datasets.
- Arrays — every array type in depth.
- Compute kernels — the full kernel catalogue and null semantics.
- Tables & joins —
RecordBatch,Tableand hash joins.