marrow marrow
  • Home
  • Tutorials
  • Guide
  • Reference
  • Benchmarks
ExperimentalOpen source, Apache 2.0

Apache Arrow,
written in Mojo.

Columnar arrays, null-aware compute kernels, Parquet and Arrow IPC, and a small query engine. Use it from Python much as you would PyArrow, or build a Mojo query into its own executable. It covers a useful part of Arrow, not all of it.

Read the quickstart GitHub
  • Shares buffers with PyArrow
  • Reads and writes Parquet
  • Answers checked against DuckDB
orders.py
import marrow as ma
from marrow import col, lit

(ma.read_parquet("orders.parquet")
   .filter(col("price") > lit(15))
   .aggregate(by=["region"], total=("sum", "price"))
   .order_by(("total", "descending"))
   .head(10)
   .optimize()
   .collect())

Nothing runs until collect(). optimize() rewrites the plan first; it is opt-in.

215 of 278 SQL queries in the test corpus answered exactly as DuckDB answers them. The other 63 are recorded too, and wait on features marrow does not have yet.
3 Arrow implementations (C++, Rust, Go) it exchanges IPC files and C Data with in Arrow's integration suite, for the layouts it implements.
3.0 MB Machine code in a compiled Parquet query on macOS arm64, held there by a CI size gate. The executable also loads the Mojo runtime libraries.
Three ways in

Call a kernel, build a plan, or compile it

The same arrays and kernels sit under all three. What changes is how much is decided before anything runs.

  • Eager Python
  • Lazy Python
  • Compiled Mojo

Names follow PyArrow’s. Each call runs its kernel straight away.

import marrow as ma

batch = ma.record_batch({
    "region": ma.array(["east", "west", "east"]),
    "price":  ma.array([10, 20, 30]),
})

hot = ma.compute.greater(batch.column("price"), ma.array([15, 15, 15]))
ma.compute.filter(batch.column("price"), hot)

Describe the query first and run it with collect(). Add optimize() to let the plan rewriter push filters down and prune columns before it runs.

import marrow as ma
from marrow import col, lit

ma.read_parquet("orders.parquet") \
  .filter(col("price") > lit(15)) \
  .aggregate(by=["region"], total=("sum", "price")) \
  .optimize() \
  .collect()

Column types are fixed when the query is compiled, so an expression can fuse into a single loop, and kernels the query never uses are not linked in. Every param in the plan becomes a command-line flag.

from marrow.dtypes import field, int64, string
from marrow.expr import DynRelation, QueryCli, col, param, scan
from marrow.schema import schema


def query() raises -> DynRelation:
    var orders = scan(
        param("src", string),
        schema(
            [field("id", int64), field("amount", int64), field("name", string)]
        ),
    )
    return orders.filter(col("amount", int64) >= param("min-amount", int64))


def main() raises:
    QueryCli(query()).run()
$ marrow compile query.mojo -o orders
$ ./orders --src orders.parquet --min-amount 250

Building takes a minute or two, so this suits a query you run many times, not one you are still shaping. --bundle copies the executable together with the runtime libraries it loads, for running it on another machine.

Where it fits

Next to the tools you already use

Each of these is more complete and better tested than marrow. The right column is what marrow offers anyway.

What it is What marrow offers
PyArrow Python bindings to Arrow C++, the reference implementation Similar method names, and arrays pass between the two without copying their buffers
Polars / DuckDB Mature query engines with far broader features A fixed query compiled into its own executable. No benchmark results are published yet, so no speed claims
arrow-rs / DataFusion Arrow and a query engine in Rust, the closest relatives in design Its IPC and C Data output is tested against arrow-rs, C++ and Go in Arrow’s integration suite
Mojo The language marrow is written in Arrow arrays, Parquet and a query engine for Mojo programs; a few element-wise kernels can also run on a GPU behind a build flag
NoteHow correctness is checked

A corpus of 278 SQL queries whose expected answers come from DuckDB, never from marrow: 215 run and match, and 63 wait on features marrow does not have yet. Arrow’s own integration suite round-trips IPC files and C Data with the C++, Rust and Go implementations, skipping the layouts marrow does not implement (unions, views, run-end encoding, extension types). See Status & limitations for the full list.

See it run

A short tour

These cells run when the site is built, so the output below is real.

Build arrays from Python data. The type is inferred, and None marks a null:

ages   = ma.array([25, 35, 45, None, 55])
names  = ma.array(["Alice", "Bob", "Carol", "Dave", "Eve"])
salary = ma.array([70_000, 90_000, 110_000, 50_000, 130_000])

print(ages)
print(names)
PrimitiveArray[int64]([25, 35, 45, NULL, 55])
StringArray([Alice, Bob, Carol, Dave, Eve])

Kernels are vectorized and null-aware; a null in either operand gives a null:

print("next year:", ma.compute.add(ages, ma.array([1] * 5)))
print("over 40:  ", ma.compute.greater(ages, ma.array([40] * 5)))
next year: PrimitiveArray[int64]([26, 36, 46, NULL, 56])
over 40:   BoolArray([False, False, True, NULL, True])

Assemble columns into a RecordBatch, then query it lazily:

people = ma.record_batch({"age": ages, "name": names, "salary": salary})

(ma.memtable(people)
   .filter(col("salary") > lit(60_000))
   .aggregate(headcount=("count", "salary"), payroll=("sum", "salary"))
   .collect()
   .to_pylist())
[{'headcount': 4, 'payroll': 400000}]

Arrays implement the Arrow PyCapsule interface, so PyArrow takes them without copying the buffers, and marrow takes PyArrow’s the same way:

import pyarrow as pa
pa_ages = pa.array(ages)          # marrow → PyArrow, no copy
back    = ma.array(pa_ages)       # PyArrow → marrow, no copy
print(pa_ages.type, "|", len(back))
int64 | 5
Keep going

Where to next

Quickstart →

Build the extension, then try arrays, kernels and tables.

Lazy queries →

Build a plan in Python, optionally optimize it, then run it.

Compile a query →

Turn a Mojo query into a native program whose parameters are command-line flags.

Status →

What is implemented, what is missing, and what is known to be wrong.

Back to top

marrow — Apache Arrow in Mojo

 
  • Quickstart

  • Guide

  • Status