Marrow reads and writes two on-disk formats: the Arrow IPC format (both the file and stream framings) for zero-overhead round trips of the in-memory layout, and Parquet for compact, columnar long-term storage.
Arrow IPC
The Arrow IPC format serialises RecordBatches with essentially no transformation of the columnar buffers, so writing and reading it back is fast and lossless across every implemented type — nested, dictionary, temporal and null columns included.
The IPC file format is seekable and carries a footer, ideal for random access. The IPC stream format is a flat sequence of batches, ideal for pipes and sockets. The API mirrors PyArrow:
Marrow ships a from-scratch Parquet reader and writer — it decodes and encodes the Parquet format itself, with no PyArrow at runtime. The Python API lives in marrow.parquet and mirrors pyarrow.parquet:
The writer also emits column statistics (min/max/null/distinct counts) and a page index, so readers on the other side can prune row groups and pages. Bloom filters are implemented but off by default and not currently reachable from Python. write_table accepts any Arrow-C-stream object too, so a PyArrow table can be written directly.
Note
The Mojo API mirrors the Python one — from marrow.parquet import read_table, write_table operate on a marrow Table, and the marrow/parquet package also exposes lower-level metadata, statistics, page-index and bloom-filter readers.
Choosing a format
Arrow IPC — you want the fastest possible read/write and are staying inside the Arrow ecosystem. The bytes on disk are the columnar buffers.
Parquet — you want compression, column pruning and row-group / page statistics for archival or interchange with data-lake tooling.