▼1.4 IO
1.4.1 CSV
1.4.2 Multiple files
On this page:
1.4.2.1 Reading into a single dataframe
1.4.2.2 Reading and processing in parallel

1.4.2 Multiple files🔗ℹ

Polars can deal with multiple files differently depending on your needs and memory strain. Let’s create some files to give us some context:

> (define df
    (dataframe (list (series '(1 2 3) #:name "foo")
                     (series (list polars-null "ham" "spam") #:name "bar"))))
> (for ([i (in-range 5)])
    (write-csv df (build-path dir (format "my_many_files_~a.csv" i))))
1.4.2.1 Reading into a single dataframe🔗ℹ

To read multiple files into a single dataframe, use a glob pattern. The files are read separately and stacked in filename order.

> (read-csv (build-path dir "my_many_files_*.csv"))

shape: (15, 2)

┌─────┬──────┐

│ foo ┆ bar  │

│ --- ┆ ---  │

│ i64 ┆ str  │

╞═════╪══════╡

│ 1   ┆ null │

│ 2   ┆ ham  │

│ 3   ┆ spam │

│ 1   ┆ null │

│ 2   ┆ ham  │

│ …   ┆ …    │

│ 2   ┆ ham  │

│ 3   ┆ spam │

│ 1   ┆ null │

│ 2   ┆ ham  │

│ 3   ┆ spam │

└─────┴──────┘

read-parquet and scan-parquet take a pattern the same way:

> (for ([i (in-range 2)])
    (write-parquet df (build-path dir (format "my_many_files_~a.parquet" i))))
> (height (read-parquet (build-path dir "my_many_files_*.parquet")))

6

API gap: no show_graph, so the query plan cannot be drawn.

1.4.2.2 Reading and processing in parallel🔗ℹ

If your files don’t have to be in a single table, build a query plan for each file.

> (for/list ([file (sort (glob (build-path dir "my_many_files_*.csv")) path<?)])
    (~> (scan-csv file)
        (group-by "bar")
        (agg (alias (count "foo") "len") (sum "foo"))
        (sort "bar")
        collect))

'(shape: (3, 3)

┌──────┬─────┬─────┐

│ bar  ┆ len ┆ foo │

│ ---  ┆ --- ┆ --- │

│ str  ┆ u32 ┆ i64 │

╞══════╪═════╪═════╡

│ null ┆ 1   ┆ 1   │

│ ham  ┆ 1   ┆ 2   │

│ spam ┆ 1   ┆ 3   │

└──────┴─────┴─────┘ shape: (3, 3)

┌──────┬─────┬─────┐

│ bar  ┆ len ┆ foo │

│ ---  ┆ --- ┆ --- │

│ str  ┆ u32 ┆ i64 │

╞══════╪═════╪═════╡

│ null ┆ 1   ┆ 1   │

│ ham  ┆ 1   ┆ 2   │

│ spam ┆ 1   ┆ 3   │

└──────┴─────┴─────┘ shape: (3, 3)

┌──────┬─────┬─────┐

│ bar  ┆ len ┆ foo │

│ ---  ┆ --- ┆ --- │

│ str  ┆ u32 ┆ i64 │

╞══════╪═════╪═════╡

│ null ┆ 1   ┆ 1   │

│ ham  ┆ 1   ┆ 2   │

│ spam ┆ 1   ┆ 3   │

└──────┴─────┴─────┘ shape: (3, 3)

┌──────┬─────┬─────┐

│ bar  ┆ len ┆ foo │

│ ---  ┆ --- ┆ --- │

│ str  ┆ u32 ┆ i64 │

╞══════╪═════╪═════╡

│ null ┆ 1   ┆ 1   │

│ ham  ┆ 1   ┆ 2   │

│ spam ┆ 1   ┆ 3   │

└──────┴─────┴─────┘ shape: (3, 3)

┌──────┬─────┬─────┐

│ bar  ┆ len ┆ foo │

│ ---  ┆ --- ┆ --- │

│ str  ┆ u32 ┆ i64 │

╞══════╪═════╪═════╡

│ null ┆ 1   ┆ 1   │

│ ham  ┆ 1   ┆ 2   │

│ spam ┆ 1   ┆ 3   │

└──────┴─────┴─────┘)

API gaps: no collect_all, so the plans run one after another rather than together on the Polars thread pool; no pl.len(), so a column’s count stands in.