On this page:
2.1 Operators and pipelines
Expr-ptr?
col
lit
dtype-spec?
all
exclude
multi-column-expr?
expr->string
meta-output-name
meta-root-names
meta-eq?
>
<
>=
<=
=
!=
and
or
not
xor
+
-
*
/
filter
sort
sort-by
select
with-columns
cast
vstack
head
tail
slice
drop
join
read-csv
scan-csv
read-parquet
scan-parquet
read-ndjson
write-csv
write-parquet
write-ndjson
lazy
collect
group-by
agg
grouped?
over
rank
gather
sum
mean
min
max
count
n-unique
median
std
var
alias
first
last
pow
round
sign
is-between
is-in
dt-year
dt-month
dt-day
dt-hour
dt-minute
dt-second
str-extract
str->date
str->datetime
2.1.1 Shadowed bindings
2.2 Series
series?
series
series->string
polars-null
polars-null?
dtype
len
null-count
series-name
rename
rename!
clone
series-clone
2.2.1 Converting to Racket values
series->list
series->vector
series->f64vector
in-series
2.2.2 Categorical, Enum and Decimal
define-enum
2.2.3 dtype promotion
2.2.4 Low-level Series API
Series-ptr?
series-new-i8
series-new-i16
series-new-i32
series-new-i64
series-new-u8
series-new-u16
series-new-u32
series-new-u64
series-new-f32
series-new-f64
series-new-bool
series-new-str
series-sum-i32
series-min-i32
series-max-i32
series-mean-i32
series-sum-f64
series-min-f64
series-max-f64
series-mean-f64
series-cast
series-sort
2.3 Data  Frames
dataframe?
dataframe
shape
shape/  values
height
width
column-names
column-name
ref
describe
2.3.1 Converting to Racket values
dataframe->columns
dataframe->hash
in-dataframe-columns
in-dataframe-rows
dataframe->rows
dataframe->f64vector
2.3.2 Low-level Data  Frame API
dataframe-new
dataframe-shape
dataframe-height
dataframe-width
dataframe-column
dataframe-column-name
dataframe-column-names
dataframe-select
display-dataframe
Data  Frame-ptr?
dataframe-vstack
dataframe-sort
2.3.3 Reading & writing
dataframe-write-csv
dataframe-write-parquet
dataframe-read-parquet
dataframe-write-json-lines
dataframe-read-json-lines
dataframe-read-csv
lazyframe-scan-csv
2.4 Lazy frames
Lazy  Frame-ptr?
lazyframe?
dataframe-lazy
lazyframe-collect
lazyframe-select
lazyframe-with-columns
lazyframe-filter
lazyframe-group-by-agg
lazyframe-sort
lazyframe-join
2.5 Low-level expression API
expr-alias
expr-col
expr-all
expr-exclude
expr-dtype-col
expr-meta-output-name
expr-meta-root-names
expr-meta-eq?
expr-over
expr-add
expr-sub
expr-mul
expr-div
expr-mod
expr-gt
expr-lt
expr-ge
expr-le
expr-eq
expr-ne
expr-and
expr-or
expr-xor
expr-not
expr-sum
expr-mean
expr-min
expr-max
expr-median
expr-count
expr-n-unique
expr-first
expr-last
expr-std
expr-var
expr-sort
expr-sort-by
2.5.1 Eager expression contexts
dataframe-select-exprs
dataframe-with-columns
dataframe-filter-expr
dataframe-group-by-agg
2.6 Generic interfaces
gen:  has-ref
has-ref?
gen:  sized
sized?
gen:  has-shape
has-shape?
gen:  has-dtype
has-dtype?
gen:  has-null-count
has-null-count?
9.3

2 Reference🔗ℹ

polars is written to read as ordinary Racket. The operators you already know — +, *, >, and, filter, sort, first — are overloaded to work on series and expressions, dispatching at runtime on what they are given and falling back to their racket/base behaviour on plain values. Operators and pipelines documents that surface and comes first because it is the one to learn.

Underneath sits a monomorphic, dtype-suffixed layer (series-new-i32, series-sum-f64, expr-add and friends). It is documented here for completeness — see Low-level expression API and the low-level subsections — but the generic spelling is the preferred one throughout.

2.1 Operators and pipelines🔗ℹ

This is the surface to write polars in. An expression describes a column computation without running it, and the operators below are the ordinary Racket ones — +, *, >, and, filter, sort — overloaded to build expressions when an operand is one, while behaving exactly as racket/base does on plain numbers and lists. Nothing about a pipeline needs a Polars-specific spelling.

Every operation takes the frame — or the expression — as its first argument, so it chains with thread-first ~> (re-provided from threading, so (require polars) is enough):

(~> df
    (filter (> (col "value") 15))
    (group-by "group")
    (agg (alias (sum (col "value")) "sum_value")))

procedure

(Expr-ptr? v) → boolean?

  v : any/c
Returns #t if v is an expression.

procedure

(col spec) → Expr-ptr?

  spec : (or/c string? regexp? dtype-spec?)

procedure

(lit v) → Expr-ptr?

  v : (or/c boolean? exact-integer? real? string? symbol?)

procedure

(dtype-spec? v) → boolean?

  v : any/c
The leaves every other operation builds on. col refers to one column or to several at once, by the shape of spec:

  • A string is the column of that name (pl.col("name")). A string of the form ^...$ is a Polars regex, in the syntax of Rust’s regex crate, as in Python.

  • A dtype is every column of that dtype (pl.col(pl.Float64)). dtype-spec? is any spelling series’ #:dtype accepts, so 'float64 and 'f64 alike. A bare 'datetime means microseconds, so match a column series built from gregor datetimes with '(datetime milliseconds) or with its dtype. 'categorical is every categorical column, and an '(enum ....) dtype the columns of exactly that Enum (Categorical, Enum and Decimal).

  • A regexp is every column whose name matches (pl.col("^sepal_.*$")), keeping the regexp’s Racket meaning: (col rx) selects exactly the names (regexp-match? rx name) accepts, in #rx and #px syntax alike. There are two exceptions. A \p{...} property class follows each side’s own version of the Unicode tables. And Racket’s own matcher misjudges some classes containing characters above U+00FF; there the selection follows the class as written. The crate has no lookaround, backreferences, atomic groups or conditionals; a regexp using them is rejected at collect. To select by such a regexp, match column-names in Racket and select the names, as in the last example below.

A multi-column col expands inside any expression to one output per matched column, in the frame’s column order, each keeping the matched column’s name; a frame with no match yields no columns.

lit lifts a Racket scalar to a literal expression: booleans, exact integers (32-bit when they fit, 64-bit otherwise), other reals (as 'float64) and strings. A symbol is the string of its name, so a categorical column compares with the symbols it reads back as. Every operator below lifts a non-expression operand with lit automatically, so it is rarely needed explicitly.

Examples:
> (define people
    (dataframe (list (series '(1 2 3) #:name "id" #:dtype 'i32)
                     (series '(57.9 72.5 53.6) #:name "weight")
                     (series '(1.56 1.77 1.65) #:name "height"))))
> (select people (* (col 'float64) 1.1))

shape: (3, 2)

┌────────┬────────┐

│ weight ┆ height │

│ ---    ┆ ---    │

│ f64    ┆ f64    │

╞════════╪════════╡

│ 63.69  ┆ 1.716  │

│ 79.75  ┆ 1.947  │

│ 58.96  ┆ 1.815  │

└────────┴────────┘

> (list (dtype-spec? 'f64) (dtype-spec? 'float))

'(#t #f)

> (select people (col "^.*ght$"))

shape: (3, 2)

┌────────┬────────┐

│ weight ┆ height │

│ ---    ┆ ---    │

│ f64    ┆ f64    │

╞════════╪════════╡

│ 57.9   ┆ 1.56   │

│ 72.5   ┆ 1.77   │

│ 53.6   ┆ 1.65   │

└────────┴────────┘

> (select people (col #rx"^he"))

shape: (3, 1)

┌────────┐

│ height │

│ ---    │

│ f64    │

╞════════╡

│ 1.56   │

│ 1.77   │

│ 1.65   │

└────────┘

> (select people (col #px"^\\w+t$"))

shape: (3, 2)

┌────────┬────────┐

│ weight ┆ height │

│ ---    ┆ ---    │

│ f64    ┆ f64    │

╞════════╪════════╡

│ 57.9   ┆ 1.56   │

│ 72.5   ┆ 1.77   │

│ 53.6   ┆ 1.65   │

└────────┴────────┘

> (select people (~> (col "id") (* 10) (alias "id10"))
                 (alias (lit 0) "zero"))

shape: (3, 2)

┌──────┬──────┐

│ id10 ┆ zero │

│ ---  ┆ ---  │

│ i32  ┆ i32  │

╞══════╪══════╡

│ 10   ┆ 0    │

│ 20   ┆ 0    │

│ 30   ┆ 0    │

└──────┴──────┘

> (select people (col #px"^(?!id)"))

lazyframe-collect: failed to collect the query: invalid

regex in selector '^(?s).*(?:^(?!id)).*$'

Resolved plan until failure:

---> FAILED HERE RESOLVING 'select' <---

DF ["id", "weight", "height"]; PROJECT */3 COLUMNS: 'select'

> (select people (filter (lambda (name) (regexp-match? #px"^(?!id)" name))
                         (column-names people)))

shape: (3, 2)

┌────────┬────────┐

│ weight ┆ height │

│ ---    ┆ ---    │

│ f64    ┆ f64    │

╞════════╪════════╡

│ 57.9   ┆ 1.56   │

│ 72.5   ┆ 1.77   │

│ 53.6   ┆ 1.65   │

└────────┴────────┘

procedure

(all) → Expr-ptr?

procedure

(exclude e name ...+) → Expr-ptr?

  e : multi-column-expr?
  name : (or/c string? regexp?)

procedure

(multi-column-expr? v) → boolean?

  v : any/c
(all) is every column (pl.all()); inside agg it is every column that is not a group key. exclude removes columns from a multi-column expression — (all), or a dtype or regexp col, or any expression built over one — by name or by regexp (.exclude), each read as col reads it, so a name of the form ^...$ is a Polars regex. A name the frame does not have is ignored, and chained excludes accumulate. multi-column-expr? recognises the expressions exclude accepts: those that expand to one output per matched column. To drop columns from a frame eagerly, drop is the direct spelling.

Examples:
> (select people (all))

shape: (3, 3)

┌─────┬────────┬────────┐

│ id  ┆ weight ┆ height │

│ --- ┆ ---    ┆ ---    │

│ i32 ┆ f64    ┆ f64    │

╞═════╪════════╪════════╡

│ 1   ┆ 57.9   ┆ 1.56   │

│ 2   ┆ 72.5   ┆ 1.77   │

│ 3   ┆ 53.6   ┆ 1.65   │

└─────┴────────┴────────┘

> (select people (exclude (all) "id"))

shape: (3, 2)

┌────────┬────────┐

│ weight ┆ height │

│ ---    ┆ ---    │

│ f64    ┆ f64    │

╞════════╪════════╡

│ 57.9   ┆ 1.56   │

│ 72.5   ┆ 1.77   │

│ 53.6   ┆ 1.65   │

└────────┴────────┘

> (select people (exclude (all) #rx"^w" "id"))

shape: (3, 1)

┌────────┐

│ height │

│ ---    │

│ f64    │

╞════════╡

│ 1.56   │

│ 1.77   │

│ 1.65   │

└────────┘

> (select people (exclude (all) "^h.*$"))

shape: (3, 2)

┌─────┬────────┐

│ id  ┆ weight │

│ --- ┆ ---    │

│ i32 ┆ f64    │

╞═════╪════════╡

│ 1   ┆ 57.9   │

│ 2   ┆ 72.5   │

│ 3   ┆ 53.6   │

└─────┴────────┘

> (select people (~> (col 'float64) (exclude "height") (* 2)))

shape: (3, 1)

┌────────┐

│ weight │

│ ---    │

│ f64    │

╞════════╡

│ 115.8  │

│ 145.0  │

│ 107.2  │

└────────┘

> (select people (~> (all) (exclude "id") (exclude "weight")))

shape: (3, 1)

┌────────┐

│ height │

│ ---    │

│ f64    │

╞════════╡

│ 1.56   │

│ 1.77   │

│ 1.65   │

└────────┘

> (~> people
      (with-columns (~> (col "height") (> 1.6) (alias "tall")))
      (group-by "tall")
      (agg (~> (all) (exclude "id") mean))
      (sort "tall"))

shape: (2, 3)

┌───────┬────────┬────────┐

│ tall  ┆ weight ┆ height │

│ ---   ┆ ---    ┆ ---    │

│ bool  ┆ f64    ┆ f64    │

╞═══════╪════════╪════════╡

│ false ┆ 57.9   ┆ 1.56   │

│ true  ┆ 63.05  ┆ 1.71   │

└───────┴────────┴────────┘

> (multi-column-expr? (col "id"))

#f

> (~> (col 'float64) (* 2) multi-column-expr?)

#t

> (exclude (col "id") "weight")

exclude: contract violation

  expected: multi-column-expr?

  given: col("id")

  in: the 1st argument of

      (->

       multi-column-expr?

       (or/c string? regexp?)

       (or/c string? regexp?)

       ...

       Expr-ptr?)

  contract from:

      <pkgs>/polars/private/generic/selectors.rkt

  blaming: top-level

   (assuming the contract is correct)

  at: <pkgs>/polars/private/generic/selectors.rkt:8:11

procedure

(expr->string e) → string?

  e : Expr-ptr?
Renders e as its plan, in the notation Polars itself uses: col("v") for a column, [(a) + (b)] for a binary operation, .alias("n") and .sum() as method suffixes. This is also what an expression prints as at the REPL and throughout this manual, so an expression is a value you can read, not an opaque pointer. A regexp col prints as the Polars pattern it is translated to.

Examples:
> (col "weight")

col("weight")

> (alias (* (col "v") 10) "v10")

[(col("v")) * (dyn int: 10)].alias("v10")

> (expr->string (> (col "v") 2))

"[(col(\"v\")) > (dyn int: 2)]"

> (~> (col "v") sum (over "k"))

col("v").sum().over([col("k")])

> (exclude (all) "id")

[cs.all() - cs.by_name('id', require_all=false)]

> (col 'float64)

cs.by_dtype([Float64])

> (col #rx"^he")

cs.matches("^(?s).*(?:^he).*$")

procedure

(meta-output-name e) → string?

  e : (or/c Expr-ptr? string?)

procedure

(meta-root-names e) → (listof string?)

  e : (or/c Expr-ptr? string?)

procedure

(meta-eq? a b) → boolean?

  a : (or/c Expr-ptr? string?)
  b : (or/c Expr-ptr? string?)
Polars’ .meta namespace: what an expression will do, read off the plan without running it. meta-output-name is the column the expression produces — the alias if it has one, else its first column, else "literal"; it raises when that cannot be known without a frame. meta-root-names lists the columns the expression reads, in tree order, duplicates included. meta-eq? is structural equality of two plans; equal? on expressions is identity. A column name is lifted with col wherever an expression is expected.

A multi-column expression is read off the plan too, before any frame says which columns it will match, so a regexp or dtype col or (all) has no root names and no output name, as in Python.

Examples:
> (define total (alias (sum (+ (col "a") (col "b"))) "total"))
> total

[(col("a")) + (col("b"))].sum().alias("total")

> (meta-output-name total)

"total"

> (meta-root-names total)

'("a" "b")

> (meta-eq? total (alias (sum (+ (col "a") (col "b"))) "total"))

#t

> (meta-eq? total (col "a"))

#f

> (equal? total (~> (+ (col "a") (col "b")) sum (alias "total")))

#f

> (meta-output-name (+ (col "a") (col "b")))

"a"

> (meta-output-name (lit 25))

"literal"

> (meta-root-names "a")

'("a")

> (meta-root-names (col #rx"^he"))

'()

> (meta-output-name (col #rx"^he"))

expr-meta-output-name: cannot determine the output name of

cs.matches("^(?s).*(?:^he).*$")

> (~> (col 'float64) (* 2) meta-root-names)

'()

> (meta-eq? (col 'float64) (col 'f64))

#t

> (meta-output-name (col 'float64))

expr-meta-output-name: cannot determine the output name of

cs.by_dtype([Float64])

> (meta-output-name "*")

expr-meta-output-name: cannot determine the output name of

cs.all()

> (meta-output-name 5)

meta-output-name: contract violation

  expected: col-expr/c

  given: 5

  in: the 1st argument of

      (-> col-expr/c string?)

  contract from:

      <pkgs>/polars/private/generic/meta.rkt

  blaming: top-level

   (assuming the contract is correct)

  at: <pkgs>/polars/private/generic/meta.rkt:9:11

procedure

(> a b ...) → any/c

  a : any/c
  b : any/c

procedure

(< a b ...) → any/c

  a : any/c
  b : any/c

procedure

(>= a b ...) → any/c

  a : any/c
  b : any/c

procedure

(<= a b ...) → any/c

  a : any/c
  b : any/c

procedure

(= a b ...) → any/c

  a : any/c
  b : any/c

procedure

(!= a b ...) → any/c

  a : any/c
  b : any/c
Overloaded comparison operators. If an operand is an expression, they build a comparison expression (scalars are lifted automatically), so (> (col "value") 15) reads like col("value") > 15. If an operand is a series, they build an eager boolean-mask series — element-wise over every numeric dtype, including 'int64 — so (> (ref df #:columns "value") 15) is a mask. Otherwise they fall back to the numeric racket/base operator and stay variadic, so (> 3 2) and (< 1 2 3) still work. != has no racket/base spelling; on numbers it is (not (= a b)). These shadow the racket/base comparisons; see Shadowed bindings.

syntax

(and expr ...)

syntax

(or expr ...)

procedure

(not x) → any/c

  x : any/c

procedure

(xor a b) → any/c

  a : any/c
  b : any/c
Overloaded boolean connectives. When an operand is an expression they build the element-wise expression (expr-and, expr-or, expr-not, expr-xor), so (and (> (col "value") 15) (< (col "cost") 3.0)) is a predicate for filter. When an operand is a series they compute an eager boolean mask. Otherwise they behave as the racket/base forms: and and or short-circuit and return the deciding value, and not negates. Note that once and or or meets an expression or series operand it evaluates its remaining operands eagerly to combine them. These shadow the racket/base bindings; see Shadowed bindings.

procedure

(+ v ...) → any/c

  v : any/c

procedure

(- v ...) → any/c

  v : any/c

procedure

(* v ...) → any/c

  v : any/c

procedure

(/ v ...) → any/c

  v : any/c
Overloaded arithmetic. When every argument is a number they are exactly the racket/base operators. Otherwise they fold left over the arguments: an expression operand yields an expression (via expr-add, expr-sub, expr-mul, expr-div), so (* (col "value") 2) reads like col("value") * 2, and a series operand yields an eagerly computed series. A single non-numeric argument is returned unchanged. These shadow the racket/base bindings; see Shadowed bindings.

procedure

(filter d predicate) → dataframe?

  d : dataframe?
  predicate : any/c
Keeps the rows of d matching predicate, which may be a boolean expression — (filter df (> (col "value") 15)) — or a precomputed boolean-mask series. Returns a new dataframe. Applied to a non-dataframe it falls back to racket/base’s filter, so (filter even? '(1 2 3 4)) is '(2 4).

procedure

(sort d    
  by    
  [#:descending descending    
  #:nulls-last nulls-last    
  #:maintain-order maintain-order]) → dataframe?
  d : dataframe?
  by : (or/c string? (non-empty-listof string?))
  descending : (or/c boolean? (listof boolean?)) = #f
  nulls-last : (or/c boolean? (listof boolean?)) = #f
  maintain-order : boolean? = #f
(sort lf    
  by    
  [#:descending descending    
  #:nulls-last nulls-last    
  #:maintain-order maintain-order]) → lazyframe?
  lf : lazyframe?
  by : (or/c string? (non-empty-listof string?))
  descending : (or/c boolean? (listof boolean?)) = #f
  nulls-last : (or/c boolean? (listof boolean?)) = #f
  maintain-order : boolean? = #f
(sort s    
  [#:descending descending    
  #:nulls-last nulls-last]) → series?
  s : series?
  descending : boolean? = #f
  nulls-last : boolean? = #f
(sort e    
  [#:descending descending    
  #:nulls-last nulls-last]) → Expr-ptr?
  e : (or/c Expr-ptr? string?)
  descending : boolean? = #f
  nulls-last : boolean? = #f
(sort lst    
  less-than?    
  [#:key extract-key    
  #:cache-keys? cache-keys?]) → list?
  lst : list?
  less-than? : (any/c any/c . -> . any/c)
  extract-key : (or/c #f (any/c . -> . any/c)) = #f
  cache-keys? : boolean? = #f
Sorts a dataframe or lazyframe by the column or columns by (df.sort), or a series by its values (Series.sort). Given an expression, or a column name lifted with col, it builds the expression that sorts that one column (Expr.sort).

Nulls come first, whatever the direction, unless nulls-last is true; NaN sorts above every other float. For a frame, descending and nulls-last are each one boolean for every key or a list of one boolean per key. Rows that tie on every key keep their input order when maintain-order is true, and are in no particular order otherwise.

A sorted expression reorders its own column only, so in with-columns it no longer lines up with the rest of its row. Use it in select or agg, and sort-by to reorder one column by others. A frame sorts by column names only: sorting by an expression (df.sort(pl.col("a") * -1)) has no spelling yet.

On a list, sort is racket/base’s, and passing it one of the Polars keywords is a contract violation.

Examples:
> (define flights
    (dataframe (list (series '("UA" "AA" "UA" "AA" "B6") #:name "carrier")
                     (series (list 12 polars-null 340 -3 polars-null) #:name "delay"))))
> (sort flights "delay" #:descending #t)

shape: (5, 2)

┌─────────┬───────┐

│ carrier ┆ delay │

│ ---     ┆ ---   │

│ str     ┆ i64   │

╞═════════╪═══════╡

│ AA      ┆ null  │

│ B6      ┆ null  │

│ UA      ┆ 340   │

│ UA      ┆ 12    │

│ AA      ┆ -3    │

└─────────┴───────┘

> (~> flights (sort "delay" #:descending #t #:nulls-last #t) (head 2))

shape: (2, 2)

┌─────────┬───────┐

│ carrier ┆ delay │

│ ---     ┆ ---   │

│ str     ┆ i64   │

╞═════════╪═══════╡

│ UA      ┆ 340   │

│ UA      ┆ 12    │

└─────────┴───────┘

> (sort flights '("carrier" "delay")
        #:descending '(#f #t) #:nulls-last '(#f #t) #:maintain-order #t)

shape: (5, 2)

┌─────────┬───────┐

│ carrier ┆ delay │

│ ---     ┆ ---   │

│ str     ┆ i64   │

╞═════════╪═══════╡

│ AA      ┆ -3    │

│ AA      ┆ null  │

│ B6      ┆ null  │

│ UA      ┆ 340   │

│ UA      ┆ 12    │

└─────────┴───────┘

> (~> flights lazy (sort "delay" #:nulls-last #t) collect)

shape: (5, 2)

┌─────────┬───────┐

│ carrier ┆ delay │

│ ---     ┆ ---   │

│ str     ┆ i64   │

╞═════════╪═══════╡

│ AA      ┆ -3    │

│ UA      ┆ 12    │

│ UA      ┆ 340   │

│ AA      ┆ null  │

│ B6      ┆ null  │

└─────────┴───────┘

> (sort (ref flights #:columns "delay") #:descending #t #:nulls-last #t)

shape: (5,)

Series: 'delay' [i64]

[

340

12

-3

null

null

]

> (select flights (sort "delay" #:nulls-last #t))

shape: (5, 1)

┌───────┐

│ delay │

│ ---   │

│ i64   │

╞═══════╡

│ -3    │

│ 12    │

│ 340   │

│ null  │

│ null  │

└───────┘

> (sort '(3 1 2) <)

'(1 2 3)

> (sort '(3 1 2) < #:descending #t)

sort: contract violation

  expected: polars-only/c

  given: #t

  in: the descending argument of

      sort/c

  contract from:

      <pkgs>/polars/private/generic/ordering.rkt

  blaming: top-level

   (assuming the contract is correct)

  at: <pkgs>/polars/private/generic/ordering.rkt:17:24

procedure

(sort-by x    
  #:by by    
  [#:descending descending    
  #:nulls-last nulls-last    
  #:maintain-order maintain-order]) → Expr-ptr?
  x : (or/c Expr-ptr? string?)
  by : (or/c Expr-ptr? string? (non-empty-listof (or/c Expr-ptr? string?)))
  descending : (or/c boolean? (listof boolean?)) = #f
  nulls-last : (or/c boolean? (listof boolean?)) = #f
  maintain-order : boolean? = #f
Builds the expression that reorders the column x by the key or keys by (Expr.sort_by); a column name is lifted with col in either place. The keywords mean what they mean for a frame sort: one boolean, or one per key. The result keeps x’s length and name. Inside agg or over it sorts each group, including when x is itself group-aware, such as a shift.

Examples:
> (define scores
    (dataframe (list (series '("ann" "bob" "cy" "dee") #:name "name")
                     (series (list 3 polars-null 1 3) #:name "score"))))
> (select scores (sort-by "name" #:by "score" #:descending #t #:nulls-last #t
                          #:maintain-order #t))

shape: (4, 1)

┌──────┐

│ name │

│ ---  │

│ str  │

╞══════╡

│ ann  │

│ dee  │

│ cy   │

│ bob  │

└──────┘

> (select scores (sort-by "name" #:by '("score" "name") #:descending '(#t #f)))

shape: (4, 1)

┌──────┐

│ name │

│ ---  │

│ str  │

╞══════╡

│ bob  │

│ ann  │

│ dee  │

│ cy   │

└──────┘

procedure

(select d spec ...) → (or/c dataframe? lazyframe?)

  d : (or/c dataframe? lazyframe?)
  spec : any/c

procedure

(with-columns d spec ...) → (or/c dataframe? lazyframe?)

  d : (or/c dataframe? lazyframe?)
  spec : any/c
The select and with_columns contexts. Each spec is a column name, a column index, an expression, or a list of those (which is spliced); names and indices are lifted with col. select returns a frame holding only the resulting columns; with-columns adds them to (or replaces them in) the existing columns. Given a dataframe they run eagerly and return a dataframe; given a lazyframe they extend the plan and return a lazyframe.

Examples:
> (~> (dataframe (list (series '(1 2 3) #:name "a")
                       (series '(4 5 6) #:name "b")))
      (select "a" (alias (+ (col "a") (col "b")) "sum")))

shape: (3, 2)

┌─────┬─────┐

│ a   ┆ sum │

│ --- ┆ --- │

│ i64 ┆ i64 │

╞═════╪═════╡

│ 1   ┆ 5   │

│ 2   ┆ 7   │

│ 3   ┆ 9   │

└─────┴─────┘

> (~> (dataframe (list (series '(1 2 3) #:name "a")))
      (with-columns (alias (* (col "a") 2) "double")))

shape: (3, 2)

┌─────┬────────┐

│ a   ┆ double │

│ --- ┆ ---    │

│ i64 ┆ i64    │

╞═════╪════════╡

│ 1   ┆ 2      │

│ 2   ┆ 4      │

│ 3   ┆ 6      │

└─────┴────────┘

procedure

(cast x dtype) → (or/c Expr-ptr? series?)

  x : (or/c Expr-ptr? series? string?)
  dtype : (or/c symbol? pair?)
Changes dtype. On an expression (or a column name, lifted with col) it builds a cast expression, matching .cast; on a series it converts eagerly and returns a series. dtype takes the same spellings as series-cast: unlike series’ #:dtype, only the canonical names ('float64, 'int32, 'string, ...) are accepted, not the short ones ('f64, 'i32, 'str); given a short name the exn:fail exception is raised. A value that does not convert becomes null, except in a cast to an Enum, which raises naming the values outside its categories, as Python’s default strict=True does (on an expression, when the plan runs).

Examples:
> (~> (dataframe (list (series '(1 2 3) #:name "v")))
      (with-columns (cast "v" 'float64)))

shape: (3, 1)

┌─────┐

│ v   │

│ --- │

│ f64 │

╞═════╡

│ 1.0 │

│ 2.0 │

│ 3.0 │

└─────┘

> (cast (series '("UA" "AA" "UA")) 'categorical)

shape: (3,)

Series: '' [cat]

[

"UA"

"AA"

"UA"

]

> (cast (series '(1 2 3)) 'f64)

->compat-dtype: unsupported cast target 'f64

> (define-enum carriers UA AA)
> (cast (series '("UA" "B6")) carriers)

series-cast: cannot convert to '(enum UA AA): conversion

from `str` to `enum` failed in column '' for 1 out of 2

values: ["B6"]

Ensure that all values in the input column are present in

the categories of the enum datatype.

procedure

(vstack top bottom) → dataframe?

  top : dataframe?
  bottom : dataframe?
Stacks the rows of bottom beneath those of top, which must have the same columns in the same order (Polars’ vstack).

Examples:
> (define top (dataframe (list (series '(1 2) #:name "v"))))
> (vstack top (dataframe (list (series '(3) #:name "v"))))

shape: (3, 1)

┌─────┐

│ v   │

│ --- │

│ i64 │

╞═════╡

│ 1   │

│ 2   │

│ 3   │

└─────┘

procedure

(head x n) → any/c

  x : (or/c series? dataframe? lazyframe? Expr-ptr?)
  n : exact-nonnegative-integer?

procedure

(tail x n) → any/c

  x : (or/c series? dataframe? lazyframe? Expr-ptr?)
  n : exact-nonnegative-integer?

procedure

(slice x offset length) → any/c

  x : (or/c series? dataframe? lazyframe? Expr-ptr?)
  offset : exact-integer?
  length : exact-nonnegative-integer?
The first n rows, the last n rows, or length rows starting at offset, of the same kind as x (Polars’ head, tail, slice). On an expression they change the length and belong inside select.

Examples:
> (define nums (dataframe (list (series '(1 2 3 4 5) #:name "v"))))
> (head nums 2)

shape: (2, 1)

┌─────┐

│ v   │

│ --- │

│ i64 │

╞═════╡

│ 1   │

│ 2   │

└─────┘

> (tail nums 2)

shape: (2, 1)

┌─────┐

│ v   │

│ --- │

│ i64 │

╞═════╡

│ 4   │

│ 5   │

└─────┘

procedure

(drop d names) → dataframe?

  d : dataframe?
  names : (or/c string? (listof string?))
Removes the named column(s) (df.drop). Applied to a list it falls back to racket/list’s drop.

procedure

(join left    
  right    
  [#:on on    
  #:left-on left-on    
  #:right-on right-on    
  #:how how]) → (or/c dataframe? lazyframe?)
  left : (or/c dataframe? lazyframe?)
  right : (or/c dataframe? lazyframe?)
  on : (or/c (listof string?) #f) = #f
  left-on : (or/c (listof string?) #f) = #f
  right-on : (or/c (listof string?) #f) = #f
  how : (or/c 'inner 'left 'outer 'cross 'semi 'anti) = 'inner
Joins right onto left on the shared key columns #:on, or on #:left-on / #:right-on (left.join(right, ...)). Eager on a dataframe, deferred on a lazyframe. As in Python Polars, only a 'cross join promises a row order; sort the result when order matters.

Examples:
> (define left (dataframe (list (series '("a" "b") #:name "k")
                                 (series '(1 2) #:name "v"))))
> (define right (dataframe (list (series '("a" "c") #:name "k")
                                  (series '(10 30) #:name "w"))))
> (~> (join left right #:on '("k") #:how 'left) (sort "k"))

shape: (2, 3)

┌─────┬─────┬──────┐

│ k   ┆ v   ┆ w    │

│ --- ┆ --- ┆ ---  │

│ str ┆ i64 ┆ i64  │

╞═════╪═════╪══════╡

│ a   ┆ 1   ┆ 10   │

│ b   ┆ 2   ┆ null │

└─────┴─────┴──────┘

> (join left right #:on '("k") #:how 'inner)

shape: (1, 3)

┌─────┬─────┬─────┐

│ k   ┆ v   ┆ w   │

│ --- ┆ --- ┆ --- │

│ str ┆ i64 ┆ i64 │

╞═════╪═════╪═════╡

│ a   ┆ 1   ┆ 10  │

└─────┴─────┴─────┘

procedure

(read-csv path 
  [#:has-header has-header 
  #:separator separator 
  #:quote-char quote-char 
  #:comment-prefix comment-prefix 
  #:skip-rows skip-rows 
  #:n-rows n-rows 
  #:null-values null-values 
  #:infer-schema-length infer-schema-length 
  #:schema-overrides schema-overrides 
  #:ignore-errors ignore-errors 
  #:try-parse-dates try-parse-dates 
  #:encoding encoding 
  #:glob glob]) 
 → dataframe?
  path : path-string?
  has-header : boolean? = #t
  separator : (or/c csv-char/c #f) = #f
  quote-char : (or/c csv-char/c #f) = #\"
  comment-prefix : (or/c non-empty-string? #f) = #f
  skip-rows : exact-nonnegative-integer? = 0
  n-rows : (or/c exact-nonnegative-integer? #f) = #f
  null-values : (or/c string? (listof string?) #f) = #f
  infer-schema-length : (or/c exact-nonnegative-integer? #f)
   = 100
  schema-overrides : (and/c (listof (cons/c string? csv-dtype/c)) distinct-names?)
   = '()
  ignore-errors : boolean? = #f
  try-parse-dates : boolean? = #f
  encoding : (or/c 'utf8 'utf8-lossy) = 'utf8
  glob : boolean? = #t
Reads CSV into a dataframe (pl.read_csv). path may be a glob pattern (*, ?, [...]): every matching file is read, in sorted filename order, and the files must share a header; #:glob #f takes the path literally, and [[] matches a literal [ in a pattern. A relative path is resolved against current-directory, whose own name is never read as a pattern. A directory is an error: name its files with a pattern.

A csv-char/c is one ASCII character other than newline or return. separator defaults to #\,; quote-char must differ from it, and #:quote-char #f turns quoting off. Lines that start with comment-prefix are skipped, as are the first skip-rows lines of each file; n-rows caps the rows read. A field equal to one of null-values reads as null. Column types are inferred from the first infer-schema-length rows: #f reads every row, and 0 makes every column a string. schema-overrides fixes the named columns’ types. A csv-dtype/c is any spelling series’ #:dtype accepts except a duration, which Polars cannot parse from CSV, and an Enum; API gap: read the column as 'categorical or 'string and cast it. Each column appears at most once (distinct-names?), and naming a column the file lacks is an error, where Python ignores the override. With #:ignore-errors #t a field that does not parse reads as null. #:try-parse-dates #t reads ISO dates, times of day and datetimes as 'date, 'time and 'datetime columns. 'utf8-lossy replaces invalid UTF-8 with U+FFFD.

The result is (collect (scan-csv path ....)) with the same keywords; as in Python, a single file is read eagerly rather than through a plan, which is faster. There is one extra check, stricter than Python’s read_csv, which returns the one column. When separator is #f and the file reads as one column whose header splits on a tab, ; or |, read-csv raises an error naming the separator to pass — unless the first row is not a string, or splits into a different number of fields. Passing a separator, even #\,, turns the check off.

A failure names the operation, the path and the cause: the operating system’s for a file that cannot be opened, Polars’ own for input it cannot parse.

The examples read "flights.tsv", 102 rows of the nycflights13 data with NA for a missing value. Without #:separator the separator check stops the read; with it, the first NA fails to parse, and #:null-values fixes that:

Examples:
> (read-csv "flights.tsv")

dataframe-read-csv: flights.tsv reads as one column whose

header splits on #\tab into 19 fields; pass #:separator

#\tab, or #:separator #\, to keep one column

> (read-csv "flights.tsv" #:separator #\tab)

dataframe-read-csv: failed to read csv from flights.tsv:

could not parse `NA` as dtype `i64` at column 'arr_delay'

(column number 9)

The current offset in the file is 648 bytes.

You might want to try:

- increasing #:infer-schema-length (e.g.

#:infer-schema-length 10000, or #f for every row),

- specifying correct dtype with #:schema-overrides

- setting #:ignore-errors to #t,

- adding `NA` to #:null-values.

Original error: ```invalid primitive value found during CSV

parsing```

> (define flights (read-csv "flights.tsv" #:separator #\tab #:null-values "NA"))
> (shape flights)

'(102 19)

> (null-count (ref flights #:columns "dep_delay"))

1

> (~> (read-csv "flights.tsv" #:separator #\tab #:null-values '("NA" ""))
      (ref #:columns "arr_delay")
      null-count)

2

Failures. A missing file or an unwritable path reports the operating system’s reason; a file in the wrong format reports Polars’:

> (read-csv "/no/such/file.csv")

dataframe-read-csv: failed to read csv from

/no/such/file.csv: cannot open file: No such file or

directory (os error 2)

> (define small (dataframe (list (series '(1 2) #:name "v"))))
> (write-parquet small "/no/such/dir/out.parquet")

dataframe-write-parquet: failed to write parquet to

/no/such/dir/out.parquet: cannot create file: No such file

or directory (os error 2)

Here not-parquet is a path in the temporary directory:

> (write-csv small not-parquet)
> (read-parquet not-parquet)

dataframe-read-parquet: failed to read parquet from

/var/tmp/polars-not-parquet.csv: parquet: File out of

specification: A Parquet file must contain a header and

footer with at least 12 bytes

> (read-ndjson not-parquet)

dataframe-read-json-lines: failed to read json lines from

/var/tmp/polars-not-parquet.csv: InternalError(TapeError) at

character 0 ('v')

Types. #:ignore-errors turns what does not parse into nulls; #:infer-schema-length widens or narrows the rows types are inferred from; #:schema-overrides and #:try-parse-dates set them outright. An override for a column the file lacks is an error:

> (~> (read-csv "flights.tsv" #:separator #\tab #:ignore-errors #t)
      (ref #:columns "dep_delay")
      null-count)

1

> (~> (read-csv "flights.tsv" #:separator #\tab #:infer-schema-length #f)
      (ref #:columns "dep_delay")
      dtype)

'string

> (~> (read-csv "flights.tsv" #:separator #\tab #:infer-schema-length 0)
      (ref #:columns "year")
      dtype)

'string

> (~> (read-csv "flights.tsv" #:separator #\tab #:null-values "NA"
                #:try-parse-dates #t
                #:schema-overrides '(("dep_delay" . f64) ("flight" . int32)))
      (select "dep_delay" "flight" "time_hour")
      (tail 3))

shape: (3, 3)

┌───────────┬────────┬─────────────────────┐

│ dep_delay ┆ flight ┆ time_hour           │

│ ---       ┆ ---    ┆ ---                 │

│ f64       ┆ i32    ┆ datetime[μs]        │

╞═══════════╪════════╪═════════════════════╡

│ -7.0      ┆ 1733   ┆ 2013-01-01 07:00:00 │

│ -5.0      ┆ 4525   ┆ 2013-01-01 15:00:00 │

│ null      ┆ 4308   ┆ 2013-01-01 16:00:00 │

└───────────┴────────┴─────────────────────┘

> (read-csv "flights.tsv" #:separator #\tab
            #:schema-overrides '(("dep_dealy" . f64)))

dataframe-read-csv: failed to read csv from flights.tsv:

schema overrides name columns not in the file: "dep_dealy"

Layout. "notes.csv" has a comment line, ; between fields and ' around a field that holds one; "latin1.csv" is not UTF-8:

> (read-csv "notes.csv" #:separator #\; #:comment-prefix "#" #:quote-char #\')

shape: (2, 2)

┌─────────┬───────────────┐

│ carrier ┆ note          │

│ ---     ┆ ---           │

│ str     ┆ str           │

╞═════════╪═══════════════╡

│ UA      ┆ late; weather │

│ AA      ┆ on time       │

└─────────┴───────────────┘

> (read-csv "notes.csv" #:separator #\; #:comment-prefix "#")

dataframe-read-csv: failed to read csv from notes.csv: found

more fields than defined in 'Schema'

Consider setting 'truncate_ragged_lines=true'.

> (read-csv "parts/part-1.csv" #:has-header #f #:skip-rows 1)

shape: (2, 3)

┌──────────┬──────────┬──────────┐

│ column_1 ┆ column_2 ┆ column_3 │

│ ---      ┆ ---      ┆ ---      │

│ str      ┆ str      ┆ i64      │

╞══════════╪══════════╪══════════╡

│ EWR      ┆ IAH      ┆ 2        │

│ LGA      ┆ IAH      ┆ 4        │

└──────────┴──────────┴──────────┘

> (~> (read-csv "flights.tsv" #:separator #\tab #:null-values "NA" #:n-rows 2)
      (select "carrier" "flight" "dep_delay"))

shape: (2, 3)

┌─────────┬────────┬───────────┐

│ carrier ┆ flight ┆ dep_delay │

│ ---     ┆ ---    ┆ ---       │

│ str     ┆ i64    ┆ i64       │

╞═════════╪════════╪═══════════╡

│ UA      ┆ 1545   ┆ 2         │

│ UA      ┆ 1714   ┆ 4         │

└─────────┴────────┴───────────┘

> (read-csv "latin1.csv")

dataframe-read-csv: failed to read csv from latin1.csv:

invalid utf-8 sequence

> (read-csv "latin1.csv" #:encoding 'utf8-lossy)

shape: (2, 2)

┌─────────┬──────────┐

│ carrier ┆ name     │

│ ---     ┆ ---      │

│ str     ┆ str      │

╞═════════╪══════════╡

│ B6      ┆ JetBlue  │

│ ZZ      ┆ Caf� Air │

└─────────┴──────────┘

Files. A pattern reads every match, in filename order; one that matches nothing, a literal path with #:glob #f, and a directory are errors:

> (read-csv "parts/*.csv")

shape: (6, 3)

┌────────┬──────┬───────────┐

│ origin ┆ dest ┆ dep_delay │

│ ---    ┆ ---  ┆ ---       │

│ str    ┆ str  ┆ i64       │

╞════════╪══════╪═══════════╡

│ EWR    ┆ IAH  ┆ 2         │

│ LGA    ┆ IAH  ┆ 4         │

│ JFK    ┆ MIA  ┆ 2         │

│ JFK    ┆ BQN  ┆ -1        │

│ LGA    ┆ ATL  ┆ -6        │

│ EWR    ┆ ORD  ┆ -4        │

└────────┴──────┴───────────┘

> (read-csv "parts/*.tsv")

dataframe-read-csv: failed to read csv from parts/*.tsv: no

files match the pattern

> (read-csv "parts/part-?.csv" #:glob #f)

dataframe-read-csv: failed to read csv from

parts/part-?.csv: cannot open file: No such file or

directory (os error 2)

> (read-csv "parts")

dataframe-read-csv: failed to read csv from parts: cannot

open file: it is a directory; pass a glob pattern such as

dir/*.csv

procedure

(scan-csv path 
  [#:has-header has-header 
  #:separator separator 
  #:quote-char quote-char 
  #:comment-prefix comment-prefix 
  #:skip-rows skip-rows 
  #:n-rows n-rows 
  #:null-values null-values 
  #:infer-schema-length infer-schema-length 
  #:schema-overrides schema-overrides 
  #:ignore-errors ignore-errors 
  #:try-parse-dates try-parse-dates 
  #:encoding encoding 
  #:glob glob]) 
 → lazyframe?
  path : path-string?
  has-header : boolean? = #t
  separator : (or/c csv-char/c #f) = #f
  quote-char : (or/c csv-char/c #f) = #\"
  comment-prefix : (or/c non-empty-string? #f) = #f
  skip-rows : exact-nonnegative-integer? = 0
  n-rows : (or/c exact-nonnegative-integer? #f) = #f
  null-values : (or/c string? (listof string?) #f) = #f
  infer-schema-length : (or/c exact-nonnegative-integer? #f)
   = 100
  schema-overrides : (and/c (listof (cons/c string? csv-dtype/c)) distinct-names?)
   = '()
  ignore-errors : boolean? = #f
  try-parse-dates : boolean? = #f
  encoding : (or/c 'utf8 'utf8-lossy) = 'utf8
  glob : boolean? = #t
Starts a lazyframe plan from CSV without reading it (pl.scan_csv); the keywords are read-csv’s. collect runs the plan, and that is where a file that cannot be read, or a glob pattern that matches no file, is reported. One thing is reported here instead: when schema-overrides is given, a column it names that the header lacks, which reads the header. The separator check does not apply, and a directory reads every file in it.

Examples:
> (~> (scan-csv "flights.tsv" #:separator #\tab #:null-values "NA")
      (filter (> (col "dep_delay") 30))
      (select "carrier" "dep_delay")
      collect)

shape: (2, 2)

┌─────────┬───────────┐

│ carrier ┆ dep_delay │

│ ---     ┆ ---       │

│ str     ┆ i64       │

╞═════════╪═══════════╡

│ UA      ┆ 47        │

│ MQ      ┆ 39        │

└─────────┴───────────┘

> (~> (scan-csv "parts/*.csv") (group-by "origin") (agg (sum "dep_delay")) collect)

shape: (3, 2)

┌────────┬───────────┐

│ origin ┆ dep_delay │

│ ---    ┆ ---       │

│ str    ┆ i64       │

╞════════╪═══════════╡

│ JFK    ┆ 1         │

│ LGA    ┆ -2        │

│ EWR    ┆ -2        │

└────────┴───────────┘

> (shape (collect (scan-csv "flights.tsv")))

'(102 1)

> (define plan (scan-csv "/no/such/file.csv"))
> (collect plan)

lazyframe-collect: failed to collect the query: No such file

or directory (os error 2): /no/such/file.csv

> (define no-match (scan-csv "parts/*.tsv"))
> (collect no-match)

lazyframe-collect: failed to collect the query: no files

match the pattern

> (scan-csv "flights.tsv" #:separator #\tab
            #:schema-overrides '(("dep_dealy" . f64)))

lazyframe-scan-csv: failed to scan flights.tsv: schema

overrides name columns not in the file: "dep_dealy"

procedure

(read-parquet path) → dataframe?

  path : path-string?

procedure

(scan-parquet path [#:n-rows n-rows]) → lazyframe?

  path : path-string?
  n-rows : (or/c exact-nonnegative-integer? #f) = #f
Read Parquet eagerly (pl.read_parquet), or start a plan from it (pl.scan_parquet). path is always a glob pattern: the matching files stack in sorted filename order, and one that matches nothing is an error, which scan-parquet leaves to collect. A directory reads every file in it and adds its key=value subdirectory names as columns; a single file or a pattern adds none, as in Python, and neither does a directory whose own path holds [, * or ?, which Polars reads as a pattern once it is escaped. read-parquet is (collect (scan-parquet path)). API gap: no #:glob, so a literal [, * or ? in a file name is spelled [[], [*] or [?] (#36).

Categorical, Enum and Decimal columns keep their dtypes (Categorical, Enum and Decimal). A file Polars cannot read raises exn:fail with Polars’ reason, even where Polars itself panics.

Examples:
> (read-parquet (build-path parquet-dir "*.parquet"))

shape: (6, 3)

┌────────┬──────┬───────────┐

│ origin ┆ dest ┆ dep_delay │

│ ---    ┆ ---  ┆ ---       │

│ str    ┆ str  ┆ i64       │

╞════════╪══════╪═══════════╡

│ EWR    ┆ IAH  ┆ 2         │

│ LGA    ┆ IAH  ┆ 4         │

│ JFK    ┆ MIA  ┆ 2         │

│ JFK    ┆ BQN  ┆ -1        │

│ LGA    ┆ ATL  ┆ -6        │

│ EWR    ┆ ORD  ┆ -4        │

└────────┴──────┴───────────┘

> (collect (scan-parquet (build-path parquet-dir "part-*.parquet") #:n-rows 3))

shape: (3, 3)

┌────────┬──────┬───────────┐

│ origin ┆ dest ┆ dep_delay │

│ ---    ┆ ---  ┆ ---       │

│ str    ┆ str  ┆ i64       │

╞════════╪══════╪═══════════╡

│ EWR    ┆ IAH  ┆ 2         │

│ LGA    ┆ IAH  ┆ 4         │

│ JFK    ┆ MIA  ┆ 2         │

└────────┴──────┴───────────┘

> (read-parquet "produce.parquet")

shape: (4, 3)

┌───────┬───────┬───────────────┐

│ item  ┆ grade ┆ price         │

│ ---   ┆ ---   ┆ ---           │

│ cat   ┆ enum  ┆ decimal[10,2] │

╞═══════╪═══════╪═══════════════╡

│ apple ┆ high  ┆ 1.25          │

│ pear  ┆ low   ┆ 0.80          │

│ apple ┆ null  ┆ null          │

│ fig   ┆ mid   ┆ 12.00         │

└───────┴───────┴───────────────┘

procedure

(read-ndjson path) → dataframe?

  path : path-string?

procedure

(write-csv d path) → void?

  d : dataframe?
  path : path-string?

procedure

(write-parquet d path) → void?

  d : dataframe?
  path : path-string?

procedure

(write-ndjson d path) → void?

  d : dataframe?
  path : path-string?
Newline-delimited JSON in (pl.read_ndjson), and a dataframe out to CSV, Parquet or newline-delimited JSON (df.write_csv and friends). API gap: there is no scan_ndjson, so read-ndjson reads one file and takes no glob pattern (#44).

procedure

(lazy d) → lazyframe?

  d : dataframe?

procedure

(collect lf) → dataframe?

  lf : lazyframe?
lazy turns a dataframe into a lazyframe — a plan that select, with-columns, filter and the other fluent operations extend without running anything — and collect executes the plan and returns the resulting dataframe. A plan that cannot run, such as one reading a column the frame lacks, is reported by collect with Polars’ reason.

Examples:
> (define four (dataframe (list (series '(1 2 3 4) #:name "v"))))
> (~> four lazy (filter (> (col "v") 2)) collect)

shape: (2, 1)

┌─────┐

│ v   │

│ --- │

│ i64 │

╞═════╡

│ 3   │

│ 4   │

└─────┘

> (~> four lazy (filter (> (col "nope") 2)) collect)

lazyframe-collect: failed to collect the query: not found:

unable to find column "nope"; valid columns: ["v"]

Resolved plan until failure:

---> FAILED HERE RESOLVING THIS_NODE <---

DF ["v"]; PROJECT */1 COLUMNS

procedure

(group-by d key ...) → grouped?

  d : dataframe?
  key : (or/c string? any/c)

procedure

(agg g agg-expr ...) → dataframe?

  g : grouped?
  agg-expr : any/c

procedure

(grouped? v) → boolean?

  v : any/c
group-by captures d and one or more group keys in a deferred grouped handle — no work happens yet — so it threads cleanly. agg consumes the handle, computing the aggregation expressions per group in a single pass, and returns a dataframe with one row per group. The split mirrors df.group_by("g").agg(...); the row order of the result is not guaranteed.

Example:
> (~> (dataframe (list (series '("x" "y" "x") #:name "k")
                       (series '(1 2 3) #:name "v")))
      (group-by "k")
      (agg (alias (sum "v") "total")))

shape: (2, 2)

┌─────┬───────┐

│ k   ┆ total │

│ --- ┆ ---   │

│ str ┆ i64   │

╞═════╪═══════╡

│ y   ┆ 2     │

│ x   ┆ 4     │

└─────┴───────┘

procedure

(over e key ...+) → Expr-ptr?

  e : (or/c Expr-ptr? string?)
  key : (or/c Expr-ptr? string?)
A window function (Expr.over): e is computed within each group of the keys — column names or expressions, as group-by takes them — and broadcast back onto the group’s rows rather than reduced to one row per group, so it belongs in with-columns where agg would collapse it. An aggregate is repeated on every row of its group; an expression that keeps one value per row, such as rank, is computed within the group and each value lands on the row it came from. Only Polars’ default group_to_rows mapping is exposed.

Examples:
> (define kv (dataframe (list (series '("x" "y" "x") #:name "k")
                              (series '(1 2 3) #:name "v"))))
> (~> kv (group-by "k") (agg (alias (sum "v") "total")))

shape: (2, 2)

┌─────┬───────┐

│ k   ┆ total │

│ --- ┆ ---   │

│ str ┆ i64   │

╞═════╪═══════╡

│ x   ┆ 4     │

│ y   ┆ 2     │

└─────┴───────┘

> (~> kv (with-columns (~> (col "v") sum (over "k") (alias "total"))))

shape: (3, 3)

┌─────┬─────┬───────┐

│ k   ┆ v   ┆ total │

│ --- ┆ --- ┆ ---   │

│ str ┆ i64 ┆ i64   │

╞═════╪═════╪═══════╡

│ x   ┆ 1   ┆ 4     │

│ y   ┆ 2   ┆ 2     │

│ x   ┆ 3   ┆ 4     │

└─────┴─────┴───────┘

> (define khv (dataframe (list (series '("a" "a" "a" "b") #:name "k")
                               (series '("x" "y" "x" "x") #:name "h")
                               (series '(1 2 3 4) #:name "v"))))
> (~> khv (with-columns (~> (col "v") sum (over "k" "h") (alias "total"))
                        (~> (col "v") mean (over (col "k")) (alias "mean"))
                        (~> (col "v") (rank #:descending #t) (over "k") (alias "rank"))
                        (~> (col "v") max (over (> (col "v") 1)) (alias "band_max"))))

shape: (4, 7)

┌─────┬─────┬─────┬───────┬──────┬──────┬──────────┐

│ k   ┆ h   ┆ v   ┆ total ┆ mean ┆ rank ┆ band_max │

│ --- ┆ --- ┆ --- ┆ ---   ┆ ---  ┆ ---  ┆ ---      │

│ str ┆ str ┆ i64 ┆ i64   ┆ f64  ┆ f64  ┆ i64      │

╞═════╪═════╪═════╪═══════╪══════╪══════╪══════════╡

│ a   ┆ x   ┆ 1   ┆ 4     ┆ 2.0  ┆ 3.0  ┆ 1        │

│ a   ┆ y   ┆ 2   ┆ 2     ┆ 2.0  ┆ 2.0  ┆ 4        │

│ a   ┆ x   ┆ 3   ┆ 4     ┆ 2.0  ┆ 1.0  ┆ 4        │

│ b   ┆ x   ┆ 4   ┆ 4     ┆ 4.0  ┆ 1.0  ┆ 4        │

└─────┴─────┴─────┴───────┴──────┴──────┴──────────┘

> (over (col "v"))

over: contract violation

  received: 1 argument

  expected: at least 2 non-keyword arguments

  in: (->*

       (col-expr/c col-expr/c)

       #:rest

       (listof col-expr/c)

       Expr-ptr?)

  contract from:

      <pkgs>/polars/private/generic/window.rkt

  blaming: top-level

   (assuming the contract is correct)

  at: <pkgs>/polars/private/generic/window.rkt:8:11

procedure

(rank x    
  [#:method method    
  #:descending descending    
  #:seed seed]) → Expr-ptr?
  x : (or/c Expr-ptr? string?)
  method : (or/c 'average 'min 'max 'dense 'ordinal) = 'average
  descending : boolean? = #f
  seed : (or/c exact-nonnegative-integer? #f) = #f

procedure

(gather x indices) → Expr-ptr?

  x : (or/c Expr-ptr? string?)
  indices : (or/c Expr-ptr? series? (listof exact-integer?))
Ordering within a column (.rank, .gather); sort-by is documented with the other sorts. rank numbers each value by its place in the sorted order, ties resolved by #:method; the result is 'uint32, or 'float64 for 'average. gather picks values by position. Each lifts a column name with col, and each is computed per group under over.

Examples:
> (define scores (dataframe (list (series '("a" "b" "c" "d") #:name "name")
                                  (series '(30 10 30 20) #:name "score"))))
> (~> scores
      (with-columns (~> (col "score") (rank #:method 'dense #:descending #t) (alias "dense"))
                    (~> (col "score") (rank #:method 'ordinal) (alias "ordinal"))))

shape: (4, 4)

┌──────┬───────┬───────┬─────────┐

│ name ┆ score ┆ dense ┆ ordinal │

│ ---  ┆ ---   ┆ ---   ┆ ---     │

│ str  ┆ i64   ┆ u32   ┆ u32     │

╞══════╪═══════╪═══════╪═════════╡

│ a    ┆ 30    ┆ 1     ┆ 3       │

│ b    ┆ 10    ┆ 3     ┆ 1       │

│ c    ┆ 30    ┆ 1     ┆ 4       │

│ d    ┆ 20    ┆ 2     ┆ 2       │

└──────┴───────┴───────┴─────────┘

> (~> scores (select (gather "name" '(3 0))))

shape: (2, 1)

┌──────┐

│ name │

│ ---  │

│ str  │

╞══════╡

│ d    │

│ a    │

└──────┘

> (rank "score" #:method 'first)

expr-rank: method must be one of 'average 'min 'max 'dense

'ordinal, got 'first

procedure

(sum v ...) → any/c

  v : any/c

procedure

(mean v ...) → any/c

  v : any/c

procedure

(min v ...) → any/c

  v : any/c

procedure

(max v ...) → any/c

  v : any/c
These dispatch on their argument. Applied to a single series, they reduce it (dispatching on its dtype): min, max and sum preserve the input dtype, while mean always returns a float64, so the mean of an integer series is a flonum (see dtype promotion). Applied to a single expression or a bare column-name string, they produce the corresponding aggregation expression — so (sum (col "value")) reads like Polars’ col("value").sum() and is used inside agg (see Operators and pipelines). Applied to anything else they fall back to the usual numeric behaviour, so (max 1 2 3) still works.

Examples:
> (define s (series '(1 2 3 4) #:name "v"))
> (list (sum s) (mean s) (min s) (max s))

'(10 2.5 1 4)

> (max 1 2 3)

3

procedure

(count x) → any/c

  x : any/c

procedure

(n-unique x) → any/c

  x : any/c

procedure

(median x) → any/c

  x : any/c

procedure

(std x [#:ddof ddof]) → any/c

  x : any/c
  ddof : exact-nonnegative-integer? = 1

procedure

(var x [#:ddof ddof]) → any/c

  x : any/c
  ddof : exact-nonnegative-integer? = 1

procedure

(alias e name) → any/c

  e : any/c
  name : string?
Aggregation-expression builders for use inside agg, alongside the expression arms of sum, mean, min and max. Each accepts an expression or a bare column-name string (lifted with col), so (count "value") and (count (col "value")) are equivalent. alias names a result, matching Polars’ .alias: (alias (sum (col "value")) "total"). std and var take a #:ddof degrees-of-freedom adjustment, defaulting to 1.

procedure

(first x) → any/c

  x : (or/c string? pair? any/c)

procedure

(last x) → any/c

  x : (or/c string? pair? any/c)
Dual-purpose. On an expression or column-name string they build the first/last-element aggregation (Polars’ .first() / .last()), for use inside agg. On a list they are the ordinary list accessors, so (first '(1 2 3)) is 1 and (last '(1 2 3)) is 3 — matching racket/list.

Name clash. racket/list also exports first and last (along with count and group-by, which polars exports too). Requiring both modules explicitly — (require racket/list polars) — is an error (identifier already required). A plain #lang racket/base program is unaffected, because racket/base does not export these names. See Shadowed bindings for how to take control.

procedure

(pow x exponent) → Expr-ptr?

  x : (or/c Expr-ptr? string?)
  exponent : (or/c Expr-ptr? real?)

procedure

(round x [#:decimals decimals]) → any/c

  x : (or/c Expr-ptr? string? number?)
  decimals : exact-nonnegative-integer? = 0
Element-wise power (**) and rounding to #:decimals places (.round). Each takes an expression or a bare column-name string (lifted with col); round on a plain number falls back to numeric rounding. Ties round to even in both cases, as in Python Polars and racket/base. See Shadowed bindings.

Examples:
> (~> (dataframe (list (series '(-2.5 -1.5 0.5 1.5 2.5) #:name "x")))
      (select (round "x")))

shape: (5, 1)

┌──────┐

│ x    │

│ ---  │

│ f64  │

╞══════╡

│ -2.0 │

│ -2.0 │

│ 0.0  │

│ 2.0  │

│ 2.0  │

└──────┘

> (round 2.5)

2.0

procedure

(sign x) → Expr-ptr?

  x : (or/c Expr-ptr? string?)
The sign of each element (.sign): -1, 0 or 1 in the column’s own dtype, so a float column gives -1.0, 0.0 and 1.0.

Example:
> (~> (dataframe (list (series '(-2 0 3) #:name "i") (series '(-2.5 0.0 3.5) #:name "f")))
      (select (sign "i") (sign "f")))

shape: (3, 2)

┌─────┬──────┐

│ i   ┆ f    │

│ --- ┆ ---  │

│ i64 ┆ f64  │

╞═════╪══════╡

│ -1  ┆ -1.0 │

│ 0   ┆ 0.0  │

│ 1   ┆ 1.0  │

└─────┴──────┘

procedure

(is-between x lower upper [#:closed closed]) → Expr-ptr?

  x : (or/c Expr-ptr? string?)
  lower : any/c
  upper : any/c
  closed : (or/c 'both 'left 'right 'none) = 'both

procedure

(is-in x rhs) → Expr-ptr?

  x : (or/c Expr-ptr? string?)
  rhs : (or/c list? series? Expr-ptr?)
Range and membership predicates (.is_between, .is_in). Bounds are lifted with lit, which has no date spelling; parse a string instead: (str->date (lit "1982-12-31")). A list rhs holds integers, reals, strings, symbols (read as strings, as lit reads them) or booleans, all of one kind.

procedure

(dt-year x) → Expr-ptr?

  x : (or/c Expr-ptr? string?)

procedure

(dt-month x) → Expr-ptr?

  x : (or/c Expr-ptr? string?)

procedure

(dt-day x) → Expr-ptr?

  x : (or/c Expr-ptr? string?)

procedure

(dt-hour x) → Expr-ptr?

  x : (or/c Expr-ptr? string?)

procedure

(dt-minute x) → Expr-ptr?

  x : (or/c Expr-ptr? string?)

procedure

(dt-second x) → Expr-ptr?

  x : (or/c Expr-ptr? string?)
Temporal component accessors on a date or datetime column (.dt.year() and friends). Also exported, with the same shape: dt-iso-year, dt-quarter, dt-week, dt-weekday, dt-ordinal-day, dt-is-leap-year, dt-date, dt-time, dt-millisecond, dt-microsecond, dt-nanosecond, dt-timestamp, dt-strftime, dt-truncate.

procedure

(str-extract x    
  pattern    
  [#:group-index group-index]) → Expr-ptr?
  x : (or/c Expr-ptr? string?)
  pattern : string?
  group-index : exact-nonnegative-integer? = 1

procedure

(str->date x    
  [#:format format    
  #:strict strict    
  #:exact exact    
  #:cache cache]) → Expr-ptr?
  x : (or/c Expr-ptr? string?)
  format : (or/c string? #f) = #f
  strict : boolean? = #t
  exact : boolean? = #t
  cache : boolean? = #t

procedure

(str->datetime x    
  [#:format format    
  #:unit unit    
  #:strict strict    
  #:exact exact    
  #:cache cache]) → Expr-ptr?
  x : (or/c Expr-ptr? string?)
  format : (or/c string? #f) = #f
  unit : (or/c 'milliseconds 'microseconds 'nanoseconds)
   = 'microseconds
  strict : boolean? = #t
  exact : boolean? = #t
  cache : boolean? = #t
str-extract returns capture group #:group-index of the first regex match (.str.extract). str->date and str->datetime parse strings with a chrono strptime #:format, inferred when omitted (.str.to_date, .str.to_datetime); #:strict #f yields null instead of raising on unparseable values.

2.1.1 Shadowed bindings🔗ℹ

(require polars) re-exports a handful of generic operations whose names also live in racket/base (min, max, sort, filter, >, <, >=, <=, =) and in racket/list (first, last, count, group-by). Under #lang racket/base this is seamless — these names are either not bound (so polars simply provides them) or bound only by the module language (which an explicit require silently shadows), and the polars versions intentionally fall back to the numeric/list behaviour for non-frame arguments.

A conflict arises only when another module providing the same name is also required explicitly — most commonly racket/list. Resolve it with the usual require sub-forms:

; keep polars' first/last/count/group-by, drop racket/list's:
(require (except-in racket/list first last count group-by) polars)
 
; keep racket/list's, reach polars' under a prefix:
(require racket/list (prefix-in pl: polars))
; then (pl:first (col "v")) for the Expr, (first '(1 2 3)) for the list
 
; keep polars', reach racket/list's under a prefix:
(require polars (prefix-in list: racket/list))

2.2 Series🔗ℹ

A series wraps a typed column and prints in the REPL the way Polars prints it; series? is its predicate. (The underlying foreign pointer is an implementation detail and not part of the public series API.)

procedure

(series? v) → boolean?

  v : any/c
Returns #t if v is a series.

procedure

(series elements [#:name name #:dtype dtype]) → series?

  elements : (or/c list? vector?)
  name : string? = ""
  dtype : (or/c #f symbol? pair?) = #f
Builds a series from a list or vector. When #:dtype is omitted the dtype is inferred from the elements; otherwise it is taken from dtype. Both short spellings ('i32, 'f64, 'str, 'bool) and canonical symbols ('int32, 'float64, 'string, 'boolean) are accepted. Use polars-null for missing values. Exact integers are coerced to flonums when the target dtype is floating point. Symbols infer 'categorical; a 'categorical or Enum (define-enum) series is built from strings or symbols alike, and an Enum raises on a value outside its categories (Categorical, Enum and Decimal).

Examples:
> (series '(1 2 3) #:name "ints")

shape: (3,)

Series: 'ints' [i64]

[

1

2

3

]

> (series '(1.5 2.5) #:name "floats" #:dtype 'f32)

shape: (2,)

Series: 'floats' [f32]

[

1.5

2.5

]

> (series (list 1 polars-null 3) #:name "with-null")

shape: (3,)

Series: 'with-null' [i64]

[

1

null

3

]

> (series '(IAH ATL IAH) #:name "dest")

shape: (3,)

Series: 'dest' [cat]

[

"IAH"

"ATL"

"IAH"

]

> (define-enum severity debug info error)
> (series '("debug" "error") #:dtype severity)

shape: (2,)

Series: '' [enum]

[

"debug"

"error"

]

> (series '(debug fatal) #:dtype severity)

series: cannot convert to '(enum debug info error):

conversion from `str` to `enum` failed in column '' for 1

out of 2 values: ["fatal"]

Ensure that all values in the input column are present in

the categories of the enum datatype.

procedure

(series->string s) → string?

  s : series?
Renders s in Polars’ series format (a shape line, a Series: ’name’ [dtype] line, then the bracketed values, truncated to the first and last five when longer than ten). This is also what a series prints as in the REPL.

value

polars-null : any/c

procedure

(polars-null? v) → boolean?

  v : any/c
polars-null is the sentinel marking a missing value: pass it among the elements given to series to produce nulls, and it is what ref returns for a null entry, and what the conversions in Converting to Racket values return by default. polars-null? tests for it.

procedure

(dtype s) → (or/c symbol? pair?)

  s : has-dtype?

procedure

(len x) → exact-nonnegative-integer?

  x : sized?

procedure

(null-count s) → exact-nonnegative-integer?

  s : has-null-count?
Generic series accessors. dtype returns the canonical dtype symbol (e.g. 'int32, 'float64, 'categorical) or list ('(datetime milliseconds #f), '(enum low mid high), '(decimal 10 2)). len returns the number of elements (and, on a dataframe, the number of rows). null-count returns the number of null entries.

procedure

(series-name s) → string?

  s : series?
Returns the name of s, as Polars’ Series.name; a column taken from a dataframe is named after the column.

> (series-name (series '(1 2) #:name "ints"))

"ints"

procedure

(rename s new-name) → series?

  s : series?
  new-name : string?

procedure

(rename! s new-name) → void?

  s : series?
  new-name : string?

procedure

(clone s) → series?

  s : series?

procedure

(series-clone s) → series?

  s : series?
rename! renames a series in place (matching Polars), returning void as is conventional for ! mutators; rename returns a renamed copy and leaves the original untouched. clone (and its series-specific alias series-clone) returns an independent copy.

2.2.1 Converting to Racket values🔗ℹ

These copy a column out of Polars in one foreign call rather than one per element: Racket allocates a buffer of the column’s native type, Rust copies the values into it, and Racket builds its values from the buffer and frees it. Nothing crosses the boundary to be freed later. At its peak a conversion holds that buffer (one native value per row, plus a byte per row when the column has nulls) beside the result it builds; in-series instead converts 4096 rows at a time. Each element comes out as ref returns it:

dtype

  

element

integer dtypes

  

exact-integer?

'float32, 'float64

  

flonum?

'boolean

  

boolean?

'string

  

string?

'categorical, '(enum cat ...)

  

symbol?

'(decimal precision scale)

  

an exact rational, as exact?

'date

  

a gregor date

'(datetime unit tz)

  

a gregor datetime, floored to the second

'(duration unit)

  

a gregor period in that unit

'time

  

a gregor time

'null

  

the null value

A null entry becomes the #:null value. A series of any other dtype raises exn:fail:contract naming the dtype, even when every entry is null. The Interoperability chapter of the guide walks through all of them.

> (series->list (series (list 1.5 polars-null)))

'(1.5 polars-null)

> (series->list (series (list "a" polars-null "")))

'("a" polars-null "")

> (series->list (series (list #t #f polars-null)))

'(#t #f polars-null)

> (series->list (series (list (datetime 2024 1 2 3 4 5) polars-null)))

'(#<datetime 2024-01-02T03:04:05> polars-null)

> (series->list (cast (series '(19724) #:dtype 'i32) 'date))

'(#<date 2024-01-02>)

> (series->list (cast (series '(11045000000000) #:dtype 'i64) 'time))

'(#<time 03:04:05>)

> (series->list (cast (series '(1500) #:dtype 'i64) '(duration milliseconds)))

'(#<period of 1500 milliseconds>)

> (series->list (cast (series (list "UA" polars-null "UA")) 'categorical))

'(UA polars-null UA)

> (series->list (ref (read-parquet "produce.parquet") "price"))

'(5/4 4/5 polars-null 12)

> (series->list (cast (series '("a") #:name "b") 'binary))

series->list: unsupported dtype

  series: "b"

  dtype: 'binary

procedure

(series->list s [#:null null-value]) → list?

  s : series?
  null-value : any/c = polars-null

procedure

(series->vector s [#:null null-value]) → vector?

  s : series?
  null-value : any/c = polars-null
Returns the elements of s in order, as a fresh list or a fresh mutable vector, with null-value in place of each null entry. Mirrors Polars’ Series.to_list(), with polars-null for None.

> (define s (series (list 3 polars-null 1) #:name "x"))
> (series->list s)

'(3 polars-null 1)

> (series->list s #:null 'missing)

'(3 missing 1)

> (series->vector s)

'#(3 polars-null 1)

> (series->vector s #:null 0)

'#(3 0 1)

procedure

(series->f64vector s [#:null null-value]) → f64vector?

  s : series?
  null-value : (or/c real? 'error) = +nan.0
Copies a numeric series into a fresh f64vector, the ffi/vector type that foreign code takes. Integers become the nearest flonum, as exact->inexact gives, and booleans become 1.0 and 0.0. A null entry becomes null-value as a flonum, as Series.to_numpy() gives nan; with 'error a null raises exn:fail:contract naming its row. Any other dtype raises exn:fail:contract naming the dtype.

The result is ordinary garbage-collected memory, which Racket CS may move: pass it to a foreign call that is not #:blocking?, and do not let foreign code keep the pointer past the call.

> (define xs (series (list 1 polars-null 3) #:name "x"))
> (f64vector->list (series->f64vector xs))

'(1.0 +nan.0 3.0)

> (f64vector->list (series->f64vector xs #:null 0))

'(1.0 0.0 3.0)

> (f64vector->list (series->f64vector (series (list #t #f))))

'(1.0 0.0)

> (f64vector->list (series->f64vector (series '(1 2)) #:null 'error))

'(1.0 2.0)

> (series->f64vector xs #:null 'error)

series->f64vector: null value

  series: "x"

  row: 1

> (series->f64vector (series '("a") #:name "s"))

series->f64vector: not a numeric series

  series: "s"

  dtype: 'string

procedure

(in-series s [#:null null-value]) → sequence?

  s : series?
  null-value : any/c = polars-null
Returns a sequence of the elements of s, converted as by series->list but 4096 rows at a time, so it never holds more than one block’s buffer and a loop that stops early converts little more than it reads. A series is itself a sequence: (for ([x s]) ....) iterates as (in-series s) does.

> (for/list ([x (in-series xs)]) x)

'(1 polars-null 3)

> (for/sum ([x (in-series xs #:null 0)]) x)

4

> (for/list ([x (series '("a" "b"))]) (string-upcase x))

'("A" "B")

> (for/first ([x (in-series (series (build-list 100000 values)))]
              #:when (> x 41))
    x)

42

2.2.2 Categorical, Enum and Decimal🔗ℹ

A 'categorical column stores each distinct string once and a code per row, as Polars’ Categorical; an Enum column (dtype '(enum cat ...), defined with define-enum) does the same over categories declared up front, in order, as pl.Enum. Both read back as symbols, which Racket interns: a symbol is already the dictionary encoding. The codes stay inside Polars. Every categorical column in the process shares them, and they restart once the last one is dropped, so each conversion fetches the strings afresh.

  • Build one with series (a list of symbols infers 'categorical), cast, or read-csv’s #:schema-overrides ('categorical only).

  • A categorical sorts and compares by its strings; an Enum by the declared order of its categories.

  • A value outside an Enum’s categories raises, whether it is built, cast or compared; a categorical takes any string.

  • Two categorical columns share one encoding, so they join, stack and compare without re-encoding; two Enums with the same categories are the same dtype.

  • describe gives a categorical or Enum column count and null_count only, as Python does.

syntax

(define-enum id category ...+)

 
category = id
  | string
Binds id to the Enum dtype with the given categories, in order: the datum '(enum category ...) that dtype reports for such a column, so equal? compares the two. A string category is the symbol of that string, for a name that is not an identifier. A duplicate category, or none, is a syntax error. The datum itself is accepted wherever a dtype is (pl.Enum([...])).

> (define-enum log-levels debug info warning error)
> log-levels

'(enum debug info warning error)

> (define levels (series '(debug info debug error) #:name "level" #:dtype log-levels))
> (equal? (dtype levels) log-levels)

#t

> (series->list (cast (series '("warning" "info")) log-levels))

'(warning info)

> (define-enum sizes small "Very High")
> sizes

'(enum small |Very High|)

> (define-enum twice debug info debug)

eval:177:0: define-enum: duplicate enum category

  at: debug

  in: (define-enum twice debug info debug)

> (series '(info fatal) #:dtype log-levels)

series: cannot convert to '(enum debug info warning error):

conversion from `str` to `enum` failed in column '' for 1

out of 2 values: ["fatal"]

Ensure that all values in the input column are present in

the categories of the enum datatype.

A '(decimal precision scale) column holds exact decimals, which ref and the conversions read as exact rationals. Decimal columns come from Parquet; cast reads them into other dtypes. API gap: no #:dtype or cast to a Decimal.

> (define logs
    (dataframe
     (list (series '(debug info debug error) #:name "level" #:dtype log-levels)
           (series '(api db api db) #:name "source"))))
> (for/list ([name (column-names logs)]) (dtype (ref logs name)))

'((enum debug info warning error) categorical)

> (filter logs (> (col "level") 'info))

shape: (1, 2)

┌───────┬────────┐

│ level ┆ source │

│ ---   ┆ ---    │

│ enum  ┆ cat    │

╞═══════╪════════╡

│ error ┆ db     │

└───────┴────────┘

> (sort logs "level")

shape: (4, 2)

┌───────┬────────┐

│ level ┆ source │

│ ---   ┆ ---    │

│ enum  ┆ cat    │

╞═══════╪════════╡

│ debug ┆ api    │

│ debug ┆ api    │

│ info  ┆ db     │

│ error ┆ db     │

└───────┴────────┘

> (sort logs "source")

shape: (4, 2)

┌───────┬────────┐

│ level ┆ source │

│ ---   ┆ ---    │

│ enum  ┆ cat    │

╞═══════╪════════╡

│ debug ┆ api    │

│ debug ┆ api    │

│ info  ┆ db     │

│ error ┆ db     │

└───────┴────────┘

> (ref (ref logs "source") 1)

'db

> (select logs (col 'categorical))

shape: (4, 1)

┌────────┐

│ source │

│ ---    │

│ cat    │

╞════════╡

│ api    │

│ db     │

│ api    │

│ db     │

└────────┘

> (select logs (> (col "level") 'fatal))

lazyframe-collect: failed to collect the query: conversion

from `str` to `enum` failed for value "fatal"

> (define produce (read-parquet "produce.parquet"))
> (dtype (ref produce "price"))

'(decimal 10 2)

> (for/sum ([price (ref produce "price")] #:unless (polars-null? price)) price)

281/20

2.2.3 dtype promotion🔗ℹ

Reductions follow a simple, predictable rule. The widening order, narrow to wide, is

  • 'int8 < 'int16 < 'int32 < 'int64

  • 'uint8 < 'uint16 < 'uint32 < 'uint64

  • any integer < 'float32 < 'float64

sum, min and max preserve the input dtype. mean promotes to 'float64. Use series-cast to change a series’ dtype explicitly.

2.2.4 Low-level Series API🔗ℹ

The generic layer is built on monomorphic, dtype-suffixed bindings that operate directly on the foreign series. They remain exported. A series wrapper is accepted anywhere one of them expects a series (the wrapper marshals transparently, and satisfies Series-ptr?), but what they return is the raw foreign pointer, not a wrapper — so the results do not print in Polars’ format and do not answer to series?. Prefer series and the generic operations above; reach for these when you need a specific dtype or a specific typed result.

procedure

(Series-ptr? v) → boolean?

  v : any/c
Recognises a foreign series pointer. Both raw pointers returned by the low-level constructors and series wrappers satisfy it.

procedure

(series-new-i8 name values) → Series-ptr?

  name : string?
  values : list?

procedure

(series-new-i16 name values) → Series-ptr?

  name : string?
  values : list?

procedure

(series-new-i32 name values) → Series-ptr?

  name : string?
  values : list?

procedure

(series-new-i64 name values) → Series-ptr?

  name : string?
  values : list?

procedure

(series-new-u8 name values) → Series-ptr?

  name : string?
  values : list?

procedure

(series-new-u16 name values) → Series-ptr?

  name : string?
  values : list?

procedure

(series-new-u32 name values) → Series-ptr?

  name : string?
  values : list?

procedure

(series-new-u64 name values) → Series-ptr?

  name : string?
  values : list?

procedure

(series-new-f32 name values) → Series-ptr?

  name : string?
  values : list?

procedure

(series-new-f64 name values) → Series-ptr?

  name : string?
  values : list?

procedure

(series-new-bool name values) → Series-ptr?

  name : string?
  values : list?

procedure

(series-new-str name values) → Series-ptr?

  name : string?
  values : list?
Build a series of the dtype named by the suffix from a list of values. polars-null among the values produces a null entry. Unlike series, no coercion happens: the integer constructors want exact integers, the float constructors want flonums (an exact 1 is rejected), series-new-bool wants booleans and series-new-str wants strings. Each constructor has a /vec sibling (series-new-i32/vec, series-new-f64/vec, …) that takes a vector instead of a list.

procedure

(series-sum-i32 s) → (or/c exact-integer? #f)

  s : Series-ptr?

procedure

(series-min-i32 s) → (or/c exact-integer? #f)

  s : Series-ptr?

procedure

(series-max-i32 s) → (or/c exact-integer? #f)

  s : Series-ptr?

procedure

(series-mean-i32 s) → (or/c flonum? #f)

  s : Series-ptr?

procedure

(series-sum-f64 s) → (or/c flonum? #f)

  s : Series-ptr?

procedure

(series-min-f64 s) → (or/c flonum? #f)

  s : Series-ptr?

procedure

(series-max-f64 s) → (or/c flonum? #f)

  s : Series-ptr?

procedure

(series-mean-f64 s) → (or/c flonum? #f)

  s : Series-ptr?
Typed reductions. The suffix names the dtype the series must have ('int32 or 'float64); applied to a series of any other dtype they return #f rather than converting, so (series-sum-i32 (series '(1 2))) is #f because series infers 'int64 for those elements. They also return #f when the reduction is undefined — the min, max or mean of a series whose entries are all null — while the sum of such a series is 0. The generic sum, min, max and mean dispatch on the dtype for you and are the preferred surface.

procedure

(series-cast s dtype) → Series-ptr?

  s : Series-ptr?
  dtype : (or/c symbol? pair?)
Returns a copy of s converted to dtype, given as a canonical dtype symbol ('int8 through 'int64, 'uint8 through 'uint64, 'float32, 'float64, 'boolean, 'string, 'binary, 'date, 'time, 'datetime, 'duration, 'null, 'categorical), as '(enum cat ...) with distinct symbols for the categories, or, for the two temporal dtypes with a time unit, as a list — '(datetime milliseconds), '(duration nanoseconds) — where the unit is one of 'nanoseconds, 'microseconds or 'milliseconds. Bare 'datetime and 'duration default to microseconds. Raises an error when Polars cannot perform the cast. The fluent cast wraps this for the generic layer.

procedure

(series-sort s    
  [#:descending descending    
  #:nulls-last nulls-last]) → Series-ptr?
  s : Series-ptr?
  descending : boolean? = #f
  nulls-last : boolean? = #f
Returns a sorted copy of s; the fluent sort on a series wraps it.

2.3 DataFrames🔗ℹ

A dataframe is a collection of equal-length named series. Like a series it is a wrapper value (dataframe?) carrying the column data; it prints as a Polars table, so display (or ~a, or the REPL) renders it with no separate display call.

procedure

(dataframe? v) → boolean?

  v : any/c
Returns #t if v is a dataframe.

procedure

(dataframe columns) → dataframe?

  columns : (listof series?)
Builds a dataframe from a list of equal-length series. The columns may be series wrappers built with series; their names become the column names.

Example:
> (dataframe (list (series '("a" "b") #:name "k")
                   (series '(1 2) #:name "v")))

shape: (2, 2)

┌─────┬─────┐

│ k   ┆ v   │

│ --- ┆ --- │

│ str ┆ i64 │

╞═════╪═════╡

│ a   ┆ 1   │

│ b   ┆ 2   │

└─────┴─────┘

procedure

(shape x) → (listof exact-nonnegative-integer?)

  x : has-shape?

procedure

(shape/values x) → 
exact-nonnegative-integer? ...
  x : has-shape?

procedure

(height d) → exact-nonnegative-integer?

  d : dataframe?

procedure

(width d) → exact-nonnegative-integer?

  d : dataframe?
shape returns the dimensions as a list — (list rows cols) for a dataframe and (list n) for a series — mirroring Polars’ shape tuples. shape/values returns the same dimensions as multiple values, for callers that want to bind them positionally with let-values or define-values. height and width return the row and column counts of a dataframe; height is also (len d).

procedure

(column-names d) → (listof string?)

  d : dataframe?

procedure

(column-name d i) → string?

  d : dataframe?
  i : exact-nonnegative-integer?
column-names returns all column names in order; column-name returns the name of the column at index i.

procedure

(ref x [key #:columns columns #:rows rows]) → any/c

  x : has-ref?
  key : (or/c exact-nonnegative-integer? string?) = absent
  columns : 
(or/c exact-nonnegative-integer? string?
      (listof (or/c exact-nonnegative-integer? string?)))
   = absent
  rows : any/c = absent
The generic element / column accessor. On a series, (ref s i) returns the element at index i. On a dataframe, a single selector — given positionally or as #:columns — returns that column (by name or index) as a series, and a list of selectors returns a column-projected dataframe. It is data-first, so it threads. #:rows is reserved for row slicing and currently raises an error. Provided by the gen:has-ref interface. For a whole column, series->list and its siblings convert in one pass where a ref loop makes one foreign call per element.

procedure

(describe x) → dataframe?

  x : (or/c series? dataframe?)
Mirrors Polars’ .describe(): returns a summary-statistics dataframe (which prints as a table). For a series the result has a "statistic" column and a "value" column, with rows adapted to the dtype — a numeric series gets "count", "null_count", "mean", "std", "min", "25%", "50%", "75%" and "max"; a temporal series (date, datetime, time, duration) drops "std"; a boolean series also drops the quantiles; a string series keeps "count", "null_count", "min" and "max"; any other dtype just the two counts. For a dataframe the result uses Polars’ fixed nine-row layout (a "statistic" column plus one column per input column), leaving a cell polars-null where a column has no value for that statistic.

Numeric, boolean, null and nested columns summarise as 'float64, every other column as strings; temporal values are written as Python prints them. Quantiles use nearest interpolation. Every statistic of every column comes from one query, so Polars computes the columns in parallel.

API gaps: a time-zone-aware datetime is written as its UTC clock time with no offset, where Python writes the local time and the offset; a binary column gets no "min" or "max".

Numeric, string and boolean columns, with nulls:

> (define flights
    (dataframe
     (list (series (list "UA" "AA" "UA" polars-null) #:name "carrier")
           (series (list 1400 733 polars-null 1089) #:name "distance")
           (series (list #t #f #t #t) #:name "on_time"))))
> (describe flights)

shape: (9, 4)

┌────────────┬─────────┬────────────┬─────────┐

│ statistic  ┆ carrier ┆ distance   ┆ on_time │

│ ---        ┆ ---     ┆ ---        ┆ ---     │

│ str        ┆ str     ┆ f64        ┆ f64     │

╞════════════╪═════════╪════════════╪═════════╡

│ count      ┆ 3       ┆ 3.0        ┆ 4.0     │

│ null_count ┆ 1       ┆ 1.0        ┆ 0.0     │

│ mean       ┆ null    ┆ 1074.0     ┆ 0.75    │

│ std        ┆ null    ┆ 333.752903 ┆ null    │

│ min        ┆ AA      ┆ 733.0      ┆ 0.0     │

│ 25%        ┆ null    ┆ 1089.0     ┆ null    │

│ 50%        ┆ null    ┆ 1089.0     ┆ null    │

│ 75%        ┆ null    ┆ 1400.0     ┆ null    │

│ max        ┆ UA      ┆ 1400.0     ┆ 1.0     │

└────────────┴─────────┴────────────┴─────────┘

Datetime, duration and date columns get a mean and quartiles; a date column’s mean is a datetime:

> (define times
    (~> (dataframe
         (list (series (list (datetime 2013 1 1 5) (datetime 2013 1 1 6)
                             (datetime 2013 1 2 7) (datetime 2013 1 3 8))
                       #:name "scheduled")
               (series (list (datetime 2013 1 1 5 12) (datetime 2013 1 1 5 57)
                             polars-null (datetime 2013 1 3 9 30))
                       #:name "departed")))
        (with-columns (alias (- (col "departed") (col "scheduled")) "delay")
                      (alias (cast "scheduled" 'date) "day"))))
> (describe times)

shape: (9, 5)

┌────────────┬─────────────────────┬─────────────────────┬──────────────────┬─────────────────────┐

│ statistic  ┆ scheduled           ┆ departed            ┆ delay            ┆ day                 │

│ ---        ┆ ---                 ┆ ---                 ┆ ---              ┆ ---                 │

│ str        ┆ str                 ┆ str                 ┆ str              ┆ str                 │

╞════════════╪═════════════════════╪═════════════════════╪══════════════════╪═════════════════════╡

│ count      ┆ 4                   ┆ 3                   ┆ 3                ┆ 4                   │

│ null_count ┆ 0                   ┆ 1                   ┆ 1                ┆ 0                   │

│ mean       ┆ 2013-01-02 00:30:00 ┆ 2013-01-01 22:53:00 ┆ 0:33:00          ┆ 2013-01-01 18:00:00 │

│ std        ┆ null                ┆ null                ┆ null             ┆ null                │

│ min        ┆ 2013-01-01 05:00:00 ┆ 2013-01-01 05:12:00 ┆ -1 day, 23:57:00 ┆ 2013-01-01          │

│ 25%        ┆ 2013-01-01 06:00:00 ┆ 2013-01-01 05:57:00 ┆ 0:12:00          ┆ 2013-01-01          │

│ 50%        ┆ 2013-01-02 07:00:00 ┆ 2013-01-01 05:57:00 ┆ 0:12:00          ┆ 2013-01-02          │

│ 75%        ┆ 2013-01-02 07:00:00 ┆ 2013-01-03 09:30:00 ┆ 1:30:00          ┆ 2013-01-02          │

│ max        ┆ 2013-01-03 08:00:00 ┆ 2013-01-03 09:30:00 ┆ 1:30:00          ┆ 2013-01-03          │

└────────────┴─────────────────────┴─────────────────────┴──────────────────┴─────────────────────┘

A series keeps only the rows its dtype has:

> (describe (series (list 3 1 polars-null 4 1 5) #:name "n"))

shape: (9, 2)

┌────────────┬──────────┐

│ statistic  ┆ value    │

│ ---        ┆ ---      │

│ str        ┆ f64      │

╞════════════╪══════════╡

│ count      ┆ 5.0      │

│ null_count ┆ 1.0      │

│ mean       ┆ 2.8      │

│ std        ┆ 1.788854 │

│ min        ┆ 1.0      │

│ 25%        ┆ 1.0      │

│ 50%        ┆ 3.0      │

│ 75%        ┆ 4.0      │

│ max        ┆ 5.0      │

└────────────┴──────────┘

> (describe (ref times "day"))

shape: (8, 2)

┌────────────┬─────────────────────┐

│ statistic  ┆ value               │

│ ---        ┆ ---                 │

│ str        ┆ str                 │

╞════════════╪═════════════════════╡

│ count      ┆ 4                   │

│ null_count ┆ 0                   │

│ mean       ┆ 2013-01-01 18:00:00 │

│ min        ┆ 2013-01-01          │

│ 25%        ┆ 2013-01-01          │

│ 50%        ┆ 2013-01-02          │

│ 75%        ┆ 2013-01-02          │

│ max        ┆ 2013-01-03          │

└────────────┴─────────────────────┘

A nested column (here the lists agg collects) and a null-dtype column report only their counts, as floats; a frame with no rows reports zero counts:

> (~> flights (group-by "carrier") (agg (col "distance")) describe)

shape: (9, 3)

┌────────────┬─────────┬──────────┐

│ statistic  ┆ carrier ┆ distance │

│ ---        ┆ ---     ┆ ---      │

│ str        ┆ str     ┆ f64      │

╞════════════╪═════════╪══════════╡

│ count      ┆ 2       ┆ 3.0      │

│ null_count ┆ 1       ┆ 0.0      │

│ mean       ┆ null    ┆ null     │

│ std        ┆ null    ┆ null     │

│ min        ┆ AA      ┆ null     │

│ 25%        ┆ null    ┆ null     │

│ 50%        ┆ null    ┆ null     │

│ 75%        ┆ null    ┆ null     │

│ max        ┆ UA      ┆ null     │

└────────────┴─────────┴──────────┘

> (~> flights (select (alias (cast "carrier" 'null) "nothing")) describe)

shape: (9, 2)

┌────────────┬─────────┐

│ statistic  ┆ nothing │

│ ---        ┆ ---     │

│ str        ┆ f64     │

╞════════════╪═════════╡

│ count      ┆ 0.0     │

│ null_count ┆ 4.0     │

│ mean       ┆ null    │

│ std        ┆ null    │

│ min        ┆ null    │

│ 25%        ┆ null    │

│ 50%        ┆ null    │

│ 75%        ┆ null    │

│ max        ┆ null    │

└────────────┴─────────┘

> (describe (head flights 0))

shape: (9, 4)

┌────────────┬─────────┬──────────┬─────────┐

│ statistic  ┆ carrier ┆ distance ┆ on_time │

│ ---        ┆ ---     ┆ ---      ┆ ---     │

│ str        ┆ str     ┆ f64      ┆ f64     │

╞════════════╪═════════╪══════════╪═════════╡

│ count      ┆ 0       ┆ 0.0      ┆ 0.0     │

│ null_count ┆ 0       ┆ 0.0      ┆ 0.0     │

│ mean       ┆ null    ┆ null     ┆ null    │

│ std        ┆ null    ┆ null     ┆ null    │

│ min        ┆ null    ┆ null     ┆ null    │

│ 25%        ┆ null    ┆ null     ┆ null    │

│ 50%        ┆ null    ┆ null     ┆ null    │

│ 75%        ┆ null    ┆ null     ┆ null    │

│ max        ┆ null    ┆ null     ┆ null    │

└────────────┴─────────┴──────────┴─────────┘

2.3.1 Converting to Racket values🔗ℹ

procedure

(dataframe->columns d 
  [#:columns columns 
  #:null null-value]) 
 → (listof (cons/c string? vector?))
  d : dataframe?
  columns : (listof string?) = (column-names d)
  null-value : any/c = polars-null
Returns each selected column, in the order of columns, paired with (series->vector column #:null null-value). Mirrors Polars’ DataFrame.to_dict(as_series=False). An unknown or repeated name, or a column of an unsupported dtype, raises exn:fail:contract naming the column.

> (define kv (dataframe (list (series '("a" "b") #:name "k")
                              (series (list 1 polars-null) #:name "v"))))
> (dataframe->columns kv)

'(("k" . #("a" "b")) ("v" . #(1 polars-null)))

> (dataframe->columns kv #:columns '("v" "k") #:null 0)

'(("v" . #(1 0)) ("k" . #("a" "b")))

> (dataframe->columns kv #:columns '("k" "k"))

dataframe->columns: duplicate column

  column: "k"

> (dataframe->columns kv #:columns '("nope"))

dataframe->columns: no such column

  column: "nope"

procedure

(dataframe->hash d 
  [#:columns columns 
  #:null null-value]) 
 → (and/c (hash/c string? vector?) immutable?)
  d : dataframe?
  columns : (listof string?) = (column-names d)
  null-value : any/c = polars-null
Like dataframe->columns, but returns an immutable hash from each selected column’s name to its vector, as Polars’ DataFrame.to_dict() gives a dict. The same names are checked and the same errors raised.

> (dataframe->hash kv)

'#hash(("k" . #("a" "b")) ("v" . #(1 polars-null)))

> (hash-ref (dataframe->hash kv #:null 0) "v")

'#(1 0)

> (dataframe->hash kv #:columns '("k"))

'#hash(("k" . #("a" "b")))

> (dataframe->hash kv #:columns '("nope"))

dataframe->hash: no such column

  column: "nope"

procedure

(in-dataframe-columns d [#:columns columns]) → sequence?

  d : dataframe?
  columns : (listof string?) = (column-names d)
Returns a sequence of the selected columns of d, in the order of columns, each as a series named after its column, as Polars’ DataFrame.iter_columns() does. Each column is fetched when the sequence reaches it, and is released once nothing refers to it. An unknown or repeated name raises exn:fail:contract naming the column. The dataframe itself is not a sequence.

> (for/list ([column (in-dataframe-columns kv)]) (series-name column))

'("k" "v")

> (for/list ([column (in-dataframe-columns kv #:columns '("v"))])
    (series->list column #:null 0))

'((1 0))

> (for/first ([column (in-dataframe-columns kv)]) column)

shape: (2,)

Series: 'k' [str]

[

"a"

"b"

]

> (in-dataframe-columns kv #:columns '("k" "k"))

in-dataframe-columns: duplicate column

  column: "k"

procedure

(in-dataframe-rows d    
  [#:columns columns    
  #:named? named?    
  #:null null-value    
  #:buffer-size buffer-size]) → sequence?
  d : dataframe?
  columns : (listof string?) = (column-names d)
  named? : boolean? = #f
  null-value : any/c = polars-null
  buffer-size : exact-positive-integer? = 512
Returns a sequence of the rows of d, as Polars’ DataFrame.iter_rows() does. Each row holds the selected columns in the order of columns: a fresh mutable vector of their values, or, when named? is true, an immutable hash from each column’s name to its value, as iter_rows(named=True) gives a dict. A value is what ref returns, as in Converting to Racket values, with null-value in place of a null. With no columns, each row is empty.

The rows are converted buffer-size at a time, as buffer_size does: each buffer is one bulk copy per column, never a foreign call per value. A loop holds one buffer’s values at a time, so memory stays bounded for any height, and one that stops early converts at most one buffer beyond the rows it reads. A larger buffer makes fewer calls and holds more values.

The columns are fetched and checked when in-dataframe-rows is called: an unknown or repeated name, or a column of an unsupported dtype, raises exn:fail:contract naming the column. The sequence holds those columns rather than d, and each iteration starts from the first row. A lazyframe is not accepted; collect it first. Python’s buffer_size=0, a row at a time, has no counterpart.

> (define trips
    (dataframe (list (series '(UA AA UA) #:name "carrier")
                     (series (list 2 polars-null -3) #:name "delay")
                     (cast (series (list (datetime 2013 1 1) (datetime 2013 1 1)
                                         (datetime 2013 1 2))
                                   #:name "day")
                           'date))))
> (for/list ([row (in-dataframe-rows trips)]) row)

'(#(UA 2 #<date 2013-01-01>)

  #(AA polars-null #<date 2013-01-01>)

  #(UA -3 #<date 2013-01-02>))

> (for/list ([row (in-dataframe-rows trips #:columns '("delay" "carrier") #:named? #t)])
    row)

'(#hash(("carrier" . UA) ("delay" . 2))

  #hash(("carrier" . AA) ("delay" . polars-null))

  #hash(("carrier" . UA) ("delay" . -3)))

> (for/sum ([row (in-dataframe-rows trips #:columns '("delay") #:null 0)])
    (vector-ref row 0))

-1

> (for/list ([row (in-dataframe-rows trips #:buffer-size 2)])
    (define-values (carrier delay day) (vector->values row))
    (list carrier (date->iso8601 day)))

'((UA "2013-01-01") (AA "2013-01-01") (UA "2013-01-02"))

> (in-dataframe-rows trips #:columns '("nope"))

in-dataframe-rows: no such column

  column: "nope"

> (in-dataframe-rows trips #:buffer-size 0)

in-dataframe-rows: contract violation

  expected: exact-positive-integer?

  given: 0

  in: the #:buffer-size argument of

      (->*

       (dataframe?)

       (#:buffer-size

        exact-positive-integer?

        #:columns

        (listof string?)

        #:named?

        boolean?

        #:null

        any/c)

       sequence?)

  contract from:

      <pkgs>/polars/private/generic/convert.rkt

  blaming: top-level

   (assuming the contract is correct)

  at: <pkgs>/polars/private/generic/convert.rkt:32:3

procedure

(dataframe->rows d 
  [#:columns columns 
  #:named? named? 
  #:null null-value]) 
 → (listof (or/c vector? (and/c hash? immutable?)))
  d : dataframe?
  columns : (listof string?) = (column-names d)
  named? : boolean? = #f
  null-value : any/c = polars-null
Returns every row of d in a list, each as in-dataframe-rows gives it: Polars’ DataFrame.rows(), and with named? true DataFrame.rows(named=True), which is DataFrame.to_dicts(). The same names are checked and the same errors raised.

> (dataframe->rows trips)

'(#(UA 2 #<date 2013-01-01>)

  #(AA polars-null #<date 2013-01-01>)

  #(UA -3 #<date 2013-01-02>))

> (dataframe->rows trips #:columns '("carrier") #:named? #t)

'(#hash(("carrier" . UA)) #hash(("carrier" . AA)) #hash(("carrier" . UA)))

> (dataframe->rows (head trips 0))

'()

> (dataframe->rows trips #:columns '("day" "day"))

dataframe->rows: duplicate column

  column: "day"

procedure

(dataframe->f64vector d 
  [#:columns columns 
  #:order order 
  #:null null-value]) 
 → 
f64vector?
exact-nonnegative-integer?
exact-nonnegative-integer?
  d : dataframe?
  columns : (listof string?) = (column-names d)
  order : (or/c 'fortran 'c) = 'fortran
  null-value : (or/c real? 'error) = +nan.0
Copies the selected columns into one fresh f64vector and returns it with its row and column counts, mirroring Polars’ DataFrame.to_numpy(). With 'fortran (column-major, the default) row i of column j is at index (+ (* j nrows) i); with 'c (row-major) it is at (+ (* i ncols) j). Each column converts as by series->f64vector. Every column’s dtype is checked before anything is copied, and a column that is not numeric raises naming the column and its dtype. With 'error, a null raises naming the first column in columns that has one, and that column’s first null row. As with series->f64vector, the buffer may move: hand it only to a foreign call that is not #:blocking?.

> (define xy (dataframe (list (series (list 1 2 polars-null) #:name "a")
                              (series '(0.5 1.5 2.5) #:name "b"))))
> (define-values (m nrows ncols) (dataframe->f64vector xy))
> (list nrows ncols)

'(3 2)

> (f64vector->list m)

'(1.0 2.0 +nan.0 0.5 1.5 2.5)

> (define-values (m/f rows/f cols/f) (dataframe->f64vector xy #:order 'fortran))
> (equal? (f64vector->list m/f) (f64vector->list m))

#t

> (define-values (m/c rows/c cols/c) (dataframe->f64vector xy #:order 'c))
> (f64vector->list m/c)

'(1.0 0.5 2.0 1.5 +nan.0 2.5)

> (define-values (b rows/b cols/b) (dataframe->f64vector xy #:columns '("b") #:null 'error))
> (f64vector->list b)

'(0.5 1.5 2.5)

> (define-values (z rows/z cols/z) (dataframe->f64vector xy #:null 0))
> (f64vector->list z)

'(1.0 2.0 0.0 0.5 1.5 2.5)

> (dataframe->f64vector xy #:null 'error)

dataframe->f64vector: null value

  column: "a"

  row: 2

> (dataframe->f64vector (dataframe (list (series '("p") #:name "s"))))

dataframe->f64vector: not a numeric column

  column: "s"

  dtype: 'string

2.3.2 Low-level DataFrame API🔗ℹ

The generic layer above is built on a set of monomorphic dataframe-* bindings that operate directly on the foreign dataframe. They remain exported and accept the dataframe wrapper (it marshals transparently); the generic operations are simply the preferred surface.

procedure

(dataframe-new columns) → dataframe?

  columns : (listof series?)

procedure

(dataframe-shape d) → 
exact-nonnegative-integer?
exact-nonnegative-integer?
  d : dataframe?

procedure

(dataframe-height d) → exact-nonnegative-integer?

  d : dataframe?

procedure

(dataframe-width d) → exact-nonnegative-integer?

  d : dataframe?

procedure

(dataframe-column d name) → series?

  d : dataframe?
  name : string?

procedure

(dataframe-column-name d i) → string?

  d : dataframe?
  i : exact-nonnegative-integer?

procedure

(dataframe-column-names d) → (listof string?)

  d : dataframe?

procedure

(dataframe-select d names) → dataframe?

  d : dataframe?
  names : (listof string?)

procedure

(display-dataframe d [out]) → void?

  d : dataframe?
  out : output-port? = (current-output-port)
The low-level dataframe operations underlying dataframe, shape, height, width, ref, column-name, and column-names. display-dataframe prints the Polars table to out; since a dataframe now prints itself, prefer plain display. dataframe-column raises an error naming the column when d has none.

Examples:
> (define scores (dataframe (list (series '(10 25 18) #:name "score" #:dtype 'i32))))
> (series-sum-i32 (dataframe-column scores "score"))

53

> (dataframe-column scores "points")

dataframe-column: no column named "points"

procedure

(DataFrame-ptr? v) → boolean?

  v : any/c
Recognises a foreign dataframe pointer. Both raw pointers returned by the low-level operations and dataframe wrappers satisfy it.

procedure

(dataframe-vstack top bottom) → DataFrame-ptr?

  top : DataFrame-ptr?
  bottom : DataFrame-ptr?
Stacks the rows of bottom beneath those of top, which must have the same columns in the same order, and returns the combined frame (Polars’ vstack). The fluent vstack is the wrapper-returning equivalent.

procedure

(dataframe-sort d 
  names 
  [#:descending descending 
  #:nulls-last nulls-last 
  #:maintain-order maintain-order]) 
 → DataFrame-ptr?
  d : DataFrame-ptr?
  names : (non-empty-listof string?)
  descending : (or/c boolean? (listof boolean?)) = #f
  nulls-last : (or/c boolean? (listof boolean?)) = #f
  maintain-order : boolean? = #f
Returns d sorted by the columns names; the fluent sort on a dataframe wraps it, and its entry describes the keywords. A column absent from d raises exn:fail.

2.3.3 Reading & writing🔗ℹ

procedure

(dataframe-write-csv d path) → void?

  d : dataframe?
  path : path-string?

procedure

(dataframe-write-parquet d path) → void?

  d : dataframe?
  path : path-string?

procedure

(dataframe-read-parquet path) → dataframe?

  path : path-string?

procedure

(dataframe-write-json-lines d path) → void?

  d : dataframe?
  path : path-string?

procedure

(dataframe-read-json-lines path) → dataframe?

  path : path-string?
Round-trip a dataframe through CSV, Parquet, or newline-delimited JSON; the fluent read-csv and friends are the surface. Like read-parquet, dataframe-read-parquet accepts a glob pattern.

procedure

(dataframe-read-csv path 
  [#:has-header has-header 
  #:separator separator 
  #:quote-char quote-char 
  #:comment-prefix comment-prefix 
  #:skip-rows skip-rows 
  #:n-rows n-rows 
  #:null-values null-values 
  #:infer-schema-length infer-schema-length 
  #:schema-overrides schema-overrides 
  #:ignore-errors ignore-errors 
  #:try-parse-dates try-parse-dates 
  #:encoding encoding 
  #:glob glob]) 
 → DataFrame-ptr?
  path : path-string?
  has-header : boolean? = #t
  separator : (or/c csv-char/c #f) = #f
  quote-char : (or/c csv-char/c #f) = #\"
  comment-prefix : (or/c non-empty-string? #f) = #f
  skip-rows : exact-nonnegative-integer? = 0
  n-rows : (or/c exact-nonnegative-integer? #f) = #f
  null-values : (or/c string? (listof string?) #f) = #f
  infer-schema-length : (or/c exact-nonnegative-integer? #f)
   = 100
  schema-overrides : (and/c (listof (cons/c string? csv-dtype/c)) distinct-names?)
   = '()
  ignore-errors : boolean? = #f
  try-parse-dates : boolean? = #f
  encoding : (or/c 'utf8 'utf8-lossy) = 'utf8
  glob : boolean? = #t

procedure

(lazyframe-scan-csv path 
  [#:has-header has-header 
  #:separator separator 
  #:quote-char quote-char 
  #:comment-prefix comment-prefix 
  #:skip-rows skip-rows 
  #:n-rows n-rows 
  #:null-values null-values 
  #:infer-schema-length infer-schema-length 
  #:schema-overrides schema-overrides 
  #:ignore-errors ignore-errors 
  #:try-parse-dates try-parse-dates 
  #:encoding encoding 
  #:glob glob]) 
 → LazyFrame-ptr?
  path : path-string?
  has-header : boolean? = #t
  separator : (or/c csv-char/c #f) = #f
  quote-char : (or/c csv-char/c #f) = #\"
  comment-prefix : (or/c non-empty-string? #f) = #f
  skip-rows : exact-nonnegative-integer? = 0
  n-rows : (or/c exact-nonnegative-integer? #f) = #f
  null-values : (or/c string? (listof string?) #f) = #f
  infer-schema-length : (or/c exact-nonnegative-integer? #f)
   = 100
  schema-overrides : (and/c (listof (cons/c string? csv-dtype/c)) distinct-names?)
   = '()
  ignore-errors : boolean? = #f
  try-parse-dates : boolean? = #f
  encoding : (or/c 'utf8 'utf8-lossy) = 'utf8
  glob : boolean? = #t
The raw-pointer reader and scan under read-csv and scan-csv, with the same keywords and checks.

Examples:
> (dataframe-height (dataframe-read-csv "flights.tsv" #:separator #\tab #:null-values "NA"))

102

> (dataframe-height (lazyframe-collect (lazyframe-scan-csv "parts/*.csv")))

6

2.4 Lazy frames🔗ℹ

A lazyframe is a query plan: a sequence of operations over a frame that Polars optimises as a whole and runs only when asked to collect. The low-level surface mirrors the eager dataframe-* bindings and, like them, returns raw foreign pointers; the fluent lazy and collect are the wrapper-returning equivalents.

procedure

(LazyFrame-ptr? v) → boolean?

  v : any/c

procedure

(lazyframe? v) → boolean?

  v : any/c
LazyFrame-ptr? recognises a foreign lazyframe pointer, raw or wrapped. lazyframe? recognises only the wrapper produced by the fluent lazy.

procedure

(dataframe-lazy df) → LazyFrame-ptr?

  df : DataFrame-ptr?

procedure

(lazyframe-collect lf) → DataFrame-ptr?

  lf : LazyFrame-ptr?
dataframe-lazy starts a plan from an in-memory frame (df.lazy()); lazyframe-collect executes a plan and returns the resulting frame (lf.collect()).

procedure

(lazyframe-select lf exprs) → LazyFrame-ptr?

  lf : LazyFrame-ptr?
  exprs : (listof Expr-ptr?)

procedure

(lazyframe-with-columns lf exprs) → LazyFrame-ptr?

  lf : LazyFrame-ptr?
  exprs : (listof Expr-ptr?)

procedure

(lazyframe-filter lf predicate) → LazyFrame-ptr?

  lf : LazyFrame-ptr?
  predicate : Expr-ptr?

procedure

(lazyframe-group-by-agg lf keys aggs) → LazyFrame-ptr?

  lf : LazyFrame-ptr?
  keys : (listof (or/c string? Expr-ptr?))
  aggs : (listof Expr-ptr?)
The lazy forms of the Eager expression contexts. Each appends a step to the plan and returns the extended plan; nothing runs until lazyframe-collect.

procedure

(lazyframe-sort lf 
  names 
  [#:descending descending 
  #:nulls-last nulls-last 
  #:maintain-order maintain-order]) 
 → LazyFrame-ptr?
  lf : LazyFrame-ptr?
  names : (non-empty-listof string?)
  descending : (or/c boolean? (listof boolean?)) = #f
  nulls-last : (or/c boolean? (listof boolean?)) = #f
  maintain-order : boolean? = #f
Appends a sort by the columns names to the plan; the fluent sort on a lazyframe wraps it. An absent column is reported at lazyframe-collect.

procedure

(lazyframe-join left    
  right    
  [#:on on    
  #:left-on left-on    
  #:right-on right-on    
  #:how how]) → LazyFrame-ptr?
  left : LazyFrame-ptr?
  right : LazyFrame-ptr?
  on : (or/c #f (listof string?)) = #f
  left-on : (or/c #f (listof string?)) = #f
  right-on : (or/c #f (listof string?)) = #f
  how : (or/c 'inner 'left 'outer 'full 'cross) = 'inner
Joins two plans. Give the key columns either as one list with #:on, when they have the same names on both sides, or as parallel #:left-on and #:right-on lists. #:how selects the join kind; 'outer and 'full are synonyms, and a 'cross join takes no keys. Omitting the keys for any other kind is an error. Collect the result with lazyframe-collect:

(lazyframe-collect
 (lazyframe-join (dataframe-lazy users) (dataframe-lazy orders)
                 #:on '("uid") #:how 'inner))

2.5 Low-level expression API🔗ℹ

The monomorphic expr-* layer that the operators in Operators and pipelines are built from. You rarely need these names directly: expr-gt underlies >, expr-add underlies +, expr-sum underlies the expression arm of sum. Reach for them when a generic name is shadowed in your module, or when you want to be explicit that an expression — rather than a number — is being built.

procedure

(expr-alias e name) → Expr-ptr?

  e : Expr-ptr?
  name : string?
Names the column an expression produces, matching .alias. The generic spelling is alias, which is the one to reach for: (alias (sum (col "value")) "total").

procedure

(expr-col name) → Expr-ptr?

  name : string?
The column reference col is built on: name is taken literally, except that a name of the form ^...$ is a regex projection, which is how the regexp arm of col is spelled.

Examples:
> (expr-col "weight")

col("weight")

> (expr-col "^.*ght$")

cs.matches("^.*ght$")

procedure

(expr-all) → Expr-ptr?

procedure

(expr-exclude e names) → Expr-ptr?

  e : multi-column-expr?
  names : (non-empty-listof (or/c string? regexp?))

procedure

(expr-dtype-col dtype) → Expr-ptr?

  dtype : dtype-spec?
The selector leaves under all, exclude and the dtype arm of col: pl.all(), .exclude(...) and pl.col(pl.Float64). expr-exclude takes its names as one list where the generic exclude is variadic. A regexp col needs no entry point of its own: it is expr-col with the pattern rendered in Polars’ ^...$ form.

Examples:
> (expr-all)

cs.all()

> (expr-exclude (expr-all) (list "id" #rx"^w"))

[cs.all() - [cs.matches("^(?s).*(?:^w).*$") | cs.by_name('id', require_all=false)]]

> (expr-dtype-col 'f64)

cs.by_dtype([Float64])

> (select people (expr-exclude (expr-dtype-col 'float64) (list "height")))

shape: (3, 1)

┌────────┐

│ weight │

│ ---    │

│ f64    │

╞════════╡

│ 57.9   │

│ 72.5   │

│ 53.6   │

└────────┘

procedure

(expr-meta-output-name e) → string?

  e : Expr-ptr?

procedure

(expr-meta-root-names e) → (listof string?)

  e : Expr-ptr?

procedure

(expr-meta-eq? a b) → boolean?

  a : Expr-ptr?
  b : Expr-ptr?
The expression-only forms of meta-output-name, meta-root-names and meta-eq?, which are the ones to write: they also accept a column name.

Examples:
> (~> (col "a") (expr-alias "b") expr-meta-output-name)

"b"

> (~> (col "a") (expr-add (col "b")) expr-meta-root-names)

'("a" "b")

> (expr-meta-eq? (col "a") (expr-col "a"))

#t

procedure

(expr-over e keys) → Expr-ptr?

  e : Expr-ptr?
  keys : (listof (or/c string? Expr-ptr?))
The window expression over is built on (Expr.over), taking its keys as one list where over is variadic.

Examples:
> (~> (col "v") expr-sum (expr-over (list "k" (col "h"))))

col("v").sum().over([col("k"), col("h")])

> (~> khv (with-columns (~> (col "v") expr-sum (expr-over (list "k")) (expr-alias "total"))))

shape: (4, 4)

┌─────┬─────┬─────┬───────┐

│ k   ┆ h   ┆ v   ┆ total │

│ --- ┆ --- ┆ --- ┆ ---   │

│ str ┆ str ┆ i64 ┆ i64   │

╞═════╪═════╪═════╪═══════╡

│ a   ┆ x   ┆ 1   ┆ 6     │

│ a   ┆ y   ┆ 2   ┆ 6     │

│ a   ┆ x   ┆ 3   ┆ 6     │

│ b   ┆ x   ┆ 4   ┆ 4     │

└─────┴─────┴─────┴───────┘

procedure

(expr-add a b) → Expr-ptr?

  a : any/c
  b : any/c

procedure

(expr-sub a b) → Expr-ptr?

  a : any/c
  b : any/c

procedure

(expr-mul a b) → Expr-ptr?

  a : any/c
  b : any/c

procedure

(expr-div a b) → Expr-ptr?

  a : any/c
  b : any/c

procedure

(expr-mod a b) → Expr-ptr?

  a : any/c
  b : any/c
Element-wise arithmetic. At least one operand is normally an expression; the other may be a scalar, which is lifted with lit. The generic +, -, * and / dispatch to these when given an expression and are the preferred surface, so write (* (col "value") 2) — which reads like col("value") * 2 — rather than calling expr-mul directly.

procedure

(expr-gt a b) → Expr-ptr?

  a : any/c
  b : any/c

procedure

(expr-lt a b) → Expr-ptr?

  a : any/c
  b : any/c

procedure

(expr-ge a b) → Expr-ptr?

  a : any/c
  b : any/c

procedure

(expr-le a b) → Expr-ptr?

  a : any/c
  b : any/c

procedure

(expr-eq a b) → Expr-ptr?

  a : any/c
  b : any/c

procedure

(expr-ne a b) → Expr-ptr?

  a : any/c
  b : any/c
Element-wise comparisons producing a boolean expression; scalars are lifted with lit. The generic >, <, >=, <=, = and != dispatch to these when given an expression and are the preferred surface: write (> (col "value") 15) for col("value") > 15.

procedure

(expr-and a b) → Expr-ptr?

  a : any/c
  b : any/c

procedure

(expr-or a b) → Expr-ptr?

  a : any/c
  b : any/c

procedure

(expr-xor a b) → Expr-ptr?

  a : any/c
  b : any/c

procedure

(expr-not e) → Expr-ptr?

  e : Expr-ptr?
Element-wise boolean logic over boolean expressions, for combining predicates. The generic and, or, xor and not dispatch to these when given an expression and are the preferred surface: (and (> (col "value") 15) (< (col "cost") 3.0)).

procedure

(expr-sum e) → Expr-ptr?

  e : Expr-ptr?

procedure

(expr-mean e) → Expr-ptr?

  e : Expr-ptr?

procedure

(expr-min e) → Expr-ptr?

  e : Expr-ptr?

procedure

(expr-max e) → Expr-ptr?

  e : Expr-ptr?

procedure

(expr-median e) → Expr-ptr?

  e : Expr-ptr?

procedure

(expr-count e) → Expr-ptr?

  e : Expr-ptr?

procedure

(expr-n-unique e) → Expr-ptr?

  e : Expr-ptr?

procedure

(expr-first e) → Expr-ptr?

  e : Expr-ptr?

procedure

(expr-last e) → Expr-ptr?

  e : Expr-ptr?

procedure

(expr-std e [#:ddof ddof]) → Expr-ptr?

  e : Expr-ptr?
  ddof : exact-nonnegative-integer? = 1

procedure

(expr-var e [#:ddof ddof]) → Expr-ptr?

  e : Expr-ptr?
  ddof : exact-nonnegative-integer? = 1
Aggregations. Each reduces the column e evaluates to — over the whole frame in a select, or per group inside group_by/agg. expr-count counts the non-null entries, as Polars’ .count() does. expr-std and expr-var take a #:ddof degrees-of-freedom adjustment, defaulting to 1. The generic sum, mean, min, max, median, count, n-unique, first, last, std and var dispatch to these when given an expression or a column name.

procedure

(expr-sort e    
  [#:descending descending    
  #:nulls-last nulls-last]) → Expr-ptr?
  e : Expr-ptr?
  descending : boolean? = #f
  nulls-last : boolean? = #f

procedure

(expr-sort-by e    
  #:by by    
  [#:descending descending    
  #:nulls-last nulls-last    
  #:maintain-order maintain-order]) → Expr-ptr?
  e : Expr-ptr?
  by : 
(or/c string? Expr-ptr?
      (non-empty-listof (or/c string? Expr-ptr?)))
  descending : (or/c boolean? (listof boolean?)) = #f
  nulls-last : (or/c boolean? (listof boolean?)) = #f
  maintain-order : boolean? = #f
The expression-only forms under sort on an expression and sort-by. Write those instead: they also accept a column name.

2.5.1 Eager expression contexts🔗ℹ

These run expressions against a dataframe and return a new frame in one step. Each is the eager convenience over the corresponding lazy operation in Lazy frames: it converts with dataframe-lazy, applies the operation, and lazyframe-collects. Like the rest of the low-level layer they accept a dataframe wrapper but return a raw DataFrame-ptr?; the fluent select, with-columns, filter and group-by/agg are the wrapper-returning equivalents.

procedure

(dataframe-select-exprs df exprs) → DataFrame-ptr?

  df : DataFrame-ptr?
  exprs : (listof Expr-ptr?)

procedure

(dataframe-with-columns df exprs) → DataFrame-ptr?

  df : DataFrame-ptr?
  exprs : (listof Expr-ptr?)

procedure

(dataframe-filter-expr df predicate) → DataFrame-ptr?

  df : DataFrame-ptr?
  predicate : Expr-ptr?
dataframe-select-exprs evaluates exprs and returns a frame containing only the resulting columns (df.select(...)). dataframe-with-columns evaluates them and adds (or replaces) the resulting columns alongside the existing ones (df.with_columns(...)). dataframe-filter-expr keeps the rows for which the boolean predicate holds (df.filter(...)).

procedure

(dataframe-group-by-agg df keys aggs) → DataFrame-ptr?

  df : DataFrame-ptr?
  keys : (listof (or/c string? Expr-ptr?))
  aggs : (listof Expr-ptr?)
Groups df by keys — column names, or expressions — and evaluates each aggregation in aggs once per group, returning a frame with one row per group (df.group_by(...).agg(...)). The row order of the result is not guaranteed.

2.6 Generic interfaces🔗ℹ

The high-level operations are small, purpose-named racket/generic interfaces. A wrapper implements the interface for each capability it has — a series and a dataframe both have a len and a shape, so both implement gen:sized and gen:has-shape; only a series has a dtype. Each interface exports its method(s) and a predicate that recognises values implementing it.

syntax

gen:has-ref

procedure

(has-ref? v) → boolean?

  v : any/c
The ref capability (method: ref). Implemented by series and dataframes.

syntax

gen:sized

procedure

(sized? v) → boolean?

  v : any/c
The len capability (method: len). Implemented by series (number of elements) and dataframes (number of rows).

syntax

gen:has-shape

procedure

(has-shape? v) → boolean?

  v : any/c
The shape capability (method: shape). Implemented by series and dataframes.

syntax

gen:has-dtype

procedure

(has-dtype? v) → boolean?

  v : any/c
The dtype capability (method: dtype). Implemented by series.

syntax

gen:has-null-count

procedure

(has-null-count? v) → boolean?

  v : any/c
The null-count capability (method: null-count). Implemented by series.

A series is also a Racket sequence (through prop:sequence): a for clause, sequence? and the racket/sequence operations see its elements, converted a block of rows at a time as by in-series, with polars-null for a null entry. A dataframe is not a sequence; iterate over its rows with in-dataframe-rows or its columns with in-dataframe-columns, or convert them with dataframe->rows or dataframe->columns.

> (define ages (series (list 34 polars-null 51) #:name "age"))
> (sequence? ages)

#t

> (for/list ([age ages]) age)

'(34 polars-null 51)

> (for/sum ([age ages] #:unless (polars-null? age)) age)

85

> (sequence? (dataframe (list ages)))

#f