2 Reference
polars is written to read as ordinary Racket. The operators
you already know —
Underneath sits a monomorphic, dtype-suffixed layer (series-new-i32,
series-sum-f64, expr-add and friends). It is documented here
for completeness —
2.1 Operators and pipelines
This is the surface to write polars in. An
expression describes a column computation without running it, and
the operators below are the ordinary Racket ones —
Every operation takes the frame —
(~> df (filter (> (col "value") 15)) (group-by "group") (agg (alias (sum (col "value")) "sum_value")))
procedure
spec : (or/c string? regexp? dtype-spec?)
procedure
v : (or/c boolean? exact-integer? real? string? symbol?)
procedure
(dtype-spec? v) → boolean?
v : any/c
A string is the column of that name (pl.col("name")). A string of the form ^...$ is a Polars regex, in the syntax of Rust’s regex crate, as in Python.
A dtype is every column of that dtype (pl.col(pl.Float64)). dtype-spec? is any spelling series’ #:dtype accepts, so 'float64 and 'f64 alike. A bare 'datetime means microseconds, so match a column series built from gregor datetimes with '(datetime milliseconds) or with its dtype. 'categorical is every categorical column, and an '(enum ....) dtype the columns of exactly that Enum (Categorical, Enum and Decimal).
A regexp is every column whose name matches (pl.col("^sepal_.*$")), keeping the regexp’s Racket meaning: (col rx) selects exactly the names (regexp-match? rx name) accepts, in #rx and #px syntax alike. There are two exceptions. A \p{...} property class follows each side’s own version of the Unicode tables. And Racket’s own matcher misjudges some classes containing characters above U+00FF; there the selection follows the class as written. The crate has no lookaround, backreferences, atomic groups or conditionals; a regexp using them is rejected at collect. To select by such a regexp, match column-names in Racket and select the names, as in the last example below.
A multi-column col expands inside any expression to one output per matched column, in the frame’s column order, each keeping the matched column’s name; a frame with no match yields no columns.
lit lifts a Racket scalar to a literal expression: booleans, exact integers (32-bit when they fit, 64-bit otherwise), other reals (as 'float64) and strings. A symbol is the string of its name, so a categorical column compares with the symbols it reads back as. Every operator below lifts a non-expression operand with lit automatically, so it is rarely needed explicitly.
> (define people (dataframe (list (series '(1 2 3) #:name "id" #:dtype 'i32) (series '(57.9 72.5 53.6) #:name "weight") (series '(1.56 1.77 1.65) #:name "height")))) > (select people (* (col 'float64) 1.1))
shape: (3, 2)
┌────────┬────────┐
│ weight ┆ height │
│ --- ┆ --- │
│ f64 ┆ f64 │
╞════════╪════════╡
│ 63.69 ┆ 1.716 │
│ 79.75 ┆ 1.947 │
│ 58.96 ┆ 1.815 │
└────────┴────────┘
> (list (dtype-spec? 'f64) (dtype-spec? 'float)) '(#t #f)
> (select people (col "^.*ght$"))
shape: (3, 2)
┌────────┬────────┐
│ weight ┆ height │
│ --- ┆ --- │
│ f64 ┆ f64 │
╞════════╪════════╡
│ 57.9 ┆ 1.56 │
│ 72.5 ┆ 1.77 │
│ 53.6 ┆ 1.65 │
└────────┴────────┘
> (select people (col #rx"^he"))
shape: (3, 1)
┌────────┐
│ height │
│ --- │
│ f64 │
╞════════╡
│ 1.56 │
│ 1.77 │
│ 1.65 │
└────────┘
> (select people (col #px"^\\w+t$"))
shape: (3, 2)
┌────────┬────────┐
│ weight ┆ height │
│ --- ┆ --- │
│ f64 ┆ f64 │
╞════════╪════════╡
│ 57.9 ┆ 1.56 │
│ 72.5 ┆ 1.77 │
│ 53.6 ┆ 1.65 │
└────────┴────────┘
> (select people (~> (col "id") (* 10) (alias "id10")) (alias (lit 0) "zero"))
shape: (3, 2)
┌──────┬──────┐
│ id10 ┆ zero │
│ --- ┆ --- │
│ i32 ┆ i32 │
╞══════╪══════╡
│ 10 ┆ 0 │
│ 20 ┆ 0 │
│ 30 ┆ 0 │
└──────┴──────┘
> (select people (col #px"^(?!id)")) lazyframe-collect: failed to collect the query: invalid
regex in selector '^(?s).*(?:^(?!id)).*$'
Resolved plan until failure:
---> FAILED HERE RESOLVING 'select' <---
DF ["id", "weight", "height"]; PROJECT */3 COLUMNS: 'select'
> (select people (filter (lambda (name) (regexp-match? #px"^(?!id)" name)) (column-names people)))
shape: (3, 2)
┌────────┬────────┐
│ weight ┆ height │
│ --- ┆ --- │
│ f64 ┆ f64 │
╞════════╪════════╡
│ 57.9 ┆ 1.56 │
│ 72.5 ┆ 1.77 │
│ 53.6 ┆ 1.65 │
└────────┴────────┘
procedure
procedure
e : multi-column-expr? name : (or/c string? regexp?)
procedure
(multi-column-expr? v) → boolean?
v : any/c
> (select people (all))
shape: (3, 3)
┌─────┬────────┬────────┐
│ id ┆ weight ┆ height │
│ --- ┆ --- ┆ --- │
│ i32 ┆ f64 ┆ f64 │
╞═════╪════════╪════════╡
│ 1 ┆ 57.9 ┆ 1.56 │
│ 2 ┆ 72.5 ┆ 1.77 │
│ 3 ┆ 53.6 ┆ 1.65 │
└─────┴────────┴────────┘
> (select people (exclude (all) "id"))
shape: (3, 2)
┌────────┬────────┐
│ weight ┆ height │
│ --- ┆ --- │
│ f64 ┆ f64 │
╞════════╪════════╡
│ 57.9 ┆ 1.56 │
│ 72.5 ┆ 1.77 │
│ 53.6 ┆ 1.65 │
└────────┴────────┘
> (select people (exclude (all) #rx"^w" "id"))
shape: (3, 1)
┌────────┐
│ height │
│ --- │
│ f64 │
╞════════╡
│ 1.56 │
│ 1.77 │
│ 1.65 │
└────────┘
> (select people (exclude (all) "^h.*$"))
shape: (3, 2)
┌─────┬────────┐
│ id ┆ weight │
│ --- ┆ --- │
│ i32 ┆ f64 │
╞═════╪════════╡
│ 1 ┆ 57.9 │
│ 2 ┆ 72.5 │
│ 3 ┆ 53.6 │
└─────┴────────┘
> (select people (~> (col 'float64) (exclude "height") (* 2)))
shape: (3, 1)
┌────────┐
│ weight │
│ --- │
│ f64 │
╞════════╡
│ 115.8 │
│ 145.0 │
│ 107.2 │
└────────┘
> (select people (~> (all) (exclude "id") (exclude "weight")))
shape: (3, 1)
┌────────┐
│ height │
│ --- │
│ f64 │
╞════════╡
│ 1.56 │
│ 1.77 │
│ 1.65 │
└────────┘
> (~> people (with-columns (~> (col "height") (> 1.6) (alias "tall"))) (group-by "tall") (agg (~> (all) (exclude "id") mean)) (sort "tall"))
shape: (2, 3)
┌───────┬────────┬────────┐
│ tall ┆ weight ┆ height │
│ --- ┆ --- ┆ --- │
│ bool ┆ f64 ┆ f64 │
╞═══════╪════════╪════════╡
│ false ┆ 57.9 ┆ 1.56 │
│ true ┆ 63.05 ┆ 1.71 │
└───────┴────────┴────────┘
> (multi-column-expr? (col "id")) #f
> (~> (col 'float64) (* 2) multi-column-expr?) #t
> (exclude (col "id") "weight") exclude: contract violation
expected: multi-column-expr?
given: col("id")
in: the 1st argument of
(->
multi-column-expr?
(or/c string? regexp?)
(or/c string? regexp?)
...
Expr-ptr?)
contract from:
<pkgs>/polars/private/generic/selectors.rkt
blaming: top-level
(assuming the contract is correct)
at: <pkgs>/polars/private/generic/selectors.rkt:8:11
procedure
(expr->string e) → string?
e : Expr-ptr?
> (col "weight") col("weight")
> (alias (* (col "v") 10) "v10") [(col("v")) * (dyn int: 10)].alias("v10")
> (expr->string (> (col "v") 2)) "[(col(\"v\")) > (dyn int: 2)]"
> (~> (col "v") sum (over "k")) col("v").sum().over([col("k")])
> (exclude (all) "id") [cs.all() - cs.by_name('id', require_all=false)]
> (col 'float64) cs.by_dtype([Float64])
> (col #rx"^he") cs.matches("^(?s).*(?:^he).*$")
procedure
(meta-output-name e) → string?
e : (or/c Expr-ptr? string?)
procedure
(meta-root-names e) → (listof string?)
e : (or/c Expr-ptr? string?)
procedure
a : (or/c Expr-ptr? string?) b : (or/c Expr-ptr? string?)
A multi-column expression is read off the plan too, before any frame says which columns it will match, so a regexp or dtype col or (all) has no root names and no output name, as in Python.
> (define total (alias (sum (+ (col "a") (col "b"))) "total")) > total [(col("a")) + (col("b"))].sum().alias("total")
> (meta-output-name total) "total"
> (meta-root-names total) '("a" "b")
> (meta-eq? total (alias (sum (+ (col "a") (col "b"))) "total")) #t
> (meta-eq? total (col "a")) #f
> (equal? total (~> (+ (col "a") (col "b")) sum (alias "total"))) #f
> (meta-output-name (+ (col "a") (col "b"))) "a"
> (meta-output-name (lit 25)) "literal"
> (meta-root-names "a") '("a")
> (meta-root-names (col #rx"^he")) '()
> (meta-output-name (col #rx"^he")) expr-meta-output-name: cannot determine the output name of
cs.matches("^(?s).*(?:^he).*$")
> (~> (col 'float64) (* 2) meta-root-names) '()
> (meta-eq? (col 'float64) (col 'f64)) #t
> (meta-output-name (col 'float64)) expr-meta-output-name: cannot determine the output name of
cs.by_dtype([Float64])
> (meta-output-name "*") expr-meta-output-name: cannot determine the output name of
cs.all()
> (meta-output-name 5) meta-output-name: contract violation
expected: col-expr/c
given: 5
in: the 1st argument of
(-> col-expr/c string?)
contract from:
<pkgs>/polars/private/generic/meta.rkt
blaming: top-level
(assuming the contract is correct)
at: <pkgs>/polars/private/generic/meta.rkt:9:11
procedure
(> a b ...) → any/c
a : any/c b : any/c
procedure
(< a b ...) → any/c
a : any/c b : any/c
procedure
(>= a b ...) → any/c
a : any/c b : any/c
procedure
(<= a b ...) → any/c
a : any/c b : any/c
procedure
(= a b ...) → any/c
a : any/c b : any/c
procedure
(!= a b ...) → any/c
a : any/c b : any/c
syntax
(and expr ...)
syntax
(or expr ...)
procedure
(not x) → any/c
x : any/c
procedure
(xor a b) → any/c
a : any/c b : any/c
procedure
(+ v ...) → any/c
v : any/c
procedure
(- v ...) → any/c
v : any/c
procedure
(* v ...) → any/c
v : any/c
procedure
(/ v ...) → any/c
v : any/c
procedure
(filter d predicate) → dataframe?
d : dataframe? predicate : any/c
procedure
(sort d by [ #:descending descending #:nulls-last nulls-last #:maintain-order maintain-order]) → dataframe? d : dataframe? by : (or/c string? (non-empty-listof string?)) descending : (or/c boolean? (listof boolean?)) = #f nulls-last : (or/c boolean? (listof boolean?)) = #f maintain-order : boolean? = #f
(sort lf by [ #:descending descending #:nulls-last nulls-last #:maintain-order maintain-order]) → lazyframe? lf : lazyframe? by : (or/c string? (non-empty-listof string?)) descending : (or/c boolean? (listof boolean?)) = #f nulls-last : (or/c boolean? (listof boolean?)) = #f maintain-order : boolean? = #f
(sort s [ #:descending descending #:nulls-last nulls-last]) → series? s : series? descending : boolean? = #f nulls-last : boolean? = #f
(sort e [ #:descending descending #:nulls-last nulls-last]) → Expr-ptr? e : (or/c Expr-ptr? string?) descending : boolean? = #f nulls-last : boolean? = #f
(sort lst less-than? [ #:key extract-key #:cache-keys? cache-keys?]) → list? lst : list? less-than? : (any/c any/c . -> . any/c) extract-key : (or/c #f (any/c . -> . any/c)) = #f cache-keys? : boolean? = #f
Nulls come first, whatever the direction, unless nulls-last is true; NaN sorts above every other float. For a frame, descending and nulls-last are each one boolean for every key or a list of one boolean per key. Rows that tie on every key keep their input order when maintain-order is true, and are in no particular order otherwise.
A sorted expression reorders its own column only, so in with-columns it no longer lines up with the rest of its row. Use it in select or agg, and sort-by to reorder one column by others. A frame sorts by column names only: sorting by an expression (df.sort(pl.col("a") * -1)) has no spelling yet.
On a list, sort is racket/base’s, and passing it one of the Polars keywords is a contract violation.
> (define flights (dataframe (list (series '("UA" "AA" "UA" "AA" "B6") #:name "carrier") (series (list 12 polars-null 340 -3 polars-null) #:name "delay")))) > (sort flights "delay" #:descending #t)
shape: (5, 2)
┌─────────┬───────┐
│ carrier ┆ delay │
│ --- ┆ --- │
│ str ┆ i64 │
╞═════════╪═══════╡
│ AA ┆ null │
│ B6 ┆ null │
│ UA ┆ 340 │
│ UA ┆ 12 │
│ AA ┆ -3 │
└─────────┴───────┘
> (~> flights (sort "delay" #:descending #t #:nulls-last #t) (head 2))
shape: (2, 2)
┌─────────┬───────┐
│ carrier ┆ delay │
│ --- ┆ --- │
│ str ┆ i64 │
╞═════════╪═══════╡
│ UA ┆ 340 │
│ UA ┆ 12 │
└─────────┴───────┘
> (sort flights '("carrier" "delay") #:descending '(#f #t) #:nulls-last '(#f #t) #:maintain-order #t)
shape: (5, 2)
┌─────────┬───────┐
│ carrier ┆ delay │
│ --- ┆ --- │
│ str ┆ i64 │
╞═════════╪═══════╡
│ AA ┆ -3 │
│ AA ┆ null │
│ B6 ┆ null │
│ UA ┆ 340 │
│ UA ┆ 12 │
└─────────┴───────┘
> (~> flights lazy (sort "delay" #:nulls-last #t) collect)
shape: (5, 2)
┌─────────┬───────┐
│ carrier ┆ delay │
│ --- ┆ --- │
│ str ┆ i64 │
╞═════════╪═══════╡
│ AA ┆ -3 │
│ UA ┆ 12 │
│ UA ┆ 340 │
│ AA ┆ null │
│ B6 ┆ null │
└─────────┴───────┘
> (sort (ref flights #:columns "delay") #:descending #t #:nulls-last #t)
shape: (5,)
Series: 'delay' [i64]
[
340
12
-3
null
null
]
> (select flights (sort "delay" #:nulls-last #t))
shape: (5, 1)
┌───────┐
│ delay │
│ --- │
│ i64 │
╞═══════╡
│ -3 │
│ 12 │
│ 340 │
│ null │
│ null │
└───────┘
> (sort '(3 1 2) <) '(1 2 3)
> (sort '(3 1 2) < #:descending #t) sort: contract violation
expected: polars-only/c
given: #t
in: the descending argument of
sort/c
contract from:
<pkgs>/polars/private/generic/ordering.rkt
blaming: top-level
(assuming the contract is correct)
at: <pkgs>/polars/private/generic/ordering.rkt:17:24
procedure
(sort-by x #:by by [ #:descending descending #:nulls-last nulls-last #:maintain-order maintain-order]) → Expr-ptr? x : (or/c Expr-ptr? string?) by : (or/c Expr-ptr? string? (non-empty-listof (or/c Expr-ptr? string?))) descending : (or/c boolean? (listof boolean?)) = #f nulls-last : (or/c boolean? (listof boolean?)) = #f maintain-order : boolean? = #f
> (define scores (dataframe (list (series '("ann" "bob" "cy" "dee") #:name "name") (series (list 3 polars-null 1 3) #:name "score"))))
> (select scores (sort-by "name" #:by "score" #:descending #t #:nulls-last #t #:maintain-order #t))
shape: (4, 1)
┌──────┐
│ name │
│ --- │
│ str │
╞══════╡
│ ann │
│ dee │
│ cy │
│ bob │
└──────┘
> (select scores (sort-by "name" #:by '("score" "name") #:descending '(#t #f)))
shape: (4, 1)
┌──────┐
│ name │
│ --- │
│ str │
╞══════╡
│ bob │
│ ann │
│ dee │
│ cy │
└──────┘
procedure
(select d spec ...) → (or/c dataframe? lazyframe?)
d : (or/c dataframe? lazyframe?) spec : any/c
procedure
(with-columns d spec ...) → (or/c dataframe? lazyframe?)
d : (or/c dataframe? lazyframe?) spec : any/c
> (~> (dataframe (list (series '(1 2 3) #:name "a") (series '(4 5 6) #:name "b"))) (select "a" (alias (+ (col "a") (col "b")) "sum")))
shape: (3, 2)
┌─────┬─────┐
│ a ┆ sum │
│ --- ┆ --- │
│ i64 ┆ i64 │
╞═════╪═════╡
│ 1 ┆ 5 │
│ 2 ┆ 7 │
│ 3 ┆ 9 │
└─────┴─────┘
> (~> (dataframe (list (series '(1 2 3) #:name "a"))) (with-columns (alias (* (col "a") 2) "double")))
shape: (3, 2)
┌─────┬────────┐
│ a ┆ double │
│ --- ┆ --- │
│ i64 ┆ i64 │
╞═════╪════════╡
│ 1 ┆ 2 │
│ 2 ┆ 4 │
│ 3 ┆ 6 │
└─────┴────────┘
> (~> (dataframe (list (series '(1 2 3) #:name "v"))) (with-columns (cast "v" 'float64)))
shape: (3, 1)
┌─────┐
│ v │
│ --- │
│ f64 │
╞═════╡
│ 1.0 │
│ 2.0 │
│ 3.0 │
└─────┘
> (cast (series '("UA" "AA" "UA")) 'categorical)
shape: (3,)
Series: '' [cat]
[
"UA"
"AA"
"UA"
]
> (cast (series '(1 2 3)) 'f64) ->compat-dtype: unsupported cast target 'f64
> (define-enum carriers UA AA) > (cast (series '("UA" "B6")) carriers) series-cast: cannot convert to '(enum UA AA): conversion
from `str` to `enum` failed in column '' for 1 out of 2
values: ["B6"]
Ensure that all values in the input column are present in
the categories of the enum datatype.
procedure
(vstack top bottom) → dataframe?
top : dataframe? bottom : dataframe?
> (define top (dataframe (list (series '(1 2) #:name "v")))) > (vstack top (dataframe (list (series '(3) #:name "v"))))
shape: (3, 1)
┌─────┐
│ v │
│ --- │
│ i64 │
╞═════╡
│ 1 │
│ 2 │
│ 3 │
└─────┘
procedure
(head x n) → any/c
x : (or/c series? dataframe? lazyframe? Expr-ptr?) n : exact-nonnegative-integer?
procedure
(tail x n) → any/c
x : (or/c series? dataframe? lazyframe? Expr-ptr?) n : exact-nonnegative-integer?
procedure
(slice x offset length) → any/c
x : (or/c series? dataframe? lazyframe? Expr-ptr?) offset : exact-integer? length : exact-nonnegative-integer?
> (define nums (dataframe (list (series '(1 2 3 4 5) #:name "v")))) > (head nums 2)
shape: (2, 1)
┌─────┐
│ v │
│ --- │
│ i64 │
╞═════╡
│ 1 │
│ 2 │
└─────┘
> (tail nums 2)
shape: (2, 1)
┌─────┐
│ v │
│ --- │
│ i64 │
╞═════╡
│ 4 │
│ 5 │
└─────┘
procedure
(drop d names) → dataframe?
d : dataframe? names : (or/c string? (listof string?))
procedure
(join left right [ #:on on #:left-on left-on #:right-on right-on #:how how]) → (or/c dataframe? lazyframe?) left : (or/c dataframe? lazyframe?) right : (or/c dataframe? lazyframe?) on : (or/c (listof string?) #f) = #f left-on : (or/c (listof string?) #f) = #f right-on : (or/c (listof string?) #f) = #f how : (or/c 'inner 'left 'outer 'cross 'semi 'anti) = 'inner
> (define left (dataframe (list (series '("a" "b") #:name "k") (series '(1 2) #:name "v"))))
> (define right (dataframe (list (series '("a" "c") #:name "k") (series '(10 30) #:name "w")))) > (~> (join left right #:on '("k") #:how 'left) (sort "k"))
shape: (2, 3)
┌─────┬─────┬──────┐
│ k ┆ v ┆ w │
│ --- ┆ --- ┆ --- │
│ str ┆ i64 ┆ i64 │
╞═════╪═════╪══════╡
│ a ┆ 1 ┆ 10 │
│ b ┆ 2 ┆ null │
└─────┴─────┴──────┘
> (join left right #:on '("k") #:how 'inner)
shape: (1, 3)
┌─────┬─────┬─────┐
│ k ┆ v ┆ w │
│ --- ┆ --- ┆ --- │
│ str ┆ i64 ┆ i64 │
╞═════╪═════╪═════╡
│ a ┆ 1 ┆ 10 │
└─────┴─────┴─────┘
procedure
(read-csv path [ #:has-header has-header #:separator separator #:quote-char quote-char #:comment-prefix comment-prefix #:skip-rows skip-rows #:n-rows n-rows #:null-values null-values #:infer-schema-length infer-schema-length #:schema-overrides schema-overrides #:ignore-errors ignore-errors #:try-parse-dates try-parse-dates #:encoding encoding #:glob glob]) → dataframe? path : path-string? has-header : boolean? = #t separator : (or/c csv-char/c #f) = #f quote-char : (or/c csv-char/c #f) = #\" comment-prefix : (or/c non-empty-string? #f) = #f skip-rows : exact-nonnegative-integer? = 0 n-rows : (or/c exact-nonnegative-integer? #f) = #f null-values : (or/c string? (listof string?) #f) = #f
infer-schema-length : (or/c exact-nonnegative-integer? #f) = 100
schema-overrides : (and/c (listof (cons/c string? csv-dtype/c)) distinct-names?) = '() ignore-errors : boolean? = #f try-parse-dates : boolean? = #f encoding : (or/c 'utf8 'utf8-lossy) = 'utf8 glob : boolean? = #t
A csv-char/c is one ASCII character other than newline or return. separator defaults to #\,; quote-char must differ from it, and #:quote-char #f turns quoting off. Lines that start with comment-prefix are skipped, as are the first skip-rows lines of each file; n-rows caps the rows read. A field equal to one of null-values reads as null. Column types are inferred from the first infer-schema-length rows: #f reads every row, and 0 makes every column a string. schema-overrides fixes the named columns’ types. A csv-dtype/c is any spelling series’ #:dtype accepts except a duration, which Polars cannot parse from CSV, and an Enum; API gap: read the column as 'categorical or 'string and cast it. Each column appears at most once (distinct-names?), and naming a column the file lacks is an error, where Python ignores the override. With #:ignore-errors #t a field that does not parse reads as null. #:try-parse-dates #t reads ISO dates, times of day and datetimes as 'date, 'time and 'datetime columns. 'utf8-lossy replaces invalid UTF-8 with U+FFFD.
The result is (collect (scan-csv path ....)) with the same
keywords; as in Python, a single file is read eagerly rather than through
a plan, which is faster. There is one extra check, stricter than Python’s
read_csv, which returns the one column. When separator is #f and the
file reads as one column whose header splits on a tab, ; or
|, read-csv raises an error naming the separator to
pass —
A failure names the operation, the path and the cause: the operating system’s for a file that cannot be opened, Polars’ own for input it cannot parse.
The examples read "flights.tsv", 102 rows of the nycflights13 data with NA for a missing value. Without #:separator the separator check stops the read; with it, the first NA fails to parse, and #:null-values fixes that:
> (read-csv "flights.tsv") dataframe-read-csv: flights.tsv reads as one column whose
header splits on #\tab into 19 fields; pass #:separator
#\tab, or #:separator #\, to keep one column
> (read-csv "flights.tsv" #:separator #\tab) dataframe-read-csv: failed to read csv from flights.tsv:
could not parse `NA` as dtype `i64` at column 'arr_delay'
(column number 9)
The current offset in the file is 648 bytes.
You might want to try:
- increasing #:infer-schema-length (e.g.
#:infer-schema-length 10000, or #f for every row),
- specifying correct dtype with #:schema-overrides
- setting #:ignore-errors to #t,
- adding `NA` to #:null-values.
Original error: ```invalid primitive value found during CSV
parsing```
> (define flights (read-csv "flights.tsv" #:separator #\tab #:null-values "NA")) > (shape flights) '(102 19)
> (null-count (ref flights #:columns "dep_delay")) 1
> (~> (read-csv "flights.tsv" #:separator #\tab #:null-values '("NA" "")) (ref #:columns "arr_delay") null-count) 2
Failures. A missing file or an unwritable path reports the operating system’s reason; a file in the wrong format reports Polars’:
> (read-csv "/no/such/file.csv") dataframe-read-csv: failed to read csv from
/no/such/file.csv: cannot open file: No such file or
directory (os error 2)
> (define small (dataframe (list (series '(1 2) #:name "v")))) > (write-parquet small "/no/such/dir/out.parquet") dataframe-write-parquet: failed to write parquet to
/no/such/dir/out.parquet: cannot create file: No such file
or directory (os error 2)
Here not-parquet is a path in the temporary directory:
> (write-csv small not-parquet) > (read-parquet not-parquet) dataframe-read-parquet: failed to read parquet from
/var/tmp/polars-not-parquet.csv: parquet: File out of
specification: A Parquet file must contain a header and
footer with at least 12 bytes
> (read-ndjson not-parquet) dataframe-read-json-lines: failed to read json lines from
/var/tmp/polars-not-parquet.csv: InternalError(TapeError) at
character 0 ('v')
Types. #:ignore-errors turns what does not parse into nulls; #:infer-schema-length widens or narrows the rows types are inferred from; #:schema-overrides and #:try-parse-dates set them outright. An override for a column the file lacks is an error:
> (~> (read-csv "flights.tsv" #:separator #\tab #:ignore-errors #t) (ref #:columns "dep_delay") null-count) 1
> (~> (read-csv "flights.tsv" #:separator #\tab #:infer-schema-length #f) (ref #:columns "dep_delay") dtype) 'string
> (~> (read-csv "flights.tsv" #:separator #\tab #:infer-schema-length 0) (ref #:columns "year") dtype) 'string
> (~> (read-csv "flights.tsv" #:separator #\tab #:null-values "NA" #:try-parse-dates #t #:schema-overrides '(("dep_delay" . f64) ("flight" . int32))) (select "dep_delay" "flight" "time_hour") (tail 3))
shape: (3, 3)
┌───────────┬────────┬─────────────────────┐
│ dep_delay ┆ flight ┆ time_hour │
│ --- ┆ --- ┆ --- │
│ f64 ┆ i32 ┆ datetime[μs] │
╞═══════════╪════════╪═════════════════════╡
│ -7.0 ┆ 1733 ┆ 2013-01-01 07:00:00 │
│ -5.0 ┆ 4525 ┆ 2013-01-01 15:00:00 │
│ null ┆ 4308 ┆ 2013-01-01 16:00:00 │
└───────────┴────────┴─────────────────────┘
> (read-csv "flights.tsv" #:separator #\tab #:schema-overrides '(("dep_dealy" . f64))) dataframe-read-csv: failed to read csv from flights.tsv:
schema overrides name columns not in the file: "dep_dealy"
Layout. "notes.csv" has a comment line, ; between fields and ' around a field that holds one; "latin1.csv" is not UTF-8:
> (read-csv "notes.csv" #:separator #\; #:comment-prefix "#" #:quote-char #\')
shape: (2, 2)
┌─────────┬───────────────┐
│ carrier ┆ note │
│ --- ┆ --- │
│ str ┆ str │
╞═════════╪═══════════════╡
│ UA ┆ late; weather │
│ AA ┆ on time │
└─────────┴───────────────┘
> (read-csv "notes.csv" #:separator #\; #:comment-prefix "#") dataframe-read-csv: failed to read csv from notes.csv: found
more fields than defined in 'Schema'
Consider setting 'truncate_ragged_lines=true'.
> (read-csv "parts/part-1.csv" #:has-header #f #:skip-rows 1)
shape: (2, 3)
┌──────────┬──────────┬──────────┐
│ column_1 ┆ column_2 ┆ column_3 │
│ --- ┆ --- ┆ --- │
│ str ┆ str ┆ i64 │
╞══════════╪══════════╪══════════╡
│ EWR ┆ IAH ┆ 2 │
│ LGA ┆ IAH ┆ 4 │
└──────────┴──────────┴──────────┘
> (~> (read-csv "flights.tsv" #:separator #\tab #:null-values "NA" #:n-rows 2) (select "carrier" "flight" "dep_delay"))
shape: (2, 3)
┌─────────┬────────┬───────────┐
│ carrier ┆ flight ┆ dep_delay │
│ --- ┆ --- ┆ --- │
│ str ┆ i64 ┆ i64 │
╞═════════╪════════╪═══════════╡
│ UA ┆ 1545 ┆ 2 │
│ UA ┆ 1714 ┆ 4 │
└─────────┴────────┴───────────┘
> (read-csv "latin1.csv") dataframe-read-csv: failed to read csv from latin1.csv:
invalid utf-8 sequence
> (read-csv "latin1.csv" #:encoding 'utf8-lossy)
shape: (2, 2)
┌─────────┬──────────┐
│ carrier ┆ name │
│ --- ┆ --- │
│ str ┆ str │
╞═════════╪══════════╡
│ B6 ┆ JetBlue │
│ ZZ ┆ Caf� Air │
└─────────┴──────────┘
Files. A pattern reads every match, in filename order; one that matches nothing, a literal path with #:glob #f, and a directory are errors:
> (read-csv "parts/*.csv")
shape: (6, 3)
┌────────┬──────┬───────────┐
│ origin ┆ dest ┆ dep_delay │
│ --- ┆ --- ┆ --- │
│ str ┆ str ┆ i64 │
╞════════╪══════╪═══════════╡
│ EWR ┆ IAH ┆ 2 │
│ LGA ┆ IAH ┆ 4 │
│ JFK ┆ MIA ┆ 2 │
│ JFK ┆ BQN ┆ -1 │
│ LGA ┆ ATL ┆ -6 │
│ EWR ┆ ORD ┆ -4 │
└────────┴──────┴───────────┘
> (read-csv "parts/*.tsv") dataframe-read-csv: failed to read csv from parts/*.tsv: no
files match the pattern
> (read-csv "parts/part-?.csv" #:glob #f) dataframe-read-csv: failed to read csv from
parts/part-?.csv: cannot open file: No such file or
directory (os error 2)
> (read-csv "parts") dataframe-read-csv: failed to read csv from parts: cannot
open file: it is a directory; pass a glob pattern such as
dir/*.csv
procedure
(scan-csv path [ #:has-header has-header #:separator separator #:quote-char quote-char #:comment-prefix comment-prefix #:skip-rows skip-rows #:n-rows n-rows #:null-values null-values #:infer-schema-length infer-schema-length #:schema-overrides schema-overrides #:ignore-errors ignore-errors #:try-parse-dates try-parse-dates #:encoding encoding #:glob glob]) → lazyframe? path : path-string? has-header : boolean? = #t separator : (or/c csv-char/c #f) = #f quote-char : (or/c csv-char/c #f) = #\" comment-prefix : (or/c non-empty-string? #f) = #f skip-rows : exact-nonnegative-integer? = 0 n-rows : (or/c exact-nonnegative-integer? #f) = #f null-values : (or/c string? (listof string?) #f) = #f
infer-schema-length : (or/c exact-nonnegative-integer? #f) = 100
schema-overrides : (and/c (listof (cons/c string? csv-dtype/c)) distinct-names?) = '() ignore-errors : boolean? = #f try-parse-dates : boolean? = #f encoding : (or/c 'utf8 'utf8-lossy) = 'utf8 glob : boolean? = #t
> (~> (scan-csv "flights.tsv" #:separator #\tab #:null-values "NA") (filter (> (col "dep_delay") 30)) (select "carrier" "dep_delay") collect)
shape: (2, 2)
┌─────────┬───────────┐
│ carrier ┆ dep_delay │
│ --- ┆ --- │
│ str ┆ i64 │
╞═════════╪═══════════╡
│ UA ┆ 47 │
│ MQ ┆ 39 │
└─────────┴───────────┘
> (~> (scan-csv "parts/*.csv") (group-by "origin") (agg (sum "dep_delay")) collect)
shape: (3, 2)
┌────────┬───────────┐
│ origin ┆ dep_delay │
│ --- ┆ --- │
│ str ┆ i64 │
╞════════╪═══════════╡
│ JFK ┆ 1 │
│ LGA ┆ -2 │
│ EWR ┆ -2 │
└────────┴───────────┘
> (shape (collect (scan-csv "flights.tsv"))) '(102 1)
> (define plan (scan-csv "/no/such/file.csv")) > (collect plan) lazyframe-collect: failed to collect the query: No such file
or directory (os error 2): /no/such/file.csv
> (define no-match (scan-csv "parts/*.tsv")) > (collect no-match) lazyframe-collect: failed to collect the query: no files
match the pattern
> (scan-csv "flights.tsv" #:separator #\tab #:schema-overrides '(("dep_dealy" . f64))) lazyframe-scan-csv: failed to scan flights.tsv: schema
overrides name columns not in the file: "dep_dealy"
procedure
(read-parquet path) → dataframe?
path : path-string?
procedure
(scan-parquet path [#:n-rows n-rows]) → lazyframe?
path : path-string? n-rows : (or/c exact-nonnegative-integer? #f) = #f
Categorical, Enum and Decimal columns keep their dtypes (Categorical, Enum and Decimal). A file Polars cannot read raises exn:fail with Polars’ reason, even where Polars itself panics.
> (read-parquet (build-path parquet-dir "*.parquet"))
shape: (6, 3)
┌────────┬──────┬───────────┐
│ origin ┆ dest ┆ dep_delay │
│ --- ┆ --- ┆ --- │
│ str ┆ str ┆ i64 │
╞════════╪══════╪═══════════╡
│ EWR ┆ IAH ┆ 2 │
│ LGA ┆ IAH ┆ 4 │
│ JFK ┆ MIA ┆ 2 │
│ JFK ┆ BQN ┆ -1 │
│ LGA ┆ ATL ┆ -6 │
│ EWR ┆ ORD ┆ -4 │
└────────┴──────┴───────────┘
> (collect (scan-parquet (build-path parquet-dir "part-*.parquet") #:n-rows 3))
shape: (3, 3)
┌────────┬──────┬───────────┐
│ origin ┆ dest ┆ dep_delay │
│ --- ┆ --- ┆ --- │
│ str ┆ str ┆ i64 │
╞════════╪══════╪═══════════╡
│ EWR ┆ IAH ┆ 2 │
│ LGA ┆ IAH ┆ 4 │
│ JFK ┆ MIA ┆ 2 │
└────────┴──────┴───────────┘
> (read-parquet "produce.parquet")
shape: (4, 3)
┌───────┬───────┬───────────────┐
│ item ┆ grade ┆ price │
│ --- ┆ --- ┆ --- │
│ cat ┆ enum ┆ decimal[10,2] │
╞═══════╪═══════╪═══════════════╡
│ apple ┆ high ┆ 1.25 │
│ pear ┆ low ┆ 0.80 │
│ apple ┆ null ┆ null │
│ fig ┆ mid ┆ 12.00 │
└───────┴───────┴───────────────┘
procedure
(read-ndjson path) → dataframe?
path : path-string?
procedure
d : dataframe? path : path-string?
procedure
(write-parquet d path) → void?
d : dataframe? path : path-string?
procedure
(write-ndjson d path) → void?
d : dataframe? path : path-string?
procedure
(lazy d) → lazyframe?
d : dataframe?
procedure
(collect lf) → dataframe?
lf : lazyframe?
> (define four (dataframe (list (series '(1 2 3 4) #:name "v")))) > (~> four lazy (filter (> (col "v") 2)) collect)
shape: (2, 1)
┌─────┐
│ v │
│ --- │
│ i64 │
╞═════╡
│ 3 │
│ 4 │
└─────┘
> (~> four lazy (filter (> (col "nope") 2)) collect) lazyframe-collect: failed to collect the query: not found:
unable to find column "nope"; valid columns: ["v"]
Resolved plan until failure:
---> FAILED HERE RESOLVING THIS_NODE <---
DF ["v"]; PROJECT */1 COLUMNS
procedure
d : dataframe? key : (or/c string? any/c)
procedure
(agg g agg-expr ...) → dataframe?
g : grouped? agg-expr : any/c
procedure
v : any/c
> (~> (dataframe (list (series '("x" "y" "x") #:name "k") (series '(1 2 3) #:name "v"))) (group-by "k") (agg (alias (sum "v") "total")))
shape: (2, 2)
┌─────┬───────┐
│ k ┆ total │
│ --- ┆ --- │
│ str ┆ i64 │
╞═════╪═══════╡
│ y ┆ 2 │
│ x ┆ 4 │
└─────┴───────┘
> (define kv (dataframe (list (series '("x" "y" "x") #:name "k") (series '(1 2 3) #:name "v")))) > (~> kv (group-by "k") (agg (alias (sum "v") "total")))
shape: (2, 2)
┌─────┬───────┐
│ k ┆ total │
│ --- ┆ --- │
│ str ┆ i64 │
╞═════╪═══════╡
│ x ┆ 4 │
│ y ┆ 2 │
└─────┴───────┘
> (~> kv (with-columns (~> (col "v") sum (over "k") (alias "total"))))
shape: (3, 3)
┌─────┬─────┬───────┐
│ k ┆ v ┆ total │
│ --- ┆ --- ┆ --- │
│ str ┆ i64 ┆ i64 │
╞═════╪═════╪═══════╡
│ x ┆ 1 ┆ 4 │
│ y ┆ 2 ┆ 2 │
│ x ┆ 3 ┆ 4 │
└─────┴─────┴───────┘
> (define khv (dataframe (list (series '("a" "a" "a" "b") #:name "k") (series '("x" "y" "x" "x") #:name "h") (series '(1 2 3 4) #:name "v"))))
> (~> khv (with-columns (~> (col "v") sum (over "k" "h") (alias "total")) (~> (col "v") mean (over (col "k")) (alias "mean")) (~> (col "v") (rank #:descending #t) (over "k") (alias "rank")) (~> (col "v") max (over (> (col "v") 1)) (alias "band_max"))))
shape: (4, 7)
┌─────┬─────┬─────┬───────┬──────┬──────┬──────────┐
│ k ┆ h ┆ v ┆ total ┆ mean ┆ rank ┆ band_max │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ str ┆ i64 ┆ i64 ┆ f64 ┆ f64 ┆ i64 │
╞═════╪═════╪═════╪═══════╪══════╪══════╪══════════╡
│ a ┆ x ┆ 1 ┆ 4 ┆ 2.0 ┆ 3.0 ┆ 1 │
│ a ┆ y ┆ 2 ┆ 2 ┆ 2.0 ┆ 2.0 ┆ 4 │
│ a ┆ x ┆ 3 ┆ 4 ┆ 2.0 ┆ 1.0 ┆ 4 │
│ b ┆ x ┆ 4 ┆ 4 ┆ 4.0 ┆ 1.0 ┆ 4 │
└─────┴─────┴─────┴───────┴──────┴──────┴──────────┘
> (over (col "v")) over: contract violation
received: 1 argument
expected: at least 2 non-keyword arguments
in: (->*
(col-expr/c col-expr/c)
#:rest
(listof col-expr/c)
Expr-ptr?)
contract from:
<pkgs>/polars/private/generic/window.rkt
blaming: top-level
(assuming the contract is correct)
at: <pkgs>/polars/private/generic/window.rkt:8:11
procedure
(rank x [ #:method method #:descending descending #:seed seed]) → Expr-ptr? x : (or/c Expr-ptr? string?) method : (or/c 'average 'min 'max 'dense 'ordinal) = 'average descending : boolean? = #f seed : (or/c exact-nonnegative-integer? #f) = #f
procedure
x : (or/c Expr-ptr? string?) indices : (or/c Expr-ptr? series? (listof exact-integer?))
> (define scores (dataframe (list (series '("a" "b" "c" "d") #:name "name") (series '(30 10 30 20) #:name "score"))))
> (~> scores (with-columns (~> (col "score") (rank #:method 'dense #:descending #t) (alias "dense")) (~> (col "score") (rank #:method 'ordinal) (alias "ordinal"))))
shape: (4, 4)
┌──────┬───────┬───────┬─────────┐
│ name ┆ score ┆ dense ┆ ordinal │
│ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ i64 ┆ u32 ┆ u32 │
╞══════╪═══════╪═══════╪═════════╡
│ a ┆ 30 ┆ 1 ┆ 3 │
│ b ┆ 10 ┆ 3 ┆ 1 │
│ c ┆ 30 ┆ 1 ┆ 4 │
│ d ┆ 20 ┆ 2 ┆ 2 │
└──────┴───────┴───────┴─────────┘
> (~> scores (select (gather "name" '(3 0))))
shape: (2, 1)
┌──────┐
│ name │
│ --- │
│ str │
╞══════╡
│ d │
│ a │
└──────┘
> (rank "score" #:method 'first) expr-rank: method must be one of 'average 'min 'max 'dense
'ordinal, got 'first
procedure
(sum v ...) → any/c
v : any/c
procedure
(mean v ...) → any/c
v : any/c
procedure
(min v ...) → any/c
v : any/c
procedure
(max v ...) → any/c
v : any/c
> (define s (series '(1 2 3 4) #:name "v")) > (list (sum s) (mean s) (min s) (max s)) '(10 2.5 1 4)
> (max 1 2 3) 3
procedure
(count x) → any/c
x : any/c
procedure
(n-unique x) → any/c
x : any/c
procedure
(median x) → any/c
x : any/c
procedure
(std x [#:ddof ddof]) → any/c
x : any/c ddof : exact-nonnegative-integer? = 1
procedure
(var x [#:ddof ddof]) → any/c
x : any/c ddof : exact-nonnegative-integer? = 1
procedure
(alias e name) → any/c
e : any/c name : string?
procedure
(first x) → any/c
x : (or/c string? pair? any/c)
procedure
(last x) → any/c
x : (or/c string? pair? any/c)
Name clash. racket/list also exports first and last (along with count and group-by, which polars exports too). Requiring both modules explicitly — (require racket/list polars) — is an error (identifier already required). A plain #lang racket/base program is unaffected, because racket/base does not export these names. See Shadowed bindings for how to take control.
procedure
x : (or/c Expr-ptr? string?) exponent : (or/c Expr-ptr? real?)
procedure
(round x [#:decimals decimals]) → any/c
x : (or/c Expr-ptr? string? number?) decimals : exact-nonnegative-integer? = 0
> (~> (dataframe (list (series '(-2.5 -1.5 0.5 1.5 2.5) #:name "x"))) (select (round "x")))
shape: (5, 1)
┌──────┐
│ x │
│ --- │
│ f64 │
╞══════╡
│ -2.0 │
│ -2.0 │
│ 0.0 │
│ 2.0 │
│ 2.0 │
└──────┘
> (round 2.5) 2.0
> (~> (dataframe (list (series '(-2 0 3) #:name "i") (series '(-2.5 0.0 3.5) #:name "f"))) (select (sign "i") (sign "f")))
shape: (3, 2)
┌─────┬──────┐
│ i ┆ f │
│ --- ┆ --- │
│ i64 ┆ f64 │
╞═════╪══════╡
│ -1 ┆ -1.0 │
│ 0 ┆ 0.0 │
│ 1 ┆ 1.0 │
└─────┴──────┘
procedure
(is-between x lower upper [#:closed closed]) → Expr-ptr?
x : (or/c Expr-ptr? string?) lower : any/c upper : any/c closed : (or/c 'both 'left 'right 'none) = 'both
procedure
x : (or/c Expr-ptr? string?) rhs : (or/c list? series? Expr-ptr?)
procedure
x : (or/c Expr-ptr? string?)
procedure
x : (or/c Expr-ptr? string?)
procedure
x : (or/c Expr-ptr? string?)
procedure
x : (or/c Expr-ptr? string?)
procedure
x : (or/c Expr-ptr? string?)
procedure
x : (or/c Expr-ptr? string?)
procedure
(str-extract x pattern [ #:group-index group-index]) → Expr-ptr? x : (or/c Expr-ptr? string?) pattern : string? group-index : exact-nonnegative-integer? = 1
procedure
(str->date x [ #:format format #:strict strict #:exact exact #:cache cache]) → Expr-ptr? x : (or/c Expr-ptr? string?) format : (or/c string? #f) = #f strict : boolean? = #t exact : boolean? = #t cache : boolean? = #t
procedure
(str->datetime x [ #:format format #:unit unit #:strict strict #:exact exact #:cache cache]) → Expr-ptr? x : (or/c Expr-ptr? string?) format : (or/c string? #f) = #f
unit : (or/c 'milliseconds 'microseconds 'nanoseconds) = 'microseconds strict : boolean? = #t exact : boolean? = #t cache : boolean? = #t
2.1.1 Shadowed bindings
(require polars) re-exports a handful of generic operations whose names also live in racket/base (min, max, sort, filter, >, <, >=, <=, =) and in racket/list (first, last, count, group-by). Under #lang racket/base this is seamless — these names are either not bound (so polars simply provides them) or bound only by the module language (which an explicit require silently shadows), and the polars versions intentionally fall back to the numeric/list behaviour for non-frame arguments.
A conflict arises only when another module providing the same name is also required explicitly — most commonly racket/list. Resolve it with the usual require sub-forms:
; keep polars' first/last/count/group-by, drop racket/list's: (require (except-in racket/list first last count group-by) polars) ; keep racket/list's, reach polars' under a prefix: (require racket/list (prefix-in pl: polars)) ; then (pl:first (col "v")) for the Expr, (first '(1 2 3)) for the list ; keep polars', reach racket/list's under a prefix: (require polars (prefix-in list: racket/list))
2.2 Series
A series wraps a typed column and prints in the REPL the way Polars prints it; series? is its predicate. (The underlying foreign pointer is an implementation detail and not part of the public series API.)
> (series '(1 2 3) #:name "ints")
shape: (3,)
Series: 'ints' [i64]
[
1
2
3
]
> (series '(1.5 2.5) #:name "floats" #:dtype 'f32)
shape: (2,)
Series: 'floats' [f32]
[
1.5
2.5
]
> (series (list 1 polars-null 3) #:name "with-null")
shape: (3,)
Series: 'with-null' [i64]
[
1
null
3
]
> (series '(IAH ATL IAH) #:name "dest")
shape: (3,)
Series: 'dest' [cat]
[
"IAH"
"ATL"
"IAH"
]
> (define-enum severity debug info error) > (series '("debug" "error") #:dtype severity)
shape: (2,)
Series: '' [enum]
[
"debug"
"error"
]
> (series '(debug fatal) #:dtype severity) series: cannot convert to '(enum debug info error):
conversion from `str` to `enum` failed in column '' for 1
out of 2 values: ["fatal"]
Ensure that all values in the input column are present in
the categories of the enum datatype.
procedure
(series->string s) → string?
s : series?
value
polars-null : any/c
procedure
(polars-null? v) → boolean?
v : any/c
procedure
s : has-dtype?
procedure
x : sized?
procedure
s : has-null-count?
procedure
(series-name s) → string?
s : series?
> (series-name (series '(1 2) #:name "ints")) "ints"
procedure
s : series? new-name : string?
procedure
s : series? new-name : string?
procedure
s : series?
procedure
(series-clone s) → series?
s : series?
2.2.1 Converting to Racket values
These copy a column out of Polars in one foreign call rather than one per element: Racket allocates a buffer of the column’s native type, Rust copies the values into it, and Racket builds its values from the buffer and frees it. Nothing crosses the boundary to be freed later. At its peak a conversion holds that buffer (one native value per row, plus a byte per row when the column has nulls) beside the result it builds; in-series instead converts 4096 rows at a time. Each element comes out as ref returns it:
dtype |
| element |
integer dtypes |
| |
'float32, 'float64 |
| |
'boolean |
| |
'string |
| |
'categorical, '(enum cat ...) |
| |
'(decimal precision scale) |
| an exact rational, as exact? |
'date |
| a gregor date |
'(datetime unit tz) |
| a gregor datetime, floored to the second |
'(duration unit) |
| a gregor period in that unit |
'time |
| a gregor time |
'null |
| the null value |
A null entry becomes the #:null value. A series of any other dtype raises exn:fail:contract naming the dtype, even when every entry is null. The Interoperability chapter of the guide walks through all of them.
> (series->list (series (list 1.5 polars-null))) '(1.5 polars-null)
> (series->list (series (list "a" polars-null ""))) '("a" polars-null "")
> (series->list (series (list #t #f polars-null))) '(#t #f polars-null)
> (series->list (series (list (datetime 2024 1 2 3 4 5) polars-null))) '(#<datetime 2024-01-02T03:04:05> polars-null)
> (series->list (cast (series '(19724) #:dtype 'i32) 'date)) '(#<date 2024-01-02>)
> (series->list (cast (series '(11045000000000) #:dtype 'i64) 'time)) '(#<time 03:04:05>)
> (series->list (cast (series '(1500) #:dtype 'i64) '(duration milliseconds))) '(#<period of 1500 milliseconds>)
> (series->list (cast (series (list "UA" polars-null "UA")) 'categorical)) '(UA polars-null UA)
> (series->list (ref (read-parquet "produce.parquet") "price")) '(5/4 4/5 polars-null 12)
> (series->list (cast (series '("a") #:name "b") 'binary)) series->list: unsupported dtype
series: "b"
dtype: 'binary
procedure
(series->list s [#:null null-value]) → list?
s : series? null-value : any/c = polars-null
procedure
(series->vector s [#:null null-value]) → vector?
s : series? null-value : any/c = polars-null
> (define s (series (list 3 polars-null 1) #:name "x")) > (series->list s) '(3 polars-null 1)
> (series->list s #:null 'missing) '(3 missing 1)
> (series->vector s) '#(3 polars-null 1)
> (series->vector s #:null 0) '#(3 0 1)
procedure
(series->f64vector s [#:null null-value]) → f64vector?
s : series? null-value : (or/c real? 'error) = +nan.0
The result is ordinary garbage-collected memory, which Racket CS may move: pass it to a foreign call that is not #:blocking?, and do not let foreign code keep the pointer past the call.
> (define xs (series (list 1 polars-null 3) #:name "x")) > (f64vector->list (series->f64vector xs)) '(1.0 +nan.0 3.0)
> (f64vector->list (series->f64vector xs #:null 0)) '(1.0 0.0 3.0)
> (f64vector->list (series->f64vector (series (list #t #f)))) '(1.0 0.0)
> (f64vector->list (series->f64vector (series '(1 2)) #:null 'error)) '(1.0 2.0)
> (series->f64vector xs #:null 'error) series->f64vector: null value
series: "x"
row: 1
> (series->f64vector (series '("a") #:name "s")) series->f64vector: not a numeric series
series: "s"
dtype: 'string
procedure
s : series? null-value : any/c = polars-null
> (for/list ([x (in-series xs)]) x) '(1 polars-null 3)
> (for/sum ([x (in-series xs #:null 0)]) x) 4
> (for/list ([x (series '("a" "b"))]) (string-upcase x)) '("A" "B")
> (for/first ([x (in-series (series (build-list 100000 values)))] #:when (> x 41)) x) 42
2.2.2 Categorical, Enum and Decimal
A 'categorical column stores each distinct string once and a code per row, as Polars’ Categorical; an Enum column (dtype '(enum cat ...), defined with define-enum) does the same over categories declared up front, in order, as pl.Enum. Both read back as symbols, which Racket interns: a symbol is already the dictionary encoding. The codes stay inside Polars. Every categorical column in the process shares them, and they restart once the last one is dropped, so each conversion fetches the strings afresh.
Build one with series (a list of symbols infers 'categorical), cast, or read-csv’s #:schema-overrides ('categorical only).
A categorical sorts and compares by its strings; an Enum by the declared order of its categories.
A value outside an Enum’s categories raises, whether it is built, cast or compared; a categorical takes any string.
Two categorical columns share one encoding, so they join, stack and compare without re-encoding; two Enums with the same categories are the same dtype.
describe gives a categorical or Enum column count and null_count only, as Python does.
syntax
(define-enum id category ...+)
category = id | string
> (define-enum log-levels debug info warning error) > log-levels '(enum debug info warning error)
> (define levels (series '(debug info debug error) #:name "level" #:dtype log-levels)) > (equal? (dtype levels) log-levels) #t
> (series->list (cast (series '("warning" "info")) log-levels)) '(warning info)
> (define-enum sizes small "Very High") > sizes '(enum small |Very High|)
> (define-enum twice debug info debug) eval:177:0: define-enum: duplicate enum category
at: debug
in: (define-enum twice debug info debug)
> (series '(info fatal) #:dtype log-levels) series: cannot convert to '(enum debug info warning error):
conversion from `str` to `enum` failed in column '' for 1
out of 2 values: ["fatal"]
Ensure that all values in the input column are present in
the categories of the enum datatype.
A '(decimal precision scale) column holds exact decimals, which ref and the conversions read as exact rationals. Decimal columns come from Parquet; cast reads them into other dtypes. API gap: no #:dtype or cast to a Decimal.
> (define logs (dataframe (list (series '(debug info debug error) #:name "level" #:dtype log-levels) (series '(api db api db) #:name "source")))) > (for/list ([name (column-names logs)]) (dtype (ref logs name))) '((enum debug info warning error) categorical)
> (filter logs (> (col "level") 'info))
shape: (1, 2)
┌───────┬────────┐
│ level ┆ source │
│ --- ┆ --- │
│ enum ┆ cat │
╞═══════╪════════╡
│ error ┆ db │
└───────┴────────┘
> (sort logs "level")
shape: (4, 2)
┌───────┬────────┐
│ level ┆ source │
│ --- ┆ --- │
│ enum ┆ cat │
╞═══════╪════════╡
│ debug ┆ api │
│ debug ┆ api │
│ info ┆ db │
│ error ┆ db │
└───────┴────────┘
> (sort logs "source")
shape: (4, 2)
┌───────┬────────┐
│ level ┆ source │
│ --- ┆ --- │
│ enum ┆ cat │
╞═══════╪════════╡
│ debug ┆ api │
│ debug ┆ api │
│ info ┆ db │
│ error ┆ db │
└───────┴────────┘
> (ref (ref logs "source") 1) 'db
> (select logs (col 'categorical))
shape: (4, 1)
┌────────┐
│ source │
│ --- │
│ cat │
╞════════╡
│ api │
│ db │
│ api │
│ db │
└────────┘
> (select logs (> (col "level") 'fatal)) lazyframe-collect: failed to collect the query: conversion
from `str` to `enum` failed for value "fatal"
> (define produce (read-parquet "produce.parquet")) > (dtype (ref produce "price")) '(decimal 10 2)
> (for/sum ([price (ref produce "price")] #:unless (polars-null? price)) price) 281/20
2.2.3 dtype promotion
Reductions follow a simple, predictable rule. The widening order, narrow to wide, is
'int8 < 'int16 < 'int32 < 'int64
'uint8 < 'uint16 < 'uint32 < 'uint64
any integer < 'float32 < 'float64
sum, min and max preserve the input dtype. mean promotes to 'float64. Use series-cast to change a series’ dtype explicitly.
2.2.4 Low-level Series API
The generic layer is built on monomorphic, dtype-suffixed bindings that operate directly on the foreign series. They remain exported. A series wrapper is accepted anywhere one of them expects a series (the wrapper marshals transparently, and satisfies Series-ptr?), but what they return is the raw foreign pointer, not a wrapper — so the results do not print in Polars’ format and do not answer to series?. Prefer series and the generic operations above; reach for these when you need a specific dtype or a specific typed result.
procedure
(Series-ptr? v) → boolean?
v : any/c
procedure
(series-new-i8 name values) → Series-ptr?
name : string? values : list?
procedure
(series-new-i16 name values) → Series-ptr?
name : string? values : list?
procedure
(series-new-i32 name values) → Series-ptr?
name : string? values : list?
procedure
(series-new-i64 name values) → Series-ptr?
name : string? values : list?
procedure
(series-new-u8 name values) → Series-ptr?
name : string? values : list?
procedure
(series-new-u16 name values) → Series-ptr?
name : string? values : list?
procedure
(series-new-u32 name values) → Series-ptr?
name : string? values : list?
procedure
(series-new-u64 name values) → Series-ptr?
name : string? values : list?
procedure
(series-new-f32 name values) → Series-ptr?
name : string? values : list?
procedure
(series-new-f64 name values) → Series-ptr?
name : string? values : list?
procedure
(series-new-bool name values) → Series-ptr?
name : string? values : list?
procedure
(series-new-str name values) → Series-ptr?
name : string? values : list?
procedure
(series-sum-i32 s) → (or/c exact-integer? #f)
s : Series-ptr?
procedure
(series-min-i32 s) → (or/c exact-integer? #f)
s : Series-ptr?
procedure
(series-max-i32 s) → (or/c exact-integer? #f)
s : Series-ptr?
procedure
(series-mean-i32 s) → (or/c flonum? #f)
s : Series-ptr?
procedure
(series-sum-f64 s) → (or/c flonum? #f)
s : Series-ptr?
procedure
(series-min-f64 s) → (or/c flonum? #f)
s : Series-ptr?
procedure
(series-max-f64 s) → (or/c flonum? #f)
s : Series-ptr?
procedure
(series-mean-f64 s) → (or/c flonum? #f)
s : Series-ptr?
procedure
(series-cast s dtype) → Series-ptr?
s : Series-ptr? dtype : (or/c symbol? pair?)
procedure
(series-sort s [ #:descending descending #:nulls-last nulls-last]) → Series-ptr? s : Series-ptr? descending : boolean? = #f nulls-last : boolean? = #f
2.3 DataFrames
A dataframe is a collection of equal-length named series. Like a series it is a wrapper value (dataframe?) carrying the column data; it prints as a Polars table, so display (or ~a, or the REPL) renders it with no separate display call.
procedure
(dataframe? v) → boolean?
v : any/c
procedure
(dataframe columns) → dataframe?
columns : (listof series?)
> (dataframe (list (series '("a" "b") #:name "k") (series '(1 2) #:name "v")))
shape: (2, 2)
┌─────┬─────┐
│ k ┆ v │
│ --- ┆ --- │
│ str ┆ i64 │
╞═════╪═════╡
│ a ┆ 1 │
│ b ┆ 2 │
└─────┴─────┘
procedure
(shape x) → (listof exact-nonnegative-integer?)
x : has-shape?
procedure
(shape/values x) →
exact-nonnegative-integer? ... x : has-shape?
procedure
d : dataframe?
procedure
d : dataframe?
procedure
(column-names d) → (listof string?)
d : dataframe?
procedure
(column-name d i) → string?
d : dataframe? i : exact-nonnegative-integer?
procedure
(ref x [key #:columns columns #:rows rows]) → any/c
x : has-ref? key : (or/c exact-nonnegative-integer? string?) = absent
columns :
(or/c exact-nonnegative-integer? string? (listof (or/c exact-nonnegative-integer? string?))) = absent rows : any/c = absent
procedure
(describe x) → dataframe?
x : (or/c series? dataframe?)
Numeric, boolean, null and nested columns summarise as 'float64, every other column as strings; temporal values are written as Python prints them. Quantiles use nearest interpolation. Every statistic of every column comes from one query, so Polars computes the columns in parallel.
API gaps: a time-zone-aware datetime is written as its UTC clock time with no offset, where Python writes the local time and the offset; a binary column gets no "min" or "max".
Numeric, string and boolean columns, with nulls:
> (define flights (dataframe (list (series (list "UA" "AA" "UA" polars-null) #:name "carrier") (series (list 1400 733 polars-null 1089) #:name "distance") (series (list #t #f #t #t) #:name "on_time")))) > (describe flights)
shape: (9, 4)
┌────────────┬─────────┬────────────┬─────────┐
│ statistic ┆ carrier ┆ distance ┆ on_time │
│ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ str ┆ f64 ┆ f64 │
╞════════════╪═════════╪════════════╪═════════╡
│ count ┆ 3 ┆ 3.0 ┆ 4.0 │
│ null_count ┆ 1 ┆ 1.0 ┆ 0.0 │
│ mean ┆ null ┆ 1074.0 ┆ 0.75 │
│ std ┆ null ┆ 333.752903 ┆ null │
│ min ┆ AA ┆ 733.0 ┆ 0.0 │
│ 25% ┆ null ┆ 1089.0 ┆ null │
│ 50% ┆ null ┆ 1089.0 ┆ null │
│ 75% ┆ null ┆ 1400.0 ┆ null │
│ max ┆ UA ┆ 1400.0 ┆ 1.0 │
└────────────┴─────────┴────────────┴─────────┘
Datetime, duration and date columns get a mean and quartiles; a date column’s mean is a datetime:
> (define times (~> (dataframe (list (series (list (datetime 2013 1 1 5) (datetime 2013 1 1 6) (datetime 2013 1 2 7) (datetime 2013 1 3 8)) #:name "scheduled") (series (list (datetime 2013 1 1 5 12) (datetime 2013 1 1 5 57) polars-null (datetime 2013 1 3 9 30)) #:name "departed"))) (with-columns (alias (- (col "departed") (col "scheduled")) "delay") (alias (cast "scheduled" 'date) "day")))) > (describe times)
shape: (9, 5)
┌────────────┬─────────────────────┬─────────────────────┬──────────────────┬─────────────────────┐
│ statistic ┆ scheduled ┆ departed ┆ delay ┆ day │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ str ┆ str ┆ str ┆ str │
╞════════════╪═════════════════════╪═════════════════════╪══════════════════╪═════════════════════╡
│ count ┆ 4 ┆ 3 ┆ 3 ┆ 4 │
│ null_count ┆ 0 ┆ 1 ┆ 1 ┆ 0 │
│ mean ┆ 2013-01-02 00:30:00 ┆ 2013-01-01 22:53:00 ┆ 0:33:00 ┆ 2013-01-01 18:00:00 │
│ std ┆ null ┆ null ┆ null ┆ null │
│ min ┆ 2013-01-01 05:00:00 ┆ 2013-01-01 05:12:00 ┆ -1 day, 23:57:00 ┆ 2013-01-01 │
│ 25% ┆ 2013-01-01 06:00:00 ┆ 2013-01-01 05:57:00 ┆ 0:12:00 ┆ 2013-01-01 │
│ 50% ┆ 2013-01-02 07:00:00 ┆ 2013-01-01 05:57:00 ┆ 0:12:00 ┆ 2013-01-02 │
│ 75% ┆ 2013-01-02 07:00:00 ┆ 2013-01-03 09:30:00 ┆ 1:30:00 ┆ 2013-01-02 │
│ max ┆ 2013-01-03 08:00:00 ┆ 2013-01-03 09:30:00 ┆ 1:30:00 ┆ 2013-01-03 │
└────────────┴─────────────────────┴─────────────────────┴──────────────────┴─────────────────────┘
A series keeps only the rows its dtype has:
> (describe (series (list 3 1 polars-null 4 1 5) #:name "n"))
shape: (9, 2)
┌────────────┬──────────┐
│ statistic ┆ value │
│ --- ┆ --- │
│ str ┆ f64 │
╞════════════╪══════════╡
│ count ┆ 5.0 │
│ null_count ┆ 1.0 │
│ mean ┆ 2.8 │
│ std ┆ 1.788854 │
│ min ┆ 1.0 │
│ 25% ┆ 1.0 │
│ 50% ┆ 3.0 │
│ 75% ┆ 4.0 │
│ max ┆ 5.0 │
└────────────┴──────────┘
> (describe (ref times "day"))
shape: (8, 2)
┌────────────┬─────────────────────┐
│ statistic ┆ value │
│ --- ┆ --- │
│ str ┆ str │
╞════════════╪═════════════════════╡
│ count ┆ 4 │
│ null_count ┆ 0 │
│ mean ┆ 2013-01-01 18:00:00 │
│ min ┆ 2013-01-01 │
│ 25% ┆ 2013-01-01 │
│ 50% ┆ 2013-01-02 │
│ 75% ┆ 2013-01-02 │
│ max ┆ 2013-01-03 │
└────────────┴─────────────────────┘
A nested column (here the lists agg collects) and a null-dtype column report only their counts, as floats; a frame with no rows reports zero counts:
> (~> flights (group-by "carrier") (agg (col "distance")) describe)
shape: (9, 3)
┌────────────┬─────────┬──────────┐
│ statistic ┆ carrier ┆ distance │
│ --- ┆ --- ┆ --- │
│ str ┆ str ┆ f64 │
╞════════════╪═════════╪══════════╡
│ count ┆ 2 ┆ 3.0 │
│ null_count ┆ 1 ┆ 0.0 │
│ mean ┆ null ┆ null │
│ std ┆ null ┆ null │
│ min ┆ AA ┆ null │
│ 25% ┆ null ┆ null │
│ 50% ┆ null ┆ null │
│ 75% ┆ null ┆ null │
│ max ┆ UA ┆ null │
└────────────┴─────────┴──────────┘
> (~> flights (select (alias (cast "carrier" 'null) "nothing")) describe)
shape: (9, 2)
┌────────────┬─────────┐
│ statistic ┆ nothing │
│ --- ┆ --- │
│ str ┆ f64 │
╞════════════╪═════════╡
│ count ┆ 0.0 │
│ null_count ┆ 4.0 │
│ mean ┆ null │
│ std ┆ null │
│ min ┆ null │
│ 25% ┆ null │
│ 50% ┆ null │
│ 75% ┆ null │
│ max ┆ null │
└────────────┴─────────┘
> (describe (head flights 0))
shape: (9, 4)
┌────────────┬─────────┬──────────┬─────────┐
│ statistic ┆ carrier ┆ distance ┆ on_time │
│ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ str ┆ f64 ┆ f64 │
╞════════════╪═════════╪══════════╪═════════╡
│ count ┆ 0 ┆ 0.0 ┆ 0.0 │
│ null_count ┆ 0 ┆ 0.0 ┆ 0.0 │
│ mean ┆ null ┆ null ┆ null │
│ std ┆ null ┆ null ┆ null │
│ min ┆ null ┆ null ┆ null │
│ 25% ┆ null ┆ null ┆ null │
│ 50% ┆ null ┆ null ┆ null │
│ 75% ┆ null ┆ null ┆ null │
│ max ┆ null ┆ null ┆ null │
└────────────┴─────────┴──────────┴─────────┘
2.3.1 Converting to Racket values
procedure
(dataframe->columns d [ #:columns columns #:null null-value]) → (listof (cons/c string? vector?)) d : dataframe? columns : (listof string?) = (column-names d) null-value : any/c = polars-null
> (define kv (dataframe (list (series '("a" "b") #:name "k") (series (list 1 polars-null) #:name "v")))) > (dataframe->columns kv) '(("k" . #("a" "b")) ("v" . #(1 polars-null)))
> (dataframe->columns kv #:columns '("v" "k") #:null 0) '(("v" . #(1 0)) ("k" . #("a" "b")))
> (dataframe->columns kv #:columns '("k" "k")) dataframe->columns: duplicate column
column: "k"
> (dataframe->columns kv #:columns '("nope")) dataframe->columns: no such column
column: "nope"
procedure
(dataframe->hash d [ #:columns columns #:null null-value]) → (and/c (hash/c string? vector?) immutable?) d : dataframe? columns : (listof string?) = (column-names d) null-value : any/c = polars-null
> (dataframe->hash kv) '#hash(("k" . #("a" "b")) ("v" . #(1 polars-null)))
> (hash-ref (dataframe->hash kv #:null 0) "v") '#(1 0)
> (dataframe->hash kv #:columns '("k")) '#hash(("k" . #("a" "b")))
> (dataframe->hash kv #:columns '("nope")) dataframe->hash: no such column
column: "nope"
procedure
(in-dataframe-columns d [#:columns columns]) → sequence?
d : dataframe? columns : (listof string?) = (column-names d)
> (for/list ([column (in-dataframe-columns kv)]) (series-name column)) '("k" "v")
> (for/list ([column (in-dataframe-columns kv #:columns '("v"))]) (series->list column #:null 0)) '((1 0))
> (for/first ([column (in-dataframe-columns kv)]) column)
shape: (2,)
Series: 'k' [str]
[
"a"
"b"
]
> (in-dataframe-columns kv #:columns '("k" "k")) in-dataframe-columns: duplicate column
column: "k"
procedure
(in-dataframe-rows d [ #:columns columns #:named? named? #:null null-value #:buffer-size buffer-size]) → sequence? d : dataframe? columns : (listof string?) = (column-names d) named? : boolean? = #f null-value : any/c = polars-null buffer-size : exact-positive-integer? = 512
The rows are converted buffer-size at a time, as buffer_size does: each buffer is one bulk copy per column, never a foreign call per value. A loop holds one buffer’s values at a time, so memory stays bounded for any height, and one that stops early converts at most one buffer beyond the rows it reads. A larger buffer makes fewer calls and holds more values.
The columns are fetched and checked when in-dataframe-rows is called: an unknown or repeated name, or a column of an unsupported dtype, raises exn:fail:contract naming the column. The sequence holds those columns rather than d, and each iteration starts from the first row. A lazyframe is not accepted; collect it first. Python’s buffer_size=0, a row at a time, has no counterpart.
> (define trips (dataframe (list (series '(UA AA UA) #:name "carrier") (series (list 2 polars-null -3) #:name "delay") (cast (series (list (datetime 2013 1 1) (datetime 2013 1 1) (datetime 2013 1 2)) #:name "day") 'date)))) > (for/list ([row (in-dataframe-rows trips)]) row)
'(#(UA 2 #<date 2013-01-01>)
#(AA polars-null #<date 2013-01-01>)
#(UA -3 #<date 2013-01-02>))
> (for/list ([row (in-dataframe-rows trips #:columns '("delay" "carrier") #:named? #t)]) row)
'(#hash(("carrier" . UA) ("delay" . 2))
#hash(("carrier" . AA) ("delay" . polars-null))
#hash(("carrier" . UA) ("delay" . -3)))
> (for/sum ([row (in-dataframe-rows trips #:columns '("delay") #:null 0)]) (vector-ref row 0)) -1
> (for/list ([row (in-dataframe-rows trips #:buffer-size 2)]) (define-values (carrier delay day) (vector->values row)) (list carrier (date->iso8601 day))) '((UA "2013-01-01") (AA "2013-01-01") (UA "2013-01-02"))
> (in-dataframe-rows trips #:columns '("nope")) in-dataframe-rows: no such column
column: "nope"
> (in-dataframe-rows trips #:buffer-size 0) in-dataframe-rows: contract violation
expected: exact-positive-integer?
given: 0
in: the #:buffer-size argument of
(->*
(dataframe?)
(#:buffer-size
exact-positive-integer?
#:columns
(listof string?)
#:named?
boolean?
#:null
any/c)
sequence?)
contract from:
<pkgs>/polars/private/generic/convert.rkt
blaming: top-level
(assuming the contract is correct)
at: <pkgs>/polars/private/generic/convert.rkt:32:3
procedure
(dataframe->rows d [ #:columns columns #:named? named? #:null null-value]) → (listof (or/c vector? (and/c hash? immutable?))) d : dataframe? columns : (listof string?) = (column-names d) named? : boolean? = #f null-value : any/c = polars-null
> (dataframe->rows trips)
'(#(UA 2 #<date 2013-01-01>)
#(AA polars-null #<date 2013-01-01>)
#(UA -3 #<date 2013-01-02>))
> (dataframe->rows trips #:columns '("carrier") #:named? #t) '(#hash(("carrier" . UA)) #hash(("carrier" . AA)) #hash(("carrier" . UA)))
> (dataframe->rows (head trips 0)) '()
> (dataframe->rows trips #:columns '("day" "day")) dataframe->rows: duplicate column
column: "day"
procedure
(dataframe->f64vector d [ #:columns columns #:order order #:null null-value])
→
f64vector? exact-nonnegative-integer? exact-nonnegative-integer? d : dataframe? columns : (listof string?) = (column-names d) order : (or/c 'fortran 'c) = 'fortran null-value : (or/c real? 'error) = +nan.0
> (define xy (dataframe (list (series (list 1 2 polars-null) #:name "a") (series '(0.5 1.5 2.5) #:name "b")))) > (define-values (m nrows ncols) (dataframe->f64vector xy)) > (list nrows ncols) '(3 2)
> (f64vector->list m) '(1.0 2.0 +nan.0 0.5 1.5 2.5)
> (define-values (m/f rows/f cols/f) (dataframe->f64vector xy #:order 'fortran)) > (equal? (f64vector->list m/f) (f64vector->list m)) #t
> (define-values (m/c rows/c cols/c) (dataframe->f64vector xy #:order 'c)) > (f64vector->list m/c) '(1.0 0.5 2.0 1.5 +nan.0 2.5)
> (define-values (b rows/b cols/b) (dataframe->f64vector xy #:columns '("b") #:null 'error)) > (f64vector->list b) '(0.5 1.5 2.5)
> (define-values (z rows/z cols/z) (dataframe->f64vector xy #:null 0)) > (f64vector->list z) '(1.0 2.0 0.0 0.5 1.5 2.5)
> (dataframe->f64vector xy #:null 'error) dataframe->f64vector: null value
column: "a"
row: 2
> (dataframe->f64vector (dataframe (list (series '("p") #:name "s")))) dataframe->f64vector: not a numeric column
column: "s"
dtype: 'string
2.3.2 Low-level DataFrame API
The generic layer above is built on a set of monomorphic dataframe-* bindings that operate directly on the foreign dataframe. They remain exported and accept the dataframe wrapper (it marshals transparently); the generic operations are simply the preferred surface.
procedure
(dataframe-new columns) → dataframe?
columns : (listof series?)
procedure
(dataframe-shape d) →
exact-nonnegative-integer? exact-nonnegative-integer? d : dataframe?
procedure
d : dataframe?
procedure
d : dataframe?
procedure
(dataframe-column d name) → series?
d : dataframe? name : string?
procedure
(dataframe-column-name d i) → string?
d : dataframe? i : exact-nonnegative-integer?
procedure
(dataframe-column-names d) → (listof string?)
d : dataframe?
procedure
(dataframe-select d names) → dataframe?
d : dataframe? names : (listof string?)
procedure
(display-dataframe d [out]) → void?
d : dataframe? out : output-port? = (current-output-port)
> (define scores (dataframe (list (series '(10 25 18) #:name "score" #:dtype 'i32)))) > (series-sum-i32 (dataframe-column scores "score")) 53
> (dataframe-column scores "points") dataframe-column: no column named "points"
procedure
(DataFrame-ptr? v) → boolean?
v : any/c
procedure
(dataframe-vstack top bottom) → DataFrame-ptr?
top : DataFrame-ptr? bottom : DataFrame-ptr?
procedure
(dataframe-sort d names [ #:descending descending #:nulls-last nulls-last #:maintain-order maintain-order]) → DataFrame-ptr? d : DataFrame-ptr? names : (non-empty-listof string?) descending : (or/c boolean? (listof boolean?)) = #f nulls-last : (or/c boolean? (listof boolean?)) = #f maintain-order : boolean? = #f
2.3.3 Reading & writing
procedure
(dataframe-write-csv d path) → void?
d : dataframe? path : path-string?
procedure
(dataframe-write-parquet d path) → void?
d : dataframe? path : path-string?
procedure
(dataframe-read-parquet path) → dataframe?
path : path-string?
procedure
(dataframe-write-json-lines d path) → void?
d : dataframe? path : path-string?
procedure
(dataframe-read-json-lines path) → dataframe?
path : path-string?
procedure
(dataframe-read-csv path [ #:has-header has-header #:separator separator #:quote-char quote-char #:comment-prefix comment-prefix #:skip-rows skip-rows #:n-rows n-rows #:null-values null-values #:infer-schema-length infer-schema-length #:schema-overrides schema-overrides #:ignore-errors ignore-errors #:try-parse-dates try-parse-dates #:encoding encoding #:glob glob]) → DataFrame-ptr? path : path-string? has-header : boolean? = #t separator : (or/c csv-char/c #f) = #f quote-char : (or/c csv-char/c #f) = #\" comment-prefix : (or/c non-empty-string? #f) = #f skip-rows : exact-nonnegative-integer? = 0 n-rows : (or/c exact-nonnegative-integer? #f) = #f null-values : (or/c string? (listof string?) #f) = #f
infer-schema-length : (or/c exact-nonnegative-integer? #f) = 100
schema-overrides : (and/c (listof (cons/c string? csv-dtype/c)) distinct-names?) = '() ignore-errors : boolean? = #f try-parse-dates : boolean? = #f encoding : (or/c 'utf8 'utf8-lossy) = 'utf8 glob : boolean? = #t
procedure
(lazyframe-scan-csv path [ #:has-header has-header #:separator separator #:quote-char quote-char #:comment-prefix comment-prefix #:skip-rows skip-rows #:n-rows n-rows #:null-values null-values #:infer-schema-length infer-schema-length #:schema-overrides schema-overrides #:ignore-errors ignore-errors #:try-parse-dates try-parse-dates #:encoding encoding #:glob glob]) → LazyFrame-ptr? path : path-string? has-header : boolean? = #t separator : (or/c csv-char/c #f) = #f quote-char : (or/c csv-char/c #f) = #\" comment-prefix : (or/c non-empty-string? #f) = #f skip-rows : exact-nonnegative-integer? = 0 n-rows : (or/c exact-nonnegative-integer? #f) = #f null-values : (or/c string? (listof string?) #f) = #f
infer-schema-length : (or/c exact-nonnegative-integer? #f) = 100
schema-overrides : (and/c (listof (cons/c string? csv-dtype/c)) distinct-names?) = '() ignore-errors : boolean? = #f try-parse-dates : boolean? = #f encoding : (or/c 'utf8 'utf8-lossy) = 'utf8 glob : boolean? = #t
> (dataframe-height (dataframe-read-csv "flights.tsv" #:separator #\tab #:null-values "NA")) 102
> (dataframe-height (lazyframe-collect (lazyframe-scan-csv "parts/*.csv"))) 6
2.4 Lazy frames
A lazyframe is a query plan: a sequence of operations over a frame that Polars optimises as a whole and runs only when asked to collect. The low-level surface mirrors the eager dataframe-* bindings and, like them, returns raw foreign pointers; the fluent lazy and collect are the wrapper-returning equivalents.
procedure
(LazyFrame-ptr? v) → boolean?
v : any/c
procedure
(lazyframe? v) → boolean?
v : any/c
procedure
(dataframe-lazy df) → LazyFrame-ptr?
df : DataFrame-ptr?
procedure
(lazyframe-collect lf) → DataFrame-ptr?
lf : LazyFrame-ptr?
procedure
(lazyframe-select lf exprs) → LazyFrame-ptr?
lf : LazyFrame-ptr? exprs : (listof Expr-ptr?)
procedure
(lazyframe-with-columns lf exprs) → LazyFrame-ptr?
lf : LazyFrame-ptr? exprs : (listof Expr-ptr?)
procedure
(lazyframe-filter lf predicate) → LazyFrame-ptr?
lf : LazyFrame-ptr? predicate : Expr-ptr?
procedure
(lazyframe-group-by-agg lf keys aggs) → LazyFrame-ptr?
lf : LazyFrame-ptr? keys : (listof (or/c string? Expr-ptr?)) aggs : (listof Expr-ptr?)
procedure
(lazyframe-sort lf names [ #:descending descending #:nulls-last nulls-last #:maintain-order maintain-order]) → LazyFrame-ptr? lf : LazyFrame-ptr? names : (non-empty-listof string?) descending : (or/c boolean? (listof boolean?)) = #f nulls-last : (or/c boolean? (listof boolean?)) = #f maintain-order : boolean? = #f
procedure
(lazyframe-join left right [ #:on on #:left-on left-on #:right-on right-on #:how how]) → LazyFrame-ptr? left : LazyFrame-ptr? right : LazyFrame-ptr? on : (or/c #f (listof string?)) = #f left-on : (or/c #f (listof string?)) = #f right-on : (or/c #f (listof string?)) = #f how : (or/c 'inner 'left 'outer 'full 'cross) = 'inner
(lazyframe-collect (lazyframe-join (dataframe-lazy users) (dataframe-lazy orders) #:on '("uid") #:how 'inner))
2.5 Low-level expression API
The monomorphic expr-* layer that the operators in
Operators and pipelines are built from. You rarely need these names directly:
expr-gt underlies >, expr-add underlies
+, expr-sum underlies the expression arm of sum.
Reach for them when a generic name is shadowed in your module, or when you
want to be explicit that an expression —
procedure
(expr-alias e name) → Expr-ptr?
e : Expr-ptr? name : string?
procedure
procedure
(expr-exclude e names) → Expr-ptr?
e : multi-column-expr? names : (non-empty-listof (or/c string? regexp?))
procedure
(expr-dtype-col dtype) → Expr-ptr?
dtype : dtype-spec?
> (expr-all) cs.all()
> (expr-exclude (expr-all) (list "id" #rx"^w")) [cs.all() - [cs.matches("^(?s).*(?:^w).*$") | cs.by_name('id', require_all=false)]]
> (expr-dtype-col 'f64) cs.by_dtype([Float64])
> (select people (expr-exclude (expr-dtype-col 'float64) (list "height")))
shape: (3, 1)
┌────────┐
│ weight │
│ --- │
│ f64 │
╞════════╡
│ 57.9 │
│ 72.5 │
│ 53.6 │
└────────┘
procedure
(expr-meta-output-name e) → string?
e : Expr-ptr?
procedure
(expr-meta-root-names e) → (listof string?)
e : Expr-ptr?
procedure
(expr-meta-eq? a b) → boolean?
a : Expr-ptr? b : Expr-ptr?
> (~> (col "a") (expr-alias "b") expr-meta-output-name) "b"
> (~> (col "a") (expr-add (col "b")) expr-meta-root-names) '("a" "b")
> (expr-meta-eq? (col "a") (expr-col "a")) #t
> (~> (col "v") expr-sum (expr-over (list "k" (col "h")))) col("v").sum().over([col("k"), col("h")])
> (~> khv (with-columns (~> (col "v") expr-sum (expr-over (list "k")) (expr-alias "total"))))
shape: (4, 4)
┌─────┬─────┬─────┬───────┐
│ k ┆ h ┆ v ┆ total │
│ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ str ┆ i64 ┆ i64 │
╞═════╪═════╪═════╪═══════╡
│ a ┆ x ┆ 1 ┆ 6 │
│ a ┆ y ┆ 2 ┆ 6 │
│ a ┆ x ┆ 3 ┆ 6 │
│ b ┆ x ┆ 4 ┆ 4 │
└─────┴─────┴─────┴───────┘
procedure
a : any/c b : any/c
procedure
a : any/c b : any/c
procedure
a : any/c b : any/c
procedure
a : any/c b : any/c
procedure
a : any/c b : any/c
procedure
a : any/c b : any/c
procedure
a : any/c b : any/c
procedure
a : any/c b : any/c
procedure
a : any/c b : any/c
procedure
a : any/c b : any/c
procedure
a : any/c b : any/c
procedure
a : any/c b : any/c
procedure
a : any/c b : any/c
procedure
a : any/c b : any/c
procedure
e : Expr-ptr?
procedure
e : Expr-ptr?
procedure
e : Expr-ptr?
procedure
e : Expr-ptr?
procedure
e : Expr-ptr?
procedure
(expr-median e) → Expr-ptr?
e : Expr-ptr?
procedure
(expr-count e) → Expr-ptr?
e : Expr-ptr?
procedure
(expr-n-unique e) → Expr-ptr?
e : Expr-ptr?
procedure
(expr-first e) → Expr-ptr?
e : Expr-ptr?
procedure
e : Expr-ptr?
procedure
e : Expr-ptr? ddof : exact-nonnegative-integer? = 1
procedure
e : Expr-ptr? ddof : exact-nonnegative-integer? = 1
procedure
(expr-sort e [ #:descending descending #:nulls-last nulls-last]) → Expr-ptr? e : Expr-ptr? descending : boolean? = #f nulls-last : boolean? = #f
procedure
(expr-sort-by e #:by by [ #:descending descending #:nulls-last nulls-last #:maintain-order maintain-order]) → Expr-ptr? e : Expr-ptr?
by :
(or/c string? Expr-ptr? (non-empty-listof (or/c string? Expr-ptr?))) descending : (or/c boolean? (listof boolean?)) = #f nulls-last : (or/c boolean? (listof boolean?)) = #f maintain-order : boolean? = #f
2.5.1 Eager expression contexts
These run expressions against a dataframe and return a new frame in one step. Each is the eager convenience over the corresponding lazy operation in Lazy frames: it converts with dataframe-lazy, applies the operation, and lazyframe-collects. Like the rest of the low-level layer they accept a dataframe wrapper but return a raw DataFrame-ptr?; the fluent select, with-columns, filter and group-by/agg are the wrapper-returning equivalents.
procedure
(dataframe-select-exprs df exprs) → DataFrame-ptr?
df : DataFrame-ptr? exprs : (listof Expr-ptr?)
procedure
(dataframe-with-columns df exprs) → DataFrame-ptr?
df : DataFrame-ptr? exprs : (listof Expr-ptr?)
procedure
(dataframe-filter-expr df predicate) → DataFrame-ptr?
df : DataFrame-ptr? predicate : Expr-ptr?
procedure
(dataframe-group-by-agg df keys aggs) → DataFrame-ptr?
df : DataFrame-ptr? keys : (listof (or/c string? Expr-ptr?)) aggs : (listof Expr-ptr?)
2.6 Generic interfaces
The high-level operations are small, purpose-named racket/generic interfaces. A wrapper implements the interface for each capability it has — a series and a dataframe both have a len and a shape, so both implement gen:sized and gen:has-shape; only a series has a dtype. Each interface exports its method(s) and a predicate that recognises values implementing it.
syntax
procedure
v : any/c
syntax
procedure
(has-shape? v) → boolean?
v : any/c
syntax
procedure
(has-dtype? v) → boolean?
v : any/c
syntax
procedure
(has-null-count? v) → boolean?
v : any/c
A series is also a Racket sequence (through prop:sequence): a for clause, sequence? and the racket/sequence operations see its elements, converted a block of rows at a time as by in-series, with polars-null for a null entry. A dataframe is not a sequence; iterate over its rows with in-dataframe-rows or its columns with in-dataframe-columns, or convert them with dataframe->rows or dataframe->columns.
> (define ages (series (list 34 polars-null 51) #:name "age")) > (sequence? ages) #t
> (for/list ([age ages]) age) '(34 polars-null 51)
> (for/sum ([age ages] #:unless (polars-null? age)) age) 85
> (sequence? (dataframe (list ages))) #f