1.2.2 Expressions and contexts
1.2.2.1 Expressions
An expression is a value; nothing runs until a context receives it.
> (define bmi-expr (/ (col "weight") (pow (col "height") 2))) > bmi-expr [(col("weight")) rust_div (col("height").pow([dyn int: 2]))]
An expression prints as its plan, in the notation Polars itself uses: the binding’s / is Polars’ RustDivide operator, shown as rust_div, and pow appears as a method suffix.
1.2.2.2 Contexts
1.2.2.2.1 select
> (~> df (select (alias bmi-expr "bmi") (~> bmi-expr mean (alias "avg_bmi")) (alias (lit 25) "ideal_max_bmi")))
shape: (4, 3)
┌───────────┬───────────┬───────────────┐
│ bmi ┆ avg_bmi ┆ ideal_max_bmi │
│ --- ┆ --- ┆ --- │
│ f64 ┆ f64 ┆ i32 │
╞═══════════╪═══════════╪═══════════════╡
│ 23.791913 ┆ 23.438973 ┆ 25 │
│ 23.141498 ┆ 23.438973 ┆ 25 │
│ 19.687787 ┆ 23.438973 ┆ 25 │
│ 27.134694 ┆ 23.438973 ┆ 25 │
└───────────┴───────────┴───────────────┘
> (~> df (select (~> bmi-expr (- (mean bmi-expr)) (/ (std bmi-expr)) (alias "deviation"))))
shape: (4, 1)
┌───────────┐
│ deviation │
│ --- │
│ f64 │
╞═══════════╡
│ 0.115645 │
│ -0.097471 │
│ -1.22912 │
│ 1.210946 │
└───────────┘
1.2.2.2.2 with-columns
> (~> df (with-columns (alias bmi-expr "bmi") (~> bmi-expr mean (alias "avg_bmi")) (alias (lit 25) "ideal_max_bmi")))
shape: (4, 7)
┌────────────────┬────────────┬────────┬────────┬───────────┬───────────┬───────────────┐
│ name ┆ birthdate ┆ weight ┆ height ┆ bmi ┆ avg_bmi ┆ ideal_max_bmi │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ date ┆ f64 ┆ f64 ┆ f64 ┆ f64 ┆ i32 │
╞════════════════╪════════════╪════════╪════════╪═══════════╪═══════════╪═══════════════╡
│ Alice Archer ┆ 1997-01-10 ┆ 57.9 ┆ 1.56 ┆ 23.791913 ┆ 23.438973 ┆ 25 │
│ Ben Brown ┆ 1985-02-15 ┆ 72.5 ┆ 1.77 ┆ 23.141498 ┆ 23.438973 ┆ 25 │
│ Chloe Cooper ┆ 1983-03-22 ┆ 53.6 ┆ 1.65 ┆ 19.687787 ┆ 23.438973 ┆ 25 │
│ Daniel Donovan ┆ 1981-04-30 ┆ 83.1 ┆ 1.75 ┆ 27.134694 ┆ 23.438973 ┆ 25 │
└────────────────┴────────────┴────────┴────────┴───────────┴───────────┴───────────────┘
1.2.2.2.3 filter
> (~> df (filter (and (is-between "birthdate" (str->date (lit "1982-12-31")) (str->date (lit "1996-01-01"))) (> (col "height") 1.7))))
shape: (1, 4)
┌───────────┬────────────┬────────┬────────┐
│ name ┆ birthdate ┆ weight ┆ height │
│ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ date ┆ f64 ┆ f64 │
╞═══════════╪════════════╪════════╪════════╡
│ Ben Brown ┆ 1985-02-15 ┆ 72.5 ┆ 1.77 │
└───────────┴────────────┴────────┴────────┘
API gaps: no date literals, so the bounds parse a string with str->date; filter takes one predicate, so combine with and.
1.2.2.2.4 group-by and aggregations
Group keys may be expressions. A bare (col "name") inside agg collects the group’s values into a list.
> (define decade (~> (col "birthdate") dt-year (/ 10) (* 10) (alias "decade"))) > (~> df (group-by decade) (agg (col "name")))
shape: (2, 2)
┌────────┬─────────────────────────────────┐
│ decade ┆ name │
│ --- ┆ --- │
│ i32 ┆ list[str] │
╞════════╪═════════════════════════════════╡
│ 1990 ┆ ["Alice Archer"] │
│ 1980 ┆ ["Ben Brown", "Chloe Cooper", … │
└────────┴─────────────────────────────────┘
> (~> df (group-by decade (alias (< (col "height") 1.7) "short?")) (agg (col "name")))
shape: (3, 3)
┌────────┬────────┬─────────────────────────────────┐
│ decade ┆ short? ┆ name │
│ --- ┆ --- ┆ --- │
│ i32 ┆ bool ┆ list[str] │
╞════════╪════════╪═════════════════════════════════╡
│ 1980 ┆ false ┆ ["Ben Brown", "Daniel Donovan"… │
│ 1980 ┆ true ┆ ["Chloe Cooper"] │
│ 1990 ┆ true ┆ ["Alice Archer"] │
└────────┴────────┴─────────────────────────────────┘
> (~> df (group-by decade (alias (< (col "height") 1.7) "short?")) (agg (~> (col "name") count (alias "len")) (~> (col "height") max (alias "tallest")) (~> (col "weight") mean (alias "avg_weight")) (~> (col "height") mean (alias "avg_height"))))
shape: (3, 6)
┌────────┬────────┬─────┬─────────┬────────────┬────────────┐
│ decade ┆ short? ┆ len ┆ tallest ┆ avg_weight ┆ avg_height │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ i32 ┆ bool ┆ u32 ┆ f64 ┆ f64 ┆ f64 │
╞════════╪════════╪═════╪═════════╪════════════╪════════════╡
│ 1980 ┆ false ┆ 2 ┆ 1.77 ┆ 77.8 ┆ 1.76 │
│ 1980 ┆ true ┆ 1 ┆ 1.65 ┆ 53.6 ┆ 1.65 │
│ 1990 ┆ true ┆ 1 ┆ 1.56 ┆ 57.9 ┆ 1.56 │
└────────┴────────┴─────┴─────────┴────────────┴────────────┘
An aggregation followed by over is computed per group but broadcast back to every row of the group, so it goes in with-columns rather than agg; the keys are whatever group-by takes.
> (~> df (with-columns (~> (col "height") mean (over decade) (alias "decade_avg_height"))))
shape: (4, 5)
┌────────────────┬────────────┬────────┬────────┬───────────────────┐
│ name ┆ birthdate ┆ weight ┆ height ┆ decade_avg_height │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ date ┆ f64 ┆ f64 ┆ f64 │
╞════════════════╪════════════╪════════╪════════╪═══════════════════╡
│ Alice Archer ┆ 1997-01-10 ┆ 57.9 ┆ 1.56 ┆ 1.56 │
│ Ben Brown ┆ 1985-02-15 ┆ 72.5 ┆ 1.77 ┆ 1.723333 │
│ Chloe Cooper ┆ 1983-03-22 ┆ 53.6 ┆ 1.65 ┆ 1.723333 │
│ Daniel Donovan ┆ 1981-04-30 ┆ 83.1 ┆ 1.75 ┆ 1.723333 │
└────────────────┴────────────┴────────┴────────┴───────────────────┘
API gaps: no pl.len(); no multi-name col("weight", "height"); no name.prefix.
1.2.2.3 Expression expansion
An expression over a multi-column col expands to one expression per matched column. (col 'float64) is every 'float64 column (pl.col(pl.Float64)), and the outputs keep the matched names.
> (define expr (* (col 'float64) 1.1)) > (select df expr)
shape: (4, 2)
┌────────┬────────┐
│ weight ┆ height │
│ --- ┆ --- │
│ f64 ┆ f64 │
╞════════╪════════╡
│ 63.69 ┆ 1.716 │
│ 79.75 ┆ 1.947 │
│ 58.96 ┆ 1.815 │
│ 91.41 ┆ 1.925 │
└────────┴────────┘
> (define df2 (dataframe (list (series '(1 2 3 4) #:name "ints") (series '("A" "B" "C" "D") #:name "letters")))) > (select df2 expr)
shape: (0, 0)
┌┐
╞╡
└┘
API gap: no name.suffix, so the expanded columns cannot be renamed weight*1.1 / height*1.1; they keep the matched names.