On this page:
1.2.2.1 Expressions
1.2.2.2 Contexts
1.2.2.2.1 select
1.2.2.2.2 with-columns
1.2.2.2.3 filter
1.2.2.2.4 group-by and aggregations
1.2.2.3 Expression expansion

1.2.2 Expressions and contexts🔗ℹ

1.2.2.1 Expressions🔗ℹ

An expression is a value; nothing runs until a context receives it.

> (define bmi-expr (/ (col "weight") (pow (col "height") 2)))
> bmi-expr

[(col("weight")) rust_div (col("height").pow([dyn int: 2]))]

An expression prints as its plan, in the notation Polars itself uses: the binding’s / is Polars’ RustDivide operator, shown as rust_div, and pow appears as a method suffix.

1.2.2.2 Contexts🔗ℹ
1.2.2.2.1 select🔗ℹ
> (~> df
      (select (alias bmi-expr "bmi")
              (~> bmi-expr mean (alias "avg_bmi"))
              (alias (lit 25) "ideal_max_bmi")))

shape: (4, 3)

┌───────────┬───────────┬───────────────┐

│ bmi       ┆ avg_bmi   ┆ ideal_max_bmi │

│ ---       ┆ ---       ┆ ---           │

│ f64       ┆ f64       ┆ i32           │

╞═══════════╪═══════════╪═══════════════╡

│ 23.791913 ┆ 23.438973 ┆ 25            │

│ 23.141498 ┆ 23.438973 ┆ 25            │

│ 19.687787 ┆ 23.438973 ┆ 25            │

│ 27.134694 ┆ 23.438973 ┆ 25            │

└───────────┴───────────┴───────────────┘

> (~> df
      (select (~> bmi-expr
                  (- (mean bmi-expr))
                  (/ (std bmi-expr))
                  (alias "deviation"))))

shape: (4, 1)

┌───────────┐

│ deviation │

│ ---       │

│ f64       │

╞═══════════╡

│ 0.115645  │

│ -0.097471 │

│ -1.22912  │

│ 1.210946  │

└───────────┘

1.2.2.2.2 with-columns🔗ℹ
> (~> df
      (with-columns (alias bmi-expr "bmi")
                    (~> bmi-expr mean (alias "avg_bmi"))
                    (alias (lit 25) "ideal_max_bmi")))

shape: (4, 7)

┌────────────────┬────────────┬────────┬────────┬───────────┬───────────┬───────────────┐

│ name           ┆ birthdate  ┆ weight ┆ height ┆ bmi       ┆ avg_bmi   ┆ ideal_max_bmi │

│ ---            ┆ ---        ┆ ---    ┆ ---    ┆ ---       ┆ ---       ┆ ---           │

│ str            ┆ date       ┆ f64    ┆ f64    ┆ f64       ┆ f64       ┆ i32           │

╞════════════════╪════════════╪════════╪════════╪═══════════╪═══════════╪═══════════════╡

│ Alice Archer   ┆ 1997-01-10 ┆ 57.9   ┆ 1.56   ┆ 23.791913 ┆ 23.438973 ┆ 25            │

│ Ben Brown      ┆ 1985-02-15 ┆ 72.5   ┆ 1.77   ┆ 23.141498 ┆ 23.438973 ┆ 25            │

│ Chloe Cooper   ┆ 1983-03-22 ┆ 53.6   ┆ 1.65   ┆ 19.687787 ┆ 23.438973 ┆ 25            │

│ Daniel Donovan ┆ 1981-04-30 ┆ 83.1   ┆ 1.75   ┆ 27.134694 ┆ 23.438973 ┆ 25            │

└────────────────┴────────────┴────────┴────────┴───────────┴───────────┴───────────────┘

1.2.2.2.3 filter🔗ℹ
> (~> df
      (filter (and (is-between "birthdate"
                               (str->date (lit "1982-12-31"))
                               (str->date (lit "1996-01-01")))
                   (> (col "height") 1.7))))

shape: (1, 4)

┌───────────┬────────────┬────────┬────────┐

│ name      ┆ birthdate  ┆ weight ┆ height │

│ ---       ┆ ---        ┆ ---    ┆ ---    │

│ str       ┆ date       ┆ f64    ┆ f64    │

╞═══════════╪════════════╪════════╪════════╡

│ Ben Brown ┆ 1985-02-15 ┆ 72.5   ┆ 1.77   │

└───────────┴────────────┴────────┴────────┘

API gaps: no date literals, so the bounds parse a string with str->date; filter takes one predicate, so combine with and.

1.2.2.2.4 group-by and aggregations🔗ℹ

Group keys may be expressions. A bare (col "name") inside agg collects the group’s values into a list.

> (define decade
    (~> (col "birthdate") dt-year (/ 10) (* 10) (alias "decade")))
> (~> df (group-by decade) (agg (col "name")))

shape: (2, 2)

┌────────┬─────────────────────────────────┐

│ decade ┆ name                            │

│ ---    ┆ ---                             │

│ i32    ┆ list[str]                       │

╞════════╪═════════════════════════════════╡

│ 1990   ┆ ["Alice Archer"]                │

│ 1980   ┆ ["Ben Brown", "Chloe Cooper", … │

└────────┴─────────────────────────────────┘

> (~> df
      (group-by decade (alias (< (col "height") 1.7) "short?"))
      (agg (col "name")))

shape: (3, 3)

┌────────┬────────┬─────────────────────────────────┐

│ decade ┆ short? ┆ name                            │

│ ---    ┆ ---    ┆ ---                             │

│ i32    ┆ bool   ┆ list[str]                       │

╞════════╪════════╪═════════════════════════════════╡

│ 1980   ┆ false  ┆ ["Ben Brown", "Daniel Donovan"… │

│ 1980   ┆ true   ┆ ["Chloe Cooper"]                │

│ 1990   ┆ true   ┆ ["Alice Archer"]                │

└────────┴────────┴─────────────────────────────────┘

> (~> df
      (group-by decade (alias (< (col "height") 1.7) "short?"))
      (agg (~> (col "name") count (alias "len"))
           (~> (col "height") max (alias "tallest"))
           (~> (col "weight") mean (alias "avg_weight"))
           (~> (col "height") mean (alias "avg_height"))))

shape: (3, 6)

┌────────┬────────┬─────┬─────────┬────────────┬────────────┐

│ decade ┆ short? ┆ len ┆ tallest ┆ avg_weight ┆ avg_height │

│ ---    ┆ ---    ┆ --- ┆ ---     ┆ ---        ┆ ---        │

│ i32    ┆ bool   ┆ u32 ┆ f64     ┆ f64        ┆ f64        │

╞════════╪════════╪═════╪═════════╪════════════╪════════════╡

│ 1980   ┆ false  ┆ 2   ┆ 1.77    ┆ 77.8       ┆ 1.76       │

│ 1980   ┆ true   ┆ 1   ┆ 1.65    ┆ 53.6       ┆ 1.65       │

│ 1990   ┆ true   ┆ 1   ┆ 1.56    ┆ 57.9       ┆ 1.56       │

└────────┴────────┴─────┴─────────┴────────────┴────────────┘

An aggregation followed by over is computed per group but broadcast back to every row of the group, so it goes in with-columns rather than agg; the keys are whatever group-by takes.

> (~> df
      (with-columns (~> (col "height") mean (over decade) (alias "decade_avg_height"))))

shape: (4, 5)

┌────────────────┬────────────┬────────┬────────┬───────────────────┐

│ name           ┆ birthdate  ┆ weight ┆ height ┆ decade_avg_height │

│ ---            ┆ ---        ┆ ---    ┆ ---    ┆ ---               │

│ str            ┆ date       ┆ f64    ┆ f64    ┆ f64               │

╞════════════════╪════════════╪════════╪════════╪═══════════════════╡

│ Alice Archer   ┆ 1997-01-10 ┆ 57.9   ┆ 1.56   ┆ 1.56              │

│ Ben Brown      ┆ 1985-02-15 ┆ 72.5   ┆ 1.77   ┆ 1.723333          │

│ Chloe Cooper   ┆ 1983-03-22 ┆ 53.6   ┆ 1.65   ┆ 1.723333          │

│ Daniel Donovan ┆ 1981-04-30 ┆ 83.1   ┆ 1.75   ┆ 1.723333          │

└────────────────┴────────────┴────────┴────────┴───────────────────┘

API gaps: no pl.len(); no multi-name col("weight", "height"); no name.prefix.

1.2.2.3 Expression expansion🔗ℹ

An expression over a multi-column col expands to one expression per matched column. (col 'float64) is every 'float64 column (pl.col(pl.Float64)), and the outputs keep the matched names.

> (define expr (* (col 'float64) 1.1))
> (select df expr)

shape: (4, 2)

┌────────┬────────┐

│ weight ┆ height │

│ ---    ┆ ---    │

│ f64    ┆ f64    │

╞════════╪════════╡

│ 63.69  ┆ 1.716  │

│ 79.75  ┆ 1.947  │

│ 58.96  ┆ 1.815  │

│ 91.41  ┆ 1.925  │

└────────┴────────┘

> (define df2
    (dataframe (list (series '(1 2 3 4) #:name "ints")
                     (series '("A" "B" "C" "D") #:name "letters"))))
> (select df2 expr)

shape: (0, 0)

┌┐

╞╡

└┘

API gap: no name.suffix, so the expanded columns cannot be renamed weight*1.1 / height*1.1; they keep the matched names.