On this page:
1.3.1.1 Function col
1.3.1.1.1 Explicit expansion by column name
1.3.1.1.2 Expansion by data type
1.3.1.1.3 Expansion by pattern matching
1.3.1.1.4 Arguments cannot be of mixed types
1.3.1.2 Selecting all columns
1.3.1.3 Excluding columns
1.3.1.4 Column renaming
1.3.1.4.1 Renaming a single column with alias
1.3.1.4.2 Prefixing and suffixing column names
1.3.1.4.3 Dynamic name replacement
1.3.1.5 Programmatically generating expressions
1.3.1.6 More flexible column selections
1.3.1.6.1 Debugging selectors

1.3.1 Expression expansion🔗ℹ

One expression can stand for several columns; it expands to one expression per matched column when a context runs it. The stock prices below are upstream’s.

> (define df
    (dataframe
     (list (series '("AAPL" "NVDA" "MSFT" "GOOG" "AMZN") #:name "ticker")
           (series '("Apple" "NVIDIA" "Microsoft" "Alphabet (Google)" "Amazon")
                   #:name "company_name")
           (series '(229.9 138.93 420.56 166.41 188.4) #:name "price")
           (series '(231.31 139.6 424.04 167.62 189.83) #:name "day_high")
           (series '(228.6 136.3 417.52 164.78 188.44) #:name "day_low")
           (series '(237.23 140.76 468.35 193.31 201.2) #:name "year_high")
           (series '(164.08 39.23 324.39 121.46 118.35) #:name "year_low"))))
> df

shape: (5, 7)

┌────────┬───────────────────┬────────┬──────────┬─────────┬───────────┬──────────┐

│ ticker ┆ company_name      ┆ price  ┆ day_high ┆ day_low ┆ year_high ┆ year_low │

│ ---    ┆ ---               ┆ ---    ┆ ---      ┆ ---     ┆ ---       ┆ ---      │

│ str    ┆ str               ┆ f64    ┆ f64      ┆ f64     ┆ f64       ┆ f64      │

╞════════╪═══════════════════╪════════╪══════════╪═════════╪═══════════╪══════════╡

│ AAPL   ┆ Apple             ┆ 229.9  ┆ 231.31   ┆ 228.6   ┆ 237.23    ┆ 164.08   │

│ NVDA   ┆ NVIDIA            ┆ 138.93 ┆ 139.6    ┆ 136.3   ┆ 140.76    ┆ 39.23    │

│ MSFT   ┆ Microsoft         ┆ 420.56 ┆ 424.04   ┆ 417.52  ┆ 468.35    ┆ 324.39   │

│ GOOG   ┆ Alphabet (Google) ┆ 166.41 ┆ 167.62   ┆ 164.78  ┆ 193.31    ┆ 121.46   │

│ AMZN   ┆ Amazon            ┆ 188.4  ┆ 189.83   ┆ 188.44  ┆ 201.2     ┆ 118.35   │

└────────┴───────────────────┴────────┴──────────┴─────────┴───────────┴──────────┘

1.3.1.1 Function col🔗ℹ
1.3.1.1.1 Explicit expansion by column name🔗ℹ

API gap: col takes one name, so there is no pl.col("price", "day_high", ...). A list of expressions is spliced into the context, so build one per name.

> (define eur-usd-rate 1.09)
> (~> df
      (with-columns
       (for/list ([name '("price" "day_high" "day_low" "year_high" "year_low")])
         (~> (col name) (/ eur-usd-rate) (round #:decimals 2)))))

shape: (5, 7)

┌────────┬───────────────────┬────────┬──────────┬─────────┬───────────┬──────────┐

│ ticker ┆ company_name      ┆ price  ┆ day_high ┆ day_low ┆ year_high ┆ year_low │

│ ---    ┆ ---               ┆ ---    ┆ ---      ┆ ---     ┆ ---       ┆ ---      │

│ str    ┆ str               ┆ f64    ┆ f64      ┆ f64     ┆ f64       ┆ f64      │

╞════════╪═══════════════════╪════════╪══════════╪═════════╪═══════════╪══════════╡

│ AAPL   ┆ Apple             ┆ 210.92 ┆ 212.21   ┆ 209.72  ┆ 217.64    ┆ 150.53   │

│ NVDA   ┆ NVIDIA            ┆ 127.46 ┆ 128.07   ┆ 125.05  ┆ 129.14    ┆ 35.99    │

│ MSFT   ┆ Microsoft         ┆ 385.83 ┆ 389.03   ┆ 383.05  ┆ 429.68    ┆ 297.61   │

│ GOOG   ┆ Alphabet (Google) ┆ 152.67 ┆ 153.78   ┆ 151.17  ┆ 177.35    ┆ 111.43   │

│ AMZN   ┆ Amazon            ┆ 172.84 ┆ 174.16   ┆ 172.88  ┆ 184.59    ┆ 108.58   │

└────────┴───────────────────┴────────┴──────────┴─────────┴───────────┴──────────┘

1.3.1.1.2 Expansion by data type🔗ℹ

A dtype selects every column of that dtype; 'f64 and 'float64 are the same selector.

> (~> df (with-columns (~> (col 'float64) (/ eur-usd-rate) (round #:decimals 2))))

shape: (5, 7)

┌────────┬───────────────────┬────────┬──────────┬─────────┬───────────┬──────────┐

│ ticker ┆ company_name      ┆ price  ┆ day_high ┆ day_low ┆ year_high ┆ year_low │

│ ---    ┆ ---               ┆ ---    ┆ ---      ┆ ---     ┆ ---       ┆ ---      │

│ str    ┆ str               ┆ f64    ┆ f64      ┆ f64     ┆ f64       ┆ f64      │

╞════════╪═══════════════════╪════════╪══════════╪═════════╪═══════════╪══════════╡

│ AAPL   ┆ Apple             ┆ 210.92 ┆ 212.21   ┆ 209.72  ┆ 217.64    ┆ 150.53   │

│ NVDA   ┆ NVIDIA            ┆ 127.46 ┆ 128.07   ┆ 125.05  ┆ 129.14    ┆ 35.99    │

│ MSFT   ┆ Microsoft         ┆ 385.83 ┆ 389.03   ┆ 383.05  ┆ 429.68    ┆ 297.61   │

│ GOOG   ┆ Alphabet (Google) ┆ 152.67 ┆ 153.78   ┆ 151.17  ┆ 177.35    ┆ 111.43   │

│ AMZN   ┆ Amazon            ┆ 172.84 ┆ 174.16   ┆ 172.88  ┆ 184.59    ┆ 108.58   │

└────────┴───────────────────┴────────┴──────────┴─────────┴───────────┴──────────┘

API gap: col takes one dtype, and selectors do not compose yet (#50), so pl.col(pl.Float32, pl.Float64) is one expression per dtype. A dtype no column has matches nothing.

> (~> df
      (with-columns
       (for/list ([type '(float32 float64)])
         (~> (col type) (/ eur-usd-rate) (round #:decimals 2)))))

shape: (5, 7)

┌────────┬───────────────────┬────────┬──────────┬─────────┬───────────┬──────────┐

│ ticker ┆ company_name      ┆ price  ┆ day_high ┆ day_low ┆ year_high ┆ year_low │

│ ---    ┆ ---               ┆ ---    ┆ ---      ┆ ---     ┆ ---       ┆ ---      │

│ str    ┆ str               ┆ f64    ┆ f64      ┆ f64     ┆ f64       ┆ f64      │

╞════════╪═══════════════════╪════════╪══════════╪═════════╪═══════════╪══════════╡

│ AAPL   ┆ Apple             ┆ 210.92 ┆ 212.21   ┆ 209.72  ┆ 217.64    ┆ 150.53   │

│ NVDA   ┆ NVIDIA            ┆ 127.46 ┆ 128.07   ┆ 125.05  ┆ 129.14    ┆ 35.99    │

│ MSFT   ┆ Microsoft         ┆ 385.83 ┆ 389.03   ┆ 383.05  ┆ 429.68    ┆ 297.61   │

│ GOOG   ┆ Alphabet (Google) ┆ 152.67 ┆ 153.78   ┆ 151.17  ┆ 177.35    ┆ 111.43   │

│ AMZN   ┆ Amazon            ┆ 172.84 ┆ 174.16   ┆ 172.88  ┆ 184.59    ┆ 108.58   │

└────────┴───────────────────┴────────┴──────────┴─────────┴───────────┴──────────┘

1.3.1.1.3 Expansion by pattern matching🔗ℹ

A string of the form ^...$ is a Polars regex, as in Python. The context takes several specs, where Python’s col takes several names.

> (select df "ticker" (col "^.*_high$") (col "^.*_low$"))

shape: (5, 5)

┌────────┬──────────┬───────────┬─────────┬──────────┐

│ ticker ┆ day_high ┆ year_high ┆ day_low ┆ year_low │

│ ---    ┆ ---      ┆ ---       ┆ ---     ┆ ---      │

│ str    ┆ f64      ┆ f64       ┆ f64     ┆ f64      │

╞════════╪══════════╪═══════════╪═════════╪══════════╡

│ AAPL   ┆ 231.31   ┆ 237.23    ┆ 228.6   ┆ 164.08   │

│ NVDA   ┆ 139.6    ┆ 140.76    ┆ 136.3   ┆ 39.23    │

│ MSFT   ┆ 424.04   ┆ 468.35    ┆ 417.52  ┆ 324.39   │

│ GOOG   ┆ 167.62   ┆ 193.31    ┆ 164.78  ┆ 121.46   │

│ AMZN   ┆ 189.83   ┆ 201.2     ┆ 188.44  ┆ 118.35   │

└────────┴──────────┴───────────┴─────────┴──────────┘

A Racket regexp is also accepted, and keeps its Racket meaning: it selects exactly the names regexp-match? accepts, so it is unanchored unless it says otherwise.

> (select df "ticker" (col #rx"_(high|low)$"))

shape: (5, 5)

┌────────┬──────────┬─────────┬───────────┬──────────┐

│ ticker ┆ day_high ┆ day_low ┆ year_high ┆ year_low │

│ ---    ┆ ---      ┆ ---     ┆ ---       ┆ ---      │

│ str    ┆ f64      ┆ f64     ┆ f64       ┆ f64      │

╞════════╪══════════╪═════════╪═══════════╪══════════╡

│ AAPL   ┆ 231.31   ┆ 228.6   ┆ 237.23    ┆ 164.08   │

│ NVDA   ┆ 139.6    ┆ 136.3   ┆ 140.76    ┆ 39.23    │

│ MSFT   ┆ 424.04   ┆ 417.52  ┆ 468.35    ┆ 324.39   │

│ GOOG   ┆ 167.62   ┆ 164.78  ┆ 193.31    ┆ 121.46   │

│ AMZN   ┆ 189.83   ┆ 188.44  ┆ 201.2     ┆ 118.35   │

└────────┴──────────┴─────────┴───────────┴──────────┘

> (select df (col #px"^\\w+_high$"))

shape: (5, 2)

┌──────────┬───────────┐

│ day_high ┆ year_high │

│ ---      ┆ ---       │

│ f64      ┆ f64       │

╞══════════╪═══════════╡

│ 231.31   ┆ 237.23    │

│ 139.6    ┆ 140.76    │

│ 424.04   ┆ 468.35    │

│ 167.62   ┆ 193.31    │

│ 189.83   ┆ 201.2     │

└──────────┴───────────┘

Polars’ engine has no lookaround or backreferences, so such a regexp fails when the query runs. Match the names in Racket instead.

> (define not-day #px"^(?!day_).*_(high|low)$")
> (select df (col not-day))

lazyframe-collect: failed to collect the query: invalid

regex in selector '^(?s).*(?:^(?!day_).*_(high|low)$).*$'

Resolved plan until failure:

---> FAILED HERE RESOLVING 'select' <---

DF ["ticker", "company_name", "price", "day_high", ...];

PROJECT */7 COLUMNS: 'select'

> (select df (filter (lambda (name) (regexp-match? not-day name))
                     (column-names df)))

shape: (5, 2)

┌───────────┬──────────┐

│ year_high ┆ year_low │

│ ---       ┆ ---      │

│ f64       ┆ f64      │

╞═══════════╪══════════╡

│ 237.23    ┆ 164.08   │

│ 140.76    ┆ 39.23    │

│ 468.35    ┆ 324.39   │

│ 193.31    ┆ 121.46   │

│ 201.2     ┆ 118.35   │

└───────────┴──────────┘

1.3.1.1.4 Arguments cannot be of mixed types🔗ℹ

Each col is one name, dtype or pattern, so the mixture is written as separate specs.

> (select df "ticker" (col 'float64))

shape: (5, 6)

┌────────┬────────┬──────────┬─────────┬───────────┬──────────┐

│ ticker ┆ price  ┆ day_high ┆ day_low ┆ year_high ┆ year_low │

│ ---    ┆ ---    ┆ ---      ┆ ---     ┆ ---       ┆ ---      │

│ str    ┆ f64    ┆ f64      ┆ f64     ┆ f64       ┆ f64      │

╞════════╪════════╪══════════╪═════════╪═══════════╪══════════╡

│ AAPL   ┆ 229.9  ┆ 231.31   ┆ 228.6   ┆ 237.23    ┆ 164.08   │

│ NVDA   ┆ 138.93 ┆ 139.6    ┆ 136.3   ┆ 140.76    ┆ 39.23    │

│ MSFT   ┆ 420.56 ┆ 424.04   ┆ 417.52  ┆ 468.35    ┆ 324.39   │

│ GOOG   ┆ 166.41 ┆ 167.62   ┆ 164.78  ┆ 193.31    ┆ 121.46   │

│ AMZN   ┆ 188.4  ┆ 189.83   ┆ 188.44  ┆ 201.2     ┆ 118.35   │

└────────┴────────┴──────────┴─────────┴───────────┴──────────┘

1.3.1.2 Selecting all columns🔗ℹ
> (select df (all))

shape: (5, 7)

┌────────┬───────────────────┬────────┬──────────┬─────────┬───────────┬──────────┐

│ ticker ┆ company_name      ┆ price  ┆ day_high ┆ day_low ┆ year_high ┆ year_low │

│ ---    ┆ ---               ┆ ---    ┆ ---      ┆ ---     ┆ ---       ┆ ---      │

│ str    ┆ str               ┆ f64    ┆ f64      ┆ f64     ┆ f64       ┆ f64      │

╞════════╪═══════════════════╪════════╪══════════╪═════════╪═══════════╪══════════╡

│ AAPL   ┆ Apple             ┆ 229.9  ┆ 231.31   ┆ 228.6   ┆ 237.23    ┆ 164.08   │

│ NVDA   ┆ NVIDIA            ┆ 138.93 ┆ 139.6    ┆ 136.3   ┆ 140.76    ┆ 39.23    │

│ MSFT   ┆ Microsoft         ┆ 420.56 ┆ 424.04   ┆ 417.52  ┆ 468.35    ┆ 324.39   │

│ GOOG   ┆ Alphabet (Google) ┆ 166.41 ┆ 167.62   ┆ 164.78  ┆ 193.31    ┆ 121.46   │

│ AMZN   ┆ Amazon            ┆ 188.4  ┆ 189.83   ┆ 188.44  ┆ 201.2     ┆ 118.35   │

└────────┴───────────────────┴────────┴──────────┴─────────┴───────────┴──────────┘

API gap: no DataFrame.equals.

1.3.1.3 Excluding columns🔗ℹ

exclude takes names, Polars regex strings and Racket regexps, read as col reads them, and applies to any multi-column expression.

> (select df (exclude (all) "^day_.*$"))

shape: (5, 5)

┌────────┬───────────────────┬────────┬───────────┬──────────┐

│ ticker ┆ company_name      ┆ price  ┆ year_high ┆ year_low │

│ ---    ┆ ---               ┆ ---    ┆ ---       ┆ ---      │

│ str    ┆ str               ┆ f64    ┆ f64       ┆ f64      │

╞════════╪═══════════════════╪════════╪═══════════╪══════════╡

│ AAPL   ┆ Apple             ┆ 229.9  ┆ 237.23    ┆ 164.08   │

│ NVDA   ┆ NVIDIA            ┆ 138.93 ┆ 140.76    ┆ 39.23    │

│ MSFT   ┆ Microsoft         ┆ 420.56 ┆ 468.35    ┆ 324.39   │

│ GOOG   ┆ Alphabet (Google) ┆ 166.41 ┆ 193.31    ┆ 121.46   │

│ AMZN   ┆ Amazon            ┆ 188.4  ┆ 201.2     ┆ 118.35   │

└────────┴───────────────────┴────────┴───────────┴──────────┘

> (select df (~> (col 'float64) (exclude #rx"^day_")))

shape: (5, 3)

┌────────┬───────────┬──────────┐

│ price  ┆ year_high ┆ year_low │

│ ---    ┆ ---       ┆ ---      │

│ f64    ┆ f64       ┆ f64      │

╞════════╪═══════════╪══════════╡

│ 229.9  ┆ 237.23    ┆ 164.08   │

│ 138.93 ┆ 140.76    ┆ 39.23    │

│ 420.56 ┆ 468.35    ┆ 324.39   │

│ 166.41 ┆ 193.31    ┆ 121.46   │

│ 188.4  ┆ 201.2     ┆ 118.35   │

└────────┴───────────┴──────────┘

API gap: no exclude by dtype.

1.3.1.4 Column renaming🔗ℹ

An expanded expression keeps each matched column’s name, so two expressions over the same column collide.

> (define gbp-usd-rate 1.31)
> (select df
          (/ (col "price") gbp-usd-rate)
          (/ (col "price") eur-usd-rate))

lazyframe-collect: failed to collect the query: duplicate:

projections contained duplicate output name 'price'. It's

possible that multiple expressions are returning the same

default column name. If this is the case, try renaming the

columns with `.alias("new_name")` to avoid duplicate column

names.

Resolved plan until failure:

---> FAILED HERE RESOLVING 'select' <---

DF ["ticker", "company_name", "price", "day_high", ...];

PROJECT */7 COLUMNS: 'select'

1.3.1.4.1 Renaming a single column with alias🔗ℹ
> (select df
          (~> (col "price") (/ gbp-usd-rate) (alias "price (GBP)"))
          (~> (col "price") (/ eur-usd-rate) (alias "price (EUR)")))

shape: (5, 2)

┌─────────────┬─────────────┐

│ price (GBP) ┆ price (EUR) │

│ ---         ┆ ---         │

│ f64         ┆ f64         │

╞═════════════╪═════════════╡

│ 175.496183  ┆ 210.917431  │

│ 106.053435  ┆ 127.458716  │

│ 321.038168  ┆ 385.834862  │

│ 127.030534  ┆ 152.669725  │

│ 143.816794  ┆ 172.844037  │

└─────────────┴─────────────┘

1.3.1.4.2 Prefixing and suffixing column names🔗ℹ

API gap: no name.prefix / name.suffix (#51). Alias each column, taking the names from the frame.

> (select df
          (for/list ([name (column-names df)]
                     #:when (regexp-match? #rx"^year_" name))
            (~> (col name) (/ eur-usd-rate) (alias (string-append "in_eur_" name))))
          (for/list ([name '("day_high" "day_low")])
            (~> (col name) (/ gbp-usd-rate) (alias (string-append name "_gbp")))))

shape: (5, 4)

┌──────────────────┬─────────────────┬──────────────┬─────────────┐

│ in_eur_year_high ┆ in_eur_year_low ┆ day_high_gbp ┆ day_low_gbp │

│ ---              ┆ ---             ┆ ---          ┆ ---         │

│ f64              ┆ f64             ┆ f64          ┆ f64         │

╞══════════════════╪═════════════════╪══════════════╪═════════════╡

│ 217.642202       ┆ 150.53211       ┆ 176.572519   ┆ 174.503817  │

│ 129.137615       ┆ 35.990826       ┆ 106.564885   ┆ 104.045802  │

│ 429.678899       ┆ 297.605505      ┆ 323.694656   ┆ 318.717557  │

│ 177.348624       ┆ 111.431193      ┆ 127.954198   ┆ 125.78626   │

│ 184.587156       ┆ 108.577982      ┆ 144.908397   ┆ 143.847328  │

└──────────────────┴─────────────────┴──────────────┴─────────────┘

1.3.1.4.3 Dynamic name replacement🔗ℹ

API gap: no name.map; the same loop applies any Racket function to the names.

> (select df (for/list ([name (column-names df)])
               (alias (col name) (string-upcase name))))

shape: (5, 7)

┌────────┬───────────────────┬────────┬──────────┬─────────┬───────────┬──────────┐

│ TICKER ┆ COMPANY_NAME      ┆ PRICE  ┆ DAY_HIGH ┆ DAY_LOW ┆ YEAR_HIGH ┆ YEAR_LOW │

│ ---    ┆ ---               ┆ ---    ┆ ---      ┆ ---     ┆ ---       ┆ ---      │

│ str    ┆ str               ┆ f64    ┆ f64      ┆ f64     ┆ f64       ┆ f64      │

╞════════╪═══════════════════╪════════╪══════════╪═════════╪═══════════╪══════════╡

│ AAPL   ┆ Apple             ┆ 229.9  ┆ 231.31   ┆ 228.6   ┆ 237.23    ┆ 164.08   │

│ NVDA   ┆ NVIDIA            ┆ 138.93 ┆ 139.6    ┆ 136.3   ┆ 140.76    ┆ 39.23    │

│ MSFT   ┆ Microsoft         ┆ 420.56 ┆ 424.04   ┆ 417.52  ┆ 468.35    ┆ 324.39   │

│ GOOG   ┆ Alphabet (Google) ┆ 166.41 ┆ 167.62   ┆ 164.78  ┆ 193.31    ┆ 121.46   │

│ AMZN   ┆ Amazon            ┆ 188.4  ┆ 189.83   ┆ 188.44  ┆ 201.2     ┆ 118.35   │

└────────┴───────────────────┴────────┴──────────┴─────────┴───────────┴──────────┘

1.3.1.5 Programmatically generating expressions🔗ℹ

Build the expressions first, then hand them to one context.

> (define (amplitude-expressions time-periods)
    (for/list ([tp (in-list time-periods)])
      (~> (col (string-append tp "_high"))
          (- (col (string-append tp "_low")))
          (alias (string-append tp "_amplitude")))))
> (~> df (with-columns (amplitude-expressions '("day" "year"))))

shape: (5, 9)

┌────────┬──────────────┬────────┬──────────┬───┬───────────┬──────────┬─────────────┬─────────────┐

│ ticker ┆ company_name ┆ price  ┆ day_high ┆ … ┆ year_high ┆ year_low ┆ day_amplitu ┆ year_amplit │

│ ---    ┆ ---          ┆ ---    ┆ ---      ┆   ┆ ---       ┆ ---      ┆ de          ┆ ude         │

│ str    ┆ str          ┆ f64    ┆ f64      ┆   ┆ f64       ┆ f64      ┆ ---         ┆ ---         │

│        ┆              ┆        ┆          ┆   ┆           ┆          ┆ f64         ┆ f64         │

╞════════╪══════════════╪════════╪══════════╪═══╪═══════════╪══════════╪═════════════╪═════════════╡

│ AAPL   ┆ Apple        ┆ 229.9  ┆ 231.31   ┆ … ┆ 237.23    ┆ 164.08   ┆ 2.71        ┆ 73.15       │

│ NVDA   ┆ NVIDIA       ┆ 138.93 ┆ 139.6    ┆ … ┆ 140.76    ┆ 39.23    ┆ 3.3         ┆ 101.53      │

│ MSFT   ┆ Microsoft    ┆ 420.56 ┆ 424.04   ┆ … ┆ 468.35    ┆ 324.39   ┆ 6.52        ┆ 143.96      │

│ GOOG   ┆ Alphabet     ┆ 166.41 ┆ 167.62   ┆ … ┆ 193.31    ┆ 121.46   ┆ 2.84        ┆ 71.85       │

│        ┆ (Google)     ┆        ┆          ┆   ┆           ┆          ┆             ┆             │

│ AMZN   ┆ Amazon       ┆ 188.4  ┆ 189.83   ┆ … ┆ 201.2     ┆ 118.35   ┆ 1.39        ┆ 82.85       │

└────────┴──────────────┴────────┴──────────┴───┴───────────┴──────────┴─────────────┴─────────────┘

1.3.1.6 More flexible column selections🔗ℹ

API gap: no polars.selectors, and selectors do not compose with set operations yet (#50). A union of disjoint selections is a list of specs; anything else is a Racket filter over the names.

> (select df (col 'string) (col #rx"_high$"))

shape: (5, 4)

┌────────┬───────────────────┬──────────┬───────────┐

│ ticker ┆ company_name      ┆ day_high ┆ year_high │

│ ---    ┆ ---               ┆ ---      ┆ ---       │

│ str    ┆ str               ┆ f64      ┆ f64       │

╞════════╪═══════════════════╪══════════╪═══════════╡

│ AAPL   ┆ Apple             ┆ 231.31   ┆ 237.23    │

│ NVDA   ┆ NVIDIA            ┆ 139.6    ┆ 140.76    │

│ MSFT   ┆ Microsoft         ┆ 424.04   ┆ 468.35    │

│ GOOG   ┆ Alphabet (Google) ┆ 167.62   ┆ 193.31    │

│ AMZN   ┆ Amazon            ┆ 189.83   ┆ 201.2     │

└────────┴───────────────────┴──────────┴───────────┘

> (select df (filter (lambda (name)
                       (and (regexp-match? #rx"_" name)
                            (~> (ref df name) dtype (eq? 'string) not)))
                     (column-names df)))

shape: (5, 4)

┌──────────┬─────────┬───────────┬──────────┐

│ day_high ┆ day_low ┆ year_high ┆ year_low │

│ ---      ┆ ---     ┆ ---       ┆ ---      │

│ f64      ┆ f64     ┆ f64       ┆ f64      │

╞══════════╪═════════╪═══════════╪══════════╡

│ 231.31   ┆ 228.6   ┆ 237.23    ┆ 164.08   │

│ 139.6    ┆ 136.3   ┆ 140.76    ┆ 39.23    │

│ 424.04   ┆ 417.52  ┆ 468.35    ┆ 324.39   │

│ 167.62   ┆ 164.78  ┆ 193.31    ┆ 121.46   │

│ 189.83   ┆ 188.44  ┆ 201.2     ┆ 118.35   │

└──────────┴─────────┴───────────┴──────────┘

1.3.1.6.1 Debugging selectors🔗ℹ

multi-column-expr? plays cs.is_selector, an expression prints as its plan, and selecting from the frame lists what a selector matches.

> (define people
    (dataframe (list (series '("Anna" "Bob") #:name "name")
                     (series '(#t #f) #:name "has_partner")
                     (series '(#f #f) #:name "has_kids")
                     (series '(#t #f) #:name "has_tattoos")
                     (series '(#t #t) #:name "is_alive"))))
> (select people (not (col #rx"^has_")))

shape: (2, 3)

┌─────────────┬──────────┬─────────────┐

│ has_partner ┆ has_kids ┆ has_tattoos │

│ ---         ┆ ---      ┆ ---         │

│ bool        ┆ bool     ┆ bool        │

╞═════════════╪══════════╪═════════════╡

│ false       ┆ true     ┆ false       │

│ true        ┆ true     ┆ true        │

└─────────────┴──────────┴─────────────┘

> (~> (col #rx"^has_") not multi-column-expr?)

#t

> (not (col #rx"^has_"))

cs.matches("^(?s).*(?:^has_).*$").not()

> (~> people (select (col #rx"^has_")) column-names)

'("has_partner" "has_kids" "has_tattoos")

meta-root-names and meta-output-name read the plan without a frame: a named expression reports its inputs and output, and a selector, by regexp or by dtype, nothing.

> (for/list ([e (amplitude-expressions '("day" "year"))])
    (list (meta-output-name e) (meta-root-names e)))

'(("day_amplitude" ("day_high" "day_low"))

  ("year_amplitude" ("year_high" "year_low")))

> (meta-root-names (col #rx"^has_"))

'()

> (meta-output-name (col #rx"^has_"))

expr-meta-output-name: cannot determine the output name of

cs.matches("^(?s).*(?:^has_).*$")

> (meta-root-names (col 'bool))

'()

> (meta-output-name (col 'bool))

expr-meta-output-name: cannot determine the output name of

cs.boolean()