1.3.1 Expression expansion
One expression can stand for several columns; it expands to one expression per matched column when a context runs it. The stock prices below are upstream’s.
> (define df (dataframe (list (series '("AAPL" "NVDA" "MSFT" "GOOG" "AMZN") #:name "ticker") (series '("Apple" "NVIDIA" "Microsoft" "Alphabet (Google)" "Amazon") #:name "company_name") (series '(229.9 138.93 420.56 166.41 188.4) #:name "price") (series '(231.31 139.6 424.04 167.62 189.83) #:name "day_high") (series '(228.6 136.3 417.52 164.78 188.44) #:name "day_low") (series '(237.23 140.76 468.35 193.31 201.2) #:name "year_high") (series '(164.08 39.23 324.39 121.46 118.35) #:name "year_low")))) > df
shape: (5, 7)
┌────────┬───────────────────┬────────┬──────────┬─────────┬───────────┬──────────┐
│ ticker ┆ company_name ┆ price ┆ day_high ┆ day_low ┆ year_high ┆ year_low │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ str ┆ f64 ┆ f64 ┆ f64 ┆ f64 ┆ f64 │
╞════════╪═══════════════════╪════════╪══════════╪═════════╪═══════════╪══════════╡
│ AAPL ┆ Apple ┆ 229.9 ┆ 231.31 ┆ 228.6 ┆ 237.23 ┆ 164.08 │
│ NVDA ┆ NVIDIA ┆ 138.93 ┆ 139.6 ┆ 136.3 ┆ 140.76 ┆ 39.23 │
│ MSFT ┆ Microsoft ┆ 420.56 ┆ 424.04 ┆ 417.52 ┆ 468.35 ┆ 324.39 │
│ GOOG ┆ Alphabet (Google) ┆ 166.41 ┆ 167.62 ┆ 164.78 ┆ 193.31 ┆ 121.46 │
│ AMZN ┆ Amazon ┆ 188.4 ┆ 189.83 ┆ 188.44 ┆ 201.2 ┆ 118.35 │
└────────┴───────────────────┴────────┴──────────┴─────────┴───────────┴──────────┘
1.3.1.1 Function col
1.3.1.1.1 Explicit expansion by column name
API gap: col takes one name, so there is no pl.col("price", "day_high", ...). A list of expressions is spliced into the context, so build one per name.
> (define eur-usd-rate 1.09)
> (~> df (with-columns (for/list ([name '("price" "day_high" "day_low" "year_high" "year_low")]) (~> (col name) (/ eur-usd-rate) (round #:decimals 2)))))
shape: (5, 7)
┌────────┬───────────────────┬────────┬──────────┬─────────┬───────────┬──────────┐
│ ticker ┆ company_name ┆ price ┆ day_high ┆ day_low ┆ year_high ┆ year_low │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ str ┆ f64 ┆ f64 ┆ f64 ┆ f64 ┆ f64 │
╞════════╪═══════════════════╪════════╪══════════╪═════════╪═══════════╪══════════╡
│ AAPL ┆ Apple ┆ 210.92 ┆ 212.21 ┆ 209.72 ┆ 217.64 ┆ 150.53 │
│ NVDA ┆ NVIDIA ┆ 127.46 ┆ 128.07 ┆ 125.05 ┆ 129.14 ┆ 35.99 │
│ MSFT ┆ Microsoft ┆ 385.83 ┆ 389.03 ┆ 383.05 ┆ 429.68 ┆ 297.61 │
│ GOOG ┆ Alphabet (Google) ┆ 152.67 ┆ 153.78 ┆ 151.17 ┆ 177.35 ┆ 111.43 │
│ AMZN ┆ Amazon ┆ 172.84 ┆ 174.16 ┆ 172.88 ┆ 184.59 ┆ 108.58 │
└────────┴───────────────────┴────────┴──────────┴─────────┴───────────┴──────────┘
1.3.1.1.2 Expansion by data type
A dtype selects every column of that dtype; 'f64 and 'float64 are the same selector.
> (~> df (with-columns (~> (col 'float64) (/ eur-usd-rate) (round #:decimals 2))))
shape: (5, 7)
┌────────┬───────────────────┬────────┬──────────┬─────────┬───────────┬──────────┐
│ ticker ┆ company_name ┆ price ┆ day_high ┆ day_low ┆ year_high ┆ year_low │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ str ┆ f64 ┆ f64 ┆ f64 ┆ f64 ┆ f64 │
╞════════╪═══════════════════╪════════╪══════════╪═════════╪═══════════╪══════════╡
│ AAPL ┆ Apple ┆ 210.92 ┆ 212.21 ┆ 209.72 ┆ 217.64 ┆ 150.53 │
│ NVDA ┆ NVIDIA ┆ 127.46 ┆ 128.07 ┆ 125.05 ┆ 129.14 ┆ 35.99 │
│ MSFT ┆ Microsoft ┆ 385.83 ┆ 389.03 ┆ 383.05 ┆ 429.68 ┆ 297.61 │
│ GOOG ┆ Alphabet (Google) ┆ 152.67 ┆ 153.78 ┆ 151.17 ┆ 177.35 ┆ 111.43 │
│ AMZN ┆ Amazon ┆ 172.84 ┆ 174.16 ┆ 172.88 ┆ 184.59 ┆ 108.58 │
└────────┴───────────────────┴────────┴──────────┴─────────┴───────────┴──────────┘
API gap: col takes one dtype, and selectors do not compose yet (#50), so pl.col(pl.Float32, pl.Float64) is one expression per dtype. A dtype no column has matches nothing.
> (~> df (with-columns (for/list ([type '(float32 float64)]) (~> (col type) (/ eur-usd-rate) (round #:decimals 2)))))
shape: (5, 7)
┌────────┬───────────────────┬────────┬──────────┬─────────┬───────────┬──────────┐
│ ticker ┆ company_name ┆ price ┆ day_high ┆ day_low ┆ year_high ┆ year_low │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ str ┆ f64 ┆ f64 ┆ f64 ┆ f64 ┆ f64 │
╞════════╪═══════════════════╪════════╪══════════╪═════════╪═══════════╪══════════╡
│ AAPL ┆ Apple ┆ 210.92 ┆ 212.21 ┆ 209.72 ┆ 217.64 ┆ 150.53 │
│ NVDA ┆ NVIDIA ┆ 127.46 ┆ 128.07 ┆ 125.05 ┆ 129.14 ┆ 35.99 │
│ MSFT ┆ Microsoft ┆ 385.83 ┆ 389.03 ┆ 383.05 ┆ 429.68 ┆ 297.61 │
│ GOOG ┆ Alphabet (Google) ┆ 152.67 ┆ 153.78 ┆ 151.17 ┆ 177.35 ┆ 111.43 │
│ AMZN ┆ Amazon ┆ 172.84 ┆ 174.16 ┆ 172.88 ┆ 184.59 ┆ 108.58 │
└────────┴───────────────────┴────────┴──────────┴─────────┴───────────┴──────────┘
1.3.1.1.3 Expansion by pattern matching
A string of the form ^...$ is a Polars regex, as in Python. The context takes several specs, where Python’s col takes several names.
> (select df "ticker" (col "^.*_high$") (col "^.*_low$"))
shape: (5, 5)
┌────────┬──────────┬───────────┬─────────┬──────────┐
│ ticker ┆ day_high ┆ year_high ┆ day_low ┆ year_low │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ f64 ┆ f64 ┆ f64 ┆ f64 │
╞════════╪══════════╪═══════════╪═════════╪══════════╡
│ AAPL ┆ 231.31 ┆ 237.23 ┆ 228.6 ┆ 164.08 │
│ NVDA ┆ 139.6 ┆ 140.76 ┆ 136.3 ┆ 39.23 │
│ MSFT ┆ 424.04 ┆ 468.35 ┆ 417.52 ┆ 324.39 │
│ GOOG ┆ 167.62 ┆ 193.31 ┆ 164.78 ┆ 121.46 │
│ AMZN ┆ 189.83 ┆ 201.2 ┆ 188.44 ┆ 118.35 │
└────────┴──────────┴───────────┴─────────┴──────────┘
A Racket regexp is also accepted, and keeps its Racket meaning: it selects exactly the names regexp-match? accepts, so it is unanchored unless it says otherwise.
> (select df "ticker" (col #rx"_(high|low)$"))
shape: (5, 5)
┌────────┬──────────┬─────────┬───────────┬──────────┐
│ ticker ┆ day_high ┆ day_low ┆ year_high ┆ year_low │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ f64 ┆ f64 ┆ f64 ┆ f64 │
╞════════╪══════════╪═════════╪═══════════╪══════════╡
│ AAPL ┆ 231.31 ┆ 228.6 ┆ 237.23 ┆ 164.08 │
│ NVDA ┆ 139.6 ┆ 136.3 ┆ 140.76 ┆ 39.23 │
│ MSFT ┆ 424.04 ┆ 417.52 ┆ 468.35 ┆ 324.39 │
│ GOOG ┆ 167.62 ┆ 164.78 ┆ 193.31 ┆ 121.46 │
│ AMZN ┆ 189.83 ┆ 188.44 ┆ 201.2 ┆ 118.35 │
└────────┴──────────┴─────────┴───────────┴──────────┘
> (select df (col #px"^\\w+_high$"))
shape: (5, 2)
┌──────────┬───────────┐
│ day_high ┆ year_high │
│ --- ┆ --- │
│ f64 ┆ f64 │
╞══════════╪═══════════╡
│ 231.31 ┆ 237.23 │
│ 139.6 ┆ 140.76 │
│ 424.04 ┆ 468.35 │
│ 167.62 ┆ 193.31 │
│ 189.83 ┆ 201.2 │
└──────────┴───────────┘
Polars’ engine has no lookaround or backreferences, so such a regexp fails when the query runs. Match the names in Racket instead.
> (define not-day #px"^(?!day_).*_(high|low)$") > (select df (col not-day)) lazyframe-collect: failed to collect the query: invalid
regex in selector '^(?s).*(?:^(?!day_).*_(high|low)$).*$'
Resolved plan until failure:
---> FAILED HERE RESOLVING 'select' <---
DF ["ticker", "company_name", "price", "day_high", ...];
PROJECT */7 COLUMNS: 'select'
> (select df (filter (lambda (name) (regexp-match? not-day name)) (column-names df)))
shape: (5, 2)
┌───────────┬──────────┐
│ year_high ┆ year_low │
│ --- ┆ --- │
│ f64 ┆ f64 │
╞═══════════╪══════════╡
│ 237.23 ┆ 164.08 │
│ 140.76 ┆ 39.23 │
│ 468.35 ┆ 324.39 │
│ 193.31 ┆ 121.46 │
│ 201.2 ┆ 118.35 │
└───────────┴──────────┘
1.3.1.1.4 Arguments cannot be of mixed types
Each col is one name, dtype or pattern, so the mixture is written as separate specs.
> (select df "ticker" (col 'float64))
shape: (5, 6)
┌────────┬────────┬──────────┬─────────┬───────────┬──────────┐
│ ticker ┆ price ┆ day_high ┆ day_low ┆ year_high ┆ year_low │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ f64 ┆ f64 ┆ f64 ┆ f64 ┆ f64 │
╞════════╪════════╪══════════╪═════════╪═══════════╪══════════╡
│ AAPL ┆ 229.9 ┆ 231.31 ┆ 228.6 ┆ 237.23 ┆ 164.08 │
│ NVDA ┆ 138.93 ┆ 139.6 ┆ 136.3 ┆ 140.76 ┆ 39.23 │
│ MSFT ┆ 420.56 ┆ 424.04 ┆ 417.52 ┆ 468.35 ┆ 324.39 │
│ GOOG ┆ 166.41 ┆ 167.62 ┆ 164.78 ┆ 193.31 ┆ 121.46 │
│ AMZN ┆ 188.4 ┆ 189.83 ┆ 188.44 ┆ 201.2 ┆ 118.35 │
└────────┴────────┴──────────┴─────────┴───────────┴──────────┘
1.3.1.2 Selecting all columns
> (select df (all))
shape: (5, 7)
┌────────┬───────────────────┬────────┬──────────┬─────────┬───────────┬──────────┐
│ ticker ┆ company_name ┆ price ┆ day_high ┆ day_low ┆ year_high ┆ year_low │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ str ┆ f64 ┆ f64 ┆ f64 ┆ f64 ┆ f64 │
╞════════╪═══════════════════╪════════╪══════════╪═════════╪═══════════╪══════════╡
│ AAPL ┆ Apple ┆ 229.9 ┆ 231.31 ┆ 228.6 ┆ 237.23 ┆ 164.08 │
│ NVDA ┆ NVIDIA ┆ 138.93 ┆ 139.6 ┆ 136.3 ┆ 140.76 ┆ 39.23 │
│ MSFT ┆ Microsoft ┆ 420.56 ┆ 424.04 ┆ 417.52 ┆ 468.35 ┆ 324.39 │
│ GOOG ┆ Alphabet (Google) ┆ 166.41 ┆ 167.62 ┆ 164.78 ┆ 193.31 ┆ 121.46 │
│ AMZN ┆ Amazon ┆ 188.4 ┆ 189.83 ┆ 188.44 ┆ 201.2 ┆ 118.35 │
└────────┴───────────────────┴────────┴──────────┴─────────┴───────────┴──────────┘
API gap: no DataFrame.equals.
1.3.1.3 Excluding columns
exclude takes names, Polars regex strings and Racket regexps, read as col reads them, and applies to any multi-column expression.
> (select df (exclude (all) "^day_.*$"))
shape: (5, 5)
┌────────┬───────────────────┬────────┬───────────┬──────────┐
│ ticker ┆ company_name ┆ price ┆ year_high ┆ year_low │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ str ┆ f64 ┆ f64 ┆ f64 │
╞════════╪═══════════════════╪════════╪═══════════╪══════════╡
│ AAPL ┆ Apple ┆ 229.9 ┆ 237.23 ┆ 164.08 │
│ NVDA ┆ NVIDIA ┆ 138.93 ┆ 140.76 ┆ 39.23 │
│ MSFT ┆ Microsoft ┆ 420.56 ┆ 468.35 ┆ 324.39 │
│ GOOG ┆ Alphabet (Google) ┆ 166.41 ┆ 193.31 ┆ 121.46 │
│ AMZN ┆ Amazon ┆ 188.4 ┆ 201.2 ┆ 118.35 │
└────────┴───────────────────┴────────┴───────────┴──────────┘
> (select df (~> (col 'float64) (exclude #rx"^day_")))
shape: (5, 3)
┌────────┬───────────┬──────────┐
│ price ┆ year_high ┆ year_low │
│ --- ┆ --- ┆ --- │
│ f64 ┆ f64 ┆ f64 │
╞════════╪═══════════╪══════════╡
│ 229.9 ┆ 237.23 ┆ 164.08 │
│ 138.93 ┆ 140.76 ┆ 39.23 │
│ 420.56 ┆ 468.35 ┆ 324.39 │
│ 166.41 ┆ 193.31 ┆ 121.46 │
│ 188.4 ┆ 201.2 ┆ 118.35 │
└────────┴───────────┴──────────┘
API gap: no exclude by dtype.
1.3.1.4 Column renaming
An expanded expression keeps each matched column’s name, so two expressions over the same column collide.
> (define gbp-usd-rate 1.31)
> (select df (/ (col "price") gbp-usd-rate) (/ (col "price") eur-usd-rate)) lazyframe-collect: failed to collect the query: duplicate:
projections contained duplicate output name 'price'. It's
possible that multiple expressions are returning the same
default column name. If this is the case, try renaming the
columns with `.alias("new_name")` to avoid duplicate column
names.
Resolved plan until failure:
---> FAILED HERE RESOLVING 'select' <---
DF ["ticker", "company_name", "price", "day_high", ...];
PROJECT */7 COLUMNS: 'select'
1.3.1.4.1 Renaming a single column with alias
> (select df (~> (col "price") (/ gbp-usd-rate) (alias "price (GBP)")) (~> (col "price") (/ eur-usd-rate) (alias "price (EUR)")))
shape: (5, 2)
┌─────────────┬─────────────┐
│ price (GBP) ┆ price (EUR) │
│ --- ┆ --- │
│ f64 ┆ f64 │
╞═════════════╪═════════════╡
│ 175.496183 ┆ 210.917431 │
│ 106.053435 ┆ 127.458716 │
│ 321.038168 ┆ 385.834862 │
│ 127.030534 ┆ 152.669725 │
│ 143.816794 ┆ 172.844037 │
└─────────────┴─────────────┘
1.3.1.4.2 Prefixing and suffixing column names
API gap: no name.prefix / name.suffix (#51). Alias each column, taking the names from the frame.
> (select df (for/list ([name (column-names df)] #:when (regexp-match? #rx"^year_" name)) (~> (col name) (/ eur-usd-rate) (alias (string-append "in_eur_" name)))) (for/list ([name '("day_high" "day_low")]) (~> (col name) (/ gbp-usd-rate) (alias (string-append name "_gbp")))))
shape: (5, 4)
┌──────────────────┬─────────────────┬──────────────┬─────────────┐
│ in_eur_year_high ┆ in_eur_year_low ┆ day_high_gbp ┆ day_low_gbp │
│ --- ┆ --- ┆ --- ┆ --- │
│ f64 ┆ f64 ┆ f64 ┆ f64 │
╞══════════════════╪═════════════════╪══════════════╪═════════════╡
│ 217.642202 ┆ 150.53211 ┆ 176.572519 ┆ 174.503817 │
│ 129.137615 ┆ 35.990826 ┆ 106.564885 ┆ 104.045802 │
│ 429.678899 ┆ 297.605505 ┆ 323.694656 ┆ 318.717557 │
│ 177.348624 ┆ 111.431193 ┆ 127.954198 ┆ 125.78626 │
│ 184.587156 ┆ 108.577982 ┆ 144.908397 ┆ 143.847328 │
└──────────────────┴─────────────────┴──────────────┴─────────────┘
1.3.1.4.3 Dynamic name replacement
API gap: no name.map; the same loop applies any Racket function to the names.
> (select df (for/list ([name (column-names df)]) (alias (col name) (string-upcase name))))
shape: (5, 7)
┌────────┬───────────────────┬────────┬──────────┬─────────┬───────────┬──────────┐
│ TICKER ┆ COMPANY_NAME ┆ PRICE ┆ DAY_HIGH ┆ DAY_LOW ┆ YEAR_HIGH ┆ YEAR_LOW │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ str ┆ f64 ┆ f64 ┆ f64 ┆ f64 ┆ f64 │
╞════════╪═══════════════════╪════════╪══════════╪═════════╪═══════════╪══════════╡
│ AAPL ┆ Apple ┆ 229.9 ┆ 231.31 ┆ 228.6 ┆ 237.23 ┆ 164.08 │
│ NVDA ┆ NVIDIA ┆ 138.93 ┆ 139.6 ┆ 136.3 ┆ 140.76 ┆ 39.23 │
│ MSFT ┆ Microsoft ┆ 420.56 ┆ 424.04 ┆ 417.52 ┆ 468.35 ┆ 324.39 │
│ GOOG ┆ Alphabet (Google) ┆ 166.41 ┆ 167.62 ┆ 164.78 ┆ 193.31 ┆ 121.46 │
│ AMZN ┆ Amazon ┆ 188.4 ┆ 189.83 ┆ 188.44 ┆ 201.2 ┆ 118.35 │
└────────┴───────────────────┴────────┴──────────┴─────────┴───────────┴──────────┘
1.3.1.5 Programmatically generating expressions
Build the expressions first, then hand them to one context.
> (define (amplitude-expressions time-periods) (for/list ([tp (in-list time-periods)]) (~> (col (string-append tp "_high")) (- (col (string-append tp "_low"))) (alias (string-append tp "_amplitude"))))) > (~> df (with-columns (amplitude-expressions '("day" "year"))))
shape: (5, 9)
┌────────┬──────────────┬────────┬──────────┬───┬───────────┬──────────┬─────────────┬─────────────┐
│ ticker ┆ company_name ┆ price ┆ day_high ┆ … ┆ year_high ┆ year_low ┆ day_amplitu ┆ year_amplit │
│ --- ┆ --- ┆ --- ┆ --- ┆ ┆ --- ┆ --- ┆ de ┆ ude │
│ str ┆ str ┆ f64 ┆ f64 ┆ ┆ f64 ┆ f64 ┆ --- ┆ --- │
│ ┆ ┆ ┆ ┆ ┆ ┆ ┆ f64 ┆ f64 │
╞════════╪══════════════╪════════╪══════════╪═══╪═══════════╪══════════╪═════════════╪═════════════╡
│ AAPL ┆ Apple ┆ 229.9 ┆ 231.31 ┆ … ┆ 237.23 ┆ 164.08 ┆ 2.71 ┆ 73.15 │
│ NVDA ┆ NVIDIA ┆ 138.93 ┆ 139.6 ┆ … ┆ 140.76 ┆ 39.23 ┆ 3.3 ┆ 101.53 │
│ MSFT ┆ Microsoft ┆ 420.56 ┆ 424.04 ┆ … ┆ 468.35 ┆ 324.39 ┆ 6.52 ┆ 143.96 │
│ GOOG ┆ Alphabet ┆ 166.41 ┆ 167.62 ┆ … ┆ 193.31 ┆ 121.46 ┆ 2.84 ┆ 71.85 │
│ ┆ (Google) ┆ ┆ ┆ ┆ ┆ ┆ ┆ │
│ AMZN ┆ Amazon ┆ 188.4 ┆ 189.83 ┆ … ┆ 201.2 ┆ 118.35 ┆ 1.39 ┆ 82.85 │
└────────┴──────────────┴────────┴──────────┴───┴───────────┴──────────┴─────────────┴─────────────┘
1.3.1.6 More flexible column selections
API gap: no polars.selectors, and selectors do not compose with set operations yet (#50). A union of disjoint selections is a list of specs; anything else is a Racket filter over the names.
> (select df (col 'string) (col #rx"_high$"))
shape: (5, 4)
┌────────┬───────────────────┬──────────┬───────────┐
│ ticker ┆ company_name ┆ day_high ┆ year_high │
│ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ str ┆ f64 ┆ f64 │
╞════════╪═══════════════════╪══════════╪═══════════╡
│ AAPL ┆ Apple ┆ 231.31 ┆ 237.23 │
│ NVDA ┆ NVIDIA ┆ 139.6 ┆ 140.76 │
│ MSFT ┆ Microsoft ┆ 424.04 ┆ 468.35 │
│ GOOG ┆ Alphabet (Google) ┆ 167.62 ┆ 193.31 │
│ AMZN ┆ Amazon ┆ 189.83 ┆ 201.2 │
└────────┴───────────────────┴──────────┴───────────┘
> (select df (filter (lambda (name) (and (regexp-match? #rx"_" name) (~> (ref df name) dtype (eq? 'string) not))) (column-names df)))
shape: (5, 4)
┌──────────┬─────────┬───────────┬──────────┐
│ day_high ┆ day_low ┆ year_high ┆ year_low │
│ --- ┆ --- ┆ --- ┆ --- │
│ f64 ┆ f64 ┆ f64 ┆ f64 │
╞══════════╪═════════╪═══════════╪══════════╡
│ 231.31 ┆ 228.6 ┆ 237.23 ┆ 164.08 │
│ 139.6 ┆ 136.3 ┆ 140.76 ┆ 39.23 │
│ 424.04 ┆ 417.52 ┆ 468.35 ┆ 324.39 │
│ 167.62 ┆ 164.78 ┆ 193.31 ┆ 121.46 │
│ 189.83 ┆ 188.44 ┆ 201.2 ┆ 118.35 │
└──────────┴─────────┴───────────┴──────────┘
1.3.1.6.1 Debugging selectors
multi-column-expr? plays cs.is_selector, an expression prints as its plan, and selecting from the frame lists what a selector matches.
> (define people (dataframe (list (series '("Anna" "Bob") #:name "name") (series '(#t #f) #:name "has_partner") (series '(#f #f) #:name "has_kids") (series '(#t #f) #:name "has_tattoos") (series '(#t #t) #:name "is_alive")))) > (select people (not (col #rx"^has_")))
shape: (2, 3)
┌─────────────┬──────────┬─────────────┐
│ has_partner ┆ has_kids ┆ has_tattoos │
│ --- ┆ --- ┆ --- │
│ bool ┆ bool ┆ bool │
╞═════════════╪══════════╪═════════════╡
│ false ┆ true ┆ false │
│ true ┆ true ┆ true │
└─────────────┴──────────┴─────────────┘
> (~> (col #rx"^has_") not multi-column-expr?) #t
> (not (col #rx"^has_")) cs.matches("^(?s).*(?:^has_).*$").not()
> (~> people (select (col #rx"^has_")) column-names) '("has_partner" "has_kids" "has_tattoos")
meta-root-names and meta-output-name read the plan without a frame: a named expression reports its inputs and output, and a selector, by regexp or by dtype, nothing.
> (for/list ([e (amplitude-expressions '("day" "year"))]) (list (meta-output-name e) (meta-root-names e)))
'(("day_amplitude" ("day_high" "day_low"))
("year_amplitude" ("year_high" "year_low")))
> (meta-root-names (col #rx"^has_")) '()
> (meta-output-name (col #rx"^has_")) expr-meta-output-name: cannot determine the output name of
cs.matches("^(?s).*(?:^has_).*$")
> (meta-root-names (col 'bool)) '()
> (meta-output-name (col 'bool)) expr-meta-output-name: cannot determine the output name of
cs.boolean()