1.2.1 Data types and structures
1.2.1.1 Series
A series is a typed, one-dimensional column. series infers a dtype or takes #:dtype; polars-null marks missing entries.
> (series '(1 2 3 4 5) #:name "ints")
shape: (5,)
Series: 'ints' [i64]
[
1
2
3
4
5
]
> (define s1 (series '(1 2 3 4 5) #:name "ints")) > (define s2 (series '(1 2 3 4 5) #:name "uints" #:dtype 'u64)) > (list (dtype s1) (dtype s2)) '(int64 uint64)
1.2.1.2 Dataframe
A dataframe is a collection of equal-length, uniquely named series.
> (define df (~> (dataframe (list (series '("Alice Archer" "Ben Brown" "Chloe Cooper" "Daniel Donovan") #:name "name") (series (list (datetime 1997 1 10) (datetime 1985 2 15) (datetime 1983 3 22) (datetime 1981 4 30)) #:name "birthdate") (series '(57.9 72.5 53.6 83.1) #:name "weight") (series '(1.56 1.77 1.65 1.75) #:name "height"))) (with-columns (cast "birthdate" 'date)))) > df
shape: (4, 4)
┌────────────────┬────────────┬────────┬────────┐
│ name ┆ birthdate ┆ weight ┆ height │
│ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ date ┆ f64 ┆ f64 │
╞════════════════╪════════════╪════════╪════════╡
│ Alice Archer ┆ 1997-01-10 ┆ 57.9 ┆ 1.56 │
│ Ben Brown ┆ 1985-02-15 ┆ 72.5 ┆ 1.77 │
│ Chloe Cooper ┆ 1983-03-22 ┆ 53.6 ┆ 1.65 │
│ Daniel Donovan ┆ 1981-04-30 ┆ 83.1 ┆ 1.75 │
└────────────────┴────────────┴────────┴────────┘
1.2.1.2.1 Inspecting a dataframe
> (head df 3)
shape: (3, 4)
┌──────────────┬────────────┬────────┬────────┐
│ name ┆ birthdate ┆ weight ┆ height │
│ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ date ┆ f64 ┆ f64 │
╞══════════════╪════════════╪════════╪════════╡
│ Alice Archer ┆ 1997-01-10 ┆ 57.9 ┆ 1.56 │
│ Ben Brown ┆ 1985-02-15 ┆ 72.5 ┆ 1.77 │
│ Chloe Cooper ┆ 1983-03-22 ┆ 53.6 ┆ 1.65 │
└──────────────┴────────────┴────────┴────────┘
> (tail df 3)
shape: (3, 4)
┌────────────────┬────────────┬────────┬────────┐
│ name ┆ birthdate ┆ weight ┆ height │
│ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ date ┆ f64 ┆ f64 │
╞════════════════╪════════════╪════════╪════════╡
│ Ben Brown ┆ 1985-02-15 ┆ 72.5 ┆ 1.77 │
│ Chloe Cooper ┆ 1983-03-22 ┆ 53.6 ┆ 1.65 │
│ Daniel Donovan ┆ 1981-04-30 ┆ 83.1 ┆ 1.75 │
└────────────────┴────────────┴────────┴────────┘
describe computes summary statistics for every column; the date column gets a mean and quartiles too:
> (describe df)
shape: (9, 5)
┌────────────┬────────────────┬─────────────────────┬───────────┬──────────┐
│ statistic ┆ name ┆ birthdate ┆ weight ┆ height │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ str ┆ str ┆ f64 ┆ f64 │
╞════════════╪════════════════╪═════════════════════╪═══════════╪══════════╡
│ count ┆ 4 ┆ 4 ┆ 4.0 ┆ 4.0 │
│ null_count ┆ 0 ┆ 0 ┆ 0.0 ┆ 0.0 │
│ mean ┆ null ┆ 1986-09-04 00:00:00 ┆ 66.775 ┆ 1.6825 │
│ std ┆ null ┆ null ┆ 13.560082 ┆ 0.097082 │
│ min ┆ Alice Archer ┆ 1981-04-30 ┆ 53.6 ┆ 1.56 │
│ 25% ┆ null ┆ 1983-03-22 ┆ 57.9 ┆ 1.65 │
│ 50% ┆ null ┆ 1985-02-15 ┆ 72.5 ┆ 1.75 │
│ 75% ┆ null ┆ 1985-02-15 ┆ 72.5 ┆ 1.75 │
│ max ┆ Daniel Donovan ┆ 1997-01-10 ┆ 83.1 ┆ 1.77 │
└────────────┴────────────────┴─────────────────────┴───────────┴──────────┘
API gaps: no glimpse; no sample / set_random_seed.
1.2.1.3 Schema
> (for ([name (column-names df)]) (printf "~a: ~a\n" name (dtype (ref df #:columns name))))
name: string
birthdate: date
weight: float64
height: float64
The #:dtype of each series plays the role of schema / schema_overrides:
> (dataframe (list (series '("Alice" "Ben" "Chloe" "Daniel") #:name "name") (series '(27 39 41 43) #:name "age" #:dtype 'u8)))
shape: (4, 2)
┌────────┬─────┐
│ name ┆ age │
│ --- ┆ --- │
│ str ┆ u8 │
╞════════╪═════╡
│ Alice ┆ 27 │
│ Ben ┆ 39 │
│ Chloe ┆ 41 │
│ Daniel ┆ 43 │
└────────┴─────┘
API gap: no schema accessor on a dataframe.
1.2.1.4 Data types
Dtype spellings accepted by series’ #:dtype:
Racket |
| Polars |
'bool |
| Boolean |
'i8 'i16 'i32 'i64 |
| Int8 … Int64 |
'u8 'u16 'u32 'u64 |
| UInt8 … UInt64 |
'f32 'f64 |
| Float32, Float64 |
'str |
| String |
'date |
| Date |
'time |
| Time |
'datetime or '(datetime milliseconds) |
| Datetime |
'categorical |
| Categorical |
'(enum low mid high) |
| Enum |
— |
| Decimal, Binary, Duration, Array, List, Struct |
Categorical and Enum values read back as symbols (Categorical, Enum and Decimal). A Decimal column, read from Parquet, has dtype '(decimal precision scale) and exact rational values, but no #:dtype spelling.
Long spellings ('int32, 'float64, 'string, …) are accepted too. Values that mix ints and floats promote to 'f64; see dtype promotion.
API gap: cast is asymmetric with series here —
> (dtype (cast (series '(1 2 3)) 'float64)) 'float64
> (dtype (cast (series '(1 2 3)) 'f64)) ->compat-dtype: unsupported cast target 'f64