On this page:
2.1 Dataset loaders
load-dataset
2.1.1 Named loaders
load-mtcars
load-iris
load-faithful
load-anscombe
load-airquality
load-us-arrests
load-plant-growth
load-tooth-growth
load-titanic
load-air-passengers
load-diabetes
load-breast-cancer
load-wine
2.2 Value representations
2.3 Polars conversion
table->polars
2.4 Matrix conversion
table->matrix
2.5 Data-frame conversion
table->data-frame
data-frame-column-names
9.3

2 Reference🔗ℹ

This section defines the loaders and conversion procedures. For worked examples, see the Guide. Contracts use the notation described in The Racket Reference.

2.1 Dataset loaders🔗ℹ

 (require datasets) package: datasets

All bindings from datasets/core are re-exported. Their definitions are in Datasets Core.

procedure

(load-dataset name    
  [#:format format    
  #:columns columns]) → any/c
  name : symbol?
  format : (or/c 'polars 'table 'matrix 'data-frame) = 'polars
  columns : (or/c #f (listof string?)) = #f
Loads name, selects columns, and converts the result to format. The return value is a Polars dataframe, an immutable dataset table, a math/matrix matrix, or a data-frame object as specified by the format.

When columns is #f, all columns are selected in source order. Otherwise it must be a nonempty list of distinct known column names. Selection preserves the requested order and all source rows, and takes place before conversion. An unknown dataset, format, or column, an empty selection, or a duplicate column raises exn:fail:contract.

Each dataframe load allocates independent storage. Mutating a returned dataframe cannot change later loads or the bundled observations. Table values are immutable. Matrix conversion raises an error for any selected nonnumeric or missing value; see Matrix conversion.

2.1.1 Named loaders🔗ℹ

Each named loader delegates to load-dataset with its corresponding dataset identifier. The keywords, defaults, selection rules, errors, and result types are the same. Variable descriptions and provenance appear in the Dataset Catalog and Provenance.

procedure

(load-mtcars [#:format format    
  #:columns columns]) → any/c
  format : symbol? = 'polars
  columns : (or/c #f (listof string?)) = #f
Loads mtcars.

procedure

(load-iris [#:format format    
  #:columns columns]) → any/c
  format : symbol? = 'polars
  columns : (or/c #f (listof string?)) = #f
Loads corrected UCI Iris.

procedure

(load-faithful [#:format format    
  #:columns columns]) → any/c
  format : symbol? = 'polars
  columns : (or/c #f (listof string?)) = #f
Loads faithful.

procedure

(load-anscombe [#:format format    
  #:columns columns]) → any/c
  format : symbol? = 'polars
  columns : (or/c #f (listof string?)) = #f
Loads the wide Anscombe quartet.

procedure

(load-airquality [#:format format    
  #:columns columns]) → any/c
  format : symbol? = 'polars
  columns : (or/c #f (listof string?)) = #f
Loads airquality, preserving missing values.

procedure

(load-us-arrests [#:format format    
  #:columns columns]) → any/c
  format : symbol? = 'polars
  columns : (or/c #f (listof string?)) = #f
Loads US arrests with state identifiers.

procedure

(load-plant-growth [#:format format    
  #:columns columns]) → any/c
  format : symbol? = 'polars
  columns : (or/c #f (listof string?)) = #f
Loads plant growth.

procedure

(load-tooth-growth [#:format format    
  #:columns columns]) → any/c
  format : symbol? = 'polars
  columns : (or/c #f (listof string?)) = #f
Loads tooth growth.

procedure

(load-titanic [#:format format    
  #:columns columns]) → any/c
  format : symbol? = 'polars
  columns : (or/c #f (listof string?)) = #f
Loads the contingency cells, including zero counts.

procedure

(load-air-passengers [#:format format    
  #:columns columns]) → any/c
  format : symbol? = 'polars
  columns : (or/c #f (listof string?)) = #f
Loads monthly counts in chronological order.

procedure

(load-diabetes [#:format format    
  #:columns columns]) → any/c
  format : symbol? = 'polars
  columns : (or/c #f (listof string?)) = #f
Loads the standardized lars representation.

procedure

(load-breast-cancer [#:format format    
  #:columns columns]) → any/c
  format : symbol? = 'polars
  columns : (or/c #f (listof string?)) = #f
Loads UCI diagnostic measurements, M/B labels, and identifiers.

procedure

(load-wine [#:format format    
  #:columns columns]) → any/c
  format : symbol? = 'polars
  columns : (or/c #f (listof string?)) = #f
Loads wine measurements and cultivar labels.

2.2 Value representations🔗ℹ

Value

  

Table

  

Polars

  

Matrix

  

data-frame

Integer

  

Exact integer

  

Int64, or Float64 in a mixed numeric column

  

Number

  

Number

Real

  

Inexact real

  

Float64

  

Number

  

Number

Category / identifier

  

String

  

String

  

Error

  

String

Missing

  

dataset-missing

  

polars-null

  

Error

  

Explicit series NA

Category levels and their order remain in the source metadata; the Polars adapter creates string columns rather than categorical or enum columns. No adapter encodes labels, imputes missing observations, or drops selected columns. The data-frame adapter records initial column order in a property because its underlying library does not define an iteration order for column names.

2.3 Polars conversion🔗ℹ

 (require datasets/polars) package: datasets

procedure

(table->polars table) → any/c

  table : dataset-table?
Constructs a Polars dataframe from the selected columns of table. Each column is newly allocated. String columns, including labels and identifiers, remain strings. A column of exact integers uses Int64; other numeric columns use Float64, with exact values converted to inexact values for that constructor. dataset-missing becomes polars-null. Column and row order are preserved. Source metadata is not attached to the dataframe; use dataset-info or dataset-table-info to retain it.

2.4 Matrix conversion🔗ℹ

 (require datasets/matrix) package: datasets

procedure

(table->matrix table) → any/c

  table : dataset-table?
Constructs a math/matrix matrix with one row per observation and one column per selected variable, in table order. Every value must be numeric and nonmissing. A nonnumeric or missing value raises exn:fail:contract identifying the column, zero-based row, and offending value. Labels and identifiers are not encoded. The matrix does not carry column names or source metadata.

2.5 Data-frame conversion🔗ℹ

 (require datasets/data-frame) package: datasets

procedure

(table->data-frame table) → any/c

  table : dataset-table?
Creates a data-frame object with a copied vector for each selected column. Each series uses dataset-missing as its NA value. The property 'datasets:columns stores the table’s ordered column names. Changing values in this frame does not change table or another load.

procedure

(data-frame-column-names frame) → (listof string?)

  frame : any/c
Returns the initial selection order stored in frame’s 'datasets:columns property. The argument must be a data-frame object created by table->data-frame or a loader using #:format 'data-frame. A frame without the property raises exn:fail:contract.

This is a snapshot of the selection at construction time. Callers that add, remove, or rename columns must update the property themselves if they continue to use it. The underlying df-series-names procedure has unspecified order.