On this page:
1.1 Installation and a first dataset
1.2 Discovering datasets and selecting columns
1.3 Grouping categorical data
1.4 Handling missing observations
1.5 Preparing a numeric matrix
1.6 Using mutable data-frame objects
1.7 Interpreting counts and time series
1.8 Using core data in another package
9.3

1 Guide🔗ℹ

This guide starts with a dataframe and then shows how to select columns, choose another representation, and interpret missing values and metadata. If you already know which procedure you need, turn to the Reference.

1.1 Installation and a first dataset🔗ℹ

Use Racket 9.3 or later. Install from the Racket catalog:

  raco pkg install --auto datasets

This also installs datasets-core and the adapter dependencies. From a checkout of the repository, install both packages from its root:

raco pkg install ./datasets-core

raco pkg install --auto ./datasets

The two top-level directories are independently installable multi-collection packages. datasets-core provides datasets/core; datasets adds the convenient loaders and adapters to the same collection. The Polars package supplies its native library. With Nix, nix develop prepares an isolated environment containing both packages and that library.

Examples:
> (require datasets (prefix-in pl: polars))
> (define iris (load-iris))
> (pl:dataframe-height iris)

150

> (pl:dataframe-column-names iris)

'("sepal-length" "sepal-width" "petal-length" "petal-width" "species")

The prefix keeps Polars operations distinct from Racket’s arithmetic and other similarly named procedures. load-iris returns 150 rows: four flower measurements and a species label. Each call creates an independent dataframe. Data is bundled with the package, so loading does not download anything.

1.2 Discovering datasets and selecting columns🔗ℹ

dataset-names lists the available identifiers. A named loader, such as load-mtcars, is shorthand for (load-dataset 'mtcars). Both accept #:columns and #:format.

Examples:
> (dataset-names)

'(air-passengers

  airquality

  anscombe

  breast-cancer

  diabetes

  faithful

  iris

  mtcars

  plant-growth

  titanic

  tooth-growth

  us-arrests

  wine)

> (define cars (load-dataset 'mtcars #:columns '("model" "mpg")))
> (pl:dataframe-column-names cars)

'("model" "mpg")

> (hash-ref (dataset-info 'mtcars) "identifiers")

'("model")

Column names are strings. A selection preserves the order you give and retains all source rows. Identifiers such as "model" are real columns, so they can be kept for display or left out of a numeric matrix. Duplicate names, empty selections, and unknown columns are reported as errors.

1.3 Grouping categorical data🔗ℹ

Category labels remain strings in Polars. For example, Iris species labels can be used directly in a group operation:

Examples:
> (define by-species
    (pl:~> iris
           (pl:group-by "species")
           (pl:agg (pl:mean (pl:col "sepal-length")))))
> (pl:dataframe-height by-species)

3

> (pl:dataframe-column-names by-species)

'("species" "sepal-length")

The result has one row for each species. Grouped output need not follow category order; the intended level order is recorded separately in metadata:

Examples:
> (define species-info
    (list-ref (hash-ref (dataset-info 'iris) "columns") 4))
> (hash-ref species-info "levels")

'("Iris-setosa" "Iris-versicolor" "Iris-virginica")

Keeping labels avoids introducing arbitrary numeric codes. See Value representations for the representation of each value type.

1.4 Handling missing observations🔗ℹ

Airquality has missing measurements. The table format uses dataset-missing; the Polars adapter translates that singleton to pl:polars-null.

Examples:
> (define air (load-airquality))
> (pl:polars-null? (pl:series-ref (pl:dataframe-column air "ozone") 4))

#t

> (define air-table (load-airquality #:format 'table))
> (define ozone (dataset-table-column air-table "ozone"))
> (for/sum ([value (in-vector ozone)])
    (if (dataset-missing? value) 1 0))

37

Missing values are distinct from zero and #f. Decide whether to exclude, replace, or otherwise account for them in your analysis. The loaders preserve them; matrix conversion reports a missing value instead of silently removing it.

1.5 Preparing a numeric matrix🔗ℹ

Choose 'matrix for a math/matrix matrix. Rows are observations, and the requested column order determines the matrix’s columns. For Iris, select the four measurements explicitly to leave out the species label:

Examples:
> (require math/matrix)
> (define measurements
    (load-iris #:format 'matrix
               #:columns '("sepal-length" "sepal-width"
                           "petal-length" "petal-width")))
> (matrix-num-rows measurements)

150

> (matrix-num-cols measurements)

4

> (matrix-ref measurements 0 0)

5.1

Trying to include a label or a missing observation produces an error naming the column and the zero-based row:

Examples:
> (load-iris #:format 'matrix #:columns '("species"))

table->matrix: requires numeric, nonmissing values

  column: "species"

  row (zero-based): 0

  value: "Iris-setosa"

> (load-airquality #:format 'matrix #:columns '("ozone"))

table->matrix: requires numeric, nonmissing values

  column: "ozone"

  row (zero-based): 4

  value: #<missing-value>

Diabetes already has ten standardized predictors in the lars representation. See its entry in the Dataset Catalog and Provenance before comparing it with another library’s dataset; raw measurements and exact scikit-learn parity are not promised.

1.6 Using mutable data-frame objects🔗ℹ

The 'data-frame format works with the Racket data-frame library. Its NA value is explicitly set to dataset-missing.

Examples:
> (require (prefix-in df: data-frame) datasets/data-frame)
> (define frame (load-airquality #:format 'data-frame
                                 #:columns '("ozone" "temp")))
> (data-frame-column-names frame)

'("ozone" "temp")

> (df:df-is-na? frame "ozone" (df:df-ref frame 4 "ozone"))

#t

The underlying library stores columns by name and does not promise an order from df:df-series-names. data-frame-column-names retrieves the initial selection order recorded by this adapter.

Changes to one frame do not affect a later load:

Examples:
> (define first (load-iris #:format 'data-frame))
> (df:df-set! first 0 -100 "sepal-length")
> (df:df-ref first 0 "sepal-length")

-100

> (df:df-ref (load-iris #:format 'data-frame) 0 "sepal-length")

5.1

1.7 Interpreting counts and time series🔗ℹ

Titanic’s rows are contingency cells, not individual passengers. Sum the "frequency" column to count people, and retain zero-frequency rows when you need the complete table of combinations:

Examples:
> (define counts (load-titanic #:format 'table))
> (dataset-table-row-count counts)

32

> (for/sum ([count (in-vector (dataset-table-column counts "frequency"))])
    count)

2201

AirPassengers contains monthly international airline passenger counts, measured in thousands. Its year and month columns make the chronology explicit:

Examples:
> (define monthly (load-air-passengers #:format 'matrix))
> (for/list ([column (in-range 3)]) (matrix-ref monthly 0 column))

'(1949 1 112)

> (for/list ([column (in-range 3)]) (matrix-ref monthly 143 column))

'(1960 12 432)

> (hash-ref (dataset-info 'air-passengers) "time-series")

'#hash(("end" . (1960 12))

       ("frequency" . 12)

       ("index" . ("year" "month"))

       ("start" . (1949 1)))

The same metadata interface records source citations, licenses, units, identifiers, and file checksums for every dataset. The catalog is generated from that registry rather than maintained as a second independent list.

1.8 Using core data in another package🔗ℹ

A package that constructs its own dataframes can depend on datasets-core and use datasets/core directly. This is particularly useful for Polars documentation: depending only on the core avoids a dependency cycle through datasets.

The following example uses core data with Polars’ constructors:

Examples:
> (require datasets/core)
> (define table (load-dataset-table 'iris #:columns '("sepal-length")))
> (define own-frame
    (pl:dataframe-new
     (list (pl:series-new-f64
            "sepal-length"
            (vector->list (dataset-table-column table "sepal-length"))))))
> (pl:dataframe-height own-frame)

150

The repository’s examples/core-with-polars.rkt demonstrates the same pattern and is tested before the adapter package is installed. For downstream Nix, fetch this repository with flake = false and install datasets-core/ from that source. This keeps the complete flake and its native dependencies out of the downstream dependency graph.