On this page:
1.1 Installing and loading data
1.2 Working with columns
1.3 Recognizing missing observations
1.4 Reading metadata and category levels
1.5 Counting people and indexing months
9.3

1 Guide🔗ℹ

A dataset is a named collection of observations. In this package, each observation occupies a row, and each variable occupies a named column. A dataset table holds those columns as immutable vectors, together with metadata describing their source and meaning.

1.1 Installing and loading data🔗ℹ

Use Racket 9.3 or later. Install the core package from the Racket catalog:

  raco pkg install --auto datasets-core

From a checkout of the repository, use raco pkg install ./datasets-core instead.

The repository contains two independently installable packages at its root: datasets-core/ and datasets/. Each has an info.rkt declaring a multi-collection package. They contribute different modules to the same datasets collection; package names and module paths need not match. The core package provides datasets/core.

Examples:
> (require datasets/core)
> (dataset-names)

'(air-passengers

  airquality

  anscombe

  breast-cancer

  diabetes

  faithful

  iris

  mtcars

  plant-growth

  titanic

  tooth-growth

  us-arrests

  wine)

> (define iris (load-dataset-table 'iris))
> (dataset-table-row-count iris)

150

> (dataset-table-names iris)

'("sepal-length" "sepal-width" "petal-length" "petal-width" "species")

Dataset identifiers are symbols, such as 'iris; column names are strings, such as "sepal-length". Loading uses bundled files and does not contact a server. There is no need to install R or run an import script.

1.2 Working with columns🔗ℹ

Use dataset-table-column to obtain a column, then ordinary vector operations to inspect its values. A vector index is zero-based.

Examples:
> (define lengths (dataset-table-column iris "sepal-length"))
> (vector-ref lengths 0)

5.1

> (for/sum ([length (in-vector lengths)]) length)

876.5000000000002

A selection keeps the source rows and places columns in the order you request:

Examples:
> (define selected
    (load-dataset-table 'iris #:columns '("species" "sepal-length")))
> (dataset-table-names selected)

'("species" "sepal-length")

> (dataset-table-row-count selected)

150

> (immutable? (dataset-table-column selected "species"))

#t

The column vectors and their strings are immutable. Keep derived results in new values rather than changing the bundled observations. Empty selections, duplicate names, and unknown names are errors; see Registry and loading for the selection rules.

1.3 Recognizing missing observations🔗ℹ

The airquality dataset contains missing ozone and solar-radiation measurements. A missing observation is represented by dataset-missing, a dedicated value recognized by dataset-missing?. It differs from #f, zero, and +nan.0.

Examples:
> (define air (load-dataset-table 'airquality))
> (define ozone (dataset-table-column air "ozone"))
> (dataset-missing? (vector-ref ozone 4))

#t

> (for/sum ([value (in-vector ozone)])
    (if (dataset-missing? value) 1 0))

37

Make any filtering decision explicit. For example, the following collects only observed ozone values; it does not modify the original table:

Examples:
> (define observed
    (for/list ([value (in-vector ozone)]
               #:unless (dataset-missing? value))
      value))
> (length observed)

116

1.4 Reading metadata and category levels🔗ℹ

dataset-info returns an immutable hash whose keys are strings. Its "columns" entry is a list in source column order. Each entry describes one variable, including its normalized name, original spelling, semantic type, units, missing-value count, and category levels.

Examples:
> (define info (dataset-info 'iris))
> (hash-ref info "rows")

150

> (define species-info (list-ref (hash-ref info "columns") 4))
> (hash-ref species-info "levels")

'("Iris-setosa" "Iris-versicolor" "Iris-virginica")

> (hash-ref species-info "original-name")

"class"

> (hash-ref info "license")

"CC-BY-4.0"

Labels remain strings in the table. The level list records their intended order; it does not assert that a categorical variable is ordinal. Metadata on a selected table still describes the complete source dataset. See Metadata fields for the field definitions.

1.5 Counting people and indexing months🔗ℹ

A row does not always represent one person. Titanic contains 32 contingency cells, each combining a class, sex, age group, and survival outcome. Its "frequency" column counts people, including zero for empty cells:

Examples:
> (define titanic (load-dataset-table 'titanic))
> (dataset-table-row-count titanic)

32

> (for/sum ([count (in-vector (dataset-table-column titanic "frequency"))])
    count)

2201

AirPassengers instead has one row per calendar month. Year and month are explicit columns, and passenger counts are in thousands:

Examples:
> (define passengers (load-dataset-table 'air-passengers))
> (for/list ([name (in-list '("year" "month" "passengers"))])
    (vector-ref (dataset-table-column passengers name) 0))

'(1949 1 112)

> (hash-ref (dataset-info 'air-passengers) "time-series")

'#hash(("end" . (1960 12))

       ("frequency" . 12)

       ("index" . ("year" "month"))

       ("start" . (1949 1)))

The Dataset Catalog and Provenance records these conventions alongside source citations and transformation notes. Consult it before interpreting a variable or comparing a dataset with another library’s version.