1 Guide
A dataset is a named collection of observations. In this package, each observation occupies a row, and each variable occupies a named column. A dataset table holds those columns as immutable vectors, together with metadata describing their source and meaning.
1.1 Installing and loading data
Use Racket 9.3 or later. Install the core package from the Racket catalog:
raco pkg install --auto datasets-core
From a checkout of the repository, use raco pkg install ./datasets-core instead.
The repository contains two independently installable packages at its root: datasets-core/ and datasets/. Each has an info.rkt declaring a multi-collection package. They contribute different modules to the same datasets collection; package names and module paths need not match. The core package provides datasets/core.
> (require datasets/core) > (dataset-names)
'(air-passengers
airquality
anscombe
breast-cancer
diabetes
faithful
iris
mtcars
plant-growth
titanic
tooth-growth
us-arrests
wine)
> (define iris (load-dataset-table 'iris)) > (dataset-table-row-count iris) 150
> (dataset-table-names iris) '("sepal-length" "sepal-width" "petal-length" "petal-width" "species")
Dataset identifiers are symbols, such as 'iris; column names are strings, such as "sepal-length". Loading uses bundled files and does not contact a server. There is no need to install R or run an import script.
1.2 Working with columns
Use dataset-table-column to obtain a column, then ordinary vector operations to inspect its values. A vector index is zero-based.
> (define lengths (dataset-table-column iris "sepal-length")) > (vector-ref lengths 0) 5.1
> (for/sum ([length (in-vector lengths)]) length) 876.5000000000002
A selection keeps the source rows and places columns in the order you request:
> (define selected (load-dataset-table 'iris #:columns '("species" "sepal-length"))) > (dataset-table-names selected) '("species" "sepal-length")
> (dataset-table-row-count selected) 150
> (immutable? (dataset-table-column selected "species")) #t
The column vectors and their strings are immutable. Keep derived results in new values rather than changing the bundled observations. Empty selections, duplicate names, and unknown names are errors; see Registry and loading for the selection rules.
1.3 Recognizing missing observations
The airquality dataset contains missing ozone and solar-radiation measurements. A missing observation is represented by dataset-missing, a dedicated value recognized by dataset-missing?. It differs from #f, zero, and +nan.0.
> (define air (load-dataset-table 'airquality)) > (define ozone (dataset-table-column air "ozone")) > (dataset-missing? (vector-ref ozone 4)) #t
> (for/sum ([value (in-vector ozone)]) (if (dataset-missing? value) 1 0)) 37
Make any filtering decision explicit. For example, the following collects only observed ozone values; it does not modify the original table:
> (define observed (for/list ([value (in-vector ozone)] #:unless (dataset-missing? value)) value)) > (length observed) 116
1.4 Reading metadata and category levels
dataset-info returns an immutable hash whose keys are strings. Its "columns" entry is a list in source column order. Each entry describes one variable, including its normalized name, original spelling, semantic type, units, missing-value count, and category levels.
> (define info (dataset-info 'iris)) > (hash-ref info "rows") 150
> (define species-info (list-ref (hash-ref info "columns") 4)) > (hash-ref species-info "levels") '("Iris-setosa" "Iris-versicolor" "Iris-virginica")
> (hash-ref species-info "original-name") "class"
> (hash-ref info "license") "CC-BY-4.0"
Labels remain strings in the table. The level list records their intended order; it does not assert that a categorical variable is ordinal. Metadata on a selected table still describes the complete source dataset. See Metadata fields for the field definitions.
1.5 Counting people and indexing months
A row does not always represent one person. Titanic contains 32 contingency cells, each combining a class, sex, age group, and survival outcome. Its "frequency" column counts people, including zero for empty cells:
> (define titanic (load-dataset-table 'titanic)) > (dataset-table-row-count titanic) 32
> (for/sum ([count (in-vector (dataset-table-column titanic "frequency"))]) count) 2201
AirPassengers instead has one row per calendar month. Year and month are explicit columns, and passenger counts are in thousands:
> (define passengers (load-dataset-table 'air-passengers))
> (for/list ([name (in-list '("year" "month" "passengers"))]) (vector-ref (dataset-table-column passengers name) 0)) '(1949 1 112)
> (hash-ref (dataset-info 'air-passengers) "time-series")
'#hash(("end" . (1960 12))
("frequency" . 12)
("index" . ("year" "month"))
("start" . (1949 1)))
The Dataset Catalog and Provenance records these conventions alongside source citations and transformation notes. Consult it before interpreting a variable or comparing a dataset with another library’s version.