On this page:
2.1 Registry and loading
dataset-names
dataset-info
load-dataset-table
2.2 Table inspection
dataset-table?
dataset-table-names
dataset-table-columns
dataset-table-column
dataset-table-row-count
dataset-table-info
2.3 Missing values
dataset-missing
dataset-missing?
2.4 Metadata fields
9.3

2 Reference🔗ℹ

 (require datasets/core) package: datasets-core

This section specifies the core API. For a step-by-step introduction, see the Guide. Contracts below describe accepted arguments and results using the notation of The Racket Reference.

2.1 Registry and loading🔗ℹ

procedure

(dataset-names) → (listof symbol?)

Returns all available dataset identifiers, sorted alphabetically. The result is stable for the bundled catalog.

procedure

(dataset-info name) → (and/c hash? immutable?)

  name : symbol?
Returns deeply immutable metadata for name. An unknown identifier raises exn:fail:contract. Metadata describes the complete source dataset; it is independent of any column selection. See Metadata fields for its fields.

procedure

(load-dataset-table name [#:columns columns]) → dataset-table?

  name : symbol?
  columns : (or/c #f (listof string?)) = #f
Reads the bundled observations for name and returns a table. Data files are resolved relative to the installed module and read when this procedure is called; the current directory does not affect loading.

If columns is #f, all columns are returned in source order. Otherwise, it must be a nonempty list of distinct known column names. Its order determines the table’s column order. Row order and row count are preserved. Unknown datasets, unknown columns, empty selections, duplicate columns, and invalid argument types raise exn:fail:contract.

The result contains immutable column vectors. Strings inside those vectors are also immutable. Missing observations are dataset-missing. Repeated loads are equal; callers should not rely on object identity.

Examples:
> (dataset-table-names
   (load-dataset-table 'mtcars #:columns '("mpg" "model")))

'("mpg" "model")

> (load-dataset-table 'iris #:columns '("unknown"))

load-dataset-table: unknown column

  column: "unknown"

  available: '("sepal-length" "sepal-width" "petal-length"

"petal-width" "species")

2.2 Table inspection🔗ℹ

procedure

(dataset-table? value) → boolean?

  value : any/c
Returns #t if value is a dataset table, #f otherwise. There is no public table constructor or mutator.

procedure

(dataset-table-names table) → (listof string?)

  table : dataset-table?
Returns the selected column names in order. Each string is immutable.

procedure

(dataset-table-columns table) → (and/c vector? immutable?)

  table : dataset-table?
Returns an immutable vector of immutable column vectors, in the same order as dataset-table-names. Each column has (dataset-table-row-count table) elements.

procedure

(dataset-table-column table name) → (and/c vector? immutable?)

  table : dataset-table?
  name : string?
Returns the column named name from table. A name that is not among the selected columns raises exn:fail:contract.

Returns the number of source observations, including rows with missing values.

procedure

(dataset-table-info table) → (and/c hash? immutable?)

  table : dataset-table?
Returns the full source metadata. Its "column-count" may exceed the number of selected columns in table.

2.3 Missing values🔗ℹ

value

dataset-missing : any/c

procedure

(dataset-missing? value) → boolean?

  value : any/c
dataset-missing is the missing-value singleton. The predicate recognizes that value by identity. It returns #f for #f, zero, NaN, strings, and symbols, including 'missing. The singleton is not a number and must be handled explicitly before arithmetic.

2.4 Metadata fields🔗ℹ

All metadata hashes, lists, and strings are immutable. Dataset hashes have these string keys:

Key

  

Meaning

name

  

Dataset identifier as a string.

rows

  

Number of source rows.

column-count

  

Number of source columns, including identifiers.

columns

  

Ordered list of column-description hashes.

identifiers

  

Column names identifying observations; an empty list if none.

suggested-targets

  

Optional response or label column names; an empty list if unspecified.

source-url

  

URL of the pinned upstream archive.

source-version

  

Upstream version or hash-pinned snapshot description.

source-sha256

  

Hexadecimal SHA-256 of the upstream archive.

sha256

  

Hexadecimal SHA-256 of the normalized .rktd file's UTF-8 bytes.

license

  

Accepted upstream redistribution terms.

citation

  

Source attribution and bibliographic citation.

notes

  

Transformations, representation choices, and historical quirks.

time-series

  

Time-series description hash, or #f.

Each column-description hash has name, original-name, type, levels, unit, description, and missing-count keys. Semantic types are strings: integer, real, categorical, identifier, or string. Category levels are ordered strings; a noncategory has an empty level list. A unit of #f means unspecified. The missing count refers to that column in the complete source dataset.

For AirPassengers, the time-series hash has frequency (12 observations per year), start and end (year/month lists), and index (the year and month column names). Other v0.1 datasets have #f for time-series.