2 Reference
| (require datasets/core) | package: datasets-core |
This section specifies the core API. For a step-by-step introduction, see the Guide. Contracts below describe accepted arguments and results using the notation of The Racket Reference.
2.1 Registry and loading
procedure
(dataset-names) → (listof symbol?)
procedure
(dataset-info name) → (and/c hash? immutable?)
name : symbol?
procedure
(load-dataset-table name [#:columns columns]) → dataset-table?
name : symbol? columns : (or/c #f (listof string?)) = #f
If columns is #f, all columns are returned in source order. Otherwise, it must be a nonempty list of distinct known column names. Its order determines the table’s column order. Row order and row count are preserved. Unknown datasets, unknown columns, empty selections, duplicate columns, and invalid argument types raise exn:fail:contract.
The result contains immutable column vectors. Strings inside those vectors are also immutable. Missing observations are dataset-missing. Repeated loads are equal; callers should not rely on object identity.
> (dataset-table-names (load-dataset-table 'mtcars #:columns '("mpg" "model"))) '("mpg" "model")
> (load-dataset-table 'iris #:columns '("unknown")) load-dataset-table: unknown column
column: "unknown"
available: '("sepal-length" "sepal-width" "petal-length"
"petal-width" "species")
2.2 Table inspection
procedure
(dataset-table? value) → boolean?
value : any/c
procedure
(dataset-table-names table) → (listof string?)
table : dataset-table?
procedure
(dataset-table-columns table) → (and/c vector? immutable?)
table : dataset-table?
procedure
(dataset-table-column table name) → (and/c vector? immutable?)
table : dataset-table? name : string?
procedure
(dataset-table-row-count table) → exact-nonnegative-integer?
table : dataset-table?
procedure
(dataset-table-info table) → (and/c hash? immutable?)
table : dataset-table?
2.3 Missing values
value
procedure
(dataset-missing? value) → boolean?
value : any/c
2.4 Metadata fields
All metadata hashes, lists, and strings are immutable. Dataset hashes have these string keys:
Key |
| Meaning |
name |
| Dataset identifier as a string. |
rows |
| Number of source rows. |
column-count |
| Number of source columns, including identifiers. |
columns |
| Ordered list of column-description hashes. |
identifiers |
| Column names identifying observations; an empty list if none. |
suggested-targets |
| Optional response or label column names; an empty list if unspecified. |
source-url |
| URL of the pinned upstream archive. |
source-version |
| Upstream version or hash-pinned snapshot description. |
source-sha256 |
| Hexadecimal SHA-256 of the upstream archive. |
sha256 |
| Hexadecimal SHA-256 of the normalized .rktd file's UTF-8 bytes. |
license |
| Accepted upstream redistribution terms. |
citation |
| Source attribution and bibliographic citation. |
notes |
| Transformations, representation choices, and historical quirks. |
time-series |
| Time-series description hash, or #f. |
Each column-description hash has name, original-name, type, levels, unit, description, and missing-count keys. Semantic types are strings: integer, real, categorical, identifier, or string. Category levels are ordered strings; a noncategory has an empty level list. A unit of #f means unspecified. The missing count refers to that column in the complete source dataset.
For AirPassengers, the time-series hash has frequency (12 observations per year), start and end (year/month lists), and index (the year and month column names). Other v0.1 datasets have #f for time-series.