Public API

Everything a caller needs. Functions the package uses internally are in Internals.

The module

OBISClient — Module
OBISClient

A Julia client for OBIS, the Ocean Biodiversity Information System, a programme of the Intergovernmental Oceanographic Commission of UNESCO.

Results are Tables.jl tables with a fixed, concretely typed schema, and they carry the licence and citation of every dataset they draw on together with the date they were retrieved.

Not an official OBIS product

This is an independent, community-maintained client. It is not affiliated with, endorsed by, or maintained by OBIS, the Intergovernmental Oceanographic Commission, or UNESCO. The name contracts Ocean Biodiversity Information System, the service the package connects to; the General registry does not accept OBIS itself, being four characters and all upper case.

Getting started

using OBISClient, DataFrames

recs = OBISClient.occurrence("Abra alba"; limit = 1000)
df = DataFrame(recs)

OBISClient.licenses(recs)                        # may these data be redistributed?
print(OBISClient.citations(recs; format = :text))  # how to credit them

Interpreting the results

Record counts measure sampling effort as much as biology, and the default view of OBIS excludes two classes of record that change what a result means. See the "Interpreting OBIS data" section of the documentation, and the disclaimer in OBIS_DISCLAIMER, before drawing conclusions.

source

Occurrences

OBISClient.occurrence — Function
occurrence(scientificname = nothing; kwargs...) -> OBISTable

Retrieve occurrence records.

Returns an OBISTable: a Tables.jl-compatible table whose columns are fixed by the schema, so the same columns appear in the same order for every query, with missing where a record carries no value. Darwin Core fields outside the core schema are preserved per row in the extra column.

Each row carries the licence and citation of the dataset it came from, so citations and licenses work on the result without a further query.

Filters

  • scientificname: taxon name at any rank. Also accepted positionally.
  • taxonid (alias aphiaid): WoRMS AphiaID. Unambiguous where a name is not.
  • geometry: WKT, for example "POLYGON ((2.3 51.8, 2.3 51.6, 2.6 51.6, 2.6 51.8, 2.3 51.8))".
  • startdate, enddate: Date, DateTime, or "YYYY-MM-DD".
  • startdepth, enddepth: metres below the surface.
  • datasetid, nodeid: UUIDs. instituteid: OceanExpert ID. areaid: OBIS area ID.
  • flags: quality flags that must be set. exclude: flags to exclude. Either case is accepted and normalized.
  • absence, dropped, event: :exclude (default), :include, or :only.
  • redlist, hab, wrims: restrict to IUCN Red List, harmful algal bloom, or WRiMS species.
  • measurementtype, measurementvalue, measurementunit and their *id forms: require a matching measurement.
  • hasextensions: require an extension, for example "DNADerivedData".

Retrieval

  • limit: stop after this many records. Records arrive ordered by UUID, which is unrelated to date, place or taxon, so a limited pull is an arbitrary subset of the matches rather than the earliest or nearest ones.
  • after: resume from a record UUID. See occurrence_pages.
  • licenses: look up dataset rights. true by default; set false to skip the extra request, which leaves the license columns missing without changing the schema.
  • check_size: estimate the query first and refuse to page through very large results. See estimate_size.

Examples

# A species, with rights attached.
recs = OBISClient.occurrence("Abra alba"; limit = 500)

# Absence records: observations where the species was looked for and not found. These are
# available through the API and not through the bulk downloads.
absences = OBISClient.occurrence("Abra alba"; absence = :only)

# Records the quality pipeline dropped, to see what a default query leaves out.
dropped = OBISClient.occurrence("Abra alba"; dropped = :only)

# An area and a period.
recs = OBISClient.occurrence(;
    scientificname = "Delphinidae",
    geometry = "POLYGON ((2.3 51.8, 2.3 51.6, 2.6 51.6, 2.6 51.8, 2.3 51.8))",
    startdate = Date(2010, 1, 1),
    enddate = Date(2020, 12, 31),
)

# Exclude records the pipeline flagged as being on land.
clean = OBISClient.occurrence("Abra alba"; exclude = "ON_LAND", limit = 1000)
source
OBISClient.occurrence_pages — Function
occurrence_pages(scientificname = nothing; kwargs...) -> OccurrencePages

Build a lazy, resumable iterator over pages of occurrence records.

Accepts the same filters as occurrence. Nothing is fetched until iteration begins, and each iteration yields one page as an OBISTable, so a query larger than memory can be processed page by page.

Extra keywords: page_size sets records per request (up to 10000); after resumes from a record UUID; licenses attaches dataset rights to every page, which costs one extra request per page and is therefore off by default.

Examples

# Summarize a large query without materializing it.
counts = Dict{String,Int}()
for page in OBISClient.occurrence_pages(scientificname = "Mollusca"; page_size = 10_000)
    for name in page.scientificName
        ismissing(name) || (counts[name] = get(counts, name, 0) + 1)
    end
end

# Checkpoint after every page so an interrupted pull can resume.
pages = OBISClient.occurrence_pages(scientificname = "Mollusca")
for page in pages
    save(page)
    write("obis.cursor", OBISClient.cursor(pages))
end
source
OBISClient.OccurrencePages — Type
OccurrencePages

A lazy iterator over pages of occurrence records.

Each iteration yields an OBISTable holding one page. Nothing is fetched until iteration starts, and only one page is held in memory at a time, so a query far larger than memory can still be processed.

The cursor is a record UUID. Read it with cursor after any page to checkpoint progress, and pass it back as after to resume — in the same session or a later one.

Construct with occurrence_pages.

Examples

# Process a large query without holding it in memory.
total = 0
for page in OBISClient.occurrence_pages(scientificname = "Mollusca")
    total += OBISClient.nrow(page)
end

# Checkpoint, then resume in another session.
pages = OBISClient.occurrence_pages(scientificname = "Mollusca")
for page in pages
    process(page)
    write("checkpoint.txt", OBISClient.cursor(pages))
    break
end

resumed = OBISClient.occurrence_pages(
    scientificname = "Mollusca", after = read("checkpoint.txt", String)
)
source
OBISClient.cursor — Function
cursor(pages::OccurrencePages) -> Union{Nothing,String}

The UUID of the last record yielded so far, or nothing before the first page.

Persist this to resume a pull later: it is the whole of the pagination state.

source
OBISClient.expected — Function
expected(pages::OccurrencePages) -> Int

How many records the API reports as matching, known after the first page is fetched.

source

Taxa and checklists

OBISClient.taxon — Function
taxon(id) -> OBISTable

Look up a taxon by WoRMS AphiaID or by exact scientific name.

An AphiaID is unambiguous; a name is not, and the API matches it exactly rather than fuzzily. There is no taxon search endpoint — use checklist to discover which taxa occur in a selection.

julia> t = OBISClient.taxon(141433);       # by AphiaID

julia> t = OBISClient.taxon("Abra alba");  # by exact name
source
OBISClient.taxon_annotations — Function
taxon_annotations(; scientificname = nothing, kwargs...) -> Vector

Annotations the WoRMS team made on scientific names that OBIS could not match automatically.

These explain why a name carries NO_MATCH or a WORMS_ANNOTATION_* flag, and whether it can be fixed.

source
OBISClient.checklist — Function
checklist(scientificname = nothing; kwargs...) -> OBISTable

Build a species checklist for a selection.

Accepts the same filters as occurrence, and returns one row per taxon with the number of records supporting it. This is the way to find out which taxa occur somewhere; the record counts are counts of observations, which reflect how much sampling has happened there as much as what lives there.

# What has been recorded in this polygon?
list = OBISClient.checklist(;
    geometry = "POLYGON ((2.3 51.8, 2.3 51.6, 2.6 51.6, 2.6 51.8, 2.3 51.8))"
)

# Red List species in the same area.
OBISClient.checklist_redlist(; geometry = "POLYGON ((2.3 51.8, 2.3 51.6, 2.6 51.6, 2.6 51.8, 2.3 51.8))")
source
OBISClient.checklist_newest — Function
checklist_newest(scientificname = nothing; kwargs...) -> OBISTable

Checklist of the most recently added species in a selection.

source

Datasets, nodes, institutes, areas, countries

OBISClient.dataset — Function
dataset(scientificname = nothing; kwargs...) -> OBISTable

Find datasets matching a query.

Accepts the same filters as occurrence, so dataset describes exactly the datasets an occurrence query draws on. Each row carries the provider's citation string, the rights statement as published, and a normalized license identifier derived from it.

This endpoint does not paginate: every match is returned in one response. An unfiltered call therefore retrieves metadata for every dataset in OBIS, currently around seven thousand.

Examples

# Datasets behind a species.
ds = OBISClient.dataset("Abra alba")

# Which licences are in play, before downloading any occurrences.
OBISClient.licenses(ds)
source
OBISClient.dataset_by_id — Function
dataset_by_id(id) -> OBISTable

Fetch one dataset by its UUID.

julia> ds = OBISClient.dataset_by_id("8acba7e7-2e50-4490-8328-b78a30472508");
source
OBISClient.dataset_errors — Function
dataset_errors(id) -> Vector

Loading errors OBIS recorded for a dataset, useful when a dataset holds fewer records than expected.

source
OBISClient.node — Function
node(id = nothing) -> OBISTable

List OBIS nodes, or fetch one by UUID.

Nodes are the regional and thematic partners that curate the data. A node UUID is the nodeid filter accepted elsewhere.

julia> nodes = OBISClient.node();          # all nodes

julia> n = OBISClient.node("4bf79a01-65a9-4db6-b37b-18434f26ddfc");
source
OBISClient.institute — Function
institute(scientificname = nothing; kwargs...) -> OBISTable

Find institutes contributing to a selection.

Identifiers are OceanExpert IDs, which is what the instituteid filter expects elsewhere.

source
OBISClient.area — Function
area(id = nothing) -> OBISTable

List the areas OBIS can filter by, or fetch one by ID.

Areas include exclusive economic zones, marine protected areas and other named regions. An area ID is what the areaid filter expects.

julia> areas = OBISClient.area();

julia> belgian = [r for r in areas.name if occursin("Belg", r)]
source
OBISClient.country — Function
country(id = nothing) -> OBISTable

List the countries OBIS can filter by, or fetch one by ID.

source
OBISClient.metrics — Function
metrics(; datasetid = nothing, nodeid = nothing) -> Any

Yearly download counts for a dataset or a node.

Provide one or the other. When both are given the API uses datasetid.

source
OBISClient.metrics_downloads — Function
metrics_downloads(datasetid; startdate = nothing, enddate = nothing, groupby = nothing)

Download counts for a dataset, optionally within a date range.

Pass groupby = "time" for one entry per download event instead of an aggregate.

source

Statistics and facets

OBISClient.statistics — Function
statistics(scientificname = nothing; kwargs...) -> Dict{String,Any}

Summary counts for a query: records, species, taxa, datasets, and the year range.

Accepts the same filters as occurrence and counts exactly the records that query would return, so it is the cheap way to size a query before running it.

julia> st = OBISClient.statistics("Abra alba");

julia> st["records"], st["datasets"]
(75356, 216)

julia> OBISClient.statistics("Abra alba"; absence = :only)["records"]
18387
source
OBISClient.statistics_years — Function
statistics_years(scientificname = nothing; kwargs...) -> Vector{Dict{String,Any}}

Number of presence records per year.

Useful for showing what record counts actually measure. A rise in a series like this tracks survey programmes, digitization projects and the growth of OBIS itself; it is not evidence that a taxon became more abundant.

julia> years = OBISClient.statistics_years("Abra alba");

julia> first(years)
Dict{String, Any}("year" => 1841, "records" => 8)
source
OBISClient.statistics_env — Function
statistics_env(scientificname = nothing; kwargs...) -> Any

Record counts per sea surface temperature, salinity, or depth bin.

source
OBISClient.statistics_qc — Function
statistics_qc(scientificname = nothing; kwargs...) -> Dict{String,Any}

Quality summary for a query: missing and invalid fields, records on land, non-marine records, and records without an AphiaID.

Worth running before an analysis. It reports what a default query silently excludes and what the records it does return are flagged for.

source
OBISClient.facet — Function
facet(facets; kwargs...) -> Dict{String,Vector}

Record counts grouped by one or more fields.

facets names the fields, for example "flags", "datasetName" or ["originalScientificName", "flags"].

The API truncates each facet's value list, so a facet result is a ranking of the most common values, not a complete enumeration. In particular the flags facet does not list every flag OBIS uses; OBISClient.KNOWN_FLAGS is the fuller reference.

julia> f = OBISClient.facet("flags"; scientificname = "Abra alba");

julia> first(f["flags"])
Dict{String, Any}("key" => "NO_DEPTH", "records" => 29772)
source
OBISClient.estimate_size — Function
estimate_size(; kwargs...) -> Int

Number of records a query would return, from one /statistics request.

julia> OBISClient.estimate_size(scientificname = "Mollusca") > 1_000_000
true
source

Results

OBISClient.OBISTable — Type
OBISTable

A typed, column-oriented table of OBIS records that implements the Tables.jl interface.

The column set is fixed by the schema for the endpoint, so two queries against the same endpoint always produce the same columns in the same order, with missing where a record carried no value. Fields outside the core schema are preserved per row in the extra column as a Dict{Symbol,Any}.

Columns are reached by property access, by index, or through any Tables.jl consumer:

tbl.scientificName          # a column
tbl[:decimalLatitude]       # the same column
DataFrame(tbl)              # with DataFrames loaded
CSV.write("out.csv", tbl)   # with CSV loaded

Use citations and licenses on a table to obtain the attribution its data requires, and metadata to recover the query and access date.

source
OBISClient.QueryMeta — Type
QueryMeta

Provenance of a result: what was asked, of which service, and when.

The access date is recorded at request time because the OBIS citation format requires Accessed: YYYY-MM-DD and nothing in the data itself supplies it. Carrying it on the result is what lets citations produce a correct citation with no bookkeeping by the caller.

Fields

  • endpoint: API path the records came from.
  • params: query parameters as sent, in canonical order.
  • accessed: date of retrieval, for citations.
  • retrieved: UTC timestamp of retrieval.
  • total: number of records the API reported as matching, which may exceed the number held here when a query was limited or paged partially.
  • base_url: service the records came from.
  • package_version: version of this package that fetched them.
source
OBISClient.metadata — Function
metadata(t::OBISTable) -> QueryMeta

Return the query, access date and reported total behind a result.

julia> meta = OBISClient.metadata(tbl);

julia> meta.accessed        # the date to put in a citation
2026-09-06
source
OBISClient.nrow — Function
nrow(t::OBISTable) -> Int

Number of records held in the table.

This can be smaller than metadata(t).total, which is how many records the API reported as matching the query.

source
OBISClient.extra_names — Function
extra_names(t::OBISTable) -> Vector{Symbol}

Names of the non-core fields present anywhere in the table, sorted.

These are the Darwin Core terms the providers supplied that fall outside the core schema. Which of them appear depends on the query, which is exactly why they are held apart from the core columns.

julia> OBISClient.extra_names(tbl)
12-element Vector{Symbol}:
 :day
 :eventTime
 ⋮
source
OBISClient.extra_column — Function
extra_column(t::OBISTable, name) -> Vector

Lift one non-core field out of extra into a plain column, with missing where a record did not carry it.

julia> OBISClient.extra_column(tbl, :waterBody)
source

Licensing and citation

OBISClient.licenses — Function
licenses(t::OBISTable) -> OBISTable

Summarize which licences a result is under.

One row per distinct licence, with how many datasets and records fall under it and what it permits. Read it before redistributing: a single CC BY-NC dataset makes the combined result non-commercial, and "unknown" covers rights statements that could not be identified — among them the literal value Restricted — which are not permission to redistribute.

julia> OBISClient.licenses(OBISClient.occurrence("Abra alba"; limit = 500))
OBISTable: 3 records, 7 columns
source
OBISClient.citations — Function
citations(t::OBISTable; format = :table)

Build the citations required for the data in a result.

Returns one entry per dataset the records came from, with the provider's citation string, its DOI where one exists, its licence, how many records in the result it contributed, and a ready-to-use citation in the format the OBIS data policy specifies — including the access date recorded when the data was retrieved.

Formats:

  • :table — an OBISTable, for further processing or writing to a file.
  • :text — the formatted citations as one string, ready to paste.
  • :bibtex — a BibTeX document, one @misc entry per dataset.

The dataset citation is not the only obligation. Where a result draws on many datasets, the policy also allows citing the OBIS database as a whole; obis_citation builds that.

Examples

recs = OBISClient.occurrence("Abra alba"; limit = 1000)

cites = OBISClient.citations(recs)          # a table, one row per dataset
print(OBISClient.citations(recs; format = :text))
write("references.bib", OBISClient.citations(recs; format = :bibtex))
source
OBISClient.obis_citation — Function
obis_citation(; description = nothing, accessed = today(), year = Dates.year(today()))

The citation for the OBIS database as a whole.

The data policy allows citing the integrated database in addition to — never instead of — the individual datasets, whose own restrictions continue to apply.

julia> OBISClient.obis_citation(; description = "Distribution records of Abra alba", accessed = Date(2026, 9, 6), year = 2026)
"OBIS (2026) Distribution records of Abra alba [Dataset] (Available: Ocean Biodiversity Information System. Intergovernmental Oceanographic Commission of UNESCO. https://obis.org. Accessed: 2026-09-06)"
source
OBISClient.normalize_license — Function
normalize_license(text) -> String

Derive a licence identifier from the free-text rights statement OBIS publishes.

Returns one of "CC0-1.0", "CC-BY-4.0", "CC-BY-NC-4.0", "CC-BY-SA-4.0", or "unknown". Both prose statements and licence URLs are accepted, so a value taken from the API and one taken from the export's licence table normalize the same way.

Non-commercial is tested before attribution: every CC-BY-NC statement also contains the word "Attribution", so the looser test would swallow the stricter licence.

julia> OBISClient.normalize_license("This work is licensed under a  Creative Commons Attribution (CC-BY) 4.0 License")
"CC-BY-4.0"

julia> OBISClient.normalize_license("This work is licensed under a  Creative Commons Attribution Non Commercial (CC-BY-NC) 4.0 License")
"CC-BY-NC-4.0"

julia> OBISClient.normalize_license("http://creativecommons.org/publicdomain/zero/1.0/legalcode")
"CC0-1.0"

julia> OBISClient.normalize_license("Restricted")
"unknown"
source
OBISClient.license_url — Function
license_url(id) -> Union{Missing,String}

Canonical URL for a licence identifier, or missing when it is not recognized.

julia> OBISClient.license_url("CC-BY-NC-4.0")
"https://creativecommons.org/licenses/by-nc/4.0/"

julia> OBISClient.license_url("unknown")
missing
source
OBISClient.permits_redistribution — Function
permits_redistribution(id) -> Bool

Whether a licence allows redistributing the data at all.

True for every recognized Creative Commons licence here. It is false for "unknown", because an unidentifiable rights statement is not permission — the observed unknowns include the literal string Restricted.

source
OBISClient.permits_commercial_use — Function
permits_commercial_use(id) -> Bool

Whether a licence permits commercial reuse.

False for CC-BY-NC-4.0 and for "unknown". A result set containing even one non-commercial dataset cannot be redistributed commercially as a whole, which is why licenses reports the mix rather than a single verdict.

source
OBISClient.requires_attribution — Function
requires_attribution(id) -> Bool

Whether a licence requires the data provider to be credited.

False only for CC0-1.0, the one licence here that waives the requirement. "unknown" counts as requiring attribution: an unidentifiable rights statement is not evidence that attribution was waived, and treating it as such is exactly the permissive guess the rest of this file avoids.

This describes the legal obligation. The OBIS data policy asks that providers be credited regardless of licence, so false is not advice to omit the credit.

julia> OBISClient.requires_attribution("CC0-1.0")
false

julia> OBISClient.requires_attribution("CC-BY-4.0")
true

julia> OBISClient.requires_attribution("unknown")
true
source
OBISClient.ACCEPTED_LICENSES — Constant
ACCEPTED_LICENSES

The three licences the OBIS data policy accepts for published datasets.

Anything else in a result is either an older record, a provider error, or a licence OBIS does not endorse, and is worth the caller's attention.

source

Quality flags

OBISClient.KNOWN_FLAGS — Constant
KNOWN_FLAGS

Quality flags documented by the OBIS QC pipeline, together with flags observed live that the pipeline reference does not list.

The vocabulary is open: OBIS adds checks over time, and the /facet endpoint truncates its flag listing, so this set cannot be treated as exhaustive. Unrecognized upper-case flags are passed through with a warning rather than rejected.

source
OBISClient.DROPPING_FLAGS — Constant
DROPPING_FLAGS

Flags whose presence causes OBIS to drop a record from the main index.

Records carrying these are reachable only with dropped = :only or dropped = :include, because a dropped record is excluded from every default query and from the bulk exports.

source
OBISClient.normalize_flag — Function
normalize_flag(flag) -> String

Return flag in the upper-case form the API matches on, warning once if the flag is not in KNOWN_FLAGS.

Unknown flags are passed through rather than rejected: OBIS adds quality checks over time, and a hard allow-list would reject a valid new flag.

julia> OBISClient.normalize_flag("on_land")
"ON_LAND"

julia> OBISClient.normalize_flag(:no_depth)
"NO_DEPTH"
source
OBISClient.normalize_flags — Function
normalize_flags(flags) -> Vector{String}

Normalize a flag, or any iterable of flags, into the upper-case forms the API expects.

Accepts a String, a Symbol, or any iterable of those. A comma-separated string is split, so both "ON_LAND,NO_DEPTH" and ["ON_LAND", "NO_DEPTH"] work.

julia> OBISClient.normalize_flags(["on_land", :NO_DEPTH])
2-element Vector{String}:
 "ON_LAND"
 "NO_DEPTH"

julia> OBISClient.normalize_flags("on_land,no_match")
2-element Vector{String}:
 "ON_LAND"
 "NO_MATCH"
source
OBISClient.drops_record — Function
drops_record(flag) -> Bool

Whether a record carrying flag is dropped from the OBIS main index.

julia> OBISClient.drops_record("NO_MATCH")
true

julia> OBISClient.drops_record("on_land")
false
source

The bulk export

OBISClient.export_covers — Function
export_covers(; absence = nothing, dropped = nothing, event = nothing, kwargs...) -> Bool
export_covers(params::QueryParams) -> Bool

Whether the bulk export can serve a query.

false only for a query that selects pure event records. The export has no column that identifies them — its event information is a _event_id carried on each occurrence, not a record class — so there is nothing in a file to select on.

Absence and dropped records are in the export, contrary to the OBIS data access page, which says exports carry neither. Verified per dataset against /statistics: across five datasets the export's absence and dropped row counts matched the API's absence = :only and dropped = :only counts exactly, and each export's total equalled the default count plus them (NOTES.md §7.3). Reading them out of an export is a filter on those two columns.

What still differs between the routes is the date. The export is regenerated periodically and the API is live, so the same query answered from each will differ by whatever OBIS has ingested since — which is why a large query raises rather than switching route on its own.

source
OBISClient.export_url — Function
export_url(dataset_id) -> String

URL of the GeoParquet export for one dataset.

julia> OBISClient.export_url("8acba7e7-2e50-4490-8328-b78a30472508")
"https://obis-open-data.s3.amazonaws.com/occurrence/8acba7e7-2e50-4490-8328-b78a30472508.parquet"
source
OBISClient.export_archive_url — Function
export_archive_url(dataset_id) -> String

URL of the source Darwin Core Archive for one dataset.

The archive holds the data exactly as the provider published it, without the fields OBIS's quality pipeline adds.

source
OBISClient.download_export — Function
download_export(dataset_id; dir = pwd(), overwrite = false) -> String

Download one dataset's GeoParquet export and return the local path.

path = OBISClient.download_export("8acba7e7-2e50-4490-8328-b78a30472508"; dir = "data")

Read it with whichever Parquet reader you prefer; the package does not impose one. The schema is nested — provider terms under source, pipeline-interpreted terms under interpreted — and so differs from the flat schema the API returns.

source
OBISClient.download_exports — Function
download_exports(scientificname = nothing; dir = pwd(), overwrite = false, kwargs...)

Download the GeoParquet exports for every dataset a query touches.

Resolves the query to its datasets with dataset, then fetches one file per dataset, serially. Returns the local paths.

One file per dataset, whole: the export is not filtered, so a query's filters are applied by you after reading. Absence and dropped records are present and carry their own columns; pure event records cannot be selected at all. export_covers reports which case a query is in.

paths = OBISClient.download_exports("Abra alba"; dir = "obis-export")
source
OBISClient.export_licenses — Function
export_licenses(; path = nothing) -> OBISTable

Fetch the licence table published alongside the bulk export.

One row per dataset, with the rights statement, its canonical URL and the provider's citation. This is a companion to the export files rather than a live view of the API: it is regenerated with the export, so it lags the current dataset list. For a live query, take licences from dataset or from an occurrence result.

Pass path to read a previously downloaded copy instead of fetching it.

source
OBISClient.read_export — Function
read_export(path; absence = :exclude, dropped = :exclude, licenses = true,
            limit = nothing, source_terms = false) -> OBISTable

Read one or more GeoParquet export files into the package's canonical occurrence schema.

path is a file, or a vector of them — so read_export(download_exports(...)) composes.

Requires DuckDB: using DuckDB loads the extension that implements this. Parquet2.jl cannot read these files, whose nested source and interpreted structs it does not support.

The result is an OBISTable with the same columns, in the same order and with the same types, as one from occurrence, so the two access routes become interchangeable in everything downstream — licenses and citations included. Two columns are always missing: the export carries marine and brackish but not freshwater or terrestrial.

absence and dropped default to :exclude, matching the API, because the export contains those records (NOTES.md §7.3) and a plain read would otherwise mix them into an ordinary count without saying so. :include and :only mean what they mean everywhere else, and the filter is pushed into the read rather than applied afterwards.

licenses = true fetches export_licenses once. Pass a table returned by it to reuse across many files, or false to leave the three rights columns missing.

source_terms = true fills the extra column with the provider's own terms from source. It is off by default: that block has 188 fields, and reading them costs far more than the core schema does.

The access date on the result is the file's modification time, not today — the data is as old as the export, and a citation built from it should say so.

source

Configuration and caching

OBISClient.config — Function
config() -> ClientConfig

Return the active client configuration.

julia> OBISClient.config().base_url
"https://api.obis.org/v3/"
source
OBISClient.configure! — Function
configure!(; kwargs...) -> ClientConfig

Change client settings. Every keyword matches a field of ClientConfig; omitted keywords keep their current value.

Raising page_size reduces the number of requests a large pull makes, which is the courteous direction. Parallelism is deliberately not offered: the OBIS data access page asks users not to parallelize downloads.

Examples

# Fewer, larger requests for a long pull.
OBISClient.configure!(page_size = 10_000, progress = true)

# Cache raw responses so an analysis can be re-run against identical data.
OBISClient.configure!(cache = OBISClient.QueryCache())

# Wait longer between requests on a shared or metered connection.
OBISClient.configure!(request_gap = 0.5)
source
OBISClient.ClientConfig — Type
ClientConfig

Runtime settings for every request the package makes. Read the active configuration with config and change it with configure!.

Fields

  • base_url: API base URL.
  • export_url: base URL of the full-export bucket.
  • user_agent: sent on every request, identifying package, version and repository.
  • timeout: per-request timeout in seconds.
  • retries: how many times to retry a retryable failure (429 and 5xx).
  • backoff_base: first backoff interval in seconds; doubles per attempt.
  • backoff_max: ceiling for a single backoff interval, in seconds.
  • page_size: records per page when paginating; capped at 10000.
  • request_gap: minimum seconds between consecutive requests.
  • api_record_limit: query size above which occurrence refuses to page and points at the export route instead.
  • cache: a QueryCache, or nothing to disable caching.
  • progress: show a progress indicator while paginating.
source
OBISClient.QueryCache — Type
QueryCache(dir = default_cache_dir(); readonly = false)

A local, content-addressed cache of raw API responses.

Each entry is keyed by a hash of the base URL, endpoint and query parameters, and stores the unparsed response body next to a metadata record giving the endpoint, the parameters, the retrieval timestamp and the response ETag. Storing the raw body means a re-run reproduces the original data even if this package's schema handling has changed since.

Entries never expire. That is the point: an expiring cache cannot support "re-run this analysis against the data as it stood in March". Remove entries explicitly with clear_cache!.

Examples

# Cache into the default location for the rest of the session.
OBISClient.configure!(cache = OBISClient.QueryCache())

# Keep an analysis's data alongside the analysis.
OBISClient.configure!(cache = OBISClient.QueryCache("data/obis-cache"))

# Re-run later against exactly what was fetched, refusing to reach the network.
OBISClient.configure!(cache = OBISClient.QueryCache("data/obis-cache"; readonly = true))
source
OBISClient.cache_entries — Function
cache_entries(cache = config().cache) -> Vector{Dict{String,Any}}

List what a cache holds: endpoint, parameters, retrieval time and size for each entry.

for e in OBISClient.cache_entries()
    println(e["endpoint"], "  ", e["retrieved"])
end
source
OBISClient.clear_cache! — Function
clear_cache!(cache = config().cache) -> Int

Delete every entry in a cache and return how many were removed.

source

Errors

OBISClient.OBISError — Type
OBISError

Abstract supertype for every error raised by OBISClient.jl. Catch this to handle any failure from the package without catching unrelated exceptions.

source
OBISClient.OBISValidationError — Type
OBISValidationError(parameter, value, message)

A request was rejected locally, before it reached the API.

The API ignores unknown query parameters and returns an empty result set for many invalid values instead of reporting an error, so a malformed request would otherwise come back as a plausible-looking empty answer. Validating locally turns that silence into this error.

source
OBISClient.OBISNameNotFoundError — Type
OBISNameNotFoundError(name)

The API matched no taxon for a scientificname.

This is reported separately because the API signals it with NAME_NOT_FOUND inside an HTTP 200 response whose body is an ordinary empty result set. Returning that empty set would let a misspelled name read as a genuine absence of records.

source
OBISClient.OBISAPIError — Type
OBISAPIError(status, endpoint, message, body)

The API reported a failure, either through an HTTP status or through an error key in an otherwise well-formed response body.

source
OBISClient.OBISConnectionError — Type
OBISConnectionError(endpoint, message)

The request never produced a response: a network failure, a timeout, or a retry budget exhausted against repeated server errors.

source
OBISClient.OBISLargeQueryError — Type
OBISLargeQueryError(estimate, threshold, suggestion)

A query was estimated to return more records than the configured API threshold.

Raised rather than silently switching access routes: the two routes differ in coverage (see export_covers), so the choice belongs to the caller.

source