Public API
Everything a caller needs. Functions the package uses internally are in Internals.
The module
OBISClient — Module
OBISClientA Julia client for OBIS, the Ocean Biodiversity Information System, a programme of the Intergovernmental Oceanographic Commission of UNESCO.
Results are Tables.jl tables with a fixed, concretely typed schema, and they carry the licence and citation of every dataset they draw on together with the date they were retrieved.
This is an independent, community-maintained client. It is not affiliated with, endorsed by, or maintained by OBIS, the Intergovernmental Oceanographic Commission, or UNESCO. The name contracts Ocean Biodiversity Information System, the service the package connects to; the General registry does not accept OBIS itself, being four characters and all upper case.
Getting started
using OBISClient, DataFrames
recs = OBISClient.occurrence("Abra alba"; limit = 1000)
df = DataFrame(recs)
OBISClient.licenses(recs) # may these data be redistributed?
print(OBISClient.citations(recs; format = :text)) # how to credit themInterpreting the results
Record counts measure sampling effort as much as biology, and the default view of OBIS excludes two classes of record that change what a result means. See the "Interpreting OBIS data" section of the documentation, and the disclaimer in OBIS_DISCLAIMER, before drawing conclusions.
Occurrences
OBISClient.occurrence — Function
occurrence(scientificname = nothing; kwargs...) -> OBISTableRetrieve occurrence records.
Returns an OBISTable: a Tables.jl-compatible table whose columns are fixed by the schema, so the same columns appear in the same order for every query, with missing where a record carries no value. Darwin Core fields outside the core schema are preserved per row in the extra column.
Each row carries the licence and citation of the dataset it came from, so citations and licenses work on the result without a further query.
Filters
scientificname: taxon name at any rank. Also accepted positionally.taxonid(aliasaphiaid): WoRMS AphiaID. Unambiguous where a name is not.geometry: WKT, for example"POLYGON ((2.3 51.8, 2.3 51.6, 2.6 51.6, 2.6 51.8, 2.3 51.8))".startdate,enddate:Date,DateTime, or"YYYY-MM-DD".startdepth,enddepth: metres below the surface.datasetid,nodeid: UUIDs.instituteid: OceanExpert ID.areaid: OBIS area ID.flags: quality flags that must be set.exclude: flags to exclude. Either case is accepted and normalized.absence,dropped,event::exclude(default),:include, or:only.redlist,hab,wrims: restrict to IUCN Red List, harmful algal bloom, or WRiMS species.measurementtype,measurementvalue,measurementunitand their*idforms: require a matching measurement.hasextensions: require an extension, for example"DNADerivedData".
Retrieval
limit: stop after this many records. Records arrive ordered by UUID, which is unrelated to date, place or taxon, so a limited pull is an arbitrary subset of the matches rather than the earliest or nearest ones.after: resume from a record UUID. Seeoccurrence_pages.licenses: look up dataset rights.trueby default; setfalseto skip the extra request, which leaves thelicensecolumnsmissingwithout changing the schema.check_size: estimate the query first and refuse to page through very large results. Seeestimate_size.
Examples
# A species, with rights attached.
recs = OBISClient.occurrence("Abra alba"; limit = 500)
# Absence records: observations where the species was looked for and not found. These are
# available through the API and not through the bulk downloads.
absences = OBISClient.occurrence("Abra alba"; absence = :only)
# Records the quality pipeline dropped, to see what a default query leaves out.
dropped = OBISClient.occurrence("Abra alba"; dropped = :only)
# An area and a period.
recs = OBISClient.occurrence(;
scientificname = "Delphinidae",
geometry = "POLYGON ((2.3 51.8, 2.3 51.6, 2.6 51.6, 2.6 51.8, 2.3 51.8))",
startdate = Date(2010, 1, 1),
enddate = Date(2020, 12, 31),
)
# Exclude records the pipeline flagged as being on land.
clean = OBISClient.occurrence("Abra alba"; exclude = "ON_LAND", limit = 1000)OBISClient.occurrence_pages — Function
occurrence_pages(scientificname = nothing; kwargs...) -> OccurrencePagesBuild a lazy, resumable iterator over pages of occurrence records.
Accepts the same filters as occurrence. Nothing is fetched until iteration begins, and each iteration yields one page as an OBISTable, so a query larger than memory can be processed page by page.
Extra keywords: page_size sets records per request (up to 10000); after resumes from a record UUID; licenses attaches dataset rights to every page, which costs one extra request per page and is therefore off by default.
Examples
# Summarize a large query without materializing it.
counts = Dict{String,Int}()
for page in OBISClient.occurrence_pages(scientificname = "Mollusca"; page_size = 10_000)
for name in page.scientificName
ismissing(name) || (counts[name] = get(counts, name, 0) + 1)
end
end
# Checkpoint after every page so an interrupted pull can resume.
pages = OBISClient.occurrence_pages(scientificname = "Mollusca")
for page in pages
save(page)
write("obis.cursor", OBISClient.cursor(pages))
endOBISClient.occurrence_by_id — Function
occurrence_by_id(id) -> OBISTableFetch a single occurrence record by its OBIS record UUID.
OBISClient.OccurrencePages — Type
OccurrencePagesA lazy iterator over pages of occurrence records.
Each iteration yields an OBISTable holding one page. Nothing is fetched until iteration starts, and only one page is held in memory at a time, so a query far larger than memory can still be processed.
The cursor is a record UUID. Read it with cursor after any page to checkpoint progress, and pass it back as after to resume — in the same session or a later one.
Construct with occurrence_pages.
Examples
# Process a large query without holding it in memory.
total = 0
for page in OBISClient.occurrence_pages(scientificname = "Mollusca")
total += OBISClient.nrow(page)
end
# Checkpoint, then resume in another session.
pages = OBISClient.occurrence_pages(scientificname = "Mollusca")
for page in pages
process(page)
write("checkpoint.txt", OBISClient.cursor(pages))
break
end
resumed = OBISClient.occurrence_pages(
scientificname = "Mollusca", after = read("checkpoint.txt", String)
)OBISClient.cursor — Function
cursor(pages::OccurrencePages) -> Union{Nothing,String}The UUID of the last record yielded so far, or nothing before the first page.
Persist this to resume a pull later: it is the whole of the pagination state.
OBISClient.fetched — Function
fetched(pages::OccurrencePages) -> IntHow many records have been yielded so far.
OBISClient.expected — Function
expected(pages::OccurrencePages) -> IntHow many records the API reports as matching, known after the first page is fetched.
Taxa and checklists
OBISClient.taxon — Function
taxon(id) -> OBISTableLook up a taxon by WoRMS AphiaID or by exact scientific name.
An AphiaID is unambiguous; a name is not, and the API matches it exactly rather than fuzzily. There is no taxon search endpoint — use checklist to discover which taxa occur in a selection.
julia> t = OBISClient.taxon(141433); # by AphiaID
julia> t = OBISClient.taxon("Abra alba"); # by exact nameOBISClient.taxon_annotations — Function
taxon_annotations(; scientificname = nothing, kwargs...) -> VectorAnnotations the WoRMS team made on scientific names that OBIS could not match automatically.
These explain why a name carries NO_MATCH or a WORMS_ANNOTATION_* flag, and whether it can be fixed.
OBISClient.checklist — Function
checklist(scientificname = nothing; kwargs...) -> OBISTableBuild a species checklist for a selection.
Accepts the same filters as occurrence, and returns one row per taxon with the number of records supporting it. This is the way to find out which taxa occur somewhere; the record counts are counts of observations, which reflect how much sampling has happened there as much as what lives there.
# What has been recorded in this polygon?
list = OBISClient.checklist(;
geometry = "POLYGON ((2.3 51.8, 2.3 51.6, 2.6 51.6, 2.6 51.8, 2.3 51.8))"
)
# Red List species in the same area.
OBISClient.checklist_redlist(; geometry = "POLYGON ((2.3 51.8, 2.3 51.6, 2.6 51.6, 2.6 51.8, 2.3 51.8))")OBISClient.checklist_redlist — Function
checklist_redlist(scientificname = nothing; kwargs...) -> OBISTableChecklist restricted to IUCN Red List species.
OBISClient.checklist_newest — Function
checklist_newest(scientificname = nothing; kwargs...) -> OBISTableChecklist of the most recently added species in a selection.
Datasets, nodes, institutes, areas, countries
OBISClient.dataset — Function
dataset(scientificname = nothing; kwargs...) -> OBISTableFind datasets matching a query.
Accepts the same filters as occurrence, so dataset describes exactly the datasets an occurrence query draws on. Each row carries the provider's citation string, the rights statement as published, and a normalized license identifier derived from it.
This endpoint does not paginate: every match is returned in one response. An unfiltered call therefore retrieves metadata for every dataset in OBIS, currently around seven thousand.
Examples
# Datasets behind a species.
ds = OBISClient.dataset("Abra alba")
# Which licences are in play, before downloading any occurrences.
OBISClient.licenses(ds)OBISClient.dataset_by_id — Function
dataset_by_id(id) -> OBISTableFetch one dataset by its UUID.
julia> ds = OBISClient.dataset_by_id("8acba7e7-2e50-4490-8328-b78a30472508");OBISClient.dataset_errors — Function
dataset_errors(id) -> VectorLoading errors OBIS recorded for a dataset, useful when a dataset holds fewer records than expected.
OBISClient.node — Function
node(id = nothing) -> OBISTableList OBIS nodes, or fetch one by UUID.
Nodes are the regional and thematic partners that curate the data. A node UUID is the nodeid filter accepted elsewhere.
julia> nodes = OBISClient.node(); # all nodes
julia> n = OBISClient.node("4bf79a01-65a9-4db6-b37b-18434f26ddfc");OBISClient.node_activities — Function
node_activities(id) -> VectorActivities recorded for a node.
OBISClient.institute — Function
institute(scientificname = nothing; kwargs...) -> OBISTableFind institutes contributing to a selection.
Identifiers are OceanExpert IDs, which is what the instituteid filter expects elsewhere.
OBISClient.institute_by_id — Function
institute_by_id(id) -> OBISTableFetch one institute by its OceanExpert ID.
OBISClient.area — Function
area(id = nothing) -> OBISTableList the areas OBIS can filter by, or fetch one by ID.
Areas include exclusive economic zones, marine protected areas and other named regions. An area ID is what the areaid filter expects.
julia> areas = OBISClient.area();
julia> belgian = [r for r in areas.name if occursin("Belg", r)]OBISClient.country — Function
country(id = nothing) -> OBISTableList the countries OBIS can filter by, or fetch one by ID.
OBISClient.metrics — Function
metrics(; datasetid = nothing, nodeid = nothing) -> AnyYearly download counts for a dataset or a node.
Provide one or the other. When both are given the API uses datasetid.
OBISClient.metrics_downloads — Function
metrics_downloads(datasetid; startdate = nothing, enddate = nothing, groupby = nothing)Download counts for a dataset, optionally within a date range.
Pass groupby = "time" for one entry per download event instead of an aggregate.
Statistics and facets
OBISClient.statistics — Function
statistics(scientificname = nothing; kwargs...) -> Dict{String,Any}Summary counts for a query: records, species, taxa, datasets, and the year range.
Accepts the same filters as occurrence and counts exactly the records that query would return, so it is the cheap way to size a query before running it.
julia> st = OBISClient.statistics("Abra alba");
julia> st["records"], st["datasets"]
(75356, 216)
julia> OBISClient.statistics("Abra alba"; absence = :only)["records"]
18387OBISClient.statistics_years — Function
statistics_years(scientificname = nothing; kwargs...) -> Vector{Dict{String,Any}}Number of presence records per year.
Useful for showing what record counts actually measure. A rise in a series like this tracks survey programmes, digitization projects and the growth of OBIS itself; it is not evidence that a taxon became more abundant.
julia> years = OBISClient.statistics_years("Abra alba");
julia> first(years)
Dict{String, Any}("year" => 1841, "records" => 8)OBISClient.statistics_env — Function
statistics_env(scientificname = nothing; kwargs...) -> AnyRecord counts per sea surface temperature, salinity, or depth bin.
OBISClient.statistics_qc — Function
statistics_qc(scientificname = nothing; kwargs...) -> Dict{String,Any}Quality summary for a query: missing and invalid fields, records on land, non-marine records, and records without an AphiaID.
Worth running before an analysis. It reports what a default query silently excludes and what the records it does return are flagged for.
OBISClient.statistics_composition — Function
statistics_composition(scientificname = nothing; kwargs...) -> AnyTaxonomic composition of a selection.
OBISClient.facet — Function
facet(facets; kwargs...) -> Dict{String,Vector}Record counts grouped by one or more fields.
facets names the fields, for example "flags", "datasetName" or ["originalScientificName", "flags"].
The API truncates each facet's value list, so a facet result is a ranking of the most common values, not a complete enumeration. In particular the flags facet does not list every flag OBIS uses; OBISClient.KNOWN_FLAGS is the fuller reference.
julia> f = OBISClient.facet("flags"; scientificname = "Abra alba");
julia> first(f["flags"])
Dict{String, Any}("key" => "NO_DEPTH", "records" => 29772)OBISClient.estimate_size — Function
estimate_size(; kwargs...) -> IntNumber of records a query would return, from one /statistics request.
julia> OBISClient.estimate_size(scientificname = "Mollusca") > 1_000_000
trueResults
OBISClient.OBISTable — Type
OBISTableA typed, column-oriented table of OBIS records that implements the Tables.jl interface.
The column set is fixed by the schema for the endpoint, so two queries against the same endpoint always produce the same columns in the same order, with missing where a record carried no value. Fields outside the core schema are preserved per row in the extra column as a Dict{Symbol,Any}.
Columns are reached by property access, by index, or through any Tables.jl consumer:
tbl.scientificName # a column
tbl[:decimalLatitude] # the same column
DataFrame(tbl) # with DataFrames loaded
CSV.write("out.csv", tbl) # with CSV loadedUse citations and licenses on a table to obtain the attribution its data requires, and metadata to recover the query and access date.
OBISClient.QueryMeta — Type
QueryMetaProvenance of a result: what was asked, of which service, and when.
The access date is recorded at request time because the OBIS citation format requires Accessed: YYYY-MM-DD and nothing in the data itself supplies it. Carrying it on the result is what lets citations produce a correct citation with no bookkeeping by the caller.
Fields
endpoint: API path the records came from.params: query parameters as sent, in canonical order.accessed: date of retrieval, for citations.retrieved: UTC timestamp of retrieval.total: number of records the API reported as matching, which may exceed the number held here when a query was limited or paged partially.base_url: service the records came from.package_version: version of this package that fetched them.
OBISClient.metadata — Function
metadata(t::OBISTable) -> QueryMetaReturn the query, access date and reported total behind a result.
julia> meta = OBISClient.metadata(tbl);
julia> meta.accessed # the date to put in a citation
2026-09-06OBISClient.nrow — Function
nrow(t::OBISTable) -> IntNumber of records held in the table.
This can be smaller than metadata(t).total, which is how many records the API reported as matching the query.
OBISClient.ncol — Function
ncol(t::OBISTable) -> IntNumber of columns, including extra.
OBISClient.extra_names — Function
extra_names(t::OBISTable) -> Vector{Symbol}Names of the non-core fields present anywhere in the table, sorted.
These are the Darwin Core terms the providers supplied that fall outside the core schema. Which of them appear depends on the query, which is exactly why they are held apart from the core columns.
julia> OBISClient.extra_names(tbl)
12-element Vector{Symbol}:
:day
:eventTime
⋮OBISClient.extra_column — Function
extra_column(t::OBISTable, name) -> VectorLift one non-core field out of extra into a plain column, with missing where a record did not carry it.
julia> OBISClient.extra_column(tbl, :waterBody)Licensing and citation
OBISClient.licenses — Function
licenses(t::OBISTable) -> OBISTableSummarize which licences a result is under.
One row per distinct licence, with how many datasets and records fall under it and what it permits. Read it before redistributing: a single CC BY-NC dataset makes the combined result non-commercial, and "unknown" covers rights statements that could not be identified — among them the literal value Restricted — which are not permission to redistribute.
julia> OBISClient.licenses(OBISClient.occurrence("Abra alba"; limit = 500))
OBISTable: 3 records, 7 columnsOBISClient.citations — Function
citations(t::OBISTable; format = :table)Build the citations required for the data in a result.
Returns one entry per dataset the records came from, with the provider's citation string, its DOI where one exists, its licence, how many records in the result it contributed, and a ready-to-use citation in the format the OBIS data policy specifies — including the access date recorded when the data was retrieved.
Formats:
:table— anOBISTable, for further processing or writing to a file.:text— the formatted citations as one string, ready to paste.:bibtex— a BibTeX document, one@miscentry per dataset.
The dataset citation is not the only obligation. Where a result draws on many datasets, the policy also allows citing the OBIS database as a whole; obis_citation builds that.
Examples
recs = OBISClient.occurrence("Abra alba"; limit = 1000)
cites = OBISClient.citations(recs) # a table, one row per dataset
print(OBISClient.citations(recs; format = :text))
write("references.bib", OBISClient.citations(recs; format = :bibtex))OBISClient.obis_citation — Function
obis_citation(; description = nothing, accessed = today(), year = Dates.year(today()))The citation for the OBIS database as a whole.
The data policy allows citing the integrated database in addition to — never instead of — the individual datasets, whose own restrictions continue to apply.
julia> OBISClient.obis_citation(; description = "Distribution records of Abra alba", accessed = Date(2026, 9, 6), year = 2026)
"OBIS (2026) Distribution records of Abra alba [Dataset] (Available: Ocean Biodiversity Information System. Intergovernmental Oceanographic Commission of UNESCO. https://obis.org. Accessed: 2026-09-06)"OBISClient.normalize_license — Function
normalize_license(text) -> StringDerive a licence identifier from the free-text rights statement OBIS publishes.
Returns one of "CC0-1.0", "CC-BY-4.0", "CC-BY-NC-4.0", "CC-BY-SA-4.0", or "unknown". Both prose statements and licence URLs are accepted, so a value taken from the API and one taken from the export's licence table normalize the same way.
Non-commercial is tested before attribution: every CC-BY-NC statement also contains the word "Attribution", so the looser test would swallow the stricter licence.
julia> OBISClient.normalize_license("This work is licensed under a Creative Commons Attribution (CC-BY) 4.0 License")
"CC-BY-4.0"
julia> OBISClient.normalize_license("This work is licensed under a Creative Commons Attribution Non Commercial (CC-BY-NC) 4.0 License")
"CC-BY-NC-4.0"
julia> OBISClient.normalize_license("http://creativecommons.org/publicdomain/zero/1.0/legalcode")
"CC0-1.0"
julia> OBISClient.normalize_license("Restricted")
"unknown"OBISClient.license_url — Function
license_url(id) -> Union{Missing,String}Canonical URL for a licence identifier, or missing when it is not recognized.
julia> OBISClient.license_url("CC-BY-NC-4.0")
"https://creativecommons.org/licenses/by-nc/4.0/"
julia> OBISClient.license_url("unknown")
missingOBISClient.permits_redistribution — Function
permits_redistribution(id) -> BoolWhether a licence allows redistributing the data at all.
True for every recognized Creative Commons licence here. It is false for "unknown", because an unidentifiable rights statement is not permission — the observed unknowns include the literal string Restricted.
OBISClient.permits_commercial_use — Function
permits_commercial_use(id) -> BoolWhether a licence permits commercial reuse.
False for CC-BY-NC-4.0 and for "unknown". A result set containing even one non-commercial dataset cannot be redistributed commercially as a whole, which is why licenses reports the mix rather than a single verdict.
OBISClient.requires_attribution — Function
requires_attribution(id) -> BoolWhether a licence requires the data provider to be credited.
False only for CC0-1.0, the one licence here that waives the requirement. "unknown" counts as requiring attribution: an unidentifiable rights statement is not evidence that attribution was waived, and treating it as such is exactly the permissive guess the rest of this file avoids.
This describes the legal obligation. The OBIS data policy asks that providers be credited regardless of licence, so false is not advice to omit the credit.
julia> OBISClient.requires_attribution("CC0-1.0")
false
julia> OBISClient.requires_attribution("CC-BY-4.0")
true
julia> OBISClient.requires_attribution("unknown")
trueOBISClient.ACCEPTED_LICENSES — Constant
ACCEPTED_LICENSESThe three licences the OBIS data policy accepts for published datasets.
Anything else in a result is either an older record, a provider error, or a licence OBIS does not endorse, and is worth the caller's attention.
OBISClient.LICENSE_URLS — Constant
LICENSE_URLSCanonical URL for each licence identifier the package recognizes.
Quality flags
OBISClient.KNOWN_FLAGS — Constant
KNOWN_FLAGSQuality flags documented by the OBIS QC pipeline, together with flags observed live that the pipeline reference does not list.
The vocabulary is open: OBIS adds checks over time, and the /facet endpoint truncates its flag listing, so this set cannot be treated as exhaustive. Unrecognized upper-case flags are passed through with a warning rather than rejected.
OBISClient.DROPPING_FLAGS — Constant
DROPPING_FLAGSFlags whose presence causes OBIS to drop a record from the main index.
Records carrying these are reachable only with dropped = :only or dropped = :include, because a dropped record is excluded from every default query and from the bulk exports.
OBISClient.normalize_flag — Function
normalize_flag(flag) -> StringReturn flag in the upper-case form the API matches on, warning once if the flag is not in KNOWN_FLAGS.
Unknown flags are passed through rather than rejected: OBIS adds quality checks over time, and a hard allow-list would reject a valid new flag.
julia> OBISClient.normalize_flag("on_land")
"ON_LAND"
julia> OBISClient.normalize_flag(:no_depth)
"NO_DEPTH"OBISClient.normalize_flags — Function
normalize_flags(flags) -> Vector{String}Normalize a flag, or any iterable of flags, into the upper-case forms the API expects.
Accepts a String, a Symbol, or any iterable of those. A comma-separated string is split, so both "ON_LAND,NO_DEPTH" and ["ON_LAND", "NO_DEPTH"] work.
julia> OBISClient.normalize_flags(["on_land", :NO_DEPTH])
2-element Vector{String}:
"ON_LAND"
"NO_DEPTH"
julia> OBISClient.normalize_flags("on_land,no_match")
2-element Vector{String}:
"ON_LAND"
"NO_MATCH"OBISClient.drops_record — Function
drops_record(flag) -> BoolWhether a record carrying flag is dropped from the OBIS main index.
julia> OBISClient.drops_record("NO_MATCH")
true
julia> OBISClient.drops_record("on_land")
falseThe bulk export
OBISClient.export_covers — Function
export_covers(; absence = nothing, dropped = nothing, event = nothing, kwargs...) -> Bool
export_covers(params::QueryParams) -> BoolWhether the bulk export can serve a query.
false only for a query that selects pure event records. The export has no column that identifies them — its event information is a _event_id carried on each occurrence, not a record class — so there is nothing in a file to select on.
Absence and dropped records are in the export, contrary to the OBIS data access page, which says exports carry neither. Verified per dataset against /statistics: across five datasets the export's absence and dropped row counts matched the API's absence = :only and dropped = :only counts exactly, and each export's total equalled the default count plus them (NOTES.md §7.3). Reading them out of an export is a filter on those two columns.
What still differs between the routes is the date. The export is regenerated periodically and the API is live, so the same query answered from each will differ by whatever OBIS has ingested since — which is why a large query raises rather than switching route on its own.
OBISClient.export_url — Function
export_url(dataset_id) -> StringURL of the GeoParquet export for one dataset.
julia> OBISClient.export_url("8acba7e7-2e50-4490-8328-b78a30472508")
"https://obis-open-data.s3.amazonaws.com/occurrence/8acba7e7-2e50-4490-8328-b78a30472508.parquet"OBISClient.export_archive_url — Function
export_archive_url(dataset_id) -> StringURL of the source Darwin Core Archive for one dataset.
The archive holds the data exactly as the provider published it, without the fields OBIS's quality pipeline adds.
OBISClient.download_export — Function
download_export(dataset_id; dir = pwd(), overwrite = false) -> StringDownload one dataset's GeoParquet export and return the local path.
path = OBISClient.download_export("8acba7e7-2e50-4490-8328-b78a30472508"; dir = "data")Read it with whichever Parquet reader you prefer; the package does not impose one. The schema is nested — provider terms under source, pipeline-interpreted terms under interpreted — and so differs from the flat schema the API returns.
OBISClient.download_exports — Function
download_exports(scientificname = nothing; dir = pwd(), overwrite = false, kwargs...)Download the GeoParquet exports for every dataset a query touches.
Resolves the query to its datasets with dataset, then fetches one file per dataset, serially. Returns the local paths.
One file per dataset, whole: the export is not filtered, so a query's filters are applied by you after reading. Absence and dropped records are present and carry their own columns; pure event records cannot be selected at all. export_covers reports which case a query is in.
paths = OBISClient.download_exports("Abra alba"; dir = "obis-export")OBISClient.export_licenses — Function
export_licenses(; path = nothing) -> OBISTableFetch the licence table published alongside the bulk export.
One row per dataset, with the rights statement, its canonical URL and the provider's citation. This is a companion to the export files rather than a live view of the API: it is regenerated with the export, so it lags the current dataset list. For a live query, take licences from dataset or from an occurrence result.
Pass path to read a previously downloaded copy instead of fetching it.
OBISClient.read_export — Function
read_export(path; absence = :exclude, dropped = :exclude, licenses = true,
limit = nothing, source_terms = false) -> OBISTableRead one or more GeoParquet export files into the package's canonical occurrence schema.
path is a file, or a vector of them — so read_export(download_exports(...)) composes.
Requires DuckDB: using DuckDB loads the extension that implements this. Parquet2.jl cannot read these files, whose nested source and interpreted structs it does not support.
The result is an OBISTable with the same columns, in the same order and with the same types, as one from occurrence, so the two access routes become interchangeable in everything downstream — licenses and citations included. Two columns are always missing: the export carries marine and brackish but not freshwater or terrestrial.
absence and dropped default to :exclude, matching the API, because the export contains those records (NOTES.md §7.3) and a plain read would otherwise mix them into an ordinary count without saying so. :include and :only mean what they mean everywhere else, and the filter is pushed into the read rather than applied afterwards.
licenses = true fetches export_licenses once. Pass a table returned by it to reuse across many files, or false to leave the three rights columns missing.
source_terms = true fills the extra column with the provider's own terms from source. It is off by default: that block has 188 fields, and reading them costs far more than the core schema does.
The access date on the result is the file's modification time, not today — the data is as old as the export, and a citation built from it should say so.
Configuration and caching
OBISClient.config — Function
config() -> ClientConfigReturn the active client configuration.
julia> OBISClient.config().base_url
"https://api.obis.org/v3/"OBISClient.configure! — Function
configure!(; kwargs...) -> ClientConfigChange client settings. Every keyword matches a field of ClientConfig; omitted keywords keep their current value.
Raising page_size reduces the number of requests a large pull makes, which is the courteous direction. Parallelism is deliberately not offered: the OBIS data access page asks users not to parallelize downloads.
Examples
# Fewer, larger requests for a long pull.
OBISClient.configure!(page_size = 10_000, progress = true)
# Cache raw responses so an analysis can be re-run against identical data.
OBISClient.configure!(cache = OBISClient.QueryCache())
# Wait longer between requests on a shared or metered connection.
OBISClient.configure!(request_gap = 0.5)OBISClient.ClientConfig — Type
ClientConfigRuntime settings for every request the package makes. Read the active configuration with config and change it with configure!.
Fields
base_url: API base URL.export_url: base URL of the full-export bucket.user_agent: sent on every request, identifying package, version and repository.timeout: per-request timeout in seconds.retries: how many times to retry a retryable failure (429 and 5xx).backoff_base: first backoff interval in seconds; doubles per attempt.backoff_max: ceiling for a single backoff interval, in seconds.page_size: records per page when paginating; capped at 10000.request_gap: minimum seconds between consecutive requests.api_record_limit: query size above whichoccurrencerefuses to page and points at the export route instead.cache: aQueryCache, ornothingto disable caching.progress: show a progress indicator while paginating.
OBISClient.QueryCache — Type
QueryCache(dir = default_cache_dir(); readonly = false)A local, content-addressed cache of raw API responses.
Each entry is keyed by a hash of the base URL, endpoint and query parameters, and stores the unparsed response body next to a metadata record giving the endpoint, the parameters, the retrieval timestamp and the response ETag. Storing the raw body means a re-run reproduces the original data even if this package's schema handling has changed since.
Entries never expire. That is the point: an expiring cache cannot support "re-run this analysis against the data as it stood in March". Remove entries explicitly with clear_cache!.
Examples
# Cache into the default location for the rest of the session.
OBISClient.configure!(cache = OBISClient.QueryCache())
# Keep an analysis's data alongside the analysis.
OBISClient.configure!(cache = OBISClient.QueryCache("data/obis-cache"))
# Re-run later against exactly what was fetched, refusing to reach the network.
OBISClient.configure!(cache = OBISClient.QueryCache("data/obis-cache"; readonly = true))OBISClient.cache_entries — Function
cache_entries(cache = config().cache) -> Vector{Dict{String,Any}}List what a cache holds: endpoint, parameters, retrieval time and size for each entry.
for e in OBISClient.cache_entries()
println(e["endpoint"], " ", e["retrieved"])
endOBISClient.clear_cache! — Function
clear_cache!(cache = config().cache) -> IntDelete every entry in a cache and return how many were removed.
OBISClient.default_cache_dir — Function
default_cache_dir() -> StringDefault cache location, honouring XDG_CACHE_HOME where it is set.
OBISClient.package_version — Function
package_version() -> VersionNumberVersion of the installed package, used in the User-Agent header.
Errors
OBISClient.OBISError — Type
OBISErrorAbstract supertype for every error raised by OBISClient.jl. Catch this to handle any failure from the package without catching unrelated exceptions.
OBISClient.OBISValidationError — Type
OBISValidationError(parameter, value, message)A request was rejected locally, before it reached the API.
The API ignores unknown query parameters and returns an empty result set for many invalid values instead of reporting an error, so a malformed request would otherwise come back as a plausible-looking empty answer. Validating locally turns that silence into this error.
OBISClient.OBISNameNotFoundError — Type
OBISNameNotFoundError(name)The API matched no taxon for a scientificname.
This is reported separately because the API signals it with NAME_NOT_FOUND inside an HTTP 200 response whose body is an ordinary empty result set. Returning that empty set would let a misspelled name read as a genuine absence of records.
OBISClient.OBISAPIError — Type
OBISAPIError(status, endpoint, message, body)The API reported a failure, either through an HTTP status or through an error key in an otherwise well-formed response body.
OBISClient.OBISConnectionError — Type
OBISConnectionError(endpoint, message)The request never produced a response: a network failure, a timeout, or a retry budget exhausted against repeated server errors.
OBISClient.OBISLargeQueryError — Type
OBISLargeQueryError(estimate, threshold, suggestion)A query was estimated to return more records than the configured API threshold.
Raised rather than silently switching access routes: the two routes differ in coverage (see export_covers), so the choice belongs to the caller.