Open Biodiversity Databases

From WikiDemocracy
Jump to navigationJump to search


    • NOTOC**

Open Biodiversity Databases

Open biodiversity databases are transforming how scientists, conservation organizations, governments, and the public understand life on Earth. Millions of observations, museum specimens, scientific names, genetic sequences, photographs, ecological traits, distribution records, conservation assessments, and historical publications that were once scattered among individual institutions can increasingly be searched and analyzed through interconnected digital systems.

These databases form part of a growing global biodiversity information infrastructure. Rather than functioning as a single universal database, the system consists of many specialized repositories connected through common standards, identifiers, software tools, and data-sharing agreements.

Major platforms include the Global Biodiversity Information Facility (GBIF), Ocean Biodiversity Information System (OBIS), iNaturalist, eBird, Catalogue of Life, Plants of the World Online, World Flora Online, Barcode of Life Data Systems (BOLD), Biodiversity Heritage Library, Atlas of Living Australia, IUCN Red List, Protected Planet, and hundreds of regional and taxonomic databases.

Together, these resources allow researchers to investigate where species occur, how their distributions are changing, which species are threatened, how organisms are related, and how ecosystems respond to environmental change.

A Global Infrastructure for Biodiversity Knowledge

The growth of open biodiversity databases represents a fundamental change in biological science. Historically, biodiversity information was distributed among museum collections, field notebooks, scientific publications, government surveys, university archives, and individual researchers.

Digitization has allowed much of this information to become searchable across institutional and national boundaries.

The Global Biodiversity Information Facility is one of the largest examples of this approach. GBIF aggregates species-occurrence records contributed by museums, research institutions, governments, citizen-science projects, universities, and conservation organizations throughout the world.

These records may document preserved specimens collected centuries ago, modern field surveys, photographs submitted through citizen-science projects, environmental observations, or other evidence that a particular organism occurred at a particular place and time.

Rather than replacing individual databases, global infrastructures frequently connect them.

National portals such as the Atlas of Living Australia, regional systems, natural-history museums, citizen-science networks, and specialist databases can publish information into larger networks while continuing to maintain their own collections and expertise.

This distributed model allows biodiversity information to remain connected to the institutions and communities that produce it while making it available for global research.

Standards Make Databases Interoperable

Open data become considerably more useful when different databases describe information in compatible ways.

One of the most important biodiversity-data standards is Darwin Core. It provides standardized terms for describing species observations, specimens, taxonomy, geographic locations, collecting events, dates, identifiers, and related information.

A museum in Australia, a citizen scientist in Kenya, and a university research project in the United States can therefore describe biodiversity records using many of the same standardized fields.

This makes it possible for large aggregators to combine records originating from thousands of independent sources.

Data standards are increasingly expanding beyond simple occurrence records. Modern biodiversity research may involve ecological surveys, DNA sequences, environmental measurements, species interactions, functional traits, genomic samples, photographs, sound recordings, and environmental DNA.

New data models and interoperability initiatives are attempting to preserve this complexity while still allowing information to be exchanged among databases.

The broader FAIR principles — making data Findable, Accessible, Interoperable, and Reusable — have consequently become central to biodiversity informatics.

Citizen Science Has Become a Major Data Source

Open biodiversity databases increasingly depend on observations collected by members of the public.

Platforms such as iNaturalist allow people to photograph organisms, record locations and dates, suggest identifications, and collaborate with other users to determine species identities.

Large numbers of qualifying iNaturalist observations become research-grade records and can subsequently be incorporated into global biodiversity infrastructures such as GBIF.

eBird, operated by the Cornell Lab of Ornithology, has developed another major model for community biodiversity monitoring. Birdwatchers around the world submit checklists recording the birds they observe, where they observed them, when they surveyed, and information about their observation effort.

These enormous datasets allow researchers to model bird distributions, migration, seasonal abundance, and long-term population trends at geographic scales that would be extremely expensive to monitor using traditional field programs alone.

Citizen science nevertheless introduces important methodological challenges. Participants do not sample landscapes randomly. People are more likely to visit accessible areas, populated regions, parks, roads, and locations known for interesting species.

Researchers therefore use statistical models, validation systems, observer calibration, sampling protocols, and data-quality filters to compensate for these biases.

Specialized Databases Document Different Dimensions of Life

No single database can adequately represent every dimension of biodiversity.

For this reason, specialized databases have developed for different taxonomic groups, ecosystems, genetic information, ecological characteristics, and research purposes.

Marine biodiversity is supported by resources such as the World Register of Marine Species (WoRMS) and Ocean Biodiversity Information System (OBIS). WoRMS provides an expert-curated taxonomic foundation for marine organisms, while OBIS aggregates marine species-occurrence records from around the world.

Plant diversity is documented through resources including Plants of the World Online, World Flora Online, the International Plant Names Index, the World Checklist of Vascular Plants, and plant-trait databases such as TRY and BIEN.

Individual animal groups have their own specialist systems. Examples include the Reptile Database, Mammal Diversity Database, World Spider Catalog, Orthoptera Species File, AmphibiaWeb, and AntWeb.

Fungi are represented through resources including Index Fungorum, UNITE, and other fungal nomenclatural and molecular databases.

These specialized databases often benefit from expert communities capable of evaluating taxonomic changes and maintaining authoritative scientific names.

DNA, Genomics and Environmental DNA

Biodiversity databases increasingly document organisms not only through physical observations but also through genetic information.

The Barcode of Life Data Systems (BOLD) connects standardized DNA barcode sequences with specimens, taxonomy, collection locations, photographs, and other biological information.

DNA barcoding can help researchers identify organisms using relatively short genetic sequences. This is particularly valuable for organisms that are difficult to distinguish visually, incomplete specimens, immature life stages, and environmental samples containing mixtures of many organisms.

Environmental DNA, or eDNA, has expanded this approach. Instead of directly observing an organism, researchers can collect water, soil, sediment, or other environmental material and search for genetic traces left behind by organisms.

Reference databases are essential to this process because unknown DNA sequences must be compared with previously identified sequences.

Resources such as BOLD, UNITE, SILVA, PR2, MIDORI, MetaZooGene, MitoFish, and the European Nucleotide Archive contribute different types of molecular reference information.

The usefulness of molecular biodiversity monitoring therefore depends heavily on the completeness and accuracy of these reference databases.

Natural-History Collections Enter the Digital Era

Natural-history museums and herbaria contain billions of physical specimens collected over centuries.

Historically, much of the information associated with these specimens could only be accessed by visiting collections or requesting information from curators.

Large digitization programs are changing this situation.

Platforms such as iDigBio, the Natural History Museum Data Portal, GBIF, Symbiota, Arctos, DiSSCo, and the Atlas of Living Australia make specimen information increasingly accessible online.

Digital records can contain specimen photographs, collection dates, geographic coordinates, collector names, taxonomy, genetic information, publications, and links to physical specimens.

Some systems also provide three-dimensional specimen imagery. MorphoSource, for example, provides digital media including 3-D scans derived from biological, paleontological, and cultural collections.

Digitized collections allow specimens collected decades or centuries ago to contribute to modern research on climate change, invasive species, geographic distributions, morphology, evolution, extinction, and environmental change.

Physical collections remain essential, however. Digital records generally represent information about specimens rather than substitutes for the specimens themselves.

Traits, Ecology and Species Interactions

Knowing that a species exists in a particular location is only one part of understanding biodiversity.

Researchers also need information about what organisms do.

Trait databases record characteristics such as body size, diet, morphology, reproductive strategies, habitat preferences, activity patterns, physiology, dispersal ability, and ecological functions.

Examples include TRY for plant traits, AVONET for birds, AmphiBIO for amphibians, Coral Trait Database, AnimalTraits, PalmTraits, EltonTraits, and numerous databases cataloged through the Open Traits Network.

Other systems document relationships among organisms.

Global Biotic Interactions (GloBI) integrates records describing interactions such as predation, parasitism, pollination, herbivory, and host relationships.

Combining occurrence records with traits and ecological interactions can help scientists move from simply mapping species toward understanding how ecosystems function.

Biodiversity Literature as Data

Centuries of biodiversity information remain embedded in books, journals, expedition reports, taxonomic descriptions, illustrations, and other scientific literature.

The Biodiversity Heritage Library (BHL) has digitized a vast collection of historical biodiversity literature and makes much of it openly accessible.

Projects such as BioStor, TreatmentBank, Global Names, Wikidata integrations, and biodiversity knowledge graphs attempt to extract structured information from this literature and connect scientific names, specimens, publications, researchers, and databases.

These efforts address an important challenge: biodiversity information is valuable not simply because it exists online but because different pieces of information can be linked together.

A species name in a nineteenth-century publication, for example, may eventually be connected with its modern accepted name, original description, museum specimens, DNA sequences, geographic observations, conservation assessment, and ecological traits.

The result is gradually becoming a network of biodiversity knowledge rather than a collection of isolated databases.

Conservation Applications

Open biodiversity databases have direct applications in conservation.

The IUCN Red List provides information about extinction risk, population trends, threats, habitats, distributions, and conservation status.

Protected Planet maintains global information about protected areas and other effective area-based conservation measures.

The World Database of Key Biodiversity Areas identifies locations considered particularly important for the persistence of global biodiversity.

BirdLife International's DataZone provides information about species, distributions, conservation assessments, and Important Bird and Biodiversity Areas.

Map of Life integrates numerous datasets to model species distributions, habitat relationships, biodiversity indicators, and conservation priorities.

When combined with satellite imagery, climate models, land-use information, and ecological monitoring, open biodiversity data can help identify areas where habitats are disappearing, species ranges are shifting, invasive species are spreading, or conservation intervention may be needed.

The Problem of Data Bias

Large databases can create an impression of comprehensive knowledge even when their records contain substantial gaps.

Biodiversity data are geographically uneven.

Regions with strong research institutions, historical collecting traditions, good transportation networks, and active citizen-science communities often have many more records than remote or economically disadvantaged regions.

Taxonomic coverage is also uneven. Birds and mammals may be documented extensively while many insects, fungi, microorganisms, marine invertebrates, and poorly studied plant groups remain inadequately represented.

Historical collecting practices can introduce additional biases. Records may cluster around cities, roads, research stations, protected areas, or locations visited by scientific expeditions.

Citizen-science databases can contain similar patterns because participants generally observe organisms where people live and travel.

The absence of a database record therefore does not necessarily indicate the absence of a species.

Understanding these biases is essential when biodiversity databases are used to estimate distributions, extinction risk, species richness, ecological change, or conservation priorities.

Data Quality and Cleaning

A database containing millions of records inevitably contains errors.

Geographic coordinates may be reversed, rounded, incorrectly entered, or accidentally placed at the headquarters of an institution rather than at the collection site.

Scientific names may be misspelled or outdated.

Dates can be incorrect.

Duplicate records may appear in multiple databases.

Species can also be misidentified.

Modern biodiversity infrastructures therefore employ automated and expert quality-control systems.

Specialized software such as CoordinateCleaner, bdclean, scrubr, rgbif, pygbif, taxize, and other biodiversity informatics tools allows researchers to retrieve, validate, clean, document, and analyze large biodiversity datasets reproducibly.

Quality control is particularly important because enormous datasets can magnify small systematic errors.

Open access does not eliminate the need for scientific judgment. Instead, it makes the methods used to evaluate and correct data increasingly important.

Open Does Not Always Mean Completely Unrestricted

The movement toward open biodiversity data generally aims to make scientific information widely accessible and reusable.

However, there are legitimate exceptions.

Publishing exact locations of extremely rare species can sometimes increase risks from poaching, illegal collection, disturbance, or wildlife trafficking.

Indigenous knowledge and data associated with Indigenous peoples may also involve governance, consent, ownership, and cultural rights that cannot be addressed simply by declaring information open.

Some records may be subject to legal, licensing, ethical, privacy, or conservation restrictions.

Modern biodiversity-data governance therefore increasingly recognizes the principle that information should be as open as possible and as restricted as necessary.

From Databases to a Global Biodiversity Knowledge Network

The future of biodiversity informatics is increasingly about connecting databases rather than simply creating larger isolated repositories.

Taxonomic databases can provide accepted species names.

Occurrence databases can show where organisms have been recorded.

Natural-history collections can preserve physical evidence.

DNA databases can provide molecular identifiers.

Trait databases can describe ecological characteristics.

Interaction databases can describe relationships among organisms.

Scientific literature can provide historical and conceptual context.

Conservation databases can identify threats and extinction risks.

When these resources use compatible standards and persistent identifiers, information from many independent systems can be linked into a larger biodiversity knowledge network.

Emerging technologies including artificial intelligence, automated species recognition, environmental DNA, remote sensing, biodiversity digital twins, knowledge graphs, and large-scale ecological modeling could dramatically increase the amount of biodiversity information available.

Their usefulness, however, will continue to depend on the quality, accessibility, transparency, and representativeness of the underlying data.

Why Open Biodiversity Databases Matter

Humanity is attempting to understand biodiversity while ecosystems are undergoing rapid environmental change.

No research institution, government, museum, or conservation organization can monitor the planet alone.

Open biodiversity databases allow information collected by thousands of organizations and millions of individuals to contribute to a shared scientific resource.

They also allow older data to acquire new value. A museum specimen collected a century ago can help document historical species distributions. A birdwatcher's checklist can contribute to population models. A DNA sequence can help identify organisms in environmental samples. A historical taxonomic publication can clarify the identity of a species described generations ago.

The greatest potential of biodiversity databases therefore lies not merely in accumulating records but in connecting observations across space, time, institutions, technologies, and scientific disciplines.

Conclusion

Open biodiversity databases are becoming part of the basic infrastructure of modern conservation and biological science.

Systems such as GBIF, OBIS, iNaturalist, eBird, Catalogue of Life, BOLD, Biodiversity Heritage Library, natural-history collection portals, trait databases, molecular repositories, and conservation platforms together provide an increasingly detailed digital representation of the living world.

Major challenges remain. Geographic and taxonomic gaps are substantial, data quality varies, historical and socioeconomic biases affect what has been recorded, and some biodiversity information requires careful governance rather than unrestricted publication.

Nevertheless, standardized and interoperable open data make it possible to combine information at a scale that would have been unimaginable only a few decades ago.

As biodiversity monitoring expands through citizen science, museum digitization, environmental DNA, automated sensors, satellites, artificial intelligence, and field research, open databases will become even more important.

Protecting biodiversity requires knowing what species exist, where they occur, how ecosystems function, and how those patterns are changing. Open biodiversity databases provide much of the information infrastructure needed to answer those questions.

    • TOC**



Global Biodiversity Databases and Data Infrastructure

| Doll | Nature Reviews Biodiversity | 2025

Discusses persistent global disparities in the collection and availability of biodiversity data and why filling geographic and taxonomic gaps matters for conservation.

| Feng et al. | Global Ecology and Biogeography | 2022

Reviews the fragmented landscape of major biodiversity databases and examines opportunities for integrating them into a more unified biodiversity knowledge base.

| Gadelha et al. | WIREs Data Mining and Knowledge Discovery | 2021

Surveys biodiversity informatics, including data collection, integration, standards, scientific workflows, data repositories, and computational approaches for analyzing biodiversity information.

| Khan, Thelwall and Kousha | Scientometrics | 2021

Investigates how biodiversity datasets are reused and cited, providing insight into the scientific impact of making primary biodiversity data openly available.

| Robertson et al. | PLOS ONE | 2014

Describes the GBIF Integrated Publishing Toolkit, which allows institutions to publish standardized biodiversity datasets through the Global Biodiversity Information Facility.

| Faith et al. | Biodiversity Informatics | 2013

Examines major gaps in available biodiversity information and recommends approaches for making biodiversity data more useful to researchers and decision-makers.

| Wieczorek et al. | PLOS ONE | 2012

Explains Darwin Core, the widely adopted community standard that allows biodiversity databases to exchange information about specimens, observations, taxa, locations, and events.

| Constable et al. | PLOS Biology | 2010

Introduces VertNet as a distributed system for making vertebrate specimen and observation records held by natural-history collections openly accessible.

| Boakes et al. | PLOS Biology | 2010

Demonstrates how biodiversity occurrence databases can contain substantial geographic and temporal biases that must be considered when interpreting their records.

| Arthur D. Chapman | GBIF Secretariat | 2005

Provides foundational guidance on biodiversity data quality, including accuracy, completeness, consistency, validation, documentation, and fitness for use.

GBIF, Darwin Core and Open Data Standards

| Lipiński | Conservation | 2026

Analyzes spatial biases in open GBIF occurrence records and shows how differences in record quality can affect biodiversity and conservation analyses.

| GBIF Secretariat | GBIF | 2026

Describes GBIF's developing data model for handling biodiversity information that is more complex than traditional species-occurrence records.

| Güntsch et al. | BioScience | 2025

Identifies essential functions that national biodiversity data infrastructures should provide to connect biodiversity information with science, policy, monitoring, and management.

| GBIF Secretariat | GBIF | 2025

Gives practical guidance for publishing structured biological survey and monitoring data so that sampling methods and observations remain interpretable after aggregation.

| Bloom et al. | GBIF Secretariat | 2021

Evaluates biodiversity information needs associated with the post-2020 global biodiversity framework and considers how open biodiversity infrastructure can address them.

| Meyer et al. | Global Ecology and Biogeography | 2016

Examines the geographic and socioeconomic factors that produce species-level biases in global occurrence databases and identifies major gaps in biodiversity knowledge.

| GBIF Secretariat | GBIF | n.d.

Explains what Darwin Core is, why biodiversity databases use it, and how standardized fields allow records from different institutions to be combined.

| GBIF Secretariat | GBIF | n.d.

Provides an overview of the standards used by GBIF to exchange occurrence, checklist, sampling-event, metadata, and other biodiversity information.

| GBIF Secretariat | GBIF | n.d.

Explains how DNA-derived biodiversity observations, metabarcoding records, and related molecular data can be published through biodiversity data infrastructures.

| GBIF Secretariat | GBIF | n.d.

Lists important quality requirements for occurrence datasets and explains how publishers can make biodiversity records easier to validate and reuse.

Marine Biodiversity Databases

| Vandepitte et al. | Hydrobiologia | 2024

Reviews the development of WoRMS and the work required to maintain an internationally coordinated, expert-curated database of marine species.

| Rabone et al. | Database | 2023

Evaluates the International Seabed Authority's DeepData biodiversity database and identifies challenges involving FAIR principles, taxonomy, duplication, and data interoperability.

| Various authors | Ecological Informatics | 2020

Uses marine mammal information to examine data quality and usability problems that can occur when researchers rely on large global marine biodiversity databases.

| Vandepitte et al. | European Journal of Taxonomy | 2018

Reviews a decade of experience operating the World Register of Marine Species and discusses the challenges of maintaining a global authoritative taxonomic database.

| Costello et al. | Nature Communications | 2017

Uses large marine biodiversity databases to delineate global marine biogeographic realms and quantify patterns of species distribution and endemicity.

| Costello et al. | PLOS ONE | 2013

Explains how the World Register of Marine Species coordinates marine taxonomy and provides standardized names that support numerous marine biodiversity databases.

| Ocean Biodiversity Information System | OBIS | n.d.

Explains the quality-control procedures OBIS uses to identify geographic, taxonomic, temporal, environmental, and other problems in marine biodiversity records.

| Ocean Biodiversity Information System | OBIS | n.d.

Describes different ways researchers can access and download the global marine species occurrence records aggregated by OBIS.

| Ocean Biodiversity Information System | OBIS | n.d.

Explains OBIS policies governing open biodiversity data, licensing, attribution, accessibility, responsibilities of publishers, and reuse of marine records.

| Ocean Biodiversity Information System | OBIS | n.d.

Provides guidance for organizations wishing to contribute marine biodiversity datasets to OBIS and make their records internationally accessible.

iNaturalist and Community Biodiversity Databases

| Mason et al. | BioScience | 2025

Reviews the rapidly expanding scientific use of iNaturalist and shows how its openly shared observations are accelerating biodiversity research across many disciplines.

| Ackland et al. | Ecology and Evolution | 2024

Develops a method for communicating confidence in iNaturalist records using observations of non-native marine species as a case study.

| Prenda et al. | Scientific Reports | 2024

Assesses the quality and coverage of millions of citizen-science bird records and identifies geographic and environmental biases relevant to biodiversity monitoring.

| Various authors | Ecological Solutions and Evidence | 2024

Develops metrics for judging whether citizen-science biodiversity datasets have sufficient completeness and geographic coverage for ecological monitoring.

| White et al. | Applications in Plant Sciences | 2023

Compares the quality of iNaturalist observations with digitized herbarium specimens and explores the strengths and limitations of both biodiversity data sources.

| Geurts et al. | Ecosphere | 2023

Investigates broad-scale spatial biases in iNaturalist and demonstrates how community-science observations become structured biodiversity datasets.

| iNaturalist | iNaturalist Help | 2023

Explains how researchers obtain and use iNaturalist observations through downloads, APIs, GBIF integration, and other biodiversity research workflows.

| Di Cecco et al. | BioScience | 2021

Examines how iNaturalist participants collect observations and how observer behavior creates spatial, temporal, and taxonomic patterns within the database.

| Kelling et al. | Ambio | 2015

Explains how large-scale citizen-science systems can use automated validation, participant calibration, and statistical modeling to improve biodiversity data quality.

| iNaturalist | iNaturalist Help | n.d.

Describes the Data Quality Assessment system used to determine whether observations qualify as Research Grade and become available to biodiversity data partners.

eBird and Open Bird Biodiversity Data

| La Sorte et al. | Diversity and Distributions | 2024

Examines coverage, environmental bias, sampling patterns, and temporal trends within eBird's worldwide network of frequently surveyed biodiversity locations.

| Fink et al. | Methods in Ecology and Evolution | 2023

Presents statistical methods designed to estimate biodiversity trends from large, heterogeneous citizen-science datasets such as eBird.

| Stuber et al. | Biological Conservation | 2022

Evaluates how semi-structured citizen-science observations can supplement traditional monitoring programs and support conservation decisions.

| Callaghan et al. | Biological Conservation | 2020

Explores the potential for massive community-science datasets to contribute to global monitoring of bird populations and biodiversity change.

| Sullivan et al. | Biological Conservation | 2014

Reviews eBird's integrated approach to citizen science, including data collection, quality control, computing infrastructure, research, and conservation applications.

| Sullivan et al. | Biological Conservation | 2009

Introduces eBird as a large citizen-based bird observation network capable of generating biodiversity information across continental scales.

| Cornell Lab of Ornithology | eBird Science | n.d.

Describes eBird's open data products, scientific tools, occurrence records, status products, and resources available for biodiversity research.

| Cornell Lab of Ornithology | eBird | n.d.

Explains how researchers can obtain the eBird Basic Dataset and other downloadable products containing global bird-observation records.

| Cornell Lab of Ornithology | eBird Status and Trends | n.d.

Documents eBird Status Data Products, which transform observations into modeled information about species abundance, distribution, and seasonal occurrence.

| Cornell Lab of Ornithology | eBird Status and Trends | n.d.

Explains eBird Trends products designed to estimate long-term changes in bird abundance using extensive community-science observations.

Taxonomic, Plant and Species Databases

| Catalogue of Life | Catalogue of Life | 2026

Provides downloadable releases of the Catalogue of Life, a major open global database integrating authoritative taxonomic checklists from specialist sources.

| Enquist et al. | Methods in Ecology and Evolution | 2026

Describes BIEN as an open biodiversity informatics ecosystem combining plant observations, inventories, plot records, taxonomy, geography, and trait information.

| Integrated Taxonomic Information System | ITIS | 2026

Provides standardized taxonomic names, classifications, synonyms, authorities, and identifiers for plants, animals, fungi, and microorganisms.

| Govaerts et al. | Scientific Data | 2021

Introduces the World Checklist of Vascular Plants as a continuously updated resource for studying global plant diversity and taxonomy.

| Kattge et al. | Global Change Biology | 2020

Describes the expanded TRY database, one of the world's largest repositories of plant functional-trait measurements.

| Maitner et al. | Methods in Ecology and Evolution | 2018

Introduces an R package that gives researchers programmatic access to the Botanical Information and Ecology Network biodiversity database.

| Ruggiero et al. | PLOS ONE | 2015

Develops a higher-level classification for living organisms using taxonomic resources including the Catalogue of Life and related biodiversity databases.

| Royal Botanic Gardens, Kew | Plants of the World Online | n.d.

Explains Plants of the World Online, Kew's global database linking accepted plant names with distributions, descriptions, images, traits, and taxonomic information.

| Royal Botanic Gardens, Kew | World Checklist of Vascular Plants | n.d.

Describes the World Checklist of Vascular Plants, an expert-curated global taxonomic resource incorporated into Plants of the World Online.

| Royal Botanic Gardens Kew, Harvard University Herbaria and Australian National Herbarium | International Plant Names Index | n.d.

Provides an authoritative open database of published scientific names and bibliographic information for seed plants, ferns, and lycophytes.

DNA Barcoding, Genomic and Molecular Biodiversity Databases

| Centre for Biodiversity Genomics | Barcode of Life Data Systems | 2026

Provides access to BOLD's global collection of DNA barcode sequences, specimen metadata, taxonomy, images, geographic information, and identification tools.

| Ratnasingham et al. | Methods in Molecular Biology | 2024

Describes BOLD version 4 and the informatics tools used to manage, analyze, identify, and disseminate DNA-based biodiversity information.

| Nakazato et al. | Frontiers in Ecology and Evolution | 2022

Examines combined use of BOLD and GenBank DNA barcoding information and discusses their value for integrating museum specimens with molecular biodiversity research.

| Droege et al. | Database | 2016

Describes the GGBN Data Standard developed to exchange information about genomic samples, tissue collections, DNA, environmental samples, and associated biodiversity records.

| Droege et al. | Biopreservation and Biobanking | 2016

Discusses the Global Genome Biodiversity Network from a botanical perspective and the importance of preserving genomic biodiversity resources for future research.

| Droege et al. | Nucleic Acids Research | 2014

Introduces the Global Genome Biodiversity Network Data Portal for discovering genomic samples maintained by biological collections around the world.

| Ratnasingham and Hebert | PLOS ONE | 2013

Introduces Barcode Index Numbers as a DNA-based system for organizing animal biodiversity and connecting sequence clusters with species-level information.

| Ratnasingham and Hebert | Molecular Ecology Notes | 2007

Introduces the Barcode of Life Data System, a global platform connecting DNA barcode sequences with taxonomy and voucher-specimen information.

| Kõljalg et al. | New Phytologist | 2005

Introduces UNITE, a web-accessible molecular database developed to help researchers identify fungi using DNA sequence information.

| Centre for Biodiversity Genomics | Barcode of Life Data Systems | n.d.

Explains how BOLD data should be cited and attributed when researchers download and reuse DNA barcode and specimen information.

Natural-History Collections and National Biodiversity Portals

| CSIRO and partner institutions | Atlas of Living Australia | 2026

Provides open access to Australian species observations, museum specimens, environmental layers, taxonomic information, images, collections, and analytical tools.

| Natural History Museum London | GBIF | 2026

Provides open access to millions of digitized collection specimens from the Natural History Museum in London through GBIF's global biodiversity infrastructure.

| Forbes, Young and Thrall | Nature Communications | 2025

Examines changes in biological collecting and warns that declining specimen acquisition could weaken the future scientific value of natural-history collections.

| Elliott, Luciano and Fortes | Biodiversity Information Science and Standards | 2024

Explores integration of large language models with the iDigBio portal to make millions of digitized natural-history records easier to search and investigate.

| Ainsa et al. | Biodiversity Information Science and Standards | 2023

Describes OpenObs, a Living Atlases platform developed to provide public access to French biodiversity occurrence information.

| Jackowiak et al. | Diversity | 2022

Proposes a comprehensive functional model for providing open digital access to biodiversity information held in natural-history collections.

| Canhos et al. | Biota Neotropica | 2022

Reviews speciesLink, Brazil's biodiversity-data infrastructure for integrating and openly distributing records from biological collections and other scientific institutions.

| Belbin et al. | Biodiversity Data Journal | 2021

Reviews the history and development of the Atlas of Living Australia and its role as a national open biodiversity-data infrastructure.

| Hedrick et al. | BioScience | 2020

Reviews the digitization of natural-history collections and explains how online specimen databases are transforming biodiversity research and access.

| Belbin and Williams | International Journal of Geographical Information Science | 2016

Uses the Atlas of Living Australia to examine how national biodiversity infrastructure can integrate biological records with environmental and geographic information.

Phylogenetic, Biodiversity Literature and Linked-Data Systems

| Biodiversity Heritage Library | Biodiversity Heritage Library | 2026

Provides openly digitized historical literature relevant to biodiversity informatics, biological databases, taxonomy, collections, and biodiversity knowledge systems.

| Penev et al. | Biodiversity Information Science and Standards | 2023

Describes the Biodiversity Knowledge Hub as infrastructure for connecting FAIR biodiversity data, publications, identifiers, and other research resources.

| Pasche et al. | Biodiversity Information Science and Standards | 2023

Presents the concept of a biodiversity-focused literature infrastructure analogous to PubMed Central that could connect publications with underlying biodiversity data.

| Dillen, Plank and Groom | Biodiversity Information Science and Standards | 2023

Explores connections between biodiversity infrastructures and Wikidata for harmonizing information about natural-history collectors and their collections.

| Global Names Architecture | Global Names | 2021

Documents open tools and services for discovering, parsing, matching, reconciling, and linking scientific names across biodiversity information systems.

| Pyle | ZooKeys | 2016

Proposes a Global Names Architecture for linking scientific names across biodiversity databases and improving discovery of biological information.

| Hinchliff et al. | Proceedings of the National Academy of Sciences | 2015

Describes the Open Tree of Life project and its synthesis of published phylogenies and taxonomy into a comprehensive, continuously updateable evolutionary tree.

| Boettiger and Temple Lang | Methods in Ecology and Evolution | 2012

Introduces an R package for discovering, downloading, and analyzing phylogenetic trees stored in the open TreeBASE database.

| Vos et al. | Nature Precedings | 2010

Discusses TreeBASE 2 and efforts to modernize the storage, discovery, exchange, and computational reuse of published phylogenetic trees.

| Page | BMC Bioinformatics | 2007

Describes TBMap, a method for linking TreeBASE phylogenetic information with taxonomic databases so evolutionary trees can be searched by organism names.

FAIR Data, Invasive Species and Emerging Biodiversity Infrastructure

| Weiland et al. | Biodiversity Information Science and Standards | 2024

Explores the use of RO-Crate and data-space technologies for integrating open agrobiodiversity information into biodiversity digital twins.

| Pagad et al. | Scientific Data | 2022

Presents a country-level compendium of the Global Register of Introduced and Invasive Species and expands access to standardized invasive-species information.

| Drucker et al. | Biodiversity Information Science and Standards | 2022

Uses plant-pollinator interaction data as a WorldFAIR case study for improving interoperability and FAIRness of complex biodiversity information.

| Groom et al. | Biodiversity Information Science and Standards | 2021

Discusses improvements to Darwin Core needed to make alien-species occurrence, introduction, establishment, and invasion information more interoperable.

| Nakazato | Biodiversity Information Science and Standards | 2020

Compares species coverage in BOLD and GenBank and examines opportunities for integrating DNA barcode and genomics databases for museomics.

| Environmental Data Initiative authors | Biodiversity Information Science and Standards | 2019

Describes how a major ecological data repository implements FAIR principles to make environmental and biodiversity research data discoverable and reusable.

| Lahti et al. | Biodiversity Information Science and Standards | 2019

Examines the principle of making biodiversity data as open as possible while allowing restrictions when legal, ethical, or ownership requirements make them necessary.

| Pagad et al. | Scientific Data | 2018

Introduces the Global Register of Introduced and Invasive Species, an open standardized database documenting alien and invasive taxa by country.

| Porter and Hajibabaei | PLOS ONE | 2018

Examines the rapid growth of COI barcode sequences in GenBank and their importance for DNA barcoding, metabarcoding, and biodiversity database development.

| GBIF Secretariat | GBIF | n.d.

Explains biodiversity data papers, a publication approach that gives scholarly credit for documenting and openly releasing reusable biodiversity datasets.

Freshwater and Regional Biodiversity Data Systems

| Dallas et al. | Freshwater Biodiversity Information System | 2026

Provides an open-access platform for hosting, visualizing, analyzing, and downloading freshwater biodiversity records from South Africa.

| Freshwater Research Centre | FBIS Africa | 2026

Expands the Freshwater Biodiversity Information System model across Africa to improve access to fish, invertebrate, plant, amphibian, and other freshwater records.

| Dallas and Shelton | South African Journal of Science | 2026

Reviews the development of FBIS and explains how open freshwater biodiversity data can support conservation assessment, monitoring, environmental management, and policy.

| BioFresh Consortium | Freshwater Biodiversity Data Portal | 2026

Provides a searchable metadatabase describing hundreds of freshwater biodiversity datasets, including their geographic coverage, taxa, accessibility, and ownership.

| Australian Antarctic Data Centre | Australian Antarctic Biodiversity Database | 2026

Provides openly accessible taxonomy, occurrence records, collections, bioregions, and alien-species information for Antarctic and sub-Antarctic biodiversity.

| National Biodiversity Network | NBN Atlas Documentation | 2026

Explains how biodiversity occurrence records and species lists from the United Kingdom's NBN Atlas can be searched and downloaded for research.

| European Commission Joint Research Centre | European Alien Species Information Network | 2026

Integrates records from multiple European databases to provide searchable information and distribution maps for alien and invasive species.

| European Commission Joint Research Centre | EASIN Web Services | 2026

Provides open REST web services allowing researchers and developers to retrieve EASIN species and geospatial biodiversity information programmatically.

| Jézéquel et al. | Scientific Data | 2020

Presents a database containing hundreds of thousands of occurrence records for freshwater fishes throughout the Amazon Basin.

| Schmidt-Kloiber et al. | Springer | 2018

Reviews freshwater biodiversity information systems, including BioFresh, species registers, occurrence databases, metadata repositories, and tools for integrating freshwater data.

Open Species-Trait Databases

| Wilman et al. | Open Traits Network | 2026

Provides access to EltonTraits, a global compilation of diet, foraging, activity, and body-mass attributes for birds and mammals.

| Gómez-Gras et al. | Scientific Data | 2025

Presents the Octocoral Trait Database, extending standardized trait-data infrastructure to soft corals, sea fans, and related octocoral groups.

| Mouret et al. | Scientific Data | 2025

Presents BeeFunc, a large species-level functional-trait database covering the wild bee fauna of France.

| Herberstein et al. | AnimalTraits | 2022

Provides a freely reusable database of body mass, metabolic rate, and brain-size measurements across terrestrial animal groups.

| Tobias et al. | Ecology Letters | 2022

Presents AVONET, a comprehensive database of morphological, ecological, geographical, and functional information for all living bird species.

| Jeliazkov et al. | Scientific Data | 2020

Introduces CESTES, a global metacommunity database integrating species occurrence, traits, environmental variables, and spatial information.

| Kissling et al. | Scientific Data | 2019

Introduces PalmTraits, a global species-level database containing functional and ecological trait information for palms.

| Oliveira et al. | Scientific Data | 2017

Introduces AmphiBIO, a global database compiling ecological, reproductive, morphological, and life-history traits for amphibians.

| Madin et al. | Scientific Data | 2016

Introduces the Coral Trait Database, an open repository of physiological, ecological, morphological, reproductive, and biogeographic traits for coral species.

| Faulwetter et al. | Biodiversity Data Journal | 2014

Describes Polytraits, an open database containing biological and ecological traits for marine polychaete worms.

Functional and Ecological Biodiversity Databases

| AusTraits Consortium | AusTraits Data Portal | 2026

Provides an interactive portal for exploring raw and summarized functional-trait measurements from Australia's diverse plant flora.

| Open Traits Network | Open Traits Network | 2026

Maintains a catalog of openly available trait databases spanning plants, mammals, birds, insects, marine organisms, microbes, and other taxonomic groups.

| Weng et al. | Scientific Data | 2025

Provides an integrated polychaete dataset combining species distributions, DNA barcodes, and functional traits.

| Saffer et al. | Scientific Data | 2024

Presents GIATAR, a spatiotemporal database combining global invasive and alien-species occurrences with species traits.

| Fu et al. | Scientific Data | 2024

Combines citizen-science observations with literature records to create a reusable functional-trait dataset for Taiwan's birds.

| Denelle et al. | Methods in Ecology and Evolution | 2023

Describes software for openly accessing the Global Inventory of Floras and Traits and integrating its checklist and trait information into research workflows.

| Tanalgo et al. | Scientific Data | 2022

Introduces DarkCideS, a global database integrating occurrence, ecological, distributional, and conservation traits for cave-dwelling bats.

| Raja et al. | Scientific Data | 2022

Introduces Ancient Reef Traits, an open database of biological and ecological traits for reef-building organisms through geological time.

| Chevalier et al. | Scientific Data | 2021

Presents WOODIV, integrating occurrence, phylogenetic, and functional-trait information for trees across the Euro-Mediterranean region.

| Weigelt et al. | Journal of Biogeography | 2019

Introduces GIFT, the Global Inventory of Floras and Traits, linking regional plant checklists with functional traits and environmental information.

Specialist Taxonomic Databases

| Uetz et al. | The Reptile Database | 2026

Maintains a continuously updated global catalog of living reptiles containing taxonomy, synonyms, distributions, type information, literature, and images.

| Natural History Museum Bern | World Spider Catalog | 2026

Provides a continuously updated global taxonomic and bibliographic catalog of recognized spider species.

| Cigliano et al. | Orthoptera Species File | 2026

Provides global taxonomy, synonyms, literature, specimen records, photographs, and sound recordings for grasshoppers, crickets, katydids, and relatives.

| Cigliano et al. | Orthoptera Species File | 2026

Describes the database's migration to TaxonWorks and its support for downloadable Darwin Core and CSV biodiversity data.

| American Society of Mammalogists | Mammal Diversity Database | 2026

Maintains a continuously updated open database of global mammal taxonomy, nomenclature, distributions, synonyms, and taxonomic changes.

| World Flora Online Consortium | World Flora Online | 2026

Provides a global online flora containing accepted plant names, synonyms, descriptions, distributions, images, references, and taxonomic classifications.

| Burgin et al. | Journal of Mammalogy | 2025

Documents major upgrades to the Mammal Diversity Database and evaluates continuing changes in global mammal taxonomy and nomenclature.

| Esposito et al. | BioScience | 2023

Uses the World Spider Catalog to reveal major historical, geographic, and taxonomic biases in efforts to document biodiversity.

| Uetz | Biodiversity Information Science and Standards | 2021

Examines how the Reptile Database curates a rapidly expanding global taxonomic literature despite relying heavily on volunteer effort.

| Borsch et al. | TAXON | 2020

Explains the architecture, expert-curation model, data standards, and FAIR principles behind World Flora Online.

Fungal, Amphibian, Marine and Paleontological Databases

| Royal Botanic Gardens Kew | Index Fungorum | 2026

Provides nomenclatural information and registration services for scientific names of fungi, lichens, yeasts, and related organisms.

| Uhen et al. | PaleoBios | 2026

Provides an extensive modern user guide to the Paleobiology Database, its data model, records, search tools, and research applications.

| SeaLifeBase Consortium | SeaLifeBase | 2026

Provides searchable biological, ecological, distributional, and taxonomic information for tens of thousands of marine species other than fishes.

| AmphibiaWeb | University of California Berkeley | 2026

Provides open species accounts, taxonomy, conservation information, distributions, life histories, photographs, and literature for the world's amphibians.

| California Academy of Sciences | AntWeb | 2026

Provides public programmatic access to the world's largest online collection of ant specimen records, images, taxonomic information, and geographic data.

| Raja et al. | Paleobiology | 2024

Examines how open paleontological databases can improve contributor recognition and more equitable citation of biodiversity data.

| Wang et al. | Database | 2023

Describes Fungal Names, a comprehensive nomenclatural repository and knowledge base containing hundreds of thousands of fungal taxon names.

| Petersen | IMA Fungus | 2016

Reviews the historical development of Index Fungorum and other major databases used to organize fungal nomenclature.

| Penev et al. | ZooKeys | 2016

Describes automated links between publications and major nomenclatural databases including IPNI, Index Fungorum, MycoBank, and ZooBank.

| Peters and McClennen | Paleobiology | 2016

Describes the Paleobiology Database API and demonstrates how open fossil biodiversity records can be incorporated into computational research workflows.

Biodiversity Literature and Knowledge Graph Infrastructure

| Biodiversity Heritage Library | BHL | 2026

Documents BHL's APIs, bulk-data exports, OAI-PMH services, cloud datasets, and tools for reusing digitized biodiversity literature.

| Page | Biodiversity Heritage Library | 2026

Explains efforts to identify individual journal articles within millions of digitized BHL pages and make them independently searchable and citable.

| Global Biotic Interactions | GloBI | 2026

Provides downloadable and API-accessible datasets describing predator-prey, host-parasite, pollination, and numerous other species interactions.

| Roderic Page | BioStor | 2026

Links article-level bibliographic records to digitized biodiversity literature contained within the Biodiversity Heritage Library.

| Page | Biodiversity Data Journal | 2023

Describes efforts to connect more than one million taxonomic names with persistent identifiers for publications and researchers.

| Kindt | Applications in Plant Sciences | 2020

Introduces WorldFlora, an open R package for matching large lists of plant names against the World Flora Online taxonomic backbone.

| Biodiversity Heritage Library | BHL Blog | 2019

Describes the integration that made tens of thousands of historical biodiversity articles in BHL discoverable through Unpaywall.

| Parr et al. | Biodiversity Data Journal | 2014

Describes Encyclopedia of Life version 2 and its architecture for aggregating openly licensed species information from many biodiversity databases.

| Poelen, Simons and Mungall | Ecological Informatics | 2014

Introduces GloBI as an open infrastructure for integrating and querying species-interaction datasets from many independent sources.

| Page | BMC Bioinformatics | 2005

Describes an early federated taxonomic search engine that combined information from multiple independent online biodiversity databases.

Natural-History Collections and Specimen Infrastructure

| Natural History Museum London | NHM Data Portal | 2026

Provides open access to digitized natural-history specimens and associated collection datasets held by the Natural History Museum.

| Integrated Digitized Biocollections | iDigBio | 2026

Aggregates digitized specimen records and media from natural-history collections throughout the United States and partner institutions.

| Symbiota Support Hub | Symbiota | 2026

Provides open-source software and biodiversity portals used by museums, herbaria, universities, and collection networks to publish specimen records.

| Symbiota Collections of Arthropods Network | SCAN | 2026

Aggregates millions of digitized arthropod specimen records from entomological collections across North America.

| Arctos Consortium | Arctos | 2026

Provides a collaborative collection-management platform linking specimen records with taxonomy, geography, agents, publications, genetics, and media.

| Duke University | MorphoSource | 2026

Provides an open digital repository for three-dimensional scans and other media derived from biological, paleontological, and cultural specimens.

| Shorthouse | Bionomia | 2026

Links natural-history specimens to the people who collected or identified them using ORCID and other persistent identifiers.

| Distributed System of Scientific Collections | DiSSCo | 2026

Develops European infrastructure for unifying and digitally accessing natural-science collection specimens across institutions.

| GBIF | Global Registry of Scientific Collections | 2026

Provides standardized information on scientific collections and institutions holding biological and geological specimens around the world.

| Scott et al. | Database | 2019

Describes the Natural History Museum Data Portal and its open architecture for specimen records, datasets, identifiers, APIs, and bulk downloads.

Software for Accessing and Cleaning Open Biodiversity Databases

| GBIF Secretariat | GBIF Technical Documentation | 2026

Explains how the pygbif Python library can search and retrieve GBIF species, occurrence, registry, and mapping data.

| Owens et al. | rOpenSci | 2026

Provides tools for querying large biodiversity occurrence databases while preserving metadata needed to correctly cite underlying datasets.

| Chamberlain et al. | rOpenSci | 2026

Provides a unified R interface for retrieving species occurrence information from multiple biodiversity databases.

| Chamberlain et al. | rOpenSci | 2026

Documents the rgbif package for programmatically searching GBIF taxonomy, occurrence records, datasets, and metadata.

| Atlas of Living Australia | galah | 2026

Provides R and Python interfaces for querying, filtering, and downloading biodiversity records from the Atlas of Living Australia.

| Biodiversity Data Cleaning Project | CRAN | 2026

Provides tools designed to identify, document, and correct errors in large biodiversity datasets before analysis.

| Chamberlain | rOpenSci | 2026

Provides functions for cleaning species-occurrence records obtained from biodiversity databases before ecological or biogeographic analysis.

| GBIF and Chamberlain | Python Package Index | 2025

Provides the maintained Python client for accessing GBIF's biodiversity APIs directly from research scripts.

| Zizka et al. | Methods in Ecology and Evolution | 2019

Introduces CoordinateCleaner, an open-source tool for automatically detecting geographic and temporal errors in large biodiversity occurrence datasets.

| Chamberlain and Szöcs | F1000Research | 2013

Introduces taxize, an R package providing reproducible programmatic access to numerous online taxonomic databases.

Molecular, eDNA and Sequence Biodiversity Databases

| MetaZooGene Consortium | MetaZooGene | 2026

Develops reference DNA-barcode and metabarcoding resources for marine animal biodiversity, particularly zooplankton.

| PR2 Consortium | Protist Ribosomal Reference Database | 2026

Provides curated ribosomal DNA sequences and taxonomy for protists and other microbial eukaryotes used in biodiversity metabarcoding.

| SILVA Consortium | SILVA Ribosomal RNA Database | 2026

Provides curated ribosomal RNA sequence datasets widely used for identifying bacterial, archaeal, and eukaryotic microbial biodiversity.

| UNITE Community | UNITE | 2026

Provides curated fungal DNA sequence reference data and species hypotheses used extensively in environmental sequencing and fungal biodiversity research.

| MitoFish Project | MitoFish | 2026

Provides a comprehensive database of fish mitochondrial genomes and associated taxonomic information.

| European Commission Joint Research Centre | FishTrace | 2026

Provides genetic reference information and biological material supporting identification and traceability of European marine fishes.

| European Molecular Biology Laboratory-EBI | European Nucleotide Archive | 2026

Provides openly accessible nucleotide sequences, environmental sequencing data, metagenomes, and associated metadata relevant to molecular biodiversity research.

| Brandt et al. | Scientific Data | 2025

Presents an openly accessible deep-sea biodiversity dataset combining metabarcoding, metagenomics, environmental DNA, and standardized metadata.

| Leray et al. | Database | 2022

Describes MIDORI2, a curated mitochondrial DNA reference database designed for taxonomic assignment in eDNA and metabarcoding studies.

| Knockaert et al. | Scientific Data | 2019

Demonstrates biodiversity-data rescue by digitizing and publishing historical marine records from decades of Kenya-Belgium scientific cooperation.

Conservation, Distribution and Biodiversity Knowledge Platforms

| International Union for Conservation of Nature | IUCN Red List API | 2026

Provides programmatic access to conservation assessments, taxonomy, threats, habitats, population trends, and distribution information from the IUCN Red List.

| UNEP-WCMC and IUCN | Protected Planet | 2026

Provides the global authoritative database of protected areas and other effective area-based conservation measures.

| UNEP-WCMC | World Database on Protected Areas | 2026

Provides downloadable spatial and descriptive information on terrestrial and marine protected areas worldwide.

| KBA Partnership | World Database of Key Biodiversity Areas | 2026

Provides information on sites identified as globally significant for the persistence of biodiversity.

| BirdLife International | BirdLife DataZone | 2026

Provides species assessments, distribution information, Important Bird and Biodiversity Areas, conservation statistics, and global bird datasets.

| Jetz Lab | Map of Life | 2026

Integrates thousands of biodiversity datasets to produce species distributions, habitat models, conservation indicators, and spatial biodiversity analyses.

| Map of Life | Map of Life Datasets | 2026

Provides access to biodiversity datasets underlying Map of Life species-distribution products, models, and conservation metrics.

| Catalogue of Life and GBIF | ChecklistBank | 2026

Provides infrastructure for storing, publishing, comparing, downloading, and reconciling taxonomic and nomenclatural checklists.

| Plazi | TreatmentBank | 2026

Extracts taxonomic treatments from scientific publications and makes species descriptions, names, figures, and references available as reusable structured data.

| Species File Group | TaxonWorks | 2026

Provides an open-source platform for managing taxonomy, nomenclature, specimens, biological associations, collecting events, literature, images, and other biodiversity data.