Back to search

George Mason University

BENCHMARKING AND CREATION OF A DEEP LEARNING COMPUTER VISION DATASET IN THE TAXONOMIC REVISION OF THE PALO SANTO TREE HERBARIUM SPECIMENS

Abstract

Premise of the studyHerbarium collections have historically been used as a manual data source for taxonomies and there are large collections that have been digitized since the start of the 21st century. Few have been curated for computer vision problems beyond image classification, and even fewer are available to the public. In this study, I investigated the steps and effort it takes both human and computer to create a publically-available data set that can be used for image classification, object detection, and segmentation. Secondly, while benchmarking the data set, I investigated the relationship between the algorithm class, data augmentation, and network size. Our hypothesis for the first research question is that, despite the tools and technology out there to help expedite the process, that it still requires a substantial amount of human effort to complete the task. Secondly, I believe that there will be signals within the data, and that there is a relationship between the class of algorithm, data augmentation, and network size. Methods: A team of three people, one botanist, a data scientist, and research assistant contributed to the creation of this benchmark data set. Data was collected both from historical and new field samples and digitized in Washington D.C. B. graveolens and B. penicillata were selected as the benchmarking species, as they are rare tree species to reduce variability and ensure and were able to collect most of the known herbarium samples around the world. Supervisely, a web-based tool to scale image annotations, was used to annotate objects such as plant features, and to track efforts [1]. Python 3.7 was used for post- processing and creation of the data sets from Supervisely [2]. Descriptive statistics for univariate and bivariate were conducted using mean (sd), frequency counts, and appropriate statistical graphing. Deep learning models were created via PyTorch and torchvision for classification (B.graveolens and B. penicillata), semantic segmentation (plant pixel versus non-plant pixel), and instance segmentation (index card, stamp, measurement bar, color bar, barcode, compound leaf, terminal leaf, and woody material), and evaluated over different augmentation strategies and network sizes [3], [4]. Results: 1,081 images were taken of 794 biologically unique samples. 115,279 actions across 1,200 human hours were taken to annotate data. 15 types of objects were annotated across the images, resulting in 19,834 objects. For classification, a model without pretrained weights based on VGG-11 performed best with an accuracy of .727, AUC of 0.685, specificity .742, and precision .265, as compared to the second best model of ResNet- 18 .565, .665, .530, and .201. FCN-100, without upsampling and a dilation factor of 1, recorded the highest dice coefficient 0.7566. MASK R-CNN V1.0 performed better for both the non-biologic and biological models with MAP full .846 and .736, respectively optimized at 50x and 25x upsampled. Classification and semantic segmentation of larger networks for ResNet and MASK R-CNN in comparison performed worse regardless of the upsample size.Conclusion: Although a substantial amount of time was spent creating this data set, it is a modular process, and the toolkits of the 21st century have made this a process tractable with a small research team. The amount of training data is different for each algorithm class despite it being a single set of images. The smaller the training data, the greater the impact of upsampling, and conversely network size. I hope that this newly-created benchmarking data and findings help researchers interested in computer vision in herbarium research make progress toward bridging known gaps in the field.

Author and committee

dc:creator, dc:contributor.*
Author
  • Valko, Matthew

Subjects

dc:subject × 6

Identifiers

dc:identifier.*
Identifier
hdl:1920/14377
OAI identifier oai:identifier
oai:MARS:1920/14377

Chain of custody

source
Harvested from
George Mason University
Base URL
mars.gmu.edu/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Valko, Matthew. BENCHMARKING AND CREATION OF A DEEP LEARNING COMPUTER VISION DATASET IN THE TAXONOMIC REVISION OF THE PALO SANTO TREE HERBARIUM SPECIMENS. 2024.