Machine-supported monitoring of plant communities and habitats

Project lead

Philipp Brun

Project staff

Thomas Grégoire Stegmüller
 

Project duration

2026 - 2029

Cooperation Financing

In Switzerland, 48% of habitats and over a third of species are threatened. However, only a minor fraction is sufficiently monitored because confident identification of habitat types and associated species is time-consuming and requires expert knowledge. We aim to expand habitat monitoring in Switzerland by developing algorithms that can semi-automatically identify plant communities from videos and link them to habitat types.

The rapid adoption of reporting platforms by citizen scientists has led to a vast increase in species observations and identifications. We recently demonstrated the impact of machine-learning support for biodiversity monitoring by integrating image-based species identification (FlorID) into FlorApp, the leading app for monitoring Swiss biodiversity. Here, we propose taking the next step: building the foundation for a service to rapidly monitor plant communities and habitats. Specifically, we aim to streamline vegetation surveys by developing the algorithmic foundations for a smartphone app that communicates with a web server to (1) derive tentative vegetation surveys from videos, (2) propose expected but not yet observed species, and (3) make suggestions for habitat types. All developments will be deployed on the SpeciesID API hosted by WSL and made freely available.

Vegetation surveys

High-quality vegetation surveys are critical for obtaining the reliable data needed to assess the state and trends in biodiversity. These surveys act as inventories of all species growing within a predefined area and may be linked to habitat types. Typically, they are conducted by professionals. Apart from FlorApp, which is used by experts and citizen scientists alike, most applications targeted at citizen scientists lack the option to record vegetation surveys. Because conducting comprehensive vegetation surveys remains challenging, we envision this project to make vegetation surveys more accessible and scalable by leveraging the recent progress in data science techniques.

Project goals

SmartHabitat's core consists of analytical pipelines of server tasks (Figure 1 and Parts I-III below), which will be deployed as new endpoints of the SpeciesID API. It builds on FlorID, our image-based identification tool that accurately distinguishes >3000 species (85% of the Swiss flora) and is integrated in the popular FlorApp.

Part I - Video interpretation: We aim to generalize identification from single photos to short 30-to-60-second videos capable of capturing entire plant communities. Such videos are expected to be taken by citizen scientists. We propose to retrain the existing FlorID vision models to recognize multiple species per frame. The model will make multi-species frame-level predictions that are aggregated to a complete species list for the video. The final deliverable for this part is an automated analytical pipeline for video-based species identification.

Part II - Predicting expected but unobserved species: We will train a recommender system on the best available vegetation survey data to identify common co-occurrences of plant species in the field. By modelling the interactions between species and sites, and exploiting the stability of plant associations within habitats, the system can take a partial list of identified species and predict others likely to be nearby. Ultimately, this model is envisioned to guide users in the field, making suggestions on which plants to look for based on their initial observations. The final deliverable for this phase is an analytical pipeline for species recommendation.

Part III - Habitat classification: To support users in the field with accurate habitat identification, we are developing machine-learning-based classifiers to identify habitats according to the Swiss TypoCH system. Specifically, we will focus on the 110 to 150 vegetated types most crucial for nature conservation. To increase accuracy, our system will combine complementary data sources using two distinct models: (1) a multimodal model that uses point observations of environmental variables (e.g., temperature and precipitation) and high-resolution remote-sensing images on vegetation structure and terrain, and (2) a specialized computer vision model that processes in-situ user photos. By integrating these models, they may provide suggestions for observed habitat type, based on location, species identified, and in situ user supplied photos. These classification tools will be detailed in a forthcoming research paper and deployed with the other pipelines on the API.

Publications