Articles
Construction of a semantic distance for inferring structure of the variability between 19th century Rosa cultivars
Article number
1384_60
Pages
477 – 484
Language
English
Abstract
Genetic resources, thanks to the genetic diversity they stand for, are of first importance to keep evolution capacity for agriculture and horticulture.
Many efforts are made to preserve and characterize them at different levels, especially: the phenotypic one for many traits as well as the genetic one with molecular markers or DNA sequencing.
Statistical tools to describe the available variability through the use of variables of different nature (qualitative, semi-ordered, quantitative) are quite sparse, especially if there are missing data.
The aim of this study is double: i) integrate different types of variables in a unique statistical analysis, such as clustering, by using the concept of ontology and defining a new distance measure based, if necessary, on expert knowledge; ii) use statistical techniques to better know the underlying reasons of the revealed structure as well as the characterization of each cluster.
Our new semantic distance based on various type of data (passport, phenotypic) was used to estimate pairwise distances between 1400 Rosa individuals.
The resulting distance matrix was projected into a new coordinates space thanks to metric multi-dimensional scaling and the projection was then used as input for clustering algorithms.
To evaluate the contribution of each rosebush characteristic to the observed structure, modalities of the variables were projected into the coordinates space and their proportion estimated for each cluster.
A similar work was done with the Gowers distance.
Data points representing individuals are wider spread into the coordinates space and projection of the modalities of the variables shows a stronger structuring when using the semantic distance.
Our distance better represents the reality and the stronger structuring of the variables modalities leads to more precise biological questions.
This semantic distance seems to be useful to investigate variability between genetic resources accessions and should be tested on more data sets.
Many efforts are made to preserve and characterize them at different levels, especially: the phenotypic one for many traits as well as the genetic one with molecular markers or DNA sequencing.
Statistical tools to describe the available variability through the use of variables of different nature (qualitative, semi-ordered, quantitative) are quite sparse, especially if there are missing data.
The aim of this study is double: i) integrate different types of variables in a unique statistical analysis, such as clustering, by using the concept of ontology and defining a new distance measure based, if necessary, on expert knowledge; ii) use statistical techniques to better know the underlying reasons of the revealed structure as well as the characterization of each cluster.
Our new semantic distance based on various type of data (passport, phenotypic) was used to estimate pairwise distances between 1400 Rosa individuals.
The resulting distance matrix was projected into a new coordinates space thanks to metric multi-dimensional scaling and the projection was then used as input for clustering algorithms.
To evaluate the contribution of each rosebush characteristic to the observed structure, modalities of the variables were projected into the coordinates space and their proportion estimated for each cluster.
A similar work was done with the Gowers distance.
Data points representing individuals are wider spread into the coordinates space and projection of the modalities of the variables shows a stronger structuring when using the semantic distance.
Our distance better represents the reality and the stronger structuring of the variables modalities leads to more precise biological questions.
This semantic distance seems to be useful to investigate variability between genetic resources accessions and should be tested on more data sets.
Authors
A. Pernet, R. Eid, C. Landès, E. Benoît, P. Santagostini, J. Marie-Magdelaine, J. Clotault, A. El Ghaziri, J. Bourbeillon
Keywords
ontology, phenotypic and genetic diversity, rose, clustering, descriptive statistical analyses, visualisation
Online Articles (68)
