Fine-Grained Visual Recognition (FGVR) tackles the problem of distinguishing highly similar categories. One of the main approaches to FGVR, namely subset learning, tries to leverage information from existing class taxonomies to improve the performance of deep neural networks. However, these methods rely on the existence of handcrafted hierarchies that are not necessarily optimal for the models. In this paper, we propose ELFIS, an expert learning framework for FGVR that clusters categories of the dataset into meta-categories using both dataset-inherent lexical and model-specific information. A set of neural networks-based experts are trained focusing on the meta-categories and are integrated into a multi-task framework. Extensive experimentation shows improvements in the SoTA FGVR benchmarks of up to +1.3% of accuracy using both CNNs and transformer-based networks. Overall, the obtained results evidence that ELFIS can be applied on top of any classification model, enabling the obtention of SoTA results. The source code will be made public soon.
ELFIS: Expert Learning for Fine-grained Image Recognition Using Subsets
Pablo J. Villacorta,Jesús M. Rodríguez-de-Vera,Marc Bolaños,Ignacio Saras'ua,Bhalaji Nagarajan,P. Radeva
Published 2023 in arXiv.org
ABSTRACT
PUBLICATION RECORD
- Publication year
2023
- Venue
arXiv.org
- Publication date
2023-03-16
- Fields of study
Computer Science
- Identifiers
- External record
- Source metadata
Semantic Scholar
CITATION MAP
EXTRACTION MAP
CLAIMS
- No claims are published for this paper.
CONCEPTS
- No concepts are published for this paper.
REFERENCES
Showing 1-56 of 56 references · Page 1 of 1
CITED BY
Showing 1-2 of 2 citing papers · Page 1 of 1