Mostrar el registro sencillo del ítem
resumen
Resumen
The annotation of repetitive sequences within plant genomes can help in the interpretation of observed phenotypes. Moreover, repeat masking is required for tasks such as whole-genome alignment, promoter analysis, or pangenome exploration. Although homology-based annotation methods are computationally expensive, k-mer strategies for masking are orders of magnitude faster. Here, we benchmarked a two-step approach, where repeats were first called by k-mer
[ver mas...]
dc.contributor.author | Contreras-Moreira, Bruno | |
dc.contributor.author | Filippi, Carla Valeria | |
dc.contributor.author | Naamati, Guy | |
dc.contributor.author | García Girón, Carlos | |
dc.contributor.author | Allen, James E. | |
dc.contributor.author | Flicek, Paul | |
dc.date.accessioned | 2021-12-10T13:45:33Z | |
dc.date.available | 2021-12-10T13:45:33Z | |
dc.date.issued | 2021-09 | |
dc.identifier.issn | 1940-3372 | |
dc.identifier.other | https://doi.org/10.1002/tpg2.20143 | |
dc.identifier.uri | http://hdl.handle.net/20.500.12123/10882 | |
dc.identifier.uri | https://acsess.onlinelibrary.wiley.com/doi/full/10.1002/tpg2.20143 | |
dc.description.abstract | The annotation of repetitive sequences within plant genomes can help in the interpretation of observed phenotypes. Moreover, repeat masking is required for tasks such as whole-genome alignment, promoter analysis, or pangenome exploration. Although homology-based annotation methods are computationally expensive, k-mer strategies for masking are orders of magnitude faster. Here, we benchmarked a two-step approach, where repeats were first called by k-mer counting and then annotated by comparison to curated libraries. This hybrid protocol was tested on 20 plant genomes from Ensembl, with the k-mer-based Repeat Detector (Red) and two repeat libraries (REdat, last updated in 2013, and nrTEplants, curated for this work). Custom libraries produced by RepeatModeler were also tested. We obtained repeated genome fractions that matched those reported in the literature but with shorter repeated elements than those produced directly by sequence homology. Inspection of the masked regions that overlapped genes revealed no preference for specific protein domains. Most Red-masked sequences could be successfully classified by sequence similarity, with the complete protocol taking less than 2 h on a desktop Linux box. A guide to curating your own repeat libraries and the scripts for masking and annotating plant genomes can be obtained at https://github.com/Ensembl/plant-scripts. | e |
dc.format | application/pdf | es_AR |
dc.language.iso | eng | es_AR |
dc.publisher | Wiley | es_AR |
dc.rights | info:eu-repo/semantics/openAccess | es_AR |
dc.rights.uri | http://creativecommons.org/licenses/by-nc-sa/4.0/ | |
dc.source | The Plant Genome 14 (3) : e20143 (November 2021) | es_AR |
dc.subject | Genomas | es_AR |
dc.subject | Genomes | eng |
dc.subject | Fitogenética | es_AR |
dc.subject | Plant Genetics | eng |
dc.subject | Genética | es_AR |
dc.subject | Genetics | eng |
dc.title | K-mer counting and curated libraries drive efficient annotation of repeats in plant genomes | es_AR |
dc.type | info:ar-repo/semantics/artículo | es_AR |
dc.type | info:eu-repo/semantics/article | es_AR |
dc.type | info:eu-repo/semantics/publishedVersion | es_AR |
dc.rights.license | Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) | |
dc.description.origen | Instituto de Biotecnología | es_AR |
dc.description.fil | Fil: Contreras-Moreira, Bruno. European Bioinformatics Institute. European Molecular Biology Laboratory; Reino Unido | es_AR |
dc.description.fil | Fil: Filippi, Carla Valeria. Instituto Nacional de Tecnología Agropecuaria (INTA). Instituto de Agrobiotecnología y Biología Molecular (IABIMO); Argentina | es_AR |
dc.description.fil | Fil: Filippi, Carla Valeria. Consejo Nacional de Investigaciones Científicas y Técnicas; Argentina | es_AR |
dc.description.fil | Fil: Filippi, Carla Valeria. European Bioinformatics Institute. European Molecular Biology Laboratory; Reino Unido | es_AR |
dc.description.fil | Fil: Naamati, Guy. European Bioinformatics Institute. European Molecular Biology Laboratory; Reino Unido | es_AR |
dc.description.fil | Fil: García Girón, Carlos. European Bioinformatics Institute. European Molecular Biology Laboratory; Reino Unido | es_AR |
dc.description.fil | Fil: Allen, James E. European Bioinformatics Institute. European Molecular Biology Laboratory; Reino Unido | es_AR |
dc.description.fil | Fil: Flicek, Paul. European Bioinformatics Institute. European Molecular Biology Laboratory; Reino Unido | es_AR |
dc.subtype | cientifico |
Ficheros en el ítem
Este ítem aparece en la(s) siguiente(s) colección(ones)
common
-
Artículos científicos [404]