当前位置: X-MOL 学术Virus Evol. › 论文详情
Our official English website, www.x-mol.net, welcomes your feedback! (Note: you will need to create a separate account there.)
LoReTTA, a user-friendly tool for assembling viral genomes from PacBio sequence data
Virus Evolution ( IF 5.5 ) Pub Date : 2021-04-22 , DOI: 10.1093/ve/veab042
Ahmed Al Qaffas 1 , Jenna Nichols 2 , Andrew J Davison 2 , Amine Ourahmane 1 , Laura Hertel 3 , Michael A McVoy 1 , Salvatore Camiolo 2
Affiliation  

Long-read, single-molecule DNA sequencing technologies have triggered a revolution in genomics by enabling the determination of large, reference-quality genomes in ways that overcome some of the limitations of short-read sequencing. However, the greater length and higher error rate of the reads generated on long-read platforms make the tools used for assembling short reads unsuitable for use in data assembly and motivate the development of new approaches. We present LoReTTA (Long Read Template-Targeted Assembler), a tool designed for performing de novo assembly of long reads generated from viral genomes on the PacBio platform. LoReTTA exploits a reference genome to guide the assembly process, an approach that has been successful with short reads. The tool was designed to deal with reads originating from viral genomes, which feature high genetic variability, possible multiple isoforms, and the dominant presence of additional organisms in clinical or environmental samples. LoReTTA was tested on a range of simulated and experimental datasets and outperformed established long-read assemblers in terms of assembly contiguity and accuracy. The software runs under the Linux operating system, is designed for easy adaptation to alternative systems, and features an automatic installation pipeline that takes care of the required dependencies. A command-line version and a user-friendly graphical interface version are available under a GPLv3 license at https://bioinformatics.cvr.ac.uk/software/ with the manual and a test dataset.

中文翻译:

LoReTTA,一种用户友好的工具,用于从 PacBio 序列数据中组装病毒基因组

长读长、单分子 DNA 测序技术通过克服短读长测序的一些限制,能够确定大型、参考质量的基因组,从而引发了基因组学的一场革命。然而,在长读长平台上生成的读长更长且错误率更高,使得用于组装短读长的工具不适合用于数据组装,并推动了新方法的开发。我们介绍了 LoReTTA(长读长模板靶向汇编器),这是一种工具,旨在对 PacBio 平台上的病毒基因组生成的长读长进行从头组装。LoReTTA 利用参考基因​​组来指导组装过程,这种方法在短读长方面取得了成功。该工具旨在处理源自病毒基因组的读取,其特点是遗传变异性高、可能存在多种同种型,以及临床或环境样本中主要存在其他生物。LoReTTA 在一系列模拟和实验数据集上进行了测试,在组装连续性和准确性方面优于已建立的长读组装器。该软件在 Linux 操作系统下运行,旨在轻松适应替代系统,并具有自动安装管道,可处理所需的依赖项。命令行版本和用户友好的图形界面版本可在 GPLv3 许可下在 https://bioinformatics.cvr.ac.uk/software/ 获得手册和测试数据集。LoReTTA 在一系列模拟和实验数据集上进行了测试,在组装连续性和准确性方面优于已建立的长读组装器。该软件在 Linux 操作系统下运行,旨在轻松适应替代系统,并具有自动安装管道,可处理所需的依赖项。命令行版本和用户友好的图形界面版本可在 GPLv3 许可下在 https://bioinformatics.cvr.ac.uk/software/ 获得手册和测试数据集。LoReTTA 在一系列模拟和实验数据集上进行了测试,在组装连续性和准确性方面优于已建立的长读组装器。该软件在 Linux 操作系统下运行,旨在轻松适应替代系统,并具有自动安装管道,可处理所需的依赖项。命令行版本和用户友好的图形界面版本可在 GPLv3 许可下在 https://bioinformatics.cvr.ac.uk/software/ 获得手册和测试数据集。并具有自动安装管道,可处理所需的依赖项。命令行版本和用户友好的图形界面版本可在 GPLv3 许可下在 https://bioinformatics.cvr.ac.uk/software/ 获得手册和测试数据集。并具有自动安装管道,可处理所需的依赖项。命令行版本和用户友好的图形界面版本可在 GPLv3 许可下在 https://bioinformatics.cvr.ac.uk/software/ 获得手册和测试数据集。
更新日期:2021-04-22
down
wechat
bug