A software tool 'CroCo' detects pervasive cross-species contamination in next generation sequencing data

Paul Simion; Khalid Belkhir; Clémentine François; Julien Veyssier; Jochen C Rink; Michaël Manuel; Hervé Philippe; Maximilian J Telford

doi:10.1186/s12915-018-0486-7

A software tool 'CroCo' detects pervasive cross-species contamination in next generation sequencing data

BMC Biol. 2018 Mar 5;16(1):28. doi: 10.1186/s12915-018-0486-7.

Authors

Paul Simion^{1

2}, Khalid Belkhir¹, Clémentine François¹, Julien Veyssier¹, Jochen C Rink³, Michaël Manuel², Hervé Philippe^{4

5}, Maximilian J Telford⁶

Affiliations

¹ Institut des Sciences de l'Evolution (ISEM), UMR 5554, CNRS, IRD, EPHE, Université de Montpellier, Montpellier, France.
² Sorbonne Université, CNRS, Institut de Biologie Paris-Seine (IBPS), Evolution Paris-Seine (UMR7138), Case 05, 7 Quai St Bernard, 75005, Paris, France.
³ Max Plank Institute of Molecular Cell Biology and Genetics, Pfotenhauerstrasse 108, 01307, Dresden, Germany.
⁴ Centre de Théorisation et de Modélisation de la Biodiversité, Station d'Ecologie Théorique et Expérimentale, UMR CNRS 5321, Moulis, 09200, France.
⁵ Département de Biochimie, Centre Robert-Cedergren, Université de Montréal, Montréal, H3C 3J7, Québec, Canada.
⁶ Centre for Life's Origins and Evolution, Department of Genetics Evolution and Environment, University College London, Darwin Building, Gower Street, London, WC1E 6BT, UK. m.telford@ucl.ac.uk.

Abstract

Background: Multiple RNA samples are frequently processed together and often mixed before multiplex sequencing in the same sequencing run. While different samples can be separated post sequencing using sample barcodes, the possibility of cross contamination between biological samples from different species that have been processed or sequenced in parallel has the potential to be extremely deleterious for downstream analyses.

Results: We present CroCo, a software package for identifying and removing such cross contaminants from assembled transcriptomes. Using multiple, recently published sequence datasets, we show that cross contamination is consistently present at varying levels in real data. Using real and simulated data, we demonstrate that CroCo detects contaminants efficiently and correctly. Using a real example from a molecular phylogenetic dataset, we show that contaminants, if not eliminated, can have a decisive, deleterious impact on downstream comparative analyses.

Conclusions: Cross contamination is pervasive in new and published datasets and, if undetected, can have serious deleterious effects on downstream analyses. CroCo is a database-independent, multi-platform tool, designed for ease of use, that efficiently and accurately detects and removes cross contamination in assembled transcriptomes to avoid these problems. We suggest that the use of CroCo should become a standard cleaning step when processing multiple samples for transcriptome sequencing.

Keywords: Contamination; Ctenophora; NGS; Phylogenomics.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Animals
Computational Biology / methods
Computational Biology / standards*
Databases, Genetic / standards*
Gene Expression Profiling / methods
Gene Expression Profiling / standards
High-Throughput Nucleotide Sequencing / methods
High-Throughput Nucleotide Sequencing / standards*
Hydrozoa
Phylogeny*
RNA, Messenger / analysis
RNA, Messenger / genetics*
Software / standards*
Species Specificity

Substances

RNA, Messenger

Abstract

Publication types

MeSH terms

Substances

Grants and funding