eCite Digital Repository

Differentially expressed genes from RNASeq and functional enrichment results are affected by the choice of single-end versus paired-end reads and stranded versus non-stranded protocols

Citation

Corley, SM and MacKenzie, KL and Beverdam, A and Roddam, LF and Wilkins, MR, Differentially expressed genes from RNASeq and functional enrichment results are affected by the choice of single-end versus paired-end reads and stranded versus non-stranded protocols, Bmc Genomics, 18 pp. 1-13. ISSN 1471-2164 (2017) [Refereed Article]


Preview
PDF (RNAseq data analysis)
17Mb
  

Copyright Statement

Copyright 2017 The Author(s). Licensed under Creative Commons Attribution 4.0 International License (CC BY 4.0) http://creativecommons.org/licenses/by/4.0/

DOI: doi:10.1186/s12864-017-3797-0

Abstract

Background: RNA-Seq is now widely used as a research tool. Choices must be made whether to use paired-end (PE) or single-end (SE) sequencing, and whether to use strand-specific or non-specific (NS) library preparation kits. To date there has been no analysis of the effect of these choices on identifying differentially expressed genes (DEGs) between controls and treated samples and on downstream functional analysis.

Results: We undertook four mammalian transcriptomics experiments to compare the effect of SE and PE protocols on read mapping, feature counting, identification of DEGs and functional analysis. For three of these experiments we also compared a non-stranded (NS) and a strand-specific approach to mapping the paired-end data. SE mapping resulted in a reduced number of reads mapped to features, in all four experiments, and lower read count per gene. Up to 4.3% of genes in the SE data and up to 12.3% of genes in the NS data had read counts which were significantly different compared to the PE data. Comparison of DEGs showed the presence of false positives (average 5%, using voom) and false negatives (average 5%, using voom) using the SE reads. These increased further, by one or two percentage points, with the NS data. Gene ontology functional enrichment (GO) of the DEGs arising from SE or NS approaches, revealed striking differences in the top 20 GO terms, with as little as 40% concordance with PE results. Caution is therefore advised in the interpretation of such results. By comparison, there was overall consistency in gene set enrichment analysis results.

Conclusions: A strand-specific protocol should be used in library preparation to generate the most reliable and accurate profile of expression. Ideally PE reads are also recommended particularly for transcriptome assembly. Whilst SE reads produce a DEG list with around 5% of false positives and false negatives, this method can substantially reduce sequencing cost and this saving could be used to increase the number of biological replicates thereby increasing the power of the experiment. As SE reads, when used in association with gene set enrichment, can generate accurate biological results, this may be a desirable trade-off.

Item Details

Item Type:Refereed Article
Keywords:RNA-Seq, Transcriptomics, Paired-end reads, Single-end reads, Differential expression, Strand-specific, Non-strand-specific
Research Division:Medical and Health Sciences
Research Group:Medical Microbiology
Research Field:Medical Bacteriology
Objective Division:Health
Objective Group:Clinical Health (Organs, Diseases and Abnormal Conditions)
Objective Field:Respiratory System and Diseases (incl. Asthma)
Author:Roddam, LF (Dr Louise Roddam)
ID Code:117257
Year Published:2017
Web of Science® Times Cited:1
Deposited By:Medicine (Discipline)
Deposited On:2017-06-05
Last Modified:2017-09-01
Downloads:7 View Download Statistics

Repository Staff Only: item control page