Home >
Community >
Removing PCR duplicates in RNA-seq Analysis
Upvote
25
Downvote
+ Bioinformatics
+ Biochemistry
+ Rna-seq
Posted by
Kifayat Kifayatkhan
Removing PCR duplicates in RNA-seq Analysis
For normal RNA-seq PCR duplicates are normally kept in, but the duplication rate can be used as a quality control: The higher the duplication rate, the lower the quality. For expression analysis, it is probably best to discard high duplication rate samples, rather than deduplicate them.
In general, the smaller the amount of RNA input into the library prepartion the worst the duplication. Many protocols for very low input quantities (such as single cell) include random barcodes called UMIs (Unique Molecular Identifiers). These allow PCR duplicates to be distinguished from genuinely independent molecules that just happen to have the sample mapping position.
For normal RNA-seq PCR duplicates are normally kept in, but the duplication rate can be used as a quality control: The higher the duplication rate, the lower the quality. For expression analysis, it is probably best to discard high duplication rate samples, rather than deduplicate them.
In general, the smaller the amount of RNA input into the library prepartion the worst the duplication. Many protocols for very low input quantities (such as single cell) include random barcodes called UMIs (Unique Molecular Identifiers). These allow PCR duplicates to be distinguished from genuinely independent molecules that just happen to have the sample mapping position.
Since the OP does not mention which RNA classes the analysis, let me add that UMIs are also extremely important for sRNA, in particular piRNAs, which quite often map to the exact same locations in the genome and have the same sequence. As afar I know current methods to mark duplicates in mapped reads would result in extreme under-reporting of these RNAs.More
Upvote
VOTE
Downvote
Generally you should just leave them as is. One does remove/mark duplicates in DNA seq.
Your answer could be improved by hinting on why, according to the paper you cite, removing "PCR duplicates" shouldn't be done in RNASeq. E.g., by citing the one crucial sentence of the abstract: "We find that a large fraction of computationally identified read duplicates are not PCR duplicates and can be explained by sampling and fragmentation bias."More
Sure..I could have excuse me for short answer.More
Upvote
VOTE
Downvote
Some of the researchers in my institute have researched on this and use UMIs (Unique Molecular Identifiers) to eliminate PCR duplicates. They have been able to increase the reproducibility of the rna seq analysis they had been working on. The UMIs capture some lengths of diverse nucleotides from the rna-sequence. I cannot validate how much UMIs can help in deduplication, as even Nature journal says that UMIs specifically work on the 3’ end. But it has been observed that eliminating duplicates doesn’t improve the accuracy of quantification and it could introduce a bias in the data. So do not eliminate the duplicates, it's just a waste of time.
Some of the researchers in my institute have researched on this and use UMIs (Unique Molecular Identifiers) to eliminate PCR duplicates. They have been able to increase the reproducibility of the rna seq analysis they had been working on. The UMIs capture some lengths of diverse nucleotides from the rna-sequence. I cannot validate how much UMIs can help in deduplication, as even Nature journal says that UMIs specifically work on the 3’ end. But it has been observed that eliminating duplicates doesn’t improve the accuracy of quantification and it could introduce a bias in the data. So do not eliminate the duplicates, it's just a waste of time.
Hello Shirin, I think it's important to provide a reference for the following statement in your answer: "But it has been observed that eliminating duplicates doesn’t improve the accuracy of quantification and it could introduce a bias in the data."More
For normal RNA-seq PCR duplicates are normally kept in, but the duplication rate can be used as a quality control: The higher the duplication rate, the lower the quality. For expression analysis, it is probably best to discard high duplication rate samples, rather than deduplicate them.
In general, the smaller the amount of RNA input into the library prepartion the worst the duplication. Many protocols for very low input quantities (such as single cell) include random barcodes called UMIs (Unique Molecular Identifiers). These allow PCR duplicates to be distinguished from genuinely independent molecules that just happen to have the sample mapping position.
For normal RNA-seq PCR duplicates are normally kept in, but the duplication rate can be used as a quality control: The higher the duplication rate, the lower the quality. For expression analysis, it is probably best to discard high duplication rate samples, rather than deduplicate them.
In general, the smaller the amount of RNA input into the library prepartion the worst the duplication. Many protocols for very low input quantities (such as single cell) include random barcodes called UMIs (Unique Molecular Identifiers). These allow PCR duplicates to be distinguished from genuinely independent molecules that just happen to have the sample mapping position.
More
VOTE
VOTE
Generally you should just leave them as is. One does remove/mark duplicates in DNA seq.
For further read check this Nature paper
Generally you should just leave them as is. One does remove/mark duplicates in DNA seq.
For further read check this Nature paper
More
VOTE
VOTE
VOTE
Some of the researchers in my institute have researched on this and use UMIs (Unique Molecular Identifiers) to eliminate PCR duplicates. They have been able to increase the reproducibility of the rna seq analysis they had been working on. The UMIs capture some lengths of diverse nucleotides from the rna-sequence. I cannot validate how much UMIs can help in deduplication, as even Nature journal says that UMIs specifically work on the 3’ end. But it has been observed that eliminating duplicates doesn’t improve the accuracy of quantification and it could introduce a bias in the data. So do not eliminate the duplicates, it's just a waste of time.
Some of the researchers in my institute have researched on this and use UMIs (Unique Molecular Identifiers) to eliminate PCR duplicates. They have been able to increase the reproducibility of the rna seq analysis they had been working on. The UMIs capture some lengths of diverse nucleotides from the rna-sequence. I cannot validate how much UMIs can help in deduplication, as even Nature journal says that UMIs specifically work on the 3’ end. But it has been observed that eliminating duplicates doesn’t improve the accuracy of quantification and it could introduce a bias in the data. So do not eliminate the duplicates, it's just a waste of time.
More
VOTE
VOTE