Mining environmental metagenomes for deeply novel biological sequence and structure
This is a metagenomic bioprospecting workflow designed to identify environmental DNA and protein sequences that remain unexplained after taxonomic, functional, and structural database searches.
Dr. Todd Treangen, Rice University
Dr. Jennifer Lu, The Johns Hopkins University
Hiba Ben Aribi, Tunis El Manar University
Felix Quintana, Rice University
Natalie Kokroko, Rice University
Jingyue Wu, Baylor College of Medicine
Mohamed Abdelrahim
- Kraken2 Database: k2_pluspf_20260626.tar.gz:
- Wastewater Metagenomic dataset 1: 456.3G bases
- Wastewater Metagenomic dataset 2: 488.7G bases
- Cacao Soil Metagenomics: 25.6G bases
prefetch \
SRR000001 \
--max-size 40g
fasterq-dump \
--temp $TMPDIR \
--split-files SRR000001
megahit \
-1 SRR000001_1.fastq.gz \
-2 SRR000001_2.fastq.gz \
-t $THREADS \
-o SRR000001_megahit
kraken2 \
--db $KRAKEN_DB \
--threads 16 \
--confidence 0.01 \
--report SRR000001_scaffolds.k2report \
--unclassified-out SRR000001_k2unclassified.fa \
--output SRR000001_scaffolds.k2
SRR000001_scaffolds.fasta
seqscreen \
--fasta SRR000001_k2unclassified.fa \
--databases SeqScreenDB_23.4 \
--working seqscreen_out \
--threads 16 \
- Run full pipeline (start to finish) on cacao soil test-set
- Run full pipeline (start to finish) on Wastewater dataset
- Incorporate human genome removal (bowtie2 against T2T) pre-assembly/classification
- Compare assembly pre/post Kraken2 classification
- Compare megahit vs. metaspades vs. ggcat for assembly
