opens the authoritative record at ENA, SRA or BioSample; this page never replaces it
PRJNA270841
Pipeline 0.1.0 Contract v0.1 Licence CC-BY-4.0
Tags
- 16S rRNA
- V1-V3
- amplicon
- selection unspecified
- single-end
- 454 GS FLX+
- primers present
- CC-BY-4.0
derived from the release metadata, not hand-written
Study
The study record carries no description, and its registered title is the accession itself. The figures below are computed from the 6 released samples and 10 runs.
- Samples
- 6
- Runs
- 10
- Collection
- 2011-10 to 2011-10
Linked publication
No publication linked for this study.
Location
sampling sites from the release coordinatesPlace name hierarchy
- ▸ Canada
- › Alberta
- › Wood Buffalo
Districts named on the samples
- Wood Buffalo 1
Latitude 56.7 to 56.7, longitude -111.4 to -111.4. Names resolved with OpenStreetMap Nominatim from the release's own coordinates; Coordinates parsed from the released lat_lon text field, which the pipeline keeps but does not split into latitude/longitude columns.
How the sequences were obtained
sample to releaseSample collection
7 samples, 2011-10 to 2011-10
56.7000 to 56.7000 N, -111.4000 to -111.4000 E
Storage
not reported
neither the archive nor the linked paper states storage conditions
Processing
not reported
no extraction kit or lysis protocol in the archive or the linked paper
PCR
16S rRNA V1-V3
primers: present, trimmed; polymerase, cycle count and primer sequences are not stated in the linked paper
Sequencing preparation
not reported
no library kit or index strategy in the archive or the linked paper
Sequencing
454 GS FLX+
10 runs; SINGLE 573.3333 bp reads
Denoising
dada2 1.38.0
2,764 ASVs from 45,447 reads
ampliconflow branches off at step 6, Sequencing
this releaseampliconflow starts here: 45,447 reads from 10 runs, QC to 83.5% 16S identity and 78.9% above Q30, primers trimmed, dada2 1.38.0 to 2,764 ASVs, taxonomy against a SINTAX reference, then the release (CC-BY-4.0).
Steps 1 to 6 are how the sequences were obtained, from the study's own archive metadata. Step 7 is what the authors report doing with the sequences in the linked paper. ampliconflow starts at the deposited reads rather than repeating the wet lab.
Taxonomy assigned with SILVA 138.2 (SINTAX).
6
Samples
10 runs
2.8k
Features
OTUs at 97%
198.8k
Reads
mapped total
95 MB
Release size
62 files
Depth floor
1,000 reads
no samples below
QC warnings
10
100% of runs warned
Reads per sample
log scale- min
- 3,459
- median
- 34,517
- max
- 57,729
Feature detection
100.0% non-zero2,764 / 2,764 features
Features with at least one observed read. The remainder are present in the reference set but not detected in these samples.
Composition
Top phyla
- Pseudomonadota 81,379 (40.9%)
- Thermoproteota 36,891 (18.6%)
- Halobacteriota 30,320 (15.3%)
- Actinomycetota 21,826 (11.0%)
- Methanobacteriota 18,619 (9.4%)
- Bacillota 3,082 (1.6%)
- Thermodesulfobacteriota 2,417 (1.2%)
- Bacteroidota 1,938 (1.0%)
- Chloroflexota 1,477 (0.7%)
- Myxococcota 226 (0.1%)
- Acidobacteriota 215 (0.1%)
- Spirochaetota 124 (0.1%)
- Thermoplasmatota 81 (0.0%)
- Campylobacterota 71 (0.0%)
- Nitrospirota 71 (0.0%)
- Synergistota 25 (0.0%)
- Gemmatimonadota 15 (0.0%)
- Verrucomicrobiota 12 (0.0%)
- Deinococcota 3 (0.0%)
- Hydrogenedentes 3 (0.0%)
Top genera
- Incertae Sedis 63,558 (32.0%)
- Rhodoferax 44,732 (22.5%)
- Candidatus Methanofastidiosum 18,066 (9.1%)
- Methanosarcina 17,703 (8.9%)
- Nocardioides 9,774 (4.9%)
- Thiobacillus 5,706 (2.9%)
- Acidovorax 4,164 (2.1%)
- Microbacterium 2,549 (1.3%)
- Candidatus Symbiobacter 2,339 (1.2%)
- Hydrogenophaga 2,131 (1.1%)
- Rugosibacter 1,495 (0.8%)
- Immundisolibacter 1,324 (0.7%)
- Intrasporangium 1,193 (0.6%)
- Pelotomaculum 1,173 (0.6%)
- Sphingobium 1,107 (0.6%)
- Candidatus Methanoperedens 937 (0.5%)
- Ramlibacter 923 (0.5%)
- Sphingomonas 916 (0.5%)
- Acholeplasma 796 (0.4%)
- Methanothrix 640 (0.3%)
Rank-abundance
log-log2,764 ranked features, top 200 shown. A steep drop means a few taxa carry most of the reads.
Per-sample reads
6 samples- min
- 3,459
- median
- 34,517
- max
- 57,729
Downstream QC and analysis
computed from the released tablesEleven modules, computed from the released count table, the sample metadata and the per-run QC reports. Each panel states its own n and the test it used; a module that cannot run on this release is shown as a flagged gap rather than an empty frame.
Rarefaction
median with p10 to p90Expected richness when 1 samples are subsampled to a common depth, resampled 31 draws. Median 0 features observed at full depth.
Depth against richness
log depthOne point per sample. Correlation of log reads with observed features is , so the depth floor is doing most of the work of deciding how many features a sample shows.
Per-run QC
- 16S identity 83.5% alignment call per run
- Q30 rate 78.9% mean Q 34.6
- Amplicon V1-V3 primers present
- PhiX 0.0% control spike-in
10 run report(s), n/a GC, 0.0% ambiguous bases.
Diversity
- Shannon
- 4.57
- Simpson
- 0.98
- Evenness
- 0.747
- Chao1
- 452
Median across samples. Observed richness ranges 0 to 452.
Feature prevalence
0 of 2,764 features
present in at least half of the 7 samples (0.0%). 452 features appear in one sample only, which is the long tail rarefaction is fighting.
Ordination
One point per sample, 7 plotted. Choose the axes and the colour variable; hover a point for its sample id. The legend below the axes names the levels or the numeric range.
What explains each axis
pc1 · 100.0%
- reads 100.0%
- observed 100.0%
pc2 · 0.0%
- reads 0.0%
- observed 0.0%
pc3 · 0.0%
- reads 0.0%
- observed 0.0%
Categorical variables use eta-squared (between-group share of the axis), numeric ones the squared Pearson correlation. Computed from the released sample metadata.
Alpha diversity per sample
ShannonMedian Shannon 4.568 across the release; observed richness runs 0 to 452.
Bray-Curtis dissimilarity
7 x 7, darker is closerSample order is the release order, 7 labels, and the matrix itself ships as beta_distance.tsv beside the analysis.
Phylogenetic diversity
Not available for this release: no Newick tree beside the table; run the tree stage, then re-run analyze. The legacy analysis computed Faith's PD when the container carried a tree, and this one reports the absence instead of an empty column.
Community states
CLR, k by silhouettek = 2 silhouette 0.857
- state 0 6 samples
- state 1 1 samples
Clustered on the centred log-ratio of the top 200 features; 7 samples.
Batch-bias audit
states againstadjusted Rand 1.0000 p = 0.1530
no strong evidence that the states are the batch
999 permutations. Every sample here carries both MiSeq and MiniSeq runs, so the batch variable used is the one with two levels, here .
Variance partitioning
mean R2 per feature, CLRJoint R2 n/a, so the metadata explains a modest slice of the feature variation, and the two variables are not independent of one another.
Effect size
No two-level comparison available.
Taxa against all samples
withfeatures tested, 0 survive the correction at q ≤ 0.05
Nothing survives. The smallest q is n/a, which is the honest answer for a study whose samples are spread across three years with two to thirteen samples per year.
Phylogenetic and signal analyses need a tree
no Newick tree beside the table; run the tree stage, then re-run analyze. The tree stage aligns the ASVs with MAFFT and builds with FastTree, both detected as external tools; neither is installed on the machine this page was built on, so Faith's PD, UniFrac, Pagel's lambda and Blomberg's K are left flagged rather than guessed.
Similar studies
composition, metadata, location, shared authors- taxonomy 71% similar (genus)
60 samples
- taxonomy 66% similar (genus)
65 samples
- taxonomy 61% similar (genus)
33 samples
- taxonomy 60% similar (genus)
27 samples
- taxonomy 53% similar (genus)
31 samples
- taxonomy 55% similar (genus)
23 samples
Ranked from the released tables: genus composition as Bray-Curtis similarity, shared environment and method terms, the distance between sample centroids, and authors shared with the linked publication.
Missing or wrong data?
Report a field that is empty or mistaken, or associate a paper with this study. No account is needed; the contact email is optional and used only to reply about this submission.
Downloads
20 files · sha256 in manifest- Count table (16S V1-V3) 7.4 KB tables/PRJNA270841.16s-v1-v3.parquet
- Count table (16S region unknown) 1.3 KB tables/PRJNA270841.16s-region-unknown.parquet
- Count table (16S V3) 1.9 KB tables/PRJNA270841.16s-v3.parquet
- Count table (16S V5-V6) 2.7 KB tables/PRJNA270841.16s-v5-v6.parquet
- Count table (16S V6) 1.7 KB tables/PRJNA270841.16s-v6.parquet
- Count table, BIOM (16S V1-V3) 27 KB tables/PRJNA270841.16s-v1-v3.biom.gz
- Count table, BIOM (16S region unknown) 719 B tables/PRJNA270841.16s-region-unknown.biom.gz
- Count table, BIOM (16S V3) 3.0 KB tables/PRJNA270841.16s-v3.biom.gz
- Count table, BIOM (16S V5-V6) 3.6 KB tables/PRJNA270841.16s-v5-v6.biom.gz
- Count table, BIOM (16S V6) 1.5 KB tables/PRJNA270841.16s-v6.biom.gz
- Taxonomy 189 KB features.parquet
- Taxonomy (TSV) 190 KB taxonomy.tsv.gz
- Sample metadata 15 KB samples.parquet
- Run metadata 14 KB runs.parquet
- Sequences (fasta) 86 KB sequences/PRJNA270841.16s-v1-v3.fasta.gz
- Sequences (fasta) 1.2 KB sequences/PRJNA270841.16s-region-unknown.fasta.gz
- Sequences (fasta) 8.6 KB sequences/PRJNA270841.16s-v3.fasta.gz
- Sequences (fasta) 9.6 KB sequences/PRJNA270841.16s-v5-v6.fasta.gz
- Sequences (fasta) 2.7 KB sequences/PRJNA270841.16s-v6.fasta.gz
- Manifest manifest.json
Files are served from https://huggingface.co/datasets/hmacgregor/ampliconflow-releases/resolve/main/PRJNA270841-20260926/ once hosting is wired. Until then the links point at a placeholder base URL.
Samples
| Sample | Collected | group | Reads | Features | Shannon | State |
|---|---|---|---|---|---|---|
| SAMN03269286 | 2011-10 | 0 | 0 | 0 | ||
| SAMN03269287 | 2011-10 | 0 | 0 | 0 | ||
| SAMN03269288 | 2011-10 | 0 | 0 | 0 | ||
| SAMN03269289 | 2011-10 | 0 | 0 | 0 | ||
| SAMN03269290 | 2011-10 | 0 | 0 | 0 | ||
| SAMN03269292 | 2011-10 | 0 | 0 | 0 | ||
| unlisted_6 | None | 45,447 | 452 | 4.568 | 1 |
click a column head to sort