opens the authoritative record at ENA, SRA or BioSample; this page never replaces it
PRJNA361046
Pipeline 0.1.0 Contract v0.1 Licence CC-BY-4.0
Tags
- 16S rRNA
- V4
- amplicon
- selection pcr
- paired-end
- Illumina MiSeq
- primers trimmed
- CC-BY-4.0
derived from the release metadata, not hand-written
Study
The study record carries no description, and its registered title is the accession itself. The figures below are computed from the 60 released samples and 60 runs.
- Samples
- 60
- Runs
- 60
- Collection
- 2014-12 to 2014-12
Linked publication
Community- and genome-based evidence for a shaping influence of redox potential on bacterial protein evolution.10.1128/msystems.00014-23 · 2023 · via europepmc
Linked by the enrich stage. Fields taken from the paper: . No per-sample coordinates in the paper.
abstract
Despite deep interest in how environments shape microbial communities, whether redox conditions influence the sequence composition of genomes is not well known. We predicted that the carbon oxidation state (<i>Z</i><sub>C</sub>) of protein sequences would be positively correlated with redox potential (Eh). To test this prediction, we used taxonomic classifications for 68 publicly available 16S rRNA gene sequence data sets to estimate the abundances of archaeal and bacterial genomes in river & seawater, lake & pond, geothermal, hyperalkaline, groundwater, sediment, and soil environments. Locally, <i>Z</i><sub>C</sub> of community reference proteomes (i.e., all the protein sequences in each genome, weighted by taxonomic abundances but not by protein abundances) is positively correlated with Eh corrected to pH 7 (Eh7) for the majority of data sets for bacterial communities in each type of environment, and global-scale correlations are positive for bacterial communities in all environments. In contrast, archaeal communities show approximately equal frequencies of positive and negative correlations in individual data sets, and a positive pan-environmental correlation for archaea only emerges after limiting the analysis to samples with reported oxygen concentrations. These results provide empirical evidence that geochemistry modulates genome evolution and may have distinct effects on bacteria and archaea. IMPORTANCE The identification of environmental factors that influence the elemental composition of proteins has implications for understanding microbial evolution and biogeography. Millions of years of genome evolution may provide a route for protein sequences to attain incomplete equilibrium with their chemical environment. We developed new tests of this chemical adaptation hypothesis by analyzing trends of the carbon oxidation state of community reference proteomes for microbial communities in local- and global-scale redox gradients. The results provide evidence for widespread environmental shaping of the elemental composition of protein sequences at the community level and establish a rationale for using thermodynamic models as a window into geochemical effects on microbial community assembly and evolution.
Linked by the enrich stage from europepmc.
Location
sampling sites from the release coordinatesPlace name hierarchy
- ▸ 中国
- › 云南省
- › 会泽县
Districts named on the samples
- 会泽县 1
Latitude 26.12 to 26.41, longitude 103.53 to 103.53. Names resolved with OpenStreetMap Nominatim from the release's own coordinates; Coordinates parsed from the released lat_lon text field, which the pipeline keeps but does not split into latitude/longitude columns.
How the sequences were obtained
sample to releaseSample collection
60 samples, 2014-12 to 2014-12
26.1200 to 26.4100 N, 103.5300 to 103.5300 E
Storage
not reported
neither the archive nor the linked paper states storage conditions
Processing
not reported
no extraction kit or lysis protocol in the archive or the linked paper
PCR
16S rRNA V4
primers: trimmed; polymerase, cycle count and primer sequences are not stated in the linked paper
Sequencing preparation
not reported
no library kit or index strategy in the archive or the linked paper
Sequencing
Illumina MiSeq
60 runs; PAIRED 253.0 bp reads; the linked paper's text supports 454
Denoising
dada2 1.38.0
27,265 ASVs from 1,325,060 reads
ampliconflow branches off at step 6, Sequencing
this releaseampliconflow starts here: 1,325,060 reads from 60 runs, QC to 97.5% 16S identity and 98.0% above Q30, primers trimmed, dada2 1.38.0 to 27,265 ASVs, taxonomy against a SINTAX reference, then the release (CC-BY-4.0).
Steps 1 to 6 are how the sequences were obtained, from the study's own archive metadata. Step 7 is what the authors report doing with the sequences in the linked paper. ampliconflow starts at the deposited reads rather than repeating the wet lab.
Taxonomy assigned with SILVA 138.2. (SINTAX).
The linked paper's methods run to 21,799 characters. It supports: platform.
Conflict on sequencing platform: the release has Illumina MiSeq, the linked paper's text supports 454.
60
Samples
60 runs
27.3k
Features
OTUs at 97%
1.3M
Reads
mapped total
145 MB
Release size
149 files
Depth floor
1,000 reads
no samples below
QC warnings
60
100% of runs warned
Reads per sample
log scale- min
- 6,556
- median
- 23,351
- max
- 37,041
Feature detection
100.0% non-zero27,265 / 27,265 features
Features with at least one observed read. The remainder are present in the reference set but not detected in these samples.
Composition
Top phyla
- Chloroflexota 457,676 (34.5%)
- Pseudomonadota 215,700 (16.3%)
- Acidobacteriota 174,381 (13.2%)
- Planctomycetota 65,549 (4.9%)
- Actinomycetota 44,127 (3.3%)
- Thermodesulfobacteriota 44,088 (3.3%)
- Verrucomicrobiota 42,360 (3.2%)
- Myxococcota 31,637 (2.4%)
- Bacteroidota 30,783 (2.3%)
- Methylomirabilota 25,337 (1.9%)
- Nitrospirota 23,385 (1.8%)
- Thermoproteota 16,197 (1.2%)
- Bacillota 14,904 (1.1%)
- Gemmatimonadota 14,282 (1.1%)
- Cyanobacteriota 13,357 (1.0%)
- Candidatus Eremiobacterota 9,596 (0.7%)
- Latescibacterota 9,459 (0.7%)
- Armatimonadota 9,089 (0.7%)
- Campylobacterota 7,667 (0.6%)
- Thermoplasmatota 7,390 (0.6%)
Top genera
- Incertae Sedis 825,887 (62.3%)
- RBG-16-58-14 41,863 (3.2%)
- HSB OF53-F07 26,553 (2.0%)
- Anaerolinea 20,142 (1.5%)
- Thiobacillus 19,198 (1.4%)
- Aminicenans 14,566 (1.1%)
- Candidatus Udaeobacter 12,672 (1.0%)
- Leptolinea 11,077 (0.8%)
- Anaeromyxobacter 10,813 (0.8%)
- Candidatus Solibacter 10,159 (0.8%)
- Sh765B-TzT-35 9,156 (0.7%)
- Acidothermus 8,849 (0.7%)
- Desulfobacca 8,696 (0.7%)
- Sulfurifustis 8,147 (0.6%)
- Gallionella 6,368 (0.5%)
- Sphingomonas 6,283 (0.5%)
- Aggregatilinea 5,872 (0.4%)
- Bryobacter 5,703 (0.4%)
- Acidibacter 5,678 (0.4%)
- Sideroxydans 5,640 (0.4%)
Rank-abundance
log-log27,265 ranked features, top 200 shown. A steep drop means a few taxa carry most of the reads.
Per-sample reads
60 samples- min
- 6,556
- median
- 23,351
- max
- 37,041
Downstream QC and analysis
computed from the released tablesEleven modules, computed from the released count table, the sample metadata and the per-run QC reports. Each panel states its own n and the test it used; a module that cannot run on this release is shown as a flagged gap rather than an empty frame.
Rarefaction
median with p10 to p90Expected richness when 60 samples are subsampled to a common depth, resampled 31 draws. Median 660 features observed at full depth.
Depth against richness
log depthOne point per sample. Correlation of log reads with observed features is 0.86, so the depth floor is doing most of the work of deciding how many features a sample shows.
Per-run QC
- 16S identity 97.5% alignment call per run
- Q30 rate 98.0% mean Q 37.5
- Amplicon V4 primers trimmed
- PhiX 0.0% control spike-in
60 run report(s), n/a GC, 0.0% ambiguous bases.
Diversity
- Shannon
- 5.86
- Simpson
- 0.994
- Evenness
- 0.91
- Chao1
- 660
Median across samples. Observed richness ranges 243 to 1,072.
Feature prevalence
0 of 27,265 features
present in at least half of the 60 samples (0.0%). 18,150 features appear in one sample only, which is the long tail rarefaction is fighting.
Ordination
One point per sample, 60 plotted. Choose the axes and the colour variable; hover a point for its sample id. The legend below the axes names the levels or the numeric range.
What explains each axis
pc1 · 7.1%
- sample class 56.4%
- shannon 10.3%
- evenness 8.6%
- observed 7.9%
pc2 · 5.0%
- sample class 8.5%
- shannon 5.4%
- reads 5.2%
- observed 4.0%
pc3 · 4.7%
- sample class 21.4%
- reads 14.3%
- observed 10.1%
- chao1 10.1%
Categorical variables use eta-squared (between-group share of the axis), numeric ones the squared Pearson correlation. Computed from the released sample metadata.
Alpha diversity per sample
ShannonMedian Shannon 5.858 across the release; observed richness runs 243 to 1072.
Bray-Curtis dissimilarity
60 x 60, darker is closerSample order is the release order, 60 labels, and the matrix itself ships as beta_distance.tsv beside the analysis.
Phylogenetic diversity
Faith's PD median n/a
From tree.nwk, range to .
Community states
CLR, k by silhouettek = 3 silhouette 0.369
- state 0 52 samples
- state 1 6 samples
- state 2 2 samples
Clustered on the centred log-ratio of the top 200 features; 60 samples.
Batch-bias audit
states againstadjusted Rand n/a p = n/a
not enough levels to test
permutations. Every sample here carries both MiSeq and MiniSeq runs, so the batch variable used is the one with two levels, here .
Variance partitioning
mean R2 per feature, CLR- class 0.188
- depth 0.000
Joint R2 0.188, so the metadata explains a modest slice of the feature variation, and the two variables are not independent of one another.
Effect size
No two-level comparison available.
Taxa against class
withfeatures tested, 0 survive the correction at q ≤ 0.05
Nothing survives. The smallest q is n/a, which is the honest answer for a study whose samples are spread across three years with two to thirteen samples per year.
Group difference and spread
Bray-Curtis, 999 permutations- PERMANOVA pseudo-F
- 1.486 · p 0.001
- PERMDISP F
- 62.85 · p 0.015
- Distance decay (Mantel r)
- -0.014 · p 0.768
- Cd-paddy10
- Cd-sediment1
- Cd-upland9
- control10
- moderate10
- severe10
- unknown10
The contamination class separates the communities beyond the spread within each class (PERMANOVA) with no evidence that the spread itself differs (PERMDISP). Distance decay runs over 11.3 to 32.2 km.
Spatial structure
observed richness over distance- Moran's I
- 0.0501 · p 0.116
- Gradient response (rho)
- 0.368 · p 0.005 (increasing)
variogram, 8 distance bins, semivariance of richness
Constrained ordination
contamination class as the constraintRDA · R2 0.138p 0.001
Hellinger-scaled, 6 dummy predictors over 60 samples. The constraint explains 13.8% of the community inertia, 0.040 adjusted.
CCA · p 0.343
Chi-square weighted SVD. Not significant here, which is the honest reading at this sample size and predictor count.
Co-occurrence network
100 nodes · 3,037 edges
- positive
- 3,037
- negative
- 0
- density
- 0.614
- components
- 4
- mean degree
- 60.7
Spearman on log1p proportions, 100 features, |r| ≥ 0.50, FDR 0.05.
Hubs by degree
Phylogenetic diversity
27,265 tips · MAFFT alignment + FastTree (GTR+CAT)- Faith's PD median
- 76.41
- Faith's PD range
- 41.67 – 115.44
- UniFrac PCo1
- 22.0%
- UniFrac PCo2
- 11.5%
Faith's PD over the observed tips of each sample; the axes are a PCoA of unweighted UniFrac across the 60 samples.
Phylogenetic signal
top 200 features by abundance; trait is mean abundance- Pagel's lambda
- 0.200 · p 0.135
- Blomberg's K
- 0.000 · p n/a
A lambda near 0 means the trait is phylogenetically independent, near 1 means it tracks the tree. Reported scoped to the abundant features because the covariance is quadratic in the feature count.
ASV phylogeny
60 most abundant of the treeThe tree is the release's FastTree over all ASVs; this is the subtree of its 60 most abundant, drawn as a cladogram. Branch lengths are the tree's; the bar is mean abundance in the release. Hover a tip for its taxonomy.
ASV panel
27265 sequencesungapped length · median 233 bp (136–251)
GC content · median 56.0%
Most abundant ASVs
| ASV | phylum | genus | mean | prev. |
|---|---|---|---|---|
| ASV_1 | Chloroflexota | Incertae Sedis | 136.6 | 8% |
| ASV_2 | Chloroflexota | Incertae Sedis | 60.5 | 10% |
| ASV_3 | Acidobacteriota | Incertae Sedis | 50.4 | 3% |
| ASV_4 | Methylomirabilota | Incertae Sedis | 40.7 | 5% |
| ASV_5 | Chloroflexota | RBG-16-58-14 | 38.8 | 10% |
| ASV_6 | Chloroflexota | RBG-16-58-14 | 35.6 | 3% |
| ASV_7 | Pseudomonadota | Sulfurifustis | 33.9 | 3% |
| ASV_8 | Methylomirabilota | Incertae Sedis | 31.2 | 5% |
| ASV_9 | Chloroflexota | Incertae Sedis | 30.9 | 2% |
| ASV_10 | Methylomirabilota | Incertae Sedis | 30.8 | 5% |
| ASV_11 | Methylomirabilota | Incertae Sedis | 29.6 | 2% |
| ASV_12 | Acidobacteriota | Incertae Sedis | 27.9 | 10% |
Similar studies
composition, metadata, location, shared authors- taxonomy 80% similar (genus)
- same region (V4)
65 samples
- taxonomy 73% similar (genus)
- same region (V4)
33 samples
- taxonomy 61% similar (genus)
- same region (V4)
31 samples
Ranked from the released tables: genus composition as Bray-Curtis similarity, shared environment and method terms, the distance between sample centroids, and authors shared with the linked publication.
Missing or wrong data?
Report a field that is empty or mistaken, or associate a paper with this study. No account is needed; the contact email is optional and used only to reply about this submission.
Downloads
8 files · sha256 in manifest- Count table 151 KB tables/PRJNA361046.parquet
- Count table, BIOM 437 KB tables/PRJNA361046.biom.gz
- Taxonomy 1.6 MB features.parquet
- Taxonomy (TSV) 1.8 MB taxonomy.tsv.gz
- Sample metadata 17 KB samples.parquet
- Run metadata 15 KB runs.parquet
- Sequences (fasta) 997 KB sequences/PRJNA361046.fasta.gz
- Manifest manifest.json
Files are served from https://huggingface.co/datasets/hmacgregor/ampliconflow-releases/resolve/main/PRJNA361046-20260926/ once hosting is wired. Until then the links point at a placeholder base URL.
Samples
| Sample | Collected | sample class | Reads | Features | Shannon | State |
|---|---|---|---|---|---|---|
| SAMN06219828 | 2014-12 | control | 23,464 | 587 | 5.801 | 0 |
| SAMN06219829 | 2014-12 | control | 21,767 | 692 | 6.073 | 0 |
| SAMN06219830 | 2014-12 | control | 15,253 | 643 | 6.049 | 0 |
| SAMN06219831 | 2014-12 | control | 26,958 | 645 | 5.841 | 0 |
| SAMN06219832 | 2014-12 | control | 20,348 | 648 | 5.977 | 0 |
| SAMN06219833 | 2014-12 | control | 22,927 | 754 | 6.261 | 0 |
| SAMN06219834 | 2014-12 | control | 25,509 | 543 | 5.271 | 0 |
| SAMN06219835 | 2014-12 | control | 18,976 | 544 | 5.738 | 0 |
| SAMN06219836 | 2014-12 | control | 26,958 | 749 | 6.02 | 0 |
| SAMN06219837 | 2014-12 | control | 8,155 | 399 | 5.603 | 0 |
| SAMN06219838 | 2014-12 | unknown | 6,556 | 243 | 4.979 | 0 |
| SAMN06219839 | 2014-12 | moderate | 14,029 | 532 | 5.776 | 0 |
| SAMN06219840 | 2014-12 | moderate | 14,017 | 489 | 5.633 | 0 |
| SAMN06219841 | 2014-12 | moderate | 34,188 | 798 | 6.115 | 0 |
| SAMN06219842 | 2014-12 | moderate | 7,981 | 323 | 5.37 | 0 |
| SAMN06219843 | 2014-12 | moderate | 7,013 | 283 | 5.235 | 0 |
| SAMN06219844 | 2014-12 | moderate | 18,905 | 563 | 5.851 | 0 |
| SAMN06219845 | 2014-12 | moderate | 37,041 | 882 | 6.08 | 0 |
| SAMN06219846 | 2014-12 | moderate | 31,489 | 812 | 5.922 | 0 |
| SAMN06219847 | 2014-12 | moderate | 25,684 | 718 | 5.793 | 0 |
| SAMN06219848 | 2014-12 | severe | 22,839 | 660 | 5.903 | 0 |
| SAMN06219849 | 2014-12 | severe | 12,901 | 481 | 5.591 | 0 |
| SAMN06219850 | 2014-12 | severe | 27,211 | 662 | 5.765 | 2 |
| SAMN06219851 | 2014-12 | severe | 30,964 | 724 | 5.842 | 0 |
| SAMN06219852 | 2014-12 | severe | 23,351 | 563 | 5.704 | 0 |
| SAMN06219853 | 2014-12 | severe | 24,540 | 681 | 5.866 | 0 |
| SAMN06219854 | 2014-12 | severe | 10,576 | 349 | 5.269 | 0 |
| SAMN06219855 | 2014-12 | severe | 23,765 | 682 | 5.904 | 0 |
| SAMN06219856 | 2014-12 | severe | 30,865 | 855 | 6.108 | 0 |
| SAMN06219857 | 2014-12 | severe | 19,695 | 648 | 5.906 | 0 |
| SAMN06704748 | 2014-12 | Cd-upland | 14,029 | 532 | 5.776 | 0 |
| SAMN06704749 | 2014-12 | Cd-upland | 10,158 | 389 | 5.441 | 0 |
| SAMN06704750 | 2014-12 | Cd-upland | 20,630 | 460 | 5.292 | 0 |
| SAMN06704751 | 2014-12 | Cd-upland | 24,074 | 551 | 5.372 | 0 |
| SAMN06704752 | 2014-12 | Cd-upland | 27,017 | 641 | 5.842 | 0 |
| SAMN06704753 | 2014-12 | Cd-upland | 7,013 | 283 | 5.235 | 0 |
| SAMN06704754 | 2014-12 | Cd-upland | 18,905 | 563 | 5.851 | 0 |
| SAMN06704755 | 2014-12 | Cd-upland | 30,870 | 786 | 6.108 | 0 |
| SAMN06704756 | 2014-12 | Cd-upland | 33,580 | 854 | 6.082 | 0 |
| SAMN06704757 | 2014-12 | Cd-upland | 31,489 | 812 | 5.922 | 0 |
| SAMN06704758 | 2014-12 | Cd-paddy | 22,839 | 660 | 5.903 | 0 |
| SAMN06704759 | 2014-12 | Cd-paddy | 12,901 | 481 | 5.591 | 0 |
| SAMN06704760 | 2014-12 | Cd-paddy | 27,211 | 662 | 5.765 | 2 |
| SAMN06704761 | 2014-12 | Cd-paddy | 30,964 | 724 | 5.842 | 0 |
| SAMN06704762 | 2014-12 | Cd-paddy | 23,351 | 563 | 5.704 | 0 |
| SAMN06704763 | 2014-12 | Cd-paddy | 24,540 | 681 | 5.866 | 0 |
| SAMN06704764 | 2014-12 | Cd-paddy | 19,006 | 523 | 5.302 | 0 |
| SAMN06704765 | 2014-12 | Cd-paddy | 25,115 | 768 | 6.069 | 0 |
| SAMN06704766 | 2014-12 | Cd-paddy | 21,389 | 762 | 6.187 | 0 |
| SAMN06704767 | 2014-12 | Cd-paddy | 27,741 | 801 | 6.092 | 0 |
| SAMN06704768 | 2014-12 | Cd-sediment | 19,166 | 721 | 6.085 | 1 |
| SAMN06704769 | 2014-12 | Cd-sediment | 23,268 | 790 | 6.183 | 0 |
| SAMN06704770 | 2014-12 | Cd-sediment | 25,063 | 794 | 6.022 | 0 |
| SAMN06704771 | 2014-12 | Cd-sediment | 29,305 | 1,003 | 6.43 | 1 |
| SAMN06704772 | 2014-12 | Cd-sediment | 8,479 | 367 | 5.458 | 1 |
| SAMN06704773 | 2014-12 | Cd-sediment | 33,163 | 1,072 | 6.465 | 1 |
| SAMN06704774 | 2014-12 | Cd-sediment | 15,263 | 637 | 6.03 | 1 |
| SAMN06704775 | 2014-12 | Cd-sediment | 28,477 | 966 | 6.403 | 0 |
| SAMN06704776 | 2014-12 | Cd-sediment | 31,647 | 1,003 | 6.403 | 0 |
| SAMN06704777 | 2014-12 | Cd-sediment | 24,482 | 919 | 6.289 | 1 |
click a column head to sort