opens the authoritative record at ENA, SRA or BioSample; this page never replaces it
PRJDB7978
Pipeline 0.1.0 Contract v0.1 Licence CC-BY-4.0
Tags
- 16S rRNA
- V4
- amplicon
- selection pcr
- paired-end
- Illumina MiSeq
- primers trimmed
- CC-BY-4.0
derived from the release metadata, not hand-written
Study
The study record carries no description, and its registered title is the accession itself. The figures below are computed from the 65 released samples and 65 runs.
- Samples
- 65
- Runs
- 65
- Collection
- 2014-06 to 2014-07
Linked publication
Microbiome analysis of the restricted bacteria in radioactive element-containing water at the Fukushima Daiichi Nuclear Power Station.10.1128/aem.02113-23 · 2024 · via europepmc
Linked by the enrich stage. Fields taken from the paper: primers, subfragment. No per-sample coordinates in the paper.
abstract
A major incident occurred at the Fukushima Daiichi Nuclear Power Station following the tsunami triggered by the Tohoku-Pacific Ocean Earthquake in March 2011, whereby seawater entered the torus room in the basement of the reactor building. Here, we identify and analyze the bacterial communities in the torus room water and several environmental samples. Samples of the torus room water (1 × 10<sup>9</sup> Bq<sup>137</sup>Cs/L) were collected by the Tokyo Electric Power Company Holdings from two sampling points between 30 cm and 1 m from the bottom of the room (TW1) and the bottom layer (TW2). A structural analysis of the bacterial communities based on 16S rRNA amplicon sequencing revealed that the predominant bacterial genera in TW1 and TW2 were similar. TW1 primarily contained the genus <i>Limnobacter</i>, a thiosulfate-oxidizing bacterium. γ-Irradiation tests on <i>Limnobacter thiooxidans</i>, the most closely related phylogenetically found in TW1, indicated that its radiation resistance was similar to ordinary bacteria. TW2 predominantly contained the genus <i>Brevirhabdus</i>, a manganese-oxidizing bacterium. Although bacterial diversity in the torus room water was lower than seawater near Fukushima, ~70% of identified genera were associated with metal corrosion. Latent environment allocation-an analytical technique that estimates habitat distributions and co-detection analyses-revealed that the microbial communities in the torus room water originated from a distinct blend of natural marine microbial and artificial bacterial communities typical of biofilms, sludge, and wastewater. Understanding the specific bacteria linked to metal corrosion in damaged plants is important for advancing decommissioning efforts.<h4>Importance</h4>In the context of nuclear power station decommissioning, the proliferation of microorganisms within the reactor and piping systems constitutes a formidable challenge. Therefore, the identification of microbial communities in such environments is of paramount importance. In the aftermath of the Fukushima Daiichi Nuclear Power Station accident, microbial community analysis was conducted on environmental samples collected mainly outside the site. However, analyses using samples from on-site areas, including adjacent soil and seawater, were not performed. This study represents the first comprehensive analysis of microbial communities, utilizing meta 16S amplicon sequencing, with a focus on environmental samples collected from the radioactive element-containing water in the torus room, including the surrounding environments. Some of the identified microbial genera are shared with those previously identified in spent nuclear fuel pools in countries such as France and Brazil. Moreover, our discussion in this paper elucidates the correlation of many of these bacteria with metal corrosion.
Linked by the enrich stage from europepmc.
Location
sampling sites from the release coordinatesNo sample coordinates in this release
The study reports no latitude or longitude, and its metadata carries no lat_lon text field either, so no map or place-name hierarchy can be drawn. Every other panel is computed from the count table and the sample metadata.
How the sequences were obtained
sample to releaseSample collection
65 samples, 2014-06 to 2014-07
Storage
not reported
no storage statement in the study metadata
Processing
not reported
no extraction kit or lysis protocol recorded
PCR
16S rRNA V4
primers: trimmed
Sequencing preparation
not reported
no library kit or index strategy recorded
Sequencing
Illumina MiSeq
65 runs; PAIRED 151.0 bp reads
Denoising
dada2 1.38.0
40,327 ASVs from 1,822,004 reads
ampliconflow branches off at step 6, Sequencing
this releaseampliconflow starts here: 1,822,004 reads from 65 runs, QC to 82.5% 16S identity and 92.1% above Q30, primers trimmed, dada2 1.38.0 to 40,327 ASVs, taxonomy against a SINTAX reference, then the release (CC-BY-4.0).
Steps 1 to 6 are how the sequences were obtained, from the study's own archive metadata. Step 7 is what the authors report doing with the sequences in the linked paper. ampliconflow starts at the deposited reads rather than repeating the wet lab.
65
Samples
65 runs
40.3k
Features
OTUs at 97%
1.8M
Reads
mapped total
70 MB
Release size
163 files
Depth floor
1,000 reads
no samples below
QC warnings
65
100% of runs warned
Reads per sample
log scale- min
- 20,050
- median
- 26,979
- max
- 39,268
Feature detection
100.0% non-zero40,327 / 40,327 features
Features with at least one observed read. The remainder are present in the reference set but not detected in these samples.
Composition
Top phyla
- Pseudomonadota 497,374 (27.3%)
- Acidobacteriota 314,304 (17.3%)
- Actinomycetota 281,967 (15.5%)
- Planctomycetota 149,875 (8.2%)
- Bacteroidota 117,440 (6.4%)
- Chloroflexota 96,683 (5.3%)
- Verrucomicrobiota 91,095 (5.0%)
- Myxococcota 80,674 (4.4%)
- Gemmatimonadota 45,061 (2.5%)
- Bacillota 38,070 (2.1%)
- Methylomirabilota 17,276 (0.9%)
- Latescibacterota 16,338 (0.9%)
- Thermodesulfobacteriota 15,446 (0.8%)
- Bdellovibrionota 9,896 (0.5%)
- Thermoproteota 6,072 (0.3%)
- Nitrospirota 5,969 (0.3%)
- Armatimonadota 5,615 (0.3%)
- Chlamydiota 4,931 (0.3%)
- Cyanobacteriota 4,862 (0.3%)
- MBNT15 3,916 (0.2%)
Top genera
- Incertae Sedis 997,235 (54.7%)
- Candidatus Udaeobacter 43,411 (2.4%)
- Bradyrhizobium 34,985 (1.9%)
- Ellin6067 31,415 (1.7%)
- Acidibacter 23,625 (1.3%)
- Acidiferrimicrobium 18,666 (1.0%)
- Gemmatimonas 18,398 (1.0%)
- Candidatus Solibacter 17,991 (1.0%)
- Rhizobacter 15,846 (0.9%)
- Acidothermus 14,452 (0.8%)
- Ferruginibacter 13,929 (0.8%)
- Nocardioides 13,857 (0.8%)
- Chthoniobacter 13,837 (0.8%)
- Gaiella 13,819 (0.8%)
- Bryobacter 12,679 (0.7%)
- Mycobacterium 12,262 (0.7%)
- Puia 12,205 (0.7%)
- Sphingomonas 10,720 (0.6%)
- Pseudolabrys 9,443 (0.5%)
- mle1-7 9,411 (0.5%)
Rank-abundance
log-log40,327 ranked features, top 200 shown. A steep drop means a few taxa carry most of the reads.
Per-sample reads
65 samples- min
- 20,050
- median
- 26,979
- max
- 39,268
Downstream QC and analysis
computed from the released tablesEleven modules, computed from the released count table, the sample metadata and the per-run QC reports. Each panel states its own n and the test it used; a module that cannot run on this release is shown as a flagged gap rather than an empty frame.
Rarefaction
median with p10 to p90Expected richness when 65 samples are subsampled to a common depth, resampled 31 draws. Median 1,275 features observed at full depth.
Depth against richness
log depthOne point per sample. Correlation of log reads with observed features is 0.755, so the depth floor is doing most of the work of deciding how many features a sample shows.
Per-run QC
- 16S identity 82.5% alignment call per run
- Q30 rate 92.1% mean Q 35.8
- Amplicon V4 primers trimmed
- PhiX 0.1% control spike-in
65 run report(s), n/a GC, 0.0% ambiguous bases.
Diversity
- Shannon
- 6.73
- Simpson
- 0.998
- Evenness
- 0.94
- Chao1
- 1,275
Median across samples. Observed richness ranges 985 to 1,566.
Feature prevalence
0 of 40,327 features
present in at least half of the 65 samples (0.0%). 29,003 features appear in one sample only, which is the long tail rarefaction is fighting.
Ordination
One point per sample, 65 plotted. Choose the axes and the colour variable; hover a point for its sample id. The legend below the axes names the levels or the numeric range.
What explains each axis
pc1 · 24.1%
- site 45.9%
- reads 15.0%
- pH 12.9%
- evenness 7.3%
pc2 · 9.9%
- site 39.9%
- reads 8.2%
- habitat 2.3%
- evenness 1.5%
pc3 · 5.0%
- site 11.3%
- observed 2.4%
- chao1 2.4%
- shannon 2.3%
Categorical variables use eta-squared (between-group share of the axis), numeric ones the squared Pearson correlation. Computed from the released sample metadata.
Alpha diversity per sample
ShannonMedian Shannon 6.729 across the release; observed richness runs 985 to 1566.
Bray-Curtis dissimilarity
65 x 65, darker is closerSample order is the release order, 65 labels, and the matrix itself ships as beta_distance.tsv beside the analysis.
Phylogenetic diversity
Faith's PD median n/a
From tree.nwk, range to .
Community states
CLR, k by silhouettek = 2 silhouette 0.510
- state 0 49 samples
- state 1 16 samples
Clustered on the centred log-ratio of the top 200 features; 65 samples.
Batch-bias audit
states againstadjusted Rand n/a p = n/a
not enough levels to test
permutations. Every sample here carries both MiSeq and MiniSeq runs, so the batch variable used is the one with two levels, here .
Variance partitioning
mean R2 per feature, CLR- habitat 0.004
- site 0.446
- pH 0.104
Joint R2 0.463, so the metadata explains a modest slice of the feature variation, and the two variables are not independent of one another.
Effect size
Shannon diversity between rhizosphere (n=33) and surface (n=32)
- Cohen's d
- -0.769 (medium)
- Cliff's delta
- -0.426
- log2 fold change
- -0.015
Means n/a and n/a. The difference is small and the spread is wide, which is what the delta says too.
Taxa against habitat
kruskal with bh200 features tested, 0 survive the correction at q ≤ 0.05
Nothing survives. The smallest q is 0.811, which is the honest answer for a study whose samples are spread across three years with two to thirteen samples per year.
- ASV_1 H 0.67 · p 0.414 · q 0.811
- ASV_2 H 0.57 · p 0.450 · q 0.811
- ASV_3 H 0.45 · p 0.504 · q 0.811
- ASV_4 H 0.72 · p 0.396 · q 0.811
- ASV_5 H 0.48 · p 0.490 · q 0.811
Group difference and spread
Bray-Curtis, 999 permutations- PERMANOVA pseudo-F
- 0.731 · p 0.809
- PERMDISP F
- 8.40 · p 0.096
- Distance decay (Mantel r)
- n/a · p n/a
- rhizosphere32
- surface33
The contamination class separates the communities beyond the spread within each class (PERMANOVA) with no evidence that the spread itself differs (PERMDISP). Distance decay runs over n/a to n/a km.
Spatial structure
observed richness over distancethe study's coordinates fall on a single site (no spread), so spatial autocorrelation is not defined
Constrained ordination
contamination class as the constraintRDA · R2 0.322p 0.001
Hellinger-scaled, 12 dummy predictors over 65 samples. The constraint explains 32.2% of the community inertia, 0.166 adjusted.
CCA · p 0.001
Chi-square weighted SVD. Not significant here, which is the honest reading at this sample size and predictor count.
Co-occurrence network
100 nodes · 4,631 edges
- positive
- 4,568
- negative
- 63
- density
- 0.936
- components
- 1
- mean degree
- 92.6
Spearman on log1p proportions, 100 features, |r| ≥ 0.50, FDR 0.05.
Hubs by degree
Differential abundance
wilcoxon · bh40,327 features tested, 0 survive q ≤ 0.05
rhizosphere against surface, {"rhizosphere":33,"surface":32}. Nothing separates the two classes after correction, so the contamination signal is a community-level shift rather than a handful of marker taxa.
- ASV_10000p 0.340 · q 0.415 · log2FC 1.00
- ASV_10001p 0.340 · q 0.415 · log2FC 1.00
- ASV_10002p 0.340 · q 0.415 · log2FC 1.00
- ASV_10003p 0.340 · q 0.415 · log2FC 1.00
- ASV_10004p 0.340 · q 0.415 · log2FC 1.00
- ASV_10005p 0.154 · q 0.415 · log2FC -1.02
R-backed alternatives kept external: ancombc, deseq2, aldex2, linda, corncob.
Feature ranking and power
by mean- ASV_1mean 120.68
- ASV_2mean 116.37
- ASV_3mean 110.77
- ASV_4mean 64.82
- ASV_5mean 63.23
- ASV_6mean 60.25
- ASV_7mean 60.18
- ASV_8mean 59.46
Power: n/a samples per group for n/a SD at power, alpha . The study was sufficiently sized for a moderate effect; the effect it actually found is small.
Phylogenetic diversity
40,327 tips · MAFFT alignment + FastTree (GTR+CAT)- Faith's PD median
- 129.53
- Faith's PD range
- 108.27 – 150.39
- UniFrac PCo1
- 10.7%
- UniFrac PCo2
- 8.1%
Faith's PD over the observed tips of each sample; the axes are a PCoA of unweighted UniFrac across the 65 samples.
Phylogenetic signal
top 200 features by abundance; trait is mean abundance- Pagel's lambda
- 0.300 · p 0.795
- Blomberg's K
- 0.000 · p n/a
A lambda near 0 means the trait is phylogenetically independent, near 1 means it tracks the tree. Reported scoped to the abundant features because the covariance is quadratic in the feature count.
ASV phylogeny
60 most abundant of the treeThe tree is the release's FastTree over all ASVs; this is the subtree of its 60 most abundant, drawn as a cladogram. Branch lengths are the tree's; the bar is mean abundance in the release. Hover a tip for its taxonomy.
ASV panel
40327 sequencesungapped length · median 250 bp (190–252)
GC content · median 55.6%
Most abundant ASVs
| ASV | phylum | genus | mean | prev. |
|---|---|---|---|---|
| ASV_1 | Pseudomonadota | Bradyrhizobium | 120.7 | 49% |
| ASV_2 | Pseudomonadota | Bradyrhizobium | 116.4 | 25% |
| ASV_3 | Verrucomicrobiota | Candidatus Udaeobacter | 110.8 | 25% |
| ASV_4 | Pseudomonadota | Rhizobacter | 64.8 | 45% |
| ASV_5 | Pseudomonadota | Incertae Sedis | 63.2 | 49% |
| ASV_6 | Pseudomonadota | Bradyrhizobium | 60.3 | 45% |
| ASV_7 | Acidobacteriota | Incertae Sedis | 60.2 | 39% |
| ASV_8 | Verrucomicrobiota | Candidatus Udaeobacter | 59.5 | 45% |
| ASV_9 | Acidobacteriota | Incertae Sedis | 58.6 | 45% |
| ASV_10 | Pseudomonadota | Incertae Sedis | 54.6 | 49% |
| ASV_11 | Chloroflexota | Incertae Sedis | 52.7 | 49% |
| ASV_12 | Pseudomonadota | Incertae Sedis | 52.6 | 25% |
Similar studies
composition, metadata, location, shared authors- taxonomy 79% similar (genus)
- same region (V4)
- shared author(s): h, t
33 samples
- taxonomy 80% similar (genus)
- same region (V4)
60 samples
- taxonomy 65% similar (genus)
- same region (V4)
31 samples
Ranked from the released tables: genus composition as Bray-Curtis similarity, shared environment and method terms, the distance between sample centroids, and authors shared with the linked publication.
Missing or wrong data?
Report a field that is empty or mistaken, or associate a paper with this study. No account is needed; the contact email is optional and used only to reply about this submission.
Downloads
8 files · sha256 in manifest- Count table 268 KB tables/PRJDB7978.parquet
- Count table, BIOM 728 KB tables/PRJDB7978.biom.gz
- Taxonomy 2.4 MB features.parquet
- Taxonomy (TSV) 2.6 MB taxonomy.tsv.gz
- Sample metadata 19 KB samples.parquet
- Run metadata 18 KB runs.parquet
- Sequences (fasta) 1.4 MB sequences/PRJDB7978.fasta.gz
- Manifest manifest.json
Files are served from https://huggingface.co/datasets/hmacgregor/ampliconflow-releases/resolve/main/PRJDB7978-20260926/ once hosting is wired. Until then the links point at a placeholder base URL.
Samples
| Sample | Collected | habitat | Reads | Features | Shannon | State |
|---|---|---|---|---|---|---|
| SAMD00160688 | 2014-06-06 | surface | 31,724 | 1,423 | 6.82 | 0 |
| SAMD00160689 | 2014-06-06 | surface | 31,899 | 1,344 | 6.748 | 0 |
| SAMD00160690 | 2014-06-06 | surface | 39,268 | 1,534 | 6.886 | 0 |
| SAMD00160691 | 2014-06-06 | rhizosphere | 33,926 | 1,349 | 6.712 | 0 |
| SAMD00160692 | 2014-06-06 | rhizosphere | 25,423 | 1,148 | 6.618 | 0 |
| SAMD00160693 | 2014-06-06 | rhizosphere | 27,808 | 1,243 | 6.673 | 0 |
| SAMD00160694 | 2014-06-06 | surface | 33,902 | 1,379 | 6.76 | 1 |
| SAMD00160695 | 2014-06-06 | surface | 37,703 | 1,483 | 6.86 | 1 |
| SAMD00160696 | 2014-06-06 | surface | 35,484 | 1,394 | 6.808 | 1 |
| SAMD00160697 | 2014-06-06 | rhizosphere | 29,118 | 1,146 | 6.587 | 1 |
| SAMD00160698 | 2014-06-06 | rhizosphere | 27,454 | 1,298 | 6.706 | 0 |
| SAMD00160699 | 2014-06-06 | rhizosphere | 33,061 | 1,258 | 6.668 | 1 |
| SAMD00160700 | 2014-06-06 | surface | 35,767 | 1,345 | 6.701 | 1 |
| SAMD00160701 | 2014-06-06 | surface | 36,647 | 1,304 | 6.643 | 1 |
| SAMD00160702 | 2014-06-06 | surface | 36,872 | 1,395 | 6.743 | 1 |
| SAMD00160703 | 2014-06-06 | rhizosphere | 26,965 | 1,240 | 6.587 | 0 |
| SAMD00160704 | 2014-06-06 | rhizosphere | 28,324 | 1,162 | 6.526 | 1 |
| SAMD00160705 | 2014-06-06 | rhizosphere | 30,086 | 1,161 | 6.575 | 1 |
| SAMD00160706 | 2014-06-28 | surface | 25,706 | 1,332 | 6.779 | 0 |
| SAMD00160707 | 2014-06-28 | surface | 29,612 | 1,479 | 6.853 | 0 |
| SAMD00160708 | 2014-06-28 | surface | 20,679 | 1,044 | 6.559 | 0 |
| SAMD00160709 | 2014-06-28 | rhizosphere | 31,331 | 1,413 | 6.815 | 0 |
| SAMD00160710 | 2014-06-28 | rhizosphere | 27,372 | 1,293 | 6.74 | 0 |
| SAMD00160711 | 2014-06-28 | rhizosphere | 26,483 | 1,176 | 6.639 | 1 |
| SAMD00160712 | 2014-06-28 | surface | 31,718 | 1,513 | 6.882 | 0 |
| SAMD00160713 | 2014-06-28 | surface | 24,184 | 1,142 | 6.662 | 0 |
| SAMD00160714 | 2014-06-28 | surface | 26,963 | 1,250 | 6.743 | 0 |
| SAMD00160715 | 2014-06-28 | rhizosphere | 26,154 | 1,270 | 6.692 | 0 |
| SAMD00160716 | 2014-06-28 | rhizosphere | 28,385 | 1,312 | 6.766 | 0 |
| SAMD00160717 | 2014-06-28 | rhizosphere | 26,391 | 1,319 | 6.788 | 0 |
| SAMD00160718 | 2014-06-28 | surface | 26,533 | 1,275 | 6.729 | 0 |
| SAMD00160719 | 2014-06-28 | surface | 25,276 | 1,272 | 6.738 | 0 |
| SAMD00160720 | 2014-06-28 | surface | 32,315 | 1,493 | 6.851 | 0 |
| SAMD00160721 | 2014-06-28 | rhizosphere | 25,016 | 1,182 | 6.676 | 0 |
| SAMD00160722 | 2014-06-28 | rhizosphere | 22,038 | 1,078 | 6.61 | 0 |
| SAMD00160723 | 2014-06-28 | rhizosphere | 21,799 | 1,092 | 6.6 | 0 |
| SAMD00160724 | 2014-06-28 | surface | 25,466 | 1,293 | 6.753 | 0 |
| SAMD00160725 | 2014-06-28 | surface | 25,904 | 1,210 | 6.669 | 0 |
| SAMD00160726 | 2014-06-28 | surface | 25,867 | 1,137 | 6.651 | 1 |
| SAMD00160727 | 2014-06-28 | rhizosphere | 24,182 | 1,232 | 6.708 | 0 |
| SAMD00160728 | 2014-06-28 | rhizosphere | 24,791 | 1,258 | 6.742 | 0 |
| SAMD00160729 | 2014-06-28 | rhizosphere | 25,153 | 1,294 | 6.756 | 0 |
| SAMD00160730 | 2014-06-25 | surface | 27,241 | 1,238 | 6.755 | 1 |
| SAMD00160731 | 2014-06-25 | surface | 28,696 | 1,498 | 6.886 | 0 |
| SAMD00160732 | 2014-06-25 | surface | 29,741 | 1,371 | 6.815 | 1 |
| SAMD00160733 | 2014-06-25 | rhizosphere | 26,979 | 1,376 | 6.829 | 0 |
| SAMD00160734 | 2014-06-25 | rhizosphere | 31,610 | 1,566 | 6.92 | 0 |
| SAMD00160735 | 2014-06-25 | rhizosphere | 29,329 | 1,260 | 6.736 | 1 |
| SAMD00160736 | 2014-06-25 | surface | 30,913 | 1,357 | 6.765 | 0 |
| SAMD00160737 | 2014-06-25 | surface | 24,866 | 1,133 | 6.619 | 0 |
| SAMD00160738 | 2014-06-25 | surface | 30,866 | 1,366 | 6.771 | 0 |
| SAMD00160739 | 2014-06-25 | rhizosphere | 28,806 | 1,281 | 6.712 | 0 |
| SAMD00160740 | 2014-06-25 | rhizosphere | 31,427 | 1,334 | 6.731 | 0 |
| SAMD00160741 | 2014-06-25 | rhizosphere | 26,113 | 1,181 | 6.608 | 0 |
| SAMD00160742 | 2014-07-19 | surface | 23,509 | 1,220 | 6.663 | 0 |
| SAMD00160743 | 2014-07-19 | surface | 22,787 | 1,231 | 6.656 | 0 |
| SAMD00160745 | 2014-07-19 | rhizosphere | 24,910 | 1,237 | 6.653 | 0 |
| SAMD00160746 | 2014-07-19 | rhizosphere | 26,438 | 1,350 | 6.744 | 0 |
| SAMD00160747 | 2014-07-19 | rhizosphere | 24,668 | 1,208 | 6.603 | 0 |
| SAMD00160748 | 2014-07-19 | surface | 25,240 | 1,285 | 6.772 | 0 |
| SAMD00160749 | 2014-07-19 | surface | 21,728 | 1,046 | 6.6 | 0 |
| SAMD00160750 | 2014-07-19 | surface | 28,128 | 1,417 | 6.862 | 0 |
| SAMD00160751 | 2014-07-19 | rhizosphere | 23,967 | 1,154 | 6.67 | 0 |
| SAMD00160752 | 2014-07-19 | rhizosphere | 23,243 | 1,022 | 6.553 | 1 |
| SAMD00160753 | 2014-07-19 | rhizosphere | 20,050 | 985 | 6.52 | 0 |
click a column head to sort