ampliconflow

opens the authoritative record at ENA, SRA or BioSample; this page never replaces it

PRJNA630822 released 25 Sept 2026

PRJNA630822

Pipeline 0.1.0 Contract v0.1 Licence CC-BY-4.0

Tags

derived from the release metadata, not hand-written

Study

The study record carries no description, and its registered title is the accession itself. The figures below are computed from the 16 released samples and 16 runs.

Samples
16
Runs
16
Collection
2015-01 to 2015-01

Linked publication

'<i>Candidatus</i> Megaira' are diverse symbionts of algae and ciliates with the potential for defensive symbiosis.

10.1099/mgen.0.000950 · 2023 · via europepmc

Linked by the enrich stage. Fields taken from the paper: primers, subfragment. No per-sample coordinates in the paper.

Linked by the enrich stage from europepmc.

Location

sampling sites from the release coordinates
16sampling sites · drag to pan, scroll or pinch to zoom

Place name hierarchy

Latitude -77.65 to -76.0663, longitude 161.007 to 163.116. Names resolved with OpenStreetMap Nominatim from the release's own coordinates; Coordinates parsed from the released lat_lon text field, which the pipeline keeps but does not split into latitude/longitude columns.

How the sequences were obtained

sample to release
01 study metadata

Sample collection

16 samples, 2015-01 to 2015-01

-77.6500 to -76.0663 N, 161.0070 to 163.1160 E

02 not reported

Storage

not reported

neither the archive nor the linked paper states storage conditions

03 not reported

Processing

not reported

no extraction kit or lysis protocol in the archive or the linked paper

04 per-run QC

PCR

16S rRNA V4

primers: present; polymerase, cycle count and primer sequences are not stated in the linked paper

05 not reported

Sequencing preparation

not reported

no library kit or index strategy in the archive or the linked paper

06 study metadata

Sequencing

Illumina MiSeq

16 runs; PAIRED 291.0 bp reads

07 this release

Denoising

dada2 1.38.0

4,770 ASVs from 488,523 reads

ampliconflow branches off at step 6, Sequencing

this release

ampliconflow starts here: 488,523 reads from 16 runs, QC to 96.0% 16S identity and 95.1% above Q30, primers trimmed, dada2 1.38.0 to 4,770 ASVs, taxonomy against a SINTAX reference, then the release (CC-BY-4.0).

Steps 1 to 6 are how the sequences were obtained, from the study's own archive metadata. Step 7 is what the authors report doing with the sequences in the linked paper. ampliconflow starts at the deposited reads rather than repeating the wet lab.

Taxonomy assigned with SILVA 138.2 (SINTAX).

The linked paper's methods run to 5,847 characters. It names no extraction kit, polymerase, cycle count or primer, so those steps stay unreported rather than guessed.

16

Samples

16 runs

4.8k

Features

OTUs at 97%

488.5k

Reads

mapped total

413 MB

Release size

74 files

Reads per sample

log scale
min
10,102
median
25,059
max
64,977

Feature detection

100.0% non-zero

4,770 / 4,770 features

Features with at least one observed read. The remainder are present in the reference set but not detected in these samples.

Composition

Taxonomy not assigned in this artifact

The count table and metadata are complete, but the reference set in this release carries empty taxonomy labels, so composition cannot be shown. The taxonomy stage of the pipeline has not been run against this build.

ranks present in features.parquet: domain, phylum, class, order, family, genus, species

Downstream QC and analysis

computed from the released tables

Eleven modules, computed from the released count table, the sample metadata and the per-run QC reports. Each panel states its own n and the test it used; a module that cannot run on this release is shown as a flagged gap rather than an empty frame.

Rarefaction

median with p10 to p90
2k 5k 10k 20k 50k 650 0

Expected richness when 16 samples are subsampled to a common depth, resampled 31 draws. Median 321 features observed at full depth.

Depth against richness

log depth
655 0 reads per sample, log scale

One point per sample. Correlation of log reads with observed features is 0.955, so the depth floor is doing most of the work of deciding how many features a sample shows.

Per-run QC

  • 16S identity 96.0% alignment call per run
  • Q30 rate 95.1% mean Q 36.7
  • Amplicon V4 primers present
  • PhiX 0.0% control spike-in

16 run report(s), n/a GC, 0.2% ambiguous bases.

Diversity

Shannon
5.46
Simpson
0.995
Evenness
0.942
Chao1
321

Median across samples. Observed richness ranges 142 to 655.

Feature prevalence

23 of 4,770 features

present in at least half of the 16 samples (0.5%). 4,440 features appear in one sample only, which is the long tail rarefaction is fighting.

Ordination

pc1 pc2

One point per sample, 16 plotted. Choose the axes and the colour variable; hover a point for its sample id. The legend below the axes names the levels or the numeric range.

What explains each axis

pc1 · 8.9%

  • pH 31.6%
  • shannon 19.6%
  • reads 17.4%
  • observed 13.8%

pc2 · 7.7%

  • pH 28.0%
  • evenness 11.5%
  • shannon 11.0%
  • observed 9.0%

pc3 · 7.4%

  • pH 17.4%
  • evenness 2.2%
  • shannon 0.2%
  • observed 0.1%

Categorical variables use eta-squared (between-group share of the axis), numeric ones the squared Pearson correlation. Computed from the released sample metadata.

Alpha diversity per sample

Shannon
6.2 16 samples

Median Shannon 5.464 across the release; observed richness runs 142 to 655.

Bray-Curtis dissimilarity

16 x 16, darker is closer

Sample order is the release order, 16 labels, and the matrix itself ships as beta_distance.tsv beside the analysis.

Phylogenetic diversity

Not available for this release: no Newick tree beside the table; run the tree stage, then re-run analyze. The legacy analysis computed Faith's PD when the container carried a tree, and this one reports the absence instead of an empty column.

Community states

CLR, k by silhouette

k = 3 silhouette 0.093

  • state 0 3 samples
  • state 1 4 samples
  • state 2 9 samples

Clustered on the centred log-ratio of the top 200 features; 16 samples.

Batch-bias audit

states against

adjusted Rand n/a p = n/a

not enough levels to test

permutations. Every sample here carries both MiSeq and MiniSeq runs, so the batch variable used is the one with two levels, here .

Variance partitioning

mean R2 per feature, CLR
  • pH 0.101
  • depth 0.000

Joint R2 0.101, so the metadata explains a modest slice of the feature variation, and the two variables are not independent of one another.

Effect size

No two-level comparison available.

Taxa against all samples

with

features tested, 0 survive the correction at q ≤ 0.05

Nothing survives. The smallest q is n/a, which is the honest answer for a study whose samples are spread across three years with two to thirteen samples per year.

Group difference and spread

Bray-Curtis, 999 permutations
PERMANOVA pseudo-F
n/a · p n/a
PERMDISP F
n/a · p n/a
Distance decay (Mantel r)
0.324 · p 0.026

    The contamination class separates the communities beyond the spread within each class (PERMANOVA) with no evidence that the spread itself differs (PERMDISP). Distance decay runs over 55.9 to 178.3 km.

    Spatial structure

    observed richness over distance
    Moran's I
    0.0164 · p 0.914
    Gradient response (rho)
    -0.377 · p 0.145 (decreasing)

    variogram, 8 distance bins, semivariance of richness

    Constrained ordination

    contamination class as the constraint

    RDA · R2 0.074p 0.003

    Hellinger-scaled, 1 dummy predictors over 16 samples. The constraint explains 7.4% of the community inertia, 0.008 adjusted.

    CCA · p 0.377

    Chi-square weighted SVD. Not significant here, which is the honest reading at this sample size and predictor count.

    Co-occurrence network

    100 nodes · 287 edges

    positive
    287
    negative
    0
    density
    0.058
    components
    18
    mean degree
    5.7

    Spearman on log1p proportions, 100 features, |r| ≥ 0.50, FDR 0.05.

    Hubs by degree

      Phylogenetic and signal analyses need a tree

      no Newick tree beside the table; run the tree stage, then re-run analyze. The tree stage aligns the ASVs with MAFFT and builds with FastTree, both detected as external tools; neither is installed on the machine this page was built on, so Faith's PD, UniFrac, Pagel's lambda and Blomberg's K are left flagged rather than guessed.

      ASV phylogeny

      0 most abundant of the tree

      No tree in this release.

      ASV panel

      no sequence file

      The representative sequences are not in this release, so length and GC cannot be drawn.

      Most abundant ASVs

      ASVphylumgenusmeanprev.
      ASV_1AcidobacteriotaBlastocatella73.869%
      ASV_2AcidobacteriotaBlastocatella71.350%
      ASV_3VerrucomicrobiotaCandidatus Udaeobacter71.144%
      ASV_4PseudomonadotaSphingomonas64.350%
      ASV_5VerrucomicrobiotaCandidatus Udaeobacter56.850%
      ASV_6BacteroidotaIncertae Sedis56.744%
      ASV_7PseudomonadotaAcidovorax56.125%
      ASV_8AcidobacteriotaBlastocatella54.219%
      ASV_9AcidobacteriotaBlastocatella52.669%
      ASV_10AcidobacteriotaBlastocatella52.313%
      ASV_11PseudomonadotaSphingomonas50.044%
      ASV_12BacteroidotaHymenobacter49.719%

      Similar studies

      composition, metadata, location, shared authors

      Ranked from the released tables: genus composition as Bray-Curtis similarity, shared environment and method terms, the distance between sample centroids, and authors shared with the linked publication.

      Missing or wrong data?

      Report a field that is empty or mistaken, or associate a paper with this study. No account is needed; the contact email is optional and used only to reply about this submission.

      Contribute to PRJNA630822

      Validated automatically where it can be, reviewed by a person where it cannot.

      Downloads

      0 files · sha256 in manifest

      Files are served from https://huggingface.co/datasets/hmacgregor/ampliconflow-releases/resolve/main/PRJNA630822-20260926/ once hosting is wired. Until then the links point at a placeholder base URL.

      Samples

      16 samples
      Sample Collected group Reads Features Shannon State
      SAMN15375003 2015-01-14 52,217 571 6.034 1
      SAMN15375004 2015-01-14 31,103 388 5.567 2
      SAMN15375005 2015-01-16 14,582 202 4.956 2
      SAMN15375006 2015-01-17 13,709 156 4.652 2
      SAMN15375007 2015-01-11 21,978 269 5.252 1
      SAMN15375008 2015-01-11 23,879 319 5.434 1
      SAMN15375009 2015-01-19 52,414 655 6.173 2
      SAMN15375010 2015-01-22 19,733 195 4.952 2
      SAMN15375011 2015-01-17 20,992 291 5.274 2
      SAMN15375012 2015-01-19 26,239 322 5.494 0
      SAMN15375013 2015-01-19 64,977 642 6.201 1
      SAMN15375014 2015-01-19 50,818 527 5.973 0
      SAMN15375015 2015-01-19 15,626 200 4.983 2
      SAMN15375016 2015-01-22 10,102 142 4.588 2
      SAMN15375017 2015-01-11 30,762 353 5.609 2
      SAMN15375018 2015-01-22 39,392 366 5.647 0

      click a column head to sort