SSU | LSU | |||
|---|---|---|---|---|
| Parc | 15,439,037 | (+ 5,969,967) | not released | (n/a) |
| Ref | 9,183,096 | (+ 6,958,406) | not released | (n/a) |
| Ref NR 99 | 905,628 | (+ 395,133) | not released | (n/a) |
Since release 115, for SSU only SSU Ref NR 99 contains a guide tree. Since release 138.1, for LSU only LSU Ref NR 99 contains a guide tree. Numbers in brackets give the difference to release 138.2.
Information about former releases can be found here.
SSU | LSU | |
|---|---|---|
| Candidates (total) | 21,073,404 | not released |
| RNAmmer | 11,093,849 | not released |
| < 300 bases | 4,001,492 | not released |
| > 2% ambiguities | 62,631 | not released |
| > 2% homopolymers | 224,299 | not released |
| > 2% vector contamination | 2,424 | not released |
| Low alignment identity | 981,731 | not released |
| Total rejected by QC | 5,634,367 | not released |
Sequences have been retrieved from EMBL-EBI/ENA and have been processed using a complex keyword search procedure and sequence based search with RNAmmer profiles. Sequences from the ENA sequence data classes EST, GSS, HTC, HTG, PAT, STD, STS, TSA, CON have been retrieved from a custom snapshot the ENA team kindly made available to us. The snapshot was finished on 19 February 2025. The WGS sequence have been retrieved from the WGS archive on ENA's FTP server. The last synchronisation for SILVA 144 was performed on 28 February 2025.
SSU Ref | SSU Ref NR | LSU Ref | LSU Ref NR | |
|---|---|---|---|---|
| Version | 144 | 144 | 144 | 144 |
| Total | 9,183,096 | 905,628 | not released | not released |
| Bacteria | 6,143,152 | 607,252 | not released | not released |
| Archaea | 171,777 | 32,179 | not released | not released |
| Eukaryota | 2,868,167 | 266,197 | not released | not released |
| Cultured | 40,088 | 40,088 | not released | not released |
| Typestrains | 32,618 | 32,618 | not released | not released |
Updated taxonomy of prokaryotic SSU sequences
Status of Eukaryotic SSU taxonomy
The content of the SILVA Ref NR 99 dataset has undergone changes for SILVA 144. Due to the large increase in the number of available genome sequences, rRNA sequences are are now clustered with 99% similarity, and individual sequences are not guaranteed to be included the SILVA Ref NR 99. Unclustering all of these sequences, as in previous releases, would greatly increase the size of SILVA Ref NR 99 without adding significant amounts of meaningful diversity to the dataset.
SSU Parc contains all aligned sequences with an alignment identity value equal and above 50, an alignment quality value equal and above 40 as well as an base pair score or sequence quality equal and above 30.
To create SSU Ref, all sequences below 1,200 bases for Bacteria and Eukarya and below 900 bases for Archaea or an alignment identity below 70 or an alignment quality value below 50 have been removed in comparison to the SSU Parc. All sequences with an alignment quality value < 75 have been assigned to color group 1 in ARB (red). All typestrains have been assigned to color group 2 in ARB (light blue).
To create SSU Ref NR 99, a 99% identity criterion to remove highly identical sequences was applied using the vsearch tool with a custom sequence order first based on presence in the last release's Ref NR 99 and second based on combination of sequence length (weighted twofold) and quality. For the sorting, the quality of a sequence is determined by ambiguities (50%), overall alignment quality (45%), and homopolymers (5%). The overall alignment quality of the sequence is calculated from its alignment score, alignment identity, and alignment percentage (all equally weighted). Sequences from cultivated species have been preserved in all cases. Detailed information about the SSU Ref NR dataset can be found here.
A guide tree was calculated by adding all sequences to the SSU Ref tree of SILVA release 138.2. For tree calculation, highly variable positions were removed for Bacteria, Archaea, and Eukarya with the respective position variability filters. Position variability filters for Bacteria, Archaea and Eukarya have been calculated and added to the dataset. The tree was extensively manually curated taking into account the latest taxonomic information.
Remark: Before using the alignment for extensive phylogenetic reconstructions all sequences should be checked carefully.
For SILVA 144, we have re-designed the process how the SILVA taxonomy is created. Please see here for a detailed description.
Besides the SILVA and EMBL-EBI/ENA taxonomy, alternative classifications taken from LPSN (03 June 2025), GTDB (version 226), and LTP (version 2024_10) databases are also available in SILVA. On the webpage, the user can switch using the taxonomy menu. In ARB, the different taxonomies can be found in the fields: tax_slv, tax_embl_ebi_ena, tax_gtdb, tax_lpsn and tax_ltp for SILVA, EMBL-EBI/ENA, GTDB, LPSN, and LTP, respectively. The corresponding *_name fields shows the respective sequence name for each entry. Please take into account that GTDB, LPSN, and LTP provide only a subset of the sequences hosted by SILVA. If no taxonomic mapping to GTDB, LPSN, or LTP was available they are assigned as "unclassified" and the respective sequence name equals EMBL-EBI/ENA.
All names of validly described species in the SSU databases have been checked for changes (basonyms, synonyms and orthographical corrections) against the LPSN - List of Prokaryotic names with Standing in Nomenclature service as of 03 June 2025.
The information if a sequence originates from a cultured or type strain has been added to the field strain and is indicated by [T] and [C]. Several sources have been used to compile the information: StrainInfo (February 2025), LPSN - List of Prokaryotic names with Standing in Nomenclature (3 June 2026), The Ribosomal Database Project II (version 11.4) and the Living Tree Project (version 2024_10).
| Source | Information | Tag | Datasets |
|---|---|---|---|
| Living Tree Project | Typestrains (curated) | l[T] | SSU, LSU |
| LPSN | Type sequences | p[T] | SSU |
| NCBI | Genomes | n[G] | SSU, LSU |
| NCBI | WGS genomes | w[G] | SSU, LSU |
| RDP II | Typestrains | r[T] | SSU |
| StrainInfo | Cultured | s[C] | SSU, LSU |
| StrainInfo | Typestrains | s[T] | SSU, LSU |
The identifiers can be used for data retrieval by searching in the strain field see FAQ.
The information if a sequence originates from a genome project has been taken from NBCI and added to the field strain. It is indicated by n[G] and w[G] for genomic entries reported in the WGS section of INSDC.
The length and colours of the bars give a first indication on the sequence and alignment quality. After downloading the sequences as an ARB file, sequences that need attention can be selected by searching for low quality (alignment, sequence) in the corresponding ARB database fields. A full description of the colour code and all database fields available in the ARB files can be found in the FAQ section. Taking into account the rich set of sequence associated information that comes along with every SILVA sequence, user designed sub-databases can be easily generated.
All rRNA sequences have been aligned based on a completely manually re-checked SEED alignment of 67,342 rRNA sequences for SSU and 4508 rRNA sequences for LSU. The SSU alignment is based on the official ssu_jan04 release of the ARB Project. The SSU SEED alignment has been considerably improved for Archaea by manual addition of more than 1,000 sequences as well as Fungi (10,000 sequences). All SSU Eukaryotic sequences (18S) have been cross-checked by Wolfgang Ludwig before their addition to the SEED. Most of the bacterial sequences have also undergone a curation process carried out by the SILVA Team. We would rate our SSU SEED alignment for all Bacteria and Archaea as good and for Eukarya as reasonable.
RNAmmer is a computational predictor for the major rRNA species (SSU, LSU) from all three domains of life. The program uses hidden Markov models trained on data from the European ribosomal RNA database project. SILVA runs the profiles of RNAmmer on all sequence entries of the EMBL-EBI/ENA archive to complement the existing predictions. All predictions are marked with RNAmmer in the ann_src_field. More information about RNAmmer can be found in the paper.
Chuvochina M, Gerken J, Frentrup M, Sandikci Y, Goldmann R, Freese HM, Göker M, Sikorski J, Yarza P, Quast C, Peplies J, Glöckner FO, Reimer LC (2026) SILVA in 2026: a global core biodata resource for rRNA within the DSMZ digital diversity. Nucleic Acids Research, Volume 54, Issue D1, 6 January 2026, Pages D334–D341, https://doi.org/10.1093/nar/gkaf1247
If you use SINA please cite:
Pruesse, E, Peplies, J and Glöckner, FO (2012) SINA: accurate high-throughput multiple sequence alignment of ribosomal RNA genes. Bioinformatics, 28, 1823-1829