Methods

Scope, data structure, normalization, and interpretation guidance for the saRNA Database.

Scope

The saRNA Database is a searchable repository of reported or tested sequences and the experimental conditions and effects associated with individual observations. It does not classify every sequence as a candidate saRNA, control, mechanistic reagent, or other functional role.

Record structure

One database record represents one reported observation or experimental context. The same nucleotide sequence may appear in multiple records when it was tested under different conditions, in different models, or in different publications.

Bibliographic metadata are stored once in the Sources table. Each record links to one source using a stable publication identifier. The Sequences table groups records that share the same normalized sense and antisense sequence strings.

Sequence normalization

Normalized sequences are trimmed and converted to uppercase. RNA U and DNA T are preserved as reported and are not interconverted. A strand for which no sequence was reported is represented as no_sequence_reported.

Sequence identity is based only on the normalized sense and antisense nucleotide strings. Chemical modifications, conjugates, delivery systems, dose, model, publication, and experimental outcome do not create separate sequence identities.

Sequence length and GC percentage are derived from normalized sequences. Missing sequences have blank derived length and GC values rather than zero.

Missing and special values

  • not_reported means that the information was unavailable or not stated. It does not mean zero or an experimentally absent result.
  • not_applicable means that the field does not logically apply to the record.
  • no_alignment_to_reference_genome means that a sequence was reported but no usable alignment to the stated reference genome is recorded.
  • not_measured means that the relevant outcome was explicitly not measured.
  • no_sequence_reported means that no nucleotide sequence was available for that strand.

Interpretation

Activity and experimental-condition fields retain heterogeneous source-reported measurements. Reported effects may be numeric values, ranges, categories, thresholds, arbitrary units, or other source-specific representations. They have not been converted to a common scoring scale and should not be compared directly without confirming compatible assays and definitions.

Source-reported values are preserved wherever practical. When a value appeared unusual but was explicitly stated by the source, it was generally retained with explanatory context rather than silently corrected.

Identifiers and corrections

Records, sequence identities, and sources use stable identifiers. Retired sequence identifiers are not reused. Corrections and additions will be documented in later versioned releases rather than silently replacing the archived release.