METHOD & INTERPRETATION
What the evidence means
Hairpin geometry and folding energy identify thermodynamically plausible local stem-loops. Many genomic sequences can form hairpins, so these signals are necessary but not sufficient.
The supervised score combines sequence and structure features learned from curated human precursors, candidate-like genomic negatives, and structurally plausible Rfam non-miRNA decoys. Logistic regression remains the selected model because it achieved the best validation PR-AUC; Random Forest, Extra Trees, and histogram gradient boosting are retained in the benchmark. It is a triage score, not proof or a calibrated probability of biological function.
Curated-reference similarity is displayed separately. A close match supports “known-like”; a distant match does not prove novelty.
Cross-species similarity compares the folded candidate with separate mouse, orangutan, chimpanzee, gorilla, and limited bottlenose-dolphin precursor panels. It can support family-level sequence similarity but is not a locus-level phylogenetic-conservation calculation and does not alter the classifier score.
Protein-region preprocessing is coordinate-aware rather than sequence-guessed. With a length-matched hg38 FASTA interval, the default masks only annotated protein-coding CDS. An optional aggressive mode masks every exon of a protein-coding transcript, but this can remove genuine miRNAs in untranslated exonic regions. Without trustworthy coordinates, no gene/exon claim is made.
Rfam conflict evidence asks whether the sequence closely resembles another structured RNA family. A close match is a warning, while a distant match cannot exclude every non-miRNA RNA class.
Protein-machinery evidence distinguishes canonical roles from candidate-specific experiments. DROSHA, DGCR8, XPO5, DICER1, TARBP2, AGO2, and TNRC6A are not inferred binding partners merely because a candidate forms a hairpin. A “supported interaction” requires RIP or biochemical binding/processing evidence for the matched external control. “Tested—not required” describes pathway dependency and does not prove that physical contact is impossible.
Literature cards are short, manually curated summaries shown only after a strong human reference match. They report evidence for the known miRNA in the cited experimental context; they are not evidence about expression or function of the submitted sample.
Missing evidence includes transcription, tissue expression, precise mature/star processing, locus-level evolutionary conservation, RISC loading, mRNA targeting, pathway effects, and disease causality.
The intended decision is: which few candidate loci should a researcher investigate experimentally first?
A reproducible local analysis
The browser uses ViennaRNA 2.7.2, scanning 60, 70, 90, and 110 nt windows. The native Python application remains the scientific reference and also supports RNALfold. The exported classifier preserves the trained scaling parameters, coefficients, and validation threshold.
Engine: ViennaRNA-2.7.2-RNAfold-windows-v1
Data release: v1-0cf79300764342d9