User Manual

Find genes, inspect datasets and run analyses in Strawberry Data Hub.

FAQ & contact

Getting started

If you have a gene ID, start with Gene Search. If you have a DNA or protein sequence, use BLAST to find matching records. To browse a cultivar or wild species, start from the genome directory on the homepage.

Choose the reference first

Gene identifiers and coordinates belong to a particular genome and annotation version. Select the same reference in each tool. A symbol such as MYB10 is not a substitute for an exact gene ID, and changing a version suffix does not convert an identifier to another assembly.

Examples for gene and expression searches
ReferenceExample identifier
F. vesca Hawaii-4 T2T v6.0FvesChr6G00063740.1
F. × ananassa Camarosa v1.0.a2FxaC_2g30690

Keep transcript suffixes, punctuation and capitalization as supplied by the source. For several genes, choose the tool that fits the task: the transcriptome heatmap accepts up to 50 IDs, Batch Sequence Retriever accepts up to 100, and enrichment accepts up to 2,000 unique IDs. Co-expression Network currently starts from one gene.

Genomes & germplasm

Genome directory and species profiles

The homepage separates cultivated strawberries, wild species, pathogens and pangenome resources. Search by cultivar, species or disease; use View all to expand a collection. Open a record to check its assembly, publication and available analysis links.

On species profiles, distinguish the primary assembly from phased haplotypes and later annotation releases. Cultivar records are separate entries in the Species menu. A T2T label describes an assembly; it does not mean that every haplotype or analysis is available.

Germplasm Resources

Open Breeding → Germplasm Resources. Search names, accessions, origins or traits, then narrow wild resources by ploidy and breeding trait. The Dataset selector also provides genome-associated materials, the public cultivar genomic panel and USDA core accessions. Follow the repository link to check availability or request material; SDH provides the records, not a distribution service.

Breeding genes and phylogeny

Open Breeding → Breeding Genes for the homepage's four trait groups, including individual genes and regulatory networks with their primary studies. The Breeding Genes page uses the same entries as the homepage. Use an internal gene link when an assembly-specific ID is supplied; otherwise consult the paper before assigning a reference gene.

The separate Species Tree presents a literature-based phylogenetic hypothesis. Read its source and branch-length notes when using it; the seven-proteome tree in Pangenome is a different analysis.

Genes & annotation

Gene Search

  1. Select the assembly, enter one exact gene ID and choose Search.
  2. Check the species and annotation version in the result. If more than one reference matches, select the intended record.
  3. Copy the CDS or protein sequence and inspect the available domains, GO, KEGG and other annotation. Follow the expression, epigenomics or genome-browser links as needed.

An empty annotation field means that no entry is recorded for that gene in the selected annotation source.

JBrowse and BLAST

JBrowse lists nuclear, chloroplast and pathogen references. Open an assembly and search for a gene or chromosome interval inside the browser; turn sequence, annotation and other supplied tracks on or off there.

BLAST opens SequenceServer. Paste or upload FASTA, select the appropriate genome, CDS or protein database, and run the compatible search program. Check identity, alignment coverage and E value before choosing a hit. Use the hit's reference assembly for subsequent coordinate searches.

Gene families, transcription factors and miRNA

Gene Family Search and Transcription Factors have separate catalogues. Choose a genome and family, then run the search; leaving both at All shows the release overview. Review domain evidence, Arabidopsis matches and any assembly-level quality warning. Click linked gene IDs to open their matched SDH records. These family assignments are computational classifications.

For miRNA, select an assembly and optionally an Rfam family, then choose Show loci. The table gives family accessions and genomic locations. Rfam links provide the family description; a predicted locus alone does not establish expression or a target gene.

Synteny & pangenome

Synteny Search

Select the input assembly, enter a gene ID and choose the target assembly. Set Flanking anchors to inspect only the matched pair or its surrounding collinear genes, then choose Search. Use the assembly-specific example if you need a starting point. The analysed assembly list is smaller than the full genome catalogue.

Pangenome

The workspace has three views:

  • SV & TE explorer: enter an Fvb1 interval or graph-site ID, or select a bin in the evidence-density plot. Filter the site table, open a site and inspect gene, CDS and TE overlaps in the reference and carrier intervals. Use the path-alignment viewer for reference windows of 5–100 kb.
  • Gene content: search a pan-gene, orthogroup or protein ID. Filter by occupancy or proteome, open group membership, and compare up to 50 selected groups. Download the current table page or selected profiles as TSV.
  • Assemblies & phylogeny: check which proteomes and genomes enter each analysis, then inspect the seven-proteome tree and its analysis information.

Gene-content analysis uses seven diploid proteomes; the sequence graph uses six genomes. Site evidence is currently limited to Fvb1, although path alignment covers Fvb1–Fvb7. A missing protein-group assignment is not proof of biological gene absence, and graph-site candidates are not all validated structural variants.

Expression & other omics

Transcriptomics

Open cultivated or wild strawberry from Omics → Transcriptomics, then select a study. The dataset description identifies the material, reference, sample groups and expression unit. Cultivated matrices use Camarosa v1.0.a2 identifiers; the wild matrices use Hawaii-4 T2T v6.0 identifiers. The sampled cultivar can differ from the reference cultivar.

  1. For Single-gene profile, enter an ID and choose Visualize gene. Use the supplied examples to inspect a known profile.
  2. For Multi-gene heatmap, enter up to 50 IDs, one per line. Choose mean or median TPM and either log expression or row Z-score, then choose Draw heatmap. Row Z-scores compare each gene's relative pattern, not absolute abundance between genes.
  3. For the cultivated eFP atlas, enter a Camarosa ID and choose Draw eFP. Hover over or focus a tissue to read its TPM; Download values saves the underlying values.

Compare tissues or treatments within the same study. The two eFP project panels have independent colour scales, and some matrices contain condition summaries rather than individual libraries. Read Dataset and methods before combining results.

Single-cell atlas

The leaf atlas loads automatically. Select a sample and cell type, choose colouring by cell type, sample or cluster, and switch between 2D and 3D UMAP. Enter a gene ID to inspect its expression. Use Reset to restore the view, Copy view link to share it and Download plot to save the figure.

Below the map, switch between Markers, Differential expression, GO terms and Defense genes; Export table saves the displayed evidence. The CRA004848 study uses F. vesca v4.0.a1 IDs. Its published results and the SDH reanalysis are separate: the regenerated map has 50,702 cells, not the paper's reported 50,327.

Epigenomics

Choose Histone modification or Chromatin accessibility from the Omics menu. Select a reference and search by gene ID or genomic interval, optionally adding 1–3 kb of flanking sequence or restricting the dataset. Open a result's dataset details to check the assay, source and type of gene or region association. Follow the related Gene record or JBrowse link to inspect the same locus.

A project record does not guarantee that all signal tracks have been processed. No matching evidence means no indexed record for that query, not an absence of chromatin activity.

Metabolomics

Use the cultivar, tissue, treatment and platform filters to find a study, then choose an assay. Inspect the abundance heatmap, PCA and metabolite distribution. In the published-statistics table, select a factor and read the reported contrasts, P values, effects, VIP scores and directions where supplied.

Search metabolites by name, identifier or m/z with a tolerance. Open a feature to view its identification and source details. Keep positive- and negative-ion assays separate; matching names or masses do not by themselves establish the same compound.

Co-expression networks

A sample network loads when Co-expression Network opens. Choose Reference genome: Hawaii-4 T2T v6.0 uses the wild tissue atlas, and Camarosa v1.0.a2 uses the fruit-ripening atlas. The expression source and library count appear below the form.

  1. Enter one exact gene ID. Choose correlation across all libraries or within tissue/condition, positive or negative relationships (or both), the minimum |r| and 10, 20 or 50 partners.
  2. Choose Find partners. The within-group view subtracts tissue or condition means before calculating correlation; it helps check whether an overall pattern is mainly explained by group differences.
  3. Drag genes to adjust the network, hover to trace connections, or select a node to compare expression. Partner links controls links among the displayed partners. Use +, − and Fit to adjust the view, and Pause, Resume or Replay to control movement.
  4. Save the figure as SVG or a 3,300-pixel-wide PNG. Export freezes the current layout. Download edges saves CSV; Download analysis saves the query, correlations and dataset information as JSON.

Correlations are calculated from log2(TPM + 1). Solid and dashed lines mark positive and negative correlations; node position is only a layout choice. The page does not accept multi-gene uploads, assign WGCNA modules or demonstrate regulatory causation. Genes removed by expression/variance filtering may still be present in Gene Search.

GWAS

In GWAS, select a study/population and phenotype or analysis. In addition to the Florida yield and fruit-size results, the catalogue includes California fruit-quality hybrids, the PG1 firmness panel, Phytophthora crown rot and angular leaf spot. The reference and statistical threshold belong to the selected study.

Select a Manhattan-plot point or table row to inspect the marker and available gene link. Search marker IDs, sort the association table or use Significant only when a numerical threshold is supplied. Download saves the filtered association results.

Genotype-group plots appear only where matched genotype and phenotype data are available. They are not available for the crown-rot panel. A study without a supplied numerical significance cutoff is not assigned one by SDH. A nearby gene is a candidate associated with the locus, not a validated causal gene.

GO & KEGG enrichment

Open GO Enrichment or KEGG Enrichment and select the annotation version matching your IDs. Paste a list or upload a plain-text file smaller than 1 MB. Submit up to 2,000 unique genes with Run enrichment.

First check how many IDs were found and how many had the relevant annotation. Review Excluded gene IDs before interpreting the results. The background is the selected assembly's GO- or pathway-annotated gene set; the form does not provide a custom background upload.

Filter GO results by BP, MF or CC, search terms, and set the adjusted-P cutoff. The tables report counts, ratios, P values and contributing genes. Download the result table or the dot and bar plots. KEGG results refer to assigned reference pathways, not a strawberry-specific experimental pathway map.

Sequence tools

Batch Sequence Retriever

Select one species/assembly, enter up to 100 IDs and choose CDS, protein or both. Click Retrieve sequences. The tool retrieves exact records from the selected reference rather than converting IDs between assemblies.

Check each row's status, sequence length and checksum. Copy or download individual records, or save Combined FASTA and Results TSV. A request is limited to 2,000,000 sequence characters in total. Missing IDs remain listed instead of being silently replaced.

If the page reports that the sequence database is unavailable, retry later. A displayed example is not confirmation that custom IDs can be retrieved while the database is offline.

Sequence Toolkit

Paste a sequence, choose DNA, RNA, protein or automatic detection, and set a translation frame if needed. Choose whether to remove gaps, then select Analyze sequence. Input can contain up to 100,000 characters. Inspect composition, reverse complement and standard-code translations; use the FASTA and statistics downloads for the returned results.

ORF Finder

Paste DNA up to 50 kb or load an SDH example. Select the genetic code (standard or plant plastid), start-codon rule, minimum length, strand and nested-ORF option, then choose Find ORFs. Inspect coordinates and translations before downloading protein FASTA, CDS FASTA or TSV. An ORF is a candidate coding region, not a confirmed gene model.

Primer design

Primer Design & Specificity accepts a pasted DNA template up to 50 kb, or an SDH gene CDS with its exact gene ID, species and assembly. Choose qPCR or conventional PCR, set the product-size range and number of pairs, then select Design primers. Advanced settings control Tm, primer length and GC content.

Review forward/reverse sequences, template positions, product size and Primer3 penalty. Download TSV or FASTA. For genomic PCR spanning introns, use the genomic template rather than a spliced CDS.

Specificity checking: Primer3 design is available, but the ipcress service for custom primer pairs is not enabled on the current deployment. Use verified pair displays a precomputed example only. It does not test your newly designed primers. Check both primers against the intended reference before ordering them.

CRISPR guide design

The SDH CRISPR tool uses FlashFry to find 20-nt SpCas9 spacers with NGG PAMs. It searches the selected genome for NGG/NAG matches on both strands, allowing up to four spacer mismatches without bulges. The indexed references are listed in the form.

  1. Choose Species and the exact Reference assembly. Use Search assemblies to narrow the list. Only indexed references are selectable.
  2. Choose the input type. For Sequence, paste genomic DNA, upload one FASTA file up to 12 KB, or load the assembly example. Use 23–2,000 bases (A/C/G/T/N), not spliced CDS or RNA. Windows containing N are skipped.
  3. For Gene ID, enter an assembly-matched gene or transcript ID, choose Locate gene, then Load region. Shorten loci longer than 2,000 bp first. For Genomic region, enter the exact chromosome/scaffold and 1-based inclusive start/end. Retrieved DNA is continuous forward-reference sequence, including introns; editing it removes its coordinate association.
  4. Choose the maximum mismatch count and select Find guides. The activity bar and timer show the job state. If the connection fails, use Check job again.
  5. Filter by GC, strand or maximum complete match count, and sort the candidates. Select a guide ID for genomic locations and annotation overlaps. Copy copies the 20-nt spacer without its PAM.
  6. Check individual guides or use Select filtered guides, then download Selected TSV or Selected JSON. Selections persist across pages and filters. Guides TSV and Results JSON export the full result.

Download before the displayed expiry time. Completed results remain available for 20 minutes and are lost on server restart. Refreshing the same browser tab restores its saved job if still available; Reset forgets it. Gene lookup is disabled where usable annotation is unavailable. Match annotations describe interval overlap across annotated isoforms, not a transcript-specific exon number.

0 MM counts every perfect spacer match, including the intended site when present. No site is automatically subtracted. A truncated search reports lower-bound counts, and at most 100 locations are returned per guide. These are genomic-match results, not editing-efficiency scores; check homoeologs and reference coverage before experimental use.

Input and reference intervals are distinct; both are 1-based inclusive and include spacer plus PAM (23 bp). Missing annotation and annotation conflicts are reported rather than silently resolved. Save the reproducibility record with the result for the reference and annotation checksums.

External workspaces

The Tools directory and Breeding menu include services maintained separately from SDH. Use the workspace's demo to check the required file layout. If an embedded window is blank or too narrow, choose Open in new window. Uploaded files go to the named service.

Inputs and uses
WorkspaceWhat to prepare
GuideRNATarget DNA and the matching genome/GTF required by HBICloud. Choose the PAM and inspect candidate guides. This is separate from SDH's indexed FlashFry tool.
MICEA table with missing values. Set the chained-equation iterations and random state in HBICloud.
KNNA numeric table with missing values. Choose the number of neighbours; review scaling before imputation.
iClusterOmics matrices with matching sample IDs and ordering. Set the distribution for each layer in HBICloud.
SCCGs Prediction ModelOne or more protein sequences in FASTA. The SCCP workspace predicts sulfur-containing-compound gene candidates.
Flowering Gene Prediction ModelProtein FASTA for the FTGDB2 classification models. Use the results to select candidates for further annotation or testing.

Imputed values are estimates, not additional measurements. Sequence classifications and guide predictions need independent validation. Check the external provider's data policy before submitting unpublished or restricted material.

Literature & Strawberry AI

The Knowledge navigation menu provides access to Strawberry AI and Knowledge Base.

Knowledge Base

Search by title, author, journal, DOI or topic, then open a paper's title to follow its source link. The catalogue lists the literature available to Strawberry AI. For example, search quasi-circadian to locate the noncanonical regulatory-network study.

Strawberry AI

Use Chat for literature questions, Agent for SDH data queries, or Auto to let the system choose. Questions can be in English or Chinese. Include exact gene IDs, species and reference versions for gene queries; several IDs can be supplied for supported batch tasks.

Open Sources to inspect citations, tool results and notes. Follow the original records before using an interpretation in a manuscript. Where sequence-download actions are returned, save the supplied sequence records rather than copying a sequence from generated prose.

The Trial control shows remaining access. To use your own provider account, choose Use my API key, select the provider, enter the model ID and key, then select Use this API. Provider charges and access rules apply. Use Clear key when finished on a shared computer.

Other community pages

Conferences lists selected upcoming strawberry and plant-genomics meetings with official sources and a verification date. Earlier records are kept separately in the historical archive.

Related databases lists external resources, and SDH news records portal updates. The homepage visitor display reports country-level traffic; it is not part of the research data.

Downloads

The Study data section provides GWAS result exports, transcriptome sample tables and original data repository links. Use Help → Data API for the OpenAPI description, resource overview, batch sequence downloads and publication-linked data.

On Downloads, search by species, assembly version or filename and narrow the collection as needed. Check the displayed reference, file type and size before downloading. Nuclear, chloroplast and pathogen resources have separate records; phased components should not be combined simply because they share a cultivar name.

The genome catalogue lists genome FASTA, GFF/GFF3 annotation, CDS and protein sequences where available. Use module-specific exports for network edges, expression values or analysis tables; they are not all files in the genome-download catalogue.

Troubleshooting & citation

A gene is not found

Check the exact identifier and reference version, including transcript suffixes. Try the selected tool's example to distinguish an input problem from a service problem. A gene can exist in the genome catalogue but be absent from an expression-filtered matrix or an annotation-specific background.

A result or workspace does not load

Use the page's retry control, or reload once. Open external workspaces in a separate window if embedding fails. For a report to the database team, include the page URL, reference, input example, error text and time of the request. Do not include API keys.

Citing and reusing data

Cite the original assembly, dataset and analysis sources, and record the reference version, study accession, settings and access date. See the FAQ for the database citation and contact details, and License & Terms of Use for reuse conditions. Source data and external tools retain their own terms.