Data API

Query registered research datasets and download machine-readable records. No account or API key is required for these read-only endpoints.

Discover, then query

The overview lists data modules, registered studies, genome assemblies and download routes. Use identifiers returned by their catalogs; gene names and assembly versions are not interchangeable.

BASE="https://sci.hainanu.edu.cn/strawberry/api/v1"
curl --fail "$BASE/catalog" -o catalog.json
curl --fail "$BASE/openapi.json" -o openapi.json
curl --fail "$BASE/genes/status"
curl --fail --get "$BASE/genes/search" \
  --data-urlencode "geneId=YOUR_GENE_ID" \
  --data-urlencode "species=SPECIES_FROM_CATALOG" \
  --data-urlencode "version=VERSION_FROM_CATALOG"

The OpenAPI document describes the public data endpoints, request parameters and response structures. Its relative server URL resolves against the document location, including the portal’s deployment prefix. Returned links without a scheme resolve against the portal root (/strawberry/), not /api/v1/.

Download data

GWAS results and original sources

GET /resources?module=gwas lists every imported study/analysis with JSON and TSV exports plus original repository or supplementary-material links. JSON includes dataset and trait metadata. TSV retains all association fields, source P-values, missing values and annotation provenance; blank cells mean missing, not zero.

curl --fail "$BASE/resources?module=gwas" -o gwas-resources.json
curl --fail "$BASE/downloads/gwas/uf-breeding/yield.tsv" -o yield.tsv
curl --fail "$BASE/downloads/gwas/uf-breeding/yield.json" -o yield.json

The original Florida population uses uf-breeding, with yield and fruit-size. Other study and analysis keys come from /gwas/studies. An external_source link is a provider’s landing page, not a promise that SDH hosts its raw files. Source-specific licensing and access terms still apply.

Data associated with a publication

curl --fail --get "$BASE/publications/data" \
  --data-urlencode "doi=10.1093/hr/uhad271"
curl --fail --get "$BASE/downloads/publication" \
  --data-urlencode "doi=10.1093/hr/uhad271" \
  --data-urlencode "format=tsv" -o publication-files.tsv
curl --fail --get "$BASE/downloads/publication.zip" \
  --data-urlencode "doi=10.1093/hr/uhad271" -o publication-data.zip

DOI matching is exact after normalization. The ZIP contains available GWAS exports and transcriptome sample metadata, plus a manifest linking other registered resources and original sources. It does not fetch external files or claim to contain every supplement from the article. If the bundle exceeds 32 MiB of uncompressed export data, download the individual resources instead.

Batch CDS and protein sequences

Choose species, assembly and example identifiers from /batch-sequences/status. Send geneIds as a string of whitespace-, comma- or semicolon-separated identifiers. The limit is 100 unique IDs and 2,000,000 sequence symbols per request.

curl --fail "$BASE/batch-sequences/status" -o sequence-catalog.json
# Requires jq; uses an example returned by this deployment.
jq -e '.available == true' sequence-catalog.json
jq '.example | {species, version, type: "both",
    geneIds: (.geneIds | join("\n"))}' sequence-catalog.json > request.json
curl --fail -X POST "$BASE/downloads/sequences?format=zip" \
  -H 'Content-Type: application/json' --data-binary @request.json \
  -o sequences.zip

ZIP includes sequences.fasta and report.json with one status per requested ID. Missing sequences are never substituted from another assembly. format=fasta is strict: incomplete requests return HTTP 409 with a report; no sequences returns HTTP 404.

Download links in Python

from urllib.parse import urljoin
import requests

portal = "https://sci.hainanu.edu.cn/strawberry/"
response = requests.get(urljoin(portal, "api/v1/resources"),
                        params={"module": "gwas"}, timeout=60)
response.raise_for_status()
for resource in response.json()["resources"]:
    for file in resource["downloads"]:
        print(resource["title"], file["format"],
              urljoin(portal, file["url"]))

Transcriptome sample metadata

/transcriptome/status lists datasets. Use a dataset slug to retrieve its complete sample/condition table without selecting a gene:

curl --fail --get "$BASE/transcriptome/samples.tsv" \
  --data-urlencode "dataset=fragaria-vesca-strawberry-atlas-2026" \
  -o samples.tsv

For JSON use /transcriptome/samples with the same parameter. Each row includes matrix-column index, source sample code, tissue/condition, within-group order and column role. TSV also carries species, accession, exact source assembly, expression unit and publication/data sources.

condition_series rows represent condition summaries, not independent biological replicates. withinGroupOrder is display order, not a validated replicate number. Stage labels are parsed from source column codes. Missing run accessions or replicate assignments are not invented.

Conventions and limits

API access does not grant new rights to third-party data. See License & Terms of Use and the original provider’s terms. This description follows the OpenAPI 3.1.1 specification.

Endpoint reference

Routes below are read from this deployment’s OpenAPI description. Paths are relative to /api/v1. Expand an endpoint for its parameters and request/response schema.

Loading endpoint reference…