Data API
Query registered research datasets and download machine-readable records. No account or API key is required for these read-only endpoints.
OpenAPI 3.1 (JSON)Resource overview (JSON)Study resources (JSON)Terms of use
Discover, then query
The overview lists data modules, registered studies, genome assemblies and download routes. Use identifiers returned by their catalogs; gene names and assembly versions are not interchangeable.
BASE="https://sci.hainanu.edu.cn/strawberry/api/v1"
curl --fail "$BASE/catalog" -o catalog.json
curl --fail "$BASE/openapi.json" -o openapi.json
curl --fail "$BASE/genes/status"
curl --fail --get "$BASE/genes/search" \
--data-urlencode "geneId=YOUR_GENE_ID" \
--data-urlencode "species=SPECIES_FROM_CATALOG" \
--data-urlencode "version=VERSION_FROM_CATALOG"
The OpenAPI document describes the public data endpoints, request parameters and response structures. Its relative server URL resolves against the document location, including the portal’s deployment prefix. Returned links without a scheme resolve against the portal root (/strawberry/), not /api/v1/.
Download data
GWAS results and original sources
GET /resources?module=gwas lists every imported study/analysis with JSON and TSV exports plus original repository or supplementary-material links. JSON includes dataset and trait metadata. TSV retains all association fields, source P-values, missing values and annotation provenance; blank cells mean missing, not zero.
curl --fail "$BASE/resources?module=gwas" -o gwas-resources.json
curl --fail "$BASE/downloads/gwas/uf-breeding/yield.tsv" -o yield.tsv
curl --fail "$BASE/downloads/gwas/uf-breeding/yield.json" -o yield.json
The original Florida population uses uf-breeding, with yield and fruit-size. Other study and analysis keys come from /gwas/studies. An external_source link is a provider’s landing page, not a promise that SDH hosts its raw files. Source-specific licensing and access terms still apply.
Data associated with a publication
curl --fail --get "$BASE/publications/data" \
--data-urlencode "doi=10.1093/hr/uhad271"
curl --fail --get "$BASE/downloads/publication" \
--data-urlencode "doi=10.1093/hr/uhad271" \
--data-urlencode "format=tsv" -o publication-files.tsv
curl --fail --get "$BASE/downloads/publication.zip" \
--data-urlencode "doi=10.1093/hr/uhad271" -o publication-data.zip
DOI matching is exact after normalization. The ZIP contains available GWAS exports and transcriptome sample metadata, plus a manifest linking other registered resources and original sources. It does not fetch external files or claim to contain every supplement from the article. If the bundle exceeds 32 MiB of uncompressed export data, download the individual resources instead.
Batch CDS and protein sequences
Choose species, assembly and example identifiers from /batch-sequences/status. Send geneIds as a string of whitespace-, comma- or semicolon-separated identifiers. The limit is 100 unique IDs and 2,000,000 sequence symbols per request.
curl --fail "$BASE/batch-sequences/status" -o sequence-catalog.json
# Requires jq; uses an example returned by this deployment.
jq -e '.available == true' sequence-catalog.json
jq '.example | {species, version, type: "both",
geneIds: (.geneIds | join("\n"))}' sequence-catalog.json > request.json
curl --fail -X POST "$BASE/downloads/sequences?format=zip" \
-H 'Content-Type: application/json' --data-binary @request.json \
-o sequences.zip
ZIP includes sequences.fasta and report.json with one status per requested ID. Missing sequences are never substituted from another assembly. format=fasta is strict: incomplete requests return HTTP 409 with a report; no sequences returns HTTP 404.
Download links in Python
from urllib.parse import urljoin
import requests
portal = "https://sci.hainanu.edu.cn/strawberry/"
response = requests.get(urljoin(portal, "api/v1/resources"),
params={"module": "gwas"}, timeout=60)
response.raise_for_status()
for resource in response.json()["resources"]:
for file in resource["downloads"]:
print(resource["title"], file["format"],
urljoin(portal, file["url"]))
Transcriptome sample metadata
/transcriptome/status lists datasets. Use a dataset slug to retrieve its complete sample/condition table without selecting a gene:
curl --fail --get "$BASE/transcriptome/samples.tsv" \
--data-urlencode "dataset=fragaria-vesca-strawberry-atlas-2026" \
-o samples.tsv
For JSON use /transcriptome/samples with the same parameter. Each row includes matrix-column index, source sample code, tissue/condition, within-group order and column role. TSV also carries species, accession, exact source assembly, expression unit and publication/data sources.
condition_series rows represent condition summaries, not independent biological replicates. withinGroupOrder is display order, not a validated replicate number. Stage labels are parsed from source column codes. Missing run accessions or replicate assignments are not invented.
Conventions and limits
- All documented operations are read-only. POST is used for structured batch queries, not data modification. Administrative, conversational and provider-configuration routes are not part of this contract.
- Anonymous cross-origin requests are allowed on documented routes, without credentials. GET, HEAD, OPTIONS and the listed POST operations are supported. No token should be sent.
- Each client address is limited to 120 requests/minute and 12 export requests/minute. Four exports can run concurrently per application instance. HTTP 429 includes
Retry-After: 60; wait before retrying. Limits are per instance and may be supplemented by deployment-level controls. - POST bodies are limited to 64 KiB; generated downloads to 32 MiB. Transcriptome comparison accepts at most 50 genes. Study catalogs are cached for up to 60 seconds. Do not treat a cached listing as proof that an external repository is currently reachable.
- Pagination is one-based where provided. Parameter names differ between existing modules (
pageSizeversussize); use the endpoint reference. Some catalogs are complete arrays rather than paginated lists. - Coordinates, reference assemblies and statistical fields retain their dataset-specific meanings. Consult module/source metadata; there is no global coordinate conversion. Overlap and nearest-gene annotations do not establish causality.
- HTTP 400: invalid input; 404: no matching resource; 409: ambiguous or incomplete request; 413: size limit; 429: rate/concurrency limit; 503: collection unavailable. Check availability fields and per-collection status even on HTTP 200.
- TSV text fields use quoted-field escaping. To preserve source values they are not prefixed or rewritten; import identifier/text columns as text in spreadsheet applications, not as formulas.
- Gene, sequence and accession identifiers are case- and assembly-sensitive. Use explicit species/version when an identifier occurs in more than one assembly. Cite the source studies and record dataset/release versions in reproducible analyses.
API access does not grant new rights to third-party data. See License & Terms of Use and the original provider’s terms. This description follows the OpenAPI 3.1.1 specification.
Endpoint reference
Routes below are read from this deployment’s OpenAPI description. Paths are relative to /api/v1. Expand an endpoint for its parameters and request/response schema.
Loading endpoint reference…