SCLib API Reference
Programmatic access to the JZIS Superconductivity Library — hybrid search, RAG Q&A, materials database, paper metadata, and more.
Quick start
1. Get your API Key — go to Dashboard → API Keys and click + New key. Copy the scl_… value.
2. Pass it in the header — every request that needs authentication should include:
# cURL curl https://api.jzis.org/sclib/v1/materials \ -H "X-API-Key: scl_YOUR_KEY"
# Python import requests API = "https://api.jzis.org/sclib/v1" headers = {"X-API-Key": "scl_YOUR_KEY"} resp = requests.get(f"{API}/materials", headers=headers) print(resp.json())
Authentication & quotas
| Identity | Auth method | Daily quota |
|---|---|---|
| Guest (no key) | None — rate-limited by IP | 3 |
| Registered user | X-API-Key: scl_… | 999 |
| Browser session | Secure HttpOnly session cookie | 999 |
API Key and JWT share the same daily quota per user. Quotas reset at 00:00 UTC. When the quota is exceeded the API returns 429 Too Many Requests.
Password resets and “revoke all sessions” invalidate browser and bearer JWT sessions. API keys remain separately revocable from the Keys dashboard.
Endpoints
Base URL: https://api.jzis.org/sclib/v1
/searchConsumes quotaOrdinary topic queries use hybrid search over retained corpus inputs, combining Vertex semantic retrieval with PostgreSQL full-text search and deterministic reranking. Scientific property or evidence conditions use the separate structured lookup described below.
POST /v1/search
Content-Type: application/json
{
"query": "iron-based superconductor pairing symmetry",
"top_k": 20,
"filters": {
"year_min": 2020,
"material_family": ["iron_based"],
"exclude_retracted": true
}
}Topic-route response: total, results[] (paper_id, title, authors, year, matched_chunk, relevance_score, material_family), query_time_ms.
Search and Ask also return scientific_query, scientific_lookup, and scientific_results[]. The bounded interpretation preserves the original query, formula notation and condition spans; unresolved scientific clauses require clarification rather than silently dropping a condition. Lookup status is not_requested, completed, unavailable, or clarification_required.
Scientific filters, including material_family in the example above, select the structured route: paper results=[] and total=0 are intentional. Read scientific_lookup.returned_count and scientific_results[] instead. Up to 20 extraction rows retain exact parent, evidence, content and generation bindings, with has_more for additional eligible rows. All requested conditions must match the same extracted record. Missing pressure is not ambient; a non-detection is not Tc equal to zero. These machine extractions are not original quotations, scientific acceptance or ML labels.
retrieval_generation declares generation_snapshot with a generation ID, activation-event ID and manifest hash, or legacy_lexical_only. Ordinary topic retrieval can use legacy lexical search without an active generation. Structured lookup requires bound derived parents in an active generation; unavailable or changed selected inputs withhold the entire structured result set, without legacy numerical fallback. An empty set is not evidence of absent superconductivity, full-corpus coverage or currentness beyond the checked snapshot.
/askConsumes quotaRoute-specific scientific retrieval and RAG question answering: structured-only lookup, separate numerical and original-source candidates, or a bounded cited answer when generation is requested. Not every question triggers model generation.
POST /v1/ask
Content-Type: application/json
{
"question": "What is the highest Tc in nickelate superconductors?",
"max_sources": 8,
"language": "auto"
}Response: answer (route-specific static notice or Markdown with [1][2] citations), sources[] (paper_id, title, year), citation_valid, citation_warnings, tokens_used, and query_time_ms.
The numerical example above is not a promise of an AI-generated maximum or a scientifically accepted record. Non-comparison structured-only Ask uses the same qualified extraction fields as Search, without embedding or generation calls. Unsupported conditions can require clarification. Read the interpretation and lookup status before consuming any numerical rows.
Mixed questions and comparisons with typed property or evidence requests return the closed scientific_mixed envelope under scientific-mixed-evidence/1.1.0. It separates scientific_results[] from original-passage sources[]. Every exact numerical-parent/original-citation pair appears once in the complete association matrix, retaining the parent revision, declared catalogue-snapshot hash, original source index, vector ID, evidence revision, evidence-record hash and full-content hash. Original citation indices are not extraction-row indices.
scientific_mixed.status=completed means bounded candidate retrieval and its joint check completed, not that a numerical explanation was established. An association is established only when the server resolves a current reviewer-owned exact result/claim/sample/passage link in the same snapshot as the selected evidence check. Otherwise it is not_established with reviewed_result_passage_bridge_missing. A reviewed link is relation metadata, not a causal conclusion or scientific acceptance.same_snapshot and not_same_snapshot describe declared Paper/catalogue metadata, not an authenticated document, the same experiment, causal support or scientific independence.scientific_acceptance=false and independent_support_count=null never provide a confidence score or an ML approval label.
Mixed max_sources is a combined limit of at most 20 numerical parents and original passages, not 20 of each. The numerical allocation is max(1, floor(max_sources / 2)); unused capacity is available to originals. With a single slot, a matching numerical row takes priority and absence of original context is explicit. The UI states “Numerical explanation not established” and displays the two inventories separately; numerical comparisons do not imply comparable experiments or that the user requested an explanation.
scientific_mixed.status=unavailable withdraws numerical rows, original citations and associations together after a failed or changed combined check. No previous eligible subset is retained. Existing routes default to not_requested; missing metadata in older responses retains legacy behavior. A malformed present envelope withholds the mixed inventories and answer prose rather than falling back to an unchecked explanation.
evidence_packing and per-source packing_info describe selected context, catalogue-source/Work diversity, heuristic roles, bounded exclusions and canonical full-payload UTF-8 bytes. Selection counts are not independent papers or experiments. The packer budgets complete retained chunks; original source cards show whitespace-normalized previews of at most 280 characters. A snippet is not the full chunk or paper; the content hash binds the retained full chunk, not only the preview.
input_budget separately reports provider token-preflight status, model, request hash, byte and input-token limits, and generation_started (null when unknown or not requested). Actual providercount_tokens observations are not billed-usage receipts or model-version attestations; UTF-8 bytes are a resource bound, not model tokens. Mixed retrieval calls neither Gemini CountTokens nor generation: input_budget.status=not_requested, tokens_used=0, assessment_scope=none and answer_mode=abstention. Its measured packing representation is not a submitted generation request. Semantic retrieval may still call embedding/vector providers, so zero generation tokens does not mean zero retrieval cost, zero API quota use or a billing guarantee.
Scientific support metadata is additive: support_policy_version, citation_indices_valid, lexical_support_checked, scientific_support_status, claim_assessments[], support_warnings, support_coverage, assessment_scope, and answer_mode. The legacy citation_valid flag is deprecated and mechanical only; valid citation indices do not demonstrate that a claim is supported.
scientific_support_status is supported, contradicted, undetermined, or not_checked. “Supported” means narrow excerpt consistency checks passed, not scientific truth, experimental confirmation, or approval as an ML label. Claim assessments include draft text, cited indices, reason codes, and attributable evidence excerpts. Coverage counts and limits describe bounded checks, not exhaustive scientific validation.
assessment_scope=generated_draft refers to an attempted generated draft, not to verification of a delivered fallback. answer_modedistinguishes synthesis, limited_synthesis, extractive_fallback, and abstention. Conflicted or unresolved draft assertions are withheld rather than shown as supported conclusions. A fallback supplies source excerpts, not verified findings. Missing or incompatible metadata is displayed as unchecked; historical saved answers do not contain this current audit envelope.
language accepts "auto", "en", or "zh". Auto selects the question language for generated answers. Static operational notices and website-owned labels remain English; original queries and source wording retain their language.
Search results and Ask sources carry evidence_provenancewith evidence kind, exact revision hashes, producer versions, source coordinates and currentness. Version rag-evidence/1.0.0leaves original roots and text permissions unreviewed; its scientific authority flags are always false. It cannot establish independent confirmation or an ML training label. Derived Facts are labeled as generated extraction text, not original quotations. Known restricted or stale excerpts are withheld.
The generation route rechecks its selected database inputs after generation. A changed or unavailable check withdraws the draft and returns an abstention. Mixed retrieval instead checks both selected inventories together before returning them, without generating a draft. A matching snapshot is not scientific acceptance or a permanent currentness guarantee. Saved provenance is historical and does not revalidate a saved answer.
Structured-only and mixed history entries cannot replay the new response-level bindings or numerical rows; they store a static interaction notice. Mixed history also omits original candidates and associations. Rerun the query for a new qualified read instead of reconstructing unsupported historical evidence.
/materialsFree · no quotaBrowse and filter the superconductor materials database.
GET /v1/materials?family=cuprate,iron_based&tc_min=50&sort=tc_max&limit=100
Filters: family (comma-separated), tc_min, ambient_sc, pressure_min, pressure_max, experimental_only, knowledge_origin, source_role, is_unconventional, has_competing_order, pairing_symmetry, structure_phase.
Sort: tc_max | tc_ambient | arxiv_year | total_papers. Pagination via limit & offset.
Family, Tc, pressure and evidence filters must match one extracted result;matching_results identifies the matching occurrences and their pressure semantics. Unknown pressure is excluded from pressure limits unlessinclude_unknown_pressure=true is explicitly requested.ambient_sc=true requires an observed positive result with explicit ambient evidence;ambient_sc=false returns 422 because absence is not a negative experiment. These references are legacy occurrence identifiers, not reviewed ML labels.
Classification filters pairing_symmetry, is_unconventional and has_competing_order use the current source-reported material summary under material-semantics/1.0.0, not a same-result Tc/state predicate. The response declares classification_filter_scope=material_reported_summary_not_joint_state. Unknown or missing values never mean false; reported false requires source-linked method and detection conditions. Family priors and stale aggregate classifications do not match these filters.
The additive material_semantics envelope separates reported properties, inferred priors, state variability, extraction conflicts and unadjudicated dispute reports. Its support counts are occurrences and bibliographic identifiers, not independent works or replications.support.count_basis explains why identifier aliases and support.legacy_total_papersmay differ, including parent rollups and other catalogue policies. Legacy responses without this envelope remain unchecked.
structure_evidence contains pending text proposals and unassigned mentions under structure-evidence/1.0.0. Structure, phase and space-group summary aliases remain null until local material/state associations and source revisions can be reviewed. Nonempty structure_phase filters return 422 rather than matching unverified catalogue labels. A source-content hash and assembled-text span are provenance proposals, not a verified publication revision, coordinate artifact or scientific approval. Structure evidence excerpts are withheld pending source access and redistribution review; raw response records do not export the new extraction-only quotation containers.
/materials/{id}Free · no quotaFull detail for a single material — Tc values, pressure, crystal structure, pairing symmetry, all source records with paper links.
GET /v1/materials/mat%3AYBa2Cu3O7
/paper/{id}Free · no quotaPaper metadata — title, authors, abstract, journal, DOI, arXiv ID, extracted materials list, retraction status.
GET /v1/paper/arxiv%3A2301.12345
/similar/{paper_id}Free · no quotaFind semantically similar papers via vector search. Returns up to 10 neighbours.
/timelineFree · no quotaReported Tc Timeline results retain result identity, revision, year basis, state, criterion, and source provenance. Deterministic stratified display budgets (max_points) precede stable offset/limit pagination.sampling describes display selection; record_summarydescribes the full filtered, unsampled dataset. Neither is an approved training dataset or a world-record chronology.
GET /v1/timeline?schema_version=1&max_points=10000&offset=0&limit=1000&compact=true
Timeline returns ETag and X-Data-Versionafter current governance checks, with Cache-Control: private, no-store. Compact mode retains result provenance. Legacy results are unreviewed;reviewed_only=true is unavailable and returns 422 until accepted result-level review data exists. The compatibility filterexperimental_only=true means reported Observed origin, not independent experimental confirmation.
/discovery/candidatesFree · no quotaVersioned, paginated candidate summaries. Fetch full evidence only when needed from /discovery/candidates/{candidate_id}.
GET /v1/discovery/candidates?schema_version=1&offset=0&limit=24
/discovery/rps/releasesFree · no quotaThe separate rps-catalog/1.3 publication catalog lists checked release items[], failed release unavailable[], status, approval_sha256 and catalog_revision. Status is published, not_published, degraded or unavailable; failed verification must not be treated as a successfully empty catalog. Snapshot hashes describe the server read point, not permanent approval or scientific validation. The current board requires catalog 1.3; release pages and details retain 1.2.
Each item has public_bundle with status, sha256 and verifier_version. Only available supplies a bundle hash and rps-public-verifier/1.0.0.not_published and unavailable retain null hash/version fields and no download. A bundle failure can degrade the catalog without discarding a separately checked score release.
Download from GET /discovery/rps/releases/{id}/bundle with both manifest_sha256 and bundle_sha256 copied from the selected item. The fixed API route rechecks current publication configuration and returns a canonical attachment; unapproved, mismatched or unavailable requests receive an error, not an unpinned substitute. The browser has not independently verified a file merely because it displays this link. Refreshing the catalog clears old scores and download pins while new checks run.
Public bundle publication needs its own administrator pin as well as the release pin. Hash integrity and deterministic RPS recomputation do not establish reviewer identity, source truth, disclosure rights, scientific acceptance or a probability of superconductivity. Public review-attestation labels remain declarations, not authenticated reviewers or redistribution permission.
/statsFree · no quotaSite-wide statistics with separate aggregate-refresh, data-snapshot, and ingestion-pipeline timestamps/status.
/health/dependenciesFree · no quotaBounded PostgreSQL and Redis dependency probes. Returns 503 when a required dependency is unavailable. Process liveness and orchestrator readiness are also exposed upstream at /livez and /readyz.
/health/dataFree · no quotaNon-gating data-health metadata for the stats cache, dataset snapshot, ingestion stages, timeline projection, and Discovery feed. Data age is reported without being treated as a readiness failure.
Error codes
| Code | Meaning |
|---|---|
| 401 | Invalid or revoked API key |
| 403 | Account inactive or insufficient permissions |
| 404 | Resource not found (material / paper ID) |
| 422 | Validation error — check request body / query params |
| 429 | Daily quota exceeded — resets at 00:00 UTC |
Every response includes X-Request-ID and X-API-Version. Error JSON preserves detailand adds error_code plus request_id; include the request ID when reporting an API problem.
Full example: Python
These general request examples keep their route-specific response handling. Scientific filters or mixed questions require the structured fields and consistency checks above; empty paper hits or citation lists alone do not describe the numerical lookup outcome.
import requests
API = "https://api.jzis.org/sclib/v1"
KEY = "scl_YOUR_KEY"
HEAD = {"X-API-Key": KEY}
# 1. List cuprate materials with Tc > 100 K (free, no quota)
mats = requests.get(
f"{API}/materials",
params={"family": "cuprate", "tc_min": 100, "limit": 50},
headers=HEAD,
).json()
for m in mats["results"]:
print(f"{m['formula']} Tc={m['tc_max']} K papers={m['total_papers']}")
# 2. Semantic search (consumes 1 quota)
hits = requests.post(
f"{API}/search",
headers={**HEAD, "Content-Type": "application/json"},
json={"query": "pressure-induced superconductivity in hydrides", "top_k": 10},
).json()
for h in hits["results"]:
print(f"[{h['relevance_score']:.2f}] {h['title']}")
# 3. Ask a question (consumes 1 quota)
ans = requests.post(
f"{API}/ask",
headers={**HEAD, "Content-Type": "application/json"},
json={"question": "What is the mechanism of high-Tc in cuprates?", "max_sources": 5},
).json()
print(ans["answer"])
for s in ans["sources"]:
print(f" [{s['index']}] {s['title']} ({s['year']})")