Data and ingestion¶
HLA-Compass separates data upload, verification, ingestion, and publication so an uploaded object cannot become a trusted Catalog Version merely because a client says the transfer completed.
These examples use SDK 5.9.2 with the Catalog lifecycle API and the
20260909_01_catalog_ingestion_tables migration. Catalog writes and scientific
reads use separate authorization checks.
The upload/ingestion methods originated in SDK 5.0.2, but the current typed-read,
Catalog lifecycle, and runtime contracts require the matching release described
here. Do not infer environment readiness from the installed package alone.
For dev/staging before public release, use the SDK candidate installation.
Before uploading¶
- Sign in with
hla-compass auth login --env dev, then check the session withhla-compass auth status --check --env dev. - Use an organization-admin or platform-admin bearer session for Catalog
creation and ingestion. Developer, module-run, and publish-only API keys
cannot perform these control-plane operations. If
HLA_API_KEYorHLA_COMPASS_API_KEYis set, it takes precedence over the stored login; remove that override from the process before using the bearer session. - Confirm the provider is registered and your organization has the intended scientific domain provisioned. Catalog creation does not create an arbitrary database schema or provision a domain.
- Discover the Catalog UUID, its
ingestionTablescontract, and the required input columns/types. Prepare the complete replacement data, not only the rows that changed. Keep the original files and stable operation keys. - Check available storage capacity and ingestion access in the target environment. A rollout or quota failure requires an administrator to resolve it; it is not a reason to write directly to Catalog storage.
Catalog lifecycle¶
Catalogs expose two separate table lists:
| Field | Meaning |
|---|---|
sourceTables |
Readable entities, including the full peptidome read surface where applicable. Reference and vocabulary entities here do not become writable. |
ingestionTables: ["samples", ...] |
Explicit writable observation destinations. Every replace ingestion must supply this entire set. |
ingestionTables: [] |
Self-service ingestion is disabled; the Catalog remains readable according to its grants. |
ingestionTables: null |
No ingestion contract is configured. An owner/operator must review it before self-service ingestion is available. |
New integrations should send ingestionTables explicitly. For compatibility,
omitting it infers a contract only when the original submitted sourceTables
list contains exclusively supported observation tables, before read-surface
expansion. A mixed historical read list does not establish a write contract and
remains unconfigured pending explicit owner/operator review. Catalog metadata
updates cannot configure or repair this contract; do not replace that review
with direct storage writes or a duplicate Catalog.
Organization administrators can create organization-private Catalogs, update
only their mutable label and description, and soft-delete an owned active
Catalog. Provider key, Catalog key, schema, read and ingestion table contracts,
ownership, and visibility remain immutable. Foreign, missing, and already deleted Catalogs
all return 404, so callers cannot use this surface to enumerate another
organization's registry.
# Replace this placeholder with the schemaName confirmed by your administrator
# from GET /v1/data-catalogs/source-tables?domain=peptidome in this organization.
observational_schema = "ADMIN_CONFIRMED_SCHEMA_NAME"
catalog = client.create_catalog(
provider_key="PROVIDER_KEY",
catalog_key="research",
domain="peptidome",
schema_name=observational_schema,
label="Research",
source_tables=["samples"],
ingestion_tables=["samples"],
)
client.update_catalog(
catalog["id"],
label="Validated research",
description=None, # Explicitly clear; omit to leave unchanged.
)
Do not create a new Catalog for each refresh. Retain the returned Catalog UUID
and ingest a new version into it. client.delete_catalog(catalog_id) is a
separate, deliberate soft-delete operation; it is not cleanup required by this
example.
description is typed as str | None: pass a string to replace it, pass
None to clear it, or omit the keyword to preserve the current value.
domain selects the provisioned scientific data domain. Obtain the actual
observation schema from the authenticated
GET /v1/data-catalogs/source-tables?domain=peptidome response's schemaName
field, or ask your administrator for that exact returned value. Use it as
schema_name; the domain key peptidome is not the tenant schema name. The
same response lists available observation tables and their columns/types.
The supported write destinations are samples, raw_files, search_sessions,
sample_hla_typing, sample_allele_scoring, sample_diseases, and
sample_peptide_mappings. A Catalog can declare a narrower subset. Reference
and vocabulary data remain operator-managed; self-service ingestion requires
the existing integer foreign-key identities in each source.
The equivalent CLI commands are hla-compass data catalog create, show,
update, delete, and versions. Create, update, and delete require an
interactive confirmation unless --yes is supplied deliberately.
On CLI create, repeat --ingestion-table for the write contract or use
--read-only to disable self-service ingestion. --source-table declares the
read surface. Omitting both ingestion options retains the legacy adapter.
Read path¶
Discover a visible catalog, then choose an immutable ready Catalog Version. For
version-sensitive typed reads, pass its currentVersion.id as the version
filter. Omitting version asks the server to resolve whichever ready version is
current at request time and is not a reproducible pin. Do not replace a Catalog
Version ID with a mutable schema name or S3 prefix.
from hla_compass.client import APIClient
discovery = APIClient(environment="dev")
catalog = next(
item
for item in discovery.list_catalogs()
if item["providerKey"] == "PROVIDER_KEY" and item["catalogKey"] == "CATALOG_KEY"
)
current_version = catalog.get("currentVersion")
if current_version is None:
raise RuntimeError("Selected catalog has no ready current version")
client = APIClient(
provider=catalog["providerKey"],
catalog=catalog["catalogKey"],
environment="dev",
)
peptides = client.get_peptides(
filters={"version": current_version["id"], "minLength": 8, "maxLength": 11},
limit=100,
)
Governed Catalog Import¶
Catalog Import uses a checksum-bound multipart lifecycle:
- Calculate the exact file size and lowercase raw-file SHA-256.
- Initialize an upload with a stable idempotency key.
- Sign consecutive parts in batches of at most 100.
- Upload the exact signed byte count and checksum for every part.
- Complete with the latest upload
stateVersionand the exact part list. - Poll until the upload is
claimedor terminal. - Submit only claimed upload UUIDs to a replace-only Catalog Ingestion Run.
- Poll the ingestion job until a terminal state.
Upload completion returns uploaded_unverified. Only the platform verifier can
move an exact S3 object version to claimed after its size and raw SHA-256
match the reservation. quarantined, aborted, and expired are terminal and
cannot be ingested.
Environment rollout gates: the SDK surface may be present while multipart
upload, reconciliation, or ingestion remains disabled in an environment. A
typed 503 rollout-gate response means the workflow is unavailable there; do
not bypass it with a legacy upload route or direct Catalog Version mutation.
SDK example¶
# Save this key before the first request and reuse it for the same file.
operation_key = "catalog-import-2026-09-09-samples"
upload = client.upload_catalog_import_file(
"CATALOG_UUID",
"samples.parquet",
idempotency_key=operation_key,
wait_until_claimed=True,
)
upload_catalog_import_file() streams the file, calculates the raw and
per-part SHA-256 values, follows the server-issued part plan, retries transient
part-transfer failures, completes with exact ETags, and polls durable status.
It never blindly replays an ambiguous completion. The lower-level initialize,
part-signing, completion, abort, and status methods remain available for clients
that need to persist each transfer step themselves. Do not log signed upload
URLs, ETags, checksums associated with sensitive content, or internal object
versions.
The equivalent CLI transfer is:
hla-compass data ingestion upload CATALOG_UUID samples.parquet \
--idempotency-key catalog-import-2026-09-09-samples --env dev
After every required upload is claimed:
job = client.data.ingestion.submit(
"CATALOG_UUID",
sources=[{"uploadId": upload["uploadId"], "targetTable": "samples"}],
idempotency_key="dataset-refresh-2026-07-14",
)
state = client.data.ingestion.status("CATALOG_UUID", job["jobId"])
Submission accepts 1–32 claimed uploads. Replace mode requires the complete
nonempty ingestionTables set declared by the Catalog. Inspect it with
client.get_catalog(catalog_id) before uploading. If the Catalog declares multiple
tables, upload all of them and include one source entry per destination in the
same submission. A single samples file is sufficient only for a Catalog whose
complete ingestionTables set is ["samples"]. Neither a null nor an empty
contract permits submission. Append mode and direct client publication
are unavailable. The retired catalog-import Module template is not an
ingestion mechanism; use this governed control plane.
The CLI equivalent is:
hla-compass data ingestion submit CATALOG_UUID \
--source CLAIMED_UPLOAD_UUID:samples \
--idempotency-key dataset-refresh-2026-09-09 --env dev
hla-compass data ingestion status CATALOG_UUID JOB_UUID --env dev
Repeat --source for every required table. The upload and submit commands ask
for confirmation; --yes explicitly suppresses it for an already-reviewed
operation.
Submission returns a queued job, not a published version. Poll
client.data.ingestion.status(catalog_id, job_id) at a bounded interval until
status is succeeded, failed, or canceled. On failure, inspect errorCode,
error, errorDetails, and the attempt before deciding to retry. On success,
retain catalogVersionId, verify that version with get_catalog_version(), and
read using that exact ID. The initial empty ready version created with a new
Catalog is not evidence that your data was ingested.
Cancellation and retry¶
Cancellation and retry are explicit control-plane operations. Fetch status
immediately before either operation and pass its top-level stateVersion as
the optimistic-concurrency value. Do not use the nested
attempt.stateVersion; that value belongs to the worker attempt, not the job.
state = client.data.ingestion.status("CATALOG_UUID", job["jobId"])
# Perform only after the user has explicitly approved cancellation.
canceled = client.data.ingestion.cancel(
"CATALOG_UUID",
job["jobId"],
expected_job_state_version=state["stateVersion"],
reason="Superseded by corrected source data",
)
Retry only when the job is failed and status reports retryable=true. Use a
new stable retry key that differs from the original submission key, and persist
it before the request:
state = client.data.ingestion.status("CATALOG_UUID", job["jobId"])
if state["status"] != "failed" or not state["retryable"]:
raise RuntimeError("This ingestion job is not retryable")
retried = client.data.ingestion.retry(
"CATALOG_UUID",
job["jobId"],
expected_job_state_version=state["stateVersion"],
retry_idempotency_key="dataset-refresh-2026-07-14-retry-1",
)
Reusing the same retry key after an ambiguous response may return
outcome=replayed; that is the prior retry result, not a newly created run.
The equivalent CLI commands are hla-compass data ingestion cancel and
hla-compass data ingestion retry. Both require confirmation unless --yes
is supplied deliberately.
Executable ingestion quickstart¶
The maintained example below exercises the high-level governed upload and the
replace-only ingestion submission for a single-destination Catalog. It stops
after submission and prints the job receipt; use the status step above to
verify publication. It mutates the selected environment, so
confirm the discovered Catalog UUID, organization, file, target table, and
stable operation ID first. Save or
open catalog_ingestion.py, then run:
python catalog_ingestion.py CATALOG_UUID samples.parquet samples \
--operation-id dataset-refresh-2026-07-15 \
--version-label 2026-07-15 \
--env dev
Ambiguous responses¶
Persist initialization and ingestion idempotency keys before the request. If a connection drops, retry with the same request and key or query status. Never change the key merely because the client did not receive the response.
Multipart completion has a durable recovery path on the platform. The high-level uploader queries durable status instead of blindly replaying an ambiguous completion call; lower-level callers must do the same while the reconciler proves the exact completed object version or safely returns the upload to an uploadable state.