Skip to content

Data and ingestion

HLA-Compass separates data upload, verification, ingestion, and publication so an uploaded object cannot become a trusted Catalog Version merely because a client says the transfer completed.

These examples use SDK 5.9.2 with the Catalog lifecycle API and the 20260909_01_catalog_ingestion_tables migration. Catalog writes and scientific reads use separate authorization checks. The upload/ingestion methods originated in SDK 5.0.2, but the current typed-read, Catalog lifecycle, and runtime contracts require the matching release described here. Do not infer environment readiness from the installed package alone. For dev/staging before public release, use the SDK candidate installation.

Before uploading

  1. Sign in with hla-compass auth login --env dev, then check the session with hla-compass auth status --check --env dev.
  2. Use an organization-admin or platform-admin bearer session for Catalog creation and ingestion. Developer, module-run, and publish-only API keys cannot perform these control-plane operations. If HLA_API_KEY or HLA_COMPASS_API_KEY is set, it takes precedence over the stored login; remove that override from the process before using the bearer session.
  3. Confirm the provider is registered and your organization has the intended scientific domain provisioned. Catalog creation does not create an arbitrary database schema or provision a domain.
  4. Discover the Catalog UUID, its ingestionTables contract, and the required input columns/types. Prepare the complete replacement data, not only the rows that changed. Keep the original files and stable operation keys.
  5. Check available storage capacity and ingestion access in the target environment. A rollout or quota failure requires an administrator to resolve it; it is not a reason to write directly to Catalog storage.
from hla_compass import APIClient

client = APIClient(environment="dev")

Catalog lifecycle

Catalogs expose two separate table lists:

Field Meaning
sourceTables Readable entities, including the full peptidome read surface where applicable. Reference and vocabulary entities here do not become writable.
ingestionTables: ["samples", ...] Explicit writable observation destinations. Every replace ingestion must supply this entire set.
ingestionTables: [] Self-service ingestion is disabled; the Catalog remains readable according to its grants.
ingestionTables: null No ingestion contract is configured. An owner/operator must review it before self-service ingestion is available.

New integrations should send ingestionTables explicitly. For compatibility, omitting it infers a contract only when the original submitted sourceTables list contains exclusively supported observation tables, before read-surface expansion. A mixed historical read list does not establish a write contract and remains unconfigured pending explicit owner/operator review. Catalog metadata updates cannot configure or repair this contract; do not replace that review with direct storage writes or a duplicate Catalog.

Organization administrators can create organization-private Catalogs, update only their mutable label and description, and soft-delete an owned active Catalog. Provider key, Catalog key, schema, read and ingestion table contracts, ownership, and visibility remain immutable. Foreign, missing, and already deleted Catalogs all return 404, so callers cannot use this surface to enumerate another organization's registry.

# Replace this placeholder with the schemaName confirmed by your administrator
# from GET /v1/data-catalogs/source-tables?domain=peptidome in this organization.
observational_schema = "ADMIN_CONFIRMED_SCHEMA_NAME"

catalog = client.create_catalog(
    provider_key="PROVIDER_KEY",
    catalog_key="research",
    domain="peptidome",
    schema_name=observational_schema,
    label="Research",
    source_tables=["samples"],
    ingestion_tables=["samples"],
)
client.update_catalog(
    catalog["id"],
    label="Validated research",
    description=None,  # Explicitly clear; omit to leave unchanged.
)

Do not create a new Catalog for each refresh. Retain the returned Catalog UUID and ingest a new version into it. client.delete_catalog(catalog_id) is a separate, deliberate soft-delete operation; it is not cleanup required by this example.

description is typed as str | None: pass a string to replace it, pass None to clear it, or omit the keyword to preserve the current value.

domain selects the provisioned scientific data domain. Obtain the actual observation schema from the authenticated GET /v1/data-catalogs/source-tables?domain=peptidome response's schemaName field, or ask your administrator for that exact returned value. Use it as schema_name; the domain key peptidome is not the tenant schema name. The same response lists available observation tables and their columns/types.

The supported write destinations are samples, raw_files, search_sessions, sample_hla_typing, sample_allele_scoring, sample_diseases, and sample_peptide_mappings. A Catalog can declare a narrower subset. Reference and vocabulary data remain operator-managed; self-service ingestion requires the existing integer foreign-key identities in each source.

The equivalent CLI commands are hla-compass data catalog create, show, update, delete, and versions. Create, update, and delete require an interactive confirmation unless --yes is supplied deliberately. On CLI create, repeat --ingestion-table for the write contract or use --read-only to disable self-service ingestion. --source-table declares the read surface. Omitting both ingestion options retains the legacy adapter.

Read path

Discover a visible catalog, then choose an immutable ready Catalog Version. For version-sensitive typed reads, pass its currentVersion.id as the version filter. Omitting version asks the server to resolve whichever ready version is current at request time and is not a reproducible pin. Do not replace a Catalog Version ID with a mutable schema name or S3 prefix.

from hla_compass.client import APIClient

discovery = APIClient(environment="dev")
catalog = next(
    item
    for item in discovery.list_catalogs()
    if item["providerKey"] == "PROVIDER_KEY" and item["catalogKey"] == "CATALOG_KEY"
)
current_version = catalog.get("currentVersion")
if current_version is None:
    raise RuntimeError("Selected catalog has no ready current version")

client = APIClient(
    provider=catalog["providerKey"],
    catalog=catalog["catalogKey"],
    environment="dev",
)
peptides = client.get_peptides(
    filters={"version": current_version["id"], "minLength": 8, "maxLength": 11},
    limit=100,
)

Governed Catalog Import

Catalog Import uses a checksum-bound multipart lifecycle:

  1. Calculate the exact file size and lowercase raw-file SHA-256.
  2. Initialize an upload with a stable idempotency key.
  3. Sign consecutive parts in batches of at most 100.
  4. Upload the exact signed byte count and checksum for every part.
  5. Complete with the latest upload stateVersion and the exact part list.
  6. Poll until the upload is claimed or terminal.
  7. Submit only claimed upload UUIDs to a replace-only Catalog Ingestion Run.
  8. Poll the ingestion job until a terminal state.

Upload completion returns uploaded_unverified. Only the platform verifier can move an exact S3 object version to claimed after its size and raw SHA-256 match the reservation. quarantined, aborted, and expired are terminal and cannot be ingested.

Environment rollout gates: the SDK surface may be present while multipart upload, reconciliation, or ingestion remains disabled in an environment. A typed 503 rollout-gate response means the workflow is unavailable there; do not bypass it with a legacy upload route or direct Catalog Version mutation.

SDK example

# Save this key before the first request and reuse it for the same file.
operation_key = "catalog-import-2026-09-09-samples"

upload = client.upload_catalog_import_file(
    "CATALOG_UUID",
    "samples.parquet",
    idempotency_key=operation_key,
    wait_until_claimed=True,
)

upload_catalog_import_file() streams the file, calculates the raw and per-part SHA-256 values, follows the server-issued part plan, retries transient part-transfer failures, completes with exact ETags, and polls durable status. It never blindly replays an ambiguous completion. The lower-level initialize, part-signing, completion, abort, and status methods remain available for clients that need to persist each transfer step themselves. Do not log signed upload URLs, ETags, checksums associated with sensitive content, or internal object versions.

The equivalent CLI transfer is:

hla-compass data ingestion upload CATALOG_UUID samples.parquet \
  --idempotency-key catalog-import-2026-09-09-samples --env dev

After every required upload is claimed:

job = client.data.ingestion.submit(
    "CATALOG_UUID",
    sources=[{"uploadId": upload["uploadId"], "targetTable": "samples"}],
    idempotency_key="dataset-refresh-2026-07-14",
)
state = client.data.ingestion.status("CATALOG_UUID", job["jobId"])

Submission accepts 1–32 claimed uploads. Replace mode requires the complete nonempty ingestionTables set declared by the Catalog. Inspect it with client.get_catalog(catalog_id) before uploading. If the Catalog declares multiple tables, upload all of them and include one source entry per destination in the same submission. A single samples file is sufficient only for a Catalog whose complete ingestionTables set is ["samples"]. Neither a null nor an empty contract permits submission. Append mode and direct client publication are unavailable. The retired catalog-import Module template is not an ingestion mechanism; use this governed control plane.

The CLI equivalent is:

hla-compass data ingestion submit CATALOG_UUID \
  --source CLAIMED_UPLOAD_UUID:samples \
  --idempotency-key dataset-refresh-2026-09-09 --env dev
hla-compass data ingestion status CATALOG_UUID JOB_UUID --env dev

Repeat --source for every required table. The upload and submit commands ask for confirmation; --yes explicitly suppresses it for an already-reviewed operation.

Submission returns a queued job, not a published version. Poll client.data.ingestion.status(catalog_id, job_id) at a bounded interval until status is succeeded, failed, or canceled. On failure, inspect errorCode, error, errorDetails, and the attempt before deciding to retry. On success, retain catalogVersionId, verify that version with get_catalog_version(), and read using that exact ID. The initial empty ready version created with a new Catalog is not evidence that your data was ingested.

Cancellation and retry

Cancellation and retry are explicit control-plane operations. Fetch status immediately before either operation and pass its top-level stateVersion as the optimistic-concurrency value. Do not use the nested attempt.stateVersion; that value belongs to the worker attempt, not the job.

state = client.data.ingestion.status("CATALOG_UUID", job["jobId"])

# Perform only after the user has explicitly approved cancellation.
canceled = client.data.ingestion.cancel(
    "CATALOG_UUID",
    job["jobId"],
    expected_job_state_version=state["stateVersion"],
    reason="Superseded by corrected source data",
)

Retry only when the job is failed and status reports retryable=true. Use a new stable retry key that differs from the original submission key, and persist it before the request:

state = client.data.ingestion.status("CATALOG_UUID", job["jobId"])
if state["status"] != "failed" or not state["retryable"]:
    raise RuntimeError("This ingestion job is not retryable")

retried = client.data.ingestion.retry(
    "CATALOG_UUID",
    job["jobId"],
    expected_job_state_version=state["stateVersion"],
    retry_idempotency_key="dataset-refresh-2026-07-14-retry-1",
)

Reusing the same retry key after an ambiguous response may return outcome=replayed; that is the prior retry result, not a newly created run. The equivalent CLI commands are hla-compass data ingestion cancel and hla-compass data ingestion retry. Both require confirmation unless --yes is supplied deliberately.

Executable ingestion quickstart

The maintained example below exercises the high-level governed upload and the replace-only ingestion submission for a single-destination Catalog. It stops after submission and prints the job receipt; use the status step above to verify publication. It mutates the selected environment, so confirm the discovered Catalog UUID, organization, file, target table, and stable operation ID first. Save or open catalog_ingestion.py, then run:

python catalog_ingestion.py CATALOG_UUID samples.parquet samples \
  --operation-id dataset-refresh-2026-07-15 \
  --version-label 2026-07-15 \
  --env dev

Ambiguous responses

Persist initialization and ingestion idempotency keys before the request. If a connection drops, retry with the same request and key or query status. Never change the key merely because the client did not receive the response.

Multipart completion has a durable recovery path on the platform. The high-level uploader queries durable status instead of blindly replaying an ambiguous completion call; lower-level callers must do the same while the reconciler proves the exact completed object version or safely returns the upload to an uploadable state.