Skip to content

Python SDK

The hla-compass package contains the Python client, CLI, module runtime, data helpers, and optional local MCP gateway. The SDK is a client of the platform; installing it does not enable a server feature that is disabled in the target environment.

SDK 5.7 keeps typed sample, peptide and protein reads compatible with older GET APIs. Each client checks public /v1/system/version metadata once per API base URL and uses POST bodies only when features.scientific_typed_post is true. Existing GET requests preserve all ID filters and explicit version pins; API errors are reported without dropping a pin or switching transport. Empty or invalid ID arrays fail locally. Recreate a client after an API upgrade to discover newly available POST support. Prepared runs, notebooks and saved views require their corresponding API release.

Install

Use Python 3.11–3.14 in a virtual environment. After SDK 5.9.2 completes backend acceptance and public release, install it from PyPI:

python -m pip install "hla-compass==5.9.2"
hla-compass --version

Pin the version tested by your application. Check the target environment's live OpenAPI document for available operations; a failed prepare request must not fall back to an unreviewed run.

Dev candidate installation

Before public release, obtain the approved SDK 5.9.2 candidate bundle for the matching validation environment. Place its wheel and release.json in sdk-wheel/. Verify the candidate's own version, filename and SHA-256 before installation; an older TestPyPI wheel does not validate this candidate.

import hashlib
import json
from pathlib import Path

release = json.loads(Path("sdk-wheel/release.json").read_text())
wheel = Path("sdk-wheel/hla_compass-5.9.2-py3-none-any.whl")
assert release["version"] == "5.9.2", "Wrong SDK candidate version"
assert release["filename"] == wheel.name, "Wrong SDK candidate filename"
assert hashlib.sha256(wheel.read_bytes()).hexdigest() == release["wheelSha256"], "SDK checksum mismatch"

Continue only if all checks succeed:

python -m pip install --index-url https://pypi.org/simple/ "./sdk-wheel/hla_compass-5.9.2-py3-none-any.whl[authoring]"
hla-compass --version

Keep this exact wheel for local module checks with hla-compass test --offline --sdk-path /path/to/sdk-wheel/hla_compass-5.9.2-py3-none-any.whl. Dependencies resolve from PyPI; do not combine TestPyPI and PyPI with --extra-index-url. Hosted checks require the matching broker/projector API.

UI kit and upgrading to 5.9

SDK 5.9.2 includes UI kit 0.2.0 in the UI and pipeline UI scaffolds: scientific charts, interactive research controls, governed host controls, and browser PDF export. The Python API and module execution contracts remain compatible.

New scaffolds include the kit automatically. For an existing module, follow the adoption checklist in its generated UI_STYLE.md; refreshing SKILL.md does not replace UI code. Pin SDK 5.9.2, regenerate the dependency lock, rebuild, and verify the published module before replacing a scientific version.

Client setup

from hla_compass.client import APIClient

client = APIClient(environment="dev")

By default the client uses the SDK credential store. Prefer that flow for user sessions. Use an API key only for a supported machine integration and keep it outside source control:

Supply HLA_API_KEY through your process's secret store. APIClient reads it automatically; an API key takes precedence over a stored login. For the user session examples below, remove both the canonical variable and its legacy alias, then sign in:

unset HLA_API_KEY HLA_COMPASS_API_KEY
hla-compass auth login --env dev

API keys, bearer sessions, and module-run credentials have different capabilities. A credential that can read one data surface is not automatically authorized for privileged control-plane operations. The SDK selects the bearer /v1/... or API-key /v1/api/... route for supported operations. See authentication for permissions and organization scope. Saving, refreshing or exporting a saved view requires a user session; API keys can read permitted saved views. Run credentials are issued by the platform to an admitted run and are not personal automation keys.

Legacy module-upload migration

SDK 5 retains the published APIClient.upload_module(module_path, module_name, version) signature only as a local migration stub. It raises APIError with status code 410 before reading module_path or making an HTTP request. Migrate source publication to APIClient.publish_module_source(manifest=..., source_zip=..., scope="org", idempotency_key="<stable-retry-key>"), or run the supported CLI flow from the module source directory:

hla-compass publish --env dev --scope org \
  --idempotency-key "$PUBLISH_OPERATION_ID" --wait

Set PUBLISH_OPERATION_ID to a new operation identifier once, then retain it for retries of the same publication. Complete the dependency-lock prerequisites in publishing prerequisites before running this command.

Both replacements enter the governed asynchronous build, scan, signature, and immutable intake workflow; neither is a direct ZIP-registration call. The source archive is uploaded directly to platform storage with a presigned PUT, so its size is not constrained by the request transport. The SDK still enforces a safe 5 MiB budget for the JSON body itself — manifest metadata and schemas — and raises APIError with status code 413 before sending an oversized request.

Against an older platform without the source-upload route the SDK falls back to embedding the archive in the request, where the historical 3 MiB compressed limit applies; the resulting error names that cause rather than implying the archive is inherently too large. Datasets and model weights are still better read through governed Catalog or storage interfaces at runtime than baked into every build.

Retries and idempotency

The client retries transient failures for HTTP-idempotent methods. An unsafe POST or PATCH is retried only when the caller supplies a stable idempotency key for an endpoint with a documented server-side replay contract. Reusing a key with a different request is an error; generating a new key after an ambiguous response may create a second operation.

For an application-level workflow:

  1. Persist the key with the local operation draft.
  2. Reuse it after timeout or connection loss.
  3. Query operation status before creating a new key.
  4. Generate a distinct key only for a deliberately new operation.

Typed data access

Discover the Catalog identity and its ready version:

hla-compass data catalogs --env dev --json > catalogs.json
python -m json.tool catalogs.json

Choose an active entry whose currentVersion is present. Set CATALOG_ID to its canonical id and CATALOG_VERSION_ID to currentVersion.id; these are different UUIDs. Use the actual values from discovery. This example prints at most 250 matching peptide records as JSON lines:

import json
import os
from hla_compass.client import APIClient

client = APIClient(environment="dev")
data_client = client.for_catalog_id(os.environ["CATALOG_ID"])
catalog_version_id = os.environ["CATALOG_VERSION_ID"]
for peptide in data_client.iter_peptides(
    filters={
        "version": catalog_version_id,
        "hla_allele": "HLA-A*02:01",
        "minLength": 8,
        "maxLength": 11,
        "sort": "id",
        "order": "asc",
    },
    page_size=100,
    max_results=250,
):
    print(json.dumps(peptide))

for_catalog_id() validates one canonical visible Catalog UUID and fails on missing, inactive, duplicate, ambiguous, or malformed discovery metadata. for_catalog(provider, catalog) accepts an exact discovered key pair with the same identity checks. Catalog binding never selects a Catalog Version. Pass the same explicit Version UUID on every page. Omitting version reads the ready current version at request time, which may change between pages.

get_peptides(), get_samples() and get_proteins() return one page as a list, not the API pagination envelope. Their iter_* helpers advance offsets until a short page or max_results; an empty list is a valid result. The example's 250-record bound is deliberate, not a complete-catalog export. Keep filters, version and deterministic sort unchanged across pages.

Typed entity IDs are numeric IDs within a Catalog, not Catalog UUIDs. ID filters accept at most 500 distinct IDs per filter; retain Catalog and Version identity alongside those IDs. An empty or invalid ID list is an error, never a request for all rows. Do not split a selection into separately charged runs to evade this limit. Larger selections need a supported saved query/materialization workflow and must fit the environment's row, byte and time budgets.

Raw SQL is a read-only compatibility surface, subject to the caller's role, Catalog access, data profile, relation allowlist and query budgets. A restricted profile may prevent saved SQL even when typed filtered reads are permitted. Do not use raw SQL or catalog-wide storage credentials as a module-runtime shortcut around governed run inputs.

Saved views and exact run inputs

On an environment with the named-source API, save a definition from Data Explorer or call client.data_views.create(...) with discovered Catalog/Version identities and SQL validated against that Catalog's schema. Saving requires a bearer user session. It saves a definition and bounded preview; it does not prove a complete export or start a module. Status, refresh and export are separate operations.

New views pin the displayed version by default. Explicit versionPinMode="current" means Follow latest until a run is prepared. Preparation freezes the resolved versions, selection and compatible module version. A mutable saved-view ID by itself is not a reproducible run input.

Multiple catalogs are independent named sources. A module must explicitly accept the named manifest and each source's entity/columns/format. A compatible single-source module can use its existing CSV/JSON input adapter; it does not automatically gain cross-catalog support. There is no implied SQL join, flattened union or multiple-run expansion. Preserve the plan, Run ID and returned Dataset Version/lineage references. Reusing an admitted plan recovers that operation; starting a new scientific replay is a separate action whose input availability and authorization must be checked.

Prepare, review, then run

First discover available modules and inspect the selected module's input schema:

hla-compass modules list --env dev --format json > modules.json
python -m json.tool modules.json

Set MODULE_ID to the selected full UUID, then inspect it:

hla-compass modules show "$MODULE_ID" --env dev --format json > module.json
python -m json.tool module.json

The list is a bounded discovery page, not an assurance that every module was returned. Fill inputs.json with a JSON object matching the selected published version's schema. Input names and example peptides are module-specific; do not assume a peptides field exists. For a saved selection, also save the supported dataContext as selection.json and pass --data-context selection.json below.

hla-compass runs prepare "$MODULE_ID" --inputs inputs.json --env dev > plan.json
python -m json.tool plan.json

Review preparedRequest (module version, parameters, mode, compute profile and any frozen data context), estimated_cost, currency and expiresAt. Only after accepting those values:

hla-compass runs confirm plan.json --env dev > run.json
python -m json.tool run.json

The equivalent SDK calls are client.prepare_module_run(module_id, parameters=...) and client.start_prepared_module_run(plan). Keep the complete plan unchanged. After a timeout, retry confirmation with that same file and the same environment, organization and user; do not generate a replacement plan or select the newest run from a list. An expired unused plan, changed price or changed request requires preparation and review again. Do not strip admission fields to force legacy start.

Dispatch is not completion. Use the exact returned Run ID to poll and retrieve the result. Set RUN_ID to that value before this standalone snippet:

import json
import os
from hla_compass.client import APIClient

client = APIClient(environment="dev")
result = client.wait_for_module_run(os.environ["RUN_ID"], timeout=600, poll_interval=5)
print(json.dumps(result, indent=2))

The result includes run/module identity, status, data and provenance context. Inspect success, error, artifactErrors and the module's own output contract; HTTP success alone does not prove usable scientific output. A polling timeout ends the wait, not the server run. Keep the Run ID for subsequent inspection.

For pipeline-specific requirements and available run paths, use the pipeline guide. Module preparation does not establish that every standalone pipeline, external runner or replay path accepts the same data context.

Personal-data export

A bearer-authenticated user can request a time-limited export of their own account data:

from hla_compass.client import APIClient

client = APIClient(environment="dev")
export = client.request_personal_data_export(
    export_format="zip",
    date_range="last90days",
    include_jobs=False,
    include_results=True,
)
print(export["download_url"])

The SDK returns signed-URL metadata and does not download the bytes. The URL is sensitive and expires after 15 minutes. API keys and module-run credentials are not accepted for personal-data export.

Local development safety

hla-compass test is offline by default and does not forward host credentials or a selected catalog into the container. Use --live only for a test that is intended to call an authenticated environment, and keep live integration jobs explicit in CI.

Files, retained outputs and reports

SDK 5.9 adds cancellable file uploads for managed organization storage and connected external S3. client.upload_storage_file(source_id, local_path, idempotency_key=..., prefix=...) returns a verified, versioned receipt. Keep the same file and idempotency key after a timeout; resume with upload_id. CLI:

hla-compass data uploads --env dev upload SOURCE_ID evidence.csv \
  --prefix analysis/ --idempotency-key evidence-upload-1
hla-compass data uploads --env dev status SOURCE_ID UPLOAD_ID

Prefix is relative to the connection; do not prepend its configured files/ again. Duplicate filenames receive distinct identities. Ctrl+C requests cancellation; pending cleanup retains its reservation until verified. Completed files are preserved. File storage does not publish a Catalog; see data and ingestion for that workflow.

Use client.runs.artifacts(run_id) and client.runs.download_artifact(...) for exact retained CSV, JSON, PDF or other registered run files. Supply the artifact, lineage and version IDs from that run's listing. Never reconstruct an S3 URL or substitute the latest run. SDK downloads verify recorded size and checksum when available; historical missing checksums are explicitly reported.

The UI kit's ReportBuilder accepts your actual figures, methods, data, filters and provenance. Print / Save as PDF is a local export. To register it, choose the saved PDF and Attach report to Results in a supporting module. The PDF (20 MiB maximum) and snapshot JSON (5 MiB maximum) become immutable user-created reports linked to the exact run. They use its existing access rules, and do not change computation outputs or credits. Results labels them user-supplied.

The SDK also provides client.runs.reports, report, attach_report and download_report; CLI uses hla-compass runs reports. MCP exposes the matching module_runs actions. Read the packaged INTEGRATIONS.md and UI_STYLE.md for complete signatures, permissions, retry behavior and examples. Use these 5.9 interfaces only after the matching API is deployed in your environment.